Compositions and methods for use in immunotherapy
The CasX:guide nucleic acid system modifies immune cells to reduce antigen processing proteins, enhancing their specificity for cancer cells and improving the therapeutic index in immunotherapy, addressing the non-specificity and side effects of current cytotoxic drugs.
Patent Information
- Application Number
- JP2022515493
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-04
- Filing Date
- 2020-09-09
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2040-09-09
AI Technical Summary
Current cytotoxic drugs used in cancer therapy are non-specific, causing damage to both normal and diseased cells, which limits treatment compatibility and increases side effects.
The use of a CasX:guide nucleic acid system for modifying target nucleic acid sequences in immune cells to reduce or eliminate proteins involved in antigen processing and presentation, thereby generating immune cells with improved specificity for cancer cells.
This approach enables the generation of immune cells with enhanced therapeutic index for immunotherapy, reducing the risk of graft-versus-host disease and improving treatment compatibility by targeting specific cancer cells while minimizing damage to normal tissues.
Smart Images

Figure 0007696335000096 
Figure 0007696335000097 
Figure 0007696335000098
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 62 / 897,947, filed on September 9, 2019, and U.S. Provisional Patent Application No. 63 / 075,041, filed on September 4, 2020, the contents of each of which are hereby incorporated by reference in their entirety.
[0002] Description of Electronically Submitted Text Files The contents of the text files electronically submitted with this specification are hereby incorporated by reference in their entirety: Computer - readable format copy of the sequence listing (file name: SCRB_016_02WO_SeqList_ST25.txt, recording date: September 9, 2020, file size 12.0 megabytes).
Background Art
[0003] Many approved therapeutic drugs, such as cancer therapeutics, are cytotoxic drugs that kill both normal and diseased cells. The therapeutic effect of these cytotoxic drugs depends on diseased cells being more sensitive than normal cells, thereby making it possible to achieve a clinical response using a dose that does not cause unacceptable side effects. However, essentially all of these non - specific drugs cause some damage, if not severe, to normal tissues, thereby often limiting treatment compatibility.
[0004] Genome engineering can provide a different approach to cytotoxic drugs in that it enables the generation of immune cells programmed to specifically bind to and kill diseased cells, such as cancer cells. The emergence of chimeric antigen receptor T cell (CAR-T) technology has brought about a new mode of therapeutic effect in certain types of cancer. By manipulating cells containing CARs to reduce HLA protein mismatches and reduce or eliminate wild-type T cell receptors or other components of the modified cells compared to the recipient subject, the possibility of graft-versus-host disease (GVHD) can be reduced or eliminated by eliminating host T cell receptor recognition of and response to mismatched (e.g., allogeneic) graft tissue (see, for example, Takahiro Kamiya, T. et al. A novel method to generate T-cell receptor-deficient chimeric antigen receptor T cells. Blood Advances 2:517 (2018)). Thus, using this approach, it has been possible to generate immune cells with an improved therapeutic index for immuno-oncology applications in subjects with diseases such as cancer, autoimmune diseases, and graft rejection.
[0005] Since the CRISPR / Cas system is suitable for genome editing in eukaryotic cells, these two technologies enable the manipulation of immune cells with potent cytotoxicity against target cells and also have the potential to reduce or eliminate cell markers that contribute to the induction of unwanted recipient immune responses to the transplantation of those cells, particularly in the case of allogeneic transplantation of such cells. Thus, there is a need for modified cells and methods of manipulating such cells to make engineered CAR-T cells for use in immunotherapy treatments, such as allogeneic-based immunotherapy treatments. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0006] In some embodiments, the present disclosure provides compositions and methods of a CasX:guide nucleic acid system (CasX:gNA system) for modifying target nucleic acid sequences of cellular genes encoding one or more proteins involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response. As described above, the proteins are selected from the group consisting of beta-2-microglobulin (B2M), T cell receptor alpha constant region (TRAC or TCRA), class II major histocompatibility complex transactivator (CIITA), T cell receptor beta constant 1 (TRBC1 or TCRB), T cell receptor beta constant 2 (TRBC2), programmed cell death 1 (PD-1), cytokine-inducible SH2 (CISH), T cell immunoreceptor with Ig and ITIM domains (TIGIT), adenosine A2a receptor (ADORA2A), killer cell lectin-like receptor C1 (NKG2A), cytotoxic T lymphocyte-associated protein 4 (CTLA-4), lymphocyte activation 3 (LAG-3), T cell immunoglobulin and mucin domain 3 (TIM-3), 2B4 (CD244), human leukocyte antigen A (HLA-A), human leukocyte antigen B (HLA-B), TGFβ receptor 2 (TGFβRII), cluster of differentiation 247 (CD247), CD3d molecule (CD3D), CD3e molecule (CD3E), CD3g molecule (CD3G), CD52 molecule (CD52), human leukocyte antigen C (HLA-C), deoxycytidine kinase (dCK), or FKBP prolyl isomerase 1A (FKBP1A). The CasX:gNA system can include a reference CasX protein, a CasX variant protein having improved properties compared to the reference CasX, a guide nucleic acid (gNA) that is a reference sequence, or a gNA variant having improved properties compared to the reference sequence, and a donor template nucleic acid that can be inserted at the cleavage site of the target nucleic acid sequence in the cell introduced by the CasX nuclease to modify the target nucleic acid sequence. Embodiments of these components are described below in this specification. In some embodiments, the present disclosure provides a gene editing pair of CasX and gNA of any of the embodiments described herein complexed as a ribonucleoprotein complex (RNP).In some embodiments, the present disclosure provides a method of modifying a gene of a cell encoding a protein involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response, wherein the gene is knocked down or knocked out from the expression of such protein.
[0007] Cells modified by the CasX:gNA system are useful, inter alia, for the preparation and use of immune cells that are also modified to express one or more chimeric antigen receptors (CARs) for immunotherapy applications, such as for the treatment of cancer or autoimmune diseases in a subject with a low likelihood of graft-versus-host disease (GVHD). Such cells are also engineered to reduce host-versus-graft complications. In other embodiments, the CasX-gNA system is used to knock in nucleic acids into cells encoding CARs and / or engineered T cell receptors (TCRs), wherein the CARs and / or TCRs include binding domains specific for tumor cell antigens including those listed hereinbelow. Such binding domains can be in the form of linear antibodies, single domain antibodies (sdAbs) such as VHHs, or single chain variable fragments (scFvs). Cells that can be used in the preparation of the modified cells include progenitor cells, hematopoietic stem cells, pluripotent stem cells, or immune cells selected from the group consisting of T cells, TREG cells, NK cells, B cells, macrophages, or dendritic cells.
[0008] In some aspects, the present disclosure provides polynucleotides and vectors encoding or comprising a CasX protein, gNA, a gene editing pair, or comprising the donor template nucleic acids described herein. In some embodiments, the vector is a viral vector such as an adeno-associated virus (AAV) vector or a lentiviral vector. In other embodiments, the vector is a non-viral particle such as a virus-like particle (VLP) or a nanoparticle.
[0009] In some embodiments, the present disclosure provides a method for modifying a target nucleic acid sequence in a cell population, comprising introducing into each cell of the population: a) a CasX:gNA system of any of the embodiments disclosed herein, or b) a nucleic acid of any of the embodiments disclosed herein, or c) a vector of any of the embodiments disclosed herein, d) a VLP of any of the embodiments disclosed herein, or e) a combination of two or more of (a)-(d) above, wherein the target nucleic acid sequence of the cell is modified by the CasX protein (e.g., single-stranded or double-stranded cleavage, or insertion, deletion, substitution, duplication, or inversion of one or more nucleotides in the target nucleic acid sequence).
[0010] In some embodiments, the present disclosure provides a cell population modified by an ex vivo method of modifying a target nucleic acid with a CasX:gNA system, vector, or VLP (or a combination thereof) of any of the embodiments described herein, wherein the expression of MHC class I molecules or T cell receptors, or the expression of proteins involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response in the modified cells is decreased or eliminated. In some embodiments, the present disclosure provides a cell population modified by an ex vivo method of modifying a target nucleic acid with a CasX:gNA system, vector, or VLP (or a combination thereof) of any of the embodiments described herein, wherein the modified cells express a CAR and / or TCR of any of the embodiments described herein at a detectable level.
[0011] In some embodiments, the present disclosure provides a method for providing anti-tumor immunity to a subject, comprising administering to the subject a therapeutically effective amount of one or more of the modified cells of any of the embodiments described herein.
[0012] In some embodiments, the present disclosure provides a method of treating a subject having a disease associated with the expression of a tumor antigen, the method comprising administering to the subject a therapeutically effective amount of one or more of the modified cells of any of the embodiments described herein.
[0013] In another aspect, the present disclosure provides a composition of immune cells modified by a gene editing pair of CasX and gNA, optionally a donor template and / or a polynucleotide encoding a CAR and / or a TCR, for use as a medicament for treating a subject having a disease associated with the expression of a tumor antigen. In the foregoing, CasX can be any of the CasX variants of any of the embodiments described herein (e.g., the sequences of Table 4), and gNA can be any of the gNA variants of any of the embodiments described herein (e.g., the sequences of Table 2). In other embodiments, the present disclosure provides a composition of cells modified by a vector comprising or encoding a gene editing pair of CasX and gNA, a donor template, and / or a polynucleotide encoding a CAR.
[0014] In some embodiments, the present disclosure provides a kit comprising a CasX:gNA system, a vector, or a VLP described herein, further comprising an excipient and a container.
[0015] In another aspect, the present disclosure provides a composition comprising a CasX:gNA system, a vector comprising or encoding a CasX:gNA system, a VLP comprising a CasX:gNA system, or a cell population edited using a CasX:gNA system for use as a medicament for treating a disease or disorder.
[0016] In another aspect, the present disclosure provides a CasX:gNA system, a composition comprising a CasX:gNA system, or a vector comprising or encoding a CasX:gNA system, a VLP comprising a CasX:gNA system, or a cell population edited using a CasX:gNA system for use in a method of treating a disease or disorder.
[0017] Incorporation by reference All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. The content of PCT / US2020 / 036505, filed on June 5, 2020, which discloses CasX variants and gNA variants, is incorporated herein in its entirety by reference.
Brief Description of the Drawings
[0018] The novel features of the present invention are specifically set forth in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description of exemplary embodiments in which the principles of the present invention are utilized, and to the accompanying drawings.
[0019]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35-1
Figure 35-2
Figure 35-3
Figure 35-4
Figure 35-5
Figure 35-6
Figure 35-7
Figure 35-8
Figure 35-9
Figure 35-10
Figure 35-11
Figure 35-12
Figure 35-13
Figure 35-14
Figure 35-15
Figure 35-16
Figure 35-17
Figure 35-18
Figure 35-19
Figure 35-20
Figure 35-21
Figure 35-22
Figure 35-23
Figure 35-24
Figure 35-25
Figure 35-26
Figure 35-27
Figure 35-28
Figure 35-29
Figure 35-30
Figure 36-1
Figure 36-2
Figure 36-3
Figure 36-4
Figure 36-5
Figure 36-6
Figure 36-7
Figure 36-8
Figure 36-9
Figure 36-10
Figure 36-11
Figure 36-12
Figure 36-13
Figure 36-14
Figure 36-15
Figure 36-16
Figure 36-17
Figure 36-18
Figure 36-19
Figure 36-20
Figure 36-21
Figure 36-22
Figure 36-23
Figure 36-24
Figure 36-25
Figure 36-26
Figure 36-27
Figure 36-28
Figure 36-29
Figure 36-30
Figure 36-31
Figure 36-32
Figure 36-33
Figure 36-34
Figure 36-35
Figure 36-36
Figure 36-37
Figure 36-38
Figure 36-39
Figure 36-40
Figure 36-41
Figure 36-42
Figure 36-43
Figure 36-44
Figure 36-45
Figure 36-46
Figure 36-47
Figure 36-48
Figure 36-49
Figure 36-50
Figure 36-51
Figure 36-52
Figure 36-53
Figure 36-54
Figure 36-55
Figure 36-56
Figure 36-57
Figure 36-58
Figure 36-59
Figure 36-60
Figure 36-61
Figure 36-62
Figure 36-63
Figure 36-64
Figure 36-65
Figure 36-66
Figure 36-67
Figure 36-68
Figure 36-69
Figure 36-70
Figure 36-71
Figure 36-72
Figure 36-73
Figure 36-74
Figure 36-75
Figure 36-76
Figure 36-77
Figure 36-78
Figure 36-79
Figure 36-80
Figure 36-81
Figure 36-82
Figure 36-83
Figure 36-84
Figure 36-85
Figure 36-86
Figure 36-87
Figure 36-88
Figure 36-89
Figure 36-90
Figure 36-91
Figure 36-92
Figure 36-93
Figure 37-1
Figure 37-2
Figure 37-3
Figure 37-4
Figure 37-5
Figure 37-6
Figure 37-7
Figure 37-8
Figure 37-9
Figure 37-10
Figure 37-11
Figure 37-12
Figure 37-13
Figure 37-14
Figure 37-15
Figure 37-16
Figure 37-17
Figure 37-18
Figure 37-19
Figure 37-20
Figure 37-21
Figure 37-22
Figure 37-23
Figure 37-24
Figure 37-25
Figure 37-26
Figure 37-27
Figure 37-28
Figure 37-29
Figure 37-30
Figure 37-31
Figure 37-32
Figure 37-33
Figure 37-34
Figure 37-35
Figure 37-36
Figure 37-37
Figure 37-38
Figure 37-39
Figure 37-40
Figure 37-41
Figure 37-42
Figure 37-43
Figure 37-44
Figure 37-45
Figure 37-46
Figure 37-47
Figure 37-48
Figure 37-49
Figure 37-50
Figure 37-51
Figure 37-52
Figure 37-53
Figure 37-54
Figure 37-55
Figure 37-56
Figure 37-57
Figure 37-58
Figure 37-59
Figure 37-60
Figure 37-61
Figure 37-62
Figure 37-63
Figure 37-64
Figure 37-65
Figure 37-66
Figure 37-67
Figure 37-68
Figure 37-69
Figure 37-70
Figure 37-71
Figure 37-72
Figure 37-73
Figure 37-74
Figure 37-75
Figure 37-76
Figure 37-77
Figure 37-78
Figure 37-79
Figure 37-80
Figure 37-81
Figure 37-82
Figure 37-83
Figure 37-84
Figure 37-85
Figure 37-86
Figure 37-87
Figure 37-88
Figure 37-89
Figure 37-90
Figure 37-91
Figure 37-92
Figure 37-93
Figure 37-94
Figure 37-95
Figure 37-96
Figure 37-97
Figure 37-98
Figure 37-99
Figure 37-100
Figure 37-101
Figure 37-102
Figure 37-103
Figure 37-104
Figure 37-105
Figure 37-106
Figure 37-107
Figure 37-108
Figure 37-109
Figure 37-110
Figure 37-111
Figure 37-112
Figure 37-113
Figure 37-114
Figure 37-115
Figure 37-116
Figure 37-117
Figure 37-118
Figure 37-119
Figure 37-120
Figure 37-121
Figure 37-122
Figure 37-123
Figure 37-124
Figure 37-125
Figure 37-126
Figure 37-127
Figure 37-128
Figure 37-129
Figure 37-130
Figure 37-131
Figure 37-132
Figure 37-133
Figure 37-134
Figure 37-135
Figure 37-136
Figure 37-137
Figure 37-138
Figure 37-139
Figure 37-140
Figure 37-141
Figure 37-142
Figure 37-143
Figure 37-144
Figure 37-145
Figure 37-146
Figure 37-147
Figure 37-148
Figure 37-149
Figure 37-150
Figure 37-151
Figure 37-152
Figure 37-153
Figure 37-154
Figure 37-155
Figure 37-156
Figure 37-157
Figure 37-158
Figure 37-159
Figure 37-160
Figure 37-161
Figure 37-162
Figure 37-163
Figure 37-164
Mode for Carrying Out the Invention
[0020] Exemplary embodiments are shown and described herein, but it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Without departing from the present invention, those skilled in the art will envision numerous variations, modifications, and substitutions. It should be understood that various alternatives to the embodiments described herein may be used in practicing the present disclosure. The following claims define the scope of the present invention and are intended to cover methods and structures within the scope of the claims and their equivalents.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, the preferred methods and materials are described below. In case of conflict, the present specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting. Without departing from the scope of the present invention, those of ordinary skill in the art will envision numerous variations, modifications, and substitutions.
[0022] Definitions As used interchangeably herein, the terms "polynucleotide" and "nucleic acid" refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. Thus, the terms "polynucleotide" and "nucleic acid" include single-stranded DNA, double-stranded DNA, multi-stranded DNA, single-stranded RNA, double-stranded RNA, multi-stranded RNA, genomic DNA, cDNA, DNA-RNA hybrids, and polymers comprising purine and pyrimidine bases, or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases.
[0023] "Hybridizable" or "complementary" are used interchangeably and mean that a nucleic acid (e.g., RNA, DNA) contains a nucleotide sequence that non-covalently binds, i.e., forms Watson-Crick base pairs and / or G / U base pairs, and is capable of "annealing" or "hybridizing" to another nucleic acid in a sequence-specific, antiparallel manner under appropriate in vitro and / or in vivo conditions of temperature and solution ionic strength (i.e., the nucleic acid specifically binds to a complementary nucleic acid). It is understood that the sequence of a polynucleotide need not be 100% complementary to its target nucleic acid sequence in order to be specifically hybridizable, and that the sequence of a polynucleotide can have at least about 70%, at least about 80%, or at least about 90%, or at least about 95% sequence identity to the target nucleic acid sequence and still be able to hybridize to the target nucleic acid sequence. Further, a polynucleotide can hybridize over one or more segments such that intervening or adjacent segments are not involved in the hybridization event (e.g., loop structures or hairpin structures, "bulges", etc.).
[0024] For purposes of the present disclosure, a "gene" includes a DNA region that encodes a gene product (e.g., a protein, RNA), as well as all DNA regions that regulate the production of the gene product, whether or not such regulatory sequences are adjacent to the coding and / or transcribed sequences. Thus, a gene can include regulatory element sequences, including, but not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, and locus control regions. The coding sequence encodes the gene product upon transcription or transcription and translation, and the coding sequences of the present disclosure can include fragments and need not include full-length open reading frames. A gene can include both the transcribed strand, e.g., the strand that includes the coding sequence, and the complementary strand.
[0025] The term "downstream" refers to a nucleotide sequence that is located 3' to a reference nucleotide sequence. In certain embodiments, the downstream nucleotide sequence is related to the sequence following the transcription start point. For example, the translation start codon of a gene is located downstream of the transcription start site.
[0026] The term "upstream" refers to a nucleotide sequence that is located 5' to a reference nucleotide sequence. In certain embodiments, the upstream nucleotide sequence is related to the sequence located 5' to the coding region or the transcription start point. For example, most promoters are located upstream of the transcription start site.
[0027] The term "regulatory element" is used interchangeably herein with the term "regulatory sequence" and is intended to include promoters, enhancers, and other expression regulatory elements (e.g., transcription termination signals such as polyadenylation signals and polyU sequences). Exemplary regulatory elements include transcription promoters such as a single transcript, metallothionein, transcription enhancer elements, transcription termination signals, polyadenylation sequences, sequences for optimization of translation initiation, and sequences that enable translation of multiple genes from a translation termination sequence, such as CMV, CMV+intron A, SV40, RSV, HIV-Ltr, elongation factor 1 alpha (EF1α), MMLV-ltr, internal ribosome entry site (IRES), or P2A peptide, but are not limited thereto. It will be understood that the selection of appropriate regulatory elements depends on the encoded component to be expressed (e.g., a protein or RNA), or whether the nucleic acid requires different polymerases, or whether it contains multiple components not intended to be expressed as a fusion protein.
[0028] The term "promoter" includes an RNA polymerase binding site, a transcription start site, a TATA box, and / or a B recognition element, and refers to a DNA sequence that includes an associated transcribable polynucleotide sequence and / or gene (or transgene) and aids or promotes the transcription and expression thereof. A promoter can be produced synthetically or can be derived from a known or naturally occurring promoter sequence or another promoter sequence. A promoter can be proximal or distal to the gene being transcribed. A promoter can also include a chimeric promoter that includes a combination of two or more heterologous sequences to confer certain properties. The promoters of the present disclosure can include variants of promoter sequences that are similar but not identical in composition to known or other promoter sequences provided herein. Promoters can be classified according to criteria related to the expression pattern of an associated coding or transcribable sequence or gene operably linked to a promoter, such as constitutive, developmental, tissue-specific, inducible, etc.
[0029] The term "enhancer" refers to a regulatory DNA sequence that regulates the expression of an associated gene when a specific protein called a transcription factor binds thereto. An enhancer can be located in an intron of a gene, or on the 5' or 3' side of the coding sequence of a gene. An enhancer can be proximal to the gene (i.e., within dozens or hundreds of base pairs (bp) of the promoter) or distal to the gene (i.e., thousands, hundreds of thousands, or even millions of bp away from the promoter). A single gene can be regulated by two or more enhancers, all of which are assumed to be within the scope of the present disclosure.
[0030] As used herein, "recombinant" means the product of various combinations of cloning, restriction, and / or ligation steps that result in a construct in which a particular nucleic acid (DNA or RNA) has a structural coding or non-coding sequence distinguishable from the endogenous nucleic acids found in nature. In general, DNA sequences encoding structural coding sequences can be assembled from cDNA fragments and short oligonucleotide linkers, or from a series of synthetic oligonucleotides, to provide synthetic nucleic acids that can be contained in cells or expressed from recombinant transcription units contained in cell-free transcription and translation systems. Such sequences can be provided in the form of open reading frames that are not interrupted by internal non-translated sequences or introns typically present in eukaryotic genes. Genomic DNA containing related sequences can also be used in the formation of recombinant genes or transcription units. Sequences of non-translated DNA may be present on the 5' or 3' side of the open reading frame, and such sequences do not interfere with the manipulation or expression of the coding region and, in fact, can play a role in regulating the production of the desired product by various mechanisms (see "enhancers" and "promoters" above).
[0031] The terms "recombinant polynucleotide" or "recombinant nucleic acid" refer to something that does not occur naturally, for example, something made by the artificial combination of two otherwise separated sequence segments through human intervention. This artificial combination is often achieved by either chemical synthesis means or the artificial manipulation of isolated segments of nucleic acids, such as genetic engineering techniques. Such artificial combinations are usually done to replace codons with redundant codons encoding the same or conserved amino acids, often while introducing or removing sequence recognition sites. Alternatively, this artificial combination is done to join together nucleic acid segments having desired functions to generate a combination of desired functions. This artificial combination is often achieved by either chemical synthesis means or the artificial manipulation of isolated segments of nucleic acids, such as genetic engineering techniques.
[0032] Similarly, the terms "recombinant polypeptide" or "recombinant protein" refer to a polypeptide or protein that is not naturally occurring and is made, for example, by an artificial combination of two separately isolated amino acid sequence segments through human intervention. Thus, for example, a polypeptide containing heterologous amino acid sequences is recombinant.
[0033] As used herein, the term "contacting" means establishing a physical association between two or more entities. For example, contacting a target nucleic acid sequence with a guide nucleic acid means that the target nucleic acid sequence and the guide nucleic acid are made to share a physical association, such that, for example, they can hybridize if their sequences share sequence similarity.
[0034] "Dissociation constant" or "K d " are used interchangeably and mean the affinity between a ligand "L" and a protein "P", i.e., how closely the ligand binds to a particular protein. This can be calculated using the equation K d = [L][P] / [LP], where [P], [L], and [LP] represent the molar concentrations of the protein, ligand, and complex, respectively.
[0035] The term "knockout" refers to the removal or expression of a gene. For example, a gene can be knocked out by either a deletion or addition of a nucleotide sequence that leads to a disruption of the reading frame. As another example, a gene can be knocked out by replacing a portion of the gene with an unrelated sequence. The term "knockdown" as used herein refers to a decrease in the expression of a gene or its gene product. As a result of gene knockdown, the activity or function of a protein can be attenuated, or the protein level can be decreased or eliminated.
[0036] As used herein, "homology-directed repair" (HDR) refers to a form of DNA repair that occurs during the repair of double-strand breaks in cells. This process requires nucleotide sequence homology and uses a donor template to repair or knockout target DNA, resulting in the transfer of genetic information from the donor to the target. If the donor template differs from the target DNA sequence and some or all of the donor template sequence is incorporated into the target DNA, homology-directed repair can result in modification of the target sequence's sequence by insertion, deletion, or mutation.
[0037] As used herein, "non-homologous end joining" (NHEJ) refers to the repair of double-strand breaks in DNA by direct ligation of the cut ends to each other, without the need for a homologous template (in contrast to homology-directed repair, which requires a homologous sequence to induce repair). NHEJ often results in the loss (deletion) of nucleotide sequence near the double-strand break site.
[0038] As used herein, "microhomology-mediated end joining" (MMEJ) refers to a mutagenic DSB repair mechanism that is always associated with deletions adjacent to the break site, without the need for a homologous template (in contrast to homology-directed repair, which requires a homologous sequence to induce repair). MMEJ often results in the loss (deletion) of nucleotide sequence near the double-strand break site.
[0039] A polynucleotide or polypeptide has a certain "sequence similarity" or "sequence identity" percentage with another polynucleotide or polypeptide, which means that when aligned, the percentage of bases or amino acids is the same and they are in the same relative positions when the two sequences are compared. Sequence similarity (sometimes referred to as percent similarity, percent identity, or homology) can be determined in several different ways. To determine sequence similarity, the sequences can be aligned using methods and computer programs known in the art, including BLAST available on the World Wide Web at ncbi.nlm.nih.gov / BLAST. The percent complementarity between specific stretches of nucleic acid sequences within a nucleic acid can be determined using any convenient method. Examples of methods include the BLAST program (Basic Local Alignment Search Tool) and the PowerBLAST program (Altschul et al., J. Mol. Biol., 1990, 215, 403-410, Zhang and Madden, Genome Res., 1997, 7, 649-656), or the use of the Gap program using the Smith-Waterman algorithm (Adv. Appl. Math., 1981, 2, 482-489) (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.), for example, using default settings.
[0040] The terms "polypeptide" and "protein" are used interchangeably herein and refer to a polymeric form of amino acids of any length that can include encoded and non-encoded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having modified peptide backbones. The term includes, but is not limited to, fusion proteins having heterologous amino acid sequences.
[0041] The term "vector" or "expression vector" refers to a replicon such as a plasmid, phage, virus, or cosmid, to which another DNA segment, i.e., an "insert", can be ligated to effect replication or expression of the ligated segment in a cell.
[0042] As used herein, the terms "naturally occurring", "unmodified", or "wild-type" when applied to a nucleic acid, polypeptide, cell, or organism refer to a nucleic acid, polypeptide, cell, or organism as found in nature.
[0043] As used herein, "mutation" refers to an insertion, deletion, substitution, duplication, or inversion of one or more amino acids or nucleotides as compared to a reference amino acid sequence or reference nucleotide sequence.
[0044] As used herein, the term "isolated" is intended to describe a polynucleotide, polypeptide, or cell that exists in an environment different from that in which it naturally occurs. An isolated genetically modified host cell may be present in a mixed population of genetically modified host cells.
[0045] As used herein, "host cell" means a cell derived from a eukaryotic cell, prokaryotic cell, or multicellular organism (e.g., cell line), and these eukaryotic or prokaryotic cells are used as recipients of nucleic acids (e.g., expression vectors) and include the progeny of the original cell that have been genetically modified by the nucleic acid. It is understood that the progeny of a single cell may not necessarily be identical in morphology or genomic or total DNA complement to the original parent due to natural, accidental, or intentional mutations. A "recombinant host cell" (also referred to as a "genetically modified host cell") is a host cell into which a heterologous nucleic acid, e.g., an expression vector, has been introduced.
[0046] The term "conservative amino acid substitution" refers to the interchangeability in proteins of amino acid residues having similar side chains. For example, the group of amino acids having aliphatic side chains consists of glycine, alanine, valine, leucine, and isoleucine; the group of amino acids having aliphatic hydroxyl side chains consists of serine and threonine; the group of amino acids having amide-containing side chains consists of asparagine and glutamine; the group of amino acids having aromatic side chains consists of phenylalanine, tyrosine, and tryptophan; the group of amino acids having basic side chains consists of lysine, arginine, and histidine; and the group of amino acids having sulfur-containing side chains consists of cysteine and methionine. Exemplary conservative amino acid substituents are valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine.
[0047] The term "chimeric antigen receptor" or "CAR", when expressed in a cell, comprises at least two domains that confer on the cell specificity for a target antigen, or a target cell having the target antigen, typically a diseased cell having a specific disease-related antigen. In some embodiments, the CAR comprises at least an extracellular antigen-binding domain (e.g., an scFv having binding specificity for a protein involved in a disease (e.g., cancer)), a transmembrane domain, and a cytoplasmic signaling domain (also referred to herein as an "intracellular signaling domain") that comprises one or more stimulatory and / or costimulatory molecules and / or a functional signaling domain derived therefrom as provided below. In some aspects, the set of polypeptides are adjacent to each other. A portion of the CAR of the present disclosure that comprises the antigen-binding domain may be present in various forms expressed as part of a polypeptide chain to which the antigen-binding domain is adjacent, including, for example, a single-domain antibody fragment (sdAb), a single-chain antibody (scFv), a humanized antibody or a bispecific antibody (Harlow et al., 1999, Using Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory Press, NY, Harlow et al., 1989, Antibodies: A Laboratory Manual, Cold Spring Harbor, N.Y., Houston et al., 1988, Proc. Natl. Acad. Sci. USA 85:5879-5883, Bird et al., 1988, Science 242:423-426), and may further comprise a hinge region, e.g., the hinge region of an immunoglobulin molecule, and a spacer that provides mobility to the receptor. The hinge, spacer, and transmembrane domain link the scFv to the activation domain and anchor the CAR to the T cell membrane. In some embodiments, the CAR compositions of the present disclosure comprise an antigen-binding domain. In further embodiments, the CAR comprises an antibody fragment that comprises an scFv.The exact amino acid sequence boundaries of a given CDR can be determined using any of several well-known schemes, including the scheme described by Kabat et al. (1991), "Sequences of Proteins of Immunological Interest," 5th Ed. Public Health Service, National Institutes of Health, Bethesda, Md. (the "Kabat" numbering scheme), the scheme described by Al-Lazikani et al., (1997) JMB 273, 927-948 (the "Chothia" numbering scheme), or a combination thereof.
[0048] The term "T cell receptor (TCR)" refers to a protein complex found on the surface of T cells that is involved in the recognition of peptide antigens bound to major histocompatibility complex (MHC) molecules. The TCR consists of multiple subunits, including the TCR alpha chain and the TCR beta chain (encoded by TRAC or TCRA, and TBRC1 or TCRB, respectively), and within these chains are complementarity-determining regions (CDRs) that determine the antigen to which the TCR binds. Additional subunits include CD-epsilon (CD3E), CD3-delta (CD3D), CD3-gamma (CD3G), and CD3-zeta (CD3Z). The extracellular domains of the TCR alpha and TCR beta subunits form the antigen-binding site of the native TCR. The CDRs of the extracellular domain of the TCR are the antigen-binding sections, and due to their diverse recognition capabilities, they are able to efficiently protect against foreign antigens or diseased cells and bring about an optimal immune response. When the TCR binds appropriately to an antigen, a conformational change in the associated CD3 chains is induced, which, together with other factors, initiates a signal transduction process and T cell activation.
[0049] As used herein, an "engineered TCR" refers to a TCR that has been engineered to include an antigen-binding domain having specificity for a target antigen, or a target cell having the target antigen, typically a diseased cell having a specific disease-related antigen. For example, an engineered TCR can include an antigen-binding domain fused to either the TCR alpha subunit or the TCR beta subunit of the TCR, or a combination thereof. For example, any antigen-binding domain, including a single domain antibody fragment (sdAb), single chain antibody (scFv), humanized antibody, or bispecific antibody, can be used with the engineered TCRs described herein. In addition to one or more subunits fused to the antigen-binding domain, an engineered TCR can also include wild-type subunits encoded by the genome of the cell. For example, an engineered TCR can include either the TCR alpha or TCR beta subunit of the TCR, as well as an antigen-binding domain fused to the wild-type CD3-delta, CD3-gamma, CD3-epsilon, and CD3-zeta subunits.
[0050] A "signaling domain" refers to the functional portion of a protein that acts by transmitting information intracellularly to regulate cell activity via a defined signaling pathway, either by generating a second messenger or by functioning as an effector in response to such a messenger.
[0051] The "intracellular signaling domain" refers to the intracellular portion of a molecule and, as used herein, is a component of a CAR. Examples of T cell-derived signaling domains are polypeptides derived from a group consisting of CD247 molecule (CD3-zeta, or CD3Z), CD27 molecule (CD27), CD28 molecule (CD28), TNF receptor superfamily member 9 (4-1BB, or 41BB), inducible T cell co-stimulator (ICOS), TNF receptor superfamily member 4 (OX40), or combinations thereof. The intracellular signaling domain generates signals that promote the immune effector functions of CAR-containing cells, such as CAR-T cells. For example, examples of immune effector functions in CAR-T cells include cytolytic activity and helper activity, including the secretion of cytokines. The intracellular signaling domain can include a signaling motif known as an immunoreceptor tyrosine-based activation motif or ITAM. Examples of primary cytoplasmic signaling sequences containing ITAM include, but are not limited to, those derived from CD3 zeta, the Fc fragment of the Ig of the IgE receptor (common FcR gamma, or FCER1G), the Fc fragment of the IgG receptor IIa (Fc gamma RIIa, or FCGR2A), Fc receptor gamma RIIB, CD3g molecule (CD3 gamma, or CD3G), CD3d molecule (CD3 delta, or CD3D), CD3e molecule (CD3 epsilon, or CD3E), CD79a, CD79b, DAP10, and DAP12.
[0052] The term "zeta", or alternatively "zeta chain", "CD3-zeta", or "TCR-zeta", is defined as the protein provided as GenBank accession number BAG36664.1, or as equivalent residues from non-human species, such as mouse, rodent, or non-human primate, and "zeta stimulatory domain", or alternatively "CD3-zeta stimulatory domain" or "TCR-zeta stimulatory domain", is defined as the amino acid residues from the cytoplasmic domain of the zeta chain that are sufficient to functionally transmit the initial signals required for T cell activation, or a functional derivative thereof. In some embodiments, the cytoplasmic domain of zeta comprises residues 52-164 of GenBank accession number BAG36664.1, or equivalent residues from non-human species that are functional orthologs thereof.
[0053] As used herein, "a protein involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response" refers to extracellular, transmembrane, and intracellular proteins or glycoproteins involved in antigen processing, presentation, recognition, and / or response. In some cases, the protein or glycoprotein is expressed on the surface of a cell and can conveniently serve as a marker for a particular cell type. For example, T cell and B cell surface proteins identify their lineage and stage in the differentiation process. In some cases, a protein involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response is a receptor having binding affinity for a ligand.
[0054] A "tumor antigen" is expressed either globally or as a fragment (e.g., an MHC peptide) on the surface of a cancer cell and is useful as a preferential target for immune cells to the cancer cell. In some embodiments, the tumor antigen is a marker expressed by both normal and cancer cells, such as CD19 on B cells. In some embodiments, the tumor antigen is a cell surface molecule that is overexpressed in cancer cells compared to normal cells.
[0055] As used herein, the term "antibody" encompasses various antibody structures including, but not limited to, monoclonal antibodies, polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies), nanobodies, single domain antibodies such as VHH antibodies, and antibody fragments, so long as they exhibit the desired antigen-binding activity or immunological activity. Antibodies represent a large family of molecules including several types of molecules such as IgD, IgG, IgA, IgM, and IgE.
[0056] A "humanized" antibody refers to an antibody that contains amino acid residues derived from non-human complementarity-determining regions (CDRs) and amino acid residues derived from human framework regions (FRs). Typically, a humanized antibody contains substantially all of the variable domain, with all or substantially all of the CDRs corresponding to the variable domain of a non-human antibody (which may include amino acid substitutions), and all or substantially all of the FRs corresponding to the variable domain of a human antibody.
[0057] As used herein, the term "monoclonal antibody" refers to an antibody obtained from a substantially homogeneous population of antibodies, which population is identical and / or binds to the same epitope. Thus, the modifier "monoclonal" indicates the characteristic that the antibody is obtained from a substantially homogeneous population of antibodies and should not be construed to require production of the antibody by any particular method.
[0058] As used herein, an "antigen-binding domain" refers to the immunologically active portion of a molecule that contains an antigen-binding site that specifically binds to an antigen (i.e., "immunoreacts with it"). An antigen-binding domain is "specifically binds to" or "specific for" an antigen if it binds with a higher affinity or binding strength than it binds to other reference antigens, including polypeptides or other substances. Examples of proteins that contain an antigen-binding domain include, but are not limited to, Fv, Fab, Fab′, Fab′-SH, F(ab′)2, diabody, linear antibody (see US 5,641,870), single domain antibody, single domain camelid antibody, single-chain fragment variable (scFv) antibody molecule, or any polypeptide chain-containing molecular structure having a specific shape that fits and recognizes and binds to an epitope.
[0059] "scFv" or "single-chain fragment variable" are used interchangeably herein and refer to an antibody fragment format that consists of the variable heavy chain ("VH") and variable light chain ("VL") of an antibody, or two copies of the VH chain or VL chain, joined together by a short flexible peptide linker that allows the scFv to form a structure that is desirable for antigen binding. An scFv is a fusion protein of the variable heavy chain (VH) and variable light chain (VL) of an immunoglobulin, each of which may be in either the VH-VL or VL-VH order and typically contains complementarity-determining regions (CDRs) joined by a linker.
[0060] The term "4-1BB" refers to a member of the TNF-R superfamily having the amino acid sequence provided as GenBank accession number AAA62478.2, or equivalent residues from non-human species, and "4-1BB co-stimulatory domain" is defined as amino acid residues 214-255 of GenBank accession number AAA62478.2, or equivalent residues from non-human species.
[0061] The term "immune effector cell" refers to a cell involved in promoting an immune response, such as an immune effector response. Examples of immune effector cells include T cells such as helper T cells and cytotoxic T cells, gamma-delta T cells, tumor-infiltrating lymphocytes, NK cells, B cells, monocytes, macrophages, or dendritic cells.
[0062] The term "immune effector function" or "immune effector response" refers to the function or response of immune effector cells, for example, that enhances or promotes an immune attack on target cells. In the context of the present disclosure, the immune effector function or response refers to the property of T cells or NK cells that promotes the killing of target cells or the inhibition of the growth or proliferation of target cells.
[0063] As used herein, the terms "treatment" or "treating" are used interchangeably herein and refer to an approach for obtaining beneficial or desired results, including, but not limited to, therapeutic and / or prophylactic benefits. Therapeutic benefit means eradication or amelioration of the underlying disorder or disease being treated. A therapeutic benefit can also be achieved by eradication or amelioration of one or more of the symptoms, or by improvement of one or more clinical parameters associated with the underlying disease such that an improvement is observed in the subject, even though the subject may still be afflicted with the underlying disorder.
[0064] As used herein, the terms "therapeutically effective amount" and "therapeutically effective dose" refer to an amount of a drug or biological agent, alone or as part of a composition, that when administered to a subject, such as a human or experimental animal, in a single dose or in multiple doses, can produce some detectable and beneficial effect on any symptom, aspect, measured parameter, or characteristic of a medical condition or state. Such an effect need not be absolutely beneficial.
[0065] As used herein, the term "administering" means a method of giving a subject a dosage of a compound (e.g., a composition of the present disclosure) or a composition (e.g., a pharmaceutical composition).
[0066] The "subject" is a mammal. Mammals include, but are not limited to, domestic animals, non-human primates, humans, rabbits, mice, rats, and other rodents.
[0067] I. General Methods The practice of the present invention uses conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, unless otherwise indicated. These can be found in standard textbooks such as Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., HaRBor Laboratory Press 2001), Short Protocols in Molecular Biology, 4th Ed. (Ausubel et al. eds., John Wiley & Sons 1999), Protein Methods (Bollag et al., John Wiley & Sons 1996), Nonviral Vectors for Gene Therapy (Wagner et al. eds., Academic Press 1999), Viral Vectors (Kaplift & Loewy eds., Academic Press 1995), Immunology Methods Manual (I. Lefkovits ed., Academic Press 1997), and Cell and Tissue Culture: Laboratory Procedures in Biotechnology (Doyle & Griffiths, John Wiley & Sons 1998), the disclosures of which are incorporated herein by reference.
[0068] When a range of values is provided, endpoints are included, and all values between the upper and lower limits of the range are included down to one-tenth of the unit of the lower limit, unless the context clearly indicates otherwise, and all other recited values and values between the recited values of the recited range are understood to be included. The upper and lower limits of these smaller ranges can independently be included in the smaller ranges and are included, subject to any specifically excluded limits in the recited range. When the recited range includes one or both of the limits, ranges excluding either or both of those included limits are also included.
[0069] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. All publications mentioned herein are incorporated herein by reference for the purpose of disclosing and describing the methods and / or materials related to what is cited in the publications.
[0070] It should be noted that, as used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise.
[0071] It will be understood that certain features of the disclosure that are described in the context of separate embodiments may sometimes be provided in combination in a single embodiment. In other instances, various features of the invention that are described in the context of a single embodiment for brevity may be provided separately or in any suitable sub-combination. All combinations of embodiments related to the invention are specifically included by the invention and are intended to be disclosed herein as if each and every combination were individually and explicitly disclosed. Additionally, all sub-combinations of various embodiments and their elements are specifically included by the invention and are disclosed herein as if each and every such sub-combination were individually and explicitly disclosed herein.
[0072] II. Systems for gene editing of proteins involved in antigen processing, presentation, recognition, and / or response In a first aspect, the present disclosure provides a system comprising a CRISPR nuclease and one or more guide nucleic acids (gNAs) that are useful for genome editing of eukaryotic cells. In some embodiments, the CRISPR nuclease is selected from the group consisting of Cas9, Cas12a, Cas12b, Cas12c, Cas12d (CasY), CasX, Cas13a, Cas13b, Cas13c, Cas13d, CasX, CasY, Cas14, Cpfl, C2cl, Csn2, and Cas Phi. In some embodiments, the CRISPR nuclease is a type V CRISPR nuclease. In some embodiments, the present disclosure provides a CasX:gNA system comprising a CasX protein and one or more guide nucleic acids (gNAs) that are specially designed to modify the target nucleic acid sequences of one or more cellular genes involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response. The gNA and CasX protein of the present disclosure form a complex, herein referred to as a ribonucleoprotein (RNP) complex, and can bind by non-covalent interactions. The use of pre-complexed CasX:gNA confers advantages in the delivery of the components of the system to cells or target nucleic acid sequences for editing of the target nucleic acid sequences. In the RNP, the gNA can provide target specificity to the complex by including a target sequence (or "spacer") having a nucleotide sequence complementary to the sequence of the target nucleic acid sequence, while the CasX protein of the pre-complexed CasX:gNA is guided (e.g., stabilized) to a target site within the target nucleic acid sequence (e.g., the B2M or TRAC gene to be modified) by association with the guide NA, providing site-specific activity. The CasX protein of the complex provides the site-specific activity of the complex, such as cleavage or nicking of the target sequence by the CasX protein, and / or the activity provided by the fusion partner in the case of a chimeric CasX protein. Additionally, the present disclosure provides methods useful for modifying the target nucleic acid sequences of a cell population to introduce or regulate the expression of one or more proteins involved in antigen processing, presentation, recognition, and / or response using the CasX:gNA system.Such a modified cell population in which proteins involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response are downregulated or eliminated is useful for immunotherapy. The CasX:gNA system of the present disclosure comprises one or more CasX proteins, one or more guide nucleic acids (gNA), and optionally one or more donor template nucleic acids comprising a nucleic acid encoding a modification of a protein involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response, which nucleic acid comprises one or more nucleotide deletions, insertions, or mutations as compared to a genomic nucleic acid sequence encoding the protein or its regulatory elements for gene function knockdown / knockout. In some embodiments, the donor polynucleotide comprises at least about 10, at least about 50, at least about 100, or at least about 200, or at least about 300, or at least about 400, or at least about 500, or at least about 600, or at least about 700, or at least about 800, or at least about 900, or at least about 1000, or at least about 10,000, or at least about 15,000 nucleotides of all or a portion of the target nucleic acid sequence of the cell gene to be modified. In other embodiments, the donor polynucleotide comprises at least about 10 to about 10,000 nucleotides, or at least about 100 to about 8000 nucleotides, or at least about 400 to about 6000 nucleotides, or at least about 600 to about 4000 nucleotides, or at least about 1000 to about 2000 nucleotides of the cell gene to be modified. In some embodiments, the donor template is a single-stranded DNA template or a single-stranded RNA template. In other embodiments, the donor template is a double-stranded DNA template.
[0073] In other embodiments, the present disclosure provides a polynucleotide encoding a chimeric antigen receptor (CAR) having binding specificity for a disease antigen, optionally a tumor cell antigen, which can be introduced into a cell modified such that the modified cell can express the CAR. In other embodiments, the present disclosure provides a polynucleotide encoding an engineered T cell receptor (TCR) having binding specificity for a disease antigen, optionally a tumor cell antigen, which can be introduced into a cell modified such that the modified cell can express the TCR.
[0074] The CasX:gNA system is useful for the treatment of subjects having certain diseases or conditions including cancer, autoimmune diseases, and transplant rejection. The components of the CasX:gNA system, as well as their use in the editing of intracellular target nucleic acids for the modification of one or more proteins involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response, as well as the use of polynucleotides encoding CARs and one or more engineered TCR subunits are described herein. The CasX:gNA system and polynucleotides described herein are useful for generating a modified cell population that efficiently kills target cells associated with diseases such as cancer, autoimmune diseases, and transplant rejection. Further, the modified cell population can be used to confer immunity to a subject having such a disease.
[0075] III. Guide Nucleic Acids of Systems for Gene Editing In another aspect, the present disclosure provides a guide nucleic acid (gNA) comprising a target sequence complementary to a target nucleic acid sequence in the target strand of a gene encoding a protein involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response, wherein the gNA can form a complex with a CRISPR protein specific for a protospacer adjacent motif (PAM) sequence comprising a TC motif in the complementary non-target strand, and the PAM sequence is located at one nucleotide 5' of the sequence in the non-target strand complementary to the target nucleic acid sequence in the target strand, with respect to the gNA.
[0076] In some embodiments, the present disclosure relates to guide nucleic acids (gNAs) utilized in a CasX:gNA system useful for genome editing of eukaryotic cells. The present disclosure provides a specially designed guide nucleic acid (a “gNA”) wherein the target sequence of the gNA (or spacer, described more fully below) is complementary to (and thus capable of hybridizing to) a target nucleic acid sequence when used as a component of a gene editing CasX:gNA system. In some embodiments, multiple gNAs are envisioned to be delivered in a CasX:gNA system for modification of a target nucleic acid sequence. For example, when knockdown / knockout of a protein-coding gene is desired, a pair of gNAs can be used to bind and cleave at two different sites within the gene.
[0077] The present disclosure provides a specially designed guide nucleic acid (a “gNA”) having a target sequence that is complementary to (and thus capable of hybridizing to) a target nucleic acid when used as a component of a gene editing CasX:gNA system. Representative but non-limiting examples of target sequences for target nucleic acid sequences of cellular genes encoding proteins involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response are presented in Tables 3A, 3B, and 3C (Tables 3A, 3B, and 3C are provided as Figures 35-37). In some embodiments, multiple gNAs are envisioned to be delivered in a CasX:gNA system for modification of a target nucleic acid sequence. For example, when knockdown / knockout of a protein-coding gene is desired, a pair of gNAs having target sequences for different or overlapping regions of the target nucleic acid sequence can be used to bind and cleave CasX at two different or overlapping sites within or proximal to the gene, which can then be edited by non-homologous end joining (NHEJ), homology-directed repair (HDR, which can include insertion of a donor template to replace all or a portion of an intron), homology-independent targeted integration (HITI), microhomology-mediated end joining (MMEJ), single-strand annealing (SSA), or base excision repair (BER).
[0078] a. Reference gNA and gNA variants In some embodiments, the gNA of the present disclosure comprises the sequence of a naturally occurring gNA (“reference gNA”). In other cases, the reference gNA of the present disclosure may be subjected to one or more mutagenesis methods described herein, such as deep mutational evolution (DME), deep mutational scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping, to generate one or more gNA variants having improved or altered properties compared to the reference gNA. gNA variants include, for example, variants that include one or more exogenous sequences fused to either the 5′ or 3′ end, or inserted internally. The activity of the reference gNA is used as a benchmark against which the activity of the gNA variant is compared, thereby allowing measurement of the improvement of the function or other properties of the gNA variant. In other embodiments, the reference gNA may be subjected to one or more intentional targeted mutations to generate a gNA variant, e.g., a rationally designed variant. As used herein, the terms gNA, gRNA, and gDNA include naturally occurring molecules and also sequence variants. In some embodiments, the gNA is a deoxyribonucleic acid molecule (“gDNA”), in some embodiments, the gNA is a ribonucleic acid molecule (“gRNA”), and in other embodiments, the gNA is chimeric and includes both DNA and RNA.
[0079] The target sequence of the gNA can bind to a coding sequence, the complement of the coding sequence, a target nucleic acid sequence including a non-coding sequence, and a regulatory element. The gNA scaffold (or "protein-binding sequence") interacts with (e.g., binds to) the CasX protein to form an RNP (more fully described below). In some embodiments, the target sequence and the scaffold each include a complementary stretch of nucleotides that hybridize to each other to form a double-stranded duplex (dsRNA duplex in the case of dgRNA). Site-specific binding and / or cleavage of a target nucleic acid sequence (e.g., genomic DNA) by the CasX protein can occur at one or more positions (e.g., the sequence of the target nucleic acid) determined by base-pairing complementarity between the target sequence of the gNA and the target nucleic acid sequence. Thus, for example, the gNA of the present disclosure is complementary to a nucleic acid in a eukaryotic cell adjacent to a sequence complementary to a TC PAM motif or PAM sequence, e.g., ATC, CTC, GTC, or TTC, and thus has a sequence and / or its regulatory sequence that can hybridize thereto and is involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response of a protein in a eukaryotic nucleic acid (e.g., a eukaryotic chromosome, chromosomal sequence, eukaryotic RNA, etc.).
[0080] In the context of nucleic acids, cleavage refers to the breaking of the covalent backbone of either a DNA or RNA nucleic acid molecule. Cleavage can be initiated by a variety of methods including, but not limited to, enzymatic or chemical hydrolysis of phosphodiester bonds. Both single-strand cleavage and double-strand cleavage are possible, and double-strand cleavage can occur as a result of two separate single-strand cleavage events. DNA cleavage can generate either blunt ends or staggered ends.
[0081] In some embodiments, the present disclosure provides a gene editing pair of CasX and gNA of any of the embodiments described herein that can be bound together prior to use in gene editing and are thus “pre-complexed” as a ribonucleoprotein complex (RNP). The use of pre-complexed RNPs confers advantages for delivery of the system components to cells or to the target nucleic acid sequence for editing of the target nucleic acid sequence. The CasX protein of the RNP provides site-specific activity that is guided (e.g., stabilized) to a target site within the target nucleic acid sequence by association with a guide RNA that includes a target sequence that can hybridize to the target nucleic acid sequence.
[0082] In some embodiments where the gNA is a gRNA, the term "targeter" or "targeter RNA" is used herein to refer to the crRNA-like molecule (crRNA: "CRISPR RNA") of the CasX dual guide RNA (and thus, for example, of the CasX single guide RNA when the "activator" and "targeter" are linked together by intervening nucleotides). Thus, for example, a CasX guide RNA (dgRNA or sgRNA) includes a guide sequence and a duplex-forming segment of the crRNA, which may also be referred to as the crRNA repeat. Since the sequence of the guide sequence hybridizes to the target nucleic acid sequence, the targeter can be modified by the user to hybridize to a specific target nucleic acid sequence, provided that the position of the PAM is taken into account. Thus, in some cases, the sequence of the targeter can be a non-naturally occurring sequence. In other cases, the sequence of the targeter can be a naturally occurring sequence derived from the gene to be edited. In the case of a dual guide RNA, the targeter and the activator each have a duplex-forming segment, and the duplex-forming segment of the targeter and the duplex-forming segment of the activator are complementary to each other and hybridize to each other to form a double-stranded duplex (dsRNA duplex in the case of a gRNA). In some embodiments, the targeter includes both the guide sequence of the guide RNA and a stretch of nucleotides that forms half of the dsRNA duplex of the protein-binding segment of the gRNA. The corresponding tracrRNA-like molecule (activator) also includes a duplex-forming stretch of nucleotides that forms the other half of the dsRNA duplex of the protein-binding segment of the CasX guide RNA. Thus, the targeter and the activator hybridize as corresponding pairs to form the CasX dual guide NA, referred to herein as "dual guide NA", "dual molecule gNA", "dgNA", "dual molecule guide NA", or "two-molecule guide NA".
[0083] In some embodiments, the activator and the targeter of the reference gNA are covalently linked to each other and include a single molecule herein referred to as "single molecule gNA", "one molecule guide NA", "single guide NA", "single guide RNA", "single molecule guide RNA", "one molecule guide RNA", "single guide DNA", "single molecule DNA", or "one molecule guide DNA" ("sgRNA", "sgRNA", or "sgDNA"). In some embodiments, the sgNA includes an "activator" or a "targeter" and can thus be "activator RNA" and "targeter RNA", respectively.
[0084] In summary, the gNA of the present disclosure includes four distinct regions or domains, an RNA triplex, a scaffold stem, an extension stem, and a target sequence that are specific for a target nucleic acid in some embodiments of the present disclosure. The RNA triplex, the scaffold stem, and the extension stem together are referred to as the "scaffold" of the reference gNA. In some embodiments, the target sequence is at the 3' end of the gNA.
[0085] b. RNA triplex In some embodiments of the guide NA (including the reference sgNA) provided herein, an RNA triplex is present, and the RNA triplex includes the sequence of the UUU--nX (about 4 - 15)--UUU stem loop (SEQ ID NO: 19) that ends with AAAG after two intervening stem loops (scaffold stem loop and extension stem loop), and forms a pseudoknot that can also extend into a duplex pseudoknot beyond the triplex. The UU-UUU-AAA sequence of the triplex is formed as a nexus between the target sequence, the scaffold stem, and the extension stem. In an exemplary reference CasX sgNA, the UUU-loop-UUU region is first encoded, then the scaffold stem loop, then the extension stem loop linked by a tetraloop, and then the triplex is terminated with AAAG before becoming a spacer.
[0086] c. Scaffold stem loop In some embodiments of the sgNAs of the present disclosure, a scaffold stem-loop follows the triplex region. The scaffold stem-loop is the region of the gNA to which the CasX protein (such as a reference or CasX variant protein) binds. In some embodiments, the scaffold stem-loop is a fairly short and stable stem-loop. In some cases, the scaffold stem-loop tolerates few changes and requires some form of RNA bubble. In some embodiments, the scaffold stem is required for CasX sgNA function. Although perhaps similar to the nexus stem of Cas9 in that it is an important stem-loop, in some embodiments, the scaffold stem of CasX sgNA has the required bulge (RNA bubble) that is different from many other stem-loops found in the CRISPR / Cas system. In some embodiments, the presence of this bulge is conserved across sgNAs that interact with different CasX proteins. Exemplary sequences of the scaffold stem-loop sequence of the gNA include the sequence CCAGCGACUAUGUCGUAUGG (SEQ ID NO: 20). In other embodiments, the present disclosure provides gNA variants in which the scaffold stem-loop is replaced with an RNA stem-loop sequence from a heterologous RNA source having proximal 5' and 3' ends, such as, but not limited to, a stem-loop sequence selected from the MS2, Qβ, U1 hairpin II, Uvsx, or PP7 stem-loop. In some cases, the heterologous RNA stem-loop of the gNA can bind to a protein, an RNA structure, a DNA sequence, or a small molecule.
[0087] d. Extension stem-loop In some embodiments of the CasX sgNA of the present disclosure, an extension stem loop follows the scaffold stem loop. In some embodiments, the extension stem comprises a synthetic tracrRNA and crRNA fusion to which most of the CasX protein is not bound. In some embodiments, the extension stem loop can be highly adaptable. In some embodiments, the single-guide gRNA is made with a GAAA tetraloop linker or a GAGAAA linker between the tracrRNA and the crRNA within the extension stem loop. In some cases, the targeter and activator of the CasX sgNA are linked to each other by intervening nucleotides, and the linker can have a length of 3 to 20 nucleotides. In some embodiments of the CasX sgNA of the present disclosure, the extension stem is a large 32bp loop located outside the CasX protein of the ribonucleoprotein complex. Exemplary sequences of the extension stem loop sequence of the sgNA include the sequence GCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGC (SEQ ID NO: 21). In some embodiments, the extension stem loop comprises a GAGAAA spacer sequence. In some embodiments, the present disclosure provides gNA variants in which the extension stem loop is replaced with an RNA stem loop sequence from a heterologous RNA source having proximal 5' and 3' ends, such as, but not limited to, a stem loop sequence selected from an MS2, Qβ, U1 hairpin II, Uvsx, or PP7 stem loop. In such cases, the heterologous RNA stem loop increases the stability of the gNA. In other embodiments, the present disclosure provides gNA variants having an extension stem loop region comprising at least 10, at least 100, at least 500, at least 1000, or at least 10,000 nucleotides.
[0088] e. Target sequence In some embodiments of the gNA of the present disclosure, a region that forms part of a triplex follows the extended stem loop, followed by a target sequence (or "spacer"). The target sequence targets the CasX ribonucleoprotein holocomplex to a specific region of the target nucleic acid sequence of the gene to be modified. Thus, for example, the gNA target sequence of the present disclosure is complementary to a portion of the B2M gene in a eukaryotic nucleic acid (e.g., eukaryotic chromosome, chromosomal sequence, eukaryotic RNA, etc.) as a component of the RNP and, therefore, has a sequence that can hybridize thereto when any one of the PAM sequences TTC, ATC, GTC, or CTC is located at one nucleotide on the 5' side of the non-target strand sequence complementary to the target sequence. The target sequence of the gNA can be modified so that the gNA can target the desired sequence of any desired target nucleic acid sequence, provided that the position of the PAM sequence is taken into account. In some embodiments, the gNA scaffold is on the 5' side of the target sequence and the target sequence is at the 3' end of the gNA. In some embodiments, the PAM sequence recognized by the RNP is TC. In other embodiments, the PAM sequence recognized by the RNP is NTC.
[0089] In some embodiments, the target sequence of the gNA is specific to and can hybridize to a portion of a gene encoding a protein involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response, including but not limited to beta-2-microglobulin (B2M), T cell receptor alpha constant region (TRAC), class II major histocompatibility complex transactivator (CIITA), T cell receptor beta constant 1 (TRBC1), T cell receptor beta constant 2 (TRBC2), human leukocyte antigen A (HLA-A), human leukocyte antigen B (HLA-B), TGFβ receptor 2 (TGFβRII), programmed cell death 1 (PD-1), cytokine-induced SH2 (CISH), lymphocyte activation 3 (LAG-3), T cell immunoreceptor with Ig and ITIM domains (TIGIT), adenosine A2a receptor (ADORA2A), killer cell lectin-like receptor C1 (NKG2A), cytotoxic T lymphocyte-associated protein 4 (CTLA-4), T cell immunoglobulin and mucin domain 3 (TIM-3), and 2B4 (CD244). In a particular embodiment, the gene is B2M. The B2M gene encodes a serum protein that is found associated with the major histocompatibility complex (MHC) class I heavy chain on the surface of almost all nucleated cells. In another particular embodiment, the gene is TRAC. The TRAC gene encodes the C-terminal constant region linked to one of the 70 variable regions of the T cell alpha receptor. After similar synthesis of the beta chain, the alpha and beta chains pair to generate an alpha-beta T cell receptor heterodimer. In another particular embodiment, the gene is CITTA. The CIITA gene provides instructions for making a protein that mainly serves to control the activity (transcription) of the genes of the major histocompatibility complex (MHC) class II. As described above, the genomic target is intended such that the encoded gene of the target is knocked out or knocked down, thereby preventing the protein (e.g., cell marker or intracellular protein) from being expressed intracellularly or being expressed at a lower level. In some embodiments, the target sequence of the gNA is specific to an exon of the gene.In other embodiments, the target sequence of the gNA is specific to an intron of a gene. In other embodiments, the target sequence of the gNA is specific to a regulatory element of a gene. In other embodiments, the target sequence of the gNA is specific to the junction of an exon, intron, and / or regulatory element of a gene. In other embodiments, the target sequence of the gNA is specific to an intergenic region. In cases where the target sequence is specific to a regulatory element, such regulatory elements include, but are not limited to, regions containing promoter regions, enhancer regions, intergenic regions, 5' untranslated regions (5' UTRs), 3' untranslated regions (3' UTRs), conserved elements, and cis-regulatory elements. The promoter region is intended to encompass nucleotides within 5 kb from the start point of the coding sequence, or in the case of gene enhancer elements or conserved elements, can be located thousands, hundreds of thousands, or even millions of base pairs away from the coding sequence of the gene of the target nucleic acid. As described above, the target is intended such that the target coding gene is knocked out or knocked down, such that the target protein is not expressed in the cell or is expressed at a lower level in the cell.
[0090] In some embodiments, the target sequence of the gNA has 14 to 35 consecutive nucleotides. In some embodiments, the target sequence has 14, 15, 16, 18, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 consecutive nucleotides. In some embodiments, the target sequence consists of 20 consecutive nucleotides. In some embodiments, the target sequence consists of 19 consecutive nucleotides. In some embodiments, the target sequence consists of 18 consecutive nucleotides. In some embodiments, the target sequence consists of 17 consecutive nucleotides. In some embodiments, the target sequence consists of 16 consecutive nucleotides. In some embodiments, the target sequence consists of 15 consecutive nucleotides. In some embodiments, the target sequence has 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 consecutive nucleotides and contains 0 to 5, 0 to 4, 0 to 3, or 0 to 2 mismatches with respect to the target nucleic acid sequence, and the RNP containing the gNA containing the target sequence can retain sufficient binding specificity to form a complementary bond to the target nucleic acid.
[0091] Representative but non-limiting examples of target sequences for inclusion in the gNA of the present disclosure are presented in Tables 3A, 3B, and 3C (included as Figures 35-37), representing the target sequences of B2M, TRAC, and CIITA, respectively.
[0092] Exemplary target sequences (spacer sequences) of gNA embodiments utilized with the CasX:gNA system for editing of the B2M gene are provided in Table 3A (SEQ ID NOs: 725-2100 and 2281-7085). In one embodiment, the target sequence of the B2M gNA comprises a sequence having at least about 65%, at least about 75%, at least about 85%, or at least about 95% identity to a sequence selected from the group consisting of the sequences set forth in Table 3A. In another embodiment, the target sequence of the gNA consists of a sequence selected from the group consisting of the sequences set forth in Table 3A. In the foregoing embodiments, for any of the target sequences, thymine (T) nucleotides are used in place of one or more or all of the uracil (U) nucleotides, whereby the gNA can be a gDNA, gRNA, or a chimera of RNA and DNA. In some embodiments, the target sequences of Table 3A have at least 1, 2, 3, 4, 5, or 6 or more thymine nucleotides in place of thymine nucleotides. In other embodiments, the gNA, gRNA, or gDNA of the present disclosure comprises one, two, three, or more target sequences of Table 3A, or a target sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identity to one or more sequences of Table 3A.
[0093] Exemplary target sequences (spacer sequences) of gNA embodiments utilized with the CasX:gNA system for editing the TRAC gene are provided in Table 3B. In one embodiment, the target sequence of the TRAC gNA comprises a sequence having at least about 65%, at least about 75%, at least about 85%, or at least about 95% identity to a sequence selected from the group consisting of the sequences set forth in Table 3B. In another embodiment, the target sequence of the gNA consists of a sequence selected from the group consisting of the sequences set forth in Table 3B. In the foregoing embodiments, for any of the target sequences, thymine (T) nucleotides are used in place of one or more or all of the uracil (U) nucleotides, whereby the gNA can be a gDNA, gRNA, or a chimeric of RNA and DNA. In some embodiments, the target sequences of Table 3B have at least 1, 2, 3, 4, 5, or 6 or more thymine nucleotides in place of uracil nucleotides. In other embodiments, the gNA, gRNA, or gDNA of the present disclosure comprises one, two, three, or more target sequences of Table 3B, or a target sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identity to one or more sequences of Table 3B.
[0094] Exemplary target sequences (spacer sequences) of gNA embodiments for use with the CasX:gNA system for editing the CIITA gene are provided in Table 3C. In one embodiment, the target sequence of the TRAC gNA comprises a sequence having at least about 65%, at least about 75%, at least about 85%, or at least about 95% identity to a sequence selected from the group consisting of the sequences set forth in Table 3C. In another embodiment, the target sequence of the gNA consists of a sequence selected from the group consisting of the sequences set forth in Table 3C. In the foregoing embodiments, for any of the target sequences, thymine (T) nucleotides are used in place of one or more or all of the uracil (U) nucleotides, whereby the gNA can be a gDNA or gRNA, or a chimera of RNA and DNA. In some embodiments, the target sequences of Table 3C have at least 1, 2, 3, 4, 5, or more than 6 thymine nucleotides in place of uracil nucleotides. In other embodiments, the gNA, gRNA, or gDNA of the present disclosure comprises 1, 2, 3, or more of the target sequences of Table 3C, or a target sequence that is at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical to one or more of the sequences of Table 3C.
[0095] In some embodiments, the CasX:gNA system comprises a first gNA and further comprises a second (and optionally a third, fourth, fifth, or more) gNA, wherein the second gNA or further gNA has a target sequence that is complementary to a different or overlapping portion of the target nucleic acid sequence as compared to the target sequence of the first gNA, whereby multiple points within the target nucleic acid are targeted, e.g., multiple cleavages are introduced into the target nucleic acid by CasX. In such cases, it will be understood that the second or further gNA complexes with additional copies of the CasX protein. By selecting the target sequences of the gNA, defined regions of the target nucleic acid sequence flanking a particular position within the target nucleic acid can be modified or edited using the CasX:gNA system described herein (including facilitating the insertion of a donor template).
[0096] f.gNA scaffold In some embodiments, the CasX reference gRNA comprises a sequence isolated from or derived from Deltaproteobacter. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary CasX reference tracrRNA sequences isolated from or derived from Deltaproteobacter can include ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 22) and ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGG (SEQ ID NO: 23). Exemplary crRNA sequences isolated from or derived from Deltaproteobacter can include the sequence of CCGAUAAGUAAAACGCAUCAAAG (SEQ ID NO: 24). In some embodiments, the reference gNA comprises a sequence that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to a sequence isolated from or derived from Deltaproteobacter.
[0097] In some embodiments, the CasX reference guide RNA comprises a sequence isolated from or derived from Planctomycetes. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary CasX reference tracrRNA sequences isolated from or derived from Planctomycetes are UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 25) and
[0098] may include UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGG (SEQ ID NO: 26). Exemplary crRNA sequences isolated from or derived from Planctomycetes may include the sequence of UCUCCGAUAAAUAAGAAGCAUCAAAG (SEQ ID NO: 27). In some embodiments, the CasX reference gNA comprises a sequence that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to a sequence isolated from or derived from Planctomycetes.
[0099] In some embodiments, the CasX reference gNA comprises a sequence isolated from or derived from Candidatus Sungbacteria. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary CasX reference tracrRNA sequences isolated from or derived from Candidatus Sungbacteria may include the sequences of GUUUACACACUCCCUCUCAUAGGGU (SEQ ID NO: 28), GUUUACACACUCCCUCUCAUGAGGU (SEQ ID NO: 29), UUUUACAUACCCCCUCUCAUGGGAU (SEQ ID NO: 30), and GUUUACACACUCCCUCUCAUGGGGG (SEQ ID NO: 31). In some embodiments, the CasX reference guide RNA comprises a sequence that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to a sequence isolated from or derived from Candidatus Sungbacteria.
[0100] Table 1 provides reference gRNA tracr sequences and scaffold sequences. In some embodiments, the present disclosure provides a scaffold having a sequence that has at least one nucleotide modification compared to a reference gNA sequence in which the gNA has any one of the sequences of SEQ ID NOs: 4-16 in Table 1, the scaffold comprising a gNA sequence. In embodiments where the vector comprises the DNA coding sequence of the gNA, or where the gNA is a chimera of gDNA or RNA and DNA, it will be understood that the thymine (T) base can be used in place of any uracil (U) base in the embodiments of the gNA sequences described herein.
Table 1
[0101] g.gNA variant In another aspect, the disclosure relates to guide nucleic acid variants (alternatively referred to herein as "gNA variants" or "gRNA variants") that include one or more modifications as compared to a reference gRNA scaffold. As used herein, "scaffold" refers to all parts of the gNA necessary for gNA function excluding the spacer sequence.
[0102] In some embodiments, the gNA variant includes one or more nucleotide substitutions, insertions, deletions, or regions of substituted or replaced nucleotides as compared to the reference gRNA sequence of the disclosure. In some embodiments, mutations can occur in any region of the reference gRNA scaffold to generate the gNA variant. In some embodiments, the scaffold of the gNA variant sequence has at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, or at least 70%, at least 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the sequence of SEQ ID NO: 4 or SEQ ID NO: 5.
[0103] In some embodiments, the gNA variant comprises one or more nucleotide changes within one or more regions of the reference gRNA scaffold that improve the properties of the reference gRNA. Exemplary regions include RNA triplexes, pseudoknots, scaffold stem-loops, and extended stem-loops. In some cases, the variant scaffold stem further comprises a bubble. In other cases, the variant scaffold further comprises a triplex loop region. In still other cases, the variant scaffold further comprises a 5' unstructured region. In one embodiment, the gNA variant scaffold comprises a scaffold stem-loop having at least 60% sequence identity with SEQ ID NO: 14. In another embodiment, the gNA variant comprises a scaffold stem-loop having the sequence CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 32). In another embodiment, the present disclosure provides a gNA scaffold that, compared to SEQ ID NO: 5, has a C18G substitution, a G55 insertion, a U1 deletion, and a modified extended stem-loop, wherein the original 6nt loop and the 13 bases proximal to the loop (32 nucleotides total) are replaced with a Uvsx hairpin (4nt loop and 5 bases proximal to the loop, 14 nucleotides total), and the bases distal to the loop of the extended stem are converted to a fully base-paired stem adjacent to the new Uvsx hairpin by a deletion of A99 and a substitution of G64U. In the foregoing embodiment, the gNA scaffold comprises the sequence ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG (SEQ ID NO: 33).
[0104] All gNA variants that have one or more improved functions or properties or add one or more new functions when the variant gNA is compared to the reference gRNA described herein are contemplated to be within the scope of this disclosure. A representative example of such a gNA variant is Guide 174 (SEQ ID NO: 2238), the design of which is described in the Examples. In some embodiments, the gNA variant adds a new function to the RNP containing the gNA variant. In some embodiments, the gNA variant has improved stability, improved solubility, improved transcription of the gNA, improved resistance to nuclease activity, increased folding rate of the gNA, decreased byproduct formation during folding, increased productive folding, improved binding affinity for the CasX protein, improved binding affinity for the target DNA when complexed with the CasX protein, improved gene editing when complexed with the CasX protein, improved editing specificity when complexed with the CasX protein, and improved ability to utilize a broader spectrum of one or more PAM sequences including ATC, CTC, GTC, or TTC in the editing of the target DNA when complexed with the CasX protein, and has improved properties selected from any combination thereof. In some cases, one or more of the improved properties of the gNA variant are improved by at least about 1.1 to about 100,000-fold compared to the reference gNA of SEQ ID NO: 4 or SEQ ID NO: 5. In other cases, one or more of the improved properties of the gNA variant are improved by at least about 1.1, at least about 10, at least about 100, at least about 1000, at least about 10,000, at least about 100,000-fold or more compared to the reference gNA of SEQ ID NO: 4 or SEQ ID NO: 5.In other cases, one or more of the improved properties of the gNA variant are improved by about 1.1 to 100,000-fold, about 1.1 to 10,000-fold, about 1.1 to 1,000-fold, about 1.1 to 500-fold, 1.1 to 100-fold, about 1.1 to 50-fold, about 1.1 to 20-fold, about 10 to 100,000-fold, about 10 to 10,000-fold, about 10 to 1,000-fold, about 10 to 500-fold, about 10 to 100-fold, about 10 to 50-fold, about 10 to 20-fold, about 2 to 70-fold, about 2 to 50-fold, about 2 to 30-fold, about 2 to 20-fold, about 2 to 10-fold, about 5 to 50-fold, about 5 to 30-fold, about 5 to 10-fold, about 100 to 100,000-fold, about 100 to 10,000-fold, about 100 to 1,000-fold, about 100 to 500-fold, about 500 to 100,000-fold, about 500 to 10,000-fold, about 500 to 1,000-fold, about 500 to 750-fold, about 1,000 to 100,000-fold, about 10,000 to 100,000-fold, about 20 to 500-fold, about 20 to 250-fold, about 20 to 200-fold, about 20 to 100-fold, about 20 to 50-fold, about 50 to 10,000-fold, about 50 to 1,000-fold, about 50 to 500-fold, about 50 to 200-fold, or about 50 to 100-fold compared to the reference gNA of SEQ ID NO: 4 or SEQ ID NO: 5. In other cases, one or more of the improved properties of the gNA variant are improved by about 1.1-fold, 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, 15-fold, 16-fold, 17-fold, 18-fold, 19-fold, 20-fold, 25-fold, 30-fold, 40-fold, 45-fold, 50-fold, 55-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, 110-fold, 120-fold, 130-fold, 140-fold, 150-fold, 160-fold, 170-fold, 180-fold, 190-fold, 200-fold, 210-fold, 220-fold, 230-fold, 240-fold, 250-fold, 260-fold, 270-fold, 280-fold, 290-fold, 300-fold, 310-fold, 320-fold, 330-fold, 340-fold, 350-fold, 360-fold, 370-fold, 380-fold, 390-fold, 400-fold, 425-fold, 450-fold, 475-fold, or 500-fold compared to the reference gNA of SEQ ID NO: 4 or SEQ ID NO: 5.
[0105] In some embodiments, the gNA variant can be made by subjecting a reference gRNA to one or more mutagenesis methods described herein, including deep mutational evolution (DME), deep mutational scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping, to generate the gNA variants of the present disclosure. The activity of the reference gRNA is used as a benchmark against which the activity of the gNA variant is compared, thereby allowing measurement of the improvement in the function of the gNA variant. In other embodiments, the reference gRNA can be subjected to one or more intentional targeted mutations, substitutions, or domain exchanges to generate gNA variants, such as rationally designed variants. Exemplary gRNA variants generated by such methods are described in the Examples, and representative sequences of the gNA scaffolds are presented in Table 2.
[0106] In some embodiments, the gNA variant comprises one or more modifications as compared to a reference guide nucleic acid scaffold sequence, and the one or more modifications are selected from at least one nucleotide substitution in a region of the gNA variant, at least one nucleotide deletion in a region of the gNA variant, at least one nucleotide insertion in a region of the gNA variant, substitution of all or a portion of a region of the gNA variant, deletion of all or a portion of a region of the gNA variant, or any combination of the foregoing. In some cases, the modification is a substitution of 1 to 15 consecutive or non-consecutive nucleotides of the gNA variant in one or more regions. In other cases, the modification is a deletion of 1 to 10 consecutive or non-consecutive nucleotides of the gNA variant in one or more regions. In other cases, the modification is an insertion of 1 to 10 consecutive or non-consecutive nucleotides of the gNA variant in one or more regions. In other cases, the modification is a substitution of the scaffold stem loop or extended stem loop with an RNA stem loop sequence from a heterologous RNA source having proximal 5' and 3' ends. In some cases, the gNA variants of the present disclosure comprise two or more modifications in one region. In other cases, the gNA variants of the present disclosure comprise modifications in two or more regions. In other cases, the gNA variant comprises any combination of the foregoing modifications described in this paragraph.
[0107] In some embodiments, when the +1 nucleotide is G, transcription from the U6 promoter is more efficient and consistent with respect to the start site, and thus a 5' G is added to the gNA variant sequence for in vivo expression. In other embodiments, since T7 polymerase strongly prefers G at the +1 position and a purine at the +2 position, two 5' Gs are added to the gNA variant sequence for in vitro transcription to increase production efficiency. In some cases, a 5' G base is added to the reference scaffold of Table 1. In other cases, a 5' G base is added to the variant scaffold of Table 2.
[0108] Table 2 provides exemplary gNA variant scaffold sequences. In Table 2, (-) indicates a deletion at the specified position relative to the reference sequence of SEQ ID NO: 5, (+) indicates an insertion of the specified base at the indicated position relative to SEQ ID NO: 5, (:) indicates the range of bases at the specified start:stop coordinates of a deletion or substitution relative to SEQ ID NO: 5, and multiple insertions, deletions, or substitutions are separated by commas, e.g., A14C, U17G. In some embodiments, the gNA variant scaffold comprises any one of the sequences of SEQ ID NOs: 2101 to 2280 listed in Table 2, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% sequence identity thereto. In embodiments where the vector comprises the DNA coding sequence of gNA, or where gNA is a chimera of gDNA or RNA and DNA, it will be understood that the thymine (T) base can be used in place of any uracil (U) base of the gNA sequence embodiments described herein.
Table 2-1
Table 2-2
Table 2-3
Table 2-4
Table 2-5
Table 2-6
Table 2-7
Table 2-8
Table 2-9
Table 2-10
Table 2-11
Table 2-12
Table 2-13
[0109] In some embodiments, the gNA variant comprises a tracrRNA stem-loop comprising the sequence -UUU-N4-25-UUU- (SEQ ID NO: 34). For example, the gNA variant comprises a scaffold stem-loop or a substitution thereof in which two triplet U motifs that contribute to the triplex region are adjacent. In some embodiments, the scaffold stem-loop or a substitution thereof comprises at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, or at least 25 nucleotides.
[0110] In some embodiments, the gNA variant comprises a crRNA sequence having -AAAG- at a position 5' to the spacer region. In some embodiments, the -AAAG- sequence is immediately 5' to the spacer region.
[0111] In some embodiments, at least one nucleotide modification to a reference gNA to produce a gNA variant comprises at least one nucleotide deletion in the CasX variant gNA compared to the reference gRNA. In some embodiments, the gNA variant comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive or non-consecutive nucleotides compared to the reference gNA. In some embodiments, at least one deletion comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more consecutive nucleotides compared to the reference gNA. In some embodiments, the gNA variant comprises a deletion of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides compared to the reference gNA, and this deletion is not in consecutive nucleotides. In embodiments where there are two or more non-consecutive deletions in the gNA variant compared to the reference gRNA, any length of deletion, and any combination of lengths of deletions, as described herein, are intended to be within the scope of the present disclosure. For example, in some embodiments, the gNA variant may comprise a first deletion of one nucleotide and a second deletion of two nucleotides, and these two deletions are not consecutive. In some embodiments, the gNA variant comprises at least two deletions in different regions of the reference gRNA. In some embodiments, the gNA variant comprises at least two deletions in the same region of the reference gRNA. For example, the region can be an extended stem-loop, a scaffold stem-loop, a scaffold stem-bubble, a triplex loop, a pseudoknot, a triplex, or the 5' end of the gNA variant. Any nucleotide deletion in the reference gRNA is intended to be within the scope of the present disclosure.
[0112] In some embodiments, at least one nucleotide modification of the reference gRNA to generate the gNA variant comprises at least one nucleotide insertion. In some embodiments, the gNA variant comprises an insertion of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 consecutive or non-consecutive nucleotides as compared to the reference gRNA. In some embodiments, the at least one nucleotide insertion comprises an insertion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more consecutive nucleotides as compared to the reference gRNA. In some embodiments, the gNA variant comprises two or more insertions as compared to the reference gRNA, and these insertions are not consecutive. In embodiments where there are two or more non-consecutive insertions in the gNA variant as compared to the reference gRNA, any length of insertion, and any combination of lengths of insertions, as described herein, are contemplated to be within the scope of the present disclosure. For example, in some embodiments, the gNA variant may comprise a first insertion of one nucleotide and a second insertion of two nucleotides, and these two insertions are not consecutive. In some embodiments, the gNA variant comprises at least two insertions in different regions of the reference gRNA. In some embodiments, the gNA variant comprises at least two insertions in the same region of the reference gRNA. For example, the region can be an extension stem-loop, a scaffold stem-loop, a scaffold stem-bubble, a triplex loop, a pseudoknot, a triplex, or the 5' end of the gNA variant. Any insertion of A, G, C, U (or T in the corresponding DNA), or combinations thereof, at any position in the reference gRNA is contemplated to be within the scope of the present disclosure.
[0113] In some embodiments, at least one nucleotide modification of the reference gRNA for generating the gNA variant comprises at least one nucleic acid substitution. In some embodiments, the gNA variant comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more consecutive or non-consecutive substituted nucleotides compared to the reference gRNA. In some embodiments, the gNA variant comprises 1 to 4 nucleotide substitutions compared to the reference gRNA. In some embodiments, at least one substitution comprises a substitution of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more consecutive nucleotides compared to the reference gRNA. In some embodiments, the gNA variant comprises two or more substitutions compared to the reference gRNA, and these substitutions are not consecutive. In embodiments where there are two or more non-consecutive substitutions in the gNA variant compared to the reference gRNA, any length of substituted nucleotides, and any combination of lengths of substituted nucleotides described herein are intended to be within the scope of the present disclosure. For example, in some embodiments, the gNA variant may comprise a first substitution of one nucleotide and a second substitution of two nucleotides, and these two substitutions are not consecutive. In some embodiments, the gNA variant comprises at least two substitutions in different regions of the reference gRNA. In some embodiments, the gNA variant comprises at least two substitutions in the same region of the reference gRNA. For example, the region can be a triplex, an extended stem-loop, a scaffold stem-loop, a scaffold stem-bubble, a triplex loop, a pseudoknot, a triplex, or the 5' end of the gNA variant. Any substitution of A, G, C, U (or T in the corresponding DNA), or combinations thereof at any position in the reference gRNA is intended to be within the scope of the present disclosure.
[0114] Combinations of any of the substitutions, insertions, and deletions described herein can be used to generate the gNA variants of the present disclosure. For example, a gNA variant can include at least one substitution and at least one deletion as compared to a reference gRNA, at least one substitution and at least one insertion as compared to a reference gRNA, at least one insertion and at least one deletion as compared to a reference gRNA, or at least one substitution, one insertion, and one deletion as compared to a reference gRNA.
[0115] In some embodiments, the gNA variant comprises a scaffold region that is at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs: 4-16. In some embodiments, the gNA variant comprises a scaffold region that is at least 60% homologous (or identical) to any one of SEQ ID NOs: 4-16.
[0116] In some embodiments, the gNA variant comprises a tracr stem-loop that is at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 14. In some embodiments, the gNA variant comprises a tracr stem-loop that is at least 60% homologous (or identical) to SEQ ID NO: 14.
[0117] In some embodiments, the gNA variant comprises an extended stem loop that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to SEQ ID NO: 15. In some embodiments, the gNA variant comprises an extended stem loop that is at least 60% homologous (or identical) to SEQ ID NO: 15.
[0118] In some embodiments, the gNA variant includes an exogenous extension stem-loop, as described below, and differs from the reference gNA in this regard. In some embodiments, the exogenous extension stem-loop has little or no identity with the reference stem-loop region disclosed herein (e.g., SEQ ID NO: 15). In some embodiments, the exogenous stem-loop is at least 10 bp, at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, at least 1,000 bp, at least 2,000 bp, at least 3,000 bp, at least 4,000 bp, at least 5,000 bp, at least 6,000 bp, at least 7,000 bp, at least 8,000 bp, at least 9,000 bp, at least 10,000 bp, at least 12,000 bp, at least 15,000 bp, or at least 20,000 bp. In some embodiments, the gNA variant includes an extension stem-loop region comprising at least 10, at least 100, at least 500, at least 1,000, or at least 10,000 nucleotides. In some embodiments, the heterologous stem-loop increases the stability of the gNA. In some embodiments, the heterologous RNA stem-loop can bind to a protein, an RNA structure, a DNA sequence, or a small molecule.In some embodiments, the exogenous stem-loop region comprises an RNA stem-loop or hairpin, such as a thermostable RNA, such as MS2 (ACAUGAGGAUUACCCAUGU (SEQ ID NO: 35)), Qβ (UGCAUGUCUAAGACAGCA (SEQ ID NO: 36)), U1 hairpin II (AAUCCAUUGCACUCCGGAUU (SEQ ID NO: 37)), Uvsx (CCUCUUCGGAGG (SEQ ID NO: 38)), PP7 (AGGAGUUUCUAUGGAAACCCU (SEQ ID NO: 39)), phage replication loop (AGGUGGGACGACCUCUCGGUCGUCCUAUCU (SEQ ID NO: 40)), kissing loop_a (UGCUCGCUCCGUUCGAGCA (SEQ ID NO: 41)), kissing loop_b1 (UGCUCGACGCGUCCUCGAGCA (SEQ ID NO: 42)), kissing loop_b2 (UGCUCGUUUGCGGCUACGAGCA (SEQ ID NO: 43)), G-quadruplex M3q (AGGGAGGGAGGGAGAGG (SEQ ID NO: 44)), G-quadruplex telomere basket (GGUUAGGGUUAGGGUUAGG (SEQ ID NO: 45)), sarcin-ricin loop (CUGCUCAGUACGAGAGGAACCGCAG (SEQ ID NO: 46)), or pseudoknot (UACACUGGGAUCGCUGAAUUAGAGAUCGGCGUCCUUUCAUUCUAUAUACUUUGGAGUUUUAAAAUGUCUCUAAGUACA (SEQ ID NO: 47)). In some embodiments, the exogenous stem-loop comprises an RNA scaffold. As used herein, "RNA scaffold" refers to a multi-dimensional RNA structure that can interact with and organize or localize one or more proteins. In some embodiments, the RNA scaffold is synthetic or non-naturally occurring. In some embodiments, the exogenous stem-loop comprises a long non-coding RNA (lncRNA). As used herein, lncRNA refers to a non-coding RNA that is approximately longer than 200 bp. In some embodiments, the 5' end and 3' end of the exogenous stem-loop form base pairs, i.e., they interact to form a region of duplex RNA.In some embodiments, the 5' end and the 3' end of the exogenous stem loop form base pairs, and one or more regions between the 5' end and the 3' end of the exogenous stem loop do not form base pairs. In some embodiments, at least one nucleotide modification comprises (a) substitution of 1 to 15 consecutive or non-consecutive nucleotides of the gNA variant in one or more regions, (b) deletion of 1 to 10 consecutive or non-consecutive nucleotides of the gNA variant in one or more regions, (c) insertion of 1 to 10 consecutive or non-consecutive nucleotides of the gNA variant in one or more regions, (d) substitution of a scaffold stem loop or an extended stem loop with an RNA stem loop sequence from a heterologous RNA source having proximal 5' and 3' ends, or any combination of (a)-(d).
[0119] In some embodiments, the gNA variant comprises a scaffold stem loop having at least 60% identity with SEQ ID NO: 14. In some embodiments, the gNA variant comprises a scaffold stem loop having at least 60% identity, at least 70% identity, at least 80% identity, at least 90% identity, at least 95% identity, at least 98% identity, or at least 99% identity with SEQ ID NO: 14. In some embodiments, the gNA variant comprises a scaffold stem loop comprising SEQ ID NO: 14.
[0120] In some embodiments, the gNA variant comprises the scaffold stem loop sequence of CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 32). In some embodiments, the gNA variant comprises the scaffold stem loop sequence of CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 32) having at least 1, 2, 3, 4, or 5 mismatches therewith.
[0121] In some embodiments, the gNA variant comprises an extended stem-loop region that comprises fewer than 32 nucleotides, fewer than 31 nucleotides, fewer than 30 nucleotides, fewer than 29 nucleotides, fewer than 28 nucleotides, fewer than 27 nucleotides, fewer than 26 nucleotides, fewer than 25 nucleotides, fewer than 24 nucleotides, fewer than 23 nucleotides, fewer than 22 nucleotides, fewer than 21 nucleotides, or fewer than 20 nucleotides. In some embodiments, the gNA variant comprises an extended stem-loop region that comprises fewer than 32 nucleotides. In some embodiments, the gNA variant further comprises a thermostable stem-loop.
[0122] In some embodiments, the sgRNA variant comprises the sequence of SEQ ID NO: 2104, SEQ ID NO: 2106, SEQ ID NO: 2163, SEQ ID NO: 2107, SEQ ID NO: 2164, SEQ ID NO: 2165, SEQ ID NO: 2166, SEQ ID NO: 2103, SEQ ID NO: 2167, SEQ ID NO: 2105, SEQ ID NO: 2108, SEQ ID NO: 2112, SEQ ID NO: 2160, SEQ ID NO: 2170, SEQ ID NO: 2114, SEQ ID NO: 2171, SEQ ID NO: 2112, SEQ ID NO: 2173, SEQ ID NO: 2102, SEQ ID NO: 2174, SEQ ID NO: 2175, SEQ ID NO: 2109, SEQ ID NO: 2176, SEQ ID NO: 2238, SEQ ID NO: 2239, SEQ ID NO: 2240, SEQ ID NO: 2241, SEQ ID NO: 2274, or SEQ ID NO: 2275.
[0123] In some embodiments, the gNA variant comprises any one of the sequences of SEQ ID NO: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259 - 2280, or has at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity thereto. In some embodiments, the gNA variant comprises one or more additional changes relative to any one of the sequences of SEQ ID NO: 2201 - 2280. In some embodiments, the gNA variant comprises any one of the sequences of SEQ ID NO: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259 - 2280.
[0124] In some embodiments, the sgRNA variant comprises one or more additional changes relative to the sequences of SEQ ID NO: 2104, SEQ ID NO: 2163, SEQ ID NO: 2107, SEQ ID NO: 2164, SEQ ID NO: 2165, SEQ ID NO: 2166, SEQ ID NO: 2103, SEQ ID NO: 2167, SEQ ID NO: 2105, SEQ ID NO: 2108, SEQ ID NO: 2112, SEQ ID NO: 2160, SEQ ID NO: 2170, SEQ ID NO: 2114, SEQ ID NO: 2171, SEQ ID NO: 2112, SEQ ID NO: 2173, SEQ ID NO: 2102, SEQ ID NO: 2174, SEQ ID NO: 2175, SEQ ID NO: 2109, SEQ ID NO: 2176, SEQ ID NO: 2238, SEQ ID NO: 2239, SEQ ID NO: 2240, SEQ ID NO: 2241, SEQ ID NO: 2274, or SEQ ID NO: 2275.
[0125] In some embodiments of the gNA variants of the present disclosure, the gNA variant comprises at least one modification, and at least one modification compared to the reference guide scaffold of SEQ ID NO: 5 is selected from one or more of the following: (a) C18G substitution in the triplex loop; (b) G55 insertion in the stem bubble; (c) U1 deletion; (d) modification of the extended stem loop, where the modification of the extended stem loop is (i) replacement of a 6-nt loop and 13 loop-proximal base pairs by the Uvsx hairpin, and (ii) deletion of A99 and substitution of G65U resulting in a fully base-paired loop-distal base. In such embodiments, the gNA variant comprises any one of the sequences of SEQ ID NO: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280.
[0126] In some embodiments, the scaffold of the gNA variant comprises any one of the sequences of SEQ ID NO: 2201-2280 in Table 2. In some embodiments, the scaffold of the gNA consists of or consists essentially of any one of the sequences of SEQ ID NO: 2201-2280. In some embodiments, the scaffold of the gNA variant sequence is at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to any one of the sequences of SEQ ID NO: 2201-2280.
[0127] In embodiments of the gNA variant, the gNA further comprises a spacer (or target sequence) region more fully described above that comprises at least 14 to about 35 nucleotides, and the spacer is designed with a sequence complementary to the target DNA. In some embodiments, the gNA variant comprises a target sequence of at least 10 to 30 nucleotides complementary to the target DNA. In some embodiments, the target sequence has 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides. In some embodiments, the gNA variant comprises a target sequence having 20 nucleotides. In some embodiments, the target sequence has 25 nucleotides. In some embodiments, the target sequence has 24 nucleotides. In some embodiments, the target sequence has 23 nucleotides. In some embodiments, the target sequence has 22 nucleotides. In some embodiments, the target sequence has 21 nucleotides. In some embodiments, the target sequence has 20 nucleotides. In some embodiments, the target sequence has 19 nucleotides. In some embodiments, the target sequence has 18 nucleotides. In some embodiments, the target sequence has 17 nucleotides. In some embodiments, the target sequence has 16 nucleotides. In some embodiments, the target sequence has 15 nucleotides. In some embodiments, the target sequence has 14 nucleotides. In some embodiments, the present disclosure provides target sequences for inclusion in gNA variants of the present disclosure that are at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% identical to the sequences of Table 3A, Table 3B, or Table 3C. In some embodiments, the target sequence of the gNA variant comprises the sequence of Table 3A, Table 3B, or Table 3C with a single nucleotide removed from the 3' end of the sequence.In other embodiments, the target sequence of the gNA variant comprises the sequence of Table 3A, Table 3B, or Table 3C with two nucleotides removed from the 3' end of the sequence. In other embodiments, the target sequence of the gNA variant comprises the sequence of Table 3A, Table 3B, or Table 3C with three nucleotides removed from the 3' end of the sequence. In other embodiments, the target sequence of the gNA variant comprises the sequence of Table 3A, Table 3B, or Table 3C with four nucleotides removed from the 3' end of the sequence. In other embodiments, the target sequence of the gNA variant comprises the sequence of Table 3 with five nucleotides removed from the 3' end of the sequence. Table 3A. gNA target sequence of B2M Table 3A is provided in Figure 35 and is referred to throughout as Table 3A. Table 3B. gNA target sequence of TRAC Table 3B is provided in Figure 36 and is referred to throughout as Table 3B. Table 3C: gNA target sequence of CIITA Table 3C is provided in Figure 37 and is referred to throughout as Table 3C.
[0128] In Tables 3A, 3B, and 3C, the left column shows the PAM sequence and the right column shows the sequence numbers of the corresponding spacer sequences (sometimes referred to herein as target sequences).
[0129] In some embodiments, the scaffold of the gNA variant is part of an RNP having a reference CasX protein comprising SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In other embodiments, the scaffold of the gNA variant is part of an RNP having a CasX variant protein comprising any one of the sequences of Tables 4, 7, 8, 9, or 11, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In the foregoing embodiments, the gNA further comprises a spacer sequence.
[0130] In some embodiments, the scaffold of the gNA variant is a variant that includes one or more additional changes relative to the sequence of a reference gRNA that includes SEQ ID NO: 4 or SEQ ID NO: 5. In embodiments where the scaffold of the reference gRNA is derived from SEQ ID NO: 4 or SEQ ID NO: 5, one or more improved or additional properties of the gNA variant are improved as compared to the same properties of SEQ ID NO: 4 or SEQ ID NO: 5.
[0131] Complex formation with h.CasX protein In some embodiments, the gNA variant has an improved ability to form a complex with a CasX protein (such as a reference CasX or a CasX variant protein) as compared to the reference gRNA. In some embodiments, the gNA variant has an improved affinity for a CasX protein (such as a reference or variant protein) as compared to the reference gRNA, thereby improving its ability to form a ribonucleoprotein (RNP) complex with the CasX protein as described in the examples. By improving ribonucleoprotein complex formation, in some embodiments, the efficiency of assembling a functional RNP can be improved. In some embodiments, more than 90%, more than 93%, more than 95%, more than 96%, more than 97%, more than 98%, or more than 99% of the RNP comprising the gNA variant and its spacer has the ability to gene edit a target nucleic acid.
[0132] Exemplary nucleotide changes that can improve the ability of a gNA variant to form a complex with a CasX protein can, in some embodiments, include replacing a scaffold stem with a thermostable stem-loop. Without wishing to be bound by any theory, replacing the scaffold stem with a thermostable stem-loop can increase the overall binding stability of the gNA variant with the CasX protein. Alternatively, or in addition, by removing most of the stem-loop, the folding dynamics of the gNA variant are changed such that a functional folded gNA is readily and rapidly produced and can be structurally assembled, for example, by reducing the degree to which the gNA variant can "entangle" itself. In some embodiments, the choice of scaffold stem-loop sequence can vary depending on the different spacers utilized in the gNA. In some embodiments, the scaffold sequence can be tailored to the spacer and thus the target sequence. Biochemical assays can be used to assess the binding affinity of the CasX protein for gNA variants for forming RNPs, including the assays of the examples. For example, one skilled in the art can measure the change in the amount of fluorescently labeled gNA bound to immobilized CasX protein in response to increasing concentrations of an additional unlabeled "cold competitor" gNA. Alternatively, or in addition, one can monitor or confirm how the fluorescence signal changes when different amounts of fluorescently labeled gNA are flowed over immobilized CasX protein. Alternatively, the ability to form an RNP can be evaluated using an in vitro cleavage assay against a defined target nucleic acid sequence.
[0133] i. gNA stability In some embodiments, the gNA variant has improved stability compared to the reference gRNA. The increased stability and efficient folding can, in some embodiments, increase the extent to which the gNA variant persists within the target cell, thereby increasing the opportunity to form a functional RNP that can perform CasX functions such as gene editing. The increased stability of the gNA variant can, in some embodiments, enable similar results with a lesser amount of gNA delivered to the cell, and then reduce the opportunity for off-target effects during gene editing.
[0134] In another aspect, the disclosure provides a gNA in which the scaffold stem loop and / or the extension stem loop are replaced with a hairpin loop or a thermostable RNA stem loop, and the resulting gNA has increased stability and can interact with certain cellular proteins or RNAs depending on the choice of loop. In some embodiments, the substituted RNA loop is selected from MS2, Qβ, U1 hairpin II, Uvsx, PP7, phage replication loop, kissing loop_a, kissing loop_b1, kissing loop_b2, G quadruplex M3q, G quadruplex telomere basket, sarcin-ricin loop, and pseudoknot. The sequences of gNA variants containing such components are shown in Table 2B.
[0135] Guide RNA stability can be evaluated in various ways, including, for example, assembling the guide in vitro, incubating for various periods in a solution mimicking the intracellular environment, and then measuring the functional activity by the in vitro cleavage assays described herein. Alternatively, or in addition, the gNA can be collected from the cells at various time points after the initial transfection / transduction of the gNA to determine how long the gNA variant persists compared to the reference gRNA.
[0136] j. Solubility In some embodiments, the gNA variant has improved solubility compared to the reference gRNA. In some embodiments, the gNA variant has improved solubility of the CasX protein:gNA RNP compared to the reference gRNA. In some embodiments, the solubility of the CasX protein:gNA RNP is improved by adding a ribozyme sequence to the 5' or 3' end of the gNA variant, e.g., the 5' or 3' side of the reference sgRNA. Some ribozymes, such as the M1 ribozyme, can increase the solubility of a protein by RNA-mediated protein folding.
[0137] The increased solubility of the CasX RNP containing the gNA variant described herein can be evaluated by various means known to those of skill in the art, e.g., by performing concentration measurement readings on a gel of the soluble fraction of lysed E. coli in which the CasX and gNA variant are expressed.
[0138] k. Resistance to nuclease activity In some embodiments, the gNA variant has improved resistance to nuclease activity compared to the reference gRNA. Without wishing to be bound by any theory, an increase in resistance to nucleases, such as nucleases found in cells, can, for example, increase the persistence of the variant gNA in the intracellular environment, thereby improving gene editing.
[0139] Many nucleases are processive and degrade RNA in a 3’ to 5’ manner. Thus, in some embodiments, addition of nuclease-resistant secondary structures to one or both ends of the gNA, or nucleotide changes that alter the secondary structure of the sgNA, can generate gNA variants with increased resistance to nuclease activity. Resistance to nuclease activity can be evaluated by a variety of methods known to those of skill in the art. For example, in vitro methods of measuring resistance to nuclease activity can include contacting a reference gNA and variant with one or more exemplary RNA nucleases and measuring degradation. Alternatively, or in addition, by using the methods described herein to measure the persistence of a gNA variant in a cellular environment, the degree to which the gNA variant is nuclease-resistant can be indicated.
[0140] l. Binding affinity for target DNA In some embodiments, the gNA variant has improved binding affinity for target DNA compared to a reference gRNA. In certain embodiments, the ribonucleoprotein complex comprising the gNA variant has improved binding affinity for target DNA compared to the affinity of the RNP comprising the reference gRNA. In some embodiments, the improved affinity of the RNP for target DNA includes improved affinity for the target sequence, improved affinity for the PAM sequence, improved ability of the RNP to locate the DNA of the target sequence, or any combination thereof. In some embodiments, the improved binding affinity for target DNA is the result of an increase in overall DNA binding affinity.
[0141] Without wishing to be bound by theory, nucleotide changes in the gNA variant that affect the function of the OBD of the CasX protein can increase the affinity of the CasX variant protein that binds to the protospacer adjacent motif (PAM), and also increase the ability to bind to or utilize an increased spectrum of PAM sequences other than the canonical TTC PAM recognized by the reference CasX protein of SEQ ID NO: 2, which includes a PAM sequence selected from the group consisting of TTC, ATC, GTC, and CTC, thereby increasing the affinity and diversity of the CasX variant protein for the target DNA sequence and resulting in a substantial increase in the target nucleic acid sequences that can be edited and / or bound compared to the reference CasX. As described more fully below, the increase in the sequences of target nucleic acids that can be edited compared to the reference CasX refers to both the PAM and the protospacer sequences, and their orientation depending on the orientation of the non-target strand. This does not mean that the PAM sequence of the non-target strand rather than the target strand determines cleavage or is mechanically involved in target recognition. For example, when referring to the TTC PAM, it can actually be the complementary GAA sequence required for target cleavage or a certain combination of nucleotides from both strands. In the case of the CasX proteins disclosed herein, the PAM is located 5' to the protospacer and at least a single nucleotide separates the PAM from the first nucleotide of the protospacer. Alternatively, or in addition, gNA changes that affect the function of helical I and / or helical II domains that increase the affinity of the CasX variant protein for the target DNA strand can increase the affinity of the CasX RNP containing the variant gNA for the target DNA.
[0142] Addition or alteration of m.gNA function In some embodiments, the gNA variant comprises a larger structural change that changes the topology of the gNA variant with respect to the reference gRNA, thereby enabling different gNA functionality. For example, in some embodiments, the gNA variant exchanges the endogenous stem loop of the reference gRNA scaffold with a stem loop that interacts with a previously identified stable RNA structure or with a protein or RNA binding partner to recruit additional moieties to CasX, or recruits CasX to a specific location such as inside a viral capsid having a binding partner to this RNA structure. In other situations, RNAs may be recruited to each other as seen in kissing loops, which may result in two CasX proteins being co-localized for more efficient gene editing at the target DNA sequence. Such RNA structures may include MS2, Qβ, U1 hairpin II, Uvsx, PP7, phage replication loops, kissing loop_a, kissing loop_b1, kissing loop_b2, G quadruplex M3q, G quadruplex telomere basket, sarcin-ricin loop, or pseudoknot.
[0143] In some embodiments, the gNA variant includes a terminal fusion partner. Exemplary terminal fusions can include fusion of the gRNA to a self-cleaving ribozyme or a protein-binding motif. As used herein, "ribozyme" refers to an RNA or segment thereof having one or more catalytic activities similar to those of protein enzymes. Exemplary ribozyme catalytic activities can include, for example, cleavage and / or ligation of RNA, cleavage and / or ligation of DNA, or peptide bond formation. In some embodiments, such fusions can improve scaffold folding or recruit DNA repair mechanisms. For example, in some embodiments, the gRNA can be fused to a hepatitis delta virus (HDV) antigenomic ribozyme, an HDV genomic ribozyme, a hatchet ribozyme (from metagenomic data), an env25 pistol ribozyme (representative from Aliistipes putredinis), an HH15 minimal hammerhead ribozyme, a tobacco ring spot virus (TRSV) ribozyme, a wild-type viral hammerhead ribozyme (and reasonable variants), or a Twister sister 1 or RBMX recruitment motif. The hammerhead ribozyme is an RNA motif that catalyzes reversible cleavage and ligation reactions at specific sites within an RNA molecule. Hammerhead ribozymes include type I, type II, and type III hammerhead ribozymes. The HDV, pistol, and hatchet ribozymes have self-cleavage activity. A gNA variant comprising one or more ribozymes can enable an extended gNA function compared to a gRNA reference. For example, a gNA comprising a self-cleaving ribozyme can, in some embodiments, be transcribed and processed into a mature gNA as part of a polycistronic transcript. Such fusions can occur at either the 5' or 3' end of the gNA. In some embodiments, the gNA variant includes fusions at both the 5' and 3' ends, and these fusions are each independently as described herein. In some embodiments, the gNA variant includes a phage replication loop or a tetraloop. In some embodiments, the gNA includes a hairpin loop that can bind to a protein.For example, in some embodiments, the hairpin loop is an MS2, Qβ, U1 hairpin II, Uvsx, or PP7 hairpin loop.
[0144] In some embodiments, the gNA variant comprises one or more RNA aptamers. As used herein, "RNA aptamer" refers to an RNA molecule that binds to a target with high affinity and high specificity.
[0145] In some embodiments, the gNA variant comprises one or more riboswitches. As used herein, "riboswitch" refers to an RNA molecule that changes its state upon binding to a small molecule.
[0146] In some embodiments, the gNA variant further comprises one or more protein-binding motifs. By adding a protein-binding motif to the reference gRNA or gNA variant of the present disclosure, in some embodiments, it may be possible for the CasX RNP to associate with additional proteins, for example, the functionality of those proteins can be added to the CasX RNP.
[0147] n. Chemically modified gNA In some embodiments, the present disclosure relates to chemically modified gNAs. In some embodiments, the present disclosure provides chemically modified gNAs that have guide RNA functionality and have reduced susceptibility to nuclease cleavage. A gNA that includes any nucleotide other than the four standard ribonucleotides A, C, G, and U, or deoxynucleotides, is a chemically modified gNA. In some cases, the chemically modified gNA includes any backbone or internucleotide linkage other than the native phosphodiester internucleotide linkage. In certain embodiments, the functionality retained includes the ability of the modified gNA to bind to CasX of any of the embodiments described herein. In certain embodiments, the functionality retained includes the ability of the modified gNA to bind to a target nucleic acid sequence. In certain embodiments, the functionality retained includes the ability to target the CasX protein or the ability of a pre-complexed CasX protein-gNA to bind to a target nucleic acid sequence. In certain embodiments, the functionality retained includes the ability to nick a target polynucleotide by CasX-gNA. In certain embodiments, the functionality retained includes the ability to cleave a target nucleic acid sequence by CasX-gNA. In certain embodiments, the functionality retained is any other known function of the gNA in a CasX system having the CasX protein of the embodiments of the present disclosure.
[0148] In some embodiments, the present disclosure provides that the nucleotide sugar modification is 2′-O-C 1-4 alkyl, e.g., 2′-O-methyl (2′-OMe), 2′-deoxy (2′-H), 2′-O-C 1-3 alkyl-O-C 1-3Provided is a chemically modified gNA incorporated into a gNA selected from the group consisting of alkyl, such as 2′-methoxyethyl (“2′-MOE”), 2′-fluoro (“2′-F”), 2′-amino (“2′-NH2”), 2′-arabinosyl (“2′-arabinose”) nucleotide, 2′-F-arabinosyl (“2′-F-arabinose”) nucleotide, 2′-locked nucleic acid (“LNA”) nucleotide, 2′-unlocked nucleic acid (“ULNA”) nucleotide, L-type sugar (“L-sugar”), and 4′-thiothribosyl nucleotide. In other embodiments, the internucleotide linkage modification incorporated into the guide RNA is phosphorothioate “P(S)” (P(S)), phosphonocarboxylate (P(CH2) n COOR), such as phosphonoacetate “PACE” (P(CH2COO - )), thiophosphonocarboxylate ((S)P(CH2) n COOR), such as thiophosphonoacetate “thioPACE” ((S)P(CH2) n COO - )), alkylphosphonate (P(C 1-3 alkyl)), such as methylphosphonate - P(CH3), boranophosphonate (P(BH3)), and phosphorodithioate (P(S)2), and is selected from the group consisting thereof.
[0149] In certain embodiments, the present disclosure provides chemically modified gNA incorporated into a gNA selected from the group consisting of nucleic acid base ( "base") modifications including 2-thiouracil ( "2-thio U"), 2-thiocytosine ( "2-thio C"), 4-thiouracil ( "4-thio U"), 6-thioguanine ( "6-thio G"), 2-aminoadenine ( "2-amino A"), 2-aminopurine, pseudouracil, hypoxanthine, 7-deazaguanine, 7-deaza-8-azaguanine, 7-deazaadenine, 7-deaza-8-azaadenine, 5-methylcytosine ( "5-methyl C"), 5-methyluracil ( "5-methyl U"), 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5,6-dihydrouracil, 5-propynylcytosine, 5-propynyluracil, 5-ethynylcytosine, 5-ethynyluracil, 5-allyluracil ( "5-allyl U"), 5-allylcytosine ( "5-allyl C"), 5-aminoallyluracil ( "5-aminoallyl U"), 5-aminoallyl-cytosine ( "5-aminoallyl C"), abasic nucleotides, Z bases, P bases, unstructured nucleic acids ( "UNA"), isoguanine ( "iso G"), isocytosine ( "iso C"), 5-methyl-2-pyrimidine, x (A, G, C, T), and y (A, G, C, T).
[0150] In other embodiments, the present disclosure provides chemically modified gNA having one or more isotope modifications introduced into a nucleotide sugar, nucleic acid base, phosphodiester bond, and / or nucleotide phosphate, the one or more isotope modifications including one or more 15 N, 13 C, 14 C, deuterium, 3 H, 32 P, 125 I, 131 I atoms, or other atoms or elements used as tracers.
[0151] In some embodiments, the "terminal" modification incorporated into the gNA is selected from the group consisting of PEG (polyethylene glycol), hydrocarbon linkers (heteroatom (O, S, N) substituted hydrocarbon spacers, halo-substituted hydrocarbon spacers, keto-, carboxyl-, amide-, thionyl-, carbamoyl-, thiocarbamoyl-containing hydrocarbon spacers), spermine linkers, dyes (e.g., fluorescein, rhodamine, cyanine) conjugated to linkers such as 6-fluorescein-hexyl, quenchers (e.g., dabcyl, BHQ), and other labels (e.g., biotin, digoxigenin, acridine, streptavidin, avidin, peptides, and / or proteins). In some embodiments, the "terminal" modification includes conjugation (or ligation) of the gNA to another molecule including oligonucleotides of deoxynucleotides and / or ribonucleotides, peptides, proteins, sugars, oligosaccharides, steroids, lipids, folic acid, vitamins, and / or other molecules. In certain embodiments, the present disclosure provides a chemically modified gNA wherein the "terminal" modification (as described above) is incorporated as a phosphodiester bond and can be incorporated anywhere between two nucleotides within the gNA, e.g., located internally within the gNA sequence via a linker such as 2-(4-butylamidofluorescein)propane-1,3-diol bis(phosphodiester) linker.
[0152] In some embodiments, the present disclosure provides amines, thiols (or sulfhydryls), hydroxyls, carboxyls, carbonyls, thionyls, thiocarbonyls, carbamoyls, thiocarbamoyls, phosphoryls, alkenes, alkynes, halogens, or fluorescent dyes, non-fluorescent labels, tags ( 14 in the case of C, e.g., biotin, avidin, streptavidin, or 15 N, 13 C, deuterium, 3 H, 32 P, 125Provide chemically modified gNA having a terminal modification including a terminal functional group such as a functional group terminal linker that can be conjugated later to a desired moiety selected from the group consisting of isotopic labels such as I, oligonucleotides (including aptamers, including deoxynucleotides and / or ribonucleotides), amino acids, peptides, proteins, sugars, oligosaccharides, steroids, lipids, folic acid, and vitamins. Conjugation uses standard chemistry well known in the art, including but not limited to coupling by any other standard method described in “Bioconjugate Techniques” by Greg T. Hermanson, Publisher Eslsevier Science, 3 rd ed. (2013) (the entire content of which is incorporated herein by reference).
[0153] IV. Proteins for Modifying Target Nucleic Acids The present disclosure provides systems comprising CRISPR nucleases useful for genome editing of eukaryotic cells. In some embodiments, the CRISPR nuclease is selected from the group consisting of Cas9, Cas12a, Cas12b, Cas12c, Cas12d (CasY), CasX, Cas13a, Cas13b, Cas13c, Cas13d, CasX, CasY, Cas14, Cpfl, C2cl, Csn2, and Cas Phi. In some embodiments, the CRISPR nuclease is a type V CRISPR nuclease. In some embodiments, the present disclosure provides a system comprising a CasX protein and one or more guide nucleic acids (gNA) that are specifically designed to modify a target nucleic acid sequence in a eukaryotic cell.
[0154] As used herein, the term "CasX protein" refers to a family of proteins and includes all naturally occurring CasX proteins, proteins that share at least 50% identity with a naturally occurring CasX protein, and CasX variants that have one or more improved properties compared to a naturally occurring reference CasX protein. CasX proteins belong to the CRISPR-Cas type V proteins. Exemplary improved properties of CasX variant embodiments include, more fully described below, improved folding of the variant, improved binding affinity for gNA, improved binding affinity for a target nucleic acid, improved ability to utilize a broader spectrum of PAM sequences in the editing and / or binding of a target DNA, improved unwinding of a target DNA, increased editing activity, improved editing efficiency, improved editing specificity, increased percentage of eukaryotic genomes that can be effectively edited, increased nuclease activity, increased target strand loading for double-strand breaks, decreased target strand loading for single-strand nicking, decreased off-target cleavage, improved binding of the non-target strand of DNA, improved protein stability, improved protein:gNA (RNP) complex stability, improved protein solubility, improved protein:gNA (RNP) complex solubility, improved protein yield, improved protein expression, and improved fusion properties, but are not limited thereto. In the foregoing embodiments, one or more of the improved properties of the RNP of the CasX variant and the gNA variant, when assayed in an equivalent manner, are improved by at least about 1.1- to about 100,000-fold compared to the RNP of the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 and the gNA of Table 1. In other instances, one or more of the improved properties of the RNP of the CasX variant and the gNA variant are improved by at least about 1.1, at least about 10, at least about 100, at least about 1000, at least about 10,000, at least about 100,000-fold, or more compared to the RNP of the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 and the gNA of Table 1.In other cases, one or more of the improved properties of the CasX variant and gNA variant RNPs, when assayed in an equivalent manner, are improved by about 1.1 to 100,000-fold, about 1.1 to 10,000-fold, about 1.1 to 1,000-fold, about 1.1 to 500-fold, about 1.1 to 100-fold, about 1.1 to 50-fold, about 1.1 to 20-fold, about 10 to 100,000-fold, about 10 to 10,000-fold, about 10 to 1,000-fold, about 10 to 500-fold, about 10 to 100-fold, about 10 to 50-fold, about 10 to 20-fold, about 2 to 70-fold, about 2 to 50-fold, about 2 to 30-fold, about 2 to 20-fold, about 2 to 10-fold, about 5 to 50-fold, about 5 to 30-fold, about 5 to 10-fold, about 100 to 100,000-fold, about 100 to 10,000-fold, about 100 to 1,000-fold, about 100 to 500-fold, about 500 to 100,000-fold, about 500 to 10,000-fold, about 500 to 1,000-fold, about 500 to 750-fold, about 1,000 to 100,000-fold, about 10,000 to 100,000-fold, about 20 to 500-fold, about 20 to 250-fold, about 20 to 200-fold, about 20 to 100-fold, about 20 to 50-fold, about 50 to 10,000-fold, about 50 to 1,000-fold, about 50 to 500-fold, about 50 to 200-fold, or about 50 to 100-fold compared to the RNP of the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 and the gNA of Table 1. In other cases, one or more of the improved properties of the CasX variant and gNA variant RNPs, when assayed in an equivalent manner, are improved by about 1.1-fold, 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, 15-fold, 16-fold, 17-fold, 18-fold, 19-fold, 20-fold, 25-fold, 30-fold, 40-fold, 45-fold, 50-fold, 55-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, 110-fold, 120-fold, 130-fold, 140-fold, 150-fold, 160-fold, 170-fold, 180-fold, 190-fold, 200-fold, 210-fold, 220-fold, 230-fold, 240-fold, 250-fold, 260-fold, 270-fold, 280-fold, 290-fold, 300-fold, 310-fold, 320-fold, 330-fold, 340-fold, 350-fold, 360-fold, 370-fold, 380-fold, 390-fold, 400-fold, 425-fold, 450-fold, 475-fold, or 500-fold compared to the RNP of the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 and the gNA of Table 1.
[0155] The term CasX variant includes variants that are fusion proteins, i.e., CasX is "fused" to a heterologous sequence. This includes CasX variants that include a CasX variant sequence and an N-terminal, C-terminal, or internal fusion to a heterologous protein or domain of CasX.
[0156] The CasX proteins of the present disclosure include at least one of the following domains: a non-target strand binding (NTSB) domain, a target strand loading (TSL) domain, a helical I domain, a helical II domain, an oligonucleotide binding domain (OBD), and an RuvC DNA cleavage domain (the last of which may be modified or deleted in catalytically dead CasX variants), as described in more detail below. In addition, the CasX variant proteins of the present disclosure have an enhanced ability to effectively edit and / or bind to target DNA using a PAM sequence selected from TTC, ATC, GTC, or CTC, as compared to the RNP of a reference CasX protein and a reference gNA. In some embodiments, the PAM sequence includes a TC motif. As described above, the PAM sequence is located at at least one nucleotide on the 5' side of the non-target strand of the protospacer having identity to the target sequence of the gNA in the assay system, as compared to the editing efficiency and / or binding of the RNP comprising the reference CasX protein and the reference gNA in an equivalent assay system. In one embodiment, the RNP of the CasX variant and the gNA variant exhibits higher editing efficiency and / or binding of the target sequence in the target DNA, as compared to the RNP comprising the reference CasX protein and the reference gNA in an equivalent assay system, and the PAM sequence of the target DNA is TTC. In another embodiment, the RNP of the CasX variant and the gNA variant exhibits higher editing efficiency and / or binding of the target sequence in the target DNA, as compared to the RNP comprising the reference CasX protein and the reference gNA in an equivalent assay system, and the PAM sequence of the target DNA is ATC. In another embodiment, the RNP of the CasX variant and the gNA variant exhibits higher editing efficiency and / or binding of the target sequence in the target DNA, as compared to the RNP comprising the reference CasX protein and the reference gNA in an equivalent assay system, and the PAM sequence of the target DNA is CTC.In another embodiment, the RNP of the CasX variant and the gNA variant exhibits higher editing efficiency and / or binding of the target sequence in the target DNA as compared to the RNP containing the reference CasX protein and the reference gNA in an equivalent assay system, and the PAM sequence of the target DNA is GTC. In the foregoing embodiments, the increase in editing efficiency and / or binding affinity for one or more PAM sequences is at least 1.5-fold higher as compared to the editing efficiency and / or binding affinity of the RNP of one of the CasX proteins of SEQ ID NOs: 1-3 and the gNA of Table 1 for those PAM sequences.
[0157] In some cases, the CasX protein is a naturally occurring protein (e.g., naturally occurring in and isolated from prokaryotic cells). In other embodiments, the CasX protein is not a naturally occurring protein (e.g., the CasX protein is a CasX variant protein, a chimeric protein, etc.). A naturally occurring CasX protein (referred to herein as a "reference CasX protein") functions as an endonuclease that catalyzes double-strand cleavage at a specific sequence of the target double-stranded DNA (dsDNA). Sequence specificity is provided by the target sequence of the associated gNA that forms the complex and hybridizes to the target sequence within the target nucleic acid.
[0158] In some embodiments, the CasX protein can bind to and / or modify (e.g., cleave, nick, methylate, demethylate, etc.) a target nucleic acid and / or a polypeptide associated with the target nucleic acid (e.g., methylation or acetylation of a histone tail). In some embodiments, the CasX protein is catalytically dead but retains the ability to bind to a target nucleic acid. Exemplary catalytically dead CasX proteins include one or more mutations in the active site of the RuvC domain of the CasX protein. In some embodiments, the catalytically dead CasX protein comprises substitutions at residues 672, 769, and / or 935 of SEQ ID NO: 1. In one embodiment, the catalytically dead CasX protein comprises D672A, E769A, and / or D935A substitutions in the reference CasX protein of SEQ ID NO: 1. In other embodiments, the catalytically dead CasX protein comprises substitutions of amino acids 659, 756, and / or 922 in the reference CasX protein of SEQ ID NO: 2. In some embodiments, the catalytically dead CasX protein comprises D659A, E756A, and / or D922A substitutions in the reference CasX protein of SEQ ID NO: 2. In further embodiments, the catalytically dead CasX protein comprises a deletion of all or part of the RuvC domain of the CasX protein. It will be understood that the same aforementioned substitutions can be introduced into the CasX variants of the present disclosure to result in dCasX variants. In one embodiment, all or part of the RuvC domain is deleted from the CasX variant to result in a dCasX variant. The catalytically inactive dCasX variant protein can be used, in some embodiments, for base editing or epigenetic modification.As the affinity for DNA increases, in some embodiments, catalytically inactive dCasX variant proteins can find their target nucleic acids faster, remain bound to the target nucleic acids for longer periods of time, bind to the target nucleic acids in a more stable manner, or combinations thereof, compared to catalytically active CasX, thereby improving those functions of the catalytically dead CasX variant proteins compared to CasX variants that retain cleavage ability.
[0159] a. Non-target strand binding domain The reference CasX protein of the present disclosure includes a non-target strand binding domain (NTSBD). The NTSBD is a domain not previously found in any Cas protein. For example, this domain is not present in Cas proteins such as Cas9, Cas12a / Cpf1, Cas13, Cas14, CASCADE, CSM, or CSY. Without being bound by theory or mechanism, the NTSBD in CasX may enable binding to non-target DNA strands and assist in unwinding of non-target and target strands. The NTSBD is presumed to be involved in unwinding or capturing the unwound non-target DNA strand. The NTSBD is in direct contact with the non-target strand of the CryoEM model structure derived so far and may contain a non-standard zinc finger domain. The NTSBD may also be involved in stabilizing DNA during unwinding, guide RNA invasion, and R-loop formation. In some embodiments, an exemplary NTSBD includes amino acids 101-191 of SEQ ID NO: 1 or amino acids 103-192 of SEQ ID NO: 2. In some embodiments, the NTSBD of the reference CasX protein includes a four-stranded beta sheet.
[0160] b. Target strand loading domain The reference CasX protein of the present disclosure includes a target strand loading (TSL) domain. The TSL domain is a domain not found in certain Cas proteins such as Cas9, Cascade, CSM, or CSY. Without wishing to be bound by theory or mechanism, the TSL domain is thought to be involved in assisting the loading of the target DNA strand into the RuvC active site of the CasX protein. In some embodiments, the TSL serves to position or capture the target strand in a folded state, such that the phosphates of the target strand DNA backbone that are cleavable are positioned at the RuvC active site. The TSL is separated by most of the TSL and includes cys4 (CXXC, CXXC zinc finger / ribbon domain (SEQ ID NO: 48)). In some embodiments, an exemplary TSL includes amino acids 825-934 of SEQ ID NO: 1 or amino acids 813-921 of SEQ ID NO: 2.
[0161] c. Helical I domain The reference CasX protein of the present disclosure includes a helical I domain. Certain Cas proteins other than CasX have domains that can be named in a similar manner. However, in some embodiments, the helical I domain of the CasX protein includes one or more unique structural features, or unique sequences, or combinations thereof, compared to non-CasX proteins. For example, in some embodiments, the helical I domain of the CasX protein includes one or more unique secondary structures compared to domains of other Cas proteins that may have similar names. For example, in some embodiments, the helical I domain of the CasX protein includes one or more alpha helices having unique structures and sequences in terms of arrangement, number, and length compared to other CRISPR proteins. In certain embodiments, the helical I domain is involved in the interaction with the spacer of the bound DNA and guide RNA. Without wishing to be bound by theory, in some cases, it is believed that the helical I domain may contribute to the binding of the protospacer adjacent motif (PAM). In some embodiments, an exemplary helical I domain includes amino acids 57-100 and 192-332 of SEQ ID NO: 1, or amino acids 59-102 and 193-333 of SEQ ID NO: 2. In some embodiments, the helical I domain of the reference CasX protein includes one or more alpha helices.
[0162] d. Helical II domain The reference CasX protein of the present disclosure includes a helical II domain. Certain Cas proteins other than CasX have domains that can be named in a similar manner. However, in some embodiments, the helical II domain of the CasX protein includes one or more unique structural features, or a unique sequence, or a combination thereof, compared to the domains of other Cas proteins that may have similar names. For example, in some embodiments, the helical II domain includes one or more unique structural alpha-helix bundles that align along the target DNA:guide RNA channel. In some embodiments, in CasX containing the helical II domain, the target strand and the guide RNA interact with the helical II (and in some embodiments, the helical I domain) to enable the RuvC domain to access the target DNA. The helical II domain is involved in binding to the guide RNA scaffold stem-loop and the bound DNA. In some embodiments, an exemplary helical II domain includes amino acids 333-509 of SEQ ID NO: 1, or amino acids 334-501 of SEQ ID NO: 2.
[0163] e. Oligonucleotide binding domain The reference CasX protein of the present disclosure includes an oligonucleotide binding domain (OBD). Certain Cas proteins other than CasX have domains that can be named in a similar manner. However, in some embodiments, the OBD comprises one or more unique functional features, or a sequence unique to the CasX protein, or a combination thereof. For example, in some embodiments, the bridging helix (BH), helical I domain, helical II domain, and oligonucleotide binding domain (OBD) together are involved in the binding of the CasX protein to the guide RNA. Thus, for example, in some embodiments, the OBD is unique to the CasX protein in that it functionally interacts with the helical I domain, or the helical II domain, or both, each of which can be unique to the CasX protein described herein. Specifically, in CasX, the OBD binds primarily to the RNA triplex of the guide RNA scaffold. The OBD may also be involved in binding to the protospacer adjacent motif (PAM). Exemplary OBD domains include amino acids 1-56 and 510-660 of SEQ ID NO: 1, or amino acids 1-58 and 502-647 of SEQ ID NO: 2.
[0164] f.RuvC DNA cleavage domain The reference CasX protein of the present disclosure includes a RuvC domain that includes two partial RuvC domains (RuvC-I and RuvC-II). The RuvC domain is an ancestral domain of all type 12 CRISPR proteins. The RuvC domain is derived from TNPB (transposase B) like a transposase. Similar to other RuvC domains, the CasX RuvC domain has a DED catalytic triad structure involved in the regulation of magnesium (Mg) ions and the cleavage of DNA. In some embodiments, RuvC is involved in the cleavage of both strands of DNA (cleaving one by one, first cleaving the non-target strand at 11-14 nucleotides (nt) in the target sequence, and then likely cleaving the target strand at 2-4 nucleotides after the target sequence) and has a DED motif active site. Particularly in CasX, the RuvC domain is unique in that it is also involved in the binding of the guide RNA scaffold stem-loop important for CasX function. Exemplary RuvC domains include amino acids 661-824 and 935-986 of SEQ ID NO: 1, or amino acids 648-812 and 922-978 of SEQ ID NO: 2.
[0165] g. Reference CasX protein The present disclosure provides a reference CasX protein. In some embodiments, the reference CasX protein is a naturally occurring protein. For example, the reference CasX protein can be isolated from naturally occurring prokaryotes such as Deltaproteobacteria species, Planctomycetes species, or Candidatus Sungbacteria species. The reference CasX protein (also referred to herein as the reference CasX protein) is a type II CRISPR / Cas endonuclease belonging to the CasX (sometimes referred to as Cas12e) family of proteins that can interact with guide RNA to form a ribonucleoprotein (RNP). In some embodiments, the RNP complex comprising the reference CasX protein can target a specific site within a target nucleic acid by base pairing between the target sequence (or spacer) of the gRNA and the target sequence within the target nucleic acid. In some embodiments, the RNP comprising the reference CasX protein can cleave the target DNA. In some embodiments, the RNP comprising the reference CasX protein can nick the target DNA. In some embodiments, the RNP comprising the reference CasX protein can edit the target DNA. For example, in embodiments where the reference CasX protein can cleave or nick DNA, non-homologous end joining (NHEJ), homology-directed repair (HDR), homology-independent target integration (HITI), microhomology-mediated end joining (MMEJ), single-strand annealing (SSA), or base excision repair (BER) follows. In some embodiments, the RNP comprising the CasX protein is a catalytically dead (catalytically inactive or having substantially no cleavage activity) CasX protein (dCasX) as described more fully above, but retains the ability to bind to the target DNA.
[0166] In some cases, the reference CasX protein is isolated from or derived from Deltaproteobacteria. In some embodiments, the CasX protein comprises a sequence that is at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to the following sequence: 1 MEKRINKIRK KLSADNATKP VSRSGPMKTL LVRVMTDDLK KRLEKRRKKP EVMPQVISNN 61 AANNLRMLLD DYTKMKEAIL QVYWQEFKDD HVGLMCKFAQ PASKKIDQNK LKPEMDEKGN 121 LTTAGFACSQ CGQPLFVYKL EQVSEKGKAY TNYFGRCNVA EHEKLILLAQ LKPEKDSDEA 181 VTYSLGKFGQ RALDFYSIHV TKESTHPVKP LAQIAGNRYA SGPVGKALSD ACMGTIASFL 241 SKYQDIIIEH QKVVKGNQKR LESLRELAGK ENLEYPSVTL PPQPHTKEGV DAYNEVIARV 301 RMWVNLNLWQ KLKLSRDDAK PLLRLKGFPS FPVVERRENE VDWWNTINEV KKLIDAKRDM 361 GRVFWSGVTA EKRNTILEGY NYLPNENDHK KREGSLENPK KPAKRQFGDL LLYLEKKYAG 421 DWGKVFDEAW ERIDKKIAGL TSHIEREEAR NAEDAQSKAV LTDWLRAKAS FVLERLKEMD 481 EKEFYACEIQ LQKWYGDLRG NPFAVEAENR VVDISGFSIG SDGHSIQYRN LLAWKYLENG 541 KREFYLLMNY GKKGRIRFTD GTDIKKSGKW QGLLYGGGKA KVIDLTFDPD DEQLIILPLA 601 FGTRQGREFI WNDLLSLETG LIKLANGRVI EKTIYNKKIG RDEPALFVAL TFERREVVDP 661 SNIKPVNLIG VDRGENIPAV IALTDPEGCP LPEFKDSSGG PTDILRIGEG YKEKQRAIQA 721 AKEVEQRRAG GYSRKFASKS RNLADDMVRN SARDLFYHAV THDAVLVFEN LSRGFGRQGK 781 RTFMTERQYT KMEDWLTAKL AYEGLTSKTY LSKTLAQYTS KTCSNCGFTI TTADYDGMLV 841 RLKKTSDGWA TTLNNKELKA EGQITYYNRY KRQTVEKELS AELDRLSEES GNNDISKWTK 901 GRRDEALFLL KKRFSHRPVQ EQFVCLDCGH EVHADEQAAL NIARSWLFLN SNSTEFKSYK 961 SGKQPFVGAW QAFYKRRLKE VWKPNA (Sequence number 1)
[0167] In some cases, the reference CasX protein is isolated from or derived from Planctomycetes. In some embodiments, the CasX protein comprises a sequence that is at least 50%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 86%, at least 87%, at least 88%, at least 89%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to the following sequence: 1 MQEIKRINKI RRRLVKDSNT KKAGKTGPMK TLLVRVMTPD LRERLENLRK KPENIPQPIS 61 NTSRANLNKL LTDYTEMKKA ILHVYWEEFQ KDPVGLMSRV AQPAPKNIDQ RKLIPVKDGN 121 ERLTSSGFAC SQCCQPLYVY KLEQVNDKGK PHTNYFGRCN VSEHERLILL SPHKPEANDE 181 LVTYSLGKFG QRALDFYSIH VTRESNHPVK PLEQIGGNSC ASGPVGKALS DACMGAVASF 241 LTKYQDIILE HQKVIKKNEK RLANLKDIAS ANGLAFPKIT LPPQPHTKEG IEAYNNVVAQ 301 IVIWVNLNLW QKLKIGRDEA KPLQRLKGFP SFPLVERQAN EVDWWDMVCN VKKLINEKKE 361 DGKVFWQNLA GYKRQEALLP YLSSEEDRKK GKKFARYQFG DLLLHLEKKH GEDWGKVYDE 421 AWERIDKKVE GLSKHIKLEE ERRSEDAQSK AALTDWLRAK ASFVIEGLKE ADKDEFCRCE 481 LKLQKWYGDL RGKPFAIEAE NSILDISGFS KQYNCAFIWQ KDGVKKLNLY LIINYFKGGK 541 LRFKKIKPEA FEANRFYTVI NKKSGEIVPM EVNFNFDDPN LIILPLAFGK RQGREFIWND 601 LLSLETGSLK LANGRVIEKT LYNRRTRQDE PALFVALTFE RREVLDSSNI KPMNLIGIDR 661 GENIPAVIAL TDPEGCPLSR FKDSLGNPTH ILRIGESYKE KQRTIQAAKE VEQRRAGGYS 721 RKYASKAKNL ADDMVRNTAR DLLYYAVTQD AMLIFENLSR GFGRQGKRTF MAERQYTRME 781 DWLTAKLAYE GLPSKTYLSK TLAQYTSKTC SNCGFTITSA DYDRVLEKLK KTATGWMTTI 841 NGKELKVEGQ ITYYNRYKRQ NVVKDLSVEL DRLSEESVNN DISSWTKGRS GEALSLLKKR 901 FSHRPVQEKF VCLNCGFETH ADEQAALNIA RSWLFLRSQE YKKYQTNKTT GNTDKRAFVE 961 TWQSFYRKKL KEVWKPAV (Sequence No. 2)
[0168] In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 2 or at least 60% similarity thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 2 or at least 80% similarity thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 2 or at least 90% similarity thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 2 or at least 95% similarity thereto. In some embodiments, the CasX protein consists of the sequence of SEQ ID NO: 2. In some embodiments, the CasX protein comprises or consists of a sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, or at least 50 mutations compared to the sequence of SEQ ID NO: 2. These mutations can be insertions, deletions, amino acid substitutions, or any combination thereof.
[0169] In some cases, the reference CasX protein is isolated from or derived from Candidatus Sungbacteria. In some embodiments, the CasX protein comprises a sequence that is at least 50%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 86%, at least 87%, at least 88%, at least 89%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to the following sequence: 1 MDNANKPSTK SLVNTTRISD HFGVTPGQVT RVFSFGIIPT KRQYAIIERW FAAVEAARER 61 LYGMLYAHFQ ENPPAYLKEK FSYETFFKGR PVLNGLRDID PTIMTSAVFT ALRHKAEGAM 121 AAFHTNHRRL FEEARKKMRE YAECLKANEA LLRGAADIDW DKIVNALRTR LNTCLAPEYD 181 AVIADFGALC AFRALIAETN ALKGAYNHAL NQMLPALVKV DEPEEAEESP RLRFFNGRIN 241 DLPKFPVAER ETPPDTETII RQLEDMARVI PDTAEILGYI HRIRHKAARR KPGSAVPLPQ 301 RVALYCAIRM ERNPEEDPST VAGHFLGEID RVCEKRRQGL VRTPFDSQIR ARYMDIISFR 361 ATLAHPDRWT EIQFLRSNAA SRRVRAETIS APFEGFSWTS NRTNPAPQYG MALAKDANAP 421 ADAPELCICL SPSSAAFSVR EKGGDLIYMR PTGGRRGKDN PGKEITWVPG SFDEYPASGV 481 ALKLRLYFGR SQARRMLTNK TWGLLSDNPR VFAANAELVG KKRNPQDRWK LFFHMVISGP 541 PPVEYLDFSS DVRSRARTVI GINRGEVNPL AYAVVSVEDG QVLEEGLLGK KEYIDQLIET 601 RRRISEYQSR EQTPPRDLRQ RVRHLQDTVL GSARAKIHSL IAFWKGILAI ERLDDQFHGR 661 EQKIIPKKTY LANKTGFMNA LSFSGAVRVD KKGNPWGGMI EIYPGGISRT CTQCGTVWLA 721 RRPKNPGHRD AMVVIPDIVD DAAATGFDNV DCDAGTVDYG ELFTLSREWV RLTPRYSRVM 781 RGTLGDLERA IRQGDDRKSR QMLELALEPQ PQWGQFFCHR CGFNGQSDVL AATNLARRAI 841 SLIRRLPDTD TPPTP (SEQ ID NO: 3)
[0170] In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 3 or at least 60% similarity thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 3 or at least 80% similarity thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 3 or at least 90% similarity thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 3 or at least 95% similarity thereto. In some embodiments, the CasX protein consists of the sequence of SEQ ID NO: 3. In some embodiments, the CasX protein comprises or consists of a sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, or at least 50 mutations compared to the sequence of SEQ ID NO: 3. These mutations can be insertions, deletions, amino acid substitutions, or any combination thereof.
[0171] h.CasX variant protein The present disclosure provides variants of a reference CasX protein (hereinafter interchangeably referred to as "CasX variant" or "CasX variant protein"), wherein the CasX variant comprises at least one modification in at least one domain of the reference CasX protein comprising the sequences of SEQ ID NOs: 1-3. In some embodiments, the CasX variant exhibits at least one improved property as compared to the reference CasX protein. All variants that improve one or more functions or properties of the CasX variant protein as compared to the reference CasX protein described herein are contemplated to be within the scope of the present disclosure. In some embodiments, the modification is a mutation of one or more amino acids of the reference CasX. In other embodiments, the modification is a substitution of one or more domains of the reference CasX with one or more domains from a different CasX. In some embodiments, the insertion comprises the insertion of part or all of a domain from a different CasX protein. The mutation can occur in any one or more domains of the reference CasX protein and can include, for example, a partial or complete deletion of one or more domains, or a substitution, deletion, or insertion of one or more amino acids in any domain of the reference CasX protein. Domains of the CasX protein include a non-target strand binding (NTSB) domain, a target strand loading (TSL) domain, a helical I domain, a helical II domain, an oligonucleotide binding domain (OBD), and a RuvC DNA cleavage domain. Any change in the amino acid sequence of the reference CasX protein that results in an improved property of the CasX protein is considered to be a CasX variant protein of the present disclosure. For example, the CasX variant can include one or more amino acid substitutions, insertions, deletions, or swapped domains, or any combination thereof, as compared to the reference CasX protein sequence.
[0172] In some embodiments, the CasX variant protein comprises at least one modification in at least each of two domains of a reference CasX protein comprising the sequences of SEQ ID NOs: 1-3. In some embodiments, the CasX variant protein comprises at least one modification in at least two domains, at least three domains, at least four domains, or at least five domains of the reference CasX protein. In some embodiments, the CasX variant protein comprises more than two modifications in at least one domain of the reference CasX protein. In some embodiments, the CasX variant protein comprises at least two modifications in at least one domain of the reference CasX protein, at least three modifications in at least one domain of the reference CasX protein, or at least four modifications in at least one domain of the reference CasX protein. In some embodiments, the CasX variant comprises more than two modifications as compared to the reference CasX protein, and each modification is made in a domain independently selected from the group consisting of the NTSBD, TSLD, helical I domain, helical II domain, OBD, and RuvC DNA cleavage domain.
[0173] In some embodiments, at least one modification of the CasX variant protein comprises a deletion of at least a portion of one domain of a reference CasX protein comprising the sequences of SEQ ID NOs: 1-3. In some embodiments, the deletion is in the NTSBD, TSLD, helical I domain, helical II domain, OBD, or RuvC DNA cleavage domain.
[0174] Suitable mutagenesis methods for generating the CasX variant proteins of the present disclosure can include, for example, deep mutational evolution (DME), deep mutational scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping. In some embodiments, the CasX variant is designed, for example, by selecting one or more desired mutations in a reference CasX. In certain embodiments, the activity of the reference CasX protein is used as a benchmark against which the activity of one or more CasX variants is compared, thereby measuring improvements in the function of the CasX variant. Exemplary improvements in the CasX variant include, but are not limited to, improved folding of the variant, improved binding affinity for gNA, improved binding affinity for target DNA, modified binding affinity for one or more PAM sequences, improved unwinding of the target DNA, increased editing activity, improved efficiency, improved editing specificity, increased nuclease activity, increased target strand loading for double-strand cleavage, decreased target strand loading for single-strand nicking, decreased off-target cleavage, improved binding of the non-target strand of DNA, improved protein stability, improved protein:gNA complex stability, improved protein solubility, improved protein:gNA complex solubility, improved protein yield, improved protein expression, and improved fusion properties, as more fully described below.
[0175] In some embodiments of the CasX variants described herein, at least one modification comprises (a) substitution of 1 to 100 consecutive or non - consecutive amino acids in the CasX variant compared to the reference CasX of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3, (b) deletion of 1 to 100 consecutive or non - consecutive amino acids in the CasX variant compared to the reference CasX, (c) insertion of 1 to 100 consecutive or non - consecutive amino acids in the CasX compared to the reference CasX, or (d) any combination of (a) to (c). In some embodiments, at least one modification comprises (a) substitution of 5 to 10 consecutive or non - consecutive amino acids in the CasX variant compared to the reference CasX of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3, (b) deletion of 1 to 5 consecutive or non - consecutive amino acids in the CasX variant compared to the reference CasX, (c) insertion of 1 to 5 consecutive or non - consecutive amino acids in the CasX compared to the reference CasX, or (d) any combination of (a) to (c).
[0176] In some embodiments, the CasX variant protein comprises, or consists of, a sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, or at least 50 mutations compared to the sequence of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. These mutations can be insertions, deletions, amino acid substitutions, or any combination thereof.
[0177] In some embodiments, the CasX variant protein comprises at least one amino acid substitution in at least one domain of the reference CasX protein. In some embodiments, the CasX variant protein comprises at least about 1 to 4 amino acid substitutions, 1 to 10 amino acid substitutions, 1 to 20 amino acid substitutions, 1 to 30 amino acid substitutions, 1 to 40 amino acid substitutions, 1 to 50 amino acid substitutions, 1 to 60 amino acid substitutions, 1 to 70 amino acid substitutions, 1 to 80 amino acid substitutions, 1 to 90 amino acid substitutions, 1 to 100 amino acid substitutions, 2 to 10 amino acid substitutions, 2 to 20 amino acid substitutions, 2 to 30 amino acid substitutions, 3 to 10 amino acid substitutions, 3 to 20 amino acid substitutions, 3 to 30 amino acid substitutions, 4 to 10 amino acid substitutions, 4 to 20 amino acid substitutions, 3 to 300 amino acid substitutions, 5 to 10 amino acid substitutions, 5 to 20 amino acid substitutions, 5 to 30 amino acid substitutions, 10 to 50 amino acid substitutions, or 20 to 50 amino acid substitutions compared to the reference CasX protein. In some embodiments, the CasX variant protein comprises at least about 100 amino acid substitutions compared to the reference CasX protein. In some embodiments, the CasX variant protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions compared to the reference CasX protein. In some embodiments, the CasX variant protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions in a single domain compared to the reference CasX protein. In some embodiments, the amino acid substitutions are conservative substitutions. In other embodiments, the substitutions are non-conservative, e.g., a polar amino acid is substituted with a non-polar amino acid or vice versa.
[0178] In some embodiments, the CasX variant protein comprises one amino acid substitution, 2 to 3 consecutive amino acid substitutions, 2 to 4 consecutive amino acid substitutions, 2 to 5 consecutive amino acid substitutions, 2 to 6 consecutive amino acid substitutions, 2 to 7 consecutive amino acid substitutions, 2 to 8 consecutive amino acid substitutions, 2 to 9 consecutive amino acid substitutions, 2 to 10 consecutive amino acid substitutions, 2 to 20 consecutive amino acid substitutions, 2 to 30 consecutive amino acid substitutions, 2 to 40 consecutive amino acid substitutions, 2 to 50 consecutive amino acid substitutions, 2 to 60 consecutive amino acid substitutions, 2 to 70 consecutive amino acid substitutions, 2 to 80 consecutive amino acid substitutions, 2 to 90 consecutive amino acid substitutions, 2 to 100 consecutive amino acid substitutions, 3 to 10 consecutive amino acid substitutions, 3 to 20 consecutive amino acid substitutions, 3 to 30 consecutive amino acid substitutions, 4 to 10 consecutive amino acid substitutions, 4 to 20 consecutive amino acid substitutions, 3 to 300 consecutive amino acid substitutions, 5 to 10 consecutive amino acid substitutions, 5 to 20 consecutive amino acid substitutions, 5 to 30 consecutive amino acid substitutions, 10 to 50 consecutive amino acid substitutions, or 20 to 50 consecutive amino acid substitutions as compared to the reference CasX protein. In some embodiments, the CasX variant protein comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive amino acid substitutions. In some embodiments, the CasX variant protein comprises at least about 100 consecutive amino acid substitutions. As used herein, "consecutive amino acids" refers to amino acids that are adjacent in the primary sequence of the polypeptide.
[0179] In some embodiments, the CasX variant protein comprises two or more substitutions compared to the reference CasX protein, and these two or more substitutions are not present in consecutive amino acids of the reference CasX sequence. For example, the first substitution may be in the first domain of the reference CasX protein, and the second substitution may be in the second domain of the reference CasX protein. In some embodiments, the CasX variant protein comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 non - consecutive substitutions compared to the reference CasX protein. In some embodiments, the CasX variant protein comprises at least 20 non - consecutive substitutions compared to the reference CasX protein. Each non - consecutive substitution can be an amino acid of any length described herein, for example, 1 - 4 amino acids, 1 - 10 amino acids, etc. In some embodiments, two or more substitutions compared to the reference CasX protein are not of the same length, for example, the first substitution is 1 amino acid and the second substitution is 3 amino acids. In some embodiments, two or more substitutions compared to the reference CasX protein are of the same length, for example, the length of both of these substitutions is 2 consecutive amino acids.
[0180] In the substitutions described herein, any amino acid can be substituted with any other amino acid. The substitution can be a conservative substitution (e.g., a basic amino acid is substituted with another basic amino acid). The substitution can be a non - conservative substitution (e.g., a basic amino acid is substituted with an acidic amino acid, or vice versa). For example, a proline of the reference CasX protein can be substituted with any one of arginine, histidine, lysine, aspartic acid, glutamic acid, serine, threonine, asparagine, glutamine, cysteine, glycine, alanine, isoleucine, leucine, methionine, phenylalanine, tryptophan, tyrosine, or valine to generate a CasX variant protein of the present disclosure.
[0181] In some embodiments, the CasX variant protein comprises at least one amino acid deletion as compared to the reference CasX protein. In some embodiments, the CasX variant protein comprises a deletion of 1 to 4 amino acids, 1 to 10 amino acids, 1 to 20 amino acids, 1 to 30 amino acids, 1 to 40 amino acids, 1 to 50 amino acids, 1 to 60 amino acids, 1 to 70 amino acids, 1 to 80 amino acids, 1 to 90 amino acids, 1 to 100 amino acids, 2 to 10 amino acids, 2 to 20 amino acids, 2 to 30 amino acids, 3 to 10 amino acids, 3 to 20 amino acids, 3 to 30 amino acids, 4 to 10 amino acids, 4 to 20 amino acids, 3 to 300 amino acids, 5 to 10 amino acids, 5 to 20 amino acids, 5 to 30 amino acids, 10 to 50 amino acids, or 20 to 50 amino acids as compared to the reference CasX protein. In some embodiments, the CasX variant comprises a deletion of at least about 100 contiguous amino acids as compared to the reference CasX protein. In some embodiments, the CasX variant protein comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or 100 contiguous amino acids as compared to the reference CasX protein. In some embodiments, the CasX variant protein comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 contiguous amino acids.
[0182] In some embodiments, the CasX variant protein comprises two or more deletions compared to the reference CasX protein, and these two or more deletions are not contiguous amino acids. For example, the first deletion may be in the first domain of the reference CasX protein, and the second deletion may be in the second domain of the reference CasX protein. In some embodiments, the CasX variant protein comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 non - contiguous deletions compared to the reference CasX protein. In some embodiments, the CasX variant protein comprises at least 20 non - contiguous deletions compared to the reference CasX protein. Each non - contiguous deletion can be an amino acid of any length described herein, for example, 1 - 4 amino acids, 1 - 10 amino acids, etc.
[0183] In some embodiments, the CasX variant protein comprises at least one amino acid insertion. In some embodiments, the CasX variant protein has an insertion of one amino acid, two to three consecutive amino acids, two to four consecutive amino acids, two to five consecutive amino acids, two to six consecutive amino acids, two to seven consecutive amino acids, two to eight consecutive amino acids, two to nine consecutive amino acids, two to ten consecutive amino acids, two to twenty consecutive amino acids, two to thirty consecutive amino acids, two to forty consecutive amino acids, two to fifty consecutive amino acids, two to sixty consecutive amino acids, two to seventy consecutive amino acids, two to eighty consecutive amino acids, two to ninety consecutive amino acids, two to one hundred consecutive amino acids, three to ten consecutive amino acids, three to twenty consecutive amino acids, three to thirty consecutive amino acids, four to ten consecutive amino acids, four to twenty consecutive amino acids, three to three hundred consecutive amino acids, five to ten consecutive amino acids, five to twenty consecutive amino acids, five to thirty consecutive amino acids, ten to fifty consecutive amino acids, or twenty to fifty consecutive amino acids, as compared to the reference CasX protein. In some embodiments, the CasX variant protein comprises an insertion of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive amino acids. In some embodiments, the CasX variant protein comprises an insertion of at least about 100 consecutive amino acids.
[0184] In some embodiments, the CasX variant protein comprises two or more insertions compared to the reference CasX protein, and these two or more insertions are not contiguous amino acids in the sequence. For example, the first insertion can be in the first domain of the reference CasX protein, and the second insertion can be in the second domain of the reference CasX protein. In some embodiments, the CasX variant protein comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 non - contiguous insertions compared to the reference CasX protein. In some embodiments, the CasX variant protein comprises at least 10 to about 20 or more non - contiguous insertions compared to the reference CasX protein. Each non - contiguous insertion can be an amino acid of any length described herein, for example, 1 to 4 amino acids, 1 to 10 amino acids, etc.
[0185] Any amino acid, or any combination of amino acids, can be inserted into the insertions described herein. For example, proline, arginine, histidine, lysine, aspartic acid, glutamic acid, serine, threonine, asparagine, glutamine, cysteine, glycine, alanine, isoleucine, leucine, methionine, phenylalanine, tryptophan, tyrosine or valine, or any combination thereof, can be inserted into the reference CasX protein of the present disclosure to generate a CasX variant protein.
[0186] Any permutation of the substitution, insertion, and deletion embodiments described herein can be combined to generate the CasX variant proteins of the present disclosure. For example, the CasX variant protein can comprise at least one substitution and at least one deletion compared to the reference CasX protein sequence, at least one substitution and at least one insertion compared to the reference CasX protein sequence, at least one insertion and at least one deletion compared to the reference CasX protein sequence, or at least one substitution, one insertion, and one deletion compared to the reference CasX protein.
[0187] In some embodiments, the CasX variant protein has at least about 60% sequence similarity, at least 70% similarity, at least 80% similarity, at least 85% similarity, at least 86% similarity, at least 87% similarity, at least 88% similarity, at least 89% similarity, at least 90% similarity, at least 91% similarity, at least 92% similarity, at least 93% similarity, at least 94% similarity, at least 95% similarity, at least 96% similarity, at least 97% similarity, at least 98% similarity, at least 99% similarity, at least 99.5% similarity, at least 99.6% similarity, at least 99.7% similarity, at least 99.8% similarity, or at least 99.9% similarity with one of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3.
[0188] In some embodiments, the CasX variant protein has at least about 60% sequence similarity with SEQ ID NO: 2 or a part thereof. In some embodiments, the CasX variant protein has a substitution of Y789T in SEQ ID NO: 2, a deletion of P793 in SEQ ID NO: 2, a substitution of Y789D in SEQ ID NO: 2, a substitution of T72S in SEQ ID NO: 2, a substitution of I546V in SEQ ID NO: 2, a substitution of E552A in SEQ ID NO: 2, a substitution of A636D in SEQ ID NO: 2, a substitution of F536S in SEQ ID NO: 2, a substitution of A708K in SEQ ID NO: 2, a substitution of Y797L in SEQ ID NO: 2, a substitution of L792G in SEQ ID NO: 2, a substitution of A739V in SEQ ID NO: 2, a substitution of G791M in SEQ ID NO: 2, an insertion of A at position 661 in SEQ ID NO: 2, a substitution of A788W in SEQ ID NO: 2, a substitution of K390R in SEQ ID NO: 2, a substitution of A751S in SEQ ID NO: 2, a substitution of E385A in SEQ ID NO: 2, an insertion of P at position 696 in SEQ ID NO: 2, an insertion of M at position 773 in SEQ ID NO: 2, a substitution of G695H in SEQ ID NO: 2, an insertion of AS at position 793 in SEQ ID NO: 2, an insertion of AS at position 795 in SEQ ID NO: 2, a substitution of C477R in SEQ ID NO: 2, a substitution of C477K in SEQ ID NO: 2, a substitution of C479A in SEQ ID NO: 2, a substitution of C479L in SEQ ID NO: 2, a substitution of I55F in SEQ ID NO: 2, a substitution of K210R in SEQ ID NO: 2, a substitution of C233S in SEQ ID NO: 2, a substitution of D231N in SEQ ID NO: 2, a substitution of Q338E in SEQ ID NO: 2, a substitution of Q338R in SEQ ID NO: 2, a substitution of L379R in SEQ ID NO: 2, a substitution of K390R in SEQ ID NO: 2, a substitution of L481Q in SEQ ID NO: 2, a substitution of F495S in SEQ ID NO: 2, a substitution of D600N in SEQ ID NO: 2, a substitution of T886K in SEQ ID NO: 2, a substitution of A739V in SEQ ID NO: 2, a substitution of K460N in SEQ ID NO: 2, a substitution of I199F in SEQ ID NO: 2, a substitution of G492P in SEQ ID NO: 2, a substitution of T153I in SEQ ID NO: 2, a substitution of R591I in SEQ ID NO: 2, an insertion of AS at position 795 in SEQ ID NO: 2, an insertion of AS at position 796 in SEQ ID NO: 2, an insertion of L at position 889 in SEQ ID NO: 2, a substitution of E121D in SEQ ID NO: 2, a substitution of S270W in SEQ ID NO: 2, a substitution of E712Q in SEQ ID NO: 2, a substitution of K942Q in SEQ ID NO: 2, a substitution of E552K in SEQ ID NO: 2, a substitution of K25Q in SEQ ID NO: 2, a substitution of N47D in SEQ ID NO: 2, an insertion of T at position 696 in SEQ ID NO: 2, a substitution of L685I in SEQ ID NO: 2, a substitution of N880D in SEQ ID NO: 2, a substitution of Q102R in SEQ ID NO: 2,Substitution of M734K at sequence number 2, substitution of A724S at sequence number 2, substitution of T704K at sequence number 2, substitution of P224K at sequence number 2, substitution of K25R at sequence number 2, substitution of M29E at sequence number 2, substitution of H152D at sequence number 2, substitution of S219R at sequence number 2, substitution of E475K at sequence number 2, substitution of G226R at sequence number 2, substitution of A377K at sequence number 2, substitution of E480K at sequence number 2, substitution of K416E at sequence number 2, substitution of H164R at sequence number 2, substitution of K767R at sequence number 2, substitution of I7F at sequence number 2, substitution of M29R at sequence number 2, substitution of H435R at sequence number 2, substitution of E385Q at sequence number 2, substitution of E385K at sequence number 2, substitution of I279F at sequence number 2, substitution of D489S at sequence number 2, substitution of D732N at sequence number 2, substitution of A739T at sequence number 2, substitution of W885R at sequence number 2, substitution of E53K at sequence number 2, substitution of A238T at sequence number 2, substitution of P283Q at sequence number 2, substitution of E292K at sequence number 2, substitution of Q628E at sequence number 2, substitution of R388Q at sequence number 2, substitution of G791M at sequence number 2, substitution of L792K at sequence number 2, substitution of L792E at sequence number 2, substitution of M779N at sequence number 2, substitution of G27D at sequence number 2, substitution of K955R at sequence number 2, substitution of S867R at sequence number 2, substitution of R693I at sequence number 2, substitution of F189Y at sequence number 2, substitution of V635M at sequence number 2, substitution of F399L at sequence number 2, substitution of E498K at sequence number 2, substitution of E386R at sequence number 2, substitution of V254G at sequence number 2, substitution of P793S at sequence number 2, substitution of K188E at sequence number 2, substitution of QT945KI at sequence number 2, substitution of T620P at sequence number 2, substitution of T946P at sequence number 2, substitution of TT949PP at sequence number 2, substitution of N952T at sequence number 2, substitution of K682E at sequence number 2, substitution of K975R at sequence number 2, substitution of L212P at sequence number 2, substitution of E292R at sequence number 2, substitution of I303K at sequence number 2, substitution of C349E at sequence number 2, substitution of E385P at sequence number 2, substitution of E386N at sequence number 2, substitution of D387K at sequence number 2, substitution of L404K at sequence number 2, substitution of E466H at sequence number 2, substitution of C477Q at sequence number 2, substitution of C477H at sequence number 2, substitution of C479A at sequence number 2Substitution of D659H at sequence number 2, substitution of T806V at sequence number 2, substitution of K808S at sequence number 2, insertion of AS at position 797 of sequence number 2, substitution of V959M at sequence number 2, substitution of K975Q at sequence number 2, substitution of W974G at sequence number 2, substitution of A708Q at sequence number 2, substitution of V711K at sequence number 2, substitution of D733T at sequence number 2, substitution of L742W at sequence number 2, substitution of V747K at sequence number 2, substitution of F755M at sequence number 2, substitution of M771A at sequence number 2, substitution of M771Q at sequence number 2, substitution of W782Q at sequence number 2, substitution of G791F at sequence number 2, substitution of L792D at sequence number 2, substitution of L792K at sequence number 2, substitution of P793Q at sequence number 2, substitution of P793G at sequence number 2, substitution of Q804A at sequence number 2, substitution of Y966N at sequence number 2, substitution of Y723N at sequence number 2, substitution of Y857R at sequence number 2, substitution of S890R at sequence number 2, substitution of S932M at sequence number 2, substitution of L897M at sequence number 2, substitution of R624G at sequence number 2, substitution of S603G at sequence number 2, substitution of N737S at sequence number 2, substitution of L307K at sequence number 2, substitution of I658V at sequence number 2, insertion of PT at position 688 of sequence number 2, insertion of SA at position 794 of sequence number 2, substitution of S877R at sequence number 2, substitution of N580T at sequence number 2, substitution of V335G at sequence number 2, substitution of T620S at sequence number 2, substitution of W345G at sequence number 2, substitution of T280S at sequence number 2, substitution of L406P at sequence number 2, substitution of A612D at sequence number 2, substitution of A751S at sequence number 2, substitution of E386R at sequence number 2, substitution of V351M at sequence number 2, substitution of K210N at sequence number 2, substitution of D40A at sequence number 2, substitution of E773G at sequence number 2, substitution of H207L at sequence number 2, substitution of T62A at sequence number 2, substitution of T287P at sequence number 2, substitution of T832A at sequence number 2, substitution of A893S at sequence number 2, insertion of V at position 14 of sequence number 2, insertion of AG at position 13 of sequence number 2, substitution of R11V at sequence number 2, substitution of R12N at sequence number 2, substitution of R13H at sequence number 2, insertion of Y at position 13 of sequence number 2, substitution of R12L at sequence number 2, insertion of Q at position 13 of sequence number 2, substitution of V15S at sequence number 2, insertion of D at position 17 of sequence number 2, or combinations thereof.,
[0189] In some embodiments, the CasX variant comprises at least one modification in the NTSB domain.
[0190] In some embodiments, the CasX variant comprises at least one modification in the TSL domain. In some embodiments, at least one modification in the TSL domain comprises one or more amino acid substitutions of amino acid Y857, S890, or S932 of SEQ ID NO: 2.
[0191] In some embodiments, the CasX variant comprises at least one modification in the Helical I domain. In some embodiments, at least one modification in the Helical I domain comprises one or more amino acid substitutions of amino acid S219, L249, E259, Q252, E292, L307, or D318 of SEQ ID NO: 2.
[0192] In some embodiments, the CasX variant comprises at least one modification in the Helical II domain. In some embodiments, at least one modification in the Helical II domain comprises one or more amino acid substitutions of amino acid D361, L379, E385, E386, D387, F399, L404, R458, C477, or D489 of SEQ ID NO: 2.
[0193] In some embodiments, the CasX variant comprises at least one modification in the OBD domain. In some embodiments, at least one modification in the OBD comprises one or more amino acid substitutions of amino acid F536, E552, T620, or I658 of SEQ ID NO: 2.
[0194] In some embodiments, the CasX variant comprises at least one modification in the RuvC DNA cleavage domain. In some embodiments, at least one modification in the RuvC DNA cleavage domain comprises one or more amino acid substitutions of amino acid K682, G695, A708, V711, D732, A739, D733, L742, V747, F755, M771, M779, W782, A788, G791, L792, P793, Y797, M799, Q804, S819, or Y857 of SEQ ID NO: 2, or a deletion of amino acid P793.
[0195] In some embodiments, the CasX variant comprises at least one modification as compared to the reference CasX sequence of SEQ ID NO: 2, and is selected from one or more of (a) an amino acid substitution of L379R, (b) an amino acid substitution of A708K, (c) an amino acid substitution of T620P, (d) an amino acid substitution of E385P, (e) an amino acid substitution of Y857R, (f) an amino acid substitution of I658V, (g) an amino acid substitution of F399L, (h) an amino acid substitution of Q252K, (i) an amino acid substitution of L404K, and (j) an amino acid deletion of P793.
[0196] In some embodiments, the CasX variant protein comprises at least two amino acid changes relative to the reference CasX protein amino acid sequence. The at least two amino acid changes can be substitutions, insertions, or deletions in the reference CasX protein amino acid sequence, or any combination thereof. The substitution, insertion, or deletion can be any substitution, insertion, or deletion in the sequence of the reference CasX protein described herein. In some embodiments, the changes are adjacent amino acid changes, non - adjacent amino acid changes, or a combination of adjacent and non - adjacent amino acid changes relative to the reference CasX protein sequence. In some embodiments, the reference CasX protein is SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 amino acid changes relative to the reference CasX protein. In some embodiments, the CasX variant protein comprises 1 - 50, 3 - 40, 5 - 30, 5 - 20, 5 - 15, 5 - 10, 10 - 50, 10 - 40, 10 - 30, 10 - 20, 15 - 50, 15 - 40, 15 - 30, 2 - 25, 2 - 24, 2 - 22, 2 - 23, 2 - 22, 2 - 21, 2 - 20, 2 - 19, 2 - 18, 2 - 17, 2 - 16, 2 - 15, 2 - 14, 2 - 12, 2 - 11, 2 - 10, 2 - 9, 2 - 8, 2 - 7, 2 - 6, 2 - 5, 2 - 4, 2 - 3, 3 - 25, 3 - 24, 3 - 22, 3 - 23, 3 - 22, 3 - 21, 3 - 20, 3 - 19, 3 - 18, 3 - 17, 3 - 16, 3 - 15, 3 - 14, 3 - 12, 3 - 11, 3 - 10, 3 - 9, 3 - 8, 3 - 7, 3 - 6, 3 - 5, 3 - 4,comprise 4 to 25, 4 to 24, 4 to 22, 4 to 23, 4 to 22, 4 to 21, 4 to 20, 4 to 19, 4 to 18, 4 to 17, 4 to 16, 4 to 15, 4 to 14, 4 to 12, 4 to 11, 4 to 10, 4 to 9, 4 to 8, 4 to 7, 4 to 6, 4 to 5, 5 to 25, 5 to 24, 5 to 22, 5 to 23, 5 to 22, 5 to 21, 5 to 20, 5 to 19, 5 to 18, 5 to 17, 5 to 16, 5 to 15, 5 to 14, 5 to 12, 5 to 11, 5 to 10, 5 to 9, 5 to 8, 5 to 7, or 5 to 6 amino acid changes. In some embodiments, the CasX variant protein comprises 15 to 20 changes relative to the reference CasX protein sequence. In some embodiments, the CasX variant protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acid changes relative to the reference CasX protein sequence. In some embodiments, at least two amino acid changes to the sequence of the reference CasX variant protein are the substitution of Y789T of SEQ ID NO: 2, the deletion of P793 of SEQ ID NO: 2, the substitution of Y789D of SEQ ID NO: 2, the substitution of T72S of SEQ ID NO: 2, the substitution of I546V of SEQ ID NO: 2, the substitution of E552A of SEQ ID NO: 2, the substitution of A636D of SEQ ID NO: 2, the substitution of F536S of SEQ ID NO: 2, the substitution of A708K of SEQ ID NO: 2, the substitution of Y797L of SEQ ID NO: 2, the substitution of L792G of SEQ ID NO: 2, the substitution of A739V of SEQ ID NO: 2, the substitution of G791M of SEQ ID NO: 2, the insertion of A at position 661 of SEQ ID NO: 2, the substitution of A788W of SEQ ID NO: 2, the substitution of K390R of SEQ ID NO: 2, the substitution of A751S of SEQ ID NO: 2, the substitution of E385A of SEQ ID NO: 2, the insertion of P at position 696 of SEQ ID NO: 2, the insertion of M at position 773 of SEQ ID NO: 2, the substitution of G695H of SEQ ID NO: 2, the insertion of AS at position 793 of SEQ ID NO: 2, the insertion of AS at position 795 of SEQ ID NO: 2, the substitution of C477R of SEQ ID NO: 2, the substitution of C477K of SEQ ID NO: 2, the substitution of C479A of SEQ ID NO: 2, the substitution of C479L of SEQ ID NO: 2, the substitution of I55F of SEQ ID NO: 2, the substitution of K210R of SEQ ID NO: 2, the substitution of C233S of SEQ ID NO: 2, the substitution of D231N of SEQ ID NO: 2, the substitution of Q338E of SEQ ID NO: 2, the substitution of Q338R of SEQ ID NO: 2, the substitution of L379R of SEQ ID NO: 2, the substitution of K390R of SEQ ID NO: 2,Substitution of L481Q at sequence number 2, substitution of F495S at sequence number 2, substitution of D600N at sequence number 2, substitution of T886K at sequence number 2, substitution of A739V at sequence number 2, substitution of K460N at sequence number 2, substitution of I199F at sequence number 2, substitution of G492P at sequence number, substitution of T153I at sequence number 2, substitution of R591I at sequence number 2, insertion of AS at position 795 of sequence number 2, insertion of AS at position 796 of sequence number 2, insertion of L at position 889 of sequence number 2, substitution of E121D at sequence number 2, substitution of S270W at sequence number 2, substitution of E712Q at sequence number 2, substitution of K942Q at sequence number 2, substitution of E552K at sequence number 2, substitution of K25Q at sequence number 2, substitution of N47D at sequence number 2, insertion of T at position 696 of sequence number 2, substitution of L685I at sequence number 2, substitution of N880D at sequence number 2, substitution of Q102R at sequence number 2, substitution of M734K at sequence number 2, substitution of A724S at sequence number 2, substitution of T704K at sequence number 2, substitution of P224K at sequence number 2, substitution of K25R at sequence number 2, substitution of M29E at sequence number 2, substitution of H152D at sequence number 2, substitution of S219R at sequence number 2, substitution of E475K at sequence number 2, substitution of G226R at sequence number 2, substitution of A377K at sequence number 2, substitution of E480K at sequence number 2, substitution of K416E at sequence number 2, substitution of H164R at sequence number 2, substitution of K767R at sequence number 2, substitution of I7F at sequence number 2, substitution of M29R at sequence number 2, substitution of H435R at sequence number 2, substitution of E385Q at sequence number 2, substitution of E385K at sequence number 2, substitution of I279F at sequence number 2, substitution of D489S at sequence number 2, substitution of D732N at sequence number 2, substitution of A739T at sequence number 2, substitution of W885R at sequence number 2, substitution of E53K at sequence number 2, substitution of A238T at sequence number 2, substitution of P283Q at sequence number 2, substitution of E292K at sequence number 2, substitution of Q628E at sequence number 2, substitution of R388Q at sequence number 2, substitution of G791M at sequence number 2, substitution of L792K at sequence number 2, substitution of L792E at sequence number 2, substitution of M779N at sequence number 2, substitution of G27D at sequence number 2, substitution of K955R at sequence number 2, substitution of S867R at sequence number 2, substitution of R693I at sequence number 2, substitution of F189Y at sequence number 2, substitution of V635M at sequence number 2, substitution of F399L at sequence number 2Substitution of E498K at sequence number 2, substitution of E386R at sequence number 2, substitution of V254G at sequence number 2, substitution of P793S at sequence number 2, substitution of K188E at sequence number 2, substitution of QT945KI at sequence number 2, substitution of T620P at sequence number 2, substitution of T946P at sequence number 2, substitution of TT949PP at sequence number 2, substitution of N952T at sequence number 2, substitution of K682E at sequence number 2, substitution of K975R at sequence number 2, substitution of L212P at sequence number 2, substitution of E292R at sequence number 2, substitution of I303K at sequence number 2, substitution of C349E at sequence number 2, substitution of E385P at sequence number 2, substitution of E386N at sequence number 2, substitution of D387K at sequence number 2, substitution of L404K at sequence number 2, substitution of E466H at sequence number 2, substitution of C477Q at sequence number 2, substitution of C477H at sequence number 2, substitution of C479A at sequence number 2, substitution of D659H at sequence number 2, substitution of T806V at sequence number 2, substitution of K808S at sequence number 2, insertion of AS at position 797 of sequence number 2, substitution of V959M at sequence number 2, substitution of K975Q at sequence number 2, substitution of W974G at sequence number 2, substitution of A708Q at sequence number 2, substitution of V711K at sequence number 2, substitution of D733T at sequence number 2, substitution of L742W at sequence number 2, substitution of V747K at sequence number 2, substitution of F755M at sequence number 2, substitution of M771A at sequence number 2, substitution of M771Q at sequence number 2, substitution of W782Q at sequence number 2, substitution of G791F at sequence number 2, substitution of L792D at sequence number 2, substitution of L792K at sequence number 2, substitution of P793Q at sequence number 2, substitution of P793G at sequence number 2, substitution of Q804A at sequence number 2, substitution of Y966N at sequence number 2, substitution of Y723N at sequence number 2, substitution of Y857R at sequence number 2, substitution of S890R at sequence number 2, substitution of S932M at sequence number 2, substitution of L897M at sequence number 2, substitution of R624G at sequence number 2, substitution of S603G at sequence number 2, substitution of N737S at sequence number 2, substitution of L307K at sequence number 2, substitution of I658V at sequence number 2, insertion of PT at position 688 of sequence number 2, insertion of SA at position 794 of sequence number 2, substitution of S877R at sequence number 2, substitution of N580T at sequence number 2, substitution of V335G at sequence number 2, substitution of T620S at sequence number 2, substitution of W345G at sequence number 2, substitution of T280S at sequence number 2Selected from the group consisting of substitution of L406P at sequence number 2, substitution of A612D at sequence number 2, substitution of A751S at sequence number 2, substitution of E386R at sequence number 2, substitution of V351M at sequence number 2, substitution of K210N at sequence number 2, substitution of D40A at sequence number 2, substitution of E773G at sequence number 2, substitution of H207L at sequence number 2, substitution of T62A at sequence number 2, substitution of T287P at sequence number 2, substitution of T832A at sequence number 2, substitution of A893S at sequence number 2, insertion of V at position 14 of sequence number 2, insertion of AG at position 13 of sequence number 2, substitution of R11V at sequence number 2, substitution of R12N at sequence number 2, substitution of R13H at sequence number 2, insertion of Y at position 13 of sequence number 2, substitution of R12L at sequence number 2, insertion of Q at position 13 of sequence number 2, substitution of V15S at sequence number 2, and insertion of D at position 17 of sequence number 2. In some embodiments, at least two amino acid changes relative to the reference CasX protein are selected from the amino acid changes disclosed in the sequences of Table 4. In some embodiments, the CasX variant comprises any combination of the foregoing embodiments of this paragraph.,
[0197] In some embodiments, the CasX variant protein comprises two or more substitutions, insertions, and / or deletions of the reference CasX protein amino acid sequence. In some embodiments, the reference CasX protein comprises or consists essentially of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions S794R and Y797L of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions K416E and A708K of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitution A708K and the deletion of P793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the deletion of P793 and the insertion of AS at position 795 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions Q367K and I425S of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitution A708K, the deletion of P at position 793, and the substitution A793V of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions Q338R and A339E of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions Q338R and A339K of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions S507G and G508R of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitution L379R, the substitution A708K, and the deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitution C477K, the substitution A708K, and the deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitution L379R, the substitution C477K, the substitution A708K, and the deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitution L379R, the substitution A708K, the deletion of P at position 793, and the substitution A739V of SEQ ID NO: 2.In some embodiments, the CasX variant protein comprises the substitutions C477K, A708K, deletion of P at position 793, and substitution A739V of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions L379R, C477K, A708K, deletion of P at position 793, and substitution A739V of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions L379R, A708K, deletion of P at position 793, and substitution M779N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions L379R, A708K, deletion of P at position 793, and substitution M771N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions L379R, 708K, deletion of P at position 793, and substitution D489S of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions L379R, A708K, deletion of P at position 793, and substitution A739T of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions L379R, A708K, deletion of P at position 793, and substitution D732N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions L379R, A708K, deletion of P at position 793, and substitution G791M of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions L379R, 708K, deletion of P at position 793, and substitution Y797L of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions L379R, C477K, A708K, deletion of P at position 793, and substitution M779N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions L379R, C477K, A708K, deletion of P at position 793, and substitution M771N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions L379R, C477K, A708K, deletion of P at position 793, and substitution D489S of SEQ ID NO: 2.In some embodiments, the CasX variant protein comprises the substitutions L379R, C477K, A708K, deletion of P at position 793, and substitution A739T of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions L379R, C477K, A708K, deletion of P at position 793, and substitution D732N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions L379R, C477K, A708K, deletion of P at position 793, and substitution G791M of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions L379R, C477K, A708K, deletion of P at position 793, and substitution Y797L of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions L379R, C477K, A708K, deletion of P at position 793, and substitution T620P of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions A708K, deletion of P at position 793, and substitution E386S of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions E386R, F399L, and deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions R581I and A739V of SEQ ID NO: 2. In some embodiments, the CasX variant comprises any combination of the foregoing embodiments of this paragraph.
[0198] In some embodiments, the CasX variant protein comprises two or more substitutions, insertions, and / or deletions of the reference CasX protein amino acid sequence. In some embodiments, the CasX variant protein comprises the substitution of A708K of SEQ ID NO: 2, the deletion of P at position 793, and the substitution of A739V. In some embodiments, the CasX variant protein comprises the substitution of L379R of SEQ ID NO: 2, the substitution of A708K, and the deletion of P at position 793. In some embodiments, the CasX variant protein comprises the substitution of C477K of SEQ ID NO: 2, the substitution of A708K, and the deletion of P at position 793. In some embodiments, the CasX variant protein comprises the substitution of L379R of SEQ ID NO: 2, the substitution of C477K, the substitution of A708K, and the deletion of P at position 793. In some embodiments, the CasX variant protein comprises the substitution of L379R of SEQ ID NO: 2, the substitution of A708K, the deletion of P at position 793, and the substitution of A739V. In some embodiments, the CasX variant protein comprises the substitution of C477K of SEQ ID NO: 2, the substitution of A708K, the deletion of P at position 793, and the substitution of A739. In some embodiments, the CasX variant protein comprises the substitution of L379R of SEQ ID NO: 2, the substitution of C477K, the substitution of A708K, the deletion of P at position 793, and the substitution of A739V. In some embodiments, the CasX variant protein comprises the substitution of L379R of SEQ ID NO: 2, the substitution of C477K, the substitution of A708K, the deletion of P at position 793, and the substitution of T620P. In some embodiments, the CasX variant protein comprises the substitution of M771A of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitution of L379R of SEQ ID NO: 2, the substitution of A708K, the deletion of P at position 793, and the substitution of D732N. In some embodiments, the CasX variant comprises any combination of the foregoing embodiments of this paragraph.
[0199] In some embodiments, the CasX variant protein comprises the substitution of W782Q of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitution of M771Q of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions of R458I and A739V of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions of L379R, A708K, deletion of P at position 793, and M771N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions of L379R, A708K, deletion of P at position 793, and A739T of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions of L379R, C477K, A708K, deletion of P at position 793, and D489S of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions of L379R, C477K, A708K, deletion of P at position 793, and D732N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitution of V711K of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions of L379R, C477K, A708K, deletion of P at position 793, and Y797L of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions of L379R, A708K, and deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions of L379R, C477K, A708K, deletion of P at position 793, and M771N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions of A708K, substitution of P at position 793, and E386S of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitutions of L379R, C477K, A708K, and deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitution of L792D of SEQ ID NO: 2.In some embodiments, the CasX variant protein comprises the substitution of G791F of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitution of A708K of SEQ ID NO: 2, the deletion of P at position 793, and the substitution of A739V. In some embodiments, the CasX variant protein comprises the substitution of L379R of SEQ ID NO: 2, the substitution of A708K, the deletion of P at position 793, and the substitution of A739V. In some embodiments, the CasX variant protein comprises the substitution of C477K of SEQ ID NO: 2, the substitution of A708K, and the substitution of P at position 793. In some embodiments, the CasX variant protein comprises the substitution of L249I of SEQ ID NO: 2 and the substitution of M771N. In some embodiments, the CasX variant protein comprises the substitution of V747K of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises the substitution of L379R of SEQ ID NO: 2, the substitution of C477, the substitution of A708K, the deletion of P at position 793, and the substitution of M779N. In some embodiments, the CasX variant protein comprises the substitution of F755M. In some embodiments, the CasX variant comprises any combination of the foregoing embodiments of this paragraph.
[0200] In some embodiments, the CasX variant protein comprises at least one modification as compared to the reference CasX sequence of SEQ ID NO: 2, and the at least one modification is selected from one or more of the following: an amino acid substitution of L379R, an amino acid substitution of A708K, an amino acid substitution of T620P, an amino acid substitution of E385P, an amino acid substitution of Y857R, an amino acid substitution of I658V, an amino acid substitution of F399L, an amino acid substitution of Q252K, and an amino acid deletion of [P793]. In some embodiments, the CasX variant protein comprises at least one modification as compared to the reference CasX sequence of SEQ ID NO: 2, and the at least one modification is selected from one or more of the following: an amino acid substitution of L379R, an amino acid substitution of A708K, an amino acid substitution of T620P, an amino acid substitution of E385P, an amino acid substitution of Y857R, an amino acid substitution of I658V, an amino acid substitution of F399L, an amino acid substitution of Q252K, an amino acid substitution of L404K, and an amino acid deletion of [P793]. In other embodiments, the CasX variant protein comprises any combination of the aforementioned substitutions or deletions as compared to the reference CasX sequence of SEQ ID NO: 2. In other embodiments, in addition to the aforementioned substitutions or deletions, the CasX variant protein can further comprise substitutions of the NTSB and / or the helical 1b domain derived from the reference CasX of SEQ ID NO: 1.
[0201] In some embodiments, the CasX variant protein comprises from 400 to 2000 amino acids, from 500 to 1500 amino acids, from 700 to 1200 amino acids, from 800 to 1100 amino acids, or from 900 to 1000 amino acids.
[0202] In some embodiments, the CasX variant protein comprises one or more modifications to regions of non - adjacent residues that form a channel where gNA:target DNA complex formation occurs. In some embodiments, the CasX variant protein comprises one or more modifications to regions of non - adjacent residues that form an interface for binding to the gNA. For example, in some embodiments of the reference CasX protein, helices I, II, and the OBD domain all contact or are in proximity to the gNA:target DNA complex, and one or more modifications to non - adjacent residues within any of these domains can improve the function of the CasX variant protein.
[0203] In some embodiments, the CasX variant protein comprises one or more modifications in a region of non-adjacent residues that form a channel for binding non-target strand DNA. For example, the CasX variant protein can comprise one or more modifications to non-adjacent residues of the NTSBD. In some embodiments, the CasX variant protein comprises one or more modifications in a region of non-adjacent residues that form an interface for binding to the PAM. For example, the CasX variant protein can comprise one or more modifications to non-adjacent residues of the helical I domain or the OBD. In some embodiments, the CasX variant protein comprises one or more modifications that include a region of non-adjacent surface-exposed residues. As used herein, "surface-exposed residue" refers to an amino acid on the surface of the CasX protein, or an amino acid at least a portion of which, e.g., a backbone or side chain portion, is on the surface of the protein. Surface-exposed residues of cellular proteins such as CasX that are exposed to the aqueous intracellular environment are often selected from positively charged hydrophilic amino acids such as arginine, asparagine, aspartic acid, glutamine, glutamic acid, histidine, lysine, serine, and threonine. Thus, for example, in some embodiments of the variants provided herein, the region of surface-exposed residues comprises one or more insertions, deletions, or substitutions as compared to the reference CasX protein. In some embodiments, one or more positively charged residues are substituted with one or more other positively charged residues, or negatively charged residues, or uncharged residues, or any combination thereof. In some embodiments, one or more amino acid residues for substitution are nearby bound nucleic acids, e.g., residues within the RuvC domain or helical I domain that contact the target DNA, or residues within the OBD or helical II domain that bind to the gNA, can be substituted with one or more positively charged amino acids or polar amino acids.
[0204] In some embodiments, the CasX variant protein comprises one or more modifications in a region of non-adjacent residues that form a core by hydrophobic packing in a domain of the reference CasX protein. Without wishing to be bound by any theory, the region that forms a core by hydrophobic packing is rich in hydrophobic amino acids such as valine, isoleucine, leucine, methionine, phenylalanine, tryptophan, and cysteine. For example, in some reference CasX proteins, the RuvC domain contains a hydrophobic pocket adjacent to the active site. In some embodiments, 2 to 15 residues of that region are charged, polar, or base-stacking. Charged amino acids (sometimes referred to herein as residues) can include, for example, arginine, lysine, aspartic acid, and glutamic acid, and the side chains of these amino acids can form salt bridges, provided that a cross-linking partner is also present. Polar amino acids can include, for example, glutamine, asparagine, histidine, serine, threonine, tyrosine, and cysteine. Polar amino acids can form hydrogen bonds as proton donors or acceptors depending on the identity of their side chains in some embodiments. As used herein, "base stacking" includes the interaction of aromatic side chains of amino acid residues (such as tryptophan, tyrosine, phenylalanine, or histidine) with stacked nucleotide bases in a nucleic acid. Any modification to a region of non-adjacent amino acids that are spatially very close to form a functional portion of the CasX variant protein is contemplated to be within the scope of the present disclosure.
[0205] i. A CasX variant protein having domains from multiple source proteins In certain embodiments, the present disclosure provides chimeric CasX proteins that include two or more different CasX proteins, such as two or more naturally occurring CasX proteins, or protein domains derived from two or more of the CasX variant protein sequences described herein. As used herein, a "chimeric CasX protein" refers to a CasX that includes at least two domains isolated from or derived from different sources, such as two naturally occurring proteins, and in some embodiments, can be isolated from different species. For example, in some embodiments, a chimeric CasX protein includes a first domain from a first CasX protein and a second domain from a second different CasX protein. In some embodiments, the first domain can be selected from the group consisting of the NTSB, TSL, Helical I, Helical II, OBD, and RuvC domains. In some embodiments, the second domain is selected from the group consisting of the NTSB, TSL, Helical I, Helical II, OBD, and RuvC domains, and the second domain is different from the aforementioned first domain. For example, a chimeric CasX protein can include the NTSB, TSL, Helical I, Helical II, OBD domains from the CasX protein of SEQ ID NO: 2 and the RuvC domain from the CasX protein of SEQ ID NO: 1, or vice versa. As a further example, a chimeric CasX protein can include the NTSB, TSL, Helical II, OBD, and RuvC domains from the CasX protein of SEQ ID NO: 2 and the Helical I domain from the CasX protein of SEQ ID NO: 1, or vice versa. Thus, in certain embodiments, a chimeric CasX protein can include the NTSB, TSL, Helical II, OBD, and RuvC domains from a first CasX protein and the Helical I domain from a second CasX protein. In some embodiments of the chimeric CasX protein, the domain of the first CasX protein is derived from the sequence of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3, the domain of the second CasX protein is derived from the sequence of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3, and the first CasX protein and the second CasX protein are not the same.In some embodiments, the domain of the first CasX protein comprises a sequence derived from SEQ ID NO: 1, and the domain of the second CasX protein comprises a sequence derived from SEQ ID NO: 2. In some embodiments, the domain of the first CasX protein comprises a sequence derived from SEQ ID NO: 1, and the domain of the second CasX protein comprises a sequence derived from SEQ ID NO: 3. In some embodiments, the domain of the first CasX protein comprises a sequence derived from SEQ ID NO: 2, and the domain of the second CasX protein comprises a sequence derived from SEQ ID NO: 3. In some embodiments, the CasX variant is selected from the group consisting of CasX variants 387, 388, 389, 390, 395, 485, 486, 487, 488, 489, 490, and 491, and their sequences are described in Table 4.
[0206] In some embodiments, the CasX variant protein comprises at least one chimeric domain comprising a first portion from a first CasX protein and a second portion from a second, different CasX protein. As used herein, a "chimeric domain" refers to a domain that comprises at least two portions isolated from or derived from different sources, such as two naturally occurring proteins, or a portion of a domain from two reference CasX proteins. The at least one chimeric domain can be any of the NTSB, TSL, Helical I, Helical II, OBD, or RuvC domains described herein. In some embodiments, the first portion of the CasX domain comprises the sequence of SEQ ID NO: 1 and the second portion of the CasX domain comprises the sequence of SEQ ID NO: 2. In some embodiments, the first portion of the CasX domain comprises the sequence of SEQ ID NO: 1 and the second portion of the CasX domain comprises the sequence of SEQ ID NO: 3. In some embodiments, the first portion of the CasX domain comprises the sequence of SEQ ID NO: 2 and the second portion of the CasX domain comprises the sequence of SEQ ID NO: 3. In some embodiments, the at least one chimeric domain comprises a chimeric RuvC domain. As an example of the foregoing, the chimeric RuvC domain comprises amino acids 661-824 of SEQ ID NO: 1 and amino acids 922-978 of SEQ ID NO: 2. As an alternative example of the foregoing, the chimeric RuvC domain comprises amino acids 648-812 of SEQ ID NO: 2 and amino acids 935-986 of SEQ ID NO: 1. In some embodiments, the CasX protein comprises a first domain from a first CasX protein, a second domain from a second CasX protein, and at least one chimeric domain comprising at least two portions isolated from different CasX proteins using the approach of the embodiments described in this paragraph. In the foregoing embodiments, the chimeric CasX protein having a domain or portion of a domain derived from SEQ ID NOs: 1, 2, and 3 can further comprise any amino acid insertion, deletion, or substitution of any of the embodiments disclosed herein.
[0207] In some embodiments, the CasX variant protein comprises the sequences set forth in Table 4, Table 7, Table 8, Table 9, or Table 11. In some embodiments, the CasX variant protein consists of the sequences set forth in Table 4. In other embodiments, the CasX variant protein comprises a sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 86%, at least 87%, at least 88%, at least 89%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the sequences set forth in Table 4, Table 7, Table 8, Table 9, or Table 11. In other embodiments, the CasX variant protein comprises the sequence set forth in Table 4 and further comprises one or more of the NLSs disclosed herein either at or near the N-terminus, C-terminus, or both. In some cases, it will be understood that the N-terminal methionine of these table CasX variants is removed from the CasX variants expressed during post-translational modification.
Table 3-1
Table 3-2
Table 3-3
Table 3-4
Table 3-5
Table 3-6
Table 3-7
Table 3-8
Table 3-9
Table 3-10
Table 3-11
Table 3-12
Table 3-13
Table 3-14
Table 3-15
Table 3-16
Table 3-17
Table 3-18
Table 3-19
Table 3-20
Table 3-21
Table 3-22
Table 3-23
Table 3-24
Table 3-25
Table 3-26
Table 3-27
Table 3-28
Table 3-29
Table 3-30
Table 3-31
Table 3-32
[0208] In some embodiments, the CasX variant protein comprises a sequence selected from the group consisting of SEQ ID NOs: 49-143, 438, 440, 442, 444, 446, 448-460, 472, 474, 478, 480, 482, 484, 486, 488, 490, 612, and 613. In some embodiments, the CasX variant protein comprises a sequence selected from the group consisting of SEQ ID NOs: 49-143, 438, 440, 442, 444, 446, 448-460, 472, 474, 478, 480, 482, 484, 486, 488, 490, 612, and 613, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 95%, or at least about 95%, or at least about 96%, or at least about 97%, or at least about 98%, or at least about 99% sequence identity thereto. In some embodiments, the CasX variant protein comprises a sequence selected from the group consisting of SEQ ID NOs: 49-143, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 95%, or at least about 95%, or at least about 96%, or at least about 97%, or at least about 98%, or at least about 99% sequence identity thereto. In some embodiments, the CasX variant protein comprises a sequence selected from the group consisting of SEQ ID NOs: 49-143.
[0209] In some embodiments, the CasX variant protein has one or more improved properties of the CasX protein as compared to a reference CasX protein, such as the reference proteins of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In some embodiments, at least one improved property of the CasX variant is improved by at least about 1.1-fold to about 100,000-fold as compared to the reference protein. In some embodiments, at least one improved property of the CasX variant is improved by at least about 1.1-fold to about 10,000-fold, at least about 1.1-fold to about 1,000-fold, at least about 1.1-fold to about 500-fold, at least about 1.1-fold to about 400-fold, at least about 1.1-fold to about 300-fold, at least about 1.1-fold to about 200-fold, at least about 1.1-fold to about 100-fold, at least about 1.1-fold to about 50-fold, at least about 1.1-fold to about 40-fold, at least about 1.1-fold to about 30-fold, at least about 1.1-fold to about 20-fold, at least about 1.1-fold to about 10-fold, at least about 1.1-fold to about 9-fold, at least about 1.1-fold to about 8-fold, at least about 1.1-fold to about 7-fold, at least about 1.1-fold to about 6-fold, at least about 1.1-fold to about 5-fold, at least about 1.1-fold to about 4-fold, at least about 1.1-fold to about 3-fold, at least about 1.1-fold to about 2-fold, at least about 1.1-fold to about 1.5-fold, at least about 1.5-fold to about 3-fold, at least about 1.5-fold to about 4-fold, at least about 1.5-fold to about 5-fold, at least about 1.5-fold to about 10-fold, at least about 5-fold to about 10-fold, at least about 10-fold to about 20-fold, at least 10-fold to about 30-fold, at least 10-fold to about 50-fold, or at least 10-fold to about 100-fold as compared to the reference CasX protein. In some embodiments, at least one improved property of the CasX variant is improved by at least about 10-fold to about 1000-fold as compared to the reference CasX protein.
[0210] In some embodiments, one or more improved properties of the CasX variant protein are improved by at least about 5, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 250, at least about 500, or at least about 1000, at least about 5,000, at least about 10,000, or at least about 100,000-fold compared to the reference CasX protein. In some embodiments, the improved properties of the CasX variant protein are improved by at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 2.1, at least about 2.2, at least about 2.3, at least about 2.4, at least about 2.5, at least about 2.6, at least about 2.7, at least about 2.8, at least about 2.9, at least about 3, at least about 3.5, at least about 4, at least about 4.5, at least about 5, at least about 5.5, at least about 6, at least about 6.5, at least about 7.0, at least about 7.5, at least about 8, at least about 8.5, at least about 9, at least about 9.5, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 500, at least about 1,000, at least about 10,000, or at least about 100,000-fold compared to the reference CasX protein.In other cases, one or more improved properties of the CasX variant are improved by about 1.1 to 100,000-fold, about 1.1 to 10,000-fold, about 1.1 to 1,000-fold, about 1.1 to 500-fold, about 1.1 to 100-fold, about 1.1 to 50-fold, about 1.1 to 20-fold, about 10 to 100,000-fold, about 10 to 10,000-fold, about 10 to 1,000-fold, about 10 to 500-fold, about 10 to 100-fold, about 10 to 50-fold, about 10 to 20-fold, about 2 to 70-fold, about 2 to 50-fold, about 2 to 30-fold, about 2 to 20-fold, about 2 to 10-fold, about 5 to 50-fold, about 5 to 30-fold, about 5 to 10-fold, about 100 to 100,000-fold, about 100 to 10,000-fold, about 100 to 1,000-fold, about 100 to 500-fold, about 500 to 100,000-fold, about 500 to 10,000-fold, about 500 to 1,000-fold, about 500 to 750-fold, about 1,000 to 100,000-fold, about 10,000 to 100,000-fold, about 20 to 500-fold, about 20 to 250-fold, about 20 to 200-fold, about 20 to 100-fold, about 20 to 50-fold, about 50 to 10,000-fold, about 50 to 1,000-fold, about 50 to 500-fold, about 50 to 200-fold, or about 50 to 100-fold compared to the reference CasX of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In other cases, one or more improved properties of the CasX variant are improved by about 1.1-fold, 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, 15-fold, 16-fold, 17-fold, 18-fold, 19-fold, 20-fold, 25-fold, 30-fold, 40-fold, 45-fold, 50-fold, 55-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, 110-fold, 120-fold, 130-fold, 140-fold, 150-fold, 160-fold, 170-fold, 180-fold, 190-fold, 200-fold, 210-fold, 220-fold, 230-fold, 240-fold, 250-fold, 260-fold, 270-fold, 280-fold, 290-fold, 300-fold, 310-fold, 320-fold, 330-fold, 340-fold, 350-fold, 360-fold, 370-fold, 380-fold, 390-fold, 400-fold, 425-fold, 450-fold, 475-fold, or 500-fold compared to the reference CasX of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3.Exemplary properties that can be improved in a CasX variant protein as compared to the same properties in the CasX protein include improved folding of the variant, improved binding affinity for gNA, improved binding affinity for target DNA, improved ability to utilize a broader spectrum of PAM sequences in the editing and / or binding of target DNA, improved unwinding of target DNA, increased editing activity, improved editing efficiency, improved editing specificity, increased nuclease activity, increased target strand loading for double-strand breaks, decreased target strand loading for single-strand nicking, decreased off-target cleavage, improved binding of the non-target strand of DNA, improved protein stability, improved CasX:gNA RNA complex stability, improved protein solubility, improved CasX:gNA RNP complex solubility, improved ability to form a gNA and cleavage-capable RNP, improved protein yield, improved protein expression, and improved fusion properties, but are not limited thereto. In some embodiments, the variant comprises at least one improved property. In other embodiments, the variant comprises at least two improved properties. In further embodiments, the variant comprises at least three improved properties. In some embodiments, the variant comprises at least four improved properties. In still further embodiments, the variant comprises at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, or more improved properties.
[0211] Exemplary improved properties include, by way of example, improved editing efficiency. In some embodiments, the RNPs of the present disclosure at a concentration of 20 pM or less, including the CasX protein and gNA, can cleave double-stranded DNA targets with an efficiency of at least 80%. In some embodiments, RNPs at a concentration of 20 pM or less can cleave double-stranded DNA targets with an efficiency of at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, or at least 95%. In some embodiments, RNPs at a concentration of 50 pM or less, 40 pM or less, 30 pM or less, 20 pM or less, 10 pM or less, or 5 pM or less can cleave double-stranded DNA targets with an efficiency of at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, or at least 95%.
[0212] These improved properties are described in more detail below.
[0213] j. Protein stability In some embodiments, the present disclosure provides CasX variant proteins having improved stability compared to a reference CasX protein. In some embodiments, the improved stability of the CasX variant protein results in higher steady-state expression of the protein, thereby improving editing efficiency. In some embodiments, the improved stability of the CasX variant protein results in a CasX protein that remains folded in a more majority functional conformation, improving editing efficiency or improving purification for manufacturing purposes. As used herein, "functional conformation" refers to a CasX protein in a conformation in which the protein can bind to gNA and target DNA. In embodiments where the CasX variant does not have one or more mutations that render the CasX variant catalytically dead, the CasX variant can cleave, nick, or otherwise modify the target DNA. For example, a functional CasX variant can, in some embodiments, be used for gene editing, and a functional conformation refers to a "editing-capable" conformation. In some exemplary embodiments, including embodiments where the CasX variant protein results in a CasX protein that remains folded in a more majority functional conformation, lower concentrations of the CasX variant are required for applications such as gene editing compared to the reference CasX protein. Thus, in some embodiments, CasX variants having improved stability have improved efficiency compared to reference CasX under one or more gene editing contexts.
[0214] In some embodiments, the present disclosure provides CasX variant proteins having improved thermal stability compared to reference CasX proteins. In some embodiments, the CasX variant proteins have improved thermal stability of the CasX variant proteins within a particular temperature range. Without wishing to be bound by any theory, some reference CasX proteins naturally function in organisms with an ecological niche in groundwater and sediments, and thus, some reference CasX proteins may have evolved to exhibit optimal function at lower or higher temperatures than may be desirable for certain applications. For example, one use of the CasX variant proteins is gene editing in mammalian cells, which is typically performed at about 37°C. In some embodiments, the CasX variant proteins described herein have improved thermal stability compared to reference CasX proteins at at least 16°C, at least 18°C, at least 20°C, at least 22°C, at least 24°C, at least 26°C, at least 28°C, at least 30°C, at least 32°C, at least 34°C, at least 35°C, at least 36°C, at least 37°C, at least 38°C, at least 39°C, at least 40°C, at least 41°C, at least 42°C, at least 44°C, at least 46°C, at least 48°C, at least 50°C, at least 52°C, or higher temperatures. In some embodiments, the CasX variant proteins have improved thermal stability and functionality compared to reference CasX proteins, thereby providing improved gene editing functionality, such as mammalian gene editing applications that may include human gene editing applications.
[0215] In some embodiments, the present disclosure provides a CasX variant protein having improved stability of a CasX variant protein:gNA complex compared to a reference CasX protein:gNA complex such that the RNP remains in a functional form. Improvements in stability can include increased thermal stability, resistance to proteolysis, enhanced pharmacokinetic properties, stability over various pH and salt conditions, and isotonicity. The improved stability of the complex can, in some embodiments, result in improved editing efficiency. In some embodiments, the RNP of the CasX variant and the gNA variant has at least 5%, at least 10%, at least 15%, or at least 20%, or at least 5-20% higher percentage of RNP with cleavage ability compared to the RNP of the reference CasX of SEQ ID NOs: 1-3 and any one of the gNAs of SEQ ID NOs: 4-16 in Table 1.
[0216] In some embodiments, the present disclosure provides a CasX variant protein having improved thermal stability of the CasX variant protein:gNA complex compared to a reference CasX protein:gNA complex. In some embodiments, the CasX variant protein has improved thermal stability compared to the reference CasX protein. In some embodiments, the CasX variant protein:gNA complex has improved thermal stability compared to a complex comprising the reference CasX protein at a temperature of at least 16°C, at least 18°C, at least 20°C, at least 22°C, at least 24°C, at least 26°C, at least 28°C, at least 30°C, at least 32°C, at least 34°C, at least 35°C, at least 36°C, at least 37°C, at least 38°C, at least 39°C, at least 40°C, at least 41°C, at least 42°C, at least 44°C, at least 46°C, at least 48°C, at least 50°C, at least 52°C, or higher. In some embodiments, the CasX variant protein has improved thermal stability of the CasX variant protein:gNA complex compared to the reference CasX protein:gNA complex, thereby resulting in improved functionality in gene editing applications such as mammalian gene editing applications which may include human gene editing applications.
[0217] In some embodiments, the improved stability and / or thermal stability of the CasX variant protein comprises faster folding kinetics of the CasX variant protein compared to the reference CasX protein, slower unfolding kinetics of the CasX variant protein compared to the reference CasX protein, greater free energy release upon folding of the CasX variant protein compared to the reference CasX protein, a higher temperature (Tm) of the CasX variant protein compared to the reference CasX protein at which 50% of the CasX variant protein is unfolded, or any combination thereof. These properties can be improved over a wide range of values, for example, at least 1.1, at least 1.5, at least 10, at least 50, at least 100, at least 500, at least 1,000, at least 5,000, or at least 10,000-fold improved compared to the reference CasX protein. In some embodiments, the improved thermal stability of the CasX variant protein comprises a higher Tm of the CasX variant protein compared to the reference CasX protein. In some embodiments, the Tm of the CasX variant protein is from about 20°C to about 30°C, from about 30°C to about 40°C, from about 40°C to about 50°C, from about 50°C to about 60°C, from about 60°C to about 70°C, from about 70°C to about 80°C, from about 80°C to about 90°C, or from about 90°C to about 100°C. Thermal stability is defined as the "melting temperature" (T m) is determined by measuring. Methods for measuring protein stability characteristics such as Tm and the free energy of unfolding are known to those skilled in the art and can be measured using standard biochemical techniques in vitro. For example, Tm can be measured using differential scanning calorimetry, a thermal analysis technique in which the difference in the amount of heat required to increase the temperature of a sample and a reference is measured as a function of temperature (Chen et al (2003) Pharm Res 20:1952-60, Ghirlando et al (1999) Immunol Lett 68:47-52). Alternatively, or in addition, the Tm of the CasX variant protein can be measured using a commercially available method such as the ThermoFisher Protein Thermal Shift system. Alternatively, or in addition, circular dichroism can be used to measure the folding and unfolding kinetics, as well as the Tm (Murray et al. (2002) J. Chromatogr Sci 40:343-9). Circular dichroism (CD) depends on the unequal absorption of left- and right-handed circularly polarized light by asymmetric molecules such as proteins. Certain structures of proteins, such as alpha helices and beta sheets, have characteristic CD spectra. Thus, in some embodiments, CD can be used to determine the secondary structure of the CasX variant protein.
[0218] In some embodiments, the improved stability and / or thermal stability of the CasX variant protein includes improved folding kinetics of the CasX variant protein as compared to the reference CasX protein. In some embodiments, the folding kinetics of the CasX variant protein are improved by at least about 5, at least about 10, at least about 50, at least about 100, at least about 500, at least about 1,000, at least about 2,000, at least about 3,000, at least about 4,000, at least about 5,000, or at least about 10,000-fold as compared to the reference CasX protein. In some embodiments, the folding kinetics of the CasX variant protein are improved by at least about 1 kJ / mol, at least about 5 kJ / mol, at least about 10 kJ / mol, at least about 20 kJ / mol, at least about 30 kJ / mol, at least about 40 kJ / mol, at least about 50 kJ / mol, at least about 60 kJ / mol, at least about 70 kJ / mol, at least about 80 kJ / mol, at least about 90 kJ / mol, at least about 100 kJ / mol, at least about 150 kJ / mol, at least about 200 kJ / mol, at least about 250 kJ / mol, at least about 300 kJ / mol, at least about 350 kJ / mol, at least about 400 kJ / mol, at least about 450 kJ / mol, or at least about 500 kJ / mol as compared to the reference CasX protein.
[0219] Exemplary amino acid changes that can increase the stability of the CasX variant protein as compared to the reference CasX protein can include, but are not limited to, amino acid changes that increase the number of hydrogen bonds within the CasX variant protein, amino acid changes that increase the number of disulfide bridges within the CasX variant protein, amino acid changes that increase the number of salt bridges within the CasX variant protein, amino acid changes that strengthen the interactions between portions of the CasX variant protein, amino acid changes that increase the buried hydrophobic surface area of the CasX variant protein, or any combination thereof.
[0220] k. Protein yield In some embodiments, the disclosure provides a CasX variant protein having improved expression and yield during purification compared to a reference CasX protein. In some embodiments, the yield of CasX variant protein purified from a bacterial host cell or a eukaryotic host cell is improved compared to a reference CasX protein. In some embodiments, the bacterial host cell is an Escherichia coli cell. In some embodiments, the eukaryotic cell is yeast, a plant (e.g., tobacco), an insect (e.g., Spodoptera frugiperda sf9 cells), a mouse, a rat, a hamster, a guinea pig, a monkey, or a human cell. In some embodiments, the eukaryotic host cell is a mammalian cell including, but not limited to, human embryonic kidney 293 (HEK293) cells, baby hamster kidney (BHK) cells, NS0 cells, SP2 / 0 cells, YO myeloma cells, P3X63 mouse myeloma cells, PER cells, PER.C6 cells, hybridoma cells, NIH3T3 cells, COS, HeLa, or Chinese hamster ovary (CHO) cells.
[0221] In some embodiments, the improved yields of CasX variant proteins are achieved by codon optimization. Cells use 64 different codons, 61 of which encode 20 standard amino acids while the other 3 function as stop codons. In some cases, a single amino acid is encoded by more than one codon. Different organisms tend to use different codons for the same naturally occurring amino acid. Thus, the choice of codons in a protein, and the choice of codons that match the organism in which the protein is expressed, can in some cases significantly affect protein translation and thus protein expression levels. In some embodiments, the CasX variant proteins are encoded by codon-optimized nucleic acids. In some embodiments, the nucleic acids encoding the CasX variant proteins are codon-optimized for expression in bacterial cells, yeast cells, insect cells, plant cells, or mammalian cells. In some embodiments, the mammalian cells are mouse, rat, hamster, guinea pig, monkey, or human. In some embodiments, the CasX variant proteins are encoded by nucleic acids that are codon-optimized for expression in human cells. In some embodiments, the CasX variant proteins are encoded by nucleic acids from which nucleotide sequences that reduce translation rates in prokaryotes and eukaryotes have been removed. For example, stretches of more than three consecutive thymine residues can reduce translation rates in certain organisms, or internal polyadenylation signals can reduce translation rates.
[0222] In some embodiments, the solubility and stability improvements described herein result in improved yields of CasX variant proteins as compared to the reference CasX protein.
[0223] Improved protein yields during expression and purification can be evaluated by methods known in the art. For example, the amount of CasX variant protein can be determined by electrophoresing the protein on an SDS-page gel and comparing the CasX variant protein to a control of known amount or concentration to determine the absolute level of the protein. Alternatively, or in addition, the purified CasX variant protein can be electrophoresed on an SDS-page gel next to a reference CasX protein undergoing the same purification process to determine the relative improvement in the yield of the CasX variant protein. Alternatively, or in addition, the level of the protein can be measured using immunohistochemical methods such as Western blot or ELISA using an antibody against CasX, or by HPLC. In the case of a protein in solution, the concentration can be determined by measuring the intrinsic UV absorbance of the protein, or by methods using protein-dependent color changes such as the Lowry assay, the Smith copper / bicinchoninic acid assay, or the Bradford dye assay. Using such methods, the total protein (e.g., total soluble protein, etc.) yield obtained by expression under certain conditions can be calculated. This can be compared, for example, to the protein yield of a reference CasX protein under similar expression conditions.
[0224] l. Protein solubility In some embodiments, the CasX variant protein has improved solubility compared to the reference CasX protein. In some embodiments, the CasX variant protein has improved solubility of the CasX:gNA ribonucleoprotein complex variant compared to a ribonucleoprotein complex comprising the reference CasX protein.
[0225] In some embodiments, the improvement in protein solubility results in a higher yield of protein from protein purification techniques such as purification from E. coli. Since higher protein solubility reduces the likelihood of aggregation within the cell, in some embodiments, the improved solubility of the CasX variant protein may enable more efficient activity within the cell. Protein aggregates can be harmful or problematic to the cell in certain embodiments, and without wishing to be bound by any theory, the increased solubility of the CasX variant protein can improve this outcome of protein aggregation. Furthermore, the improved solubility of the CasX variant protein may enable enhancement of formulations that allow delivery of higher effective amounts of the functional protein, for example, in desired gene editing applications. In some embodiments, the improved solubility of the CasX variant protein compared to the reference CasX protein results in an improved yield of the CasX variant protein during purification of at least about 5, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 250, at least about 500, or at least about 1000-fold.In some embodiments, due to the improved solubility of the CasX variant protein compared to the reference CasX protein, the activity of the CasX variant protein in cells is improved by at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 2.1, at least about 2.2, at least about 2.3, at least about 2.4, at least about 2.5, at least about 2.6, at least about 2.7, at least about 2.8, at least about 2.9, at least about 3, at least about 3.5, at least about 4, at least about 4.5, at least about 5, at least about 5.5, at least about 6, at least about 6.5, at least about 7.0, at least about 7.5, at least about 8, at least about 8.5, at least about 9, at least about 9.5, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, or at least about 15-fold.
[0226] Methods for measuring the solubility of CasX proteins and its improvement in CasX variant proteins will be readily apparent to those skilled in the art. For example, the solubility of CasX variant proteins can, in some embodiments, be measured by performing concentration measurement readings on gels of the soluble fraction of lysed E. coli. Alternatively, or in addition, the improvement in the solubility of CasX variant proteins can be measured by measuring the maintenance of soluble protein products throughout the process of complete protein purification. For example, soluble protein products can be measured at one or more steps among gel affinity purification, tag cleavage, cation exchange purification, and electrophoresis of proteins on size exclusion columns. In some embodiments, concentration measurements of all bands of the protein on the gel are read after each step in the purification process. CasX variant proteins with improved solubility can, in some embodiments, maintain higher concentrations compared to the reference CasX protein in terms of mg / L of protein during protein purification, while insoluble protein variants can be lost at one or more steps due to buffer exchange, filtration steps, interaction with purification columns, etc.
[0227] In some embodiments, improving the solubility of CasX variant proteins results in a higher yield compared to the reference CasX protein in terms of mg / L of protein during protein purification.
[0228] In some embodiments, improving the solubility of CasX variant proteins allows for a greater number of editing events when evaluated in editing assays such as the EGFP disruption assay described herein, compared to less soluble proteins.
[0229] Protein affinity for m.gNA In some embodiments, the CasX variant protein has an improved affinity for gNA compared to the reference CasX protein, resulting in the formation of a ribonucleoprotein complex. The increased affinity of the CasX variant protein for gNA can result in, for example, a lower K d for the production of the RNP complex and, in some cases, can result in more stable ribonucleoprotein complex formation. In some embodiments, the increased affinity of the CasX variant protein for gNA results in an increase in the stability of the ribonucleoprotein complex when delivered to human cells. This increase in stability can not only affect the function and utility of the complex in the cells of the subject, but can also result in improved pharmacokinetic properties in the blood when delivered to the subject. In some embodiments, the increased affinity of the CasX variant protein, and the resulting increase in the stability of the ribonucleoprotein complex, allows for the delivery of a lower dose of the CasX variant protein to the subject or cells while still having the desired activity, such as in vivo or in vitro gene editing.
[0230] In some embodiments, the higher affinity (tighter binding) of the CasX variant protein for gNA allows for a greater number of editing events when both the CasX variant protein and gNA remain in the RNP complex. The increase in editing events can be evaluated using an editing assay such as the EGFP disruption assay described herein.
[0231] In some embodiments, the K dis increased by at least about 1.1-fold, at least about 1.2-fold, at least about 1.3-fold, at least about 1.4-fold, at least 1.5-fold, at least about 1.6-fold, at least about 1.7-fold, at least about 1.8-fold, at least about 1.9-fold, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40-fold, at least about 45-fold, at least about 50-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 90-fold, or at least about 100-fold as compared to the reference CasX protein. In some embodiments, the CasX variant has a binding affinity for gNA that is increased by about 1.1- to about 10-fold as compared to the reference CasX protein of SEQ ID NO: 2.
[0232] Without wishing to be bound by theory, in some embodiments, amino acid changes in the helical I domain can increase the binding affinity of the CasX variant protein for the gNA target sequence, while changes in the helical II domain can increase the binding affinity of the CasX variant protein for the gNA scaffold stem-loop, and changes in the oligonucleotide binding domain (OBD) can increase the binding affinity of the CasX variant protein for the gRNA triplex.
[0233] Methods for measuring the CasX protein binding affinity for gNA include in vitro methods using purified CasX protein and gNA. When the gNA or CasX protein is labeled with a fluorophore, the binding affinity for the reference CasX and variant proteins can be measured by fluorescence polarization. Alternatively, or in addition, the binding affinity can be measured by biolayer interferometry, electrophoretic mobility shift assay (EMSA), or filter binding. Additional standard techniques for quantifying the absolute affinity of RNA-binding proteins such as the reference CasX and variant proteins of the present disclosure for specific gNAs such as reference gNAs and their variants include, but are not limited to, isothermal titration calorimetry (ITC) and surface plasmon resonance (SPR), as well as the methods of the examples.
[0234] n. Affinity for target DNA In some embodiments, the CasX variant protein has an improved binding affinity for the target nucleic acid as compared to the binding affinity of the reference CasX protein for the target nucleic acid. In some embodiments, the improved affinity for the target nucleic acid includes an improved affinity for the target nucleic acid sequence, an improved affinity for the PAM sequence, an improved ability to find DNA for a target nucleic acid sequence, or any combination thereof. Without wishing to be bound by theory, it is believed that CRISPR / Cas system proteins such as CasX can find their target nucleic acid sequences by one-dimensional diffusion along a DNA molecule. This process is thought to include (1) binding of the ribonucleoprotein to the DNA molecule, followed by (2) stopping at the target nucleic acid sequence, both of which, in some embodiments, may be affected by the improved affinity of the CasX protein for the target nucleic acid sequence, thereby improving the function of the CasX variant protein as compared to the reference CasX protein.
[0235] In some embodiments, CasX variant proteins having improved target nucleic acid affinity have increased overall affinity for DNA. In some embodiments, CasX variant proteins having improved target nucleic acid affinity have increased affinity for specific PAM sequences other than the canonical TTC PAM recognized by the reference CasX protein of SEQ ID NO: 2, including binding affinity for PAM sequences selected from the group consisting of TTC, ATC, GTC, and CTC. Without wishing to be bound by theory, these protein variants interact more strongly with DNA overall and have an increased ability to bind to additional PAM sequences beyond the capabilities of wild-type CasX, resulting in an increased ability to access and edit sequences within the target DNA, thereby potentially enabling a more efficient search process for the CasX protein for the target sequence. A higher overall affinity for DNA may also, in some embodiments, increase the frequency with which the CasX protein can efficiently initiate and terminate the binding and rewinding steps, thereby promoting invasion of the target strand and R-loop formation and ultimately promoting cleavage of the target nucleic acid sequence.
[0236] Amino acid changes in the NTSBD that increase the efficiency of unwinding or capture of the unwound non-target DNA strand, without wishing to be bound by theory, may be able to increase the affinity of the CasX variant protein for the target DNA. Alternatively, or in addition, amino acid changes in the NTSBD that increase the ability of the NTSBD to stabilize DNA during unwinding can increase the affinity of the CasX variant protein for the target DNA. Alternatively, or in addition, amino acid changes in the OBD can increase the affinity of the CasX variant protein that binds to the protospacer adjacent motif (PAM), thereby increasing the affinity of the CasX variant protein for the target nucleic acid. Alternatively, or in addition, amino acid changes in helical I and / or II, RuvC, and the TSL domain that increase the affinity of the CasX variant protein for the target nucleic acid strand can increase the affinity of the CasX variant protein for the target nucleic acid.
[0237] In some embodiments, the CasX variant protein has an increased binding affinity for the target nucleic acid sequence as compared to the reference protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In some embodiments, the affinity of the CasX variant protein of the present disclosure for the target nucleic acid molecule is at least about 1.1, at least about 1.2, at least about 1.3, at least 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, or at least about 100-fold increased as compared to the reference CasX protein.
[0238] In some embodiments, the CasX variant protein has improved binding affinity for the non-target strand of the target nucleic acid. As used herein, the term "non-target strand" refers to the strand of a DNA target nucleic acid sequence that does not form Watson-Crick base pairs with the target sequence of the gNA and is complementary to the target strand.
[0239] Methods for measuring the affinity of a CasX protein (such as a reference or variant) for a target nucleic acid molecule can include electrophoretic mobility shift assay (EMSA), filter binding, isothermal titration calorimetry (ITC), and surface plasmon resonance (SPR), fluorescence polarization, and biolayer interferometry (BLI). Additional methods for measuring the affinity of a CasX protein for a target include in vitro biochemical assays that measure DNA cleavage events over time.
[0240] In some embodiments, a CasX variant protein having a higher affinity for a target nucleic acid can cleave the target nucleic acid sequence more rapidly than a reference CasX protein that does not have an increased affinity for the target nucleic acid.
[0241] In some embodiments, the CasX variant protein is catalytically dead (dCasX). In some embodiments, the present disclosure provides an RNP comprising a catalytically dead CasX protein that retains the ability to bind to target DNA. Exemplary catalytically dead CasX variant proteins include one or more mutations in the active site of the RuvC domain of the CasX protein. In some embodiments, the catalytically dead CasX variant protein comprises substitutions at residues 672, 769, and / or 935 of SEQ ID NO: 1. In some embodiments, the catalytically dead CasX variant protein comprises the substitutions D672A, E769A, and / or D935A in the reference CasX protein of SEQ ID NO: 1. In some embodiments, the catalytically dead CasX protein comprises substitutions at residues 659, 765, and / or 922 of SEQ ID NO: 2. In some embodiments, the catalytically dead CasX protein comprises the substitutions D659A, E756A, and / or D922A in the reference CasX protein of SEQ ID NO: 2. In further embodiments, the catalytically dead CasX variant protein comprises a deletion of all or part of the RuvC domain of the reference CasX protein.
[0242] In some embodiments, the improved affinity of the CasX variant protein for DNA also improves the function of the catalytically inactive version of the CasX variant protein. In some embodiments, the catalytically inactive version of the CasX variant protein contains one or more mutations in the DED motif in RuvC. The catalytically dead CasX variant protein can, in some embodiments, be used for base editing or epigenetic modification. The higher the affinity for DNA, in some embodiments, the catalytically dead CasX variant protein can find its target nucleic acid faster, can remain bound to the target nucleic acid for a longer period of time, can bind to the target nucleic acid in a more stable manner, or a combination thereof compared to catalytically active CasX, thereby improving the function of the catalytically dead CasX variant protein.
[0243] o. Improved specificity for the target site In some embodiments, the CasX variant protein has improved specificity for a target nucleic acid sequence compared to a reference CasX protein. As used herein, "specificity," also sometimes referred to as "target specificity," refers to the degree to which a CRISPR / Cas-based ribonucleoprotein complex cleaves an off-target sequence that is similar but not identical to the target DNA sequence. For example, a CasX variant RNP with higher specificity exhibits a reduction in off-target cleavage of sequences compared to a reference CasX protein. The specificity of CRISPR / Cas-based proteins and the reduction of potentially harmful off-target effects can be extremely important for achieving an acceptable therapeutic index for use in mammalian subjects.
[0244] In some embodiments, the CasX variant protein has improved specificity for a target site within a target sequence that is complementary to the target sequence of the gNA.
[0245] Without wishing to be bound by theory, amino acid changes in helical I and II domains that increase the specificity of the CasX variant protein for the target nucleic acid strand may be able to increase the specificity of the CasX variant protein for the entire target DNA. In some embodiments, the amino acid changes that increase the specificity of the CasX variant protein for the target DNA may also result in a decrease in the affinity of the CasX variant protein for DNA.
[0246] Methods for testing the target specificity of a CasX protein (such as a variant or reference) may include guide and Circularization for In vitro Reporting of Cleavage Effects by Sequencing (CIRCLE-seq), or similar methods. Briefly, in the CIRCLE-seq technique, genomic DNA is sheared and circularized by ligation of stem-loop adapters, and these stem-loop adapters are nicked in the stem-loop region to expose a 4-nucleotide palindromic overhang. Subsequently, intramolecular ligation and degradation of the remaining linear DNA are performed. Thereafter, circular DNA molecules containing the CasX cleavage site are linearized with CasX, ligated to the exposed ends with adapter adapters, and then high-throughput sequencing is performed to generate paired-end reads containing information on off-target sites. Additional assays that can be used to detect off-target events and thus CasX protein specificity include assays used to detect and quantify indels (insertions and deletions) formed at selected off-target sites, such as mismatch detection nuclease assays and next-generation sequencing (NGS). Exemplary mismatch detection assays include nuclease assays in which genomic DNA from cells treated with CasX and sgNA is PCR amplified, denatured, and re-hybridized to form heteroduplex DNA containing one wild-type strand and one strand with an indel. The mismatch is recognized and cleaved by a mismatch detection nuclease such as Surveyor nuclease or T7 endonuclease I.
[0247] Unwinding of p.DNA In some embodiments, the CasX variant protein has an improved ability to unwind DNA as compared to the reference CasX protein. In some embodiments, the CasX variant protein has enhanced DNA unwinding properties. Insufficient dsDNA unwinding has previously been shown to impair or prevent the ability of the CRISPR / Cas system proteins anaCas9 or Cas14s to cleave DNA. Thus, without wishing to be bound by any theory, it is likely that the increased DNA cleavage activity by some CasX variant proteins is at least partially due to an increased ability to find and unwind dsDNA at the target site.
[0248] Without wishing to be bound by theory, it is believed that amino acid changes in the NTSB domain can produce CasX variant proteins with increased DNA unwinding properties. Alternatively, or in addition, amino acid changes in the OBD or helical domain regions that interact with the PAM can also produce CasX variant proteins with increased DNA unwinding properties.
[0249] Methods for measuring the ability of a CasX protein (such as a variant or reference) to unwind DNA include, but are not limited to, in vitro assays that observe an increase in the rate of a dsDNA target in fluorescence polarization or biolayer interferometry.
[0250] q. Catalytic activity The CasX:gNA-based ribonucleoprotein complex disclosed herein includes a reference CasX protein or variant that binds to a target nucleic acid sequence and cleaves the target nucleic acid sequence. In some embodiments, the CasX variant protein has improved catalytic activity compared to the reference CasX protein. Without wishing to be bound by theory, in some cases, cleavage of the target strand is thought to be a limiting factor for Cas12-like molecules in causing dsDNA cleavage. In some embodiments, the CasX variant protein improves the bending of the DNA target strand and cleavage of this strand, resulting in an improvement in the overall efficiency of dsDNA cleavage by the CasX ribonucleoprotein complex.
[0251] In some embodiments, the CasX variant protein has increased nuclease activity compared to the reference CasX protein. Variants having increased nuclease activity can be generated, for example, by amino acid changes in the RuvC nuclease domain. In one embodiment, the CasX variant includes a nuclease domain having nickase activity. In the foregoing embodiment, the CasX nickase of the CasX:gNA system produces a single-strand break within 10-18 nucleotides 3' of the PAM site of the non-target strand. In another embodiment, the CasX variant includes a nuclease domain having double-strand break activity. In the foregoing embodiment, the CasX of the CasX:gNA system produces a double-strand break within 18-26 nucleotides 5' of the PAM site on the target strand and within 10-18 nucleotides 3' of the non-target strand. Nuclease activity can be assayed by various methods including the methods of the examples. In one embodiment, the CasX variant has a K cleavage constant that is at least 2-fold, or at least 3-fold, or at least 4-fold, or at least 5-fold, or at least 6-fold, or at least 7-fold, or at least 8-fold, or at least 9-fold, or at least 10-fold greater than that of the reference wild-type CasX.
[0252] In some embodiments, the CasX variant protein has increased target strand loading for double-strand cleavage. Variants having increased target strand loading activity can be generated, for example, by amino acid changes in the TLS domain.
[0253] Without wishing to be bound by theory, amino acid changes in the TSL domain can result in a CasX variant protein having improved catalytic activity. Alternatively, or in addition, amino acid changes around the binding channel of the RNA:DNA duplex can also improve the catalytic activity of the CasX variant protein.
[0254] In some embodiments, the CasX variant protein has increased collateral cleavage activity compared to a reference CasX protein. As used herein, "collateral cleavage activity" refers to the recognition of a target nucleic acid sequence and additional non-target cleavage of the nucleic acid after cleavage. In some embodiments, the CasX variant protein has decreased collateral cleavage activity compared to a reference CasX protein.
[0255] In some embodiments, for example, embodiments that include applications where target DNA cleavage is not a desired outcome, improving the catalytic activity of the CasX variant protein includes changing, decreasing, or abolishing the catalytic activity of the CasX variant protein. In some embodiments, a ribonucleoprotein complex comprising a CasX variant protein binds to target DNA and does not cleave the target DNA.
[0256] In some embodiments, a CasX ribonucleoprotein complex comprising a CasX variant protein binds to target DNA but generates a single-strand nick in the target DNA. In some embodiments, particularly embodiments where the CasX protein is a nickase, the CasX variant protein has decreased target strand loading for single-strand nicking. Variants having decreased target strand loading can be generated, for example, by amino acid changes in the TSL domain.
[0257] Exemplary methods for characterizing the catalytic activity of the CasX protein can include, but are not limited to, in vitro cleavage assays. In some embodiments, electrophoresis of DNA products on an agarose gel can be used to examine the kinetics of strand cleavage.
[0258] r. Affinity for target DNA and RNA In some embodiments, a ribonucleoprotein complex comprising a reference CasX protein or a variant thereof binds to and cleaves target DNA. In some embodiments, a variant of the reference CasX protein increases the specificity of the CasX variant protein for target RNA and increases the activity of the CasX variant protein for target RNA as compared to the reference CasX protein. For example, the CasX variant protein may exhibit increased binding affinity for target RNA or increased cleavage of target RNA as compared to the reference CasX protein. In some embodiments, a ribonucleoprotein complex comprising the CasX variant protein binds to and / or cleaves target RNA. In one embodiment, the CasX variant has a binding affinity for the target nucleic acid sequence that is at least about 2-fold to about 10-fold increased as compared to the reference protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3.
[0259] s. Combinations of mutations In some embodiments, the present disclosure provides variants that are combinations of mutations from distinct CasX variant proteins. In some embodiments, any variant to any of the domains described herein can be combined with any other variant described herein. In some embodiments, any variant within any of the domains described herein can be combined with any other variant described herein within the same domain. By combining different combinations of amino acid changes, in some embodiments, new optimized variants can be generated in which the function is further improved by the combination of amino acid changes. In some embodiments, the effect of combining amino acid changes on CasX protein function is linear. As used herein, a linear combination refers to a combination in which the effect on the function is equal to the sum of the effects of each of the individual amino acid changes when assayed alone. In some embodiments, the effect of combining amino acid changes on CasX protein function is synergistic. As used herein, a synergistic combination refers to a combination in which the effect on the function is greater than the sum of the effects of each of the individual amino acid changes when assayed alone. In some embodiments, by combining amino acid changes, a CasX variant protein is generated in which two or more functions of the CasX protein are improved compared to the reference CasX protein.
[0260] t.CasX fusion protein In some embodiments, the present disclosure provides a CasX protein comprising a heterologous protein fused to CasX. In some cases, CasX is the reference CasX protein. In other cases, CasX is any of the CasX variants of any of the embodiments described herein.
[0261] In some embodiments, the CasX variant protein is fused to (i.e., is part of) one or more proteins or domains thereof having different activities of interest. For example, in some embodiments, the CasX variant protein is fused to a protein (or domain thereof) that inhibits transcription, modifies a target nucleic acid sequence, or modifies a polypeptide associated with a nucleic acid (e.g., histone modification).
[0262] In some embodiments, a heterologous polypeptide (or a heterologous amino acid such as a cysteine residue or a non-natural amino acid) can be inserted at one or more positions within the CasX protein to generate a CasX fusion protein. In other embodiments, cysteine residues are inserted at one or more positions within the CasX protein, followed by conjugation of a heterologous polypeptide as described below. In some alternative embodiments, a heterologous polypeptide or heterologous amino acid can be added to the N-terminus or C-terminus of a reference or CasX variant protein. In other embodiments, a heterologous polypeptide or heterologous amino acid can be inserted within the sequence of the CasX protein.
[0263] In some embodiments, the reference CasX or variant fusion protein retains RNA guide sequence-specific target nucleic acid binding and cleavage activity. In some cases, the reference CasX or variant fusion protein has (retains) at least 50% of the activity (e.g., cleavage and / or binding activity) of the corresponding reference CasX or variant protein without the insertion of a heterologous protein. In some cases, the reference CasX or variant fusion protein retains at least about 60%, or at least about 70% or more, at least about 80%, or at least about 90%, or at least about 92%, or at least about 95%, or at least about 98%, or at least about 100% of the activity (e.g., cleavage and / or binding activity) of the corresponding CasX protein without the insertion of a heterologous protein.
[0264] In some cases, the reference CasX or variant fusion polypeptide retains (has) target nucleic acid binding activity as compared to the activity of the CasX protein in which no heterologous amino acid or heterologous polypeptide is inserted. For example, in some cases, the reference CasX or variant fusion polypeptide has (retains) at least 50% of the binding activity of the corresponding CasX protein (CasX protein without the insertion). For example, in some cases, the reference CasX or variant polypeptide has (retains) at least 60% (at least 70%, at least 80%, at least 90%, at least 92%, at least 95%, at least 98%, or 100%) of the binding activity of the corresponding parental CasX protein (CasX protein without the insertion).
[0265] In some cases, the reference CasX or variant fusion polypeptide retains (has) target nucleic acid binding and / or cleavage activity as compared to the activity of the parental CasX protein in which no heterologous amino acid or heterologous polypeptide is inserted. For example, in some cases, the reference CasX or variant fusion polypeptide has (retains) at least 50% of the binding and / or cleavage activity of the corresponding parental CasX protein (CasX protein without the insertion). For example, in some cases, the reference CasX or variant fusion polypeptide has (retains) at least 60% (at least 70%, at least 80%, at least 90%, at least 92%, at least 95%, at least 98%, or 100%) of the binding and / or cleavage activity of the corresponding CasX parental polypeptide (CasX protein without the insertion). Methods for measuring the cleavage and / or binding activity of the CasX protein and / or CasX fusion polypeptide are known to those skilled in the art, and any convenient method can be used.
[0266] A variety of heterologous polypeptides are suitable for inclusion into the reference CasX or CasX variant fusion proteins of the present disclosure. In some cases, the fusion partner can regulate the transcription of the target DNA (e.g., inhibit transcription, increase transcription). For example, in some cases, the fusion partner is a protein (or protein-derived domain) that inhibits transcription (e.g., a transcriptional repressor, recruitment of a transcription inhibitory protein, modification of the target DNA such as methylation, recruitment of a DNA modification factor, regulation of histones associated with the target DNA, recruitment of histone modification factors such as those that modify histone acetylation and / or methylation, etc.). In some cases, the fusion partner is a protein (or protein-derived domain) that increases transcription (e.g., a transcriptional activator, recruitment of a transcription activating protein, modification of the target DNA such as ...
Claims
A CasX:gRNA system comprising a chimeric CasX variant protein and a first guide nucleic acid (gRNA), wherein the first gRNA comprises a target sequence complementary to a target nucleic acid sequence of a gene encoding a first protein involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response, wherein the chimeric CasX variant protein comprises (a) a sequence having at least 90% sequence identity to SEQ ID NO: 138, (b) the sequence of amino acids 101-191 of the NSTB domain of SEQ ID NO: 1, or a sequence having at least 90% sequence identity thereto, and (c) the sequence of amino acids 192-332 of the helical I domain of SEQ ID NO: 1, or a sequence having at least 90% sequence identity thereto, wherein the chimeric CasX variant protein and the first gRNA can form a ribonucleoprotein complex (RNP), and the RNP exhibits at least one or more improved properties compared to an RNP comprising the reference CasX protein of SEQ ID NO: 2 and any one of the reference gRNAs of SEQ ID NOs: 4-16 in an equivalent assay system, wherein the at least one improved property comprises higher editing efficiency and / or binding of the RNP to the target sequence, a CasX:gRNA system. **Claim 2** The first protein is (a) an immune cell surface marker or immune checkpoint protein; (b) an intracellular protein; or (c) A CasX:gRNA system according to claim 1, selected from the group consisting of beta-2-microglobulin (B2M), T cell receptor alpha chain constant region (TRAC), class II major histocompatibility complex transactivator (CIITA), T cell receptor beta constant 1 (TRBC1), T cell receptor beta constant 2 (TRBC2), human leukocyte antigen A (HLA-A), human leukocyte antigen B (HLA-B), TGFβ receptor 2 (TGFβRII), programmed cell death 1 (PD-1), cytokine-induced SH2 (CISH), lymphocyte activation 3 (LAG-3), T cell immunoreceptor with Ig and ITIM domains (TIGIT), adenosine A2a receptor (ADORA2A), killer cell lectin-like receptor C1 (NKG2A), cytotoxic T lymphocyte-associated protein 4 (CTLA-4), T cell immunoglobulin and mucin domain 3 (TIM-3), and 2B4 (CD244). **Claim 3** The CasX:gRNA system according to claim 1 or 2, further comprising a second gRNA comprising a target sequence complementary to a target nucleic acid sequence of an immune cell gene encoding a second protein selected from the group consisting of beta-2-microglobulin (B2M), T cell receptor alpha chain constant region (TRAC), class II major histocompatibility complex transactivator (CIITA), T cell receptor beta constant 1 (TRBC1), T cell receptor beta constant 2 (TRBC2), human leukocyte antigen A (HLA-A), human leukocyte antigen B (HLA-B), TGFβRII, PD-1, CISH, LAG-3, TIGIT, ADORA2A, NKG2A, CTLA-4, TIM-3, and CD244, wherein the second protein is different from the first protein. **Claim 4** The CasX:gRNA system according to claim 3, wherein the first gRNA and / or the second gRNA is a guide RNA (gRNA). **Claim 5** The first gRNA and / or the second gRNA includes a scaffold stem-loop sequence of CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 32), and includes the sequence of SEQ ID NO: 2238 or a sequence having at least 70% sequence identity thereto, the CasX:gRNA system according to claim 3 or 4.
6. The first gRNA and / or the second gRNA is chemically modified, the CasX:gRNA system according to any one of claims 3 to 5.
7. The chimeric CasX variant protein is (a) a tag strand loading (TSL) domain; (b) a helical II domain; and (c) an oligonucleotide binding domain (OBD), and wherein the chimeric CasX variant protein includes the amino acids 59 to 102 of the helical I domain of SEQ ID NO: 2 or a sequence having at least 90% sequence identity thereto, and the amino acids 648 to 812 and 922 to 978 of the RuvC domain of SEQ ID NO: 2 or a sequence having at least 90% sequence identity thereto, the CasX:gRNA system according to any one of claims 1 to 6.
8. The chimeric CasX variant protein, compared to the reference CasX protein of SEQ ID NO: 2, is as follows: (a) an amino acid substitution of L379R; (b) an amino acid substitution of A708K; (c) an amino acid substitution of T620P; (d) an amino acid substitution of E385P; (e) an amino acid substitution of Y857R; (f) an amino acid substitution of I658V; (g) an amino acid substitution of F399L; (h) an amino acid substitution of Q252K; (i) an amino acid substitution of L404K; and (j) an amino acid deletion of P793, The CasX:gNA system according to claim 7, comprising one modification selected from the group consisting of **Claim 9** The CasX:gNA system according to any one of claims 1 to 8, wherein the chimeric CasX variant protein comprises one or more nuclear localization signals (NLSs). **Claim 10** The CasX:gNA system according to any one of claims 1 to 9, wherein the RNP has a cleavage ability percentage that is at least 20% higher compared to the RNP of the reference CasX of SEQ ID NO: 2 and the reference gNA comprising any one of the sequences of SEQ ID NOs: 4 to 16. **Claim 11** The CasX:gNA system according to any one of claims 1 to 9, wherein the chimeric CasX variant protein is a catalytically inactive CasX (dCasX) protein, and the dCasX and the gNA retain the ability to bind to a target nucleic acid. **Claim 12** The CasX:gNA system according to any one of claims 1 to 9, further comprising a donor template nucleic acid. **Claim 13** (a) a sequence encoding the chimeric CasX variant protein of the CasX:gNA system according to any one of claims 1 to 12; and (b) a gNA comprising a target sequence complementary to the target nucleic acid sequence of a gene encoding a first protein involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response. **Claim 14** A vector comprising the one or more polynucleotides according to claim 13. **Claim 15** The vector according to claim 14, wherein the vector is selected from the group consisting of a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated virus (AAV) vector, a virus-like particle (VLP), a herpes simplex virus (HSV) vector, a plasmid, a minicircle, a nanoplasmid, a nanoparticle, a DNA vector, and an RNA vector. **Claim 16** A method for modifying a target nucleic acid sequence of a gene in an in vitro or ex vivo cell population, wherein the cell population excludes human germ cells or human embryonic cells, and the gene encodes a protein involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response, wherein the method comprises introducing into the cells (a) the CasX:gRNA system according to any one of claims 1 to 12, (b) one or more polynucleotides according to claim 13, (c) the vector according to claim 14 or 15, or (d) a combination of two or more of (a) to (c), wherein the target nucleic acid sequence of the cells is modified by the chimeric CasX variant protein. **Claim 17** The method according to claim 16, wherein the CasX:gRNA system is introduced into the cells as an RNP and / or the cells are modified by introduction of a polynucleotide encoding a chimeric antigen receptor (CAR) having binding affinity for a disease antigen. **Claim 18** The method according to claim 17, wherein the disease antigen is a tumor cell antigen. **Claim 19** The method according to any one of claims 16 to 18, wherein the cell population is suitable for use in providing anti-tumor immunity in a subject. **Claim 20** The method according to any one of claims 16 to 19, wherein the cell population is suitable for use in the treatment of a subject. **Claim 21** The method according to claim 20, wherein the subject has cancer or an autoimmune disease. **Claim 22** The method according to claim 21, wherein the cancer expresses a tumor cell antigen and the CAR has specific binding affinity for the tumor cell antigen. **Claim 23** A composition comprising a CasX:gNA system for use in a method of preparing cells for immunotherapy in a subject, wherein the method comprises modifying immune cells by reducing or eliminating the expression of one or more proteins involved in antigen processing, antigen presentation, antigen recognition and / or antigen response, wherein the cells exclude human germ cells or human embryonic cells, the method comprising contacting the target nucleic acid sequence of the immune cells with the composition, the CasX:gNA system comprising a chimeric CasX variant protein and one or more gNAs, wherein the chimeric CasX variant protein, has at least 90% identity to SEQ ID NO: 138, has the sequence of amino acids 101-191 of the NSTB domain of SEQ ID NO: 1, or a sequence having at least 90% sequence identity thereto, and has the sequence of amino acids 192-332 of the helical I domain of SEQ ID NO: 1, or a sequence having at least 90% sequence identity thereto, and wherein each gNA comprises a target sequence complementary to the target nucleic acid sequence of one or more genes encoding the one or more proteins involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response.
24. The composition according to claim 23, wherein the one or more proteins are selected from the group consisting of B2M, TTRAC, CII TA, TRBC1, TRBC2, HLA-A, HLA-B, TGFβRII, PD-1, CISH, LAG-3, TIGIT, ADORA2A, NKG2A, CTLA-4, TIM-3, and CD244.
25. The method, (a) a chimeric antigen receptor (CAR) having specific binding affinity for a tumor cell antigen, and / or (b) an engineered T cell receptor (TCR) comprising a binding domain having binding affinity for a disease antigen, wherein the disease antigen is a tumor cell antigen. The composition according to claim 23 or 24, comprising introducing into the immune cell a polynucleic acid encoding **Claim 26** wherein the cell is (a) autologous to the subject receiving the cell, or (b) allogeneic to the subject receiving the cell, The composition according to any one of claims 23 to 25. **Claim 27** The composition according to any one of claims 23 to 26, wherein the subject has cancer or an autoimmune disease. **Claim 28** The composition according to claim 27, wherein the cancer expresses a tumor cell antigen and the CAR has specific binding affinity for the tumor cell antigen. **Claim 29** By administration of a therapeutically effective amount of the cell, improvement of clinical parameters or clinical evaluation items related to the disease in the subject, selected from one or more of complete response, partial response, or tumor shrinkage as incomplete response; progression-free period, treatment success period, biomarker response; progression-free survival period; disease-free survival period; recurrence-free period; metastasis-free period; overall survival period; improvement of quality of life; and improvement of symptoms, is brought about. The composition according to any one of claims 23 to 28. **Claim 30** A guide nucleic acid (gNA) comprising a target sequence complementary to a target nucleic acid sequence in the target strand of a gene encoding a protein involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response, wherein the gNA (a) comprises the scaffold stem-loop sequence of CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 32), and (b) comprises at least 70% identity to the sequence of SEQ ID NO: 2238, Here, the gNA can form a ribonucleoprotein complex (RNP) with a CasX variant protein specific for a protospacer adjacent motif (PAM) sequence containing a TC motif in the complementary non-target strand, the PAM sequence is located at one nucleotide on the 5' side of the sequence in the non-target strand complementary to the target nucleic acid sequence in the target strand, the target sequence is located at the 3' end of the gNA, and here, the CasX variant protein contains a sequence having at least 90% sequence identity with SEQ ID NO: 138, gNA.
Citation Information
Patent Citations
RNA-guided nucleic acid modifying enzymes and methods of use thereof
WO2018064371A1
Gene editing therapy for HIV infection via dual targeting of HIV genome and CCR5
WO2018152418A1