Identifying tissue-specific extragenic safe harbors for gene therapy
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- REGENERON PHARMACEUTICALS INC
- Filing Date
- 2023-04-28
- Publication Date
- 2026-05-12
AI Technical Summary
The safety and effectiveness of existing gene therapy methods in humans have not been fully verified, especially because different chromatin statuses in different tissues, the commonly used gene safe harbor may be silent in some tissues.
Methods and compositions are provided for inserting nucleic acids encoding products of interest into gene safe havens in human or animal cells, or expressing these nucleic acids in cells, populations or receptors. The method includes the use of a nuclease agent or a nucleic acid encoded therein, the target site is in a gene safe harbor, and the nucleic acid encoding a product of interest is operatively connected to the promoter.
By inserting nucleic acids encoding products of interest into the gene safe harbor, the stable and reliable expression of these products in specific cells or tissues is achieved, improving the safety and effectiveness of gene therapy.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Patent Application No. 63 / 336,663, filed April 29, 2022, which is incorporated herein by reference in its entirety for all purposes.
[0002] Reference to a sequence listing submitted as an XML file via EFS WEB The sequence listing set forth in file 591806SEQLIST.xml is 504 kilobytes, was created on April 27, 2023, and is incorporated herein by reference. [Background technology]
[0003] Current gene therapy approaches rely on episomal expression of transgenes and / or their insertion into specific genomic loci. Episomal methods have proven to be liver-restricted due to dilution or silencing. Integration at specific loci allows for sustained transgene expression. However, this approach has yet to be proven effective and safe in human settings. Canonical genomic safe harbor loci in humans, such as AAVS1, CCR5, and Rosa26, are all intragenic and are less well studied than mouse genomic safe harbor loci. Additionally, because different tissues have different chromatin states for a given locus, canonical genomic safe harbor loci may be silenced in some tissues. Therefore, tissue-specific genomic safe harbor loci are needed. Summary of the Invention [Means for solving the problem]
[0004] Compositions and methods are provided for inserting a nucleic acid encoding a product of interest into a genomic safe harbor locus in a cell, a population of cells, or a subject, or for expressing a nucleic acid encoding a product of interest from a genomic safe harbor locus in a cell, a population of cells, or a subject. Also provided are cells or populations of cells that contain a nucleic acid construct that includes a coding sequence for a product of interest inserted into a genomic safe harbor locus. Also provided are methods for identifying genomic safe harbor loci for use in specific cell or tissue types.
[0005] In one aspect, provided are methods of incorporating a nucleic acid construct into a genomic safe harbor locus in a cell, such as a human cell (e.g., a mammalian cell), methods of expressing a product of interest from a genomic safe harbor locus in a cell, such as a human cell (e.g., a mammalian cell), methods of incorporating a nucleic acid construct into a genomic safe harbor locus in a cell (e.g., a mammalian cell) of a subject, such as a human cell of a human subject (e.g., a mammalian subject), and methods of expressing a product of interest from a genomic safe harbor locus in a cell (e.g., a mammalian cell) of a subject, such as a human cell of a human subject (e.g., a mammalian subject).
[0006] Methods are provided for integrating a nucleic acid construct into a genomic safe harbor locus in a cell (eg, a mammalian cell), such as a human cell. Such methods can include administering to a cell (e.g., a human cell) (a) a nuclease agent or one or more nucleic acids encoding a nuclease agent, wherein the nuclease agent targets a nuclease target site in a genomic safe harbor locus, the genomic safe harbor locus being selected from the following genomic locations: (i) genomic coordinates of about 77460242 to about 77460537 on human chromosome 13; (ii) genomic coordinates of about 170031084 to about 170031382 on human chromosome 6; and (iii) genomic coordinates of about 25207412 to about 25207703 on human chromosome 9; and (b) a nucleic acid construct, wherein the nucleic acid construct comprises a nucleic acid operably linked to a promoter, the nucleic acid encoding a product of interest, wherein the nuclease agent cleaves the nuclease target site and the nucleic acid construct is inserted into the genomic safe harbor locus. Also provided are methods for expressing a product of interest from a genomic safe harbor locus in a cell, such as a human cell (e.g., a mammalian cell).Such methods include administering to a cell (e.g., a human cell) (a) a nuclease agent or one or more nucleic acids encoding a nuclease agent, wherein the nuclease agent targets a nuclease target site at a genomic safe harbor locus, the genomic safe harbor locus being located at the following genomic locations: (i) genomic coordinates of about 77460242 to about 77460537 on human chromosome 13; (ii) genomic coordinates of about 170031084 to about 170031382 on human chromosome 6; and (iii) genomic coordinates of about 252074 on human chromosome 9. The method may include administering (a) a nuclease agent or one or more nucleic acids encoding the nuclease agent, wherein the nucleic acid construct comprises a nucleic acid operably linked to a promoter, the nucleic acid encoding a product of interest, wherein the nuclease agent cleaves the nuclease target site, the nucleic acid construct is inserted into the genomic safe harbor locus to create a modified genomic safe harbor locus, and the product of interest is expressed from the modified genomic safe harbor locus. In some such methods, the cell (e.g., a human cell) is a liver cell. In some such methods, the cell (e.g., a human cell) is a liver cell. In some such methods, the cell (e.g., a human cell) is in vitro or ex vivo. In some such methods, the cell (e.g., a human cell) is present in a subject in vivo. Also provided are methods of integrating a nucleic acid construct into a genomic safe harbor locus in a cell (e.g., a mammalian cell) of a subject (e.g., a mammalian subject), such as a human cell of a human subject.Such methods may include administering to a subject (e.g., a human subject) (a) a nuclease agent or one or more nucleic acids encoding a nuclease agent, wherein the nuclease agent targets a nuclease target site in a genomic safe harbor locus, the genomic safe harbor locus being selected from the following genomic locations: (i) genomic coordinates of about 77460242 to about 77460537 on human chromosome 13; (ii) genomic coordinates of about 170031084 to about 170031382 on human chromosome 6; and (iii) genomic coordinates of about 25207412 to about 25207703 on human chromosome 9; and (b) a nucleic acid construct, wherein the nucleic acid construct comprises a nucleic acid operably linked to a promoter, the nucleic acid encoding a product of interest, wherein the nuclease agent cleaves the nuclease target site and the nucleic acid construct is inserted into the genomic safe harbor locus. Also provided are methods for expressing a product of interest from a genomic safe harbor locus in a cell (e.g., a mammalian cell) of a subject (e.g., a mammalian subject), such as a human cell of a human subject. Such methods include administering to the subject (e.g., a human subject) (a) a nuclease agent or one or more nucleic acids encoding a nuclease agent, wherein the nuclease agent targets a nuclease target site in the genomic safe harbor locus, the genomic safe harbor locus being located at the following genomic locations: (i) genomic coordinates from about 77460242 to about 77460537 on human chromosome 13, (ii) genomic coordinates from about 170031084 to about 170031382 on human chromosome 6, and (iii) genomic coordinates from about 252074 on human chromosome 9. The method may include administering (a) a nuclease agent or one or more nucleic acids encoding the nuclease agent, selected from genomic coordinates 12 to about 25,207,703; and (b) a nucleic acid construct, wherein the nucleic acid construct comprises a nucleic acid operably linked to a promoter, the nucleic acid encoding a product of interest, wherein the nuclease agent cleaves the nuclease target site, and the nucleic acid construct is inserted into the genomic safe harbor locus to create a modified genomic safe harbor locus, and wherein the product of interest is expressed from the modified genomic safe harbor locus.In some such methods, the cell (e.g., a human cell) is a liver cell. In some such methods, the cell (e.g., a human cell) is a liver cell.
[0007] In some such methods, the genomic safe harbor locus is selected from the following genomic locations: (i) human chromosome 13, coordinates 77460242-77460537; (ii) human chromosome 6, coordinates 170031084-170031382; and (iii) human chromosome 9, coordinates 25207412-25207703. In some such methods, the genomic safe harbor locus is at genomic coordinates from about 77460242 to about 77460537 on human chromosome 13. In some such methods, the genomic safe harbor locus is human chromosome 13, coordinates 77460242-77460537, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 39. In some such methods, the genomic safe harbor locus is at genomic coordinates from about 170031084 to about 170031382 on human chromosome 6. In some such methods, the genomic safe harbor locus is human chromosome 6, coordinates 170031084-170031382, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 40. In some such methods, the genomic safe harbor locus is genomic coordinates from about 25207412 to about 25207703 on human chromosome 9. In some such methods, the genomic safe harbor locus is human chromosome 9, coordinates 25207412-25207703, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 41.
[0008] In some such methods, the nuclease agent comprises (a) a zinc finger nuclease (ZFN), (b) a transcription activator-like effector nuclease (TALEN), or (c) (i) a Cas protein or a nucleic acid encoding a Cas protein, and (ii) a guide RNA or one or more DNAs encoding the guide RNA, wherein the guide RNA comprises a DNA-targeting segment that targets the guide RNA target sequence, and the guide RNA binds to the Cas protein and targets the Cas protein to the guide RNA target sequence.
[0009] In some such methods, the nuclease agent comprises (a) a Cas protein or a nucleic acid encoding a Cas protein, and (b) a guide RNA or one or more DNAs encoding the guide RNA, wherein the guide RNA comprises a DNA-targeting segment that targets the guide RNA target sequence, and the guide RNA binds to the Cas protein and targets the Cas protein to the guide RNA target sequence. In some such methods, the method comprises administering the guide RNA in the form of RNA. In some such methods, the guide RNA comprises at least one modification. In some such methods, the at least one modification comprises a 2'-O-methyl modified nucleotide. In some such methods, the at least one modification comprises an internucleotide phosphorothioate bond. In some such methods, the guide RNA is a single guide RNA (sgRNA). In some such methods, the Cas protein is a Cas9 protein. In some such methods, the Cas protein is a CasX protein. In some such methods, the Cas protein is a CasΦ protein. In some such methods, the Cas protein is a Cpfl protein. In some such methods, the Cas9 protein is derived from a Streptococcus pyogenes Cas9 protein, a Staphylococcus aureus Cas9 protein, a Campylobacter jejuni Cas9 protein, a Streptococcus thermophilus Cas9 protein, or a Neisseria meningitidis Cas9 protein. In some such methods, the Cas protein is derived from a Streptococcus pyogenes Cas9 protein. In some such methods, a nucleic acid encoding the Cas protein is codon-optimized for expression in a mammalian cell or a human cell. In some such methods, the method comprises administering a nucleic acid encoding the Cas protein, wherein the nucleic acid comprises an mRNA encoding the Cas protein.In some such methods, the mRNA encoding the Cas protein includes at least one modification. In some such methods, the Cas protein or a nucleic acid encoding the Cas protein and the guide RNA or one or more DNAs encoding the guide RNA are associated with a lipid nanoparticle. In some such methods, the genomic safe harbor locus is selected from the following genomic locations: (i) human chromosome 13, coordinates 77460242-77460537; (ii) human chromosome 6, coordinates 170031084-170031382; and (iii) human chromosome 9, coordinates 25207412-25207703. In some such methods, the genomic safe harbor locus is at genomic coordinates of about 77460242 to about 77460537 on human chromosome 13. In some such methods, the genomic safe harbor locus is human chromosome 13, coordinates 77460242-77460537, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 39. In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 25, 45, and 228-256, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 25, 45, and 228-256, and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 25, 45, and 228-256, and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 25, 45, and 228-256.In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 25, 45, 235, 237, and 246; and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 25, 45, 235, 237, and 246; and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 25, 45, 235, 237, and 246; and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 25, 45, 235, 237, and 246. In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in SEQ ID NO:25, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in SEQ ID NO:25. In some such methods, the DNA-targeting segment comprises SEQ ID NO:25. In some such methods, the DNA-targeting segment consists of SEQ ID NO:25. In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in SEQ ID NO:45, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in SEQ ID NO:45. In some such methods, the DNA-targeting segment comprises SEQ ID NO:45. In some such methods, the DNA-targeting segment consists of SEQ ID NO:45. In some such methods, the genomic safe harbor locus is at genomic coordinates from about 170031084 to about 170031382 on human chromosome 6. In some such methods, the genomic safe harbor locus is human chromosome 6, coordinates 170031084 to 170031382, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO:40.In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 26, 46, and 257-285, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 26, 46, and 257-285, and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 26, 46, and 257-285, and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 26, 46, and 257-285. In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 26, 46, 268, 271, and 280; and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 26, 46, 268, 271, and 280; and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 26, 46, 268, 271, and 280; and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 26, 46, 268, 271, and 280. In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in SEQ ID NO: 26, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in SEQ ID NO: 26. In some such methods, the DNA-targeting segment comprises SEQ ID NO: 26. In some such methods, the DNA-targeting segment consists of SEQ ID NO: 26.In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in SEQ ID NO: 46, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in SEQ ID NO: 46. In some such methods, the DNA-targeting segment comprises SEQ ID NO: 46. In some such methods, the DNA-targeting segment consists of SEQ ID NO: 46. In some such methods, the genomic safe harbor locus is genomic coordinates from about 25207412 to about 25207703 on human chromosome 9. In some such methods, the genomic safe harbor locus is human chromosome 9, coordinates 25207412 to 25207703, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 41. In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 27, 47, and 286-314, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 27, 47, and 286-314, and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 27, 47, and 286-314, and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 27, 47, and 286-314.In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 27, 47, 288, 296, 305, 306, and 310; and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 27, 47, 288, 296, 305, 306, and 310; and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 27, 47, 288, 296, 305, 306, and 310; and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 27, 47, 288, 296, 305, 306, and 310. In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in SEQ ID NO:27, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in SEQ ID NO:27. In some such methods, the DNA-targeting segment comprises SEQ ID NO:27. In some such methods, the DNA-targeting segment consists of SEQ ID NO:27. In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in SEQ ID NO:47, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in SEQ ID NO:47. In some such methods, the DNA-targeting segment comprises SEQ ID NO:47. In some such methods, the DNA-targeting segment consists of SEQ ID NO:47.
[0010] Also provided are methods of integrating a nucleic acid construct into a genomic safe harbor locus in a mouse cell. Some such methods include introducing into a mouse cell (a) a nuclease agent or one or more nucleic acids encoding a nuclease agent, wherein the nuclease agent targets a nuclease target site at a genomic safe harbor locus, the genomic safe harbor locus being located at the following genomic locations: (i) genomic coordinates of about 103,450,397 to about 103,451,396 on mouse chromosome 14; (ii) genomic coordinates of about 15,226,387 to about 15,227,386 on mouse chromosome 17; and (iii) genomic coordinates from about 92,827,563 to about 92,828,592 on mouse chromosome 4. The method may include administering (a) a nuclease agent or one or more nucleic acids encoding the nuclease agent, and (b) a nucleic acid construct, wherein the nucleic acid construct comprises a nucleic acid operably linked to a promoter, the nucleic acid encoding a product of interest, wherein the nuclease agent cleaves the nuclease target site, and the nucleic acid construct is inserted into the genomic safe harbor locus. Also provided is a method for expressing a product of interest from a genomic safe harbor locus in a mouse cell. Some such methods include administering to a mouse cell (a) a nuclease agent or one or more nucleic acids encoding a nuclease agent, wherein the nuclease agent targets a nuclease target site at a genomic safe harbor locus, the genomic safe harbor locus being located at the following genomic locations: (i) genomic coordinates of about 103,450,397 to about 103,451,396 on mouse chromosome 14; (ii) genomic coordinates of about 15,226,387 to about 15,227,386 on mouse chromosome 17; and (iii) genomic coordinates of about 92 on mouse chromosome 4. The method includes administering (a) a nuclease agent or one or more nucleic acids encoding the nuclease agent, selected from genomic coordinates of about 92,827,563 to about 92,828,592; and (b) a nucleic acid construct, wherein the nucleic acid construct comprises a nucleic acid operably linked to a promoter, the nucleic acid encoding a desired product, wherein the nuclease agent cleaves the nuclease target site, and the nucleic acid construct is inserted into the genomic safe harbor locus to create a modified genomic safe harbor locus, and the desired product is expressed from the modified genomic safe harbor locus.In some such methods, the mouse cell is a liver cell. In some such methods, the mouse cell is a hepatocyte. In some such methods, the mouse tissue is in vitro or ex vivo. In some such methods, the mouse cell is in vivo in a subject. Also provided are methods of integrating a nucleic acid construct into a genomic safe harbor locus in a mouse cell of a mouse subject. Some such methods include introducing into the mouse subject (a) a nuclease agent or one or more nucleic acids encoding a nuclease agent, wherein the nuclease agent targets a nuclease target site in the genomic safe harbor locus, the genomic safe harbor locus being located at the following genomic locations: (i) genomic coordinates of about 103,450,397 to about 103,451,396 on mouse chromosome 14; (ii) genomic coordinates of about 15,226,387 to about 15,227,386 on mouse chromosome 17; (b) a nuclease agent or one or more nucleic acids encoding the nuclease agent, wherein the nuclease agent cleaves the nuclease target site and the nucleic acid construct is inserted into the genomic safe harbor locus. Also provided is a method for expressing a product of interest from a genomic safe harbor locus in a mouse cell of a mouse subject, the method comprising administering to the mouse subject a nuclease agent or one or more nucleic acids encoding the nuclease agent, and (c) a nucleic acid construct, the nucleic acid construct comprising a nucleic acid operably linked to a promoter, the nucleic acid encoding a product of interest, wherein the nuclease agent cleaves the nuclease target site and the nucleic acid construct is inserted into the genomic safe harbor locus.Some such methods include administering to a mouse subject (a) a nuclease agent or one or more nucleic acids encoding a nuclease agent, wherein the nuclease agent targets a nuclease target site at a genomic safe harbor locus, the genomic safe harbor locus being located at the following genomic locations: (i) genomic coordinates of about 103,450,397 to about 103,451,396 on mouse chromosome 14; (ii) genomic coordinates of about 15,226,387 to about 15,227,386 on mouse chromosome 17; and (iii) genomic coordinates of about 92 on mouse chromosome 4. The method includes administering (a) a nuclease agent or one or more nucleic acids encoding the nuclease agent, wherein the nuclease agent is selected from genomic coordinates of about 92,827,563 to about 92,828,592; and (b) a nucleic acid construct, wherein the nucleic acid construct comprises a nucleic acid operably linked to a promoter, the nucleic acid encoding a product of interest, wherein the nuclease agent cleaves the nuclease target site, and the nucleic acid construct is inserted into the genomic safe harbor locus to create a modified genomic safe harbor locus, and the product of interest is expressed from the modified genomic safe harbor locus. In some such methods, the mouse cell is a liver cell. In some such methods, the mouse cell is a liver cell.
[0011] In some such methods, the genomic safe harbor locus is selected from the following genomic locations: (i) mouse chromosome 14, coordinates 103,450,397-103,451,396; (ii) mouse chromosome 17, coordinates 15,226,387-15,227,386; and (iii) human chromosome 4, coordinates 92,827,563-92,828,592. In some such methods, the genomic safe harbor locus is at genomic coordinates of about 103,450,397 to about 103,451,396 on mouse chromosome 14. In some such methods, the genomic safe harbor locus is mouse chromosome 14, coordinates 103,450,397-103,451,396, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO:405. In some such methods, the genomic safe harbor locus is at genomic coordinates from about 15,226,387 to about 15,227,386 on mouse chromosome 17. In some such methods, the genomic safe harbor locus is at mouse chromosome 17, coordinates 15,226,387 to 15,227,386, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 406. In some such methods, the genomic safe harbor locus is at genomic coordinates from about 92,827,563 to about 92,828,592 on mouse chromosome 4. In some such methods, the genomic safe harbor locus is at mouse chromosome 4, coordinates 92,827,563 to 92,828,592, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 407.
[0012] In some such methods, the nuclease agent comprises (a) a zinc finger nuclease (ZFN), (b) a transcription activator-like effector nuclease (TALEN), or (c) (i) a Cas protein or a nucleic acid encoding a Cas protein, and (ii) a guide RNA or one or more DNAs encoding the guide RNA, wherein the guide RNA comprises a DNA-targeting segment that targets the guide RNA target sequence, and the guide RNA binds to the Cas protein and targets the Cas protein to the guide RNA target sequence. In some such methods, the nuclease agent comprises (a) a Cas protein or a nucleic acid encoding a Cas protein, and (b) a guide RNA or one or more DNAs encoding the guide RNA, wherein the guide RNA comprises a DNA-targeting segment that targets the guide RNA target sequence, and the guide RNA binds to the Cas protein and targets the Cas protein to the guide RNA target sequence. In some such methods, the method comprises administering the guide RNA in the form of RNA. In some such methods, the guide RNA comprises at least one modification. In some such methods, at least one modification comprises a 2'-O-methyl modified nucleotide. In some such methods, at least one modification comprises an internucleotide phosphorothioate linkage. In some such methods, the guide RNA is a single guide RNA (sgRNA). In some such methods, the Cas protein is a Cas9 protein. In some such methods, the Cas9 protein is derived from a Streptococcus pyogenes Cas9 protein, a Staphylococcus aureus Cas9 protein, a Campylobacter jejuni Cas9 protein, a Streptococcus thermophilus Cas9 protein, or a Neisseria meningitidis Cas9 protein. In some such methods, the Cas protein is derived from a Streptococcus pyogenes Cas9 protein.In some such methods, the nucleic acid encoding the Cas protein is codon-optimized for expression in mammalian or mouse cells. In some such methods, the method includes administering a nucleic acid encoding a Cas protein, wherein the nucleic acid includes an mRNA encoding the Cas protein. In some such methods, the mRNA encoding the Cas protein includes at least one modification. In some such methods, the Cas protein or nucleic acid encoding the Cas protein and the guide RNA or one or more DNAs encoding the guide RNAs are associated with a lipid nanoparticle. In some such methods, the genomic safe harbor locus is selected from the following genomic locations: (i) mouse chromosome 14, coordinates 103,450,397-103,451,396; (ii) mouse chromosome 17, coordinates 15,226,387-15,227,386; and (iii) human chromosome 4, coordinates 92,827,563-92,828,592. In some such methods, the genomic safe harbor locus is at genomic coordinates of about 103,450,397 to about 103,451,396 on mouse chromosome 14. In some such methods, the genomic safe harbor locus is mouse chromosome 14, coordinates 103,450,397 to 103,451,396, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO:405. In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 315-344, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 315-344, and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 315-344, and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 315-344.In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 318, 320, 321, and 341, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 318, 320, 321, and 341, and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 318, 320, 321, and 341, and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 318, 320, 321, and 341. In some such methods, the genomic safe harbor locus is at genomic coordinates from about 15,226,387 to about 15,227,386 on mouse chromosome 17. In some such methods, the genomic safe harbor locus is mouse chromosome 17, coordinates 15,226,387-15,227,386, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 406. In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 345-374, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 345-374, and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 345-374, and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 345-374.In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 347, 360, 369, and 370, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 347, 360, 369, and 370, and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 347, 360, 369, and 370, and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 347, 360, 369, and 370. In some such methods, the genomic safe harbor locus is at genomic coordinates from about 92,827,563 to about 92,828,592 on mouse chromosome 4. In some such methods, the genomic safe harbor locus is mouse chromosome 4, coordinates 92,827,563-92,828,592, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 407. In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 375-404, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 375-404, and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 375-404, and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 375-404.In some such methods, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 379, 380, and 388; and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 379, 380, and 388; and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 379, 380, and 388; and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 379, 380, and 388.
[0013] In some such methods, the nucleic acid construct is administered simultaneously with the nuclease agent or one or more nucleic acids encoding the nuclease agent. In some such methods, the nucleic acid construct is not administered simultaneously with the nuclease agent or one or more nucleic acids encoding the nuclease agent. In some such methods, the nucleic acid construct is administered before the nuclease agent or one or more nucleic acids encoding the nuclease agent. In some such methods, the nucleic acid construct is administered after the nuclease agent or one or more nucleic acids encoding the nuclease agent.
[0014] In some such methods, the product of interest is a polypeptide of interest. In some such methods, the polypeptide of interest comprises a therapeutic polypeptide. In some such methods, the polypeptide of interest is a secreted polypeptide. In some such methods, the polypeptide of interest is an intracellular polypeptide.
[0015] In some such methods, the promoter is active in liver cells. In some such methods, the promoter is a tissue-specific promoter. In some such methods, the promoter is a constitutive promoter. In some such methods, the promoter is an inducible promoter.
[0016] In some such methods, the nucleic acid construct does not comprise homology arms. In some such methods, the nucleic acid construct is inserted into the target genomic locus via non-homologous end joining. In some such methods, the nucleic acid construct comprises homology arms. In some such methods, the nucleic acid construct is inserted into the target genomic locus via homology directed repair. In some such methods, the nucleic acid construct is single-stranded DNA or double-stranded DNA. In some such methods, the nucleic acid construct is single-stranded DNA.
[0017] In some such methods, the nucleic acid construct is in a nucleic acid vector or lipid nanoparticle. In some such methods, the nucleic acid construct is in a nucleic acid vector. In some such methods, the nucleic acid vector is a viral vector. In some such methods, the nucleic acid vector is an adeno-associated virus (AAV) vector. In some such methods, the AAV vector is a single-stranded AAV (ssAAV) vector. In some such methods, the AAV vector is derived from an AAV8 vector, an AAV3B vector, an AAV5 vector, an AAV6 vector, an AAV7 vector, an AAV9 vector, an AAVrh.74 vector, an AAV-DJ vector, or an AAVhu.37 vector. In some such methods, the AAV vector is a recombinant AAV8 (rAAV8) vector. In some such methods, the AAV vector is a single-stranded rAAV8 vector.
[0018] In another aspect, a cell (e.g., a mammalian cell, such as a human cell) produced by any of the above methods is provided. In another aspect, a cell (e.g., a mammalian cell, such as a human cell) is provided that comprises a nucleic acid construct integrated into a genomic safe harbor locus. In some such cells, the nucleic acid construct comprises a nucleic acid operably linked to a promoter, the nucleic acid encoding a product of interest, and the genomic safe harbor locus is selected from the following genomic locations: (i) genomic coordinates from about 77460242 to about 77460537 on human chromosome 13; (ii) genomic coordinates from about 170031084 to about 170031382 on human chromosome 6; and (iii) genomic coordinates from about 25207412 to about 25207703 on human chromosome 9. In some such cells, the nucleic acid construct comprises a nucleic acid operably linked to a promoter, the nucleic acid encoding a product of interest, and the genomic safe harbor locus is selected from the following genomic locations: (i) genomic coordinates from about 103,450,397 to about 103,451,396 on mouse chromosome 14; (ii) genomic coordinates from about 15,226,387 to about 15,227,386 on mouse chromosome 17; and (iii) genomic coordinates from about 92,827,563 to about 92,828,592 on mouse chromosome 4.
[0019] In some such methods, the cell is a human cell. In some such methods, the cell is a mouse cell. In some such methods, the cell is a liver cell (e.g., a human liver cell). In some such methods, the cell is a liver cell (e.g., a human hepatocyte).
[0020] In some such cells, a product of interest is expressed. In some such cells, the product of interest is a polypeptide of interest. In some such cells, the polypeptide of interest includes a therapeutic polypeptide. In some such cells, the polypeptide of interest is a secreted polypeptide. In some such cells, the polypeptide of interest is an intracellular polypeptide. In some such cells, the promoter is active in liver cells. In some such cells, the promoter is a tissue-specific promoter. In some such cells, the promoter is a constitutive promoter. In some such cells, the promoter is an inducible promoter.
[0021] In some such cells, the genomic safe harbor locus is selected from the following genomic locations: (i) human chromosome 13, coordinates 77460242-77460537; (ii) human chromosome 6, coordinates 170031084-170031382; and (iii) human chromosome 9, coordinates 25207412-25207703. In some such cells, the genomic safe harbor locus is at genomic coordinates from about 77460242 to about 77460537 on human chromosome 13. In some such cells, the genomic safe harbor locus is human chromosome 13, coordinates 77460242-77460537, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 39. In some such cells, the genomic safe harbor locus is at genomic coordinates from about 170031084 to about 170031382 on human chromosome 6. In some such cells, the genomic safe harbor locus is human chromosome 6, coordinates 170031084 to 170031382, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 40. In some such cells, the genomic safe harbor locus is genomic coordinates from about 25207412 to about 25207703 on human chromosome 9. In some such cells, the genomic safe harbor locus is human chromosome 9, coordinates 25207412 to 25207703, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 41.
[0022] In some such cells, the genomic safe harbor locus is selected from the following genomic locations: (i) mouse chromosome 14, coordinates 103,450,397-103,451,396; (ii) mouse chromosome 17, coordinates 15,226,387-15,227,386; and (iii) human chromosome 4, coordinates 92,827,563-92,828,592. In some such cells, the genomic safe harbor locus is at genomic coordinates of about 103,450,397 to about 103,451,396 on mouse chromosome 14. In some such cells, the genomic safe harbor locus is mouse chromosome 14, coordinates 103,450,397-103,451,396, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO:405. In some such cells, the genomic safe harbor locus is at genomic coordinates of about 15,226,387 to about 15,227,386 on mouse chromosome 17. In some such cells, the genomic safe harbor locus is at mouse chromosome 17, coordinates 15,226,387 to 15,227,386, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 406. In some such cells, the genomic safe harbor locus is at genomic coordinates of about 92,827,563 to about 92,828,592 on mouse chromosome 4. In some such cells, the genomic safe harbor locus is at mouse chromosome 4, coordinates 92,827,563 to 92,828,592, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 407.
[0023] In another aspect, a composition is provided comprising a guide RNA or DNA encoding a guide RNA, wherein the guide RNA comprises a DNA-targeting segment that targets a guide RNA target sequence within a genomic safe harbor locus and a protein-binding segment that binds to a Cas protein, wherein the genomic safe harbor locus is selected from the following genomic locations: (i) genomic coordinates from about 77460242 to about 77460537 on human chromosome 13; (ii) genomic coordinates from about 170031084 to about 170031382 on human chromosome 6; and (iii) genomic coordinates from about 25207412 to about 25207703 on human chromosome 9. In another aspect, a composition is provided comprising a guide RNA or DNA encoding a guide RNA, wherein the guide RNA comprises a DNA-targeting segment that targets a guide RNA target sequence within a genomic safe harbor locus and a protein-binding segment that binds to a Cas protein, wherein the genomic safe harbor locus is selected from the following genomic locations: (i) genomic coordinates of about 103,450,397 to about 103,451,396 on mouse chromosome 14; (ii) genomic coordinates of about 15,226,387 to about 15,227,386 on mouse chromosome 17; and (iii) genomic coordinates of about 92,827,563 to about 92,828,592 on mouse chromosome 4.
[0024] In some such compositions, the genomic safe harbor locus is selected from the following genomic locations: (i) human chromosome 13, coordinates 77460242-77460537; (ii) human chromosome 6, coordinates 170031084-170031382; and (iii) human chromosome 9, coordinates 25207412-25207703. In some such compositions, the genomic safe harbor locus is at genomic coordinates from about 77460242 to about 77460537 on human chromosome 13. In some such compositions, the genomic safe harbor locus is human chromosome 13, coordinates 77460242-77460537, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO:39. In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 25, 45, and 228-256; and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 25, 45, and 228-256; and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 25, 45, and 228-256; and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 25, 45, and 228-256. In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 25, 45, 235, 237, and 246; and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 25, 45, 235, 237, and 246; and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 25, 45, 235, 237, and 246; and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 25, 45, 235, 237, and 246.In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in SEQ ID NO:25, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in SEQ ID NO:25. In some such compositions, the DNA-targeting segment comprises SEQ ID NO:25. In some such compositions, the DNA-targeting segment consists of SEQ ID NO:25. In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in SEQ ID NO:45, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in SEQ ID NO:45. In some such compositions, the DNA-targeting segment comprises SEQ ID NO:45. In some such compositions, the DNA-targeting segment consists of SEQ ID NO:45. In some such compositions, the genomic safe harbor locus is at genomic coordinates from about 170031084 to about 170031382 on human chromosome 6. In some such compositions, the genomic safe harbor locus is human chromosome 6, coordinates 170031084 to 170031382, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO:40. In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 26, 46, and 257-285, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 26, 46, and 257-285, and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 26, 46, and 257-285, and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 26, 46, and 257-285.In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 26, 46, 268, 271, and 280; and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 26, 46, 268, 271, and 280; and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 26, 46, 268, 271, and 280; and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 26, 46, 268, 271, and 280. In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in SEQ ID NO:26, and / or (II) the DNA-targeting segment is at least 90%, or at least 95%, identical to the sequence set forth in SEQ ID NO:26. In some such compositions, the DNA-targeting segment comprises SEQ ID NO:26. In some such compositions, the DNA-targeting segment consists of SEQ ID NO:26. In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in SEQ ID NO:46, and / or (II) the DNA-targeting segment is at least 90%, or at least 95% identical to the sequence set forth in SEQ ID NO:46. In some such compositions, the DNA-targeting segment comprises SEQ ID NO:46. In some such compositions, the DNA-targeting segment consists of SEQ ID NO:46. In some such compositions, the genomic safe harbor locus is at genomic coordinates from about 25207412 to about 25207703 on human chromosome 9. In some such compositions, the genomic safe harbor locus is human chromosome 9, coordinates 25207412 to 25207703, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO:41.In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 27, 47, and 286-314, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 27, 47, and 286-314, and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 27, 47, and 286-314, and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 27, 47, and 286-314. In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 27, 47, 288, 296, 305, 306, and 310; and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 27, 47, 288, 296, 305, 306, and 310; and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 27, 47, 288, 296, 305, 306, and 310; and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 27, 47, 288, 296, 305, 306, and 310. In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in SEQ ID NO: 27, and / or (II) the DNA-targeting segment is at least 90%, or at least 95% identical to the sequence set forth in SEQ ID NO: 27. In some such compositions, the DNA-targeting segment comprises SEQ ID NO: 27. In some such compositions, the DNA-targeting segment consists of SEQ ID NO: 27.In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in SEQ ID NO: 47, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in SEQ ID NO: 47. In some such compositions, the DNA-targeting segment comprises SEQ ID NO: 47. In some such compositions, the DNA-targeting segment consists of SEQ ID NO: 47.
[0025] In some such cells, the genomic safe harbor locus is selected from the following genomic locations: (i) mouse chromosome 14, coordinates 103,450,397-103,451,396; (ii) mouse chromosome 17, coordinates 15,226,387-15,227,386; and (iii) human chromosome 4, coordinates 92,827,563-92,828,592. In some such compositions, the genomic safe harbor locus is at genomic coordinates from about 103,450,397 to about 103,451,396 on mouse chromosome 14. In some such compositions, the genomic safe harbor locus is mouse chromosome 14, coordinates 103,450,397-103,451,396, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO:405. In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 315-344, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 315-344, and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 315-344, and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 315-344. In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 318, 320, 321, and 341; and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 318, 320, 321, and 341; and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 318, 320, 321, and 341; and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 318, 320, 321, and 341.In some such compositions, the genomic safe harbor locus is at genomic coordinates from about 15,226,387 to about 15,227,386 on mouse chromosome 17. In some such compositions, the genomic safe harbor locus is mouse chromosome 17, coordinates 15,226,387 to 15,227,386, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO:406. In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 345-374, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 345-374, and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 345-374, and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 345-374. In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 347, 360, 369, and 370, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 347, 360, 369, and 370, and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 347, 360, 369, and 370, and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 347, 360, 369, and 370. In some such compositions, the genomic safe harbor locus is at genomic coordinates from about 92,827,563 to about 92,828,592 on mouse chromosome 4. In some such compositions, the genomic safe harbor locus is mouse chromosome 4, coordinates 92,827,563 to 92,828,592, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO:407.In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 375-404, and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 375-404, and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 375-404, and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 375-404. In some such compositions, (I) the DNA-targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in any one of SEQ ID NOs: 379, 380, and 388; and / or (II) the DNA-targeting segment is at least 90% or at least 95% identical to the sequence set forth in any one of SEQ ID NOs: 379, 380, and 388; and / or (III) the DNA-targeting segment comprises any one of SEQ ID NOs: 379, 380, and 388; and / or (IV) the DNA-targeting segment consists of any one of SEQ ID NOs: 379, 380, and 388.
[0026] In some such compositions, the composition comprises DNA encoding the guide RNA. In some such compositions, the DNA encoding the guide RNA is in a nucleic acid vector. In some such compositions, the nucleic acid vector is a viral vector. In some such compositions, the nucleic acid vector is an adeno-associated virus (AAV) vector. In some such compositions, the AAV vector is a single-stranded AAV (ssAAV) vector. In some such compositions, the AAV vector is derived from an AAV8 vector, an AAV3B vector, an AAV5 vector, an AAV6 vector, an AAV7 vector, an AAV9 vector, an AAVrh.74 vector, an AAV-DJ vector, or an AAVhu.37 vector. In some such compositions, the AAV vector is a recombinant AAV8 (rAAV8) vector. In some such compositions, the AAV vector is a single-stranded rAAV8 vector. In some such compositions, the composition comprises a guide RNA in the form of RNA. In some such compositions, the guide RNA comprises at least one modification. In some such compositions, the at least one modification comprises a 2'-O-methyl modified nucleotide. In some such compositions, at least one modification comprises a phosphorothioate internucleotide linkage. In some such compositions, the guide RNA is a single guide RNA (sgRNA).
[0027] In some such compositions, the composition further comprises a Cas protein or a nucleic acid encoding a Cas protein. In some such compositions, the composition comprises a Cas protein. In some such compositions, the composition comprises a nucleic acid encoding a Cas protein. In some such compositions, the nucleic acid encoding the Cas protein is codon-optimized for expression in mammalian or human cells. In some such compositions, the nucleic acid encoding the Cas protein comprises DNA encoding the Cas protein. In some such compositions, the DNA encoding the guide RNA is in a nucleic acid vector. In some such compositions, the nucleic acid vector is a viral vector. In some such compositions, the nucleic acid vector is an adeno-associated virus (AAV) vector. In some such compositions, the AAV vector is a single-stranded AAV (ssAAV) vector. In some such compositions, the AAV vector is derived from an AAV8 vector, an AAV3B vector, an AAV5 vector, an AAV6 vector, an AAV7 vector, an AAV9 vector, an AAVrh.74 vector, an AAV-DJ vector, or an AAVhu.37 vector. In some such compositions, the AAV vector is a recombinant AAV8 (rAAV8) vector. In some such compositions, the AAV vector is a single-stranded rAAV8 vector. In some such compositions, the nucleic acid encoding the Cas protein comprises an mRNA encoding the Cas protein. In some such compositions, the mRNA encoding the Cas protein comprises at least one modification. In some such compositions, the Cas protein or nucleic acid encoding the Cas protein and the guide RNA or one or more DNAs encoding the guide RNAs are associated with a lipid nanoparticle. In some such compositions, the Cas protein is a Cas9 protein. In some such compositions, the Cas protein is a CasX protein. In some such compositions, the Cas protein is a CasΦ protein. In some such compositions, the Cas protein is a Cpf1 protein.In some such compositions, the Cas9 protein is derived from a Streptococcus pyogenes Cas9 protein, a Staphylococcus aureus Cas9 protein, a Campylobacter jejuni Cas9 protein, a Streptococcus thermophilus Cas9 protein, or a Neisseria meningitidis Cas9 protein. In some such compositions, the Cas protein is derived from a Streptococcus pyogenes Cas9 protein.
[0028] In some such compositions, the composition further comprises a nucleic acid construct, the nucleic acid construct comprising a nucleic acid operably linked to a promoter, the nucleic acid encoding a product of interest. In some such compositions, the product of interest is a polypeptide of interest. In some such compositions, the polypeptide of interest comprises a therapeutic polypeptide. In some such compositions, the polypeptide of interest is a secreted polypeptide. In some such compositions, the polypeptide of interest is an intracellular polypeptide. In some such compositions, the promoter is active in liver cells. In some such compositions, the promoter is a tissue-specific promoter. In some such compositions, the promoter is a constitutive promoter. In some such compositions, the promoter is an inducible promoter. In some such compositions, the nucleic acid construct does not comprise homology arms. In some such compositions, the nucleic acid construct comprises homology arms. In some such compositions, the nucleic acid construct is single-stranded DNA or double-stranded DNA. In some such compositions, the nucleic acid construct is single-stranded DNA.
[0029] In some such compositions, the nucleic acid construct is in a nucleic acid vector or lipid nanoparticle. In some such compositions, the nucleic acid construct is in a nucleic acid vector. In some such compositions, the nucleic acid vector is a viral vector. In some such compositions, the nucleic acid vector is an adeno-associated virus (AAV) vector. In some such compositions, the AAV vector is a single-stranded AAV (ssAAV) vector. In some such compositions, the AAV vector is derived from an AAV8 vector, an AAV3B vector, an AAV5 vector, an AAV6 vector, an AAV7 vector, an AAV9 vector, an AAVrh.74 vector, an AAV-DJ vector, or an AAVhu.37 vector. In some such compositions, the AAV vector is a recombinant AAV8 (rAAV8) vector. In some such compositions, the AAV vector is a single-stranded rAAV8 vector.
[0030] In another aspect, a nucleic acid comprising a genomic safe harbor locus is provided, comprising an integrated nucleic acid construct. In some such nucleic acids, the nucleic acid construct comprises a nucleic acid operably linked to a promoter, the nucleic acid encoding a product of interest, and the genomic safe harbor locus is selected from the following genomic locations: (i) genomic coordinates of about 77460242 to about 77460537 on human chromosome 13, (ii) genomic coordinates of about 170031084 to about 170031382 on human chromosome 6, and (iii) genomic coordinates of about 25207412 to about 25207703 on human chromosome 9. In some such nucleic acids, the nucleic acid construct comprises a nucleic acid operably linked to a promoter, the nucleic acid encoding a product of interest, and the genomic safe harbor locus is selected from the following genomic locations: (i) genomic coordinates from about 103,450,397 to about 103,451,396 on mouse chromosome 14; (ii) genomic coordinates from about 15,226,387 to about 15,227,386 on mouse chromosome 17; and (iii) genomic coordinates from about 92,827,563 to about 92,828,592 on mouse chromosome 4.
[0031] In some such nucleic acids, the product of interest is a polypeptide of interest. In some such nucleic acids, the polypeptide of interest includes a therapeutic polypeptide. In some such nucleic acids, the polypeptide of interest is a secreted polypeptide. In some such nucleic acids, the polypeptide of interest is an intracellular polypeptide. In some such nucleic acids, the promoter is active in liver cells. In some such nucleic acids, the promoter is a tissue-specific promoter. In some such nucleic acids, the promoter is a constitutive promoter. In some such nucleic acids, the promoter is an inducible promoter.
[0032] In some such nucleic acids, the genomic safe harbor locus is selected from the following genomic locations: (i) human chromosome 13, coordinates 77460242-77460537; (ii) human chromosome 6, coordinates 170031084-170031382; and (iii) human chromosome 9, coordinates 25207412-25207703. In some such nucleic acids, the genomic safe harbor locus is at genomic coordinates from about 77460242 to about 77460537 on human chromosome 13. In some such nucleic acids, the genomic safe harbor locus is human chromosome 13, coordinates 77460242-77460537, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 39. In some such nucleic acids, the genomic safe harbor locus is at genomic coordinates from about 170031084 to about 170031382 on human chromosome 6. In some such nucleic acids, the genomic safe harbor locus is human chromosome 6, coordinates 170031084 to 170031382, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 40. In some such nucleic acids, the genomic safe harbor locus is genomic coordinates from about 25207412 to about 25207703 on human chromosome 9. In some such nucleic acids, the genomic safe harbor locus is human chromosome 9, coordinates 25207412 to 25207703, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 41.
[0033] In some such nucleic acids, the genomic safe harbor locus is selected from the following genomic locations: (i) mouse chromosome 14, coordinates 103,450,397-103,451,396; (ii) mouse chromosome 17, coordinates 15,226,387-15,227,386; and (iii) human chromosome 4, coordinates 92,827,563-92,828,592. In some such nucleic acids, the genomic safe harbor locus is at genomic coordinates of about 103,450,397 to about 103,451,396 on mouse chromosome 14. In some such nucleic acids, the genomic safe harbor locus is mouse chromosome 14, coordinates 103,450,397-103,451,396, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO:405. In some such nucleic acids, the genomic safe harbor locus is at genomic coordinates of about 15,226,387 to about 15,227,386 on mouse chromosome 17. In some such nucleic acids, the genomic safe harbor locus is at mouse chromosome 17, coordinates 15,226,387 to 15,227,386, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 406. In some such nucleic acids, the genomic safe harbor locus is at genomic coordinates of about 92,827,563 to about 92,828,592 on mouse chromosome 4. In some such nucleic acids, the genomic safe harbor locus is at mouse chromosome 4, coordinates 92,827,563 to 92,828,592, or comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 407.
[0034] In another aspect, methods for identifying one or more genomic safe harbor loci in a tissue or cell type of interest are provided. Some such methods include: (a) identifying accessible genomic loci in the tissue or cell type of interest; (b) selecting the genomic loci identified in step (a) based on safety criteria, functional silencing criteria, and / or structural accessibility criteria; and (c) selecting the genomic loci identified in step (b) based on the availability, efficacy, and specificity of guide RNAs. In some such methods, step (a) includes identifying accessible genomic loci using an assay for transposase-accessible chromatin by high-throughput sequencing. In some such methods, step (a) includes identifying accessible genomic loci using DNase I hypersensitive site sequencing. In some such methods, step (a) includes identifying accessible genomic loci using an assay for transposase-accessible chromatin by high-throughput sequencing and DNase I hypersensitive site sequencing. In some such methods, step (b) comprises selecting the genomic loci identified in step (a) based on safety criteria, functional silencing criteria, and structural accessibility criteria. In some such methods, the safety criteria of step (b) comprise selecting the genomic locus only if it is more than 300 kb from any cancer-associated gene, more than 300 kb from any miRNA or small RNA, and more than 50 kb from the 5' end of any gene. In some such methods, the functional silencing criteria of step (b) comprise selecting the genomic locus only if it is more than 50 kb from any origin of replication and more than 50 kb from any ultraconserved element. In some such methods, the structural accessibility criteria of step (b) comprise selecting the genomic locus only if it is not in a copy number variable region. In some such methods, the efficacy of step (c) comprises editing efficacy in a tissue or cell type of interest.In some such methods, the method further comprises analyzing the chromatin environment of the genomic loci selected in step (c) for markers to disqualify any genomic loci within regions predicted to be regulatory regions, heterochromatic regions, regions involved in chromatin tertiary organization, or transcriptionally active regions. In some such methods, markers for regulatory regions include H3K4me1, H3K27ac, and H3K4me3. In some such methods, markers for heterochromatic regions include H3K9me3. In some such methods, markers for regions involved in chromatin tertiary organization include CTCF. In some such methods, markers for transcriptionally active regions include H3K36me3, PolR2A, RNASeq-, and RNASeq+.In some such methods, step (a) comprises identifying accessible genomic loci using high-throughput sequencing and an assay for transposase-accessible chromatin by DNase I hypersensitive site sequencing; step (b) comprises selecting the genomic loci identified in step (a) based on a safety criterion, a functional silencing criterion, and a structural accessibility criterion; the safety criterion of step (b) comprises selecting the genomic locus only if it is more than 300 kb from any cancer-associated gene, more than 300 kb from any miRNA or small RNA, and more than 50 kb from the 5' end of any gene; and the functional silencing criterion of step (b) comprises selecting the genomic locus only if it is more than 50 kb from any origin of replication and more than 50 kb from any ultraconserved element. and wherein the structural accessibility criterion of step (b) comprises selecting a genomic locus only if the genomic locus is not in a copy number variable region, and the method further comprises analyzing the chromatin environment of the genomic locus selected in step (c) for markers to disqualify any genomic locus within a region predicted to be a regulatory region, a heterochromatic region, a region involved in chromatin 3D organization, or a transcriptionally active region, wherein markers for regulatory regions include H3K4me1, H3K27ac, and H3K4me3, markers for heterochromatic regions include H3K9me3, markers for regions involved in chromatin 3D organization include CTCF, and markers for transcriptionally active regions include H3K36me3, PolR2A, RNASeq-, and RNASeq+. In some such methods, the method is for identifying one or more genomic safe harbor loci in a human tissue or cell type of interest. In some such methods, the tissue or cell type of interest is liver. In some such methods, the tissue or cell type of interest is a hematopoietic cell. [Brief explanation of the drawings]
[0035] [Figure 1] FIG. 1 shows the systematic approach used to identify liver-specific extragenic genomic safe harbor loci.
[0036] [Figure 2] Figure 1 shows the editing efficiencies of 33 gRNAs covering 20 loci after screening in primary human hepatocytes derived from three different donors. Also shown are the editing efficiencies of control gRNAs targeting AAVS1, ROSA26, and CCR5.
[0037] [Figure 3A] Analyzing the chromatin environment based on Chip-Seq data for chromatin marks, we present manual curation of six potential liver-specific extragenic genomic safe harbor loci (L-SH4, L-SH11, L-SH17, L-SH5, L-SH18, and L-SH20, respectively) to disqualify any potential safe harbors within regulatory regions (H3K4me1, H3K27ac, H3K4me3), heterochromatic regions (H3K9me3), or regions predicted to be involved in chromatin organization (CTCF signaling). [Figure 3B] Analyzing the chromatin environment based on Chip-Seq data for chromatin marks, we present manual curation of six potential liver-specific extragenic genomic safe harbor loci (L-SH4, L-SH11, L-SH17, L-SH5, L-SH18, and L-SH20, respectively) to disqualify any potential safe harbors within regulatory regions (H3K4me1, H3K27ac, H3K4me3), heterochromatic regions (H3K9me3), or regions predicted to be involved in chromatin organization (CTCF signaling). [Figure 3C]Analyzing the chromatin environment based on Chip-Seq data for chromatin marks, we present manual curation of six potential liver-specific extragenic genomic safe harbor loci (L-SH4, L-SH11, L-SH17, L-SH5, L-SH18, and L-SH20, respectively) to disqualify any potential safe harbors within regulatory regions (H3K4me1, H3K27ac, H3K4me3), heterochromatic regions (H3K9me3), or regions predicted to be involved in chromatin organization (CTCF signaling). [Figure 3D] Analyzing the chromatin environment based on Chip-Seq data for chromatin marks, we present manual curation of six potential liver-specific extragenic genomic safe harbor loci (L-SH4, L-SH11, L-SH17, L-SH5, L-SH18, and L-SH20, respectively) to disqualify any potential safe harbors within regulatory regions (H3K4me1, H3K27ac, H3K4me3), heterochromatic regions (H3K9me3), or regions predicted to be involved in chromatin organization (CTCF signaling). [Figure 3E] Analyzing the chromatin environment based on Chip-Seq data for chromatin marks, we present manual curation of six potential liver-specific extragenic genomic safe harbor loci (L-SH4, L-SH11, L-SH17, L-SH5, L-SH18, and L-SH20, respectively) to disqualify any potential safe harbors within regulatory regions (H3K4me1, H3K27ac, H3K4me3), heterochromatic regions (H3K9me3), or regions predicted to be involved in chromatin organization (CTCF signaling). [Figure 3F]Analyzing the chromatin environment based on Chip-Seq data for chromatin marks, we present manual curation of six potential liver-specific extragenic genomic safe harbor loci (L-SH4, L-SH11, L-SH17, L-SH5, L-SH18, and L-SH20, respectively) to disqualify any potential safe harbors within regulatory regions (H3K4me1, H3K27ac, H3K4me3), heterochromatic regions (H3K9me3), or regions predicted to be involved in chromatin organization (CTCF signaling).
[0038] [Figure 4A] Figure 4A shows the editing efficiency at the L-SH5, L-SH18, and L-SH20 genomic loci in primary human hepatocytes in 96-well plates 96 hours after transfection with 100 ng of Cas9 mRNA and 25 nM of sgRNA (Figure 4A) or 96 hours after administration of Cas9 mRNA and sgRNA via lipid nanoparticles (1 μg / mL dose) (Figure 4B). To assess editing efficiency, next-generation sequencing (NGS) was used to determine the percentage of cells with insertions / deletions (indels). [Figure 4B] Figure 4A shows the editing efficiency at the L-SH5, L-SH18, and L-SH20 genomic loci in primary human hepatocytes in 96-well plates 96 hours after transfection with 100 ng of Cas9 mRNA and 25 nM of sgRNA (Figure 4A) or 96 hours after administration of Cas9 mRNA and sgRNA via lipid nanoparticles (1 μg / mL dose) (Figure 4B). To assess editing efficiency, next-generation sequencing (NGS) was used to determine the percentage of cells with insertions / deletions (indels).
[0039] [Figure 5]Figure 1 shows editing efficiency at the L-SH5, L-SH18, and L-SH20 genomic loci in HepG2 cells after LNP-mediated delivery of Cas9 mRNA and sgRNA and co-delivery of AAV-DJ containing a firefly luciferase (FLuc) coding sequence driven by a CMV promoter. To assess editing efficiency, next-generation sequencing (NGS) was used to determine the percentage of cells with insertions / deletions (indels).
[0040] [Figure 6] Figure 1 shows FLuc signal in HepG2 cells after LNP-mediated delivery of Cas9 mRNA and sgRNA (targeting L-SH5, L-SH18, or L-SH20) and delivery of AAV-DJ carrying a FLuc coding sequence driven by a CMV promoter. Negative controls included untreated samples, AAV-DJ-only samples (no integration), and samples in which the sgRNA was a non-targeting sgRNA (no integration). After 23 passages, episomal AAV-DJ FLuc is diluted out, and only safe-harbor integrated AAV-DJ is maintained.
[0041] [Figure 7] Figure 1 shows editing efficiency at the L-SH5, L-SH18, and L-SH20 genomic loci in primary human hepatocytes after delivery of 1 μg / mL of LNPs containing AAV-DJ carrying a FLuc coding sequence driven by a CMV promoter and Cas9 mRNA and sgRNA. To assess editing efficiency, next-generation sequencing (NGS) was used to determine the percentage of cells with insertions / deletions (indels).
[0042] [Figure 8]Figure 1 shows FLuc signals in primary human hepatocytes after delivery of 1 μg / mL LNP containing Cas9 mRNA and sgRNA (targeting L-SH5, L-SH18, or L-SH20) and AAV-DJ carrying a FLuc coding sequence driven by a CMV promoter at a multiplicity of infection (MOI) of 10, 10, or 10. Samples in which the sgRNA was a non-targeting sgRNA were used as controls. FLuc signals were assessed 72 hours after delivery of the CRISPR / Cas9 and FLuc nucleic acid constructs.
[0043] [Figure 9] FIG. 1 shows a schematic for testing sgRNAs targeting L-SH5, L-SH18, and L-SH20 for CRIS PR / Cas9-mediated insertion of a CMV-FLuc donor in a humanized liver mouse model.
[0044] [Figure 10] A CMV promoter-driven transgene (FLuc) is shown inserted into human primary hepatocytes using an AAV-DJ vector.
[0045] [Figure 11] FIG. 1 shows a schematic for testing the safety profile of targeting potential safe harbor loci in a humanized liver mouse model.
[0046] [Figure 12] 1 shows the levels of human albumin (hALb) detected by ELISA in serum from immunodeficient FRG mice 25 weeks after engraftment of primary human hepatocytes.
[0047] [Figure 13]Long-term expression of FLuc in a humanized liver mouse model. IVIS imaging was performed to assay for FLuc expression in FRG mice 12 months after engraftment of primary human hepatocytes. Nucleic acid constructs for insertion of the FLuc transgene into potential safe harbor loci L-SH5, L-SH18, and L-SH20 were delivered to primary human hepatocytes using AAV-DJ vectors. Images were reconstructed from IVIS analysis.
[0048] [Figure 14A] This study demonstrates the safety of targeting the safe harbor loci L-SH5, L-SH18, and L-SH20 in a humanized liver mouse model. No obvious dysregulation of liver enzymes was observed in the serum of immunodeficient FRG mice after engraftment of primary human hepatocytes. Liver markers ALT (Figure 14A), AST (Figure 14B), and ALP (Figure 14C) were consistent between treatment groups. Bilirubin levels (Figure 14D) were reduced in the treatment group. Body weight remained consistent between treatment groups (Figure 14E). [Figure 14B] This study demonstrates the safety of targeting the safe harbor loci L-SH5, L-SH18, and L-SH20 in a humanized liver mouse model. No obvious dysregulation of liver enzymes was observed in the serum of immunodeficient FRG mice after engraftment of primary human hepatocytes. Liver markers ALT (Figure 14A), AST (Figure 14B), and ALP (Figure 14C) were consistent between treatment groups. Bilirubin levels (Figure 14D) were reduced in the treatment group. Body weight remained consistent between treatment groups (Figure 14E). [Figure 14C] This study demonstrates the safety of targeting the safe harbor loci L-SH5, L-SH18, and L-SH20 in a humanized liver mouse model. No obvious dysregulation of liver enzymes was observed in the serum of immunodeficient FRG mice after engraftment of primary human hepatocytes. Liver markers ALT (Figure 14A), AST (Figure 14B), and ALP (Figure 14C) were consistent between treatment groups. Bilirubin levels (Figure 14D) were reduced in the treatment group. Body weight remained consistent between treatment groups (Figure 14E). [Figure 14D]This study demonstrates the safety of targeting the safe harbor loci L-SH5, L-SH18, and L-SH20 in a humanized liver mouse model. No obvious dysregulation of liver enzymes was observed in the serum of immunodeficient FRG mice after engraftment of primary human hepatocytes. Liver markers ALT (Figure 14A), AST (Figure 14B), and ALP (Figure 14C) were consistent between treatment groups. Bilirubin levels (Figure 14D) were reduced in the treatment group. Body weight remained consistent between treatment groups (Figure 14E). [Figure 14E] This study demonstrates the safety of targeting the safe harbor loci L-SH5, L-SH18, and L-SH20 in a humanized liver mouse model. No obvious dysregulation of liver enzymes was observed in the serum of immunodeficient FRG mice after engraftment of primary human hepatocytes. Liver markers ALT (Figure 14A), AST (Figure 14B), and ALP (Figure 14C) were consistent between treatment groups. Bilirubin levels (Figure 14D) were reduced in the treatment group. Body weight remained consistent between treatment groups (Figure 14E).
[0049] [Figure 15] Figure 1 shows liver tissue from humanized mice stained for H&E, human FAH, human ASGR1, and Ki67. No significant staining was observed for H&E or Ki67, a marker of proliferation in the liver, suggesting the absence of tumor formation or active oncogenic transformation. Staining for human FAH and human ASGR1 demonstrates the high degree of humanization of the mouse liver.
[0050] [Figure 16] Alignment blocks are shown between the human chromosomal region containing the human safe harbor locus L-SH5 (indicated by an arrow) and the corresponding mouse chromosomal block with the same alignment order.
[0051] [Figure 17] Alignment blocks are shown between the human chromosomal region containing the human safe harbor locus L-SH18 (indicated by an arrow) and the corresponding mouse chromosomal block with the same alignment order.
[0052] [Figure 18] Alignment blocks are shown between the human chromosomal region containing the human safe harbor locus L-SH20 (indicated by an arrow) and the corresponding mouse chromosomal block with the same alignment order. DETAILED DESCRIPTION OF THE INVENTION
[0053] definition The terms "protein," "polypeptide," and "peptide," used interchangeably herein, include polymeric forms of amino acids of any length, including coded and non-coded amino acids, and amino acids that are chemically or biochemically modified or derivatized. These terms also include modified polymers, such as polypeptides having modified peptide backbones. The term "domain" refers to any portion of a protein or polypeptide having a specific function or structure.
[0054] Proteins are said to have an "N-terminus" and a "C-terminus". The term "N-terminus" refers to the beginning of the protein or polypeptide, which ends with an amino acid having a free amine group (-NH2). The term "C-terminus" refers to the end of the amino acid chain (protein or polypeptide) terminated by a free carboxyl group (-COOH).
[0055] The terms "nucleic acid" and "polynucleotide," used interchangeably herein, include polymeric forms of nucleotides of any length, containing ribonucleotides, deoxyribonucleotides, or analogs or modified versions thereof. These include single-, double-, and multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, and polymers that contain purine bases, pyrimidine bases, or other natural, chemically modified, biochemically modified, non-natural, or derivatized nucleotide bases.
[0056] Nucleic acids are said to have a "5' end" and a "3' end" because mononucleotides react to form oligonucleotides in a manner such that the 5' phosphate of one mononucleotide pentose ring is unidirectionally linked to the 3' oxygen of its neighbor via a phosphodiester bond. An end of an oligonucleotide is referred to as the "5' end" if its 5' phosphate is not linked to the 3' oxygen of a mononucleotide pentose ring. An end of an oligonucleotide is referred to as the "3' end" if its 3' oxygen is not linked to the 5' phosphate of another mononucleotide pentose ring. A nucleic acid sequence, even if it is internal to a larger oligonucleotide, can also be said to have 5' and 3' ends. In either a linear or circular DNA molecule, distinct elements are referred to as being "upstream" or "downstream" 5' or 3' elements.
[0057] The term "genomically integrated" refers to a nucleic acid that has been introduced into the genome of a cell such that the nucleotide sequence is integrated into the genome of the cell. Any protocol can be used for stable integration of a nucleic acid into the genome of a cell.
[0058] The term "viral vector" refers to a recombinant nucleic acid that contains at least one element of viral origin and contains elements sufficient for or that allow packaging into a viral vector particle. The vector and / or particle can be used to transfer DNA, RNA, or other nucleic acids into cells in vitro, ex vivo, or in vivo. Many forms of viral vectors are known.
[0059] The term "isolated," with respect to cells, tissues (e.g., liver samples), proteins, and nucleic acids, includes substantially pure preparations of cells, tissues (e.g., liver samples), proteins, and nucleic acids, including cells, tissues (e.g., liver samples), proteins, and nucleic acids that are relatively purified with respect to other bacteria, viruses, cells, or other components that may normally be present in situ. The term "isolated" also includes cells, tissues, proteins, and nucleic acids that have no naturally occurring counterpart and have been chemically synthesized and thereby substantially free from other contaminating cells, tissues (e.g., liver samples), proteins, and nucleic acids, or that have been separated or purified from most other components (e.g., cellular components) that naturally accompany them (e.g., other cellular proteins, polynucleotides, or other components).
[0060] The term "wild-type" includes entities having structure and / or activity as found in a normal state or context (as opposed to mutant, diseased, altered, etc.). Wild-type genes and polypeptides often exist in multiple alternative forms (e.g., alleles).
[0061] The term "endogenous sequence" refers to a nucleic acid sequence that occurs naturally within a cell or animal. For example, an endogenous human Rosa26 sequence refers to the native Rosa26 sequence that occurs naturally at the human Rosa26 locus.
[0062] An "exogenous" molecule or sequence includes a molecule or sequence that is not normally present in a cell in that form. Normal presence includes presence with respect to a particular developmental stage and environmental conditions of the cell. An exogenous molecule or sequence can include, for example, a mutated version of a corresponding endogenous sequence in a cell, such as a humanized version of an endogenous sequence, or can include a sequence that corresponds to an endogenous sequence within the cell but in a different form (i.e., not within a chromosome). In contrast, an endogenous molecule or sequence includes a molecule or sequence that is normally present in that form in a particular cell, at a particular developmental stage, and under particular environmental conditions.
[0063] The term "heterologous" when used in the context of a nucleic acid or a protein indicates that the nucleic acid or protein comprises at least two segments that are not naturally found together in the same molecule. For example, when used with respect to a segment of a nucleic acid or a segment of a protein, the term "heterologous" indicates that the nucleic acid or protein comprises two or more subsequences that are not found in the same relationship to each other (e.g., linked together) in nature. As an example, a "heterologous" region of a nucleic acid vector is a segment of nucleic acid within or attached to another nucleic acid molecule that is not found in association with that other molecule in nature. For example, a heterologous region of a nucleic acid vector can include a coding sequence that is adjacent to sequences not found in association with the coding sequence in nature. Similarly, a "heterologous" region of a protein is a segment of amino acids within or attached to another peptide molecule that is not found in association with that other peptide molecule in nature (e.g., a fusion protein or a tagged protein). Similarly, a nucleic acid or protein can include a heterologous tag or a heterologous secretion or localization sequence.
[0064] "Codon optimization" (i.e., a "codon-optimized" sequence) involves the process of modifying a nucleic acid sequence for enhanced expression in a particular host cell by taking advantage of codon degeneracy, as indicated by the diversity of three-base-pair codon combinations that specify an amino acid, generally by replacing at least one codon of the native sequence with a codon used more frequently or most frequently in the host cell's genes, while maintaining the native amino acid sequence. For example, a nucleic acid encoding a polypeptide of interest can be modified, compared to the naturally occurring nucleic acid sequence, to use alternative codons that are more frequently used in a given prokaryotic or eukaryotic cell, including bacterial cells, yeast cells, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, hamster cells, or any other host cell. Codon usage tables are readily available, for example, in "codon usage databases." These tables can be applied in a variety of ways. See Nakamura et al. (2000) Nucleic Acids Res. 28(1):292, incorporated herein by reference in its entirety for all purposes. Computer algorithms are also available for codon optimization of a particular sequence for expression in a particular host (see, eg, Gene Forge).
[0065] The term "locus" refers to the specific location of a gene (or significant sequence), DNA sequence, polypeptide-coding sequence, or chromosomal position in the genome of an organism. For example, a "Rosa26 locus" can refer to the specific location of the Rosa26 gene, a Rosa26 DNA sequence, or a chromosomal Rosa26 location in the genome of an organism identified as containing such a sequence. A "Rosa26 locus" can include regulatory elements of the Rosa26 gene, including, for example, an enhancer, promoter, 5' and / or 3' untranslated regions (UTRs), or a combination thereof.
[0066] The term "gene" refers to a DNA sequence in a chromosome that, when naturally occurring, may contain at least one coding region and at least one non-coding region. Such a DNA sequence in a chromosome that encodes a product (e.g., without limitation, an RNA product and / or a polypeptide product) may include coding regions interrupted by non-coding introns and sequences located adjacent to the coding region on both the 5' and 3' ends, such that a gene corresponds to a full-length RNA (including 5' and 3' untranslated sequences). In addition, other non-coding sequences, including regulatory sequences (e.g., without limitation, promoters, enhancers, and transcription factor binding sites), polyadenylation signals, internal ribosome entry sites, silencers, insulating sequences, and matrix attachment regions, may also be present in a gene. These sequences may be located adjacent (e.g., without limitation, within 10 kb) or distant from the coding region of a gene, and they affect the level or rate of transcription and translation of the gene.
[0067] The term "allele" refers to a variant form of a gene. Some genes have different forms that are located at the same position, or locus, on a chromosome. Diploid organisms have two alleles at each locus. Each pair of alleles represents a genotype at a particular locus. A genotype is described as homozygous if there are two identical alleles at a particular locus, and heterozygous if the two alleles are different.
[0068] A "promoter" is a regulatory region of DNA that typically contains a TATA box that can direct RNA polymerase II to begin RNA synthesis at the appropriate transcription start site for a particular polynucleotide sequence. A promoter may additionally contain other regions that affect the rate of transcription initiation. The promoter sequences disclosed herein regulate transcription of an operably linked polynucleotide. The promoter may be active in one or more of the cell types disclosed herein (e.g., human cells, human liver cells, or human liver hepatocytes). The promoter may be, for example, a constitutively active promoter, a conditional promoter, an inducible promoter, a temporally restricted promoter (e.g., a developmentally regulated promoter), or a spatially restricted promoter (e.g., a cell-specific or tissue-specific promoter). Examples of promoters can be found, for example, in International Publication No. WO 2013 / 176772, which is incorporated by reference in its entirety for all purposes.
[0069] "Operable linkage" or "operably linked" includes the juxtaposition of two or more components (e.g., a promoter and another sequence element) that allows for the normal function of both components and the potential for at least one of the components to mediate the function of at least one of the other components. For example, a promoter can be operably linked to a coding sequence if the promoter controls the level of transcription of the coding sequence depending on the presence or absence of one or more transcriptional regulatory factors. Operable linkage can include proximity of such sequences to each other or acting in trans (e.g., regulatory sequences can act at a distance to control transcription of the coding sequence).
[0070] The methods and compositions provided herein employ a variety of different components. Some components throughout the description may have active variants and fragments. The term "functional" refers to the inherent ability of a protein or nucleic acid (or a fragment or variant thereof) to exhibit biological activity or function. The biological function of a functional fragment or variant may be the same as compared to the original molecule, or may actually be altered (e.g., with respect to their specificity, selectivity, or efficacy), while retaining the basic biological function of the molecule.
[0071] The term "variant" refers to a nucleotide sequence that differs (e.g., by one nucleotide) from the most common sequence in a population, or a protein sequence that differs (e.g., by one amino acid) from the most common sequence in a population.
[0072] The term "fragment", when referring to a protein, refers to a protein that is shorter or has fewer amino acids than the full-length protein. The term "fragment", when referring to a nucleic acid, refers to a nucleic acid that is shorter or has fewer nucleotides than the full-length nucleic acid. A fragment, for example, when referring to a protein fragment, can be an N-terminal fragment (i.e., removal of a portion of the C-terminus of the protein), a C-terminal fragment (i.e., removal of a portion of the N-terminus of the protein), or an internal fragment (i.e., removal of a portion of each of the N-terminus and C-terminus of the protein). A fragment, for example, when referring to a nucleic acid fragment, can be a 5' fragment (i.e., removal of a portion of the 3' terminus of the nucleic acid), a 3' fragment (i.e., removal of a portion of the 5' terminus of the nucleic acid), or an internal fragment (i.e., removal of a portion of each of the 5' terminus and 3' terminus of the nucleic acid).
[0073] "Sequence identity" or "identity" in the context of two polynucleotide or polypeptide sequences refers to the residues of the two sequences that are the same when aligned for maximum correspondence over a specified comparison window. When using percentage sequence identity with respect to proteins, non-identical residue positions often differ by conservative amino acid substitutions, in which an amino acid residue is replaced with another amino acid residue having similar chemical properties (e.g., charge or hydrophobicity) and thus does not alter the functional properties of the molecule. When sequences differ by conservative substitutions, the percent sequence identity may be adjusted upward to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity." Means for making this adjustment are well known. Typically, this involves scoring conservative substitutions as partial rather than complete mismatches, thereby increasing the percentage sequence identity. Thus, for example, where identical amino acids are given a score of 1 and non-conservative substitutions are given a score of 0, conservative substitutions are given a score of 0 to 1. Scoring of conservative substitutions is calculated, for example, as implemented in the program PC / GENE (Intelligenetics, Mountain View, Calif.).
[0074] "Percentage of sequence identity" includes values determined by comparing two optimally aligned sequences (maximum number of perfectly matched residues) over a comparison window, where the portion of the polynucleotide sequence in the comparison window may contain additions or deletions (i.e., gaps) when compared to a reference sequence (no additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions where the same nucleic acid base or amino acid residue occurs in both sequences to obtain the number of matching positions, dividing the number of matching positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. Unless otherwise specified (e.g., the shorter sequence includes concatenated non-homologous sequences), the comparison window is the entire length of the shorter of the two sequences being compared.
[0075] Unless otherwise specified, sequence identity / similarity values include values obtained using GAP version 10 with the following parameters: % identity and % similarity for nucleotide sequences using a GAP weight of 50, a length weight of 3, and the nwsgapdna.cmp scoring matrix; % identity and % similarity for amino acid sequences using a GAP weight of 8, a length weight of 2, and the BLOSUM62 scoring matrix; or any equivalent program thereof. "Equivalent program" includes any sequence comparison program that produces alignments with identical nucleotide or amino acid residue matches and identical percent sequence identity for any two sequences at issue when compared to corresponding alignments produced by GAP version 10.
[0076] The term "conservative amino acid substitution" refers to the substitution of an amino acid normally present in a sequence with a different amino acid of similar size, charge, or polarity. Examples of conservative substitutions include the substitution of a non-polar (hydrophobic) residue such as isoleucine, valine, or leucine for another non-polar residue. Similarly, examples of conservative substitutions include the substitution of one polar (hydrophilic) residue for another, such as between arginine and lysine, between glutamine and asparagine, or between glycine and serine. Additionally, the substitution of a basic residue such as lysine, arginine, or histidine for another, or the substitution of one acidic residue such as aspartic acid or glutamic acid for another, are further examples of conservative substitutions. Examples of non-conservative substitutions include the substitution of a non-polar (hydrophobic) amino acid residue, e.g., isoleucine, valine, leucine, alanine, or methionine, for a polar (hydrophilic) residue, e.g., cysteine, glutamine, glutamic acid, or lysine, and / or the substitution of a polar residue for a non-polar residue. Exemplary amino acid classes are summarized below.
[0077] [Table 1]
[0078] A "homologous" sequence (e.g., a nucleic acid sequence) includes a sequence that is identical to or substantially similar to a known reference sequence, e.g., at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the known reference sequence. Homologous sequences can include, for example, orthologous and paralogous sequences. For example, homologous genes typically derive from a common ancestral DNA sequence through either a speciation event (orthologous genes) or a gene duplication event (paralogous genes). "Orthologous" genes include genes in different species that evolved from a common ancestral gene through speciation. Orthologs typically retain the same function during evolution. "Paralogous" genes include genes that are related by duplication within a genome. Paralogs can evolve new functions during evolution.
[0079] The term "in vitro" includes an artificial environment and processes or reactions that occur within an artificial environment (e.g., a test tube or an isolated cell or cell line). The term "in vivo" includes a natural environment (e.g., a cell or organism or body) and processes or reactions that occur within a natural environment. The term "ex vivo" includes cells removed from an individual's body and processes or reactions that occur within such cells.
[0080] A composition or method "comprising" or "including" one or more recited elements may include other elements not specifically recited. For example, a composition "comprises" or "includes" a protein may contain the protein alone or in combination with other components. The transitional phrase "consisting essentially of" means that the scope of the claim shall be construed to include the specific elements recited in the claim as well as elements that do not materially affect the basic and novel characteristics of the claimed invention. Thus, the term "consisting essentially of," when used in the claims of the present invention, is not intended to be construed as equivalent to "comprising."
[0081] "Optional" or "optionally" means that the subsequently described event or circumstance may or may not occur, and that the description includes examples where the event or circumstance occurs and examples where it does not occur.
[0082] Designation of a range of values includes all integers within or defining that range, and all subranges defined by integers within that range, e.g., 5-10 nucleotides is understood to mean 5, 6, 7, 8, 9, or 10 nucleotides, while 5-10% is understood to include all possible values between 5% and 10%.
[0083] At least 17 nucleotides of a 20 nucleotide sequence is understood to include 17, 18, 19, or 20 nucleotides of the provided sequence, thereby providing an upper limit even if one is not specifically provided, as is clearly understood. Similarly, up to 3 nucleotides is understood to encompass 0, 1, 2, or 3 nucleotides, providing a lower limit even if one is not specifically provided. When "at least," "up to," or other similar language modifies a number, it can be understood to modify each number in the series.
[0084] As used herein, "less than" or "less than" is understood as the value adjacent to the phrase and any logically lower value or integer logically from the context, up to 0. For example, a duplex region of "two or fewer nucleotide base pairs" has 2, 1, or 0 nucleotide base pairs. When "less than" or "less than" appears before a series of numbers or ranges, it is understood that each of the numbers in the series or range is modified.
[0085] As used herein, when a maximum value is expressed as 100% (e.g., 100% inhibition), it is understood that the value is limited by the detection method. For example, 100% inhibition is understood as inhibition relative to a level below the detection level of the assay.
[0086] Unless otherwise clear from the context, the term "about" includes values that are ±5% of the stated value. In certain embodiments, the term "about" is understood to include variations or errors accepted within the art, such as two standard deviations from the mean, or the sensitivity of the method used to make the measurement, or percentages of values accepted in the art, such as those with age. When "about" is present before the first value in a series, it can be understood to modify each value in the series.
[0087] The term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative ("or").
[0088] The term "or" refers to any one member of a particular list and also includes any combination of members of that list.
[0089] The singular articles "a," "an," and "the" include plural references unless the context clearly dictates otherwise. For example, the term "protein" or "at least one protein" can include a plurality of proteins, including mixtures thereof.
[0090] Statistically significant means p≦0.05.
[0091] In the event of a discrepancy between the sequence in this application and the indicated accession number or position in the accession number, the sequence in this application will control.
[0092] I. Overview Current gene therapy approaches rely on episomal expression of transgenes and / or insertion into specific genomic loci. Episomal methods have proven to be liver-restricted due to dilution or silencing. Integration at specific loci allows for sustained transgene expression. However, this approach remains to be proven effective and safe in human settings. Canonical genomic safe harbor loci in humans, such as AAVS1, CCR5, and Rosa26, are all intragenic and are less well studied than mouse genomic safe harbor loci. In addition, because different tissues have different chromatin states for a given locus, canonical genomic safe harbor loci may be silenced in some tissues. All of the canonical genomic safe harbor loci in humans have additional drawbacks. Methylation mechanisms can silence transgenes at the AAVS1 locus in some cell lines, knockout of CCR5 can result in increased susceptibility to infection with West Nile virus and Japanese encephalitis, and the human Rosa26 locus is less well studied than its mouse ortholog. Therefore, there is a need for tissue-specific genomic safe harbor loci.
[0093] Compositions and methods are provided for inserting a nucleic acid encoding a product of interest into a genomic safe harbor locus in a cell, a population of cells, or a subject (e.g., a subject in need thereof), or for expressing a nucleic acid encoding a product of interest from a genomic safe harbor locus in a cell, a population of cells, or a subject (e.g., a subject in need thereof). Also provided are cells or populations of cells, or subjects, comprising a nucleic acid construct comprising a coding sequence for a product of interest inserted into a genomic safe harbor locus. Also provided herein are methods for identifying genomic safe harbor loci (e.g., extragenic genomic safe harbor loci) for use in specific cell or tissue types.
[0094] II. Compositions for inserting nucleic acid constructs into genomic safe harbor loci and for expressing products of interest from genomic safe harbor loci in cells and subjects Provided herein are nucleic acid constructs and compositions that enable insertion of a coding sequence for a product of interest into a genomic safe harbor locus and / or expression of the coding sequence for a product of interest from a genomic safe harbor locus. The nucleic acid constructs and compositions can be used in methods for integration into and / or expression from a genomic safe harbor locus in a cell or subject. Nuclease agents (e.g., targeted to a genomic safe harbor locus) or nucleic acids encoding the nuclease agents are also provided to facilitate integration of a nucleic acid construct into a genomic safe harbor locus. Nuclease agents that target a genomic safe harbor locus or a nucleic acid encoding a nuclease agent are also provided to facilitate integration of a nucleic acid construct into a genomic safe harbor locus.
[0095] A. Genomic Safe Harbor Loci Method for Identifying Genomic Safe Harbor Loci Interactions between the integrated exogenous DNA and the host genome can limit the reliability and safety of integration, and can result in obvious phenotypic effects not due to targeted gene modification, but instead due to the unintended effects of integration on surrounding endogenous genes. For example, randomly inserted transgenes are susceptible to position effects and silencing, which can make their expression unreliable and unpredictable. Similarly, integration of exogenous DNA into a chromosomal locus can affect surrounding endogenous genes and chromatin, thereby changing the behavior and phenotype of the cell.
[0096] As used herein, a target genomic locus may be a genomic safe harbor locus. A genomic safe harbor locus includes a chromosomal locus at which a transgene or other exogenous nucleic acid insert can be stably and reliably expressed in a tissue of interest without overtly altering cellular behavior or phenotype (i.e., without adversely affecting the host cell). For example, a genomic safe harbor locus may be a locus in which expression of an inserted gene sequence is not perturbed by read-through expression from adjacent genes. For example, a genomic safe harbor locus may include a chromosomal locus into which exogenous DNA can integrate and function in a predictable manner without adversely affecting the structure or expression of endogenous genes. Genomic safe harbor loci can be targeted with high efficiency, and safe harbor loci can be disrupted without any obvious phenotype. A genomic safe harbor locus can include extragenic or intragenic regions, e.g., intragenic loci that are nonessential, unnecessary, or can be disrupted without any obvious phenotypic consequences.
[0097] The genomic safe harbor loci described herein can be genomic loci that, when targeted for integration in a subject, do not alter liver function. The genomic safe harbor loci described herein can be genomic loci that, when targeted for integration in a subject, do not alter alanine aminotransferase (alanine transaminase or ALT) levels. The genomic safe harbor loci described herein can be genomic loci that, when targeted for integration in a subject, do not alter aspartate aminotransferase (AST) levels. The genomic safe harbor loci described herein can be genomic loci that, when targeted for integration in a subject, do not alter alkaline phosphatase (ALP) levels. The genomic safe harbor loci described herein can be genomic loci that, when targeted for integration in a subject, do not alter body weight. The genomic safe harbor loci described herein can be genomic loci that, when targeted for integration in a subject, do not alter proliferation (e.g., as assessed by Ki67 staining) in a target organ, such as the liver, etc. A genomic safe harbor locus as described herein can be a genomic locus that, when targeted for integration in a subject, does not cause oncogenic transformation in a target organ, such as the liver (e.g., as assessed by H&E staining).
[0098] A genomic safe harbor locus as described herein can be a genomic locus that has an open chromatin configuration in the liver such that an exogenous nucleic acid insert can be stably and reliably expressed in the liver. Alternatively, a genomic safe harbor locus can be a genomic locus that has an open chromatin configuration in another tissue or cell type (e.g., a hematopoietic cell such as a hematopoietic stem cell, T cell, B cell, and / or macrophage) such that an exogenous nucleic acid insert can be stably and reliably expressed in that tissue or cell type.
[0099] The genomic safe harbor loci described herein may be extragenic genomic safe harbor loci (i.e., located outside of genes). In a specific example, the genomic safe harbor loci described herein are extragenic genomic safe harbor loci with an open chromatin configuration in the liver.
[0100] In specific examples, genomic safe harbor loci can be more than 300 kb from any cancer-associated gene (e.g., to prevent insertional oncogenesis), more than 300 kb from any miRNA or small RNA (e.g., to preserve regulation of gene expression and cell development), more than 50 kb from the 5' end of any gene (e.g., to avoid perturbing endogenous gene expression), more than 50 kb from any origin of replication, more than 50 kb from any ultraconserved element (e.g., non-coding intragenic or intergenic regions that are completely conserved in the human, mouse, and rat genomes), outside of regions of copy number variation, and open chromatin (e.g., as determined by ATAC-Seq analysis (e.g., in human liver biopsy samples)). Additionally, genomic safe harbor loci may not overlap with regulatory regions (e.g., H3K4me1, H3K27ac, and / or H3K4me3 markers), heterochromatic regions (e.g., H3K9me3 markers), or regions predicted to be involved in chromatin organization (e.g., CTCF signaling).
[0101] For example, a method for identifying genomic safe harbor loci (e.g., extragenic genomic safe harbor loci) may include: (a) identifying accessible genomic loci (i.e., chromatin sites) in a tissue or cell type of interest (e.g., based on an ATAC-Seq dataset); (b) eliminating the genomic loci identified in step (a) based on safety criteria, functional silencing criteria, and / or structural accessibility criteria; and (c) eliminating the loci identified in step (b) based on gRNA availability, efficacy (editing efficiency), and specificity (off-target analysis). Such a method may further include analyzing the chromatin environment for chromatin marks to disqualify any potential safe harbors contained in regulatory regions (e.g., H3K4me1, H3K27ac, and / or H3K4me3), heterochromatic regions (e.g., H3K9me3), or regions predicted to be involved in chromatin 3D organization (e.g., CTCF signaling).
[0102] Eukaryotic chromatin is tightly packaged into arrays of nucleosomes, each consisting of a histone octamer core wrapped around DNA and separated by linker DNA. The nucleosome core is composed of histone proteins, which can be post-translationally modified by covalent modifications or replaced by histone variants. The positioning of nucleosomes throughout the genome has important regulatory functions by altering the in vivo availability of binding sites for transcription factors and the basal transcription machinery, thus affecting DNA-dependent processes such as transcription, DNA repair, replication, and recombination. Accessible genomic loci are regions of open chromatin. Open chromatin regions are nucleosome-depleted regions that can be bound by protein factors and play various roles in DNA replication, nuclear organization, and gene transcription. Step (a) can include identifying accessible genomic loci using an assay for transposase-accessible chromatin, such as ATAC-Seq analysis. ATAC-Seq represents an assay for transposase-accessible chromatin by high-throughput sequencing. See, for example, Buenrostro et al. (2013) Nat. Methods 10(12):1213-1218 and Buenrostro et al. (2015) Curr. Protoc. Mol. Biol. 109:21.29.1-21.29.9, each of which is incorporated herein by reference in its entirety for all purposes. The ATAC-Seq method relies on next-generation sequencing (NGS) library construction using the hyperactive transposase Tn5. NGS adapters are loaded onto the transposase, which simultaneously fragments chromatin and integrates these adapters into open chromatin regions. The resulting libraries can be sequenced by NGS, and regions of the genome with open or accessible chromatin are analyzed using bioinformatics. As a first step, cells are harvested. After harvesting, the cells are lysed with a non-ionic detergent to obtain pure nuclei.The resulting chromatin is then fragmented and simultaneously tagged with sequencing adapters using Tn5 transposase to generate an ATAC-Seq library. After purification, the library can be amplified by PCR using barcoded primers. The resulting library can then be analyzed by qPCR or next-generation sequencing. ATAC-seq identifies accessible ATAC regions by probing open chromatin with a hyperactive mutant Tn5 transposase, which inserts sequencing adapters into open regions of the genome. While naturally occurring transposases have low levels of activity, ATAC-seq uses a mutated, hyperactive transposase. In a process called tagmentation, Tn5 transposase cleaves double-stranded DNA and tags it with sequencing adapters. The tagged DNA fragments are then purified, PCR-amplified, and sequenced using next-generation sequencing. Sequencing reads can then be used to infer regions of increased accessibility and to map regions of transcription factor binding sites and nucleosome positioning. The number of reads in a region correlates with how open the chromatin is at single nucleotide resolution.
[0103] Step (a) can also include identifying accessible genomic loci using, for example, DNase I hypersensitive site sequencing (DNase-Seq). DNase-Seq is a method used to identify the location of regulatory regions based on genome-wide sequencing of regions sensitive to DNase I cleavage. This method utilizes DNase I to selectively digest nucleosome-depleted DNA, whereas DNA regions tightly surrounded by nucleosomes and higher-order structures are more resistant. High-throughput methods identify DNase I hypersensitive sites across the entire genome by capturing DNase-digested fragments and sequencing them using high-throughput next-generation sequencing.
[0104] In step (b), the safety criterion may include selecting a genomic locus only if it is more than 300 kb from any cancer-associated gene (e.g., to prevent insertional carcinogenesis), more than 300 kb from any miRNA or small RNA (e.g., to preserve the regulation of gene expression and cell development), and / or more than 50 kb from the 5' end of any gene (e.g., to avoid perturbing endogenous gene expression). The functional silencing criterion may include selecting a genomic locus only if it is more than 50 kb from any origin of replication and / or more than 50 kb from any ultraconserved element (e.g., a non-coding intragenic or intergenic region that is completely conserved in the human, mouse, and rat genomes). The structural accessibility criterion may include selecting a genomic locus only if it is not in a copy number variable region.
[0105] In step (c), loci can be filtered based on gRNA availability, efficacy (editing efficiency), and specificity (off-target analysis). gRNA availability refers to the presence of a suitable target sequence for the guide RNA, taking into account PAM requirements. Efficacy refers to the editing efficiency of the gRNA in the tissue or cell type of interest. Any appropriate threshold for editing efficiency can be set. For example, a locus or gRNA can be selected if the editing efficiency is at least about 10%, at least about 11%, at least about 12%, at least about 13%, at least about 14%, at least about 15%, at least about 16%, at least about 17%, at least about 18%, at least about 19%, or at least about 20%. In one example, gRNA efficacy is measured in primary cells (e.g., primary hepatocytes). In another example, gRNA efficacy is measured in the tissue of interest in vivo. In a specific example, gRNA efficacy is measured in primary cells from multiple different donors (e.g., primary hepatocytes from multiple different donors, such as two or three different donors). Any appropriate threshold for gRNA specificity can be used. For example, a guide RNA can be selected if there are no other sequences in the genome that perfectly match the guide RNA target sequence or have only one mismatch. In another example, a guide RNA can be selected if there are no other sequences in the genome that perfectly match the guide RNA target sequence or have only one or two mismatches.
[0106] Such methods may include analyzing the chromatin environment for markers (e.g., signals or chromatin marks) and disqualifying any potential safe harbors contained within regions predicted to be regulatory regions (e.g., H3K4me1, H3K27ac, and / or H3K4me3), heterochromatic regions (e.g., H3K9me3), chromatin organization (e.g., CTCF signaling), or transcriptionally active regions (e.g., H3K36me3, PolR2A, RNASeq-, and RNASeq+). For example, ChIP-Seq data on transcription factor binding, genome-wide DNA methylation, promoter / enhancer signatures inferred by histone marks, and chromatin accessibility can be used. Post-translational modifications on histone tails closely correlate with transcriptional state. For example, trimethylation of histone H3 lysine 4 (H3K4me3) indicates active gene promoters. Monomethylation of histone 3 on lysine 4 (H3K4me1) is a mark linked to enhancers. Identifying regions enriched in H3K4me1 and depleted in H3K4me3, or regions enriched in both H3K4me1 and H3K27ac, has proven to be a viable method for discovering enhancers. H3K27ac is an activation mark that distinguishes active from primed enhancers. H3K9me3 indicates regions undergoing long-term repression. CTCF's primary role is thought to be in regulating the 3D structure of chromatin. CTCF binds strands of DNA together, thus forming chromatin loops and anchoring DNA to cellular structures such as the nuclear lamina. It also defines the boundary between active and heterochromatic DNA. Because the 3D structure of DNA influences gene regulation, CTCF activity influences gene expression. CTCF is thought to be a key component of insulator activity, a sequence that blocks interactions between enhancers and promoters. CTCF binding has also been shown to promote and repress gene expression. It is unknown whether CTCF affects gene expression solely through its looping activity, or whether it has some other, unknown activity.H3K36me3 marks the gene body and experimentally indicates the absence of interfering transcription units. PolR2A indicates transcriptional activity and is used to indicate the absence of transcripts originating from this region. RNASeq- indicates transcriptional activity on the negative strand of DNA, and RNASeq+ indicates transcriptional activity on the positive strand of DNA; both are used to indicate the absence of transcription from that region. RNA-Seq (RNA sequencing) is a sequencing technique that uses next-generation sequencing (NGS) to reveal the presence and amount of RNA in a biological sample. mRNA is extracted from the sample, fragmented, and copied into stable ds-cDNA. The ds-cDNA is sequenced using high-throughput short-read sequencing. These sequences can then be aligned with a reference genome sequence to reconstruct which genomic regions are transcribed.
[0107] In some embodiments, integration of a nucleic acid construct into a genomic safe harbor locus described herein does not cause liver toxicity. In some embodiments, integration of a nucleic acid construct into a genomic safe harbor locus described herein does not alter expression in adjacent genes. In some embodiments, integration of a nucleic acid construct into a genomic safe harbor locus described herein does not cause liver toxicity and does not alter expression in adjacent genes.
[0108] In specific examples, genomic safe harbor loci are located at the following genomic locations: (i) human chromosome 13, coordinates 77460242 to 77460537 (referred to herein as L-SH5), or the corresponding region (e.g., orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse; (ii) human chromosome 6, coordinates 170031084 to 170031382 (referred to herein as L-SH18); is selected from (iii) a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse, and (iv) human chromosome 9, coordinates 25207412-25207703 (referred to herein as L-SH20), or a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse. Throughout this application, genomic coordinates referenced are based on the genome annotation in the GRCh38 (also referred to as hg38) assembly of the human genome from the Genome Reference Consortium, available at the National Center for Biotechnology Information website. Exemplary sequences of L-SH5, L-SH18, and L-SH20, based on genome annotations in the GRCh38 (also referred to as hg38) assembly of the human genome from the Genome Reference Consortium, are set forth in SEQ ID NOs: 39, 40, and 41, respectively.Tools and methods for converting genomic coordinates between one assembly and another are known in the art and can be used to convert the genomic coordinates provided herein to corresponding coordinates in another assembly of the human genome, such as to a previous assembly produced by the same organization or using the same algorithm (e.g., GRCh38 to GRCh37), and to an assembly produced by a different organization or algorithm (e.g., GRCh38 produced by the International Human Genome Sequencing Consortium to NCBI33). Available methods and tools known in the art include, but are not limited to, the NCBI Genome Remapping Service available at the National Center for Biotechnology Information website, UCSC LiftOver available at the UCSC Genome Brower website, and Assembly Converter available at the Ensembl.org website.
[0109] In specific examples, genomic safe harbor loci are located at the following genomic coordinates: (i) from about 77460242 to about 77460537 on human chromosome 13 (corresponding to L-SH5) or a corresponding region (e.g., an orthologous region or a syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse; (ii) from about 170031084 to about 170031382 on human chromosome 6 (corresponding to L-SH18); and (iii) about 25207412 to about 25207703 of human chromosome 9 (corresponding to L-SH20), or the corresponding region (e.g., orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., non-human primate), or rodent such as a rat or mouse. The term "about" when referring to genomic coordinates means ±20 base pairs. In other examples, genomic safe harbor loci are near the region identified by the above coordinates. The term "near" when referring to genomic coordinates means ±5 kb, ±4 kb, ±3 kb, ±2 kb, ±1 kb, ±0.5 kb, ±0.4 kb, ±0.3 kb, ±0.2 kb, or ±0.1 kb.
[0110] In one specific example, the genomic safe harbor locus is human L-SH5 (chromosome 13, coordinates 77460242-77460537) or a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse. For example, the genomic safe harbor locus can include the sequence set forth in SEQ ID NO: 39 or a variant thereof located at the same position or locus in a chromosome of a human, or an orthologous or syntenic region in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse. Syntenic regions are derived from a single ancestral genomic region. For example, syntenic regions can be derived from different organisms and result from speciation.
[0111] In another specific example, the genomic safe harbor locus is human L-SH18 (chromosome 6, coordinates 170031084-170031382), or the corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse. For example, the genomic safe harbor locus can include the sequence set forth in SEQ ID NO: 40, or a variant thereof, located at the same position or locus in a chromosome of a human, or an orthologous or syntenic region in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse.
[0112] In another specific example, the genomic safe harbor locus is human L-SH20 (chromosome 9, coordinates 25207412-25207703), or the corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse. For example, the genomic safe harbor locus can include the sequence set forth in SEQ ID NO: 41, or a variant thereof, located at the same position or locus in a chromosome of a human, or an orthologous or syntenic region in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse.
[0113] In one specific example, the genomic safe harbor locus corresponds to human L-SH5 (coordinates from about 77460242 to about 77460537 on chromosome 13), or a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse, or a variant thereof located at the same position or locus in a human chromosome, or an orthologous or syntenic region in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse. The term "about," when referring to genomic coordinates, means ±20 base pairs. In other examples, the genomic safe harbor locus is near the region identified by the above coordinates. The term "nearby" when referring to genomic coordinates means ±5 kb, ±4 kb, ±3 kb, ±2 kb, ±1 kb, ±0.5 kb, ±0.4 kb, ±0.3 kb, ±0.2 kb, or ±0.1 kb.
[0114] In another specific example, the genomic safe harbor locus corresponds to human L-SH18 (coordinates from about 170031084 to about 170031382 on chromosome 6), or a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse, or a variant thereof located at the same position or locus in a human chromosome, or an orthologous or syntenic region in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse. The term "about," when referring to genomic coordinates, means ±20 base pairs. In other examples, the genomic safe harbor locus is near the region identified by the above coordinates. The term "nearby" when referring to genomic coordinates means ±5 kb, ±4 kb, ±3 kb, ±2 kb, ±1 kb, ±0.5 kb, ±0.4 kb, ±0.3 kb, ±0.2 kb, or ±0.1 kb.
[0115] In another specific example, the genomic safe harbor locus corresponds to human L-SH20 (coordinates of about 25207412 to about 25207703 on chromosome 9), or a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse, or a variant thereof located at the same position or locus in a human chromosome, or an orthologous or syntenic region in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse. The term "about," when referring to genomic coordinates, means ±20 base pairs. In other examples, the genomic safe harbor locus is near the region identified by the above coordinates. The term "nearby" when referring to genomic coordinates means ±5 kb, ±4 kb, ±3 kb, ±2 kb, ±1 kb, ±0.5 kb, ±0.4 kb, ±0.3 kb, ±0.2 kb, or ±0.1 kb.
[0116] In specific examples, genomic safe harbor loci are located at the following genomic locations: (i) mouse chromosome 14, coordinates 103,450,397-103,451,396 (referred to herein as mouse L-SH5), or the corresponding region (e.g., orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., non-human primate), or rodent, such as a rat; (ii) mouse chromosome 17, coordinates 15,226,387-15,227,386 (referred to herein as mouse L-SH18). and (iii) mouse chromosome 4, coordinates 92,827,563-92,828,592 (referred to herein as mouse L-SH20), or the corresponding region (e.g., orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., non-human primate), or rodent such as rat. Throughout this application, genomic coordinates referenced are based on the genome annotation in the GRCm38 (also referred to as mm10) assembly of the mouse genome from the Genome Reference Consortium, available at the National Center for Biotechnology Information website. Exemplary sequences of L-SH5, L-SH18, and L-SH20, based on genome annotations in the GRCm38 (also referred to as mm10) assembly of the mouse genome from the Genome Reference Consortium, are set forth in SEQ ID NOs: 405, 406, and 407, respectively. Tools and methods for converting genome coordinates between one assembly and another are known in the art and can be used to convert the genome coordinates provided herein to corresponding coordinates in another assembly of the mouse genome, including to a previous assembly produced by the same institution or using the same algorithm, and to an assembly produced by a different institution or algorithm.Available methods and tools known in the art include, but are not limited to, the NCBI Genome Remapping Service available at the National Center for Biotechnology Information website, UCSC LiftOver available at the UCSC Genome Brower website, and Assembly Converter available at the Ensembl.org website.
[0117] In specific examples, genomic safe harbor loci are located at the following genomic coordinates: (i) mouse chromosome 14 from about 103,450,397 to about 103,451,396 (corresponding to mouse L-SH5) or the corresponding region (e.g., an orthologous region or a syntenic region) in a non-human animal, a non-human mammal (e.g., a non-human primate), or a rodent such as a rat; (ii) mouse chromosome 17 from about 15,226,387 to about 15,227,386 (corresponding to mouse L-SH18); and (iii) from about 92,827,563 to about 92,828,592 of mouse chromosome 4 (corresponding to mouse L-SH20) or the corresponding region (e.g., orthologous region or syntenic region) in a non-human animal, non-human mammal (e.g., non-human primate), or rodent such as a rat. The term "about" when referring to genomic coordinates means ±20 base pairs. In other examples, genomic safe harbor loci are near the region identified by the above coordinates. The term "near" when referring to genomic coordinates means ±5 kb, ±4 kb, ±3 kb, ±2 kb, ±1 kb, ±0.5 kb, ±0.4 kb, ±0.3 kb, ±0.2 kb, or ±0.1 kb.
[0118] In one specific example, the genomic safe harbor locus is mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) or a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat. For example, the genomic safe harbor locus can include the sequence set forth in SEQ ID NO: 405 or a variant thereof located at the same position or locus in a mouse chromosome or an orthologous or syntenic region in a human or non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat. Syntenic regions are derived from a single ancestral genomic region. For example, syntenic regions can be derived from different organisms and result from speciation.
[0119] In another specific example, the genomic safe harbor locus is mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386), or the corresponding region (e.g., an orthologous region or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat. For example, the genomic safe harbor locus can include the sequence set forth in SEQ ID NO: 406, or a variant thereof, located at the same position or locus in a chromosome of a mouse, or an orthologous region or syntenic region in a human or non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat.
[0120] In another specific example, the genomic safe harbor locus is mouse L-SH20 (chromosome 4, coordinates 92,827,563-92,828,592), or the corresponding region (e.g., an orthologous region or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat. For example, the genomic safe harbor locus can include the sequence set forth in SEQ ID NO: 407, or a variant thereof, located at the same position or locus in a chromosome of a mouse, or an orthologous region or syntenic region in a human or non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat.
[0121] In one specific example, the genomic safe harbor locus corresponds to mouse L-SH5 (coordinates of about 103,450,397 to about 103,451,396 on mouse chromosome 14), or a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat, or a variant thereof located at the same position or locus in a chromosome of a mouse, or an orthologous or syntenic region in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat. The term "about," when referring to genomic coordinates, means ±20 base pairs. In other examples, the genomic safe harbor locus is near the region identified by the above coordinates. The term "nearby" when referring to genomic coordinates means ±5 kb, ±4 kb, ±3 kb, ±2 kb, ±1 kb, ±0.5 kb, ±0.4 kb, ±0.3 kb, ±0.2 kb, or ±0.1 kb.
[0122] In another specific example, the genomic safe harbor locus corresponds to mouse L-SH18 (coordinates from about 15,226,387 to about 15,227,386 on mouse chromosome 17), or a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat, or a variant thereof located at the same position or locus in a chromosome of a mouse, or an orthologous or syntenic region in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat. The term "about," when referring to genomic coordinates, means ±20 base pairs. In other examples, the genomic safe harbor locus is near the region identified by the above coordinates. The term "nearby" when referring to genomic coordinates means ±5 kb, ±4 kb, ±3 kb, ±2 kb, ±1 kb, ±0.5 kb, ±0.4 kb, ±0.3 kb, ±0.2 kb, or ±0.1 kb.
[0123] In another specific example, the genomic safe harbor locus corresponds to mouse L-SH20 (coordinates from about 92,827,563 to about 92,828,592 on mouse chromosome 4), or a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat, or a variant thereof located at the same position or locus in a chromosome of a mouse, or an orthologous or syntenic region in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat. The term "about," when referring to genomic coordinates, means ±20 base pairs. In other examples, the genomic safe harbor locus is near the region identified by the above coordinates. The term "nearby" when referring to genomic coordinates means ±5 kb, ±4 kb, ±3 kb, ±2 kb, ±1 kb, ±0.5 kb, ±0.4 kb, ±0.3 kb, ±0.2 kb, or ±0.1 kb.
[0124] B. Nucleic Acid Constructs Encoding Products of Interest The compositions and methods described herein involve the use of nucleic acid constructs comprising a coding sequence for a product of interest (e.g., a polypeptide of interest) operably linked to a promoter. Such nucleic acid constructs can be for insertion into a target genomic locus (e.g., a genomic safe harbor locus described elsewhere herein) or into a cleavage site created by a nuclease agent or CRISPR / Cas system disclosed elsewhere herein. The term cleavage site includes a DNA sequence at which a nick or double-stranded break is created by a nuclease agent (e.g., a Cas9 protein complexed with a guide RNA). In some embodiments, the double-stranded break is created by a Cas9 protein complexed with a guide RNA, e.g., an SpCas9 protein complexed with an SpCas9 guide RNA.
[0125] The length of the nucleic acid constructs disclosed herein can vary. The constructs can be, for example, about 1 kb to about 5 kb, such as about 1 kb to about 4.5 kb or about 1 kb to about 4 kb. Exemplary nucleic acid constructs are about 1 kb to about 5 kb in length, or about 1 kb to about 4 kb in length. Alternatively, the nucleic acid constructs can be about 1 kb to about 1.5 kb, about 1.5 kb to about 2 kb, about 2 kb to about 2.5 kb, about 2.5 kb to about 3 kb, about 3 kb to about 3.5 kb, about 3.5 kb to about 4 kb, about 4 kb to about 4.5 kb, or about 4.5 kb to about 5 kb in length. Alternatively, the nucleic acid constructs can be, for example, 5 kb or less, 4.5 kb or less, 4 kb or less, 3.5 kb or less, 3 kb or less, or 2.5 kb or less in length.
[0126] The construct may comprise deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), may be single-stranded, double-stranded, or partially single-stranded and partially double-stranded, and may be introduced into a host cell in linear or circular (e.g., minicircle) form. See, e.g., U.S. Patent Application Publication Nos. 2010 / 0047805, 2011 / 0281361, and 2011 / 0207221, each of which is incorporated by reference in its entirety for all purposes. If introduced in linear form, the ends of the construct may be protected (e.g., from exonuclease degradation) by known methods. For example, one or more dideoxynucleotide residues may be added to the 3' end of the linear molecule, and / or self-complementary oligonucleotides may be attached to one or both ends. See, for example, Chang et al. (1987) Proc. Natl. Acad. Sci. USA 84:4959-4963 and Nehls et al. (1996) Science 272:886-889, each of which is incorporated herein by reference in its entirety for all purposes. Additional methods for protecting exogenous polynucleotides from degradation include, but are not limited to, the addition of terminal amino group(s) and the use of modified internucleotide linkages such as phosphorothioates, phosphoramidates, and O-methylribose or deoxyribose residues. Constructs can be introduced into cells as part of vector molecules containing additional sequences, such as, for example, origins of replication, promoters, and genes encoding antibiotic resistance. Constructs may also omit viral elements. Furthermore, the constructs can be introduced as naked nucleic acid, as nucleic acid complexed with agents such as liposomes, polymers, or poloxamers, or can be delivered by viral vectors (e.g., adenovirus, adeno-associated virus (AAV), herpesvirus, retrovirus, lentivirus).
[0127] The constructs disclosed herein can be modified at either or both termini, as needed, to include one or more suitable structural features and / or to confer one or more functional benefits. For example, structural modifications can vary depending on the method used to deliver the constructs disclosed herein to host cells (e.g., using viral vector delivery or packaging in lipid nanoparticles for delivery). Such modifications include, for example, terminal structures such as inverted terminal repeats (ITRs), hairpins, loops, and other structures such as toroids. For example, the constructs disclosed herein can include one, two, or three ITRs, or can include two or fewer ITRs. Various methods of structural modification are known.
[0128] The construct includes a promoter and / or enhancer that drives expression of the product of interest, such as a constitutive promoter or an inducible or tissue-specific (e.g., liver-specific) promoter that drives expression of the product of interest in an episome or upon integration. Non-limiting exemplary constitutive promoters include the cytomegalovirus (CMV) immediate-early promoter, the simian virus (SV40) promoter, the adenovirus major late (MLP) promoter, the Rous sarcoma virus (RSV) promoter, the mouse mammary tumor virus (MMTV) promoter, the phosphoglycerate kinase (PGK) promoter, the elongation factor-alpha (EF1a) promoter, the ubiquitin promoter, the actin promoter, the tubulin promoter, the immunoglobulin promoter, functional fragments thereof, or any combination of the foregoing. For example, the promoter can be a CMV promoter or a truncated CMV promoter. In another example, the promoter can be an EF1a promoter. Liver-suitable promoters include, for example, the albumin (ALB) promoter or the transthyretin (TTR) promoter. Liver-suitable enhancers include, for example, the SERPINA1 enhancer. Non-limiting exemplary inducible promoters include promoters that are inducible by heat shock, light, chemicals, peptides, metals, steroids, antibiotics, or alcohol. Inducible promoters may also have a low basal (uninduced) expression level, such as the Tet-On® promoter (Clontech).
[0129] In some examples, the nucleic acid construct functions for homology-independent insertion of a nucleic acid encoding a product of interest (e.g., a polypeptide of interest). Such nucleic acid constructs can function, for example, in non-dividing cells (e.g., cells in which non-homologous end joining (NHEJ), rather than homologous recombination (HR), is the primary mechanism by which double-stranded DNA breaks are repaired) or dividing cells (cells that are actively dividing). Such constructs can be, for example, homology-independent donor constructs. In preferred embodiments, the promoter and other regulatory sequences are suitable for use in humans, e.g., recognized by regulatory elements in human cells (e.g., human liver cells) and acceptable to regulatory authorities for use in humans.
[0130] The constructs disclosed herein can be modified to include or exclude any suitable structural features as needed for any particular use and / or to confer one or more desired functions. For example, some constructs disclosed herein do not include homology arms. Some constructs disclosed herein are capable of insertion into a target genomic locus or cleavage site in a target DNA sequence of a nuclease agent by non-homologous end joining (e.g., insertion into a genomic safe harbor locus). For example, such constructs can be inserted into a blunt-ended double-stranded break following cleavage by a nuclease agent disclosed herein (e.g., a CRISPR / Cas system, e.g., a SpyCas9 CRISPR / Cas system). In a specific example, the construct can be delivered via AAV and can be capable of insertion by non-homologous end joining (e.g., the construct can be free of homology arms).
[0131] In certain instances, the construct can be inserted via homology-independent targeted integration. For example, the nucleic acid construct or construct's target product coding sequence (e.g., target polypeptide coding sequence) and promoter can be flanked on each side by nuclease agent target sites (e.g., the same target sites in the target DNA sequence for targeted insertion (e.g., in a genomic safe harbor locus) and the same nuclease agent used to cleave the target DNA sequence for targeted insertion). The nuclease agent can then cleave the adjacent target sites. In a specific example, the construct is delivered by AAV-mediated delivery, and cleavage of the adjacent target sites can remove the inverted terminal repeats (ITRs) of the AAV. In some instances, the target DNA sequence for targeted insertion (e.g., a target DNA sequence in a safe harbor locus, such as a gRNA target sequence including flanking protospacer adjacent motifs) is no longer present when the product coding sequence of interest (e.g., a polypeptide coding sequence of interest) and promoter are inserted in one orientation into the cleavage site or target DNA sequence, but is re-formed when the product coding sequence of interest (e.g., a polypeptide coding sequence of interest) is inserted in the opposite orientation into the cleavage site or target DNA sequence.
[0132] The constructs disclosed herein may include a polyadenylation sequence (e.g., downstream or 3' of the product coding sequence of interest). Methods for designing suitable polyadenylation tail sequences are well known. A polyadenylation tail sequence may be encoded, for example, as a "polyA" stretch downstream of the product of a polypeptide coding sequence of interest. The polyA tail may contain, for example, at least 20, 30, 40, 50, 60, 70, 80, 90, or 100 adenines, and optionally up to 300 adenines. In specific examples, the polyA tail contains 95, 96, 97, 98, 99, or 100 adenine nucleotides. Methods for designing suitable polyadenylation tail sequences and / or polyadenylation signal sequences are well known. For example, the polyadenylation signal sequence AAUAAA is commonly used in mammalian systems, although variants such as UAUAAA or AU / GUAAA have been identified. See, e.g., Proudfoot (2011) Genes & Dev. 25(17):1770-82, incorporated herein by reference in its entirety for all purposes. The term polyadenylation signal sequence refers to any sequence that directs the termination of transcription and the addition of a poly(A) tail to an mRNA transcript. In eukaryotes, transcription terminators are recognized by protein factors, and termination is followed by polyadenylation, the process of adding a poly(A) tail to an mRNA transcript in the presence of poly(A) polymerase. Mammalian poly(A) signals typically consist of a core sequence approximately 45 nucleotides long, which may be flanked by various auxiliary sequences that serve to increase the efficiency of cleavage and polyadenylation. The core sequence, called the polyA recognition motif or sequence, consists of a highly conserved upstream element (AATAAA or AAUAAA) in the mRNA that is recognized by the cleavage and polyadenylation-specificity factor (CPSF), and a poorly defined downstream region (rich in U or G and U) that is bound by the cleavage stimulation factor (CstF).Examples of transcription terminators that can be used include, for example, the human growth hormone (HGH) polyadenylation signal, the simian virus 40 (SV40) late polyadenylation signal, the rabbit beta globin polyadenylation signal, the bovine growth hormone (BGH) polyadenylation signal, the phosphoglycerate kinase (PGK) polyadenylation signal, the AOX1 transcription termination sequence, the CYC1 transcription termination sequence, or any transcription termination sequence known to be suitable for regulating gene expression in eukaryotic cells. In one example, the polyadenylation signal is the simian virus 40 (SV40) late polyadenylation signal. In another example, the polyadenylation signal is the bovine growth hormone (BGH) polyadenylation signal.
[0133] (1) The product and polypeptide of interest Any product of interest can be encoded by the nucleic acid constructs disclosed herein. For example, the product of interest can be a therapeutic product of interest, such as a therapeutic RNA or a therapeutic polypeptide.
[0134] In one example, the product of interest is an RNA of interest, such as an miRNA, an antisense oligonucleotide, an RNAi agent, or a guide RNA for use in a CRISPR / Cas system. For example, the RNA of interest can be a therapeutic RNA.
[0135] An "RNAi agent" is a composition comprising small double-stranded RNA or RNA-like (e.g., chemically modified RNA) oligonucleotide molecules that can facilitate the degradation or inhibition of translation of a target RNA, such as messenger RNA (mRNA), in a sequence-specific manner. The oligonucleotides of an RNAi agent are polymers of linked nucleosides, each of which can be independently modified or unmodified. RNAi agents function via the RNA interference mechanism (i.e., induce RNA interference through interaction with the RNA interference pathway machinery (RNA-induced silencing complex or RISC) in mammalian cells). Although RNAi agents, as that term is used herein, are believed to act primarily via the RNA interference mechanism, the disclosed RNAi agents are not constrained or limited to any particular pathway or mechanism of action. RNAi agents disclosed herein include sense and antisense strands, and include, but are not limited to, small interfering RNA (siRNA), double-stranded RNA (dsRNA), microRNA (miRNA), short hairpin RNA (shRNA), and Dicer substrates. The antisense strand of an RNAi agent described herein is at least partially complementary to the sequence (i.e., the sequence or order of nucleobases or nucleotides described by a series of letters using standard nomenclature) of a target RNA.
[0136] Single-stranded ASOs and RNA interference (RNAi) share the fundamental principle of oligonucleotides binding to target RNA via Watson-Crick base pairing. Without wishing to be bound by theory, during RNAi, a small RNA duplex (RNAi agent) binds to an RNA-induced silencing complex (RISC), where one strand (the passenger strand) is lost and the remaining strand (the guide strand) binds to complementary RNA in cooperation with the RISC. Argonaute 2 (Ago2), a catalytic component of RISC, then cleaves the target RNA. The guide strand is always associated with either the complementary sense strand or a protein (RISC). In contrast, ASOs must survive and function as single strands. ASOs bind to target RNA and either block other factors, such as ribosomes or splicing factors, from binding to the RNA or recruit proteins such as nucleases. Various modifications and target regions are selected for ASOs based on the desired mechanism of action. Gapmers are ASO oligonucleotides containing 2-5 chemically modified nucleotides (e.g., LNA or 2'-MOE) at each end flanking a central 8-10 base gap in the DNA. After binding to the target RNA, the DNA-RNA hybrid serves as a substrate for RNase H.
[0137] In another example, the product of interest is a polypeptide of interest. In one example, the polypeptide of interest is a therapeutic polypeptide. For example, a therapeutic polypeptide can be a polypeptide that is lacking or defective in a subject. In one example, the polypeptide of interest is an enzyme.
[0138] In one example, the polypeptide of interest is an antibody or antigen-binding protein. In another example, the polypeptide of interest is an exogenous T cell receptor or chimeric antigen receptor (CAR). In another example, the polypeptide of interest is a Cas protein (e.g., Cas9) for use in a CRISPR / Cas system.
[0139] As disclosed herein, an "antigen-binding protein" includes any protein that binds to an antigen. Examples of antigen-binding proteins include antibodies, antigen-binding fragments of antibodies, multispecific antibodies (e.g., bispecific antibodies), scFvs, bis-scFvs, diabodies, triabodies, tetrabodies, V-NARs, VHHs, VLs, F(ab)s, F(ab)s, DVDs (dual variable domain antigen-binding proteins), SVDs (single variable domain antigen-binding proteins), bispecific T-cell engagers (BiTEs), or Davis bodies (U.S. Patent No. 8,586,713, incorporated herein by reference in its entirety for all purposes).
[0140] The term "antibody" includes immunoglobulin molecules comprising four polypeptide chains, two heavy (H) chains and two light (L) chains, interconnected by disulfide bonds. Each heavy chain contains a heavy chain variable domain and a heavy chain constant region (C H The constant region of the heavy chain contains C H 1. C H 2, and C H Each light chain contains three domains: a light chain variable domain and a light chain constant region (C L ). Heavy and light chain variable domains can be further divided into regions of hypervariability called complementarity-determining regions (CDRs), interspersed with more conserved regions called framework regions (FRs). Each heavy and light chain variable domain contains three CDRs and four FRs arranged from amino-terminus to carboxy-terminus in the following order: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4 (heavy chain CDRs may be abbreviated as HCDR1, HCDR2, and HCDR3, and light chain CDRs may be abbreviated as LCDR1, LCDR2, and LCDR3). The term "high affinity" antibody refers to an antibody having a high affinity of about 10 -9 M or less (e.g., 1×10 -9 M, 1 x 10 -10 M, about 1 x 10 -11 M, or approximately 1 x 10 -12 M) K for its target epitope D In one embodiment, K Dis measured by surface plasmon resonance, e.g., BIACORE™. D is measured by ELISA.
[0141] The antigen-binding protein or antibody may be, for example, a neutralizing antigen-binding protein or antibody or a broadly neutralizing antigen-binding protein or antibody. Neutralizing antibodies protect cells from antigens or infectious agents by neutralizing all biological effects of the cell. Broadly neutralizing antibodies (bNAbs) affect multiple strains of a particular bacterium or virus. For example, broadly neutralizing antibodies can focus on conserved functional targets and attack vulnerable sites on conserved bacterial or viral proteins (e.g., the vulnerable site of the influenza virus protein hemagglutinin). Antibodies produced by the immune system during infection or vaccination tend to focus on easily accessible loops on the surface of bacteria or viruses, which often have large sequence and conformational variability. This is problematic for two reasons: populations of bacteria or viruses can quickly evade these antibodies, and the antibodies attack parts of proteins that are not essential for function. Broadly neutralizing antibodies, termed "broad spectrum" because they attack many strains of bacteria or viruses and "neutralizing" because they attack key functional sites of bacteria or viruses to prevent infection, can overcome these problems. Unfortunately, however, these antibodies are usually too slow to provide effective protection from disease.
[0142] The antigen-binding proteins disclosed herein can target any antigen. The term "antigen" refers to a substance, whether the entire molecule or a domain within the molecule, that can induce the production of antibodies with binding specificity to that substance. The term antigen also includes substances that do not induce antibody production through self-recognition in wild-type host organisms, but can induce such a response in host animals using appropriate genetic engineering to break immunological tolerance.
[0143] As an example, the target antigen may be a disease-associated antigen. The term "disease-associated antigen" refers to an antigen whose presence is correlated with the onset or progression of a particular disease. For example, the antigen may be present in a disease-associated protein (i.e., a protein whose expression is correlated with the onset or progression of a disease). Optionally, the disease-associated protein may be a protein that is expressed in a particular type of disease but is not normally expressed in healthy adult tissues (i.e., a protein with disease-specific or disease-restricted expression). However, a disease-associated protein need not exhibit disease-specific or disease-restricted expression.
[0144] As an example, the disease-associated antigen may be a cancer-associated antigen. The term "cancer-associated antigen" refers to an antigen whose presence is correlated with the development or progression of one or more types of cancer. For example, the antigen may be present in a cancer-associated protein (i.e., a protein whose expression is correlated with the development or progression of one or more types of cancer). For example, a cancer-associated protein may be an oncogenic protein (i.e., a protein with an activity that may contribute to the progression of cancer, such as a protein that regulates cell growth) or a tumor suppressor protein (i.e., a protein that acts to reduce the likelihood of cancer formation, typically by negatively regulating the cell cycle or promoting apoptosis). Optionally, a cancer-associated protein may be a protein that is expressed in a particular type of cancer but not normally expressed in healthy adult tissues (i.e., a protein with cancer-specific, cancer-restricted, tumor-specific, or tumor-restricted expression). However, a cancer-associated protein need not have cancer-specific, cancer-restricted, tumor-specific, or tumor-restricted expression. Examples of proteins considered cancer-specific or cancer-restricted are cancer-testis antigen or carcinoembryonic antigen. Cancer-testis antigens (CTAs) are a large family of tumor-associated antigens expressed in human tumors of different histological origins but not in normal tissues, except in male germ cells. In cancer, these onset antigens can be re-expressed and serve as sites of immune activation. Carcinoembryonic antigens (OFAs) are proteins normally present only during fetal development but are found in adults with certain types of cancer.
[0145] As another example, the disease-associated antigen may be an infectious disease-associated antigen. The term "infectious disease-associated antigen" refers to an antigen whose presence is correlated with the onset or progression of a particular infectious disease. For example, the antigen may be present in an infectious disease-associated protein (i.e., a protein whose expression is correlated with the onset or progression of an infectious disease). Optionally, the infectious disease-associated protein may be a protein that is expressed in a particular type of infectious disease but is not normally expressed in healthy adult tissues (i.e., a protein with infectious disease-specific or infectious disease-restricted expression). However, infectious disease-associated proteins do not need to exhibit infectious disease-specific or infectious disease-restricted expression. For example, the antigen may be a viral antigen or a bacterial antigen. Such antigens include, for example, molecular structures on the surface of viruses or bacteria (e.g., viral proteins or bacterial proteins) that can be recognized by the immune system and induce an immune response.
[0146] The term "epitope" refers to a site on an antigen to which an antigen-binding protein (e.g., an antibody) binds. Epitopes can be formed from contiguous or noncontiguous amino acids juxtaposed by tertiary folding of one or more proteins. Epitopes formed from contiguous amino acids (also known as linear epitopes) are typically retained upon exposure to denaturing solvents, whereas epitopes formed by tertiary folding (also known as conformational epitopes) are typically lost upon treatment with denaturing solvents. Epitopes typically comprise at least three, more usually at least five, or 8-10 amino acids in a unique spatial conformation. Methods for determining the spatial conformation of epitopes include, for example, X-ray crystallography and two-dimensional nuclear magnetic resonance. See, for example, "Epitope Mapping Protocols," in Methods in Molecular Biology, Vol. 66, Glenn E. Morris, Ed. (1996), incorporated herein by reference in its entirety for all purposes.
[0147] The term "heavy chain" or "immunoglobulin heavy chain" includes immunoglobulin heavy chain sequences, including immunoglobulin heavy chain constant region sequences, from any organism. A heavy chain variable domain, unless otherwise specified, contains three heavy chain CDRs and four FR regions. Fragments of heavy chains include CDRs, CDRs and FRs, and combinations thereof. A typical heavy chain contains (from N-terminus to C-terminus) a CDR following the variable domain. H 1 domain, hinge, C H 2 domain, and C H3 A functional fragment of a heavy chain can specifically recognize an epitope (e.g., in the micromolar, nanomolar, or picomolar range). KD The heavy chain variable domain is a V that recognizes an epitope in a V region present in the germline, and can be expressed and secreted from cells and contains a fragment that contains at least one CDR. H , D H , and J H V derived from segmental repertoire H , D H , and J H The sequences, locations, and nomenclature of the V, D, and J heavy chain segments of various organisms can be found in the IMGT database, accessible via the Internet on the World Wide Web (www) at the URL "imgt.org."
[0148] The term "light chain" includes immunoglobulin light chain sequences from any organism, including human kappa (κ) and lambda (λ) light chains and VpreB, as well as surrogate light chains, unless otherwise specified. A light chain variable domain typically includes three light chain CDRs and four framework (FR) regions, unless otherwise specified. Generally, a full-length light chain includes, from the amino terminus to the carboxyl terminus, a variable domain including FR1-CDR1-FR2-CDR2-FR3-CDR3-FR4, and a light chain constant region amino acid sequence. The light chain variable domain is derived from a repertoire of light chain V and J gene segments present in the germline. L and light chain J LLight chain variable region nucleotide sequences generally comprise gene segments. The sequences, locations, and nomenclature of light chain V and J gene segments from various organisms can be found in the IMGT database, accessible via the World Wide Web (www) Internet at the URL "imgt.org." Light chains include, for example, those that do not selectively bind to either the first or second epitopes selectively bound by the epitope-binding protein in which they appear. Light chains also include those that bind to and recognize, or assist heavy chains in binding and recognizing, one or more epitopes selectively bound by the epitope-binding protein in which they appear.
[0149] As used herein, the term "complementarity-determining region" or "CDR" includes an amino acid sequence encoded by a nucleic acid sequence of an organism's immunoglobulin genes that normally (i.e., in a wild-type animal) appears between two framework regions within the variable region of a light or heavy chain of an immunoglobulin molecule (e.g., an antibody or T-cell receptor). CDRs can be encoded, for example, by germline sequences or by sequences rearranged, for example, by naive or mature B or T cells. CDRs can be modified by somatic mutation (e.g., different from the sequence encoded in the animal's germline), humanization, and / or amino acid substitution, addition, or deletion. In some situations (e.g., in the case of a CDR3), CDRs can be encoded by two or more sequences (e.g., germline sequences) that are not contiguous (e.g., in an unrearranged nucleic acid sequence) but are contiguous in a B-cell nucleic acid sequence, for example, as a result of splicing or joining of sequences (e.g., formation of a heavy chain CDR3 by VDJ recombination).
[0150] The term "unrearranged" includes a state of the immunoglobulin locus in which the V and J gene segments (as well as the D gene segments in the case of heavy chains) are maintained separately but can combine to form rearranged V(D)J genes composed of a single V, (D), and J of the V(D)J repertoire. The term "rearranged" refers to a state in which the V segments form essentially complete V(D)J genes. H or V L It includes the organization of heavy or light chain immunoglobulin loci in which the DJ or J segments are arranged immediately adjacent in an arrangement that encodes the domains, respectively.
[0151] The antigen-binding protein may be a single-chain antigen-binding protein such as scFv. Alternatively, the antigen-binding protein is not a single-chain antigen-binding protein. For example, the antigen-binding protein may comprise separate light and heavy chains. The heavy chain coding sequence may be upstream of the light chain coding sequence, or the light chain coding sequence may be upstream of the heavy chain coding sequence. In one specific example, the heavy chain coding sequence is upstream of the light chain coding sequence. For example, the heavy chain coding sequence may be V H , D H , and J H The light chain coding sequence can comprise a light chain V L and light chain J L The nucleic acid construct may comprise a gene segment. The antigen-binding protein coding sequence may be operably linked to an exogenous promoter. Similarly, the antigen-binding protein coding sequence of the nucleic acid construct may comprise an exogenous signal sequence for secretion. In a specific example, the antigen-binding protein comprises separate light and heavy chains, each chain operably linked to a separate exogenous signal sequence.
[0152] Signal sequences (i.e., N-terminal signal sequences) mediate targeting of nascent secretory and membrane proteins to the endoplasmic reticulum (ER) in a signal recognition particle (SRP)-dependent manner. Typically, signal sequences are cleaved cotranslationally to generate a signal peptide and mature protein. Examples of exogenous signal sequences or peptides that can be used include signal sequences / peptides such as mouse albumin, human albumin, mouse ROR1, human ROR1, human azurocidin, Cricetulus griseus Ig kappa chain VIII region MOPC63-like, and human Ig kappa chain VIII region VG. Other known signal sequences / peptides can also be used. In a specific example, the ROR1 signal sequence is used.
[0153] One or more nucleic acids in a sequence encoding an antigen-binding protein (e.g., a heavy chain-encoding sequence and a light chain-encoding sequence) can be combined in a multicistronic expression construct. For example, nucleic acids encoding heavy and light chains can be combined in a bicistronic expression construct. A multicistronic expression vector simultaneously expresses two or more separate proteins from the same mRNA (i.e., transcripts generated from the same promoter). Suitable strategies for multicistronic expression of proteins include, for example, the use of 2A peptides and internal ribosome entry sites (IRES). As one example, such a multicistronic vector can use one or more internal ribosome entry sites (IRES) to enable translation initiation from internal regions of the mRNA. As another example, such a multicistronic vector can use one or more 2A peptides. These peptides are small, "self-cleaving" peptides, generally 18-22 amino acids in length, that produce equimolar levels of multiple genes from the same mRNA. The ribosome skips the synthesis of the glycyl-prolyl peptide bond at the C-terminus of the 2A peptide, resulting in a "cleavage" between the 2A peptide and the peptide immediately downstream. See, e.g., Kim et al. (2011) PLoS One 6(4):e18556, incorporated herein by reference in its entirety for all purposes. The "cleavage" occurs between the C-terminal glycine and proline residues, and the upstream cistron adds several residues to its end, while the downstream cistron begins with a proline. As a result, the "cleaved" downstream peptide has a proline at its N-terminus. 2A-mediated cleavage is a universal phenomenon in all eukaryotic cells. 2A peptides have been identified in picornaviruses, insect viruses, and type C rotaviruses. See, e.g., Szymczak et al. (2005) Expert Opin Biol Ther 5:627-638, incorporated herein by reference in its entirety for all purposes.Examples of 2A peptides that can be used include Thosea asigna virus 2A (T2A), porcine teschovirus-1 2A (P2A), equine rhinitis A virus (ERAV) 2A (E2A), and FMDV 2A (F2A). Exemplary T2A, P2A, E2A, and F2A sequences include the following: T2A (EGRGSLLTCGDVEENPGP, SEQ ID NO: 31), P2A (ATNFSLLKQAGDVEENPGP, SEQ ID NO: 32), E2A (QCTNYALLKLAGDVESNPGP, SEQ ID NO: 33), and F2A (VKQTLNFDLLKLAGDVESNPGP, SEQ ID NO: 34). A GSG residue can be added to the 5' end of any of these peptides to improve cleavage efficiency.
[0154] In some nucleic acid constructs, a nucleic acid encoding a furin cleavage site is included between the light chain coding sequence and the heavy chain coding sequence. In some nucleic acid constructs, a nucleic acid encoding a linker (e.g., GSG) is included between the light chain coding sequence and the heavy chain coding sequence (e.g., immediately upstream of the 2A peptide coding sequence). For example, a furin cleavage site can be included upstream of the 2A peptide, with both the furin cleavage site and the 2A peptide located between the light chain and the heavy chain (i.e., upstream chain-furin cleavage site-2A peptide-downstream chain). During translation, the first cleavage event occurs at the 2A peptide sequence. However, most of the 2A peptide remains attached as a remnant to the C-terminus of the upstream chain (e.g., the light chain if the light chain is upstream of the heavy chain, or the heavy chain if the heavy chain is upstream of the light chain), with one amino acid added to the N-terminus of the downstream chain (or the N-terminus of the signal sequence if a signal sequence is included upstream of the downstream chain). A second cleavage event, initiated at the furin cleavage site, generates the upstream chain without the 2A remnant for post-translational processing to yield the more native heavy or light chain.
[0155] The term "chimeric antigen receptor" (CAR) refers to a molecule that combines a binding domain for a component present on a target cell, e.g., an antibody-based specificity for a desired antigen, with a T cell receptor activating intracellular domain to generate a chimeric protein that exhibits specific anti-target cellular immune activity. For example, a CAR may comprise an extracellular single-chain antibody binding domain (scFv) fused to the intracellular signaling domain of the T cell antigen receptor complex zeta chain, and when expressed in a T cell, may have the ability to redirect antigen recognition based on the specificity of a monoclonal antibody.
[0156] The polypeptide of interest can be a secreted polypeptide (e.g., a protein that is secreted by the cell and / or a protein that is functionally active as a soluble extracellular protein), or an intracellular polypeptide (e.g., a protein that is not secreted by the cell but is functionally active intracellularly, including soluble cytosolic polypeptides).
[0157] The polypeptide of interest can be a wild-type polypeptide. Alternatively, the polypeptide of interest can be a variant or mutant polypeptide.
[0158] In one example, the polypeptide of interest is a liver protein (e.g., a protein that is endogenously produced in the liver and / or functionally active in the liver). In another example, the polypeptide of interest can be a circulating protein produced by the liver. In another example, the polypeptide of interest can be a non-liver protein.
[0159] A polypeptide of interest can be an exogenous polypeptide. An "exogenous" polypeptide coding sequence can refer to a coding sequence that has been introduced from an exogenous source into a site within a host cell genome (e.g., at a genomic locus, such as a genomic safe harbor locus described herein). That is, an exogenous polypeptide coding sequence is exogenous with respect to its insertion site, and a polypeptide expressed from such an exogenous coding sequence is referred to as an exogenous polypeptide. A heterologous coding sequence can be naturally occurring or engineered, and can be wild-type or mutant. An exogenous coding sequence can include nucleotide sequences other than the sequence encoding the exogenous polypeptide (e.g., an internal ribosome entry site). An exogenous coding sequence can be a coding sequence that naturally occurs in the host genome as a wild-type or variant (e.g., mutant). For example, a host cell contains a coding sequence of interest (as a wild-type or variant), but the same coding sequence or a variant thereof can be introduced as an exogenous source (e.g., for expression at a highly expressed locus). An exogenous coding sequence can also be a coding sequence that does not naturally occur within the host genome, or that expresses an exogenous polypeptide that does not naturally occur within the host genome. An exogenous coding sequence can comprise an exogenous nucleic acid sequence (e.g., a nucleic acid sequence that is not endogenous to the recipient cell) or can be exogenous with respect to its insertion site and / or with respect to the recipient cell.
[0160] The coding sequence of a polypeptide of interest can be codon-optimized for expression in a host cell. For example, the coding sequence can be codon-optimized or can use one or more alternative codons for one or more amino acids of the polypeptide of interest (i.e., the same amino acid sequence). Alternative codons, as used herein, refer to variations in codon usage for a given amino acid, which may or may not be preferred or optimized codons (codon optimization) for a given expression system. Preferred codon usage, or codons that are well tolerated in a given expression system, are known.
[0161] (2) Vector The nucleic acid constructs disclosed herein can be provided in vectors for expression or for integration into and expression from a target genomic locus (e.g., a genomic safe harbor locus). The vectors can include additional sequences, such as an origin of replication, a promoter, and a gene encoding antibiotic resistance. The vectors can also include a nuclease agent component disclosed elsewhere herein. For example, the vector can include a nucleic acid construct encoding a product of interest (e.g., a polypeptide of interest), a CRISPR / Cas system (a nucleic acid encoding a Cas protein and a gRNA), one or more components of a CRISPR / Cas system, or a combination thereof (e.g., a nucleic acid construct and a gRNA). In some cases, a vector including a nucleic acid construct encoding a product of interest (e.g., a polypeptide of interest) does not include any of the nuclease agent components described herein (e.g., does not include a nucleic acid encoding a Cas protein and does not include a nucleic acid encoding a gRNA). Some such vectors include homology arms corresponding to a target site in the target genomic locus. Other such vectors do not include any homology arms.
[0162] Some vectors may be circular. Alternatively, vectors may be linear. Vectors may be packaged for delivery via lipid nanoparticles, liposomes, non-lipid nanoparticles, or viral capsids. Non-limiting exemplary vectors include plasmids, phagemids, cosmids, artificial chromosomes, minichromosomes, transposons, viral vectors, and expression vectors.
[0163] The vector may be a viral vector, such as an adeno-associated viral (AAV) vector. The AAV may be of any suitable serotype and may be single-stranded AAV (ssAAV) or self-complementary AAV (scAAV). Other exemplary viruses / viral vectors include retroviruses, lentiviruses, adenoviruses, vaccinia viruses, poxviruses, and herpes simplex viruses. Viruses can infect dividing cells, non-dividing cells, or both dividing and non-dividing cells. Viruses may integrate into the host genome, or alternatively, not integrate into the host genome. Such viruses may also be engineered to reduce immunity. Viruses may be replication-competent or replication-deficient (e.g., defective in one or more genes required for additional rounds of virion replication and / or packaging). Viruses may induce transient or longer-lasting expression. Viral vectors may be genetically modified from their wild-type counterparts. For example, a viral vector may contain one or more nucleotide insertions, deletions, or substitutions to facilitate cloning or to alter one or more characteristics of the vector. Such characteristics may include packaging capacity, transduction efficiency, immunogenicity, genome integration, replication, transcription, and translation. In some examples, a portion of the viral genome may be deleted to allow the virus to package exogenous sequences with a larger size. In some examples, the viral vector may have enhanced transduction efficiency. In some examples, the immune response induced by the virus in the host may be reduced. In some examples, a viral gene (such as integrase) that promotes integration of viral sequences into the host genome may be mutated to render the virus non-integrating. In some examples, the viral vector may be replication-deficient. In some examples, the viral vector may contain exogenous transcriptional or translational control sequences to drive expression of coding sequences on the vector. In some examples, the virus may be helper-dependent.For example, a virus may require one or more helper components to provide viral components (such as viral proteins) needed to amplify and package a vector into a viral particle. In such cases, one or more helper components, including one or more vectors encoding the viral components, can be introduced into a host cell or population of host cells along with the vector system described herein. In other examples, the virus can be helper-free. For example, the virus can amplify and package a vector without a helper virus. In some examples, the vector system described herein may also encode viral components needed for viral amplification and packaging.
[0164] Exemplary viral titers (e.g., AAV titers) are about 10 12 ~about 10 16 Other exemplary viral titers (e.g., AAV titers) include titers of about 10 vg / mL of body weight. 12 ~about 10 16 Examples include vg / kg.
[0165] Adeno-associated viruses (AAVs) are endemic to multiple species, including humans and non-human primates (NHPs). To date, at least 12 natural serotypes and hundreds of natural variants have been isolated and characterized. See, for example, Li et al. (2020) Nat. Rev. Genet. 21:255-272, incorporated herein by reference in its entirety for all purposes. AAV particles are naturally composed of a non-enveloped icosahedral protein capsid containing a single-stranded DNA (ssDNA) genome. The DNA genome is flanked by two inverted terminal repeats (ITRs) that serve as viral origins of replication and packaging signals. The rep gene encodes four proteins required for viral replication and packaging, while the cap gene encodes three structural capsid subunits that define AAV serotypes and an assembly-activating protein (AAP) that promotes virion assembly in some serotypes.
[0166] Recombinant AAV (rAAV) is currently one of the most commonly used viral vectors in gene therapy, treating human diseases by delivering therapeutic transgenes to target cells in vivo. rAAV vectors are composed of an icosahedral capsid similar to native AAV, but rAAV virions do not encapsidate AAV protein coding or AAV replication sequences. These viral vectors are non-replicative. The only viral sequences required in rAAV vectors are the two ITRs, which are required to direct genome replication and packaging during rAAV vector production. The rAAV genome lacks the AAV rep and cap genes, making them non-replicative in vivo. rAAV vectors are generated by expressing the rep and cap genes in trans in combination with the intended transgene cassette flanked by the AAV ITRs, along with additional viral helper proteins.
[0167] In the rAAV genome, a gene expression cassette can be placed between ITR sequences. Typically, the rAAV genome cassette contains a promoter for driving the expression of the transgene, followed by a polyadenylation sequence. The ITRs flanking the rAAV expression cassette are usually derived from AAV2, the first serotype isolated and converted into a recombinant viral vector. Since then, most rAAV production methods rely on the AAV2Rep-based packaging system. See, for example, Colella et al. (2017) Mol. Ther. Methods Clin. Dev. 8:87-104, the entire contents of which are incorporated herein by reference for all purposes.
[0168] The specific serotype of a recombinant AAV vector influences its in vivo tropism for specific tissues. AAV capsid proteins mediate attachment and entry into target cells, followed by endosomal escape and transport to the nucleus. Therefore, the choice of serotype when developing an rAAV vector influences which cell types and tissues the vector is most likely to bind to and transduce when injected in vivo. Some serotypes of rAAV, including rAAV8, can transduce the liver when delivered systemically in mice, NHPs, and humans. See, e.g., Li et al. (2020) Nat. Rev. Genet. 21:255-272, incorporated herein by reference in its entirety for all purposes.
[0169] Upon entering the nucleus, the ssDNA genome is released from the virion, and a complementary DNA strand is synthesized to generate a double-stranded DNA (dsDNA) molecule. The double-stranded AAV genome naturally circularizes via its ITRs, becoming an episome that persists extrachromosomally in the nucleus. Therefore, for episomal gene therapy programs, rAAV-delivered rAAV episomes provide long-term promoter-driven gene expression in non-dividing cells. However, this rAAV-delivered episomal DNA is diluted as cells divide. In contrast, the gene therapy described herein is based on gene insertion, which allows for long-term gene expression.
[0170] The ssDNA AAV genome consists of two open reading frames, Rep and Cap, flanked by two inverted terminal repeats that allow synthesis of a complementary DNA strand. When constructing an AAV transfer plasmid, the transgene is placed between the two ITRs, and Rep and Cap can be supplied in trans. In addition to Rep and Cap, AAV may require a helper plasmid containing genes from adenovirus. These genes (E4, E2a, and VA) mediate AAV replication. For example, the transfer plasmid, Rep / Cap, and helper plasmid can be transfected into HEK293 cells containing the adenovirus gene E1+ to produce infectious AAV particles. Alternatively, Rep, Cap, and adenovirus helper genes can be combined into a single plasmid. Similar packaging cells and methods can be used for other viruses, such as retroviruses.
[0171] Multiple serotypes of AAV have been identified. These serotypes differ in the types of cells they infect (i.e., their tropism), allowing for preferential transduction of certain cell types. The term AAV includes, for example, AAV1, AAV2, AAV3, AAV3B, AAV4, AAV5, AAV6, AAV6.2, AAV7, AAVrh.64R1, AAVhu.37, AAVrh.8, AAVrh.32.33, AAV8, AAV9, AAV-DJ, AAV2 / 8, AAVrh10, AAVLK03, AV10, AAV11, AAV12, rh10, and hybrids thereof, avian AAV, bovine AAV, canine AAV, equine AAV, primate AAV, non-primate AAV, and ovine AAV. The genomic sequences of various serotypes of AAV, as well as the sequences of the natural terminal repeats (TRs), Rep proteins, and capsid subunits, are known in the art. Such sequences can be found in the literature or public databases such as GenBank. "AAV vector," as used herein, refers to an AAV vector that contains heterologous sequences that are not of AAV origin (i.e., nucleic acid sequences that are heterologous to AAV), and typically contains a sequence encoding a heterologous polypeptide of interest. The constructs may include AAV1, AAV2, AAV3, AAV3B, AAV4, AAV5, AAV6, AAV6.2, AAV7, AAVrh.64R1, AAVhu.37, AAVrh.8, AAVrh.32.33, AAV8, AAV9, AAV-DJ, AAV2 / 8, AAVrh10, AAVLK03, AV10, AAV11, AAV12, rh10, and hybrids thereof, avian AAV, bovine AAV, canine AAV, equine AAV, primate AAV, non-primate AAV, and ovine AAV capsid sequences. Generally, the heterologous nucleic acid sequence (transgene) is flanked by at least one, and typically two, AAV inverted terminal repeat (ITR) sequences. AAV vectors can be either single-stranded (ssAAV) or self-complementary (scAAV). Examples of serotypes for liver tissue include AAV3B, AAV5, AAV6, AAV7, AAV8, AAV9, AAVrh.74, AAV-DJ, and AAVhu.37, particularly AAV8. In a specific example, the AAV vector comprising the nucleic acid construct can be recombinant AAV8 (rAAV8).The rAAV8 vectors described herein are vectors whose capsid is derived from AAV8. For example, an AAV vector that uses ITRs from AAV2 and the capsid of AAV8 is considered to be an rAAV8 vector herein.
[0172] Tropism can be further refined through pseudotyping, which is a mixture of capsids and genomes from different viral serotypes. For example, AAV2 / 5 refers to a virus containing a serotype 2 genome packaged in a serotype 5 capsid. The use of pseudotyped viruses not only improves transduction efficiency but can also alter tropism. Hybrid capsids derived from different serotypes can also be used to modify viral tropism. For example, AAV-DJ contains hybrid capsids from eight serotypes and exhibits high infectivity across a wide range of cell types in vivo. AAV-DJ8 is another example that exhibits the properties of AAV-DJ but with enhanced brain uptake. AAV serotypes can also be modified by mutations. Examples of mutational modifications in AAV2 include Y444F, Y500F, Y730F, and S662V. Examples of mutational modifications in AAV3 include Y705F, Y731F, and T492V. Examples of mutational modifications of AAV6 include S663V and T492V. Other pseudotyped / modified AAV variants include AAV2 / 1, AAV2 / 6, AAV2 / 7, AAV2 / 8, AAV2 / 9, AAV2.5, AAV8.2, and AAV / SASTG.
[0173] To accelerate transgene expression, self-complementary AAV (scAAV) variants can be used. Because AAV relies on the cell's DNA replication machinery to synthesize the complementary strand of its single-stranded DNA genome, transgene expression can be delayed. To address this delay, scAAV can be used, which contain complementary sequences that can spontaneously anneal upon infection, eliminating the need for host cell DNA synthesis. However, single-stranded AAV (ssAAV) vectors can also be used.
[0174] To increase packaging capacity, a long transgene can be split between two AAV transfer plasmids, one containing a 3' splice donor and the other a 5' splice acceptor. Upon co-infection of cells, these viruses can form concatemers that can be spliced together to express the full-length transgene. This allows for expression of longer transgenes, but at a reduced efficiency. A similar method for increasing capacity utilizes homologous recombination. For example, the transgene can be split between two transfer plasmids, but with substantial sequence overlap, such that co-expression induces homologous recombination and expression of the full-length transgene.
[0175] C. Nuclease Agents and CRISPR / Cas Systems The methods and compositions disclosed herein may utilize nuclease agents, such as clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR-associated (Cas) systems, zinc finger nuclease (ZFN) systems, or transcription activator-like effector nuclease (TALEN) systems, or components of such systems, to modify a target genomic locus within a target locus, such as a genomic safe harbor locus, for insertion of a nucleic acid construct disclosed herein. Generally, the nuclease agent involves the use of an engineered cleavage system to induce a double-stranded break or nick (i.e., a single-stranded break) at the nuclease target site. Cleavage or nicking can occur through the use of a specific nuclease, such as an engineered ZFN, TALEN, or CRISPR / Cas system, with an engineered guide RNA to induce specific cleavage or nicking of the nuclease target site. Any nuclease agent that induces a nick or double-stranded break at the desired target sequence can be used in the methods and compositions disclosed herein. The nuclease agent can be used to create a site of insertion at a desired locus (a genomic safe harbor locus) within the host genome, at which a nucleic acid construct is inserted to express a product of interest (e.g., a polypeptide of interest). The product of interest (e.g., a polypeptide of interest) can be exogenous with respect to the insertion site or locus, such as an extragenic genomic safe harbor locus where the product of interest (e.g., a polypeptide of interest) is normally expressed.
[0176] In one example, the nuclease agent is a CRISPR / Cas system. In another example, the nuclease agent comprises one or more ZFNs. In yet another example, the nuclease agent comprises one or more TALENs. In a specific example, the CRISPR / Cas system or components of such a system targets a genomic safe harbor locus described elsewhere herein in the cell. In a more specific example, the CRISPR / Cas system or components of such a system targets an L-SH5, L-SH18, or L-SH20 genomic safe harbor locus described herein (e.g., a human L-SH5, L-SH18, or L-SH20 genomic safe harbor locus) in the cell. In a more specific example, the CRISPR / Cas system or components of such a system targets a human L-SH5, L-SH18, or L-SH20 genomic safe harbor locus described herein in the cell. In specific examples, the CRISPR / Cas system or components of such a system are targeted in cells to the mouse L-SH5, L-SH18, or L-SH20 genomic safe harbor loci described elsewhere herein.
[0177] CRISPR / Cas systems include transcripts and other elements involved in the expression of or directing the activity of Cas genes. CRISPR / Cas systems can be, for example, Type I, Type II, Type III, or Type V systems (e.g., subtype VA or subtype VB). The methods and compositions disclosed herein can employ CRISPR / Cas systems by utilizing a CRISPR complex (including a guide RNA (gRNA) complexed with a Cas protein) for site-specific binding or cleavage of nucleic acids. A CRISPR / Cas system that targets a genomic safe harbor locus includes a Cas protein (or a nucleic acid encoding a Cas protein) and one or more guide RNAs (or DNA encoding one or more guide RNAs), each of which targets a different guide RNA target sequence in the target genomic locus.
[0178] The CRISPR / Cas systems used in the compositions and methods disclosed herein may not be naturally occurring. Non-naturally occurring systems include any that exhibit human involvement, such as one or more components of the system being modified or mutated from their naturally occurring state, being at least substantially free of at least one other component with which they are naturally associated in nature, or being associated with at least one other component with which they are not naturally associated. For example, some CRISPR / Cas systems use non-naturally occurring CRISPR complexes that include a non-naturally occurring gRNA and Cas protein together, use non-naturally occurring Cas proteins, or use non-naturally occurring gRNAs.
[0179] (1) Target genomic locus Any target genomic locus capable of expressing a gene can be used, such as the safe harbor loci described elsewhere herein. As used herein, a target genomic locus can be a genomic safe harbor locus. A genomic safe harbor locus includes a chromosomal locus at which a transgene or other exogenous nucleic acid insert can be stably and reliably expressed in a tissue of interest without overtly altering cellular behavior or phenotype (i.e., without adversely affecting the host cell). For example, a genomic safe harbor locus can be a locus in which expression of an inserted gene sequence is not perturbed by read-through expression from adjacent genes. For example, a genomic safe harbor locus can include a chromosomal locus at which exogenous DNA can integrate and function in a predictable manner without adversely affecting the structure or expression of endogenous genes. Genomic safe harbor loci can be targeted with high efficiency, and safe harbor loci can be disrupted without any obvious phenotype. A genomic safe harbor locus can include extragenic or intragenic regions, e.g., intragenic loci that are nonessential, dispensable, or that can be disrupted without any obvious phenotypic consequences.
[0180] The genomic safe harbor loci described herein can be genomic loci that, when targeted for integration in a subject, do not alter liver function. The genomic safe harbor loci described herein can be genomic loci that, when targeted for integration in a subject, do not alter alanine aminotransferase (alanine transaminase or ALT) levels. The genomic safe harbor loci described herein can be genomic loci that, when targeted for integration in a subject, do not alter aspartate aminotransferase (AST) levels. The genomic safe harbor loci described herein can be genomic loci that, when targeted for integration in a subject, do not alter alkaline phosphatase (ALP) levels. The genomic safe harbor loci described herein can be genomic loci that, when targeted for integration in a subject, do not alter body weight. The genomic safe harbor loci described herein can be genomic loci that, when targeted for integration in a subject, do not alter proliferation (e.g., as assessed by Ki67 staining) in a target organ, such as the liver, etc. A genomic safe harbor locus as described herein can be a genomic locus that, when targeted for integration in a subject, does not cause oncogenic transformation in a target organ, such as the liver (e.g., as assessed by H&E staining).
[0181] A genomic safe harbor locus as described herein can be a genomic locus that has an open chromatin configuration in the liver such that an exogenous nucleic acid insert can be stably and reliably expressed in the liver. Alternatively, a genomic safe harbor locus can be a genomic locus that has an open chromatin configuration in another tissue or cell type (e.g., a hematopoietic cell such as a hematopoietic stem cell, T cell, B cell, and / or macrophage) such that an exogenous nucleic acid insert can be stably and reliably expressed in that tissue or cell type.
[0182] The genomic safe harbor loci described herein may be extragenic genomic safe harbor loci (i.e., located outside of genes). In a specific example, the genomic safe harbor loci described herein are extragenic genomic safe harbor loci with an open chromatin configuration in the liver.
[0183] In specific examples, genomic safe harbor loci can be more than 300 kb from any cancer-associated gene (e.g., to prevent insertional oncogenesis), more than 300 kb from any miRNA or small RNA (e.g., to preserve regulation of gene expression and cell development), more than 50 kb from the 5' end of any gene (e.g., to avoid perturbing endogenous gene expression), more than 50 kb from any origin of replication, more than 50 kb from any ultraconserved element (e.g., non-coding intragenic or intergenic regions that are completely conserved in the human, mouse, and rat genomes), outside of regions of copy number variation, and open chromatin (e.g., as determined by ATAC-Seq analysis (e.g., in human liver biopsy samples)). Additionally, genomic safe harbor loci may not overlap with regulatory regions (e.g., H3K4me1, H3K27ac, and / or H3K4me3 markers), heterochromatic regions (e.g., H3K9me3 markers), or regions predicted to be involved in chromatin organization (e.g., CTCF signaling).
[0184] In specific examples, genomic safe harbor loci are located at the following genomic locations: (i) human chromosome 13, coordinates 77460242 to 77460537 (referred to herein as L-SH5), or the corresponding region (e.g., orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse; (ii) human chromosome 6, coordinates 170031084 to 170031382 (referred to herein as L-SH18); is selected from (iii) a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or a rodent such as a rat or mouse, and (iv) human chromosome 9, coordinates 25207412-25207703 (referred to herein as L-SH20), or a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or a rodent such as a rat or mouse.
[0185] In specific examples, genomic safe harbor loci are located at the following genomic coordinates: (i) from about 77460242 to about 77460537 on human chromosome 13 (corresponding to L-SH5), or the corresponding region (e.g., an orthologous region or a syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse; (ii) from about 170031084 to about 170031382 on human chromosome 6 (corresponding to L-SH18), or the corresponding region (e.g., an orthologous region or a syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse; and (iii) about 25207412 to about 25207703 of human chromosome 9 (corresponding to L-SH20), or the corresponding region (e.g., orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., non-human primate), or rodent such as a rat or mouse. The term "about" when referring to genomic coordinates means ±20 base pairs. In other examples, genomic safe harbor loci are near the region identified by the above coordinates. The term "near" when referring to genomic coordinates means ±5 kb, ±4 kb, ±3 kb, ±2 kb, ±1 kb, ±0.5 kb, ±0.4 kb, ±0.3 kb, ±0.2 kb, or ±0.1 kb.
[0186] In one specific example, the genomic safe harbor locus is human L-SH5 (chromosome 13, coordinates 77460242-77460537) or the corresponding region (e.g., an orthologous region or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse. Syntenic regions are derived from a single ancestral genomic region. For example, syntenic regions can be derived from different organisms and result from speciation.
[0187] In another specific example, the genomic safe harbor locus is human L-SH18 (chromosome 6, coordinates 170031084-170031382), or the corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, a non-human mammal (e.g., a non-human primate), or a rodent such as a rat or mouse.
[0188] In another specific example, the genomic safe harbor locus is human L-SH20 (chromosome 9, coordinates 25207412-25207703), or the corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, a non-human mammal (e.g., a non-human primate), or a rodent such as a rat or mouse.
[0189] In one specific example, the genomic safe harbor locus corresponds to human L-SH5 (coordinates from about 77460242 to about 77460537 on chromosome 13), or a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse, or a variant thereof located at the same position or locus in a human chromosome, or an orthologous or syntenic region in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse. The term "about," when referring to genomic coordinates, means ±20 base pairs. In other examples, the genomic safe harbor locus is near the region identified by the above coordinates. The term "nearby" when referring to genomic coordinates means ±5 kb, ±4 kb, ±3 kb, ±2 kb, ±1 kb, ±0.5 kb, ±0.4 kb, ±0.3 kb, ±0.2 kb, or ±0.1 kb.
[0190] In another specific example, the genomic safe harbor locus corresponds to human L-SH18 (coordinates from about 170031084 to about 170031382 on chromosome 6), or a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse, or a variant thereof located at the same position or locus in a human chromosome, or an orthologous or syntenic region in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse. The term "about," when referring to genomic coordinates, means ±20 base pairs. In other examples, the genomic safe harbor locus is near the region identified by the above coordinates. The term "nearby" when referring to genomic coordinates means ±5 kb, ±4 kb, ±3 kb, ±2 kb, ±1 kb, ±0.5 kb, ±0.4 kb, ±0.3 kb, ±0.2 kb, or ±0.1 kb.
[0191] In another specific example, the genomic safe harbor locus corresponds to human L-SH20 (coordinates of about 25207412 to about 25207703 on chromosome 9), or a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse, or a variant thereof located at the same position or locus in a human chromosome, or an orthologous or syntenic region in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat or mouse. The term "about," when referring to genomic coordinates, means ±20 base pairs. In other examples, the genomic safe harbor locus is near the region identified by the above coordinates. The term "nearby" when referring to genomic coordinates means ±5 kb, ±4 kb, ±3 kb, ±2 kb, ±1 kb, ±0.5 kb, ±0.4 kb, ±0.3 kb, ±0.2 kb, or ±0.1 kb.
[0192] In specific examples, genomic safe harbor loci are located at the following genomic locations: (i) mouse chromosome 14, coordinates 103,450,397-103,451,396 (referred to herein as mouse L-SH5), or the corresponding region (e.g., orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., non-human primate), or rodent, such as a rat; (ii) mouse chromosome 17, coordinates 15,226,387-15,227,386 (referred to herein as mouse L-SH18). and (iii) mouse chromosome 4, coordinates 92,827,563 to 92,828,592 (referred to herein as mouse L-SH20), or the corresponding region (e.g., orthologous region or syntenic region) in a non-human animal, non-human mammal (e.g., non-human primate), or rodent such as a rat.
[0193] In specific examples, genomic safe harbor loci are located at the following genomic coordinates: (i) mouse chromosome 14 from about 103,450,397 to about 103,451,396 (corresponding to mouse L-SH5) or the corresponding region (e.g., an orthologous region or a syntenic region) in a non-human animal, a non-human mammal (e.g., a non-human primate), or a rodent such as a rat; (ii) mouse chromosome 17 from about 15,226,387 to about 15,227,386 (corresponding to mouse L-SH18); and (iii) from about 92,827,563 to about 92,828,592 of mouse chromosome 4 (corresponding to mouse L-SH20) or the corresponding region (e.g., orthologous region or syntenic region) in a non-human animal, non-human mammal (e.g., non-human primate), or rodent such as a rat. The term "about" when referring to genomic coordinates means ±20 base pairs. In other examples, genomic safe harbor loci are near the region identified by the above coordinates. The term "near" when referring to genomic coordinates means ±5 kb, ±4 kb, ±3 kb, ±2 kb, ±1 kb, ±0.5 kb, ±0.4 kb, ±0.3 kb, ±0.2 kb, or ±0.1 kb.
[0194] In one specific example, the genomic safe harbor locus is mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) or the corresponding region (e.g., an orthologous region or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat. Syntenic regions are derived from a single ancestral genomic region. For example, syntenic regions can be derived from different organisms and result from speciation.
[0195] In another specific example, the genomic safe harbor locus is mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386), or the corresponding region (e.g., an orthologous region or syntenic region) in a non-human animal, a non-human mammal (e.g., a non-human primate), or a rodent such as a rat.
[0196] In another specific example, the genomic safe harbor locus is mouse L-SH20 (chromosome 4, coordinates 92,827,563-92,828,592), or the corresponding region (e.g., an orthologous region or syntenic region) in a non-human animal, a non-human mammal (e.g., a non-human primate), or a rodent such as a rat.
[0197] In one specific example, the genomic safe harbor locus corresponds to mouse L-SH5 (coordinates of about 103,450,397 to about 103,451,396 on mouse chromosome 14), or a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat, or a variant thereof located at the same position or locus in a chromosome of a mouse, or an orthologous or syntenic region in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat. The term "about," when referring to genomic coordinates, means ±20 base pairs. In other examples, the genomic safe harbor locus is near the region identified by the above coordinates. The term "nearby" when referring to genomic coordinates means ±5 kb, ±4 kb, ±3 kb, ±2 kb, ±1 kb, ±0.5 kb, ±0.4 kb, ±0.3 kb, ±0.2 kb, or ±0.1 kb.
[0198] In another specific example, the genomic safe harbor locus corresponds to mouse L-SH18 (coordinates from about 15,226,387 to about 15,227,386 on mouse chromosome 17), or a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat, or a variant thereof located at the same position or locus in a chromosome of a mouse, or an orthologous or syntenic region in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat. The term "about," when referring to genomic coordinates, means ±20 base pairs. In other examples, the genomic safe harbor locus is near the region identified by the above coordinates. The term "nearby" when referring to genomic coordinates means ±5 kb, ±4 kb, ±3 kb, ±2 kb, ±1 kb, ±0.5 kb, ±0.4 kb, ±0.3 kb, ±0.2 kb, or ±0.1 kb.
[0199] In another specific example, the genomic safe harbor locus corresponds to mouse L-SH20 (coordinates from about 92,827,563 to about 92,828,592 on mouse chromosome 4), or a corresponding region (e.g., an orthologous or syntenic region) in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat, or a variant thereof located at the same position or locus in a chromosome of a mouse, or an orthologous or syntenic region in a non-human animal, non-human mammal (e.g., a non-human primate), or rodent such as a rat. The term "about," when referring to genomic coordinates, means ±20 base pairs. In other examples, the genomic safe harbor locus is near the region identified by the above coordinates. The term "nearby" when referring to genomic coordinates means ±5 kb, ±4 kb, ±3 kb, ±2 kb, ±1 kb, ±0.5 kb, ±0.4 kb, ±0.3 kb, ±0.2 kb, or ±0.1 kb.
[0200] (2) Cas protein Cas proteins generally contain at least one RNA recognition or binding domain capable of interacting with a guide RNA. Cas proteins may also contain a nuclease domain (e.g., a DNase domain or an RNase domain), a DNA-binding domain, a helicase domain, a protein-protein interaction domain, a dimerization domain, and other domains. Some such domains (e.g., a DNase domain) may be derived from naturally occurring Cas proteins. Other such domains may be added to create modified Cas proteins. Nuclease domains have catalytic activity for nucleic acid cleavage, including covalent cleavage of nucleic acid molecules. Cleavage can generate blunt or sticky ends and may be single-stranded or double-stranded. For example, wild-type Cas9 proteins typically generate blunt cleavage products. Alternatively, wild-type Cpf1 proteins (e.g., FnCpf1) may result in cleavage products with a 5-nucleotide 5' overhang, with cleavage occurring after the 18th base pair from the PAM sequence on the non-targeted strand and after the 23rd base on the targeted strand. The Cas protein may have full cleavage activity and create a double-stranded break at the target genomic locus (e.g., a double-stranded break with blunt ends), or it may be a nickase that creates a single-stranded break at the target genomic locus.
[0201] Examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (Csn1 or Csx12), Cas10, Cas10d, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (CasA), Cse2 (CasB), and Cse3 (CasE). , Cse4 (CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, and Cu1966, as well as homologous or modified versions thereof.
[0202] An exemplary Cas protein is a Cas9 protein or a protein derived from a Cas9 protein. Cas9 proteins are from type II CRISPR / Cas systems and typically share four key motifs with a conserved architecture: motifs 1, 2, and 4 are RuvC-like motifs, and motif 3 is an HNH motif. Exemplary Cas9 proteins are Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicellulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohlobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Neisseria meningitidis, or Campylobacter jejuni. Additional examples of Cas9 family members are described in International Publication No. WO 2014 / 131833, which is incorporated by reference in its entirety for all purposes. Cas9 from Streptococcus pyogenes (SpCas9) (e.g., assigned UniProt accession number Q99ZW2) is an exemplary Cas9 protein. An exemplary SpCas9 protein sequence is set forth in SEQ ID NO: 1 (encoded by the DNA sequence set forth in SEQ ID NO: 2). Smaller Cas9 proteins (e.g., Cas9 proteins whose coding sequences are compatible with maximum AAV packaging capacity when combined with the guide RNA coding sequence and regulatory elements of Cas9 and guide RNA, such as SaCas9, CjCas9, and Nme2Cas9) are other exemplary Cas9 proteins. For example, Cas9 from S. aureus (SaCas9) (e.g., assigned UniProt accession number J7RUA5) is another exemplary Cas9 protein. Similarly, Cas9 from Campylobacter jejuni (CjCas9) (e.g., assigned UniProt accession number Q0P897) is another exemplary Cas9 protein. See, e.g., Kim et al. (2017) Nat. Commun. 8:14500, incorporated herein by reference in its entirety for all purposes. SaCas9 is smaller than SpCas9, and CjCas9 is smaller than both SaCas9 and SpCas9. Cas9 from Neisseria meningitidis (Nme2Cas9) is another exemplary Cas9 protein. See, e.g., Edraki et al. (2019) Mol.See Cell 73(4):714-726, incorporated herein by reference in its entirety for all purposes. Other exemplary Cas9 proteins include Streptococcus thermophilus Cas9 proteins (e.g., Streptococcus thermophilus LMD-9 Cas9 encoded by the CRISPR1 locus (St1Cas9) or Streptococcus thermophilus Cas9 encoded by the CRISPR3 locus (St3Cas9). Other exemplary Cas9 proteins include Francisella novicida Cas9 (FnCas9) or an RHA Francisella novicida Cas9 mutant that recognizes an alternative PAM (E1369R / E1449H / R1556A substitutions). For these and other exemplary Cas9 proteins, see, for example, Cebrian-Serrano and Davies (2017) Mamm. Genome 28(7):247-261, incorporated herein by reference in its entirety for all purposes. Examples of Cas9 coding sequences, Cas9 mRNA, and Cas9 protein sequences are provided in WO 2013 / 176772, WO 2014 / 065596, WO 2016 / 106121, WO 2019 / 067910, WO 2020 / 082042, U.S. Patent Application Publication Nos. 2020 / 0270617, WO 2020 / 082041, U.S. Patent Application Publication Nos. 2020 / 0268906, WO 2020 / 082046, and U.S. Patent Application Publication Nos. 2020 / 0289628, each of which is incorporated by reference in its entirety for all purposes. Specific examples of ORFs and Cas9 amino acid sequences are provided in Table 30 of paragraph
[0449] of WO 2019 / 067910, and specific examples of Cas9 mRNAs and ORFs are provided in paragraphs
[0214] to
[0234] of WO 2019 / 067910. See also WO 2020 / 082046(A2) (pp. 84-85) and Table 24 in WO 2020 / 069296, each of which is incorporated by reference in its entirety for all purposes.
[0203] Another example of a Cas protein is the Cpf1 (CRISPR from Prevotella and Francisella 1, Cas12a) protein. Cpf1 is a large protein (approximately 1300 amino acids) that contains a RuvC-like nuclease domain homologous to the corresponding domain in Cas9, along with a counterpart of Cas9's characteristic arginine-rich cluster. However, Cpf1 lacks the HNH nuclease domain present in Cas9 proteins, and in contrast to Cas9, which contains long inserts containing the HNH domain, the RuvC-like domain is contiguous in the Cpf1 sequence. See, e.g., Zetsche et al. (2015) Cell 163(3):759-771, incorporated herein by reference in its entirety for all purposes. Exemplary Cpf1 proteins are Francisella tularensis 1, Francisella tularensis subsp.novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp.SCADC, Acidaminococcus sp.BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis Cpf1 from Francisella novicida U112 (FnCpf1, assigned UniProt accession number A0Q7Q2) is an exemplary Cpf1 protein.
[0204] Another example of a Cas protein is CasX (Cas12e). CasX is an RNA-guided DNA endonuclease that generates cohesive double-strand breaks in DNA. CasX is less than 1,000 amino acids in size. Exemplary CasX proteins are from Deltaproteobacteria (DpbCasX or DpbCas12e) and Planctomycetes (PlmCasX or PlmCas12e). Like Cpf1, CasX uses a single RuvC active site for DNA cleavage. See, e.g., Liu et al. (2019) Nature 566(7743):218-223, incorporated herein by reference in its entirety for all purposes.
[0205] Another example of a Cas protein is CasΦ (CasPhi or Cas12j), which is uniquely found in bacteriophages. CasΦ is less than 1000 amino acids in size (e.g., 700-800 amino acids). CasΦ cleavage generates a cohesive 5' overhang. A single RuvC active site in CasΦ is capable of crRNA processing and DNA cleavage. See, e.g., Pausch et al. (2020) Science 369(6501):333-337, incorporated herein by reference in its entirety for all purposes.
[0206] The Cas protein can be a wild-type protein (i.e., one occurring in nature), a modified Cas protein (i.e., a Cas protein variant), or a fragment of a wild-type or modified Cas protein. The Cas protein can also be a catalytically active mutant or fragment of the wild-type or modified Cas protein. A catalytically active mutant or fragment can comprise at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to the wild-type or modified Cas protein or a portion thereof, and an active mutant retains the ability to cleave at the desired cleavage site and thus retains nick-inducing or double-strand break-inducing activity. Assays for nick-inducing or double-strand break-inducing activity are known and generally measure the overall activity and specificity of the Cas protein on a DNA substrate containing the cleavage site.
[0207] One example of a modified Cas protein is the modified SpCas9-HF1 protein, which is a high-fidelity mutant of Streptococcus pyogenes Cas9 with modifications (N497A / R661A / Q695A / Q926A) designed to reduce nonspecific DNA contacts. See, e.g., Kleinstiver et al. (2016) Nature 529(7587):490-495, incorporated herein by reference in its entirety for all purposes. Another example of a modified Cas protein is the modified eSpCas9 mutant (K848A / K1003A / R1060A) designed to reduce off-target effects. See, e.g., Slaymaker et al. (2016) Science 351(6268):84-88, incorporated herein by reference in its entirety for all purposes. Other SpCas9 mutants include K855A and K810A / K1003A / R1060A. For these and other modified Cas proteins, see, e.g., Cebrian-Serrano and Davie (2017) Mamm. Genome 28(7):247-261, incorporated herein by reference in its entirety for all purposes. Another example of a modified Cas9 protein is xCas9, an SpCas9 mutant that can recognize an expanded range of PAM sequences. See, e.g., Hu et al. (2018) Nature 556:57-63, incorporated herein by reference in its entirety for all purposes.
[0208] Cas proteins can be modified to increase or decrease one or more of nucleic acid binding affinity, nucleic acid binding specificity, and enzymatic activity. Cas proteins can also be modified to alter other activities or properties of the protein, such as stability. For example, one or more nuclease domains of a Cas protein can be modified, deleted, or inactivated, or the Cas protein can be truncated to remove domains that are not essential for protein function, or to optimize (e.g., enhance or decrease) an activity or property of the Cas protein.
[0209] Cas proteins can contain at least one nuclease domain, such as a DNase domain. For example, wild-type Cpf1 proteins generally contain a RuvC-like domain, likely in a dimeric conformation, that cleaves both strands of target DNA. Similarly, CasX and CasΦ generally contain a single RuvC-like domain that cleaves both strands of target DNA. Cas proteins can also contain at least two nuclease domains, such as a DNase domain. For example, wild-type Cas9 proteins generally contain a RuvC-like nuclease domain and an HNH-like nuclease domain. The RuvC domain and the HNH domain can each cleave different strands of double-stranded DNA to create a double-strand break in the DNA. See, e.g., Jinek et al. (2012) Science 337(6096):816-821, incorporated herein by reference in its entirety for all purposes.
[0210] One or more of the nuclease domains can be deleted or mutated so that they are no longer functional or have reduced nuclease activity. For example, if one of the nuclease domains is deleted or mutated in a Cas9 protein, the resulting Cas9 protein can be called a nickase and can generate single-strand breaks in double-stranded target DNA but not double-strand breaks (i.e., it can cleave either the complementary strand or the non-complementary strand, but not both). If none of the nuclease domains is deleted or mutated in a Cas9 protein, the Cas9 protein retains double-strand break-inducing activity. An example of a mutation that converts Cas9 into a nickase is the D10A (alanine to aspartic acid at position 10 of Cas9) mutation in the RuvC domain of Cas9 from S. pyogenes. Similarly, H939A (histidine to alanine at amino acid position 839), H840A (histidine to alanine at amino acid position 840), or N863A (asparagine to alanine at amino acid position N863) in the HNH domain of Cas9 from S. pyogenes can convert Cas9 into a nickase. Other examples of mutations that convert Cas9 into a nickase include corresponding mutations in Cas9 from S. thermophilus. See, e.g., Sapranauskas et al. (2011) Nucleic Acids Res. 39(21):9275-9282 and WO 2013 / 141680, each of which is incorporated by reference in its entirety for all purposes. Such mutations can be generated using methods such as site-directed mutagenesis, PCR-mediated mutagenesis, or total gene synthesis. Other examples of nickase-generating mutations can be found, for example, in WO 2013 / 176772 and WO 2013 / 142578, each of which is incorporated by reference in its entirety for all purposes.
[0211] Examples of inactivating mutations in the catalytic domain of xCas9 are the same as those described above for SpCas9. Examples of inactivating mutations in the catalytic domain of the Staphylococcus aureus Cas9 protein are also known. For example, the Staphylococcus aureus Cas9 enzyme (SaCas9) may include a substitution at position N580 (e.g., an N580A substitution) or at position D10 (e.g., a D10A substitution) to generate a Cas nickase. See, for example, International Publication No. WO 2016 / 106236, which is incorporated by reference in its entirety for all purposes. Examples of inactivating mutations in the catalytic domain of Nme2Cas9 are also known (e.g., D16A or H588A). Examples of inactivating mutations in the catalytic domain of St1Cas9 are also known (e.g., D9A, D598A, H599A, or N622A). Examples of inactivating mutations in the catalytic domain of St3Cas9 are also known (e.g., D10A or N870A). Examples of inactivating mutations in the catalytic domain of CjCas9 are also known (e.g., the combination of D8A and H559A). Examples of inactivating mutations in the catalytic domain of FnCas9 and RHA FnCas9 are also known (e.g., N995A).
[0212] Examples of inactivating mutations in the catalytic domain of the Cpf1 protein are also known. For the Cpf1 proteins from Francisella novicida U112 (FnCpf1), Acidaminococcus species BV3L6 (AsCpf1), Lachnospiraceae bacterium ND2006 (LbCpf1), and Moraxella bovoculi 237 (MbCpf1 Cpf1), such mutations can include mutations at positions 908, 993, or 1263 in AsCpf1 or corresponding positions in Cpf1 orthologs, or at positions 832, 925, 947, or 1180 in LbCpf1 or corresponding positions in Cpf1 orthologs. Such mutations can include, for example, one or more of the mutations D908A, E993A, and D1263A in AsCpf1 or corresponding mutations in Cpf1 orthologs, or D832A, E925A, D947A, and D1180A in LbCpf1 or corresponding mutations in Cpf1 orthologs. See, e.g., U.S. Patent Application Publication No. 2016 / 0208243, which is incorporated by reference in its entirety for all purposes.
[0213] Examples of inactivating mutations in the catalytic domain of the CasX protein are also known. For CasX proteins from Deltaproteobacteria, D672A, E769A, and D935A (individually or in combination) or corresponding positions in other CasX orthologs are inactivating. See, e.g., Liu et al. (2019) Nature 566(7743):218-223, incorporated herein by reference in its entirety for all purposes.
[0214] Examples of inactivating mutations in the catalytic domain of the CasΦ protein are also known. For example, D371A and D394A, alone or in combination, are inactivating mutations. See, e.g., Pausch et al. (2020) Science 369(6501):333-337, the entire contents of which are incorporated herein by reference for all purposes.
[0215] Cas proteins can also be operably linked to heterologous polypeptides as fusion proteins. For example, Cas proteins can be fused to a cleavage domain. See WO 2014 / 089290, incorporated herein by reference in its entirety for all purposes. Cas proteins can also be fused to heterologous polypeptides that provide increased or decreased stability. The fusion domain or heterologous polypeptide can be located at the N-terminus, C-terminus, or internally within the Cas protein.
[0216] For example, Cas proteins can be fused to one or more heterologous polypeptides that provide subcellular localization. Such heterologous polypeptides can include one or more nuclear localization signals (NLSs), such as a monopartite SV40 NLS and / or a bipartite alpha importin NLS for targeting the nucleus, a mitochondrial localization signal for targeting mitochondria, an ER retention signal, etc. See, e.g., Lange et al. (2007) J. Biol. Chem. 282(8):5101-5105, incorporated herein by reference in its entirety for all purposes. Such subcellular localization signals can be located at the N-terminus, C-terminus, or anywhere within the Cas protein. The NLS can include a stretch of basic amino acids and can be a monopartite or bipartite sequence. Optionally, the Cas protein can include two or more NLSs, including an NLS at the N-terminus (e.g., an alpha importin NLS or a monopartite NLS) and an NLS at the C-terminus (e.g., an SV40 NLS or a bipartite NLS). The Cas protein may also contain two or more NLSs at the N-terminus and / or two or more NLSs at the C-terminus.
[0217] For example, a Cas protein can be fused with 1 to 10 NLSs (e.g., fused with 1 to 5 NLSs, fused with 1 NLS). When a single NLS is used, the NLS can be linked at the N-terminus or C-terminus of the Cas protein sequence. The NLS can also be inserted within the Cas protein sequence. Alternatively, a Cas protein can be fused with two or more NLSs. For example, a Cas protein can be fused with 2, 3, 4, or 5 NLSs. In a specific example, a Cas protein can be fused with two NLSs. In certain circumstances, the two NLSs can be the same (e.g., two SV40 NLSs) or different. For example, a Cas protein can be fused to two SV40 NLS sequences linked at the carboxy termini. Alternatively, a Cas protein can be fused with two NLSs, one linked at the N-terminus and one linked at the C-terminus. In other examples, a Cas protein can be fused with three NLSs or no NLSs. The NLS can be a single-part sequence, such as the SV40 NLS, PKKKRKV (SEQ ID NO: 3) or PKKKRRV (SEQ ID NO: 4). The NLS can be a bipartite sequence, such as the nucleoplasmin NLS, KRPAATKKAGQAKKKK (SEQ ID NO: 5). In a specific example, a single PKKKRKV (SEQ ID NO: 3) NLS can be linked at the C-terminus of the Cas protein. One or more linkers are optionally included in the fusion site.
[0218] The Cas protein can also be operably linked to a cell penetration domain or protein transduction domain. For example, the cell penetration domain can be derived from the HIV-1 TAT protein, the TLM cell penetration motif from human hepatitis B virus, MPG, Pep-1, VP22, a cell penetration peptide from herpes simplex virus, or a polyarginine peptide sequence. See, e.g., WO 2014 / 089290 and WO 2013 / 176772, each of which is incorporated by reference in its entirety for all purposes. The cell penetration domain can be located at the N-terminus, C-terminus, or anywhere within the Cas protein.
[0219] The Cas protein may also be operably linked to a heterologous polypeptide, such as a fluorescent protein, a purification tag, or an epitope tag, to facilitate tracking or purification. Examples of fluorescent proteins include green fluorescent protein (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, Emerald, Azami Green, Monomeric Azami). Green, CopGFP, AceGFP, ZsGreenl), yellow fluorescent proteins (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, ZsYellowl), blue fluorescent proteins (e.g., eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire), cyan fluorescent proteins (e.g., eCFP, Cerulean, CyPet, AmCyanl, Midoriishi-Cyan), red fluorescent proteins (e.g., mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRedl, AsRed2, eqFP611, mRaspberry, mStrawberry, Jred), orange fluorescent proteins (e.g., mOrange, mKO, Kusabira-Orange, Monomeric Kusabira-Orange, mTangerine, tdTomato), and any other suitable fluorescent proteins.Examples of tags include glutathione-S-transferase (GST), chitin binding protein (CBP), maltose binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, hemagglutinin (HA), nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, histidine (His), biotincarboxyl carrier protein (BCCP), and calmodulin.
[0220] Cas proteins can also be tethered to labeled nucleic acids. Such tethering (i.e., physical linkage) can be achieved through covalent or non-covalent interactions, and tethering can be achieved directly (e.g., via direct fusion or chemical conjugation, which can be achieved by modification of cysteine or lysine residues on the protein or intein modifications) or via one or more intervening linker or adapter molecules, such as streptavidin or aptamers. See, e.g., Pierce et al. (2005) Mini Rev. Med. Chem. 5(1):41-55, Duckworth et al. (2007) Angew. Chem. Int. Ed. Engl. 46(46):8819-8822, Schaeffer and Dixon (2009) Australian J. Chem. 62(10):1328-1332, Goodman et al. (2009) Chembiochem. 10(9):1551-1557, and Khatwani et al. (2012) Bioorg. Med. Chem. 20(14):4532-4539, each of which is incorporated by reference in its entirety for all purposes. Non-covalent strategies for synthesizing protein-nucleic acid conjugates include biotin-streptavidin and nickel-histidine methods. Covalent protein-nucleic acid conjugates can be synthesized by connecting appropriately functionalized nucleic acids and proteins using a variety of chemistries. Some of these chemistries involve direct attachment of oligonucleotides to amino acid residues on the protein surface (e.g., lysine amines or cysteine thiols), while other, more complex schemes require post-translational modifications of the protein or the involvement of catalytic or reactive protein domains. Methods for covalently attaching proteins to nucleic acids can include, for example, chemical crosslinking of oligonucleotides to protein lysine or cysteine residues, expressed protein ligation, chemoenzymatic methods, and the use of photoaptamers. Labeled nucleic acids can be tethered to the C-terminus, N-terminus, or internal regions of the Cas protein. In one example, the labeled nucleic acid is tethered to the C-terminus or N-terminus of the Cas protein.Similarly, the Cas protein can be tethered to the 5' end, 3' end, or an internal region of the labeled nucleic acid. That is, the labeled nucleic acid can be tethered in any orientation and polarity. For example, the Cas protein can be tethered to the 5' end or the 3' end of the labeled nucleic acid.
[0221] The Cas protein can be provided in any form. For example, the Cas protein can be provided in the form of a protein, such as a Cas protein complexed with a gRNA. Alternatively, the Cas protein can be provided in the form of a nucleic acid encoding the Cas protein, such as RNA (e.g., messenger RNA (mRNA)) or DNA. Optionally, the nucleic acid encoding the Cas protein can be codon-optimized for efficient translation into protein in a particular cell or organism. For example, the nucleic acid encoding the Cas protein can be modified to use alternative codons more frequently used in bacterial cells, yeast cells, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, or any other host cell of interest, compared to the naturally occurring polynucleotide sequence. Once the nucleic acid encoding the Cas protein is introduced into a cell, the Cas protein can be expressed transiently, conditionally, or constitutively within the cell.
[0222] The nucleic acid encoding the Cas protein can be stably integrated into the genome of the cell and operably linked to a promoter active in the cell. Alternatively, the nucleic acid encoding the Cas protein can be operably linked to a promoter in an expression construct. An expression construct includes any nucleic acid construct capable of directing the expression of a gene or other nucleic acid sequence of interest (e.g., a Cas gene) and transferring such a nucleic acid sequence of interest into a target cell. For example, the nucleic acid encoding the Cas protein can be present in a vector containing DNA encoding a gRNA. Alternatively, the nucleic acid can be in a vector or plasmid that is separate from the vector containing DNA encoding the gRNA. Promoters that can be used in expression constructs include, for example, promoters active in human cells, human liver cells, or human hepatocytes. Such promoters can be, for example, conditional promoters, inducible promoters, constitutive promoters, or tissue-specific promoters. Optionally, the promoter can be a bidirectional promoter that drives expression of both the Cas protein in one direction and the guide RNA in the other direction. Such a bidirectional promoter can consist of (1) a complete conventional unidirectional Pol III promoter containing three external control elements: a distal sequence element (DSE), a proximal sequence element (PSE), and a TATA box; and (2) a second basic Pol III promoter containing a PSE and a TATA box fused in reverse orientation to the 5' end of the DSE. For example, in the H1 promoter, the DSE is adjacent to the PSE and TATA box, and the promoter can be made bidirectional by adding a PSE and a TATA box from the U6 promoter to create a hybrid promoter in which transcription is controlled in reverse orientation. See, for example, U.S. Patent Application Publication No. 2016 / 0074535, which is incorporated herein by reference in its entirety for all purposes. Using a bidirectional promoter to simultaneously express genes encoding Cas proteins and guide RNAs allows for the generation of compact expression cassettes for easy delivery. In a preferred embodiment, the promoter is approved by regulatory authorities for use in humans.In certain embodiments, the promoter drives expression in liver cells.
[0223] Various promoters can be used to drive Cas or Cas9 expression. In some methods, small promoters are used so that the Cas or Cas9 coding sequence can fit into an AAV construct. For example, Cas or Cas9 and one or more gRNAs (e.g., one gRNA, two gRNAs, three gRNAs, or four gRNAs) can be delivered via LNP-mediated delivery (e.g., in the form of RNA) or adeno-associated virus (AAV)-mediated delivery (e.g., AAV8-mediated delivery). For example, the nuclease agent can be CRISPR / Cas9, and Cas9 mRNA and gRNA (e.g., targeting the human L-SH5, L-SH18, or L-SH20 genomic safe harbor locus described herein) can be delivered via LNP-mediated delivery or AAV-mediated delivery. For example, the nuclease agent may be CRISPR / Cas9, and Cas9 mRNA and gRNA (e.g., targeting the mouse L-SH5, L-SH18, or L-SH20 genomic safe harbor loci described herein) may be delivered via LNP-mediated delivery or AAV-mediated delivery. The Cas or Cas9 and gRNA(s) may be delivered in a single AAV or via two separate AAVs. For example, a first AAV may carry a Cas or Cas9 expression cassette, and a second AAV may carry a gRNA expression cassette. Similarly, a first AAV may carry a Cas or Cas9 expression cassette, and a second AAV may carry two or more gRNA expression cassettes. Alternatively, a single AAV may carry a Cas or Cas9 expression cassette (e.g., a Cas or Cas9 coding sequence operably linked to a promoter) and a gRNA expression cassette (e.g., a gRNA coding sequence operably linked to a promoter). Similarly, a single AAV can carry a Cas or Cas9 expression cassette (e.g., a Cas or Cas9 coding sequence operably linked to a promoter) and two or more gRNA expression cassettes (e.g., a gRNA coding sequence operably linked to a promoter). Various promoters can be used to drive the expression of the gRNA, such as the U6 promoter or the small tRNA Gln. Similarly, various promoters can be used to drive the expression of Cas9.For example, small promoters are used so that the Cas9 coding sequence can fit into the AAV construct, as well as small Cas9 proteins (e.g., SaCas9 or CjCas9 are used to maximize AAV packaging capacity).
[0224] Cas proteins provided as mRNA can be modified to improve stability and / or immunogenicity. Modifications can be made to one or more nucleosides within the mRNA. The mRNA encoding the Cas protein can also be capped. The Cas mRNA can further comprise a polyadenylation (polyA or poly(A) or polyadenine) tail. For example, the Cas mRNA can comprise modifications to one or more nucleosides within the mRNA, the Cas mRNA can be capped, or the Cas mRNA can comprise a poly(A) tail.
[0225] (3) Guide RNA A "guide RNA" or "gRNA" is an RNA molecule that binds to a Cas protein (e.g., a Cas9 protein) and targets the Cas protein to a specific location within a target DNA. A guide RNA may contain two segments: a "DNA-targeting segment" (also called a "guide sequence") and a "protein-binding segment." A "segment" comprises a section or region of a molecule, e.g., a contiguous stretch of nucleotides in an RNA. Some gRNAs, such as those for Cas9, may contain two separate RNA molecules: an "activator RNA" (e.g., tracrRNA) and a "targeter RNA" (e.g., CRISPR RNA or crRNA). Other gRNAs are single RNA molecules (single RNA polynucleotides), which may also be called "single-molecule gRNA," "single guide RNA," or "sgRNA." See, for example, International Publication Nos. WO 2013 / 176772, WO 2014 / 065596, WO 2014 / 089290, WO 2014 / 093622, WO 2014 / 099750, WO 2013 / 142578, and WO 2014 / 131833, each of which is incorporated by reference in its entirety for all purposes. Guide RNA can refer to either CRISPR RNA (crRNA) or a combination of crRNA and trans-activating CRISPR RNA (tracrRNA). The crRNA and tracrRNA can be associated as a single RNA molecule (single guide RNA or sgRNA) or in two separate RNA molecules (dual guide RNA or dgRNA). For example, in the case of Cas9, a single guide RNA can include a crRNA fused to a tracrRNA (e.g., via a linker). For example, in the case of Cpfl and CasΦ, only the crRNA is required to achieve binding to the target sequence. The terms "guide RNA" and "gRNA" include both double-molecule (i.e., modular) gRNAs and single-molecule gRNAs. In some of the methods and compositions disclosed herein, the gRNA is an S. pyogenes Cas9 gRNA or its equivalent.In some of the methods and compositions disclosed herein, the gRNA is a S. aureus Cas9 gRNA or an equivalent thereof.
[0226] Exemplary bimolecular gRNAs include a crRNA-like ("CRISPR RNA" or "targeter RNA" or "crRNA" or "crRNA repeat") molecule and a corresponding tracrRNA-like ("trans-activating CRISPR RNA" or "activator RNA" or "tracrRNA") molecule. The crRNA comprises both the DNA-targeting segment (single strand) of the gRNA and a stretch of nucleotides that forms one half of the dsRNA duplex of the protein-binding segment of the gRNA. An example of a crRNA tail (e.g., for use with S. pyogenes Cas9) located downstream (3') of the DNA-targeting segment comprises, consists essentially of, or consists of GUUUUAGAGCUAUGCU (SEQ ID NO: 6) or GUUUUAGAGCUAUGCUGUUUUG (SEQ ID NO: 7). Any of the DNA-targeting segments disclosed herein can be attached to the 5' end of SEQ ID NO: 6 or 7 to form a crRNA.
[0227] The corresponding tracrRNA (activator-RNA) contains a stretch of nucleotides that forms the other half of the dsRNA duplex of the protein-binding segment of the gRNA. The stretch of nucleotides in the crRNA is complementary to and hybridizes with the stretch of nucleotides in the tracrRNA, forming the dsRNA duplex of the protein-binding domain of the gRNA. Thus, each crRNA can be said to have a corresponding tracrRNA. Exemplary tracrRNA sequences (e.g., for use with S. pyogenes Cas9) comprise, consist essentially of, or consist of any one of AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUU (SEQ ID NO: 8), AAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (SEQ ID NO: 9), or GUUGGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 10).
[0228] In systems where both a crRNA and a tracrRNA are required, the crRNA and the corresponding tracrRNA hybridize to form a gRNA. In systems where only a crRNA is required, the crRNA can be the gRNA. The crRNA additionally provides a single-stranded DNA targeting segment that hybridizes to the complementary strand of the target DNA. When used for intracellular modification, the exact sequence of a given crRNA or tracrRNA molecule can be designed to be specific for the species in which the RNA molecule is used. See, e.g., Mali et al. (2013) Science 339(6121):823-826, Jinek et al. (2012) Science 337(6096):816-821, Hwang et al. (2013) Nat. Biotechnol. 31(3):227-229, Jiang et al. (2013) Nat. Biotechnol. 31(3):233-239, and Cong et al. (2013) Science 339(6121):819:823, each of which is incorporated by reference in its entirety for all purposes.
[0229] The DNA-targeting segment (crRNA) of a given gRNA contains a nucleotide sequence that is complementary to a sequence on the complementary strand of the target DNA, as described in more detail below. The DNA-targeting segment of the gRNA interacts with the target DNA in a sequence-specific manner through hybridization (i.e., base pairing). Thus, the nucleotide sequence of the DNA-targeting segment can be varied to determine the location within the target DNA where the gRNA and target DNA interact. The DNA-targeting segment of a given gRNA can be modified to hybridize to any desired sequence within the target DNA. Naturally occurring crRNAs vary depending on the CRISPR / Cas system and organism, but often contain a targeting segment that is 21-72 nucleotides long, flanked by two direct repeats (DRs) that are 21-46 nucleotides long (see, e.g., WO 2014 / 131833, incorporated by reference in its entirety for all purposes). In S. pyogenes, the DRs are 36 nucleotides long, and the targeting segment is 30 nucleotides long. The 3'-located DR is complementary to and hybridizes with the corresponding tracrRNA, which then binds to the Cas protein.
[0230] A DNA-targeting segment can have a length of, for example, at least about 12, at least about 15, at least about 17, at least about 18, at least about 19, at least about 20, at least about 25, at least about 30, at least about 35, or at least about 40 nucleotides. Such a DNA-targeting segment can have a length of, for example, about 12 to about 100, about 12 to about 80, about 12 to about 50, about 12 to about 40, about 12 to about 30, about 12 to about 25, or about 12 to about 20 nucleotides. For example, a DNA-targeting segment can be about 15 to about 25 nucleotides (e.g., about 17 to about 20 nucleotides, or about 17, 18, 19, or 20 nucleotides). See, e.g., U.S. Patent Application Publication No. 2016 / 0024523, incorporated herein by reference in its entirety for all purposes. For Cas9 derived from S. pyogenes, a typical DNA-targeting segment is 16-20 nucleotides or 17-20 nucleotides in length. For Cas9 derived from S. aureus, a typical DNA-targeting segment is 21-23 nucleotides in length. For Cpf1, a typical DNA-targeting segment is at least 16 nucleotides in length or at least 18 nucleotides in length.
[0231] In one example, the DNA-targeting segment can be about 20 nucleotides in length. However, shorter and longer sequences can also be used for the targeting segment (e.g., 15-25 nucleotides in length, such as 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length). The degree of identity between the DNA-targeting segment and the corresponding guide RNA target sequence (or the degree of complementarity between the DNA-targeting segment and the other strand of the guide RNA target sequence) can be, for example, about 75%, about 80%, about 85%, about 90%, or 100%. The DNA-targeting segment and the corresponding guide RNA target sequence can contain one or more mismatches. For example, the DNA-targeting segment of the guide RNA and the corresponding guide RNA target sequence can contain 1-4, 1-3, 1-2, 1, 2, 3, or 4 mismatches (e.g., the total length of the guide RNA target sequence is at least 17, at least 18, at least 19, or at least 20 or more nucleotides). For example, the DNA-targeting segment of a guide RNA and the corresponding guide RNA target sequence can contain 1 to 4, 1 to 3, 1 to 2, 1, 2, 3, or 4 mismatches, with the total length of the guide RNA target sequence being 20 nucleotides.
[0232] As an example, a guide RNA that targets a genomic safe harbor locus described herein may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25-27, 45-47, and 228-314. Alternatively, a guide RNA that targets a genomic safe harbor locus described herein may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25-27, 45-47, and 228-314. Alternatively, guide RNAs targeting genomic safe harbor loci described herein may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25-27, 45-47, and 228-314. Alternatively, guide RNAs targeting genomic safe harbor loci described herein may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25-27, 45-47, and 228-314. Alternatively, guide RNAs targeting genomic safe harbor loci described herein can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence set forth in any one of SEQ ID NOs: 25-27, 45-47, and 228-314 (the DNA-targeting segment).Alternatively, guide RNAs targeting genomic safe harbor loci described herein can comprise a DNA-targeting segment that is at least 90%, or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence set forth in any one of SEQ ID NOs: 25-27, 45-47, and 228-314 (the DNA-targeting segment). Alternatively, guide RNAs targeting genomic safe harbor loci described herein can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from a sequence set forth in any one of SEQ ID NOs: 25-27, 45-47, and 228-314 (the DNA-targeting segment). Alternatively, guide RNAs targeting genomic safe harbor loci described herein can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence set forth in any one of SEQ ID NOs: 25-27, 45-47, and 228-314 (the DNA-targeting segment).
[0233] As an example, a guide RNA targeted to a genomic safe harbor locus described herein may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 315-404. Alternatively, a guide RNA targeted to a genomic safe harbor locus described herein may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 315-404. Alternatively, a guide RNA targeted to a genomic safe harbor locus described herein may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 315-404. Alternatively, guide RNAs targeting genomic safe harbor loci described herein may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to a sequence set forth in any one of SEQ ID NOs: 315-404 (the DNA-targeting segment). Alternatively, guide RNAs targeting genomic safe harbor loci described herein may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of a sequence set forth in any one of SEQ ID NOs: 315-404 (the DNA-targeting segment). Alternatively, guide RNAs targeting genomic safe harbor loci described herein may comprise a DNA-targeting segment that is at least 90% or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of a sequence set forth in any one of SEQ ID NOs: 315-404 (the DNA-targeting segment).Alternatively, guide RNAs targeting genomic safe harbor loci described herein can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from a sequence set forth in any one of SEQ ID NOs: 315-404 (DNA-targeting segment). Alternatively, guide RNAs targeting genomic safe harbor loci described herein can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides (DNA-targeting segment) of a sequence set forth in any one of SEQ ID NOs: 315-404.
[0234] As an example, a guide RNA that targets a genomic safe harbor locus described herein can comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25-27 and 45-47. As an example, a guide RNA that targets a genomic safe harbor locus described herein can comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25-27. Alternatively, a guide RNA that targets a genomic safe harbor locus described herein can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides (DNA-targeting segment) of a sequence set forth in any one of SEQ ID NOs: 25-27 and 45-47. Alternatively, guide RNAs targeting genomic safe harbor loci described herein may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to a sequence set forth in any one of SEQ ID NOS: 25-27 and 45-47 (the DNA-targeting segment). Alternatively, guide RNAs targeting genomic safe harbor loci described herein may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to a sequence set forth in any one of SEQ ID NOS: 25-27 and 45-47 (the DNA-targeting segment). Alternatively, guide RNAs targeting genomic safe harbor loci described herein may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence set forth in any one of SEQ ID NOS: 25-27 and 45-47 (the DNA-targeting segment).Alternatively, guide RNAs targeting genomic safe harbor loci described herein can comprise a DNA-targeting segment that is at least 90% or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence set forth in any one of SEQ ID NOs: 25-27 and 45-47 (the DNA-targeting segment). Alternatively, guide RNAs targeting genomic safe harbor loci described herein can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from a sequence set forth in any one of SEQ ID NOs: 25-27 and 45-47 (the DNA-targeting segment). Alternatively, guide RNAs targeting genomic safe harbor loci described herein can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence set forth in any one of SEQ ID NOs: 25-27, and 45-47 (the DNA-targeting segment).
[0235] As another example, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25, 45, and 228-256. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25, 45, and 228-256. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25, 45, and 228-256. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25, 45, and 228-256. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25, 45, and 228-256.Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides to the sequence set forth in any one of SEQ ID NOs: 25, 45, and 228-256 (the DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence set forth in any one of SEQ ID NOs: 25, 45, and 228-256 (the DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence set forth in any one of SEQ ID NOs: 25, 45, and 228-256 (the DNA-targeting segment), differing by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide.
[0236] As another example, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25, 45, 235, 237, and 246. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25, 45, 235, 237, and 246. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25, 45, 235, 237, and 246. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25, 45, 235, 237, and 246. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of a sequence set forth in any one of SEQ ID NOs: 25, 45, 235, 237, and 246 (the DNA-targeting segment).Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25, 45, 235, 237, and 246. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 25, 45, 235, 237, and 246. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence set forth in any one of SEQ ID NOs: 25, 45, 235, 237, and 246 (the DNA-targeting segment), differing by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide.
[0237] As another example, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of the sequence (DNA-targeting segment) set forth in SEQ ID NO: 25 or 45. As another example, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of the sequence (DNA-targeting segment) set forth in SEQ ID NO: 25. As another example, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of the sequence (DNA-targeting segment) set forth in SEQ ID NO: 45. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in SEQ ID NO: 25 or 45. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in SEQ ID NO: 25. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in SEQ ID NO: 45 (DNA-targeting segment).Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequence (DNA-targeting segment) set forth in SEQ ID NO: 25 or 45. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequence (DNA-targeting segment) set forth in SEQ ID NO: 25. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequence set forth in SEQ ID NO: 45 (DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to the sequence set forth in SEQ ID NO: 25 or 45 (DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to the sequence set forth in SEQ ID NO: 25 (DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to the sequence set forth in SEQ ID NO: 45 (DNA-targeting segment).Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides to the sequence (DNA-targeting segment) set forth in SEQ ID NO: 25 or 45. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides to the sequence (DNA-targeting segment) set forth in SEQ ID NO: 25. Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides to the sequence set forth in SEQ ID NO: 45 (the DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides to the sequence set forth in SEQ ID NO: 25 or 45 (the DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in SEQ ID NO: 25 (DNA-targeting segment).Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in SEQ ID NO: 45 (the DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence set forth in SEQ ID NO: 25 or 45 (the DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence set forth in SEQ ID NO: 25 (DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence set forth in SEQ ID NO: 45 (DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in SEQ ID NO: 25 or 45 (the DNA-targeting segment), differing by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide.Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in SEQ ID NO: 25 (the DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH5 (chromosome 13, coordinates 77460242-77460537) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in SEQ ID NO: 45 (the DNA-targeting segment).
[0238] As another example, a guide RNA targeting mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 315-344. Alternatively, a guide RNA targeting mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 315-344. Alternatively, a guide RNA targeting mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 315-344. Alternatively, a guide RNA targeting mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 315-344. Alternatively, a guide RNA targeting mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 315-344.Alternatively, a guide RNA targeting mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 315-344 (the DNA-targeting segment). Alternatively, a guide RNA targeting mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence set forth in any one of SEQ ID NOs: 315-344 (the DNA-targeting segment). Alternatively, a guide RNA targeting mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence set forth in any one of SEQ ID NOs: 315-344 (the DNA-targeting segment), differing by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide.
[0239] As another example, a guide RNA targeting mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 318, 320, 321, and 341. Alternatively, a guide RNA targeting mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 318, 320, 321, and 341. Alternatively, a guide RNA targeting mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 318, 320, 321, and 341. Alternatively, a guide RNA targeting mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 318, 320, 321, and 341. Alternatively, a guide RNA targeting mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of a sequence set forth in any one of SEQ ID NOs: 318, 320, 321, and 341 (the DNA-targeting segment).Alternatively, a guide RNA targeting mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 318, 320, 321, and 341. Alternatively, a guide RNA targeting mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 318, 320, 321, and 341. Alternatively, a guide RNA targeting mouse L-SH5 (chromosome 14, coordinates 103,450,397-103,451,396) can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence set forth in any one of SEQ ID NOs: 318, 320, 321, and 341 (the DNA-targeting segment), differing by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide.
[0240] As another example, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 26, 46, and 257-285. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 26, 46, and 257-285. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 26, 46, and 257-285. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 26, 46, and 257-285. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 26, 46, and 257-285.Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 26, 46, and 257-285. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 26, 46, and 257-285. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence set forth in any one of SEQ ID NOs: 26, 46, and 257-285 (the DNA-targeting segment), differing by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide.
[0241] As another example, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 26, 46, 268, 271, and 280. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 26, 46, 268, 271, and 280. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 26, 46, 268, 271, and 280. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 26, 46, 268, 271, and 280. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 26, 46, 268, 271, and 280.Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 26, 46, 268, 271, and 280. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 26, 46, 268, 271, and 280. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence set forth in any one of SEQ ID NOs: 26, 46, 268, 271, and 280 (the DNA-targeting segment), differing by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide.
[0242] As another example, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of the sequence (DNA-targeting segment) set forth in SEQ ID NO: 26 or 46. As another example, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of the sequence (DNA-targeting segment) set forth in SEQ ID NO: 26. As another example, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of the sequence (DNA-targeting segment) set forth in SEQ ID NO: 46. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in SEQ ID NO: 26 or 46. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in SEQ ID NO: 26. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in SEQ ID NO: 46 (DNA-targeting segment).Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequence (DNA-targeting segment) set forth in SEQ ID NO: 26 or 46. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequence (DNA-targeting segment) set forth in SEQ ID NO: 26. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequence set forth in SEQ ID NO: 46 (DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to the sequence set forth in SEQ ID NO: 26 or 46 (DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to the sequence set forth in SEQ ID NO: 26 (DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to the sequence set forth in SEQ ID NO: 46 (DNA-targeting segment).Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides to the sequence (DNA-targeting segment) set forth in SEQ ID NO: 26 or 46. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides to the sequence (DNA-targeting segment) set forth in SEQ ID NO: 26. Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides to the sequence set forth in SEQ ID NO: 46 (the DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides to the sequence set forth in SEQ ID NO: 26 or 46 (the DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in SEQ ID NO: 26 (DNA-targeting segment).Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence set forth in SEQ ID NO: 46 (the DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence set forth in SEQ ID NO: 26 or 46 (the DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence set forth in SEQ ID NO: 26 (DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence set forth in SEQ ID NO: 46 (DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in SEQ ID NO: 26 or 46 (the DNA-targeting segment), differing by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide.Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in SEQ ID NO: 26 (the DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH18 (chromosome 6, coordinates 170031084-170031382) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in SEQ ID NO: 46 (the DNA-targeting segment).
[0243] As another example, a guide RNA targeting mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 345-374. Alternatively, a guide RNA targeting mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 345-374. Alternatively, a guide RNA targeting mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 345-374. Alternatively, a guide RNA targeting mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386) may comprise a DNA-targeting segment that is at least 90%, or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 345-374. Alternatively, a guide RNA targeting mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 345-374.Alternatively, a guide RNA targeting mouse L-SH5 (chromosome 17, coordinates 15,226,387-15,227,386) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 345-374. Alternatively, a guide RNA targeting mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 345-374. Alternatively, a guide RNA targeting mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence set forth in any one of SEQ ID NOs: 345-374 (the DNA-targeting segment), differing by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide.
[0244] As another example, a guide RNA targeting mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 347, 360, 369, and 370. Alternatively, a guide RNA targeting mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 347, 360, 369, and 370. Alternatively, a guide RNA targeting mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 347, 360, 369, and 370. Alternatively, a guide RNA targeting mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 347, 360, 369, and 370. Alternatively, a guide RNA targeting mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of a sequence set forth in any one of SEQ ID NOs: 347, 360, 369, and 370 (the DNA-targeting segment).Alternatively, a guide RNA targeting mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 347, 360, 369, and 370. Alternatively, a guide RNA targeting mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 347, 360, 369, and 370. Alternatively, a guide RNA targeting mouse L-SH18 (chromosome 17, coordinates 15,226,387-15,227,386) can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence set forth in any one of SEQ ID NOs: 347, 360, 369, and 370 (the DNA-targeting segment), differing by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide.
[0245] As another example, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 27, 47, and 286-314. Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 27, 47, and 286-314. Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 27, 47, and 286-314. Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 27, 47, and 286-314. Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 27, 47, and 286-314.Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 27, 47, and 286-314 (the DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence set forth in any one of SEQ ID NOs: 27, 47, and 286-314 (the DNA-targeting segment). Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence set forth in any one of SEQ ID NOs: 27, 47, and 286-314 (the DNA-targeting segment), differing by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide.
[0246] As another example, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 27, 47, 288, 296, 305, 306, and 310. Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 27, 47, 288, 296, 305, 306, and 310. Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 27, 47, 288, 296, 305, 306, and 310. Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 27, 47, 288, 296, 305, 306, and 310. Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 27, 47, 288, 296, 305, 306, and 310.Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment that is at least 90% or at least 95% identical for at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 27, 47, 288, 296, 305, 306, and 310. Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 27, 47, 288, 296, 305, 306, and 310. Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a sequence set forth in any one of SEQ ID NOs: 27, 47, 288, 296, 305, 306, and 310 (the DNA-targeting segment), differing by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide.
[0247] As another example, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of the sequence (DNA-targeting segment) set forth in SEQ ID NO: 27 or 47. As another example, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of the sequence (DNA-targeting segment) set forth in SEQ ID NO: 27. As another example, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of the sequence (DNA-targeting segment) set forth in SEQ ID NO: 47. Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in SEQ ID NO: 27 or 47. Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in SEQ ID NO: 27. Alternatively, a guide RNA targeting human L-SH20 (chromosome 9, coordinates 25207412-25207703) may comprise a DNA-targeting segment com...
Claims
1. A method for incorporating a nucleic acid construct into a genome-safe harbor locus in a human cell or for expressing a target product therefrom, wherein the human cell contains (a) A nuclease agent or one or more nucleic acids encoding the nuclease agent, wherein the nuclease agent targets a nuclease target site in the genome-safe harbor locus, and the genome-safe harbor locus is located at the following genomic location, i.e., (i) Genomic coordinates of approximately 77460242 to 77460537 on human chromosome 13, (ii) Genomic coordinates of approximately 170031084 to approximately 170031382 on human chromosome 6, and (iii) A nuclease agent or one or more nucleic acids encoding the nuclease agent, selected from genomic coordinates approximately 25207412 to approximately 25207703 on human chromosome 9, and (b) The nucleic acid construct comprising a nucleic acid operably linked to a promoter, wherein the nucleic acid encodes the product of the object of the object, and the administration of the nucleic acid construct The nuclease agent cleaves the nuclease target site, and the nucleic acid construct is inserted into the genome-safe harbor locus. The aforementioned human cells are liver cells, The method involves the use of human cells in vitro or ex vivo.
2. The method according to claim 1, wherein the human cells are liver cells.
3. The method according to claim 1, wherein the genome-safe harbor locus is human chromosome 13, coordinates 77460242 to 77460537, or includes or consists of the sequence described in Sequence ID No.
39.
4. The method according to claim 1, wherein the genome-safe harbor locus is on human chromosome 6, coordinates 170031084 to 170031382, or includes or consists of the sequence described in Sequence ID No.
40.
5. The method according to claim 1, wherein the genome-safe harbor locus is on human chromosome 9, coordinates 25207412 to 25207703, or includes or consists of the sequence described in Sequence ID No.
41.
6. The nuclease agent is (a) Zinc finger nuclease (ZFN), (b) Transcription activator-like effector nuclease (TALEN), or (c) (i) Cas protein or nucleic acid encoding the said Cas protein, and (ii) The method according to any one of claims 1 to 5, comprising a guide RNA or one or more DNAs encoding the guide RNA, wherein the guide RNA includes a DNA targeting segment that targets a guide RNA target sequence, and the guide RNA binds to the Cas protein, causing the Cas protein to target the guide RNA target sequence.
7. The nuclease agent is (a) Cas9 protein or nucleic acid encoding the Cas9 protein, (b) The method according to any one of claims 1 to 5, comprising: a guide RNA or one or more DNAs encoding the guide RNA, wherein the guide RNA includes a DNA targeting segment that targets a guide RNA target sequence, and the guide RNA binds to the Cas9 protein, causing the Cas9 protein to target the guide RNA target sequence;
8. The method according to claim 7, wherein the Cas9 protein is derived from Streptococcus pyogenes Cas9 protein.
9. The genome-safe harbor locus is on human chromosome 13, coordinates 77460242 to 77460537, or includes or comprises the sequence described in Sequence ID No.
39. (I) The DNA targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence described in any one of SEQ ID NOs. 25, 45, and 228-256, and / or (II) The DNA targeting segment is at least 90% or at least 95% identical to the sequence described in any one of SEQ ID NOs. 25, 45, and 228-256, and / or (III) The DNA targeting segment includes one of SEQ ID NOs. 25, 45, and 228-256, and / or (IV) The method according to claim 7, wherein the DNA targeting segment comprises one of sequence numbers 25, 45, and 228-256.
10. The method according to claim 9, wherein the DNA targeting segment includes or consists of SEQ ID NO:
25.
11. The genome-safe harbor locus is located on human chromosome 6, coordinates 170031084 to 170031382, or includes or comprises the sequence described in Sequence ID No.
40. (I) The DNA targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence described in any one of SEQ ID NOs: 26, 46, and 257-285, and / or (II) The DNA targeting segment is at least 90% or at least 95% identical to the sequence described in any one of SEQ ID NOs: 26, 46, and 257-285, and / or (III) The DNA targeting segment includes one of SEQ ID NOs: 26, 46, and 257-285, and / or (IV) The method according to claim 7, wherein the DNA targeting segment comprises one of sequence numbers 26, 46, and 257-285.
12. The method according to claim 11, wherein the DNA targeting segment includes or consists of SEQ ID NO:
26.
13. The genome-safe harbor locus is on human chromosome 9, coordinates 25207412 to 25207703, or includes or comprises the sequence described in Sequence ID No.
41. (I) The DNA targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence described in any one of SEQ ID NOs: 27, 47, and 286-314, and / or (II) The DNA targeting segment is at least 90% or at least 95% identical to the sequence described in any one of SEQ ID NOs: 27, 47, and 286-314, and / or (III) The DNA targeting segment comprises one of sequence numbers 27, 47, and 286-314, and / or (IV) The method according to claim 7, wherein the DNA targeting segment comprises one of sequence numbers 27, 47, and 286-314.
14. The method according to claim 13, wherein the DNA targeting segment includes or consists of SEQ ID NO:
27.
15. A composition for use in a method of incorporating a nucleic acid construct into a genome-safe harbor gene locus in human cells of human interest, or expressing a product of interest therefrom, The composition comprises a nuclease agent or one or more nucleic acids encoding the nuclease agent, The above method applies to the human subject, (a) The nuclease agent or one or more nucleic acids encoding the nuclease agent, wherein the nuclease agent targets a nuclease target site in the genome-safe harbor locus, and the genome-safe harbor locus is located at the following genomic location, i.e., (i) Genomic coordinates of approximately 77460242 to 77460537 on human chromosome 13, (ii) Genomic coordinates of approximately 170031084 to approximately 170031382 on human chromosome 6, and (iii) The nuclease agent or one or more nucleic acids encoding the nuclease agent, selected from the genomic coordinates of approximately 25207412 to approximately 25207703 on human chromosome 9, and (b) The nucleic acid construct comprising a nucleic acid operably linked to a promoter, wherein the nucleic acid encodes the product of the object of the object, and the administration of the nucleic acid construct The nuclease agent cleaves the nuclease target site, and the nucleic acid construct is inserted into the genome-safe harbor locus. The human cells mentioned above are liver cells, and the composition is for use.
16. The composition for use according to claim 15, wherein the human cells are hepatocytes.
17. The composition for use according to claim 15, wherein the genome-safe harbor locus is human chromosome 13, coordinates 77460242 to 77460537, or comprises or consists of the sequence described in Sequence ID No.
39.
18. The composition for use according to claim 15, wherein the genome-safe harbor locus is on human chromosome 6, coordinates 170031084 to 170031382, or comprises or consists of the sequence described in Sequence ID No.
40.
19. The composition for use according to claim 15, wherein the genome-safe harbor locus is on human chromosome 9, coordinates 25207412 to 25207703, or comprises or consists of the sequence described in Sequence ID No.
41.
20. The aforementioned nuclease agent is (a) Zinc finger nuclease (ZFN), (b) Transcription activator-like effector nuclease (TALEN), or (c) (i) Cas protein or nucleic acid encoding the said Cas protein, and (ii) A composition for use according to any one of claims 15 to 19, comprising a guide RNA or one or more DNAs encoding the guide RNA, wherein the guide RNA includes a DNA targeting segment that targets a guide RNA target sequence, and the guide RNA binds to the Cas protein, causing the Cas protein to target the guide RNA target sequence.
21. The aforementioned nuclease agent is (a) Cas9 protein or nucleic acid encoding the Cas9 protein, (b) A composition for use according to any one of claims 15 to 19, comprising: a guide RNA or one or more DNAs encoding the guide RNA, wherein the guide RNA includes a DNA targeting segment that targets a guide RNA target sequence, and the guide RNA binds to the Cas9 protein, causing the Cas9 protein to target the guide RNA target sequence.
22. The composition for use according to claim 21, wherein the Cas9 protein is derived from the Streptococcus pyogenes Cas9 protein.
23. The genome-safe harbor locus is on human chromosome 13, coordinates 77460242 to 77460537, or includes or comprises the sequence described in Sequence ID No.
39. (I) The DNA targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence described in any one of SEQ ID NOs. 25, 45, and 228-256, and / or (II) The DNA targeting segment is at least 90% or at least 95% identical to the sequence described in any one of SEQ ID NOs. 25, 45, and 228-256, and / or (III) The DNA targeting segment includes one of SEQ ID NOs. 25, 45, and 228-256, and / or (IV) The composition for use according to claim 21, wherein the DNA targeting segment comprises one of sequence numbers 25, 45, and 228-256.
24. The composition for use according to claim 23, wherein the DNA targeting segment comprises or consists of Sequence ID No.
25.
25. The genome-safe harbor locus is located on human chromosome 6, coordinates 170031084 to 170031382, or includes or comprises the sequence described in Sequence ID No.
40. (I) The DNA targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence described in any one of SEQ ID NOs: 26, 46, and 257-285, and / or (II) The DNA targeting segment is at least 90% or at least 95% identical to the sequence described in any one of SEQ ID NOs: 26, 46, and 257-285, and / or (III) The DNA targeting segment includes one of SEQ ID NOs: 26, 46, and 257-285, and / or (IV) The composition for use according to claim 21, wherein the DNA targeting segment comprises one of sequence numbers 26, 46, and 257-285.
26. The composition for use according to claim 25, wherein the DNA targeting segment comprises or consists of SEQ ID NO:
26.
27. The genome-safe harbor locus is on human chromosome 9, coordinates 25207412 to 25207703, or includes or comprises the sequence described in Sequence ID No.
41. (I) The DNA targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence described in any one of SEQ ID NOs: 27, 47, and 286-314, and / or (II) The DNA targeting segment is at least 90% or at least 95% identical to the sequence described in any one of SEQ ID NOs: 27, 47, and 286-314, and / or (III) The DNA targeting segment comprises one of sequence numbers 27, 47, and 286-314, and / or (IV) The composition for use according to claim 21, wherein the DNA targeting segment comprises one of sequence numbers 27, 47, and 286-314.
28. The composition for use according to claim 27, wherein the DNA targeting segment comprises or consists of Sequence ID No.
27.
29. Human cells containing nucleic acid constructs integrated into genome-safe harbor loci, The nucleic acid construct comprises a nucleic acid operably linked to a promoter, the nucleic acid encoding the desired product, The aforementioned genome-safe harbor locus is located at the following genomic location, namely, (i) Genomic coordinates of approximately 77460242 to 77460537 on human chromosome 13, (ii) Genomic coordinates of approximately 170031084 to approximately 170031382 on human chromosome 6, and (iii) Selected from approximately 25207412 to 25207703 genomic coordinates on human chromosome 9, The aforementioned human cells are liver cells, human cells.
30. The cell according to claim 29, wherein the human cell is a liver cell.
31. A composition comprising guide RNA or DNA encoding guide RNA, wherein the guide RNA comprises a DNA targeting segment that targets a guide RNA target sequence within a genome-safe harbor locus and a protein-binding segment that binds to the Cas9 protein, and the genome-safe harbor locus is located at the following genomic location, i.e., (i) Genomic coordinates of approximately 77460242 to 77460537 on human chromosome 13, (ii) Genomic coordinates of approximately 170031084 to approximately 170031382 on human chromosome 6, and (iii) A composition selected from genomic coordinates approximately 25207412 to 25207703 on human chromosome 9.
32. The composition according to claim 31, wherein the Cas9 protein is derived from the Streptococcus pyogenes Cas9 protein.
33. The genome-safe harbor locus is located on human chromosome 13, coordinates 77460242-77460537, or contains or comprises the sequence described in Sequence ID No.
39. (I) The DNA targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence described in any one of SEQ ID NOs. 25, 45, and 228-256, and / or (II) The DNA targeting segment is at least 90% or at least 95% identical to the sequence described in any one of SEQ ID NOs. 25, 45, and 228-256, and / or (III) The DNA targeting segment includes one of SEQ ID NOs. 25, 45, and 228-256, and / or (IV) The composition according to claim 31, wherein the DNA targeting segment comprises one of sequence numbers 25, 45, and 228-256.
34. The composition according to claim 33, wherein the DNA targeting segment includes or consists of Sequence ID No.
25.
35. The genome-safe harbor locus is located on human chromosome 6, coordinates 170031084-170031382, or contains or consists of the sequence described in Sequence ID No.
40. (I) The DNA targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence described in any one of SEQ ID NOs: 26, 46, and 257-285, and / or (II) The DNA targeting segment is at least 90% or at least 95% identical to the sequence described in any one of SEQ ID NOs: 26, 46, and 257-285, and / or (III) The DNA targeting segment includes one of SEQ ID NOs: 26, 46, and 257-285, and / or (IV) The composition according to claim 31, wherein the DNA targeting segment comprises one of sequence numbers 26, 46, and 257-285.
36. The composition according to claim 35, wherein the DNA targeting segment includes or consists of SEQ ID NO:
26.
37. The genome-safe harbor locus is located on human chromosome 9, coordinates 25207412-25207703, or contains or comprises the sequence described in Sequence ID No.
41. (I) The DNA targeting segment comprises at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence described in any one of SEQ ID NOs: 27, 47, and 286-314, and / or (II) The DNA targeting segment is at least 90% or at least 95% identical to the sequence described in any one of SEQ ID NOs: 27, 47, and 286-314, and / or (III) The DNA targeting segment comprises one of sequence numbers 27, 47, and 286-314, and / or (IV) The composition according to claim 31, wherein the DNA targeting segment comprises one of sequence numbers 27, 47, and 286 to 314.
38. The composition according to claim 37, wherein the DNA targeting segment includes or consists of Sequence ID No.
27.
39. The composition according to any one of claims 31 to 38, further comprising a nucleic acid construct, wherein the nucleic acid construct comprises a nucleic acid operably linked to a promoter, and the nucleic acid codes for the product of the choice.