Composition and method for improved gene editing

By using a base editing enzyme and specific guide polynucleotides to introduce site-specific mutations and selectively enrich edited cells with a cytotoxic agent, the method addresses the inefficiencies and errors in current base editing technologies, achieving improved editing efficiency and accuracy.

JP2025081589APending Publication Date: 2025-05-27ASTRAZENECA AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025026734
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-04-12
Filing Date
2025-02-21
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Current base editing technologies face challenges with low efficiency and high error rates, particularly in achieving biallelic integration and gene silencing, which are essential for therapeutic applications in genetic diseases.

Method used

The method involves introducing a base editing enzyme into a cell population, along with specific guide polynucleotides that target both a gene encoding a cytotoxic agent receptor and a target polynucleotide. This enables site-specific mutations to be introduced, and by using a cytotoxic agent, the edited cells can be selectively enriched, thereby improving editing efficiency and accuracy.

Benefits of technology

This approach significantly enhances the efficiency of site-specific mutations and allows for the precise selection of edited cells, thereby overcoming the limitations of existing base editing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025081589000034
    Figure 2025081589000034
  • Figure 2025081589000035
    Figure 2025081589000035
  • Figure 2025081589000036
    Figure 2025081589000036
Patent Text Reader

Abstract

To provide a method for introducing site-specific mutation into a target cell, and a method for determining effectiveness of enzyme which can be introduced.SOLUTION: A method for introducing site-specific mutation in a target polynucleotide includes: (a) introducing into a cell population a base editor, a first guide polynucleotide which hybridizes with a gene encoding a cytotoxic agent receptor to form a first composite together with the base editor and provides mutation in the gene encoding a cytotoxic agent receptor, and a second guide polynucleotide which hybridizes with a target polynucleotide to form a second composite together with the base editor and provides mutation in the target polynucleotide; (b) bringing the cell population in contact with the cytotoxic agent; and (c) enriching a target cell including the mutation in the target polynucleotide by selecting a cytotoxic agent resistant cell from the cell population.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure provides a method for introducing a site-specific mutation into a target cell and a method for determining the efficacy of an enzyme capable of introducing a site-specific mutation. The present disclosure also provides a method for providing a two-allele sequence integration, a method for integrating a target sequence into a locus of a cell genome, and a method for introducing a stable episomal vector into a cell. The present disclosure further provides a method for generating human cells resistant to diphtheria toxin.

Background Art

[0002] Targeted nucleic acid modification by programmable site-specific nucleases such as zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), and RNA-guided Cas9 is a very promising approach for studying gene function and has great potential for providing new therapeutics for genetic diseases. Typically, a programmable nuclease generates a double-strand break (DSB) in a target sequence. Subsequently, the DSB may be repaired by mutation via the non-homologous end joining (NHEJ) pathway, or the DNA around the cleavage site may be replaced with a template simultaneously introduced via the homologous recombination repair (HDR) pathway. For an overview of targeted nucleic acid modification, see, for example, Non-Patent Document 1, Non-Patent Document 2, and Non-Patent Document 3.

[0003] Disadvantages of relying on NHEJ and HDR include, for example, the low efficiency of HDR and the undesirable off-target activity by NHEJ. The low efficiency of HDR makes the selection of accurate on-target modification particularly difficult (see, for example, Non-Patent Document 1, Non-Patent Document 4, Non-Patent Document 5). Various attempts to boost HDR over NHEJ include, for example, generating one or more single-strand nicks in the target DNA rather than DSB (see, for example, Non-Patent Document 6, Non-Patent Document 7). However, in the art, there remains a need for improved selection of HDR events, for example, when biallelic integration or gene silencing typically achieved by an HDR template is desired.

[0004] HDR is less error-prone than NHEJ, but still HDR tends to generate unwanted modifications that compete with the targeted modification. Therefore, base editing has emerged in recent years as a powerful and accurate gene editing technology that promotes single base pair substitutions at specific positions in the genome. Compared with HDR-based methods for site-specific modification, base editing provides a more effective method for introducing single nucleotide mutations and overcomes some of the limitations associated with HDR. Base editing involves site-specific modification of single DNA bases, along with manipulation of the natural DNA repair machinery to avoid faithful repair of the modified base. Base editors are typically chimeric proteins that include a DNA targeting module and a catalytic domain capable of deaminating, for example, cytosine bases to thymine or adenine bases to guanine. For example, the DNA targeting module can be based on catalytically inactive Cas9 (dCas9) or a Cas9 nickase variant (Cas9n) guided by a guide RNA molecule (sgRNA or gRNA). The catalytic domain can be a cytidine deaminase or an adenine deaminase. It is not necessary to generate DSBs to edit DNA bases and limit the generation of insertions and deletions (indels) at the target site and off-target sites. Therefore, base editing does not depend on the cellular HDR machinery and is thus more efficient than HDR and has fewer inaccurate modifications by NHEJ. Modified base editing systems are described, for example, in Non-Patent Documents 8, 9, 10, and 11. For an overview of base editing, see, for example, Non-Patent Documents 12, 13, and 14.

Prior Art Documents

Non-Patent Documents

[0005]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Non-Patent Document 4

Non-Patent Document 5

Non-Patent Document 6

Non-Patent Document 7

Non-Patent Document 8

Non-Patent Document 9

Non-Patent Document 10

Non-Patent Document 11

Non-Patent Document 12

Non-Patent Document 13

Non-Patent Document 14

Summary of the Invention

Problems to be Solved by the Invention

[0006] Since many genetic diseases can be caused by changes in specific nucleotides at specific positions in the genome (e.g., a change from C to T in a specific codon of a gene associated with the disease), base editing can serve as a promising therapeutic approach for treating hereditary disorders based on single nucleotide variants. However, despite improvements beyond conventional CRISPR / Cas9 editing, base editing efficiency still remains low to moderate and, in addition, is plagued by mismatches across the genome. Accordingly, there remains a need in the art for more efficient and improved base editing systems.

[0007] Various publications, the entire disclosures of which are incorporated herein by reference, are cited herein.

Means for Solving the Problems

[0008] In some embodiments, the present disclosure provides a method for introducing a site-specific mutation into a target polynucleotide in a target cell in a cell population, the method comprising: (a) introducing into the cell population: (i) a base editing enzyme; (ii) a first guide polynucleotide that (1) hybridizes to a gene encoding a cytotoxic agent (CA) receptor and (2) forms a first complex with the base editing enzyme, wherein the base editing enzyme of the first complex provides a mutation in the gene encoding the CA receptor, and the mutation in the gene encoding the CA receptor forms CA-resistant cells in the cell population; and (iii) a second guide polynucleotide that (1) hybridizes to the target polynucleotide and (2) forms a second complex with the base editing enzyme, wherein the base editing enzyme of the second complex provides a mutation in the target polynucleotide; (b) contacting the cell population with CA; and (c) enriching for target cells containing a mutation in the target polynucleotide by selecting CA-resistant cells from the cell population.

[0009] In some embodiments, the present disclosure provides a method for determining the potency of a base editing enzyme in a cell population, comprising: (a) introducing into the cell population: (i) a base editing enzyme; (ii) a first guide polynucleotide that hybridizes to (1) a gene encoding a cytotoxic agent (CA) receptor and (2) forms a first complex with the base editing enzyme, wherein the base editing enzyme of the first complex introduces a mutation in the gene encoding the CA receptor, and the mutation in the gene encoding the CA receptor forms CA-resistant cells in the cell population; and (iii) a second guide polynucleotide that hybridizes to (1) a target polynucleotide and (2) forms a second complex with the base editing enzyme, wherein the base editing enzyme of the second complex introduces a mutation in the target polynucleotide; (b) contacting the cell population with CA to isolate CA-resistant cells; and (c) determining the potency of the base editing enzyme by determining the ratio of CA-resistant cells to the total cell population.

[0010] In some embodiments, the base editing enzyme comprises a DNA targeting domain and a DNA editing domain.

[0011] In some embodiments, the DNA targeting domain comprises Cas9. In some embodiments, Cas9 comprises a mutation in the catalytic domain. In some embodiments, the base editing enzyme comprises catalytically inactive Cas9 and a DNA editing domain. In some embodiments, the base editing enzyme comprises Cas9 (nCas9) capable of generating a single-stranded DNA break and a DNA editing domain. In some embodiments, nCas9 comprises a mutation relative to wild-type Cas9 at amino acid residue D10 or H840 (numbering is relative to SEQ ID NO: 3). In some embodiments, Cas9 is at least 90% identical to SEQ ID NO: 3 or 4.

[0012] In some embodiments, the DNA editing domain comprises a deaminase. In some embodiments, the deaminase is a cytidine deaminase or an adenosine deaminase. In some embodiments, the deaminase is a cytidine deaminase. In some embodiments, the deaminase is an adenosine deaminase. In some embodiments, the deaminase is an apolipoprotein B mRNA editing complex (APOBEC) deaminase, activation-induced cytidine deaminase (AID), ACF1 / ASE deaminase, ADAT deaminase, or ADAR deaminase. In some embodiments, the deaminase is an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the deaminase is APOBEC1.

[0013] In some embodiments, the base editing enzyme further comprises a DNA glycosylase inhibitor domain. In some embodiments, the DNA glycosylase is an uracil DNA glycosylase inhibitor (UGI). In some embodiments, the base editing enzyme comprises nCas9 and a cytidine deaminase. In some embodiments, the base editing enzyme comprises nCas9 and an adenosine deaminase. In some embodiments, the base editing enzyme comprises a polypeptide sequence that is at least 90% identical to SEQ ID NO: 6. In some embodiments, the base editing enzyme is BE3.

[0014] In some embodiments, the first and / or second guide polynucleotide is an RNA polynucleotide. In some embodiments, the first and / or second guide polynucleotide further comprises a tracrRNA sequence.

[0015] In some embodiments, the cell population is a human cell.

[0016] In some embodiments, the mutation in the gene encoding the CA receptor is a point mutation from cytidine (C) to thymine (T). In some embodiments, the mutation in the gene encoding the CA receptor is a point mutation from adenine (A) to guanine (G).

[0017] In some embodiments, the CA is diphtheria toxin. In some embodiments, the cytotoxic agent (CA) receptor is the receptor for diphtheria toxin. In some embodiments, the CA receptor is heparin-binding EGF-like growth factor (HB-EGF). In some embodiments, HB-EGF comprises the polypeptide sequence of SEQ ID NO: 8.

[0018] In some embodiments, the base editing enzyme of the first complex provides a mutation at one or more of amino acids 107-148 of HB-EGF. In some embodiments, the base editing enzyme of the first complex provides a mutation at one or more of amino acids 138-144 of HB-EGF. In some embodiments, the base editing enzyme of the first complex provides a mutation at amino acid 141 of HB-EGF. In some embodiments, the base editing enzyme of the first complex provides a mutation from GLU141 to LYS141 in the amino acid sequence of HB-EGF.

[0019] In some embodiments, the base editing enzyme of the first complex provides a mutation in a region of HB-EGF that binds diphtheria toxin. In some embodiments, the base editing enzyme of the first complex provides a mutation in HB-EGF that makes the target cell resistant to diphtheria toxin. In some embodiments, the mutation in the target polynucleotide is a point mutation from cytosine (C) to thymine (T) in the target polynucleotide. In some embodiments, the mutation in the target polynucleotide is a point mutation from adenine (A) to guanine (G) in the target polynucleotide.

[0020] In some embodiments, the base editing enzyme is introduced into the cell population as a polynucleotide encoding the base editing enzyme. In some embodiments, the polynucleotide encoding the base editing enzyme, the first guide polynucleotide of (ii), and the second guide polynucleotide of (iii) are on a single vector. In some embodiments, the polynucleotide encoding the base editing enzyme, the first guide polynucleotide of (ii), and the second guide polynucleotide of (iii) are on one or more vectors. In some embodiments, the vector is a viral vector. In some embodiments, the viral vector is an adenovirus, a lentivirus, or an adeno-associated virus.

[0021] In some embodiments, the present disclosure provides a method for providing a biallelic integration of a sequence of interest (SOI) at a toxin-sensitive gene (TSG) locus in the genome of a cell, the method comprising: (a) introducing into a cell population (i) a nuclease capable of generating a double-strand break, (ii) a guide polynucleotide capable of forming a complex with the nuclease and hybridizing to the TSG locus, and (iii) a donor polynucleotide comprising (1) a 5' homology arm, a 3' homology arm, and a mutation in the native coding sequence of the TSG, the mutation conferring resistance to a toxin, and (2) the SOI, wherein as a result of introducing (i), (ii), and (iii), the donor polynucleotide is integrated into the TSG locus; (b) contacting the cell population with a toxin; and (c) selecting one or more cells resistant to the toxin, wherein the one or more cells resistant to the toxin comprise a biallelic integration of the SOI.

[0022] In some embodiments, the donor polynucleotide is integrated by homologous recombination repair (HDR). In some embodiments, the donor polynucleotide is integrated by non-homologous end joining (NHEJ).

[0023] In some embodiments, the TSG locus comprises introns and exons. In some embodiments, the donor polynucleotide further comprises a splicing acceptor sequence. In some embodiments, the nuclease capable of generating a double-strand break generates the break within an intron. In some embodiments, the mutation in the native coding sequence of the TSG is in an exon of the TSG locus.

[0024] In some embodiments, the present disclosure provides a method of integrating a sequence of interest (SOI) into a target locus in the genome of a cell, the method comprising: (a) introducing into a cell population (i) a nuclease capable of generating a double-strand break, (ii) a guide polynucleotide capable of forming a complex with the nuclease and hybridizing to a toxin-sensitive gene (TSG) locus in the genome of the cell, wherein the TSG is an essential gene, and (iii) a donor polynucleotide comprising (1) a functional TSG gene comprising a mutation in the native coding sequence of the TSG, the mutation conferring resistance to a toxin, (2) the SOI, and (3) a sequence for genomic integration at the target locus, wherein introduction of (i), (ii), and (iii) results in inactivation of the TSG in the genome of the cell by the nuclease and integration of the donor polynucleotide at the target locus; (b) contacting the cell population with a toxin; and (c) selecting one or more cells resistant to the toxin, wherein the one or more cells resistant to the toxin comprise the SOI integrated at the target locus.

[0025] In some embodiments, the sequence for genomic integration is obtained from a transposon or a retroviral vector.

[0026] In some embodiments, the functional TSG of the donor polynucleotide or episomal vector is resistant to nuclease inactivation. In some embodiments, mutations in the native coding sequence of the TSG remove the protospacer adjacent motif from the native coding sequence. In some embodiments, the guide polynucleotide cannot hybridize to the functional TSG of the donor polynucleotide or episomal vector.

[0027] In some embodiments, the nuclease capable of generating a double-strand break is Cas9. In some embodiments, Cas9 is capable of generating sticky ends. In some embodiments, Cas9 comprises the polypeptide sequence of SEQ ID NO: 3 or 4.

[0028] In some embodiments, the guide polynucleotide is an RNA polynucleotide. In some embodiments, the guide polynucleotide further comprises a tracrRNA sequence.

[0029] In some embodiments, the donor polynucleotide is a vector. In some embodiments, the mutation in the native coding sequence of the TSG is a substitution mutation, an insertion, or a deletion. In some embodiments, the mutation in the native coding sequence of the TSG is a mutation in the toxin-binding region of the protein encoded by the TSG. In some embodiments, the TSG locus comprises a gene encoding heparin-binding EGF-like growth factor (HB-EGF). In some embodiments, the TSG encodes HB-EGF (SEQ ID NO: 8).

[0030] In some embodiments, the mutation in the native coding sequence of the TSG is a mutation in one or more of amino acids 107-148 of HB-EGF (SEQ ID NO: 8). In some embodiments, the mutation in the native coding sequence of the TSG is a mutation in one or more of amino acids 138-144 of HB-EGF (SEQ ID NO: 8). In some embodiments, the mutation in the native coding sequence of the TSG is a mutation at amino acid 141 of HB-EGF (SEQ ID NO: 8). In some embodiments, the mutation in the native coding sequence of the TSG is a mutation from GLU141 to LYS141 of HB-EGF (SEQ ID NO: 8).

[0031] In some embodiments, the toxin is diphtheria toxin. In some embodiments, the mutation in the native coding sequence of the TSG confers resistance to diphtheria toxin on the cell. In some embodiments, the toxin is an antibody-drug conjugate and the TSG encodes a receptor for the antibody-drug conjugate.

[0032] In some embodiments, the present disclosure provides a method of conferring resistance to diphtheria toxin in human cells, the method comprising introducing into the cells (i) a base editing enzyme and (ii) a guide polynucleotide that targets the heparin-binding EGF-like growth factor (HB-EGF) receptor in human cells, wherein the base editing enzyme forms a complex with the guide polynucleotide, the base editing enzyme is targeted to HB-EGF, and the base editing enzyme provides site-specific mutations in HB-EGF to confer resistance to diphtheria toxin in human cells.

[0033] In some embodiments, the base editing enzyme comprises a DNA targeting domain and a DNA editing domain.

[0034] In some embodiments, the DNA targeting domain comprises Cas9. In some embodiments, Cas9 comprises a mutation in the catalytic domain. In some embodiments, the base editing enzyme comprises catalytically inactive Cas9 and a DNA editing domain. In some embodiments, the base editing enzyme comprises Cas9 (nCas9) capable of generating a single-stranded DNA break and a DNA editing domain. In some embodiments, nCas9 comprises a mutation relative to wild-type Cas9 at amino acid residue D10 or H840 (numbering is relative to SEQ ID NO: 3). In some embodiments, Cas9 is at least 90% identical to SEQ ID NO: 3 or 4.

[0035] In some embodiments, the DNA editing domain comprises a deaminase. In some embodiments, the deaminase is selected from cytidine deaminase and adenosine deaminase. In some embodiments, the deaminase is cytidine deaminase. In some embodiments, the deaminase is adenosine deaminase. In some embodiments, the deaminase is selected from apolipoprotein B mRNA editing complex (APOBEC) deaminases, activation-induced cytidine deaminase (AID), ACF1 / ASE deaminase, ADAT deaminases, and TadA deaminases. In some embodiments, the deaminase is an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the cytidine deaminase is APOBEC1. In some embodiments, the base editing enzyme further comprises a DNA glycosylase inhibitor domain. In some embodiments, the DNA glycosylase is uracil DNA glycosylase inhibitor (UGI).

[0036] In some embodiments, the base editing enzyme comprises nCas9 and cytidine deaminase. In some embodiments, the base editing enzyme comprises nCas9 and adenosine deaminase. In some embodiments, the base editing enzyme comprises a polypeptide sequence that is at least 90% identical to SEQ ID NO: 6. In some embodiments, the base editing enzyme is BE3.

[0037] In some embodiments, the guide polynucleotide is an RNA polynucleotide. In some embodiments, the guide polynucleotide further comprises a tracrRNA sequence.

[0038] In some embodiments, the site-specific mutation is in one or more of amino acids 107-148 of HB-EGF (SEQ ID NO: 8). In some embodiments, the site-specific mutation is in one or more of amino acids 138-144 of HB-EGF (SEQ ID NO: 8). In some embodiments, the site-specific mutation is at amino acid 141 of HB-EGF (SEQ ID NO: 8). In some embodiments, the site-specific mutation is a mutation from GLU141 to LYS141 of HB-EGF (SEQ ID NO: 8). In some embodiments, the site-specific mutation is in the region of HB-EGF that binds to diphtheria toxin.

[0039] In some embodiments, the present disclosure provides a method for integrating and enriching a target sequence (SOI) into a target locus in the genome of a cell, comprising: (a) introducing into a cell population (i) a nuclease capable of generating a double-strand break, (ii) a guide polynucleotide capable of forming a complex with the nuclease and hybridizing to an essential gene (ExG) locus in the genome of the cell, and (iii) a donor polynucleotide comprising (1) a functional ExG gene comprising a mutation in the native coding sequence of ExG, the mutation conferring resistance to inactivation by the guide polynucleotide, (2) the SOI, and (3) a sequence for genomic integration at the target locus, wherein introducing (i), (ii), and (iii) results in inactivation of ExG in the genome of the cell by the nuclease and integration of the donor polynucleotide at the target locus; (b) culturing the cells; and (c) selecting one or more surviving cells, wherein the one or more surviving cells comprise the SOI integrated at the target locus.

[0040] In some embodiments, the present disclosure provides a method of introducing a stable episomal vector into a cell, the method comprising: (a) introducing into a cell population: (i) a nuclease capable of generating a double-strand break; (ii) a guide polynucleotide capable of forming a complex with the nuclease and hybridizing to an essential gene (ExG) locus in the genome of the cell, wherein introduction of (i) and (ii) results in inactivation of the ExG in the genome of the cell by the nuclease; and (iii) an episomal vector comprising: (1) a functional ExG comprising a mutation in the native coding sequence of the ExG, the mutation conferring resistance to inactivation by the nuclease; and (2) an autonomous DNA replication sequence; (b) culturing the cell; and (c) selecting one or more viable cells, wherein the one or more viable cells comprise the episomal vector.

[0041] In some embodiments, the mutation in the native coding sequence of the ExG removes the protospacer adjacent motif from the native coding sequence. In some embodiments, the guide polynucleotide is unable to hybridize to the functional ExG of the donor polynucleotide or episomal vector.

[0042] In some embodiments, the nuclease capable of generating a double-strand break is Cas9. In some embodiments, Cas9 is capable of generating sticky ends. In some embodiments, Cas9 comprises the polypeptide sequence of SEQ ID NO: 3 or 4.

[0043] In some embodiments, the guide polynucleotide is an RNA polynucleotide. In some embodiments, the guide polynucleotide further comprises a tracrRNA sequence.

[0044] In some embodiments, the donor polynucleotide is a vector. In some embodiments, the mutation in the native coding sequence of the ExG is a substitution mutation, an insertion, or a deletion.

[0045] In some embodiments, the sequences for genomic integration are obtained from transposons or retroviral vectors. In some embodiments, the episomal vector is an artificial chromosome or a plasmid.

[0046] In some embodiments, two or more guide polynucleotides are introduced into a cell population, each guide polynucleotide forms a complex with a nuclease, and each guide polynucleotide hybridizes to a different region of ExG.

[0047] In some embodiments, the method further comprises introducing the nuclease of (a)(i) and the guide polynucleotide of (a)(ii) into a living cell to enrich for living cells containing the SOI integrated into the target locus. In some embodiments, the method further comprises introducing the nuclease of (a)(i) and the guide polynucleotide of (a)(ii) into a living cell to enrich for living cells containing an episomal vector. In some embodiments, the nuclease of (a)(i) and the guide polynucleotide of (a)(ii) are introduced into living cells for multiple enrichment rounds. BRIEF DESCRIPTION OF THE DRAWINGS

[0048]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Modes for Carrying Out the Invention

[0049] The present disclosure provides methods for introducing site-specific mutations into target cells and for determining the efficacy of enzymes capable of introducing site-specific mutations. The present disclosure also provides methods for providing integration of sequences of two alleles, for integrating a target sequence into a genomic locus of a cell, and for introducing a stable episomal vector into a cell. The present disclosure further provides a method for generating human cells resistant to diphtheria toxin.

[0050] Definitions As used herein, "a" or "an" can mean one or more. When used in this specification and the claims herein and in conjunction with the word "comprising", the word "a" or "an" can mean one or two or more. As used herein, "another" or "further" can mean at least a second or more.

[0051] Throughout this application, the term "about" is used to indicate that a value includes the inherent variability of error for the method / device used to determine that value, or the variability that exists between test subjects. Typically, this term means that, depending on the situation, it encompasses variability of approximately or less than 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19% or 20%.

[0052] The use of the term "or" in the claims is used to mean "and / or" unless explicitly stated to refer only to alternatives or unless the alternatives are mutually exclusive, but this disclosure supports definitions that refer only to alternatives and "and / or".

[0053] As used in this specification and the claims, the term "comprising" (and any form of "comprising", such as "comprise" and "comprises"), "having" (and any form of "having", such as "have" and "has"), "including" (and any form of "including", such as "includes" and "include") or "containing" (and any form of "containing", such as "contains" and "contain") is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. Any embodiment contemplated herein is intended to be executable with respect to any method, system, host cell, expression vector and / or composition of the present disclosure. Further, the methods and proteins of the present disclosure can be obtained using the compositions, systems, host cells and / or vectors of the present disclosure.

[0054] The use of the term "for example" and its corresponding abbreviation "e.g." (whether italicized or not) means that the recited term is representative examples and embodiments of the present disclosure that are not limited to the specifically recited examples or references unless otherwise expressly stated.

[0055] "Nucleic acid", "nucleic acid molecule", "nucleotide", "nucleotide sequence", "oligonucleotide", or "polynucleotide" means a polymeric compound containing covalently bonded nucleotides. The term "nucleic acid" includes ribonucleic acid (RNA) and deoxyribonucleic acid (DNA), both of which may be single-stranded or double-stranded. Examples of DNA include, but are not limited to, complementary DNA (cDNA), genomic DNA, plasmid or vector DNA, and synthetic DNA. In some embodiments, the present disclosure provides a polynucleotide encoding any one of the polypeptides disclosed herein, for example, targeting a polynucleotide encoding a Cas protein or a variant thereof.

[0056] "Gene" refers to a collection of nucleotides encoding a polypeptide, which includes cDNA and genomic DNA nucleic acid molecules. "Gene" also refers to a nucleic acid fragment that can act as regulatory sequences preceding (5' non-coding sequence) and following (3' non-coding sequence) the coding sequence.

[0057] A nucleic acid molecule is "hybridizable" or "hybridized" to another nucleic acid molecule, such as cDNA, genomic DNA, or RNA, when the single-stranded form of the nucleic acid molecule can anneal to the other nucleic acid molecule under appropriate conditions of temperature and solution ionic strength. Hybridization and washing conditions are known and are exemplified in Sambrook et al., Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (1989), particularly Chapter 11 and Table 11.1 therein. The conditions of temperature and ionic strength determine the "stringency" of hybridization. Stringency conditions can be adjusted to screen for moderately similar fragments, such as homologous sequences from distantly related organisms, against highly similar fragments, such as genes replicating functional enzymes from closely related organisms. For a preliminary screening of homologous nucleic acids, T at 55 °C mCorresponding low stringency hybridization conditions can be used, such as 5×SSC, 0.1% SDS, 0.25% milk, and no formamide; or 30% formamide, 5×SSC, 0.5% SDS. Medium stringency hybridization conditions correspond to a higher T m , for example, corresponding to 40% formamide and 5× or 6× SCC. High stringency hybridization conditions correspond to a maximum T m , for example, corresponding to 50% formamide, 5× or 6× SCC. Hybridization requires that two nucleic acids contain complementary sequences, but base mismatches are possible depending on the stringency of the hybridization.

[0058] The term "complementary" is used to describe the relationship between nucleotide bases that can hybridize to each other. For example, with respect to DNA, adenosine is complementary to thymine, and cytosine is complementary to guanine. Accordingly, the present disclosure also includes isolated nucleic acid fragments complementary to the complete sequences disclosed or used herein and nucleic acid sequences substantially similar thereto.

[0059] A DNA "coding sequence" is a double-stranded DNA sequence that, when placed under the control of appropriate regulatory sequences, is transcribed and translated into a polypeptide in cells, either in vitro or in vivo. "Suitable regulatory sequences" refer to nucleotide sequences that are located upstream (5' non-coding sequence), within, or downstream (3' non-coding sequence) of the coding sequence and affect transcription, RNA processing or stability, or translation of the associated coding sequence. Examples of regulatory sequences include promoters, translational leader sequences, introns, polyadenylation recognition sequences, RNA processing sites, effector binding sites, and stem-loop structures. The boundaries of the coding sequence are determined by the start codon at the 5' (amino) terminus and the translation stop codon at the 3' (carboxyl) terminus. Examples of coding sequences include, but are not limited to, prokaryotic sequences, cDNA from mRNA, genomic DNA sequences, and synthetic DNA sequences. When the coding sequence is intended for expression in eukaryotic cells, polyadenylation signals and transcription termination sequences are typically located 3' to the coding sequence.

[0060] A "native coding sequence" typically refers to the wild-type sequence in the genome, and a "native coding sequence" can also refer to a sequence that is substantially similar to the wild-type sequence, having, for example, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence similarity to the wild-type sequence.

[0061] An "open reading frame", abbreviated as ORF, means a nucleic acid sequence of a length that contains a translation initiation signal or start codon, such as ATG or AUG, and a stop codon and can potentially be translated into a polypeptide sequence, either DNA, cDNA, or RNA.

[0062] The term "homologous recombination" refers to the insertion of a foreign DNA sequence into another DNA molecule, for example, the insertion of a vector into a chromosome. Optionally, the vector targets a chromosomal site specific for homologous recombination. For specific homologous recombination, the vector typically contains a region of sufficient length of homology with the chromosomal sequence to allow complementary binding and integration of the vector into the chromosome. Longer regions of complementarity and a greater degree of sequence similarity can increase the efficiency of homologous recombination.

[0063] Polynucleotides can be propagated according to the disclosure herein using methods known in the art. Once suitable host systems and growth conditions are established, recombinant expression vectors can be propagated and prepared in large quantities. As described herein, expression vectors that can be used include, but are not limited to, the following vectors or derivatives thereof: human or animal viruses, such as vaccinia virus or adenovirus; insect viruses, such as baculovirus; yeast vectors; bacteriophage vectors (e.g., lambda); and plasmid and cosmid DNA vectors.

[0064] As used herein, "operably linked" means that a polynucleotide of interest, for example, a polynucleotide encoding a Cas9 protein, is linked to regulatory elements so as to enable expression of the polynucleotide sequence. In some embodiments, the regulatory element is a promoter. In some embodiments, the polynucleotide of interest is operably linked to a promoter on an expression vector.

[0065] As used herein, "promoter", "promoter sequence", or "promoter region" refers to a DNA regulatory region / sequence that can bind to RNA polymerase and participate in the initiation of transcription of downstream coding or non-coding sequences. In some examples of the present disclosure, the promoter sequence includes the transcription start site and extends upstream to include the minimum number of bases or elements used to initiate transcription at a detectable level above background. In some embodiments, the promoter sequence includes the transcription start site and a protein binding domain responsible for binding of RNA polymerase. Eukaryotic promoters often, but not always, contain a "TATA" box and a "CAT" box. Various promoters, such as inducible promoters, can be used to drive the various vectors of the present disclosure.

[0066] A "vector" is any means for cloning and / or transferring nucleic acids into a host cell. A vector can be a replicon that can attach another DNA segment and cause replication of the attached segment. A "replicon" is any genetic element (e.g., plasmid, phage, cosmid, chromosome, virus) that functions as an autonomous unit of DNA replication in vivo, i.e., can replicate under its own control. In some embodiments of the present disclosure, this vector is an episomal vector, i.e., a non-integrating extrachromosomal plasmid capable of autonomous replication. In some embodiments, the episomal vector contains an autonomous DNA replication sequence, i.e., a sequence that enables replication of the vector, typically including an origin of replication (OriP). In some embodiments, the autonomous DNA replication sequence is a scaffold / matrix attachment region (S / MAR). In some embodiments, the autonomous DNA replication sequence is a viral OriP. The episomal vector can be removed or lost from the cell population after many cell generations, e.g., by asymmetric segregation. In some embodiments, the episomal vector is a stable episomal vector that remains in the cell, i.e., does not disappear from the cell. In some embodiments, the episomal vector is an artificial chromosome or a plasmid. In some embodiments, the episomal vector contains an autonomous DNA replication sequence. Examples of episomal vectors used in genome engineering and gene therapy include the papovaviridae virus family, including simian virus 40 (SV40) and BK virus, bovine papillomavirus type 1 (BPV-1), Kaposi's sarcoma-associated herpesvirus (KSHV), and the Herpesviridae virus family, including Epstein-Barr virus (EBV), and are derived from the S / MAR region of the human interferon-β gene. In some embodiments, the episomal vector is an artificial chromosome. In some embodiments, the episomal vector is a minichromosome.Episomal vectors are further described, for example, in Van Craenenbroeck et al., Eur J Biochem 267:5665-5678(2000), and Lufino et al., Mol Ther 16(9):1525-1538(2008).

[0067] The term "vector" includes both viral and non-viral means for introducing nucleic acids into cells in vitro, ex vivo, or in vivo. Many vectors known in the art can be used for manipulating nucleic acids, such as for incorporating response elements and promoters into genes. Possible vectors include, for example, plasmids or modified viruses, including bacteriophages such as lambda derivatives, or plasmids such as PBR322 or pUC plasmid derivatives, or Bluescript vectors. For example, insertion of a DNA fragment corresponding to a response element and a promoter into a suitable vector can be achieved by ligating the appropriate DNA fragment into a selected vector having complementary sticky ends. Alternatively, the ends of the DNA molecule can be enzymatically modified, or any site can be generated by ligating a nucleotide sequence (linker) into the DNA ends. Such vectors can be modified to contain a selectable marker gene that provides for the selection of cells that have incorporated the marker into the cell genome. Such markers enable the identification and / or selection of host cells that have incorporated the marker and express the protein encoded by the marker.

[0068] Viral vectors, particularly retroviral vectors, are used in a wide range of gene delivery applications in cells and live animals. Viral vectors that can be used include, but are not limited to, retroviruses, adenoviruses, adeno-associated viruses, poxviruses, baculoviruses, vaccinia viruses, herpes simplex viruses, Epstein-Barr viruses, adenoviruses, geminiviruses, and calimoviruses vectors. Retroviral vectors emerged as tools for gene therapy by facilitating genomic insertion of a desired sequence. The retroviral genome (e.g., murine leukemia virus (MLV), feline leukemia virus (FLV), or any virus belonging to the Retroviridae viral family) contains long terminal repeat (LTR) sequences flanking the viral genes. When the host is infected with the virus, the LTRs are recognized by integrase, which integrates the viral genome into the host genome. A retroviral vector for targeted gene insertion has the desired sequence inserted between the LTRs instead of having any of the viral genes. The LTRs are recognized by integrase and integrate the desired sequence into the genome of the host cell. Further details regarding retroviral vectors can be found, for example, in Kurian et al., Mol Pathol 53(4):173-176, and Vargas et al., J Transl Med 14:288 (2016).

[0069] Non-viral vectors include, but are not limited to, plasmids, liposomes, charged lipids, DNA-protein complexes, and biopolymers. The vector may also include, in addition to the nucleic acid, one or more regulatory regions and / or selectable markers useful in the selection, measurement, and monitoring of nucleic acid transfer results (such as the tissue to be transferred, the duration of expression, etc.).

[0070] Transposons and transposable elements can be included on a vector. A transposon is a mobile genetic element that contains flanking repeat sequences recognized by a transposase, which subsequently removes the transposon from its locus in the genome and inserts it into another genomic locus (commonly referred to as the "cut-and-paste" mechanism). Transposons are engineered for genome engineering by flanking a desired sequence to be inserted with repeat sequences recognizable by a transposase. The repeat sequences can collectively be referred to as the "transposon sequence". In some embodiments, the transposon sequence and the desired sequence to be inserted are included on a vector, the transposon sequence is recognized by a transposase, and subsequently the desired sequence can be integrated into the genome by the transposase. Transposons are described, for example, in Pray, Nature Education 1(1):204, (2008), Vargas et al., J Transl Med 14:288 (2016), and VandenDriessche et al., Blood 114(8):1461-1468 (2009). Non-limiting examples of transposon sequences include sleeping beauty (SB), piggyBac (PB), and Tol2 transposons.

[0071] Vectors can be introduced into a desired host cell by known methods such as, but not limited to, transfection, transduction, cell fusion, and lipofection. Vectors can contain various regulatory elements such as a promoter. In some embodiments, vector design can be based on constructs designed by Mali et al., Nature Methods 10:957-63 (2013). In some embodiments, the present disclosure provides an expression vector comprising any of the polynucleotides described herein, for example, an expression vector comprising a polynucleotide encoding a Cas protein or a variant thereof. In some embodiments, the present disclosure provides an expression vector comprising a polynucleotide encoding Cas9 protein or a variant thereof.

[0072] The term "plasmid" often carries genes that are not part of the central metabolism of the cell and typically refers to an extrachromosomal element in the form of a circular double-stranded DNA molecule. Such an element can be an autonomously replicating sequence, a genomic integration sequence, a phage or nucleotide sequence, linear, circular, or supercoiled single-stranded or double-stranded DNA or RNA derived from any source, in which a number of nucleotide sequences are joined or recombined in a unique construct that can introduce into the cell a promoter fragment and a DNA sequence for a selected gene product along with an appropriate 3' untranslated sequence.

[0073] As used herein, "transfection" means the introduction of an exogenous nucleic acid molecule, such as a vector, into a cell. A "transfected" cell contains an exogenous nucleic acid molecule inside the cell, and a "transformed" cell is a cell in which the exogenous nucleic acid molecule inside the cell induces a phenotypic change in the cell. The transfected nucleic acid molecule can be integrated into the genomic DNA of the host cell and / or can be maintained extrachromosomally by the cell for a transient or long period of time. A host cell or organism that expresses an exogenous nucleic acid molecule or fragment is referred to as a "recombinant", "transformed", or "transgenic" organism. In some embodiments, the present disclosure provides a host cell comprising any of the expression vectors described herein, for example, an expression vector comprising a polynucleotide encoding a Cas protein or a variant thereof. In some embodiments, the present disclosure provides a host cell comprising an expression vector comprising a polynucleotide encoding Cas9 protein or a variant thereof.

[0074] The term "host cell" refers to a cell into which a recombinant expression vector has been introduced. The term "host cell" refers not only to the cell (the "parent" cell) into which the expression vector has been introduced, but also to the progeny of such a cell. For example, such progeny may not be identical to the parent cell because certain modifications may occur in later generations due to mutations or environmental influences, but are still included within the scope of the term "host cell".

[0075] The terms "peptide", "polypeptide", and "protein" are used interchangeably herein and refer to a polymeric form of amino acids of any length, including coded and non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having modified peptide backbones.

[0076] The starting point of a protein or polypeptide is known as the "N-terminus" (or amino terminus, NH 2 -terminus, N-terminal portion, or amine terminus), which refers to the free amine (-NH 2 ) group of the first amino acid residue of the protein or polypeptide. The end point of a protein or polypeptide is known as the "C-terminus" (or carboxy terminus, carboxyl terminus, C-terminal portion, or COOH terminus), which refers to the free carboxyl group (-COOH) of the last amino acid residue of the protein or polypeptide.

[0077] As used herein, "amino acid" refers to a compound containing both a carboxyl (-COOH) and an amino (-NH 2 ) group. "Amino acid" refers to both natural and non-natural (i.e., synthetic) amino acids. Natural amino acids having three-letter and one-letter abbreviations include alanine (Ala; A); arginine (Arg, R); asparagine (Asn; N); aspartic acid (Asp; D); cysteine (Cys; C); glutamine (Gln; Q); glutamic acid (Glu; E); glycine (Gly; G); histidine (His; H); isoleucine (Ile; I); leucine (Leu; L); lysine (Lys; K); methionine (Met; M); phenylalanine (Phe; F); proline (Pro; P); serine (Ser; S); threonine (Thr; T); tryptophan (Trp; W); tyrosine (Tyr; Y); and valine (Val; V).

[0078] "Amino acid substitution" refers to a polypeptide or protein that contains one or more substitutions at an amino acid residue of a wild-type or naturally occurring amino acid with a different amino acid relative to that wild-type or naturally occurring amino acid. The substituted amino acid can be a synthetic or naturally occurring amino acid. In some embodiments, the substituted amino acid is a naturally occurring amino acid selected from the group consisting of A, R, N, D, C, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, and V. Substitution variants can be described using an abbreviation system. For example, a substitution variant in which the 5th amino acid residue is substituted can be abbreviated as "X5Y", where "X" is the wild-type or naturally occurring amino acid to be replaced, "5" is the amino acid residue position within the amino acid sequence of the protein or polypeptide, and "Y" is the substitution, or non-wild-type or non-naturally occurring amino acid.

[0079] An "isolated" polypeptide, protein, peptide, or nucleic acid is a molecule that has been removed from its natural environment. It should also be understood that an "isolated" polypeptide, protein, peptide, or nucleic acid can be formulated with excipients, such as diluents or adjuvants, and still be considered isolated.

[0080] The term "recombinant", when used with respect to a nucleic acid molecule, peptide, polypeptide, or protein, means those that are of a new combination of genetic material not known to occur naturally, or those that result from such a combination. Recombinant molecules can be produced by any of the well-known techniques available in the field of recombinant technology, by way of example and not limitation, polymerase chain reaction (PCR), gene splicing (e.g., using restriction endonucleases), and solid-phase synthesis of nucleic acid molecules, peptides, or proteins.

[0081] When used with respect to a polypeptide or protein, the term "domain" means a distinct functional and / or structural unit within the protein. A domain may be responsible for a particular function or interaction that contributes to the overall role of the protein. Domains can exist in a variety of biological contexts. Similar domains can be found in proteins with different functions. Alternatively, domains with low sequence identity (i.e., less than about 50%, less than about 40%, less than about 30%, less than about 20%, less than about 10%, less than about 5%, or less than about 1% sequence identity) can have the same function. In some embodiments, the DNA targeting domain is Cas9, or a Cas9 domain. In some embodiments, the Cas9 domain is the RuvC domain. In some embodiments, the Cas9 domain is the HNH domain. In some embodiments, the Cas9 domain is the Rec domain. In some embodiments, the DNA editing domain is a deaminase, or a deaminase domain.

[0082] When used with respect to a polypeptide or protein, the term "motif" generally refers to a set of conserved amino acid residues typically less than 20 amino acids in length that can be important for the function of the protein. Specific sequence motifs can mediate general functions such as binding or targeting of the protein to a particular intracellular location in various proteins. Examples of motifs include, but are not limited to, nuclear localization signals, microbody targeting motifs, motifs that prevent or promote secretion, and motifs that facilitate recognition and binding of proteins. Motif databases and / or motif search tools are known to those of skill in the art and include, for example, PROSITE (expasy.ch / sprot / prosite.html), Pfam (pfam.wustl.edu), PRINTS (biochem.ucl.ac.uk / bsm / dbbrowser / PRINTS / PRINTS.html), and Minimotif Miner (cse-mnm.engr.uconn.edu:8080 / MNM / SMSSearchServlet).

[0083] As used herein, the term "modified" protein means a protein that contains one or more modifications to achieve a desired property in the protein. Exemplary modifications include, but are not limited to, insertions, deletions, substitutions, or fusions with another domain or protein. Modified proteins of the present disclosure include modified Cas9 proteins.

[0084] In some embodiments, the modified protein is generated from a wild-type protein. As used herein, a "wild-type" protein or nucleic acid is a natural, unmodified protein or nucleic acid. For example, a wild-type Cas9 protein can be isolated from the organism Streptococcus pyogenes. Wild-type is contrasted with "mutant", which contains one or more modifications in the amino acid and / or nucleotide sequence of the protein or nucleic acid.

[0085] The term "sequence similarity" or "% similarity" as used herein refers to the degree of identity or correspondence between nucleic acid sequences or between amino acid sequences. "Sequence similarity" as used herein refers to a nucleic acid sequence in which one or more nucleotide base changes result in one or more amino acid substitutions but do not affect the functional properties of the protein encoded by the DNA sequence. "Sequence similarity" also refers to modifications of nucleic acids that do not substantially affect the functional properties of the resulting transcript, for example, deletions or insertions of one or more nucleotide bases. Thus, it is understood that the present disclosure encompasses more than specific exemplary sequences. Methods for making nucleotide base substitutions are known, as are methods for determining retention of the biological activity of the encoded product.

[0086] Furthermore, one of ordinary skill in the art will recognize that similar sequences encompassed by the present disclosure are also defined by their ability to hybridize to the sequences exemplified herein under stringent conditions. Similar nucleic acid sequences of the present disclosure are nucleic acids in which the DNA sequence is at least 70%, at least 80%, at least 90%, at least 95% or at least 99% identical to the DNA sequence of the nucleic acids disclosed herein. Similar nucleic acid sequences of the present disclosure are nucleic acids in which the DNA sequence is about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 99%, at least about 99% or about 100% identical to the DNA sequence of the nucleic acids disclosed herein.

[0087] As used herein, "sequence similarity" refers to two or more amino acid sequences in which more than about 40% of the amino acids are identical, or more than about 60% of the amino acids are functionally identical. Functionally identical or functionally similar amino acids have chemically similar side chains. For example, amino acids can be grouped as follows according to functional similarity: Positively charged side chains: Arg, His, Lys; Negatively charged side chains: Asp, Glu; Polar uncharged side chains: Ser, Thr, Asn, Gln; Hydrophobic side chains: Ala, Val, Ile, Leu, Met, Phe, Tyr, Trp; Others: Cys, Gly, Pro.

[0088] In some embodiments, similar amino acid sequences of the present disclosure have amino acids that are at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90% or at least 99% identical.

[0089] In some embodiments, similar amino acid sequences of the present disclosure have amino acids that are at least 60%, at least 70%, at least 80%, at least 90% or at least 95% functionally identical. In some embodiments, similar amino acid sequences of the present disclosure have amino acids that are about 40%, at least about 40%, about 45%, at least about 45%, about 50%, at least about 50%, about 55%, at least about 55%, about 60%, at least about 60%, about 65%, at least about 65%, about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 97%, at least about 97%, about 98%, at least about 98%, about 99%, at least about 99% or about 100% identical.

[0090] In some embodiments, similar amino acid sequences of the present disclosure have amino acids that are about 60%, at least about 60%, about 65%, at least about 65%, about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 97%, at least about 97%, about 98%, at least about 98%, about 99%, at least about 99% or about 100% functionally identical.

[0091] As used herein, the term "same protein" refers to a protein having a structure or amino acid sequence that is substantially similar to a reference protein and that performs the same biochemical function as the reference protein, and includes proteins that differ from the reference protein by substitution or deletion of one or more amino acids at one or more positions in the amino acid sequence, i.e., deletion of amino acids that are at least about 60%, at least about 60%, about 65%, at least about 65%, about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 97%, at least about 97%, about 98%, at least about 98%, about 99%, at least about 99%, or about 100% identical. In one aspect, "same protein" refers to a protein having the same amino acid sequence as the reference protein.

[0092] Sequence similarity can be determined by sequence alignment using standard methods in the art, such as BLAST, MUSCLE, Clustal (e.g., ClustalW and ClustalX), and T-Coffee (including variants such as M-Coffee, R-Coffee, and Expresso).

[0093] The terms "sequence identity" or "% identity" with respect to a nucleic acid sequence or an amino acid sequence refer to the percentage of residues in a comparison sequence that are identical when the sequences are aligned over a defined comparison window. In some embodiments, only a particular portion of two or more sequences is aligned to determine sequence identity. In some embodiments, only a particular domain of two or more sequences is aligned to determine sequence similarity. The comparison window can be a segment of at least 10 to over 1000 residues, at least 20 to about 1000 residues, or at least 50 to 500 residues that can be aligned and compared. Methods of alignment for determining sequence identity are well-known and can be performed using publicly available databases such as BLAST. "Percent identity" or "% identity", when referring to an amino acid sequence, can be determined by methods known in the art. For example, in some embodiments, the "percent identity" of two amino acid sequences is determined using the algorithm of Karlin and Altschul, Proc Nat Acad Sci USA 87:2264-2268(1990) as modified in Karlin and Altschul, Proc Nat Acad Sci USA 90:5873-5877(1993). Such algorithms are incorporated into the BLAST programs, such as BLAST+ or NBLAST and XBLAST programs described in Altschul et al., Journal of Molecular Biology, 215:403-410(1990). A BLAST protein search can be performed using programs such as the XBLAST program, score = 50, word length = 3, etc., to obtain an amino acid sequence that is homologous to the protein molecules of the present disclosure. When there are gaps between two sequences, Gapped BLAST can be utilized as described in Altschul et al., Nucleic Acids Research 25(17):3389-3402(1997). When using the BLAST and Gapped BLAST programs, the default parameters of each program (e.g., XBLAST and NBLAST) can be used.

[0094] In some embodiments, the polypeptide or nucleic acid molecule has 70%, at least 70%, 75%, at least 75%, 80%, at least 80%, 85%, at least 85%, 90%, at least 90%, 95%, at least 95%, 97%, at least 97%, 98%, at least 98%, 99%, at least 99% or 100% sequence identity with a reference polypeptide or nucleic acid molecule (or a fragment of the reference polypeptide or nucleic acid molecule), respectively. In some embodiments, the polypeptide or nucleic acid molecule has about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 97%, at least about 97%, about 98%, at least about 98%, about 99%, at least about 99% or about 100% sequence identity with a reference polypeptide or nucleic acid molecule (or a fragment of the reference polypeptide or nucleic acid molecule), respectively.

[0095] As used herein, "base edit" or "base editing" refers to the conversion of one nucleotide base pair to another. For example, base editing can convert cytosine (C) to thymine (T), or adenine (A) to guanine (G). Thus, base editing can swap a C-G base pair for an A-T base pair in a double-stranded polynucleotide, i.e., base editing generates a point mutation in the polynucleotide. Base editing is typically performed by a base editing enzyme, which in some embodiments includes a DNA targeting domain and a catalytic domain capable of base editing, i.e., a DNA editing domain. In some embodiments, the DNA targeting domain is Cas9, e.g., catalytically inactive Cas9 (dCas9), or Cas9 capable of generating a single-strand break (nCas9). In some embodiments, the DNA editing domain is a deaminase domain. The term "deaminase" refers to an enzyme that catalyzes a deamination reaction.

[0096] Base editing typically occurs via deamination, which refers to the removal of an amine group from a molecule, such as cytosine or adenosine. Deamination converts cytosine to uracil and adenosine to inosine. Exemplary cytidine deaminases include, for example, apolipoprotein B mRNA editing complex (APOBEC) deaminases, activation-induced cytidine deaminase (AID), and ACF1 / ASE deaminases. Exemplary adenosine deaminases include, for example, ADAR deaminases and ADAT deaminases (such as TadA).

[0097] In an exemplary base editing process, the base editing enzyme includes a modified Cas9 domain (nCas9) capable of generating a single-stranded DNA break (i.e., a "nick"), a cytidine deaminase domain, and a uracil DNA-glycosylase inhibitor domain (UGI). nCas9 is directed by a guide RNA to a target polynucleotide containing a "C-G" base pair, where the cytidine deaminase converts the cytosine in "C-G" to uracil to generate a "U-G" mismatch. nCas9 also generates a nick in the non-edited strand of the target polynucleotide. UGI inhibits the natural cellular repair that would convert the newly converted uracil back to cytosine, and the natural cellular mismatch repair mechanism activated by the nicked DNA strand converts the "U-G" mismatch to a "U-A" mismatch. Further DNA replication and repair convert uracil to thymine, and base editing of the target polynucleotide is complete. An example of a base editing enzyme is BE3 described in Komor et al., Nature 533(7603):420-424 (2016). Further exemplary base editing processes are described, for example, in Eid et al., Biochem J 475:1955-1964 (2018).

[0098] Methods for generating catalytically dead Cas9 domains (dCas9) are known (see, e.g., Jinek et al., Science 337:816-821 (2012); Qi et al., Cell 152(5):1173-1183 (2013)). For example, the DNA cleavage domain of Cas9 is known to contain two subdomains, an HNH nuclease subdomain and an RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, and the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9.

[0099] Non-limiting examples of base editing enzymes are described, for example, in U.S. Patent Nos. 9,068,179, 9,840,699, 10,167,457, and Eid et al., Biochem J 475(11):1955-1964 (2018), Gehrke et al., Nat Biotechnol 36:977-982 (2018), Hess et al., Mol Cell 68:26-43 (2017), Kim et al., Nat Biotechnol 35:435-437 (2017), Komor et al., Nature 533:420-424 (2016), Komor et al., Science Adv 3(8):eaao4774 (2017), Nishida et al., Science 353:aaf8729 (2016), Rees et al., Nat Commun 8:15790 (2017), Shimatani et al., Nat Biotechnol 35:441-443 (2017).

[0100] "Cytotoxic agent" or "cytotoxin," as used herein, typically refers to any agent that causes cell death by impairing or inhibiting one or more essential cellular processes. For example, cytotoxins such as diphtheria toxin, Shiga toxin, and Pseudomonas aeruginosa exotoxin function by impairing or inhibiting ribosome function, which halts protein synthesis and leads to cell death. Cytotoxins such as dolastatin, auristatin, and maytansine target microtubule function, which disrupts cell division and leads to cell death. Cytotoxins such as duocarmycin or calicheamicin target DNA directly and kill cells at any point in the cell cycle. Often, cytotoxic agents are introduced into cells by binding to receptors on the cell surface. The cytotoxic agent may be a natural compound or a derivative thereof, or the cytotoxic agent may be a synthetic molecule or peptide. In one embodiment, the cytotoxic agent may be an antibody-drug conjugate (ADC) comprising a monoclonal antibody (mAb) conjugated to a biologically active agent using a labile linker. The ADC combines the specificity of the mAb with the efficacy of the agent in the targeted killing of specific cells, such as cancer cells. ADCs (also referred to as "immunotoxins") are further described, for example, in Srivastava et al., Biomed Res Ther 2(1):169-183 (2015), and Grawunder and Barth (Eds.), Next Generation Antibody Drug Conjugates (ADCs) and Immunotoxins, Springer, 2017; doi:10.1007 / 978-3-319-46877-8.

[0101] The term "biallelic" site, as used herein, refers to a locus in the genome that contains two observed alleles. Thus, "biallelic" modification refers to the modification of both alleles in the genome of mammalian cells. For example, biallelic mutation means that mutations are present in both copies (i.e., the maternal copy and the paternal copy) of a particular gene.

[0102] Method for Introducing Site-Specific Mutations and Determining Their Efficacy In some embodiments, the present disclosure provides a method for introducing a site-specific mutation into a target polynucleotide in a target cell within a cell population, the method comprising: (a) introducing into the cell population: (i) a base editing enzyme; (ii) a first guide polynucleotide that: (1) hybridizes with a gene encoding a cytotoxic agent (CA) receptor and (2) forms a first complex with the base editing enzyme, wherein the base editing enzyme of the first complex provides a mutation in the gene encoding the CA receptor, and the mutation in the gene encoding the CA receptor forms CA-resistant cells in the cell population; and (iii) a second guide polynucleotide that: (1) hybridizes with the target polynucleotide and (2) forms a second complex with the base editing enzyme, wherein the base editing enzyme of the second complex provides a mutation in the target polynucleotide; (b) contacting the cell population with CA; and (c) enriching for target cells containing the mutation in the target polynucleotide by selecting CA-resistant cells from the cell population.

[0103] In some embodiments, the present disclosure provides a method for determining the efficacy of a base editing enzyme in a cell population, the method comprising: (a) introducing into the cell population: (i) a base editing enzyme; (ii) a first guide polynucleotide that: (1) hybridizes with a gene encoding a cytotoxic agent (CA) receptor and (2) forms a first complex with the base editing enzyme, wherein the base editing enzyme of the first complex introduces a mutation in the gene encoding the CA receptor, and the mutation in the gene encoding the CA receptor forms CA-resistant cells in the cell population; and (iii) a second guide polynucleotide that: (1) hybridizes with the target polynucleotide and (2) forms a second complex with the base editing enzyme, wherein the base editing enzyme of the second complex introduces a mutation in the target polynucleotide; (b) contacting the cell population with CA to isolate CA-resistant cells; and (c) determining the efficacy of the base editing enzyme by determining the ratio of the CA-resistant cells to the total cell population.

[0104] The methods of the present disclosure provide an efficient way to introduce a single nucleotide mutation (e.g., a C:G to T:A mutation) into various cell lines. Conventional limitations of genome engineering and gene editing strategies have been plagued by the inability to successfully distinguish edited cells from unedited cells, for example, because one or more of the editing components may not have been properly introduced into the cell or may not be expressed within the cell. Thus, there is a need in the art for increasing editing efficiency by selection and enrichment of edited cells.

[0105] The present disclosure also provides a rapid and accurate method for determining editing efficacy in a cell population. Such methods can facilitate determination of whether editing has occurred without the need for large-scale sequence analysis of target cells. The method can further enable evaluation of multiple guide polynucleotide sequences to determine the most effective guide polynucleotide sequence for a particular purpose. The methods of the present disclosure are “co-targeting enrichment” strategies that dramatically improve the editing efficiency of base editing enzymes. In the “co-targeting enrichment” strategy, two guide polynucleotides are introduced into the cell: a first guide polynucleotide, e.g., a “selection” polynucleotide that guides a base editing enzyme to a “selection” site, and a second guide polynucleotide, e.g., a “target” polynucleotide that guides the base editing enzyme to a “target” site. In some embodiments, successful editing of the “selection” site results in cells that survive under specific selection conditions (e.g., exposure to a cytotoxic agent, high or low temperature, medium lacking one or more nutrients, etc.). Figure 1A shows an embodiment of the present disclosure, depicting a starting cell population having “target” and “selection” sites. Under conditions where no selection is present, only a small percentage of cells have the desired “edited” site. Under “co-targeting HB-EGF + diphtheria toxin selection,” a much higher percentage of cells have the desired “edited” target site.

[0106] In some embodiments, successful editing of the "selection" site enables easy separation of edited cells based on physical or chemical properties (e.g., changes in cell shape or size and / or the ability to generate fluorescence, chemiluminescence, etc.) from unedited cells. In some embodiments, cells having an edited "selection" site are likely to also have an edited "target" site (which may be due to, for example, successful introduction and / or expression of one or more of the editing components). Thus, selection of cells having an edited "selection" site enriches cells having an edited "target" site and increases editing efficiency.

[0107] As used herein, "site-specific mutation" includes a single nucleotide substitution, e.g., a conversion of cytosine to thymine, or vice versa, or adenine to guanine, or vice versa, in a polynucleotide sequence. In some embodiments, site-specific mutations are generated by base editing enzymes. In some embodiments, site-specific mutations occur via deamination, e.g., by a deaminase of a nucleotide in a target polynucleotide. In some embodiments, the base editing enzyme includes a deaminase.

[0108] In some embodiments, a site-specific mutation in a target polynucleotide results in a change in the polypeptide sequence encoded by the polynucleotide. In some embodiments, a site-specific mutation in a target polynucleotide changes the expression of downstream polynucleotide sequences in a cell. For example, the expression of a downstream sequence polynucleotide sequence may be inactivated, such that the sequence is not transcribed, the encoded protein is not produced, or the sequence does not function as a wild-type sequence. For example, a protein or miRNA coding sequence may be inactivated, such that the protein is not produced.

[0109] In some embodiments, the site-specific mutation in the regulatory sequence increases the expression of the downstream polynucleotide. In some embodiments, the site-specific mutation inactivates the regulatory sequence such that the regulatory sequence no longer functions as a regulatory sequence. Non-limiting examples of regulatory sequences include the promoters, transcription terminators, enhancers, and other regulatory elements described herein. In some embodiments, the site-specific mutation results in a "knockout" of the target polynucleotide.

[0110] In some embodiments, the target cell is a eukaryotic cell. In some embodiments, the eukaryotic cell is an animal or human cell. In some embodiments, the target cell is a human cell. In some embodiments, the human cell is a stem cell. Stem cells can be, for example, pluripotent stem cells, such as embryonic stem cells (ESCs), adult stem cells, induced pluripotent stem cells (iPSCs), tissue-specific stem cells (e.g., hematopoietic stem cells), and mesenchymal stem cells (MSCs). In some embodiments, the human cell is a differentiated form of any of the cells described herein. In some embodiments, the eukaryotic cell is a cell derived from a primary cell under culture. In some embodiments, the cell is a stem cell or a stem cell line.

[0111] In some embodiments, the eukaryotic cell is a hepatocyte, such as a human hepatocyte, an animal hepatocyte, or a non-parenchymal cell. For example, the eukaryotic cell can be a human hepatocyte for adherent metabolism testing, a human hepatocyte for adherent induction testing, a human hepatocyte certified by QUALYST TRANSPORTER for adherent use, a human hepatocyte for suspension testing (e.g., 10-donor and 20-donor pooled hepatocytes), a human liver Kupffer cell, a human liver stellate cell, a canine hepatocyte (e.g., single and pooled beagle hepatocytes), a mouse hepatocyte (e.g., CD-1 and C57BI / 6 hepatocytes), a rat hepatocyte (e.g., Sprague-Dawley, Wistar Han, and Wistar hepatocytes), a monkey hepatocyte (e.g., cynomolgus or rhesus hepatocytes), a cat hepatocyte (e.g., domestic short hair hepatocytes), and a rabbit hepatocyte (e.g., New Zealand white hepatocytes).

[0112] In some embodiments, the methods of the present disclosure include introducing a base editing enzyme into a cell population. In some embodiments, the base editing enzyme includes a DNA targeting domain and a DNA editing domain. In some embodiments, the DNA targeting domain includes Cas9. In some embodiments, Cas9 includes a mutation in the catalytic domain. In some embodiments, the base editing enzyme includes catalytically inactive Cas9 and a DNA editing domain. In some embodiments, the base editing enzyme includes Cas9 (nCas9) capable of generating a single-stranded DNA break and a DNA editing domain. In some embodiments, nCas9 includes a mutation relative to wild-type Cas9 at amino acid residue D10 or H840 (numbering is relative to SEQ ID NO: 3). In some embodiments, Cas9 includes a polypeptide having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence identity with SEQ ID NO: 3. In some embodiments, Cas9 includes a polypeptide that is at least 90% identical to SEQ ID NO: 3. In some embodiments, Cas9 includes a polypeptide having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence identity with SEQ ID NO: 4. In some embodiments, Cas9 includes a polypeptide that is at least 90% identical to SEQ ID NO: 4.

[0113] The CRISPR-Cas system is a prokaryotic adaptive immune system that has recently been discovered and modified to enable robust and site-specific genome engineering in various organisms and cell lines. Generally, the CRISPR-Cas system is a protein-RNA complex that uses an RNA molecule (e.g., guide RNA) as a guide to localize the complex to a target DNA sequence through base pairing of the guide RNA to the target DNA sequence. Typically, Cas9 may also require a short protospacer adjacent motif (PAM) sequence adjacent to the target DNA sequence to bind to the DNA. Once the complex with the guide RNA is formed, Cas9 "searches" for the target DNA sequence by binding to a sequence that matches the PAM sequence. When Cas9 properly recognizes the pair of PAM and guide RNA using the target sequence, subsequently the Cas9 protein acts as an endonuclease to cleave the targeted DNA sequence. Cas9 proteins from different bacterial species can recognize different PAM sequences. For example, Cas9 from S. pyogenes (SpCas9) recognizes a PAM sequence of 5'-NGG-3', where N is any nucleotide. Cas9 proteins can also be modified to recognize PAMs different from wild-type Cas9. See, for example, Sternberg et al., Nature 507(7490):62-67(2014), Kleinstiver et al., Nature 523:481-485(2015), and Hu et al., Nature 556:57-63(2018).

[0114] Among known Cas proteins, SpCas9 has been the most widely used as a tool for genome engineering. The SpCas9 protein is a large multi-domain protein containing two different nuclease domains. As used herein, "Cas9" includes, for example, codon-optimized variants and modified Cas9 described in U.S. Patent Nos. 9,944,912, 9,512,446, and 10,093,910, and Cas9 variants of U.S. Patent Application No. 62 / 728,184, filed September 7, 2018, and encompasses any Cas9 protein and its variants. Point mutations can be introduced into Cas9 to obtain catalytically inactive Cas9, or dead Cas9 (dCas9), which abolishes nuclease activity but still retains its ability to bind DNA in a manner programmed by a guide RNA. In principle, when fused to another protein or domain, dCas9 can target that protein to substantially any DNA sequence simply by co-expressing it with an appropriate guide RNA. See, for example, Mali et al., Nat Methods 10(10):957-963 (2013), Horvath et al., Nature 482:331-338 (2012), Qi et al., Cell 152(5):1173-1183 (2013). In embodiments, the point mutations include mutations at positions D10 and H840 of wild-type Cas9 (numbering is relative to the amino acid sequence of wild-type SpCas9). In embodiments, dCas9 includes D10A and H840A mutations.

[0115] The wild-type Cas9 protein may also be modified such that the Cas9 protein has nickase activity capable of cleaving only one strand of double-stranded DNA, rather than nuclease activity that generates double-stranded breaks. Cas9 nickase (nCas9) is described, for example, in Cho et al., Genome Res 24:132-141 (2013), Ran et al., Cell 154:1380-1389 (2013), and Mali et al., Nat Biotechnol 31:833-838 (2013). In some embodiments, the Cas9 nickase comprises a single amino acid substitution relative to wild-type Cas9. In some embodiments, the single amino acid substitution is at position D10 of Cas9 (numbering is relative to SEQ ID NO: 3). In some embodiments, the single amino acid substitution is H10A (numbering is relative to SEQ ID NO: 3). In some embodiments, the single amino acid substitution is at position H840 of Cas9 (numbering is relative to SEQ ID NO: 3). In some embodiments, the single amino acid substitution is H840A (numbering is relative to SEQ ID NO: 3).

[0116] In some embodiments, the base editing enzyme comprises a DNA targeting domain and a DNA editing domain. In some embodiments, the DNA editing domain comprises a deaminase. In some embodiments, the deaminase is a cytidine deaminase or an adenosine deaminase. In some embodiments, the deaminase is a cytidine deaminase. In some embodiments, the deaminase is an adenosine deaminase. In some embodiments, the deaminase is an apolipoprotein B mRNA editing complex (APOBEC) deaminase, activation-induced cytidine deaminase (AID), ACF1 / ASE deaminase, ADAT deaminase, or ADAR deaminase. In some embodiments, the deaminase is an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the deaminase is APOBEC1.

[0117] As described herein, deaminase enzymes catalyze deamination, for example, of cytosine or adenosine. An exemplary family of cytosine deaminases is the APOBEC family, which includes 11 proteins that act to initiate mutagenesis in a controlled and beneficial manner (Conticello et al., Genome Biol 9(6):229(2008)). Activation-induced cytidine deaminase (AID), a member of one family, mediates antibody diversification by converting cytosine to uracil in ssDNA in a transcription-dependent strand-biased manner (Reynaud et al., Nat Immunol 4(7):631-638(2003)). APOBEC3 protects human cells from specific HIV-1 strains via deamination of cytosine in reverse-transcribed viral ssDNA (Bhagwat et al., DNA Repair (Amst) 3(1):85-89(2004)). All of these proteins require Zn 2+ coordinating motifs (His-X-Glu-X 23-26 -Pro-Cys-X 2-4-Cys) and bound water molecules. The Glu residue of the motif acts to activate a water molecule to zinc hydroxide for nucleophilic attack in the deamination reaction. Each member of the family preferentially deaminates at its own specific "hot spot" that ranges from WRC (where W is A or T and R is A or G) in the case of hAID to TTC in the case of hAPOBEC3F (Navaratnam et al., Int J Hematol 83(3):195 - 200(2006)). From the recent crystal structure of the catalytic domain of APOBEC3G, a secondary structure composed of a five-stranded β-sheet core flanked by six α-helices has been revealed, which is thought to be conserved across the family (Holden et al., Nature 456:121 - 124(2008)). The active center loop has been shown to govern both ssDNA binding and determination of the identity of the "hot spot" (Chelico et al., J Biol Chem 284(41):27761 - 27765(2009)). Overexpression of these enzymes is associated with genomic instability and cancer, and thus the importance of sequence-specific targeting has been emphasized (Pham et al., Biochemistry 44(8):2703 - 2715(2005)).

[0118] Another exemplary suitable type of nucleic acid editing enzyme and domain is adenosine deaminase. Examples of adenosine deaminases include adenosine deaminase acting on tRNA (ADAT), and the adenosine deaminase acting on RNA (ADAR) family. As deaminases of the ADAT family, there is TadA of tRNA adenosine deaminase that shares sequence similarity with APOBEC enzymes. As an ADAR family deaminase, there is ADAR2 that converts adenosine in double-stranded RNA to inosine to enable base editing of RNA. See, for example, Gaudelli et al., Nature 551:464 - 471(2017), Cox et al., Science 358:1019 - 1027(2017).

[0119] In some embodiments, the base editing enzyme further comprises a DNA glycosylase inhibitor domain. In some embodiments, the DNA glycosylase is a uracil DNA glycosylase inhibitor (UGI). Generally, DNA glycosylases such as uracil DNA glycosylase are part of the base excision repair pathway and, when detecting a U:G mismatch (where "U" is generated from the deamination of cytosine), perform error-free repair, converting the U back to the wild-type sequence and effectively "canceling" the base editing. Thus, by adding a DNA glycosylase inhibitor (e.g., a uracil DNA glycosylase inhibitor), the base excision repair pathway is inhibited and the base editing efficiency is increased. Non-limiting examples of DNA glycosylases include OGG1, MAG1, and UNG. The DNA glycosylase inhibitor can be a small molecule or a protein. For example, protein inhibitors of uracil DNA glycosylase are described in Mol et al., Cell 82:701-708 (1995), Serrano-Heras et al., J Biol Chem 281:7068-7074 (2006), and New England Biolabs Catalog No. M0281S and M0281L (neb.com / products / m0281-uracil-glycosylase-inhibitor-ugi). Small molecule inhibitors of DNA glycosylase are described, for example, in Huang et al., J Am Chem Soc 131(4):1344-1345 (2009), Jacobs et al., PLoS One 8(12):e81667 (2013), Donley et al., ACS Chem Biol 10(10):2334-2343 (2015), Tahara et al., J Am Chem Soc 140(6):2105-2114 (2018).

[0120] Thus, in some embodiments, the base editing enzyme of the present disclosure comprises a single-strand cleavage and Cas9 capable of producing cytidine deaminase. In some embodiments, the base editing enzyme of the present disclosure comprises nCas9 and cytidine deaminase. In some embodiments, the base editing enzyme of the present disclosure comprises a single-strand cleavage and Cas9 capable of producing adenosine deaminase. In some embodiments, the base editing enzyme of the present disclosure comprises nCas9 and adenosine deaminase. In some embodiments, the base editing enzyme is at least 90% identical to SEQ ID NO: 6. In some embodiments, the base editing enzyme comprises a polypeptide having at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, or at least 90% sequence identity with SEQ ID NO: 6. In some embodiments, the base editing enzyme comprises a polypeptide having at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence identity with SEQ ID NO: 6. In some embodiments, the polynucleotide encoding the base editing enzyme is at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identical to SEQ ID NO: 5. In some embodiments, the base editing enzyme is BE3.

[0121] In some embodiments, the method of the present disclosure comprises introducing into a cell population a first guide polynucleotide that hybridizes to a gene encoding a cytotoxic agent (CA) receptor and forms a first complex with a base editing enzyme, wherein the base editing enzyme of the first complex provides a mutation in the gene encoding the CA receptor, and the mutation in the gene encoding the CA receptor forms CA-resistant cells in the cell population.

[0122] In some embodiments, the first guide polynucleotide is an RNA molecule. RNA molecules that bind to CRISPR-Cas components and target them to specific locations within the target DNA are referred to herein as "RNA guide polynucleotides", "guide RNAs", "gRNAs", "small guide RNAs", "single guide RNAs", or "sgRNAs", and may also be referred to herein as "DNA-targeting RNAs". The guide polynucleotide may be introduced into the target cell as an isolated molecule, e.g., an RNA molecule, or introduced into the cell using an expression vector containing DNA encoding the guide polynucleotide, e.g., an RNA guide polynucleotide. In some embodiments, the guide polynucleotide is 10 to 150 nucleotides. In some embodiments, the guide polynucleotide is 20 to 120 nucleotides. In some embodiments, the guide polynucleotide is 30 to 100 nucleotides. In some embodiments, the guide polynucleotide is 40 to 80 nucleotides. In some embodiments, the guide polynucleotide is 50 to 60 nucleotides. In some embodiments, the guide polynucleotide is 10 to 35 nucleotides. In some embodiments, the guide polynucleotide is 15 to 30 nucleotides. In some embodiments, the guide polynucleotide is 20 to 25 nucleotides.

[0123] In some embodiments, the RNA guide polynucleotide comprises at least two nucleotide segments: at least one "DNA-binding segment" and at least one "polypeptide-binding segment". "Segment" means a part, section, or region of a molecule, e.g., a continuous stretch of nucleotides of a guide polynucleotide molecule. The definition of "segment" is not limited to a specific number of total base pairs unless specifically defined otherwise.

[0124] In some embodiments, the guide polynucleotide comprises a DNA-binding segment. In some embodiments, the DNA-binding segment of the guide polynucleotide comprises a nucleotide sequence complementary to a specific sequence within the target polynucleotide. In some embodiments, the DNA-binding segment of the guide polynucleotide hybridizes to a gene encoding a cytotoxic agent (CA) receptor in the target cell. In some embodiments, the DNA-binding segment of the guide polynucleotide hybridizes to a target polynucleotide sequence in the target cell. Target cells including various types of eukaryotic cells are described herein.

[0125] In some embodiments, the guide polynucleotide comprises a polypeptide-binding segment. In some embodiments, the polypeptide-binding segment of the guide polynucleotide binds to the DNA targeting domain of the base editing enzyme of the present disclosure. In some embodiments, the polypeptide-binding segment of the guide polynucleotide binds to Cas9 of the base editing enzyme. In some embodiments, the polypeptide-binding segment of the guide polynucleotide binds to dCas9 of the base editing enzyme. In some embodiments, the polypeptide-binding segment of the guide polynucleotide binds to nCas9 of the base editing enzyme. Various RNA guide polynucleotides that bind to the Cas9 protein are described, for example, in US Patent Application Publication Nos. 2014 / 0068797, 2014 / 0273037, 2014 / 0273226, 2014 / 0295556, 2014 / 0295557, 2014 / 0349405, 2015 / 0045546, 2015 / 0071898, 2015 / 0071899, and 2015 / 0071906.

[0126] In some embodiments, the guide polynucleotide further comprises a tracrRNA. A "tracrRNA" or trans-activating CRISPR-RNA forms an RNA duplex with a pre-crRNA, or pre-CRISPR-RNA, which is then cleaved by the RNA-specific ribonuclease RNase III to form a crRNA / tracrRNA hybrid. In some embodiments, the guide polynucleotide comprises a crRNA / tracrRNA hybrid. In some embodiments, the tracrRNA component of the guide polynucleotide activates the Cas9 protein. In some embodiments, activation of the Cas9 protein comprises activating the nuclease activity of Cas9. In some embodiments, activation of the Cas9 protein comprises a Cas9 protein that binds to a target polynucleotide sequence.

[0127] In some embodiments, the sequence of the guide polynucleotide is designed to target a base editing enzyme to a specific location in the target polynucleotide sequence. Various tools and programs are available to facilitate the design of such guide polynucleotides (see, e.g., the Benchling Base Editor Design Guide (benchling.com / editor#create / crispr), as well as the BE-Designer and BE-Analyzer from CRISPR RGEN Tools (see Hwang et al., bioRxiv dx.doi.org / 10.1101 / 373944, first published on Jul. 22, 2018)).

[0128] In some embodiments, the DNA-binding segment of the first guide polynucleotide hybridizes with the gene encoding the cytotoxic agent (CA) receptor, and the polypeptide-binding segment of the first guide polynucleotide binds to the DNA targeting domain of the base editing enzyme to form a first complex with the base editing enzyme. In some embodiments, the DNA-binding segment of the first guide polynucleotide hybridizes with the gene encoding the cytotoxic agent (CA) receptor, and the polypeptide-binding segment of the first guide polynucleotide binds to Cas9 of the base editing enzyme to form a first complex with the base editing enzyme. In some embodiments, the DNA-binding segment of the first guide polynucleotide hybridizes with the gene encoding the cytotoxic agent (CA) receptor, and the polypeptide-binding segment of the first guide polynucleotide binds to dCas9 of the base editing enzyme to form a first complex with the base editing enzyme. In some embodiments, the DNA-binding segment of the first guide polynucleotide hybridizes with the gene encoding the cytotoxic agent (CA) receptor, and the polypeptide-binding segment of the first guide polynucleotide binds to nCas9 of the base editing enzyme to form a first complex with the base editing enzyme.

[0129] In some embodiments, the first complex is targeted to the gene encoding the CA receptor by the first guide polynucleotide, and the base editing enzyme of the first complex introduces a mutation into the gene encoding the CA receptor. In some embodiments, the mutation in the gene encoding the CA receptor is introduced by the base editing domain of the base editing enzyme of the first complex. In some embodiments, the mutation in the gene encoding the CA receptor forms CA-resistant cells in the cell population. In some embodiments, the mutation is a point mutation from cytosine (C) to thymine (T). In some embodiments, the mutation is a point mutation from adenine (A) to guanine (G). The specific location of the mutation in the CA receptor can be induced, for example, by the design of the first guide polynucleotide using tools such as the Benchling base editor design guides, BE-Designer, and BE-Analyzer described herein. In some embodiments, the first guide polynucleotide is an RNA polynucleotide. In some embodiments, the first guide polynucleotide further comprises a tracrRNA sequence.

[0130] In some embodiments, the CA is a compound that causes or promotes cell death, as described herein. In some embodiments, the CA is a toxin. In some embodiments, the CA is a natural toxin. In some embodiments, the CA is a synthetic poison. In some embodiments, the CA is a small molecule, peptide, or protein. In some embodiments, the CA is an antibody-drug conjugate. In some embodiments, the CA is a monoclonal antibody conjugated to a biologically active agent via a labile linker. In some embodiments, the CA is a biotoxin. In some embodiments, the toxin is produced by cyanobacteria (cyanotoxin), dinoflagellates (dinotoxin), spiders, snakes, scorpions, frogs, jellyfish, sea mammals, poisonous fish, corals, or octopuses. Examples of toxins include, for example, diphtheria toxin, botulinum toxin, ricin, apitoxin, saxitoxin, Pseudomonas exotoxin, and mycotoxin. In some embodiments, the CA is diphtheria toxin. In some embodiments, the CA is an antibody-drug conjugate. In some embodiments, the antibody-drug conjugate comprises an antibody conjugated to a toxin. In some embodiments, the toxin is a small molecule, RNase, or pro-apoptotic protein.

[0131] In some embodiments, the CA is toxic to one organism, such as a human, but not to another organism, such as a mouse. In some embodiments, the CA is toxic to an organism during a certain stage of its life cycle (e.g., fetal stage), but not during another life cycle stage of the organism (e.g., adult stage). In some embodiments, the CA is toxic in one organ of an animal, but not to another organ of the same animal. In some embodiments, the CA is toxic to a subject (e.g., a human or an animal) in a certain condition or state (e.g., being ill), but not to the same subject in another condition or state (e.g., being healthy). In some embodiments, the CA is toxic to one cell type, but not to another cell type. In some embodiments, the CA is toxic to cells in a certain cell state (e.g., differentiated), but not to the same cells in another cell state (e.g., undifferentiated). In some embodiments, the CA is toxic to cells in one environment (e.g., low temperature), but not to the same cells in another environment (e.g., high temperature). In some embodiments, the toxin is toxic to human cells, but not to mouse cells.

[0132] In some embodiments, the CA receptor is a biological receptor that binds to the CA. The CA receptor is a protein molecule that binds to the CA and is typically located on the cell membrane. For example, diphtheria toxin binds to the human heparin-binding EGF-like growth factor (HB-EGF). The CA receptor may be specific for one CA, or the CA receptor may bind to two or more CAs. For example, monosialoganglioside (GM 1) can act as a receptor for both cholera toxin and the heat-labile enterotoxin of E. coli (E. coli). Alternatively, two or more CA receptors may bind to one CA. For example, botulinum toxin is thought to bind to different receptors on nerve cells and epithelial cells. In some embodiments, the CA receptor is a receptor that binds to CA. In some embodiments, the CA receptor is a G protein-coupled receptor. In some embodiments, the CA receptor is a receptor for an antibody, such as an antibody of an antibody-drug conjugate. In some embodiments, the CA receptor is a receptor for diphtheria toxin. In some embodiments, the CA receptor is HB-EGF.

[0133] In some embodiments, one or more mutations in the polynucleotide encoding the CA receptor protein confer CA resistance. In some embodiments, mutations in the CA binding region of the CA receptor confer CA resistance. In some embodiments, charge reversal mutations of amino acids at or near the CA binding site of the CA receptor confer CA resistance. Examples of charge reversal mutations include substitution of a negatively charged amino acid such as Glu or Asp with a positively charged amino acid such as Lys or Arg, or vice versa. In some embodiments, polarity reversal mutations of amino acids at or near the CA binding site of the CA receptor confer CA resistance. Examples of polarity reversal mutations include substitution of a polar amino acid such as Gln or Asn with a non-polar amino acid such as Val or Ile, or vice versa. In some embodiments, by substituting a relatively small amino acid residue at or near the CA binding site of the CA receptor with a "bulky" amino acid residue, the binding pocket is blocked, preventing the binding of CA, and thus conferring CA resistance. Examples of small amino acids include Gly or Ala, while Trp is generally considered a bulky amino acid.

[0134] In some embodiments, one or more mutations in the polynucleotide encoding the CA receptor alter one or more codons in the amino acid sequence of the CA receptor. In some embodiments, one or more mutations in the polynucleotide encoding the CA receptor alter a single codon in the amino acid sequence of the CA receptor. In some embodiments, a single nucleotide mutation in the polynucleotide encoding the CA receptor protein confers CA receptor resistance. In some embodiments, the single nucleotide mutation is a point mutation from cytosine (C) to thymine (T) in the polynucleotide sequence encoding the CA receptor. In some embodiments, the single nucleotide mutation is a point mutation from adenine (A) to guanine (G) in the polynucleotide sequence encoding the CA receptor. In some embodiments, one or more mutations in the CA receptor are provided by the base editing enzymes described herein. The base editing enzyme is specifically targeted to the CA receptor by a DNA targeting domain (e.g., a Cas9 domain), followed by the base editing domain (e.g., a deaminase domain) providing the mutation in the CA receptor. In some embodiments, one or more mutations in the CA receptor are provided by a base editing enzyme comprising nCas9 and a cytidine deaminase. In some embodiments, one or more mutations in the CA receptor are provided by a base editing enzyme comprising nCas9 and an adenosine deaminase. In some embodiments, one or more mutations in the CA receptor are provided by a base editing enzyme comprising a polypeptide having at least 90% sequence identity with SEQ ID NO: 6. In some embodiments, the base editing enzyme is BE3.

[0135] In some embodiments, the CA receptor is the receptor for diphtheria toxin. In some embodiments, the diphtheria toxin receptor is human HB-EGF. Unless otherwise indicated, "HB-EGF" as used herein without an organism modifier refers to human HB-EGF. HB-EGF proteins from other organisms such as mice are specifically described as "mouse HB-EGF".

[0136] Diphtheria toxin, a "A-B" toxin that is a two-component protein forming a complex with two subunits, typically binds with a disulfide bridge, the "A" subunit is typically considered the "active" part, and the "B" subunit is generally the "binding" part. Diphtheria toxin is known to bind to the EGF-like domain of HB-EGF, which is widely expressed in various tissues. Figure 3A shows an exemplary mechanism of action of the A-B diphtheria toxin on its receptor. As shown in Figure 3A, the diphtheria subunit B is responsible for the binding of the membrane-bound receptor HB-EGF. Upon binding, the diphtheria toxin enters the cell via receptor-mediated endocytosis. Subsequently, the catalytic subunit A is cleaved from subunit B through the reduction of the disulfide bond between the two subunits, leaving the endocytic vesicle and catalyzing the addition of ADP-ribose to ribosomal elongation factor 2 (EF2). The ADP-ribosylation of EF2 stops protein synthesis and leads to cell death.

[0137] Unlike human HB-EGF, mouse HB-EGF is resistant to binding and, thus, mice are resistant to diphtheria toxin. Figure 3B shows significant differences in the amino acid sequences of human and mouse HB-EGF proteins. Thus, in some embodiments, one or more mutations in the polynucleotide encoding the HB-EGF protein confer diphtheria toxin resistance. In some embodiments, one or more mutations in the polynucleotide encoding HB-EGF alter one or more codons in the amino acid sequence of HB-EGF. In some embodiments, one or more mutations in the polynucleotide encoding HB-EGF alter a single codon in the amino acid sequence of HB-EGF. In some embodiments, a single nucleotide mutation in the polynucleotide encoding the HB-EGF protein confers diphtheria toxin resistance. In some embodiments, the single nucleotide mutation is a point mutation from cytosine (C) to thymine (T) in the polynucleotide sequence encoding HB-EGF. In some embodiments, the single nucleotide mutation is a point mutation from adenine (A) to guanine (G) in the polynucleotide sequence encoding HB-EGF.

[0138] In some embodiments, mutations in the diphtheria toxin binding region of HB-EGF confer diphtheria toxin resistance. In some embodiments, mutations in the EGF-like domain of HB-EGF confer diphtheria toxin resistance. In some embodiments, charge reversal mutations of amino acids in or near the diphtheria toxin binding site of HB-EGF confer diphtheria toxin resistance. In some embodiments, the charge reversal mutation is a substitution from a negatively charged residue such as Glu or Asp to a positively charged residue such as Lys or Arg. In some embodiments, the charge reversal mutation is a substitution from a positively charged residue such as Lys or Arg to a negatively charged residue such as Glu or Asp. In some embodiments, polarity reversal mutations of amino acids in or near the diphtheria toxin binding site of HB-EGF confer diphtheria toxin resistance. In some embodiments, the polarity reversal mutation is a substitution from a polar amino acid residue such as Gln or Asn to a non-polar amino acid residue such as Ala, Val, or Ile. In some embodiments, the polarity reversal mutation is a substitution from a non-polar amino acid residue such as Ala, Val, or Ile to a polar amino acid residue such as Gln or Asn. In some embodiments, the mutation is a substitution from a relatively small amino acid residue such as Gly or Ala in or near the diphtheria toxin binding site of HB-EGF to a "bulky" amino acid residue such as Trp. In some embodiments, the mutation from a small residue to a bulky residue blocks the binding pocket and prevents the binding of diphtheria toxin, thereby conferring resistance.

[0139] In some embodiments, mutations in one or more of amino acids 100 to 160 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, mutations in one or more of amino acids 105 to 150 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, mutations in one or more of amino acids 107 to 148 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, mutations in one or more of amino acids 120 to 145 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, mutations in one or more of amino acids 135 to 143 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, mutations in one or more of amino acids 138 to 144 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, a mutation at amino acid 141 of wild-type HB-EGF (SEQ ID NO: 8) confers diphtheria toxin resistance. In some embodiments, the mutation at amino acid 141 of wild-type HB-EGF (SEQ ID NO: 8) is from GLU141 to ARG141. In some embodiments, the mutation at amino acid 141 of wild-type HB-EGF (SEQ ID NO: 8) is from GLU141 to HIS141. In some embodiments, the mutation at amino acid 141 of wild-type HB-EGF (SEQ ID NO: 8) is from GLU141 to LYS141. In some embodiments, the mutation from GLU141 to LYS141 in wild-type HB-EGF (SEQ ID NO: 8) confers diphtheria toxin resistance.

[0140] In some embodiments, one or more mutations in HB-EGF are provided by a base editing enzyme described herein. The base editing enzyme is specifically targeted to HB-EGF by a DNA targeting domain (e.g., a Cas9 domain), and subsequently a base editing domain (e.g., a deaminase domain) provides a mutation in HB-EGF. In some embodiments, one or more mutations in HB-EGF are provided by a base editing enzyme comprising nCas9 and a cytidine deaminase. In some embodiments, one or more mutations in HB-EGF are provided by a base editing enzyme comprising nCas9 and an adenosine deaminase. In some embodiments, one or more mutations in HB-EGF are provided by a base editing enzyme comprising a polypeptide having at least 90% sequence identity with SEQ ID NO: 6. In some embodiments, the base editing enzyme is BE3.

[0141] In some embodiments, the DNA binding segment of the second guide polynucleotide hybridizes to a target polynucleotide in a target cell, and the polypeptide binding segment of the second guide polynucleotide binds to the DNA targeting domain of the base editing enzyme to form a second complex with the base editing enzyme. In some embodiments, the DNA binding segment of the second guide polynucleotide hybridizes to a target polynucleotide in a target cell, and the polypeptide binding segment of the second guide polynucleotide binds to Cas9 of the base editing enzyme to form a second complex with the base editing enzyme. In some embodiments, the DNA binding segment of the second guide polynucleotide hybridizes to a target polynucleotide in a target cell, and the polypeptide binding segment of the second guide polynucleotide binds to dCas9 of the base editing enzyme to form a second complex with the base editing enzyme. In some embodiments, the DNA binding segment of the second guide polynucleotide hybridizes to a target polynucleotide in a target cell, and the polypeptide binding segment of the second guide polynucleotide binds to nCas9 of the base editing enzyme to form a second complex with the base editing enzyme.

[0142] In some embodiments, the second complex is targeted to the target polynucleotide by the second guide polynucleotide, and the base editing enzyme of the second complex introduces a mutation in the target polynucleotide. In some embodiments, the mutation in the target polynucleotide is introduced by the base editing domain of the base editing enzyme of the second complex. In some embodiments, the mutation in the target polynucleotide is a point mutation from cytosine (C) to thymine (T). In some embodiments, the mutation in the target polynucleotide is a point mutation from adenine (A) to guanine (G). The specific position of the mutation in the target polynucleotide can be induced, for example, by the design of the second guide polynucleotide using tools such as the Benchling base editor design guide, BE-Designer, and BE-Analyzer described herein. In some embodiments, the second guide polynucleotide is an RNA polynucleotide. In some embodiments, the second guide polynucleotide further comprises a tracrRNA sequence.

[0143] In some embodiments, the C-to-T mutation in the target polynucleotide inactivates the expression of the target polynucleotide in the target cell. In some embodiments, the A-to-G mutation in the target polynucleotide inactivates the expression of the target polynucleotide in the target cell. In some embodiments, the target polynucleotide encodes a protein or miRNA. In some embodiments, the target polynucleotide is a regulatory sequence, and the C-to-T mutation changes the function of the regulatory sequence. In some embodiments, the target polynucleotide is a regulatory sequence, and the A-to-G mutation changes the function of the regulatory sequence.

[0144] In some embodiments, the base editing enzyme of the present disclosure is introduced into a cell population as a polynucleotide encoding the base editing enzyme. In some embodiments, the first and / or second guide polynucleotide is introduced into the cell population as one or more polynucleotides encoding the first and / or second guide polynucleotide. In some embodiments, the base editing enzyme, the first guide polynucleotide, and the second guide polynucleotide are introduced into the cell population via a vector. In some embodiments, the polynucleotide encoding the base editing enzyme, the first guide polynucleotide, and the second guide polynucleotide are on a single vector. In some embodiments, the vector is a viral vector. In some embodiments, the polynucleotide encoding the base editing enzyme, the first guide polynucleotide, and the second guide polynucleotide are on one or more vectors. In some embodiments, the one or more vectors are viral vectors. In some embodiments, the viral vector is an adenovirus, an adeno-associated virus, or a lentivirus. Viral transduction (administration can be local, targeted, or systemic) by adenovirus, adeno-associated virus (AAV), and lentiviral vectors has been used as a delivery method for in vivo gene therapy. Methods for introducing a vector such as a viral vector into a cell (e.g., transfection) are described herein.

[0145] In some embodiments, the base editing enzyme, the first guide polynucleotide, and / or the second guide polynucleotide are introduced into the cell population via delivery particles. In some embodiments, the base editing enzyme, the first guide polynucleotide, and / or the second guide polynucleotide are introduced into the cell population via vesicles.

[0146] In some embodiments, the efficacy of a base editing enzyme can be determined by calculating the ratio of CA-resistant cells to the total cell population. In some embodiments, the number of CA-resistant cells can be counted using techniques known in the art, such as, for example, counting using a hemocytometer, measuring absorbance at a specific wavelength (e.g., 580 nm or 600 nm), and / or measuring the fluorescence of a phosphor for detecting the cell population. In some embodiments, the total cell population is determined and the ratio of CA-resistant cells to the total cell population is calculated by dividing the total cell population by the CA-resistant cells. In some embodiments, the ratio of CA-resistant cells to the total cell population approximates the base editing efficacy at the target polynucleotide.

[0147] Site-specific integration method As described herein, HDR-based DNA double-strand break repair can provide site-specific integration, such as biallelic integration of a sequence of interest (SOI) at a target locus. In applications such as the correction of genetic variants, gene therapy, and the generation of transgenic animals, site-specific integration of the genetic modification of interest, particularly biallelic integration, is highly desirable. Unfortunately, due to the low efficiency of HDR-based DNA double-strand break repair, screening and isolation of site-specific integration, particularly biallelic integration, are often difficult and cumbersome and may require expensive and time-consuming sequencing and analysis. The methods of the present disclosure apply the "co-targeting enrichment" strategy described herein to generate site-specific integration of a sequence of interest and provide a simple and efficient screening method for cells having the desired integration. In some embodiments, the site-specific integration is biallelic integration.

[0148] In some embodiments, the present disclosure provides a method for providing a biallelic integration of a sequence of interest (SOI) at a toxin-sensing gene (TSG) locus in the genome of a cell, the method comprising: (a) introducing into a cell population: (i) a nuclease capable of generating a double-strand break; (ii) a guide polynucleotide capable of forming a complex with the nuclease and hybridizing to the TSG locus; and (iii) a donor polynucleotide comprising: (1) a 5' homology arm, a 3' homology arm, and a mutation in the native coding sequence of the TSG, the mutation conferring resistance to a toxin; and (2) the SOI, wherein as a result of introducing (i), (ii), and (iii), the donor polynucleotide is integrated into the TSG locus; (b) contacting the cell population with a toxin; and selecting one or more cells resistant to the toxin, wherein the one or more cells resistant to the toxin comprise a biallelic integration of the SOI.

[0149] Figure 10A shows an embodiment of the method provided herein. In Figure 10A, the wild-type sequence of HB-EGF is sensitive to diphtheria toxin. The solid-bordered boxes in the sequence indicate exons, and the double lines represent introns. The Cas9 nuclease is targeted to an intron of HB-EGF by a guide polynucleotide of the CRISPR-Cas complex to generate a double-strand break. The HDR template is introduced into cells having a splicing acceptor sequence to ligate an exon on the HDR template to an adjacent genomic exon, a diphtheria toxin resistance mutation in the exon immediately preceding the double-strand break, and a gene of interest (GOI). HDR repairs the double-strand break and inserts the splicing acceptor sequence, the diphtheria toxin resistance mutation, and the GOI at the cleavage site. Thus, only cells having a biallelic integration of the HDR template (and thus the GOI) are resistant to diphtheria toxin, and cells that are monoallelic or not repaired by HDR are sensitive to the toxin. Thus, cells that survive contact with the toxin have a biallelic integration of the GOI.

[0150] In some embodiments, the TSG locus encodes HB-EGF and the toxin is diphtheria toxin. In some embodiments, the nuclease capable of generating double-strand breaks is Cas9. In some embodiments, the guide polynucleotide is guide RNA. In some embodiments, the donor polynucleotide is an HDR template. In some embodiments, the SOI is the target gene. In some embodiments, the integration of the donor polynucleotide within the TSG locus is a biallelic integration.

[0151] In some embodiments, the present disclosure provides a method for integrating a sequence of interest (SOI) into a target locus in the genome of a cell, the method comprising: (a) introducing into a cell population (i) a nuclease capable of generating double-strand breaks, (ii) a guide polynucleotide capable of forming a complex with the nuclease and hybridizing with a toxin-sensitive gene (TSG) locus in the genome of the cell, wherein the TSG is an essential gene, and (iii) a donor polynucleotide comprising (1) a functional TSG gene comprising a mutation in the native coding sequence of the TSG, the mutation conferring resistance to the toxin, (2) the SOI, and (3) a sequence for genomic integration at the target locus, wherein introducing (i), (ii), and (iii) results in inactivation of the TSG in the genome of the cell by the nuclease and integration of the donor polynucleotide at the target locus; (b) contacting the cell population with a toxin; and (c) selecting one or more cells resistant to the toxin, wherein the one or more cells resistant to the toxin comprise the SOI integrated into the pre-target locus.

[0152] In some embodiments, the present disclosure provides a method for introducing a stable episomal vector into a cell, the method comprising: (a) introducing into a cell population: (i) a nuclease capable of generating a double-strand break; (ii) a guide polynucleotide capable of forming a complex with the nuclease and hybridizing to a toxin-sensitive gene (TSG) locus in the genome of the cell, wherein introduction of (i) and (ii) results in inactivation of the TSG in the genome of the cell by the nuclease; and (iii) an episomal vector comprising: (1) a functional TSG comprising a mutation in the native coding sequence of the TSG, the mutation conferring resistance to the toxin; (2) a sequence of interest (SOI); and (3) an autonomous DNA replication sequence; (b) contacting the cell population with a toxin; and (c) selecting one or more cells resistant to the toxin, wherein the one or more cells resistant to the toxin comprise the episomal vector. In some embodiments, the TSG is an essential gene.

[0153] In some embodiments, the nuclease capable of generating a double-strand break is Cas9. As described herein, Cas9 is a monomeric protein comprising a DNA targeting domain (which interacts with a guide polynucleotide, such as a guide RNA) and a nuclease domain (which cleaves a target polynucleotide, such as a TSG locus). The Cas9 protein generates site-specific cleavage in a nucleic acid. In some embodiments, the Cas9 protein generates site-specific double-strand breaks in DNA. The ability of Cas9 to target a specific sequence in a nucleic acid (i.e., site specificity) is achieved by Cas9 complexing with a guide polynucleotide (such as a guide RNA) that hybridizes to a specified sequence (such as a TSG locus). In some embodiments, Cas9 is a Cas9 variant described in U.S. Patent Application No. 62 / 728,184, filed Sep. 7, 2018.

[0154] In some embodiments, Cas9 is capable of generating sticky ends. Cas9 capable of generating sticky ends is described, for example, in International Publication PCTUS No. 2018 / 061680 pamphlet filed on November 16, 2018. In some embodiments, Cas9 capable of generating sticky ends is a dimeric Cas9 fusion protein. In some embodiments, it is advantageous to use a dimeric nuclease, i.e., a nuclease that is inactive until both monomers of the dimer are present at the target sequence, to achieve a higher degree of targeting specificity. The binding domains and cleavage domains of natural nucleases (e.g., Cas9, etc.), as well as modular binding domains and cleavage domains that can be fused to generate nuclease-binding specific target sites, are well known to those skilled in the art. For example, a Cas9 protein having a binding domain of an RNA-programmable nuclease (e.g., Cas9) or an inactive DNA cleavage domain can be used as a binding domain (e.g., binds to a gRNA to direct binding to the target site) to specifically bind to a desired target site, and fused or conjugated to a cleavage domain, e.g., the cleavage domain of endonuclease FokI, to generate a modified nuclease that cleaves the target region. The Cas9-FokI fusion protein is further described, for example, in U.S. Patent Application Publication No. 2015 / 0071899, and Guilinger et al., "Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification" Nature Biotechnology 32:577-582 (2014).

[0155] In some embodiments, Cas9 comprises the polypeptide of SEQ ID NO: 3 or 4. In some embodiments, Cas9 comprises a polypeptide having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence identity with SEQ ID NO: 3 or 4. In some embodiments, Cas9 is SEQ ID NO: 3 or 4.

[0156] In some embodiments, the guide polynucleotide is an RNA polynucleotide. RNA molecules that bind to CRISPR-Cas components and target them to specific positions within the target DNA are referred to herein as "RNA guide polynucleotides", "guide RNAs", "gRNAs", "small guide RNAs", "single guide RNAs", or "sgRNAs", and may also be referred to herein as "DNA-targeting RNAs". The guide polynucleotide may be introduced into the target cell as an isolated molecule, e.g., an RNA molecule, or may be introduced into the cell using an expression vector that contains DNA encoding the guide polynucleotide, e.g., an RNA guide polynucleotide. In some embodiments, the guide polynucleotide is 10 to 150 nucleotides. In some embodiments, the guide polynucleotide is 20 to 120 nucleotides. In some embodiments, the guide polynucleotide is 30 to 100 nucleotides. In some embodiments, the guide polynucleotide is 40 to 80 nucleotides. In some embodiments, the guide polynucleotide is 50 to 60 nucleotides. In some embodiments, the guide polynucleotide is 10 to 35 nucleotides. In some embodiments, the guide polynucleotide is 15 to 30 nucleotides. In some embodiments, the guide polynucleotide is 20 to 25 nucleotides.

[0157] In some embodiments, the RNA-guided polynucleotide comprises at least two nucleotide segments: at least one "DNA-binding segment" and at least one "polypeptide-binding segment". A "segment" means a part, section, or region of a molecule, e.g., a continuous stretch of nucleotides of a guide polynucleotide molecule. The definition of "segment" is not limited to a specific number of total base pairs unless specifically defined otherwise.

[0158] In some embodiments, the guide polynucleotide comprises a DNA-binding segment. In some embodiments, the DNA-binding segment of the guide polynucleotide comprises a nucleotide sequence complementary to a specific sequence within the target polynucleotide. In some embodiments, the DNA-binding segment of the guide polynucleotide hybridizes to a toxin-sensitive gene (TSG) locus in a cell. Various types of cells, such as eukaryotic cells, are described herein.

[0159] In some embodiments, the guide polynucleotide comprises a polypeptide-binding segment. In some embodiments, the polypeptide-binding segment of the guide polynucleotide binds to the DNA targeting domain of the nuclease of the present disclosure. In some embodiments, the polypeptide-binding segment of the guide polynucleotide binds to Cas9. In some embodiments, the polypeptide-binding segment of the guide polynucleotide binds to dCas9. In some embodiments, the polypeptide-binding segment of the guide polynucleotide binds to nCas9. Various RNA-guided polynucleotides that bind to the Cas9 protein are described, for example, in U.S. Patent Application Publication Nos. 2014 / 0068797, 2014 / 0273037, 2014 / 0273226, 2014 / 0295556, 2014 / 0295557, 2014 / 0349405, 2015 / 0045546, 2015 / 0071898, 2015 / 0071899, and 2015 / 0071906.

[0160] In some embodiments, the guide polynucleotide further comprises a tracrRNA. A "tracrRNA" or trans-activating CRISPR-RNA forms an RNA duplex with a pre-crRNA, or pre-CRISPR-RNA, which is then cleaved by the RNA-specific ribonuclease RNase III to form a crRNA / tracrRNA hybrid. In some embodiments, the guide polynucleotide comprises a crRNA / tracrRNA hybrid. In some embodiments, the tracrRNA component of the guide polynucleotide activates the Cas9 protein. In some embodiments, activation of the Cas9 protein comprises activating the nuclease activity of Cas9. In some embodiments, activation of the Cas9 protein comprises a Cas9 protein that binds to a target polynucleotide sequence, such as a TSG locus.

[0161] In some embodiments, the guide polynucleotide guides a nuclease to the TSG locus, and the nuclease generates a double-strand break at the TSG locus. In some embodiments, the guide polynucleotide is a guide RNA. In some embodiments, the nuclease is Cas9. In some embodiments, the double-strand break at the TSG locus inactivates the TSG. In some embodiments, inactivation of the TSG locus confers toxin resistance to the cell. In some embodiments, inactivation of the TSG locus confers toxin resistance to the cell, but also disrupts the normal cellular function of the TSG locus. In some embodiments, the TSG locus encodes a gene that performs a cellular function unrelated to toxin sensitivity. For example, the TSG locus can encode a protein that promotes cell growth or division, a receptor for a signaling molecule (e.g., a molecule in the vicinity of the cell), or a protein that interacts with another protein, organelle, or biomolecule to perform a normal cellular function.

[0162] In some embodiments, the TSG is an essential gene. An essential gene is a gene of an organism that is thought to be indispensable for survival under certain conditions. In some embodiments, disruption or deletion of the TSG results in cell death. In some embodiments, the TSG is a auxotrophic gene, i.e., a gene that produces a specific compound necessary for growth or survival. Examples of auxotrophic genes include genes involved in nucleotide biosynthesis such as adenine, cytosine, guanine, thymine, or uracil, or in amino acid biosynthesis such as histidine, leucine, lysine, methionine, or tryptophan. In some embodiments, the TSG is a gene in a metabolic pathway. In some embodiments, the TSG is a gene in the autophagy pathway. In some embodiments, the TSG is a gene in cell division such as mitosis, cytoskeletal organization, or response to stress or stimuli. In some embodiments, the TSG encodes a protein that promotes cell growth or division, a receptor for signaling molecules (e.g., molecules in the vicinity of the cell), or a protein that interacts with another protein, organelle, or biomolecule. Exemplary essential genes include, but are not limited to, the genes listed in FIG. 23. Further examples of essential genes are provided, for example, in Hart et al., Cell 163:1515-1526 (2015), Zhang et al., Microb Cell 2(8):280-287 (2015), and Fraser, Cell Systems 1:381-382 (2015).

[0163] Thus, in some embodiments, inactivation of a native TSG (i.e., a TSG in the genome of a cell), e.g., a double-strand break in a sequence generated by a nuclease, results in an adverse effect on the cell. In some embodiments, inactivation of the native TSG results in cell death. In such cases, an “exogenous” TSG or a portion thereof can be introduced into the cell to compensate for the inactivated native TSG. In some embodiments, a portion of the TSG encodes a polypeptide that performs substantially the same function as the native protein encoded by the TSG. In some embodiments, a portion of the TSG is introduced to complement a partially inactivated TSG. In some embodiments, the nuclease inactivates a portion of the native TSG (e.g., by disrupting a portion of the coding sequence of the TSG), and the exogenous TSG includes the disrupted portion of the coding sequence that can be transcribed with the non-disrupted portion of the native sequence to form a functional TSG. In some embodiments, the exogenous TSG or a portion thereof is integrated into the native TSG locus in the genome of the cell. In some embodiments, the exogenous TSG or a portion thereof is integrated into a genomic locus different from the TSG locus. In some embodiments, the exogenous TSG or a portion thereof is integrated by sequences for genomic integration. In some embodiments, the sequences for genomic integration are obtained from a retroviral vector. In some embodiments, the sequences for genomic integration are obtained from a transposon. In some embodiments, the TSG encodes a CA receptor. In some embodiments, the TSG encodes HB-EGF. In some embodiments, the TSG encodes a receptor for an antibody, e.g., an antibody of an antibody-drug conjugate.

[0164] In some embodiments, the exogenous TSG is introduced into the exogenous polynucleotide of the cell. In some embodiments, the exogenous TSG is expressed from the exogenous polynucleotide. In some embodiments, the exogenous polynucleotide is a plasmid. In some embodiments, the exogenous polynucleotide is a donor polynucleotide. In some embodiments, the donor polynucleotide is a vector. Exemplary vectors are provided herein.

[0165] In some embodiments, the exogenous polynucleotide is an episomal vector. In some embodiments, the episomal vector is a stable episomal vector, i.e., an episomal vector that remains in the cell. As described herein, the episomal vector contains an autonomous DNA replication sequence that allows the episomal vector to replicate and remain in the cell. In some embodiments, the episomal vector is an artificial chromosome. In some embodiments, the episomal vector is a plasmid.

[0166] In some embodiments, the donor polynucleotide comprises 5' and 3' homology arms. In some embodiments, the donor polynucleotide is a donor plasmid. In some embodiments, the 5' and 3' homology arms of the donor polynucleotide are complementary to a portion of the TSG locus in the genome of the cell. Thus, when optimally aligned, the donor polynucleotide overlaps with one or more nucleotides of the TSG (e.g., about or at least about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 or more nucleotides). In some embodiments, when the donor polynucleotide and a portion of the TSG locus are optimally aligned, the nearest nucleotide of the donor polynucleotide is within about 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 100, 1500, 2000, 2500, 5000, 10000 or more nucleotides from the TSG locus. In some embodiments, a donor polynucleotide comprising an SOI flanked by 5' and 3' homology arms is introduced into the cell, and the 5' and 3' homology arms share sequence similarity with either side of the integration site at the TSG locus. In some embodiments, the 5' and 3' homology arms share at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence similarity with either side of the integration site at the TSG locus. In some embodiments, the TSG encodes a CA receptor. In an embodiment, the TSG encodes HB-EGF. In some embodiments, the TSG encodes a receptor for an antibody, such as an antibody of an antibody-drug conjugate.

[0167] In some embodiments, the 5' and 3' homology arms in the donor polynucleotide facilitate the integration of the donor polynucleotide into the genome by homologous recombination repair (HDR). In some embodiments, the donor polynucleotide is integrated by HDR. In some embodiments, the donor polynucleotide is an HDR template. The HDR pathway is an endogenous DNA repair pathway capable of repairing double-strand breaks. Repair by the HDR pathway typically has high fidelity and relies on homologous recombination with an HDR template having homologous regions (e.g., 5' and 3' homology arms) to the repair site. In some embodiments, the TSG locus is cleaved by a nuclease in a manner that promotes HDR, such as by generating sticky ends. In some embodiments, the TSG locus is cleaved by a nuclease in a manner that promotes HDR across a low-fidelity repair pathway such as non-homologous end joining (NHEJ).

[0168] In some embodiments, the donor polypeptide is integrated by NHEJ. The NHEJ pathway is an endogenous DNA repair pathway capable of repairing double-strand breaks. Generally, NHEJ has high repair efficiency compared to HDR, but has low fidelity, although the error is reduced when the double-strand break in the DNA has compatible sticky ends or overhangs. In some embodiments, the TSG locus is cleaved by a nuclease in a manner that reduces the error of NHEJ repair. In some embodiments, the cleavage in the TSG locus includes sticky ends.

[0169] In some embodiments, the donor polynucleotide comprises sequences for genomic integration. In some embodiments, the sequences for genomic integration at the target locus are obtained from a transposon. As described herein, a transposon comprises transposon sequences that are recognized by a transposase, which subsequently inserts a transposon comprising the transposon sequences and the sequence of interest (SOI) into the genome. In some embodiments, the target locus is any genomic locus capable of expressing the SOI without disrupting normal cellular function. Exemplary transposons are described herein. Thus, in some embodiments, the donor polynucleotide comprises a functional TSG comprising a mutation in the native coding sequence of the TSG, wherein the mutation confers resistance to a toxin, the SOI, and transposon sequences for genomic integration at the target locus. In some embodiments, the native TSG of the cell is inactivated by a nuclease, and the donor polynucleotide provides a functional TSG that is toxin resistant and capable of compensating for the native cellular function of the native TSG. In some embodiments, the TSG encodes a CA receptor. In an embodiment, the TSG encodes HB-EGF. In some embodiments, the TSG encodes a receptor for an antibody, such as an antibody of an antibody-drug conjugate.

[0170] In some embodiments, the donor polynucleotide comprises sequences for genomic integration. In some embodiments, the sequences for genomic integration at the target locus are obtained from a retroviral vector. As described herein, a retroviral vector contains sequences recognized by integrase, typically LTRs, and subsequently integrase inserts the retroviral vector containing the LTR and SOI into the genome. In some embodiments, the target locus is any genomic locus capable of expressing the SOI without disrupting normal cellular function. Exemplary retroviral vectors are described herein. Thus, in some embodiments, the donor polynucleotide comprises a functional TSG that contains a mutation in the native coding sequence of the TSG, wherein the mutation confers resistance to a toxin, the SOI, and a retroviral vector for genomic integration at the target locus. In some embodiments, the native TSG of the cell is inactivated by a nuclease, and the donor polynucleotide provides a functional TSG that can compensate for the native cellular function of the native TSG while being toxin resistant. In some embodiments, the TSG encodes a CA receptor. In an embodiment, the TSG encodes HB-EGF. In some embodiments, the TSG encodes a receptor for an antibody, such as an antibody of an antibody-drug conjugate.

[0171] In some embodiments, an episomal vector is introduced into a cell. In some embodiments, the episomal vector comprises a functional TSG that contains a mutation in the native coding sequence of the TSG, the mutation confers resistance to a toxin, a functional TSG, an SOI, and an autonomous DNA replication sequence. As described herein, an episomal vector is a non-integrating extrachromosomal plasmid capable of autonomous replication. In some embodiments, the autonomous DNA replication sequence is derived from a viral genomic sequence. In some embodiments, the autonomous DNA replication sequence is derived from a mammalian genomic sequence. In some embodiments, the episomal vector is an artificial chromosome or a plasmid. In some embodiments, the plasmid is a viral plasmid. In some embodiments, the viral plasmid is an SV40 vector, a BKV vector, a KSHV vector, or an EBV vector. Thus, in some embodiments, the native TSG of the cell is inactivated by a nuclease, and the episomal vector provides a functional TSG that is toxin-resistant and capable of compensating for the native cellular functions of the native TSG. In some embodiments, the TSG encodes a CA receptor. In an embodiment, the TSG encodes HB-EGF. In some embodiments, the TSG encodes a receptor for an antibody, such as an antibody-drug conjugate.

[0172] In some embodiments, a toxin-sensitive gene (TSG) confers toxin sensitivity to a cell, i.e., the cell is prone to harmful reactions such as growth arrest or death by the toxin. In some embodiments, the TSG encodes a receptor that binds to the toxin. In some embodiments, the receptor is a CA receptor. A CA receptor is a protein molecule that binds to CA and is typically located on the cell membrane. For example, diphtheria toxin binds to human heparin-binding EGF-like growth factor (HB-EGF). A CA receptor may be specific for one CA, or a CA receptor may bind to two or more CAs. For example, monosialoganglioside (GM 1) can act as a receptor for both cholera toxin and E. coli heat-labile enterotoxin. Alternatively, two or more CA receptors may bind to one CA. For example, botulinum toxin is thought to bind to different receptors on nerve cells and epithelial cells. In some embodiments, the CA receptor is a receptor that binds to CA. In some embodiments, the CA receptor is a G protein-coupled receptor. In some embodiments, the CA receptor binds to diphtheria toxin. In some embodiments, the CA receptor is a receptor for an antibody, such as an antibody of an antibody-drug conjugate. In some embodiments, the TSG locus contains a gene encoding heparin-binding EGF-like growth factor (HB-EGF). The mechanisms by which HB-EGF and diphtheria toxin cause cell death are described herein and are illustrated, for example, in FIG. 3A.

[0173] In some embodiments, the TSG locus contains introns and exons. In some embodiments, the double-strand break is generated by a nuclease in an intron. In some embodiments, the double-strand break is generated by a nuclease in an exon. In some embodiments, for example, a mutation in the native coding sequence of a TSG that confers toxin resistance is in an exon. In some embodiments, the donor polynucleotide contains the native coding sequence of a TSG that contains a mutation conferring toxin resistance. In some embodiments, "native coding sequence" refers to a sequence that is substantially similar to the wild-type sequence encoding a polypeptide, e.g., having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence similarity to the wild-type sequence.

[0174] In some embodiments, the donor polynucleotide comprises an exon of the native coding sequence of a TSG, the exon comprises a mutation that confers toxin resistance, and the donor polynucleotide further comprises a splicing acceptor sequence. As used herein, "splicing acceptor" or "splicing acceptor sequence" refers to a sequence at the 3' end of an intron that facilitates the joining of two exons flanking the intron. In some embodiments, the splicing acceptor sequence has at least about 90% sequence identity with the splicing acceptor sequence of the TSG locus in the genome of the cell. In some embodiments, the exon integrated into the TSG locus from the donor polynucleotide binds to an adjacent exon in the genome of the cell when the TSG is transcribed for expression. In some embodiments, the splicing acceptor sequence integrated into the TSG locus from the donor polynucleotide facilitates the binding of an exon adjacent to the exon integrated into the TSG locus from the donor polynucleotide in the genome of the cell.

[0175] In some embodiments, the 5’ and 3’ homology arms of the donor polynucleotide are complementary to a portion of the TSG locus in the genome of the cell. Thus, when optimally aligned, the donor polynucleotide overlaps with one or more nucleotides of the TSG (e.g., about or at least about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 or more nucleotides). In some embodiments, when the donor polynucleotide and the portion of the TSG locus are optimally aligned, the nearest nucleotide of the donor polynucleotide is within about 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 100, 1500, 2000, 2500, 5000, 10000 or more nucleotides from the TSG locus. In some embodiments, a donor polynucleotide comprising an SOI flanked by 5’ and 3’ homology arms is introduced into the cell, and the 5’ and 3’ homology arms share sequence similarity with the sides of either integration site at the TSG locus. In some embodiments, the integration site at the TSG locus is a nuclease cleavage site, i.e., a double-strand break site. In some embodiments, the 5’ and 3’ homology arms share at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence similarity with the sides of either integration site at the TSG locus. In some embodiments, the integration site at the TSG locus is a nuclease cleavage site. In some embodiments, the TSG encodes a CA receptor. In embodiments, the TSG encodes HB-EGF.

[0176] In some embodiments, the TSG encodes HB-EGF and a double-strand break is generated in the intron of the HB-EGF gene. In some embodiments, the TSG encodes HB-EGF and a double-strand break is generated in the exon of the HB-EGF gene. In some embodiments, the double-strand break is in the intron of the HB-EGF gene and the mutation in the native coding sequence of the HB-EGF gene is in the exon of the HB-EGF gene. In some embodiments, the double-strand break is in the intron of the HB-EGF gene and the mutation in the native coding sequence of the HB-EGF gene is in the exon immediately following the cleaved intron. In some embodiments, the double-strand break is in the exon of the HB-EGF gene and the mutation in the native coding sequence of the HB-EGF gene is in the same exon of the HB-EGF gene. In some embodiments, the double-strand break is in the exon of the HB-EGF gene and the mutation in the native coding sequence of the HB-EGF gene is in a different exon of the HB-EGF gene.

[0177] In some embodiments, the 5' and 3' homology arms of the donor polynucleotide share sequence similarity with HB-EGF at the nuclease cleavage site. In some embodiments, the double-strand break is in an intron of HB-EGF and the 5' and 3' homology arms include homology to the intron sequence. In some embodiments, the double-strand break is in an exon of HB-EGF and the 5' and 3' homology arms include homology to the exon sequence. In some embodiments, the 5' and 3' homology arms of the donor polynucleotide are designed to insert the donor polynucleotide at the double-strand break site, for example, by HDR. In some embodiments, the 5' and 3' homology arms have at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence similarity with either side of a nuclease (e.g., Cas9) cleavage site in HB-EGF.

[0178] In some embodiments, the native coding sequence comprises one or more alterations to the wild-type sequence, but the polypeptide encoded by the native coding sequence is substantially similar to the polypeptide encoded by the wild-type sequence. For example, the amino acid sequence of the polypeptide is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identical. In some embodiments, the polypeptides encoded by the native coding sequence and the wild-type sequence have similar structures, such as similar overall shapes and folds, as determined by one of ordinary skill in the art. In some embodiments, the native coding sequence comprises a portion of the wild-type sequence. For example, the native coding sequence is substantially similar to one or more exons and / or one or more introns of the wild-type sequence that encode a protein, such that the exons and / or introns of the native coding sequence can replace the corresponding wild-type exons and / or introns for encoding a polypeptide with a sequence and / or structure that is substantially identical to the wild-type polypeptide. In some embodiments, the native coding sequence comprises a mutation to the wild-type sequence. In some embodiments, the mutation in the native coding sequence of the TSG is in an exon.

[0179] In some embodiments, the donor polynucleotide is a functional TSG comprising a mutation in the native coding sequence of the TSG, the mutation conferring resistance to a toxin, a SOI, and a sequence for genomic integration at the target locus. The term "functional" TSG refers to a TSG that encodes a polypeptide that is substantially similar to the polypeptide encoded by the native coding sequence. In some embodiments, the functional TSG comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence similarity to the native coding sequence of the TSG, and further comprises a mutation in the native coding sequence of the TSG that confers toxin resistance. In some embodiments, the polypeptide encoded by the functional TSG has substantially the same structure as the polypeptide encoded by the native coding sequence and performs the same cellular functions, except that the polypeptide encoded by the functional TSG is toxin resistant. In some embodiments, the polypeptide encoded by the functional TSG loses its ability to bind to the toxin. In some embodiments, the polypeptide encoded by the functional TSG loses its ability to transport and / or move the toxin intracellularly.

[0180] In some embodiments, the mutation in the native coding sequence of the TSG is a substitution mutation, an insertion, or a deletion. In some embodiments, the mutation is a nucleotide substitution in the coding sequence of the TSG that changes a single amino acid in the encoded polypeptide sequence. In some embodiments, the mutation is a substitution of one or more nucleotides that changes one or more amino acids in the encoded polypeptide sequence. In some embodiments, the mutation is a substitution of one or more nucleotides that changes an amino acid codon to a stop codon. In some embodiments, the mutation is a nucleotide insertion in the coding sequence of the TSG that results in the insertion of one or more amino acids in the encoded polypeptide sequence. In some embodiments, the mutation is a nucleotide deletion in the coding sequence of the TSG that results in the deletion of one or more amino acids in the encoded polypeptide sequence.

[0181] In some embodiments, the mutation in the native coding sequence of the TSG is a mutation in the toxin-binding region of the protein encoded by the TSG. In some embodiments, as a result of the mutation in the toxin-binding region, the protein loses its ability to bind the toxin. In some embodiments, the protein encoded by the functional TSG has substantially the same structure as the protein encoded by the native coding sequence and performs the same cellular functions, except that the protein encoded by the functional TSG containing the mutation is toxin-resistant. In some embodiments, the protein encoded by the functional TSG loses its ability to bind the toxin. In some embodiments, the protein encoded by the functional TSG loses its ability to transport and / or move the toxin intracellularly.

[0182] In some embodiments, the TSG encodes a receptor that binds to a toxin. In some embodiments, the receptor is a CA receptor. In some embodiments, the TSG encodes a receptor that binds to diphtheria toxin. In some embodiments, the TSG encodes heparin-binding EGF-like growth factor (HB-EGF). In some embodiments, a mutation in the native coding sequence of the TSG renders cells resistant to diphtheria toxin.

[0183] In some embodiments, the toxin is a natural toxin. In some embodiments, the toxin is a synthetic poison. In some embodiments, the toxin is a small molecule, peptide, or protein. In some embodiments, the toxin is an antibody-drug conjugate. In some embodiments, the toxin is a monoclonal antibody conjugated to a biologically active agent via a labile linker. In some embodiments, the toxin is a biotoxin. In some embodiments, the toxin is produced by cyanobacteria (cyanotoxin), dinoflagellates (dinotoxin), spiders, snakes, scorpions, frogs, jellyfish, etc., sea mammals, poisonous fish, corals, or octopuses. Examples of toxins include, for example, diphtheria toxin, botulinum toxin, ricin, apitoxin, saxitoxin, Pseudomonas exotoxin, and mycotoxin. In some embodiments, the toxin is diphtheria toxin. In some embodiments, the toxin is an antibody-drug conjugate.

[0184] In some embodiments, the toxin is toxic to one organism, such as a human, but not to another organism, such as a mouse. In some embodiments, the toxin is toxic to an organism at a certain stage of its life cycle (e.g., fetal stage), but not at another stage of the organism's life cycle (e.g., adult stage). In some embodiments, the toxin is toxic within one organ of an animal, but not to another organ of the same animal. In some embodiments, the toxin is toxic to a subject (e.g., a human or an animal) in a certain condition or state (e.g., being ill), but not to the same subject in another condition or state (e.g., being healthy). In some embodiments, the toxin is toxic to one cell type, but not to another cell type. In some embodiments, the toxin is toxic to cells in a certain cell state (e.g., differentiated), but not to the same cells in another cell state (e.g., undifferentiated). In some embodiments, the toxin is toxic to cells in one environment (e.g., low temperature), but not to the same cells in another environment (e.g., high temperature). In some embodiments, the toxin is toxic to human cells, but not to mouse cells.

[0185] In some embodiments, mutations in one or more of amino acids 100 to 160 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, mutations in one or more of amino acids 105 to 150 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, mutations in one or more of amino acids 107 to 148 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, mutations in one or more of amino acids 120 to 145 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, mutations in one or more of amino acids 135 to 143 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, mutations in one or more of amino acids 138 to 144 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, a mutation at amino acid 141 of wild-type HB-EGF (SEQ ID NO: 8) confers diphtheria toxin resistance. In some embodiments, the mutation at amino acid 141 of wild-type HB-EGF (SEQ ID NO: 8) is from GLU141 to ARG141. In some embodiments, the mutation at amino acid 141 of wild-type HB-EGF (SEQ ID NO: 8) is from GLU141 to HIS141. In some embodiments, the mutation at amino acid 141 of wild-type HB-EGF (SEQ ID NO: 8) is from GLU141 to LYS141. In some embodiments, the mutation from GLU141 to LYS141 in wild-type HB-EGF (SEQ ID NO: 8) confers diphtheria toxin resistance.

[0186] Thus, in some embodiments, the mutation in the native coding sequence of the TSG is a mutation in one or more of amino acids 100-160 of HB-EGF (SEQ ID NO: 8). In some embodiments, the mutation in the native coding sequence of the TSG is a mutation in one or more of amino acids 105-150 of HB-EGF (SEQ ID NO: 8). In some embodiments, the mutation in the native coding sequence of the TSG is a mutation in one or more of amino acids 107-148 of HB-EGF (SEQ ID NO: 8). In some embodiments, the mutation in the native coding sequence of the TSG is a mutation in one or more of amino acids 120-145 of HB-EGF (SEQ ID NO: 8). In some embodiments, the mutation in the native coding sequence of the TSG is a mutation in one or more of amino acids 135-143 of HB-EGF (SEQ ID NO: 8). In some embodiments, the mutation in the native coding sequence of the TSG is a mutation in one or more of amino acids 138-144 of wild-type HB-EGF (SEQ ID NO: 8). In some embodiments, the mutation in the native coding sequence of the TSG is a mutation at amino acid 141 of HB-EGF (SEQ ID NO: 8). In some embodiments, the mutation in the native coding sequence of the TSG is a mutation from GLU141 to LYS141 in HB-EGF (SEQ ID NO: 8). In some embodiments, the mutation in the native coding sequence of the TSG is a mutation from GLU141 to HIS141 in HB-EGF (SEQ ID NO: 8). In some embodiments, the mutation in the native coding sequence of the TSG is a mutation from GLU141 to ARG141 in HB-EGF (SEQ ID NO: 8). In some embodiments, the mutation from GLU141 to LYS141 in HB-EGF (SEQ ID NO: 8) confers diphtheria toxin resistance.

[0187] In some embodiments, the functional TSG in the donor polynucleotide or episomal vector is resistant to nuclease inactivation. In some embodiments, the functional TSG comprises one or more mutations in the native coding sequence of the TSG, and the one or more mutations confer resistance to nuclease inactivation. In some embodiments, the functional TSG does not bind to the nuclease. In some embodiments, a TSG that does not bind to the nuclease is less susceptible to cleavage by the nuclease. As described herein, nucleases such as certain types of Cas9 may require a PAM sequence in the target sequence, or in the vicinity thereof, in addition to recognition of the target sequence by a guide polynucleotide (e.g., guide RNA) via hybridization. In some embodiments, Cas9 binds to the PAM sequence prior to initiating nuclease activity. In some embodiments, a target sequence that does not contain a PAM in the target sequence, or in an adjacent or neighboring region, does not bind to the nuclease. Thus, in some embodiments, a target sequence that does not contain a PAM in the target sequence, or in an adjacent or neighboring region, is not cleaved by the nuclease and is therefore resistant to nuclease inactivation. In some embodiments, the functional TSG does not contain a PAM sequence. In some embodiments, a TSG that does not contain a PAM sequence is resistant to nuclease inactivation.

[0188] In some embodiments, the PAM is within about 30 to about 1 nucleotide of the target sequence. In some embodiments, the PAM is within about 20 to about 2 nucleotides of the target sequence. In some embodiments, the PAM is within about 10 to about 3 nucleotides of the target sequence. In some embodiments, the PAM is within about 10, about 9, about 8, about 7, about 6, about 5, about 4, about 3, about 2, or about 1 nucleotide of the target sequence. In some embodiments, the PAM is present upstream (i.e., in the 5' direction) of the target sequence. In some embodiments, the PAM is present downstream (i.e., in the 3' direction) of the target sequence. In some embodiments, the PAM is located within the target sequence.

[0189] In some embodiments, the polypeptide encoded by the functional TSG cannot hybridize with the guide polynucleotide. In some embodiments, a TSG that does not hybridize with the guide polynucleotide is difficult to cleave by a nuclease such as Cas9. As described herein, the guide polynucleotide can hybridize with the target sequence, i.e., is "recognized" by the guide polynucleotide for cleavage by a nuclease such as Cas9. Thus, a sequence that does not hybridize with the guide polynucleotide is not recognized for cleavage by a nuclease such as Cas9. In some embodiments, a sequence that does not hybridize with the guide polynucleotide is resistant to nuclease-mediated inactivation. In some embodiments, the guide polynucleotide can hybridize with the TSG in the genome of the cell, and the functional TSG on the donor polynucleotide or episomal vector contains one or more mutations in the native coding sequence of the TSG, such that the guide polynucleotide can (1) hybridize with the TSG in the genome of the cell and (2) cannot hybridize with the functional TSG on the donor polynucleotide or episomal vector. In some embodiments, a functional TSG that is resistant to nuclease-mediated inactivation is introduced into the cell simultaneously with a nuclease that targets the ExG in the genome of the cell.

[0190] In some embodiments, the SOI comprises a polynucleotide encoding a protein. In some embodiments, the SOI comprises a mutant gene. In some embodiments, the SOI comprises a non-coding sequence, such as a microRNA. In some embodiments, the SOI is operably linked to a regulatory element. In some embodiments, the SOI is a regulatory element. In some embodiments, the SOI comprises a resistance cassette, such as a gene conferring resistance to an antibiotic. In some embodiments, the SOI comprises a marker, such as a selectable marker or a selectable marker. In some embodiments, the SOI comprises a marker, such as a restriction site, a fluorescent protein, or a selectable marker.

[0191] In some embodiments, the SOI contains a mutation of the wild-type gene in the genome of the cell. In some embodiments, the mutation includes a point mutation, i.e., a substitution of a single nucleotide. In some embodiments, the mutation includes a substitution of multiple nucleotides. In some embodiments, the mutation introduces a stop codon. In some embodiments, the mutation includes an insertion of a nucleotide in the wild-type sequence. In some embodiments, the mutation includes a deletion of a nucleotide in the wild-type sequence. In some embodiments, the mutation includes a frameshift mutation.

[0192] In some embodiments, after introducing the nuclease, the guide polynucleotide, and the donor polynucleotide or episomal vector, the cell population is contacted with a toxin. Examples of toxins are provided herein. In some embodiments, the toxin is a natural toxin. In some embodiments, the toxin is a synthetic poison. In some embodiments, the toxin is a small molecule, a peptide, or a protein. In some embodiments, the toxin is an antibody-drug conjugate. In some embodiments, the toxin is a monoclonal antibody conjugated to a biologically active agent via a labile linker. In some embodiments, the toxin is a biotoxin. In some embodiments, the toxin is produced by cyanobacteria (cyanotoxin), dinoflagellates (dinotoxin), spiders, snakes, scorpions, frogs, jellyfish, sea mammals, poisonous fish, corals, or octopuses. Examples of toxins include, for example, diphtheria toxin, botulinum toxin, ricin, apitoxin, saxitoxin, Pseudomonas exotoxin, and mycotoxin. In some embodiments, the toxin is diphtheria toxin. In some embodiments, the toxin is an antibody-drug conjugate.

[0193] In some embodiments, the toxin is toxic to one organism, such as a human, but not to another organism, such as a mouse. In some embodiments, the toxin is toxic to an organism during a certain stage of its life cycle (e.g., fetal stage), but not during another life cycle of the organism (e.g., adult stage). In some embodiments, the toxin is toxic within one organ of an animal, but not to another organ of the same animal. In some embodiments, the toxin is toxic to a subject (e.g., a human or an animal) in a certain condition or state (e.g., being ill), but not to the same subject in another condition or state (e.g., being healthy). In some embodiments, the toxin is toxic to one cell type, but not to another cell type. In some embodiments, the toxin is toxic to cells in a certain cell state (e.g., differentiated), but not to the same cells in another cell state (e.g., undifferentiated). In some embodiments, the toxin is toxic to cells in a certain environment (e.g., low temperature), but not to the same cells in another environment (e.g., high temperature). In some embodiments, the toxin is toxic to human cells, but not to mouse cells. In some embodiments, the toxin is diphtheria toxin. In some embodiments, the toxin is an antibody-drug conjugate.

[0194] In some embodiments, after contacting a cell population with a toxin, one or more cells resistant to the toxin are selected. In some embodiments, the one or more cells resistant to the toxin are viable cells. In some embodiments, the viable cells have (1) an inactivated native TSG (e.g., inactivated by double-strand breaks generated by a nuclease), and (2) a functional TSG containing a mutation that confers toxin resistance. Cells that meet only one of the above two conditions are prone to cell death: when the native TSG is inactivated, the cells are sensitive to the toxin and die upon contact with the toxin; when the functional TSG is not introduced, the cells lack the normal cell functions of the TSG and die due to the absence of normal cell functions.

[0195] In embodiments involving the introduction of a donor polynucleotide comprising 5' and 3' homology arms (e.g., the homologous sequences for HDR), viable cells comprise the integration of two alleles of the donor polynucleotide comprising the SOI at the native TSG locus, the native TSG is disrupted by the integration of the donor polynucleotide, and the cells comprise a functional toxin-resistant TSG. Thus, in such embodiments, one or more cells resistant to the toxin comprise the integration of two alleles of the SOI. In embodiments involving the introduction of a donor polynucleotide comprising a sequence for genomic integration (e.g., a transposon, a lentiviral vector sequence, or a retroviral vector sequence) at the target locus, viable cells comprise the integration of the donor polynucleotide comprising an inactivated native TSG, as well as a functional toxin-resistant TSG and the SOI at the target locus. In such embodiments, one or more cells resistant to the toxin comprise the SOI integrated into the target locus. In embodiments involving the introduction of an episomal vector, viable cells comprise an inactivated native TSG, as well as a stable episomal vector comprising a functional toxin-resistant TSG and the SOI. In such embodiments, one or more cells resistant to the toxin comprise the episomal vector.

[0196] Method for providing diphtheria toxin resistance In some embodiments, the disclosure provides a method for providing resistance to diphtheria toxin in a human cell, the method comprising introducing into the cell: (i) a base editing enzyme; and (ii) a guide polynucleotide that targets the heparin-binding EGF-like growth factor (HB-EGF) receptor in the human cell, wherein the base editing enzyme forms a complex with the guide polynucleotide, the base editing enzyme is targeted to HB-EGF, and the base editing enzyme provides site-specific mutations in HB-EGF to provide resistance to diphtheria toxin in the human cell.

[0197] In some embodiments, the human cells are a human cell line. In some embodiments, the human cells are stem cells. The stem cells can be, for example, pluripotent stem cells, such as embryonic stem cells (ESCs), adult stem cells, induced pluripotent stem cells (iPSCs), tissue-specific stem cells (e.g., hematopoietic stem cells), and mesenchymal stem cells (MSCs). In some embodiments, the human cells are any differentiated form of the cells described herein. In some embodiments, the eukaryotic cells are cells derived from primary cells under culture. In some embodiments, the cells are stem cells or a stem cell line. In some embodiments, the human cells are hepatocytes, such as human hepatocytes, animal hepatocytes, or non-parenchymal cells. For example, the eukaryotic cells can be human hepatocytes for adherent metabolism assays, human hepatocytes for adherent induction assays, adherent QUALYST TRANSPORTER CERTIFIED human hepatocytes, human hepatocytes for suspension assays (e.g., 10 donor and 20 donor pool hepatocytes), human liver Kupffer cells, or human liver stellate cells. In some embodiments, the human cells are immune cells. In some embodiments, the immune cells are granulocytes, mast cells, monocytes, dendritic cells, natural killer cells, B cells, primary T cells, cytotoxic T cells, helper T cells, CD8+ T cells, CD4+ T cells, or regulatory T cells.

[0198] In some embodiments, human cells are xenografted or transplanted into non-human animals. In some embodiments, the non-human animal is a mouse, rat, hamster, guinea pig, rabbit, or pig. In some embodiments, the human cells are cells in a humanized organ of a non-human animal. In some embodiments, a "humanized" organ refers to a human organ grown in an animal. In some embodiments, a "humanized" organ refers to an organ generated by an animal, from which its animal-specific cells have been removed and which has been grafted with human cells. A humanized organ can be immunocompatible with a human. In some embodiments, the humanized organ is a liver, kidney, pancreas, heart, lung, or stomach. Humanized organs are very useful for the study and modeling of human diseases. However, most gene selection tools cannot be converted into humanized organs in host animals because most selection markers are harmful to the host animal. Humanized organs are further described, for example, in Garry et al., Regen Med 11(7):617-619, Garry et al., Circ Res 124:23-25(2019), and Nguyen et al., Drug Discov Today 23(11):1812-1817(2018).

[0199] The present disclosure provides a highly advantageous selection method that can be used for humanized cells in an animal host by using diphtheria toxin, which is toxic to humans but not to mice. However, the method of the present invention is not limited to diphtheria toxin and can be used in conjunction with any compound with differential toxicity, i.e., toxic to one organism but not to another. The method of the present invention further provides diphtheria toxin resistance by engineering the receptor for the toxin, which may be desirable in situations where the toxin does not enter the cell, in contrast to conventional methods that focus on diphthamide biosynthesis protein 2 (DPH2) (see, for example, Picco et al., Sci Rep 5:14721).

[0200] In some embodiments, the humanized organ is generated by transplanting human cells into an animal. In some embodiments, the animal is an immunodeficient mouse. In some embodiments, the animal is an immunodeficient adult mouse. In some embodiments, the humanized organ is generated by suppressing one or more animal genes and expressing one or more human genes in the animal's organ. In some embodiments, the humanized organ is a liver. In some embodiments, the humanized organ is a pancreas. In some embodiments, the humanized organ is a heart. In some embodiments, the humanized organ expresses a human gene encoding a receptor for a cytotoxic agent, namely the CA receptor described herein. In some embodiments, the humanized organ is toxin-sensitive, but the rest of the animal is toxin-resistant. In some embodiments, the humanized organ expresses human HB-EGF. In some embodiments, the humanized organ is diphtheria toxin-sensitive, but the rest of the animal is diphtheria toxin-resistant. In some embodiments, the humanized organ is a humanized liver in a mouse, the humanized liver expresses human HB-EGF and is sensitive to diphtheria toxin, but the rest of the mouse is HB-EGF-resistant. Thus, upon exposure to diphtheria toxin, only the humanized cells in the mouse liver die.

[0201] In some embodiments, the base editing enzyme comprises a DNA targeting domain and a DNA editing domain. In some embodiments, the DNA targeting domain comprises Cas9. The Cas9 protein is described herein. In some embodiments, Cas9 comprises a mutation in the catalytic domain. In some embodiments, the base editing enzyme comprises catalytically inactive Cas9 (dCas9) and a DNA editing domain. In some embodiments, nCas9 comprises mutations relative to wild-type Cas9 at amino acid residues D10 and H840 (numbering is relative to SEQ ID NO: 3). In some embodiments, the base editing enzyme comprises Cas9 (nCas9) capable of generating a single-stranded DNA break, and a DNA editing domain. In some embodiments, nCas9 comprises a mutation relative to wild-type Cas9 at amino acid residue D10 or H840 (numbering is relative to SEQ ID NO: 3). In some embodiments, Cas9 comprises a polypeptide having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence identity with SEQ ID NO: 3. In some embodiments, Cas9 comprises a polypeptide having at least 90% sequence identity to SEQ ID NO: 3. In some embodiments, Cas9 comprises a polypeptide having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence identity with SEQ ID NO: 4. In some embodiments, Cas9 comprises a polypeptide having at least 90% sequence identity to SEQ ID NO: 4.

[0202] In some embodiments, the DNA editing domain comprises a deaminase. In some embodiments, the deaminase is a cytidine deaminase or an adenosine deaminase. In some embodiments, the deaminase is a cytidine deaminase. In some embodiments, the deaminase is an adenosine deaminase. In some embodiments, the deaminase is an apolipoprotein B mRNA editing complex (APOBEC) deaminase, activation-induced cytidine deaminase (AID), ACF1 / ASE deaminase, ADAT deaminase, or ADAR deaminase. In some embodiments, the deaminase is an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the deaminase is APOBEC1.

[0203] In some embodiments, the base editing enzyme further comprises a DNA glycosylase inhibitor domain. In some embodiments, the DNA glycosylase is uracil DNA glycosylase inhibitor (UGI). Generally, DNA glycosylases such as uracil DNA glycosylase are part of the base excision repair pathway and, when detecting a U:G mismatch (where "U" is generated from the deamination of cytosine), perform error-free repair, converting the U back to the wild-type sequence and effectively "undoing" the base editing. Thus, by adding a DNA glycosylase inhibitor (e.g., uracil DNA glycosylase inhibitor), the base excision repair pathway is inhibited and the base editing efficiency is increased. Non-limiting examples of DNA glycosylases include OGG1, MAG1, and UNG. The DNA glycosylase inhibitor can be a small molecule or a protein. For example, protein inhibitors of uracil DNA glycosylase are described in Mol et al., Cell 82:701-708 (1995), Serrano-Heras et al., J Biol Chem 281:7068-7074 (2006), and New England Biolabs Catalog No. M0281S and M0281L (neb.com / products / m0281-uracil-glycosylase-inhibitor-ugi). Small molecule inhibitors of DNA glycosylases are described, for example, in Huang et al., J Am Chem Soc 131(4):1344-1345 (2009), Jacobs et al., PLoS One 8(12):e81667 (2013), Donley et al., ACS Chem Biol 10(10):2334-2343 (2015), Tahara et al., J Am Chem Soc 140(6):2105-2114 (2018).

[0204] Thus, in some embodiments, the base editing enzyme of the present disclosure comprises nCas9 and a cytidine deaminase. In some embodiments, the base editing enzyme of the present disclosure comprises nCas9 and an adenosine deaminase. In some embodiments, the base editing enzyme comprises a polypeptide having at least 90% sequence identity to SEQ ID NO: 6. In some embodiments, the base editing enzyme comprises a polypeptide having at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, or at least 90% sequence identity to SEQ ID NO: 6. In some embodiments, the base editing enzyme is at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identical to SEQ ID NO: 6. In some embodiments, the polynucleotide encoding the base editing enzyme is at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identical to SEQ ID NO: 5. In some embodiments, the base editing enzyme is BE3.

[0205] In some embodiments, the method of the present disclosure comprises introducing into a human cell a guide polynucleotide that targets the HB-EGF receptor in the human cell. In some embodiments, the guide polynucleotide forms a complex with the base editing enzyme, and the base editing enzyme is targeted to HB-EGF by the guide polynucleotide and provides diphtheria toxin resistance in the human cell by providing a site-specific mutation in HB-EGF.

[0206] In some embodiments, the guide polynucleotide is an RNA molecule. The guide polynucleotide may be introduced into a target cell as an isolated molecule, e.g., an RNA molecule, or may be introduced into the cell using an expression vector containing DNA encoding the guide polynucleotide, e.g., an RNA guide polynucleotide. In some embodiments, the guide polynucleotide is 10 to 150 nucleotides. In some embodiments, the guide polynucleotide is 20 to 120 nucleotides. In some embodiments, the guide polynucleotide is 30 to 100 nucleotides. In some embodiments, the guide polynucleotide is 40 to 80 nucleotides. In some embodiments, the guide polynucleotide is 50 to 60 nucleotides. In some embodiments, the guide polynucleotide is 10 to 35 nucleotides. In some embodiments, the guide polynucleotide is 15 to 30 nucleotides. In some embodiments, the guide polynucleotide is 20 to 25 nucleotides.

[0207] In some embodiments, the RNA guide polynucleotide comprises at least two nucleotide segments: at least one "DNA binding segment" and at least one "polypeptide binding segment". A "segment" means a part, section, or region of a molecule, e.g., a continuous stretch of nucleotides of a guide polynucleotide molecule. The definition of "segment" is not limited to a specific number of total base pairs unless specifically defined otherwise.

[0208] In some embodiments, the guide polynucleotide comprises a DNA binding segment. In some embodiments, the DNA binding segment of the guide polynucleotide comprises a nucleotide sequence complementary to a specific sequence within the target polynucleotide. In some embodiments, the DNA binding segment of the guide polynucleotide hybridizes to a gene encoding a cytotoxic agent (CA) receptor in the target cell. In some embodiments, the DNA binding segment of the guide polynucleotide hybridizes to the gene encoding HB-EGF. In some embodiments, the DNA binding segment of the guide polynucleotide hybridizes to a target polynucleotide sequence in the target cell. Target cells including various types of eukaryotic cells are described herein.

[0209] In some embodiments, the guide polynucleotide comprises a polypeptide binding segment. In some embodiments, the polypeptide binding segment of the guide polynucleotide binds to the DNA targeting domain of the base editing enzyme of the present disclosure. In some embodiments, the polypeptide binding segment of the guide polynucleotide binds to Cas9 of the base editing enzyme. In some embodiments, the polypeptide binding segment of the guide polynucleotide binds to dCas9 of the base editing enzyme. In some embodiments, the polypeptide binding segment of the guide polynucleotide binds to nCas9 of the base editing enzyme. Various RNA guide polynucleotides that bind to the Cas9 protein are described, for example, in US Patent Application Publication Nos. 2014 / 0068797, 2014 / 0273037, 2014 / 0273226, 2014 / 0295556, 2014 / 0295557, 2014 / 0349405, 2015 / 0045546, 2015 / 0071898, 2015 / 0071899, and 2015 / 0071906.

[0210] In some embodiments, the guide polynucleotide further comprises a tracrRNA. A "tracrRNA" or trans-activating CRISPR-RNA forms an RNA duplex with a pre-crRNA, or pre-CRISPR-RNA, which is then cleaved by the RNA-specific ribonuclease RNase III to form a crRNA / tracrRNA hybrid. In some embodiments, the guide polynucleotide comprises a crRNA / tracrRNA hybrid. In some embodiments, the tracrRNA component of the guide polynucleotide activates the Cas9 protein. In some embodiments, activation of the Cas9 protein comprises activating the nuclease activity of Cas9. In some embodiments, activation of the Cas9 protein comprises a Cas9 protein that binds to a target polynucleotide sequence.

[0211] In some embodiments, the sequence of the guide polynucleotide is designed to target a base editing enzyme to a specific location in the target polynucleotide sequence. Various tools and programs are available to facilitate the design of such guide polynucleotides (see, e.g., the Benchling Base Editor Design Guide (benchling.com / editor#create / crispr), as well as the BE-Designer and BE-Analyzer from CRISPR RGEN Tools (see Hwang et al., bioRxiv dx.doi.org / 10.1101 / 373944, first published July 22, 2018)).

[0212] In some embodiments, the DNA-binding segment of the guide polynucleotide hybridizes with the gene encoding HB-EGF, and the polypeptide-binding segment of the guide polynucleotide forms a complex with the base editing enzyme by binding to the DNA targeting domain of the base editing enzyme. In some embodiments, the DNA-binding segment of the guide polynucleotide hybridizes with the gene encoding HB-EGF, and the polypeptide-binding segment of the guide polynucleotide forms a complex with the base editing enzyme by binding to Cas9 of the base editing enzyme. In some embodiments, the DNA-binding segment of the guide polynucleotide hybridizes with the gene encoding HB-EGF, and the polypeptide-binding segment of the guide polynucleotide forms a complex with the base editing enzyme by binding to dCas9 of the base editing enzyme. In some embodiments, the DNA-binding segment of the guide polynucleotide hybridizes with the gene encoding HB-EGF, and the polypeptide-binding segment of the guide polynucleotide forms a complex with the base editing enzyme by binding to nCas9 of the base editing enzyme.

[0213] In some embodiments, the complex is targeted to HB-EGF by the guide polynucleotide, and the base editing enzyme of the complex introduces a mutation into HB-EGF. In some embodiments, the mutation in HB-EGF is introduced by the base editing domain of the base editing enzyme of the complex. In some embodiments, the mutation in HB-EGF forms diphtheria toxin-resistant cells. In some embodiments, the mutation is a point mutation from cytosine (C) to thymine (T). In some embodiments, the mutation is a point mutation from adenine (A) to guanine (G). The specific location of the mutation in HB-EGF can be induced, for example, by the design of the guide polynucleotide using tools such as the Benchling base editor design guides, BE-Designer, and BE-Analyzer described herein. In some embodiments, the guide polynucleotide is an RNA polynucleotide. In some embodiments, the guide polynucleotide further comprises a tracrRNA sequence.

[0214] In some embodiments, the site-specific mutation is in the region of HB-EGF that binds to diphtheria toxin. In some embodiments, the mutation in the EGF-like domain of HB-EGF confers diphtheria toxin resistance. In some embodiments, a charge-reversal mutation of an amino acid at or near the diphtheria toxin binding site of HB-EGF confers diphtheria toxin resistance. In some embodiments, the charge-reversal mutation is a substitution from a negatively charged residue such as Glu or Asp to a positively charged residue such as Lys or Arg. In some embodiments, the charge-reversal mutation is a substitution from a positively charged residue such as Lys or Arg to a negatively charged residue such as Glu or Asp. In some embodiments, a polarity-reversal mutation of an amino acid at or near the diphtheria toxin binding site of HB-EGF confers diphtheria toxin resistance. In some embodiments, the polarity-reversal mutation is a substitution from a polar amino acid residue such as Gln or Asn to a nonpolar amino acid residue such as Ala, Val, or Ile. In some embodiments, the polarity-reversal mutation is a substitution from a nonpolar amino acid residue such as Ala, Val, or Ile to a polar amino acid residue such as Gln or Asn. In some embodiments, the mutation is a substitution of a relatively small amino acid residue such as Gly or Ala at or near the diphtheria toxin binding site of HB-EGF with a "bulky" amino acid residue such as Trp. In some embodiments, the mutation from a small residue to a bulky residue blocks the binding pocket and prevents the binding of diphtheria toxin, thereby conferring resistance.

[0215] In some embodiments, mutations in one or more of amino acids 100 - 160 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, mutations in one or more of amino acids 105 - 150 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, mutations in one or more of amino acids 107 - 148 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, mutations in one or more of amino acids 120 - 145 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, mutations in one or more of amino acids 135 - 143 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, mutations in one or more of amino acids 138 - 144 of wild-type HB-EGF (SEQ ID NO: 8) confer diphtheria toxin resistance. In some embodiments, a mutation at amino acid 141 of wild-type HB-EGF (SEQ ID NO: 8) confers diphtheria toxin resistance. In some embodiments, the mutation at amino acid 141 of wild-type HB-EGF (SEQ ID NO: 8) is from GLU141 to ARG141. In some embodiments, the mutation at amino acid 141 of wild-type HB-EGF (SEQ ID NO: 8) is from GLU141 to HIS141. In some embodiments, the mutation at amino acid 141 of wild-type HB-EGF (SEQ ID NO: 8) is from GLU141 to LYS141. In some embodiments, the mutation from GLU141 to LYS141 in wild-type HB-EGF (SEQ ID NO: 8) confers diphtheria toxin resistance.

[0216] Thus, in some embodiments, the site-specific mutation is in one or more of amino acids 100-160 of HB-EGF (SEQ ID NO: 8). In some embodiments, the site-specific mutation is in one or more of amino acids 105-150 of HB-EGF (SEQ ID NO: 8). In some embodiments, the site-specific mutation is in one or more of amino acids 107-148 of HB-EGF (SEQ ID NO: 8). In some embodiments, the site-specific mutation is in one or more of amino acids 120-145 of HB-EGF (SEQ ID NO: 8). In some embodiments, the site-specific mutation is in one or more of amino acids 135-143 of HB-EGF (SEQ ID NO: 8). In some embodiments, the site-specific mutation is in one or more of amino acids 138-144 of wild-type HB-EGF (SEQ ID NO: 8). In some embodiments, the site-specific mutation is at amino acid 141 of HB-EGF (SEQ ID NO: 8). In some embodiments, the site-specific mutation is a mutation from GLU141 to LYS141 in HB-EGF (SEQ ID NO: 8). In some embodiments, the site-specific mutation is a mutation from GLU141 to HIS141 in HB-EGF (SEQ ID NO: 8). In some embodiments, the site-specific mutation is a mutation from GLU141 to ARG141 in HB-EGF (SEQ ID NO: 8). In some embodiments, the mutation from GLU141 to LYS141 in HB-EGF (SEQ ID NO: 8) confers diphtheria toxin resistance.

[0217] Selection method using essential genes The methods of the present disclosure are not necessarily limited to selection using toxin sensitivity genes. Essential genes are genes of an organism that are considered indispensable for survival under specific conditions. In embodiments, the essential genes are used as the "selection" site in the co-targeting enrichment strategy described herein.

[0218] In some embodiments, the present disclosure provides a method for integrating and enriching a sequence of interest (SOI) into a mammalian genomic target locus in the genome of a cell, the method comprising: (a) introducing into a cell population: (i) a nuclease capable of generating a double-strand break; (ii) a guide polynucleotide capable of forming a complex with the nuclease and hybridizing to and inactivating an essential gene (ExG) locus in the genome of the cell; and (iii) a donor polynucleotide comprising: (1) a functional ExG gene containing a mutation in the native coding sequence of the ExG, the mutation conferring resistance to inactivation by the guide polynucleotide; (2) the SOI; and (3) a sequence for genomic integration at the target locus, wherein introducing (i), (ii), and (iii) results in inactivation of the ExG in the genome of the cell by the nuclease and integration of the donor polynucleotide at the target locus; (b) culturing the cells; and (c) selecting one or more viable cells, wherein the one or more viable cells comprise the SOI integrated at the target locus.

[0219] FIG. 13 shows an embodiment of the method of the present disclosure. In FIG. 13, a CRISPR-Cas complex is introduced into cells that target ExG, an essential gene for cell survival. A vector containing a target gene (GOI) and a modified ExG* that is resistant to targeting by the CRISPR-Cas complex is also introduced into the cells. As a result, cells that cleave ExG (indicated by an asterisk in the ExG sequence) and vectors that have successfully introduced ExG* can survive, and cells that do not have this vector die as a result of lacking ExG. The guide RNA of the CRISPR-Cas complex can be designed and selected such that it is nearly 100% efficient against ExG in the cell genome and / or multiple guide RNAs can be used to target the same ExG. Alternatively, or additionally, multiple rounds of selecting surviving cells and introducing the CRISPR-Cas complex may be performed such that the surviving cells have an increased tendency to lack genomic copies of ExG and survive due to the presence of ExG* (and thus GOI). Thus, the surviving cells are enriched for having GOI.

[0220] In some embodiments, essential genes are genes necessary for an organism to survive. In some embodiments, disruption or deletion of an essential gene results in cell death. In some embodiments, essential genes are auxotrophic genes, i.e., genes that produce specific compounds necessary for growth or survival. Examples of auxotrophic genes include genes involved in nucleotide biosynthesis such as adenine, cytosine, guanine, thymine, or uracil, or in amino acid biosynthesis such as histidine, leucine, lysine, methionine, or tryptophan. In some embodiments, essential genes are genes in metabolic pathways. In some embodiments, essential genes are genes in the autophagy pathway. In some embodiments, essential genes are genes in cell division, such as in mitosis, cytoskeletal organization, or response to stress or stimuli. In some embodiments, essential genes encode proteins that promote cell growth or division, receptors for signaling molecules (e.g., molecules near the cell), or proteins that interact with another protein, organelle, or biomolecule. Exemplary essential genes include, but are not limited to, the genes listed in FIG. 23. Further examples of essential genes are provided, for example, in Hart et al., Cell 163:1515-1526 (2015), Zhang et al., Microb Cell 2(8):280-287 (2015), and Fraser, Cell Systems 1:381-382 (2015).

[0221] In some embodiments, the nuclease capable of generating a double-strand break is Cas9. In some embodiments, the Cas9 protein generates a site-specific cleavage in a nucleic acid. In some embodiments, the Cas9 protein generates a site-specific double-strand break in DNA. The ability of Cas9 to target a specific sequence in a nucleic acid (i.e., site specificity) is achieved by Cas9 complexed with a guide polynucleotide (e.g., guide RNA) that hybridizes to a designated sequence (e.g., ExG locus). In some embodiments, Cas9 is a Cas9 variant described in U.S. Patent Application No. 62 / 728,184, filed September 7, 2018.

[0222] In some embodiments, Cas9 is capable of generating sticky ends. Cas9 capable of generating sticky ends is described, for example, in the pamphlet of International Publication No. WO 2018 / 061680, filed on November 16, 2018. In some embodiments, Cas9 capable of generating sticky ends is a dimeric Cas9 fusion protein. The binding domains and cleavage domains of natural nucleases (such as Cas9, etc.), as well as modular binding domains and cleavage domains that can be fused to generate nuclease-binding specific target sites, are well-known to those skilled in the art. For example, the binding domain of an RNA programmable nuclease (such as Cas9), or a Cas9 protein having an inactive DNA cleavage domain, can be used as a binding domain (for example, binding to a gRNA to direct binding to a target site) to specifically bind to a desired target site, and fused or conjugated to a cleavage domain, such as the cleavage domain of endonuclease FokI, to generate a modified nuclease that cleaves the target region. The Cas9-FokI fusion protein is further described, for example, in U.S. Patent Application Publication No. 2015 / 0071899, and Guilinger et al., "Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification" Nature Biotechnology 32:577-582 (2014).

[0223] In some embodiments, Cas9 comprises the polypeptide sequence of SEQ ID NO: 3 or 4. In some embodiments, Cas9 has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence identity with SEQ ID NO: 3 or 4. In some embodiments, Cas9 is SEQ ID NO: 3 or 4.

[0224] In some embodiments, the guide polynucleotide is an RNA polynucleotide. RNA molecules that bind to CRISPR-Cas components and target them to specific positions within the target DNA are referred to herein as "RNA guide polynucleotides", "guide RNAs", "gRNAs", "small guide RNAs", "single guide RNAs", or "sgRNAs", and may also be referred to herein as "DNA-targeting RNAs". The guide polynucleotide may be introduced into the target cell as an isolated molecule, e.g., an RNA molecule, or introduced into the cell using an expression vector containing DNA encoding the guide polynucleotide, e.g., an RNA guide polynucleotide. In some embodiments, the guide polynucleotide is 10 to 150 nucleotides. In some embodiments, the guide polynucleotide is 20 to 120 nucleotides. In some embodiments, the guide polynucleotide is 30 to 100 nucleotides. In some embodiments, the guide polynucleotide is 40 to 80 nucleotides. In some embodiments, the guide polynucleotide is 50 to 60 nucleotides. In some embodiments, the guide polynucleotide is 10 to 35 nucleotides. In some embodiments, the guide polynucleotide is 15 to 30 nucleotides. In some embodiments, the guide polynucleotide is 20 to 25 nucleotides.

[0225] In some embodiments, the RNA-guided polynucleotide comprises at least two nucleotide segments: at least one "DNA-binding segment" and at least one "polypeptide-binding segment". A "segment" means a part, section, or region of a molecule, e.g., a continuous stretch of nucleotides of a guide polynucleotide molecule. The definition of "segment" is not limited to a specific number of total base pairs unless specifically defined otherwise.

[0226] In some embodiments, the guide polynucleotide comprises a DNA-binding segment. In some embodiments, the DNA-binding segment of the guide polynucleotide comprises a nucleotide sequence complementary to a specific sequence within the target polynucleotide. In some embodiments, the DNA-binding segment of the guide polynucleotide hybridizes to an essential gene locus (ExG) in a cell. Various types of cells, such as eukaryotic cells, are described herein.

[0227] In some embodiments, the guide polynucleotide comprises a polypeptide-binding segment. In some embodiments, the polypeptide-binding segment of the guide polynucleotide binds to the DNA targeting domain of a nuclease of the present disclosure. In some embodiments, the polypeptide-binding segment of the guide polynucleotide binds to Cas9. In some embodiments, the polypeptide-binding segment of the guide polynucleotide binds to dCas9. In some embodiments, the polypeptide-binding segment of the guide polynucleotide binds to nCas9. Various RNA-guided polynucleotides that bind to the Cas9 protein are described, for example, in U.S. Patent Application Publication Nos. 2014 / 0068797, 2014 / 0273037, 2014 / 0273226, 2014 / 0295556, 2014 / 0295557, 2014 / 0349405, 2015 / 0045546, 2015 / 0071898, 2015 / 0071899, and 2015 / 0071906.

[0228] In some embodiments, the guide polynucleotide further comprises a tracrRNA. A "tracrRNA" or trans-activating CRISPR-RNA forms an RNA duplex with a pre-crRNA, or pre-CRISPR-RNA, which is then cleaved by the RNA-specific ribonuclease RNase III to form a crRNA / tracrRNA hybrid. In some embodiments, the guide polynucleotide comprises a crRNA / tracrRNA hybrid. In some embodiments, the tracrRNA component of the guide polynucleotide activates the Cas9 protein. In some embodiments, activation of the Cas9 protein comprises activating the nuclease activity of Cas9. In some embodiments, activation of the Cas9 protein comprises a Cas9 protein that binds to a target polynucleotide sequence, such as the ExG locus.

[0229] In some embodiments, the guide polynucleotide guides a nuclease to the ExG locus, and the nuclease generates a double-strand break at the ExG locus. In some embodiments, the guide polynucleotide is a guide RNA. In some embodiments, the nuclease is Cas9. In some embodiments, the double-strand break at the ExG locus inactivates ExG. In some embodiments, inactivation of the ExG locus disrupts an essential cellular function. In some embodiments, inactivation of the ExG locus prevents cell division. In some embodiments, inactivation of the ExG locus results in cell death.

[0230] In some embodiments, an "exogenous" ExG or a portion thereof can be introduced into a cell to compensate for an inactivated native ExG. In some embodiments, the exogenous ExG is a functional ExG. The term "functional" ExG refers to an ExG that encodes a polypeptide that is substantially similar to the polypeptide encoded by the native coding sequence. In some embodiments, the functional ExG comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence similarity to the native coding sequence of the ExG, and further comprises a mutation in the native coding sequence of the ExG that confers resistance to inactivation by a nuclease. In some embodiments, the functional ExG is resistant to inactivation by a nuclease, and the polypeptide encoded by the functional ExG has substantially the same structure as the polypeptide encoded by the native coding sequence and performs the same cellular function.

[0231] In some embodiments, a portion of the ExG encodes a polypeptide that performs substantially the same function as the native protein encoded by the ExG. In some embodiments, a portion of the ExG is introduced to complement a partially inactivated ExG. In some embodiments, the nuclease inactivates a portion of the native ExG (e.g., by disrupting a portion of the coding sequence of the ExG), and the exogenous ExG comprises the disrupted portion of the coding sequence that can be transcribed with the non-disrupted portion of the native sequence to form a functional ExG. In some embodiments, the exogenous ExG or a portion thereof is integrated into the native ExG locus in the genome of the cell. In some embodiments, the exogenous ExG or a portion thereof is integrated into a genomic locus different from the ExG locus.

[0232] In some embodiments, the functional ExG does not bind to the nuclease. In some embodiments, an ExG that does not bind to the nuclease is resistant to cleavage by the nuclease. As described herein, nucleases such as certain types of Cas9 may require a PAM sequence in or near the target sequence in addition to recognition of the target sequence by a guide polynucleotide (e.g., guide RNA) via hybridization. In some embodiments, Cas9 binds to the PAM sequence before initiating nuclease activity. In some embodiments, a target sequence, or a target sequence that does not contain a PAM in the adjacent or neighboring region, does not bind to the nuclease. Thus, in some embodiments, a target sequence, or a target sequence that does not contain a PAM in the adjacent or neighboring region, is not cleaved by the nuclease and is therefore resistant to nuclease-mediated inactivation. In some embodiments, mutations in the native coding sequence of ExG remove the PAM sequence. In some embodiments, an ExG that does not contain a PAM sequence is resistant to nuclease-mediated inactivation.

[0233] In some embodiments, the PAM is within about 30 to about 1 nucleotide of the target sequence. In some embodiments, the PAM is within about 20 to about 2 nucleotides of the target sequence. In some embodiments, the PAM is within about 10 to about 3 nucleotides of the target sequence. In some embodiments, the PAM is within about 10, about 9, about 8, about 7, about 6, about 5, about 4, about 3, about 2, or about 1 nucleotide of the target sequence. In some embodiments, the PAM is present upstream (i.e., in the 5' direction) of the target sequence. In some embodiments, the PAM is present downstream (i.e., in the 3' direction) of the target sequence. In some embodiments, the PAM is located within the target sequence.

[0234] In some embodiments, the polypeptide encoded by the functional ExG cannot hybridize with the guide polynucleotide. In some embodiments, an ExG that does not hybridize with the guide polynucleotide is resistant to cleavage by a nuclease such as Cas9. As described herein, the guide polynucleotide can hybridize with the target sequence, i.e., the target sequence can be "recognized" by the guide polynucleotide for cleavage by a nuclease such as Cas9. Thus, a sequence that does not hybridize with the guide polynucleotide is not recognized for cleavage by a nuclease such as Cas9. In some embodiments, a sequence that does not hybridize with the guide polynucleotide is resistant to nuclease-mediated inactivation. In some embodiments, the guide polynucleotide can hybridize with the ExG in the genome of the cell, and the functional ExG on the donor polynucleotide or episomal vector contains a mutation in the native coding sequence of the ExG, such that the guide polynucleotide can (1) hybridize with the ExG in the genome of the cell and (2) not hybridize with the functional ExG on the donor polynucleotide or episomal vector. In some embodiments, a functional ExG that is resistant to nuclease-mediated inactivation is introduced into the cell simultaneously with a nuclease that targets the ExG in the genome of the cell.

[0235] In some embodiments, the functional ExG contains one or more mutations relative to the wild-type sequence, but the polypeptide encoded by the native coding sequence is substantially similar to the polypeptide encoded by the wild-type sequence. For example, the amino acid sequence of the polypeptide is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identical. In some embodiments, the polypeptides encoded by the functional ExG and the wild-type ExG have similar structures, such as similar overall shapes and folds, as determined by one of ordinary skill in the art. In some embodiments, the functional ExG contains a portion of the wild-type sequence. In some embodiments, the functional ExG contains a mutation relative to the wild-type sequence. In some embodiments, the functional ExG contains a mutation in the native coding sequence of the ExG, and the mutation confers resistance to inactivation by a nuclease.

[0236] In some embodiments, the mutations in the native coding sequence of ExG are substitution mutations, insertions, or deletions. In some embodiments, a substitution mutation is a substitution of one or more nucleotides in a polynucleotide sequence, but the encoded amino acid sequence remains unchanged. In some embodiments, a substitution mutation replaces one or more nucleotides to change the codon for an amino acid to a degenerate codon for the same amino acid. For example, the native coding sequence may include the sequence "CAT" that encodes histidine, and the mutation may change this sequence to "CAC", which also encodes histidine. In some embodiments, a substitution mutation replaces one or more nucleotides to change an amino acid to a different amino acid, but this different amino acid has similar properties so that the overall structure of the encoded polypeptide or the overall function of the protein is not affected. For example, a substitution mutation may result in a change from leucine to isoleucine, from glutamine to asparagine, from glutamate to aspartate, from serine to threonine, and so on.

[0237] In some embodiments, an exogenous ExG or a portion thereof (e.g., an ExG that includes a mutation in the native coding sequence of ExG, where the mutation confers resistance to nuclease inactivation) is introduced into cells in an exogenous polynucleotide. In some embodiments, the exogenous ExG is expressed from the exogenous polynucleotide. In some embodiments, the exogenous polynucleotide is a plasmid. In some embodiments, the exogenous polynucleotide is a donor polynucleotide. In some embodiments, the donor polynucleotide is a vector. Exemplary vectors are provided herein.

[0238] In some embodiments, the exogenous ExG or a portion thereof on the donor polynucleotide is integrated into the genome of the cell by sequences for genome integration. In some embodiments, the sequences for genome integration are obtained from a retroviral vector. In some embodiments, the sequences for genome integration are obtained from a transposon.

[0239] In some embodiments, the donor polynucleotide comprises sequences for genomic integration. In some embodiments, the sequences for genomic integration at the target locus are obtained from a transposon. As described herein, a transposon comprises transposon sequences that are recognized by a transposase, which subsequently inserts a transposon comprising the transposon sequences and the sequence of interest (SOI) into the genome. In some embodiments, the target locus is any genomic locus capable of expressing the SOI without disrupting normal cellular function. Exemplary transposons are described herein. Thus, in some embodiments, the donor polynucleotide comprises a functional ExG comprising a mutation in the native coding sequence of the ExG, wherein the mutation confers resistance to nuclease inactivation, the SOI, and transposon sequences for genomic integration at the target locus. In some embodiments, the native ExG of the cell is inactivated by a nuclease, and the donor polynucleotide provides a functional ExG that is resistant to nuclease inactivation and capable of compensating for the native cellular function of the native ExG.

[0240] In some embodiments, the donor polynucleotide comprises sequences for genomic integration. In some embodiments, the sequences for genomic integration at the target locus are obtained from a retroviral vector. As described herein, a retroviral vector contains sequences recognized by integrase, typically LTRs, and subsequently integrase inserts the retroviral vector containing the LTR and the SOI into the genome. In some embodiments, the target locus is any genomic locus capable of expressing the SOI without disrupting normal cellular function. Exemplary retroviral vectors are described herein. Thus, in some embodiments, the donor polynucleotide comprises a functional ExG containing a mutation in the native coding sequence of the ExG, wherein the mutation confers resistance to nuclease inactivation, the SOI, and a retroviral vector for genomic integration at the target locus. In some embodiments, the native ExG of the cell is inactivated by a nuclease, and the donor polynucleotide provides a functional ExG that can compensate for the native cellular function of the native ExG while being resistant to nuclease inactivation.

[0241] In some embodiments, the exogenous polynucleotide is an episomal vector. In some embodiments, the episomal vector is a stable episomal vector, i.e., an episomal vector that persists in the cell. As described herein, an episomal vector contains an autonomous DNA replication sequence that allows the episomal vector to replicate and persist in the cell. In some embodiments, the episomal vector is an artificial chromosome. In some embodiments, the episomal vector is a plasmid.

[0242] In some embodiments, an episomal vector is introduced into a cell. In some embodiments, the episomal vector comprises a functional ExG that contains a mutation in the native coding sequence of ExG, the mutation conferring resistance to nuclease inactivation, an SOI, and an autonomous DNA replication sequence. As described herein, an episomal vector is a non-integrating extrachromosomal plasmid capable of autonomous replication. In some embodiments, the autonomous DNA replication sequence is derived from a viral genomic sequence. In some embodiments, the autonomous DNA replication sequence is derived from a mammalian genomic sequence. In some embodiments, the episomal vector is an artificial chromosome or a plasmid. In some embodiments, the plasmid is a viral plasmid. In some embodiments, the viral plasmid is an SV40 vector, a BKV vector, a KSHV vector, or an EBV vector. Thus, in some embodiments, the native ExG of the cell is inactivated by a nuclease, and the episomal vector provides a functional ExG that is resistant to nuclease inactivation yet capable of compensating for the native cellular functions of the native ExG.

[0243] In some embodiments, the SOI comprises a polynucleotide encoding a protein. In some embodiments, the SOI comprises a mutant gene. In some embodiments, the SOI comprises a non-coding sequence, such as a microRNA. In some embodiments, the SOI is operably linked to a regulatory element. In some embodiments, the SOI is a regulatory element. In some embodiments, the SOI comprises a resistance cassette, such as a gene conferring resistance to an antibiotic. In some embodiments, the SOI comprises a marker, such as a selectable marker or a selectable marker. In some embodiments, the SOI comprises a marker, such as a restriction site, a fluorescent protein, or a selectable marker.

[0244] In some embodiments, the SOI contains a mutation of the wild-type gene in the genome of the cell. In some embodiments, the mutation includes a point mutation, i.e., a substitution of a single nucleotide. In some embodiments, the mutation includes a substitution of multiple nucleotides. In some embodiments, the mutation introduces a stop codon. In some embodiments, the mutation includes an insertion of a nucleotide in the wild-type sequence. In some embodiments, the mutation includes a deletion of a nucleotide in the wild-type sequence. In some embodiments, the mutation includes a frameshift mutation.

[0245] In some embodiments, the guide polynucleotide has a targeting efficiency of more than 80%, more than 85%, more than 90%, more than 95%, or about 100% for ExG in the genome of the cell. The targeting efficiency can be measured, for example, by the proportion of cells having inactivated ExG in the cell population. The guide polynucleotide can be designed and selected to increase efficiency using design tools such as, for example, Chop Chop (chopchop.cbu.uib.no), CasFinder (arep.med.harvard.edu / CasFinder), E-CRISP (e-crisp.org / E-CRISP / designcrispr.html), CRISPR-ERA (crispr-era.stanford.edu / index.jsp).

[0246] In some embodiments, two or more guide polynucleotides are introduced into a cell population, each guide polynucleotide forms a complex with a nuclease, and each guide polynucleotide hybridizes to a different region of the ExG. In some embodiments, multiple guide polynucleotides can be used to increase the efficiency of inactivating ExG in the genome of a cell. For example, a first guide polynucleotide can target the 5' region of the ExG, a second guide polynucleotide can target the internal region of the ExG, and a third guide polynucleotide can target the 3' region of the ExG. The targeting efficiency of each guide polynucleotide can vary, but nuclease cleavage in any of the 5', 3', or internal regions inactivates the ExG, so the overall efficiency can be increased by using two or more guide polynucleotides that target the same gene. In some embodiments, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, or at least 20 different guide polynucleotides are introduced into the cell population.

[0247] In some embodiments, viable cells include a mixture of cells containing ExG* and SOI integrated at a target locus or on an episomal vector and cells containing ExG that has not been inactivated by a nuclease, for example, due to inherent inefficiencies in the nuclease or failure to introduce the nuclease and / or guide polynucleotide into the cell. Thus, in some embodiments, one or more steps of the method are repeated to enrich for viable cells containing the desired SOI. Repeated introduction of the nuclease and guide polynucleotide can increase the likelihood that ExG in the genome of the cell is inactivated, thereby enriching for viable cells containing ExG* and SOI integrated at a target locus or on an episomal vector.

[0248] Thus, in embodiments of a method for integrating an SOI into a target locus, the method comprises introducing into one or more selected viable cells a nuclease capable of generating a double-strand break and a guide polynucleotide capable of forming a complex with and hybridizing to an ExG in the genome of the cell, to enrich for viable cells comprising the SOI integrated into the target locus. In embodiments of a method for introducing a stable episomal vector into a cell, the method comprises introducing into one or more selected viable cells a nuclease capable of generating a double-strand break and a guide polynucleotide capable of forming a complex with and hybridizing to an ExG in the genome of the cell, to enrich for viable cells comprising the episomal vector.

[0249] In some embodiments, the nuclease and the guide polynucleotide are introduced into the viable cells for multiple rounds of enrichment. In some embodiments, the nuclease and the guide polynucleotide are introduced for 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more than 20 rounds of enrichment. Each round of targeting increases the likelihood that the viable cells contain the SOI. That is, this enriches for viable cells comprising the SOI integrated into the target locus or episomal vector.

[0250] Sequence Sequences of various polynucleotides and polypeptides are provided herein.

[0251] Polynucleotide sequence of the Cas9 protein (SpCas9; SEQ ID NO: 1) from Streptococcus pyogenes:

Chemical formula

Chemical formula

[0252] Polynucleotide sequence of the Cas9 protein (FnCas9; SEQ ID NO: 2) derived from Francisella novicida:

Chem.

Chem.

Chem.

[0253] Polypeptide sequence of SpCas9 (SEQ ID NO: 3):

Chem.

[0254] Polypeptide sequence of FnCas9 (SEQ ID NO: 4):

Chem.

[0255] Polynucleotide sequence of BE3 (SEQ ID NO: 5):

Chem.

Chem.

Chem.

[0256] Polypeptide sequence of BE3 (SEQ ID NO: 6):

Chem.

[0257] Polynucleotide sequence of the HB-EGF locus (SEQ ID NO: 7):

Chem.

[0258] Polypeptide sequence of HB-EGF protein (SEQ ID NO: 8): [Chemistry]

[0259] All references cited in this specification, such as patents, patent applications, papers, textbooks, etc., and the references cited therein, are hereby incorporated by reference in their entirety into this specification to the extent that they are not already cited. [Examples]

[0260] Example 1. Experimental protocol In this example, a protocol for co-target enrichment is provided.

[0261] Maintain the expression of heparin-binding EGF receptor every 2 - 3 days during culture and subculture until the cell line is transfected. The cells should be >80% confluent on the day of transfection.

[0262] Transfect the cells with a plasmid encoding a base editor or Cas9, and / or a plasmid encoding a guide RNA targeting HB-EGF and a plasmid encoding the gene of interest. Prepare the DNA-lipid complex for transfection according to the manufacturer's protocol. Alternatively, mRNA and RNP complexes may be used.

[0263] Add the complex to plates containing freshly trypsinized cells seeded the day before.

[0264] Remove the medium 72 hours after transfection, trypsinize the cells, and then re-seed them onto new plates with twice the surface area of the previous plates.

[0265] The next day, add diphtheria toxin at a concentration of 20 ng / mL to the wells. Two days later, perform a new diphtheria toxin treatment.

[0266] Monitor cell growth and, if necessary, transfer the cells to larger plates or flasks until all cells undergoing negative selection are dead.

[0267] After 1 - 2 weeks, analyze the cells by next-generation sequencing to determine the editing efficiency.

[0268] Example 2. Screening of guide RNAs In this example, guide RNAs (gRNAs) were screened to identify gRNAs that confer diphtheria toxin resistance when co-transfected with BE3. A panel of gRNAs was designed to tile the EGF-like domain of HB-EGF (see Figure 4C). Each gRNA was co-transfected with BE3 into HEK293 or HCT116 cells at a transfection weight ratio of 1:4.

[0269] The cells were treated with 20 ng / mL of diphtheria toxin on the third day after transfection and again on the fifth day after transfection. Cell growth was measured by confluence using INCUCYTE ZOOM.

[0270] The results shown in Figures 4A and 4B respectively indicate that HEK293 and HCT116 cells co-transfected with HB-EGF gRNA16 and BE3 had the highest growth levels among all transfected cells. The results of Sanger sequencing and next-generation sequencing analysis shown in Figures 5B - 5D revealed that the diphtheria toxin resistance in gRNA16-transfected cells was a result of the E141K mutation introduced by BE3 base editing. The sequence of gRNA16 is shown in Figure 5A.

[0271] Example 3. Co-targeted enrichment by BE3 and Cas9 In this example, to generate diphtheria toxin-resistant cells, co-targeted enrichment using diphtheria toxin selection was tested using BE3 and Cas9 by co-transfection of the targeted gRNA and gRNA16 identified in Example 2.

[0272] Plasmid construction Cas9 plasmid: The DNA sequence encoding SpCas9, T2A self-cleaving peptide, and puromycin N-acetyltransferase was synthesized by GeneArt and cloned into an expression vector containing the CMV promoter and BGH polyA tail. See the plasmid map in Figure 15.

[0273] BE3 plasmid. The DNA sequence of base editor 3 was synthesized by GeneArt using the restriction sites BamHI and XhoI and cloned into pcDNA3.1(+). See the plasmid map in Figure 14.

[0274] gRNA plasmid. The target sequence of gRNA was introduced into the AarI cleavage site of the template plasmid using complementary primer pairs (5’-AAAC-N20-3’ and 5’-ACCG-N20-3’). The template plasmid was synthesized by GeneArt. It contains a U6 promoter driving the gRNA expression cassette, and the rpsL-BSD selection cassette was cloned in the region of the gRNA target sequence flanked by two AarI restriction sites. The primers can be seen in Table 1. The plasmids for gRNAs targeting BFP and EGFR are described in Coelho et al., BMC Biology 16:150 (2018) and shown in Figures 17 - 23.

[0275]

Table 1

[0276]

Table 2

[0277] Cell culture and transfection HEK293T and HCT116 cells obtained from ATCC were maintained in Dulbecco's modified Eagle's medium (DMEM) supplemented with 10% fetal bovine serum (FBS). PC9 - BFP cells were maintained in DMEM medium containing 10% FBS.

[0278] Transfection was performed using FUGENE HD transfection reagent (Promega) with a 3:1 transfection reagent to DNA ratio according to the manufacturer's instructions. Transfection in this study was performed in 24 - well plates and 48 - well plates. 24 hours before transfection, 1.25×10 5 and 6.75×10 4Individual cells were seeded into 24-well plates and 48-well plates respectively. Transfection was performed using a total of 500 ng and 250 ng of DNA for 24-well plates and 48-well plates respectively.

[0279] For co-target enrichment, Cas9 or BE3 plasmid DNA, targeting gRNA plasmid DNA, and selection gRNA plasmid DNA were transfected at a weight ratio of 8:1:1. The sequences of the targeting gRNAs for the PCKS9 site are shown in Figure 7C, and the sequences of the targeting gRNAs for the DPM2, EGFR, EMX1, and Yas85 sites are shown in Figure 7E. Cells were treated with 20 ng / ml of diphtheria toxin 3 days after transfection and then again 5 days after transfection. Cells were harvested for downstream applications when they reached >80% confluence. For all cell types used in this study, cells were harvested 7 days after transfection for genomic extraction. For other different cell lines or primary cells, different diphtheria toxin doses and treatment times may be applied to kill all wild-type cells.

[0280] Next-generation sequencing and data analysis Genomic DNA was extracted from cells 72 hours after transfection or treatment using QUICKEXTRACT DNA Extraction Solution (Lucigen) according to the instructions. NGS libraries were prepared by two-step PCR. The first PCR was performed using NEBNEXT Q5 Hot Start HiFi PCR Master Mix (New England Biolabs) according to the instructions. The second PCR was performed using 1 ng of the product from the first PCR with the KAPA HiFi PCR Kit (KAPABIOSYSTEMS). PCR products were purified using Agencourt AMPure XP (Beckman Coulter) and analyzed by Fragment analyzer.

[0281] The results in FIGS. 7A and 7B show the BE3 base editing efficiencies of different cytosines at the PCSK9 target sites in HCT116 and HEK293 cells, respectively. The "control" condition shows a relatively low base editing efficiency without diphtheria toxin selection, and the "enrichment" condition shows a significantly higher base editing efficiency when diphtheria toxin selection is used. The results in FIG. 7D show an increase in base editing efficiency at different cytosines in the DPM2, EGFR, EMX1, and Yas85 target sites when diphtheria toxin selection is used ("enrichment") compared to the "control" condition without diphtheria toxin.

[0282] The results in FIG. 8A show the Cas9 editing efficiency by measuring the proportion of indels generated at the PCSK9 target site in HEK293 and HCT116 cells. Similar to the case of base editing, the Cas9 editing efficiency was significantly increased above the "control" condition without diphtheria toxin selection under the "enrichment" condition using diphtheria toxin selection. The results in FIG. 8B show a similar increase in Cas9 editing efficiency at the DPM2, EXM1, and Yas85 target sites.

[0283] Example 4.2 Allele Integration In this example, diphtheria toxin selection was tested to improve the knock-in (insertion) efficiency by which the gene of interest achieves bi-allelic integration.

[0284] Donor plasmid for knock-in. A knock-in plasmid for mCherry was synthesized by Genescript. See FIG. 23 for the plasmid map. Also see FIG. 10A for the experimental design.

[0285] For the knock-in experiment, transfection was performed in a 24-well plate format. Cas9 plasmid DNA, gRNA plasmid DNA, and mCherry knock-in (KI) or control plasmid DNA were transfected at different weight ratios under different conditions shown in Table 2. Cells were treated with 20 ng / ml of diphtheria toxin 3 days after transfection and then again 5 days after transfection. Subsequently, the cells were maintained in fresh medium without diphtheria toxin. Genomes of all samples were harvested for PCR analysis 13 days after transfection. 22 days after transfection, cells from transfection condition 3, transfection negative controls 1 and 2, and the mCherry positive control cell line were resuspended and analyzed by FACS.

[0286]

Table 3

[0287] Cells in which the insertion was successful translate mCherry containing the mutant HB-EGF gene, and the cells show mCherry fluorescence. As shown in Figure 10B, almost all cells transfected with Cas9, gRNA SaW10, and the mCherry HDR template were mCherry positive after diphtheria toxin selection, and cells without the mCherry donor plasmid showed no mCherry fluorescence at all. Figure 10C shows that the expression of mCherry is uniform throughout the population (Figure 10C).

[0288] Figures 10E and 10F show the results of PCR analysis using the strategy outlined in Figure 10D. The first PCR reaction (PCR1) amplifies the binding region, with the forward primer (PCR1_F primer) binding to a sequence in the genome and the reverse primer (PCR1_R primer) binding to a sequence in the GOI. Thus, only cells in which the GOI has been integrated will show a positive band in PCR1. The second PCR reaction (PCR2) amplifies the insertion region, with the forward primer (PCR2_F primer) binding to a sequence in the 5' end of the insertion and the reverse primer (PCR2_R primer) binding to a sequence in the 3' end of the insertion. Thus, amplification by PCR2 occurs only if all alleles in the cell have successfully inserted the GOI, and the amplification product is shown as a single integration band. If any wild-type alleles are present, a WT band will be shown.

[0289] Figure 10E shows positive bands for all tested conditions including the introduction of Cas9, gRNA, and mCherry donor plasmid, indicating that the insertion was successfully achieved. The single integration band for all three conditions in Figure 10F indicates that no wild-type alleles were present in the tested cells, i.e., a two-allele integration was achieved.

[0290] Example 5. Detailed Experimental Protocol An experimental protocol for subsequent examples is provided.

[0291] Construction of Plasmids and Template DNA A plasmid expressing S. pyogenes Cas9 (SpCas9) was constructed by cloning a codon-optimized SpCas9 sequence fused to a nuclear localization signal (NLS) and a self-cleaving puromycin resistance protein (T2A-Puro), which was synthesized by GeneArt, into the pVAX1 vector. Two versions of the SpCas9 plasmid were constructed to drive the expression of SpCas9 under the control of the CMV promoter (CMV-SpCas9) or the EF1α promoter (EF1α-SpCas9). The cytidine base editor 3 (CBE3) was synthesized by GeneArt using its published sequence and cloned into the pcDNA3.1(+) vector. Two versions of the plasmid were constructed to control the expression of CBE3 under the CMV promoter (CMV-CBE3) or the EF1α promoter (EF1α-CBE3). Similarly, the adenine base editor 7.10 (ABE7.10) was synthesized using its published sequence and cloned into the pcDNA3.1(+) vector. Two versions of the plasmid were constructed to control the expression of ABE7.10 under the CMV promoter (CMV-ABE7.10) or the EF1α promoter (EF1α-ABE7.10). Individual sequence components were obtained from Integrated DNA Technolgies (IDT) and assembled using Gibson assembly (New England Biolabs).

[0292] Plasmids expressing different sgRNAs were cloned by replacing the target sequence of the template plasmid. Complementary primer pairs containing the target sequence (5’-AAAC-N20-3’ and 5’-ACCG-N20-3’) were annealed (5 minutes at 95 °C. Subsequently cooled to 25 °C at 1 °C / min) and assembled with the AarI-digested template using T4 ligase. All primer pairs are listed in Table 3A. Plasmids expressing sgRNAs targeting BFP, as well as plasmids expressing sgRNAs targeting EGFR and CBE3, have been described in previous publications.

[0293] A plasmid that acts as a repair template for the HBEGF or HIST2BC locus was obtained from GenScript or modified using Gibson assembly. Individual sequence components were obtained from IDT. The template plasmid for the HBEGF locus was designed to contain a strong splicing acceptor sequence encoded by the polyA sequence, followed by a mutant CDS sequence of HBEGF starting from exon 4, and a self-cleaving mCherry coding sequence. The template plasmid for HIST2BC was designed to contain a GFP coding sequence, followed by a self-cleaving blasticidin resistance protein coding sequence. For both loci, pHMEJ and pHR were designed to contain left and right homology arms flanking the inserted sequence, and pNHEJ was designed not to contain homology arms. pHMEJ was designed to contain one sgRNA cleavage site flanking each homology arm, while pHR did not contain this site. To compare puromycin selection by DT selection, a self-cleaving puromycin resistance protein coding sequence was inserted between the HBEGF exon sequence and the self-cleaving mCherry coding sequence (pHMEJ_PuroR).

[0294] The double-stranded DNA (dsDNA) template was prepared by PCR amplification of plasmid pHMEJ using the primers listed in Table 3B, followed by purification using MAGBIO magnetic SPRI beads. The PCR amplification was performed using high-fidelity PHUSION polymerase. The ssDNA template was prepared using the GUIDE-IT™ Long ssDNA Production System (Takara Bio Inc.) with the primers listed in Tables 3A - 3E. The final product was purified using MAGBIO magnetic SPRI beads and analyzed using a Fragment Analyzer (Agilent). The template for the CD34 locus was obtained from IDT as a PAGE-purified oligonucleotide.

[0295]

Table 4

[0296]

Table 5

[0297]

Table 6

[0298]

Table 7

[0299]

Table 8

[0300]

Table 9

[0301]

Table 10

[0302]

Table 11

[0303]

Table 12

[0304] Cell culture HEK293 (ATCC, CRL-1573), HCT116 (ATCC, CCL-247), and PC9-BFP cells were maintained in Dulbecco's modified Eagle's medium (DMEM) supplemented with 10% fetal bovine serum. Human induced pluripotent stem cells (iPSCs) were maintained in Cellartis DEF-CS 500 Culture System (Takara Bio Inc.) according to the manufacturer's instructions. All cell lines were cultured at 37 °C, 5% CO 2 2. Cells were authenticated by STR profiling and were negative for mycoplasma testing.

[0305] Isolation, activation, and expansion of T cells Blood from healthy donors was obtained from AstraZeneca’s blood donation center (Mölndal, Sweden). Peripheral blood mononuclear cells were isolated from fresh blood using Lympoprep (STEMCELL Technologies) density gradient centrifugation, and total CD4+ T cells were enriched by negative selection using the EasySep Human CD4+ T Cell Enrichment Kit (STEMCELL Technologies). Enriched CD4+ T cells were further purified to a mean purity of 98% by fluorescence-activated cell sorting (FACSAria III, BD Biosciences) based on the exclusion of CD8+CD14+CD16+CD19+CD25+ cell surface markers. The following antibodies were purchased from BD Biosciences: CD4-PECF594 (RPA-T4), CD25-PECy7 (M-A251), CD8-APC-Cy7 (RPA-T8), CD14-APC-Cy7 (MφP-9), CD16-APC-Cy7 (3G8), CD19-APC-Cy7 (SJ25-C1), CD45RO-BV510. (UCHL1). Cell sorting was performed using a FACSAria III (BD Biosciences).

[0306] CD4+ T cells were propagated in RPMI-1640 medium containing the following supplements: 1% (v / v) GlutaMAX-I, 1% (v / v) non-essential amino acids, 1 mM sodium pyruvate, 1% (v / v) L-glutamine, 50 U / mL penicillin and streptomycin, and 10% heat-inactivated FBS (all from Gibco, life Technologies). T cells were activated using the T Cell Activation / Expansion kit (130-091-441, Miltenyi). 1×10 6 cells / mL were activated at a bead-to-cell ratio of 1:2, and 2×10 5 cells were seeded per well into round-bottom tissue culture-treated 96-well plates for 24 hours. Cells were pooled prior to electroporation.

[0307] Cell Transfection Twenty-four hours prior to transfection, 1.25×10 5 or 6.75×10 4 HEK293, HCT116, and PC9-BFP cells were seeded into 24-well or 48-well plates, respectively. Transfection was performed using FuGENE HD Transfection Reagent (Promega) at a transfection reagent-to-plasmid DNA ratio of 3:1. For the 24-well plate format, the amount and weight ratio of transfected DNA are listed in Tables 4 and 5. For the 48-well plate format, the amount of DNA was reduced by half.

[0308] [Table 13]

[0309] [Table 14]

[0310] iPSCs were transfected with FuGENE HD using a 2.5:1 transfection reagent to DNA ratio and a reverse transfection protocol. For transfection, 4.2×10 4 cells per well in a 48-well format were seeded directly onto the prepared transfection complexes described in Table 6.

[0311]

Table 15

[0312] CD4+ T cells were electroporated with ribonucleoprotein complexes (RNPs) using 10 μL of the Neon transfection kit (MPK1096, ThermoFisher). CD3 protein was generated using the method described above. An additional purification step was performed on a HiLoad 26 / 600 Superdex 200 pg column (GE Healthcare) with a mobile phase containing 20 mM Tris-Cl pH 8.0, 200 mM NaCl, 10% glycerol, and 1 mM TCEP. The purified CBE3 protein was concentrated to 5 mg / mL in a Vivaspin protein concentration spin column (GE Healthcare) at 4 °C and then snap-frozen in small aliquots in liquid nitrogen. RNPs were prepared as follows: 20 μg of CBE3 protein, 2 μg of target sgRNA, and 2 μg of selection sgRNA (TrueGuide Synthetic gRNA, Life Technologies), and 2.4 μg of electroporation enhancer oligonucleotide (Sigma) (Table 3E) were mixed and incubated for 15 minutes. Cells were washed with PBS and resuspended in Buffer R at a concentration of 5×10 7 cells / mL. 5×10 5Cells were electroporated with RNP using the following settings: voltage: 1600 V, width: 10 ms, pulse number: 3. After electroporation, cells were incubated overnight in 1 mL of RPMI medium supplemented with 10% heat-inactivated FBS in a 24-well plate. The next day, cells were harvested, centrifuged at 300×g for 5 minutes, resuspended in 1 mL of complete growth medium containing 500 U / mL of IL-2 (Prepotech), and divided into 5 wells of a round-bottom 96-well plate.

[0313] Diphtheria toxin (DT) treatment Transfected HEK293, HCT116, and PC9-BFP cells were selected with 20 ng / mL of DT on days 3 and 5 after transfection. iPSCs were treated with 20 ng / mL of DT on day 3 after transfection. The growth medium supplemented with DT was changed daily until the negative control cells died. Transfected CD4+ T cells were treated with 1000 ng / mL of DT on days 1, 4, and 7 after electroporation.

[0314] Alamar Blue assay Cell viability was analyzed using AlamarBlue cell viability reagent (ThermoFisher) according to the manufacturer's instructions.

[0315] PCR analysis PCR analysis was performed to distinguish successful knock-in into HBEGF intron 3 (PCR1) from the wild-type sequence (PCR2). The PCR reaction was performed in a 20 μL volume using 1.5 μL of extracted genomic DNA as a template. PHUSION (ThermoFisher) was used at a primer concentration of 0.5 μM according to the manufacturer's recommended protocol. Primer pairs PCR1_fwd and PCR1_rev were used for PCR1 to detect the knock-in junction (annealing temperature; 62 °C, extension time: 1 minute), and primer pairs PCR2_fwd and PCR2_rev were used for PCR2 to detect the wild-type HBEGF intron (annealing temperature; 64.5 °C, extension time: 5 seconds). The sequences of the primer pairs are provided in Table 3D. For PCR2, the extension time was set to 5 seconds to support amplification of the wild-type HBEGF intron 3 product (280 bp) over the integrated PCT product (2229 bp).

[0316] Flow cytometry analysis The frequency of cells expressing mCherry and GFP was evaluated with a BD Fortessa flow cytometer (BD Biosciences), and the flow cytometry data were analyzed with FlowJo software (Three Star).

[0317] Genomic DNA extraction and next-generation amplicon sequencing Genomic DNA was extracted from cells 3 days after transfection or complete DT selection using QuickExtract DNA Extraction Solution (Lucigen) according to the manufacturer's instructions. The target amplicons were analyzed from genomic DNA samples on the NextSeq platform (Illumina). The target genomic sites were amplified in the first round of PCR using primers containing NGS forward and reverse adapters (Table 3C). The first PCR was set up with 0.5 μM primers and 1.5 μL of genomic DNA using NEBNext Q5 Hot Start HiFi PCR Master Mix (New England Biolabs) in a 15 μL reaction. PCR was performed under the following cycle conditions: 2 minutes at 98°C, 10 seconds at 98°C, 20 seconds at the annealing temperature of each pair of primers (calculated using the NEB Tm Calculator), and 10 seconds at 65°C for 5 cycles, followed by 10 seconds at 98°C, 20 seconds at 98°C, and 10 seconds at 65°C for 25 cycles, followed by a final extension at 65°C for 5 minutes. The PCR products were purified using the HighPre PCR Clean-up System (MAGBIO Genomics), and the size and DNA concentration of the correct PCR products were analyzed using the Fragment Analyzer (Agilent). In the second round of PCR, unique Illumina indexes were added to the PCR products using KAPA HiFi HotStart Ready Mix (Roche). Index primers were added in the second PCR step, and 1 ng of the purified PCR product from the first PCR was used as a template in a 50 μL reaction. PCR was performed under the following cycle conditions: 3 minutes at 72°C, 30 seconds at 98°C, followed by 10 seconds at 98°C, 30 seconds at 63°C, and 3 minutes at 72°C for 10 cycles, followed by a final extension at 72°C for 5 minutes. The final PCR products were purified using the HighPre PCR Clean-up System (MAGBIO Genomics) and analyzed using the Fragment analyzer (Agilent).The library was quantified using a Qubit 4 Fluorometer (Life Technologies) and pooled for sequencing on a NextSeq instrument (Illumina).

[0318] Bioinformatics NGS sequencing data was demultiplexed using bcl2fastq software, and individual FASTQ files were analyzed using a Perl implementation of a Matlab script described in previous publications. For quantification of indel or base editing frequencies, sequencing reads were scanned for matches to two 10-bp sequences flanking an intervening window where an indel or base edit could occur. If no match was found (allowing a maximum of 1 bp mismatch on each side), the read was excluded from the analysis. If the length of the intervening window was longer or shorter than the reference sequence, the sequencing reads were classified as insertions or deletions, respectively. The frequency of insertions or deletions was calculated as the percentage of reads classified as insertions or deletions within the total number of reads analyzed. If the length of this intervening window exactly matched the reference sequence, the read was classified as not containing an indel. For these reads, the frequency of each base at each locus was calculated within the intervening window and used as the frequency of base editing.

[0319] Cytidine base editing and DT treatment in humanized mice for hHBEGF expression All mouse experiments were approved by the AstraZeneca internal committee regarding animal research and by the Gothenburg Ethics Committee for Experimental Animals regarding experimental animals (license number: 162-2015+), and these comply with the EU directive on the protection of animals used for scientific purposes. Experimental mice were generated as double heterozygotes by breeding Alb-Cre mice (The Jackson Laboratory) onto an iDTR mouse (expression of the transgene human HBEGF is blocked by a STOP sequence flanked by loxP) on a C57BL / 6NCrl genetic background. Mice were housed in negative pressure IVC cages in a temperature-controlled room (21 °C) that had a 12:12 h light-dark cycle (dawn: 5:30 am, lights on: 6:00 am, dusk: 5:30 pm, lights off: 6 pm) and was humidity-controlled (45 - 55%). Mice had free access to a normal solid diet (R36, Lactamin AB) and water.

[0320] For base editing, six 6-month-old male and six 6-month-old female mice were randomized into two groups, with an equal number of male and female mice in each group. An adenoviral vector expressing a sgRNA targeting the CBE3, sgRNA10, and mouse Pcsk9 (1 × 10 per mouse) 9 was intravenously injected. Two weeks after virus administration, all mice were intraperitoneally administered DT (200 ng / kg). Control mice were sacrificed 24 hours after DT injection. Experimental mice were sacrificed 11 days after DT injection. Four mice were sacrificed prior to the experimental endpoint as they reached the humane endpoint of the ethics license. At autopsy, liver tissue was collected for morphological and molecular analysis.

[0321] Example 6. Amino Acid Substitutions in HBEGF In this example, base editing was used to scan for mutations in the human EGF-like domain that confer resistance to diphtheria toxin (DT) in cells.

[0322] A detailed experimental protocol is described in Example 5. Briefly, to screen for sgRNAs, each sgRNA was co-transfected at a weight ratio of 1:4 with CBE3 or ABE7.10. Transfection was performed using FuGENE HD transfection reagent (Promega) according to the manufacturer's instructions, using a 3:1 transfection reagent to plasmid DNA ratio. Cells were treated with 20 ng / mL of diphtheria toxin 3 days after transfection and then again 5 days after transfection. Cell viability was analyzed using AlamarBlue cell viability reagent (Thermo Fisher) according to the manufacturer's instructions. Genomic DNA was extracted from viable cells and analyzed by Amplicon-Seq using next-generation sequencing (NGS).

[0323] Fourteen single-guide RNAs (sgRNAs) (Figure 24A) that tile the exon sequences encoding the human EGF-like domain, covering all regions encoding amino acids different from the mouse EGF-like domain. Each sgRNA was transiently expressed in HEK293 cells together with either cytidine base editor 3 (CBE3) or adenosine base editor 7.10 (ABE7.10). The corresponding mutations, C to T (by CBE3) or A to G (by ABE7.10), were introduced into the editing window of each sgRNA. Edited cells were treated with a lethal dose of DT (20 ng / μl in HEK293 cells) 72 hours after transfection, and cell proliferation was monitored. The results in Figure 24B show that CBE3 in combination with sgRNA7 or sgRNA10 induced effective resistance mutations to DT in HBEGF, and ABE7.10 induced resistance in combination with sgRNA5 or sgRNA10.

[0324] The combination of ABE7.10 / sgRNA5 or CBE3 / sgRNA10 was selected for further analysis. Genomic DNA was harvested from the resistant cells, and its corresponding targeted locus was analyzed by Amplicon-Seq using next-generation sequencing (NGS). The majority of the mutations introduced by the combination of CBE3 and sgRNA10 in the resistant cells resulted in a Glu141Lys substitution in HBEGF. Approximately 90% of the variants introduced by the combination of ABE7.10 / sgRNA5 resulted in a Tyr123Cys conversion in HBEGF (see Figures 24C and 25A - C). No impaired proliferation was observed in the edited cells compared to wild-type cells, indicating that there is no deleterious effect introduced by the edited HBEGF variants (Figure 25D).

[0325] Collectively, these data showed that DT resistance can be introduced without changing cell proliferation by modifying a single amino acid in the HBEGF protein using base editing. Therefore, the DT-HBEGF system can be effectively applied to select genomic editing events in cells.

[0326] Example 7. Enrichment of Cytidine and Adenosine Base Editing In this example, the DT-HBEGF selection system was tested for enrichment of base editing events at a second, unrelated genomic locus. Figure 26A provides an overview of the DT-HBEGF co-selection strategy.

[0327] A detailed experimental protocol is described in Example 5. Briefly, for co-targeting enrichment, Cas9 / CBE3 / ABE7.10 plasmid DNA, targeting sgRNA plasmid DNA, and selection sgRNA plasmid DNA were transfected at a weight ratio of 8:1:1. Transfection was carried out using FuGENE HD transfection reagent (Promega) according to the manufacturer's instructions, using a transfection reagent to plasmid DNA ratio of 3:1. Cells were treated with 20 ng / mL of diphtheria toxin 3 days after transfection and then again 5 days after transfection. Genomic DNA was extracted from viable cells and analyzed by Amplicon-Seq using next-generation sequencing (NGS).

[0328] First, CBE co-selection in HEK293 cells was performed. sgRNAs targeting five different genomic loci were tested: DPM2 (dolichyl-phosphate mannosyltransferase subunit 2), EGFR (epidermal growth factor receptor), EMX1 (empty spiracles homeobox 1), PCSK9 (proprotein convertase subtilisin / kexin type 9), and DNMT3B (DNA methyltransferase 3β). Each of these sgRNAs was co-transfected into cells together with CBE3 and sgRNA10 as described in Example 6, and the selected cells were enriched with DT (20 ng / μL) starting 72 hours after transfection. Thereafter, genomic DNA was harvested from the cells with or without selection and analyzed by NGS.

[0329] Notably, a significant increase in the C-T conversion rate was observed across all test sites in the cells selected by DT, compared to the unselected cells, and the fold change ranged from 4.1-fold to 7.0-fold (Figure 26B). In the case of the DPM2 site, the total conversion rate increased from 20% to 94% by DT selection (Figure 26B). Similar improvements in editing efficiency were also observed when this method was applied to other cell lines. A 12.8-fold increase in the C-T conversion rate at the PCSK9 locus in HCT116 cells and a 4.9-fold increase at the integrated BFP locus in DT-treated PC9 cells, compared to untreated cells (Figure 26C).

[0330] Similar co-selection experiments were performed for the enrichment of ABE editing events. Five sgRNAs were tested, including one targeting EMX1 and the other four targeting different two sites of novel genomic loci (CTLA4 (cytotoxic T lymphocyte-associated protein 4), IL2RA (interleukin 2 receptor subunit α), and AAVS1 (adeno-associated virus integration site 1)). Each of these sgRNAs was co-transfected into HEK293 cells together with ABE7.10 and sgRNA5 as described in Example 6. After 72 hours, the selected cells were treated with DT (20 ng / μl). Genomic DNA was extracted from both the selected and unselected cells and analyzed by Amplicon-Seq. A dramatic improvement in the A-G conversion rate from 5.7-fold to 12.7-fold was observed across all tested targets in the selected cells, compared to the unselected cells. At the CTLA4 and IL2RA of the targeted loci, the total conversion rate increased from 4.6% to 39% and from 11.5% to 77.4%, respectively (Figure 26D).

[0331] In addition to co-selection for base editing events, the possibility of co-selection indels generated by SpCas9 was also tested. Four sgRNAs used for CBE co-selection (targeting DPM2, EMX1, PCSK9, and DNMT3B, respectively) were tested in experiments on genome editing co-selection. Each sgRNA was co-transfected into HEK293 cells together with the SpCas9 / sgRNA10 combination (described above in Example 6) to generate indels, and Amplicon-Seq was performed after selection. It was observed that the indel rate increased to over 90% across all four targets (DPM2, EMX1, PCSK9, and DNMT3B). In particular, the editing efficiency at the PCKS9 site increased from 30% to 98% through DT selection (Figure 26E).

[0332] Example 8. Efficient Enrichment of Biallelic Knock-In Events at the HBEGF Locus In this example, experiments were conducted to enhance the knock-in efficiency of the target gene or to achieve biallelic knock-in of the target gene.

[0333] The detailed experimental protocol is described in Example 5. Briefly, for the knock-in experiment, Cas9 plasmid DNA, sgRNAIn3 plasmid DNA, and template DNA were transfected at a weight ratio of 4:1:10. Transfection was performed using FuGENE HD transfection reagent (Promega) according to the manufacturer's instructions using a transfection reagent to plasmid DNA ratio of 3:1. Twenty-two days after transfection, the cells were evaluated with a BD Fortessa (BD Biosciences), and the flow cytometry data were analyzed with FlowJo software (Three Star). Genomic DNA was also extracted from the cells, and PCR analysis was performed to distinguish successful knock-in into HBEGF intron 3 (PCR1) from the wild-type sequence (PCR2).

[0334] It was hypothesized that cells could be rendered DT-resistant by knock-in in intron 3 of HBEGF with a cassette containing a strong splicing acceptor combined with a cDNA sequence that contains all of the exon remaining downstream of exon 3 and contains mutations that prevent the ligation of DT. Based on the base editing screening described in Example 6 and the presence of similar substitutions in mouse HBEGF, a Glu141Lys amino acid substitution was inserted (see Figure 25A). To further rule out the possibility of a deleterious effect of this substitution on cell viability, recombinant Glu141Lys-substituted HBEGF protein was shown to still be functional in inducing p44 / p42 MAPK phosphorylation and no significant difference was observed compared to wild-type HBEGF, suggesting that its major function in EGFR activation is maintained (Figure 27A).

[0335] Subsequently, a knock-in strategy was designed to introduce DT-resistant HBEGF bound to the gene of interest. First, an sgRNA (sgRNAIn3) targeting the middle region of intron 3 of HBEGF, which has low predicted off-target sites and is efficient in inducing indels at the target site, was selected. The repair template was also designed to contain a splice acceptor and encode the Glu141Lys substitution and the remainder of the mutant HBEGF exon sequence linked to the gene of interest (e.g., mCherry or GFP) by a T2A self-cleaving peptide (Figure 27B). In this design, wild-type cells or edited cells with small indels in intron 3 do not acquire DT resistance, while cells containing the desired knock-in become DT-resistant.

[0336] Repair templates were tested in various forms including plasmid, double-stranded DNA (dsDNA), and single-stranded DNA (ssDNA) to determine knock-in efficiency. The templates were designed with or without homology arms or flanking sgRNAs and were expected to be integrated into the HBEGF locus by non-homologous end joining (NHEJ), homologous recombination (HR), or homology-mediated end joining (HMEJ) (Figure 27C). Each template was co-transfected into HEK293 cells with SpCas9 and sgRNAIn3 to generate knock-in cells. Selection was performed as described above. Since the expression of the mCherry or GFP gene is linked to the mutant HBEGF gene, only cells containing the correct insertion were expected to express the functional fluorescent protein. The percentage of knock-in cells (fluorescent cells) was quantified by flow cytometry analysis.

[0337] Notably, mCherry or GFP positive cells were generated regardless of the template applied, and the proportion of knock-in cells was observed to increase dramatically after selection under all conditions (Figure 27C). In particular, cells repaired with plasmid templates containing homology arms and sgRNA (pHMEJ) or plasmid templates containing only homology arms (pHR) achieved nearly 100% knock-in after selection (Figure 27C). Among all the templates tested, pHMEJ was shown to be the most efficient, with only 34.8% of the knock-in cells obtained without selection (Figure 27C). These observations are consistent with additional results showing bi-allelic mutations in base editing selection (Figure 24B), suggesting that cells may require bi-allelic knock-in to survive DT treatment. Two pairs of primers were designed to check the genomic status of the edited cells, one pair amplifying the 5’ junction of the knock-in sequence (PCR1) and the other pair amplifying the wild-type sequence of the HBEGF intron (PCR2). PCR analysis was performed on cells repaired with the pHMEJ template with or without selection. Despite both samples showing bands of homologous knock-in (PCR1), only the wild-type band was detected in the unselected samples (Figure 27E), indicating that all cells obtained bi-allelic knock-in after DT selection.

[0338] The DT selection method was further compared with the conventional antibiotic-dependent selection method with respect to enrichment of knock-in events. A novel pHMEJ template was designed to contain both the DT resistance mutation and the puromycin resistance gene, and the expression of these two selectable markers was linked by a P2A self-cleaving peptide (Figure 27D). This novel template for knock-in was tested, and knock-in cells were enriched with either DT or puromycin, followed by flow cytometry analysis. Interestingly, nearly 100% of mCherry-positive cells were observed in both populations, but DT-enriched cells showed a dramatically higher mean fluorescence intensity compared to puromycin-enriched cells (Figure 27D). This observation, together with PCR analysis (Figure 27E), indicated that DT selection enriched cells containing a two-allele knock-in, while puromycin selection did not.

[0339] This genetic engineering strategy is referred to herein as "Xential" (recombination (X) in an essential locus for cell survival).

[0340] Example 9. Enrichment of knockout and knock-in events by Xential co-selection In this example, Xential knock-in was tested for enrichment of knockout or knock-in events at a second, unrelated locus.

[0341] A detailed experimental protocol is described in Example 5. Briefly, for the Xential co-selection experiment, the amounts of each transfected plasmid are listed in Table 7 below. Transfection was performed using FuGENE HD transfection reagent (Promega) according to the manufacturer's instructions, with a 3:1 transfection reagent to plasmid DNA ratio. Cells were treated with 20 ng / ml of diphtheria toxin 3 days after transfection and then again 5 days after transfection. Twenty-two days after transfection, cells were evaluated on a BD Fortessa (BD Biosciences), and flow cytometry data were analyzed with FlowJo software (Three Star). Genomic DNA was also extracted from the cells, and the same PCR analysis and Amplicon-Seq analysis were performed as described for the previous examples.

[0342]

Table 16

[0343] First, enrichment of knockout events was tested. The same four sgRNAs (targeting DPM2, EMX1, PCSK9, and DNMT3B, respectively) that were tested in the indel enrichment experiment described in Example 7 (Figure 26E) were used. Each sgRNA was co-delivered into HEK293 cells together with SpCas9, sgRNAIn3, and the pHMEJ template, and DT selection was performed as described in Figure 28A. Genomic DNA was extracted from these cells and analyzed by Amplicon-Seq. A significant improvement in editing efficiency across all targets (ranging from 4.4-fold to 14.3-fold improvement) was observed in the selected cells compared to the unselected cells. In particular, the editing efficiency at the EMX1 locus increased from 22% to 88% by DT selection (Figure 28B). All surviving cells maintained mCherry expression, indicating that the edited cells maintained accurate knock-in at the HBEGF locus (Figure 28D).

[0344] Subsequently, Xential was tested for co-selection of knock-in events. Two forms of repair template plasmids were designed (one with pHR and one with pHMEJ), and the same sgRNA was used to introduce a C-terminal GFP tag into the histone protein H2B (HIST2BC). SpCas9, sgRNA, and two templates targeting HIST2BC and HBEGF were co-delivered into HEK293 cells, and the knock-in efficiency was analyzed by the ratio of GFP (HIST2BC) or mCherry (HBEGF). A significantly improved knock-in efficiency was obtained after DT selection in either form of the provided template. For the pHR template, the efficiency was improved up to 6.4-fold, and for the pHMEJ template, the efficiency was improved up to 5.3-fold and reached 48% (Figure 28C). By reducing the ratio of the amount of sgRNA and template at the HBEGF locus to that at the HIST2BC locus, the knock-in efficiency at the HIST2BC locus in the selected cells can be increased, indicating that the enrichment magnification is adjustable (Figure 28C). The percentage of GFP-positive cells in the enriched cells increased from 23% to 42% to 48% by applying increasing weight ratios of the repair plasmid at the HIST2BC locus to that at the HBEGF locus, 1:1 to 3:1 and 1:1 to 4:1, respectively, while the percentage of mCherry-positive cells was maintained at approximately 100% (Figure 28E). This method also demonstrated the enrichment of the efficiency of oligo-mediated knock-in at the CD34 locus. When co-selection was applied, a 26-fold increase in the percentage of knock-in cells was observed, suggesting the flexibility of template use in knock-in-mediated co-selection (Figure 28F).

[0345] Example 10. Enrichment of Base Editing and Knock-In Events in iPSCs In this example, experiments were conducted using DT-HBEGF selection to enrich base editing events and precise knock-in events in iPSCs.

[0346] A detailed experimental protocol is described in Example 5. Briefly, for co-selection of CBE / ABE in iPSCs, iPSC / CBE3 / ABE7.10 plasmid DNA, targeting sgRNA plasmid DNA, and selection sgRNA plasmid DNA were transfected at a weight ratio of 8:1:1. For Xential knock-in in iPSCs, Cas9 plasmid DNA, sgRNAIn3 plasmid DNA, and template plasmid DNA were transfected at a weight ratio of 4:1:10. Transfection was performed using FuGENE HD transfection reagent (Promega) according to the manufacturer's instructions, with a transfection reagent to plasmid DNA ratio of 2.5:1 and a reverse transfection protocol. Three days after transfection, the cells were treated with 20 ng / ml of diphtheria toxin. The growth medium supplemented with DT was changed daily until the negative control cells died. Xential knock-in cells were evaluated with BD Fortessa (BD Biosciences), and the flow cytometry data were analyzed with FlowJo software (Three Star). Genomic DNA was also extracted from the cells, and the same PCR analysis and Amplicon-Seq analysis were performed as described for the previous examples.

[0347] Two sgRNAs were selected for co-selection of CBE and ABE. One targets EMX1 at a locus that has been widely tested in other genome editing studies, and the other targets CTLA4 of a gene that has been extensively studied for its role in immune signaling. Each sgRNA was co-transfected into iPSCs with the CBE3 / sgRNA10 or ABE7.10 / sgRNA5 pair. Selection was performed by DT treatment (20 ng / μL) starting 72 hours after transfection. Genomic DNA was extracted at the confluence time point, and the target locus was analyzed by Amplicon-Seq using NGS. Notably, a dramatic increase in editing efficiency at the time of DT selection was observed at all test sites for both CBE and ABE. The increase in CBE editing efficiency ranged from 19-fold to 60-fold across these two sites, and the increase in ABE editing efficiency was approximately 24-fold at both sites. Through DT selection, the C-T conversion rate at the EMX1 site increased from 5% to 91%, and the A-G conversion rate at the CTLA4 site increased from 0.8% to 19% (Figures 29A, B).

[0348] Next, Xential was tested in iPSCs. The iPSCs were provided with the pHMEJ template together with SpCas9 and sgRNAIn3, and the knock-in efficiency was 25.6% without selection. The knock-in efficiency increased to nearly 100% after DT selection (Figure 29C). To detect accurate insertion and the wild-type HBEGF intron, the same PCR analysis as in Example 8 was performed. After DT selection, no wild-type band remaining in the targeted HBEGF was detected, suggesting complete biallelic knock-in in the selected pool of iPSCs (Figure 29D).

[0349] Example 11. Enrichment of Base Editing Events in Primary T Cells In this example, experiments were performed using DT-HBEGF selection to enrich for cytidine base editing events in primary T cells at a second, unrelated genomic locus. Further, experiments were performed using the DT-HBEGF selection system to enrich for knock-in events at the HBEGF locus.

[0350] A detailed experimental protocol is described in Example 5. Briefly, for CBE co-selection in primary T cells, 20 μg of CBE3 protein, 2 μg of target sgRNA, and 2 μg of selection sgRNA (TrueGuide Synthetic gRNA, Life Technologies), and 2.4 μg of electroporation enhancer oligonucleotide (HPLC purified, Sigma) (Table 3E) were mixed, incubated for 15 minutes, and subsequently electroporated into primary T cells. Transfected CD4+ T cells were treated with 1000 ng / mL of DT on days 1, 4, and 7 after electroporation. As described in the foregoing examples, genomic DNA was also extracted from the cells and Amplicon-Seq analysis was performed. For the Xential experiment in primary T cells, 5 μg of SpCas9 protein (Life Technologies), 1.2 μg of dual gRNA In3 (Alt-R CRISPR-Cas9 crRNA, Alt-R CRISPR-Cas9 tracrRNA, IDT) were mixed, incubated for 15 minutes, and subsequently electroporated into primary T cells together with 1 μg of dsDNA template. Transfected CD4+ T cells were treated with 1000 ng / mL of DT on days 1, 4, 6, and 8 after electroporation. On day 10 after electroporation, the cells were analyzed by flow cytometry.

[0351] Three sgRNAs were designed to introduce premature stop codons into PCDC1 (programmed cell death protein 1), CTLA4, and IL2RA, respectively, due to their important roles in immune regulation. Each sgRNA was co-electroporated into isolated CD4+ T cells together with purified CBE3 protein and synthetic sgRNA10. Primary T cells were selected with 1000 ng / μL DT starting 24 hours after electroporation, and genomic DNA from unselected and selected cells was analyzed 9 days after transfection. An increase in base editing efficiency from 1.7- to 1.8-fold was observed at all three loci compared to unselected cells (Figure 30). Three different forms of dsDNA (dsHR, dsHMEJ, dsHR2) described in Figure 3 were applied as repair templates. Each template was electroporated into primary CD4+ T cells together with premixed SpCas9 protein and synthetic dual gRNA In3 complex. Primary T cells containing 1000 ng / μL DT were selected starting 24 hours after electroporation, and the knock-in efficiency of unselected and selected cells was analyzed 10 days after transfection. An increase in knock-in efficiency from 3- to 8-fold was observed for all three versions of the template in selected cells compared to unselected cells.

[0352] Example 12. Enrichment of Base Editing Events In Vivo by Co-Selection In this example, experiments were performed using DT-HBEGF selection to enrich for cytosine base editing events in a humanized mouse model at a second, irrelevant genomic locus.

[0353] A detailed experimental protocol is described in Example 5 (see the section "Cytosine Base Editing and DT Treatment of Humanized Mice for hHBEGF Expression").

[0354] In a humanized mouse model expressing human HBEGF (hHBEGF), cotransselection of cytosine base editing events was tested under the hepatocyte-specific albumin promoter. The mouse Pcsk9 gene was selected as the target locus, and the sgRNA was designed to introduce premature stop codons into Pcsk9 together with CBE3 by an adenovirus (AdV8) delivering CBE3, an sgRNA targeting Pcsk9, and an sgRNA targeting human HBEGF. Two weeks after AdV8 injection, the mice were treated with DT (200 ng / kg, intraperitoneally). The mice were divided into two groups, and the non-enriched control was terminated in 24 hours before DT could exert toxicity. The enriched group was terminated 11 days after DT treatment (Figure 31A). Amplicon-Seq analysis of the genome from mouse liver showed a 2.8-fold increase in base editing efficiency at the selected locus as a result of DT selection (Figure 31B). Notably, in the enriched group, a 2.5-fold improvement in Pcsk9 editing was also confirmed compared to the control group (Figure 31C), which demonstrated for the first time that genome editing events can be cotransselected in vivo using toxin-mediated selection.

[0355] Example 13 Enrichment of Prime Editing Events by Cotransselection In this experiment, the DT-HBEGF selection system was used for enrichment of prime editing events at a second unrelated genomic locus.

[0356] For co-targeted enrichment, the PE2 plasmid DNA, the targeted pegRNA plasmid DNA, and the selection pegRNA_HBEGF12 plasmid DNA were transfected at a weight ratio of 8:1:1. Transfection was performed using the FuGENE HD transfection reagent (Promega) at a transfection reagent to plasmid DNA ratio of 3:1. The cells were treated with 20 ng / ml of diphtheria toxin 3 days after transfection and then again 5 days after transfection. Genomic DNA was extracted from the surviving cells and analyzed by Amplicon-Seq using next-generation sequencing (NGS).

[0357] Prime editing co-selection in HEK293 cells was tested. Four prime editing guide RNAs (pegRNAs) were used to target three different genomic loci: EMX1 (empty spiracles homeobox 1), FANCF (Fanconi anemia complementation group F), and HEK3. Each of these pegRNAs was co-transfected into cells with prime editor 2 (PE2) and pegRNA_HBEGF12 (designed to introduce an E141H resistant mutation at the HBEGF locus), and the selected cells were enriched with DT (20 ng / mL) starting 72 hours after transfection. Genomic DNA was then harvested from the cells with or without selection and analyzed by NGS. A significant increase in prime editing efficiency at the HBEGF locus (from approximately 1% to over 99%) was observed. At all co-selected target loci, editing efficiencies above the average were observed in DT-selected cells compared to unselected cells, with fold increases ranging from 1.5-fold to 44-fold.

[0358] Example 14 Enrichment of Cas9 editing events by co-selection with an anti-CD52 antibody-drug maytansinoid (DM1) conjugate (anti-CD52-DM1) In this experiment, an anti-CD52-DM1 antibody conjugate drug was used for the selection of SpCas9 editing events at a second, unrelated genomic locus.

[0359] SpCas9 editing co-selection in primary CD4+ T cells was tested. Three sgRNAs were used to target three different genomic loci: PDCD1, CTLA4, and IL2RA, respectively.

[0360] For co-selection of SpCas9 in primary T cells, 5 μg of TrueCut Cas9 Protein v2 (Life Technologies), 0.6 μg of target sgRNA, and 0.6 μg of selection sgRNA (TrueGuide Synthetic gRNA, Life Technologies) and 0.8 μg of electroportation enhancer oligonucleotide for Cas9 (HPLC purified, Sigma) (Table S1) were mixed, incubated for 15 minutes, and subsequently electrotransfected into primary T cells. Transfected CD4+ T cells were treated with 2.5 μg / ml of anti-CD52-DM1, 2.5 μg / ml of NIP228-DM1, and PBS on days 2, 4, and 6 after electroportation, respectively. Genomic DNA was also extracted from the cells and Amplicon-Seq analysis was performed.

[0361] The anti-CD52, alemtuzumab, (Campath-1) antibody sequence was read from the Drugbank database (https: / / www.drugbank.ca / drugs / DB00087), and the antibody variable light and heavy gene segments were designed and obtained from Thermofisher for cloning into the in-house pOE IgG1 antibody expression vector. The cloned pOE-anti-CD52.IgG1 expression construct was transfected into CHO-G22 cells and cultured for 14 days. The conditioned medium was harvested, filtered (0.2 μM filter), and purified via Protein A using an Aligent Pure FPLC instrument. The antibody was dialyzed against 1×PBS at pH 7.2, and binding to the human CD52 antigen (Abcam) was confirmed via SPR using an Octet and compared to commercially available Campath-1. Furthermore, the molecular weight was verified using mass spectrometry, and the monomer content was determined by size exclusion chromatography. The anti-CD52 and negative control (NIP228) mAbs were buffer-exchanged into 1×borate buffer at pH 8.5, and 40 mg of each antibody was incubated with 4.5 molar equivalents of the SMCC-DM1 payload. The degree of drug conjugation was determined by reduced reverse phase mass spectrometry, and the reaction was stopped by adding 10% v / v of 1M Tris-HCl. Ceramic hydroxyapatite chromatography was used to simultaneously remove free or unconjugated SMCC-DM-1 payload and protein aggregates. Subsequently, the ADC was dialyzed against PBS pH 7.2. The concentration and endotoxin level were measured using a nanodrop (Thermofisher) and an Endosafe (Charles Rivers) instrument, respectively.

[0362] Each synthetic sgRNA was co-electroporated into isolated CD4+ T cells together with the synthetic sgRNA targeting SpCas9 protein and CD52. The electroporated T cells were treated with 2.5 μg / ml of anti-CD52-DM1, 2.5 μg / ml of NIP228-DM1 (negative control antibody-drug conjugate), and PBS (untreated) starting 48 hours after electroporation, and genomic DNA from the treated cells was analyzed 7 days after the first treatment. Thereafter, genomic DNA was harvested from the cells with or without selection and analyzed by NGS. An increase in the indel rate was observed in the samples treated with anti-CD52-DM1 compared to the samples treated with Nip228-DM1 or PBS (untreated). A two-tailed paired t test was performed to compare the difference between the indel rates of anti-CD52-DM1-treated cells and Nip228-DM1-treated cells, which showed that the increase in the indel rate at the targeted loci (IL2RA, CTLA4, PDCD1) was significant (P = 0.0044). The same analysis comparing the indel rate of cells treated with anti-CD52-DM1 and the indel rate of untreated cells also showed that the increase in the indel rate at the targeted loci was significant (P = 0.0008).

Claims

1. 1. A method for introducing site-specific mutations into a target polynucleotide in a target cell in a cell population, comprising: (a) introducing into the cell population (i) a base editing enzyme; and (ii) (1) hybridizing with a gene encoding a cytotoxic agent (CA) receptor; (2) A first guide polynucleotide that forms a first complex with the base editing enzyme, a first guide polynucleotide, wherein the base editing enzyme of the first complex provides a mutation in the gene encoding the CA receptor, and the mutation in the gene encoding the CA receptor forms CA resistant cells in the cell population; (iii) (1) hybridizing with the target polynucleotide; (2) A second guide polynucleotide that forms a second complex with the base editing enzyme, The base editing enzyme of the second complex provides a mutation in the target polynucleotide; and Introducing (b) contacting the cell population with the CA; (c) enriching the target cells containing the mutation in the target polynucleotide by selecting the CA resistant cells from the cell population; A method comprising:

2. 1. A method for determining the efficacy of a base editing enzyme in a cell population, comprising: (a) introducing into the cell population (i) a base editing enzyme; and (ii) (1) hybridizing with a gene encoding a cytotoxic agent (CA) receptor; (2) A first guide polynucleotide that forms a first complex with the base editing enzyme, a first guide polynucleotide, wherein the base editing enzyme of the first complex introduces a mutation in the gene encoding the CA receptor, and the mutation in the gene encoding the CA receptor forms CA resistant cells in the cell population; (iii) (1) hybridizing with the target polynucleotide; (2) A second guide polynucleotide that forms a second complex with the base editing enzyme, a second guide polynucleotide, wherein the base editing enzyme of the second complex introduces a mutation in the target polynucleotide; and Introducing (b) contacting the cell population with the CA to isolate CA-resistant cells; (c) determining the ratio of the CA resistant cells to the total cell population to determine the potency of the base editing enzyme; and A method comprising:

3. The method of claim 1 or 2, wherein the base editing enzyme comprises a DNA targeting domain and a DNA editing domain.

4. 4. The method of claim 3, wherein the DNA targeting domain comprises Cas9.

5. 5. The method of claim 4, wherein the Cas9 comprises a mutation in a catalytic domain.

6. The method of any one of claims 1 to 5, wherein the base editing enzyme comprises a catalytically inactive Cas9 and a DNA editing domain.

7. The method of any one of claims 1 to 5, wherein the base editing enzyme comprises Cas9 (nCas9) capable of generating single-stranded DNA breaks and a DNA editing domain.

8. 8. The method of claim 7, wherein the nCas9 comprises a mutation relative to wild-type Cas9 at amino acid residue D10 or H840 (numbering is relative to SEQ ID NO: 3).

9. The method of any one of claims 4 to 8, wherein the Cas9 is at least 90% identical to SEQ ID NO: 3 or 4.

10. The method of any one of claims 3 to 9, wherein the DNA editing domain comprises a deaminase.

11. The method of claim 10, wherein the deaminase is a cytidine deaminase or an adenosine deaminase.

12. 12. The method of claim 11, wherein the deaminase is a cytidine deaminase.

13. 12. The method of claim 11, wherein the deaminase is adenosine deaminase.

14. The method of any one of claims 10 to 13, wherein the deaminase is apolipoprotein B mRNA editing complex (APOBEC) deaminase, activation-induced cytidine deaminase (AID), ACF1 / ASE deaminase, ADAT deaminase, or ADAR deaminase.

15. 15. The method of claim 14, wherein the deaminase is an apolipoprotein B mRNA editing complex (APOBEC) family deaminase.

16. 16. The method of claim 15, wherein the deaminase is APOBEC1.

17. The method of any one of claims 3 to 16, wherein the base editing enzyme further comprises a DNA glycosylase inhibitor domain.

18. 18. The method of claim 17, wherein the DNA glycosylase is uracil DNA glycosylase inhibitor (UGI).

19. The method of any one of claims 1 to 4 or 6 to 18, wherein the base editing enzyme comprises nCas9 and cytidine deaminase.

20. The method of any one of claims 1 to 4 or 6 to 18, wherein the base editing enzyme comprises nCas9 and adenosine deaminase.

21. The method of any one of claims 1 to 12 or 13 to 19, wherein the base editing enzyme comprises a polypeptide sequence that is at least 90% identical to SEQ ID NO:

6.

22. The method of any one of claims 1 to 12 or 13 to 19, wherein the base editing enzyme is BE3.

23. 23. The method of any one of claims 1 to 22, wherein the first and / or second guide polynucleotide is an RNA polynucleotide.

24. 24. The method of any one of claims 1 to 23, wherein the first and / or second guide polynucleotide further comprises a tracrRNA sequence.

25. The method of any one of claims 1 to 24, wherein the cell population is human cells.

26. The method of any one of claims 1 to 25, wherein the mutation in the gene encoding the CA receptor is a cytidine (C) to thymine (T) point mutation.

27. The method of any one of claims 1 to 25, wherein the mutation in the gene encoding the CA receptor is an adenine (A) to guanine (G) point mutation.

28. The method of any one of claims 1 to 27, wherein the CA is diphtheria toxin.

29. 29. The method of claim 28, wherein the cytotoxic agent (CA) receptor is a receptor for diphtheria toxin.

30. 30. The method of claim 29, wherein the CA receptor is heparin-binding EGF-like growth factor (HB-EGF).

31. The method of claim 30, wherein the HB-EGF comprises the polypeptide sequence of SEQ ID NO:

8.

32. 32. The method of claim 31, wherein the base editing enzyme of the first complex provides a mutation in one or more of amino acids 107-148 of HB-EGF (SEQ ID NO:8).

33. 33. The method of claim 32, wherein the base editing enzyme of the first complex provides a mutation in one or more of amino acids 138-144 of HB-EGF (SEQ ID NO: 8).

34. 34. The method of claim 33, wherein the base editing enzyme of the first complex provides a mutation at amino acid 141 of HB-EGF (SEQ ID NO: 8).

35. 35. The method of claim 34, wherein the base editing enzyme of the first complex provides a GLU141 to LYS141 mutation in the amino acid sequence of HB-EGF (SEQ ID NO: 8).

36. 36. The method of any one of claims 1 to 35, wherein the base editing enzyme of the first complex provides a mutation in a region of HB-EGF that binds to diphtheria toxin.

37. The method of any one of claims 1 to 36, wherein the base editing enzyme of the first complex provides a mutation in HB-EGF that renders the target cell resistant to diphtheria toxin.

38. 38. The method of any one of claims 1 to 37, wherein the mutation in the target polynucleotide is a point mutation from cytidine (C) to thymine (T) in the target polynucleotide.

39. 38. The method of any one of claims 1 to 37, wherein the mutation in the target polynucleotide is a point mutation from adenine (A) to guanine (G) in the target polynucleotide.

40. The method of any one of claims 1 to 39, wherein the base editing enzyme is introduced into the cell population as a polynucleotide encoding the base editing enzyme.

41. 41. The method of Claim 40, wherein the polynucleotide encoding the base editing enzyme, the first guide polynucleotide of (ii), and the second guide polynucleotide of (iii) are on a single vector.

42. 41. The method of Claim 40, wherein the polynucleotide encoding the base editing enzyme, the (ii) first guide polynucleotide, and the (iii) second guide polynucleotide are on one or more vectors.

43. 43. The method of claim 41 or 42, wherein the vector is a viral vector.

44. 44. The method of claim 43, wherein the viral vector is an adenovirus, a lentivirus, or an adeno-associated virus.

45. 1. A method for providing biallelic integration of a sequence of interest (SOI) into a toxin sensitivity gene (TSG) locus in the genome of a cell, comprising: (a) within the cell population, (i) a nuclease capable of generating a double-stranded break; and (ii) a guide polynucleotide capable of forming a complex with the nuclease and hybridizing to the TSG locus; (iii) (1) a mutation in a 5' homology arm, a 3' homology arm, and the native coding sequence of the TSG, the mutation conferring resistance to the toxin; (2) the SOI; A donor polynucleotide comprising: To introduce introducing (i), (ii), and (iii) such that the donor polynucleotide is integrated into the TSG locus; (b) contacting the cell population with the toxin; (c) selecting one or more cells that are resistant to the toxin; and Including, The one or more cells that are resistant to the toxin comprise the biallelic integration of the SOI.

46. 46. ​​The method of claim 45, wherein the donor polynucleotide is integrated by homology directed repair (HDR).

47. 46. ​​The method of claim 45, wherein the donor polynucleotide is integrated by non-homologous end joining (NHEJ).

48. 48. The method of any one of claims 45 to 47, wherein the TSG locus comprises an intron and an exon.

49. 49. The method of claim 48, wherein the donor polynucleotide further comprises a splicing acceptor sequence.

50. 50. The method of claim 48 or 49, wherein the nuclease capable of generating a double-stranded break generates a break in the intron.

51. 51. The method of any one of claims 48 to 50, wherein the mutation in the native coding sequence of the TSG is in the exon of the TSG locus.

52. 1. A method for integrating a sequence of interest (SOI) into a target locus in the genome of a cell, comprising: (a) within the cell population, (i) a nuclease capable of generating a double-stranded break; and (ii) a guide polynucleotide capable of forming a complex with the nuclease and hybridizing to a toxin sensitivity gene (TSG) locus in the genome of the cell, wherein the TSG is an essential gene; and (iii) (1) a functional TSG gene comprising a mutation in a native coding sequence of the TSG, the mutation conferring resistance to the toxin; and (2) the SOI; (3) a sequence for genomic integration at the target locus; and A donor polynucleotide comprising: To introduce As a result of introducing (i), (ii), and (iii), inactivation of the TSG in the genome of the cell by the nuclease; and introducing, which results in integration of the donor polynucleotide into the target locus; (b) contacting the cell population with the toxin; (c) selecting one or more cells that are resistant to the toxin, selecting the one or more cells that are resistant to the toxin, the one or more cells containing the SOI integrated into the target locus; A method comprising:

53. 53. The method of claim 52, wherein the sequences for genomic integration are obtained from a transposon or a retroviral vector.

54. 54. The method of any one of claims 45 to 53, wherein the functional TSG of the donor polynucleotide is resistant to inactivation by the nuclease.

55. 55. The method of any one of claims 45 to 54, wherein the mutation in the native coding sequence of the TSG removes a protospacer adjacent motif from the native coding sequence.

56. 56. The method of any one of claims 45 to 55, wherein the guide polynucleotide is not capable of hybridizing to the functional TSG of the donor polynucleotide.

57. 57. The method of any one of claims 45 to 56, wherein the nuclease capable of generating a double stranded break is Cas9.

58. 58. The method of claim 57, wherein the Cas9 is capable of generating sticky ends.

59. 59. The method of claim 57 or 58, wherein the Cas9 comprises the polypeptide sequence of SEQ ID NO: 3 or 4.

60. 60. The method of any one of claims 45 to 59, wherein the guide polynucleotide is an RNA polynucleotide.

61. 61. The method of any one of claims 45 to 60, wherein the guide polynucleotide further comprises a tracrRNA sequence.

62. The method of any one of claims 45 to 61, wherein the donor polynucleotide is a vector.

63. 63. The method of any one of claims 45 to 62, wherein the mutation in the native coding sequence of the TSG is a substitution mutation, an insertion or a deletion.

64. 64. The method of any one of claims 45 to 63, wherein the mutation in the native coding sequence of the TSG is a mutation in a toxin-binding domain of a protein encoded by the TSG.

65. 65. The method of any one of claims 45 to 64, wherein the TSG locus comprises a gene encoding heparin-binding EGF-like growth factor (HB-EGF).

66. The method of any one of claims 45 to 65, wherein the TSG encodes HB-EGF (SEQ ID NO: 8).

67. 67. The method of any one of claims 45 to 66, wherein the mutation in the native coding sequence of the TSG is a mutation in one or more of amino acids 107 to 148 of HB-EGF (SEQ ID NO: 8).

68. 68. The method of claim 67, wherein the mutation in the native coding sequence of the TSG is a mutation in one or more of amino acids 138-144 of HB-EGF (SEQ ID NO:8).

69. 69. The method of claim 68, wherein the mutation in the native coding sequence of the TSG is a mutation at amino acid 141 of HB-EGF (SEQ ID NO:8).

70. 70. The method of claim 69, wherein the mutation in the native coding sequence of the TSG is a GLU141 to LYS141 mutation in HB-EGF (SEQ ID NO:8).

71. 71. The method of any one of claims 65 to 70, wherein the toxin is diphtheria toxin.

72. 72. The method of any one of claims 65 to 71, wherein the mutation in the native coding sequence of the TSG renders the cell resistant to diphtheria toxin.

73. 73. The method of any one of claims 45 to 72, wherein the toxin is an antibody-drug conjugate and the TSG encodes a receptor for the antibody-drug conjugate.

74. 1. A method of providing resistance to diphtheria toxin in a human cell, comprising administering to the cell (i) a base editing enzyme; and (ii) a guide polynucleotide that targets a heparin-binding EGF-like growth factor (HB-EGF) receptor in the human cell; and Including the introduction of The base editing enzyme complexes with the guide polynucleotide, and the base editing enzyme is targeted to the HB-EGF and provides diphtheria toxin resistance in the human cell by providing a site-specific mutation in the HB-EGF.

75. 75. The method of claim 74, wherein the base editing enzyme comprises a DNA targeting domain and a DNA editing domain.

76. 76. The method of Claim 75, wherein the DNA targeting domain comprises Cas9.

77. 77. The method of claim 76, wherein the Cas9 comprises a mutation in the catalytic domain.

78. The method of any one of claims 74 to 77, wherein the base editing enzyme comprises a catalytically inactive Cas9 and a DNA editing domain.

79. The method of any one of claims 74 to 77, wherein the base editing enzyme comprises Cas9 (nCas9) capable of generating single-stranded DNA breaks, and a DNA editing domain.

80. 80. The method of claim 79, wherein the nCas9 comprises a mutation relative to wild-type Cas9 at amino acid residue D10 or H840 (numbering is relative to SEQ ID NO: 3).

81. The method of any one of claims 76 to 80, wherein the Cas9 is at least 90% identical to SEQ ID NO: 3 or 4.

82. The method of any one of claims 75 to 81, wherein the DNA editing domain comprises a deaminase.

83. 83. The method of claim 82, wherein the deaminase is selected from cytidine deaminase and adenosine deaminase.

84. 84. The method of claim 83, wherein the deaminase is a cytidine deaminase.

85. 84. The method of claim 83, wherein the deaminase is adenosine deaminase.

86. 86. The method of any one of claims 82 to 85, wherein the deaminase is selected from apolipoprotein B mRNA editing complex (APOBEC) deaminase, activation-induced cytidine deaminase (AID), ACF1 / ASE deaminase, ADAT deaminase, and TadA deaminase.

87. 87. The method of claim 86, wherein the deaminase is an apolipoprotein B mRNA editing complex (APOBEC) family deaminase.

88. 88. The method of claim 87, wherein the cytidine deaminase is APOBEC1.

89. The method of any one of claims 74 to 88, wherein the base editing enzyme further comprises a DNA glycosylase inhibitor domain.

90. 90. The method of claim 89, wherein the DNA glycosylase is a uracil DNA glycosylase inhibitor (UGI).

91. The method of any one of claims 74 to 84 or 86 to 90, wherein the base editing enzyme comprises nCas9 and cytidine deaminase.

92. The method of any one of claims 74 to 83 or 85 to 90, wherein the base editing enzyme comprises nCas9 and adenosine deaminase.

93. The method of any one of claims 74-83 or 86-91, wherein the base editing enzyme comprises a polypeptide sequence that is at least 90% identical to SEQ ID NO:

6.

94. The method of any one of claims 74 to 83 or 86 to 93, wherein the base editing enzyme is BE3.

95. 95. The method of any one of claims 74 to 94, wherein the guide polynucleotide is an RNA polynucleotide.

96. 96. The method of any one of claims 74 to 95, wherein the guide polynucleotide further comprises a tracrRNA sequence.

97. 97. The method of any one of claims 74 to 96, wherein the site-specific mutation is in one or more of amino acids 107 to 148 of the HB-EGF (SEQ ID NO: 8).

98. 98. The method of claim 97, wherein the site-specific mutation is in one or more of amino acids 138-144 of the HB-EGF (SEQ ID NO:8).

99. 99. The method of claim 98, wherein the site-specific mutation is at amino acid 141 of the HB-EGF (SEQ ID NO:8).

100. 100. The method of claim 99, wherein the site-specific mutation is a GLU141 to LYS141 mutation of the HB-EGF (SEQ ID NO: 8).

101. 101. The method of any one of claims 74 to 100, wherein the site-specific mutation is in a region of the HB-EGF that binds to diphtheria toxin.

102. 1. A method for integrating and enriching a sequence of interest (SOI) into a target locus in the genome of a cell, comprising: (a) within the cell population, (i) a nuclease capable of generating a double-stranded break; and (ii) a guide polynucleotide capable of forming a complex with the nuclease and hybridizing to an essential gene (ExG) locus in the genome of the cell; (iii) (1) a functional ExG gene comprising a mutation in a native coding sequence of the ExG, the mutation conferring resistance to inactivation by the guide polynucleotide; and (2) the SOI; (3) a sequence for genomic integration at the target locus; and A donor polynucleotide comprising: To introduce As a result of introducing (i), (ii), and (iii), inactivation of the ExG in the genome of the cell by the nuclease; and resulting in integration of the donor polynucleotide into the target locus; To introduce and (b) cultivating the cells; and (c) selecting one or more surviving cells, selecting said one or more surviving cells, said one or more surviving cells comprising said SOI integrated at said target locus; A method comprising:

103. 1. A method for introducing a stable episomal vector into a cell, comprising: (a) within the cell population, (i) a nuclease capable of generating a double-stranded break; and (ii) a guide polynucleotide capable of forming a complex with the nuclease and hybridizing to an essential gene (ExG) locus in the genome of the cell, a guide polynucleotide, the introduction of which results in the inactivation of the ExG in the genome of the cell by the nuclease; (iii) (1) a functional ExG comprising a mutation in the native coding sequence of the ExG, the mutation conferring resistance to the inactivation by the nuclease; and (2) an autonomous DNA replication sequence; An episomal vector comprising Introducing (b) cultivating the cells; and (c) selecting one or more surviving cells; The one or more viable cells comprise the episomal vector.

104. 104. The method of claim 102 or 103, wherein the mutation in the native coding sequence of the ExG removes a protospacer adjacent motif from the native coding sequence.

105. 105. The method of any one of claims 102 to 104, wherein the guide polynucleotide is not capable of hybridizing to the functional ExG of the donor polynucleotide or the episomal vector.

106. The method of any one of claims 102 to 105, wherein the nuclease capable of generating a double-stranded break is Cas9.

107. 107. The method of claim 106, wherein the Cas9 is capable of generating sticky ends.

108. The method of claim 104 or 107, wherein the Cas9 comprises the polypeptide sequence of SEQ ID NO: 3 or 4.

109. The method of any one of claims 102 to 108, wherein the guide polynucleotide is an RNA polynucleotide.

110. The method of any one of claims 102 to 109, wherein the guide polynucleotide further comprises a tracrRNA sequence.

111. The method of any one of claims 102 to 110, wherein the donor polynucleotide is a vector.

112. 112. The method of any one of claims 102 to 111, wherein the mutation in the native coding sequence of the ExG is a substitution mutation, an insertion, or a deletion.

113. 113. The method of any one of claims 102 or 104 to 112, wherein the sequences for genomic integration are obtained from a transposon or a retroviral vector.

114. The method of any one of claims 103 to 112, wherein the episomal vector is an artificial chromosome or a plasmid.

115. 115. The method of any one of claims 102-114, wherein two or more guide polynucleotides are introduced into the cell population, each guide polynucleotide forming a complex with the nuclease, and each guide polynucleotide hybridizing to a different region of the ExG.

116. 116. The method of any one of claims 102, 104-113, or 115, further comprising introducing the nuclease of (a)(i) and the guide polynucleotide of (a)(ii) into living cells to enrich for living cells that contain the SOI integrated into the target locus.

117. 116. The method of any one of claims 103-112, 114, or 115, further comprising introducing the nuclease of (a)(i) and the guide polynucleotide of (a)(ii) into viable cells to enrich for viable cells containing the episomal vector.

118. 118. The method of claim 116 or 117, wherein the nuclease of (a)(i) and the guide polynucleotide of (a)(ii) are introduced into the viable cell for multiple rounds of enrichment.