Synthetic intron screening system, components thereof, and methods of using same to enrich for base editing activity

WO2026106966A1PCT designated stage Publication Date: 2026-05-21THE BROAD INST INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
THE BROAD INST INC
Filing Date
2025-11-11
Publication Date
2026-05-21

Smart Images

  • Figure US2025054980_21052026_PF_FP_ABST
    Figure US2025054980_21052026_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are polynucleotides, vectors, complexes, compositions, systems, kits, methods and uses for enriching for gene editing in a cell. Some aspects of the disclosure relate to the use of a polynucleotide cassette comprising a coding sequence of a selection marker gene that is disrupted by a synthetic intron sequence to prevent gene expression of the selection marker. The synthetic intron further comprises a defective splice site that is correctable by a base edit such that when the splice site is corrected by a gene editor, the intron is removed by the endogenous slicing system of the cell and expression of the selection marker gene can occur. Methods described herein screen for the presence of active gene editors within the cell by subjecting the cell to the selection pressure of the selection marker.
Need to check novelty before this filing date? Find Prior Art

Description

SYNTHETIC INTRON SCREENING SYSTEM, COMPONENTS THEREOF, AND METHODS OF USING SAME TO ENRICH FOR BASE EDITING ACTIVITYRELATED APPLICATIONS

[0001] This application claims priority under 35 U.S.C. § 119(e) to U.S. Provisional Patent Application, U.S.S.N. 63 / 719,638, filed November 12, 2024, and U.S. Provisional Patent Application, U.S.S.N. 63 / 868,864, filed August 22, 2025, each of which is incorporated herein by reference.REFERENCE TO AN ELECTRONIC SEQUENCE LISTING

[0002] The contents of the electronic sequence listing (B119570207WO00-SEQ-GJM.xml; Size: 107,109 bytes; and Date of Creation: November 11, 2025) are herein incorporated by reference in their entirety.BACKGROUND

[0003] The functional characterization of genetic variants proves a continuing challenge in the genomics field. As of October 2024, over half of annotated mutations (~1.6 million) in ClinVar are classified as “variants of unknown significance” (Henrie et al., 2018). A variety of technologies have been deployed to address this problem, starting with deep mutational scanning (DMS) (Araya et al., 2011). This approach expresses an exogenous open reading frame containing the mutated gene of interest, allowing for interrogation of precise mutations. However, by removing genes from their native context, DMS changes the transcriptional profile, and thereby may be changing the activity of the protein. Another approach for functional characterization of genetic variants is saturation genome editing (SGE). SGE addresses the concern posed by DMS by permitting precision nucleotide changes in the endogenous context (Findlay et al., 2014). SGE employs the homology-directed repair pathway in combination with oligonucleotides containing the edit of interest. Due to challenges with efficiency, this technology is generally limited to use in haploid cell models.

[0004] More recent technologies, such as CRISPR knockout (CRISPRko) have allowed for high-throughput interrogation of the functionally important protein regions (Shi et al., 2015; He et al., 2019; Munoz et al., 2016; Schoenenberg et al., 2018; Herman et al. 2022).However, approximately a third of the signal from such screens is reported to generate the inframe mutations. Out-of-frame mutations do not permit evaluating the local effects of the perturbation, since they affect folding regardless of the functional significance of the targeted region. As a more precise high-throughput alternative, base editing can install A>G or C>TB1195.70207WO00 1 / 102edits within an approximate four nucleotide-long window (Komor et al., 2016; Gaudelli et al., 2017; Neugebauer et al., 2023). Most recently, CRISPR prime editing has been developed to install precise mutations-insertions, deletions, or single-base pair changes, but at lower efficiency than base editors (Anzalone et al., 2019). Base editing has been used for a variety of genome scanning purposes, including interrogating post-translational modifications, inhibition mechanisms, and non-coding regulatory elements (Hanna et al., 2021; Perner et al., 2023; Martin-Rufino et al., 2023; Lue et al., 2023; Li et al., 2023; Kennedy et al., 2024).

[0005] One current limitation of precision editing systems — including base editing — in the context of functional characterization of genetic variants is cell-to-cell editing variability. This cell-to-cell range of activity may be due to the semi-random nature of lentiviral integration (Serrao et al., 2016; Shao et al., 2022), transgene silencing (Cabrera et al., 2022), or the cellular immune response to Cas9 (Chew et al., 2016). Traditional indirect selection methods fail to account for this variation, increasing noise and decreasing the ability to detect biologically relevant areas of interest in high-throughput endogenous screens.SUMMARY

[0006] The present invention stems from the recognition that gene editing activity (e.g., by any nucleic acid programmable gene editor) varies depending on the cell being edited. Thus, the present disclosure describes polynucleotides, vectors, complexes, compositions, systems, kits, methods, and uses of novel synthetic introns to enrich for cells in which gene editing is active. In particular, the present disclosure describes a polynucleotide cassette comprising a coding sequence of a selection marker gene (e.g., antibiotic resistance gene) that is disrupted by a synthetic intron sequence that comprises a defective 5'-splice donor site that is correctable by a single base edit (e.g., by a base editor). The defective 5'-splice donor site prevents a cell’s RNA splicing machinery from recognizing and removing the intron from the selection marker gene; however, when the 5 '-splice donor site is corrected by a gene editor (e.g., a base editor or prime editor), the intron is removed by the endogenous slicing system of the cell, and the selection marker gene is restored. This relationship is leveraged to identify the presence of active gene editors in the cell by subjecting the cell to the selection pressure of the selection marker. Cell survival indicates the desired base edit was incorporated at the 5 '-splice donor site of the synthetic intron, therefore indicating the cells contain active gene editors.

[0007] The present disclosure further contemplates the use of the activity-based selection methods described herein for high-throughput mutational screens of disease-associated genesB1195.70207WO00 2 / 102(e.g., cancer genes) to identify regions of interest for further interrogation, and / or to observe improved identification of known relevant residues and domains in said genes, as well as possible novel gain- and loss-of-function mutations.

[0008] In one aspect, the present disclosure describes polynucleotides comprising a nucleotide sequence encoding a synthetic intron. In certain embodiments, the present disclosure describes polynucleotides comprising a nucleotide sequence encoding a synthetic intron having at least 80% sequence identity to a nucleic acid sequence set forth in SEQ ID NO: 31, wherein the synthetic intron comprises a defective 5'-splice donor site which is correctable by a gene editor. In certain embodiments, the synthetic intron further comprises a protospacer adjacent motif (PAM) sequence having the sequence NGN, wherein N is A, G, T, or C. The synthetic intron may further comprise at least one, at least two, or at least three stop codons, such that the intron contains a stop codon in every reading frame. The disclosure further contemplates synthetic introns derived from the SV40 intron (SEQ ID NO: 31). In certain embodiments, the synthetic intron comprises nucleic acid sequence set forth in SEQ ID NOs.: 12-22.

[0009] In certain embodiments, the polynucleotides comprising the synthetic intron disclosed herein contain a defective 5'-splice donor site that comprises an “AT” or “GC” nucleotide sequence in place of the canonical “GT” nucleotide sequence of SEQ ID NO: 31. The canonical “GT” sequence are the first two nucleotides of the intron. In certain embodiments, the defective 5 '-splice donor site is corrected back to the canonical GT sequence of SEQ ID NO: 31 by a gene editor, for example, an adenine or cytidine base editor.

[0010] In another aspect, the present disclosure describes splice-intron guide RNAs (sigRNA, also referred to herein as “splice guide RNA”). As referred to herein, a sigRNA is a guide RNA suitable for base editing or prime editing, that targets a portion of a synthetic intron. In certain embodiments, the sigRNA comprises a first portion comprising a region of complementarity to a synthetic intron, and a second portion comprising a trans-activating CRISPR RNA(tracrRNA). In certain embodiments, the tracrRNA facilitates the binding of the splice-intron gRNA to a napDNAbp, for example, a Cas9 protein (e.g., a Cas9 protein as part of a base editor). In certain embodiments, the sigRNA comprises a sequence set forth in any one of SEQ ID NOs: 1-11. In certain embodiments, the spacer sequence of the sigRNA comprises sequence set forth in any one of SEQ ID NOs: 1-11.

[0011] In another aspect, the present disclosure describes a complex comprising: (a) a gene editor comprising a DNA binding protein, and (b) a sigRNA disclosed herein. The present disclosure contemplates the use of the sigRNA to direct (e.g., target) a gene editor to theB1195.70207WO00 3 / 102synthetic intron within the polynucleotide cassettes described herein, specifically to the defective 5'-splice donor site, to allow for gene editing. In certain embodiments, the gene editor is a base editor. In certain embodiments, the base editor is an adenine base editor. In certain embodiments, the base editor is a cytidine base editor. In certain embodiments, the defective 5'-splice donor site is restored (e.g., edited) by making a single nucleobase edit within the 5 '-splice donor site. In certain embodiments, an adenine base editor corrects an “AT” sequence within the 5 '-splice donor site to a “GT” sequence. In certain embodiments, a cytidine base editor corrects a “GC” sequence within the 5'-splice donor site to a “GT” sequence.

[0012] In a further aspect, the present disclosure describes polynucleotide cassettes comprising a nucleotide sequence encoding a selection marker disrupted by a synthetic intron with a defective 5' donor splice site that is correctable by a gene editor. In certain embodiments, the selection marker is an antibiotic resistance gene (e.g., a puromycin resistance gene). In certain embodiments, the presence of the synthetic intron within the selection marker prevents translation of a functional selection marker protein (e.g., a puromycin resistance protein).

[0013] In certain embodiments, the present disclosure presents polynucleotide cassettes comprising a coding sequence of a selection marker gene (e.g., antibiotic resistance gene) that is disrupted by a synthetic intron sequence that comprises a defective 5 '-splice donor site that is correctable by a base edit (e.g., by a base editor or prime editor). The defective 5'-splice donor site prevents a cell’s RNA splicing machinery from recognizing and removing the intron from the selection marker gene. When the 5 '-splice donor site is corrected by a gene editor (e.g., a base editor or prime editor), the intron is removed by the endogenous slicing system of the cell, and the selection marker gene is restored. The presence of active gene editors within a cell is then identified by subjecting the cell to the selection pressure of the selection marker. Without wishing to be bound by any particular theory, if the desired base edit was incorporated at the 5 '-splice donor site of the synthetic intron, the cell will survive the selection pressure, thus indicating the cell contains active gene editors.

[0014] In some aspects, the disclosure provides nucleic acids encoding the polynucleotides described herein, the sigRNAs described herein, and / or the complexes described herein. In another aspect, the disclosure provides expression vectors comprising such nucleic acids. In yet another aspect, the disclosure provides cells (e.g., bacterial cells) that comprise a polynucleotide described herein, a complex described herein, a nucleic acid described herein, and / or a vector described herein.B1195.70207WO00 4 / 102

[0015] In a further aspect, the present disclosure describes compositions of nucleic acids comprising: (a) a first nucleic acid encoding a selection marker disrupted by a synthetic intron, wherein the synthetic intron comprises a defective 5' donor splice site that is correctable by a gene editor; (b) a second nucleic acid encoding a splice-intron guide RNA (sigRNA) comprising a spacer sequence that has complementarity to the defective 5' donor splice site of the synthetic intron of (a), and a trans-activating CRISPR RNA (tracrRNA); and / or (c) a third nucleic acid encoding a guide RNA (gRNA) comprising a spacer sequence comprising a region of complementarity to a target gene of interest; and / or (d) a fourth nucleic acid encoding a gene expression marker; and / or (e) a fifth nucleic acid encoding a gene editor.

[0016] In another aspect, the present disclosure provides systems for enriching for gene editing activity in a cell comprising: (a) a nucleic acid encoding a selection marker disrupted by a synthetic intron, wherein the synthetic intron comprises a defective 5' donor splice site that is correctable by a gene editor; (b) a splice-intron guide RNA (sigRNA), or a nucleic acid encoding a splice-intron guide RNA, comprising a spacer sequence that has complementarity to the defective 5' donor splice site of the synthetic intron of (a), and a trans-activating CRISPR RNA(tracrRNA); (c) a gene editor or a nucleic acid encoding a gene editor; and (d) a guide RNA, or a nucleic acid encoding a guide RNA, comprising a spacer sequence comprising a region of complementarity to a target gene of interest, wherein gene editingdependent correction of the defective 5' donor splice site results in the enrichment of editing of the target gene of interest. In certain embodiments, the nucleic acid of (a), the nucleic acid of (b), the nucleic acid of (c), and / or the nucleic acid of (d) are located on the same nucleic acid construct or different nucleic acid constructs. In certain embodiments, the gene editor is a nucleic acid programable gene editor. Without wishing to be bound by any particular theory, the present disclosure contemplates that if the gene editor successfully edits the 5'-splice donor site of the synthetic intron, that the gene editor will also successfully edit the target gene corresponding to the gRNA of (d). In some embodiments, the system is used for profiling gene editing activity in a cell.

[0017] In yet another aspect, the disclosure provides a kit comprising: (i) the polynucleotides described above, the sigRNAs described above, the complexes described above, the nucleic acids described above, the vectors described above, the cells described above, the compositions described above, or the systems described above; and (ii) a set of instructions for enriching for gene editing in a cell.B1195.70207WO00 5 / 102

[0018] In a further aspect, the present disclosure provides methods for enriching for gene editing activity in a cell. In certain embodiments, the method comprises: (a) introducing into a cell: (i) a nucleic acid encoding a selection marker disrupted by a synthetic intron with a defective 5' donor splice site that is correctable by a nucleic acid programmable gene editor, (ii) a sigRNA comprising a spacer sequence that has complementarity to the defective 5' donor splice site of the synthetic intron of (i), and a trans-activating CRISPR RNA(tracrRNA), (iii) a gene editor, and / or (iv) a guide RNA comprising a spacer sequence comprising a region of complementarity to a target gene of interest (that is not the 5 '-splice donor site); (b) subjecting the cell to the selection pressure of the selection marker of (i), wherein cell survival indicates the gene editor made the desired edit at the 5 '-splice donor site; and / or (c) sequencing the target DNA sequence to confirm presence of edit.

[0019] In another embodiment, the present disclosure presents the use of the compositions and / or systems described herein for gene editing a target nucleobase within a synthetic intron, wherein editing the target nucleobase restores the 5 '-splice donor site of the synthetic intron. In other embodiments, the uses described herein are for modifying a target nucleic acid sequence for research purposes (e.g., measuring the activity of a base editor in a cell).

[0020] The foregoing concepts, and additional concepts discussed below, may be arranged in any suitable combination, as the present disclosure is not limited in this respect. Further, other advantages and novel features of the present disclosure will become apparent from the following detailed description of various non-limiting embodiments when considered in conjunction with the accompanying figures.BRIEF DESCRIPTION OF DRAWINGS

[0021] FIGs. 1A-1G show Cas9-NG base editor tiling of TP53 and CBE optimization. FIG.1A shows a schematic of a Cas9-NG base editor tiling screen (top) and a schematic of Cas9-NG base editor tiling screens targeting TP53, with either ABE8e (ABE) or rat APOBEC (CBE) deaminases (bottom). FIG. IB shows the mean z-score of drug arms relative to dropout along length of TP53 by median amino acid residue for adenine (left) and cytosine (right) base editors. FIGs. 1C and ID show Sanger sequencing-derived change in abundance from day 7 to day 21 of the edited TP53 allele targeted by sigRNA selected for validation. The most frequent ABE (FIG. 1C) or CBE (FIG. ID) amino acid mutation is shown. FIG. IE shows a Western blot validation of sgTP53 1, 2, and 3, showing p53 versus A40p53 levels. Vinculin used as loading control. D = day sample was taken; N = Nutlin treatment; * = non-specific band of the anti-p53 antibody.FIG. IF shows a GFP activity assay evaluatingB1195.70207WO00 6 / 102activity differences between Cas9-SpG and NG base editors in both A549 and MelJuSo cells as observed by flow cytometry. FIG. 1G shows a GFP activity assay benchmarking common Cas9-SpG CBEs in both A549 and MelJuSo cells as observed by flow cytometry.

[0022] FIGs. 2A-2F show Cas9-NG TP53 screen quality and validation. FIG. 2A shows replicate arm correlations of ABE (top) and CBE (bottom) arms across all drug conditions. Pearson correlation reported. FIGs. 2B and 2C show scatterplots depicting sigRNA distribution between drug arms for ABE (FIG. 2B) and CBE (FIG. 2C) arms, highlighting sigRNAs nominated for validation (squares). FIG. 2D shows scatterplots depicting sgRNA z-scores in etoposide versus Nutlin arms for ABE (left) and CBE (right) screens. The sgRNAs nominated for validation are indicated by squares. FIG. 2E shows annotated TP53 domains. Domain start and end positions are labeled. TAD = transactivation domains, OD = oligomerization domain. FIG. 2F shows validation of target allele abundance of sgTP53 1, 2, and 3. The sgRNAs were transduced into cells expressing Cas9-NG-ABE and treated with Nutlin starting on Day 7. Allelic fractions at each time point were determined by PCR of the target site and Illumina sequencing. LFC in abundance from Day 7 to Day 21 is shown and color coded. Wild type sequence is shown in bold. Mutated amino acids are underlined. Percent reads are normalized to total reads with >2% frequency.

[0023] FIGs. 3A-3K show the activity-based selection method development. FIG. 3A shows a schematic depicting base editor activity-based selection screen design. To the right, it is shown that base editors can restore the intron splice donor sequence to the recognized mammalian consensus sequence, allowing for splicing and functional protein expression. If no base editing occurs, lack of a recognized splice site and thus splicing event prevents production of a functional protein. To the left, it is shown that base editors can disrupt the endogenous CD274 splice site, knocking out CD274 expression. FIG. 3B displays a vector schematic showing splice guide variation and corresponding splice target variation within primary screen libraries. Underlined nucleotides can be base edited to the splice donor consensus sequence. Nucleotides in bold are part of the consensus sequence. FIG. 3C shows timelines of ABE (left) and CBE (right) primary screens. Screens followed the same timeline beginning with the splice target library transduction. FIGs. 3D and 3E show scatterplots of guide-level z-scores ranked by average guide z-score for CBE (FIG. 3E) and ABE (FIG. 3D) screens. Log-fold changes were calculated relative to pDNA. The black line represents the average guide z-score. FIGs. 3F and 3G show validation of selected primary screen guides. Bars represent CD274 depletion as measured by APC positivity using presence-based selection (white bars) and activity-based selection (stippled bars) across CBE (FIG. 3G) andB1195.70207WO00 7 / 102ABE (FIG. 3F) modalities. FIG. 3H shows vector schematics showing variation in introns (top) and corresponding splice-targeting sgRNAs (bottom). Underlined nucleotides can be base edited to the splice donor consensus sequence (GT). Nucleotides in in italics are part of the consensus sequence. The blue box depicts nucleotides within the puromycin resistance exon. The pink box depicts nucleotides within the intron. PAM sequence is shown in bold. FIG. 31 shows distribution of maximum replicate sgRNA z-scores for ABE (left) and CBE (right) screens. Horizontal dashed line represents validation cutoff of z>=1.96. Squares represent sgRNAs nominated for validation. FIG. 3J is a schematic of splice-targeting sgRNA validation timeline. FIG. 3K includes exemplary flow cytometry plots of splicetargeting sgRNA validation for ABE (top) and CBE (bottom) sgRNAs. CD274 editing was measured by staining with an APC antibody. Both presence- and activity-based selection conditions were gated to include only GFP+ (transduced) cells. The percentage indicates the percent of APC+ (CD274 unedited) cells.

[0024] FIGs. 4A-4J show the activity-based selection method development and extension.FIG. 4A shows replicate correlations between the splice guide library (SGL) and the splice target library (STL) replicates for ABEs (right) and CBEs (left). Note that the ABE SGL Rep A + Intron Rep B sample was lost to contamination. FIGs. 4B and 4C show Pearson correlations between replicates from CBE (FIG. 4B) and ABE (FIG. 4C) arms. Gray dots represent sigRNAs nominated for validation. Dotted line represents the z-score cutoff of 1.96.FIG. 4D shows the percent base editing by nucleotide as determined by Sanger sequencing of targeted CD274 locus. Only the within-window edits are shown. FIG. 4E shows a GFP-on activity assay to evaluate possible endogenous splicing by comparing the GFP-intron construct in cells with and without base editors. FIG. 4F shows the editing efficiency of presence- versus activity-based selection conditions with an intron inserted in the +1 reading frame of GFP. FIG. 4G shows the editing efficiency of presence- versus activity-based selection conditions with an intron inserted in the +0 reading frame of hygromycin resistance.FIG. 4H shows the editing efficiency of presence- versus activity-based selection conditions with a puromycin resistance intron and SpRY-Cas9 base editors. FIG. 41 shows SpRY-Cas9 versus SpG-Cas9 self-editing at integrated sigRNA locus. Two splice-targeting sgRNAs were used for both ABE and CBE, and the average sgCD274 self-editing is plotted. Individual sgCD274 self-editing rates are shown as dots. This only considers editing within three base pairs of the standard base editing window. FIG. 4J shows a representative example of flow cytometry gating strategy shown with ABE sgRNA 1 activity-based selection condition.B1195.70207WO00 8 / 102Populations were gated for live cells, then single cells, then transduced cells via GFP-FITC signal, and finally editing efficiency of CD274 was determined via APC signal.

[0025] FIGs. 5A-5F show the activity-based selection base editor tiling of TP53. FIG. 5A shows a schematic of the Cas9-SpG base editor tiling screen. FIG. 5B shows a basic depiction of selection method constructs. The presence-based selection construct (top) contains an intact puromycin resistance gene. The activity-based selection construct (bottom) contains an intron-disrupted puromycin resistance gene and a splice guide targeting the mutated 5'-splice donor site of the intron. FIG. 5C shows self-editing analysis: percent edited reads by position along TP53 sigRNA. Boxes depict 25th and 75th percentiles as minima and maxima and the center represents the mean; whiskers depict 10th and 90th percentile. FIG.5D shows average etoposide z-scores relative to the no-drug arm by BE and selection condition. Guides are colored by predicted mutation and plotted by the median targeted residue. Guides targeting coding regions of TP53 are depicted. FIGs. 5E and 5F show average etoposide z-scores relative to no-drug arm of all ABE (FIG. 5E) and CBE (FIG. 5F) TP53-targeting mutations separated by predicted mutation. Boxes depict 25th and 75th percentiles as minima and maxima and the center represents the mean; whiskers depict 10th and 90th percentile. Significance is calculated between selection methods within each mutation type using the Mann-Whitney- Wilcoxon two-sided test. Stars depict p-value significance according to the following: p-value <= 0.001: ***, <= 0.01: **, <= 0.05: *, >0.05: n.s.

[0026] FIGs. 6A-6K show the activity-based selection base editor tiling of TP53. FIG. 6A shows the NGS read match rate with expected guides after deconvolution with PoolQ. FIG.6B shows the potential self-editing mechanism. Rather than identifying and editing the endogenous target, the sigRNA may bind to itself by forming a one-nucleotide bulge to recognize the NGN PAM using the first G of the tracer sequence. FIG. 6C shows the selfediting rate by selection method and base editor. Shown in gray is the baseline pDNA selfediting rate with no editors. FIG. 6D shows the screen replicate Pearson correlations before and after applying a self-editing in-silico correction. FIG. 6E shows ROC / AUC of TP53 screen control sigRNA. True positives are splice-site control sigRNAs. False positives are intergenic control sigRNAs. FIG. 6F shows sgRNA LFC distribution graphs for ABE (left) and CBE (right) with presence-based selection (top) and activity-based selection (bottom) conditions. FIGs. 6G and 6H show the sgRNA distribution between selection conditions. Pearson correlations are as follows: r(ABE)=0.88, r(CBE)=0.89. Dashed line indicates y = x.FIG. 61 shows prime editing rate at HEK3 locus by selection method and prime editor. TheB1195.70207WO00 9 / 102HEK3 epegRNA was used in all vectors. Three puromycin resistance-targeting epegRNAs were designed and tested for activity-based selection conditions. FIG. 6 J shows self-editing rate by tracrRNA of presence- and activity-based selection methods across two sgRNAs via Sanger sequencing of the sgRNA cassette. Bars depict the average of two replicates, individually shown as dots. FIG. 6K shows A>G editing efficiency by tracrRNA of presence- versus activity-based selection methods across two sgRNAs via Illumina sequencing of the TP53 target loci. Bars depict the average of two replicates, individually shown as dots.

[0027] FIG. 7 shows how the heterogeneity of base editing activity impacts screens.

[0028] FIGs. 8A-8D show TP53 screen validation. FIG. 8A shows difference in means (Ax ) between z-scores within annotated domain versus interdomain regions of TP53 by selection method and technology. Domain regions are outlined in FIG. 2E. Black dots depict the mean z-score. FIG. 8B shows the fraction of guides predicted to introduce ClinVar pathogenic TP53 variants whose z-score (etoposide arm) falls below the indicated threshold as compared to guides that only introduce benign variants. Analysis is repeated with z-scores from the use of activity-based and presence-based selection to support interpretation of effect size with the use of each approach. FIG. 8C shows ROC-AUC analysis to characterize the separation of ClinVar pathogenic and benign variants at a range of effect sizes in deep mutational scanning (DMS), prime editing (PE), and base editing (BE) TP53 screens. The table reports the counts of variants considered in the analysis. PE guides are filtered by efficiency of editing at the sensor locus at various thresholds. FIG. 8D shows validation of target allele abundance of sgTP534 and 5 with activity-based selection vectors. The sgRNAs were transduced into cells expressing SpG-Cas9- ABE and treated with etoposide starting on Day 11. Allelic fractions at each time point were determined by PCR of the target site and Illumina sequencing. Log fold changes (LFC) in abundance from Day 11 to Day 26 is shown and color coded. The wild type sequence is shown in bold. Mutated amino acids are underlined. Percent reads are normalized to total reads with >1.5% frequency. Lowercase letters represent nucleotides depicting splice acceptor site disruption. Uppercase letters represent amino acids.

[0029] FIGs. 9A-9F show identification of guide features that influence on-target activity in TP53 tiling screen etoposide arm. FIG. 9A shows Rule Set 3 Sequence score of TP53 tiling screen positive control guides, namely those that introduce missense, nonsense, or splice-site mutations not previously identified as benign, discretized by activity. The results of this analysis are presented separately with the use of a z-score cutoff of -2, -3, or -4 to define guides as active. Significance is calculated between active and inactive guides using theB1195.70207WO00 10 / 102Mann-Whitney one-sided test. This analysis incorporates data with the use of both the ABE8e and TadCBEd editor. FIG. 9B shows the same as FIG. 9A, but with the use of the BE-Hive model to predict base editing guide efficacy instead of Rule Set 3 Sequence score. FIG. 9C shows the same as FIG. 9A, but with the use of the FORECasT-BE model to predict base editing guide efficacy with ABE8e. FIG. 9D shows the same as FIG. 9C, but estimating model performance at predicting base editing guide activity with TadCBEd. FIG. 9E shows sequence motifs representing nucleotide identities at each position relative to the editable nucleotide that are enriched among active positive control guides. Active guides are defined as those with a z-score below -2, whereas inactive guides represent those with a z-score between -2 and 2. Analysis is stratified by base editor. FIG. 9F shows sequence motifs representing nucleotide identities at each position relative to the editable nucleotide that are enriched among active positive control guides. Active guides are defined as those with a z-score below -4, whereas inactive guides represent those with a z-score between -1 and 1.DEFINITIONS

[0030] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this invention belongs. The following references provide one of skill with a general definition of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them unless specified otherwise.

[0031] The term “amino acid” refers to organic compounds that serve as the building blocks of proteins. Twenty amino acids play essential roles in protein synthesis and function and differ based on their side chain structures, which contribute to their distinct properties and functions. Amino acids are often grouped by the chemistry of the side chain. These groups are polar-uncharged, polar-charged and non-polar. Eight of the 20 amino acids are polar-uncharged: asparagine (N), cysteine (C), glutamine (Q), histidine (H), serine (S), threonine (T), tryptophan (W) and tyrosine (Y). Eight of the 20 amino acids are non-polar: alanine (A), glycine (G), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), proline (P), and valine (V). The final four are polar-charged: arginine (R), aspartic acid (D), glutamic acid (E), and lysine (L).B1195.70207WO00 11 / 102

[0032] The terms “administer,” “administering,” and “administration” refer to implanting, absorbing, ingesting, injecting, inhaling, or otherwise introducing a treatment or therapeutic agent, or a composition of treatments or therapeutic agents, in or on a cell or subject.

[0033] The terms “composition” and “formulation” are used interchangeably.

[0034] The term “biomolecule” or “biological molecule” refers to any substance produced by cells or living organisms and includes carbohydrates, lipids, nucleic acids, proteins, and vitamins.

[0035] The terms “condition,” “disease,” and “disorder” are used interchangeably.

[0036] A “cell,” as used herein, may be present in a population of cells (e.g., in a tissue, a sample, a biopsy, an organ, or an organoid). In some embodiments, a population of cells is composed of a plurality of different cell types. Cells for use in the methods and systems of the present disclosure can be present within an organism, a single cell type derived from an organism, or a mixture of cell types. Included are naturally occurring cells and cell populations, genetically engineered cell lines, cells derived from transgenic animals, cells from a subject, etc. In some embodiments, the cells are mammalian cells (e.g., complex cell populations such as naturally occurring tissues). In some embodiments, the cells are from a human. In some embodiments, the cells are collected from a subject (e.g., a human) through a medical procedure, such as a biopsy. Alternatively, the cells may be a cultured population (e.g., a culture derived from a complex population, or a culture derived from a single cell type where the cells have differentiated into multiple lineages). The cells may also be provided in situ in a tissue sample.

[0037] As used herein, “complementarity” refers to the ability of a nucleic acid to form hydrogen bond(s) with another nucleic acid by traditional Watson-Crick base-pairing. A percent complementarity indicates the percentage of residues in a nucleic acid molecule which can form hydrogen bonds (i.e., Watson-Crick base pairing) with a second nucleic acid (e.g., about 5, 6, 7, 8, 9, 10 out of 10, being about 50%, 60%, 70%, 80%, 90%, and 100% complementary respectively). “Perfectly complementary” means that all the contiguous residues of a nucleic acid sequence form hydrogen bonds with the same number of contiguous residues in a second nucleic acid sequence. “Substantially complementary” as used herein refers to a degree of complementarity that is at least about any one of 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of about 40, 50, 60, 70, 80, 100, 150, 200, 250, or more nucleotides, or refers to two nucleic acids that hybridize under stringent conditions.B1195.70207WO00 12 / 102

[0038] In the context of the present application, “region of complementarity” refers to a nucleic acid sequence to which a gRNA sequence is designed to have perfect complementarity or substantial complementarity, and hybridization between the target sequence (e.g., 5'-splice donor site) and the gRNA forms a double stranded RNA (dsRNA) region containing a target base (e.g., an adenosine or cytidine), which recruits an base editor that deaminates the target base.

[0039] The term “deaminase” or “deaminase domain,” as used herein, refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase, catalyzing the hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase domain, catalyzing the hydrolytic deamination of cytosine to uracil. In some embodiments, the deaminase is an adenine deaminase, which catalyzes the hydrolytic deamination of adenine or adenine. In some embodiments, the deaminase or deaminase domain is an adenine deaminase, catalyzing the hydrolytic deamination of adenosine or deoxy adenosine to inosine or deoxyinosine, respectively. In some embodiments, the adenine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA).

[0040] The adenine deaminases (e.g. engineered adenine deaminases, evolved adenine deaminases) or cytidine deaminases (e.g. engineered cytidine deaminases, evolved cytidine deaminases) provided herein may be from any organism, such as a bacterium. In some embodiments, the deaminase or deaminase domain is a variant of a naturally-occurring deaminase from an organism that does not occur in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75% at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally-occurring deaminase from an organism.

[0041] The term “gene editing” refers to a type of genetic engineering in which DNA is inserted, deleted, modified, or replaced in the genome of a living organism by a “gene editor.” Accordingly, as referred to herein, the term “gene editor” may refer to any known gene editing tool, such as CRISPR / Cas9 technology (Clustered regularly interspaced short palindromic repeats), TALENs (Transcription activator-like effector nucleases) and ZFNs (Zinc finger nucleases). In certain embodiments, the gene editor is a base editor or prime editor. In certain embodiments, the base editor is an adenine base editor. In certain embodiments, the base editor is a cytidine base editor.B1195.70207WO00 13 / 102

[0042] The term “base editor (BE)” or “nucleobase editor (NBE)” refers to an agent comprising a polypeptide that is capable of making a modification to a base (e.g., A, T, C, G, or U) within a nucleic acid sequence (e.g., DNA or RNA). In some embodiments, the base editor is capable of deaminating a base within a nucleic acid. In some embodiments, the base editor is capable of deaminating a base within a DNA molecule. In some embodiments, the base editor is capable of deaminating an adenine (A) in DNA. In some embodiments, the base editor is a fusion protein comprising a nucleic acid programmable DNA binding protein (napDNAbp) fused to an adenine deaminase. In some embodiments, the base editor is a Cas9 protein fused to an adenine deaminase. In some embodiments, the base editor is a Cas9 nickase (nCas9) fused to an adenine deaminase. In some embodiments, the base editor is a nuclease-inactive Cas9 (dCas9) fused to an adenine deaminase. In some embodiments, the base editor is a Cas9 protein fused to a cytidine deaminase. In some embodiments, the base editor is a Cas9 nickase (nCas9) fused to a cytidine deaminase. In some embodiments, the base editor is a nuclease-inactive Cas9 (dCas9) fused to a cytidine deaminase. In some embodiments, the base editor is fused to an inhibitor of base excision repair, for example, a UGI domain, or a dISN domain. In some embodiments, the fusion protein comprises a Cas9 nickase fused to a deaminase and an inhibitor of base excision repair, such as a UGI or dISN domain. In some embodiments, the dCas9 domain of the fusion protein comprises one or more mutations, which inactivates the nuclease activity of the Cas9 protein. In some embodiments, the fusion protein comprises one or more mutations, which renders Cas9 capable of cleaving only one strand of a nucleic acid duplex. In some embodiments, the adenine base editor is selected from TadA-8e, ABE8.0, ABE8e, AYBE, ABE9, and variants thereof. In some embodiments, the cytidine base editor is selected from evoCDAmax, CBE6, CGBE, BE4max, TadCBE, and variants thereof.

[0043] As used herein, the term “prime editing” refers to an approach for gene editing using napDNAbps, a polymerase (e.g., a reverse transcriptase), and specialized guide RNAs that include a DNA synthesis template for encoding desired new genetic information (or deleting genetic information) that is then incorporated into a target DNA sequence. Classical prime editing is described in Anzalone et al. Search-and-replace genome editing without doublestrand breaks or donor DNA. Nature 576, 149-157 (2019), which is incorporated herein by reference.

[0044] The term “prime editor” or “PE” refers to fusion constructs comprising a napDNAbp (e.g., Cas9 nickase) and a reverse transcriptase that is capable of carrying out prime editing on a target nucleotide sequence in the presence of a pegRNA (or “extended guide RNA”). InB1195.70207WO00 14 / 102some embodiments, simple edits (e.g., modifying the anticodon sequence) may be performed using standard prime editing (e.g., PE2 / PE3) machinery. General disclosure of prime editing are described in Chen et al., “Prime editing for precise and highly versatile genome manipulation,” Nature Reviews Genetics volume 24, pages 161-177 (2023), which is incorporated herein by reference.

[0045] The term “Cas9” or “Cas9 nuclease” refers to an RNA-guided nuclease comprising a Cas9 domain, or a fragment thereof (e.g., a protein comprising an active or inactive DNA cleavage domain of Cas9, and / or the gRNA binding domain of Cas9). A “Cas9 domain,” as used herein, is a protein fragment comprising an active or fully or partly inactive cleavage domain of Cas9 and / or the gRNA binding domain of Cas9. A “Cas9 protein” is a full length Cas9 protein. A Cas9 nuclease is also referred to sometimes as a casnl nuclease or a CRISPR (Clustered Regularly Interspaced Short Palindromic Repeat) -associated nuclease. CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain spacers, sequences complementary to antecedent mobile elements, and target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, correct processing of pre-crRN A requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (me), and a Cas9 domain. The tracrRNA serves as a guide for ribonuclease 3-aided processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves a linear or circular dsDNA target complementary to the spacer. The strand in the target DNA not complementary to crRNA is first cut endonucleolytically, then trimmed 3 '-5' exonucleolytically. In nature, DNA-binding and cleavage typically requires protein and both RNAs. However, single guide RNAs (“sgRNA”, or simply “gRNA”) can be engineered so as to incorporate aspects of both the crRNA and tracrRNA into a single RNA species. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821(2012), the contents of which are incorporated herein by reference. Cas9 recognizes a short motif in the CRISPR repeat sequences (the PAM or protospacer adjacent motif) to help distinguish self versus non-self. Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., “Complete genome sequence of an Ml strain of Streptococcus pyogenes."' Ferretti et al., J. J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C., Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y, Jia H.G., Najar F.Z., Ren Q„ Zhu H„ Song L„ White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663(2001); “CRISPR RNA maturation by trans-encoded small RNA and host factorB1195.70207WO00 15 / 102RNase III.” Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y, Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471:602-607(2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816- 821(2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including, but not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference. In some embodiments, a Cas9 nuclease comprises one or more mutations that partially impair or inactivate the DNA cleavage domain.

[0046] As used herein, the term “nucleic acid programmable DNA binding protein” or “napDNAbp,” of which Cas proteins such as Cas9 and variants thereof are examples, refers to a protein that uses RNA:DNA hybridization to target and bind to specific sequences in a DNA molecule. Each napDNAbp is associated with at least one guide nucleic acid (e.g., guide RNA), which localizes the napDNAbp to a DNA sequence that comprises a DNA strand (i.e., a target strand) that is complementary to the guide nucleic acid, or a portion thereof (e.g., the protospacer of a guide RNA). In other words, the guide nucleic acid “programs” the napDNAbp (e.g., Cas9, or a variant thereof) to localize and bind to a complementary sequence.

[0047] Without being bound by theory, the binding mechanism of a napDNAbp-guide RNA complex, in general, includes the step of forming an R-loop whereby the napDNAbp induces the unwinding of a double-strand DNA target, thereby separating the strands in the region bound by the napDNAbp. The guide RNA protospacer then hybridizes to the “target strand.” This displaces a “non-target strand” that is complementary to the target strand, which forms the single strand region of the R-loop. In some embodiments, the napDNAbp includes one or more nuclease activities, which then cut the DNA, leaving various types of lesions. For example, the napDNAbp may comprise a nuclease activity that cuts the non-target strand at a first location, and / or cuts the target strand at a second location. Depending on the nuclease activity, the target DNA can be cut to form a “double- stranded break” whereby both strands are cut. In other embodiments, the target DNA can be cut at only a single site, i.e., the DNA is “nicked” on one strand.B1195.70207WO00 16 / 102

[0048] In some embodiments, the napDNAbp of a base editor is a Cas9 domain. In some embodiments, the base editor comprises a Cas9 protein fused to an adenine or a cytidine deaminase. In some embodiments, the base editor comprises a Cas9 nickase (nCas9) fused to an adenine or a cytidine deaminase. In some embodiments, the base editor comprises a nuclease-inactive Cas9 (dCas9) fused to an adenine or cytidine deaminase. In some embodiments, the base editor comprises a Cas9 nickase (nCas9) fused to an adenine or cytidine deaminase.

[0049] As used herein, the term “guide RNA” is a particular type of nucleic acid which is commonly associated with a Cas protein (e.g., a Cas9 protein), directing the Cas protein to a specific sequence in a DNA molecule that includes complementarity to the protospacer sequence of the guide RNA.

[0050] The term “synthetic intron” refers to a noncoding DNA sequence that has been modified from the wildtype sequence. In some embodiments, a synthetic intron has been modified to have a defective 5' splice donor sequence, therefore preventing the intron from being spliced and preventing transgene expression. Those skilled in the art will recognize or be able to ascertain additional or alternative wildtype introns (e.g., SV40) that can be modified into synthetic introns to be used in the context of the invention. Non-limiting examples of such wildtype introns include the human beta-globin intron and human EFlalpha intron.

[0051] The term “defective 5' splice site” refers to modifications to the canonical GT sequence, the first two nucleotides of an intron, such that either or both nucleotides are altered to another nucleotide, preventing efficient splicing.

[0052] The term “defective 3' splice site” refers to modifications to the canonical AG sequence, the first two nucleotides of an intron, such that either or both nucleotides are altered to another nucleotide, preventing efficient splicing.

[0053] As used herein, the term “splice-intron guide RNA” (sigRNA) is a type of guide RNA that targets a portion of the 5' splice donor site or 3' splice acceptor site of a synthetic intron described herein. A sigRNA may have the structure of a “traditional guide RNA” or “traditional single guide RNA” (e.g., to be used with a base editor) or may have a modified gRNA structure of a “prime editing guide RNAs” (or “pegRNAs”), to be used in prime editing methods.

[0054] As used herein, the term “guide RNA” or “gRNA” is a particular type of guide nucleic acid which is mostly commonly associated with a Cas protein of a CRISPR-Cas9 andB1195.70207WO00 17 / 102which associates with Cas9, directing the Cas9 protein to a specific sequence in a DNA molecule that includes complementarity to the protospacer sequence of the guide RNA. However, this term also embraces the equivalent guide nucleic acid molecules that associate with Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), and which otherwise program the Cas9 equivalent to localize to a specific target nucleotide sequence. The Cas9 equivalents may include other napDNAbp from any type of CRISPR system (e.g., type II, V, VI), including Cpfl (a type-V CRISPR-Cas systems), C2cl (a type V CRISPR-Cas system), C2c2 (a type VI CRISPR-Cas system) and C2c3 (a type V CRISPR-Cas system). Further Cas-equivalents are described in Makarova et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector,” Science 2016; 353(6299), the contents of which are incorporated herein by reference. Exemplary sequences are and structures of guide RNAs are provided herein. As used herein, the “guide RNA” may also be referred to as a “traditional guide RNA” to contrast it with the modified forms of guide RNA termed “prime editing guide RNAs” (or “pegRNAs”) which have been invented for the prime editing methods and composition disclosed herein.

[0055] The terms “prime editing guide RNA” or “pegRNA” or “extended guide RNA” refer to a specialized form of a guide RNA that has been modified to include one or more additional sequences for implementing the prime editing methods and compositions described herein. As described herein, the prime editing guide RNA comprise one or more “extended regions” of nucleic acid sequence. The extended regions may comprise, but are not limited to, single-stranded RNA or DNA. Further, the extended regions may occur at the 3' end of a traditional guide RNA. In other arrangements, the extended regions may occur at the 5' end of a traditional guide RNA. In still other arrangements, the extended region may occur at an intramolecular region of the traditional guide RNA, for example, in the gRNA core region which associates and / or binds to the napDNAbp. The extended region comprises a “reverse transcriptase template sequence” or “DNA synthesis template” which encodes (by the polymerase of the prime editor) a single- stranded DNA which, in turn, has been designed to be (a) homologous with the endogenous target DNA to be edited, and (b) which comprises at least one desired nucleotide change (e.g., a transition, a transversion, a deletion, or an insertion) to be introduced or integrated into the endogenous target DNA. The extended region may also comprise other functional sequence elements, such as, but not limited to, a “primer binding site” and a “spacer or linker” sequence, or other structural elements, such as, but not limited to aptamers, stem loops, hairpins, toe loops, or an RNA-protein recruitmentB1195.70207WO00 18 / 102domain. General disclosure of pegRNAs are described in Nelson et al., “Engineered pegRNAs improve prime editing efficiency,” Nature Biotechnology, volume 40, pages402-410 (2022), the contents of which are incorporated herein by reference.

[0056] As referred to herein, gRNAs and pegRNAs may comprise various structural elements that include, but are not limited to: one or more spacer sequence, one or more gRNA cores, one or more extension arms, and one or more transcription terminators.

[0057] For example, a gRNA may direct a Cas protein (e.g., as part of a base editor) to a target site in the gene encoding a splice donor site of a synthetic intron (e.g., a modified version of the SV40 intron). However, this term also embraces the equivalent guide nucleic acid molecules that associate with Cas protein equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), and which otherwise program the Cas protein equivalent to localize to a specific target nucleotide sequence. The Cas protein equivalents may include other napDNAbps from any type of CRISPR system (e.g., type II, V, VI), including Cpfl (a type-V CRISPR-Cas system), C2cl (a type V CRISPR-Cas system), C2c2 (a type VI CRISPR-Cas system), and C2c3 (a type V CRISPR-Cas system). Further Cas-equivalents are described in Makarova et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector,” Science 2016; 353(6299), which is incorporated herein by reference. Exemplary sequences and structures of guide RNAs are provided herein.

[0058] Functionally, guide RNAs (gRNAs) associate with a Cas protein, directing (or programming) the Cas protein to a specific sequence in a DNA molecule that includes a sequence complementary (e.g., a region of complementarity) to the protospacer sequence for the guide RNA (gRNA). A gRNA is a component of the CRISPR / Cas system. The sequence specificity of a Cas DNA-binding protein is determined by gRNAs, which have nucleotide base-pairing complementarity to target DNA sequences (also referred to herein as a region of complementarity). For example, in some embodiments, a gRNA comprises a region of complementarity to a target DNA sequence (e.g., a target gene). The native gRNA comprises a 20 nucleotide (nt) Specificity Determining Sequence (SDS), or spacer, which specifies the DNA sequence to be targeted, and is immediately followed by an ~80 nt scaffold sequence, which associates the gRNA with the Cas protein. In some embodiments, an SDS of the present disclosure comprises a region of complementarity to a target DNA sequence. In some embodiments, an SDS of the present disclosure has a length of 15 to 100 nucleotides, or more. For example, an SDS may have a length of 15 to 90, 15 to 85, 15 to 80, 15 to 75, 15 to 70, 15 to 65, 15 to 60, 15 to 55, 15 to 50, 15 to 45, 15 to 40, 15 to 35, 15 to 30, or 15 to 20B1195.70207WO00 19 / 102nucleotides. In some embodiments, the SDS is 20 nucleotides long. For example, the SDS may be 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides long. At least a portion of the target DNA sequence is complementary to the SDS of the gRNA. For a Cas protein to successfully bind to the DNA target sequence, a region of the target sequence is complementary to the SDS of the gRNA sequence and is immediately followed by the correct protospacer adjacent motif (PAM) sequence. In some embodiments, an SDS is 100% complementary to its target sequence. In some embodiments, the SDS sequence is less than 100% complementary to its target sequence and is, thus, considered to be partially complementary to its target sequence. For example, a targeting sequence may be 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% complementary to its target sequence. In some embodiments, the SDS of template DNA or target DNA may differ from a complementary region of a gRNA by 1, 2, 3, 4, or 5 nucleotides.

[0059] In some embodiments, the guide RNA is about 15-120 nucleotides long and comprises a sequence of at least 10 contiguous nucleotides that is complementary to a target sequence (e.g., a target gene sequence). In some embodiments, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, or 120 nucleotides long. In some embodiments, the guide RNA comprises a sequence of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more contiguous nucleotides that is complementary to a target sequence. Sequence complementarity refers to distinct interactions between adenine and thymine (DNA) or uracil (RNA), and between guanine and cytosine.

[0060] The term “gene” refers to a nucleic acid fragment that expresses a specific protein, including regulatory sequences preceding (5' non-coding sequences) and following (3' noncoding sequences) the coding sequence. “Native gene” refers to a gene as found in nature with its own regulatory sequences. “Chimeric gene” or “chimeric construct” refers to any gene or a construct, not a native gene, comprising regulatory and coding sequences that are not found together in nature. Accordingly, a chimeric gene or chimeric construct may comprise regulatory sequences and coding sequences that are derived from different sources, or regulatory sequences and coding sequences derived from the same source, but arranged in a manner different than that found in nature. “Endogenous gene” refers to a native gene in its natural location in the genome of an organism. A “foreign” gene refers to a gene not normallyB1195.70207WO00 20 / 102found in the host organism, but which is introduced into the host organism by gene transfer. Foreign genes can comprise native genes inserted into a non- native organism, or chimeric genes. A “transgene” is a gene that has been introduced into the genome by a transformation procedure.

[0061] The term “fusion protein” as used herein refers to a hybrid polypeptide that comprises protein domains from at least two different proteins. One protein may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxy-terminal (C-terminal) protein, thus forming an “amino-terminal fusion protein” or a “carboxy-terminal fusion protein,” respectively. A protein may comprise different domains, for example, a Cas9 protein fused to a deaminase domain (z.e., a base editor). Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced via recombinant protein expression and purification, which is especially suited for fusion proteins comprising a peptide linker. Methods for recombinant protein expression and purification are well known, and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), the entire contents of which is incorporated herein by reference. Fusion proteins useful for the methods disclosed herein include cytidine base editors (CBEs), in which the deaminase domain is a cytidine deaminase. Other fusion proteins useful for the methods disclosed herein include adenine base editors (ABEs), in which the deaminase domain is an adenine deaminase.

[0062] The term “mutation” as used herein, refers to a substitution, insertion, or deletion of a single residue or a combination of residues within a sequence, e.g., a nucleic acid or amino acid sequence, with another residue, or a deletion or insertion of one or more residues within a sequence. Mutations are typically described herein by identifying the original residue followed by the position of the residue within the sequence and by the identity of the newly substituted residue. Various methods for making the amino acid substitutions (mutations) provided herein are well known in the art, and are provided by, for example, Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)).

[0063] The terms “polynucleotide,” “nucleotide sequence,” “nucleic acid,” “nucleic acid molecule,” “nucleic acid sequence,” and “oligonucleotide” refer to a series of nucleotide bases (also called “nucleotides”) in DNA and RNA and mean any chain of two or more nucleotides. The polynucleotides can be chimeric mixtures or derivatives or modified versions thereof, and single-stranded or double- stranded. The oligonucleotide can beB1195.70207WO00 21 / 102modified at the base moiety, sugar moiety, or phosphate backbone, for example, to improve stability of the molecule, its hybridization parameters, etc.

[0064] The term “ribonucleotide” refers to a nucleotide containing ribose as its pentose component. It is considered a molecular precursor of nucleic acids. Nucleotides are the basic building blocks of DNA and RNA. Ribonucleotides themselves are basic monomeric building blocks for RNA. In living organisms, the most common bases for ribonucleotides are adenine (A), guanine (G), cytosine (C), or uracil (U).

[0065] A “protein,” “peptide,” or “polypeptide” comprises a polymer of amino acid residues linked together by peptide bonds. The term refers to proteins, polypeptides, and peptides of any size, structure, or function. Typically, a protein will be at least three amino acids long. A protein may refer to an individual protein or a collection of proteins. Proteins may contain only natural amino acids, although non-natural amino acids (z.e., compounds that do not occur in nature but that can be incorporated into a polypeptide chain) and / or amino acid analogs as are known in the art may alternatively be employed. Also, one or more of the amino acids in a protein may be modified, for example, by the addition of a chemical entity such as a carbohydrate group, a hydroxyl group, a phosphate group, a farnesyl group, an isofarnesyl group, a lipid, a fatty acid group, a linker for conjugation or functionalization, or other modification. A protein may also be a single molecule or may be a multi-molecular complex. A protein may be a fragment of a naturally occurring protein or peptide. A protein may be naturally occurring, recombinant, synthetic, or any combination of these. A protein may also be a therapeutic protein administered as a treatment for a disease or disorder (e.g., one that is associated with a change in the RNA expression and / or translation profile of a cell taken from a subject).

[0066] A “subject” to which administration is contemplated refers to a human (i.e., male or female of any age group, e.g., pediatric subject (e.g., infant, child, or adolescent) or adult subject (e.g., young adult, middle-aged adult, or senior adult)) or non-human animal. In some embodiments, the non-human animal is a mammal (e.g., primate (e.g., cynomolgus monkey or rhesus monkey) or mouse). The term “patient” refers to a subject in need of treatment of a disease. In some embodiments, the subject is human. In some embodiments, the patient is human. The human may be a male or female at any stage of development. A subject or patient “in need” of treatment of a disease or disorder includes, without limitation, those who exhibit any risk factors or symptoms of a disease or disorder. In some embodiments, a subject is a non-human experimental animal (e.g., a mouse, rat, dog, pig, or non-human primate).B1195.70207WO00 22 / 102

[0067] As used herein, the term “variant” should be taken to mean the exhibition of qualities that have a pattern that deviates from what occurs in nature comprising one or more changes in amino acid residues (i.e., “substitutions”, “insertions, or “deletions”) as compared to a wild type amino acid sequence. The term “variant” encompasses homologous proteins having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with a reference sequence and having the same or substantially the same functional activity or activities as the reference sequence. The term also encompasses mutants, truncations, or domains of a reference sequence that display the same or substantially the same functional activity or activities as the reference sequence.

[0068] As used herein the term “wild type” is a term of the art understood by skilled persons and means the typical form of an organism, strain, gene, or characteristic as it occurs in nature as distinguished from mutant or variant forms.

[0069] Throughout the present disclosure, when a range of values is listed, it is intended to encompass each value and sub-range within the range. Where ranges are given, endpoints are included.

[0070] Unless otherwise required by context, singular terms shall include pluralities, and plural terms shall include the singular.

[0071] The details of certain embodiments of the invention are set forth in the Detailed Description of Certain Embodiments, as described below. Other features, objects, and advantages of the invention will be apparent from the Drawings, Definitions, Examples, and Claims.DETAILED DESCRIPTION

[0072] The aspects described herein are not limited to specific embodiments, compositions, complexes, systems, methods, uses, or configurations, and as such can, of course, vary. The terminology used herein is for the purpose of describing particular aspects only and, unless specifically defined herein, is not intended to be limiting.

[0073] In the present disclosure, defective splicing has been used to connect the expression of a selection marker gene to the activity of a gene editor within a cell to enrich for gene editing activity in the cell. For example, the present disclosure describes the use of a polynucleotide cassette comprising a coding sequence of a selection marker gene that is disrupted by a synthetic intron sequence to prevent gene expression of the selection marker. The synthetic intron further comprises a defective 5'-splice donor site that is correctable by a base edit (e.g., by a base editor) such that when the 5 '-splice donor site is corrected by a gene editor, theB1195.70207WO00 23 / 102intron is removed by the endogenous slicing system of the cell and expression of the selection marker gene can occur. The methods described herein apply this relationship to identify the presence of active gene editors (e.g., base editors) in the cell by subjecting the cell to the selection pressure of the selection marker. Accordingly, cell survival indicates the desired base edit was incorporated at the 5 '-splice donor site of the synthetic intron, thus indicating the cells contain active gene editors.Synthetic introns

[0074] The present disclosure contemplates, without limitation, any suitable intron sequence for use in the context of the present invention, such as an intron from a human gene (e.g. human beta globin) or a virus that is known to splice efficiently in human cells (e.g., SV40). The present disclosure contemplates designing and verifying synthetic introns from naturally occurring introns for use in the context of the present invention. For example, naturally occurring intron sequences may be modified (e.g., at the 5’ or 3’ splice site) to render them “defective” such that they cannot be recognized by the endogenous slicing system of the cell. A person of ordinary skill in the art would be able to recognize and identify suitable intronic sequences for use in the context of the present invention.

[0075] These intronic sequences may be modified to include mutations within the recognition sequence of the 5' or 3' splice site of the intron that may be correctable by a gene editor (e.g., a base editor or prime editor) to restore the 5' or 3' splice site such that the intronic sequence can be recognized and consequently removed by the splicing system of the cell. The present disclosure further contemplates the use of these synthetic introns for enriching for cells comprising active gene editors.

[0076] In one aspect, the present disclosure describes a polynucleotide comprising a nucleotide sequence encoding a synthetic intron having at least 80% sequence identity to the nucleic acid sequence of the SV40 intron (SEQ ID NO: 31). In certain embodiments, the present disclosure describes a polynucleotide comprising a nucleotide sequence encoding a synthetic intron having at least 80% sequence identity to the nucleic acid sequence of the SV40 intron (SEQ ID NO: 31), further comprising at least one nucleic acid mutation relative to the naturally occurring SV40 intron.

[0077] Nucleic acid sequence of SV40 intron

[0078] GTAAGTTTAGTCTTTTTGTCTTTTATTTCAGGTCCCGGATCCGGTGGTGGTG CAAATCAAAGAACTGCTCCTCAGTGGATGTTGCCTTTACTTCTAG (SEQ ID NO: 31)B1195.70207WO00 24 / 102

[0079] The bolding in SEQ ID NO: 31 indicates the “GT” pair modified to either an “AT” pair or a “GC” pair to render the 5’ splice donor site defective.

[0080] In certain embodiments, the polynucleotide comprises a synthetic intron that is at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% identical to the nucleic acid sequence of the SV40 intron (SEQ ID NO: 31). In certain embodiments, the polynucleotide comprises a synthetic intron that is at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% identical to the nucleic acid sequence of the SV40 intron (SEQ ID NO: 31), further comprising at least one nucleic acid mutation relative to the naturally occurring SV40 intron. The present disclosure describes polynucleotides comprising a nucleotide sequence encoding a synthetic intron having at least one, at least two, at least three, at least four, at least five, at least six, or at least seven nucleic acid substitutions compared to a nucleic acid sequence of the SV40 intron (SEQ ID NO: 31). In certain embodiments, the synthetic intron has at least one nucleic acid substitution compared to a nucleic acid sequence of the SV40 intron. In certain embodiments, the synthetic intron has at least two nucleic acid substitutions compared to a nucleic acid sequence of the SV40 intron. In certain embodiments, the synthetic intron has at least three nucleic acid substitutions compared to a nucleic acid sequence of the SV40 intron. In certain embodiments, the synthetic intron has at least four nucleic acid substitutions compared to a nucleic acid sequence of the SV40 intron. In certain embodiments, the synthetic intron has at least five nucleic acid substitutions compared to a nucleic acid sequence of the SV40 intron. In certain embodiments, the synthetic intron has at least six nucleic acid substitutions compared to a nucleic acid sequence of the SV40 intron. In certain embodiments, the synthetic intron has at least seven nucleic acid substitutions compared to a nucleic acid sequence of the SV40 intron.

[0081] As described herein, the term “substitution” refers to the addition of a nucleobase, the deletion of a nucleobase, or the “swapping” of the wildtype nucleobase for different nucleobase (e.g., A for a G) not found in the wild type sequence.

[0082] In certain embodiments, the synthetic intron comprises a defective 5' splice donor site which is correctable by a gene editor, such as a base editor or primer editor. As defined herein, the term “defective 5' splice donor site” refers to a 5' splice donor site of a synthetic intron which contains one or more nucleic acid mutations relative to a naturally occurring intron (e.g., SV40 intron) that prevents the intronic sequence from being removed (e.g., spliced) from a coding sequence by a spliceosome.

[0083] The present disclosure also contemplates the use of a synthetic intron comprising a defective 3' splice acceptor site which is correctable by a gene editor, such as a base editor orB1195.70207WO00 25 / 102primer editor. As defined herein, the term “defective 3' splice acceptor site” refers to a 3' splice acceptor site of a synthetic intron which contains one or more nucleic acid mutations relative to a naturally occurring intron that prevents the intronic sequence from being removed (e.g., spliced) from a coding sequence by a spliceosome.

[0084] In certain embodiments, the defective 5' splice donor site comprises a single nucleic acid mutation within the canonical “GT” recognition sequence of the SV40 intron. In certain embodiments, the defective 5' splice donor site comprises a “AT” or “GC” (positions underlined in Table 1) in place of the canonical “GT” recognition sequence of the SV40 intron (first two nucleotides in SEQ ID NO: 31). In certain embodiments, the defective 5' splice donor site comprises a “AT” in in place of the canonical “GT” recognition sequence of the SV40 intron (SEQ ID NO: 31). In certain embodiments, the defective 5' splice donor site comprises a “GC” in in place of the canonical “GT” recognition sequence of the SV40 intron (SEQ ID NO: 31). In certain embodiments, the defective 5' splice donor site is restored (e.g., corrected) by converting the “AT” mutation back to the canonical “GT” recognition sequence of the SV40 intron via a single base edit. In certain embodiments, the defective 5' splice donor site is restored by converting the “AT” mutation back to the canonical “GT” recognition sequence of the SV40 intron by a base editor (e.g., an adenine base editor). In certain embodiments, the defective 5' splice donor site is restored (e.g., corrected) by converting the “GC” mutation back to the canonical “GT” recognition sequence of the SV40 intron via a single base edit. In certain embodiments, the defective 5' splice donor site is restored by converting the “GC” mutation back to the canonical “GT” recognition sequence of the SV40 intron by a base editor (e.g., a cytidine base editor). In certain embodiments, the synthetic intron is removed from a coding sequence if the defective 5' splice donor site is corrected by a base editor.

[0085] In certain embodiments, the synthetic introns described herein further comprises a protospacer adjacent motif (PAM) sequence. In certain embodiments, the synthetic intron comprises the PAM sequence “NGN”, wherein N is defined as an A, G, T, or C. In certain embodiments, the PAM sequence is located downstream of the mutated recognition sequence of the synthetic intron. In certain embodiments, the polynucleotides described herein are DNA sequences.

[0086] In certain embodiments, the synthetic intron comprises a nucleic acid sequence selected from the nucleic acid sequences listed in Table 1.Table 1.B1195.70207WO00 26 / 102Bolding = nucleotides part of the exon.Underline = target sequence in defective 5 ’donor site; AT = correctable by ABE; CG = correctable by CBE.

[0087] In some embodiments, the polynucleotide comprising the synthetic intron is inserted into the protein coding region of a selection marker gene. In certain embodiments, the first three nucleotides of the synthetic intron are complementary to the coding region of the selection marker (see Table 1). In certain embodiments, the synthetic intron is further modified to include at least one, at least two, or at least three stop codons that are in frame with the selection marker gene. In certain embodiments, the synthetic intron is modified to include a stop codon that is in every reading frame of the selection marker gene to prevent translation of functional selection marker gene product (e.g., puromycin resistance protein). Nucleic acid sequences of exemplary synthetic introns comprising at least one stop codon are listed in Table 2.B1195.70207WO00 27 / 102Table 2.Polynucleotide cassettes

[0088] In another aspect, the present disclosure describes polynucleotide cassettes comprising a nucleotide sequence encoding a selection marker disrupted by any synthetic intron described herein or elsewhere in the art (FIG 3 A and 3B). In certain embodiments, the polynucleotide cassette comprises a nucleotide sequence encoding a selection marker disrupted by a synthetic intron having at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% sequence identity to a nucleic acid sequence of the SV40 intron (SEQ ID NO: 31), wherein the synthetic intron comprises a defective 5' donor splice site that is correctable by a gene editor. In certain embodiments, the polynucleotide cassette comprises a nucleotide sequence encoding a selection marker disrupted by a synthetic intron having at least one, at least, two, at least three, at least four, at least five, at least six, or at least seven nucleic acid substitutions compared to a nucleic acid sequence of the SV40 intron (SEQ ID NO: 31), wherein the synthetic intron comprises a defective 5' donor splice site that is correctable by a gene editor.

[0089] As defined herein, a “selection marker gene” is a gene introduced into cells, such as bacteria or cells in culture, which confers one or more traits suitable for artificial selection. The skilled artisan would be able to identify suitable selection marker genes to be utilized in the context of the invention. In certain embodiments, the selection marker gene is an antibiotic resistance gene, an antimetabolite marker gene, or a herbicide resistance gene.

[0090] In certain embodiments, the selection marker gene is an antibiotic resistance gene. In certain embodiments, the antibiotic resistance gene is selected from the group consisting of a puromycin resistance gene, a blasticidin resistance gene, an ampicillin resistance gene, a chloramphenicol resistance gene, a streptomycin resistance gene, or a kanamycin resistance gene. In certain embodiments, the selection marker gene is a puromycin resistance gene.B1195.70207WO00 28 / 102

[0091] In certain embodiments, the polynucleotide cassette may comprise the structure selected from:5'-[selection marker-N']- [synthetic intron] -[selection marker-C']-3'5'-[expression marker] -[selection marker-N']-[synthetic intron] -[selection marker-cos'wherein each instance of “]-[” comprises an optional linker, e.g. a peptide linker.

[0092] In certain embodiments, the expression marker encodes for a fluorescent protein (FP). In certain embodiments, the gene expression marker encodes for Green Fluorescent Protein (GFP). One of ordinary skill in the art would be able to recognize and identify other expression marker genes that may be utilized within the context of the invention.Splice-intron guide RNA (sigRNA)

[0093] In some aspects, the present disclosure describes a splice-intron guide RNA (sigRNA, also referred to as a splice guide RNA) comprising a first portion comprising a region of complementarity (e.g., the spacer sequence of a gRNA) to a nucleic acid encoding a 5 ’-splice donor site of a SV40 intron or a variant thereof (e.g., SEQ ID NOs: 12-22).

[0094] In certain embodiments, the sigRNA is designed for use with a base editor. In certain embodiments, the sigRNA is designed for use with a primer editor.

[0095] In some embodiments, the splice-intron guide RNAs comprise a region of complementarity to a nucleic acid encoding ay one of the synthetic introns described herein or elsewhere. In some embodiments, the provided splice-intron gRNAs target a gene editor to a site in a synthetic intron described herein or elsewhere. In some embodiments, the spliceintron gRNAs target a base editor to a site in a synthetic intron such that the base editor can correct a spliceosome recognition sequence within the defective 5’-splice donor site.

[0096] In some embodiments, the region of complementarity comprises a nucleic acid encoding a 5’splice donor site of the synthetic intron that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a nucleic acid sequence set forth in SEQ ID NO.: 31, or comprises at least one, at least two, at least three, at least four, at least five, at least six, or at least seven substitutions relative to the sequence set forth in SEQ ID NO: 31. In certain embodiments, the region of complementarity comprises a nucleic acid sequence selected Table 3.Table 3.B1195.70207WO00 29 / 102

[0097] In other embodiments, the splice-intron gRNA further comprises a second portion comprising a trans-activating crRNA (tracrRNA). In certain embodiments, the tracrRNA facilitates the binding of the splice-intron gRNA to a napDNAbp, for example, a Cas9 protein (e.g., a Cas9 protein as part of a base editor).

[0098] In some embodiments, the sigRNA comprises the structure 5'-[spacer]-[tracrRNA]-3'. The present disclosure contemplates the use of any suitable tracrRNA that facilitates the binding of the splice-intron gRNA to a napDNAbp, for example, a Cas9 protein as part of a base editor or primer editor. For example, in some embodiments, the tracrRNA comprises a sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to, or comprising at least one, at least two, at least three, at least four, at least five, at least six, or at least seven substitutions relative to any one of the sequences set forth in SEQ ID NOs: 32-34. In certain embodiments, the splice-intron gRNA comprises the sequence set forth in SEQ ID NOs: 32-34.

[0099] GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTT GAAAAAGTGGCACCGAGTCGGTGC (SEQ ID NO: 32)

[0100] GTTTAAGAGCTATGCTGGAAACAGCATAGCAAGTTTAAATAAGGCTAGTCC GTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC (SEQ ID NO: 33)

[0101] GTTTAAGAGCTATGCTGGAAACAGCATAGCAAGTTTAAATAAGGCTAGTCC GTTATCAACTTTGCTGGAAACAGCAAAGTGGCACCGAGTCGGTGC (SEQ ID NO: 34)

[0102] In certain embodiments, the splice-intron gRNA is comprised of ribonucleotides. In some embodiments, a splice-intron gRNA is about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 75, about 100, or more nucleotides in length. In some embodiments, a splice-intron gRNA is about 50-150, about 60-140, about 70-130, about 80-120, or about 90-110 nucleotides in length. In some embodiments, the first portion of theB1195.70207WO00 30 / 102splice-intron gRNAis about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30 nucleotides in length.

[0103] In some embodiments, a splice-intron gRNA comprises an optional linker sequence. For example, the splice-intron gRNAs provided herein may comprise an optional linker sequence between the first portion of the splice-intron gRNA and the tracrRNA sequence. In certain embodiments, the optional linker sequence is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, or at least 50 nucleotides in length.

[0104] In general, a guide RNA spacer is any RNA sequence having sufficient complementarity with a target polynucleotide sequence (e.g., synthetic intron) to hybridize with the target sequence and direct sequence-specific binding of a napDNAbp (e.g., Cas9, which may be part of a base editor) to the target sequence. In some embodiments, the degree of complementarity between the spacer and its corresponding target sequence in the synthetic intron described herein, when optimally aligned using a suitable alignment algorithm, is about or more than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more (or the spacer and the corresponding target sequence comprise one, two, three, four, five, six, seven, eight, nine, or ten nucleic acid differences). In certain embodiments, the spacer is 100% complementary to its corresponding target sequence in the synthetic intron. Optimal alignment may be determined with the use of any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (e.g., the Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net).

[0105] In certain embodiments, the polynucleotide cassette described herein may comprise the structure:5'-[sigRNA]-[selection marker-N']-[synthetic intron] -[selection marker-C']-3' 5'-[sigRNA]-[expression marker] -[selection marker-N']-[synthetic intron]-[selection marker-C']-3'B1195.70207WO00 31 / 102wherein each instance of “]-[” comprises an optional linker, e.g., a peptide linker.Compositions

[0106] The present disclosure describes compositions that may comprise any of the nucleic acids encoding the polynucleotides, the polynucleotide cassettes, splice-intron guide RNAs, and / or gene editors described herein. In certain embodiments, a composition may further comprise a nucleic acid encoding a guide RNA (gRNA) that comprises a spacer sequence comprising a region of complementarity to a target gene of interest, a nucleic acid encoding a gene expression marker, and / or a nucleic acid encoding a gene editor.

[0107] In one aspect, the present disclosure presents a composition of nucleic acids, the composition comprising:(a) a first nucleic acid encoding a selection marker disrupted by a synthetic intron having at least 80% sequence identity to a nucleic acid sequence set forth in SEQ ID NO: 31, wherein the synthetic intron comprises a defective 5' donor splice site that is correctable by a gene editor;(b) a second nucleic acid encoding a splice-intron guide RNA (sigRNA) comprising a spacer sequence that has complementarity to the defective 5’ donor splice site of the synthetic intron of (a), and a trans-activating CRISPR RNA(tracrRNA); and / or(c) a third nucleic acid encoding a guide RNA (gRNA) comprising a spacer sequence comprising a region of complementarity to a target gene of interest; and / or(d) a fourth nucleic acid encoding a gene expression marker; and / or(e) a fifth nucleic acid encoding a gene editor.

[0108] In certain embodiments, the gene editor is a nucleic acid programmable gene editor.

[0109] In certain embodiments, the gRNA directs the gene editor (e.g., base editor) to a specific sequence in a DNA molecule that includes complementarity to the spacer sequence of the guide RNA. In certain embodiment, the gRNA may direct a base editor to a target DNA molecule (e.g., a gene) within the genome of a cell. In certain embodiment, the gRNA is able to bind to the same base editor that the sigRNA.

[0110] In certain embodiments, the expression marker encodes for a fluorescent protein (FP). In certain embodiments, the gene expression marker encodes for Green Fluorescent Protein (GFP). One of ordinary skill in the art would be able to recognize and identify other expression marker genes that may be utilized within the context of the invention.B1195.70207WO00 32 / 102

[0111] In certain embodiments, the nucleic acid of (a), the nucleic acid of (b), the nucleic acid of (c), and / or the nucleic acid of (d) are located on the same nucleic acid construct or different nucleic acid constructs.

[0112] In certain embodiments the nucleic acid programmable gene editor is a base editor or a prime editor. Accordingly, the sigRNA may be a guide RNA molecule that is suitable for base editor or prime editing. In certain embodiments, the sigRNA is suitable for base editing. In certain embodiments, the sigRNA is suitable for prime editing.Complexes

[0113] In another aspect, the present disclosure provides complexes comprising a gene editor (e.g., a base editor or prime editor) and a guide RNA (e.g., a sigRNA), optionally bound to a nucleic acid programmable DNA binding protein (napDNAbp) of the gene editor. In some embodiments, described herein are compositions comprising a base editor and a sigRNA, optionally bound to a Cas9 domain (e.g., a dCas9, a nuclease active Cas9, or a Cas9 nickase) of the base editor. In some embodiments, described herein are compositions comprising a prime editor and a sigRNA, optionally bound to a Cas9 domain (e.g., a dCas9, a nuclease active Cas9, or a Cas9 nickase) of the prime editor. In some embodiments, the disclosure describes complexes comprising a gene editor and a sigRNA bound to napDNAbp of the gene editor (e.g., a base editor or prime editor). In some embodiments, the disclosure provides complexes comprising any of the base editors or prime editors provided herein, and a sigRNA bound to a Cas9 domain (e.g., a dCas9, a nuclease active Cas9, or a Cas9 nickase) of the base editor or prime editor, respectively.

[0114] In certain embodiments, the composition or complex further comprises a target nucleic acid (e.g., a synthetic intron). In certain embodiments, the gene editor is a fusion protein capable of base editing. In some embodiments, the fusion protein comprises a deaminase and a napDNAbp. In some embodiments, the deaminase is a cytidine deaminase or an adenosine deaminase. In some embodiments, the deaminase is a cytidine deaminase. In some embodiments, the deaminase is an adenosine deaminase.napDNAbp

[0115] The gene editors (e.g., base editors or primer editors) described herein may comprise a nucleic acid programmable DNA binding protein (napDNAbp). In one aspect, a napDNAbp can be associated with or complexed with at least one polynucleotide described herein (e.g.,B1195.70207WO00 33 / 102guide RNA or a splice-intron guide RNA), which localizes the napDNAbp to a DNA sequence that comprises a DNA strand (i.e., a target nucleic acid sequence) that is complementary to the guide nucleic acid, or a portion thereof (e.g., the spacer of a guide RNA which anneals to the protospacer of the DNA target). In other words, the guide nucleic acid “programs” the napDNAbp (e.g., Cas9 or equivalent) to localize and bind to complementary sequence of the protospacer in the DNA.

[0116] As noted herein, Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., “Complete genome sequence of an Ml strain of Streptococcus pyogenes.” Ferretti et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C., Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y, Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663(2001); “CRISPR RNA maturation by transencoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y, Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471:602-607(2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E.Science 337:816-821(2012), the entire contents of each of which are incorporated herein by reference).

[0117] Examples of Cas9 and Cas9 equivalents are provided as follows; however, these specific examples are not meant to be limiting. The prime editors and base editors of the present disclosure may use any suitable napDNAbp, including any suitable Cas9 or Cas9 equivalent. In various embodiments, the napDNAbp may be any Class 2 CRISPR-Cas system, including any type II, type V, or type VI CRISPR-Cas enzyme. In some embodiments, Cas9 refers to the wild type Cas9 from Streptococcus pyogenes (SpCas9; NCBI Reference Sequence: NC_017053.1, Uniport Reference Sequence: Q99ZW2, SEQ ID NO: 27), or a variant Cas9 protein having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to SpCas9.

[0118] Streptococcus pyogenes Cas9:MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGET AEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERH PIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLN PDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEK KNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLF LAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKB1195.70207WO00 34 / 102EIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFD NGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWM TRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNE LTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEI SGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTY AHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQ LIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGR HKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKL YLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKS DNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVE TRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNY HHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYF FYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIV KKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKG KSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRK RMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLD EIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYF DTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO: 27)

[0119] The base editors or prime editors may comprise the canonical SpCas9, or any ortholog Cas9 protein, or any variant Cas9 protein — including any naturally occurring variant, mutant, or otherwise engineered version of Cas9 — that is known or which can be made or evolved through a directed evolutionary or otherwise mutagenic process. In various embodiments, the Cas9 or Cas9 variants have a nickase activity, i.e., only cleave of strand of the target DNA sequence.

[0120] Any suitable napDNAbp may be used in the base editors or prime editors described herein. The skilled person will be able to identify the specific CRISPR-Cas enzyme being referenced in this Application based on the nomenclature that is used, whether it is old (i.e., “legacy”) or new nomenclature. The particular CRISPR-Cas nomenclature used in any given instance in this Application is not limiting in any way and the skilled person will be able to identify which CRISPR-Cas enzyme is being referenced. CRISPR-Cas nomenclature is extensively discussed in Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?,” The CRISPR Journal, Vol.l. No.5, 2018, the entire contents of which are incorporated herein by reference. In various embodiments, the napDNAbp may be any Class 2 CRISPR-Cas system, including any type II, type V, or type VI CRISPR-CasB1195.70207WO00 35 / 102enzyme. In certain embodiments, the napDNAbp is selected from the group consisting of Cas9, CasX, CasY, Cpfl, C2cl, C2c2, C2C3, Sp-Cas9, SpRY, SpG-Cas9, NG-Cas9, NRRH-Cas9, spCas9, geoCas9, saCas9, Nme2Cas9, Casl2(a-i), Casl4, Argonaute, and variants thereof. In certain embodiments, the nucleic acid programmable DNA binding proteins include, without limitation, Cas9 (e.g., dCas9 and nCas9), Casl2e (CasX), Casl2d (CasY), Casl2a (Cpfl), Casl2bl (C2cl), andCasl2c (C2c3).

[0121] In various embodiments, the Cas9 or Cas9 variants have a nickase activity (nCas9), and only cleave one strand of the target DNA sequence. In certain embodiments, the nCas9 is a Cas9 domain comprising a mutation corresponding to the D10A mutation of the wild type SpCas9 polypeptide of SEQ ID NO: 27. In other embodiments, the Cas9 or Cas9 variants have inactive nucleases, i.e., are “dead” Cas9 proteins (dCas9). In some embodiments, the napDNAbp is a dCas9, and corresponds to a Cas9 amino acid sequence having one or more mutations that inactivate the Cas9 nuclease activity (e.g., a Cas9 domain comprising a mutation corresponding to the D10A and / or H840A mutation of the wild type SpCas9 polypeptide of SEQ ID NO: 27).

[0122] Other variant Cas9 proteins that may be used are those having a smaller molecular weight than the canonical SpCas9 (e.g., for easier delivery) or having modified or rearranged primary amino acid structure (e.g., the circular permutant formats). The base editors or prime editors described herein may also comprise Cas9 equivalents, including Casl2a / Cpfl and Casl2b proteins which are the result of convergent evolution. The napDNAbps used herein (e.g., SpCas9, Cas9 variant, or Cas9 equivalents) may also contain various modifications that alter / enhance their PAM specificities (e.g., SpRY).

[0123] Lastly, the application contemplates any Cas9, Cas9 variant, or Cas9 equivalent which has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identity to a reference Cas9 sequence, such as SpCas9 or a reference Cas9 equivalent (e.g., Casl2a / Cpfl).Prime editors

[0124] Gene editors useful for the methods disclosed herein include prime editors. Any prime editor known in the art can be used in the methods of the present disclosure. In some embodiments, the prime editor comprises a nucleic acid-programmable DNA-binding protein (napDNAbp) and a polymerase. In some embodiments, the prime editor comprises a napDNAbp (e.g., a Cas9 protein, such as SpCas9, or a variant thereof, such as nCas9 orB1195.70207WO00 36 / 102dCas9) and a polymerase (e.g., a reverse transcriptase, such as an MMLV reverse transcriptase, a Tfl reverse transcriptase, or a variant thereof). General disclosure of prime editing is described in Anzalone, A. V. et al., Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149-157 (2019), which is incorporated herein by reference.

[0125] Some aspects of the present disclosure provide methods of prime editing a synthetic intron within a coding gene. In one aspect, the present disclosure provides methods of prime editing a synthetic intron g comprising contacting a nucleic acid sequence encoding the synthetic intron gene with a prime editor and a stable pegRNA. In certain embodiments, the pegRNA comprising a space sequence that have a region of complementary to a defective 5' splice donor site or a 3' splice acceptor site within a synthetic intron. In some embodiments, the step of contacting installs an edit that removes the mutation from the 5' splice donor site or a 3' splice acceptor site within a synthetic intron.Base editors

[0126] Gene editors useful for the methods disclosed herein include adenine base editors (ABEs), in which the deaminase domain is an adenine deaminase. The adenine deaminases (e.g. engineered adenine deaminases, evolved adenine deaminases) provided herein may be enzymes that convert adenine (A) to guanine (G) in DNA, leading to an A:T to G:C base pair conversion. In some embodiments, the adenine deaminase is derived from a bacterium, such as E.coli, S. aureus, B. subtilis, S. typhimurium, S. putrefaciens, H. influenzae, C. crescentus. G. sulfurreducens, or S. pyogenes. In some embodiments, the adenine deaminase is a TadA deaminase. It should be appreciated, however, that additional adenosine deaminases useful in the present application would be apparent to the skilled artisan and are within the scope of this disclosure. For example, the adenosine deaminase may be a homolog of an adenosine deaminase. In certain embodiments, the adenine deaminase is TadA-8e or variants thereof. In certain embodiments, the adenine base editor may be selected from the group consisting of ABE8.0, ABE8e, AYBE, ABE9, and variants thereof. In some embodiments, the gene editor is an ABE8e adenine base editor.

[0127] Exemplary TadA amino acid sequence:MRRAFITGVFFLSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGW NRPIGRHDPTAHAEIMAERQGGEVMQNYREIDATEYVTEEPCVMCAGAMIHSRIGRV VFGARDAKTGAAGSEMDVEHHPGMNHRVEITEGIEADECAAEESDFFRMRRQEIKA QKKAQSSTD (SEQ ID NO: 28 - EcTadA)B1195.70207WO00 37 / 102

[0128] ABE8e adenine deaminase domain amino acid sequence:SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAH AEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGWRNSKRGA AGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN(SEQ ID NO: 35)

[0129] In some embodiments, the gene editors useful for the methods disclosed herein include cytidine base editors (CBEs), in which the deaminase domain is a cytidine deaminase. The cytidine deaminase protein can be any cytidine deaminase protein known in the art. In some embodiments, the cytidine deaminase is a deaminase from the apolipoprotein B mRNA-editing complex (APOBEC) family deaminase. Examples of wildtype cytidine deaminase proteins include rat apolipoprotein B mRNA editing catalytic subunit 1 (rAPOBECl; Uniport Reference Sequence: P38483), which is represented by the amino acid sequence set forth in SEQ ID NO: 29 and human APOBEC 1 (Uniport reference Sequence: P41238), which is represented by the amino acid sequence set forth in SEQ ID NO: 30.

[0130] rAPOBECl MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNT NKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIAR LYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWV RLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK (SEQ ID NO: 29)

[0131] hAPOBECl MTSEKGPSTGDPTLRRRIEPWEFDVFYDPRELRKEACLLYEIKWGMSRKIWRSSGKN TTNHVEVNFIKKFTSERDFHPSMSCSITWFLSWSPCWECSQAIREFLSRHPGVTLVIYV ARLFWHMDQQNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYP PLWMMLYALELHCIILSLPPCLKISRRWQNHLTFFRLHLQNCHYQTIPPHILLATGLIHP SVAWR (SEQ ID NO: 30)

[0132] In some embodiments, a cytidine deaminase domain is an apolipoprotein B mRNA-editing complex (APOBEC) family deaminase variant. In some embodiments, the deaminase is an APOBEC 1, APOBEC 1, APOBEC3, APOBEC3A-APOBEC3F, or APOBEC4 deaminase protein variant. In some embodiments, the fusion protein comprising a cytidine deaminase further comprises one or more Uracil-DNA glycosylase inhibitor (UGI) domains. In some embodiments, the UGI proteins include fragments of UGI proteins and proteins homologous to a UGI or a UGI fragments. In some embodiments, the cytidine base editor may be selectedB1195.70207WO00 38 / 102from the group consisting of evoCDAmax, CBE6, CGBE, BE4max, TadCBEd, and variants thereof. In some embodiments, the gene editor is an TadCBEd cytidine base editor.

[0133] TadCBEd cytidine deaminase domain amino acid sequence:SEVEFSHEYWMRHALTLAKRARDERKAPVGAVLVLNNRVIGEGWNRAIGLHDPTAH AEIIALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMINSRIGRVVFGVRNSKRGAA GSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN( SEQ ID NO: 36)Systems

[0134] The present disclosure describes systems of polynucleotide cassettes comprising novel synthetic introns to identify gene expression and activity of a base editor within a cell. The present disclosure contemplates the use of the systems and methods described herein for base editing a target nucleic acid within a target nucleic acid sequence.

[0135] In one aspect, the present disclosure provides systems for enriching for base editing in a cell. In some embodiments, such a system comprises:(a) a nucleic acid encoding a selection marker disrupted by a synthetic intron having at least 80% sequence identity to a nucleic acid sequence set forth in SEQ ID NO: 31, wherein the synthetic intron comprises a defective 5' donor splice site that is correctable by a gene editor;(b) a splice-intron guide RNA (sigRNA) or a nucleic acid encoding a splice-intron guide RNA comprising a spacer sequence that has complementarity to the defective 5' donor splice site of the synthetic intron of (a), and a trans-activating CRISPR RNA (tracrRNA);(c) a gene editor or a nucleic acid encoding a gene editor;(d) a guide RNA or a nucleic acid encoding a guide RNA comprising a spacer sequence comprising a region of complementarity to a target gene of interest, wherein gene editing-dependent correction of the defective 5' donor splice site results in the enrichment of editing of the target gene of interest.

[0136] In certain embodiments, the gene editor is a nucleic acid programmable gene editor. In certain embodiments, the nucleic acid programmable gene editor is a base editor or a prime editor. Accordingly, the sigRNA may be a guide RNA molecule that is suitable for base editor or prime editing. In certain embodiments, the sigRNA is suitable for base editing. In certain embodiments, the sigRNA is suitable for prime editing.

[0137] Any of the polynucleotides, polynucleotide cassettes, sigRNAs, or gene editors described herein may be used in the systems of the present disclosure.B1195.70207WO00 39 / 102Methods and Uses

[0138] In another aspect, the present disclosure provides methods for enriching for base editing in a cell. For example, the methods for enriching for base editing in a cell use a polynucleotide cassette comprising a coding sequence of a selection marker gene that is disrupted by a synthetic intron sequence to prevent gene expression of the selection marker. The synthetic intron further comprises a defective 5 '-splice donor site that is correctable by a single base edit (e.g., by a base editor) such that when the 5'-splice donor site is corrected by a base editor, the intron is removed by the endogenous slicing system of the cell and expression of the selection marker gene can occur. The presence of active base editors within the cell is then determined by subjecting the cell to the selection pressure of the restored selection marker. Cell survival indicates the desired base edit was incorporated at the 5 '-splice donor site of the synthetic intron and indicates the cells are enriched with active gene editors (e.g., base editors or prime editors).

[0139] In some embodiments, such a method comprises the steps:(a) introducing into a cell:(i) a nucleic acid encoding a selection marker disrupted by a synthetic intron having at least 80% sequence identity to a nucleic acid sequence set forth in SEQ ID NOs.: 31, wherein the synthetic intron comprises a defective 5' donor splice site that is correctable by a nucleic acid programmable gene editor;(ii) a splice-intron guide RNA (sigRNA) or a nucleic acid encoding a spliceintron guide RNA comprising a spacer sequence that has complementarity to the defective 5' donor splice site of the synthetic intron of (i), and a trans-activating CRISPR RNA(tracrRNA);(iii) a gene editor or a nucleic acid encoding a gene editor; and / or (iv) a guide RNA or a nucleic acid encoding a guide RNA comprising a spacer sequence comprising a region of complementarity to a target gene of interest;(b) subjecting the cell to the selection pressure of the selection marker of (i), wherein cell survival indicates the gene editor made the desired edit at the splice donor site; and / or (c) sequencing the target DNA sequence to confirm presence of edit.

[0140] In certain embodiments, the gene editor of (ii) can form a complex with the guide RNA and the splice-intron guide RNA. In certain embodiments, the ability for the base editor to correct the defective 5' splice donor site in the synthetic intron, correlates to the base editing of the target gene.B1195.70207WO00 40 / 102

[0141] In another embodiment, the method may comprises an initial step of (a) introducing to the cell a nucleic acid encoding the gene editor of (ii), the guide RNA of (iv), and a second selection marker, and (b) subjecting the cell to the selection pressure of the second selection marker, wherein cell survival indicates the presence of the nucleic acid encoding the gene editor is in the cell. The present disclosure contemplates the use of this initial step to determine the presence of the nucleic acid encoding the base editor within the cell to be used as a comparison metric to the presence of active base editors in the cell. See FIG. 3C.

[0142] In certain embodiments, the gene editor is a base editor or a prime editor.

[0143] Any of the polynucleotides, polynucleotide cassettes, sigRNAs, or gene editors described herein may be used in the methods of the present disclosure. Accordingly, the sigRNA may be a guide RNA molecule that is suitable for base editor or prime editing. In certain embodiments, the sigRNA is suitable for base editing. In certain embodiments, the sigRNA is suitable for prime editing.Nucleic acids, Vectors, Cells, and Kits

[0144] The present disclosure provides, in some aspects, nucleic acids and vectors encoding any of the polynucleotides (e.g., sigRNAs), complexes (e.g., sigRNAs and base editors), and / or systems described herein. In some aspects, the present disclosure provides nucleic acids and vectors encoding a polynucleotide and a base editor as disclosed herein. In some embodiments, the nucleic acids and vectors provided herein comprise DNA (e.g., plasmid DNA or viral DNA). In some embodiments, the nucleic acids and vectors provided herein comprise RNA (e.g., mRNA or viral RNA).

[0145] Cells that may contain any of the polynucleotides (e.g., sigRNAs), vectors, complexes (e.g., sigRNAs and base editors), and / or systems described herein are also provided by the present disclosure. The methods described herein may be used to deliver a sigRNA and base editor into a eukaryotic cell (e.g., a mammalian cell, such as a human cell). In some embodiments, the cell is in vitro (e.g., a cultured cell). In some embodiments, the cell is in vivo (e.g., in a subject, such as a human subject). In some embodiments, the cell is ex vivo (e.g., isolated from a subject and may be administered back to the same or a different subject).

[0146] In some embodiments, a host cell is transiently or non-transiently transfected with one or more vectors described herein. In some embodiments, a cell is transfected as it naturally occurs in a subject. In some embodiments, a cell that is transfected is taken from a subject. In some embodiments, the cell is derived from cells taken from a subject, such as a cell line. InB1195.70207WO00 41 / 102some embodiments, a cell transfected with one or more vectors described herein is used to establish a new cell line comprising one or more vector-derived sequences. In some embodiments, a cell transiently transfected with the components of a base editing system as described herein (such as by transient transfection of one or more vectors, or transfection with RNA), and modified through the activity of a base editing complex, is used to establish a new cell line comprising cells containing the modification but lacking any other exogenous sequence.

[0147] The polynucleotides (e.g., splice-intron RNAs, gRNAs), nucleic acids, vectors, complexes, compositions, and / or cells, described herein may also be assembled into kits. In some embodiments, the kit comprises nucleic acids or vectors for expression of the polynucleotides (e.g., sigRNA), complexes (e.g., sigRNA and base editors), and / or systems described herein. In some embodiments, the kit comprises appropriate vectors for the expression of such polynucleotides to target the Cas9 protein of a base editor to a desired target sequence, e.g., a 5'-splice donor site of a synthetic intron (e.g., a variant of the SV40 intron). In some embodiments, the sigRNAs in the kit are useful for causing a mutation in a synthetic intron.

[0148] The kits described herein may include one or more containers housing components for performing the methods described herein, and optionally instructions for use. Any of the kits described herein may further comprise components needed for performing the base editing methods described herein. Each component of the kits, where applicable, may be provided in liquid form e.g., in solution) or in solid form, (e.g., a dry powder). In certain cases, some of the components may be reconstitutable or otherwise processible (e.g., to an active form), for example, by the addition of a suitable solvent or other species (for example, water), which may or may not be provided with the kit.

[0149] In some embodiments, the kits may optionally include instructions and / or promotion for use of the components provided. As used herein, “instructions” can define a component of instruction and / or promotion, and typically involve written instructions on or associated with packaging of the disclosure. Instructions also can include any oral or electronic instructions provided in any manner such that a user will clearly recognize that the instructions are to be associated with the kit, for example, audiovisual (e.g., videotape, DVD, etc.), Internet, and / or web-based communications, etc. The written instructions may be in a form prescribed by a governmental agency regulating the manufacture, use, or sale of pharmaceuticals or biological products, which can also reflect approval by the agency of manufacture, use, or sale for animal administration. As used herein, “promoted” includes all methods of doingB1195.70207WO00 42 / 102business including methods of education, hospital and other clinical instruction, scientific inquiry, drug discovery or development, academic research, pharmaceutical industry activity including pharmaceutical sales, and any advertising or other promotional activity including written, oral, and electronic communication of any form, associated with the disclosure. Additionally, the kits may include other components depending on the specific application, as described herein.

[0150] The kits may contain any one or more of the components described herein in one or more containers. The components may be prepared sterilely, packaged in a syringe, and shipped refrigerated. Alternatively, they may be housed in a vial or other container for storage. A second container may have other components prepared sterilely. Alternatively, the kits may include the active agents premixed and shipped in a vial, tube, or other container.

[0151] The kits may have a variety of forms, such as a blister pouch, a shrink-wrapped pouch, a vacuum sealable pouch, a sealable thermoformed tray, or a similar pouch or tray form, with the accessories loosely packed within the pouch, one or more tubes, containers, a box, or a bag. The kits may be sterilized after the accessories are added, thereby allowing the individual accessories in the container to be otherwise unwrapped. The kits can be sterilized using any appropriate sterilization techniques, such as radiation sterilization, heat sterilization, or other sterilization methods known in the art. The kits may also include other components, depending on the specific application, for example, containers, cell media, salts, buffers, reagents, syringes, needles, a fabric, such as gauze, for applying or removing a disinfecting agent, disposable gloves, a support for the agents prior to administration, etc.EXAMPLES

[0152] In order that the present disclosure may be more fully understood, the following examples are set forth. The synthetic and biological examples described in this application are offered to illustrate the nucleic acids, polypeptides, compositions, methods, uses, and systems provided herein and are not to be construed in any way as limiting in their scope.

[0153] Base editing (BE) is a CRISPR-based technology that enables high-throughput, nucleotide-level functional interrogation of the genome, which is essential for understanding the genetic basis of human disease and informing therapeutic development. BE screens have emerged as a powerful experimental approach, yet significant cell-to-cell variability in editing efficiency introduces noise that may obscure meaningful results. The present disclosure addresses this limitation by describing a co-selection method that enriches for cells with high base editing activity, substantially increasing editing efficiency at a target locus.B1195.70207WO00 43 / 102Example 1. Cas9-NG BE tiling of TP 53

[0154] This example describes that the canonical tumor suppressor gene TP53 encodes the protein p53, which serves a variety of growth control functions, including DNA damage repair and cell cycle regulation via apoptosis (Wang et al., 2023). p53 mutations occur in 50-60% of cancers (Baugh et al., 2018), making critical the functional characterization of genetic variants of this gene.

[0155] A base editing screen was conducted to detect gain-of-function and loss-of-function mutations along the length of TP53. The library was designed to include 923 NGG-PAM sigRNAs targeting TP53, 75 non-targeting sigRNAs, 75 intergenic sgRNAs, and 32 pan-lethal splice donor positive control sigRNAs (Hanna et al., 2021). A549 cells were first established to express Cas9-NG ABE8e or APOBEC with blasticidin resistance (FIG. 1A) (Richter et al., 2020; Neugebauer et al., 2023). The sigRNA library was then transduced in duplicate and selected on puromycin resistance for transduced cells. After selection was complete on day 7, the cells were split into three arms per base editor: no drug, Nutlin-3, and etoposide. The anti-cancer drug Nutlin-3 acts as an inhibitor of the p53 downregulator MDM2 (Arya et al., 2010), while etoposide induces DNA damage. Twenty-one days posttransduction, the sigRNA locus was sequenced for each condition.

[0156] After sequencing deconvolution, log fold changes were calculated from the pDNA for the dropout arm, and from the dropout condition for the drug arms. The average replicate z-scores were then calculated per sigRNA using the intergenic controls (FIG. IB, FIG. 2A). As expected, enrichment in the Nutlin-3 arms and depletion in the etoposide arms were mainly observed. In both drug conditions, the ABE arms showed greater magnitude of signal than the CBE arms.

[0157] Next, ten sigRNAs were selected with a range of enrichment and depletion from the ABE arms for validation (FIGs. 2B-2C). All sigRNAs selected can be edited with both base editors. To validate sigRNAs, the 10 single-guide constructs were transduced into A549s expressing Cas9- NG ABE8e or APOBEC, selected on puromycin, then split between no drug, Nutlin-3, and etoposide conditions on day 7. The sigRNA-targeted TP53 locus on days 7 and 21 were PCR amplified and Illumina sequenced and then allele frequencies were analyzed with CRISPResso2 (Clement et al., 2019). Using the most frequently observed edited allele in the drug arms, the early and late time points in terms of allele abundance relative to the pDNA were compared. From the primary screen it was expected sigRNA 1-3 would show enrichment of the most commonly edited allele in both drug arms, but onlyB1195.70207WO00 44 / 102enrichment for sigRNA 3 in the Nutlin-3 condition in the ABE arm was observed (FIG. 1C).In the CBE arm, enrichment in the Nutlin-3 condition for all 3 sigRNAs, and mild depletion of sigRNA 1 in the etoposide condition was observed (FIG. ID). sigRNA 4-10 was similarly assessed and found similar levels of variation from the primary screen.

[0158] The potential gain-of-function mutations seen from the primary screen in sigRNA 1 and 3 were further evaluated. Both sigRNAs mutate the TP53 start codon, possibly leading to the use of the alternative start site at amino acid 40 and creating the truncated isoform A40p53 (Hafsi et al., 2013). This isoform cannot bind the downregulator MDM2, and therefore may result in increased p53 expression levels, aiding cell survival. To functionally explore this Metl mutation, a Western blot of sigRNA 1-3 screened with ABEs was performed (FIG. IE). As expected, increased expression of the A40p53 isoform with both sigRNA 1 and 3 was observed, while sigRNA 2 -validated in the ABE arm as a D21G edit-showed no shortened isoform. Overall, base editor tiling of TP53 identified regions of interest such as start codon mutations.CBE optimization

[0159] Given the relatively poor performance of the CBEs compared with the ABEs, it was determined if an alternative Cas9 or CBE variant could increase editing efficiency and therefore screen signal. The performance of the PAM-flexible Cas9 variants Cas9-NG and Cas9-SpG with both APOBEC and TadCBEd in both A549s and MelJuSos was evaluated via an activity assay with an sigRNA targeting EGFP that introduces two stop codons into the coding sequence, disrupting fluorescence. Via flow cytometry on day 14 post-transduction, the highest EGFP editing with Cas9-SpG-TadCBEd across both cell types was observed (FIG. IF). TadCBEd out-performed APOBEC with both Cas9-SpG and Cas9-NG. Therefore Cas9-SpG as the PAM-flexible Cas variant was further tested.

[0160] With Cas9-SpG, the performance of three OT editors (TadCBEd, CBEmax, CBE-T1.52), two dual OT and A>G base editors (TadDE, CABE-T3.1), and one OG editor (CGBE) were benchmarked (FIG. 1G) (Neugebauer et al., 2023; Kurt et al., 2021; Lam et al., 2023; ). The EGFP activity assay was conducted in two cell lines, A549s and MelJuSos. On day 14, EGFP knockout was flowed for and the EGFP locus was Sanger sequenced for possible edits undetectable by flow. Across both cell types and readouts, consistently highest editing efficiency with TadCBEd was observed, with approximately 19% editing in A549s and 23% editing in MelJuSos as observed by flow cytometry. However, even with this most efficient CBE, over 75% of cells remained unedited in the post-selection population.B1195.70207WO00 45 / 102Developing a co-selection method to enrich for BE activity

[0161] The standard selection method for base editors, referred to as presence-based selection, relies on the co-expression of a drug resistance marker. This indirect approach does not account for expression heterogeneity, possibly due to differing levels of base editor and drug expression required for efficacy. A co- selection method was developed, referred to here as activity-based selection, that directly enriches for cells that actively base edit at the splice donor site of a selectable marker, rendering the marker functional.Primary screen

[0162] This direct co-selection method was developed by interrupting the puromycin resistance gene with a synthetic intron containing a defective splice donor site (FIG. 3A). This intron is derived from the SV40 intron, commonly spliced in human cells and having been shown to increase transgene expression (Xu et al., 2018). Additionally, the intron contains a stop codon in every reading frame that prevents translation of functional puromycin resistance protein. The splicing mechanism in mammalian cells relies on U1 small nuclear ribonucleoprotein (snRNP) recognition of a two-nucleotide sequence (GT) at the 5' end of the intron (Roca et al., 2013). Upon binding, additional spliceosome elements are recruited, allowing for targeted intron excision. The synthetic intron’s 5 '-splice donor site consists of either AT or GC, neither of which can be recognized by the U1 snRNP. However, when base editing occurs to restore the 5' splice site to the recognized GT, the U1 snRNP can recognize its target and initiate splicing, allowing for the translation of functional puromycin resistance protein. When placed under the selective pressure of puromycin, only cells that effectively base edit at the splice site can survive.

[0163] To identify guides that base edit and intron variants that splice efficiently, two libraries were designed containing 200 ABE and 200 CBE splice guides targeting corresponding sequences at the 5' splice site (FIG. 3B). For use with Cas9, guides are 20-nucleotides long, with 3 nucleotides targeting the exon portion of puromycin resistance and the other 17 targeting the intron. In order to maximize the likelihood of splicing, guides were designed using common motifs and variations derived from a sequence consensus collated from over 200,000 mammalian splice sites (Roca et al., 2013). To maximize the likelihood of base editing, CRISPick was used to select guides with high Rule Set 3 scores, balancing on-target efficiency with off-target activity predictions (Doench et al., 2016; Sanson et al., 2018). Finally, two additional libraries were designed — one for each BE — that contain theB1195.70207WO00 46 / 102corresponding intron sequences targeted by each splice guide. A secondary guide was also included upstream targeting the wild type 5' splice site of nonessential cell-surface marker CD274 (gene ID: PD -LI) in order to assess if base editing activity at one locus corresponds to editing activity at another locus. When base editing occurs on the antisense strand of CD274, the recognized GT splice site is disrupted and PD-L1 expression decreases.

[0164] The screen timeline differed for ABEs and CBEs given difficulty with sustained expression of TadCBEd (FIG. 3C). Each splice guide library (SGL) was screened in duplicate by each splice target library in duplicate (STL). For the ABE arm, ABE8e was transduced into A549s and selected with blasticidin, then the SGL was transduced and selected with hygromycin. For the CBEs, the SGL was first transduced and selected with hygromycin, then the TadCBEd was transduced and selected with blasticidin. Lastly for both arms, the STL was transduced, selected on puromycin, and the splice guide locus was sequenced. Log fold changes were then calculated relative to the pDNA and the log fold changes were z-scored relative to the whole population (FIGs. 3D and 3E). TadCBEd guidelevel LFC replicate arms correlated highly (r=0.65-0.95), while ABE replicates did not (r=0.2-0.41) (FIG. 2A). This difference is attributed to the differing timelines, as ABEs were removed from selective pressure for three weeks longer than CBEs, possibly allowing for an increase in the heterogeneity of ABE expression.Splice guide validation

[0165] For sigRNA validation, all sigRNA with at least one replicate z-score greater than or equal to 1.96 was nominated, resulting in 5 CBE and 6 ABE sigRNAs (FIGs. 4B and 4C).11 new constructs were assembled, each containing a unique splice guide and its corresponding intron sequence disrupting the puromycin resistance gene. Base editing at the CD274 splice site was evaluated in both presence- and activity-based selection conditions by staining and flowing for CD274 knockout. To help ensure that results would generalize, splice guide validation was performed in MelJuSo cells. After establishing the base editor lines on blasticidin, the 11 splice guide validation constructs were separately transduced, and cells were split between puromycin (activity-based selection) and no-drug (presence-based selection) conditions on day 4. On day 7, cells were dosed with interferon-gamma to induce CD274 expression. On day 9, after puromycin selection finished, cells were stained and flowed for CD274 with an APC fluorophore. For both conditions, a co-expressed protein, GFP, was gated on to capture transduced cells. When comparing presence- and activity-based selection methods, activity-based selection was observed to decrease the number of uneditedB1195.70207WO00 47 / 102cells in the final selected population from -40% to -10% with the most active guides (FIGs.3F and 3G). This corresponded to an increase in editing efficiency at the CD274 locus from -60% to -90%. The levels of editing of the activity-based selection condition was confirmed via Sanger sequencing of the CD274 locus with the EditR analysis tool (Kluesner et al., 2018) (FIG. 4F). Greatly increased editing at C6 along sgCD274, compared to C5, was observed; likely due to the lower likelihood of editing observed in TadCBEd when the cytosine follows an adenine (Neugebauer et al., 2023). The two guides were selected with the highest editing efficiency from each BE arm.Co-selection system modularity

[0166] The modularity of this co- selection system given that seventeen out of twenty nucleotides targeted by the splice guide lie within the intron sequence, anchored to the exon by just three nucleotides: CAG. Additionally, while the intron was inserted within the +2 reading frame of the puromycin resistance gene, given the presence of a stop codon within every reading frame, the intron can be inserted after a CAG within any selectable gene of interest and still use the previously validated splice guides. The intron was inserted in GFP (+1 reading frame) to evaluate if endogenous splicing may be occurring in the system. This construct was placed into cells, with and without ABE and CBEs, and observed minimal GFP positivity (<1%) in cells without base editors, indicating that the editing and splicing occurring was solely due to the exogenous mechanism (FIG. 4E). Insertion of the intron within GFP showed lower editing efficiencies of CD274 with both selection conditions compared to puromycin selection; with activity-based selection moderately decreasing the number of unedited cells in the final population (FIG. 4F). When the intron was inserted within the hygromycin resistance gene (+0 reading frame), similar levels of editing was observed as with the puromycin conditions and over a four-fold change decrease in unedited cells was observed with activity-based selection (FIG. 4G), highlighting the modularity of this approach.Evaluating activity-based selection with SpRY-Cas9

[0167] SpRY-Cas9, a nearly PAM-less variant of SpCas9 (Walton et al., 2020), shows decreased editing efficiencies compared to its more PAM-specific counterparts. To evaluate if activity-based selection could increase SpRY-Cas9 editing efficiency, two validated splice guide vectors per BE were transduced with a disrupted puromycin resistance gene into SpRY-Cas9-BE-expressing cells, and then CD274 editing was evaluated via flow cytometryB1195.70207WO00 48 / 102(FIG. 4H). Low levels of editing at the target locus (<20%) were observed and no increase in editing efficiency was observed with activity-based selection. This low efficiency editing may be due to a phenomenon seen previously in plant genome editing with SpRY-Cas9: selfediting (Ren et al., 2021). Due to its lack of a PAM requirement, SpRY-Cas9 sigRNAs can recognize and edit the integrated guide locus, creating a mutated guide that can no longer recognize its target sequence, thus resulting in decreased editing efficiency. To test this, the sgCD274 locus of the activity-based selection gDNA was Sanger sequenced and high levels of self-editing with both A>G and OT conversions were found. The sgCD274 locus was also sequenced with SpG-Cas9, which should not be able to edit given the PAM restrictions. Interestingly, some degree of self-editing with SpG-Cas9 at A4 and C6 was also observed, the most frequently edited nucleotide sequences as observed with previous Sanger sequencing results (FIG. 4D). The self-editing rate may be higher with activity-based selection given that by enriching for editing at the target loci, it is likely also enriching for editing at the integrated guide locus. Given the high rate of guide mutation, it is believed that self-editing inhibits the use of activity-based selection with SpRY-Cas9.Pooled screening with TP53

[0168] Given the success of SpG-Cas9 base editor co-selection at increasing editing efficiency at a single locus, this increased efficiency was further tested to see if it generalized across a library of guides. TP53 was re-screened with a 3,094-guide library to compare this selection approach with traditional presence-based selection. This library, designed using BEAGLE (Hanna et al., 2021), includes 2,785 NNNN-PAM guides tiling along the length of TP53, 206 intergenic controls, and 103 pan-lethal splice site positive controls. This library was cloned into two vectors: a presence-based selection vector with an intact puromycin resistance gene (FIG. 5B) and an activity-based selection vector with an intron-disrupted puromycin resistance gene and its corresponding splice guide (FIG. 5C). To conduct this screen, A549 cells containing SpG-Cas9-ABE8e or -TadCBEd on blasticidin were first established (FIG. 5A). The TP53 tiling library was then transduced and selected on puromycin starting on day 2 for the presence-based selection version and day 4 for the activity-based selection versions to allow time for editing to occur at the puromycin resistance intron splice site. Upon finishing puromycin selection, cells were split between etoposide and no-drug conditions and passaged for 19 days after initial drug selection. A negative- selection screen was used given the decreased sensitivity of positive-selectionsB1195.70207WO00 49 / 102screens due to growth- advantage mutations overtaking the population. Finally, the guide library was sequenced.Developing a self-editing correction pipeline

[0169] Upon deconvolution of the next generation sequencing data, an approximately 20% lower mapping rate was observed between the observed and expected sigRNA sequences with activity-based selection than with presence-based selection (FIG. 6A). This high rate of unmapped sequences suggested mutation of the integrated guides, possibly again due to selfediting. This may be possible if the sigRNA forms a one-nucleotide bulge to recognize the first guanine of the tracrRNA sequence as the middle guanine of the NGN PAM used by SpG-Cas9 (FIG. 6B); alternatively, GTT, the beginning of the tracrRNA sequence, may serve as a PAM for SpG-Cas9 under these conditions. To further evaluate this phenomenon, a raw reads file was generated with the read counts from the top one million unexpected sequences. These unexpected sequences were mapped back to their original guides given any A>G or C>T edits. The percent of unexpected reads out of the total reads for each guide were calculated, and it was observed that activity-based selection results in substantially higher self-editing rates overall (FIG. 6C). For each guide, the percent of base-edited nucleotides (G or T) out of the total reads for that nucleotide along the sigRNA were also calculated. Low rates of editing in the presence-based selection condition were observed, indicating that typical selection methods do not result in notable self-editing. This editing occurs mainly within the expected base editor window of 4-8 nucleotides along the sigRNA (FIG. 5C), further supporting the self-editing theory. Activity-based selection results in higher selfediting rates due to enriching for active base editors both at the endogenous and the integrated guide loci.

[0170] To account for self-editing, an in-silico correction to add the read counts from the unexpected sequences to the read counts of their mapped original guides was applied. Logfold changes of the drop-out arms relative to the pDNA were then calculated, and log-fold changes of the etoposide arms relative to the drop-out arms were calculated. When comparing pre-correction to post-correction Pearson correlations of the etoposide arms of each condition, it was observed that the in-silico correction increased correlations (FIG. 6D) in all conditions, but the most increase with the activity-based selection conditions was observed as expected given their higher rates of self-editing.B1195.70207WO00 50 / 102TP53 screen results

[0171] After applying the self-editing correction, the log-fold changes were z-scored in the etoposide arm relative to the intergenic controls. Analysis was restricted to guides using an NGN PAM given that this library was screened with SpG-Cas9. The median residue target was then targeted given all possible A>G or C>T edits within the 4-8 nucleotide base editing window. When comparing the etoposide arm z-scores in presence- versus activity-based selection conditions relative to the intergenic control sigRNAs, substantially stronger signals were observed, including both enrichment and depletion, within the activity-based selection conditions (FIG. 5D). Whereas the range of z-scores in the presence-based selection conditions is -12.5 to 4.2 and -17.4 to 8.4 in the ABEs and CBEs, respectively, activity-based selection showed a wider range of -36.5 to 15.9 for ABEs and -52.9 to 17.7 for CBEs.Overall, both selection methods similarly delineated positive versus negative controls as shown in ROC-AUC plots, showing ABE activity-based selection outperformed presencebased selection (r=0.88 vs. r=0.82) and CBE presence-based selection slightly outperformed activity-based selection (r=0.76 vs. r=0.74) (FIG. 6E). Additionally, the no-drug arm guide distributions showed a wider dynamic range with the positive control splice-site guides, while the intergenic controls showed little variation between conditions (FIG. 6F).

[0172] To evaluate if activity-based selection was erroneously amplifying signal, the distribution of guides in both selection conditions were compared and observed similar patterns, but with stronger magnitude of signal for activity-based selection (ABE r= 0.89, CBE r=0.90) (FIGs. 6G and 6H). Additionally, guides were divided into their predicted mutation bins based on BEAGLE predictions (Hanna et al., 2021). For mutations that were expected to show loss-of-function phenotypes, such as missense and splice site-targeting mutations, significant depletion in the activity-based selection condition was observed compared to the presence-based selection condition (FIGs. 5E and 5F). For mutations which were not expected to see a growth phenotype, such as silent or no edit mutations, no significant difference was observed between the selection methods, ultimately demonstrating that activity-based selection increases signal-to-noise ratio in pooled screens.Discussion

[0173] Base editing screens offer a scalable approach to interrogating endogenous genetic elements, but variation in editing activity on a cell-to-cell basis limits screen resolution. The present disclosure addresses this limitation and presents a splice-based selection method to enrich for actively editing cells, reducing the number of unedited cells from over 40% to lessB1195.70207WO00 51 / 102than 10% at a single locus. Additionally, the method demonstrated an increased signal-to-noise in a pooled, negative selection screen tiling TP53, identifying specific point mutations and domains of functional importance. Given the modularity of this splice guide and intron system, activity-based selection can be easily incorporated into most base editor experiments to more easily and clearly detect potential gain and loss-of-function mutations. Beyond drug selection, for example, the intron could be inserted within a fluorescent reporter such that actively-editing cells could be enriched using FACS. The present disclosure contemplates that the presented approach will also be especially useful for models that rely on transient base editor delivery, such as via protein or mRNA, as such non-integrating delivery methods currently have no straightforward means of selecting for cells that received sufficient levels of base. Activity-based selection mitigates this transient delivery challenge for base editors by creating a stable selection marker. The present disclosure also contemplates the use of this approach with other Cas proteins by altering the PAM sequence in the synthetic intron.

[0174] The present disclosure also contemplates that the increased sensitivity of the base editing screen may serve as a valuable tool for uncovering the functional roles of poorly characterized genomic regions, including non-coding areas with regulatory potential and underexplored genes. The base editing screen described herein can easily be applied to models that have proven scalable for genome- wide interrogation via other CRISPR techniques, such as knockout, interference, or activation, enabling the investigation of tens to hundreds of thousands of guide RNAs. Further, screens focused on a specific gene or sets of genes can identify drug resistance mechanisms, map functional domains, and rare gain-of-function mutations. Following a base editor screen, genomic regions of interest can be further refined with more precise editing technologies that are best deployed at focused loci, such as saturation genome editing or prime editing.References for Example 11. Anzalone, Andrew V., Peyton B. Randolph, Jessie R. Davis, Alexander A. Sousa, Luke W. Koblan, Jonathan M. Levy, Peter J. Chen, et al. 2019. “Search-and-Replace Genome Editing without Double-Strand Breaks or Donor DNA.” Nature 576 (7785): 149-57.2. Araya, Carlos L., and Douglas M. Fowler. 2011. “Deep Mutational Scanning:Assessing Protein Function on a Massive Scale.” Trends in Biotechnology 29 (9): 435-42.B1195.70207WO00 52 / 1023. Arya, A. K., A. El-Fert, T. Devling, R. M. Eccles, M. A. Aslam, C. P. Rubbi, N.Vlatkovic, et al. 2010. “Nutlin-3, the Small-Molecule Inhibitor of MDM2, Promotes Senescence and Radiosensitises Laryngeal Carcinoma Cells Harbouring Wild-Type P53.” British Journal of Cancer 103 (2): 186-95.4. Baugh, Evan H., Hua Ke, Arnold J. Levine, Richard A. Bonneau, and Chang S. Chan.2018. “Why Are There Hotspot Mutations in the TP53 Gene in Human Cancers?” Cell Death and Differentiation 25 (1): 154-60.5. Cabrera, Alan, Hailey I. Edelstein, Fokion Glykofrydis, Kasey S. Love, Sebastian Palacios, Josh Tycko, Meng Zhang, et al. 2022. “The Sound of Silence: Transgene Silencing in Mammalian Cell Engineering.” Cell Systems 13 (12): 950-73.6. Chew, Wei Leong, Mohammadsharif Tabebordbar, Jason K. W. Cheng, Prashant Mali, Elizabeth Y. Wu, Alex H. M. Ng, Kexian Zhu, Amy J. Wagers, and George M. Church. 2016. “A Multifunctional AAV-CRISPR-Cas9 and Its Host Response.” Nature Methods 13 (10): 868-74.7. Clement, Kendell, Holly Rees, Matthew C. Canver, Jason M. Gehrke, Rick Farouni, Jonathan Y. Hsu, Mitchel A. Cole, et al. 2019. “CRISPResso2 Provides Accurate and Rapid Genome Editing Sequence Analysis.” Nature Biotechnology 37 (3): 224-26. 8. Findlay, Gregory M., Evan A. Boyle, Ronald J. Hause, Jason C. Klein, and Jay Shendure. 2014. “Saturation Editing of Genomic Regions by Multiplex Homology- Directed Repair.” Nature 513 (7516): 120-23.9. Hafsi, Hind, Daniela Santos-Silva, Stephanie Courtois-Cox, and Pierre Hainaut. 2013.“Effects of A40p53, an Isoform of P53 Lacking the N-Terminus, on Transactivation Capacity of the Tumor Suppressor Protein P53.” BMC Cancer 13 (1): 134.10. Henrie, Alex, Sarah E. Hemphill, Nicole Ruiz-Schultz, Brandon Cushman, Marina T.DiStefano, Danielle Azzariti, Steven M. Harrison, Heidi L. Rehm, and Karen Eilbeck.2018. “ClinVar Miner: Demonstrating Utility of a Web-Based Tool for Viewing and Filtering ClinVar Data.” Human Mutation 39 (8): 1051-60.11. Jinek, Martin, Krzysztof Chylinski, Ines Fonfara, Michael Hauer, Jennifer A. Doudna, and Emmanuelle Charpentier. 2012. “A Programmable Dual-RNA-Guided DNA Endonuclease in Adaptive Bacterial Immunity.” Science (New York, N.Y.) 337 (6096): 816-21.12. Kennedy, Patrick H., Amin Alborzian Deh Sheikh, Matthew Balakar, Alexander C.Jones, Meagan E. Olive, Mudra Hegde, Maria I. Matias, et al. 2024. “Post-B1195.70207WO00 53 / 102Translational Modification-Centric Base Editor Screens to Assess Phosphorylation Site Functionality in High Throughput.” Nature Methods 21 (6): 1033-43.13. Kluesner, Mitchell G., Derek A. Nedveck, Walker S. Lahr, John R. Garbe, Juan E.Abrahante, Beau R. Webber, and Branden S. Moriarity. 2018. “EditR: A Method to Quantify Base Editing from Sanger Sequencing.” The CR1SPR Journal 1 (June): 239- 50.14. Kurt, Ibrahim C., Ronghao Zhou, Sowmya Iyer, Sara P. Garcia, Bret R. Miller, Lukas M. Langner, Julian Griinewald, and J. Keith Joung. 2021. “CRISPR C-to-G Base Editors for Inducing Targeted DNA Transversions in Human Cells.” Nature Biotechnology 39 (1): 41-46.15. Lam, Dieter K., Patricia R. Feliciano, Amena Arif, Tanggis Bohnuud, Thomas P.Fernandez, Jason M. Gehrke, Phil Grayson, et al. 2023. “Improved Cytosine Base Editors Generated from TadA Variants.” Nature Biotechnology 41 (5): 686-97.16. Li, Haoxin, Tiantai Ma, Jarrett R. Remsberg, Sang Joon Won, Kristen E. DeMeester, Evert Njomen, Daisuke Ogasawara, et al. 2023. “Assigning Functionality to Cysteines by Base Editing of Cancer Dependency Genes.” Nature Chemical Biology 19 (11): 1320-30.17. Lue, Nicholas Z., Emma M. Garcia, Kevin C. Ngan, Ceejay Lee, John G. Doench, and Brian B. Liau. 2023. “Base Editor Scanning Charts the DNMT3A Activity Landscape.” Nature Chemical Biology 19 (2): 176-86.18. Martin-Rufino, Jorge D., Nicole Castano, Michael Pang, Emanuelle I. Grody, Samantha Joubran, Alexis Caulier, Lara Wahlster, et al. 2023. “Massively Parallel Base Editing to Map Variant Effects in Human Hematopoiesis.” Cell 186 (11): 2456- 2474.e24.19. Neugebauer, Monica E., Alvin Hsu, Mandana Arbab, Nicholas A. Krasnow, Amber N. McElroy, Smriti Pandey, Jordan L. Doman, et al. 2023. “Evolution of an Adenine Base Editor into a Small, Efficient Cytosine Base Editor with Low off-Target Activity.” Nature Biotechnology 41 (5): 673-85.20. Perner, Florian, Eytan M. Stein, Daniela V. Wenge, Sukrit Singh, Jeonghyeon Kim, Athina Apazidis, Homa Rahnamoun, et al. 2023. “MEN1 Mutations Mediate Clinical Resistance to Menin Inhibition.” Nature 615 (7954): 913-19.21. Ren, Qiurong, Simon Sretenovic, Shishi Liu, Xu Tang, Lan Huang, Yao He, Li Liu, et al. 2021. “P AM-Less Plant Genome Editing Using a CRISPR-SpRY Toolbox.” Nature Plants 7 (1): 25-33.B1195.70207WO00 54 / 10222. Richter, Michelle F., Kevin T. Zhao, Elliot Eton, Audrone Lapinaite, Gregory A. Newby, B. W. Thuronyi, Christopher Wilson, et al. 2020. “Phage- Assisted Evolution of an Adenine Base Editor with Improved Cas Domain Compatibility and Activity.” Nature Biotechnology 38 (7): 883-91.23. Roca, Xavier, Adrian R. Krainer, and Ian C. Eperon. 2013. “Pick One, but Be Quick:5' Splice Sites and the Problems of Too Many Choices.” Genes & Development 27 (2): 129-44.24. Serrao, Erik, and Alan N. Engelman. 2016. “Sites of Retroviral DNA Integration:From Basic Research to Clinical Applications.” Critical Reviews in Biochemistry and Molecular Biology 51 (1): 26-42.25. Shao, Lipei, Rongye Shi, Yingdong Zhao, Hui Liu, Alexander Lu, Jinxia Ma, Yihua Cai, et al. 2022. “Genome-Wide Profiling of Retroviral DNA Integration and Its Effect on Clinical Pre-Infusion CAR T-Cell Products.” Journal of Translational Medicine 20 (1): 514.26. Walton, Russell T., Kathleen A. Christie, Madelynn N. Whittaker, and Benjamin P.Kleinstiver. 2020. “Unconstrained Genome Targeting with Near-PAMless Engineered CRISPR-Cas9 Variants.” Science (New York, N.Y.) 368 (6488): 290-96.27. Wang, Haolan, Ming Guo, Hudie Wei, and Yongheng Chen. 2023. “Targeting P53 Pathways: Mechanisms, Structures and Advances in Therapy.” Signal Transduction and Targeted Therapy 8 (1): 1-35.28. Xu, Dan-Hua, Xiao-Yin Wang, Yan-Long Jia, Tian-Yun Wang, Zheng-Wei Tian, Xin Feng, and Yin-Na Zhang. 2018. “SV40 Intron, a Potent Strong Intron Element That Effectively Increases Transgene Expression in Transfected Chinese Hamster Ovary Cells.” Journal of Cellular and Molecular Medicine 22 (4): 2231-39.Example 2. Activity-based selection for enhanced base editor mutational scanning

[0175] The functional characterization of genetic elements is one of the core challenges of modern biology, leading to the development of numerous experimental approaches. Deep mutational scanning (DMS) relies on overexpression of a library of exogenous open reading frames containing all variants of a gene of interest, allowing for precise mutational interrogation1,2. DMS, however, often examines genes outside of their native context and does not account for natural transcriptional and splicing regulation. This approach is also often limited by the need for cell line engineering to remove endogenous gene expression3. Likewise, massively parallel reporter assays (MPRAs), used to interrogate variants of non-B1195.70207WO00 55 / 102coding regions, also remove genetic elements from their endogenous context4 6. Saturation genome editing (SGE) addresses this concern by enabling precise endogenous mutations via homology-directed repair with oligonucleotide templates containing the edit of interest7,8. Efficiency limitations, however, have generally required haploid cell models, limiting the generalizability and scalability of this technology. More recently, prime editing has been developed to install a range of precise mutations-insertions, deletions, or single-base pair changes-but currently operates at low efficiency, limiting its use in high-throughput screens912

[0176] Base editing offers a high-throughput, endogenous alternative in which an adenine or cytosine deaminase is fused to a Cas protein, enabling A>G or C>T editing within a window of approximately four nucleotides13 |. Base editing has been used for a variety of genome scanning purposes, including interrogating drug resistance mechanisms, post-translational modifications, and non-coding regulatory elements16 22. Base editing screens, however, are currently hindered by variation in editing activity across a population of cells. In large-scale screens, base editors are frequently transduced into cells via lend virus, resulting in a semirandom locus of integration that may impact editing activity depending on the genomic context23,24. While it may be possible to reduce this heterogeneity via single-cell cloning to generate stable base editor cell lines, this process is time-consuming, costly, and labor-intensive. Additional factors such as transgene silencing or the cellular immune response to Cas9 may also contribute to this heterogeneity25,26. Regardless of the underlying cause, this variation in activity adds noise to screens given that some cells edit highly efficiently and others not at all. Traditional indirect selection methods fail to account for this variation, limiting the resolution of base editing screens and their ability to distinguish biological signal. Presented herein is an activity-based co-selection method to improve the resolution of base editing screens.ResultsCas9-NG base editor tiling of TP53

[0177] To benchmark the high-throughput mutational scanning ability of base editing, a tiling screen was conducted to identify loss-of-function mutations along the highly- studied gene TP5327~31. Encoding the tumor suppressor protein p53, TP 53 serves a variety of growth control functions, including DNA damage repair and cell cycle regulation via apoptosis32. The A549 lung cancer cell line enables both positive and negative selection screens for p53 activity with Nutlin-3 and etoposide, respectively. Nutlin acts as an inhibitor of the p53B1195.70207WO00 56 / 102downregulator MDM2, leading to the stabilization of p53 and enrichment of p53 loss-of-function mutations under normal growth conditions33-35. Conversely, etoposide induces DNA damage such that loss-of-function mutations deplete given the decreased ability of p53 to induce DNA damage repair34,36’37.

[0178] A base editing library was designed to include 923 single guide RNAs (sgRNAs) targeting TP53, 75 non-targeting sgRNAs, 75 intergenic sgRNAs, and 32 sgRNAs targeting splice donors of pan-lethal genes to serve as positive controls, all using NGN PAMs or known-active PAMs with Cas9-NG38. These libraries were cloned into all-in-one base editor lentiviral vectors that contain the PAM-flexible variant Cas9-NG fused to either the adenine base editor (ABE) or cytosine base editor (CBE) machinery, using the ABE8e and rat APOB EC deaminases, respectively, and a 2A peptide links a puromycin resistance cassette (FIG. 1A)15’39’40. The sgRNA tiling library was transduced into A549 cells in duplicate and selected with puromycin, then split between no drug, Nutlin, and etoposide conditions. After two weeks of drug treatment, the integrated sgRNA library was retrieved via PCR of the genomic DNA.

[0179] After sequencing deconvolution, log fold changes (LFCs) were calculated from the pDNA for the no drug (dropout) arm, and for Nutlin and etoposide from the no drug arm. Z-scores were then calculated relative to the intergenic control sgRNAs and averaged between replicates (FIG. IB, FIG. 2A). Much of the observed signal fell within known TP53 domains, especially the DNA binding domain (FIG. 2E). Notably, in both drug conditions, ABE showed greater magnitude of signal than CBE. Furthermore, most sgRNAs were negatively correlated across the two conditions, enriching with Nutlin and depleting with etoposide, although some sgRNAs deviated from this pattern (FIG. 2D). Three (3) such sgRNAs were selected for validation, hereby referred to as sgTP53s: sgTP53 1 and 3, predicted to generate a start codon MetlThr mutation with ABE, and sgTP53 2, predicted to generate an Asp21Gly mutation with ABE.

[0180] The selected 7P53-targeting sgRNAs were individually delivered into A549 cells expressing Cas9-NG-ABE8e, selected on puromycin, then split between no drug and Nutlin arms (for sgTP53 sequences, shown in Table A).Table A.B1195.70207WO00 57 / 102

[0181] The identity of the mutations introduced by these sgRNAs were verified by PCR and sequencing of the target locus, analyzing the results with CRISPResso2 software41. Validation of sgTP53 2 (FIG. 2F) shows that it introduces a Asp21Gly mutation that depletes in the presence of Nutlin; Asp21 is both part of the transactivation domain and the region known to interact with MDM2. For sgTP53 1 and 3, MetlThr alleles enrich in the presence of Nutlin, indicating an advantageous mutation (FIG. 2F). Previous work has shown that utilization of an alternative downstream start codon at amino acid 40 creates a truncated isoform, A40p53, that cannot bind MDM2 and that lacks most of the transactivation domain (TAD), which is consistent with the enrichment phenotype observed in the screen42,43. To explore the effect of these base edits at the protein level, a Western blot was performed with early and late timepoints (FIG. IE). An increase was observed in a shorter p53 isoform for sgRNAs predicted to cause a Metl mutation (sgTP53 1 and sgTP53 3), which was consistent with the hypothesis that mutations at the canonical start codon result in the use of an alternative start site.

[0182] CBE optimization

[0183] Given the relatively poor performance of the APOBEC-based CBE compared with the ABE in the prior screen, it was asked whether alternative Cas9 or CBE variants could increase editing efficiency. Of particular interest were CBEs evolved from ABEs, such as TadCBEd, which have been shown to have higher on-target activity and lower off-target activity than APOBEC15. Using the Fragmid vector assembly system44, the performance of the PAM-flexible Cas9 variants Cas9-NG and SpG-Cas945were evaluated with APOBEC and TadCBEd in A549 and MelJuSo cells. An activity assay was conducted in which GFP is targeted by an sgRNA that can introduce two stop codons into the coding sequence, disrupting fluorescence. The highest GFP depletion was observed with SpG-Cas9-TadCBEd across both cell types (FIG. IF). With both CBEs tested, SpG-Cas9 resulted in greater GFP disruption than Cas9-NG, and therefore SpG-Cas9 was chosen as the PAM-flexible Cas variant.

[0184] With SpG-Cas9, the performance of three T editors (TadCBEd, CBEmax, CBE-T1.52), two dual OT and A>G editors (TadDE, CABE-T3.1), and one OG editor (CGBE)B1195.70207WO00 58 / 102were then benchmarked (FIG. 1G)15,46~48; given that TadCBEd outperformed APOB EC in the previous experiment, APOB EC was not included in this second round of experimentation. With the same GFP activity assay construct, the highest editing efficiency was consistently observed with TadCBEd, with approximately 20% editing in A549 cells and 30% editing in MelJuSo cells, as determined via Sanger sequencing of the target locus and analysis with EditR49. Therefore SpG-Cas9-TadCBEd was selected as the CBE for future experiments. Even with this more efficient CBE, however, approximately 70% of cells remained unedited in the post-selection population.Developing a co-selection method to enrich for base editor activity

[0185] A standard method for establishing base editor cell models relies on a co-expressed drug resistance or fluorescent marker, an indirect measure of base editor activity hereafter referred to as presence-based selection. This approach does not necessarily result in a high fraction of editing activity as demonstrated above, and the presence of cells in a pooled population that contain an sgRNA but do not edit the target site obscures signal. Inspired by co- selection approaches that enrich for cells with increased homology-directed repair efficiency50,51, a co-selection method was developed, referred to as activity-based selection, that directly enriches for cells that have base editing activity at the splice donor site of a selectable marker, rendering the marker functional.

[0186] This base editing co-selection method was initially developed by interrupting the puromycin resistance gene with a synthetic intron containing a defective splice donor site (FIG. 3A). This sequence was derived from the SV40 intron, which is efficiently spliced in human cells and has been shown to increase transgene expression52. As the canonical splice donor site is GT, the 5’ splice site sequence was designed to comprise either AT or GC, which can be edited by an ABE or CBE, respectively, to restore splicing. The splice donor sequence corresponds to nucleotides 4 and 5 of the sgRNA, as these positions are highly effective for base editing. The total variable target site is 20-nucleotides long, with 3 nucleotides targeting the exon portion of the puromycin resistance gene and the other 17 targeting the intron, followed by an NGG PAM (FIG. 3H). In order to maximize the likelihood of splicing upon successful editing, variations of the SV40 intron were designed using common motifs derived from a sequence consensus collated from over 200,000 mammalian splice sites53,54. Further, the intron variants avoid adenines and cytosines in positions 6-9 to mitigate the possible confounder of multiple edits. To maximize the likelihood of base editing activity, these intron sequences were chosen such that their corresponding targeting sgRNAs have high Rule Set 3B1195.70207WO00 59 / 102scores as determined by CRISPick55. The intron also contains a stop codon such that unspliced transcripts will result in premature termination of translation. With these parameters, 200 sequences were selected for each base editor to compose the intron variant libraries.

[0187] Accordingly, two more libraries were designed-one for each base editor - each containing 200 sgRNAs that target the corresponding intron sequences at the 5’ splice site of the puromycin resistance gene. These two splice- site-targeting sgRNA libraries were cloned into vectors that also include an sgRNA targeting the wildtype 5’ splice site of the nonessential cell surface marker CD274 (gene ID: PD-L1), in order to assess if base editing activity at one locus corresponds to base editing activity at a distinct endogenous locus (FIG.3H). When base editing occurs on the antisense strand of CD274, the 5’ splice site is disrupted and PD-L1 expression decreases.

[0188] Each combination of splice-site-targeting sgRNA library and corresponding intron variant library were screened in duplicate in A549 cells. For the ABE arm, stable expression of ABE8e was first established then the splice-site-targeting sgRNA library was transduced. For the CBE arm, the splice- site-targeting sgRNA library was first transduced and then TadCBEd was introduced. Lastly, for both arms, the corresponding intron library was transduced, selected with puromycin, and at the end of the screen PCR was used to retrieve the enriched splice-site-targeting sgRNAs (FIG. 3C, FIG. 4A). Replicates were more correlated for CBE (r=0.65-0.95) than ABE (r=0.2-0.41). This difference was attributed to the differing timelines, as the ABE was removed from selective pressure for three weeks longer than the CBE, possibly allowing for an increase in the heterogeneity of ABE expression.Splice-targeting sgRNA validation

[0189] To validate the screen and identify top-performing sgRNA - intron pairs, all sgRNAs with at least one replicate z-score greater than or equal to 1.96 were nominated, resulting in 6 ABE and 5 CBE sgRNAs (FIG. 31, Table B).Table B.B1195.70207WO00 60 / 102*(+ stop codons in all reading frames, used for hygro cassette)B1195.70207WO00 61 / 102

[0190] To minimize the chance of off-target activity, it was ensured that all nominated sgRNAs did not map to a perfect or 1 -nucleotide mismatch from another sequence in the genome by using Cas-OFFinder56. Each of the constructs contained a nominated splicetargeting sgRNA and its corresponding intron sequence disrupting the puromycin resistance gene, as well as an intact, upstream GFP marker. To evaluate the generalizability of this approach, this validation experiment was performed in MelJuSo cells.

[0191] After establishing stable expression of ABEs and CBEs in cells, the validation constructs were separately transduced and cells were divided between no drug (presencebased selection) and puromycin (activity-based selection) conditions. After selection, flow cytometry was performed to assess CD274 protein levels (FIG. 3J, FIG. 4J). For both selection conditions, GFP expression was gated to identify cells transduced with the sgRNA. When comparing presence- versus activity-based selection methods, it was observed that activity-based selection decreased the fraction of unedited cells in the final selected population from 41.5% to 9.9% with the most active ABE sgRNA, and from 46.6% to 11.0% with the most active CBE sgRNA (FIG. 3K). Across both base editors, activity-based selection resulted in a substantial decrease in the number of unedited cells in the final population for every sgRNA - intron pair nominated, indicating the robustness of this approach (FIGs. 3F-3G). The levels of editing of the activity-based selection condition were confirmed by sequencing the endogenous CD274 locus and analyzing per-base editing efficiency with the EditR analysis tool49(FIG. 4D).

[0192] The modular design of the co-selection system allows for easy adaptation to other selectable markers given that 17 out of 20 nucleotides targeted by the splice sgRNA lie within the intron sequence, anchored to the exon by just three nucleotides: CAG. To demonstrate this flexibility, the intron was inserted within the hygromycin resistance gene and observed similarly-improved performance of activity-based selection with SpG-Cas9 base editors (FIG. 4G). Additionally, the intron was inserted into GFP to evaluate if cryptic splicing is occurring in the system, which would bypass the need for base editors. This construct was transduced into cells with and without base editors and observed minimal GFP positivity (<1%) in cells without base editors, indicating that base editing activity is required for splicing to occur (FIG. 4E).

[0193] It was next assessed if activity-based selection could also increase the editing efficiency of SpRY-Cas9, a nearly PAM-less Cas9 variant45'57,58. Using the previously validated sgRNAs, CD274 editing was evaluated by ABEs and CBEs via flow cytometry.B1195.70207WO00 62 / 102Most cells (>80%) remained unedited after selection, and no improvement was noticed with the activity-based selection approach (FIG. 4H). It was hypothesized that this low efficiency editing may be due to self-editing of the sgRNA, as previously observed59. Due to its lack of a PAM requirement, SpRY-Cas9 sgRNAs may recognize and edit the integrated guide cassette, creating a mutated sgRNA that can no longer recognize its target sequence, thus resulting in decreased editing efficiency. To test this hypothesis, Sanger sequencing of the integrated CD274 sgRNA expression cassette was performed. For comparison, the same was done for cells with SpG-Cas9, which should not be able to self-edit given PAM restrictions. For all samples, the most edited nucleotides fell within the canonical base editing window of 4 - 8 nucleotides along the integrated sgRNA. SpRY-Cas9 showed the highest self-editing at positions 4 and 6, with 87.5% A>G editing at nucleotide 4 along the guide and 86.0% C>T editing at nucleotide 6 (FIG. 41). Surprisingly, cells with SpG-Cas9 also showed detectable editing of the integrated sgRNA (editing atA4: 28.5%, C6: 38.0%). Given the high rate of sgRNA mutations, it is concluded that self-editing inhibits the use of activity-based selection with SpRY-Cas9 under these experimental conditions.Pooled screening with TP53

[0194] Given the success of SpG-Cas9 base editor co-selection at increasing editing efficiency at a single locus, it was next asked if this approach could generalize across a library of sgRNAs. TP53 was re-screened to directly compare this selection approach with traditional presence-based selection (FIG. 5A). A new library was designed to include 2,785 sgRNAs tiling along the length of TP53 with no PAM restriction (i.e. NNN PAM; it was hoped that SpRY-Cas9 would prove effective), 206 intergenic controls, and 103 sgRNAs targeting the splice sites of pan-lethal genes as positive controls, all of which utilize an NGG PAM. Although there exist numerous tools for predicting sgRNA on-target editing activity, the library was not filtered for predicted active sgRNAs in order to maximize sgRNA coverage across the gene. Some of these predictive approaches in Supplementary Note 1, below. This library was cloned into two vectors per BE: a presence-based selection vector with an intact puromycin resistance gene and an activity-based selection vector with an intron-disrupted puromycin resistance gene and its corresponding splice-targeting ABE or CBE sgRNA (FIG. 5B). The two sgRNA - intron pairs with the highest activity-based selection editing efficiency for each BE were selected (ABE: sgRNA 1, 4; CBE: sgRNA 2, 3, FIGs. 3F-3G ) and used these as quasi-replicates.B1195.70207WO00 63 / 102

[0195] To conduct this screen, A549 cells containing SpG-Cas9-ABE8e or -TadCBEd were first established on blasticidin. The TP53 tiling library was then transduced and selected on puromycin starting on day 2 for the presence-based selection version and day 4 for the activity-based selection versions, giving extra time for the latter to allow for editing to occur at the puromycin resistance intron. Following selection, cells were split between etoposide and no drug conditions; a negative-selection condition was selected because data quality of such screens is more reliant on high editing activity. After two weeks of additional passaging, cells were harvested and the sgRNA library was sequenced.

[0196] Upon sequencing deconvolution, an approximately 20% lower match rate was observed between the observed and expected sgRNA sequences with activity-based selection than with presence-based selection (FIG. 6A), which was attributed to self-editing, as observed above with SpG-Cas9. This may be possible if GTT, the beginning of the tracrRNA sequence, can serve as a PAM for SpG-Cas9 under these conditions. Alternatively, the sgRNA could form a one-nucleotide bulge, thereby shifting the first guanine of the tracrRNA sequence to be the middle guanine of an NGN PAM that is recognized by SpG-Cas9 (FIG.6B). Regardless of the mechanism, to further explore this phenomenon, the top one million unexpected sequencing reads were obtained and mapped back to their original sgRNAs given any A>G or OT editing. The percent of unexpected reads out of the total reads was then calculated for each sgRNA and it was observed that activity-based selection resulted in substantially higher self-editing rates overall (FIG. 6C). For each sgRNA, the percent of base-edited nucleotides (G or T) out of the total reads for that nucleotide across all screening conditions was also calculated. With both selection methods, the most-edited bases fell within the expected base editing window of 4 - 8 nucleotides along the sgRNA (FIG. 5C), further supporting the hypothesis of self-editing rather than sequencing error or other sources of technical noise. Much lower rates of self-editing were observed in the presence-based selection condition, indicating that typical selection methods do not result in substantial selfediting, which likely explains why this phenomenon has not, been previously reported with SpG-Cas9.

[0197] To account for this self-editing, a computational correction was applied to add the read counts from the unexpected sequences to the read counts of the original sgRNAs to which they mapped. The LFC of the no drug arms was then calculated relative to the pDNA, and the LFC of the etoposide arms relative to the no drug arms. It was observed that the computational correction improved replicate correlations (FIG. 6D), with the greatest increase occurring under the activity-based selection conditions, as expected given theirB1195.70207WO00 64 / 102higher rates of self-editing. The analysis was therefore continued using the corrected LFC values.

[0198] For sgRNAs targeting TP53, the subsequent analysis was restricted to sgRNAs with an NGN PAM given that this library was screened with SpG-Cas9. LFCs were averaged across replicates, but slightly stronger signal was observed with ABE sgRNA 1 and CBE sgRNA 2, leading to the recommendation of these options for future screens. Notably, the positive control pan-lethal-targeting sgRNAs showed a wider dynamic range of LFC values with activity-based selection compared to presence-based selection, while the negative control intergenic sgRNA distribution remained similar between methods (FIG. 6F). LFC values were then z-scored for sgRNAs targeting TP53, relative to the intergenic controls, and saw a substantially wider dynamic range of z-scores with activity-based selection relative to presence-based selection (FIG. 5D, FIG. 6G). Whereas the z-scores in the presence-based selection conditions range from -12.5 to 4.2 and -17.4 to 8.4 for ABE and CBE, respectively, activity-based selection showed a wider range of -36.5 to 15.9 for ABE and -52.9 to 17.7 for CBE. When discretized by predicted mutation bin, it was observed that sgRNAs predicted to introduce either a silent edit or no edit showed no significant difference between the selection regimes for both ABE and CBE, whereas missense (ABE) or splice site (CBE) mutations were significantly more depleted with activity-based selection compared to presence-based selection (FIGs. 5E-5F), indicating that activity-based selection did not indiscriminately increase the range of the screen but rather substantially increased the ratio of signal-to-noise.

[0199] The ability of base editing was then assessed in comparison to other genome engineering technologies to recapitulate known domains and mutations of TP53. The ability of BE was first compared against CRISPR knockout - another widely used method for endogenous high-throughput interrogation - to identify known functional domains of TP5360~64. With CRISPR knockout, for every insertion or deletion created, there is only, on average, a % chance that the repair outcome results in an in-frame mutation. Out-of-frame mutations, regardless of whether they fall within functional domains, will likely lead to a non-functional protein, and thus the overall resolution of this approach is limited by the need to distinguish the signal produced by % of mutations from the noise produced by the other %. To compare the performance of these approaches for identifying domains, the same TP53 library was screened in cells expressing SpG-Cas9-nuclease. The greatest difference in mean z-score was observed between domain versus interdomain regions for the base editor activity-based selection methods (ABE: x = 6.36; CBE: x = 6.27) (FIG. 8A). Comparatively, presencebased selection methods showed less distinction between domain and interdomain regionsB1195.70207WO00 65 / 102(ABE: x = 2.09; CBE: x = 3.02). Cas9-nuclease showed a similar distinction as the presencebased selection method between domain and interdomain regions (x = 2.94), but also showed depletion within interdomain regions (x = -1.68), likely due to frameshift mutations in interdomain regions disrupting the overall protein structure. Conversely, for base editors across all selection methods, the means of interdomain regions centered near 0 (ABE presence: x = -0.06; CBE presence: x = 0.42; ABE activity: x = -0.05; CBE activity: x = -0.03). Overall, base editing outperformed Cas9-nuclease at distinguishing domain versus interdomain regions, and activity-based selection more clearly delineated these functional regions compared to traditional selection methods.

[0200] Next, the accuracy of base editing selection methods to recapitulate known TP 53 variant effects was explored by evaluating guides that introduce ClinVar pathogenic or benign mutations. For presence and activity-based selection, respectively, 68% and 71% of pathogenic-associated guides deplete (z < -2) in contrast to 5% and 6% for benign- associated guides. The enhanced signal with activity-based selection becomes more apparent at more stringent z-score thresholds, with activity-based selection detecting 22% more pathogenic-associated guides at a z-score cutoff of -5 (FIG. 8B).

[0201] This performance was next compared to TP53 screens employing deep mutational scanning via ORF overexpression34and prime editing12, and saw that base editing screens with activity-based selection more effectively distinguished between pathogenic and benign variants (FIG. 8C). Notably, limited efficacy of prime editing guides was observed, potentially due to low editing efficiency; while performance improved when restricting the analysis to prime editing guide RNAs (pegRNAs) with editing efficiency above 25%, as detected via a sensor assay employed in that study, this filtering step eliminates 93.8% of guides from the analysis. Thus, base editing with the use of activity-based selection stands as a leading approach to reliably achieve high efficiency of variant introduction at endogenous loci.

[0202] Co-selection methods have been previously suggested for prime editing. Therefore, it was next explored whether adapting the activity-based selection approach could enrich for cells with active prime editors and increase signal-to-noise ratio65,66. Constructs with a puromycin-resistance-disrupted- splice- site-targeting engineered prime editing guide RNA (epegRNA) was designed, paired with a previously validated HEK3 epegRNA (Table C)67 fi9. Table C.B1195.70207WO00 66 / 102B1195.70207WO00 67 / 102

[0203] Each epegRNA includes a unique pseudoknot, either mpknot or evopreQl, to minimize deterioration69. When evaluated against the HEK3 epegRNA coupled with an intact puromycin resistance cassette, as would be used with presence-based selection, little difference in editing was observed at the HEK3 locus with the co-selection method, using both PE7 and PEmax (FIG. 6I)10,67. While further modifications - such as scaffold improvements, or extended culturing times - may somewhat enhance editing at the target locus, activity-based selection currently remains of limited utility for prime editing applications, at least for the conditions tested here.

[0204] To further evaluate the performance of base editor activity-based selection, two strongly-depleted sgRNAs were selected from the TP53 screen for additional validation. One of the mutations introduced (by sgTP53 4; predicted mutation: Tyrl26His) is designated by ClinVar as a variant with Conflicting Classifications of Pathogenicity, while the other mutation (by sgTP53 5; predicted mutation: Tyr236His) is designated as Pathogenic / Likely Pathogenic. Each sgRNA was introduced individually and the screening conditions were repeated, this time reading out the impact of the base editor by sequencing the target site for each sgRNA and examining changes in allelic abundance. For sgTP534 and 5, the wild type allele represented 1% and 2% of reads prior to etoposide (day 11), highlighting the high efficiency of activity-based selection, which then enriched to 56% and 54% of the population following etoposide treatment (day 26), corroborating the results from the primary screen (FIG. 8D). With sgTP534, it was observed that alleles that disrupted the canonical splice acceptor site (AG) deceased in abundance (93% to 23%), while the Tyrl26His mutation on its own increased in abundance (6% to 21%), although not to the same degree as the wildtype allele (1% to 56%), indicating that this variant is indeed deleterious to TP53 function. For sgTP53 5, the Tyr236His mutation is likewise less fit than the wildtype allele, consistent with its classification as Likely Pathogenic. Compound mutants, in which a Met237Thr mutation was also generated, were more deleterious. Although both individual sgRNAs created deleterious mutations as identified in the primary screen, base editors frequently introduce a variety of mutations, underscoring the importance of validation at the allelic level and theB1195.70207WO00 68 / 102value of examining specific mutations in a controlled setting via techniques such as saturation mutagenesis and prime editing.

[0205] Using these same sgRNAs, the mitigation of self-editing by using an alternative tracrRNA was also explored. If self-editing is occurring due to recognition of an NGN PAM using the 5' G of the tracrRNA, as previously described, an alternative 5' nucleotide of the tracrRNA may correct for this problem. An experiment was thereby conducted to evaluate the self-editing and on-target editing rates of the previously-used tracrRNA (5' G) against variant tracrRNAs that have a 5' A or 5' C70. Allelic level editing of the no drug arm was determined via NGS on genomic DNA extracted 26 days post-transduction. With both sgTP534 and 5, a drastic reduction in self-editing rate was observed with tracrRNAs that do not contain a G in the first position (FIG. 6J). However, a complete loss of editing efficiency was also observed at the target locus (~0% A>G editing with tracrRNA 5'A or C) (FIG. 6K). While further experimentation may discover a tracrRNA with low self-editing and high on-target editing activity, it’s recommend using the canonical 5’G tracrRNA and implementing the self-editing correction.Discussion

[0206] Base editing screens offer a scalable approach to interrogating endogenous genetic elements, but variation in editing activity on a cell-to-cell basis limits screen resolution. Presented herein is a splice-based selection method to enrich for actively editing cells, reducing the number of unedited cells from over 40% to less than 10% at a single locus. Additionally, increased signal-to-noise ratio in a pooled, negative selection screen tiling TP53 was demonstrated, identifying specific point mutations and domains of functional importance.

[0207] Previous base editing co-enrichment methods involve the design and optimization of sgRNAs specific to the selectable marker of interest, such as sgRNAs targeting a start codon of a specific gene, a GFP to BFP conversion, or Diphtheria Toxin receptor71 B. In contrast, the intron-based selection method disclosed herein is adaptable to almost any selection marker without requiring additional design and validation of sgRNAs, allowing for swift incorporation into most base editing experiments. For example, the intron could be inserted within a fluorescent reporter such that actively-editing cells could be enriched for using fluorescence-activated cell sorting. This approach could be especially useful for models that rely on transient base editor delivery, such as via protein or mRNA, which currently have no straightforward means of selecting for cells that received sufficient levels of the baseB1195.70207WO00 69 / 102editor19,74. Activity-based selection mitigates this transient delivery challenge for base editors by generating a stable selection marker. Additionally, simply by altering the PAM sequence in the synthetic intron, this approach could be repurposed for use with other Cas proteins.

[0208] One limitation of activity-based selection when using the SpG variant of Cas9 is the increased propensity for self-editing given the sequence requirements of the tracrRNA.Although a computational approach was developed to rescue the sequencing read counts from self-edited sgRNAs, it is noted that self-editing may generate novel guides and thus novel off-target sites. These sites, however, are of minimal concern because they no longer contain editable nucleotides. For example, with the use of an ABE, self-editing introduces Gs within the guide at positions 4 - 8. The guide may therefore match novel off-target sites with Gs in the editing window, but these are, by definition, not editable by an ABE. Further, while more apparent with activity-based selection, self-editing activity was also observed with traditional selection, underscoring the need to correct for this phenomenon when screening with PAM-flexible Cas9 varieties, lest the cells with the most-active base editing activity fall out of the analysis due to the sgRNA no longer being recognized during sequencing deconvolution. As an additional technical consideration, adequate pre-selection screening coverage is 2 - 3 times larger for the activity-based selection method than for traditional selection methods, as cells must also actively base edit to survive. Once selection is complete, screening can continue at lOOOx coverage, as is typical for pooled screens.

[0209] The increased sensitivity of this selection method will enhance the ability of base editors to uncover the functional roles of poorly characterized genomic regions, including underexplored genes and non-coding areas with regulatory potential. Base editing screens are easily applied to models that have been proven scalable for genome-wide interrogation by other CRISPR techniques, such as knockout, interference, or activation, enabling the investigation of tens to hundreds of thousands of sgRNAs, with no requirement that the target sites be contiguous, as is necessary for prime editing-based screens that read out the target site75. Further, screens focused on a specific gene or gene set can uncover drug resistance mechanisms, map functional domains, and identify rare gain-of-function mutations1820’76 7S. Following a base editing screen, genomic regions of interest can be further refined with more precise editing technologies that are best deployed at focused loci, such as saturation genome editing or prime editing. Together, this suite of tools offers great potential for understanding disease mechanisms and informing drug discovery efforts.B1195.70207WO00 70 / 102

[0210] Approaches for predicting on-target base editing activity There exist numerous approaches to predict a guide’s propensity for on-target activity. Herein, the Rule Set 3 Sequence score was leveraged, which was developed to optimize predictions of on-target activity in SpyoCas9 CRISPRko screens, to estimate guide potential for editing activity to select splice guides55. External works have introduced models that incorporate nucleotide motifs that influence base editing activity in an editor- specific manner79,80. Here, the ability of these models to distinguish effective guides in the presented TP53 tiling screen with etoposide was evaluated.

[0211] For this analysis, all ABE and CBE guides expected to introduce a missense, nonsense, or splice site mutation in TP53 were collated; guides that only introduced benign edits in the 3-9 window, as annotated by ClinVar, were excluded from this set. Since guides that introduce loss-of- function variants should deplete in the etoposide arm, define inactive guides were defined as those that do not deplete in this arm. While this method may falsely classify guides as inactive if they introduce benign missense mutations not yet annotated by ClinVar, this outcome is preferable to the alternative of the entire analysis being restricted to the small number of guides that introduce known pathogenic variants.

[0212] The ability for Rule Set 3 Sequence (RS3) scores to predict the likelihood of base editing activity55was first evaluated. Since this model was agnostic to the base editor in use, guides with conflicting categorizations across editors (n=56) were removed from the analysis, leaving a total of 212 active and 311 inactive guides. It was found that RS3 scores among inactive guides are significantly lower than those of active guides (Mann- Whitney- Wilcoxon (M.W.W) 1-sided U- test, p=0.017) (FIG. 9A). Modifying the z-score cutoff to recognize active guides to z<-3 and z<-4 (and adjusting the inactive guide accordingly) yields a more significant difference in RS3 score comparing active to inactive guides (M.W.W 1-sided U-test p=0.0030 and p=0.00026 respectively). This analysis suggests that the RS3 score identifies features of base editing guides likely to be active.

[0213] The BE-Hive editing efficiency model predicts base editing efficiency both from learned motifs in the editing window that determine activity of base editors as well as features that influence Cas9 guide efficiency in general79. The model returns a logit value, which is scaled to be centered at 0 with a standard deviation of 2, and represents the likelihood of guide activity relative to the average sgRNA-target pair from the training data. This model was applied to the data using the source Github (github.com / maxwshen / be_predict_efficiency), indicating the use of the mES cell line as recommended by the developers and the use of ABE8 and BE4Max editors in the absence ofB1195.70207WO00 71 / 102model specifications for ABE8e and TadCBEd. The same sets of active and inactive guides were employed as above to identify the performance of this model, the only differences being that active and inactive CBE and ABE guides were kept separate and were allowed to repeat in this analysis since BE-Hive reports editor-specific predicted activity scores. A significant difference was observed in logit scores of predicted activity by the BE-Hive editing efficiency model between active and inactive guides (M.W.W 1-sided U-test, p=0.016) (FIG.9B). As with RS3, this difference increases at lower z-score thresholds to define a guide as active (p=0.0076, p=0.0056).

[0214] Another approach, FORECasT-BE, uses a gradient boosted tree model that predicts base editing frequency at guide positions 3-10 by leveraging sequence features and the melting temperature of the guide to the target DNA80. The identity of the base preceding the edited one is the feature with the greatest weight; for both editors, greater activity is most likely when the editable position is preceded by a thymine. The FORECasT-BE web tool was used to obtain the predicted editing rates of each sequence in comparison to a normal distribution of editing rates. FORECasT-BE scores predict the activity for guides with the use of ABE8e, effectively separating active and inactive positive control guides (p=0.00012, 0.00039, 0.0033 at a z-score activity cutoff of -2, -3 -4 respectively, M.W.W. one-sided U test) (FIG. 9CIn contrast, there is no significant difference in FORECasT-BE scores for active and inactive guides with the use of TadCBEd (p=0.49, 0.57, 0.17) (FIG. 9D). This difference in performance is likely due to the inclusion of ABE8e but not TadCBEd in the FORECasT-BE training data.

[0215] Lastly, the impact of nucleotide identity surrounding the editable position on guide activity was directly evaluated. Previously, it has been reported that ABEs and TadCBE favor a 5’ T or C whereas a leading A discourages activity15. To assess if this pattern was present in the data, for each editor, active (z<-2) and inactive (-2<=z<=2) guides were split. Within each group, the frequency of nucleotides were identified at each distance relative to the editable nucleotide within the 4-8 window. For guides with multiple editable nucleotides in the window, each served as a separate entry in determining such frequencies. The ratio of the frequencies was then identified in active against inactive guides, such that a value greater than one indicates overrepresentation of the particular nucleotide at a given position in active guides. This analysis is presented as logo plots (FIG. 9E). As anticipated, the presence of a T or C preceding the editable nucleotide is associated with greater activity of ABE8e. In contrast, nucleotide identity preceding the edited position has negligible influence onB1195.70207WO00 72 / 102TadCBEd in this context, perhaps explaining the poor performance of models that leverage nucleotide identity surrounding the editable position to predict activity.

[0216] Overall, although some of these approaches are successful in delineating between active and inactive guides on this data, it is not recommended to filter out guides on the basis of predicted on-target activity. The use of all possible guides to introduce the edit of interest is encouraged, as these predictors are not sufficiently robust to ensure elimination of only inactive guides, and thus potential discoveries could be missed by preemptively applying such filters, especially in positive selection screens.MethodsTable of ReagentsB1195.70207WO00 73 / 102Tiling library designs

[0217] The sgRNA sequences for Cas9-NG (expanded NGN PAM) and SpG-Cas9 (NNN PAM) TP53 tiling libraries were designed using TP53 Ensembl transcript ENST00000269305.9. Using BEAGLE, all sgRNAs targeting coding sequences were included as well as all sgRNAs for which the start was up to 50 nucleotides into the intron and UTRs. NGN PAM libraries included most NGN PAMs as well as other designated active PAMs observed with Cas9-NG38. sgRNAs that created BsmBI restriction sites or sgRNAs that had more than three consecutive T’s were filtered out.Base editor tiling library annotation

[0218] All adenines or cytosines in positions 4-8 of the sgRNA (where 1 is the most PAM-distal position, and positions 21-23 are the PAM), were considered to be edited for ABE and CBE conditions, respectively. The assigned amino acid residue for an sgRNA, regardless of the predicted edit, was calculated based on the position of nucleotides 4-8 of each sgRNA within the coding sequence. Nucleotide edits were used to predict amino acid changes of each sgRNA. Mutation bins were designated by predicted amino acid mutation type. The sgRNAs containing multiple mutation types were binned as the most severe mutation type given the following order: Nonsense > Splice site > Missense > Intron > Silent > UTR. The sgRNAs with no adenines or cytosines within the editing window for ABEs or CBEs, respectively, were binned as “No edits.”B1195.70207WO00 74 / 102TP53 domain annotation

[0219] Annotations for the three primary functional domains of TP53aligned with those provided by Gould and colleagues12: the N-terminal transactivation domains (TAD) span amino acids 1-61, the central DNA-binding domain (DBD) encompasses amino acids 96-292, and the C-terminal tetramerization domain (TD), also known as the oligomerization domain, includes amino acids 324-356.Splice-targeting sgRNA library design

[0220] Splice-targeting sgRNAs that target the disrupted intron splice site were designed following the consensus sequence CAGATGKGTANNNHNHNCGG (ABE) (SEQ ID NO: 79) or CAGGCGKGTANNNHNHNCGG (CBE) (SEQ ID NO: 80), with the bold letter representing the base edit necessary to occur for spliceosome recognition. Nucleotides 1-10 are derived from the human 5’ splice site consensus sequence in order to maximize chances for splicing with the following exceptions: nucleotide 6 was assigned a G and nucleotide 7 was assigned a K such that only one nucleotide is editable within the standard base editing window54. Nucleotides 11-20 were designed to balance sequence variability with parameters defined by Rule Set 3 to maximize sgRNA efficiency55. Parameters included the following: nucleotides 14 and 16 are anything but G, nucleotides 18, 19, and 20 are assigned CGG, respectively. The sgRNAs that created BsmBI restriction sites or sgRNAs that had more than three consecutive T’s were removed. Five thousand of the remaining sgRNAs were run through CRISPick to characterize on and off-target activity predictions. The sgRNAs were filtered for an on-target score >=0.7 and an off-target rank <500. Of the remaining 857 sgRNAs, 200 ABE and 200 CBE sgRNAs were randomly selected to compose the splicetargeting sgRNA libraries.Intron design

[0221] Synthetic introns were derived from the SV40 intron and inserted after base 310 of the puromycin resistance gene. The primary intron was modified to contain a stop codon beginning at 23 nucleotides from the start of the codon. For the primary screen, intron libraries corresponding to the splice-targeting sgRNA libraries were created for both ABE and CBE conditions. A second intron was adapted from the previous intron to include additional stop codons beginning at 27 and 31 nucleotides from the start of the intron, such that there is a stop codon within every reading frame. All intron sequences corresponding to splicetargeting sgRNAs used after the primary SpG-Cas9 TP53 screen are listed in Table A. TheB1195.70207WO00 75 / 102intron was inserted in the +2 reading frame of the puromycin resistance gene, the +0 reading frame of the hygromycin resistance gene, and the +1 reading frame of GFP.Modular vector design

[0222] All pF plasmids were made by gene synthesis into the EcoRV site of pUC57-Kan (Genscript).Modular vector destination vector pre-digest

[0223] Destination vectors were pre-digested using BbsI and NEBuffer 2.0 (New England Biolabs) at 37°C for 2 h. Linearized destination vectors were gel purified using 0.7% agarose gels and extracted with the Monarch DNA Gel Extraction Kit (New England Biolabs), before further purification by isopropanol precipitation.Modular vector Golden Gate assembly

[0224] The pF vectors were diluted to 10 nM in sterile water and cloned into a pre-digested destination vector via Golden Gate cloning. Each reaction contained 3 pL BbsI (New England Biolabs), 1.25 pL T4 ligase (New England Biolabs), 3 pL of lOx T4 ligase buffer (New England Biolabs), 75 ng destination vector, and a 1:1 molar ratio of fragments :destination vector. Reactions were carried out under the following thermocycler conditions: (1) 37 °C for 5 min; (2) 16°C for 5 min; (3) go to (1), xlOO; (4) 37°C for 30 min; (5) 65°C for 20 min. The Golden Gate product was treated with Exonuclease V (New England Biolabs) at 37°C for 30 min before enzyme inactivation with the addition of EDTA to 11 mM. Per reaction, 10 pL of product was transformed into Stbl3 chemically competent E. coli (Invitrogen) via heat shock, and grown at 37 °C for 16 h on agar with 100 pg / mL carbenicillin. Colonies were picked and grown at 37 °C for 16 h in 5 mL Luria-Bertani (LB) broth with 100 pg / mL carbenicillin.Plasmid DNA (pDNA) was prepared (QIAprep Spin Miniprep Kit, Qiagen). Purified plasmids were verified by restriction enzyme digest and whole plasmid sequencing through Plasmidsaurus.Library production

[0225] Oligonucleotide pools for CP 1609 and CP1610 were synthesized by TWIST; CP 1986, CP 1987, CP 1988, and CP 1989 were synthesized by IDT; pools for CP2087, CP2088, CP2089, CP2090, and CP1845 were synthesized by Genscript. BsmBI recognition sites were appended to each sgRNA sequence along with the appropriate forward and reverse overhang sequences (bold italic) for cloning into the sgRNA expression plasmids, as well as primer sites to allow differential amplification of subsets from the same synthesis pool. For CP 1609,B1195.70207WO00 76 / 102CP1610, CP2087, CP2088, CP2089, and CP2090, the final oligonucleotide sequence structure was thus:

[0226] 5'-[Forward Primer]CGTCTCACACCG (SEQ ID NO: 81) [sgRNA, 20 nt]G7 CGAGACG (SEQ ID NO: 82) [Reverse Primer]-3'. For CP1986 and CP1987, the sequence structure was the same except the forward and reverse overhang sequences were replaced with the following: TCCC, GTTT For CP1988 and CP1989, the forward and reverse overhang sequences were as follows: CTGG, GGGT.

[0227] Primers were used to amplify individual subpools using 25 pL 2x NEBnext PCR master mix (New England Biolabs), 2 pL of oligonucleotide pool (~30-300 ng), 5 pL of primer mix at a final concentration of 0.5 pM, and water to a final volume of 50 pL. PCR cycling conditions: (1) 98°C for 1 min; (2) 98°C for 30 sec; (3) 53°C for 30 s; (3) 72°C for 30 s; (4) go to step 2, x6-14 depending on library size; (5) 72°C for 5 min.

[0228] The resulting amplicons were PCR-purified (Qiagen) and cloned into their respective library vector via Golden Gate cloning with Esp3I (Fisher Scientific) and T7 ligase (Epizyme) under the following thermocycler conditions: (1) 37°C for 5 min; (2) 20°C for 5 min; (3) go to step 1, xlOO; (4) 37°C for 30 min; (5) 65°C for 10 min. The ligated product was isopropanol precipitated and electroporated into Stbl4 electrocompetent cells (Invitrogen) and grown at 37 °C for 16 h on agar with 100 pg / mL carbenicillin. Colonies were scraped and plasmid DNA (pDNA) was prepared (HiSpeed Plasmid Maxi, Qiagen). To confirm library representation and distribution, the pDNA was sequenced by Illumina MiSeq.Lentivirus production

[0229] For small-scale virus production, the following procedure was used: 18 h before transfection, HEK293T cells were seeded in 6- well dishes at a density of 1x106cells per well in 1 mL of DMEM + 10% heat-inactivated FBS. Transfection was performed using the TransIT-LTl transfection reagent (Mirus) according to the manufacturer’s protocol. Briefly, for each construct, 324pL of Opti-MEM (Corning) and 17pL LT1 was combined with a DNA mixture of the packaging plasmid pCMV_VSVG (125 ng; Addgene 8454), psPAX2 (1250 ng; Addgene 12260), and the transfer vector (312.5 ng). The solutions were incubated at room temperature for 30 min and added dropwise to cells. Plates were then transferred to a 37°C incubator for 6-8 h, after which the media was removed and replaced with DMEM + 10% FBS media supplemented with 1% BSA. Virus was harvested and filtered 40 hours after this media change.B1195.70207WO00 77 / 102

[0230] A larger-scale procedure was used for pooled library production. 18 h before transfection, 18xl06HEK293T cells were seeded in a 175 cm2tissue culture flask and the transfection was performed the same as for small-scale production using 6 mL of Opti-MEM, 305 pL of LT1, and a DNA mixture of pCMV_VSVG (5 pg), psPAX2 (50 pg), and 10 pg of the transfer vector. Following addition of the transfection mix, flasks were transferred to a 37°C incubator for 6-8 h, then the media was aspirated and replaced with B SA- supplemented media; virus was harvested and fdtered 40 h after this media change.Determination of lentiviral titer

[0231] To determine lentiviral titer for transductions, cell lines were transduced in 12- well plates with a range of virus volumes (e.g., 0, 150, 300, 500, and 800 pL virus) with 3e6 cells per well in the presence of polybrene. For transduction, the plates were centrifuged at 821 x g for 2 h, after which 2 mL of warm media was added to reduce viral toxicity. Plates were then transferred to a 37°C incubator for 4-6 h. Each well was then trypsinized and pooled. Two days post-transduction, an equal number of cells were seeded into two wells of a 6-well plate, and puromycin was added to one well. The following passage, both wells were counted for viability. A viral dose resulting in -30% transduction efficiency, corresponding to an MOI of ~0.35, was used for subsequent library screening.Small molecule dosages

[0232] The dosages for the selection drugs puromycin, blasticidin, and hygromycin were as follows for the relevant cell lines: A549: puromycin 1.5 pg / mL, blasticidin 5 pg / mL, hygromycin 400 pg / mL; MelJuSo: puromycin 1 pg / mL, blasticidin 4 pg / mL, hygromycin 100 pg / mL.

[0233] Puromycin selection was completed over 5-7 days, while blasticidin and hygromycin selection were completed over 12-14 days. For drug screens in A549 cells, etoposide and Nutlin-3 were dosed at 5 pM and 2.5 pM respectively over the course of the experiment.Lentiviral transduction to establish stable cell lines

[0234] In order to establish stable Cas-expressing cell lines for screens, MelJuSo cells or A549 cells were transduced with pRDA_867, pRDB_092, pRDB_268, or pRDB_270 in the presence of polybrene at a dosage of 4 pg / mL for MelJuSo cells and 1 pg / mL for A549 cells. Cells were centrifuged at 821 x g for 2 h in 12-well plates, after which 2 mL of warm media was added to reduce viral toxicity. Plates were then transferred to a 37 °C incubator for 4-6 h.B1195.70207WO00 78 / 102Replicate wells were then trypsinized and pooled. Successfully infected cells were selected for with blasticidin as described above.Pooled screens

[0235] For pooled screens, cells were transduced in 2 biological replicates with a lentiviral library. Transductions were performed at a low multiplicity of infection (MOI -0.35), using enough cells to achieve a representation of at least 1,000 transduced cells per sgRNA assuming a 20% - 40% transduction efficiency. The transduction protocol was the same as listed above. Puromycin was added 2 days post-transduction for presence-based selection conditions and 4 days post-transduction for activity-based selection conditions to allow time for editing to occur. Cells were passaged on puromycin for 2-3 passages to ensure complete removal of non-transduced cells. When selection was complete, if drugs were used in the screen, cells were split to any drug arms (each at a representation of at least 1,000 cells per sgRNA) and passaged every 2-3 days for an additional 2 weeks to allow sgRNAs to enrich or deplete; cell counts were taken at each passage to monitor growth. At the conclusion of each screen, cells were pelleted by centrifugation, resuspended in PBS, and frozen promptly for genomic DNA isolation.Genomic DNA isolation, PCR, and sequencing

[0236] Genomic DNA (gDNA) was isolated using either the KingFisher Flex Purification System with the Mag-Bind Blood & Tissue DNA HDQ Kit (Omega Bio-Tek), or the Macherey Nagel NucleoSpin Blood Maxi (2e7-le8 cells), Midi (5e6-2e7 cells), or Mini (< 5e6 cells) kits, per the manufacturer’s instructions. The gDNA concentrations were measured by Qubit. For samples where gDNA was limited, gDNA was purified prior to PCR using the Zymo OneStep PCR Inhibitor Removal Kit (Zymo), per the manufacturer’s instructions.

[0237] For PCR amplification, unless otherwise noted, gDNA was divided into 100 pL reactions such that each well had at most 10 pg of gDNA. Plasmid DNA (pDNA) was also included at a maximum of 100 pg per well. Each well of a 96- well PCR plate contained 1.5 pL of Titanium Taq (Takara), 10 pL of Titanium Taq buffer, 8 pL of dNTPs, 5 pL of DMSO, 0.5 pL of P5 primer at 100 pM stock, 10 pL of P7 primer at 5 pM stock, and sterile water added to 100 pL. PCR cycling conditions were as follows: (1) 95 °C for 1 min; (2) 94 °C for 30 s, (3) 52 °C for 30 s, (4) 72 °C for 30 s, (5) go to step 1, x28; (6) 72 °C for 10 min. PCR products were purified with Agencourt AMPure XP SPRI beads according to manufacturer’s instructions (Beckman Coulter, A63880). Samples were sequenced using Illumina technology (either MiSeq, HiSeq, or NovaSeq) with a 5% spike-in of PhiX.B1195.70207WO00 79 / 102Validation experiments

[0238] For validation experiments, all sgTP53 sequences can be found in Table A. Individual sgRNA vectors were assembled and made into lentivirus as described above. At least 3xl06cells were transduced in duplicate with a virus volume to obtain -30% transduction efficiency and were selected with puromycin to remove uninfected cells; puromycin doses and selection timeline were as described above.

[0239] For Cas9-NG base editor validation experiments (related to FIGs. 1A-1B, 1E-1G, 2D), after puromycin selection was completed, cells were split between no drug and Nutlin conditions and cultured for an additional 14 days. Genomic DNA was isolated using plate lysis. Two days before lysis, cells were seeded on a 96-well plate. On the day of lysis, media was aspirated, 25 p L Lucigen QuickExtract Buffer was added to wells, and the plate was vortexed. Cells were then transferred to a 96-well PCR plate and heated at 65°C for 15 minutes. The plate was then heated at 95 °C for 5 minutes. Target sites were amplified using a 2-step PCR. In the first round of PCR, genomic DNA was amplified using custom primers designed to amplify each target site. Each well contained 50 p L of NEBNext High Fidelity 2X PCR Master Mix (New England Biolabs), 0.5 pL of each primer at 100 pM, and 49 pL of gDNA. A touchdown PCR with the following cycling conditions was used: (1) 98°C for 1 min; (2) 98°C for 30 s; (3) 68°C for 30 s (- 1°C per cycle); (4) 72°C for 1 min; (5) Go to step 2, x 15; (6) 72°C for 10 mins. The second round of PCR appended Illumina adapters and well barcodes for sequencing. Each well contained 1.5 pl of Titanium Taq (Takara), 10 pL of Titanium Taq buffer, 8 pL of dNTPs, 5 pL of DMSO, 0.5 pL of P5 primer at 100 pM, 10 pL of P7 primer at 5 pM, 55 pL of water, and 10 pL of PCR product from the first PCR. The following cycling conditions were used: (1) 95 °C for 1 min; (2) 94 °C for 30 s; (3) 52.5 °C for 30 s; (4) 72 °C for 30 s; (5) go to step 2, x 15; (6) 72 °C for 10 mins. Each well was separately purified with Agencourt AMPure XP SPRI beads according to the manufacturer’s instructions (Beckman Coulter, A63880), using a 1:1 ratio of beads to PCR product. Sample concentrations were quantified by Qubit. Samples were then sequenced using Illumina MiSeq300 and a 10% PhiX spike-in. FastQ files were then processed with CRISPResso2 to obtain allele-level abundances using a minimum allele frequency cutoff of 2%.

[0240] For SpG-Cas9 base editor splice-targeting sgRNA validation experiments, after puromycin selection was completed, MelJuSo cells were passaged with interferon-y to stimulate CD274 expression at a concentration of 0.1 pg / mL. Two days later, cells were stained with APC anti-human CD274 antibody (BioLegend 393610). Staining was performedB1195.70207WO00 80 / 102by seeding 150 pL of cells per well in a 96-well U-bottom plate, centrifuging at 1000 x g for 5 min, resuspending in 99 pLflow buffer (PBS +2% FBS, 1% EDTA) and 1 pL antibody. The plate was incubated on ice for 20-30 mins, then cells were washed 3x by centrifugation and resuspension in 200 pL of flow buffer per well. Flow cytometry was performed using a Beckman Coulter CytoFLEX, and the acquired FCS files were analyzed using FlowJo™ Software.

[0241] For SpG-Cas9 and SpRY-Cas9 editing efficiency validation experiments, the CD274 locus was amplified using two custom primers ordered from IDT with sequences as follows: 5'-TTAGAACCACCAAGTCCCAT-3' (SEQ ID NO: 83); 5'-AGAAGACTTTGCCATTGTGT-3' (SEQ ID NO: 84). For self-editing experiments, the following custom primers ordered from IDT were used to amplify the integrated sgRNA: 5'-AGCAGAGATCCAGTTTGGTT-3' (SEQ ID NO: 85);5'-TGCAATATTTGCATGTCGCT-3' (SEQ ID NO: 86). PCR amplification was performed using a one-step PCR protocol. Each well contained 50 pL of NEBNext High Fidelity 2X PCR Master Mix (New England Biolabs), 0.5 pL of each primer at 100 pM, and 49 pL of gDNA. A touchdown PCR with the following cycling conditions was used: (1) 98 °C for 1 min; (2) 98 °C for 20 s; (3) 68 °C for 30 s (- 1°C per cycle); (4) 72 °C for 1 min; (5) Go to step 2, x 15; (6) 98 °C for 20 s; (7) 53 °C for 30 s; (8) 72 °C for 1 min (9) Go to step 6, x 15; (10) 72 °C for 10 mins. Each well was separately purified with Agencourt AMPure XP SPRI beads according to the manufacturer’s instructions (Beckman Coulter, A63880), using a 1:1 ratio of beads to PCR product and sent for Sanger sequencing at Azenta. Reads were overlaid and compared with the native CD274 locus or the original sgRNA sequence using SnapGene.

[0242] For TP53 validation experiments with SpG-Cas9, the different tracrRNAs used with sgTP534 and sgTP535 are handles from70and designated as “high functional.” Sequences are below.

[0243] 5’A:ATCTGAGAGCCAAAAATGGCAAGTTCAGATAAGGCCAGACCGTTACCAGCTTAAA TAAGCGATCCTAAAGCCCCGAA (SEQ ID NO: 87)

[0244] 5’C:CTTACCGAACTAGGAATAGTAAGTGGTAAGAAGGCCTGACCGTAATAAGCCTGAA AAGGCGACCAAAAAGGGGGGATTTTATCTCCCCTTTAA (SEQ ID NO: 88)

[0245] In these TP53 validation experiments with SpG-Cas9, puromycin was added on days 4 and 7. Cells were passaged with no drug on day 9 to expand the populations. Cells were split between no drug and etoposide (5uM) conditions on day 11 and an early time pointB1195.70207WO00 81 / 102pellets were collected. Cells were passaged with their drug condition until day 26 at which point final time point pellets were collected. Cells then underwent gDNA isolation using Machery-Nagel’s NucleoSpin Blood Mini Kit per the manufacturer’s instructions. 3 PCRs were then performed using 50ng / well given limited cell numbers. The first PCR to amplify the sgRNA cassette for Sanger sequencing was performed using the following custom primers: 5’-AGCAGAGATCCAGTTTGGTT-3’ (SEQ ID NO: 85); 5’-TGCAATATTTGCATGTCGCT-3’ (SEQ ID NO: 86) and the one-step PCR protocol described above for SpG-Cas9 and SpRY-Cas9 editing efficiency. Each well was then separately purified with Agencourt AMPure XP SPRI beads according to the manufacturer’s instructions (Beckman Coulter, A63880), using a 1:1 ratio of beads to PCR product and sent for Sanger sequencing at Azenta. For the PCR amplifying sgTP534, a two-step PCR was performed as described in the Cas9-NG base editor validation methods section. The following primers were used: 5’-TTGTGGAAAGGACGAAA CACCGACTCCACACGCAAATTTCCT-3’ (SEQ ID NO: 89); 5’-TCTACTATTCTTTCCCCTGCACTGTtgtgccc tgactttcaactc-3’ (SEQ ID NO: 90).

[0246] For amplifying sgTP53 5, a two-step PCR was performed using the following primers for step 1: 5’-TTGTGGAAAGGACGAAACACCGAGGTTGGCTCTGACTGTACC-3’ (SEQ ID NO: 91) ; 5’-TCTACTATTCTTTCCCCTGCACTGTAGCAGTAAGGAGATTCCCCG-3’ (SEQ ID NO: 92). For step 1 of each PCR, each well contained 1.5 pl of Titanium Taq (Takara), 10 pL of Titanium Taq buffer, 8 pL of dNTPs, 5 pL of DMSO, 0.5 pL of each primer at lOOpM, 74.5 pL of water, and 50ng of gDNA. The following cycling conditions were used for the first PCR: (1) 95 °C for 5 min; (2) 94 °C for 30 s; (3) 53 °C for 30 s; (4) 72 °C for 20 s; (5) go to step 2, x 17; (6) 72 °C for 10 mins. The second PCR used the same conditions as described in previous second step PCRs. For both sgTP534 and sgTP535 PCRs, cells were pooled at equal ratios and purified with Agencourt AMPure XP SPRI beads according to the manufacturer’s instructions (Beckman Coulter, A63880), using a 0.7:1 ratio of beads to PCR product. The purified product for sgTP534 was sent for Illumina Sequencing using MiSeq500 and the purified product for sgTP535 was sent for MiSeq300, both with PhiX at 10%. Anon-edited vector was also sequenced in order to assess that the cells contained the expected baseline TP53 sequence. FastQ files were then processed with CRISPResso2 to obtain allele-level abundances using a minimum allele frequency cutoff of 1.5% to include the wildtype sequence.B1195.70207WO00 82 / 102Western blot

[0247] Cells were prepared for Western blotting by lifting via scraping and lysing with NP-40 lysis buffer (ThermoFisher, FNN002) supplemented with protease inhibitor (Roche, 11836170001). The extracted protein samples were quantified using the Pierce™ BCA Protein Assay kit (ThermoFisher, 23225) to determine loading volume. Samples were loaded at 15-20 pg per well onto a NuPAGE 4-12% Bis-Tris Acrylamide midi gel along with Precision Plus dual color ladder (BioRad, 1610374). The gel was run on an XCell4 Surelock Midi-Cell (ThermoFisher WR0100) with MES buffer (Life Technologies, NP0002) at 120V for 60 mins. The gel was transferred to a membrane using the iBlot2 transfer stack package (ThermoFisher, IB23001). Blocking was done by rocking for 1 h at room temperature in blocking buffer (LLCOR, 927-70050). Primary (Abeam, ab28 (anti-p53) and ab219649 (anti-Vinculin)) and secondary (LLCOR, 925-32210 and 925-68071) antibody solutions were made in blocking buffer supplemented with Tween-20 (VWR, 100216-360). The blot was incubated in the primary antibody solution while rocking at 4°C overnight, followed by 3 x 10 min washes in TBST. Secondary antibody solution was added and the blot was incubated for 1 h at room temperature then washed 3x for 10 min in TBST. The prepared blot was imaged on the LLCOR Odyssey. The resulting image was transformed into greyscale and inverted for easier band visualization.Prime editing activity-based selection

[0248] Dual-epegRNA constructs were designed with a puromycin resistance disrupted splice site-targeting epegRNA and a previously validated HEK3 epegRNA (Table C). For the puromycin resistance-targeting epegRNAs, three epegRNAs were designed using PrimeDesign to introduce a two-nucleotide substitution: AC to GT68. Both the primer binding site (PBS) and reverse transcription template (RTT) lengths were varied to optimize efficiency of the co-selecting epegRNA. The HEK3 epegRNA includes the pseudoknot evopreQl, while the co-selection epegRNA includes the pseudoknot mpknot69. To evaluate this method for prime editors, PE7 and PEmax were electroporated into MelJuSo cells and selected with hygromycin for 14 days. The three activity-based pegRNA vectors were then transduced via lentivirus as well as a control presence-based pegRNA vector, containing the HEK3 edit and an intact puromycin resistance cassette. Puromycin was selected on and the HEK3 locus was sequenced on day 14 post-transduction.B1195.70207WO00 83 / 102Electroporation of prime editors

[0249] MelJuSo cells were electroporated using the 4D-Nucleofector™ System (4D-Nucleofector™ Core and X Unit) and the Amaxa SF Cell Line 4D-Nucleofector™ X Kit (Cat. V4XC-2032). Electroporation was performed using the program CM-130. Each electroporation well contained 200,000 cells, 400 ng of either the PEmax or PE7 PiggyBac plasmid (pRDA_960 and pRDB_494, respectively), and 160 ng of the transposase-containing plasmid pRDA_179, in a total volume of 20 pl of electroporation solution from the kit. Immediately after electroporation, 100 pl of fresh, penicillin-streptomycin-free media was added to each well, and the cells were incubated at 37°C for 15 minutes. Wells electroporated with the same prime editor were then pooled and seeded in penicillin-streptomycin-free media. 24h post-electroporation, the culture media was replaced with media containing 1% penicillin-streptomycin. 72h post-electroporation, the cells were subjected to hygromycin selection (100 pg / ml) for 14 days.Quantification and statistical analysisSelf-editing analysis pipeline

[0250] For the SpG-Cas9 NNN-PAM screen, a dictionary was created with all sgRNA sequences as keys and all possible A>G or OT edits as values. The top 1 million unexpected sequences in terms of total read count frequency were then mapped to their corresponding original sgRNA sequence had any self-editing occurred. Finally, read counts from these unexpected sequences were added to their mapped sgRNA sequences and proceeded with analysis.Screen analysis

[0251] sgRNA sequences were extracted from sequencing reads by running PoolQ with the search prefix “CACCG”, except for the splice-targeting sgRNA screen, which used the prefix “TCCCG”. Reads were counted by alignment to a reference file of all possible sgRNAs present in the library. The read was then assigned to a condition (e.g., a well on the PCR plate) on the basis of the 8 nt index included in the P7 primer. After deconvolution, read counts were log-normalized by the following formula:Log Norm Read Count = log2( -■■■■ - — — * le6 + 1)Total Read Count In Condition

[0252] The LFCs were then calculated. All no drug (dropout) conditions were compared to the plasmid DNA (pDNA); drug-treated conditions were compared to the time-matched dropout sample. The correlation between LFCs of replicates was assessed. Prior to furtherB1195.70207WO00 84 / 102analysis, sgRNAs for which the log-normalized reads per million of the pDNA was > 3 standard deviations below the mean were filtered out. For all pooled screens except the splice-targeting sgRNA screen, z-scores were calculated relative to the intergenic controls. For the splice-targeting sgRNA screen, z-scores were calculated relative to the mean and standard deviation of all replicates.

[0253] For integration of base editing guides with ClinVar TP53 variant information, classifications are inclusive of those recognized as “likely” benign or pathogenic. Guides were associated with pathogenicity if they were predicted to introduce a ClinVar TP 53 pathogenic mutation, considering activity in the 4-8 editing window. Benign guides indicate those not predicted to introduce pathogenic missense or nonsense variants, considering activity to be possible in the 3-9 editing window. These classifications are determined separately for the context of ABE and CBE, and the total counts from each context are combined to represent the full set of potential edits from base editing guides.

[0254] For the analysis of external TP53 screens, the Giacomelli deep mutational scanning screen z-scores are obtained from MaveDB (mavedb.org / score-sets / urn:mavedb:00000068-a-l)79.The prime editing TP53 screen Nutlin arm data was obtained from the github (github.com / samgould2 / p53-prime-editing-sensor), modifying figure6.ipynb to retrieve log2 fold changes between D34 and D4 as calculated by MAGeCK before pegRNAs introducing silent edits are excluded from the dataset. These guides were additionally filtered to those with at least 10% or 25% editing efficiency to evaluate how restricting the analysis to effective guides improves screening performance. It is noted that filtering by 10% and 25% editing efficiency eliminates 85.4% and 93.8% of guides from analysis.Statistics

[0255] For all experiments, replicates were distinct and not re-sampled. Statistical tests, central tendencies, p-values, and sample sizes are described in figures or brief descriptions. For figures featuring boxplots, boxes depict 25th (QI) and 75th (Q3) percentiles as minima and maxima and the center represents the median. Whiskers identify outlier points by depicting QI - 1.5*IQR and Q3 + 1.5IQR, where IQR represents the range between QI and Q3, and outliers are shown as diamonds.Data visualization

[0256] Figures were created with Python and FlowJo™ Software. Schematics were created with BioRender.com.B1195.70207WO00 85 / 102References for Example 21. Araya, C. L. & Fowler, D. M. Deep mutational scanning: assessing protein function on a massive scale. Trends BiotechnoL 29, 435-442 (2011).2. Fowler, D. M. & Fields, S. Deep mutational scanning: a new style of protein science.Nat. Methods 11, 801-807 (2014).3. Wei, H. & Li, X. Deep mutational scanning: A versatile tool in systematically mapping genotypes to phenotypes. Front. Genet. 14, 1087267 (2023).4. Patwardhan, R. P. et al. High-resolution analysis of DNA regulatory elements by synthetic saturation mutagenesis. Nat. BiotechnoL 27, 1173-1175 (2009).5. Patwardhan, R. P. et al. Massively parallel functional dissection of mammalian enhancers in vivo. Nat. BiotechnoL 30, 265-270 (2012).6. Melnikov, A. et aL Systematic dissection and optimization of inducible enhancers in human cells using a massively parallel reporter assay. Nat. BiotechnoL 30, 271-277 (2012).7. Findlay, G. M., Boyle, E. A., Hause, R. J., Klein, J. C. & Shendure, J. Saturation editing of genomic regions by multiplex homology-directed repair. Nature 513, 120-123 (2014).8. Findlay, G. M. et aL Accurate classification of BRCA1 variants with saturation genome editing. Nature 562, 217-222 (2018).9. Anzalone, A. V. et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149-157 (2019).10. Yan, J. et al. Improving prime editing with an endogenous small RNA-binding protein.Nature 628, 639-647 (2024).11. Erwood, S. et al. Saturation variant interpretation using CRISPR prime editing. Nat.BiotechnoL 40, 885-895 (2022).12. Gould, S. I. et al. High-throughput evaluation of genetic variants with prime editing sensor libraries. Nat. BiotechnoL 1-15 (2024).13. Komor, A. C., Kim, Y. B., Packer, M. S., Zuris, J. A. & Liu, D. R. Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature 533, 420-424 (2016).14. Gaudelli, N. M. et aL Programmable base editing of A»T to G*C in genomic DNA without DNA cleavage. Nature 551, 464-471 (2017).15. Neugebauer, M. E. et al. Evolution of an adenine base editor into a small, efficient cytosine base editor with low off-target activity. Nat. BiotechnoL 41, 673-685 (2023).B1195.70207WO00 86 / 10216. Hanna, R. E. et al. Massively parallel assessment of human variants with base editor screens. Cell 184, 1064-1080.e20 (2021).17. Sanchez-Rivera, F. J. et al. Base editing sensor libraries for high-throughput engineering and functional analysis of cancer-associated single nucleotide variants. Nat. Biotechnol. (2022) doi:10.1038 / s41587-021-01172-3.18. Pemer, F. et al. MEN1 mutations mediate clinical resistance to menin inhibition. Nature 615, 913-919 (2023).19. Martin-Rufino, J. D. et al. Massively parallel base editing to map variant effects in human hematopoiesis. Cell 186, 2456-2474.e24 (2023).20. Eue, N. Z. et al. Base editor scanning charts the DNMT3A activity landscape. Nat.Chem. Biol. 19, 176-186 (2023).21. Kennedy, P. H. et al. Post- translational modification-centric base editor screens to assess phosphorylation site functionality in high throughput. Nat. Methods 21, 1033-1043 (2024).22. Ei, H. et al. Assigning functionality to cysteines by base editing of cancer dependency genes. Nat. Chem. Biol. 19, 1320-1330 (2023).23. Serrao, E. & Engelman, A. N. Sites of retroviral DNA integration: From basic research to clinical applications. Crit. Rev. Biochem. Mol. Biol. 51, 26-42 (2016).24. Shao, L. et al. Genome- wide profiling of retroviral DNA integration and its effect on clinical pre-infusion CAR T-cell products. J. Transl. Med. 20, 514 (2022).25. Cabrera, A. et al. The sound of silence: Transgene silencing in mammalian cell engineering. Cell Syst. 13, 950-973 (2022).26. Chew, W. L. et al. A multifunctional AAV-CRISPR-Cas9 and its host response. Nat.Methods 13, 868-874 (2016).27. Baugh, E. H., Ke, H., Levine, A. J., Bonneau, R. A. & Chan, C. S. Why are there hotspot mutations in the TP53 gene in human cancers? Cell Death Differ. 25, 154-160 (2018).28. Olivier, M., Hollstein, M. & Hainaut, P. TP53 mutations in human cancers: origins, consequences, and clinical use. Cold Spring Harb. Perspect. Biol. 2, a001008 (2010). 29. The TP53 database, https: / / tp53.cancer.gov / .30. Petitjean, A., Achatz, M. I. W., Borresen-Dale, A. L., Hainaut, P. & Olivier, M. TP53 mutations in human cancers: functional selection and impact on cancer prognosis and outcomes. Oncogene 26, 2157-2165 (2007).31. Donehower, L. A. et al. Integrated analysis of TP53 gene and pathway alterations in the cancer genome atlas. Cell Rep. 28, 3010 (2019).B1195.70207WO00 87 / 10232. Wang, H., Guo, M., Wei, H. & Chen, Y. Targeting p53 pathways: mechanisms, structures and advances in therapy. Signal Transduct. Target. Then 8, 1-35 (2023).33. Arya, A. K. et al. Nutlin-3, the small-molecule inhibitor of MDM2, promotes senescence and radiosensitises laryngeal carcinoma cells harbouring wild-type p53. Br. J. Cancer 103, 186-195 (2010).34. Giacomelli, A. O. et al. Mutational processes shape the landscape of TP53 mutations in human cancer. Nat. Genet. 50, 1381-1387 (2018).35. Kucab, J. E., Hollstein, M., Arlt, V. M. & Phillips, D. H. Nutlin-3a selects for cells harbouring TP53 mutations. Int. J. Cancer 140, 877-887 (2017).36. Montecucco, A., Zanetta, F. & Biamonti, G. Molecular mechanisms of etoposide. EXCL1 J. 14, 95-108 (2015).37. Menendez, D. et al. Etoposide-induced DNA damage is increased in p53 mutants:identification of ATR and other genes that influence effects of p53 mutations on Top2- induced cytotoxicity. Oncotarget 13, 332-346 (2022).38. Sangree, A. K. et al. Benchmarking of SpCas9 variants enables deeper base editor screens of BRCA1 and BCL2. Nat. Commun. 13, 1-17 (2022).39. Richter, M. F. et al. Phage-assisted evolution of an adenine base editor with improved Cas domain compatibility and activity. Nat. Biotechnol. 38, 883-891 (2020).40. Nishimasu, H. et al. Engineered CRISPR-Cas9 nuclease with expanded targeting space.Science 361, 1259-1262 (2018).41. Clement, K. et al. CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat. Biotechnol. 37, 224-226 (2019).42. Hafsi, H., Santos-Silva, D., Courtois-Cox, S. & Hainaut, P. Effects of A40p53, an isoform of p53 lacking the N-terminus, on transactivation capacity of the tumor suppressor protein p53. BMC Cancer 13, 134 (2013).43. Joruiz, S. M. & Bourdon, J.-C. P53 isoforms: Key regulators of the cell fate decision.Cold Spring Harb. Perspect. Med. 6, (2016).44. McGee, A. V. et al. Modular vector assembly enables rapid assessment of emerging CRISPR technologies. bioRxiv (2023) doi:10.1101 / 2023.10.25.564061.45. Walton, R. T., Christie, K. A., Whittaker, M. N. & Kleinstiver, B. P. Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9 variants. Science 368, 290-296 (2020).46. Kurt, I. C. et al. CRISPR C-to-G base editors for inducing targeted DNA transversions in human cells. Nat. Biotechnol. 39, 41-46 (2021).B1195.70207WO00 88 / 10247. Lam, D. K. et al. Improved cytosine base editors generated from TadA variants. Nat. Biotechnol. 41, 686-697 (2023).48. Huang, T. P. et al. Circularly permuted and PAM-modified Cas9 variants broaden the targeting scope of base editors. Nat. Biotechnol. 37, 626-631 (2019).49. Kluesner, M. G. et al. EditR: A method to quantify base editing from Sanger sequencing.CRISPR J. 1, 239-250 (2018).50. Shy, B. R., MacDougall, M. S., Clarke, R. & Merrill, B. J. Co-incident insertion enables high efficiency genome engineering in mouse embryonic stem cells. Nucleic Acids Res.44, 7997-8010 (2016).51. Agudelo, D. et al. Marker- free coselection for CRISPR-driven genome editing in human cells. Nat. Methods 14, 615-620 (2017).52. Xu, D.-H. et al. SV40 intron, a potent strong intron element that effectively increases transgene expression in transfected Chinese hamster ovary cells. J. Cell. Mol. Med. 22, 2231-2239 (2018).53. Roca, X. et al. Widespread recognition of 5’ splice sites by noncanonical base-pairing to U1 snRN A involving bulged nucleotides. Genes Dev. 26, 1098-1109 (2012).54. Roca, X., Krainer, A. R. & Eperon, I. C. Pick one, but be quick: 5’ splice sites and the problems of too many choices. Genes Dev. 27, 129-144 (2013).55. DeWeirdt, P. C. et al. Accounting for small variations in the tracrRNA sequence improves sgRNA activity predictions for CRISPR screening. Nat. Commun. 13, 5255 (2022).56. Bae, S., Park, J. & Kim, J.-S. Cas-OFFinder: a fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases. Bioinformatics 30, 1473-1475 (2014).57. Hibshman, G. N. et al. Unraveling the mechanisms of PAMless DNA interrogation by SpRY-Cas9. Nat. Commun. 15, 3663 (2024).58. Shi, H. et al. Rapid two-step target capture ensures efficient CRISPR-Cas9-guided genome editing. Mol. Cell 85, 1730-1742. e9 (2025).59. Ren, Q. et al. PAM-less plant genome editing using a CRISPR-SpRY toolbox. Nat.Plants 7, 25-33 (2021).60. Shi, J. et al. Discovery of cancer drug targets by CRISPR-Cas9 screening of protein domains. Nat. Biotechnol. 33, 661-667 (2015).61. He, W. et al. De novo identification of essential protein domains from CRISPR-Cas9 tiling-sgRNA knockout screens. Nat. Commun. 10, 4541 (2019).B1195.70207WO00 89 / 10262. Munoz, D. M. et al. CRISPR screens provide a comprehensive assessment of cancer vulnerabilities but generate false-positive hits for highly amplified genomic regions. Cancer Discov. 6, 900-913 (2016).63. Schoenenberg, V. A. C. et al. CRISPRO: identification of functional protein coding sequences based on genome editing dense mutagenesis. Genome Biol. 19, 169 (2018).64. Herman, J. A. et al. Functional dissection of human mitotic genes using CRISPR-Cas9 tiling screens. Genes Dev. 36, 495-510 (2022).65. Levesque, S. et al. Marker-free co-selection for successive rounds of prime editing in human cells. Nat. Commun. 13, 5909 (2022).66. Herger, M. et al. High-throughput screening of human genetic variants by pooled prime editing. Cell Genom. 5, 100814 (2025).67. Chen, P. J. et al. Enhanced prime editing systems by manipulating cellular determinants of editing outcomes. Cell 184, 5635-5652.e29 (2021).68. Hsu, J. Y. et al. PrimeDesign software for rapid and simplified design of prime editing guide RNAs. Nat. Commun. 12, 1034 (2021).69. Nelson, J. W. et al. Engineered pegRNAs improve prime editing efficiency. Nat.Biotechnol. 40, 402-410 (2022).70. Reis, A. C. et al. Simultaneous repression of multiple bacterial genes using nonrepetitive extra-long sgRNA arrays. Nat. Biotechnol. 37, 1294-1301 (2019).71. Katti, A. et al. GO: a functional reporter system to identify and enrich base editing activity. Nucleic Acids Res. 48, 2841-2852 (2020).72. Coelho, M. A. et al. BE-FLARE: a fluorescent reporter of base editing activity reveals editing characteristics of APOBEC3A and APOBEC3B. BMC Biol. 16, 150 (2018). 73. Li, S. et al. Universal toxin-based selection for precise genome engineering in human cells. Nat. Commun. 12, 497 (2021).74. Schmidt, R. et al. Base-editing mutagenesis maps alleles to tune human T cell functions.Nature 625, 805-812 (2024).75. Kim, Y, Oh, H.-C., Lee, S. & Kim, H. H. Saturation profiling of drug-resistant genetic variants using prime editing. Nat. Biotechnol. (2024) doi:10.1038 / s41587-024-02465-z.76. Coelho, M. A. et al. Base editing screens map mutations affecting interferon-y signaling in cancer. Cancer Cell 41, 288-303.e6 (2023).77. Pablo, J. L. B. et al. Scanning mutagenesis of the voltage-gated sodium channel NaV1.2 using base editing. Cell Rep. 42, 112563 (2023).78. Coelho, M. A. et al. Base editing screens define the genetic landscape of cancer drugB1195.70207WO00 90 / 102resistance mechanisms. Nat. Genet. (2024) doi:10.1038 / s41588-024-01948-8.79. Rubin, A. F. et al. MaveDB 2024: a curated community database with over seven million variant effects from multiplexed functional assays. Genome Biol. 26, 13 (2025).INCORPORATED BY REFERENCE

[0257] All of the U.S. patents, U.S. patent application publications, U.S. patent applications, foreign patents, foreign patent applications, and non-patent publications referred to in this specification are incorporated herein by reference in their entireties.

[0258] Although the foregoing methods have been described in some detail to facilitate understanding, it will be apparent that certain changes and modifications may be practiced within the scope of the appended claims. Accordingly, the described embodiments are to be considered as illustrative and not restrictive, and the claimed invention is not to be limited to the details given herein, but may be modified within the scope and equivalents of the appended claims.EQUIVALENTS AND SCOPE

[0259] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. The scope of the present invention is not intended to be limited to the above description, but rather is as set forth in the appended claims.

[0260] In the claims articles such as “a,” “an,” and “the” may mean one or more than one unless indicated to the contrary or otherwise evident from the context. Claims or descriptions that include “or” between one or more members of a group are considered satisfied if one, more than one, or all of the group members are present in, employed in, or otherwise relevant to a given product or process unless indicated to the contrary or otherwise evident from the context. The invention includes embodiments in which exactly one member of the group is present in, employed in, or otherwise relevant to a given product or process. The invention also includes embodiments in which more than one, or all of the group members are present in, employed in, or otherwise relevant to a given product or process.

[0261] Furthermore, it is to be understood that the invention encompasses all variations, combinations, and permutations in which one or more limitations, elements, clauses, descriptive terms, etc., from one or more of the claims or from relevant portions of the description is introduced into another claim. For example, any claim that is dependent on another claim can be modified to include one or more limitations found in any other claimB1195.70207WO00 91 / 102that is dependent on the same base claim. Furthermore, where the claims recite a composition, it is to be understood that methods of using the composition for any of the purposes disclosed herein are included, and methods of making the composition according to any of the methods of making disclosed herein or other methods known in the art are included, unless otherwise indicated or unless it would be evident to one of ordinary skill in the art that a contradiction or inconsistency would arise.

[0262] Where elements are presented as lists, e.g., in Markush group format, it is to be understood that each subgroup of the elements is also disclosed, and any element(s) can be removed from the group. It is also noted that the term “comprising” is intended to be open and permits the inclusion of additional elements or steps. It should be understood that, in general, where the invention, or aspects of the invention, is / are referred to as comprising particular elements, features, steps, etc., certain embodiments of the invention or aspects of the invention consist, or consist essentially of, such elements, features, steps, etc. For purposes of simplicity those embodiments have not been specifically set forth in haec verba herein. Thus, for each embodiment of the invention that comprises one or more elements, features, steps, etc., the invention also provides embodiments that consist or consist essentially of those elements, features, steps, etc.

[0263] Where ranges are given, endpoints are included. Furthermore, it is to be understood that unless otherwise indicated or otherwise evident from the context and / or the understanding of one of ordinary skill in the art, values that are expressed as ranges can assume any specific value within the stated ranges in different embodiments of the invention, to the tenth of the unit of the lower limit of the range, unless the context clearly dictates otherwise. It is also to be understood that unless otherwise indicated or otherwise evident from the context and / or the understanding of one of ordinary skill in the art, values expressed as ranges can assume any subrange within the given range, wherein the endpoints of the subrange are expressed to the same degree of accuracy as the tenth of the unit of the lower limit of the range.

[0264] In addition, it is to be understood that any particular embodiment of the present invention may be explicitly excluded from any one or more of the claims. Where ranges are given, any value within the range may explicitly be excluded from any one or more of the claims. Any embodiment, element, feature, application, or aspect of the compositions, complexes, systems, methods, and / or uses of the invention can be excluded from any one or more claims. For purposes of brevity, all of the embodiments in which one or more elements, features, purposes, or aspects is excluded are not set forth explicitly herein.B1195.70207WO00 92 / 102

Claims

CLAIMSWhat is claimed is:

1. A polynucleotide comprising a nucleotide sequence encoding a synthetic intron having at least 80% sequence identity to a nucleic acid sequence set forth in SEQ ID NO: 31, wherein the synthetic intron comprises a defective 5' donor site which is correctable by a gene editor.

2. The polynucleotide of claim 1, wherein the synthetic intron further comprises a protospacer adjacent motif (PAM) sequence.

3. The polynucleotide of claim 2, wherein the PAM sequence is NGN, and N is A, G, T, or C.

4. The polynucleotide of any one of claims 1-3, wherein the synthetic intron further comprises at least one, at least two, or at least three stop codons.

5. The polynucleotide of any one of claims 1-4, wherein the synthetic intron is at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% identical to the nucleic acid sequence of SEQ ID NO: 31.

6. The polynucleotide of any one of claims 1-4, wherein the synthetic intron comprises at least 1, at least two, at least three, at least 4, at least 5, at least 6, at least 7 nucleic acid substitutions relative to the nucleic acid sequence of SEQ ID NO: 31.

7. The polynucleotide of any one of claims 1-6, wherein the defective 5' donor splice site comprises an AT pair in place of the canonical GT pair of SEQ ID NO: 31.

8. The polynucleotide of claim 7, wherein the AT pair can be converted to the canonical GT pair with an adenine base editor (ABE) edit.

9. The polynucleotide of any one of claims 1-6, wherein the defective 5' donor splice site comprises an GC pair in place of the canonical GT pair of SEQ ID NO: 31.B1195.70207WO00 93 / 10210. The polynucleotide of claim 9, wherein the GC pair can be converted to the canonical GT pair with a cytidine base editor (CBE) edit.

11. The polynucleotide of any one of claims 1-10, wherein the synthetic intron comprises the nucleic acid sequence of any one of SEQ ID NOs: 12-22.

12. A splice-intron guide RNA (sigRNA) comprising a first portion comprising a region of complementarity to a polynucleotide of any one of claims 1-11, and a second portion comprising a trans-activating CRISPR RNA(tracrRNA).

13. The sigRNA of claim 12, wherein the region of complementarity comprises between 15 and 25 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 12-22.

14. The sigRNA of claim 12 or 13, wherein the region of complementarity comprises 17 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 12-22.

15. The sigRNA of any one of claims 12-14, wherein the sigRNA comprises a sequence set forth in any one of SEQ ID NOs: 1-11.

16. The sigRNA of any one of claims 12-15, wherein the sigRNA further comprises at least 1, at least 2, at least 3 nucleic acid substitutions relative to the nucleic acid sequence of any one of SEQ ID NOs: 1-11.

17. A complex comprising: (a) a gene editor comprising a DNA binding protein, and (b) the sigRNA of any one of claims 13-16.

18. The complex of claim 17, wherein the DNA binding protein is a nucleic acid programmable DNA binding protein (napDNAbp).

19. The complex of claim 17 or 18, wherein the napDNAbp is selected from the group consisting of Cas9, CasX, CasY, Cpfl, C2cl, C2c2, C2C3, Sp-Cas9, SpRY, SpG-Cas9, NG-Cas9, NRRH-Cas9, spCas9, geoCas9, saCas9, Nme2Cas9, Casl2(a-i), Casl4, Argonaute, and variants thereof.B1195.70207WO00 94 / 10220. The complex of any one of claims 17-19, wherein the napDNAbp is a Cas9 protein, a dCas9 protein, or a nCas9 protein.

21. The complex of any one of claims 17-20, wherein the gene editor further comprises a deaminase.

22. The complex of claim 21, wherein the deaminase is a cytidine deaminase or an adenosine deaminase.

23. The complex of claim 21 or 22, wherein the gene editor is a cytidine base editor selected from the group consisting of evoCDAmax, CBE6, CGBE, BE4max, TadCBEd, and variants thereof.

24. The complex of claim 22 or 23, wherein the cytidine base editor is TadCBEd.

25. The complex of claim 21 or 22, wherein the gene editor is an adenosine base editor selected from the group consisting of TadA-8e, ABE8.0, ABE8e, AYBE, ABE9, and variants thereof.

26. The complex of claim 22 or 25, wherein the adenosine base editor is ABE8e.

27. A polynucleotide cassette comprising a nucleotide sequence encoding a selection marker disrupted by a synthetic intron having at least 80% sequence identity to a nucleic acid sequence set forth in SEQ ID NO: 31, wherein the synthetic intron comprises a defective 5' donor splice site that is correctable by a gene editor.

28. The polynucleotide cassette of claim 27, wherein the selection marker is an antibiotic resistance gene.

29. The polynucleotide cassette of claim 28, wherein the antibiotic resistance gene is selected from the group consisting of a puromycin resistance gene, a blasticidin resistance gene, an ampicillin resistance gene, a chloramphenicol resistance gene, a streptomycin resistance gene, or a kanamycin resistance gene.B1195.70207WO00 95 / 10230. A nucleic acid encoding the polynucleotide of any one of claims 1-11, the sigRNA of any one of claims 12-16, the complex of any one of claims 17-26, and / or polynucleotide cassette of any one of claims 27-29.

31. An expression vector comprising a nucleic acid encoding the polynucleotide of any one of claims 1-11, the sigRNA of any one of claims 12-16, the complex of any one of claims 17-26, the polynucleotide cassette of any one of claims 27-29, and / or the nucleic acid of claim 30.

32. A cell comprising the polynucleotide of any one of claims 1-11, the sigRNA of any one of claims 12-16, the complex of any one of claims 17-26, the polynucleotide cassette of any one of claims 27-29, the nucleic acid of claim 30, and / or the expression vector of claim 31.

33. The cell of claim 32, wherein the cell is a bacterial cell.

34. The cell of claim 32, wherein the cell is a eukaryotic cell.

35. A composition of nucleic acids, the composition comprising:(a) a first nucleic acid encoding a selection marker disrupted by a synthetic intron having at least 80% sequence identity to a nucleic acid sequence set forth in SEQ ID NO: 31, wherein the synthetic intron comprises a defective 5' donor splice site that is correctable by a nucleic acid programmable gene editor;(b) a second nucleic acid encoding a splice-intron guide RNA (sigRNA) comprising a spacer sequence that has complementarity to the defective 5' donor splice site of the synthetic intron of (a), and a trans-activating CRISPR RNA(tracrRNA); and / or(c) a third nucleic acid encoding a guide RNA (gRNA) comprising a spacer sequence comprising a region of complementarity to a target gene of interest; and / or(d) a fourth nucleic acid encoding a gene expression marker; and / or(e) a fifth nucleic acid encoding a gene editor.

36. The composition of claim 35, wherein the synthetic intron further comprises a protospacer adjacent motif (PAM) sequence.B1195.70207WO00 96 / 10237. The composition of claim 36, wherein the PAM sequence is NGN, and N is A, G, T, or C.

38. The composition of any one of claims 35-37, wherein the synthetic intron further comprises at least one, at least two, or at least three stop codons.

39. The composition of any one of claims 35-38, wherein the synthetic intron is at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% identical to the nucleic acid sequence of SEQ ID NO: 31.

40. The composition of any one of claims 35-39, wherein the synthetic intron comprises at least 1, at least two, at least three, at least 4, at least 5, at least 6, or at least 7 nucleic acid substitutions relative to the nucleic acid sequence of SEQ ID NO: 31.

41. The composition of any one of claims 35-40, wherein the defective 5' donor splice site comprises an AT pair in place of the canonical GT pair of SEQ ID NO: 31.

42. The composition of claim 41, wherein the AT pair can be converted to the canonical GT pair with an adenine base editor (ABE) edit.

43. The composition of any one of claims 35-40, wherein the defective 5' donor splice site comprises an GC pair in place of the canonical GT pair of SEQ ID NO: 31.

44. The composition of claim 43, wherein the GC pair can be converted to the canonical GT pair with a cytidine base editor (CBE) edit.

45. The composition of any one of claims 35-44, wherein the selection marker is an antibiotic resistance gene.

46. The composition of claim 45, wherein the antibiotic resistance gene is selected from the group consisting of a puromycin resistance gene, a blasticidin resistance gene, an ampicillin resistance gene, a chloramphenicol resistance gene, a streptomycin resistance gene, or a kanamycin resistance gene.B1195.70207WO00 97 / 10247. The composition of claim 35, wherein the spacer sequence of the sigRNA comprises between 15 and 25 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 12-22.

48. The composition of claim 35 or 47, wherein the spacer sequence of the sigRNA comprises 17 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 12-22.

49. The composition of any one of claims 35, 47, or 48, wherein the sigRNA comprises a sequence set forth in any one of SEQ ID NOs: 1-11.

50. The composition of any one of claims 35 or 47-49, wherein the sigRNA further comprises at least 1, at least 2, or at least 3 nucleic acid substitutions relative to the nucleic acid sequence of any one of SEQ ID NOs: 1-11.

51. The composition of claim 35, wherein the target gene of interest is in the genome of the cell.

52. The composition of claim 35, wherein the gene expression marker encodes for a fluorescent protein (FP).

53. The composition of claim 35 or 52, wherein the gene expression marker encodes for Green Fluorescent Protein (GFP).

54. The composition of claim 35, wherein the gene editor is a base editor.

55. The composition of claim 54, wherein the base editor comprises a DNA binding protein is a nucleic acid programmable DNA binding protein (napDNAbp).

56. The composition of claim 55, wherein the napDNAbp is selected from the group consisting of Cas9, CasX, CasY, Cpfl, C2cl, C2c2, C2C3, Sp-Cas9, SpRY, SpG-Cas9, NG-Cas9, NRRH-Cas9, spCas9, geoCas9, saCas9, Nme2Cas9, Casl2(a-i), Casl4, Argonaute, and variants thereof.B1195.70207WO00 98 / 10257. The composition of claim 55 or 56, wherein the napDNAbp is a Cas9 protein, a dCas9 protein, or a nCas9 protein.

58. The composition of any one of claims 35 or 54-57, wherein the gene editor further comprises a deaminase.

59. The composition of claim 58, wherein the deaminase is a cytidine deaminase or an adenosine deaminase.

60. The composition of claim 58 or 59, wherein the gene editor is a cytidine base editor selected from the group consisting of evoCDAmax, CBE6, CGBE, BE4max, TadCBEd, and variants thereof.

61. The composition of claim 60, wherein the cytidine base editor is TadCBEd.

62. The composition of claim 58 or 59, wherein the gene editor is an adenosine base editor selected from the group consisting of TadA-8e, ABE8.0, ABE8e, AYBE, ABE9, and variants thereof.

63. The composition of claim 62, wherein the adenosine base editor is ABE8e.

64. A system for enriching for gene editing activity in a cell comprising:(a) a nucleic acid encoding a selection marker disrupted by a synthetic intron having at least 80% sequence identity to a nucleic acid sequence set forth in SEQ ID NO: 31, wherein the synthetic intron comprises a defective 5' donor splice site that is correctable by a nucleic acid programmable gene editor;(b) a splice-intron guide RNA (sigRNA) or a nucleic acid encoding a splice-intron guide RNA comprising a spacer sequence that has complementarity to the defective 5' donor splice site of the synthetic intron of (a), and a trans-activating CRISPR RNA(tracrRNA);(c) a gene editor or a nucleic acid encoding a gene editor; and / or(d) a guide RNA or a nucleic acid encoding a guide RNA comprising a spacer sequence comprising a region of complementarity to a target gene of interest;B1195.70207WO00 99 / 102wherein gene editing-dependent correction of the defective 5' donor splice site results in the enrichment of editing of the target gene of interest.

65. The system of claim 64, wherein the nucleic acid of (a), the nucleic acid of (b), the nucleic acid of (c), and / or the nucleic acid of (d) are located on the same nucleic acid construct or different nucleic acid constructs.

66. The system of claim 64 or 65, wherein the gene editor is a base editor or prime editor.

67. The system of any one of claims 64- 66, wherein the target gene of interest is in the genome of the cell.

68. A kit comprising (i) the polynucleotide of any one of claims 1-11, the sigRNA of any one of claims 12-16, the complex of any one of claims 17-26, the polynucleotide cassette of any one of claims 27-29, the nucleic acid of claim 30, the expression vector of claim 31, the cells of any one of claims 32-34, the composition of any one of claims 35-63, and / or the system of any one of claims 64-66, and (ii) a set of instructions for conducting gene editing.

69. A method for enriching for gene editing activity in a cell, the method comprising:(a) introducing into a cell:(i) a nucleic acid encoding a selection marker disrupted by a synthetic intron having at least 80% sequence identity to a nucleic acid sequence set forth SEQ ID NO: 31, wherein the synthetic intron comprises a defective 5' donor splice site that is correctable by a nucleic acid programmable gene editor;(ii) a splice-intron guide RNA (sigRNA) or a nucleic acid encoding a spliceintron guide RNA comprising a spacer sequence that has complementarity to the defective 5' donor splice site of the synthetic intron of (i), and a trans-activating CRISPR RNA(tracrRNA);(iii) a gene editor or a nucleic acid encoding a gene editor; and / or (iv) a guide RNA or a nucleic acid encoding a guide RNA comprising a spacer sequence comprising a region of complementarity to a target gene of interest;(b) subjecting the cell to the selection pressure of the selection marker of (i), wherein cell survival indicates the gene editor made the desired edit at the splice donor site; and / orB1195.70207WO00 100 / 102(c) sequencing the target DNA sequence to confirm presence of edit.

70. The method of claim 69, wherein the gene editor of (ii) can form a complex with the guide RNA and the splice-intron guide RNA.

71. The method of claim 69 or 70, wherein the method comprises an initial step of (a) introducing to the cell a nucleic acid encoding the gene editor of (ii), the guide RNA of (iv), and a second selection marker, and (b) subjecting the cell to the selection pressure of the second selection marker, wherein cell survival indicates the presence of the nucleic acid encoding the gene editor is in the cell.

72. Use of the polynucleotide of any one of claims 1-11, the sigRNA of any one of claims 12-16, the complex of any one of claims 17-26, the polynucleotide cassette of any one of claims 27-29, the nucleic acid of claim 30, the expression vector of claim 31, the cells of any one of claims 32-34, the composition of any one of claims 35-63, the system of any one of claims 64-66, the kit of claim 67, and / or the method of any one of claims 68-71 for ex vivo gene editing.B1195.70207WO00 101 / 102