High-throughput identification of nuclease cut sites using randomized target sequences
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-02-03
- Publication Date
- 2026-08-13
Smart Images

Figure US2026013774_13082026_PF_FP_ABST
Abstract
Description
[0001] HIGH-THROUGHPUT IDENTIFICATION OF NUCLEASE CUT SITES USING RANDOMIZED TARGET SEQUENCES
[0002] RELATED APPLICATIONS
[0003] This application claims the benefit under 35 U.S.C. § 119(e) to U.S. Provisional Application No. 63 / 753,902, filed February 4, 2025, entitled “HIGH-THROUGHPUT IDENTIFICATION OF NUCLEASE CUT SITES USING RANDOMIZED TARGET SEQUENCES”, the entire contents of which are incorporated herein by reference.
[0004] REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0005] The contents of the electronic sequence listing (S232070001WO00-SEQ-KTW.xml; Size: 46,361 bytes; Date of Creation: February 3, 2026) are herein incorporated by reference.
[0006] FEDERALLY SPONSORED RESEARCH
[0007] This invention was made with government support under grant number Al 176471 awarded by the National Institutes of Health. The government has certain rights in the invention.
[0008] BACKGROUND
[0009] Genome editing technology has tremendous potential to become a transformative therapy for a wide range of genetic diseases. The recent regulatory approval of exagamglogene autotemcel (exa-cel) CASGEVY® — a CRISPR / Cas9 autologous genome edited hematopoietic stem cell therapy used to induce fetal hemoglobin production to treat sickle cell disease — in both the US and Europe marks a major milestone in this exciting field. Numerous other genome editing therapies are currently undergoing clinical trials and pre-clinical development.
[0010] However, genome editors can also introduce unintended modifications at off-target sites, potentially confounding biological research and posing key safety risks in therapeutic applications. Large-scale chromosomal translocations are associated with off-target DSBs. Of primary concern for clinical gene editing is that off-target activity may result in cells with a proliferative advantage, increasing the potential for malignant transformation — a risk observed in early and recent gene therapy studies with viral vectors.
[0011] SUMMARY
[0012] Genome editing enzymes can introduce targeted changes to the DNA in living cells, transforming biological research and enabling the first approved gene editing therapy for sickle cell disease. However, their genome- wide activity and / or specificity can be altered by geneticvariation at on- or off-target sites, potentially impacting both their precision and therapeutic safety. Due to a lack of scalable methods to measure genome-wide editing activity in cells from large populations and diverse target libraries, the frequency and extent of these variant effects on editing remains unknown. To understand common features of high-impact variants, the present disclosure describes a new massively parallel biochemical assay which may be referred to herein as “CHANCE-seq” and was developed to measure genome editor activity across millions of mismatched target sites. These methods provide a way to account for genetic variation when designing genome editing strategies for research and therapeutics.
[0013] In some embodiments, this disclosure provides a method for determining an activity and / or specificity of a genome editor for a target polynucleotide, comprising: (a) obtaining a library comprising circular DNAs, each circular DNA comprising a randomized target polynucleotide; (b) cleaving randomized target polynucleotides of circular DNAs of the library using the genome editor, thereby producing linearized DNAs; and (c) determining the activity and / or specificity of the genome editor for the target polynucleotide using the cleaved randomized target polynucleotides of the linearized DNAs.
[0014] In some embodiments, the genome editor comprises a base editor, an RNA-guided nuclease (RGN) complex, a zinc-finger nuclease (ZNF), a transcription activator-like effector nuclease (TALEN), a CRISPR-associated transpose, a prime editor, or a meganuclease.
[0015] In some embodiments, the RGN complex comprises a CRISPR nuclease and a guide RNA. In some embodiments, the CRISPR nuclease is selected from the group consisting of Cas3, Cas8a, Cas5, Cas8b, Cas8c, CaslOd, Csel, Cse2, Csyl, Csy2, Csy3, GSU0054, CaslO, Csm2, Cmr5, CaslO, Csxll, CsxlO, Csfl, Cas9, Csn2, Cas4, Casl2, Casl2a (Cpfl), Casl2b (C2cl), Casl2c (C2c3), Casl2d (CasY), Casl2e (CasX), Casl2f (Casl4, C2cl0), Casl2g, Casl2h, Casl2i, Casl2k (C2c5), C2c4, C2c8, and C2c9, including any naturally occurring or engineered variants of any the above CRISPR nucleases.
[0016] In some embodiments, cleaving the randomized target polynucleotides of the circular DNAs using the genome editor comprises cleaving once. In some embodiments, the CRISPR nuclease is Cas9. In some embodiments, the CRISPR nuclease is a naturally occurring or engineered variant of Cas9. In some embodiments, the randomized target polynucleotide comprises a portion that is complementary to a homology region of the guide RNA and a protospacer adjacent motif (PAM).
[0017] In some embodiments, the base editor is an adenosine base editor (ABE) or a cytosine base editor (CBE) and cleaving comprises: nicking the randomized target polynucleotides ofcircular DNAs using an ABE or CBE; and cleaving the nicked randomized target polynucleotide using a nuclease optionally EndoV (ABE) or USER (CBE) or EndoQ (ABE or CBE).
[0018] In some embodiments, the library further comprises circular DNAs comprising target polynucleotides; and cleaving further comprises cleaving target polynucleotides of the circular DNAs thereby producing target polynucleotide linearized DNAs. In some embodiments, the method comprises determining an activity of the genome editor for the target polynucleotide using the target polynucleotide linearized DNAs.
[0019] In some embodiments, the method comprises determining an activity of the genome editor for a randomized target polynucleotide using the cleaved randomized target polynucleotides of the linearized DNAs. In some embodiments, the method comprises comparing the activity of the genome editor for the target polynucleotide to the activity of the genome editor for the randomized target polynucleotide.
[0020] In some embodiments, the target polynucleotide and the randomized target polynucleotide differ from one another by 1-6 nucleotides. In some embodiments, the randomized target polynucleotides are partially randomized target polynucleotides. In some embodiments, the partially randomized target polynucleotides comprise fewer polynucleotides that will not be cleaved by the genome editor compared to completely randomized target polynucleotides.
[0021] In some embodiments, the partially randomized target polynucleotides are generated by: (a) at each nucleotide of a sequence of the target polynucleotide, changing the nucleotide to a different nucleotide in 10%-50% of the partially randomized target polynucleotides; and (b) synthesizing the sequences of the partially randomized target polynucleotides.
[0022] In some embodiments, synthesizing the sequences of the partially randomized target polynucleotides comprises synthesizing using polymerase chain reaction with a degenerate primer. In some embodiments, synthesizing the sequences of the partially randomized target polynucleotides comprises synthesizing using partially degenerate oligonucleotide synthesis. In some embodiments, each partially randomized target polynucleotide comprises no more than 10 substitutions relative to the target polynucleotide. In some embodiments, each partially randomized target polynucleotide comprises no more than 5 substitutions relative to the target polynucleotide. In some embodiments, each circular DNA of the library comprises a unique molecule identifier (UMI).
[0023] In some embodiments, identifying the randomized target polynucleotides in the cleaved linearized DNAs comprises sequencing the cleaved linearized DNAs. In some embodiments, sequencing comprises next- generation sequencing. In some embodiments, the targetpolynucleotide comprises a portion of a gene. In some embodiments, the portion of the gene comprises a disease-associated mutation.
[0024] In some embodiments, the method further comprises generating the library, the generating comprising: (a) amplifying a portion of a template DNA using: a plurality of first primers that each comprise (i) a polynucleotide that is complementary to a first portion of the template DNA, (ii) a partially randomized target polynucleotide, (iii) a unique molecular index, and (iv) a universal primer sequence; and a second primer that is complementary to a second portion of the template DNA, to produce linear precursor polynucleotides comprising a portion of the template DNA and different partially randomized target polynucleotides; and (b) circularizing the linear precursor polynucleotides to produce the circular DNAs.
[0025] In some embodiments, generating the library comprises introducing a restriction enzyme site into the linear precursor polynucleotides. In some embodiments, generating the library comprises introducing an Uracil DNA glycosylase and a DNA glycosylase-lyase Endonuclease VIII (USER) enzyme substrate (uracil) via PCR onto the ends of the linear precursor polynucleotides and cleaving the USER enzyme substrate using a USER enzyme prior to circularizing the linear precursor polynucleotides to form the circular DNAs. In some embodiments, the method further comprises (c) degrading residual linear DNA.
[0026] In some embodiments, obtaining the library comprises bottlenecking the number of different randomized target polynucleotides in the library. In some embodiments, bottlenecking comprises reducing the number of different randomized target polynucleotide in the library to less than 10 million different randomized target polynucleotides. In some embodiments, bottlenecking reduces the number of different randomized target polynucleotide in the library to less than 3 million different randomized target polynucleotides.
[0027] In some embodiments, determining the activity of the genome editor for the target polynucleotide comprises determining the efficiency with which the target polynucleotide is cleaved using the cleaved target polynucleotide linearized DNAs. In some embodiments, determining the activity of the genome editor for a randomized target polynucleotide comprises determining the efficiency with which the randomized target polynucleotide is cleaved using the cleaved randomized target polynucleotides of the linearized DNAs.
[0028] In some embodiments, determining the specificity of the genome editor comprises determining an activity of the genome editor for one or more of the randomized target polynucleotides, optionally wherein determining the activity of the genome editor for one or more of the randomized target polynucleotides comprises determining the relative activity of the genome editor for one or more of the randomized target polynucleotides.In some embodiments, determining the specificity of the genome editor comprises determining the efficiency with which a randomized target polynucleotide of the circular DNAs is cleaved and / or the number of different randomized target polynucleotides of the circular DNAs that are cleaved.
[0029] In some embodiments, the method further comprises making a copy of the library. In some embodiments, the method further comprises digesting circular DNAs of the copy of the library using a restriction enzyme to produce control linearized DNAs. In some embodiments, determining the activity of the genome editor for the target polynucleotide comprises determining a relative amount of the cleaved target polynucleotides of the linearized DNAs and the control linearized DNAs comprising the target polynucleotide.
[0030] In some embodiments, determining the activity of the genome editor for a randomized target polynucleotide comprises determining a relative amount of the cleaved randomized target polynucleotide of the linearized DNAs and the control linearized DNAs comprising the randomized target polynucleotide.
[0031] In some embodiments, determining the specificity of the genome editor for the target polynucleotide comprises comparing the activity of the genome editor for the target polynucleotide and the activity of the genome editor for one or more of the randomized target polynucleotides. In some embodiments, each circular DNA of the library comprises only one randomized target polynucleotide or only one target polynucleotide.
[0032] In some embodiments, the method further comprises: determining an activity and / or specificity of a genome editor for a second target polynucleotide, the method comprising: (a) obtaining a library comprising circular DNAs, each circular DNA comprising a second randomized target polynucleotide; (b) cleaving second randomized target polynucleotides of circular DNAs of the library using the genome editor, thereby producing different linearized DNAs; and (c) determining the activity and / or specificity of the genome editor for the second target polynucleotide using the cleaved second randomized target polynucleotides of the linearized DNAs.
[0033] In some embodiments, the method further comprises comparing the activity and / or specificity for the genome editor for the target polynucleotide with the activity and / or specificity of the genome editor for the second target polynucleotide, optionally wherein comparing comprises comparing using a ratio or difference.
[0034] In some embodiments, this disclosure provides a method for determining an activity and / or specificity of a genome editor for a target polynucleotide, comprising: (a) obtaining a library comprising circular DNAs, each circular DNA comprising a target polynucleotide orpartially randomized target polynucleotide and a restriction site, wherein obtaining the library comprises obtaining a bottlenecked library that comprises no more than 10 million different partially randomized target polynucleotides; (b) obtaining a copy of the library; (c) cleaving partially randomized target polynucleotides of circular DNAs of the library using the genome editor, thereby producing cleaved linearized DNAs; (d) digesting circular DNAs of the copy of the library using a restriction enzyme to produce control linearized DNAs; and (e) determining the activity and / or specificity of the genome editor for the target polynucleotide using the cleaved linearized DNAs and the control linearized DNAs.
[0035] In some embodiments, the partially randomized target polynucleotides are generated by: (a) at each nucleotide of a sequence of the target polynucleotide, changing the nucleotide to a different nucleotide in 3%-50% of the partially randomized target polynucleotides; and (b) synthesizing the sequences via PCR of the partially randomized target polynucleotides.
[0036] In some embodiments, this disclosure provides a method for determining an activity of a genome editor for one or more randomized target polynucleotides, comprising: (a) obtaining a library comprising circular DNAs, each circular DNA comprising a randomized target polynucleotide; (b) cleaving randomized target polynucleotides of circular DNAs of the library using the genome editor, thereby producing linearized DNAs; and (c) determining the activity of the genome editor for the one or more randomized target polynucleotides using the cleaved randomized target polynucleotides of the linearized DNAs.
[0037] In some embodiments, this disclosure provides a method for determining an activity of a genome editor for one or more partially randomized target polynucleotides, comprising: (a) obtaining a library comprising circular DNAs, each circular DNA comprising a partially randomized target polynucleotide and a restriction site, wherein obtaining the library comprises obtaining a bottlenecked library that comprises no more than 10 million different partially randomized target polynucleotides; (b) obtaining a copy of the library; (c) cleaving partially randomized target polynucleotides of circular DNAs of the library using the genome editor, thereby producing cleaved linearized DNAs; (d) digesting circular DNAs of the copy of the library using a restriction enzyme to produce control linearized DNAs; and (e) determining the activity of the genome editor for the one or more target polynucleotides or randomized target polynucleotides using the cleaved linearized DNAs and the control linearized DNAs.
[0038] In some embodiments, this disclosure provides a method for determining an activity and / or specificity of a guide RNA (gRNA) for a target polynucleotide, comprising: (a) obtaining a library comprising circular DNAs, each circular DNA comprising a randomized target polynucleotide; (b) cleaving randomized target polynucleotides of circular DNAs of the libraryusing an RNA-guided nuclease and the gRNA, thereby producing linearized DNAs; and (c) determining the activity and / or specificity of the gRNA for the target polynucleotide using the cleaved randomized target polynucleotides of the linearized DNAs.
[0039] In some embodiments, the method further comprises determining a cleavage site of the genome editor in one or more randomized target polynucleotides of circular DNAs of the library or in a target polynucleotide of circular DNAs of the library. In some embodiments, the method further comprises determining whether the cleavage site is a staggered cleavage site or a blunt end cleavage site. In some embodiments, the method further comprises determining the overhang length of the staggered cleavage site. In some embodiments, the method further comprises determining an activity of the genome editor for at least two randomized target polynucleotides. In some embodiments, the method comprises determining whether the cleavage site is a staggered cleavage site or a blunt end cleavage site for at least 10, at least 50, at least 100, at least 1000, at least 10,000, or at least 100,000 randomized target polynucleotides.
[0040] In some embodiments, the method comprises determining an amount of randomized target polynucleotide whose cleavage results in a staggered cleavage site, and / or determining an amount of randomized target polynucleotide whose cleavage results in a blunt end cleavage site. In some embodiments, the method comprises determining a ratio between the amount of randomized target polynucleotide whose cleavage results in a staggered cleavage site and the amount of randomized target polynucleotide whose cleavage results in a blunt end cleavage site In some embodiments, the method further comprises comparing the activity of the genome editor for the at least two randomized target polynucleotides.
[0041] In some embodiments, the sequences of the least two randomized target polynucleotides are within a hamming distance of 6 or less. In some embodiments, the sequences of the least two randomized target polynucleotides are within a hamming distance of 2 or less. In some embodiments, the sequences of the least two randomized target polynucleotides differ by a hamming distance of 1 or less. In some embodiments, the method further comprises determining an effect of a specific mutation on the activity of the genome editor for the at least two randomized target polynucleotide using the activity of the genome editor for the at least two randomized target polynucleotides.
[0042] In some embodiments, determining activity comprises determining relative activity. In some embodiments, the this disclosure provides a method for determining an activity and / or specificity of a non-nicking base editor for a target polynucleotide, comprising: (a) obtaining a library comprising circular DNAs, each circular DNA comprising a randomized target polynucleotide; (b) editing a strand of the randomized target polynucleotides of circular DNAsof the library using the genome editor, thereby producing a mismatched strands; (c) cleaving the mismatched strands using a mismatch endonuclease I or nuclease that recognizes mismatched DNA bases, thereby producing linearized DNAs; and (d) determining the activity and / or specificity of the genome editor for the target polynucleotide using the cleaved randomized target polynucleotides of the linearized DNAs.
[0043] BRIEF DESCRIPTION OF THE DRAWINGS FIG. 1 shows an overview of the experimental design to quantify the frequency and magnitude of effects of individual genetic variation on CRISPR-Cas9 nuclease genome editing activity. First, an optimized, streamlined, high-throughput GUIDE- seq-2 method was developed to rapidly profile cellular on- and off-target editing activity. Second, GUIDE-seq-2 was applied on a population scale to discover Cas9 genome-wide editing activity associated with 6 gRNA targets in cells from 95 individuals from 4 populations characterized by the 1000 Genomes Project. Third, multiplexed targeted sequencing was performed to validate and quantify editing at the GUIDE-seq-2 sites identified in a population-scale survey. Fourth, by integrating orthogonal measures of variant effects, a high confidence set of variant effects on Cas9 editing activity was defined. Fifth, to more comprehensively survey the off-target landscape, a massively parallel biochemical assay, CHANCE-seq, was developed to interrogate the biochemical activity of Cas9 across millions of mismatched off-target sequences.
[0044] FIGs. 2A-2E show development and optimization of GUIDE-seq-2, a highly scalable, sensitive, and unbiased method to define cellular genome-wide activity of genome editing nucleases. FIG. 2A shows a schematic overview of streamlined GUIDE-seq-2 library preparation workflow. A short end-protected double-stranded oligodeoxynucleotide (dsODN) tag is integrated into sites of Cas9-induced DNA double-stranded breaks in living cells.
[0045] Genomic DNA isolated from these cells is tagmented to provide a universal priming site. A single round of tag-specific PCR is performed to amplify tag-containing DNA for high-throughput sequencing. FIG. 2B shows scatterplots of GUIDE-seq and GUIDE-seq-2 read counts (log 10) from experiments performed on primary human T cells for 8 target sites, r is Pearson’s correlation coefficient. FIG. 2C shows scatterplots of GUIDE-seq-2 read counts (log 10) between two GUIDE-seq-2 libraries independently prepared from experiments performed in primary human T cells for 8 target sites, r is Pearson’s correlation coefficient. FIG. 2D shows Manhattan plots showing the genome-wide distribution of GUIDE-seq-2 detected off-target sites with chromosomal position (x-axis) and read counts (y-axis). Arrow indicates the intended on-target site. FIG. 2E shows the sensitivity of GUIDE-seq-2 establishedby determining lower limit of detection (LLOD) and quantification (LLOQ) X-axis: expected fraction of off-target in sample. Y-axis: GUIDE-seq-2 readcount. n=3.
[0046] FIGs. 3A-3O show cellular off-target editing sites detected by GUIDE-seq-2 frequently overlap genetic variants. FIG. 3A shows a stacked bar plot of number of GUIDE-seq-2 on- or off-targets sites for each of 6 gRNA targets and proportion of sites that overlap with genetic variation, n=95 individuals. FIG. 3B shows a pie chart of genomic distribution of off-targets detected in our study cohort. Untranslated Region (UTR). FIG. 3C shows a Manhattan plot showing off-targets with chromosomal position versus GUIDE-seq-2 read count, highlighted by presence of genetic variation. Off-target sites overlapping with genetic variants are represented by dark circles. On-target sites are highlighted as squares, on-target that overlaps genetic variation as a circled square. FIG. 3D shows a heatmap of genotypes of variant-containing off-targets found in the study cohort. Horizontal rows represent individuals, Vertical columns represent one on- or off-target site. Off-target genotypes shown: off-target of homozygous reference (0), heterozygous (1), homozygous alternative (2), multiple variants with multiple combinations (3). African Ancestry in Southwest USA (ASW), Northern Europeans from Utah (CEU), Han Chinese in Beijing, China (CHB), Mexican Ancestry in Los Angeles, CA (MXL). FIG. 3E shows a pie chart representing fraction of cellular off-target sites that overlap genetic variants and bar plot showing number of genetic variants per population in this study cohort (n=95 individuals) FIG. 3F shows the same analysis as FIG. 3E, extending to all variants in the 1,000 Genomes Project (n=3,202 individuals). South Asian (SAS), East Asian (EAS), (EUR), Admixed American (AMR), African / African American (AFR). FIG. 3G shows the same analysis as FIG. 3E, extending to all variants in gnomAD (n>80,000 individuals), Amish (AMI), Middle Eastern (MID), Ashkenazi Jewish (ASJ), Finnish (FIN), Non-Finnish European (NFE). FIG. 3H shows a bar plot indicating the proportion of genetic variants that contained deletions, insertions or SNPs (n=212, total overlapped variants), 1,000 Genomes Project (n=622, total overlapped variants), and gnomAD database (n=6,311, total overlapped variants). FIG. 31 shows a bar plot indicating the proportion of SNPs that were transitions or transversions (n=194), 1,000 Genomes project (n=579), and gnomAD database (n=5,567). FIG. 3J shows a bar plot indicating the proportion of genetic variants that were common or rare (n=212), in 1,000 Genomes project(n=622), and gnomAD database (n=6311), rare=minor allele frequency (MAF) less than 1%. FIG. 3K is a bar plot showing on- / off-target sites overlapping with different numbers of variants. FIG. 3L is a bar plot showing distribution of variants across the protospacer and PAM sequence. Horizontal dotted line represents average number of variants across all positions. FIG. 3M is a bar plot showing an example of a variant found predominantlywithin one genetic superpopulation. Visualization showing the on-target, reference variant and non-reference variant at position 2. FIG. 3N is a violin plot showing an estimated percentage of people affected by common variant-containing off-targets. FIG. 30 is a line plot showing percentage of variants covered based on bootstrap sampling from the 1000 Genomes Project (n=3202 individuals). Lines represent different variant set: “All variants” is the total 622 variants overlapped with 1,115 GUIDE-seq-2 sites. “0.001” is the set of 294 variants with MAF>0.1%. “0.01” is the set of 131 variants with MAF>1%. “0.05” is the set of 81 variants with MAF>5%. The horizontal dotted line represents 95% variant coverage. The vertical solid line marks the study size (n = 95). The vertical dotted line marks the sample size required (n = 55) to achieve 95% variant coverage for common variants. X-axis represents sampling size (i.e., number of individuals) drawn on log 10 scale.
[0047] FIGs. 4A-4M show overlapping genetic variants frequently affect editing activity. FIG.
[0048] 4A is a schematic of validation experiments. On- and off-targets sites were identified by GUIDE-seq-2, a subset were validated using multiplex targeted sequencing. Statistical analyses were performed to quantify variant effect across orthogonal measures. FIG. 4B shows a heatmap of GUIDE-seq-2 activity data, colored by Z-score, for six gRNA targets in 95 individuals across four 1000 Genomes Project populations: ASW, CEU, CHB, and MXL. Each row represents an individual and each column represents an on- or off-target site. FIG. 4C shows a scatterplot showing GUIDE-seq-2 limits of detection across 95 individuals. Each dot represents an on / off-target site, y-axis represents percentage of individuals having the on- / off-target. x-axis represents normalized GUIDE-seq-2 activity based on on-target indel frequency. Box represents the GUIDE-seq-2 limits of detection for off-target sites (without genetic variants) detected in 95% of individuals with -0.2% indel frequency. FIG. 4D shows volcano plots showing variant effects on GUIDE-seq-2 read counts, indels, integration, allelic imbalance (x-axis). Y-axis is statistical significance (negative log false discovery rate, FDR). White dots represent significantly increased variant activity (FDR<=0.1) and dark grey dots represent significantly decreased variant activity (FDR<=0.1). Absolute difference >=0.5 threshold is used in addition to FDR for GUIDE-seq-2 read counts, indels, and integration. FIG. 4E is a pie chart showing the proportion of high confidence variants that increase or decrease editing activity out of the 167 variants. FIG. 4F shows a detailed view of the 21 high-confidence variants, defined based on both GUIDE-seq-2 and multiplexed targeted sequencing data, sorted based on indel% from increasing to decreasing. High-confidence variant effects have significant differences in two or more variant effect measures, and variant positions are outlined in black. Off-targets sequence alignment is shown to the left of the on-target gRNA names. Shaded boxes representmismatches to on-target sequence. The alternative (non-reference) allele is shown on the bottom. Variants are labelled from 1 to 21, next to the alignment plot. Sequence number 9 has a 1 bp A insertion, represented by a smaller square at position 13. *On-target sequence contains a variant. Bar plots showing variant effects of GUIDE-seq-2 read counts (log), indel% (log), integration* ? (log), and variant allelic imbalance for indel-reads and tag-reads in the heterozygous individuals. n=95 for GUIDE-seq-2, indel, integration analyses. Sample size varies based on the number of heterozygous individuals for variant allele imbalance analyses. Increased effect and decreased effect are indicated. FIG. 4G shows detailed bar plots of unexpected and expected variant effect examples. On-target site shown above, mismatched nucleotides below. Variant nucleotide boxed. Student’s t-test and fold change are calculated between homozygous reference and homozygous alternative. ***p <0.0001. FIG. 4H shows a bar plot of variant effects in human primary T cells. Fold change between homozygous reference (bottom) and homozygous alternative (top) are shown to the right. **p < 0.001 ***p < 0.0001. FIG. 41 shows a scatter plot showing allele frequency of the 212 variants detected by GUIDE-seq-2, indicated by variants that are de novo, increase, or decrease editing activity r=Pearson’s. FIG. 4J Pie chart showing proportion of 212 variants detected by GUIDE-seq-2 that are de novo, increase, or decrease editing activity. FIG. 4K shows a detailed view of variants, defined by GUIDE-seq-2 and multiplexed targeted sequencing data, sorted based on GUIDE-seq-2 read counts. Off-target sequence alignment is shown to the right of the on-target gRNA names. Shaded boxes represent mismatches to on-target sequence. Variant positions are boxed in black with alternative (nonreference) allele on the bottom. De novo and high-confidence variants are shown with a check mark. Variant 12 has a 1 bp insertion, represented by a smaller square. On-target sequence contains a variant. Bar plots show variant effects in the heterozygous individuals. n=95 for GUIDE-seq-2, indel, integration analyses. Sample size varies based on the number of heterozygous individuals for variant allele imbalance analyses. Increased effect and decreased effect are indicated. FIG. 4L shows correlation between human primary T cells, (3 donors, 2 replicates each) and LCLs, 95 donors, for off-targets without variants, r= Pearson correlation coefficient. FIG. 4M shows a detailed view of variants, defined by GUIDE-seq-2 and multiplexed targeted sequencing data, sorted based on GUIDE-seq-2 read counts. Similar to FIG. 4K but with variants added only identified by GUIDE-seq-2 and not multiplexed targeted sequencing. Also includes variants with deletions indicated by black lines in the sequence alignment.
[0049] FIGs. 5A-5Q show a massively parallel biochemical assay, CHANCE-seq, reveals key genetic features that impact Cas9 activity. FIG. 5A is a schematic of CHANCE-seq assay whichcreates libraries of partially randomized target sites and selectively sequences those that are cut by Cas9 compared to a control. Gray indicates target, black is plasmid backbone, squares in gray region are randomized mismatches in protospacer and PAM. The same partially randomized target sequences are analyzed across control and Cas9 treatments. A control restriction enzyme cleaves a site adjacent to all randomized targets, whereas only select molecules are cleaved with Cas9. FIG. 5B shows a histogram of frequency of mismatches in CHANCE-seq libraries treated with Cas9 versus control libraries. Average of two Cas9 technical replicates for each target and five control libraries is shown. FIG. 5C shows a binned scatterplot of % relative activity between replicate 1 and replicate 2 for Cas9 CHANCE-seq library for CTLA4, indicated by density, r = Pearson’s. FIG. 5D shows a heatmap of inferred PAM specificity from CHANCE-seq, colored by percent relative activity. FIG. 5E shows a bar plot showing number of unique sequences at each mismatch level in CHANCE-seq, the line shows percent coverage of unique sequences. Error bars are standard deviation of 5 targets in control libraries. FIG. 5F shows violin plots showing relative activity % at each mismatch level 0-6 for CTLA4 target using CHANCE-seq. FIG. 5G shows a line graph showing distribution of mismatch counts with CHANCE-seq for CTLA4 target along the protospacer with NGG PAMs. FIG. 5H shows a Stripplot illustrating contiguous mismatch start positions (2-6 mismatches) with % relative activity on a log scale. Only longest contiguous block (length of block equal to number of mismatches) is shown. FIG. 51 shows a bar plot showing effect of mutation type on activity. All CHANCE-seq off-targets were divided into 5 groups based on relative activity. For each activity group, the percentage of off-targets falling into 1 of 5 mismatch categories is shown. FIG. 5J shows a global analysis showing loglO relative activity for CTLA4, using CHANCE-seq, stratified by mismatch. The average of 2 technical Cas9 replicates was taken to determine relative activity percentage. Plots are divided into bins based on activity. Transitions and transversions are are indicated, with transparency at each position proportional to position mismatch frequency within each bin. On-target sequence is shown at the bottom of the plot. FIG. 5K shows a visualization of sequences that had 10 highest and lowest relative activity for CTLA4 using CHANCE-seq Sequences are stratified by mismatch number and the average of 2 technical Cas9 replicates was taken to determine relative activity. The top sequence is the on-target site, each subsequent sequence shows where mismatches occur. The number to the right shows the percent relative activity for each sequence. FIG. 5L shows a schematic overview of CHANCE-seq variant pairs analyzed. FIGs. 5M-5N show heatmaps showing differences in activity when increasing the number of mismatches from 1 to 2 for CFD (FIG. 5M) and CHANCE-seq (FIG. 5N). Both reference and variant sequences have one prior mismatch ateach of 19 other positions, but variant sequences have one additional variant mismatch at pos I6:G->A (top) or pos I :G->A (bottom) (boxed). The percent activity change is calculated as (ref-var) / ref. FIG. 50 shows a MA plot illustrating synergistic effects of combined mismatches. Each dot represents a pair of mismatches. The x-axis shows the mean of the expected and observed activity changes, and the y-axis shows the log2(expected / observed) ratio. A pseudocount of 0.001 was added to both expected and observed activity changes. Positive values indicate synergistic effects, where combined mismatches lead to a greater reduction in activity than expected under the independence assumption, while points near zero represent neutral effects. FIG. 5P shows an example CTLA4 sequences of mismatches with high synergy verses those with neutral effects with activity values. Sequences chosen from FIG. 51. FIG. 5Q shows a bar plot with expected activity change vs observed when there is high synergy vs neutral mismatch effects.
[0050] FIGs. 6A-6E show genetic features that impact effect of overlapping genetic variants on off-target activity. FIG. 6A is a schematic overview of CHANCE-seq variant pairs analyzed. FIG. 6B is a scatterplot of effect of single mismatches on relative activity in variant pairs.
[0051] Reference is the off-target in pair with less mismatches, variant is the off-target with increased mismatches. FIG. 6C shows changes in relative activity (delta activity) for variants with an increase in edit distance. Mismatch numbers (off-target a variant) are indicated. A histogram on the right displays the distribution of delta activity on a logarithmic scale. FIG. 6D shows an identification of high-impact, single edit distance, off-target pairs within CHANCE-seq datasets for the CTLA4 target. Stratified by mismatch, variants increase mismatch number by 1. The top sequence is the on-target, followed by off-target / variants. Boxes with no nucleotide are the original off-target sequence, variants are represented by nucleotide in the box. Absolute relative activity between on-target and variant is shown to the right of each sequence. Organized by variants with highest delta activity between off-target and variant. FIG. 6E shows positions and mismatch types with high effect on activity within the CHANCE-seq datasets. Stratified by mismatch number, context of off-target is shown on the left, variant is shown on the right. Size of x’s and circles represent the percentage of off-target contexts or variants at each position containing the transition mismatches. The % of each mismatch or variant that is a transition mutation is indicated.
[0052] FIGs. 7A-7B show a schematic overview of GUIDE-seq and GUIDE-seq-2. FIG. 7A shows both GUIDE-seq and GUIDE-seq-2 begin with cells transfected with a nuclease genome editor and a GUIDE-seq double-stranded oligodeoxynucleotide (dsODN) tag. GUIDE-seq requires fragmentation of genomic DNA (typically by physical shearing with sonication)followed by end-repair / A-tailing, ligation, and nested PCR. GUIDE-seq-2 eliminates the requirement for specialized equipment for physical DNA shearing along with 6 additional molecular biology or purification steps by leveraging Tn5 transposase for library preparation and eliminating the requirement for nested PCR. First, genomic DNA from GUIDE-seq dsODN tag integrated cells is tagmented with unique-molecular index containing transposomes.
[0053] Tagmented gDNA is then subjected to a single round of PCR amplification, using a tag-specific primer combined with the i7 barcode primer, obviating the need for a second round of PCR. FIG. 7B shows GUIDE-seq and GUIDE-seq-2 forward and reverse tag- specific primers are shown. Use of a single PCR step for library preparation enabled a primer design capable of distinguishing mispriming artifacts without changing the length of GUIDE-seq tag and affecting the tag integration efficiencies. Bases expected to be observed with the new primer design in the GUIDE-seq-2 are underlined. The simplified GUIDE-seq-2 workflow substantially streamlines the process and enables high-throughput experiments, while also decreasing the requirement of input genomic DNA for library preparation by approximately 4-fold.
[0054] FIGs. 8A-8E show automation of library preparation and optimization for lymphoblastoid cell lines. FIG. 8A shows scatterplots of GUIDE-seq-2 read counts (log scale) between beads-based library normalization and regular SPRI beads purification. FIG. 8B shows scatterplots of GUIDE-seq-2 read counts (log scale) between two independently prepared beads-based library normalization libraries. FIG. 8C is a schematic representation of automated GUIDE-seq-2 library preparation workflow. FIG. 8D shows scatterplots of GUIDE-seq-2 read counts (log scale) between GUIDE-seq-2 libraries prepared using the automated workflow and the standard manual workflow. FIG. 8E shows scatterplots of GUIDE-seq-2 read counts (log scale) between two independently prepared GUIDE-seq-2 libraries using the automated workflow.
[0055] FIGs. 9A-9E show GUIDE-seq-2 optimization for lymphoblastoid cell lines. Viability (FIG. 9A) and live cell count (FIG. 9B), of lymphoblastoid cell line (LCL) population assessed 3 days post nucleofection with different doses of dsODN (n=2). FIG. 9C shows indels rates at the intended target site 3 days post nucleofection with different doses of dsODN. FIG. 9D shows dsODN integration rates 3 days post nucleofection with different doses of dsODN. FIG.
[0056] 9E shows scatterplots of GUIDE-seq-2 read counts (log scale) between two independently prepared GUIDE-seq-2 libraries for two sgRNA target sites across three LCL donors, showing GUIDE-seq-2 technical reproducibility, r is Pearson’s correlation coefficient.
[0057] FIGs. 10A-10D show GUIDE-seq-2 visualization plots for 3 targets. The intended on-target site is shown at the top of the alignment visualization. Each row shows a genomic off-target site with average of GUIDE-seq-2 read counts and standard deviation (n=95) shown to the right. Mismatches to the intended target site are shown as boxed nucleotides, matches are indicated with a dot. The first nucleotide of the NGG PAM sequence is shown without shading. The same genomic off-target site but with an overlapping variant sequence is shown further to the right. The average and standard deviation for the variant group (i.e., donors of heterozygous or homozygous alternative genotype) is shown next to it.
[0058] FIGs. 11A-11B show correlations of GUIDE-seq-2 read counts with indel mutations in LCLs and between donors in human primary T cells. FIG. 11A shows a correlation of on-target GUIDE-seq-2 read counts and indel (%) for the six sgRNAs target sites evaluated across the 95 lymphoblastoid cell line (LCL) donors, r is Pearson’s correlation coefficient. FIG. 11B shows a correlation of GUIDE-seq-2 read counts between different human primary T cell donors, dots labeled “OT-#” are the examples of off-targets containing SNPs that are visualized in detail in FIG. 4H.
[0059] FIG. 12 shows visualizations of off-targets with high-confidence variant effects. Each off-target visualization panel contains: a sequence-alignment plot, a bar plot of GUIDE-seq-2 read counts, a bar plot of indel percentage, and a bar plot of dsODN integration percentage (from left to right). From top to bottom in the sequence- alignment plot shows the on-target sequence (i.e., ON) and the off-targets sequences corresponding to homozygous reference genotype (i.e., 0), heterozygous genotype (i.e., 1, if exist), and homozygous alternative genotype (i.e., 2, if exist). Fold change comparing homozygous reference genotype and homozygous alternative genotype is shown. Student’s t-test significance is shown. *p-value<=0.01. **p-value<=0.001. ***p-value<=0.0001. inf= fold change infinity. Variant ID 19 and 20 occur in the same off-target, therefore, multiple genotypes exist. K: T or G. M: A or C. S: C or G. Y: C or T. R: A or G. Off-target annotation is shown on the bottom, including variant ID (corresponding to FIGs.
[0060] 4A-4H), variant location (hg38), off-target coordinate, and on-target name.
[0061] FIGs. 13A-13F show CHANCE-seq detailed schematic and plots. FIG. 13A shows a detailed schematic of CHANCE-seq assay. Composition of primers are color coded. In the circles, gray represents the target site, black is the plasmid backbone, squares in the gray region represent randomized mismatches in protospacer and PAM. PCR1 creates a diverse library using mixed base primers and plasmid DNA as template. After PCR1, the linear randomized target library is bottlenecked to ~3E6 unique molecules to copy in PCR2. PCR2 products are circularized, and the library is then split between Cas9 and restriction enzyme control ensuring the same composition in both samples. Restriction enzyme cleaves all randomized targets, select molecules are cleaved with Cas9. Only cleaved circles are prepared for Illumina sequencing andactivity for each sequence is calculated. FIG. 13B shows a hexbin plot of relative activity between replicate 1 and replicate 2 for Cas9 CHANCE-seq library using gRNA. Each dot is an off-target. 4 targets are shown. FIG. 13C shows violin plots showing % relative activity at each mismatch level 0-6 for AAVS1, CCR5, LAG3 and TRAC target using CHANCE-seq. . FIG. 13D is a line graph showing distribution of mismatch counts with CHANCE-seq for AAVS1, CCR5, LAG3 and TRAC targets along the protospacer with NGG PAMs, by density. FIG. 13E shows a strip plot illustrating contiguous mismatch start positions (2-6 mismatches) with % relative activity on a log scale. Values plotted for AAVS1, CCR5, LAG3 and TRAC targets. FIG. 13F shows a heatmap of mutation types by activity level and protospacer position. Circle size represents frequency of any mismatch at each position within each activity group and color indicates the percentage of specific mismatch type at each position.
[0062] FIGs. 14A-14D show CHANCE-seq global visualizations for 4 targets. Global analysis showing loglO relative activity for, AAVS1 (FIG. 14A), CCR5 (FIG. 14B), LAG3 (FIG. 14C) and, TRAC (FIG. 14D) using CHANCE-seq, stratified by mismatch. The average of 2 technical Cas9 replicates was taken to determine relative activity percentage. Plots are divided into bins based on activity. Transitions and transversions are indicated, with transparency at each position proportional to position mismatch frequency within each bin. On-target sequence is shown at the bottom of the plot.
[0063] FIGs. 15A-15D show a visualization of sequences that had the top 10 highest and lowest relative activity for, AAVS1 (FIG. 15A), CCR5 (FIG. 15B), LAG3 (FIG. 15C), TRAC (FIG.
[0064] 15D) using CHANCE-seq. Sequences are stratified by mismatch number and the average of 2 technical Cas9 replicates was taken to determine relative activity. The top sequence is the on-target, each subsequent sequence shows where mismatches occur. The number to the right shows the percent relative activity for each sequence.
[0065] FIG. 16 shows GUIDE-seq-2 profile correlates well between inhouse Tn5 and commercially available Tn5 Scatterplots of GUIDE-seq-2 read counts (loglO) from experiments performed on primary human T cells for 4 target sites. Genomic DNA were tagmented using inhouse Tn5 (x-axis) and seqWell Tagify-UMI (y-axis) before GUIDE-seq-2 PCR. GUIDE-seq-2 library prepared with each method was sequenced individually with NextSeq2000. r = Pearson’s correlation coefficient.
[0066] FIGs. 17A-17B show visualizations of off-targets with variant effects defined by GUIDE-seq-2 and / or multiplexed targeted sequencing. Each off-target visualization panel contains: a sequence- alignment plot, bar plots of GUIDE-seq-2 read counts, indel percentage, dsODN integration percentage, and indel percentage from a no editor control (from left to right).From top to bottom in the sequence-alignment plot shows the on-target sequence (i.e., ON) and the off-targets sequences corresponding to homozygous reference genotype (0), and may contain heterozygous genotype (1), and homozygous alternative genotype (2). Fold change comparing homozygous reference genotype and homozygous alternative genotype is shown. Student’s t-test significance is shown. *p-value<=0.01. **p-value<=0.001. ***p-value<=0.0001. Variant ID S2 and S3 occur in the same off-target, therefore, multiple genotypes exist. IUPAC DNA codes are: K: T or G. M: A or C. S: C or G. Y: C or T. R: A or G. N: A, C, G, or T. Off-target annotation is shown on the bottom, including variant ID (corresponding to figure 4f), variant location (hg38), off-target coordinate, DNA strand (+ or -), and on-target name.
[0067] FIGs. 18A-18E show visualization of sequences that had the top 10 highest and lowest relative activity for AAVS1 (FIG. 18A), CCR5 (FIG. 18B), CTLA4 (FIG. 18C), LAGS (FIG.
[0068] 18D), and TRAC (FIG. 18E), using CHANCE-seq. Sequences are stratified by mismatch number and the average of 2 technical Cas9 replicates was taken to determine relative activity. The top sequence is the on-target, each subsequent sequence shows where mismatches occur. The number to the right shows the percent relative activity for each sequence.
[0069] FIGs. 19A-19C shows GUIDE-seq identified off-targets applied to 1,000 Genomes Project for variant discovery. FIG. 19A shows a violin plot showing percentage of population affected by a given off-target affecting variant. Each dot is a variant. FIG. 19B shows a violin plot showing percentage of population affected stratified by on-target specificity. The three groups are based on the number of off-targets per on-target. Low group contains the most specific on-targets, defined as off-targets per on-target less than 20. The total number of off-target affecting variants is 26. Median group contains on-target with off-targets numbers range from 20 to 200. The total number of off-target affecting variants is 62. High group contains on-target with off-targets numbers more than 200. The total number of off-target affecting variants is 118. Significance determined by 2-way ANOVA. FIG. 19C shows a violin plot showing percentage of off-targets overlapping with variants on a per individual basis. Each dot is a donor from the 1000 Genome Project.
[0070] FIGs. 20A-20F. FIG. 20A shows a schematic of CHANCE-seq ABE assay which creates libraries of partially randomized target sites via PCR and selectively sequences those that are modified by ABE compared to a control. DNA is circularized by intramolecular ligation. Residual linear DNA is enzymatically degraded with an exonuclease cocktail, leaving highly pure DNA circles. Gray indicates on- or off-target, black is plasmid backbone, squares in gray region are randomized mismatches in protospacer and PAM. A control restriction enzyme cleaves a site adjacent to all randomized targets, whereas only select molecules are cleaved withnCas9-ABE. The same partially randomized target sequences are analyzed across control and Cas9 treatments. nCas9-ABE nicks the target DNA strand and deaminates adenine bases to inosine (within the base-editing window) on the nontarget strand at on-target and off-target sites. Nicked inosine-containing DNA circles are then processed with endonuclease V, which cleaves DNA adjacent to inosines. Cleaved circles are prepped for next generation sequencing. FIG. 20B shows a histogram of frequency of mismatches in CHANCE-seq ABE libraries treated with Cas9-ABE8e versus control libraries. FIG. 20C shows a bar plot of number of unique sequences at each mismatch level in CHANCE-seq ABE, the line shows percent coverage of unique sequences. Error bars are standard deviation of 2 targets in control libraries. FIG. 20D shows a violin plot for CBLB and CD7 targets showing relative activity % at each mismatch level 0-6 using CHANCE-seq ABE. FIG. 20E shows a visualization of truncated list of off-target sites identified by CHANCE-seq ABE aligned against the intended target site of CBLB or CD7. Activity is calculated relative to the on- target which is 100%. The intended target sequence is shown in the top line and off-target sites are ordered from top to bottom by CHANCE-seq ABE relative activity %. Mismatches to the intended target sequence are indicated by colored nucleotides; matches are shown as dots. FIG. 20F shows a representative alignment of CHANCE-seq ABE reads of on-targets for CBLB and CD7, as visualized by the Integrative Genomics Viewer (IGV). The target adenine base is underlined, the protospacer-adjacent motif (NGG) is shown in bold and the single-stranded DNA breaks are marked with colored arrows by nCas9 (bottom) and endonuclease V (top).
[0071] DETAILED DESCRIPTION
[0072] Provided herein, in some aspects, are methods for determining an activity and / or specificity of a genome editor for a target polynucleotide. Genome editors are typically used to alter a gene sequence in a genome, but they can also be used on non-genomic DNAs. Genome editors typically recognize a target polynucleotide and edit (e.g., change the sequence of) that target polynucleotide. For example, RNA guided nucleases (RGNs) bind to target polynucleotide via a guide RNA, which comprises a sequence that is complementary to the target polynucleotide, leading to cleavage of the target polynucleotide by the nuclease. Repair of the cleaved target nucleotide can result in an edit to the target polynucleotide. In another example, base editors comprise a catalytically inactive RGN that binds to target polynucleotide via a guide RNA, which can lead to editing of the target polynucleotide by the base editor.
[0073] However, genome editors can have off-target effects where the gene editor edits a polynucleotide that is not the target polynucleotide. Off-target editing is typically undesirable inboth experimental and clinical settings. Thus, there is a need for methods that determine the activity (e.g., cleaving target polynucleotides) and / or specificity (e.g., cleaving off-target polynucleotides) of a gene editor for a target polynucleotide.
[0074] Previous methods for determining the activity and / or specificity for a genome editor have significant drawbacks that include but are not limited to relatively high expense, complexity of methodologies, and time-consuming molecular biology protocols, and / or limitations in application between different genomes. For example, genomic heterogeneity between different subjects (e.g., humans) can result in different potential off-target sites associated with the same gene editor targeting the same target sequence. However, systematically and accurately checking for all of these potential off-target effects in vivo using current techniques is impractical given the diversity of the human genome and the need to obtain relevant cell types from many individuals. To address this problem, this disclosure includes unbiased, high throughput methods for determining genome editor activity and specificity across millions of diverse mismatched target sites. The results of such methods can be used to make unbiased determinations about potential off-target genome editing for a given genome editor (e.g., a Cas9 RNP) across a variety of different genomes.
[0075] In some embodiments, this disclosure provides a method for determining an activity and / or specificity of a genome editor for a target polynucleotide, comprising: (a) obtaining a library comprising circular DNAs, each circular DNA comprising a randomized target polynucleotide; (b) cleaving randomized target polynucleotides of circular DNAs of the library using the genome editor, thereby producing linearized DNAs; and (c) determining the activity and / or specificity of the genome editor for the target polynucleotide using the cleaved randomized target polynucleotides of the linearized DNAs.
[0076] A “genome editor” includes a molecule (e.g., a protein enzyme) that can cleave a polynucleotide (e.g., a single strand cleavage or a double strand cleavage). In some embodiments, a genome editor refers to a ribonucleotide protein complex (e.g., a Cas9-guide RNA complex). A protein component of a gene editor is referred to as the gene editor protein component (e.g., a Cas9). A nucleotide (e.g., RNA / DNA) component of a gene editor is referred to as the gene editor nucleotide component (e.g., a guide RNA). In some embodiments, a genome editor can introduce one or more of an insertion, deletion, substitution, transversion, or transition into a polynucleotide. In some embodiments, the genome editor comprises a base editor e.g., an adenosine base editor or a cytosine base editor or equivalents, derivatives, or functional variants thereof. Base editors typically comprise a nucleic acid guided nuclease (e.g., a nickase), a nucleic acid that guides the nucleic acid guided nuclease (e.g., a guide RNA) to atarget polynucleotide (e.g., a target DNA), and a deaminase (e.g., an adenosine deaminase or a cytidine deaminase) In some embodiments, the genome editor comprises a zinc-finger nuclease (ZNF) or equivalents, derivatives, or functional variants thereof. In some embodiments, the genome editor comprises a transcription activator-like effector nuclease (TALEN) or equivalents, derivatives, or functional variants thereof. In some embodiments, the genome editor comprises a meganuclease or equivalents, derivatives, or functional variants thereof. In some embodiments, the genome editor is a prime editor (e.g., PEI, PE2, PE3, PE4, PE5, or a Nuclease Prime Editor or equivalents, derivatives, or functional variants thereof). A prime editor typically comprises a nucleic acid guided nuclease (e.g., a nickase), a reverse transcriptase, and a prime editing guide RNA, which comprises a guide sequence, a primer binding site, and a template sequence with the desired edit for incorporation into the target polynucleotide. In some embodiments, the genome editor is a CRISPR-associated transpose (e.g., Type IF-3 or Type V-K) e.g., as described in Tenjo-Castano F et al., Mol Cell. 2024 Jun 20;84(12):2353-2367.e5, PMID: 38834066. In some embodiments, the genome editor is an RNA-guided nuclease (RGN) or equivalents, derivatives, or functional variants thereof.
[0077] An “RNA-guided nuclease (RGN)” includes a nuclease that is capable of binding to a gRNA (or alternatively a DNA / RNA chimera guide RNA) and that is directed to a target polynucleotide by the gRNA. In some embodiments, an RGN comprises a CRISPR-associated protein (Cas protein or CRISPR nuclease) or a variant thereof. In some embodiments, a Cas protein is Cas3, Cas8a, Cas5, Cas8b, Cas8c, CaslOd, Csel, Cse2, Csyl, Csy2, Csy3, GSU0054, CaslO, Csm2, Cmr5, CaslO, Csxll, CsxlO, Csfl, Cas9, Csn2, Cas4, Casl2, Casl2a (Cpfl), Casl2b (C2cl), Casl2c (C2c3), Casl2d (CasY), Casl2e (CasX), Casl2f (Casl4, C2cl0), Casl2g, Casl2h, Casl2i, Casl2k (C2c5), C2c4, C2c8, and C2c9 or equivalents, derivatives, or functional variants thereof. In some embodiments, a Cas protein is Cas9. In some embodiments, a Cas protein is a naturally occurring or engineered variant of Cas9. In some embodiments, the RGN is catalytically inactive (e.g., a dCas9). In some embodiments, any of the herein disclosed RNA-guided nucleases may be configured as a nickase, i.e., wherein the RNA-guided nuclease is engineered to cut only one strand of a DNA target sequence. In other embodiments, any of the herein disclosed RNA-guided nucleases may be configured as a nuclease-inactive or “dead” RNA-guided nuclease wherein the enzyme retains its ability to bind to a guide RNA and be guided to and to bind to a target DNA, but lacks or is engineered to lack any nuclease activity, i.e., wherein a “dead” RNA-guided nuclease binds to but does not cleave either strand of a target DNA. One of ordinary skill in the art will appreciate that nickase and / or dead variants of known RNA-guided nucleases are available and / or can be prepared using known methods.A “guide RNA” includes an RNA polynucleotide that is capable of directing an RNA-guided nuclease (RGN) to a target polynucleotide (e.g., a portion of genomic locus). A gRNA typically comprises to short RNA components: (1) a CRISPR RNA (crRNA) that contains a sequence (a homology region (also called a spacer)) of about 20 nucleotides (though it can be longer or shorter) that is complementary (e.g., wholly complementary) to a target DNA sequence, and (2) a trans-activating crRNA (tracrRNA) that binds to the crRNA and to a CRISPR-associated (Cas) protein, for example, Cas9. In practice, these two components are often combined into a single, chimeric RNA molecule referred to as a single-guide RNA (sgRNA). In some embodiments, a gRNA is a sgRNA. A gRNA directs the Cas protein to the specific DNA or RNA sequence, where the Cas protein can induce a double strand break (DSB). Herein, the term “gRNA” includes a two-part gRNA as well as a sgRNA unless stated otherwise. Typically, homology regions (e.g., within a crRNA) are designed to be complementary to a specific polynucleotide sequence (also referred to as an on-target polynucleotide sequence) and not complementary to other polynucleotide sequences (also referred to as an off-target polynucleotide sequence).
[0078] In some embodiments, a gRNA is operably linked to a promoter. In some embodiments, a promoter is a constitutive promoter (e.g., a SV40, CMV, UBC, U6, EF1A, PGK or CAGG promoter). In some embodiments, a promoter is an inducible promoter (e.g., a TET promoter). In some embodiments, a promoter is a U6 promoter. In some embodiments, a promoter is a tRNA promoter. In some embodiments, the tRNA promoter is a eukaryotic tRNA promoter. In some embodiments, the tRNA promoter is a prokaryotic tRNA promoter. In some embodiments, the tRNA promoter is human tRNA promoter. In some embodiments, the tRNA promoter is an Arabidopsis tRNA promoter. In some embodiments, the tRNA promoter is a glycine or alanine promoter. In some embodiments, a promoter is a human cysteine tRNA (hCtRNA) promoter. In other embodiments, a promoter is a cell-specific or tissue- specific protein, such as a promoter that may be active in the transcription of genes only in certain cell types (e.g., lymphocytes, stem cells, or cancer cells).
[0079] “Complementary” refers to the relationship between two polynucleotides (DNA or RNA) in which each nucleotide on one strand pairs (binds) specifically with a corresponding nucleotide on another strand. This complementary base pairing is driven by hydrogen bonds: A forms two hydrogen bonds with T (or U in RNA), and C forms three hydrogen bonds with G. In some embodiments, a gRNA is 100% complementary to a target gene sequence. That is, each nucleotide of the gRNA is paired (bound) to the target gene sequence. In other embodiments, agRNA is less than 100% complementary to target gene sequence. For example, there can be one or more mismatches between a gRNA and a target gene sequence.
[0080] A “target polynucleotide” includes a polynucleotide that has been selected for cleavage using a genome editor (e.g., an RGN-guide RNA complex). In some embodiments, the target polynucleotide is a single stranded polynucleotide. In some embodiments, the target polynucleotide comprises at least 25 nucleotides (e.g., at least 50 nucleotides, at least 75 nucleotides, at least 100 nucleotides, at least 150 nucleotides, at least 200 nucleotides, at least 250 nucleotides, at least 500 nucleotides, or at least 1000 nucleotides). In some embodiments, the target polynucleotide comprises a target portion (e.g., 10-30 nucleotides, 15-30 nucleotides, 15-35 nucleotides, 10-40 nucleotides; or e.g., at least 10 nucleotides, at least 15 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides; or, e.g., up to 10 nucleotides, up to 15 nucleotides, up to 20 nucleotides, up to 30 nucleotides, up to 40 nucleotides, up to 50 nucleotides) that are complementary to the homology region of a gRNA of the genome editor. In some embodiments, the target polynucleotide is a double stranded polynucleotide. In some embodiments, the target portion of a double stranded target polynucleotide is adjacent to a protospacer adjacent motif (PAM) sequence. The PAM sequence is typically located on the reverse complement strand of the target sequence. PAM sequences are commonly required by CRISPR-based RGNs for cleaving a double stranded target polynucleotide, but RGNs have been shown to be cleave single stranded polynucleotides. PAM sequences for different CRISPR RGNs are known in the art. For example, a PAM sequence for Cas9 is 5’-NGG-3’. In some embodiments, the target polynucleotide is a portion of a gene. In some embodiments, the gene is a mammalian gene. In some embodiments, the mammalian gene is a human gene, a mouse gene, a non-human primate gene, a rat gene, or a ferret gene. In some embodiments, the gene is a plant gene. In some embodiments, the gene is an insect gene. In some embodiments, the gene is a microorganism gene (e.g., a parasitic gene, a viral gene, or a bacterial gene).
[0081] Obtaining a Library
[0082] In some embodiments, a method for determining an activity and / or specificity of a genome editor for a target polynucleotide comprises obtaining a library comprising circular DNAs, each circular DNA comprising a randomized target polynucleotide.
[0083] A “library” includes polynucleotides (e.g., circular DNAs). In some embodiments, the library comprises circular DNAs, each comprising a randomized target polynucleotide. In some embodiments, the library comprises additional circular DNAs, comprising a targetpolynucleotide, and not comprising a randomized target polynucleotide. In some embodiments, the library further comprises one or more circular polynucleotides (e.g., RNA or DNA) that do not comprise a randomized target polynucleotide and / or a target polynucleotide. In some embodiments, the library further comprises one or more linear polynucleotides (e.g., RNA or DNA). In some embodiments, the library comprises at least 10 (e.g., at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, at least IO10, at least 1011, at least 1012, at least 1013, or at least 1014) circular DNAs that each comprise a randomized target polynucleotide. In some embodiments, the library comprises at least 10 (e.g., at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least IO10, at least 1011, at least 1012, at least 1013, or at least 1014) circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises less than 108(e.g., less than 107, less than 106, less than 105, or less than 104) circular DNAs that each comprise a different randomized target polynucleotide.
[0084] In some embodiments, the library comprises 100-10,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 100-5,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 100-3,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 100-2,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 100-1,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 100-500,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 100-100,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 100-50,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 100-10,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 100-1000 circular DNAs that each comprise a different randomized target polynucleotide.
[0085] In some embodiments, the library comprises 1000-10,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 1000-5,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 1000-3,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, thelibrary comprises 1000-2,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 1000-1,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 1000-500,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 1000-100,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 1000-50,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 1000-10,000 circular DNAs that each comprise a different randomized target polynucleotide.
[0086] In some embodiments, the library comprises 10,000-10,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 10,000-5,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 10,000-3,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 10,000-2,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 10,000-1,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 10,000-500,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 10,000-100,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 10,000-50,000 circular DNAs that each comprise a different randomized target polynucleotide.
[0087] In some embodiments, the library comprises 100,000-10,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 100,000-5,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 100,000-3,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 100,000-2,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 100,000-1,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 100,000-500,000 circular DNAs that each comprise a different randomized target polynucleotide.
[0088] In some embodiments, the library comprises 1,000,000-10,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the librarycomprises 1,000,000-5,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 1,000,000-3,000,000 circular DNAs that each comprise a different randomized target polynucleotide. In some embodiments, the library comprises 1,000,000-2,000,000 circular DNAs that each comprise a different randomized target polynucleotide.
[0089] A “circular DNA” includes a DNA that forms a closed loop and has no ends. In some embodiments, a circular DNA comprises a target polynucleotide. In some embodiments, a circular DNA is single stranded. In some embodiments, a circular DNA is double stranded. In some embodiments, a circular DNA comprises a randomized target polynucleotide. In some embodiments, a circular DNA comprises one randomized target polynucleotide. In some embodiments, a circular DNA comprises one target polynucleotide. In some embodiments, a circular DNA comprises one randomized target polynucleotide and / or one target polynucleotide. In some embodiments, a circular DNA comprises a unique molecular identifier (UMI). In some embodiments, a circular DNA comprises at least 25 nucleotides (e.g., at least 50 nucleotides, at least 75 nucleotides, at least 100 nucleotides, at least 150 nucleotides, at least 200 nucleotides, at least 250 nucleotides, at least 500 nucleotides, or at least 1000 nucleotides). In some embodiments, the circular DNA comprises about 200 to about 1000 nucleotides. In some embodiments, the circular DNA comprises about 300 to about 600 nucleotides.
[0090] “About” as used herein refers to at most + / - 5% of a given value (e.g., at most 5%, at most 3%, at most 2%, or at most 1%).
[0091] In some embodiments, a circular DNA comprises a randomized target polynucleotide. A “randomized target polynucleotide” includes a polynucleotide that is a random variant of the target polynucleotide (e.g., an off-target mismatch of the target polynucleotide). In some embodiments, the randomized target polynucleotide comprises a randomized target portion of the target polynucleotide and / or a randomized PAM sequence that is adjacent to the target portion in the target polynucleotide. In some embodiments, the randomized target portion is not completely randomized. In some embodiments, the randomized target polynucleotide comprises 15-30 nucleotides. In some embodiments, the randomized target polynucleotide comprises at least 10 nucleotides (e.g., at least 25 nucleotides, at least 50 nucleotides, at least 75 nucleotides, at least 100 nucleotides, at least 150 nucleotides, at least 200 nucleotides, at least 250 nucleotides, at least 500 nucleotides, or at least 1000 nucleotides). In some embodiments, the randomized target polynucleotide comprises at least 1 nucleotide substitution (e.g., at least 1 nucleotide substitution, at least 2 nucleotide substitutions, at least 3 nucleotide substitutions, at least 4 nucleotide substitutions, at least 5 nucleotide substitutions, at least 6 nucleotidesubstitutions, at least 7 nucleotide substitutions, at least 8 nucleotide substitutions, at least 9 nucleotide substitutions, at least 10 nucleotide substitutions, at least 11 nucleotide substitutions, at least 12 nucleotide substitutions, at least 13 nucleotide substitutions, at least 14 nucleotide substitutions, at least 15 nucleotide substitutions, at least 16 nucleotide substitutions, at least 17 nucleotide substitutions, at least 18 nucleotide substitutions, at least 19 nucleotide substitutions, or at least 20 nucleotide substitutions) relative to the target portion of the target polynucleotide (e.g., the target gene). In some embodiments, the random variant of the target polynucleotide comprises no more than 3 nucleotide substitutions (e.g., no more than 3 nucleotide substitutions, no more than 4 nucleotide substitutions, no more than 5 nucleotide substitutions, no more than 6 nucleotide substitutions, no more than 7 nucleotide substitutions no more than 8 nucleotide substitutions no more than 9 nucleotide substitutions or no more than 10 nucleotide substitutions) relative to the target portion of the target polynucleotide. In some embodiments, the random variant of the target polynucleotide comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotide substitutions. In some embodiments, the random variant of the target polynucleotide comprises 1-6 nucleotide substitutions. In some embodiments, the random variant of the target polynucleotide comprises 1-10 nucleotide substitutions.
[0092] In some embodiments, the randomized target polynucleotides comprise substitutions, insertions, deletions, or combinations thereof relative to a target polynucleotide. In some embodiments, such randomized target polynucleotides correspond to sequence variants observed across different individuals or genomes. In some embodiments, such diversity allows assessment of genome editor activity and or specificity across a range of sequence variants.
[0093] In some embodiments, the randomized target polynucleotide is a partially randomized target polynucleotide. A “partially randomized target polynucleotide” includes a target polynucleotide having one or more mutations (e.g., substitutions) relative to the target polynucleotide that are determined using biased randomization (e.g., as described in Grasas et al. Computers & Industrial Engineering 110 (2017): 216-228). In some embodiments, the partially randomized target polynucleotides comprise fewer polynucleotides that will not be cleaved by the RGN complex compared to completely randomized target polynucleotides. In some embodiments, the sequences of partially randomized target polynucleotides may be generated by, for each nucleotide of the target portion of the target polynucleotide, randomly substituting that nucleotide to another nucleotide some fraction of the time. For example, for each nucleotide of the target portion of the target polynucleotide, substituting the nucleotide about 30% of the time, and keeping the nucleotide unchanged about 70% of the time. In anotherexample, for each nucleotide of the target portion of the target polynucleotide, substituting the nucleotide 1 %-50% of the time, and keeping the nucleotide unchanged 50%-99% of the time. In another example, for each nucleotide of the target portion of the target polynucleotide, substituting the nucleotide 3%-25% of the time, and keeping the nucleotide unchanged 75%-97% of the time. In another example, for each nucleotide of the target portion of the target polynucleotide, substituting the nucleotide 5%- 15% of the time, and keeping the nucleotide unchanged 85%-95% of the time. In another example, for each nucleotide of the target portion of the target polynucleotide, substituting the nucleotide 6%-l 1% of the time, and keeping the nucleotide unchanged 89%-94% of the time. In another example, for each nucleotide of the target portion of the target polynucleotide, substituting the nucleotide 7%-10% of the time, and keeping the nucleotide unchanged 90%-93% of the time. In another example, for each nucleotide of the target portion of the target polynucleotide, substituting the nucleotide 21%-30% of the time, and keeping the nucleotide unchanged 70%-79% of the time. In another example, for each nucleotide of the target portion of the target polynucleotide, substituting the nucleotide 18%-32% of the time, and keeping the nucleotide unchanged 68%-78% of the time. In another example, for each nucleotide of the target portion of the target polynucleotide, substituting the nucleotide 5%-40% of the time, and keeping the nucleotide unchanged 60%-95% of the time. In another example, for each nucleotide of the target portion of the target polynucleotide, substituting the nucleotide 20%-40% of the time, and keeping the nucleotide unchanged 60%-80% of the time. In another example, for each nucleotide of the target portion of the target polynucleotide, substituting the nucleotide 25%-35% of the time, and keeping the nucleotide unchanged 65%-75% of the time.
[0094] In some embodiments, the randomized target polynucleotide comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23 nucleotide substitutions relative to the target portion of the target polynucleotide. In some embodiments, the randomized target polynucleotide comprises 1-10 nucleotide substitutions relative to the target portion of the target polynucleotide. In some embodiments, the randomized target polynucleotide comprises 1-5 nucleotide substitutions relative to the target portion of the target polynucleotide. In some embodiments, circular DNAs of the library each comprise a randomized target polynucleotide, and the mean, median, and / or mode number of substitution mutations in the target portion of the randomized target polynucleotide of the circular DNAs is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23 nucleotide substitutions relative to the target portion of the target polynucleotide. In some embodiments, circular DNAs of the library each comprise a randomized target polynucleotide, and the mean, median, and / or mode number of substitutionmutations in the target portion of the randomized target polynucleotide of the circular DNAs is 1-10 nucleotide substitutions relative to the target portion of the target polynucleotide. In some embodiments, circular DNAs of the library each comprise a randomized target polynucleotide, and the mean, median, and / or mode number of substitution mutations in the target portion of the randomized target polynucleotide of the circular DNAs is 1-5 nucleotide substitutions relative to the target portion of the target polynucleotide.
[0095] In some embodiments, a method comprises synthesizing the randomized target polynucleotide. The randomized target polynucleotide may be synthesized in any suitable way. In some embodiments, the randomized target polynucleotide is synthesized using polymerase chain reaction with a degenerate primer. For example, a degenerate primer may comprise a first portion that is random or partially random (in reference to a target polynucleotide) and a second portion that is complementary to a template polynucleotide. In some embodiments, a method comprises synthesizing the sequences of the partially randomized target polynucleotides comprising synthesizing using partially degenerate oligonucleotide synthesis.
[0096] In some embodiments, obtaining the library comprises further comprises generating the library, the generating comprising: (a) amplifying a portion of a template DNA (e.g., a template plasmid) using: a plurality of first primers (e.g., degenerate primers) that (i) each comprises a polynucleotide that is complementary to a first portion of the template DNA, and (ii) a randomized target polynucleotide; and (b) a second primer that is complementary to a second portion of the template DNA, to produce linear precursor polynucleotides comprising a portion of the template DNA and different partially randomized target polynucleotides. In some embodiments, a method further comprises (b) circularizing the linear polynucleotides to produce the circular DNAs.
[0097] In some embodiments, generating the library comprises: (a) amplifying a portion of a template DNA (e.g., a template plasmid) using: a plurality of first primers that each comprise (i) a polynucleotide that is complementary to a first portion of the template DNA, (ii) a randomized target polynucleotide, (iii) a unique molecular index, and (iv) a universal primer sequence; and a second primer that is complementary to a second portion of the template DNA, to produce linear precursor polynucleotides comprising a portion of the template DNA and different partially randomized target polynucleotides.
[0098] A “template DNA” may be any suitable template DNA including a template plasmid. The portion of the template DNA typically does not comprise a sequence that can be cleaved by the RNP-guide RNA complex. For example, the portion of the template DNA typically does not comprise a sequence of the target portion of the target polynucleotide adjacent to a PAM site.Additionally, the portion of the template DNA typically does not comprise a randomized target sequence adjacent to a PAM sequence. In some embodiments, amplifying the portion of the template DNA using the first primer and the second primer produces linear DNAs that is about 25 to about 2000 nucleotides in length.
[0099] A “unique molecular identifier” or “UMI” includes short sequences that are added to DNA fragments (e.g., circular DNAs) to identify the input DNA molecule in next generation sequencing. In some embodiments, UMIs are random sequences. In some embodiments, UMIs are 6-30 nucleotides in length. In some embodiments, UMIs are 10-25 nucleotides in length. In some embodiments, UMIs are 15-25 nucleotides in length. In some embodiments, UMIs are 5-15 nucleotides in length. Typically, UMIs are not adjacent to a PAM sequence.
[0100] A ’’universal primer sequence” includes a sequence used in a downstream molecular biology process (e.g., NGS or PCR). For example, a universal primer sequence may be a sequence added to or included in a polynucleotide (e.g., a target polynucleotide or randomized target polynucleotide) during a first polymerase chain reaction that can be used as a primer binding site during a second polymerase chain reaction. In another example, a universal primer sequence is compatible with preparation for next generation sequencing (e.g., a sequence that is complementary to a primer used in preparing a polynucleotide for next generation sequencing).
[0101] In some embodiments, generating the library comprises introducing a restriction enzyme site into the linear precursor polynucleotides. The restriction enzyme site may be any suitable restriction enzyme site. In some embodiments, the restriction enzyme site introduced is not found in another location in linear precursor polynucleotide. In some embodiments, the restriction enzyme site is a site for a type I, type II, type III, or type IV restriction enzyme. In some embodiments, the restriction enzyme site is a site for EcoRi, BamHl, Hindlll, Notl, Sau3AI, PvuII, Smal, TaqI, Kpnl, Sphl, or Xbal. In some embodiments, the restriction enzyme site is introduced via a PCR reaction. For example, the restriction enzyme site may be included the first primer or the second primer used to produce linear precursor polynucleotides.
[0102] In some embodiments, generating the library comprises introducing an Uracil DNA glycosylase (UDG) and a DNA glycosylase-lyase Endonuclease VIII (USER) enzyme substrate, uracil onto one or both ends of the linear precursor polynucleotides. In some embodiments, the USER enzyme substrate is introduced via PCR.
[0103] In some embodiments, a method comprising bottlenecking circular DNAs of the library and / or the linear DNAs that are precursor’s to the circular DNAs. In some embodiments, prior to circularizing the linear precursor DNAs, a method comprising bottlenecking the linear DNAs. In some embodiments, a method comprising obtaining a library (e.g., a library described herein)and bottlenecking the library prior to cleaving randomized target polynucleotides of circular DNAs of the library.
[0104] “Bottlenecking” polynucleotides (e.g., linear DNAs or circular DNAs) includes reducing the number of different polynucleotides in the library (e.g., by random sub- sampling). For example, the linear precursor DNAs may comprise about 1013different randomized polynucleotides before bottlenecking and about 107different randomized polynucleotides after bottlenecking. In some embodiments, bottlenecking comprises reduces the number of different randomized target polynucleotides (e.g., in the circular DNAs of the library of the linear DNAs) by at least 2-fold (e.g., at least 5-fold, at least 10-fold, at least 20-fold, at least 50-fold, at least 1000-fold, at least 10,000-fold, at least 100,000-fold, or at least 1,000,000-fold). In some embodiments, bottlenecking comprising bottlenecking the linear precursor DNAs or the circular DNAs to collectively comprise no more than 10 million (e.g., no more than 5 million, no more than 3 million, no more than 2 million, or no more than 1 million, no more than 500,000, no more than 250,000, no more than 100,000, or no more than 50,000) different randomized target polynucleotides. Bottlenecking may be achieved using any suitable means. In some embodiments, bottlenecking comprises obtaining a solution comprising the linear precursor DNAs or the circular polynucleotides, and extracting a fraction of the solution with a molar amount of polynucleotides that corresponds to the number of different polynucleotides desired for the library. In some embodiments, bottlenecking comprises transfecting or transforming the linear precursor DNAs or circular polynucleotides into cells and then extracting the linear precursor DNAs or circular polynucleotides from the subset of cells.
[0105] Bottlenecking may be advantageous when the total number of different polynucleotides in a library (e.g., the number of different randomized target polynucleotides) is so great that detecting cleavage would be unduly expensive or infeasible given current sequencing technology and / or when activity and / or specificity can be determined by measuring cleavage of a subset of the randomized target polynucleotides. Quantitative measurements of activity in a defined set of mismatched targets may be more informative than ensemble measurements of a larger set, where activity cannot be quantified for any individual library member.
[0106] In some embodiments, generating the library comprises circularizing the linear precursor DNAs. Circularizing the linear precursor DNAs may be performed using any suitable method. In some embodiments, circularizing the linear precursor DNAs comprises processing the USER enzyme sites (Uracil bases) with a USER enzyme to produce sticky ends, processing the ends to be compatible for ligation (e.g. T4 PNK), and then ligating the sticky ends using a ligase (e.g.,T4 ligase). In some embodiments, a method comprises degrading residual linear precursor polynucleotides that are not circularized (e.g., using exonucleases).
[0107] In some embodiments, a method comprises making a copy of the library. In some embodiments, a method comprises making a copy of library prior to circularizing the library (e.g., a copy of a library comprising linearized DNAs having restriction enzyme sites and partially randomized sequences). In some embodiments, a method comprises making a copy of library after circularization of the library (e.g., a copy of a library comprising circular DNAs having restriction enzyme sites and partially randomized sequences). In some embodiments, a method comprises making at least 1 copy of library. In some embodiments, a method comprises making 1, 2, 3, 4 ,56, 7, 8, 9, or 10 copies of library. In some embodiments, making a copy of library comprises making a copy of the circular DNAs of the library. Making a copy of the library may be performed using any suitable means. For example, a copy of the library may be made using a polymerase chain reaction that amplifies the circular DNAs, and then separating the DNA from the PCR reaction into two or more different containers. In some embodiments, a copy of the library may be made by separating the library (e.g., a liquid phase library) into two or more different containers).
[0108] Cleaving Circularized DNAs
[0109] In some embodiments, a method for determining an activity and / or specificity of a genome editor for a target polynucleotide comprises cleaving randomized target polynucleotides of circular DNAs of the library (e.g., as described herein) using a genome editor (e.g., an RGN, a base editor, or another editor described herein), thereby producing linearized DNAs. In some embodiments, a method comprising digesting circular DNAs of the copy of the library using a restriction enzyme to produce control linearized DNAs. The control linearized DNAs may be used positive control for cleavage of the circularized DNA.
[0110] In some embodiments, a method for determining an activity and / or specificity of a genome editor for a target polynucleotide comprises cleaving randomized target polynucleotides of circular DNAs of the library (e.g., as described herein) using an RNA-guided nuclease (e.g., as described herein) and the gRNA (e.g., as described herein), thereby producing linearized DNAs. In some embodiments, cleaving randomized target polynucleotides of circular DNAs of the library comprising contacting the randomized target polynucleotides of circular DNAs of the library with a genome editor (e.g., an RNA-guided nuclease and a gRNA). In some embodiments, contacting comprises contacting in condition suitable for cleavage by the genome editor (e.g., physiological conditions). In some embodiments, cleaving randomized target polynucleotides of circular DNAs of the library comprising cleaving single stranded circularDNAs. In some embodiments, cleaving randomized target polynucleotides of circular DNAs of the library comprises cleaving doubled stranded circular DNAs. In some embodiments, cleaving randomized target polynucleotides comprises making a double strand break in the randomized target polynucleotide. In some embodiments, cleaving randomized target polynucleotides comprises making only one double strand break in the randomized target polynucleotide thereby producing one linear DNA for each cleaved circular DNA. In some embodiments, cleaving randomized target polynucleotides comprises cleaving using a base editor (e.g., an ABE or a CBE) (e.g., see FIG. 20A). ABEs nick the target DNA strand and deaminate adenine bases to inosine (within the base editing window) on the non-target strand at on- and off-target sites. Then Endonuclease V may be used on nicked and inosine-containing genomic DNA which cleaves DNA adjacent to inosines to produce linear DNA with 5’ staggered ends. A-to-I deamination as inosine bases are converted to guanine after PCR amplification. For CBEs nick the target DNA strand and deaminate cytosine bases to Uracil (within the base editing window) on the non-target strand at on- and off-target sites. Then USER enzyme may be used on nicked and Uracil containing genomic DNA which excises Uracil the base which, in combination with the nicked target strand, creates a double stranded break. In some embodiments, instead of relying on the base editors nicking of the target DNA strand to create a full double-stranded break, enzymes such as mismatch repair endonucleases are used to process DNA heteroduplexes containing deaminated bases into DNA double- stranded breaks.
[0111] In some embodiments, a method for determining an activity and / or specificity of a genome editor for a target polynucleotide comprises cleaving randomized target polynucleotides of circular DNAs of the library (e.g., as described herein), by contacting the circular DNAs with a base editor and cleaving the circularized DNAs with a nuclease, thereby producing linearized DNAs. In some embodiments, the method comprises the procedure shown in FIG. 20A. In some embodiments, the base editor comprises a deaminase. In some embodiments, the deaminase is an adenine deaminase (e.g., evolved or engineered TadA variants). In some embodiments, the deaminase is a cytidine deaminase. In some embodiments, the deaminase is engineered (e.g., via substitutions or truncations). In some embodiments, the base editor is BE1, BE2, HF2-BE2, BE3, HF-BE3, YE1-BE3, EE-BE3, YEE-BE3, VQR-BE3, EQR-BE3, VRER-BE3, SaKKH-BE3, FNLS-BE3, RA-BE3, A3A-BE3, eA3A-HFl-BE3-2xUGI, eA3A-Hypa-BE3-2xUGI, hA3A-BE3, hA3A-BE3-Y130F, hA3A-BE3-D131Y, hA3A-BE3-Y132D, SaCas9-BE3, xCas9-BE3, ScCas9-BE3, SniperCas9-BE3, iSpyMac-BE3, Target-AID, Target-AID-NG, BE-PLUS, BE4, SaCas9-BE4, SaCas9-BE4-Gam, BE4-Gam, BE4-Max, AncBE4-Max, evoBE4max, evoFERNY-BE4max, Casl2a-BE, TaC9-CBE, DdCBE, CRISPR-X, TAM, ABE7.10, xCas9-ABE7.10, VQR-ABE, Sa(KKH)-ABE, ABEmax, ABE7.10max, ABE8e, ABE8s, ABE9, GBE, CGBE1, CGBE, GBE2.0, AYBE, AXBE, gGBE, gTBE, SPACE, A&C-BEmax, ACBE, Target-ACEmax, or STEMEE In some embodiments, the base editor is any suitable base editor known in the art (e.g., a base editor disclosed in Gao et al., Mol Ther Nucleic Acids. 2025 Nov 10;36(4):102771, which is hereby incorporated by reference.) In some embodiments, the base editor catalyzes deamination of a base within the DNA to generate a modified base. In some embodiments the modified base is inosine (e.g., produced by adenine deamination). In some embodiments the modified base is uracil e.g., produced by cytidine deamination).
[0112] In some embodiments, the modified base is selectively recognized by a nuclease that cleaves (e.g., nicks) at or near the modified base. In some embodiments, the nuclease is endonuclease V (EndoV). In some embodiments, the nuclease is Uracil-Specific Excision Reagent (USER). In some embodiments, the nuclease is endonuclease Q (EndoQ). In some embodiments, the nuclease is any suitable nuclease capable of producing a nick at or near a modified base generated by the base editor. In some embodiments, the method comprises introducing a nick using a programmable nickase. In some embodiments, the programmable nickase is derived from a Cas protein. In some embodiments, the nickase is nCas9. In some embodiments, the nickase is any suitable programmable nickase known in the art. In some embodiments, the nuclease-generated nick and the nickase-generated nick are positioned on opposite strands of the DNA. In some embodiments, the nuclease- generated nick and the nickase-generated nick are sufficiently proximate and / or offset to generate a double-strand break. In some embodiments, the method comprises contacting nicked DNA with a DNA polymerase. In some embodiments, the DNA polymerase lacks 3’ to 5’ exonuclease activity. In some embodiments, the DNA polymerase is a Klenow fragment lacking 3’ to 5’ exonuclease activity. In some embodiments, the DNA polymerase extends the DNA from a nick, thereby producing a linearized DNA.
[0113] In some embodiments, determining the activity and / or specificity of a guide RNA for a target polynucleotide comprises sequencing the cleaved randomized target polynucleotides of the linearized DNA and / or sequencing cleaved target polynucleotide linearized DNAs. Unless stated otherwise, “cleaved polynucleotides” (e.g., cleaved target polynucleotides, cleaved randomized polynucleotides, cleaved partially randomized polynucleotides) refer to polynucleotides that have been cleaved using a genome editor. Unless states otherwise, “control polynucleotides” or “digested polynucleotides” refer to polynucleotide that have been digested using a control enzyme (e.g., a restriction enzyme). In some embodiments, determining the activity and / or specificity of a guide RNA for a target polynucleotide comprises sequencing thecontrol linearized DNAs. Sequencing may be performed by any suitable means. In some embodiments, sequencing is next- generation sequencing e.g., ILLUMINA, PACBIO, OXFORD NANOPORE, ELEMENT, or ION TORRENT next- generation sequencing. In some embodiments, determining the activity and / or specificity of a guide RNA for a target polynucleotide comprises determining an amount (e.g., fraction, percentage, and / or number) of target polynucleotides that are linearized DNAs (e.g., via sequencing). In some embodiments, determining the activity and / or specificity of a guide RNA for a target polynucleotide comprises determining an amount of cleaved randomized target polynucleotides of linearized DNA (e.g., via sequencing). In some embodiments, determining the activity and / or specificity of a guide RNA for a target polynucleotide comprises sequencing the cleaved randomized target polynucleotides of the linear DNA to identify an amount of randomized target polynucleotides cleaved. In some embodiments, determining the activity and / or specificity of a guide RNA for a target polynucleotide comprises sequencing the cleaved randomized target polynucleotides of the linear DNA to identify an amount of different randomized target polynucleotides cleaved (e.g., polynucleotides with different randomized target portions).
[0114] Determining Activity and / or Specificity
[0115] In some embodiments, a method comprises determining the activity of the genome editor for the target polynucleotide using the cleaved randomized target polynucleotides of the linearized DNAs. In some embodiments, a method comprises determining the activity of the genome editor for a randomized target polynucleotide using the cleaved randomized target polynucleotides of the linearized DNAs.
[0116] “Activity” of a genome editor for a target polynucleotides includes how frequently the genome editor, edits (e.g., cleaves) the target polynucleotide. In some embodiments, a genome editor that cleaves a target polynucleotide in 50 out of 100 target polynucleotides has a 50% activity. In some embodiments, determining the activity of the genome editor for a target polynucleotide comprises determining an activity of an amount of the genome editor per unit time. For example, determining that 1 nM of RGN-guide RNA complex may cleave 1000 target polynucleotides per second.
[0117] In some embodiments, determining the activity of the genome editor for the target polynucleotide comprises determining the efficiency with which the target polynucleotide is cleaved using the cleaved target polynucleotide linearized DNAs (e.g. produced using the genome editor) and the digested target polynucleotide in the control linearized DNAs (e.g., produced using the restriction enzyme). In some embodiments, determining the efficiency with which the target polynucleotide is cleaved comprises determining an amount of cleaved targetpolynucleotides of the linearized DNAs and an amount of digested target polynucleotide in the control linearized DNAs (e.g., using next generation sequencing). In some embodiments, determining the efficiency with which the target polynucleotide is cleaved comprises determining a ratio between an amount of cleaved target polynucleotides of the linearized DNAs and an amount of digested target polynucleotide in the control linearized DNAs (e.g., using next generation sequencing). In some embodiments, determining a relative amount comprises sequencing cleaved target polynucleotides of the linearized DNAs and digested target polynucleotide in the control linearized DNAs and determining a number of sequencing reads of cleaved target polynucleotides of the linearized DNAs and a number of sequencing reads of digested target polynucleotide in the control linearized DNAs. For example, a target polynucleotide appears in the control linearized DNAs (digested by the restriction enzyme) with 5 sequencing reads and in the cleaved target polynucleotides of the linearized DNA (cleaved with the genome editor) has 200 reads. Activity may be determined by taking a ratio of the target polynucleotide reads in the sequenced control linearized DNAs and in the cleaved target polynucleotides of the linearized DNA. In this example, the activity would be 200 / 5. This indicates that the genome editor has very high activity for the given randomized target polynucleotide. In another example, the given target polynucleotide appears with 200 reads in the control linearized DNAs and with 0 reads in the cleaved target polynucleotides. This indicates that the genome editor has low to no activity for the given target polynucleotide. In some embodiments, determining the activity of a guide RNA for a target polynucleotide comprises determining using a number of target polynucleotides cleaved (e.g., as identified using NGS). In some embodiments, determining the activity of a guide RNA for a target polynucleotide comprises determining using a number of target polynucleotides cleaved and a number of randomized target polynucleotides cleaved (e.g., as determined by NGS). In some embodiments, determining the activity of the guide RNA for a target polynucleotide comprises determining using a number of target polynucleotides cleaved compared to the total number of target polynucleotide the library (e.g., when a known amount of target polynucleotide in in the library).
[0118] In some embodiments, determining the activity of the genome editor for a randomized target polynucleotide comprises determining the efficiency with which the randomized target polynucleotide is cleaved using the cleaved randomized target polynucleotides of the linearized DNAs. In some embodiments, determining the efficiency with which the randomized target polynucleotide is cleaved comprises determining an amount of the cleaved randomized target polynucleotides of the linearized DNAs and an amount of the control linearized DNAscomprising the randomized target polynucleotide (e.g., using NGS). In some embodiments, determining the efficiency with which the randomized target polynucleotide is cleaved comprises determining a ratio between an amount of the cleaved randomized target polynucleotides of the linearized DNAs and an amount of the control linearized DNAs comprising the randomized target polynucleotide using next generation sequencing. In some embodiments, determining the activity of a genome editor for a randomized target polynucleotide comprises determining using an amount of the randomized target polynucleotide that are cleaved (e.g., as identified using NGS).
[0119] In some embodiments, a method comprises comparing the activity of the genome editor for two or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more) different randomized target sequences. In some embodiments, the method comprises comparing the activity of the genome editor for at least 2 (e.g., at least 2, at least 5, at least 10, at least 1000, at least 1000, at least 10,000, at least 100,000, at least 1,000,000, or at least 10,000,000 different randomized target sequences). In some embodiments, a method comprises comparing the activity of the genome editor for a first randomized target sequence and the activity of the genome editor for a second randomized target sequence. In some embodiments, the two or more target sequences differ by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 different nucleotides. In some embodiments, the two or more target sequences differ by 1-6 different nucleotides. In some embodiments, a method comprises comparing the activity of the genome editor for two or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more) different randomized target sequences. In some embodiments, the method comprises determining which of the two or more different randomized polynucleotides the editor has the highest activity for. In some embodiments, the method comprises determining a magnitude in the change (e.g., a fold change) in activity of the genome editor for the two or more different randomized polynucleotides. In some embodiments, a method comprises comparing the activity of the genome editor for two or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more) different randomized target sequences. In some embodiments, the method comprises determining which of the two or more different randomized polynucleotides the editor has similar activity for (e.g., within 5%).
[0120] “Specificity” of a genome editor for a target polynucleotide relates to how often the genome editor edits (e.g., cleaves) an off-target polynucleotide (e.g., a randomized target polynucleotide). In some embodiments, determining the specificity of a guide RNA for a target polynucleotide comprises determining the activity of the genome editor for a randomized target polynucleotide. In some embodiments, determining the specificity of a guide RNA for a target polynucleotide comprises determining the activity of the genome editor for a randomized targetpolynucleotide (e.g., an off-target polynucleotide) and determining the activity of the genome editor for the target polynucleotide (e.g., an on-target polynucleotide). In some embodiments, determining the specificity of a guide RNA for a target polynucleotide comprises determining a ratio between the activity of the genome editor for a randomized target polynucleotide and determining the activity of the genome editor for the target polynucleotide.
[0121] In some embodiments, a method comprises determining the specificity of a guide RNA for a target polynucleotide using the cleaved randomized target polynucleotides of the linear DNA. In some embodiments, determining the specificity of a guide RNA for a target polynucleotide comprises determining an amount of randomized target polynucleotides cleaved (e.g., using NGS). In some embodiments, determining the specificity of a guide RNA for a target polynucleotide comprises determining an amount of different randomized target polynucleotides cleaved (e.g., using NGS). In some embodiments, determining the specificity of a guide RNA for a target polynucleotide comprises determining an amount of randomized target polynucleotides cleaved (e.g., using NGS) relative to an amount of target polynucleotides cleaved. In some embodiments, determining the specificity of a guide RNA for a target polynucleotide comprises determining a ratio between the amount of randomized target polynucleotides cleaved and an amount of target polynucleotides cleaved.
[0122] In some embodiments, determining the specificity of the genome editor comprises determining the efficiency of the genome editor for a randomized target polynucleotide. In some embodiments, determining the specificity of the genome editor comprises determining the efficiency of the genome editor for at least 1 (e.g., at least 5, at least 10, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 104, at least 105, at least 106, or at least 107) randomized target polynucleotides. In some embodiments, determining the specificity of the genome editor comprises determining the efficiency of the genome editor for the target polynucleotide. In some embodiments, determining the specificity of the genome editor for the target polynucleotide comprises determining a ratio between the efficiency of the genome editor for the target polynucleotide and the efficiency of the genome editor for one or more randomized target polynucleotides.
[0123] In some embodiments, determining the specificity of the genome editor for the target polynucleotide comprises comparing the activity of the genome editor for the target polynucleotide and the activity of the genome editor for one or more of the randomized target polynucleotides. In some embodiments, determining the specificity of the genome editor for the target polynucleotide comprises determining a ratio between the activity of the genome editor for the target polynucleotide and the activity of the genome editor for one or more of therandomized target polynucleotides. In some embodiments, determining an activity of the genome editor for a target polynucleotide and / or a randomized target polynucleotide comprises determining the relative activity of the genome editor for the polynucleotide or / and the randomized target polynucleotide. In some embodiments, a method comprises determining relative activity of the genome editor. Percent relative activity of a genome editor may be determined using (a) an amount of off-target edits by the genome editor, and (b) an amount of on-target edits by the genome editor. An amount of off-target edits by the genome editor may be determined by determining a fold change in an amount of off-target edits in a genome editor treated condition and an amount of off-target edits in a control condition (also called fold change off-target). An amount of on-target edits by the genome editor may be determined by determining a fold change in an amount of on-target edits in a genome editor treated condition and an amount of off-target edits in a control condition (also called fold-change on-target). In some embodiments, determining relative activity of the genome editor comprises determining a ratio of an amount of off-target edits and an amount of on-target edits. In some embodiments, determining relative activity of the genome editor comprises determining fold change off-target and fold-change on-target.
[0124] In some embodiments, genome editor activity differs between randomized target polynucleotides that differ by one or more nucleotides. In some embodiments, such differences are reflected by differences in cleavage efficiency, activity and / or specificity for target polynucleotides comprising different nucleotides at a given position. In some embodiments, the effect of multiple mismatches within a randomized target polynucleotide on cleavage activity depends on the identity, position, and / or combination of the mismatches present.
[0125] In some aspects, a method comprises determining a cleavage site (e.g., the cleavage sequence and type of cleavage (e.g., blunt end or staggered)) of the genome editor (e.g., Cas9) in one or more randomized target polynucleotides of circular DNAs of the library or in a target polynucleotide of circular DNAs of the library. For example, in some embodiments, the method comprises determining a cleavage of the genome editor in at least 2 (e.g., at least 5, at least 10, at least 50, least 100, at least 500, at least 1000, at least 5000, at least 10000, at least 25000, at least 50000, or at least 100000) randomized target polynucleotides of circular DNAs of the library or in a target polynucleotide of circular DNAs of the library. In some embodiments, determining a cleavage site comprises sequencing (e.g., NGS) one or both ends of a cleaved linearized circular polynucleotide of a library. In some embodiments, the at least two polynucleotide are within a hamming distance of 10 (e.g., within a hamming distance of 9, within a hamming distance of 8, within a hamming distance of 7, within a hamming distance of65within a hamming distance of 55within a hamming distance of 43within a hamming distance of 35within a hamming distance of 2, within a hamming distance of 1).
[0126] In some embodiments, the method comprises comparing an activity of the genome editor between at least 2 (e.g., at least 5, at least 10, at least 50, least 100, at least 500, at least 1000, at least 5000, at least 10000, at least 25000, at least 50000, or at least 100000) different randomized target polynucleotides. In some embodiments, the at least two polynucleotide are within a hamming distance of 10 (e.g., within a hamming distance of 9, within a hamming distance of 8, within a hamming distance of 7, within a hamming distance of 65within a hamming distance of 53within a hamming distance of 43within a hamming distance of 3 within a hamming distance of 2, within a hamming distance of 1). In some embodiments, the method comprises determining an effect of a substitution in a randomized target polynucleotide by comparing the activity of the genome editor between at least 2 different randomized target polynucleotides (e.g., at least 2 different randomized target polynucleotides that differ by a hamming distance of 1). In some embodiments, the method comprises determining an effect of at least two substitutions in a randomized target polynucleotide by comparing the activity of the genome editor between at least 2 (e.g., at least 3, at least 4, at least 5, at least 6) different randomized target polynucleotides (e.g., at least 2 different randomized target polynucleotides that differ by a hamming distance of at least 2 (e.g., at least 3, at least 4, at least 5, at least 6)).
[0127] Additional Embodiments
[0128] 1. A method for determining an amount of off-target edits of a genome editor, comprising:
[0129] (a) obtaining a library comprising circular DNAs, each circular DNA comprising a randomized target polynucleotide or a target polynucleotide;
[0130] (b) cleaving randomized target polynucleotides and / or target polynucleotides of circular DNAs of the library using the genome editor, thereby producing linearized DNAs;
[0131] (c) identifying the cleaved randomized target polynucleotides of the linearized DNAs thereby determining an amount of off-target edits of the genome editor.
[0132] 2. The method of embodiment 1, further comprising determining an amount of on-target edits of the genome editor, the determining comprising identifying the cleaved target polynucleotides of the linearized DNAs.
[0133] 3. The method of embodiment 2, further comprising determining a relative activity of the genome editor for the target polynucleotide using the amount of off-target edits of the genome editor and the amount of on-target edits of the genome editor.4. The method of any previous embodiment, wherein the identifying comprises nextgeneration sequencing.
[0134] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.
[0135] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”
[0136] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.
[0137] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively.
[0138] Wherever a specific protein is identified by name (e.g., Cas9), the application also contemplates variants of a known sequence having an amino acid sequence having at least 80%, 85%, 90%, 95%, 99%, or up to 100% percent sequence identity with a known amino acid sequence of any such specific protein.
[0139] EXAMPLES
[0140] Example 1. Population-scale cellular GUIDE-seq-2 and biochemical CHANCE-seq profiles reveal human genetic variation frequently affects Cas9 off-target activity.
[0141] Main
[0142] An ongoing challenge in genome editing is the possibility of detrimental off-target edits in the genome. Given that individual genomes differ from the human reference genome at approximately 4-5 million positions, genetic variations such as single-nucleotide variants (SNVs), insertions, and deletions can overlap with Cas9 target-complementary or protospacer adjacent motif (PAM) sites. Such overlap can alter activity at existing off-target sites or create new ones. Thus, for exa-cel and similar therapies, a critical question is how genetic variation within patient populations may influence the risk of unwanted off-target editing. As genome editing therapies expand to broader patient groups, accounting for human genetic diversity will be essential to ensure their consistent safety and efficacy across populations. Despite intense interest, the frequency and extent to which genetic variation affect cellular editing remains largely unknown. This major gap is due to the lack of scalable, unbiased methods to experimentally measure genome- wide on- and off-target editing activity in cells from large, diverse populations, as well as tools capable of comprehensively assessing diverse off-target sites to acquire fundamental understanding. Although previous studies have shown that genetic variation can affect genome editing and off-target activity, they have primarily focused on in silico predictions of variant-containing off-target sites rather than unbiased, genome- wide experimental assessments.
[0143] Unbiased methods (that do not depend on a priori pre-selection of candidate sites) to define the frequency and location of unintended genome- wide off-target activity fall broadly into two categories: cellular methods that are more direct, and biochemical approaches that offerhigher sensitivity. GUIDE- seq is a broadly adopted cellular method to define genome- wide on-and off-target activity of genome editing nucleases such as Cas9 which was instrumental in defining the specificity of exa-cel. It is based on the principle of efficiently integrating a short double-stranded oligodeoxynucleotide (dsODN) tag into nuclease-induced double-stranded breaks (DSBs) followed by tag-specific amplification of surrounding genomic DNA. CHANGE-seq is a biochemical method for selective sequencing of editor-modified DNA that is more sensitive than GUIDE- seq, but requires further validation in cells. Improved versions of these assays with increased scalability or off-target diversity would enable improved understanding of the impact of genetic variation on genome editing.
[0144] Although early studies demonstrated that SNVs could influence Cas9 activity, they could not systematically assess how prevalent or impactful these variants are in a population- scale cellular context. Previously, CHANGE-seq was used to demonstrate that SNVs can significantly impact Cas9 biochemical in vitro cleavage in six donors. Computational studies leveraging data from the 1000 Genomes Project and Human Genetic Diversity Project estimated the theoretical prevalence of variant effects using off-target scoring algorithms. Cancellierri et al. developed CRISPRme, a variant- aw are off-target search tool, to identify off-target sites for the BCL11A enhancer target in exa-cel, and validated a variant- specific off-target in human hematopoietic stem cells that resulting in allele- specific indels and pericentric inversions8. A common limitation of in silico off-target studies is their reliance on genomic databases predominantly composed of individuals of European ancestry, potentially underestimating off-target risks in other populations.
[0145] To address this critical gap for better understanding the impact of individual genetic variation on Cas9 activity, GUIDE-seq-2 was developed, a new streamlined, high-throughput method for population-scale, genome- wide profiling of Cas9 off-target activity. GUIDE-seq-2 was applied to six gRNAs in lymphoblastoid cell lines from 95 individuals characterized by the 1000 Genomes Project and identify a high-confidence set of variants that affect off-target activity by multiplex targeted sequencing. These results reveal that genetic variants frequently overlap Cas9 off-target sites and alter cellular off-target activity. To further investigate these effects, a novel massively parallel biochemical assay, CHANCE-seq, was developed and enabled the identification of genetic features associated with off-target cleavage in vitro (FIG.
[0146] 1).
[0147] Results
[0148] Development of high-throughput method to evaluate cellular genome-wide activity of genome editorsTo enable analysis of Cas9 nuclease genome- wide off-target activity in a large number of well-characterized lymphoblastoid cell lines from genetically diverse individuals, it was reasoned that GUIDE-seq could be streamlined, as it would be challenging to scale the original method to many samples.
[0149] To substantially increase scalability, subtractive changes were searched and resulted in a streamlined and improved, high-throughput method called GUIDE-seq-2 which substantially reduces processing steps, time, and cost compared to the original method. GUIDE-seq-2 utilizes Tn5-transposon-based tagmentation and a new library design that eliminates the need for nested rounds of PCR (FIG. 2A), integrates GUIDE-seq improvements without changing the original GUIDE-seq tag sequence to maintain full compatibility, and simplifies both library preparation and sequencing. The redesigned PCR strategy simultaneously improves detection accuracy by enabling mispriming detection, and sequencing without custom primers and sequencer configuration that were previously required. By introducing tagmentation which simultaneously fragments genomic DNA and adds required sequencing adapters, multiple library preparation and purification steps were eliminated, including random genomic DNA shearing, end repair / A-tailing and adapter ligation, and reduced genomic DNA input requirements by 4-fold (FIGs. 7A-7B and Protocol 1). All experiments for GUIDE-seq-2 development were performed in primary human T cells.
[0150] The sensitivity of GUIDE-seq-2 was directly compared with GUIDE-seq at 8 Cas9 target sites across 5 therapeutically relevant loci in CD4+ / CD8+human primary T cells (CCR5, TRAC, PDCD1, AAVS1 and CTI 4) with variable numbers of known off-targets identified in a previous study. A strong correlation was found between GUIDE-seq and GUIDE-seq-2 read counts at these 8 sites (r = 0.8 to 1.0, r is Pearson’s correlation coefficient) (FIG. 2B). For each of the 8 targets, independent GUIDE-seq-2 library preparations were performed and it was found that replicate read counts were strongly correlated (r = 0.99 to 1.0) (FIG. 2C), demonstrating high technical reproducibility. The positions of on and off-target sites were located genomewide (Fig 2D).
[0151] To filter out mispriming artifacts, in the GUIDE-seq-2 design the original tag sequence was kept but the length of the primer annealing to the dsODN tag region was reduced, allowing for the observation of expected sequences of 5 to 12 bases of the dsODN region, and the proper identification of expected junction between dsODN and flanking genomic DNA in the sequencing data (FIG. 7B) An advantage of shortening the tag-specific primer (instead of increasing the dsODN tag length), is that tag integration rates upon nucleofection remain unchanged, in contrast to the decreases that can occur when the tag length is increased.To further streamline the GUIDE-seq-2 workflow, a bead-based normalization step was introduced, using limiting amount of beads for single-step PCR purification and DNA normalization, avoiding labor-intensive serial dilution step for library quantification.
[0152] Additionally, GUIDE-seq-2 was adapted and optimized for an automated liquid handling platform, demonstrating high technical reproducibility and comparability to manual GUIDE-seq-2, and expanding its potential for high-throughput applications (FIGs. 8A-8E). In summary, GUIDE-seq-2 method is comparable to original GUIDE-seq in sensitivity, is highly reproducible, automatable, requires less genomic DNA for library preparation, reduces processing steps and time to 3 hours, and has improved detection accuracy (Protocol 1).
[0153] CRISPR-Cas9 cellular off-target editing sites frequently overlap human genetic variants.
[0154] GUIDE-seq-2 enables efficient and precise assessment of Cas9 off-target activity, allowing investigation of the impact of human genetic variation at unprecedented scale.
[0155] Specifically, cellular on- and off-target sites associated with six gRNA targets and one control from 95 donors across four populations, in a total of 665 GUIDE-seq-2 libraries, were analyzed (ASW, African Ancestry in Southwest USA; CEU, Utah residents with Northern and Western European ancestry; CHB, Han Chinese in Beijing; and MXL, Mexican Ancestry in Los Angeles,). The six gRNA target sites (AAVS1 site 14, CCR5 site 8, CTLA4 site 9, CXCR4 site 8, LAG3 site 9 and TRAC site 1) were selected due to their intermediate levels of off-target activity as measured in a previous study . GUIDE-seq-2 dsODN integration was optimized for lymphoblastoid cell lines (LCLs) characterized by the Genome-in-a-bottle consortium (GIAB) and evaluated at two targets with GUIDE-seq-2 (FIGs. 9A-9E).
[0156] 1,115 unique on- and off-target sites were identified genome- wide ranging from 56 to 333 per target (FIG. 3A). These sites were distributed mostly in intergenic and intronic areas that comprise the majority of the human genome, and found also in promoter, 5’ and 3’ UTR, and exonic regions in expected proportions ( FIG. 3B).
[0157] Notably, cellular off-targets detected by GUIDE-seq-2 frequently overlapped with human genetic variants ( FIGs. 3C-3D, FIGs. 10A-10C), at frequencies much higher than previous estimates of 1.5-8.5%. Specifically, 16.6% of on- or off-target sites (185 of 1,115) overlapped with one or more genetic variants found in the cohort (n=212 overlapping variants), with individuals of ASW ancestry having the highest frequencies of off-targets overlapping genetic variation (13%), and those of CHB ancestry having the lowest (7.8%) (FIG. 3E). When considering the entire 1000 Genomes Project population of 3,202 individuals with high-coverage whole genome sequencing (n=622 overlapping variants), the overlap percentage increased to 41.5% (463 of 1,115 sites) (FIG. 3F). Finally, when considering all variants ingnomAD (Genome Aggregation Database, n=6,311 overlapping variants), currently the largest collection of human genetic variation data, it was found that nearly all (i.e., 99% or 1,104 of 1,115) of sites overlap with at least one annotated genetic variant (FIG. 3G).
[0158] Among these variants, SNPs are the most frequent type, followed by insertions then deletions (FIG. 3H). Among the SNPs, transition mutations outnumber trans version mutations by an approximately a 2:1 ratio, consistent with expectations for genome- wide studies (FIG. 31). Overlapping variants have a wide range of frequencies across populations. In the cohort, 60% variants are common variants (defined as minor allele frequency or MAF>1%) (FIG. 3 J). More rare variants were found as variant database size increased, with 79% rare variants in the 1000 Genomes Project and 98% in gnomAD.
[0159] Most GUIDE-seq-2 sites (n=161) overlapped with only one variant, although 24 off-target sites harbored two or three (FIG. 3K). It was observed that overlapping SNP variants were evenly distributed across the protospacer and PAM positions (FIG. 3L) with an average of 7.6 SNPs per position. Interestingly, some variants were highly population-specific, suggesting the importance of population-specific, variant-aware analyses. For example, SNP rs78326679, occurring at the second position in LAG3 site 9 off-target, is much more likely to occur in people of African American ancestry (FIG. 3m). It was estimated that the average variant-affecting off-target will be found in 15.6% of individuals, however, this number will depend on the variant allele frequencies (VAF) observed in each specific population (FIG. 3N).
[0160] Overall, the study of 95 individuals efficiently surveyed 97% of common variants and 34.1% of all variants in the 1000 Genomes Project (with MAF down to 0.03%) that overlapped the GUIDE-seq-2 cellular off-target sites discovered herein(FIG. 3o). Based on these results, it is estimated that on average, a sample of 55 individuals would be sufficient to survey 95% of common variants that overlap cellular off-target sites (FIG. 30).
[0161] Validation of GUIDE-seq-2 discovered off-target sites.
[0162] To validate GUIDE-seq-2 discovered off-target sites and identify a high-confidence set of variant effects on off-target editing activity, multiplexed targeted sequencing was conducted to measure editing frequencies at 283 on- and off-target sites associated with the 6 targets presented here. This panel was designed to cover as many variant-containing off-targets as possible while encompassing a broad range of editing activities. Among the 185 total variantcontaining off-targets, the multiplex targeted sequencing panel successfully covered 148 sites, corresponding to 167 variants. Then, the significance of variant effects on editing activity was assessed, as measured by indel and GUIDE-seq-2 tag integration frequency (FIG. 4A).Using GUIDE-seq-2, a broad range of off-target activities associated with 6 gRNAs targets were quantified that varied substantially between individuals from the ASW, CEU, CHB, and MXL populations (FIG. 4B). It was confirmed that GUIDE-seq-2 read counts were strongly correlated with indel and GUIDE-seq-2 tag integration frequencies as measured by multiplex targeted sequencing (FIGs. 11A-1 IB). It was estimated that GUIDE-seq-2 lower limits of detection in the studies were -0.2%, based on the lowest estimated indel frequencies of off-target sites that could be detected in 95% of samples (estimated by the product of GUIDE-seq-2 relative off-target activity and on-target editing percentage) (FIG. 4C).
[0163] Overlapping genetic variants frequently affect off-target editing in human cells.
[0164] To define a set of high-confidence variant effects on off-target activity, five variant effect measures were integrated. Specifically, the effect of overlapping genetic variation was modeled on normalized GUIDE-seq-2 read counts, indel frequencies, tag integration frequencies, and allele-specific editing. If Cas9 exhibits a preference for reference or alternative genotypes, this should be reflected as an imbalance in the frequency of edited alleles (FIG. 4D). A high-confidence variant effect was defined as one that was significant in two or more of these measures (see Methods).
[0165] A high proportion of variants (12.6%) measured by both GUIDE-seq and targeted sequencing significantly impacted off-target activity, of which 6.6% (n=ll) and 6% (n=10) of variants led to increased or decreased editing, respectively (FIG. 4E). Variant effects on indel percentage ranged from -25% to 32% and were corroborated by comparable changes in GUIDE-seq-2 read counts and tag -integration percentages (FIG. 4F, FIG. 12). Notably, variant effects on TRAC site 1 on-target activity (FIG. 4F, high-confidence variant #21) was only identified by targeted amplicon sequencing, as it was at the upper limits of quantification by GUIDE-seq-2.
[0166] Certain variant effects were predictable, such as those that revert a mismatch to the canonical S. pyogenes Cas9 NGG PAM sequence. For example, a T>G variant that altered an NGG PAM-containing off-target sequence (associated with a gRNA targeted to CXCR4) to NTG significantly reduced Cas9 editing activity, consistent with the importance of PAM recognition to license Cas9 target DNA cleavage (FIG. 4G).
[0167] Others were unexpected, such as a T>G variant at the first position of a Cas9 off-target sequence associated with a CCR5-targeting gRNA, which surprisingly led to a substantial 70-fold increase in Cas9 indels to -10% (FIG. 4G). Thes findings underscore the complexity of predicting Cas9 activity on variant-containing genomic sites effects and highlight the value of high-throughput profiling techniques.
[0168] Strong variant effects on off-target editing observed in human primary T cells.To determine whether these results extended to therapeutically relevant primary human cells, GUIDE-seq-2 was performed in primary human T cells from three healthy donors and evaluated Cas9 off-target activity for two sgRNAs target sites (CXCR4 site 8 and LAG3 site 9). By directly evaluating GUIDE-seq-2 read counts correlations between the three donors (FIG. 1 IB), strong evidence was found that the genome- wide off-target activity of Cas9 in these human primary T cells is affected by the presence of SNVs (FIG. 4H) with approximately 8 to 3,000-fold changes in GUIDE-seq-2 read counts.
[0169] Massively parallel biochemical assay to quantify activity at diverse mismatched off-targets.
[0170] The population-scale GUIDE-seq-2 analysis presented herein identified, to Applicant’s knowledge, the largest set of genetic variants known to affect Cas9 genome- wide off-target activity and revealed the first estimates of the frequency with which overlapping off-targets modulate Cas9 activity. However, it remained challenging to infer general characteristics of genetic variants likely to impact Cas9 activity based on these sites alone.
[0171] To better understand influential genetic context and variants, a high-throughput biochemical assay was designed that could measure Cas9 activity within highly diverse, partially randomized, off-target libraries called CHANGE-seq Randomized (CHANCE-seq) (FIG. 5A, FIG. 13A). CHANCE-seq builds on principles of CHANGE-seq, a method previously developed to selectively sequence DNA cleaved and linearized by Cas9 from populations of circularized genomic DNA molecules, but operates on partially randomized off-target libraries. It was reasoned that quantitative measurements of enrichment of millions of variations of target sequences between control and Cas9 treatment conditions could be more informative than ensemble measurements of larger libraries, where library members are unlikely to be shared between treatment conditions.
[0172] Using CHANCE-seq, five of the targets (AAVS1, CCR5, CTLA4, LAG3, and TRAC) that were profiled with population- scale GUIDE-seq-2 were surveyed. First, highly diverse libraries (IO10or more) where each position of a 23 bp Cas9 target and PAM sequence has been partially randomized by PCR, were generated, resulting in unique molecular index (UMI) barcoded libraries of off-targets with an average of 6 mismatched to the original on-target sequence (FIG. 5A). The presence of UMIs can be used to correct for sequencing errors and other artifacts that could confound analyses. Second, importantly, to generate circularized DNA libraries containing multiple copies of the same randomized target, a subsample of approximately 3-5 million sequences was bottlenecked and amplified. Third, these circularized DNA libraries were split and treated with either Cas9 that can recognize and cleave similar on-or off-target sites, or a control restriction enzyme that cuts a site adjacent to the randomizedtarget sequence, cleaving all molecules. Only cleaved DNA circles have free ends for subsequent adapter ligation and high-throughput sequencing. For each on- or off-target sequence, the ratio of normalized read counts between Cas9 and control libraries were calculated as a measure of Cas9 activity (FIG. 5A, FIG. 13A, Note 2 and Protocol 2).
[0173] To determine whether the CHANCE-seq experiments effectively enriched for Cas9-cleaved DNA molecules, the distribution of mismatches within the library when cleaved by Cas9 versus a control restriction enzyme that would cut universally, was measured. It was observed that Cas9-cleaved libraries were significantly enriched for sites with lower numbers of mismatches compared to control libraries (X2, p < 0.05), indicating that the assay was working as intended (FIG. 5B).
[0174] Relative off-target activity measured between CHANCE-seq technical replicates is strongly correlated (r = 0.89 to 0.98) (FIG. 5C, FIG. 13B). PAM preferences for all measured targets were as expected, with the highest activity (84.1-99.3%) at canonical NGG PAM, some activity at NAG (53.5-72.3%) and NGA (11.8-51.3%), and minimal activity at other non-canonical PAM sequences (FIG. 5D). Each level of mismatch from 1-6 is covered by 0.6 to 0.9 million sequences for each target, for a total of 4 million unique sequences after filtering for adequate representation in control libraries (FIG. 5E). Average off-target activity generally decreases with increased mismatch number, although a number of exceptions were observed where high mismatch sites retain high activity (FIG. 5F, FIG. 13C).
[0175] At lower total mismatch numbers (1 to 2), mismatches are often observed in off-target sites at PAM-distal and proximal regions, specifically positions 1 and 17-19 of the target sequence (where position 21 is the first PAM position), whereas high-activity sites with higher total mismatches (3-6), tend to aggregate mismatches around positions 13-15. At all mismatch levels, mismatches at position 16 are poorly tolerated (FIG. 5G, FIG. 13D). The longest block of contiguous mismatches were found in PAM-distal regions (starting at 5’ positions 1-2 as expected), but also in the middle (starting at positions 9-11) (FIG. 5H, FIG. 13E). These two regions where Cas9 is most prone to off-target site mismatches coincide with those where structural studies have found proteimDNA phosphate backbone contacts (between target strand positions 1-2, 8-9, 10-13) that stabilize off-target site unwinding and recognition.
[0176] Transition mutations such as G>A or T>C that result in wobble base pairings (rG:dT or rU:dG) are generally well tolerated and have the smallest effect on activity (FIG. 51), specifically at positions 4, 16 and 22 (FIG. 13F), whereas Hoogsteen and transversion mutations are most deleterious for Cas9 activity (FIG. 51) at positions 1, 22, and 23 for Hoogsteen mutations (FIG. 13F).Interesting patterns emerge upon stratifying the global activity data by mismatch number and relative activity. In addition to the PAM, positions 16 and 18 appear the least tolerant to harboring mismatches while maintaining activity. At position 16, there is a stark contrast between mismatches that result in high editing activity (transitions) and lower activity (transversions) (FIGs. 5J-5K, FIGs. 14A-14D). At higher off-target mismatch numbers, high-activity sites often accumulate more PAM-distal mismatches (FIG. 5K, FIGs. 15A-15D).
[0177] Features of high- and low- impact variants and contexts
[0178] CHANCE-seq enables quantitative assessment of variant effects across diverse contexts by comparing activity differences between off-targets. Specifically, all pairs of off-targets with up to 6 mismatches that differed by a single mismatched nucleotide were enumerated and single nucleotide variants (SNVs) were defined as those that increase mismatches to their on-target sequence (FIG. 6A). At this point, there were pairs of ‘off-target context’ and ‘off-target variant’ sequences. This approach enables to the identification of important mismatches which when reverted to wild-type would increase editing activity. To understand variant effects, high-activity context sequences were explored, grouped by number of mismatches and sorted based on activity changes (denoted as A), classifying variants with minimal activity changes as neutral or well-tolerated, and those with substantial activity changes as deleterious high-impact variants.
[0179] The majority of variants that increase the number of mismatches relative to the on-target site reduce Cas9 activity to some degree (FIG. 6B). Variants have a greater impact on activity (‘leverage’) when the total numbers of “off-target context” mismatches is low (e.g., 0-2) compared to those with higher mismatches (3-5). This reduction in activity is less pronounced at PAM-distal positions (1 and 2) (FIG. 6C).
[0180] Next, the genetic features associated with deleterious variants that substantially diminished highly active off-target effects were investigated. Specifically, for each number of mismatches, the top 1000 most active sites from each of the five targets in CHANCE-seq data were sorted by variant effect (i.e., A), and the top 100 deleterious variants (including their paired off-target context) were selected (FIG. 6D, ).
[0181] High-impact variants were enriched in the PAM-proximal (i.e., position 16 to 20) and PAM regions (FIG. 6E). As the number of mismatches increased, deleterious variants became more widely distributed across different positions; however, their frequencies in the PAM region declined. Position 1 consistently exhibited the lowest likelihood of harboring deleterious variants, regardless of the number of mismatches. Notably, most deleterious variants were trans versions. Distinct and almost complementary patterns in the off-target context also emerged. As mismatches increased within context sequences, they tended to cluster in the distal(position 1 to 3) and middle (position 8-15) regions of the spacer. Additionally, most mismatches in high-activity off-target context sequences were transition mutations (FIG. 6E).
[0182] Discussion
[0183] Presented herein is the first experimental, population-scale view of the impact of genetic variation on editing activity and suggests that nearly every cellular off-target site for a given CRISPR gRNA target may overlap with variation in some individuals, and that these overlapping variants frequently alter unintended Cas9 off-target activity. In a therapeutic context, of greatest concern would be the subset of variants that increase off-target editing and safety risks to patients (although it should be noted only some unintended off-target sites may have an adverse biological consequence). For biological research, genetic variation between the genomes of donors of cell lines or primary cells may confound genome editing outcomes. The estimate of frequent variant-induced effects on editing suggests that they should be carefully accounted for in the early stages of therapeutic lead target discovery or basic research, in a population-specific manner.
[0184] GUIDE-seq-2 can simplify the cellular off-target analysis of genome editing for many labs, and enables large-scale experimental design as well as evaluating large number of gRNA targets. This novel and streamlined GUIDE-seq-2 approach enabled the survey of 97% of common genetic variants found with the 1,000 Genomes Project resulting in 21 high-confidence variant effects. However, developing general understanding of how genetic variants influence off-target activity requires high diversity datasets with even representation of possible mismatches. This gap motivated the exploration of complementary alternatives.
[0185] Massively parallel biochemical profiling with CHANCE-seq has advantages over existing approaches. In contrast to biochemical methods like standard CHANGE-seq and Digenome-seq which characterize off-target activity in genomic DNA, CHANCE-seq interrogates highly diverse, partially randomized target libraries. These libraries are intrinsically better suited for understanding how different mismatches may affect CRISPR-Cas9 activity. Protocol 1: GUIDE-seq-2
[0186] Reagents
[0187] • IDTE pH 8.0 (IX TE Solution) (Integrated DNA Technologies, 11050204) • IDT custom oligonucleotides
[0188] • Synthetic chemically modified sgRNA (Synthego)
[0189] • Tn5 transposase
[0190] • Proteinase K (NEB, P8107S)• Ultrapure Nuclease Free water (Invitrogen 10977023)
[0191] • Magnum FLX magnetic rack (Alpaqua, A00400)
[0192] • dNTPs (NEB, cat.no. N0447L)
[0193] • Qubit dsDNA HS Assay Kit (Thermo Fisher Scientific, Q32854)
[0194] • Lib Quant Kit (Illumina / Uni) (Kapa Biosystems, KK4824)
[0195] • Illumina
[0196] • PhiX Control V3 KIT (Illumina, FC- 110-3001)
[0197] • Ultra pure EDTA 0.5 M (Invitrogen, 15575038)
[0198] • Ethanol (Sigma, cat.no. E7023)
[0199] • Sera-Mag Magnetic Beads; Carboxyl, Speedbeads; hydrophobic; 5 solids (Fisher / GE,. 9981123)
[0200] • Guanidine thiocyanate (Sigma, G9277)
[0201] • Sodium Chloride 5 M Sterile (Invitrogen, AM9760G)
[0202] • TRIS Buffer 1.0 M solution, pH 8.0 (Fisher, 50146868)
[0203] • Polyethylene Glycol 8000 (Fisher, 507516674)
[0204] • IM Magnesium chloride (Invitrogen, AM9530G)
[0205] • Tween-20 (Sigma, P7949-500ML)
[0206] • Sodium Hydroxide solution (Sigma, 72068- 100ML)
[0207] Reagent setup
[0208] SPRI-guanidine binding buffer. SPRLguanidine binding buffer is composed by 4 M guanidine thiocyanate, 40 mM TRIS, 17.6 mM EDTA, pH 8.0. To prepare 100 ml of SPRI-guanidine binding buffer, weight 47.26 g of guanidine thiocyanate, add water and keep it on a magnetic stirrer. Once guanidine thiocyanate is dissolved, add 4 ml of TRIS 1 M pH 8 and 3.52 ml of EDTA 0.5 M pH 8. Bring the pH to 8 with HC1.
[0209] Sera-Mag Magnetic Beads preparation. Add 1 ml of Sera-Mag Magnetic Beads (Fisher / GE) to a 1.5 ml Eppendorf tube. Place in a magnetic rack. Remove the liquid. Remove the tube from the rack. Add 1 ml of TE and homogenize. Place back in the magnetic rack and remove the liquid. Repeat this step for a total of two TE pH 8.0 washes. Then, add 1 ml of TE pH 8.0. Note: this beads preparation step is required for preparing SPRLRNA beads, SPRLDNA beads and SPRI-guanidine beads.
[0210] SPRI-guanidine beads preparation. Add 10 ml of 5 M NaCl to 9 g of PEG 8000 and then add SPRI-guanidine binding buffer (prepared as described above) up to 49 ml. Homogenizeduring 5 min. Add 1 ml of Sera-Mag Magnetic Beads in TE (prepared as described above) and homogenize. Keep at 4 °C. SPRI-guanidine beads can be stored at 4°C for up to 6 months.
[0211] SPRI-DNA beads preparation. Add 10 ml of 5 M NaCl, 500 pl of 1 M TRIS and 100 pl of 0.5 M EDTA to 9 g of PEG 8000. Complete the volume to 49 ml with ultra-pure water. Add 1 ml of Sera-Mag Magnetic Beads in TE (prepared as described above) and homogenize. Add 27.5 pl of Tween-20 and homogenize. Keep at 4°C. SPRI- beads can be stored at 4 °C for up to 6 months. SPRI-DNA beads can be replaced by AMPure XP beads.
[0212] Tn5
[0213] pTXBl-Tn5 (addgene # 60240) was expressed in Rosetta(DE3)pLysS (millipore sigma #71403). After cell lysis and clarification, Tn5 was purified with chitin beads on HiPrep Q Sepharose 16 / 10 column (Cytiva) and cleaved with 20 mM HEPES, pH 7.2; 200 mM NaCl, 1 mM EDTA, 10 % glycerol, 0.2 % Triton X-100, 50 mM DTT. Eluted Tn5 was further purified with size exclusion and dialyzed into 100 mM HEPES pH7.2, 200 mM NaCl, 0.2 mM EDTA, 2 mM TCEP, 0.2% TritonX-100, 20% glycerol. To make stocks, purified Tn5 was diluted to 4.5 mg / mL with 2x Tn5 buffer (below) and further mixed with Glycerol and 2x Tn5 buffer at 1 : 1.1 : 0.33 to make 1.85 mg / mL Tn5 final. Stocks were stored at -20C until transposome assembly.
[0214] 2X Tn5 dialysis buffer
[0215] To prepare 500 ml of 2X Tn5 dialysis buffer add as follows:
[0216] Component Volume (ml) Final Concentration HEPES-KOH (1 M) pH 50 100 mM
[0217] 7.2
[0218] NaCl (5 M) 20 200 mM
[0219] EDTA (0.5 M) 0.2 0.2 mM
[0220] DTT (I M) 1 2 mM
[0221] TritonX-100 1 0.2% (vol / vol) Glycerol 100 20% (vol / vol)
[0222] Nuclease-free water to 500 ml
[0223] Total 500
[0224] 2X Tn5 dialysis buffer can be stored at 4 °C for up to 6 months.
[0225] 5X TAPS-DMF buffer: 50 mM TAPS-NaOH pH 8.5, 25 mMMgCh, 50% (vol / vol) DMF.
[0226] Prepare fresh.
[0227] Component Volume (pl) Final Concentration TAPS-NaOH (0.5 M) 100 50 mM
[0228] MgCl2(I M) 25 25 mM
[0229] DMF 500 50% (vol / vol)
[0230] Nuclease-free water 375
[0231] Total 1000Primers and oligos
[0232] 15 oligos
[0233]
[0234] Primers for PCR
[0235]
[0236] i7 dsODN primers
[0237]
[0238]
[0239]
[0240] Protocol:
[0241] Tn5 oligos annealing
[0242] Resuspend oligos in IDTE pH 8, to 100 pM.
[0243] Prepare i5 oligos:
[0244]
[0245] • Each top i5 oligo is paired with the same bottom oligo (Tn5-A-bottom)
[0246] • On a thermocycler, set up the follow annealing program: 95 °C for 5 min, -1 °C / 30 seconds for 70 cycles, hold at 4 °C. After annealing, add 100 pl of IDTE pH 8.0 to bring theconcentration of the annealed oligonucleotides to 25 |aM. Keep the annealed oligonucleotides at -20 °C. The annealed adapters will be used for transposome assembly.
[0247] Transposome assembly
[0248] Prepare Tn5 complex with each annealed i5 oligo:
[0249]
[0250] • Incubate at RT for 1 hour.
[0251] • Store assembled Tn5 at -20 °C.
[0252] Tagmentation
[0253] Prepare a tagmentation reaction for each gDNA sample. 100 ng of gDNA is tagmented with Tn5 complexed with a unique i5 barcode.
[0254]
[0255] • Incubate the reaction at 55 °C for 7 minutes.
[0256] • Dilute proteinase K (NEB) 1:1 in water (2.5 pl of proteinase K and 2.5 pl of water) and add 5 pl of the dilution to tagmentation reaction. Incubate at 55 °C for 15 minutes.
[0257] • Purify with 1.8X (81 pL) of SPRI-Guanidine beads, elute in 11 pL of IDTE.
[0258] GUIDE-seq-2 PCR
[0259] Prepare PCR reaction. Each tagmented gDNA is split into + and - group (5 pL of input each). Each + and - sample will contain the same i7 barcode, the only difference is that primers will be specific for + and - group.
[0260] For the i7 dsODN primers, make a plate with 5 pM of the primers aliquoted. Example is shown below.
[0261]
[0262]
[0263] Run the PCR thermocycler program as follows (70 °C A-l.OC / cycle), for 15 cycles decrease the temperature for one degree per cycle starting with 70°C:
[0264] Step Temperature Time
[0265] Denaturation 95 °C 5 min
[0266] 15 cycles: 95°C 30 s
[0267]
[0268] 20 cycles: 95°C 30 s
[0269] 55°C 1 min
[0270] 72°C 30 s
[0271] extension 72°C 5 min
[0272] End 4°C Hold
[0273] • Purify with 1.8X (54 pl) SPRI beads, elute in 30 pl of TE.
[0274] qPCR library quantification
[0275] • Make 1:10 dilution of each + and - library from 101to 10’5with IDTE in total volume of 50 pl.
[0276] • 10’5diluted samples will be quantified in duplicates.
[0277] Prepare qPCR master mix:
[0278]
[0279] Set up qPCR reaction:
[0280]
[0281] • Pool library to achieve 8 x 109molecules in 5 pl or to 2 nM for NextSeq2000
[0282] SequencingSequence based on the information listed below specific to each sequencer. Index2 is 18nt: MiniSeq: 146, 8, 18, 146
[0283] 50 pl of denatured library + 6uL of 20 pM PhiX + 444 pl of buffer
[0284] MiSeq: 146, 8, 18, 146 (GUIDEseq_v2 sample sheet)
[0285] 10 pl of denatured library + 100 pl of 12.5 pM PhiX + 940 pl of buffer
[0286] NextSeq 550: 146, 8, 18, 146
[0287] 120 pl of denatured library + 15 pl of 20 pM PhiX + 1165 pl of buffer
[0288] NextSeq 2000: 146, 8, 18, 146
[0289] 550 pM of denatured library + 20% (vol / vol) of 550 pM PhiX (Load 20 uL) Tagmentation protocol using commercially available Tn5
[0290] Pre-assembled Tagify™ i5 UMI (seqWell 301210 or 301230) can be used in replacement of the above steps 1) Tn5 oligos annealing through 3) Tagmentation with the following protocol.
[0291] Reagents
[0292] • Pre-assembled Tn5: Tagify™ i5 UMI (seqWell 301210 or 301230) custom ordered for custom volume packaging.
[0293] • 3x coding buffer: seqWell 101000
[0294] • X solution: seqWell 101001
[0295] • MAGwise Magnetic Purification Beads: seqWell 101003
[0296] Tagmentation with Tagify™ i5 UMI
[0297] Prepare tagmentation reaction for each gDNA sample. 100 ng of gDNA is tagmented with Tn5 complexed with a unique i5 barcode.
[0298]
[0299] • Incubate the reaction at 55 °C for 7 minutes.
[0300] • Add 20 pl of X solution to the 40 uL tagmentation reaction. Incubate at 68 °C for 10 minutes.
[0301] • Purify with 1.8X (108 pL) of MAGwise beads and perform the wash / elution step the same as above.
[0302] Note 2: CHANCE-seq developmentCHANGE-Seq R
[0303] End-repair. After cleavage of the circularized, partially-randomized libraries with RNP complex, library preparation of resulting linear DNA was initially performed without end-repair. Consequently, data analysis showed unexpected enrichment profiles. Since the library is comprised of large numbers of off-targets with a range of mismatches, it was hypothesized that staggered cleavage was likely occurring frequently for some targets, resulting in 5’ overhangs and therefore, end-repair was necessary to observe the full spectrum of cleavage outcomes during library preparation for Illumina sequencing.
[0304] Assay Development. The initial experiments for CHANGE-Seq R utilized primers that did not have a unique molecular index (UMI) incorporated next to the randomized target site. Only one PCR was used to create a diverse library and then produced the circularized library from the products of this PCR. Without an initial bottlenecking step, this resulted in a library with hundreds of millions of unique molecules which made it challenging to obtain adequate sequencing depth in the control samples. Because there was not an added PCR step to make copies of the library, many off-targets that were present in the Cas9 cleaved were not present in the control sample and vice versa. Additionally, the end repair process resulted in base insertions in some Cas9 cleaved off-targets but not in control (Srfl produces blunt ends). Because there were no UMIs in the primers, this made it difficult to confidently locate the same off-target sequence in both the Cas9 and control sample and therefore difficult to calculate enrichment of a particular off-target. To overcome these issues, 15 bp UMIs were added to the primers and an additional PCR step to both bottleneck and copy the randomized linear DNA library.
[0305] The CHANCE-seq approach has three key advantages compared to earlier approaches described by Pattanayak et al.: 1) it is performed on circularized, partially randomized target libraries such that both parts of the cut and linearized target sequence can be readily sequenced after cleavage, 2) its read out is dependent on a single (rather than multiple) Cas9 cleavage events and 3) a controlled library size enables quantitative measurements of the enrichment of specific off-target sequences between control and nuclease treatments.
[0306] Number of Unique Molecules. After adding UMIs to the primers it was important to determine how many unique molecules could be surveyed that would allow adequate sequence depth of at least lOx coverage and ensure no collision issues, i.e., one UMI assigned to multiple off-target sequences. A 15bp UMI with an NNWNNW pattern, to reduce GC content, allows for approximately 33 million unique combinations. It was determined that 3-4 million unique molecules limited collision effects and allowed an achievable sequencing target of approximately 50 million reads per sample.Protocol 2: CHANCE-seq
[0307] Supplies
[0308] • IDTE pH 8.0 (IX TE Solution) (Integrated DNA Technologies, 11050204)
[0309] • Plasmid DNA p2T-CAG-eGFP-BlastR (modified from Addgene #107190)
[0310] • MilliporeSigma custom oligonucleotides
[0311] • Synthego custom sgRNA
[0312] • Proteinase K (NEB, P8107S)
[0313] • KAPA HiFi HotStart + ReadyMix (250 x 50 pl reactions) (Roche, KK2602)
[0314] • KAPA HiFi HotStart Uracil-i- ReadyMix (250 x 50 pl reactions) (Roche, KK2802) • Dpnl (NEB R0176L)
[0315] • Ultrapure Nuclease Free water (Invitrogen 10977023)
[0316] • Magnum FLX magnetic rack (Alpaqua, A00400)
[0317] • T4 Polynucleotide Kinase (PNK) (NEB, M0201L)
[0318] • T4 DNA Ligase (NEB, M0202L)
[0319] • 10X T4 DNA ligase Buffer (NEB), supplied with T4 DNA Ligase
[0320] • USER Enzyme (NEB, M5505L)
[0321] • Exonuclease I (E. coli) (NEB, M0293L)
[0322] • Lambda Exonuclease (NEB, M0262L)
[0323] • Exonuclease III (E. coli) (NEB, M0206L)
[0324] • Plasmid-Safe ATP-dependent DNase (Epicentre, E3110K)
[0325] • 10X Ampligase Buffer (Lucigen, A1905B)
[0326] • 25mM ATP solution (Epicentre), supplied with Plasmid-Safe ATP-dependent DNase • Srfl (NEB, R0629S)
[0327] • Quick CIP (NEB, M0525L)
[0328] • rCutsmart Reaction buffer (NEB, B6004SVIAL, supplied with Quick CIP and Srfl) • SpCas9 (NEB, M0386T)
[0329] • Klenow Fragment (3' -> 5' exo-) (NEB, M0212L)
[0330] • NEBuffer 2 (NEB, B7002SVIAL supplied with Klenow Fragment)
[0331] • dNTPs (NEB, cat.no. N0447L)
[0332] • NEBNext® Ultra™ II Ligation Module (NEB E7595L)
[0333] • NEBNext® dA-Tailing Module (NEB E6053L)
[0334] • Kapa PEG / NaCl SPRI solution (Roche, 7961928001)• NEBNext® Multiplex Oligos for Illumina® (Dual Index Primers Set 1) (New England BioLabs, E7600S)
[0335] • NEBNext adapter for Illumina (NEB), supplied with NEBNext® Multiplex Oligos for Illumina®
[0336] • Qubit dsDNA HS Assay Kit (Thermo Fisher Scientific, Q32854)
[0337] • NextSeq2000® kit 200-cycles (Illumina)
[0338] • Flow Cell, supplied with NextSeq2000® Reagent Kit
[0339] • NextSeq2000 RSB Buffer, supplied NextSeq2000® Reagent Kit
[0340] • PhiX Control V3 KIT (Illumina, FC- 110-3001)
[0341] • Ultra pure EDTA 0.5 M (Invitrogen, 15575038)
[0342] • Ethanol (Sigma, cat.no. E7023)
[0343] • Sera-Mag Magnetic Beads; Carboxyl, Speedbeads; hydrophobic; 5 solids (Fisher / GE,.9981123)
[0344] • Guanidine thiocyanate (Sigma, G9277)
[0345] • Sodium Chloride 5 M Sterile (Invitrogen, AM9760G)
[0346] • TRIS Buffer 1.0 M solution, pH 8.0 (Fisher, 50146868)
[0347] • Polyethylene Glycol 8000 (Fisher, 507516674)
[0348] • Kappa Library Quantification Kit (Illumina / Uni) (Roche, KK4824)
[0349] • IM Magnesium chloride (Invitrogen, AM9530G)
[0350] • Tween-20 (Sigma, P7949-500ML)
[0351] • HEPES (Fisher, BP310-1)
[0352] • Sodium Hydroxide solution (Sigma, 72068- 100ML)
[0353] • 12well, 2% Agarose gel Cassettes (Yourgene Health, CG- 10600- 13-200(16))
[0354] • Ranger MQ Dual Dye Loading Buffer, 300 bp + Ik bp Marker (Yourgene Health, CG- 14000-12-21)
[0355] Reagent Setup
[0356] Resuspend the CHANCE-seq Randomized custom oligonucleotides. For PCR1, forward mixed base primers were designed with a universal primer sequence, 15 base-pair unique molecular index (UMI) and partially randomized target sites to produce amplicons with a range of alterations at protospacer and PAM positions for each of the six target sites. For PCR2, Forward and Reverse ACG / deoxyUridine primers were designed to introduce a Srfl restriction site and ACGU sequences. Oligonucleotides were resuspended to 100 pM in TE pH 8.0.
[0357] Resuspended oligonucleotides were kept at -20°C.N = 25,25,25,25(GACT)
[0358] W=50.50(AT)
[0359] N1 = ACT (10%each) G=70%
[0360] N2 = GCT (10%each) A=70%
[0361] N3 = AGT (10%each) C=70%
[0362] N4 = ACG (10%each) T=70%
[0363] oAF01_CTLA4.s9_FWD ACGCGAGCTGCATGTGTCAGANNWNNWNNWNNWNNW(N1)(N1)(N2)(N3)(N4)(N1)(N 2)(Nl)(Nl)(Nl)(N3)(N3)(N2)(N4)(Nl)(Nl)(N2)(N3)(N2)(N3)(Nl)(Nl)(Nl)cttcttcaagtccgccatg c (SEQ ID NO: 35)
[0364] oAF02_AAVSl.sl4_FWD
[0365] ACGCGAGCTGC ATGTGTC AGANNWNNWNNWNNWNNW(N 1 ) (N 1 ) (N 1 ) (N 1 )(N3 )(N3)(N2)(N3)(N4)(N2)(Nl)(Nl)(Nl)(N2)(N3)(N2)(Nl)(Nl)(N2)(N4)(N4)(Nl)(Nl)cttcttcaagt ccgccatgc (SEQ ID NO: 36)
[0366] oAF04_LAG3.s9_FWD
[0367] ACGCGAGCTGC ATGTGTC AGANNWNNWNNWNNWNNW(N 1 ) (N2) (N2) (N 1 )(N 1 )(N3)(N4)(Nl)(N2)(Nl)(N2)(N4)(N3)(N3)(N4)(Nl)(Nl)(N2)(NI)(Nl)(Nl)(Nl)(NI)cttcttcaagtc cgccatgc (SEQ ID NO: 37)
[0368] oAF05_TRAC.sl_FWD ACGCGAGCTGCATGTGTCAGANNWNNWNNWNNWNNW(N1)(N4)(N3)(N2)(N1 )(Nl)(Nl)(N4)(N4)(N3)(N4)(Nl)(Nl)(N2)(N4)(N2)(N4)(N3)(N4)(Nl)(N4)(Nl)(Nl)cttcttcaagt ccgccatgc (SEQ ID NO: 38)
[0369] oAF06_CCR5.s8_FWD ACGCGAGCTGCATGTGTCAGANNWNNWNNWNNWNNW(N1)(NI)(N2)(N3)(N2) (Nl)(N4)(N2)(N2)(Nl)(N2)(N2)(Nl)(Nl)(N2)(N2)(N2)(N2)(N2)(N3)(N2)(Nl)(NI)cttcttcaagtcc gccatgc (SEQ ID NO: 39)
[0370] o.AF07_PCRl_REV- tgtgatcgcgcttctcgtt (SEQ ID NO: 40)
[0371] o.AF08_PCR2_FWD - ACG / deoxyUridine / acgcgagctgcatgtgtcaga (SEQ ID NO: 41) oAF09_PCR2_REV- ACG / deoxyUridine / GCCCGGGCtctcgttggggtctttgctc (SEQ ID NO: 42)Resuspend 1.5nmol gRNA from Synthego - Resuspend in 15ul lx TE buffer (lOmM Tris, ImM EDTA pH8.0 supplied with guides from Synthego). Aliquot and freeze at -80C. Avoid freeze / thaw cycles.
[0372] CTLA4.S9 - GGACTGAGGGCCATGGACACGGG (SEQ ID NO: 43) AAVS1.S14 - GGGGCCACTAGGGACAGGATTGG (SEQ ID NO: 44) LAG3.S9 - GAAGGCTGAGATCCTGGAGGGGG (SEQ ID NO: 45)
[0373] TRAC.sl - GTCAGGGTTCTGGATATCTGTGG (SEQ ID NO: 46)
[0374] CCR5.S8 - GGACAGTAAGAAGGAAAAACAGG (SEQ ID NO: 47)
[0375] Prepare Plasmid DNA - Prepare plasmid DNA p2T-CAG-eGFP-BlastR via miniprep and dilute to ~lng / ul.
[0376] SPRI-guanidine binding buffer 4M guanidine thiocyanate, 40mM TRIS, 17.6mM EDTA, pH 8.0. TRIS IM pH 8 and EDTA 0.5M pH 8 can be added to the 4M guanidine (after the guanidine is solubilized in water - add the proper volume for getting the right final concentration) and then the pH will be very close to 8. Bring the pH to 8 with HC1.
[0377] Sera-Mag Magnetic Beads preparation Add 1 ml of Sera-Mag Magnetic Beads (Fisher / GE) to a 1.5 ml Eppendorf tube. Place in a magnetic rack. Remove the liquid. Remove the tube from the rack. Add 1 ml of TE and homogenize. Place back in the magnetic rack and remove the liquid. Repeat this step for a total of two TE pH 8.0 washes. Then, add 1 ml of TE pH 8.0. Note: this beads preparation step is required for preparing SPRI-guanidine beads and SPRI-beads.
[0378] SPRI-guanidine beads preparation Add 10 ml of 5M NaCl to 9 g of PEG 8000 and then add SPRI-guanidine binding buffer (prepared as described above) up to 49 ml. Homogenize for 5 min. Add 1 ml of Sera-Mag Magnetic Beads in TE (prepared as described above) and homogenize. Keep at 4°C.
[0379] SPRI-beads preparation Add 10 ml of 5M NaCl, 500 pl of IM TRIS and 100 pl of 0.5M EDTA to 9 g of PEG 8000. Complete the volume to 49 ml with ultra-pure water. Add 1 ml of Sera-Mag Magnetic Beads in TE (prepared as described above) and homogenize. Add 27.5 pl of Tween-20 and homogenize. Keep at 4°C.
[0380] 10X Cas9 Nuclease Reaction Buffer - 200mM HEPES, IM NaCl, lOmM MgC12,lmM EDTA; bring to pH 6.5 at 25°C with NaOH.
[0381] Procedure
[0382] PCR1 -Creating highly diverse, partially randomized library for each target site
[0383] 1| Perform PCR1 for each target using the following components and PCR conditions.
[0384] Component _ Volume (pl) _
[0385] 2x Kapa HiFi HotStart -i-Ready Mix 251 OuM Mixed B ase+eGFP_Fwd Primer 1.5
[0386] lOuM O.AF07 PCR1 Rev Primer 1.5
[0387] p2T-CAG-eGFP-BlastR plasmid (Ing) x
[0388] Nuclease-free water x
[0389] Total 50
[0390] Step Temperature Time Cycles Denaturation 95 °C 3min 1
[0391] Denaturation 98 °C 20 s 20
[0392] Annealing 73 °C 15 s 20
[0393] Extension 73°C 30 s 20
[0394] Extension 72 °C 3min 1
[0395] Hold 4°C 1
[0396] 2| Add lul Dpnl directly to PCR reaction. Incubate at 37°C for 15 minutes. Hold at 10°C 3| Dilute proteinase K 1: 1 in water (2.5 pl of proteinase K and 2.5 pl of water) and add 5 pl of the dilution to PCR reaction. Incubate at 55 °C for 15 minutes.
[0397] 4| Add 1.8X volumes (90 pl) of SPRI-guanidine beads to the PCR reaction, mix thoroughly by pipetting 10 times. Incubate at room temperature for 10 minutes. Place the reaction plate onto a Magnum FLX magnetic rack for 5 minutes. Remove the cleared solution from the reaction plate and discard. Add 200 pl of 80% ethanol, incubate for 30 seconds and remove the supernatant. Repeat this step for a total of two ethanol washes. Remove ethanol completely and let the samples air dry for 3 minutes on the magnetic rack. Remove the plate from the magnetic rack and add 40 pl of TE pH 8.0, and pipette 10 times to mix. Incubate at room temperature for 2 minutes. Place the reaction plate back to the magnetic rack for 1 minute. Transfer the eluted DNA to a new plate.
[0398] Note: using SPRI-Guanidine beads in this purification step to completely inactivate and remove Kapa HiFi HotStart + polymerase will help increase yield of circularized DNA.
[0399] Carryover of this enzyme will remove the 3' overhangs generated in step 9, due to its strong 3'-5' exonuclease activity.
[0400] 5| Quantify PCR products with Qubit. Dilute each sample to a concentration of Ipg / ul.
[0401] *Safe Stop - store purified PCR products at 4°C or -20°C*
[0402] PCR2-Bottlenecking and making copies of library
[0403] 6| Perform PCR2 with Ipg / ul dilutions of PCR 1 products for each target using the following components and PCR conditions. Use Uracil tolerant polymerase to synthesize ACGU sequences.
[0404] Component Volume (pl)
[0405] 2x Kapa HiFi HotStart Uracil -i-Ready Mix 25
[0406] lOuM O.AF08 PCR2 Fwd Primer 1.5lOuM O.AF09 PCR2 Rev Primer 1.5
[0407] PCR1 product (Ipg) x
[0408] Nuclease-free water x Total
[0409]
[0410] Step Temperature Time Cycles Denaturation 95 °C 3min 1
[0411] Denaturation 98 °C 20 s 30
[0412] Annealing 72 °C 15 s 30
[0413] Extension 72°C 15 s 30
[0414] Extension 72 °C 3min 1
[0415] Hold 4°C 1
[0416] 7| Dilute proteinase K 1:1 in water (2.5 pl of proteinase K and 2.5 pl of water) and add 5 pl of the dilution to PCR reaction. Incubate at 55 °C for 15 minutes.
[0417] 8| Purify - Add 1.8X volumes (90 pl) of SPRI-guanidine beads to the PCR reaction, and follow purification as described in step 4. Add 40 pl of TE pH 8.0 to elute and transfer the eluted DNA to a new plate.
[0418] 9| USER / PNK. Set up the USER / PNK reaction as follows.
[0419] Component Volume (pl) T4 DNA Ligase Buffer ( 1 OX) 5
[0420] USER Enzyme (1 U / pl) 3
[0421] T4 Polynucleotide Kinase (10 U / pl) 2
[0422] Purified DNA (no more than 8ug) 40
[0423] Total 50
[0424] Incubate in thermocycler at 37 °C for 1 hour.
[0425] 10| Purifty - Add 1.8X volumes (90 pl) of SPRI beads to the PCR reaction, and follow purification as described in step 4. Add 35 pl of TE pH 8.0 to elute and transfer the eluted DNA to a new plate.
[0426] 11| Run each sample on a QIAxcel capillary electrophoresis instrument, in a 0.2 ml thinwalled 12-well strip tube with a QIAxcel DNA High Resolution Kit (Qiagen), QX Alignment Marker 50 bp to 1.5kb (Qiagen) and QX Size Marker 15bp - 3kb (Qiagen), following manufacturer’s instruction. The average size should of DNA should be -466 bp.
[0427] 12| Quantify by Qubit dsDNA High sensitivity assay. Calculate 500ng of DNA per reaction. Use TE pH8.0 to bring each reaction volume to 88ul. Save any remaining DNA not used for circularization to run as control on QIAxcel in step 19.
[0428] Note: Number of samples will likely double at this step.
[0429] 13| Intramolecular circularization. Set up the DNA circularization as follows:
[0430] Component Volume (pl)T4 DNA Ligase Buffer ( 1 OX) 10
[0431] T4 DNA Ligase (400 U / pl) 2
[0432] USER / PNK treated DNA (500 ng) variable
[0433] IDTE pH 8.0 variable
[0434] Total 100
[0435] Incubate in thermocycler at 16°C for at least 16 hours (overnight).
[0436] 14| Purify the circularized DNA reactions as previously described in step 4 by adding IX volumes (lOOpl) of SPRI beads to the DNA. Elute in 37 pl of TE pH 8.0. Transfer eluted DNA to a clean plate.
[0437] 15| Plasmid-Safe ATP-dependent DNase / Lambda Exo / ExoVExoIII treatment:
[0438] Component Volume (pl) Ampligase Buffer (10X) 5
[0439] ATP (25 mM) 2
[0440] Plasmid-Safe ATP-Dependent DNase (10 U / pl) 2
[0441] Lambda Exonuclease (5 U / pl) 2
[0442]
[0443] Incubate in a thermocycler at 37 °C for 1 h. Add 5 pl of EDTA 0.5 M, incubate at 70 °C for 30 min, hold at 4 °C.
[0444] 16| Purify the circularized, exonuclease treated DNA reactions as previously described in step 4 by adding IX volumes (50 pl) of SPRI beads to the DNA. Elute with 43 pl of TE pH 8.0. Transfer the supernatant to a new plate.
[0445] 17| Quick CIP treatment of circularized exonuclease-treated DNA. Set up Quick CIP reaction as follows:
[0446] Component Volume (pl) CutSmart Reaction Buffer (1 OX) 5
[0447] Quick CIP (5 U / pl) 2
[0448] Exonuclease Treated DNA 43
[0449] Total 50
[0450] Incubate at 37 °C for 10 min. Heat inactivate at 80 °C for 2 min.
[0451] 18| Purify CIP treated, circularized DNA as previously described in step 4 by adding 1.8X volumes (90 pl) of SPRI-beads to the DNA. Elute with lOpl of TE pH 8.0. Transfer the supernatant to a new plate, pool all reactions of the same sample type, and quantify by Qubit HS assay.19| QC Step: Run each sample on QIAxcel with 6ul circularized DNA. Run saved USER / PNK treated DNA as a size control. Circularized DNA should run faster on the gel. *Safe-Stop-Circularized, Exonuclease and CIP treated DNA can be stored at -20°C* Cleavage of enzymatically purified, circularized plasmid DNA
[0452] 20| sgRNA dilution and re-fold. Dilute the sgRNA to 9 pM in lOul nuclease-free water and use the follow program on a thermocycler for sgRNA re-fold:
[0453] Step Temperature Time Cycles
[0454] 1 90 °C 5 min 1
[0455] 2 90-25 °C Ramp rate
[0456] 2%
[0457] Hold 4°C 1
[0458] 21| In vitro cleavage with S.py Cas9. Perform step 21 concurrently with step 22 (Srfl cleavage).
[0459] A) Dilute Cas9 to luM
[0460] Component Volume (pl)
[0461] Cas9 Nuclease Reaction Buffer 6
[0462] (10X)
[0463] NEB S.py Cas9 (20uM) 3
[0464] Nuclease free H20 51
[0465] Total 60
[0466] B) Prepare RNP complex with Cas9 and sgRNA. Setup in vitro master-mix:
[0467] Component Volume (pl)
[0468] Cas9 Nuclease Reaction Buffer 5
[0469] (10X)
[0470] S.py Cas9 dilution mix (A) (luM) 4.5
[0471] sgRNA-folded (9 pM) 1.5
[0472] Total cleavage master-mix 11
[0473] Incubate at room temperature for 10 min.
[0474] C) In vitro cleavage. Add circularized DNA (125 ng, volume varies) and nuclease-free water to complete the 50 pl:
[0475] Cleavage master-mix 11
[0476] Exonuclease Treated DNA (125 ng) X
[0477] Nuclease-free water _ X_
[0478] Total 50
[0479] Incubate in a thermocycler at 37 °C for 1 h, hold at 4 °C.
[0480] 22| Prepare reaction for control restriction enzyme Srfl cleavage.
[0481] A) Prepare Master mix.Component Volume (pl)
[0482] Neb rCutsmart Buffer (10X) 5
[0483] NEB Srfl 1
[0484] NFW 5
[0485] Total cleavage master-mix 11
[0486] Incubate at room temperature for 10 min.
[0487] DNA should come from same pooled circularized sample as step 21 for each target. (125 ng, volume varies) and nuclease-free water to complete the 39 pl.
[0488] Component Volume (pl)
[0489] Cleavage Master-mix 11
[0490] Circularized DNA (125ng) X
[0491] Nuclease-free water X
[0492] Total volume 50
[0493] Incubate in thermocycler at 37 °C for 40 minutes, then 65 °C for 20 minutes, hold at 4 °C.
[0494] 19| Dilute proteinase K 1:5 in water (1 pl of proteinase K and 4 pl of water) and add 5 pl of the dilution to the in vztra-cleaved DNA and incubate in a thermocycler at 37 °C for 15 min.
[0495] 20| Purify cleaved (Cas9 and Srfl) DNA as previously described in step 4 by adding IX volumes (55 pl) of SPRI-beads to the DNA. Elute with 40pl of TE pH 8.0. Transfer the supernatant to a new plate.
[0496] 211 End repair. Setup the end-repair master mix:
[0497] Component Volume (pl) NEB Buffer 2 (10X) 5
[0498] Klenow fragment (3'->5' exo-) (5 U / pl) 1
[0499] dNTP 4
[0500] Total end-repair master- mix 10
[0501] Add 10 pl end-repair master- mix to each eluted DNA sample.
[0502] End-repair master-mix 10
[0503] cleaved DNA 40
[0504] Total 50
[0505] Incubate on a thermocycler at 37 °C for 30 min, then 75 °C for 20 min, hold at 4 °C.
[0506] 22| Purify end-repaired DNA as previously described in step 4 by adding IX volumes (50 pl) of SPRI-beads to the DNA. Elute with 42pl of TE pH 8.0. Transfer the supernatant to a new plate. Keep the beads.
[0507] 23| A-tailing. Setup the A-tailing master mix
[0508] Component _ Volume (pl) NEBNext dA-Tailing Reaction Buffer _ 5 _>
[0509] Klenow Fragment
[0510]
[0511] " > " exo-) 3
[0512] Total A-tailing master-mix 8
[0513] Add 8 pl of A-tailing master-mix to each eluted DNA sample with beads.
[0514] A-tailing master-mix 8
[0515] Cleaved DNA / beads 42
[0516] Total 50
[0517] Incubate on a thermocycler at 37 °C for 30 min, hold at 4 °C.
[0518] 24| Purify A-tailed DNA as previously described in step 4 by adding 1.8X volumes (90 pl) of SPRI-beads to the DNA. Elute with 60pl of TE pH 8.0. Transfer the supernatant to a new plate. Keep the beads.
[0519] 25| Adapter ligation. Setup the adapter ligation master-mix
[0520] Component _ Volume (pl) NEBNext Ultra II Ligation Master Mix 30
[0521] NEBNext Ligation Enhancer 1
[0522] Total MM 31 End Prep Reaction Mixture / beads 60
[0523] NEB N ext Adapter for Illumina ( 15 p M) 2.5
[0524] Total reaction volume 93.5
[0525] Incubate on a thermocycler at 20 °C for 15 minutes with heated lid off, hold at 4 °C. Note'. Prepare single-use aliquots of NEB adapters to avoid adapter dimer formation due to freeze-thaw hydrolysis of the 3' T'
[0526] 26| Purify Add IX volumes (50 pl) of PEG / NaCl solution to the adapter-ligated DNA and purify DNA as described in step 4. Elute in 47 pl of TE pH 8.0 and keep the beads.
[0527] 27| USER enzyme. Add 3 pl of USER enzyme, provided with NEBNext® Multiplex Oligos for Illumina® (Dual Index Primers Set 1) to the adapter ligated DNA with beads.
[0528] Incubate at 37 °C for 15 min.
[0529] 28| Purify Add 0.7X volumes (35 pl) of PEG / NaCl solution to the USER Enzyme treated DNA and purify as previously described in step 4. Elute in 20 pl of TE pH 8.0. Transfer the supernatant to a new semi-skirted PCR plate and quantify by Qubit dsDNA HS assay and proper Qubit assay tubes (usually about 1-5 ng / pl).
[0530] 29| PCR. Prepare PCR master-mix for adding dual-index barcodes:
[0531] A)
[0532] Component Volume (pl)
[0533] 2x Kapa HiFi HotStart Ready Mix 25NEB next i5 Primer (lOuM)
[0534] NEB next i7 Primer (lOuM) USER treated DNA (10-20ng) Nuclease-free water Total
[0535]
[0536] Note: Carefully record which primers were used for each
[0537] B) Perform the PCR using the following thermocycling conditions:
[0538] Step Temperature Time Cycles Denaturation 98 °C 45 s 1
[0539] Denaturation 98 °C 15 s 20
[0540] Annealing 65 °C 30 s 20
[0541] Extension 72 °C 30 s 20
[0542] Extension 72 °C 1 min 1
[0543] Hold 4°C 1
[0544] *Safe-Stop. PCR can hold at 4°C overnight *
[0545] 30| Purification. Add 0.7X volumes of SPRI-beads to the PCR and purify as previously described in step 4. Elute in 30 pl of TE pH 8.0. Transfer the supernatant to a new semi-skirted plate.
[0546] Note: Can run samples on Qiaxcel to determine purity and size of NGS library prepared samples.
[0547] 31| Quantify library for sequencing. Make 1:10 serial dilutions of 20 pl from 101to 10’7dilution of each sample from the library (PCR), starting with 2pl of DNA and 18 pl of TE, and mix well.
[0548] A) Assemble qPCR with Kapa library quantification kit. Prepare master-mix solution as follows:
[0549] Component _ 1 reaction (pl) Final Concentration KAPA SYBR FAST qPCR Master 12 IX
[0550] Mix (2X) + Primer Premix (10X)
[0551] Nuclease-free water _ 4 _
[0552] Total qPCR mix 16
[0553] B) Assay 4 different dilution factors (4 pl) for each sample (IO-4,to 10’7from the library) in duplicate (in an appropriate 96-well plate). A standard curve (provided with Kapa Library Quantification Kit) and a non-template control (NTC) are required. Add 4 pl of each standard in duplicate, and nuclease-free water in the NTC. Add 16 pl of qPCR master-mix to each sample.Component Volume (pl) Final concentration
[0554] qPCR mix 16
[0555] Sample (add nuclease-free water 4 variable
[0556] into the NTC well)
[0557] Total 20
[0558] C) Seal the plate and spin down. Run qPCR in appropriate thermocycler with the following program:
[0559] Cycling step Temperature Time Cycles
[0560] Initial denaturation 95 °C 5 min 1
[0561] Denaturation 95 °C 30 s 35 Annealing / extension / data 60 °C 45 sec 35
[0562] acquisition
[0563] Melt curve analysis 60-95 °C
[0564] D) Add the appropriate DNA copies for each standard when setting up the qPCR plate in the qPCR program, as follows:
[0565] Standard dsDNA molecules / pl
[0566] Standard 1 1.2xl07
[0567] Standard 2 1.2xl06
[0568] Standard s 1.2xl05
[0569] Standard 4 1.2xl04
[0570] Standard s 1.2xl03
[0571] Standard 6 1.2xl02
[0572] 32| Analyze qPCR results. Multiply the average of duplicate values by the dilution factor and by the five-fold dilution factor of the qPCR reaction, as follows: Total copies / pl = # * dilution factor.
[0573] 33| Pool library for NextSeq2000. Pool all the samples in one library at equimolar concentrations. IX pooled library should be in a total volume of 5 pl, ~ 8 x 109molecules.
[0574] 34| Size selection - to remove smaller DNA fragments, use 2% agarose gels and 300- Ikb marker from Yourgene Health. Size select 500~900bp fragment of pooled library with LightBench automated size selection instrument.
[0575] 35| Confirm size of library by running pooled, sized selected sample on Qiaxcel.
[0576] 36| Check concentration with Qubit High Sensitivity Assay. Use the size of the library calculated from step 35. Concentration should be around 2nM.
[0577] 37| Loading the sample for sequencing.
[0578] A) Dilute library to 625pM with RSB buffer (supplied with Nextseq2000 kits) for a total of 24ul.B) Prepare the Phix control V3 (PhiX Control V3 KIT) by mixing 2ul of lOnM PhiX control with 8ul RSB buffer for a total of lOul
[0579] C) Take 7.8 ul of PhiX mix from step B and mix with 16.2ul RSB for a total of 24ul D) Take 20.4ul of library mix from step A and mix with 3.6ul PhiX mix from step C.
[0580] 35| Load and sequence library using a NextSeq2000200-cycles kit according to manufacturer’s instructions using NextSeq 2000 system. Sequencing is performed with 100 bp paired-end reads and 8 bp dual-index reads. At least lOx coverage of each sample is required.
[0581] Example 2. Frequent context-dependent effects of common genetic variants on Cas9 editing activity revealed by population-scale GUIDE-seq-2 and massively parallel CHANCE-seq profiling (Update to Example 1 )
[0582] Genome editing technology has tremendous potential to become transformative therapies for a wide range of genetic diseases. The recent regulatory approval of exagamglogene autotemcel (exa-cel) — a Cas9 autologous genome edited hematopoietic stem cells to induce fetal hemoglobin to treat sickle cell disease — in the US and Europe1011marks a major milestone in this exciting field. Several other genome editing therapies are currently undergoing clinical trials12.
[0583] However, genome editors can also introduce unintended modifications at off-target sites13, such as large-scale chromosomal translocations or deletions associated with DSBs9 15 16. This unintended, off-target editing can potentially confound biological research and pose safety risks in therapeutic applications14. For example, Lorenzini et al. describe an off-target site in the intron of USP9X gene that resulted in disrupted expression of this gene and over 150 downstream genes. Of primary concern for clinical gene editing is that off-target activity may result in cells with a proliferative advantage, increasing the potential for malignant transformation — a risk observed in early and recent gene therapy studies with viral vectors17 l9.
[0584] Individual genomes differ from the human reference genome at approximately 4-5 million positions20. Furthermore, intraindividual somatic variation also occur, although the frequency of these spontaneous variants is generally orders of magnitude lower than that of germline variants that are present in every cell within that individual. Genetic variations such as single-nucleotide variants (SNVs), insertions, and deletions can overlap with Cas9 target-complementary or protospacer adjacent motif (PAM) sites. Such overlap can alter activity at existing off-target sites or create new ones. Thus, for exa-cel and similar therapies, a critical question is how genetic variation within patient populations may influence the risk of unwanted off-target editing.
[0585] As genome editing therapies expand to broader patient groups and different targets, accounting for human genetic diversity will be essential to ensure their consistent safety and efficacy acrosspopulations. Despite intense interest, the frequency and extent to which genetic variation generally affect cellular editing remains largely unknown. This major gap in knowledge is due to the lack of scalable, unbiased methods to experimentally measure genome-wide on- and off-target editing activity in cells from large, diverse populations, as well as tools capable of comprehensively assessing sufficiently diverse off-target sites to acquire fundamental understanding. Although previous studies have shown that genetic variation can affect genome editing and off-target activity, they have primarily focused on in silico predictions of variant-containing off-target sites rather than unbiased, genome- wide experimental assessments. Misek et al. found that genetic variation can confound the analysis of ancestry-associated genetic dependencies for cancer cell survival. However, the study presented in this example does not directly measure on- or off-target mutation frequencies, but rather indirectly quantifies changes in gRNA abundance as a proxy for on-target editing activity21. A variant (rsll4518452-C) in the intron of CPS1 was identified by Cancelierri et al. to cause off-target editing associated with the exa-cel target at that genomic location. Recently, Yen et al. confirmed that low-frequency off-target editing in the presence of this variant was detectable in two SCD patients treated with exa-cel, underscoring the importance of variation- aware off-target analysis.
[0586] Unbiased methods (those that do not depend on a priori pre-selection of candidate sites) to define the frequency and location of unintended genome- wide off-target acti v i ty9 / 22 26fall broadly into two categories: cellular methods that are more direct, and biochemical approaches that offer higher sensitivity. GUIDE-seq9,27is a broadly adopted cellular method to define genome- wide on-and off-target activity of genome editing nucleases such as Cas9 which was instrumental in defining the specificity of exa-cel28. It is based on the principle of efficiently integrating a short double- stranded oligodeoxynucleotide (dsODN) tag into nuclease-induced double- stranded breaks (DSBs) followed by tag-specific amplification of surrounding genomic DNA. CHANGE-seq is a biochemical method for selective sequencing of editor-modified DNA that is more sensitive than GUIDE-seq22, but requires further validation in cells. Improved versions of these assays with increased scalability or off-target diversity would enable improved understanding of the impact of genetic variation on genome editing.
[0587] Although early studies demonstrated that SNVs could influence Cas9 activity29, they could not systematically discern how prevalent or impactful these variants are in a population-scale cellular context. Previously, CHANGE-seq was used to demonstrate that SNVs can significantly impact Cas9 biochemical in vitro cleavage in six donors22. Computational studies leveraging data from the 1000 Genomes Project and Human Genetic Diversity Project estimated the theoretical prevalence of variant effects using off-target scoring algorithms6,7. Cancellierri et al. developedCRISPRme, a variant-aware off-target search tool, to identify off-target sites for the BCL11A enhancer target in exa-cel, and validated a variant-specific off-target in human hematopoietic stem cells that resulting in allele- specific indels and pericentric inversions8. A common limitation of in silico off-target studies is their reliance on genomic databases predominantly composed of individuals of European ancestry, potentially underestimating off-target risks in other populations30.
[0588] To address this critical gap in understanding the impact of individual genetic variation on Cas9 activity, GUIDE-seq-2 was developed, a new, streamlined, high-throughput method for population- scale, genome-wide profiling of Cas9 off-target activity. GUIDE-seq-2 was applied to six gRNAs in lymphoblastoid cell lines from 95 individuals characterized by the 1000 Genomes Project20,31. A high-confidence set of variants that affect off-target activity was identified by multiplex targeted sequencing. These results reveal that genetic variants frequently overlap Cas9 off-target sites and alter cellular off-target activity. To further investigate these effects, a novel massively parallel biochemical assay, CHANCE-seq, was developed that enabled the identification of sequence contexts and features associated with off-target cleavage in vitro. Results
[0589] Development of a high-throughput method to evaluate cellular genome-wide activity of genome editors
[0590] To enable the analysis of Cas9 nuclease genome- wide off-target activity in a large number of well-characterized lymphoblastoid cell lines from genetically diverse individual31, it was reasoned that GUIDE-seq could be streamlined9,27, as it would be challenging to scale the original method to many samples.
[0591] To substantially increase scalability, subtractive changes32were searched for that resulted in a streamlined and improved, high-throughput method called GUIDE-seq-2 which substantially reduces processing steps, time, and cost, compared to the original method. GUIDE-seq-2 utilizes Tn5-transposon-based tagmentation33 36and a new library design that eliminates the need for nested rounds of PCR (FIG. 2A), integrates GUIDE-seq improvements37without changing the original GUIDE-seq tag sequence to maintain full compatibility, and simplifies both library preparation and sequencing.
[0592] The redesigned GUIDE-seq-2 PCR strategy simultaneously improves detection accuracy by enabling mispriming detection, and sequencing without custom primers and sequencer configuration that were previously required. By introducing tagmentation which simultaneously fragments genomic DNA and adds required sequencing adapters, multiple library preparation was eliminated and purification steps that include random genomic DNA shearing, end repair / A-tailingand adapter ligation, and reduced genomic DNA input requirements by 4-fold (FIGs.7A-7B). All experiments for GUIDE-seq-2 development were performed in primary human T cells.
[0593] GUIDE-seq-2 was performed with GUIDE-seq at 8 Cas9 target sites across 5 therapeutically relevant loci in CD4+ / CD8+human primary T cells (CCR5, TRAC, PDCD1, AAVS1 and CTLA4) with variable numbers of known off-targets identified in an earlier study22. A strong correlation was found between GUIDE-seq and GUIDE-seq-2 read counts at these 8 sites (r = 0.8 to 1.0, r is Pearson’s correlation coefficient) (FIG. 2B). For each of the 8 targets, independent GUIDE-seq-2 library preparations were performed and it was found that replicate read counts were strongly correlated (r = 0.99 to 1.0) (FIG.2C), demonstrating high technical reproducibility. The positions of on- and off-target sites were genome- wide as expected (FIG. 2D). GUIDE-seq-2 sensitivity was further estimated using defined mixtures of genomic DNA from ‘on-target’ single-cell clones with only on-target GUIDE-seq-2 tag integrations with DNA from ‘off-target’ clones with both on- and off-target tag integrations. It was found that the lower limit of detection (LLOD) , was between 0.01 and 0.1% of dsODN integrated cells (n=3). The lower limit of quantification (LLOQ) was approximately 0.1% of dsODN integrated cells (FIG. 2E).
[0594] To filter out mispriming artifacts, in the GUIDE-seq-2 design the original tag sequence was kept but the length of the primer annealing to the dsODN tag region was reduced, allowing for observation of expected sequences of 5 to 12 bases of dsODN region beyond the primer, and properly identify the expected junction between dsODN and flanking genomic DNA in the sequencing data (FIG. 7B). An advantage of shortening the tag- specific primer (instead of increasing the tag length37), is that tag integration rates upon nucleofection remain unchanged, in contrast to the decreases in tag integration and sensitivity that can occur when tag length is increased38.
[0595] To further streamline the GUIDE-seq-2 workflow, a bead-based normalization step was introduced, using limiting amount of beads for single-step PCR purification and DNA normalization39, avoiding labor-intensive serial dilution step for library quantification. Additionally, we adapted and optimized GUIDE-seq-2 for an automated liquid handling platform, demonstrating high technical reproducibility and comparability to manual GUIDE-seq-2, and expanding its potential for high-throughput applications (FIGs.8A-8E). In summary, the GUIDE-seq-2 approach is comparable to original GUIDE-seq in sensitivity, is highly reproducible, automatable, requires less genomic DNA for library preparation, reduces processing steps and time to 3 hours, and has improved detection accuracy.
[0596] CR1SPR-Cas9 cellular off-target editing sites frequently overlap human genetic variantsGUIDE-seq-2 enables efficient and precise assessment of cellular Cas9 off-target activity, allowing investigation of the impact of human genetic variation at unprecedented scale. Specifically, cellular on- and off-target sites associated with six gRNA targets and one control were analyzed from 95 donors across four populations, in a total of 665 GUIDE-seq-2 libraries (ASW, African Ancestry in Southwest USA; CEU, Utah residents with Northern and Western European ancestry; CHB, Han Chinese in Beijing; and MXL, Mexican Ancestry in Los Angeles, cell lines20). The six gRNA target sites (AAVS1 site 14, CCR5 site 8, CTLA4 site 9, CXCR4 site 8, LAG3 site 9 and TRAC site 1) were selected due to their intermediate levels of off-target activity as measured in a previous study22(gRNAs). GUIDE-seq-2 dsODN integration was optimized for lymphoblastoid cell lines (LCLs) characterized by the Genome-in- a-bottle consortium (GIAB)40and evaluated at two targets with GUIDE-seq-2 (FIGs. 9A-9E). As a control, to identify the background of spontaneous, naturally occurring DNA DSBs, cells from each donor were transfected with dsODN tag only (without Cas9 or gRNA).
[0597] 1,115 unique on- and off-target sites were identified genome-wide ranging from 56 to 333 per target (FIG. 3A), using stringent criteria that require at least two molecularly distinct GUIDE-seq-2 dsODN tag integration events for off-target site calling. These sites were distributed mostly in intergenic and intronic areas that comprise the majority of the human genome, and found also in promoter, 5’ and 3’ UTR, and exonic regions in expected proportions (FIG. 3B).
[0598] Cellular off-targets detected by GUIDE-seq-2 frequently overlapped with human genetic variants (FIGs. 3C-3D, FIGs. 10A-10C). Specifically, 16.6% of on- or off-target sites (185 of 1,115) overlapped with one or more genetic variants found in the cohort (n=212 overlapping variants), with individuals of ASW ancestry having the highest frequencies of off-targets overlapping genetic variation (13%), and those of CHB ancestry having the lowest (7.8%) (FIG.
[0599] 3E). When considering the entire 1000 Genomes Project population of 3,202 individuals with high-coverage whole genome sequencing (n=622 overlapping variants), the overlap percentage increased to 41.5% (463 of 1,115 sites) (FIG. 3F). Finally, when considering all variants in gnomAD (Genome Aggregation Database, n=6,311 overlapping variants), currently the largest collection of human genetic variation data41, it was found that nearly all (i.e., 99% or 1,104 of 1,115) of sites overlap with at least one annotated genetic variant (FIG. 3G).
[0600] Among these variants, SNPs are the most frequent type, followed by insertions then deletions (FIG. 3H). Among the SNPs, transition mutations outnumber transversion mutations by an approximately a 2:1 ratio, consistent with expectations for genome-wide studies42(FIG. 31).
[0601] Overlapping variants have a wide range of frequencies across populations. In the cohort, 60% variants are common variants (defined as minor allele frequency or MAF>1%) (FIG. 3J). Asexpected, more rare variants were found as variant database size increased, with 79% rare variants in the 1000 Genomes Project and 98% in gnomAD.
[0602] Most GUIDE-seq-2 detected sites (n= 161) overlapped with only one variant, although 24 off-target sites harbored two or three (FIG. 3K). It was observed that overlapping SNP variants were evenly distributed across the protospacer and PAM positions (FIG. 3L) with an average of 7.6 SNPs per position. Interestingly, some variants were highly population-specific, suggesting the importance of population-specific, variant-aware analyses. For example, SNP rs78326679, occurring at the second position in LAG3 site 9 off-target, is much more likely to occur in people of African American ancestry (FIG. 3M). It was estimated that the average common off-target affecting variant will be found in 16.2% of individuals, however, this number will depend on the variant allele frequencies (VAF) observed in each specific population (FIG. 3N). For a more comprehensive estimate, this analysis was expanded to 58 targets and 1403 off-targets analyzed by GUIDE-seq, then looked for overlap of common variants in the 1000 Genomes Project. The average off-target affecting variant was found in 22.1% of individuals (FIG. 19A). When stratifying by targets with high (>100), medium (20 to 100) and low (<20) numbers of off-targets, no significant differences were found between groups. (FIG. 19B). On a per individual basis, it is estimated that each person will have an average of 3% off-targets overlapping with common variants, on average (FIG. 19C).
[0603] Overall, the study of 95 individuals efficiently surveyed 97% of common variants and 34.1% of all variants in the 1000 Genomes Project (with MAF down to 0.03%) that overlapped the GUIDE-seq-2 cellular off-target sites discovered herein (FIG. 30). Based on these results, it is estimated that on average, a sample of 55 individuals would be sufficient to survey 95% of common variants that overlap cellular off-target sites (FIG. 30).
[0604] Validation of GUIDE-seq-2 discovered off-target sites
[0605] To validate GUIDE-seq-2 discovered off-target sites and identify a high-confidence set of variant effects on off-target editing activity, multiplexed targeted sequencing was conducted to measure editing frequencies at 283 on- and off-target sites associated with the 6 studied targets . This panel was designed to cover as many variant-containing off-targets as possible while encompassing a broad range of editing activities (based on GUIDE-seq-2 read counts). Out of the total 283 on-and off- targets sites, 185 harbored genetic variants. The multiplex targeted sequencing panel successfully covered 148 sites, corresponding to 167 variants or 90% of variant containing sites. At least 1000 sequencing reads per amplicon was required for downstream analysis. Unedited controls were included for each panel (FIGs. 17A-17B). Then, the significance of variant effectson editing activity was assessed as measured by indel and GUIDE-seq-2 tag integration frequency (FIG. 4A).
[0606] Using GUIDE-seq-2, a broad range of off-target activities were quantified associated with 6 gRNAs targets that varied substantially between individuals from the ASW, CEU, CHB, and MXL populations (FIG. 4B). It was confirmed that GUIDE-seq-2 read counts were strongly correlated with indel and GUIDE-seq-2 tag integration frequencies as measured by multiplex targeted sequencing (FIG. 11A). It was estimated thatGUIDE-seq-2 lower limits of detection to be -0.2%, based on the 930 sites not overlapping with any variants in the 95 donors and the lowest estimated indel frequencies of off-target sites could be detected in 95% of samples (estimated by the product of GUIDE-seq-2 relative off-target activity and on-target editing percentage) (FIG. 4C).
[0607] Overlapping genetic variants frequently affect off-target editing in human cells
[0608] To verify variant effects on off-target activity, five variant effect measures were integrated. Specifically, the effect of overlapping genetic variation was modeled on normalized GUIDE-seq-2 read counts, indel frequencies, tag integration frequencies, and allele-specific editing. If Cas9 exhibits a preference for reference or alternative genotypes, this should be reflected as an imbalance in the frequency of edited alleles (FIG. 4D). A high-confidence variant effect was defined as one that was significant in two or more of these measures (see Methods).
[0609] There was no significant correlation between variant allele frequency and GUIDE-seq-2 variant effect (r=0.002, p-value=0.9) (FIG. 41). A high proportion of variants (14.1%) measured by GUIDE-seq-2 significantly impacted off-target activity, of which 6.6% (n=14) and 7.5% (n=16) of variants led to increased or decreased editing, respectively. Of the 14 variants that increased editing, 4 were de novo sites in which variants created new off-target sites not detected in the reference genome (FIG. 4J). Variant effects on GUIDE-seq-2 reads ranged from -2.3 to 7.0 (log2-transformed GUIDE-seq-2 read counts) and were corroborated by comparable changes in indel and tag-integration percentages (FIG. 4K, FIG. 4M). Notably, variant effects on TRAC site 1 on-target activity (high-confidence variant #15) and variant #13, 14 and 16 were only identified by targeted amplicon sequencing, as it was at the upper limits of quantification by GUIDE-seq-2 (FIG.4K, FIGs. 17A-17B). Additional variants that had significant GUIDE-seq-2 read counts but no targeted sequencing data available, as well as those variants with deletions can be seen in FIGs. 4M and 17A-17B.
[0610] Certain variant effects were unexpected, such as a T>G variant at the first position of a Cas9 off-target sequence associated with a CC75-targeting gRNA (high-confidence variant #1), which surprisingly led to a substantial 70-fold increase in Cas9 indels to -10% (FIG. 4K, FIG. 4G,FIGs. 17A-17B). These findings underscore the complexity of predicting Cas9 activity on variantcontaining genomic sites effects and highlight the value of high-throughput profiling techniques.
[0611] Others were more predictable, such as those that revert a mismatch to the canonical .S'. pyogenes Cas9 NGG PAM sequence. For example, a T>G variant that altered an NGG PAM-containing off-target sequence (associated with a gRNA targeted to CXCR4, high-confidence variant #4) to NTG significantly reduced Cas9 editing activity, consistent with the importance of PAM recognition to license Cas9 target DNA cleavage43(FIG. 4K, FIG. 4G, FIGs. 17A-17B).
[0612] Strong variant effects on off-target editing observed in human primary T cells
[0613] To determine whether these results extended to therapeutically relevant primary human cells, GUIDE- seq-2 was performed in primary human T cells from three healthy donors and evaluated Cas9 off-target activity for two gRNAs target sites (CXCR4 site 8 and LAG3 site 9). By directly evaluating GUIDE-seq-2 read counts correlations between the three donors (FIG. 11B), strong evidence was found that the genome-wide off-target activity of Cas9 in these human primary T cells is affected by the presence of SNVs (FIG. 4H) with approximately 8 to 3,000-fold changes in GUIDE-seq-2 read counts. To understand how cell-type influenced editing outcomes, GUIDE-seq-2 on- and off-target read counts were then compared from the 95 LCL donors and 3 T-cell donors for 3 target sites (CXCR4 site 8, LAG3 site 9 and CCR5 site 8) at off-target sites that did not overlap with genetic variants. GUIDE-seq-2 read counts were strongly correlated between LCLs and T cells (Pearson correlation coefficient of 0.92), suggesting that data generated in LCLs are generalizable to other cell types (FIG. 4L).
[0614] Massively parallel biochemical assay to quantify activity at diverse mismatched off-targets The population-scale GUIDE-seq-2 analysis identified the largest set of genetic variants known to affect Cas9 genome-wide off-target activity and revealed the first estimates of the frequency with which overlapping off-targets modulate Cas9 activity. However, it would be challenging to infer general characteristics of genetic variants likely to impact Cas9 activity based on 21 variant- affected off-target sites alone.
[0615] To better understand influential genetic context and mismatches, a high-throughput, cost-effective, biochemical assay was designed that could measure Cas9 activity within highly diverse, partially randomized, off-target libraries called Combinatorial High-throughput Analysis of Nuclease Cleavage Efficiency by Sequencing or CHANCE-seq (FIG. 5A, FIG. 13A). It was reasoned that quantitative measurements of enrichment of millions of variations of target sequences between control and Cas9 treatment conditions could be more informative than earlier ensemble measurements of larger libraries44, where library members are unlikely to be shared between treatment conditions.CHANCE-seq builds on principles of CHANGE-seq, a previously developed method to selectively sequence DNA cleaved and linearized by Cas9 from populations of circularized genomic DNA molecules22, but operates on large, partially randomized off-target libraries instead of genomic DNA. CHANCE-seq enables profiling of activity of Cas9 at millions of mismatched targets with many higher-order combinations of 3 or more mismatches. CHANCE-seq libraries can be orders of magnitude larger than earlier oligonucleotide array-based approaches45,46, overcoming limitations that off-target sites identified from genomic DNA are usually small in number.
[0616] Using CHANCE-seq, five of the targets (AAVS1, CCR5, CTLA4, LAG3, and TRAC) that were profiled with population- scale GUIDE-seq-2 were surveyed. First, highly diverse libraries (IO10or more) were generated where each position of a 23 bp Cas9 target and PAM sequence has been partially randomized by PCR, resulting in unique molecular index (UMI) barcoded libraries of off-targets with an average of 6 mismatches when compared to the original on-target sequence (FIG. 5A). The presence of UMIs can be used to correct for sequencing errors and other artifacts that could confound analyses. Second, critically, to generate circularized DNA libraries containing multiple copies of the same randomized target, a subsample of approximately 1 million sequences were bottlenecked and amplified. Third, these circularized DNA libraries were split and treated with either Cas9 that can recognize and cleave similar on- or off-target sites, or a control restriction enzyme that cuts a site adjacent to the randomized target sequence, cleaving all molecules. Only cleaved DNA circles have free ends for subsequent adapter ligation and high-throughput sequencing. For each on- or off-target sequence, the ratio of normalized CHANCE-seq read counts were calculated between Cas9 and control libraries as a measure of Cas9 activity (FIG. 5A, FIG.
[0617] 13A).
[0618] Massively parallel CHANCE-seq measurements are globally consistent with Cas9 activity To determine whether these CHANCE-seq experiments effectively enriched for Cas9-cleaved DNA molecules, the frequency of mismatches within the libraries were compared when cleaved by Cas9 versus a control restriction enzyme that cuts all library molecules. It was observed that Cas9-cleaved libraries were significantly enriched for sites with fewer mismatches compared to control libraries (X2, p < 0.05), indicating that the assay was functioning as intended (FIG. 5B).
[0619] Relative off-target activity measured between CHANCE-seq technical replicates was strongly correlated (r = 0.89 to 0.97) (FIG. 5C, FIG. 13B). Cas9 PAM preferences determined by CHANCE-seq showed highest activity (94.9% ± 8.5%) at canonical NGG PAM, some activity at NAG (69.3% ± 7.8%) and NGA (35.5% ± 17.2%), and minimal activity at other non-canonical PAMs (FIG. 5D). The coverage of sequences with up to 6 mismatches ranged from 0.9 to 1.2million sequences for each target, after filtering for adequate representation in control libraries (FIG. 5E). Off-target activity typically decreased with increased mismatch number, although a long tail of exceptions where high mismatch sites retain high activity were observed (FIG. 5F, FIG. 13C).
[0620] The longest blocks of contiguous mismatches that preserved high activity were found in PAM-distal regions (starting at 5’ positions 1-2), but also in the middle (starting at positions 9-11) (FIG.
[0621] 5H, FIG. 13D). These two regions where Cas9 is most tolerant to contiguous off-target site mismatches coincide with those where structural studies have found proteimDNA phosphate backbone contacts (between target strand positions 1-2, 8-9, 10-13) that stabilize off-target site unwinding and recognition47.
[0622] Transition mutations such as G>A or T>C that result in wobble base pairings (rG:dT or rU:dG) were generally well tolerated and had the smallest effect on activity (FIG. 13E), specifically at positions 4, 16 and 22 (FIG. 13F), whereas Hoogsteen and transversion mutations were most deleterious for Cas9 activity (FIG. 13E) particularly at positions 1, 22, and 23 (FIG.
[0623] 13F). In addition to the PAM, positions 16 and 18 appeared least tolerant to mismatches while maintaining activity. At position 16, there is a stark contrast between transition mismatches that resulted in high editing activity and transversions that lowered activity (FIGs. 14A-14D, ISA-ISE). At higher off-target mismatch numbers, high-activity sites often accumulate more PAM-distal mismatches and mismatches become less common in the seed region (FIGs. 18A-18E).
[0624] CHANCE-seq enables quantitative assessment of variant effects in diverse contexts by comparing activity differences between homologous off-targets. Specifically, all pairs of off-targets were enumerated qith up to 6 mismatches that differed by a single mismatched nucleotide and defined single nucleotide variants (SNVs) as those that increase mismatches to their on-target sequence, resulting in 7.36 million variant pairs (FIG. 5L). The majority of variants that increase the number of mismatches relative to the on-target site reduce Cas9 activity (FIG. 6B). Variants have a greater impact on activity (‘leverage’) when the total numbers of “off-target context” mismatches is low (e.g. 0-2) compared to those with higher mismatches (3-5) (FIG. 6B).
[0625] CHANCE-seq reveals novel, context-dependent, synergistic effects of mismatches on Cas9 activity
[0626] By increasing coverage of activity at mismatched targets by more than X-fold compared to earlier methods (FIG. 5E), CHANCE-seq enables the quantification of mismatch-associated activity changes in many different sequence contexts. Cutting frequency determination (CFD) score, a leading method used to predict off-target activity, was trained on an average of 152 single-mismatched gRNAs per target. CHANCE-seq experiments quantify on average 1 millionmismatched sites per target, an increase of more than 6000-fold, across targets with mismatches ranging up to 7 and beyond. CFD score relies on a penalty matrix that varies by position and mismatch identity, but assumes penalties are fully independent. The penalty for a mismatch at a given position does not change, regardless of whether there are neighboring mismatches. For example, for every type of mismatch at every protospacer position, CFD predicts that an additional position 16 G^A or position 1 G^T mismatch will have no impact on activity (FIG.
[0627] 5M). In contrast, for any mismatch, CHANCE-seq activity values dynamically change based on position and mismatch identity of contextual mismatches (FIG. 5K). For the same mismatches that CFD predicts has no effect, activity changes were measure at many mismatch combinations (FIG. 5K).
[0628] Quantifying mismatch activity in diverse sequence contexts enables identification of synergistic versus neutral mismatch combinations. Many mismatches, particularly those with modest effects, can exert synergistically greater effects in combination (FIG. 50). For example, a mismatch can have strongly synergistic effects witha nearby mismatch, but largely neutral effects with a distant mismatch (FIGs. 5P-5Q).
[0629] Discussion
[0630] Presented herein is the first experimental, population-scale view of the impact of genetic variation on editing activity and suggests that nearly every cellular off-target site for a given CRISPR gRNA target may overlap with variation in some individuals, and that these overlapping variants frequently alter unintended Cas9 off-target activity. In a therapeutic context, of greatest concern would be the subset of variants that increase off-target editing and safety risks to patients54,55(although it should be noted only some unintended off-target sites may have an adverse biological consequence). For biological research, genetic variation between the genomes of donors of cell lines or primary cells may confound genome editing outcomes. The estimate of frequent variant-induced effects on editing described herein suggests that they should be carefully accounted for in the early stages of therapeutic lead target discovery or basic research, in a population- specific manner.
[0631] GUIDE-seq-2 can simplify the cellular off-target analysis of genome editing for many labs, and enables large-scale experimental design as herein as well as evaluating large number of gRNA targets. This novel and streamlined GUIDE-seq-2 approach enabled for the efficient survey 97% of common genetic variants found with the 1,000 Genomes Project resulting in 21 high-confidence variants and 13 additional variants that had significant effects when measured by GUIDE-seq-2 or targeted sequencing. However, developing general understanding of how genetic variants influence off-target activity requires high-diversity datasets with even representation of possiblemismatches. This gap motivated the exploration of alternatives solutions that could not be found by genomic off-target profiling.
[0632] Massively parallel biochemical profiling with CHANCE-seq has advantages over existing approaches. In contrast to biochemical methods like CHANGE-seq and Digenome-seq which characterize off-target activity in genomic DNA, CHANCE-seq interrogates large, highly diverse, partially randomized target libraries. These libraries are intrinsically better suited for understanding how different mismatches may affect CRISPR-Cas9 activity, because they quantify the impact of every mismatch in every possible position, embedded within millions of diverse sequence contexts.
[0633] References
[0634] 1. Umov, F. D. et al. Highly efficient endogenous human gene correction using designed zinc-finger nucleases. Nature 435, 646-651 (2005).
[0635] 2. Anzalone, A. V., Koblan, L. W. & Liu, D. R. Genome editing with CRISPR-Cas nucleases, base editors, transposases and prime editors. Nat. BiotechnoL 38, 824-844 (2020). 3. Mali, P. et al. RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013).
[0636] 4. Cong, L. et al. Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013).
[0637] 5. Frangoul, H. et al. Exagamglogene Autotemcel for Severe Sickle Cell Disease. N. Engl. J. Med. 390, 1649-1662 (2024).
[0638] 6. Scott, D. A. & Zhang, F. Implications of human genetic variation in CRISPR-based therapeutic genome editing. Nat. Med. 23, 1095-1101 (2017).
[0639] 7. Lessard, S. et al. Human genetic variation alters CRISPR-Cas9 on- and off-targeting specificity at therapeutically implicated loci. Proc. Natl. Acad. Sci. U. S. A. 114, El 1257-E11266 (2017).
[0640] 8. Cancellieri, S. et al. Human genetic diversity alters off-target outcomes of therapeutic gene editing. Nat. Genet. 55, 34-43 (2023).
[0641] 9. Tsai, S. Q. et al. GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases. Nat. BiotechnoL 33, 187-197 (2015).
[0642] 10. Sheridan, C. The world’s first CRISPR therapy is approved: who will receive it? Nat. BiotechnoL 42, 3-4 (2023).11. Mullard, A. 2023 FDA approvals. Nat. Rev. Drug Discov. (2024) doi:10.1038 / 41573-024-00001-x.
[0643] 12. Schambach, A. et al. A new age of precision gene therapy. Lancet 403, 568-582 (2024).
[0644] 13. Fu, Y. et al. High-frequency off-target mutagenesis induced by CRISPR-Cas nucleases in human cells. Nat. Biotechnol. 31, 822-826 (2013).
[0645] 14. Yin, J. et al. Cas9 exo-endonuclease eliminates chromosomal translocations during genome editing. Nat. Commun. 13, 1204 (2022).
[0646] 15. Frock, R. L. et al. Genome- wide detection of DNA double- stranded breaks induced by engineered nucleases. Nat. Biotechnol. 33, 179-186 (2014).
[0647] 16. Center for Biologies Evaluation & Research. Human Gene Therapy Products Incorporating Human Genome Editing. U.S. Food and Drug Administration https: / / www.fda.gov / regulatory-information / search-fda-guidance-documents / human-gene-therapy-products-incorporating-human-genome-editing.
[0648] 17. Hacein-Bey-Abina, S. et al. LM02-associated clonal T cell proliferation in two patients after gene therapy for SCID-X1. Science 302, 415-419 (2003).
[0649] 18. McCormack, M. P. & Rabbitts, T. H. Activation of the T-cell oncogene LMO2 after gene therapy for X-linked severe combined immunodeficiency. N. Engl. J. Med. 350, 913-922 (2004).
[0650] 19. Duncan, C. N. et al. Hematologic cancer after gene therapy for cerebral adrenoleukodystrophy. N. Engl. J. Med. 391, 1287-1301 (2024).
[0651] 20. 1000 Genomes Project Consortium et al. A global reference for human genetic variation. Nature 526, 68-74 (2015).
[0652] 21. Misek, S. A. et al. Germline variation contributes to false negatives in CRISPR-based experiments with varying burden across ancestries. bioRxiv 2022.11.18.517155 (2022) doi:10.1101 / 2022.11.18.517155.
[0653] 22. Lazzarotto, C. R. et al. CHANGE-seq reveals genetic and epigenetic effects on CRISPR-Cas9 genome- wide activity. Nat. Biotechnol. 38, 1317-1327 (2020).
[0654] 23. Tsai, S. Q. et al. CIRCLE-seq: a highly sensitive in vitro screen for genome-wide CRISPR-Cas9 nuclease off-targets. Nat. Methods 14, 607-614 (2017).
[0655] 24. Kim, D. et al. Digenome-seq: genome- wide profiling of CRISPR-Cas9 off-target effects in human cells. Nat. Methods 12, 237-43, 1 p following 243 (2015).
[0656] 25. Wienert, B. et al. Unbiased detection of CRISPR off-targets in vivo using DISCOVER-Seq. Science 364, 286-289 (2019).26. Cameron, P. et al. Mapping the genomic landscape of CRISPR-Cas9 cleavage. Nat. Methods 14, 600-606 (2017).
[0657] 27. Malinin, N. L. et al. Defining genome-wide CRISPR-Cas genome-editing nuclease activity with GUIDE-seq. Nat. Protoc. 16, 5592-5615 (2021).
[0658] 28. Yen, A. et al. Specificity of CRISPR-Cas9 editing in exagamglogene autotemcel. N. Engl. J. Med. 390, 1723-1725 (2024).
[0659] 29. Yang, L. et al. Targeted and genome- wide sequencing reveal single nucleotide variations impacting specificity of Cas9 in human stem cells. Nat. Commun. 5, 1-6 (2014).
[0660] 30. Fatumo, S. et al. A roadmap to increase diversity in genomic studies. Nat. Med. 28, 243-250 (2022).
[0661] 31. Byrska-Bishop, M. et al. High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell 185, 3426-3440. el9 (2022).
[0662] 32. Adams, G. S., Converse, B. A., Hales, A. H. & Klotz, L. E. People systematically overlook subtractive changes. Nature 592, 258-261 (2021).
[0663] 33. Berg, D. E., Davies, J., Allet, B. & Rochaix, J. D. Transposition of R factor genes to bacteriophage lambda. Proc. Natl. Acad. Sci. U. S. A. 72, 3628-3632 (1975).
[0664] 34. Reznikoff, W. S. Transposon Tn5. Annu. Rev. Genet. 42, 269-286 (2008).
[0665] 35. Adey, A. et al. Rapid, low-input, low-bias construction of shotgun fragment libraries by high-density in vitro transposition. Genome Biol. 11, R119 (2010).
[0666] 36. Giannoukos, G. et al. UDiTaS™, a genome editing detection method for indels and genome rearrangements. BMC Genomics 19, 212 (2018).
[0667] 37. Nobles, C. L. et al. iGUIDE: an improved pipeline for analyzing CRISPR cleavage specificity. Genome Biol. 20, 14 (2019).
[0668] 38. Liang, S.-Q. et al. Genome-wide detection of CRISPR editing in vivo using GUIDE-tag. Nat. Commun. 13, 437 (2022).
[0669] 39. Hosomichi, K., Mitsunaga, S., Nagasaki, H. & Inoue, I. A Bead-based Normalization for Uniform Sequencing depth (BeNUS) protocol for multi-samples sequencing exemplified by HLA-B. BMC Genomics 15, 645 (2014).
[0670] 40. Zook, J. M. et al. An open resource for accurately benchmarking small variant and reference calls. Nat. BiotechnoL 37, 561-566 (2019).
[0671] 41. Karczewski, K. J. et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 581, 434-443 (2020).
[0672] 42. DePristo, M. A. et al. A framework for variation discovery and genotyping using nextgeneration DNA sequencing data. Nat. Genet. 43, 491-498 (2011).43. Jinek, M. et al. A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science 337, 816-821 (2012).
[0673] 44. Pattanayak, V. et al. High-throughput profiling of off-target DNA cleavage reveals RNA-programmed Cas9 nuclease specificity. Nat. Biotechnol. 31, 839-843 (2013).
[0674] 45. Petri, K. et al. Global-scale CRISPR gene editor specificity profiling by ONE-seq identifies population-specific, variant off-target effects. bioRxiv 2021.04.05.438458 (2021) doi: 10.1101 / 2021.04.05.438458.
[0675] 46. Jones, S. K. et al. Massively parallel kinetic profiling of natural and engineered CRISPR nucleases. Nat. Biotechnol. 39, 84-93 (2020).
[0676] 47. Pacesa, M. et al. Structural basis for Cas9 off-target activity. Cell 185, 4067-408 l.e21 (2022).
[0677] 48. Shrikumar, A., Greenside, P. & Kundaje, A. Learning important features through propagating activation differences. ICML 70, 3145-3153 (2017).
[0678] 49. Sternberg, S. H., LaFrance, B., Kaplan, M. & Doudna, J. A. Conformational control of DNA target cleavage by CRISPR-Cas9. Nature 527, 110-113 (2015).
[0679] 50. Lin, J., Zhang, Z., Zhang, S., Chen, J. & Wong, K.-C. CRISPR-net: A recurrent convolutional network quantifies CRISPR off-target activities with mismatches and indels. Adv. Sci. (Weinh.) 7, 1903562 (2020).
[0680] 51. Alkan, F., Wenzel, A., Anthon, C., Havgaard, J. H. & Gorodkin, J. CRISPR-Cas9 off-targeting assessment with nucleic acid duplex energy parameters. Genome Biol. 19, 177 (2018).
[0681] 52. Doench, J. G. et al. Optimized sgRNA design to maximize activity and minimize off-target effects of CRISPR-Cas9. Nat. Biotechnol. 34, 184-191 (2016).
[0682] 53. Chen, Q. et al. Genome- wide CRISPR off-target prediction and optimization using RNA-DN A interaction fingerprints. Nat. Commun. 14, 7521 (2023).
[0683] 54. Cheng, Y. & Tsai, S. Q. Illuminating the genome-wide activity of genome editors for safe and effective therapeutics. Genome Biol. 19, 226 (2018).
[0684] 55. Saha, K. Accounting for diversity in the design of CRISPR-based therapeutic genome editing. Nat. Genet. 55, 6-7 (2023).
[0685] 56. Komor, A. C., Kim, Y. B., Packer, M. S., Zuris, J. A. & Liu, D. R. Programmable editing of a target base in genomic DNA without double- stranded DNA cleavage. Nature 533, 420-424 (2016).
[0686] 57. Gaudelli, N. M. et al. Programmable base editing of A»T to G*C in genomic DNA without DNA cleavage. Nature 551, 464-471 (2017).58. Anzalone, A. V. et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149-157 (2019).
Claims
CLAIMSWhat is claimed is:
1. A method for determining an activity and / or specificity of a genome editor for a target polynucleotide, comprising:(a) obtaining a library comprising circular DNAs, each circular DNA comprising a randomized target polynucleotide;(b) cleaving randomized target polynucleotides of circular DNAs of the library using the genome editor, thereby producing linearized DNAs; and(c) determining the activity and / or specificity of the genome editor for the target polynucleotide using the cleaved randomized target polynucleotides of the linearized DNAs.
2. The method of claim 1, wherein the genome editor comprises a base editor, an RNA-guided nuclease (RGN) complex, a zinc-finger nuclease (ZNF), a transcription activator-like effector nuclease (TALEN), a CRISPR-associated transpose, a prime editor, or a meganuclease.
3. The method of claim 2, wherein the RGN complex comprises a CRISPR nuclease and a guide RNA.
4. The method of claim 3, wherein the CRISPR nuclease is selected from the group consisting of Cas3, Cas8a, Cas5, Cas8b, Cas8c, CaslOd, Csel, Cse2, Csyl, Csy2, Csy3, GSU0054, CaslO, Csm2, Cmr5, CaslO, Csxll, CsxlO, Csfl, Cas9, Csn2, Cas4, Casl2, Casl2a (Cpfl), Casl2b (C2cl), Casl2c (C2c3), Casl2d (CasY), Casl2e (CasX), Casl2f (Casl4, C2cl0), Casl2g, Casl2h, Casl2i, Casl2k (C2c5), C2c4, C2c8, and C2c9.
5. The method of claim 1, wherein cleaving the randomized target polynucleotides of the circular DNAs using the genome editor comprises cleaving once.
6. The method of claim 4 or claim 5, wherein the CRISPR nuclease is Cas9.
7. The method of any one of claims 3-6, wherein the randomized target polynucleotide comprises a portion that is complementary to a homology region of the guide RNA and a protospacer adjacent motif (PAM).
8. The method of claim 2, wherein the base editor is an adenosine base editor (ABE) or a cytosine base editor (CBE) and cleaving comprises:nicking the randomized target polynucleotides of circular DNAs using an ABE or CBE; andcleaving the nicked randomized target polynucleotide using an nuclease optionally EndoV (ABE) or USER (CBE) or EndoQ (ABE or CBE).
9. The method of any one of claims 1-8, wherein the library further comprises circular DNAs comprising target polynucleotides; and cleaving further comprises cleaving target polynucleotides of the circular DNAs thereby producing target polynucleotide linearized DNAs.
10. The method of claim 9, comprising determining an activity of the genome editor for the target polynucleotide using the target polynucleotide linearized DNAs.
11. The method of any one of claims 1-10, comprising determining an activity of the genome editor for a randomized target polynucleotide using the cleaved randomized target polynucleotides of the linearized DNAs.
12. The method of claim 11, comprising comparing the activity of the genome editor for the target polynucleotide to the activity of the genome editor for the randomized target polynucleotide.
13. The method of claim 12, wherein the target polynucleotide and the randomized target polynucleotide differ from one another by 1-6 nucleotides.
14. The method of any one of claims 1-13, wherein the randomized target polynucleotides are partially randomized target polynucleotides.
15. The method of claim 14, wherein the partially randomized target polynucleotides comprise fewer polynucleotides that will not be cleaved by the genome editor compared to completely randomized target polynucleotides.
16. The method of claim 15, wherein the partially randomized target polynucleotides are generated by:(a) at each nucleotide of a sequence of the target polynucleotide, changing the nucleotide to a different nucleotide in 10%-50% of the partially randomized target polynucleotides; and (b) synthesizing the sequences of the partially randomized target polynucleotides.
17. The method of claim 16, wherein synthesizing the sequences of the partially randomized target polynucleotides comprises synthesizing using polymerase chain reaction with a degenerate primer.
18. The method of any one of claims 16 or 17, wherein synthesizing the sequences of the partially randomized target polynucleotides comprises synthesizing using partially degenerate oligonucleotide synthesis.
19. The method of any one of claims 1-18, wherein each partially randomized target polynucleotide comprises no more than 10 substitutions relative to the target polynucleotide.
20. The method of any one of claims 1-18, wherein each partially randomized target polynucleotide comprises no more than 5 substitutions relative to the target polynucleotide.
21. The method of claim any one of claims 1-20, wherein each circular DNA of the library comprises a unique molecule identifier (UMI).
22. The method of any one of claims 1-21, wherein identifying the randomized target polynucleotides in the cleaved linearized DNAs comprises sequencing the cleaved linearized DNAs.
23. The method of claim 22, wherein sequencing comprises next-generation sequencing.
24. The method of any one of claims 1-23, wherein the target polynucleotide comprises a portion of a gene.
25. The method of claim 24, wherein the portion of the gene comprises a diseases associated mutation.
26. The method of any one of claims 1-25, further comprising generating the library, the generating comprising:(a) amplifying a portion of a template DNA using: a plurality of first primers that each comprise (i) a polynucleotide that is complementary to a first portion of the template DNA, (ii) a partially randomized target polynucleotide, (iii) a unique molecular index, and (iv) a universal primer sequence; and a second primer that is complementary to a second portion of the template DNA, to produce linear precursor polynucleotides comprising a portion of the template DNA and different partially randomized target polynucleotides; and(b) circularizing the linear precursor polynucleotides to produce the circular DNAs.
27. The method of claim 26, wherein generating the library comprises introducing a restriction enzyme site into the linear precursor polynucleotides.
28. The method of claim 26 or claim 27, wherein generating the library comprises introducing an Uracil DNA glycosylase and a DNA glycosylase-lyase Endonuclease VIII (USER) enzyme substrate (uracil) via PCR onto the ends of the linear precursor polynucleotides and cleaving the USER enzyme substrate using a USER enzyme prior to circularizing the linear precursor polynucleotides to form the circular DNAs.
29. The method of any one of claims 26-28, further comprising (c) degrading residual linear DNA.
30. The method any one of claims 1-29, wherein obtaining the library comprises bottlenecking the number of different randomized target polynucleotides in the library.
31. The method of claim 30, wherein bottlenecking comprises reducing the number of different randomized target polynucleotide in the library to less than 10 million different randomized target polynucleotides.
32. The method of claim 30, wherein bottlenecking reduces the number of different randomized target polynucleotide in the library to less than 3 million different randomized target polynucleotides.
33. The method of any one of claims 9-32, wherein determining the activity of the genome editor for the target polynucleotide comprises determining the efficiency with which the target polynucleotide is cleaved using the cleaved target polynucleotide linearized DNAs.
34. The method of any one of claims 11-33, wherein determining the activity of the genome editor for a randomized target polynucleotide comprises determining the efficiency with which the randomized target polynucleotide is cleaved using the cleaved randomized target polynucleotides of the linearized DNAs.
35. The method of any one of claims 1-34, wherein determining the specificity of the genome editor comprises determining an activity of the genome editor for one or more of the randomized target polynucleotides, optionally wherein determining the activity of the genome editor for one or more of the randomized target polynucleotides comprises determining the relative activity of the genome editor for one or more of the randomized target polynucleotides.
36. The method of any one of claims 1-35, wherein determining the specificity of the genome editor comprises determining the efficiency with which a randomized target polynucleotide of the circular DNAs is cleaved and / or the number of different randomized target polynucleotides of the circular DNAs that are cleaved.
37. The method of any one of claims 1-36, further comprising making a copy of the library.
38. The method of claim 37, further comprising digesting circular DNAs of the copy of the library using a restriction enzyme to produce control linearized DNAs.
39. The method of claim 38, wherein determining the activity of the genome editor for the target polynucleotide comprises determining a relative amount of the cleaved target polynucleotides of the linearized DNAs and the control linearized DNAs comprising the target polynucleotide.
40. The method of claim 38 or claim 39, wherein determining the activity of the genome editor for a randomized target polynucleotide comprises determining a relative amount of the cleaved randomized target polynucleotide of the linearized DNAs and the control linearized DNAs comprising the randomized target polynucleotide.
41. The method of any one of claims 38-40, wherein determining the specificity of the genome editor for the target polynucleotide comprises comparing the activity of the genome editor for the target polynucleotide and the activity of the genome editor for one or more of the randomized target polynucleotides.
42. The method of any one of claims 1-41, wherein each circular DNA of the library comprises only one randomized target polynucleotide or only one target polynucleotide.
43. The method of any one of claims 1-42, further comprising:determining an activity and / or specificity of a genome editor for a second target polynucleotide, the method comprising:(a) obtaining a library comprising circular DNAs, each circular DNA comprising a second randomized target polynucleotide;(b) cleaving second randomized target polynucleotides of circular DNAs of the library using the genome editor, thereby producing different linearized DNAs; and(c) determining the activity and / or specificity of the genome editor for the second target polynucleotide using the cleaved second randomized target polynucleotides of the linearized DNAs.
44. The method of claim 43, further comprising comparing the activity and / or specificity for the genome editor for the target polynucleotide with the activity and / or specificity of the genome editor for the second target polynucleotide, optionally wherein comparing comprises comparing using a ratio or difference.
45. A method for determining an activity and / or specificity of a genome editor for a target polynucleotide, comprising:(a) obtaining a library comprising circular DNAs, each circular DNA comprising a target polynucleotide or partially randomized target polynucleotide and a restriction site, wherein obtaining the library comprises obtaining a bottlenecked library that comprises no more than 10 million different partially randomized target polynucleotides;(b) obtaining a copy of the library;(c) cleaving partially randomized target polynucleotides of circular DNAs of the library using the genome editor, thereby producing cleaved linearized DNAs;(d) digesting circular DNAs of the copy of the library using a restriction enzyme to produce control linearized DNAs;and(e) determining the activity and / or specificity of the genome editor for the target polynucleotide using the cleaved linearized DNAs and the control linearized DNAs.
46. The method of claim 45, wherein the partially randomized target polynucleotides are generated by:(a) at each nucleotide of a sequence of the target polynucleotide, changing the nucleotide to a different nucleotide in 3%-50% of the partially randomized target polynucleotides; and (b) synthesizing the sequences via PCR of the partially randomized target polynucleotides.
47. A method for determining an activity of a genome editor for one or more randomized target polynucleotides, comprising:(a) obtaining a library comprising circular DNAs, each circular DNA comprising a randomized target polynucleotide;(b) cleaving randomized target polynucleotides of circular DNAs of the library using the genome editor, thereby producing linearized DNAs; and(c) determining the activity of the genome editor for the one or more randomized target polynucleotides using the cleaved randomized target polynucleotides of the linearized DNAs.
48. A method for determining an activity of a genome editor for one or more partially randomized target polynucleotides, comprising:(a) obtaining a library comprising circular DNAs, each circular DNA comprising a partially randomized target polynucleotide and a restriction site, wherein obtaining the library comprises obtaining a bottlenecked library that comprises no more than 10 million different partially randomized target polynucleotides;(b) obtaining a copy of the library;(c) cleaving partially randomized target polynucleotides of circular DNAs of the library using the genome editor, thereby producing cleaved linearized DNAs;(d) digesting circular DNAs of the copy of the library using a restriction enzyme to produce control linearized DNAs;and(e) determining the activity of the genome editor for the one or more target polynucleotides or randomized target polynucleotides using the cleaved linearized DNAs and the control linearized DNAs.
49. A method for determining an activity and / or specificity of a guide RNA (gRNA) for a target polynucleotide, comprising:(a) obtaining a library comprising circular DNAs, each circular DNA comprising a randomized target polynucleotide;(b) cleaving randomized target polynucleotides of circular DNAs of the library using an RNA-guided nuclease and the gRNA, thereby producing linearized DNAs; and(c) determining the activity and / or specificity of the gRNA for the target polynucleotide using the cleaved randomized target polynucleotides of the linearized DNAs.
50. The method of any one of claims 1-49, further comprising determining a cleavage site of the genome editor in one or more randomized target polynucleotides of circular DNAs of the library or in a target polynucleotide of circular DNAs of the library.
51. The method of claim 50, further comprising determining whether the cleavage site is a staggered cleavage site or a blunt end cleavage site.
52. The method of claim 51, further comprising determining the overhang length of the staggered cleavage site.
53. The method of any one of claims 47-52, further comprising determining an activity of the genome editor for at least two randomized target polynucleotides.
54. The method of any one of claims 47-53, comprising determining whether the cleavage site is a staggered cleavage site or a blunt end cleavage site for at least 10, at least 50, at least 100, at least 1000, at least 10,000, or at least 100,000 randomized target polynucleotides.
55. The method of claim 54, comprising determining an amount of randomized target polynucleotide whose cleavage results in a staggered cleavage site, and / or determining an amount of randomized target polynucleotide whose cleavage results in a blunt end cleavage site.
56. The method of claim 55, comprising determining a ratio between the amount of randomized target polynucleotide whose cleavage results in a staggered cleavage site and the amount of randomized target polynucleotide whose cleavage results in a blunt end cleavage site57. The method of any one of claims 47-56, further comprising comparing the activity of the genome editor for the at least two randomized target polynucleotides.
58. The method of any one of claim 53-57, wherein the sequences of the least two randomized target polynucleotides are within a hamming distance of 6 or less.
59. The method of claim 58, wherein the sequences of the least two randomized target polynucleotides are within a hamming distance of 2 or less.
60. The method of claim 58, wherein the sequences of the least two randomized target polynucleotides differ by a hamming distance of 1 or less.
61. The method of any one of claims 53-60, further comprising determining an effect of a specific mutation on the activity of the genome editor for the at least two randomized target polynucleotide using the activity of the genome editor for the at least two randomized target polynucleotides.
62. The method of any one of claims 1-61, wherein determining activity comprises determining relative activity.
63. A method for determining an activity and / or specificity of a non-nicking base editor for a target polynucleotide, comprising:(a) obtaining a library comprising circular DNAs, each circular DNA comprising a randomized target polynucleotide;(b) editing a strand of the randomized target polynucleotides of circular DNAs of the library using the genome editor, thereby producing a mismatched strands;(c) cleaving the mismatched strands using a mismatch endonuclease I or nuclease that recognizes mismatched DNA bases, thereby producing linearized DNAs; and(d) determining the activity and / or specificity of the genome editor for the target polynucleotide using the cleaved randomized target polynucleotides of the linearized DNAs.