Indel frequency prediction
A data analysis pipeline for quantifying indels in polynucleotide sequences addresses the underestimation issue by aligning sequencing reads to a reference sequence and applying quality and alignment scores, enhancing the accuracy of indel frequency estimation for both known and unknown nucleases.
Patent Information
- Application Number
- JP2025533054
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-05-02
- Filing Date
- 2023-12-05
- Publication Date
- 2026-01-21
AI Technical Summary
Existing computational methods for quantifying insertions and deletions (indels) in polynucleotide sequences cleaved by nucleic acid-guided nucleases, such as CRISPR-associated proteins, often underestimate the frequency of indels, particularly for novel or poorly characterized nucleases with unknown cleavage sites, due to the assumption of known cleavage sites.
A data analysis pipeline is developed to quantify indels by aligning sequencing reads from targeted amplicon sequencing (TAS) to a reference sequence, defining a window based on the nucleic acid guide sequence, and estimating indel frequency using parameters like read quality scores and alignment scores, applicable to both characterized and uncharacterized nucleases.
Accurately quantifies indels in polynucleotide sequences, including those caused by novel nucleases, improving the estimation of editing efficiency and effectiveness of nucleic acid-guided nucleases.
Smart Images

Figure 2026502071000012 
Figure 2026502071000013 
Figure 2026502071000014
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 430,827, filed December 7, 2022, and European Patent Application No. 23315141.4, filed May 2, 2023, the contents of which are incorporated herein by reference in their entireties.
[0002] The present disclosure relates to quantification of insertions and / or deletions near a cleavage site within a polynucleotide target sequence. [Background technology]
[0003] Nucleic acid-guided nucleases can be used to edit polynucleotide sequences, such as the genome of an organism, at target locations with high precision. Nucleases are enzymes that can cleave phosphodiester bonds between nucleotides in nucleic acids. Genome editing methods involve using CRISPR (clustered regularly interspaced short palindromic repeats)-associated proteins or similar nucleic acid-guided nucleases to induce DNA double-strand breaks (DSBs) at predictable genomic locations relative to user-specified target sequences. DNA DSBs are repaired by intracellular mechanisms, such as non-homologous end joining (NHEJ). The repair process can result in sequence variants, including, for example, insertions and deletions (indels). Quantifying the frequency of indels at the target cleavage site within an edited cell population is important for evaluating the effectiveness of a nucleic acid-guided nuclease editing system.
[0004] Currently, a variety of nucleic acid-guided nucleases are well known, and additional naturally occurring and engineered nucleases are being discovered and characterized. Novel nucleases may not be fully characterized, i.e., their cleavage site and editing window for a user-specified target sequence are unknown. Existing indel quantification data analysis pipelines for next-generation sequencing (NGS) rely on the assumption that the cleavage site and editing window are established, which is not true for novel, poorly characterized nucleases. As a result, existing computational tools may undercount indels at the target cleavage site in edited cell populations. Few computational methods have been developed to date to address this need. Summary of the Invention [Means for solving the problem]
[0005] The present disclosure is based in part on the discovery that a data analysis pipeline can be configured to quantify insertions and deletions in a target polynucleotide sequence cleaved by a nucleic acid-guided nuclease, even when the cleavage site of the nuclease relative to a user-specified target sequence is unknown. The methods disclosed herein include computational pipelines for indel quantification of polynucleotide sequences cleaved by nucleic acid-guided nucleases, including uncharacterized nucleases. The methods disclosed herein include parameters for alignment of next-generation sequencing (NGS) reads from targeted amplicon sequencing (TAS) of a nuclease-edited sample or a control sample that have been obtained and aligned to a reference amplicon sequence. Exemplary experimental and computational data are disclosed herein that demonstrate that the methods can be used to accurately quantify indels from novel, uncharacterized nucleases, for example, nucleases with unknown cleavage sites relative to a user-specified target sequence.
[0006] In some embodiments, provided herein are methods for quantifying insertions and / or deletions in a polynucleotide sequence caused by cleavage of a target polynucleotide sequence by a nucleic acid-guided nuclease. In some embodiments, the method includes receiving, in a computer system, sample sequence data including a plurality of sequencing reads; filtering the plurality of sequencing reads; aligning the plurality of sequencing reads to a reference sequence; defining a window based on a sequence position within the nucleic acid guide sequence and the position of an end of the nucleic acid guide sequence; determining the number of sequencing reads containing insertions or deletions within the window relative to the reference sequence based on the alignment of each sequencing read of the plurality of sequencing reads to the reference sequence; and estimating the amount of insertions and / or deletions in the polynucleotide sequence mediated by the nucleotide-directed nuclease based on the number of sequencing reads containing insertions or deletions. In some embodiments, the nucleic acid-guided nuclease is a class 2 nuclease. In some embodiments, the nucleic acid-guided nuclease is a type II nuclease. In some embodiments, the nucleic acid-guided nuclease is SpCas9. In some embodiments, the nucleic acid-guided nuclease is AsCasl2a. In some embodiments, the nucleic acid guided nuclease is a V-type nuclease. In some embodiments, the nucleic acid guide is a guide RNA (gRNA). In some embodiments, the nucleic acid guide is a single guide RNA (sgRNA).
[0007] In some embodiments, the center of the window is located at the center of the nucleic acid guide sequence. In some embodiments, the center of the window is located at the site where the nuclease cleaved the polynucleotide sequence. In some embodiments, the length of the window is equal to the length of the nucleic acid guide sequence. In some embodiments, the 5' end of the window extends 50 base pairs 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 50 base pairs 3' to the 3' end of the nucleic acid guide. In some embodiments, the 5' end of the window extends 40 base pairs 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 40 base pairs 3' to the 3' end of the nucleic acid guide. In some embodiments, the 5' end of the window extends 30 base pairs 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 30 base pairs 3' to the 3' end of the nucleic acid guide. In some embodiments, the 5' end of the window extends 20 base pairs 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 20 base pairs 3' to the 3' end of the nucleic acid guide. In some embodiments, the 5' end of the window extends 10 base pairs 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 10 base pairs 3' to the 3' end of the nucleic acid guide. In some embodiments, the 5' end of the window extends 5 base pairs 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 5 base pairs 3' to the 3' end of the nucleic acid guide.
[0008] In some embodiments, the method further comprises trimming adapter sequences from the plurality of sequencing reads. In some embodiments, the plurality of sequencing reads are generated from targeted amplicon sequencing. In some embodiments, the plurality of sequencing reads are generated from targeted amplicon sequencing of DNA isolated from cells edited by nucleic acid-guided nucleases. In some embodiments, the plurality of sequencing reads are paired-end sequencing reads. In some embodiments, the method further comprises read-merging the paired-end sequencing reads to generate a single read for alignment to a reference sequence. In some embodiments, the read-merging of the paired-end reads further comprises applying a minimum paired-end read overlap score to the plurality of sequencing reads. In some embodiments, the minimum paired-end read overlap score is 0, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100. In some embodiments, the minimum paired-end read overlap score is 10. In some embodiments, the read-merging of the paired-end reads further comprises applying a maximum paired-end read overlap score to the plurality of sequencing reads. In some embodiments, the maximum paired-end read overlap score is 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, or 300. In some embodiments, the maximum paired-end read overlap score is 100. In some embodiments, the filtering step further comprises applying a minimum average read quality score to the plurality of sequencing reads. In some embodiments, the minimum average read quality score is 0, 5, 10, 15, 20, 25, 30, 35, or 40. In some embodiments, the filtering step further comprises applying a minimum single base pair score to the plurality of sequencing reads. In some embodiments, the minimum single base pair score is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20.In some embodiments, the aligning step further comprises applying an amplicon minimum alignment score to the plurality of sequencing reads. In some embodiments, the amplicon minimum alignment score is 0, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100.
[0009] In some embodiments, the nucleic acid-guided nuclease is a FokI nuclease. In some embodiments, the FokI nuclease is fused to a transcription activator-like (TAL) protein. In some embodiments, the nucleic acid-guided nuclease is a zinc finger nuclease.
[0010] In some embodiments, described herein is a computer program product tangibly embodied on a computer-readable medium, the computer program product comprising instructions that, when executed by one or more processors, are configured to: receive sample sequence data comprising a plurality of sequencing reads; filter the plurality of sequencing reads; align the plurality of sequencing reads to a reference sequence; define a window based on a sequence position within the nucleic acid guide sequence and a position of an end of the nucleic acid guide sequence; determine a number of sequencing reads that contain insertions or deletions within the window relative to the reference sequence based on the alignment of each sequencing read of the plurality of sequencing reads to the reference sequence; and estimate an amount of insertions and / or deletions in the polynucleotide sequence mediated by the nucleotide-directed nuclease based on the number of sequencing reads that contain insertions or deletions.
[0011] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Methods and materials are described herein for use in the present invention; other suitable methods and materials known in the art can also be used. The materials, methods, and examples are illustrative only and are not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control.
[0012] Other features and advantages of the invention will become apparent from the following detailed description and drawings, and from the claims. [Brief explanation of the drawings]
[0013] [Figure 1A] Schematic diagram of a Type II nuclease featuring an sgRNA interacting with a target DNA sequence, showing the PAM (NGG) sequence, guide sequence, target DNA sequence, and cleavage site. [Figure 1B] Schematic diagram of a Type V nuclease featuring a gRNA interacting with a target DNA sequence, showing the PAM (TTTN) sequence, guide sequence, target DNA sequence, and cleavage site. [Figure 2A]
[0023] Figure 1 shows the upstream molecular biology workflow of the methods disclosed herein. The steps include designing primers to amplify the editing target, isolating cellular DNA, PCR-amplifying the editing target with adapters, adding next-generation sequencing (NGS) adapters to the PCR products, pooling barcoded amplicons, bead-purifying the amplicons, spiking in diverse DNA, and performing NGS on a sequencing instrument. [Figure 2B] 1 is a diagram of the computational pipeline disclosed herein. The steps include quality control of received sequencing reads, alignment of sequencing reads to target sequences, and quantification of indels according to the methods disclosed herein. [Figure 3A]Schematic diagram of received sequencing reads for a nucleic acid-guided nuclease-edited sample, including multiple sequencing reads containing indels as a result of cleavage by a nucleic acid-guided nuclease. The nuclease is Cas9, a type II nuclease with a known cleavage site relative to the guide sequence. For each type of indel, the number of sequencing reads containing that indel is shown on the right. Previously known cleavage sites are indicated by rectangular boxes. [Figure 3B]
[0023] Figure 1 is a schematic diagram of received sequencing reads for a nucleic acid-guided nuclease-edited sample, including a subset of sequencing reads that contain indels as a result of cleavage by a nucleic acid-guided nuclease. The nuclease is a novel V-type nuclease. For each type of indel, the number of sequencing reads containing that indel is shown on the right. Some indels are missed by standard computational methods. [Figure 3C]
[0023] Figure 1 is a schematic diagram of received sequencing reads for a nucleic acid-guided nuclease-edited sample, including a subset of sequencing reads that contain indels as a result of cleavage by a nucleic acid-guided nuclease. The nuclease is a novel V-type nuclease. For each type of indel, the number of sequencing reads containing that indel is shown on the right. A custom data processing method captures all indels generated by cleavage by the nuclease. [Figure 4A] FIG. 1 is a plot comparing a standard data processing pipeline for the known nuclease Cas9 with the data processing pipeline disclosed herein for Cas9. [Figure 4B] FIG. 1 is a schematic diagram of a dilution experiment to validate the data processing pipeline disclosed herein, in which nucleic acid-guided nuclease-edited target sequences are serially diluted with unedited target sequences to assess the effectiveness of the data processing methods disclosed herein. [Figure 4C] FIG. 4C is a plot of the expected percentage of indels (x-axis) versus the observed percentage of indels (y-axis) for the experiment shown in the schematic diagram of FIG. 4B. [Figure 5]FIG. 1 is a diagram of computer system components that can be used to implement a computational pipeline for indel quantification of polynucleotide sequences cleaved by nucleic acid-guided nucleases, including uncharacterized nucleases. [Figure 6] Plot of the percentage of indels (y-axis) for two nucleases (x-axis): Cas9 (reference) and novel type V nuclease with increasing doses of the novel nuclease using two different guide sequences, A1 and F2. [Figure 7A] FIG. 1 is a schematic diagram of an experiment in which the frequency of 21 nucleotide insertions was quantified by the computational pipeline disclosed herein. [Figure 7B] 7B is a schematic diagram of received sequencing reads of the amplicon shown in FIG. 7A, including multiple sequencing reads, a subset of which contain a 21-nucleotide insertion in a spike-in dilution experiment. [Figure 7C] 7B is a schematic diagram of received sequencing reads of the amplicon shown in FIG. 7A, including multiple sequencing reads, a subset of which contain a 21-nucleotide insertion in a spike-in dilution experiment. [Figure 7D] 7B is a schematic diagram of received sequencing reads of the amplicon shown in FIG. 7A, including multiple sequencing reads, a subset of which contain a 21-nucleotide insertion in a spike-in dilution experiment. DETAILED DESCRIPTION OF THE INVENTION
[0014] definition As used herein, the terms "nucleic acid," "polynucleotide," and "oligonucleotide" are interchangeable and refer to deoxyribonucleotide or ribonucleotide polymers in either single- or double-stranded form, in linear or cyclic conformation. For purposes of this disclosure, these terms should not be construed as limiting with respect to the length of the polymer. These terms can encompass known analogues of natural nucleotides as well as nucleotides modified in the base, sugar, and / or phosphate moieties (e.g., phosphorothioate backbones). Generally, unless otherwise specified, analogues of a particular nucleotide have the same base-pairing specificity; i.e., an analogue of A will base pair with T.
[0015] As used herein, the terms "polypeptide," "peptide," and "protein" are used interchangeably to refer to a polymer of amino acid residues. The terms also apply to amino acid polymers in which one or more amino acids are chemical analogues or modified derivatives of a corresponding naturally occurring amino acid.
[0016] As used herein, "CRISPR" refers to either clustered regularly interspaced short palindromic repeats or a DNA locus that serves to guide CRISPR-associated proteins or similar nucleotide-directed nucleases. It also describes artificial, constructed, or selected systems that use these frameworks or proteins. CRISPR systems and associated proteins differ among the currently described type I, type II, and type III systems, although other similar systems may yet be described.
[0017] As used herein, "CRISPR system" collectively refers to the transcripts and other elements involved in directing the expression or activity of CRISPR-associated ("Cas") genes, including sequences encoding Cas genes, tracr (transactivating CRISPR) sequences (e.g., tracrRNA or active partial tracrRNA), tracr-mate sequences (including "direct repeats" and tracrRNA-processed partial direct repeats associated with endogenous CRISPR systems), guide sequences (also referred to as "spacers" associated with endogenous CRISPR systems), and other sequences and transcripts from the CRISPR locus. One or more tracr-mate sequences operably linked to a guide sequence (e.g., direct repeat-spacer-direct repeat) may also be referred to as a pre-processed precrRNA (pre-CRISPR RNA) or a nuclease-processed crRNA. CRISPR systems can also include modified, replaced, or engineered guide, tracr, or chimeric RNA sequences and the proteins they interact with (e.g., Briner, et al., Mal Cell 56(2)333-9(2014)). The methods disclosed herein can also be applied to other non-CRISPR nucleotide-directed nucleases.
[0018] As used herein, the term "guide sequence" refers to, for example, the portion of a guide RNA (gRNA) or single guide RNA (sgRNA) that confers specificity to a nucleic acid-guided nuclease for its target and mediates the formation of an RNA-DNA duplex between the target RNA and the target DNA sequence. For example, the targeting specificity of a CRISPR-Cas9 complex is determined by the sequence of approximately 20 nt at the 5' end of the gRNA. The length of a guide sequence is typically 17-24 bp. As used herein, the "center of a guide sequence" refers to the midpoint of the guide sequence. For example, if the guide sequence is 20 nucleotides long, the "center of a guide sequence" would be between nucleotides 10 and 11, numbered from the 5' end of the guide sequence to the 3' end of the guide sequence.
[0019] As used herein, the term "cleavage" or "splitting" of a nucleic acid refers to the cleavage of the covalent backbone of a nucleic acid molecule. Cleavage can be initiated by a variety of methods, including, but not limited to, enzymatic or chemical hydrolysis of a phosphodiester bond. Both single-strand and double-strand cleavage are possible, and double-strand cleavage can occur as a result of two different single-strand cleavage events. DNA cleavage can result in the generation of either blunt ends or staggered "sticky" ends. In certain embodiments, cleavage refers to a double-strand break between nucleic acids within a double-stranded DNA or RNA strand.
[0020] As used herein, the terms "genomic region" or "genomic segment," as used interchangeably herein, refer to a contiguous stretch of nucleotides within the genome of an organism. A genomic region can be as long as several kb (e.g., at least 5 kb, at least 10 kb, or at least 20 kb), up to or exceeding an entire chromosome.
[0021] In some cases, nucleotide sequences are provided using the letter conventions recommended by the International Union of Pure and Applied Chemistry (IUPAC), or a subset thereof. The IUPAC nucleotide code used herein includes: A = adenine, C = cytosine, G = guanine, T = thymine, U = uracil, R = A or G, Y = C or T, S = G or C, W = A or T, K = G or T, M = A or C, B = C or G or T, D = A or G or T, H = A or C or T, V = A or C or G, N = any base, "." or "-" = gap. In some embodiments, the set {A, C, G, T, U} is adenosine, cytidine, guanosine, thymidine, and uridine, respectively. In some embodiments, the set of letters is {A, C, G, T, U, I, X, W} for adenosine, cytidine, guanosine, thymidine, uridine, inosine, uridine, xanthosine, pseudouridine, respectively. In some embodiments, the set of letters is {A, C, G, T, U, I, X, P, R, Y, N} for adenosine, cytidine, guanosine, thymidine, uridine, inosine, uridine, xanthosine, pseudouridine, unspecified purine, unspecified pyrimidine, and unspecified nucleotide, respectively. Modified sequences, non-natural sequences, or sequences with modified linkages can be within a genomic sequence, a guide sequence, or a tracr sequence.
[0022] The percent (%) of nucleotide and / or amino acid sequence identity is understood as the percentage of nucleotides or amino acid residues that are identical to the nucleotides or amino acid residues in a candidate sequence compared to a reference sequence when the two sequences are aligned. To determine percent identity, the sequences are aligned and, if necessary, gaps are introduced to achieve the maximum percent identity. Sequence alignment procedures for determining percent identity are well known to those skilled in the art. Many publicly available computer software programs, such as BLAST, BLAST2, ALIGN2, or MEGALIGN (DNASTAR) software, can be used to align sequences. Those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithms required to achieve maximum alignment over the entire length of the sequences being compared. When sequences are aligned, the percent sequence identity of a given sequence A to, with, or relative to a given sequence B (which can alternatively be translated as a given sequence A having or containing a particular percent sequence identity to, with, or relative to a given sequence B) can be calculated as follows: Percent sequence identity = X / Y100, where X is the number of residues scored as identical matches by a sequence alignment program or algorithm aligning A and B, and Y is the total number of residues in B. If the length of sequence A is not equal to the length of sequence B, the percent sequence identity of A to B will not be equal to the percent sequence identity of B to A. Mismatches can similarly be defined as differences between the natural binding partners of nucleotides. The number, position, and type of mismatches can be calculated and used for identification or ranking purposes.
[0023] As used herein, "mutation" encompasses any change in a DNA, RNA, or protein sequence from a wild-type sequence or some other reference, including, but not limited to, a point mutation, transition, insertion, transversion, translocation, deletion, inversion, duplication, rearrangement, or combinations thereof.
[0024] As used herein, in relation to a polynucleotide sequence cleaved by an RNA-guided nuclease, the term "insertion" is used when the polynucleotide sequence has one or more extra bases compared to the polynucleotide sequence before cleavage by the RNA-guided nuclease. Similarly, in relation to a polynucleotide sequence cleaved by an RNA-guided nuclease, the term "deletion" is used when the polynucleotide sequence has one or more missing bases compared to the polynucleotide sequence before cleavage by the RNA-guided nuclease. In relation to a polynucleotide sequence cleaved by an RNA-guided nuclease, the term "indel" refers to either an insertion or a deletion. Cleavage by an RNA-guided nuclease can result in multiple indels, multiple insertions, multiple deletions, or a combination of one or more nucleotide insertions and one or more nucleotide deletions.
[0025] CRISPR CRISPR (clustered regularly interspaced short palindromic repeats) is an acronym for DNA loci containing multiple short direct repeats of a base sequence. Prokaryotic CRISPR / Cas systems have been adapted for use in eukaryotes as gene editors (silencing, enhancing, or altering specific genes) (see, e.g., Cong, Science, 15:339(6121):819-823(2013) and Jinek, et al., Science, 337(6096):816-21(2012)). By transfecting cells with the necessary elements, including cas genes and specifically designed CRISPRs, polynucleotide sequences can be cleaved and modified at virtually any desired location, for example, by specific targeting with guide RNAs that confer specificity to the nuclease. Numerous methods exist for expressing the guide strand or the Cas protein, including inducible expression of one or both. There are many methods for introducing guide strands and Cas proteins into cells, including viral transduction, injection or microinjection, nanoparticle or other delivery, protein uptake, RNA or DNA uptake, and uptake of a combination of protein and RNA or DNA. A combination of methods can also be used simultaneously or sequentially. Multiple deliveries of RNA, DNA, or protein can occur with or without further protein expression. Methods for preparing compositions for use in genome editing using the CRISPR / Cas system are described in detail in International Publication Nos. WO 2013 / 176772 and WO 2014 / 018423, which are specifically incorporated herein by reference in their entirety.
[0026] nuclease In some embodiments, a nuclease for use in the methods described herein is a Class 2 Cas nuclease. In some embodiments, the nuclease has double-stranded endonuclease activity. In some embodiments, the nuclease comprises a Cas nuclease, such as a Class 2 Cas nuclease (which can be, for example, a Type II, Type V, or Type VI Cas nuclease). Class 2 Cas nucleases include, for example, Cas9, Cpfl, C2cl, C2c2, and C2c3 proteins and modifications thereof.
[0027] Examples of Cas9 nucleases include those in the type II CRISPR systems of Streptococcus pyogenes (S. pyogenes), Staphylococcus aureus (S. aureus), and other prokaryotes (see, e.g., the list in the next paragraph), as well as modified (e.g., engineered or mutant) versions thereof. See, e.g., U.S. Patent Application Publication Nos. 2016 / 0312198A1 and 2016 / 0312199A1. Figure 1A shows a type II nuclease featuring an sgRNA that interacts with a target DNA sequence. The PAM (NGG) sequence, guide sequence, target DNA sequence, and cleavage site are shown. Other examples of Cas nucleases include the Csm or Cmr complexes of type III CRISPR systems or their Cas10, Csml, or Cmr2 subunits, and the Cascade complex of type I CRISPR systems or their Cas3 subunits. Figure 1B shows a Type V nuclease featuring a gRNA interacting with a target DNA sequence. The PAM (TTTN) sequence, guide sequence, target DNA sequence, and cleavage site are shown. In some embodiments, the Cas nuclease can be derived from a Type IIA, Type 11B, or Type IIC system. For a discussion of various CRISPR systems and Cas nucleases, see, e.g., Makarova et al., Nat. Rev. Microbiol. 9:467-477 (2011); Makarova et al., Nat. Rev. Microbiol, 13:722-36 (2015); Shmakov et al., Molecular Cell, 60:385-397 (2015). In some embodiments, the RNA-guided DNA binding agent is a Cas nickase, such as a Cas9 nickase. In some embodiments, the RNA-guided DNA binding agent is a Streptococcus pyogenes (S. pyogenes) Cas9 nuclease.
[0028] Non-limiting exemplary species from which nucleases (e.g., Cas nucleases) may be derived include, but are not limited to, Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Listeria innocua, Lactobacillus gasseri, Francisella novicida, Wolinella succinogenes, Sutterella wadsworthensis, Gammaproteobacterium, Neisseria meningitidis, and the like. meningitidis, Campylobacter jejuni, Pasteurella multocida, Fibrobacter succinogene, Rhodospirillum rubrum, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum), Alicyclobacillus acidocaldarius, Bacillus pseudomycoidesPseudomonas pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Lactobacillus buchneri, Treponema denticola, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus foreoxidansferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Streptococcus pasteurianus, Neisseria cinerea, Campylobacter zari, Parvibaculum lavamentivorans, Corynebacterium diphtheriae diphtheria, Acidaminococcus sp., Lachnospiraceae bacterium ND2006, and Acaryochloris marina.
[0029] In some embodiments, the Cas nuclease is a Cas9 nuclease from Streptococcus pyogenes. In some embodiments, the Cas nuclease is a Cas9 nuclease from Streptococcus thermophilus. In some embodiments, the Cas nuclease is a Cas9 nuclease from Neisseria meningitidis. In some embodiments, the Cas nuclease is a Cas9 nuclease from Staphylococcus aureus. In some embodiments, the Cas nuclease is a Cpfl nuclease from Francisella novicida. In some embodiments, the Cas nuclease is a Cpfl nuclease from Acidaminococcus sp. In some embodiments, the Cas nuclease is Cpfl nuclease from the Lachnospiraceae bacterium ND2006.In some embodiments, the Cas nuclease is selected from the group consisting of Francisella tularensis, Lachnospiraceae bacteria, Butyrivibrio proteoclasticus, Peregrinibacteria bacteria, Parcubacteria bacteria, Smithella, Acidaminococcus, Candidatus, Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi, Leptospira inadai, Porphyromonas crevicularis, and the like. In some embodiments, the Cas nuclease is a Cpfl nuclease from Acidaminococcus or Lachnospiraceae.
[0030] Wild-type Cas9 has two nuclease domains: RuvC and HNH. The RuvC domain cleaves the non-target DNA strand, and the HNH domain cleaves the target strand of DNA. In some embodiments, the Cas9 nuclease comprises two or more RuvC domains and / or two or more HNH domains. In some embodiments, the Cas9 nuclease is wild-type Cas9. In some embodiments, Cas9 is capable of inducing a double-stranded break in the target DNA. In certain embodiments, the Cas nuclease is capable of cleaving one or both strands of dsDNA. In some embodiments, the Cas nuclease is capable of cleaving a single strand of DNA. In some embodiments, the Cas nuclease may not have DNA nickase activity.
[0031] In some embodiments, chimeric Cas nucleases are used, in which one domain or region of a protein is replaced with a portion of a different protein. In some embodiments, the Cas nuclease domain can be replaced with a domain from a different nuclease, such as Fok1. In some embodiments, the Cas nuclease can be an engineered nuclease, in which the polypeptide sequence of the nuclease has been modified, in some instances to confer advantageous properties to the nuclease. In some embodiments, the cleavage site of the nuclease relative to the location of the user-specified target sequence is unknown.
[0032] In other embodiments, the Cas nuclease may be derived from a type I CRISPR / Cas system. In some embodiments, the Cas nuclease may be a component of a cascade complex of a type I CRISPR / Cas system. In some embodiments, the Cas nuclease may be a Cas3 protein. In some embodiments, the Cas nuclease may be derived from a type III CRISPR / Cas system. In some embodiments, the Cas nuclease may have RNA cleavage activity.
[0033] Data Processing Pipeline In some embodiments, disclosed herein is a data processing pipeline that includes receiving sample sequence data, applying quality control filters to the sample sequence data, aligning sequencing reads of the sample sequence data to a reference sequence, defining a window based on a sequence position within the nucleic acid guide and the end position of the nucleic acid guide sequence, and quantifying indels in the sample sequence data. In some embodiments, the plurality of sequencing reads are obtained from a next-generation sequencing (NGS) instrument. In some embodiments, the NGS sequencing instrument is an Illumina MiSeq™ instrument. In some embodiments, sequencing adapters and amplification adapters are trimmed by default, such that the returned NGS read data does not have adapter sequences within the reads. In some embodiments, the sequencing adapters and amplification adapters are trimmed as part of the data processing pipeline of the methods disclosed herein.
[0034] 2A shows steps in an upstream molecular biology workflow that can be performed to generate data that is subsequently processed by the data processing pipeline disclosed herein. The steps in the upstream molecular biology workflow can include designing primers to amplify editing targets, isolating cellular DNA, PCR-amplifying the editing targets with adapters, adding next-generation sequencing (NGS) adapters to the PCR products, pooling the barcoded amplicons, bead-purifying the amplicons, spiking in diverse DNA, and performing next-generation sequencing (NGS) on a sequencing instrument.
[0035] 2B shows the steps of the computational pipeline disclosed herein. The steps may include quality control of received sequencing reads, alignment of sequencing reads to target sequences, and quantification of indels according to the methods disclosed herein. These steps are described in further detail below.
[0036] Previously available calculation methods may underestimate the frequency of indels in a target polynucleotide after cleavage by a nucleic acid-guided nuclease, for example, a type II nuclease with a known cleavage site relative to the guide sequence. Figure 3A is a schematic diagram of received sequencing reads of a nucleic acid-guided nuclease-edited sample, including multiple sequencing reads containing indels as a result of cleavage by a nucleic acid-guided nuclease. The nuclease is Cas9, a type II nuclease with a known cleavage site relative to the guide sequence. For each type of indel, the number of sequencing reads containing that indel is shown on the right. Previously known cleavage sites are indicated by rectangular boxes.
[0037] Previously available computational methods may underestimate the frequency of indels in a target polynucleotide after cleavage by a nucleic acid-guided nuclease, for example, a V-type nuclease with a known cleavage site relative to the guide sequence. Figure 3B is a schematic diagram of received sequencing reads for a nucleic acid-guided nuclease-edited sample, including multiple sequencing reads containing indels as a result of cleavage by the nucleic acid-guided nuclease. The nuclease is a novel V-type nuclease. For each type of indel, the number of sequencing reads containing that indel is shown on the right. Some indels are missed by standard computational methods.
[0038] The computational pipeline disclosed herein, and described in more detail below, improves the accuracy of quantifying indel frequency in target polynucleotides after cleavage by a nucleic acid-guided nuclease. Figure 3C is a schematic diagram of received sequencing reads for a nucleic acid-guided nuclease-edited sample, including a plurality of sequencing reads, a subset of which contains indels as a result of cleavage by the nucleic acid-guided nuclease. The nuclease is a novel type V nuclease. For each type of indel, the number of sequencing reads containing that indel is shown on the right. Custom data processing methods capture all indels generated by nuclease cleavage.
[0039] In some embodiments, the data processing pipeline disclosed herein includes one or more modules implemented by the CRISPResso software tools package (Clement K, et al. CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat Biotechnol. 2019 Mar;37(3):224-226 and Canver MC, et al. Integrated design, execution, and analysis of arrayed and pooled CRISPR genome-editing experiments. Nat Protoc. 2018 May;13(5):946-986, incorporated herein by reference in their entireties). In some embodiments, one or more read filtering parameters are applied to the sample sequence data using CRISPResso to remove potentially false positive indels from the sample sequence data to improve the accuracy of estimates of the frequency of indels in the sample sequence data. Read filtering parameters are described in further detail below. In some embodiments, read filtering is performed based on the PHRED quality score, as described in Ewing B, Green P. Base-calling of automated sequencer traces using phred. II. Error probabilities. Genome Res. 1998 Mar;8(3):186-94, incorporated herein by reference in its entirety. The PHRED quality score measures the quality of nucleotide base call identifications in sequencing reads generated by an automated DNA sequence instrument.
[0040] In some embodiments, a minimum average read quality ("q" or "min_average_read_quality") is applied to filter sample sequence data to remove potentially false positive indels. This parameter allows specification of a minimum average quality score for including reads in subsequent analyses. The PHRED score represents the confidence in the assignment of specific nucleotides in a read. A maximum score of 40 corresponds to an error rate of 0.01%. This average quality of reads is useful for filtering out low-quality reads. In some embodiments, a "min_average_read_quality" value of 0, 5, 10, 15, 20, 25, 30, 35, or 40 is applied to sample sequence data.
[0041] In some embodiments, a minimum single base pair score ("s" or "min_single_bp_quality") is applied to filter sample sequence data to remove potentially false positive indels. This parameter allows for specification of a minimum single bp score for including reads in subsequent analyses. This parameter provides more stringent filtering; any reads with a single bp quality below the threshold are discarded. In some embodiments, a "min_single_bp_quality" value of 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 is applied to sample sequence data.
[0042] In some embodiments, an amplicon minimum alignment score ("amas" or "amplicon_min_alignment_score") is applied to filter sample sequence data. After reads are aligned to a reference sequence, homology is calculated as the number of base pairs they have in common. This is useful for filtering erroneous reads that do not align to the target sequence, for example, arising from alternative primer positions. In some embodiments, an "amplicon_min_alignment_score" value of 0, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 is applied to sample sequence data.
[0043] In some embodiments, a combination of the "min_average_read_quality" score, the "min_single_bp_quality" score, and the "amplicon_min_alignment_score" is applied to filter the sample sequence data to improve the estimation of indel frequency in the sample sequence data after cleavage by a target polynucleotide sequence with a nuclease. In some embodiments, a combination of the "min_average_read_quality" score, the "min_single_bp_quality" score, and the "amplicon_min_alignment_score", further combined with a window size based on sequence position within a guide sequence, is applied to filter the sample sequence data to improve the estimation of indel frequency in the sample sequence data after cleavage by a target polynucleotide sequence with a nuclease. As described in Table 1 below, combinations of these parameters can be used to increase the accuracy of indel quantification according to the methods disclosed herein.
[0044] In some embodiments, a combination of the "min_average_read_quality" score, the "min_single_bp_quality" score, and the "amplicon_min_alignment_score" is applied to filter the sample sequence data to improve the estimation of indel frequency in the sample sequence data after cleavage by a target polynucleotide sequence with a nuclease. In some embodiments, a combination of the "min_average_read_quality" score, the "min_single_bp_quality" score, and the "amplicon_min_alignment_score," further in combination with a window size based on the sequence position of a known insertion site, is applied to filter the sample sequence data to improve the estimation of indel frequency in the sample sequence data after cleavage by a target polynucleotide sequence with a nuclease and insertion of the polynucleotide sequence. As described in Table 1 below, combinations of these parameters can be used to increase the accuracy of indel quantification according to the methods disclosed herein.
[0045] In some embodiments, when the sample sequence data includes paired-end reads, the data processing pipeline disclosed herein includes one or more modules implemented by the FLASH software tool (Magoc T, Salzberg SL. FLASH: fast length adjustment of short reads to improve genome assemblies. Bioinformatics. 2011 Nov 1;27(21):2957-63, incorporated herein by reference in its entirety). FLASH is a fast and accurate software tool for merging paired-end reads from next-generation sequencing experiments, designed to merge pairs of reads when the original DNA fragments are shorter than twice the length of the reads. The resulting longer reads can significantly improve genome assembly. FLASH calculates the mismatch rate within two overlapping regions. If the mismatch ratio exceeds a threshold mismatch ratio, FLASH determines that the reads are erroneously overlapping. In some embodiments, paired-end reads are merged using FLASH to generate a single read for alignment to a target reference sequence and reduce sequencing errors that may exist at the end of the sequencing read.
[0046] In some embodiments, a maximum paired-end read overlap ("max_paired_end_reads_overlap") is applied to paired-end sample sequence data for the FLASH read merging step of the data processing pipeline. This parameter represents the maximum overlap length expected for approximately 90% of read pairs. In some embodiments, a "max_paired_end_reads_overlap" value of 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, or 300 is applied to sample sequence data.
[0047] In some embodiments, a minimum paired-end read overlap ("min_paired_end_reads_overlap") is applied to paired-end sample sequence data for the FLASH read merging step of the data processing pipeline. This parameter represents the minimum required overlap length between two reads to provide a reliable overlap. In some embodiments, a "max_paired_end_reads_overlap" value of 0, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 is applied to sample sequence data.
[0048] In some embodiments, a combination of the "min_average_read_quality" score, "min_single_bp_quality" score, "min_paired_end_reads_overlap" and "max_paired_end_reads_overlap" is applied to filter the sample sequence data to improve estimation of indel frequency in the sample sequence data after cleavage by a target polynucleotide sequence with a nuclease. In some embodiments, a combination of the "min_average_read_quality" score, "min_single_bp_quality" score, "min_paired_end_reads_overlap" and "max_paired_end_reads_overlap", further combined with a window size based on sequence position within a guide sequence, is applied to filter the sample sequence data to improve estimation of indel frequency in the sample sequence data after cleavage by a target polynucleotide sequence with a nuclease. As described in Table 1 below, combinations of these parameters can be used to increase the accuracy of indel quantification according to the methods disclosed herein.
[0049] In some embodiments, a combination of the "min_average_read_quality" score, "min_single_bp_quality" score, "min_paired_end_reads_overlap," and "max_paired_end_reads_overlap" is applied to filter the sample sequence data to improve estimation of indel frequency in the sample sequence data after cleavage by a target polynucleotide sequence with a nuclease. In some embodiments, a combination of the "min_average_read_quality" score, "min_single_bp_quality" score, "min_paired_end_reads_overlap," and "max_paired_end_reads_overlap," further in combination with a window size based on the sequence position of a known insertion site, is applied to filter the sample sequence data to improve estimation of indel frequency in the sample sequence data after cleavage by a target polynucleotide sequence with a nuclease and insertion of the polynucleotide sequence. As described in Table 1 below, combinations of these parameters can be used to increase the accuracy of indel quantification according to the methods disclosed herein.
[0050] In some embodiments, the quantification window extends 50, 40, 30, 20, 10, or 5 nucleotides 5' to the 5' end of the guide sequence. In some embodiments, the quantification window extends 50, 40, 30, 20, 10, or 5 nucleotides 3' to the 3' end of the guide sequence. In some embodiments, the quantification window extends 50, 40, 30, 20, 10, or 5 nucleotides 5' to the known insertion site. In some embodiments, the quantification window extends 50, 40, 30, 20, 10, or 5 nucleotides 3' to the known insertion site. In some embodiments, the quantification window is 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, or 500 nucleotides in length.
[0051] In some embodiments, the detected indels are about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides in length. In some embodiments, the detected indels are about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more nucleotides in length. In some embodiments, the indels comprise CRISPR-mediated donor insertions. In some embodiments, the indels are the result of homology-directed repair (HDR). In some embodiments, the indels are the result of insertions by CRISPR-associated transposase (CAST). In some embodiments, the insertions comprise sequences from a genomic library.
[0052] [Table 1]
[0053] [Table 2]
[0054] [Table 3]
[0055] [Table 4]
[0056] [Table 5]
[0057] [Table 6]
[0058] [Table 7]
[0059] [Table 8]
[0060] [Table 9]
[0061] [Table 10]
[0062] [Table 11]
[0063] Computer implementation of the method 5 is a diagram of computer system 500 components that can be used to implement a computational pipeline for indel quantification of polynucleotide sequences cleaved by nucleic acid-guided nucleases, including uncharacterized nucleases. Computer system 500 can be used to implement methods that include parameters for alignment of next-generation sequencing (NGS) reads from targeted amplicon sequencing (TAS) of nuclease-edited or control samples obtained and aligned to a reference amplicon sequence.
[0064] Computing device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 550 is intended to represent various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, and other similar computing devices. Additionally, computing device 500 or 550 may include a universal serial bus (USB) flash drive. The USB flash drive may store an operating system and other applications. The USB flash drive may include input / output components, such as a wireless transmitter or a USB connector that can be inserted into a USB port of another computing device. The components, their connections and relationships, and their functions illustrated herein are intended to be exemplary only and are not intended to limit the implementation of the methods and configurations described and / or claimed herein.
[0065] The computing device 500 includes a processor 502, memory 504, a storage device 506, a high-speed interface 508 connecting to the memory 504 and a high-speed expansion port 510, a low-speed bus 514, and a low-speed interface 512 connecting to the storage device 506. The components 502, 504, 506, 508, 510, and 512 are interconnected using various buses and may be implemented on a common motherboard or in other suitable manners. The processor 502 processes instructions for execution within the computing device 500, including instructions stored in the memory 504 or the storage device 506, and can display graphical information for a GUI on an external input / output device, such as a display 516 coupled to the high-speed interface 508. Other implementations may use multiple processors and / or multiple buses, along with multiple memories and types of memory, as appropriate. Multiple computing devices 500 may also be connected, each providing a portion of the necessary operations, such as, for example, a server bank, a cluster of blade servers, or a multiprocessor system.
[0066] The memory 504 stores information within the computing device 500. In one implementation, the memory 504 is one or more volatile memory units. In another embodiment, the memory 504 is one or more non-volatile memory units. The memory 504 may also be another form of computer-readable medium, such as a magnetic or optical disk.
[0067] The storage device 506 can provide mass storage for the computing device 500. In one embodiment, the storage device 506 can be or include a computer-readable medium such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices including devices in a storage area network or other configuration. A computer program product can be tangibly embodied on an information carrier. The computer program product can also include instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as memory 504, the storage device 506, or memory on the processor 502.
[0068] The high-speed controller 508 manages bandwidth-intensive operations of the computing device 500, while the low-speed controller 512 manages less bandwidth-intensive operations. This division of functions is merely an example. In one embodiment, the high-speed controller 508 is coupled to the memory 504, the display 516, via a graphics processor or accelerator, for example, and to a high-speed expansion port 510 that can accept various expansion cards (not shown). In this embodiment, the low-speed controller 512 is coupled to the storage device 506 and the low-speed expansion port 514. The low-speed expansion port may include various communication ports, such as USB, Bluetooth, Ethernet, wireless Ethernet, etc., but may also be coupled via a network adapter to one or more input / output devices, such as a keyboard, a pointing device, a microphone / speaker pair, a scanner, or a networking device, such as a switch or router. The computing device 500, as shown, may be implemented in several different forms. For example, it may be implemented as a standard server 520 or multiple times in a group of such servers. It may also be implemented as part of a rack server system 524. Furthermore, it may be implemented in a personal computer such as a laptop computer 522. Alternatively, components from computing device 500 may be combined with other components in a mobile device (not shown), such as device 550. Each such device may include one or more computing devices 500, 550, and the overall system may consist of multiple computing devices 500, 550 in communication with each other.
[0069] Computing device 500, as shown, may be implemented in several different forms. For example, it may be implemented as a standard server 520 or multiple times in a group of such servers. It may also be implemented as part of a rack server system 524. Furthermore, it may be implemented in a personal computer, such as a laptop computer 522. Alternatively, components from computing device 500 may be combined with other components in a mobile device (not shown), such as device 550. Each such device may include one or more computing devices 500, 550, and the entire system may be made up of multiple computing devices 500, 550 communicating with each other.
[0070] Computing device 550 includes, among other components, a processor 552, memory 564, input / output devices such as a display 554, a communication interface 566, and a transceiver 568. Device 550 may be provided with a storage device such as a microdrive or other device to provide additional storage. Each of the components 550, 552, 564, 554, 566, 568 are interconnected using various buses, and some of the components may be implemented on a common motherboard or in any other suitable manner.
[0071] Processor 552 can execute instructions within computing device 550, including instructions stored in memory 564. The processor can be implemented as a chipset of chips including separate analog and digital processors. Additionally, the processor can be implemented using any of a number of architectures. For example, processor 510 can be a CISC (Complex Instruction Set Computer) processor, a RISC (Reduced Instruction Set Computer) processor, or a MISC (Minimum Instruction Set Computer) processor. The processor can coordinate other components of device 550, such as controlling a user interface, applications run by device 550, and wireless communication by device 550.
[0072] Processor 552 can communicate with a user via control interface 558 and display interface 556 coupled to display 554. Display 554 can be, for example, a TFT (thin film transistor liquid crystal display) display, an OLED (organic light emitting diode) display, or other suitable display technology. Display interface 556 can include appropriate circuitry for driving display 554 to present graphics and other information to the user. Control interface 558 can receive commands from the user and convert them for submission to processor 552. Additionally, an external interface 562 can be in communication with processor 552 to enable short-range communication between device 550 and other devices. External interface 562 can provide, for example, for wired communication in some implementations or wireless communication in other implementations; multiple interfaces can also be used.
[0073] Memory 564 stores information within computing device 550. Memory 564 may be implemented as one or more of a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. Expansion memory 574 may also be provided and connected to device 550 via expansion interface 572, which may include, for example, a SIMM (single in-line memory module) card interface. Such expansion memory 574 may provide additional storage space for device 550 or may store applications or other information for device 550. Specifically, expansion memory 574 may include instructions for implementing or supplementing the processes described above, and may also include secure information. Thus, for example, expansion memory 574 may be provided as a security module for device 550 and programmed with instructions that enable secure use of device 550. Additionally, secure applications may be provided via SIMM cards, along with additional information, such as placing identifying information on the SIMM card in an unhackable manner.
[0074] The memory may include, for example, flash memory and / or NVRAM memory, as described below. In one embodiment, a computer program product is tangibly embodied on an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is, for example, a computer-readable or machine-readable medium, such as memory 564, expansion memory 574, or memory on processor 552, receivable via transceiver 568 or external interface 562.
[0075] Device 550 can communicate wirelessly via communication interface 566, which may include digital signal processing circuitry as needed. Communication interface 566 can provide for communication in a variety of modes or protocols, such as GSM voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS. Such communication can occur, for example, via radio frequency transceiver 568. Additionally, short-range communication can occur, such as using Bluetooth, Wi-Fi, or other such transceivers (not shown). Additionally, GPS (Global Positioning System) receiver module 570 can provide additional navigation- and location-related wireless data to device 550, which can be used as appropriate by applications executing on device 550.
[0076] Device 550 can also communicate voice using audio codec 560, which can receive voice information from a user and convert it into usable digital information. Audio codec 560 can also generate audible sounds for the user, for example, through a speaker in the handset of device 550. Such sounds can include sounds from a voice call, recorded sounds such as voice messages, music files, and the like, and can also include sounds generated by applications running on device 550.
[0077] The computing device 550, as shown, can be implemented in many different forms, for example, as a mobile phone 580. It can also be implemented as part of a smartphone 582, personal digital assistant, or other similar mobile device.
[0078] Various implementations of the systems and methods described herein may be realized in digital electronic circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations of such implementations. These various implementations may include implementation in one or more computer programs executable and / or interpretable by a programmable system including at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0079] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language and / or assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal, such as a magnetic disk, optical disk, memory, or programmable logic device (PLD). The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0080] To provide for user interaction, the systems and techniques described herein can be implemented in a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to a user, and a keyboard and pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices can be used to provide for user interaction as well; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0081] The systems and techniques described herein may be implemented in a computing system that includes back-end components, such as data servers, middleware components, e.g., application servers, or front-end components, such as client computers having a graphical user interface or web browser that allows users to interact with implementations of the systems and techniques described herein, or any combination of such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, e.g., a communications network. Examples of communications networks include a local area network ("LAN"), a wide area network ("WAN"), and the Internet.
[0082] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communications network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. Numerous embodiments have been described. Nevertheless, it will be understood that various modifications are possible without departing from the spirit and scope of the invention. In addition, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. In addition, other steps may be provided or deleted from the described flows, and other components may be added to or removed from the described systems. Accordingly, other embodiments are within the scope of the following claims.
[0083] All of the embodiments and functional operations of the present disclosure described herein may be implemented in digital electronic circuitry, or computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in one or more combinations thereof. Embodiments of the methods and compositions may be implemented as one or more computer program products, e.g., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or to control the operation of a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter providing a machine-readable propagated signal, or one or more combinations thereof. The term "data processing apparatus" encompasses all apparatuses, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus may include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to an appropriate receiver apparatus.
[0084] A computer program (also known as a program, software, software application, script, or code) may be written in any type of programming language, including compiled or interpreted languages, and may be deployed in any form, such as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple associated files (e.g., files storing one or more modules, subprograms, or portions of code). A computer program may be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communications network.
[0085] The processes and logic flows described herein may be performed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and apparatus may be implemented as, special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0086] Processors suitable for executing a computer program include, by way of example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will include one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, or both, for storing data, or will be operatively coupled to receive data from or transfer data to them. However, a computer need not have such devices. Furthermore, a computer can be incorporated into another device, such as a tablet computer, a mobile phone, a personal digital assistant (PDA), a portable audio player, or a global positioning system (GPS) receiver, to name a few. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0087] To provide for user interaction, embodiments of the present disclosure can be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices can be used to provide for user interaction as well; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0088] Embodiments of the present disclosure may be implemented in a computing system including back-end components, e.g., a data server or middleware components, a computing system including an application server, or a front-end component, e.g., a computing system including a client computer having a graphical user interface or web browser through which a user can interact with method embodiments, or any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, e.g., a communications network. Examples of communications networks include a local area network ("LAN"), a wide area network ("WAN") such as the Internet, or the like.
[0089] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communications network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. While this specification contains many details, these should not be construed as limitations on the scope of the invention or what may be claimed, but rather as descriptions of features particular to particular embodiments of the invention. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as acting in a particular combination, and even initially claimed as such, one or more features from a claimed combination may, in some cases, be carved out of that combination, and the claimed combination may be directed to a subcombination or variations of the subcombination.
[0090] Similarly, while operations are shown in the figures in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown, or in any sequential order, or that all of the illustrated operations be performed, to achieve desirable results. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated into a single software product or packaged into multiple software products.
[0091] In each instance where an HTML file is mentioned, other file types or formats can be substituted. For example, the HTML file can be replaced with an XML, JSON, plain text, or other type of file. Additionally, where a table or hash table is mentioned, other data structures (such as a spreadsheet, relational database, or structured file) can be used. [Example]
[0092] The invention is further described in the following examples, which do not limit the scope of the invention described in the claims.
[0093] Example 1: Benchmarking of indel quantification methods using well-characterized nucleases To evaluate the effectiveness of the indel quantification method disclosed herein, we apply the data processing pipeline to data generated by next-generation sequencing of mammalian cells edited with SpCas9, a type II nuclease with a known cleavage site and editing pattern. The percentage of cells estimated to contain indels in the edited target DNA sequence is determined using standard CRISPResso2 processing parameters and the method disclosed herein. As shown in Figure 4A, the method disclosed herein performs similarly to standard CRISPResso2 processing parameters, indicating that the method does not overestimate or underestimate the percentage of cells estimated to contain indels in the edited target DNA sequence using a well-characterized nuclease with a known cleavage site and editing pattern.
[0094] Example 2: Dilution experiments to quantify indels To verify the accuracy of the indel quantification method disclosed herein, we applied the data processing pipeline to data generated by next-generation sequencing of mammalian cells edited by a V-type nuclease with an unknown cut site in a serial dilution experiment. DNA isolated from the V-type nuclease-edited mammalian cells was serially diluted with DNA isolated from unedited cells, e.g., 0% unedited / 100% edited, 25% unedited / 75% edited, 50% unedited / 50% edited, 75% unedited / 25% edited, and 0% unedited / 100% edited (Figure 4B). The serially diluted DNA mixtures were sequenced to generate sequencing reads containing the edited target sequence. As shown in Figure 4C, the observed indel percentage (y-axis) was plotted against the expected indel percentage (x-axis) for each serial dilution concentration. The linearity of the plot, with an R-squared value of approximately 0.99, confirmed the effectiveness of the indel quantification data processing pipeline.
[0095] Example 3: Quantification of indels after editing with novel nucleases To validate the effectiveness of the indel quantification method disclosed herein for estimating indels in target sequences edited with a novel nuclease and variable guide sequence, the data processing pipeline is applied to data generated by next-generation sequencing of mammalian cells edited with a novel V-type nuclease using increasing concentrations and two different guides: guide sequences A1 and F2. As shown in Figure 6, the data processing pipeline accurately estimates the percentage of indels in edits with the novel V-type nuclease at nuclease concentrations of 5 μM, 11 μM, and 22 μM for guide sequences A1 and F2. Furthermore, the data processing pipeline is applied to data generated by next-generation sequencing of mammalian cells edited with the nuclease AsCas12a to validate the effectiveness of the indel quantification method disclosed herein.
[0096] Example 4: Dilution experiments to quantify larger insertions To verify the accuracy of the indel quantification method disclosed herein for quantifying larger insertions, i.e., insertions greater than approximately 10 nucleotides, we applied the data processing pipeline to data generated by next-generation sequencing of mammalian cells in which a sequence of known length, 21 nucleotides, was introduced into a specific insertion site in a serial dilution spike-in experiment (see Figure 7A). DNA isolated from the insertion-containing mammalian cells was serially diluted with DNA isolated from cells without the insertion, e.g., 0% insertion / no insertion, 25% insertion / 75% no insertion, 50% insertion / 50% no insertion, 75% insertion / 25% no insertion, and 0% insertion / 100% no insertion (Figure 4B). The serially diluted DNA mixtures were sequenced to generate sequencing reads containing the insertion site. As shown in Figure 7B, in the experiment where the insertion-containing DNA was spiked at 25%, the sequence spike was detected in approximately 23% of cases. As shown in Figure 7C, in the experiment where the insertion-containing DNA was spiked at 50%, the sequence spike was detected in approximately 47% of cases. As shown in Figure 7D, in an experiment where DNA containing an insertion was spiked at 75%, sequence spikes were detected at approximately 76%. These results demonstrate that the data processing pipeline disclosed herein is effective for quantifying, for example, CRISPR-mediated donor insertions, insertions mediated by homology-directed repair (HDR), insertions introduced by CRISPR-associated transposase (CAST) and other genome engineering methods, and insertions in genomic libraries used in high-throughput screening (HTS).
[0097] Other embodiments While the present invention has been described with reference to its detailed description, it will be understood that the foregoing description is for illustrative purposes only and is not intended to limit the scope of the invention as defined by the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.
Claims
1. 1. A method for quantifying insertions and / or deletions in a polynucleotide sequence caused by cleavage of a target polynucleotide sequence by a nucleic acid-guided nuclease, comprising: receiving, in a computer system, sample sequence data comprising a plurality of sequencing reads; filtering the plurality of sequencing reads; aligning the plurality of sequencing reads to a reference sequence; defining a window based on a sequence position within the nucleic acid guide sequence and the positions of the ends of the nucleic acid guide sequence; determining the number of sequencing reads containing insertions or deletions within the window relative to the reference sequence based on the alignment of each sequencing read among the plurality of sequencing reads to the reference sequence; estimating the amount of insertions and / or deletions in the polynucleotide sequence mediated by the nucleotide-directed nuclease based on the number of sequencing reads containing insertions or deletions. A method comprising:
2. 10. The method of claim 1, wherein the nucleic acid-guided nuclease is a class 2 nuclease.
3. 3. The method of claim 1 or 2, wherein the nucleic acid-guided nuclease is a type II nuclease.
4. 4. The method of claim 3, wherein the nucleic acid-guided nuclease is SpCas9.
5. 4. The method of claim 3, wherein the nucleic acid-guided nuclease is AsCasl2a.
6. 3. The method of claim 1 or 2, wherein the nucleic acid-guided nuclease is a type V nuclease.
7. 7. The method of any one of claims 1 to 6, wherein the nucleic acid guide is a guide RNA (gRNA).
8. 7. The method of any one of claims 1 to 6, wherein the nucleic acid guide is a single guide RNA (sgRNA).
9. 9. The method of any one of claims 1 to 8, wherein the center of the window is located at the center of the nucleic acid guide sequence.
10. The method of any one of claims 1 to 8, wherein the center of the window is located at the site where the nuclease cleaved the polynucleotide sequence.
11. 11. The method of any one of claims 1 to 10, wherein the length of the window is equal to the length of the nucleic acid guide sequence.
12. 12. The method of any one of claims 1 to 11, wherein the 5' end of the window extends 50 base pairs 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 50 base pairs 3' to the 3' end of the nucleic acid guide.
13. 13. The method of any one of claims 1 to 12, wherein the 5' end of the window extends 40 base pairs 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 40 base pairs 3' to the 3' end of the nucleic acid guide.
14. 14. The method of any one of claims 1 to 13, wherein the 5' end of the window extends 30 base pairs 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 30 base pairs 3' to the 3' end of the nucleic acid guide.
15. 15. The method of any one of claims 1 to 14, wherein the 5' end of the window extends 20 base pairs 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 20 base pairs 3' to the 3' end of the nucleic acid guide.
16. 16. The method of any one of claims 1 to 15, wherein the 5' end of the window extends 10 base pairs 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 10 base pairs 3' to the 3' end of the nucleic acid guide.
17. 17. The method of any one of claims 1 to 16, wherein the 5' end of the window extends 5 base pairs 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 5 base pairs 3' to the 3' end of the nucleic acid guide.
18. 18. The method of any one of claims 1 to 17, further comprising trimming adapter sequences from the plurality of sequencing reads.
19. 19. The method of any one of claims 1 to 18, wherein the plurality of sequencing reads are generated from targeted amplicon sequencing.
20. 20. The method of any one of claims 1-19, wherein the plurality of sequencing reads are generated from targeted amplicon sequencing of DNA isolated from cells edited by the nucleic acid-guided nuclease.
21. 21. The method of any one of claims 1 to 20, wherein the plurality of sequencing reads are paired-end sequencing reads.
22. 22. The method of claim 21, further comprising read-merging the paired-end sequencing reads to generate a single read for alignment to the reference sequence.
23. 23. The method of claim 22, wherein the read merging of the paired-end reads further comprises applying a minimum paired-end read overlap score to the plurality of sequencing reads.
24. 24. The method of claim 23, wherein the minimum paired-end read overlap score is 0, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100.
25. 25. The method of claim 23 or 24, wherein the minimum paired-end read overlap score is 10.
26. 26. The method of any one of claims 22 to 25, wherein the read merging of paired-end reads further comprises applying a maximum paired-end read overlap score to the plurality of sequencing reads.
27. 27. The method of claim 26, wherein the maximum paired-end read overlap score is 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, or 300.
28. 28. The method of claim 26 or 27, wherein the maximum paired-end read overlap score is 100.
29. 29. The method of any one of claims 1 to 28, wherein the filtering step further comprises applying a minimum average read quality score to the plurality of sequencing reads.
30. 30. The method of claim 29, wherein the minimum average read quality score is 0, 5, 10, 15, 20, 25, 30, 35, or 40.
31. 31. The method of any one of claims 1 to 30, wherein the filtering step further comprises applying a minimum single base pair score to the plurality of sequencing reads.
32. 32. The method of claim 31 , wherein the minimum single base pair score is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20.
33. 33. The method of any one of claims 1 to 32, wherein the aligning step further comprises applying an amplicon minimum alignment score to the plurality of sequencing reads.
34. 34. The method of claim 33, wherein the amplicon minimum alignment score is 0, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100.
35. 35. The method of any one of claims 1 to 34, wherein the nucleic acid-guided nuclease is a FokI nuclease.
36. 36. The method of claim 35, wherein the FokI nuclease is fused to a transcription activator-like (TAL) protein.
37. 37. The method of any one of claims 1 to 36, wherein the nucleic acid-guided nuclease is a zinc finger nuclease.
38. A computer program product tangibly embodied on a computer-readable medium, the computer program product comprising instructions that, when executed by one or more processors, receiving sample sequence data including a plurality of sequencing reads; filtering the plurality of sequencing reads; aligning the plurality of sequencing reads to a reference sequence; defining a window based on a sequence position within the nucleic acid guide sequence and the positions of the ends of the nucleic acid guide sequence; determining the number of sequencing reads containing insertions or deletions within the window relative to the reference sequence based on the alignment of each sequencing read among the plurality of sequencing reads to the reference sequence; estimating the amount of insertions and / or deletions in the polynucleotide sequence mediated by the nucleotide-directed nuclease based on the number of sequencing reads containing insertions or deletions; 1. A computer program product configured to: