Predicting insertion missing frequency
By receiving and filtering the sequencing reads, comparing and defining windows, the problem of difficulty in quantifying the frequency of uncharacterized nuclease insertion and deletion is solved, and an accurate evaluation of the uncharacterized nuclease editing system is achieved.
Patent Information
- Application Number
- CN202380088821.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-02
- Filing Date
- 2023-12-05
- Publication Date
- 2025-08-08
AI Technical Summary
Existing computing tools are difficult to accurately quantify the frequency of insertion and deletion of under-characterized nucleases in polynucleotide sequences, especially when nuclease cleavage sites relative to user-specified target sequences are unknown.
A calculation process is provided to estimate insertion and deletion frequencies in polynucleotide sequences by receiving and filtering multiple sequencing reads, performing alignment, and estimating the frequency of insertion and deletion in polynucleotide sequences based on the position definition window of the nucleic acid guide sequence, suitable for uncharacterized nucleases.
The ability to accurately quantify the insertion and deletion frequency of uncharacterized nucleases in polynucleotide sequences improves the accuracy of evaluating the efficacy of nuclease editing systems.
Smart Images

Figure BDA0005464797810000171 
Figure BDA0005464797810000181 
Figure BDA0005464797810000191
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 430,827, filed on December 7, 2022, and European Patent Application No. EP 23315141.4, filed on May 2, 2023, the contents of which are incorporated herein by reference in their entirety. Technical Field
[0003] The present disclosure relates to the quantification of insertions and / or deletions near cleavage sites within a polynucleotide target sequence. Background Art
[0004] Nucleic acid-guided nucleases can be used to edit polynucleotide sequences, such as the genome of an organism, at a targeted location with high precision. Nucleases are enzymes that can cut the phosphodiester bonds between the nucleotides of a nucleic acid. Genome editing methods include using CRISPR (clustered regularly interspaced short palindromic repeats)-related proteins or similar nucleic acid-guided nucleases to induce DNA double-strand breaks (DSBs) at predictable genomic positions relative to a user-specified target sequence. DNA DSBs are repaired by intracellular mechanisms, such as by non-homologous end joining (NHEJ). The repair process may result in sequence variants, including, for example, insertions and deletions (indels). Quantification of the frequency of insertions and deletions at the target cleavage site within the editing cell population is important for evaluating the efficacy of the nucleic acid-guided nuclease editing system.
[0005] Currently, a variety of nucleic acid-guided nucleases are known, and many more naturally occurring and engineered nucleases are being discovered and characterized. New nucleases may be poorly characterized, meaning their cleavage sites and editing windows relative to the user-specified target sequence are unknown. Existing indel quantification data analysis processes for next-generation sequencing (NGS) rely on the assumption that the cleavage sites and editing windows are well-defined, which is not applicable to new nucleases that are not well characterized. Therefore, existing computational tools may underestimate indels at target cleavage sites within edited cell populations. To date, few computational methods have been developed to address this need. Summary of the Invention
[0006] The present disclosure is based in part on the discovery that a data analysis process can be configured to quantify insertions and deletions of a target polynucleotide sequence that has been cleaved by a nucleic acid-guided nuclease, even if the cleavage site of the nuclease relative to a user-specified target sequence is unknown. The methods disclosed herein include computational processes for quantifying insertions and deletions of a polynucleotide sequence that has been cleaved by a nucleic acid-guided nuclease (including an uncharacterized nuclease). The methods disclosed herein include parameters for aligning next-generation sequencing (NGS) reads of targeted amplicon sequencing (TAS) of a nuclease-edited sample or control sample that has been obtained and aligned to a reference amplicon sequence. Exemplary experimental and computational data are disclosed herein that demonstrate that these methods can be used to accurately quantify insertions and deletions of novel uncharacterized nucleases, e.g., nucleases with unknown cleavage sites relative to a user-specified target sequence.
[0007] In some embodiments, provided herein are methods for quantifying insertions and / or deletions in a polynucleotide sequence caused by cleavage of a target polynucleotide sequence by a nucleic acid-guided nuclease. In some embodiments, these methods include, in a computer system, receiving sample sequence data comprising a plurality of sequencing reads; filtering the plurality of sequencing reads; aligning the plurality of sequencing reads to a reference sequence; defining a window based on the sequence position within the nucleic acid guide sequence and the position of the ends of the nucleic acid guide sequence; determining the number of sequencing reads comprising insertions or deletions relative to the reference sequence within the window based on the alignment of each sequencing read in the plurality of sequencing reads to the reference sequence; and estimating the number of insertions and / or deletions in the polynucleotide sequence mediated by the nucleotide-guided nuclease based on the number of sequencing reads comprising insertions or deletions. In some embodiments, the nucleic acid-guided nuclease is a Class 2 nuclease. In some embodiments, the nucleic acid-guided nuclease is a Type II nuclease. In some embodiments, the nucleic acid-guided nuclease is SpCas9. In some embodiments, the nucleic acid-guided nuclease is AsCas12a. In some embodiments, the nucleic acid-guided nuclease is a Type V nuclease. In some embodiments, the nucleic acid guide is a guide RNA (gRNA). In some embodiments, the nucleic acid guide is a single guide RNA (sgRNA).
[0008] In some embodiments, the center of the window is located at the center of the nucleic acid guide sequence. In some embodiments, the center of the window is located at the site where the nuclease cleaves the polynucleotide sequence. In some embodiments, the length of the window is equal to the length of the nucleic acid guide sequence. In some embodiments, the 5′ end of the window extends 50 base pairs from the 5′ end to the 5′ end of the nucleic acid guide, and the 3′ end of the window extends 50 base pairs from the 3′ end to the 3′ end of the nucleic acid guide. In some embodiments, the 5′ end of the window extends 40 base pairs from the 5′ end to the 5′ end of the nucleic acid guide, and the 3′ end of the window extends 40 base pairs from the 3′ end to the 3′ end of the nucleic acid guide. In some embodiments, the 5′ end of the window extends 30 base pairs from the 5′ end to the 5′ end of the nucleic acid guide, and the 3′ end of the window extends 30 base pairs from the 3′ end to the 3′ end of the nucleic acid guide. In some embodiments, the 5′ end of the window extends 20 base pairs from the 5′ end to the 5′ end of the nucleic acid guide, and the 3′ end of the window extends 20 base pairs from the 3′ end to the 3′ end of the nucleic acid guide. In some embodiments, the 5′ end of the window extends 10 base pairs from the 5′ end to the 5′ end of the nucleic acid guide, and the 3′ end of the window extends 10 base pairs from the 3′ end to the 3′ end of the nucleic acid guide. In some embodiments, the 5′ end of the window extends 5 base pairs from the 5′ end to the 5′ end of the nucleic acid guide, and the 3′ end of the window extends 5 base pairs from the 3′ end to the 3′ end of the nucleic acid guide.
[0009] In some embodiments, the method further comprises trimming adapter sequences from a plurality of sequencing reads. In some embodiments, the plurality of sequencing reads are generated by targeted amplicon sequencing. In some embodiments, the plurality of sequencing reads are generated by targeted amplicon sequencing of DNA isolated from cells edited by a nucleic acid-guided nuclease. In some embodiments, the plurality of sequencing reads are end-paired sequencing reads. In some embodiments, the method further comprises merging reads of the end-paired sequencing reads to produce a single read for alignment with a reference sequence. In some embodiments, the read merging of the end-paired reads further comprises applying a minimum end-paired read overlap score to the plurality of sequencing reads. In some embodiments, the minimum end-paired read overlap score is 0, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100. In some embodiments, the minimum end-paired read overlap score is 10. In some embodiments, the read merging of the end-paired reads further comprises applying a maximum end-paired read overlap score to the plurality of sequencing reads. In some embodiments, the maximum end-paired read overlap score is 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, or 300. In some embodiments, the maximum end-paired read overlap score is 100. In some embodiments, the filtering step further comprises applying a minimum average read quality score to the plurality of sequencing reads. In some embodiments, the minimum average read quality score is 0, 5, 10, 15, 20, 25, 30, 35, or 40. In some embodiments, the filtering step further comprises applying a minimum single base pair score to the plurality of sequencing reads. In some embodiments, the minimum single base pair score is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. In some embodiments, the aligning step further comprises applying an amplicon minimum alignment score to the plurality of sequencing reads. In some embodiments, the amplicon minimum alignment score is 0, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100.
[0010] In some embodiments, the nucleic acid-guided nuclease is a FokI nuclease. In some embodiments, the FokI nuclease is fused to a transcription activator-like (TAL) protein. In some embodiments, the nucleic acid-guided nuclease is a zinc finger nuclease.
[0011] In some embodiments, described herein is a computer program product tangibly embodied on a computer-readable medium, comprising instructions that, when executed by one or more processors, are configured to: receive sample sequence data comprising a plurality of sequencing reads; filter the plurality of sequencing reads; align the plurality of sequencing reads to a reference sequence; define a window based on sequence position within a nucleic acid guide sequence and the position of an end of the nucleic acid guide sequence; determine, based on alignment of each sequencing read in the plurality of sequencing reads to the reference sequence, a number of sequencing reads comprising insertions or deletions within the window relative to the reference sequence; and estimate, based on the number of sequencing reads comprising insertions or deletions, a number of insertions and / or deletions in the polynucleotide sequence mediated by a nucleotide-guided nuclease.
[0012] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Methods and materials for use in the present invention are described herein; other suitable methods and materials known in the art may also be used. These materials, methods, and examples are illustrative only and are not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated by reference in their entirety. In the event of a conflict, the present specification, including definitions, will control.
[0013] Other features and advantages of the invention will be apparent from the following detailed description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1A A cartoon representation of a type II nuclease, featuring an sgRNA interacting with a target DNA sequence. The PAM (NGG) sequence, guide sequence, target DNA sequence, and cleavage site are indicated.
[0015] Figure 1B A cartoon representation of a V-type nuclease, featuring a gRNA interacting with a target DNA sequence. The PAM (TTTN) sequence, guide sequence, target DNA sequence, and cleavage site are indicated.
[0016] Figure 2A Figure 1 is a diagram of the upstream molecular biology workflow for the methods disclosed herein. The steps include: designing primers to amplify the editing target; isolating cellular DNA; amplifying the editing target using adapter PCR; adding next-generation sequencing (NGS) adapters to the PCR products; mixing barcoded amplicons; bead purification of amplicons; spiking into diversity DNA; and performing NGS on a sequencing instrument.
[0017] Figure 2Bis a diagram of the computational process disclosed herein. The steps include quality control of received sequencing reads; alignment of sequencing reads to target sequences; and quantification of indels according to the methods disclosed herein.
[0018] Figure 3A Figure 2 is a schematic diagram of received sequencing reads for a nucleic acid-guided nuclease-edited sample, comprising a plurality of sequencing reads, a subset of which contain indels resulting from nucleic acid-guided nuclease cleavage. The nuclease is Cas9, a type II nuclease with a known cleavage site relative to the guide sequence. For each type of indel, the number of sequencing reads containing that indel is indicated to the right. The previously known cleavage site is indicated by a rectangular box.
[0019] Figure 3B This figure shows a schematic representation of the sequence reads received for a nucleic acid-guided nuclease-edited sample, including multiple sequence reads, a subset of which contain indels resulting from nucleic acid-guided nuclease cleavage. The nuclease is a novel V-type nuclease. For each type of indel, the number of sequence reads containing that indel is indicated on the right. Several indels were missed by standard computational processing methods.
[0020] Figure 3C This figure shows a schematic representation of received sequencing reads for a nucleic acid-guided nuclease-edited sample. The sample contains multiple sequencing reads, a subset of which contains indels resulting from nucleic acid-guided nuclease cleavage. The nuclease is a novel V-type nuclease. For each type of indel, the number of sequencing reads containing that indel is indicated to the right. A custom data processing method captures all indels generated by nuclease cleavage.
[0021] Figure 4A is a diagram comparing the standard data processing pipeline of the known nuclease Cas9 with the data processing pipeline of Cas9 disclosed in this article.
[0022] Figure 4B Schematic diagram of a dilution experiment used to validate the data processing pipeline disclosed herein, in which a target sequence edited by a nucleic acid-guided nuclease was serially diluted with an unedited target sequence to assess the efficacy of the data processing method disclosed herein.
[0023] Figure 4C yes Figure 4B Schematic diagram of the expected percentage of indels (x-axis) versus the observed percentage of indels (y-axis) for the experiment depicted in FIG.
[0024] Figure 5 is a diagram of computer system components that can be used to implement a computational process for quantifying indels in polynucleotide sequences that have been cleaved by nucleic acid-guided nucleases, including uncharacterized nucleases.
[0025] Figure 6 is a graph of the percentage of indels (y-axis) for two nucleases (x-axis): Cas9 (reference) and a novel V-type nuclease using two different guide sequences, A1 and F2, at increasing doses.
[0026] Figure 7A is a schematic diagram of an experiment in which the frequency of 21 nucleotide insertions was quantified by the computational pipeline disclosed herein.
[0027] Figure 7B yes Figure 7A Schematic diagram of the received sequencing reads for the amplicon depicted in , which comprises a plurality of sequencing reads, a subset of which comprises an insertion of 21 nucleotides spiked into the dilution experiment.
[0028] Figure 7C yes Figure 7A Schematic diagram of the received sequencing reads for the amplicon depicted in , which comprises a plurality of sequencing reads, a subset of which comprises an insertion of 21 nucleotides spiked into the dilution experiment.
[0029] Figure 7D yes Figure 7A Schematic diagram of the received sequencing reads for the amplicon depicted in , which comprises a plurality of sequencing reads, a subset of which comprises an insertion of 21 nucleotides spiked into the dilution experiment. DETAILED DESCRIPTION
[0030] definition
[0031] As used herein, the terms "nucleic acid," "polynucleotide," and "oligonucleotide" are interchangeable and refer to deoxyribonucleotide or ribonucleotide polymers, in linear or cyclic configurations, and in single-stranded or double-stranded form. For the purposes of this disclosure, these terms should not be construed as limiting the length of the polymer. These terms can encompass known analogs of natural nucleotides, as well as nucleotides modified in base, sugar, and / or phosphate moieties (e.g., phosphorothioate backbones). In general, unless otherwise indicated, analogs of a particular nucleotide have the same base pairing specificity; i.e., an analog of A will base pair with T.
[0032] As used herein, the terms "polypeptide," "peptide," and "protein" are used interchangeably to refer to a polymer of amino acid residues. The terms also apply to amino acid polymers in which one or more amino acids are chemical analogs or modified derivatives of corresponding naturally occurring amino acids.
[0033] As used herein, "CRISPR" refers to clustered regularly interspaced short palindromic repeats or any DNA locus used to guide CRISPR-associated proteins or similar nucleotide-guided nucleases. It also describes artificial, constructed, or selected systems derived using these frameworks or proteins. CRISPR systems and associated proteins vary among the currently described Type I, Type II, and Type III systems, although other similar systems may not yet be described.
[0034] As used herein, "CRISPR system" collectively refers to transcripts and other elements involved in the expression or directing the activity of CRISPR-associated ("Cas") genes, including sequences encoding Cas genes, tracr (trans-activating CRISPR) sequences (e.g., tracrRNA or active partial tracrRNA), tracr-pairing sequences (encompassing "direct repeats" and tracrRNA-processed partial direct repeats in the context of endogenous CRISPR systems), guide sequences (also referred to as "spacers" in the context of endogenous CRISPR systems), and other sequences and transcripts from the CRISPR locus. One or more tracr-pairing sequences (e.g., direct repeat-spacer-direct repeat) operably linked to a guide sequence may also be referred to as pre-crRNA (precursor CRISPR RNA) before treatment or as crRNA after nuclease treatment. CRISPR systems may also include modified, exchanged, or engineered, guide, tracr, or chimeric RNA sequences and their interacting proteins (e.g., Briner et al., Mal Cell 56(2):333-9 (2014)). The methods disclosed herein may also be applicable to other non-CRISPR nucleotide-guided nucleases.
[0035] As used herein, the term "guide sequence" refers to a portion of, for example, a guide RNA (gRNA) or a single guide RNA (sgRNA), that confers specificity to a nucleic acid guide nuclease for its target and mediates the formation of an RNA-DNA duplex between the targeting RNA and the target DNA sequence. For example, the targeting specificity of the CRISPR-Cas9 complex is determined by the approximately 20 nt sequence at the 5' end of the gRNA. The length of the guide sequence is typically between 17-24 bp. As used herein, the "center of the guide sequence" refers to the midpoint of the guide sequence. For example, if the guide sequence is 20 nucleotides in length, the "center of the guide sequence" will be located between nucleotides 10 and 11, numbered from the 5' end of the guide sequence to the 3' end of the guide sequence.
[0036] As used herein, the term "cleavage" or "cleaving" of nucleic acid refers to the rupture of the covalent backbone of a nucleic acid molecule. Cutting can be initiated by a variety of methods, including but not limited to enzymatic or chemical hydrolysis of phosphodiester bonds. Both single-stranded and double-stranded cleavage are possible, and double-stranded cleavage may be the result of two different single-stranded cleavage events. DNA cleavage can result in the production of blunt ends or staggered "sticky" ends. In certain embodiments, cutting refers to double-stranded cleavage between nucleic acids within a double-stranded DNA or RNA chain.
[0037] As used herein, the terms "genomic region" or "genomic fragment," used interchangeably herein, refer to a contiguous length of nucleotides in the genome of an organism. A genomic region can be as small as a few kb (e.g., at least 5 kb, at least 10 kb, or at least 20 kb) in length, up to an entire chromosome, or larger.
[0038] In some cases, the nucleotide sequence is provided using the character representation or a subset thereof recommended by the International Union of Pure and Applied Chemistry (IUPAC). The IUPAC nucleotide code used herein includes, A=adenine, C=cytosine, G=guanine, T=thymine, U=uracil, R=A or G, Y=C or T, S=G or C, W=A or T, K=G or T, M=A or C, B=C or G or T, D=A or G or T, H=A or C or T, V=A or C or G, N=any base, "." or "-" = space. In some embodiments, the set {A, C, G, T, U} is used for adenosine, cytidine, guanosine, thymidine and uridine, respectively. In some embodiments, the set {A, C, G, T, U, I, X, W} is used for adenosine, cytidine, guanosine, thymidine, uridine, inosine, uridine, xanthine nucleoside, pseudouridine, respectively. In some embodiments, the character set is {A, C, G, T, U, I, X, P, R, Y, N} for adenosine, cytidine, guanosine, thymidine, uridine, inosine, uridine, xanthosine, pseudouridine, unspecified purine, unspecified pyrimidine, and unspecified nucleotide, respectively. The modified sequence, non-natural sequence, or sequence with a modification incorporated therein can be in the genome, guide, or tracr sequence.
[0039] Nucleotide and / or amino acid sequence identity percentage (%) is understood to be the nucleotide or amino acid residue percentage of the identical nucleotide or amino acid residue in the candidate sequence compared to the reference sequence when two sequences are compared. In order to determine the identity percentage, the sequences are aligned, and if necessary, gaps are introduced to achieve maximum sequence identity percentage. Sequence alignment programs that determine identity percentage are well known to those skilled in the art. Generally, publicly available computer software such as BLAST, BLAST2, ALIGN2 or MEGALIGN (DNASTAR) software are used to align sequences. Those skilled in the art can determine the appropriate parameters for measuring alignment, including any algorithm that needs to achieve maximum alignment over the full length of the compared sequences. When aligning sequences, the sequence identity percentage of a given sequence A and a given sequence B (which can also be expressed as a given sequence A having or comprising a specific sequence identity percentage with a given sequence B) can be calculated as: sequence identity percentage=X / Y100, where X is the number of residues for which the alignment of A and B is scored as identical matches by a sequence alignment program or algorithm, and Y is the total number of residues in B. If the length of sequence A is not equal to the length of sequence B, then the percent sequence identity of A to B will not be equal to the percent sequence identity of B to A. Mismatches can similarly be defined as differences between the natural binding partners of a nucleotide. The number, position, and type of mismatches can be calculated and used for identification or ranking purposes.
[0040] As used herein, "mutation" encompasses any change in a DNA, RNA, or protein sequence from the wild-type sequence or some other reference, including but not limited to point mutations, transitions, insertions, transversions, translocations, deletions, inversions, duplications, recombination, or combinations thereof.
[0041] As used herein, in the context of a polynucleotide sequence cleaved by an RNA-guided nuclease, the term "insertion" is used when the polynucleotide sequence has one or more additional bases compared to the polynucleotide sequence before cleavage by the RNA-guided nuclease. Similarly, in the context of a polynucleotide sequence cleaved by an RNA-guided nuclease, the term "deletion" is used when the polynucleotide sequence has one or more missing bases compared to the polynucleotide sequence before cleavage by the RNA-guided nuclease. In the context of a polynucleotide sequence cleaved by an RNA-guided nuclease, the term "indel" refers to an insertion or a deletion. Cleavage by an RNA-guided nuclease can result in multiple indels, multiple insertions, multiple deletions, or a combination of insertions of one or more nucleotides and deletions of one or more nucleotides.
[0042] CRISPR
[0043] CRISPR (clustered regularly spaced short palindromic repeats) is an acronym for a DNA locus comprising multiple short direct repeats of a base sequence. Prokaryotic CRISPR / Cas systems have been adapted for use as gene editing (silencing, enhancing or changing specific genes) for eukaryotic organisms (see, for example, Cong, Science [Science], 15: 339 (6121): 819-823 (2013) and Jinek et al., Science [Science], 337 (6096): 816-21 (2012)). By transfecting cells with the desired elements (including cas genes and specially designed CRISPR), polynucleotide sequences can be cut and modified at almost any desired position by, for example, unique targeting of guide RNAs that confer nuclease specificity. There are a variety of methods for expressing guide chains or Cas proteins, including inducible expression of one or both. There are a variety of methods for introducing guide strands and Cas proteins into cells, including viral transduction, injection or microinjection, nanoparticle or other delivery, protein uptake, RNA or DNA uptake, protein and RNA or DNA combination uptake. Combinations of methods can also be used simultaneously or sequentially. Multiple rounds of delivery of RNA, DNA or protein can occur with or without further protein expression. Methods for preparing compositions for genome editing using the CRISPR / Cas system are described in detail in WO 2013 / 176772 and WO 2014 / 018423, which are hereby incorporated herein by reference in their entirety.
[0044] Nuclease
[0045] In some embodiments, the nuclease used in the methods described herein is a Class 2 Cas nuclease. In some embodiments, the nuclease has double-stranded endonuclease activity. In some embodiments, the nuclease includes a Cas nuclease, such as a Class 2 Cas nuclease (which can be, for example, a Type II, V, or VI Cas nuclease). Class 2 Cas nucleases include, for example, Cas9, Cpf1, C2cl, C2c2, and C2c3 proteins and modifications thereof.
[0046] Examples of Cas9 nucleases include nucleases of type II CRISPR systems of Streptococcus pyogenes, Staphylococcus aureus, and other prokaryotes (see, e.g., the list in the next paragraph) and modified (e.g., engineered or mutated) versions thereof. See, e.g., US 2016 / 0312198 A1; US 2016 / 0312199 A1. Figure 1AA type II nuclease is shown, characterized by interaction of an sgRNA with a target DNA sequence. The PAM (NGG) sequence, guide sequence, target DNA sequence, and cleavage site are indicated. Other examples of Cas nucleases include the Csm or Cmr complexes of type III CRISPR systems, or their Cas 10, Csml, or Cmr2 subunits; and the Cascade complex of type I CRISPR systems, or their Cas3 subunits. Figure 1B A V-type nuclease is shown, characterized by the interaction of a gRNA with a target DNA sequence. The PAM (TTTN) sequence, guide sequence, target DNA sequence, and cleavage site are indicated. In some embodiments, the Cas nuclease can be from a type IIA, type IIB, or type IIC system. For discussion of various CRISPR systems and Cas nucleases, see, for example, Makarova et al., Nat. Rev. Microbiol. 9:467-477 (2011); Makarova et al., Nat. Rev. Microbiol., 13:722-36 (2015); Shmakov et al., Molecular Cell, 60:385-397 (2015). In some embodiments, the RNA-guided DNA binder is a Cas nickase, such as a Cas9 nickase. In some embodiments, the RNA-guided DNA binder is a Streptococcus pyogenes Cas9 nuclease.
[0047] Non-limiting exemplary species from which nucleases (e.g., Cas nucleases) can be derived include, but are not limited to, Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp.), Staphylococcus aureus, Listeria innocua, Lactobacillus gasseri, Francisella novicida, Wolinella succinogenes, Sutterella wadsworthensis, Gammaproteobacterium, Neisseria meningitidis, Campylobacter Jejuni, Pasteurella multocida, Fibrobacter succinogene, Rhodospirillum rubrum, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis), Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Lactobacillus buchneri, Treponema denticola, Microscilla marina, and Burkholderiales bacterium), naphthalene-degrading Polaromonas naphthalenivorans, Polaromonas sp.), Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans ferrooxidans), Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleuschthonoplastes, Oscillatoria sp.), Petrotoga mobilis, Thermosipho africanus, Streptococcus pasteurianus, Neisseria cinerea, Campylobacter Zari, Parvibaculum lavamentivorans, Corynebacterium diphtheria, Acidaminococcus sp., Lachnospiraceae bacterium ND2006, and Acaryochloris marina.
[0048] In some embodiments, the Cas nuclease is a Cas9 nuclease from Streptococcus pyogenes. In some embodiments, the Cas nuclease is a Cas9 nuclease from Streptococcus thermophilus. In some embodiments, the Cas nuclease is a Cas9 nuclease from Neisseria meningitidis. In some embodiments, the Cas nuclease is a Cas9 nuclease from Staphylococcus aureus. In some embodiments, the Cas nuclease is a Cpf1 nuclease from Francisella novicida. In some embodiments, the Cas nuclease is a Cpf1 nuclease from an Acidaminococcus species. In some embodiments, the Cas nuclease is a Cpf1 nuclease from the Lachnospiraceae bacterium ND2006. In some embodiments, the Cas nuclease is a Cpf1 nuclease from Francisella tularensis, Lachnospiraceae, Butyrivibrioproteoclasticus, Peregrinibacteria bacterium, Parcubacteria bacterium, Smithella, Acidaminococcus, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi, Leptospira inadai, Porphyromonas crevioricanis, Prevotella disiens, or Porphyromonas macacae. In some embodiments, the Cas nuclease is a Cpf1 nuclease from Acidaminococcus or Lachnospiraceae.
[0049] Wild-type Cas9 has two nuclease domains: RuvC and HNH. The RuvC domain cleaves non-target DNA strands, and the HNH domain cleaves target DNA strands. In some embodiments, the Cas9 nuclease comprises more than one RuvC domain and / or more than one HNH domain. In some embodiments, the Cas9 nuclease is wild-type Cas9. In some embodiments, Cas9 is capable of inducing double-strand breaks in target DNA. In certain embodiments, Cas nucleases can cleave one or both strands of dsDNA. In some embodiments, Cas nucleases can cleave single strands of DNA. In some embodiments, Cas nucleases may not have DNA nickase activity.
[0050] In some embodiments, chimeric Cas nucleases are used in which one domain or region of a protein is replaced by a portion of a different protein. In some embodiments, the Cas nuclease domain can be replaced with a domain from a different nuclease, such as Fok 1. In some embodiments, the Cas nuclease can be a modified nuclease in which the polypeptide sequence of the nuclease has been modified to confer advantageous properties on the nuclease in some instances. In some embodiments, the cleavage site of the nuclease relative to the position of the user-specified target sequence is unknown.
[0051] In other embodiments, the Cas nuclease may be from a type I CRISPR / Cas system. In some embodiments, the Cas nuclease may be a component of the Cascade complex of a type I CRISPR / Cas system. In some embodiments, the Cas nuclease may be a Cas3 protein. In some embodiments, the Cas nuclease may be from a type III CRISPR / Cas system. In some embodiments, the Cas nuclease may have RNA cleavage activity.
[0052] Data processing flow
[0053] In some embodiments, a data processing flow is disclosed herein that includes the following steps: receiving sample sequence data; applying quality control filters to the sample sequence data; aligning sequencing reads of the sample sequence data with a reference sequence; defining a window based on the sequence position within the nucleic acid guide and the position of the end of the nucleic acid guide sequence; and quantifying indels in the sample sequence data. In some embodiments, the plurality of sequencing reads are obtained from a next generation sequencing (NGS) instrument. In some embodiments, the NGS sequencing instrument is an Illumina MiSeq TM In some embodiments, sequencing and amplification adapters are pruned by default, so that the returned NGS read data does not have adapter sequences in the reads. In some embodiments, sequencing and amplification adapters are pruned as part of the data processing flow of the methods disclosed herein.
[0054] Figure 2A The steps of an upstream molecular biology workflow are shown, which can be performed to generate data that is subsequently processed by the data processing pipeline disclosed herein. The steps of the upstream molecular biology workflow can include designing primers to amplify editing targets; isolating cellular DNA; amplifying editing targets using adapter PCR; adding next-generation sequencing (NGS) adapters to PCR products; mixing barcoded amplicons; bead purification of amplicons; spiking diversity DNA; and performing next-generation sequencing (NGS) on a sequencing instrument.
[0055] Figure 2BThe steps of the computational workflow disclosed herein are shown. The steps may include performing quality control on received sequence reads; aligning the sequence reads to the target sequence; and quantifying indels according to the methods disclosed herein. These steps are described in more detail below.
[0056] Previously available computational methods can underestimate the frequency of indels in a target polynucleotide following nuclease cleavage of a nucleic acid guide, for example, following cleavage with a type II nuclease with a known cleavage site relative to the guide sequence. Figure 3A Figure 2 is a schematic diagram of received sequencing reads for a nucleic acid-guided nuclease-edited sample, comprising a plurality of sequencing reads, a subset of which contain indels resulting from nucleic acid-guided nuclease cleavage. The nuclease is Cas9, a type II nuclease with a known cleavage site relative to the guide sequence. For each type of indel, the number of sequencing reads containing that indel is indicated to the right. The previously known cleavage site is indicated by a rectangular box.
[0057] Previously available computational methods can underestimate the frequency of indels in a target polynucleotide following nuclease cleavage of a nucleic acid guide, for example, following cleavage with a V-type nuclease with a known cleavage site relative to the guide sequence. Figure 3B This figure shows a schematic representation of the sequence reads received for a nucleic acid-guided nuclease-edited sample, including multiple sequence reads, a subset of which contain indels resulting from nucleic acid-guided nuclease cleavage. The nuclease is a novel V-type nuclease. For each type of indel, the number of sequence reads containing that indel is indicated on the right. Several indels were missed by standard computational processing methods.
[0058] The computational pipeline disclosed herein and described in further detail below improves the accuracy of quantifying indel frequencies in a target polynucleotide following cleavage by a nucleic acid-guided nuclease. Figure 3C This figure shows a schematic representation of received sequencing reads for a nucleic acid-guided nuclease-edited sample. The sample contains multiple sequencing reads, a subset of which contains indels resulting from nucleic acid-guided nuclease cleavage. The nuclease is a novel V-type nuclease. For each type of indel, the number of sequencing reads containing that indel is indicated to the right. A custom data processing method captures all indels generated by nuclease cleavage.
[0059] In some embodiments, the data processing flows disclosed herein include one or more modules implemented by the CRISPResso software toolkit (Clement K, et al. CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat Biotechnol. 2019 Mar;37(3):224-226., and Canver MC, et al. Integrated design, execution, and analysis of arrayed and pooled CRISPR genome-editing experiments. Nat Protoc. 2018 May;13(5):946-986., which are incorporated herein by reference in their entireties). In some embodiments, CRISPResso is used to apply one or more read filtering parameters to the sample sequence data to remove potential false positive indels from the sample sequence data, thereby improving the accuracy of indel frequency estimates in the sample sequence data. The read filtering parameters are described in more detail below. In some embodiments, read filtering is performed based on PHRED quality scores, which are described in: Ewing B, Green P. Base-calling of automated sequencertraces using phred.II. Error probabilities. Genome Res. 1998 Mar;8(3):186-94, which is incorporated herein by reference in its entirety. PHRED quality scores measure the quality of identification of nucleotide base calls in sequence reads generated by automated DNA sequencers.
[0060] In some embodiments, a minimum average read quality ("q" or "min_average_read_quality") is applied to filter the sample sequence data to remove potential false positive indels. This parameter allows for the specification of a minimum average quality score for including reads in subsequent analysis. The PHRED score represents the confidence level in the assignment of a particular nucleotide in a read. A maximum score of 40 corresponds to an error rate of 0.01%. This average read quality is used to filter out low-quality reads. In some embodiments, a "min_average_read_quality" value of 0, 5, 10, 15, 20, 25, 30, 35, or 40 is applied to the sample sequence data.
[0061] In some embodiments, a minimum single base pair score ("s" or "min_single_bp_quality") is applied to filter the sample sequence data to remove potential false positive indels. This parameter allows for the specification of a minimum single base pair score for including a read in subsequent analysis. This parameter provides stricter filtering; any reads with a single base pair quality below the threshold are discarded. In some embodiments, a "min_single_bp_quality" value of 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 is applied to the sample sequence data.
[0062] In some embodiments, an amplicon minimum alignment score ("amas" or "amplicon_min_alignment_score") is applied to filter the sample sequence data. After the reads are aligned to the reference sequence, homology is calculated as the number of base pairs they share. This is used to filter erroneous reads that do not align to the target sequence, such as those caused by alternative primer positions. In some embodiments, an "amplicon_min_alignment_score" value of 0, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 is applied to the sample sequence data.
[0063] In some embodiments, a combination of the "min_average_read_quality" score, the "min_single_bp_quality" score, and the "amplicon_min_alignment_score" is applied to filter sample sequence data to improve the estimate of indel frequency in the sample sequence data after nuclease cleavage of the target polynucleotide sequence. In some embodiments, a combination of the "min_average_read_quality" score, the "min_single_bp_quality" score, and the "amplicon_min_alignment_score" is applied, further combined with a window size based on sequence position within the guide sequence, to filter sample sequence data to improve the estimate of indel frequency in the sample sequence data after nuclease cleavage of the target polynucleotide sequence. As described in Table 1 below, the combination of these parameters can be used to enhance the accuracy of indel quantification according to the methods disclosed herein.
[0064] In some embodiments, a combination of the "min_average_read_quality" score, the "min_single_bp_quality" score, and the "amplicon_min_alignment_score" is applied to filter the sample sequence data to improve the estimate of the frequency of indels in the sample sequence data after the nuclease has cleaved the target polynucleotide sequence. In some embodiments, a combination of the "min_average_read_quality" score, the "min_single_bp_quality" score, and the "amplicon_min_alignment_score" is applied, further combined with a window size based on the sequence position of known insertion sites, to filter the sample sequence data to improve the estimate of the frequency of indels in the sample sequence data after the nuclease has cleaved the target polynucleotide sequence and inserted into the polynucleotide sequence. As described in Table 1 below, the combination of these parameters can be used to enhance the accuracy of indel quantification according to the methods disclosed herein.
[0065] In some embodiments, if the sample sequence data includes paired-end reads, the data processing workflow disclosed herein includes one or more modules ( T, Salzberg SL. FLASH: fast length adjustment of short reads to improve genome assemblies. Bioinformatics. 2011 Nov 1; 27(21): 2957-63., which is incorporated herein by reference in its entirety. FLASH is a fast and accurate software tool for merging end-paired reads from next-generation sequencing experiments, designed to merge read pairs when the original DNA fragments are shorter than twice the read length. The resulting longer reads can significantly improve genome assembly. FLASH calculates the mismatch rate within two overlapping regions. If the mismatch rate exceeds a threshold mismatch rate, FLASH determines that the read is an incorrect overlap. In some embodiments, FLASH is used to merge end-paired reads to produce a single read for alignment to a target reference sequence and to reduce sequencing errors that may be present at the ends of sequenced reads.
[0066] In some embodiments, for the FLASH read merging step of the data processing workflow, a maximum paired end read overlap ("max_paired_end_reads_overlap") is applied to the paired end sample sequence data. This parameter represents the maximum overlap length expected in about 90% of the read pairs. In some embodiments, a "max_paired_end_reads_overlap" value of 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, or 300 is applied to the sample sequence data.
[0067] In some embodiments, for the FLASH read merging step of the data processing workflow, a minimum paired end read overlap ("min_paired_end_reads_overlap") is applied to the paired end sample sequence data. This parameter represents the minimum overlap length required between two reads to provide a reliable overlap. In some embodiments, a "max_paired_end_reads_overlap" value of 0, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 is applied to the sample sequence data.
[0068] In some embodiments, a combination of the "min_average_read_quality" score, the "min_single_bp_quality" score, the "min_paired_end_reads_overlap," and the "max_paired_end_reads_overlap" is applied to filter sample sequence data to improve the estimate of indel frequency in the sample sequence data after nuclease cleavage of the target polynucleotide sequence. In some embodiments, a combination of the "min_average_read_quality" score, the "min_single_bp_quality" score, the "min_paired_end_reads_overlap," and the "max_paired_end_reads_overlap" is applied, further combined with a window size based on sequence position within the guide sequence, to filter sample sequence data to improve the estimate of indel frequency in the sample sequence data after nuclease cleavage of the target polynucleotide sequence. As described in Table 1 below, combinations of these parameters can be used to enhance the accuracy of indel quantification according to the methods disclosed herein.
[0069] In some embodiments, a combination of the "min_average_read_quality" score, the "min_single_bp_quality" score, the "min_paired_end_reads_overlap" score, and the "max_paired_end_reads_overlap" score is applied to filter the sample sequence data to improve the estimate of the frequency of indels in the sample sequence data after the nuclease cleaves the target polynucleotide sequence. In some embodiments, a combination of the "min_average_read_quality" score, the "min_single_bp_quality" score, the "min_paired_end_reads_overlap" score, and the "max_paired_end_reads_overlap" score is applied, further combined with a window size based on the sequence position of known insertion sites, to filter the sample sequence data to improve the estimate of the frequency of indels in the sample sequence data after the nuclease cleaves the target polynucleotide sequence and inserts into the polynucleotide sequence. As described in Table 1 below, the combination of these parameters can be used to enhance the accuracy of indel quantification according to the methods disclosed herein.
[0070] In some embodiments, the quantitative window extends 50, 40, 30, 20, 10, or 5 nucleotides from 5' to the 5' end of the guide sequence. In some embodiments, the quantitative window extends 50, 40, 30, 20, 10, or 5 nucleotides from 3' to the 3' end of the guide sequence. In some embodiments, the quantitative window extends 50, 40, 30, 20, 10, or 5 nucleotides from 5' to a known insertion site. In some embodiments, the quantitative window extends 50, 40, 30, 20, 10, or 5 nucleotides from 3' to a known insertion site. In some embodiments, the length of the quantitative window is 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, or 500 nucleotides.
[0071] In certain embodiments, the length of the insertion / deletion to be detected is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides. In certain embodiments, the length of the insertion / deletion to be detected is about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more nucleotides. In certain embodiments, the insertion / deletion comprises a CRISPR-mediated donor insertion. In certain embodiments, the insertion / deletion is the result of homology-directed repair (HDR). In certain embodiments, the insertion / deletion is the result of CRISPR-associated transposase (CAST). In certain embodiments, the insertion comprises a sequence from a genomic library.
[0072] Table 1:
[0073]
[0074]
[0075]
[0076]
[0077]
[0078]
[0079]
[0080]
[0081]
[0082]
[0083]
[0084]
[0085]
[0086]
[0087] Computer implementation of the method
[0088] Figure 5 5 is a diagram of components of a computer system 500 that can be used to implement a computational process for quantifying indels in a polynucleotide sequence that has been cleaved by a nucleic acid-guided nuclease (including uncharacterized nucleases). The computer system 500 can be used to implement methods that include parameters for aligning next-generation sequencing (NGS) reads from targeted amplicon sequencing (TAS) of a nuclease-edited sample or a control sample that has been obtained and aligned to a reference amplicon sequence.
[0089] Computing device 500 is intended to represent various forms of digital computers (such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers). Computing device 550 is intended to represent various forms of mobile devices (such as personal digital assistants, mobile phones, smartphones, and other similar computing devices). In addition, computing device 500 or 550 may include a universal serial bus (USB) flash drive. A USB flash drive can store an operating system and other applications. A USB flash drive may include input / output components, such as a wireless transmitter or a USB connector that can be plugged into a USB port of another computing device. The components shown herein, their connections and relationships, and their functions are intended to be exemplary only and are not intended to limit the implementation of the methods and compositions described and / or claimed in this document.
[0090] Computing device 500 includes a processor 502, a memory 504, a storage device 506, a high-speed interface 508 connected to the memory 504 and a high-speed expansion port 510, and a low-speed interface 512 connected to a low-speed bus 514 and the storage device 506. The various components 502, 504, 506, 508, 510, and 512 are interconnected using a variety of buses and can be mounted on a shared motherboard or otherwise as appropriate. The processor 502 can process instructions executed within the computing device 500 (including instructions stored in the memory 504 or on the storage device 506) to display graphical information for a GUI on an external input / output device (such as a display 516 coupled to the high-speed interface 508). In other embodiments, multiple processors and / or multiple buses, as well as multiple memories and multiple memory types, can be used as appropriate. In addition, multiple computing devices 500 can be connected, each device providing a portion of the necessary operations (for example, as a server bank, a blade server group, or a multi-processor system).
[0091] Memory 504 stores information within computing device 500. In one embodiment, memory 504 is one or more volatile memory units. In another embodiment, memory 504 is one or more non-volatile memory units. Memory 504 may also be another form of computer-readable medium (e.g., a magnetic disk or optical disk).
[0092] Storage device 506 can provide mass storage for computing device 500. In one embodiment, storage device 506 can be or contain a computer-readable medium, such as a floppy disk drive, a hard disk drive, an optical disk drive, or a magnetic tape drive, a flash memory or other similar solid-state memory device or device array (including a storage area network or other configured device). A computer program product can be tangibly embodied in an information carrier. A computer program product can also include instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable medium or machine-readable medium (such as memory 504, storage device 506, or memory on processor 502).
[0093] High-speed controller 508 manages bandwidth-intensive operations of computing device 500, while low-speed controller 512 manages less bandwidth-intensive operations. This functional allocation is merely an example. In one embodiment, high-speed controller 508 is coupled to memory 504, display 516 (e.g., via a graphics processor or accelerator), and high-speed expansion ports 510, which can accept a variety of expansion cards (not shown). In this embodiment, low-speed controller 512 is coupled to storage device 506 and low-speed expansion ports 514. The low-speed expansion ports, which can include a variety of communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices (e.g., a keyboard, a pointing device, a microphone / speaker combination, a scanner, or a networking device (e.g., a switch or router)) via a network adapter. Computing device 500 can be implemented in a number of different forms, as shown in the figure. For example, it can be implemented as a standard server 520, or multiple implementations of such servers can be implemented. It can also be implemented as part of a rack-mounted server system 524. Furthermore, it can be implemented in a personal computer such as laptop computer 522. Alternatively, components from computing device 500 may be combined with other components in a mobile device (not shown), such as device 550. Each such device may contain one or more of computing devices 500, 550, and the entire system may be comprised of multiple computing devices 500, 550 in communication with each other.
[0094] Computing device 500 can be implemented in a number of different forms, as shown in the figure. For example, it can be implemented as a standard server 520, or multiple implementations of a group of such servers. It can also be implemented as part of a rack-mounted server system 524. In addition, it can be implemented in a personal computer such as laptop computer 522. Alternatively, components from computing device 500 can be combined with other components in a mobile device (not shown) such as device 550. Each such device can contain one or more of computing devices 500, 550, and the entire system can be composed of multiple computing devices 500, 550 in communication with each other.
[0095] Computing device 550 includes a processor 552, a memory 564 and input / output devices (such as a display 554), a communication interface 566 and a transceiver 568, among other components. Device 550 may also be equipped with a storage device (such as a microdrive or other device) to provide additional storage. The various components 550, 552, 564, 554, 566, and 568 are interconnected using a variety of buses, and several components may be mounted on a shared motherboard or in other ways as appropriate.
[0096] The processor 552 can execute instructions within the computing device 550 (including instructions stored in the memory 564). The processor can be implemented as a chipset including single or multiple analog processors and digital processors. In addition, the processor can be implemented using any of a variety of architectures. For example, the processor 510 can be a CISC (Complex Instruction Set Computer) processor, a RISC (Reduced Instruction Set Computer) processor, or a MISC (Minimum Instruction Set Computer) processor. The processor can provide, for example, coordination of other components of the device 550 (such as control of a user interface, applications executed by the device 550, and wireless communications performed by the device 550).
[0097] The processor 552 can communicate with the user through a control interface 558 and a display interface 556 coupled to the display 554. The display 554 can be, for example, a TFT (thin film transistor liquid crystal display) display or an OLED (organic light emitting diode) display, or other appropriate display technology. The display interface 556 can include appropriate circuitry for driving the display 554 to present graphics and other information to the user. The control interface 558 can receive commands from the user and convert them for submission to the processor 552. In addition, an external interface 562 can be provided to communicate with the processor 552 so that the device 550 can perform near-field communication (near area communication) with other devices. The external interface 562 can provide, for example, wired communication in some embodiments, or wireless communication in other embodiments, and multiple interfaces can also be used.
[0098] Memory 564 stores information within computing device 550. Memory 564 can be implemented as one or more computer-readable media, one or more volatile memory units, or one or more non-volatile memory units. Expansion memory 574 can also be provided and connected to device 550 via expansion interface 572, which can include, for example, a SIMM (Single In-line Memory Module) card interface. Such expansion memory 574 can provide additional storage space for device 550, or can also store applications or other information for device 550. In particular, expansion memory 574 can include instructions for executing or supplementing the processes described above, and can also include security information. Thus, for example, expansion memory 574 can be provided as a security module for device 550 and can be programmed with instructions that allow secure use of device 550. In addition, secure applications and additional information (such as placing identification information on a SIMM card in an unbreakable manner) can be provided via a SIMM card.
[0099] The memory may include, for example, flash memory and / or NVRAM memory, as discussed below. In one embodiment, the computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable medium or machine-readable medium, such as memory 564, expansion memory 574, or memory on processor 552 that can be received, for example, via transceiver 568 or external interface 562.
[0100] The device 550 can communicate wirelessly via a communication interface 566, which may include digital signal processing circuitry, if necessary. The communication interface 566 can provide communication in various modes or protocols, such as GSM voice calls, SMS, EMS, or MMS messaging services, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS. Such communication can occur, for example, via a radio frequency transceiver 568. In addition, short-range communication can be performed using, for example, Bluetooth, Wi-Fi, or other such transceivers (not shown). In addition, a GPS (Global Positioning System) receiver module 570 can provide additional navigation-related and positioning-related wireless data to the device 550, which can be used as appropriate by applications running on the device 550.
[0101] Device 550 may also communicate audibly using an audio codec 560 that receives voice information from a user and converts it into usable digital information. Audio codec 560 may also generate audible sounds for the user, such as through a speaker (e.g., in the handheld portion of device 550). Such sounds may include sounds from voice phone calls, may include recorded sounds (e.g., voice messages, music files, etc.), and may also include sounds generated by applications operating on device 550.
[0102] Computing device 550 can be implemented in a number of different forms, as shown in the figure. For example, it can be implemented as a mobile phone 580. It can also be implemented as part of a smart phone 582, a personal digital assistant, or other similar mobile device.
[0103] Various embodiments of the systems and methods described herein can be implemented in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations of such embodiments. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system that includes at least one programmable processor that can be coupled for special or general purposes to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.
[0104] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or in assembly / machine language. As used herein, the terms "machine-readable medium," "computer-readable medium," and "machine-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0105] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input from the user can be received in any form, including auditory, voice, or tactile input.
[0106] The systems and techniques described herein can be implemented in a computing system that includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer with a graphical user interface or a web browser through which a user can interact with an embodiment of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any digital data communication form or medium (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), and the Internet.
[0107] A computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0108] A number of embodiments have been described. However, it will be appreciated that numerous changes may be made without departing from the spirit and scope of the invention. Furthermore, the logic flows depicted in the accompanying drawings do not require the particular order shown or ordered sequence to achieve the desired results. Additionally, other steps may be provided or eliminated from the described flows, and other components may be added to or removed from the described systems. Accordingly, other embodiments are within the scope of the appended claims.
[0109] The embodiments of the present disclosure and all functional operations described in this specification can be implemented in digital electronic circuit systems, or in computer software, firmware or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them. The embodiments of these methods and compositions can be implemented as one or more computer program products, for example, one or more modules of computer program instructions encoded on a computer-readable medium, for being executed by a data processing device or controlling the operation of the data processing device. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a material composition that affects a machine-readable propagation signal, or a combination of one or more of them. The term "data processing device" covers all devices, devices and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the device can also include code that creates an execution environment for the computer program in question, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagation signal is an artificially generated signal, such as an electrical signal, an optical signal, or an electromagnetic signal generated by a machine, the purpose of which is to encode information for transmission to a suitable receiver device.
[0110] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing multiple portions of one or more modules, subroutines, or code). A computer program can be deployed to be executed on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0111] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform the functions of operating on input data and generating output. These processes and logic flows can also be performed by a device, and the device can also be implemented as a special-purpose logic circuit, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0112] Processors suitable for executing computer programs include, by way of example only, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or be operatively coupled to receive data from or transmit data to or from them, or both. However, a computer need not have such devices. Furthermore, a computer may be embedded in another device, such as a tablet computer, mobile phone, personal digital assistant (PDA), mobile audio player, Global Positioning System (GPS) receiver, etc. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CDROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0113] To provide interaction with a user, embodiments of the present disclosure may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form, including auditory, voice, or tactile input.
[0114] Embodiments of the present disclosure may be implemented in a computing system that includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer with a graphical user interface or a web browser through which a user can interact with an embodiment of the method), or any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected via any digital data communication form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), such as the Internet.
[0115] A computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0116] Although this specification contains many details, these should not be interpreted as limitations on the scope of the invention or the scope that may be claimed, but rather as descriptions of features specific to particular embodiments of the invention. Certain features described in this specification in the context of separate embodiments may also be implemented in a single embodiment in combination. Conversely, different features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable sub-combination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed as such, one or more features from a claimed combination may in some cases be deleted from that combination, and a claimed combination may involve sub-combinations or variations of sub-combinations.
[0117] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in an ordered order, or that all of the operations shown can be performed to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. Moreover, the separation of different system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0118] In each instance where an HTML file is mentioned, other file types or formats may be substituted. For example, an HTML file may be replaced by XML, JSON, plain text, or other types of files. Additionally, where a table or hash table is mentioned, other data structures (such as a spreadsheet, a relational database, or a structured file) may be used.
[0119] Examples
[0120] The present invention is further illustrated in the following examples, which do not limit the scope of the invention described in the claims.
[0121] Example 1: Benchmarking Indel Quantification Methods Using Well-Characterized Nucleases
[0122] To evaluate the efficacy of the indel quantification method disclosed herein, the data processing pipeline was applied to data generated by next-generation sequencing of mammalian cells edited with SpCas9 (a type II nuclease with known cleavage sites and editing patterns). The percentage of cells containing indels on the target DNA sequence after editing was estimated using standard CRISPResso2 processing parameters and the methods disclosed herein. Figure 4A As shown, the methods disclosed herein perform similarly to standard CRISPResso2 processing parameters, indicating that the method does not over- or underestimate the percentage of cells that contain indels at the target DNA sequence after editing with well-characterized nucleases with known cleavage sites and editing patterns.
[0123] Example 2: Dilution experiment to quantify indels
[0124] To validate the accuracy of the indel quantification method disclosed herein, the data processing pipeline was applied to data generated by next-generation sequencing of mammalian cells edited with V-type nucleases whose cleavage sites were unknown in a serial dilution experiment. DNA isolated from mammalian cells edited with V-type nucleases was serially diluted with DNA isolated from unedited cells at a ratio of, for example, 0% unedited / 100% edited; 25% unedited / 75% edited; 50% unedited / 50% edited; 75% unedited / 25% edited; 0% unedited / 100% edited ( Figure 4B ). The serially diluted DNA mixture is sequenced to generate sequencing reads containing the edited target sequence. Figure 4C As shown, for each serial dilution concentration, the observed indel percentage (y-axis) is plotted against the expected indel percentage (x-axis). The linearity of the plot (R-squared value of approximately 0.99) confirms the efficacy of the indel quantification data processing pipeline.
[0125] Example 3: Quantification of indels after editing with a novel nuclease
[0126] To validate the efficacy of the indel quantification method disclosed herein for estimating indels in target sequences edited by novel nucleases and different guide sequences, the data processing pipeline was applied to data generated by next-generation sequencing of mammalian cells edited by the novel V-type nuclease at increasing concentrations and using two different guides: guide sequences A1 and F2. Figure 6 As shown, the data processing pipeline accurately estimates the percentage of indels edited by the novel V-type nuclease using guide sequences A1 and F2 at nuclease concentrations of 5 μM, 11 μM, and 22 μM. In addition, the data processing pipeline was applied to data generated by next-generation sequencing of mammalian cells edited by the nuclease AsCas12a to validate the efficacy of the indel quantification method disclosed herein.
[0127] Example 4: Dilution Experiment to Quantify Large Inserts
[0128] To verify the accuracy of the indel quantification method disclosed herein for quantifying larger insertions (i.e., insertions greater than about 10 nucleotides), the data processing pipeline was applied to data generated by next-generation sequencing of mammalian cells in a serial dilution spike-in experiment in which a sequence of known length, 21 nucleotides, was introduced at a specific insertion site (see Figure 7A DNA isolated from mammalian cells containing an insertion is serially diluted with DNA isolated from cells not containing an insertion at ratios such as 0% insertion / no insertion; 25% insertion / 75% no insertion; 50% insertion / 50% no insertion; 75% insertion / 25% no insertion; 0% insertion / 100% no insertion ( Figure 4B ). The serially diluted DNA mixture is sequenced to generate sequencing reads containing the insertion site. Figure 7B As shown, for the experiment containing the inserted DNA at 25% incorporation, the incorporation sequence was detected at approximately 23%. Figure 7C As shown, for an experiment containing 50% incorporation of inserted DNA, approximately 47% of the incorporation sequence was detected. Figure 7D As shown, for an experiment in which the DNA containing the insertion was incorporated at 75%, the incorporation sequence was detected at approximately 76%. These results demonstrate that the data processing pipeline disclosed herein can effectively quantify insertions introduced by, for example, CRISPR-mediated donor insertions, homology-directed repair (HDR)-mediated insertions, CRISPR-associated transposase (CAST) and other genome engineering methods, as well as insertions in genomic libraries used in high-throughput screening (HTS).
[0129] Other embodiments
[0130] It should be understood that although the invention has been described in conjunction with its detailed description, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages and modifications are within the scope of the following claims.
Claims
1. A method for quantifying insertions and / or deletions in a polynucleotide sequence caused by cleavage of a target polynucleotide sequence by a nucleic acid-guided nuclease, the method comprising: In a computer system, sample sequence data comprising a plurality of sequencing reads is received; filtering the plurality of sequencing reads; aligning the plurality of sequencing reads to a reference sequence; defining a window based on sequence position within the nucleic acid guide sequence and the position of the ends of the nucleic acid guide sequence; determining a number of sequencing reads comprising an insertion or deletion within the window relative to the reference sequence based on an alignment of each sequencing read in the plurality of sequencing reads to the reference sequence; Based on the number of sequencing reads comprising insertions or deletions, the number of insertions and / or deletions in the polynucleotide sequence mediated by the nucleotide-guided nuclease is estimated.
2. The method of claim 1, wherein the nucleic acid-guided nuclease is a Class 2 nuclease.
3. The method of any one of claims 1-2, wherein the nucleic acid-guided nuclease is a type II nuclease.
4. The method of claim 3, wherein the nucleic acid-guided nuclease is SpCas9.
5. The method of claim 3, wherein the nucleic acid-guided nuclease is AsCas12a.
6. The method of any one of claims 1-2, wherein the nucleic acid-guided nuclease is a V-type nuclease.
7. The method of any one of claims 1-6, wherein the nucleic acid guide is a guide RNA (gRNA).
8. The method of any one of claims 1-6, wherein the nucleic acid guide is a single guide RNA (sgRNA).
9. The method of any one of claims 1-8, wherein the center of the window is located at the center of the nucleic acid guide sequence.
10. The method of any one of claims 1-8, wherein the center of the window is located at the site where the nuclease cleaves the polynucleotide sequence.
11. The method of any one of claims 1-10, wherein the length of the window is equal to the length of the nucleic acid guide sequence.
12. The method of any one of claims 1-11, wherein the 5' end of the window extends 50 base pairs from 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 50 base pairs from 3' to the 3' end of the nucleic acid guide.
13. The method of any one of claims 1-12, wherein the 5' end of the window extends 40 base pairs from 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 40 base pairs from 3' to the 3' end of the nucleic acid guide.
14. The method of any one of claims 1-13, wherein the 5' end of the window extends 30 base pairs from 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 30 base pairs from 3' to the 3' end of the nucleic acid guide.
15. The method of any one of claims 1-14, wherein the 5' end of the window extends 20 base pairs from 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 20 base pairs from 3' to the 3' end of the nucleic acid guide.
16. The method of any one of claims 1-15, wherein the 5' end of the window extends 10 base pairs from 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 10 base pairs from 3' to the 3' end of the nucleic acid guide.
17. The method of any one of claims 1-16, wherein the 5' end of the window extends 5 base pairs from 5' to the 5' end of the nucleic acid guide, and the 3' end of the window extends 5 base pairs from 3' to the 3' end of the nucleic acid guide.
18. The method of any one of claims 1-17, wherein the method further comprises trimming adapter sequences from the plurality of sequencing reads.
19. The method of any one of claims 1-18, wherein the plurality of sequencing reads are generated by targeted amplicon sequencing.
20. The method of any one of claims 1-19, wherein the plurality of sequencing reads are generated by targeted amplicon sequencing of DNA isolated from a cell edited by the nucleic acid-guided nuclease.
21. The method of any one of claims 1-20, wherein the plurality of sequencing reads are paired-end sequencing reads.
22. The method of claim 21, wherein the method further comprises read merging of the paired-end sequencing reads to generate a single read for alignment to the reference sequence.
23. The method of claim 22, wherein read merging of the paired-end reads further comprises applying a minimum paired-end read overlap score to the plurality of sequencing reads.
24. The method of claim 23, wherein the minimum paired-end read overlap score is 0, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100.
25. The method of any one of claims 23-24, wherein the minimum paired-end read overlap score is 10.
26. The method of any one of claims 22-25, wherein read merging of the paired-end reads further comprises applying a maximum paired-end read overlap score to the plurality of sequencing reads.
27. The method of claim 26, wherein the maximum paired-end read overlap score is 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, or 300.
28. The method of any one of claims 26-27, wherein the maximum paired-end read overlap score is 100.
29. The method of any one of claims 1-28, wherein the filtering step further comprises applying a minimum average read quality score to the plurality of sequencing reads.
30. The method of claim 29, wherein the minimum average read quality score is 0, 5, 10, 15, 20, 25, 30, 35, or 40.
31. The method of any one of claims 1-30, wherein the filtering step further comprises applying a minimum single base pair score to the plurality of sequencing reads.
32. The method of claim 31 , wherein the minimum single base pair score is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20.
33. The method of any one of claims 1-32, wherein the aligning step further comprises applying an amplicon minimum alignment score to the plurality of sequencing reads.
34. The method of claim 33, wherein the amplicon minimum alignment score is 0, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100.
35. The method of any one of claims 1-34, wherein the nucleic acid-guided nuclease is FokI nuclease.
36. The method of claim 35, wherein the FokI nuclease is fused to a transcription activator-like (TAL) protein.
37. The method of any one of claims 1-36, wherein the nucleic acid-guided nuclease is a zinc finger nuclease.
38. A computer program product tangibly embodied on a computer-readable medium, comprising instructions that, when executed by one or more processors, are configured to: receiving sample sequence data comprising a plurality of sequencing reads; filtering the plurality of sequencing reads; aligning the plurality of sequencing reads to a reference sequence; defining a window based on sequence position within the nucleic acid guide sequence and the position of the ends of the nucleic acid guide sequence; determining a number of sequencing reads comprising an insertion or deletion within the window relative to the reference sequence based on an alignment of each sequencing read in the plurality of sequencing reads to the reference sequence; and Based on the number of sequencing reads comprising insertions or deletions, the number of insertions and / or deletions in the polynucleotide sequence mediated by the nucleotide-guided nuclease is estimated.
Citation Information
Patent Citations
Methods and compositions for RNA-directed target DNA modification and for RNA-directed modulation of transcription
WO2013176772A1
Inducible DNA binding proteins and genome perturbation tools and applications thereof
WO2014018423A2