Identification, characterization, and quantification of double-strand DNA break repair introduced by CRISPR.

JP7898488B2Active Publication Date: 2026-07-31INTEGRATED DNA TECHNOLOGIES INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
INTEGRATED DNA TECHNOLOGIES INC
Filing Date
2024-10-28
Publication Date
2026-07-31

Smart Images

  • Figure 0007898488000005
    Figure 0007898488000005
  • Figure 0007898488000006
    Figure 0007898488000006
  • Figure 0007898488000007
    Figure 0007898488000007
Patent Text Reader

Abstract

To provide a system and process for identifying and characterizing double-stranded DNA break repair sites that are based on biological information, and have improved accuracy.SOLUTION: A process is configured to: receive sample sequence data including a plurality of sequences; analyze and merge the sample sequence data; when a single-stranded or double-stranded DNA oligonucleotide donor is provided, output a target site sequence including a prediction result of a repair event; bine the merged sequence with the target site sequence or an arbitrary target prediction result; output a target-read realignment; realign a binned target-read realignment with a target site; generate a final alignment; analyze the final alignment; identify and quantify mutations within a pre-defined sequence distance window deriving from a canonical enzyme specific cut site; and output the final alignment, and data on results of the analysis and quantification as tables or graphics.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to related applications This application claims priority to U.S. Provisional Patent Application Nos. 62 / 870,426 and 62 / 870,471, both filed on July 3, 2019, and U.S. Provisional Patent Application Nos. 62 / 952,603 and 62 / 952,598, both filed on December 23, 2019, and their contents are hereby incorporated by reference in their entirety.

[0002] Systems and processes for identifying and characterizing double - strand DNA break repair sites with improved accuracy based on biological information are described herein. A sequence alignment process is also described that uses biological data to inform an alignment matrix for position - specific alignment scoring, thereby resulting in accurate identification of non - standard target sites.

Background Art

[0003] Genome editing has been transformed by the use of targeted nucleases such as CRISPR proteins. CRISPR enzymes form ribonucleoprotein (RNP) complexes when hybridized with either a two - part crRNA and tracrRNA or a single - guide RNA (sgRNA). By either approach, a short protospacer sequence (guide RNA or "gRNA") targets a specific sequence in a complementary molecule. When a match is found, these enzymes introduce a break into one or both DNA (or RNA) strands. CRISPR enzymes that target DNA (e.g., Cas9, Cas12a / Cpf1) introduce double - strand breaks (DSBs) at predictable genomic locations relative to the hybridization target of the gRNA. DNA DSBs are repaired by intracellular mechanisms, but the repair process often results in insertions and deletions (indels), substitutions, and other sub - optimal allelic variants.

[0004] Each cell in the affected population must repair itself independently of neighboring cells, and a particular outcome may involve different resulting alleles; therefore, a population of cells may contain multiple alleles at the targeted site. In addition, the targeting ability of these nucleases is often somewhat nonspecific, which can lead to undesirable mutations at other off-target genomic locations.

[0005] Characterizing and quantifying multiple alleles at both on-target and off-target sites is highly desirable. Researchers often use DNA sequencing (e.g., Illumina's next-generation sequencing; NGS) to observe the diversity of the resulting alleles. Multiplexed polymerase chain reaction (PCR) can be performed to amplify and enrich all targeted sites. The resulting amplicons can be sequenced. Multiple alleles can then be characterized and counted using specialized software.

[0006] Numerous specialized software tools have been developed to characterize allele variants arising from DSBs. Previous tools include CRISPResso[1], crispRvariants[2], and Amplican[3]. These tools generally work by aligning each sequence read to the expected amplicon target using Needleman-Wunsch, bwa, or a custom-ordered alignment algorithm. The algorithm considers possible reads: A list of target alignments is created. Each alignment is scored based on the number of nucleotide matches, mismatches, and missing values ​​(gaps). The best-scoring alignment is used for downstream data processing.

[0007] Alignment algorithms may create a target alignment for a query that is of equal importance, which is most likely to occur if the query contains insertions or deletions. From equally important choices, the alignment method will either return all of them or select one choice. When selecting, some methods make a random selection. Without a good predictive model or set of heuristic rules for making the selection, the selection of the alignment is variable, which can lead to incorrect indel annotation and lower accuracy results.

[0008] There is a need for algorithms and processes that identify and characterize double-strand DNA break repair sites with improved accuracy, based on biological information. [Overview of the project]

[0009] One embodiment described herein is a computer implementation process for identifying and characterizing double-strand DNA break repair sites with improved accuracy, comprising the steps of: receiving sample sequence data comprising multiple sequences; analyzing and merging the sample sequence data and outputting the merged sequence; when a single-strand or double-strand DNA oligonucleotide donor is provided, developing a target site sequence containing predicted repair event results and outputting the target prediction results; using a mapper to bin the merged sequence by the target site sequence or any target prediction results and outputting the target read alignment; using an enzyme-specific site-specific scoring matrix derived from biological data applied based on guide sequences and the locations of standard enzyme-specific break sites to realign the binned target read alignment with the target site and generate a final alignment; analyzing the final alignment and identifying and quantifying mutations within a predetermined sequence distance window derived from standard enzyme-specific break sites; and outputting the final alignment, analysis, and quantification results as a table or graphic on a processor. In one embodiment, the sequence data comprises sequences derived from a population of cells or a target. In another embodiment, the enzyme-specific cleavage site includes one or more of Cas9, Cas12a, or other Cas enzymes. In another embodiment, the given sequence distance window is enzyme-specific and includes 1 nt to about 15 nt. In another embodiment, the results show edit percentage, insertion percentage, deletion percentage, or a combination thereof. In another embodiment, the accuracy of identifying variant target sites is improved by about 15 to about 20% compared to an equivalent process.

[0010] Another embodiment described herein is a computer implementation process for aligning biological sequences, comprising the steps of: receiving sample sequence data comprising: aligning the sequence data with a predicted target sequence using a matrix based on enzyme-specific site-specific scoring of specific nuclease target sites; and outputting the alignment results as a table or graphic on a processor. In one embodiment, the sequence data comprises sequences derived from a population of cells or a target. In another embodiment, the specific nuclease target sequence comprises a target site for one or more of Cas9, Cas12a, or other Cas enzymes. In another embodiment, the matrix uses site-specific gap start and elongation penalties.

[0011] Another embodiment described herein is a method for identifying and characterizing double-strand DNA break repair sites with improved accuracy, comprising: extracting genomic DNA from a population or tissue of cells derived from a target; amplifying the genomic DNA using multiplex PCR to identify the target site. The method comprises the steps of: generating enriched amplicons for a column; sequencing the amplicons to obtain sample sequence data; then receiving sample sequence data containing multiple sequences; analyzing and merging the sample sequence data and outputting the merged sequences; when a single-stranded or double-stranded DNA oligonucleotide donor is provided, developing a target site sequence containing predicted repair event results and outputting target prediction results; using a mapper to binn the merged sequences by the target site sequence or any target prediction results and outputting target read alignments; using an enzyme-specific site-specific scoring matrix derived from biological data applied based on the locations of guide sequences and standard enzyme-specific cleavage sites to realign the binned target read alignments with the target sites and generate a final alignment; analyzing the final alignment and identifying and quantifying mutations within a predetermined sequence distance window derived from standard enzyme-specific cleavage sites; and outputting the final alignment, analysis, and quantification results as a table or graphic on a processor. In one embodiment, the enzyme-specific cleavage sites include one or more of Cas9, Cas12a, or other Cas enzymes. In another embodiment, the given sequence distance window is enzyme-specific and includes 1 nt to approximately 15 nt. In another embodiment, the results show edit percentage, insertion percentage, deletion percentage, or a combination thereof. In another embodiment, the accuracy of identifying variant target sites is improved by approximately 15 to 20 percent compared to an equivalent process.

[0012] This patent or application includes at least one drawing drawn in color. A copy of this application or the patent accompanied by the color drawing will be provided by the Patent Office upon application and payment of the required fees. [Brief explanation of the drawing]

[0013] [Figure 1]Complete workflow for CRISPAltRations. Edited genomic DNA is extracted and amplified using targeted multiplex PCR to enrich on-target and predicted off-target loci. Amplicons are sequenced in Illumina MiSeq. Read pairs are merged into single fragments (FLASH), mapped to the genome (minimap2), and binned by their alignment to predict amplicon locations. Reads in each bin are identified for cleavage sites, a position-specific gap start / extension bonus matrix is ​​created, and then realigned with the predicted amplicon sequence to preferentially align indels closer to the cleavage site / predicted indel profile for each enzyme (CRISPAltRations code +psnw). Indels intersecting the upstream or downstream window of the cleavage site are annotated. The edit percentage is the sum of reads containing indels / observed totals. [Figure 2] A directed acyclic image of the CRISPRAltRations pipeline. Dashed boxes represent steps in the pipeline, each step may include one or more software tools. Lines and arrows indicate the flow of information through the pipeline. Two key steps are minimap2_orig_reads (orange), which uses minimap2[4] to align sequence reads against a reference genome; this is an optional step. minimap2 is then used to align sequence reads against their expected target regions. The CRISPy Python tool (developed internally) aligns sequence reads against those target regions by calling a specially ordered modified version of psnw, which performs cleavage realignment against the target regions. CRISPy also takes part in characterizing detected indels in the aligned reads. [Figure 3]MinION sequencing read data aligned to the reference target demonstrates blunt insertions. Drops within the expected range (highlighted in gray) indicate large insertions. Mismatches at the inner edges of the gray highlighting indicate mismatches between the observed read data and the expected reference. [Figure 4] Examples of schematic diagrams illustrating the types of structural variants that may occur when using DNA oligos as templates for homologous recombination repair (HDR). These examples are not exhaustive. In the schematic diagrams, blue represents the reference sequence, green represents the homologous arm, and orange represents the desired insertion sequence. (A) and (B) are examples of double-stranded and single-stranded template oligos, respectively. (1) [Complete Repair] The region containing the DSB is repaired without any introduced structural variants, even in the presence of a single-stranded or double-stranded template oligo. (2) [HDR-mediated Repair] The template oligo directs the HDR, resulting in the desired insertion. Only the desired insertion sequence is observed in the repaired DNA. (3) [Non-homologous End-joining (NHEJ) Repair] The template oligo is inserted blunt-end after the DSB. (4) [NHEJ Repair with Duplicate Insertion] The template oligo is inserted blunt-end multiple times after the DSB. Examples 3 and 4 are most likely to occur with double-stranded oligonucleotides used as repair templates, and homologous arms present in the donor sequence are also inserted into the genome. [Figure 5] Location of editing event occurrences in Cas9 data. Using a set of 274 Cas9 nuclease cleavage sites, the locations of indel generation guided the editing of specific genomic targets in the Jurcut cell line. Both (A) deletion and (B) insertion site events were normalized by quantifying them as the total percentage of each indel event in the sample. Only sites with >50 reads and >5% indels were used for analysis to limit low-confidence signals and remove noise. Indels were quantified within a 20 bp window from the cleavage site. Outliers were mainly found to be sites with non-reference indels present in the Jurcut cell line that were not thought to be caused by DSB activity. [Figure 6] Location of editing event occurrences in Cas12a data. Using a set of 199 Cas12a, the locations of indel generation from nuclease cleavage sites guide the editing of specific genomic targets in Jurcut cell lines. Both (A) deletion and (B) insertion site events were normalized by quantifying them as the total percentage of each indel event in the sample. Only sites with >50 reads and >5% indels were used for analysis to limit noise from low-confidence signals. Indels were quantified within a 20 bp window from the cleavage site. Outliers were mainly found to be sites with non-reference indels present in Jurcut cell lines that were not thought to be caused by DSB activity. [Figure 7] Vectors of gap start / elongation penalties near nuclease-introduced DSBs. To positively weight indels near cleavage sites in sequence alignment, we use a position-specific matrix of values ​​representing variable gap start or elongation penalties, where the array length is equal to the integer position of each nucleic acid in the target (thick blue line). The vector values ​​(red circles and blue diamonds) are varied based on their proximity to the nuclease cleavage site (vertical black dashed line). Thus, indels closest to the cleavage site have the smallest gap start or elongation penalty. [Figure 8] Selection of the optimal window for observing true CRISPR-Cas9 edit events. (A) The edit window is the distance of nucleotides around the CRISPR cleavage site used to identify indels. True and false edit events were calculated as the percentage of indels from targeted sequencing of treated or untreated cells for (B) Cas9 (n=263 pairs of treated, vs. control) or (C) Cas12a (n=384 pairs of treated, vs. control), respectively. The most true edit events are collected in a 4nt window for Cas9 and a 7nt window for Cas12a, but expanding the window will only further collect additional false edits. [Figure 9]Screenshot of deduplication-removed read alignment by observed frequency. The total region is indicated by the height of the vertical gray bar (top), and the reads are the horizontal colored bars. Lighter colored reads indicate indels observed more frequently. Thin horizontal lines indicate deletions, vertical purple "I" symbols indicate insertions, and colored bars within the reads indicate mismatched bases. [Figure 10] CRISPAltRations reliably finds the correct indels. Each bar reports the percentage of indels that were accurately returned. Synthetically generated data representing 12,060 unique indel events was uniformly distributed across bins of each size (x-axis) and 603 unique amplicons. Error bars represent a 95% confidence interval across the entire target. [Figure 11] CRISPRAltRations more accurately reports the total edit percentage in 603 synthetic targets edited by simulated (A)Cas9 or (B)Cas12a. Dots represent the edit percentage observed at each target site, and horizontal lines represent the percentage of indels introduced at each target in the synthesized data created. CRISPResso2Pooled and CRISPResso1Pooled were performed for each panel using default pipeline parameters, with the minimum read depth required for quantification removed (to prevent incorrect dropout of amplicons). [Modes for carrying out the invention]

[0014] One embodiment described herein is an analysis pipeline called CRISPAltRations. See Figures 1-2. Briefly, this pipeline takes in a FASTQ file and uses FLASH to build a merged R1 / R2 consensus. Simultaneously, a target site reference is built that describes the sequences for all expected on-target locations. Optionally, targets containing the expected outcomes of homologous recombination repair (HDR) events are built. Next, the merged sequence reads are aligned with the target reference sequences using minimap2 (originally developed for rapid alignment of long reads, e.g., those created by Oxford Nanopore Technologies MinION). Then, the reads aligned with each target are realigned using a modified Smith-Waterman aligner. The modified aligner can improve the detection of insertions and deletions resulting from DSB repair. All observed variants within a given distance of the DSB location are characterized and quantified. Finally, the results are summarized in tables and graphs. The various programs, tools, and file types described (and those listed below) are well known and readily accessible to those skilled in the art. It should be understood that these programs, tools, and file types are illustrative and not intended to be limiting. Other tools and file types may be used to perform the described processing and analysis.

[0015] In this analysis pipeline, the following improvements over previous methods are described. First, the use of minimap2 [4] enables the alignment of reads created from both short and long read arrays. Second, the ability to fully characterize (i.e., correctly occur) HDR events is improved by constructing the expected outcomes of homologous recombination repair events. Third, the use of a modified Needleman - Wunsch aligner that can accept Cas - specific bonus matrices enables significantly improved indel characterization and quantification of the percentage of editing (%) over previous methods. Fourth, the graphical visualization of the introduced allelic variants is improved. Fifth, the predicted repair events described in previous tools [5] can be compared against the observed repair, and the molecular pathways involved in repair can be described.

[0016] In one embodiment, the process described herein has the following advantageous uses: · Accurate characterization of indel profiles arising from DSBs. · The percentage of reads containing indels after DSB repair is used to calculate the percentage of editing. This metric (% editing) is used to determine the effectiveness of gRNAs for use in CRISPR - Cas gene editing. · Accurate characterization of the resulting indels similarly improves the ability to identify the percentage of a cell's chromosomes in a population of cells containing frame - shifting mutations. Frame - shifting mutations alter the proteins encoded by the affected genes. · Accurate characterization of inserted sequences. · Accurate characterization of multiple mutations resulting from the delivery of multiple gRNA / Cas9 (i.e., ribonucleoprotein complexes) or double guide region modifications. · Analysis of indels sequenced on long read platforms such as MinION. Additionally, this enables the step - by - step characterization of both ends of large (>400nt) insertions that occur as deletion events after DSB repair. · Improved visualization of results.

[0017] One embodiment described herein is a computer implementation process for identifying and characterizing double-strand DNA break repair sites with improved accuracy, comprising the steps of: receiving sample sequence data comprising multiple sequences; analyzing and merging the sample sequence data and outputting the merged sequences; given a single-strand or double-strand DNA oligonucleotide donor, developing a target site sequence containing predicted repair events and outputting the target prediction results; using a mapper to bin the merged sequences by the target site sequence or any target prediction results and outputting a target read alignment; using a guide sequence and an enzyme-specific site-specific scoring matrix derived from biological data applied based on the locations of standard enzyme-specific break sites to realign the binned target read alignment with the target site and generate a final alignment; analyzing the final alignment and identifying and quantifying mutations within a predetermined sequence distance window derived from standard enzyme-specific break sites; and outputting the final alignment, analysis, and quantification results as a table or graphic on a processor.

[0018] In one embodiment, the edited genomic DNA is extracted and amplified using targeted multiplex PCR to enrich for on-target and predicted off-target loci. The amplicon is sequenced on an Illumina MiSeq. Read pairs are merged into single fragments (FLASH), mapped to the genome (minimap2), and binned by their alignment to predict the amplicon position. Reads within each bin find the cleavage sites, create a position-specific gap start / extension bonus matrix, and then realign to the predicted amplicon sequence to preferentially align indels closer to the cleavage site / predicted indel profile for each enzyme (CRISPAltRations code + psnw). Indels intersecting a window upstream or downstream of the cleavage site are annotated. The percent edited is the sum of reads containing indels / observed total.

[0019] In some embodiments, the process described herein uses minimap2 [4] which allows alignment of reads created from both short and long read sequences. Previous tools typically accept only short read sequencing data, such as that created by Illumina sequencers. Others have used long read sequencing data to examine large insertions or deletions [6 - 8], but no publicly available tool exists that does this alone. Handling of long read data is made possible, in part, by the use of the minimap2 aligner. For example, the results of the alignment can be visualized, which shows the identification of insertions of blunt-ended molecules in the DNA after DSB repair (FIG. 3). Another embodiment sorts between real indels and those derived from noise using a method similar to a previously published tool [7] where small indels are ignored. That is (FIG. 3). Another embodiment sorts between real indels and those derived from noise using a method similar to a previously published tool [7] where small indels are ignored.

[0020] In another embodiment, the ability to characterize a complete HDR event is improved by constructing the expected outcome of the HDR event. The reference file contains, in FASTA format, the respective expected sequence targets and modified sequence targets as well. A first step aimed at constructing this file involves creating a reference sequence index that allows reads to align with their respective expected structural variants. For example, when examining the targeted region for a DSB and the double-stranded DNA donor oligo that enables HDR, there are several different possible biological repair outcomes: complete repair (Figure 4-1), HDR-mediated repair (Figure 4-2), NHEJ repair (Figure 4-3), and NHEJ repair with duplicate insertion (Figure 4-4). Other outcomes such as template fragments or triple template insertions are also possible (not shown). Similar reference file construction techniques have been used by other tools such as UDiTaS®[9].

[0021] In another embodiment, a modified version of the Needleman-Wunsch algorithm is used to realign the reads to their expected targets. The method described herein increases the accuracy of alignments containing indels (as annotations in the CIGAR string of the alignment). This significantly improves the characterization of indels and the quantification of edit % compared to previous methods. DNA sequencing aligners such as minimap2 and the Needleman-Wunsch method consider indel alignments using fixed penalties for gap initiation and extension. This method is improved by realigning the reads to their targets using position-specific gap initiation and extension penalties (possible in a tool called "psnw") so that alignments with indels are advantageous in that they overlap or are located near the predicted DSBs. This position-specific matrix is ​​a set that reflects the precisely characterized indel profiles of the specific Cas enzyme used for editing (Figures 5-6). Therefore, indel-based alignment is most favorable at or near the predicted target cleavage site (variable scoring strategy; Figure 7). This method allows for precise realignment of indels, particularly those occurring in repetitive regions within the reference sequence. This technique improves the ability to identify the most biologically probable outcomes.

[0022] A recently developed tool (CRISPResso2

[11] ) uses a detailed alignment strategy for the cut sites. However, the process described herein is tailored to the actual edit data at the Cas9 / Cas12a sites and implements a detailed alignment method for the cut sites using a full gap start / extension matrix during the alignment performed in C++. In contrast, CRISPResso2 allows only a single bonus at the cut sites and uses the method performed in Python.

[0023] In another embodiment, the process described herein collects indels near the nuclease cleavage site, as well as tag indels that intersect the cleavage site or within a fixed distance. Several published accounts suggest a fixed distance of 1–2 nt, but data supporting these selections are limited. In developing the embodiments described herein, the optimal distance around the cleavage site (i.e., window size) was studied using a set of treated Cas9-RNP and paired untreated control samples. It was observed that a 4 nt window for Cas9, or a 7 nt window for Cas12a, provided the highest sensitivity and acceptable specificity (Figure 4). The requirement for a larger window for Cas12a may be due to the mechanism of action, Ca s12a performs double-strand breaks by generating two single-strand breaks separated by 5 bp (leaving a “sticky” end)

[12] . Thus, the processes described herein may be extended to other nucleases (e.g., CasX) for which biological data inform the target window size and enzymatic mechanism of action

[13] .

[0024] In another embodiment, graphical visualization of allele variants is improved. Downstream of the alignment step, several other analyses specific to the described method are performed. To create improved visualization, reads are deduplication based on the identification of identified indel sequences in the post-alignment CRISPR editing window. The deduplication reads are written back to a BAM file, and the frequency of each deduplication read in the original population of reads is written to the associated BAM tag. After the file is indexed, the indels in the deduplication reads and their associated frequencies can be visualized using a commonly available IGV tool

[10] (Figure 9).

[0025] In another embodiment, the predicted repair events described in the previous tool [5] may be compared to the observed repair and used to determine the molecular pathways involved in the repair. The systems described herein also add the ability to compare the observed indel profile to the predicted indel profile, which allows for rapid identification of whether the experimental treatment has altered the intracellular mechanisms of DNA repair.

[0026] The practicality of the systems and methods described herein is demonstrated by creating a synthetic set of 603 gRNA:amplicon pairs. For each target, 4000 read pairs (2 × 150 bp) were synthesized using the simulated Illumina MiSeq. The indels are synthetically created using error profiles from the v3 platform. In half of the reads, random indels are introduced based on models created from observed editing profiles for Cas9 and Cas12a (Figures 4-5). The synthetic data are analyzed using the CRISPRAltRations system described herein, as well as the previously published CRISPResso1 and CRISPResso2 tools

[11] . By implementing the method herein, the ability to correctly characterize the indels is improved by approximately 15-20% (Figure 10). The algorithm described herein increases accuracy because it provides a biologically informed selection of the best alignment for targets where multiple equivalent scored alignments are possible. In addition, the method herein more accurately calculates the percentage of modified DNA molecules (Figure 11). The processes and strategies described herein represent a significant enhancement to the characterization and quantification of indels introduced after DSB repair.

[0027] Another embodiment described herein is a computer implementation process for aligning biological sequences, comprising the steps of: receiving sample sequence data comprising: aligning the sequence data with a predicted target sequence using a matrix based on enzyme-specific site-specific scoring of specific nuclease target sites; and outputting the alignment results as a table or graphic on a processor. In one embodiment, the sequence data comprises sequences derived from a population of cells or a target. In another embodiment, the specific nuclease target sequence comprises a target site for one or more of Cas9, Cas12a, or other Cas enzymes. In another embodiment, the matrix uses site-specific gap start and elongation penalties.

[0028] Another embodiment described herein is a method for identifying and characterizing double-strand DNA break repair sites with improved accuracy, comprising: extracting genomic DNA from a population or tissue of cells derived from the target; amplifying the genomic DNA using multiplex PCR to generate amplicons enriched with respect to the target site sequence; sequencing the amplicons and sample distribution The method comprises the steps of: obtaining column data; then receiving sample sequence data containing multiple sequences; analyzing and merging the sample sequence data and outputting the merged sequences; when a single-stranded or double-stranded DNA oligonucleotide donor is provided, developing a target site sequence containing predicted repair events and outputting target prediction results; using a mapper to binn the merged sequences by the target site sequence or any target prediction results and outputting target read alignments; using an enzyme-specific site-specific scoring matrix derived from biological data applied based on the locations of guide sequences and standard enzyme-specific cleavage sites to realign the binned target read alignments with the target sites and generate a final alignment; analyzing the final alignment and identifying and quantifying mutations within a predetermined sequence distance window derived from standard enzyme-specific cleavage sites; and outputting the final alignment, analysis, and quantification results as a table or graphic on a processor.

[0029] Many different arrangements of the various components and processes described herein, as well as many other components or processes not shown herein, are possible without departing from the spirit and scope of this disclosure. It should be understood that embodiments or aspects may include, or otherwise be performed by, various combinations of hardware, software, or electronic components. For example, various microprocessors and application-specific integrated circuits ("ASICs") may be used, as well as software in various languages. Servers and various computer devices may also be used, and may include one or more processing units, one or more computer-readable media, one or more input / output interfaces, and various connections (e.g., system buses) for connecting components.

[0030] It will be apparent to those skilled in the art that appropriate modifications and adaptations to the compositions, formulations, methods, processes, and applications described herein can be made without departing from the scope of any of their embodiments or aspects. The compositions and methods provided are illustrative and are not intended to limit the scope of any particular embodiment. All the various embodiments, aspects, and options disclosed herein can be combined with any variations or iterations. The scope of the methods and processes disclosed herein includes all actual or possible combinations of the embodiments, aspects, options, examples, and preferences described herein. The methods disclosed herein may exclude any component or step, replace any component or step disclosed herein, or include any component or step disclosed herein anywhere. If the meaning of any term in any of the patents or publications incorporated by reference conflicts with the meaning of a term used herein, the meaning of the term or expression herein shall prevail. Furthermore, this specification discloses and describes only illustrative embodiments. All patents and publications cited herein are incorporated herein by reference with respect to their specific teachings. The following is a description of the claims as they were at the time of filing the application. [Claim 1] A computer implementation process for identifying and characterizing double-strand DNA break repair sites with improved accuracy, A step of receiving sample sequence data containing multiple sequences; A step to analyze and merge sample sequence data and output the merged sequence; Given a single-stranded or double-stranded DNA oligonucleotide donor, the process involves developing a target site sequence containing predicted repair events and outputting the target prediction results; The step involves using a mapper to binn the merged sequences by target site sequences or any target prediction results, and outputting the target read alignment; A step of realigning the binned target read alignment with the target site and generating the final alignment using an enzyme-specific site-specific scoring matrix derived from biological data applied based on the guide sequence and the location of standard enzyme-specific cleavage sites; A step of analyzing the final alignment to identify and quantify mutations within a predetermined sequence distance window derived from standard enzyme-specific cleavage sites; Steps to output the final alignment, analysis, and quantification results as a table or graphic. A process that involves executing on a processor. [Claim 2] The process according to claim 1, wherein the sequence data includes sequences derived from a population of cells or a target. [Claim 3] The process according to claim 1, wherein the enzyme-specific cleavage site comprises one or more of Cas9, Cas12a, or other Cas enzymes. [Claim 4] The process according to claim 1, wherein a predetermined sequence distance window is enzyme-specific and includes 1 nt to about 15 nt. [Claim 5] The process according to claim 1, wherein the result shows an edit percentage, an insertion percentage, a deletion percentage, or a combination thereof. [Claim 6] The process according to claim 1, wherein the accuracy of identifying variant target sites is improved by approximately 15 to 20 percent compared to an equivalent process. [Claim 7] A computer implementation process for aligning biological sequences, A step of receiving sample sequence data containing multiple sequences; A step of aligning sequence data with a predicted target sequence using a matrix based on enzyme-specific site-specific scoring of specific nuclease target sites; Steps to output alignment results as a table or graphic. A process that involves executing on a processor. [Claim 8] The process according to claim 7, wherein the sequence data includes sequences derived from a population of cells or a target. [Claim 9] The process according to claim 7, wherein the specific nuclease target sequence includes a target site for one or more of Cas9, Cas12a, or other Cas enzymes. [Claim 10] The process according to claim 7, wherein the matrix uses position-specific gap start and stretching penalties. [Claim 11] A method for identifying and characterizing double-strand DNA break repair sites with improved accuracy, Extracting genomic DNA from a population or tissue of cells derived from the target; Amplifying genomic DNA using multiplex PCR to generate amplicons enriched with respect to target site sequences; To determine the sequence of the amplicon and obtain sample sequence data; after that, A step of receiving sample sequence data containing multiple sequences; A step to analyze and merge sample sequence data and output the merged sequence; Given a single-stranded or double-stranded DNA oligonucleotide donor, the process involves developing a target site sequence containing predicted repair events and outputting the target prediction results; The step involves using a mapper to binn the merged sequences by target site sequences or any target prediction results, and outputting the target read alignment; A step of realigning the binned target read alignment with the target site and generating the final alignment using an enzyme-specific site-specific scoring matrix derived from biological data applied based on the guide sequence and the location of standard enzyme-specific cleavage sites; A step of analyzing the final alignment to identify and quantify mutations within a predetermined sequence distance window derived from standard enzyme-specific cleavage sites; Steps to output the final alignment, analysis, and quantification results as a table or graphic. Execute on the processor Methods that include... [Claim 12] The process according to claim 1, wherein the enzyme-specific cleavage site comprises one or more of Cas9, Cas12a, or other Cas enzymes. [Claim 13] The process according to claim 1, wherein a predetermined sequence distance window is enzyme-specific and includes 1 nt to about 15 nt. [Claim 14] The process according to claim 1, wherein the result shows an edit percentage, an insertion percentage, a deletion percentage, or a combination thereof. [Claim 15] The process according to claim 1, wherein the accuracy of identifying variant target sites is improved by approximately 15 to 20 percent compared to an equivalent process.

[0031] References 1.Pinello, L. et al., “Analyzing CRISPR genome-editing experiments with CRISPResso.” Nat Biotechnol. 34(7): 695-697 (2016). 2. Lindsay, H. et al., “CrispRVariants: precisely charting the mutation spectrum in genome engineering experiments,” Nat. Biotechnol. 34(7): 701-703 (2015). 3.Labun, K. et al., “Accurate analysis of genuine CRISPR editing events with ampliCan Kornel,” bioRxiv 249474 (2018); now published in Genome Research 29: 843-847 (2019) 4.Li, H., “Minimap2: Pairwise alignment for nucleotide sequences,” Bioin formatics 34(18): 3094-3100 (2018). 5.Shen, M. W. et al., “Predictable and precise template-free CRISPR editing of pathogenic variants,” Nature 563 (7733): 646-651 (2018). 6.Hendel, A. et al., “Quantifying genome-editing outcomes at endogenous loci with SMRT sequencing.” Cell Rep. 7(1): 293-305 (2014). 7.Iyer, S. et al., “Precise therapeutic gene correction by a simple nuclease-induced double-stranded break,” Nature 568 (7753): 561-565 (2019). 8.Vu, G. T. H. et al., “Endogenous sequence patterns predispose the repair modes of CRISPR / Cas9-induced DNA double-stranded breaks in Arabidopsis thaliana,” Plant J. 92(1): 57-67 (2017). 9.Giannoukos, G. et al., “UDiTaS TM , a genome editing detection method for indels and genome rearrangements,” BMC Genomics 19: 212 (2018). 10. Robinson, J., “Integrated genomics viewer,” Nat. Biotechnol. 29(1), 24-26 (2012). 11. Clement, K. et al., “Analysis and comparison of genome editing using CRISPResso2,” bioRxiv 1-20 (2018). Now published in Nat. Biotechnol. 37(3): 224-226 (2019) 12. Zetsche, B. et al., “Cpf1 Is a Single RNA-Guided Endonuclease of a Class 2 CRISPR-Cas System,” Cell 163(3): 759-771 (2015). 13. Liu, JJ et al., “CasX enzymes comprise a distinct family of RNA-guided genome editors,” Nature 566(7743): 218-223 (2019).

[0032] Computer code An example code used to generate a 1D scoring matrix for gap start bonuses during alignment using psnw.

number

number

number

[0033] An example code used to generate a 1D scoring matrix for gap stretching bonuses during alignment using psnw.

number

Claims

1. A method for selecting a guide RNA by identifying and characterizing double-stranded DNA (dsDNA) break repair sites with improved accuracy, comprising the following steps: (a) A step of receiving two or more guide RNAs; (b) The step of receiving genomic sample sequence data comprising a plurality of sequences enriched with respect to a target site sequence, wherein the genomic sample sequence data enriched with respect to the target site sequence is at least partially based on CRISPR-Cas editing performed in a population of cells or tissue using two or more guide RNAs that target the target site sequence; (c) A step of merging enriched genome sample sequence data with respect to the target site sequence; (d) A step of generating a predicted target site sequence of the genome, including predicted dsDNA repair events, based on the provided single-stranded or double-stranded DNA oligonucleotide donor; (e) Using a mapper, binning the merged sequences based on their alignment to the genome and outputting the binned target read alignment; (f) A step of realigning the binned target read alignment derived from step (e) with the predicted target site sequence derived from step (d), wherein the binned target read alignment is realigned using an aligner weighted by multiple bonus scoring matrices of Cas enzyme-specific, position-specific, full-gap initiation and gap extension, the aligner simultaneously and preferentially aligns multiple edit events within a predetermined sequence distance window of predicted dsDNA break repair events for each Cas enzyme and each guide RNA, thereby generating a final alignment; (g) Using the final alignment, identify and quantify the editing events within a predetermined sequence distance window of the predicted dsDNA break repair events for each Cas enzyme and each guide RNA; and (h) A step of selecting one or more guide RNAs having effective CRISPR-Cas editing from two or more guide RNAs based on one or more quantitative data selected from the group consisting of editing percentage, insertion percentage, and deletion percentage.

2. The method according to claim 1, further comprising the following steps: (i) A step in which one or more selected guide RNAs are used in a further CRISPR-Cas editing test.

3. The method according to claim 1, further comprising the following steps: (j) A step of outputting the final alignment, edit percentage, insertion percentage, deletion percentage, or a combination thereof, as a table or graphic.

4. The method according to claim 1, wherein multiple bonus scoring matrices optimize the alignment of multiple edit events occurring near or in the vicinity of a predicted dsDNA break repair event, using position-specific gap start and gap extension variable penalty vectors.

5. The method according to claim 4, wherein multiple bonus scoring matrices are derived from the biological editing data of standard Cas enzyme cleavage sites and the positions of each guide RNA.

6. The method according to claim 1, wherein the genome sample sequence data includes sequences derived from a population of cells or a target.

7. The method according to claim 1, wherein the Cas enzyme is Cas9 or Cas12a.

8. The method according to claim 1, wherein the sequence distance window for predicted dsDNA cleavage repair events for each Cas enzyme is 1 nt to 15 nt.

9. The method according to claim 7, wherein the predetermined sequence distance window for predicted dsDNA cleavage repair events for Cas9 is 8 nt.

10. The method according to claim 7, wherein a predetermined sequence distance window for a predicted dsDNA break repair event for Cas12a is 9 nt from the offset center point and -3 bp from the PAM distal cut.

11. The method according to claim 1, wherein the Cas enzyme-specific position-specific matrix is ​​weighted according to a predetermined matrix to the editing event closest to the predicted dsDNA break repair event having the lowest gap initiation or gap elongation penalty.

12. The method according to claim 1, wherein the Cas enzyme-specific position-specific matrix has a gap start penalty of 0 to +4.

13. The method according to claim 1, wherein the Cas enzyme-specific position-specific matrix has a gap extension penalty of 0 to +1.

14. Genome sample sequence data, Extracting edited genomic DNA from a population of cells or tissue; Amplifying edited genomic DNA using multiplex PCR to generate amplicons enriched with respect to target site sequences; Performing next-generation sequencing on amplicons rich in target site sequences to obtain genome sample sequence data rich in target site sequences; The method according to claim 1, obtained by...