A method for specifying nucleazeon / off-target editing positions, called "CTL-seq" (CRISPR Tag Linear-seq).

The method improves the detection of CRISPR editing sites by co-delivering unique tag sequences and using specific primers for genomic amplification, addressing the limitations of current detection methods and enhancing accuracy and sensitivity.

JP2026048971APending Publication Date: 2026-03-17INTEGRATED DNA TECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Current methods for identifying CRISPR editing sites lack accuracy and sensitivity, leading to incomplete detection of on- and off-target editing sites, which can induce mutations and alter gene function.

Method used

A method involving the co-delivery of guide RNA and unique tag sequences to cells, followed by genomic DNA isolation, fragmentation, and amplification using specific primers, allowing for improved detection of CRISPR editing sites through sequencing and alignment to a reference genome.

Benefits of technology

Enhances the accuracy and sensitivity of identifying on- and off-target CRISPR editing sites, reducing false positives and improving the reliability of genomic editing analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026048971000017
    Figure 2026048971000017
  • Figure 2026048971000018
    Figure 2026048971000018
  • Figure 2026048971000019
    Figure 2026048971000019
Patent Text Reader

Abstract

This provides a method for identifying and specifying on / off-target CRISPR editing sites with improved accuracy and sensitivity. [Solution] (a) The steps of isolating genomic DNA from a cell having one or more tag sequences incorporated into a target site in the cell's genome, fragmenting the genomic DNA, and linking the fragmented genomic DNA to a universal adapter sequence containing a unique molecular index (UMI), (b) A step of generating an amplified sequence by amplifying a ligated DNA fragment using multiplex PCR containing an RNaseH2 cleavage enzyme, a tag-specific primer and a universal adapter sequence-specific primer, and (c) A method comprising the step of sequencing the amplified sequence using a universal sequencing primer that targets the 5' universal tail sequence of a tag-specific primer, thereby identifying the on / off target CRISPR editing site.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 055,460, filed Jul. 23, 2020, the entire disclosure of which is incorporated herein by reference.

[0002] Reference to Sequence Listing This application is filed with a computer - readable form of the sequence listing in accordance with 37 C.F.R. § 1.821(c). The text file, “013670 - 9056 - WO01_sequence_listing_19 - JUL - 2021_ST25.txt”, submitted via EFS contains 273 sequences, was created on Jul. 19, 2021, and has a file size of 153 kilobytes, the entire disclosure of which is incorporated herein by reference.

[0003] Technical Field Described herein are methods for identifying and specifying on - and off - target CRISPR editing sites with improved accuracy and sensitivity.

Background Art

[0004] CRISPR (Clustered and Regularly Arranged Short Palindromic Sequence Repeats) has revolutionized genomics by enabling the easy introduction of changes into the genetic code. CRISPR systems, such as the Cas9 and Cas12a proteins, are induced to targets by RNA oligonucleotide sequences (forming ribonucleoproteins; RNPs) to which Cas proteins are bound, where the enzymes generate double-strand breaks (DSBs) in the DNA sequence. Natural cellular mechanisms typically repair DSBs using non-homologous end joining (NHEJ) or homologous recombination repair (HDR) molecular pathways. DNA repaired via NHEJs occurring at on- and off-target sites often contains indels (insertions / deletions), which can induce mutations and alter the function of the encoded gene. Therefore, identifying these sites is crucial for analyzing the effects of on- and off-target editing on biological phenotypes.

[0005] Currently, there is no “absolute standard” for identifying or specifying off-target editing sites for CRISPR or other nucleases. Many methods have been developed. These methods employ a variety of strategies, including detection of endogenous repair mechanisms assembled with DSBs (Discover-Seq[1]), incorporation of DNA tag sequences into the host cell genome (GUIDE-Seq; see U.S. Patent No. 9,822,407, iGUIDE[2, 3]), or in vitro DNA cleavage (BLISS[4], CIRCLE-Seq[5], SiteSeq[6]).

[0006] Cellular or cell-based (sometimes called in vivo) and biochemical (sometimes called in vitro) off-target assay designation systems each have their advantages. Because DNA-bound proteins and epigenetic marks alter the function of nuclease activity, it has been suggested that cellular or cell-based methods may better identify actual editing targets [7]. However, because biochemical methods specify sites not identified by cellular or cell-based methods, it has been suggested that biochemical methods may be more comprehensive [5, 6]. Nevertheless, these current tools tend to have incomplete sensitivity [5, 6] (see Figure 1).

[0007] What is needed is a method for detecting and specifying on- and off-target CRISPR editing sites with improved accuracy and sensitivity. [Overview of the project]

[0008] One embodiment described herein is a method for identifying and designating on- and off-target CRISPR editing sites with improved accuracy and sensitivity, comprising: (a) guide sequence RNA (sgRNA) or two parts CRISPR The method comprises the steps of: (b) simultaneously delivering RNA:transactivating crRNA (crRNA:tracrRNA) double helix, one or more tag sequences, and an RNA guide endonuclease to cells; (c) incubating the cells for a time sufficient to allow double-strand breaks to occur; (d) isolating genomic DNA from the cells, fragmenting the genomic DNA, and ligating the fragmented genomic DNA to a unique molecular index containing a universal adapter sequence; (e) amplifying the ligated DNA fragments using primers targeting the tag and the universal adapter sequence to generate a first set of amplified sequences; (f) amplifying the first set of amplified sequences using a universal sequencing primer targeting the tail of a Tag-pTOP or Tag-pBOT primer to generate a second set of amplified sequences; (g) sequencing the pooled sequences and obtaining sequencing data; and (f) identifying on / off-target CRISPR editing loci. In one embodiment, a universal sequencing primer targets the SP1 or SP2 sequence (SEQ ID NOs. 7, 8) tail of a Tag-pTOP or Tag-pBOT primer to generate a second set of amplified sequences. In another embodiment, a universal sequencing primer targets the pre-designed non-homologous sequence (SEQ ID NOs. 269-273) tail of a Tag-pTOP or Tag-pBot primer to generate a second set of amplified sequences. In yet another embodiment, a universal sequencing primer targets the pre-designed 13-mer tail of a Tag-pTOP or Tag-pBot primer to generate a second set of amplified sequences. In yet another embodiment, step (g) includes the processor performing (i) aligning the sequence data to a reference genome, (ii) identifying on / off-target CRISPR editing loci, and (iii) aligning, analyzing, and outputting the resulting data as a file, table, or figure in a custom format.In another embodiment, the method further includes the step of normalizing a second set of (e1) amplified sequences after step (e) to produce a concentration-normalized library, pooling the normalized library with other samples to produce a pooled library, and continuing through steps (f) to (i). In another embodiment, step (d) uses suppression PCR. In another embodiment, the RNA guide endonuclease comprises an endogenously expressed Cas enzyme, a Cas expression vector, a Cas protein, or a CasRNP complex. In another embodiment, the RNA guide endonuclease comprises an endogenously expressed Cas9 enzyme, a Cas9 expression vector, a Cas9 protein, or a Cas9RNP complex. In another embodiment, the cells comprise human or mouse cells. In another embodiment, the time is approximately 24 hours to approximately 96 hours. In another embodiment, multiple tag sequences are delivered simultaneously. In another embodiment, the tag sequence comprises a double-stranded deoxyribooligonucleotide (dsDNA) containing 52 base pairs. In another embodiment, the tag sequence includes a 5' terminal phosphate group, as well as phosphorothioate bonds between the 1st and 2nd, 2nd and 3rd, 50th and 51st, and 51st and 52nd nucleotides. In yet another embodiment, the tag sequence includes double-stranded DNA comprising complementary upper and lower strand pairs of sequence numbers 1-2 or 7-268.

[0009] Other embodiments described herein are on- and off-target CRISPR editing sites identified or designated using the methods described herein.

[0010] Another embodiment described herein is a method for designing a 52-base pair tag sequence, wherein the processor has (a) a GC content of 40-90%, and a maximum homopolymer length of A:2, C:3. G:2, T:2, weighted homopolymer ratio <20, self-folding T m <50°C, and self-dimer T mA method comprising the steps of: (b) randomly generating 13-nucleotide sequences at <50°C; (c) removing sequences that perfectly align to a particular genome, or sequences that are homopolymers or GG or CC dinucleotide motifs to obtain a set of 13-mers; (d) selecting a subset of 13-mer sequences containing one or fewer CC or GG dinucleotide motifs; (e) concatenating four of the 13-mer subset sequences to form a random 52-mer sequence; (f) aligning the random 52-mer sequence to a genome; (h) removing random 52-mer sequences that are similar to the genome to generate a subset of 52-mer sequences; and (e) outputting the subset of 52-mer sequences and generating a complementary strand to produce a double-stranded 52-base pair tag sequence. In one embodiment, the genome is human or mouse. In another embodiment, the 52-base pair tag sequence is not complementary to the genome. In another embodiment, the method further includes the step of designing primers for a 52-base pair tag sequence. In another embodiment, the 52-base pair tag sequence includes a 5' terminal phosphate group, as well as phosphorothioate bonds between the 1st and 2nd, 2nd and 3rd, 50th and 51st, and 51st and 52nd nucleotides of the 52-base pair tag sequence. In another embodiment, the method further includes the step of synthesizing oligonucleotides containing a 52-base pair tag sequence, a complement of a 52-base pair tag sequence, or primers for a 52-base pair tag sequence.

[0011] Other embodiments described herein are one or more 52-base-pair tag sequences designed using the methods described herein. In one embodiment, the 52-base-pair tag sequence comprises double-stranded DNA including the upper and lower strand pairs of SEQ ID NOs. 1-2 or 7-268.

[0012] Another embodiment described herein is a method for designing primers and adapter primers partially complementary to a 52-base pair tag sequence described in claim 23, comprising, in a processor, the steps of (a) designing tag primers partially complementary to the upper and lower strands of the tag sequence, and (b) designing adapter primers partially complementary to the upper strand of the adapter sequence, wherein the tag primers comprise a 5' universal tail sequence, and the adapter primers comprise a sequence complementary to the tail of the Tag-pTOP or Tag-pBOT primer. In one embodiment, the 5' universal tail sequence is complementary to an SP1 or SP2 sequence (SEQ ID NOs. 7, 8), a locus-specific segment, six ribonucleotides (rN) from the 3' end, a 3' end mismatch, a 3' end block (3'-C3 spacer), a pre-designed non-homologous sequence (SEQ ID NOs. 269-273), or a pre-designed 13-mer sequence. In another embodiment, primers partially complementary to the upper and lower strands of the tag sequence contain a tail sequence complementary to the SP1 sequence (SEQ ID NO: 7), and the adapter primer contains a sequence complementary to the SP2 sequence (SEQ ID NO: 8) tail of the Tag-pTOP or Tag-pBOT primer, or, primers partially complementary to the upper and lower strands of the tag sequence contain a tail sequence complementary to the SP2 sequence (SEQ ID NO: 8), and the adapter primer contains a sequence complementary to the SP1 sequence (SEQ ID NO: 7) tail of the Tag-pTOP or Tag-pBOT primer. In another embodiment, amplification of a nucleic acid molecule using primers complementary to the upper and lower strands of the tag sequence and a primer complementary to the upper strand of the adapter sequence produces a PCR product containing a portion of the tag sequence, an sgDNA sequence, and the adapter sequence. In another embodiment, the method further includes the step of synthesizing oligonucleotides containing the sequences of the forward and reverse tag primers and the adapter primer. In another embodiment, a 52-base pair tag sequence and primers partially complementary to the 52-base pair tag sequence are designed and selected using an algorithm that predicts whether the primers are likely to be partially complementary or whether they tend to form primer dimers.

[0013] Other embodiments described herein are 5 designed using the methods described herein. The present invention relates to one or more primers and one or more adapter primers that are partially complementary to a two-base pair tag sequence. In one embodiment, the primers include the sequences of SEQ ID NOs: 3 and 4 and the adapter primer, and the adapter primer includes the sequence of SEQ ID NO: 5.

[0014] Another embodiment described herein involves the use of one or more double-stranded 52-base pair tag sequences to identify on- and off-target CRISPR editing sites. [Brief explanation of the drawing]

[0015] [Figure 1] The figure shows the percentage of reads common to three biological replicas in white regions, while reads common to two replicas or present in one replica are shown in black regions. Table 1 shows GUIDE-seq[3]-based designations of four different gRNAs in 96-well format in 3-sequence format. gRNA complexes were generated by mixing equimolar amounts of Alt-R crRNA-XT and Alt-R tracrRNA. Using the Nucleofector® system (Lonza), HEK293 cells stably expressing Cas9 were transfected with 10 μM gRNA and 0.5 μM dsODN GUIDE-seq tag. After 72 hours, genomic DNA (gDNA) was isolated. Genomic DNA was fragmented and adapters were ligated using the LotusDNA Library Preparation Kit (IDT). Libraries were generated by amplification from the inserted tag to the ligated adapter[3]. The libraries were then sequenced paired-end on an Illumina® platform. [Figure 2]This figure shows that GUIDE-Seq finds more off-target sites than can be verified by rhAmpSeq targeted amplification. The results presented are a collection of 331 GUIDE-Seq designated sites when delivering gRNA sequences (internal names: AR, CTNNB1, EMX1, GRHPR, HPRT38087, HPRT38285, VEGFA) to HEK293 cells stably expressing WT Cas9. Each guide was designed and targeted by a single rhAmpSeq panel aligned with the entire reference genome, so GUIDE-Seq designated off-targets assigned to ≥0.1% of the entire read aligned to the reference genome. In subsequent experiments, the gRNA was again delivered to the same cells and editing was assayed with rhAmpSeq. Targets were called “edited” if ≥1% or more indels were observed in the treated state than in the untreated control sample. [Figure 3] This figure illustrates the variation in GUIDE-Seq tag integration rates. The figure shows the percentage of tag integration (normalized to edit %) for 118 unique Cas9 on / off-target sites with InDel editing in an rhAmpSeq panel targeting GUIDE-Seq-designated on / off-target loci of guide sequences targeting the RAG1, RAG2, and EMX1 genes. Each guide, along with a 34-base pair GUIDE-Seq, dsODN tag, was co-delivered by nucleofection to HEK293 cells stably expressing Cas9. After 72 hours, DNA was extracted, amplified by rhAmpSeq multiplex PCR, sequenced by Illumina® MiSeq, and analyzed by a custom pipeline. Normalized tag integration rates are calculated as the percentage obtained by dividing the number of reads sequenced at each target containing the tag sequence by the total number of reads containing alleles different from the reference genome (indicating Cas9 editing). [Figure 4]This figure shows the design of rhAmpSeq primers for heterogeneous sequence tags. The illustration illustrates the steps of the design process using the rhAmpSeq design pipeline, including the design of forward primers for the upper (1) and lower (2) chains, discarding unwanted primers, and selecting a tag-target primer (3) that has a 5' overlap but no 3' overlap sequence, and therefore the upper / lower chain primer dimer forms a hairpin. [Figure 5] This figure outlines the rhAmpSeq design pipeline used to construct overlapping primer designs. In the pipeline, known sequences are added to the 5' and 3' ends of each tag sequence, the inputs are quality-controlled, and assays (shown in Figure 4A) are designed for the upper and lower strands of each tag. Primers targeting each tag strand are paired so that at least four nucleotides of the 3' end of the RNA nucleotides overlap between primers targeting the same tag, and primer pairs are ranked and selected. The acronyms hg38 and mm38 represent the human and mouse genome types, respectively. [Figure 6] This figure illustrates hairpin formation when overlapping primers generate PCR amplicons. The figure shows representative target sequences and hairpin PCR products of undesirable short amplicons from overlapping primer regions having a 5' primer tail end complementary to the 3' and 5' ends of the PCR product. [Figure 7] This figure shows the number of target sites (black bars) that incorporate either an explicitly specified single tag (sequence number 9-40) or a pool of tags listed in Table 5 (sequence numbers 9-40, 45-268). The striped bars (CTLmax) indicate the maximum number of target sites that can theoretically be found using a combination of single tags (sequence numbers 9-40) (23 out of a maximum of 32 sites). Pool A1 contains all single tags (sequence numbers 9-40). Pools B1-B6 each contain 16 different tags (sequence numbers 45-268). Pool C1 contains all tested tags (sequence numbers 9-40, 45-268). Embedding phenomena were determined using an in-house data analysis tool. [Figure 8]This figure shows the number of target sites (black bars) in which an explicitly specified single tag (sequence number 9-40) or a pool of tags listed in Table 5 (sequence numbers 9-40, 45-268) is incorporated. The striped bars (CTLmax) indicate the maximum number of target sites that can theoretically be found using a combination of single tags (sequence numbers 9-40) (47 out of a maximum of 53 sites). Pool A1 contains all single tags (sequence numbers 9-40). Pools B1-B6 each contain 16 different tags (sequence numbers 45-268). Pool C1 contains all tested tags (sequence numbers 9-40, 45-268). Incorporation phenomena were determined using an in-house data analysis tool. [Modes for carrying out the invention]

[0016] Described herein are methods for detecting and specifying on- and off-target CRISPR editing sites with improved accuracy and sensitivity. Information on the intracellular state is maintained by constructing based on conventional in vivo specification methods. Sensitivity is increased by co-delivering a set of unique pre-defined sequence tags. In one aspect, the co-delivery of a set of pre-defined unique tags may be in the range of 13 to 80 base pairs. In another aspect, the co-delivered set of pre-defined tags may be composed of 13-base pair tag array tags, 26-base pair tag array tags, 39-base pair tag array tags, 52-base pair tag array tags, 65-base pair tag array tags, or 78-base pair tag array tags. In another aspect, the unique pre-defined tag is a set of 52-base pair tag array tags (as the length of the array tag increases, the ability of rhPrimer to find good primer landing sites improves). This limitation is thought to be alleviated by using diverse tag sequences different from the human and mouse genomes. Specificity is improved by constructing the rhAmp technology of Integrated DNA Technologies (IDT) that uses RNAaseH2 (Pyrococcus abyssi) to release blocks of primers accurately annealed to the target, thereby reducing the incidence of false priming. Specificity can be further enhanced by simply using reads containing the tag sequence expected at the 5' end to specify the target. Incorporating suppression PCR into this method makes it more affordable. Conventional in vivo methods (e.g., GUIDE-seq and iGUIDE) require parallel PCR reactions (2-pool amplification) to amplify by annealing and extending to the upper and lower strands of the tag. Here, suppression PCR can be used to amplify both pools simultaneously without causing problematic dimer sequences.

[0017] The GUIDE-Seq dsDNA tag was co-delivered with one guide RNA to HEK293 cells constitutively expressing Cas9 using nucleofection. For teachings such as this, see U.S. Patent No. 9,822,407, which is incorporated herein by reference. All four different guide RNAs were tested in this manner. In cells, a ribonucleoprotein complex (RNP) forms between the expressed Cas9 and the guide RNA, introducing a double-strand break. The repaired break can contain the co-delivered tag. After delivery, the cells were incubated and the resulting DNA was extracted. Target amplification was performed according to the GUIDE-Seq protocol and assayed with a modified version of the GUIDE-Seq analysis pipeline (github.com / aryeelab / guideseq). The designated targets were compared across three biological replicates (co-delivery of unique guide RNA + tag). Not all of the designated targets were common to all biological replicates (common / designated targets overall: 7 / 31, 6 / 19, 2 / 4, 3 / 5 respectively, see Table 1). However, more than 90% of all reads attributable to any target were attributable to common targets (average, see Figure 1).

[0018]

Table 1-1

[0019]

Table 1-2

[0020] Furthermore, designated targets may not be able to be replicated or detected using orthogonal methods. Using the GUIDE-Seq method, GUIDE-Seq DNA tags were co-delivered with each of six guides to HEK293 cells that constitutively express Cas9 using nucleofection (each tag was delivered with one guide RNA). The rhAmpSeq multiplex amplicon panel was designed to amplify the designated targets and the inventors quantified editing in biological replicates. Of the 331 targets designated by GUIDE-Seq, only 41 (12%) could be verified by rhAmpSeq (see Figure 2).

[0021] dsDNA tag sequences used in NHEJ repair, co-delivered with guide RNA to stably expressing CRISPR cell lines, are incorporated at varying rates. Here, GUIDE-Seq dsDNA tags were co-delivered to HEK293 cells constitutively expressing Cas9, along with each of the six guides. In another embodiment, dsDNA tag sequences used in NHEJ repair, co-delivered with CRISPR RNPs, are incorporated at varying rates. Here, GUIDE-Seq dsDNA tags were co-delivered to HEK293 cells constitutively expressing Cas9, along with each of the six guides. rhAmpSeq panels were developed to amplify specified targets, and tag integration rates were analyzed in biological replicas using a custom analysis pipeline. These results indicate that, depending on the target, tags are incorporated in 0–85% of edited genomic copies (see Figure 3). While not bound by any theory, this rate is assumed to vary depending on the sequence context.

[0022] This specification describes a method for improving the signal-to-noise ratio by combining Integrated DNA Technology's rhAmpSeq® technology, suppression PCR, and novel heterogeneous DNA sequence design that specifies off-target editing sites for nucleases within the host genome.

[0023] In this method, Cas9, sgRNA or bipartite CRISPR RNA:transactivating crRNA (crRNA:tracrRNA) double helix, and one or more double-stranded DNA (dsDNA) tag sequences are delivered to cells. Simultaneous delivery of multiple tags can improve tag integration at off-target sites (see below). The tag sequences have sequence contents that are significantly different (i.e., heterogeneous) from the host genome. After a nuclease introduces a DSB, NHEJ repair inserts the tag sequence into the target site, forming a known primer landing site. After the cells have time to repair the DSB and possibly divide further (e.g., after 72 hours), the genomic DNA is isolated and fragmented (e.g., Covaris® shear, enzyme-based shear, Tn5, etc.), and a universal adapter sequence containing a unique molecular index (UMI) is ligated to the fragmented DNA, with unligated material removed. The DNA fragments are then amplified by targeting primers to the tag and universal adapter sequence (first PCR). Using universal primers, a sample index (PCR2) is added, the amplified material is concentration-normalized and pooled with other samples, and the pooled material is sequenced using Illumina® (or similar) instruments. The sequenced reads are aligned to a reference genome, and the loci mapped by numerous reads can be used to specify on / off target locations.

[0024] Heterogeneous sequences have a GC content of 40-90%, maximum homopolymer length A:2, C:3, G:2, T:2, weighted homopolymer ratio <20, and self-folding T m <50°C, and self-dimer T m The design was achieved by generating 1M random 13-mer sequences with <50°C>. From the list of sequences, sequences that were perfectly aligned to a human (GRCh38.p2;hg38) or mouse (GRCh38.p4;mm38) reference genome, or sequences containing problematic motifs (homopolymers, most GG or CC dinucleotide motifs), were removed to obtain 479 sequences.

[0025] To design the 52-base pair tag sequences described herein, 49 13-mer oligo sequences containing ≤1 C or G dinucleotides were selected, generating 10,000 unique combinations of four 13-mer sequences. The length of each concatenated sequence (e.g., pasting four 13-mer sequences in a single line using software) is 52 nucleotides. Each 52-nucleotide tag sequence was then aligned to the human (GRCh38.p2) and mouse (GRChm38.p4) genomes using an internally modified version of bwa ​​called bwa-psm. Running bwa-psm returns all possible secondary matches up to a defined threshold. A set of tag sequences (SEQ ID NOs. 1-2) intended to function as a group, with no similarity to the human or mouse genome, was designed (Max Seed Size: 7, Seed Edit Distance: 2, Max Edit Distance: 21, Max Gap Open: 2, Max Gap Expansion: 3, Mismatch Penalty: 1, Gap Open Penalty: 1, Gap Expansion Penalty: 1).

[0026] The overlapping rhAmpSeqV1 primers (SEQ ID NOs. 3-4) are designed to complement the upper and lower strands of the tag, as well as the 5' end of the adapter sequence (SEQ ID NOs. 6) (Figure 4). The tag-specific primers (SEQ ID NOs. 3-4) contain a 5' universal tail sequence matching the SP1 and SP2 primer sequences (SEQ ID NOs. 7-8), a locus-specific segment, six ribonucleotides (rN) from the 3' end, a 3' end mismatch, and a 3' end block (3'-C3 spacer). The adapter-specific primer (SEQ ID NOs. 5) targets the 5' end of the 5'-P5 adapter sequence (SEQ ID NOs. 6), which contains an intrinsic molecular index (UMI) sequence (Table 2). The primers are designed to target the positive and negative strands of such annealed tags, and if these primers unexpectedly dimerize, the resulting product forms a hairpin, removing the oligo from the available reaction template (e.g., suppression PCR) (Figure 6A-B). The primer sequence that targets the tag is an IDT (an algorithm with a public UI). Based on a proprietary design algorithm designed and executed by (internal copy of the idtdna: www.idtdna.com / site / account?ReturnURL= / site / order / designtool / index / RHAMPSEQ), the primer pairs are selected to best perform in amplifying the target template sequence (Figure 5). The primer sequences were evaluated for nonspecific binding to all other tag sequences as well as to both human and mouse primary genome assemblies, and were found to be less likely to form off-target amplicons in combination with the presence of a universal adapter sequence and human or mouse genomic DNA.

[0027] The primers were desired to function in pairs, with one tag-specific primer (upper or lower strand) paired with an adapter-specific primer (SEQ ID NO: 5). This would result in amplification of a molecule containing the tag, gDNA, and a portion of the adapter sequence when amplified using suppression PCR (Figure 4).

[0028] [Table 2]

[0029] One embodiment described herein is a method for identifying and designating on- and off-target CRISPR editing sites with improved accuracy and sensitivity, comprising: (a) co-delivering a guide sequence RNA (sgRNA) or a two-part CRISPR RNA:trans-activated crRNA (crRNA:tracrRNA) double helix and one or more tag sequences to cells; (b) incubating the cells for a set period of time; (c) isolating genomic DNA from the cells, fragmenting the genomic DNA, and ligating the fragmented genomic DNA to a unique molecular index containing a universal adapter sequence; and (d) tag The method comprises the steps of (e) amplifying a ligated DNA fragment using a primer that targets a universal adapter sequence to generate a first set of amplified sequences, (f) amplifying the first set of amplified sequences using a universal sequencing primer that targets the tail of a Tag-pTOP or Tag-pBOT primer to generate a second set of amplified sequences, (g) sequencing the pooled sequences to obtain sequencing data, and (g) identifying on / off target CRISPR editing loci. In one embodiment, the universal sequencing primer targets the SP1 or SP2 sequence (SEQ ID NOs. 7, 8) tail of a Tag-pTOP or Tag-pBOT primer to generate a second set of amplified sequences. In another embodiment, the universal sequencing primer targets a pre-designed non-homologous sequence (Table 6, SEQ ID NOs. 269-273) tail of a Tag-pTOP or Tag-pBot to generate a second set of amplified sequences. In yet another embodiment, a universal primer targets a pre-designed 13-mer tail of a Tag-pTOP or Tag-pBot primer to generate a second set of amplified sequences. In one embodiment, step (g) includes performing on a processor (i) aligning sequence data to a reference genome, (ii) identifying on / off target CRISPR editing loci, and (iii) aligning, analyzing, and outputting the resulting data as a table or figure. In another embodiment, the method further includes, after step (e), (e1) normalizing the second set of amplified sequences to generate a concentration-normalized library, pooling the normalized library with other samples to generate a pooled library, and continuing through steps (f) to (i). In one embodiment, step (d) uses suppression PCR. In another embodiment, cells constitutively express the Cas enzyme, or are co-delivered with a Cas expression vector, or co-delivered with a Cas protein, or co-delivered with a CasRNP complex. In another embodiment, cells constitutively express the Cas9 enzyme, are co-delivered with a Cas9 expression vector, are co-delivered with the Cas9 protein, or are co-delivered with the Cas9RNP complex.In another embodiment, the cells include human or mouse cells. In another embodiment, the time is approximately 24 to approximately 96 hours. In another embodiment, multiple tag sequences are delivered simultaneously. In another embodiment, the tag sequence includes a double-stranded deoxyribooligonucleotide (dsDNA) containing 52 base pairs. In another embodiment, the tag sequence includes a 5' terminal phosphate group, as well as phosphorothioate bonds between the 1st and 2nd, 2nd and 3rd, 50th and 51st, and 51st and 52nd nucleotides. In another embodiment, the tag sequence includes double-stranded DNA containing the upper and lower strand pairs of sequence numbers 9-40 or 45-268.

[0030] Another embodiment described herein is an on- and off-target CRISPR editing site identified or designated using the method described herein.

[0031] Another embodiment described herein is a method for designing a 52-base pair tag sequence, wherein the processor has (a) a GC content of 40-90%, a maximum homopolymer length of A:2, C:3, G:2, T:2, a weighted homopolymer ratio of <20, and a self-folding T m <50°C, and self-dimer T m A method comprising the steps of: (b) randomly generating 13-nucleotide sequences at <50°C; (c) removing sequences that perfectly align to a particular genome, or sequences that are homopolymers or GG or CC dinucleotide motifs to obtain a set of 13-mers; (d) selecting a subset of 13-mer sequences containing one or fewer CC or GG dinucleotide motifs; (e) concatenating four of the 13-mer subset sequences to form a random 52-mer sequence; (f) aligning the random 52-mer sequence to a genome; (g) removing random 52-mer sequences that are similar to the genome to generate a subset of 52-mer sequences; and (h) outputting the subset of 52-mer sequences and generating a complementary strand to produce a double-stranded 52-base pair tag sequence. In one embodiment, the genome is human or mouse. In one embodiment, the 52-base pair tag sequence is not complementary to the genome. In another embodiment, this method The method further includes the step of designing primers for a 52-base pair tag sequence. In another embodiment, the 52-base pair tag sequence includes a 5' terminal phosphate group, as well as phosphorothioate bonds between the 1st and 2nd, 2nd and 3rd, 50th and 51st, and 51st and 52nd nucleotides of the 52-base pair tag sequence. In another embodiment, the method further includes the step of synthesizing oligonucleotides containing a 52-base pair tag sequence, a complement of a 52-base pair tag sequence, or primers for a 52-base pair tag sequence.

[0032] Another embodiment described herein is one or more 52-base-pair tag sequences designed using the method described herein. In one embodiment, the 52-base-pair tag sequence comprises double-stranded DNA including complementary upper and lower strand pairs of sequence numbers 9-40 or 45-268.

[0033] Another embodiment described herein is a method for designing primers and adapter primers partially complementary to the 52-base pair tag sequences described herein, comprising the steps of: (a) designing primers partially complementary to the upper and lower strands of the tag sequence; and (b) designing an adapter primer partially complementary to the upper strand of the adapter sequence, wherein the tag primer comprises a 5' universal tail sequence, a locus-specific segment, 6 ribonucleotides (rN) at the 3' end, a 3' end mismatch, and a 3' end block (3'-C3 spacer) complementary to the SP1 or SP2 sequence (SEQ ID NOs. 7, 8), and the adapter primer comprises a sequence complementary to the SP1 or SP2 sequence (SEQ ID NOs. 7, 8). In one embodiment, primers partially complementary to the upper and lower strands of the tag sequence contain a sequence complementary to the SP1 sequence, and the adapter primer contains a sequence complementary to the SP2 sequence; or, primers partially complementary to the upper and lower strands of the tag sequence contain a sequence complementary to the SP2 sequence, and the adapter primer contains a sequence complementary to the SP1 sequence. In another embodiment, amplification of a nucleic acid molecule using primers complementary to the upper and lower strands of the tag sequence and a primer complementary to the upper strand of the adapter sequence produces a PCR product containing a portion of the tag sequence, an sgDNA sequence, and an adapter sequence. In yet another embodiment, the method further includes the step of synthesizing oligonucleotides containing the sequences of the forward and reverse tag primers and the adapter primer.

[0034] In another embodiment described herein, a 52-base pair tag sequence and primers partially complementary to the 52-base pair tag sequence are designed and selected using an algorithm that predicts whether the primers are likely to be partially complementary and whether they tend to form primer dimers.

[0035] Another embodiment described herein is one or more primers and one or more adapter primers that are partially complementary to a 52-base pair tag sequence designed using the method described herein. In one embodiment, the primers partially complementary to the 52-base pair tag sequence include the sequences of SEQ ID NOs: 3, 4, and the adapter primer includes the sequence of SEQ ID NO: 5.

[0036] Another embodiment described herein involves the use of one or more double-stranded 52-base pair tag sequences to identify on- and off-target CRISPR editing sites.

[0037] It will be apparent to those skilled in the art that the compositions, formulations, methods, processes, and uses described herein can be appropriately modified and adapted without departing from the scope of any of their embodiments or aspects. The compositions and methods provided are illustrative and are not intended to limit the scope of any particular embodiment. All of the various embodiments, aspects, and options disclosed herein can be combined in any variation or iteration. The scope of the methods and processes described herein is as described herein. This disclosure includes all actual or possible combinations of embodiments, aspects, options, examples, and selections. Methods described herein may omit any component or step, substitute any component or step disclosed herein, or include any component or step disclosed elsewhere herein. It should also be understood that embodiments include, and otherwise may be performed by, various combinations of hardware, software, and electronic components. For example, various microprocessors and application-specific integrated circuits ("ASICs") may be used, as well as software in various languages. Servers and various computing devices may also be used, and may include one or more processing units, one or more computer-readable media, one or more input / output interfaces, and various connections (e.g., system buses) for connecting components. If the meaning of any term in any patent or publication incorporated by reference conflicts with the meaning of a term used herein, the meaning of the term or phrase in this disclosure shall prevail. Furthermore, this disclosure merely discloses and describes exemplary embodiments. All patents and publications cited herein are incorporated herein by reference with respect to their specific teachings.

[0038] The various embodiments and aspects of the present invention described herein are summarized in the following sections:

[0039] Item 1. A method for identifying and specifying on- and off-target CRISPR editing sites with improved accuracy and sensitivity, (a) A step of simultaneously delivering a guide sequence RNA (sgRNA) or a bipartite CRISPR RNA:trans-activated crRNA (crRNA:tracrRNA) double helix, one or more tag sequences, and an RNA guide endonuclease to a cell, (b) A step of incubating the cells for a sufficient amount of time for double-strand breaks to occur, (c) The steps of isolating genomic DNA from cells, fragmenting the genomic DNA, and linking the fragmented genomic DNA to a unique molecular index containing a universal adapter sequence, (d) A step of amplifying the ligated DNA fragment using primers that target the tag and universal adapter sequence to generate a first set of amplified sequences, (e) A step of amplifying a first set of amplified sequences using a universal sequencing primer that targets the tail of a Tag-pTOP or Tag-pBOT primer to generate a second set of amplified sequences, (f) The steps of determining the sequence of the pooled sequences and obtaining the sequence determination data, (g) A method comprising the step of identifying an on / off target CRISPR editing locus. Item 2. The method according to Item 1, wherein a universal sequencing primer targets the SP1 or SP2 sequence (SEQ ID NOs. 7, 8) tail of a Tag-pTOP or Tag-pBOT primer to generate a second set of amplified sequences. Item 3. The method according to item 1 or 2, wherein a universal sequencing primer targets a pre-designed non-homologous sequence (SEQ ID NOs. 269-273) tail of a Tag-pTOP or Tag-pBot primer to generate a second set of amplified sequences. Item 4. The method according to any one of items 1 to 3, wherein a universal sequencing primer targets a pre-designed 13-mer tail of a Tag-pTOP or Tag-pBot primer to generate a second set of amplified sequences. Item 5. Step (g) is performed by the processor, Section 6. Aligning sequence data to the reference genome. (a)(ii) Identify on / off-target CRISPR editing loci, and (b)(iii) Arrange, analyze, and store the resulting data in a custom format file, table, or The method described in any one of items 1 to 4, which includes performing the output as a figure. Section 7. After step (e), (a)(e1) The method according to any one of claims 1 to 5, further comprising the steps of specifying a second set of amplified sequences to generate a concentration-normalized library, and pooling the normalized library with other samples to generate a pooled library, and continuing through steps (f) to (i). Item 8. The method described in any one of items 1 to 6, wherein step (d) uses the suppression PCR method. Item 9. The method according to any one of items 1 to 7, wherein the RNA guide endonuclease comprises an endogenously expressed Cas enzyme, a Cas expression vector, a Cas protein, or a CasRNP complex. Item 10. The method according to any one of items 1 to 8, wherein the RNA guide endonuclease comprises an endogenously expressed Cas9 enzyme, a Cas9 expression vector, a Cas9 protein, or a Cas9RNP complex. Item 11. The method according to any one of items 1 to 9, wherein the cells include human or mouse cells. Item 12. The method described in any one of items 1 to 10, wherein the duration is approximately 24 hours to approximately 96 hours. Item 13. The method described in any one of items 1 to 11, wherein multiple tag sequences are delivered simultaneously. Item 14. The method according to any one of items 1 to 12, wherein the tag sequence comprises a double-stranded deoxyribooligonucleotide (dsDNA) containing 52 base pairs. Item 15. The method according to any one of items 1 to 13, wherein the tag sequence includes a 5' terminal phosphate group and phosphorothioate bonds between the 1st and 2nd, 2nd and 3rd, 50th and 51st, and 51st and 52nd nucleotides. Item 16. The method according to any one of items 1 to 14, wherein the tag sequence comprises double-stranded DNA containing complementary upper and lower strand pairs of sequence numbers 1 to 2 or 7 to 268. Section 17. On- and off-target CRSPR editing sites identified or specified using any one of the methods in Sections 1 through 15.

[0040] Section 18.52 A method for designing a base-pair tag sequence, wherein a processor, (a) GC content 40-90%, maximum homopolymer length A:2, C:3, G:2, T:2, weighted homopolymer ratio <20, self-folding T m <50°C, and self-dimer T m A step of randomly generating a 13-nucleotide sequence at <50°C, (b) A step of removing sequences that are a perfect match to a specific genome, or sequences that are homopolymers or GG or CC dinucleotide motifs, to obtain a set of 13-mers, (c) A step of selecting a subset of 13-mer sequences containing one or fewer CC or GG dinucleotide motifs, (d) A step of concatenating four 13-mer subset sequences to form a random 52-mer sequence, (e) A step of aligning random 52-mer sequences in the genome, (f) A step of generating a subset of 52-mer sequences by removing random 52-mer sequences that are similar to the genome, A method comprising the steps of (g) outputting a subset of 52-mer sequences, generating a complementary strand, and generating a double-stranded 52-base pair tag sequence. Item 19. The method described in Item 17, wherein the genome is human or mouse. The method described in section 20.52, wherein the 2-base pair tag sequence is not complementary to the genome. The method according to any one of sections 17 to 19, further comprising the step of designing primers for a 2-base pair tag sequence. The method according to any one of the claims 17-20, wherein the 22.52 base pair tag sequence includes a 5' terminal phosphate group and a phosphorothioate bond between the 1st and 2nd, 2nd and 3rd, 50th and 51st, and 51st and 52nd nucleotides of the 52 base pair tag sequence. The method according to any one of the claims 17 to 21, further comprising the step of synthesizing an oligonucleotide comprising a 52-base pair tag sequence, a complement of a 52-base pair tag sequence, or a primer of a 52-base pair tag sequence. Section 24. One or more 52-base pair tag sequences designed using the methods of Sections 17-22. The 52-base-pair tag sequence described in item 23, wherein the 52-base-pair tag sequence contains double-stranded DNA including the upper and lower strand pairs of sequence numbers 1-2 or 7-268.

[0041] Section 26. A method for designing primers and adapter primers that are partially complementary to the 52-base pair tag sequences described in Section 23, wherein a processor is used. (a) A step of designing tag primers that are partially complementary to the upper and lower strands of the tag sequence, (b) the step of designing an adapter primer that is partially complementary to the upper chain of the adapter sequence, (c) Here, (d) Tag primer contains a 5' universal tail sequence, (e) A method in which the adapter primer contains a sequence complementary to the tail of the Tag-pTOP or Tag-pBOT primer. The method according to item 25, wherein the 5' universal tail sequence is complementary to an SP1 or SP2 sequence (sequences 7, 8), a locus-specific segment, six ribonucleotides (rN) from the 3' end, a 3' end mismatch, a 3' end block (3'-C3 spacer), a pre-designed non-homologous sequence (sequences 269-273), or a pre-designed 13-mer sequence. Item 28. The method according to Item 25 or 26, wherein a primer partially complementary to the upper and lower strands of the tag sequence contains a tail sequence complementary to the SP1 sequence (SEQ ID NO: 7), and the adapter primer contains a sequence complementary to the SP2 sequence (SEQ ID NO: 8) tail of the Tag-pTOP or Tag-pBOT primer, or a primer partially complementary to the upper and lower strands of the tag sequence contains a tail sequence complementary to the SP2 sequence (SEQ ID NO: 8), and the adapter primer contains a sequence complementary to the SP1 sequence (SEQ ID NO: 7) tail of the Tag-pTOP or Tag-pBOT primer. Item 29. The method according to any one of items 25 to 27, wherein amplification of a nucleic acid molecule using primers complementary to the upper and lower strands of the tag sequence and a primer complementary to the upper strand of the adapter sequence generates a PCR product containing a portion of the tag sequence, an sgDNA sequence, and an adapter sequence. Item 30. The method according to any one of items 25 to 28, further comprising the step of synthesizing oligonucleotides containing sequences of forward and reverse tag primers and adapter primers. The method according to any one of sections 17-21 and 25-29, wherein primers partially complementary to 52-base pair tag sequences and 52-base pair tag sequences are designed and selected using an algorithm that predicts whether the primers are likely to be partially complementary or whether they tend to form primer dimers.

[0042] Section 32. One or more primers and one or more adapter primers that are partially complementary to a 52-base pair tag sequence designed using the methods described in Sections 22-25. Item 33. The primer includes the sequences of SEQ ID NOs: 3, 4 and the adapter primer, The primer according to item 32, wherein the adapter primer contains the sequence of Sequence ID No. 5. Item 34. Use of one or more double-stranded 52-base pair tag sequences to identify on- and off-target CRISPR editing sites.

[0043] References 1.Wienert et al., “Unbiased detection of CRISPR off-targets in vivo using DISCOVER-seq,” Science 364(6437): 286-289 (2019). 2.Nobles et al., “IGUIDE: An improved pipeline for analyzing CRISPR cleavage specificity,” Genome Biol. 20(14): 4-9 (2019). 3. Tsai et al., “GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases,” Nature Biotechnol. 33(2): 187-197 (2015). 4. Yan et al., “BLISS is a versatile and quantitative method for genome-wide profiling of DNA double-strand breaks,” Nature Commun. 8: 15058 (2017). 5. Tsai et al., “CIRCLE-seq: a highly sensitive in vitro screen for genome-wide CRISPR-Cas9 nuclease off-targets,” Nature Methods 14(6): 607-614 (2017). 6. Cameron et al., “Mapping the genomic landscape of CRISPR-Cas9 cleavage,” Nature Methods 14(6): 600-606 (2017). 7. Char and Moosburner, “Unraveling CRISPR-Cas9 genome engineering parameters via a library-on-library approach,” Nature Methods 12(9): 823-826 (2015). 8. Rand et al., “Headloop suppression PCR and its application to selective amplification of methylated DNA sequences,” Nucleic Acids Res. 33(14): e127 (2005).<Xref

Example

[0044] Example 1 This experiment demonstrates that using 52-base pair long double-stranded DNA tags and various gene sequences increases the efficiency of tag integration. The sequences used are shown in Tables 3–5. The double-stranded tags were generated by hybridization of the upper strand and the complementary lower strand (Tables 3–4; SEQ ID NOs. 9–40 or 45–268). Sixteen different tag designs were separately introduced into HEK293 cells constitutively expressing Cas9, along with guide RNA targeting the EMX1 locus. Alternatively, either a pool of 16 tags or a single pool of 112 tags was introduced into HEK293 cells constitutively expressing Cas9, along with guide RNA targeting the EMX1 locus. Guide RNA was electroporated at a concentration of 10 μM, while single or pooled tags were delivered at a final concentration of 0.5 μM. The level of tag integration was determined by targeted amplification using rhAmpSeq primers (SEQ ID NOs: 3-4) to enrich known on- and off-target sites of the EMX1 guide RNA. The rhAmpSeq pool for EMX1 consists of 32 sites representing empirically determined on- and off-target loci. The amplified products were sequenced using Illumina® MiSeq, and the level of tag integration was determined using custom software. This example demonstrates that tag integration efficiency differs individually among single-tag constructs, and is therefore sequence-dependent, ranging from 6 (CTL021) to 13 (CTL169, CTL079, CTL002) out of a maximum of 32 sites (single tag, Figure 7). By mathematically combining the single-tag results, a hypothetical number of 23 sites was calculated (CTLmax, Figure 7). The hypothesis that combining the tag pool increases the likelihood of tag integration was tested and demonstrated (pooled tag, Table, Figure 7). Pool A1 consisted of tags represented by single tags (see Table 5), and tag embedding was detected in 21 out of a maximum of 32 locations, demonstrating a higher rate than that achieved by any single tag. Similarly, Pool B3 showed tag embedding in 21 out of a maximum of 32 locations. Again, variability between pools was demonstrated (pooled tags, Figure 7), suggesting that optimizing tag design can potentially maximize tag embedding.

[0045] [Table 3-1]

[0046] [Table 3-2]

[0047] Example 2 This experiment demonstrates that using 52-base pair long double-stranded DNA tags and various gene sequences increases the efficiency of tag integration. The sequences used are shown in Tables 3–5. The double-stranded tags were generated by hybridization of the upper strand and the complementary lower strand (SEQ ID NOs. 9–40 or 45–268). Sixteen different tag designs were separately introduced into HEK293 cells constitutively expressing Cas9, along with guide RNA targeting the AR locus. Alternatively, either a pool of 16 tags or a single pool of 112 tags was introduced into HEK293 cells constitutively expressing Cas9, along with guide RNA targeting the AR locus. Guide RNA was electroporated at a concentration of 10 μM, while single or pooled tags were delivered at a final concentration of 0.5 μM. Tag integration levels were determined by targeted amplification using rhAmpSeq primers (SEQ ID NOs. 3–4) to enrich known on- and off-target sites of the AR guide RNA. The AR rhAmpSeq pool consists of 53 sites representing empirically determined on and off target gene loci. The amplified products were sequenced with Illumina® MiSeq, and the tag integration level was determined using custom software. In this example, the tag integration efficiency was 35 out of 53 sites (CTL085). We show that the single-tag constructs differ individually and are therefore sequence-dependent in the range from CTL134 to 41 (CTL002) (single tag, Table 5, Figure 8).

[0048] By mathematically combining the results of single tags, a virtual number of 47 sites was calculated (CTLmax, Figure 8). The hypothesis that combining tag pools increases the likelihood of tag embedding was tested and demonstrated (pooled tags, Table 5, Figure 8). Pool B4 (see Table 5) showed 44 tag embedding phenomena out of a maximum of 53 sites, demonstrating a higher rate than achieved by any single tag. Again, variability between pools was demonstrated (pooled tags, Table 5, Figure 8), suggesting that optimizing tag design can potentially maximize tag embedding.

[0049] [Table 4-1]

[0050] [Table 4-2]

[0051] [Table 4-3]

[0052] [Table 4-4]

[0053] [Table 4-5]

[0054] [Table 4-6]

[0055] [Table 4-7]

[0056] [Table 4-8]

[0057] Table 4-9

[0058] Table 5

[0059] Table 6

Claims

1. A method for identifying and specifying on- and off-target CRISPR editing sites with improved accuracy and sensitivity, (a) A step of simultaneously delivering a guide sequence RNA (sgRNA) or a bipartite CRISPR RNA:trans-activated crRNA (crRNA:tracrRNA) double helix, one or more tag sequences, and an RNA guide endonuclease to a cell, (b) The step of incubating the cells for a sufficient time for double-strand breaks to occur, (c) The steps of isolating genomic DNA from the cells, fragmenting the genomic DNA, and linking the fragmented genomic DNA to a unique molecular index containing a universal adapter sequence, (d) A step of amplifying the ligated DNA fragment using primers that target the tag and the universal adapter sequence to generate a first set of amplified sequences, (e) A step of generating a second set of amplified sequences by amplifying the first set of amplified sequences using a universal sequencing primer that targets the tail of a Tag-pTOP or Tag-pBOT primer, (f) The steps of determining the sequence of the pooled sequences and obtaining the sequence determination data, (g) A step of identifying on / off target CRISPR editing loci, The method, including the method described above.

2. The method according to claim 1, wherein the universal sequencing primer targets the SP1 or SP2 sequence (SEQ ID NOs) tail of the Tag-pTOP or Tag-pBOT primer to generate a second set of amplified sequences.

3. The method according to claim 1 or 2, wherein the universal sequencing primer targets a pre-designed non-homologous sequence (SEQ ID NOs. 269-273) tail of the Tag-pTOP or Tag-pBot primer to generate a second set of amplified sequences.

4. The method according to any one of claims 1 to 3, wherein the universal sequencing primer targets a pre-designed 13-mer tail of the Tag-pTOP or Tag-pBot primer to generate a second set of amplified sequences.

5. Step (g) is performed by the processor. (i) Aligning the sequence data with the reference genome, (ii) Identifying on / off-target CRISPR editing loci, and (iii) The method according to any one of claims 1 to 4, comprising performing arrangement, analysis, and outputting the resulting data as a file, table, or figure in a custom format.

6. After step (e), The method according to any one of claims 1 to 5, further comprising the step of (e1) normalizing the second set of amplified sequences to generate a concentration-normalized library, and pooling the normalized library with other samples to generate a pooled library, and continuing through steps (f) to (i).

7. The method according to any one of claims 1 to 6, wherein step (d) uses a suppression PCR method.

8. Any one of claims 1 to 7, wherein the RNA guide endonuclease comprises an endogenously expressed Cas enzyme, a Cas expression vector, a Cas protein, or a CasRNP complex. The method described in item 1.

9. The method according to any one of claims 1 to 8, wherein the RNA guide endonuclease comprises an endogenously expressed Cas9 enzyme, a Cas9 expression vector, a Cas9 protein, or a Cas9 RNP complex.

10. The method according to any one of claims 1 to 9, wherein the cells include human or mouse cells.

11. The method according to any one of claims 1 to 10, wherein the aforementioned time is approximately 24 hours to approximately 96 hours.

12. The method according to any one of claims 1 to 11, wherein multiple tag sequences are delivered simultaneously.

13. The method according to any one of claims 1 to 12, wherein the tag sequence comprises a double-stranded deoxyribooligonucleotide (dsDNA) containing 52 base pairs.

14. The method according to any one of claims 1 to 13, wherein the tag sequence includes a 5' terminal phosphate group and phosphorothioate bonds between the 1st and 2nd, 2nd and 3rd, 50th and 51st, and 51st and 52nd nucleotides.

15. The method according to any one of claims 1 to 14, wherein the tag sequence comprises double-stranded DNA including complementary upper and lower strand pairs of sequence numbers 1 to 2 or 7 to 268.

16. On and off-target CRISPR editing sites identified or designated using the method of any one of claims 1 to 15.

17. A method for designing a 52-base pair tag sequence, using a processor, (a) GC content 40-90%, maximum homopolymer length A: 2, C: 3, G: 2, T: 2, weighted homopolymer ratio < 20, self-folding T m <50°C, and self-dimer T m <A step of randomly generating a 13-nucleotide sequence at 50°C, (b) A step of obtaining a set of 13-mers by removing sequences that are perfectly aligned to a specific genome, or sequences that are homopolymers or GG or CC dinucleotide motifs, (c) Selecting a subset of the 13-mer sequence containing one or less CC or GG dinucleotide motifs, (d) The step of concatenating four of the 13-mer subset sequences to form a random 52-mer sequence, (e) The step of aligning the random 52-mer sequences in the genome, (f) A step of generating a subset of 52-mer sequences by removing random 52-mer sequences that are similar to the genome, (h) A step of outputting the subset of the 52-mer sequence, generating a complementary strand, and generating a double-stranded 52-base pair tag sequence, The method, which includes performing the following:

18. The method according to claim 17, wherein the genome is human or mouse.

19. The method according to claim 17 or 18, wherein the 52-base pair tag sequence is not complementary to the genome.

20. The method according to any one of claims 17 to 19, further comprising the step of designing primers for the 52 base pair tag sequence.

21. The method according to any one of claims 17 to 20, wherein the 52-base pair tag sequence includes a 5' terminal phosphate group and a phosphorothioate bond between the 1st and 2nd, 2nd and 3rd, 50th and 51st, and 51st and 52nd nucleotides of the 52-base pair tag sequence.

22. The method according to any one of claims 17 to 21, further comprising the step of synthesizing an oligonucleotide comprising the 52 base pair tag sequence, a complement of the 52 base pair tag sequence, or a primer of the 52 base pair tag sequence.

23. One or more 52-base pair tag sequences designed using the methods described in claims 17 to 22.

24. The 52-base-pair tag sequence according to claim 23, wherein the 52-base-pair tag sequence includes double-stranded DNA comprising the upper and lower strand pairs of sequence numbers 1-2 or 7-268.

25. A method for designing primers and adapter primers that are partially complementary to the 52-base pair tag sequence described in claim 23, comprising a processor, (a) The step of designing tag primers that are partially complementary to the upper and lower strands of the tag sequence, (b) The step of designing an adapter primer that is partially complementary to the upper chain of the adapter array, This includes performing the following: The tag primer includes a 5' universal tail sequence, The method wherein the adapter primer includes an array complementary to the tail of the Tag-pTOP or Tag-pBOT primer.

26. The 5' universal tail sequence consists of an SP1 or SP2 sequence (SEQ ID NOs: 7, 8), a locus-specific segment, six ribonucleotides (rN) from the 3' end, a 3' end mismatch, and a 3' end block (3'-C). 3 The method according to claim 25, wherein the spacer is complementary to a pre-designed non-homologous sequence (sequence numbers 269-273) or a pre-designed 13-mer sequence.

27. The method according to claim 25 or 26, wherein the primer partially complementary to the upper and lower chains of the tag sequence includes a tail sequence complementary to the SP1 sequence (SEQ ID NO: 7), and the adapter primer includes a sequence complementary to the SP2 sequence (SEQ ID NO: 8) tail of the Tag-pTOP or Tag-pBOT primer, or the primer partially complementary to the upper and lower chains of the tag sequence includes a tail sequence complementary to the SP2 sequence (SEQ ID NO: 8), and the adapter primer includes a sequence complementary to the SP1 sequence (SEQ ID NO: 7) tail of the Tag-pTOP or Tag-pBOT primer.

28. The method according to any one of claims 25 to 27, wherein amplification of a nucleic acid molecule using primers complementary to the upper and lower strands of the tag sequence and a primer complementary to the upper strand of the adapter sequence generates a PCR product comprising a portion of the tag sequence, an sgDNA sequence, and the adapter sequence.

29. The arrangement of the forward and reverse tag primers and the adapter primer The method according to any one of claims 25 to 28, further comprising the step of synthesizing an oligonucleotide comprising a column.

30. The method according to any one of claims 17-21 and 25-29, wherein a 52-base pair tag sequence and a primer partially complementary to the 52-base pair tag sequence are designed and selected using an algorithm that predicts whether the primers are likely to be partially complementary or whether they tend to form a primer dimer.

31. One or more primers and one or more adapter primers that are partially complementary to the 52-base pair tag sequence designed using the methods described in claims 22 to 25.

32. The primer according to claim 31, wherein the primer comprises the sequences of sequence numbers 3 and 4 and the adapter primer, and the adapter primer comprises the sequence of sequence number 5.

33. Use of one or more double-stranded 52-base pair tag sequences to identify on- and off-target CRISPR editing sites.