A method for assessing off-target risk of gene editing based on targeted enrichment sequencing

By screening off-target sequences with a variety of prediction software and combining it with targeted enrichment sequencing technology, the problems of insufficient sensitivity and specificity of gene editing off-target detection methods were solved, and efficient and accurate off-target site detection was achieved, ensuring the safety and reliability of gene editing.

CN119207553BActive Publication Date: 2025-10-10SHANGHAI WEIKE MEDICAL TESTING LABORATORY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411391939.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-08
Publication Date
2025-10-10
Estimated Expiration
2044-10-08

AI Technical Summary

Technical Problem

Existing gene editing off-target detection methods lack sensitivity and specificity, making it difficult to achieve efficient and accurate off-target site detection, especially in large-scale sample analysis, which affects the safety and reliability of gene therapy.

Method used

A variety of prediction software are used to screen off-target sequences, combined with targeted enrichment sequencing technology, to perform high-depth sequencing on the edited samples to be tested. Through bioinformatics analysis and molecular biology techniques, off-target sites are accurately identified and an off-target risk assessment report is generated.

Benefits of technology

It significantly improves the sensitivity and specificity of off-target site detection, reduces sequencing costs, ensures the safety and reliability of gene editing, can accurately identify low-frequency mutations and chromosomal translocation events, and provide a comprehensive off-target effect assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119207553B_ABST
    Figure CN119207553B_ABST
Patent Text Reader

Abstract

The application provides a method for evaluating off-target risk of gene editing based on targeted enrichment sequencing, comprising: screening off-target sequences and collecting them by using multiple prediction software respectively to obtain multiple off-target sequence lists, and selecting off-target sequences appearing in multiple off-target sequence lists and having high potential off-target risk as off-target sites; performing high-depth sequencing on the off-target sites in genomes of a to-be-tested sample and a control sample by using a TES method, comparing sequencing results with a reference genome, performing mutation detection and structural variation detection on the comparison results, obtaining mutation sites and structural variation sites existing only in the editing sample, and identifying authenticity of the off-target sites based on the mutation sites and the structural variation sites; evaluating off-target risk and editing effect; and visualizing the off-target site information and automatically generating an off-target risk evaluation report. The evaluation method has high sensitivity and high specificity, can improve accuracy and sensitivity of off-target detection of gene editing, and significantly reduces overall sequencing cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of biotechnology and genome editing technology, and particularly relates to a method for assessing off-target risks of gene editing based on targeted enrichment sequencing. Background Art

[0002] Since its introduction in 2012, CRISPR-Cas9 gene editing technology has garnered significant attention from researchers and research institutions worldwide due to its efficient editing capabilities and unique mechanism. CRISPR-Cas (clustered regularly interspaced short palindromic repeats and CRISPR associated) is a naturally occurring immune defense mechanism in bacteria and archaea, primarily used to combat viral invasion. Due to the programmable nature of its guide RNA and the pronounced activity of the Cas nuclease in various cells and tissues, it has been widely used in genome editing. These properties give the CRISPR-Cas system significant advantages and application value in genome editing.

[0003] The application of CRISPR-Cas9 technology in gene therapy is maturing, demonstrating revolutionary potential. Through precise gene modification or repair, this technology has opened up new avenues for treating genetic diseases. CRISPR-Cas9 plays a crucial role in the targeted repair of genetic diseases. It can precisely edit specific genetic defects, not only repairing disease-related genes but also enabling fine-tuning of gene function, bringing hope to patients with genetic diseases. For example, gene therapy research for blood disorders such as β-thalassemia and sickle cell anemia has made progress, with some entering clinical trials and demonstrating promising therapeutic prospects. CRISPR-Cas9 technology also demonstrates tremendous potential in cancer treatment research. By targeting oncogenes or tumor suppressor genes, researchers can gain deeper insights into the mechanisms of cancer development and progression. This in-depth molecular understanding provides a solid foundation for the development of new cancer treatment strategies, driving the advancement of personalized and precision medicine. With continued technological advancements and in-depth clinical trials, CRISPR-Cas9 is expected to revolutionize cancer treatment.

[0004] Although CRISPR technology has been hailed as a revolutionary tool in the field of gene editing, its application still faces some significant limitations. These limitations not only affect the editing accuracy and specificity of CRISPR technology, but also pose challenges to editing efficiency and delivery systems.

[0005] While the CRISPR-Cas9 system is capable of precise cutting at specific DNA sequences, it can also cause nonspecific cutting, inducing double-strand breaks in DNA at non-target sites, known as "off-target effects." These off-target effects can lead to unintended gene editing, potentially affecting cell function or triggering unforeseen consequences. In clinical applications, off-target effects can affect the safety and controllability of therapeutic effects, increasing uncertainty and risk during treatment. Therefore, researchers and scientists are actively exploring various methods to reduce the off-target effects of the CRISPR-Cas9 system to improve its safety and reliability in gene editing, such as optimizing the design of the Cas9 enzyme, improving the design of guide RNA (gRNA), double-strand repair inhibition, dual-guide RNA strategies, chemical modifications and regulatory systems, and the development of new gene editing tools.

[0006] While reducing off-target effects, detecting them is equally important. Existing technical solutions for detecting off-target effects in gene editing mainly include the following: Although whole genome sequencing (WGS) can fully sequence the genome, it is expensive when analyzing multiple samples or a large number of samples, it is difficult to perform high-depth sequencing, it is not easy to detect low-frequency off-target mutations, and the data analysis is complex and requires a lot of computing resources. Therefore, although WGS can provide comprehensive off-target analysis, it is difficult to achieve high-throughput detection of large-scale samples, which limits its use in certain applications. Technologies such as Digenome-seq, Circle-seq, and SITE-seq utilize the in vitro nuclease properties of Cas9 to cut genomic DNA in vitro. After the products are processed, they are sequenced or other means to screen for off-target sites. However, their operation methods are complex, there may be interference from background signals, the detection accuracy is low, and it may not be able to detect all types of off-target events, especially those off-target sites that do not rely on the NHEJ repair pathway, which will lead to incomplete detection.

[0007] Therefore, developing highly sensitive and specific off-target detection and analysis methods that can quickly and efficiently obtain accurate off-target effect site information is crucial to ensuring the safety and reliability of gene therapy. Summary of the Invention

[0008] In order to solve the problem of insufficient sensitivity and specificity of existing gene editing off-target detection methods, the present invention provides a method for assessing the off-target risk of gene editing based on targeted enrichment sequencing. This method has high sensitivity and high specificity, can improve the accuracy and sensitivity of gene editing off-target detection, and significantly reduce the overall sequencing cost, thereby achieving optimal resource allocation.

[0009] The present invention is achieved through the following technical solutions:

[0010] The present invention provides a method for assessing off-target risks of gene editing based on targeted enrichment sequencing, the method comprising:

[0011] The prediction software Cas-Offinder, CRISPOR, COSMID, CCTop, and GUIDES were used to screen sequences with the same length and PAM as the on-target sgRNA across the entire genome. The screened sequences were then aligned with the on-target sgRNA to screen out off-target sequences.

[0012] The off-target sequences screened by Cas-Offinder, CRISPOR, COSMID, CCTop, and GUIDES were summarized to obtain an off-target sequence table, and the off-target sequences that appeared in different off-target sequence tables and had a high potential off-target risk were selected as off-target sites;

[0013] A targeted enrichment sequencing method is used to perform high-depth sequencing of the off-target sites in the genomes of the edited samples to be tested and the unedited samples to be controlled, thereby obtaining the off-target site sequences of the edited samples to be tested and the off-target site sequences of the unedited samples to be controlled;

[0014] Compare the off-target site sequences of the edited sample to be tested and the off-target site sequences of the unedited control sample to the reference genome to obtain markdup (marked duplicate) bam files;

[0015] Perform mutation detection, mutation information annotation, and mutation site filtering on the markup bam file to obtain mutation sites that only exist in the edited sample to be tested;

[0016] Performing structural variation detection, structural variation information annotation, and structural variation site filtering on the markdup bam file to obtain structural variation sites that only exist in the edited sample to be tested;

[0017] Integrate data on the off-target site, the mutation site that only exists in the edited sample to be tested, and the structural variation site that only exists in the edited sample to be tested, and identify the authenticity of the off-target site through comparative analysis;

[0018] Evaluate the off-target risk and editing effect of sgRNA;

[0019] The off-target site information is visualized to display the distribution, frequency and relationship of the off-target sites with the target editing sites, and an off-target risk assessment report is automatically generated.

[0020] Furthermore, the prediction software Cas-Offinder, CRISPOR, COSMID, CCTop, and GUIDES are used to screen sequences with the same length and PAM as the on-target sgRNA across the entire genome, and the screened sequences are compared with the on-target sgRNA to screen out off-target sequences, specifically including:

[0021] Enter the PAM sequence, length, and species information of the on-target sgRNA;

[0022] The prediction software Cas-Offinder, CRISPOR, COSMID, CCTop, and GUIDES were used to screen sequences with the same length and PAM as the on-target sgRNA across the entire genome;

[0023] The screened sequences were aligned with the on-target sgRNA, and the threshold information mismatch (number of mismatches), insertion (insertion site), and deletion (deletion site) were calculated.

[0024] Input threshold information mismatch, insertion and deletion, filter out the off-target sequences that meet the conditions according to the input threshold information and output them to obtain the off-target sequences.

[0025] Furthermore, the off-target sequences screened by Cas-Offinder, CRISPOR, COSMID, CCTop and GUIDES are summarized to obtain an off-target sequence table, and the off-target sequences that appear in different off-target sequence tables and have a high potential off-target risk are selected as off-target sites, specifically including:

[0026] The off-target sequences screened by Cas-Offinder, CRISPOR, COSMID, CCTop and GUIDES were summarized to obtain an off-target sequence list;

[0027] Remove the repeated or obviously erroneous off-target sequences in each off-target sequence list;

[0028] The prediction software Cas-Offinder, CRISPOR, COSMID, CCTop and GUIDES were used to calculate and score the potential off-target risk of the off-target sequence based on sequence similarity, PAM sequence and target sequence position;

[0029] The off-target sequences that appear in different off-target sequence lists and have high potential off-target risks are selected as off-target sites.

[0030] Furthermore, the targeted enrichment sequencing method is used to perform high-depth sequencing on the off-target sites in the genomes of the edited sample to be tested and the unedited control sample to obtain the off-target site sequences of the edited sample to be tested and the off-target site sequences of the unedited control sample, specifically including:

[0031] Designing and synthesizing specific capture probes labeled with biotin or other tags for the off-target sites;

[0032] Extracting genomic DNA from the edited sample to be tested and genomic DNA from the unedited control sample, and evaluating the integrity and size distribution of the genomic DNA from the edited sample to be tested and the genomic DNA from the unedited control sample using agarose gel electrophoresis or a bioanalyzer;

[0033] Fragmenting the genomic DNA of the edited sample to be tested and the genomic DNA of the control unedited sample, respectively, and constructing a library suitable for sequencing;

[0034] Mixing the constructed library with the specific capture probe, performing a hybridization reaction in a hybridization buffer, capturing the target off-target site DNA fragments, and obtaining a hybridization mixture;

[0035] Separating the target off-target site DNA fragments from the hybridization mixture using a magnetic bead purification method;

[0036] Performing PCR amplification on the target off-target site DNA fragments to obtain an enriched library;

[0037] Using QPCR or a bioanalyzer to detect and quantify the enriched library to ensure the quality and concentration of the enriched library;

[0038] The enriched library is sequenced at high depth using a high-throughput sequencing platform to obtain the off-target site sequences of the edited sample to be tested and the off-target site sequences of the control unedited sample.

[0039] Furthermore, the off-target site sequences of the edited sample to be tested and the off-target site sequences of the control unedited sample are respectively aligned with the reference genome to obtain a markup bam file, specifically comprising:

[0040] performing quality control, low-quality read removal, and sequencing error correction on the off-target site sequences of the edited sample to be tested and the off-target site sequences of the control unedited sample, respectively, to obtain two sets of processed sequencing data;

[0041] The two sets of processed sequencing data were aligned with the hg38 reference genome downloaded from the Ensembl database using the BWA-MEM mode in the BWA software to obtain bam files;

[0042] The bam file was marked for duplication using Picard software to obtain a markup bam file after duplication was marked.

[0043] Furthermore, performing mutation detection, mutation information annotation, and mutation site filtering on the markup bam file to obtain mutation sites that exist only in the edited sample to be tested specifically includes:

[0044] Use GATK-Mutect2 mutation detection software to simultaneously perform mutation detection analysis on the data of the edited sample to be tested and the control unedited sample in the markdup bam file, identify the mutation sites on the genome of the edited sample to be tested and the control unedited sample, and obtain a vcf result file containing the frequency of the mutation site allele, the number of reads supporting the mutation, the depth (total coverage depth of the site), chromosome information, site location, and mutation direction through statistical analysis;

[0045] The transcript version number of the probe-capable region gene is pre-determined to form a standard file, and the mutation information of the vcf result file of the mutation detection is annotated using ANNOVAR annotation software to obtain multi-dimensional annotation results, which include mutation type, gene name, transcript ID, amino acid change, possible biological function and disease association;

[0046] First, the filtering module of the GATK-Mutect2 mutation detection software was used to filter based on base quality score and coverage to remove low-quality mutation sites and mutation sites with coverage below the threshold in the vcf result file of the mutation detection. Secondly, based on the allele frequency of the mutation site, the mutation sites below the frequency threshold were filtered. Then, based on the Fisher test, the significance of the mutation sites in the edited sample to be tested and the control unedited sample was evaluated, and the mutation sites that were significantly present in the edited sample to be tested were retained. Finally, the mutation sites were filtered based on information of the agreed thresholds of position and depth, and the high-quality mutation sites that existed only in the edited sample to be tested were retained.

[0047] Furthermore, the markup bam file is subjected to structural variation detection, structural variation information annotation, and structural variation site filtering to obtain structural variation sites that exist only in the edited sample to be tested, specifically including:

[0048] Perform structural variation detection and analysis on the markdup bam file using Manta software, including deletion, duplication, insertion, chromosome inversion, sequence translocation within the chromosome, and sequence translocation between chromosomes, to obtain structural variation raw detection information;

[0049] Annotate the structural variation raw detection information using AnnotSV software, including structural variation breakpoint site information, breakpoint position feature annotation, variation quality score, and variation type, to obtain annotated structural variation site information;

[0050] Filter out structural variation sites with low credibility from the annotated structural variation site information, and retain high-quality structural variation sites. Further filter out structural variation sites detected in the control unedited sample, and retain structural variation sites only present in the to-be-tested edited sample.

[0051] Further, the off-target sites, mutation sites only present in the to-be-tested edited sample, and structural variation sites only present in the to-be-tested edited sample are integrated, compared and analyzed to identify the authenticity of the off-target sites, specifically including:

[0052] Integrate the off-target sites, mutation sites only present in the to-be-tested edited sample, and structural variation sites only present in the to-be-tested edited sample, compare the off-target sites with the mutation sites only present in the to-be-tested edited sample and the structural variation sites only present in the to-be-tested edited sample, determine which off-target sites are verified in actual detection, thereby identifying the authenticity of the off-target sites and improving the sensitivity of off-target detection results, and the verified off-target sites are summarized to obtain an off-target site list.

[0053] Further, the evaluation of the off-target risk and editing effect of sgRNA specifically includes:

[0054] Using the constructed interpretation database resources (constructed based on GeneCards, Ensembl, Reactome, HGMD, etc., which contain descriptions of the effects of gene function, site biological function, etc. on detection sites), the role of the off-target site in gene expression regulation, signal transduction pathway and cell function is explored, and the biological function influence of the off-target site is comprehensively evaluated to evaluate the off-target risk and editing effect of sgRNA.

[0055] Further, the off-target site information is visualized to show the distribution, frequency and relationship with the target editing site of the off-target site, and an off-target risk evaluation report is automatically generated, specifically including:

[0056] Visualizing the off-target site information using a variety of result charts and / or result graphs to display the distribution, frequency, and relationship of the off-target sites with the target editing site, thereby obtaining a visualization chart;

[0057] Automatically generate an off-target risk assessment report, the off-target risk assessment report including:

[0058] a. a list of off-target sites;

[0059] b. the biological function impact of the off-target sites;

[0060] c. Assessment of off-target risk: Based on a and b, conduct a comprehensive assessment of the potential risks of off-target sites.

[0061] One or more technical solutions in the embodiments of the present invention have at least the following technical effects or advantages:

[0062] 1. The present invention provides a method for assessing the off-target risk of gene editing based on targeted enrichment sequencing. The method uses bioinformatics methods to predict off-target sites in advance, and uses targeted enrichment sequencing technology to perform high-depth sequencing only on the predicted off-target sites. This analysis strategy can not only improve the detection sensitivity of off-target sites under the conditions of the same data volume, greatly enhance the ability to capture low-frequency mutations, but also reduce sequencing costs. Through advanced bioinformatics analysis and precise molecular biology techniques, the present invention can accurately identify and quantify rare mutations generated during the gene editing process and chromosomal translocation events caused by double-stranded DNA breaks, significantly enhancing the ability to assess the off-target risk of gene editing results, and greatly improving the safety and reliability of gene editing operations.

[0063] 2. The present invention provides a method for assessing the off-target risk of gene editing based on targeted enrichment sequencing. The method first predicts off-target sites. Since a single software is difficult to comprehensively predict off-target sites, multiple prediction software are used for comprehensive analysis, and then TES technology is used to perform high-depth sequencing of the predicted off-target sites. TES is a high-throughput sequencing technology for specific genomic regions, which has the advantages of high sensitivity, high specificity and cost-effectiveness. When designing TES probes, only these predicted off-target sites are targeted and extended to the 10kb region around the target site. TES technology can efficiently enrich the predicted off-target site regions and can focus on these regions during high-throughput sequencing, thereby significantly improving the detection sensitivity of off-target sites without increasing the amount of data. Even low-frequency mutations can be accurately detected. At the same time, by reducing the sequencing of non-target areas, resources are saved and costs are reduced, making the assessment of off-target effects more economical and efficient.

[0064] 3. The present invention provides a method for assessing the off-target risk of gene editing based on targeted enrichment sequencing. The application of TES technology in this method can accurately detect mutations at off-target sites, even if they appear at a very low frequency in the genome. This highly sensitive verification method provides us with a deep insight into the off-target effects that may occur during the gene editing process, ensures the accuracy of the evaluation results, and greatly reduces the detection cost. After high-depth sequencing, the sequencing data of the edited sample to be tested and the control unedited sample are aligned to the reference genome, mutations are detected and annotated, and the mutation frequency of the off-target site and its biological impact are calculated. The mutation sites that specifically exist in the edited sample to be tested are obtained, and the sites present in the control unedited sample are eliminated, ensuring that the mutations we detect are truly derived from the gene editing process, rather than natural mutations or sequencing errors. At the same time, this detection method also ensures that rare mutations and long-segment chromosomal translocation events caused by double-strand breaks can be detected, thereby achieving a comprehensive off-target effect assessment to ensure the safety and specificity of gene editing. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0066] Figure 1 Flowchart for off-target site prediction.

[0067] Figure 2 Flowchart for off-target site validation.

[0068] Figure 3 Flowchart for off-target effect assessment.

[0069] Figure 4 Comparison chart of predicted off-target sites and off-target sites detected by WGS and TES.

[0070] Figure 5 Comparison of the detection frequencies of off-target sites between WGS and TES. DETAILED DESCRIPTION

[0071] The present invention will be described in detail below in conjunction with specific embodiments and examples, and the advantages and various effects of the present invention will be more clearly presented. It should be understood by those skilled in the art that these specific embodiments and examples are for illustrating the present invention, rather than for limiting the present invention.

[0072] Throughout this specification, unless otherwise specified, the terms used herein should be understood as having the same meaning as commonly used in the art. Therefore, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. In the event of any conflict, the present specification shall take precedence.

[0073] Unless otherwise specified, various raw materials, reagents, instruments and equipment used in the present invention can be purchased from the market or prepared by existing methods.

[0074] The following is a detailed description of a method for assessing off-target risks of gene editing based on targeted enrichment sequencing in conjunction with examples and experimental data.

[0075] The specific meanings of the terms used in the present invention are as follows:

[0076] WGS: Whole Genome Sequencing

[0077] TES: Target Enrichment Sequencing

[0078] AF: Allele Frequency, allele frequency

[0079] BND: Breakend, chromosome translocation

[0080] Example 1

[0081] This embodiment provides a method for assessing off-target risks of gene editing based on targeted enrichment sequencing, specifically comprising:

[0082] 1. Off-target site prediction

[0083] 1) Input the PAM sequence, length, and species information of the on-target sgRNA;

[0084] 2) Use prediction software Cas-Offinder, CRISPOR, COSMID, CCTop, and GUIDES to screen for sequences with the same PAM and length as the on-target sgRNA across the entire genome;

[0085] 3) Align the found sequence with the on-target sgRNA and calculate mismatch, insertion, and deletion information;

[0086] 4) Filter the qualified sequences (off-target) output according to the input threshold information (mismatch, insertion and deletion) (such as Figure 1 ).

[0087] 2. Off-target site verification

[0088] 1) Aggregate the off-target sequences provided by multiple prediction software and algorithms (Cas-Offinder, CRISPOR, COSMID, CCTop, GUIDES), obtain the off-target sequence list, remove the repeated or obviously incorrect prediction sites in each off-target sequence list, compare the prediction results from different sources, find the off-target sites with high consistency, that is, the sites that appear in the off-target sequence lists of multiple prediction software, and then select the potential high-risk off-target sites according to the consistency and predicted off-target probability;

[0089] The specific steps for predicting off-target probability are as follows:

[0090] The off-target prediction software Cas-Offinder, CRISPOR, COSMID, CCTop, GUIDES determines the potential off-target risk of the site mainly based on a series of calculations and scores related to sequence similarity, PAM sequence, target sequence position and other biophysical factors. The following are the key factors and methods for these tools to determine off-target risk:

[0091] (1). Sequence similarity

[0092] Base matching: Through off-target prediction software, compare the sequence similarity of the target sequence with other potential off-target sites in the genome. We mainly focus on the base matching of the 20-base guide sequence and the PAM sequence.

[0093] Mismatch site: Off-target detection software will record the number and position of mismatched bases at each potential off-target site compared to the target sequence. The fewer the number of mismatches and the closer the position to the PAM sequence, the higher the risk of off-target.

[0094] (2). PAM sequence

[0095] PAM sequence specificity: Different Cas nucleases have different requirements for PAM sequences. Off-target prediction tools will look for sequences similar to the target PAM sequence, and the closer to the ideal PAM sequence position, the higher the risk of off-target.

[0096] Mismatch near PAM: If the mismatch occurs near the PAM sequence (the first 3 base positions of PAM), the risk of off-target will be higher.

[0097] (3). Target sequence position

[0098] Background sequence of target position: in the region of highly repetitive sequence and high GC content, the risk of off-target is higher.

[0099] Genomic structure: the sequence near the splicing site may have multiple splicing variants or incomplete sequence characteristics, which increases the chance of non-specific matching with the target sequence. Especially near the splicing site, there is higher sequence variability, which increases the risk of off-target. Regulatory elements such as enhancers, promoters and transcription factor binding sites usually have important functional roles, and their sequences may be more conserved or have special sequence characteristics. The sequence characteristics of these regions have higher similarity with the target sequence, which also increases the risk of off-target.

[0100] By combining the above factors, off-target prediction software can evaluate potential off-target sites in the genome, and by appearing in different off-target sequence lists, determine the sites with higher off-target risk.

[0101] 2) Design specific capture probes for the screened off-target sites, synthesize TES probes, and label with biotin or other labels.

[0102] 3) Extract DNA from the sample to be tested and the control unedited sample, and use gel electrophoresis (agarose gel) or bioanalyzer (Agilent 2100 Bioanalyzer) to evaluate the integrity and size distribution of the DNA;

[0103] 4) Fragment the DNA sample and perform end repair, adapter ligation and other library construction steps to construct a library suitable for sequencing;

[0104] 5) Mix the constructed library with biotin-labeled specific probes and perform hybridization reaction in hybridization buffer to specifically bind the target region to the probe and capture the target off-target site;

[0105] 6) Purify the captured DNA fragments from the hybridization mixture by magnetic bead purification method and remove the unbound non-target sequences by elution;

[0106] 7) Perform PCR amplification on the captured DNA fragments to increase the concentration of the target region;

[0107] 8) Use QPCR or bioanalyzer to detect and quantify the enriched library to ensure the quality and concentration of the library;

[0108] 9) Use high-throughput sequencing platform to perform high-depth sequencing (e.g. Figure 2 ) on the enriched off-target sites.

[0109] 3. Off-target effect evaluation

[0110] 1) Use high-throughput sequencing technology to perform targeted sequencing of off-target regions of the sequencing library to obtain raw genome sequencing data (fastq data);

[0111] 2) performing quality control, removal of low-quality reads, and sequencing error correction on the raw data obtained in 1) to obtain processed sequencing data (clean data);

[0112] 3) Use BWA software in BWA-MEM mode to align the processed sequencing data obtained in 2) with the reference genome (hg38 downloaded from the Ensembl database) to obtain a bam file;

[0113] 4) Mark duplicates in the bam file. Use Picard software to analyze the position and sequence of each read, identify duplicate reads with the same sequence at the same position, and mark them as MarkDuplicatesreads. After obtaining the marked duplicates, markdup bam;

[0114] 5) GATK-Mutect2 mutation detection software was used to simultaneously perform mutation detection analysis on the data of the test edited samples (i.e., samples that have undergone gene editing) and the control unedited samples, identifying mutation sites on the genome, including insertions (ins), deletions (del), and substitutions (snv). The software assessed the credibility of the mutation by statistically analyzing the sequencing read information of each potential mutation position, including base quality, alignment quality, and read coverage. Statistical analysis was performed to obtain a vcf result file containing information such as the frequency of the mutant allele, the number of reads supporting the mutation, the total coverage depth of the site, chromosome information, site location, and mutation direction. The software was debugged, and ultimately low-frequency mutations were accurately detected, with a mutation frequency detection limit of 0.1%;

[0115] 6) Predetermine the transcript version number of the probe-capable region gene to form a standard file. Use ANNOVAR annotation software and set the corresponding parameters to perform mutation annotation on the mutation detection vcf file results to obtain multi-dimensional annotation results, including mutation type (such as missense, nonsense, splice site variation, etc.), gene name, transcript ID, amino acid changes, possible biological functions and disease associations;

[0116] 7) By filtering the mutation detection sites, first, using the filtering module of the mutation detection software GATK-Mutect2, based on base quality score and coverage, remove low-quality mutation sites and mutation sites with coverage below the threshold in the mutation detection vcf result file. Based on the mutation site frequency, sites below the frequency threshold are filtered. Based on the Fisher test, the significance of the sites in the sample and the control unedited sample is evaluated, and off-target sites that are significantly present in the sample are retained. Secondly, the mutation sites are filtered based on information such as position and depth, and finally, high-quality mutation sites that are only present in the edited sample to be tested are retained;

[0117] 8) For the original call of structural variations, Manta software is used to combine paired-end, split read and coverage depth information to perform a comprehensive analysis of structural variations, including deletions, duplications, insertions, inversions, and translocations within or between chromosomes.

[0118] 9) Annotate the structural variation results using AnnotSV software to obtain information such as structural variation breakpoint site information, variation quality score, variation type, and breakpoint location feature annotations;

[0119] 10) Structural variant site filtering is performed based on the alignment coverage and alignment quality of the breakpoints, variant characteristics, variant frequency, and the processing of repetitive regions and low-complexity regions. Low-confidence variants are filtered out, while high-quality off-target structural variant sites are retained. Variant call sites in the control unedited sample are filtered out, and only structural variant sites present in the edited sample to be tested are retained;

[0120] 11) Integrate and compare the off-target site prediction information with the mutation site and structural variation site data that are only present in the edited sample to be tested in the above steps to determine which predicted off-target sites are verified in the actual detection. The authenticity of the off-target sites is identified by comparing the dual results of the algorithm and the experiment. The verified off-target sites are summarized to obtain a list of off-target sites;

[0121] 12) First, we thoroughly documented and tracked the origins of sgRNAs to ensure that each sgRNA design had a clear provenance and detailed design strategy, including the algorithms and tools used, as well as the specific sequence information of the sgRNA. Next, we focused on the phenomenon of mismatch_target_sgRNA, which refers to the potential incomplete match between the sgRNA and the target sequence. Finally, we utilized our constructed interpretation database resources (including gene function, site biological function, and other impacts) to explore the role of off-target sites in gene expression regulation, signal transduction pathways, and cell function. We also comprehensively evaluated the biological functional impact of off-target sites to assess off-target risk and editing efficacy.

[0122] 13) The off-target site information was visualized and analyzed using R language, using a variety of result charts and graphs to show the distribution, frequency, and relationship of off-target sites to the target editing site;

[0123] 14) The final off-target analysis results and visualization results are automatically generated into a report through a Python language program. After the off-target analysis is completed, the analysis results and visualization charts are automatically sorted to generate a detailed report. This report not only includes the off-target site list obtained in step 11) and the biological function impact of the off-target sites in step 12), but also includes a comprehensive assessment of the potential risks of the off-target sites (such as Figure 3 ).

[0124] Example 2

[0125] The main purpose of this example is to compare the quality control results, detection results, and low-frequency detection performance of WGS sequencing and TES sequencing of the targeted region of predicted off-target sites, and to compare the sensitivity of the two sequencing methods for off-target detection analysis.

[0126] First, we compared the differences in quality control results between the two sequencing technologies, including but not limited to data quality, sequencing depth, and coverage uniformity. The main implementation steps are as follows:

[0127] 1) Sample selection:

[0128] Three parallel experiments were performed using one test edited sample (UCAR-T cell) and one unedited control sample (control cell) (the test edited sample was a sample that had undergone CRISPR-Cas9 gene editing, and the TRAC gene sgRNA was AGAGTCTCTCAGCTGGTACA). DNA was extracted from each sample, and the samples were processed and grouped identically.

[0129] 2) Experimental grouping: Each sample was divided into two groups.

[0130] Group 1: Whole genome sequencing (WGS) was performed on the edited sample to be tested and the unedited control sample, named WGS-Sample.

[0131] Group 2: The edited samples to be tested and the unedited control samples were subjected to targeted enrichment sequencing (TES). The predicted off-target sites were subjected to targeted enrichment sequencing using the method in Example 1. This group was named TES-Sample.

[0132] 3) Sequencing technology and parallel experiments:

[0133] Each group conducted three parallel experiments, named WGS-Sample1, WGS-Sample2, WGS-Sample3, TES-Sample1, TES-Sample2 and TES-Sample3, to ensure the reliability and repeatability of the data.

[0134] All sequencing experiments were performed using high-throughput sequencing technology (Illumina) to ensure high data quality and comparability.

[0135] The off-target site information predicted by Group 2 using the method of Example 1 is shown in Table 1, and the sequence information of the designed TES probe is shown in Table 2.

[0136] Table 1: List of predicted off-target sites

[0137]

[0138]

[0139] Table 2: TES probe sequence information

[0140]

[0141]

[0142]

[0143]

[0144] 4) Data analysis and comparison:

[0145] Whole Genome Sequencing (WGS) Panel:

[0146] The whole genome sequencing data obtained from each experiment were quality controlled by fastqQC. The quality controlled data were aligned with the reference genome (hg38 downloaded from the Ensembl database) using BWA. Quality control information such as the average sequencing depth of each sample was calculated using bamdst software. All variants, including single nucleotide polymorphisms (SNPs) and insertion / deletion variants (InDels) in the targeted and off-target regions, were identified and annotated using GATK-Mutect2 mutation detection software and ANNOVAR mutation annotation software. Targeted region sequencing (TES) group:

[0147] The sequencing data of the targeted region obtained in each experiment were quality controlled by fastqQC. The quality-controlled data were aligned with the reference genome by BWA, and the average sequencing depth of the targeted region and other quality control information were calculated by bamdst software. The detected mutation information was analyzed and annotated by GATK-Mutect2 mutation detection software and ANNOVAR mutation annotation software.

[0148] 5) Comparison and statistical analysis of quality control results:

[0149] Comparing the quality control results of WGS sequencing and targeted region sequencing, as shown in Table 3, reveals that both methods achieved very high mapping rates, at 99.80% for WGS and 99.77% for TES, respectively. This demonstrates that both sequencing methods achieve excellent data mapping. Although the data volume of WGS sequencing (30GB) far exceeds that of TES sequencing (2GB), the average sequencing depth of TES sequencing in the target region (approximately 12,000x) is significantly higher than that of WGS sequencing (approximately 100x). Furthermore, TES sequencing significantly outperforms WGS sequencing in terms of high-depth coverage at thresholds >= 300x, >= 400x, and >= 500x. These quality control results demonstrate TES sequencing's superiority in providing both deep and high coverage. For accurate detection of off-target sites, TES sequencing is a more suitable option.

[0150] Table 3: Comparison of quality control results of WGS sequencing and TES targeted region sequencing

[0151]

[0152]

[0153] In Table 3:

[0154] QC: Quality Control

[0155] Sample: Sample

[0156] [Total]Raw Reads (All reads): Total number of raw sequencing reads

[0157] [Total]Raw Data(Mb): The total amount of raw data (in Mb)

[0158] [Total]Mapped Reads: The total number of reads mapped to the reference genome

[0159] [Total]Fraction of Mapped Reads: The proportion of mapped reads to the total reads.

[0160] [Target]Average depth: Average coverage depth of the target area

[0161] [Target]Coverage(>0x): The proportion of the target area that has at least 0x coverage

[0162] [Target]Coverage (>=4x): The target area has at least 4x coverage.

[0163] [Target]Coverage (>=10x): The proportion of the target area that has at least 10x coverage

[0164] [Target]Coverage (>=30x): The target area has at least 30x coverage

[0165] [Target]Coverage (>=100x): The proportion of the target area that has at least 100x coverage

[0166] [Target]Coverage (>=200x): The target area has at least 200x coverage

[0167] [Target]Coverage (>=300x): The target area has at least 300x coverage

[0168] [Target]Coverage (>=400x): The target area has at least 400x coverage

[0169] [Target]Coverage (>=500x): The target area has at least 500x coverage

[0170] [Target]Coverage (>=600x): The proportion of the target area with at least 600x coverage

[0171] [Target]Coverage (>=700x): The target area has at least 700x coverage

[0172] [Target]Coverage (>=800x): The target area has at least 800x coverage.

[0173] 6) Comparison and statistical analysis of test results:

[0174] Comparing the off-target site detection results of predicted off-target sites in WGS sequencing and targeted region sequencing, the results are shown in Table 4. We can conclude that:

[0175] The results of TES and WGS for predicting off-target sites show that TES (targeted sequencing) detected the predicted off-target site mutations in all samples, while WGS (whole genome sequencing) detected only one mutation site among all the listed mutation sites, and the other mutation sites were not detected (marked as "None"). The number of detected sites accounted for 4.8% of the predicted number of sites, while TES could detect all 21 detection sites, and the number of detected sites accounted for 100% of the predicted number of sites. The comparison of their detection rates can be seen. Figure 4 , which indicates that TES has high resolution and sensitivity in detecting specific mutations.

[0176] The types of variations detected by TES include insertions (ins), deletions (del), and single nucleotide mutations (snv). This shows that TES sequencing can identify multiple types of genomic variations. In addition, the variation frequencies in TES samples showed a certain consistency among the three samples, indicating that TES sequencing can provide reproducible results between repeated samples.

[0177] Due to its high depth of coverage, TES samples detected mutation frequencies ranging from 0.21% to 5.13%. However, due to the low sequencing depth of WGS, WGS samples can only detect more than 5% of mutation sites. The frequency comparison is obvious. Figure 5 ,The diversity results from the mutation frequencies indicate that TES can detect mutations ranging from low ,frequency to high frequency.

[0178] For example, multiple variants were detected in the ARHGEF17 and DCDC1 genes in TES samples, but not in WGS samples. This indicates that the regions specifically designed for TES can be detected unbiasedly and accurately, while although WGS provides genome-wide coverage, the vast majority of off-target sites are not detected.

[0179] In summary, TES (targeted region sequencing) has significant advantages in detecting variants in non-target regions (off-target), mainly reflected in its excellent sensitivity and accuracy. In contrast, although WGS (whole genome sequencing) provides comprehensive genome coverage, in this experiment, the predicted variant detection rate of off-target sites is relatively low, only 4.8%, which is far lower than the 100% detection rate of TES sequencing. In addition, the detection limit of WGS is a variant frequency greater than 5%, while TES sequencing can detect low-frequency variants as low as 0.2%. The high-depth sequencing technology of TES can improve the detection sensitivity while also improving the detection sensitivity of rare mutations. Therefore, for off-target variant detection that requires high sensitivity and high accuracy, TES sequencing is a more suitable choice.

[0180] Table 4: Comparison of off-target detection results between WGS sequencing and TES targeted region sequencing

[0181]

[0182]

[0183]

[0184] In Table 4:

[0185] Variant: mutation form, gene: base mutation or mutation position: reference sequence / replacement sequence

[0186] VarType: mutation type

[0187] None: Not detected.

[0188] Example 3

[0189] The main purpose of this example is to verify the ability of targeted enrichment sequencing (TES) to successfully identify chromosomal translocations under a large-area coverage strategy, and whether these translocation events can be detected if specific regions are not covered by TES. The main implementation steps are as follows:

[0190] 1) Enrichment probe design:

[0191] Design specific probes or primers to capture and enrich specific genomic regions of chromosomal translocations, ensuring that the probes cover sufficient areas to include possible translocation breakpoints.

[0192] 2) Sample selection:

[0193] DNA was extracted from the sample to be tested and the control unedited sample, and appropriate quality control was performed to ensure that the DNA was suitable for subsequent enrichment and sequencing. The TRAC gene sgRNA was AGAGTCTCTCAGCTGGTACA.

[0194] 3) Sample grouping:

[0195] The sample to be tested was divided into two groups, and each group was analyzed in triplicate.

[0196] Group 1: The sample to be tested and the control unedited sample were subjected to targeted enrichment sequencing, and the detection probes in Table 2 were added (the conventional probes refer to detection probes that do not cover specific genes of chromosomal translocation--i.e., all probes in Table 2 except probes chr19:50217800-50219120, chr14:22546384-22548207, and chr17:82315941-82317441). The sample was named Sample-1.

[0197] Group 2: The sample to be tested and the control unedited sample were subjected to targeted enrichment sequencing according to the steps in Example 1, and the detection probes in Table 2 that cover specific genes of chromosomal translocation were added. The sample was named Sample-2.

[0198] 4) Target region enrichment:

[0199] Hybridization capture was performed using the designed probes to enrich the specific genomic regions of chromosomal translocation, ensuring that the target region was fully covered.

[0200] 5) Sequencing:

[0201] High-throughput sequencing was performed on the enriched target region to generate a large amount of sequencing data.

[0202] 6) Data analysis:

[0203] Data alignment and structural variation analysis were performed according to the steps 3.1-3.4, 3.8, and 3.9 in Example 1 to identify the translocation events captured by the specific probes in Sample-2, and the results were compared with those of the conventional probes in Sample-1.

[0204] 7) Result comparison:

[0205] By comparing the test results of the two groups of samples in Table 5, it can be seen that no structural variation was detected in the three parallel samples of Sample 1, Sample 1-1, Sample 1-2, and Sample 1-3, while all structural variations were detected in Sample 2-1, Sample 2-2, and Sample 2-3. This indicates that when the targeted enrichment sequencing (TES) method is used and the region where the translocation occurs is successfully covered (Sample 2), chromosomal translocation and large-fragment deletion can be accurately detected.

[0206] Table 5. TES detection Translocation comparison results

[0207]

[0208]

[0209] In Table 5:

[0210] Start_feature:Gene name

[0211] End_feature: gene name

[0212] chrstart: starting breakpoint position

[0213] chr2End: end breakpoint position

[0214] Freq: mutation frequency

[0215] Qual: quality score of the variant

[0216] Type: mutation type, such as DEL (deletion), BND (chromosomal translocation)

[0217] CT: breakpoint link mode

[0218] SV_len: length of the variant or number of bases affected

[0219] Finally, it should be noted that the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0220] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0221] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A method for assessing off-target risk of gene editing based on targeted enrichment sequencing, characterized in that: The method comprises: The prediction software Cas-Offinder, CRISPOR, COSMID, CCTop, and GUIDES were used to screen sequences with the same length and PAM as the on-targets gRNA across the entire genome. The screened sequences were compared with the on-targets gRNA to screen out off-target sequences. The off-target sequences screened by Cas-Offinder, CRISPOR, COSMID, CCTop, and GUIDES were summarized to obtain an off-target sequence table, and the off-target sequences that appeared in different off-target sequence tables and had a high potential off-target risk were selected as off-target sites; A targeted enrichment sequencing method is used to perform high-depth sequencing of the off-target sites in the genomes of the edited samples to be tested and the unedited samples to be controlled, thereby obtaining the off-target site sequences of the edited samples to be tested and the off-target site sequences of the unedited samples to be controlled; The off-target site sequences of the edited sample to be tested and the off-target site sequences of the control unedited sample are respectively aligned with the reference genome to obtain a markup bam file; Perform mutation detection, mutation information annotation, and mutation site filtering on the markup bam file to obtain mutation sites that only exist in the edited sample to be tested; Performing structural variation detection, structural variation information annotation, and structural variation site filtering on the markdup bam file to obtain structural variation sites that only exist in the edited sample to be tested; Integrate data on the off-target site, the mutation site that only exists in the edited sample to be tested, and the structural variation site that only exists in the edited sample to be tested, and identify the authenticity of the off-target site through comparative analysis; Assess the off-target risk and editing effect of sgRNA; The off-target site information is visualized to display the distribution, frequency and relationship of the off-target sites with the target editing sites, and an off-target risk assessment report is automatically generated.

2. The method for assessing off-target risk of gene editing based on targeted enrichment sequencing according to claim 1, characterized in that: The prediction software Cas-Offinder, CRISPOR, COSMID, CCTop, and GUIDES are used to screen sequences with the same length and PAM as the on-target sgRNA across the entire genome, and the screened sequences are compared with the on-target sgRNA to screen out off-target sequences, specifically including: Enter the PAM sequence, length, and species information of the on-target sgRNA; The prediction software Cas-Offinder, CRISPOR, COSMID, CCTop, and GUIDES were used to screen sequences with the same length and PAM as the on-target sgRNA across the entire genome; Align the screened sequences with the on-target sgRNA and calculate the threshold information mismatch, insertion, and deletion; Input threshold information mismatch, insertion, and deletion, filter out the off-target sequences that meet the conditions according to the input threshold information, and output them to obtain the off-target sequences.

3. The method for assessing off-target risk of gene editing based on targeted enrichment sequencing according to claim 1, characterized in that: The off-target sequences screened by Cas-Offinder, CRISPOR, COSMID, CCTop and GUIDES are summarized to obtain an off-target sequence table, and the off-target sequences that appear in different off-target sequence tables and have a high potential off-target risk are selected as off-target sites, specifically including: The off-target sequences screened by Cas-Offinder, CRISPOR, COSMID, CCTop and GUIDES were summarized to obtain an off-target sequence list; Remove the repeated or obviously erroneous off-target sequences in each off-target sequence list; The prediction software Cas-Offinder, CRISPOR, COSMID, CCTop and GUIDES were used to calculate and score the potential off-target risk of the off-target sequence based on sequence similarity, PAM sequence and target sequence position; The off-target sequences that appear in different off-target sequence lists and have high potential off-target risks are selected as off-target sites.

4. The method for assessing off-target risk of gene editing based on targeted enrichment sequencing according to claim 1, characterized in that: The targeted enrichment sequencing method is used to perform high-depth sequencing of the off-target sites in the genomes of the edited sample to be tested and the unedited sample to be controlled, to obtain the off-target site sequences of the edited sample to be tested and the off-target site sequences of the unedited sample to be controlled, specifically comprising: Designing and synthesizing specific capture probes labeled with biotin or other tags for the off-target sites; Extracting genomic DNA from the edited sample to be tested and genomic DNA from the unedited control sample, and evaluating the integrity and size distribution of the genomic DNA from the edited sample to be tested and the genomic DNA from the unedited control sample using agarose gel electrophoresis or a bioanalyzer; Fragmenting the genomic DNA of the edited sample to be tested and the genomic DNA of the control unedited sample, respectively, and constructing a library suitable for sequencing; Mixing the constructed library with the specific capture probe, performing a hybridization reaction in a hybridization buffer, capturing the target off-target site DNA fragments, and obtaining a hybridization mixture; Separating the target off-target site DNA fragments from the hybridization mixture using a magnetic bead purification method; Performing PCR amplification on the target off-target site DNA fragments to obtain an enriched library; Using QPCR or a bioanalyzer to detect and quantify the enriched library to ensure the quality and concentration of the enriched library; The enriched library is sequenced at high depth using a high-throughput sequencing platform to obtain the off-target site sequences of the edited sample to be tested and the off-target site sequences of the control unedited sample.

5. The method for assessing off-target risk of gene editing based on targeted enrichment sequencing according to claim 1, characterized in that: The step of respectively aligning the off-target site sequence of the edited sample to be tested and the off-target site sequence of the unedited control sample with the reference genome to obtain a markup bam file specifically includes: performing quality control, low-quality read removal, and sequencing error correction on the off-target site sequences of the edited sample to be tested and the off-target site sequences of the control unedited sample, respectively, to obtain two sets of processed sequencing data; The two sets of processed sequencing data were aligned with the hg38 reference genome downloaded from the Ensembl database using the BWA-MEM mode in the BWA software to obtain bam files; The bam file was marked for duplication using Picard software to obtain a markup bam file after duplication was marked.

6. The method for assessing off-target risk of gene editing based on targeted enrichment sequencing according to claim 1, characterized in that: The markup bam file is subjected to mutation detection, mutation information annotation, and mutation site filtering to obtain mutation sites that exist only in the edited sample to be tested, specifically comprising: Use GATK-Mutect2 mutation detection software to simultaneously perform mutation detection analysis on the data of the edited sample to be tested and the control unedited sample in the markdup bam file, identify the mutation sites on the genome of the edited sample to be tested and the control unedited sample, and obtain a vcf result file containing the frequency of the mutation site allele, the number of reads supporting the mutation, the depth, chromosome information, the site location, and the mutation direction through statistical analysis; The transcript version number of the probe-capable region gene is pre-determined to form a standard file, and the mutation information of the vcf result file of the mutation detection is annotated using ANNOVAR annotation software to obtain multi-dimensional annotation results, which include mutation type, gene name, transcript ID, amino acid change, possible biological function and disease association; First, the filtering module of the GATK-Mutect2 mutation detection software was used to filter based on base quality score and coverage to remove low-quality mutation sites and mutation sites with coverage below the threshold in the vcf result file of the mutation detection. Secondly, based on the allele frequency of the mutation site, the mutation sites below the frequency threshold were filtered. Then, based on the Fisher test, the significance of the mutation sites in the edited sample to be tested and the control unedited sample was evaluated, and the mutation sites that were significantly present in the edited sample to be tested were retained. Finally, the mutation sites were filtered based on information of the agreed thresholds of position and depth, and the high-quality mutation sites that existed only in the edited sample to be tested were retained.

7. The method for assessing off-target risk of gene editing based on targeted enrichment sequencing according to claim 1, characterized in that: The step of performing structural variation detection, structural variation information annotation, and structural variation site filtering on the markup bam file to obtain structural variation sites that exist only in the edited sample to be tested specifically includes: Manta software was used to detect and analyze structural variations in the markup bam file, including deletions, duplications, insertions, chromosomal inversions, intrachromosomal sequence translocations, and interchromosomal sequence translocations, to obtain the original detection information of structural variations; AnnotSV software was used to annotate the original structural variation detection information, including structural variation breakpoint site information, breakpoint position feature annotation, variation quality score and variation type, to obtain annotated structural variation site information; The structural variation sites with low credibility in the annotated structural variation site information are filtered out, and high-quality structural variation sites are retained. Then, the structural variation sites detected in the control unedited sample are filtered out, and the structural variation sites that only exist in the edited sample to be tested are retained.

8. The method for assessing off-target risk of gene editing based on targeted enrichment sequencing according to claim 1, characterized in that: The data integration of the off-target site, the mutation site that exists only in the edited sample to be tested, and the structural variation site that exists only in the edited sample to be tested, and comparative analysis to identify the authenticity of the off-target site specifically includes: The off-target sites, the mutation sites that only exist in the edited sample to be tested, and the structural variation sites that only exist in the edited sample to be tested are integrated, and the off-target sites are compared with the mutation sites that only exist in the edited sample to be tested and the structural variation sites that only exist in the edited sample to be tested to determine which of the off-target sites have been verified in actual detection, thereby identifying the authenticity of the off-target sites and improving the sensitivity of the off-target detection results. The verified off-target sites are summarized to obtain a list of off-target sites.

9. The method for assessing off-target risk of gene editing based on targeted enrichment sequencing according to claim 8, characterized in that: The evaluation of the off-target risk and editing effect of sgRNA specifically includes: By utilizing the constructed interpretation database resources, the role of the off-target sites in gene expression regulation, signal transduction pathways and cell functions is explored, and the biological functional impact of the off-target sites is comprehensively evaluated to assess the off-target risk and editing effect of sgRNA.

10. The method for assessing off-target risk of gene editing based on targeted enrichment sequencing according to claim 9, characterized in that: The visualization of the off-target site information, displaying the distribution, frequency and relationship of the off-target sites with the target editing sites, and automatically generating an off-target risk assessment report, specifically includes: Visualizing the off-target site information using a variety of result charts and / or result graphs to display the distribution, frequency, and relationship of the off-target sites with the target editing site, thereby obtaining a visualization chart; Automatically generate an off-target risk assessment report, the off-target risk assessment report including: a. a list of off-target sites; b. the biological function impact of the off-target sites; c. Assessment of off-target risk: Based on a and b, conduct a comprehensive assessment of the potential risks of off-target sites.

Citation Information

Patent Citations

  • CRISPR assisted DNA target enrichment method and application thereof

    CN109837273A

  • Use of off-target sequences for DNA analysis

    CN110475874A