An off-target detection method for adenine base editor based on inosine enrichment and its application

By directly capturing the enzymatic products of ABE using the Ino-seq method, the insufficient sensitivity and specificity of existing ABE off-target detection methods are resolved, enabling high-sensitivity and high-specificity whole-genome detection in live cells, and supporting the clinical safety assessment and optimization of ABE technology.

CN122081448APending Publication Date: 2026-05-26ZHUHAI SHU TONG MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610058365.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-16
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing adenine base editor (ABE) off-target detection methods suffer from low sensitivity and poor specificity, making it impossible to directly detect sgRNA-dependent and non-dependent off-target events in physiologically relevant cellular environments. Furthermore, they fail to reflect the influence of native chromatin states, making it difficult to meet the clinical safety assessment and optimization needs of ABE technology.

Method used

An inosine-enriched adenine base editor-based off-target detection method (Ino-seq) was adopted. By introducing the adenine base editor system into cells, genomic DNA was extracted, inosine-containing DNA fragments were enriched, a sequencing library was constructed, and high-throughput sequencing was performed. The unique position signal characteristics and dual enrichment strategy were used to identify off-target sites. Combined with multi-time point analysis, high sensitivity and high specificity of whole-genome detection were achieved.

Benefits of technology

It enables unbiased, high-sensitivity (F1 score > 0.89) and high-specificity (validation accuracy > 95%) detection of ABE-induced genomic inosine in living cells, and can simultaneously identify sgRNA-dependent and non-dependent off-target sites, providing a reliable tool for the clinical safety assessment and optimization of ABE technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122081448A_ABST
    Figure CN122081448A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of gene editing technology and discloses an inosine-enriched adenine base editor off-target detection method (Ino-seq) and its applications. The method of this invention utilizes endonuclease V to specifically cleave inosine-containing DNA, combined with the stabilization of a thermostable single-stranded DNA-binding protein and the high affinity enrichment of streptavidin-biotin, to achieve direct capture of ABE enzymatic products. This method can simultaneously detect sgRNA-dependent and independent off-target effects, with systematic comparisons showing F1 scores of 0.892–0.931. Targeted deep sequencing validation showed an overall accuracy of 95.3%. Ino-seq can also detect endogenous genomic inosine, observing a variation range of >1,000-fold in 10 cell lines. This invention provides a crucial safety assessment tool for the clinical translation of ABE technology, supporting comprehensive off-target analysis in preclinical and clinical studies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene editing technology, specifically to an inosine-enriched adenine base editor (ABE) off-target detection method (Ino-seq) and its applications. Specifically, it relates to applications in ABE safety assessment, base editor optimization, and endogenous genomic inosine detection. Background Technology

[0002] Adenine base editor (ABE) is an innovative genome editing technology composed of a catalytically inactivated variant of the CRISPR-Cas9 system (dCas9 or nCas9) fused with adenine deaminase (TadA). ABE converts adenine (A) at the target site to inosine (I) via adenine deamination. Inosine is then recognized by the cell as guanine (G) during DNA replication or repair, thus achieving precise A·T to G·C base pair conversion without the need for double-strand DNA breaks (DSBs) or donor templates.

[0003] Base editing (ABE) technology has shown great potential in the treatment of genetic diseases. According to Gaudelli et al. (Nature, 2017), approximately 48% of known pathogenic point mutations can be repaired using base editing technology. ABE has entered clinical trials for the treatment of genetic diseases such as CPS1 deficiency. However, genome-wide detection of off-target effects is a crucial prerequisite for the clinical translation of ABE technology. Off-target editing may introduce mutations at unexpected genomic locations, leading to potential safety risks such as tumorigenesis and immunogenicity.

[0004] Existing ABE off-target detection methods each have their inherent limitations: (1) GUIDE-seq (Tsai et al., Nat Biotechnol. 2015; 33(2):187-197): This method relies on DSB-mediated double-stranded oligonucleotide (dsODN) integration for labeling. Since ABE works through a base deamination mechanism without generating DSB, GUIDE-seq has an inherent incompatibility with ABE detection and has extremely low detection sensitivity (F1 score: 0.043-0.213).

[0005] (2) Tracking-seq (Yin et al., Nat Biotechnol. 2022; 40(10):1455-1464): detects off-target events by capturing DNA repair intermediates, but has moderate sensitivity (F1 score: 0.032), may miss fast repair editing events, and cannot distinguish between different types of off-target events.

[0006] (3) Selict-seq (Li et al., Nucleic Acids Res. 2023; 51(8):e47): Although it can enrich DNA containing inosine, its sensitivity and specificity fluctuate significantly. For example, at the VEGFA site3, Selict-seq detected only 36 off-target sites, of which only 4 overlapped with GUIDE-seq, while Ino-seq detected 66 sites, 25 of which overlapped with GUIDE-seq, showing a detection efficiency difference of >6-fold.

[0007] (4) Whole genome sequencing (WGS): requires ultra-high sequencing depth (usually >100×) to distinguish real editing events from sequencing errors, and has an extremely low signal-to-noise ratio (F1 score: 0.001-0.006), is costly and difficult to detect low-frequency events.

[0008] (5) CHANGE-seq-BE (Lazzarotto et al., Nat Biotechnol. 2020; 38(11):1327-1337): It has high sensitivity to sgRNA-dependent off-targets in cell-free systems, but cannot detect sgRNA-independent events (which can account for 12-100% of total off-targets) and may not reflect the effect of native chromatin accessibility on off-target activity (F1 score: 0.038).

[0009] Table 1 Performance comparison of existing ABE off-target detection methods These limitations highlight the urgent need for novel methods that can directly, with high sensitivity and specificity, detect ABE-induced inosine in physiologically relevant cellular environments. An ideal detection method should be able to: 1) directly capture the enzymatic products of ABE; 2) simultaneously detect sgRNA-dependent and non-dependent off-target effects; 3) reflect the influence of native chromatin state in living cells; 4) possess high sensitivity and specificity; and 5) enable dynamic analysis at multiple time points. Summary of the Invention

[0010] The purpose of this invention is to provide an inosine-enriched adenine base editor-based off-target detection method (Ino-seq) that overcomes the shortcomings of existing technologies. This method enables unbiased, high-sensitivity (F1 score > 0.89) and high-specificity (validation accuracy > 95%) whole-genome detection of ABE-induced genomic inosine in live cells, including sgRNA-dependent and non-dependent off-target sites. This provides a reliable tool for the clinical safety assessment of ABE technology and the optimization of base editors.

[0011] To achieve the above objectives, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides an off-target detection method for adenine base editor based on inosine enrichment, comprising the following steps: S1: The adenine base editor system was introduced into cells, and genomic DNA was extracted after culture and inosine site-specific cleavage was performed. S2: Enrichment of DNA fragments containing inosine; S3: Construct the gene library enriched with inosine-containing DNA fragments and perform high-throughput sequencing; S4: Identify off-target editing sites in the high-throughput sequencing based on location features and count the reads with the location features; the location features are reads showing A→G mutations.

[0012] As a preferred embodiment of the inosine-enriched adenine base editor-based off-target detection method of the present invention, the inosine-enriched adenine base editor-based off-target detection method includes the following steps: (1) Sample preparation: The adenine base editor system was introduced into the cells, and genomic DNA was extracted after culturing; (2) Sequencing library construction: a) Fragment the genomic DNA, perform end repair, and add an A tail; b) Connecting a double-stranded DNA linear adapter; preferably, the double-stranded DNA linear adapter is modified with phosphate thioester and biotin; c) Perform DNA damage repair treatment; d) Treatment with exonuclease; preferably, treatment with exonuclease is repeated at least three times; to remove DNA fragments without ligated adapters; e) Perform specific cleavage; preferably, the cleavage is performed using endonuclease V, and the cleavage site is the 3' end of the inosine site; f) Denature the cut DNA and stabilize the single-stranded DNA using a single-stranded DNA binding protein; preferably, the single-stranded DNA binding protein is a heat-resistant single-stranded DNA binding protein. g) Incubation enrichment treatment, retaining the enriched single-stranded DNA fragments in the supernatant; preferably, the incubation is performed using streptavidin magnetic beads; h) Connecting a bridging adapter to the single-stranded DNA fragment; preferably, the bridging adapter contains a unique molecular identifier (UMI). (3) Library amplification and high-throughput sequencing: The sequencing library was amplified by PCR and then subjected to paired-end sequencing; (4) Data analysis and off-target site identification: a) Align the sequencing reads to the reference genome; preferably, use UMI to remove PCR duplicates to obtain unique reads; b) Filtering based on positional features: Only reads that show an A→G mutation at the second base position of the P7 primer are retained; Preferably, the reference genome at this position is adenine, has a base quality ≥ Q30, and has ≥ 3 unique reads; c) Filter background signals in control samples to identify off-target sites.

[0013] In some embodiments, in step (1), the adenine base editor system is selected from ABE7.10, ABE8e, ABE8e-WQ or variants thereof; And / or, the extraction time point is at least one of 24 hours, 48 ​​hours, 72 hours, and 96 hours after cell introduction; And / or, further include extracting genomic DNA at multiple time points to analyze the temporal dynamics of off-target accumulation; And / or, the concentration of the genomic DNA is ≥20 ng / µL, and the purity A260 / A280 ratio is 1.8-2.0.

[0014] In some embodiments, in step (2), the adapter sequence of the double-stranded DNA linear adapter includes P5-ps6(-) and P5-ps6(+)-biotin; The nucleotide sequence of the P5-ps6(-) is A*C*A*CTCTTTCCCTACACGACGCTCTTCCGA*T*C*T; The nucleotide sequence of the P5-ps6(+)-biotin is / 5Phos / G*A*T*CGGAAGAGCGTCGTGTAGGGAAAGAG* / iBiodT / *G*T; And / or, the nucleotide sequence of the positive strand of the bridging linker is / 5Phos / CCACGCGTGCTCTACANNNTNNNNTNNNAGATCGGAAGAGCACACGTCTGAACTCCAGT-NH2; the nucleotide sequence of the negative strand of the bridging linker is TGTAGAGCACGCGTGGNNNNNN-NH2. The asterisk (*) represents a thiophosphate bond; / 5Phos / represents 5' phosphorylation; / iBiodT / represents internal biotin-dT modification; NNNTNNNNTNNN represents the unique molecular identifier UMI; and NH2 represents 3' amino modification.

[0015] And / or, the fragmentation results in DNA fragment sizes ranging from 400 bp to 700 bp, with a peak at 500 bp; And / or, the enzymes used in the DNA damage repair treatment include endonuclease IV, Bst full-length polymerase, and Taq DNA ligase; the treatment is carried out in a system containing NAD⁺ and dNTPs, incubated at 37°C for 60 minutes and then at 45°C for 60 minutes. And / or, the exonuclease treatment is repeated at least three times, the exonuclease including exonuclease I and exonuclease III; the treatment is performed by incubation at 37°C for 2 hours followed by inactivation at 75°C for 10 minutes; And / or, the specific cleavage uses endonuclease V, preferably Escherichia coli endonuclease V, with a concentration ≥10 U / µL; the cleavage condition is incubation at 20°C for 1 hour; And / or, the denaturation is performed using a thermostable single-stranded DNA-binding protein derived from thermophilic bacteria (…). Thermus aquaticus The ET SSB protein was incubated at a concentration of 1 µM to 5 µM under the following denaturation conditions: 95 °C for 10 minutes followed by a rapid ice bath for 10 minutes. And / or, the incubation enrichment is performed using streptavidin magnetic beads, preferably with the incubation time controlled at 10 minutes to avoid the single-stranded DNA renaturation caused by excessive incubation time.

[0016] In some implementations, in step (3), the amplification includes a first round of amplification and a second round of amplification; The first round of amplification used MightyAmp DNA polymerase for two cycles of PCR amplification; the PCR amplification program was 98°C for 2 minutes; 98°C for 10 seconds, 68°C for 75 seconds, for 2 cycles; The second round of amplification used FastPfu DNA polymerase for 8-12 cycles of PCR amplification; the PCR amplification was performed at 95°C for 2 minutes; 95°C for 30 seconds, 58°C for 45 seconds, 72°C for 1 minute, for 8-12 cycles; 72°C for 5 minutes. And / or, the high-throughput sequencing uses paired-end 150 bp sequencing (PE150) with a sequencing depth ≥30 Gb.

[0017] In some implementations, in step (4), the position feature filtering is based on the following principle: due to the library construction design, the sequencing reads of the real inosine site show an A→G mutation at the second base of the P7 primer with 100% consistency, while the mutations of non-specific signals at this position are randomly distributed; the non-specific signals include sequencing errors, PCR artifacts, and other DNA damage; And / or, the threshold for identifying off-target sites is ≥4 unique reads; And / or, the control samples include at least one of sgRNA-only treatment, ABE-only treatment, and untreated blank controls.

[0018] In some embodiments, the inosine-enriched adenine base editor-based off-target detection method of the present invention further includes at least one of the following steps: a. Integrate chromatin state data to analyze the genomic characteristics of off-target sites; b. The chromatin state data includes H3K27ac, H3K36me3, H3K9me3 histone modification markers and / or ATAC-seq chromatin accessibility data; c. Calculate the enrichment fold of off-target sites in active enhancer regions, transcriptionally active regions, heterochromatin regions, and open chromatin regions.

[0019] Secondly, the present invention provides applications of the aforementioned adenine base editor-based off-target detection method based on inosine enrichment in at least the following aspects: 1) Evaluate the genome-wide off-target profile of the adenine base editor across different genomic loci and cell types, including sgRNA-dependent and non-sgRNA-dependent off-target events; 2) Screening and optimizing base editor variants with higher specificity; 3) Preclinical and clinical safety assessments of ABE-based gene therapy products, including off-target detection of patient-derived cells or clinical trial samples; 4) Detect endogenous genomic inosine in cells or tissues.

[0020] Thirdly, the present invention provides a method for verifying the off-target detection results of an adenine base editor, comprising the following steps: (1) Use the method described above to detect off-target sites and obtain a list of candidate off-target sites; (2) The candidate off-target sites are targeted and captured using a dual capture strategy based on UMI; (3) Perform deep sequencing to achieve a coverage of >10,000×; (4) Analyze sequencing data and the verification criteria are: the number of reads after UMI deduplication is >100 and the editing rate is significantly increased compared with the control.

[0021] Fourthly, the present invention provides a kit for off-target detection of adenine base editors, comprising: a. DNA fragmentation reagents; b. Terminal repair and A-tail addition reagents; c. The aforementioned double-stranded DNA linear adapters and bridging adapters; d. DNA ligase and its buffer solution; e. A mixture of DNA damage repair enzymes, the mixture comprising endonuclease IV, Bst full-length polymerase, Taq DNA ligase, NAD⁺, and dNTPs; f. DE exonuclease I with a concentration ≥20 U / μL, DE exonuclease III with a concentration ≥100 U / μL and their corresponding buffer solutions; g. Escherichia coli endonuclease V with a concentration ≥10 U / μL and its reaction buffer; h. A thermostable single-stranded DNA-binding protein with a concentration ≥5 μM; preferably derived from thermophilic bacteria ( Thermus aquaticus ET SSB protein; i. Streptavidin magnetic beads and binding / washing buffer; preferably, the binding / washing buffer comprises 5 mM Tris-HCl (pH 7.5), 0.5 mM EDTA and 1 M NaCl; j. Single-stranded DNA ligation reagents, including T4 DNA ligase and 50% PEG8000; k. PCR amplification reagents and primers; l. Magnetic bead purification reagent; Preferably, it also includes an experimental operation manual and a data analysis procedure description.

[0022] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) Excellent detection performance Systematic comparative studies show that Ino-seq achieves an F1 score of 0.892–0.931, significantly outperforming all existing methods: a 148–931-fold improvement compared to WGS (F1: 0.001–0.006); a 4.2–21.7-fold improvement compared to GUIDE-seq (F1: 0.043–0.213); a 27.9–29.1-fold improvement compared to Tracking-seq (F1: 0.032); and a 23.5–24.5-fold improvement compared to CHANGE-seq-BE (F1: 0.038). Targeted deep sequencing validation (>10,000× coverage) shows that, using a reads ≥4 threshold, Ino-seq achieves an overall validation rate of 95.3% (1,330 / 1,396), with validation rates for individual targets ranging from 80.0% to 100%.

[0023] (2) Comprehensive testing capabilities Ino-seq can simultaneously detect sgRNA-dependent and sgRNA-independent off-target events. In the analysis of eight genomic loci, the proportion of sgRNA-independent off-target events showed significant target-specific variation (12-100%): low proportions for HEK293site4 (12%), CD7 (33-37%); high proportions for B2M (94-96%), RNF2 (77-93%), and CBLB (91-100%). Crucially, Ino-seq identified 892 sites missed by all other methods, including treatment-related high-risk sites with editing rates >5%, which may pose clinical safety risks.

[0024] (3) Unique technical principles 1) Direct detection of inosine intermediates: Unlike methods that rely on Cas9 binding or DNA repair intermediates, Ino-seq directly captures the enzymatic product of ABE activity—inosine, thus enabling the detection of all types of ABE-induced editing, without being limited by sgRNA binding affinity or DNA repair kinetics.

[0025] 2) Characteristic positional signals: Library construction design generates unique positional signal characteristics—the actual inosine site consistently shows an A→G mutation at the second base of the P7 primer (100% positional consistency), while non-specific signals (sequencing errors, PCR artifacts, other DNA damage) are randomly distributed. P <0.001, χ ²Verification). This positional constraint, combined with UMI, provides powerful background filtering capabilities (signal-to-noise ratio > 50:1).

[0026] 3) Dual enrichment strategy: EndoV specific cleavage and streptavidin-biotin high affinity enrichment (Kd ≈ 10⁻¹) 5 By combining M) to achieve efficient capture of inosine-containing DNA fragments (enrichment factor >100-fold).

[0027] (4) Physiologically related cell detection Ino-seq, performed in live cells, can reflect the impact of native chromatin accessibility on off-target activity. Integrated cut & tag histone modification analysis showed that off-target sites were enriched 3.7-fold (95% CI: 3.4–4.0) at H3K27ac-labeled activity enhancers. P <0.001); enriched 2.6-fold (95% CI: 2.4–2.8) in H3K36me3-tagged transcriptional regions. P <0.001); only 1.4-fold enrichment in H3K9me3-labeled heterochromatin regions ( P= 0.08, not statistically significant; ATAC-seq analysis showed a 2.5-fold enrichment in open chromatin regions ( P <0.001). These findings indicate that ABE off-target activity is not only determined by sgRNA-DNA complementarity, but is also significantly affected by chromatin accessibility.

[0028] (5) Time dynamic analysis capability Multi-timepoint (24h, 48h, 72h, 96h) analysis revealed a complex time-dependent pattern of off-target accumulation: CD7 sites gradually increased from 27 sites at 24h to 122 sites at 96h (a 4.5-fold increase); HEK293 site4 peaked at 273 sites at 48h; only 17-22% of sites were consistently detected across all timepoints. This temporal heterogeneity reflects the dynamic editing process and the transient detection window of inosine intermediates, indicating that the cumulative off-target load substantially exceeds single-timepoint measurements (average 4.5-5.8-fold), highlighting the importance of multi-timepoint analysis.

[0029] (6) Supports base editor optimization Ino-seq was used to compare three variants of ABE7.10, ABE8e, and ABE8e-WQ (at CD7, CIITA, and PDCD1 targets, at 48h and 72h time points, for a total of 6 conditions): ABE8e showed a 2.6-fold increase in median off-target counts compared to ABE7.10 (33.5 vs 13). P =0.002); ABE8e-WQ reduced the number of off-target attacks by 64% compared to ABE8e (median 12). P = 0.001); ABE8e-WQ is equivalent to ABE7.10 ( P = 0.352). These results demonstrate that: 1) increased on-target activity does not guarantee a reduction in off-target effects; 2) off-target effects can be successfully reduced through rational engineering; and 3) comprehensive off-target analysis is required for the development of base editors.

[0030] (7) Expanded detection range In addition to ABE-induced inosine, Ino-seq can also detect endogenous genomic inosine (possibly derived from spontaneous deamination, ADAR enzyme activity, or oxidative stress). A variation range of >1,000-fold was observed in 10 human cell lines: high inosine cell lines HCT116 (12,765 sites), HeLa (9,234 sites), K562 (8,891 sites), HaCaT (7,136 sites); and low inosine cell lines HepG2 (12 sites), A549 (34 sites), THP-1 (56 sites), and HEK293T (78 sites). GO analysis showed that affected genes in high inosine cell lines were enriched in neural signaling pathways. P<0.01 (FDR corrected), providing a tool for studying the biological functions of endogenous DNA modifications.

[0031] (8) Clinical translational value Ino-seq provides a crucial preclinical and clinical safety assessment tool for ABE-based gene therapy: it can assess off-target effects in patient-derived iPSCs or primary cells; detect treatment-related high-risk sites (editing rate >5%); integrate chromatin state to assess off-target risk; and support regulatory safety reviews. This method has been successfully applied to the safety assessment of eight clinically relevant targets (CD7, PDCD1, B2M, CBLB, CIITA, VEGFA, etc.). Attached Figure Description

[0032] Figure 1 To explain the working principle and verification of the Ino-seq method; The figure shows (a) a schematic diagram of the Ino-seq workflow, which shows the complete steps from ABE cell treatment to off-target site identification, including: cell transfection → genomic DNA extraction → library construction (fragmentation, end repair, adapter ligation, damage repair, exonuclease treatment) → EndoV cleavage → DNA denaturation → streptavidin magnetic bead enrichment → bridging adapter ligation → PCR amplification → high-throughput sequencing → bioinformatics analysis; (b) Genome browser (IGV) visualization of HEK293 site4 showed strong signal enrichment (peak >200 reads) in the ABE+sgRNA group, while no detectable signal (<5 reads) was found in the three control groups (sgRNA-only, ABE-only, Blank), with a signal-to-noise ratio >40:1; (c) Position-specific signal verification: all reads showed an A→G mutation at +2 bp at the P7 primer end; 100% positional consistency. n =1,148 off-target sites). (d) Analysis of the top 50 off-target sites shows the number of reads, the number of mismatches, and the type of off-target; the editing rate is not strictly correlated with sgRNA sequence similarity, and the highest off-target site (397 reads) has 16 mismatches; (e) Off-target cumulative dynamic curves of 8 targets at 4 time points (24h, 48h, 72h, 96h).

[0033] Figure 2 A systematic comparison of Ino-seq with five orthogonal methods; The figure shows signal enrichment analysis (ab), which displays the signal-to-noise ratio of Ino-seq and WGS at off-target sites. (cf) Venn diagram showing overlap between methods targeting different targets; (gh)sgRNA-dependent and non-dependent off-target sequence motif analysis.

[0034] Figure 3 For targeted deep sequencing validation and performance evaluation; In the figure, (a) shows the validation rate statistics for 6 targets (overall validation rate 95.3%). (b) Scatter plot of the correlation between reads and edit rate ( R ² value annotation); (cd) Comparison of F1 scores and precision-recall rates for different methods.

[0035] Figure 4 For the integrated analysis of chromatin characteristics and off-target distribution; The figure shows the enrichment heatmap of off-target sites on H3K27ac, H3K36me3, H3K9me3 and ATAC-seq signals (±2 kb region, 50 bp resolution).

[0036] Figure 5 Off-target comparative analysis of ABE variants; In the figure, (a) box plot of the number of off-target sites for ABE7.10, ABE8e, and ABE8e-WQ under 6 conditions; (b) Statistical classification of sgRNA-dependent and non-dependent off-target effects; (c) Whole-genome Circos visualization shows the chromosomal distribution of off-target sites; (di) sgRNA-dependent and non-dependent off-target sequence motifs of the three variants.

[0037] Figure 6 For the detection of endogenous genomic inosine across cell lines; In the figure, (a) is a bar chart of the number of inosine sites in 10 cell lines (logarithmic scale, variation range >1,000-fold). (b) Box plot of the number of inosine sites between cell lines; (c) Pie chart of genomic annotation distribution of inosine sites. Detailed Implementation

[0038] To better illustrate the objectives, technical solutions, and advantages of this invention, the invention will be further described below with reference to specific embodiments. Those skilled in the art should understand that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0039] Unless otherwise specified, the experimental methods used in the examples are conventional methods; the materials and reagents used are commercially available unless otherwise specified.

[0040] Example 1: Establishment and Verification of the Ino-seq Method 1. Cell culture and ABE system introduction HEK293T cells were used as the model system. Four parallel experiments were designed, as shown in Table 2: Table 2. Four sets of parallel experimental setups Necessity of the control group: sgRNA control distinguishes DNA damage caused by Cas9 binding from ABE deamination; ABE control identifies global off-target editing independent of sgRNA; blank control identifies endogenous genomic inosine levels. All groups were cultured under the same conditions for 48 hours and harvested simultaneously. Genomic DNA was extracted and used for parallel Ino-seq library construction and sequencing.

[0041] (1) Cell culture HEK293T cells (ATCC CRL-3216) were cultured in DMEM medium (Gibco) containing 10% FBS (Gibco) and maintained at 37°C, 5% CO2, and saturated humidity. Cell density was maintained at 30-70% confluence, and cells were passaged every 2-3 days.

[0042] (2) Plasmid transfection 24 hours before transfection, HEK293T cells were inoculated at a rate of 2×10⁻⁶. 5 Cells were seeded at a density of 6-well plates. Cell confluence should reach 60%-70% at transfection. Transfection was performed using Lipofectamine 3000 (Invitrogen). Transfection system per well: ABE8e expression plasmid (pCMV-ABE8e-P2A-GFP): 1.5 µg; sgRNA expression plasmid (pU6-sgRNA-HEK293site4): 0.5 µg; Lipofectamine 3000: 5 µL; P3000 reagent: 4 µL; Opti-MEM: 250 µL.

[0043] Follow the manufacturer's instructions. Replace with complete culture medium 6 hours after transfection. Control group transfection setup: sgRNA control (1.5 µg empty vector + 0.5 µg sgRNA plasmid); ABE control (1.5 µg ABE8e plasmid + 0.5 µg empty vector); blank control (2 µg empty vector).

[0044] (3) Transfection efficiency detection and cell harvesting 24 hours after transfection, GFP-positive cells were observed under a fluorescence microscope to assess transfection efficiency. Typical transfection efficiency should be >50%. 48 hours after transfection (when ABE editing reaches a stable level), cells were harvested according to standard procedures.

[0045] (4) Genomic DNA extraction Genomic DNA was extracted using the DNeasy Blood & Tissue Kit (Qiagen). DNA quality testing: NanoDrop was used to determine concentration and purity (A260 / A280 should be 1.8-2.0); Qubit dsDNA HS Assay was used for accurate quantification; agarose gel electrophoresis was used to check integrity (high molecular weight bands, no degradation). DNA requirements: concentration ≥20 ng / µL; purity A260 / A280 = 1.8-2.0, A260 / A230 >2.0; integrity with no significant degradation; total amount ≥10 µg per sample (for subsequent library construction).

[0046] 2. Ino-seq library construction (1) Annealing of P5-ps6-Biotin connector Prepare an annealing system with a total volume of 50 µL, comprising 5 µL of 10×STE, 5 µL of H₂O, 20 µL of P₅-ps₆(-) (100 µM), and 20 µL of P₅-ps₆(+)-Biotin (100 µM). The annealing program was as follows: incubation at 95 °C for 5 minutes, followed by cooling at a rate of -1 °C per minute for 70 cycles, and finally holding at 4 °C.

[0047] The P5-ps6(-) sequence is as follows: A*C*A*CTCTTTCCCTACACGACGCTCTTCCGA*T*C*T (SEQ ID NO: 1), where * represents a thiophosphate bond.

[0048] The sequence of the P5-ps6(+)-Biotin is as follows: / 5Phos / G*A*T*CGGAAGAGCGTCGTGTAGGGAAAGAG* / iBiodT / *G*T (SEQ ID NO: 2), where / 5Phos / represents 5' phosphorylation, / iBiodT / represents internal biotin-dT modification, and * represents a thiophosphate bond.

[0049] (2) DNA disruption The sonication fragmentation volume was calculated at 10 µg / reaction. 10 µg of genomic DNA was fragmented using a Covaris M220 sonicator: peak power 50 W, duty cycle 10%, 200 cycles per cycle, 30-second on-off time, 6 cycles. Electrophoresis detection: 1 µL of sample was run on a 1.5% agarose gel to observe the DNA fragment size distribution. The fragmented library size should be 400-700 bp, with the main band at 500 bp. If the fragments were too large, fragmentation was continued. Purification of the fragmented product: Purification was performed using 1× volume of AMPure XP magnetic beads, eluted with 50 µL of TE buffer, and concentration measured using a quorum analyzer.

[0050] (3) End repair and A-tailing End repair was performed using the NEBNext Ultra II End Repair / dA-Tailing Module (NEB) at 20°C for 30 minutes, followed by A-tail addition at 65°C for 30 minutes. See the kit instructions for reaction details.

[0051] (4) Connector connection (add P5-ps6-Biotin connector) The annealed bilinker (P5-ps6(-) and P5-ps6(+)-Biotin, final concentration 40 µM) prepared in step (1) was ligated by incubation at 20 °C for 60 min. Ligation product purification: purified with 1× volume of purification magnetic beads (95 µL), eluted with 34 µL of TE buffer, and the concentration was determined by Qubit.

[0052] (5) Damage repair In a system containing endonuclease IV, Bst full-length polymerase, Taq DNA ligase, NAD⁺, and dNTPs, the mixture was incubated at 37°C for 60 minutes and then at 45°C for 60 minutes.

[0053] Prepare the following reaction system in a sterile PCR tube with a total volume of 50µL: 38µL of adapter ligation product DNA from step (4), 5µL of NEBuffer 3.0 (10×), 1µL of 50mM NAD⁺, 1µL of 2.5mM dNTPs, 2µL of Endo IV, 1µL of Bstfull-length polymerase, and 2µL of Taq DNA ligase.

[0054] The mixture was thoroughly mixed using a pipette before incubation. The PCR incubation procedure was as follows: heated lid: 50°C; incubation at 37°C for 60 minutes; incubation at 45°C for 60 minutes; hold at 4°C. Reaction principle: Endonuclease IV removes the 5' phosphate group and 3' abase site at the DNA damage site; Bst full-length polymerase fills the single-base gap; Taq DNA ligase ligates the nick; the 37°C step repairs most of the DNA damage; the 45°C step optimizes Taq ligase activity.

[0055] After the above reaction is complete, add 2× (100µL) magnetic beads for purification (total volume after addition is 150µL). Elute with 43µL of enzyme-free water, and take 1µL for Qubit measurement. After damage repair, subsequent operations should not involve vortexing the DNA sample; instead, gently mix with a pipette to prevent new DNA damage.

[0056] (6) Exonuclease treatment (repeated three times) The enzyme digestion reaction system was prepared with a total volume of 50 µL: exonuclease buffer (10×): 5 µL, exonuclease I: 2 µL, exonuclease III: 1 µL, and DNA from step (5) for damage repair: 42 µL. The enzyme reaction program was as follows: incubation at 37 °C for 2 hours (heated lid 42 °C); incubation at 75 °C for 10 minutes (heated lid 80 °C); and maintenance at 4 °C.

[0057] Exonuclease I is used to remove single-stranded DNA (DNA fragments without adapters), while exonuclease III digests the 3' end of double-stranded DNA. Together, they ensure thorough removal of DNA without adapters. New England Biolabs (NEB)'s M0293 (Exo I) and M0206 (Exo III) were used. After exonuclease digestion, the product was purified using magnetic beads under the following conditions: 1× beads (50µL), 80% ethanol wash ×2, TE buffer elution volume: 43µL. 1µL of the elution buffer was used to measure the Qubit concentration to determine the adapter ligation efficiency (amount of DNA after exonuclease digestion / amount of DNA added to the library). This step needs to be repeated three times to thoroughly remove any DNA fragments without adapter ligation.

[0058] (7) EndoV digestion Prepare the following reaction system in a sterile PCR tube with a total volume of 50µL: DNA product from exonuclease treatment in step (6): 42µL, NEBuffer 4 (10×): 5µL, Endonuclease V: 1µL, and enzyme-free water: 2µL.

[0059] The concentration of the endonuclease V is 10 U / µL. Endo V is an inosine-specific nuclease that recognizes and cleaves inosine (adenine deamination product) in DNA, thereby creating a nick at the ABE editing site.

[0060] After thoroughly mixing with a pipette, incubate using the following PCR procedure: hot cap: 42℃; incubate at 20℃ for 1 hour; hold at 4℃.

[0061] After the above reaction was completed, 2× (100µL) magnetic beads were added for purification (the total volume after addition was 150µL). 43µL of enzyme-free water was used for elution, and 1µL was taken to measure the Qubit concentration.

[0062] (8) DNA denaturation and streptavidin capture First, remove the Dynabeads® MyOne Streptavidin T1 magnetic beads from 4°C, gently mix them, and let them sit at room temperature for half an hour to reach room temperature.

[0063] Denaturation and Cooling of DNA: Add the thermostable single-stranded DNA binding protein ET SSB (derived from thermophilic bacteria) to 42 µL of double-stranded DNA (dsDNA) digested by EndoV in step (7) above. Thermus aquaticus To achieve a final ET SSB concentration of 2 µM (recommended range 1-5 µM), add 10 µM of ET SSB stock solution to a stock solution of 10 µM. Example calculation: DNA volume 42 µL, target final concentration 2 µM, ET SSB stock solution concentration 10 µM, required ET SSB volume: 42 × 2 / 10 = 8.4 µL (add enzyme-free water to a total volume of 50 µL). After thorough mixing, place the reaction tube in a PCR instrument: denature at 95°C for 10 minutes (to completely denature double-stranded DNA into single-stranded DNA); quickly transfer to ice and incubate for 10 minutes (ET SSB binds to single-stranded DNA, preventing annealing). Mechanism of action of ET SSB: This protein maintains high affinity for single-stranded DNA even at 95°C, forming a stable protein-DNA complex, effectively preventing single-stranded DNA annealing at room temperature. If ET SSB is not added or the concentration is insufficient, single-stranded DNA will anneal rapidly, leading to a significant decrease in streptavidin magnetic bead enrichment efficiency (from >80% to <20%).

[0064] Wash the magnetic beads three times: Take 25 µL of streptavidin magnetic beads (Dynabeads MyOne Streptavidin T1) and add 5 volumes (125 µL) of binding / washing buffer (B&W Buffer: 5 mM Tris-HCl pH 7.5, 0.5 mM EDTA, 1 M NaCl) for washing. After mixing, rotate for 5 minutes, discard the supernatant, and repeat the washing process three times. Alternatively, the Binding & Washing Buffer that comes with DynabeadsMyOne Streptavidin T1 can be used. Resuspend the magnetic beads in 42 µL of B&W Buffer and add it to 50 µL of denatured DNA (total volume approximately 92 µL). Mix thoroughly by pipetting (keeping on ice), then transfer to a rotary mixer at room temperature and incubate for 10 minutes.

[0065] The incubation time should be strictly controlled within 10 minutes. Incubation time that is too long will cause single-stranded DNA to renature, reducing the concentration of single-stranded DNA.

[0066] After incubation, place the sample on a magnetic rack for 5 minutes, retain the supernatant (containing enriched single-stranded DNA fragments), add 2 µL of proteinase K (20 mg / mL), and incubate at 37°C for 15 minutes to remove ET SSB protein. After proteinase K digestion, add 1 µL of carrier RNA (1 µg / µL) to the sample, purify the DNA using the QIAquick PCR Purification Kit, and wash with 45 µL of enzyme-free water. Take 1 µL of each sample for Qubit ssDNA and dsDNA concentration detection to confirm the single-stranded DNA enrichment efficiency.

[0067] (9) Single-stranded DNA bridging adapter connection Prepare the reaction system in a sterile PCR tube: streptavidin capture DNA from step (8): 42.4 µL, 10× T4 ligase buffer: 8 µL, Bridge adapter (50 µM, sense and antisense strands 1:1): 1.6 µL, T4 DNA ligase (400 U / µL): 4 µL.

[0068] The bridging adapter includes a positive linker (Bridge adapter-upper-12) and an antisense linker (Bridge adapter-lower-12). The positive linker sequence is: / 5Phos / CCACGCGTGCTCTACANNNTNNNNTNNNAGATCGGAAGAGCACACGTCTGAACTCCAGT-NH2 (SEQ ID NO: 3), where / 5Phos / represents 5' phosphorylation, NNNTNNNNTNNNN represents 12 random base sequences as a unique molecular identifier (UMI), and -NH2 represents 3' amino modification. The antisense linker sequence is TGTAGAGCACGCGTGGNNNNNN-NH2 (SEQ ID NO: 4), where NNNNNN represents 6 random base sequences, and -NH2 represents 3' amino modification.

[0069] After thoroughly mixing by blowing and blowing, add 24µL of 50% (by weight / volume) PEG8000, bringing the total volume to 80µL. Since the mixture becomes very viscous after adding PEG8000, it needs to be slowly blown and blown 10-15 times with a wide-mouth pipette tip until fully mixed.

[0070] The above reaction system was placed at room temperature and incubated at the lowest speed of a rotary mixer for 14 hours. During the incubation process, the mixture was gently blown with a wide-mouth pipette tip every 2 hours to improve the joint connection efficiency.

[0071] After the reaction was completed, 1 µL of carrier RNA (1 µg / µL) was added to the sample, and the DNA was purified using the QIAquick PCRPurification Kit. After elution with 22 µL of enzyme-free water, 1 µL of the purified DNA was used to detect the concentrations of Qubit ssDNA and dsDNA.

[0072] (10) Library expansion First round of amplification: Prepare the reaction system on ice with a total volume of 50 µL, including: DNA ligated by the adapter in step (9): 20 µL, I5 primer (10 µM): 2 µL, I7 primer (10 µM): 2 µL, 2× MightyAmp Buffer Ver. 3 (Mg²⁺, dNTP plus): 25 µL, and MightyAmp DNA polymerase Ver. 3: 1 µL.

[0073] After thoroughly mixing with pipette tip, incubate using the following PCR program: 98℃ for 2 minutes; 98℃ for 10 seconds, 68℃ for 75 seconds, 2 cycles; hold at 4℃.

[0074] After the reaction, add 1 µL of carrier RNA, then purify using a PCR product purification kit, eluting with 64 µL of enzyme-free water. When using the QIAquick PCR Purification Kit, after adding Buffer PB, add 10 µL of 3M sodium acetate to adjust the solution color back to yellow. After purification, take 1 µL to determine the Qubit concentration.

[0075] Second round of amplification: Prepare the reaction system on ice with a total volume of 100 µL, including: DNA from the first round of amplification: 63 µL, 5× FastPfu buffer: 20 µL, dNTPs (2.5 mM for each): 8 µL, P5 primer (10 µM): 4 µL, P7 primer (10 µM): 4 µL, and FastPfu DNA polymerase (2.5 U / µL): 1 µL.

[0076] Perform 8-12 cycles of PCR amplification according to the following procedure: 95℃ for 2 minutes; 95℃ for 30 seconds, 58℃ for 45 seconds, 72℃ for 1 minute, 8-12 cycles; 72℃ for 5 minutes; hold at 10℃.

[0077] Gel extraction: After PCR, perform 2% agarose gel electrophoresis directly to recover the 300-700 bp bands. Purify the bands using the QIAquick Gel Extraction Kit to obtain the final sequencing library.

[0078] 3. High-throughput sequencing The prepared sequencing library was sequenced using the MGI T7 sequencing platform, with a sequencing data volume of 30G (150bp sequencing at both ends) to ensure sufficient sequencing depth to detect low-frequency off-target events.

[0079] 4. Data Analysis (1) Quality control of raw data Quality filtering was performed using fastp (version 0.23.2) with the following conditions: removal of low-quality bases (Q<20), removal of adapter sequences, and retention of reads ≥50 bp in length.

[0080] (2) Sequence alignment Clean reads were aligned to the human reference genome (hg38) using BWA-MEM (version 0.7.17), and then sorted and indexed using samtools.

[0081] (3) UMI treatment and PCR duplicate removal Use UMI-tools (version 1.1.2) to extract UMI sequences and remove PCR duplicates.

[0082] (4) Location-specific filtering Only reads showing an A→G mutation at the second base position of the P7 primer are retained. Specifically, the reference genome must contain adenine (A) at this position, the sequencing read must contain guanine (G) at this position, and the base quality must be ≥Q30; the number of unique reads must be ≥3. This filtering strategy is based on the ABE editing mechanism: ABE deaminates adenine (A) to inosine (I), which is read as guanine (G) during PCR amplification and sequencing. Therefore, the A→G mutation is a characteristic signal of ABE editing.

[0083] (5) Off-target site identification For each candidate site, the number of unique reads supporting that site (after deduplication of UMI) is counted. A threshold of ≥4 unique reads is set. Background signals are filtered in the control samples, and a list of off-target sites is output, including chromosomal location, number of unique reads, editing frequency, number of mismatches with sgRNA sequence, and off-target type.

[0084] 5. Results like Figure 1 As shown in (b), 1,148 off-target sites were detected at 48 hours against the HEK293 site4. IGV visualization revealed strong signal enrichment at the on-target sites (chr20: 32761949-32761972), while no detectable signal was observed in any of the control groups. Figure 1 As shown in (c), all captured reads exhibited a characteristic A→G substitution at the second base of the P7 primer, confirming the specific detection of inosine-containing DNA fragments. Figure 1 (d) shows the Top 50 off-target site analysis: 71 reads (0 mismatches) at the on-target site; 397 reads (16 mismatches, sgRNA-independent) at the most significant off-target site; and 100-214 reads (6-8 mismatches, sgRNA-dependent) at moderately strong off-target sites. This pattern demonstrates that the editing rate is not strictly correlated with sgRNA sequence similarity, highlighting the ability of Ino-seq to comprehensively detect both sgRNA-dependent and non-sgRNA-dependent editing events.

[0085] This embodiment successfully established the Ino-seq method and validated it in HEK293T cells. This method can detect ABE-induced genomic inosine with high sensitivity, providing a reliable tool for ABE safety assessment. Compared with existing off-target detection methods, Ino-seq has the following advantages: 1) Direct detection of the enzymatic product of ABE, inosine; 2) Simultaneous detection of sgRNA-dependent and non-dependent off-target effects; 3) A filtering strategy based on characteristic position signals provides strong background removal capabilities; 4) The use of UMI eliminates PCR amplification bias and accurately calculates editing frequencies. Key technical points summarized: DNA fragment size should be strictly controlled between 400-700 bp, with the major band at 500 bp; exonuclease treatment must be repeated three times to ensure thorough removal of unligated adapter DNA; the recommended ET SSB protein concentration is 1-5 µM to prevent single-stranded DNA renaturation; streptavidin magnetic bead incubation time should be strictly controlled to 10 minutes, as prolonged incubation will lead to DNA renaturation; after magnetic bead enrichment, retain the supernatant (not the magnetic beads), as the supernatant contains enriched single-stranded DNA fragments; the bridging adapter becomes very viscous after adding PEG8000, requiring slow pipetting with a wide-mouth pipette tip; after damage repair, do not vortex the DNA sample to prevent new DNA damage. Following these operational guidelines ensures experimental reproducibility and accurate results.

[0086] Example 2: Multi-target, multi-time point analysis Off-target sites at eight genomic loci (CD7, HEK293 site4, PDCD1, VEGFA site3, B2M, CBLB, CIITA, and RNF2) were analyzed using the Ino-seq method described in Example 1. Among these targets, HEK293 site4, RNF2, and VEGFA site3 were baseline sites; CD7, PDCD1, B2M, CBLB, and CIITA were clinically relevant therapeutic sites. Samples were collected at 24 h, 48 h, 72 h, and 96 h post-transfection to capture the temporal dynamics of off-target accumulation.

[0087] The test results are as follows: (1) Time-dependent off-target accumulation like Figure 1 As shown in (e): CD7 sites gradually increased from 27 sites at 24h to 122 sites at 96h (reaching a plateau between 72 and 96h); HEK293 site4 peaked at 273 sites at 48h; PDCD1 and VEGFA site3 showed lower off-target loads (6–22 and 48–67 sites, respectively). Key finding: Only 17–22% of sites were consistently detected across all time points, indicating that the cumulative off-target load substantially exceeded single-time-point measurements, highlighting the importance of multi-time-point analysis.

[0088] (2) sgRNA-independent off-target variants Significant target-specific variations were observed across different targets: major sgRNA-dependent HEK293 site4 (12%), CD7 (33-37%); high non-dependent proportions B2M (94-96%), RNF2 (77-93%), CBLB (91-100%); time-dependent changes VEGFA site3 decreased from 58% to 25%, while PDCD1 increased from 55% to 81%.

[0089] (3) Comparison with calculated predictions The consistency between Cas-OFFinder predictions (≤8 mismatches) and experimentally detected off-target sites is limited. For example, with CD7, only 9-33 off-target sites were predicted in the experimental tests; 18-94 additional sites were not predicted. This result indicates that computational prediction tools based solely on sequence matching cannot comprehensively identify actual off-target sites, highlighting the necessity of experimental detection.

[0090] Example 3: Systematic Comparison with Five Orthogonal Methods To comprehensively evaluate the performance of Ino-seq, a systematic comparison was conducted with five established off-target detection methods. The five established off-target detection methods include: GUIDE-seq, CHANGE-seq-BE, Tracking-seq, Selict-seq, and whole-genome sequencing (WGS) (100X depth sequencing).

[0091] (1) Analyze signal enrichment at HEK293 site4 and CD7 off-target sites. like Figure 2 As shown in (ab): Ino-seq shows sharp enrichment peaks at the edit sites (normalized signal ~6-8); WGS shows near-baseline signals, making it difficult to distinguish from random genomic sampling. The significant difference in signal-to-noise ratio reflects the advantage of Ino-seq's targeted enrichment strategy, directly capturing inosine-containing fragments rather than relying on sequence coverage.

[0092] (2) Inter-method overlap analysis reveals limited overlap between different methods. like Figure 2(cf) shows that: HEK293 site4 (72h) Ino-seq has 1,234 candidate sites (unique reads≥3), of which 1,148 pass the threshold verification (unique reads≥4), with 59 overlapping with GUIDE-seq, 19 overlapping with Tracking-seq, and 205 overlapping with WGS; CD7 (72h) Ino-seq has 274 sites, with 58 overlapping with CHANGE-seq-BE, 4 overlapping with GUIDE-seq, and 12 overlapping with WGS; VEGFA site3 (72h) compared with Selict-seq, Ino-seq has 66 sites, of which 25 overlap with GUIDE-seq, while Selict-seq has 36 sites, with only 4 overlapping with GUIDE-seq. The two EndoV-based methods show a >6-fold difference in GUIDE-seq consistency. These results indicate that Ino-seq provides the most comprehensive off-target detection, particularly for sgRNA-independent modifications that are undetectable by binding-based methods.

[0093] (3) Sequence motif analysis of off-target sites distinguished between sgRNA-dependent and non-dependent events. like Figure 2 As shown in (gh): Dependent off-target sites have strong sequence similarity to on-target sites, especially in the seed region (positions 1-12), which is conserved and is typical of NGG PAM motifs; Independent off-target sites lack sequence similarity to on-target sites, showing adenine enrichment, consistent with the editing window preference of TadA8e.

[0094] Example 4: Targeted deep sequencing verification To assess the specificity of Ino-seq, targeted deep sequencing was used to validate the predicted off-target sites of six targets (HEK293 site4, B2M, PDCD1, RNF2, CIITA, and CD7).

[0095] Site selection consisted of all Ino-seq predicted sites with ≥3 reads, plus representative low-read sites. Validation was performed using a dual capture method based on UMI, achieving >10,000X coverage. Validation criteria were >100 reads after UMI deduplication and an edit rate >0.01% relative to the control.

[0096] Background levels were determined using control sample analysis, with the mean background set at 1.1 reads (standard deviation = 0.9); mean + 2.5 times standard deviation = 3.4; a recommended threshold of reads ≥ 4 was suggested. AUPRC and AUROC analyses confirmed that this threshold achieved the optimal balance between accuracy and sensitivity.

[0097] like Figure 3 As shown in (a), using a reads ≥4 threshold, Ino-seq achieved high validation rates at each target: HEK293site4 1,148 / 1,192 (96.3%); CIITA 36 / 38 (94.7%); B2M 7 / 7 (100%); CD7 203 / 229 (88.6%); PDCD1 28 / 33 (84.8%); RNF2 12 / 15 (80.0%). Overall precision: 95.3% (1,330 / 1,396). Here, 1,192 represent all candidate sites (unique reads ≥3), and 1,148 represent sites after threshold screening (unique reads ≥4). All statistical analyses were based on the validation threshold (reads ≥4).

[0098] Figure 3 The scatter plot showing the correlation between the number of reads and the edit rate, as shown in (b), indicates a high correlation. R (²>0.85). For example... Figure 3 As shown in (cd), a systematic comparison with five orthogonal methods demonstrates the superior performance of Ino-seq: F1 scores Ino-seq 0.892–0.931, WGS 0.001–0.006, GUIDE-seq 0.043–0.213, Tracking-seq 0.032, and CHANGE-seq-BE 0.038. Ino-seq identified 892 sites missed by all other methods, including treatment-related sites with edit rates >5%, which may constitute clinical safety concerns.

[0099] Example 5: Integrative Analysis of Chromatin Features To investigate the impact of chromatin accessibility on ABE off-target activity, Ino-seq data were integrated with CUT&Tag histone modification profiles for analysis: H3K27ac (activation enhancer marker), H3K36me3 (activation transcription marker), H3K9me3 (heterochromatin marker), and ATAC-seq (chromatin accessibility). CUT&Tag analysis was performed using HEK293T cells, following standard procedures.

[0100] like Figure 4 As shown, histone modification signals were analyzed in the ±2 kb region around the off-target site, revealing a significant enrichment pattern: H3K27ac was enriched 3.7-fold (95% CI: 3.4–4.0). P <0.001 (active enhancer); H3K36me3 2.6-fold enrichment (95% CI: 2.4-2.8, P<0.001 (transcriptional region); H3K9me3 1.4-fold enrichment ( P = 0.08, no significant difference (heterochromatin, minimum); ATAC-seq 2.5-fold enrichment ( P <0.001 (open chromatin region). These findings indicate that ABE off-target activity is determined not only by sgRNA-DNA complementarity but also by chromatin accessibility. This explains why computational prediction tools based solely on sequence matching (such as Cas-OFFinder) have limited consistency with experimentally detected sites.

[0101] Genome annotation analysis of off-target sites showed that they comprised 37-44% of intronic regions, 15-29% of intergenic regions, and 2-7% (least) of exon regions. Repetitive element analysis revealed that off-target sites were enriched in repetitive regions, particularly LINE and SINE elements, reflecting high copy numbers and sequence similarity of these elements.

[0102] Example 6: Comparative Analysis of ABE Variants Ino-seq was used to evaluate the specificity improvements of three ABE variants: ABE7.10 (an earlier version), ABE8e (improved on-target activity), and ABE8e-WQ (engineered variants with W90Y and Q154R mutations). Analysis was performed at 48h and 72h time points for the three targets (CD7, CIITA, PDCD1) (a total of six conditions).

[0103] like Figure 5 As shown in (a): ABE8e vs ABE7.10, ABE8e has a median of 33.5 loci, while ABE7.10 has a median of 13 loci, an increase of 2.6 times. P = 0.002, Wilcoxon rank-sum test); ABE8e-WQ vs ABE8e, ABE8e-WQ median 12 loci, comparable to ABE7.10 ( P = 0.352), a 64% reduction compared to ABE8e ( P = 0.001). Example of target specificity CIITA-72h: 47 sites of ABE8e, 10 sites of ABE7.10, and 2 sites of ABE8e-WQ.

[0104] like Figure 5 As shown in (di), sequence motif analysis of the three variants revealed that: for sgRNA-dependent off-target, all variants maintained similar recognition patterns, the proximal seed region of PAM was conserved, and the typical NGG PAM motif was observed; for sgRNA-independent off-target, there was a lack of on-target sequence similarity, but adenine enrichment was observed in the editing window. Figure 5(b) shows the classification statistics of sgRNA-dependent and non-dependent off-target effects, revealing the distribution of off-target types for different variants. Figure 5 (c) shows the chromosomal distribution of off-target sites in the whole genome Circos visualization.

[0105] These results demonstrate that: 1) Increased on-target activity (ABE8e vs ABE7.10) does not guarantee a reduction in off-target effects; 2) Off-target effects can be successfully reduced through rational engineering (ABE8e-WQ); 3) Comprehensive off-target analysis is required during the development of base editors; and 4) In clinical applications, a trade-off between on-target efficiency and off-target load should be made based on target specificity analysis.

[0106] Example 7: Detection of endogenous genomic inosine In addition to ABE-induced inosine, Ino-seq can also detect endogenous genomic inosine (possibly derived from spontaneous deamination or enzymatic modification). Ten human cell lines were selected for analysis: A549 (lung adenocarcinoma), C33A (cervical cancer), HaCaT (keratinocytes), HCT116 (colon cancer), HEK293T (renal epithelial cells), HeLa (cervical cancer), HepG2 (liver cancer), K562 (leukemia), SiHa (cervical cancer), and THP-1 (monocytes).

[0107] like Figure 6 As shown in (a), the results showed a total of >1,000-fold variation in endogenous inosine levels: high inosine cell lines HCT116 (12,765 sites), HeLa (9,234 sites), K562 (8,891 sites), HaCaT (7,136 sites); and low inosine cell lines HepG2 (12 sites), A549 (34 sites), THP-1 (56 sites), HEK293T (78 sites), and C33A (95 sites). Figure 6 (b) shows a box plot of the number of inosine sites among cell lines, revealing significant differences. Figure 6 (c) shows the distribution of endogenous genomic inosine in the genome: 38-44% in intronic regions; 15-30% in intergenic regions; and <7% in exon regions. GO analysis of affected genes in high inosine cell lines showed enrichment in neural signaling pathways (P<0.01, FDR corrected), suggesting potential association with specific biological functions or stress responses. These findings demonstrate that Ino-seq can detect genomic inosine from multiple sources, providing a tool for studying endogenous DNA modifications and expanding the application of this method beyond base editing.

[0108] Example 8: Clinically Relevant Application Examples (1) Preclinical safety assessment For ABE treatment strategies intended for clinical use, Ino-seq methods can be used for comprehensive off-target assessment. This includes: performing ABE editing in relevant cell types (such as patient-derived iPSCs or primary cells); detecting off-target effects using Ino-seq at multiple time points (24h-96h); integrating chromatin state data to assess off-target risks in open chromatin regions; functionally annotating detected off-target sites to assess potential pathogenic risks; and validating high-risk sites using targeted deep sequencing.

[0109] (2) Base editor optimization Ino-seq can be used to guide optimization when developing novel ABE variants. This includes: comparing the off-target profiles of candidate variants across multiple targets; assessing the relative proportions of sgRNA-dependent and non-sgRNA-dependent off-targets; selecting variants that minimize off-targets while maintaining on-target activity; and iterative testing until an acceptable safety threshold is reached.

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A method for off-target detection of adenine base editor based on inosine enrichment, characterized in that, Includes the following steps: S1: The adenine base editor system was introduced into cells, and genomic DNA was extracted after culture and inosine site-specific cleavage was performed. S2: Enrichment of DNA fragments containing inosine; S3: Construct the gene library enriched with inosine-containing DNA fragments and perform high-throughput sequencing; S4: Identify off-target editing sites in the high-throughput sequencing based on location features and count the reads with the location features; the location features are reads showing A→G mutations.

2. The off-target detection method for adenine base editor based on inosine enrichment according to claim 1, characterized in that, Includes the following steps: (1) Sample preparation: The adenine base editor system was introduced into the cells, and genomic DNA was extracted after culturing; (2) Sequencing library construction: a) Fragment the genomic DNA, perform end repair, and add an A tail; b) Connecting linear adapters to double-stranded DNA; c) Perform DNA damage repair treatment; d) Use exonuclease to remove DNA fragments without ligated adapters; e) Perform specific cutting; f) Denature the cut DNA and stabilize the single-stranded DNA using a single-stranded DNA binding protein; g) Incubation enrichment treatment, retaining the single-stranded DNA fragments enriched in the supernatant; h) Connect the bridging adapter to the single-stranded DNA fragment; (3) Library amplification and high-throughput sequencing: The sequencing library was amplified by PCR and then subjected to paired-end sequencing; (4) Data analysis and off-target site identification: a) Align the sequencing reads to the reference genome; b) Filtering based on positional features: Only reads that show an A→G mutation at the second base position of the P7 primer are retained; c) Filter background signals in control samples to identify off-target sites.

3. The off-target detection method for adenine base editor based on inosine enrichment according to claim 2, characterized in that, In step (1), the adenine base editor system is selected from ABE7.10, ABE8e, ABE8e-WQ or variants thereof; And / or, the extraction time point is at least one of 24 hours, 48 ​​hours, 72 hours, and 96 hours after cell introduction; And / or, further include extracting genomic DNA at multiple time points to analyze the temporal dynamics of off-target accumulation; And / or, the concentration of the genomic DNA is ≥20 ng / µL, and the purity A260 / A280 ratio is 1.8-2.

0.

4. The off-target detection method for adenine base editor based on inosine enrichment according to claim 2, characterized in that, In step (2), the adapter sequence of the double-stranded DNA linear adapter includes P5-ps6(-) and P5-ps6(+)-biotin; The nucleotide sequence of the P5-ps6(-) is A*C*A*CTCTTTCCCTACACGACGCTCTTCCGA*T*C*T; The nucleotide sequence of the P5-ps6(+)-biotin is / 5Phos / G*A*T*CGGAAGAGCGTCGTGTAGGGAAAGAG* / iBiodT / *G*T; And / or, the nucleotide sequence of the positive strand of the bridging linker is / 5Phos / CCACGCGTGCTCTACANNNTNNNNTNNNAGATCGGAAGAGCACACGTCTGAACTCCAGT-NH2; the nucleotide sequence of the negative strand of the bridging linker is TGTAGAGCACGCGTGGNNNNNN-NH2. The asterisk (*) indicates a thiophosphate bond; / 5Phos / indicates 5' phosphorylation; / iBiodT / indicates internal biotin-dT modification; NNNTNNNNTNNN indicates a unique molecular identifier (UMI); NH2 indicates 3' amino modification. And / or, the fragmentation results in DNA fragment sizes ranging from 400 bp to 700 bp, with a peak at 500 bp; And / or, the enzymes used in the DNA damage repair treatment include endonuclease IV, Bst full-length polymerase, and Taq DNA ligase; the treatment is carried out in a system containing NAD⁺ and dNTPs, incubated at 37°C for 60 minutes and then at 45°C for 60 minutes. And / or, the exonuclease treatment is repeated at least three times, the exonuclease including exonuclease I and exonuclease III; the treatment is performed by incubation at 37°C for 2 hours followed by inactivation at 75°C for 10 minutes; And / or, the specific cleavage is performed using endonuclease V at a concentration ≥10 U / µL; the cleavage conditions are incubation at 20°C for 1 hour. And / or, the denaturation is performed using a thermostable single-stranded DNA-binding protein derived from the ET SSB protein of thermophilic bacteria; its concentration is 1µM-5µM; the denaturation conditions are incubation at 95°C for 10 minutes, followed by a rapid ice bath for 10 minutes; And / or, the enrichment process is performed using streptavidin magnetic beads.

5. The off-target detection method for adenine base editor based on inosine enrichment according to claim 1, characterized in that, In step (3), the amplification includes a first round of amplification and a second round of amplification; The first round of amplification used MightyAmp DNA polymerase for two cycles of PCR amplification; the PCR amplification program was 98°C for 2 minutes; 98°C for 10 seconds, 68°C for 75 seconds, for two cycles; The second round of amplification used FastPfu DNA polymerase for 8-12 cycles of PCR amplification; the PCR amplification was performed at 95°C for 2 minutes; 95°C for 30 seconds, 58°C for 45 seconds, 72°C for 1 minute, for 8-12 cycles; 72°C for 5 minutes. And / or, the high-throughput sequencing uses paired-end 150 bp sequencing (PE150) with a sequencing depth ≥30 Gb.

6. The off-target detection method for adenine base editor based on inosine enrichment according to claim 1, characterized in that, In step (4), the position feature filtering is based on the following principle: due to the library construction design, the sequencing reads of the real inosine site show an A→G mutation at the second base of the P7 primer with 100% consistency, while the mutations of non-specific signals at this position are randomly distributed; the non-specific signals include sequencing errors, PCR artifacts, and other DNA damage; And / or, the threshold for identifying off-target sites is ≥4 unique reads; And / or, the control samples include at least one of sgRNA-only treatment, ABE-only treatment, and untreated blank controls.

7. The off-target detection method for adenine base editor based on inosine enrichment according to any one of claims 1-6, characterized in that, It also includes at least one of the following steps: a. Integrate chromatin state data to analyze the genomic characteristics of off-target sites; b. The chromatin state data includes H3K27ac, H3K36me3, H3K9me3 histone modification markers and / or ATAC-seq chromatin accessibility data; c. Calculate the enrichment fold of off-target sites in active enhancer regions, transcriptionally active regions, heterochromatin regions, and open chromatin regions.

8. The off-target detection method of the adenine base editor based on inosine enrichment as described in any one of claims 1-7 is used in at least one of the following aspects: 1) Evaluate the genome-wide off-target profile of the adenine base editor across different genomic loci and cell types, including sgRNA-dependent and non-sgRNA-dependent off-target events; 2) Screening and optimizing base editor variants with higher specificity; 3) Preclinical and clinical safety assessments of ABE-based gene therapy products, including off-target detection of patient-derived cells or clinical trial samples; 4) Detect endogenous genomic inosine in cells or tissues.

9. A method for verifying the off-target detection results of an adenine base editor, characterized in that, Includes the following steps: (1) Use the method described in any one of claims 1-7 to detect off-target sites and obtain a list of candidate off-target sites; (2) The candidate off-target sites are targeted and captured using a dual capture strategy based on UMI; (3) Perform deep sequencing to achieve a coverage of >10,000×; (4) Analyze sequencing data and the verification criteria are: the number of reads after UMI deduplication is >100 and the editing rate is significantly increased compared with the control.

10. A kit for off-target detection of adenine base editor, characterized in that, include: a. DNA fragmentation reagents; b. Terminal repair and A-tail addition reagents; c. The double-stranded DNA linear adapter and bridging adapter as described in claim 4; d. DNA ligase and its buffer solution; e. A mixture of DNA damage repair enzymes, the mixture comprising endonuclease IV, Bst full-length polymerase, Taq DNA ligase, NAD⁺, and dNTPs; f. DE exonuclease I with a concentration ≥20 U / μL, DE exonuclease III with a concentration ≥100 U / μL and their corresponding buffer solutions; g. Escherichia coli endonuclease V with a concentration ≥10 U / μL and its reaction buffer; h. Heat-resistant single-stranded DNA binding proteins with a concentration ≥5 μM; i. Streptavidin magnetic beads and binding / washing buffer; j. Single-stranded DNA ligation reagents, including T4 DNA ligase and 50% PEG8000; k. PCR amplification reagents and primers; l. Magnetic bead purification reagent.