TIGR-tas cytosine base editing system and application thereof

By fusing cytosine deaminase with the TIGR-Tas system effector protein parTasR, and optimizing TIGR-Tas-CBE-V3 using tigRNA mismatch programming and a fluorescent reporter system, the problems of PAM dependence and low editing efficiency were solved, achieving efficient and specific cytosine base editing, which is suitable for disease models and gene therapy.

CN122629031APending Publication Date: 2026-08-25AGSINO GENSOURCES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610808785.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing RNA-guided nucleases rely on specific protospacer sequence neighbor motifs (PAMs) to recognize limited target ranges. The TIGR-Tas framework has low editing efficiency and produces many byproducts in mammalian cells, making it difficult to meet the needs of clinical applications.

Method used

Cytosine deaminase was fused with the TIGR-Tas system effector protein parTasR. ParTasR was induced to a nickase-like catalytic state through tigRNA mismatch programming. Compatible deaminases were screened using a fluorescent reporter system, and the variant TIGR-Tas-CBE-V3 was optimized.

Benefits of technology

It overcomes the limitations of PAM, expands the target range, improves editing efficiency, reduces double-strand break-related byproducts, and maintains genome and transcriptome specificity, making it suitable for disease model construction and gene therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122629031A_ABST
    Figure CN122629031A_ABST
Patent Text Reader

Abstract

The application belongs to the field of bioengineering and gene editing technology, and particularly relates to a TIGR-Tas cytosine base editing system and application thereof. The application provides a PAM-independent, efficient controllable and low byproduct TIGR-Tas cytosine base editor and an optimized variant thereof by fusing cytosine deaminase with a TIGR-Tas system effector protein parTasR and combining mismatch programming of tigRNA to induce parTasR to be in a nickase-like catalytic state, so as to realize windowed C→T base conversion in mammalian cells and provide better technical support for disease model construction and gene therapy and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bioengineering and gene editing technology, specifically relating to a TIGR-Tas cytosine base editing system and its applications. Background Technology

[0002] Genome editing technology enables targeted modification at the genome level and is widely used in basic research and biotechnology applications. Single-base mutations (SNVs) are an important molecular basis for hereditary diseases and complex trait differences. Base editing technology, by achieving single-base substitution without introducing double-strand breaks (DSBs), reduces the risk of byproducts such as insertions / deletions caused by DSBs, and therefore has important value in point mutation construction, functional screening, and genetic variation research. Cytosine base editors (CBEs) are typically composed of a fusion of cytosine deaminase and RNA-guided nuclease (or its catalytically modified form), and can mediate C→T conversion within a certain editing window.

[0003] However, most existing RNA-guided nucleases rely on specific protospacer-adjacent motifs (PAMs) for targeted binding and cleavage / cutting. The PAM requirement, along with editing window constraints, limits the range of editable targets, making it difficult to effectively cover many target sites. Secondly, efficient in vivo delivery remains a major obstacle to clinical translation. While widely used effectors such as SpCas9 exhibit potent activity and broad compatibility with base editing enzymes (BEs), their large molecular weight (approximately 1,368 amino acids) complicates the packaging and delivery of clinically relevant vectors. Furthermore, in mammalian cell environments, the compatibility of different deaminases with nuclease frameworks varies significantly, potentially leading to insufficient editing efficiency, suboptimal window conditions, decreased product purity, or increased indel byproducts. Therefore, a trade-off often exists between "improving editing efficiency" and "maintaining a good product profile and specificity."

[0004] Recent studies have reported a class of TIGR-Tas systems, whose effector protein TasR possesses programmable nucleic acid recognition capabilities and is considered to have PAM-independent potential, thus offering a possible solution to overcome PAM limitations. Moreover, this effector protein TasR consists of only about 300 amino acids, providing a feasible approach for its efficient in vivo delivery. Meanwhile, the catalytic behavior of nucleases in eukaryotic cells can be regulated through the design of guide RNA structures or sequences. For example, by introducing specific site mismatches into the spacer region of the guide RNA, the catalytic activity of the nuclease can be biased towards single-strand cleavage (nickase-like), thereby synergizing with deamination reactions and reducing double-strand break-related byproducts.

[0005] Despite this, constructing an efficient, controllable, and mammalian cell-compatible cytosine base editing system targeting the TIGR-Tas framework still faces several key challenges. For example, it requires screening cytosine deaminases and linker conformations compatible with the TIGR-Tas nuclease framework; establishing stable and reproducible mismatch programming strategies to induce nickase-like states and achieve windowed C→T transitions; and addressing issues such as low editing efficiency and significant differences between sites at most endogenous loci. Furthermore, whether enhanced editor activity will lead to off-target editing, genome-wide variation burden, or increased transcriptome-specific RNA editing requires systematic evaluation. Therefore, there is an urgent need to provide a PAM-independent TIGR-Tas cytosine base editor and its optimized variants that can expand target coverage, improve endogenous site editing efficiency, and maintain low indel byproducts and good genome and transcriptome specificity to meet practical application needs. Summary of the Invention

[0006] To address the problems in the prior art, this invention provides a TIGR-Tas cytosine base editing system and its applications, as well as optimized variants, guide RNAs, polynucleotides, constructs, expression systems, and related vectors. The aim is to provide a PAM-independent, highly efficient, controllable, and low-byproduct TIGR-Tas cytosine base editor and its optimized variants.

[0007] Technical solution: The first aspect of this invention proposes a TIGR-Tas cytosine base editing system, the base editing system comprising: an editing protein, the editing protein being a fusion protein of cytosine deaminase and parTasR; and tigRNA, the tigRNA being specifically bound to the parTasR portion of the editing protein and guiding the editing protein to a target DNA sequence.

[0008] Furthermore, the fusion protein is an N-terminal fusion configuration of deaminase-XTEN-parTasR; wherein the amino acid sequence of the parTasR protein is as shown in SEQ ID NO:1, or an amino acid sequence having more than 90% sequence similarity to SEQ ID NO:1, and having the function of the amino acid shown in SEQ ID NO:1; the amino acid sequence of XTEN is as shown in SEQ ID NO:2; the amino acid sequence of the deaminase is as shown in SEQ ID NO:3 to SEQ ID NO:16, or an amino acid sequence having more than 90% sequence similarity to SEQ ID NO:3 to SEQ ID NO:16, and having the function of the amino acid shown in SEQ ID NO:3 to SEQ ID NO:16; the cytosine deaminase is selected from hA3A, hA3A variants, the APOBEC family, the AICDA family, or eCDA1.

[0009] Preferably, the cytosine deaminase is selected from Homo-AICDA, Mus-AICDA, or eCDA1 of the AICDA family.

[0010] Furthermore, the tigRNA contains at least one mismatched base that is not complementary to the target DNA, and the mismatch site and number are set to enable parTasR to exhibit single-strand cleavage enzyme activity; the mismatch site is located at position 7 or 12 of the spacer.

[0011] Furthermore, the parTasR is an engineered variant containing at least one beneficial amino acid substitution, wherein the beneficial amino acid substitution is the replacement of a non-arginine residue with arginine.

[0012] Preferably, the engineered variant comprises a double mutant or a triple mutant, wherein the triple mutant (V3) variant comprises a combination of mutations of L15, T106, and E141 (e.g., L15R, T106R, E141R) and is named TIGR-Tas-CBE-V3.

[0013] A second aspect of the present invention provides an isolated polynucleotide encoding the fusion protein.

[0014] A third aspect of the present invention provides a construct containing the isolated polynucleotides described above.

[0015] A fourth aspect of the present invention provides an expression system containing the aforementioned construct, or having an exogenous polynucleotide integrated into its genome.

[0016] A fifth aspect of the present invention provides a vector combination comprising: (a) an expression vector for an editor containing an amino acid sequence encoding the fusion protein; and (b) an expression vector for tigRNA containing a nucleotide sequence encoding the tigRNA.

[0017] Furthermore, the nucleotide sequence of the expression vector of the editor is SEQ ID NO:18, or a nucleotide sequence having at least 90% sequence similarity to SEQ ID NO:18, and having the function of the nucleotide shown in SEQ ID NO:18; the nucleotide sequence of the expression vector of the tigRNA is SEQ ID NO:19, or a nucleotide sequence having at least 90% sequence similarity to SEQ ID NO:19, and having the function of the nucleotide shown in SEQ ID NO:19.

[0018] Furthermore, the expression vector of the editor and / or the expression vector of the tigRNA also includes a fluorescent labeling element; the fluorescent labeling element is used for enrichment and sorting of transfected positive cells.

[0019] The sixth aspect of this invention proposes a gene editing method for non-disease diagnosis and treatment purposes, comprising: using the system, the polynucleotide, the construct, the expression system, and the vector combination to perform C→T base editing on target DNA.

[0020] The seventh aspect of the present invention provides a method for evaluating the editing specificity of the TIGR-Tas cytosine base editing system, comprising one or more of the following steps: (1) verifying candidate off-target sites based on off-target prediction and targeted amplicon sequencing; (2) evaluating SNV / indel burden and distribution based on whole genome sequencing; and (3) evaluating transcriptome RNA SNV burden based on RNA sequencing.

[0021] The eighth aspect of this invention proposes the application of the TIGR-Tas cytosine base editing system in the preparation of reagents for editing clinically relevant pathogenic SNVs, the system having PAM-independent properties.

[0022] The ninth aspect of this invention provides a kit for evaluating cytosine base editing activity, comprising: the TIGR-Tas cytosine base editing system, and a fluorescent reporter system, wherein the fluorescent reporter system carrier is used to indicate whether C→T editing has occurred and to achieve quantitative detection.

[0023] Furthermore, the fluorescent reporter system vector is a eukaryotic expression vector containing a fluorescent protein coding sequence and its expression regulatory elements. The fluorescent protein coding sequence contains a preset mutation site, so that the fluorescent phenotype expressed by the vector is different from the wild-type fluorescent phenotype. When the mutation site is mediated by a cytosine base editor to undergo base conversion and revert to the wild-type sequence, a detectable change in the fluorescent phenotype occurs, which is used to characterize the editing activity. The nucleotide sequence of the fluorescent reporter system vector is as shown in SEQ ID NO:17, or a nucleotide sequence with more than 90% sequence similarity to SEQ ID NO:17, and has the function of the nucleotide shown in SEQ ID NO:17.

[0024] Furthermore, the fluorescent reporter system vector is used for: (1) screening cytosine deaminases compatible with the TIGR-Tas framework; and (2) screening and optimizing engineered variants of parTasR.

[0025] Beneficial effects This invention achieves windowed C→T base conversion in mammalian cells and reduces double-strand break-related byproducts by fusing cytosine deaminase with the TIGR-Tas system effector protein parTasR and combining it with mismatch programming of tigRNA to induce parTasR to be in a nickase-like catalytic state.

[0026] (1) Breaking through PAM limitations and expanding the target range: The system of this invention has PAM-independent characteristics. Under the unified in silico framework, the theoretical reachability of ClinVar pathogenic SNVs is significantly higher than that of PAM-dependent systems such as SpCas9 and LbCas12a, and it can cover almost all pathogenic SNVs evaluated.

[0027] (2) Windowing editing to reduce byproducts: by programming tigRNA mismatch, parTasR is induced to be in a nickase-like catalytic state, thereby achieving windowed C→T conversion and significantly reducing indel byproducts associated with double strand breaks.

[0028] (3) Significantly improved editing efficiency: By iteratively scanning arginine and combining mutations to engineer parTasR, the optimized variant TIGR-Tas-CBE-V3 was obtained, which showed higher windowed C→T editing efficiency in multiple endogenous loci and multiple human cell lines, while maintaining high product purity.

[0029] (4) Good genome and transcriptome specificity: After multi-level evaluation by off-target prediction and targeted amplicon sequencing, whole genome sequencing (WGS) and RNA sequencing (RNA-seq), TIGR-Tas-CBE-V3 enhances targeting activity without increasing the burden of detectable genome-wide variation and can maintain or reduce the burden of transcriptome RNA SNVs.

[0030] (5) Outstanding clinical application potential: The system of this invention can effectively edit and verify the representative pathogenic SNVs that are exclusively accessible by TIGR-Tas, providing better technical support for disease model construction and gene therapy. Attached Figure Description

[0031] Figure 1 This is a schematic diagram illustrating the working principle of the TIGR-Tas cytosine base editor. Figure 2 A schematic diagram illustrating the construction and detection principle of the BFP→GFP fluorescent reporter system; Figure 3 The results of fluorescence reporter screening of different pyrimidine deaminases under the TIGR-Tas framework (with GFP positive rate / intensity as readout) were used to screen for compatible deaminases; Figure 4 The results of amplicon sequencing screening of different deaminases paired with tigRNA-7mis or tigRNA-12mis in the endogenous locus environment are used to determine compatible deaminases and mismatch schemes. Figure 5 To compare the editing efficiency of tigRNA-12mis and tigRNA-7mis at multiple endogenous sites, and to compare deaminases that can achieve higher editing efficiency at specific sites; Figure 6 To characterize the editing effect at more endogenous sites using the optimal configuration (defined as TIGR-Tas-CBE), the range of C→T conversion efficiency at different sites is shown; Figure 7 This is a schematic diagram of the overall engineering strategy for parTasR (an iterative process of arginine scanning + combinatorial mutation). Figure 8 The results of the V1 single-point arginine scanning mutant library screening are shown, with each point representing a V1.n single-point mutant, in terms of editing output (GFP) relative to wild-type (WT). + Quantify the doubling change in readings; Figure 9 The results of V2 double mutant construction and screening are shown. Each point represents a V2.n single-point mutant, which is the editing output (GFP) of the V2.n variants combined with the top V1.n variants relative to the optimal single-point mutant V1.1. +Quantify the doubling change of the reading; Figure 10 Based on the construction and screening results of the V3 triple mutant, the optimal triple mutant combination was determined and named TIGR-Tas-CBE-V3; the V3.n variant formed by combining the top V2.n variant with a compatible single-point mutant was compared with the editing output (GFP) of the optimal double mutant V2.1. + Quantify the doubling change of the reading; Figure 11 The results show the parallel validation of TIGR-Tas-CBE (WT), intermediate variants (V1.1, V2.1), and TIGR-Tas-CBE-V3 in the BFP→GFP reporter system. Figure 12 To compare the editing efficiency of TIGR-Tas-CBE and TIGR-Tas-CBE-V3 at endogenous target sites in HEK293T cells (example of representative sites). Figure 13 A summary statistical chart of all endogenous site results is provided, showing the average efficiency improvement factor of TIGR-Tas-CBE-V3 compared to TIGR-Tas-CBE; Figure 14 The layout of the editing windows for TIGR-Tas-CBE and TIGR-Tas-CBE-V3 is shown below. Figure 15 A comparison of indel levels and product purity between TIGR-Tas-CBE and TIGR-Tas-CBE-V3, where... Figure 15 In the figure, a represents the comparison results of indel levels, and b represents the comparison results of product purity. Figure 16 To demonstrate the improved cross-cell line efficiency of TIGR-Tas-CBE-V3, and to compare the editing efficiency of the target sites, the following was shown: Figure 16 In the figure, a represents cells in HeLa cells, and b represents cells in HCT116 cells; Figure 17 The statistical results of product purity and indel level for TIGR-Tas-CBE and TIGR-Tas-CBE-V3 are shown below. Figure 17 In the table, a represents the statistical results of indel levels in HeLa cells, b represents the product purity in HeLa cells, c represents the statistical results of indel levels in HCT116 cells, and d represents the product purity in HCT116 cells. Figure 18 The editing window distribution results for TIGR-Tas-CBE and TIGR-Tas-CBE-V3 are shown below. Figure 18 In the figure, a represents cells in HeLa cells, and b represents cells in HCT116 cells; Figure 19 This is a diagram showing the target validation results of candidate off-target sites predicted by Cas-OFFinder. Figure 19 In the table, a represents the targeting and off-target efficiency of candidate site 1, b represents the targeting and off-target efficiency of candidate site 2, c represents the targeting and off-target efficiency of candidate site 3, d represents the targeting and off-target efficiency of candidate site 4, and e represents the targeting and off-target efficiency of candidate site 5. Figure 20 A statistical plot of whole-genome sequencing (WGS) variant profiles for TIGR-Tas-CBE and TIGR-Tas-CBE-V3 edited cells, where... Figure 20 In this context, 'a' represents the cumulative frequency of a specific SNV, and 'b' represents the frequency (proportion) of a specific SNV. Figure 21 RNA sequencing (RNA-seq) transcriptome RNA SNV statistics for TIGR-Tas-CBE and TIGR-Tas-CBE-V3 edited cells, including RNA SNV substitution profiles and total burden comparison; Figure 22 The statistical plot of variant types and substitution patterns in the ClinVar pathogenic SNV dataset shows that pathogenic SNVs are mainly single base substitutions and have a high proportion of transformation mutations such as C→T / G→A. Figure 23 This is a comparison graph showing the theoretical reachability of PAM-independent TIGR-Tas and PAM-dependent systems (such as SpCas9, LbCas12a, Un1Cas12f1, OgeuIscB) to ClinVar pathogenic SNVs within the unified in silico framework. Figure 24 A schematic / statistical chart for screening TIGR-Tas-exclusive pathogenic SNVs and classifying them according to application purpose; Figure 25 The figure shows the experimental validation results of TIGR-Tas-CBE-V3 against representative TIGR-Tas-exclusive pathogenic SNVs (including amplicon sequencing validation of PCCB c.683C>T and G6PC c.1039C>T sites). Detailed Implementation

[0032] To better illustrate the objectives, technical solutions, and advantages of this invention, the invention will be further described below with reference to specific embodiments. Those skilled in the art should understand that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0033] Unless otherwise specified, the experimental methods used in the examples are conventional methods; the materials and reagents used are commercially available unless otherwise specified.

[0034] General Method 1: Constructing tigRNA expression vectors targeting specific sites (1) Design the spacer sequence corresponding to the target site and synthesize a pair of complementary oligonucleotides (Forward and Reverse). In one embodiment, an overhang sequence (adapter sequence) compatible with the vector BsaI site is introduced at the 5' end of the oligonucleotide, wherein the adapter sequence can be represented by lowercase letters and the spacer sequence by uppercase letters. The tigRNA targeting spacer sequence and the annealed oligonucleotide sequence used for Golden Gate cloning are shown in Table 1.

[0035] Table 1. tigRNA targeting spacer sequence and annealed oligonucleotides used in Golden Gate cloning.

[0036]

[0037]

[0038]

[0039] (2) Anneal the forward and reverse oligonucleotides to obtain double-stranded spacer inserts. The annealing buffer may contain NaCl, Tris-HCl (pH 7.5-9.0) and ddH2O. The annealing procedure may include: incubating at 95°C for 3-10 min, then slowly cooling to 25°C, and further cooling to 4-10°C for storage.

[0040] (3) The tigRNA / sgRNA expression vector containing the U6 promoter is linearized by BsaI restriction enzyme to obtain compatible sticky ends. The linearized product can be purified and recovered.

[0041] (4) Ligate the double-stranded spacer insert from step (2) with the linearized vector from step (3). The ligation method can be Golden Gate (IIS digestion and ligation cycle reaction) or T4 DNA ligase.

[0042] (5) Transform the ligation product into competent E. coli (e.g., DH5α), select single clones for colony PCR or enzyme digestion screening, and verify the correctness of the spacer sequence by Sanger sequencing on positive clones. Amplify and extract plasmids with correct sequencing to obtain tigRNA expression plasmids targeting the target site.

[0043] General Method 2: Mammalian cell transfection, sorting, genomic DNA preparation, and target site amplification sequencing to statistically analyze editing efficiency. (1) Cell resuscitation and plate culture Resuscitate HEK293T cells and culture them routinely in culture dishes or plates containing complete culture medium at 37°C and 5% CO2. Plate the cells into 24-well or 6-well plates, ensuring that the cell confluence reaches 60%–90%, preferably 75%–85%, at the time of transfection.

[0044] (2) Instantaneous transfection When the cell fusion reaches the transfection requirement, the editor expression vector and tigRNA expression vector are mixed at a preset mass ratio, and a transfection reagent (such as EZ trans or a similar liposome transfection reagent) is added to form a transfection complex before being added to the cells.

[0045] In one embodiment, the total amount of transfected DNA can be 0.2–2.0 μg / well (24-well plate), and the mass ratio of editor plasmid to guide RNA plasmid can be 1:0.1–1:1. The medium is replaced with fresh complete medium 6–12 hours after transfection for further culture.

[0046] (3) Cell collection and sorting Cells were cultured for 24–96 hours after transfection, with the best results obtained after 72 hours. Fluorescent positive cells were then gated and sorted by flow cytometry to enrich the successfully transfected cell population.

[0047] In one embodiment, the number of fluorescently positive cells obtained through sorting can be 0.5 × 10⁻⁶. 4 ~2×10 5 indivual.

[0048] (4) Genomic DNA preparation After centrifugation, the sorted cells are collected and genomic DNA is prepared using cell lysis buffer, or lysis products that can be directly used for PCR are obtained. The lysis buffer can be a commercially available rapid genome lysis system or a lysis system containing proteinase K. The resulting lysis products can be directly used for subsequent target site PCR amplification.

[0049] (5) Target site PCR amplification Using the obtained genomic DNA (or lysis products) as a template, the target genomic region is amplified using a high-fidelity DNA polymerase (e.g., Phanta Super-Fidelity DNA Polymerase or a similar high-fidelity enzyme). The PCR reaction system and procedure can be set according to the primer Tm and the length of the amplified fragment, and typically include: denaturation, annealing, and extension cycles, with the number of cycles ranging from 20 to 35. The endogenous site targeting sequences used in this example are shown in Table 2 below.

[0050] Table 2 Target Sequences and Amplification Primers

[0051]

[0052]

[0053] (6) Sequencing and editing efficiency calculation After purification, PCR products can be subjected to Sanger sequencing or high-throughput sequencing (amplicon deep sequencing). When using Sanger sequencing, peak decomposition algorithms or commercial software can be used to quantify the mixed peaks of the edited sites to obtain the target base conversion ratio. When using amplicon deep sequencing, the C→T conversion ratio within the target window is used as an indicator of editing efficiency, and the indel ratio and product purity are statistically analyzed when necessary.

[0054] General Method 3: Candidate Off-Target Site Prediction and Targeted Amplicon Deep Sequencing Validation On-target site selection and off-target prediction In one implementation, one or more on-target sites to be evaluated (preferably 3 to 10; more preferably 5) are selected, and an off-target prediction tool is used to calculate and predict the candidate off-target sites. The off-target prediction tool may be Cas-OFFinder or equivalent software / algorithm.

[0055] In a preferred embodiment, several candidate off-target sites are screened for each on-target site for verification. The number of candidate off-target sites can be 3 to 20 per site, preferably 5 to 10 per site. For example, 8 candidate off-target sites are screened for each on-target site. The sequence information of the target site and potential off-target sites used in this case is shown in Table 3 below.

[0056] Table 3. Endogenous site target sequences and potential off-target sites

[0057]

[0058] (2) Cell processing and sample preparation Co-transfect cells with the editor (e.g., TIGR-Tas-CBE or a variant thereof) and the corresponding tigRNA. Cell transfection, culture, (optional) flow cytometry sorting, and genomic DNA preparation are performed according to "General Method 2".

[0059] In one implementation, GFP can be... + / mCherry +Double-positive populations were enriched using FACS to improve the proportion of transfected positive cells and the sequencing signal-to-noise ratio.

[0060] (3) Target site amplification and amplicon deep sequencing PCR primers were designed for the on-target site and the predicted candidate off-target site, respectively. The target region was amplified, and the amplification and sequencing were performed according to "General Method 2" (including high-fidelity PCR, product purification and amplicon deep sequencing).

[0061] (4) Editing efficiency and off-target signal determination The base conversion ratio within the target window is used as an indicator of editing efficiency. Preferably, the C→T conversion ratio within the 6th to 12th position window of the protospacer is used as the primary indicator. The editing signals of off-target sites are compared with those of on-target sites to determine whether there are detectable editing signals at candidate off-target sites.

[0062] General Method 4: Whole Genome Sequencing (WGS) to Assess Genome-Wide SNV Burden (1) Sample preparation and genomic DNA extraction In one implementation, cells are co-transfected with the editor to be evaluated and tigRNA, cultured for 48–120 h (preferably 72 h), and then collected. Optionally, FACS enrichment of transfected positive cells is performed first. High-quality genomic DNA (gDNA) is then extracted for WGS library construction.

[0063] (2) WGS library construction and sequencing The gDNA is fragmented, end-repaired, adapter-ligated, and amplified to obtain a WGS sequencing library, which is then subjected to high-throughput sequencing. The sequencing depth can be selected according to the experimental purpose, for example, 20× to 60× (preferably 30× or higher).

[0064] (3) Data processing and mutation detection After quality control and filtering of the sequencing data, the reads are aligned to a reference genome, followed by SNV variant detection. Variant detection can employ any commonly used workflow or software combination (e.g., workflows based on GATK, bcftools, Strelka, etc.), with appropriate filtering conditions set to remove low-quality and suspected false-positive variants.

[0065] (4) Genome-wide distribution, substitution patterns and functional region annotation Analyze the distribution characteristics of SNVs throughout the genome (e.g., whether there is regional clustering or hotspot enrichment) and calculate the SNV substitution spectrum (e.g., the proportion of substitution types such as C→T / G→A).

[0066] In one implementation, SNVs are annotated with genomic functional regions, their distribution in exons, introns and intergenic regions is statistically analyzed, and whether there is an enrichment trend in coding regions is assessed.

[0067] (5) Comparison of variation burden Compare editor-treated groups with control groups (e.g., wild-type editor / empty vector / untreated group) in terms of total SNVs, substitution profile, and genome distribution to assess whether editor engineering or activity enhancement causes detectable changes in genome-wide variation burden.

[0068] General Method 5: RNA Sequencing (RNA-seq) to Assess Transcriptome RNA SNV Burden and Lineage (1) Sample preparation and total RNA extraction In one implementation, cells are co-transfected with the editor to be evaluated and tigRNA, cultured for 24–120 h (preferably 72 h), and then the cells are collected and total RNA is extracted. Optionally, poly(A) enrichment or rRNA removal is performed to construct an RNA-seq library.

[0069] (2) RNA-seq library construction and sequencing RNA libraries are constructed and high-throughput sequencing is performed to obtain transcriptome sequencing data. The sequencing depth can be selected according to the experimental purpose, for example, 20M~100M reads / sample.

[0070] (3) Data processing and RNA SNV detection After quality control of the RNA-seq data, the reads are aligned to a reference genome or transcriptome and RNA SNVs are detected. Any commonly used workflow or software combination can be used for RNA SNV detection, and filtering conditions can be set to reduce false positives caused by alignment errors and technical noise.

[0071] (4) RNA SNV lineage and burden statistics The substitution profile of RNA SNVs (e.g., the proportion of types such as C→T / G→A) and the total burden (total number of RNA SNVs or burden per unit length) were statistically analyzed and compared with the control group to assess whether the editor caused an increase or decrease in transcriptome-wide nonspecific RNA editing.

[0072] The first aspect of this invention proposes the TIGR-Tas cytosine base editing system, as detailed below: Example 1: Construction of PAM-independent TIGR-Tas cytosine base editor 1.1 Construction of the TIGR-Tas cytosine base editor expression vector (deaminase-parTasR fusion) In one embodiment, a base editor is obtained by fusing cytosine deaminase with parTasR protein (amino acid sequence as shown in SEQ ID NO:1) as a nuclease backbone.

[0073] The cytosine deaminase may be selected from: hA3A, hA3A-Y130F, hA3A-Y132D, Mus-APOBEC1, APOBEC3, APOBEC3b, APOBEC3c, APOBEC3d, APOBEC3g, APOBEC3h, Catfish-AICDA, Homo-AICDA, Mus-AICDA, and eCDA1 (corresponding amino acid sequences as shown in SEQ ID NO:3 to SEQ ID NO:16). In a preferred embodiment, the deaminase is Homo-AICDA, Mus-AICDA, or eCDA1. In one embodiment, the deaminase is linked to the N-terminus of parTasR using the flexible linker peptide XTEN (amino acid sequence as shown in SEQ ID NO:2) to obtain a deaminase-XTEN-parTasR fusion protein.

[0074] In one embodiment, the above-described fusion expression framework is cloned into the mammalian expression vector backbone ancBE4-NG (Addgene#226853) to obtain the TIGR-Tas cytosine base editor expression plasmid. Cloning methods may include homologous recombination (such as Gibson Assembly or In-Fusion) or restriction endonuclease ligation. The resulting recombinant plasmid is initially screened by colony PCR and then confirmed by full-length sequencing. After confirming the absence of unexpected mutations, it is used for subsequent experiments. The nucleotide sequence of the editor expression vector is shown in SEQ ID NO:18.

[0075] 1.2 Achieving the nickase-like state of TasR through mismatch programming of tigRNA Co-transfecting cells with the TIGR-Tas cytosine base editor expression vector and the aforementioned tigRNA expression vector enables TIGR-Tas base editing. The working principle is illustrated below. Figure 1 As shown.

[0076] In one embodiment, by introducing a mismatched base at a predetermined position into the spacer of the TIGR-Tas system guide RNA (tigRNA), the cleavage behavior of TasR is changed from a double-strand cleavage tendency to a predominantly single-strand cleavage (nickase-like) tendency. This synergizes with the cytosine deamination reaction to achieve windowed C→T conversion and reduce byproducts associated with double-strand breaks (DSB). In a preferred embodiment, the mismatched position is set at position 7 or 12 of the spacer (denoted as tigRNA-7mis and tigRNA-12mis, respectively), to compare its impact on editing efficiency and product profile.

[0077] In one embodiment, the tigRNA expression vector is constructed using pGL3-U6-sgRNA-EGFP (Addgene#176544) as a backbone, and the fluorescent marker gene mCherry (Addgene#176016) is introduced to indicate transfected positive cells. The tigRNA spacer is cloned into a BsaI-linearized vector using the annealed oligonucleotide and BsaI / Golden Gate assembly method described in "General Method 1" to obtain tigRNA expression plasmids containing different mismatch designs (tigRNA-7mis or tigRNA-12mis). The nucleotide sequence of the tigRNA expression vector is shown in SEQ ID NO:19.

[0078] 1.3 Screening for compatible deaminases using a fluorescent reporter system In one embodiment, a fluorescent reporter plasmid for detecting cytosine base editing activity is constructed. The principle is illustrated below. Figure 2 As shown, the reporter plasmid is constructed based on pCMV-T7-EGFP (Addgene#133962). A single codon substitution is introduced at amino acid position 66 of the EGFP coding sequence through site-directed mutagenesis, changing its phenotype from green to blue fluorescence, resulting in the pCMV-CBE-BFP reporter plasmid. Preferably, this mutation replaces the codon TAC encoding Tyr with the codon CAC encoding His. After restoring the original EGFP sequence at this site through cytosine base editing, the green fluorescence is restored, thus enabling the detection and quantification of cytosine base editing activity.

[0079] In one embodiment, cell plating, co-transfection, culture, and flow cytometry detection for fluorescent reporter screening are performed according to "General Method 2". Preferably, co-transfection includes: (1) a fluorescent reporter plasmid; (2) a TIGR-Tas cytosine base editor expression plasmid; and (3) a tigRNA expression plasmid; and transfection is performed using EZ transfection reagent. For flow cytometry analysis, the mCherry / BFP double-positive population is used as the transfection positive gate, and the GFP positivity rate and / or GFP fluorescence intensity are used as indicators of editing activity readout.

[0080] In one specific embodiment, 14 cytosine deaminases were initially screened in combination with tigRNA-7mis. Flow cytometry results showed that Homo-AICDA, Mus-AICDA, and eCDA1 produced a higher proportion of GFP-positive cells, while most APOBEC family deaminases had low activity. Figure 3 ).

[0081] 1.4 Validation and Determination of Endogenous Locus Configuration To verify the transferability of the screening results of the report system in the endogenous locus environment, in one embodiment, multiple endogenous target sites were selected, and TIGR-Tas cytosine base editors containing 14 deaminases were constructed respectively. Each deaminase was paired with tigRNA-7mis or tigRNA-12mis respectively to obtain two sets of candidate editor combinations.

[0082] In one implementation, cell co-transfection, culture, genomic DNA preparation, target site PCR amplification, and amplicon sequencing are performed according to "General Method 2". The C→T conversion ratio within the target window is used as an indicator of editing efficiency. Figure 4 Sequencing results showed that in both tigRNA-7mis and tigRNA-12mis configurations, Homo-AICDA, Mus-AICDA, and eCDA1 all exhibited high C→T editing activity, consistent with the results of the reporter system. Further comparisons showed that, for example... Figure 5 As shown, tigRNA-12mis and tigRNA-7mis can effectively achieve targeted editing at multiple endogenous sites, and can stably improve editing efficiency. Among them, Mus-AICDA can achieve a C→T conversion rate of about 30% at a specific site.

[0083] Based on the above results, tigRNA-12mis and tigRNA-7mis have similar functions. In this embodiment, tigRNA-12mis was selected for subsequent experiments. tigRNA-7mis also has similar technical effects, which will not be repeated here. The N-terminal Mus-AICDA-XTEN-parTasR was selected to be combined with tigRNA-12mis, and this configuration was defined as TIGR-Tas-CBE. Using this TIGR-Tas-CBE, more endogenous sites were characterized. The overall C→T conversion efficiency varied with different target sites, such as... Figure 6 As shown, it is approximately 1% to 20%.

[0084] Example 2: Iterative protein engineering of parTasR to obtain the enhanced TIGR-Tas cytosine base editor TIGR-Tas-CBE-V3 2.1 Engineering Design Approach and Mutation Library Construction In one embodiment, given that the editing efficiency of TIGR-Tas-CBE at most endogenous sites in HEK293T cells is limited, the effector protein parTasR is systematically engineered to improve its base editing activity.

[0085] In one implementation, referring to the strategy that "increasing local positive charge can enhance the activity of RNA-guided nucleases in eukaryotic cells," parTasR is modified using an iterative arginine-scanning and combinatorial mutagenesis strategy: all native arginine residues in the parTasR protein are retained unchanged, and single-point arginine substitutions are performed at each of the remaining amino acid positions to construct a single-point mutant library (V1.n). The number of single-point mutants can be more than 300.

[0086] In one implementation, each parTasR variant is combined with the compatible deaminase fusion expression framework determined in Example 1, preferably the N-terminal Mus-AICDA-XTEN-parTasR structure, and is compatible with the tigRNA-12mis configuration determined in Example 1; in the specific screening stage, the BFP→GFP reporter system can also be used to improve the screening throughput.

[0087] In one implementation, to improve the editing efficiency of the basic TIGR-Tas-CBE at multiple endogenous sites, TasR is systematically engineered. Preferably, a strategy of "arginine scanning + combinatorial engineering" is adopted: such as... Figure 7As shown, the engineering modification includes iterative arginine scanning and combined mutation: retaining the natural arginine residues, replacing arginine at the remaining sites one by one to construct a single-point mutation library (V1), and after screening to obtain gain sites, further constructing double mutants (V2) and triple mutants (V3) to obtain an editor with significantly enhanced activity.

[0088] 2.2 Iterative screening and acquisition of combinatorial mutants based on a fluorescence reporter system In one embodiment, the editing activity of each parTasR variant was quantitatively assessed using a BFP→GFP fluorescent reporter system. Cell transfection, culture, and flow cytometry were performed according to "General Method 2"; and the reporter editing activity of each variant was normalized to wild-type TIGR-Tas-CBE (WT) as a control.

[0089] (1) V1 single-point mutation screening In one specific embodiment, approximately 300 V1 single-point mutants were screened. For example... Figure 8 The results showed that most single-point substitutions did not improve or even reduce activity, but about 70 variants showed improved activity relative to WT, of which 12 variants achieved an improvement of ≥2.5 times (denoted as topV1, V1.1-V1.12), which were used for subsequent combination construction.

[0090] (2) Construction and screening of V2 double mutants To evaluate the synergistic or antagonistic effects (epistasis) among single-point mutations, one implementation combines topV1 single-point mutations in pairs to construct 78 double mutants (V2.n), and uses the best single-point mutant V1.1 as a benchmark for screening. Figure 9 The results showed that 30 double mutants had activities exceeding V1.1, and further screening yielded the 10 double mutants (V2.1-V2.10) with the highest activities for subsequent construction.

[0091] (3) Construction and screening of V3 triple mutants In one implementation, combinatorial compatibility analysis is performed on the mutation sites in the high-performance double mutants, and several mutation sites with high recurrence frequency and good combinatorial behavior are selected for the construction of triple mutants. Preferably, the mutation sites include T106, A173, L15, E141, V148, and S280. Forty-five triple mutants (V3.n) are constructed by combining topV2 with the above sites, and the best double mutant V2.1 is used as a control for screening. Figure 10 The results showed that a total of 8 triple mutants were obtained, achieving an improvement of ≥1.25-fold relative to V2.1, with the L15-T106-E141 combination exhibiting the highest activity. This triple mutant was named TIGR-Tas-CBE-V3.

[0092] In a preferred embodiment, V3 contains L15R, T106R, and E141R mutations located in the RuvC and coiled-coil related regions. TIGR-Tas-CBE (WT), intermediate variants (V1.1, V2.1), and TIGR-Tas-CBE-V3 (V3) were compared in parallel, and BFP→GFP reporter editing was evaluated using fluorescence microscopy. Figure 11 The results showed that, compared to WT and intermediate variants, V3 exhibited significantly enhanced editing activity in the reporter system. Therefore, TIGR-Tas-CBE-V3 obtained through the iterative engineering strategy can serve as an enhanced base editor for subsequent endogenous locus editing and systematic characterization.

[0093] Because the TIGR-Tas framework has PAM-independent characteristics and combines mismatch programming to achieve windowed editing, the TIGR-Tas-CBE and TIGR-Tas-CBE-V3 proposed in this invention can significantly expand target coverage and improve editing efficiency, thus making them more suitable for a wide range of gene editing scenarios.

[0094] Example 3: TIGR-Tas-CBE-V3 enhances C→T base editing and maintains editing window and product profile in multiple cell lines and multiple endogenous loci. 3.1 Comparison of editing efficiency of TIGR-Tas-CBE and TIGR-Tas-CBE-V3 at multiple endogenous sites in HEK293T cells In one implementation, to systematically evaluate the endogenous locus editing performance of the optimized variant TIGR-Tas-CBE-V3, TIGR-Tas-CBE (WT) and TIGR-Tas-CBE-V3 were compared in parallel in HEK293T cells.

[0095] In one implementation, a workflow combining fluorescence enrichment and amplicon deep sequencing is used for detection: cell transfection, culture, flow cytometry sorting (FACS), genomic DNA preparation, target site PCR amplification, and amplicon sequencing are all performed according to "General Method 2". Preferably, GFP is used. + / mCherry + Double-positive cells were used as a sorting gating population to enhance signal enrichment in successfully transfected cells. The C→T conversion ratio within the target window was used as an indicator of editing efficiency.

[0096] In one specific embodiment, approximately 30 endogenous genomic target sites were selected in HEK293T cells for detection, such as... Figure 12The results showed that TIGR-Tas-CBE-V3 had significantly higher C→T editing efficiency than TIGR-Tas-CBE at most sites. At representative sites, the total C→T editing efficiency of TIGR-Tas-CBE-V3 increased from 0.28% to 12.11%, from 1.73% to 26.63%, from 4.97% to 26.48%, from 19.13% to 40.58%, and from 24.43% to 82.45%. A summary analysis of all detected sites was performed, as follows... Figure 13 The average editing efficiency of TIGR-Tas-CBE-V3 is about 35.30 times higher than that of TIGR-Tas-CBE, indicating that the combined engineering of parTasR can stably enhance the base editing activity mediated by TIGR-Tas in different genomic backgrounds.

[0097] 3.2 Editing window, product purity and indel byproduct analysis In one implementation, the editing windows, product profiles, and indel byproducts of TIGR-Tas-CBE and TIGR-Tas-CBE-V3 were characterized. Amplicon sequencing data analysis was performed according to "General Method 2".

[0098] like Figure 14 The results showed that the editing windows of both editors were mainly concentrated in positions 6–12 of the protospacer, with the peak position preferably appearing in position 7. Figure 15 Compared with TIGR-Tas-CBE, TIGR-Tas-CBE-V3 exhibits reduced indel formation and improved product purity.

[0099] 3.3 Generalization validation in HeLa and HCT116 cells In one implementation, to verify the cross-cell line applicability of TIGR-Tas-CBE-V3, several validated endogenous target sites were selected in HeLa and HCT116 cells for detection. Cell transfection, sorting, and sequencing analysis were performed according to "General Method 2".

[0100] In one specific embodiment, 12 validation target sites were tested in HeLa and HCT116 cells, such as Figure 16 The results showed that TIGR-Tas-CBE-V3 improved the efficiency of windowed C→T editing by an average of approximately 2.95 times and 3.63 times, respectively. Meanwhile, as... Figure 17 TIGR-Tas-CBE-V3 maintained high product purity and low indel levels in the aforementioned cell lines, such as Figure 18 It also maintains the feature that the editing window is focused on positions 6 to 12.

[0101] In summary, TIGR-Tas-CBE-V3 can achieve reproducible and robust windowed C→T editing enhancement on various human cell lines and multiple endogenous loci, while maintaining a superior product profile and low indel byproducts while improving efficiency. This verifies the effectiveness of arginine scan-guided combinatorial engineering strategies for optimizing the TIGR-Tas base editor.

[0102] Example 4: TIGR-Tas-CBE-V3 maintains genome and transcriptome specificity while enhancing targeted editing activity. 4.1 Candidate Off-Target Site Screening and Target Validation Based on Computational Prediction In one implementation, to assess the potential off-target effects of TIGR-Tas-CBE and its highly active variant TIGR-Tas-CBE-V3, their specificity is evaluated by combining computational prediction and target validation. Candidate off-target site prediction and amplicon deep sequencing validation are performed according to "General Method 3". Preferably, five representative on-target sites are selected, and eight candidate off-target sites are screened for validation at each on-target site; the C→T conversion ratio within the 6th to 12th window of the protospacer is used as the editing signal.

[0103] like Figure 19 The results showed that both editors exhibited significant windowed C→T enrichment at on-target sites, while the editing signal at candidate off-target sites was close to background levels. Notably, although TIGR-Tas-CBE-V3 showed significantly enhanced on-target activity, no detectable increase in editing was observed at candidate off-target sites, indicating that enhanced activity was not accompanied by a detectable decrease in specificity.

[0104] 4.2 Whole-genome sequencing (WGS) assesses genome-wide variation burden. In one implementation, to unbiasedly assess the consequences of genome-wide editing, whole-genome sequencing (WGS) is performed on the edited cells, and the genome-wide distribution and burden of single nucleotide variants (SNVs) and indels are detected. WGS data processing, variant detection, and distribution statistics are performed according to the "General Methods".

[0105] like Figure 20 The results showed that SNV substitutions included multiple categories, with C→T / G→A substitutions accounting for a significant proportion, consistent with cytosine deamination-related mutations. These results indicate that the engineering of TIGR-Tas-CBE-V3 did not lead to a detectable increase in genome-wide variation burden.

[0106] 4.3 RNA sequencing (RNA-seq) to assess RNA SNV burden at the transcriptome level In one implementation, considering that some cytosine deaminase-associated editors may induce transcriptome nonspecific RNA editing, RNA sequencing (RNA-seq) is performed on edited cells to analyze the lineage and burden of RNA SNVs. RNA-seq data processing and RNA SNV detection are performed according to "General Method 5".

[0107] like Figure 21 The results showed that transcriptome SNV replacement types included multiple categories, with C→T / G→A replacements accounting for a certain proportion. Notably, under the tested conditions, the total RNA SNV burden of TIGR-Tas-CBE-V3 was lower than that of TIGR-Tas-CBE, suggesting that enhanced DNA-targeted editing activity did not increase non-specific RNA editing, and may even reduce transcriptome-level non-specific variation under certain conditions.

[0108] In summary, through computational prediction-targeting validation using "General Method 3" and unbiased analysis using WGS and RNA-seq using "General Method 4" and "General Method 5", it is demonstrated that TIGR-Tas-CBE-V3 significantly improves targeted base editing activity while maintaining good genome and transcriptome specificity profiles.

[0109] In summary, TIGR-Tas-CBE-V3 significantly enhances on-target activity without increasing the burden of detectable genome-wide variants, and can maintain or reduce the burden of transcriptome RNA SNVs.

[0110] Example 5: Experimental validation of the widespread reachability of PAM-independent TIGR-Tas-CBE to pathogenic SNVs in ClinVar and the exclusive sites of TIGR-Tas. 5.1 Construction of ClinVar pathogenic SNV dataset and statistics of variant types In one implementation, to evaluate the targeting potential of PAM-independent TIGR-Tas-CBE for clinically relevant human variants, disease-related single nucleotide variants (SNVs) are aggregated from the ClinVar database, and preferably only the set of variants with high-confidence pathogenic annotations are included for subsequent analysis.

[0111] like Figure 22 Statistical analysis of the substitution types of the pathogenic SNVs showed that pathogenic mutations were mainly single-base substitutions, with transition mutations being the most prevalent, especially C→T and G→A substitutions, which are consistent with the transition range characteristics of the cytosine base editor.

[0112] 5.2 Comparison of the theoretical accessibility of pathogenic SNVs on different editing platforms under the unified in silico framework In one implementation, a unified in-silico reachability assessment framework is constructed, integrating the editor's PAM requirements with editing window constraints to assess whether a specific SNV is "reachable" under a given editing platform. In a preferred implementation, the determination criteria include at least: (i) editor-specific PAM constraints (if present); (ii) the target base is located within a preset editing window (e.g., the 6th–12th position window of the protospacer or a window range matching a specific editor); and (iii) evaluation can be performed in both positive and complementary strand directions (e.g., C→T and complementary strand G→A).

[0113] In one specific embodiment, the theoretical reachability of the PAM-independent TIGR-Tas framework to pathogenic SNVs is compared with that of several representative PAM-dependent systems (e.g., SpCas9, LbCas12a, Un1Cas12f1, OgeuIscB) under the same decision criteria. Figure 23 The results showed that, due to the absence of PAM constraints, the PAM-independent TIGR-Tas framework could cover almost all evaluated pathogenic SNVs, significantly outperforming the PAM-dependent system. This advantage primarily stems from the absence of PAM constraints, allowing for targeted placement at sites inaccessible under traditional PAM limitations.

[0114] 5.3 Classification and experimental validation of TIGR-Tas-exclusively reachable pathogenic SNVs In one implementation, pathogenic SNVs (TIGR-Tas-exclusive variant set) that are accessible only within the TIGR-Tas framework under the aforementioned unified criteria are further screened and classified according to their application purpose, such as... Figure 24 This includes, but is not limited to: (i) sites used for disease model construction (e.g., C→T / G→A substitution related sites); and (ii) sites used for potential correction guidance analysis (e.g., A→G / T→C substitution related sites).

[0115] In a preferred embodiment, several representative disease model loci are selected from the TIGR-Tas-exclusive variant set for experimental validation. The selected loci are edited using TIGR-Tas-CBE-V3, and cell transfection, sorting, genomic DNA preparation, target site PCR amplification, and amplicon deep sequencing are performed according to "General Method 2".

[0116] In one specific embodiment, multiple TIGR-Tas-exclusive pathogenic SNVs were validated in human cells, such as... Figure 25 Amplicon sequencing results showed that target C→T editing could be detected at all tested sites, including the PCCB gene variant c.683C>T (p.Pro228Leu) and the G6PC gene variant c.1039C>T (p.Gln347Ter).

[0117] In summary, the PAM-independent TIGR-Tas-CBE framework demonstrates significantly higher theoretical reachability to pathogenic SNVs in ClinVar under unified constraints, and can effectively validate C→T editing for pathogenic SNVs exclusively reachable by TIGR-Tas. This indicates that the system can be used to expand the targetable range of clinically relevant variants and for disease model construction applications.

[0118] Based on the TIGR-Tas cytosine base editing system proposed in the first aspect of this invention, the second aspect of this invention proposes an isolated polynucleotide encoding a fusion protein in the TIGR-Tas cytosine base editing system.

[0119] The third aspect proposes a construct containing the aforementioned polynucleotides.

[0120] The fourth aspect proposes an expression system containing the above-described constructs, or having exogenous polynucleotides integrated into its genome.

[0121] A fifth aspect of this invention provides a vector combination comprising: (a) an expression vector for an editor, comprising a nucleotide sequence encoding the aforementioned fusion protein; and (b) an expression vector for tigRNA, comprising a nucleotide sequence encoding the aforementioned tigRNA. The nucleotide sequence of the expression vector for the editor is SEQ ID NO:18, or a nucleotide sequence having at least 90% sequence similarity to SEQ ID NO:18, and having the function of the nucleotide shown in SEQ ID NO:18; the nucleotide sequence of the expression vector for the tigRNA is SEQ ID NO:19, or a nucleotide sequence having at least 90% sequence similarity to SEQ ID NO:19, and having the function of the nucleotide shown in SEQ ID NO:19. The expression vector for the editor and / or the expression vector for the tigRNA further comprises a fluorescent labeling element; the fluorescent labeling element is used for enrichment and sorting of transfected positive cells.

[0122] The sixth aspect of this invention provides a gene editing method for non-disease diagnosis and treatment purposes, comprising: using the above-described system, the above-described polynucleotide, the above-described construct, the above-described epithelial system, and the above-described vector combination to perform C→T base editing on target DNA.

[0123] The seventh aspect of the present invention proposes a method for evaluating the editing specificity of the above-mentioned TIGR-Tas cytosine base editing system, comprising one or more of the following steps: (1) verifying candidate off-target sites based on off-target prediction and targeted amplicon sequencing; (2) evaluating SNV / indel burden and distribution based on whole genome sequencing; and (3) evaluating transcriptome RNA SNV burden based on RNA sequencing.

[0124] The eighth aspect of the present invention proposes the application of the above-described TIGR-Tas cytosine base editing system in the preparation of reagents for editing clinically relevant pathogenic SNVs, said system having PAM-independent properties.

[0125] A ninth aspect of this invention provides a kit for evaluating cytosine base editing activity, comprising: the aforementioned TIGR-Tas cytosine base editing system, and a fluorescent reporter system. The fluorescent reporter system vector is used to indicate whether C→T editing has occurred and to achieve quantitative detection. The fluorescent reporter system vector is a eukaryotic expression vector containing a fluorescent protein coding sequence and its expression regulatory elements. The fluorescent protein coding sequence contains a preset mutation site, such that the fluorescent phenotype expressed by the vector differs from the wild-type fluorescent phenotype. When the mutation site undergoes a base conversion mediated by the cytosine base editor and reverts to the wild-type sequence, a detectable change in the fluorescent phenotype occurs, which is used to characterize the editing activity. The nucleotide sequence of the fluorescent reporter system vector is as shown in SEQ ID NO:17, or an amino acid sequence having more than 90% sequence similarity to SEQ ID NO:17, and having the function of the amino acid shown in SEQ ID NO:17. The fluorescent reporter system vector is used for: (1) screening cytosine deaminases compatible with the TIGR-Tas framework; and (2) screening and optimizing engineered variants of parTasR (including single-point mutants, double mutants and triple mutants).

[0126] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A TIGR-Tas cytosine base editing system, characterized in that, The base editing system includes: The editing protein is a fusion protein of cytosine deaminase and parTasR; The tigRNA specifically binds to the parTasR portion of the edited protein and guides the edited protein to the target DNA sequence.

2. The TIGR-Tas cytosine base editing system according to claim 1, characterized in that, The fusion protein is the N-terminal fusion configuration of deaminase-XTEN-parTasR; The amino acid sequence of the parTasR protein is shown in SEQ ID NO:1, or an amino acid sequence that has more than 90% sequence similarity to SEQ ID NO:1 and has the function of the amino acid shown in SEQ ID NO:

1. The amino acid sequence of XTEN is shown in SEQ ID NO:2; The amino acid sequence of the deaminase is shown in SEQ ID NO:3 to SEQ ID NO:16, or an amino acid sequence that has more than 90% sequence similarity to SEQ ID NO:3 to SEQ ID NO:16 and has the function of the amino acids shown in SEQ ID NO:3 to SEQ ID NO:16; The cytosine deaminase is selected from hA3A, hA3A variants, the APOBEC family, the AICDA family, or eCDA1.

3. The TIGR-Tas cytosine base editing system according to claim 2, characterized in that, The cytosine deaminase is selected from Homo-AICDA, Mus-AICDA, or eCDA1 of the AICDA family.

4. The TIGR-Tas cytosine base editing system according to claim 1, characterized in that, The tigRNA contains at least one mismatched base that is not complementary to the target DNA, and the mismatch site and number are set to enable parTasR to exhibit single-strand cleavage enzyme activity; the mismatch site is located at position 7 or 12 of the spacer.

5. The TIGR-Tas cytosine base editing system according to claim 1, characterized in that, The parTasR is an engineered variant containing at least one beneficial amino acid substitution, wherein the beneficial amino acid substitution is the replacement of a non-arginine residue with arginine.

6. The TIGR-Tas cytosine base editing system according to claim 5, characterized in that, The engineered variants include double or triple mutants, with the triple mutant variant containing a combination of mutations in L15, T106, and E141, and named TIGR-Tas-CBE-V3.

7. An isolated polynucleotide, characterized in that, The polynucleotide encodes the fusion protein according to any one of claims 1-6.

8. A construct, characterized in that, The construct contains the isolated polynucleotide as described in claim 7.

9. An expression system, characterized in that, The expression system contains the construct of claim 8, or has an exogenous polynucleotide of claim 8 integrated into its genome.

10. A carrier assembly, characterized in that, Include: (a) An expression vector for an editor, comprising an amino acid sequence encoding the fusion protein of any one of claims 1-2; (b) An expression vector for tigRNA containing a nucleotide sequence encoding the tigRNA.

11. The carrier assembly according to claim 10, characterized in that, The nucleotide sequence of the expression vector of the editor is SEQ ID NO:18, or a nucleotide sequence that has at least 90% sequence similarity to SEQ ID NO:18, and has the function of the nucleotide shown in SEQ ID NO:18; The nucleotide sequence of the expression vector of the tigRNA is SEQ ID NO:19, or a nucleotide sequence that has at least 90% sequence similarity to SEQ ID NO:19, and has the function of the nucleotide shown in SEQ ID NO:

19.

12. The carrier assembly according to claim 10, characterized in that, The expression vector of the editor and / or the expression vector of the tigRNA further comprises a fluorescent labeling element; the fluorescent labeling element is used for enrichment and sorting of transfected positive cells.

13. A gene editing method for purposes other than disease diagnosis and treatment, characterized in that, include: C→T base editing of target DNA was performed using the system of any one of claims 1-6, the polynucleotide of claim 7, the construct of claim 8, the expression system of claim 9, or the vector combination of any one of claims 10-12.

14. A method for evaluating the editing specificity of the TIGR-Tas cytosine base editing system according to any one of claims 1-6, characterized in that, Includes one or more of the following steps: (1) Validation of candidate off-target sites based on off-target prediction and targeted amplicon sequencing; (2) Assessment of SNV / indel burden and distribution based on whole-genome sequencing; (3) Assessment of transcriptome RNA SNV burden based on RNA sequencing.

15. The use of the TIGR-Tas cytosine base editing system according to any one of claims 1-6 in the preparation of reagents for editing clinically relevant pathogenic SNVs, characterized in that, The system has PAM-independent characteristics.

16. A kit for evaluating cytosine base editing activity, characterized in that, It includes: the TIGR-Tas cytosine base editing system as described in claims 1-6, and a fluorescent reporter system, wherein the fluorescent reporter system carrier is used to indicate whether C→T editing has occurred and to achieve quantitative detection.

17. The kit according to claim 16, characterized in that, The fluorescent reporter system vector is a eukaryotic expression vector containing a fluorescent protein coding sequence and its expression regulatory elements. The fluorescent protein coding sequence contains a preset mutation site, so that the fluorescent phenotype expressed by the vector is different from the wild-type fluorescent phenotype. When the mutation site is mediated by a cytosine base editor to undergo base conversion and revert to the wild-type sequence, a detectable change in the fluorescent phenotype occurs, which is used to characterize the editing activity. The nucleotide sequence of the fluorescent reporter system vector is as shown in SEQ ID NO:17, or a nucleotide sequence with more than 90% sequence similarity to SEQ ID NO:17, and has the function of the nucleotide shown in SEQ ID NO:

17.

18. The kit according to claim 17, characterized in that, The fluorescent reporter system carrier is used for: (1) Screening for cytosine deaminases compatible with the TIGR-Tas framework; (2) Screening and optimizing the engineered variants of parTasR.