Novel tandem repeat-containing polypeptides that bind to DNA

Novel TALE-like polypeptides with enhanced affinity and specificity address the limitations of current TALENs in genome editing, achieving more efficient targeted genome modification.

WO2025108365A1PCT designated stage expired Publication Date: 2025-05-30INST OF ZOOLOGY CHINESE ACAD OF SCI +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/133476
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-21
Filing Date
2024-11-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Current genome editing technologies, such as TALENs, face challenges in achieving high specificity and binding affinity, limiting their efficiency in targeted genome modification.

Method used

Development of novel TALE-like polypeptides with a higher affinity and smaller size, utilizing specific amino acid sequences in tandem repeats to enhance DNA binding specificity.

Benefits of technology

The novel TALE-like polypeptides demonstrate improved DNA binding affinity and specificity, enabling more efficient targeted genome modification compared to conventional TALENs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2024133476-FTAPPB-I100001
    Figure PCTCN2024133476-FTAPPB-I100001
  • Figure PCTCN2024133476-FTAPPB-I100002
    Figure PCTCN2024133476-FTAPPB-I100002
  • Figure PCTCN2024133476-FTAPPB-I100003
    Figure PCTCN2024133476-FTAPPB-I100003
Patent Text Reader

Abstract

Provided herein are novel TRPs comprising a N-terminal region, two or more tandem repeats and a C-terminal region. Provided herein are also a fusion polypeptide comprising the TRP fused to a fusion partner, or a gene editing system comprising the fusion polypeptide, as well as a method for gene editing with the fusion polypeptide.
Need to check novelty before this filing date? Find Prior Art

Description

Novel Tandem Repeat-containing Polypeptides that bind to DNATechnical Field

[0001] The present invention relates to molecular biology. In particular, the present invention provides novel TR-containing proteins (TRPs) and the use thereof in molecular biology, such as gene editing and artificial transcription factor (TF) library.Background

[0002] The modification of genome at a predetermined site has been enabled by employing site-specific systems. Genome-editing techniques such as meganucleases, designer zinc finger nucleases (ZFNs) , or transcription activator-like effector nucleases (TALENs) , and CRISPR / Cas systems are available for producing targeted genome modification.

[0003] The core of gene editing technologies lies a DNA-binding protein coupled with functional domains such as nucleases, deaminases, or reverse transcriptase. Among the four technologies, three were grounded in protein-DNA recognition, with ZNF and TALE belonging to the category of tandem repeat (TR) proteins. Among the large number of TRs that exist in nature, only a very small portion have been experimentally studied, and even fewer have been developed into biotechnological tools.

[0004] Each of these technologies has its own set of components and mechanisms. Although the CRISPR / Cas systems became popular recently, each of the technologies has its advantages and disadvantages, making them suitable for different applications. Therefore, there remains a need of developing improved TALENs, such as seeking novel TALE-like polypeptides with improved specificity and / or binding affinity.Summary of the Invention

[0005] To meet the need above, the inventors identified a number of TR-containing proteins (TRPs) that binds to DNA with sequence specificity, among which some TRPs from STAR family exhibit a TALE-like structure, and thus, referred to a TALE-like polypeptide hereinafter. The identified TALE-like polypeptides of the present disclosure provide a higher affinity than the canonical TALE polypeptide with a smaller size and less TRs.

[0006] In a first aspect, the present disclosure provides a TRP, which can bind to DNA with sequence specificity. In some embodiment, the TRP is a programmed / programmable TALE-like polypeptide comprising a N-terminal region, two or more tandem repeats and a C-terminal region, wherein each of the tandem repeats comprises an amino acid sequence selected from the group consisting of Formulae I to VII and XXII to XXIV

[0007] FX1NDNLVKVAAX2X3GX4X5X6ALQX7LLDX8GPALRQAG      (I)

[0008] where

[0009] X1 is G or S, X2X3 are repeat variable di-residue (RVD) , X4 is G or S, X5 is A or Q, X6 is H or Q, X7 is A or T, and X8 is K or R;

[0010] F X1H X2QIV X3IAS X4X5GGSQAL X6X7VL X8X9X10A X11L X12X13X14G     (II)

[0011] where

[0012] X1 is T or K, X2 is Q, E or R, X3 is A or G, X4X5 are RVD, X6 is N or D, X7 is T or K, X8 is A or V, X9 is T, R or K, X10 is H or Y, X11 is A, P or Q, X12 is T or R, X13 is A, D or T, and X14 is A or V.

[0013] GGREQVIKIAA X1X2GGKQALQALLDKSPALRQAG     (III)

[0014] where X1X2 are RVD,

[0015] FSNDNLVRIGG X1X2GAKKTLDTLLQVYPKLTQGG     (IV)

[0016] where X1X2 are RVD,

[0017] ILSGQETNRIK X1X2GGAKALETLSEKAEALHRAG     (V)

[0018] where X1X2 are RVD,

[0019] FSKQEAVAIAS X1X2GGSQALNTVLATHATLTAAG     (VI)

[0020] where X1X2 are RVD,

[0021] FAVEDVSAIAA X1X2GGAPALQAVVDHLELLMTRH     (VII)

[0022] where X1X2 are RVD,

[0023] FX1X2DNLX3KVAAX4X5GGX6QALLDKX7PX8LRX9AG      (XXII)

[0024] where X1 is G or S, X2 is N or P, X3 is V or I, X4X5 are RVD, X6 is A or Q, X7 is G or S, X8 is A or T, and X9 is Q or N,

[0025] GSREQVIKIAAX1X2GGQQALQALLDKGPALRNAG     (XXIII)

[0026] where X1X2 are RVD,

[0027] and

[0028] FSNDNLVRIGGX1X2GAKKTLDTLLQVYPQLTQGG     (XXIV)

[0029] where X1X2 are RVD.

[0030] In some embodiments, the TRP is a DNA binding polypeptide comprising an amino acid sequence selected from the group consisting of SEQ ID NOs: 3-6 and 9-22.

[0031] In a second aspect, the present disclosure provides a fusion polypeptide comprising a TRP, which can bind to DNA with sequence specificity. In some embodiment, the TRP is the programmed / programmable TALE-like or DNA binding polypeptide of the present disclosure fused to a fusion partner.

[0032] In a third aspect, the present disclosure provides a recombinant gene editing system comprising a fusion polypeptide comprising the TRP of the present disclosure fused to a fusion partner, or a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide, wherein the fusion partner is a polypeptide that provides an activity of nuclease.

[0033] In a fourth aspect, the present disclosure provides a composition comprising a fusion polypeptide comprising the TRP of the present disclosure fused to a fusion partner, or a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide.

[0034] In a fifth aspect, the present disclosure provides a method of introducing a double-strand break into a polynucleotide of interest comprising a step of contacting the polynucleotide with a recombinant gene editing system comprising a fusion polypeptide comprising the TRP of the present disclosure fused to a fusion partner, or a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide, wherein the fusion partner is a polypeptide that provides an activity of nuclease.

[0035] The present disclosure also provides a method of modifying a genomic sequence in a cell comprising a step of introducing into the cell a recombinant gene editing system comprising a fusion polypeptide comprising the TRP of the present disclosure fused to a fusion partner, or a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide, wherein the fusion partner is a polypeptide that provides an activity of nuclease.

[0036] In a sixth aspect, the present disclosure provides a randomized library of artificial transcription factors (TFs) , comprising a plurality of cells, each of which harboring a vector comprising a nucleotide sequence encoding a polypeptide of the present disclosure, wherein the tandem repeats are randomized between cells.

[0037] In a seventh aspect, the present disclosure provides a method for screening DNA-binding TRPs, comprising

[0038] i) retrieving proteins in TR families from a database;

[0039] ii) predicting TRs in the proteins and screening the TRPs; and

[0040] iii) establishing a model and analyzing the screened TRPs with the model to obtain TRPs that are predicted to be capable of binding to DNA.Brief Description of the Drawings

[0041] Fig. 1 shows the identification and characterization of tandem repeat proteins.

[0042] A) Computational pipeline for the identification of both known and novel TRs within the nonredundant protein database. B) Schematic illustration of the TR-related terms and parameters. C) Distribution of unit numbers and periods for all TRs with periods ≤ 80 and repeat units ≤ 80. Dashed box indicates the proportion of short TRs with periods ≤ 10 and repeat units ≤ 10. The upper box displays the taxonomic distribution for different periods. The arrow indicates the location of ZNFs. D) Distribution of known and novel TR counts at both the protein and cluster levels. The counts of proteins and clusters were converted to log10 scale. TRs with different period lengths are color coded. E) Cluster size distribution of known and novel TRs. Each color represents a distinct cluster size range. The absolute numbers of clusters were illustrated on the pie plot. F) The left panel illustrates the distribution of cluster sizes and the counts of genus-level assemblies. Specifically, for TRs within each cluster size range, we retrieved the count of genus-level assemblies accessible at NCBI for the originating species of each protein and then depicted the corresponding distribution. The right panel represents the Pearson correlation analysis between cluster size and the median statistics for the count of genus-level assemblies.

[0043] Fig. 2 shows the DNA binding prediction model and strategy of candidate prioritization.

[0044] A) The generation of training datasets and PLM-DBPPred architecture integrates three NLP models (ProteinBERT, ProtTrans, and ESM) to predict DBPs. B) Comparing PLM-DBPPred with other DBP classifiers via ROC-AUC analysis utilizing the PDB600 test dataset. AUC values are indicated for each classifier. C) Evaluation results of various DBP prediction tools. Each column was individually normalized, and the performance is depicted using a color key. D) Strategies to select TR clusters and candidates. The Venn plot illustrates the enriched cluster list generated by different functional annotation methods, wherein DBP, DRD and DBGO represent DNA binding proteins predicted by PLM-DBPPred, DNA-related domains and DNA binding-related GO annotations, respectively.

[0045] Fig. 3 shows the experimental screening and validation design.

[0046] A) Experimental screening pipeline. For the in vivo B1H screening, all 100 candidate genes were cloned and inserted into B1H protein expression vectors and subsequently screened. For the in vitro screening, all 100 candidate genes were initially assessed for their expression levels in E. coli. Subsequently, proteins with high expression were purified. The purified proteins were then subjected to a BLI-based screen to test their DNA binding activity. Candidates showing DNA-binding activities were further subjected to SELEX, the enriched libraries were sequenced, and the enriched binding motifs were obtained. The numbers in each step indicate the corresponding candidates identified and selected during the screening process. The Venn plot illustrates the positive candidates identified by different screening methods. The light blue and dark blue colors represent positive candidates with DNA binding activity and specific DNA binding activity, respectively. B) Positive candidates identified by the B1H platform. For each candidate, we provide information including the protein accession (acustomized name was given to each, shown in parenthesis) , resource species, repeat weblogo and enriched binding motif. C) Positive candidates identified by the SELEX platform. D) Experimental design for the validation and characterization of positive candidates. We established four independent validation methods to verify protein binding activity toward the enriched motifs. These methods included one in vitro platform (EMSA) , two GFP-based platforms in E. coli, and the CUT&Tag platform in 293T cells (Material and Methods) . Proteins that were validated in two or more independent assays were deemed true positives. The analysis of positive candidates was then extended to three further analyses: sequence and structure feature analysis, family-level characterization and native biological function prediction.

[0047] Fig. 4 shows the characterization of the STAR family.

[0048] A) EMSA validation results for PqSTAR1 and AspSTAR1. The protein concentration was varied, while the probe content was kept constant at 40 nM. The abbreviation “C” denotes the competitor probe, which was used at a concentration 100-fold higher than that of the specific probe. B) GFP activation results for PqSTAR1 and AspSTAR1. MutODD is a non-DNA binding protein that served as a negative control. “P” and “R” represent protein and reporter plasmids, respectively. Two-sided Student’s t-test, n = 3 biologically independent samples. Data are shown as mean ± s.d. ***, p value < 0.001. C) BLI binding results for PqSTAR1 and AspSTAR1. The PqSTAR1 protein concentrations were 0.25 nM, 0.5 nM, 1 nM, 2 nM, and 4 nM, and the AspSTAR1 protein concentrations were 3.125 nM, 6.25 nM, 12.5 nM, 25 nM, and 50 nM. D) EM micrographs of negatively stained PqSTAR1-DNA complexes, AspSTAR1-DNA complex and their representative 2D classification averages. Upper panel scale bar is 50 nm, bottom panel scale bar is 10 nm. E) Protein architecture of PqSTAR1, AspSTAR1 and AvrBs3. TS, T3SS. NLS, nuclear localization signal. TAD, transcriptional activation domain. The gene encoding the predicted PqSTAR1 lacks a start codon; we therefore placed a “? ” at the N-terminus of STAR1. F) Sequence comparison of PqSTAR1, AspSTAR1 and AvrBs3. G) Protein tertiary structure of PqSTAR1, AspSTAR1 and AvrBs3. Structures of PqSTAR1 and AspSTAR1 were predicted by AlphaFold2. The structure of AvrBs3 was derived from the PDB database (2YPF) . N and C represent N-and C-terminal ends. The repeat region and RVDs are colored blue and red, respectively. The expanded dashed box illustrates the structure of a single repeat unit. H) Multiple sequence alignment and phylogenetic tree of PqSTAR homologs and AspSTAR homologs. The phylogenetic tree was constructed using full-length protein sequences. In each aligned column, the degree of conservation is represented by the undertone of the amino acid. The bootstrap confidence scores are represented by circle sizes. Squares with different colors in the middle represent different taxa. The color and identity coordinates are shown on the right. I) Distribution of unit numbers of canonical TALEs, PqSTARs and AspSTARs. canonical TALEs represent TALEs from Xanthomonas, PqSTARs represent STARs from Pseudomonas, and AspSTARs represent STARs from Apophysomyces. Statistical analysis was by one-way ANOVA analysis. *, p value < 0.05; ****, p value < 0.0001. J) Reprogramming DNA binding specificity of PqSTAR1 and AspSTAR1. For PqSTAR1, WT, 23m and 67m represent the wild-type protein, a variant with modified RVD2 and RVD3 and a variant with modified RVD6 and RVD7, respectively. For AspSTAR1, WT, 3m and 678m represent the wild-type protein, a variant with modified RVD3, and a variant with modified RVD6, RVD7 and RVD8, respectively. RVD alterations are highlighted in red. The nucleic acid bases under the corresponding RVDs are derived following the binding code of canonical TALEs. In the EMSA assay, probe 1 corresponds to the target sequence of the wild-type protein, while probes 2 and 3 represent the target sequences of the two variants. K) GFP activation results for artificial STARs. MutODD is a non-DNA binding protein serving as a negative control. “P” and “R” represent protein and reporter plasmids, respectively. Two-sided Student’s t-test, n = 3 biologically independent samples. Data are shown as mean ± s.d. **, p value < 0.01; ***, p value < 0.001; ****, p value < 0.0001. L) CUT&Tag results of PqSTAR1 and AspSTAR1 in 293T cells. Two biologically independent samples were employed.

[0049] Fig. 5 shows the characterization of STAR-based transcriptional regulator.

[0050] A) Illustration of STAR-based transcriptional regulator. B) Reported TF binding motif and RVD design of STARs. C) Principal component analysis (PCA) for RNA-seq samples. PCA was performed using rlog-transformed count data from DESeq2. D) Significantly up-regulated (red) and down-regulated (blue) transcripts after STAR-based ATFs transduction (p value < 0.05 and a log2foldchange>|1|) . E) Motifs enriched in the upregulated gene promoter regions. F) Gene set enrichment analysis (GSEA) plots of 2 previously characterized NF-κB regulated gene sets. Gene set 1 and 2 derived from Cormier et al. 2023 (NF-κB signaling activation and roles in thyroid cancers: implication of MAP3K14 / NIK. Oncogenesis, 12, 55) and the TFlink database (Liska et al., 2022, TFLink: an integrated gateway to access transcription factor–target gene interactions for multiple species. Database, 2022, baac083) , respectively. NES, normalized enrichment score. G) Heatmap of normalized reads count for genes listed in f (combination of gene set 1 and 2) . Statistical significance analysis was employed by the Wilcoxon signed-rank test. *, p value < 0.05, **, p value < 0.01. H) Gene set enrichment analysis (GSEA) plots of 2 previously characterized SMAD4 regulated gene sets. Gene set 1 and 2 resourced from the MSigDB database (SMAD4_Q6) (Liberzon et al., 2015, The molecular signatures database hallmark gene set collection. Cell systems, 1, 417-425) and the TFlink database, respectively. I) Heatmap of normalized reads count for genes listed in h (combination of gene set 1 and 2) . Statistical significance analysis was employed by the Wilcoxon signed-rank test. **, p value < 0.01; ***, p value < 0.001.

[0051] Fig. 6 shows the characterization of the MOON family.

[0052] A) EMSA results for SpMOON1. The GC-rich sequence served as a control. The protein concentration was varied, while the probe content was kept constant at 40 nM. B) GFP activation validation for SpMOON1. MutODD is a non-DNA binding protein that served as a negative control. “P” and “R” represent protein and reporter plasmids, respectively. Two-sided Student’s t-test, n = 3 biologically independent samples. Data are shown as mean ± s.d. *, p value < 0.05; **, p value < 0.01. C) BLI results for SpMOON1. Protein concentrations were set as follows: 9.8 nM, 14.8 nM, 22.2 nM, 33.3 nM, 50 nM. D) An EM micrograph of negatively stained SpMOON1-DNA complex and its representative 2D classification averages. Upper panel scale bar is 50 nm, bottom panel scale bar is 10 nm. E) Protein architecture of SpMOON1, KI67_HUMAN and KI67_MOUSE protein. The architecture of KI67_HUMAN and KI67_MOUSE was derived from the literature (Sobecki et al., 2016, The cell proliferation antigen Ki-67 organises heterochromatin. elife, 5, e13722) . F) Sequence comparison of SpMOON1, KI67_HUMAN and KI67_MOUSE. G) Predicted protein tertiary structure of SpMOON1. The FHA, PP1_bind domain and repeat region are colored yellow, orange and blue, respectively. The expanded dashed box illustrates the structure of a single repeat unit. H) Domain architecture, repeat and secondary structure weblogos of the MOON family. The functional domains, FHA and PP1_bind, are illustrated. The C-terminal region, indicated by the red rectangle, contains repeat units. The unit number within the MOON family ranges from 6 to 36. I) Multiple sequence alignment and phylogenetic tree of the proteins in the MOON family. The phylogenetic tree was constructed using full-length protein sequences. The names of experimentally tested homologs of SpMOON1 are labeled. In each aligned column, the degree of conservation is represented by the undertone of the amino acid. The bootstrap confidence scores are represented by circle sizes. Squares with different colors in the middle represent different taxa. The color and identity coordinates are shown on the right. J) Species, repeat logos and enriched motifs of SpMOON2. K) Assessment of SpMOON2 binding to AT / GC-rich probes using BLI assay. Protein concentrations were set as follows: 9.8 nM, 14.8 nM, 22.2 nM, 33.3 nM, 50 nM.

[0053] Fig. 7 shows the characterization of the pTERF family.

[0054] A) EMSA results for pTERF1. The protein concentration was varied, while the probe content was kept constant at 40 nM. B) GFP activation validation for pTERF1. MutODD is a non-DNA binding protein that served as a negative control. “P” and “R” represent protein and reporter plasmids, respectively. Two-sided Student’s t-test, n = 3 biologically independent samples. Data are shown as mean ± s.d. ***, p value < 0.001. C) BLI results for pTERF1. Protein concentrations were set as follows: 1.48 nM, 2.22 nM, 3.33 nM, 5 nM, 7.5 nM. D) An EM micrograph of negatively stained pTERF1-DNA complex and its representative 2D classification averages. Upper panel scale bar is 50 nm, bottom panel scale bar is 10 nm. E) Protein architecture of the pTERF1, MTEF1_HUMAN and MTTF_DROME proteins. The architecture of MTEF1_HUMAN and MTTF_DROME was derived from the literature (Roberti et al., 2009, The MTERF family proteins: mitochondrial transcription regulators and beyond. Biochimica et Biophysica Acta (BBA) -Bioenergetics, 1787, 303-311) . F) Sequence comparison of pTERF1, MTEF1_HUMAN and MTTF_DROME. G) Protein tertiary structure of pTERF1 and MTEF1_HUMAN (PDB: 3MVA) . The expanded box illustrates the structure of a single repeat unit or an mTERF motif. H) Domain architecture, repeat and secondary structure weblogos of the pTERF family. Red rectangles represent the repeat units within pTERF proteins, ranging from 3 to 11. I) Multiple sequence alignment and phylogenetic tree of the pTERF family. The names of experimentally tested homologs of pTERF1 are labeled. In each aligned column, the degree of conservation is represented by the undertone of the amino acid. The bootstrap confidence scores are represented by circle sizes. The color and identity coordinates are shown on the right. J) Species, repeat logos and enriched motifs of pTERF2. K) EMSA results for pTERF2. The protein concentration was varied, while the probe content was kept constant at 40 nM. L) BLI results for pTERF2. Protein concentrations were set as follows: 0.625 nM, 1.25 nM, 2.5 nM, 5 nM, 10 nM.

[0055] Fig. 8 shows the characterization of PqSTAR4.

[0056] A) Species, repeat logos and enriched motifs of PqSTAR4. B) GFP activation validation for PqSTAR4. MutODD is a non-DNA binding protein that served as a negative control. “P” and “R” represent protein and reporter plasmids, respectively. Two-sided Student’s t-test, n = 3 biologically independent samples. Data are shown as mean ± s.d. ***, p value < 0.001.

[0057] Fig. 9 shows the characterization of the TRP of SEQ ID NO: 10.

[0058] A) Species, repeat logos and enriched motifs of SEQ ID NO: 10 A0A662FLR2. B) GFP activation validation for SEQ ID NO: 10 A0A662FLR2. MutODD is a non-DNA binding protein that served as a negative control. “P” and “R” represent protein and reporter plasmids, respectively. Two-sided Student’s t-test, n = 3 biologically independent samples. Data are shown as mean ± s.d. *, p value < 0.05. C) EMSA and BLI results for SEQ ID NO: 10 A0A662FLR2. For EMSA assay, the protein concentration was varied, while the probe content was kept constant at 40 nM. For BLI assay, protein concentrations were set as follows: 62.5 nM, 125 nM, 250 nM, 500 nM, 1000 nM.Detailed Description of the Invention

[0059] 1. Definitions

[0060] The present application refers to tandem repeat containing proteins / polypeptides (TRPs) that binds to DNA, preferably with a sequence specificity. As used herein, the terms “tandem repeat containing protein / polypeptide” and “TRP” refers to a polypeptide that comprises an N terminal region, a C terminal region, and a middle region consisting of tandem repeats (TRs) . As an Example, canonical TALE (such as those derived from AvrBs3 (UniProt No. P14727) ) is one of the known DNA-binding TRPs, in which each of the TRs comprises repeat variable di-residue (RVD) recognizing and binding to a nucleotide. The other residues in the repeat do not participate in the binding to DNA, and can be considered as non-RVD scaffolds. For canonical TALE, the RVDs are selected from the group consisting of HD (recognizing C) , NN (recognizing A and G) , NG (recognizing T) , NS / NI (recognizing A) , NK (recognizing G) , KS / KG (recognizing T) , HI (recognizing A and G) .

[0061] The term “TALE-like polypeptide” refers to a polypeptide having a similar structure to the known TALEs, i.e., comprising TRs, such as those identified in the present disclosure.

[0062] In the context of the present application, a TR in a programmed / programmable TALE-like polypeptide “derived from” a native TR in a TRP or “derived from” a TRP refers to a TR obtained by the modification to the RVD, preferably without the modification to the other residues in the native TR.

[0063] The terms “polypeptide” and “protein” are used interchangeably herein and refer to a polymer of amino acids and includes full-length proteins and fragments thereof.

[0064] Endonucleases are enzymes that cleave the phosphodiester bond within a polynucleotide chain, and include restriction endonucleases that cleave DNA at specific sites without damaging the bases. Examples of endonucleases include, but are not limited to, restriction endonucleases, meganucleases, TAL effector nucleases (TALENs) , zinc finger nucleases, and Cas (CRISPR-associated) effector endonucleases.

[0065] As used herein, "nucleic acid" means a polynucleotide and includes a single or a double-stranded polymer of deoxyribonucleotide or ribonucleotide bases. Nucleic acids may also include fragments and modified nucleotides. Thus, the terms "polynucleotide" , "nucleic acid sequence" , "nucleotide sequence" and "nucleic acid fragment" are used interchangeably to denote a polymer of RNA and / or DNA and / or RNA-DNA that is single-or double-stranded, optionally comprising synthetic, non-natural, or altered nucleotide bases. Nucleotides (usually found in their 5'-monophosphate form) are referred to by their single letter designation as follows: "A" for adenosine or deoxyadenosine (for RNA or DNA, respectively) , "C" for cytosine or deoxycytosine, "G" for guanosine or deoxyguanosine, "U" for uridine, "T" for deoxythymidine, "R" for purines (A or G) , "Y" for pyrimidines (C or T) , "K" for G or T, "H" for A or C or T, "I" for inosine, and "N" for any nucleotide.

[0066] The term "genome" as it applies to a prokaryotic and eukaryotic cell or organism cells encompasses not only chromosomal DNA found within the nucleus, but organelle DNA found within subcellular components (e.g., mitochondria, or plastid) of the cell. The term "genome" refers to the entire complement of genetic material (genes and non-coding sequences) that is present in each cell of an organism, or virus or organelle; and / or a complete set of chromosomes inherited as a (haploid) unit from one parent.

[0067] The term "homology" refers to DNA sequences that are similar. For example, a "region of homology to a genomic region" that is found on the donor DNA is a region of DNA that has a similar sequence to a given "genomic region" in the cell or organism genome. A region of homology can be of any length that is sufficient to promote homologous recombination at the cleaved target site. For example, the region of homology can comprise 5-3000 or more bases, such as at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, or 3000 in length to enable the homologous recombination with the corresponding genomic region.

[0068] As used herein, a "genomic region" is a segment on a chromosome or organelle DNA of a cell. that is present either upstream or downstream of the target site or, alternatively, also comprises a portion (at either 5’ or 3’ end) of the target site. The genomic region can comprise can comprise 5-3000 or more bases, such as at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, or 3000 in length to enable the homologous recombination with the corresponding region of homology.

[0069] The term "homologous recombination" (HR) means the exchange of DNA fragments between two DNA molecules at the sites of homology. The frequency of homologous recombination is influenced by a number of factors. The amount of homologous recombination and the relative proportion of homologous to non-homologous recombination vary in different organisms. Generally, the length of the region of homology affects the frequency of homologous recombination events: the longer the region of homology, the greater the frequency. Further, the homologous recombination needs a certain length of the homologous region, which is species-variable. See, for example, Singer et al., (1982) Cell 31: 25-33; Shen and Huang, (1986) Genetics 112: 441-57; Watt et al., (1985) Proc. Natl. Acad. Sci. USA 82: 4768-72, Sugawara and Haber, (1992) Mol Cell Biol 12: 563-75, Rubnitz and Subramani, (1984) Mol Cell Biol 4: 2253-8; Ayares et al., (1986) Proc. Natl. Acad. Sci. USA 83: 5199-203; Liskay et al., (1987) Genetics 115: 161-7.

[0070] "Sequence identity" or "identity" in the context of nucleotide or amino acid sequences refers to the nucleotide bases or amino acid residues in two sequences that are the same when aligned for maximum correspondence over a specified comparison window.

[0071] The term "percentage of sequence identity" refers to the value determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the nucleotide or amino acid sequence in the comparison window may comprise additions or deletions (i.e., gaps) as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences. The percentage sequence identity is calculated by dividing the number of matched positions (i.e., positions at which the nucleotide bases or amino acid residues in the two sequences are identical) by the total number of positions in the window of comparison and multiplying the results by 100. For example, when aligning two sequences, if 950 positions in two sequences, which are optimally aligned in a comparison window of 1000 positions, are identical, the sequences are 95%identical to each other.

[0072] A variety of comparison methods have been designed for sequence alignments and the calculations of percent identity or similarity, including, but not limited to, the MegAlign. TM. program of the LASERGENE bioinformatics computing suite (DNASTAR Inc., Madison, Wis. ) . Within the context of this application it will be understood that where sequence analysis software is used for analysis, the results of the analysis will be based on the "default values" of the program referenced, unless otherwise specified. As used herein "default values" will mean any set of values or parameters that originally load with the software when first initialized.

[0073] "BLAST" is a searching algorithm provided by NCBI used to find regions of similarity between biological sequences. The program compares nucleotide or protein sequences to sequence databases and calculates the statistical significance of matches to identify sequences having sufficient similarity to a query sequence such that the similarity would not be predicted to have occurred randomly. BLAST reports the identified sequences and their local alignment to the query sequence. It is well understood by one skilled in the art that many levels of sequence identity are useful in identifying polypeptides from other species or modified naturally or synthetically wherein such polypeptides have the same or similar function or activity. Useful examples of percent identities include, but are not limited to, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%or 95%, or any percentage from 50%to 100%. Indeed, any amino acid identity from 50%to 100%may be useful in describing the present disclosure, such as 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%or 99%.

[0074] An "isolated" polynucleotide, polypeptide, or protein is substantially or essentially free from components that normally accompany or interact with the polynucleotide, polypeptide, or protein as found in its naturally occurring environment. Thus, an isolated polynucleotide or polypeptide or protein is substantially free of other cellular material, or culture medium when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized. Preferably, an "isolated" polynucleotide is free of sequences that naturally flank the polynucleotide (i.e., sequences located at the 5' and 3' ends of the polynucleotide) in the genomic DNA of the organism from which the polynucleotide is derived. Isolated polynucleotides and polypeptides may be purified from a cell in which they naturally occur. The methods for isolating or purifying polynucleotides or polypeptides are known to a person skilled in the art. The term also embraces recombinant or chemically synthesized polynucleotides and polypeptides.

[0075] The term "fragment" refers to a contiguous set of nucleotides or amino acids. In one embodiment, a fragment comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or greater than 20 contiguous nucleotides. In one embodiment, a fragment comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or greater than 20 contiguous amino acids. A fragment may or may not exhibit the function of a sequence sharing some percent identity over the length of said fragment.

[0076] The term "functional fragment" refers to a portion of an isolated polynucleotide or polypeptide that displays the same activity or function as the longer or full-length sequence from which it derives.

[0077] The term "gene" includes a nucleic acid fragment that expresses a functional molecule such as, but not limited to, a specific protein, including regulatory sequences preceding (5' non-coding sequences) and following (3' non-coding sequences) the coding sequence.

[0078] The term "endogenous" means a sequence or other molecule that naturally occurs in a cell or organism. An endogenous polynucleotide is normally found in the genome of a cell; that is, not heterologous.

[0079] The term "heterologous" refers to the difference between the original environment, location, or composition of a particular polynucleotide or polypeptide and its current environment, location, or composition. Non-limiting examples include differences in taxonomic derivation (e.g., a polynucleotide obtained from species A would be heterologous if inserted into the genome of species B, or of a different variety or cultivar of species A; or a polynucleotide obtained from a bacterium was introduced into a cell of a plant or an animal) , or sequence (e.g., a polynucleotide obtained from species A, isolated, modified, and re-introduced into a plant of species A) .

[0080] "Coding sequence" refers to a nucleotide sequence which codes for a specific amino acid sequence. "Regulatory sequences" refer to nucleotide sequences located upstream (5' non-coding sequences) , within, or downstream (3' non-coding sequences) of a coding sequence, which influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences include, but are not limited to, promoters, translation leader sequences, 5' untranslated sequences, 3' untranslated sequences, introns, polyadenylation signal sequences, RNA processing sites, effector binding sites, and stem-loop structures.

[0081] A "mutated gene" is a gene that has been altered through human intervention. A "mutated gene" has a sequence that differs from the sequence of the corresponding non-mutated gene by the addition, deletion, insertion or substitution of at least one nucleotide. In the present disclosure, the mutated gene comprises an alteration that results from the fusion polypeptide as disclosed herein, such as that comprises a fusion partner providing an activity of nuclease. A mutated organism is an organism comprising a mutated gene.

[0082] As used herein, a "targeted mutation" is a mutation in a gene that is made in a target sequence within the gene using any method known to a person skilled in the art, including a method involving a TALEN comprising the TALE-like polypeptide as disclosed herein fused to an endonuclease.

[0083] The term "knock-out" refers to a DNA sequence in a cell that has been rendered partially or completely inoperative, e.g., by targeting with a TALEN comprising the TALE-like polypeptide as disclosed herein fused to an endonuclease; for example, a DNA sequence prior to knock-out could have encoded an amino acid sequence, or could have had a regulatory function (e.g., promoter) .

[0084] The term "knock-in" represents the replacement or insertion of a DNA sequence at a specific site in the genome of a cell by targeting with a TALEN (for example by homologous recombination (HR) , wherein a suitable donor DNA polynucleotide is also used) . The knock-in can be a specific insertion of a heterologous nucleotide sequence that encodes an amino acid sequence or a functional RNA, or a specific insertion of a transcriptional regulatory element.

[0085] The term "domain" means a contiguous stretch of nucleotides (that can be RNA, DNA, and / or RNA-DNA-combination sequence) or contiguous or non-contiguous amino acids.

[0086] A "conserved domain" or "motif" means a set of nucleotides or amino acids conserved at specific positions along an aligned sequence of evolutionarily related genes or proteins. While nucleotides or amino acids at other positions can vary between homologous proteins, nucleotides or amino acids that are highly conserved at specific positions indicate amino acids that are essential for the structure, the stability, or the function of a polynucleotide or protein.

[0087] A "codon-optimized" nucleotide sequence is a nucleotide sequence having its frequency of codon usage designed to mimic the frequency of preferred codon usage of the host cell. An "optimized" polynucleotide comprises a nucleotide sequence that has been optimized for improved expression in a particular heterologous host cell.

[0088] A "promoter" is a nucleotide sequence involved in recognition and binding of RNA polymerase and other proteins to initiate transcription. Promoters may be derived in their entirety from a native gene, or be composed of different elements derived from different promoters found in nature, and / or comprise synthetic DNA segments. It is understood by those skilled in the art that different promoters may direct the expression of a gene in different tissues or cell types, or at different stages of development, or in response to different environmental conditions. It is further recognized that since in most cases the exact boundaries of regulatory sequences have not been completely defined, DNA fragments of some variation may have identical promoter activity.

[0089] A promoter that causes a gene to be expressed in most tissues or cell types at most times are commonly referred to as "constitutive promoter" . The term "inducible promoter" or “regulated promoter” refers to a promoter that selectively express a coding sequence or functional RNA in response to the presence of an endogenous or exogenous stimulus, for example by chemical compounds (chemical inducers) or in response to environmental, hormonal, chemical, and / or developmental signals. Inducible or regulated promoters include, for example, promoters induced or regulated by light, heat, stress, flooding or drought, salt stress, osmotic stress, phytohormones, wounding, or chemicals such as ethanol, abscisic acid (ABA) , jasmonate, salicylic acid, or safeners.

[0090] An "enhancer" is a nucleotide sequence that can stimulate promoter activity, and may be an innate element of the promoter or a heterologous element inserted to enhance the activity or tissue-specificity of a promoter.

[0091] The term "translation leader sequence" refers to a nucleotide sequence located between the promoter sequence and the coding sequence. The translation leader sequence is present in the mRNA upstream of the start codon. The translation leader sequence may affect processing of the primary transcript to mRNA, mRNA stability or translation efficiency.

[0092] The term "3' non-coding sequences" , which can be exchanged with "transcription terminator" or "termination sequences" refer to nucleotide sequences located downstream of a coding sequence and include polyadenylation signal sequences and other sequences encoding regulatory signals capable of affecting mRNA processing or gene expression. The polyadenylation signal is usually characterized by affecting the addition of polyadenylic acid tracts to the 3' end of the mRNA precursor.

[0093] The term "host" refers to an organism or cell into which a heterologous component (polynucleotide, polypeptide, other molecule, cell) has been introduced. As used herein, a "host cell" refers to an in vivo or isolated eukaryotic cell, prokaryotic cell (e.g., bacterial or archaeal cell) , or cell from a multicellular organism (e.g., a cell line) cultured as a unicellular entity, into which a heterologous polynucleotide or polypeptide has been introduced. The cell is selected from the group consisting of: an archaeal cell, a bacterial cell, a eukaryotic cell, a eukaryotic single-cell organism, a somatic cell, a germ cell, a stem cell, a plant cell, an algal cell and an animal cell, such as an invertebrate cell, a vertebrate cell, a fish cell, a frog cell, a bird cell, an insect cell, a mammalian cell, a pig cell, a cow cell, a goat cell, a sheep cell, a rodent cell, a rat cell, a mouse cell, a non-human primate cell, and a human cell. In some cases, the cell is isolated. In some cases, the cell is in vivo.

[0094] The term "recombinant" refers to an artificial combination of two otherwise separated segments of sequence, e.g., by chemical synthesis, or by genetic engineering techniques.

[0095] The terms "plasmid" and "vector" refer to a linear or circular extra chromosomal element often carrying genes that are not part of the central metabolism of the cell, and usually in the form of double-stranded DNA. Such elements may be autonomously replicating sequences, genome integrating sequences, phage, or nucleotide sequences, in linear or circular form, of a single-or double-stranded DNA or RNA, derived from any source, in which a number of nucleotide sequences have been joined or recombined into a unique construction which is capable of introducing a polynucleotide of interest into a cell.

[0096] The term "construct" , when referring to nucleic acid molecules, comprises an artificial combination of nucleic acid sequences, e.g., regulatory and coding sequences that are not all found together in nature. When the nucleic acid construct contains the control sequences required to express the coding sequence of the present invention, the term is synonymous with the term “expression cassette” . For example, a nucleic acid construct may comprise regulatory sequences and coding sequences that are derived from different sources, or regulatory sequences and coding sequences derived from the same source, but arranged in a manner different than that found in nature. Such a construct may be used by itself or may be used in conjunction with a vector. If a vector is used, then the choice of vector is dependent upon the method that will be used to introduce the vector into the host cells as is well known to those skilled in the art. The vector for expressing a coding sequence (e.g., comprising an expression construct) is referred to as “expression vector” .

[0097] The term "expression" , as used herein, refers to the production of a functional end-product (e.g., an mRNA, guide RNA, or a protein) in either precursor or mature form.

[0098] As used herein, an "effector" or "effector protein" is a protein that encompasses an activity including recognizing, binding to, and / or cleaving or nicking a polynucleotide target. An effector, or effector protein, may also be an endonuclease, such as the TALE-like polypeptide or the fusion polypeptide of the present disclosure.

[0099] A "functional fragment" of a TALE-like polypeptide refers to a portion of the TALE-like polypeptide of the present disclosure in which the ability to recognize and bind to the target site is retained. The "functional variant" of a TALE-like polypeptide refers to a variant of the TALE-like polypeptide disclosed herein in which the ability to recognize and bind to a target sequence is retained.

[0100] The terms "targeting domain" and "targeting region" are used interchangeably herein and includes a nucleotide sequence that can hybridize (is complementary) to one strand (nucleotide sequence) of a double strand DNA target site. The percent complementation between the targeting region and the target sequence can be at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 63%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%or 100%. The variable targeting region can be at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides in length. In some embodiments, the variable targeting domain comprises a contiguous stretch of 12 to 30 nucleotides. The targeting domain can be composed of a DNA sequence, a RNA sequence, a modified DNA sequence, a modified RNA sequence, or any combination thereof.

[0101] The terms "target site" , "target sequence" , and "target region" are used interchangeably herein and refer to a nucleotide sequence on a chromosome, episome, a locus, or any other DNA molecule in the genome (including chromosomal, chloroplastic, mitochondrial DNA, plasmid DNA) of a cell, at which the TALE-like polypeptide of the present disclosure can recognize and bind to. The target site can be an endogenous site in the genome of a cell, or alternatively, the target site can be heterologous to the cell and thereby not be naturally occurring in the genome of the cell, or the target site can be found in a heterologous genomic location compared to where it occurs in nature.

[0102] An "altered target site" , "altered target sequence" , "modified target site" , "modified target sequence" are used interchangeably herein and refer to a target sequence as disclosed herein that comprises at least one alteration when compared to non-altered target sequence. Such "alteration" includes, for example: (i) replacement of at least one nucleotide, (ii) a deletion of at least one nucleotide, (iii) an insertion of at least one nucleotide, (iv) a chemical alteration of at least one nucleotide, or (v) any combination of (i) - (iv) .

[0103] A "modified nucleotide" or "edited nucleotide" refers to a nucleotide sequence of interest that comprises at least one alteration when compared to its non-modified nucleotide sequence. Such "alterations" include, for example: (i) replacement of at least one nucleotide, (ii) a deletion of at least one nucleotide, (iii) an insertion of at least one nucleotide, (iv) a chemical alteration of at least one nucleotide, or (v) any combination of (i) - (iv) .

[0104] Methods for "modifying a target site" and "altering a target site" are used interchangeably herein and refer to methods for producing an altered target site.

[0105] As used herein, "donor DNA" is a DNA construct that comprises a polynucleotide of interest to be inserted into the target site of the TALE-like polypeptide of the present disclosure.

[0106] The term "polynucleotide modification template" includes a polynucleotide that comprises at least one nucleotide modification when compared to the nucleotide sequence to be edited. A nucleotide modification can be the substitution, addition, insertion or deletion of at least one nucleotide. Optionally, the polynucleotide modification template can further comprise homologous nucleotide sequences flanking the at least one nucleotide modification, wherein the flanking homologous nucleotide sequences provide sufficient homology to the desired nucleotide sequence to be edited.

[0107] As used herein, the term "before" , in reference to a sequence position, refers to an occurrence of one sequence upstream of another sequence (at the 5’ end for nucleotide sequences, or at the N terminus for the amino acid sequences) . The term “after” in reference to a sequence position, refers to an occurrence of one sequence downstream of another sequence (at the 3’ end for nucleotide sequences, or at the C terminus for the amino acid sequences) .

[0108] 2. TR-containing polypeptide capable of binding to DNA with sequence specificity

[0109] The inventors have identified a number of TRPs, which can bind to DNA with sequence specificity. Upon the sequence analysis, the inventors classified the TRPs as families STAR (Short TALE-like Repeat proteins) , MOON (Marine Organism-Originated DNA binding protein) , pTERF (prokaryotic mTERF-like protein) , among which the STAR family includes two TRPs PqSTAR1 (SEQ ID NO: 1) and AspSTAR1 (SEQ ID NO: 2) , each comprising nine TRs showing similar structures with canonical TALE (AvrBs3) , and the same TALE code (nucleotide recognition by RVD) .

[0110] Therefore, the present disclosure provides a TRP, which can bind to DNA with sequence specificity. In some embodiment, the TRP is a programmed / programmable TALE-like polypeptide derived from the TRP in the STAR family, in which the TR domain is programmed / programmable by inserting, deleting and / or reordering TRs, and / or substituting the RVD in a TR.

[0111] In some embodiments, the programmed / programmable TALE-like polypeptide comprises a N-terminal region, two or more tandem repeats and a C-terminal region, wherein each of the tandem repeats comprises an amino acid sequence of

[0112] FX1X2DNLX3X4X5X6X7X8X9GX10X11X12X13LX14X15LLX16X17X18PX19LX20X21X22G

[0113] where

[0114] X1 is G or S, X2 is N or P, X3 is V or I, X4 is K or R, X5 is V or I, X6 is A or G, X7 is A or G, X8X9 are repeat variable di-residue (RVD) , X10 is G or S or A, X11 is A or Q or K, X12 is Q or H or K, X13 is A or T, X14 is Q or D, X15 is A or T, X16 is D or Q, X17 is K or V or R, X18 is G or Y or S, X19 is A or K or Q or R or T, X20 is R or A or T, X21 is Q or N, and X22 is A or G.

[0115] In some embodiments, the programmed / programmable TALE-like polypeptide comprises a N-terminal region, two or more tandem repeats and a C-terminal region, wherein each of the tandem repeats comprises an amino acid sequence selected from the group consisting of Formulae I to VII and XXII to XXIV

[0116] FX1NDNLVKVAAX2X3GX4X5X6ALQX7LLDX8GPALRQAG      (I)

[0117] where

[0118] X1 is G or S, X2X3 are repeat variable di-residue (RVD) , X4 is G or S, X5 is A or Q, X6 is H or Q, X7 is A or T, and X8 is K or R;

[0119] F X1H X2QIV X3IAS X4X5GGSQAL X6X7VL X8X9X10A X11L X12X13X14G     (II)

[0120] where

[0121] X1 is T or K, X2 is Q, E or R, X3 is A or G, X4X5 are RVD, X6 is N or D, X7 is T or K, X8 is A or V, X9 is T, R or K, X10 is H or Y, X11 is A, P or Q, X12 is T or R, X13 is A, D or T, and X14 is A or V.

[0122] GGREQVIKIAA X1X2GGKQALQALLDKSPALRQAG     (III)

[0123] where X1X2 are RVD,

[0124] FSNDNLVRIGG X1X2GAKKTLDTLLQVYPKLTQGG     (IV)

[0125] where X1X2 are RVD,

[0126] ILSGQETNRIK X1X2GGAKALETLSEKAEALHRAG     (V)

[0127] where X1X2 are RVD,

[0128] FSKQEAVAIAS X1X2GGSQALNTVLATHATLTAAG     (VI)

[0129] where X1X2 are RVD,

[0130] FAVEDVSAIAA X1X2GGAPALQAVVDHLELLMTRH     (VII)

[0131] where X1X2 are RVD,

[0132] FX1X2DNLX3KVAAX4X5GGX6QALLDKX7PX8LRX9AG      (XXII)

[0133] where X1 is G or S, X2 is N or P, X3 is V or I, X4X5 are RVD, X6 is A or Q, X7 is G or S, X8 is A or T, and X9 is Q or N,

[0134] GSREQVIKIAAX1X2GGQQALQALLDKGPALRNAG     (XXIII)

[0135] where X1X2 are RVD,

[0136] and

[0137] FSNDNLVRIGGX1X2GAKKTLDTLLQVYPQLTQGG     (XXIV)

[0138] where X1X2 are RVD.

[0139] In some embodiments, the RVD is selected from the group consisting of HD (recognizing C) , NN (recognizing A and G) , NG (recognizing T) , NS / NI (recognizing A) , NK (recognizing G) , KS / KG (recognizing T) , HI (recognizing A and G) .

[0140] In some embodiments, each of the tandem repeats comprises an amino acid sequence selected from the group consisting of the Formulae III to XXI and XXIII to XXXV

[0141] GGREQVIKIAA X1X2GGKQALQALLDKSPALRQAG     (III)

[0142] FSNDNLVRIGG X1X2GAKKTLDTLLQVYPKLTQGG     (IV)

[0143] ILSGQETNRIK X1X2GGAKALETLSEKAEALHRAG     (V)

[0144] FSKQEAVAIAS X1X2GGSQALNTVLATHATLTAAG     (VI)

[0145] FAVEDVSAIAA X1X2GGAPALQAVVDHLELLMTRH     (VII)

[0146] FGNDNLVKVAA X1X2GGAQALQALLDKGPALRQAG     (VIII)

[0147] FGNDNLVKVAA X1X2GSQHALQALLDKGPALRQAG     (IX)

[0148] FGNDNLVKVAA X1X2GGAQALQALLDKGPALRQAG     (X)

[0149] FGNDNLVKVAA X1X2GSQQALQALLDKGPALRQAG     (XI)

[0150] FGNDNLVKVAA X1X2GGAQALQALLDRGPALRQAG     (XII)

[0151] FSNDNLVKVAA X1X2GGAHALQALLDKGPALRQAG     (XIII)

[0152] FSNDNLVKVAA X1X2GGQQALQTLLDKGPALRQAG     (XIV)

[0153] FTHQQIVAIAS X1X2GGSQALNTVLATHAALTAAG     (XV)

[0154] FTHQQIVAIAS X1X2GGSQALDKVLATHAPLTAAG     (XVI)

[0155] FTHRQIVGIAS X1X2GGSQALDTVLVRYAPLRDAG     (XVII)

[0156] FKHEQIVGIAS X1X2GGSQALDKVLATHAQLTAVG     (XVIII)

[0157] FKHEQIVAIAS X1X2GGSQALDKVLVKYAPLTAAG     (XIX)

[0158] FTHQQIVAIAS X1X2GGSQALDTVLATHAQLTTAG     (XX)

[0159] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG     (XXI)

[0160] GSREQVIKIAAX1X2GGQQALQALLDKGPALRNAG     (XXIII)

[0161] FSNDNLVRIGGX1X2GAKKTLDTLLQVYPQLTQGG     (XXIV)

[0162] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXV)

[0163] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXVI)

[0164] FSPDNLIKVAAX1X2GGAQALQALLDKSPALRQAG     (XXVII)

[0165] FGPDNLVKVAAX1X2GGQQALQALLDKGPALRQAG    (XXVIII)

[0166] FGPDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXIX)

[0167] FGPDNLVKVAAX1X2GGAQALQALLDKGPTLRQAG    (XXX)

[0168] FSPDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXXI)

[0169] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXXII)

[0170] FGNDNLVKVAAX1X2GGQQALQALLDKGPALRNAG    (XXXIII)

[0171] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRNAG    (XXXIV)

[0172] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXXV)

[0173] where X1X2 are RVD.

[0174] The N-terminal region and C-terminal region can be those from the known TALEs in the art, such as AvrBs3. The N-terminal region and C-terminal region can be those from the TRPs in STAR family of the present disclosure.

[0175] In some embodiments, the N-terminal region comprises an amino acid sequence of positions 1-53 of SEQ ID NO: 1 (optionally comprising an additional N-terminal M, if needed) , positions 1-97 of SEQ ID NO: 2, or positions 1-39 of SEQ ID NO: 27. In some embodiments, the N-terminal region can be truncated by removing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more residues from the N terminus. For example, the N-terminal region comprises an amino acid sequence of positions 5-53, 10-53, or 15-53 of SEQ ID NO: 1, the N-terminal region comprises an amino acid sequence of positions 5-97, 10-97, 15-97, or 20-97 of SEQ ID NO: 2, or the N-terminal region comprises an amino acid sequence of positions 5-39, 10-39 or 15-39 of SEQ ID NO: 27.

[0176] In some embodiments, the C-terminal region comprises an amino acid sequence of positions 351-382 of SEQ ID NO: 1, positions 392-762 of SEQ ID NO: 2, or positions 469-500 of SEQ ID NO: 27. In some embodiments, the C-terminal region can be truncated by removing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more residues from the C terminus. For example, the C-terminal region comprises an amino acid sequence of positions 351-380, 351-375 or 351-370 of SEQ ID NO: 1, the N-terminal region comprises an amino acid sequence of positions 392-750, 392-700 or 392-650 of SEQ ID NO: 2, or the N-terminal region comprises an amino acid sequence of positions 469-495, 469-490 or 469-485 of SEQ ID NO: 27.

[0177] In some embodiments, the N-terminal region comprises an amino acid sequence at least 60%, 70%, 80%, 90%, 92%, 94%, 96%or 98%identical to the amino acid sequence of positions 1-53 of SEQ ID NO: 1 (optionally comprising an additional N-terminal M, if needed) , at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%or 98%identical to the amino acid sequence of positions 1-97 of SEQ ID NO: 2, or at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%or 97%identical to the amino acid sequence of positions 1-39 of SEQ ID NO: 27.

[0178] In some embodiments, the C-terminal region comprises an amino acid sequence at least 60%, 70%, 80%, 90%, 93%or 96%identical to the amino acid sequence of positions 351-382 of SEQ ID NO: 1, at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%or 99.5%identical to the amino acid sequence of positions 392-762 of SEQ ID NO: 2, or at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, or 96%identical to the amino acid sequence of positions 469-500 of SEQ ID NO: 27.

[0179] Surprisingly, the inventors found that the STARs can achieve a higher affinity than canonical TALE with less TRs. In particular, both PqSTAR1 and AspSTAR1 comprise nine TRs, recognizing and binding to a sequence of nine nucleotides, while the AvrBs3 comprises eighteen TRs, but the binding affinities of PqSTAR1 and AspSTAR1 to DNA are higher than or comparable with AvrBs3. Therefore, the target sequence of the TALE-like polypeptide can be shorter than canonical TALE, offering a more flexible design of the TALE-like polypeptide.

[0180] In some embodiments, the programmed / programmable TALE-like polypeptide comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more TRs. Preferably, the programmed / programmable TALE-like polypeptide comprises 3-17, 4-17, 5-17, 6-17, 7-17, 8-17 or 9-17 TRs.

[0181] In some embodiments, the TRs are identical except the RVD. In some embodiments, the TRs are different from each other in the residues other than RVD.

[0182] In some embodiments, the TRs are derived from native TRs from a single TRP. In some embodiments, the TRs are derived from SEQ ID NO: 1. In some embodiments, the TRs are derived from SEQ ID NO: 2. In some embodiments, the TRs are derived from SEQ ID NO: 27.

[0183] In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae I, III and IV. In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae II and V to VII. In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae XXII to XXIV.

[0184] In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae III, IV and VIII to XIV. In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae V to VII and XV to XX. In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae XXIII to XXXV.

[0185] In some embodiments, the programmed / programmable TALE-like polypeptide comprises an amino acid sequence of SEQ ID NO: 25, 26 or 28,

[0186] SEQ ID NO: 25

[0187] SEQ ID NO: 26

[0188] SEQ ID NO: 28

[0189] wherein X1X2 are RVDs selected from the group consisting of HD (recognizing C) , NN (recognizing A and G) , NG (recognizing T) , NS / NI (recognizing A) , NK (recognizing G) , KS / KG (recognizing T) , HI (recognizing A and G) .

[0190] In some embodiments, the programmed / programmable TALE-like polypeptide comprises an amino acid sequence of SEQ ID NO: 1, 2 or 27 or an amino acid sequence at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%or 99.5%identical to any one of SEQ ID NO: 1, 2 or 27, e.g., with varied RVDs.

[0191] The inventors have also identified DNA binding TRPs of SEQ ID NOs: 3-6 and 9-22. Therefore, the present disclosure further provides a TRP comprising an amino acid sequence selected from the group consisting of SEQ ID NOs: 3-6 and 9-22 or an amino acid sequence at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%or 99.5%identical to any one of SEQ ID NOs: 3-6 and 9-22, preferably with the TRs unchanged.

[0192] 3. Fusion polypeptide

[0193] The present disclosure provides a fusion polypeptide comprising a TRP, which can bind to DNA with sequence specificity, fused to a fusion partner.

[0194] In some embodiments, the TRP is a programmed / programmable TALE-like polypeptide. In some embodiments, the programmed / programmable TALE-like polypeptide comprises a N-terminal region, two or more tandem repeats and a C-terminal region, wherein each of the tandem repeats comprises an amino acid sequence of

[0195] FX1X2DNLX3X4X5X6X7X8X9GX10X11X12X13LX14X15LLX16X17X18PX19LX20X21X22G

[0196] where

[0197] X1 is G or S, X2 is N or P, X3 is V or I, X4 is K or R, X5 is V or I, X6 is A or G, X7 is A or G, X8X9 are repeat variable di-residue (RVD) , X10 is G or S or A, X11 is A or Q or K, X12 is Q or H or K, X13 is A or T, X14 is Q or D, X15 is A or T, X16 is D or Q, X17 is K or V or R, X18 is G or Y or S, X19 is A or K or Q or R or T, X20 is R or A or T, X21 is Q or N, and X22 is A or G.

[0198] In some embodiments, the programmed / programmable TALE-like polypeptide comprises a N-terminal region, two or more tandem repeats and a C-terminal region, wherein each of the tandem repeats comprises an amino acid sequence selected from the group consisting of Formulae I to VII and XXII to XXIV

[0199] FX1NDNLVKVAAX2X3GX4X5X6ALQX7LLDX8GPALRQAG      (I)

[0200] where

[0201] X1 is G or S, X2X3 are repeat variable di-residue (RVD) , X4 is G or S, X5 is A or Q, X6 is H or Q, X7 is A or T, and X8 is K or R;

[0202] F X1H X2QIV X3IAS X4X5GGSQAL X6X7VL X8X9X10A X11L X12X13X14G     (II)

[0203] where

[0204] X1 is T or K, X2 is Q, E or R, X3 is A or G, X4X5 are RVD, X6 is N or D, X7 is T or K, X8 is A or V, X9 is T, R or K, X10 is H or Y, X11 is A, P or Q, X12 is T or R, X13 is A, D or T, and X14 is A or V.

[0205] GGREQVIKIAA X1X2GGKQALQALLDKSPALRQAG     (III)

[0206] where X1X2 are RVD,

[0207] FSNDNLVRIGG X1X2GAKKTLDTLLQVYPKLTQGG     (IV)

[0208] where X1X2 are RVD,

[0209] ILSGQETNRIK X1X2GGAKALETLSEKAEALHRAG     (V)

[0210] where X1X2 are RVD,

[0211] FSKQEAVAIAS X1X2GGSQALNTVLATHATLTAAG     (VI)

[0212] where X1X2 are RVD,

[0213] FAVEDVSAIAA X1X2GGAPALQAVVDHLELLMTRH     (VII)

[0214] where X1X2 are RVD,

[0215] FX1X2DNLX3KVAAX4X5GGX6QALLDKX7PX8LRX9AG      (XXII)

[0216] where X1 is G or S, X2 is N or P, X3 is V or I, X4X5 are RVD, X6 is A or Q, X7 is G or S, X8 is A or T, and X9 is Q or N,

[0217] GSREQVIKIAAX1X2GGQQALQALLDKGPALRNAG     (XXIII)

[0218] where X1X2 are RVD,

[0219] and

[0220] FSNDNLVRIGGX1X2GAKKTLDTLLQVYPQLTQGG     (XXIV)

[0221] where X1X2 are RVD.

[0222] In some embodiments, the RVD is selected from the group consisting of HD (recognizing C) , NN (recognizing A and G) , NG (recognizing T) , NS / NI (recognizing A) , NK (recognizing G) , KS / KG (recognizing T) , HI (recognizing A and G) .

[0223] In some embodiments, each of the tandem repeats comprises an amino acid sequence selected from the group consisting of the Formulae III to XXI and XXIII to XXXV

[0224] GGREQVIKIAA X1X2GGKQALQALLDKSPALRQAG     (III)

[0225] FSNDNLVRIGG X1X2GAKKTLDTLLQVYPKLTQGG     (IV)

[0226] ILSGQETNRIK X1X2GGAKALETLSEKAEALHRAG     (V)

[0227] FSKQEAVAIAS X1X2GGSQALNTVLATHATLTAAG     (VI)

[0228] FAVEDVSAIAA X1X2GGAPALQAVVDHLELLMTRH     (VII)

[0229] FGNDNLVKVAA X1X2GGAQALQALLDKGPALRQAG     (VIII)

[0230] FGNDNLVKVAA X1X2GSQHALQALLDKGPALRQAG     (IX)

[0231] FGNDNLVKVAA X1X2GGAQALQALLDKGPALRQAG     (X)

[0232] FGNDNLVKVAA X1X2GSQQALQALLDKGPALRQAG     (XI)

[0233] FGNDNLVKVAA X1X2GGAQALQALLDRGPALRQAG     (XII)

[0234] FSNDNLVKVAA X1X2GGAHALQALLDKGPALRQAG     (XIII)

[0235] FSNDNLVKVAA X1X2GGQQALQTLLDKGPALRQAG     (XIV)

[0236] FTHQQIVAIAS X1X2GGSQALNTVLATHAALTAAG     (XV)

[0237] FTHQQIVAIAS X1X2GGSQALDKVLATHAPLTAAG     (XVI)

[0238] FTHRQIVGIAS X1X2GGSQALDTVLVRYAPLRDAG     (XVII)

[0239] FKHEQIVGIAS X1X2GGSQALDKVLATHAQLTAVG     (XVIII)

[0240] FKHEQIVAIAS X1X2GGSQALDKVLVKYAPLTAAG     (XIX)

[0241] FTHQQIVAIAS X1X2GGSQALDTVLATHAQLTTAG     (XX)

[0242] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG     (XXI)

[0243] GSREQVIKIAAX1X2GGQQALQALLDKGPALRNAG     (XXIII)

[0244] FSNDNLVRIGGX1X2GAKKTLDTLLQVYPQLTQGG     (XXIV)

[0245] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXV)

[0246] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXVI)

[0247] FSPDNLIKVAAX1X2GGAQALQALLDKSPALRQAG     (XXVII)

[0248] FGPDNLVKVAAX1X2GGQQALQALLDKGPALRQAG    (XXVIII)

[0249] FGPDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXIX)

[0250] FGPDNLVKVAAX1X2GGAQALQALLDKGPTLRQAG    (XXX)

[0251] FSPDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXXI)

[0252] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXXII)

[0253] FGNDNLVKVAAX1X2GGQQALQALLDKGPALRNAG    (XXXIII)

[0254] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRNAG    (XXXIV)

[0255] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXXV)

[0256] where X1X2 are RVD.

[0257] The N-terminal region and C-terminal region can be those from the known TALEs in the art, such as AvrBs3. The N-terminal region and C-terminal region can be those from the TRPs in STAR family of the present disclosure.

[0258] In some embodiments, the N-terminal region comprises an amino acid sequence of positions 1-53 of SEQ ID NO: 1 (optionally comprising an additional N-terminal M, if needed) , positions 1-97 of SEQ ID NO: 2, or positions 1-39 of SEQ ID NO: 27. In some embodiments, the N-terminal region can be truncated by removing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more residues from the N terminus. For example, the N-terminal region comprises an amino acid sequence of positions 5-53, 10-53, or 15-53 of SEQ ID NO: 1, the N-terminal region comprises an amino acid sequence of positions 5-97, 10-97, 15-97, or 20-97 of SEQ ID NO: 2, or the N-terminal region comprises an amino acid sequence of positions 5-39, 10-39 or 15-39 of SEQ ID NO: 27.

[0259] In some embodiments, the C-terminal region comprises an amino acid sequence of positions 351-382 of SEQ ID NO: 1, positions 392-762 of SEQ ID NO: 2, or positions 469-500 of SEQ ID NO: 27. In some embodiments, the C-terminal region can be truncated by removing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more residues from the C terminus. For example, the C-terminal region comprises an amino acid sequence of positions 351-380, 351-375 or 351-370 of SEQ ID NO: 1, the N-terminal region comprises an amino acid sequence of positions 392-750, 392-700 or 392-650 of SEQ ID NO: 2, or the N-terminal region comprises an amino acid sequence of positions 469-495, 469-490 or 469-485 of SEQ ID NO: 27.

[0260] In some embodiments, the N-terminal region comprises an amino acid sequence at least 60%, 70%, 80%, 90%, 92%, 94%, 96%or 98%identical to the amino acid sequence of positions 1-53 of SEQ ID NO: 1 (optionally comprising an additional N-terminal M, if needed) , at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%or 98%identical to the amino acid sequence of positions 1-97 of SEQ ID NO: 2, or at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%or 97%identical to the amino acid sequence of positions 1-39 of SEQ ID NO: 27.

[0261] In some embodiments, the C-terminal region comprises an amino acid sequence at least 60%, 70%, 80%, 90%, 93%or 96%identical to the amino acid sequence of positions 351-382 of SEQ ID NO: 1, at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%or 99.5%identical to the amino acid sequence of positions 392-762 of SEQ ID NO: 2, or at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, or 96%identical to the amino acid sequence of positions 469-500 of SEQ ID NO: 27.

[0262] In some embodiments, the programmed / programmable TALE-like polypeptide comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more TRs. Preferably, the programmed / programmable TALE-like polypeptide comprises 3-17, 4-17, 5-17, 6-17, 7-17, 8-17 or 9-17 TRs.

[0263] In some embodiments, the TRs are identical except the RVD. In some embodiments, the TRs are different from each other in the residues other than RVD.

[0264] In some embodiments, the TRs are derived from native TRs from a single TRP. In some embodiments, the TRs are derived from SEQ ID NO: 1. In some embodiments, the TRs are derived from SEQ ID NO: 2. In some embodiments, the TRs are derived from SEQ ID NO: 27.

[0265] In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae I, III and IV. In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae II and V to VII. In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae XXII to XXIV.

[0266] In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae III, IV and VIII to XIV. In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae V to VII and XV to XX. In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae XXIII to XXXV.

[0267] In some embodiments, the programmed / programmable TALE-like polypeptide comprises an amino acid sequence of SEQ ID NO: 25, 26 or 28,

[0268] SEQ ID NO: 25

[0269] SEQ ID NO: 26

[0270] SEQ ID NO: 28

[0271] wherein X1X2 are RVDs selected from the group consisting of HD (recognizing C) , NN (recognizing A and G) , NG (recognizing T) , NS / NI (recognizing A) , NK (recognizing G) , KS / KG (recognizing T) , HI (recognizing A and G) .

[0272] In some embodiments, the programmed / programmable TALE-like polypeptide comprises an amino acid sequence of SEQ ID NO: 1, 2 or 27 or an amino acid sequence at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%or 99.5%identical to any one of SEQ ID NO: 1, 2 or 27, e.g., with varied RVDs.

[0273] In some embodiments, the TRP comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 3-6 and 9-22 or an amino acid sequence at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%or 99.5%identical to any one of SEQ ID NOs: 3-6 and 9-22, preferably with the TRs unchanged.

[0274] In some embodiments, the fusion partner is a polypeptide that provides an activity of nuclease, including, for example, an active (double strand break creating) , partially active (nickase) nuclease. The nuclease can be an endonuclease. In some embodiments, the fusion partner is another polypeptide or domain, for example Clo51 or FokI nuclease, to generate double-strand breaks (Guilinger et al. Nature Biotechnology, volume 32, number 6, June 2014) .

[0275] In some embodiments, the fusion partner is a polypeptide that provides an activity that indirectly increases transcription by acting directly on the target DNA or on a polypeptide (e.g., a histone or other DNA-binding protein) associated with the target DNA.

[0276] In further embodiments, the fusion partner is a polypeptide that provides for methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, ribosylation activity, deribosylation activity, myristoylation activity, or demyristoylation activity.

[0277] In further embodiments, the fusion partner is a polypeptide that directly provides for increased transcription of the target nucleic acid (e.g., a transcription activator or a fragment thereof, a protein or fragment thereof that recruits a transcription activator, a small molecule / drug-responsive transcription regulator, etc. ) .

[0278] In some embodiments, the fusion partner is a polypeptide that directs editing of single or multiple bases in a polynucleotide sequence, for example a site-specific deaminase that can change the identity of a nucleotide, for example from C-G to T-A or an A-T to G-C (Gaudelli et al., 2017, Programmable base editing of A-T to G-C in genomic DNA without DNA cleavage. Nature, 551 (7681) : 464-471; Nishida et al., 2016, Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems. Science, 353 (6305) : 1248; Komor et al., 2016, Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature, 533 (7603) : 420-424) .

[0279] The fusion polypeptide may comprise, for example, a deaminase (such as, but not limited to, a cytidine deaminase, an adenine deaminase, APOBEC1, APOBEC3A, BE2, BE3, BE4, ABEs, or the like) . In some embodiments, the fusion partner includes base edit repair inhibitors and glycosylase inhibitors (e.g., uracil glycosylase inhibitor (to prevent uracil removal) ) .

[0280] The fusion polypeptide may also comprise a heterologous nuclear localization sequence (NLS) . A heterologous NLS herein may be of sufficient strength to drive accumulation of the fusion polypeptide in a detectable amount in the nucleus of a eukaryotic cell. An NLS may comprise one (monopartite) or more (e.g., bipartite) short sequences (e.g., 2 to 20 residues) of basic, positively charged residues (e.g., lysine and / or arginine) . An NLS may be present at the N-terminus or C-terminus of the fusion polypeptide, for example. Two or more NLS sequences can be present, for example, on both the N-and C-termini of the fusion polypeptide.

[0281] 4. Polynucleotide and construct for expressing the TRP or fusion polypeptide

[0282] The TRP or fusion polypeptide of the present disclosure can be isolated from a recombinant source where the host cell is genetically modified to express the nucleotide sequence encoding the polypeptide. Alternatively, the TRP or fusion polypeptide can be produced using cell free protein expression systems, or be synthetically produced.

[0283] Therefore, the present disclosure also provides an isolated polynucleotide comprising a nucleotide sequence encoding the TRP or fusion polypeptide of the present disclosure.

[0284] The TRP polypeptide or fusion polypeptide can be expressed in a cell. Cells include, but are not limited to, human, non-human, animal, bacterial, fungal, insect, yeast, and plant cells.

[0285] Standard recombinant DNA and molecular cloning techniques used herein are well known in the art and are described more fully in Sambrook et al., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory: Cold Spring Harbor, N. Y. (1989) . Transformation methods are well known to those skilled in the art and are described infra.

[0286] Provided are also vectors and constructs including circular plasmids, and linear polynucleotides, comprising a polynucleotide of interest and optionally other components including linkers, adapters, regulatory sequences.

[0287] In some embodiments, the vector comprises an expression cassette encoding the TRP or fusion polypeptide.

[0288] In some embodiments, the expression of the TRP or fusion polypeptide is driven by a constitutive promoter, an inducible promoter, or a spatio-temporal specific promoter.

[0289] 5. Gene editing with the fusion polypeptide

[0290] The present disclosure provides a recombinant gene editing system comprising a fusion polypeptide a fusion polypeptide comprising a TRP, which can bind to DNA with sequence specificity, fused to a fusion partner, or a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide.

[0291] The present disclosure provides a composition comprising a fusion polypeptide comprising a TRP, which can bind to DNA with sequence specificity, fused to a fusion partner, or a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide.

[0292] The present disclosure provides method of introducing a double-strand break into a polynucleotide of interest, comprising a step of contacting the polynucleotide with a recombinant gene editing system comprising a fusion polypeptide comprising a TRP, which can bind to DNA with sequence specificity, fused to a fusion partner.

[0293] The present disclosure provides method of modifying a genomic sequence in a cell such as eukaryotic cell, comprising a step of introducing into the cell a recombinant gene editing system comprising a fusion polypeptide comprising a TRP, which can bind to DNA with sequence specificity, fused to a fusion partner, or a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide.

[0294] In some embodiments, the TRP is a programmed / programmable TALE-like polypeptide. In some embodiments, the programmed / programmable TALE-like polypeptide comprises a N-terminal region, two or more tandem repeats and a C-terminal region, wherein each of the tandem repeats comprises an amino acid sequence of

[0295] FX1X2DNLX3X4X5X6X7X8X9GX10X11X12X13LX14X15LLX16X17X18PX19LX20X21X22G

[0296] where

[0297] X1 is G or S, X2 is N or P, X3 is V or I, X4 is K or R, X5 is V or I, X6 is A or G, X7 is A or G, X8X9 are repeat variable di-residue (RVD) , X10 is G or S or A, X11 is A or Q or K, X12 is Q or H or K, X13 is A or T, X14 is Q or D, X15 is A or T, X16 is D or Q, X17 is K or V or R, X18 is G or Y or S, X19 is A or K or Q or R or T, X20 is R or A or T, X21 is Q or N, and X22 is A or G.

[0298] In some embodiments, the programmed / programmable TALE-like polypeptide comprises a N-terminal region, two or more tandem repeats and a C-terminal region, wherein each of the tandem repeats comprises an amino acid sequence selected from the group consisting of Formulae I to VII and XXII to XXIV

[0299] FX1NDNLVKVAAX2X3GX4X5X6ALQX7LLDX8GPALRQAG      (I)

[0300] where

[0301] X1 is G or S, X2X3 are repeat variable di-residue (RVD) , X4 is G or S, X5 is A or Q, X6 is H or Q, X7 is A or T, and X8 is K or R;

[0302] F X1H X2QIV X3IAS X4X5GGSQAL X6X7VL X8X9X10A X11L X12X13X14G     (II)

[0303] where

[0304] X1 is T or K, X2 is Q, E or R, X3 is A or G, X4X5 are RVD, X6 is N or D, X7 is T or K, X8 is A or V, X9 is T, R or K, X10 is H or Y, X11 is A, P or Q, X12 is T or R, X13 is A, D or T, and X14 is A or V.

[0305] GGREQVIKIAA X1X2GGKQALQALLDKSPALRQAG     (III)

[0306] where X1X2 are RVD,

[0307] FSNDNLVRIGG X1X2GAKKTLDTLLQVYPKLTQGG     (IV)

[0308] where X1X2 are RVD,

[0309] ILSGQETNRIK X1X2GGAKALETLSEKAEALHRAG     (V)

[0310] where X1X2 are RVD,

[0311] FSKQEAVAIAS X1X2GGSQALNTVLATHATLTAAG     (VI)

[0312] where X1X2 are RVD,

[0313] FAVEDVSAIAA X1X2GGAPALQAVVDHLELLMTRH     (VII)

[0314] where X1X2 are RVD,

[0315] FX1X2DNLX3KVAAX4X5GGX6QALLDKX7PX8LRX9AG      (XXII)

[0316] where X1 is G or S, X2 is N or P, X3 is V or I, X4X5 are RVD, X6 is A or Q, X7 is G or S, X8 is A or T, and X9 is Q or N,

[0317] GSREQVIKIAAX1X2GGQQALQALLDKGPALRNAG     (XXIII)

[0318] where X1X2 are RVD,

[0319] and

[0320] FSNDNLVRIGGX1X2GAKKTLDTLLQVYPQLTQGG     (XXIV)

[0321] where X1X2 are RVD.

[0322] In some embodiments, the RVD is selected from the group consisting of HD (recognizing C) , NN (recognizing A and G) , NG (recognizing T) , NS / NI (recognizing A) , NK (recognizing G) , KS / KG (recognizing T) , HI (recognizing A and G) .

[0323] In some embodiments, each of the tandem repeats comprises an amino acid sequence selected from the group consisting of the Formulae III to XXI and XXIII to XXXV

[0324] GGREQVIKIAA X1X2GGKQALQALLDKSPALRQAG     (III)

[0325] FSNDNLVRIGG X1X2GAKKTLDTLLQVYPKLTQGG     (IV)

[0326] ILSGQETNRIK X1X2GGAKALETLSEKAEALHRAG     (V)

[0327] FSKQEAVAIAS X1X2GGSQALNTVLATHATLTAAG     (VI)

[0328] FAVEDVSAIAA X1X2GGAPALQAVVDHLELLMTRH     (VII)

[0329] FGNDNLVKVAA X1X2GGAQALQALLDKGPALRQAG     (VIII)

[0330] FGNDNLVKVAA X1X2GSQHALQALLDKGPALRQAG     (IX)

[0331] FGNDNLVKVAA X1X2GGAQALQALLDKGPALRQAG     (X)

[0332] FGNDNLVKVAA X1X2GSQQALQALLDKGPALRQAG     (XI)

[0333] FGNDNLVKVAA X1X2GGAQALQALLDRGPALRQAG     (XII)

[0334] FSNDNLVKVAA X1X2GGAHALQALLDKGPALRQAG     (XIII)

[0335] FSNDNLVKVAA X1X2GGQQALQTLLDKGPALRQAG     (XIV)

[0336] FTHQQIVAIAS X1X2GGSQALNTVLATHAALTAAG     (XV)

[0337] FTHQQIVAIAS X1X2GGSQALDKVLATHAPLTAAG     (XVI)

[0338] FTHRQIVGIAS X1X2GGSQALDTVLVRYAPLRDAG     (XVII)

[0339] FKHEQIVGIAS X1X2GGSQALDKVLATHAQLTAVG     (XVIII)

[0340] FKHEQIVAIAS X1X2GGSQALDKVLVKYAPLTAAG     (XIX)

[0341] FTHQQIVAIAS X1X2GGSQALDTVLATHAQLTTAG     (XX)

[0342] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG     (XXI)

[0343] GSREQVIKIAAX1X2GGQQALQALLDKGPALRNAG     (XXIII)

[0344] FSNDNLVRIGGX1X2GAKKTLDTLLQVYPQLTQGG     (XXIV)

[0345] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXV)

[0346] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXVI)

[0347] FSPDNLIKVAAX1X2GGAQALQALLDKSPALRQAG     (XXVII)

[0348] FGPDNLVKVAAX1X2GGQQALQALLDKGPALRQAG    (XXVIII)

[0349] FGPDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXIX)

[0350] FGPDNLVKVAAX1X2GGAQALQALLDKGPTLRQAG    (XXX)

[0351] FSPDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXXI)

[0352] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXXII)

[0353] FGNDNLVKVAAX1X2GGQQALQALLDKGPALRNAG    (XXXIII)

[0354] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRNAG    (XXXIV)

[0355] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXXV)

[0356] where X1X2 are RVD.

[0357] The N-terminal region and C-terminal region can be those from the known TALEs in the art, such as AvrBs3. The N-terminal region and C-terminal region can be those from the TRPs in STAR family of the present disclosure.

[0358] In some embodiments, the N-terminal region comprises an amino acid sequence of positions 1-53 of SEQ ID NO: 1 (optionally comprising an additional N-terminal M, if needed) , positions 1-97 of SEQ ID NO: 2, or positions 1-39 of SEQ ID NO: 27. In some embodiments, the N-terminal region can be truncated by removing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more residues from the N terminus. For example, the N-terminal region comprises an amino acid sequence of positions 5-53, 10-53, or 15-53 of SEQ ID NO: 1, or the N-terminal region comprises an amino acid sequence of positions 5-97, 10-97, 15-97, or 20-97 of SEQ ID NO: 2, or the N-terminal region comprises an amino acid sequence of positions 5-39, 10-39 or 15-39 of SEQ ID NO: 27.

[0359] In some embodiments, the C-terminal region comprises an amino acid sequence of positions 351-382 of SEQ ID NO: 1, positions 392-762 of SEQ ID NO: 2, or positions 469-500 of SEQ ID NO: 27. In some embodiments, the C-terminal region can be truncated by removing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more residues from the C terminus. For example, the C-terminal region comprises an amino acid sequence of positions 351-380, 351-375 or 351-370 of SEQ ID NO: 1, the N-terminal region comprises an amino acid sequence of positions 392-750, 392-700 or 392-650 of SEQ ID NO: 2, or the N-terminal region comprises an amino acid sequence of positions 469-495, 469-490 or 469-485 of SEQ ID NO: 27.

[0360] In some embodiments, the N-terminal region comprises an amino acid sequence at least 60%, 70%, 80%, 90%, 92%, 94%, 96%or 98%identical to the amino acid sequence of positions 1-53 of SEQ ID NO: 1 (optionally comprising an additional N-terminal M, if needed) , at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%or 98%identical to the amino acid sequence of positions 1-97 of SEQ ID NO: 2, or at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%or 97%identical to the amino acid sequence of positions 1-39 of SEQ ID NO: 27.

[0361] In some embodiments, the C-terminal region comprises an amino acid sequence at least 60%, 70%, 80%, 90%, 93%or 96%identical to the amino acid sequence of positions 351-382 of SEQ ID NO: 1, at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%or 99.5%identical to the amino acid sequence of positions 392-762 of SEQ ID NO: 2, or at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, or 96%identical to the amino acid sequence of positions 469-500 of SEQ ID NO: 27.

[0362] In some embodiments, the programmed / programmable TALE-like polypeptide comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more TRs. Preferably, the programmed / programmable TALE-like polypeptide comprises 3-17, 4-17, 5-17, 6-17, 7-17, 8-17 or 9-17 TRs.

[0363] In some embodiments, the TRs are identical except the RVD. In some embodiments, the TRs are different from each other in the residues other than RVD.

[0364] In some embodiments, the TRs are derived from native TRs from a single TRP. In some embodiments, the TRs are derived from SEQ ID NO: 1. In some embodiments, the TRs are derived from SEQ ID NO: 2. In some embodiments, the TRs are derived from SEQ ID NO: 27.

[0365] In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae I, III and IV. In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae II and V to VII. In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae XXII to XXIV.

[0366] In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae III, IV and VIII to XIV. In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae V to VII and XV to XX. In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae XXIII to XXXV.

[0367] In some embodiments, the programmed / programmable TALE-like polypeptide comprises an amino acid sequence of SEQ ID NO: 25, 26 or 28,

[0368] SEQ ID NO: 25

[0369] SEQ ID NO: 26

[0370] SEQ ID NO: 28

[0371] wherein X1X2 are RVDs selected from the group consisting of HD (recognizing C) , NN (recognizing A and G) , NG (recognizing T) , NS / NI (recognizing A) , NK (recognizing G) , KS / KG (recognizing T) , HI (recognizing A and G) .

[0372] In some embodiments, the programmed / programmable TALE-like polypeptide comprises an amino acid sequence of SEQ ID NO: 1, 2 or 27 or an amino acid sequence at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%or 99.5%identical to any one of SEQ ID NO: 1, 2 or 27, e.g., with varied RVDs.

[0373] In some embodiments, the TRP comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 3-6 and 9-22 or an amino acid sequence at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%or 99.5%identical to any one of SEQ ID NOs: 3-6 and 9-22, preferably with the TRs unchanged.

[0374] In some embodiments, the fusion partner is a polypeptide that provides an activity of nuclease, including, for example, an active (double strand break creating) , partially active (nickase) nuclease. The nuclease can be an endonuclease. In some embodiments, the fusion partner is another polypeptide or domain, for example Clo51 or FokI nuclease, to generate double-strand breaks (Guilinger et al. Nature Biotechnology, volume 32, number 6, June 2014) .

[0375] In some embodiments, the fusion partner is a polypeptide that provides an activity that indirectly increases transcription by acting directly on the target DNA or on a polypeptide (e.g., a histone or other DNA-binding protein) associated with the target DNA.

[0376] In further embodiments, the fusion partner is a polypeptide that provides for methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, ribosylation activity, deribosylation activity, myristoylation activity, or demyristoylation activity.

[0377] In further embodiments, the fusion partner is a polypeptide that directly provides for increased transcription of the target nucleic acid (e.g., a transcription activator or a fragment thereof, a protein or fragment thereof that recruits a transcription activator, a small molecule / drug-responsive transcription regulator, etc. ) .

[0378] In some embodiments, the fusion partner is a polypeptide that directs editing of single or multiple bases in a polynucleotide sequence, for example a site-specific deaminase that can change the identity of a nucleotide, for example from C-G to T-A or an A-T to G-C (Gaudelli et al., Programmable base editing of A-T to G-C in genomic DNA without DNA cleavage. "Nature (2017) ; Nishida et al. "Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems. "Science 353 (6305) (2016) ; Komor et al. "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. "Nature 533 (7603) (2016) : 420-4.

[0379] The fusion polypeptide may comprise, for example, a deaminase (such as, but not limited to, a cytidine deaminase, an adenine deaminase, APOBEC1, APOBEC3A, BE2, BE3, BE4, ABEs, or the like) . In some embodiments, the fusion partner includes base edit repair inhibitors and glycosylase inhibitors (e.g., uracil glycosylase inhibitor (to prevent uracil removal) ) .

[0380] The fusion polypeptide may also comprise a heterologous nuclear localization sequence (NLS) . A heterologous NLS herein may be of sufficient strength to drive accumulation of the fusion polypeptide in a detectable amount in the nucleus of a eukaryotic cell. An NLS may comprise one (monopartite) or more (e.g., bipartite) short sequences (e.g., 2 to 20 residues) of basic, positively charged residues (e.g., lysine and / or arginine) . An NLS may be present at the N-terminus or C-terminus of the fusion polypeptide, for example. Two or more NLS sequences can be present, for example, on both the N-and C-termini of the fusion polypeptide.

[0381] In some embodiments, the recombinant gene editing system further comprises a heterologous polynucleotide, such as an expression cassette, a transgene, a donor DNA, or a polynucleotide modification template.

[0382] Methods for introducing polynucleotides or polypeptides or a polynucleotide-protein complex into cells or organisms are known in the art including, but not limited to, microinjection, electroporation, stable transformation methods, transient transformation methods, ballistic particle acceleration (particle bombardment) , whiskers mediated transformation, Agrobacterium-mediated transformation, direct gene transfer, viral-mediated introduction, transfection, transduction, cell-penetrating peptides, mesoporous silica nanoparticle (MSN) -mediated direct protein delivery, topical applications, sexual crossing, sexual breeding, and any combination thereof.

[0383] 6. Artificial transcription factors (TFs)

[0384] The present disclosure provides randomized libraries of artificial TFs, comprising a plurality of cells, each of which harboring a vector comprising a nucleotide sequence encoding a polypeptide comprising a N-terminal region, two or more tandem repeats and a C-terminal region.

[0385] In some embodiments, each of the TRs comprises an amino acid sequence of FX1X2DNLX3X4X5X6X7X8X9GX10X11X12X13LX14X15LLX16X17X18PX19LX20X21X22G where

[0386] X1 is G or S, X2 is N or P, X3 is V or I, X4 is K or R, X5 is V or I, X6 is A or G, X7 is A or G, X8X9 are repeat variable di-residue (RVD) , X10 is G or S or A, X11 is A or Q or K, X12 is Q or H or K, X13 is A or T, X14 is Q or D, X15 is A or T, X16 is D or Q, X17 is K or V or R, X18 is G or Y or S, X19 is A or K or Q or R or T, X20 is R or A or T, X21 is Q or N, and X22 is A or G.

[0387] In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of Formulae I to VII and XXII to XXIV

[0388] FX1NDNLVKVAAX2X3GX4X5X6ALQX7LLDX8GPALRQAG      (I)

[0389] where

[0390] X1 is G or S, X2X3 are repeat variable di-residue (RVD) , X4 is G or S, X5 is A or Q, X6 is H or Q, X7 is A or T, and X8 is K or R;

[0391] F X1H X2QIV X3IAS X4X5GGSQAL X6X7VL X8X9X10A X11L X12X13X14G     (II)

[0392] where

[0393] X1 is T or K, X2 is Q, E or R, X3 is A or G, X4X5 are RVD, X6 is N or D, X7 is T or K, X8 is A or V, X9 is T, R or K, X10 is H or Y; X11 is A, P or Q, X12 is T or R, X13 is A, D or T, and X14 is A or V.

[0394] GGREQVIKIAA X1X2GGKQALQALLDKSPALRQAG     (III)

[0395] where X1X2 are RVD,

[0396] FSNDNLVRIGG X1X2GAKKTLDTLLQVYPKLTQGG     (IV)

[0397] where X1X2 are RVD,

[0398] ILSGQETNRIK X1X2GGAKALETLSEKAEALHRAG     (V)

[0399] where X1X2 are RVD,

[0400] FSKQEAVAIAS X1X2GGSQALNTVLATHATLTAAG     (VI)

[0401] where X1X2 are RVD,

[0402] FAVEDVSAIAA X1X2GGAPALQAVVDHLELLMTRH     (VII)

[0403] where X1X2 are RVD,

[0404] FX1X2DNLX3KVAAX4X5GGX6QALLDKX7PX8LRX9AG      (XXII)

[0405] where X1 is G or S, X2 is N or P, X3 is V or I, X4X5 are RVD, X6 is A or Q, X7 is G or S, X8 is A or T, and X9 is Q or N,

[0406] GSREQVIKIAAX1X2GGQQALQALLDKGPALRNAG     (XXIII)

[0407] where X1X2 are RVD,

[0408] and

[0409] FSNDNLVRIGGX1X2GAKKTLDTLLQVYPQLTQGG     (XXIV)

[0410] where X1X2 are RVD,

[0411] wherein the tandem repeats are randomized between cells.

[0412] In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae III to XXI and XXIII to XXXV

[0413] GGREQVIKIAA X1X2GGKQALQALLDKSPALRQAG     (III)

[0414] FSNDNLVRIGG X1X2GAKKTLDTLLQVYPKLTQGG     (IV)

[0415] ILSGQETNRIK X1X2GGAKALETLSEKAEALHRAG     (V)

[0416] FSKQEAVAIAS X1X2GGSQALNTVLATHATLTAAG     (VI)

[0417] FAVEDVSAIAA X1X2GGAPALQAVVDHLELLMTRH     (VII)

[0418] FGNDNLVKVAA X1X2GGAQALQALLDKGPALRQAG     (VIII)

[0419] FGNDNLVKVAA X1X2GSQHALQALLDKGPALRQAG     (IX)

[0420] FGNDNLVKVAA X1X2GGAQALQALLDKGPALRQAG     (X)

[0421] FGNDNLVKVAA X1X2GSQQALQALLDKGPALRQAG     (XI)

[0422] FGNDNLVKVAA X1X2GGAQALQALLDRGPALRQAG     (XII)

[0423] FSNDNLVKVAA X1X2GGAHALQALLDKGPALRQAG     (XIII)

[0424] FSNDNLVKVAA X1X2GGQQALQTLLDKGPALRQAG     (XIV)

[0425] FTHQQIVAIAS X1X2GGSQALNTVLATHAALTAAG     (XV)

[0426] FTHQQIVAIAS X1X2GGSQALDKVLATHAPLTAAG     (XVI)

[0427] FTHRQIVGIAS X1X2GGSQALDTVLVRYAPLRDAG     (XVII)

[0428] FKHEQIVGIAS X1X2GGSQALDKVLATHAQLTAVG     (XVIII)

[0429] FKHEQIVAIAS X1X2GGSQALDKVLVKYAPLTAAG     (XIX)

[0430] FTHQQIVAIAS X1X2GGSQALDTVLATHAQLTTAG     (XX)

[0431] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG     (XXI)

[0432] GSREQVIKIAAX1X2GGQQALQALLDKGPALRNAG     (XXIII)

[0433] FSNDNLVRIGGX1X2GAKKTLDTLLQVYPQLTQGG     (XXIV)

[0434] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXV)

[0435] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXVI)

[0436] FSPDNLIKVAAX1X2GGAQALQALLDKSPALRQAG     (XXVII)

[0437] FGPDNLVKVAAX1X2GGQQALQALLDKGPALRQAG    (XXVIII)

[0438] FGPDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXIX)

[0439] FGPDNLVKVAAX1X2GGAQALQALLDKGPTLRQAG    (XXX)

[0440] FSPDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXXI)

[0441] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXXII)

[0442] FGNDNLVKVAAX1X2GGQQALQALLDKGPALRNAG    (XXXIII)

[0443] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRNAG    (XXXIV)

[0444] FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG    (XXXV)

[0445] where X1X2 are RVD.

[0446] In some embodiments, each of the tandem repeats comprises an amino acid sequence selected from the group consisting of the Formulae I, III and IV.

[0447] In some embodiments, each of the tandem repeats comprises an amino acid sequence selected from the group consisting of the Formulae II and V to VII.

[0448] In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae XXII to XXIV.

[0449] In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae III, IV and VIII to XIV.

[0450] In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae V to VII and XV to XX.

[0451] In some embodiments, each of the TRs comprises an amino acid sequence selected from the group consisting of the Formulae XXIII to XXXV.

[0452] In some embodiments, the polypeptide comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more TRs. Preferably, the polypeptide comprises 6-8 TRs.

[0453] In some embodiments, the programmed / programmable TALE-like polypeptide comprises an amino acid sequence of SEQ ID NO: 25, 26 or 28,

[0454] SEQ ID NO: 25

[0455] SEQ ID NO: 26

[0456] SEQ ID NO: 28

[0457] wherein X1X2 are RVDs selected from the group consisting of HD (recognizing C) , NN (recognizing A and G) , NG (recognizing T) , NS / NI (recognizing A) , NK (recognizing G) , KS / KG (recognizing T) , HI (recognizing A and G) .

[0458] 7. Method for screening DNA-binding TRPs

[0459] The present disclosure provides a method for screening DNA-binding TRPs, comprising

[0460] i) retrieving proteins in TR families from a database;

[0461] ii) predicting TRs in the proteins and screening the TRPs; and

[0462] iii) establishing a model and analyzing the screened TRPs with the model to obtain TRPs that are predicted to be capable of binding to DNA.

[0463] The database in step i) can be UniProtKB / Swiss-Prot. The predicting of step ii) can be conducted with XSTREAM and using the repeat region sequence as queries to comprehensively search all repeat regions of TRs. TR clusters can be obtained using MMseqs25.

[0464] In step iii) , the model can be established by in integrating two or more pretrained transformer models, such as ProteinBERT, ProtTans and ESM, and then, training the integrated model with datasets generated from a database such as UniProtKB / Swiss-Prot. For example, the process to generate the positive dataset were as follows: 1) Retrieval of all DNA binding proteins (DBPs) from UniProtKB / Swiss-Prot through a keyword search for “DNA binding” ; 2) Filtering proteins without GO term containing “DNA binding” ; 3) Elimination of sequences that had less than 60 amino acids (aa) or greater than 4000 aa 4) Elimination of sequences that contain “X|x” and “J|j” characters; 5) Grouping the remaining sequences that shared a sequence similarity of ≥ 50%by MMseqs2. The non-DNA binding proteins (NDBPs) dataset was generated by the following steps: 1) Retrieval of sequences that had a < 25%sequence similarity to any sequences in positive dataset. 2) Filtering proteins with any description that containing “DNA binding” . 3) Elimination of sequences that contain “X|x” and “J|j” characters; 4) Grouping the remaining sequences that shared a sequence similarity of ≥ 50%by MMseqs2. Totally, 12, 989 DBPs and 12, 1455 NDBPs were generated. To balance the number of DBPs, 12, 989 NDBPs were randomly selected for training. The resulting dataset was denoted as UniSwiss25978 dataset.

[0465] The method may further comprise a step of generating test datasets for testing the established model.

[0466] In some embodiments, the method further comprises:

[0467] iv) testing the obtained TRPs for binding to DNA.

[0468] In some embodiments, the binding can be tested with conventional assays, such as an assay using B1H system, BLI system and / or SELEX system.

[0469] The present disclosure further provides a device for performing the method for screening DNA-binding TRPs of the present disclosure, comprising

[0470] i) a unit for retrieving proteins in TR families from a database;

[0471] ii) a unit for predicting TRs in the proteins and screening the TRPs; and

[0472] iii) a unit for establishing a model and analyzing the screened TRPs with the model to obtain TRPs that are predicted to be capable of binding to DNA.

[0473] The present disclosure further provides a device for performing the method for screening DNA-binding TRPs of the present disclosure, comprising

[0474] a) a processor,

[0475] b) a storer coupled to the processor, and

[0476] c) a computer program stored in the storer,

[0477] wherein the processor executes the computer program to carry out the method for screening DNA-binding TRPs of the present disclosure.

[0478] The present disclosure provides a computer readable storage medium storing an executable instruction which makes a processor to execute the method for screening DNA-binding TRPs of the present disclosure.

[0479] The present disclosure provides a computer program product comprising a computer program which is executed by a processor to carry out the method for screening DNA-binding TRPs of the present disclosure.

[0480] Examples

[0481] Example 1. Material and Methods

[0482] 1.1. Computational pipeline for characterizing repeat proteins

[0483] All unique proteins were downloaded from UniRef100 and NCBI-nr database (frozen in August 2022) . A three-step pipeline was implemented to characterize proteins with tandem repeats (TRs) .

[0484] Firstly, proteins containing periodic repeats were identified by XSTREAM which utilizes the short seed extension method to detect protein TRs (Newman and Cooper, 2007, XSTREAM: a practical algorithm for identification and architecture modeling of tandem repeats in protein sequences. BMC Bioinformatics, 8, 382) .

[0485] Secondly, to classify novel and known TRs, we first collected the well-known TR families from literatures and retrieved curated proteins within these families in the UniProtKB / Swiss-Prot database (Bairoch and Apweiler, 2000, The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000. Nucleic acids research, 28, 45-48; Chakrabarty and Parekh, 2022, DbStRiPs: Database of structural repeats in proteins. Protein Science, 31, 23-36; Andrade et al., 2001, Protein repeats: structures, functions, and evolution. Journal of structural biology, 134, 117-131; and Kamel et al., 2021, REP2: A Web Server to Detect Common Tandem Repeats in Protein Sequences. Journal of Molecular Biology, 433, 166895) .

[0486] Next, we extracted repeat region sequences from these known TR proteins. Subsequently, these sequences were employed as queries for a comprehensive search across all repeat regions of the TRs identified in the initial step. Hits exhibiting at least 30%identity and 70%coverage were designated as putative known TRs and were further validated through domain annotation. Finally, both known and novel TR clusters were obtained using MMseqs2 (Steinegger and Soding, 2017, MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nature Biotechnology, 35, 1026-1028) with the following parameters: -c 0.7 --min-seq-id 0.3 --cov-mode 1 --cluster-mode 1. Considering that repeats can occur in non-integer multiples and their boundaries often do not coincide, we utilized two instances of the consensus sequence for the clustering process.

[0487] To investigate the underlying reasons for the significant prevalence of orphan TRs in novel type, we performed the following analysis. Firstly, we analyzed the count of genus-level assemblies through a two-step process: 1) identifying the originating species and their corresponding genus. 2) using the Entrez Direct E-utilities’ esearch commands to retrieve the count of genus-level assemblies (Kans, 2023, Entrez programming utilities help [Internet] . National Center for Biotechnology Information (US) ) . Secondly, we assessed the completeness of genome assemblies to which TRs belong. Specifically, for TRs within each cluster size range, we retrieved the corresponding genome assemblies and then randomly selected 2,000 of them for assessing genome completeness using BUSCO ( et al., 2015, BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics, 31, 3210-3212) .

[0488] 1.2. Generation of protein identifier

[0489] We developed a hierarchical naming scheme for the “TRs of interest” dataset, which incorporates three levels of classification: period, cluster size and unit number. Specifically, 1) all TRs were initially categorized into 21 major groups based on their period length (15, 16, ..., 35) . 2) Within each period length group, proteins were further sorted by cluster size (C1 to C100) . 3) In cases where two clusters shared identical period lengths and cluster sizes, their ranking was then determined by the median of unit number. Taking the unique identifier “TR_15_C1_1” as an example, TR represents tandem repeat, 15 signifies the period length of 15, C1 indicates “Cluster 1” , and 1 represents the TR’s position within its respective cluster, determined by sorting the unit numbers in ascending order.

[0490] 1.3. Correlation analysis

[0491] The correlation analysis in this study was calculated using Pearson correlation coefficients and was performed in the RStudio environment (Allaire, 2012, RStudio: integrated development environment for R. Boston, MA, 770, 165-171) .

[0492] 1.4. Evaluation of sequence complexity

[0493] Sequence complexity within a repeat was represented by normalized Shannon’s entropy score (NSS) (Sander and Schneider, 1991 Database of homology-derived protein structures and the structural meaning of sequence alignment. Proteins, 9, 56-68; and Shannon, 1948, A mathematical theory of communication. The Bell system technical journal, 27, 379-423) . Consensus sequences of TRs were used for Shannon score calculating. Specifically, the Shannon entropy score is defined as the negative of a sum of the products of amino acid frequencies in a typical repeat sequence (Pi) and binary logarithms of those frequencies (log2 (Pi) ) . The calculated Shannon score was then normalized by the length of consensus sequences.

[0494] 1.5. Identity cutoff selection

[0495] To determine a valid identity threshold for distinguishing different TR families, we collected 100 members of eight well-known TRP families, namely ZNF, ANK, ARM, LRR, PPR, TPR, WD40 and TALE. Inter-and intra-repeats represent repeats between different TRP families and within the same TRP family, respectively. Identity values between inter-and intra-repeats were calculated by using the EMBOSS needle program with default parameters (Rice et al., 2000, EMBOSS: the European molecular biology open software suite. Trends in genetics, 16, 276-277) . Two copies of consensus sequences were used as input.

[0496] 1.6. Training datasets generation

[0497] A training dataset was constructed to train the proposed predictor based on the UniProtKB / Swiss-Prot database (Bairoch and Apweiler, 2000, The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000. Nucleic acids research, 28, 45-48) . The positive dataset was generated as follows: 1) Retrieval of all DBPs from UniProtKB / Swiss-Prot through a keyword search for “DNA binding” ; 2) Filtering proteins without a GO term containing “DNA binding” ; 3) Elimination of sequences that are < 60 amino acids (aa) or >4,000 aa; 4) Elimination of sequences that contain “X|x” and “J|j” characters; 5) Grouping the remaining sequences that shared a sequence similarity of ≥ 50%by MMseqs2. The non-DNA binding proteins (NDBPs) dataset was generated by the following steps: 1) Retrieval of sequences that had a < 25%sequence similarity to any sequences in positive dataset. 2) Filtering proteins with any description that contains “DNA binding” . 3) Elimination of sequences that contain “X|x” and “J|j” characters; 4) Grouping the remaining sequences that share a sequence similarity of ≥ 50%by MMseqs2. Totally, 12, 989 DBPs and 121, 455 NDBPs were generated. To balance the number of DBPs, 12, 989 NDBPs were randomly selected for training, denoting as UniSwiss25978 dataset.

[0498] 1.7. Test datasets generation

[0499] An independent test dataset was procured from the Protein Data Bank (PDB) (Sussman et al., 1998, Protein Data Bank (PDB) : database of three-dimensional structural information of biological macromolecules. Acta Crystallographica Section D: Biological Crystallography, 54, 1078-1084) which satisfied the following criteria similar as the training dataset generation, except for the following additions: 1) Filtering out sequences that had a ≥ 25%sequence identity to any other sequences in the training datasets utilized by the evaluated tools; 2) Grouping the remaining sequences that shared a sequence identity of ≥ 25%by MMseqs2. A total of 359 representative DBPs and 364 representative NDBPs were generated after fulfilling the filtration criteria. Subsequently, 300 DBPs and NDBPs were randomly chosen, resulting in the PDB600 dataset. The performance of each tool was evaluated using several parameters, including Accuracy, Specificity, Recall, Precision, F1 score, and Area under the Curve (AUC) . The Receiver operating characteristic (ROC) curves were plotted by using the R “pROC” package (Robin et al., 2011, pROC: an open-source package for R and S+ to analyze and compare ROC curves. BMC bioinformatics, 12, 1-8) .

[0500] 1.8. DNA binding protein prediction model construction

[0501] The DNA binding protein prediction model, PLM-DBPPred, was constructed by integrating the ProteinBERT, ProtTans and ESM pre-trained transformer models (Brandes et al., 2022, ProteinBERT: a universal deep-learning model of protein sequence and function. Bioinformatics, 38, 2102-2110; Elnaggar et al., 2021, Prottrans: Toward understanding the language of life through self-supervised learning. IEEE transactions on pattern analysis and machine intelligence, 44, 7112-7127; and Rives et al., 2021, Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National Academy of Sciences, 118, e2016239118) . More specifically, the predicted results from these three models are obtained independently, and the average of the three probabilities is calculated as the final result. The cutoff was set as 0.5.

[0502] ProteinBERT

[0503] ProteinBERT is pre-trained in a self-supervised manner using ~106 million protein sequences from the UniRef90 database, along with their corresponding functional annotations from the gene ontology database. Initially, the model takes protein sequences as input and processes them through a feature encoding step. The resulting embedding vectors are then subjected to attention calculations through a dedicated attention layer. To reduce feature dimensionality, the output of the attention layer is transformed via a fully connected layer, followed by a sigmoid activation function. Furthermore, additional layers of attention and fully connected components are consecutively stacked after the initial one. In addition to employ a two-layer attention for classifier, we also utilized fully connected layers for comparison. The final output of this model is a probability score that signifies whether a given protein exhibits DNA-binding capabilities or not.

[0504] In the model training process, all layers of the pre-trained model were initially frozen, with the exception of the attention and fully connected layers, which were trained for up to 10 epochs. Subsequently, all layers were unfrozen and trained for an additional three epochs. To ensure optimal learning, the dynamic learning rate adjustment technique, ReduceLROnPlateau, was employed. The loss was calculated using binary cross-entropy, and an early stopping policy was applied to prevent overfitting. Utilizing the pre-trained model, the entire training procedure was completed on a single GPU, Tesla P100-PCIE-12GB, with the learning rate and batch size set to 0.0001 and 32, respectively.

[0505] ProtTrans

[0506] For ProtTrans-based architectures, six models were trained by combining two architectures from ProtTrans (ProtT5-XL-UniRef50 and ProtBERT-BFD) with three different classifiers. The ProtT5-XL-UniRef50 is a variant of the T5 model (Raffel et al., 2020, Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21, 5485-5551) , designed as an encoder-only architecture. It was initially trained on the BFD dataset and then fine-tuned on the UniRef50 dataset. The ProtBERT-BFD is based on the BERT architecture and trained on the BFD dataset. Following the PLMs, three classifiers were used to process the embedding information. The first one is a vanilla MLP, serving as a baseline for classification performance. The second one is the Light Attention (LA) classifier, which employs a pair of 1D convolution to extract pivotal information ( et al., 2021, Light attention predicts protein location from the language of life. Bioinformatics Advances, 1, vbab035) . The third classifier is biLSTM_TextCNN, which is known for its effectiveness in sentiment classification (Jiang et al., 2022, Research on sentiment classification for netizens based on the BERT-BiLSTM-TextCNN model. PeerJ Computer Science, 8, e1005) . Finally, probabilities signify whether a given protein exhibits DNA-binding capabilities or not were obtained.

[0507] In the model training phase, all layers of the pre-trained ProtT5-XL-UniRef50 model were kept frozen. The classifier was trained for a duration of 10 epochs. The batch size was set to 64. The learning rate was initialized at 0.00005 and optimization was performed using the Adam with a weight decay of 0.001. Similar to the operations in ProteinBERT, the learning rate was dynamically adjusted based on validation set performance using the ReduceLROnPlateau. The entire fine-tuning procedure was completed utilizing a single GPU on Tesla P100-PCIE-12GB.

[0508] ESM

[0509] As for ESM-based models, we select the esm2_t30_150M_UR50D and esm2_t33_650M_UR50D model for encoding. After obtaining protein sequence embeddings through the processing phase, these embeddings are subsequently directed to the MLP, LA and biLSTM_TextCNN classifier, following the methodology outlined in ProtTrans.

[0510] During the model training process, we performed fine-tuning by unfreezing the embedding norm before layers, embedding norm after layers, and the Roberta head for masked language modeling layers within the pre-trained ESM models, in combination with three different classifiers. The training spanned 10 epochs, with the batch size set to 120 and the learning rate initialized at 0.01. Optimization was performed using the AdamW with a weight decay of 0.0001. We also employed the ReduceLROnPlateau to dynamically adjust the learning rate based on validation set results. The entire fine-tuning procedure was completed on 8 Tesla P100-PCIE-12GB GPUs.

[0511] 1.9. Gene ontology (GO) enrichment analysis of DBPs and NDBPs

[0512] The dataset used for DBP GO enrichment were generated similar to the test dataset used for evaluating the DBP prediction tool, with the exception that sequences sharing similarity with any sequences in the training datasets used by the evaluated tools were not removed. PANNZER2 ( et al., 2018, PANNZER2: a rapid functional annotation web server. Nucleic acids research, 46, W84-W88) was applied for GO term identification and terms with a PPV greater than 0.6 were subsequently submitted to clusterProfiler for GO enrichment (Wu et al., 2021, clusterProfiler 4.0: A universal enrichment tool for interpreting omics data. The Innovation, 2, 100141) .

[0513] 1.10. DNA binding domains (DBDs) investigation and protein domain annotation

[0514] DNA binding domains (DBDs) were obtained from various sources, including databases such as AnimalTFDB 4.0 (Shen et al., 2023, AnimalTFDB 4.0: a comprehensive animal transcription factor database updated with variation and expression annotations. Nucleic Acids Research, 51, D39-D45) , PlantTFDB 5.0 (Tian et al., 2020, PlantRegMap: charting functional regulatory maps in plants. Nucleic acids research, 48, D1104-D1113) and the DNA-binding domain (DBD) database (Wilson et al., 2008, DBD––taxonomically broad transcription factor predictions: new content and functionality. Nucleic acids research, 36, D88-D92) . Additionally, a manual keyword search was conducted in the Pfam database for terms such as “DNA binding” and “transcription factor” to ensure comprehensive coverage. Partner domains aside from the DBDs were extracted based on the domain architecture file from the Pfam FTP site. The top 10 enriched partner domains of each DBD were collected and GO terms derived from Pfam2GO were used for conducting GO enrichment (http:  / / www. geneontology. org / external2go / pfam2go) . DNA-related domain (DRD) accessions were extracted and manually verified according to the following keywords: DNA binding, RNA binding, nucleic acid binding, transcription factor, nuclease, helicase, deaminase, integrase, ligase, transposase, polymerase, methylase, recombinase. Proteins annotated with any DNA related domains were added into DRD list.

[0515] All repeat proteins were scanned for functional domain entries in the Pfam database (V35.1) by using hmmsearch (Finn et al., 2011, HMMER web server: interactive sequence similarity searching. Nucleic acids research, 39, W29-W37) with the “--cut_ga” option.

[0516] 1.11. Enrichment analysis of different function annotation

[0517] Enrichment analysis of different function was referred to the GO enrichment (Zheng and Wang, 2008, GOEAST: a web-based software toolkit for Gene Ontology enrichment analysis. Nucleic acids research, 36, W358-W363) . Take DBP enrichment as example, where A is the count of DBPs in clusterX, B is total number of proteins in clusterX, C is the count of DBPs outside of clusterX, and D is total number of proteins outside of clusterX. The odds ratio was calculated as (A / B)  /  (C / D) for a specific function annotation. Subsequently, the log2 odds ratio (LR) was determined to represent the enrichment score of a particular function for each cluster (Zheng and Wang, 2008 above) . A larger LR indicates a higher level of enrichment of the function within the cluster compared to the overall sample, and vice versa.

[0518] 1.12. Protein structure predictions

[0519] Protein structure model predictions were conducted by using the ColabFold v1.5.2-patch platform with default parameters (https:  / / colab. research. google. com / github / sokrypton / ColabFold / blob / main / AlphaFold2. ipynb) (Mirdita et al., 2022, ColabFold: making protein folding accessible to all. Nature methods, 19, 679-682) .

[0520] 1.13. Construction of phylogenetic trees

[0521] Repeat region sequences were used to construct the phylogenetic tree for all TRP family, with the exception of MOON and STAR families, which were generated using full length sequences. All phylogenetic trees were constructed by FastTree program (Price et al., 2010, FastTree 2–approximately maximum-likelihood trees for large alignments. PloS one, 5, e9490) .

[0522] 1.14. Transcription activation domain prediction

[0523] The transcription activation domain was predicted by ADpred (Erijman et al., 2020, A high-throughput screen for transcription activation domains reveals their sequence features and permits prediction by deep learning. Molecular cell, 78, 890-902. e896) .

[0524] 1.15. Type III secretion signal prediction

[0525] The Type III secretion signal (T3SS) was predicted by EffectiveDB software suite (http:  / / effectivedb. org) (Eichinger et al., 2016, EffectiveDB-updates and novel features for a better annotation of bacterial secreted proteins and Type III, IV, VI secretion systems. Nucleic acids research, 44, D669-D674) . The cutoff was set as 0.99.

[0526] 1.16. Structure comparison

[0527] Structural comparison was conducted through the computation of the root mean-square deviation (RMSD) matrix, utilizing the R “Bio3d” package (Grant et al., 2021, The Bio3D packages for structural bioinformatics. Protein Sci, 30, 20-30) . The resulting matrix was subsequently visualized by the headmap. 2 function in R “gplots” package (Warnes et al., 2009, gplots: Various R programming tools for plotting data. R package version, 2, 1) .

[0528] 1.17. Target gene analysis

[0529] Putative target genes of identified DNA-binding TRPs were determined through the following steps: 1) Extracting the promoter sequence (-1,000 bp to the TSS site) . 2) Searching the promoter region for enriched motif by using the program FIMO in the MEME Suite. 3) Extracting the putative target genes from the target hits. 4) Performing GO annotation for all proteins in genome by PANNZER2 ( et al., 2018 above) . 5) Performing GO enrichment for potential target genes by clusterProfiler (Wu et al., 2021 above) .

[0530] 1.18. Taxonomic breadth of TR clusters

[0531] For estimation of the taxonomic breadth of each TR cluster, we calculated the last common ancestor (LCA) of their members using Taxonkit v0.13.0 (Shen and Ren, 2021, TaxonKit: A practical and efficient NCBI taxonomy toolkit. Journal of genetics and genomics, 48, 844-850) . Specifically, we first clustered all identified 4, 575, 091 TRs by using MMseqs2 (Steinegger and Soding, 2017 above) with the following parameters: -c 0.7 --min-seq-id 0.3 --cov-mode 1 --cluster-mode 1. Subsequently, the computation of the LCA was executed for all TR cluster containing a minimum of 10 members.

[0532] 1.19. Protein expression and purification

[0533] The pET28a vector in frame with the N-terminal His tag and SUMO tag were used to express and purify protein in vitro. Expression constructs for all candidates were synthetized by BGI after codon optimization for E. coli. The assembled genes were constructed into a pET28a expression vectors and then transformed into Rosetta E. coli cells. For protein expression, 0.1 mM IPTG were added when OD600 reached to 0.6-0.8 and then incubated at 16℃ for 18 h. Cell pellets were resuspended in binding buffer (50 mM Tris-HCl, 500 mM NaCl) before lysis by sonication (200W, 3s on / 3s off on ice for 10 min) . The supernatant was loaded onto a HisTrap HP column (GE Healthcare) after column washing with binding buffer with 5 column volumes. The protein was eluted with binding buffer containing 300 mM imidazole. Molecular sieve chromatography with Superose 6 or Superose 12 HR16 / 50 was then performed to obtain protein of high purify.

[0534] 1.20. Bio-Layer Interferometry (BLI) -based screen

[0535] The Bio-Layer Interferometry (BLI) -based screening method was developed based on the findings of Marklund et al. that DNA-binding proteins rapidly associate and dissociate across various sequences but efficiently rebind to the target sequence via searching, leading to macroscopic specific binding (Marklund et al., 2022, Sequence specificity in DNA binding is mainly governed by association. Science, 375, 442-445) . Accordingly, the association of a DBP with a random dsDNA library could occur, and the binding signal could be reflected by the response captured by the BLI experiment.

[0536] Feasibility testing of BLI-based screening began with comparisons of responses between well-established DBPs and NDBPs. The DBPs included 1) T_AAVS1, a TALE designed to target the AAVS1 locus (Hockemeyer et al., 2011, Genetic engineering of human pluripotent cells using TALE nucleases. Nature biotechnology, 29, 731-734) , and 2) Zif268, a natural zinc-finger protein (Christy and Nathans, 1989, DNA binding site of the growth factor-inducible protein Zif268. Proceedings of the National Academy of Sciences, 86, 8737-8741) . The NDBPs included 1) PUM1, the human Pumilio homolog 1 protein (Spassov and Jurecic, 2002, Cloning and comparative sequence analysis of PUM1 and PUM2 genes, human members of the Pumilio family of RNA-binding proteins. Gene, 299, 195-204) , 2) SUMO, a solubility tag (Kim et al., 2002, Versatile protein tag, SUMO: its enzymology and biological function. Journal of cellular physiology, 191, 257-268) , and 3) ULP1, ubiquitin-like-specific protease 1. Two different loading densities were tested (1 or 5 nm) . The experimental protocol was set as follows: BLI-based screen was performed on the Octet RED system (Fortebio) at 25 ℃. The Ni-NTA (NTA) biosensors were dipped in the BLI assay buffer (PBS, 0.02%Tween, pH7.4) for 10 min before using. The 88-bp dsDNA library with 60-nt randomized regions were annealed and diluted to100 nM. The dsDNA product was purified by PAGE-gel extraction. Three replicate series were used for each protein. The procedures steps included: baseline for 60s, loading His-tagged protein, another baseline for 60s, association for 60s, dissociation for 30s and regeneration for 30s. The response scores were collected and used for further data processing.

[0537] 1.21. SELEX screen

[0538] SELEX assay was performed according to several previous studies (Bouvet, 2009, DNA-Protein Interactions. Springer, pp. 139-150; and Miller et al., 2011, A TALE nuclease architecture for efficient genome editing. Nature biotechnology, 29, 143-148) . The dsDNA library contains 20-nt random sequences. For each round of selection, 200 ng purified protein with 3 μL DynabeadsTM His-Tag Isolation and Pulldown beads (Thermo, #10103D) were incubated for 30 min at 25 ℃. After the unbounded protein removal, 2 μg dsDNA library was added to protein-beads complex and incubated with 100 μL SELEX buffer (50 mM Tris, 150 mM NaCl, 20 mM KCl, 2.5 mM MgCl2, 10 μM ZnCl2, 0.05 %Tween20, 0.01 %BSA, 20 μg / mL dIdC) . After 1h incubation at room temperature, unbounded dsDNA was washed away with SELEX buffer five times. Bounded dsDNA was then amplified for the next round selection. After five cycles, recovered DNA fragments were cloned and sequenced.

[0539] 1.22. B1H screen

[0540] The plasmids for B1H system were purchased from Addgene (#12609 as the reporter plasmid and #18039 for the expression of the protein to be tested) . The 18-nt random sequences library reporter of B1H screening were constructed by T4 ligase with the sticky ends being generated by NotI and EcoRI. To conduct the B1H assay, a reporter vector with 18-nt random sequences library and a protein expression vector were co-transformed into US0 E. coli strain by electroporation. The successful binding events were enriched on plates with 10 mM 3-AT. The enriched cells were scraped from the plates, followed by plasmid extraction and sequencing.

[0541] 1.23. GFP activation validation

[0542] The GFP activation validation system was constructed based on the B1H screen system. Specifically, the HIS3 marker gene was substituted with the GFP reporter gene (Fig. 3D) . Moreover, the 18-nt random sequence upstream of the promoter was changed to the enriched motif of interest. Subsequently, the binding affinity of tested motifs is indicated by the GFP signal which could be detected by flow cytometry.

[0543] 1.24. GFP repression validation

[0544] The GFP repression system was the modified version of the PAM-SCANR system which was first established to identify functional PAM diversity across CRISPR-Cas systems (Leenay et al., 2016, Identifying and visualizing functional PAM diversity across CRISPR-Cas systems. Molecular cell, 62, 137-147) . Specifically, the lacI and promoter of lacI were deleted and the PAM sequence on reporter plasmid was substituted into the enriched motif of interest (Fig. 3D) . After the co-transformation of protein expression plasmid and the reporter plasmid, the successful binding of tested motifs with protein can block the expression of GFP, leading to a decreased level of GFP signal detected by flow cytometry.

[0545] 1.25. EMSA

[0546] For EMSA assay, both FAM labeled or unlabeled probes were generated by annealing oligos. Different protein concentration series were designed for each reaction with the last two lanes added with specific / non-specific unlabeled probes, which set as binding competitors. Binding reactions were incubated at room temperature for 30 min and resolved on a 6%native PAGE gel for 1 h at 80 V. The gels were visualized with a UV transilluminator (Bio-Rad) .

[0547] 1.26. Negative staining sample preparation, data collection and 2D classification average Individual proteins were purified following the procedures described above. All the complexes were reconstituted by incubating protein and DNA in a 1: 10 ratio on ice for 30 minutes in a binding buffer containing 10 mM Tris-HCl (pH 8.0) and 150 mM NaCl.

[0548] For the preparation of negative staining samples, all specimens were diluted to a final concentration of 0.5 μM and subjected to negative staining in a 2% (w / v) uranyl acetate solution, following the standard deep-stain protocol on holey-carbon coated EM copper grids covered with a thin layer of continuous carbon (Liu and Wang, 2011, Single particle electron microscopy reconstruction of the exosome complex using the random conical tilt method. JoVE (Journal of Visualized Experiments) , e2574) . The negatively stained specimens were subsequently imaged using a FEI Tecnai-F20 electron microscope operating at an acceleration voltage of 200 kV. Images were captured at a nominal magnification of 50,000x, with a defocus range of 2.5 to 3.5 μm.These electron micrographs were recorded using a Gatan Ultrascan4000 4k × 4k CCD camera.

[0549] Subsequently, the acquired micrographs were processed as negative stain data in CryoSPARC (Punjani et al., 2017, cryoSPARC: algorithms for rapid unsupervised cryo-EM structure determination. Nature methods, 14, 290-296) . Manual Picker was utilized to select individual particles, and 2D reference-free classification was employed to determine the average particle size.

[0550] 1.27. BLI for binding kinetics

[0551] BLI experiments were performed on the Octet RED system (Fortebio) at 25 ℃. The streptavidin (SA) sensors were dipped in the assay buffer (PBS, 0.02%Tween, pH7.4) for 10 min before use. Complementary pairs of labelled oligonucleotides (5’ biotin) were annealed in 10 × annealing buffer (Thermo, TECH TIP #45) . The dsDNA product was purified by PAGE-gel extraction. A six-point concentration series were designed for each protein with the last one diluted with PBST buffer, which set as reference. The procedures were set as follows: baseline for 60 s, loading the biotin-conjugated dsDNA for 120 s, another baseline for 120 s, association for 120 s, and dissociation for 120 s. Dissociation (kdis) and association rate constants (kon) were determined with the Octet Data Analysis Software, as a result of a global fit considering the entire step times, and assuming a 1: 1 binding model.

[0552] 1.28. Assembly of artificial STAR proteins

[0553] The one-step construction of artificial STAR proteins followed the previously reported Golden Gate method, with several refinements (Cermak et al., 2011, Efficient design and assembly of custom TALEN and other TAL effector-based constructs for DNA targeting. Nucleic acids research, 39, e82-e82) . Specifically, the repeat module on pHD-1 was replaced with modules derived from PqSTAR1, each possessing a unique overhang. Additionally, we implemented modifications to the pFUS_A vector by inserting the N-and C-terminal regions of PqSTAR1 at both ends of LacZ, respectively. Two internal BsaI sites were strategically placed at the 5’ and 3’ termini of LacZ to facilitate vector linearization with the enzyme, resulting in the generation of suitable overhangs for the incorporation of the repeat modules.

[0554] 1.29. Cell lines

[0555] Human 293T cells were cultured in DMEM medium containing 10%fetal bovine serum (Invitrogen, Carlsbad, USA) with 1%penicillin / streptomycin (Millipore, TMS-AB2-C) . All cells were incubated in a humidified incubator at 37 ℃ with 5%carbon dioxide.

[0556] 1.30. CUT&Tag experimental procedure

[0557] Plasmids vectors for protein expression in human 293T cells were constructed with an Ef1a promoter and a 3X FLAG tag at C-terminal. A P2A-GFP sequence following the FLAG tag was used to detect the plasmid transfection efficiency. The 293T cells were seeded into 24-well plates at a density of 1.2 × 105 cells / well. For transfection, 1 μg of plasmid was used with Lipofectamine2000 (Invitrogen) . Transfected cells were then cultured at 37 ℃ for another 24 h before conducting the Western Blot and CUT&Tag assays. CUT&Tag assay was performed with CUT&Tag Assay Kit (TD903, Vazyme Biotech) following the manufacturer’s instructions.

[0558] 1.31. RNA-seq experimental procedure

[0559] Plasmids vectors for STAR-based ATFs expression in human 293T cells were constructed with a CMV promoter in pVAX1 (Snapgene) . A P2A-GFP sequence following the VPR activator domain was used to detect the plasmid transfection efficiency. The 293T cells were seeded into 6-well plates at a density of 8 × 105 cells / well. For transfection, 2.5 μg of plasmid was used with Lipofectamine2000 (Invitrogen) . Transfected cells were then cultured at 37 ℃ for another 48 h before RNA extraction. Total RNA was extracted by using TRIzol regent and sequencing was done by the Annoroad Gene Tech. (Beijing) Co., Ltd.

[0560] 1.32. Sequencing and data processing

[0561] SELEX, B1H and CUT&TAG samples were sequenced on the Illumina NovaSeq PE150 platform. RNA-seq samples were sequenced on the MGIDNBSEQ T7 platform. Read trimming and filtering were performed using Trimmomatic version 0.33 (Bolger et al., 2014, Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics, 30, 2114-2120) .

[0562] For SELEX and B1H data, after separating samples by barcodes, random DNA library region of each sample was extracted by mapping the fixed flanking sequences.

[0563] For CUT&Tag data, paired-end reads were aligned to the hg19 genome assembly by using BWA-MEM (Li, 2013, Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv preprint arXiv: 1303.3997) . Then the peaks were called by MACS2.0 (Zhang et al., 2008, Model-based analysis of ChIP-Seq (MACS) . Genome biology, 9, 1-9) with the parameter of -p 0.01. The control library was generated by treating WT 293T cells with FLAG-tag antibody. Overlap analysis was performed using bedtools (Quinlan and Hall, 2010, BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics, 26, 841-842) . Enriched motifs for all types of data were discovered and visualized by HOMER program (Heinz et al., 2010, Simple combinations of lineage-determining transcription factors prime cis-regulatory elements required for macrophage and B cell identities. Mol Cell, 38, 576-589) .

[0564] For RNA-seq data, paired-end reads were aligned to the human hg19 genome by using hisat2 (Kim et al., 2019, Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nature biotechnology, 37, 907-915) . StringTie was used to generate gene counts which were fed into DESeq2 for expression analysis (Pertea et al., 2015, StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nature biotechnology, 33, 290-295; and Love et al., 2014, Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome biology, 15, 1-21) . Genes were considered differentially expressed for an adjusted p value < 0.05 and a log2 fold change > |1|. The promoter region between nucleotide positions -2,000 and +500 from the transcriptional start site (TSS) were used to perform motif enrichment. Gene set enrichment analysis was performed by using clusterProfiler R package in R studio (Wu et al., 2021 above) .

[0565] 1.33. Quantification and statistical analysis

[0566] Datasets were assessed using GraphPad Prism 8. All numerical values are presented as means ± s.d. Statistical significance for B1H validation analyses was determined using two-sided Student’s t-tests, while the comparison of RNA-seq read counts employed the Wilcoxon signed-rank test. For all situation, p value < 0.05 was considered statistically significant, p value < 0.01 was considered statistically extremely significant.

[0567] Example 2: Identification and characterization of tandem repeat proteins

[0568] This Example was conducted to identify and characterize proteins comprising tandem repeats (TRs) , which bind to DNA with sequence specificity.

[0569] As shown in Fig. 1A, proteins that containing periodic repeats were identified by XSTREAM with the default parameters (unit number ≥ 2 and period ≥ 3) by searching the UniRef100 and NCBI-nr databases and 43, 648, 764 TRs were identified. Upon aligning all repeat units, the assessment of the conservation of each column enables the derivation of variable amino acids (VAAs) (Fig. 1B) .

[0570] For the classification of known and novel TRs, we first collected the well-known TR families from literatures and retrieved curated proteins within these families in the UniProtKB / Swiss-Prot database.

[0571] TRs that exhibited the following characteristics were defined as low quality and excluded from further analysis: 1) incomplete TRs, 2) composite repeats and 3) TRs consisting of only two repeat units. In total, 4, 575, 091 TRs passed this filtration procedure, and the unit numbers and periods of these proteins were plotted (see Fig. 1C) .

[0572] Next, we predicted TRs within these proteins by XSTREAM and utilized the repeat region sequence as queries to comprehensively search all repeat regions of TRs identified in the initial step. The repeat region sequences of those proteins were utilized as queries to search for homologs among all identified TRs. Hits exhibiting at least 30%identity and 70%coverage were designated as putative known TRs and were further validated through domain annotation. Finally, both known and novel TR clusters were obtained using MMseqs2 with the following parameters: -c 0.7 --min-seq-id 0.3 --cov-mode 1 --cluster-mode 1. Considering that repeats can occur in non-integer multiples and their boundaries often do not coincide, the clustering process utilized two instances of the consensus sequences.

[0573] Impressively, among the 4, 575, 091 TRs, only 4.4% (199, 240) were categorized as known TRs, and the remaining 95.6% (4, 375, 851) were novel.

[0574] Next, we clustered the known and novel TRs using a 30%identity and 70%coverage cutoff (Material and Methods) , and they showed distinct patterns. As indicated in Fig. 1D, known TRs were grouped into a much smaller number of unique clusters, and their periods were concentrated between 15-55 aa. In comparison, the protein numbers and cluster numbers of novel TRs were much closer to each other (Fig. 1D) . On average, the number of proteins within each cluster was much smaller for novel TRs than for known TRs (Fig. 1E) . Notably, a significant proportion of the novel clusters (~71%) consisted of only one protein member, while no orphan TRs were found among the known TR clusters (Fig. 1E) . The significant prevalence of orphan TRs in novel types could be attributed to limited sequencing coverage of genomic diversity or poor genome assembly quality. To identify the primary reason, we quantified two indicators across different cluster sizes: (1) the count of genus-level assemblies accessible from NCBI for the originating species of each protein and (2) the completeness of the genome assembly to which the TRs belonged. As the cluster size increased, we observed a corresponding rise in the count of the median assembly at the genus level (Fig. 1F, PCC=0.97, p value < 0.01) . However, the pattern of assembly completeness exhibited only a weak trend (PCC=0.66, p value=0.2285) . These findings indicate that the high percentage of orphan TRs is primarily due to inadequate genome sequencing data from neighboring species, and enormously diverse novel TR families remain to be explored.

[0575] Taken together, these analyses show that only a tiny portion of TR diversity has been studied to date, and our current understanding of TRs is deeply biased toward a few well-studied families. A comprehensive exploration of TRs is needed.

[0576] Example 3: Development of the model for predicting the DNA binding

[0577] This Example was conducted to develop the model for predicting the DNA binding of the TR-containing proteins (TRPs) .

[0578] Initially, we employed several filtering criteria to characterize TRs with potential programmability in binding DNA sequences. Specifically, based on known nucleic acid binding families with important biological functions, we focused on repeats of 15-35 amino acids long that recur 6-40 times (15 ≤ period ≤ 35, 6 ≤ unit numbers ≤ 40) within a protein to enable proper folding with a reasonably complex binding interface. To ensure diversity within clusters, we excluded clusters if the repeat regions of all members were identical. To assess within-repeat diversity, the normalized Shannon’s entropy score (NSS) was introduced by calculating the amino acid diversity of the consensus sequence (Material and Methods) . In total, 125, 624 TRs with a defined period and unit number, as well as an NSS exceeding 0.6, passed the filtration procedure and were named “TRs of interest” . We developed a hierarchical naming scheme for the “TRs of interest” proteins that incorporates several levels of classification: period, cluster size and unit number.

[0579] For predicting DNA binding proteins (DBPs) , multiple computational techniques have been developed by incorporating diverse information sources such as sequences, structural features, and physicochemical properties. We conducted a comprehensive analysis of all available DBP prediction models developed since 2010 and found that 78%of the models were based on traditional machine learning (TML) , while 22%were based on deep learning (DL) . Furthermore, we found that TML-based models tended to have smaller training datasets compared to DL-based models. Notably, approximately 67%of TML models utilized the PDB1075 dataset (Liu et al., 2014, iDNA-Prot| dis: identifying DNA-binding proteins by incorporating amino acid distance-pairs and reduced alphabet profile into the general pseudo amino acid composition. PloS one, 9, e106691) . Therefore, our endeavors concentrated on two aspects: 1) generating high-quality datasets and 2) developing a robust model.

[0580] Firstly, we assembled a high-quality training dataset from the UniProtKB / Swiss-Prot database, encompassing 12, 989 DBPs and an equal number of non-DBPs (NDBPs) , collectively referred to as the UniSwiss25978 dataset (version 2022.10, Fig. 2A) . Additionally, an independent test dataset, named PDB600, was generated, consisting of 300 DBPs and 300 NDBPs sourced from the PDB database. Subsequently, we investigated three prominent PLMs: ProteinBERT (Brandes et al., 2022 above) , ProtTrans (Elnaggar et al., 2021 above) , and ESM (Rives et al., 2021 above) , each distinguished by unique architectures and pre-training on diverse databases. We integrated these PLMs with various classification layers. For the ProteinBERT-based model, we evaluated two architectures: one incorporating two attention layers, and another integrating only fully connected layers. For the ProtTrans-based and ESM-based models, we combined each with three distinct classifiers: MLP, light attention (LA) , and biLSTM_TextCNN. Finally, probabilities were obtained for each model through a feed-forward network (Material and Methods) . We fine-tuned these models using the UniSwiss25978 dataset and subsequently evaluated their performance using the PDB600 dataset.

[0581] Based on the test results, the optimal combination for each language model was selected, namely ProteinBERT with two attention layers, ProtT5 with LA classifier and ESM2_30 with LA classifier. ROC-AUC analysis illustrated robust predictive capabilities of all three models in identifying DBPs, and the ProtTrans-based model generated the best outcomes, with an impressive AUC of 0.88. To further improve model performance, we adopted an ensemble learning strategy to integrate all three models. Intriguingly, combining these models enhanced classification performance, resulting in an AUC score of 0.90 (Fig. 2B) . This newly developed transformer-based model combining three PLMs was designated PLM-DBPPred (protein language model-enhanced DNA binding protein prediction) (Fig. 2A) . Next, we compared it with existing state-of-the-art methods. PLM-DBPPred showed the best performance, demonstrating superior performance in AUC, accuracy and F1 score, with values of 0.90, 0.82, and 0.80, respectively (Fig. 2C) .

[0582] Using PLM-DBPPred, we predicted DBPs for the full-length proteins of the “TRs of interest” dataset, resulting in a candidate list that consists of 8, 865 TRPs, distributed across 1, 640 clusters, with potential DNA-binding capability.

[0583] Example 4: Screening of DNA binding TRPs

[0584] 4.1. Selecting TRP candidates

[0585] To further prioritize candidates with a higher likelihood of being DBPs, we conducted a survey of the well-studied DBPs to gain deeper insights into the biological processes in which they participate and their domain characteristics. We first performed Gene Ontology (GO) analysis using a dataset containing 1,000 DBPs and 1,000 NDBPs in the PDB database. As expected, DBPs were highly enriched in DNA-related metabolism and functions compared to NDBPs. According to the enriched GO terms, we traced them to higher-level hierarchy, including GO: 0090304 (nucleic acid metabolic process) , GO: 0003700 (DNA-binding transcription factor activity) , GO: 0003676 (nucleic acid binding) , GO: 0140640 (catalytic activity, acting on a nucleic acid) , and GO: 0005634 (nucleus) . TRPs with annotations to these terms and their subterms were categorized into the DNA-binding related GO (DBGO) list.

[0586] Considering that DBPs are often composed of DNA binding domains (DBDs) and specific functional domains, we conducted an analysis of the partner domains associated with all previously identified DBDs and found that they were enriched in nucleic acid metabolic processes and occurred in various forms, such as hydrolases, nucleases, and methyltransferases, and TRPs annotated with these domains were categorized into the DRD list.

[0587] Given that proteins grouped in clusters often share similar functional annotations and features, we established several selection strategies at the cluster level. For strategies based on functional annotation, we calculated an enrichment score for each functional annotation (see Material and Methods) . A cluster with a specific function was designated when the corresponding log2 odds ratio (LR) score exceeded 0. Accordingly, an overlap cluster between DBP, DBGO and DRD was generated (Fig. 2D) . We first selected 51 clusters with DBP function enrichment. We next selected 15 clusters with DRD / DBGO functional enrichment. In addition, considering that all of these functional annotations relied on existing knowledge, we further selected 9 clusters without any functional annotation and based solely on basic protein features such as small size. Due to the variations in cluster sizes and the difficulty of synthesizing genes containing tandem repeats, a variable number (1~3) of proteins within each cluster were chosen for gene synthesis. Overall, genes encoding 100 members distributed in 75 clusters were successfully synthesized. It is worth noting that in some cases when the protein contains DRDs, only gene fragment encoding the repeat region was synthesized.

[0588] 4.2. Experimental screening and validation

[0589] Considering the unique characteristics of different DBPs and their variable compatibility with specific functional assays, we employed both in vivo and in vitro screening strategies for identifying novel DNA-binding TRPs (Fig. 3A) .

[0590] For the in vivo platform, the well-established B1H (bacteria-one-hybrid) system was employed, which converts DNA binding events into the survival of bacteria. We cloned all 100 candidate genes into the B1H vector and performed screening as previously described (Meng et al., 2005, A bacterial one-hybrid system for determining the DNA-binding specificity of transcription factors. Nature biotechnology, 23, 988-994) . Among all 100 candidates, we successfully identified specific binding motifs for four different proteins (Figs. 3A, B) .

[0591] For the in vitro platform, we first assessed the expression level of each protein in E. coli and further purified the proteins with high expression. Out of the 100 candidates, we produced 28 proteins with high purity (Fig. 3A) . Subsequently, we developed a Bio-Layer Interferometry (BLI) -based screening method, validated the method by comparing the responses of well-known DBPs and NDBPs, and determined the appropriate experimental conditions (see Material and Methods) . Accordingly, we conducted BLI-based screen for 28 purified candidates, and identified eight proteins that displayed DNA binding activity under a 0.1 nm binding response (nm shift in BLI curve) cutoff. Next, we performed SELEX (systematic evolution of ligands by exponential enrichment) analysis (Riley et al., 2014, SELEX-seq: a method for characterizing the complete repertoire of binding site preferences for transcription factor complexes. Hox Genes: Methods and Protocols, 255-278) of these eight candidates to determine their binding specificity and successfully identified enriched DNA sequence motifs for two candidates (Figs. 3A, C) .

[0592] In total, we identified 11 proteins with DNA-binding activity, six of which displayed sequence specificity of DNA binding under the specific screening conditions (SEQ ID NOs: 1-6, Fig. 3A) , which were classified as families STAR (Short TALE-like Repeat proteins) , MOON (Marine Organism-Originated DNA binding protein) and pTERF (prokaryotic mTERF-like protein) .

[0593] To further validate the specific binding activity of these six positive hits obtained from the screen, we implemented four different validation assays (Fig. 3D) . To validate protein-DNA interactions in vitro, EMSAs (electrophoretic mobility shift assays) were used (Hellman and Fried, 2007, Electrophoretic mobility shift assay (EMSA) for detecting protein–nucleic acid interactions. Nature protocols, 2, 1849-1861) . Two other methods were modified from published studies (Leenay et al., 2016 and Meng et al., 2005 above) , in which protein binding to a specific DNA sequence induces or represses GFP signal in E. coli (see Material and Methods) . Additionally, CUT&Tag (Cleavage Under Targets and Tagmentation) technology was used to identify specific DNA binding in mammalian cells (Kaya-Okur et al., 2019, CUT&Tag for efficient epigenomic profiling of small samples and single cells. Nature communications, 10, 1930) . All six positives were validated in two or more independent assays and were therefore deemed true positives, and a customized name was given to each based on their features (Figs. 3B, C) , which will be described in later sections.

[0594] For each protein, we analyzed the sequence features, predicted the tertiary structure by Alphafold2 and identified homologs. We then characterized the common features for each family, including the verification of repeat boundaries and overall protein architecture, visualization of repeat units and secondary structure patterns, and construction of a phylogenetic tree. In certain cases, additional homologs were synthesized to further elucidate the binding properties of the family (Fig. 3D) .

[0595] Example 5: Characterization of the STAR family

[0596] This Example was conducted to characterize the STAR family (PqSTAR1 (SEQ ID NO: 1) and AspSTAR1 (SEQ ID NO: 2) ) .

[0597] 5.1. Characterization of STAR proteins

[0598] Through B1H screening, we identified DNA motif enrichment for PqSTAR1 and AspSTAR1 (Fig. 3B) . Subsequently, we confirmed their specific binding to the identified DNA sequences through EMSA and GFP activation assays (Fig. 4 A and B) . The binding affinities of PqSTAR1 and AspSTAR1 were further quantified using BLI assay and found to be approximately 3.8 nM and 224 nM, respectively (Fig. 4C) . Notably, the negative stain electron microscopy (EM) analysis (see 1.26 in Example 1) revealed a more uniform protein structure in the presence of DNA, as compared to the condition with protein alone (Fig. 4D) . The 2D classification results revealed an average particle size of approximately and for PqSTAR1-DNA complex and AspSTAR1-DNA complex, respectively (bottom panel of Fig. 4D) .

[0599] Comparative analysis of these sequences revealed some degree of conservation at specific amino acid positions (Fig. 3B) . Intriguingly, we observed high variability at positions 12 and 13, similar to the repeat variable di-residues (RVDs) of TALE-like repeats (Fig. 3B) . Therefore, we conducted a comparative analysis between the two newly identified proteins and canonical TALEs (AvrBs3 was chosen as a reference for comparison) . The overall architectures of PqSTAR1, AspSTAR1, and AvrBs3 shared certain features, such as an N-terminal type III secretion signal (T3SS) and a central DNA-binding repeat domain (Fig. 4E) . However, differences were evident: the lengths of the whole proteins and the non-repetitive regions of PqSTAR1 and AspSTAR1 were shorter (note: the gene encoding the predicted PqSTAR1 lacks a start codon) , there were fewer repeat units, and both lacked a C-terminal transcriptional activation domain. Pairwise comparisons across full-length proteins and repeat regions revealed low sequence identity with canonical TALE (Fig. 4F) , which explained why these two proteins were not initially identified as known TRs. Multiple sequence alignment showed that only nine positions out of 34 were highly conserved across all three proteins, while the RVDs of PqSTAR1 and AspSTAR1 were largely the same as those of canonical TALEs, and the DNA motif recognized by each protein correlated well with the order of RVDs (Fig. 4E) , indicating that they shared the TALE code of nucleotide sequence recognition.

[0600] Next, we subjected these two proteins to protein disorder and structure prediction using DISOPRED3 and AlphaFold2. The repeat regions exhibit well-ordered structures, with an average pLDDT (predicted local distance difference test) score exceeding 80, indicating a confident 3D conformation. The overall structures exhibit a helix-loop-helix architecture, with the RVD positions within the loop region, resembling the canonical TALE structure represented by AvrBs3 (Fig. 4G) . Notably, quantification using BLI indicated that PqSTAR1 exhibited a high DNA-binding affinity with only 9 repeats (Kd = 3.8 nM, Fig. 4C) , surpassing that of the canonical TALE with 10 repeats (Kd = 400 nM) by 100-fold (Rinaldi et al, 2017, The effect of increasing numbers of repeats on TAL effector DNA binding specificity. Nucleic acids research, 45, 6960-6970) . These data collectively suggest that despite some similarity in the overall structure, these two proteins distantly related to TALEs show a higher or comparable DNA-binding affinity, even with a lower unit number. Hence, we designated these two proteins as Short TALE-like Repeat proteins (STAR) , named Pseudomonas quercus STAR homolog 1 (PqSTAR1) and Apophysomyces sp. STAR homolog 1 (AspSTAR1) , respectively, based on the species of origin and the order of characterization.

[0601] Subsequently, we identified all homologous proteins within the originating genus of PqSTAR1 and AspSTAR1. Overall, we identified seven homologs with clear repeat boundaries for PqSTAR1 and two for AspSTAR1. We constructed an evolutionary tree using the full-length sequences of these proteins (Fig. 4H) . The consensus sequence alignments revealed a high degree of amino acid conservation within groups, particularly among PqSTARs (Fig. 4H) . Intriguingly, compared to the canonical TALEs, these two types of STAR proteins overall had smaller number of repeats (Fig. 4I) , suggesting different DNA binding properties.

[0602] The identification of clear binding sequences motivated us to search for potential target genes of PqSTAR1 and AspSTAR1. The T3SS prediction (see 1.15 in Example 1) suggested that PqSTAR1 and AspSTAR1 can be secreted by type III secretion system, implying a possible role in the host genome. The originating species of PqSTAR1, Pseudomonas quercus, was extracted from Quercus mongolica leaf spots, suggesting Quercus mongolica as a potential host. We performed GO enrichment analysis of potential target genes in original genome (Pseudomonas quercus) and potential host genome (Quercus mongolica) . Interestingly, enrichment of GO terms is exclusively observed within the gene set of the host genome, predominantly participating in stress response processes, which suggests PqSTAR1’s role as an effector in regulating host stress-related genes. Following the recognition patterns of RVDs, we conducted a similar analysis of another target gene PqSTAR4 (SEQ ID NO: 27) . The results (see Fig. 8) were consistent with those observed for PqSTAR1, showing that PqSTAR4 specifically recognizes and binds to DNA in a manner similar to PqSTAR1, and strongly suggesting that PqSTARs participate in the infection process of Pseudomonas quercus on Quercus mongolica by regulating host stress-related genes.

[0603] To test the DNA-binding specificity of PqSTAR1 and AspSTAR1, we designed two variants of each by modifying the RVDs of 1-3 repeat units (Fig. 4J) . Both EMSA and GFP activation assays indicated that PqSTAR1 RVD variants showed high specificity to new target sequences with only two RVDs modified (Fig. 4J) . The AspSTAR1 RVD variant showed similar specific binding activity, albeit with lower activity. Due to the high specificity and strong affinity of PqSTAR1 in binding DNA, we aimed to assess its potential as a programmable DNA binding tool. To this end, we first constructed a collection of plasmids encoding diverse repeat modules derived from the PqSTAR1 backbone and utilized the Golden Gate method to assemble customized repeat arrays (Cermak et al., 2011 above) . Subsequently, we assembled three artificial STAR proteins (SEQ ID NOs: 29-31) targeting different sequences, each comprising nine repeat units. In the GFP activation assay, all artificial STARs activated GFP expression, and exhibited binding activity independent of the 5’ thymine (T0) at the binding site (Fig. 4K) , which is preferred by TALEs (Boch et al., 2009, Breaking the code of DNA binding specificity of TAL-type III effectors. Science, 326, 1509-1512) . Collectively, these findings indicate that PqSTAR1 can be readily designed to target specific DNA sequences without T0 dependency.

[0604] Next, we expressed PqSTAR1 and AspSTAR1 with a 3X FLAG tag in 293T cells and conducted a CUT&Tag assay. Western blotting indicated that both PqSTAR1 and AspSTAR1 were expressed efficiently. Furthermore, the motifs enriched in the CUT&Tag assay were similar to both the B1H screen results and the predicted motifs (Fig. 4L) . These data collectively indicate that PqSTAR1 and AspSTAR1 bind to specific DNA sequences in the human genome, supporting their potential application in human cells.

[0605] 5.2. Gene activation by STAR-based transcriptional regulators

[0606] Artificial transcription factors (ATFs) are DNA binding regulators designed to control the expression of a specific gene or a set of genes in a predetermined manner (Miller et al., 2011 above) . While ATFs based on TALE and CRISPR have been developed, they usually regulate one target gene, due to the relatively long target sequence required for binding.

[0607] To compare the capability of STAR and TALE in binding short sequence motifs, we first constructed artificial STARs and TALEs targeting the binding motifs of two well-known TFs, namely NF-κB and SMAD4. Both the gene activation assay and the EMSA assay indicated that the canonical TALE-based ATFs containing only 9 repeats lack binding activity. In contrast, STAR-based ATFs bind to 9 bp target sequences effectively, demonstrating the unique advantages of STARs in binding short DNA motifs. For NF-κB, the target sequence is GGGAATCCC, and the STAR-based ATF has an amino acid sequence of SEQ ID NO: 23; and for SMAD4, the target sequence is GGCCAGACA, and the STAR-based ATF has an amino acid sequence of SEQ ID NO: 24.

[0608] Next, as described in 1.31 of Example 1, we fused the aforementioned STARs with VPR activation domain at the C-terminus, expressed in human 293T cells (Fig. 5A, B) , and performed RNA-seq analysis for four groups of samples, including a wild-type (WT) group without plasmid transduction, the VPR-only group (VPR-only) , STAR targets NF-κB binding sites (STAR_NF-κB) and STAR targets SMAD4 binding sites (STAR_SMAD4) . A robust reproducibility was observed across biological replicates, wherein the VPR-only group closely resembling the WT group (Fig. 5C) . In contrast, the groups treated with STAR-based ATFs exhibit a discernible deviation from the control group, highlighting distinctive transcriptional alterations induced by STAR-based ATFs (Fig. 5C) . Compared to the VPR-only group, transfection with STAR_NF-κB and STAR_SMAD4 resulted in the upregulation of 1, 338 and 2, 489 differentially expressed genes (DEGs, padj < 0.05 and fold change > 2) , accounting for the majority of DEGs (> 70%) (Fig. 5D) .

[0609] Next, we performed motif enrichment and gene set enrichment analysis (GSEA, see 1.32 in Example 1) to determine whether these up-regulated DEGs are truly targets of the designed STAR-based ATFs. Notably, the promoter region of these up-regulated DEGs exhibited enrichment for motifs closely resembling the binding motifs of NF-κB and SMAD4 (Fig. 5E) . Moreover, previously published gene sets of NF-κB / SMAD4-targeted genes were significantly enriched in the up-regulated genes of STAR-based ATF groups (Fig. 5F, H) . The expression levels of the reported target genes were significantly higher than those in the VPR-only groups (Fig. 5G, I) . Collectively, these data indicated that STAR-based ATFs effectively enhanced the expression of a large number of endogenous genes by targeting a specific regulatory motif shared by these genes, establishing a proof of concept for STAR as a platform for constructing ATFs that regulate the transcription network.

[0610] Example 6: Characterization of the MOON family

[0611] This Example was conducted to characterize the MOON family.

[0612] The in vitro screening identified two proteins that bind to specific DNA sequences. One of them (XP_022797784.1, SEQ ID NO: 3) , identified from Stylophora pistillata, showed a binding preference toward AT-rich sequences (Fig. 3C) . Accordingly, we named it Marine Organism-Originated DNA binding protein (MOON) , with the specific name Stylophora pistillata MOON homolog 1 (SpMOON1) . We first confirmed its binding activity of the enriched motif by EMSA (Fig. 6A) , and quantified the binding capacity against GC-rich sequences as a reference using both GFP activation and BLI assays. The results demonstrated that SpMOON1 bound to AT-rich sequences with an approximately 100-fold higher affinity than to GC-rich sequences (Figs. 6B, C) . Negative stain EM analysis showed that AT-rich DNA can stabilize and homogenize protein particles, implying the DNA-binding activity (Fig. 6D) . Further 2D classification results revealed an average particle size of approximately  (bottom panel of Fig. 6D) . These data collectively suggest that SpMOON1 is a DNA-binding protein with a preference for AT-rich sequences.

[0613] The overall architecture of SpMOON1 includes a forkhead-associated (FHA) domain at the N-terminus, a protein phosphatase 1 (PP1) in the middle, and a repeat region at the C-terminus (Fig. 6E) . As we synthesized only the TR region of SpMOON1, the DNA-binding ability was not influenced by FHA and PP1 domains. Notably, a similar architecture has been reported for the Ki67 protein in vertebrates, which is a widely used proliferation marker. Human Ki67 and SpMOON1 exhibit considerable conservation in their N-terminal regions, both containing FHA and PP1 domains. However, notable differences exist in their C-terminal regions. For instance, human Ki67 encodes 16 repeats of approximately 120 amino acids, including a highly conserved 22 amino acid sequence (TPKEKAQALEDLAGFKELFQTP) known as the Ki67 motif. The Ki67 motif and the repeat unit of SpMOON1 are similar in length but share low identity (Fig. 6F) , and no evidence has been presented regarding the DNA binding activity of these Ki67 repeats. Interestingly, there is a leucine / arginine-rich (LR) domain at Ki67 C-terminus that has been experimentally demonstrated to bind AT-rich DNA in vitro. Consequently, we compared the repeat region of SpMOON1 with the LR repeat region of human Ki67, revealing a shared identity of only 22%. Structural comparison was impeded due to the limited confidence in the predicted structure of SpMOON1 (Fig. 6G) . Taken together, these findings propose a shared N-terminal characteristic between SpMOON1 and the Ki67 protein, while significant variations exist within the DNA-binding activity-conferring region. The conservation of the overall protein architecture and binding preferences provides clues that could help elucidate the potential function of SpMOON1 in Stylophora pistillata. Further in-depth research is necessary to elucidate the evolutionary connection between SpMOON1 and human Ki67 repeats.

[0614] Next, we retrieved other proteins from the cluster to which SpMOON1 belongs and expanded our search to include unannotated genome data obtained from NCBI, resulting in the identification of a total of 36 homologs, primarily from the Anthozoa class (31 / 36) . The domain architecture of these homologs was predominantly characterized by an FHA domain at the N-terminus, a PP1 domain in the middle, and varied repeat units at the C-terminus (Fig. 6H) . Multiple sequence alignment of the repeat units of MOON proteins revealed a high degree of conservation, except for a few positions in the middle (Figs. 6H, I) . We selected several homologs with varying evolutionary distances and different VAA sets within the hypervariable positions. Specifically, two proteins came from species distantly related to Stylophora pistillata (Acropora digitifera and Montipora capitate) , and one was from Stylophora pistillat. Accordingly, we named them AdMOON1 (SEQ ID NO: 7) , McMOON1 (SEQ ID NO: 8) , and SpMOON2 (SEQ ID NO: 4) , respectively. After purifying these proteins, we conducted the BLI-based binding assay and all three homologs exhibited DNA-binding activity, while only SpMOON2 showed enrichment of a specific motif in the subsequent SELEX screen (Fig. 6J) . Notably, the enriched motif displayed an AT-rich pattern similar to that observed in SpMOON1. Further BLI experiments revealed a 10-fold increase in the binding affinity of AT-rich sequences compared to GC-rich sequences, indicating a preference for AT-rich binding (Fig. 6K) . The similar binding pattern observed between SpMOON1 and SpMOON2 can be attributed to their shared VAA sets, suggesting that the binding preference could be potentially determined by these VAAs.

[0615] Collectively, these findings demonstrate broad DNA binding activity within the MOON protein family, with certain members exhibiting a preference for AT-rich sequences.

[0616] Example 7: Characterization of the pTERF family

[0617] This Example was conducted to characterize the pTERF family.

[0618] The other DNA-binding TRP identified via the in vitro screen was from the ruminant gut metagenome, which binds the motif ACTNNNAGTC (Fig. 3C) . The resource metagenomic-assembled genome was taxonomically classified as Clostridia bacterium, an unclassified bacterium within the Bacillota phylum. We first confirmed the binding activity of this TRP by EMSA and GFP activation assays (Fig. 7A, B) . BLI assays were further performed to quantify the binding affinity, revealing a Kd value of 1.87 nM (Fig. 7C) . Moreover, we observed a relatively uniform distribution of protein particles in the presence of DNA (Fig. 7D) . Further 2D classification results revealed an average particle size of approximately  (bottom panel of Fig. 7D) .

[0619] Notably, several mTERF (mitochondrial transcription termination factor) motifs were annotated within the repeat region. We have accordingly named it as prokaryotic mTERF-like protein, specifically designated as pTERF homolog 1 (pTERF1) . mTERF motifs are non-tandem repeats and have been well documented as encoding functional nucleic acid-binding proteins in eukaryotes. However, currently, there is no evidence suggesting the existence of a prokaryotic homolog for the mTERF motif. Thus, we conducted a comparative analysis between pTERF1 and human mTERF1 (UniProt entry name: MTERF1_HUMAN) and Drosophila DmTTF (UniProt entry name: MTTF_DROME) . The repeats within the MTERF1_HUMAN and MTTF_DROME proteins are interspersed throughout the protein sequence, while the repeats in pTERF1 are arranged in a tandem manner (Fig. 7E) . Remarkable sequence variability was evident among these three proteins (Fig. 7F) , consistent with a previous study. Although MTERF1_HUMAN and MTTF_DROME shared low identity of their sequences, they retained conserved mTERF motif characteristics, including the preservation of a proline at position 8, along with a leucine or another hydrophobic amino acid at positions 11, 18, and 25, forming three leucine zipper (LZ) -like heptads X3LX3. Notably, these conserved features were not observed in pTERF1 (Fig. 3C) . Given the absence of obvious sequence similarity, we conducted a structural-level comparative analysis between pTERF1 and MTERF1_HUMAN (PDB accession: 3MVA) . The repeat region of pTERF1 showed well-ordered structure with an average pLDDT score exceeding 80, and the general structures of the pTERF1 repeat unit and the single mTERF motif were similar, with each being composed of three α-helices. Nevertheless, differences were apparent, including variations in the number of helical turns and the local conformation of the individual unit (Fig. 7G) .

[0620] To elucidate the evolutionary relationship between pTERF1 and the eukaryotic mTERF family, we conducted a comparative analysis encompassing all experimentally characterized mTERF homologs documented within the UniProtKB / Swiss-Prot database. A total of 30 mTERF protein sequences and their corresponding tertiary structures were retrieved for investigation. Among them, the tertiary structures of two human mTERF proteins, MTERF1_HUMAN and MTERF3_HUMAN, were extracted from the PDB database, while the rest were obtained from the Alphafold2 database. A sequence-level comparison revealed the presence of several mTERF subtypes, including mTERF1, mTERF2, mTERF3 and mTERF4. Notably, plant mTERF4 and mammalian mTERF4 were clustered in different groups, revealing distinct evolutionary paths. Notably, pTERF1 displayed limited sequence identity with all eukaryotic-origin mTERF proteins, ranging from 9%to 18%. This observation potentially implies a unique evolutionary trajectory of pTERF1 compared with eukaryotic mTERFs. In contrast to sequence-level comparisons, structural analysis, as depicted by the root mean square deviation (RMSD) , revealed a relatively conserved relationship among distinct subtypes. These findings support the existence of a certain level of functional conservation among these proteins despite limited sequence similarity. Importantly, both mammalian and plant mTERFs are nucleus-encoded, although they function in mitochondria or chloroplasts. Emerging evidence supports the idea that bacterial genes contribute to the origination and evolution of mitochondria. Hence, the identification of pTERF1 raises the exciting possibility of the presence of the first prokaryotic-origin mTERF-like protein.

[0621] To further characterize the pTERF family, we investigated the homologs of pTERF1 within the same cluster and recovered five proteins. Deep homolog searching in the NCBI database identified 14 additional homologs. Most of them had a unit number below 6 and therefore were not included in the original list of TRs of interest (Fig. 7H) . All of these proteins were obtained from the metagenomic-assembled genomes (MAGs) . The MAG assembly completeness levels ranged from approximately 70%to 100%, and contamination levels from 0%to 10%, indicating a medium to high quality of the assembly. Additionally, all the MAGs were taxonomically classified within the Bacteria kingdom, with the majority falling within the Bacillota phylum. In order to confirm the widespread presence of pTERFs in metagenomic ecosystems, we further conducted a search for homologous sequences in the MGnify database, and successfully identifying an additional 27 homologs. To investigate whether other proteins within the pTERF family also possessed DNA binding ability, we selected several genes for analysis from various positions on the phylogenetic tree, one of which (pTERF2) was successfully synthesized (Fig. 7I) . The protein was purified, and a subsequent SELEX screen was conducted. Interestingly, a motif different from that of pTERF1 was enriched (Fig. 7J) . We further validated DNA binding activity by EMSAs and BLI assays (Fig. 7K, L) . The obtained data indicated that distinct repeat arrangements of pTERF1 and pTERF2 resulted in different binding specificities, suggesting potential reprogrammability.

[0622] To gain initial insights into the functions of pTERF proteins, we conducted an analysis to predict the potential gene functions targeted by pTERF1 and pTERF2 (see 1.17 in Example 1) . It was shown that pTERF1 and pTERF2 were associated with similar GO terms related to RNA metabolic processes. Collectively, these data revealed the binding characteristics of pTERF family, and highlight the discovery of the first class of distant mTERF homologs in prokaryotes, shedding light on their evolutionary significance and functional diversity. Furthermore, the recognition of diverse DNA sequences is facilitated by the different arrangements of repeats within the protein family, suggesting the intriguing possibility of reprogrammability.

[0623] Example 8: Characterization of additional TRP

[0624] This Example is conducted to characterize other TRPs.

[0625] The TRP of SEQ ID NO: 10 was tested by BLI assay, GFP activation and EMSA assay as described above.

[0626] As shown in Fig. 9, the TRP can specifically recognize the motif of AATAGCTTTT.

[0627] Example 9: Randomized libraries of artificial transcription factors (TFs)

[0628] The one-step construction of randomized libraries followed the previously reported Golden Gate method, with several refinements (Cermak et al., 2011, Efficient design and assembly of custom TALEN and other TAL effector-based constructs for DNA targeting. Nucleic acids research 39, e82-e82) . Specifically, the repeat module on pHD-1 was replaced with modules derived from PqSTAR1, each possessing a unique overhang. Additionally, we implemented modifications to the pFUS_A vector by inserting the N-and C-terminal regions of PqSTAR1 at both ends of LacZ, respectively. Two internal BsaI sites were strategically placed at the 5’ and 3’ termini of LacZ to facilitate vector linearization with the enzyme, resulting in the generation of suitable overhangs for the incorporation of the repeat modules. The products of these ligation reactions were used to transform E. coli to yield independent transformants. Plasmids were purified from these transformants, which constituted a randomized library.

[0629] Results:

[0630] a) Tool development: We generated a pool of distinct STAR repeat modules, each possessing unique DNA-binding specificities. These modules were subsequently randomized and assembled into a library of artificial STARs, designed to target arbitrary N-mer DNA sequences. These proteins were then fused to different regulation elements, such as transcription activator, transcription repressor, epigenetic regulatory elements.

[0631] b) Phenotype screening: The aforementioned tools could be utilized to screen functional phenotypes in bacteria, mammalian cells and plant cells.

[0632] Sequences of the Identified TR-containing proteins

[0633] The tandem repeats are underlined and the RVDs are bold.

[0634] SEQ ID NO: 1 PqSTAR1 (WP_178089108.1)

[0635] SEQ ID NO: 2 AspSTAR1 (KAG0189736.1)

[0636] SEQ ID NO: 3 SpMOON1 (XP_022797784.1)

[0637] SEQ ID NO: 4 SpMOON2 (XP_022797764.1)

[0638] SEQ ID NO: 5 pTERF1 (MBQ3414883.1)

[0639] SEQ ID NO: 6 pTERF2 (MBE6149963.1)

[0640] SEQ ID NO: 7 AdMOON1:

[0641] SEQ ID NO: 8 McMOON1:

[0642] SEQ ID NO: 9 A0A7G6T5J7 (cluster_2)

[0643] SEQ ID NO: 10 A0A662FLR2 (cluster_4)

[0644] SEQ ID NO: 11 EKD25150.1 (cluster_4)

[0645] SEQ ID NO: 12 WP_274699397.1 (cluster_5)

[0646] SEQ ID NO: 13 CAH8242182.1 (cluster_5)

[0647] SEQ ID NO: 14 WP_277432871.1 (cluster_5)

[0648] SEQ ID NO: 15 K1X4V8 (cluster_8)

[0649] SEQ ID NO: 16 A0A835Z2I7 (cluster_10)

[0650] SEQ ID NO: 17 KAG5183916.1 (cluster_10)

[0651] SEQ ID NO: 18 A0A1G0X562 (cluster_12)

[0652] SEQ ID NO: 19 MCH8959980.1 (cluster_15)

[0653] SEQ ID NO: 20 MGYP003304872820 (cluster_16)

[0654] SEQ ID NO: 21 V4B0R4 (cluster_18)

[0655] SEQ ID NO: 22 A0A3S0VJV9 (cluster_19)

[0656] SEQ ID NO: 23 STAR_NF-κB

[0657] SEQ ID NO: 24 STAR_SMAD4

[0658] SEQ ID NO: 25 PqSTAR_based TALE-like polypeptide

[0659] SEQ ID NO: 26 AspSTAR1_based TALE-like polypeptide

[0660] SEQ ID NO: 27 PqSTAR4 WP_178089118.1

[0661] SEQ ID NO: 28 PqSTAR4 based TALE-like polypeptide

[0662] SEQ ID NO: 29 Artificial STAR1

[0663] SEQ ID NO: 30 Artificial STAR2

[0664] SEQ ID NO: 31 Artificial STAR3

Claims

1.A programmed / programmable TALE-like polypeptide comprising a N-terminal region, two or more tandem repeats and a C-terminal region, wherein each of the tandem repeats comprises an amino acid sequence selected from the group consisting of Formulae I to VIIFX1NDNLVKVAAX2X3GX4X5X6ALQX7LLDX8GPALRQAG     (I)whereX1 is G or S,X2X3 are repeat variable di-residue (RVD) ,X4 is G or S,X5 is A or Q,X6 is H or Q,X7 is A or T, andX8 is K or R;F X1H X2QIV X3IAS X4X5GGSQAL X6X7VL X8X9X10A X11L X12X13X14G     (II)whereX1 is T or K,X2 is Q, E or R,X3 is A or G,X4X5 are RVD,X6 is N or D,X7 is T or K,X8 is A or V,X9 is T, R or K,X10 is H or Y;X11 is A, P or Q,X12 is T or R,X13 is A, D or T, andX14 is A or V.GGREQVIKIAA X1X2GGKQALQALLDKSPALRQAG     (III)where X1X2 are RVD,FSNDNLVRIGG X1X2GAKKTLDTLLQVYPKLTQGG     (IV)where X1X2 are RVD,ILSGQETNRIK X1X2GGAKALETLSEKAEALHRAG     (V)where X1X2 are RVD,FSKQEAVAIAS X1X2GGSQALNTVLATHATLTAAG     (VI)where X1X2 are RVD,andFAVEDVSAIAA X1X2GGAPALQAVVDHLELLMTRH     (VII)where X1X2 are RVD.2.The programmed / programmable TALE-like polypeptide of claim 1, wherein each of the tandem repeats comprises an amino acid sequence selected from the group consisting of the Formulae III to XXIGGREQVIKIAA X1X2GGKQALQALLDKSPALRQAG     (III)FSNDNLVRIGG X1X2GAKKTLDTLLQVYPKLTQGG     (IV)ILSGQETNRIK X1X2GGAKALETLSEKAEALHRAG     (V)FSKQEAVAIAS X1X2GGSQALNTVLATHATLTAAG     (VI)FAVEDVSAIAA X1X2GGAPALQAVVDHLELLMTRH     (VII)FGNDNLVKVAA X1X2GGAQALQALLDKGPALRQAG     (VIII)FGNDNLVKVAA X1X2GSQHALQALLDKGPALRQAG     (IX)FGNDNLVKVAA X1X2GGAQALQALLDKGPALRQAG     (X)FGNDNLVKVAA X1X2GSQQALQALLDKGPALRQAG     (XI)FGNDNLVKVAA X1X2GGAQALQALLDRGPALRQAG     (XII)FSNDNLVKVAA X1X2GGAHALQALLDKGPALRQAG     (XIII)FSNDNLVKVAA X1X2GGQQALQTLLDKGPALRQAG     (XIV)FTHQQIVAIAS X1X2GGSQALNTVLATHAALTAAG     (XV)FTHQQIVAIAS X1X2GGSQALDKVLATHAPLTAAG     (XVI)FTHRQIVGIAS X1X2GGSQALDTVLVRYAPLRDAG     (XVII)FKHEQIVGIAS X1X2GGSQALDKVLATHAQLTAVG     (XVIII)FKHEQIVAIAS X1X2GGSQALDKVLVKYAPLTAAG     (XIX)FTHQQIVAIAS X1X2GGSQALDTVLATHAQLTTAG     (XX)FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG     (XXI)where X1X2 are RVD.3.The programmed / programmable TALE-like polypeptide of claim 1, wherein each of the tandem repeats comprises an amino acid sequence selected from the group consisting of the Formulae I, III and IV.4.The programmed / programmable TALE-like polypeptide of claim 1, wherein each of the tandem repeats comprises an amino acid sequence selected from the group consisting of the Formulae II and V to VII.5.The programmed / programmable TALE-like polypeptide of claim 2, wherein each of the tandem repeats comprises an amino acid sequence selected from the group consisting of the Formulae III, IV and VIII to XIV.6.The programmed / programmable TALE-like polypeptide of claim 2, wherein each of the tandem repeats comprises an amino acid sequence selected from the group consisting of the Formulae V to VII and XV to XX.7.A fusion polypeptide comprising the programmed / programmable TALE-like polypeptide of any one of claims 1-6 fused to a fusion partner.8.The fusion polypeptide of claim 7, wherein the fusion partner is selected from the group consisting ofi) a polypeptide that provides an activity of nuclease,ii) a polypeptide that provides an activity that indirectly increases transcription by acting directly on the target DNA or on a polypeptide (e.g., a histone or other DNA-binding protein) associated with the target DNA,iii) a polypeptide that provides for methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, ribosylation activity, deribosylation activity, myristoylation activity, or demyristoylation activity.9.A fusion polypeptide comprising a DNA binding polypeptide and a fusion partner, wherein the DNA binding polypeptide comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-6 and 9-28.10.The fusion polypeptide of claim 9, wherein the fusion partner is selected from the group consisting ofi) a polypeptide that provides an activity of nuclease,ii) a polypeptide that provides an activity that indirectly increases transcription by acting directly on the target DNA or on a polypeptide (e.g., a histone or other DNA-binding protein) associated with the target DNA,iii) a polypeptide that provides for methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, ribosylation activity, deribosylation activity, myristoylation activity, or demyristoylation activity.11.The fusion polypeptide of any one of claims 7-10, wherein the fusion partner comprises a nuclear localization sequence (NLS) .12.A recombinant gene editing system comprising a fusion polypeptide comprising the programmed / programmable TALE-like polypeptide of any one of claims 1-6 fused to a fusion partner, or a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide, wherein the fusion partner is a polypeptide that provides an activity of nuclease.13.The recombinant gene editing system of claim 12, comprising a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide.14.A composition comprising a fusion polypeptide comprising the programmed / programmable TALE-like polypeptide of any one of claims 1-6 fused to a fusion partner, or a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide, wherein the fusion partner is a polypeptide that provides an activity of nuclease.15.A method of introducing a double-strand break into a polynucleotide of interest comprising a step of contacting the polynucleotide with a recombinant gene editing system comprising a fusion polypeptide comprising the programmed / programmable TALE-like polypeptide of any one of claims 1-6 fused to a fusion partner, or a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide, wherein the fusion partner is a polypeptide that provides an activity of nuclease.16.A method of modifying a genomic sequence in a cell comprising a step of introducing into the cell a recombinant gene editing system comprising a fusion polypeptide comprising the programmed / programmable TALE-like polypeptide of any one of claims 1-6 fused to a fusion partner, or a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide, wherein the fusion partner is a polypeptide that provides an activity of nuclease.17.The method of claim 16, wherein the recombinant gene editing system comprising a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide.18.A recombinant gene editing system comprising a fusion polypeptide comprising a DNA binding polypeptide fused to a fusion partner, or a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide, wherein the DNA binding polypeptide comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-6 and 9-28, wherein the fusion partner is a polypeptide that provides an activity of nuclease.19.The recombinant gene editing system of claim 18, comprising a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide.20.A composition comprising a fusion polypeptide comprising a DNA binding polypeptide fused to a fusion partner, or a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide, wherein the DNAbinding polypeptide comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-6 and 9-28, wherein the fusion partner is a polypeptide that provides an activity of nuclease.21.A method of introducing a double-strand break into a polynucleotide of interest comprising a step of contacting the polynucleotide with a recombinant gene editing system comprising a fusion polypeptide comprising a DNA binding polypeptide fused to a fusion partner, or a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide, wherein the DNA binding polypeptide comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-6 and 9-28, wherein the fusion partner is a polypeptide that provides an activity of nuclease.22.A method of modifying a genomic sequence in a cell comprising a step of introducing into the cell a recombinant gene editing system comprising a fusion polypeptide comprising a DNA binding polypeptide fused to a fusion partner, or a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide, wherein the DNA binding polypeptide comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-6 and 9-28, wherein the fusion partner is a polypeptide that provides an activity of nuclease.23.The method of claim 22, wherein the recombinant gene editing system comprising a polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide.24.The method of claim 22 or 23, wherein the cell is a eukaryotic cell.25.A randomized library of artificial transcription factors (TFs) , comprising a plurality of cells, each of which harboring a vector comprising a nucleotide sequence encoding a polypeptide comprising a N-terminal region, two or more tandem repeats and a C-terminal region, wherein each of the tandem repeats comprises an amino acid sequence selected from the group consisting of Formulae I to VIIFX1NDNLVKVAAX2X3GX4X5X6ALQX7LLDX8GPALRQAG      (I)whereX1 is G or S,X2X3 are repeat variable di-residue (RVD) ,X4 is G or S,X5 is A or Q,X6 is H or Q,X7 is A or T, andX8 is K or R;F X1H X2QIV X3IAS X4X5GGSQAL X6X7VL X8X9X10A X11L X12X13X14G     (II)whereX1 is T or K,X2 is Q, E or R,X3 is A or G,X4X5 are RVD,X6 is N or D,X7 is T or K,X8 is A or V,X9 is T, R or K,X10 is H or Y;X11 is A, P or Q,X12 is T or R,X13 is A, D or T, andX14 is A or V.GGREQVIKIAA X1X2GGKQALQALLDKSPALRQAG     (III)where X1X2 are RVD,FSNDNLVRIGG X1X2GAKKTLDTLLQVYPKLTQGG     (IV)where X1X2 are RVD,ILSGQETNRIK X1X2GGAKALETLSEKAEALHRAG     (V)where X1X2 are RVD,FSKQEAVAIAS X1X2GGSQALNTVLATHATLTAAG     (VI)where X1X2 are RVD,andFAVEDVSAIAA X1X2GGAPALQAVVDHLELLMTRH     (VII)where X1X2 are RVD,wherein the tandem repeats are randomized between cells.26.The randomized library of claim 25, wherein each of the tandem repeats comprises an amino acid sequence selected from the group consisting of the Formulae III to XXIGGREQVIKIAA X1X2GGKQALQALLDKSPALRQAG     (III)FSNDNLVRIGG X1X2GAKKTLDTLLQVYPKLTQGG     (IV)ILSGQETNRIK X1X2GGAKALETLSEKAEALHRAG     (V)FSKQEAVAIAS X1X2GGSQALNTVLATHATLTAAG     (VI)FAVEDVSAIAA X1X2GGAPALQAVVDHLELLMTRH     (VII)FGNDNLVKVAA X1X2GGAQALQALLDKGPALRQAG     (VIII)FGNDNLVKVAA X1X2GSQHALQALLDKGPALRQAG     (IX)FGNDNLVKVAA X1X2GGAQALQALLDKGPALRQAG     (X)FGNDNLVKVAA X1X2GSQQALQALLDKGPALRQAG     (XI)FGNDNLVKVAA X1X2GGAQALQALLDRGPALRQAG     (XII)FSNDNLVKVAA X1X2GGAHALQALLDKGPALRQAG     (XIII)FSNDNLVKVAA X1X2GGQQALQTLLDKGPALRQAG     (XIV)FTHQQIVAIAS X1X2GGSQALNTVLATHAALTAAG     (XV)FTHQQIVAIAS X1X2GGSQALDKVLATHAPLTAAG     (XVI)FTHRQIVGIAS X1X2GGSQALDTVLVRYAPLRDAG     (XVII)FKHEQIVGIAS X1X2GGSQALDKVLATHAQLTAVG     (XVIII)FKHEQIVAIAS X1X2GGSQALDKVLVKYAPLTAAG     (XIX)FTHQQIVAIAS X1X2GGSQALDTVLATHAQLTTAG     (XX)FSNDNLVKVAAX1X2GGAQALQALLDKGPALRQAG     (XXI)where X1X2 are RVD.27.The randomized library of claim 25, wherein each of the tandem repeats comprises an amino acid sequence selected from the group consisting of the Formulae I, III and IV.28.The randomized library of claim 25, wherein each of the tandem repeats comprises an amino acid sequence selected from the group consisting of the Formulae II and V to VII.29.The randomized library of claim 26, wherein each of the tandem repeats comprises an amino acid sequence selected from the group consisting of the Formulae III, IV and VIII to XIV.30.The randomized library of claim 26, wherein each of the tandem repeats comprises an amino acid sequence selected from the group consisting of the Formulae V to VII and XV to XX.31.The randomized library of any one of claims 25-30, wherein the polypeptide comprises 6-8 TRs.

Citation Information

Patent Citations

  • Method for building TALE (transcription activator-like effector) repeated sequences

    CN102787125A

  • Novel DNA-binding proteins and uses thereof

    CN103025344A

  • Drug-induced fusion protein for genome editing, coding gene and application thereof

    CN109206520A

  • Nucleic Acid Binding Domains and Methods of Use Thereof

    US20220306699A1