High throughput screening and sequencing methods

By employing high-throughput screening and sequencing of DNA-modifying enzymes, and utilizing expression vector libraries and nanopore sequencing technology, the time-consuming and labor-intensive generation of site-specific recombinases in existing technologies has been resolved. This approach enables highly efficient screening and sequencing of DNA-modifying enzymes and target sites, thereby improving the efficiency and accuracy of genomic DNA manipulation.

CN121152876APending Publication Date: 2025-12-16TECHNISCHE UNIVERSITAT DRESDEN
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202480029004.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-04
Filing Date
2024-04-17
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing technologies are time-consuming and labor-intensive in generating site-specific recombinases (Y-SSRs) with custom specificity, and it is difficult to efficiently screen and sequence target sites of various DNA modification enzymes, resulting in low efficiency of genomic DNA manipulation.

Method used

High-throughput methods were used to screen and sequence DNA-modifying enzymes. By providing an expression vector library containing multiple vectors, host cells were introduced, and DNA-modifying enzymes were cultured and expressed. Plasmid DNA was isolated, and nanopore sequencing was used to determine whether the DNA sequence was modified. Clustering and common sequence polishing were performed using unique molecular identifiers (UMIs) to quickly identify valid DNA modification events.

Benefits of technology

This technology enables efficient screening and sequencing of various DNA modifying enzymes and their target sites, significantly improving the efficiency and accuracy of DNA manipulation while reducing screening time and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present invention is in the field of DNA modifying enzymes and provides high throughput methods for characterizing and sequencing a variety of DNA modifying enzymes or target sites for DNA modifying enzymes. The present invention provides a method for screening and sequencing a plurality of DNA modifying enzymes, the method comprising the steps of: providing an expression vector library comprising a plurality of different vectors wherein each vector comprises a first region encoding a DNA modifying enzyme variant and a second region comprising one or more target sites of a DNA modifying enzyme; introducing the expression vector library into a host cell; culturing the host cell and expressing the DNA modification enzyme; separating plasmid DNA (Deoxyribose Nucleic Acid) from the host cell culture; sequencing the first region and the second region of the expression vector; and determining whether the DNA sequence of the second region on the expression vector is changed by the DNA modifying enzyme or not based on the sequencing result. The invention further provides methods for screening and sequencing a plurality of target sites of a DNA modifying enzyme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a high-throughput method for characterizing and sequencing a variety of DNA-modifying enzymes or their target sites. Background Technology

[0002] DNA-modifying enzymes, particularly site-specific recombinase (SSR) systems, allow for precise manipulation of DNA without triggering endogenous DNA repair pathways. They possess the unique ability to perform cleavage and immediate rejoining of processed DNA within the body.

[0003] Tyrosine-type site-specific recombinases (Y-SSRs) are widely used tools in genome engineering. Besides the Cre / loxP and Flp / FRT systems, other recombinase systems are known in the art. For example, US 7,422,889 and US 7,915,037 disclose the so-called Dre / rox system, which contains the Dre recombinase isolated from enterobacterial bacteriophage D6, with a recognition site called the rox site. Other known recombinase systems are the VCre / VloxP system isolated from Vibrio plasmid p0908, and the sCre / SloxP system (WO 2010 / 143606 A1; Suzuki and Nakayama, 2011). Other site-specific DNA recombinase systems are the Nigri / nox system disclosed in EP 2877585 B1, the Vika / vox system disclosed in EP 2690177 B1, and the Panto / pox system disclosed in EP 3263708 B1.

[0004] Y-SSRs can exchange DNA strands between their target DNA sequences, facilitating controlled excision, inversion, insertion, or exchange of DNA. Because this process is highly precise and independent of DNA repair mechanisms, it provides seamless manipulation of genomic DNA without side effects (Meinke et al. 2016). These characteristics distinguish Y-SSRs from nuclease-based genome engineering methods, such as the CRISPR-Cas system. All nuclease-based genome editing tools rely on the host cell's DNA repair pathways, which can ultimately lead to unintended editing results (reviewed in Anzalone et al. 2020). However, the widely used CRISPR-Cas9 system has the advantage of being rapidly and efficiently programmed to specifically edit new target sites. In this respect, Y-SSRs lag behind, and currently require significant time and effort to generate Y-SSRs with customized specificity. This bottleneck represents a considerable obstacle to leveraging the full potential of designer-recombinases as versatile genome editing tools.

[0005] Current methods for adapting site-specific recombinases to novel target sequences use directed molecular evolution. A powerful method for engineered Y-SSR evolution utilizes plasmid-based bacterial applications and is called substrate-associated directed evolution (SLiDE) (Buchholz and Stewart 2001; Buchholz and Hauber 2011; Lansing et al. 2019; Hoersten et al. 2021; Lansing et al. 2022).

[0006] Substrate-associated directed evolution allows recombinase libraries to evolve gradually by screening for random mutations combined with a selection scheme targeting novel, progressively altered target sequence activities. Several recombinases have been developed in this manner (Buchholz and Stewart 2001; Sarkar et al. 2007; Karpinski et al. 2016; Lansing et al. 2019; Lansing et al. 2022), demonstrating the applicability of substrate-associated directed evolution. Improvements to this approach have also been reported (Lansing et al. 2019), but the generation of new designer Y-SSRs still requires significant resources. Large libraries of gene variants (around 10) are generated during multiple rounds of mutagenesis and selection. 5 and 10 8 (Lawrence et al., 2013; Rognes et al., 2016; Vaser et al., 2017). Since screening libraries to identify valid variants is usually a manual process, it is laborious and time-consuming, and typically requires six to twelve months of work.

[0007] These drawbacks are a direct consequence of the stochastic approach employed when creating next-generation recombinases using substrate-associated directed evolution. At the start of a new project, it is unknown which sequence modifications will result in protein variants exhibiting recombinant activity at predefined target sequences. Improvements based on protein modeling of designer recombinases active at defined target sites have yielded some success (Abi-Ghanem et al. 2012). However, the evolution and characterization of DNA-modifying enzymes remains a cumbersome and time-consuming process due to the complex nature of the entire enzymatic reaction, including DNA binding, DNA bending, and catalysis (reviewed in Meinke et al. 2016). Furthermore, the limited number of testable variants reduces the likelihood of identifying the optimal variant.

[0008] Therefore, one object of the present invention is to provide a high-throughput method for characterizing and sequencing a variety of DNA-modifying enzymes. Another object of the present invention is to provide a high-throughput method for characterizing and sequencing target sites of DNA-modifying enzymes. Summary of the Invention

[0009] The objective of this invention is achieved by providing the method of this invention.

[0010] According to a first aspect, the present invention provides a method for screening and sequencing multiple DNA modifying enzymes, the method comprising the following steps: Provides an expression vector library containing a variety of different vectors, wherein each vector contains a first region encoding a variant of a DNA-modifying enzyme and a second region containing one or more target sites of the DNA-modifying enzyme; The expression vector library was introduced into the host cell; The host cells are cultured and the DNA-modifying enzyme is expressed; Plasmid DNA was isolated from host cell cultures; Sequencing of the first and second regions of the expression vector; and Based on the sequencing results, it is determined whether the DNA sequence of the second region on the expression vector has been altered by the DNA-modifying enzyme.

[0011] According to a second aspect, the present invention provides a method for screening and sequencing multiple target sites of DNA modifying enzymes, the method comprising the following steps: Provides an expression vector library containing a variety of different vectors, wherein each vector contains a first region encoding a DNA-modifying enzyme and a second region containing a variant of one or more target sites of the DNA-modifying enzyme; The expression vector library was introduced into the host cell; The host cells are cultured and the DNA-modifying enzyme is expressed; Plasmid DNA was isolated from host cell cultures; Sequencing of the first and second regions of the expression vector; and Based on the sequencing results, it is determined whether the DNA sequence of the second region on the expression vector has been altered by the DNA-modifying enzyme.

[0012] According to one embodiment, in the method of the present invention, the first region further comprises a unique molecular identifier (UMI). According to a preferred embodiment, the unique molecular identifier is an oligonucleotide comprising at least 50 random nucleotides.

[0013] According to a preferred embodiment, the unique molecular identifier is located in a first region of the expression vector adjacent to the sequence encoding the DNA-modifying enzyme.

[0014] According to another embodiment, the method further includes the following steps: Cluster the unique molecular identifiers. Generate and polish the common sequence. Determine the number of DNA modification events for each DNA modifying enzyme; and Determine the activity rate of each DNA-modifying enzyme.

[0015] According to a further implementation, the first and second regions of the vector are sequenced in a single step.

[0016] According to one embodiment, sequencing of the first and second regions of the expression vector includes nanopore sequencing.

[0017] According to a preferred embodiment, the DNA modifying enzyme is selected from recombinases, integrases, adenosine base editors (ABEs), zinc finger nucleases, transcription activator-like effector nucleases, and Cas nucleases.

[0018] According to yet another embodiment, the DNA-modifying enzyme comprises more than one subunit.

[0019] According to a further embodiment, the DNA-modifying enzyme comprises at least two distinct subunits.

[0020] According to yet another embodiment, prior to the sequencing, the first and second regions of the expression vector are excised from the expression vector.

[0021] Other aspects and embodiments of the invention will become apparent from the appended claims and the following detailed description. Attached Figure Description

[0022] The present invention is further illustrated by the following figures and examples, but is not limited thereto.

[0023] Figure 1 An overview of the workflow of a particularly preferred method according to the invention is shown. An evolved recombinase, along with a unique molecular identifier (UMI), is cloned into a vector containing a loxF8 target site. *E. coli* is transformed using the designed vector. E. coli Cells. Transformed bacteria were cultured to express the enzyme. To screen for the same enzyme variant at different target sites, plasmid DNA was isolated, and the DNA-modifying enzyme, along with its UMI, was subcloned into additional target sites (HG1, HG2, and HG2L) and cultured as previously described. From all plasmids isolated from cell cultures, regions of interest (including regions 1 and 2) were excised and sequenced via nanopore sequencing. For further evaluation, UMIs were clustered, allowing for concordant sequence polishing and used to identify recombination events.

[0024] Figure 2 A shows the results of DNA editing quantitative sequencing screening of the loxF8 recombinase at four target sites: loxF8 (SEQ ID NO: 11), HG1 (SEQ ID NO: 12), HG2 (SEQ ID NO: 13), and HG2L (SEQ ID NO: 14). UMI clusters containing evolved recombinases are marked with gray dots, and D7 control clusters (SEQ ID NO: 19 and 20) are marked with black dots. Three selected clusters are highlighted: 138 (SEQ ID NO: 21 and 22), 181 (SEQ ID NO: 23 and 24), and 1244 (SEQ ID NO: 25 and 26). Figure 2 B shows the median recombination rate of recombinase D7. Figure 2 C shows that 52 non-D7 clusters identified by the method of the present invention have less than 10% off-target activity on three off-targets and more than 25% activity on the target.

[0025] Figure 3 A is a schematic diagram of the recombination assay (Examples 5 and 10). Recombinant and non-recombinant plasmids (top circles) are digested with a restriction enzyme that excises the recombinase gene. Because the plasmid sizes differ between the two versions, the resulting DNA fragments containing the plasmid backbone also differ in size. This is visualized by agarose gel electrophoresis (bottom schematic). The mixture will contain both fragments, with the intensity of each band corresponding to the amount of the corresponding fragment, making it possible to quantify the recombination rate. A pictogram is labeled on the right side of the gel scheme to indicate the fragment corresponding to each band. Two triangles represent the non-recombinant fragment, and one triangle represents the recombinant fragment. Figure 3 B shows the corresponding agarose gel analysis of recombinant assays performed on four different DNA modifying enzymes (cluster 138, 181, 1244 and D7 control) at four different target sites: loxF8 (SEQ ID NO: 11), HG1 (SEQ ID NO: 12), HG2 (SEQ ID NO: 13) and HG2L (SEQ ID NO: 14). Figure 3 C shows the recombination assay performed in three replicates ( Figure 3 B) and quantitative results. The band intensities of recombinant and non-recombinant products for selected recombinase clusters (138, 181, 1244) and the control recombinase D7 were determined using the image analysis software Fiji (Schindelin et al., 2012). The band intensity value of the recombinant product was divided by the combined value of the recombinant and non-recombinant bands to determine the fraction of recombinant DNA, which was then converted to a percentage value by multiplying by 100. Each point represents one replicate of the assay for the corresponding variant.

[0026] Figure 4 The adapted substrate-associated directed evolution (SLiDE) method used in Examples 4 and 11 is illustrated schematically. The gene for a DNA-editing enzyme used for base editing is amplified using error-prone PCR or DNA rearrangement. This results in multiple copies of the mutated gene. This gene library is then cloned into a bacterial expression vector containing the target site for the base editor, where the restriction enzyme (RE) site is located where the base change should occur. The expression vector is transformed into bacteria expressing the DNA-modifying enzyme. The sgRNA translated from the same vector guides the base-editing enzyme to the target site and induces editing, depending on the enzyme's activity. An example of editing can be seen in the bubble at the top right of the figure. Here, A is modified to G, resulting in the loss of the RE site. Digestion with the corresponding RE results in two possible outcomes: the plasmid is digested twice or once. Only the plasmid obtained from a single digestion is a valid template for error-prone PCR (performed with primers as indicated by the arrows) to begin a new evolutionary cycle.

[0027] Figure 5 A is a schematic diagram of the base editing determination explained in detail in Example 8. Figure 5 B shows the agarose gel analysis of base editing assays performed on Cas12f-ABE (WT, SEQ ID NO: 27) and evolutionary libraries derived from WT at target sites 1 to 3 (SEQ ID NO: 15, 16 and 6). Figure 5 C shows the base editing assay results of the evolved Cas12f-ABE library without mutations at different protein expression levels (induced by L-arabinose) after four evolutionary cycles.

[0028] Figure 6 An overview of the workflow of another particularly preferred method according to the invention is shown. An evolutionary variant of Cas12f-ABE, along with a unique molecular identifier (UMI), is cloned into a vector containing three target sites. After transformation into bacteria, the transformed bacteria are cultured to express the enzyme. Plasmid DNA is isolated, the regions of interest (including regions 1 and 2) are excised, and sequencing is performed. Clustering of the UMIs allows for concordant sequence polishing and is used to identify base editing events.

[0029] Figure 7A shows the results of screening for Cas12f-ABE variants at three target sites: target site 1 (SEQ ID NO: 15), target site 2 (SEQ ID NO: 6), and target site 3 (SEQ ID NO: 16). UMI clusters containing Cas12f-ABE variants are marked with gray dots, and the WT (SEQ ID NO: 27) control cluster is marked with black dots. The three selected clusters (2 (SEQ ID NO: 28), 3030 (SEQ ID NO: 29), and 3301 (SEQ ID NO: 30)) are highlighted. Figure 7 B indicates the median edit rate of the WT cluster. Figure 7 C shows that the 58 non-WT clusters identified using the method of the present invention have an editing rate of over 90% at all three target sites. Figure 7 D shows the three target sites used in the Cas12f-ABE screening of the WT cluster and three variants from clusters 2, 3030, and 3301. Figure 6 and 7 The percentage of read segments with correct editing, no editing, or other editing results on A (SEQ ID NO: 15, 16, and 6).

[0030] Figure 8 A shows the corresponding agarose gel analysis of base editing plasmid assays performed simultaneously in triplicate (rep. 1-3) on three different target sites (target site 1 (SEQ ID NO: 15), target site 2 (SEQ ID NO: 6), and target site 3 (SEQ ID NO: 16)) for four different DNA modifying enzymes (cluster 2, 3030, 3301, and WT control). Figure 8 B shows the determination of base editing plasmids ( Figure 8 A) Quantitative results. The band intensities of edited and unedited products of selected Cas12f-ABE clusters (2, 3030, 3301) and the WT control were determined using the image analysis software Fiji. The band intensity value of the edited product was divided by the combined value of the edited and unedited bands to calculate the fraction, which was then converted to a percentage value. Each point represents one replicate of the determination for the corresponding variant.

[0031] Figure 9Figure A shows the editing results of DNA-modifying enzymes from clusters 2, 3030, and 3301, and WT Cas12f-ABE at three E. coli genomic target sites (SEQ ID NO: 8-10), analyzed by Sanger sequencing. A negative control is also included, in which the DNA-modifying enzymes are not expressed in the bacteria. Arrows indicate the locations where base editing should have occurred. The sequence of unedited genomic DNA is shown above the curves. Bases are also placed directly above each line plot at the location where editing should have occurred. In cases where there are two bases at this location, the upper base corresponds to a higher curve. Figure 9 B shows the use of the EditR program (Kluesner et al., 2018) for... Figure 9 The Sanger sequencing results of A quantify the base editing rate.

[0032] Figure 10 An overview of the workflow of another particularly preferred method according to the invention is shown. Sequences encoding different variants of DNA-modifying enzymes are cloned into plasmids containing the corresponding enzyme target sites. *E. coli* cells are transformed with the vector. The transformed bacteria are cultured to express the encoded DNA-modifying enzyme. Plasmid DNA is isolated. From all plasmids isolated from the cell culture, regions of interest (including regions 1 and 2) are excised and sequenced using nanopore sequencing.

[0033] Figure 11 This section shows the screening results of naturally occurring tyrosine site-specific recombinases at multiple target sites. Recombinases and target sites had previously been identified and screened without the use of UMI. Recombination events are displayed as a heatmap of the percentage of recombination for each possible combination. Target sites are shown at the level of similarity and ranked according to their similarity. The recombinases tested are aligned on the vertical axis based on their homology. sequence

[0034] The sequences mentioned in this specification are disclosed in the accompanying sequence listing. Detailed Implementation

[0035] Before describing the invention in detail below, it should be understood that the invention is not limited to the specific methods, schemes, and reagents described herein, as these can vary. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the invention, which is limited only by the appended claims. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.

[0036] Preferably, the terms used herein are defined as described in "A multilingual glossary of biotechnological terms: (IUPAC Recommendations)", Leuenberger, HGW, Nagel, B. and Klbl, H.eds. (1995), Helvetica Chimica Acta, CH-4010 Basel, Switzerland.

[0037] Throughout this specification and the following claims, unless the context otherwise requires, the word "comprise" and variations such as "comprises" and "comprising" shall be understood to imply inclusion of the stated integer or step or group of integers or steps, but not to exclude any other integer or step or group of integers or steps. In the following paragraphs, different aspects of the invention are defined in more detail. Each aspect so defined may be combined with any other one or more aspects unless explicitly stated otherwise. Any feature designated as optional, preferred, or advantageous may be combined with any other one or more features designated as optional, preferred, or advantageous.

[0038] Numerous documents are referenced throughout this specification. Each document referenced herein (including all patents, patent applications, scientific publications, manufacturer's specifications, instructions for use, etc.), whether mentioned above or below, is incorporated herein in its entirety. Nothing herein should be construed as an admission that the invention is not entitled to claim prior invention by virtue of such prior disclosure. Some of the documents referenced herein are characterized by being "incorporated." In the event of a conflict between the definitions or teachings of such incorporated references and those listed in this specification, the text of this specification shall prevail.

[0039] The elements of the invention will be described below. These elements are listed together with specific embodiments; however, it should be understood that they can be combined in any manner and in any number to create additional embodiments. The various examples and preferred embodiments described should not be construed as limiting the invention to the explicitly described embodiments. This specification should be understood to support and cover embodiments that combine the explicitly described embodiments with any number of the disclosed and / or preferred elements. Furthermore, unless the context otherwise requires, any permutation and combination of all elements described in this application should be considered as disclosed by the description of this application. definition

[0040] The following provides definitions for some terms that are frequently used in this specification. These terms, in each instance of their use, have their respective defined and preferred meanings in the remainder of the specification.

[0041] As used in this specification and the appended claims, the singular forms “a,” “an,” and “this” include the plural designation, unless otherwise expressly provided.

[0042] The term "sequence comparison" is used herein to refer to the process in which one sequence acts as a reference sequence for comparison with a test sequence. When using a sequence comparison algorithm, the test and reference sequences are input into a computer, and subsequence coordinates are specified if necessary, along with the sequence algorithm program parameters. Default program parameters are typically used, or alternative parameters may be specified. The sequence comparison algorithm then calculates the percentage sequence similarity of the test sequence relative to the reference sequence based on the program parameters. When comparing two sequences and no reference sequence is specified for comparison to calculate the percentage sequence similarity, the longer of the two sequences to be compared should be used to calculate the sequence similarity unless otherwise explicitly stated. If a reference sequence is specified, the sequence similarity is determined based on the full length of the reference sequence indicated by one of the SEQ ID NOs of this invention unless otherwise explicitly stated.

[0043] The sequence alignment methods used for comparison are well known in the art. The optimal alignment of sequences for comparison can be performed, for example, by Smith and Waterman’s local homology algorithm (Adv. Appl. Math. 2:482, 1970), by Needleman and Wunsch’s homology alignment algorithm (J. Mol. Biol. 48:443, 1970), by Pearson and Lipman’s similarity search method (Proc. Natl. Acad. Sci. USA 85:2444, 1988), by computerized implementations of these algorithms (e.g., GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics software package, Genetics Computer Group, 575 Science Dr., Madison, Wis.), or by manual alignment and visual inspection (see, for example, Ausubel et al., Current Protocols in Molecular Biology (1995 supplement). Suitable algorithms for determining percentage sequence similarity and sequence identity are BLAST and BLAST. Algorithm 2.0 is described in Altschul et al. (Nuc. Acids Res. 25:3389-402, 1977) and Altschul et al. (J. Mol. Biol. 215:403-10, 1990), respectively. Software for BLAST analysis is publicly available from the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ). This algorithm involves first identifying high-scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which match or satisfy a threshold score T when compared to words of the same length in the database sequence. T is called the neighbor word score threshold (Altschul et al., ibid.). These initial neighbor word hits are used as seeds to start the search to find longer HSPs containing them. Word hits extend in both directions along each sequence, as long as the accumulated alignment score can increase. For nucleotide sequences, parameters M (reward score for matching residue pairs; always >0) and N are used. (Penalty for mismatched residues; always <0) Calculate the cumulative score. For an amino acid sequence, the cumulative score is calculated using a scoring matrix. Word hit extension in each direction stops when: the cumulative alignment score decreases by an amount X from its maximum value; the cumulative score becomes 0 or below due to the accumulation of one or more negatively scored residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the alignment sensitivity and speed.The BLASTN program (for nucleotide sequences) uses a word length of 11 (W), an expected value of 10 (E), M=5, N=-4, and a comparison of the two strands as default values. For amino acid sequences, the BLASTP program uses a word length of 3 and an expected value of 10 (E), and a BLOSUM62 score matrix (see Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89:10915, 1989) with an alignment of 50 (B), an expected value of 10 (E), M=5, N=-4, and a comparison of the two strands as default values. The BLAST algorithm also performs statistical analysis of the similarity between two sequences (see, for example, Karlin and Altschul, Proc. Natl. Acad. Sci. USA 90:5873-87, 1993). One measure of similarity provided by the BLAST algorithm is the minimum total probability (P(N)), which provides an indication of the probability that a match will occur by chance between two nucleotide or amino acid sequences. For example, if the minimum total probability in a comparison of the test nucleic acid with a reference nucleic acid is less than about 0.2, typically less than about 0.01, and more typically less than about 0.001, then the nucleic acid is considered to be similar to the reference sequence.

[0044] The terms “nucleic acid” and “nucleic acid molecule” are used synonymously herein and are understood to refer to, as is generally accepted in the art, a single-stranded or double-stranded oligomer or polymer of deoxyribonucleotide or ribonucleotide bases, or both. As used herein, the term “nucleic acid” includes not only deoxyribonucleic acid (DNA) and ribonucleic acid (RNA), but also all other linear polymers in which the bases adenine (A), cytosine (C), guanine (G), and thymine (T) or uracil (U) are arranged in their corresponding sequences (nucleic acid sequences). The invention also includes the corresponding RNA sequence (in which thymine is replaced by uracil), complementary sequence, and sequence having a modified nucleic acid backbone or 3' or 5' end. However, nucleic acids in the form of DNA are preferred.

[0045] As used herein, the term "DNA-modifying enzyme" describes any enzyme capable of manipulating the structure of nucleic acids in the genome. More specifically, this term refers to the corresponding enzyme capable of catalyzing recombination, selected from excision, integration, inversion, and translocation reactions. This term encompasses proteins selected from recombinases, integrases, adenosine base editors (ABEs), zinc finger nucleases, transcription activator-like effector nucleases, and Cas nucleases. Such enzymes typically exist in dimer or tetramer form, meaning that the enzyme contains two or more subunits that may be identical or different. Specific preferred examples of DNA-modifying enzymes are recombinases, specifically including site-specific recombinases (SSRs) and, more specifically, tyrosine recombinases (Y-SSRs). Recombinases can exist in monomeric, dimeric, or tetrameric form. Dimers and tetramers may each contain two or four monomers of the same protein with recombinase activity (homodimers or homotetramers), or alternatively, two or more different monomers with recombinase activity (heterodimers or heterotetramers).

[0046] As used herein, the term "target site" (sometimes also called "recognition site") refers to a specific nucleotide sequence that a DNA editing enzyme recognizes and where DNA modifications (such as breaks and strand exchanges) occur, either in or near the site. An example of a target site is a recombinase target site. The target or recognition site of a recombinase is typically 30 to 200 base pairs in length and consists of two inverted repeat recombinase-binding regions flanking a central spacer sequence (Meinke et al., 2016). An example of such a recognition site can be seen in the SSR Cre / loxP binding complex, where the Cre recombinase binds to a 34-base-pair loxP target sequence. The LoxP recognition site contains two 13-base-pair inverted repeat Cre binding elements flanking an 8-base-pair spacer region. The left half of the site is the 13-base-pair binding element to the left of the spacer region, and the right half is the 13-base-pair binding element to the right of the spacer region. Depending on the number and relative orientation of the recognition site and its spacer region, DNA recombinases can excise, integrate, invert, or replace gene content (reviewed in Meinke et al., 2016). Therefore, the target or recognition site according to the invention is preferably a nucleotide sequence comprising a first half-site, a second half-site, and a spacer region separating the first and second half-sites. Another example of a target site is the target site of a Cas-like gene editor. The target site of the Un1Cas12f1 enzyme, for example, has a 4 bp PAM site (TTTR), followed by a 20 bp recognition site for sgRNA. Other target sites of DNA-modifying enzymes are, for example, the Un1Cas12f1 target site (e.g., disclosed in Xin et al., 2022), the Cas9 target site (e.g., disclosed in Cong et al., 2013), the TALEN target site (e.g., disclosed in Gaj et al., 2013), and the homing endonuclease target site (e.g., disclosed in Stoddard, 2006). All scientific references cited herein are incorporated herein by reference.

[0047] For a recombination event to occur, the recombinase complex recognizes a first recognition site and a second recognition site on the DNA double helix. These recognition sites are also called upstream and downstream recognition sites, depending on their location on the DNA double helix.

[0048] In symmetric target sites, the first half-site (e.g., the left half-site) and the second half-site (e.g., the right half-site) are identical and palindromic (reverse complementary). In asymmetric target sites, the first half-site (e.g., the left half-site) and the second half-site (e.g., the right half-site) are not identical and are not palindromic, meaning they differ from each other at least one nucleotide.

[0049] As used herein, the term "variant of DNA modifying enzyme" refers to a DNA modifying enzyme that carries modifications in its amino acid sequence compared to a reference DNA modifying enzyme. Preferably, the reference DNA modifying enzyme is a naturally occurring DNA modifying enzyme or a known DNA modifying enzyme, such as those described in the art. The variant of the DNA modifying enzyme preferably exhibits at least one amino acid modification compared to the reference DNA modifying enzyme. The term "at least one amino acid modification" as used in the context of a variant of the DNA modifying enzyme is not limited to a specific number of amino acid modifications, but preferably includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more amino acid modifications compared to the reference DNA modifying enzyme. Amino acid modifications as used herein include substitution, insertion, and excision. Examples of various expression libraries of DNA-modifying enzymes are disclosed in Buchholz and Stewart, 2001, Lansing et al., 2019, Hoersten et al., 2022, and Lansing et al., 2022. Similarly, as used herein, the term “variant of one or more target sites” means a target site carrying at least one modification in its nucleic acid sequence compared to a reference target site. The terms “at least one nucleic acid modification” or “at least one nucleotide modification” (the two terms are used interchangeably herein) as used in the context of variants of target sites are not limited to a specific number of nucleotide modifications, but include at least one, and preferably 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotide modifications compared to a reference target site. Nucleotide modifications as used herein include substitution, insertion, and excision. Preferably, the reference target site is a naturally occurring target site or a known target site, such as the target site of a DNA-modifying enzyme as described in the art.

[0050] As used herein, the term “unique molecular identifier” (“UMI”) refers to a type of molecular barcoding. A molecular barcode is a short sequence that uniquely marks a molecule in a sample library. UMIs are known to those skilled in the art and are described in detail in Zurek et al., 2020 or Karst et al., 2021, both of which are incorporated herein by reference in their entirety. According to a preferred embodiment of the invention, a UMI is an oligonucleotide comprising at least 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 random nucleotides. According to a particularly preferred embodiment, a UMI is an oligonucleotide comprising at least 50 random nucleotides. The UMI may preferably form part of a UMI tag, which may further comprise one or more sequences downstream and / or upstream (e.g., flanking the random nucleotides) of the UMI that can serve as one or more primer binding sites. The UMI tag may also include one or more, preferably at least two, restriction sites, preferably downstream and / or upstream of the random nucleotide of the UMI. According to a preferred embodiment, the UMI tag includes a first primer-binding site, a first restriction site, the random nucleotide of the UMI, a second restriction site, and a second primer-binding site.

[0051] As used herein, the term "cell" or "host cell" refers to an intact cell, that is, a cell with an intact membrane that has not yet released its normal intracellular components, such as enzymes, organelles, or genetic material. An intact cell is preferably a living cell, that is, a living cell capable of performing its normal metabolic functions. Preferably, the term refers to any cell that can be transfected or transformed with exogenous nucleic acids.

[0052] As used herein, the term "regulatory nucleic acid sequence" refers to the gene regulatory region of DNA. In addition to the promoter region, this term encompasses the operator gene region further away from the gene, as well as nucleic acid sequences that affect gene expression, such as cis-elements, enhancers, or silencers. As used herein, the term "promoter region" refers to the nucleotide sequence on DNA that allows for the regulated expression of a gene. The promoter region allows for the regulated expression of nucleic acids encoding the corresponding protein. The promoter region is located at the 5' end of the gene, and therefore precedes the coding region. Both bacterial and eukaryotic promoters are applicable to this invention.

[0053] The term “identical” in the context of two or more nucleic acid or polypeptide sequences herein means identical, i.e., two or more sequences or subsequences containing the same nucleotide or amino acid sequence. If sequences have the same nucleotide or amino acid residue sequence, they are “identical” to each other (= 100% identity). The term “substantially identical” as used in the context of two nucleic acid or polypeptide sequences means that, when compared and aligned to the maximum correspondence within a comparison window or specified region, the two compared sequences share at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with the specified sequence, as measured by manual alignment and visual inspection. The percentage identity of “substantially identical” also applies to sequences that are “substantially anticomplementary” to each other.

[0054] As used herein, the term "one or more" refers to the number of corresponding entities represented by this term. Specifically, it encompasses one, two, three, four, five, six, seven, eight, nine, ten, or more than ten, preferably two, three, or four, and more preferably two or three. For example, the term "one or more target sites" preferably refers to one target site, two target sites, three target sites, four target sites, or more than four target sites, such as five, six, or seven target sites. According to a preferred embodiment, the term "one or more target sites" refers to two target sites or three target sites. According to a particularly preferred embodiment, this term refers to two target sites.

[0055] The terms “expression vector library” and “expression library” are used interchangeably herein. An expression library generally refers to a collection of plasmids (or phages) containing a representative sample of DNA or genomic fragments constructed in a manner that enables transcription and translation by a host organism into which the expression library is introduced. In the context of this invention, an expression library is used for expression cloning of variants of the DNA-modifying enzymes disclosed herein, or for expression cloning of variants of the DNA-modifying enzymes and one or more target sites of the DNA-modifying enzymes disclosed herein. Expression cloning is a technique in DNA cloning that uses an expression vector to produce a clonal library, each clone expressing one protein, in this case a DNA-modifying enzyme (a variant). The expression library is screened for properties of interest, and clones of interest are recovered for further analysis. Expression cloning is a method well known in the art and is described in further detail, for example, in Lodes et al., 2004, which is incorporated herein by reference. Implementation Plan Description

[0056] According to the present invention, a method for screening and sequencing multiple DNA modifying enzymes includes the following steps: Provides an expression vector library containing a variety of different vectors, wherein each vector contains a first region encoding a variant of a DNA-modifying enzyme and a second region containing one or more target sites of the DNA-modifying enzyme; Introduce the expression vector library into the host cell; Culture host cells and express DNA-modifying enzymes; Plasmid DNA was isolated from host cell cultures; At least the first and second regions of the expression vector were sequenced; and Based on the sequencing results, determine whether the DNA sequence of the second region on the expression vector has been altered by DNA-modifying enzymes.

[0057] According to the method described and as defined above, the DNA-modifying enzyme is selected from recombinases, integrases, adenosine base editors (ABEs), zinc finger nucleases, transcription activator-like effector nucleases, and Cas nucleases. The DNA-modifying enzyme can also be a broad range of nucleases (also known as homing endonucleases). According to a preferred embodiment, the DNA-modifying enzyme is a recombinase or an integrase. According to a particularly preferred embodiment, the DNA-modifying enzyme is a recombinase.

[0058] According to one embodiment of the invention, the DNA-modifying enzyme is a monomer. According to other preferred embodiments of the invention, the DNA-modifying enzyme is a dimer comprising at least two protein monomers (also referred to as subunits). The terms monomer and subunit are used interchangeably herein. The dimer may comprise two monomers of the same type (i.e., two identical protein monomers, homodimer) or two monomers of different types (i.e., two different protein monomers, heterodimer). According to other preferred embodiments of the invention, the DNA-modifying enzyme comprises at least four protein monomers, i.e., it is in tetrameric form (tetramer). Such a tetramer may comprise four monomers of the same type (i.e., four identical protein monomers, homotetramer), or monomers of different types, such as two, three, or four different monomers (heterotetromer). According to a particularly preferred embodiment, the DNA-modifying enzyme comprises a heterodimer or heterotetramer containing different protein monomers, i.e., a heterodimer or heterotetramer.

[0059] According to the present invention, each expression vector may encode one variant or multiple variants of a DNA-modifying enzyme. Consistently, each expression vector may encode a first monomer and a second monomer of the DNA-modifying enzyme, which may be the same or different. Similarly, each expression vector may encode a first, second, third, and fourth monomer of the DNA-modifying enzyme, which may be the same or different. In the case where the vector encodes multiple variants of the DNA-modifying enzyme, the vector preferably encodes two or four variants (or different monomers) of the DNA-modifying enzyme. In this case, the encoded variants of the DNA-modifying enzyme preferably act together to modify DNA. The two or four variants of the DNA-modifying enzyme preferably act together in the form of a dimer or tetramer as defined herein.

[0060] The DNA-modifying enzyme can be a naturally occurring DNA-modifying enzyme, or it can be a variant of a naturally occurring DNA-modifying enzyme. According to a preferred embodiment of the invention, the DNA-modifying enzyme is an evolved enzyme. Preferably, as described, for example, in Buchholz and Stewart 2001, Buchholz and Hauber 2011, Lansing et al. 2019, Hoersten et al. 2021, and Lansing et al. 2022, the DNA-modifying enzyme has been evolved using substrate-associated directed evolution (SLiDE), all of which are incorporated herein by reference in their entirety.

[0061] According to a preferred embodiment, the variant of the DNA-modifying enzyme is based on a naturally occurring or known DNA-modifying enzyme and contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more amino acid modifications compared to the naturally occurring or known DNA-modifying enzyme (also referred to herein as a reference DNA-modifying enzyme). When evolving the DNA-modifying enzyme, the variant of the DNA-modifying enzyme in the first expression library preferably contains a small number of amino acid modifications, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 amino acid modifications compared to the reference DNA-modifying enzyme. During the evolution of the DNA-modifying enzyme, the number of amino acid modifications is preferably increased in each additional expression library compared to the reference DNA-modifying enzyme. Examples of the corresponding evolutionary processes are described in Buchholz and Stewart 2001, Buchholz and Hauber 2011, Lansing et al. 2019, Hoersten et al. 2021 and Lansing et al. 2022, all of which are cited in whole and incorporated herein by reference.

[0062] According to a particularly preferred embodiment of the present invention, the protein having recombinase activity is a recombinase, more preferably a site-specific recombinase, and even more preferably a tyrosine site-specific recombinase.

[0063] According to an optional preferred embodiment, the DNA-modifying enzyme is an integrase.

[0064] According to the present invention and as defined above, the target site of a DNA-modifying enzyme is a nucleotide sequence. Particularly in the case of a recombinase as a DNA-modifying enzyme, the target site comprises a first half-site, a second half-site, and a spacer region separating the first and second half-sites as defined above. According to a preferred embodiment, the target site may be a known target site of the corresponding DNA-modifying enzyme. However, the target site is not limited to specific target sites that may be known in the art for a DNA-modifying enzyme. The target site may be a modified version of a known target site. Thus, the target site may be a naturally occurring target site of a DNA-modifying enzyme, or it may be an artificially created target site, such as a modified version of a known target site, preferably a modified version of a known naturally occurring target site. Such modified versions of known target sites preferably differ from known target sites in 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides. Optionally, the modified target site may differ from a known target site at about 2%, about 4%, about 6%, about 8%, about 10%, about 12%, about 14%, about 16%, about 18%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, or about 50% of the nucleotides.

[0065] According to a preferred embodiment, the variant of the target site is based on a naturally occurring or known target site of the DNA-modifying enzyme, and contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotide modifications compared to a naturally occurring or known target site of the DNA-modifying enzyme (also referred to herein as a reference target site).

[0066] The expression vectors according to the invention are not limited to any particular expression vector. Typically, an expression vector contains an origin of replication, a promoter, and a specific gene sequence that allows phenotypic selection of the host cell containing the vector. According to the invention, the vector contains a first region encoding a variant of a DNA-modifying enzyme, and a second region containing at least one, at least two, at least three, or more target sites of the DNA-modifying enzyme, as defined above. If more than one target site is present in the vector, the target sites are preferably substantially identical or substantially anticomplementary to each other.

[0067] According to a particularly preferred embodiment, the vector contains two target sites of a DNA-modifying enzyme. In this case, the DNA-modifying enzyme is preferably a recombinase or an integrase. According to a particularly preferred embodiment, when two target sites are present in the vector, the DNA recombinase is a recombinase.

[0068] A particularly preferred expression vector used in this invention is the pEVO vector as described in Buchholz and Stewart 2001.

[0069] In the case of using evolved DNA-modifying enzymes in the method of the present invention, the gene encoding the DNA-modifying enzyme is preferably excised from the corresponding evolutionary library, such as... Figure 1 and 6 As shown. The corresponding evolutionary libraries are described, for example, in Buchholz and Stewart 2001 and Lansing et al. 2022.

[0070] According to a particularly preferred embodiment of the invention, the expression vector further comprises a unique molecular identifier (UMI) as defined above. The location of the UMI within the vector is not particularly limited. However, according to a preferred embodiment, the unique molecular identifier is located in a first region of the expression vector adjacent to the sequence encoding a DNA-modifying enzyme. According to a particularly preferred embodiment, the UMI is located downstream of the sequence encoding the DNA-modifying enzyme, and therefore upstream of a second region containing one or more target sites of the DNA-modifying enzyme, such as, for example… Figure 1 and Figure 6As shown. According to a particularly preferred embodiment, the UMI forms part of a UMI tag, further comprising, in addition to the random nucleotide sequence of the UMI, additional sequences such as one or more primer binding sites and / or additional sequences such as one or more restriction sites. According to one embodiment, the primer binding site and restriction site are preferably located upstream and downstream of the random nucleotide, for example, a UMI tag according to a preferred embodiment of the invention has the following structure: primer binding site 1 – restriction site 1 – random nucleotide – restriction site 2 – primer binding site 2. Exemplary and preferred UMI tags according to the invention have sequences according to SEQ ID NO: 17 or SEQ ID NO: 18. These exemplary sequences contain 50 random nucleotides at positions 43 to 92, referred to as UMI. In the sequences, the random nucleotides are represented by the letters “N”, “Y”, and “R”, where “N” represents any nucleotide selected from A, C, G, and T; where “Y” represents any nucleotide selected from C and T; and where “R” represents any nucleotide selected from A and G. Primer binding sites are located at positions 1 to 20 and 101 to 120. In SEQ ID NO: 17, restriction site 1 (BsiWI) is located at positions 21 to 26, and restriction site 2 (SbfI) is located at positions 93 to 101. In SEQ ID NO: 18, restriction site 1 (XbaI) is located at positions 21 to 26, and restriction site 2 (SbfI) is located at positions 93 to 100.

[0071] Expression vector libraries can contain multiple different libraries. Each of the multiple libraries is preferably different from the other libraries in the number of modifications introduced into a variant of a DNA-modifying enzyme encoded by the respective expression vector or a variant of a target site. For example, a first library of multiple different libraries can contain multiple different vectors, each encoding a variant of a DNA-modifying enzyme that has one, two, or three modifications compared to a reference DNA-modifying enzyme (such as a naturally occurring or known DNA-modifying enzyme). A second library of multiple different libraries can contain multiple different vectors, each encoding a variant of a DNA-modifying enzyme that has one, two, or three modifications compared to the variant of the DNA-modifying enzyme encoded in the first library. Then, a third library can encode a variant of a DNA-modifying enzyme that has one, two, or three modifications compared to the variant of the DNA-modifying enzyme encoded in the second library, or in other words, seven, eight, or nine modifications compared to a reference DNA-modifying enzyme, and so on. It should be understood that the number of modifications referred to in this example is not particularly limited and is for illustrative purposes only. This also applies to variations of the target sites according to the method of the present invention, i.e., the first library contains target sites with a limited number of nucleotide modifications compared to the reference target site, the second library contains target sites with the same or similar number of nucleotide modifications compared to the target sites of the first library, and so on.

[0072] According to the present invention, an expression vector library is introduced into a host cell. The host cell, as understood in this invention, is a naturally occurring cell or cell line (optionally transformed or genetically modified) containing at least one vector as described above. Therefore, the present invention includes host cells containing at least one expression vector according to the present invention as a plasmid. According to a preferred embodiment, the host cell contains an expression vector according to the present invention as a plasmid. This embodiment is particularly useful if bacteria are used as the host cell.

[0073] According to one embodiment of the invention, the host cell is selected from bacterial cells, yeast cells, fungal cells, and mammalian cells. Suitable bacterial cells include strains of Gram-negative bacteria (such as Escherichia coli). Escherichia coli ), Proteus spp. Proteus ) and Pseudomonas spp. Pseudomonas ) strains) and Gram-positive bacterial strains (such as Bacillus spp.) Bacillus Streptomyces ( Streptomyces Staphylococcus spp. Staphylococcus ) and Lactococcus spp. Lactococcus Cells of strains of Trichoderma. Suitable fungal cells include those from the genus Trichoderma. Trichoderma ), Neurospores ( Neurospora ) and Aspergillus ( AspergillusCells from species. Suitable yeast cells include those from the genus *Saccharomyces* (Yeast). Saccharomyces (e.g., brewer's yeast) Saccharomyces cerevisiae )), genus *Fissionyomyces* ( Schizosaccharomyces (For example, *Schizosaccharomyces cerevisiae*) Schizo saccharomyces pombe )), Pichia pastoris ( Pichia (For example, Pichia pastoris) Pichia pastoris ) and Pichia pastoris ( Pichia methanolicd )) and Hansenula genus ( Hansenula The host cell is a bacterial cell. Suitable mammalian cells include, for example, CHO cells, BHK cells, HeLa cells, COS cells, 293 HEK cells, etc. However, amphibian cells, insect cells, plant cells, and any other cells used in the art for expression can also be used. According to a preferred embodiment, the host cell is a bacterial cell. According to a particularly preferred embodiment, the host cell is an *Escherichia coli* cell.

[0074] Nucleic acids are introduced into host cells using gene manipulation techniques known to those skilled in the art. Suitable methods include cell transformation, transfection, or viral infection, thereby introducing a nucleic acid sequence encoding a DNA-modifying enzyme as a component of a vector into the cells. Cell culture is performed using methods known to those skilled in the art for culturing the appropriate cells. Therefore, it is preferable to transfer cells to a conventional culture medium and culture them in a temperature and gas atmosphere that favors cell survival and allows for the expression of the encoded protein. A preferred method for introducing an expression vector library into host cells is transformation, preferably by electroporation. As an example of the steps for introducing an expression vector library into host cells, E. coli cells can be transformed using an electroporation vector.

[0075] According to one embodiment of the method of the invention, host cells containing the expression vector are plated on a suitable culture medium (such as an agar plate) to allow for the growth and selection of individual cell communities. This step can also be used to select a specific number of individual communities instead of culturing all cells simultaneously, thereby reducing the total number of library members and the number of enzyme variants to be screened. Optionally, the host cells containing the expression vector can be briefly cultured in a suitable culture medium (e.g., for a limited time, such as 0.5 to 1 hour), and an amount of culture medium equivalent to the desired number of variants can be removed from this medium for further culture, as described below. The number of transformed bacteria present per μl of culture medium can be estimated based on the number of communities on the plate. If, for example, 10 μl of culture medium is plated on a plate, and the cultured plate shows 100 different communities, then 1 μl of culture medium contains 10 transformed bacteria.

[0076] According to the invention, host cells containing an expression vector library or cells selected as described above are cultured to allow plasmid expression. Culture conditions depend on the cell line used and are not particularly limited. However, it is not necessary to culture each selected community separately. The selected communities can be cultured together in a single cell culture. Optionally, host cells containing an expression vector library or cells selected as described above are cultured in more than one separate cell culture. To activate the expression of the nucleic acid encoding a DNA-modifying enzyme, the nucleic acid encoding the DNA-modifying enzyme further includes a regulatory nucleic acid sequence, preferably a promoter region. Thus, the expression of the nucleic acid encoding the DNA-modifying enzyme is initiated or regulated by activating the regulatory nucleic acid sequence. Cells are preferably cultured under conditions that allow for cell growth and protein expression. Preferably, cell growth is allowed for more than 2 to 12 hours or longer, allowing for more than 2, preferably more than 3, more than 4, more than 5, more than 6, more than 7, more than 8, more than 9, or more than 10 cell doublings.

[0077] Once a DNA-modifying enzyme is expressed in a host cell culture, if the enzyme recognizes a corresponding target site, it begins to modify DNA at one or more target sites in the second region of the expression vector. For example, in the case of a recombinase, the recombinase complex recognizes a first and a second target site on the plasmid to initiate a recombination event. In the case of a DNA segment excision event, the segment is flanked by two target sites oriented in the same direction and mediated by a corresponding recombinase protein encoded on the same vector. If such a recombination event occurs, the DNA segment between the two target sites is excised during cell culture. This invention also covers alternative modifications to the DNA in the second region of the vector, such as, but not limited to, inversions, integration events, or single-nucleotide level events in the second region of the vector. To determine whether any modification has occurred in the second region of the vector, plasmid DNA is isolated from the host cell culture after culturing host cells and expressing a library encoding a DNA-modifying enzyme. Isolation of plasmid DNA from the cell culture can be performed by any conventional method well known in the art. After isolation, the plasmid DNA is sequenced to determine its nucleotide sequences in at least the first and second regions. It should be understood that, in addition to the first and second regions containing sequences encoding one or more DNA-modifying enzymes and one or more target sites, more or all of the plasmid can be sequenced. However, for high-throughput screening, sequencing of the first and second regions is preferred because the information obtained by sequencing the first and second regions is sufficient for the method of the present invention. This sequencing step allows determination of whether the sequence of a specific DNA-modifying enzyme, the sequence of the target region, and the DNA sequence in the second region on the expression vector have been altered by the DNA-modifying enzyme. According to a preferred embodiment, the first and second regions of the vector are sequenced in a single method step, i.e., they are sequenced together rather than sequentially.

[0078] There are no particular limitations on the sequencing method. However, long-read sequencing methods are preferred, especially for high-throughput methods. A particularly preferred sequencing method according to the invention is as described in Wang et al., 2021 (incorporated herein), or nanopore sequencing, as provided by, for example, Oxford Nanopore Technologies, UK. In short, nanopore sequencing works by monitoring changes in ionic current as nucleic acids pass through protein nanopores. The resulting signal is decoded to provide a specific DNA sequence. Based on the sequencing results, the activity rate of each DNA-modifying enzyme at its corresponding target site can be determined.

[0079] In cases where the sequence of the DNA-modifying enzyme is known, screening sequence data is compared with the known sequence of the DNA-modifying enzyme to identify which sequence reads belong to which DNA-modifying enzyme. The DNA-modifying enzyme sequences may be known because they have been previously sequenced or because they were synthesized according to defined sequences. Sequence reads are separated based on their sequence matches and processed individually from there for generating consensus sequences and determining activity rates. In this case, it is preferable not to use UMIs and UMI tags in the method of the present invention.

[0080] In cases where UMIs or UMI tags are used, the method according to the invention may further include the following steps: clustering unique molecular identifiers, generating and polishing common sequences, determining the number of DNA modification events for each DNA modifying enzyme, and determining the activity rate of each DNA modifying enzyme. The clustering and polishing of individual sequencing results are described in Zurek et al., 2020 and Karst et al., 2021, both of which are incorporated herein by reference in their entirety. For example, UMI sequences of sequencing reads can be identified by comparing upstream and / or downstream sequences of the UMI. In cases where upstream sequences are used to identify UMIs, DNA sequences with a UMI length after matching are collected together with the sequence read ID. In cases where downstream sequences are used to identify UMIs, DNA sequences with a UMI length before matching are collected together with the sequence read ID. In cases where both downstream and upstream DNA are used to identify UMIs, DNA sequences between matches are collected together with the sequence read ID. The collected UMI sequences are then clustered based on sequence identity while maintaining their association with their read IDs, for example using vsearch (Rognes et al., 2016). Based on these clusters and their associated read IDs, the sequence reads are then divided into separate files, each containing reads corresponding to a UMI cluster. From each of these clusters of separated reads, a portion of the sequence containing the DNA-modifying enzyme gene is aligned with a reference gene containing high similarity to the screened variant. Typically, such genes are identified by sequencing the screened DNA-modifying enzyme library or by selecting DNA-modifying enzyme genes modified to generate the screened variant library. Using this alignment, software for shared sequence generation and sequence polishing can be used to determine the most probable gene sequence of the DNA-modifying enzyme associated with this cluster based on all sequence reads. In this respect, polishing refers to performing additional rounds of processing on the reads to improve sequence accuracy.

[0081] There are no particular limitations on the methods used to determine the activity rate. For example, particularly for DNA-modifying enzymes that cause five or more base alterations in DNA, sequencing reads that have been isolated by sequence comparison or UMI clustering are compared with expected sequences from a second region in both altered and unaltered sequences. Based on this sequence comparison, the number of reads matching the altered DNA can be counted. The change in count divided by the total number of reads matching the second region yields the activity rate of the DNA-modifying enzyme associated with the isolated reads. A particularly preferred example of a five or more base alteration in DNA is the excision of 741 bp in the pEVO expression vector using a recombinase (Examples 2 and 5).

[0082] In cases where DNA-modifying enzymes cause changes smaller than 5 bp, the location of the expected change in the second region of the sequencing read can be identified by comparing the sequence with that of the second region. Sequences at the change locations in the sequencing read are collected to count all observed results. The count of the expected DNA changes is divided by the total number of reads matching the second region to obtain the activity rate of the DNA-modifying enzyme associated with the isolated read.

[0083] In cases where DNA-modifying enzymes are screened at multiple target sites, additional references to a second region can be provided to correlate the activity rate of the identified DNA-modifying enzyme with multiple target sites. The additional references are sequences of the second region where the target sites are replaced by alternative target sites used in the screening. These references can also be provided in altered and unaltered versions to determine the activity rates as described above.

[0084] According to a preferred embodiment of the invention, prior to sequencing, the first and second regions of the expression vector are excised from the expression vector, and the remaining portion of the vector is discarded. Excision can be performed using any conventional method known in the art. A preferred method for excising portions of the vector is to use restriction enzymes. Excision of the first and second regions of the expression vector allows the sequencing reaction to occur only on the relevant portions of the expression vector, i.e., those portions containing sequences encoding DNA-modifying enzymes, optionally present UMIs, target sequences, and segments readily modifiable by DNA-modifying enzymes.

[0085] According to a particularly preferred embodiment of the invention, a gene encoding a DNA-modifying enzyme is cloned into an expression vector along with a unique molecular identifier (UMI), the expression vector further containing one or more target sites of interest. The barcoded plasmid is transformed into appropriate cells, and the cells are plated on one or more culture plates. Based on the number of colonies on the plates, a defined number of clones are cultured in a medium that induces expression of the gene encoding the DNA-modifying enzyme. This results in multiple copies of the transformed plasmid, a portion of which is modified by the DNA-modifying enzyme. The plasmid is isolated and sequenced, preferably using nanopore sequencing. The sequences are then clustered using the UMI. These clusters are used to construct an accurate common sequence for the variant gene. Thus, a cluster refers to an enzyme variant identified using the method of the invention. The counting of recombinant plasmids yields the editing rate of a specific variant of the DNA-modifying enzyme at the corresponding target site.

[0086] The present invention, as described above, can also be used to screen and sequence multiple target sites of DNA-modifying enzymes. Therefore, all embodiments and descriptions provided above are applicable to further aspects of the present invention, which provides a method for screening and sequencing multiple target sites of DNA-modifying enzymes, the method comprising the following steps: Provides an expression vector library containing a variety of different vectors, wherein each vector contains a first region encoding a DNA-modifying enzyme and a second region containing a variant of one or more target sites of the DNA-modifying enzyme; Introduce the expression vector library into the host cell; Culture host cells and express DNA-modifying enzymes; Plasmid DNA was isolated from host cell cultures; Sequencing of the first and second regions of the expression vector; and Based on the sequencing results, determine whether the DNA sequence of the second region on the expression vector has been altered by DNA-modifying enzymes.

[0087] In this second aspect of the invention, the target site in the second region of the vector can be a different naturally occurring target site of a different DNA modifying enzyme. Alternatively, the target site can be an artificially created target site, such as a modified version of a known target site. Such a modified version of a known target site preferably differs from the known target site at 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. Optionally, the modified target site may differ from the known target site at about 2%, about 4%, about 6%, about 8%, about 10%, about 12%, about 14%, about 16%, about 18%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, or about 50% of the nucleotides.

[0088] This invention provides a novel sequencing and screening method for acquiring sequence information and recombination rates of library variants in a high-throughput manner. Specifically, the high-throughput screening method enables the characterization of thousands of DNA-editing enzyme variants at multiple target sites, and also characterizes thousands of target sites of DNA-editing enzyme variants. This method preferably utilizes nanopore technology to sequence full-length enzyme variants in a fast turnaround time. By capturing target sites and enzyme sequences, the DNA editing rate of each enzyme variant can be quantified. Using unique molecular identifiers, the sequencing results can be clustered into highly accurate common sequences. The applicability of the method of the present invention is demonstrated in the following non-limiting examples using two different DNA-modifying enzymes: a variant of the Cre designer recombinase D7 (Lansing et al., 2022) and an evolved Cas12f-derived mini-ABE.

[0089] All references and patent documents cited in this article are included in their entirety. Example Example 1: Preparation of UMI fragments

[0090] To generate the DNA fragments for ligation, single-stranded DNA oligonucleotides containing primer binding sites, restriction sites, and 50 random bases were ordered. Two variants of this oligonucleotide were used, depending on which pEVO plasmid it was intended for (Buchholz and Stewart, 2001). The “Cas12f UMI-tag” (SEQ ID NO: 18) was used to screen for the evolved Cas12f-ABE variant, while the “loxF8 UMI-tag” (SEQ ID NO: 17) was used to screen for the loxF8 recombinase variant. To make these oligonucleotides double-stranded, 50 μl PCR was performed using 20 μM primers UMIprimer F and UMIprimer R (SEQ ID NO: 32 and 33), 10 μM oligonucleotides, 10 μl 5x MyTaq buffer, and 1 μl MyTaq polymerase (Bioline). The PCR cycler was set to 94°C for 90 seconds, followed by 10 cycles: 15 seconds at 94°C, 15 seconds at 54°C, and 15 seconds at 72°C. For the Cas12f UMI-tag, the PCR product was digested with SbfI and XbaI, or for the loxF8 UMI-tag, the PCR product was digested with BsiWI and SbfI. The digests were then purified again using the Isolate II PCR and Gel Kit (Bioline) and measured on a Qubit 2.0 (Thermo Fisher Scientific) using the Qubit HS dsDNA Kit. Example 2: Barcoding of Enzyme Variant

[0091] Evolutionary libraries of Cas12f-ABE and loxF8 recombinases were obtained through directed evolution of the fusion protein Un1Cas12f1-Tad8e (abbreviated as Cas12f-ABE, SEQ ID NO: 27) and the tyrosine site-specific recombinase Cre (SEQ ID NO: 31). The resulting enzyme variants contained mutations randomly obtained compared to their sources. DNA editing enzyme gene variants were obtained from the pEVO plasmid used in evolution by digesting the evolved Cas12f-ABE library with XbaI and BsrGI-HF or digesting the loxF8 library with SacI and BsiWI. The Cas12f-ABE library and the Cas12f-ABE control were ligated at a ratio of 60 ng Cas12f-ABE fragment, 4.8 ng UMI tag, and 100 ng pEVO-BE plasmid digested with BsrgI-HF and SbfI. The loxF8 library and D7 control (Lansing et al. 2022, SEQ ID NO: 19 and 20) were ligated at a ratio of 40 ng recombinase gene fragment, 1.9 ng loxF8 UMI tag and 120 ng SacI and SbfI digested pEVO-loxF8 plasmid.

[0092] The ligated plasmid was desalted in distilled water using an MF-Millipore membrane filter (Merck) for 30 minutes and then transformed into XL-1 Blue Escherichia coli (Agilent) via electroporation. The transformed bacteria were cultured in SOC medium at 37°C for 30 minutes. 2 μl of this culture was plated onto an agarose plate containing 15 mg / ml chloramphenicol and incubated overnight at 37°C. The colony count on the plate was performed to determine the number of transformed bacteria per μl of SOC culture.

[0093] To specify the number of variants used for screening, an amount of SOC culture equivalent to the desired number of variants was cultured overnight in 100 ml LB medium containing 25 mg / ml chloramphenicol and a defined amount of L-arabinose. For Cas12f-ABE screening, approximately 4,000 transformed bacteria from each library and approximately 100 transformed bacteria from each control were cultured with 10 μg / ml L-arabinose. For loxF8 recombinase screening, approximately 4,000 transformed bacteria from the library and 50 transformed bacteria from the control were cultured with 1 μg / ml L-arabinose. For each sequencing run, different libraries and controls were cultured together, and plasmid DNA was extracted from these cultures using the GeneJet Plasmid Miniprep kit (Thermo Fisher Scientific).

[0094] To test the same recombinase variant at multiple target sites, recombinase DNA selected from loxF8 was digested with SacI and SbfI and isolated from agarose gels using Isolate II PCR and a gel kit (Bioline). The recombinase dimer fragment was then cloned into a pEVO plasmid containing the off-targets of interest (HG1, HG2, HG2L, SEQ ID NO: 12-14; these specific off-targets were identified in the human genome as having high similarity to the target site). The ligated plasmid was desalted using an MF-Millipore membrane filter (Merck) and transformed into XL-1 Blue *E. coli* (Agilent) via electroporation. The transformed bacteria were cultured in SOC medium at 37°C for 30 min and then transferred to 100 ml LB medium containing 25 mg / ml chloramphenicol and 100 μg / ml L-arabinose. Plasmid DNA was subsequently extracted from these cultures using the GeneJet Plasmid Miniprep kit (Thermo Fisher Scientific).

[0095] Barcoded plasmid extracts from the Cas12f-ABE library were digested with ScaI and BsrgI, while barcoded plasmids selected from loxF8 were digested with ScaI and SacI. Fragments containing evolutionary genes, UMI tags, and target sites were excised from the isolated fragments using an Isolate II PCR and Bioline gel excision kit via agarose gel excision. Example 3: Nanopore sequencing and library screening processing

[0096] The DNA concentration of the digested and isolated plasmid extracts (Example 2) was measured using the Qubit dsDNA HS Assay kit on a Qubit 2.0 fluorometer (Thermo Fisher Scientific). Nanopore sequencing libraries for the loxF8 library were prepared according to the "Amplicons by Ligation (SQK-LSK110)" protocol from Oxford Nanopore Technologies, while the Cas12f-ABE library was prepared using the "Amplicons by Ligation (SQK-LSK112)" protocol. The prepared loxF8 library was then loaded into a MinION FLO-MIN106 flow cell (Oxford Nanopore Technologies) with r9.4.1 wells, while the Cas12f library was loaded into a MinION FLO-MIN110 flow cell (Oxford Nanopore Technologies) with r10.4 wells. Sequencing was performed for 72 hours. Each screening was performed on a single flow cell.

[0097] Base calling of sequence data was performed using a high-precision model of the loxF8 library and an ultra-precision model of the Cas12f-ABE library (Oxford Nanopore Technologies) on guppy version 6.0.1. Sequence data processing was performed on a custom-developed pipeline (available at https: / / github.com / ltschmitt / DEQSeq). First, reads were filtered using Filtlong v0.2.1 (https: / / github.com / rrwick / Filtlong) with a minimum length of 2,900 bp for Cas12f-ABE and a minimum length of 3,000 bp for loxF8. Additionally, Filtlong was used to filter reads with a minimum average Phred quality of 10 for loxF8 and a minimum average Phred quality of 18 for Cas12f-ABE. Then, minimap2 (Li, 2021) was used to align the sequence with reference sequences containing Un1Cas12f1-ABE8e (Cas12f-ABE filter, SEQ ID NO: 34) or D7 (F8 dimer filter, SEQ ID NO: 35) and a UMI consisting of 50 random bases. To ensure gene and UMI coverage, samtools (Danecek et al., 2021) was used to filter aligned reads based on the coordinates of the beginning of the enzyme gene and the end of the UMI.

[0098] The UMIs were then extracted from the filtered alignments using the `stackStringsFromBam` function in the R package `GenomicAlignments` (Lawrence et al., 2013). The UMIs were then clustered using `VSEARCH` (Rognes et al., 2016) with a `cluster_identity` value of 0.7. Sequence reads from clusters with a minimum size of 50 reads were then transferred to separate files and aligned with gene-UMI reference sequences. These separate read files and alignments were used to construct a shared sequence using racon (Vaser et al. 2017), followed by further polishing with medaka (https: / / github.com / nanoporetech / medaka), both using standard settings. The polishing process was run in parallel with GNU (Tange 2023). Finally, the gene sequences were extracted and translated into amino acids using the R package `GenomicAlignments`.

[0099] For loxF8 screening, the DNA excision rate of the enzyme is determined by aligning cluster reads with reference sequences containing target site regions that represent both unrecombined and recombined variants. These references begin 100 bp upstream of the first target site and end 100 bp downstream of the last target site. For each target site sequence, a separate reference is provided to identify different targets. The additional references are essentially the same as the loxF8 recombined and unrecombined references (SEQ ID NO: 36 and 37), the only difference being that the loxF8 target sites are replaced by HG1, HG2, or HG2L (SEQ ID NO: 38 to 43). The recombination rate of the variants is determined based on the read counts from the target site region alignments.

[0100] For Cas12f-ABE screening, the base editing rate of the enzyme variant was determined by comparing cluster reads with a reference containing unedited target sites (SEQ ID NO: 44). Using the GenomicAlignments R package (Lawrence et al., 2013), four-base alignments from positions two to five, counting from the TTTG PAM sequence at all three target sites, were extracted. Base editing in Cas12f-ABE is expected at position three or four after the PAM sequence. Therefore, comparing this region can provide information about potential edits at adjacent positions. The editing rate was then determined by counting correctly edited reads, unedited reads, and other resulting edits.

[0101] The results from each filter were combined and filtered to obtain clusters with 100 or more reads. All further data processing and visualization were performed in R using the tidyverse (Wickham et al. 2019) and stringdist packages (van der Loo, 2014). Example 4: Extraction of enzyme variants

[0102] To verify the screening obtained as described in Example 3, enzyme variants were extracted from the screened library using PCR. Reverse primers specific to the UMI of the clusters of interest (“loxF8 UMI-138 R” (SEQ ID NO: 45), “loxF8 UMI-181 R” (SEQ ID NO: 46), “loxF8 UMI-1244 R” (SEQ ID NO: 47), “Cas12f UMI-2 R” (SEQ ID NO: 48), “Cas12f UMI-3030 R” (SEQ ID NO: 49), “Cas12f UMI-3301 R” (SEQ ID NO: 50)) were designed and used in conjunction with universal forward primers (bound to plasmids (“loxF8 universal F” (SEQ ID NO: 51) for loxF8 selection and “Evolution F” (SEQ ID NO: 52) for Cas12f-ABE selection) before the enzyme-encoding sequence in region 1) to amplify the enzyme gene. PCR was performed using high-fidelity polymerase (Herculase II Fusion DNA polymerase, Agilent) and XbaI and BsrgI. PCR products were digested with Cas12f-ABE enzyme or SacI and BsiWI (loxF8 recombinase) for further cloning. Example 5: Recombinant assay

[0103] A schematic diagram of the recombinant assay used is shown in Figure 3 In section A, pEVO vectors with different target sites were previously disclosed (Lansing et al., 2022). Recombinases were cloned into pEVO vectors with corresponding target sites using SacI and BsiWI restriction enzymes. Recombinase expression was controlled by the L-arabinose inducible promoter system (araBAD). Recombination at each target site on the evolved plasmid resulted in the excision of a 741 bp fragment from the plasmid. The resulting size differences were mediated by recombinase activity and could be detected by restriction digestion followed by gel electrophoresis. Example 6: pEVO carrier of Un1Cas12f1-ABE8e

[0104] To generate plasmids containing target sites, oligonucleotides (SEQ ID NO: 54 and 55) containing the target sites were annealed to form DNA fragments, which were then cloned into a BglII-digested pEVO backbone using Cold Fusion (System Biosciences). SgRNA backbone fragments were synthesized (Twist Biosciences) and cloned into pEVO vectors containing target sites using NsiI and NotI restriction enzymes. The pEVO plasmids containing the sgRNA backbone and target sites were then used as templates in PCR reactions, with the reverse primers containing three distinct sgRNA spacer regions. In this manner, the generated sgRNA fragments were progressively introduced into the pEVO vectors: sgRNA1 was cloned using NotI and NsiI, sgRNA2 using NsiI, and sgRNA3 using XhoI and SalI. The Un1Cas12f1 fragment was generated by Twist Biosciences and amplified by PCR, and the TadA gene was prepared by PCR using the pABE8e-protein midparticle as a template (Addgene plasmid #161788). To create Cas12f-ABE, two PCR products were used as templates for overlapping PCR. Cas12f-ABE was then cloned into the pEVO-TS-sg1-sg2-sg3 vector using BsrGI and XbaI restriction enzymes. Expression of Cas12f-ABE was controlled by the araBAD L-arabinose inducible promoter system. Example 7: Evolution of Cas12f-ABE

[0105] A schematic diagram of the process is shown below. Figure 4 In the first step, a Cas12f-ABE library was generated using error-prone PCR with low-fidelity DNA polymerases (MyTaq, Bioline, and Primers Evolution F and Evolution R (SEQ ID NO: 52 and 53)) and cloned into a vector using BsrGI and XbaI restriction enzymes. After transformation into XL-1 blue E. coli, enzyme expression was induced with 200 μg / ml L-arabinose. Following base-editing expression, plasmids were isolated and digested with restriction enzymes (REs) targeting the sgRNA and its target sites. Because base editing at the target sites occurs at the restriction enzyme sites, digestion with these enzymes linearizes the edited plasmids, while the unedited plasmids are cleaved into two fragments ( Figure 4 This process can be visualized and quantified using agarose gel electrophoresis. Figure 5 A). The next round of directed evolution begins with error-prone PCR on enzymatically digested DNA. Only the edited plasmid is replicated because of the primers ( Figure 4The small arrows indicate that only the edited DNA fragments are valid templates for PCR. The PCR products were then cloned into unedited pEVO vectors, marking the start of a new evolutionary cycle. To increase selection pressure, L-arabinose levels and the resulting induced enzyme expression were reduced from 200 μg / ml to 10 μg / ml. Finally, to enrich libraries for DNA editing quantitative sequencing screening, four cycles of evolution were performed using high-fidelity polymerase (Herculase II Fusion DNA polymerase, Agilent) at 10 μg / ml L-arabinose. Figure 5 C). Example 8: Base Editing Determination

[0106] A schematic diagram of base editing determination is shown in Figure 5 In section A, edited and unedited pEVO plasmids (circles at the top) for base editing were digested with restriction enzymes (REs) specific to the sequences in the target site (boxes) and the sgRNA array (grey). If the target site was edited (boxes with lines), the RE site was lost. The number and size of fragments changed due to the varying number of cuts. Digestion of the unedited plasmid produced two DNA fragments (the two smallest fragments), and digestion of the edited plasmid produced one DNA fragment (the largest fragment). These fragments were visualized using agarose gel electrophoresis (schematic diagram shown at the bottom). A mixture of edited and unedited plasmids is shown in the right lane of the gel schematic. On the right side of the gel, pictograms are used to indicate which editing state a fragment belongs to. Lines with boxes filled with horizontal lines indicate a single edited DNA fragment, and shorter lines with half-boxes indicate two unedited fragments that have been cut. For the assay, the same pEVO plasmid used in the previous examples was used. The Cas12f-ABE variant was cloned using BsrGI and XbaI restriction enzymes. Variant expression is controlled by the L-arabinose-inducible promoter system (araBAD). Base editing at the target site results in the loss of the restriction enzyme site. Therefore, restriction digestion will linearize the edited plasmid, while the unedited plasmid will produce two fragments, which can be detected by gel electrophoresis. Figure 5 A). Example 9: Modification of the Escherichia coli genome

[0107] Three distinct sgRNAs targeting *E. coli* genomic DNA were cloned into pEVO plasmids. Cas12f-ABE WT and three variants were cloned into these vectors using BsrGI and XbaI restriction enzymes. The modified vectors were then transformed into *E. coli* cells. After overnight culture, the cells were centrifuged and resuspended in 200 μl ddH2O. This suspension was then heated to 95°C for 10 minutes and centrifuged again. The supernatant containing gDNA was used for PCR, producing DNA fragments containing all three sgRNA target sites (primers “E_Coli gDNA F” (SEQ ID NO: 56) and “E_Coli gDNA R” (SEQ ID NO: 57)). The PCR products were sequenced using Sanger sequencing, and the editing rate was analyzed using EditR (Kluesner et al., 2018). Example 10: Screening for site-specific recombinases in evolution

[0108] The screening method was performed on an evolved site-specific recombinase library targeting a sequence in the human factor VIII gene (loxF8) (Lansing et al., 2022). The protocol for this method is shown in... Figure 1 Lansing et al., 2022 used 96 randomly selected D7 variants (SEQ ID NO: 19 and 20) identified by recombinases as controls. The aim of the screening was to identify variants with lower off-target activity compared to D7 while maintaining similar on-target activity. To this end, in addition to the loxF8 target site, evolved recombinase libraries were simultaneously screened at three off-target sites recognized by D7, namely HG1, HG2, and HG2L (SEQ ID NO: 12, 13, and 14, respectively) (Lansing et al., 2022). Figure 2 A). Screening yielded a total of 2,515 UMI clusters with 50 or more reads, of which 53 clusters were identified as D7 controls. Analysis of the polished D7 recombinase sequences showed no sequence errors, indicating near 100% accuracy. The median recombination rate of D7 was 80.2% on its expected target, 5.8% on HG1, 57.7% on HG2, and 72.6% on HG2L. Figure 2 B). Of the 2,476 non-D7 clusters, 70 clusters were identified that exhibited less than 10% off-target activity on three off-targets and more than 25% activity on the target. Figure 2 C).

[0109] To validate the results, three variants (clusters 138 (SEQ ID NO: 21 and 22), 181 (SEQ ID NO: 23 and 24), and 1244 (SEQ ID NO: 25 and 26)) were extracted from the screened loxF8 library by PCR amplification using primers specific to the UMI of the variants and a universal forward primer (binding to the plasmid before the sequence encoding the enzyme in region 1) (SEQ ID NO: 51). These variants were selected due to their varying levels of on-target activity while maintaining low off-target activity. The established plasmid-based recombination assay was used. Figure 3 A) The evaluation of variants confirmed that the recombination rate of the tested variants was comparable in the assay. Figure 3 (B, 3C). Therefore, the method of the present invention identifies DNA-modifying enzymes with much lower off-target activity than recombinase D7. Example 11: Screening of Evolved Cas12f-ABE

[0110] The method of the present invention was further evaluated on a CRISPR-based system, a schematic diagram of which is shown in [illustration]. Figure 6 In this study, the directed evolution of substrate associations, as described in Buchholz and Stewart, 2001, was adapted to methods such as Sheriff et al., 2022 (CaSLiDE). Figure 4 The evolution of the ABE system described in [reference needed]. Using this method, 46 directed evolutionary cycles were performed to produce the Un1Cas12f1-ABE8e (Cas12f-ABE) library, which has improved editing efficiency compared to the original Cas12f-ABE. Figure 5 B). The evolved library was then further enriched for four cycles using SLiDE with high-fidelity PCR. Figure 5 C). This library was then screened using the method of the present invention, with the original Un1Cas12f1-ABE8e (wild-type, WT) used as a control. The screening yielded a total of 3,606 UMI clusters with 50 or more reads, of which 123 clusters were identified as WT controls (…). Figure 7 A). The median edit rates of the identified WTs were 0.5% (target site 1), 16% (target site 2), and 6.2% (target site 3). Figure 7 B). Among 3,483 non-WT clusters, 58 achieved an editing rate exceeding 90% across all three target sites. Figure 7 C).

[0111] To validate these results, three variants (cluster 2 (SEQ ID NO: 28), 3030 (SEQ ID NO: 29), and 3301 (SEQ ID NO: 30)) with high activity against all three target sites were extracted from the screened Cas12f-ABE library. Figure 7 A), and evaluated using base editing assays, in Figure 5 Figure A illustrates this schematically. In short, this assay is performed by digesting the CaSLiDE evolution plasmid with a restriction enzyme that recognizes the target site (box). Figure 5 A). Successfully edited plasmid (with a dashed box). Figure 5 A) No restriction sites. Unedited target sites were cleaved by restriction enzymes (REs), while edited target sites were not cleaved due to the alteration of the restriction sites. This resulted in distinct DNA fragments, visualized using agarose gel electrophoresis. The edit-dependent differences in the fragments allowed for the quantification of the amount of edited and unedited fragments. This assay showed that the variant simultaneously exhibited editing rates of 91.1%, 60.8%, and 82.6% at all three target sites, while for WT, it recorded 0% (…). Figure 8 A and 8B). Variant was also tested on the genomic DNA of *E. coli*. Three different sgRNAs were tested with the selected variants or the WT control. The results were edited by Sanger sequencing analysis. Figure 9 (A and 9B). Quantitative base editing rates further confirm that the editing efficiency of the variants identified using the method of this invention is superior to WT. Example 12: Screening for evolved tyrosine site-specific recombinases

[0112] High-throughput sequencing was performed on known tyrosine site-specific recombinases (Panto, Dre, Cre, Vika, and VCRe) and novel tyrosine site-specific recombinases (called YR1, YR2, etc.). All combinations of the 13 tested recombinases and their corresponding 13 target sites were generated on 169 (13x13) independent vectors. Specifically, the 13 target site sequences were cloned individually into expression vectors (pEVO vectors described in Buchholz and Stewart, 2001). The resulting constructs were then pooled and linearized for cloning in a single ligation reaction within the pool of 13 recombinase-coding sequences. After overnight culture and induction of recombinase expression, plasmid DNA was recovered, and fragments carrying the recombinase and target site sequences were excised. Nanopore sequencing yielded a total of 417,769 reads containing the specified recombinase and target sites. All 169 possible combinations of recombinase and target sites were identified, with a minimum coverage of 224 reads. Using this data, the recombination rate of a single recombinase at all target sites is calculated, providing the specificity profile of each recombinase. Figure 11The detailed recombination rates are shown in Table 1 below. Table 1: Recombinase rates at 13 target sites for the 13 tested recombinases

Claims

1. A method for screening and sequencing multiple DNA modifying enzymes, the method comprising the following steps: Provides an expression vector library containing a variety of different vectors, wherein each vector contains a first region encoding a variant of a DNA-modifying enzyme and a second region containing one or more target sites of the DNA-modifying enzyme; The expression vector library was introduced into the host cell; The host cells are cultured and the DNA-modifying enzyme is expressed; Plasmid DNA was isolated from host cell cultures; Sequencing was performed on the first and second regions of the expression vector. as well as Based on the sequencing results, it is determined whether the DNA sequence of the second region on the expression vector has been altered by the DNA-modifying enzyme.

2. A method for screening and sequencing multiple target sites of DNA-modifying enzymes, the method comprising the following steps: Provides an expression vector library containing a variety of different vectors, wherein each vector contains a first region encoding a DNA-modifying enzyme and a second region containing a variant of one or more target sites of the DNA-modifying enzyme; The expression vector library was introduced into the host cell; The host cells are cultured and the DNA-modifying enzyme is expressed; Plasmid DNA was isolated from host cell cultures; Sequencing was performed on the first and second regions of the expression vector. as well as Based on the sequencing results, it is determined whether the DNA sequence of the second region on the expression vector has been altered by the DNA-modifying enzyme.

3. The method of claim 1 or 2, wherein the first region further comprises a unique molecular identifier.

4. The method of claim 3, wherein the unique molecular identifier is an oligonucleotide comprising at least 50 random nucleotides.

5. The method of claim 3 or 4, wherein the unique molecular identifier is located in a first region of the expression vector adjacent to the sequence encoding the DNA-modifying enzyme.

6. The method of any one of claims 3 to 5, further comprising the following steps: Cluster the unique molecular identifiers. Generate and polish the common sequence. Determine the number of DNA modification events for each DNA modifying enzyme; and Determine the activity rate of each DNA-modifying enzyme.

7. The method of any one of claims 1 to 6, wherein the first and second regions of the vector are sequenced in a single step.

8. The method of any one of claims 1 to 7, wherein sequencing of the first and second regions of the expression vector comprises nanopore sequencing.

9. The method of any one of claims 1 to 8, wherein the DNA modifying enzyme is selected from recombinases, integrases, adenosine base editors (ABEs), zinc finger nucleases, transcription activator-like effector nucleases, and Cas nucleases.

10. The method of any one of claims 1 to 9, wherein the DNA modifying enzyme comprises more than one subunit.

11. The method of claim 10, wherein the DNA modifying enzyme comprises at least two distinct subunits.

12. The method of any one of claims 1 to 11, wherein, prior to the sequencing, a first region and a second region of the expression vector are excised from the expression vector.

Citation Information

Patent Citations

  • Protein with recombinase activity for site-specific DNA-recombination

    EP2690177B1

  • Protein with recombinase activity for site-specific DNA-recombination

    EP2877585B1

  • Protein with recombinase activity for site-specific DNA-recombination

    EP3263708B1

  • Dre recombinase and recombinase systems employing Dre recombinase

    US7422889B2

  • Dre recombinase and recombinase systems employing Dre recombinase

    US7915037B2