A directed evolution method based on primary and secondary replicons of geminiviruses

By using geminivirus replicon vectors in plant cells for directed evolution of genetic elements, the problem of low efficiency in directed evolution systems in plant cells has been solved. This enables efficient screening of mutants of genetic elements with desired functions and provides high-throughput genetic element screening capabilities.

CN115516089BActive Publication Date: 2025-12-19SUZHOU QI BIODESIGN BIOTECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180033728.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-07
Filing Date
2021-03-30
Publication Date
2025-12-19
Estimated Expiration
2041-03-30

AI Technical Summary

Technical Problem

Existing directed evolution systems are difficult to work efficiently in plant cells, mainly due to the unique anatomical structure and differences in cellular regulatory networks of plant cells, which prevent existing tools from effectively achieving efficient and high-throughput directed evolution within plants.

Method used

Directed evolution of genetic elements is achieved in plant cells using geminivirus replicons. By inserting mutants of genetic elements into the vector of geminivirus replicons, the activity of geminivirus Rep/RepA proteins is coupled with the function of genetic elements, thereby achieving efficient amplification and enrichment of genetic elements.

Benefits of technology

It enables efficient and rapid directed evolution in plant cells, allowing for the screening of mutants of genetic elements with desired functions. It overcomes the barriers of the unique structure and regulatory network of plant cells, providing high-throughput genetic element screening capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003930670210000181
    Figure BDA0003930670210000181
  • Figure BDA0003930670210000182
    Figure BDA0003930670210000182
  • Figure BDA0003930670210000183
    Figure BDA0003930670210000183
Patent Text Reader

Abstract

A directed evolution method based on geminiviruses is provided, which is a directed evolution method using in vivo screening of genetic elements by primary and secondary replicons of geminiviruses in plant cells.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of genetic engineering. Specifically, the present application relates to a directed evolution method based on geminivirus. More specifically, the present application relates to a directed evolution method using primary and secondary replicons of geminivirus to screen genetic elements in vivo in plant cells. BACKGROUND

[0003] In the long history of life, organisms constantly produce variations, some of which are beneficial to the survival of organisms, and some of which are harmful to the survival of organisms. Under the pressure of natural selection, the variations beneficial to the survival of organisms are retained and enriched, and the unfavorable ones are eliminated, which is the process of evolution.

[0004] According to this principle, in the field of modern molecular biology, researchers simulate this process in the laboratory, artificially create a large number of mutations, and give selection pressure according to the required function and purpose, and screen out genotypes with desired characteristics. This simulation of evolution at the molecular level is called directed evolution. Directed evolution can be used to modify proteins without knowing the structure and mechanism of the target protein. Therefore, directed evolution is an effective method for obtaining new functional proteins in current molecular biology research.

[0005] So far, there have been many directed evolution systems reported. The main ones are the MAGE (multiplex automated genome engineering) system developed by George M. Church et al. in 2009, the PACE (phage-assisted continuous evolution) system developed by David. Liu et al. in 2011, the EvolvR system developed by John E. Dueber et al. in 2018, etc. These systems can work efficiently and have completed the evolution of many genetic tools such as xCas9. However, there is currently no efficient and high-throughput directed evolution system based on higher plants.

[0006] Several factors make it urgent to develop a directed evolution system working in plant cells. First, different working temperature and pH are needed in plants. Most of the tools currently used in plants are derived from E. coli or mammals, and the living temperature of E. coli or mammals is generally 37°C, and the optimum working temperature of the proteins derived from them is also generally 37°C, but the optimum growth temperature of most plants is 20-25°C; similarly, the pH of animal cells is generally 7.2-7.4, the pH of E. coli is 7.0-7.5, and the pH of the cytoplasmic matrix of plant cells is 5.6-5.9. These make the proteins derived from E. coli and mammals or the products obtained by the directed evolution system relying on E. coli and mammals not work as efficiently as before in plants due to thermodynamic and chemical factors. Second, plant cells have special anatomical structures. Compared with E. coli, as eukaryotes, plant cells have structures such as nucleus and endoplasmic reticulum; compared with animal cells, plant cells have structures such as chloroplasts and cell walls, and it is difficult to evolve some proteins related to these structures in E. coli or mammalian systems. Third, plant cells have a unique cellular regulatory network. During the long evolution process, a complex regulatory network has been formed among various genetic elements inside the cell, and there are great differences between prokaryotes and eukaryotes, and between plant cells and animal cells. Since any element is unlikely to be independent of the network, the products obtained by the directed evolution system outside the plant cell system may not work effectively in the plant system. And for the elements in the plant cell regulatory network, it is impossible to achieve directed evolution in other biological systems. Due to the above reasons, some elements that can work efficiently in mammalian or E. coli systems, such as eAPOBEC3A reported by J. Keith Joung et al., do not show activity in plant cells. Therefore, there is a need in the art to develop a directed evolution system based on plant cells.

[0007] However, at present, although researchers have made attempts to perform directed evolution in plants in several works, these systems are difficult to apply and popularize. The reason is that some inherent factors of plants hinder the development of related research. First, many existing directed evolution systems rely on high-frequency homologous recombination in vivo, but the homologous recombination efficiency in plants is less than 0.1%; second, bacteria, yeast, and animal cells can be efficiently prepared into cell lines, but at present only a few plant species can be prepared into cell lines, and the repeatability is also low. Third, the lack of high-efficiency transformation system in plants makes high-throughput research in plants require a great amount of work. These factors greatly hinder the development of directed evolution systems in plants, and it is difficult to have substantial changes in a short time.

[0008] Geminiviruses are the largest group of plant single-stranded DNA viruses, with virions in doublet structure, monomer or dimmer, and each DNA molecule size of 2.5-3.0 kb, total genome size of 2.5-5.2 kb. After invading plant cells, the viruses of this family will first form a double-stranded DNA replication intermediate in the nucleus, and then under the action of the virus-encoded replication initiation protein Rep / RepA and the endogenous DNA polymerase in the cell, the viral genome is amplified in the form of rolling circle replication (such as Figure 1 ).

[0009] So far, more than 500 species of Geminiviridae have been found; according to the genomic structure, Geminiviridae is divided into 9 genera. Among them, only the genus of maize streak virus (Mastrevirus) can infect monocotyledonous plants, and its structure is the simplest and its research is the most thorough. The members of the genus of maize streak virus are monomer viruses, with a genome size of 2500-2800 bp, encoding 4 proteins, respectively, moving protein (MP), capsid protein (CP), replication initiation protein Rep and RepA (such as Figure 2 ). Among them, Rep and RepA are encoded by a common piece of DNA, and two transcripts are obtained by alternative splicing. Rep / RepA is related to the initiation and termination of viral rolling circle replication, inhibition of plant immunity, and regulation of viral gene expression. MP and CP are related to the movement and packaging of virions, respectively, and are not related to replication. In addition, the genome of the members of the genus of maize streak virus also contains two intergenic regions: LIR (large intergenic region) and SIR (small intergenic region). The former is a bidirectional promoter and contains a stable stem-loop structure, which can be recognized by Rep / RepA and is the replication origin of rolling circle replication; the latter is a bidirectional terminator and is related to the formation of double-stranded DNA intermediates during replication. Since LIR and SIR are the only two cis-acting elements required for viral replication, and Rep / RepA is the only trans-acting factor required, researchers have developed a disassembled viral replicon (such as Figure 3 ), which only contains LIR and SIR, and the rest can be any sequence, and Rep / RepA can be expressed in situ or ex situ, driving the rolling circle replication of the replicon.

[0010] SUMMARY

[0011] The present application provides a method for directed evolution of genetic elements to obtain mutants of the genetic elements with desired functions, the method comprising:

[0012] i) providing a library of mutants of said genetic element comprising a plurality of mutants of said genetic element inserted individually into a vector comprising a geminivirus replicon, wherein said mutants are inserted into the geminivirus replicon, whereby said mutants are amplified when the geminivirus replicon replicates,

[0013] ii) transforming a population of plant cells with said library, and

[0014] iii) cultivating said population of plant cells, detecting and selecting genetic element mutants that are enriched in said population of plant cells,

[0015] wherein the level of replication of said geminivirus replicon in said plant cells is set in relation to the desired function of said genetic element mutants.

[0016] In some embodiments, said genetic element is selected from the group consisting of a protein coding sequence; a functional RNA coding sequence, such as a tRNA, siRNA coding sequence; an expression regulatory sequence such as a promoter sequence, enhancer sequence, terminator sequence.

[0017] In some embodiments, said genetic element is derived from a plant, or is desired for use in a plant.

[0018] In some embodiments, said library of mutants of said genetic element is obtained by inserting a plurality of mutants of said genetic element individually into a vector comprising a geminivirus replicon.

[0019] In some embodiments, said plurality of mutants of said genetic element is generated by random mutagenesis of said genetic element.

[0020] In some embodiments, said library is generated by random mutagenesis of said genetic element that has been inserted into a vector comprising a geminivirus replicon.

[0021] In some embodiments, said vector comprising a geminivirus replicon is a circular DNA, such as a plasmid or a minicircle DNA.

[0022] In some embodiments, said vector comprising a geminivirus replicon comprises at least one LIR, for example, said LIR comprises the nucleotide sequence set forth in SEQ ID NO: 1.

[0023] In some embodiments, said vector comprising a geminivirus replicon further comprises at least one SIR, for example, said SIR comprises the nucleotide sequence set forth in SEQ ID NO: 2.

[0024] In some embodiments, said vector comprising a geminivirus replicon comprises one LIR.

[0025] In some embodiments, the vector comprising a geminivirus replicon comprises two LIRs.

[0026] In some embodiments, in the vector comprising a geminivirus replicon, the inserted mutant of the genetic element is operably linked to an expression control sequence.

[0027] In some embodiments, the vector comprising a geminivirus replicon further comprises an expression cassette for a geminivirus Rep and / or RepA protein.

[0028] In some embodiments, the vector comprising a geminivirus replicon does not comprise an expression cassette for a geminivirus Rep and / or RepA protein.

[0029] In some embodiments, the method further comprises introducing into the plant cells a further vector for expression of a geminivirus Rep and / or RepA protein.

[0030] In some embodiments, the further vector for expression of a geminivirus Rep and / or RepA protein is co-transformed with the library into the population of plant cells.

[0031] In some embodiments, the plant cells already comprise a vector for expression of a geminivirus Rep and / or RepA protein, and / or the genome of the plant cells has integrated an expression cassette for a geminivirus Rep and / or RepA protein.

[0032] In some embodiments, the geminivirus Rep protein comprises the amino acid sequence of SEQ ID NO: 3, or an amino acid sequence comprising the amino acid substitution K229E or Y20C relative to SEQ ID NO: 3, preferably the amino acid sequence of SEQ ID NO: 4.

[0033] In some embodiments, the geminivirus RepA protein comprises the amino acid sequence of SEQ ID NO: 5, or an amino acid sequence comprising the amino acid substitution K229E or Y20C relative to SEQ ID NO: 5, preferably the amino acid sequence of SEQ ID NO: 6.

[0034] In some embodiments, wherein in the transformation of step ii), the number of vector molecules comprising a mutant in the library is 10 3 to 10 5 times the number of cells in the population of plant cells.

[0035] In some embodiments, wherein in step iii), detecting and selecting the genetic element mutants that are enriched in the population of plant cells can be performed by high-throughput sequencing.

[0036] In some embodiments, further comprising step iv) identifying the function of said enriched genetic element mutant.

[0037] In some embodiments, the plant is a monocot or a dicot, for example selected from the group consisting of maize, wheat, rice, barley, sorghum, bean, sugar beet, tomato, cassava, cucumber, Arabidopsis thaliana and tobacco.

[0038] In some embodiments, the directed evolution of a genetic element is accomplished by coupling the expression or activity of a geminivirus Rep and / or RepA protein in said plant cell to the desired function of said genetic element mutant.

[0039] In some embodiments, the directed evolution of a genetic element is accomplished by having a functional genetic element activate Rep / RepA expression, thereby driving rolling circle replication, enabling self-enrichment; a non-functional genetic element is unable to activate Rep / RepA expression, thereby failing to enrich.

[0040] In some embodiments, the genetic element is a promoter.

[0041] In some embodiments, the method i) further comprises placing a library of promoters to be evolved upstream of Rep / RepA in the replicon.

[0042] In some embodiments, the genetic element is the Cauliflower Mosaic Virus (CaMV) 35S promoter TATA-box.

[0043] In some embodiments, the genetic element is a sequence encoding a transcriptional activator.

[0044] In some embodiments, the method i) further comprises inserting a recognition sequence of a transcriptional activator upstream of Rep / RepA and a minimal transcription initiation element between said recognition sequence and Rep / RepA; and placing a library of transcriptional activators to be evolved in the replicon.

[0045] In some embodiments, the genetic element is a DNA binding domain.

[0046] In some embodiments, the method i) further comprises inserting a target binding sequence of said DNA binding domain upstream of Rep / RepA and a minimal transcription initiation element between said recognition sequence and Rep / RepA; and placing a DNA binding domain to be evolved in the replicon as a fusion protein with a transcriptional activator that is sequence nonspecific.

[0047] In some embodiments, the genetic element is a sequence encoding a recombinase.

[0048] In some embodiments, the method i) further comprises separating Rep / RepA into two parts, placing them on the two wings of the recombinase recognition sequence; and placing the sequence encoding the recombinase to be evolved in the replicon.

[0049] In some embodiments, the method i) further comprises adding a 5’ intron and a 3’ intron between Rep / RepA and the recombinase recognition sequence.

[0050] In some embodiments, the genetic element is a prime editing guide RNA (pegRNA).

[0051] In some embodiments, the method i) further comprises inserting a target site at the N-terminus of Rep / RepA, and causing a frameshift of the open reading frame of Rep / RepA; and inserting an expression cassette of pegRNA into the geminivirus replicon, and inserting a fluorescent reporter system on its two wings.

[0052] In some embodiments, the desired function of the genetic element is coupled to the expression of a nuclease.

[0053] In some embodiments, the nuclease is a sequence-specific nuclease.

[0054] In some embodiments, the directed evolution of the genetic element is achieved by: a functional genetic element drives self-enrichment by activating the expression of a nuclease or directing the nuclease to cleave its recognition site, thus driving rolling circle replication; a non-functional genetic element is unable to make the nuclease cleave its recognition site, thus failing to enrich.

[0055] In some embodiments, the genetic element is a DNA-binding domain.

[0056] In some embodiments, the method i) further comprises fusing a library of DNA-binding domains to be evolved with a non-sequence-specific nuclease, and placing them together with their recognition sequences in the replicon.

[0057] In some embodiments, the genetic element is a sequence encoding a non-sequence-specific nuclease.

[0058] In some embodiments, the genetic element is a sequence encoding a transcriptional activator.

[0059] In some embodiments, the method i) further comprises inserting a recognition sequence of a transcriptional activator upstream of the nuclease, inserting a minimal transcription initiation element between the recognition sequence and the nuclease; and co-locating a library of transcriptional activators to be evolved with the recognition sequence of the nuclease in a replicon.

[0060] In some embodiments, the genetic element is a sequence encoding a recombinase.

[0061] In some embodiments, the method i) further comprises co-locating a library of recombinases to be evolved with the recognition sequence of the nuclease in a replicon; and separating the nuclease into two parts flanking the recombinase recognition sequence.

[0062] In some embodiments, the method i) further comprises adding a 5' intron and a 3' intron between the nuclease and the recombinase recognition sequence.

[0063] In some embodiments, the genetic element is a PAM (protospacer adjacent motif) of a Cas protein.

[0064] In some embodiments, the method further comprises co-locating a PAM to be evolved with a target sequence of the Cas protein in a replicon.

[0065] In some embodiments, the genetic element is an sgRNA.

[0066] In some embodiments, the method further comprises co-locating an sgRNA to be evolved with a target sequence of the Cas protein in a replicon.

[0067] The present application also provides a kit for carrying out the method of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0069] Figure 1 A rolling circle replication model of geminiviruses is shown.

[0070] Figure 2 The genomic structure of a Mastrevirus virus is shown.

[0071] Figure 3 A deconstructed viral replicon strategy for geminiviruses is shown.

[0072] Figure 4The basic principle of the in-plant directed evolution system based on the primary replicon is shown: different alleles (mutants) of the gene of interest (GOI) are placed in the geminivirus replicon to form a library to be screened; the desired function of the GOI is coupled to the expression of Rep / RepA, i.e. the functional alleles with the desired function are able to cause Rep / RepA expression in plant cells, while the non-functional alleles without the desired function are not able to cause Rep / RepA expression in plant cells.

[0073] Figure 5 The idea of implementing promoter directed evolution using the primary replicon directed evolution system is shown.

[0074] Figure 6 The idea of implementing transcriptional activator directed evolution using the primary replicon directed evolution system is shown.

[0075] Figure 7 The idea of implementing DNA binding domain directed evolution using the primary replicon directed evolution system is shown.

[0076] Figure 8 The idea of implementing recombinase directed evolution using the primary replicon directed evolution system is shown.

[0077] Figure 9 The principle of forming the secondary replicon is shown.

[0078] Figure 10 The basic principle of the in-plant directed evolution system based on the secondary replicon is shown.

[0079] Figure 11 The idea of implementing DNA binding domain directed evolution using the secondary replicon directed evolution system is shown.

[0080] Figure 12 The idea of implementing non-sequence specific nuclease directed evolution using the secondary replicon directed evolution system is shown.

[0081] Figure 13 The idea of implementing transcriptional activator directed evolution using the secondary replicon directed evolution system is shown.

[0082] Figure 14 The idea of implementing recombinase directed evolution using the secondary replicon directed evolution system is shown.

[0083] Figure 15 The construction of the screening library in Example 1 is shown.

[0084] Figure 16 The Rep Y20 screening results without the addition of the replication enhancer in Example 1 are shown.

[0085] Figure 17 Rep Y20 screening results with the addition of replication enhancer in Example 1 are shown.

[0086] Figure 18 Construction of the screening library in Example 2 is shown.

[0087] Figure 19 Screening results in Example 2 are shown.

[0088] Figure 20 Principle of the experiment and vector construction in Example 3 are shown.

[0089] Figure 21 Screening results in Example 3 are shown.

[0090] Figure 22 Construction of the screening library in Example 4 is shown.

[0091] Figure 23 Principle of the screening of PAM library in Example 4 is shown.

[0092] Figure 24 Screening results of 3 bases at 3' end of PAM library sequence in Example 4 are shown.

[0093] Figure 25 Sequence diagram of 6 bases in PAM library sequence in Example 4 is shown.

[0094] Figure 26 Principle of the experiment and vector construction in Example 5 are shown.

[0095] Figure 27 Screening results in Example 5 are shown.

[0096] Figure 28 Principle of directed evolution of base editor is shown in the schematic diagram, in which the base editor mutant expression library, the Rep / RepA expression vector in inactivated form, the sgRNA expression construct against the Rep / RepA coding sequence in inactivated form are co-transformed into plant cells. When the base editor mutant has the desired base editing activity, the Rep / RepA in inactivated form can be corrected to the Rep / RepA in active form, thereby obtaining enrichment.

[0097] Figure 29 Principle of directed evolution of recombinase is shown in the schematic diagram, in which the recombinase mutant expression library and the Rep / RepA gene and promoter in reverse expression vector are co-transformed into plant cells. When the recombinase mutant has the desired activity, the Rep / RepA gene in reverse can be inverted, thereby being driven to express by the promoter, and the enrichment of the recombinase mutant is achieved. DETAILED DESCRIPTION

[0099] In one aspect, the present application provides a method of directed evolution of a genetic element to obtain a mutant of said genetic element having a desired function, said method comprising:

[0100] i) providing a library of mutants of said genetic element comprising a plurality of mutants of said genetic element inserted individually into a vector comprising a geminivirus replicon, wherein said mutants are inserted into the geminivirus replicon, whereby said mutants are amplified when the geminivirus replicon replicates,

[0101] ii) transforming a population of plant cells with said library, and

[0102] iii) cultivating said population of plant cells, detecting and selecting genetic element mutants that are enriched in said population of plant cells,

[0103] wherein the level of replication of said geminivirus replicon in said plant cells is set to be associated with the desired function of said genetic element mutants.

[0104] As used herein, the term "genetic element" refers to a nucleotide sequence / nucleic acid molecule that is capable of performing a specific function within a cell, preferably within a plant cell. Examples of genetic elements include, but are not limited to, protein coding sequences, functional RNA (e.g. tRNA, siRNA, etc.) coding sequences, expression regulatory sequences such as promoter sequences, enhancer sequences, terminator sequences, etc. In some preferred embodiments, the genetic element is derived from a plant, or is desired to be applied to a plant.

[0105] In the context of the present specification, the term "library" is used in its meaning known in the art of cell biology and molecular biology, and it refers to a collection of different nucleic acid fragments / nucleic acid molecules. One particular type of library is a library comprising random mutants generated by random mutagenesis. Another example is a designed (or synthetic) library comprising different nucleic acid fragments / nucleic acid molecules that are specifically engineered.

[0106] In some embodiments, the library of mutants of a genetic element is obtained by inserting a plurality of mutants of said genetic element individually into a vector comprising a geminivirus replicon. In some embodiments, the plurality of mutants of said genetic element is generated by random mutagenesis.

[0107] In some embodiments, the library can be generated by random mutagenesis of said genetic element inserted into a vector comprising a geminivirus replicon.

[0108] In the context of the present specification, the term "random mutagenesis" is used in its known meaning in the field of cell biology and molecular biology; it refers to a process in which DNA is mutated randomly to produce mutant genes and proteins. A number of these mutant genes can then be compiled into a library. Non-limiting examples of random mutagenesis methods are error-prone PCR, UV radiation, and chemical mutagens.

[0109] A "geminivirus" is a DNA virus that infects plants, which is a virus having one or two single-stranded circular DNA molecules. Exemplary geminiviruses include, but are not limited to, viruses belonging to the genus Maize streak virus, such as MSV (maize streak virus), WDV (Wheat dwarf virus), BeYDV (Bean yellow dwarf virus), and the like, viruses belonging to the genus Beet curly top virus, such as BCTV (Beet curly top virus), and the like, viruses belonging to the genus Tomato pseudo-curl virus, such as TPCTV (Tomato pseudo-curl virus), and the like, and viruses such as BGMV (Bean golden mosaic virus), ACMV (African cassava mosaic virus), SLCV (Squash leaf curl virus), TGMV (Tomato golden mosaic virus), and TYLCV (Tomato Yellow Leaf Curl Virus), and the like. In some preferred embodiments, the geminivirus is WDV.

[0110] The plant of the present application can be a monocotyledonous or dicotyledonous plant, as long as the geminivirus replicon is able to replicate in its cells. Suitable plants include, but are not limited to, maize, wheat, rice, barley, sorghum, bean, beet, tomato, cassava, cucumber, Arabidopsis, tobacco, and the like.

[0111] In some embodiments, the plant cell is an isolated plant cell. In some preferred embodiments, the plant cell is a protoplast cell.

[0112] In some embodiments, the plant cell is a cell in a plant tissue or plant organ or plant body, i.e., the cell is not isolated from a plant tissue or plant organ or plant body. For example, the plant cell can be a leaf cell.

[0113] The "replicative level" of a geminivirus replicon can be determined by detecting the copy number of the geminivirus replicon. Methods of detecting the copy number of a geminivirus replicon are known in the art, including but not limited to PCR (e.g. quantitative PCR) methods or sequencing (e.g. deep sequencing) methods.

[0114] In some embodiments, the vector comprising a geminivirus replicon is a circular DNA, e.g. a double-stranded or single-stranded circular DNA. In some embodiments, the vector comprising a geminivirus replicon is a plasmid. In some embodiments, the vector comprising a geminivirus replicon is a minicircle DNA.

[0115] In some embodiments, the vector comprising a geminivirus replicon comprises at least one LIR (large intergenic region).

[0116] In some embodiments, the vector comprising a geminivirus replicon further comprises at least one, e.g. one, SIR (small intergenic region).

[0117] In some embodiments, the vector comprising a geminivirus replicon comprises one LIR. In this case, the entire vector is replicated as a geminivirus replicon.

[0118] In some preferred embodiments, the vector comprising a geminivirus replicon comprises two LIRs. In some embodiments, a SIR is comprised between the two LIRs. In this case, the first LIR is replicated as a geminivirus replicon up to the sequence of the second LIR (comprising the SIR). Preferably, the genetic element variant is located between the two LIRs.

[0119] In some embodiments, the LIR comprises the nucleotide sequence set forth in SEQ ID NO: 1. In some embodiments, the SIR comprises the nucleotide sequence set forth in SEQ ID NO: 2.

[0120] In some embodiments, in the vector comprising a geminivirus replicon, the inserted mutant of the genetic element is operably linked to an expression control sequence.

[0121] "Expression control sequence" and "expression control element" are used interchangeably and refer to nucleotide sequences located upstream (5' non-coding sequences), within, or downstream (3' non-coding sequences) of a coding sequence, and which influence the transcription, RNA processing or stability, or translation of the associated coding sequence. A plant expression control element is a nucleotide sequence capable of controlling the transcription, RNA processing or stability, or translation of a nucleotide sequence of interest in a plant. Expression control sequences can include, but are not limited to, promoters, translation leader sequences, introns, and polyadenylation recognition sequences. A "promoter" refers to a nucleic acid segment that functions to control transcription of another nucleic acid segment. In some embodiments of the application, a promoter is a promoter that is capable of controlling gene transcription in a plant cell, whether or not it is derived from a plant cell. A promoter can be a constitutive promoter or a tissue-specific promoter or a developmentally-regulated promoter or an inducible promoter.

[0122] In some embodiments, the vector comprising a geminivirus replicon further comprises an expression cassette for a geminivirus Rep and / or RepA protein.

[0123] An expression cassette for a geminivirus Rep and / or RepA protein typically comprises a nucleotide sequence encoding a geminivirus Rep and / or RepA protein and expression control elements operably linked thereto.

[0124] In some embodiments, the vector comprising a geminivirus replicon does not comprise an expression cassette for a geminivirus Rep and / or RepA protein. Thus, the geminivirus Rep and / or RepA protein needs to be provided in trans.

[0125] In some embodiments, the method further comprises introducing an additional vector for expressing a geminivirus Rep and / or RepA protein into the plant cell. A vector for expressing a geminivirus Rep and / or RepA protein typically comprises an expression cassette for a geminivirus Rep and / or RepA protein. In some embodiments, the additional vector for expressing a geminivirus Rep and / or RepA protein is co-transformed into the population of plant cells with the library.

[0126] In some embodiments, the plant cell already comprises a vector for expressing a geminivirus Rep and / or RepA protein, and / or the plant cell genome has integrated an expression cassette for a geminivirus Rep and / or RepA protein.

[0127] In some embodiments, the geminivirus Rep protein comprises the amino acid sequence set forth in SEQ ID NO: 3, or an amino acid sequence comprising the amino acid substitution K229E or Y20C relative to SEQ ID NO: 3, for example SEQ ID NO: 4. In some embodiments, the geminivirus RepA protein comprises the amino acid sequence set forth in SEQ ID NO: 5 or an amino acid sequence comprising the amino acid substitution K229E or Y20C relative to SEQ ID NO: 5, for example SEQ ID NO: 6. In some preferred embodiments, the geminivirus Rep protein comprises the amino acid sequence set forth in SEQ ID NO: 4. In some preferred embodiments, the geminivirus RepA protein comprises the amino acid sequence set forth in SEQ ID NO: 6.

[0128] By "the level of replication of the geminivirus replicon in the plant cell is set in relation to the desired function of the genetic element mutant" is meant that the level of replication of the geminivirus replicon in the plant cell is such that the genetic element mutant having the desired function is able to cause replication of the geminivirus replicon in the plant cell, preferably at a higher, more preferably at a significantly higher level, than the level of replication of the geminivirus replicon in the plant cell by the genetic element mutant not having the desired function. For example, the genetic element mutant having the desired function can be able to cause replication of the geminivirus replicon in the plant cell, while the genetic element variant not having the desired function causes no replication of the geminivirus replicon in the plant cell; or preferably the genetic element mutant having the desired function is able to cause high level replication of the geminivirus replicon in the plant cell, while the genetic element variant not having the desired function causes low level or no replication of the geminivirus replicon in the plant cell. The geminivirus replicon comprising the genetic element mutant having the desired function is amplified or significantly amplified due to the replication or high level replication, when compared to other geminivirus replicons not comprising the genetic element mutant having the desired function, which allows for the enrichment of the genetic element mutant having the desired function.

[0129] Rep and / or RepA proteins are replication initiation proteins of geminiviruses, the activity or expression level of which is usually positively correlated with the replication level (e.g. copy number) of the geminivirus replicon within a certain range. Thus, in some embodiments, the "activity or expression level of a geminivirus Rep and / or RepA protein in said plant cell can be set to be associated with the desired function of said genetic element mutant". For example, the activity or expression level of a geminivirus Rep and / or RepA protein in a plant cell comprising a genetic element mutant having the desired function can be higher, preferably significantly higher, than the activity or expression level of a geminivirus Rep and / or RepA protein in a plant cell comprising a genetic element mutant not having the desired function. In some embodiments, the activity of the geminivirus Rep and / or RepA protein is the activity of mediating (initiating) replication of the geminivirus replicon, which can be determined, for example, by detecting the replication level of the geminivirus replicon. For example, a genetic element mutant having the desired function can be enabled to cause expression of a geminivirus Rep and / or RepA protein in a plant cell, while a genetic element mutant not having the desired function causes no expression of a geminivirus Rep and / or RepA protein in a plant cell; or a genetic element mutant having the desired function can be enabled to cause high level expression of a geminivirus Rep and / or RepA protein in a plant cell, while a genetic element mutant not having the desired function causes low level expression or no expression of a geminivirus Rep and / or RepA protein in a plant cell. Alternatively, a genetic element mutant having the desired function can be enabled to cause activity of a geminivirus Rep and / or RepA protein in a plant cell, while a genetic element mutant not having the desired function causes no activity of a geminivirus Rep and / or RepA protein in a plant cell; or a genetic element mutant having the desired function can be enabled to cause high activity of a geminivirus Rep and / or RepA protein in a plant cell, while a genetic element mutant not having the desired function causes low activity or no activity of a geminivirus Rep and / or RepA protein in a plant cell. The expression or activity or high level expression or high activity of a geminivirus Rep and / or RepA protein in said plant cell will result in amplification or significant amplification of said geminivirus replicon, thereby achieving enrichment of the genetic element mutant having the desired function.

[0130] The "low level" or "low activity" described herein is relative to "high level" or "high activity", and does not necessarily mean that it is lower than the normal level or normal activity.

[0131] The replication level of a geminivirus in a plant cell or the activity or expression level of a Rep and / or RepA protein of a geminivirus in a plant cell can be set in correlation with the desired function of the genetic element mutant, either directly or indirectly. The person skilled in the art is able to achieve such a correlation depending on the type of genetic element and the desired specific function of its mutant.

[0132] In the context of the present specification, the term "expression level" is used in its meaning known in the art of cell biology and molecular biology; it refers to the transcriptional and / or translational level of a DNA fragment and its derived mRNA, respectively.

[0133] For example, when the genetic element is an expression regulatory element (such as a promoter, an enhancer, etc.), the coding sequence of a geminivirus Rep and / or RepA protein can be placed directly under the control of the expression regulatory element mutant (such as a promoter mutant, an enhancer mutant, etc.). If the mutant is able to enhance gene expression, it will lead to an increased expression of Rep and / or RepA protein, which in turn leads to an increased replication of the geminivirus replicon and the corresponding expression regulatory element mutant (such as a promoter mutant). By detecting significantly enriched mutant sequences, expression regulatory element mutants that enhance gene expression, i.e. evolved expression regulatory elements, can be obtained.

[0134] When the genetic element is a protein coding sequence, the expected function of the protein encoded thereby can be correlated with the activity or expression level of a geminivirus Rep and / or RepA protein.

[0135] For example, if directed evolution of a base editor is desired, a library of base editor mutants constructed in a vector comprising a geminivirus replicon, an expression vector comprising a coding sequence for a Rep / RepA protein in an inactivated form specifically designed for the desired functional feature of the base editor, an sgRNA expression construct for the Rep / RepA coding sequence in the inactivated form can be co-transformed into plant cells. When a base editor mutant in the plant cell has the desired base editing activity, the specifically designed inactivated form of Rep / RepA can be corrected to an active form of Rep / RepA, which in turn induces replication of the geminivirus replicon, so that the mutant is enriched.

[0136] Alternatively, if directed evolution of the recombinase is desired, a library of the recombinase mutant constructed in a vector containing a geminivirus replicon, along with an expression vector in which the Rep / RepA coding sequence is reversed with the promoter, can be co-transformed into plant cells. When the recombinase mutant in the plant cells exhibits the desired activity, it can invert the reversed Rep / RepA gene, thereby enabling it to be expressed by the promoter and inducing the replication of the geminivirus replicon, thus enriching the recombinase mutant.

[0137] In some embodiments, during the protoplast transformation in step ii), the number of vector molecules containing the mutant in the library is 10 times the number of cells in the population of plant cells. 3 Up to 10 5 This ratio can reduce the probability of multiple different vector molecules transforming into the same cell while ensuring transformation efficiency, thus reducing the background of screening.

[0138] In the context of this specification, the term "primary replicon" refers to a replicon formed by a tandem LIR on a vector in a plant geminivirus system that can be recognized and circularized by Rep / RepA. Primary replicons are amplified via rolling circle replication.

[0139] In some embodiments, directed evolution of genetic elements is accomplished by coupling the expression or activity of the geminivirus Rep and / or RepA proteins in the plant cells with the desired function of the genetic element mutant. In some embodiments, directed evolution of genetic elements is accomplished by: functional genetic elements activating Rep / RepA expression, thereby driving rolling circle replication and achieving self-enrichment; and non-functional genetic elements failing to activate Rep / RepA expression, thus failing to achieve enrichment.

[0140] In some embodiments, the genetic element is a promoter. In some embodiments, method i) further includes placing the promoter library to be evolved upstream of Rep / RepA in the replicon. A functional promoter can drive the expression of downstream Rep / RepA, drive its own rolling circle replication, increase the copy number, and achieve self-enrichment; a non-functional promoter cannot drive the expression of downstream Rep / RepA and cannot achieve enrichment, thereby achieving directed evolution of the promoter.

[0141] In some implementations, the genetic element is the CaMV 35S promoter TATA-box.

[0142] In the context of this specification, the term "transcription activator" is a DNA-binding protein capable of activating gene expression. Transcription activators regulate the transcription process by binding to upstream promoter elements.

[0143] In some embodiments, the genetic element is a sequence encoding a transcriptional activator. In some embodiments, the method i) further comprises inserting a recognition sequence of the transcriptional activator upstream of Rep / RepA and a minimal transcription initiation element between the recognition sequence and Rep / RepA; and placing a library of transcriptional activators to be evolved in the replicon.

[0144] In some embodiments, the genetic element is a DNA binding domain. In some embodiments, the method i) further comprises inserting a target binding sequence of the DNA binding domain upstream of Rep / RepA and a minimal transcription initiation element between the recognition sequence and Rep / RepA; and placing a DNA binding domain to be evolved in a fusion protein with a sequence-unspecific transcriptional activator in the replicon. A functional transcriptional activator can bind to its recognition sequence and activate the expression of Rep / RepA downstream, thus driving rolling circle replication and self-enrichment; a non-functional transcriptional activator cannot activate the expression of Rep / RepA downstream and thus cannot be enriched, thus achieving directed evolution of the transcriptional activator.

[0145] In the context of the present specification, the term "recombinase" refers to an enzyme involved in the process of genetic location recombination. It is responsible for recognizing and cutting specific recombination sites, and connecting two molecules involved in recombination. In some embodiments, the genetic element is a sequence encoding a recombinase. In some embodiments, the method i) further comprises separating Rep / RepA into two parts, flanking the recombinase recognition sequence; and placing a sequence encoding a recombinase to be evolved in the replicon. In some embodiments, the method i) further comprises adding a 5' intron and a 3' intron between Rep / RepA and the recombinase recognition sequence. A functional recombinase can recognize its specific recognition site, mediate DNA recombination, and Rep / RepA can be normally expressed, thus driving rolling circle replication and self-enrichment; a non-functional recombinase cannot mediate DNA recombination and thus cannot express Rep / RepA, and thus cannot be enriched, thus achieving directed evolution of the recombinase

[0146] In some embodiments, the genetic element is a prime editing guide pegRNA. In some embodiments, the method i) further comprises inserting a target site at the N-terminus of Rep / RepA and frameshifting the open reading frame of Rep / RepA; and inserting an expression cassette of pegRNA into the geminivirus replicon and inserting a fluorescent reporter system at both flanks. If the viral replicon undergoes rolling circle replication under the action of pegRNA, a fluorescent signal is reported; if the viral replicon does not undergo rolling circle replication, there is no fluorescent signal. In some embodiments, when tobacco leaves are transformed with a low concentration library, active pegRNAs are significantly enriched because a low concentration ensures that most cells only enter one vector, meeting the screening requirements.

[0147] In the context of the present specification, the term "nuclease" refers to a class of enzymes that catalyze the hydrolysis of phosphodiester bonds with nucleic acid as substrate. In some embodiments, the desired function of the genetic element is coupled to the expression of a nuclease. In some embodiments, the nuclease is a sequence-specific nuclease.

[0148] In the context of the present specification, the term "secondary replicon" refers to a replicon formed as follows: the primary replicon of a geminivirus produces a double-stranded DNA break (DSB) under the action of a sequence-specific nuclease, which can be connected to the right border (RB) of the plasmid under the guidance of VirD2, and then form a replicon under the action of Rep / RepA. In some embodiments, the directed evolution of the genetic element is achieved by the following: a functional genetic element forms a secondary replicon by activating the expression of a nuclease or guiding the nuclease to cut its recognition site, thereby driving rolling circle replication and achieving self-enrichment; a non-functional genetic element cannot make the nuclease cut its recognition site, so it cannot form a secondary replicon and cannot be enriched.

[0149] In some embodiments, the genetic element is a DNA binding domain. In some embodiments, the method i) further comprises fusing a library of DNA binding domains to be evolved with a non-sequence-specific nuclease, and placing the recognition sequence of the nuclease together in the replicon. A functional DNA binding domain can guide the nuclease to cut the target sequence, and under the action of virD2, generate a secondary replicon; a non-functional DNA binding domain cannot guide the nuclease to cut the target sequence, so it does not generate a secondary replicon. By detecting the secondary replicon, the directed evolution of the DNA binding domain is achieved.

[0150] In some embodiments, the genetic element is a sequence encoding a non-sequence-specific nuclease.

[0151] In some embodiments, the genetic element is a sequence encoding a transcriptional activator. In some embodiments, the method i) further comprises inserting a recognition sequence of the transcriptional activator upstream of the nuclease, inserting a minimal transcription initiation element between the recognition sequence and the nuclease; and co-locating a library of transcriptional activators to be evolved with the recognition sequence of the nuclease in a replicon. A functional transcriptional activator can activate the expression of the nuclease and cleave its recognition sequence to form a secondary replicon under the action of virD2; a non-functional transcriptional activator cannot activate the expression of the nuclease and thus cannot generate a secondary replicon. By detecting the secondary replicon, directed evolution of the transcriptional activator can be achieved.

[0152] In some embodiments, the genetic element is a sequence encoding a recombinase. In some embodiments, the method i) further comprises co-locating a library of recombinases to be evolved with the recognition sequence of the nuclease in a replicon; and separating the nuclease into two parts flanking the recognition sequence of the recombinase. In some embodiments, the method i) further comprises adding a 5' intron and a 3' intron between the nuclease and the recognition sequence of the recombinase. A functional recombinase can mediate DNA recombination, the nuclease is expressed and cleaves its recognition site to generate a secondary replicon; a non-functional recombinase cannot mediate DNA recombination, the nuclease is not normally expressed and no secondary replicon is generated. By detecting the secondary replicon, directed evolution of the recombinase can be achieved.

[0153] In some embodiments, the genetic element is a PAM (protospacer adjacent motif) of a Cas protein. In some embodiments, the method further comprises co-locating a PAM to be evolved with a target sequence of the Cas protein in a replicon. A PAM that can be recognized by Cas will generate a DSB at the target region to form a secondary replicon, and the information of the PAM will be retained in the secondary replicon; a PAM that cannot be recognized by Cas will not generate a DSB and thus cannot form a secondary replicon. By detecting the secondary replicon, directed evolution of the PAM can be achieved.

[0154] In some embodiments, the genetic element is an sgRNA. In some embodiments, the method further comprises co-locating an sgRNA to be evolved with a target sequence of the Cas protein in a replicon. An active sgRNA can guide Cas9 to cleave the target site downstream of it to form a secondary replicon; a non-active sgRNA cannot generate a DSB and thus cannot generate a secondary replicon. By detecting the secondary replicon, directed evolution of the sgRNA can be achieved.

[0155] In some embodiments, in step iii), detecting and selecting the genetically enriched element mutants that are enriched in the population of plant cells can be performed by high-throughput sequencing. For example, total DNA of the population of plant cells can be extracted and subjected to high-throughput sequencing against the genetic element.

[0156] In some embodiments, the method further comprises step iv) identifying the function of the enriched genetically element mutants.

[0157] In one aspect, the present application provides genetically element mutants or the encoded products thereof obtained by the method of the present application, and the use of the obtained genetically element mutants or the encoded products thereof in plants, particularly in plant genetic engineering.

[0158] In one aspect, the present application provides a kit for performing the method of the present application. The kit may, for example, comprise a vector comprising a geminivirus replicon, and / or a vector for expressing a geminivirus Rep and / or RepA protein. The kit can further comprise instructions for performing the method of the present application. Examples

[0159] Further understanding of the present application can be obtained by reference to the specific examples given herein, which are intended for purposes of illustration only and are not intended to limit the scope of the present application. Obviously, many modifications and variations of this application are possible in light of its teachings, and it is intended that the scope of the application encompass these modifications and variations.

[0160] A wheat dwarf virus (WDV) and bean yellow dwarf virus (BeYDV) replicon system was developed. Both viruses belong to the genus Polerovirus and have very similar genomic structures, and can achieve efficient genome amplification in monocotyledonous and dicotyledonous plants, respectively.

[0161] 1. Directed evolution system based on primary replicon

[0162] The existing directed evolution systems have various methods, but the core idea is similar, that is, to enrich the functional target gene (Gene of Interest, GOI), and to filter the non-functional target gene. However, compared with bacteria and yeast, the division (i.e. genome DNA replication) of plant cells is very slow, which is difficult to meet the needs, therefore, it is hoped to use the replication of viruses to replace the division of plant cells, so as to realize the enrichment of target genes.

[0163] In the geminivirus system, the tandem LIRs on the vector can be recognized by Rep / RepA and circularized into primary replicon (PR), which can then undergo rolling circle replication, and the copy number can be increased by about 3 orders of magnitude. Among them, Rep / RepA is the only protein required for this process. Based on this principle, a screening library of target genes (which can be generated by error-prone PCR or saturation mutation) can be cloned into the geminivirus replicon, and other elements can be added to make the desired function of the target gene coupled with the expression of Rep / RepA, so as to construct a geminivirus primary replicon-based in vivo directed evolution system (such as Figure 4 ) in plants. In this system, target gene alleles with desired functions can directly or indirectly drive the expression of Rep / RepA, so as to enrich themselves; and target gene alleles without desired functions cannot initiate the expression of Rep / RepA, and cannot be enriched. Then, through deep sequencing, it can be inferred which alleles are functional, so as to achieve the purpose of evolution.

[0164] Using the primary replicon, the inventors expect to realize the directed evolution of genetic elements such as promoters, transcription activators, DNA binding proteins, and recombinases. To realize the directed evolution of promoters, the promoter library to be evolved can be placed upstream of Rep / RepA. Functional promoters can drive the expression of downstream Rep / RepA, drive the rolling circle replication of themselves, increase the copy number, and realize enrichment; non-functional promoters cannot drive the expression of downstream Rep / RepA, and cannot realize enrichment, thereby realizing the directed evolution of promoters (such as Figure 5 ). In the directed evolution of transcription activators, the recognition sequence of the transcription activator can be inserted upstream of Rep / RepA, and a mini-promoter can be added; the transcription activator library to be evolved can be inserted into the replicon. Functional transcription activators can bind to their recognition sequences and activate the expression of downstream Rep / RepA, thereby driving rolling circle replication and realizing self-enrichment; non-functional transcription activators cannot activate downstream Rep / RepA, and cannot realize enrichment, thereby realizing the directed evolution of transcription activators (such as Figure 6). In the directed evolution of DNA binding domain, the DNA binding domain to be evolved can be fused with a transcriptional activator without sequence specificity to form a fusion protein, which is placed in the replicon together; the target binding sequence of the DNA binding domain is inserted upstream of Rep / RepA, and is also assisted by a minimal transcription initiation element. In this way, the functional DNA binding domain can bind to its target sequence and bring the transcriptional activator to the minimal transcription initiation element to drive the expression of downstream Rep / RepA, drive the rolling circle replication, and achieve self-enrichment; the DNA binding domain without function cannot bind to the target sequence, cannot activate the downstream Rep / RepA, and cannot achieve enrichment, thereby realizing the directed evolution of the DNA binding domain (such as Figure 7 ). To realize the directed evolution of recombinase, the recombinase to be evolved can be placed in the replicon; Rep / RepA is divided into two parts and placed on the two wings of the recombinase recognition sequence. In order to ensure that Rep / RepA can normally function after recombination, 5' intron and 3' intron can be added between Rep / RepA and the recombinase recognition sequence, so that the recombinase recognition sequence can be cut off after transcription, and Rep / RepA can be normally translated. In this way, the functional recombinase can recognize its specific recognition site, mediate DNA recombination, and Rep / RepA can be normally expressed, thereby driving the rolling circle replication and realizing its self-enrichment; the recombinase without function cannot mediate DNA recombination and cannot make Rep / RepA express, thereby cannot achieve enrichment, thereby realizing the directed evolution of the recombinase (such as Figure 8 ).

[0165] 2. Directed evolution system based on secondary replicon

[0166] In the Agrobacterium-mediated plant genetic transformation system, a series of Vir proteins encoded by Agrobacterium can recognize the right border (RB) sequence on the Ti plasmid and produce a single-stranded DNA break nick at a specific position thereon, and then the VirD2 protein can covalently bind to the 5' DNA end of the nick, release the T-DNA sequence, and be transported into the plant cell nucleus. In the nucleus, VirD2 can recognize the double-stranded DNA break (DSB) spontaneously generated on the plant genome and connect it through non-homologous end joining (NHEJ) and other ways under the action of a series of host factors, and insert the T-DNA sequence into the plant genome.

[0167] Based on this principle, the inventors found in experiments that if a DSB is artificially generated on the geminivirus replicon by a sequence-specific nuclease, the RB region can be ligated to the break under the guidance of VirD2, and then form a secondary replicon (SR) under the action of Rep / RepA (as shown in Figure 9 ). The inventors subsequently found that this ligation can be divided into two modes, cis-ligation and trans-ligation, and the former is dominant. In summary, whether a secondary replicon is generated depends entirely on whether a DSB is generated at a specific position by a sequence-specific nuclease.

[0168] In addition, compared with other directed evolution methods, the directed evolution relying on secondary replicon has the following advantages: first, in the evolution process, in order to ensure that most cells only enter one vector, the concentration of the target gene library must be kept low, but this also means that the initial expression of the target gene is low, which may not be enough to meet the screening needs. In the directed evolution relying on secondary replicon, the target gene will first undergo the first round of rolling circle replication under the action of Rep to form a primary replicon, and in this process, the copy number of the target gene can be increased by three orders of magnitude in a short time, and the expression is greatly improved, so that it is enough to meet the following screening steps. Second, in the process of directed evolution, the secondary replicon has undergone two enrichments: one is the generation of the secondary replicon under the action of the sequence-specific nuclease; the other is the second round of rolling circle replication of the secondary replicon under the action of Rep / RepA, which greatly improves the copy number. These make the directed evolution relying on secondary replicon have extraordinary advantages.

[0169] According to this principle, the secondary replicon can be used to high-throughput screen or evolve sequence-specific nucleases (such as Figure 10 ), and can also be used to study the cleavage mode of sequence-specific nucleases (such as PAM of Cas nuclease) and guide RNA (such as sgRNA of Cas9 or crRNA of Cas12a). In addition, genetic elements that can be coupled with nuclease expression can also be evolved in high throughput.

[0170] For example, DNA binding domains can be evolved. The inventors fuse a library of DNA binding domains to be evolved with a non-sequence-specific nuclease, and place them together with the target sequence in the replicon. In this way, the functional DNA binding domain can guide the nuclease to cut the target sequence, and under the action of virD2, a secondary replicon is generated; the non-functional DNA binding domain cannot guide the nuclease to cut the target sequence, so that no secondary replicon is generated. By detecting the secondary replicon, the directed evolution of the DNA binding domain is achieved (such as Figure 11). Similarly, directed evolution of non-sequence-specific nucleases can also be achieved (e.g. Figure 12 ). In addition, directed evolution of transcriptional activators can also be achieved using secondary replicon. The recognition sequence of the transcriptional activator can be placed upstream of the nuclease, and assisted by a mini-promoter; the library of transcriptional activators to be evolved is placed in the replicon together with the recognition sequence of the nuclease. In this way, the functional transcriptional activator can activate the expression of the nuclease, and cut its recognition sequence to form the secondary replicon under the action of virD2; the non-functional transcriptional activator cannot activate the expression of the nuclease, and thus cannot generate the secondary replicon. By detecting the secondary replicon, the directed evolution of the transcriptional activator can be achieved (e.g. Figure 13 ). Using the secondary replicon-dependent directed evolution system, the directed evolution of recombinase can also be achieved. The library of recombinase to be evolved is placed in the replicon together with the recognition sequence of the nuclease, and the nuclease is divided into two parts and placed on both sides of the recognition sequence of the recombinase. In order to ensure that the recombined nuclease can normally function, 5' intron and 3' intron can be added between the nuclease and the recognition sequence of the recombinase, so that the recognition sequence of the recombinase can be cut off after transcription, and the nuclease can be normally translated. In this way, the functional recombinase can mediate DNA recombination, the nuclease can be expressed, and its recognition site can be cut, and the secondary replicon can be generated; the non-functional recombinase cannot mediate DNA recombination, the nuclease cannot be normally expressed, and the secondary replicon cannot be generated. By detecting the secondary replicon, the directed evolution of the recombinase can be achieved (e.g. Figure 14 ).

[0171] Experimental materials and methods

[0172] 1. Cultivation of wheat seedlings:

[0173] Wheat seeds were planted in a culture room under the conditions of temperature 25±2℃, light intensity 1000Lx, and light illumination 14-16h / d, and cultivated for about 1-2 weeks.

[0174] 2. Protoplast isolation:

[0175] 1) The middle part of the young leaves of wheat was cut into 0.5-1mm filaments, placed in 0.6M Mannitol solution for 10 minutes of light-avoiding treatment, and then filtered with a filter screen, and placed in 50ml enzyme solution for 5 hours of light-avoiding digestion at 20-25℃ (first 0.5h of static enzyme digestion, and then 4.5h of slow shaking at 10rpm).

[0176] 2) 10ml W5 solution was added; the pH value was 5.7, the enzyme digestion product was diluted, and the enzyme digestion solution was filtered with a 75um nylon filter membrane in a 50ml round-bottom centrifuge tube.

[0177] 3) 100 g, 23°C, 3 min, discard supernatant.

[0178] 4) 10 ml W5 solution, gently suspend the precipitate, place on ice for 30 min to allow protoplasts to gradually settle, discard supernatant.

[0179] 5) Add an appropriate amount of MMG solution to suspend, adjust the concentration of protoplasts to 2 x 10 5 / ml-1 x 10 6 / ml, place on ice, and wait for transformation.

[0180] 3. Wheat protoplast transformation

[0181] 1) Add 20 μg plasmid to a 2 ml centrifuge tube, add 200 μl protoplasts with a blunt-ended gun, gently mix, stand for 3-5 min, add 250 μl PEG solution, gently mix, avoid light, induce transformation for 30 min.

[0182] 2) Add 900 μl W5 solution, mix well at room temperature, 80 g centrifuge for 3 min, discard supernatant.

[0183] 3) Add 1 ml W5 solution, mix well, gently transfer to a 6-well plate with 1 ml W5 solution added in advance, 23°C culture for 24-48 h.

[0184] 4. Fluorescent quantitative PCR detection of amplicon copy number

[0185] After 24-48 h of culture, wheat protoplast DNA was extracted. DpnI was used for treatment to digest residual plasmid DNA. The PCR system was as follows: BIO-RAD iTaq Universal SYBR Green Mix 10 μL, diluted DNA template 8.4 μL, F primer 0.8 μL, R primer 0.8 μL. The qPCR program was as follows: 95°C for 30 s, 95°C for 10 s, 60°C for 15 s, 38 cycles. Primer WDVLIR-qF / R was used to amplify WDV replicon, and primer TaPDS-qF / R was used to amplify genomic DNA. The Ct value of qPCR results was converted into absolute concentration using a standard curve, and the ratio of the two was calculated to obtain the copy number of WDV amplicon.

[0186] 5. Tobacco plant culture

[0187] A layer of filter paper was placed in a culture dish and soaked with water, and tobacco seeds were scattered on the filter paper. The culture was carried out at 22°C under light for about 5 days. The germinated seedlings were transplanted into culture pots, and the culture was carried out under the conditions of temperature 22 ± 2°C, light intensity 1000 Lx, and light illumination 14-16 h / d for 4 weeks.

[0188] 6. Agrobacterium-mediated tobacco transient transformation.

[0189] Agrobacterium with target plasmid was inoculated in LB medium containing kanamycin, rifampicin, and cultured at 28°C overnight. Next day, 0.3ml of the turbid Agrobacterium culture was re-inoculated into 6ml fresh medium and cultured at 28°C for 4-6 hours. When the culture reached OD 0.6-1.0, it was centrifuged, the pellet was re-suspended in tobacco infiltration solution, and the OD was adjusted to the target concentration (not to exceed 1.6). The suspension was incubated at room temperature in the dark for 30min to 3 hours. Flat and healthy tobacco leaves were selected and the incubated suspension was injected into them using a syringe. Samples were taken for analysis 48h-96h later.

[0190] 7. Deep sequencing method

[0191] 1) Extract protoplast DNA, then treat with DpnI to digest residual plasmid DNA.

[0192] 2) In Example 1, the DNA was subjected to first-round PCR amplification using primers 35Sp-200F and WDV-Rep-150R. The first-round PCR product was subjected to second-round amplification using barcode primers ngs35Sp-300F and ngsWDV-Rep-100R. In Example 2, the DNA was subjected to first-round PCR amplification using primers 35Sp-200F and WDV-Rep-100R. The first-round PCR product was subjected to second-round amplification using barcode primers ngs35Sp-250F and ngsWDV-Rep-50R.

[0193] 3) The second-round PCR products were gel-recovered, mixed in equal proportions, and sent to a company for library construction and deep sequencing.

[0194] The compositions of the solutions used in the above method are as follows:

[0195] 50ml enzyme solution

[0196]

[0197] 500ml W5

[0198]

[0199] 10ml MMG solution

[0200] Reagent Amount added Final concentration Mannitol (0.8 M) 5ml 0.4M MgCl2(1M) 0.15ml 15 mM MES (200 mM) 0.2ml 4 mM DDW To 10 ml

[0201] 4ml PEG solution

[0202] Reagent Amount added Final concentration PEG 4000 1.6g 40% Mannitol (0.8 M) 1ml 0.2M CaCl2(1M) 0.4ml 0.1M DDW To 4 ml

[0203] 5. Tobacco infiltration solution

[0204]

[0205]

[0206] Example 1: Evolution of the 20th amino acid of Rep / RepA using the geminivirus-assisted plant directed evolution system

[0207] The 20th amino acid of wild-type wheat dwarf virus Rep / RepA is tyrosine, which is highly conserved in the genus. Previous studies have shown that the mutant Rep / RepAY20C of Rep does not have the initial replication activity. In order to verify whether the directed evolution system of the present application can work, the evolution of the 20th amino acid of Rep / RepA is first attempted.

[0208] First, the codon of the 20th amino acid of Rep / RepA is changed from TAT to NNN by PCR, and then a library with a diversity of 64 (4 3 ) is obtained by cloning the PCR fragment into the geminivirus vector (as Figure 15 ). Then the library is transformed into wheat protoplasts at different concentration gradients (10 μg-0.1 ng, divided into 10 concentration gradients). After 48 h, the protoplast DNA is extracted, and the site is deeply sequenced. The sequencing results of each concentration are compared with the results of the initial library.

[0209] It is found that among the 64 alleles contained in the library, 12 alleles encoding proline, glutamine, arginine, leucine, tyrosine, lysine, and alanine (as Figure 16 ) are enriched. 35% of the amino acids are positively selected, which does not conform to the initial expectation. The presumed reason is that the expression amount of Rep / RepA is positively correlated with the copy number of rolling circle replication, and when the use amount of the library in protoplast transformation is gradually reduced, the proportion of cells into which only one molecule is transformed increases, and when the expression amount of Rep / RepA expressed by a single molecule is relatively low, it is not enough to drive efficient rolling circle replication of the replicon. If the use amount of the library is higher, the probability of co-transformation of functional alleles and non-functional alleles into the same cell increases, which will cause the amplification of non-functional alleles. Such a result is to make the background noise very high.

[0210] Based on such results, the inventors hope to find a replication enhancer. Such replication enhancer cannot initiate rolling circle replication independently, but on the other hand, when the expression level of Rep / RepA is relatively low, the presence of the replication enhancer can greatly increase the copy number of replicon. Previous studies have shown that the Rep / RepA of the geminivirus is a multifunctional protein, which can initiate rolling circle replication, is a post-transcriptional gene silencing (PTGS) inhibitor, a transcription activator of viral encoded genes, can interact with endogenous proteins to make mature cells regain the ability of high level DNA replication, etc. Based on such assumption, the inventors try to find a mutant of Rep / RepA, which on the one hand can no longer initiate rolling circle replication, and on the other hand still retains other functions except for initiating rolling circle replication, so as to be used as a replication enhancer.

[0211] Firstly, the inventors constructed several Rep / RepA mutants, including Y106H, K229E, Y20C, H59R, E198A and H91R, and then detected the replication initiation activity of these mutants by fluorescence quantitative PCR. The results showed that, except for H91R, the other mutants all completely had no replication initiation activity. Then, in order to screen the appropriate replication enhancer, the added amount of Rep / RepA plasmid was reduced from the original 10 μg to 100 ng, and under this concentration, the copy number level of replicon could only reach a low level. In this case, the addition of replication enhancer should greatly improve the level of copy number. After detection, it was found that RepA, Rep / RepA K229E and Y20C could be used as replication enhancers. In the subsequent experiments, Rep / RepA Y20C was used as the replication enhancer.

[0212] Evolution experiments of the 20th codon of Rep / RepA were carried out again. It was found that after the addition of the replication enhancer, four alleles were enriched, which respectively encoded phenylalanine and tyrosine (such as Figure 17 ). In addition to the expected tyrosine, phenylalanine also occurred enrichment. The specific reason is that the motif in which Rep / RepA Y20 is located needs to interact with DNA, and phenylalanine has aromatic residues like tyrosine, so it can also form π-π stacking with DNA bases to function. This work shows that the system of the present application can work.

[0213] In this experiment, it was found that the amount of the added library was also a key factor affecting the evolution system. In a series of concentration gradients, it was found that when the dilution multiple reached between 5-10, that is, the amount of the added was about 1×10 -13 mol to 1×10 -15 mol, which was 103 up to 10 5 The effect of screening is best when the ratio is 2-10.

[0214] Example 2, Evolution of CaMV 35S promoter TATA-box using a geminivirus-assisted plant directed evolution system

[0215] Cauliflower mosaic virus (CaMV) 35S promoter is a commonly used constitutive promoter in plants. It has a TATA-box motif at about 30bp upstream of its transcription initiation site, with the sequence of 5'-ctatataag-3'. Most eukaryotic pol II promoters have a TATA-box element, which is essential for the transcriptional activity of the promoter. In this example, we plan to evolve the CaMV 35S promoter TATA-box, and at the same time study the effect of the sequence of this element on the activity of the CaMV 35S promoter.

[0216] First, we couple the activity of CaMV 35S promoter with the expression of Rep / RepA. We use CaMV 35S promoter to drive WDV Rep / RepA, and put both of them into the replicon.

[0217] Second, we construct the screening library. We change the TATA-box of CaMV 35S promoter from CTATATAAG to CNNNNNNNG by PCR, and then clone this PCR fragment into the geminivirus vector, obtaining a library with a theoretical diversity of 16384 (4 7 ) (e.g. Figure 18 ). Then we transform wheat protoplasts with this library at different concentrations. After 48h, we extract the protoplast DNA and sequence this site. We compare the sequencing results of each concentration with the results of the initial library.

[0218] We find that, through Sanger sequencing and amplicon sequencing, the library with the sequence of CNNNNNNNG, after screening, the sequence corresponding to the original TATA-box is enriched, so that the sequence almost all changes to CTATATAAG (e.g. Figure 19 ). This is consistent with the expected result.

[0219] Example 3, Rapid screening of pegRNA using a directed evolution system based on primary replicon

[0220] The prime editing system is a gene editing system that can make arbitrary base substitutions and small fragment deletions or insertions, which includes two parts, namely the nCas9 (H840A) protein fused with the M-MLV reverse transcriptase, and the prime editing guide RNA (pegRNA). The prime editing system can work relatively efficiently in yeast and animal cells, but it is extremely inefficient in plants and has very strong site specificity. The pegRNA includes four parts, the spacer part is responsible for recognizing the target point, the scaffold part is responsible for binding with nCas9, the primer binding site (PBS) is responsible for complementing the sequence at the 5' end of the nCas nick as a primer, and the reverse transcription template (RT) part is responsible for repairing the 3' end of the nick to the established sequence as a reverse transcription template. For a target point, the spacer and scaffold regions are fixed, but the PBS region and the RT region have great variability, and the length and sequence of the two have a great influence on the efficiency of prime editing. Therefore, it is desired to establish a high-throughput plant system capable of screening pegRNA.

[0221] First, the function of pegRNA needs to be coupled with the expression of Rep / RepA. The target point is inserted at the N-terminus of Rep / RepA, and the open reading frame of Rep / RepA is shifted. When pegRNA is absent or has no efficiency, Rep / RepA cannot be expressed. However, if pegRNA is active, it can introduce a short insertion or deletion at the target point, so that Rep / RepA can be correctly expressed. Based on this principle, the expression cassette of pegRNA is inserted into the geminivirus replicon, and a fluorescent reporter system is inserted on both sides (such as Figure 20 ). If the viral replicon undergoes rolling circle replication under the action of pegRNA, the fluorescent signal reports; if the viral replicon does not undergo rolling circle replication, there is no fluorescent signal.

[0222] To verify whether the system works, two vectors are constructed, one of which contains a pegRNA with known activity, and the other introduces a mutation in the PBS of pegRNA, making it lose activity. The two vectors are mixed in equal proportions to form a library, and are transformed into tobacco leaves at different concentrations. After 6 days, DNA is extracted and detected, and the applicant found that when the library is transformed into tobacco leaves at a high concentration, the active pegRNA is not enriched; when the library is transformed into tobacco leaves at a low concentration, the active pegRNA is significantly enriched (such as Figure 21 ), which is consistent with expectations and proves that the system can work.

[0223] Example 4 Rapid identification of PAM of Cas protein using directed evolution system based on secondary replicon

[0224] In the field of genome editing, researchers have discovered many Cas proteins with nuclease activity (including Cas9, Cas12a, Cas12b, etc.). One of the characteristics of these Cas proteins is the need for a specific sequence, called PAM (protospacer adjacent motif), upstream or downstream of the cleavage target. Different Cas proteins have different PAM sequences.

[0225] In previous studies, bioinformatics analysis was used to find a Cas12a protein in Flavobacterium branchiophilum, called FbCas12a. In order to make it applicable in biotechnology, its PAM sequence needs to be determined first. For this purpose, a library containing 4096 PAMs was constructed (such as Figure 22 ).

[0226] If a PAM can be recognized by FbCas12a, the latter will produce DSB in the target region, and then form a secondary replicon, and the information of the PAM will be left in the secondary replicon; if a PAM cannot be recognized by FbCas12a, DSB will not be produced, and the PAM will not be left in the secondary replicon. Therefore, as long as the secondary replicon is specifically detected, it can be known which PAM can be recognized by FbCas12a (such as Figure 23 ).

[0227] Two target points, OsEPSPSc3 and c5, were selected for testing. After testing, it was found that FbCas12a can recognize the PAM 'TTT' (such as Figure 24 ). And in other positions, no obvious base preference was found (such as Figure 25 ).

[0228] Example 5 Rapid screening of sgRNA using directed evolution system based on secondary replicon

[0229] For Cas proteins, a piece of RNA is needed to guide its nucleic acid cleavage, called sgRNA (Cas9) or crRNA (Cas12a, Cas12b). Its sequence and structure have a great influence on the activity of Cas protein. Therefore, it is hoped to establish a system that can high-throughput and rapidly screen sgRNA.

[0230] Similar to Example 4, the vector shown in the figure was designed (such as Figure 26). If a sgRNA is active, it can guide Cas9 to cut the target downstream, thus forming a secondary replicon; if a sgRNA is not active, it cannot generate a DSB, and thus cannot generate a secondary replicon. Based on this principle, two vectors were constructed, one of which contains an active sgRNA, and the other contains a non-active sgRNA. After 5 days of screening, it was found that in the secondary replicon, the active sgRNA was significantly screened, and was related to the library concentration (e.g. Figure 27 ) This proves that the system is working.

[0231] SEQUENCE LISTING

[0232] >SEQ ID NO 1 WDV-LIR

[0233] GGTAGTGAACAGAAGTCCGGCAGGTCCTTAGCGAAAAAACGGGGTGTGCCAGAAAACTCTATCCTCTACCCTGCGTGGAGGTGTGAATTCTGCACACTGCAAATGCAATGTGTCCAATGCTTTATATAGGGCAGGTTTTGGCGGGAGAACAGGGCCCTAGTGTTCCCACGGTAGCGTAGCGAATCGTGTGGGCCCTGTTCGGTGTGCGGTCGGGGGGCCTCCACGCGGGTTATAATATTACCCCGCGTGGTGGCCCCCGACGCGCACTCGGCTTTTCGTGAGTGCGCGGAGGCTTTTGGACCACATCTTTTCTGATCACTTTCGTGGAAGATGTTGATTTATCACACTTTTGACGGGGAAATCTGTGCCATGCCTTAGCTTATAAGGAAGTGCGTGGTAGCCCATCTCG

[0234] >SEQ ID NO 2 WDV-SIR

[0235] TAAAATAATATTTTATTTATCTCATGTCATTCGATTACAGAGGCTCGGCTACGAGCAAAGACAAACCAAATATAACAAACAACAACCCTTACACAATGACATCGGAAAACGAAATACAACACCCTGAGATATTACATTTATAGAAACTGTACGCCGTCCGCGCTAGGACAG

[0236] >SEQ ID NO 3 WDV-Rep

[0237] MASSSAPRFRVYSKYLFLTYPQCTLEPQYALDSLRTLLNKYEPLYIAAVRELHEDGSPHLHVLVQNKLRASITNPNALNLRMDTSPFSIFHPNIQAAKDCNQVRDYITKEVDSDVNTAEWGTFVAVSTPGRKDRDADMKQIIESSSSREEFLSMVCNRFPFEWSIRLKDFEYTARHLFPDPVATYTPEFPTESLICHETIESWKNEHLYSESPGRHKSIYICGPTRTGKTSWARSLGTHNYYNSLVDFTTYDVNAKYNIIDDIPFKFTPNWKCFVGAQRDFTVNPKYGKRKVIRGGIPCIILVNPDEDWLKDMTPEQSDYMYSNTVVHYMYEGETFINYSFASGEDVTASQ*

[0238] >SEQ ID NO 4 WDV-Rep Y20C

[0239] MASSSAPRFRVYSKYLFLTCPQCTLEPQYALDSLRTLLNKYEPLYIAAVRELHEDGSPHLHVLVQNKLRASITNPNALNLRMDTSPFSIFHPNIQAAKDCNQVRDYITKEVDSDVNTAEWGTFVAVSTPGRKDRDADMKQIIESSSSREEFLSMVCNRFPFEWSIRLKDFEYTARHLFPDPVATYTPEFPTESLICHETIESWKNEHLYSESPGRHKSIYICGPTRTGKTSWARSLGTHNYYNSLVDFTTYDVNAKYNIIDDIPFKFTPNWKCFVGAQRDFTVNPKYGKRKVIRGGIPCIILVNPDEDWLKDMTPEQSDYMYSNTVVHYMYEGETFINYSFASGEDVTASQ*

[0240] >SEQ ID NO 5 WDV-RepA

[0241] MASSSAPRFRVYSKYLFLTYPQCTLEPQYALDSLRTLLNKYEPLYIAAVRELHEDGSPHLHVLVQNKLRASITNPNALNLRMDTSPFSIFHPNIQAAKDCNQVRDYITKEVDSDVNTAEWGTFVAVSTPGRKDRDADMKQIIESSSSREEFLSMVCNRFPFEWSIRLKDFEYTARHLFPDPVATYTPEFPTESLICHETIESWKNEHLYSVSLESYILCTSTPADQAQSDLEWMDDYSRSHRGGISPSTSAGQPEQERLPGQGL

[0242] >SEQ ID NO 6 WDV-RepA Y20C

[0243] MASSSAPRFRVYSKYLFLTCPQCTLEPQYALDSLRTLLNKYEPLYIAAVRELHEDGSPHLHVLVQNKLRASITNPNALNLRMDTSPFSIFHPNIQAAKDCNQVRDYITKEVDSDVNTAEWGTFVAVSTPGRKDRDADMKQIIESSSSREEFLSMVCNRFPFEWSIRLKDFEYTARHLFPDPVATYTPEFPTESLICHETIESWKNEHLYSVSLESYILCTSTPADQAQSDLEWMDDYSRSHRGGISPSTSAGQPEQERLPGQGL SEQUENCE LISTING <110> Beijing Qihuo Biotechnology Co., Ltd. <120> A directed evolution method based on geminivirus primary and secondary replicons <130> P2020TC1345 <160> 6 <170> PatentIn version 3.5 <210> 1 <211> 409 <212> DNA <213> Wheat dwarf virus <220> <223> WDV-LIR <400> 1 ggtagtgaac agaagtccgg caggtcctta gcgaaaaaac ggggtgtgcc agaaaactct 60 atcctctacc ctgcgtggag gtgtgaattc tgcacactgc aaatgcaatg tgtccaatgc 120 tttatatagg gcaggttttg gcgggagaac agggccctag tgttcccacg gtagcgtagc 180 gaatcgtgtg ggccctgttc ggtgtgcggt cggggggcct ccacgcgggt tataatatta 240 ccccgcgtgg tggcccccga cgcgcactcg gcttttcgtg agtgcgcgga ggcttttgga 300 ccacatcttt tctgatcact ttcgtggaag atgttgattt atcacacttt tgacggggaa 360 atctgtgcca tgccttagct tataaggaag tgcgtggtag cccatctcg 409 <210> 2 <211> 171 <212> DNA <213> Wheat dwarf virus <220> <223> WDV-SIR <400> 2 taaaataata ttttatttat ctcatgtcat tcgattacag aggctcggct acgagcaaag 60 acaaaccaaa tataacaaac aacaaccctt acacaatgac atcggaaaac gaaatacaac 120 accctgagat attacattta tagaaactgt acgccgtccg cgctaggaca g 171 <210> 3 <211> 351 <212> PRT <213> Wheat dwarf virus <220> <223> WDV-Rep <400> 3 Met Ala Ser Ser Ser Ala Pro Arg Phe Arg Val Tyr Ser Lys Tyr Leu 1 5 10 15 Phe Leu Thr Tyr Pro Gln Cys Thr Leu Glu Pro Gln Tyr Ala Leu Asp 20 25 30 Ser Leu Arg Thr Leu Leu Asn Lys Tyr Glu Pro Leu Tyr Ile Ala Ala 35 40 45 Val Arg Glu Leu His Glu Asp Gly Ser Pro His Leu His Val Leu Val 50 55 60 Gln Asn Lys Leu Arg Ala Ser Ile Thr Asn Pro Asn Ala Leu Asn Leu 65 70 75 80 Arg Met Asp Thr Ser Pro Phe Ser Ile Phe His Pro Asn Ile Gln Ala 85 90 95 Ala Lys Asp Cys Asn Gln Val Arg Asp Tyr Ile Thr Lys Glu Val Asp 100 105 110 Ser Asp Val Asn Thr Ala Glu Trp Gly Thr Phe Val Ala Val Ser Thr 115 120 125 Pro Gly Arg Lys Asp Arg Asp Ala Asp Met Lys Gln Ile Ile Glu Ser 130 135 140 Ser Ser Ser Arg Glu Glu Phe Leu Ser Met Val Cys Asn Arg Phe Pro 145 150 155 160 Phe Glu Trp Ser Ile Arg Leu Lys Asp Phe Glu Tyr Thr Ala Arg His 165 170 175 Leu Phe Pro Asp Pro Val Ala Thr Tyr Thr Pro Glu Phe Pro Thr Glu 180 185 190 Ser Leu Ile Cys His Glu Thr Ile Glu Ser Trp Lys Asn Glu His Leu 195 200 205 Tyr Ser Glu Ser Pro Gly Arg His Lys Ser Ile Tyr Ile Cys Gly Pro 210 215 220 Thr Arg Thr Gly Lys Thr Ser Trp Ala Arg Ser Leu Gly Thr His Asn 225 230 235 240 Tyr Tyr Asn Ser Leu Val Asp Phe Thr Thr Tyr Asp Val Asn Ala Lys 245 250 255 Tyr Asn Ile Ile Asp Asp Ile Pro Phe Lys Phe Thr Pro Asn Trp Lys 260 265 270 Cys Phe Val Gly Ala Gln Arg Asp Phe Thr Val Asn Pro Lys Tyr Gly 275 280 285 Lys Arg Lys Val Ile Arg Gly Gly Ile Pro Cys Ile Ile Leu Val Asn 290 295 300 Pro Asp Glu Asp Trp Leu Lys Asp Met Thr Pro Glu Gin Ser Asp Tyr 305 310 315 320 Met Tyr Ser Asn Thr Val Val His Tyr Met Tyr Glu Gly Glu Thr Phe 325 330 335 Ile Asn Tyr Ser Phe Ala Ser Gly Glu Asp Val Thr Ala Ser Gin 340 345 350 <210> 4 <211> 351 <212> PRT <213> Artificial Sequence <220> <223> WDV-Rep Y20C <400> 4 Met Ala Ser Ser Ser Ala Pro Arg Phe Arg Val Tyr Ser Lys Tyr Leu 1 5 10 15 Phe Leu Thr Cys Pro Gin Cys Thr Leu Glu Pro Gin Tyr Ala Leu Asp 20 25 30 Ser Leu Arg Thr Leu Leu Asn Lys Tyr Glu Pro Leu Tyr Ile Ala Ala 35 40 45 Val Arg Glu Leu His Glu Asp Gly Ser Pro His Leu His Val Leu Val 50 55 60 Gln Asn Lys Leu Arg Ala Ser Ile Thr Asn Pro Asn Ala Leu Asn Leu 65 70 75 80 Arg Met Asp Thr Ser Pro Phe Ser lie Phe His Pro Asn lie Gin Ala 85 90 95 Ala Lys Asp Cys Asn Gin Val Arg Asp Tyr lie Thr Lys Glu Val Asp 100 105 110 Ser Asp Val Asn Thr Ala Glu Trp Gly Thr Phe Val Ala Val Ser Thr 115 120 125 Pro Gly Arg Lys Asp Arg Asp Ala Asp Met Lys Gin lie lie Glu Ser 130 135 140 Ser Ser Ser Arg Glu Glu Phe Leu Ser Met Val Cys Asn Arg Phe Pro 145 150 155 160 Phe Glu Trp Ser lie Arg Leu Lys Asp Phe Glu Tyr Thr Ala Arg His 165 170 175 Leu Phe Pro Asp Pro Val Ala Thr Tyr Thr Pro Glu Phe Pro Thr Glu 180 185 190 Ser Leu lie Cys His Glu Thr lie Glu Ser Trp Lys Asn Glu His Leu 195 200 205 Tyr Ser Glu Ser Pro Gly Arg His Lys Ser lie Tyr lie Cys Gly Pro 210 215 220 Thr Arg Thr Gly Lys Thr Ser Trp Ala Arg Ser Leu Gly Thr His Asn 225 230 235 240 Tyr Tyr Asn Ser Leu Val Asp Phe Thr Thr Tyr Asp Val Asn Ala Lys 245 250 255 Tyr Asn Ile Ile Asp Asp Ile Pro Phe Lys Phe Thr Pro Asn Trp Lys 260 265 270 Cys Phe Val Gly Ala Gln Arg Asp Phe Thr Val Asn Pro Lys Tyr Gly 275 280 285 Lys Arg Lys Val Ile Arg Gly Gly Ile Pro Cys Ile Ile Leu Val Asn 290 295 300 Pro Asp Glu Asp Trp Leu Lys Asp Met Thr Pro Glu Gln Ser Asp Tyr 305 310 315 320 Met Tyr Ser Asn Thr Val Val His Tyr Met Tyr Glu Gly Glu Thr Phe 325 330 335 Ile Asn Tyr Ser Phe Ala Ser Gly Glu Asp Val Thr Ala Ser Gln 340 345 350 <210> 5 <211> 264 <212> PRT <213> Wheat dwarf virus <220> <223> WDV‑RepA <400> 5 Met Ala Ser Ser Ser Ala Pro Arg Phe Arg Val Tyr Ser Lys Tyr Leu 1 5 10 15 Phe Leu Thr Tyr Pro Gin Cys Thr Leu Glu Pro Gin Tyr Ala Leu Asp 20 25 30 Ser Leu Arg Thr Leu Leu Asn Lys Tyr Glu Pro Leu Tyr Ile Ala Ala 35 40 45 Val Arg Glu Leu His Glu Asp Gly Ser Pro His Leu His Val Leu Val 50 55 60 Gln Asn Lys Leu Arg Ala Ser Ile Thr Asn Pro Asn Ala Leu Asn Leu 65 70 75 80 Arg Met Asp Thr Ser Pro Phe Ser Ile Phe His Pro Asn Ile Gin Ala 85 90 95 Ala Lys Asp Cys Asn Gin Val Arg Asp Tyr Ile Thr Lys Glu Val Asp 100 105 110 Ser Asp Val Asn Thr Ala Glu Trp Gly Thr Phe Val Ala Val Ser Thr 115 120 125 Pro Gly Arg Lys Asp Arg Asp Ala Asp Met Lys Gin Ile Ile Glu Ser 130 135 140 Ser Ser Ser Arg Glu Glu Phe Leu Ser Met Val Cys Asn Arg Phe Pro 145 150 155 160 Phe Glu Trp Ser Ile Arg Leu Lys Asp Phe Glu Tyr Thr Ala Arg His 165 170 175 Leu Phe Pro Asp Pro Val Ala Thr Tyr Thr Pro Glu Phe Pro Thr Glu 180 185 190 Ser Leu Ile Cys His Glu Thr Ile Glu Ser Trp Lys Asn Glu His Leu 195 200 205 Tyr Ser Val Ser Leu Glu Ser Tyr Ile Leu Cys Thr Ser Thr Pro Ala 210 215 220 Asp Gln Ala Gln Ser Asp Leu Glu Trp Met Asp Asp Tyr Ser Arg Ser 225 230 235 240 His Arg Gly Gly Ile Ser Pro Ser Thr Ser Ala Gly Gln Pro Glu Gln 245 250 255 Glu Arg Leu Pro Gly Gln Gly Leu 260 <210> 6 <211> 264 <212> PRT <213> Artificial Sequence <220> <223> WDV-RepA Y20C <400> 6 Met Ala Ser Ser Ser Ala Pro Arg Phe Arg Val Tyr Ser Lys Tyr Leu 1 5 10 15 Phe Leu Thr Cys Pro Gln Cys Thr Leu Glu Pro Gln Tyr Ala Leu Asp 20 25 30 Ser Leu Arg Thr Leu Leu Asn Lys Tyr Glu Pro Leu Tyr Ile Ala Ala 35 40 45 Val Arg Glu Leu His Glu Asp Gly Ser Pro His Leu His Val Leu Val 50 55 60 Gln Asn Lys Leu Arg Ala Ser Ile Thr Asn Pro Asn Ala Leu Asn Leu 65 70 75 80 Arg Met Asp Thr Ser Pro Phe Ser Ile Phe His Pro Asn Ile Gln Ala 85 90 95 Ala Lys Asp Cys Asn Gln Val Arg Asp Tyr Ile Thr Lys Glu Val Asp 100 105 110 Ser Asp Val Asn Thr Ala Glu Trp Gly Thr Phe Val Ala Val Ser Thr 115 120 125 Pro Gly Arg Lys Asp Arg Asp Ala Asp Met Lys Gln Ile Ile Glu Ser 130 135 140 Ser Ser Ser Arg Glu Glu Phe Leu Ser Met Val Cys Asn Arg Phe Pro 145 150 155 160 Phe Glu Trp Ser Ile Arg Leu Lys Asp Phe Glu Tyr Thr Ala Arg His 165 170 175 Leu Phe Pro Asp Pro Val Ala Thr Tyr Thr Pro Glu Phe Pro Thr Glu 180 185 190 Ser Leu Ile Cys His Glu Thr Ile Glu Ser Trp Lys Asn Glu His Leu 195 200 205 Tyr Ser Val Ser Leu Glu Ser Tyr Ile Leu Cys Thr Ser Thr Pro Ala 210 215 220 Asp Gln Ala Gln Ser Asp Leu Glu Trp Met Asp Asp Tyr Ser Arg Ser 225 230 235 240 His Arg Gly Gly Ile Ser Pro Ser Thr Ser Ala Gly Gln Pro Glu Gln 245 250 255 Glu Arg Leu Pro Gly Gln Gly Leu 260

Claims

1. A method of directed evolution of a genetic element to obtain a mutant of the genetic element having a desired function, the method comprising: i) providing a library of mutants of the genetic element comprising a plurality of mutants of the genetic element, each inserted into a vector comprising a geminivirus replicon; the mutants being inserted into the geminivirus replicon whereby the mutants are amplified when the geminivirus replicon replicates, ii) transforming a population of plant cells with the library, and iii) culturing the population of plant cells, detecting and selecting mutants of the genetic element that are enriched in the population of plant cells; wherein the level of replication in plant cells of a geminivirus replicon having a mutant of the genetic element having the desired function is higher than the level of replication in plant cells of a geminivirus replicon having a mutant of the genetic element not having the desired function; wherein the method complies with any one of (i)-(iii): (i) the vector comprising the geminivirus replicon comprises an expression cassette for geminivirus Rep and RepA proteins; or (ii) the vector comprising the geminivirus replicon comprises an expression cassette for geminivirus Rep or RepA proteins, and the method further comprises introducing into the plant cells a vector for expressing the other of geminivirus Rep or RepA proteins, the vector for expressing the other of geminivirus Rep or RepA proteins being co-transformed with the library into the population of plant cells; or, the plant cells already comprise a vector for expressing the other of geminivirus Rep or RepA proteins, or the genome of the plant cells already has integrated an expression cassette for the other of geminivirus Rep or RepA proteins; or (iii) the vector comprising the geminivirus replicon does not comprise an expression cassette for geminivirus Rep and RepA proteins, and the method further comprises introducing into the plant cells a vector for expressing the other of geminivirus Rep and RepA proteins, the vector for expressing the other of geminivirus Rep and RepA proteins being co-transformed with the library into the population of plant cells; or, the plant cells already comprise a vector for expressing the other of geminivirus Rep and RepA proteins, or the genome of the plant cells already has integrated an expression cassette for the other of geminivirus Rep or RepA proteins; and the vector comprising the geminivirus replicon further comprises at least one LIR, at least one SIR; the geminivirus belongs to the genus Maize streak virus; the genetic element is derived from a plant, or is desired for use in a plant; the genetic element is selected from a protein-coding sequence, a functional RNA-coding sequence, or an expression regulatory sequence; and the plant is a monocot or a dicot.

2. The method of claim 1, wherein the library of mutants of the genetic element is obtained by inserting a plurality of mutants of the genetic element, each, into a vector comprising a geminivirus replicon.

3. The method of claim 2, wherein the plurality of mutants of the genetic element is generated by random mutagenesis of the genetic element.

4. The method of claim 3, wherein the library is generated by random mutagenesis of the genetic element that has been inserted into a vector comprising a geminivirus replicon.

5. The method of any one of claims 1-4, wherein the geminivirus is Wheat dwarf virus (WDV).

6. The method of any one of claims 1-4, wherein the geminivirus is Bean yellow dwarf virus (BeYDV).

7. The method of any one of claims 1-4, wherein the vector comprising a geminivirus replicon is a circular DNA.

8. The method of claim 7, wherein the circular DNA is a plasmid or a minicircle DNA.

9. The method of claim 8, wherein the nucleotide sequence of the LIR is set forth in SEQ ID NO:

1.

10. The method of claim 9, wherein the nucleotide sequence of the SIR is set forth in SEQ ID NO:

2.

11. The method of claim 10, wherein the vector comprising a geminivirus replicon comprises one LIR.

12. The method of claim 10, wherein the vector comprising a geminivirus replicon comprises two LIRs.

13. The method of any one of claims 11 or 12, wherein the inserted mutant of the genetic element is operably linked to an expression control sequence in the vector comprising a geminivirus replicon.

14. The method of claim 1, wherein the amino acid sequence of the geminivirus Rep protein is set forth in SEQ ID NO: 3; or the amino acid sequence having the amino acid substitution Y20C relative to SEQ ID NO: 3 is set forth in SEQ ID NO:

4.

15. The method of claim 14, wherein the amino acid sequence of the geminivirus RepA protein is set forth in SEQ ID NO: 5; or the amino acid sequence having the amino acid substitution Y20C relative to SEQ ID NO: 5 is set forth in SEQ ID NO:

6.

16. The method of claim 15, wherein the activity or expression level of the geminivirus Rep and / or RepA protein in a plant cell comprising a mutant of the genetic element having the desired function is higher than the activity or expression level of the geminivirus Rep and / or RepA protein in a plant cell comprising a mutant of the genetic element not having the desired function.

17. The method of claim 1, wherein in the transformation of step ii) the number of vector molecules comprising the mutant in the library is 10 3 times the number of cells in the population of plant cells. 5 times the number of cells in the population of plant cells.

18. The method of any one of claims 16 or 17, wherein the directed evolution of the genetic element is accomplished by coupling the expression or activity of the geminivirus Rep and RepA proteins in the plant cell to the desired function of the mutant of the genetic element.

19. The method of claim 18, wherein the directed evolution of the genetic element is accomplished by the functional genetic element activating Rep / RepA expression, thereby driving rolling circle replication, enabling self-enrichment; the non-functional genetic element failing to activate Rep / RepA expression, thereby failing to enable enrichment.

20. The method of claim 1, wherein the genetic element is a promoter.

21. The method of claim 20, said method i) further comprising placing the library of promoters to be evolved upstream of Rep / RepA in the replicon.

22. The method of claim 20 or 21, said genetic element is the TATA-box of the Cauliflower Mosaic Virus (CaMV) 35S promoter.

23. The method of claim 1, said genetic element is a sequence encoding a transcriptional activator.

24. The method of claim 23, said method i) further comprising inserting a recognition sequence of the transcriptional activator upstream of Rep / RepA and a minimal transcription initiation element between the recognition sequence and Rep / RepA; and placing a library of transcriptional activators to be evolved in the replicon.

25. The method of claim 1, said genetic element is a DNA-binding domain.

26. The method of claim 25, said method i) further comprising inserting a target binding sequence of the DNA-binding domain upstream of Rep / RepA and a minimal transcription initiation element between the target binding sequence and Rep / RepA; and placing a DNA-binding domain to be evolved in a fusion protein with a non-sequence specific transcriptional activator in the replicon.

27. The method of claim 1, said genetic element is a sequence encoding a recombinase.

28. The method of claim 27, said method i) further comprising separating Rep / RepA into two parts flanking the recognition sequence of the recombinase; and placing a sequence encoding a recombinase to be evolved in the replicon.

29. The method of claim 27 or 28, said method i) further comprising adding a 5’ intron and a 3’ intron between Rep / RepA and the recognition sequence of the recombinase.

30. The method of claim 1, said genetic element is a prime editing guide RNA (pegRNA).

31. The method of claim 30, said method i) further comprising inserting a target site at the N-terminus of Rep / RepA and frameshifting the open reading frame of Rep / RepA; and inserting an expression cassette of the pegRNA in the geminivirus replicon flanking a fluorescent reporter system.

32. The method of claim 1, the desired function of the genetic element is coupled to the expression of a nuclease.

33. The method of claim 32, said nuclease is a specific nuclease.

34. The method of claim 32 or 33, the directed evolution of the genetic element is achieved by: a functional genetic element drives self-enrichment by activating the expression of a nuclease or directing the nuclease to cleave its recognition site, thereby driving rolling circle replication; a non-functional genetic element fails to enrich because it cannot make the nuclease cleave its recognition site.

35. The method of claim 34, said genetic element is a DNA-binding domain.

36. The method of claim 35, said method i) further comprising fusing a library of DNA-binding domains to be evolved to a non-sequence specific nuclease and placing them together with their recognition sequences in the replicon.

37. The method of claim 34, wherein the genetic element is a sequence encoding a non-sequence specific nuclease.

38. The method of claim 34, wherein the genetic element is a sequence encoding a transcriptional activator.

39. The method of claim 38, wherein the method of i) further comprises inserting a recognition sequence of the transcriptional activator upstream of the nuclease, inserting a minimal transcription initiation element between the recognition sequence and the nuclease; and co-locating a library of transcriptional activators to be evolved with the recognition sequence of the nuclease in the replicon.

40. The method of claim 34, wherein the genetic element is a sequence encoding a recombinase.

41. The method of claim 40, wherein the method of i) further comprises co-locating a library of recombinases to be evolved with the recognition sequence of the nuclease in the replicon; and separating the nuclease into two parts flanking the recognition sequence of the recombinase.

42. The method of claim 40 or 41, wherein the method of i) further comprises adding a 5' intron and a 3' intron between the nuclease and the recognition sequence of the recombinase.

43. The method of claim 34, wherein the genetic element is a PAM (protospacer adjacent motif) of a Cas protein.

44. The method of claim 43, further comprising co-locating a PAM to be evolved with a target sequence of the Cas protein in the replicon.

45. The method of claim 34, wherein the genetic element is an sgRNA.

46. The method of claim 45, further comprising co-locating an sgRNA to be evolved with a target sequence of the Cas protein in the replicon.

47. The method of claim 1, wherein in step iii), detecting and selecting the genetic element variant that is enriched in the population of plant cells is performed by high-throughput sequencing.

48. The method of claim 1, further comprising step iv) identifying the function of the enriched genetic element mutant.

49. The method of claim 1, wherein the plant is selected from the group consisting of maize, wheat, rice, barley, sorghum, bean, sugar beet, tomato, cassava, cucumber, Arabidopsis, and tobacco.

Citation Information

Patent Citations

  • Methods for isolating cells without using transgene marker sequences

    CN117051035A