Novel reverse transcriptase xrt1 and its use in prime editing
Patent Information
- Application Number
- CN202610984039.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-07-03
AI Technical Summary
[0004]尽管工程化MMLV-RT取得了巨大成功,但它仍存在固有的局限:其一,保真度瓶颈,其固有的错配倾向可能导致非目标核苷酸掺入,影响编辑纯度;其二,结构耐受性差,当pegRNA模板区或基因组靶点存在稳定二级结构(如发夹)时,其合成效率骤降;其三,尺寸较大,与Cas9融合后可能对病毒载体(如AAV)的包装递送构成挑战
[0007]本发明的目的是提供一种新型逆转录酶XRT1及其在引导编辑中的应用。
Smart Images

Figure CN122503350B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene editing, specifically to a novel reverse transcriptase XRT1 and its application in guided editing. Background Technology
[0002] Prime editing is a next-generation precision genome editing technology, with reverse transcriptase (RT) as one of its core components. Currently, this system primarily relies on multiple engineered Moroni murine leukemia virus reverse transcriptases (MMLV-RT). For example, the MMLV-RT used in the basic PE2 system has undergone evolutionary modification to acquire two key mutations (D200N / L603W), significantly enhancing its fusion compatibility with nCas9 (H840A), continuous synthesis capability, and stability at 37°C. Subsequent optimized versions, such as PEmax and the latest PE5max system, have further modified the RT (e.g., by adding nuclear localization signals and optimizing linkers) to improve its nuclear concentration, editing efficiency, and performance in complex genomic environments.
[0003] Discovering novel reverse transcriptases is crucial for advancing guided editing technology: First, it ensures independent control over core technologies, preventing dependence on foreign suppliers for fundamental components; second, it allows my country to seize early opportunities in the biopharmaceutical industry, helping it develop technological advantages in gene therapy, agricultural breeding, and other fields, thereby enhancing its international competitiveness; and third, it lays a leading foundation for cutting-edge research, promoting original discoveries in basic research and disease model construction. In summary, the continuous discovery and modification of novel reverse transcriptases is not only key to driving the iteration of guided editing technology but also an important measure for my country to achieve parity and even leadership in global biotechnology competition.
[0004] Despite the great success of engineered MMLV-RT, it still has inherent limitations: First, fidelity bottleneck, its inherent mismatch tendency may lead to the incorporation of non-target nucleotides, affecting editing purity; second, poor structural tolerance, its synthesis efficiency drops sharply when there are stable secondary structures (such as hairpins) in the pegRNA template region or genomic target site; third, its large size, after fusion with Cas9, may pose a challenge to the packaging and delivery of viral vectors (such as AAV).
[0005] Given the aforementioned shortcomings, the need to explore novel radiofrequency ablation (RT) methods is particularly urgent. Their core value lies in: first, finding RTs with high fidelity, strong strand substitution capabilities, or the ability to read complex templates, which can expand the range of editable sequences, reduce byproduct generation, and achieve cleaner, more predictable editing; second, exploring miniaturized RTs with good thermal stability, which can help develop more efficient and tissue-specific delivery systems and simplify editor configurations; and third, more efficient and safer RTs, which are the foundation for guiding editing towards clinical treatment (such as genetic disease repair and somatic cell therapy), and can improve the success rate and safety of human cell editing.
[0006] Therefore, there is an urgent need in the field for a novel reverse transcriptase, especially a miniaturized reverse transcriptase and a guided editing system containing such reverse transcriptase. Summary of the Invention
[0007] The purpose of this invention is to provide a novel reverse transcriptase XRT1 and its application in guided editing.
[0008] In a first aspect, the present invention provides a reverse transcriptase selected from the group consisting of: (1) It has the amino acid sequence shown in SEQ ID NO:1; (2) An amino acid sequence having ≥80%, ≥85%, ≥90%, ≥91%, ≥92%, ≥93%, ≥94%, ≥95%, ≥96%, ≥97%, ≥98%, ≥99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:1.
[0009] In a second aspect, the present invention provides a fusion protein comprising the reverse transcriptase and Cas protein described in the first aspect of the present invention.
[0010] In another preferred embodiment, the fusion protein has the structure shown in formula (I): Z1-Z2-Z3-Z4-Z5-Z6 (I), In the formula, "-" represents a peptide bond or a linking peptide. Z1 is either a nuclear localization signal sequence or none; Z2 is a Cas protein; Z3 is a connection sequence; Z4 is a nuclear localization signal sequence or none; Z5 is a reverse transcriptase; Z6 is either a nuclear positioning signal sequence or none.
[0011] In another preferred embodiment, the nuclear localization signal sequence has an amino acid sequence as shown in SEQ ID NO:2.
[0012] In another preferred embodiment, the linker sequence has an amino acid sequence as shown in SEQ ID NO:3.
[0013] In another preferred embodiment, the amino acid sequence of the reverse transcriptase is selected from the group consisting of: (a1) Has the amino acid sequence shown in SEQ ID NO:1; or (a2) An amino acid sequence having ≥80%, ≥85%, ≥90%, ≥91%, ≥92%, ≥93%, ≥94%, ≥95%, ≥96%, ≥97%, ≥98%, ≥99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:1.
[0014] In another preferred embodiment, the Cas protein includes: Cas9 or a variant thereof, Cas12 or a variant thereof, Cas13 or a variant thereof, Cas14 or a variant thereof, and TnpB.
[0015] In another preferred embodiment, the Cas protein is Cas9 or a variant thereof.
[0016] In another preferred embodiment, the amino acid sequence of the Cas protein is selected from the group consisting of: (b1) Having the amino acid sequence shown in SEQ ID NO:4; or (b2) An amino acid sequence having ≥80%, ≥85%, ≥90%, ≥91%, ≥92%, ≥93%, ≥94%, ≥95%, ≥96%, ≥97%, ≥98%, ≥99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:4.
[0017] In another preferred embodiment, the amino acid sequence of the fusion protein is selected from the group consisting of: (c1) Has the amino acid sequence shown in SEQ ID NO:6; or (c2) An amino acid sequence having ≥80%, ≥85%, ≥90%, ≥91%, ≥92%, ≥93%, ≥94%, ≥95%, ≥96%, ≥97%, ≥98%, ≥99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:6.
[0018] In another preferred embodiment, the fusion protein further includes a tag sequence.
[0019] A third aspect of the present invention provides a CRISPR-Cas composition comprising: (1) Guide RNA, or one or more DNA molecules encoding said guide RNA; and (2) The fusion protein described in the second aspect of the present invention, or the nucleic acid molecule encoding the fusion protein described in the second aspect of the present invention.
[0020] In another preferred embodiment, the CRISPR-Cas composition is used to reverse transcribe base G into base C.
[0021] In a fourth aspect, the present invention provides a polynucleotide encoding the reverse transcriptase described in the first aspect of the present invention or the fusion protein described in the second aspect of the present invention.
[0022] In another preferred embodiment, the polynucleotide is a DNA molecule codon-optimized according to the codon preference of the host cell.
[0023] In another preferred embodiment, the host cell includes eukaryotic cells and prokaryotic cells.
[0024] In a fifth aspect, the present invention provides a carrier containing the polynucleotide described in the fourth aspect of the present invention.
[0025] A sixth aspect of the present invention provides a delivery system comprising the reverse transcriptase described in the first aspect of the present invention, the fusion protein described in the second aspect of the present invention, the CRISPR-Cas composition described in the third aspect of the present invention, the polynucleotide described in the fourth aspect of the present invention, or the vector described in the fifth aspect of the present invention, and a delivery medium.
[0026] In another preferred embodiment, the delivery medium includes nanoparticles, plasmids, exosomes, microbubbles, gene guns, or electroporation devices.
[0027] A seventh aspect of the present invention provides a host cell expressing the reverse transcriptase described in the first aspect of the present invention, the fusion protein described in the second aspect of the present invention, or the CRISPR-Cas composition described in the third aspect of the present invention, or having a genome integrated with the polynucleotide described in the fourth aspect of the present invention, or containing the vector described in the fifth aspect of the present invention, or the delivery system described in the sixth aspect of the present invention.
[0028] In an eighth aspect, the present invention provides a kit comprising the reverse transcriptase described in the first aspect of the present invention, the fusion protein described in the second aspect of the present invention, the CRISPR-Cas composition described in the third aspect of the present invention, or the host cell described in the seventh aspect of the present invention.
[0029] In another preferred embodiment, the components of the kit are in the same or different containers.
[0030] In another preferred embodiment, the kit includes one or more buffer solutions.
[0031] In another preferred embodiment, the kit also includes a label or instructions.
[0032] In another preferred embodiment, the kit is used for one or more of the following: gene or genome editing, plant breeding, targeting a target gene, cutting a target gene or a non-target gene, or single base substitution.
[0033] A ninth aspect of the present invention provides a method for editing a target nucleic acid, the method comprising contacting the target nucleic acid with the CRISPR-Cas composition described in the third aspect of the present invention.
[0034] The tenth aspect of the present invention provides the use of the fusion protein described in the first aspect of the present invention, the guide RNA described in the second aspect of the present invention, the CRISPR-Cas composition described in the third aspect of the present invention, the polynucleotide described in the fourth aspect of the present invention, the vector described in the fifth aspect of the present invention, the delivery system described in the sixth aspect of the present invention, the host cell described in the seventh aspect of the present invention, or the kit described in the eighth aspect of the present invention for guided editing and plant breeding.
[0035] It should be understood that, within the scope of this invention, the above-described technical features of this invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be described in detail here. Attached Figure Description
[0036] Figure 1 A schematic diagram of the bootloader editor structure is shown.
[0037] Figure 2 The novel reverse transcriptase XRT1 was shown to have activity in guided editing. Detailed Implementation
[0038] Through extensive and in-depth research, the inventors have obtained a novel reverse transcriptase, XRT1, for the first time. The amino acid sequence length of this XRT reverse transcriptase is significantly shorter than that of the commonly used MMLV reverse transcriptase. Experimental results demonstrate that the guided editing system constructed by combining this novel reverse transcriptase with Cas protein possesses precise base editing activity, converting base G to base C. This novel reverse transcriptase can be used for precise gene editing. Based on this, the present invention was completed.
[0039] It should be understood that the specific methods and experimental conditions of the invention described below in varying degrees of detail are intended to provide a substantive understanding of the invention. Definitions of certain terms used in this specification are provided below. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0040] the term As used herein, the terms “containing” or “including (comprise)” can be open-ended, semi-closed, or closed-ended. In other words, the terms also include “consistently made of” or “made of”.
[0041] As used herein, the term “and / or” refers to and covers any and all possible combinations of one or more of the related listed items.
[0042] As used in this article, the term "significant" means that, in a hypothesis test, the observed effect (such as the difference between the experimental and control groups) is unlikely to be caused solely by random error. A hypothesis test includes: the null hypothesis (H0), which assumes that the observed effect does not exist (such as no difference between the experimental and control groups); the p-value, which is the probability of observing the current or more extreme effect when H0 is true; and the significance threshold (α). The significance threshold is typically used to determine whether a hypothesis test is significant. Generally, the significance threshold is 0.05. If the p-value ≤ α, then H0 is rejected, meaning the observed effect exists, and the result is called "significant."
[0043] As used herein, a protein “fragment,” “variant,” or “homologous” may optionally be characterized as having at least 60%, preferably 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with a reference protein (such as a reference isoform). In some preferred embodiments, fragments, variants, isoforms, and homologs of a reference protein may be characterized by their ability to perform the functions performed by the reference protein.
[0044] Typically, protein derivatization does not adversely affect the protein's desired activity (e.g., activity binding to guide RNA, endonuclease activity, activity of binding to and cleaving a target sequence at a specific site guided by guide RNA); that is, the protein derivative has the same activity as the original protein. Modified forms of "derivatives" include the deletion, insertion, modification, and / or substitution of one or more amino acids in the protein. The term "artificial evolutionary modification" indicates artificial involvement.
[0045] Those skilled in the art will recognize that the structure of a protein can be altered without adversely affecting its activity and functionality, for example, by introducing one or more conserved amino acid substitutions into the protein's amino acid sequence without adversely affecting the protein molecule's activity and / or three-dimensional structure.
[0046] Those skilled in the art will recognize examples and implementations of conserved amino acid substitutions. Specifically, an amino acid residue can be substituted with another amino acid residue belonging to the same group as the site to be substituted; that is, a nonpolar amino acid residue can replace another nonpolar amino acid residue, a polar, uncharged amino acid residue can replace another polar, uncharged amino acid residue, a basic amino acid residue can replace another basic amino acid residue, and an acidic amino acid residue can replace another acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. Conservative substitutions, where an amino acid is replaced by another amino acid belonging to the same group, fall within the scope of this disclosure, provided that the substitution does not result in the inactivation of the protein's biological activity. Therefore, the proteins of this disclosure may contain one or more conserved substitutions in their amino acid sequences, preferably generated by substitutions according to Table 1. Additionally, this disclosure also covers proteins that also contain one or more other nonconservative substitutions, provided that such nonconservative substitutions do not significantly affect the desired function and biological activity of the proteins of this disclosure.
[0047] Conserved amino acid substitutions can occur at one or more predicted non-essential amino acid residues. “Non-essential” amino acid residues are those that can be altered (deleted, substituted, or replaced) without changing biological activity, while “essential” amino acid residues are required for biological activity. A “conserved amino acid substitution” is a substitution in which an amino acid residue is replaced by an amino acid residue with a similar side chain. Amino acid substitutions can occur in non-conserved regions of Cas enzymes. Generally, such substitutions are not performed on conserved amino acid residues, or on amino acid residues located within conserved motifs, where such residues are required for protein activity. However, those skilled in the art will understand that functional variants may have fewer conserved or non-conserved alterations in conserved regions.
[0048] In some implementations, the selected group of amino acids considered to be mutually conserved substitutions includes: Table 1 Those skilled in the art will recognize that one or more amino acid residues can be altered (replaced, deleted, truncated, or inserted) from the N and / or C ends of a protein while retaining its functional activity. Therefore, proteins that have one or more amino acid residues altered from the N and / or C ends of the fusion proteins of this invention, or fusion proteins including Cas proteins, while retaining their desired functional activity, are also within the scope of this disclosure. These alterations may include those introduced by modern molecular methods such as PCR, which includes PCR amplification that alters or lengthens the protein-coding sequence by means of oligonucleotides containing amino acid-coding sequences used in PCR amplification.
[0049] It should be recognized that proteins can be altered in various ways, including amino acid substitution, deletion, truncation, and insertion, and methods for such operations are generally known to those skilled in the art. For example, amino acid sequence variants of proteins can be prepared by mutating DNA. This can also be accomplished through other forms of mutagenesis and / or directed evolution, for example, by using known mutagenesis, recombination, and / or shuffling methods, combined with relevant screening methods, to perform one or more amino acid substitutions; or the deletion and / or insertion of one or more amino acids.
[0050] Those skilled in the art will understand that minor amino acid changes (e.g., naturally occurring mutations) or (e.g., using r-DNA technology) may occur in the fusion proteins of this application or fusion proteins including Cas proteins without loss of protein function or activity. If these mutations occur in the catalytic domain, active site, or other functional domains of the protein, the properties of the polypeptide may be altered, but the polypeptide may retain its activity. If the mutations are not located near the catalytic domain, active site, or other functional domains, a smaller impact can be expected.
[0051] Those skilled in the art can identify the essential amino acids of fusion proteins using methods known in the art, such as localized mutagenesis, protein evolution, or bioinformatics analysis. The catalytic domains, active sites, or other functional domains of a protein can also be determined through physical structural analysis, such as by techniques like nuclear magnetic resonance, crystallography, electron diffraction, or photoaffinity labeling, combined with mutations in presumed key site amino acids.
[0052] reverse transcriptase This invention discovers a series of novel reverse transcriptases, XRT, which are significantly smaller in size than existing MMLV reverse transcriptases. Furthermore, these novel reverse transcriptases possess the activity of converting base G to base C. Preferably, the reverse transcriptase of this invention refers to XRT1 reverse transcriptase.
[0053] Preferably, the reverse transcriptase is selected from the group consisting of: (1) It has the amino acid sequence shown in SEQ ID NO:1; (2) An amino acid sequence having ≥80%, ≥85%, ≥90%, ≥91%, ≥92%, ≥93%, ≥94%, ≥95%, ≥96%, ≥97%, ≥98%, ≥99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:1.
[0054] Fusion protein This invention provides a fusion protein comprising the reverse transcriptase and Cas protein of this invention. As used herein, the term "Cas protein" is used in the broadest sense, including wild-type Cas protein, its derivatives or variants, analogs, and functional fragments such as oligonucleotide-binding fragments.
[0055] The term “wild type” has the meaning commonly understood by those skilled in the art as referring to the typical form of an organism, strain, gene, or protein, or the characteristic that distinguishes it from mutant or variant forms when it exists in nature, which can be isolated from its natural source and has not been intentionally modified by humans.
[0056] Preferably, the Cas protein may include: Cas9 or its variants, Cas12 or its variants, Cas13 or its variants, Cas14 or its variants, TnpB, etc. Preferably, the Cas protein is Cas9 or its variants. Preferably, the Cas protein used in this invention is a type of Cas9, nCas9, whose amino acid sequence is shown in SEQ ID NO:4. A "variant" or "homolog" of a Cas protein refers to a polypeptide that substantially retains the function or activity of the Cas protein.
[0057] Preferably, the amino acid sequence of the Cas protein has one or more amino acid substitutions, deletions, or additions compared to the amino acid sequence shown in SEQ ID NO:4, and substantially retains the biological function of its source sequence; the substitutions, deletions, or additions of one or more amino acids include one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) amino acid substitutions, deletions, or additions.
[0058] Preferably, the amino acid sequence of the fusion protein has the amino acid sequence shown in SEQ ID NO:6; or an amino acid sequence having ≥80%, ≥85%, ≥90%, ≥91%, ≥92%, ≥93%, ≥94%, ≥95%, ≥96%, ≥97%, ≥98%, ≥99%, or 100% sequence identity with the amino acid sequence shown in SEQ ID NO:6.
[0059] Guide RNA As used herein, the terms “guide RNA,” “guide sequence,” and “sgRNA” are used interchangeably and have the meanings commonly understood by those skilled in the art. Generally, a guide RNA may comprise a direct repeat (DR) sequence and a spacer sequence, or consist substantially of or composed of a direct repeat sequence and a spacer sequence.
[0060] In some cases, a guide sequence is any polynucleotide sequence that is sufficiently complementary to the target sequence / gene / nucleotide to hybridize with said target sequence / gene / nucleotide and guide the CRISPR-Cas complex to specifically bind to said target sequence / gene / nucleotide. Generally, when optimally aligned, the complementarity between the guide sequence and its corresponding target sequence / gene / nucleotide is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%.
[0061] Preferably, more than 80% of the spacer sequence is complementary to the target nucleic acid. More preferably, more than 90%, more preferably more than 95%, even more preferably more than 99%, and still more preferably 100% complementary to the target sequence / target gene / target nucleic acid.
[0062] Preferably, the spacer sequence is about or at least about 16 nucleotides in length, for example, about 16-100, about 16-90, about 17-70, about 17-50, or about 18-41 consecutive nucleotides in length; optionally, the spacer sequence is about 20 consecutive nucleotides in length. Preferably, the spacer sequence is about 20 nt in length.
[0063] carrier As used in this article, a "carrier" is a nucleic acid molecule that can transport another nucleic acid molecule linked to it.
[0064] Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules including one or more free ends, or without free ends (e.g., circular); nucleic acid molecules including DNA, RNA, or both; and a wide variety of other polynucleotides known in the art. Vectors can be introduced into host cells through transformation, transduction, or transfection, thereby enabling the expression of their carried genetic material elements in the host cells. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including proteins, fusion proteins, isolated nucleic acid molecules, etc., as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector may contain a variety of elements controlling expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. Vectors may also contain a replication initiation site.
[0065] Vectors include plasmids and viral vectors. A plasmid is a circular double-stranded DNA loop in which another DNA fragment can be inserted, for example, using standard molecular cloning techniques. A viral vector contains a virus-derived DNA or RNA sequence within a vector used to package the virus; viruses include, for example, retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses. Viral vectors also contain polynucleotides carried by a virus intended for transfection into a host cell. Some vectors (e.g., bacterial vectors with bacterial origins of replication and augmented mammalian vectors) are capable of autonomous replication in the host cells into which they are introduced.
[0066] Other vectors (e.g., non-attachment mammalian vectors) integrate into the host cell's genome after introduction and thereby replicate along with the host genome. Furthermore, some vectors can direct the expression of genes they are operatively linked to. Such vectors are called "expression vectors."
[0067] Those skilled in the art will understand that the design of expression vectors can depend on factors such as the selection of host cells to be transformed and the desired expression level.
[0068] Reagent test kit The present invention provides a kit comprising the reverse transcriptase, fusion protein, CRISPR-Cas composition, host cell, or any two or all of the components thereof.
[0069] Preferably, the kit of this disclosure further includes a label or instructions for use of one or more components contained therein, and / or a label or instructions for use in combination with one or more additional components that may be available elsewhere or are required.
[0070] Preferably, the kit further includes one or more buffer solutions that can be used to dissolve any of the one or more components contained therein, and / or to provide suitable reaction conditions for one or more of the components. Such buffer solutions may include one or more of the following: PBS, HEPES, Tris, MOPS, Na₂CO₃, NaHCO₃, NaB, or combinations thereof. In some embodiments, reaction conditions include an appropriate pH, such as an alkaline pH. In some embodiments, the pH is between 7 and 10.
[0071] Preferably, any one or more of the reagent kit components can be stored in a suitable container or at a suitable temperature, such as 4 degrees Celsius.
[0072] Target sequence, target gene or target nucleic acid As used herein, the terms “target sequence,” “target gene,” or “target nucleic acid” are used interchangeably and refer to a specific nucleic acid containing a nucleic acid sequence that is wholly or partially complementary to a spacer sequence in the guide RNA. A “target sequence” is a polynucleotide targeted by a spacer sequence in the guide RNA, such as a sequence complementary to that spacer sequence, wherein hybridization between the target sequence and the spacer sequence will promote the formation of a CRISPR-Cas complex (including the Cas protein and the guide RNA). Perfect complementarity between sequences is not required, as long as sufficient complementarity exists to induce hybridization and promote the formation of a CRISPR-Cas complex. Preferably, the target nucleic acid contains a non-coding region (e.g., a promoter or terminator). Preferably, the target nucleic acid is single-stranded or double-stranded.
[0073] Sequence identity As used herein, “sequence identity” refers to the percentage of nucleotide / amino acid residues in the main sequence that are identical to those in the reference sequence after aligning sequences and, where necessary, introducing gaps to achieve the maximum percentage of sequence identity between sequences. To determine the percentage of sequence identity between two or more nucleic acid or amino acid sequences, those skilled in the art are familiar with various methods for performing double and multiple sequence alignments, for example, using publicly available computer software such as ClustalOmega (Söding, J. Bioinformatics (2005) 21, 951-960), T-coffee (Notredame et al. J. Mol. Biol. (2000) 302, 205-217), Kalign (Lassmann and Sonnhammer 2005, BMC Bioinformatics, 6(298)), and MAFFT (Katoh and Standley, Molecular Biology and Evolution (2013) 30(4) 772–780). When using this software, use the default parameters, preferably using features such as open space penalty and extended penalty.
[0074] The main advantages of this invention include: (1) The present invention has discovered a novel reverse transcriptase XRT1, which is much smaller in size than the existing MMLV reverse transcriptase.
[0075] (2) The guided editing system based on XRT1 can be used for precise base editing, and then can be used in fields such as gene-based theoretical research, plant breeding, and gene therapy.
[0076] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Experimental methods in the following embodiments, unless otherwise specified, are generally performed under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or as recommended by the manufacturer. Unless otherwise stated, percentages and parts are weight percentages and parts by weight.
[0077] Example 1: Discovering more reverse transcriptases active in guided editing systems To obtain more active reverse transcriptases, 17 potentially active reverse transcriptases were first identified using bioinformatics methods and named XRT. These XRT reverse transcriptases all have an amino acid length of less than 460 aa, significantly shorter than the existing MMLV reverse transcriptase (the full-length version of MMLV is 692 aa, and the truncated version is 515 aa).
[0078] Based on the commonly used bootloader editor PE2, the original engineered MMLV-RT was replaced with a newly discovered reverse transcriptase. Figure 1 ).
[0079] Example 2: Study on the editing activity of XRT reverse transcriptase The editing activity of XRT reverse transcriptase was verified by transient transformation in 293T cells, with PE2 cells carrying MMLV-RT serving as a positive control. The target editing design was to set a target site on the endogenous gene FANCF and introduce precise G-to-C editing. Sequencing results after transformation showed that XRT1 had editing activity, with an editing efficiency of 7.42% (…). Figure 2 In addition, XRT1 has an amino acid length of 451 aa, which is significantly shorter than the existing MMLV reverse transcriptase (the full-length version is 692 aa and the truncated version is 515 aa).
[0080] sequence All documents mentioned in this invention are incorporated herein by reference as if each document were individually incorporated by reference. Furthermore, it should be understood that after reading the foregoing teachings of this invention, those skilled in the art can make various alterations or modifications to this invention, and these equivalent forms also fall within the scope defined by the appended claims.
Claims
1. A reverse transcriptase, characterized in that, The amino acid sequence of the reverse transcriptase is shown in SEQ ID NO:
1.
2. A fusion protein, characterized in that, The fusion protein includes the reverse transcriptase and Cas protein as described in claim 1.
3. A CRISPR-Cas composition, characterized in that, The CRISPR-Cas composition comprises: (1) Guide RNA, or one or more DNA molecules encoding said guide RNA; and (2) The fusion protein of claim 2, or the nucleic acid molecule encoding the fusion protein of claim 2.
4. A polynucleotide, characterized in that, The polynucleotide encodes the reverse transcriptase of claim 1 or the fusion protein of claim 2.
5. A carrier, characterized in that, The carrier contains the polynucleotide as described in claim 4.
6. A delivery system, characterized in that, The delivery system includes the reverse transcriptase of claim 1, the fusion protein of claim 2, the CRISPR-Cas composition of claim 3, the polynucleotide of claim 4, or the vector of claim 5, and the delivery medium.
7. A host cell, characterized in that, The host cell expresses the reverse transcriptase of claim 1, the fusion protein of claim 2, or the CRISPR-Cas composition of claim 3, or has the polynucleotide of claim 4 integrated into its genome, or contains the vector of claim 5, or the delivery system of claim 6.
8. A reagent kit, characterized in that, The kit comprises the reverse transcriptase of claim 1, the fusion protein of claim 2, the CRISPR-Cas composition of claim 3, or the host cell of claim 7.
9. A method for editing target nucleic acids, characterized in that, The method includes contacting the target nucleic acid with the CRISPR-Cas composition of claim 3.
10. Use of the reverse transcriptase of claim 1, the fusion protein of claim 2, the CRISPR-Cas composition of claim 3, the polynucleotide of claim 4, the vector of claim 5, the delivery system of claim 6, the host cell of claim 7, or the kit of claim 8, characterized in that, Used to guide editing.
Citation Information
Patent Citations
Pilot editing system based on PERV reverse transcriptase
CN118995664A
Reverse prime editing system
WO2026045988A1