Nucleotide sequence editing method and fusion protein with editing effect
By developing a fusion protein that combines nuclease and helicase, the problems of high operational complexity and low editing efficiency in existing gene editing technologies are solved, and efficient and accurate nucleotide sequence editing is achieved, which is suitable for genetic improvement and molecular biology research in a variety of biological systems.
Patent Information
- Application Number
- CN202511022781.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-08-22
AI Technical Summary
The existing gene editing technology has problems such as high operational complexity, high time cost, immunogenicity risk and low editing efficiency. Especially in the CRISPR/Cas9 system, it is difficult to achieve efficient and accurate base replacement, insertion and deletion.
Develop a fusion protein that combines nucleases and helicases with the functions of scanning ssDNA single strands and identifying, aligning, annealing, and annealing 3’DNA microhomologous ends to achieve efficient editing of target nucleotides by integrating nuclease activity and guiding RNA specific localization. The fusion protein is composed of the helicase domain of DNA polymerase θ or DNA helicase HELQ and nuclease, which can efficiently perform base replacement, deletion or insertion in the cell.
It significantly improves the accuracy and feasibility of gene editing, reduces operational complexity and time cost, improves editing efficiency, has the advantages of safety, reliability and convenience of use, and is suitable for genetic improvement and molecular biology research of a variety of biological systems.
Smart Images

Figure CN120518784A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of genetic engineering technology, and in particular to a nucleotide sequence editing method and a fusion protein with editing function. Background Art
[0002] Gene editing refers to the efficient and precise insertion, deletion, or substitution of bases at specific locations in the genome through gene editing techniques. Modern gene editing techniques primarily target DNA double-strand breaks (DSBs), further repairing DSBs through homologous recombination (HDR) and nonhomologous end joining (NHEJ) to achieve gene editing. Currently, the CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) / Cas9 (CRISPR-associated protein 9) system has rapidly developed and become a leading tool for targeted genome editing worldwide. To mitigate off-target effects caused by Cas9 nuclease activity, Cas9 nickases were developed. These nickases are generated by mutating one of the two nuclease active regions, RuvC1 and HNH. This mutation creates a single-strand nick rather than a double-strand break at the target site.
[0003] Prime editor (PE) technology is a versatile and precise gene editing technique capable of achieving arbitrary base substitutions, small insertions, and deletions. Existing PE systems include traditional RNA reverse transcriptases and emerging DNA polymerase-mediated approaches. The original PE1 utilizes a reverse transcriptase (RT) fused to nCas9 (H840A) and a first-generation editing guide RNA (pegRNA) to directly extend the desired genetic information from the pegRNA to the target genomic site. PE2 utilizes an engineered M-MLV reverse transcriptase to further improve editing efficiency. The DNA polymerase-mediated PE system utilizes the DNA polymerase phi29 from the Bacillus subtilis phage, using the MS2 viral coat protein MCP to connect the polymerase to the repair template, and Ecklenow DNA polymerase I from Escherichia coli, using the HuH endonuclease to bind to ssDNA containing a specific recognition sequence, further enabling precise editing. Current PE systems rely on genetically engineered viruses or bacteria, which carries potential risks such as immunogenicity.
[0004] Based on this, the present invention intends to develop a new nucleotide sequence editing method, hoping to use nuclease and helicase with the function of scanning ssDNA single strands and identifying, aligning and annealing 3'DNA microhomologous ends to form a fusion protein, so as to achieve accurate, efficient and safe editing of nucleotide sequences. Summary of the Invention
[0005] The purpose of the present invention is to provide a nucleotide sequence editing method and a fusion protein with editing function to address the problems existing in the above-mentioned prior art. By integrating nuclease activity and the functions of scanning single-stranded ssDNA and recognizing, aligning, and annealing 3' DNA microhomologous ends, the fusion protein can efficiently complete base substitution, deletion, or insertion of target nucleotides under the specific positioning of the guide RNA, thereby significantly improving the accuracy and feasibility of gene editing, greatly reducing the complexity and time cost of the operation, and having outstanding advantages such as safety, reliability, high editing efficiency, and ease of use.
[0006] To achieve the above object, the present invention provides the following solutions:
[0007] The present invention provides a fusion protein with nucleotide sequence editing function, wherein the fusion protein is obtained by connecting and combining a protein with the functions of scanning ssDNA single strands and recognizing, aligning, and annealing 3'DNA microhomologous ends with a nuclease;
[0008] The protein having the functions of scanning ssDNA single strands and identifying, aligning and annealing 3'DNA micro-homologous ends is a protein containing a helicase domain.
[0009] Furthermore, the protein containing the helicase domain includes at least one of (1)-(2);
[0010] (1) The helicase domain of DNA polymerase θ;
[0011] (2) DNA helicase HELQ.
[0012] Furthermore, the fusion protein also includes DNA polymerase.
[0013] Furthermore, the DNA polymerase is POLA2, POLA4, POLD2, POLD3, POLD4, POLE3 or POLB.
[0014] Furthermore, the nuclease is Cas9 or a nucleotide editing enzyme with similar cutting function to Cas9.
[0015] In some preferred embodiments, the nucleotide editing enzyme is at least one of nCas9, ZFN, TALEN, cjCas9, SaCas9, NmeCas9, Cas12a, Cas12b, Cas12j, Cas13, Cas14, IscB, and TnpB. In some more preferred embodiments, the nucleotide editing enzyme is nCas9. In some further preferred embodiments, the nCas9 is nCas9 containing the H840A mutation.
[0016] Furthermore, the nucleotide sequence of the helicase domain of the DNA polymerase θ is shown in SEQ ID NO.1.
[0017] Furthermore, the nuclease and the protein containing the helicase domain are connected and combined via a connecting peptide. This approach allows the fusion protein to maintain its conformational stability, prevents cross-domain interactions, and allows their respective functions to proceed normally.
[0018] Furthermore, the nuclease is located at the N-terminal side or the C-terminal side of the protein containing the helicase domain.
[0019] The present invention also provides a gene encoding the above fusion protein.
[0020] The present invention also provides a biological material comprising the above-mentioned encoding gene, wherein the biological material is any one of (a) to (b);
[0021] (a) an expression vector comprising the encoding gene;
[0022] (b) A cell comprising the expression vector.
[0023] The present invention also provides a kit for gene editing, comprising the above-mentioned fusion protein or biological material.
[0024] In some specific embodiments, those skilled in the art may also add necessary additives to the kit. These additives may include buffers (such as Tris-HCl buffer), ions (such as potassium ions), cosolvents or protective agents (such as glycerol or dimethyl sulfoxide), protein stabilizers (such as bovine serum albumin, gelatin), and symmetric compounds (such as polyethylene glycol). The addition of these additives can help maintain the stability and activity of the Cas9 system and improve its efficiency during experimental operations.
[0025] In a specific implementation, the components involved in the kit of the present invention can be independently delivered to the cell in the form of ribonucleoprotein, mRNA or plasmid. That is, the complex consisting of the fusion protein and the guide RNA can be delivered directly into the cell, thereby achieving rapid and effective genome editing; in addition, the mRNA of the fusion protein can also be delivered into the cell, which will be translated into the fusion protein for editing; in addition, the fusion protein encoding plasmid can also be delivered into the cell so that the fusion protein is expressed in the cell for genome editing. The above-mentioned delivery form is well known in the art, and those skilled in the art can set the specific form of the components involved in the kit in combination with actual needs, so that each component can be delivered into the cell by the same or different delivery methods, and the target nucleotide is contacted with the nuclease, guide RNA, ssDNA, and an optional protein containing a helicase domain.
[0026] The present invention also provides the use of the above-mentioned fusion protein, encoding gene or biological material in nucleotide sequence editing.
[0027] The present invention also provides a nucleotide sequence editing method, comprising the steps of contacting a target nucleotide with a nuclease, a guide RNA, and a DNA repair template in a cell, and performing nucleotide sequence editing under the action of a protein having the function of scanning ssDNA single strands and recognizing, aligning, and annealing 3' DNA microhomologous ends.
[0028] Furthermore, the nuclease and the protein having the functions of scanning ssDNA single strands and recognizing, aligning, and annealing 3' DNA microhomologous ends participate in nucleotide sequence editing in the form of the above-mentioned fusion protein.
[0029] Furthermore, the nucleotide sequence editing is also performed with the participation of a DNA polymerase, which is derived from eukaryotic cells, preferably mammalian cells, and more preferably human cells.
[0030] In some embodiments, the guide RNA is sgRNA.
[0031] In some embodiments, the cell is a eukaryotic cell, preferably a mammalian cell, more preferably a human cell.
[0032] In some embodiments, the editing comprises a substitution, deletion, or insertion of a base.
[0033] The present invention discloses the following technical effects:
[0034] Through in-depth research, the present invention discovered that proteins containing helicase domains (such as the helicase domain of DNA polymerase θ or DNA helicase HELQ) have the function of scanning ssDNA single strands and recognizing, aligning, and annealing 3'DNA microhomologous ends. After fusing them with nucleases, they can edit the nucleotide sequence of the target site. This mechanism is based on the full contact of the nuclease, target nucleotide, guide RNA, and DNA repair template in the cell, and under the action of proteins with the function of scanning ssDNA single strands and recognizing, aligning, and annealing 3'DNA microhomologous ends, accurate and efficient editing of the nucleotide sequence is achieved, thereby significantly improving the overall efficiency of gene editing.
[0035] Based on the above findings, the present invention further developed a fusion protein with editing function and a matching gene editing product. The fusion protein integrates nuclease activity with scanning ssDNA single strands and the functions of identifying, aligning, and annealing 3'DNA micro-homologous ends. Under the specific positioning of the guide RNA, it can efficiently complete the base replacement, deletion, or insertion of the target nucleotide. The technical solution of the present invention not only significantly improves the accuracy and feasibility of gene editing, but also greatly reduces the operational complexity and time cost. It has outstanding advantages such as safety, reliability, high editing efficiency, and ease of use. It provides new tools and technical support for the field of genetic engineering, is suitable for genetic improvement and molecular biology research in a variety of biological systems, and has broad application prospects and promotion value. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0037] Figure 1 Statistical graph showing the gene editing efficiency after transfection of nCas9-HD fusion protein expression vector, different endogenous site editing (HEXA+5GtoC, HEXA+5Cins, FANCF del6G) repair template ssDNA with 5'loop and corresponding sgRNA expression vectors in HEK293T cells;
[0038] Figure 2 This is a statistical graph of the gene editing efficiency after transfection of the nCas9-HD fusion protein expression vector, the endogenous site editing B3GAT3+2TCins-5'loop repair template, and the corresponding sgRNA expression vector in porcine kidney cells PK15.
[0039] Figure 3This is a statistical graph showing the gene editing efficiency after transfection of the nCas9-HELQ fusion protein expression vector, the sgRNA expression vector for repairing the mCherry fluorescent protein, and the repair template mCherry-stop-correction-5'loop into mCherry-stop Reporter cells.
[0040] Figure 4 Statistical graph showing the gene editing efficiency after transfection of HEK293T cells with nCas9-HD-POLA2 fusion protein expression vector, different endogenous site editing (HEXA+5GtoC, HEXA+5Cins, FANCF del6G) repair templates ssDNA with 5'loop, and corresponding sgRNA expression vectors;
[0041] Figure 5 Statistical graph showing the gene editing efficiency after transfection of HEK293T cells with nCas9-HD-POLA4 fusion protein expression vector, different endogenous site editing (HEXA+5GtoC, HEXA+5Cins, FANCF del6G) repair templates ssDNA with 5'loop, and corresponding sgRNA expression vectors;
[0042] Figure 6 Statistical graph showing the gene editing efficiency after transfection of HEK293T cells with nCas9-HD-POLD2 fusion protein expression vector, different endogenous site editing (HEXA+5GtoC, HEXA+5Cins, FANCF del6G) repair template ssDNA with 5'loop, and corresponding sgRNA expression vectors;
[0043] Figure 7 Statistical graph showing the gene editing efficiency after transfection of HEK293T cells with nCas9-HD-POLD3 fusion protein expression vector, different endogenous site editing (HEXA+5GtoC, HEXA+5Cins, FANCF del6G) repair template ssDNA with 5'loop, and corresponding sgRNA expression vectors;
[0044] Figure 8 Statistical graph showing the gene editing efficiency after transfection of HEK293T cells with nCas9-HD-POLD4 fusion protein expression vector, different endogenous site editing (HEXA+5GtoC, HEXA+5Cins, FANCF del6G) repair template ssDNA with 5'loop, and corresponding sgRNA expression vectors;
[0045] Figure 9Statistical graph showing the gene editing efficiency after transfection of HEK293T cells with nCas9-HD-POLE3 fusion protein expression vector, different endogenous site editing (HEXA+5GtoC, HEXA+5Cins, FANCF del6G) repair template ssDNA with 5'loop, and corresponding sgRNA expression vectors;
[0046] Figure 10 This is a statistical graph showing the gene editing efficiency after K562 cells were transfected with an nCas9-HD-POLA2 fusion protein expression vector, different endogenous site editing (HEXA+5Cins, FANCF del6G) repair templates ssDNA with 5'loop, and the corresponding sgRNA expression vectors.
[0047] Figure 11 Statistical graph showing the gene editing efficiency after transfection of nCas9-HD-POLA4 fusion protein expression vector, different endogenous site editing (RUNX1+5GtoT, FANCF del6G) repair template ssDNA with 5'loop and corresponding sgRNA expression vectors in K562 cells;
[0048] Figure 12 This is a statistical graph showing the gene editing efficiency after transfection of nCas9-HD-POLD3 fusion protein expression vector, different endogenous site editing (HEXA+1AtoG, RUNX1+5GtoT) repair template ssDNA with 5'loop, and corresponding sgRNA expression vectors in K562 cells;
[0049] Figure 13 This is a statistical graph showing the gene editing efficiency after transfection of nCas9-HD-POLD4 fusion protein expression vector, different endogenous site editing (FANCF del6G, RUNX1+5GtoT) repair template ssDNA with 5'loop and corresponding sgRNA expression vectors in K562 cells;
[0050] Figure 14 Statistical graph showing the gene editing efficiency after transfection of nCas9-HD-POLB fusion protein expression vector, different endogenous site editing (HEXA+1AtoG, RUNX1+5GtoT) repair template ssDNA with 5'loop, and corresponding sgRNA expression vectors in K562 cells;
[0051] Figure 15Statistical graph showing the gene editing efficiency after transfection of nCas9-HD-POLA2 fusion protein expression vector, different endogenous site editing (RUNX1+5GtoT, FANCF del6G) repair template ssDNA with 5'loop and corresponding sgRNA expression vectors in A549 cells;
[0052] Figure 16 Statistical graph showing the gene editing efficiency after transfection of nCas9-HD-POLA4 fusion protein expression vector, different endogenous site editing (HEXA+5Cins, RUNX1+5GtoT) repair template ssDNA with 5'loop, and corresponding sgRNA expression vectors in A549 cells;
[0053] Figure 17 This is a statistical graph showing the gene editing efficiency after transfection of nCas9-HD-POLD2 fusion protein expression vector, different endogenous site editing (HEXA+5Cins, HEXA+1AtoG) repair template ssDNA with 5'loop, and corresponding sgRNA expression vectors in A549 cells;
[0054] Figure 18 This is a statistical graph showing the gene editing efficiency after transfection of nCas9-HD-POLD3 fusion protein expression vector, different endogenous site editing (HEXA+5Cins, RUNX1+5GtoT) repair template ssDNA with 5'loop and corresponding sgRNA expression vectors in A549 cells;
[0055] Figure 19 This is a statistical graph showing the gene editing efficiency after transfection of A549 cells with nCas9-HD-POLD4 fusion protein expression vector, different endogenous site editing (HEXA+5Cins, HEXA+1AtoG) repair templates ssDNA with 5'loop, and corresponding sgRNA expression vectors;
[0056] Figure 20 Statistical graph of gene editing efficiency after transfection of nCas9-HD-POLB fusion protein expression vector, different endogenous site editing (FANCF del6G, HEXA+1AtoG) repair template ssDNA with 5'loop and corresponding sgRNA expression vector in A549 cells. DETAILED DESCRIPTION
[0057] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as limiting the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.
[0058] It should be understood that the terms described herein are intended only to describe particular embodiments and are not intended to limit the present invention. In addition, for numerical ranges herein, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. The intermediate value within any stated value or stated range, and each smaller range between any other stated value or intermediate value within the stated range, is also encompassed within the present invention. The upper and lower limits of these smaller ranges may be independently included or excluded within the scope.
[0059] Unless otherwise indicated, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. Although only preferred methods and materials are described herein, any methods and materials similar or equivalent to those described herein may also be used in the practice or testing of the present invention. All documents mentioned in this specification are incorporated by reference to disclose and describe the methods and / or materials associated with the documents. In the event of any conflict with any incorporated document, the contents of this specification shall prevail.
[0060] It will be apparent to those skilled in the art that various modifications and variations may be made to the specific embodiments described herein without departing from the scope or spirit of the invention. Other embodiments will be apparent to those skilled in the art from the description of the invention. The description and examples are intended to be exemplary only.
[0061] The words “include,” “including,” “have,” “contain,” etc. used in this document are open-ended terms, meaning including but not limited to.
[0062] In the present invention, the term "exogenous DNA polymerase POLQ helicase domain or DNA helicase HELQ; helicase fused to other DNA polymerases or variants thereof" refers to introduction into cells by external action methods such as genetic engineering.
[0063] As used herein, the term "POLQ" is an abbreviation for DNA polymerase theta, a member of the DNA polymerase A family. It comprises an N-terminal helicase-like domain, a central domain of unknown function, and a C-terminal polymerase domain homologous to DNA polymerase A. The term "exogenous DNA polymerase POLQ helicase domain" refers to the helicase domain of DNA polymerase theta introduced from an external source (e.g., a transfection vector or exogenous protein). DNA helicase HELQ is an important DNA repair enzyme and a member of the helicase family. It possesses ATP-dependent DNA helicase activity. Its core domain comprises the helicase domain and a C-terminal extension domain.
[0064] As used herein, the term "Cas9" refers to a nuclease in the CRISPR system, primarily comprising two key functional regions: 1. a targeting domain, which recognizes and binds to a target DNA sequence; and 2. a cleavage domain, which possesses nuclease activity and cleaves the recognized DNA sequence. Cas9 as defined herein includes Cas9 from various sources. In some embodiments, the Cas9 is selected from SpyCas9 (Streptococcus pyogenes Cas9), SaCas9 (Staphylococcus aureus Cas9), NmeCas9 (Neisseria meningitidis Cas9), FnCas9 (Francisella novicida Cas9), and CjCas9 (Campylobacter jejuni Cas9). These different Cas9 types have varying sizes and properties, such as higher specificity, more efficient DNA cleavage activity, or advantages in gene editing in specific cell types or tissues. The appropriate Cas9 type can be selected based on specific research needs. Cas9 as defined herein also includes Cas9 variants. In some embodiments, the Cas9 variants include ncas9 (nickase Cas9), Cas9-HF (high-fidelity Cas9), dCas9 (dead Cas9), etc.
[0065] In the present invention, the term "nickase" refers to a class of enzymes that participate in the cleavage of DNA or RNA molecules in organisms, also known as nucleases. Nickases can cleave the phosphodiester bonds of DNA or RNA molecules through specific enzymatic activity, thereby generating broken base chains.
[0066] In this invention, the term "nCas9" refers to "nickase Cas9," a Cas9 variant modified to only trigger single-strand cleavage on one DNA strand, rather than double-strand cleavage on both strands. This modification is intended to increase the precision of gene editing and reduce the possibility of unintended cleavage. "nCas9 (H840A)" herein refers to nCas9 containing the H840A mutation.
[0067] In the present invention, the term "guide RNA" refers to an RNA molecule used to guide the Cas9 protein to achieve targeted genome editing in the CRISPR / Cas9 system. Guide RNA is typically composed of a 20-base sequence, which contains a sequence that pairs with the target DNA sequence so that the Cas9 protein can bind to the target DNA and achieve specific editing. "sgRNA" refers to a single guide RNA (single guide RNA). sgRNA combines the previously separated CRISPR RNA (crRNA) and transcription initiation subsequence (tracrRNA) into a single RNA molecule. This simplifies experimental operations and increases the ease of use of the CRISPR / Cas9 system. sgRNA is typically composed of a guide sequence and a tracrRNA binding sequence. The guide sequence pairs with the DNA sequence of the target gene, thereby guiding the Cas9 protein to target the target gene and achieve genome editing.
[0068] In this application, the term "ssDNA with 5' loop" refers to single-stranded DNA with a GGGGTTTCCCC / CCCCTTTGGGG sequence at its 5' end, i.e., a looped structure at the 5' end of the ssDNA. In the field of gene editing, ssDNA with 5' loop can be used as a repair template to achieve precise genome editing in the CRISPR / Cas9 system. The ssDNA with 5' loop repair template mediates the repair of double-stranded DNA breaks caused by Cas9 cleavage in cells through homologous recombination, thereby repairing or modifying the target gene.
[0069] In the present invention, the term "recognition" refers to the process by which gene editing tools can accurately find the target DNA sequence; the term "alignment" means that when cells perform DNA repair, the repair template needs to be highly consistent with the DNA sequence at the break in order to insert new genetic information; the term "annealing" refers to the process by which the single-stranded DNA repair template and the DNA gap formed by nuclease cutting are complementary sequences, and the base pairing principle of DNA molecules is used to further form hydrogen bonds and ultimately form a double-stranded DNA structure.
[0070] Example 1: Endogenous precise editing in mammalian HEK293T and PK15 cells using nCas9 (H840A) protein, ssDNA with 5' loop, and the helicase domain of exogenous POLQ
[0071] Purpose: To demonstrate the effectiveness of a gene editing method that exposes a target nucleotide to nCas9, guide RNA, ssDNA, and the helicase domain of exogenous POLQ. This example demonstrates its application in cells for precise endogenous editing.
[0072] S1. Construction of a fusion expression vector of nCas9 (H840A) protein and the helicase domain of human DNA polymerase POLQ (named nCas9-HD fusion protein expression vector):
[0073] 1. The DNA polymerase sequence was synthesized by GenScript Biotech Co., Ltd.; using the synthesized POLQ sequence as a template, primers were designed to amplify the helicase domain of POLQ, designated L1. The nucleotide sequence is shown in SEQ ID NO. 1.
[0074] 2. Using CMV-PE2 as a template, corresponding primers were designed to amplify two fragments, designated L2 and L3, with nucleotide sequences shown in SEQ ID NO. 2 and SEQ ID NO. 3, respectively.
[0075] CMV-4380bp-F: CGCAAATGGGCGGTAGGCGTG (SEQ ID NO.4);
[0076] Linker-4380bp-R:GCTGCTGCCGCCGCTGCTGC (SEQ ID NO.5);
[0077] SV40-3320bp-F: CCCAAGAAGAAGAGGAAAGT (SEQ ID NO.6);
[0078] CMV-3320bp-R: ACGCCTACCGCCATTTGCG (SEQ ID NO. 7).
[0079] 3. Homologously recombine the three fragments (L1, L2, and L3) to generate an nCas9-HD fusion protein expression vector. The amino acid sequence of the nCas9-HD fusion protein is shown in SEQ ID NO. 8.
[0080] SEQ ID NO.8: DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGE TAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYP TIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQ YADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAG YIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNR EKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSL LYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNA SLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRK LINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTV KVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQ NGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLIT QRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKD FQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNF FKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIK LPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIIEQI SEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIH QSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSSGGSSGGSSSMNLLRRSGKRRRSESGSDSFSGSG GDSSASPQFLSGSVLSPPPGLGRCLKAAAAGECKPTVPDYEIDKLLLANWGLPKAVLEKYHSFGVKKMFEWQAECL LLGQVLEGKNLVYSAPTSAGKTLVAELLILKRVLEMRKKALFILPFVSVAKEKKYYLQSLFQEVGIKVDGYMGSTS PSRHFSSLDIAVCTIERANGLINRLIEENKMDLLGMVVVDELHMLGDSHRGYLLELLLTKICYITRKSASCQADLA SSLSNAVQIVGMSATLPNLELVASWLNAELYHTDFRPVPLLESVKVGNSIYDSSMKLVREFPMLQVKGDEDHVVS LCYETICDNHSVLLFCPSKKWCEKLADIIAREFYNLHHQAEGLVKPSECPPVILEQKELLEVMDQLRRLPSGLDSV LQKTVPWGVAFHHAGLTFEERDIIEGAFRQGLIRVLAATSTLSSGVNLPARRVIIRTPIFGGRPLDILTYKQMVGR AGRKGVDTVGESILICKNSEKSKGIALLQGSLKPVRSCLQRREGEEVTGSMIRAILEIIVGGVASTSQDMHTYAAC TFLAASMKEGKQGIQRNQESVQLGAIEACVMWLLENEFIQSTEASDGTEGKVYHPTHLGSATLSSSLSPADTLDIF ADLQRAMKGFVLENDLHILYLVTPMFEDWTTIDWYRFFCCLWEKLPTSMKRVAELVGVEEGFLARCVKGKVVARTER QHRQMAIHKRFFTSLVLLDLISEVPLREINQKYGCNRGQIQSLQQSAAVYAGMITVFSNRLGWHNMELLLSQFQKR LTFGIQRELCDLVRVSLLNAQRARVLYASGFHTVADLARANIVEVEVILKNAVPFKSARKAVDEEEEAVEERRNMR TIWVTGRKGLTEREAAALIVEEARMILQQDLVEMGVQWN ; Among them, the wavy line indicates the nCas9 protein, the single underline indicates the connecting peptide, and the double underline indicates the helicase domain protein of the DNA polymerase POLQ.
[0081] S2. Design and construct sgRNA expression vector targeting endogenous editing sites:
[0082] The sgRNA was designed as follows:
[0083] HEXA-sgRNA: GAACGTGTTCCACTGGCATC (SEQ ID NO.9);
[0084] FANCF-sgRNA: GGAATCCCTTCTGCAGCACC (SEQ ID NO. 10);
[0085] B3GAT3-sgRNA: GGACCGCCAGATGTGTGAAG (SEQ ID NO. 11).
[0086] Design and synthesize ssDNA with 5'loop for precise editing of endogenous sites:
[0087] HEXA+5GtoC-5'loop:GGGGTTTCCCCAGGAAGGATCATCTACGAGATGCCAGTGGAACACGTT (SEQ ID NO. 12);
[0088] HEXA+5Cins-5'loop:GGGGTTTCCCCAGGAAGGATCATCTACGCAGATGCCAGTGGAACACGTT (SEQ ID NO. 13);
[0089] FANCF del6G-5'loop:GGGGTTTCCCCGGAAAAGCGATGCAGGTGCTGCAGAAGGGAT (SEQ IDNO.14);
[0090] B3GAT3+2TCins-5'loop: GGGGTTTCCCCTGCCTCTGGCCTCCTCTGATCACACATCTGGCG (SEQ ID NO. 15).
[0091] S3. HEK293T cells were transfected with an nCas9-HD fusion protein expression vector, along with ssDNA templates containing 5'loops for different endogenous site editing (HEXA+5GtoC, HEXA+5Cins, FANCF del6G), and corresponding sgRNA expression vectors. Porcine kidney PK15 cells were transfected with an nCas9-HD fusion protein expression vector, along with a B3GAT3+2TCins-5'loop repair template for endogenous site editing, and corresponding sgRNA expression vectors. Transfection conditions were performed according to the jetPRIME transfection reagent manufacturer's instructions. Cells were cultured at 37°C, 5% CO2, and maintained in DMEM supplemented with 10% fetal bovine serum (FBS). Twenty-four hours after transfection, positive cells were cultured for another four days in complete medium supplemented with 2 μg / mL puromycin, and harvested.
[0092] S4. The positive cell population obtained by lysis is used as a template for PCR amplification. The targeted editing region is amplified, purified, and sequenced using primers with sequencing barcodes.
[0093] S5. Compare and analyze the amplicon sequencing file with the reference sequence.
[0094] like Figure 1As shown in the figure, in HEK293T cells, the nCas9-HD fusion protein and the editing template ssDNA with a 5' loop can achieve precise editing of different endogenous sites. The precise editing efficiency is the ratio of the precise target base sequence to the total base sequence at the site after precise repair. Among them, at the HEXA+5GtoC site, the precise editing efficiency of nCas9 when alone contacting sgRNA and ssDNAwith 5'loop single-stranded template was 5.29%. When nCas9 was fused with the helicase domain, the precise editing efficiency was 42.41%, an increase of 7.02 times; at the FANCF del6G site, the precise editing efficiency of nCas9 when alone contacting sgRNA and ssDNA with 5'loop single-stranded template was 0.80%. When nCas9 was fused with the helicase domain, the precise editing efficiency was 4.83%, an increase of 5.04 times; at the HEXA+5Cins site, the precise editing efficiency of nCas9 when alone contacting sgRNA and ssDNA with 5'loop single-stranded template was 5.47%. When nCas9 was fused with the helicase domain, the precise editing efficiency was 28.80%, an increase of 4.27 times.
[0095] like Figure 2 As shown, in pig kidney PK15 cells, the editing efficiency of the B3GAT3+2TCins site was 0.22% when nCas9 alone contacted sgRNA and ssDNA with 5'loop single-stranded template. When nCas9 was fused with the helicase domain, the editing efficiency was 2.08%, an increase of 8.45 times.
[0096] In summary, in mammalian cells, the fusion of nuclease domains to helicases can significantly improve the efficiency of precise editing of endogenous sites.
[0097] Example 2: Precise editing of nucleotide sequences using nCas9 (H840A) protein, ssDNA with 5'loop, and exogenous DNA helicase HELQ
[0098] Objective: To demonstrate the effectiveness of a gene editing approach that exposes target nucleotides to nCas9, guide RNA, ssDNA, and exogenous DNA helicase HELQ using an mCherry-stop Reporter cell line, enabling precise nucleotide editing.
[0099] S1. Construction of nCas9 (H840A) protein and human DNA helicase HELQ fusion expression vector (named nCas9-HELQ fusion protein expression vector):
[0100] 1. The DNA helicase HELQ sequence was amplified using the pEnCMV-HELQ(human)-3×FLAG-SV40-Neo vector (purchased from Miaoling Biotechnology) as a template and is denoted as H1. Its nucleotide sequence is shown in SEQ ID NO. 16.
[0101] 2. Amplify the L2 and L3 fragments according to the method of Example 1.
[0102] 3. Homologously recombine the three fragments (H1, L2, and L3) to generate an nCas9-HELQ fusion protein expression vector. The amino acid sequence of the nCas9-HELQ fusion protein is shown in SEQ ID NO. 17.
[0103] SEQ ID NO.17: DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSG ETAEATRLKRTARRRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKY PTIYHLRKKLVDSTDKADLRIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDA KAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLLAQIGD QYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYA GYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDN REKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHS LLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFN ASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSR KLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQT VKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYL QNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLI TQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLSKLVSDFRK DFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMN FFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLII KLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHCHYLDEIIEQ ISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLI HQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSMDECGSRIRRVSLPKRNRPSLG CIFGAPTAAELVPGDEGKEEEEMVAENRRRKTAGVLPVEVQPLLLSDSPECLVLGGGDTNPDLLRHMPTDRGVGDQ PNDSEVDMFGDYDSFTENSFIAQVDDLEQKYMQLPEHKKHATDFATENLCSESKNKLSITTIGNLTELQTDKHTE NQSGYEGVTIEPGADALLYDVPSSQAIYFENLQNSSNDLGDHSMKERDWKSSSNHTVNEELPHNCIEQPQQNDESSS KVRTSSDNMNRRKSIKDHLKNAMTGNAKAKQTPIFSRSKQLKDTLLSEEINVAKKTVESSSNDLGPFYSLPSKVRDLY AQFKGIEKLYEWQHTCLTLNSVQERKNLIYSLPTSSGGMTLVAEILMLQELLCCRKDVLMILPYVAIVQEKISGLSS FGIELGFFVEEYAGSKGRFPPTKRREKKSLYIATIEKGHSLVNSLIETGRIDSLGLVVVDELHMIGEGSRGATLEM TLAKILYTSKTTQIIGMSATLNNVEDLQKFLQAEYYTSQFRPVELKEYLKINDTIYEVDSKAENGMTFSRLLNYKY SDTLKKMDPDHLVALVTEVIPNYSCLVFCPSKKNCENVAEMICKFLSKEYLKHKEKEKCEVIKNLKNIGNGNLCPV LKRTIPFGVAYHHSGLTSDERKLLEEAYSTGVLCLFTCTSTLAAWGNLPARRVILRAPYVAKEFLKRNQYKQMIGR AGRAGIDTIGESILILQEKDKQQVLELITKPLENCYSHLVQEFTKGIQTLFLSLIGLKIATNLDDIYHFMNGTFFG VQQKVLLKEKSLWEITVESLRYLTEKGLLQKDTIYKSEEEVQYNFHITKLGRASFKGTIDLAYCDILYRDLKKGLE GLVLESLLHLIYLTTPYDLVSQCNPDWMIYFRQFSQLSPAEQNVAAILGVSESFIGKKASGQAIGKKVDKNVVNRL YLSFWLYTLLKETNIWTVSEKFNMPRGYIQNLLTGTASFSSCVLHFCEELEEFWVYRALLVELTKKLTYCVKAELI PLMEVTGVLEGRAKQLYSAGYKSLMHLANANPEVLVRTIDHLSRRQAKQIVSSAKMLLHEKAEALQEEVEELLRLP SDFPGAVASSTDKA ; Among them, the wavy line indicates the nCas9 protein, the single underline indicates the connecting peptide, and the double underline indicates the HELQ protein.
[0104] S2. Design and construct an sgRNA expression vector for repairing mCherry fluorescent protein:
[0105] sgRNA sequence: CACCTTCAGCTTGGCGGTCT (SEQ ID NO. 18).
[0106] mCherry-stop-correction-5'loop:GGGGTTTCCCCTACGAGGGCACTCAAACCGCCAAGCTGAAG (SEQ ID NO. 19).
[0107] S3. Gene editing was performed using the 293T-mCherry-stop-reporter cell line. A premature stop codon, TAG, exists in the mCherry protein coding region. When precise nucleotide editing is achieved, TAG becomes CAA, and the mCherry protein is expressed normally. Under the correct excitation light, the entire cell can be observed to glow red.
[0108] mCherry-stop Reporter cells were transfected with an nCas9-HELQ fusion protein expression vector, an sgRNA expression vector for repairing the mCherry fluorescent protein, and a repair template, mCherry-stop-correction-5'loop. Cell culture and transfection conditions were the same as in Example 1. The sgRNA expression vector contained a GFP sequence, meaning that when the plasmid was successfully transfected into the cells and expressed normally, the cells could be observed to emit green light under the correct excitation light, which served as an enrichment marker for flow cytometry. When the plasmid was transfected and expressed normally, the cells as a whole emitted green light. When the premature terminator of the mCherry-stop Reporter cells was precisely edited, the cells as a whole emitted green and red light. Otherwise, the cells did not emit any light.
[0109] Flow cytometry was used for detection. The total number of particles in the experiment was 10,000. The proportion of cells emitting red light in green light was selected as the standard for accurate editing. Figure 3 As shown, the repair efficiency of precise editing using the nCas9-HELQ fusion protein expression vector was 25.4%.
[0110] Example 3: Endogenous precise editing in mammalian HEK293T cells using nCas9 (H840A) protein, ssDNA with 5' loop, and the helicase domain of exogenous POLQ fused with other DNA polymerases
[0111] Objective: To demonstrate the effectiveness of a gene editing method that involves contacting the target nucleotide with nCas9, guide RNA, ssDNA with a 5' loop, and the helicase domain of exogenous POLQ with the remaining DNA polymerase. This example demonstrates its true application to endogenous precise editing in cells.
[0112] In this example, nCas9 (H840A) protein and the helicase domain of human DNA polymerase POLQ were fused to different DNA polymerases to construct an nCas9-HD-DNA polymerase fusion protein expression vector, and then the efficiency of endogenous gene editing in cells was tested. The following uses the fusion protein nCas9-HD-POLD4 constructed using DNA polymerase POLD4 as an example to illustrate the process of constructing the fusion protein expression vector and testing the gene editing efficiency:
[0113] S1. Construction of a fusion protein expression vector containing the nCas9 (H840A) protein and the helicase domain of the human DNA polymerase POLQ fused to the mammalian DNA polymerase POLD4 (named nCas9-HD-POLD4 fusion protein expression vector):
[0114] 1. DNA polymerase sequences were synthesized by GenScript Biotech Co., Ltd.
[0115] 2. Design primers to amplify the corresponding DNA polymerase POLD4 sequence, denoted as P1, whose nucleotide sequence is shown in SEQ ID NO. 20.
[0116] 3. Amplify the L1, L2, and L3 fragments according to the method of Example 1.
[0117] 4. Homologously recombine the four fragments (P1, L1, L2, and L3) to generate an nCas9-HD-POLD4 fusion protein expression vector. The amino acid sequence of the nCas9-HD-POLD4 fusion protein is shown in SEQ ID NO. 21.
[0118] SEQ ID NO.21: DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSG ETAEATRLKRTARRRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKY PTIYHLRKKLVDSTDKADLRIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDA KAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLLAQIGD QYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYA GYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDN REKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHS LLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFN ASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYALFDDKVMKQLKRRRYTGWGRLSR KLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQT VKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHPVENTQLQNEKLYLYYL QNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLI TQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRK DFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMN FFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLII KLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHCHYLDEIIEQ ISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLI HQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSSGGSSGGSSSMNLLRRSGKRRRSESGSDSFSGS GGDSSASPQFLSGSVLSPPPGLGRCLKAAAAGECKPTVPDYEIDKLLLANWGLPKAVLEKYHSFGVKKMFEWQAEC LLLGQVLEGKNLVYSAPTSAGKTLVAELLILKRVLEMRKKALFILPFVSVAKEKKYYLQSLFQEVGIKVDGYMGST SPSRHFSSLDIAVCTIERANGLINRLIEENKMDLLGMVVVDELHMLGDSHRGYLLELLLTKICYITRKSASCQADL ASSLSNAVQIVGMSATLPNLELVASWLNAELYHTDFRPPVLLESVKVGNSIYDSSMKLVREFPMLQVKGDEDHVV SLCYETICDNHSVLLFCPSKKWCEKLADIIAREFYNLHHQAEGLVKPSECPPVILEQKELLEVMDQLRRLPSGLDS VLQKTVPWGVAFHHAGLTFEERDIIEGAFRQGLIRVLAATSTLSSGVNLPARRVIIRTPIFGGRPLDILTYKQMVG RAGRKGVDTVGESILICKNSEKSKGIALLQGSLKPVRSCLQRREGEEVTGSMIRAILEIIVGGVASTSQDMHTYAA CTFLAASMKEGKQGIQRNQESVQLGAIEACVMWLLENEFIQSTEASDGTEGKVYHPTHLGSATLSSSLSPADTLDIFADLQRAMKGFVLENDLHILYLVTPMFEDWTTIDWYRFFCLWEKLPTSMKRVAELVGVEEGFLARCVKGKVVARTE RQHRQMAIHKRFFTSLVLLDLISEVPLREINQKYGCNRGQIQSLQQSAAVYAGMITVFSNRLGWHNMELLLSQFQK RLTFGIQRELCDLVRVSLLNAQRARVLYASGFHTVADLARANIVEVEVILKNAVPFKSARKAVDEEEEAVEERRNM RTIWVTGRKGLTEREAAALIVEEARMILQQDLVEMGVQWNSGGSSGGSSGSETPGTSESATPESSGGSSGGSSMGR KRLITDSYPVVKRREGPAGHSKGELAPELGEEPQPRDEEEAELELLRQFDLAWQYGPCTGITRLQRWCRAKQMGLE PPPEVWQVLKTHPGDPRFQCSLWHLYPL ; Among them, the single wavy line indicates the nCas9 (H840A) protein, the single underline indicates the connecting peptide, the double underline indicates the helicase domain protein of DNA polymerase POLQ, and the double wavy line indicates the POLD4 protein.
[0119] S2. HEK293T cells were transfected with an nCas9-HD-DNA polymerase fusion protein expression vector (using an nCas9 (H840A) protein expression vector as a control), ssDNA with a 5'-loop repair template for different endogenous site editing (HEXA+5GtoC, HEXA+5Cins, FANCFdel6G), and corresponding sgRNA expression vectors. The construction methods for the ssDNA with a 5'-loop and corresponding sgRNA expression vectors, transfection and cell culture conditions, screening conditions for positive cell populations, lysis and amplification, and analysis and comparison procedures were the same as in Example 1.
[0120] The results are as follows Figures 4 - 9 As shown, in HEK293T cells, the fusion protein constructed by fusing the helicase domain of nCas9 (H840A) protein and human DNA polymerase POLQ with mammalian cell DNA polymerase and the editing template ssDNA with 5'loop can perform precise editing of different types of endogenous sites. Figures 4 - 9 The authors demonstrated that the addition of the HD domain to several DNA polymerases (including POLA2, POLA4, POLD2, POLD3, POLD4, and POLE3) improved the editing efficiency of endogenous sites. For example, nCas9-POLA2 had a precise editing efficiency of 3.24% at HEXA+5GtoC, but the addition of the HD domain increased this efficiency to 11.05%. The precise editing efficiency at HEXA+5Cins was 3.08%, but the addition of the HD domain increased this efficiency to 9.11%. The precise editing efficiency at FANCF del6G was 0.94%, but the addition of the HD domain increased this efficiency to 1.14%. The addition of the HD domain to the polymerase enabled precise editing through contact with the nuclease-cleaving protein, sgRNA, and editing template.
[0121] Example 4: Endogenous precise editing can be achieved using the nCas9-HD-DNA polymerase-mediated PE editing system in different mammalian cells
[0122] Objective: To investigate whether nCas9-HD-POLQ fusion protein, in combination with sgRNA and ssDNA with 5'loop, can achieve precise endogenous editing in different mammalian cells.
[0123] Different nCas9-HD-DNA polymerase fusion protein expression vectors constructed in Example 3 were used to perform gene editing on K562 and A549 cells, respectively. The method was the same as that in Example 3, and new endogenous site-precisely edited ssDNA with 5'loop and corresponding sgRNA expression vectors were introduced:
[0124] sgRNA expression vector for endogenous editing sites:
[0125] RUNX1-sgRNA: GCATTTTCAGGAGGAAGCGA (SEQ ID NO. 22).
[0126] Precise editing of ssDNA with 5'loop at endogenous sites:
[0127] RUNX1+5GtoT-5'loop:CCCCTTTGGGGTGTCTGAAGCAATCGCTTCCTCCTGAAAAT (SEQ ID NO. 23).
[0128] HEXA+1AtoG-5'loop:CCCCTTTGGGGAGGAAGGATCATCTACCAGACCGCCAGTGGAACACGTTCAA (SEQ ID NO. 24).
[0129] The results showed that in K562 cells and A549 cells, the nCas9-HD-DNA polymerase fusion protein expression vector, the contact sgRNA expression vector and the corresponding repair template ssDNA with 5'loop can all perform precise editing of endogenous sites.
[0130] Figures 10 - 14 Results for endogenous site editing in K562 cells. For example, at the insertion editing HEXA+5Cins site, nCas9-POLA2 achieved a precise editing efficiency of 0.87%. After incorporating the HD domain, this efficiency increased to 1.09%. At the deletion editing FANCF del6G site, nCas9-POLA2 achieved a precise editing efficiency of 0.18%. After incorporating the HD domain, this efficiency increased to 0.48%.
[0131] Figures 15 - 20Results for endogenous site editing in A549 cells. For example, in single-base substitution editing of the RUNX1+5GtoT site, the precise editing efficiency of nCas9-POLA2 was 0.08%, but after fusion with the HD domain, the efficiency increased to 0.22%. In deletion editing of the FANCF del6G site, the editing efficiency of nCas9-POLA2 was 0.94%, but after fusion with the HD domain, the efficiency increased to 4.49%.
[0132] In summary, the HD fusion of multiple polymerase-mediated PE editing systems of the present invention can effectively perform various types of precise editing.
[0133] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A fusion protein with nucleotide sequence editing function, characterized in that: The fusion protein is obtained by combining a protein with the functions of scanning ssDNA single strands and recognizing, aligning, and annealing 3'DNA micro-homologous ends with a nuclease; The protein having the functions of scanning ssDNA single strands and identifying, aligning and annealing 3'DNA micro-homologous ends is a protein containing a helicase domain.
2. The fusion protein according to claim 1, characterized in that The protein containing the helicase domain includes at least one of (1)-(2); (1) The helicase domain of DNA polymerase θ; (2) DNA helicase HELQ.
3. The fusion protein according to claim 2, characterized in that The fusion protein also includes a DNA polymerase.
4. The fusion protein according to claim 3, characterized in that The DNA polymerase is POLA2, POLA4, POLD2, POLD3, POLD4, POLE3 or POLB.
5. The fusion protein according to claim 1, characterized in that The nuclease is Cas9 or a nucleotide editing enzyme with similar cleavage function to Cas9; and / or The nucleotide editing enzyme is at least one of nCas9, ZFN, TALEN, cjCas9, SaCas9, NmeCas9, Cas12a, Cas12b, Cas12j, Cas13, Cas14, IscB and TnpB.
6. The fusion protein according to claim 1, characterized in that The nuclease and the protein containing the helicase domain are connected and combined via a connecting peptide.
7. A gene encoding the fusion protein according to any one of claims 1 to 6.
8. A biological material comprising the encoding gene according to claim 7, characterized in that: The biological material is any one of (a)-(b); (a) an expression vector comprising the encoding gene; (b) A cell comprising the expression vector.
9. A kit for gene editing, characterized in that: The method comprises the fusion protein according to any one of claims 1 to 6 or the biomaterial according to claim 8.
10. Use of the fusion protein according to any one of claims 1 to 6, the encoding gene according to claim 7, or the biomaterial according to claim 8 in nucleotide sequence editing.
11. A method for editing a nucleotide sequence, characterized in that: The method includes contacting the target nucleotide with a nuclease, a guide RNA, and a DNA repair template in the cell, using the nuclease to cut a DNA gap, and performing nucleotide sequence editing under the action of a protein that has the function of scanning ssDNA single strands and recognizing, aligning, and annealing 3'DNA microhomologous ends.
12. The nucleotide sequence editing method according to claim 11, characterized in that The nuclease and the protein having the functions of scanning ssDNA single strands and identifying, aligning, and annealing 3'DNA microhomologous ends participate in nucleotide sequence editing in the form of the fusion protein according to any one of claims 1 to 6.
13. The nucleotide sequence editing method according to claim 11, characterized in that The nucleotide sequence editing is also performed with the participation of DNA polymerase.
14. The nucleotide sequence editing method according to claim 13, characterized in that: The DNA polymerase is POLA2, POLA4, POLD2, POLD3, POLD4, POLE3 or POLB.
Citation Information
Patent Citations
DNA polymerase mediated nucleotide sequence editing method and composition
CN118581153A
DNA polymerase mediated genome editing
CN118715317A
Compositions and methods for precise genome editing using retrons
WO2025010350A2