Targeting nucleases, targeting nuclease systems, polynucleotides, and methods of gene editing
By developing a novel targeted nuclease BaCas9 and a guide RNA system, the limitations of the CRISPR-Cas system in selecting gene editing targets have been overcome, enabling efficient gene editing of AT-rich regions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CROPEDIT BIOTECHNOLOGY INC
- Filing Date
- 2025-09-26
- Publication Date
- 2026-06-19
AI Technical Summary
Existing CRISPR-Cas systems require searching for proto-interval neighboring motifs (PAMs) within the target gene's nucleic acid sequence when selecting gene editing targets, which is insufficient to meet the need for targeting any location in the genome of an organism.
A novel targeted nuclease, BaCas9, with the PAM sequence NNAAW, is provided. It has higher cleavage efficiency in recognizing NNAAW, is suitable for editing AT-rich regions, and can guide the targeted nuclease system for genome editing via guide RNA.
It improves the efficiency and flexibility of gene editing, especially in regions rich in AT bases, and is applicable to genome editing of bacteria, fungi, insects, animals, and plant cells.
Smart Images

Figure CN122235115A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of molecular biology, and specifically relates to a targeted nuclease, a targeted nuclease system, a polynucleotide, and a gene editing method for cellular genomes. More specifically, this invention relates to a novel guide RNA / Cas9 endonuclease system for altering cellular genomes, including compositions and methods of use. Background Technology
[0002] Targeted nucleases are core components for nucleic acid detection and genome modification, and there are numerous literature reports on them. Based on the differences in the nucleases used in various reported genome editing technologies, they can be divided into the following four categories: (a) large-scale homing nucleases (HE), (b) zinc finger nucleases (ZFN), (c) transcription activator effector nucleases (TALEN), and (d) CRISPR-Cas systems.
[0003] The first to emerge was the zinc finger nuclease (ZFN). ZFN technology uses the zinc finger domain of zinc finger proteins to recognize specific DNA sequences and recombine them with the nuclease active domain of the type II restriction endonuclease FokI to construct artificial recombinant zinc finger nucleases. These recombinant zinc finger nucleases bind to specific genomic DNA sites through their zinc finger domains, and then the nuclease active domain from FokI can cleave the DNA near those sites.
[0004] Shortly after ZFN in 2012, transcription activator-like effector nucleases (TALENs) emerged. Based on a similar principle to ZFNs, they combine the DNA base-specific recognition domain of the TALE (transcription activator-like effector) protein with the cleavage domain of the FokI nuclease to create recombinant proteins (TALE nucleases, TALENs). Thus, TALEs have the ability to specifically recognize specific DNA sites in the genome, while the FokI nuclease's active domain can cleave DNA near that site. Unlike ZFNs, TALENs are easier to artificially modify in their ability to recognize and bind DNA base sequences, allowing for the design of recombinant nucleases that cleave different DNA sites in the genome.
[0005] By 2013, another technology capable of cutting genomic DNA emerged: CRISPR / Cas (clustered regularly interspaced short palindromic DNA).
[0006] CRISPR / Cas is an adaptive immune defense system found in bacteria and archaea, specifically designed to combat the exogenous DNA of bacteriophages invading bacteria. The CRISPR / Cas system integrates fragments of the invading bacteriophage's DNA into a CRISPR stream and uses corresponding CRISPR RNAs (crRNAs) to guide the degradation of homologous sequences in the invading bacteriophage's exogenous DNA, thereby eliminating the invading bacteriophage. Based on this principle, one such technology developed for genomic DNA cutting in higher organisms is CRISPR / Cas9. Its function is the same as the ZFN and TALEN technologies mentioned above, but it is simpler to operate and more flexible in use. Therefore, the CRISPR-Cas system has many advantages, including simple and convenient operation, high genome editing efficiency, low cytotoxicity, and wide applicability.
[0007] In 2015, a new technology called Cpf1 emerged based on CRISPR / Cas. It is very similar to CRISPR / Cas9, both of which use transcribed RNA to bind to specific DNA and then cut the target DNA.
[0008] There are six main types of CRISPR-Cas systems, but current research on their use as gene editing technology mainly focuses on types II and V. As a genome editing technology, CRISPR-Cas systems have a technical limitation. When selecting a target gene editing site, both types II and V require searching for a suitable protospacer adjacent motif (PAM) within the target gene's nucleic acid sequence. Although CRISPR-Cas systems capable of recognizing different PAMs have been developed—for example, SpCas9 can recognize NGG, SaCas9 can recognize NNGRRT, LbCpf1 can recognize TTTN, and Cas12f strictly recognizes TTN—new CRISPR / Cas systems still need to be developed to meet the need for targeting any location within the genome of an organism. Summary of the Invention
[0009] To address the aforementioned technical problems, this invention provides a novel targeted nuclease, a targeted nuclease system, a polynucleotide, and a gene editing method using the aforementioned system.
[0010] Specifically, the present invention provides:
[0011] 1. A targeted nuclease comprising an amino acid sequence having more than 70% identity with SEQ ID NO.22 in the sequence listing and having endonuclease activity.
[0012] Optionally, the amino acid sequence is the amino acid sequence shown in SEQ ID NO.22 or SEQ ID NO.23 in the sequence listing, or the amino acid sequence is encoded by the nucleotide sequence shown in SEQ ID NO.5 in the sequence listing.
[0013] Optionally, SEQ ID NO.22 in the sequence listing is used as a reference sequence, which has a site-directed mutation at position 8.
[0014] Optionally, the amino acid sequence of the targeted editing enzyme is the amino acid sequence shown in SEQ ID NO.23 of the sequence listing.
[0015] 2. A targeted nuclease system comprising any of the targeted nucleases described above and guide RNA for guiding the targeted nuclease to bind to a target sequence.
[0016] Optionally, the guide RNA comprises a nucleotide sequence having more than 65% identity with SEQ ID NO.4 in the sequence listing, or
[0017] The nucleotide sequence of the guide RNA contains spacers of 18-20 nt in length.
[0018] Optionally, the guide RNA comprises a non-naturally occurring crRNA linked to tracrRNA, wherein the nucleotide sequence of the tracrRNA is SEQ ID NO.3, and the nucleotide sequence of the crRNA is SEQ ID NO.2, or
[0019] The nucleotide sequence of the guide RNA is SEQ ID NO.4 in the sequence listing.
[0020] 3. A polynucleotide, characterized in that it comprises a nucleotide sequence selected from any one of (i) and (ii):
[0021] (i) the nucleotide sequence encoding any of the aforementioned targeted nucleases; and
[0022] (ii) A nucleotide sequence that hybridizes with the nucleotide sequence in (i) and has CRISPR / Cas enzyme activity.
[0023] Optionally, the nucleotide sequence of the polynucleotide is SEQ ID NO.1 or SEQ ID NO.18 in the sequence listing.
[0024] 4. A gene editing method for cell genomes, comprising the following steps:
[0025] (i) Provide a DNA sequence comprising the coding sequence of the targeted nuclease system described above.
[0026] (ii) The DNA sequence encoding the target nuclease system is transferred into a cell via a transgenic method, thereby editing the genome of the cell.
[0027] Compared with the prior art, the present invention has the following advantages and positive effects:
[0028] The PAM sequence of the targeted nuclease (BaCas9) of this invention can be NNAAW. NNAAW is a specific sequence, and BaCas9 recognizes NNAAW with higher substrate cleavage efficiency than NGAGW, NNGAW, and NNGGW. Therefore, BaCas9's PAM prefers NNAAW, and the PAM is rich in AT bases, which is also suitable for editing AT-rich regions.
[0029] Furthermore, both in vitro and in vivo E. coli experiments have demonstrated that the BaCas9 nuclease PAM recognition and guide RNA described in this invention is reliable. BaCas9 possesses endonuclease activity and has great potential in gene editing applications. Attached Figure Description
[0030] Figure 1 A schematic diagram of a CRISPR-Cas system derived from Bacillus sp. PS06 is shown.
[0031] Figure 2 The diagram shows the secondary structure of gRNA composed of crRNA and tracrRNA derived from Bacillus sp. PS06.
[0032] Figure 3 The image shows a PAGE electrophoresis diagram of the purified BaCas9 protein expressed in prokaryotes.
[0033] Figure 4 The diagram shows the structure of Seq-900, which includes the target sequence TS and the 10N randomized sequence.
[0034] Figure 5 The sequencing peak diagram of the Seq-900 sequence is shown.
[0035] Figure 6 The image shows an electrophoresis diagram of the products of BaCas9 digested with Seq-900 in vitro.
[0036] Figure 7 The diagram shows the sequencing peaks of the product digested by Seq-900 in vitro using BaCas9.
[0037] Figure 8 Electrophoresis diagrams of BaCas9 digested Seq-900 at different in vitro temperatures are shown.
[0038] Figure 9 Electrophoresis images of BaCas9 digested Seq-900 at different reaction times in vitro are shown.
[0039] Figure 10 Electrophoresis diagrams of BaCas9 digested with Seq-900 at different spacer lengths are shown.
[0040] Figure 11 A schematic diagram of the PY80-RH18 gRNA vector is shown.
[0041] Figure 12 The image shows an electrophoresis diagram of the products from the in vitro digestion of Seq-900ABRRW by BaCas9.
[0042] Figure 13 The image shows the electrophoresis diagrams of the products of LacZ-1369 digested by BaCas9 in vitro. From left to right, the images show the enzyme system containing the target sequences Lg3, Lg4, and Lg5, and the electrophoresis diagrams of the products of LacZ-1369 digested by control example 1.
[0043] Figure 14 This diagram illustrates the BaCas9-mediated HDR knockout of the LacZ gene.
[0044] Figure 15 A schematic diagram of the PY86 carrier is shown.
[0045] Figure 16 A schematic diagram of the PY86-Lg3 Escherichia coli HDR editing vector is shown.
[0046] Figure 17 A schematic diagram of the PY86-Lg6 Enterobacter HDR editing vector is shown.
[0047] Figure 18 The diagram shows the LacZ gene knockout colonies: A: a plate containing the LacZ-Lg3 target in the vector; B: a plate containing the LacZ-Lg6 target in the vector; and C: a control plate.
[0048] Figure 19 The image shows PCR electrophoresis images of colonies picked from plate A (containing LacZ-Lg3 target), plate B (containing LacZ-Lg6 target), and plate C (CK control without target sequence) after LacZ gene HDR knockout. In the image, 1-4 are colonies picked from plate A and plate B, respectively, and PCR was repeated four times.
[0049] Figure 20 A schematic diagram of the P15A-LacZst vector is shown.
[0050] Figure 21A schematic diagram of the BaCas9-based ABE editing vector is shown.
[0051] Figure 22 The peak diagram of LacZst ABE base editing sequencing is shown. Detailed Implementation
[0052] The present invention will be further described below with reference to the accompanying drawings through specific embodiments. However, this is not a limitation of the present invention. Those skilled in the art can make various modifications or improvements based on the basic idea of the present invention, but as long as they do not depart from the basic idea of the present invention, they are all within the scope of the present invention.
[0053] In a first aspect, the present invention provides a targeted nuclease comprising an amino acid sequence having more than 70% identity with SEQ ID NO.22 in the sequence listing and having endonuclease activity.
[0054] Preferably, the targeted nuclease comprises an amino acid sequence that has 80% or more, or 90% or more, or 95% or more, or even 99% or more, or even 100% identity with SEQ ID NO.22 in the sequence listing and has endonuclease activity.
[0055] Preferably, the targeted nuclease is derived from the Cas9 endonuclease of Bacillus sp. PS06.
[0056] Those skilled in the art can readily mutate the amino acid or nucleotide sequences of the present invention using known methods, such as directed evolution and point mutation. Those artificially modified amino acids or nucleotides having 70% or more, 80% or more, 90% or more, or 95% or more identity with the amino acid or nucleotide sequences isolated from the present invention, are derived from and equivalent to the sequences of the present invention, provided they maintain the activity of expressing the target gene. The term "identity" as used herein refers to the similarity of natural amino acid or nucleic acid sequences. "Identity" includes nucleotide sequences whose promoter nucleotide sequences of the present invention have 70% or more, 75% or more, 85% or more, 90% or more, or 95% or more identity. Identity can be evaluated visually or using computer software. Using computer software, the identity between two or more sequences can be expressed as a percentage (%), which can be used to evaluate the identity between related sequences.
[0057] Preferably, the amino acid sequence of the targeted nuclease of the present invention is the amino acid sequence shown in SEQ ID NO.22 in the sequence listing.
[0058] Preferably, the targeted nuclease of the present invention is a Cas9 endonuclease, which has at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% identity with the amino acid sequence of SEQ ID NO. 22 in the sequence listing.
[0059] In one specific embodiment, the present invention also provides a mutated targeted nuclease, with SEQ ID NO.22 in the sequence listing as a reference sequence, having a site-directed mutation at position 8 in its amino acid sequence.
[0060] Preferably, using SEQ ID NO.22 in the sequence listing as a reference sequence, the 8th amino acid codon D GCG is mutated to codon A GAG.
[0061] Preferably, the amino acid sequence of the mutated targeted nuclease is the amino acid sequence shown in SEQ ID NO.23 in the sequence listing.
[0062] Preferably, the mutated targeting nuclease is an nBaCas9(D8A) endonuclease, which has at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% identity with the amino acid sequence of SEQ ID NO.23 in the sequence listing.
[0063] In a second aspect, the present invention provides a targeted nuclease system comprising the aforementioned targeted nuclease and guide RNA for guiding the targeted nuclease to bind to a target sequence.
[0064] The targeted nuclease system can form a guide RNA / BaCas9 endonuclease complex.
[0065] Preferably, the guide RNA can be a single guide RNA capable of forming a guide RNA / BaCas9 endonuclease complex, wherein the guide RNA / Cas9 endonuclease complex can recognize and bind to the target sequence and generate a nick or cleave the target sequence within the target sequence, wherein the single guide RNA contains a nucleotide sequence having at least 65%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% identity with SEQ ID NO.4 in the sequence listing.
[0066] In one embodiment, the single guide RNA comprises a non-naturally occurring crRNA linked to tracrRNA, wherein the nucleotide sequence of said tracrRNA is SEQ ID NO.3.
[0067] In another embodiment, the single-guide RNA comprises a non-naturally occurring crRNA linked to tracrRNA, wherein the nucleotide sequence of said crRNA is SEQ ID NO.2.
[0068] Preferably, the guide RNA comprises a non-naturally occurring crRNA linked to tracrRNA, wherein the nucleotide sequence of the tracrRNA is SEQ ID NO.3, and the nucleotide sequence of the crRNA is SEQ ID NO.2.
[0069] Optionally, the nucleotide sequence of the guide RNA is SEQ ID NO.4 of the sequence listing.
[0070] In one embodiment, the guide RNA includes a spacer, and the spacer length affects the activity of the targeted nuclease. The spacer length is preferably 12-32 nt, such as 12 nt, 14 nt, 16 nt, 18 nt, 20 nt, 22 nt, 24 nt, 26 nt, 28 nt, 30 nt and 32 nt, more preferably 18-20 nt.
[0071] Preferably, the guide RNA comprises a nucleotide sequence complementary to SEQ ID NO.10, SEQ ID NO.11, SEQ ID NO.12, or SEQ ID NO.13 or SEQ ID NO.17 in the sequence listing, and more preferably comprises a nucleotide sequence complementary to SEQ ID NO.10 in the sequence listing.
[0072] Preferably, the targeted nuclease or enzyme system of the present invention can recognize any different PAM in the genome of an organism, and the characteristic PAM sequence to be recognized is NNRRW.
[0073] In a third aspect, the present invention provides a polynucleotide comprising a nucleotide sequence selected from any one of (i) and (ii):
[0074] (i) the nucleotide sequence encoding any of the aforementioned targeted nucleases; and
[0075] (ii) A nucleotide sequence that hybridizes with the nucleotide sequence in (i) and has CRISPR / Cas enzyme activity.
[0076] Preferably, the nucleotide sequence of the polynucleotide is SEQ ID NO.1 or SEQ ID NO.18 in the sequence listing.
[0077] In a fourth aspect, the present invention provides a gene editing method for cell genomes, comprising the following steps:
[0078] (i) Provide a DNA sequence containing the coding sequence of the aforementioned targeted nuclease system.
[0079] (ii) Editing the cell genome by binding the DNA sequence containing the coding sequence of the targeted nuclease system to the target sequence in the cell genome through a transgenic method.
[0080] In one embodiment, the present invention provides a method for editing target sequences or target sites in a cell genome, the method comprising providing at least one Cas9 endonuclease and / or other functional fragments, at least one guide RNA, wherein the guide RNA and Cas9 endonuclease can form a complex capable of recognizing, binding to a target site or target sequence, and creating a cut in the target sequence or cleaving the target sequence.
[0081] Preferably, the method of modifying the target sequence or target site can be carried out at a temperature of 20-45°C, and more preferably at a temperature of 30-37°C.
[0082] Methods for modifying target sequences or target sites may include performing enzymatic digestion reactions on the cell genome (e.g., nucleic acid substrates) using the aforementioned targeted nuclease system. Preferably, the duration of the digestion reaction is 15-180 min, more preferably 75-180 min.
[0083] Preferably, the target sequence or target site can be located in the cell genome.
[0084] Preferably, the cell or target site can be a bacterial, fungal, insect, animal, or plant cell.
[0085] In a fifth aspect, the present invention provides a recombinant expression vector comprising the polynucleotides of the third aspect.
[0086] In a sixth aspect, the present invention provides a host cell comprising the recombinant expression vector of the fifth aspect.
[0087] The present invention also provides a novel CRISPR-Cas system and method of use, comprising BaCas9 endonuclease, guide RNA, and a recognized PAM sequence.
[0088] In one embodiment of the present invention, the guide RNA refers to RNA capable of forming a guide RNA / Cas9 endonuclease complex, wherein the guide RNA comprises a double-stranded molecule of non-naturally occurring chimeric crRNA and tracrRNA. The non-naturally occurring chimeric crRNA contains a variable targeting domain capable of targeting and binding to the nucleotide sequence of a target gene in the cell genome, and the non-naturally occurring chimeric crRNA contains at least one crRNA fragment derived from Bacillus sp. PS06. The tracrRNA is derived from Bacillus sp. PS06.
[0089] In one embodiment of the present invention, the guide RNA comprises a double-stranded molecule of non-naturally occurring chimeric crRNA and tracrRNA. The guide RNA can link the two strands of crRNA and tracrRNA together through an oligonucleotide chain to form sgRNA. The sgRNA can form a guide RNA / Cas9 endonuclease complex with BaCas9 endonuclease. The guide RNA / Cas9 endonuclease complex can recognize and bind to the target sequence and cause a nick or cleave of the target sequence.
[0090] The targeted nuclease or nuclease system of the present invention or the above-described editing method can cause double-stranded nucleotide splicing of the target gene and changes in the nucleotide sequence of the target gene target site.
[0091] sequence
[0092] The nucleic acid and amino acid sequences SEQ ID NO. and their lengths provided or used in this invention are shown in Table 1 below.
[0093] Table 1
[0094]
[0095]
[0096] The following examples further explain or illustrate the content of the present invention, but these examples should not be construed as limiting the scope of protection of the present invention.
[0097] example
[0098] Unless otherwise specified, the experimental methods used in the following examples are conventional methods. Unless otherwise specified, the experimental materials used in the following examples are commercially available biochemical reagents.
[0099] Transetta(DE3) E. coli competent cells and prokaryotic expression vectors were purchased from Beijing TransGen Biotech Co., Ltd. High-fidelity polymerase KOD ONE used for PCR was purchased from Toyobo Biotechnology Co., Ltd. Antibiotics, arabinose IPTG and x-gal, and Ni-NTA agarose resin pre-packed gravity columns were all purchased from Shanghai Sangon Biotech Co., Ltd. Plasmid purification kits, gel extraction kits, and PCR product purification kits were purchased from OMEGA Bio-tek. Restriction endonucleases, ligases, and the T7 in vitro transcription kit were also used. The T7 Quick High Yield RNA Synthesis Kit (NEB#E2050S) was purchased from New England Biolabs. PCR primers were synthesized by Beijing Qingke Biotechnology Co., Ltd. Nucleic acid sequences were synthesized by Nanjing Genscript Biotech Co., Ltd.
[0100] Unless otherwise specified, the routine molecular biology experimental procedures used in the following factual cases are based on Molecular Cloning: A Laboratory Manual (Second Edition).
[0101] In the following example, the present invention provides a Cas9 endonuclease derived from Bacillus sp. PS06, named BaCas9, whose gRNA and PAM sequences were analyzed and identified, and whose functions were identified in vitro and in microorganisms.
[0102] Example 1: Obtaining the BaCas9 nucleic acid sequence
[0103] Information on multiple strains containing the CRISPR-Cas system was collected from the NCBI database. Bioinformatics analysis was used to identify the Cas9 nuclease gene sequence. Considering the evolutionary similarity and coding region length of Cas9 nucleases, the Cas9 nuclease from Bacillus sp. PS06 (see SEQ ID NO.1) was selected and named BaCas9; and the Cas9 nuclease from Bacillus massilionigeriensis strain Marseille-P2384 (see SEQ ID NO.21) was named BmCas9. As a candidate gene, since BmCas9 did not show activity in E. coli tests, the BaCas9 nuclease from Bacillus sp. PS06 was chosen as the focus of the study.
[0104] Example 2: Obtaining endogenous gRNA from BaCas9
[0105] By analyzing the genome of Bacillus sp. PS06, its CRISPR tandem repeat domain was identified, and the crRNA sequence was extracted from it (see SEQ ID NO.2). Then, sequences that are partially complementary to the crRNA sequence and have termination signals were searched upstream and downstream of the Cas9 gene and CRISPR structure. The secondary structure formed by these sequences and crRNA was analyzed by RNA structure software, and the tracrRNA sequence was deduced (see SEQ ID NO.3).
[0106] The T7 promoter, target sequence, crRNA sequence, and tracrRNA sequence were sequentially ligated and synthesized using an in vitro oligonucleotide chain synthesis method. The resulting products were then amplified and enriched using PCR, and the PCR products were isolated and purified. The T7 Quick High Yield RNA Synthesis Kit (NEB#E2050S) was used to synthesize gRNA via in vitro transcription (see SEQ ID NO. 4).
[0107] Example 3: Prokaryotic expression and purification of BaCas9 nuclease
[0108] To achieve high expression of the BaCas9 nuclease in *E. coli*, codon optimization of BaCas9 was performed according to the codons preferred by *E. coli* (SEQ ID NO.5). The optimized BaCas9 gene was artificially synthesized, named OBaCas9, and ligated into full-length gold... The Blunt E2 prokaryotic expression vector was inserted, and sequencing analysis confirmed that OBaCas9 was correctly ligated into the vector, which was named E2-OBaCas9. The E2-OBaCas9 vector was transformed into Transetta(DE3) competent E. coli cells, and recombinant bacteria were obtained after positive clone selection.
[0109] Recombinant bacterial monoclonal cultures were inoculated into 50 mL LB liquid culture medium containing Amp and cultured overnight at 37°C with a shaker at 220 rpm. The next day, the overnight bacteria were scaled up to 1 L of LB liquid medium at a 1:100 ratio and cultured at 37°C until the D600 reached 0.5-0.6. 1.0 mM IPTG was added to induce expression, and the culture was transferred to 16°C with a shaker at 150 rpm and cultured overnight. The bacterial cells were collected by centrifugation, resuspended in lysis buffer, and sonicated (60% amplitude, 3s On / 4s Off, 10 minutes, Ningbo Xinzhi JY92-IID ultrasonic cell disruptor). The supernatant was collected by centrifugation and added to a equilibrated Ni-NTA agarose resin gravity column for gradient elution. The eluent was collected in fractions and analyzed by SDS-PAGE electrophoresis (see Appendix). Figure 3 The protein had a molecular weight of 125 kDa. The eluent with the highest protein content was then purified by dialysis desalting to obtain BaCas9 protein with high purity (amino acid sequence SEQ ID NO.22).
[0110] Example 4: In vitro enzyme digestion assay of BaCas9 protein (SEQ ID NO.22) to detect its activity.
[0111] First, a nucleic acid substrate, Seq-900, for BaCas9 restriction enzyme digestion assay was designed and constructed. It is 900 bp long and contains 10 random nucleotides. (See attached diagram for structural illustration.) Figure 4 The sequence is shown in SEQ ID NO.6, and the sequencing peak diagram is attached. Figure 5 Then, a gRNA transcription template targeting RH18 was designed and synthesized using Seq-900. The RH18 target sequence is shown in SEQ ID NO.7. The gRNA transcription template targeting RH18 was designed and synthesized in vitro using a (NEB#E2050S) kit, and named RH18-gRNA. The in vitro transcription conditions are shown in Table 2. BaCas9 in vitro digestion was then performed. The reaction buffer, test substrate Seq-900, BaCas9 nuclease, and RH18-gRNA were added and mixed according to the conditions shown in Table 3. The digestion reaction was carried out at 37℃ for 3 hours, followed by inactivation of the nuclease at 70℃ for 5 minutes. The digestion reaction mixture was analyzed by agarose gel electrophoresis. The electrophoresis results are shown in the appendix. Figure 6 Enzyme digestion yielded two nucleic acid fragments, 570bp and 320bp. The cleavage products were recovered and sequenced. See the attached sequencing peak diagram. Figure 7 In vitro enzyme digestion and sequencing results showed that BaCas9 has endonuclease activity in vitro.
[0112] Next, following the above method, the activity of BaCas9 at different temperatures was tested. Ten portions of the enzyme digestion reaction system were mixed and digested at 20℃, 22℃, 25℃, 28℃, 30℃, 32℃, 35℃, 37℃, 40℃, and 42℃ for 2-3 hours, respectively. The nuclease was then inactivated by high-temperature treatment at 70℃ for 5 minutes. The digestion reaction mixture was then analyzed by agarose gel electrophoresis (see attached image). Figure 8 The results showed that BaCas9 was active in vitro from 20℃ to 42℃, with the strongest activity at 37℃.
[0113] Next, the rate of BaCas9 digestion was tested. Following the same method, a 130 μL digestion reaction mixture was prepared on ice, and 10 μL was aliquoted into 12 tubes. All tubes were placed at 37°C for incubation. The reaction mixtures were removed at 15 min, 30 min, 45 min, 60 min, 75 min, 90 min, 105 min, 120 min, 135 min, 150 min, 165 min, and 180 min, and then inactivated by heating at 70°C for 5 min. Each of the 12 digestion reaction mixtures was analyzed by agarose gel electrophoresis (see attached image). Figure 9 The results showed that BaCas9 produced enzyme digestion products as early as 15 min, reached its peak at 75 min, and the brightness of the enzyme digestion products did not increase from 75 to 180 min.
[0114] The effect of spacer length on BaCas9 activity was investigated. Primers (T7-Spacer-GTTTCGGTACTCTCAGAGAAACCT) were designed with spacer lengths of 12 nt, 14 nt, 16 nt, 18 nt, 20 nt, 22 nt, 24 nt, 26 nt, 28 nt, 30 nt, and 32 nt, respectively. High-fidelity enzyme PCR was used to amplify sgRNA-transcribed DNA templates, and the corresponding BaCas9 sgRNA was synthesized in vitro. The test substrate, BaCas9 nuclease, sgRNA, and reaction buffer were added and mixed. The enzyme digestion reaction was carried out at 37°C for 2-3 hours, followed by a high-temperature treatment at 70°C for 5 minutes. The digestion reaction mixture was analyzed by agarose gel electrophoresis (see appendix). Figure 10 The results showed that BaCas9 was active in vitro when the spacer length was 18-20 nt.
[0115] The above results indicate that BaCas9 has endonuclease activity in vitro.
[0116] Example 5: Identification of the PAM sequence of the BaCas9 nuclease
[0117] PAM library consumption assay identified the PAM sequence of the BaCas9 nuclease.
[0118] BaCAS9 Editing Vector Construction: The J23119 promoter was ligated into the E2-BaCas9 vector obtained in Example 3 to construct the editing vector PY80. Next, the RH18 gRNA sequence was designed and synthesized based on the target sequence of the PAM library, and ligated downstream of the J23119 promoter in PY80 to complete the construction of the editing vector PY80-RH18 gRNA (see Appendix). Figure 11 Next, the vector PY80-RH18 gRNA and -Blunt E2 was transformed into competent Transetta(DE3) Escherichia coli cells, and positive colonies were identified. The identified positive colonies were then cultured by shaking to prepare electrocompetent cells.
[0119] The PAM library was constructed by artificially synthesizing the oligonucleotide sequence RH18-NNNNNN shown in SEQ ID NO.8, which contains the target sequence (SEQ ID NO.7) and 6 random bases. SEQ ID NO.8 was ligated into the pUC19S vector (obtained by modifying pUC19, replacing the Amp resistance gene with the Spe resistance gene aada) and transformed into DE3 electrocompetent cells containing PY80-RH18gRNA and pEASY-Blunt E2, respectively. 1 mL LSOC medium and 1.0 mM MIPTG were added, and the cells were treated at 37°C for 1-2 hours. Plasmids were then extracted, and PCR amplification of the PAM region sequence was performed using primers designed according to NGS sequencing requirements, with fewer than 24 cycles. The products were then subjected to next-generation sequencing.
[0120] PAM analysis was performed, counting the occurrences of 4096 PAM sequence combinations in both the experimental and control groups, and then standardizing the results using the total number of PAM sequences in each group. For any given PAM sequence, a PAM sequence was considered significantly consumed if log2 (standardized value in the control group / standardized value in the experimental group) was greater than 3.5. The characteristic PAM sequence identified by BaCas9 was NNRRW.
[0121] Example 6: In vitro verification of BaCas9 nuclease PAM sequence
[0122] After the PAM sequence was determined, a subset of the PAM sequence was selected for in vitro validation. Twenty-four BaCas9 digestion assay substrates, Seq-900ABRRW, approximately 900 bp in length, were designed and constructed, each containing 5 nucleotides of PAM (ABRRW). This was achieved by replacing the 10N in the Seq-900 sequence from Example 4 with ABRRW. Twenty-four sequences were constructed separately for each substrate. Then, a gRNA transcription template targeting RH18 was designed and synthesized. The RH18 target sequence is shown in SEQ ID NO. 7. The RH18 target gRNA transcription template was designed and synthesized, and the gRNA was synthesized in vitro using a (NEB#E2050S) kit, named RH18-gRNA. Then, in vitro BaCas9 digestion assay was performed. Reaction buffer, equal amounts of the test substrate Seq-900ABRRW, BaCas9 nuclease, and RH18-gRNA were added and mixed according to the conditions shown in Table 3. The digestion reaction was carried out at 37°C for 3 hours, followed by inactivation of the nuclease at 70°C for 5 minutes. The digestion reaction mixture was then analyzed by agarose gel electrophoresis (see attached image). Figure 12 Enzymatic digestion yielded two nucleic acid fragments, 570 bp and 320 bp. In vitro digestion results showed that BaCas9 recognized NNAAW with higher substrate cleavage efficiency than NNAGW, NNGAW, and NNGGW, indicating a stronger preference for NNAAW.
[0123] Next, the LacZ gene of Escherichia coli Transetta (DE3) was selected, primers were designed, and a 1369bp fragment of the LacZ gene was amplified from the genome of Transetta (DE3) using the high-fidelity DNA polymerase KOD ONE. This fragment was named LacZ-1369 (SEQ ID NO.9) and used as a substrate for BaCas9 digestion.
[0124] BaCas9 target sites, Lg3 (SEQ ID NO.10), Lg4 (SEQ ID NO.11), and Lg5 (SEQ ID NO.12), were designed and synthesized as sgRNA transcription DNA templates. LacZ-1369 was then digested in vitro using the sgRNA in vitro transcription and BaCas9 in vitro digestion methods described in Example 4. The digestion reaction mixture was analyzed by agarose gel electrophoresis (see [link to analysis]). Figure 13 If DNA fragments are generated below the LacZ-1369 fragment targeting Lg3, Lg4, and Lg5 on the gRAN, it indicates that BaCas9 is active in in vitro digestion. Figure 13 As can be seen, DNA fragments were generated below the LacZ-1369 fragment with Lg3, Lg4 and Lg5 target gRAN, which were consistent with the expected size, indicating that BaCas9 is active in in vitro cleavage.
[0125] Example 7: BaCas9 E. coli genome editing
[0126] In the genome of a living cell, once a double-strand break occurs in DNA, the cell initiates self-repair mechanisms, including non-homologous end joining (NHEJ) and homologous directed repair (HDR). Among these, NHEJ is the most common mechanism of cellular self-repair and a major cause of DNA insertion, deletion, and chromosomal translocation. It is also the basis for our gene editing.
[0127] BaCas9 activity was assessed by knocking out the LacZ gene in the *E. coli* genome and observing the presence of white colonies on x-gal-containing medium. The activity of BaCas9 in *E. coli* was verified by knocking out the LacZ gene using the homology-directed repair (HDR) method (see [link to relevant documentation]). Figure 14 ).
[0128] A donor vector was constructed, and the Donor sequence of LacZ (SEQ ID NO.14) was designed, amplified by PCR, and ligated into the PUC19S vector to construct the LacZ donor vector PUC19S-Donor. The constructed donor vector PUC19S-Donor was transformed into Transetta(DE3) E. coli competent cells and cultured overnight on LB solid medium containing Spe. Positive colonies were identified. Single Transetta(DE3) positive colonies containing the PUC19S-Donor plasmid were inoculated into SOC liquid medium containing Spe to prepare electroporation competent cells, named Transetta(DE3)-LacZD, and stored at -80℃ for later use.
[0129] The bacterial editing vector was constructed by ligating the arabinose-inducible promoter ParaBAD, along with the expression cassette ParaBAD-Gam-Beta-Exo containing the three HDR-related genes Gam, Beta, and Exo, to the vector E2-BaCas9 (the expression vector constructed in Example 3), naming it E2-BaCas9-HDR. Then, the J23119 promoter was ligated to the vector E2-BaCas9-HDR to construct the bacterial HDR editing vector PY86 (see Appendix). Figure 14 The LacZ gene targets Lg3 (SEQ ID NO. 10) and Lg6 (SEQ ID NO. 13) were then designed, and the target sequences and sgRNA were ligated into the PY86 vector to construct the BaCas9 bacterial editing vectors PY86-Lg3 and PY86-Lg6 (see appendix for details). Figure 16 and attached Figure 17 ).
[0130] The constructed vectors PY86-Lg3, PY86-Lg6, and the control vector PY86 were transformed into Transetta(DE3)-LacZD electroporated competent cells, respectively. The cells were then plated onto LB solid medium containing Amp and Spe and incubated overnight at 37°C. Positive bacteria containing both the donor and editing vectors were screened. These positive bacteria were then inoculated into SOC liquid medium containing Amp and Spe and incubated overnight at 37°C with shaking at 220 rpm. The bacterial culture was then inoculated at a 1:100 ratio into 5 ml of SOC liquid medium containing Amp and Spe. After shaking to achieve an OD600 of 0.5-0.6, 1.0 mM IPTG and 1.0 mM arabinose were added for induction, and the cells were further incubated at 37°C with shaking at 220 rpm for 6-8 hours. The bacterial suspension was then diluted at different gradients and spread onto SOC solid medium containing Amp, Spe, 1.0 mM MIPTG, 1.0 mM arabinose and x-gal. The culture was then incubated at 37°C overnight, and the colony growth was observed.
[0131] The results show (attached) Figure 18 Plate A, containing the Lg3 target vector, produced white colonies with an editing efficiency of approximately 40%. Plate B, containing the LacZ-Lg6 target vector, also produced white colonies with an editing efficiency of approximately 15%. The difference in editing efficiency may be due to the different target sequences. Control plate C, without the target vector, produced no white colonies. Four white colonies were picked from plate A, four from plate B, and one blue colony from plate C for PCR. Primers PF and DR were used to amplify the LacZ gene target region. Agarose gel electrophoresis was performed (see attached image). Figure 19 The PCR product from white colonies was 220 bp smaller than that from blue colonies. The PCR products were recovered and sequenced. Sequencing results showed that a precise 220 bp sequence (SEQ ID NO. 15) was deleted from the LacZ gene in the white colonies, consistent with the designed HDR deletion genotype. This demonstrates that BaCas9 can achieve precise gene editing in E. coli.
[0132] Example 8: BaCas9 E. coli gene base editing
[0133] Base editing technology is a novel target gene modification technology developed based on the CRISPR / Cas system. It can achieve single nucleotide site-directed mutations without cutting the nucleic acid backbone and can directly chemically modify target nuclear bases during genome and transcriptome editing.
[0134] DNA base editors fuse deaminases with nCas9 to make irreversible and permanent changes to the genome. DNA base editors are divided into cytosine base editors (CBE) and adenine base editors (ABE). CBE can convert target sites of the four DNA bases (A, T, C, and G) from C·G to T·A, while ABE can convert target sites of A·T to G·C. We used an ABE base editor to verify BaCas9 activity. We selected the LacZ gene, searched for the TGG codon, mutated it to TAG, causing premature termination of LacZ, and then edited it using the BaCas9 ABE base editor to restore the codon to TGG.
[0135] Construction of the LacZ expression vector: Primers were designed to amplify the LacZ gene from the Transetta(DE3) genome. The LacZ gene was mutated to LacZst (sequence SEQ ID NO.16) using overlap extension PCR. LacZst was then ligated into the low-copy plasmid P15A to construct P15A-LacZst (see appendix). Figure 20 ).
[0136] Construction of the BaCas9 base editing vector: The Tada8e gene was synthesized artificially. Primers were designed to mutate the D codon (GCG) of the 8th amino acid of BaCAS9 to the A codon (GAG), thus mutating BaCAS9 into nBaCAS9(D8A). Then, Tada8e, the linker (32aa), and nBaCAS9(D8A) were ligated into the vector using the Gibson assembly method. -Blunt E2 vector to obtain E2-nBaCAS9(D8A) vector. Then, primers were designed to connect J23119 promoter to vector E2-nBaCAS9(D8A) to complete the construction of E. coli BaCas9 base editing vector, named PY100.
[0137] Next, a target site stg2 was designed at the location where the TGG mutation in the LacZst sequence was changed to TAG (see SEQ ID NO.17). Then, primers were designed to amplify and fuse the target site stg2 with the BaCas9 gRNA scaffold. The scaffold was then ligated into the PY100 vector using the Gibson assembly method, thus constructing the base editing vector PY100-stg2 (see [link to original text]). Figure 21 ).
[0138] Next, the control group CK (PY100 and P15A-LacZst) and the experimental group PT (PY100-stg2 and P15A-LacZst), respectively, were co-transformed into S2060 *E. coli* heat-shock competent cells. The cells were then plated onto LB solid medium containing Amp and Kan and incubated overnight at 37°C. Positive colonies were then selected. The selected positive colonies containing both vectors were inoculated into LB liquid medium containing Amp and Kan and incubated overnight at 37°C with shaking at 220 rpm. The activated bacterial culture was then inoculated at a 1:100 ratio into 5 ml of LB liquid medium containing Amp and Kan. After shaking to bring the OD600 to 0.5-0.6, 1.0 mM IPTG was added to induce expression, and the cells were further incubated at 37°C with shaking at 220 rpm for 3-5 hours. The bacterial suspension was then diluted at different gradients and spread onto LB solid medium containing Amp, Kan, and 1.0 mM IPTG, and incubated at 37°C overnight to observe colony growth.
[0139] Twenty-four colonies were picked from plates of both the control group (CK) and the experimental group (PT). Primers were designed to amplify the upstream and downstream sequences of the LacZst target site, and sequencing was performed. Sequencing peak diagrams are attached. Figure 22In the control group (CK), the five A's at the stg2 target site in the LacZst sequencing peak diagram remained unchanged, and no LacZst editing occurred in any of the 24 colonies. In the experimental group (PT), four of the five A's at the stg2 target site in the LacZst sequencing peak diagram showed double peaks, indicating a mutation of A to G. Among these, TAG mutated back to TGG, and base editing occurred in 23 colonies of LacZst. The experimental results demonstrate that the BaCas9-based ABE base editor possesses base editing capabilities in *E. coli*.
[0140] In summary, both in vitro and in vivo experimental results in E. coli demonstrate that the BaCas9 nuclease PAM recognition and guide RNA described in this invention is reliable. BaCas9 possesses endonuclease activity and has great potential for gene editing applications.
[0141] Table 2. In vitro transcription reaction system of gRNA
[0142]
[0143] Table 3. Cas9 in vitro enzyme digestion reaction system
[0144] Components Volume μL Seq-900 (200ng / μL) 3.0 Cas9 (0.5 mg / mL) 1.0 sgRNA (0.5 μg / μL) 1.0 10×BaCas9 buffer* 1.0 <![CDATA[ddH2O]]> 4.0 Total volume 10.0
Claims
1. A targeted nuclease, characterized in that It contains an amino acid sequence that has more than 70% identity with SEQ ID NO.22 in the sequence listing and has endonuclease activity.
2. The targeted nuclease of claim 1, wherein, The amino acid sequence is the amino acid sequence shown in SEQ ID NO.22 or SEQ ID NO.23 in the sequence listing, or the amino acid sequence is encoded by the nucleotide sequence shown in SEQ ID NO.5 in the sequence listing.
3. The targeted nuclease of claim 1, wherein Using SEQ ID NO.22 in the sequence listing as a reference sequence, this amino acid sequence has a site-directed mutation at position 8.
4. The targeted nuclease according to claim 3, characterized in that, The amino acid sequence of the targeted editing enzyme is the amino acid sequence shown in SEQ ID NO.23 in the sequence listing.
5. A targeted nuclease system, characterized in that... It comprises the targeted nuclease as described in any one of claims 1-4 and guide RNA for guiding the targeted nuclease to bind to the target sequence.
6. The targeted nuclease system according to claim 5, characterized in that the guide RNA comprises a nucleotide sequence having more than 65% identity with SEQ ID NO.4 in the sequence listing, or The nucleotide sequence of the guide RNA contains spacers of 18-20 nt in length.
7. The targeted nuclease system according to claim 5 or 6, characterized in that... The guide RNA comprises a non-naturally occurring crRNA linked to tracrRNA, wherein the nucleotide sequence of the tracrRNA is SEQ ID NO.3, and the nucleotide sequence of the crRNA is SEQ ID NO.2, or... The nucleotide sequence of the guide RNA is SEQ ID NO.4 in the sequence listing.
8. A polynucleotide, characterized in that... Contains a nucleotide sequence selected from either (i) or (ii) below: (i) the nucleotide sequence encoding the targeted nuclease of any one of claims 1 to 4; and (ii) A nucleotide sequence that hybridizes with the nucleotide sequence in (i) and has CRISPR / Cas enzyme activity.
9. The polynucleotide according to claim 8, characterized in that... The nucleotide sequence of the polynucleotide is SEQ ID NO.1 or SEQ ID NO.18 in the sequence listing.
10. A gene editing method for cell genomes, comprising the following steps: (i) Providing a DNA sequence comprising the coding sequence of the targeted nuclease system according to any one of claims 5 to 7. (ii) The genome of the cell is edited by transferring the DNA sequence containing the coding sequence of the targeted nuclease system according to any one of claims 5 to 7 into the cell via a transgenic method.