Cas protein, gene editing system comprising same, and use thereof
By modifying the endogenous CRISPR-Cas9 system of Streptococcus equi subsp. veterinaria (CGMCC NO.26991), a Cas protein and its CRISPR-Cas system were developed, enabling efficient genome editing of Streptococcus equi. This solved the problems of long cycle, low intensity, and difficult isolation and purification in HA production, and provided an efficient gene editing tool.
Patent Information
- Application Number
- PCT/CN2025/107574
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-08
- Filing Date
- 2025-07-08
- Publication Date
- 2026-01-15
AI Technical Summary
In the existing technology, gene editing tools for Streptococcus equi cannot be effectively applied, and traditional strains have long production cycles, low intensity, and are difficult to isolate and purify when producing HA.
Develop a Cas protein and its CRISPR-Cas system, specifically targeting Streptococcus equi subsp. veterinaria (CGMCC NO.26991), and achieve seamless genome editing by modifying its endogenous CRISPR-Cas9 system.
Efficient gene editing of the genome of Streptococcus equi subsp. veterinaria was successfully achieved, solving the problems of long cycle, low intensity and difficult isolation and purification in HA production, and providing an efficient gene editing tool.
Smart Images

Figure CN2025107574_15012026_PF_FP_ABST
Abstract
Description
A Cas protein, its gene editing system, and its applications
[0001] This application claims priority to Chinese Patent Application No. 202410909454.4, filed on July 8, 2024, entitled “A Cas protein and its gene editing system and application”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of bioengineering technology, and in particular to a Cas protein, its gene editing system, and its applications. Background Technology
[0003] Hyaluronic acid (HA) is an acidic polysaccharide widely distributed in the human body, found in tissues such as skin, eyes, and joints, and plays a role in lubrication and cell adhesion. With the development of synthetic biology, the method of enhancing the metabolic pathways of target products through rational design has become increasingly common. Currently, HA is widely expressed heterologously, and many generally recognized safe (GRAS) strains, such as Bacillus subtilis, Lactococcus lactis, and Corynebacterium glutamicum, have been designed as alternative HA-producing strains. These strains share the characteristics of relative safety, clear genetic background, and mature genetic manipulation techniques, but compared to the traditional Streptococcus equi fermentation cycle (approximately 50-100 hours), production intensity is lower, and subsequent isolation and purification are difficult. In actual production, Streptococcus equi is favored due to its short fermentation cycle (approximately 24-26 hours) and high production intensity. Although Streptococcus equi is widely used in the HA industry, highly efficient genome editing tools are not yet available for this species.
[0004] Bacteriophage infection has long been a problem plaguing the fermentation industry. Over hundreds of millions of years of evolution, bacteria have developed an adaptive immune system. This system, guided by RNA-guided nucleases, can cleave foreign nucleic acid elements, thus helping microorganisms resist the invasion of bacteriophages and other pathogens. This system is known as the CRISPR-Cas (Clustered Regularly Interspaced Short Palindromic Repeats and CRISPR-Associated Protein) system, composed of clusters of regularly interspaced short palindromic repeats and CRISPR-associated proteins. This system can be divided into two classes: Class I and Class II. Class I effector proteins are complexes composed of multiple Cas proteins, while Class II effector proteins contain only a single Cas protein with multiple domains.
[0005] In summary, there is an urgent need to develop a gene-editing tool that can be used for Streptococcus equi. Summary of the Invention
[0006] This application discovers a Cas protein and its CRISPR-Cas system, which solves the problem of the difficulty in applying genome editing tools to Streptococcus equi. Furthermore, the application modifies Streptococcus equi to achieve seamless editing of the genome of Streptococcus equi subsp. zooepidemicus CGMCC NO.26991.
[0007] On the one hand, this application provides a Cas protein, the amino acid sequence of which contains SEQ ID NO:5 or has at least 90% identity with SEQ ID NO:5;
[0008] Optionally, the amino acid sequence of the Cas protein has at least 90% identity with SEQ ID NO:5, which is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity.
[0009] Alternatively, the amino acid sequence of the Cas protein comprises the amino acid sequence shown in SEQ ID NO:5;
[0010] Optionally, the protein has Cas effector protein activity;
[0011] The Cas effector protein activity refers to the Cas protein having nuclease properties, which enables it to recognize and cleave specific DNA sequences.
[0012] Optionally, the protein is an effector protein in a CRISPR / Cas system;
[0013] In one alternative implementation, the effector protein in the CRISPR / Cas system refers to the Cas protein forming a complex with the guide polynucleotide in the CRISPR / Cas system, or the Cas protein specifically binding to the guide polynucleotide to the target nucleic acid.
[0014] Optionally, the PAM sequence (5'→3') recognized by the Cas protein is selected from NTN, ANN, TTN, ATC, NAC, AGA, TGC, TCT, NGN, CGC, NTC, GCA, TCG, TTT, CCG, GGG, NAG, ACA, CGG, CNG, ACN, GTG, CNT, TTG, TCN, GGT, TNC, CCN, CGT, TGG, CGA, NGG, TCC, AGT, NCA, CAN, TCA, NNG, TAC, CCT, NTG, CGN, TGN, CAT, NGC, GNG, GNC, NNA, GAA, TTC, CTT, ATA, TAT, GCT, NCC, TTA, AGN, GNN, CAA, CAA, and CAA. C, AGG, NTT, ANG, GNA, GTT, NGA, TAA, GTA, GGN, GNT, NCG, ATT, CCA, CNN, AAA, AAC, ATN, GAG, CTG, ACG, NAA, TAN, NAT, CNA, GCN, GTC, NCN, CTN, CNC, ANT, NNC, CAG, NAN , any one or more of ATG, NCT, CCC, AAN, TGT, TNA, ACC, GAT, ACT, AAT, GGA, GAN, ANC, GAC, NNT, CTA, TNN, GCG, GTN, TNT, AAG, TAG, NGT, NTA, ANA, CTC, GCC, TGA, GGC, AGC, TNG;
[0015] Alternatively, the PAM sequence recognized by the Cas protein is NGG.
[0016] N can be any base.
[0017] Optionally, the Cas protein forms a complex with the guide polynucleotide, or the Cas protein specifically binds to the target nucleic acid with the guide polynucleotide.
[0018] On the other hand, this application also provides a guide polynucleotide comprising: a repeat sequence and / or a guide sequence engineered to hybridize with a target nucleic acid;
[0019] Optionally, the repetitive sequence comprises SEQ ID NO:6 or a nucleotide sequence having at least 80% identity with SEQ ID NO:6;
[0020] Optionally, the same-direction repeat sequence having at least 80% identity with SEQ ID NO:6 is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity.
[0021] Alternatively, the nucleotide sequence of the same-direction repeat sequence is as shown in SEQ ID NO:6;
[0022] Optionally, the guide sequence comprises 15-60 nucleotides;
[0023] Alternatively, the guide sequence comprises 15-35 nucleotides;
[0024] The guiding sequence may contain 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 37, 38, 39, 40, 45, 50, 55, or 60 nucleotides.
[0025] In one alternative implementation, the guiding sequence may be a spacer sequence, which is 15-60 nucleotides upstream of the PAM sequence in the target nucleic acid.
[0026] Optionally, the same-direction repeating sequence is connected to the guiding sequence;
[0027] Alternatively, the same-direction repeating sequence is located at the 3' end of the guiding sequence;
[0028] Alternatively, the same-direction repeating sequence is located at the 5' end of the guiding sequence;
[0029] Alternatively, the same-direction repeating sequence is located at the 5' end and the 3' end of the guiding sequence;
[0030] Optionally, the guide polynucleotide forms a complex with the Cas protein and guides the complex to bind to the target nucleic acid in a sequence-specific manner;
[0031] Alternatively, the Cas protein is an amino acid sequence containing SEQ ID NO:5 or having at least 90% identity with SEQ ID NO:5.
[0032] Alternatively, the amino acid sequence of the Cas protein comprises the amino acid sequence shown in SEQ ID NO:5.
[0033] The engineering includes any technique that manipulates genes at the molecular level, such as mutating or knocking out a foreign gene and / or gene, then introducing it into a recipient cell after in vitro recombination, so that the gene can be replicated, transcribed, translated and expressed in the recipient cell.
[0034] On the other hand, this application also provides a fusion protein or conjugate comprising: the Cas protein and / or homologous or heterologous functional domains;
[0035] Optionally, the homologous or heterologous functional domains are connected to the N-terminus or C-terminus of the Cas protein via a linker or are internally fused or conjugated.
[0036] Optionally, the homologous or heterologous functional domains are fused to the N-terminus or C-terminus of the Cas protein, or internally fused or conjugated.
[0037] Optionally, the homologous or heterologous functional domains are selected from one or more of the following: subcellular localization signal, DNA binding domain, protease domain, transcription activation domain, transcription repression domain, nuclease domain, deaminase domain, uracil DNA glycosylase domain (UDG), uracil DNA glycosylase repression domain (UGI), methylase, demethylase, transcription release factor, histone acetylase domain, histone deacetylase domain, DNA ligase, affinity tag, reporter tag, affinity domain, and reporter domain.
[0038] Optionally, the nuclease domain includes a polypeptide with ssDNA cleavage activity and / or a polypeptide with dsDNA cleavage activity.
[0039] On the other hand, this application also provides a nucleic acid molecule that encodes a protein as follows (A1) or (A2):
[0040] The Cas protein described in (A1);
[0041] (A2) The fusion protein or conjugate described therein;
[0042] Optionally, the nucleic acid molecule is expressed in cells;
[0043] Optionally, the nucleic acid molecule is expressed in prokaryotic cells;
[0044] Optionally, the nucleic acid molecule is expressed in eukaryotic cells;
[0045] Optionally, the nucleic acid is expressed in eukaryotes, mammals such as humans or non-human mammals, plants, insects, birds, reptiles, rodents, fish, worms / nematodes, or yeast;
[0046] Optionally, the nucleic acid molecule is expressed in Streptococcus equi;
[0047] Optionally, the nucleic acid molecule is expressed in Streptococcus equi subsp. zooepidemicus, with accession number CGMCC No. 26991;
[0048] Optionally, the nucleic acid molecule is expressed in Escherichia coli; alternatively, the nucleic acid molecule contains a nucleotide sequence as shown in SEQ ID NO:4.
[0049] Optionally, the nucleic acid molecule contains SEQ ID NO:4 or a nucleotide sequence having at least 90% identity with SEQ ID NO:4;
[0050] More preferably, the nucleic acid molecule has at least 90% identity with SEQ ID NO:4, which is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity.
[0051] Alternatively, the nucleotide sequence of the nucleic acid molecule is shown in SEQ ID NO:4.
[0052] Those skilled in the art will understand that nucleic acid molecular sequences can be selectively optimized to improve their expression.
[0053] In one optional embodiment, the Streptococcus equi subsp. zooepidemicus is deposited at the China General Microbiological Culture Collection Center with accession number CGMCC No. 26991.
[0054] In one alternative embodiment, the Escherichia coli is Escherichia coli MG1655 or DH5α.
[0055] On the other hand, this application also provides a CRISPR-Cas protein system, which comprises the following (B1) and / or (B2):
[0056] (B1) the Cas protein, or the fusion protein or conjugate, or the nucleic acid molecule described herein;
[0057] (B2) The guiding polynucleotide, or the nucleotide sequence encoding the guiding polynucleotide;
[0058] Optionally, the Cas protein domain or the fusion protein or conjugate forms a complex with the guiding polynucleotide; the guiding polynucleotide contains a guiding sequence, which is engineered to guide the complex to bind to the target nucleic acid in a sequence-specific manner.
[0059] Optionally, the guiding polynucleotide comprises a homologous repeat sequence linked to the guiding sequence; more preferably, the homologous repeat sequence has at least 80% identity with SEQ ID NO:6.
[0060] More preferably, the same-direction repeat sequence having at least 80% identity with SEQ ID NO:6 is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity.
[0061] Optionally, the guide sequence comprises 15-35 nucleotides;
[0062] The guiding sequence may contain 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides.
[0063] Optionally, the guide sequence hybridizes with the target nucleic acid; more preferably, the guide sequence and the target nucleic acid are 90%-100% complementary.
[0064] The guide sequence and the target nucleic acid can be 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% complementary.
[0065] Optionally, the guiding sequence is located at the 3' end of the same-direction repeating sequence;
[0066] Optionally, the guiding sequence is located at the 5' end of the same-direction repeating sequence;
[0067] Optionally, the guiding sequence is located at the 5' end and the 3' end of the same-direction repeating sequence;
[0068] Optionally, the target nucleic acid is DNA or RNA; more preferably, the target nucleic acid is dsDNA or ssDNA.
[0069] Optionally, the DNA is eukaryotic DNA; more preferably, the eukaryotic DNA is non-human mammal DNA, non-human primate DNA, human DNA, plant DNA, insect DNA, bird DNA, reptile DNA, rodent DNA, fish DNA, worm / nematode DNA, or yeast DNA.
[0070] Optionally, the target nucleic acid is an endogenous gene of Streptococcus equi.
[0071] Optionally, the target nucleic acid is a hyaluronic acid lyase-related gene.
[0072] In an alternative embodiment, the target nucleic acid is the hylb HA lyase gene, as shown in SEQ ID NO:7, and the guide sequence is shown in SEQ ID NO:8.
[0073] Those skilled in the art will understand that the target nucleic acid can be determined according to the needs and the guide sequence can be designed based on the specific sequence of the target nucleic acid; therefore, no further restrictions are placed on its specific sequence here.
[0074] Furthermore, the CRISPR-Cas protein system further includes a fusion protein or conjugate comprising: the Cas protein and / or homologous or heterologous functional domains.
[0075] In a preferred embodiment, the CRISPR-Cas protein system is an endogenous CRISPR-Cas protein system of Streptococcus equi, and the Streptococcus equi is the zooepidemic subsp. zooepidemicus, which is deposited at the China General Microbiological Culture Collection Center with accession number CGMCC No. 26991.
[0076] The Cas protein, guide polynucleotide, fusion protein, conjugate, or nucleic acid molecule in the CRISPR-Cas protein system can be expressed in either integrated or free form. Those skilled in the art can modify the system based on existing technology.
[0077] In a preferred embodiment, the method for constructing the CRISPR-Cas protein system includes introducing a guide polynucleotide into Streptococcus equi.
[0078] Those skilled in the art can choose the appropriate introduction method based on the actual situation, as long as it can guide the expression and function of polynucleotides; no specific requirements are made here.
[0079] On the other hand, this application also provides a vector system comprising one or more recombinant vectors, the recombinant vectors comprising the nucleic acid molecule and the CRISPR-Cas protein system;
[0080] Optionally, the recombinant vector further comprises a regulatory sequence;
[0081] Optionally, the polynucleotide sequence encoding the Cas protein, fusion protein, or conjugate is operatively linked to the regulatory sequence, and / or the polynucleotide sequence encoding the guide polynucleotide is operatively linked to the regulatory sequence; more preferably, the regulatory sequence is selected from one or more of promoters, enhancers, internal ribosome entry sites, and transcription termination signals; more preferably, the promoter is selected from one or more of constitutive promoters, inducible promoters, broad-spectrum promoters, and tissue-specific promoters; more preferably, the transcription termination signal is selected from polyadenylation signals or polyU sequences;
[0082] Optionally, the isolated nucleic acid is linked to an aptamer sequence;
[0083] The aptamer sequences shown contain all sequences that can be linked to isolated nucleic acids.
[0084] On the other hand, this application also provides a recombinant cell comprising the Cas protein, the guide polynucleotide, the fusion protein or conjugate, the nucleic acid, the CRISPR-Cas protein system, or the vector system.
[0085] Optionally, the cells are prokaryotic cells;
[0086] Optionally, the cells are Streptococcus spp.;
[0087] Alternatively, the cells are Streptococcus equi.
[0088] In one optional embodiment, the Streptococcus equi subsp. zooepidemicus is deposited at the China General Microbiological Culture Collection Center with accession number CGMCC No. 26991.
[0089] On the other hand, this application also provides a method for binding, cutting or modifying genes, or altering cell states, the method comprising contacting the target nucleic acid and / or cells with the Cas protein, the guide polynucleotide, the fusion protein or conjugate, the nucleic acid, the CRISPR-Cas protein system, or the vector system described herein.
[0090] Optionally, the method results in one or more of the following: increased or decreased expression of a specific gene, induction of cell senescence in vitro or in vivo, cell cycle arrest in vitro or in vivo, promotion and / or inhibition of cell growth in vitro or in vivo, induction of non-responsiveness in vitro or in vivo, induction of apoptosis in vitro or in vivo, and induction of necrosis in vitro or in vivo.
[0091] Optionally, the method is for non-diagnostic and / or therapeutic purposes.
[0092] On the other hand, this application also provides a gene editing method for Streptococcus equi, which involves contacting Streptococcus equi with the Cas protein, the guide polynucleotide, the fusion protein or conjugate, the nucleic acid, the CRISPR-Cas protein system, or the vector system.
[0093] In a preferred embodiment, this application provides a method for editing endogenous genes in Streptococcus equi, the method comprising: introducing a homologous repeat sequence, an upstream homologous arm of a target nucleic acid, a downstream homologous arm of a target nucleic acid, and a guide sequence into Streptococcus equi to achieve editing of endogenous genes in Streptococcus equi.
[0094] Optionally, the unidirectional repeat sequence, the guide sequence, and the unidirectional repeat sequence are sequentially linked and inserted into the upstream homologous arm and the downstream homologous arm of the target nucleic acid.
[0095] Those skilled in the art can use the above methods to introduce the target gene or knock out the target gene as needed.
[0096] Optionally, the import includes importing using the pCas and pTargetF or pSET4s-EM vectors.
[0097] In one optional embodiment, the Streptococcus equi subsp. zooepidemicus is deposited at the China General Microbiological Culture Collection Center with accession number CGMCC No. 26991.
[0098] On the other hand, this application also provides compositions comprising the Cas protein, the guide polynucleotide, the fusion protein or conjugate, the nucleic acid, the CRISPR-Cas protein system, the vector system and / or the recombinant cells.
[0099] On the other hand, this application also provides Streptococcus equi subsp. zooepidemicus, which is deposited at the China General Microbiological Culture Collection Center with accession number CGMCC No. 26991;
[0100] Optionally, the Streptococcus equi subsp. veterinaria contains one or more nucleotide sequences shown in SEQ ID NO:1-3 and / or nucleotide sequences having at least 80% identity with SEQ ID NO:1-3;
[0101] The Streptococcus equi subsp. veterinaryis contains any one or more of SEQ ID NO:1, SEQ ID NO:2, and SEQ ID NO:3.
[0102] The Streptococcus equi subsp. veterinaria has a nucleotide sequence that is at least 80%, 80%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 10% identical to any one or more of SEQ ID NO:1-3.
[0103] On the other hand, this application also provides the application of the Cas protein, the guide polynucleotide, the fusion protein or conjugate, the nucleic acid, the CRISPR-Cas protein system, the vector system, the recombinant cell, the method, the gene editing method, or the composition in the preparation of hyaluronic acid, increasing hyaluronic acid production, and / or in Streptococcus equi gene editing.
[0104] Optionally, the Streptococcus equi gene editing can be scarless editing. Optionally, scarless editing includes scarless insertion or scarless knockout, which can be selected by those skilled in the art according to their needs.
[0105] It is understood that those skilled in the art can select appropriate gene editing systems and gene editing methods according to the actual situation to complete the assembly and construction of the CRISPR-Cas protein system, the vector system, and the recombinant cells.
[0106] Those skilled in the art can use conventional methods to edit the endogenous genes of Streptococcus equi using the Cas protein, the guide polynucleotide, the fusion protein or conjugate, the nucleic acid, the CRISPR-Cas protein system, the vector system, the recombinant cell, the method described, the gene editing method described, or the composition described.
[0107] In one optional embodiment, this application provides a method for editing endogenous genes in Streptococcus equi, the method comprising: ligating a homologous repeat sequence, an upstream homologous arm of a target nucleic acid, a downstream homologous arm of a target nucleic acid, and a guide sequence to construct a plasmid; and transferring the plasmid into Streptococcus equi subsp. CGMCC NO.26991 to achieve editing of endogenous genes in Streptococcus equi.
[0108] Those skilled in the art can use the above methods to introduce the target gene or knock out the target gene as needed.
[0109] This application has the following beneficial effects:
[0110] In this application, a strain of Streptococcus equi subsp. veterinaria was screened and deposited at the China General Microbiological Culture Collection Center, accession number CGMCC No. 26991.
[0111] Furthermore, through genomic information analysis of Streptococcus equi subsp. veterinaria (CGMCC No. 26991), a novel Cas protein, Cas9, and its CRISPR-Cas9 system were discovered. The PAM sequence of this Streptococcus equi subsp. veterinaria was also determined, proving that its endogenous CRISPR-Cas9 system is active, thus solving the problem of not being able to use genome editing tools in Streptococcus equi.
[0112] Further research on the endogenous CRISPR-Cas9 system of Streptococcus equi subsp. Pleurotus equine (CGMCC No. 26991) demonstrated that the endogenous CRISPR-Cas9 system of Streptococcus equi subsp. Pleurotus equine (CGMCC No. 26991) has exogenous gene cutting activity. Finally, a Streptococcus equine CRISPR-Cas9 gene editing system was successfully developed, which can complete the genome editing of Streptococcus equi subsp. Pleurotus equine without leaving a trace. Attached Figure Description
[0113] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0114] Figure 1 is a schematic diagram of the endogenous CRISPR-Cas9 gene editing system of Streptococcus equi subsp. CGMCC No. 26991, including repetitive sequences, spacer sequences and related Cas proteins. In the schematic diagram, diamonds represent repetitive sequences, rectangles represent spacer sequences, and repetitive sequences and spacer sequences alternate to form a CRISPR array (short palindromic repeats that are regularly spaced and clustered). Other shapes represent related Cas proteins.
[0115] Figure 2 shows the results of agarose gel electrophoresis of the genome of the transformant strain in Example 4. Lanes 1-8 from left to right are the genome amplification products of the transformant strain, and lane 9 is the genome amplification product of strain CGMCC NO.26991.
[0116] Biological Preservation Information:
[0117] Streptococcus equi subsp. zooepidemicus is deposited at the China General Microbiological Culture Collection Center (CGMCC) with accession number CGMCC No. 26991 and deposit date of March 31, 2023. Detailed Implementation
[0118] Technical terms:
[0119] "Identity" refers to the degree of similarity between the nucleotide sequences of two nucleic acid molecules or the amino acid sequences of two protein molecules in molecular evolution studies.
[0120] In a broad sense, "recombination" refers to any gene exchange process that results in a change in genotype.
[0121] "Spacer" is a special DNA sequence derived from the genome of bacteria or archaea. It is a guide sequence formed during the process of cutting and integrating the DNA of invading bacteriophages or plasmids into the CRISPR spacer region. When editing with the CRISPR / Cas9 system, it can match or pair with the target nucleic acid sequence being edited.
[0122] The PAM sequence plays a crucial role in the CRISPR / Cas9 system. It serves as a binding signal for the Cas9 protein, guiding Cas9 to bind to genomic DNA.
[0123] "Repetitive sequences" are about 20-50 bp in length and contain 5-7 bp palindromic sequences. The transcription products can form hairpin structures to stabilize the overall secondary structure of RNA.
[0124] "Cas protein" refers to CRISPR-related protein (Cas) (also known as "CRISPR-related protein", "CRISPR effector", "effector", "Cas protein", "Cas enzyme" or "CRISPR enzyme"), which is a protein that performs enzymatic activity and / or binds to target sites on nucleic acids specified by RNA guides. Cas protein has endonuclease activity, nicking enzyme activity, exonuclease activity, transposase activity and / or excision activity; it may also be nuclease inactivated or partially inactivated.
[0125] "Guide RNA (also known as guide RNA, gRNA, crRNA or sgRNA)" refers to any RNA molecule that facilitates the targeting of the Cas protein of this application to a target nucleic acid (such as DNA and / or RNA). "Guide RNA" includes one or more guide RNAs and their equivalents known to those skilled in the art, including but not limited to RNA-based molecules (e.g., direct repeat (DR) sequences) capable of forming a complex with the Cas protein, and containing a sequence (e.g., a spacer sequence) that is sufficiently complementary to the target nucleic acid sequence to hybridize with the target nucleic acid sequence and guide the complex to bind specifically to the target nucleic acid sequence.
[0126] The “CRISPR system” is collectively referred to as transcripts and other elements involved in the expression of CRISPR-related (“Cas”) genes or directing their activity, including sequences encoding Cas genes, tracr (trans-activating CRISPR) sequences (e.g., tracrRNA or active tracrRNA), tracr pairing sequences (covering “homogeneous repeats” and partial holotypes processed by tracrRNA in the context of an endogenous CRISPR system), directing sequences (also known as “spacers” in the context of an endogenous CRISPR system), or other sequences and transcripts derived from CRISPR loci.
[0127] The terms "polynucleotide," "nucleotide," "nucleotide sequence," "nucleic acid," and "oligonucleotide" are used interchangeably. They refer to polymeric forms of nucleotides of any length, which are deoxyribonucleotides or ribonucleotides, or similar substances. Polynucleotides can have any three-dimensional structure and can perform any function, known or unknown. Polynucleotides may contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, the nucleotide structure can be modified before or after polymer assembly. The sequence of nucleotides can be interrupted by non-nucleotide components. Polynucleotides can be further modified after polymerization, such as by conjugation with labeled components.
[0128] "Wild type" is a term as understood by those skilled in the art and refers to the typical form of an organism, strain, or gene, or the characteristic that distinguishes it from mutant or variant forms when it exists in nature.
[0129] "Complementarity" refers to the ability of a nucleic acid to form one or more hydrogen bonds with another nucleic acid sequence via traditional Watson-Crick base pairing or other non-traditional types. The complementarity percentage indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary, respectively). "Complete complementarity" means that all consecutive residues in a nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in a second nucleic acid sequence. As used herein, “substantially complementary” refers to a complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% in a region having 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids hybridizing under stringent conditions.
[0130] "Expression" refers to the process of transcription from a DNA template into polynucleotides (such as mRNA or other RNA transcripts) and / or the subsequent translation of transcribed mRNA into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as "gene products." If the polynucleotides originate from genomic DNA, expression can include the splicing of mRNA in eukaryotic cells.
[0131] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably herein to refer to polymers containing amino acids of any length. The polymer may be linear or branched, may contain modified amino acids, and may be interrupted by non-amino acid components. These terms also cover polymers containing modified amino acids; such modifications include disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as binding to a labeled component. As used herein, the term “amino acid” includes natural and / or non-natural or synthetic amino acids, including glycine and its D and L optical isomers, as well as amino acid analogs and peptide mimics.
[0132] A "vector" or "expression vector" is a replicon, such as a plasmid, bacteriophage, virus, or kinase, to which another DNA segment (i.e., an "insertion") may attach to the replicon in order to induce replication of the attached segment in the cell.
[0133] An “expression cassette” contains a DNA coding sequence operatively linked to a promoter. “Operably linked” means that the component is in a relationship that allows it to function in the intended manner. For example, if a promoter affects its transcription or expression, the promoter is operatively linked to the coding sequence.
[0134] "Scarless editing" refers to gene editing that leaves no detectable markers or scars. This technology avoids introducing unintended insertions / deletions or editing scars into the genome, thereby reducing potential stimulation of the cell's natural defense mechanisms and minimizing interference with gene expression.
[0135] The terms “recombinant expression vector” or “DNA construct” are used interchangeably herein and refer to a DNA molecule comprising a vector and at least one insert. Recombinant expression vectors are typically created for the purpose of expressing and / or propagating the insert or for constructing other recombinant nucleotide sequences. The insert may or may not be operatively ligated to a promoter sequence and may or may not be operatively ligated to a DNA regulatory sequence.
[0136] Several aspects of this application relate to vector systems comprising one or more vectors, or the vectors themselves. Vectors can be designed for the expression of CRISPR transcripts (e.g., nucleic acid transcripts, proteins, or enzymes) in prokaryotic or eukaryotic cells. For example, CRISPR transcripts can be expressed in bacterial cells such as *Escherichia coli*, insect cells (using baculovirus expression vectors), yeast cells, or mammalian cells.
[0137] Prokaryotes use vectors to amplify multiple copies of a vector and express one or more nucleic acids, such as providing a source for delivery to one or more proteins into a host cell or host organism. Protein expression in prokaryotes is most frequently performed in *E. coli* using vectors containing constitutive or inducible promoters that direct the expression of fusion or non-fusion proteins. Fusion vectors add multiple amino acids to the protein encoded therein, such as to the amino terminus of the recombinant protein. Such fusion vectors can be used for one or more purposes, such as: (i) increasing the expression of the recombinant protein; (ii) increasing the solubility of the recombinant protein; and (iii) assisting in the purification of the recombinant protein by acting as a ligand in affinity purification. Typically, in fusion expression vectors, protein cleavage sites are introduced at the junction of the fusion moiety and the recombinant protein to allow the recombinant protein to be separated from the fusion moiety after purification of the fusion protein.
[0138] To more clearly illustrate the overall concept of this application, a detailed description is provided below with reference to the accompanying drawings and embodiments. Numerous specific details are set forth in the following description to provide a more thorough understanding of this application. However, it will be apparent to those skilled in the art that this application can be implemented without one or more of these details. In other instances, to avoid confusion with this application, some technical features well-known in the art have not been described.
[0139] Unless otherwise specified, all reagents or instruments used in the following embodiments, unless otherwise indicated by the manufacturer, are commercially available products. Where specific conditions are not specified in the embodiments, they are performed under standard conditions or conditions recommended by the manufacturer.
[0140] The plasmids, restriction enzymes, PCR enzymes, column DNA extraction kits, and DNA gel recovery kits used in the following examples are commercial products. The specific operations were performed according to the kit instructions.
[0141] Unless otherwise stated, the practice of this application employs conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, which are within the scope of skill in this field. See Sambrook, Fritsch, and Maniatis, *Molecular Cloning: A Laboratory Manual*, 2nd Edition (1989).
[0142] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. While any methods and materials similar to or equivalent to those described herein may also be used in the practice or testing of this application, alternative methods and materials are described hereafter. All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials relevant to those publications.
[0143] All publications and patents referenced in this specification are incorporated herein by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference, and are incorporated herein by reference to disclose and describe the methods and / or materials relating to the publications referenced.
[0144] Example 1: Identification of the activity of the endogenous CRISPR-Cas9 system
[0145] In this embodiment, a strain of Streptococcus equi subsp. zooepidemicus capable of producing high levels of hyaluronic acid was screened and deposited at the China General Microbiological Culture Collection Center (CGMCC) with accession number CGMCC No. 26991.
[0146] Analysis of the genome information of Streptococcus equi subsp. CGMCC NO.26991 confirmed that its genome contains the CRISPR-Cas9 system (Figure 1). Based on three spacer sequences (Spacer sequence, Spacer1 sequence is SEQ ID NO:1: atcctctgtattacttgatcaaacagatct; Spacer2 sequence is SEQ ID NO:2: taattgcaatacgttagtatcaatttcaag; Spacer3 sequence is SEQ ID NO:3: ttgaggaaggtttctgttgccaattgtctt), the specific recognition sequence (PAM sequence) of CRISPR-Cas9 in this Streptococcus equi subsp. CGMCC is inferred to be NGG (N can be any base). This demonstrates that its endogenous CRISPR-Cas9 system is active. The proof process is as follows.
[0147] (1) Determination of PAM sequence
[0148] The PAM sequence is a short sequence immediately adjacent to the spacer. Different Cas9s recognize different PAMs. For example, Cas9 in Streptococcus pyogenes recognizes the PAM NGG, which is adjacent to the spacer and located downstream of it. Therefore, when developing CRISPR-Cas9, it is necessary to first find the spacer sequence, obtain its neighboring sequences, and then summarize the PAM sequence.
[0149] Because the spacer sequence in *Streptococcus equi* subsp. *porphyria* CGMCC NO.26991 is relatively small, searching for PAM sequences is not accurate enough. Since PAM sequences are specific recognition sequences for Cas9, similar Cas9 structures should lead to similar PAM sequences. Therefore, we searched for similar neighboring PAM sequences along with spacer sequences from the NCTC12090 / NCTC11606 genomes, which have high Cas9 similarity to *Streptococcus equi* subsp. *porphyria* CGMCC NO.26991. The results are shown in Table 1. The pattern shows that the PAM sequence is NGG, where N can be any of the four bases A, T, C, or G. However, it is not 100% conserved, and structures similar to NGG, such as TTG in NCTC11606, may appear.
[0150] Table 1. Identification of the PAM sequence in CGMCC NO.26991
[0151] (2) Construction of exogenous plasmids in Escherichia coli DH5α
[0152] Using the *E. coli*-streptococcus* shuttle vector pDL278 as the backbone, homologous arms were designed on both sides of the lactose-inducible promoter. These homologous arms, along with Spacer1 (SEQ ID NO:1 sequence) or Spacer1-PAM sequence (Spacer1-NGG) from CGMCC NO.26991, were used as primers to synthesize primers SY-SPF / R (homologous arm + Spacer1) or SY-SP-PAMF / R (homologous arm + Spacer1-NGG). The primer sequences are shown in Table 2 below, with the underlined portions representing the Spacer1 sequence or Spacer1-PAM sequence. PCR amplification was performed using pDL278 as a template, with the primers amplifying using Spacer1 or Spacer1-NGG instead of the lactose-inducible promoter.
[0153] The PCR product was digested and amplified by DpnⅠ□, then transformed into E. coli DH5α, plated on LB plates containing a final concentration of 50 mg / L spectinomycin, and the transformants were picked and sent for sequencing. The sequencing results showed that the plasmid pDL278-Spacer / pDL278-Spacer-NGG was successfully constructed.
[0154] Table 2
[0155] The above-mentioned spectinomycin solution is prepared as follows: Under aseptic conditions, accurately weigh 1g of spectinomycin powder, add 10mL of sterile water to the powder using a pipette, and vortex until the powder dissolves to obtain a 100mg / mL spectinomycin solution. When using, dilute according to the required multiple.
[0156] (3) Transformation of exogenous plasmid into Streptococcus equi subsp. Pesticide CGMCC NO.26991
[0157] The exogenous plasmids pDL278 (blank control) / pDL278-Spacer / pDL278-Spacer-NGG were transformed into CGMCCNO.26991, respectively.
[0158] The transformation method was as follows: 1 μg of plasmid was added to competent cells of *Streptococcus equi* subsp. *porphyria* (CGMCC NO.26991), gently mixed, and transferred to a 0.22 mm electroporation cuvette. The cells were incubated on ice for 10 min. The cuvette was then placed in an electroporator and electroporated at a voltage of 200 Ω. 1 mL of seed culture medium was added, and the cells were incubated at 37°C and 200 rpm for a certain period. After incubation, the bacterial culture was centrifuged at 6000 rpm for 2 min to collect the cells. The collected cells were then plated on plates containing spectinomycin (final concentration 100 μg / mL) and incubated upside down at 37°C.
[0159] The above seed culture medium is prepared as follows: 5 g / L glucose, 10 g / L peptone, 5 g / L yeast extract, 1 g / L potassium dihydrogen phosphate, 1 g / L magnesium sulfate heptahydrate, with the remainder being water. If solid culture medium is required, add 2% agar. After the culture medium is prepared, sterilize it at 115℃ for 20 minutes.
[0160] The number of transformants on the plate was counted, with the number of pDL278 transformants set at 100%. The relative transformation efficiencies of the other two plasmids were calculated as follows.
[0161] Relative transformation efficiency % = number of transformants in the transforming plasmid / number of transformants in the blank control PDL278 plasmid * 100%.
[0162] Spacer-NGG is the recognition / cleavage site of Cas9. When the plasmid pDL278 containing the spacer-NGG phage is transformed into CGMCC NO.26991, the plasmid is cleaved at this site, making it impossible for the strain to survive on the resistance plate. The resistance plate should show little or no growth of transformants.
[0163] Table 3 Relative transformation efficiency of exogenous plasmids
[0164] (4) Results Analysis
[0165] As can be seen from Table 3 above, the transformation efficiency of the plasmid without the PAM sequence was not significantly different from that of the control. After adding the PAM sequence, no transformants grew. It can be preliminarily determined that the CGMCC NO.26991 endogenous CRISPR-Cas9 system has cleavage activity.
[0166] In this embodiment, to verify the activity of the endogenous CRISPR-Cas9 system in *Streptococcus equi* subsp. *porphyria* CGMCC NO.26991, plasmids pDL278-Spacer (without PAM sequence) and pDL278-Spacer-NGG (with PAM sequence) were constructed using the pDL278 shuttle vector as the backbone. After transformation of CGMCC NO.26991, the transformation efficiency of pDL278-Spacer was not significantly different from that of the control, while the transformation efficiency of pDL278-Spacer-NGG was 0, indicating that the endogenous CRISPR-Cas9 system is active and can recognize the PAM sequence and cleave its adjacent sequences.
[0167] Example 2: Cuttering the endogenous gene of Streptococcus equi subsp. Pesticide CGMCC NO.26991 using the CRISPR-Cas9 system.
[0168] (1) Construction of plasmid pSET4s-repeat-hylb Spacer-repeat
[0169] The endogenous target was selected as the hylb HA lyase gene, whose nucleotide sequence is shown in SEQ ID NO:7. The first 33 bp of any NGG sequence in the endogenous target gene was selected as the spacer sequence of the target gene. The spacer sequence used in this embodiment is CAGACAACAACTACACTCTTTTTATCAATGGGC (SEQ ID NO:8).
[0170] In this embodiment, the *E. coli*-streptococcus* shuttle vector pSET4s was used as the backbone. Primers were used to linearize the pSET4s vector backbone. The repeat sequence (as shown in SEQ ID NO:6) and the spacer sequence were designed together as primers repeat-spacer-repeat-F / R, and the two fragments were ligated. The resulting plasmid was transformed into *E. coli* DH5α, plated on LB agar plates containing 100 μg / mL of erythromycin, and transformed individuals were picked and sequenced. Sequencing results showed that the plasmid pSET4s-repeat-hylb Spacer-repeat was successfully constructed. The plasmid pSET4s-repeat-repeat was constructed using the same method.
[0171] The primers mentioned above are shown in Table 4.
[0172] Table 4
[0173] The above-mentioned erythromycin solution is prepared as follows: Under sterile conditions, accurately weigh 1g of erythromycin powder, add 10mL of anhydrous ethanol to the powder using a pipette, and vortex until the powder dissolves to obtain a 100mg / mL erythromycin solution. When using, dilute according to the required multiple.
[0174] (2) Transformation of exogenous plasmid into Streptococcus equi subsp. Pesticide CGMCC NO.26991
[0175] The exogenous plasmids pSET4s-repeat-hylb Spacer-repeat and pSET4s-repeat-repeat were transformed into CGMCCNO.26991, respectively.
[0176] The transformation method was as follows: 1 μg of plasmid was added to competent cells of *Streptococcus equi* subsp. *porphyria* (CGMCC NO.26991), gently mixed, and transferred to a 0.22 mm electroporation cuvette. The cells were then incubated on ice for 10 min. The cuvette was placed in an electroporator and electroporated at a voltage of 200 Ω. 1 mL of seed culture medium was added, and the cells were incubated at 37°C and 200 rpm for a certain period. After incubation, the bacterial culture was centrifuged at 6000 rpm for 2 min to collect the cells. The collected cells were then plated on plates containing erythromycin (final concentration 10 μg / mL) and incubated upside down at 37°C.
[0177] The above seed culture medium is prepared as follows: 5 g / L glucose, 10 g / L peptone, 5 g / L yeast extract, 1 g / L potassium dihydrogen phosphate, 1 g / L magnesium sulfate heptahydrate, with the remainder being water. If solid culture medium is required, add 2% agar. After the culture medium is prepared, sterilize it at 115℃ for 20 minutes.
[0178] The number of transformants on the plates was counted, and the relative transformation efficiency of pSET4s-repeat-hylb Spacer-repeat was calculated with the number of transformants from pSET4s-repeat-repeat as 100%. Theoretically, after transformation with plasmid pSET4s-repeat-hylb Spacer-repeat, the resistant plates should show little or no transformant growth. The experimental results are shown in Table 5.
[0179] Relative transformation efficiency % = (Number of transformants in the transformation plasmid / Number of transformants in the blank control pSET4s-repeat-repeat plasmid) * 100%
[0180] Table 5. Relative transformation efficiency of exogenous plasmids
[0181] (3) Relative transformation rate of exogenous plasmids and analysis of results
[0182] As shown in Table 5 above, the transformation efficiency of exogenous plasmids with spacer sequence targeting is significantly lower than that of exogenous plasmids without spacer targeting. It can be preliminarily concluded that the endogenous CRISPR-Cas9 system of *Streptococcus equi* subsp. *epidemicus* CGMCC NO.26991 can provide target sites through exogenous plasmids to cleave endogenous genomic genes.
[0183] Example 3: Comparison of cas9 protein cleavage activity with other cas9 proteins from Streptococcus equi subspecies of veterinary disease.
[0184] (1) Plasmid construction
[0185] We selected pCas vectors commonly used for Escherichia coli gene editing to express Cas9 proteins of different Streptococcus equi subsp. equi and Streptococcus equi subsp. equi CGMCC NO.26991, and compared the cleavage efficiency of the E. coli genome. For the Cas9 protein cleavage activity comparison experiment, we used pCas and pTargetF vectors for verification. The pCas and pTargetF vectors were provided by the team of Qi Qingsheng at Shandong University. The source and construction method of the original plasmids are referenced in Jiang Y, Chen B, Duan C, Sun B, Yang J, Yang S. Multigene editing in the Escherichia coli genome via the CRISPR-Cas9 system. Appl Environ Microbiol. 2015 Apr; 81(7):2506-14. The pTargetF vector contains an aspartate kinase thrA gene recognition site. The pCas plasmids for different Cas9 proteins were synthesized by Beijing Qingke Biotechnology Co., Ltd.
[0186] The Cas9 protein of Streptococcus equi subsp. CGMCC NO.26991 (nucleotide sequence as shown in SEQ ID NO:4, amino acid sequence as shown in SEQ ID NO:5), the Cas9 protein of the pCas vector (nucleotide sequence as shown in SEQ ID NO:11, amino acid sequence as shown in SEQ ID NO:12), and the Cas9 protein of NCTC11606 (nucleotide sequence as shown in SEQ ID NO:13, amino acid sequence as shown in SEQ ID NO:14) were named plasmids pCas-HX(26991), pCas, and pCas-11606 respectively according to the source of the Cas9 sequence. At the same time, a blank plasmid without the Cas9 sequence was constructed as a control group and named pCas-blank control. The pTargetF vector, commonly used for gene editing in Escherichia coli, was used to express sgRNA sequences designed with Escherichia coli, Corynebacterium glutamicum, Lactobacillus, or Bacillus subtilis genome information as target genes. The construction method was the same as described in the above literature: Jiang Y, Chen B, Duan C, Sun B, Yang J, Yang S. Multigene editing in the Escherichia coli genome via the CRISPR-Cas9 system. Appl Environ Microbiol. 2015 Apr; 81(7):2506-14.
[0187] (2) Plasmid transformation of E. coli
[0188] Using *E. coli* MG1655 as the substrate cell line, 100 ng of pCas plasmids containing different Cas9 proteins were added to 50 μL of commercially available MG1655 competent cells. Transformation was performed according to the MG1655 competent cell instructions, and the cells were finally plated on LB agar containing kanamycin (50 mg / L) and incubated upside down at 37°C. Transformants were picked, cultured on LB agar, and then treated with calcium chloride solution to prepare competent cells. The transformation process was repeated, transforming pTargetF plasmid, and the cells were plated on LB agar containing kanamycin (50 mg / L) and streptomycin (50 mg / L) and incubated upside down at 37°C. The number of transformants was counted, and the results are shown in Table 6.
[0189] The above LB medium is prepared as follows: 5 g / L sodium chloride, 5 g / L yeast extract, 10 g / L tryptone, with the remainder being water. If LB solid medium is to be prepared, add 2% agar powder.
[0190] The preparation method for the above-mentioned kanamycin solution is as follows: Under aseptic conditions, accurately weigh 0.5g of kanamycin powder, add 10mL of sterile water to the powder using a pipette, and vortex until the powder dissolves. This yields a 50mg / mL spectinomycin solution. When using, dilute as needed. The preparation method for streptomycin solution is the same as that for kanamycin solution.
[0191] Table 6. Number of transformants from different sources of Cas9 protein pCas plasmids in three batches of transformation experiments
[0192] (3) Analysis of the number of transformants from pCas plasmids containing Cas9 protein from different sources
[0193] As shown in Table 6 above, the pCas-pTargetF editing system expressing the endogenous Cas9 protein of Streptococcus equi subsp. Pleurotus equine (CGMCC NO.26991) produced the fewest transformants and exhibited the highest cleavage efficiency of the strain's genome. Therefore, it can be concluded that the endogenous Cas9 protein of Streptococcus equi subsp. Pleurotus equine (CGMCC NO.26991) exhibits the best cleavage activity compared to other Cas9 proteins.
[0194] Example 4: Scarless knockout of the endogenous gene in Streptococcus equi subsp. Pleurotus equine, CGMCC NO.26991, using the CRISPR-Cas9 system.
[0195] Based on Example 2, the plasmid pSET4s-repeat-hylb Spacer-repeat was constructed. In this example, upstream and downstream homologous arms of the hylb HA lyase gene were inserted into this plasmid to construct the plasmid pSET4s-hylb-up / down, which removes the hylb gene without scarring. The upstream homologous arm of the hylb HA lyase gene is shown in SEQ ID NO:9, and the downstream homologous arm is shown in SEQ ID NO:10.
[0196] (1) Construction of pSET4s-hylb-up / down plasmid
[0197] The upstream and downstream homologous arms of the hylb HA lyase gene were amplified by PCR using up-F / R and down-F / R primers, respectively. The upstream homologous arm of the hylb HA lyase gene is shown in SEQ ID NO:9, and the downstream homologous arm is shown in SEQ ID NO:10. After being ligated into an up-down fragment by ligation PCR, it was ligated with the pSET4s-repeat-hylb Spacer-repeat fragment linearized using pSETxxh-F / R primers to construct the pSET4s-hylb-up / down plasmid.
[0198] The exogenous plasmid pSET4s-hylb-up / down was transformed into CGMCC NO.26991 using the same method as in Example 2. After transformation, the number of transformants on the plate was counted, and the Cas9 system editing efficiency was calculated. Cas9 system editing efficiency % = Number of successfully edited transformants / Total number of plasmid transformants edited * 100%.
[0199] The plasmid construction method, plasmid transformation method for Streptococcus equi subsp. veterinaria var. veterinaria CGMCC NO.26991, strain culture medium formulation, and antibiotic concentration used in this example are all the same as in Example 2.
[0200] The primers mentioned above are shown in Table 7.
[0201] Table 7
[0202] (2) Transformed strains were passaged under antibiotic-free conditions to lose the exogenous plasmid pSET4s-hylb-up / down, thus avoiding exogenous gene contamination of the genome information of Streptococcus equi subsp. veterinary.
[0203] (3) Genome extraction method of transformant strains: The genomes of Streptococcus equi subsp. CGMCC NO.26991 and transformants were extracted using the FastPure Bacteria DNA Isolation Mini Kit-BOX 2 gene extraction kit from Nanjing Novizan Biotechnology Co., Ltd.
[0204] (4) Results and analysis of agarose gel electrophoresis
[0205] The genomes of the transformant strains were subjected to agarose gel electrophoresis, and the results are shown in Figure 2. The pSET4s-hylb-up / down plasmid with introduced homologous arms was able to repair the genome of *Streptococcus equi* subsp. *porphyria* CGMCC NO.26991, which had been cleaved by the CRISPR-Cas9 system, achieving scarless knockout of endogenous genes in *Streptococcus equi* subsp. *porphyria* CGMCC NO.26991. Based on the number of transformants, the scarless editing efficiency of *Streptococcus equi* subsp. *porphyria* CGMCC NO.26991 in this embodiment was 6%.
[0206] in conclusion
[0207] This experiment is the first to achieve efficient and scarless knockout of target genes in the genome of Streptococcus equi subsp. equine using an endogenous CRISPR-Cas9 system, laying the foundation for synthetic biology modification of Streptococcus equi subsp. equine.
[0208] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. Cas protein, characterized in that, The amino acid sequence of the Cas protein contains SEQ ID NO:5 or has at least 90% identity with SEQ ID NO:
5.
2. A CRISPR-Cas protein system, characterized in that, The CRISPR-Cas protein system comprises the Cas protein as described in claim 1.
3. The CRISPR-Cas protein system according to claim 2, characterized in that, It also includes a guide polynucleotide comprising: a direct repeat sequence and / or a guide sequence engineered to hybridize with a target nucleic acid; Optionally, the guide polynucleotide forms a complex with the Cas protein and guides the complex to bind to the target nucleic acid in a sequence-specific manner; Alternatively, the Cas protein is the Cas protein of claim 1; Optionally, the repetitive sequence comprises SEQ ID NO:6 or a nucleotide sequence having at least 80% identity with SEQ ID NO:6; Optionally, the guide sequence comprises 15-60 nucleotides; Optionally, the same-direction repeating sequence is connected to the guiding sequence; Alternatively, the same-direction repeating sequence is located at the 3' end of the guiding sequence; Alternatively, the same-direction repeating sequence is located at the 5' end of the guiding sequence; Alternatively, the same-direction repeating sequence is located at the 5' and 3' ends of the guiding sequence.
4. The CRISPR-Cas protein system according to claim 2, characterized in that, It also includes a fusion protein or conjugate comprising: the Cas protein as described in claim 1 and / or homologous or heterologous functional domains.
5. A nucleic acid molecule, characterized in that, The nucleic acid molecule encodes the following protein: (A1) or (A2): (A1) The Cas protein according to claim 1; (A2) The fusion protein or conjugate as described in claim 4.
6. A carrier system, characterized in that, The vector system comprises one or more recombinant vectors, the recombinant vectors comprising nucleic acid molecules encoding the CRISPR-Cas protein system as described in any one of claims 2-4, or the nucleic acid molecule as described in claim 5.
7. A recombinant cell, characterized in that, The recombinant cells comprise nucleic acid molecules encoding the Cas protein of claim 1, nucleic acid molecules of the CRISPR-Cas protein system of any one of claims 2-4, nucleic acid molecules of claim 5, or vector systems of claim 6; Optionally, the cells are prokaryotic cells; Optionally, the cells are Streptococcus spp.; Alternatively, the cells are Streptococcus equi.
8. A method for binding, cutting, or modifying genes, or altering cell states, characterized in that, The method includes contacting the target nucleic acid and / or cell with the Cas protein as described in claim 1, the CRISPR-Cas protein system as described in any one of claims 2-4, the nucleic acid molecule as described in claim 5, or the vector system as described in claim 6; Optionally, the method results in one or more of the following: increased or decreased expression of a specific gene, induction of cell senescence in vitro or in vivo, cell cycle arrest in vitro or in vivo, promotion and / or inhibition of cell growth in vitro or in vivo, induction of non-responsiveness in vitro or in vivo, induction of apoptosis in vitro or in vivo, and induction of necrosis in vitro or in vivo. Optionally, the method is for non-diagnostic and / or therapeutic purposes.
9. The application of the Cas protein of claim 1, the CRISPR-Cas protein system of any one of claims 2-4, the nucleic acid molecule of claim 5, or the vector system of claim 6, the recombinant cell of claim 7, or the method of claim 8 in the preparation of hyaluronic acid, increasing hyaluronic acid production, and / or gene editing of Streptococcus equi.
10. A composition, characterized in that, The composition contains the Cas protein as described in claim 1, the CRISPR-Cas protein system as described in any one of claims 2-4, the nucleic acid molecule as described in claim 5, the vector system as described in claim 6, and / or the recombinant cell as described in claim 7.
Citation Information
Patent Citations
Methods, cells and organisms
CN105637087A
Genetically engineered bacterium for producing hyaluronic acid and application of genetically engineered bacterium
CN116004496A
Novel cas9 proteins and guiding features for DNA targeting and genome editing
US20170275648A1
Novel crispr enzymes, methods, systems and uses thereof
WO2023114953A2