Novel CRISPR enzymes and systems and their applications
By developing a novel Cas protein that forms a complex with guide RNA, the shortcomings of existing CRISPR/Cas systems have been addressed, enabling more efficient gene editing and nucleic acid detection, and improving targeting and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG SHUNFENG BIOTECH CO LTD
- Filing Date
- 2022-10-21
- Publication Date
- 2026-07-31
AI Technical Summary
Existing CRISPR/Cas systems each have their own advantages and disadvantages, and there is a need to develop a more robust new CRISPR/Cas system with good performance in many aspects to improve the efficiency and accuracy of gene editing.
A novel Cas protein, including Cas-sf0005, Cas-sf9417, and Cas-sf8553, was developed. By using genetic engineering techniques to make minor changes to its amino acid sequence, its biological function was preserved, and it formed a complex with guide RNA to recognize and cleave specific target sequences.
It enables more efficient gene editing and nucleic acid detection, improves targeting and reduces off-target effects, and enhances the robustness of the system.
Smart Images

Figure CN118726314B_ABST
Abstract
Description
[0001] This invention is a divisional application of Chinese patent application CN202211296617.3, entitled "Novel CRISPR Enzymes and Systems and Applications", filed on October 21, 2022.
[0002] This application claims priority to Chinese patent application CN202111244315.7, filed on October 26, 2021. The entire contents of the aforementioned Chinese patent application are incorporated herein by reference. Technical Field
[0003] This invention relates to the field of gene editing, particularly to the field of regularly clustered short palindromic repeats (CRISPR) technology. Specifically, this invention relates to a novel CRISPR enzyme (or, as referred to as a CRISPR protein, Cas effector protein, Cas enzyme, or Cas protein), fusion proteins comprising such proteins, and nucleic acid molecules encoding them. This invention also relates to complexes and compositions for nucleic acid editing (e.g., gene or genome editing) comprising the Cas protein or fusion protein of this invention, or nucleic acid molecules encoding them. Background Technology
[0004] CRISPR / Cas technology is a widely used gene editing technology that uses RNA to specifically bind to target sequences on the genome and cut DNA to create double-strand breaks, using biological non-homologous end joining or homologous recombination for site-specific gene editing.
[0005] The CRISPR / Cas9 system is the most commonly used type II CRISPR system. It recognizes the 3'-NGG PAM motif and performs blunt-end cleavage on the target sequence. CRISPR / Cas Type V systems are a newly discovered class of CRISPR systems with a 5'-TTN motif, performing sticky-end cleavage on the target sequence; examples include Cpf1, C2c1, CasX, and CasY. However, the different CRISPR / Cas systems currently available each have their own advantages and disadvantages. For example, Cas9, C2c1, and CasX all require two guide RNAs, while Cpf1 only requires one and can be used for multiplex gene editing. CasX is 980 amino acids in size, while common systems like Cas9, C2c1, CasY, and Cpf1 are typically around 1300 amino acids. Furthermore, the PAM sequences of Cas9, Cpf1, CasX, and CasY are relatively complex and diverse, while C2c1 recognizes the strict 5'-TTN, making its target site easier to predict than other systems and reducing potential off-target effects.
[0006] In conclusion, given the limitations of currently available CRISPR / Cas systems, developing a more robust new CRISPR / Cas system with superior performance in multiple aspects is of great significance to the development of biotechnology. Summary of the Invention
[0007] Through extensive experimentation and repeated exploration, the inventors of this application unexpectedly discovered a novel endonuclease (Cas enzyme). Based on this discovery, the inventors developed a new CRISPR / Cas system, as well as gene editing methods and nucleic acid detection methods based on this system.
[0008] Cas effector protein
[0009] On one hand, this invention provides a Cas protein, which is an effector protein in the CRISPR / Cas system. In this invention, it is referred to as Cas-sf0005 (amino acid sequence as shown in SEQ ID No. 1), Cas-sf9417 (amino acid sequence as shown in SEQ ID No. 4), Cas-sf8553 (amino acid sequence as shown in SEQ ID No. 7), and Cas-sf1846 (amino acid sequence as shown in SEQ ID No. 10). Phylogenetic analysis shows that the above-mentioned Cas proteins belong to... Similar proteins.
[0010] In one embodiment, the amino acid sequence of the Cas protein has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID No. 1, SEQ ID No. 4, SEQ ID No. 7, or SEQ ID No. 10, and substantially retains the biological function of the sequence from which it is derived.
[0011] In one embodiment, the Cas protein amino acid sequence has one or more amino acid substitutions, deletions, or additions compared to SEQ ID No. 1, SEQ ID No. 4, SEQ ID No. 7, or SEQ ID No. 10; and substantially retains the biological function of its derived sequence; the one or more amino acid substitutions, deletions, or additions include substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids.
[0012] Those skilled in the art will understand that the structure of a protein can be altered without adversely affecting its activity and function. For example, one or more conserved amino acid substitutions can be introduced into the amino acid sequence of a protein without adversely affecting the activity and / or three-dimensional structure of the protein molecule. Examples and implementations of conserved amino acid substitutions are familiar to those skilled in the art. Specifically, an amino acid residue can be substituted with another amino acid residue belonging to the same group as the site to be substituted, i.e., replacing another nonpolar amino acid residue with a nonpolar amino acid residue, replacing another polar uncharged amino acid residue with a polar uncharged amino acid residue, replacing another basic amino acid residue with a basic amino acid residue, and replacing another acidic amino acid residue with an acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. Conservative substitutions, where an amino acid is replaced by another amino acid belonging to the same group, fall within the scope of this invention, provided that the substitution does not lead to inactivation of the protein's biological activity. Therefore, the proteins of this invention can contain one or more conserved substitutions in their amino acid sequence, preferably generated by substitutions according to Table 1. Furthermore, this invention also covers proteins that also contain one or more other nonconservative substitutions, provided that such nonconservative substitutions do not significantly affect the desired function and biological activity of the proteins of this invention.
[0013] Conserved amino acid substitutions can occur at one or more predicted non-essential amino acid residues. “Non-essential” amino acid residues are those that can be altered (deleted, substituted, or replaced) without changing their biological activity, while “essential” amino acid residues are required for biological activity. A “conserved amino acid substitution” is a substitution in which an amino acid residue is replaced by an amino acid residue with a similar side chain. Amino acid substitutions can occur in the non-conserved regions of the aforementioned Cas protein. Generally, such substitutions are not performed on conserved amino acid residues, or on amino acid residues located within conserved motifs, where such residues are required for protein activity. However, those skilled in the art will understand that functional variants may have fewer conserved or non-conserved alterations in conserved regions.
[0014] Table 1
[0015] The initial residues Representative substitution Preferred replacement Ala(A) Val; Leu; Ile Val Arg(R) Lys;Gln;Asn Lys Asn(N) Gln; His; Lys; Arg Gln Asp(D) Glu Glu Cys(C) Ser Ser Gln(Q) Asn Asn Glu(E) Asp Asp Gly(G) Pro; Ala Ala His(H) Asn; Gln; Lys; Arg Arg Ile(I) Leu; Val; Met; Ala; Phe Leu Leu(L) Ile; Val; Met; Ala; Phe Ile Lys(K) Arg;Gln;Asn Arg Met(M) Leu; Phe; Ile Leu Phe(F) Leu; Val; Ile; Ala; Tyr Leu Pro(P) Ala Ala Ser(S) Thr Thr Thr(T) Ser Ser Trp(W) Tyr; Phe Tyr Tyr(Y) Trp; Phe; Thr; Ser Phe Val(V) Ile; Leu; Met; Phe; Ala Leu
[0016] As is well known in the art, one or more amino acid residues can be altered (replaced, deleted, truncated, or inserted) from the N and / or C ends of a protein while retaining its functional activity. Therefore, proteins that have one or more amino acid residues altered from their N and / or C ends while retaining their desired functional activity are also within the scope of this invention. These alterations can include those introduced by modern molecular methods such as PCR, which includes PCR amplification that alters or lengthens the protein-coding sequence by means of oligonucleotides containing amino acid-coding sequences used in the PCR amplification.
[0017] It should be recognized that proteins can be altered in various ways, including amino acid substitutions, deletions, truncations, and insertions, and methods for such operations are generally known in the art. For example, amino acid sequence variants of the aforementioned proteins can be prepared by mutating DNA. This can also be accomplished through other forms of mutagenesis and / or directed evolution, for example, using known mutagenesis, recombination, and / or shuffling methods, combined with relevant screening methods, to perform single or multiple amino acid substitutions, deletions, and / or insertions.
[0018] Those skilled in the art will understand that these minor amino acid changes in the Cas protein of this invention can occur (e.g., naturally occurring mutations) or be generated (e.g., using r-DNA technology) without loss of protein function or activity. If these mutations occur in the catalytic domain, active site, or other functional domains of the protein, the properties of the polypeptide may be altered, but the polypeptide may retain its activity. If the mutations are not located near the catalytic domain, active site, or other functional domains, a smaller impact can be expected.
[0019] Those skilled in the art can identify the essential amino acids of the Cas protein of the present invention using methods known in the art, such as localized mutagenesis, protein evolution, or bioinformatics analysis. The catalytic domains, active sites, or other functional domains of the protein can also be determined through physical structural analysis, such as by techniques like nuclear magnetic resonance, crystallography, electron diffraction, or photoaffinity labeling, combined with mutations in presumed key site amino acids.
[0020] In one embodiment, the Cas protein contains the amino acid sequence shown in SEQ ID No. 1, SEQ ID No. 4, SEQ ID No. 7, or SEQ ID No. 10.
[0021] In one embodiment, the Cas protein has the amino acid sequence shown in SEQ ID No. 1, SEQ ID No. 4, SEQ ID No. 7, or SEQ ID No. 10.
[0022] In one embodiment, the Cas protein is a derivative protein with the same biological function as a protein having the sequence shown in SEQ ID No. 1, SEQ ID No. 4, SEQ ID No. 7, or SEQ ID No. 10.
[0023] The biological functions include, but are not limited to, activities that bind to guide RNA, endonuclease activities, and activities that bind to and cleave target sequences at specific sites under the guidance of guide RNA, including but not limited to Cis cleavage activities and Trans cleavage activities.
[0024] The present invention also provides a fusion protein comprising the Cas protein as described above and other modified portions.
[0025] In one embodiment, the modified portion is selected from other proteins or peptides, detectable markers, or any combination thereof.
[0026] In one embodiment, the modified portion is selected from epitope tags, reporter gene sequences, nuclear localization signal (NLS) sequences, targeting portions, transcriptional activation domains (e.g., VP64), transcriptional repression domains (e.g., KRAB or SID domains), nuclease domains (e.g., Fok1), and domains having activities selected from: nucleotide deaminase, methyltransferase activity, demethylase, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity; and any combination thereof. The NLS sequences are well known to those skilled in the art, and examples include, but are not limited to, the SV40 large T antigen, EGL-13, c-Myc, and TUS protein.
[0027] In one embodiment, the NLS sequence is located at, near, or close to the end (e.g., N-terminus, C-terminus, or both ends) of the Cas protein of the present invention.
[0028] The epitope tag is well known to those skilled in the art, including but not limited to His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art can choose other suitable epitope tags (e.g., for purification, detection, or tracing).
[0029] The reporter gene sequences are well known to those skilled in the art, and examples include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.
[0030] In one embodiment, the fusion protein of the present invention includes a domain capable of binding to DNA molecules or intracellular molecules, such as maltose-binding protein (MBP), the DNA-binding domain (DBD) of Lex A, the DBD of GAL4, etc.
[0031] In one embodiment, the fusion protein of the present invention contains a detectable marker, such as a fluorescent dye, such as FITC or DAPI.
[0032] In one embodiment, the Cas protein of the present invention is optionally coupled, conjugated, or fused to the modified portion via a linker.
[0033] In one embodiment, the modified portion is directly connected to the N-terminus or C-terminus of the Cas protein of the present invention.
[0034] In one embodiment, the modified portion is attached to the N-terminus or C-terminus of the Cas protein of the present invention via a linker. Such linkers are well known in the art, and examples include, but are not limited to, linkers containing one or more (e.g., 1, 2, 3, 4, or 5) amino acids (e.g., Glu or Ser) or amino acid derivatives (e.g., Ahx, β-Ala, GABA, or Ava), or PEG, etc.
[0035] The Cas protein, protein derivative, or fusion protein of the present invention is not limited by the manner of its production. For example, it can be produced by genetic engineering methods (recombinant technology) or by chemical synthesis methods.
[0036] Nucleic acid of Cas protein
[0037] On the other hand, the present invention provides an isolated polynucleotide comprising:
[0038] (a) A polynucleotide sequence encoding the Cas protein or fusion protein of the present invention;
[0039] (b) Polynucleotides with sequences as shown in SEQ ID No. 2, 3, 5, 6, 8, 9, 11 or 12;
[0040] (c) A sequence having one or more base substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 11, or 12) compared to the sequence shown in SEQ ID No. 2, 3, 5, 6, 7, 8, 9, or 10 base substitutions, deletions, or additions.
[0041] (d) A polynucleotide whose nucleotide sequence has ≥80% homology (preferably ≥90%, more preferably ≥95%, most preferably ≥98%) to the sequence shown in SEQ ID No. 2, 3, 5, 6, 8, 9, 11 or 12, and encodes the polypeptide shown in SEQ ID No. 1 or SEQ ID No. 4 or SEQ ID No. 7 or SEQ ID No. 10; or,
[0042] (e) and any of the polynucleotides complementary to the polynucleotides described in (a)-(d).
[0043] In one embodiment, the nucleotide sequence described in any one of (a)-(e) is codon-optimized for expression in prokaryotic cells. In one embodiment, the nucleotide sequence described in any one of (a)-(e) is codon-optimized for expression in eukaryotic cells.
[0044] In one embodiment, the polynucleotide is preferably single-stranded or double-stranded.
[0045] Direct Repeat sequence
[0046] On the other hand, the present invention provides an engineered unidirectional repeat sequence that forms a complex with the above-mentioned Cas protein.
[0047] The unidirectional repeat sequence is linked with a guide sequence that can hybridize with the target sequence to form a guide RNA (or gRNA).
[0048] The hybridization of the target sequence with the gRNA represents at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity between the nucleic acid sequences of the target sequence and the gRNA, thereby enabling hybridization to form a complex; or represents that the nucleic acid sequences of the target sequence and the gRNA have at least 12, 15, 16, 17, 18, 19, 20, 21, 22, or more bases that can complement each other to form a complex.
[0049] In some embodiments, the repetitive sequence has at least 90% sequence identity with any one of the sequences in SEQ ID No. 13-16. In some embodiments, the repetitive sequence has one or more base substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases) compared to any one of the sequences shown in SEQ ID No. 13-16.
[0050] In some embodiments, the same-direction repeating sequence is as shown in any of SEQ ID No. 13-16.
[0051] Guide RNA (gRNA)
[0052] On the other hand, the present invention provides a gRNA comprising a first segment and a second segment; the first segment is also referred to as a "backbone region", "protein binding region", "protein binding sequence", or "direct repeat sequence"; the second segment is also referred to as a "target sequence for targeting nucleic acids", "target segment for targeting nucleic acids", or "guide sequence for targeting target sequences".
[0053] The first segment of the gRNA can interact with the Cas protein of the present invention, thereby enabling the Cas protein and gRNA to form a complex.
[0054] In a preferred embodiment, the first segment is a repeating sequence in the same direction as described above.
[0055] The target sequence or target region of the nucleic acid targeted by this invention comprises a nucleotide sequence complementary to a sequence in the target nucleic acid. In other words, the target sequence or target region of the nucleic acid targeted by this invention interacts with the target nucleic acid in a sequence-specific manner through hybridization (i.e., base pairing). Therefore, the target sequence or target region of the nucleic acid can be altered or modified to hybridize with any desired sequence within the target nucleic acid. The nucleic acid is selected from DNA or RNA.
[0056] The percentage of complementarity between the target sequence or target region of the target nucleic acid and the target sequence of the target nucleic acid may be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%).
[0057] The "backbone region," "protein-binding region," "protein-binding sequence," or "direct repeat sequence" of the gRNA of this invention can interact with CRISPR proteins (or Cas proteins). The gRNA of this invention guides the interacting Cas protein to a specific nucleotide sequence within the target nucleic acid through the targeting sequence of the target nucleic acid.
[0058] Preferably, the guide RNA comprises a first segment and a second segment in the 5' to 3' direction.
[0059] In this invention, the second segment can also be understood as a guide sequence for hybridization with the target sequence.
[0060] The gRNA of the present invention can form a complex with the Cas protein.
[0061] The gRNA of the Cas-sf0005 protein of the present invention contains a guide sequence for hybridization with a target nucleic acid, wherein the target nucleic acid includes a sequence located at the 3' end of the adjacent motif (PAM) in the prototype spacer region; the aforementioned PAM sequence is 5'TBN-3', wherein B = T / C / G and N = A / T / C / G.
[0062] The gRNA of the Cas-sf9417 protein of the present invention contains a guide sequence for hybridization with a target nucleic acid, wherein the target nucleic acid includes a sequence located at the 3' end of the adjacent motif (PAM) in the prototype spacer region; the aforementioned PAM sequence is 5'-TTN-3', wherein N = A / T / C / G.
[0063] The gRNA of the Cas-sf8553 protein of the present invention contains a guide sequence for hybridization with a target nucleic acid, wherein the target nucleic acid includes a sequence located at the 3' end of the adjacent motif (PAM) in the prototype spacer region; the aforementioned PAM sequence is 5'-TTN-3', wherein N = A / T / C / G.
[0064] carrier
[0065] The present invention also provides a carrier comprising, as described above, a Cas protein, an isolated nucleic acid molecule or a polynucleotide; preferably, it further comprises a regulatory element operatively linked thereto.
[0066] In one embodiment, the regulatory element is selected from one or more of the following: enhancers, transposons, promoters, terminators, leader sequences, polyadenylation sequences, and marker genes.
[0067] In one embodiment, the vector includes a cloning vector, an expression vector, a shuttle vector, and an integration vector.
[0068] In some implementations, the vectors included in the system are viral vectors (e.g., retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated vectors, and herpes simplex vectors), and may also be plasmids, viruses, granules, bacteriophages, etc., which are well known to those skilled in the art.
[0069] CRISPR system
[0070] The present invention provides an engineered, non-naturally occurring vector system, or a CRISPR-Cas system, comprising a Cas protein or a nucleic acid sequence encoding the Cas protein and a nucleic acid encoding one or more guide RNAs.
[0071] In one embodiment, the nucleic acid sequence encoding the Cas protein and the nucleic acid encoding one or more guide RNAs are artificially synthesized.
[0072] In one embodiment, the nucleic acid sequence encoding the Cas protein and the nucleic acid encoding one or more guide RNAs do not coexist naturally.
[0073] The one or more guide RNAs target one or more target sequences in the cell. The one or more target sequences hybridize to the genomic loci of a DNA molecule encoding one or more gene products and guide the Cas protein to the genomic locus of the DNA molecule of the one or more gene products. Once the Cas protein reaches the target sequence location, it modifies, edits, or cuts the target sequence, thereby altering or modifying the expression of the one or more gene products.
[0074] The cells of this invention include one or more of animals, plants, or microorganisms.
[0075] In some embodiments, the Cas protein is codon-optimized for expression in cells.
[0076] In some embodiments, the Cas protein cleaves one or both strands at the target sequence location.
[0077] In some embodiments, the Cas protein cleaves the complementary and / or non-complementary strands of the target nucleic acid under the mediation of gRNA.
[0078] Preferably, the Cas protein simultaneously cleaves the complementary and non-complementary strands of the target nucleic acid.
[0079] In this invention, the gRNA guides the Cas protein to recognize and bind to the complementary strand, and the non-complementary strand is a nucleic acid strand paired with the complementary strand. The PAM sequence is located on the non-complementary strand, and the complementary strand contains a PAM complementary sequence that pairs with the aforementioned PAM sequence.
[0080] In one embodiment, the cleavage site of Cas-sf0005 on the complementary strand of the target sequence is between the 21st and 22nd nt of the 5' end of the PAM complementary sequence, and the cleavage site of Cas-sf0005 on the non-complementary strand of the target sequence is between the 15th and 16th nt of the 3' end of the PAM sequence. The gRNA guides the Cas-sf0005 protein to recognize and bind to the aforementioned complementary strand, wherein the aforementioned non-complementary strand is the DNA strand paired with the complementary strand.
[0081] In one embodiment, the cleavage site of Cas-sf9417 on the complementary strand of the target sequence is between the 21st and 22nd nt of the 5' end of the PAM complementary sequence, and the cleavage site of Cas-sf9417 on the non-complementary strand of the target sequence is between the 14th and 16th nt of the 3' end of the PAM sequence. The gRNA guides the Cas-sf9417 protein to recognize and bind to the aforementioned complementary strand, wherein the aforementioned non-complementary strand is the DNA strand paired with the complementary strand.
[0082] In one embodiment, the cleavage site of the complementary strand of the target sequence by Cas-sf8553 is between the 21st and 23rd nt of the 5' end of the PAM complementary sequence, and the cleavage site of the non-complementary strand of the target sequence by Cas-sf8553 is between the 14th and 15th nt of the 3' end of the PAM sequence. The gRNA guides the Cas-sf8553 protein to recognize and bind to the aforementioned complementary strand, wherein the aforementioned non-complementary strand is the DNA strand paired with the complementary strand.
[0083] The present invention also provides an engineered, non-naturally occurring carrier system, which may include one or more carriers, the one or more carriers comprising:
[0084] a) A first regulatory element, which is operatively linked to the gRNA.
[0085] b) A second regulatory element operatively linked to the Cas protein;
[0086] Components (a) and (b) are located on the same or different carriers in the system.
[0087] The first and second regulatory elements include promoters (e.g., constitutive or inducible promoters), enhancers (e.g., 35S promoters or 35S enhanced promoters), internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and polyU sequences).
[0088] In some implementations, the vector in the system is a viral vector (e.g., a retroviral vector, lentiviral vector, adenovirus vector, adeno-associated vector, and herpes simplex vector), or it can be a plasmid, virus, granule, bacteriophage, or other type known to those skilled in the art.
[0089] In some embodiments, the system provided herein is a delivery system. In some embodiments, the delivery system is a nanoparticle, liposome, exosome, microbubble, or gene gun.
[0090] In one embodiment, the target sequence is a DNA or RNA sequence derived from prokaryotic or eukaryotic cells. In another embodiment, the target sequence is a non-naturally occurring DNA or RNA sequence.
[0091] In one embodiment, the target sequence is present within the cell. In another embodiment, the target sequence is present in the cell nucleus or cytoplasm (e.g., organelles). In one embodiment, the cell is a eukaryotic cell. In other embodiments, the cell is a prokaryotic cell.
[0092] In one embodiment, the Cas protein is linked to one or more NLS sequences. In one embodiment, the fusion protein comprises one or more NLS sequences. In one embodiment, the NLS sequence is linked to the N-terminus or C-terminus of the protein. In one embodiment, the NLS sequence is fused to the N-terminus or C-terminus of the protein.
[0093] On the other hand, the present invention relates to an engineered CRISPR system comprising the aforementioned Cas protein and one or more guide RNAs, wherein the guide RNA comprises a homologous repeat sequence and a spacer sequence capable of hybridizing with a target nucleic acid, and the Cas protein is capable of binding the guide RNA and targeting a target nucleic acid sequence complementary to the spacer sequence.
[0094] In one embodiment, the Cas enzyme is Cas-sf0005, the target nucleic acid is DNA (preferably double-stranded DNA), the target nucleic acid is located at the 3' end of the protospacer adjacent motif (PAM), and the PAM has the sequence shown in 5'TBN-3', wherein B = T / C / G and N = A / T / C / G.
[0095] In one embodiment, when the Cas enzyme is Cas-sf9417 or Cas-sf8553, and the target nucleic acid is DNA (preferably double-stranded DNA), the target nucleic acid is located at the 3' end of the adjacent motif (PAM) of the original spacer sequence, and the PAM is 5'-TTN-3', where N = A / T / C / G.
[0096] Protein-nucleic acid complexes / compositions
[0097] On the other hand, the present invention provides a complex or composition comprising:
[0098] (i) Protein components selected from: the aforementioned Cas proteins, derived proteins, or fusion proteins, and any combination thereof; and
[0099] (ii) A nucleic acid component comprising (a) a guide sequence capable of hybridizing with a target sequence; and (b) a unidirectional repeat sequence capable of binding to the Cas protein of the present invention.
[0100] The protein components and nucleic acid components combine to form a complex.
[0101] In one embodiment, the nucleic acid component is a guide RNA in a CRISPR-Cas system.
[0102] In one embodiment, the complex or composition is non-natural or modified. In one embodiment, at least one component of the complex or composition is non-natural or modified. In one embodiment, the first component is non-natural or modified; and / or, the second component is non-natural or modified.
[0103] Activated CRISPR complex
[0104] On the other hand, the present invention also provides an activated CRISPR complex comprising: (1) a protein component selected from the Cas protein, derived protein, or fusion protein of the present invention, and any combination thereof; (2) a gRNA comprising (a) a guide sequence capable of hybridizing with a target sequence; and (b) a homologous repeat sequence capable of binding to the Cas protein of the present invention; and (3) a target sequence bound to the gRNA. Preferably, the binding is a binding between the target sequence of the target nucleic acid on the gRNA and the target nucleic acid.
[0105] The term “activated CRISPR complex”, “activated complex”, or “ternary complex” used in this article refers to the complex formed by the binding or modification of the Cas protein, gRNA, and target nucleic acid in the CRISPR system.
[0106] The Cas protein and gRNA of this invention can form a binary complex, which is activated upon binding to a nucleic acid substrate to form an activated CRISPR complex. The nucleic acid substrate is complementary to the spacer sequence (or, in other words, the guide sequence for hybridization with the target nucleic acid) in the gRNA. In some embodiments, the spacer sequence of the gRNA perfectly matches the target substrate. In other embodiments, the spacer sequence of the gRNA partially (continuously or discontinuously) matches the target substrate.
[0107] In a preferred embodiment, the activated CRISPR complex can exhibit side-branch nuclease cleavage activity, which refers to the non-specific or random cleavage activity of the activated CRISPR complex on single-stranded nucleic acids, also known in the art as trans cleavage activity.
[0108] Delivery and delivery composition
[0109] The Cas proteins, gRNAs, fusion proteins, nucleic acid molecules, vectors, systems, complexes, and compositions of the present invention can be delivered by any method known in the art. Such methods include, but are not limited to, electroporation, lipid transfection, nuclear transfection, microinjection, acoustic pore effect, gene gun, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendritic transfection, heat shock transfection, nuclear transfection, magnetic transfection, lipid transfection, puncture transfection, optical transfection, reagent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial viruses, etc.
[0110] Therefore, in another aspect, the present invention provides a delivery composition comprising a delivery vector and selected from one or more of the following: the Cas protein, fusion protein, nucleic acid molecule, vector, system, complex and composition of the present invention.
[0111] In one embodiment, the delivery carrier is a particle.
[0112] In one embodiment, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, gene guns, or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).
[0113] host cells
[0114] The present invention also relates to an in vitro, ex vivo, or in vivo cell or cell line or its progeny comprising: the Cas protein, fusion protein, nucleic acid molecule, protein-nucleic acid complex, activated CRISPR complex, vector, or delivery composition of the present invention.
[0115] In some implementations, the cell is a prokaryotic cell.
[0116] In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a non-human mammalian cell, such as cells of non-human primates, cattle, sheep, pigs, dogs, monkeys, rabbits, or rodents (e.g., rats or mice). In some embodiments, the cell is a non-mammalian eukaryotic cell, such as cells of poultry (e.g., chickens), fish, or crustaceans (e.g., clams, shrimp). In some embodiments, the cell is a plant cell, such as cells of monocotyledonous or dicotyledonous plants, or cells of cultivated plants or food crops such as cassava, corn, sorghum, soybeans, wheat, oats, or rice, such as algae, trees, or productive plants, fruits, or vegetables (e.g., trees such as citrus trees, nut trees; nightshade plants, cotton, tobacco, tomatoes, grapes, coffee, cocoa, etc.).
[0117] In some implementations, the cell is a stem cell or stem cell line.
[0118] In some cases, the host cells of the present invention contain genetic or genomic modifications that are not present in their wild type.
[0119] Gene editing methods and applications
[0120] The Cas protein, nucleic acid, the above-described composition, the above-described CRISPR / Cas system, the above-described vector system, the above-described delivery composition, or the above-described activated CRISPR complex or the above-described host cell of the present invention can be used for any or more of the following purposes: targeting and / or editing target nucleic acids; cleaving double-stranded DNA, single-stranded DNA, or single-stranded RNA; non-specifically cleaving and / or degrading side-branched nucleic acids; non-specifically cleaving single-stranded nucleic acids; nucleic acid detection; detecting nucleic acids in target samples; specifically editing double-stranded nucleic acids; base editing double-stranded nucleic acids; base editing single-stranded nucleic acids. In other embodiments, they can also be used to prepare reagents or kits for any or more of the above purposes.
[0121] The present invention also provides the use of the above-mentioned Cas protein, nucleic acid, the above-mentioned composition, the above-mentioned CIRSPR / Cas system, the above-mentioned vector system, the above-mentioned delivery composition, or the above-mentioned activated CRISPR complex in gene editing, gene targeting, or gene cutting; or, in the preparation of reagents or kits for gene editing, gene targeting, or gene cutting.
[0122] In one embodiment, the gene editing, gene targeting, or gene cutting is performed intracellularly and / or extracellularly.
[0123] The present invention also provides a method for editing, targeting, or cleaving a target nucleic acid, the method comprising contacting the target nucleic acid with the aforementioned Cas protein, nucleic acid, the aforementioned composition, the aforementioned CIRSPR / Cas system, the aforementioned vector system, the aforementioned delivery composition, or the aforementioned activated CRISPR complex. In one embodiment, the method comprises editing, targeting, or cleaving the target nucleic acid intracellularly or extracellularly.
[0124] The gene editing or editing of target nucleic acids includes modifying genes, knocking out genes, altering the expression of gene products, repairing mutations, and / or inserting polynucleotides, and gene mutations.
[0125] The editing can be performed in prokaryotic and / or eukaryotic cells.
[0126] On the other hand, the present invention also provides the use of the above-mentioned Cas protein, nucleic acid, the above-mentioned composition, the above-mentioned CIRSPR / Cas system, the above-mentioned carrier system, the above-mentioned delivery composition or the above-mentioned activated CRISPR complex in nucleic acid detection, or in the preparation of reagents or kits for nucleic acid detection.
[0127] On the other hand, the present invention also provides a method for cleaving single-stranded nucleic acids, the method comprising contacting a nucleic acid population with the aforementioned Cas protein and gRNA, wherein the nucleic acid population comprises a target nucleic acid and a plurality of non-target single-stranded nucleic acids, and the Cas protein cleaving the plurality of non-target single-stranded nucleic acids.
[0128] The gRNA can bind to the Cas protein.
[0129] The gRNA can target the target nucleic acid.
[0130] The contact can be outside the body, outside the body, or inside the cells within the body.
[0131] Preferably, the cleavage of the single-stranded nucleic acid is non-specific.
[0132] On the other hand, the present invention also provides the use of the above-described Cas protein, nucleic acid, the above-described composition, the above-described CIRSPR / Cas system, the above-described carrier system, the above-described delivery composition, or the above-described activated CRISPR complex in the non-specific cleavage of single-stranded nucleic acids, or in the preparation of reagents or kits for the non-specific cleavage of single-stranded nucleic acids.
[0133] On the other hand, the present invention also provides a kit for gene editing, gene targeting or gene cutting, the kit comprising the above-mentioned Cas protein, gRNA, nucleic acid, the above-mentioned composition, the above-mentioned CIRSPR / Cas system, the above-mentioned vector system, the above-mentioned delivery composition, the above-mentioned activated CRISPR complex or the above-mentioned host cell.
[0134] On the other hand, the present invention also provides a kit for detecting target nucleic acids in a sample, the kit comprising: (a) a Cas protein, or a nucleic acid encoding the Cas protein; (b) a guide RNA, or a nucleic acid encoding the guide RNA, or a precursor RNA containing the guide RNA, or a nucleic acid encoding the precursor RNA; and (c) a single-stranded nucleic acid detector that is single-stranded and does not hybridize with the guide RNA.
[0135] As is known in the art, precursor RNA can be cleaved or processed into the aforementioned mature guide RNA.
[0136] On the other hand, the invention provides the use of the above-mentioned Cas protein, nucleic acid, composition, CRISPR / Cas system, vector system, delivery composition, activated CRISPR complex, or host cell in the preparation of formulations or kits, wherein the formulations or kits are used for:
[0137] (i) Gene or genome editing;
[0138] (ii) Target nucleic acid detection and / or diagnosis;
[0139] (iii) Editing target sequences in target loci to modify biological or non-human organisms;
[0140] (iv) Treatment of the disease;
[0141] (iv) Targeting target genes.
[0142] Preferably, the above-mentioned gene or genome editing is performed intracellularly or extracellularly.
[0143] Preferably, the target nucleic acid detection and / or diagnosis is performed in vitro.
[0144] Preferably, the treatment of the disease is to treat symptoms caused by defects in the target sequence at the target locus.
[0145] In another aspect, the present invention provides a method for detecting target nucleic acids in a sample, the method comprising contacting the sample with the Cas protein, gRNA (guide RNA) and a single-stranded nucleic acid detector, the gRNA including a region binding to the Cas protein and a guide sequence for hybridization with the target nucleic acid; detecting a detectable signal generated by the single-stranded nucleic acid detector by the Cas protein cleaving the single-stranded nucleic acid detector, thereby detecting the target nucleic acid; the single-stranded nucleic acid detector not hybridizing with the gRNA.
[0146] Methods for specifically modifying target nucleic acids
[0147] On the other hand, the present invention also provides a method for specifically modifying target nucleic acids, the method comprising: contacting the target nucleic acid with the above-mentioned Cas protein, nucleic acid, the above-mentioned composition, the above-mentioned CIRSPR / Cas system, the above-mentioned vector system, the above-mentioned delivery composition or the above-mentioned activated CRISPR complex.
[0148] This specific modification can occur in vivo or in vitro.
[0149] This specific modification can occur either inside or outside the cell.
[0150] In some cases, the cells are selected from prokaryotic or eukaryotic cells, such as animal cells, plant cells, or microbial cells.
[0151] In one embodiment, the modification refers to a break in the target sequence, such as a single-strand / double-strand break in DNA or a single-strand break in RNA.
[0152] In some cases, the method further includes contacting the target nucleic acid with a donor polynucleotide, wherein the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide, or a portion of a copy of the donor polynucleotide is integrated into the target nucleic acid.
[0153] In one embodiment, the modification further includes inserting an editing template (e.g., exogenous nucleic acid) into the break.
[0154] In one embodiment, the method further includes contacting the editing template with the target nucleic acid or delivering it to a cell containing the target nucleic acid. In this embodiment, the method repairs the broken target gene by homologous recombination with a foreign template polynucleotide; in some embodiments, the repair results in a mutation, including the insertion, deletion, or substitution of one or more nucleotides of the target gene; in other embodiments, the mutation results in a change in one or more amino acids in a protein expressed from a gene containing the target sequence.
[0155] Detection (non-specific cutting)
[0156] On the other hand, the present invention provides a method for detecting target nucleic acids in a sample, the method comprising contacting the sample with the above-mentioned Cas protein, nucleic acid, the above-mentioned composition, the above-mentioned CIRSPR / Cas system, the above-mentioned carrier system, the above-mentioned delivery composition or the above-mentioned activated CRISPR complex and a single-stranded nucleic acid detector; detecting a detectable signal generated by the single-stranded nucleic acid detector by the Cas protein cleaving the single-stranded nucleic acid, thereby detecting the target nucleic acid.
[0157] In this invention, the target nucleic acid includes ribonucleotides or deoxyribonucleotides; including single-stranded nucleic acids and double-stranded nucleic acids, such as single-stranded DNA, double-stranded DNA, single-stranded RNA, and double-stranded RNA.
[0158] In one embodiment, the target nucleic acid is derived from samples such as viruses, bacteria, microorganisms, soil, water sources, humans, animals, and plants. Preferably, the target nucleic acid is a product enriched or amplified by methods such as PCR, NASBA, RPA, SDA, LAMP, HAD, NEAR, MDA, RCA, LCR, and RAM.
[0159] In one embodiment, the target nucleic acid is viral nucleic acid, bacterial nucleic acid, disease-related specific nucleic acid, such as specific mutation sites or SNP sites, or nucleic acids that differ from controls; preferably, the virus is a plant virus or animal virus, such as papillomavirus, hepatocyte DNA virus, herpesvirus, adenovirus, poxvirus, parvovirus, or coronavirus; preferably, the virus is a coronavirus, preferably SARS, SARS-CoV2 (COVID-19), HCoV-229E, HCoV-OC43, HCoV-NL63, HCoV-HKU1, or Mers-Cov.
[0160] In this invention, the gRNA has at least a 50% match with the target sequence on the target nucleic acid, preferably at least 60%, preferably at least 70%, preferably at least 80%, and preferably at least 90%.
[0161] In one implementation, when the target sequence contains one or more signature sites (such as specific mutation sites or SNPs), the signature sites perfectly match the gRNA.
[0162] In one embodiment, the detection method may include one or more gRNAs with different guide sequences that target different target sequences.
[0163] In this invention, the single-stranded nucleic acid detector includes, but is not limited to, single-stranded DNA, single-stranded RNA, DNA-RNA hybrids, nucleic acid analogs, base modifiers, and single-stranded nucleic acid detectors containing base-free spacers; "nucleic acid analogs" include, but are not limited to: locked nucleic acids, bridging nucleic acids, morpholine nucleic acids, ethylene glycol nucleic acids, hexitol nucleic acids, threonine nucleic acids, arabinose nucleic acids, 2'-oxymethyl RNA, 2'-methoxyacetyl RNA, 2'-fluoroRNA, 2'-aminoRNA, 4'-thioRNA, and combinations thereof, including optional ribonucleotide or deoxyribonucleotide residues.
[0164] In this invention, the detectable signal is achieved through the following methods: vision-based detection, sensor-based detection, color detection, fluorescence signal-based detection, gold nanoparticle-based detection, fluorescence polarization, and colloidal phase transition / decomposition.
[0165] Dispersion, electrochemical detection and semiconductor-based detection.
[0166] In this invention, preferably, a fluorescent group and a quenching group are respectively disposed at both ends of the single-stranded nucleic acid detector, so that the single-stranded nucleic acid detector can exhibit a detectable fluorescent signal after being cleaved. The fluorescent group is selected from one or any combination of FAM, FITC, VIC, JOE, TET, CY3, CY5, ROX, Texas Red, or LC RED460; the quenching group is selected from one or any combination of BHQ1, BHQ2, BHQ3, Dabcy1, or Tamra.
[0167] In other embodiments, different labeling molecules are set at the 5' and 3' ends of the single-stranded nucleic acid detector, and the colloidal gold test results of the single-stranded nucleic acid detector before and after being cleaved by the Cas protein are detected by colloidal gold detection. The single-stranded nucleic acid detector will show different color development results on the colloidal gold detection line and control line before and after being cleaved by the Cas protein.
[0168] In some implementations, the method for detecting target nucleic acids may further include comparing the level of a detectable signal with a reference signal level, and determining the amount of target nucleic acid in the sample based on the level of the detectable signal.
[0169] In some implementations, the method for detecting target nucleic acids may also include using RNA reporter nucleic acids and DNA reporter nucleic acids (e.g., fluorescence color) on different channels and determining the level of a detectable signal by measuring the signal levels of the RNA and DNA reporter molecules and by measuring the amount of target nucleic acid in the RNA and DNA reporter molecules, sampling based on the combined (e.g., using minimum or product) level of the detectable signal.
[0170] In one embodiment, the target gene is present within the cell.
[0171] In one embodiment, the cell is a prokaryotic cell.
[0172] In one embodiment, the cell is a eukaryotic cell.
[0173] In one embodiment, the cell is an animal cell.
[0174] In one embodiment, the cell is a human cell.
[0175] In one embodiment, the cell is a plant cell, such as the cell of a cultivated plant (e.g., cassava, corn, sorghum, wheat, or rice), algae, tree, or vegetable.
[0176] In one embodiment, the target gene is present in an in vitro nucleic acid molecule (e.g., a plasmid).
[0177] In one embodiment, the target gene is present in a plasmid.
[0178] Terminology Definition
[0179] In this invention, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the operational steps used herein, such as molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics, and recombinant DNA, are all conventional steps widely used in their respective fields. To better understand this invention, definitions and explanations of relevant terms are provided below.
[0180] Cas protein
[0181] In this invention, Cas protein, Cas enzyme, and Cas effector protein can be used interchangeably; the inventors have for the first time discovered and identified a Cas effector protein having an amino acid sequence selected from the following:
[0182] (i) SEQ ID No. 1 or SEQ ID No. 4 or SEQ ID No. 7 or SEQ ID No. 10;
[0183] (ii) A sequence having one or more amino acid substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids) compared to the sequence shown in SEQ ID No. 1, SEQ ID No. 4, SEQ ID No. 7, or SEQ ID No. 10; or
[0184] (iii) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the sequence shown in SEQ ID No. 1, SEQ ID No. 4, SEQ ID No. 7, or SEQ ID No. 10.
[0185] Nucleic acid cleavage or cleavage of nucleic acids as used herein includes: DNA or RNA breaks in target nucleic acids produced by the Cas enzyme described herein (Cis cleavage), and DNA or RNA breaks in side-branched nucleic acid substrates (single-stranded nucleic acid substrates) (i.e., nonspecific or non-targeted, trans cleavage). In some embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break or a single-stranded RNA break.
[0186] CRISPR system
[0187] As used herein, the terms “regularly clustered short palindromic repeats (CRISPR)-CRISPR-related (Cas) (CRISPR-Cas) system” or “CRISPR system” are used interchangeably and have the meaning commonly understood by those skilled in the art, which typically includes transcripts or other elements relating to the expression of CRISPR-related (“Cas”) genes, or transcripts or other elements capable of directing the activity of said Cas genes.
[0188] CRISPR / Cas complex
[0189] As used herein, the term “CRISPR / Cas complex” refers to a complex formed by the binding of guide RNA or mature crRNA to the Cas protein, which contains a guide sequence that hybridizes to the target sequence and a homologous repeat sequence that binds to the Cas protein. This complex is capable of recognizing and cleaving polynucleotides that hybridize with the guide RNA or mature crRNA.
[0190] Guide RNA (gRNA)
[0191] As used herein, the terms “guide RNA (gRNA),” “mature crRNA,” and “guide sequence” are used interchangeably and have the meanings commonly understood by those skilled in the art. Generally, guide RNA may comprise a direct repeat sequence and a guide sequence, or consist substantially of or composed of a direct repeat sequence and a guide sequence.
[0192] In some cases, the guide sequence is any polynucleotide sequence that is sufficiently complementary to the target sequence to hybridize with the target sequence and guide the specific binding of the CRISPR / Cas complex to the target sequence. In one embodiment, the complementarity between the guide sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% when optimal alignment is achieved. Determining the optimal alignment is within the capabilities of a person skilled in the art. For example, publicly available and commercially available alignment algorithms and programs exist, such as, but not limited to, ClustalW, the Smith-Waterman algorithm in MATLAB, Bowtie, Geneious, Biopython, and SeqMan.
[0193] target sequence
[0194] A "target sequence" refers to a polynucleotide targeted by a guide sequence in the gRNA, such as a sequence complementary to that guide sequence, where hybridization between the target and guide sequences will promote the formation of a CRISPR / Cas complex (including the Cas protein and gRNA). Perfect complementarity is not required, as long as sufficient complementarity exists to induce hybridization and promote the formation of a CRISPR / Cas complex.
[0195] The target sequence can contain any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located inside or outside the cell. In some cases, the target sequence is located in the cell nucleus or cytoplasm. In some cases, the target sequence may be located in an organelle of a eukaryotic cell, such as a mitochondrion or chloroplast. The sequence or template that can be used for recombination into a target locus containing the target sequence is referred to as an "edit template," "edit polynucleotide," or "edit sequence." In one embodiment, the edit template is a foreign nucleic acid. In one embodiment, the recombination is homologous recombination.
[0196] In this invention, the "target sequence," "target polynucleotide," or "target nucleic acid" can be any endogenous or exogenous polynucleotide for a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or useless DNA). In some cases, the target sequence should be associated with a protospacer adjacent motif (PAM).
[0197] Single-stranded nucleic acid detector
[0198] The single-stranded nucleic acid detector described in this invention refers to a sequence containing 2-200 nucleotides, preferably 2-150 nucleotides, more preferably 3-100 nucleotides, more preferably 3-30 nucleotides, more preferably 4-20 nucleotides, and even more preferably 5-15 nucleotides. It is preferably a single-stranded DNA molecule, a single-stranded RNA molecule, or a single-stranded DNA-RNA hybrid.
[0199] The single-stranded nucleic acid detector has different reporter groups or label molecules at both ends. When it is in its initial state (i.e., uncut state), it does not present a reporter signal. When the single-stranded nucleic acid detector is cut, it presents a detectable signal, that is, it shows a detectable difference after cutting compared to before cutting.
[0200] In one embodiment, the reporter group or labeled molecule includes a fluorescent group and a quencher group, wherein the fluorescent group is selected from one or any of FAM, FITC, VIC, JOE, TET, CY3, CY5, ROX, Texas Red or LCRED460; and the quencher group is selected from one or any of BHQ1, BHQ2, BHQ3, Dabcy1 or Tamra.
[0201] In one embodiment, the single-stranded nucleic acid detector has a first molecule (such as FAM or FITC) connected to the 5' end and a second molecule (such as biotin) connected to the 3' end. The reaction system containing the single-stranded nucleic acid detector is used in conjunction with a flow strip to detect target nucleic acids (preferably, colloidal gold detection). The flow strip is designed with two capture lines, with an antibody binding the first molecule (i.e., the first molecule antibody) at the sample contact end (colloidal gold), an antibody binding the first molecule antibody at the first line (control line), and an antibody binding the second molecule antibody (i.e., the second molecule antibody, such as avidin) at the second line (test line). As the reaction flows along the strip, the first molecule antibody binds to the first molecule, carrying cleaved or uncleaved oligonucleotides to the capture lines. Cleaved reporters bind to the first molecule antibody at the first capture line, while uncleaved reporters bind to the second molecule antibody at the second capture line. The binding of the reporter group at each line results in a strong readout / signal (e.g., color). As more reporters are cleaved, more signal accumulates at the first capture line, and less signal appears at the second line. In some aspects, the present invention relates to the use of flow strips as described herein for the detection of nucleic acids. In some aspects, the present invention relates to methods for detecting nucleic acids using flow strips as defined herein, such as (lateral)flow assays or (lateral)flow immunochromatographic assays. In some aspects, the molecules in the single-stranded nucleic acid detector may be interchanged or their positions altered, provided that their reporting principle is the same as or similar to that of the present invention, and any modifications thereof are also included in the present invention.
[0202] The detection method described in this invention can be used for the quantitative detection of target nucleic acids. The quantitative detection index can be determined based on the signal strength of the reporter group, such as the luminescence intensity of the fluorescent group or the width of the colored band.
[0203] wild type
[0204] As used herein, the term “wildtype” has the meaning commonly understood by those skilled in the art as referring to the typical form of an organism, strain, or gene, or the characteristic that distinguishes it from mutant or variant forms when it exists in nature, is separable from its natural source and has not been intentionally modified by humans.
[0205] Derivatization
[0206] As used herein, the term "derivation" refers to the chemical modification of an amino acid, polypeptide, or protein in which one or more substituents are covalently linked to the amino acid, polypeptide, or protein. Substituents may also be referred to as side chains.
[0207] A derivatized protein is a derivative of the original protein. Generally, the derivatization of a protein does not adversely affect its desired activity (e.g., activity to bind to guide RNA, endonuclease activity, activity to bind to and cleave a target sequence at a specific site under the guidance of guide RNA). In other words, the derivative of a protein has the same activity as the original protein.
[0208] Derivatized proteins
[0209] Also known as "protein derivatives," these are modified forms of proteins, where one or more amino acids of the protein may be deleted, inserted, modified, and / or substituted.
[0210] Not naturally occurring
[0211] As used herein, the terms “non-naturally occurring” or “engineered” are used interchangeably and indicate artificial involvement. When these terms are used to describe nucleic acid molecules or peptides, they indicate that the nucleic acid molecule or peptide is at least substantially free from at least one other component bound to it, either naturally occurring or found in nature.
[0212] Orthologue (ortholog)
[0213] As used herein, the term "orthologue" has the meaning commonly understood by those skilled in the art. As further guidance, an "orthologue" of a protein, as described herein, refers to a protein belonging to a different species that performs the same or similar function as the protein that is its orthologue.
[0214] identity
[0215] As used herein, the term "identity" refers to the sequence matching between two polypeptides or two nucleic acids. Two compared sequences are identical at a position when the same base or amino acid monomeric subunit occupies the same location (e.g., a position in each of two DNA molecules is occupied by adenine, or a position in each of two polypeptides is occupied by lysine). The "percentage identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions compared × 100. For example, if six out of ten positions in two sequences match, then the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGGTT share 50% identity (three out of six positions match). Typically, two sequences are compared to produce the maximum identity. Such comparisons can be made using methods readily available, for example, computer programs such as the Align program (DNAstar, Inc.) Needleman et al. (1970) J. Mol. Biol. 48: 443-453. The percentage identity between two amino acid sequences can also be determined using the algorithm of E. Meyers and W. Miller (Comput. Appl Biosci., 4:11-17 (1988)) integrated into the ALIGN program (version 2.0), which uses a PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4. Alternatively, the percentage identity between two amino acid sequences can be determined using the Needleman and Wunsch algorithm (J MoIBiol. 48:444-453 (1970)) in the GAP program integrated into the GCG software package (available at www.gcg.com), which uses a Blossum 62 matrix or a PAM250 matrix, along with gap weights of 16, 14, 12, 10, 8, 6, or 4, and length weights of 1, 2, 3, 4, 5, or 6.
[0216] carrier
[0217] The term "vector" refers to a nucleic acid molecule capable of delivering another nucleic acid molecule linked to it. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules including one or more free ends, or without free ends (e.g., circular); nucleic acid molecules including DNA, RNA, or both; and a wide variety of other polynucleotides known in the art. A vector can be introduced into a host cell through transformation, transduction, or transfection, thereby enabling the expression of its carried genetic material elements in the host cell. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including proteins, fusion proteins, isolated nucleic acid molecules, etc., as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector may contain a variety of elements controlling expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. Additionally, the vector may contain a replication initiation site.
[0218] One type of vector is a "plasmid," which is a circular double-stranded DNA loop into which another DNA fragment can be inserted, for example, using standard molecular cloning techniques.
[0219] Another type of vector is the viral vector, in which a virus-derived DNA or RNA sequence is present in a vector used to package the virus (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses). Viral vectors also contain polynucleotides carried by the virus used for transfection into a host cell. Some vectors (e.g., bacterial vectors with bacterial origins of replication and episodic mammalian vectors) are capable of autonomous replication in the host cells into which they are introduced.
[0220] Other vectors (e.g., non-attachment mammalian vectors) integrate into the host cell's genome upon introduction and thereby replicate along with the host genome. Furthermore, some vectors are capable of directing the expression of genes they are operatively linked to. Such vectors are referred to herein as "expression vectors."
[0221] host cells
[0222] As used herein, the term “host cell” refers to a cell that can be used to introduce a vector, including but not limited to prokaryotic cells such as Escherichia coli or Bacillus subtilis, and eukaryotic cells such as microbial cells, fungal cells, animal cells, and plant cells.
[0223] Those skilled in the art will understand that the design of expression vectors can depend on factors such as the selection of host cells to be transformed and the desired expression level.
[0224] Control element
[0225] As used herein, the term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences), for which detailed descriptions can be found in Goeddel, *Gene Expression Technology: Methods in Enzymology*, 185, Academic Press, San Diego, California (1990). In some cases, regulatory elements include those sequences that direct constitutive expression of a nucleotide sequence in many types of host cells and those sequences that direct expression of that nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes). In some cases, regulatory elements can also be directed to express in a time-dependent manner (such as in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue or cell type specific. In some cases, the term "regulatory element" covers enhancer elements such as WPRE; CMV enhancer; R-U5' fragment in the LTR of HTLV-I (Mol. Cell. Biol., Vol. 8(1), pp. 466-472, 1988); SV40 enhancer; and intron sequence between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), pp. 1527-31, 1981).
[0226] promoter
[0227] As used herein, the term "promoter" has the meaning known to those skilled in the art, referring to a non-coding nucleotide sequence located upstream of a gene that initiates the expression of a downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell under most or all physiological conditions of the cell. An inducible promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when an inducer corresponding to the promoter is present in the cell. A tissue-specific promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when the cell is a cell of the tissue type corresponding to that promoter.
[0228] NLS
[0229] A “nuclear localization signal” or “nuclear localization sequence” (NLS) is an amino acid sequence that “tags” a protein to allow it to be transported to the nucleus via nuclear transport; that is, a protein with an NLS is transported to the nucleus. Typically, an NLS contains positively charged Lys or Arg residues exposed on the protein surface. Exemplary nuclear localization sequences include, but are not limited to, NLS from the following: SV40 large T antigen, EGL-13, c-Myc, and TUS protein. In some embodiments, the NLS contains the PKKKRKV sequence. In some embodiments, the NLS contains the AVKRPAATKKAGQAKKKKLD sequence. In some embodiments, the NLS contains the PAAKRVKLD sequence. In some embodiments, the NLS contains the MSRRRKANPTKLSENAKKLAKEVEN sequence. In some embodiments, the NLS contains the KLKIKRPVK sequence. Other nuclear localization sequences include, but are not limited to, the acidic M9 domain of hnRNP A1, the KIPIK sequence in the yeast transcriptional repressor Matα2, and PY-NLS.
[0230] Operable connection
[0231] As used herein, the term “operably linked” is intended to mean that the nucleotide sequence of interest is linked to one or more regulatory elements in a manner that allows the expression of that nucleotide sequence (e.g., in an in vitro transcription / translation system or in the host cell when the vector is introduced into the host cell).
[0232] Complementarity
[0233] As used herein, the term "complementarity" refers to the ability of a nucleic acid to form one or more hydrogen bonds with another nucleic acid sequence via conventional Watson-Crick or other non-conventional types. The percentage of complementarity indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Complete complementarity" means that all consecutive residues in a nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in a second nucleic acid sequence. As used herein, “substantially complementary” refers to a complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% in a region having 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids hybridizing under stringent conditions.
[0234] Strict conditions
[0235] As used herein, “strict conditions” for hybridization refer to conditions under which a nucleic acid complementary to the target sequence hybridizes primarily with the target sequence and substantially does not hybridize to non-target sequences. Strict conditions are typically sequence-dependent and vary depending on many factors. Generally, the longer the sequence, the higher the temperature at which it specifically hybridizes to its target sequence.
[0236] Hybridization
[0237] The terms “hybridization” or “complementary” or “substantially complementary” refer to nucleic acids (such as RNA, DNA) containing nucleotide sequences that enable them to bind non-covalently, that is, to form base pairs and / or G / U base pairs with another nucleic acid in a sequence-specific, antiparallel manner (i.e., nucleic acid-specific binding of complementary nucleic acids), also known as “annealing” or “hybridization”.
[0238] Hybridization requires two nucleic acids to contain complementary sequences, although mismatches between bases may exist. Suitable conditions for hybridization between two nucleic acids depend on their length and degree of complementarity, which are well-known variables in the art. Typically, hybridizable nucleic acids are 8 nucleotides or longer (e.g., 10 nucleotides or longer, 12 nucleotides or longer, 15 nucleotides or longer, 20 nucleotides or longer, 22 nucleotides or longer, 25 nucleotides or longer, or 30 nucleotides or longer).
[0239] It should be understood that the sequence of a polynucleotide does not need to be 100% complementary to the sequence of its target nucleic acid for specific hybridization. The polynucleotide may contain 60% or higher, 65% or higher, 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 98% or higher, 99% or higher, 99.5% or higher, or have 100% sequence complementarity with the target region of the target nucleic acid sequence it hybridizes with.
[0240] Hybridization of the target sequence with gRNA means that at least 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the nucleic acid sequences of the target sequence and gRNA can hybridize to form a complex; or it means that at least 12, 15, 16, 17, 18, 19, 20, 21, 22, or more bases of the nucleic acid sequences of the target sequence and gRNA can be complementary and hybridize to form a complex.
[0241] Express
[0242] As used herein, the term "expression" refers to the process by which a DNA template is transcribed into polynucleotides (such as mRNA or other RNA transcripts) and / or the transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as "gene products." If the polynucleotides are derived from genomic DNA, expression can include the splicing of mRNA in eukaryotic cells.
[0243] connector
[0244] As used herein, the term "linker" refers to a linear polypeptide formed by the linkage of multiple amino acid residues via peptide bonds. The linkers of this invention can be synthetically produced amino acid sequences or naturally occurring polypeptide sequences, such as polypeptides with hinge region functions. Such linker polypeptides are well known in the art (see, for example, Holliger, P. et al. (1993) Proc. Natl. Acad. Sci. USA 90:6444-6448; Poljak, RJ et al. (1994) Structure 2:1121-1123).
[0245] treat
[0246] As used in this article, the term "treatment" means to treat or cure a disease, to delay the onset of symptoms of a disease, and / or to slow the progression of a disease.
[0247] Subjects
[0248] As used herein, the term “subject” includes, but is not limited to, various animals, plants and microorganisms.
[0249] animal
[0250] For example, mammals, such as bovids, equines, sheep, suidae, canids, felines, lagos, rodents (e.g., mice or rats), non-human primates (e.g., macaques or cynomolgus monkeys), or humans. In some embodiments, the subject (e.g., a human) suffers from a condition (e.g., a condition caused by a disease-related gene defect).
[0251] plant
[0252] The term "plant" should be understood as any differentiated multicellular organism capable of photosynthesis, including crop plants at any stage of maturity or development, particularly monocotyledonous or dicotyledonous plants, vegetable crops including artichokes, kohlrabi, arugula, leeks, asparagus, lettuce (e.g., head lettuce, leaf lettuce, romaine lettuce), bok choy, taro, cucurbits (e.g., melons, watermelons, crenshaw melons, cantaloupes, Roman melons), and rapeseed crops (e.g., Brussels sprouts, cabbage, cauliflower, broccoli, kale, headless cabbage, Chinese cabbage). Bok choy, artichokes, carrots, napa cabbage, okra, onions, celery, parsley, chickpeas, parsnips, chicory, peppers, potatoes, gourds (e.g., zucchini, cucumber, baby zucchini, squash, pumpkin), radishes, dried onions, rutabaga, purple eggplant (also known as eggplant), ginseng, lettuce, scallions, endive, garlic, spinach, green onions, squash, leafy greens, beets (sugar beets and fodder beets), sweet potatoes, romaine lettuce, wasabi, tomatoes, turnips, and spices; fruits and / or vine crops such as apples, apricots, cherries, and nectarines. Fruits including peaches, pears, plums, prunes, cherries, quince, almonds, chestnuts, hazelnuts, pecans, pistachios, walnuts, citrus fruits, blueberries, boysenberries, cranberries, currants, raspberries, strawberries, blackberries, grapes, avocados, bananas, kiwifruit, persimmons, pomegranates, pineapples, tropical fruits, pears, melons, mangoes, papayas, and lychees; field crops such as clover, alfalfa, evening primrose, silvergrass, corn / maize (feed corn, sweet corn, popcorn), hops, jojoba, peanuts, rice, safflower, and small grain crops (barley, oats, rye, wheat, etc.). ), sorghum, tobacco, kapok, legumes (beans, lentils, peas, soybeans), oil plants (rapeseed, mustard, poppy, olive, sunflower, coconut, castor oil plants, cocoa beans, peanuts), Arabidopsis, fiber plants (cotton, flax, hemp, jute), Lauraceae (cinnamon, camphor), or a plant such as coffee, sugarcane, tea, and natural rubber plants; and / or bedding plants, such as flowering plants, cacti, succulents and / or ornamental plants, and trees such as forests (broadleaf trees and evergreen trees, such as conifers), fruit trees, ornamental trees, and nut-bearing trees, as well as shrubs and other seedlings.
[0253] Beneficial effects of the invention
[0254] This invention discovers a novel Cas enzyme. Blast results show that the Cas enzyme of this application has low similarity to previously reported Cas enzymes. It can exhibit nuclease activity in vivo and in vitro and has broad application prospects.
[0255] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings and examples. However, those skilled in the art will understand that the following drawings and examples are for illustrative purposes only and are not intended to limit the scope of the invention. Various objects and advantages of the present invention will become apparent to those skilled in the art from the following detailed description of the drawings and preferred embodiments. Attached Figure Description
[0256] Figure 1 The PAM structure of .Cas-sf0005.
[0257] Figure 2 The cleavage results of double-stranded nucleic acids by Cas-sf0005.
[0258] Figure 3 Sequencing results of target genes edited in eukaryotic cells using Cas-sf0005.
[0259] Figure 4 The cutting efficiency of .Cas-sf0005 on complementary and non-complementary chains.
[0260] Figure 5 The location where .Cas-sf0005 cleaves the target nucleic acid.
[0261] Figure 6 The PAM structure of Cas-sf9417.
[0262] Figure 7 The cleavage results of Cas-sf9417 on double-stranded nucleic acids.
[0263] Figure 8 Sequencing results of target genes edited in eukaryotic cells using Cas-sf9417.
[0264] Figure 9 The cutting efficiency of .Cas-sf9417 on complementary and non-complementary chains.
[0265] Figure 10 The location where .Cas-sf9417 cuts the target nucleic acid.
[0266] Figure 11 The PAM structure of Cas-sf8553.
[0267] Figure 12 The cleavage results of Cas-sf8553 on double-stranded nucleic acids.
[0268] Figure 13 The cleavage efficiency of .Cas-sf8553 on complementary and non-complementary chains.
[0269] Figure 14 The location where Cas-sf8553 cleaves the target nucleic acid.
[0270] Figure 15 Fluorescence results of Cas-sf1846 for nucleic acid detection.
[0271] Sequence information
[0272] SEQ ID No. describe 1 amino acid sequence of Cas-sf0005 2 Nucleic acid sequence encoding Cas-sf0005 3 Humanized Cas-sf0005 nucleic acid sequence 4 Amino acid sequence of Cas-sf9417 5 Nucleic acid sequence encoding Cas-sf9417 6 Humanized Cas-sf9417 nucleic acid sequence 7 amino acid sequence of Cas-sf8553 8 Nucleic acid sequence encoding Cas-sf8553 9 Humanized Cas-sf8553 nucleic acid sequence 10 amino acid sequence of Cas-sf1846 11 Nucleic acid sequence encoding Cas-sf1846 12 Humanized Cas-sf1846 nucleic acid sequence 13 DR region of Cas-sf0005 gRNA 14 DR region of Cas-sf9417 gRNA 15 DR region of Cas-sf8553 gRNA 16 DR region of Cas-sf1846 gRNA Detailed Implementation
[0273] The following examples are for illustrative purposes only and are not intended to limit the invention. Unless otherwise specified, the experiments and methods described in the examples are generally performed according to conventional methods well known in the art and described in various references. For example, conventional techniques such as immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA used in this invention can be found in Sambrook, Fritsch, and Maniatis, *Molecular Cloning: A Laboratory Manual*, 2nd edition (1989); *Current Protocols in Molecular Biology* (edited by F.M. Ausubel et al., (1987)); and the *Methods in Enzyme* series (academic publishing company): *PCR 2: A PRACTICAL*. APPROACH (edited by MJ MacPherson, BD Hames and GR Taylor (1995)), Harlow and Lane (1988) Antibodies, ALABORATORY MANUAL, and Animal Cell Culture (edited by R.R. Freshney (1987)).
[0274] Furthermore, unless specific conditions are specified in the examples, conventional conditions or conditions recommended by the manufacturer should be followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products. Those skilled in the art will understand that the examples are described by way of illustration and are not intended to limit the scope of protection claimed by the invention. All disclosures and other references mentioned herein are incorporated herein by reference in their entirety.
[0275] Example 1. Obtaining the Cas protein
[0276] The inventors analyzed the metagenomics of uncultured organisms and identified four novel Cas enzymes through redundancy removal and protein clustering analysis. Blast analysis showed that the Cas proteins had low sequence similarity to previously reported Cas proteins, and phylogenetic analysis indicated that they belonged to [a specific group / organism / element]. Similar proteins are named Cas-sf0005, Cas-sf9417, Cas-sf8553 and Cas-sf1846 in this invention.
[0277] The amino acid sequence, encoding nucleic acid sequence, and nucleic acid sequence optimized with human codons of the above protein are shown below:
[0278] SEQ ID No. describe 1 amino acid sequence of Cas-sf0005 2 Nucleic acid sequence encoding Cas-sf0005 3 Humanized Cas-sf0005 nucleic acid sequence 4 Amino acid sequence of Cas-sf9417 5 Nucleic acid sequence encoding Cas-sf9417 6 Humanized Cas-sf9417 nucleic acid sequence 7 amino acid sequence of Cas-sf8553 8 Nucleic acid sequence encoding Cas-sf8553 9 Humanized Cas-sf8553 nucleic acid sequence 10 amino acid sequence of Cas-sf1846 11 Nucleic acid sequence encoding Cas-sf1846 12 Humanized Cas-sf1846 nucleic acid sequence
[0279] Analysis of the homologous repeat sequences of the gRNAs corresponding to the above proteins showed the following results:
[0280] The direct repeat sequence of the gRNA corresponding to the Cas-sf0005 protein is CCGUCAACGUUCAACGCUUGCUCGGUUCGCCGAGAC (SEQ ID No. 13) or CCGUCAACGUUCAACGCUUGCUCGGUACGCCGAGAC.
[0281] The homologous repeat sequence of the gRNA corresponding to the Cas-sf9417 protein is GUGGGAACCCUUCCUGAUGGCUCGAUCCGUCGAGAC (SEQ ID No. 14).
[0282] The homologous repeat sequence of the gRNA corresponding to the Cas-sf8553 protein is GUGAGGACGACAACCGAAGGCUCGGUGCGCCGAGAC (SEQ ID No. 15);
[0283] The direct repeat sequence of the gRNA corresponding to the Cas-sf1846 protein is GUUGCAAGCCUUACUGAUUGCUCCUCACGAGGAGAC (SEQ ID No. 16).
[0284] Example 2. PAM identification of Cas-sf0005 protein
[0285] Construct a PAM library and synthesize the following sequence: CGTGTTTCGTAAAGTCTGGAAACGCGGAAGCCCCCAGCGCTTCAGCGTTCNNNNNN TCCCCTACGTGCTGCTGAAGTTGCCCGCAA, where N represents a random deoxyribonucleotide, and the underlined sequence is the target sequence. After Klenow enzyme completion, it is ligated into the pcyc184 vector. After transformation into *E. coli*, plasmids are extracted to form a PAM library. The 20bp sequence of gRNA Cas-sf0005-5'spacer1 is: CCGUCAACGUUCAACGCUUGCUCGGUUCGCCGAGAC UCCCCUACGUGC UGCUGAAG (The underlined area is the target area);
[0286] Primer sequence: TK-117: CGGCATTCCTGCTGAACCGCTCTTCCGATCT;
[0287] TK-111:GATCGGAAGAGCGGTTCAGCAGGAATGCCG;
[0288] PAM-after-F:ACTCAGGGGTCTTCGGTTTCCGTGTT;
[0289] S6-PAM-after:ACTCAGCTGAACCGTCCTTCCG;
[0290] Obtain the Cas-sf0005 protein-preferred PAM library: 50 nM Cas-sf0005 protein and 50 nM gRNA were incubated at 25°C for 10 min in a buffer solution of Tris-HAC 40 mM, Mg(AC)2 30 mM, BSA 120 μg / ml, DTT 12 mM, and pH 7.0. The PAM library plasmid (10 ng / μL) was added, and the mixture was incubated at 37°C for 1 h and then at 85°C for 20 min. 2.5 U DreamTaq DNA Polymerase (5 U / μL) (Thermo Fisher Scientific) and 4 μL 2.5 mM dNTP Mix (TransGold) were added to the above system, and the mixture was incubated at 72°C for 30 min to perform end-closing and 3' A addition. The product was purified using an Omega Gel Extraction Kit D2500. After annealing, primers TK-117 and TK-111 were ligated with the above product using T4 ligase. The resulting ligation product was subjected to PCR with primers PAM-after-F and S6-PAM-after to obtain the Cas-sf0005 protein-preferred PAM library. The PCR product was then subjected to next-generation sequencing to obtain the PAM sequence. A Weblogo plot was used to create the results, as shown below. Figure 1 As shown, the PAM sequence identified by Cas-sf0005 is 5'TBN-3', where B = T / C / G and N = A / T / C / G.
[0291] Example 3. Application of Cas-sf0005 protein in double-stranded nucleic acid editing
[0292] This embodiment tests the cleavage activity of the Cas-sf0005 protein on double-stranded DNA in vitro. In this embodiment, gRNA that can pair with the target nucleic acid guides the Cas-protein to recognize and bind to the double-stranded target nucleic acid; subsequently, the Cas protein is activated to cleave the double-stranded target nucleic acid, thereby cleaving the double-stranded target nucleic acid in the system. The cleaved double-stranded target nucleic acid is then detected by agarose gel electrophoresis.
[0293] In this embodiment, the target nucleic acid is selected as double-stranded DNA, and its sequence is as follows: The vector T-Vector-pEASY-Blunt Simple Cloning Vector is used for ligation; the italicized portion TTA represents the PAM sequence, and the underlined region is the target region. gRNA-Cas-sf0005-5'spacer1-20bp: CCGUCAACGUUCAACGCUUGCUCGGUUCGCCGAGAC UCCC CUACGUGCUGCUGAAG (The underlined area is the target region); Simultaneously, the above fragment was amplified by PCR and used as the target nucleic acid for detection; the following reaction system was used: 20 μL system, Tris-HAC 40 mM, Mg(AC)2 30 mM, BSA 120 μg / ml, DTT 12 mM, pH 7.0, Cas-sf0005 50 nM, gRNA 500 nM, with a final concentration of double-stranded target nucleic acid of 5 ng / μL. Cas protein and gRNA were incubated at 25℃ for 10 min; double-stranded target nucleic acid was added, and incubation was carried out at 37℃ for 1 h. Detection was performed by 1.0% agarose gel electrophoresis. The experimental group had target nucleic acid and gRNA added, while the control group did not have gRNA added.
[0294] like Figure 2 As shown in Figure A (lanes 1, 2, and 3 from left to right), lane 1 is the experimental group, lane 2 is the control group, and lane 3 is the Trans2K plus DNA marker. Compared with the control group without added gRNA, the Cas-sf0005 in the experimental group can effectively cleave double-stranded nucleic acids in the system when the target nucleic acid (plasmid) and gRNA are present. Figure 2 As shown in Figure B, lane 1 is the Trans2K plus DNA marker, lanes 2-5 are the experimental groups with different PAMs, and lane 3 is the control group. Compared with the control group without gRNA, the Cas-sf0005 in the experimental group can effectively cleave double-stranded nucleic acids in the system containing target nucleic acids (PCR products) of TTN PAM and gRNA. The experimental results indicate that Cas-sf0005 can be used for the cleavage and editing of double-stranded target nucleic acids.
[0295] Example 4. Editing efficiency of Cas-sf0005 protein in animal cells
[0296] The gene-editing activity of the Cas-sf0005 protein was verified in animal cells, and a target was designed for the FUT8 gene in Chinese hamster ovary cells (CHO). The vector pcDNA3.3 was modified to carry EGFP fluorescent protein and the PuroR resistance gene. The SV40 NLS-Cas-sf0005 fusion protein was inserted via the BsmB1 restriction site; the U6 promoter and gRNA sequence were inserted via the Mfe1 restriction site. The CMV promoter initiated the expression of the SV40 NLS-Cas-sf0005-NLS-GFP fusion protein. The Cas-sf0005-NLS protein and the GFP protein were linked by the linker peptide T2A. The EF-1α promoter initiated the expression of the puromycin resistance gene.
[0297] Plating: CHO cells are plated when the confluence reaches 70-80%, with a cell number of 8*10^4 cells / well in a 12-well plate.
[0298] Transfection: Transfect after 6-8 hours of plating. Add 3.25 μl of Lipo3000 to 125 μl of opti-MEM and mix well; add 3 μg of plasmid and 10 μl of P3000 to 125 μl of opti-MEM and mix well. Mix the diluted Lipo3000 with the diluted plasmid thoroughly and incubate at room temperature for 5 minutes. Add the incubated mixture to the culture medium containing cells for transfection.
[0299] Selection with puromycin: Add puromycin 24 h after transfection, to a final concentration of 10 ng / ml. After 24 h of puromycin treatment, replace with normal culture medium and continue culturing for another 24 h.
[0300] DNA extraction, PCR amplification of the area near the editing region, and hiTOM sequencing: Cells were collected after trypsin digestion, and genomic DNA was extracted using a cell / tissue genomic DNA extraction kit (Biotech). The genomic DNA was amplified near the target site using primers NAFUT8-JC-HITOM-1F1: GGAGTGAGTACGGTGTGCTCCTCCTTACTTACCCTTGG; and NAFUT8-JC-HITOM-1R1: GAGTTGGATGCTGGATGGTTGTTCTTTGGTGGGACTATG. The PCR products were then sequenced using hiTOM (http: / / 121.40.237.174 / HiTOM / Sampleacceptance_sanyang.php).
[0301] Sequencing data analysis was conducted to statistically analyze the types and proportions of sequences within 15 nt upstream and 10 nt downstream of the target site, thereby obtaining the editing efficiency of the Cas-sf0005 protein at the target site.
[0302] CHO cell FUT8 gene target sequence: gR1-FUT8: The italicized portion is the PAM sequence, and the underlined area is the target region. The gRNA sequence is GCCGUCAACGUUCAACGCUUGCUCGGUUCGCCGAGAC GACAAACUGGGAUACCCACCACA The underlined area is the target area.
[0303] Analysis showed that Cas-sf0005 achieved an editing efficiency of 22.49% in the target site gR1-FUT8 of CHO cells, with all editing types being InDel. Partial sequencing results of the edited target nucleic acids are shown below. Figure 3 As shown.
[0304] Example 5. Cleavage characteristics of Cas-sf0005 protein - NTS / TS cleavage efficiency
[0305] This embodiment assesses the cleavage efficiency of Cas-sf0005 on the complementary strand (TS) and non-complementary strand (NTS) of target nucleotides (double-stranded DNA) in vitro. 5'6-FAM labels the non-complementary strand (NTS), and 5'ROX labels the complementary strand (TS). gRNA guides the Cas-sf0005 protein to recognize and bind to the complementary strand (TS) of the target nucleic acid, thereby cleaving the complementary strand (TS) and non-complementary strand (NTS) of the target nucleic acid in the system. The cleaved target nucleic acids are then detected by capillary electrophoresis (ABI 3730xl genetic analyzer). DNA fragments migrate from the cathode to the anode in the gel, arranged according to fragment length. When they reach the scanning window of the laser scanner at the anode end, the fluorescent dye is excited, emitting light of a specific wavelength, which is recorded according to the fluorescence intensity. The electrophoretic trajectory of each DNA fragment carrying the fluorescent dye is recorded according to the actual time it takes to pass through the laser scanning window, represented by a fluorescence absorption peak. A higher peak indicates a greater quantity of the fragment; the time of peak appearance is directly related to fragment size—the smaller the fragment, the earlier the peak appears. In the FAM fluorescence channel, the uncleaved NTS fragment size of Cas-sf0005 is 380 nt, and the fragment size after cleavage of the NTS by Cas-sf0005 is approximately 126 nt. In the ROX fluorescence channel, the uncleaved TS fragment size of Cas-sf0005 is 380 nt, and the fragment size after cleavage of the TS by Cas-sf0005 is approximately 254 nt. The fragment cleavage efficiency is calculated as: Cleavage efficiency = Cleavage peak area / (Cleavage peak area + Uncleaved peak area).
[0306] In this embodiment, the target nucleic acid was selected as double-stranded DNA (PCR product), and the primers were:
[0307] XQ0001-5FAM:GTATGTTGTGTGGAATTGTG 5'6-FAM;
[0308] XQ0002-5ROX:GCTGCGCGTAACCACCACAC 5'ROX
[0309] The amplified product sequence is as follows:
[0310]
[0311] The italicized portion represents the PAM sequence, and the underlined area represents the target region.
[0312] Cas-sf0005-5spacer1:
[0313] GCCGUCAACGUUCAACGCUUGCUCGGUUCGCCGAGAC UCCCCUACGUGCUGCUGAAG (The underlined area is the target area)
[0314] The following reaction system was used: 20 μL system containing Tris-HAC 40 mM, Mg(AC)2 30 mM, BSA 120 μg / ml, DTT 12 mM, pH 7.0, Cas-sf0005 50 nM, gRNA 100 nM, and 1 μL of double-stranded target nucleic acid (PCR product). Incubation was performed at 37℃ for 5 min, 15 min, 30 min, and 60 min, followed by incubation at 85℃ for 20 min. Products in the FAM and ROX channels were detected by capillary electrophoresis (ABI 3730xl genetic analyzer). Data analysis was performed using Gene Mapper 4.1 software to calculate the NTS / TS cleavage efficiency.
[0315] The results are as follows Figure 4 As shown, Cas-sf0005 cleaves both NTS and TS simultaneously when cleaving double-stranded target nucleic acids, with essentially the same cleavage efficiency.
[0316] Example 6. Cleavage characteristics of Cas-sf0005 protein - cleavage site
[0317] This embodiment uses in vitro detection to determine the cleavage sites of the Cas-sf0005 protein on the complementary and non-complementary strands of double-stranded target nucleic acids. In this embodiment, gRNA guides the Cas-sf0005 protein to recognize and bind to the double-stranded target nucleic acid; the Cas-sf0005 protein is activated to cleave the double-stranded target nucleic acid, thereby cleaving the double-stranded target nucleic acid in the system. The cleaved double-stranded target nucleic acid is then padded with A and ligated with a T-linked adapter. The ligation product is enriched by PCR and then Sanger sequenced.
[0318] In this embodiment, the target nucleic acid was selected as double-stranded DNA (plasmid), and the synthesized sequence was: The vector T-Vector-pEASY-Blunt Simple Cloning Vector is inserted. The italicized part is the PAM sequence, and the underlined area is the target region.
[0319] gRNA: Cas-sf0005-5spacer1:
[0320] GCCGUCAACGUUCAACGCUUGCUCGGUUCGCCGAGAC UCCCCUACGUGCUGCUGAAG (The underlined area is the target area).
[0321] The following reaction system was used: 50 μL system, buffer: Tris-HAC 40 mM, Mg(AC)2 30 mM, BSA 120 μg / ml, DTT 12 mM, pH 7.0, final Cas concentration 100 nM, final gRNA concentration 250 nM, double-stranded target nucleic acid 10 ng / μL (plasmid). Cas protein and gRNA were incubated at 25°C for 10 min; double-stranded target nucleic acid was added, and the mixture was incubated at 37°C for 1 h, followed by incubation at 85°C for 5 min; 50 μL of 2X Taq DNA Polymerase Mix (Novizan) (1:1) was added to the above system, and the reaction was carried out at 72°C for 30 min; the above reaction solution was recovered; 2 μL of annealed primers (TK-117: CGGCATTCCTGCTGAACCGCTCTTCCGATCT, TK-111: GATCGGAAGAGCGGTTCAGCAGGAATGCCG) were added to the recovered liquid, and the mixture was incubated with T4 (NEB) ligase at 22°C for 1 h. 10 μL of the ligation product was taken, and PCR was performed with primers S1-PAM-after: ACTCAGCGGCATTCCTGCTGAACCGC, PQ0275-F: CCGTATTACCGCCTTTGAG, and 2X Taq DNA Polymerase Mix (Novizan). The PCR product was then subjected to Sanger sequencing. The results are as follows. Figure 5 As shown, when the Cas-sf0005 protein cleaves the target nucleic acid, the cleavage site is between 15-16 nt of the non-complementary strand NTS and between 21-22 nt of the complementary strand TS. That is, the cleavage site of Cas-sf0005 on the complementary strand of the target sequence is between 21 and 22 nt from the 5' end of the PAM complementary sequence, and the cleavage site of Cas-sf0005 on the non-complementary strand of the target sequence is between 15 and 16 nt from the 3' end of the PAM sequence. The gRNA guides the Cas-sf0005 protein to recognize and bind to the aforementioned complementary strand, which is the DNA strand paired with the complementary strand.
[0322] Example 7. PAM identification of Cas-sf9417 protein
[0323] Construction of the Cas-sf9417 protein expression plasmid: After optimizing the nucleic acid sequence with human codons, the gene was synthesized and ligated into the *E. coli* expression vector PeT28(a)+. The JM23119 promoter was added to the PeT28(a)+-Cas-sf9417 vector to initiate Cas-sf9417 CrRNA transcription. The resulting vector was: PeT28(a)+-Cas-sf9417-JM23119-crRNA, with the crRNA sequence: GUGGGAACCCUUCCUGAUGGCUCGAUCCGUCGAGAC UCCCCUACGUGCUGCUGAAGUUGC Underlined sequences represent target sequences; PAM library construction: Synthetic sequence CGTGTTTCGTAAAGTCTGGAAACGCGGAAGCCCCCAGCGCTTCAGCGTTCNNNNNN TCCCCTACGTGCTGCTGAAGTTGC CCGCAA, where N represents a random deoxyribonucleotide, and the underlined sequence is the target sequence. After being filled with Klenow enzyme, it was ligated into the pcyc184 vector. Following transformation into E. coli, plasmids were extracted to form a PAM library.
[0324] PAM library subtraction experiment: Competent cells were prepared: BL21(DE3)-PeT28(a)+-Cas-sf9417-JM23119-crRNA. The PAM library plasmid was transformed into competent cells: BL21(DE3)-PeT28(a)+-Cas-sf9417-JM23119-crRNA was plated on LB agar plates containing kanamycin and chloramphenicol, incubated overnight at 37°C, and the cells were collected. The bacterial concentration was adjusted to OD600 ≤ 0.6-0.8, and 0.2 mM IPTG was added, followed by induction at 37°C for 4 h. Plasmid extraction was performed using FastPure EndoFree Plasmid MaxiKit (vazyme) to obtain the subtracted PAM library. Primers: PAM-F: GGTCTTCGGTTTCCGTGTT; PAM-R: TGGCGTTGACTCTCAGTCAT. PCR was performed using a 30 ng / μL plasmid (PAM library) as template primers to obtain control group samples, and PCR was performed using a 30 ng / μL plasmid (subtracted PAM library) as template to obtain experimental group samples. Both control and experimental group samples were sent for next-generation sequencing data analysis. Weblogo was used for plotting. The PAM structure recognized by Cas-sf9417 was 5'-TTN-3', where N = A / T / C / G. Figure 6 As shown.
[0325] Example 8. Application of Cas-sf9417 protein in double-stranded nucleic acid editing
[0326] This embodiment describes the in vitro detection of the cis cleavage activity of Cas-sf9417 on double-stranded DNA. In this embodiment, gRNA that can pair with the target nucleic acid guides the Cas-sf9417 protein to recognize and bind to the target nucleic acid, thereby cleaving the target nucleic acid in the system. The cleaved target nucleic acid is then detected by agarose gel electrophoresis.
[0327] In this embodiment, the target nucleic acid was selected as double-stranded DNA (plasmid), 5spacer1-PAM, whose sequence is as follows: The vector T-Vector-pEASY-Blunt Simple Cloning Vector is inserted; the italicized part is the PAM sequence, N = A / T / C / G, and the underlined area is the target region.
[0328] gRNA:Cas-sf9417-5spacer1:
[0329] GUGGGAACCCUUCCUGAUGGCUCGAUCCGUCGAGAC UCCCCUACGUGCUGCUGAAG (The underlined area is the target area)
[0330] The following reaction system was used: 20 μL system, buffer: Tris-HAC 40 mM, Mg(AC)2 30 mM, BSA 120 μg / ml, DTT 12 mM, pH 7.0, final concentration of Cas-sf9417 100 nM, final concentration of gRNA 200 nM, final concentration of double-stranded target nucleic acid 5 ng / μL. Incubation was performed at 37℃ for 1 h and then at 85℃ for 20 min. The cleavage products were analyzed by agarose gel electrophoresis to detect the Cas-sf9417 cleavage ability. The experimental group received Cas-sf9417 protein, gRNA, and target nucleic acid, while the control group (CK) received no gRNA.
[0331] The results are as follows Figure 7 As shown, compared with the control group without gRNA, Cas-sf9417 in the PAM-TTN experimental group was able to cleave double-stranded nucleic acids in the system, exhibiting obvious cleavage bands. This indicates that Cas-sf9417 can be used for the cleavage and editing of double-stranded target nucleic acids with PAM-TTN.
[0332] Example 9. Editing efficiency of Cas-sf9417 protein in animal cells
[0333] The gene-editing activity of the Cas-sf9417 protein was validated in animal cells, with a target designed for the FUT8 gene in Chinese hamster ovary cells (CHO). The vector pcDNA3.3 was modified to carry EGFP fluorescent protein and the PuroR resistance gene. The SV40 NLS-Cas-sf9417-NLS fusion protein was inserted via the BsmB1 restriction site; the U6 promoter and gRNA sequence were inserted via the Mfe1 restriction site. The CMV promoter initiated the expression of the SV40 NLS-Cas-sf9417-NLS-GFP fusion protein. The Cas-sf9417-NLS and GFP proteins were linked using the linker peptide T2A. The EF-1α promoter initiated the expression of the puromycin resistance gene.
[0334] Plating: CHO cells are plated when the confluence reaches 70-80%, with a cell number of 8*10^4 cells / well in a 12-well plate.
[0335] Transfection: Transfect after 12-24 hours of plating. Add 2 μg of plasmid to 100 μL of opti-MEM and mix well; add 4 μL of diluted plasmid to... EL Transfection Reagent (TRAN) was incubated at room temperature for 15-20 minutes. The incubated mixture was then added to a culture medium containing cells for transfection.
[0336] Selection with puromycin: Add puromycin 24 h after transfection, to a final concentration of 10 ng / ml. After 24 h of puromycin treatment, replace with normal culture medium and continue culturing for another 24 h.
[0337] DNA extraction, PCR amplification of the region near the editing site, and hiTOM sequencing: Cells were collected after trypsin digestion, and genomic DNA was extracted using a cell / tissue genomic DNA extraction kit (Biotech). The genomic DNA was amplified near the target site using primers NAFUT8-JC-Hitom-1F2: GGAGTGAGTACGGTGTGCAGTCCATGTCAGACGCACTG; PQ0106-FUT8-HiTom-R5: GAGTTGGATGCTGGATGGTACAGAACCACTTGTTGGTC. The PCR products were then sequenced using hiTOM (http: / / 12140237174 / HiTOM / Sample_acceptance_sanyang.php).
[0338] Sequencing data analysis was conducted to statistically analyze the types and proportions of sequences within a 20bp range upstream and downstream of the target site -18, thereby obtaining the editing efficiency of the Cas-sf9417 protein at the target site.
[0339] CHO cell FUT8 gene target sequence: gR16-FUT8: The italicized portion is the PAM sequence, and the underlined area is the target region. The gRNA sequence is GUGGGAACCCUUCCUGAUGGCUCGAUCCGUCGAGAC AACAAAGAAGGGUCAUCAGUGG The underlined area is the target area.
[0340] Analysis showed that Cas-sf9417 achieved an editing efficiency of 11.39% in the target site gR16-Cas12i3-target-FUT8 of CHO cells, with the editing type being InDel. Partial sequencing results of the edited target nucleic acid are shown below. Figure 8 As shown.
[0341] Example 10. Cleavage characteristics of Cas-sf9417 protein - NTS / TS cleavage efficiency
[0342] This embodiment describes the in vitro assay of the cleavage efficiency of Cas-sf9417 on the complementary (TS) and non-complementary (NTS) strands of target nucleotide double-stranded DNA. The non-complementary (NTS) strand was labeled with 5'6-FAM, and the complementary (TS) strand with 5'ROX. gRNA guided the Cas-sf9417 protein to recognize and bind to the complementary strand of the target nucleic acid, thereby cleaving the target nucleic acid in the system. The cleaved target nucleic acid was then detected by capillary electrophoresis (ABI 3730xl genetic analyzer). DNA fragments migrate from the cathode to the anode in the gel, arranged according to fragment length. When they reach the scanning window of the laser scanner at the anode end, the fluorescent dye is excited, emitting light of a specific wavelength, which is recorded according to the fluorescence intensity. The electrophoretic trajectory of each DNA fragment carrying the fluorescent dye is recorded according to the actual time it takes to pass through the laser scanning window, and each fragment is represented by a fluorescence absorption peak. The higher the peak value, the greater the amount of that fragment; the time of peak appearance is directly related to fragment size, with smaller fragments appearing earlier. In the FAM fluorescence channel, the uncleaved NTS fragment size of Cas-sf9417 is 380 nt, and the fragment size after cleavage of the NTS by Cas-sf9417 is approximately 126 nt. In the ROX fluorescence channel, the uncleaved TS fragment size of Cas-sf9417 is 380 nt, and the fragment size after cleavage of the TS by Cas-sf9417 is approximately 254 nt. The fragment cleavage efficiency is calculated as: Cleavage efficiency = Cleavage peak area / (Cleavage peak area + Uncleaved peak area).
[0343] In this embodiment, the target nucleic acid was selected as double-stranded DNA (PCR product), and the primers were:
[0344] XQ0001-5FAM:GTATGTTGTGTGGAATTGTG 5'6-FAM;
[0345] XQ0002-5ROX:GCTGCGCGTAACCACCACAC 5'ROX
[0346] The amplified product sequence is as follows:
[0347]
[0348] The italicized portion represents the PAM sequence, and the underlined area represents the target region.
[0349] gRNA: Cas-sf9417-5spacer1:
[0350] GUGGGAACCCUUCCUGAUGGCUCGAUCCGUCGAGAC UCCCCUACGUGCUGCUGAAG (The underlined area is the target area)
[0351] The following reaction system was used: 20 μL system, buffer: Tris-HAC 40 mM, Mg(AC)2 30 mM, BSA 120 μg / ml, DTT 12 mM, pH 7.0, Cas-sf9417 50 nM, gRNA 100 nM, and 1 μL of double-stranded target nucleic acid (PCR product). Incubation was performed at 37℃ for 5 min, 15 min, 30 min, and 60 min, and at 85℃ for 20 min. Products in the FAM and ROX channels were detected by capillary electrophoresis (ABI 3730xl genetic analyzer). Data analysis was performed using Gene Mapper 4.1 software to calculate the cleavage efficiency of NTS and TS.
[0352] The results are as follows Figure 9 As shown: Cas-sf9417 preferentially cleaves NTS when cleaving target nucleic acids.
[0353] Example 11. Cleavage characteristics of Cas-sf9417 protein - cleavage site
[0354] This embodiment uses in vitro detection to determine the cleavage sites of the Cas-sf9417 protein on the complementary and non-complementary strands of double-stranded target nucleic acids. In this embodiment, gRNA guides the Cas-sf9417 protein to recognize and bind to the double-stranded target nucleic acid; the Cas protein then activates its Cis cleavage activity on the double-stranded target nucleic acid, thereby cleaving the double-stranded target nucleic acid in the system. The cleaved double-stranded target nucleic acid is then padded with A and ligated with a T-linked adapter. The ligation product is enriched by PCR and then Sanger sequenced.
[0355] In this embodiment, the target nucleic acid is selected as double-stranded DNA (plasmid), sequence: The vector T-Vector-pEASY-Blunt Simple Cloning Vector is inserted. The italicized part is the PAM sequence, and the underlined area is the target region.
[0356] gRNA: Cas-sf9417-5spacer1:
[0357] GUGGGAACCCUUCCUGAUGGCUCGAUCCGUCGAGAC UCCCCUACGUGCUGCUGAAG (The underlined area is the target area).
[0358] The following reaction system was used: 50 μL system, buffer: Tris-HAC 40 mM, Mg(AC)2 30 mM, BSA 120 μg / ml, DTT 12 mM, pH 7.0, Cas-sf9417 100 nM, gRNA 250 nM, double-stranded target nucleic acid 10 ng / μL (plasmid). Cas protein and gRNA were incubated at 25°C for 10 min; double-stranded target nucleic acid was added, and the mixture was incubated at 37°C for 1 h, followed by incubation at 85°C for 5 min; 50 μL of 2X Taq DNA Polymerase Mix (Novizan) (1:1) was added to the above system, and the reaction was carried out at 72°C for 30 min; the above reaction solution was recovered; 2 μL of annealed primers (2 μM, TK-117: CGGCATTCCTGCTGAACCGCTCTTCCGATCT, TK-111: GATCGGAAGAGCGGTTCAGCAGGAATGCCG) were added to the recovered liquid, and the mixture was incubated with T4 (NEB) ligase at 22°C for 1 h. 10 μL of the ligation product was taken, and PCR was performed with primers S1-PAM-after: ACTCAGCGGCATTCCTGCTGAACCGC, PQ0275-F: CCGTATTACCGCCTTTGAG, and 2X Taq DNA Polymerase Mix (Novizan). The PCR product was then subjected to Sanger sequencing. The results are as follows. Figure 10 As shown, when the Cas-sf9417 protein cleaves the target nucleic acid, the cleavage site is between 14-16 nt of the NTS and between 21-22 nt of the TS. That is, the cleavage site of Cas-sf9417 on the complementary strand of the target sequence is between 21 nt and 22 nt at the 5' end of the PAM complementary sequence, and the cleavage site of Cas-sf9417 on the non-complementary strand of the target sequence is between 14 nt and 16 nt at the 3' end of the PAM sequence. The gRNA guides the Cas-sf9417 protein to recognize and bind to the aforementioned complementary strand, which is the DNA strand that pairs with the complementary strand.
[0359] Example 12. PAM identification of Cas-sf8553 protein
[0360] Construction of the Cas-sf8553 protein expression plasmid: After optimizing the nucleic acid sequence with human codons, the gene was synthesized and ligated into the *E. coli* expression vector PeT28(a)+. The JM23119 promoter was added to the PeT28(a)+-Cas-sf8553 vector to initiate Cas-sf8553 CrRNA transcription. The resulting vector was: PeT28(a)+-Cas-sf8553-JM23119-crRNA, with the crRNA sequence: GUGAGGACGACAACCGAAGGCUCGGUGCGCCGAGAC UCCCCUACGUGCUGCUGAAGUUGC Underlined sequences represent target sequences; PAM library construction: Synthetic sequence CGTGTTTCGTAAAGTCTGGAAACGCGGAAGCCCCCAGCGCTTCAGCGTTCNNNNNN TCCCCTACGTGCTGCTGAAGTTGC CCGCAA, where N represents a random deoxyribonucleotide, and the underlined sequence is the target sequence. After being filled with Klenow enzyme, it was ligated into the pcyc184 vector. Following transformation into E. coli, plasmids were extracted to form a PAM library.
[0361] PAM library subtraction experiment: Competent cells were prepared: BL21(DE3)-PeT28(a)+-Cas-sf8553-JM23119-crRNA. The PAM library plasmid was transformed into competent cells: BL21(DE3)-PeT28(a)+-Cas-sf8553-JM23119-crRNA, plated on LB agar plates containing kanamycin and chloramphenicol, and incubated overnight at 37°C. The cells were then collected, and the bacterial concentration was adjusted to OD600 ≤ 0.6-0.8. 0.2 mM IPTG was added, and the cells were induced at 37°C for 4 h. Plasmid extraction was performed using a large-scale extraction kit to obtain the subtracted PAM library. Primers: PAM-F: GGTCTTCGGTTTCCGTGTT; PAM-R: TGGCGTTGACTCTCAGTCAT. Using a 30 ng / μL plasmid (PAM library) as a template, PCR was performed to obtain control group samples, and using a 30 ng / μL plasmid (subtracted PAM library) as a template, PCR was performed to obtain experimental group samples. Both control and experimental group samples were sent for next-generation sequencing data analysis. Weblogo was used for plotting. The PAM structure recognized by Cas-sf8553 was 5'-TTN-3', N=A / T / C / G, as shown below. Figure 11 As shown.
[0362] Example 13. Application of Cas-sf8553 protein in double-stranded nucleic acid editing
[0363] This embodiment detects the cis cleavage activity of Cas-sf8553 on double-stranded DNA. In this embodiment, gRNA that can pair with the target nucleic acid guides the Cas-sf8553 protein to recognize and bind to the target nucleic acid, thereby cleaving the target nucleic acid in the system. The cleaved target nucleic acid is then detected by agarose gel electrophoresis.
[0364] In this embodiment, the target nucleic acid was selected as double-stranded DNA (plasmid), 5spacer1-PAM, whose sequence is as follows: Linked to carrier T-
[0365] Vector-pEASY-Blunt Simple Cloning Vector; the italicized part is the PAM sequence, N = A / T / C / G, and the underlined area is the target region.
[0366] Cas-sf8553-5spacer1:
[0367] GUGAGGACGACAACCGAAGGCUCGGUGCGCCGAGAC UCCCCUACGUGCUGCUGAAG (The underlined area is the target area)
[0368] The following reaction system was used: 20 μL system, buffer: Tris-HAC 40 mM, Mg(AC)2 30 mM, BSA 120 μg / ml, DTT 12 mM, pH 7.0, Cas-sf8553 100 nM, gRNA 200 nM, and the final concentration of double-stranded target nucleic acid was 5 ng / μL. Incubation was performed at 37℃ for 1 h and then at 85℃ for 20 min. The cleavage products were subjected to electrophoresis with 1% agarose gel to detect the Cas-sf8553 cleavage ability. The experimental group was supplemented with Cas-sf8553 protein, gRNA, and target nucleic acid, while the control group (CK) was not supplemented with gRNA.
[0369] The results are as follows Figure 12 As shown, compared with the control group without gRNA, Cas-sf8553 was able to cleave double-stranded nucleic acids in the PAM-TTN experimental group, exhibiting obvious cleavage bands. This indicates that Cas-sf8553 can be used for the cleavage and editing of double-stranded target nucleic acids with PAM-TTN.
[0370] Example 14. Editing efficiency of Cas-sf8553 protein in animal cells
[0371] The gene-editing activity of the Cas-sf8553 protein was verified in animal cells, and a target was designed for the FUT8 gene in Chinese hamster ovary cells (CHO). The vector pcDNA3.3 was modified to carry EGFP fluorescent protein and the PuroR resistance gene. The SV40 NLS-Cas-sf8553-NLS fusion protein was inserted via the BsmB1 restriction site; the U6 promoter and gRNA sequence were inserted via the Mfe1 restriction site. The CMV promoter initiated the expression of the SV40 NLS-Cas-sf8553-NLS-GFP fusion protein. The Cas-sf8553-NLS protein and the GFP protein were linked by the linker peptide T2A. The EF-1α promoter initiated the expression of the puromycin resistance gene.
[0372] Plating: CHO cells are plated when the confluence reaches 70-80%, with a cell number of 8*10^4 cells / well in a 12-well plate.
[0373] Transfection: Transfect after 12-24 hours of plating. Add 2 μg of plasmid to 100 μL of opti-MEM and mix well; add 4 μL of diluted plasmid to... EL Transfection Reagent (TRAN) was incubated at room temperature for 15-20 minutes. The incubated mixture was then added to a culture medium containing cells for transfection.
[0374] Selection with puromycin: Add puromycin 24 h after transfection, to a final concentration of 10 ng / ml. After 24 h of puromycin treatment, replace with normal culture medium and continue culturing for another 24 h.
[0375] DNA extraction, PCR amplification of the area near the editing region, and hiTOM sequencing: Cells were collected after trypsin digestion, and genomic DNA was extracted using a cell / tissue genomic DNA extraction kit (Biotech). The genomic DNA was amplified near the target region using primers PQ0106-FUT8-HiTom-F1: GGAGTGAGTACGGTGTGCGAGTTCTGTTGCATGGTAGG; and PQ0106-FUT8-HiTom-R1: GAGTTGGATGCTGGATGGGCCAAGCTTCTTGGTGGTTTC. The PCR products were then sequenced using hiTOM (http: / / 121.40.237.174 / HiTOM / Sample_acceptance_sanyang.php).
[0376] Sequencing data analysis was conducted to statistically analyze the sequence types and proportions within a 20bp range upstream and downstream of the target site -18, thereby obtaining the editing efficiency of the Cas-sf0005 protein at the target site.
[0377] CHO cell FUT8 gene target sequence: gR3-FUT8: The italicized portion represents the PAM sequence, and the underlined area is the target region. The gRNA sequence is GUGAGGACGACAACCGAAGGCUCGGUGCGCCGAGAC CAGCCAAGGUUGUGGACGGAUCA The underlined area is the target area.
[0378] The analysis results showed that the editing efficiency of Cas-sf8553 in the target gR3-FUT8 of CHO cells was 1.05%, and the editing type was InDel.
[0379] Example 15. Cleavage characteristics of Cas-sf8553 protein - NTS / TS cleavage efficiency
[0380] This embodiment examines the cleavage efficiency of Cas-sf8553 on the complementary strand (TS) and non-complementary strand (NTS) of target nucleotides (double-stranded DNA). The non-complementary strand (NTS) is labeled with 5'6-FAM, and the complementary strand (TS) is labeled with 5'ROX. gRNA guides the Cas-sf8553 protein to recognize and bind to the target nucleic acid, thereby cleaving the target nucleic acid in the system. The cleaved target nucleic acid is then detected by capillary electrophoresis (ABI 3730xl genetic analyzer). DNA fragments migrate from the cathode to the anode in the gel, arranged according to fragment length. When they reach the scanning window of the laser scanner at the anode end, the fluorescent dye is excited, emitting light of a specific wavelength, which is recorded according to the fluorescence intensity. The electrophoretic trajectory of each DNA fragment carrying the fluorescent dye is recorded according to the actual time it takes for it to pass through the laser scanning window, and each fragment is represented by a fluorescence absorption peak. The higher the peak value, the greater the quantity of the fragment; the time of peak appearance is directly related to fragment size, with smaller fragments appearing earlier. In the FAM fluorescence channel, the uncleaved NTS fragment size of Cas-sf8553 is 380 nt, and the fragment size after cleavage of the NTS by Cas-sf8553 is approximately 126 nt. In the ROX fluorescence channel, the uncleaved TS fragment size of Cas-sf8553 is 380 nt, and the fragment size after cleavage of the TS by Cas-sf8553 is approximately 254 nt. The fragment cleavage efficiency is calculated as: Cleavage efficiency = Cleavage peak area / (Cleavage peak area + Uncleaved peak area).
[0381] In this embodiment, the target nucleic acid was selected as double-stranded DNA (PCR product), and the primers were:
[0382] XQ0001-5FAM:GTATGTTGTGTGGAATTGTG 5'6-FAM;
[0383] XQ0002-5ROX:GCTGCGCGTAACCACCACAC 5'ROX
[0384] The amplified product sequence is as follows:
[0385]
[0386] The italicized portion represents the PAM sequence, and the underlined area represents the target region.
[0387] gRNA: Cas-sf8553-5spacer1:
[0388] GUGAGGACGACAACCGAAGGCUCGGUGCGCCGAGAC UCCCCUACGUGCUGCUGAAG (The underlined area is the target area).
[0389] The following reaction system was used: 20 μL system, buffer: Tris-HAC 40 mM, Mg(AC)2 30 mM, BSA 120 μg / ml, DTT 12 mM, pH 7.0, Cas-sf8553 50 nM, gRNA 100 nM, and 1 μL of double-stranded target nucleic acid (PCR product). Incubation was performed at 37℃ for 5 min, 15 min, 30 min, and 60 min, and at 85℃ for 20 min. Products in the FAM and ROX channels were detected by capillary electrophoresis (ABI 3730xl genetic analyzer). Data analysis was performed using Gene Mapper 4.1 software to calculate the NTS / TS cleavage efficiency.
[0390] The results are as follows Figure 13 As shown, Cas-sf8553 cleaves both NTS and TS simultaneously when cleaving the target nucleic acid, with essentially the same cleavage efficiency.
[0391] Example 16. Cleavage characteristics of Cas-sf8553 protein - cleavage site
[0392] This embodiment uses in vitro detection to determine the cleavage sites of the complementary and non-complementary strands of the Cas-sf8553 protein-bound double-stranded target nucleic acid. In this embodiment, gRNA guides the Cas-sf8553 protein to recognize and bind to the double-stranded target nucleic acid; the Cas-sf8553 protein then activates its Cis cleavage activity against the double-stranded target nucleic acid, thereby cleaving the double-stranded target nucleic acid in the system. The cleaved double-stranded target nucleic acid is then filled with an A-linked adapter, followed by PCR enrichment and Sanger sequencing.
[0393] In this embodiment, the target nucleic acid is selected as double-stranded DNA (plasmid), sequence: The vector T-Vector-pEASY-Blunt Simple Cloning Vector is inserted. The italicized part is the PAM sequence, and the underlined area is the target region.
[0394] gRNA: Cas-sf8553-5spacer1:
[0395] GUGAGGACGACAACCGAAGGCUCGGUGCGCCGAGAC UCCCCUACGUGCUGCUGAAG (The underlined area is the target area)
[0396] The following reaction system was used: 50 μL system, buffer: Tris-HAC 40 mM, Mg(AC)2 30 mM, BSA 120 μg / ml, DTT 12 mM, pH 7.0, Ca-sf8553 100 nM, gRNA 250 nM, double-stranded target nucleic acid 10 ng / μL (plasmid). Cas protein and gRNA were incubated at 25°C for 10 min; double-stranded target nucleic acid was added, and the mixture was incubated at 37°C for 1 h, followed by incubation at 85°C for 5 min; 50 μL of 2X Taq DNA Polymerase Mix (Novizan) (1:1) was added to the above system, and the reaction was carried out at 72°C for 30 min; the above reaction solution was recovered; 2 μL of annealed primers (2 μM, TK-117: CGGCATTCCTGCTGAACCGCTCTTCCGATCT, TK-111: GATCGGAAGAGCGGTTCAGCAGGAATGCCG) were added to the recovered liquid, and the mixture was incubated with T4 (NEB) ligase at 22°C for 1 h. 10 μL of the ligation product was taken, and PCR was performed with primers S1-PAM-after: ACTCAGCGGCATTCCTGCTGAACCGC, PQ0275-F: CCGTATTACCGCCTTTGAG, and 2X Taq DNA Polymerase Mix (Novizan). The PCR product was then subjected to Sanger sequencing. The results are as follows. Figure 14 As shown, when the Cas-sf8553 protein cleaves the target nucleic acid, the cleavage sites are between NTS 14-15 nt and TS 21-23 nt. That is, the cleavage site of Cas-sf8553 on the complementary strand of the target sequence is between 21 nt and 23 nt at the 5' end of the PAM complementary sequence, and the cleavage site of Cas-sf8553 on the non-complementary strand of the target sequence is between 14 nt and 15 nt at the 3' end of the PAM sequence. The gRNA guides the Cas-sf8553 protein to recognize and bind to the aforementioned complementary strand, which is the DNA strand paired with the complementary strand.
[0397] Example 17. Application of Cas-sf1846 protein in nucleic acid detection
[0398] This embodiment verifies the trans-cleavage activity of Cas-sf1846 through in vitro detection experiments. In this embodiment, gRNA that can pair with the target nucleic acid guides the Cas-sf1846 protein to recognize and bind to the target nucleic acid; subsequently, the Cas-sf1846 protein is activated to trans-cleave any single-stranded nucleic acid, thereby cleaving the single-stranded nucleic acid detector in the system; the two ends of the single-stranded nucleic acid detector are respectively provided with fluorescent and quenching groups, and fluorescence is excited if the single-stranded nucleic acid detector is cleaved; in other embodiments, the two ends of the single-stranded nucleic acid detector can also be provided with labels that can be detected by colloidal gold.
[0399] In this embodiment, the target nucleic acid was selected as single-stranded DNA, NB-i3g1-ssDNA0, whose sequence is: CGACATTCCGAAGAACGCTGAAGCGCTGGGGGCAAATTGTGCAATTTGCGGC.
[0400] The gRNA sequence corresponding to Cas-sf1846 is: GUUGCAAGCCUUACUGAUUGCUCCUCACGAGGAGAC CCC CCAGCGCUUCAGCGUUC (The underlined area is the target region of the gRNA);
[0401] The single-stranded nucleic acid detector sequence is FAM-TTATT-BHQ1;
[0402] The following reaction system was used: Cas-sf1846 final concentration 50 nM, gRNA final concentration 50 nM, target nucleic acid final concentration 50 nM, and single-stranded nucleic acid detector final concentration 200 nM. Incubation was performed at 37°C, and FAM fluorescence was read every 1 min. The control group did not contain any target nucleic acid.
[0403] like Figure 15 As shown, compared to the control without target nucleic acid, the detection of single-stranded nucleic acids in the Cas-sf1846 cleavage system in the presence of target nucleic acid resulted in rapid fluorescence reporting. These experiments demonstrate that, when used with a single-stranded nucleic acid detector, Cas-sf1846 can be used for the detection of target nucleic acids. Figure 15 In the figures, 1 represents the experimental results with the addition of the target nucleic acid, and 2 represents the control group without the addition of the target nucleic acid.
[0404] Although specific embodiments of the invention have been described in detail, those skilled in the art will understand that various modifications and variations can be made to the details based on all the published teachings, and all such changes are within the scope of protection of the invention. The entire scope of the invention is given by the appended claims and any equivalents thereof.
Claims
1. A Cas protein, characterized in that, The amino acid sequence of the Cas protein is shown in SEQ ID No.
10.
2. An isolated polynucleotide, comprising, The polynucleotide is a polynucleotide sequence encoding the Cas protein of claim 1.
3. A carrier, characterized in that, The vector comprises the polynucleotide of claim 2 and a regulatory element operatively linked thereto.
4. A CRISPR-Cas system, characterized in that, The system includes the Cas protein of claim 1 and at least one gRNA; the gRNA includes a unidirectional repeat sequence capable of binding to the Cas protein of claim 1 and a guide sequence capable of targeting a target sequence.
5. A composition, characterized in that, The composition comprises: (i) A protein component selected from the Cas protein of claim 1; (ii) A nucleic acid component selected from: gRNA, or nucleic acid encoding gRNA, or precursor RNA of gRNA, or precursor RNA nucleic acid encoding gRNA, wherein the gRNA comprises a repetitive sequence capable of binding to the Cas protein of claim 1 and a guide sequence capable of targeting a target sequence. The protein components and nucleic acid components combine to form a complex.
6. An engineered host cell, characterized in that, The host cell comprises the Cas protein of claim 1, or the polynucleotide of claim 2, or the vector of claim 3, or the CRISPR-Cas system of claim 4, or the composition of claim 5.
7. The use of the Cas protein of claim 1, or the polynucleotide of claim 2, or the vector of claim 3, or the CRISPR-Cas system of claim 4, or the composition of claim 5, or the host cell of claim 6 in any one or more of the following, wherein the use is for non-disease diagnosis and treatment purposes: gene editing, gene targeting, gene cutting; Alternatively, in the preparation of formulations or kits for use in gene editing, gene targeting, or gene cutting.
8. A method for detecting a target nucleic acid in a sample, the method being a method for non-disease diagnosis and treatment purposes, the method comprising contacting the sample with the Cas protein of claim 1, gRNA, and a single-stranded nucleic acid detector, the gRNA comprising a region binding to the Cas protein and a guide sequence for hybridization with the target nucleic acid; detecting a detectable signal generated by the single-stranded nucleic acid detector cleaved by the Cas protein, thereby detecting the target nucleic acid; the single-stranded nucleic acid detector not hybridizing with the gRNA.
9. A kit for gene editing, gene targeting, or gene cutting, the kit comprising the Cas protein of claim 1, or the polynucleotide of claim 2, or the vector of claim 3, or the CRISPR-Cas system of claim 4, or the composition of claim 5, or the host cell of claim 6.