Cas protein composition capable of being efficiently edited in rice
By using a combination of Cas protein with a specific nucleic acid sequence and guide RNA in rice, the problem of unstable editing effect of CRISPR/Cas system in plant cells was solved, and stable editing of rice OsGW2 and OsPDS genes was achieved.
Patent Information
- Application Number
- CN202511701008.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-19
- Publication Date
- 2026-01-09
AI Technical Summary
Existing CRISPR/Cas systems are not very effective at editing in plant cells, making it difficult to achieve efficient gene editing.
A composition of Cas protein and guide RNA containing a specific nucleic acid sequence has been developed for rice gene editing. By including a direct repeat sequence and a guide sequence in the guide RNA, efficient hybridization with the target sequence of the plant genome is achieved, and gene editing is performed by binding to a specific Cas protein.
Stable gene editing effects were achieved in rice cells, and the OsGW2 and OsPDS genes were successfully knocked out, demonstrating significant editing efficiency.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of nucleic acid editing. Specifically, it relates to compositions comprising a first nucleic acid and a second nucleic acid encoding a Cas effector protein and a guide RNA, respectively, as well as vectors and host cells comprising the first and second nucleic acids. The invention also relates to complexes for nucleic acid editing (e.g., gene or genome editing), and methods for using said compositions and complexes for nucleic acid editing (e.g., gene or genome editing). Background Technology
[0002] Rice is my country's primary staple food crop, with about 60% of the population relying on it as their main food source. Annual rice production exceeds 200 million tons, ensuring a stable rice supply is crucial for the basic livelihoods of both urban and rural residents. As one of the cereal crops with the highest yield potential, rice plays a vital role in ensuring food security, regulating planting structure, and increasing farmers' income. In scientific research, rice is a "model grass" for functional genomics research. Its compact genome, efficient transformation system, and high-density genetic map make it an ideal platform for analyzing complex agronomic traits, validating gene-editing tools, and conducting synthetic biology research. From "green super rice" to "salt-tolerant rice," each technological breakthrough provides solutions for improving arable land quality, adapting to climate change, and enhancing nutritional quality, propelling my country from a "major rice producer" to a "leading rice science and technology powerhouse."
[0003] The rice OsPDS (Phytoene Desaturase) gene encodes a redox enzyme involved in the biosynthesis of photosynthetic pigments. When the OsPDS gene is disrupted or mutated, it leads to albino and dwarf phenotypes in rice. This is because the synthesis of carotenoids is inhibited, affecting normal photosynthesis. In gene editing research, OsPDS is often used as a target gene because its mutant phenotypes are easily identifiable. Furthermore, the application prospects of this nuclease in rice gene editing can be inferred based on the OsPDS gene editing results.
[0004] Gene editing technology is one of the most commonly used biotechnologies in the biological field, with applications spanning areas such as breeding, disease treatment, drug development, and molecular diagnostics. Currently, CRISPR / Cas systems are mainly classified into two categories based on the composition of their effector proteins: Class 2 systems composed of a single effector protein and Class 1 systems composed of multiple effector proteins. Each category is further divided into three types and multiple subtypes based on evolutionary and functional diversity. The well-known Cas9 system belongs to Type II of the Class 2 system. The CRISPR / Cas9 system requires the Cas9 protein and sgRNA composed of tracrRNA and crRNA to function and achieve cleavage of the target site. Due to its simple design and efficient cleavage, the CRISPR / Cas9 system has become the most widely used gene editing system.
[0005] CRISPR / Cas Type V systems possess a 5' TTN motif and employ sticky-end cleavage of target sequences, such as Cpf1, Cas12i, Cas12j, and Casλ. Unlike Cas9, Cas12i and Cas12j do not require tracrRNA; only a guide RNA is needed to cleave the target sequence. Compared to Cas9 and Cpf1, Cas12i and Cas12j proteins are smaller, less than 1000 amino acids, making them easier to deliver to cells. Casδ also belongs to the Type V CRISPR / Cas system. Chinese patent (202411237279.5, publication date: 20250701) discloses a Type V Cas enzyme (Casδ-1) and its truncated form. The inventors detected the editing of target sequences by Casδ-1 and its truncated form in animal cells and maize protoplasts, but have not established a stable and efficient gene editing system in the staple crop rice.
[0006] Given that Cas protein compositions exhibit good editing effects in animal cells, their editing effectiveness in plant cells is often significantly reduced (e.g., they are uneditable or unstable). Therefore, there is a need to provide a set of Cas protein compositions that offer good editing efficiency and stable editing in plant cells. Summary of the Invention
[0007] Through extensive experimentation and repeated exploration, the inventors of this application unexpectedly discovered a combination of Cas proteins and guide RNAs that exhibit good editing efficiency and stable editing in plant cells. Based on this discovery, the inventors developed a new CRISPR / Cas system composition and a gene editing method based on this composition.
[0008] Composition
[0009] Therefore, in a first aspect, the present invention provides a composition comprising:
[0010] (i) The first nucleic acid, which is the nucleotide sequence encoding the Cas protein; and
[0011] (ii) The second nucleic acid, which is a nucleotide sequence encoding or expressing a guide RNA;
[0012] The guide RNA comprises, from 5' to 3', a unidirectional repeating sequence as shown in SEQ ID NO: 3 or SEQ ID NO: 5, and a guide sequence; wherein the guide sequence is capable of hybridizing with a target sequence derived from a plant genome.
[0013] In some implementations, the target sequence is derived from the plant genome.
[0014] In some implementations, the target sequence is derived from the rice genome.
[0015] In some implementations, the target sequence is derived from the rice OsGW2 gene or the OsPDS gene.
[0016] In some embodiments, the target sequence is as shown in SEQ ID NO: 16 or SEQ ID NO: 12.
[0017] In some embodiments, the guiding sequence is as shown in SEQ ID NO: 14 or SEQ ID NO: 18.
[0018] In some embodiments, the target sequence is derived from the rice OsPDS gene. In some embodiments, the guide RNA comprises a unidirectional repeat sequence as shown in SEQ ID NO: 3 and a guide sequence as shown in SEQ ID NO: 14 from the 5' to 3' direction.
[0019] In some embodiments, the target sequence is derived from the rice OsGW2 gene. In some embodiments, the guide RNA comprises a unidirectional repeat sequence as shown in SEQ ID NO: 3 and a guide sequence as shown in SEQ ID NO: 18 from the 5' to 3' direction.
[0020] In some embodiments, the target sequence is derived from the rice OsPDS gene. In some embodiments, the guide RNA comprises a unidirectional repeat sequence as shown in SEQ ID NO: 5 and a guide sequence as shown in SEQ ID NO: 14 from the 5' to 3' direction.
[0021] In some embodiments, the target sequence is derived from the rice OsGW2 gene. In some embodiments, the guide RNA comprises a unidirectional repeat sequence as shown in SEQ ID NO: 5 and a guide sequence as shown in SEQ ID NO: 18 from the 5' to 3' direction.
[0022] In some implementations, the Cas protein is an effector protein in the CRISPR / Cas system.
[0023] The proteins of the present invention can be derivatized, for example, by being linked to another molecule (e.g., another polypeptide or protein). Generally, protein derivatization (e.g., labeling) does not adversely affect the protein's desired activity (e.g., activity binding to guide RNA, endonuclease activity, activity of binding to and cleaving a target sequence at a specific site guided by guide RNA). Therefore, the proteins of the present invention are also intended to include such derivatized forms. For example, the proteins of the present invention can be functionally linked (by chemical coupling, gene fusion, non-covalent linkage, or other means) to one or more other molecular groups, such as another protein or polypeptide, a detection reagent, a pharmaceutical reagent, etc.
[0024] In particular, the protein of the present invention can be linked to other functional units. For example, it can be linked to a nuclear localization signal (NLS) sequence to enhance the ability of the protein of the present invention to enter the cell nucleus. For example, it can be linked to a targeting moiety to make the protein of the present invention targeted. For example, it can be linked to a detectable tag to facilitate the detection of the protein of the present invention. For example, it can be linked to an epitope tag to facilitate the expression, detection, tracing, and / or purification of the protein of the present invention.
[0025] In some embodiments, the Cas protein is as shown in SEQ ID NO: 1.
[0026] In some implementations, the first nucleic acid and the second nucleic acid are present on the same or different vectors.
[0027] In some embodiments, the first nucleic acid is operatively linked to a first regulatory element (e.g., a promoter). In some embodiments, the second nucleic acid is operatively linked to a second regulatory element (e.g., a promoter).
[0028] In some embodiments, the composition does not contain trans-acting crRNA (tracrRNA).
[0029] In some embodiments, the composition is non-natural or modified. In some embodiments, at least one component of the composition is non-natural or modified.
[0030] In some embodiments, when the target sequence is DNA, the target sequence is located at the 3' end of the protospacer adjacent motif (PAM), and the PAM has a 5'-RYR sequence, where R is A or G and Y is T or C.
[0031] In some implementations, the sequence of the PAM is selected from ATG, ACG, GTG, ATA, ACA, GCA, GTA and / or GCG.
[0032] In some implementations, when the target sequence is RNA, the target sequence does not have a PAM domain restriction.
[0033] In some embodiments, the target sequence is a DNA or RNA sequence derived from prokaryotic or eukaryotic cells. In some embodiments, the target sequence is a non-naturally occurring DNA or RNA sequence.
[0034] In some embodiments, the target sequence is present within the cell. In some embodiments, the target sequence is present in the cell nucleus or cytoplasm (e.g., organelles). In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a prokaryotic cell.
[0035] In some embodiments, the protein is linked to one or more NLS sequences. In some embodiments, the conjugate or fusion protein comprises one or more NLS sequences. In some embodiments, the NLS sequence is linked to the N-terminus or C-terminus of the protein. In some embodiments, the NLS sequence is fused to the N-terminus or C-terminus of the protein.
[0036] Vector and host cell
[0037] In a second aspect, the present invention also provides a carrier comprising a first nucleic acid and a second nucleic acid in the composition as described in the first aspect.
[0038] The vectors of the present invention can be cloning vectors or expression vectors. In some embodiments, the vectors of the present invention are, for example, plasmids, granules, bacteriophages, Cosmids, etc. In some embodiments, the vectors are capable of expressing the proteins, protein truncated forms, fusion proteins, isolated nucleic acid molecules as described in the fifth aspect, or complexes as described in the sixth aspect in a subject (e.g., a mammal, such as a human).
[0039] In a third aspect, the present invention also provides a host cell comprising the vector described above.
[0040] Such host cells include, but are not limited to, prokaryotic cells such as Escherichia coli cells, and eukaryotic cells such as yeast cells, insect cells, plant cells (e.g., cassava, corn, sorghum, soybean, wheat, oat, or rice cells), and animal cells (e.g., mammalian cells, such as mouse cells, human cells, etc.). The cells of the present invention can also be cell lines, such as 293T cells.
[0041] In some implementations, the host cell is a rice cell.
[0042] In a third aspect, the present invention also provides a composite comprising:
[0043] (i) Protein components, which are Cas proteins.
[0044] (ii) A nucleic acid component, which is a guide RNA, wherein the guide RNA comprises, from the 5' to the 3' direction, a homologous repeat sequence as shown in SEQ ID NO: 3 or SEQ ID NO: 5, and a guide sequence; wherein the guide sequence is capable of hybridizing with a target sequence derived from a plant genome;
[0045] The protein component and the nucleic acid component combine to form a complex.
[0046] In some implementations, the target sequence is derived from the plant genome.
[0047] In some implementations, the target sequence is derived from the rice genome.
[0048] In some implementations, the target sequence is derived from the rice OsGW2 gene or the OsPDS gene.
[0049] In some embodiments, the target sequence is as shown in SEQ ID NO: 16 or SEQ ID NO: 12.
[0050] In some embodiments, the guiding sequence is as shown in SEQ ID NO: 14 or SEQ ID NO: 18.
[0051] In some embodiments, the Cas protein is as shown in SEQ ID NO: 1.
[0052] In some implementations, the first nucleic acid and the second nucleic acid are present on the same or different vectors.
[0053] In some implementations, the first nucleic acid is operatively linked to a first regulatory element (e.g., a promoter).
[0054] In some implementations, the second nucleic acid is operatively linked to a second regulatory element (e.g., a promoter).
[0055] In some embodiments, the target sequence is derived from the rice OsPDS gene. In some embodiments, the guide RNA comprises a unidirectional repeat sequence as shown in SEQ ID NO: 3 and a guide sequence as shown in SEQ ID NO: 14 from the 5' to 3' direction.
[0056] In some embodiments, the target sequence is derived from the rice OsGW2 gene. In some embodiments, the guide RNA comprises a unidirectional repeat sequence as shown in SEQ ID NO: 3 and a guide sequence as shown in SEQ ID NO: 18 from the 5' to 3' direction.
[0057] In some embodiments, the target sequence is derived from the rice OsPDS gene. In some embodiments, the guide RNA comprises a unidirectional repeat sequence as shown in SEQ ID NO: 5 and a guide sequence as shown in SEQ ID NO: 14 from the 5' to 3' direction.
[0058] In some embodiments, the target sequence is derived from the rice OsGW2 gene. In some embodiments, the guide RNA comprises a unidirectional repeat sequence as shown in SEQ ID NO: 5 and a guide sequence as shown in SEQ ID NO: 18 from the 5' to 3' direction.
[0059] Delivery and delivery composition
[0060] The compositions and complexes of the present invention can be delivered by any method known in the art. Such methods include, but are not limited to, electroporation, lipid transfection, nuclear transfection, microinjection, acoustic pore effect, gene gun, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendritic transfection, heat shock transfection, nuclear transfection, magnetic transfection, lipid transfection, puncture transfection, optical transfection, reagent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial viruses, etc.
[0061] Therefore, in another aspect, the present invention provides a delivery composition comprising a delivery carrier, and the composition or complex described above.
[0062] In some implementations, the delivery carrier is a particle.
[0063] In some embodiments, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, gene guns, or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).
[0064] Methods and uses
[0065] In another aspect, the present invention provides a method for modifying a target gene, comprising: contacting the target gene with a complex or composition as described above, or delivering it to a cell containing the target gene; wherein the target sequence is present in the target gene.
[0066] In some implementations, the target sequence is derived from the plant genome.
[0067] In some implementations, the target sequence is derived from the rice genome.
[0068] In some implementations, the target sequence is derived from the rice OsGW2 gene or the OsPDS gene.
[0069] In some embodiments, the target sequence is as shown in SEQ ID NO: 16 or SEQ ID NO: 12.
[0070] In some implementations, the modification refers to a break in the target sequence, such as a double-strand break in DNA or a single-strand break in RNA.
[0071] In some embodiments, the modification further includes inserting a foreign nucleic acid into the break.
[0072] In some embodiments, the method is used to modify a target gene in vitro or ex vivo. In some embodiments, the method is not a method for treating humans or animals as a therapy. In some embodiments, the method does not include the step of modifying human germline genetic characteristics.
[0073] In some embodiments, the target gene is present in an in vitro nucleic acid molecule (e.g., a plasmid).
[0074] In some embodiments, the method results in a break in the target sequence (e.g., a double-strand break in DNA or a single-strand break in RNA). In some embodiments, the break results in a reduction in transcription of the target gene.
[0075] In some embodiments, the method further includes contacting the target gene with an editing template (e.g., a foreign nucleic acid) or delivering it to a cell containing the target gene. In such embodiments, the method repairs the broken target gene through homologous recombination with the editing template (e.g., a foreign nucleic acid), wherein the repair results in a mutation, including the insertion, deletion, or substitution of one or more nucleotides of the target gene. In some embodiments, the mutation results in a change in one or more amino acids in a protein expressed from a gene containing the target sequence.
[0076] Therefore, in some embodiments, the modification further includes inserting an editing template (e.g., exogenous nucleic acid) into the break.
[0077] In another aspect, the present invention provides a plant cell or its progeny obtained by the method described above, wherein the plant cell contains modifications not present in its wild type.
[0078] In some embodiments, the plant cells are rice cells.
[0079] Terminology Definition
[0080] In this invention, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the operational steps used herein, such as molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics, and recombinant DNA, are all conventional steps widely used in their respective fields. To better understand this invention, definitions and explanations of relevant terms are provided below.
[0081] As used herein, the terms “guide RNA” and “crRNA” are used interchangeably and have the meanings commonly understood by those skilled in the art. Generally, a guide RNA may comprise a direct repeat sequence and a guide sequence, or consist substantially of or composed of a direct repeat sequence and a guide sequence (also referred to as a spacer sequence in the context of an endogenous CRISPR system). In some cases, the guide sequence is any polynucleotide sequence that is sufficiently complementary to the target sequence to hybridize with said target sequence and guide the specific binding of the CRISPR / Cas complex to said target sequence. In some embodiments, the complementarity between the guide sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% when optimal alignment is achieved. Determining optimal alignment is within the capabilities of those skilled in the art. For example, publicly available and commercially available alignment algorithms and programs exist, such as, but not limited to, ClustalW, the Smith-Waterman algorithm in MATLAB, Bowtie, Geneious, Biopython, and SeqMan.
[0082] As used herein, the term "complex" refers to a ribonucleoprotein complex formed by the binding of guide RNA or crRNA to the Cas protein, which contains a guide sequence that hybridizes to a target sequence and binds to the Cas protein. This ribonucleoprotein complex is capable of recognizing and cleaving polynucleotides that hybridize with the guide RNA or mature crRNA.
[0083] Therefore, in the formation of a CRISPR / Cas complex, a "target sequence" refers to a polynucleotide targeted by a guide sequence designed to be targeted, such as a sequence complementary to the guide sequence, wherein hybridization between the target sequence and the guide sequence will promote the formation of a CRISPR / Cas complex. Perfect complementarity is not required, as long as sufficient complementarity exists to induce hybridization and promote the formation of a CRISPR / Cas complex. The target sequence can contain any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located in the cell nucleus or cytoplasm. In some cases, the target sequence may be located in an organelle of a eukaryotic cell, such as a mitochondrion or chloroplast. The sequence or template that can be used for recombination into a target locus containing the target sequence is called an "edit template," "edit polynucleotide," or "edit sequence." In some embodiments, the edit template is a foreign nucleic acid. In some embodiments, the recombination is homologous recombination.
[0084] In this invention, the term "target sequence" or "target polynucleotide" can refer to any endogenous or exogenous polynucleotide for a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or useless DNA). In some cases, the target sequence is believed to be associated with a protospacer adjacent motif (PAM). The precise sequence and length requirements for the PAM vary depending on the Cas effector enzyme used, but the PAM is typically a 2-5 base pair sequence adjacent to the protospacer sequence (i.e., the target sequence). Those skilled in the art can identify the PAM sequence used with a given Cas effector protein. In this document, "specific motif sequence recognized by the Cas protein" or "motif sequence" refers to the PAM sequence.
[0085] As used herein, the term "vector" refers to a nucleic acid delivery vehicle into which polynucleotides can be inserted. When a vector enables the expression of a protein encoded by the inserted polynucleotide, it is called an expression vector. Vectors can be introduced into host cells through transformation, transduction, or transfection, allowing the genetic material elements they carry to be expressed in the host cells. Vectors are well-known to those skilled in the art and include, but are not limited to: plasmids; phage particles; Cos plasmids; artificial chromosomes, such as yeast artificial chromosomes (YAC), bacterial artificial chromosomes (BAC), or P1-derived artificial chromosomes (PAC); bacteriophages such as λ phage or M13 phage; and animal viruses. Animal viruses that can be used as vectors include, but are not limited to, retrotranscriptoviruses (including lentiviruses), adenoviruses, adeno-associated viruses, herpesviruses (such as herpes simplex virus), poxviruses, baculoviruses, papillomaviruses, and papillomaviruses (such as SV40). A vector may contain multiple elements controlling expression, including but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. Additionally, a vector may contain a replication initiation site.
[0086] As used herein, the term "host cell" refers to a cell that can be used to introduce a vector, including but not limited to prokaryotic cells such as Escherichia coli or Bacillus subtilis, fungal cells such as yeast cells or Aspergillus, insect cells such as S2 Drosophila cells or Sf9, or animal cells such as fibroblasts, CHO cells, COS cells, NSO cells, HeLa cells, BHK cells, HEK 293 cells, or human cells.
[0087] Those skilled in the art will understand that the design of expression vectors can depend on factors such as the selection of host cells to be transformed and the desired expression level. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including proteins, fusion proteins, isolated nucleic acid molecules, etc. (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes) as described herein.
[0088] As used herein, the term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences), for which detailed descriptions can be found in Goeddel, *Gene Expression Technology: Methods in Enzymology*, 185, Academic Press, San Diego, California (1990). In some cases, regulatory elements include those sequences that direct constitutive expression of a nucleotide sequence in many types of host cells and those sequences that direct expression of that nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes). In some cases, regulatory elements can also be directed to express in a time-dependent manner (such as in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue or cell type specific. In some cases, the term "regulatory element" covers enhancer elements such as WPRE; CMV enhancer; R-U5' fragment in the LTR of HTLV-I (Mol. Cell. Biol., Vol. 8(1), pp. 466-472, 1988); SV40 enhancer; and intron sequence between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), pp. 1527-31, 1981).
[0089] As used herein, the term "promoter" has the meaning known to those skilled in the art, referring to a non-coding nucleotide sequence located upstream of a gene that initiates the expression of a downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell under most or all physiological conditions of the cell. An inducible promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when an inducer corresponding to the promoter is present in the cell. A tissue-specific promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when the cell is a cell of the tissue type corresponding to that promoter.
[0090] As used herein, the term “operably linked” is intended to mean that the nucleotide sequence of interest is linked to one or more regulatory elements in a manner that allows the expression of that nucleotide sequence (e.g., in an in vitro transcription / translation system or in the host cell when the vector is introduced into the host cell).
[0091] As used herein, the term "hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonding between the bases of these nucleotide residues. Hydrogen bonding can occur via Watson-Crick base pairing, Hoogstein binding, or any other sequence-specific mechanism. The complex can consist of two strands forming a duplex, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination thereof. Hybridization can constitute a step in a broader process, such as the initiation of PCR or the cleavage of a polynucleotide by an enzyme. A sequence capable of hybridizing with a given sequence is called the "complement" of that given sequence.
[0092] As used herein, the term "expression" refers to the process by which a DNA template is transcribed into polynucleotides (such as mRNA or other RNA transcripts) and / or the transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as "gene products." If the polynucleotides are derived from genomic DNA, expression can include the splicing of mRNA in eukaryotic cells.
[0093] As used in this article, the term "treatment" means to treat or cure a disease, to delay the onset of symptoms of a disease, and / or to slow the progression of a disease.
[0094] As used herein, the term "subject" includes, but is not limited to, various animals, such as mammals, including bovines, equines, sheep, suidae, canines, felines, lagomorphs, rodents (e.g., mice or rats), non-human primates (e.g., macaques or cynomolgus monkeys), or humans. In some embodiments, the subject (e.g., a human) suffers from a condition (e.g., a condition caused by a disease-related gene defect).
[0095] Beneficial effects of the invention
[0096] Given that the editing effect of Cas protein and gRNA combinations, which are effective in animal cells, is often greatly reduced in plant cells (e.g., unable to edit, or unstable), the inventors further explored the combination of Cas protein and gRNA in rice to obtain Cas protein compositions with good editing efficiency and stable editing in plant cells, and discovered two sets of Cas protein compositions capable of stably editing rice genes.
[0097] Specifically, in the functional gene OsGW2 in rice, the combination of Casδ protein with gRNA1 or gRNA5 successfully achieved gene knockout, while other gRNA combinations (gRNA2, gRNA3, gRNA4, gRNA6) did not. Phenotypic experiments further verified that knocking out the OsGW2 gene with the combination of Casδ protein with gRNA1 or gRNA5 resulted in larger rice grains. Similarly, the above combinations yielded similar results in the rice model gene OsPDS.
[0098] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings and examples. However, those skilled in the art will understand that the following drawings and examples are for illustrative purposes only and are not intended to limit the scope of the invention. Various objects and advantages of the present invention will become apparent to those skilled in the art from the following detailed description of the drawings and preferred embodiments. Attached Figure Description
[0099] Figure 1 This is a schematic diagram of the structure of the Casδ plant genome-directed modification backbone vector used in embodiments of the present invention. It also shows the structure of the NLS Casδ NLS expression unit initiated by pZmUbiI and the backbone vector of crRNA initiated by pOsU3.
[0100] Figure 2 To analyze the experimental results of the editing activity of Casδ on the OsPDS site in stable genetic transformation of rice.
[0101] Figure 3 Phenotypic analysis of the OsPDS gene after Casδ editing mutation.
[0102] Sequence information
[0103] The descriptions of the sequences involved in this application are provided in the table below.
[0104] Table 1: Sequence Information Detailed Implementation
[0105] The present invention will now be described with reference to the following embodiments, which are intended to illustrate the invention (and not limit it). Unless otherwise specified, specific conditions in the embodiments are performed under conventional conditions or conditions recommended by the manufacturer. Reagents or instruments used, unless otherwise specified, are all commercially available conventional products. Those skilled in the art will understand that the embodiments are described by way of example and are not intended to limit the scope of protection claimed by the present invention.
[0106] The rice material used for rice conversion in the following examples is Nipponbare (Oryza sativa L.), provided by Wuhan Aidijing Biotechnology Co., Ltd.
[0107] The vector backbone used was pBUE411-Cas9, described in the literature “Xing HL, Dong L, Wang ZP, et al. A CRISPR / Cas9 toolkit for multiplex genome editing in plants. BMC Plant Biol. 2014;14:327. Published 2014 Nov 29. doi:10.1186 / s12870-014-0327-y”, and can also be purchased from Shanghai Qincheng Biotechnology Co., Ltd.
[0108] The primers, DNA synthesis, and sequencing in the following examples were all performed by Sangon Biotech (Shanghai) Co., Ltd.
[0109] Example 1. Combination of Cas protein and gRNA
[0110] The inventors previously discovered a Casδ protein and studied it in animal cells. Given that Cas protein compositions showing good editing effects in animal cells typically exhibit significantly reduced editing efficacy in plant cells (e.g., unable to edit, or unstable), the inventors further explored combinations of Cas protein and gRNA in rice to obtain Cas protein compositions with good editing efficiency and stable editing in plant cells. Unexpectedly, they discovered two sets of Cas protein compositions capable of stably editing rice genes.
[0111] Therefore, this application describes the construction of vectors containing different combinations of Cas proteins and gRNAs for rice. Details are as follows:
[0112] Two types of Cas proteins were selected: the Casδ protein shown in SEQ ID NO: 1, and the truncated Casδ protein shown in SEQ ID NO: 2 (31 amino acids were deleted from the N-terminus compared to the Casδ protein).
[0113] The gRNA (crRNA) contains a direct repeat sequence and a target sequence from the 5' end to the 3' end. Among them, the selected gRNA direct repeat sequences have nine different lengths, as shown in SEQ ID NO: 3 to SEQ ID NO: 11.
[0114] Based on this, taking advantage of the Casδ protein’s ability to recognize the 5'-RYR PAM site, two target fragments were selected in the OsPDS gene, as shown in SEQ ID NO: 12 and SEQ ID NO: 13, respectively. The target sequences of the corresponding gRNAs of specific lengths (26bp) are shown in SEQ ID NO: 14 and SEQ ID NO: 15, respectively.
[0115] Two target fragments were also selected from the OsGW2 gene, as shown in SEQ ID NO:16 and SEQ ID NO:17, respectively, and the target sequences of the corresponding gRNAs of specific lengths (26bp) are shown in SEQ ID NO:18 and SEQ ID NO:19, respectively.
[0116] Example 2. Construction of pBUE411-Casδ-OsPDS editing vector
[0117] To explore the editing efficiency in rice, this embodiment selected the model gene OsPDS (Phytoene Desaturase) in rice for research and constructed the pBUE411-Casδ-OsPDS editing vector.
[0118] Construction of pBUE411-Casδ backbone vector (Casδ protein expression unit)
[0119] Figure 1 This is a schematic diagram of the structure of the Casδ plant genome-directed modification backbone vector used in embodiments of the present invention (the crRNA therein is also referred to as gRNA in this document). The diagram shows the structure of the backbone vector for the NLS Casδ NLS expression unit initiated by pZmUbiI and the crRNA initiated by pOsU3.
[0120] Specifically, plasmid pBUE411-Cas9 was digested with the restriction endonuclease EcoRI to obtain a vector backbone of approximately 13633 bp. Then, the Ubi promoter fragment was amplified using primers Ubi-p-EcoRI-F and Ubi-p-EcoRI-R, the Casδ fragment containing homologous arms was amplified using Casδ-Ubi-pF and Casδ-Ubi-pR, and E9-ter was amplified using primers E9-ter-F and E9-ter-R. The specific primer sequences are shown in Tables 1 and 2. Finally, vector backbone 1, the Ubi promoter fragment, the Casδ fragment containing homologous arms, and the E9-ter fragment (transcription terminator) were subjected to three-fragment homologous recombination. Colony PCR and first-generation sequencing confirmed the successful ligation of the pBUE411-Casδ backbone vector. Ultimately, the pBUE411-Casδ backbone vector, containing Casδ protein expression units, was successfully constructed and used in subsequent experiments.
[0121] Table 2
[0122] Construction of pBUE411-Casδ-OsPDS vector (gRNA expression unit)
[0123] (1) The plasmid pBUE411-Casδ was digested with the restriction endonuclease BsaI to obtain a vector backbone of about 17167 bp.
[0124] (2) Denaturation, annealing and extension were performed using Casδ-OsPDS-F1 and Casδ-OsPDS-R1 primers at a concentration of 100 μmol. The OsPDS target 1 fragment (containing the coding sequence of gRNA) was recovered. The OsPDS target 2 fragment (containing the coding sequence of gRNA) was obtained in the same way. The specific sequences are shown in Table 1 and Table 3.
[0125] (3) Single-fragment homologous recombination was performed between the vector backbone 2 and the OsPDS target fragment. The success of OsPDS target ligation was confirmed by colony PCR and first-generation sequencing. Successful ligation yielded the pBUE411-Casδ-OsPDS vector, which can express gRNA and Casδ protein.
[0126] (4) After culturing the recombinant E. coli containing the gene-editing vector from step (3), extract plasmid DNA and add it to Agrobacterium competent cells. Incubate on ice for 30 min, in liquid nitrogen for 1 min, and incubate in water at 37°C for 2 min. Add 700 μl of culture medium (antibiotic-free), shake and incubate for 3–5 h, and spread it on a culture plate containing the appropriate antibiotic. Incubate upside down in an incubator for two days to obtain recombinant Agrobacterium containing the gene-editing vector.
[0127] (5) The recombinant Agrobacterium was handed over to Wuhan Aidijing Biotechnology Co., Ltd. for stable genetic transformation in rice.
[0128] Table 3
[0129] Example 3. Screening and editing verification of Cas protein and gRNA combinations
[0130] To further explore the editing efficiency of the combination of Cas protein and gRNA in rice, this example selected the functional gene OsGW2 in rice for research (the rice grains become larger after the functional gene OsGW2 is knocked out).
[0131] Two target fragments were selected from the OsGW2 gene, as shown in SEQ ID NO:16 and SEQ ID NO:17, respectively, and the corresponding gRNA target sequences are shown in SEQ ID NO:18 and SEQ ID NO:19, respectively. The specifically designed gRNAs are as follows:
[0132] (1) gRNA1 contains, from the 5' end to the 3' end, a 36 bp (SEQ ID NO: 3) homologous repeat sequence and a 26 bp (SEQ ID NO: 18) target sequence.
[0133] (2) gRNA2 contains, from the 5' end to the 3' end, a 36 bp (SEQ ID NO: 3) homologous repeat sequence and a 26 bp (SEQ ID NO: 19) target sequence.
[0134] (3) gRNA3 contains, from the 5' end to the 3' end, a 34 bp (SEQ ID NO: 4) homologous repeat sequence and a 26 bp (SEQ ID NO: 18) target sequence.
[0135] (4) gRNA4 contains, from the 5' end to the 3' end, a 34 bp (SEQ ID NO: 4) homologous repeat sequence and a 26 bp (SEQ ID NO: 19) target sequence.
[0136] (5) gRNA5 contains, from the 5' end to the 3' end, a 32 bp (SEQ ID NO: 5) homologous repeat sequence and a 26 bp (SEQ ID NO: 18) target sequence.
[0137] (6) gRNA6 contains, from the 5' end to the 3' end, a 32 bp (SEQ ID NO: 5) homologous repeat sequence and a 26 bp (SEQ ID NO: 19) target sequence.
[0138] The specific process is as follows: The pBUE411-Casδ-OsGW2 editing vector was constructed according to the method described in Example 2, enabling it to express the Casδ protein shown in SEQ ID NO: 1 and one of gRNA1 to gRNA6 as described above. The specific primer sequences for vector construction are shown in Tables 1 and 4. The pBUE411-Casδ-OsGW2 editing vector was transformed into rice plants using Agrobacterium tumefaciens to obtain T0 generation pBUE411-Casδ-OsGW2 transformed rice plants.
[0139] Genomic DNA was extracted from T0 generation transgenic rice plants using the CTAB method, and PCR amplification was performed using OsGW2-jc-F and OsGW2-jc-R primers (primer sequences are shown in Table 4) to detect whether the functional gene OsGW2 had been edited.
[0140] Genotyping of the obtained PCR fragments was performed using first-generation sequencing. The sequencing results showed that the combination of Casδ protein with gRNA1 or gRNA5 successfully introduced a small deletion mutation at the target site of the rice OsGW2 gene, leading to premature termination of protein translation and achieving effective gene knockout. Other gRNA combinations (gRNA2, gRNA3, gRNA4, gRNA6) did not show successful editing events in this experiment.
[0141] Phenotypic verification: The edited positive plants had grains with significantly larger lengths and widths than wild-type rice. This phenotype is completely consistent with known research findings—knocking out the OsGW2 gene leads to larger rice grains—proving the success of the editing at a functional level.
[0142] Effects of gRNA Structure: Experimental results show that the length of the direct repeat sequence in the gRNA has a crucial impact on editing efficiency. For example, targeting the same sequence as SEQ ID NO:16, editing was successful when using a 36 bp direct repeat sequence (gRNA1) or a 32 bp direct repeat sequence (gRNA5), but unsuccessful when using a 34 bp direct repeat sequence (gRNA3). This suggests that the Casδ protein may have a specific preference for gRNA structure.
[0143] Furthermore, when targeting the target sequence of SEQ ID NO:17, neither gRNA2, gRNA4, nor gRNA6 were edited, demonstrating that the choice of target sequence also has a significant impact on the editing results.
[0144] Table 4
[0145] Example 4. Validation of the combination of Cas protein and gRNA
[0146] This application utilizes the model gene OsPDS to further test whether different combinations of Cas proteins and gRNAs can be successfully edited, and the stability of successful editing.
[0147] Two target fragments were selected from the OsPDS gene, as shown in SEQ ID NO:12 and SEQ ID NO:13, respectively, and the corresponding gRNA target sequences are shown in SEQ ID NO:14 and SEQ ID NO:15, respectively. The specifically designed gRNAs are as follows:
[0148] (1) gRNA1 contains, from the 5' end to the 3' end, a 36 bp (SEQ ID NO: 3) homologous repeat sequence and a 26 bp (SEQ ID NO: 14) target sequence.
[0149] (2) gRNA2 contains, from the 5' end to the 3' end, a 36 bp (SEQ ID NO: 3) homologous repeat sequence and a 26 bp (SEQ ID NO: 15) target sequence.
[0150] (3) gRNA3 contains, from the 5' end to the 3' end, a 32 bp (SEQ ID NO: 5) homologous repeat sequence and a 26 bp (SEQ ID NO: 14) target sequence.
[0151] (4) gRNA4 contains, from the 5' end to the 3' end, a 34 bp (SEQ ID NO: 5) homologous repeat sequence and a 26 bp (SEQ ID NO: 15) target sequence.
[0152] The specific process is as follows: The pBUE411-Casδ-OsPDS editing vector was constructed according to the method described in Example 2, enabling it to express the Casδ protein shown in SEQ ID NO: 1 and one of gRNA1 to gRNA4 as described above. The specific primer sequences for vector construction are shown in Tables 1 and 5. The pBUE411-Casδ-OsGW2 editing vector was transformed into rice plants using Agrobacterium tumefaciens to obtain T0 generation OsPDS transgenic rice plants. Genomic DNA was extracted from the T0 generation transgenic rice plants using the CTAB method, and PCR amplification was performed using OsPDS-jc-F and OsPDS-jc-R primers. The primer sequences are shown in Table 5.
[0153] Genotyping of the obtained PCR fragments was performed using first-generation sequencing. The sequencing results showed that the combination of Casδ protein with gRNA1 or gRNA3 could successfully introduce a small deletion mutation at the target site of the rice OsPDS gene, leading to premature termination of protein translation. Figure 2 No plants with OsPDS gene editing were detected at target 2 (a combination of Casδ protein and gRNA2 or gRNA4).
[0154] Phenotypic observation of gene-edited positive plants showed that the leaves exhibited leukoplakia ( Figure 3 ).
[0155] Table 5
[0156] Ultimately, the two optimal combinations of Cas protein and gRNA were selected from a variety of combinations: (1) a combination of gRNA with a 36 bp (SEQ ID NO: 3) and a target sequence with a length of 26 bp (e.g., SEQ ID NO: 14) and (2) a combination of gRNA with a 32 bp (SEQ ID NO: 5) and a target sequence with a length of 26 bp (e.g., SEQ ID NO: 14). These combinations not only successfully edited the target sequence but also exhibited significantly higher editing stability among the many combinations.
[0157] Although specific embodiments of the invention have been described in detail, those skilled in the art will understand that various modifications and variations can be made to the details based on all the published teachings, and all such changes are within the scope of protection of the invention. The entire scope of the invention is given by the appended claims and any equivalents thereof.
Claims
1. A composition comprising: (i) The first nucleic acid, which is the nucleotide sequence encoding the Cas protein; and (ii) The second nucleic acid, which is a nucleotide sequence encoding or expressing a guide RNA; in, The guide RNA comprises, from 5' to 3', a unidirectional repeat sequence as shown in SEQ ID NO: 3 or SEQ ID NO: 5, and a guide sequence; wherein the guide sequence is capable of hybridizing with a target sequence derived from a plant genome.
2. The composition of claim 1, wherein, The target sequence is derived from the rice genome; Preferably, the target sequence is derived from the rice OsGW2 gene or the OsPDS gene; Preferably, the target sequence is as shown in SEQ ID NO: 16 or SEQ ID NO:
12.
3. The composition according to claim 1 or 2, having one or more of the following characteristics: (1) The guiding sequence is as shown in SEQ ID NO: 14 or SEQ ID NO: 18; (2) The Cas protein is shown in SEQ ID NO: 1; (3) The first nucleic acid and the second nucleic acid exist on the same or different vectors; (4) The first nucleic acid is operatively connected to a first regulatory element (e.g., a promoter); (5) The second nucleic acid is operatively linked to a second regulatory element (e.g., a promoter).
4. A vector comprising the first nucleic acid and the second nucleic acid in any one of claims 1-3.
5. A host cell comprising the vector of claim 4; Preferably, the host cell is a plant cell; for example, a rice cell.
6. A complex comprising: (i) Protein components, which are Cas proteins; and (ii) A nucleic acid component, which is a guide RNA, wherein the guide RNA comprises, from the 5' to the 3' direction, a unidirectional repeat sequence as shown in SEQ ID NO: 3 or SEQ ID NO: 5, and a guide sequence; wherein, The guide sequence is capable of hybridizing with a target sequence derived from the plant genome; The protein component and the nucleic acid component combine to form a complex.
7. The complex of claim 6, wherein, The target sequence is derived from the rice genome; Preferably, the target sequence is derived from the rice OsGW2 gene or the OsPDS gene; Preferably, the target sequence is as shown in SEQ ID NO: 16 or SEQ ID NO: 12; Preferably, the complex has one or more of the following characteristics: (1) The guiding sequence is as shown in SEQ ID NO: 14 or SEQ ID NO: 18; (2) The Cas protein is shown in SEQ ID NO: 1; (3) The first nucleic acid and the second nucleic acid exist on the same or different vectors; (4) The first nucleic acid is operatively connected to a first regulatory element (e.g., a promoter); (5) The second nucleic acid is operatively linked to a second regulatory element (e.g., a promoter).
8. A method for modifying a target gene, comprising: The composition of any one of claims 1-3 or the complex of claim 6 or 7 is contacted with the target gene or delivered to a cell containing the target gene; wherein the target sequence is present in the target gene.
9. The method of claim 8, wherein the method has one or more features selected from the following: (1) The target sequence is derived from the plant genome; Preferably, the target sequence is derived from the rice genome; Preferably, the target sequence is derived from the rice OsGW2 gene or the OsPDS gene; Preferably, the target sequence is as shown in SEQ ID NO: 16 or SEQ ID NO: 12; (2) The modification refers to the breakage of the target sequence, such as a double-strand break in DNA or a single-strand break in RNA; (3) The modification further includes inserting exogenous nucleic acid into the break; (4) The delivery refers to Agrobacterium-mediated transformation.
10. A plant cell or its progeny obtained by the method of claim 9, wherein the plant cell contains modifications not present in its wild type; Preferably, the plant is rice.
Citation Information
Patent Citations
Novel CRISPR-Casdelta enzymes and systems
CN119193540A