Cas protein compositions edited in maize

By using a combination of nucleic acid and guide RNA encoding Cas protein in maize cells, the problem of low gene editing efficiency in plant cells was solved, achieving stable target site editing and gene knockout, especially in the ZmTS1, ZmPSY1 and ZmDGK1 genes.

CN122405682APending Publication Date: 2026-07-17CHINA AGRI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA AGRI UNIV
Filing Date
2026-05-29
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In plant cells, existing Cas protein compositions have low gene editing efficiency and are unstable, making it difficult to achieve effective target sequence editing.

Method used

A composition is provided comprising nucleic acid encoding a Cas protein and guide RNA, wherein the guide RNA comprises a specific homologous repeat sequence and a target sequence, is capable of hybridizing with a target sequence in a plant genome, and is delivered to plant cells via a vector to cleave the target sequence by binding to endonuclease activity.

Benefits of technology

Stable gene editing was achieved in maize cells, and the editing of target sites was successfully performed, especially the gene knockout of ZmTS1, ZmPSY1 and ZmDGK1 genes, with the editing type being small fragment deletion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

This application discloses a Cas protein composition for editing in maize, belonging to the fields of peptide and genetic engineering technology. The technical problem this application aims to solve is how to improve the gene editing efficiency of plant cells. To this end, this application provides a composition comprising: (i) a first nucleic acid, which is a nucleotide sequence encoding a Cas protein; and (ii) a second nucleic acid, which is a nucleotide sequence encoding or expressing a guide RNA; wherein the guide RNA contains a unidirectional repeat sequence as shown in SEQ ID NO:3 from the 5' to 3' direction, and a guide sequence; wherein the guide sequence is capable of hybridizing with a target sequence derived from the plant genome. In maize, the combination of the Casσ protein with gRNA targeting the target gene successfully achieved gene knockout, and the editing type was as expected, a small fragment deletion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of peptide and genetic engineering technology, and specifically relates to Cas protein compositions edited in maize. Background Technology

[0002] As an important food and feed crop, maize plays an increasingly vital role in global agricultural production. As one of the cereal crops with the highest yield potential, maize holds a crucial position in ensuring food security, adjusting planting structures, and increasing farmers' income. In scientific research, maize is a "model grass" for functional genomics research. Its compact genome, efficient transformation system, and high-density genetic map make it an ideal platform for analyzing complex agronomic traits, validating gene-editing tools, and conducting synthetic biology research.

[0003] Gene editing technology is one of the most commonly used biotechnologies in the biological field, with applications spanning areas such as breeding, disease treatment, drug development, and molecular diagnostics. Currently, CRISPR / Cas systems are mainly classified into two categories based on the composition of their effector proteins: Class 2 systems composed of a single effector protein and Class 1 systems composed of multiple effector proteins. Each category is further divided into three types and multiple subtypes based on evolutionary and functional diversity. The well-known Cas9 system belongs to Type II of the Class 2 system. The CRISPR / Cas9 system requires the Cas9 protein and sgRNA composed of tracrRNA and crRNA to function and achieve cleavage of the target site. Due to its simple design and efficient cleavage, the CRISPR / Cas9 system has become the most widely used gene editing system.

[0004] CRISPR / Cas Type V systems possess a 5'-TTN motif and employ sticky-end cleavage of target sequences, such as Cpf1, Cas12i, Cas12j, and Casλ. Unlike Cas9, Cas12i and Cas12j do not require tracrRNA; only a guide RNA is needed to cleave the target sequence. Compared to Cas9 and Cpf1, Cas12i and Cas12j proteins are smaller, less than 1000 amino acids, making them easier to deliver to cells. Casσ also belongs to the Type V CRISPR / Cas system. Chinese patent (ZL2024 1 1237286.5, publication date: 20250812) discloses a Type V Cas enzyme (Casσ-1). Editing of target sequences by Casσ-1 has been detected in animal cells, but a gene-editing system has not been established in the staple crop maize.

[0005] In the past, genetically modified technology solved the problems of "food security" and "survival" in corn production, achieving stable and guaranteed yields through insect resistance and herbicide tolerance. Gene editing technology, on the other hand, focuses on the future, addressing the issues of "moderate prosperity" and "quality" in corn production through precision breeding. It will cultivate "all-rounder" varieties that are higher-yielding, of higher quality, and more resilient, providing solid technical support for ensuring my country's food security.

[0006] Given that Cas protein compositions exhibit editing effects in animal cells, their editing effectiveness in plant cells is often significantly reduced (e.g., they are uneditable or unstable). Therefore, there is a need to provide a set of Cas protein compositions that demonstrate editing efficiency and stable editing in plant cells. Summary of the Invention

[0007] The technical problem this application aims to solve is how to improve the gene editing efficiency of plant cells. To solve this technical problem, this application provides the following technical solution: In a first aspect, this application provides a composition comprising: (i) The first nucleic acid, which is the nucleotide sequence encoding the Cas protein; and (ii) The second nucleic acid, which is a nucleotide sequence encoding or expressing a guide RNA; The guide RNA comprises a unidirectional repeat sequence as shown in SEQ ID NO:3 from 5' to 3', and a guide sequence; wherein the guide sequence is capable of hybridizing with a target sequence derived from a plant genome.

[0008] In some implementations, the target sequence is derived from the plant genome.

[0009] In some implementations, the target sequence is derived from the maize genome.

[0010] In some implementations, the target sequence is derived from the maize ZmTS1, ZmPSY1, or ZmDGK1 gene.

[0011] In some embodiments, the target sequence is as shown in SEQ ID NO:4, SEQ ID NO:5 or SEQ ID NO:6.

[0012] In some embodiments, the guiding sequence is as shown in SEQ ID NO:7, SEQ ID NO:8 or SEQ ID NO:9.

[0013] In some embodiments, the target sequence is derived from the maize ZmTS1 gene. In some embodiments, the guide RNA comprises a unidirectional repeat sequence as shown in SEQ ID NO:3 and a guide sequence as shown in SEQ ID NO:7 from the 5' to 3' direction.

[0014] In some embodiments, the target sequence is derived from the maize ZmPSY1 gene. In some embodiments, the guide RNA comprises a unidirectional repeat sequence as shown in SEQ ID NO:3 and a guide sequence as shown in SEQ ID NO:8 from the 5' to 3' direction.

[0015] In some embodiments, the target sequence is derived from the maize ZmDGK1 gene. In some embodiments, the guide RNA comprises a unidirectional repeat sequence as shown in SEQ ID NO:3 and a guide sequence as shown in SEQ ID NO:9 from the 5' to 3' direction.

[0016] In some implementations, the Cas protein is an effector protein in the CRISPR / Cas system.

[0017] The protein of this application can be derivatized, for example, by being linked to another molecule (e.g., another polypeptide or protein). Generally, protein derivatization (e.g., labeling) does not adversely affect the protein's desired activity (e.g., activity binding to guide RNA, endonuclease activity, activity of binding to and cleaving a target sequence at a specific site guided by guide RNA). Therefore, the protein of this application is also intended to include such derivatized forms. For example, the protein of this application can be functionally linked (through chemical coupling, gene fusion, non-covalent linkage, or other means) to one or more other molecular groups, such as another protein or polypeptide, a detection reagent, a pharmaceutical reagent, etc.

[0018] Specifically, the protein of this application can be linked to other functional units. For example, it can be linked to a nuclear localization signal (NLS) sequence to enhance the protein's ability to enter the cell nucleus. For example, it can be linked to a targeting region to make the protein targeted. For example, it can be linked to a detectable label to facilitate the detection of the protein. For example, it can be linked to an epitope tag to facilitate the expression, detection, tracing, and / or purification of the protein.

[0019] In some embodiments, the Cas protein is as shown in SEQ ID NO:1.

[0020] In some embodiments, the Cas protein, as shown in SEQ ID NO:2, is fused with a nuclear localization signal and a Flag tag.

[0021] In some implementations, the first nucleic acid and the second nucleic acid are present on the same or different vectors.

[0022] In some embodiments, the first nucleic acid is operatively linked to a first regulatory element (e.g., a promoter). In some embodiments, the second nucleic acid is operatively linked to a second regulatory element (e.g., a promoter).

[0023] In some implementations, the gene editing system does not contain trans-acting crRNA (tracrRNA).

[0024] In some embodiments, the gene-editing system is non-natural or modified. In some embodiments, at least one component of the composition is non-natural or modified.

[0025] In some embodiments, when the target sequence is DNA, the target sequence is located at the 3' end of the protospacer adjacent motif (PAM), and the PAM has a sequence represented by 5'-NTN, where N is A, T, C, or G.

[0026] In some embodiments, the sequence of the PAM is selected from ATG, TTG, CTG, GTG, ATT, TTT, CTT, GTT, ATA, TTA, CTA, GTA, ATC, TTC, CTC and / or GTC.

[0027] In some implementations, when the target sequence is RNA, the target sequence does not have a PAM domain restriction.

[0028] In some embodiments, the target sequence is a DNA or RNA sequence derived from prokaryotic or eukaryotic cells. In some embodiments, the target sequence is a non-naturally occurring DNA or RNA sequence.

[0029] In some embodiments, the target sequence is present within the cell. In some embodiments, the target sequence is present in the cell nucleus or cytoplasm (e.g., organelles). In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a prokaryotic cell.

[0030] In some embodiments, the protein is linked to one or more NLS sequences. In some embodiments, the conjugate or fusion protein comprises one or more NLS sequences. In some embodiments, the NLS sequence is linked to the N-terminus or C-terminus of the protein. In some embodiments, the NLS sequence is fused to the N-terminus or C-terminus of the protein.

[0031] Vector and host cell In a second aspect, this application also provides a carrier comprising a first nucleic acid and a second nucleic acid in the composition as described in the first aspect.

[0032] The vectors used in this application can be cloning vectors or expression vectors. In some embodiments, the vectors used in this application are, for example, plasmids, granules, bacteriophages, Cosmids, etc. In some embodiments, the vectors are capable of expressing the proteins, protein truncated forms, fusion proteins, isolated nucleic acid molecules as described in the fifth aspect, or complexes as described in the sixth aspect in a subject (e.g., a mammal, such as a human). In a third aspect, this application also provides a host cell comprising the vector described above.

[0033] Such host cells include, but are not limited to, prokaryotic cells such as Escherichia coli cells, and eukaryotic cells such as yeast cells, insect cells, plant cells (e.g., cassava, corn, sorghum, soybean, wheat, oat, or rice cells), and animal cells (e.g., mammalian cells, such as mouse cells, human cells, etc.). The cells in this application can also be cell lines, such as 293T cells.

[0034] In some implementations, the host cell is a corn cell.

[0035] This application also provides a complex comprising: (i) Protein components, which are Cas proteins. (ii) A nucleic acid component, which is a guide RNA, wherein the guide RNA comprises, from the 5' to the 3' direction, a homologous repeat sequence as shown in SEQ ID NO:3, and a guide sequence; wherein the guide sequence is capable of hybridizing with a target sequence derived from a plant genome; The protein component and the nucleic acid component combine to form a complex.

[0036] In some implementations, the target sequence is derived from the plant genome.

[0037] In some implementations, the target sequence is derived from the maize genome.

[0038] In some implementations, the target sequence is derived from the maize ZmTS1, ZmPSY1, or ZmDGK1 gene.

[0039] In some embodiments, the target sequence is as shown in SEQ ID NO:4, SEQ ID NO:5 or SEQ ID NO:6.

[0040] In some embodiments, the guiding sequence is as shown in SEQ ID NO:7, SEQ ID NO:8 or SEQ ID NO:9.

[0041] In some embodiments, the Cas protein is as shown in SEQ ID NO:1 or SEQ ID NO:2.

[0042] In some implementations, the first nucleic acid and the second nucleic acid are present on the same or different vectors.

[0043] In some implementations, the first nucleic acid is operatively linked to a first regulatory element (e.g., a promoter).

[0044] In some implementations, the second nucleic acid is operatively linked to a second regulatory element (e.g., a promoter).

[0045] In some embodiments, the target sequence is derived from the maize ZmTS1 gene. In some embodiments, the guide RNA comprises a unidirectional repeat sequence as shown in SEQ ID NO:3 and a guide sequence as shown in SEQ ID NO:7 from the 5' to 3' direction.

[0046] In some embodiments, the target sequence is derived from the maize ZmPSY1 gene. In some embodiments, the guide RNA comprises a unidirectional repeat sequence as shown in SEQ ID NO:3 and a guide sequence as shown in SEQ ID NO:8 from the 5' to 3' direction.

[0047] In some embodiments, the target sequence is derived from the maize ZmDGK1 gene. In some embodiments, the guide RNA comprises a unidirectional repeat sequence as shown in SEQ ID NO:3 and a guide sequence as shown in SEQ ID NO:9 from the 5' to 3' direction.

[0048] Delivery and delivery composition The compositions and complexes of this application can be delivered by any method known in the art. Such methods include, but are not limited to, electroporation, lipid transfection, nuclear transfection, microinjection, acoustic pore effect, gene gun, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendritic transfection, heat shock transfection, nuclear transfection, magnetic transfection, lipid transfection, puncture transfection, optical transfection, reagent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial viruses, etc.

[0049] Therefore, in another aspect, this application provides a delivery composition comprising a delivery carrier and the composition or complex described above.

[0050] In some implementations, the delivery carrier is a particle.

[0051] In some embodiments, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, gene guns, or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).

[0052] Methods and uses In another aspect, this application provides a method for modifying a target gene, comprising: contacting the target gene with the complex or composition described above, or delivering it to a cell containing the target gene; wherein the target sequence is present in the target gene.

[0053] In some implementations, the target sequence is derived from the plant genome.

[0054] In some implementations, the target sequence is derived from the maize genome.

[0055] In some implementations, the target sequence is derived from the maize ZmTS1, ZmPSY1, or ZmDGK1 gene.

[0056] In some embodiments, the target sequence is as shown in SEQ ID NO:4, SEQ ID NO:5 or SEQ ID NO:6.

[0057] In some implementations, the modification refers to a break in the target sequence, such as a double-strand break in DNA or a single-strand break in RNA.

[0058] In some embodiments, the modification further includes inserting a foreign nucleic acid into the break. In some embodiments, the method is used to modify a target gene in vitro or ex vivo. In some embodiments, the method is not a method for treating humans or animals as a therapy. In some embodiments, the method does not include the step of modifying human germline genetic characteristics.

[0059] In some embodiments, the target gene is present in an in vitro nucleic acid molecule (e.g., a plasmid).

[0060] In some embodiments, the method results in a break in the target sequence (e.g., a double-strand break in DNA or a single-strand break in RNA). In some embodiments, the break results in a reduction in transcription of the target gene.

[0061] In some embodiments, the method further includes contacting the target gene with an editing template (e.g., a foreign nucleic acid) or delivering it to a cell containing the target gene. In such embodiments, the method repairs the broken target gene through homologous recombination with the editing template (e.g., a foreign nucleic acid), wherein the repair results in a mutation, including the insertion, deletion, or substitution of one or more nucleotides of the target gene. In some embodiments, the mutation results in a change in one or more amino acids in a protein expressed from a gene containing the target sequence.

[0062] Therefore, in some embodiments, the modification further includes inserting an editing template (e.g., exogenous nucleic acid) into the break.

[0063] In another aspect, this application provides a plant cell or its progeny obtained by the method described above, wherein the plant cell contains modifications not present in its wild type.

[0064] In some embodiments, the plant cells are corn cells.

[0065] The beneficial effects achieved by this application are as follows: While Casσ can effectively edit target sites in animal cells, its editing efficiency is low, and its editing effect in plant cells is often greatly reduced (e.g., unable to edit or unstable). Therefore, in order to obtain a Cas protein composition with better editing efficiency and stable editing in plant cells, the combination of Cas protein and gRNA was further explored in maize protoplasts. A maize editing optimization vector was constructed, and effective editing of target sites was successfully achieved in maize.

[0066] Specifically, in the functional genes ZmTS1, ZmPSY1, or ZmDGK1 in maize, the combination of Casσ protein with gRNA targeting the ZmTS1, ZmPSY1, or ZmDGK1 genes successfully achieved gene knockout, and the editing type was as expected, which was a small fragment deletion. Attached Figure Description

[0067] Figure 1 This is a schematic diagram of the structure of the Casσ plant genome-directed modification backbone vector used in the embodiments of this application. It also shows the structure of the backbone vector for the NLS Casσ NLS expression unit initiated by pZmUbiI and the crRNA initiated by pOsU3.

[0068] Figure 2 Analysis of experimental results on the editing activity of Casσ on the ZmTS1 site during transient protoplast transformation.

[0069] Figure 3 Analysis of experimental results on the editing activity of Casσ on the ZmPSY1 site during transient protoplast transformation.

[0070] Figure 4 Analysis of experimental results on the editing activity of Casσ on the ZmDGK1 site during transient protoplast transformation. Detailed Implementation

[0071] I. Terms used in this application: Examples of resources describing many of the molecular biology-related terms used in this article can be found in the following literature: Alberts et al., Molecular Biology of The Cell, 5th ed., Garland Science Publishing, Inc.: New York, 2007; Rieger et al., Glossary of Genetics: Classical and Molecular, 5th ed., Springer-Verlag: New York, 1991; King et al., A Dictionary of Genetics, 6th ed., Oxford University Press: New York, 2002; and Lewin, GenesIX, Oxford University Press: New York, 2007.

[0072] Any references cited in this article, including, for example, all patents, published patent applications and non-patent publications, are incorporated in their entirety by reference.

[0073] For ease of understanding this application, several terms and abbreviations used herein are defined as follows: In this application, "identity" refers to the similarity of amino acid or nucleotide sequences. The similarity of amino acid sequences (or nucleotide sequences) can be determined using homology search sites on the Internet, such as the BLAST page on the NCBI homepage. For example, in Advanced BLAST 2.1, by using blastp as the program, setting the Expect value to 10, setting all filters to OFF, using BLOSUM62 as the matrix, and setting the Gap existence cost, Perresidue gap cost, and Lambda ratio to 11, 1, and 0.85 (default values) respectively, and performing a search for the similarity of a pair of amino acid sequences, the similarity value (%) can be obtained.

[0074] Specifically, the consistency of 70% or more can be 75% or more. Specifically, the consistency of 75% or more can be 80% or more. Specifically, the consistency of 80% or more can be 85% or more. Specifically, the consistency of 85% or more can be 90% or more. Specifically, the consistency of 90% or more can be 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more. More specifically, the consistency of 70% or more can be at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% consistency.

[0075] When used in a list of two or more items, the term "and / or" means that any of the listed items can be used alone or in combination with any one or more of the listed items. For example, the expression "A and / or B" is intended to mean either or both of A and B, i.e., A alone, B alone, or a combination of A and B. The expression "A, B and / or C" means A alone, B alone, C alone, a combination of A and B, a combination of A and C, a combination of B and C, or a combination of A, B and C.

[0076] The term "comprising" is not intended to be restrictive, but rather inclusive and implies the presence of other elements besides those listed, and can be interpreted as "including but not limited to". The term "comprising" also encompasses the terms "consisting of" and "substantially consisting of". In this document, the terms "including" and "comprise" are used interchangeably.

[0077] The terms “protein,” “peptide,” and “polypeptide” are used interchangeably herein and refer to polymers of amino acid residues linked together by peptide (amide) bonds. These terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, proteins, peptides, or polypeptides are at least 3 amino acids in length. Proteins, peptides, or polypeptides can refer to a single protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide can be modified, for example, by adding chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, linkers for conjugation, functionalization, or other modifications. Proteins, peptides, or polypeptides can also be single molecules or can be multi-molecular complexes. Proteins, peptides, or polypeptides can simply be fragments of naturally occurring proteins or peptides. Proteins, peptides, or polypeptides can be naturally occurring, recombinant, or synthetic, or any combination thereof. Any protein provided herein can be produced by any method known in the art. For example, the proteins provided herein can be produced by recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers.

[0078] The terms "nuclear localization signal," "nuclear localization sequence," or "NLS," used interchangeably in this article, refer to the amino acid sequence that "tags" a protein for transport into the nucleus via nuclear transport. Typically, this signal consists of a short sequence of one or more positively charged lysine or arginine residues exposed on the protein surface. Different nuclear localization proteins may share the same NLS. The NLS functions in the opposite way to the nuclear export signal, aiming to expel the protein from the nucleus.

[0079] As used in this article, the term "fusion protein" refers to a hybrid polypeptide containing protein domains from at least two different proteins. One protein may be located at the N-terminal (N-terminal) portion or the C-terminal (C-terminal) portion of the fusion protein, thus forming an "N-terminal fusion protein" or a "C-terminal fusion protein," respectively.

[0080] The term "biomaterial" refers to any material that carries genetic information and is capable of self-replication or replication within a biological system, such as genes, plasmids, microorganisms, animals, and plants.

[0081] As used in this article, "plant" includes explants, plant parts, seedlings, plantlets, or whole plants at any stage of regeneration or development.

[0082] The term "operationally ligated" can refer to a functional connection between a promoter and transcribed DNA, enabling the promoter to function and initiate transcription of the transcribed DNA. The term "operationally ligated" can also refer to a functional connection between other regulatory elements and a target gene to regulate the transcription and / or expression of the target gene.

[0083] As used herein, an "expression cassette" refers to a cassette containing at least transcribed DNA operatively linked to one or more regulatory elements, typically at least a promoter and a 3' UTR (such as a terminator).

[0084] As used herein, the term "vector" refers to any construct that can be used for transformation purposes, i.e., to introduce heterologous DNA into a host cell. Examples include plasmids, granules, viruses, bacteriophages, or linear or circular DNA.

[0085] In this application, "editing" or "genome editing" means using targeted genome editing technology to produce a targeted mutation, deletion, inversion, or substitution of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1000, at least 2500, at least 5000, or at least 10,000 nucleotides of endogenous plant genome nucleic acid sequence.

[0086] In this application, “editing” or “genome editing” also covers the use of targeted genome editing technology to target or integrate at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 75, at least 100, at least 250, at least 500, at least 750, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, or at least 10,000 nucleotides into the endogenous genome of a plant.

[0087] In this application, a “target site” for genome editing refers to a location within a plant genome of a polynucleotide sequence that is targeted and cleaved by a site-specific nuclease, thereby introducing a double-strand break (or single-strand nick) into the nucleic acid backbone and / or its complementary DNA strand. The site-specific nuclease may bind to the target site, for example, via a non-coding guide RNA (e.g., but not limited to CRISPR RNA (crRNA) or single-strand guide RNA (sgRNA)). The non-coding guide RNA provided herein may be complementary to the target site (e.g., complementary to the strand of a double-stranded nucleic acid molecule or the chromosome of the target site). A “target site” also refers to a location within the plant genome of a polynucleotide sequence that is bound and cleaved by another site-specific nuclease, which may not be guided by a non-coding RNA molecule, such as a broad-spectrum nuclease, zinc finger nuclease (ZFN), or transcription activator-like effector nuclease (TALEN), to introduce a double-strand break (or single-strand nick) into the polynucleotide sequence and / or its complementary DNA strand.

[0088] In this application, the term "guide RNA" or "gRNA" is a short RNA sequence comprising (1) a structural or scaffold RNA sequence required to bind or interact with RNA-guided nucleases and / or other RNA molecules (e.g., tracrRNA), and (2) an RNA sequence that is identical or complementary to a target sequence or target site (referred to herein as the "guide sequence"). A "single-stranded guide RNA" (or "sgRNA") is an RNA molecule comprising tracrRNA and crRNA covalently linked by a linker sequence, which may be expressed as a single RNA transcript or molecule. The guide RNA comprises a guide or target sequence ("guide sequence") identical or complementary to a target site within the plant genome, for example, at or near a GA oxidase gene. An interphase sequence adjacent motif (PAM) may be present immediately adjacent to the 5' end of a genomic target site sequence complementary to the guide RNA's target sequence and upstream of it in the genome, i.e., downstream (3') of the sense (+) strand immediately adjacent to the genomic target site (relative to the guide RNA's target sequence), as is known in the art. The genomic PAM sequence (relative to the target sequence of the guide RNA) on the sense (+) strand adjacent to the target site may contain 5'-NGG-3'. However, the corresponding sequence of the guide RNA (i.e., immediately downstream (3') of the target sequence of the guide RNA) is typically not complementary to the genomic PAM sequence. The guide RNA can usually be a non-coding RNA molecule that does not encode a protein.

[0089] II. Implementation Examples The present application will now be described in further detail with reference to specific embodiments. The embodiments given are merely illustrative of the present application and are not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the present application in any way.

[0090] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.

[0091] The backbone vector of the recombinant Casσ protein expression vector, pBUE411(Bar), was kindly provided by Professor Chen Qijun's research group at the College of Biological Sciences, China Agricultural University, and is disclosed in the literature "Xing HL, Dong L, Wang ZP, et al. A CRISPR / Cas9toolkit for multiplex genome editing in plants. BMC Plant Biol. 2014;14:327.Published 2014 Nov 29. doi:10.1186 / s12870-014-0327-y." The public can obtain the above-mentioned biological material from the applicant. The obtained biological material is only for repeating the experiments of this application and may not be used for other purposes.

[0092] The primers, DNA synthesis, and sequencing in the following examples were all performed by Sangon Biotech (Shanghai) Co., Ltd.

[0093] Example 1. Combination of Cas protein and gRNA We previously discovered a Casσ protein and studied it in animal cells. However, Cas protein compositions that showed good editing effects in animal cells often exhibited significantly reduced editing efficacy in plant cells (e.g., inability to edit, or instability). Therefore, to obtain Cas protein compositions with good editing efficiency and stable editing in plant cells, we further explored combinations of Casσ protein and gRNA in maize and unexpectedly discovered a group of Casσ protein compositions (gene editing systems) capable of editing maize genes.

[0094] Therefore, this application describes a gene editing system for maize that includes the Casσ protein and different combinations of gRNAs. Details are as follows: The selected Cas protein is the Casσ protein shown in SEQ ID NO:1, and the Casσ fusion protein with the Casσ fusion nuclear localization signal is shown in SEQ ID NO:2. From the N-terminus, positions 1 to 9 are the amino acid sequence of SV40-NLS, positions 10 to 32 are the amino acid sequence of 3×FLAG, positions 33 to 914 are the amino acid sequence of Casσ-1, positions 915 to 918 are the linker peptide, and positions 919 to 933 are the amino acid sequence of the nucleoplasmin NLS signal peptide.

[0095] The coding sequence of the Casσ fusion protein is shown in SEQ ID NO:10. In the sequence shown in SEQ ID NO:10, positions 1 to 27 from the 5' end are the nucleotide sequence of SV40-NLS, positions 28 to 96 are the nucleotide sequence of 3×FLAG, positions 97 to 2742 are the nucleotide sequence of Casσ-1, positions 2743 to 2754 are the linker peptide, and positions 2755 to 2802 are the nucleotide sequences of the nucleoplasminNLS signal peptide.

[0096] The gRNA (crRNA) contains a direct repeat sequence and a target sequence from the 5' end to the 3' end. The selected direct repeat sequence of the gRNA is shown in SEQ ID NO:3.

[0097] Based on this, taking advantage of the Casσ protein’s ability to recognize the 5'-NTN PAM site, one target fragment was selected from each of the ZmTS1, ZmPSY1 and ZmDGK1 genes, as shown in SEQ ID NO:4-6, and the corresponding target sequences of the specific length (26bp) gRNA are shown in SEQ ID NO:7-9.

[0098] Example 2. Construction of the pBUE411-Casσ-ZmTS1 editing vector To explore the editing efficiency in maize, this embodiment selected the model gene ZmTS1 in maize to study the gene editing efficiency of the Casσ gene editing system on plant cells.

[0099] The genomic sequence of the ZmTS1 gene is SEQ ID NO:15, and the target sequence is located in the first intron.

[0100] This embodiment also selected the ZmPSY1 and ZmDGK1 genes from maize to study the gene editing efficiency of the Casσ gene editing system in plant cells.

[0101] The genomic sequence of the ZmPSY1 gene is SEQ ID NO:16, and the target sequence is located in the 3rd exon.

[0102] The genomic sequence of the ZmDGK1 gene is SEQ ID NO:17, and the target sequence is located in the first exon.

[0103] 2.1 Construction of pBUE411-Casσ backbone vector (Casσ protein expression unit) Figure 1 This is a schematic diagram of the structure of the Casσ plant genome-directed modification backbone vector used in the embodiments of this application (the crRNA therein is also referred to as gRNA herein). The diagram shows the structure of the backbone vector for the NLS Casσ NLS expression unit initiated by pZmUbiI and the crRNA initiated by pOsU3.

[0104] An expression cassette (DNA fragment) of Casσ was constructed, and the nucleotide sequence of the Casσ expression cassette is shown in SEQ ID NO:11. Specifically, positions 1 to 1992 of SEQ ID NO:11 are the nucleotide sequence of the Ubi promoter, positions 2005 to 4809 are the coding sequence of the Casσ fusion protein, and positions 4816 to 5450 are the nucleotide sequence of the E9-ter terminator.

[0105] The structure of the pBUE411-Casσ backbone vector is to insert a DNA molecule with a nucleotide sequence as shown in SEQ ID NO:11 between the EcoRI restriction sites of plasmid pBUE411(Bar).

[0106] 2.2 Construction of the pBUE411-Casσ-ZmTS1 vector (gRNA expression unit) (1) The plasmid pBUE411-Casσ was digested with the restriction endonuclease BsaI to obtain a vector backbone of about 18469 bp.

[0107] (2) Annealing synthesizes ZmTS1 target fragment (containing the coding sequence of gRNA), ZmPSY1 target fragment and ZmDGK1 target fragment.

[0108] The nucleotide sequence of the ZmTS1 target fragment is shown in SEQ ID NO:12, where positions 21 to 56 of SEQ ID NO:12 are gRNA repetitive sequences, and positions 57 to 82 are the nucleotide sequences corresponding to the target sequence of ZmTS1 gRNA.

[0109] The nucleotide sequence of the ZmPSY1 target fragment is shown in SEQ ID NO:13, where positions 21 to 56 of SEQ ID NO:13 are gRNA repetitive sequences, and positions 57 to 82 are the nucleotide sequences corresponding to the target sequence of ZmPSY1 gRNA.

[0110] The nucleotide sequence of the ZmDGK1 target fragment is shown in SEQ ID NO:14, where positions 21 to 56 of SEQ ID NO:14 are gRNA repetitive sequences, and positions 57 to 82 are the nucleotide sequences corresponding to the target sequence of ZmDGK1 gRNA.

[0111] (3) Single-fragment homologous recombination was performed between vector backbone 2 and the ZmTS1 target fragment, ZmPSY1 target fragment, or ZmDGK1 target fragment, respectively. The success of ligation of the corresponding target fragment was confirmed by colony PCR and first-generation sequencing. Successful ligation yielded the pBUE411-Casσ-ZmTS1, pBUE411-Casσ-ZmPSY1, and pBUE411-Casσ-ZmDGK1 vectors.

[0112] The pBUE411-Casσ-ZmTS1 vector is constructed by inserting the ZmTS1 target fragment between the BsaI restriction site of plasmid pBUE411-Casσ, while keeping the other nucleotide sequences of plasmid pBUE411-Casσ unchanged. The pBUE411-Casσ-ZmTS1 vector can express the Casσ fusion protein and gRNA hybridized to SEQ ID NO:4. The Casσ fusion protein and gRNA form a complex with gene-editing activity.

[0113] The pBUE411-Casσ-ZmPSY1 vector is constructed by inserting the ZmPSY1 target fragment between the BsaI restriction site of plasmid pBUE411-Casσ, while keeping the other nucleotide sequences of plasmid pBUE411-Casσ unchanged. The pBUE411-Casσ-ZmPSY1 vector can express the Casσ fusion protein and gRNA hybridized to SEQ ID NO:5. The Casσ fusion protein and gRNA form a complex with gene-editing activity.

[0114] The pBUE411-Casσ-ZmDGK1 vector is constructed by inserting the ZmDGK1 target fragment between the BsaI restriction site and the target site of plasmid pBUE411-Casσ, while maintaining the other nucleotide sequences of plasmid pBUE411-Casσ unchanged. The pBUE411-Casσ-ZmDGK1 vector can express the Casσ fusion protein and gRNA hybridized to SEQ ID NO:6. The Casσ fusion protein and gRNA form a complex with gene-editing activity.

[0115] (4) After culturing the recombinant Escherichia coli containing the gene editing vector in step (3), a large amount of plasmid DNA was extracted and protoplast transformation was performed.

[0116] 2.3 Preparation and transformation of maize protoplasts The extraction and transformation methods for maize protoplasts were based on the experimental methods of Yoo et al. (Yoo et al., 2007), as follows: Preparation of protoplasts: (1) Cultivate LH244 etiolated corn seedlings. Corn can grow in the dark for 10-14 days. Take the middle part of the leaves with good growth as the material for preparing protoplasts.

[0117] (2) Prepare the required enzyme hydrolysate and 0.6 M mannitol equilibrium solution according to the experimental needs.

[0118] (3) After removing the leaves, stack them neatly together, cut them into thin strips with a sharp blade, and immerse them in 0.6 M mannitol for 10 min to equilibrate.

[0119] (4) After equilibration, filter through a 200-mesh sieve, place the cut leaves into the prepared enzymatic hydrolysate, and vacuum for 30 minutes. Incubate at 28℃ in a constant temperature shaker at 40 rpm in the dark for 3-4 hours. After enzymatic hydrolysis, shake at 80 rpm for 5 minutes to separate the protoplasts as much as possible. The enzymatic hydrolysate will turn yellow.

[0120] (5) Filter the enzyme-digested solution through a 200-mesh sieve, slowly collect it into a 50 mL centrifuge tube, centrifuge at 110 g at room temperature for 5 min with an upward speed of 1 g and a downward speed of 1 g, and carefully remove the supernatant.

[0121] (6) Slowly add an appropriate amount of W5 solution along the wall, gently resuspend the precipitate, centrifuge at 110 g at room temperature for 5 min, with an ascending speed of 1 g and a descending speed of 1 g, and carefully remove the supernatant.

[0122] (7) Add an appropriate amount of W5 solution slowly along the wall again, gently resuspend the precipitate, and place it on ice for 45-60 min.

[0123] (8) Use a pipette or tube to aspirate as much of the supernatant as possible, and add an appropriate amount of MMG until the number of protoplasts is 12 × 10⁻⁶. 6 / mL.

[0124] The specific preparation of the buffer solution used in the above experiment is as follows: Table 1. Formulas for relevant mother liquors (stored at 4 ℃) (solvent is deionized water)

[0125] Table 2 W5 (stored at 4 ℃) solution

[0126] Table 3 MMG (stored at 4 ℃) solution

[0127] Table 4 40% PEG (prepare fresh before use) solution

[0128] Table 5. Preparation of Enzyme Hydrolysate (15 mL):

[0129] Transformation of protoplasts: (1) Add 60 µg of the plasmid to be transformed (pBUE411-Casσ-ZmTS1, pBUE411-Casσ-ZmPSY1 or pBUE411-Casσ-ZmDGK1 vector, with a concentration greater than 1 µg / µL) to a 2 mL centrifuge tube, then add 200 µL of protoplast and 40% PEG (the amount of PEG is the sum of the volumes of the plasmid and protoplast), gently mix, and react at room temperature in the dark for 18 min.

[0130] (2) After the reaction is complete, add twice the volume of W5 solution to terminate the reaction. Centrifuge at 110 g at room temperature for 5 min with an upward speed of 1 g and a downward speed of 1 g. Carefully remove the supernatant.

[0131] (3) Add an appropriate amount of W5 to gently resuspend the precipitate, centrifuge at 110 g at room temperature for 5 min with an upward speed of 1 g and a downward speed of 1 g, and carefully remove the supernatant.

[0132] (4) Add an appropriate amount of W5 solution to gently resuspend the centrifuge tube, place it flat, and incubate at 37°C in the dark for 44-48 h.

[0133] (5) Centrifuge at 110 g at room temperature for 5 min, with an upward speed of 1 g and a downward speed of 1 g. Carefully remove the supernatant and collect the protoplasts for further experiments.

[0134] Two positive protoplasts were prepared by transforming the pBUE411-Casσ-ZmTS1 vector, two positive protoplasts were prepared by transforming the pBUE411-Casσ-ZmPSY1 vector, and two positive protoplasts were prepared by transforming the pBUE411-Casσ-ZmDGK1 vector.

[0135] 2.4 Editing Efficiency Test (6) DNA was extracted from maize protoplasts, and the sequences containing the target site (200 bp) were amplified. Specifically, the amplified ZmTS1 gene fragment containing the target site (200 bp) was located at positions 467 to 671 of SEQ ID NO:15; the amplified ZmPSY1 gene fragment containing the target site (200 bp) was located at positions 1123 to 1335 of SEQ ID NO:16; and the amplified ZmDGK1 gene fragment containing the target site (200 bp) was located at positions 75 to 300 of SEQ ID NO:16. The PCR products were sent to Beijing Geneplus Medical Laboratory Co., Ltd. for next-generation library construction and sequencing.

[0136] The results are as follows Figures 2 to 4 As shown. The results indicate that the codon-optimized Casσ for maize can effectively edit endogenous genes. Editing efficiency is the ratio of the number of reads that underwent editing (base deletion, insertion, or mutation) to the total number of reads. Specifically, ZmTS1 had an editing efficiency of 10.74%, with a total of 769,247 reads and 80,540 edited reads. Some editing types are shown below. Figure 2 Among them, ZmPSY1 has an editing efficiency of 6.15%, a total of 899,955 reads, and 55,347 edited reads. Some editing types include... Figure 3 The ZmDGK1 editing efficiency was 7.51%, with a total of 784,513 reads and 58,917 edited reads. Some editing types, such as... Figure 4 .

[0137] Partial sequences in this application: SEQ ID NO:1, the amino acid sequence of the Casσ protein: MSNYKNIKFKLVPFSQKDLINMQLNVNLHQQCYREFVEQFCVLCNIPFPGLSKDQIEQKRKQLNLSEDDEKDINYIKDLVKNKNNIGNSIYAFFTGTKKEMPSRKTDLTPLYRLLKANILPFSLLKGRENYKKSIFQTVINQTLEKFKSYFKCNESVENNFKLSLNKDSNEEQVLNESEMKDLQNLFENLSKNQSFSFFNFNKNWFSKDKIKTKLLNNETNKIKSLSSEEIDLILSYKDKLYSNEFDLISMFVEFNLQKQKAESLKSQADLNLFKNNNYSFRIGSNYENFNLTQNNKDILLEINSSMGEKITFKIIPHKKTQIWNLEKNNVKITSGENLGNYKSVDVIKMKRPADIKAKLLKTSELNIEIKNNQIYCNFIYEYKCSDHGVYFFHCSGNKKPDEKNENILKERERTFSFIDLGLFPMYSISTFKYNNKSNDGEILVKSGSGNEKLDFGSAFKIHSIQIGKNSTNLNKIKQLLEKLKDLKTYLKFSKSISSFDENSYQRQLKTGVEISELNSLSFQKISEIKSINLGFNESFNKEYFLKLIENQTFTQKELLLLNCKIKDLFKILYKEYSNIKNSRIFKFNKEDDLICDGYYWLQVIDEIINIKKSLTYFNSKPSEKGNKSKFIFLKDFNYKNNFANNYAKIAASRLKKYCLEHKVDVCVFEKNLNNFLQSKDNDKKTNKTLINWANRNLFEKIKLALEEHDICVSEVDGKHSSQLDPQTMNWGARDNLNGNGNKEKIFFERNGQIIQQNADLSASEVLAKRFFTRYEDIVHIYIDQKIKDDKTILKLVKGKVRVESYLKKTINSCYAIVDENGFLKPISKKDYNKFQELPSKPRTDIKSNEMYRHGSKWYHFQQHREFQQDLLARGRELKKIA。

[0138] SEQ ID NO:2, Amino acid sequence of the Casσ fusion protein: 。

[0139] SEQ ID NO:3, gRNA direct repeat sequence: 5'-AGTGCAATAGTTACAGAATAGTAATTATATTCGCAC-3'.

[0140] SEQ ID NO:4, ZmTS1 target: 5'-CAGGACGCACGCCAAGTCGCCACCAA-3'.

[0141] SEQ ID NO:5, ZmPSY1 target: 5'-AGCTTGTAGATGGGCCAAACGCCAAC-3'.

[0142] SEQ ID NO:6, ZmDGK1 target: 5'-TGGGATGCAGCGGGCACCTTGTCCTT-3'.

[0143] SEQ ID NO:7, Guide sequence of ZmTS1 gRNA: 5'-CAGGACGCACGCCAAGUCGCCACCAA-3'.

[0144] SEQ ID NO:8, Guide sequence of ZmPSY1 gRNA: 5'-AGCUUGUAGAUGGGCCAAACGCCAAC-3'.

[0145] SEQ ID NO:9, Guide sequence of ZmDGK1 gRNA: 5'-UGGGAUGCAGCGGGCACCUUGUCCUU-3'.

[0146] SEQ ID NO:10, coding sequence of the Casσ fusion protein:

[0147] The nucleotide sequence of SEQ ID NO:11, Ubi-Casσ-E9-ter:

[0148] SEQ ID NO:12, ZmTS1 target fragment: 5'-CAGCTGCGACCACCGTCCgAGTGCAATAGTTACAGAATAGTAATTATATTCGCACCAGGACGCACGCCAAGTCGCCACCAATTTTTTTAAGCTTCTGCAGTG-3'.

[0149] SEQ ID NO:13, ZmPSY1 target fragment: 5'-CAGCTGCGACCACCGCTCCgAGTGCAATAGTTACAGAATAGTAATTATATTCGCACAGCTTGTAGATGGGCCAAACGCCAACTTTTTTTAAGCTTCTGCAGTG-3'.

[0150] SEQ ID NO:14, ZmDGK1 target fragment: 5'-CAGCTGCGACCACCGCTCCgAGTGCAATAGTTACAGAATAGTAATTATATTCGCACTGGGATGCAGCGGGCACCTTGTCCTTTTTTTTTAAGCTTCTGCAGTG-3'.

[0151] SEQ ID NO:15, ZmTS1 gene:

[0152] SEQ ID NO:16, ZmPSY1 gene:

[0153] SEQ ID NO:17, ZmDGK1 gene:

[0154] The present application has been described in detail above. Those skilled in the art will recognize that the present application can be implemented in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments are given in this application, it should be understood that further modifications can be made to the present application. In summary, in accordance with the principles of this application, this application is intended to include any changes, uses, or improvements to the present application, including changes made using conventional techniques known in the art that depart from the scope disclosed herein.

Claims

1. A composition comprising: (i) The first nucleic acid, which is the nucleotide sequence encoding the Cas protein; and (ii) A second nucleic acid, which is a nucleotide sequence encoding or expressing a guide RNA; in, The guide RNA comprises a unidirectional repeat sequence as shown in SEQ ID NO:3 from 5' to 3', and a guide sequence; wherein the guide sequence is capable of hybridizing with a target sequence derived from a plant genome.

2. The composition of claim 1, wherein, The target sequence is derived from the maize genome; Preferably, the target sequence is derived from the maize ZmTS1 gene, ZmPSY1 gene, or ZmDGK1 gene; Preferably, the target sequence is as shown in SEQ ID NO:4, SEQ ID NO:5 or SEQ ID NO:

6.

3. The composition according to claim 1 or 2, having one or more of the following characteristics: (1) The guiding sequence is as shown in SEQ ID NO:7, SEQ ID NO:8 or SEQ ID NO:9; (2) The Cas protein is shown in SEQ ID NO:1; (3) The first nucleic acid and the second nucleic acid exist on the same or different vectors; (4) The first nucleic acid is operatively connected to a first regulatory element (e.g., a promoter); (5) The second nucleic acid is operatively linked to a second regulatory element (e.g., a promoter).

4. A vector comprising the first nucleic acid and the second nucleic acid in any one of claims 1-3.

5. A host cell comprising the vector of claim 4; Preferably, the host cell is a plant cell; for example, a corn cell.

6. A complex comprising: (i) Protein components, which are Cas proteins; and (ii) A nucleic acid component, which is a guide RNA, wherein the guide RNA comprises, from the 5' to the 3' direction, a unidirectional repeat sequence as shown in SEQ ID NO:3, and a guide sequence; wherein, The guide sequence is capable of hybridizing with a target sequence derived from the plant genome; The protein component and the nucleic acid component combine to form a complex.

7. The complex of claim 6, wherein, The target sequence is derived from the maize genome; Preferably, the target sequence is derived from the maize ZmTS1 gene, ZmPSY1 gene, or ZmDGK1 gene; Preferably, the target sequence is as shown in SEQ ID NO:4, SEQ ID NO:5 or SEQ ID NO:6; Preferably, the complex has one or more of the following characteristics: (1) The guiding sequence is as shown in SEQ ID NO:7, SEQ ID NO:8 or SEQ ID NO:9; (2) The Cas protein is shown in SEQ ID NO:1; (3) The first nucleic acid and the second nucleic acid exist on the same or different vectors; (4) The first nucleic acid is operatively connected to a first regulatory element (e.g., a promoter); (5) The second nucleic acid is operatively linked to a second regulatory element (e.g., a promoter).

8. A method for modifying a target gene, comprising: The composition of any one of claims 1-3 or the complex of claim 6 or 7 is contacted with the target gene or delivered to a cell containing the target gene; wherein the target sequence is present in the target gene.

9. The method of claim 8, wherein the method has one or more features selected from the following: (1) The target sequence is derived from the plant genome; Preferably, the target sequence is derived from the maize genome; Preferably, the target sequence is derived from the maize ZmTS1 gene, ZmPSY1 gene, or ZmDGK1 gene; Preferably, the target sequence is as shown in SEQ ID NO:4, SEQ ID NO:5 or SEQ ID NO:6; (2) The modification refers to the breakage of the target sequence; (3) The modification also includes inserting exogenous nucleic acid into the break.

10. A plant cell or its progeny obtained by the method of claim 9, wherein the plant cell contains modifications not present in its wild type; Preferably, the plant is corn.

Citation Information

Patent Citations

  • Novel CRISPR-Cassigma enzymes and systems

    CN119193541A