Novel CRISPR-CAS system for genome editing
The novel Cas-alpha endonuclease addresses the limitations of existing genome editing technologies by providing a PAM-dependent, cost-effective, and specific CRISPR-Cas system for precise genome editing in eukaryotic cells.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- PIONEER HI BREED INTERNATIONAL INC
- Filing Date
- 2025-12-18
- Publication Date
- 2026-04-22
AI Technical Summary
Existing genome editing technologies, such as ZF N, TALEN, and homing meganucleases, suffer from low specificity and require significant time and expense to redesign for each target site, necessitating the development of novel CRISPR-Cas systems with improved targeting capabilities across various organisms, particularly in eukaryotes like animals and plants.
A novel CRISPR-Cas endonuclease, termed Cas-alpha, is developed with specific domains and motifs, enabling targeted double-strand DNA cleavage in a PAM-dependent manner, optimized for expression in eukaryotic cells, including plant, fungal, and animal cells, using guide polynucleotides for precise genome editing.
Cas-alpha endonucleases demonstrate high specificity and efficiency in cutting double-strand DNA across different kingdoms, facilitating precise genome editing in eukaryotic cells with reduced design and production costs.
Smart Images

Figure 2026068736000024 
Figure 2026068736000025 
Figure 2026068736000026
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application is based on U.S. Provisional Application No. 62 / 779989, filed on December 14, 2018. U.S. Provisional Application No. 62 / 794427, filed on January 18, 2019, and due March 1, 2019 U.S. Provisional Application No. 62 / 819409, filed on the 5th, was filed on May 24, 2019. U.S. Provisional Application No. 62 / 852788, and U.S. Provisional Application No. 10, 2019 This asserts the interests of application No. 62 / 913492, all of which are by reference. The entirety of this is incorporated herein.
[0002] Reference to electronically submitted sequence listings A formal copy of the sequence listing, having a size of 714,386 bytes, is submitted concurrently with this specification. This was created on December 9, 2019, and is named RTS21920B_SequenceLi An array table in ASCII format with the filename sting_ST25.txt and It is submitted electronically via EFS-Web. This ASCII format document includes The sequence listings contained herein are part of this specification and are incorporated herein by reference as a whole. .
[0003] This disclosure relates to the field of molecular biology, and in particular to novel RNA-induced Cas endonuclei. The present invention relates to a composition of a ze system, and to compositions and methods for editing or modifying the genome of a cell. do. [Background technology]
[0004] Recombinant DNA technology allows for the insertion and / or identification of DNA sequences at target genomic locations. It has become possible to modify the endogenous chromosome sequence using a site-directed recombination system. Site-specific insertion technology, like other recombination technologies, can be used to target genes in various organisms. It has been used for generating target insertions for offspring. Designer zinc finger nuclease (ZF N), transcription-activating effector nuclease (TALEN), or homing meganuclease Genome editing technologies such as rease can be used to generate targeted genome perturbations. However, these systems have low specificity and need to be redesigned for each target site, therefore Therefore, there is a tendency to use specially designed nucleases that require significant expense and time to produce.
[0005] Effector tans encompass diverse activities (DNA recognition, binding, and selective cleavage). CRISPR (Clustered Regular Arrangement Short Palinus Sequence Repeating) contains various domains of proteins. Newer techniques using archaeal or bacterial adaptive immune systems, referred to as (T), have been identified.
[0006] Despite some identification and characterization of these systems, endogenous and previously introduced To identify novel effectors and systems for editing heterologous polynucleotides. Furthermore, it remains necessary to demonstrate the activity in eukaryotes, particularly animals and plants. It is being done.
[0007] Novel Cas endonuclease, "Cas-alpha", exemplary protein, and Methods and compositions for their use are described herein. [Overview of the project] [Means for solving the problem]
[0008] This specification discloses a composition of a novel Cas endonuclease and methods of using the same. These novel Cas-alpha class endonucleases target double-stranded DNA and can be induced by a guide polynucleotide to cut it in a PAM-dependent manner, as demonstrated in three different kingdoms of prokaryotes (E. coli) and eukaryotes, namely the plant kingdom, the animal kingdom, and the fungal kingdom.
[0009] In one aspect, a synthetic composition is provided that includes a CRISPR-Cas endonuclease comprising at least one zinc finger-like domain, at least one bridging helix-like domain, and a three-split type RuvC domain (including discontinuous RuvC-I domain, RuvC-II domain, and RuvC-III domain), and optionally includes a heterologous polynucleotide.
[0010] In any aspect, in either the composition or the method, at least one component optimized for expression in eukaryotic cells, particularly plant cells, fungal cells, or animal cells, is provided.
[0011] In one aspect, a synthetic composition is provided that includes a polynucleotide encoding a CRISPR-Cas effector protein derived from an organism selected from the group consisting of: Acidibacillus sulfuroxidans, Alicyclobacillus acidoterrestris, Aneurinibacillus danicus, archaea and a heterologous polynucleotide. rchaea), Bacillus genus, Bacill us cereus, Bacillus megateri um, Bacillus pseudomycoid es, Bacillus sp., Bacillus thuringiensis B acillus thuringiensis, Bacillus toyonensis Bacil lus toyonensis, Bacillus wiedmannii B iedmannii, Bacteroides prevotii ebeius, Bos taurus, Brevibacillus centrosporus Brevibacillus centrosporus, Candidatus Aureabacteria bac terium, Candidatus Aureabacteria bacterium terium, Candidatus Levybacteria bacterium s Levybacteria bacterium, Candidatus Micrarchaeota archae on, Candidatus Micrarchaeota archaeon on, Cellulosilyticum ru minicola, Clostridioides difficile des difficile, Clostridium botulinum m botulinum, Clostridium fallax fallax, Clostridium hiranonis Clostridium nonis, Clostridium humii Clostridium novyi, Clostridium Clostridium paraputrificum Clostridium pasteuria Clostridium perfringens (num), Clostridium perfringens Clostridium sp., Clostridium Clostridium tetani, Clostridium veen Clostridium ventriculi, desulfovibrio phlebosum Lactosivorans (Desulfovibrio fructosivorans), Dorea longicatena, Eubacterium silla Eubacterium siraeum, Flavobacterium thermophilum Flavobacterium thermophilum, Red Junglefowl (Gallus gallus), Hepatitis delta virus (Homo sapiens), human herpesvirus Human betaherpesvirus 5, Hydrogenivirga genus Hydrogenivirga sp.), House mouse (Mus musculus) Parageobacillus thermoglucosidasius ermoglucosidasius), Peptoclostridium sp. stridium sp.), Phascolarctobacterium genus Tobacterium sp.), Prevotella copri (Prevotella copri) pri), Ruminiclostridium fungatei (Hungatei), Ruminococcus albus (Ruminococcus alb) (us), Ruminococcus genus (Ruminococcus sp.), Saccharomyces se Saccharomyces cerevisiae, Simian virus 4 0 (Simian virus 40), potato (Solanum tuberos um), sulfurihydrogenibium azolense ium azorense), Syntrophomonas palmitatica (Syntroph Omonas palmitatica), Tobacco E-virus TCH virus, and Zea mays.
[0012] In one embodiment, the synthetic composition includes eukaryotic cells and heterologous CRISPR-Cas effectors. The aforementioned heterologous CRISPR-Cas effector proteins were less than 800, 790~ 800, less than 790, 780-790, less than 780, 770-780, less than 770, 76 0-770, less than 760, 750-760, less than 750, 740-750, less than 740, 730-740, less than 730, 720-730, less than 720, 710-720, less than 710 Amino acids in the range of 700-710, or less than 700, for example, less than 700, less than 790, 7 Less than 80, less than 750, less than 700, less than 650, less than 600, less than 550, less than 500 A synthetic composition containing amino acids less than 450, less than 400, less than 350, or less than 350 It will be provided.
[0013] In one embodiment, a synthetic composition comprising a CRISPR-Cas endonuclease, The CRISPR-Cas endonuclease, when aligned with SEQ ID NO: 17, For the amino acid position number of column number 17, at least one, at least two, and at least one of the following: A synthetic composition containing 3, at least 4, at least 5, at least 6, or 7 of each. The items offered are: Glycine (G) at position 337, Glycine (G) at position 341, and G at position 430. Rutamic acid (E), leucine (L) at position 432, cysteine (C) at position 487, 490 Cysteine (C) at position 507, cysteine (C) at position 507, and / or cysteine ( C) or histidine (H).
[0014] In one embodiment, a synthetic composition comprising a CRISPR-Cas endonuclease, The CRISPR-Cas endonuclease is composed of one, two, or three of the following motifs. A synthetic composition containing: GxxxG, ExL, and / or one or more Cx is provided. n (C ,H)(where n = one or more amino acids).
[0015] In one embodiment, a synthetic composition comprising a CRISPR-Cas endonuclease, The CRISPR-Cas endonuclease contains one or more zinc finger motifs. A synthetic composition containing the above is provided.
[0016] In one embodiment, sequence numbers 17, 18, 19, 20, 32, 33, 34, 35, 36, 37 ,38,254,255,256,257,258,259,260,261,262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, From 363, 364, 365, 366, 367, 368, 369, 370, and 371 At least 250, 250-300, and at least 300 sequences selected from the group , 300-350 pieces, at least 350 pieces, 350-400 pieces, at least 400 pieces, also It consists of over 400 consecutive amino acids, and at least 50%, 50%-55%, and at least 5 5%, 55%-60%, at least 60%, 60%-65%, at least 65%, 65% ~70%, at least 70%, 70%~75%, at least 75%, 75%~80%, little At least 80%, 80% to 85%, at least 85%, 85% to 90%, at least 90 %, 90%~95%, at least 95%, 95%~96%, at least 96%, 96%~ 97%, at least 97%, 97%~98%, at least 98%, 98%~99%, less CRISPR- A synthetic composition containing Cas effector proteins is provided.
[0017] In one embodiment, sequence numbers 17, 18, 19, 20, 32, 33, 34, 35, 36, 37 ,38,254,255,256,257,258,259,260,261,262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, From 363, 364, 365, 366, 367, 368, 369, 370, and 371 At least 250, 250-500, and at least 500 sequences selected from the group , 500-600 pieces, at least 600 pieces, 600-700 pieces, at least 700 pieces, 7 00-750 pieces, at least 750 pieces, 750-800 pieces, at least 800 pieces, 800 ~850 pieces, at least 850 pieces, 850~900 pieces, at least 900 pieces, 900~9 50, at least 950, 950-1000, at least 1000, or 100 More than 0 amino acids, and at least 50%, 50%~55%, at least 55%, 55%~ 60%, at least 60%, 60% to 65%, at least 65%, 65% to 70%, less At least 70%, 70% to 75%, at least 75%, 75% to 80%, at least 80% , 80%~85%, at least 85%, 85%~90%, at least 90%, 90%~9 5%, at least 95%, 95%-96%, at least 96%, 96%-97%, less 97%, 97%-98%, at least 98%, 98%-99%, at least 99%, CRISPR-Cas effects that share 99% to 100%, or 100%, sequence identity. A synthetic composition containing a polynucleotide encoding a protein is provided.
[0018] In one embodiment, sequences 57, 58, 59, 64, 65, 66, 67, 68, 73, 74 , 75, 76, 77, 102, 103, 104, 105, 177, 178, 179, 18 0, 181, 182, 185, 186, 187, 188, 189, 190, 191, 19 2, 193, 194, 195, 196, 197, 198, 204, 205, 206, 20 7, 208, 209, 210, 211, 212, 213, 214, 215, 216, 21 7, 218, 219, 220, 221, 222, 223, 224, 230, 231, 23 2, 233, 234, 238, 240, 241, 245, 246, 247, 248, 25 At least one, two, three, or four RNA sequences selected from the group consisting of 2 and 253. pieces, 5 pieces, 6 pieces, 7 pieces, 8 pieces, 9 pieces, 10 pieces, 11 pieces, 12 pieces, 13 pieces, 14 pieces, 15 pieces , 16 pieces, 17 pieces, 18 pieces, 19 pieces, 20 pieces, 21 pieces, 22 pieces, 23 pieces, 24 pieces, 25 pieces , 26, 27, 28, 29, 30, or more than 30 consecutive nucleotides, and At least 50%, 50% to 55%, at least 55%, 55% to 60%, at least 60 %, 60%~65%, at least 65%, 65%~70%, at least 70%, 70%~ 75%, at least 75%, 75%~80%, at least 80%, 80%~85%, less At least 85%, 85% to 90%, at least 90%, 90% to 95%, at least 95% 95%~96%, at least 96%, 96%~97%, at least 97%, 97%~9 8%, at least 98%, 98% to 99%, at least 99%, 99% to 100%, or C can hybridize with polynucleotides that share 100% sequence identity. A synthetic set containing polynucleotides encoding RISPR-Cas effector proteins The finished product is provided.
[0019] Any method or composition described herein may further comprise heterogeneous polynucleotides. Heterogeneous polynucleotides can be selected from the following group: promoter, intron Non-coding expression regulatory elements such as enhancers or terminators; Dona - Polynucleotides; at least one modification compared to the polynucleotide sequence in cells Selectively included polynucleotide modification templates; trans genes; guide RNA; Guide DNA; Guide RNA-DNA hybrid; Endonucleases; Nuclear localization SIG Null; and cellular transit peptides.
[0020] In one embodiment, a method for using any of the compositions disclosed herein is provided. In some embodiments, for example, Cas- is used in the genome of a cell or in vitro. A method has been proposed for binding alpha-endonucleases to target sequences of polynucleotides. In some embodiments, Cas-alpha-endonuclease is used as a guide polymer. It forms a complex with nucleotides, for example, guide RNA. In some embodiments, the complex The body creates a nick (single strand) or break (2 strands) in the polynucleotide at or near the target sequence. It recognizes, binds to, and optionally generates (this strand). In some embodiments, nicks or Breaks are repaired by non-homologous end joining (NHEJ). In some embodiments, A break or crack is performed using a polynucleotide modification template or donor DNA molecule. It is repaired by homologous recombination repair (HDR) or homologous recombination (HR).
[0021] The novel Cas endonucleases described herein can be used in any prokaryotic or eukaryotic cell. The target polynucleotide, which contains the appropriate PAM and is directed by the guide polynucleotide In Ochid, or adjacent to it, a double-strand break can be generated. In this example, the cells are plant cells, or animal cells, or fungal cells. In some examples, The plants are selected from the following group: corn, soybeans, cotton, wheat, cabbage. Nora, rapeseed, sorghum, rice, rye, barley, millet, wild oats, sugarcane Birch, grass, switchgrass, alfalfa, sunflower, tobacco, peanuts, potato Moss, tobacco, Arabidopsis, safflower, and tomato.
[0022] Brief description of the drawings and sequence list This disclosure includes the following detailed description which forms part of this application, as well as the drawings and drawings accompanying this specification. You can gain a better understanding from the sequence list. [Brief explanation of the drawing]
[0023] [Figure 1-1] Figure 1: Figures 1A-1D show the complete CRISPR-Cas system, including all the components necessary for acquisition and interference. These include genes encoding all the proteins necessary for spacer acquisition and integration (Cas1 and Cas2) in an operon-like structure adjacent to the CRISPR array, as well as a novel protein containing a DNA cleavage domain, Cas-alpha(α). Furthermore, a gene encoding a protein homologous to Cas4 was also encoded at this locus. Figure 1A shows the locus structures of the Cas-alpha1, Cas-alpha3, and Cas-alpha4 systems. Figure 1B shows the locus structure of the Cas-alpha2 system. Figure 1C shows the locus structure of the Cas-alpha6 system. Figure 1D shows the locus structures of the Cas-alpha5, 7, 8, 9, 10, and 11 systems. [Figure 1-2] (As stated above.) [Figure 1-3] (As stated above.) [Figure 1-4] (As stated above.) [Figure 2] Figure 2 shows a detailed structural analysis of the Cas-alpha protein, highlighting its clear differences from previously described class 2 endonucleases. Conserved residues are shown. Key residues involved in DNA cleavage are marked with an asterisk. The numbers correspond to the Cas-alpha-1 protein. [Figure 3] Figure 3 outlines a method for detecting double-stranded DNA targets and cleavage using cell lysates expressing Cas-alpha endonuclease. [Figure 4-1]Figure 4: Figures 4A-4E show the cleavage of the target polynucleotide by Cas-alpha-1 endonuclease at nucleotide position 21. Figure 4A shows data for Cas-alpha-1 negative controls, Figure 4B shows data for Cas-alpha-1 using the entire (complete) CRISPR locus with the CRISPR array modified to cleave at the target polynucleotide, Figure 4C shows data for the Cas-alpha-1 complete locus plus when expression is enhanced using the T7 promoter, Figure 4D shows data for the Cas-alpha-1 minimal locus when expression is enhanced using the T7 promoter, and Figure 4E shows response data for the remaining CRISPR locus when Cas-alpha-1 is absent but expression is enhanced with the T7 promoter. [Figure 4-2] (As stated above.) [Figure 4-3] (As stated above.) [Figure 4-4] (As stated above.) [Figure 4-5] (As stated above.) [Figure 5-1] Figure 5: Figures 5A-5B show schematic diagrams for determining the orientation of PAM recognition relative to spacer recognition, where the guide RNA is designed to base-pair with the sense or antisense strand of the T2 target. If the guide RNA designed to base-pair with the sense strand results in a restoration of PAM preference and a cleavage signal, the protospacer is on the antisense strand and PAM recognition occurs 3' relative to it (Figure 5A). Conversely, if the guide RNA designed to base-pair with the antisense strand results in PAM preference and a cleavage signal, the protospacer is on the sense strand and PAM recognition occurs 5' relative to it (Figure 5B). [Figure 5-2] (As stated above.) [Figure 6-1]Figure 6: Figures 6A-6E show the cleavage of target polynucleotides at nucleotide position 24 by Cas-alpha4 endonuclease. Figure 6A shows data for Cas-alpha4 negative controls. Figure 6B shows data for Cas-alpha4 plus T2-1 sgRNA. Figure 6C shows data for Cas-alpha4 plus T2-2 sgRNA. Figure 6D shows data for Cas-alpha4 plus T2-1 crRNA / tracrRNA. Figure 6E shows data for Cas-alpha4 plus T2-2 crRNA / tracrRNA. [Figure 6-2] (As stated above.) [Figure 6-3] (As stated above.) [Figure 6-4] (As stated above.) [Figure 6-5] (As stated above.) [Figure 7-1] Figure 7: Figures 7A-7K show representative Cas-alpha loci, endonucleases, proteins, guide RNA components, and other sequences identified from various bacteria and archaea, including: Candidatus Micrarchaeota archaeon (Figures 7A, 7B, 7E), Candidatus Aureabacteria bacterium (Figure 7C), various non-cultured bacteria (Figures 7D, 7F), Parageobacillus thermoglucosidasius (Figure 7G), Acidibacillus sulfuroxidans (Figure 7H), Ruminococcus sp. (Figure 7I), Syntrophomonas palmitatica *Clostridium palmitatica* (Figure 7J) and *Clostridium novyi* (Figure 7K). [Figure 7-2] (As stated above.) [Figure 7-3] (As stated above.) [Figure 7-4] (As stated above.) [Figure 7-5] (As stated above.) [Figure 7-6] (As stated above.) [Figure 7-7] (As stated above.) [Figure 7-8] (As stated above.) [Figure 7-9] (As stated above.) [Figure 7-10] (As stated above.) [Figure 7-11] (As stated above.) [Figure 8-1]Figure 8: Figures 8A-8K show the distinct structural features of representative Cas-alpha proteins, with the protein sequence shown in bold. The non-bold letters below each amino acid residue indicate possible secondary structural features, where C represents an unstructured element or coil, E represents a beta chain, and H represents an alpha helix. The zinc finger domain is shown by a dashed box, and the asterisk indicates key amino acid residues involved in zinc ion binding. The RuvC subdomain of the split RuvC domain is shown by a solid box. The bridge helix is shown by a dashed-dotted box. The coiled coil is depicted by a solid cylinder. The solid plus sign indicates key catalytic residues characteristic of the RuvC domain motif.Figure 8A shows Cas-alpha 1 (SEQ ID NO: 17) derived from Candidatus Micrarchaeota archaeon, Figure 8B shows Cas-alpha 2 derived from Candidatus Micrarchaeota archaeon (SEQ ID NO: 18), Figure 8C shows Cas-alpha 3 (SEQ ID NO: 19) derived from Candidatus Aureabacteria bacterium, Figure 8D shows Cas-alpha 4 (SEQ ID NO: 20) derived from non-cultured bacteria, and Figure 8E shows Cas-alpha 1 (SEQ ID NO: 17) derived from Candidatus Micrarchaeota Figure 8F shows Cas-alpha 5 (SEQ ID NO: 32) derived from archaeon, Figure 8G shows Cas-alpha 7 (SEQ ID NO: 34) derived from Parageobacillus thermoglucosidasius, Figure 8H shows Cas-alpha 8 (SEQ ID NO: 35) derived from Acidibacillus sulfuroxidans, Figure 8I shows Cas-alpha 9 (SEQ ID NO: 36) derived from Ruminococcus sp., Figure 8J shows Cas-alpha 10 (SEQ ID NO: 37) derived from Syntrophomonas palmitatica, which features a unique motif of three zinc finger domains, and Figure 8K shows Clostridium novii This shows Cas-alpha-11 (sequence number 38) derived from novyi. Whole-genome sequencing of organisms containing Cas-alpha-11 has shown that the Cas-alpha locus is the only CRISPR system in that organism. [Figure 8-2] (As stated above.) [Figure 8-3] (As stated above.) [Figure 8-4] (As stated above.) [Figure 8-5] (As stated above.) [Figure 8-6] (As stated above.) [Figure 8-7] (As stated above.) [Figure 8-8] (As stated above.) [Figure 8-9] (As stated above.) [Figure 8-10] (As stated above.) [Figure 8-11] (As stated above.) [Figure 9-1] Figure 9: Figure 9A shows how the Cas-alpha protein subunit interacts with the hybrid double helix of target DNA and guide RNA. Figure 9B is a three-dimensional model of the C-terminal half of Cas-alpha 4, showing the helical hairpin / bridge helix region, RuvC domain, and regions identified as zinc finger motifs common to Cas proteins. [Figure 9-2] (As stated above.) [Figure 10-1] Figure 10: Figures 10A-10D show examples of expression constructs for using Cas-alpha endonuclease in eukaryotic cells. Figure 10A is an example of a human cell Cas-alpha DNA expression construct. Figure 10B is an example of a plant cell Cas-alpha DNA expression construct. Figure 10C is an example of a yeast (Saccharomyces cerevisiae) Cas-alpha DNA expression construct. Figure 10D is an example of a yeast (Saccharomyces cerevisiae) Cas-alpha DNA expression construct. [Figure 10-2] (As stated above.) [Figure 10-3] (As stated above.) [Figure 10-4] (As stated above.) [Figure 11-1]Figure 11: Figures 11A-11D show examples of eukaryotic optimized Cas-alpha guide RNA expression constructs. Figure 11A is an example of a human cell single guide RNA (sgRNA) DNA expression construct. Figure 11B is an example of a plant cell single guide RNA (sgRNA) DNA expression construct. Figure 11C is an example of a yeast (Saccharomyces cerevisiae) single guide RNA (sgRNA) DNA expression construct. Figure 11D is another example of a plant cell single guide RNA (sgRNA) DNA expression construct. [Figure 11-2] (As stated above.) [Figure 11-3] (As stated above.) [Figure 11-4] (As stated above.) [Figure 12] Figure 12 shows an example of a recombinant gene for the recombinant expression and purification of Cas-alpha endonuclease in E. coli. [Figure 13] Figure 13 shows double-strand break repair mutations in plant cells derived from Cas-alpha endonuclease activity. It shows mutations caused by Cas-alpha 4 in Zea mays. The WT reference is SEQ ID NO: 120, mutation 1 is SEQ ID NO: 121, mutation 2 is SEQ ID NO: 122, mutation 3 is SEQ ID NO: 123, and mutation 4 is SEQ ID NO: 124. [Figure 14-1] Figure 14: Figures 14A-14B show double-strand break repair mutations in animal cells derived from Cas-alpha endonuclease activity. Figure 14A shows indel mutations obtained from Cas-alpha 4 RNP electroporation (VEGFA target 2 mutations 1-5, given as SEQ ID NOs. 127-131 compared to WT reference SEQ ID NO. 126; VEGFA target 3 (mutation), given as SEQ ID NO. 133 compared to WT reference SEQ ID NO. 132). Figure 14B shows indel mutations obtained from Cas-alpha 4 and sgRNA DNA expression cassette lipofection, VEGFA target 3 (mutations 1 and 2, given as SEQ ID NOs. 134-135 compared to WT reference SEQ ID NO. 132). [Figure 14-2](As stated above.) [Figure 15-1] Figure 15: Figures 15A-15D show Cas-alpha4 double-stranded DNA targeted cleavage. Figure 15A shows that supercoiled (SC) plasmid DNA containing a guide RNA target (approximately 20 bp) adjacent to the 3' end of the PAM (5'-TTTR-3', where R represents A or Gbp) was completely converted to a linear morphology (FLL), thus indicating the formation of a dsDNA break. Furthermore, the cleavage of the linear DNA produced DNA fragments of the expected size, further confirming Cas-alpha4 mediated dsDNA break formation. Figure 15B shows that Cas-alpha4 requires PAM and guide RNA to cleave the dsDNA target. Figure 15C shows that Cas-alpha4 produces a 5' overhang DNA cleavage site, and the cleavage mainly occurs around a position 20-24 bp from the PAM sequence. Figure 15D shows the trans-acting ssDNase activity of Cas-alpha4 activated by dsDNA only in the presence of guide RNA. [Figure 15-2] (As stated above.) [Figure 15-3] (As stated above.) [Figure 15-4] (As stated above.) [Figure 16-1]Figure 16: Figures 16A-16T show the double-strand DNA target cleavage activity of all Cas-alpha endonucleases except Cas-alpha 5. Figure 16A is the negative control (-IPTG). Figure 16B is the negative control (+IPTG). Figure 16C shows the cleavage of the double-strand DNA target by Cas-alpha 2 (-IPTG) at protospacer position 21. Figure 16D shows the cleavage of the double-strand DNA target by Cas-alpha 2 (+IPTG) at protospacer position 21. Figure 16E shows no cleavage of the double-strand DNA target by Cas-alpha 3 (-IPTG). Figure 16F shows the cleavage of the double-strand DNA target by Cas-alpha 3 (+IPTG) at protospacer position 21. Figure 16G shows no cleavage of the double-strand DNA target by Cas-alpha 5 (-IPTG). Figure 16H shows no cleavage of the double-stranded DNA target by Cas-alpha 5 (-IPTG). Figure 16I shows cleavage of the double-stranded DNA target by Cas-alpha 6 (-IPTG). Figure 16J shows no cleavage of the double-stranded DNA target by Cas-alpha 6 (+IPTG) at protospacer position 24. Figure 16K shows cleavage of the double-stranded DNA target by Cas-alpha 7 (-IPTG) at protospacer position 24. Figure 16L shows cleavage of the double-stranded DNA target by Cas-alpha 7 (+IPTG) at protospacer position 24. Figure 16M shows no cleavage of the double-stranded DNA target by Cas-alpha 8 (-IPTG). Figure 16N shows cleavage of the double-stranded DNA target by Cas-alpha 8 (+IPTG) at protospacer position 24. Figure 16O shows cleavage of the double-stranded DNA target by Cas-alpha 9 (-IPTG) at protospacer position 24. Figure 16P shows the cleavage of a double-stranded DNA target by Cas-alpha-9 (+IPTG) at protospacer position 24. Figure 16Q shows the cleavage of a double-stranded DNA target by Cas-alpha-10 (-IPTG) at protospacer position 24. Figure 16R shows the cleavage of a double-stranded DNA target by Cas-alpha-10 (+IPTG) at protospacer position 24. Figure 16S shows the cleavage of a double-stranded DNA target by Cas-alpha-11 (-IPTG) at protospacer position 24.Figure 16T shows the cleavage of a double-stranded DNA target by Cas-alpha-11 (+IPTG) at the protospacer position 24. [Figure 16-2] (As stated above.) [Figure 16-3] (As stated above.) [Figure 16-4] (As stated above.) [Figure 16-5] (As stated above.) [Figure 16-6] (As stated above.) [Figure 16-7] (As stated above.) [Figure 16-8] (As stated above.) [Figure 16-9] (As stated above.) [Figure 16-10] (As stated above.) [Figure 16-11] (As stated above.) [Figure 16-12] (As stated above.) [Figure 16-13] (As stated above.) [Figure 16-14] (As stated above.) [Figure 16-15] (As stated above.) [Figure 16-16] (As stated above.) [Figure 16-17] (As stated above.) [Figure 16-18] (As stated above.) [Figure 16-19] (As stated above.) [Figure 16-20] (As stated above.) [Figure 17-1]Figure 17: Figure 17A shows one method for evaluating Cas-alpha double-strand DNA targeted breaks in E. coli cells. Figures 17B–17E show double-strand DNA targeted breaks in E. coli. The “no target” experiment provides a baseline for transformation efficiency in the absence of double-strand DNA targeted breaks. The “target” experiment, PAM+T2, was performed both with and without IPTG (0.5 mM) to test targeted breaks under different Cas-alpha endonuclease and guide RNA expression conditions. Figure 17B shows the results for Cas-alpha 2 and Cas-alpha 3. Figure 17C shows the results for Cas-alpha 6 and Cas-alpha 7. Figure 17D shows the results for Cas-alpha 8 and Cas-alpha 9. Figure 17E shows the results for Cas-alpha 10 and Cas-alpha 11. [Figure 17-2] (As stated above.) [Figure 17-3] (As stated above.) [Figure 17-4] (As stated above.) [Figure 17-5] (As stated above.) [Figure 18-1] Figure 18: Figures 18A-18B show double-strand break repair mutations in plant cells from Cas-alpha endonuclease activity for particle gun experiments delivering Cas-alpha 10 DNA expression constructs to Zea-mays immature embryos. Figure 18A shows the recovery of target deletions occurring at or near the nuclease cleavage site of the nptII target site. Figure 18B shows the recovery of target deletions occurring at or near the nuclease cleavage site of the ms26 target site. [Figure 18-2] (As stated above.) [Figure 19-1]Figure 19: Figure 19A shows the experimental design for homologous recombination repair in eukaryotic cells, Saccharomyces cerevisiae. An exogenously supplied DNA repair template (double-stranded) with homology adjacent to the Cas-alpha-10 target site was used to introduce one or two immature stop codons (depending on the DNA repair outcome) into the ade2 gene following a Cas-alpha-10 induced double-strand break (DSB). To avoid targeting of the repair template, it also included a T-to-A change in the PAM region of Cas-alpha-10. Figure 19B shows that when both the repair template and Cas-alpha-10 and sgRNA expression constructs were transformed, and a double-strand break was created by Cas-alpha-endonuclease and repaired with the template (HDR), a red cell phenotype indicating ade2 gene disruption was recovered. Figure 19C shows the sequencing results of the Cas-alpha-10 ade2 gene target site, confirming the introduction of at least one stop codon in three independent red colonies (labeled "1", "2", and "3"). The stop codon was introduced into the antisense frame. Sequence ID 170 is shown as the reference DNA sequence from Saccharomyces cerevisiae, the repair template DNA is Sequence ID 171, repair result 1 in red colony 1 is Sequence ID 172, repair result 2 in red colony 1 is Sequence ID 173, repair result 1 in red colony 2 is Sequence ID 174, repair result 1 in red colony 3 is Sequence ID 175, and repair result 2 in red colony 3 is Sequence ID 176. [Figure 19-2] (As stated above.) [Figure 19-3] (As stated above.) [Figure 20]Figure 20 shows the phylogenetic relationships among several Cas-alpha orthologs. Three supergroups were identified (I, II, and III). Group I included clade 1 (Candidate Archaea and Aureabacteria (where Cas1, Cas2, and Cas4 are usually encoded at the locus)). Group II included clade 2 (Aquificae (genera Sulfurihydrogenibium and Hydrogenivirga) and Deltaproteobacteria (genus Desulfovibrio)), clade 3 (Candidate Archaea (where Cas1, Cas2, and Cas4 are usually encoded at the locus)), clade 4 (Bacteroidetes (genera Prevotella and Bacteroides)), clade 5 (Candidate Levybacterium), and clade 6 (Clostridia (genera Dorea, Ruminococcus, Clostridium, Clostridioides, Peptocolstridium, Cellulosilyticym, Eubacterium)). Group III included clade 7 (Bacilli (genera Bacillus, Acidibacillus, Aneurinibacillus, Brevibacillus, Parageobacillus, Alicyclobacillus)), clade 8 (Negativicutes (genus Phascolarctobacterium)), and clade 9 (Flavobacteriia (genus Flavobacterium)).The diamond symbol indicates Cas-alpha 1-11 endonucleases as described herein. [Figure 21-1] Figure 21: Figure 21A shows a transposase (Tnp)-associated Cas-alpha CRISPR system. In both cases, the Tnp-like protein encodes Cas-alpha endonuclease and upstream of the CRISPR array. Figure 21B shows the target site and the Cas-alpha endonuclease and guide RNA complexed with the Tnp-like protein positioned to incorporate a DNA payload (dashed circle) within or near the Cas-alpha double-stranded DNA target site. [Figure 21-2] (As stated above.) [Modes for carrying out the invention]
[0024] The sequence description and the sequence listing attached to this specification are from 37C.FR§§1.821 and 1 Disclosure of nucleotide and amino acid sequences in a patent application as stipulated in .825 It conforms to the applicable rules. The sequence description is incorporated herein by reference 37 Includes three-letter codes for amino acids as defined in CFR §§ 1.821 and 1.825. .
[0025] Sequence ID 1 is Candidatus miclalcaeota alcaeon. Cas-alpha 1 gene derived from *Micrarchaeota archaeon* This is Cas1 encoded in the PRT sequence.
[0026] Sequence ID 2 is Candidatus miclalcaeota alcaeon. Cas-alpha 2 gene derived from *Micrarchaeota archaeon* This is Cas1 encoded in the PRT sequence.
[0027] Sequence ID 3 is Candidatus aurea bacterium. Cas-alpha-3 derivative of *Aureabacteria bacterium* (us). This is Cas1 encoded in the PRT sequence of the Dendrobium constellation.
[0028] Sequence ID 4 is encoded in the Cas-alpha4 locus PRT sequence derived from non-cultured archaea. This is Cas1.
[0029] Sequence ID 5 is Candidatus miclalcaeota alkaeon. Cas-alpha 1 gene derived from *Micrarchaeota archaeon* This is Cas2 encoded in the PRT sequence.
[0030] Sequence ID 6 is Candidatus miclalcaeota alcaeon. Cas-alpha 2 gene derived from *Micrarchaeota archaeon* This is Cas2 encoded in the PRT sequence.
[0031] Sequence ID 7 is Candidatus aurea bacterium. Cas-alpha-3 derivative of *Aureabacteria bacterium* (us). This is Cas2 encoded in the PRT sequence of the Dendrobium constellation.
[0032] Sequence ID 8 is encoded in the Cas-alpha4 locus PRT sequence derived from non-cultured archaea. This is Cas2.
[0033] Sequence ID 9 is Candidatus miclalcaeota alcaeon. Cas-alpha 1 gene derived from *Micrarchaeota archaeon* This is Cas4 encoded in the PRT sequence.
[0034] Sequence ID 10 is Candidatus miclalcaeota alcaeon. Cas-alpha 2 gene derived from *Micrarchaeota archaeon* (us). This is Cas4 encoded in the PRT sequence of the gestational symmetry.
[0035] Sequence ID 11 is Candidatus aurea bacterium. Cas-alpha-3 derived from *Aureabacteria bacterium* (tus). This is Cas4 encoded by the PRT sequence at the gene locus.
[0036] Sequence ID 12 encodes the Cas-alpha 4 locus PRT sequence derived from uncultured archaea. This is Cas4.
[0037] Sequence ID 13 is Candidatus miclalcaeota alcaeon. Cas-alpha-1 enzyme derived from *Micrarchaeota archaeon* (us). This is the DNA sequence of the donuclease gene.
[0038] Sequence ID 14 is Candidatus micralcaeota alcaeon. Cas-alpha-2 enzyme derived from *Micrarchaeota archaeon* (us). This is the DNA sequence of the donuclease gene.
[0039] Sequence ID 15 is Candidatus aurea bacterium. Cas-alpha-3 derived from *Aureabacteria bacterium* (tus). This is the DNA sequence of the endonuclease gene.
[0040] Sequence ID 16 is a Cas-alpha-4 endonuclease gene D derived from uncultured archaea. It is an NA sequence.
[0041] Sequence ID 17 is Candidatus micralcaeota alcaeon. Cas-alpha-1 enzyme derived from *Micrarchaeota archaeon* (us). This is a donuclease (Cas14b4) PRT sequence.
[0042] Sequence ID 18 is Candidatus miclalcaeota alcaeon. Cas-alpha-2 enzyme derived from *Micrarchaeota archaeon* (us). This is a donuclease PRT sequence.
[0043] Sequence ID 19 is Candidatus aurea bacterium. Cas-alpha-3 e (derived from *Aureabacteria bacterium*) This is an endonuclease PRT sequence.
[0044] Sequence ID 20 is a Cas-alpha-4 endonuclease (Cas) derived from non-cultured archaea. 14a1) This is a PRT sequence.
[0045] Sequence ID 21 is Candidatus micralcaeota alcaeon. Cas-alpha 1 gene derived from *Micrarchaeota archaeon* (us). This is a coccus DNA sequence.
[0046] Sequence ID 22 is Candidatus micralcaeota alcaeon. Cas-alpha 2 gene derived from *Micrarchaeota archaeon* (us). This is a coccus DNA sequence.
[0047] Sequence ID 23 is Candidatus aurea bacterium. Cas-alpha-3 derived from *Aureabacteria bacterium* (tus). This is the DNA sequence of the gene locus.
[0048] Sequence ID 24 is a Cas-alpha4 locus DNA sequence derived from uncultured archaea.
[0049] Sequence ID 25 is Candidatus micralcaeota alkaeon. Cas-alpha-5 e-ene (derived from *Micrarchaeota archaeon*) This is the DNA sequence of the donuclease gene.
[0050] Sequence ID 26 is a Cas-alpha-6 endonuclease gene D derived from non-cultured archaea. It is an NA sequence.
[0051] Sequence ID 27 is Parageobacillus thermoglucosidasius. Cas-alpha-7 enzyme derived from illus thermoglucosidasius This is the DNA sequence of the donuclease gene.
[0052] Sequence ID 28 is Acidobacillus sulfuroxidance. Cas-alpha-8 endonuclease gene D (derived from sulfuroxidans) It is an NA sequence.
[0053] Sequence ID 29 is derived from the genus Ruminococcus sp. This is the DNA sequence of the s-alpha9 endonuclease gene.
[0054] Sequence ID 30 is Syntrophomonas palmitatica. Cas-alpha-10 endonuclease gene D (derived from *S. palmitatica*) It is an NA sequence.
[0055] Sequence ID 31 is Clostridium novyi. i) The derived Cas-alpha-11 endonuclease gene DNA sequence.
[0056] Sequence ID 32 is Candidatus miclalcaeota alcaeon. Cas-alpha-5 e-ene (derived from *Micrarchaeota archaeon*) This is a donuclease PRT sequence.
[0057] Sequence ID 33 is a Cas-alpha-6 endonuclease PRT derived from non-cultured archaea. It is a row.
[0058] Sequence ID 34 is Parageobacillus thermoglucosidasius. Cas-alpha-7 enzyme derived from illus thermoglucosidasius This is a donuclease PRT sequence.
[0059] Sequence ID 35 is Acidobacillus sulfuroxidance. Cas-alpha-8 endonuclease PRT derived from sulfuroxidans It is a row.
[0060] Sequence ID 36 is derived from the genus Ruminococcus sp. This is the s-alpha9 endonuclease PRT sequence.
[0061] Sequence ID 37 is Syntrophomonas palmitatica. Cas-alpha-10 endonuclease PRT derived from *S. palmitatica* It is a row.
[0062] Sequence ID 38 is Clostridium novyi i) It is a Cas-alpha-11 endonuclease PRT sequence derived from [source].
[0063] Sequence ID 39 is Candidatus miclalcaeota alcaeon. Cas-alpha 5 gene derived from *Micrarchaeota archaeon* (us). This is a coccus DNA sequence.
[0064] Sequence ID 40 is a Cas-alpha6 locus DNA sequence derived from uncultured archaea.
[0065] Sequence ID 41 is Parageobacillus thermoglucosidasius. Cas-alpha 7 gene derived from *Illus thermoglucosidasius* This is a coccus DNA sequence.
[0066] Sequence ID 42 is Acidobacillus sulfuroxidance. This is the Cas-alpha 8 locus DNA sequence derived from sulfuroxidans.
[0067] Sequence ID 43 is derived from the genus Ruminococcus sp. This is the s-alpha9 locus DNA sequence.
[0068] Sequence ID 44 is Syntrophomonas palmitatica. This is the Cas-alpha-10 locus DNA sequence derived from *S. palmitatica*.
[0069] Sequence ID 45 is Clostridium novyi i) The derived Cas-alpha-11 locus DNA sequence.
[0070] Sequence ID 46 is Candidatus miclalcaeota alcaeon. Cas-alpha-1 repeat (derived from *Micrarchaeota archaeon*) This is the consensus DNA sequence.
[0071] Sequence ID 47 is Candidatus micralcaeota alcaeon. Cas-alpha 2-replica (derived from *Micrarchaeota archaeon*) This is the consensus DNA sequence.
[0072] Sequence ID 48 is Candidatus aurea bacterium. Cas-alpha-3 derived from *Aureabacteria bacterium* (tus). This is a repeat consensus DNA sequence.
[0073] Sequence ID 49 is Cas-alpha-4 repeat consensus DNA derived from uncultured archaea. It is an array.
[0074] Sequence ID 50 is Candidatus miclalcaeota alcaeon. Cas-alpha 5-repeat (derived from *Micrarchaeota archaeon*) This is the consensus DNA sequence.
[0075] Sequence ID 51 is Cas-alpha-6 repeat consensus DNA derived from uncultured archaea. It is an array.
[0076] Sequence ID 52 is Parageobacillus thermoglucosidasius. Cas-alpha 7-repeat (derived from Illus thermoglucosidasius) This is the consensus DNA sequence.
[0077] Sequence ID 53 is Acidobacillus sulfuroxidance. Cas-alpha-8 repeat consensus DNA derived from sulfuroxidans It is an array.
[0078] Sequence ID 54 is derived from the genus Ruminococcus (Ruminococcus sp.) Ca This is the s-alpha9 repeat consensus DNA sequence.
[0079] Sequence ID 55 is Syntrophomonas palmitatica. Cas-alpha-10 repeat consensus DNA derived from *S. palmitatica* It is an array.
[0080] Sequence ID 56 is Clostridium novyi i) This is the Cas-alpha-11 repeat consensus DNA sequence derived from [the source].
[0081] Sequence ID 57 is a Cas-alpha-1 crRNA derived from an artificial product (where N is arbitrary). This is an RNA sequence (representing nucleotides).
[0082] Sequence ID 58 is a Cas-alpha-2 crRNA derived from an artificial product (where N is arbitrary). This is an RNA sequence (representing nucleotides).
[0083] Sequence ID 59 is a Cas-alpha-4 crRNA derived from an artificial product (where N is arbitrary). This is an RNA sequence (representing nucleotides).
[0084] Sequence ID 60 is Candidatus micralcaeota alcaeon. Cas-alpha-1 t (derived from *Micrarchaeota archaeon*) This is the racrRNA version 1 RNA sequence.
[0085] Sequence ID 61 is Candidatus miclalcaeota alcaeon. Cas-alpha-1 t (derived from *Micrarchaeota archaeon*) This is the racrRNA version 2 RNA sequence.
[0086] Sequence ID 62 is Candidatus micralcaeota alcaeon. Cas-alpha-1 t (derived from *Micrarchaeota archaeon*) This is the racrRNA version 3 RNA sequence.
[0087] Sequence ID 63 is Candidatus micralcaeota alcaeon. Cas-alpha-1 t (derived from *Micrarchaeota archaeon*) This is the racrRNA version 4 RNA sequence.
[0088] Sequence ID 64 is Candidatus micralcaeota alcaeon. Cas-alpha-2 t (derived from *Micrarchaeota archaeon*) This is the racrRNA version 1 RNA sequence.
[0089] Sequence ID 65 is Candidatus micralcaeota alkaeon. Cas-alpha-2 t (derived from *Micrarchaeota archaeon*) This is the racrRNA version 2 RNA sequence.
[0090] Sequence ID 66 is Candidatus miclalcaeota alcaeon. Cas-alpha-2 t (derived from *Micrarchaeota archaeon*) This is the racrRNA version 3 RNA sequence.
[0091] Sequence ID 67 is Candidatus micralcaeota alcaeon. Cas-alpha-2 t (derived from *Micrarchaeota archaeon*) This is the racrRNA version 4 RNA sequence.
[0092] Sequence ID 68 is a Cas-alpha-4 tracrRNA version derived from non-cultured archaea. This is an RNA sequence.
[0093] Sequence ID 69 is a Cas-alpha-1 sgRNA version 1 RN derived from an artificial product. It is an A array.
[0094] Sequence ID 70 is a Cas-alpha-1 sgRNA version 2 RN derived from an artificial product. It is an A array.
[0095] Sequence ID 71 is a Cas-alpha-1 sgRNA version 3 RN derived from an artificial product. It is an A array.
[0096] Sequence ID 72 is a Cas-alpha-1 sgRNA version 4 RN derived from an artificial product. It is an A array.
[0097] Sequence ID 73 is a Cas-alpha-2 sgRNA version 1 RN derived from an artificial product. It is an A array.
[0098] Sequence ID 74 is a Cas-alpha-2 sgRNA version 2 RN derived from an artificial product. It is an A array.
[0099] Sequence ID 75 is a Cas-alpha-2 sgRNA version 3 RN derived from an artificial product. It is an A array.
[0100] Sequence ID 76 is a Cas-alpha-2 sgRNA version 4 RN derived from an artificial product. It is an A array.
[0101] Sequence ID 77 is a Cas-alpha4 sgRNA version 1 RN derived from an artificial product. It is an A array.
[0102] Sequence ID 78 is a T2 spacer DNA sequence derived from an artificial product.
[0103] Sequence ID 79 is a synthetically derived, engineered to target the T2 DNA sequence. This is the entire Cas-alpha1 gene locus.
[0104] Sequence ID 80 is a synthetically derived, engineered to target the T2 DNA sequence. This is the minor Cas-alpha1 gene locus.
[0105] Sequence ID 81 is a 10× histidine-tagged PRT sequence derived from an artificial product.
[0106] Sequence ID 82 is a 6× histidine-tagged PRT sequence derived from an artificial product.
[0107] Sequence ID 83 is a maltose-binding protein tag PRT sequence derived from an artificial product.
[0108] Sequence ID 84 is derived from the Tobacco etch virus. Tobacco etch virus (PRT) puncture site It is a row.
[0109] Sequence ID 85 is an A1 oligonucleotide DNA sequence derived from an artificial product.
[0110] Sequence number 86 is an A2 oligonucleotide DNA sequence derived from an artificial product.
[0111] Sequence number 87 is an R0 oligonucleotide DNA sequence derived from an artificial product.
[0112] Sequence number 88 is a C0 oligonucleotide DNA sequence derived from an artificial product.
[0113] Sequence number 89 is an F1 oligonucleotide DNA sequence derived from an artificial product.
[0114] Sequence number 90 is an R1 oligonucleotide DNA sequence derived from an artificial product.
[0115] Sequence number 91 is the bridge amplification part of the F1 oligonucleotide DNA sequence derived from an artificial product portion.
[0116] Sequence number 92 is the bridge amplification part of the R1 oligonucleotide DNA sequence derived from an artificial product portion.
[0117] Sequence number 93 is an F2 oligonucleotide DNA sequence derived from an artificial product.
[0118] Sequence number 94 is an R2 oligonucleotide DNA sequence derived from an artificial product.
[0119] Sequence number 95 is a C1 oligonucleotide DNA sequence derived from an artificial product.
[0120] Sequence number 96 is a sequence obtained from cleavage at position 21 of the target DNA sequence derived from an artificial product and adapter ligation ation.
[0121] Sequence number 97 is the adapter part of the DNA sequence of sequence number 96 derived from an artificial product.
[0122] Sequence number 98 is the target portion of the DNA sequence of sequence number 96 derived from an artificial product.
[0123] Sequence number 99 is the 5' sequence of the PAM DNA sequence derived from an artificial product.
[0124] Sequence number 100 is a fixed double-stranded DNA target DNA sequence derived from an artificial product.
[0125] Sequence number 101 is a T2 target sequence DNA sequence derived from an artificial product.
[0126] Sequence number 102 is a Cas-alpha4 T2-1 sgRNA RN A sequence.
[0127] Sequence number 103 is a Cas-alpha4 T2-2 sgRNA RN A sequence.
[0128] Sequence number 104 is a Cas-alpha4 T2-1 crRNA RN A sequence.
[0129] Sequence number 105 is a Cas-alpha4 T2-2 crRNA RN A sequence.
[0130] Sequence number 106 is the ST- LS1 intron 2 DNA sequence derived from Solanum tuberosum (potato).
[0131] Sequence number 107 is from Simian virus 40 SV40 NLS PRT sequence.
[0132] Sequence number 108 is the Nuc NLS PRT sequence derived from Mus musculus (mouse).
[0133] Array number 109 is the maize UBI pro from Zea mays (maize) and is a
[0134] motor DNA sequence. Array number 110 is the beta-actin promoter DNA sequence
[0135] from Gallus gallus (chicken). Array number 111 is the CMV enhancer DNA
[0136] sequence from human beta-herpesvirus 5. Array number 112 is the maize UBI 5
[0137] prime untranslated region DNA sequence from Zea mays (maize). Array number 113 is the maize UBI int
[0138] ron 1 DNA sequence from Zea mays (maize).
[0139] Array number 114 is a hybrid intron DNA sequence from an artificial product.
[0140] Array number 115 is the maize U6 polymerase III promoter DNA sequence from Zea mays (maize).
[0141] Array number 116 is the human U6 polymerase III promoter DNA sequence
[0142] from Homo sapiens (human). Array number 117 is a Strep II tag PRT sequence
[0143] from an artificial product. Array number 118 is the bGH poly(A) terminator DNA sequence
[0143] Sequence ID 119 is derived from potato (Solanum tuberosum). This is the protease inhibitor II (Pin II) terminator DNA sequence.
[0144] Sequence ID 120 is derived from Zea mays. Mays)Wt reference (Liguleless target 2 and 3) DNA sequences.
[0145] Sequence ID 121 is a variant 1 (Ligulel) derived from Zea mays. These are ess target 2 and 3-DNA (Exp.) DNA sequences.
[0146] Sequence ID 122 is a variant 2 (Ligulel) derived from Zea mays. These are ess target 2 and 3-DNA (Exp.) DNA sequences.
[0147] Sequence ID 123 is a variant 3 (Ligulel) derived from Zea mays. These are ess target 2 and 3-DNA (Exp.) DNA sequences.
[0148] Sequence ID 124 is a variant 4 (Ligulel) derived from Zea mays. These are ess target 2 and 3-DNA (Exp.) DNA sequences.
[0149] Sequence ID 125 is a variant 5 (Ligulel) derived from Zea mays. These are ess target 2 and 3-DNA (Exp.) DNA sequences.
[0150] Sequence ID 126 is HEK29, derived from Homo sapiens. 3. This is the Wt reference (VEGFA target 2) DNA sequence.
[0151] Sequence ID 127 is a variant of Homo sapiens, specifically variant 1(V). This is an EGFA-targeted 2-RNP (DNA) sequence.
[0152] Sequence ID 128 is a variant of Homo sapiens, specifically variant 2(V). This is an EGFA-targeted 2-RNP (DNA) sequence.
[0153] Sequence ID 129 is a variant of Homo sapiens, specifically mutation 3(V). This is an EGFA-targeted 2-RNP (DNA) sequence.
[0154] Sequence ID 130 is a variant of Homo sapiens, specifically mutation 4(V). This is an EGFA-targeted 2-RNP (DNA) sequence.
[0155] Sequence ID 131 is a variant of Homo sapiens, specifically mutation 5(V). This is an EGFA-targeted 2-RNP (DNA) sequence.
[0156] Sequence ID 132 is HEK29, derived from Homo sapiens. 3 is the Wt reference (VEGFA target 3) DNA sequence.
[0157] Sequence ID 133 is a variant of Homo sapiens (V) 1. This is an EGFA-targeted 3-RNP DNA sequence.
[0158] Sequence ID 134 is a variant of Homo sapiens, specifically variant 1(V). This is an EGFA-targeted 3-DNA (Exp)DNA sequence.
[0159] Sequence ID 135 is a variant of Homo sapiens, specifically variant 2(V). This is an EGFA-targeted 3-DNA (Exp)DNA sequence.
[0160] Sequence ID 136 is Saccharomyces cerevisiae. This is the ROX3 promoter DNA sequence derived from revisiae.
[0161] Sequence ID 137 is Saccharomyces cerevisiae. This is a GAL promoter DNA sequence derived from revisiae.
[0162] Sequence ID 138 is an artificially derived HH ribozyme (where N is the 6-nucleus of the ribozyme). This is a DNA sequence that shows a nucleotide complementary to rheotide 3'.
[0163] Sequence ID 139 is Hepatitis delta virus. This is the HDV ribozyme DNA sequence derived from s).
[0164] Sequence ID 140 is Saccharomyces cerevisiae. This is the SNR52 promoter DNA sequence derived from revisiae.
[0165] Sequence ID 141 is Saccharomyces cerevisiae. This is the SUP4 terminator DNA sequence derived from revisiae.
[0166] Sequence ID 142 is a DNA sequence derived from an artificial product, as shown at the top of Figure 15C.
[0167] Sequence ID 143 is a DNA sequence derived from an artificial product, as shown at the bottom of Figure 15C.
[0168] Sequence ID 144 is a reference DN from Zea Mays, shown in Figure 18A. It is an A array.
[0169] Sequence ID 145 is a mutant DNA sequence derived from Zea mays. ru.
[0170] Sequence ID 146 is a variant DNA sequence derived from Zea mays. ru.
[0171] Sequence ID 147 is a variant DNA sequence derived from Zea mays. ru.
[0172] Sequence ID 148 is a variant DNA sequence derived from Zea mays. ru.
[0173] Sequence ID 149 is a variant DNA sequence derived from Zea mays. ru.
[0174] Sequence ID 150 is a variant 6 DNA sequence derived from Zea mays. ru.
[0175] Sequence ID 151 is a variant DNA sequence derived from Zea mays. ru.
[0176] Sequence ID 152 is a variant DNA sequence derived from Zea mays. ru.
[0177] Sequence ID 153 is a variant DNA sequence derived from Zea mays. ru.
[0178] Sequence ID 154 is a mutated DNA sequence derived from Zea mays. be.
[0179] Sequence ID 155 is a mutant 11 DNA sequence derived from Zea mays. be.
[0180] Sequence ID 156 is a mutated 12 DNA sequence derived from Zea mays. be.
[0181] Sequence ID 157 is a mutated 13 DNA sequence derived from Zea mays. be.
[0182] Sequence ID 158 is a mutated 14 DNA sequence derived from Zea mays. be.
[0183] Sequence ID 159 is a variant DNA sequence derived from Zea mays. be.
[0184] Sequence ID 160 is a mutated 16 DNA sequence derived from Zea mays. be.
[0185] Sequence ID 161 is a variant 17 DNA sequence derived from Zea mays. be.
[0186] Sequence ID 162 is a mutant 18 DNA sequence derived from Zea mays. be.
[0187] Sequence ID 163 is a variant 19 DNA sequence derived from Zea mays. be.
[0188] Sequence ID 164 is a reference DN shown in Figure 18B, derived from Zea Mays. It is an A array.
[0189] Sequence ID 165 is a mutant DNA sequence derived from Zea mays. ru.
[0190] Sequence ID 166 is a variant DNA sequence derived from Zea mays. ru.
[0191] Sequence ID 167 is a variant DNA sequence derived from Zea mays. ru.
[0192] Sequence ID 168 is a variant DNA sequence derived from Zea mays. ru.
[0193] Sequence ID 169 is a variant DNA sequence derived from Zea mays. ru.
[0194] Sequence ID 170 is Saccharomyces cerevisiae. This is the reference DNA sequence shown in Figure 19C, derived from revisiae.
[0195] Sequence ID 171 is a repair template DNA sequence derived from an artificial product.
[0196] Sequence ID 172 is Saccharomyces cerevisiae. This is a DNA sequence resulting from repair (revisiae).
[0197] Sequence ID 173 is Saccharomyces cerevisiae. This is a DNA sequence resulting from repair (revisiae).
[0198] Sequence ID 174 is Saccharomyces cerevisiae. This is a DNA sequence resulting from repair (revisiae).
[0199] Sequence ID 175 is Saccharomyces cerevisiae. This is a DNA sequence resulting from repair (revisiae).
[0200] Sequence ID 176 is Saccharomyces cerevisiae. This is a DNA sequence resulting from repair (revisiae).
[0201] Sequence ID 177 is a Cas-alpha-3 crRNA derived from an artificial product (where N is arbitrary). This is an RNA sequence (showing the nucleotides of a particular type).
[0202] Sequence ID 178 is a Cas-alpha 5 crRNA derived from an artificial product (where N is arbitrary). This is an RNA sequence (showing the nucleotides of a particular type).
[0203] Sequence ID 179 is a Cas-alpha-6 crRNA derived from an artificial product (where N is arbitrary). This is an RNA sequence (showing the nucleotides of a particular type).
[0204] Sequence ID 180 is a Cas-alpha-7 crRNA derived from an artificial product (where N is arbitrary). This is an RNA sequence (showing the nucleotides of a particular type).
[0205] Sequence ID 181 is a Cas-alpha-8 crRNA derived from an artificial product (where N is arbitrary). This is an RNA sequence (showing the nucleotides of a particular type).
[0206] Sequence ID 182 is a Cas-alpha-9 crRNA derived from an artificial product (where N is arbitrary). This is an RNA sequence (showing the nucleotides of a particular type).
[0207] Sequence ID 183 is a Cas-alpha-10 crRNA derived from an artificial product (where N is This is an RNA sequence (representing any nucleotide).
[0208] Sequence ID 184 is a Cas-alpha-11 crRNA derived from an artificial product (where N is This is an RNA sequence (representing any nucleotide).
[0209] Sequence ID 185 is Candidatus miclalcaeota alcaeon. Cas-alpha-2 derived from tus Micrarchaeota archaeon This is the tracrRNA version 5 RNA sequence.
[0210] Sequence ID 186 is Candidatus miclalcaeota alcaeon. Cas-alpha-2 derived from tus Micrarchaeota archaeon This is the tracrRNA version 6 RNA sequence.
[0211] Sequence ID 187 is Candidatus micralcaeota alcaeon. Cas-alpha-2 derived from tus Micrarchaeota archaeon This is the tracrRNA version 7 RNA sequence.
[0212] Sequence ID 188 is a Cas-alpha-6 tracrRNA variant derived from uncultured archaea. This is the 1 RNA sequence.
[0213] Sequence ID 189 is a Cas-alpha-6 tracrRNA variant derived from non-cultured archaea. This is the 2 RNA sequence.
[0214] Sequence ID 190 is a Cas-alpha-6 tracrRNA variant derived from non-cultured archaea. This is the 3 RNA sequence.
[0215] Sequence ID 191 is a Cas-alpha-6 tracrRNA variant derived from non-cultured archaea. This is a 4 RNA sequence.
[0216] Sequence ID 192 is Parageobacillus thermoglucosidasius. Cas-alpha-7 derived from *Cillus thermoglucosidasius* This is the tracrRNA version 1 RNA sequence.
[0217] Sequence ID 193 is Parageobacillus thermoglucosidasius. Cas-alpha-7 derived from *Cillus thermoglucosidasius* This is the tracrRNA version 2 RNA sequence.
[0218] Sequence ID 194 is Acidobacillus sulfuroxidance. Cas-alpha-8 tracrRNA variant derived from sulfuroxidans This is the 1 RNA sequence.
[0219] Sequence ID 195 is Acidobacillus sulfuroxidance. Cas-alpha-8 tracrRNA variant derived from sulfuroxidans This is the 2 RNA sequence.
[0220] Sequence ID 196 is Acidobacillus sulfuroxidance. Cas-alpha-8 tracrRNA variant derived from sulfuroxidans This is the 3 RNA sequence.
[0221] Sequence ID 197 is derived from the genus Ruminococcus sp. This is the s-alpha9 tracrRNA version 1 RNA sequence.
[0222] Sequence ID 198 is derived from the genus Ruminococcus sp. This is the s-alpha9 tracrRNA version 2 RNA sequence.
[0223] Sequence ID 199 is Syntrophomonas palmitatica. Cas-alpha-10 tracrRNA variant derived from *As palmitatica* This is the 1 RNA sequence.
[0224] Sequence ID 200 is Syntrophomonas palmitatica. Cas-alpha-10 tracrRNA variant derived from *As palmitatica* This is the 2 RNA sequence.
[0225] Sequence ID 201 is Syntrophomonas palmitatica. Cas-alpha-10 tracrRNA variant derived from *As palmitatica* This is the 3 RNA sequence.
[0226] Sequence ID 202 is Syntrophomonas palmitatica. Cas-alpha-10 tracrRNA variant derived from *As palmitatica* This is a 4 RNA sequence.
[0227] Sequence ID 203 is Syntrophomonas palmitatica. Cas-alpha-10 tracrRNA variant derived from *As palmitatica* This is the 5 RNA sequence.
[0228] Sequence ID 204 is Clostridium novii. This is the Cas-alpha-11 tracrRNA version 1 RNA sequence derived from yi). .
[0229] Sequence ID 205 is Clostridium novii. This is the Cas-alpha-11 tracrRNA version 2 RNA sequence derived from yi). .
[0230] Sequence ID 206 is Clostridium novii. This is the Cas-alpha-11 tracrRNA version 3 RNA sequence derived from yi). .
[0231] Sequence ID 207 is Clostridium novii. This is the Cas-alpha-11 tracrRNA version 4 RNA sequence derived from yi. .
[0232] Sequence ID 208 is a synthetic Cas-alpha-2 sgRNA version 5 R. It is an NA sequence.
[0233] Sequence ID 209 is a synthetic Cas-alpha-2 sgRNA version 6 R. It is an NA sequence.
[0234] Sequence ID 210 is a synthetic Cas-alpha-2 sgRNA version 7 R. It is an NA sequence.
[0235] Sequence ID 211 is a synthetic Cas-alpha-6 sgRNA version 1 R. It is an NA sequence.
[0236] Sequence ID 212 is a synthetic Cas-alpha-6 sgRNA version 2 R. It is an NA sequence.
[0237] Sequence ID 213 is a synthetic Cas-alpha-6 sgRNA version 3 R. It is an NA sequence.
[0238] Sequence ID 214 is a synthetic Cas-alpha-6 sgRNA version 4 R. It is an NA sequence.
[0239] Sequence ID 215 is a synthetic Cas-alpha7 sgRNA version 1 R. It is an NA sequence.
[0240] Sequence ID 216 is a synthetic Cas-alpha7 sgRNA version 2 R. It is an NA sequence.
[0241] Sequence ID 217 is a synthetic Cas-alpha7 sgRNA version 3 R. It is an NA sequence.
[0242] Sequence ID 218 is a synthetic Cas-alpha-8 sgRNA version 1 R. It is an NA sequence.
[0243] Sequence ID 219 is a synthetic Cas-alpha-8 sgRNA version 2 R. It is an NA sequence.
[0244] Sequence ID 220 is a synthetic Cas-alpha-8 sgRNA version 3 R. It is an NA sequence.
[0245] Sequence ID 221 is a synthetic Cas-alpha-8 sgRNA version 4 R. It is an NA sequence.
[0246] Sequence ID 222 is a synthetic Cas-alpha-9 sgRNA version 1 R. It is an NA sequence.
[0247] Sequence ID 223 is a synthetic Cas-alpha-9 sgRNA version 2 R. It is an NA sequence.
[0248] Sequence ID 224 is a synthetic Cas-alpha-9 sgRNA version 3 R. It is an NA sequence.
[0249] Sequence ID 225 is a Cas-alpha-10 sgRNA version 1 derived from an artificial product. This is an RNA sequence.
[0250] Sequence ID 226 is a Cas-alpha-10 sgRNA version 2 derived from an artificial product. This is an RNA sequence.
[0251] Sequence ID 227 is a Cas-alpha-10 sgRNA version 3 derived from an artificial product. This is an RNA sequence.
[0252] Sequence ID 228 is a Cas-alpha-10 sgRNA version 4 derived from an artificial product. This is an RNA sequence.
[0253] Sequence ID 229 is a Cas-alpha-10 sgRNA version 5 derived from an artificial product. This is an RNA sequence.
[0254] Sequence ID 230 is a Cas-alpha-11 sgRNA version 1 derived from an artificial product. This is an RNA sequence.
[0255] Sequence ID 231 is a Cas-alpha-11 sgRNA version 2 derived from an artificial product. This is an RNA sequence.
[0256] Sequence ID 232 is a Cas-alpha-11 sgRNA version 3 derived from an artificial product. This is an RNA sequence.
[0257] Sequence ID 233 is a Cas-alpha-11 sgRNA version 4 derived from an artificial product. This is an RNA sequence.
[0258] Sequence ID 234 is a Cas-alpha-11 sgRNA version 5 derived from an artificial product. This is an RNA sequence.
[0259] Sequence ID 235 is derived from the artificial product Cas-alpha-4 zea mize (Zea may s) This is a codon-optimized gene DNA sequence.
[0260] Sequence ID 236 is derived from the artificial product Cas-alpha-10 zea mize (Zea ma ys) This is a codon-optimized gene DNA sequence.
[0261] Sequence ID 237 is derived from the artificial product Cas-alpha-10 Saccharomyces cerevisiae. (Saccharomyces cerevisiae) This is a codon-optimized gene DNA sequence.
[0262] Sequence ID 238 is a Cas-alpha-4 sgRNA backbone derived from an artificial product. It is an NA sequence.
[0263] Sequence ID 239 is a Cas-alpha-10 sgRNA backbone derived from an artificial product. This is an RNA sequence.
[0264] Sequence ID 240 is derived from the artificial product Cas-alpha-4 Liguleless2 s The gRNA target sequence is an RNA sequence.
[0265] Sequence ID 241 is derived from the artificial product Cas-alpha-4 Liguleless3 s The gRNA target sequence is an RNA sequence.
[0266] Sequence ID 242 is a Cas-alpha-10 nptII sgRNA variant derived from an artificial product. The sequence is an RNA sequence.
[0267] Sequence ID 243 targets the artificially derived Cas-alpha-10 ms26 sgRNA. The sequence is an RNA sequence.
[0268] Sequence ID 244 targets a Cas-alpha-10 ade2 sgRNA derived from an artificial product. The sequence is an RNA sequence.
[0269] Sequence ID 245 is a Cas-alpha-4 VEGFA2 sgRNA variant derived from an artificial product. The sequence is an RNA sequence.
[0270] Sequence ID 246 is a Cas-alpha-4 VEGFA3 sgRNA variant derived from an artificial product. The sequence is an RNA sequence.
[0271] Sequence ID 247 is a Cas-alpha-4 sgRNA targeting agent derived from an artificial product. This is the Liguleless2 RNA sequence.
[0272] Sequence ID 248 is a Cas-alpha-4 sgRNA targeting agent derived from an artificial product. This is the Liguleless3 RNA sequence.
[0273] Sequence ID 249 is a Cas-alpha-10 sgRNA target derived from an artificial product. This is the nptII RNA sequence.
[0274] Sequence ID 250 is a Cas-alpha-10 sgRNA target derived from an artificial product. This is the ms26 RNA sequence.
[0275] Sequence ID 251 is a Cas-alpha-10 sgRNA target derived from an artificial product. This is the ngade2 RNA sequence.
[0276] Sequence ID 252 is a Cas-alpha-4 sgRNA targeting agent derived from an artificial product. This is the VEGFA2 RNA sequence.
[0277] Sequence ID 253 is a Cas-alpha-4 sgRNA targeting agent derived from an artificial product. This is the VEGFA3 RNA sequence.
[0278] Sequence ID 254 is Clostridioides difficile. Cas-alpha-12 endonuclease PRT derived from (des difficile) It is a row.
[0279] Sequence ID 255 is Clostridium paraptriphycum. Cas-alpha-13 endonuclease PR derived from Paraputrificum This is a T-array.
[0280] Sequence ID 256 is Clostridium novii. This is a Cas-alpha-14 endonuclease PRT sequence derived from (yi).
[0281] Sequence ID 257 is Ruminococcus albus. This is a Cas-alpha-15 endonuclease PRT sequence derived from (s).
[0282] Sequence ID 258 is Clostridium hiranonis. This is a Cas-alpha-16 endonuclease PRT sequence derived from anonis.
[0283] Sequence ID 259 is Clostridium ifmii. i) This is a Cas-alpha-17 endonuclease PRT sequence derived from [source].
[0284] Sequence ID 260 is Cellulosilyticum luminicola. Cas-alpha-18 endonuclease PRT derived from *Cum ruminicola* It is an array.
[0285] Sequence ID 261 is Eubacterium siraeum. This is a Cas-alpha-19 endonuclease PRT sequence derived from aeum.
[0286] Sequence ID 262 is Clostridium botulinum. This is a Cas-alpha 20 endonuclease PRT sequence derived from *Ulinum*.
[0287] Sequence ID 263 is Clostridium botulinum. This is a Cas-alpha-21 endonuclease PRT sequence derived from *Ulinum*.
[0288] Sequence ID 264 is Luminiclostridium fungatei. Cas-alpha-22 endonuclease PRT derived from *Idium hungatei* It is an array.
[0289] Sequence ID No. 265 is Desulfovibio fructosiborans. Cas-alpha-23 endonuclear derived from Rio fructosivorans This is a ZE-PRT sequence.
[0290] Sequence ID 266 is Bacillus toyonensi This is a Cas-alpha 24 endonuclease PRT sequence derived from s).
[0291] Sequence ID 267 is Clostridium paraptriphycum. Cas-alpha-25 endonuclease PR derived from Paraputrificum This is a T-array.
[0292] Sequence ID 268 is Clostridium ventrilix. It is a Cas-alpha26 endonuclease PRT sequence derived from (entriculi). .
[0293] Sequence ID 269 is derived from the genus Ruminococcus sp. This is the as-alpha27 endonuclease PRT sequence.
[0294] Sequence ID 270 is derived from the genus Ruminococcus sp. This is the as-alpha28 endonuclease PRT sequence.
[0295] Sequence ID 271 is a species of the genus Peptoclostridium. This is a Cas-alpha 29 endonuclease PRT sequence derived from sp. (a specific gene).
[0296] Sequence ID 272 is a Cas-alpha derived from the genus Bacillus sp. This is a 30-endonuclease PRT sequence.
[0297] Sequence ID 273 is Clostridioides difficile. Cas-alpha-31 endonuclease PRT derived from (des difficile) It is a row.
[0298] Sequence ID 274 is Clostridioides difficile. Cas-alpha-32 endonuclease PRT derived from (des difficile) It is a row.
[0299] Sequence ID No. 275 is a Cas-alpha-33 endonuclease PR derived from non-cultured archaea. This is a T-array.
[0300] Sequence ID No. 276 is a Cas-alpha-34 endonuclease PR derived from non-cultured archaea. This is a T-array.
[0301] Sequence ID No. 277 is a Cas-alpha-35 endonuclease PR derived from non-cultured archaea. This is a T-array.
[0302] Sequence ID No. 278 is a Cas-alpha-36 endonuclease PR derived from non-cultured archaea. This is a T-array.
[0303] Sequence ID No. 279 is a Cas-alpha-37 endonuclease PR derived from non-cultured archaea. This is a T-array.
[0304] Sequence ID No. 280 is a Cas-alpha-38 endonuclease PR derived from non-cultured archaea. This is a T-array.
[0305] Sequence ID No. 281 is a Cas-alpha-39 endonuclease PR derived from non-cultured archaea. This is a T-array.
[0306] Sequence ID 282 is Cas-alpha-40 endonuclease P derived from uncultured archaea. It is an RT sequence.
[0307] Sequence ID No. 283 is a Cas-alpha-41 endonuclease PR derived from non-cultured archaea. This is a T-array.
[0308] Sequence ID 284 is Clostridioides difficile. Cas-alpha-42 endonuclease PRT derived from (des difficile) It is a row.
[0309] Sequence ID No. 285 is Desulfovibio fructosiborans. Cas-alpha 43 endonuclear derived from rio fructosivorans This is a ZE-PRT sequence.
[0310] Sequence ID 286 is Clostridium botulinum. This is a Cas-alpha 44 endonuclease PRT sequence derived from *Ulinum*.
[0311] Sequence ID 287 is Clostridioides difficile. Cas-alpha-45 endonuclease PRT derived from (des difficile) It is a row.
[0312] Sequence ID 288 is Clostridioides difficile. Cas-alpha-46 endonuclease PRT derived from (des difficile) It is a row.
[0313] Sequence ID 289 is Clostridioides difficile. Cas-alpha-47 endonuclease PRT derived from (des difficile) It is a row.
[0314] Sequence ID 290 is Clostridioides difficile. Cas-alpha-48 endonuclease PRT derived from (des difficile) It is a row.
[0315] Sequence ID 291 is Clostridioides difficile. Cas-alpha-49 endonuclease PRT derived from (des difficile) It is a row.
[0316] Sequence ID 292 is Clostridioides difficile. Cas-alpha 50 endonuclease PRT derived from (des difficile) It is a row.
[0317] Sequence ID 293 is Clostridioides difficile. Cas-alpha 51 endonuclease PRT derived from (des difficile) It is a row.
[0318] Sequence ID 294 is Clostridioides difficile. Cas-alpha-52 endonuclease PRT derived from (des difficile) It is a row.
[0319] Sequence ID 295 is Clostridioides difficile. Cas-alpha 53 endonuclease PRT derived from (des difficile) It is a row.
[0320] Sequence ID 296 is Clostridioides difficile. Cas-alpha-54 endonuclease PRT derived from (des difficile) It is a row.
[0321] Sequence ID 297 is Clostridium hiranonis. This is a Cas-alpha 55 endonuclease PRT sequence derived from anonis.
[0322] Sequence ID 298 is Clostridioides difficile. Cas-alpha-56 endonuclease PRT derived from (des difficile) It is a row.
[0323] Sequence ID 299 is Aneurinibacillus danix. This is a Cas-alpha 57 endonuclease PRT sequence derived from *Danicus s*. .
[0324] Sequence ID 300 is Parageobacillus thermoglucosidasius. Cas-alpha 58 derived from *Cillus thermoglucosidasius* This is an endonuclease PRT sequence.
[0325] Sequence ID 301 is Brevibacillus centrosporus. Cas-alpha 59 endonuclease PRT derived from Centrosporus It is a row.
[0326] Sequence ID 302 is Clostridium pastorianum. The Cas-alpha-60 endonuclease PRT sequence derived from *Asteurianum* be.
[0327] Sequence ID 303 is Eubacterium siraeum. This is a Cas-alpha 61 endonuclease PRT sequence derived from aeum.
[0328] Sequence ID 304 is Bacillus toyonensi This is a Cas-alpha 62 endonuclease PRT sequence derived from s).
[0329] Sequence ID 305 is derived from the genus Ruminococcus sp. This is the as-alpha63 endonuclease PRT sequence.
[0330] Sequence ID 306 is derived from the genus Ruminococcus sp. This is the as-alpha64 endonuclease PRT sequence.
[0331] Sequence ID 307 is Clostridium perfringens. The Cas-alpha65 endonuclease PRT sequence derived from *Perfringens* be.
[0332] Sequence ID 308 is Bacillus thuringiensis. This is a Cas-alpha 66 endonuclease PRT sequence derived from *Giensis*.
[0333] Sequence ID 309 is Clostridium perfringens. The Cas-alpha67 endonuclease PRT sequence derived from *Perfringens* be.
[0334] Sequence ID 310 is derived from Bacillus cereus. This is the as-alpha68 endonuclease PRT sequence.
[0335] Sequence ID 311 is Bacillus toyonensi This is a Cas-alpha 69 endonuclease PRT sequence derived from s).
[0336] Sequence ID 312 is Bacillus toyonensi This is a Cas-alpha 70 endonuclease PRT sequence derived from s).
[0337] Sequence ID 313 is Bacillus toyonensi This is a Cas-alpha 71 endonuclease PRT sequence derived from s).
[0338] Sequence ID 314 is Alicyclobacillus acidoteles oris. Cas-alpha-72 endothelial (derived from *Cillus acidoterrestris*) This is the Clease PRT sequence.
[0339] Sequence ID 315 is Clostridium tetanus. i) The Cas-alpha 73 endonuclease PRT sequence derived from [source].
[0340] Sequence ID 316 is Candidatus levibacteria (Candida). Cas-Alpha 74 derived from tus Levybacteria bacterium This is an endonuclease PRT sequence.
[0341] Sequence ID 317 is derived from Bacillus cereus. This is the as-alpha75 endonuclease PRT sequence.
[0342] Sequence ID 318 is derived from Bacillus cereus. This is the as-alpha76 endonuclease PRT sequence.
[0343] Sequence ID 319 is derived from Bacillus cereus. This is the as-alpha77 endonuclease PRT sequence.
[0344] Sequence ID 320 is Clostridium paraptriphycum. Cas-alpha-78 endonuclease PR derived from Paraputrificum This is a T-array.
[0345] Sequence ID 321 is derived from Bacillus cereus. This is the as-alpha79 endonuclease PRT sequence.
[0346] Sequence ID 322 is Bacillus thuringensis. This is a Cas-alpha 80 endonuclease PRT sequence derived from *Giensis*.
[0347] Sequence ID 323 is derived from Bacillus cereus. This is the as-alpha81 endonuclease PRT sequence.
[0348] Sequence ID 324 is Bacillus toyonensi This is a Cas-alpha 82 endonuclease PRT sequence derived from s).
[0349] Sequence ID 325 is derived from Bacillus cereus. This is the as-alpha83 endonuclease PRT sequence.
[0350] Sequence ID 326 is Bacillus toyonensi This is a Cas-alpha 84 endonuclease PRT sequence derived from s).
[0351] Sequence ID 327 is Bacillus wiedmannii. This is a Cas-alpha 85 endonuclease PRT sequence derived from nii).
[0352] Sequence ID 328 is derived from Bacillus cereus. This is the as-alpha86 endonuclease PRT sequence.
[0353] Sequence ID 329 is derived from Bacillus cereus. This is the as-alpha87 endonuclease PRT sequence.
[0354] Sequence ID 330 is Bacillus toyonensi This is a Cas-alpha 88 endonuclease PRT sequence derived from s).
[0355] Sequence ID 331 is derived from Bacillus cereus. This is the as-alpha89 endonuclease PRT sequence.
[0356] Sequence ID 332 is Bacillus toyonensi This is a Cas-alpha 90 endonuclease PRT sequence derived from (s).
[0357] Sequence ID 333 is Bacillus thuringiensis. This is a Cas-alpha 91 endonuclease PRT sequence derived from *Giensis*.
[0358] Sequence ID 334 is derived from Bacillus cereus. This is the as-alpha92 endonuclease PRT sequence.
[0359] Sequence ID 335 is derived from Bacillus cereus. This is the as-alpha93 endonuclease PRT sequence.
[0360] Sequence ID 336 is derived from Bacillus cereus. This is the as-alpha94 endonuclease PRT sequence.
[0361] Sequence ID 337 is Bacillus thuringensis. This is a Cas-alpha 95 endonuclease PRT sequence derived from *Giensis*.
[0362] Sequence ID 338 is a Cas-alpha derived from the genus Bacillus sp. This is a 96-endonuclease PRT sequence.
[0363] Sequence ID 339 is derived from Bacillus cereus. This is the as-alpha97 endonuclease PRT sequence.
[0364] Sequence ID 340 is derived from Bacillus cereus. This is the as-alpha98 endonuclease PRT sequence.
[0365] Sequence ID 341 is Bacillus thuringensis. This is a Cas-alpha 99 endonuclease PRT sequence derived from *Giensis*.
[0366] Sequence ID 342 is a Cas-alpha derived from the genus Bacillus sp. This is a 100-endonuclease PRT sequence.
[0367] Sequence ID 343 is derived from Prevotella copri. This is the Cas-alpha-101 endonuclease PRT sequence.
[0368] Sequence ID 344 is derived from Prevotella copri. This is the Cas-alpha-102 endonuclease PRT sequence.
[0369] Sequence ID 345 is Clostridioides difficile. Cas-alpha-103 endonuclease PRT (derived from des difficile) It is an array.
[0370] Sequence ID 346 is Clostridioides difficile. Cas-alpha-104 endonuclease PRT (derived from des difficile) It is an array.
[0371] Sequence ID 347 is Clostridioides difficile. Cas-alpha-105 endonuclease PRT (derived from des difficile) It is an array.
[0372] Sequence ID 348 is Clostridioides difficile. Cas-alpha-106 endonuclease PRT (derived from des difficile) It is an array.
[0373] Sequence ID 349 is Clostridioides difficile. Cas-alpha-107 endonuclease PRT (derived from des difficile) It is an array.
[0374] Sequence ID 350 is Clostridioides difficile. Cas-alpha-108 endonuclease PRT (derived from des difficile) It is an array.
[0375] Sequence ID 351 is Clostridioides difficile. Cas-alpha-109 endonuclease PRT (derived from des difficile) It is an array.
[0376] Sequence ID 352 is Flavobacterium thermofilm. Cas-alpha-110 endonuclease P derived from *Tetramorium thermophilum* It is an RT sequence.
[0377] Sequence ID 353 is a bacterium belonging to the genus Phascolarctoba Cas-alpha-111 endonuclease PRT sequence derived from *Cterium sp.* That is the case.
[0378] Sequence ID 354 is Bacillus pseudomycoides. It is a Cas-alpha-112 endonuclease PRT sequence derived from mycoides. .
[0379] Sequence ID 355 is Bacteroides plebeius. This is a Cas-alpha-113 endonuclease PRT sequence derived from Beius.
[0380] Sequence ID 356 is Clostridium botulinum. This is a Cas-alpha-114 endonuclease PRT sequence derived from *Ulinum*.
[0381] Sequence ID 357 is Bacillus pseudomycoides. It is a Cas-alpha-115 endonuclease PRT sequence derived from mycoides. .
[0382] Sequence ID 358 is Bacillus pseudomycoides. It is a Cas-alpha-116 endonuclease PRT sequence derived from mycoides. .
[0383] Sequence ID 359 is Clostridium botulinum. This is a Cas-alpha-117 endonuclease PRT sequence derived from *Ulinum*.
[0384] Sequence ID 360 is Clostridium botulinum. This is a Cas-alpha-118 endonuclease PRT sequence derived from *Ulinum*.
[0385] Sequence ID 361 is Clostridium botulinum. This is a Cas-alpha-119 endonuclease PRT sequence derived from *Ulinum*.
[0386] Sequence ID 362 is from the genus Hydrogenivirga sp. This is the original Cas-alpha 120 endonuclease PRT sequence.
[0387] Sequence ID 363 is Bacillus megateriu This is a Cas-alpha-121 endonuclease PRT sequence derived from m).
[0388] Sequence ID 364 is Clostridium fa This is a Cas-alpha-122 endonuclease PRT sequence derived from llax.
[0389] Sequence ID 365 is Bacteroides plebeius. This is a Cas-alpha-123 endonuclease PRT sequence derived from Beius.
[0390] Sequence ID 366 is Bacillus thuringensis. This is a Cas-alpha-124 endonuclease PRT sequence derived from *Giensis*.
[0391] Sequence ID 367 is derived from Bacillus cereus. This is the as-alpha-125 endonuclease PRT sequence.
[0392] Sequence ID 368 is a C from the genus Clostridium sp. This is the as-alpha-126 endonuclease PRT sequence.
[0393] Sequence ID 369 is Bacteroides plebeius. This is a Cas-alpha-127 endonuclease PRT sequence derived from Beius.
[0394] Sequence ID 370 is derived from Dorea longicatena. This is the original Cas-alpha 128 endonuclease PRT sequence.
[0395] Sequence ID No. 371 is sulfurihydrogenibium azolense. Cas-alpha-129 endonuclea derived from ogenibium azorense. This is the -zePRT sequence.
[0396] Novel CRISPR effects systems, and elements including such systems For example, novel guide polynucleotides / endonucleases, but not limited to the following. The complex, guide polynucleotide, guide RNA element, Cas protein, and Endonucleases, and proteins containing endonuclease functionalities (domains) Compositions and methods for this purpose are provided. Also provided are endonucleases, cleavage-ready complexes, and A combination for direct delivery of id RNA and guide RNA / Cas endonuclease complex Products and methods are provided. This disclosure further describes the modification of target sequences in the genome of cells, and Compositions and methods for gene editing and insertion of target polynucleotides into the genome of cells include.
[0397] Terms used in the claims and specification shall, unless otherwise specified, be defined as follows: As defined below, when used in this specification and the appended claims. The singular forms "a," "an," and "the" are used depending on the context. Unless otherwise specified, it should be noted that the instructions include multiple targets.
[0398] definition As used herein, "nucleic acid" means polynucleotide, deoxyribonucleotide. Contains single-stranded or double-stranded polymers of nucleotide bases or ribonucleotide bases. Nucleic acids may also include fragments and modified nucleotides. Therefore, the term "polynucleotide" is used. "Otid," "nucleic acid sequence," "nucleotide sequence," and "nucleic acid fragment" are single-stranded or double-stranded. To show RNA and / or DNA and / or RNA-DNA polymers, interchangeable Used for optional synthetic nucleotide bases, non-natural nucleotide bases, or modified nucleotides Contains rheotide bases. Nucleotides (usually found in 5'-monophosphate form) are simple In single-letter abbreviations, it is referred to as follows: in contrast to adenosine or deoxyadenosine ( "A" for RNA or DNA, and "C" for cytosine or deoxycytosine, respectively. "G" for guanosine or deoxyguanosine, "U" for uridine, deoxy "T" for thymidine, "R" for purines (A or G), and "C" or "T" for pyrimidines. For "Y", "K" for G or T, "H" for A, C or T, and "I" for inosine. , and "N" for any nucleotide.
[0399] The term "genome," when applied to prokaryotic and eukaryotic cells, or to living cells, refers to the genome found within the nucleus. Not only chromosomal DNA, but also cellular components of the cell (for example, mitochondria or positive This also includes organelle DNA found within (Cydro).
[0400] "Open Reading Frame" is abbreviated as ORF.
[0401] The term "selective hybridization" refers to stringent hybridization. Under these conditions, hybridization of nucleic acid sequences to substantially eliminate non-target nucleic acid sequences and non-target nucleic acids Compared to grading, a larger degree of detection is possible for specific nucleic acid target sequences (e.g., background Includes mention of hybridization (at least twice the number of rounds). Sequences that hybridize effectively are typically at least 80% or 90% identical to each other. They possess up to 100% sequence identity (i.e., complete complementarity).
[0402] The term "stringent conditions" or "stringent hybridization" The "conditions" include the probe targeting its target in an in vitro hybridization assay. This includes mention of the conditions under which the sequence will selectively hybridize. The specific conditions will depend on the array and will likely differ depending on the environment. By controlling the stringency of the scrubbing conditions and / or washing conditions, the process It is possible to identify a target sequence that is 100% complementary to the target (homologous probing). This means that sequence mismatches are tolerated, and sequence mismatches are allowed so that lower similarity is detected. In general, stringency conditions can be adjusted (non-homologous probing). The probe is less than approximately 1000 nucleotides in length, and can be selectively divided into 500 nucleotides. It is less than 1. Typically, stringent conditions are pH 7.0-8.3, and a short probe... (For example, 10-50 nucleotides), at least at approximately 30°C, and also long For probes (e.g., more than 50 nucleotides), at a temperature of at least approximately 60°C The salt concentration is approximately 1.5 M Na ion concentration, typically about 0.01-1.0 M Na ions. It is the concentration. Stringent conditions also include destabilization by formamide, etc. It can be formed by adding substances. An example of low stringency conditions is 37 30-35% formamide, 1M NaCl, 1% SDS(sodium dodecyl sulfate) at ℃ Hybridization with a buffer consisting of (aluminum), and 1-2 × S at 50-55°C. Washing in SC (20 × SSC = 3.0M NaCl / 0.3M sodium citrate) Purification is one example. An example of moderate stringency conditions is 37°C with 40-45% humidity. Hybridization of formamide, 1M NaCl, 1% SDS, and 55 Washing at ~60°C with 0.5-1×SSC is an example of high stringency conditions. An example is the situation at 37°C with 50% formamide, 1M NaCl, and 1% SDS. Hybridization and washing with 0.1×SSC at 60-65°C are mentioned. ru.
[0403] "Homologousity" refers to similar DNA sequences. For example, similar sequences found in donor DNA. A "homologous region to a genomic region" is a region that is analogous to a given "genomic region" of a cell or organism's genome. A homologous region is a region of DNA that has a similar sequence. It may be long enough to promote the recombination. For example, homologous regions, To have sufficient homology to undergo homologous recombination with the corresponding genomic region, at least 5-10, 5-15, 5-20, 5-25, 5-30, 5-35, 5-40, 5-45 5-50, 5-55, 5-60, 5-65, 5-70, 5-75, 5-80, 5-85 5-90, 5-95, 5-100, 5-200, 5-300, 5-400, 5-500 , 5-600, 5-700, 5-800, 5-900, 5-1000, 5-1100, 5 ~1200, 5~1300, 5~1400, 5~1500, 5~1600, 5~1700 , 5-1800, 5-1900, 5-2000, 5-2100, 5-2200, 5-23 00, 5~2400, 5~2500, 5~2600, 5~2700, 5~2800, 5~ It can contain lengths of 2900, 5-3000, 5-3100 or more bases. "Sufficient homology" means that the two polynucleotide sequences act as substrates for homologous recombination. To demonstrate that they have sufficient structural similarity to do so. This structural similarity includes each polynucle This includes the full length of the otidol fragment and the sequence similarity of the polynucleotide. Sequence similarity is determined by the Sequence identity percentage across the entire length of the column, and / or columns with 100% sequence identity. Conserved regions containing localized similarities such as subnucleotides and distributions over a portion of the sequence length This can be explained by the column identity percentage.
[0404] As used herein, “genomic region” refers to a region located on either side of the target site, This is a segment of the chromosome in the cell's genome that also includes a portion of the target site. The Nomu range is at least 5-10, 5-15, 5-20, 5-25, 5-30, 5-35 5-40, 5-45, 5-50, 5-55, 5-60, 5-65, 5-70, 5-75 , 5-80, 5-85, 5-90, 5-95, 5-100, 5-200, 5-300, 5 ~400, 5~500, 5~600, 5~700, 5~800, 5~900, 5~100 0, 5-1100, 5-1200, 5-1300, 5-1400, 5-1500, 5-1 600, 5-1700, 5-1800, 5-1900, 5-2000, 5-2100, 5 ~2200, 5~2300, 5~2400, 5~2500, 5~2600, 5~2700 , 5-2800, 5-2900, 5-3000, 5-3100 or more bases It can contain, and as a result, this genomic region undergoes homologous recombination with the corresponding homologous region. It has sufficient homology to do so.
[0405] As used herein, "homologous recombination" (HR) refers to recombination between two DNA molecules at a homologous site. This involves the exchange of DNA fragments. The frequency of homologous recombination is influenced by many factors. The amount of recombination, and the relative ratio of homologous to non-homologous recombination, vary among different organisms. Generally speaking, The length of the homologous region affects the frequency of homologous recombination events; the longer the homologous region, the higher the frequency of recombination events. The frequency increases. The length of the homologous region required to observe homologous recombination also varies by species. In many examples, homology of at least 5kb is used, but homologous recombination is performed with homology of 25-50bp. This has been observed. For example, Singer et al., (1982) Cell 3 1:25-33;Shen and Huang,(1986)Genetics 11 2:441-57;Watt et al.,(1985)Proc.Natl.Aca d.Sci.USA 82:4768-72, Sugawara and Haber, (1992) Mol Cell Biol 12:563-75, Rubnitz an d Subramani, (1984) Mol Cell Biol 4:2253-8 ;Ayares et al.,1986)Proc.Natl.Acad.Sci.U SA 83:5199-203;Liskay et al.,(1987)Genet See ics 115:161-7.
[0406] In the context of nucleic acid sequences or polypeptide sequences, "sequence identity" or "identity" is specific. Two identical items when aligned for the greatest match across the entire comparison window This refers to nucleic acid bases or amino acid residues in the sequence.
[0407] The term "percentage of sequence identity" refers to two optimal values across the entire comparison window. This refers to the value determined by comparing the lined arrays, in the comparison window. The polynucleotide or polypeptide sequence portion is optimized by combining these two sequences. In order to insert, compare the reference sequence (which does not include additions or deletions) with the added or deleted elements (sun This may include gaps. The percentage is the same nucleic acid base within both sequences. Alternatively, determine the number of positions where amino acid residues occur, obtain the number of matched positions, and then... Divide the number of positions by the total number of positions in the comparison window, and multiply the result by 100 to determine if the arrays are identical. It is calculated by obtaining the percentage of sex. Useful example of sequence identity percentage For example, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90% Other examples include 95%, or any percentage between 50% and 100%, but these are not limited to these. These identities cannot be determined. Their identities cannot be determined using any of the programs described herein. It can be determined.
[0408] Sequence alignment and the calculation of percentages of identity or similarity are performed by LASERGENE. Bioinformatics Computing Suite (DNASTAR Inc., M adison, WI)'s MegAlign® program (trademark) This can be determined using various comparison methods designed to detect homologous sequences. In the context of this application, when sequence analysis software is used for analysis, unless otherwise specifically stated. Unless otherwise understood, the analysis results will be based on the "default values" of the program mentioned. Yes. When used in this specification, "default value" refers to the software during initial initialization. This refers to any set of values or parameters that are initially loaded.
[0409] The "Clustal V method of alignment" is Clustal V (Higgins and Sharp, (1989) CABIOS 5:151-153, Higgin s et al.,(1992)Comput Appl Biosci 8:189- (as explained in 191) and LASERGENE Bioinformatics Computing suite (DNASTAR Inc., Madison, WI) Corresponds to alignment methods found within the MegAlign™ program. For alignment, the default values are GAP PENALTY=10 and GAP L Corresponds to ENGTH PENALTY=10. Protein using the Clustal method. Default parameters for calculating pairwise alignment and identity percentage of sequences The data is KTUPLE=1, GAP PENALTY=3, WINDOW=5 and DIA GONALS SAVED=5. For nucleic acids, these parameters are KTUPL E=2, GAP PENALTY=5, WINDOW=4, and DIAGONALS SA VED=4. After array alignment using the Clustal V program. By examining the "array distance" table in the same program, the "identity percentage" can be obtained. This can be done. The "Clustal W method of alignment" is Clustal W (Hi ggins and Sharp, (1989) CABIOS 5:151-153, H iggins et al.,(1992)Comput Appl Biosci 8 (Explained in 189-191) and LASERGENE Bioinf Omatics Computing Suite (DNASTAR Inc., Madison) Alignment found in the MegAlign(trademark) v6.1 program of WI) Complies with the law. Default parameters for multiple alignment (GAP PENALTY =10, GAP LENGTH PENALTY=0.2, Delay Diverge n Seqs(%)=30, DNA Transition Weight=0.5, P rotein Weight Matrix=Gonnet Series, DNA W eight Matrix (IUB). Array using the Clustal W program After alignment, by examining the "sequence distance" table in the same program, the "identity pattern" can be determined. -cent can be obtained. Unless otherwise specified, the sequence identity / The similarity value is calculated using the following parameters: GAP Version 10 (GCG, Acc This refers to values obtained using elrys (San Diego, CA): nucleotide pairing Column identity % and similarity % are calculated with a gap generation penalty weight of 50 and gap length extension. Penalty weight 3, and the nwsgapdna.cmp scoring matrix Usage; Amino acid sequence identity % and similarity % are subject to a gap generation penalty weight of 8. Use a gap length extension penalty of 2, and the BLOSUM62 scoring matrix. (Henikoff and Henikoff,(1989)Proc.Natl.A (cad.Sci.USA 89:10915). GAP is Needleman and Wunsch, (1970) J Mol Biol 48:443-53 algorithm Using the `m` method, we maximize the number of matches and minimize the number of gaps across the two arrays. Find the alignment. GAP considers all possible alignments and gap positions. Then, using a gap creation penalty and a gap extension penalty for each matching base unit, Then, create an alignment with the maximum number of matched bases and the minimum gap. "LAST" is used to find similarity regions between biological sequences, and is a national biotechnological organization. This is a search algorithm provided by the National Center for Biotechnology Information (NCBI). The program compares a nucleotide sequence or protein sequence with a sequence database to determine if it matches. The statistical significance of the sequence is calculated, and sequences that are sufficiently similar to the query sequence are selected based on their similarity. Identify the sequence in a way that is not expected to have occurred. BLAST identifies the sequence and its intervals. We report local alignments for aligned sequences. Sequence identity at many levels Polypeptides from other species, or modified naturally or synthetically (the Such polypeptides are useful in identifying those with the same or similar function or activity. This will be well understood by those skilled in the art. As a useful example of identity percentage, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% Any percentage between 50% and 100% can be cited, but is not limited to these. In fact, Any amino acid identity between 50% and 100% may be useful in describing this disclosure, for example, 51% 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61% 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71% 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81% 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91% 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% may be useful. ru.
[0410] Polynucleotide sequences and polypeptide sequences, their variants and their sequences Structural relationships are referred to as "homonymy," "homonymous," and "substantial" in this specification, and are interchangeable terms. It can be described as "identical," "substantially similar," and "substantially corresponding." These are changes in one or more amino acids or nucleotide bases that affect the function of the molecule, for example. affecting the ability to mediate gene expression or produce a particular phenotype. This refers to polypeptide sequences or nucleic acid sequences that lack [something]. These terms also refer to the function of the resulting nucleic acid. This refers to a modification of the nucleic acid sequence in which the properties remain substantially unchanged compared to the original, unmodified nucleic acid. These modifications include the deletion, substitution, and / or insertion of one or more nucleotides in a nucleic acid fragment. It includes. Substantially similar nucleic acid sequences are included (moderate stringency - Under conditions, for example, 0.5 × SSC, 0.1% SDS, at 60°C) the formulations exemplified herein Ability to hybridize with a column or any portion of a nucleotide sequence disclosed herein. They can be defined by and they are functional with any of the nucleic acid sequences disclosed herein. It is equivalent to adjusting the stringency condition to obtain fragments with a moderate degree of similarity, for example. Homologous sequences from distantly related organisms are used with highly similar fragments, such as functional yeasts from closely related organisms. It can be used to screen for genes that replicate basic elements, etc. Post-hybridization The cleaning process determines the stringency conditions.
[0411] A "centimorgan" (cM) or "map unit" is a combination of two polynucleotide sequences. The distance between any pair of syngenes, markers, target sites, gene loci, or any pair thereof, where 1% of the results of meiosis are recombination. Therefore, Centi Morgan states that two linked recombinations are possible. Equivalent to an average recombination frequency of 1% between genes, markers, target sites, loci, or any pair thereof. It is equivalent to a certain distance.
[0412] "Isolated" or "purified" nucleic acid molecules, polynucleotides, polypeptides or Proteins, or their biologically active portions, are usually found in their naturally occurring environments. such as polynucleotides or proteins that occur simultaneously with or interact with them It substantially or essentially does not contain the components. Therefore, isolated or purified polynucleotides If the cytoplasm or protein is produced by recombinant technology, other cellular materials may also be cultured. It does not contain any other substances, and if it is chemically synthesized, the chemical precursors are also other chemicals. It also contains virtually no chemical substances. Ideally, it should be an "isolated" polynucleotide (ideally The protein-coding sequence is the genome DN of the organism from which its polynucleotides originate. In A, there is a sequence that is naturally adjacent to that polynucleotide (i.e., that polynucleotide It does not include sequences located at the 5' and 3' ends of Otide. For example, in various embodiments In the case of isolated polynucleotides, the genome of the cell from which those polynucleotides originated is formed. The nucleotide sequences contained in DNA that are naturally adjacent to the polynucleotides are It could be approximately 5kb, 4kb, 3kb, 2kb, 1kb, 0.5kb, or less than 0.1kb. The isolated polynucleotides can be purified from cells in which they naturally occur. Using conventional nucleic acid purification methods known to the industry, we obtain isolated polynucleotides. This can be done. This term includes recombinant polynucleotides and chemically synthesized polynucleotides. Chido is also included.
[0413] The term "fragment" refers to a sequence of nucleotides or amino acids. In this state, the fragments are 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 1 It consists of 5, 16, 17, 18, 19, 20, or more than 20 consecutive nucleotides. Morphologically, the fragments are 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, It is a sequence of 15, 16, 17, 18, 19, 20, or more than 20 consecutive amino acids. The fragment is, Even if the sequence exhibits the function of sharing some percentage of identity over the entire length of the aforementioned fragment It is not necessary to show it.
[0414] The terms "functionally equivalent fragment" and "functionally equivalent fragment" are used herein. These terms are used interchangeably. These terms refer to the same activity or function as the longer sequence from which they are derived. This refers to an isolated nucleic acid fragment or part or partial sequence of a polypeptide that exhibits the following characteristics. In this context, the fragment, regardless of whether it encodes an active protein or not, influences gene expression. It retains the ability to change or generate a specific phenotype. For example, a fragment can be modified. It can be used to design genes that produce a desired phenotype in a plant. The gene is The nucleic acid fragment, whether or not it encodes an active enzyme, is part of the plant promoter sequence. It is used in a suppressive manner by coupling it with sense orientation or antisense orientation. It can be designed in this way.
[0415] "Genes" include, but are not limited to, regulatory sequences preceding the coding sequence (5' A specific sequence including a non-coding sequence and a subsequent regulatory sequence (3' non-coding sequence). It contains nucleic acid fragments that express functional molecules such as proteins. "Native genes" are This refers to genes found in their naturally occurring endogenous locations, along with their own regulatory sequences.
[0416] The term "endogenous" refers to sequences or other molecules that are naturally present in cells or organisms. In some embodiments, endogenous polynucleotides are typically found in the cellular genome; that is, different It's not a seed.
[0417] An "allele" is one of several alternative forms of a gene that occupies a given gene locus on a chromosome. If all alleles present at a given gene locus on a chromosome are the same, then the plant body is... They are homozygous at the gene locus. The alleles present at a given gene locus on the chromosome are different. If so, the plant is heterozygous at that gene locus.
[0418] A "coding sequence" refers to a polynucleotide sequence that codes for a specific amino acid sequence. It tastes good. The "adjustment array" is the upstream (5' non-coding array) of the coding array, coding Nuclei located within the coding sequence or downstream of the coding sequence (3' non-coding sequence) This refers to the ocidal sequence, and the transcription, RNA processing, or stabilization of the associated coding sequence. It affects sex or translation. Regulatory sequences include, but are not limited to, promoters. Translation leader sequence, 5' untranslated sequence, 3' untranslated sequence, intron, polyadenylated label Examples include the target sequence, RNA processing site, effector binding site, and stem-loop structure. It can be done.
[0419] A "mutated gene" is a gene that has been altered through human intervention. A "mutant gene" is a gene that has been altered by the addition, deletion, or substitution of at least one nucleotide. It has a sequence different from the sequence of the non-mutant gene. In certain embodiments of this disclosure, this mutation The gene is a guide polynucleotide / Cas endonuclease disclosed herein. This includes modifications resulting from the Tem. A mutant plant is a plant containing a mutant gene.
[0420] As used herein, “targeted mutation” refers to the inducible Cas mutation disclosed herein. Using any method known to those skilled in the art, including a method involving a donuclease system, the target Genes (targets) are created by modifying target sequences within genes, including natural genes. It is a mutation in a gene (called a target gene).
[0421] The terms "knockout," "gene knockout," and "genetic knockout" are, in fact, In the detailed documentation, they are used interchangeably. Knockout is a targeting by the Cas protein. This refers to the DNA sequence of a cell that has been partially or completely disabled by a genotype, for example, a genotype. The DNA sequence before coding may have encoded an amino acid sequence, and may have had a regulatory function (for example) They may have had a promoter.
[0422] The terms "knock-in," "gene knock-in," "gene insertion," and "genetic knock-in" The terms "and" are used interchangeably in this specification. Knock-in is a type of terminating process that uses the Cas protein. By getting (for example, by homologous recombination (HR)), suitable donor DNA polynucleotides (Rheotides are also used) Substitution or insertion of DNA sequences in specific DNA sequences within cells This means that an example of knock-in is the heterologous amino acid coding in the coding region of a gene. This involves the specific insertion of a ligature sequence, or the specific insertion of a transcriptional regulatory element at a gene locus.
[0423] A "domain" is a continuous stretch of nucleotides (RNA, DNA and / or RNA- It can refer to a DNA combination sequence, or a continuous stretch of amino acids.
[0424] The terms "conserved domain" or "motif" refer to a series of polynucleotides or evolutionarily related domains. This refers to a set of amino acids conserved at specific positions in the aligned sequence of a protein. While amino acids at other positions may vary among homologous proteins, those at specific positions can be highly maintained. The amino acids referred to are those essential for the structure, stability, or activity of proteins. These are due to the high degree of conservation of the aligned sequences of their protein homolog family. To be identified, proteins with newly determined sequences are identified as previously identified proteins. It can be used as an identifier or "signature" to determine whether or not it belongs to a family. can.
[0425] "Codon-modified gene," "codon-prioritized gene," or "codon-optimized gene" refers to: It has codon usage frequencies designed to mimic the preferred codon usage frequencies of chief cells. It refers to the gene that...
[0426] "Optimized" polynucleotides are those whose expression is improved in specific heterologous host cells. This refers to an array that has been optimized for a specific purpose.
[0427] "Plant-optimized nucleotide sequences" refer to sequences that improve expression in plants, particularly in plants. This is a nucleotide sequence optimized for current growth. The plant-optimized nucleotide sequence is , including codon-optimized genes. Plant-optimized nucleotide sequences are one for improving expression. Using the above plant-preferred codons, for example, the Cas endonuclea disclosed herein It can be synthesized by modifying the nucleotide sequence that codes for proteins such as enzymes. This is possible. For example, in the consideration of the use of host-preferred codons, Campbell a See nd Gowri (1990) Plant Physiol. 92:1-11. sea bream.
[0428] The promoter recognizes RNA polymerase and other proteins that initiate transcription. This is the DNA region involved in the synthesis. The promoter sequence is the proximal and more distal upstream element. It consists of elements, and the latter element is often called an enhancer. " is a DNA sequence that can stimulate promoter activity, and is unique to the promoter. It may be an element, or an insertion to enhance the promoter level or tissue specificity. The introduced heterogeneous elements may also be present. The promoter as a whole is derived from natural genes. It may be done with different elements derived from different promoters found in nature. It may be composed of and / or include synthetic DNA segments. Various promotions The organism, in various tissue or cell types, or at various developmental stages, or under various environmental conditions It is understood by those skilled in the art that gene expression can be induced accordingly. In some cases, the exact boundaries of the regulatory array are not fully defined, which can lead to some variations. It has also been recognized that DNA fragments from different ethers may possess identical promoter activity. It is being done.
[0429] In most cell types, the promoter that causes the most gene expression is generally the "promoter". It is called a "chemical promoter." The term "inducible promoter" refers to, for example, a chemical promoter. In response to the presence of endogenous or exogenous stimuli, or the environment, In response to hormonal, chemical, and / or developmental signals, coding sequences or mechanisms This refers to a promoter that selectively expresses active RNA. It is different from inductive or regulatory promoters. For example, light, heat, stress, flooding or drought, salt stress, osmotic pressure Stress, plant hormones, wounds, or ethanol, abscisic acid (ABA), jasmon Promo is induced or modulated by chemical substances such as phosphate, salicylic acid, or toxicity mitigators. One example is [the person who is responsible for this].
[0430] The "translation leader sequence" is located between the promoter sequence and the coding sequence of a gene. This refers to a polynucleotide sequence. The translation leader sequence is upstream of the translation initiation sequence of mRNA. It is located in [location]. The translation leader sequence is involved in the processing of the primary transcript into mRNA, and the mRNA This may affect stability or translation efficiency. Examples of translation reader sequences have been reported (e.g., Turner and Foster,(1995)Mol Biotechnol 3 :225-236).
[0431] "3' non-coding sequence", "transcription terminator", or "termination sequence" is a coding sequence. This refers to the DNA sequence located downstream of the ping sequence, including the polyadenylation recognition sequence and the mRNA programming sequence. Other sequences that encode regulatory signals capable of affecting cessation or gene sequencing. It includes. The polyadenylation signal is usually the polyadenylation of the 3' end of the mRNA precursor. Characterized by its influence on the addition of nilate systems. Use of different 3' non-coding sequences. Ingelbrecht et al., (1989) Plant Cell 1: Examples are given in pages 671-680.
[0432] RNA transcripts are produced by the transcription of DNA sequences catalyzed by RNA polymerase. It means the product of. If the RNA transcript is a completely complementary copy of the DNA sequence, It is called the primary transcript or pre-mRA. The RNA transcript is produced after the transcription of the primary transcript preRNA. When the RNA sequence is obtained through processing, it is called mature RNA or mRNA. Messenger RNA or mRNA does not have introns, and is used by cells This refers to RNA that can be translated into protein. "cDNA" is complementary to the mRNA template. It is the target, and refers to DNA synthesized from an mRNA template using reverse transcriptase. DNA can be single strands or two strands using a Klenow fragment of DNA polymerase I. It may be converted to a strand form. "Sense" RNA refers to an RNA transcript containing mRNA. It can be translated into proteins either intracellularly or in vitro. "Antisense RNA" is , complementary to all or part of the target primary transcript or mRNA, and expressing the target gene expression This refers to RNA transcripts that block (for example, U.S. Patent No. 5,107,065). (See reference). The complementarity of antisense RNA is such that any part of a specific gene transcript can be replaced. In other words, 5' non-coding sequences, 3' non-coding sequences, introns or coding sequences Complementarity with the sequence is acceptable. "Functional RNA" refers to antisense RNA, ribozyme This refers to RNA, or other RNA that is not translated but affects intracellular processes. The terms "complement" and "reverse complement" are interchangeable in this specification with respect to mRNA transcripts. It is intended to be used to define the antisense RNA of the message.
[0433] The term "genome" refers to the genetic material present in each cell of an organism, or in a virus or organelle. The entire constituent material (genes and non-coding sequences); and / or from one parent ( A haploid refers to a set of chromosomes inherited as a unit.
[0434] The term "operably linked" refers to a single nucleic acid in which the function of one is regulated by the function of the other. This refers to the association of nucleic acid sequences on a fragment. For example, a promoter controls the expression of a coding sequence. If it is possible to control this coding array, it is operably linked to this coding array (that is, (Note: This coding sequence is under the transcriptional control of the promoter.) The coding sequence is, It can be functionally linked to an adjustable array in the sense or antisense direction. The complementary RNA region is directly or indirectly located at the 5' end of the target mRNA, or at the target mR It can bind operably to the 3' end of NA, or within the target mRNA, or the first The complementary region is the 5' end of the target mRNA, and its complement is the 3' end of the target mRNA. ru.
[0435] Generally, a "host" is a heterogeneous component (polynucleotide, polypeptide, other molecules, cell). This refers to an organism or cell into which a heterologous polymorph has been introduced. As used herein, "host cell" means a heterologous polymorph. In vitro cultured as single-celled organisms into which nucleotides or polypeptides have been introduced. In vitro or in vitro eukaryotic cells, prokaryotic cells (e.g., bacterial or archaeal cells), or multicellular organisms This refers to cells derived from a substance (e.g., cell lines). In some embodiments, the cells are derived from the following: Selected from the following groups: archaeal cells, bacterial cells, eukaryotic cells, eukaryotic unicellular organisms, somatic cells, reproductive cells. cells, stem cells, plant cells, algae cells, animal cells, invertebrate cells, vertebrate cells, fish cells Cells, frog cells, bird cells, insect cells, mammalian cells, pig cells, bovine cells, goat cells, Sheep cells, rodent cells, rat cells, mouse cells, non-human primate cells, and human cells. In some cases, the cells are in vitro. In some cases, the cells are in vivo. .
[0436] The term "recombinant" refers to the isolation of nucleic acids, for example, by chemical synthesis or genetic engineering techniques. Segment manipulation creates an artificial pair of two otherwise separate sequence segments. It refers to a combination or pairing.
[0437] The terms "plasmid," "vector," and "cassette" refer to cells that are not part of the central metabolism of a cell. This refers to additional linear or circular chromosomal elements that often carry genes, usually of double-stranded DNA. It takes shape. Such elements are single-stranded or double-stranded DNA or R in a linear or circular form. NA, of any origin, self-replicating sequence, genome integration sequence, phage or nucleus These could be ocidal sequences, and among them, many nucleotide sequences are the target polynucleotides. It is either bound to or recombinant with a specific construct that allows the cytoplasm to be introduced into cells. A "transformation cassette" contains genes, and in addition to those genes, it also contains traits specific to a host cell. This refers to a specific vector that contains elements that promote conversion. An "expression cassette" contains genes. Furthermore, in addition to the gene, a specific vector having an element that causes the host to express that gene. It refers to.
[0438] Terms: "recombinant DNA molecule," "recombinant DNA construct," "expression construct," "construct," The terms "recombinant construct" and "recombinant construct" are used interchangeably in this specification. Recombinant DNA construct Lacto consists of nucleic acid fragments, such as regulatory sequences and co- This includes artificial combinations of sequencing sequences. For example, recombinant DNA constructs may have different origins. It contains regulatory and coding sequences derived from the source, or originates from the same source but is not found in nature. The methods may include adjustment sequences and code sequences arranged in different ways. The structure can be used on its own or in conjunction with a vector. In such cases, the selection of the vector is, as is well known to those skilled in the art, to introduce the vector into the host cell. It depends on the method used for introduction. For example, a plasmid vector can be used. Skilled technicians can successfully transform, select, and proliferate host cells. Therefore, they have a thorough understanding of the genetic elements that must be present on the vector. Skilled technicians. Furthermore, different independent transformation events result in different expression levels and patterns. (Jones et al.,(1985)EMBO J 4:2411-2418;D e Almeida et al.,(1989)Mol Gen Genetics 218:78-86), therefore, in order to obtain a strain exhibiting the desired expression level and pattern, Therefore, it should be recognized that multiple events are usually screened. Cleaning involves Southern blot analysis of DNA, Northern blot analysis of mRNA expression, PCR, and real-time analysis. Quantitative PCR (qPCR), reverse transcription PCR (RT-PCR), immunoblotting of protein expression Standard molecular biological, biological analysis, including enzyme or activity analysis and / or phenotypic analysis. This can be done by chemical and other analytical methods.
[0439] The term "heterogeneous" refers to the original environment or context of a particular polynucleotide or polypeptide sequence. This refers to the difference between a place or composition and its current environment, location, or composition. Non-restrictive examples include: Therefore, differences in taxonomic origin (for example, derived from Zea mays) The polynucleotide sequence is the genome of the rice plant (Oryza sativa), or the z When inserted into a different subspecies or variety of Zea mays, it is a different species. (or polynucleotides obtained from bacteria have been introduced into plant cells), or the sequence Differences (for example, derived from Zea mays, isolated, modified, and Examples include polynucleotide sequences reintroduced into sorghum plants. In this case, the "different species" associated with the sequence refers to sequences originating from different species, subspecies, or introduced species. , or if from the same species, the composition and / or genomic loci are due to intentional human intervention. This can refer to sequences that have been substantially altered from their natural form. For example, heterologous polynucleotides The promoter operably linked to the nucleotide is different from the species from which this polynucleotide originated. If it is from the same / similar species, then one or both are from the original species. The morphology and / or genomic locus are substantially modified, or this promoter —However, it is not a naturally occurring promoter of operable, linked polynucleotides. Or, this One or more regulatory regions and / or polynucleotides shown in the specification are entirely synthetic. In another example, the target polynucleotide for cleavage by Cas endonuclease is Cas endonuclease may belong to a different organism. In another example, Cas endonuclease Nucleases and guide RNAs are templates for insertion into target polynucleotides. Alternatively, it may be introduced into the target polynucleotide along with additional polynucleotides that act as donors. Here, additional polynucleotides are the target polynucleotide and / or Casuen. It is a different species from dunuclease.
[0440] The term "expression," as used herein, refers to a precursor or a mature functional end product (e.g., For example, it refers to the production of mRNA, guide RNA, or protein.
[0441] "Mature" proteins are polypeptides that have undergone post-translational processing (i.e., primary translation). This refers to a polypeptide from which all prepeptides or propeptides present in the product have been removed. .
[0442] "Precursor" proteins are the primary products of mRNA translation (i.e., prepeptides and prepeptides). This means that the propeptides are still present. Prepeptides and propeptides are This could be, but is not limited to, an intracellular localization signal.
[0443] The CRISPR (Clustered Regular Arrangement Short Palitic Sequence Repeat) gene locus is, for example, DNA cleavage is used by bacterial and archaeal cells to destroy foreign DNA. This refers to a specific gene locus that codes for a component of the TEM (Horvath and Barran). gou,2010,Science 327:167-170;Published March 1, 2007 (International Publication No. 2007 / 025097 pamphlet). The CRISPR locus is short Short direct repeats separated by variable DNA sequences (called spacers) It can be composed of a CRISPR array containing CRISPR repeats, which include various Ca The s(CRISPR-related) gene may be located adjacent to it.
[0444] As used herein, "effector" or "effector protein" refers to poly This includes recognizing, binding to, and / or cleaving or nicking nucleotide targets. It is a protein that contains activity. Effectors or effector proteins are also called It can be an endonuclease. The "effector complex" of the CRISPR system contains c It contains Cas proteins involved in the recognition and binding of rRNA and targets. Some of the proteins may further contain domains involved in target polynucleotide cleavage. .
[0445] The term "Cas protein" is encoded by the Cas (CRISPR-related) gene. It refers to a protein. Cas proteins are encoded by genes at the Cas locus. It contains proteins, as well as adaptation molecules and interference molecules. Interfering molecules in the combination include endonucleases. Cas end nucleases as described herein A nuclease contains one or more nuclease domains. Cas endonuclease and Therefore, the novel Cas-alpha protein and Cas9 protein disclosed herein Cpf1 (Cas12) protein, C2c1 protein, C2c2 protein, C2 c3 protein, Cas3, Cas3-HD, Cas5, Cas7, Cas8, Cas1 Examples include, but are not limited to, 0, or combinations or complexes thereof. C When the as protein is a complex with a suitable polynucleotide component, a specific polynucleotide Recognizes, binds to, and selectively nicks or cleaves all or part of the rheotide target sequence. "Cas endonuclease" or "Cas effector protein" that can perform this function. It is possible that the Cas-alpha-endonuclease of this disclosure contains one or more RuvC nucleases. This includes those that have a crease domain. Cas proteins are natural Cas proteins. Functional fragments or functional variants of the protein, or at least 5 of the natural Cas protein 0, 50-100, at least 100, 100-150, at least 150, 150-2 00, at least 200, 200-250, at least 250, 250-300, less 300, 300-350, at least 350, 350-400, at least 400, 400-450, at least 500, or more than 500 consecutive amino acids and at least 5 0%, 50%-55%, at least 55%, 55%-60%, at least 60%, 60% ~65%, at least 65%, 65%~70%, at least 70%, 70%~75%, small At least 75%, 75%-80%, at least 80%, 80%-85%, at least 85 %, 85%~90%, at least 90%, at least 90%~95%, at least 95% 95%~96%, at least 96%, 96%~97%, at least 97%, 97%~9 8%, at least 98%, 98% to 99%, at least 99%, 99% to 100%, or A protein that has 100% sequence identity and retains the activity of at least a portion of the natural sequence. It is further defined as quality.
[0446] "Functional fragments," "functionally equivalent fragments," and "functional" of Cas endonucleases. "Equivalent fragments" are used interchangeably in this specification to recognize, combine, and arbitrarily identify target sites. The ability to selectively unravel, nicking, or sever (single strand within the target site) is retained. (or introduces a double-strand break) a portion or partial sequence of the Cas endonuclease sequence of the present disclosure This refers to a part or partial sequence of this Cas endonuclease, one of its domains. A single complete or partial (functional) peptide, for example, Cas3, but not limited to the following: All functional parts of the HD domain, all functional parts of the Cas3 helicase domain, protein All functional parts of quality (but not limited to Cas5, Cas5d, Cas7 and Cas It can include (e.g., 8b1).
[0447] Cas endonuclease or Cas alpha containing Cas-alpha as described herein "Functional variants" of effector proteins, "functionally equivalent variants", and The terms "functionally equivalent variant" and "functionally equivalent variant" are used interchangeably in this specification. It recognizes, binds to, selectively unravels, nicks, or cleaves all or part of the target sequence. The ability to do so is retained in the variants of the Cas effector protein disclosed herein. To point.
[0448] Cas endonucleases may also include multifunctional Cas endonucleases. The terms "multifunctional Cas endonuclease" and "multifunctional Cas endonuclease" The term "polypeptide" is used interchangeably in this specification and refers to Cas endonuclease function (C At least one protein that can act as an endonuclease is a domei. (including ), and, but not limited to, functions that form complexes (complexes with other proteins). Few include at least a second protein domain capable of forming a fusion. This includes reference to a single polypeptide having at least one other function. In one embodiment, multiple Functional Cas endonucleases have a domain typical of Cas endonucleases. (internal, upstream (5'), downstream (3'), or both internal 5' and 3', or these) (any combination of the above), including at least one further protein domain.
[0449] The terms “cascade” and “cascade complex” are used interchangeably herein. They assemble with polynucleotides to form polynucleotide-protein complexes (PNPs). This includes references to the multi-subunit protein complex obtained. The cascade is the assembly of the complex. Furthermore, for stability and the identification of target nucleic acid sequences, PNPs are polynucleotide-dependent. The cascade is complementary to the variable targeting domain of the guide polynucleotide. It functions as a surveillance complex that identifies target nucleic acids and binds to them selectively.
[0450] Terms: "Ready to cut cascade", "CR cascade", "Ready to cut cascade complex "Combination", "CR Cascade Complex", "Cutting Ready Cascade System", "CRC" And “cr cascade system” is used interchangeably in this specification, and polynucleotide It can assemble with polynucleotides to form nucleotide-protein complexes (PNPs). This includes references to multi-subunit protein complexes, such as cascade tans. One type of protein recognizes, binds to, and selectively unwinds all or part of a target sequence. It is a Cas endonuclease that can cut or cleave.
[0451] The terms "5'-cap" and "7-methylguanylate (m7G) cap" refer to the following terms in this specification. In the text, they are used interchangeably. The 7-methylguanylate residue is a messenger in eukaryotic cells. It is located at the 5' end of RNA (mRNA). RNA polymerase II (Pol II) It transcribes mRNA within eukaryotic cells. Messenger RNA capping is generally performed as follows: It occurs as follows: The terminal 5' phosphate group of the mRNA transcript is affected by the RNA terminal phosphatase It is removed by guanosine monophosphate (GMP), leaving two terminal phosphate groups. Nilyltransferase adds to the terminal phosphate group of the transcript, and the 5' at the end of the transcript -5' triphosphate-linked guanine remains. Finally, the 7-nitrogen of this terminal guanine becomes methylated. It is methylated by lanceferase.
[0452] In this specification, the term "without a 5'-cap" means "without a 5'-cap." For example, it is used to mean RNA that has a 5'-hydroxyl group. RNA can be called, for example, "uncapped RNA." 5'-capped RNA Since it undergoes nuclear export, uncapped RNA accumulates more sufficiently in the nucleus after transcription. This is possible. One or more RNA components in this specification are not capped.
[0453] As used herein, the term “guide polynucleotide” means “Cas endonucleus.” It can form a complex with an enzyme, for example, the Cas endonuclease described herein. Cas endonucleases recognize DNA target sites, bind to them selectively, and perform various actions. This refers to a polynucleotide sequence that allows for selective cleavage. A DNA sequence is an RNA sequence, a DNA sequence, or a combination thereof (RNA-DNA combination). It may be a mixed arrangement.
[0454] A "functional fragment" of guide RNA, crRNA, or tracrRNA, "functionally equivalent to" The terms “a fragment” and “functionally equivalent fragment” are used interchangeably in this specification. Each has the ability to function as a guide RNA, crRNA, or tracrRNA, respectively. A portion of the guide RNA, crRNA, or tracrRNA of this disclosure that retains or Refers to a subarray.
[0455] "Functional variants" of guide RNA, crRNA, or tracrRNA (each, respectively) "Functionally equivalent variant" and "functionally equivalent variant" are defined herein. These are used interchangeably and function as guide RNA, crRNA, or tracrRNA, respectively. The guide RNA, crRNA, or tracrRNA of this disclosure retains the ability to perform the function. This refers to a variant of [the original].
[0456] The terms "single guide RNA" and "sgRNA" are used interchangeably in this specification. , fused to tracrRNA (trans-activated CRISPR RNA) (tracrR Variable targeting (bound to tracr mate sequence that hybridizes to NA) crRNA (CRISPR RNA) containing lin is involved in the synthetic fusion of two RNA molecules. Regarding this, single guide RNA forms a complex with type II Cas endonuclease. crRNA or crRNA fragments of a type II CRISPR / Cas system that can, may contain tracrRNA or tracrRNA fragments, and the guide RNA / Cas enzyme The donuclease complex can guide Cas endonuclease to the DNA target site. Cas endonucleases can recognize the target site on the DNA and bind to it selectively. Furthermore, it allows for selective nicking or cleavage (introducing single-strand or double-strand breaks). ru.
[0457] The terms “variable targeting domain” or “VT domain” are interchangeable in this specification. It is used to hybridize one strand (nucleotide sequence) of the double-stranded DNA target site. It contains complementary nucleotide sequences. The first nucleotide sequence domain (VT) The complementarity rate between the domain and the target sequence is at least 50%, 51%, 52%, or 53%. 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63% 63%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73% 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83% 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93% It could be 94%, 95%, 96%, 97%, 98%, 99%, or 100%. The length of the getting domain is at least 12, 13, 14, 15, 16, 17, 18 , 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 Nucle It may be ocid. In some embodiments, the variable targeting domain is 12-3 It contains a stretch of 0 nucleotides. The variable targeting domain is DNA Sequences, RNA sequences, modified DNA sequences, modified RNA sequences, or any combination thereof It can be constructed.
[0458] The term (of guide polynucleotides) "Cas endonuclease recognition domain" or " The term "CER domain" is used interchangeably in this specification and refers to Cas endonuclease polyp It contains a nucleotide sequence that interacts with ptide. The CER domain is (trans action) t The racr nucleotide mate sequence is followed by the tracr nucleotide sequence. The main types are DNA sequences, RNA sequences, modified DNA sequences, and modified RNA sequences (for example, 201 The specification of U.S. Patent Application Publication No. 2015 / 0059010A1 was published on February 26, 2005. (See [reference]), or can consist of any combination thereof.
[0459] As used herein, the term "guide polynucleotide / Cas endonuclease" is used with respect to the present invention. "Complex", "Guide polynucleotide / Cas endonuclease system", "Guide "Polynucleotide / Cas complex", "Guide polynucleotide / Cas system", "Inducible Cas system," "Polynucleotide-induced endonuclease," "PGEN" These are interchangeable and can form a complex with at least one gas used interchangeably in this specification. This refers to an idiopolynucleotide and at least one Cas endonuclease. Here, the guide polynucleotide / Cas endonuclease complex is Cas endo Nucleases can be guided to DNA target sites, and Cas endonucleases can target DNA A. Recognizes, binds to, and selectively nicks or cleaves the target site (single-strand or double-stranded). This enables the introduction of a main-strand break. Guide polynucleotide / C in this specification The as endonuclease complex is a known CRISPR system (Horvath an d Barrangou,2010,Science 327:167-170;Mak arova et al.2015,Nature Reviews Microbio logy Vol.13:1-15;Zetsche et al.,2015,Cel l 163,1-13;Shmakov et al.,2015,Molecular One of the Cas proteins and suitable polynucleotides (Cell 60, 1-13) May contain taint.
[0460] Terms: "guide RNA / Cas endonuclease complex", "guide RNA / Cas e "Endonuclease system", "Guide RNA / Cas complex", "Guide RNA / Ca "s system", "gRNA / Cas complex", "gRNA / Cas system", "RNA In this specification, "inducible endonuclease" and "RGEN" are used interchangeably and are complex. at least one RNA component and at least one Cas end that can form This refers to a nuclease, and here it refers to the guide RNA / Cas endonuclease complex. The body can induce Cas endonucleases at DNA target sites, and Cas endonucleases The crease recognizes, binds to, and selectively nicks or cleaves a target site on the DNA. This makes it possible to introduce single-strand or double-strand breaks.
[0461] Terms: "target site," "target sequence," "target site sequence," "target DNA," "target gene" "Locus", "Genome target site", "Genome target sequence", "Genome target gene locus", and "Proto The term "spacer" is used interchangeably in this specification and is not limited to the following, but applies to guide lines. The creotide / Cas endonuclease complex recognizes, binds to, and selectively nitrates. A cell's chromosome, episome, gene locus, or gene that can be cut or cleaved. Any other DNA molecules in the chromosome (e.g., chromosomal DNA, chloroplast DNA, mitochondrial DNA) This refers to polynucleotide sequences, such as nucleotide sequences on plasmid DNA (NA). The target site may be an endogenous site in the cell's genome, or the target site may be heterogeneous to the cell. Therefore, it does not naturally exist in the cell's genome, or the target site does not occur naturally. It can be found at different genomic locations relative to the site. When used herein, the term "endogenous mark" is used. The terms "target sequence" and "natural target sequence" are used interchangeably herein and refer to sequences inherent in the cellular genome. It either occurs in the cell's genome, and the intrinsic location or origin of its target sequence within the cell's genome. This refers to the target sequence present at the biological site. "Artificial target site" or "artificial target sequence" is defined herein. In other words, it is used interchangeably and refers to a target sequence introduced into the genome of a cell. The target sequence is identical in sequence to the endogenous target sequence or natural target sequence in the cell's genome, but the details... They can be located at different positions within the cytoplasm's genome (i.e., non-intrinsic or non-developmental locations).
[0462] In this specification, “protospacer adjacent motif” (PAM) refers to the guide described herein. Recognized (targeted) by the dopolynucleotide / Cas endonuclease system This refers to a short nucleotide sequence adjacent to the target sequence (protospacer). Without the PAM sequence, the target DNA sequence cannot be correctly recognized. The sequence and length of PAM in the details refer to the Cas protein used, or the Cas protein It can vary depending on the complex. The PAM sequence can be of any length, but generally it is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18 It is 19 or 20 nucleotides long.
[0463] "Modified target site," "Modified target sequence," "Modified target site," and "Modified target sequence" are all terms used in the original text. The detailed documentation uses compatible terms and includes at least one modification compared to the unmodified target sequence. This refers to target sequences disclosed herein, including the following: Such “modifications” include, for example, (i (ii) substitution of at least one nucleotide, (ii) deletion of at least one nucleotide (iii) insertion of at least one nucleotide, (iv) at least one nucleo This includes chemical modifications of the cydore, or any combination of (v), (i) to (iv).
[0464] A "modified nucleotide" or "edited nucleotide" is a modified nucleotide with its unmodified nucleotide sequence. This refers to the target nucleotide sequence that includes at least one modification when compared. For example, (i) substitution of at least one nucleotide, (ii) at least (iii) deletion of one nucleotide, (iv) insertion of at least one nucleotide Chemical modification of at least one nucleotide, or any combination of (v)(i)~(iv) A combination was mentioned.
[0465] The methods for "modifying a target site" and "altering a target site" are interchangeable in this specification. It refers to a method used to generate modified target sites.
[0466] As used herein, “donor DNA” refers to the target site of Cas endonuclease. It is a DNA construct containing the polynucleotides intended for insertion.
[0467] The term "polynucleotide modification template" refers to a comparison with the nucleotide sequence to be edited. and contains polynucleotides with at least one nucleotide modification. Modification may be the substitution, addition, or deletion of at least one nucleotide. (Optional) Furthermore, polynucleotide modification templates have at least one nucleotide modification. It may contain adjacent homologous nucleotide sequences, and these adjacent homologous nucleotide sequences are editable. It provides sufficient homology to the desired nucleotide sequence of the elephant.
[0468] The term “plant-optimized Cas endonuclease” as used herein refers to plant cells or plant cells. Multifunctional, encoded by nucleotide sequences optimized for expression in organisms. This refers to Cas proteins, including Cas proteins.
[0469] "Plant-optimized nucleotide sequence encoding Cas endonuclease", "Cas E "Plant optimization constructs encoding endonucleases" and "Cas endonuclease encoding The plant-optimized polynucleotides used herein are interchangeable and apply to plant cells or This is a nucleotide sequence encoding a Cas protein optimized for expression in plants. Refers to a row, or its variant or functional fragment. Plant optimization Cas endonuclear Plants containing Cas include plants containing nucleotide sequences that encode Cas sequences and / or Cas sequences. This includes plants containing endonuclease proteins. In one embodiment, plant optimization Cas end Nuclease nucleotide sequences are optimized for corn, rice, wheat, and d It is an optimized Cas endonuclease, either optimized for Iz, cotton, or canola.
[0470] The term "plant" generally includes whole plants, plant organs, plant tissues, seeds, plant cells, seeds, and These include their descendants. Plants are either monocots or dicots. As for plant cells... While not limited to seeds, suspension cultures, embryos, meristematic tissue regions, callus tissue, leaves, roots, and scutellaria are also included. These include cells derived from pollen, gametophyte, sporophyte, pollen, and microspores. "Plant elements" refer to the whole plant. It is intended to refer to a substance or plant component, which includes differentiated and / or undifferentiated tissues, for example, This may include, but is not limited to, plant tissues, parts, and cell types. In one embodiment, Each plant element is one of the following: whole plant, seedling, meristematic tissue, basic tissue, vascular tissue, epidermal tissue. Weave, seeds, leaves, roots, shoots, stems, flowers, fruits, stolons, bulbs, tubers, corms, keiki, shi Tubes, buds, tumor tissue, and various forms of cells and cultures (e.g., single cells, protops) Last, embryo, callus tissue). Protoplasts lack a cell wall, therefore protoplasts (In a technically "intact" plant cell, as if it were naturally occurring with all its components intact) It should be noted that the term "plant organ" does not refer to plant tissue, or even to the morphological organs of a plant. This refers to a group of tissues that constitute functionally independent parts. When used herein, "plant tissues" refers to a group of tissues that constitute functionally independent parts. The term "element" is synonymous with "part" of a plant, referring to any part of a plant, a separate tissue and / or organ. It may include "official" and throughout, it can be used interchangeably with the term "organization." Similarly, "plant reproduction." An "element" generally refers to the initiation of another plant through the sexual or asexual reproduction of that plant. Any part of a plant that can be grown, for example, but not limited to, seeds, seedlings, roots, shoots, It refers to cuttings, scions, grafts, stolons, bulbs, tubers, corms, keikis, or buds. As illustrated, plant elements are found in plants, or in plant organs, tissue cultures, or cell cultures. obtain.
[0471] "Offspring" includes subsequent generations of the plant.
[0472] As used herein, the term "plant part" refers to plant cells, plant protoplasts, and plant cells. Plant cell tissue cultures, plant callus, plant masses, and plants or plant parts that can regenerate intact embryo, pollen, ovule, seed, leaf, flower, branch, fruit, grain, ear, rachis, pod, stalk, root, root tip, This refers to plant cells such as anthers, and parts thereof themselves. The term "cereal" refers to a plant used for purposes other than the cultivation or propagation of seeds. It is intended to mean mature seeds produced by growers. Descendants of regenerated plants, Bali Ants and mutants are also included in the scope of the present invention if some of them contain introduced polynucleotides. It is included in the enclosure.
[0473] "Monocotyledonous" or "monocotyledonous plants" The term "cot" refers to the "monocotyledonae" class. This refers to a subclass of angiosperms, whose seeds typically consist of a single initial leaf or cotyledon. This term includes only the whole plant, plant elements, plant organs (e.g., leaves, stems, roots, etc.). This includes references to seeds, plant cells, and their offspring.
[0474] The term "of dicotyledonous plants" or "of dicotyledonous plants (dico t) is also known as the "dicotyledonae" class of plants. This refers to a subclass of an organism, whose seeds generally contain only two initial leaves or cotyledons. The words include whole plant, plant elements, plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, and This includes references to their descendants.
[0475] As used herein, "male-sterile plant" means a plant that is viable or otherwise capable of fertilization. These are plants that do not produce viable male gametes. In this specification, "female-sterile plants" refers to plants that do not produce viable male gametes. A plant that does not produce viable, or otherwise fertilizable, female gametes. It is recognized that both the organism and the female-sterile plant can be female-fertile and male-twisted, respectively. Furthermore, male-fertile (other than female-fertile) plants, when hybridized with female-fertile plants, can survive. Female-fertile plants (other than male-fertile plants) can produce offspring, and female-fertile plants can interbreed with male-fertile plants. When mated, they are known to be able to produce viable offspring.
[0476] In this specification, the term "unconventional yeast" refers to Saccharomyces. ) Yeast species of the genus (for example, S. cerevisiae) also include schizosaccharo This refers to any yeast species that is not of the genus Schizosaccharomyces. . (“Non-Conventional Yeasts in Genetics,B iochemistry and biotechnology:Practical Protocols”, K. Wolf, KDBreunig, G. Barth, Ed. Refer to s., Springer-Verlag, Berlin, Germany, 2003. (see).
[0477] In relation to this disclosure, the terms "crossbred" or "crossbred (cr)" "Ossing" is the process of pollination to produce offspring (i.e., cells, seeds, or plants). This refers to the fusion of gametes. This term is used in the context of sexual hybridization (the fusion of gametes from one plant to another). (Pollination) and self-pollination (self-pollination, i.e., pollen and ovules (endoplasmic reticulum and megaspores) of the same plant) This includes both (or plants that are genetically identical).
[0478] The term "gene transfer" refers to the desired transfer of a gene locus from one genetic background to another. This refers to the transfer of a rell. For example, the transfer of a desired allele to a specific locus involves two steps. It can be transmitted to at least one offspring plant through sexual cross-pollination between the parent plants. However, in that case, at least one parent plant has the desired allele in its genome. Alternatively, for example, allele transmission could involve, for instance, two donor genomes in a fused protoplast. This can be done by recombination between them, but in that case, at least one of the donor protocols Last has the desired allele in its genome. The desired allele is, for example, a transgene. Selection of modified (mutated or edited) natural alleles, markers, or QTLs. It could be a specific allele.
[0479] The term "isogenetic lineage" is a comparative term, referring to individuals who are genetically identical but have undergone different treatments. It is a reference organism. For example, two genetically identical maize plant embryos are used, one of which is a modified organism. After undergoing the necessary procedures (such as the introduction of CRISPR-Cas effector endonucleases), another 1 One can be divided into two different groups of controls that do not undergo such treatment. Thus, the phenotypic differences between the two groups are attributable solely to the treatment, and not to the endogenous genes of the plants. This can sometimes be unrelated to any inherent characteristics of the structure.
[0480] "Introduction" is a method by which an ingredient enters the inside of a biological cell or the cell itself, or Polynucleotides, polypeptides, or polynucleotide-proteins are applied to targets such as living organisms. It is intended to mean providing a qualitative complex.
[0481] "Target polynucleotides" improve the desirability of crops, i.e., the traits of agricultural benefit. Contains a nucleotide sequence that encodes a protein or polypeptide. For Chido, important traits from an agricultural economic perspective include herbicide resistance, insecticide resistance, disease resistance, and nematode resistance. Sex, herbicide resistance, microbial resistance, fungal resistance, viral resistance, fertility or sterility, cereal characteristics, city Key traits for products, phenotypic markers, or other traits of agricultural or commercial importance Examples of polynucleotides encoding the target polynucleotide include, but are not limited to, these. Otid can also be used in sense or antisense orientation. Furthermore, for two or more purposes Polynucleotides are used together, or "stacked," to provide further benefits. It is possible.
[0482] A "complex trait locus" is a gene that has multiple transgenes that are genetically linked to each other. It contains the gene locus of Mu.
[0483] The compositions and methods described herein provide plants with "agricultural traits" or "agriculturally important forms." It may provide improvements in "quality" or "traits of agricultural benefit," and such traits include the following: This may include, but is not limited to, the methods or compositions described herein, and does not include any modifications made by such methods or compositions. Compared to plants of isogenic lineage, it exhibits disease resistance, drought tolerance, heat tolerance, cold tolerance, salt tolerance, and Salt tolerance, metal tolerance, herbicide tolerance, improved water use efficiency, improved nitrogen utilization, improved nitrogen fixation, harmful effects Insect resistance, herbivore resistance, pathogen resistance, improved yield, enhanced health, improved vitality, improved growth, photosynthesis Improvement of performance, nutritional enhancement, modification of protein content, modification of oil content, increase of biomass Increased root length, improved root structure, regulation of metabolites, regulation of proteome, increased seed weight, Modification of seed carbohydrate composition, modification of seed oil composition, modification of seed protein composition, seed nutrition composition Modification of the original.
[0484] "Agricultural trait potential" refers to the phenotype, preference, and characteristics at a certain point in the life cycle. Alternatively, it may exhibit improved agricultural traits, or the same plant may exhibit the aforementioned phenotype in another related plant element. It is intended to mean the ability of plant elements to transmit information.
[0485] As used herein, “decreased,” “less,” “slower,” and “increased” are used to mean “increased.” The terms "greater," "faster," "enhanced," and "larger" have not been altered. The characteristics of the modified plant element or the resulting plant compared to the original plant element or the resulting plant. This refers to a decrease or increase. For example, a decrease in a characteristic is at least 1%, at least 2%, or less. 3%, at least 4%, at least 5%, 5% to 10%, at least 10%, 10% ~20%, at least 15%, at least 20%, 20%~30%, at least 25%, At least 30%, 30%-40%, at least 35%, at least 40%, 40%-5 0%, at least 45%, at least 50%, 50%~60%, at least about 60%, 6 0%~70%, 70%~80%, at least 75%, at least about 80%, 80%~90 %, at least about 90%, 90% to 100%, at least 100%, 100% to 200% , at least 200%, at least about 300%, at least about 400%) or more, It may be lower than the untreated control, and the increase in the property may be at least 1%, at least 2%, or less. 3%, at least 4%, at least 5%, 5% to 10%, at least 10%, 10% ~20%, at least 15%, at least 20%, 20%~30%, at least 25%, At least 30%, 30%-40%, at least 35%, at least 40%, 40%-5 0%, at least 45%, at least 50%, 50%~60%, at least about 60%, 6 0%~70%, 70%~80%, at least 75%, at least about 80%, 80%~90 %, at least about 90%, 90% to 100%, at least 100%, 100% to 200% , at least 200%, at least about 300%, at least about 400% or more, un It may be higher than the target of the treatment.
[0486] As used herein, the term “before” means, with respect to the arrangement position, one arrangement, another arrangement. This refers to the upstream or 5' side of the sequence.
[0487] The meanings of the abbreviations are as follows: "sec" means seconds, and "min" means minutes. "h" means time, "d" means day, and "μL" means microliter, mL stands for milliliter, L stands for liter, and μM stands for micromoles. "M" means "of", and "mM" means "of millimoles", and "M" means "of moles", and "m "Mol" means millimoles, while "μmole" or "umole" means micromoles. "g" means gram, "μg" or "ug" means microgram, and "ng" means microgram. " means nanogram, "U" means unit, "bp" means base pair, and "kb" This means kilobase.
[0488] Classification of CRISPR-Cas systems CRISPR-Cas systems are classified according to the sequence and structural analysis of their components. It has a multi-subunit effector complex (including Type I, Type III, and Type IV). Class 1 systems and single protein effectors (including type II, type V, and type VI). Multiple CRISPR / Cas systems are described, including a Class 2 system having ). (Makarova et al. 2015, Nature Reviews Mi) crobiology Vol.13:1-15;Zetsche et al.,20 15,Cell 163,1-13;Shmakov et al.,2015,Mol ecular Cell 60,1-13;Haft et al.,2005,Com putational Biology,PLoS Comput Biol 1(6) :e60; and Koonin et al. 2017, Curr Opinion Mi Crobiology 37:67-78).
[0489] The CRISPR-Cas system uses at least CRISPR RNA (crRNA). The molecule contains at least one CRISPR-related (Cas) protein and crRNA ribo It forms a nuclear protein (crRNP) effector complex. CRISPR-Cas genetics A columb is a single location where DNA targeting spacers encoding crRNA components are scattered. An array consisting of repeats and an operon-like unit of the Cas gene encoding the Cas protein component. It contains nitrates. The resulting ribonucleoprotein complex is polynucleotide by sequence-specific methods. Recognizing Tide (Jore et al., Nature Structural & Molecular Biology 18,529-536(2011)). crRN A moves the non-complementary strand while forming a base pair with the complementary DNA strand, creating what is known as an R loop. By forming this, the effector (protein or complex) is placed on the double-stranded DNA sequence. It functions as a guide RNA for column-specific binding. (Jore et al., 201 1.Nature Structural & Molecular Biology 18,529-536).
[0490] RNA transcripts (pre-crRNA) from the CRISPR locus are type I and type III. By CRISPR-related (Cas) endoribonuclease in the stem, or II In the type system, RNase III specifically cleaves the repeat sequence. The number of CRISPR-related genes at a given CRISPR locus varies by species. It can change.
[0491] Different Cas genes encoding proteins with different domains, different CRI It exists in the SPR system. The CAS operon is one or more effector end nuclei. It contains genes encoding -ase and other Cas proteins. Makarova et al.2011, Nat Rev Microbiol. 2011 9(6):467-477;Makarova et al.2015,Na Nature Reviews Microbiology Vol.13:1-15; and Koonin et al.2017,Current Opinion Microb (Includes those described in iology 37:67-78). This refers to those involved in expression (pre-crRNA processing, e.g., Cas 6 or R Nase III), substances involved in buffering (crRNA and effectors for target binding) - Modules, and domains for targeted cleavage, involved in adaptation The components (spacer insertion, e.g., Cas1 or Cas2), and the components involved in the support (adjustment). It includes sections, helpers, or unknown functions. Some domains have two or more purposes. It can function in this way, and among them, for example, Cas9 has a domain and target for endonuclease function. Includes domains for target cutting.
[0492] Cas endonucleases directly form RNA-DNA base pairs, resulting in a single CR Induced by ISPR RNA (crRNA), the protospacer adjacent motif (PAM) ) recognizes DNA target sites adjacent to (Jore, MM et al., 2011, Nat.Struct.Mol.Biol.18:529-536, Westra,E. R. et al., 2012, Molecular Cell 46:595-605, and Sinkunas, T. et al., 2013, EMBO J.32:385-3 94).
[0493] Class I CRISPR-Cas system Class I CRISPR-Cas systems include Type I, Type III, and Type IV. A unique feature of the LASS I system is that instead of a single protein, it uses effector end nuclei. The presence of an ase complex. The cascade complex is an RNA recognition motif (RRM). ) and diverse RAMP (Repeat-Related Mysterious Protein) protein superfa Contains the nucleic acid binding domain, which is the core fold of Millie (Makarova et al. l.2013,Biochem Soc Trans 41,1392-1400;Ma karova et al.2015,Nature Reviews Microbi (ology Vol.13:1-15). The RAMP protein subunit is (crR NA-effector complex includes Cas5 and Cas7 (including the framework of the NA effector complex) (where Cas The 5 subunit binds to the 5' handle of crRNA and interacts with the larger subunit. They also often loosely associate with effector complexes and are generally pre-crRNA progenitor Includes Cas6, which functions as a repeat-specific RNase in securing (Charp entier et al.,FEMS Microbiol Rev 2015,39 :428-441;Niewoehner et al.,RNA 2016,22:3 18-329).
[0494] The Type I CRISPR-Cas system includes at least Cas5 and Cas7, Effector tranquilizer called CAD (CRISPR-associated complex for antiviral defense) It contains a protein complex. The effector complex is a single CRISPR RNA (crRN). A) Works with Cas3 to defend against invading viral DNA (Brouns ,SJJet al.Science 321:960-964;Makarov a et al.2015,Nature Reviews Microbiology Vol.13:1-15). The type I CRISPR-Cas gene locus is double-stranded DNA (d sDNA) and single-stranded DNA (ss) have been shown to have the ability to unwind RNA-DNA double strands. Encodes metal-dependent nucleases with DNA-stimulated superfamily 2 helicases. Includes the signature gene cas3 (or variant cas3' or cas3) Makarova et al.2015, Nature Reviews Micro Biology Vol.13:1-15). Following target recognition, Cas3 endonucleation The ase is recruited to the cascade-crRNA-target DNA complex to cleave the DNA target and Decompose (Westra, E et al. (2012) Molecular Ce ll 46:595-605, Sinkunas, T. et al. (2011) EMB O J.30:1335-1342, and Sinkunas, T. et al. (201 3) EMBO J.32:385-394). In some type I systems, Cas6 is It may be an active endonuclease involved in crRNA processing, and Cas5 and Ca s7 functions as a non-catalytic RNA-binding protein, but in the type I-C system, crR NA processing can be catalyzed by Cas5 (Makarova et al. 2) 015,Nature Reviews Microbiology Vol.13:1 -15). Type I systems can be divided into seven subtypes (Makarova et al.). al. 2011, Nat Rev Microbiol.2011 9(6):467 -477;Koonin et al.2017,Curr Opinion Micr (obiology 37:67-78). At least the protein subunit Cas7, Modified type I CRISPR-related complex for adaptive antiviral protection including Cas5 and Cas6 It is a cascade, and one of these subunits is Cas3 endonuclease. Alternatively, a compound in which the modified restricted endonuclease FokI is synthetically fused is described. (From the international publication no. 2013 / 098244, released on July 4, 2013) to).
[0495] The type III CRISPR-Cas system, which includes multiple Cas7 genes, uses ssRNA or It targets ssDNA and either RNase or target RNA-activating DNA nuclease. It functions as (Tamulaitis et al., Trends in Mic robiology 25(10)49-61,2017). Csm(III-A type) and The Cmr(III-B) complex binds / cleaves target RNA and links it to ssDNA degradation. It functions as an RNA-activated single-stranded (ss) DNase. When it infects with foreign DNA, C Novel transcripts of RISPR RNA (crRNA)-induced Csm or Cmr complexes Binding to Cas10 recruits Cas10 DNase to actively transcribed phage DNA. As a result, both the transcript and phage DNA are degraded, but the host DNA is not. No. The Cas10 HD domain is involved in ssDNase activity and Csm3 / Cmr4 The b unit is involved in the endoribonuclease activity of the Csm / Cmr complex. Target RN The 3' flanking sequence of A is important for ssDNase activity of Csm / Cmr:cr Base pairing with the 5' handle of RNA protects host DNA from degradation.
[0496] Type IV systems include typical Type I CAS5 and CAS7 domains in addition to CAS8-like domains. This includes the main feature, but is a characteristic of most other CRISPR-Cas systems: CRISP The R array may be missing.
[0497] Class II CRISPR-Cas system Class II CRISPR-Cas systems include Type II, Type V, and Type VI. A unique feature of the LASS II system is that it uses a single Cas effect pedal instead of a complex effects pedal. - The presence of the protein. Type II and type V Cas proteins are RNase It contains a RuvC endonuclease domain that employs an H-fold.
[0498] The type II CRISPR / Cas system induces Cas endonucleases to target DNA. To facilitate this, crRNA and tracrRNA (trans-activated CRISPR RNA) ) is used. crRNA has a spacer region complementary to one strand of the double-stranded DNA target, It base pairs with tracrRNA (trans-activated CRISPR RNA) and Casen It forms an RNA double strand that allows donuclease to cleave the DNA target, leaving a region that retains a blunt end. It includes. The spacer is associated with the Cas1 and Cas2 proteins, and is not well understood. Obtained by the process. The type II CRISPR / Cas locus is generally known as cas. In addition to the 9 genes, it also includes the cas1 and cas2 genes (Chylinski et al. .,2013,RNA Biology 10:726-737;Makarova e t al.2015,Nature Reviews Microbiology Vo l.13:1-15). The type II CRISPR-Cas locus is located at each CRISPR It can encode tracrRNA that is partially complementary to the repeats in the array, Cs It can contain other proteins such as n1 and Csn2. cas1 and cas2 genes The presence of Cas9 in the vicinity of the offspring is a characteristic of the type II locus (Makarova et al.2015,Nature Reviews Microbiology V (ol.13:1-15).
[0499] The V-type CRISPR / Cas system is a single Cas system including Cpf1 (Cas12). Contains endonuclease (Koonin et al., Curr Opinion M (Icrobiology 37:67-78, 2017), this is different from Cas9. Furthermore, additional trans-activated CRISPR(tracr)RNA is required for targeted cleavage. It is an active RNA-induced endonuclease that does not require any additional processing.
[0500] The VI-type CRISPR-Cas system is a two-HEPN (higher eukaryotes and prokaryotes) system. It has a nucleotide-binding domain, but it does not have an HNH domain or a RuvC domain. Contains the cas13 gene, which is independent of tracrRNA activity. Most of the HEPN domain. It contains a conserved motif that constitutes the metal-independent endoRNase active site (Anant Haram et al., Biol Direct 8:15, 2013). This characteristic Therefore, the Type VI system has a DNA target that is not common to other CRISPR-Cas systems. It is thought to act on RNA targets to a certain extent.
[0501] Novel CRISPR-Cas system This specification describes a novel CRISPR-Cas system, its components, and a method of using said components. This is disclosed in [the document]. This system uses a novel Cas effector protein, Cas-A Includes rufa.
[0502] The novel CRISPR-Cas system components described herein are different Cas system components. One or more subunits derived from Tem, or two or more different bacteria or archaeoprokaryotes. Subunits derived from or modified from there, and / or synthesized or manipulated It may contain the ingredients listed.
[0503] This newly identified CRISPR-Cas system, including a novel sequence of the cas gene, is based on... This will be described in the specification. Furthermore, novel Cas genes and proteins will be described.
[0504] One of the features of the novel Cas-alpha system is shown in Figures 1A-1D. This is the locus structure. In some aspects, the Cas-alpha genome locus is Cas1 gene, Cas2 gene, Cas4 gene, and effector protein Cas -Contains the cas-alpha gene which codes for alpha. The CRISPR array contains the genes encoding Cas-alpha endonuclease. It may be found before or after. In some embodiments, the cas-alpha locus is the effect The cas-alpha gene encoding the tar protein, and CRISPR containing repeats. Includes array, but any of the cas1 gene, cas2 gene, and / or cas4 gene It may not contain one or more of these.
[0505] CRISPR-Cas system components Cas protein Adaptation (spacer insertion), interference (effector module target coupling, target) Nicking or cleavage (e.g., endonuclease activity), expression (pre-crRNA plutonium Numerous proteins, including those involved in reduction, regulation, and other processes, are classified as CRIS. This can be coded into a PR cas operon.
[0506] The two proteins, Cas1 and Cas2, are conserved among many CRISPR systems. (For example, Koonin et al., Curr Opinion Mic (described in Robiology 37:67-78, 2017). Cas1 is two It is a metal-dependent DNA-specific endonuclease that produces main-strand DNA fragments. In this system, Cas1 forms a stable complex with Cas2, which is CRISPR cis Essential for obtaining and inserting spacers for the system (Nunez et al., Nat ure Str Mol Biol 21:528-534,2014).
[0507] Many other tans containing Cas4 (which may have similarities to RecB nuclease) The protein has been identified in various systems, and is a new one for integration into CRISPR arrays. It is thought to play a role in capturing viral DNA sequences (Zhang et al.). ,PLOS One 7(10):e47232,2012).
[0508] Some proteins can encompass multiple functions. For example, class 2 type II systemic proteins... Cas9, the signature protein of the genus, is involved in pre-crRNA processing and target binding. It has been demonstrated that it is involved in target cleavage and other related processes.
[0509] The novel Cas-alpha proteins disclosed herein are effector proteins Contains proteins (endonucleases) and adaptation proteins. Cas Endonuclease Aase has been identified from several bacterial and archaeal sources, as shown in Figures 7A-7K. These are some examples.
[0510] Cas Endonuclease and Effects Endonucleases are enzymes that cleave phosphate diester bonds in polynucleotide chains. There are restriction endonucleases that cut DNA at specific sites without damaging the bases. Includes. Examples of endonucleases include restriction endonucleases, meganucleases, TAL effector nuclease (TALEN), zinc finger nuclease, and One example is the CAS (CRISPR-related) effector end nuclease.
[0511] Cas endonuclease can be used as a single effector protein or in combination with other components. It forms an effector complex, unwinds the DNA double strand at the target sequence, and the Cas effector Polynucleotides that form complexes with proteins (not limited to the following, but including crRN) When mediated by the recognition of the target sequence by A or guide RNA (such as), it is selectively reduced. At the very least, it cuts one DNA strand. Generally, Cas endonucleases cut the target sequence Such recognition and cutting are precise protospacer adjacent motifs (PAMs) in DNA. If located at the 3' end of the target sequence, or adjacent to the 3' end of the DNA target sequence, it will occur. Alternatively, the Cas endonucleases described herein, when combined with a suitable RNA component, In combination, it may lack DNA cleavage activity or nicking activity, but it still targets the DNA sequence. It can then bind specifically. (U.S. Patent Application Publication No. 1, published March 19, 2015) Specification No. 2015 / 0082478 and U.S. Patent Application No. 2 published on February 26, 2015. (See also Specification No. 015 / 0059010).
[0512] Cas endonuclease is used in individual effects pedals (Class 2 CRISPR systems). ) as, or as part of a larger effects unit complex (Class I CRISPR system). It can occur as a part.
[0513] Examples of Cas endonucleases mentioned include Cas3 (Class I I). Characteristics of Class II systems), Cas9 (characteristics of Class II systems), and Cas12 (C Examples include, but are not limited to, pf1) (characteristics of Class 2 V systems).
[0514] Cas3 (and its variants Cas3' and Cas3") are single-stranded DNA nuclei. It functions as an enzyme (HD domain) and an ATP-dependent helicase. Cas3 endonucleus The rease variant is a derivative of one or both of the Cas3 endonuclease polypeptides. This can be achieved by disabling the main functional activity. ATPase-dependent Licase activity (by deletion or knockout of the Cas3-helicase domain, or important This can be achieved by mutagenesis of certain residues, or, as previously described, in the absence of ATP. By constructing the response (Sinkunas, T, et al., 2013, EMBO (J.32:385-394) By disabling it, the modified Cas3 endonuclease is included. The cut-ready cascade can be converted to nickase (the HD domain is still (Because it functions as such). Inactivation of HD endonuclease activity is not limited to the following, but This can be achieved by methods known in the art, such as mutagenesis of key residues in the HD domain. It is possible to cut a ready cascade containing a modified Cas3 endonuclease into a helical cascade. It can be converted to -ase. Cas helicase and Cas3 HD endonuclease Inactivation of both activity is important for both the helicase and HD domains, but is not limited to the following. This can be achieved by methods known in the art, such as mutagenesis of certain residues, and modified C The cleavage-ready cascade containing the as3 endonuclease is bound to the target sequence by a binding agent. It can be converted into protein.
[0515] "Cas9" (formerly called Cas5, Csn1, or Csx12) is a cr nucleus Rheotide and tracr nucleotides, or complexes with single guide polynucleotides. Cas endogenucleates which form a Cas endogenucleus that specifically recognizes and cleaves all or part of the DNA target sequence. It is a rease. Cas9 recognizes the 3' GC-rich PAM sequence of the target dsDNA. The Cas9 protein has an HNH(HNH) nucleus adjacent to the RuvC-II domain. Contains RuvC nuclease along with ase. RuvC nuclease and HNH nuclease Each of the zes can cleave a single DNA strand at its target sequence (coordination of both domains). The activity results in DNA double-strand breaks, while the activity of one domain produces a nick. Generally, the RuvC domain includes subdomains I, II, and III, and the domain Subdomain I is located near the N-terminus of Cas9, and subdomains II and III are part of the protein It is located in the center and adjacent to the HNH domain (Hsu et al, Cell 157). (1262-1278). Cas9 endonucleases typically contain at least one poly(1 / 2) DNA cleavage using Cas9 endonuclease, which forms a complex with nucleotide components. It originates from Type II CRISPR systems, including the CRISPR system. For example, Cas9 is a CRISPR system. PR RNA (crRNA) and transactivated CRISPR RNA (tracrRN) A) can be combined with another example. Cas9 can be combined with a single guide RNA. It can be synthesized (Makarova et al. 2015, Nature R eviews Microbiology Vol.13:1-15).
[0516] Cas12 (formerly Cpf1, and variants c2c1, c2c3, CasX, and (formerly known as CasY) contains a RuvC nuclease domain and targets dsDNA. A 5' overhang was generated. Some variants differ from Cas9 functionality, tr It does not require acrRNA. Cas12 and its variants are located on the target dsDNA. 'Recognizes AT-rich PAM sequences. The insertion called Nuc in the Cas12a protein...' The main role has been demonstrated in target chain cleavage (Yamano et al., Cell 2016, 165:949-962). Other Cas12 proteins Research into these mutations, along with the RuvC domain which is responsible for cleavage in the Nuc domain, has shown that induction and target mutations are important. It was demonstrated that it contributes to target binding (Swarts et al., Mol Cell 2 017,66:221-233 e224).
[0517] Cas endonucleases and effector proteins are used in targeted genome editing (single and (Through multiple double-strand breaks and nicks), and targeted genome regulation (Cas protein or sg (By binding the epigenetic effector domain to RNA) This can be done. Cas endonuclease also functions as an RNA-inducing recombinase. It can be manipulated to allow the assembly of multiple proteins and nucleic acid complexes via RNA tethers. - Can function as a scaffold (Mali et al., 2013, Nature Met hods Vol.10:957-963).
[0518] Cas-alphaendonuclease Cas-alpha endonuclease is a functional RNA consisting of fewer than 800 amino acids. It is defined as an induced PAM-dependent dsDNA cleavage protein and is divided into three subdomains. Furthermore, the C-terminal Ru includes a bridge helix and one or more zinc finger motifs. vC catalytic domain; and helical bundle, WED wedge (or "oligonucleotide bond"). The domains include a "combined domain," an OBD domain, and optionally a zinc finger motif. Includes the N-terminal Rec subunit.
[0519] When Cas-alpha-endonuclease is aligned with SEQ ID NO: 17, it corresponds to SEQ ID NO: 1 For each of the 7 amino acid position numbers, at least 1, at least 2, at least 3, and at least 3 of the following Including at least 4, 5, 6, or 7: Glycine (G) at position 337, 3 Glycine (G) is ranked 41st, glutamic acid (E) is ranked 430th, and leucine (L) is ranked 432nd. Cysteine (C) ranked 487th, cysteine (C) ranked 490th, cysteine (C) ranked 507th ), and / or cysteine (C) or histidine (H) at position 512. Cas-alph Aendonuclease is, Next motif: GxxxG, ExL, Cx n C, Cx n (C or H) (where n=1) (The above amino acids).
[0520] The RuvC domain has been demonstrated in the literature to encompass the functionality of an endonuclease. Yes. Cas-alpha-endonuclease encodes an effector protein. Isolation or identification from a locus containing an as-alpha gene and an array of multiple repeats. It is possible. In some embodiments, the cas-alpha locus is a part of the cas1 gene. Or all of them may further include the Cas2 gene and / or the Cas4 gene.
[0521] The zinc finger motif typically contains one or more zinc ions, usually cysteine and histite. This is a domain that coordinates with the din side chain and stabilizes those folds. The Kufinger gene is formed by a pattern of cysteine and histidine residues that coordinate to the zinc ion. (For example, C4 means that the zinc ion is coordinated to four cysteine residues.) In C3H, the zinc ion is coordinated to three cysteine residues and one histidine residue. (This means that...)
[0522] The Cas-alpha protein contains one or more zinc fins that can form a zinc-binding domain. Contains a gar (ZFN) coordination motif. The zinc finger-like motif is present in both the target and non-target chains. It can assist in the isolation and loading of guide RNA to the DNA target. One or more Cas-alpha proteins containing the zinc finger motif target polynucleotides This may provide further stability to the ribonucleoprotein complex on the cas-alpha protein. The substance contains a C4 or C3H zinc-binding domain.
[0523] Several Cas-alpha proteins and polynucleotides are shown in Figures 7A-7K. The important structural motifs of endonuclease proteins are shown in Figures 8A to 8K.
[0524] Cas-alpha endonuclease (1) is mediated by the nucleotide sequence of the guide RNA. (2) Binding to and cleaving double-stranded DNA targets containing same-sex sequences and PAM sequences. It is an RNA-induced endonuclease that can produce T-cells. In some embodiments, PAM can produce T-cells. It is C-rich. In some embodiments, PAM is C-rich.
[0525] Cas-alpha-endonuclease functions as a double-strand break inducer, and also Nikka -ase, or single-strand break inducer. In some embodiments, catalytically inert Cas - Alpha-endonucleases target or transfer to target DNA sequences. It can be used for induction, but does not induce cleavage. In some embodiments, it catalytically inactivates. The Cas-alpha protein is co-administered with a functional endonuclease that cleaves the target sequence. It can be used for the following. In some embodiments, catalytically inactive Cas-alpha protein Base editing molecules such as deaminase can be combined. In some embodiments, deaminase The enzyme can be cytidine deaminase. In some embodiments, the deaminase is adenine It may be a deaminase. In some embodiments, the deaminase is ADAR-2. obtain.
[0526] Cas-alpha-endonuclease is further represented by SEQ ID NOs. 17, 18, 19, 20, 3 2, 33, 34, 35, 36, 37, 38, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, The apparatus that retains any or at least some of the activity of 369, 370, and 371 At least 50, 50-100, or at least 100 functional fragments or functional variants , 100-150, at least 150, 150-200, at least 200, 200-2 50, at least 250, 250-300, at least 300, 300-350, less 350, 350-400, at least 400, 400-450, at least 500, Or more than 500 consecutive amino acids and at least 50%, 50%-55%, and at least 5 5%, 55%-60%, at least 60%, 60%-65%, at least 65%, 65% ~70%, at least 70%, 70%~75%, at least 75%, 75%~80%, little At least 80%, 80% to 85%, at least 85%, 85% to 90%, at least 90 %, at least 90%~95%, at least 95%, 95%~96%, at least 96% 96%~97%, at least 97%, 97%~98%, at least 98%, 98%~9 RNA with 9%, at least 99%, 99% to 100%, or 100% sequence identity It is defined as an inducible double-strand DNA break protein. Cas-alpha endonuclease The "functional fragment" recognizes or binds to one strand of a double-stranded polynucleotide, if The ability to nicking, or the ability to cleave both strands of a double-stranded polynucleotide, or It holds any combination of them.
[0527] Cas-alpha-endonuclease is represented by SEQ ID NOs: 13, 14, 15, 16, 25, 2 At least 50, 50-100 of any of 6, 27, 28, 29, 30, or 31, At least 100, 100-150, at least 150, 150-200, at least 2 00, 200-250, at least 250, 250-300, at least 300, 300 ~350, at least 350, 350~400, at least 400, 400~450, little At least 500, 500-550, at least 600, 600-650, at least 65 0, 650-700, at least 700, 700-750, at least 750, 750- 800, at least 800, 800-850, at least 850, 850-900, less At least 900, 900-950, at least 950, 950-1000, at least 10 00 or more consecutive nucleotides, and at least 50%, 50% to 5 5%, at least 55%, 55% to 60%, at least 60%, 60% to 65%, less 65%, 65%-70%, at least 70%, 70%-75%, at least 75%, 75%~80%, at least 80%, 80%~85%, at least 85%, 85%~90 %, at least 90%, 90%~95%, at least 95%, 95%~96%, at least 96%, 96%~97%, at least 97%, 97%~98%, at least 98%, 9 8% to 99%, at least 99%, 99% to 100%, or 100% sequence identity Polynucleotides containing, or SEQ ID NOs: 17, 18, 19, 20, 32, 33, 34, 3 5, 36, 37, 38, 254, 255, 256, 257, 258, 259, 260, 2 61, 262, 263, 264, 265, 266, 267, 268, 269, 270, 2 71, 272, 273, 274, 275, 276, 277, 278, 279, 280, 2 81, 282, 283, 284, 285, 286, 287, 288, 289, 290, 2 91, 292, 293, 294, 295, 296, 297, 298, 299, 300, 3 01, 302, 303, 304, 305, 306, 307, 308, 309, 310, 3 11, 312, 313, 314, 315, 316, 317, 318, 319, 320, 3 21, 322, 323, 324, 325, 326, 327, 328, 329, 330, 3 31, 332, 333, 334, 335, 336, 337, 338, 339, 340, 3 41, 342, 343, 344, 345, 346, 347, 348, 349, 350, 3 51, 352, 353, 354, 355, 356, 357, 358, 359, 360, 3 61, 362, 363, 364, 365, 366, 367, 368, 369, 370 and It can be encoded by a polynucleotide that encodes any one of 371.
[0528] Cas endonuclease, effector protein, and for use in the method of disclosure. Its functional fragments are obtained from natural sources, or from genetically modified host cells that produce the protein. It can be isolated from recombinant sources that have been modified to express the encoding nucleic acid sequence. It is possible. Alternatively, the Cas protein can be produced using a cell-free protein expression system. Alternatively, it can be produced synthetically. The effector Cas nuclease can be isolated and differentiated. It may be introduced into seed cells, and due to its natural form, it may have a different form or activity than its natural source. It may be modified to indicate the size. Such modifications include fragments, variants, and placements. Examples include, but are not limited to, replacement, deletion, and insertion.
[0529] Cas endonuclease and Cas effector protein fragments and variants are Endonucle Methods for measuring ase activity are well known in the art, for example, in 2013. International publication pamphlet No. 2013 / 166113, released on January 7, 2016 International publication pamphlet No. 2016 / 186953, released on January 24, and 201 The brochure International Publication No. 2016 / 186946, released on November 24, 2016, is cited as an example. These are possible, but are not limited to these.
[0530] Cas endonucleases may include modified forms of Cas polypeptides. Modified forms of as polypeptides include naturally occurring nucleases of Cas proteins. Examples of amino acid changes that reduce activity (e.g., deletion, insertion, or substitution) include those that reduce activity. For example, in some cases, the modified form of the Cas protein corresponds to the wild-type Cas protein. Less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5% of the peptide It has less than 1% nuclease activity (US patent application published March 6, 2014). (Patent Publication No. 2014 / 0068797). In some examples, the Cas polypeptide The modified form is substantially devoid of nuclease activity and is catalytically "inactivated Cas" or It is called "inactivated Cas (dCas)". Inactivated Cas / Inactivated Cas is Contains inactivated Cas endonuclease (dCas). Catalytically inactive Cas efferent Inductor proteins can be fused to heterologous sequences to induce or modify their activity.
[0531] Cas endonucleases are composed of one or more heterologous protein domains (e.g., Cas tan A fusion protein containing one, two, three or more domains in addition to the protein. It may be a part. Such a fusion protein may have any further protein sequences, and In some cases, this can occur between any two domains, for example, between Cas and a first heterogeneous domain. May contain a linker sequence. Proteins that can be fused to Cas proteins in this specification. Examples of tags include epitope tags (e.g., histidine [His], V5, FLAG, Influenza hemagglutinin [HA], myc, VSV-G, thioredoxin [Trx] ), reporter (e.g., glutathione-5-transferase [GST], horseradish) Biperoxidase [HRP], Chloramphenicol acetyltransferase [C AT], beta-galactosidase, beta-glucuronidase [GUS], lucifera -se, green fluorescent protein [GFP], HcRed, DsRed, cyan fluorescent protein [CFP], yellow fluorescent protein [YFP], blue fluorescent protein [BFP]), and Domains having one or more of the following activities, but not limited to them: methylation Enzyme activity, demethylase activity, transcriptional activation activity (e.g., VP16 or VP64), Repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, and nucleic acid binding activity. Cas proteins can also bind to DNA molecules or other molecules, such as maltose-binding proteins. (MBP), S-tag, Lex A DNA-binding domain (DBD), GAL4A DN A-binding domain and protein that binds to herpes simplex virus (HSV) VP16 They can merge.
[0532] Catalytically active and / or inactive Cas endonucleases can fuse with heterologous sequences. (U.S. Patent Application Publication No. 2014 / 0068797, published on March 6, 2014) (Written). Suitable fusion partners include, but are not limited to, target DNA or target D. Directly to polypeptides related to NA (e.g., histones or other DNA-binding proteins) Examples include polypeptides that, through contact, provide activity that indirectly increases transcription. Further preferred fusion partners include, but are not limited to, methyltran. Spherase activity, demethylase activity, acetyltransferase activity, deacetylase activity enzyme activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination Activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, riboshi Provides ribosylation activity, drimosylation activity, myristoylation activity, or demyristoylation activity. Polypeptides are examples. More suitable fusion partners are not limited to the following: However, polypeptides that directly cause increased transcription of target nucleic acids (e.g., transcription activators or similar) Proteins that induce fragments, transcription activators, small molecule / drug-responsive transcription regulators, etc. Examples include fragments of Cas-alphae. Partially active or catalytically inactive Cas-alphae Nucleases also produce double-strand breaks by transferring another protein or domain to it. For example, it can fuse with Clo51 nuclease or FokI nuclease (Guili). Nger et al.Nature biotechnology,volume 3 (2, number 6, June 2014).
[0533] Catalytically active or inactive proteins, such as Cas-alpha proteins, as described herein. Sexual Cas proteins also refer to the editing of one or more bases in a polynucleotide sequence. The identity of the molecule being shown, for example, a nucleotide, is, for example, from C·G to T·A or A·T. It can fuse with site-specific deaminases that can be converted to G·C (Gaudel li et al.,Programmable base editing of A ·T to G·C in genomic DNA without DNA cle avage.”Nature(2017);Nishida et al.”Targe ted nucleotide editing using hybrid prok aryotic and vertebrate adaptive immune system ystems.”Science 353(6305)(2016); al.”Programmable editing of a target ba se in genomic DNA without double-strande d DNA cleavage.”Nature 533(7603)(2016):4 20-4. Base editing fusion proteins are, for example, active (double-strand cleavage), partially active. (Nickase), or inactivated (catalytically inactive) Cas-alphaendonucleus Reases and deaminases (for example, cytidine deaminase, etc., but not limited to the following) Denine deaminase, APOBEC1, APOBEC3A, BE2, BE3, BE4, A May include BE, etc. Base editing repair inhibitors and glycosylase inhibitors (e.g.) For example, how many uracil glycosylase inhibitors (which prevent the removal of uracil) are there? In that embodiment, it is intended as another component of the base editing system.
[0534] The Cas endonucleases described herein are derived from methods known in the art. For example, Pan International Publication No. 2016 / 186953, published on November 24, 2016. It can be expressed and purified by the methods described in the frets.
[0535] To date, it has recognized specific PAM sequences (as published internationally on November 24, 2016). Pamphlet No. 2016 / 186953, published internationally on November 24, 2016. Brochure No. 2016 / 186946, and Zetsche B et al. (2015.Cell 163,1013), it is possible to cleave target DNA at a specific location. Many Cas endonucleases have been reported. Naturally, those skilled in the art will know that Based on the methods and embodiments described herein using a novel inductive Cas system, These methods are suitable for use with any inducible endonuclease system. It is understood that it is possible to combine them.
[0536] Cas effector proteins may contain heterologous nuclear localization sequences (NLS). The heterologous NLS amino acid sequence contained herein is, for example, in an amount detectable in the nucleus of yeast cells as described herein. It can have sufficient strength to drive the accumulation of Cas protein. NLS is basic, positive One (segment) or more (e.g., one) of the charge-bearing residues (e.g., lysine and / or arginine) For example, it may contain a short sequence (e.g., 2-20 residues) of two segments, and the protein If exposed to the surface, it can be positioned anywhere in the Cas amino acid sequence. LS is operably ligated, for example, to the N-terminus or C-terminus of the Cas protein as described herein. It is acceptable. For example, two or more NLS sequences may be present in a Cas protein, for example, a Cas protein It can be linked to the N-terminus and C-terminus of the protein. The Cas endonuclease gene is Cas co Upstream of the SV40 nuclear target signal in the don region, and bifid VirD2 nuclear localization in the Cas codon region. Signal (Tinland et al. (1992) Proc. Natl. Acad.) (Sci.USA 89:7442-6) Can be operably connected downstream. Non-limiting examples of suitable NLS sequences include U.S. Patent No. 6,660,830 and the same. Examples include those disclosed in Specification No. 7,309,576.
[0537] Guide polynucleotides Guide polynucleotides are used for target recognition and binding to the target by Cas endonucleases. It enables the cleavage of targets by combination and optional selection, and can be a single molecule or a double molecule. A polynucleotide sequence is an RNA sequence, a DNA sequence, or a combination thereof (RNA- It may be a DNA combination sequence. Optionally, the guide polynucleotide may be small. at least one nucleotide, phosphodiester bond, or linkage modification, for example, limited to the following: However, locked nucleic acid (LNA), 5-methyl dC, 2,6-diaminopurine, 2'- FluoroA, 2'-FluoroU, 2'-O-methylRNA, phosphorothioate bond, co Linking to sterol molecules, linking to polyethylene glycol molecules, spacer 18 (he Covalent linkage from 5' to 3' to the xaethylene glycol chain molecule, resulting in linkage or cyclization. It may include linking by fusion. Guide polynucleotides containing only ribonucleic acid are "ga Also called "gRNA" or "gRNA" (US Special Classification of RNA published on March 19, 2015) Approved application publication No. 2015 / 0082478 and published on February 26, 2015. (U.S. Patent Application Publication No. 2015 / 0059010). Guide polynucleotides are, They may be created through genetic engineering or by synthesis.
[0538] Guide polynucleotides contain regions not found together in nature, resulting in chimeric non-natural molars. It contains id RNA (i.e., they are heterogeneous). For example, Cas end Target DNA linked to a second nucleotide sequence capable of recognizing a nuclease. The first nucleotide sequence domain that can hybridize to the nucleotide sequence inside Chimeric non-natural gas containing a variable targeting domain (called a VT domain) In id RNA, the first and second nucleotide sequences are naturally linked together. It was searched and not found.
[0539] Guide polynucleotides are cr nucleotide sequences (such as crRNA) and tracr Double molecules containing nucleotide sequences (such as tracrRNA) (double-stranded guide polynucleotides) It can be (also called tide). In some cases, crRNA and tracrRNA are linked. There is a linker polynucleotide that forms a single guide, such as sgRNA.
[0540] cr nucleotides can hybridize to nucleotide sequences in target DNA. The first nucleotide sequence domain (variable targeting domain or VT domain) (called), and the second part of the Cas endonuclease recognition (CER) domain It contains the nucleotide sequence (also called the tracr-mate sequence). The column hybridizes along the complementary region of the tracr nucleotide, and Cas endonucleus They can together form a crease-recognizing domain, i.e., a CER domain. The domain can interact with Cas endonuclease polypeptides. The cr nucleotides and tracr nucleotides of the chain guide polynucleotides are RNA, This may be DNA and / or RNA-DNA combination sequences. In some embodiments, this may be a combination of DNA and / or RNA-DNA sequences. The cr nucleotide molecule of the double-stranded guide polynucleotide is called "crDNA" (DNA nucleotide). (When composed of a continuous stretch of creotide), or "crNNA" (RNA nucleo (When composed of a continuous stretch of DNA), or "crDNA-RNA" (DNA nucleus) This is called a cr nucleus (when it is composed of a combination of an oside and an RNA nucleotide). Otide is naturally present in bacteria and archaea. It may contain fragments of crRNA present in the crnucleotides disclosed herein. It is naturally present in bacteria and archaea that may exist. The sizes of the crRNA fragments are 2, 3, 4, 5, 6, 7, 8, 9, and 1. 0 pieces, 11 pieces, 12 pieces, 13 pieces, 14 pieces, 15 pieces, 16 pieces, 17 pieces, 18 pieces, 19 pieces, 2 It can vary from zero or more nucleotides, but is not limited to these. In some embodiments, the crRNA molecule is selected from the group consisting of SEQ ID NOs: 57, 58, and 59. Selected.
[0541] In some embodiments, the tracr nucleotide sequence is "tracrRNA" (RN (When it consists of a continuous stretch of A nucleotides), or "tracrDNA" (D (When composed of a continuous stretch of NA nucleotides), or "tracrDNA-R NA (when composed of a combination of DNA nucleotides and RNA nucleotides) It is called [name]. In one embodiment, the RNA / Cas9 endonuclease complex is induced by R NA is a double-stranded RNA containing double-stranded crRNA-tracrRNA. NA (trans-activated CRISPR RNA) has (i)CRISPR RNA in the direction from 5' to 3'. The sequence that anneals with the repeat region of PR type II crRNA, and (ii) stem Contains the portion containing the oat (Deltcheva et al., Nature 471:60) 2-607). Double-stranded guide polynucleotides complex with Cas endonuclease. The guide polynucleotide / Cas endonuclease complex can be formed. The guide polynucleotide (also known as the Cas endonuclease system) is Cas Endonucleases can be induced to target genomic sites, and Cas endonucleases It recognizes the target site, binds to the target site, and selectively nicks the target site or This makes it possible to break (introduce single-strand or double-strand breaks). (March 19, 2015) The published U.S. Patent Application Publication No. 2015 / 0082478 and February 2015 (U.S. Patent Application Publication No. 2015 / 0059010, published on the 6th).
[0542] In some embodiments, the tracrRNA molecule is a group consisting of SEQ ID NOs. 60-68. They are selected.
[0543] In one embodiment, the guide polynucleotide forms a PGEN as described herein. The guide polynucleotide is a guide polynucleotide that can be used as a target DNA A first nucleotide sequence domain complementary to the creotide sequence, and the Cas endonucleus It includes a second nucleotide sequence domain that interacts with the ase polypeptide.
[0544] In one embodiment, the guide polynucleotide is a guide polynucleotide as described herein. The first nucleotide sequence and the second nucleotide sequence domain are DNA distributions. Guide polynucleos selected from a group consisting of columns, RNA sequences, and combinations thereof. It's Chido.
[0545] In one embodiment, the guide polynucleotide is a guide polynucleotide as described herein. The first nucleotide sequence and the second nucleotide sequence domain enhance stability. RNA backbone modifications that enhance stability, DNA backbone modifications that enhance stability, and these It is a guide polynucleotide selected from a group of combinations (Kanasty et al.,2013,Common RNA-backbone modifica tions,Nature Materials 12:976-977;20153 The specification of U.S. Patent Application Publication No. 2015 / 0082478, published on the 19th of the month, 2015 Refer to U.S. Patent Application Publication No. 2015 / 0059010, published on February 26. (I want to be treated that way)
[0546] The guide RNA is a chimeric non-natural crR linked to at least one tracrRNA. It contains a bimolecular molecule containing NA. Chimeric non-natural crRNAs are found in regions not naturally occurring together. It contains crRNAs that include the region (i.e., they are heterogeneous). For example, the second nu Nucleotide sequences (also called tracrmate sequences) linked to target DNA A first nucleotide sequence domain that can hybridize to a creotide sequence (possibly) crRNA containing a variable targeting domain (also called a VT domain), the first The first sequence and the second sequence are not found linked together in nature.
[0547] Guide polynucleotides are also cr nucleos ligated to tracr nucleotide sequences. It may be a single molecule containing a tide sequence (also called a single guide polynucleotide). The single guide polynucleotide hybridizes with the nucleotide sequence in the target DNA. The first nucleotide sequence domain that can be modified (variable targeting domain or (This is called the VT domain), and interacts with Cas endonuclease polypeptides. It includes the Cas endonuclease recognition domain (CER domain). Several implementation forms In this state, the sgRNA molecule is selected from the group consisting of sequence numbers 69 to 77.
[0548] The VT domain and / or CER domain of a single guide polynucleotide are RNA It may include sequences, DNA sequences, or RNA-DNA combination sequences. cr nucleo Single guide polynucleotides composed of sequences derived from tide and tracr nucleotides. Tide is "single guide RNA" (composed of a continuous stretch of RNA nucleotides) (if present), or "single guide DNA" (composed of a continuous stretch of DNA nucleotides) (If applicable), or "Single Guide RNA-DNA" (RNA nucleotides and DNA It can be called a single nucleotide (when it is composed of a combination of nucleotides). Idopolynucleotides can form complexes with Cas endonucleases, Guide polynucleotide / Cas endonuclease complex (guide polynucleotide) The Cas endonuclease system (also known as the Cas endonuclease system) is a system that uses Cas endonucleases. It can be guided to the target site, and Cas endonuclease can recognize the target site. , binds to the target site and optionally nicks or cleaves the target site (single-strand or double-strand) (Introduces chain severance) which enables this. (U.S. patent application published March 19, 2015) Publication No. 2015 / 0082478 and the U.S. Patent Publication published on February 26, 2015 (Publication No. 2015 / 0059010).
[0549] Chimeric non-natural single guide RNA (sgRNA) is a region not found together in nature. It contains sgRNAs that include (i.e., they are heterogeneous). For example, in nature there is one The second nucleotide sequence that is not found to be linked together (also called the tracr-mate sequence) It can hybridize to a nucleotide sequence in target DNA that is linked to it. The first nucleotide sequence domain (called the variable targeting domain or VT domain) sgRNA containing (which can be detected).
[0550] Linking cr nucleotides and tracr nucleotides of single-stranded guide polynucleotides. The nucleotide sequence is an RNA sequence, a DNA sequence, or an RNA-DNA combination sequence. It may include. In one embodiment, a single guide polynucleotide cr nucleotide A nucleotide sequence (also called a "loop") that links do and tracr nucleotides. at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 1 6, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 ,30,31,32,33,34,35,36,37,38,39,40,41,42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 5 6, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 , 70, 71, 72, 73, 74, 75, 76, 77, 78, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 9 The nucleotide length may be 5, 96, 97, 98, 99, or 100. In another embodiment, This involves using single guide polynucleotides, specifically Cr nucleotides and TraCr nucleotides. The nucleotide sequences to be linked are not limited to the following, but include GAAA tetraloop sequences, etc. It may contain a tetraloop arrangement.
[0551] Guide polynucleotides are chemically synthesized (limited to the following): However, Hendel et al. 2015, Nature Biotechnolo In vitro generation of guide polynucleotides (such as gy33, 985-989), and / or self-splicing of guide RNA (but not limited to Xie et al.) This includes the relevant technical fields (e.g., ., (2015), PNAS 112:3570-3575). It can be manufactured using the known method.
[0552] Protospacer Adjacent Motif (PAM) In this specification, “protospacer adjacent motif” (PAM) refers to a guide polynucleotide. Target sequences recognized (targeted) by the Cas endonuclease system (protosperm This refers to a short nucleotide sequence adjacent to the target DN. Cas endonucleases target DN If there is no PAM sequence after the A sequence, the target DNA sequence cannot be correctly recognized. i. The sequence and length of PAM in this specification are determined by the Cas protein or Cas protein used. It can vary depending on the protein complex. The PAM sequence can be of any length, but generally, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 1 It is 7, 18, 19, or 20 nucleotides long.
[0553] "Randomized PAM" and "Randomized Protospacer Adjacent Motif" are used herein. They are used interchangeably and by the guide polynucleotide / Cas endonuclease system Random DNA adjacent to the target sequence (protospacer) that is recognized (targeted) as such. This refers to a sequence. Randomized PAM sequences can be of any length, but are generally 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 1 It is 9 or 20 nucleotides long. Randomized nucleotides are nucleotide A, Includes either C, G, or T.
[0554] Guide polynucleotide / Cas endonuclease complex The guide polynucleotide / Cas endonuclease complex described herein is a target Recognize all or part of the sequence, join them, and optionally, nicking, untying, or It can be cut.
[0555] Guide polynucleotides / CasE can cleave both strands of a DNA target sequence. The endonuclease complex typically has all of its endonuclease domains in a functional state. Cas proteins possessing (e.g., wild-type endonuclease domain, or each endonuclease) This includes variants that retain some or all of the activity of the clearase domain. Therefore, wild-type Cas protein (for example, the Cas protein disclosed herein) (Quality), or maintaining the activity of some or all of the endonuclease domains of the Cas protein. The variant possessed by Cas can cleave both strands of the DNA target sequence. A good example is a protein.
[0556] Guide polynucleotides / Cas enzymes that can cleave single strands of DNA target sequences. In this specification, a donuclease complex is defined as having nickas activity (e.g., partial cleavage ability). It can be characterized as having the force. Cas niccas is typically characterized by Cas being the DNA target sequence A single-function end that allows cutting only one strand (i.e., inserting a nick). Contains a nuclease domain. For example, Cas9 niccas is (i) mutated, functionally impaired. (ii) the entire RuvC domain, and (ii) the functional HNH domain (e.g., wild-type HNH domain) (i) may include. As another example, Cas9 nickase may include (i) a functional RuvC domain. (e.g., wild-type RuvC domain), and (ii) mutated, dysfunctional HNH domain This may include. Non-limiting examples of Cas9 nickase suitable for use herein include 2014 Disclosed in U.S. Patent Application Publication No. 2014 / 0189896, published on July 3rd. It is being used. A pair of Cas Nickers are being used to increase the specificity of DNA targeting. Ze can be used. Generally, this involves binding with RNA components that have different guide sequences. By being targeted, DN on the reverse chain within that region for desired targeting. To provide two Cas nickases that target the A sequence and make nicks in close proximity. Therefore, it can be done. Such breaks near each DNA strand are double-strand breaks (i.e., single-strand breaks). This results in a DSB with an overhang, which subsequently leads to a non-homologous terminal bond NH It is recognized as a substrate for EJ (which is prone to incomplete repair leading to mutations) or homologous recombination HR. Each of these embodiments has, for example, at least about 5, 5-10, At least 10, 10-15, at least 15, 15-20, at least 20, 20-3 0, at least 30, 30-40, at least 40, 40-50, at least 50, 50 ~60, at least 60, 60~70, at least 70, 70~80, at least 80, 80-90, at least 90, 90-100, or 100 or more (or 5- A base fraction of any integer (100), which may be separated from each other. One or two of the following in this specification One Cas nickase protein can be used in a Cas nickase pair. For example, Cas9 niccas with a mutated RuvC domain but a functional HNH domain. (That is, Cas9 HNH+ / RuvC-) can be used (for example, Leptococcus pyogenes (Streptococcus pyogenes) Cas 9 HNH+ / RuvC-). Each Cas9 nickase (e.g., Cas9 HNH+ / R uvC-) is a guide that directs each nickase to target its respective specific DNA site. By using preferred RNA components in this specification that have NA sequences, it is possible to obtain RNAs that are close to each other (most It can be induced to specific DNA sites (up to 100 base pairs away).
[0557] In certain embodiments, the guide polynucleotide / Cas endonuclease complex is It can bind to the DNA target site sequence, but does not cleave any strands at the target site sequence. Such a complex is dysfunctional when all of its nuclease domains are mutated. It may contain s proteins. For example, it can bind to DNA target site sequences, but the target The Cas9 protein, which does not cleave any strands at any site, is a mutated, dysfunctional RuvC protein. It may contain both the main and mutated, dysfunctional HNH domains. It binds to the target DNA sequence. However, the Cas protein in this specification does not cleave it, in order to regulate gene expression. It can be used, for example, in which case the Cas protein is a transcription factor (or part thereof) (For example, a repressor or activator, for example, any of those disclosed herein) and They can merge.
[0558] In one embodiment, the guide polynucleotide / Cas endonuclear described herein The PGEN complex is PGEN, and the Cas endonuclease is optional. Specifically, at least one protein subunit or its functional fragment is covalently or non-covalently bonded. This is a PGEN that is linked or assembled using a coupled mechanism.
[0559] In one embodiment of the present disclosure, a guide polynucleotide / Cas endonuclease complex This includes at least one guide polynucleotide and at least one Cas end nucleus. Guide polynucleotide / Cas endonuclease complex containing ase polypeptide ( PGEN) wherein the Cas endonuclease polypeptide is at least one The protein subunit or a functional fragment thereof is a chimeric polynucleotide. It is a non-natural guide polynucleotide, and the guide polynucleotide / Cas end nucleotide The rease complex recognizes and binds to all or part of the target sequence, and selectively performs the following action: It can be kinged, unraveled, or cut.
[0560] Cas effector proteins are Cas-alpha effectors disclosed herein. - It could be a protein.
[0561] In one embodiment of the present disclosure, the guide polynucleotide / Cas effector complex is small It contains at least one guide polynucleotide and Cas-alpha effector protein. The guide polynucleotide / Cas effector protein complex (PGEN) The guide polynucleotide / Cas effector protein complex controls the entire target sequence. Alternatively, it may be partially recognized, joined, and optionally nicked, untied, or cut. Cut.
[0562] PGEN may be a guide polynucleotide / Cas effector protein complex. Here, the Cas effector protein is at least one protein subunit or further comprising one or more copies of the functional fragment. Some embodiments So, the protein subunits are the Cas1 protein subunit and the Cas2 protein subunit. Protein subunits, Cas4 protein subunits, and any combination thereof. It is selected from the group consisting of the following. PGEN is a guide polynucleotide / Cas effector It may be a protein complex, where the Cas effector protein is Cas1, Ca At least two different protein subunits selected from the group consisting of s2 and Cas4. This also includes knitwear.
[0563] PGEN may be a guide polynucleotide / Cas effector protein complex. Here, the Cas effector protein is Cas1, Cas2, and optionally C At least three selected from the group consisting of one additional Cas protein including as4. Further comprising different protein subunits, or functional fragments thereof.
[0564] In one embodiment, the guide polynucleotide / Cas endonuclear described herein The ze complex (PGEN) is PGEN, and the Cas endonuclease is at least It is covalently or noncovalently linked to another protein subunit or its functional fragment. PGEN is a guide polynucleotide / Cas effector tangent. It may be a protein complex, where the Cas effector protein polypeptide is C as1 protein subunit, Cas2 protein subunit, and optionally Cas One further Cas protein subunit containing 4 proteins, and any of these At least one protein subunit selected from a group consisting of combinations, or It is linked to one or more copies of that functional fragment by covalent or non-covalent bonds. or assembled. PGEN is a guide polynucleotide / Cas effector It may be a protein complex, where the Cas effector protein is Cas1, A group consisting of Cas2 and one additional Cas protein, optionally including Cas4. Covalent or non-covalent bonding to at least two different protein subunits selected from They are linked or assembled. PGEN is a guide polynucleotide / Ca It may be an s-effector protein complex, where the Cas effector protein This includes Cas1, Cas2, and optionally Cas4, one further Cas protein At least three different tans selected from the group consisting of tannins and combinations thereof. It is linked to the protein subunit by covalent or non-covalent bonds.
[0565] Any component of the guide polynucleotide / Cas effector protein complex, guide The polynucleotide / Cas effector protein complex itself, as well as the polynucleotide The modified template and / or donor DNA are obtained by any method known in the art. It can then be introduced into different cells or organisms.
[0566] Recombinant constructs for cell transformation Disclosed guide polynucleotides, Cas endonucleases, polynucleotide repair Decorative template, donor DNA, guide polynucleotide / Cas disclosed herein An endonuclease system, and optionally one or more polynucleotides of interest. Any combination of these, including further, can be introduced into cells. This includes, but is not limited to, humans, non-humans, animals, bacteria, fungi, insects, yeasts, and non-humans. Conventional yeast and plant cells, and plant bodies and the method described herein. Seeds are one example.
[0567] The standard recombinant DNA and molecular cloning techniques used herein are subject to the regulations of the said technology. It is well known in the field, as described by Sambrook et al., Molecular Cl oning:A Laboratory Manual, Cold Spring Ha rbor Laboratory: Cold Spring Harbor, NY (19 This is explained in more detail in 89). The transformation method is well known to those skilled in the art, and is described below. I will explain it here.
[0568] The vectors and constructs include circular plasmids and linear polynucleotides, and these include The target polynucleotide, and optionally, linkers, adapters, regulatory elements or Other components, including analytical elements, are included. In some embodiments, the recognition site and / or target The target sites are introns, coding sequences, 5'UTR, 3'UTR, and / or regulatory regions. It can be contained within.
[0569] For the expression and utilization of a novel CRISPR-Cas system in prokaryotic and eukaryotic cells. Ingredients The present invention further relates to all or one of the target sequences in prokaryotic cells / organisms or eukaryotic cells / organisms. It can recognize, join, and optionally nick, untie, or cut parts. This provides an expression construct for expressing a guide RNA / Cas system.
[0570] In one embodiment, the expression construct of the present disclosure is a Cas gene (or as described herein). Nucleotides encoding optimized plants, including the Cas endonuclease gene A promoter operably ligated to a sequence, and a guide RNA operably ligated to the guide RNA of this disclosure. Includes a promoter. The promoter is a prokaryotic cell / organism or eukaryotic cell / organism odor. This allows for the driving of the expression of a functionally linked nucleotide sequence.
[0571] Nucleotide arrangement of guide polynucleotides, VT domain, and / or CER domain The column modification involves a 5' cap, a 3' polyadenylated tail, a riboswitch arrangement, and a stability control arrangement. The sequence, the dsRNA double-stranded sequence, and the target of the guide polynucleotide are positioned within the cell. Modifications or sequences that define a specific type of modification or sequence, modifications or sequences that result in tracking, and binding for proteins. Modifications or sequences resulting from site changes, locked nucleic acids (LNA), 5-methyl dC nucleotides , 2,6-diaminopurine nucleotide, 2'-fluoroA nucleotide, 2'-fluorine U nucleotide; 2'-O-methylRNA nucleotide, phosphorothioate bond, Linking to sterol molecules, linking to polyethylene glycol molecules, 18 spacer molecules A group consisting of a connection to, a covalent connection from 5' to 3', or any combination thereof. These modifications may be selected from, but are not limited to, at least one additional This can result in advantageous characteristics, and these additional advantageous characteristics can alter or regulate stability, or intracellular activity. Getting, tracking, fluorescent labeling, binding sites for proteins or protein complexes Changes in binding affinity to complementary target sequences, changes in resistance to cytodegradation, and fine Selected from the group exhibiting increased vesicular permeability.
[0572] In order to perform Cas9-mediated DNA targeting, RNA such as gRNA is used in eukaryotic cells. The method for expressing the component involves using the RNA polymerase III (Pol III) promoter. This involves using the precisely defined unmodified 5' and 3' ends. Enables transcription of RNA containing (DiCarlo et al., Nucleic Acids Res.41:4336-4343;Ma et al.,Mol.The r. Nucleic Acids 3:e161). This strategy is used for maize and daikon radish. It has been successfully applied to several different cell species, including Z. (Published March 19, 2015) (U.S. Patent Application Publication No. 2015 / 0082478). Without a 5' cap. Methods for expressing RNA components are described (International publication released February 18, 2016) (Publication No. 2016 / 025131)
[0573] It has a target polynucleotide that is inserted into the target site for Cas endonucleases. Various methods and compositions can be used to obtain cells or organisms. One such method involves using homologous recombination (HR) to incorporate the target polynucleotide into the target site. This can be done. In one method described herein, the target polynucleotide is donor —It is introduced into living cells by DNA constructs.
[0574] The donor DNA construct further includes first and second homologous regions adjacent to the target polynucleotide. Includes the region. The first and second homologous regions of donor DNA are, respectively, the genome of a cell or organism. For the first and second genomic regions located within or adjacent to the target site, They have the same sex.
[0575] Donor DNA can be ligated to guide polynucleotides. DNA is a useful target and donor for genome editing, gene insertion, and regulation of target genomes. This may enable DNA co-localization, and if the function of the endogenous HR mechanism is significantly reduced, This may be useful for targeting cells that have likely terminated cell division (Mali et al.). ,2013,Nature Methods Vol.10:957-963).
[0576] The amount of homology or sequence identity shared by target and donor polynucleotides varies. It is possible that it will be around 1-20 bp, 20-50 bp, 50-100 bp, 75-150 bp, 100~250bp, 150~300bp, 200~400bp, 250~500 bp, 300~600bp, 350~750bp, 400~800bp, 450~900 bp, 500~1,000bp, 600~1,250bp, 700~1,500bp, 8 00~1,750bp, 900~2,000bp, 1~2.5kb, 1.5~3kb, 2 ~4kb, 2.5~5kb, 3~6kb, 3.5~7kb, 4~8kb, 5~10kb, or includes the entire length and / or the entire region having a unit integer value within the range of the total length of the target site. These ranges include all integers within this range; for example, the range 1 to 20bp includes 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17 It includes 18, 19, and 20 bp. The amount of homology is also the completeness of the two polynucleotides. It can be described as the sequence identity percentage over the entire aligned length, At least approximately 50%, 55%, 60%, 65%, 70%, 71%, 72%, and 73%. 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%-99%, 99%, 99%-100% or 100% sequence identity percentage. Sufficient homology is polynucleotide. The length, the overall sequence identity percentage, and the conserved region of a sequence of nucleotides selected from among the optional sequences. or any combination with local sequence identity percentage, for example, sufficient homology This has at least 80% sequence identity with respect to the region of the target gene locus, ranging from 75 to 150 b. This can be explained as the region of p. Sufficient homology also exists under high stringency conditions. This can be explained by the predictive ability of two polynucleotides that hybridize specifically. Yes, for example, Sambrook et al., (1989) Molecular C loning:A Laboratory Manual, (Cold Spring Harbor Laboratory Press, NY);Current Prot. ocols in Molecular Biology,Ausubel et al .,Eds(1994)Current Protocols,(Greene Pub lishing Associates, Inc. and John Wiley &S ons, Inc.; and Tijssen (1993) Laboratory Tec hniques in Biochemistry and Molecular Bi ology--Hybridization with Nucleic Acid P See Robes (Elsevier, New York).
[0577] Structural similarity between a given genomic region and a corresponding homologous region found on donor DNA. Sex can be any degree of sequence identity that allows homologous recombination to occur. For example. Homology shared by "homologous regions" of donor DNA and "genomic regions" of the biological genome. Alternatively, the amount of sequence identity should be at least 50%, 55%, such that the sequence undergoes homologous recombination. 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, Sequence identity can be 96%, 97%, 98%, 99%, or 100%.
[0578] Homologous regions on donor DNA can be homologous to any sequence adjacent to the target site. In some cases, homologous regions are equivalent to genomic sequences directly adjacent to the target site. They share column homology, but this homologous region may be further 5' or 3' relative to the target site. It is recognized that it is possible to design it to have sufficient homology with respect to the region. This region can also exhibit homology to fragments of target sites in addition to downstream genomic regions. ru.
[0579] In one embodiment, the first homologous region further includes a first fragment of the target site, and the second homologous region The region includes a second fragment of the target site, and these first and second fragments are distinct.
[0580] Target polynucleotide This specification provides a detailed description of the target polynucleotide, including its commercial market and It contains polynucleotides that reflect the commercial market interests related to crop development. As goods and markets change, and developing countries expand into global markets, new crops and technologies also emerge. They will appear. Furthermore, regarding agricultural traits and characteristics such as yield and hybrid vigor As our understanding increases, the selection of genes for genetic modification will change accordingly. It will probably happen.
[0581] Common categories of target polynucleotides include, for example, zinc fingers. Genes involved in information transfer, genes involved in signal transduction such as kinases, and heat shock proteins Examples include genes involved in housekeeping, such as those related to blood circulation. More specific objectives: Rheotides include, but are not limited to, crop yield, grain quality, and nutrient content of grains. Quantity, quality and quantity of starch and carbohydrates, as well as grain size, sucrose load, protein Factors affecting quality and quantity, nitrogen fixation and / or utilization, fatty acids and oil composition, etc. , genes involved in traits of agricultural benefit, abiotic stress (drought, nitrogen, temperature, salinity, To confer resistance to toxic metals or trace elements, or toxins such as insecticides and herbicides. Genes that encode proteins that confer resistance to (things, etc.), biological stress ( Attacks by fungi, viruses, bacteria, insects and nematodes, and diseases associated with these organisms. Examples include genes that encode proteins that confer resistance to disease progression, etc. , but not limited to these.
[0582] Agriculturally important traits such as oil, starch, and protein content are different from those obtained through conventional breeding methods. In addition to use, it can be genetically modified. Modifications include oleic acid, saturated oil, and unsaturated oil. Increased content of Japanese oil, increased lysine and sulfur concentration, provision of essential amino acids, and starch Modifications are included. The modification of the folidthionine protein is protected by U.S. Patent No. 5,703,049. The specifications, specifications No. 5,885,801, specifications No. 5,885,802, and the same It is stated in the 5,990,389 item.
[0583] The target polynucleotide sequence is a protein involved in the induction of disease resistance or pest resistance. You may also code "disease resistance" or "pest resistance" as the interaction between plants and pathogens. This means that plants avoid the harmful symptoms that result from pests. Pest resistance genes are genes that prevent pests from stimulating the growth of necrotic organisms. For pests that hinder high yields, such as worms, armyworms, and European corn borers. It may encode resistance to diseases such as lysozyme or cecropine, which protect antibacterial properties. and insect resistance genes, or defensins, glucanases or kitchen that protect antifungal properties Bacillus thuringiensis controls proteins such as enzymes, or nematodes or insects. Bacillus thuringiensis endotoxin, protease inhibitor Harmful agents, collagenases, lectins, or glucosidases are all useful gene products. For example, genes encoding disease resistance traits include those for detoxification against fumonisin. Genes (U.S. Patent No. 5,792,931), non-pathogenic (AVR) genes and disease resistance Sex (R) gene (Jones et al. (1994) Science 266:78 9; Martin et al. (1993) Science 262:1432; and Mindrinos et al. (1994) Cell 78:1089) are cited as examples. It is possible.
[0584] Insect resistance genes are used for high yields against cutworms, armyworms, European corn borers, etc. It may encode resistance to pests that inhibit it. Such genes include, for example, Bacillus thuringiensis (toxic) Protein genes (U.S. Patent No. 5,366,892; U.S. Patent No. 5,747,450) Specification; Specification No. 5,736,514; Specification No. 5,723,756; Specification No. 5, Specification No. 593,881; and Geiser et al. (1986) Gene 48 Examples include: 109).
[0585] Obtained by the expression of "herbicide-resistant proteins" or "nucleic acid molecules encoding herbicide resistance." As for proteins that can be affected, cells that do not express that protein are subjected to higher concentrations of herbicide. The ability to withstand a certain concentration for a longer period of time than cells that do not express that protein. One example is a protein that confers the ability to withstand herbicides to cells. Herbicide resistance traits include acetaminophen. The function of lactate synthase (ALS, also known as acetohydroxy acid synthase or AHAS) Herbicides that act to inhibit this, especially sulfonylurea. (UK: Resistance to sulfonylurea-type herbicides) Herbicides that work by inhibiting the function of the gene that controls glutamine synthesis, for example, A gene that encodes resistance to sphinotricin or basta (e.g., bar (genes), genes encoding resistance to glyphosate (e.g., EPSP synthase gene) Genes and GAT genes), genes encoding resistance to HPPD inhibitors (e.g., H The plant body (PPD gene), or other such genes known in the art It can be introduced into, for example, U.S. Patent No. 7,626,077, Sections 5, 3. Specification No. 10,667, Specification No. 5,866,775, Specification No. 6,225,114 Specification, Specification No. 6,248,876, Specification No. 7,169,970, No. 6,8 See Specification No. 67,293 and Specification No. 9,187,762. The gene codes for resistance to the herbicide Basta, and the nptII gene codes for resistance to the antibiotic Kana Encoding resistance to mycin and geneticin, the ALS gene mutant is used for weed control. This codes for resistance to the agent chlorsulfuron.
[0586] Furthermore, the target polynucleotide also acts as a messenger for the target gene sequence of the desired purpose. - RNA (mRNA) may contain a complementary antisense sequence in at least a portion of it. Sense nucleotides are constructed to hybridize with the corresponding mRNA. The cystic sequence hybridizes with the corresponding mRNA, thereby interfering with its expression. It may be modified as long as it is within limits. In this way, 70% and 80% for the corresponding antisense array. Antisense constructs with % or 85% sequence identity can be used. The use of antisense nucleotide portions to interfere with the expression of target genes is possible. Yes, it is possible. Generally, at least 50 nucleotides, 100 nucleotides, 200 nucleotides You can use a sequence of 1 or more characters.
[0587] Furthermore, the target polynucleotide also senses the suppression of endogenous gene expression in plants. It can be used for orientation. Sense-oriented polynucleotides can be used for plant genes. Methods for suppressing expression are known in the art. These methods generally involve endogenous Plants that have operably ligated at least a portion of the nucleotide sequence corresponding to gene transcription. Transforming plants with DNA constructs containing promoters that drive expression within the body It is included. Typically, such nucleotide sequences are equivalent to the transcription sequence of the endogenous gene. Sequence identity is generally defined as sequence identity exceeding approximately 65%, or sequence identity exceeding approximately 85%. It has sequence identity of approximately 95% or more. U.S. Patent No. 5,283,184 See also Specification No. 5,034,323.
[0588] The target polynucleotide may also be a phenotypic marker. The phenotypic marker is positive. Whether they are selection markers or negative selection markers, visual markers and selection markers A screening marker or selection marker containing a marker. Any phenotypic marker It can be used. In detail, selection markers or screening markers are many In some cases, under specific conditions, it is possible to identify a single molecule or a cell containing this molecule, or DNA segments that enable selection favorably or unfavorably for molecules or cells. These markers include, but are not limited to, RNA, peptides, or proteins. It can encode activities such as the production of proteins, or RNA, peptides, proteins, It can provide bonding sites for inorganic and organic compounds or compositions.
[0589] Examples of selection markers include, but are not limited to, DNA segmentation including restriction enzyme sites. Spectinomycin, ampicillin, kanamycin, tetracycline, Bast a. Neomycin phosphotransferase II (NEO) and hygromycin phospho In contrast to antibiotics such as transferase (HPT), and other compounds that are toxic in other respects. DNA segment encoding a product that provides resistance; other aspects in recipient cells Then, the DNA segment that codes for the missing product (e.g., tRNA gene, nutrient Psychophilic markers); DNA segments encoding easily identifiable products (e.g., β-gas). Phenotypic markers such as lactosidase and GUS; green fluorescent protein (GFP), cyanide Fluorescent proteins such as (CFP), yellow (YFP), and red (RFP) fluorescent proteins, and cell surface proteins); novel primer sites for PCR (e.g., previously common) Generation of two previously unplaced DNA sequences (parallel), restriction endonuclease or other The presence of DNA sequences that have not been affected by DNA modifying enzymes, chemicals, etc., or that have been affected by such substances, and The presence of DNA sequences necessary for specific modifications (e.g., methylation) that enable their identification is a key factor. It can be done.
[0590] Further selection markers include sulfonylurea, glufosinate ammonium, and bron. Moxynyl, imidazolinone, and 2,4-dichlorophenoxyacetate (2,4-D) Examples of genes that confer resistance to herbicides include sulfonylurea. imidazolinone, triazolopyrimidine sulfonamide, pyrimidinyl salicylate and Acetolactate synthase (AL) resistant to sulfonylaminocarbonyltriazolinone S)(Shaner and Singh,1997,Herbicide Activation ity:Toxicol Biochem Mol Biol 69-110); Sar resistant 5-enolpyruvir schimate-3-phosphate (EPSPS) (Sar oha et al., 1998, J. Plant Biochemistry & B Please refer to iotechnology Vol 7:65-72;
[0591] The target polynucleotide may have other traits, such as herbicide resistance or as described herein. It can be stacked or used in combination with any other traits (but not limited to these). It contains genes that can be used. The target polynucleotide and / or trait was made public on October 3, 2013. The published U.S. Patent Application Publication No. 2013 / 0263324 and dated August 1, 2013 As described in the international publication pamphlet No. 2013 / 112686, which was published in [location], They can be stacked together at the synphenotypic locus.
[0592] The target polypeptide is encoded by the target polynucleotide described herein. It contains any protein or polypeptide.
[0593] Furthermore, the genome contains the target polynucleotide incorporated into the target site, and less A method for identifying a single type of plant cell is provided. Insertion into a target site or its vicinity in the genome. Various methods can be used to identify plant cells that have undergone testing. Such methods include This can be considered a direct analysis of the target sequence to detect changes within the target sequence, and is limited to the following: Although not performed, PCR method, sequencing method, nuclease digestion method, Southern blotting method and this Any combination of these is possible. For example, the US Special Measures released on May 21, 2009 See Patent Application Publication No. 2009 / 0133152. The method also applies to the genome. This involves recovering a plant body from plant cells containing the incorporated target polynucleotide. The substance may be sterile or fertile. Any polynucleotide of any purpose is provided. It was confirmed that the molecule could be incorporated into a target site in the plant genome and expressed in the plant.
[0594] Sequence optimization for expression in plants Methods for synthesizing plant-preferential genes are available in the art. For example, U.S. patent Specification No. 5,380,831 and Specification No. 5,436,391, and Murra y et al. (1989) Nucleic Acids Res.17:477-4 See page 98. Further sequence modifications to enhance gene expression in plant hosts are These modifications include, for example, coding a pseudo-polyadenylation signal. Removal of one or more sequences, one or more exon-intron splice site signals Removal, removal of one or more transposon-like repeats, and potential adverse effects on gene expression. One example is the removal of other well-characterized sequences. The GC content of a sequence is The average level of a given plant host (calculated by referring to known genes expressed in that host plant cell) It can be adjusted to (this). If possible, the predicted hairparticles of one or more mRNAs. The sequence is modified to avoid secondary structures. Therefore, the "plant-optimized nucleo" of this disclosure A "CID sequence" contains one or more modifications of such sequences.
[0595] Expression element The Cas protein or other CRISPR system components disclosed herein may be coded Any polynucleotides that facilitate transcription or regulation in host cells are different It can be functionally linked to the expression elements of a species. Such expression elements include pro Examples include Motors, Leaders, Introns, and Terminators, but are not limited to these. i. The expression element may be "minimal," and this may be an expression regulator or modifier. It refers to a short sequence derived from a natural source that still functions. Alternatively, it refers to an expression element. The nucleotide can be "optimized," meaning its polynucleotide sequence is optimized for a particular host cell. It is modified from its natural state in order to function with more desirable characteristics. This means (for example, improving its expression in maize plants, but not limited to the following) To achieve this, the bacterial promoter may be "maize-optimized." Alternatively, the expression element The expression can be "synthetic," meaning the expression element is designed in silico and the host cell This means that it is synthesized for use in cells. Synthetic expression elements are entirely synthetic. However, it is also acceptable if it is partially synthetic (including fragments of naturally occurring polynucleotide sequences). stomach.
[0596] Certain promoters can induce RNA synthesis at a higher rate than others. This has been shown to be true. These are called "strong promoters." Lomotors induce high levels of RNA synthesis only in specific cell or tissue types. This shows that the promoter favorably induces RNA synthesis in certain tissues, while in other tissues it does not. To induce RNA synthesis at a high level, use a "tissue-specific promoter" or a "tissue-priority promoter". It is often called a "motor."
[0597] A plant promoter contains a promoter that can initiate transcription in a plant cell. Hmm. For an explanation of plant promoters, see Potenza et al., 2004, I n vitro Cell Dev Biol 40:1-22;Porto et a l.,2014,Molecular Biotechnology(2014),56 See (1), 38-49.
[0598] Examples of constitutive promoters include the core CaMV 35S promoter (Odel). l et al., (1985) Nature 313:810-2); Rineactin ( McElroy et al., (1990) Plant Cell 2:163-71 ); Ubiquitin (Christensen et al., (1989) Plant M ol Biol 12:619-32; ALS promoter (US 5,659,0 Examples include Specification No. 26.
[0599] Tissue-preferred promoters target enhanced expression within specific plant tissues. It can be used for the purpose of [details omitted]. As an organization's preferred promoter, for example, July 2013 International release pamphlet No. 2013 / 103367, published on the 11th of the month, Kawamata et al.,(1997)Plant Cell Physiol 38:792-8 03;Hansen et al.,(1997)Mol Gen Genet 254 :337-43;Russell et al.,(1997)Transgenic Res 6:157-68;Rinehart et al.,(1996)Plant Physiol 112:1331-41;Van Camp et al.,(19 96)Plant Physiol 112:525-35;Canevascini et al.,(1996)Plant Physiol 112:513-524;L am,(1994)Results Probl Cell Differ 20:18 1-96; and Guevara-Garcia et al., (1993) Plant J 4:495-505 is cited as an example. A leaf-preferred promoter is, for example, Yam. amoto et al.,(1997)Plant J 12:255-65;Kwo n et al., (1994) Plant Physiol 105:357-67; Yamamoto et al.,(1994)Plant Cell Physiol 35:773-8;Gotor et al.,(1993)Plant J 3:5 09-18;Orozco et al.,(1993)Plant Mol Biol 23:1129-38;Matsuoka et al.,(1993)Proc.N atl.Acad.Sci.USA90:9586-90;Simpson et al .,(1958)EMBO J 4:2723-9;Timko et al.,(19 88) Nature 318:57-8 is cited. Examples of root-preferential promoters include For example, Hire et al., (1992) Plant Mol Biol 20:2 07-18 (Soybean root-specific glutamine synthase gene); Miao et al., ( 1991) Plant Cell 3:11-22 (Cytosol glutamine synthase (G S));Keller and Baumgartner, (1991) Plant C ell 3:1051-61 (Root-specific control of the GRP 1.8 gene in green beans) The element of nature; Sanger et al., (1990) Plant Mol Bi ol 14:433-43 (A. tumefaciens man) Root-specific promoter of nopine synthase (MAS); Bogusz et al., 1990) Plant Cell 2:633-41 (Parasponia andersonii) Parasponia andersonii and Trema tomentosa A root-specific promoter isolated from tomentosa); Leach and A oyagi, (1991) Plant Sci 79:69-76 (A. rhisogenes (A rhizogenes) rolC and rolD root-inducing genes); Teeri et al. l., (1989) EMBO J 8:343-50 (Agrobacterium (Agrob Acterium) Wound-inducing TR1' and TR2' genes; VfENOD-GRP3 gene Genetic promoter (Kuster et al., (1995) Plant Mol B iol 29:759-72); and rolB promoter (Capana et al. .,(1994) Plant Mol Biol 25:681-91; Pazeolin Gene (Murai et al., (1983) Science 23:476-82; Sengopta-Gopalen et al.,(1988)Proc.Natl. Acad.Sci.USA82:3320-4) is cited. U.S. 5,837, Specification No. 876, Specification No. 5,750,386, Specification No. 5,633,363, Specification No. 5,459,252, Specification No. 5,401,836, Specification No. 5,110, See also Specification No. 732 and Specifications No. 5,023,179.
[0600] As a seed-preferential promoter, a seed-specific promoter that is active during seed development, and seed Both types of seed germination promoters that are active during germination can be cited. (Thompson et al.) See al., (1989) BioEssays 10:108. Seed priority As a romor, Cim1 (cytokinin induction message); cZ19B1 (to Sorghum 19kDa zein; and milps (myo-inositol-1-phosphate synthase). ); and, for example, International Publication No. 2000 / 01117, published on March 2, 2000. Examples include those disclosed in Pamphlet No. 7 and U.S. Patent No. 6,225,529. However, it is not limited to these. Suitable seed promoters for dicotyledonous plants include the following: While not limited to these, mame-β-phaseolin, napin, β-conglycinin, and soybean leukate are also used. Examples include tin and cruciferin. Seed preference promoters for monocotyledonous plants include , but not limited to the following, corn 15kDa zein, 22kDa zein, 27k Da gammazein, waxy, shrunken 1, shrunken 2, globulin 1, oleo Syn and nuc1 are examples. International release No. 2000 / was released on March 9, 2000. Please also refer to pamphlet number 012733. This document contains information on END1 and END2 genes. A seed-preferential promoter derived from offspring has been disclosed.
[0601] The application of exogenous chemical modifiers regulates gene expression in prokaryotic and eukaryotic cells or organisms. Therefore, a chemically inducible (regulating) promoter can be used. The promoter is chemical Even if the application of a chemical substance induces gene expression through a chemically inducible promoter, the application of the chemical substance The promoter used may be a chemosuppressive promoter that suppresses gene expression. As for the ter, it is not limited to the following, but is a benzenesulfonamide herbicide toxicity mitigator. The maize In2-2 promoter is activated by (De Veylder et al.) al., (1997) Plant Cell Physiol 38:568-77), Corn is activated by hydrophobic electrophilic compounds used as pre-germination herbicides. GST Promoter (GST-II-27, International Release No. 1, released January 21, 1993) (Pamphlet No. 993 / 001294), and Tobacco P activated by salicylic acid R-1a promoter (Ono et al., (2004) Biosci Biote (Chonol Biochem 68:803-7) is one example. Other chemical regulation promotions As a promoter, a steroid-responsive promoter (for example, a glucocorticoid-inducible promoter) is used. Motor (Schena et al., (1991) Proc. Natl. Acad.) Sci.USA88:10421-5;McNellis et al.,(1998) See Plant J 14:247-257); Tetracycline-inducible and Tetracycline-inhibiting promoter (Gatz et al., (1991) Mol Gen Genet 227:229-37; U.S. Patent No. 5,814,618 and Examples include the specifications of the same patent application No. 5,789,156.
[0602] Pathogen-inducible promoters induced after pathogen infection are not limited to the following: However, PR protein, SA protein, beta-1,3-glucanase, chitinase, etc. Which of the following substances regulate expression?
[0603] As for stress-inducible promoters, the RD29A promoter (Kasuga et al.) al. (1999) Nature Biotechnol. 17:287-91) was cited. Those skilled in the art can provide stress such as drought, osmotic stress, salt stress and temperature stress. A protocol for simulating the response conditions, and the simulated or A protocol for evaluating the stress tolerance of plants placed under naturally occurring stress conditions. He is very knowledgeable about [the subject].
[0604] Another example of an inducible promoter useful for plant cells was published on November 21, 2013. The ZmCAS1 prototype described in U.S. Patent Application Publication No. 2013 / 0312137 It's a motor.
[0605] Many new types of promoters useful for plant cells are constantly being discovered, and there are many examples. However, The Biochemist, edited by Okamuro and Goldberg (1989) stry of Plants,Vol.115,Stumpf and Conn,e See ds (New York, NY: Academic Press), pp. 1-82. It can be released.
[0606] Genome modification by novel CRISPR-Cas system components As described herein, guided Cas endonucleases target DNA targets It can recognize sequences, join them, and introduce single-strand (nick) or double-strand breaks. When a double-strand break is introduced into the DNA, the cell's DNA repair mechanism attempts to repair the break. It is activated. The error-prone DNA repair mechanism creates mutations at double-strand break sites. This is possible. The most common repair mechanism for joining damaged ends together is non-homologous ends. The NHEJ binding pathway (Bleuyard et al., (2006) DNA Repair 5:1-12). Chromosomal structural integrity is typically maintained through repair. However, deletions, insertions, or other rearrangements (such as chromosomal translocations) are also possible (Siebert and Puchta,2002,Plant Cell 14:1121-31;Pa cher et al., 2007, Genetics 175:21-9).
[0607] DNA double-strand breaks appear to be effective factors for stimulating the homologous recombination pathway. Puchta et al.,(1995)Plant Mol Biol 28:28 1-92;Tzfira and White,(2005)Trends Biote chnol 23:567-9;Puchta,(2005)J Exp Bot 56 (1-14) Using DNA cleavage agents, homologous DNA artificially constructed in plants can be created. A 2- to 9-fold increase in homologous recombination was observed between repeats (Puchta et al., (1995) Plant Mol Biol 28:281-92). Corn Pro Experiments using linear DNA molecules within toplasts revealed enhanced homology between plasmids. Recombination was proven (Lyznik et al., (1991) Mol Gen Gen et 230:209-18).
[0608] Homologous recombination repair (HDR) is a mechanism that repairs double-strand and single-strand DNA breaks within cells. Yes, homologous recombination repair includes homologous recombination (HR) and single-strand annealing (SSA). (Lieber.2010 Annu.Rev.Biochem.79:1) 81-211). The most common form of HDR is the combination of donor DNA and acceptor DNA. This is called homologous recombination (HR) that has the longest sequence homology requirement between the sequences. Other forms of HDR include This includes single-strand annealing (SSA) and cleavage-induced replication, which are compared to HR. It requires shorter sequence homology. Homologous recombination repair for nicking (single-strand breaks) , can occur through a different mechanism than HDR for double-strand breaks (Davis and Maiz els.PNAS(0027-8424),111(10),p.E924-E932) .
[0609] For example, homologous recombination (HR) can modify the genomes of prokaryotic and eukaryotic living cells or other biological cells. Homologous recombination is a powerful tool for genetic engineering. (Halfter et al.) al., (1992) Mol Gen Genet 231:186-93) and insects ( Dray and Gloor, 1997, Genetics 147:689-99) This has been demonstrated in [location]. Homologous recombination has also been achieved in other organisms. For example, Due to homologous recombination in the parasitic protozoan genus Leishmania. This required at least 150-200 bp of homology (Papadopoulo u and Dumas,(1997) Nucleic Acids Res 25:4 278-86). The filamentous fungus Aspergillus n In idulans, gene substitution is achieved using a small adjacent homology of 50 bp. (Chaveroche et al., (2000) Nucleic Acids Res 28:e97). Targeted gene replacement is also used in the ciliate Tetrahymena. This has also been demonstrated in thermophila (Tetrahymena thermophila). (Gaertig et al., (1994) Nucleic Acids R (es 22:5391-8). In mammals, homologous recombination is performed by growing organisms in culture and transforming them. Using pluripotent embryonic stem cell (ES) lines that can be selected and introduced into mouse embryos The most successful in [the field] (Watson et al., 1992, Recombination) nant DNA,2nd Ed.,Scientific American Boo ks distributed by W.H. Freeman & Co.).
[0610] Gene targeting The guide polynucleotide / Cas systems described herein are gene targeting It can be used for ing.
[0611] Generally, DNA targeting is suitable for polynucleotide components of Cas proteins. Related to this, cleaving one or both strands at a specific polynucleotide sequence in a cell. This can be performed by inducing single-strand or double-strand breaks in the DNA, which can then be performed on the cell's DN. The A repair mechanism is activated, and non-homologous end joining (NHEJ) or homologous recombination repair (HDR) is performed. The process can repair the damage and lead to modification of the target site.
[0612] The length of the DNA sequence at the target site can vary, for example, the length may be at least 12, 13, or 14 , 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, The target site contains 28, 29, 30, or 30 or more nucleotides. It can be a sentence structure, or in other words, the sequence on one chain reads the same thing in the opposite direction on the complementary chain. Further extraction is possible. Nick / cleavage sites can be located within the target sequence. Alternatively, the nick / cleavage site may be located outside the target sequence. In another variation... The cleavage can occur at nucleotide positions directly opposite each other, potentially resulting in a blunt end cleavage. It has properties, and in other cases, the notches are misaligned, resulting in a single-stranded o This can result in an overhang (which may be a 5' overhang or a 3' overhang). It is possible. Active variants of genomic target sites can also be used. The active variant is at least 65%, 70%, 75%, and 80% of the given target site. 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% It can contain 99% or more sequence identity, and the active variant has biological activity. It retains these molecules and can therefore be recognized and cleaved by Cas endonucleases.
[0613] Assay for measuring single-strand or double-strand breaks at target sites by endonucleases This is known in the field of technology, and generally, the entire amount of the active substance on the DNA substrate including the recognition site. The physiological activity and specificity are measured.
[0614] The targeting method described herein involves, for example, targeting two or more DNA target regions in this manner. The procedure can be implemented so that the position is targeted. Such methods can be selectively performed using multiple methods. It can be characterized as follows. In certain embodiments, 2, 3, 4, 5, 6, 7, 8, 9, 10 Or, more target sites may be targeted simultaneously. Multiplexing typically involves targeting multiple different RNAs. Each component solidifies the guide polynucleotide / Cas endonuclease complex. A (designed to guide to a specific DNA target site) is provided herein. It will be implemented using the getting method.
[0615] gene editing The genome sequence editing process, which combines DSBs and modification templates, is generally performed by staining. DSB inducers that can recognize target sequences in body sequences and introduce DSBs into genome sequences. Nucleic acids encoding a quality or DSB inducer, and a smaller amount compared to the nucleotide sequence to be edited. A polynucleotide modification template containing at least one nucleotide modification. This involves introducing the template into host cells. The polynucleotide modification template further It may include a nucleotide sequence adjacent to at least one nucleotide modification. The adjacent sequences are substantially homologous to the chromosomal regions adjacent to the DSB. DSB inducer For example, genome editing using the Cas-gRNA complex, for example, March 19, 2015 The specification of U.S. Patent Application Publication No. 2015 / 0082478, published on February 2015. International publication pamphlet No. 2015 / 026886, released on the 26th, January 2016. International publication brochure No. 2016 / 007347, released on the 14th, and 2016 As stated in the international publication pamphlet No. 2016 / 025131, released on February 18th Yes, they are.
[0616] Several uses of guide RNA / Cas endonuclease systems have been described. (For example, U.S. Patent Application Publication No. 2015 / 00824, published on March 19, 2015) Specification No. 78A1, International Publication No. 2015 / 02688, published on February 26, 2015. Pamphlet No. 6, and U.S. Patent Application Publication No. 2015, published on February 26, 2015. (See Specification No. 0059010), which includes the target nucleotide sequence (modification) Modification or substitution of elements, insertion of the target polynucleotide, gene knockout gene knock-in, modification of splicing sites and / or introduction of alternative splicing sites Modification of the nucleotide sequence encoding the target protein, amino acids and / or protein Quality fusion, and gene silencing by expressing reverse repeats in the target gene. This includes, but is not limited to, ing.
[0617] Proteins undergo various processes including amino acid substitution, deletion, truncation, and insertion. It can be modified. Methods for such operations are generally known. For example, Protein amino acid sequence variants can be prepared by mutations in DNA. Methods for mutagenesis and nucleotide sequence modification include, for example, Kunkel, (198 5)Proc.Natl.Acad.Sci.USA 82:488-92;Kunke l et al.,(1987)Meth Enzymol 154:367-82;Rice Japanese Patent No. 4,873,192; Walker and Gaastra, eds. (1983) Techniques in Molecular Biology(M acMillan Publishing Company, New York) and so The cited literature is listed here. It is unlikely that it will affect the biological activity of the protein. Guidance on mino acid substitution can be found, for example, in Dayhoff et al. (1978) )Atlas of Protein Sequence and Structure (Natl Biomed Res Found, Washington, DC) This can be found in models, such as by replacing one amino acid with another amino acid that has similar properties. Conservative substitutions would be preferable. Conservative deletions, insertions, and amino acid substitutions are protein It is expected that this will not cause radical changes to the properties, such as substitution, deletion, insertion, or combination thereof. The impact of combinations can be evaluated with routine screening analyses. Double-strand breaks Assays for inductive activity are known, and generally involve DNA substrates containing the target site. The overall activity and specificity of the above-mentioned active substances are measured.
[0618] Cas endonuclease, and Cas endonuclease and guide polynucleotide This specification describes a genome editing method using a complex containing tide. Guide RNA and PA Following the characterization of the M sequence, the endonuclease and associated CRISPR RNA (c Components of rRNA are used to modify chromosomal DNA in other organisms, including plants. It may contain a complex to facilitate optimal expression and nuclear localization (for eukaryotic cells). The gene was published in International Publication No. 2016 / 186953 on November 24, 2016. Optimized as indicated on the frets, and then by a method known in the art, D It can be delivered to cells as an NA expression cassette. The components necessary to include the active complex are also, RNA with or without modifications that protect RNA from degradation, i.e., caps Capped or uncapped mRNA (Zhang, Y. et al.) (2016, Nat.Commun.7:12617), or Cas protein guidepost As a renucleotide complex (International Publication No. 2017 / 070, published April 27, 2017) It may be delivered as pamphlet No. 032, or any combination thereof. Furthermore, Some or more parts of the complex and crRNA can be expressed from a DNA construct. On the other hand, the other components are RNA with or without modifications that protect the RNA from degradation. In other words, as capped or uncapped mRNA (Zhang et al. (al.2016Nat.Commun.7:12617), or Cas protein guide As a dopolynucleotide complex (International Publication No. 2017 / 0, published April 27, 2017) Pamphlet No. 70032 may be delivered, or any combination thereof. To produce RNA in vivo, for example, the international publicly released on June 22, 2017 As described in Patent Publication No. 2017 / 105991, using tRNA-derived elements This process cleaves the crRNA transcript into a mature form capable of inducing a complex to form at its DNA target site. Endogenous RNAse can also be recruited for this purpose. The nickase complex can be recruited individually or cooperatively. To use in combination to generate one or more DNA nicks on one or both of the DNA strands. Furthermore, the cleavage activity of Cas endonuclease is important in its cleavage domain. It can be inactivated by modifying the catalytic residue (Sinkunas, Te t al, 2013, EMBO J. 32:385-394), as a result, homologous recombination repair It can be used to enhance reproduction, induce transcriptional activation, or reconstruct DNA local structures. It generates RNA-induced helicase. Furthermore, the Cas cleavage domain and the helicase domain It knocks out both of the other DNA cleavage, DNA nicking, DNA binding, and transcription activities. Activation, transcriptional repression, DNA rearrangement, DNA deamination, DNA unwinding, DNA recombination amplification It can be used in combination with strong DNA integration, DNA inversion, and DNA repair agents.
[0619] tracrRNA (if present) and CRISPR of the CRISPR-Cas system -Other components of the Cas system (e.g., variable targeting domain, crRNA lipo The transfer directions (of loops, anti-repeats) were published internationally on November 24, 2016. Pamphlet No. 2016 / 186946 and published on November 24, 2016 It can be presumed that this is as described in the International Publication Pamphlet No. 2016 / 186953.
[0620] Once appropriate guide RNA requirements are established as described herein, they will be opened to the public. The PAM priority for each of the new systems presented can be examined. The cleavage complex is random When degradation of the PAM library occurs, it may be due to mutagenesis of key residues, or As previously mentioned, by constructing the reaction in the absence of ATP (Sink unas, T. et al., 2013, EMBO J.32:385-394), AT By inactivating Pase-dependent helicase activity, the complex is converted to nicasse. It is possible. Two PAM randomizations separated by two protospacer targets By utilizing the region, double-stranded DNA breaks can be generated, which can then be captured and sequenced. This allows us to examine the PAM sequences that support cleavage by each complex.
[0621] In one embodiment, the present invention describes a method for modifying a target site in the genome of a cell. The method involves introducing at least one PGEN described herein into cells, and the This includes identifying at least one cell having a modification in the target site. Modifications include (i) substitution of at least one nucleotide, and (ii) substitution of at least one nucleo. (iii) deletion of a nucleotide, (iii) insertion of at least one nucleotide, at least one nucleotide Chemical modifications of ocide, and selection from the group consisting of any combination of (v)(i)~(iv) It will be selected.
[0622] The nucleotide to be edited is recognized and cleaved by Cas endonuclease. It may be located inside or outside the target area. In one embodiment, at least one Nu The modification of the creotide is recognized and cleaved at the target site by Cas endonuclease. This is not a revision. In another practical form, at least one nucleotide and geno should be edited. Between the target site and the target site, there are at least 1, 2, 3, 4, 5, 6, 7, 8, and 9 pieces, 10 pieces, 11 pieces, 12 pieces, 13 pieces, 14 pieces, 15 pieces, 16 pieces, 17 pieces, 18 pieces, 19 pieces pieces, 20 pieces, 21 pieces, 22 pieces, 23 pieces, 24 pieces, 25 pieces, 26 pieces, 27 pieces, 30 pieces, 40 pieces pieces, 50 pieces, 100 pieces, 200 pieces, 300 pieces, 400 pieces, 500 pieces, 600 pieces, 700 pieces There are 900 or 1000 nucleotides.
[0623] Knockout is the insertion of nucleotide bases into the target DNA sequence by indels (NHEJ). (Inclusion or deletion), or reduction or complete loss of function of the target site or its vicinity. It can be generated by the specific removal of the sequence.
[0624] Guide polynucleotide / Cas endonuclease-induced targeted mutations are Cas endonuclease-induced mutations. Nucleases located inside or outside the genomic target site that are recognized and cleaved by nucleases This can occur within the rheotide sequence.
[0625] Methods for editing nucleotide sequences in a cell's genome can change the function of non-functional gene products. By restoring the condition, a method that does not use exogenous selective markers can be avoided.
[0626] In one embodiment, the present invention describes a method for modifying a target site in the genome of a cell. This method involves at least one PGEN described herein, and at least one D The process includes introducing donor DNA into cells, wherein the donor DNA contains the target polynucleotide. Furthermore, the target polynucleotide is selectively incorporated into the target site or its vicinity. This further includes identifying at least one cell.
[0627] In one embodiment, the method disclosed herein involves combining a target polynucleotide at a target site. Homologous recombination (HR) can be used to provide the necessary components.
[0628] The activity of the CRISPR-Cas system components described herein allows for insertion into the target site. Various methods and combinations are used to produce cells or organisms having the target polynucleotides. The product may be used. In one method described herein, the target polynucleotide is It is introduced into living cells by a donor DNA construct. When used herein, "donor" DNA is the target polynucleotide that is inserted into the target site of Cas endonuclease. It is a DNA construct that includes the target polynucleotide. The donor DNA construct further contains adjacent to the target polynucleotide. It includes the first and second homologous regions. The first and second homologous regions of the donor DNA are each The first and located within or adjacent to the target site of the genome of a cell or organism. It exhibits homology to the second genomic region.
[0629] Donor DNA can be bound to guide polynucleotides. DNA is a useful target and donor for genome editing, gene insertion, and regulation of target genomes. This may enable DNA co-localization, and if the function of the endogenous HR mechanism is significantly reduced, This may be useful for targeting cells that have likely terminated cell division (Mali et al.). ,2013,Nature Methods Vol.10:957-963).
[0630] The amount of homology or sequence identity shared by target and donor polynucleotides varies. It is possible that it will be around 1-20 bp, 20-50 bp, 50-100 bp, 75-150 bp, 100~250bp, 150~300bp, 200~400bp, 250~500 bp, 300~600bp, 350~750bp, 400~800bp, 450~900 bp, 500~1,000bp, 600~1,250bp, 700~1,500bp, 8 00~1,750bp, 900~2,000bp, 1~2.5kb, 1.5~3kb, 2 ~4kb, 2.5~5kb, 3~6kb, 3.5~7kb, 4~8kb, 5~10kb, or includes the entire length and / or the entire region having a unit integer value within the range of the total length of the target site. These ranges include all integers within this range; for example, the range 1 to 20bp includes 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17 It includes 18, 19, and 20 bp. The amount of homology is also the completeness of the two polynucleotides. It can be described as the sequence identity percentage over the entire aligned length, At least approximately 50%, 55%, 60%, 65%, 70%, 71%, 72%, and 73%. 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, Sequence identity percentages of 94%, 95%, 96%, 97%, 98%, 99%, or 100% It contains. Sufficient homology is determined by the length of the polynucleotide and the overall sequence identity pars. The sequence identity percentage is determined by selecting a conserved region or local sequence identity percentage of a sequence of nucleotides. This includes any combination of, for example, sufficient homology is less for the region of the target gene locus. It can be described as a region of 75-150 bp with at least 80% sequence identity. The significant homology also indicates that the two particles specifically hybridize under high stringency conditions. This can be explained by the predictive power of renucleotides, for example, Sambrook e t al.,(1989)Molecular Cloning: A Laborato ry Manual, (Cold Spring Harbor Laboratory Press, NY);Current Protocols in Molecule r Biology,Ausubel et al.,Eds(1994)Curren t Protocols, (Greene Publishing Associate s, Inc. and John Wiley & Sons, Inc.); and Tijs sen(1993) Laboratory Techniques in Bioche mistry and Molecular Biology--Hybridizat ion with Nucleic Acid Probes,(Elsevier,N See (New York).
[0631] Episomal DNA molecules also ligate within double-strand breaks, for example, chromosome 2 T-DNA can be incorporated into the main strand break (Chilton and Que, 2003)Plant Physiol 133:956-65;Salomon an d Puchta, (1998) EMBO J 17:6086-95). double strand break The surrounding sequence is modified, for example, by exonuclease activity involved in the maturation of double-strand breaks. When altered, the gene conversion pathway is affected, for example, by non-divisible somatic cells or sister staining after DNA replication. If homologous sequences such as split cells can be used, the original structure can be recovered (Molin ier et al., (2004) Plant Cell 16:342-52). different Local and / or epigenetic DNA sequences can also be used as DNA repair templates for homologous recombination. It can function (Puchta, (1999) Genetics 152:11) 73-81).
[0632] In one embodiment, the disclosure includes a method for editing nucleotide sequences in the genome of a cell. This method involves at least one PGEN described herein, and polynucleotide modification. This includes introducing into a template (where the polynucleotide modified template is (including at least one nucleotide modification of the nucleotide sequence), optionally edited The process further includes selecting at least one cell containing the collected nucleotide sequence.
[0633] The guide polynucleotide / Cas endonuclease system targets the target genome nucleotide. To enable editing (modification) of the ocidal sequence, at least one polynucleotide modification is required. It can be used in combination with templates. (Published in the US on March 19, 2015) Patent application publication No. 2015 / 0082478 and the country published on February 26, 2015 (See also the brochure No. 2015 / 026886.)
[0634] The target polynucleotide and / or trait is from International Publication No. 20, published on September 27, 2012. Pamphlet No. 12 / 129373, and International Release No. 2013, published August 1, 2013. As described in pamphlet No. 112686, they are stacked together at complex trait loci. This is possible. The guide polynucleotide / Cas9 endonuclear described herein The Xe system provides a system for efficiently generating double-strand breaks at complex trait loci. It enables the accumulation of traits.
[0635] Guide polynucleotides / Ca as described herein mediate gene targeting. The s system is a method for inducing heterogeneous gene insertion, and / or a complex comprising multiple heterogenes. In a method for generating trait loci, the International Publication No. 2, published on September 27, 2012, describes the method for generating trait loci. It may be used in the same manner as disclosed in pamphlet No. 012 / 129373. In this method, instead of using double-strand break inducers to introduce the target gene, The guide polynucleotide / Cas system disclosed herein is used. The input genes are separated by 0.1, 0.2, 0.3, 0.4, 0.5, 1.0, 2, or 5 centimeters. By inserting it within the thymorgan (cM) region, the introduced gene can be propagated as a single locus. It can be done (for example, U.S. Patent Application Publication No. 2, published on October 3, 2013). Specification No. 013 / 0263324 or International Publication No. 20 published on March 14, 2013 (Please refer to pamphlet No. 12 / 129373). After selecting a plant containing the introduced gene... Cross plants containing (at least) one introduced gene, and produce F1 hybrids containing both introduced genes. It can be formed. Of the offspring from these F1 (F2 or BC1), 1 / 500 The offspring will have two different transgenes recombined on the same chromosome. Subsequently, this composite locus can reproduce as a single locus along with both introduced traits. By repeating this process, any desired trait can be accumulated.
[0636] Further use of guide RNA / Cas endonuclease systems is described ( For example, U.S. Patent Application Publication No. 2015 / 008247, published on March 19, 2015. Specification No. 8, International Publication No. 2015 / 026886, published on February 26, 2015. Nfret, U.S. Patent Application Publication No. 2015 / 0059, published on February 26, 2015. Specification No. 010, International Publication No. 2016 / 007347, published on January 14, 2016. This information is from the brochure and the International Publication No. 201 of the PCT application, published on February 18, 2016. (Please refer to the 6 / 025131 pamphlet), which contains the target nucleotide sequence Modification or substitution of regulatory elements, insertion of target polynucleotides, gene knocking Out, gene knock-in, modification of splicing site and / or alternative splicing site Introduction of positions, modification of the nucleotide sequence encoding the target protein, amino acids and / or Genetic fusion, and the expression of reverse repeats in the target gene, This includes, but is not limited to, lending.
[0637] The characteristics obtained from the gene editing compositions and methods described herein can be evaluated. Chromosome spacings that correlate with the target phenotype or trait can be identified. To identify it, various methods well known in the relevant field can be used. These chromosome boundaries will be linked to genes that control the desired trait. It is depicted so as to encompass the marker. In other words, the chromosome spacing is the spacing (spacing) Any marker present within (including terminal markers that define boundaries) of a specific trait It is determined so that it can be used as a marker. In one embodiment, the chromosome spacing is at least It may contain one more QTL, and in fact, more than one QTL within the same interval. When multiple QTLs are very close together, one marker can be linked to more than one QTL. Therefore, the relationship between a specific marker and a specific QTL may become unclear. Conversely, for example, non If two markers that are always in close proximity are co-separated from the desired phenotypic trait, then these Whether each marker identifies the same QTL, or two different QTLs. , often becoming unclear. The term “quantitative trait locus” or “QTL” refers to at least one The genetic background, for example, the expression of quantitative phenotypic traits in at least one breeding population. This refers to the DNA region associated with the difference. QTL regions contain genes that influence the trait in question. or is closely associated with such genes. "QTL alleles" are haplotypes. A gene or other genetic element containing multiple genes or other genetic factors within a continuous genomic region or linkage group, such as a p. It is possible that a QTL allele can mean a haplotype within a specific window. The window may be defined by one or more sets of polymorphic markers, and A haplotype is a continuous genomic region that can be tracked. The car position can be defined by the allele's unique fingerprint.
[0638] Introduction of CRISPR-Cas system components into cells The methods and compositions disclosed herein are for specific methods of introducing sequences into living organisms or cells. It is not dependent on the presence of polynucleotides or polynucleotides inside at least one cell of this organism. The process simply involves introducing the lipeptide. The introduction involves integrating the nucleic acid into the cell's genome. This includes references to the incorporation of nucleic acids into eukaryotic or prokaryotic cells, and also includes references to nucleic acids, The transfer of proteins or polynucleotide-protein complexes (PGEN, RGEN) to cells. This includes references to transient (direct) provision.
[0639] Polynucleotides or polypeptides or polynucleotide-protein complexes Methods for introducing into cells or organisms are known in the art, and microinjection , electroporation, stable transformation method, transient transformation method, ballistic particle acceleration method (particles) (Impact method), whisker-mediated transformation, Agrobacterium genus m) Veterinary transformation, direct gene transfer, viral-mediated transfer, transfection, trait transfer Cell-permeable peptides, mesoporous silica nanoparticles (MSNs) mediated direct protein delivery. This includes, local application, sexual hybridization, sexual reproduction, and any combination thereof. It is not limited to them.
[0640] For example, guide polynucleotides (guide RNA, crnucleotide + tracr Nucleotides, guide DNA and / or guide RNA-DNA molecules) are single-stranded or double-stranded. It can be directly (transiently) introduced into cells as a chain polynucleotide molecule. Guide RNA (also (crRNA + tracrRNA) also exists within the cell, and within the cell, guide RNA (cr Activated by specific promoters capable of transcribing RNA (tracrRNA molecules). A heterogeneous nucleus that encodes a guide RNA (or crRNA + tracrRNA) linked to it. This can be indirectly introduced by introducing recombinant DNA molecules containing acid fragments. . Specific promoters are not limited, but are strictly defined and unaltered 5' end. RNA polymerase III promo enables transcription of RNA with terminal and 3' ends. It could be a ter (Ma et al., 2014, Mol.Ther.Nucleic Acids 3:e161;DiCarlo et al.,2013,Nucleic Acids Res. 41:4336-4343; Published February 26, 2015 (International Publication No. 2015 / 026887 pamphlet). Transcribing guide RNA within cells. Any promoter can be used, which includes calling the guide RNA Includes a thermal shock / thermal induction promoter operably linked to the nucleotide sequence. Born.
[0641] Plant cells are similar to animal cells (such as human cells), fungal cells (such as yeast cells), and protoplastic cells. Unlike plants, for example, plant cells contain a plant cell wall that can act as a barrier to the delivery of components. nothing.
[0642] Cas endonuclease, and / or guide RNA, and / or ribonucleoprotein compound Combination and / or transfer of polynucleotides encoding one or more of the aforementioned into plant cells Delivery is by methods known in the art, for example, but not limited to, the Rhizobium group (Rhizobiales)-mediated transformation (e.g., Agrobacterium genus (Agrob acterium), Ochrobacterium, particle-mediated transport Particle impact method (PEG), polyethylene glycol (PEG) (for example, onto protoplasts) ) mediated transfection, electroporation, cell-permeable peptides, or meso This can be achieved by porous silica nanoparticle (MSN)-mediated direct protein delivery. ru.
[0643] Cas endonucleases such as the Cas endonucleases described herein are Ca The s polypeptide itself (called direct delivery of Cas endonuclease), Cas protein mRNA encoding a nucleotide, and / or guide polynucleotide / Cas endonucleus. By directly introducing the enzyme complex itself using any method known in the art, Cas endonuclease can be introduced into cells. By introducing a recombinant DNA molecule encoding rease into the cell, it is introduced indirectly. Endonucleases can be processed using any method known in the art. It can be transiently introduced into cells or integrated into the host cell's genome. The uptake of ases and / or derived polynucleotides into cells was observed on May 12, 2016. As described in the published international publication No. 2016 / 073433, cells This can be promoted using permeable peptides (CPPs). Cas endonuclei in cells Any promoter capable of expressing the enzyme can be used, including C A thermal shock absorber operably linked to the nucleotide sequence encoding the as endonuclease This includes a thermally inductive promoter.
[0644] Direct delivery of polynucleotide modification templates to plant cells is possible via particle-mediated delivery. It can be achieved, and is not limited to, polyethylene glycol to protoplasts ( PEG-mediated transfection, whisker-mediated transformation, electroporation , particle impact method, cell-permeable peptide, or mesoporous silica nanoparticle (MSN) mediated direct Other direct delivery methods such as contact protein delivery, in polypropylene in eukaryotic cells such as plant cells. It can be successfully used for the delivery of nucleotide modification templates.
[0645] Donor DNA can be introduced by any means known in the art. It is possible. Donor DNA is, for example, from the genus Agrobacterium. Transmutation or bioristic particle impact methods, among others, are known in the art. It can be provided by a transformation method of choice. Donor DNA is transiently present in cells. It may be done, and it could also be introduced by a virus replica. In the presence of crease and the target site, donor DNA is inserted into the transformed plant genome. It can be done.
[0646] Direct delivery of any one of the inducible Cas system components is guided by polynucleotides / Ca To enhance the enrichment and / or visualization of cells that receive s-endonuclease complex components. This can be accompanied by the direct delivery (co-delivery) of other mRNAs. For example, guide polymerase Cleotide / Cas endonuclease component (and / or guide polynucleotide / Ca The s-endonuclease complex itself) is used as a phenotypic marker (not limited to, but CRC(B) ruce et al.2000 The Plant Cell 12:65-79) By directly co-delivering together with mRNA encoding transcription activators (such as), In the international publication brochure No. 2017 / 070032, released on April 27, 2017... As described, the function of non-functional gene products is restored, and exogenous selective markers are used. This allows for the selection and enrichment of cells without the need for cell selection.
[0647] The guide RNA / Cas endonuclease complex described herein (as described herein) To introduce the cleavage-ready complex (represented by the image) into cells, the individual components of the complex are separated. Individually or collectively into cells, directly (RNA and Cas endonuclease for guidance) Direct delivery as proteins and protein subunits, or functional fragments thereof. ), or recombinant constructs expressing components (guide RNA, Cas endonuclease, tan This includes introduction by protein subunits (or functional fragments thereof). For the introduction of the doRNA / Cas endonuclease complex (RGEN) into cells, ribonuclide The rheotide protein guides the guide RNA / Cas endonuclease complex into the cell. This includes the inclusion of ribonucleotide-proteins as described herein. It can be assembled before being introduced into the cell. Guide RNA / Cas endonuclease Components that make up a bonucleotide protein (at least one Cas endonuclease) (at least one guide RNA, at least one protein subunit) In vitro or introduced into cells (targeted by genome modifications as described herein) It can be assembled by any means known in the art beforehand.
[0648] Direct delivery of RGEN ribonucleoprotein is followed by rapid degradation of the complex within the cell. Genome editing at target sites of the cell genome in a manner that makes the presence of the complex temporary. This makes it possible. The transient presence of this RGEN complex leads to a reduction in off-target effects. This is possible. In contrast, the RGEN component (guide RNA, Cas) determined by the plasmid DNA sequence... Delivery of 9 endonucleases leads to sustained expression of RGEN from these plasmids. This can lead to an increase in off-target effects (Cradick, TJet al.) (2013) Nucleic Acids Res 41:9584-9592;Fu ,Y et al.(2014)Nat.Biotechnol.31:822-826 ).
[0649] Direct delivery is via guide RNA / Cas endonuclease complex (RGEN) (as specified herein). Any one component (e.g., at least one) of the cleavage-ready complex described in A guide RNA, at least one Cas protein, and optionally one further Proteins) are used to create microparticles (not limited to, but including gold particles, tungsten particles, and silicon carbide). This can be achieved by combining it with a delivery matrix containing elementary whisker particles, etc. (This is possible (International Publication No. 2017 / 070032, published on April 27, 2017) (See also the documentation). The delivery matrix is any one of the components, e.g., Cas. It may also contain a ndonuclease, which adheres to a solid matrix (e.g., impact particles). He is wearing it.
[0650] In one embodiment, the guide polynucleotide / Cas endonuclease complex is guide R Guide RNA and Cas endonuclease that form the NA / Cas endonuclease complex This is a complex in which ase proteins are introduced into cells as RNA and protein, respectively. .
[0651] In one embodiment, the guide polynucleotide / Cas endonuclease complex is guide R Guide RNA and Cas endonuclease that form the NA / Cas endonuclease complex The ase protein, and at least one protein subunit of the complex, It is a complex that is introduced into cells as RNA and protein.
[0652] In one embodiment, the guide polynucleotide / Cas endonuclease complex is guide R Guide RN that forms NA and Cas endonuclease complex (ready-to-cleave complex) A and Cas endonuclease proteins, and at least one protein of the complex The quality subunit is pre-assembled in vitro, forming a ribonucleotide-protein complex. It is a complex that is introduced into cells.
[0653] Polynucleotides, polypeptides, or polynucleotide-protein complexes (PGENs) Protocols for introducing RGEN into plants or eukaryotic cells such as plant cells are not known. Microinjection (Crossway et al., (1986)B iotechniques 4:320-34 and U.S. Patent No. 6,300,543 (Book), meristematic tissue transformation (U.S. Patent No. 5,736,369), electroporation Riggs et al., (1986) Proc. Natl. Acad. Sc i.USA 83:5602-6, Agrobacterium genus ) Transformation by vector (U.S. Patent No. 5,563,055 and No. 5,981,840) (Specification), Whisker-mediated transformation (Ainley et al. 2013, Plant Biotechnology Journal 11:1126-1134;Shah een A.and M. Arshad 2011 Properties and Applications of Silicon Carbide(2011),34 5-358 Editor(s):Gerhardt,Rosario.Publish er:InTech,Rijeka,Croatia.CODEN:69PQBP;IS BN: 978-953-307-201-2), direct hereditary import (Paszkowski) et al., (1984) EMBO J 3:2717-22), and particle acceleration French (US Patent No. 4,945,050; same as US Patent No. 5,879,918; same as US Patent No. 5) Inventory No. 886,244; same as Inventory No. 5,932,782; Tomes et al ., (1995) “Direct DNA Transfer into Intact Plant Cells via Microprojectile Bombard ment” in Plant Cell,Tissue,and Organ Cul ture:Fundamental Methods,ed.Gamborg&Phil lips(Springer-Verlag,Berlin);McCabe et a l., (1988) Biotechnology 6:923-6; Weissinge r et al.,(1988)Ann Rev Genet 22:421-77;S anford et al.,(1987)Particulate Science and Technology5:27-37(タマネギ);Christou et al.,(1988)Plant Physiol87:671-4(ダイズ);Fin er and McMullen,(1991)In vitro Cell Dev Biol 27P:175-82 (Soybeans); Singh et al., (1998) Theor Appl Genet 96:319-24 (Soybeans); Datta et al., (1990) Biotechnology 8:736-40 (Rice); Kl ein et al.,(1988)Proc.Natl.Acad.Sci.USA 85:4305-9 (Corn); Klein et al., (1988) Bio Technology 6:559-63 (Corn); U.S. 5,240,8 Specification No. 55; Specifications No. 5,322,783 and No. 5,324,646; Klein et al.,(1988)Plant Physiol 91:440- 4 (Corn); Fromm et al., (1990) Biotechnolo gy 8:833-9 (corn); Hooykaas-Van Slogtere n et al., (1984) Nature 311:763-4; U.S. 5,7 Specification No. 36,369 (cereals); Bytebier et al., (1987) Pr oc.Natl.Acad.Sci.USA 84:5345-9(Liliaceae) ceae));De Wet et al,(1985)in The Experim ental Manipulation of Ovule Tissues, ed.C hapman et al.,(Longman, New York),pp.197- 209(Pollen);Kaeppler et al.,(1990)Plant Cell Rep 9:415-8) and Kaeppler et al., (1992) The or Appl Genet 84:560-6 (Whisker-mediated transformation); D'Ha lluin et al.,(1992)Plant Cell4:1495-505( Electroporation; Li et al., (1993) Plant Cell Rep 12:250-5;Christou and Ford(1995)Anna ls Botany 75:407-13 (rice) and Osjoda et al., (1996) Nat Biotechnol 14:745-50 (Agrobacterium • via Agrobacterium tumefaciens It contains sorghum.
[0654] Alternatively, polynucleotides can bring cells or organisms into contact with viruses or viral nucleic acids. This allows for introduction into plants or plant cells. Generally, such methods involve poly This involves incorporating nucleotides into viral DNA or RNA molecules. In this example, the target polypeptide was first synthesized as part of a viral polyprotein. This is then subjected to proteolytic degradation in vivo or in vitro to obtain the desired recombinant protein. It can produce a substance. Polynucleotides containing viral DNA or RNA molecules can be introduced into the plant body. Methods for introducing and expressing the encoded protein are known, for example, U.S. Patent No. 5,889,191, U.S. Patent No. 5,889,190, Specification No. 66,785, Specification No. 5,589,367, and Specification No. 5,316,931 Please refer to the specifications.
[0655] Polynucleotides or recombinant DNA constructs were transformed using various transient transformation methods, resulting in prokaryotic transformations. and can be provided to or introduced into eukaryotic cells or organisms. Such transient transformation methods are limited This does not involve the direct introduction of polynucleotide constructs into plants, but it does include that.
[0656] Removal of any or all components (proteins and / or nucleic acids) of the inducible Cas system Methods using molecules that facilitate absorption, such as cell-permeable peptides and nanocarriers. Nucleic acids and proteins can be supplied to cells by any method, including 2011. The specification of U.S. Patent Application Publication No. 2011 / 0035836, published on the 10th of the month, and 20 See also European Patent Application Publication No. 2821486A1, published on January 7, 2015. I want to be treated that way.
[0657] Others who introduce polynucleotides into prokaryotic and eukaryotic cells or parts of organisms or plants. Methods such as plastid transformation and the extraction of polynucleotides from seedlings or mature seeds into tissues. You can use the method of implementing it.
[0658] Stable transformation is the process in which a nucleotide construct introduced into an organism is integrated into the organism's genome. It is intended to mean that it can be passed down to its offspring. Transient transformation is, Cleotides are introduced into organisms, but they are either not incorporated into the organism's genome or are polypeptides. This is intended to mean that the drug is introduced into the organism. Transient transformation is introduced This means that the composition is expressed or present only temporarily within the organism.
[0659] Modified genome at or near the target site without using screening marker phenotypes Various methods can be used to identify those cells that possess this characteristic. This method can be considered a direct analysis of the target sequence to detect changes within the target sequence. The following are some, but are not limited to: PCR method, sequencing method, nuclease digestion method, Southern Block Examples include the T method and any combination thereof.
[0660] Cells and plants The polynucleotides and polypeptides of this disclosure can be introduced into cells. This includes, but is not limited to, humans, non-humans, animals, mammals, bacteria, fungi, insects, Yeast, unconventional yeast and plant cells, and plants prepared by the method described herein. Examples include objects and seeds. Monocotyledonous plants and dicotyledonous plants, as well as any plant containing plant elements. The substance can be used in conjunction with the compositions and methods described herein.
[0661] Examples of usable monocotyledonous plants include, but are not limited to, corn (there is a type of corn). (Zea mays), rice (Oryza sativa), La Imugi (Secale cereale), sorghum (sorghum) Sorghum bicolor, Sorghum vulgaris vulgare), barnyard millet (for example, pearl millet, Pennisetum glaucum (Penn isetum glaucum), Kibi (Panicum myriaceum (Panicum miliaceum), Awa (Setaria italica) Finger millet (Eleusine coracana), Wheat (species of the genus Triticum, e.g., Triticum aestibum) Triticum aestivum, Triticum monococum monococcum), sugarcane (species of the genus Saccharum (Saccharum s pp.)), wild oats (Avena genus), barley (Hordeum genus) ordeum), switchgrass (Panicum virgatum) gatum), pineapple (Ananas comosus) ), bananas (species of the genus Musa (Musa spp.)), palms, ornamental plants, turfgrass and other grasses It can be listed.
[0662] Examples of dicotyledonous plants that can be used include, but are not limited to, soybeans (glycine • Glycine max, a species of the Brassica genus (for example) However, it is not limited to Brassica napus (Brassica napus) a napus), B. campestris, Brassica rapa (Brassica rapa), Brassica.junc ea)), alfalfa (Medicago sativa) ), tobacco (Nicotiana tabacum), Arabic Dopsis (Arabidopsis thaliana) , sunflower (Helianthus annuus), wa Gossypium arboreum, Gossypium • Barbadense (Gossypium barbadense), and peanut (A Arachis hypogaea, tomato (Solanum lycoides) Solanum lycopersicum, potato (Solanum... One example is Solanum tuberosum.
[0663] Further plants that can be used include safflower (Carthamus tinctorius). Sweet potato (Carthamus tinctorius), sweet potato (Ipomoea batatas) Cassava (Ipomoea batatus), Cassava (Manihot esculenta) Manihot esculenta, coffee (Coffea species) , coconut (Cocos nucifera), citrus tree ( Citrus species), cocoa (Theobroma c acao)), tea plant (Camellia sinensis) )), banana (Musa genus), avocado (Persea americana (Persea americana) Ficus americana), Fig (Ficus casi ca)), guava (Psidium guajava), man Goat (Mangifera indica), olive ( Olea europaea, papaya (Calica papaya (C) arica papaya), cashew (anacardium occidental (Anac (Ardiia occidentale), Macadamia (Macadamia integrifolia) (Macadamia integrifolia), almond (Prunus a) Migdalus (Prunus amygdalus), sugar beet (Betta bulgaricus) Examples include sorghum (Beta vulgaris), vegetables, ornamental plants, and conifers.
[0664] Vegetables that can be used include tomatoes (Lycopersicon esculentum), lettuce (for example, lettuce (Lactuca sativa) ), sayamame (green bean (Phaseolus vulgaris)), lima bean ( Lima bean (Phaseolus limensis), pea (Lathyrus genus) Lathyrus species, as well as cucumber (Cucumber (C. sativus)), can Talop (C. cantalupensis) and muskme Members of the genus Cucumis, such as C. melo, are mentioned. It can be grown. As an ornamental plant, azalea (Rhododendron genus) n) species), hydrangea (Macrophylla hyd rangea)), Hibiscus (Hibiscus rosa sinensis (Hibiscus Rosa sanensis), rose (Rosa genus), tulip (Tulip) -Tulipa species), trumpet daffodil (Narcissus genus) (seed), petunia (petunia hybrida), ca Dianthus caryophyll us)), Poinsettia (Euphorbia pulcherrima (Euphorbia pulc Examples include herrima and chrysanthemum.
[0665] Coniferous trees that can be used include the loblolly pine (Pinus taeda). aeda)), Slash pine (Pinus ellioti) i)) Ponderosa pine (Pinus ponderosa) ), Lodgepole pine (Pinus contorta), and Pine trees such as the Monterey pine (Pinus radiata); Douglas fir (Pseudotsuga menziesi) ii) American hemlock (Tsuga canadensis) ;Spear (Picea glauca); American cedar (Cedar) Sequoia sempervirens (European mo Abies amabilis and balsam fir (Abies amabilis) Abies balsamea and other fir trees; and Western red cedar. (Thuja plicata) and Alaska yellow cedar ( Chamaecyparis nootkatensis Examples include the Himalayan cedar (Cedar sis)).
[0666] In certain embodiments of this disclosure, fertile plants produce viable male and female gametes. It is a plant, and it is a self-fertile plant. Such self-fertile plants do not use gametes from other plants, or It is possible to produce offspring plants without relying on the contribution of the genetic material contained therein. Other applications of this disclosure The application method involves the plant producing viable, or otherwise fertilizable, male or female gametes. Because they do not produce gametes or both, non-self-fertile plants are sometimes used.
[0667] This disclosure is used in the breeding of plants that include one or more introduced traits.
[0668] How can two traits be stacked in the genome, for example, at a genetic distance of 5 cM from each other? A non-restrictive example of this is described as follows: the first DSB mark within the genome window. It includes a first transgenic target site integrated into the target site, and the first target genome inheritance The first plant, which does not have stromas, has the target genome inserted into different genomic insertion sites within the genome window. The first plant is crossed with a second transgenic plant containing the chromosome locus. The second plant is then crossed with the first chromosome. It does not contain lancegenic target sites. Approximately 5% of plant offspring from this cross are first DS The first transgenic target site integrated into the B target site, and within the genome window It has both the first target genomic locus integrated into different genomic insertion sites. Descendant plants having both sites within the defined genome window are incorporated into the second DSB target site. The embedded second transgenic target site and / or within the defined genomic window It includes a second target genome locus, and the first transgenic target site and the first It can be further crossbred with a third transgenic plant that lacks the target genome locus. Following this, the first transgenic target site, the first target genomic gene locus, and the genome... It has a second target genome locus integrated into a different genomic insertion site within the window. Select offspring. Using this method, select at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 or more DSB pairs Transgenic target sites incorporated into elephant regions and / or different within the genome window Transgenic gene containing a complex trait locus having a target genome locus incorporated into the site Nick plants can be created. In this way, various complex trait loci can be generated. It is possible.
[0669] Cells and animals The polynucleotides and polypeptides of this disclosure can be introduced into animal cells. As for phytocellular organisms, they are not limited to the following, but include chordates, arthropods, mollusks, annelids, Organisms of the phylum Cnidarians or Echinoderms; or mammals, insects, birds, amphibians, reptiles This includes organisms of the class, or the class that includes fish. In some embodiments, animals include humans, mammoths, and mammoths. Rat, C. elegans, rat, fruit fly (Drosophila) Drosophila species, zebrafish, chickens, dogs, cats, mo Lumot, hamster, chicken, chicken, southern medaka, sea lamprey, pufferfish, Amaga This refers to animals such as monkeys (e.g., species of the genus Xenopus), monkeys, or chimpanzees. The specific cell types shown include haploid cells, diploid cells, germ cells, neurons, and muscle cells. Measuring cells, endocrine or exocrine cells, epithelial cells, muscle cells, tumor cells, embryonic cells (embryo nic cells), hematopoietic cells, bone cells, germ cells, somatic cells, stem cells Examples include cells, pluripotent stem cells, induced pluripotent stem cells, progenitor cells, meiotic cells, and mitotic cells. In some embodiments, multiple cells of biological origin may be used.
[0670] The disclosed novel Cas9 orthologues can be used to edit the genomes of animal cells in various ways. It can be used for this purpose. In one embodiment, it may be desirable to delete one or more nucleotides. In another embodiment, it may be desirable to insert one or more nucleotides. It may be desirable to substitute one or more nucleotides. In another embodiment, another atom or It modifies one or more nucleotides through covalent or non-covalent interactions with molecules. It may be desirable to do so.
[0671] Genome modification with Cas9 orthologs alters the genotype and / or phenotype of the target organism. It can be used to bring about such a change. Such a change is preferably the target phenotype or raw Improvement of scientifically important characteristics, correction of intrinsic defects, or expression of certain types of expression markers. Related to this. In some aspects, the target phenotype or physiologically important trait is the whole of the animal. The animal's health, adaptability or reproductive capacity, its ecological adaptability, or its movement in its environment. It relates to the relationship or interaction betwee...
Claims
1. (a) (i) C-terminal trisplit RuvC domain, (ii) The following amino acid motifs: GxxxG, ExL, Cx n C, and Cx n (C or H) (where G = glycine, E = glutamic acid, C = cysteine, H = histidine, x = Any amino acid, and n = an integer from 0 to 11, (iii) Alpha helix, and (iv) Multiple beta sheets forming a wedge-shaped domain Cas endonuclease containing; (b) Target double-stranded DNA polynucleotide (wherein the target double-stranded DNA polynucleotide) (Cydo is a different species from the aforementioned Cas endonuclease source), and (c) Variable targeting including a region complementary to the target double-stranded DNA polynucleotide. Guide polynucleotides containing the g domain Includes: The Cas endonuclease is the PAM sequence on the target double-stranded DNA polynucleotide. The guide polynucleotide and the Cas endonuclease recognize the mark It forms a complex that binds to double-stranded DNA polynucleotides. Synthetic composition.
2. The Cas endonuclease described in claim 1 contains fewer than 800 amino acids. Synthetic composition.
3. The Cas endonuclease is a polynuclease encoding the Cas endonuclease. The synthetic composition according to claim 1, provided as a creotide.
4. The aforementioned Cas endonuclease cleaves the double-stranded DNA polynucleotide. The synthetic composition described in item 1.
5. The synthetic composition according to claim 1, further comprising heterogeneous polynucleotides.
6. The synthetic composition according to claim 6, wherein the heterogeneous polynucleotide is an expression element.
7. The synthetic composition according to claim 6, wherein the heterogeneous polynucleotide is a transgene.
8. The synthetic composition according to claim 6, wherein the heterogeneous polynucleotide is a donor DNA molecule. 。
9. The heterogeneous polynucleotide is a polynucleotide modification template, as described in claim 6. The synthetic composition shown.
10. The CRISPR-Cas endonuclease is catalytically inactive, as described in claim 1. A synthetic composition of [the substance].
11. The Cas endonuclease recognizes a PAM sequence containing multiple T or C nucleotides. The synthetic composition according to claim 1.
12. The PAM arrays are TTAT, TTTR, N(T>V)TTR, N(W>S)TTTR , N(Y>R)N(Y>S>R)TTN(A>G>Y), N(W>S)N(Y>R)TT A claim selected from the group consisting of TR, CTT, N(T>W>C)TTC, and CCD. The synthetic composition described in 10.
13. The Cas endonuclease is part of the fusion protein as described in claim 1. composition.
14. The synthetic composition according to claim 11, further comprising a deaminase.
15. The fusion protein further comprises heterologous nuclease domains, according to claim 11. composition.
16. The synthetic composition according to claim 1, further comprising eukaryotic cells.
17. The synthesis according to claim 16, wherein the eukaryotic cell is a plant cell, an animal cell, or a fungal cell. composition.
18. The plant cell is a monocotyledonous plant cell or a dicotyledonous plant cell, according to claim 17. composition.
19. The aforementioned plant cells are from corn, soybeans, cotton, wheat, canola, rapeseed, moro. Koshi rice, rice, rye, barley, millet, wild oats, sugarcane, turfgrass, switchgrass Su, alfalfa, sunflower, tobacco, peanut, potato, Arabidopsis (A It is of biological origin selected from the group consisting of *Rabidopsis*, safflower, and tomato. The synthetic composition according to claim 17.
20. Polynucleotide encoding Cas endonuclease of the synthetic composition according to claim 1 Do.
21. The polynucleotide according to claim 20, further comprising at least one further polynucleotide Cleotide.
22. Claim 2, the at least one further polynucleotide is an expression element. The polynucleotide described in 1.
23. The at least one further polynucleotide is a gene, as per claim 21. Polynucleotides.
24. The synthetic combination according to claim 1, wherein at least one component is attached to a solid matrix. Finished product.
25. (a) Sequence IDs 17, 18, 19, 20, 32, 33, 34, 35, 36, 37, 38 、254、255、256、257、258、259、260、261、262、263 、264、265、266、267、268、269、270、271、272、273 、274、275、276、277、278、279、280、281、282、283 、284、285、286、287、288、289、290、291、292、293 、294、295、296、297、298、299、300、301、302、303 、304、305、306、307、308、309、310、311、312、313 、314、315、316、317、318、319、320、321、322、323 、324、325、326、327、328、329、330、331、332、333 、334、335、336、337、338、339、340、341、342、343 、344、345、346、347、348、349、350、351、352、353 、354、355、356、357、358、359、360、361、362、363 A group consisting of 364, 365, 366, 367, 368, 369, 370, and 371. A Cas endonuclease, or its mechanism, that is at least 80% identical to the selected sequence. A functional fragment or variant; (b) Target double-stranded DNA polynucleotide (wherein the target double-stranded DNA polynucleotide) (Cydo is different from the aforementioned Cas endonuclease source); and (c) Variable targeting including a region complementary to the target double-stranded DNA polynucleotide. Guide polynucleotides containing the g domain Includes: The Cas endonuclease is the PAM sequence on the target double-stranded DNA polynucleotide. Recognizing the target, the guide polynucleotide and the Cas endonuclease target two It forms a complex that binds to the main strand of DNA polynucleotides. Synthetic composition.
26. A method for introducing targeted editing into a target polynucleotide, (a) (i) C-terminal trisplit RuvC domain, (ii) The following amino acid motifs: GxxxG, ExL, Cx n C, and Cx n (C or H) (where G = glycine, E = glutamic acid, C = cysteine, H = histidine, x = Any amino acid, and n = an integer from 0 to 11, (iii) Alpha helix, and (iv) Multiple beta sheets forming a wedge-shaped domain Cas endonuclease containing (Here, the Cas endonuclease performs the PAM sequence on the target polynucleotide) (to recognize); (b) A variable targeting domain substantially complementary to a portion of the target polynucleotide Guide polynucleotide containing yl (wherein the guide polynucleotide and Cas- Alpha-endonuclease can recognize and bind to the target polynucleotide. (Forming a complex) A method comprising providing a heterogeneous composition containing the above.
27. The composition further includes cells, and before introducing the heterologous composition, the genotype of at least one biological cell In addition, at least one nucleotide modification is introduced compared to the target sequence of the cell's genome. Further including, incubating the cells, and generating a whole organism from the cells, Before introducing the heterologous composition, the genome of at least one cell of the organism contains the heterologous composition. The presence of at least one nucleotide modification compared to the target sequence of the cell genome indicates that The method according to claim 26, further comprising confirming that
28. The Cas endonuclease recognizes a PAM sequence containing multiple T or C nucleotides. The method according to claim 26.
29. The PAM arrays are TTAT, TTTR, N(T>V)TTR, N(W>S)TTTR , N(Y>R)N(Y>S>R)TTN(A>G>Y), N(W>S)N(Y>R)TT A claim selected from the group consisting of TR, CTT, N(T>W>C)TTC, and CCD. The method described in 26.
30. The method according to claim 27, wherein the cell is a eukaryotic cell.
31. The eukaryotic cells are derived from or obtained from animals, fungi, or plants, according to the claim. The method described in 30.
32. The method according to claim 31, wherein the plant is a monocotyledonous plant or a dicotyledonous plant.
33. The aforementioned plants are corn, soybeans, cotton, wheat, canola, rapeseed, and sorghum. Rice, rye, barley, millet, wild oats, sugarcane, turfgrass, switchgrass, Alfalfa, sunflower, tobacco, peanuts, potato, Arabidopsis (Ara Selected from the group consisting of bidopsis, safflower, and tomato, as described in claim 31. Method of loading.
34. The method according to claim 27, further comprising introducing heterologous polynucleotides.
35. The method according to claim 34, wherein the heterogeneous polynucleotide is a donor DNA molecule.
36. The heterologous polynucleotide contains a sequence that is at least 50% identical to the sequence of the cell. The method according to claim 34, wherein the polynucleotide modification template is...
37. Offspring of an organism obtained by the method of claim 27, comprising at least one cell A progeny that retains at least one of the aforementioned nucleotide modifications.