Protein related to soybean plant height, pod number per plant and grain number per plant, and biological material and application thereof
By regulating the expression of genes encoding GmCOL2a or GmCOL2b proteins using CRISPR/Cas9 technology, soybean plant height was reduced and the number of pods and seeds per plant was increased, solving the problem of plant type improvement in soybean breeding and enhancing lodging resistance and yield.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INSTITUTE OF CROP SCIENCE CHINESE ACADEMY OF AGRICULTURAL SCIENCES
- Filing Date
- 2023-11-08
- Publication Date
- 2026-07-21
Smart Images

Figure HDA0004537692550000011 
Figure HDA0004537692550000012 
Figure HDA0004537692550000021
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology, specifically relating to proteins and their biomaterials and applications related to soybean plant height, number of pods per plant, and number of grains per plant. Background Technology
[0002] Soybeans are an important dual-purpose crop for both food and oil, and a major source of high-quality protein for humans and livestock feed. With the global population continuing to grow at an unprecedented rate, the demand for food is surging, making increasing crop yields a pressing challenge for modern crop breeding technology. Soybeans, corn, and wheat are the four main food crops that meet the dietary needs of most of the world's population. Through research efforts aimed at significantly increasing crop yields, researchers have made significant discoveries about the advantages of semi-dwarf crops. These crops possess many excellent characteristics, such as enhanced lodging resistance, efficient fertilizer use, and improved light energy utilization. Furthermore, semi-dwarf crops generally have higher resistance to various plant diseases and pests, which to some extent ensures stable yields under high-density planting conditions. Soybeans are an important dual-purpose crop for both food and oil, and a major source of high-quality protein for humans and livestock feed. Soybean pods mainly grow on stem nodes. Plant height, number of pods per plant, and number of seeds per plant are key traits determining lodging resistance and yield, and are important targets for breeding new high-yielding soybean varieties with ideal plant types. Because there are few reports on genes and mechanisms of action that can simultaneously improve soybean plant height, number of pods per plant, and number of seeds, the molecular breeding of ideal soybean plant architecture is limited. Therefore, the ability to simultaneously reduce soybean plant height and increase the number of pods and seeds per plant is of significant theoretical and practical importance for improving soybean lodging resistance and yield. Summary of the Invention
[0003] The technical problem to be solved by this invention is how to improve soybean plant height, number of pods per plant, and / or number of seeds. The technical problem to be solved is not limited to the described technical subject matter; other technical subject matter not mentioned herein will be clearly understood by those skilled in the art through the following description.
[0004] To solve the above-mentioned technical problems, the present invention provides the following technical solutions:
[0005] The present invention provides the use of a protein, or a substance that inhibits, reduces, or downregulates the expression of a gene encoding said protein, or a substance that inhibits, reduces, or downregulates the activity and / or content of said protein, in any of the following:
[0006] D1) Increase soybean yield,
[0007] D2) Reduces soybean plant height;
[0008] D2) Prepare products that reduce soybean plant height;
[0009] D3) Cultivating semi-dwarf soybeans;
[0010] D4) Preparation of products from semi-dwarf soybeans;
[0011] D5) Improved semi-dwarf soybeans or products made from semi-dwarf soybeans;
[0012] D6) Increase the number of pods per soybean plant;
[0013] D7) Prepare products that increase the number of pods per soybean plant;
[0014] D8) Cultivate soybeans with an increased number of pods per plant;
[0015] D9) Prepare soybean products with an increased number of pods per plant;
[0016] D10) increases the number of soybean grains per plant;
[0017] D11) Prepare products that increase the number of soybean grains per plant;
[0018] D12) Cultivate soybeans with increased number of seeds per plant;
[0019] D13) Prepare soybean products with increased number of grains per plant;
[0020] D14) Soybean breeding;
[0021] The protein is any one of the following:
[0022] A1) The amino acid sequence of this protein is SEQ ID No. 2;
[0023] A2) A protein that has more than 80% identity with and has the same function as the protein shown in A1) obtained by substituting and / or deleting and / or adding amino acid residues of the amino acid sequence shown in SEQ ID No. 2.
[0024] A3) A fusion protein with the same function is obtained by attaching a tag to the N-terminus and / or C-terminus of A1) or A2).
[0025] To facilitate the purification or detection of the protein in A1), a tag protein can be attached to the amino or carboxyl terminus of the protein, which consists of the amino acid sequence shown in SEQ ID No. 2 in the sequence listing.
[0026] The tagged proteins include, but are not limited to: GST (glutathione thiotransferase) tagged protein, His6 tagged protein (His-tag), MBP (maltose-binding protein) tagged protein, Flag tagged protein, SUMO tagged protein, HA tagged protein, Myc tagged protein, eGFP (enhanced green fluorescent protein), eCFP (enhanced cyan fluorescent protein), eYFP (enhanced yellow-green fluorescent protein), mCherry (monomer red fluorescent protein), or AviTag tagged protein.
[0027] Those skilled in the art can readily mutate the nucleotide sequences encoding proteins GmCOL2a and / or GmCOL2b of the present invention using known methods, such as directed evolution or point mutation. Artificially modified nucleotides that possess 75% or more of the nucleotide sequence identity with the proteins GmCOL2a and / or GmCOL2b isolated in the present invention, provided they encode and function as proteins GmCOL2a and / or GmCOL2b, are derived from and equivalent to the nucleotide sequences of the present invention.
[0028] The aforementioned 75% or higher identity can be 80%, 85%, 90%, or 95% or higher. In this text, identity refers to the identity of amino acid sequences or nucleotide sequences. The identity of amino acid sequences can be determined using homology search sites on the Internet, such as the BLAST page on the NCBI homepage. For example, in Advanced BLAST 2.1, by using blastp as the procedure, setting the Expect value to 10, setting all filters to OFF, using BLOSUM62 as the matrix, and setting the Gap existence cost, Per residue gap cost, and Lambda ratio to 11, 1, and 0.85 (default values) respectively, a search can be performed to calculate the identity of amino acid sequences, and then the identity value (%) can be obtained. In this document, the 80% or more identity can be at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity.
[0029] In this document, the substance regulating the activity and / or content of the protein may be a substance regulating gene expression, wherein the gene encodes the protein GmCOL2a or / and GmCOL2b. The substance regulating gene expression may be a substance performing at least one of the following six types of regulation: 1) regulation at the transcriptional level of the gene; 2) post-transcriptional regulation of the gene (i.e., regulation of splicing or processing of the primary transcript of the gene); 3) regulation of RNA transport of the gene (i.e., regulation of the transport of mRNA of the gene from the nucleus to the cytoplasm); 4) regulation of the translation of the gene; 5) regulation of mRNA degradation of the gene; and 6) post-translational regulation of the gene (i.e., regulation of the activity of the protein translated from the gene).
[0030] The purpose of plant breeding mentioned above may include cultivating plants with reduced plant height and increased number of pods and seeds per plant.
[0031] The protein described in A2) above is the protein whose amino acid sequence is SEQ ID No.2.
[0032] Furthermore, the substance described in the above applications is a biomaterial, and the biomaterial is any one of the following:
[0033] B1) Nucleic acid molecules;
[0034] B2) Genes that express the nucleic acid molecules described in B1);
[0035] B3), an expression cassette containing the gene described in B2);
[0036] B4), a recombinant vector containing the gene described in B2), or a recombinant vector containing the expression cassette described in B3);
[0037] B5) Recombinant microorganisms containing the gene described in B2), or recombinant microorganisms containing the expression cassette described in B3), or recombinant microorganisms containing the recombinant vector described in B4);
[0038] B6) A transgenic plant cell line containing the gene described in B2), or a transgenic plant cell line containing the expression cassette described in B3), or a transgenic plant cell line containing the recombinant vector described in B4);
[0039] B7), transgenic plant tissue containing the gene described in B2), or transgenic plant tissue containing the expression cassette described in B3), or transgenic plant tissue containing the recombinant vector described in B4);
[0040] B8) a transgenic plant organ containing the gene described in B2), or a transgenic plant organ containing the expression cassette described in B3), or a transgenic plant organ containing the recombinant vector described in B3).
[0041] B1) The nucleic acid molecule may be DNA or RNA. The RNA may be gRNA or shRNA.
[0042] In this article, the nucleic acid molecule described in B1) is gRNA, and the target of the gRNA is the gene encoding the protein.
[0043] The protein-coding gene may be the protein-coding gene shown in C1), C2), or C3).
[0044] C1) The coding sequence is a cDNA molecule or DNA molecule of SEQ ID No. 1 or SEQ ID No. 3;
[0045] A cDNA or DNA molecule that hybridizes with a cDNA or DNA molecule defined by C2 and C1 and encodes a protein with the same function.
[0046] The DNA molecule shown in SEQ ID No. 1 encodes the amino acid sequence of protein GmCOL2a, which is SEQ ID No. 2. The DNA molecule shown in SEQ ID No. 3 encodes the amino acid sequence of protein GmCOL2b, which is SEQ ID No. 4. The nucleotide sequence shown in SEQ ID No. 1 is the nucleotide sequence of the gene encoding protein GmCOL2a (CDS). The nucleotide sequence shown in SEQ ID No. 3 is the nucleotide sequence of the gene encoding protein GmCOL2b (CDS).
[0047] The genes for proteins GmCOL2a and / or GmCOL2b described in this invention (GmCOL2a and / or GmCOL2b genes) can be any nucleotide sequence capable of encoding proteins GmCOL2a and / or GmCOL2b. Considering codon degeneracy and codon preferences among different species, those skilled in the art can use codons suitable for expression in specific species as needed.
[0048] B1) The nucleic acid molecule may also include nucleic acid molecules obtained by codon preference modification based on the nucleotide sequence shown in SEQ ID No. 1 or / and SEQ ID No. 3.
[0049] The nucleic acid molecules also include those that have a nucleotide sequence identity of more than 95% with the nucleotide sequence shown in SEQ ID No. 1 or / and SEQ ID No. 3 and originate from the same species. The nucleic acid molecules described herein can be DNA, such as cDNA, genomic DNA, or recombinant DNA; the nucleic acid molecules can also be RNA, such as gRNA, mRNA, siRNA, shRNA, sgRNA, miRNA, or antisense RNA.
[0050] The vectors described herein are well-known to those skilled in the art and include, but are not limited to: plasmids, bacteriophages (such as λ phage or M13 filamentous phage), granules (i.e., Cosmids), Ti plasmids, or viral vectors. Specifically, it may be PTF101-SpCas9.
[0051] Recombinant expression vectors containing the GmCOL2a or GmCOL2b gene can be constructed using existing plant expression vectors. These plant expression vectors include, but are not limited to, binary Agrobacterium vectors and vectors suitable for plant microbombardment. The plant expression vectors may also contain the 3' untranslated region of the exogenous gene, i.e., containing a polyadenylate signal and any other DNA fragment involved in mRNA processing or gene expression. The polyadenylate signal can guide the addition of polyadenylate to the 3' end of the mRNA precursor; similar functions exist for the untranslated regions transcribed at the 3' end of genes including, but not limited to, Agrobacterium crown gall-inducing (Ti) plasmids (such as the Nos gene for lipase synthesis) and plant genes (such as the soybean storage protein gene).
[0052] When constructing recombinant plant expression vectors using the GmCOL2a and / or GmCOL2b genes, any enhancing or constitutive promoter can be added before the transcription initiation nucleotide, including but not limited to the cauliflower mosaic virus (CAMV) 35S promoter and the maize ubiquitin promoter. These can be used alone or in combination with other plant promoters. Furthermore, when constructing plant expression vectors using the genes of this invention, enhancers, including translational enhancers or transcriptional enhancers, can also be used. These enhancer regions can be ATG start codons or adjacent region start codons, but they must be identical to the reading frame of the coding sequence to ensure correct translation of the entire sequence. The sources of the translation control signals and start codons are broad, and they can be natural or synthetic. The translation initiation region can originate from the transcription initiation region or structural genes.
[0053] To facilitate the identification and screening of transgenic plant cells or plants, the plant expression vectors used can be processed, such as by adding genes that can be expressed in plants, encoding enzymes or luminescent compounds that produce color changes (GUS genes, luciferase genes, etc.), antibiotic resistance markers (gentamicin markers, kanamycin markers, etc.), or chemical reagent resistance marker genes (such as herbicide resistance genes). From a safety perspective, transgenic plants can be screened directly under stress without adding any selective marker genes.
[0054] By using any vector capable of guiding the expression of exogenous genes in plants, introducing a fragment targeting the knockout of the GmCOL2a or / and GmCOL2b genes provided in this invention into plant cells or recipient plants can result in plants with reduced plant height and increased number of pods and seeds per plant. Expression vectors carrying the GmCOL2a or / and GmCOL2b genes can be transformed into plant cells or tissues using conventional biological methods such as Ti plasmids, Ri plasmids, plant virus vectors, direct DNA transformation, microinjection, electrocoagulation, and Agrobacterium-mediated transformation, and the transformed plant tissues can be cultured into plants.
[0055] The microorganisms described in this article may be yeast, bacteria, algae, or fungi. Among them, bacteria may originate from genera such as *Escherichia*, *Erwinia*, *Agrobacterium*, *Flavobacterium*, *Alcaligenes*, *Pseudomonas*, and *Bacillus*. Specifically, they may be *Escherichia coli* DH5α and / or *Agrobacterium tumefaciens* EHA105.
[0056] The target site of the gRNA may be nucleotides 26 to 45 of SEQ ID No. 1 or / and nucleotides 32 to 51 of SEQ ID No. 3.
[0057] The present invention also provides a method for increasing soybean yield, reducing soybean plant height, increasing the number of pods per soybean plant and / or increasing the number of soybean grains, comprising downregulating or reducing or weakening the expression level of the encoding gene of the aforementioned protein in the target soybean or the content of the protein to obtain transgenic soybean, wherein the transgenic soybean has a lower plant height than the target soybean, and the transgenic soybean has a higher number of pods per plant and a higher number of grains per plant than the target soybean.
[0058] The aforementioned soybeans with reduced plant height and increased number of pods and seeds per plant are understood to include not only first-generation transgenic plants obtained by transforming the target soybean with the GmCOL2a or / and GmCOL2b genes, but also their progeny. This gene can be propagated within this species, or it can be transferred into other varieties of the same species using conventional breeding techniques, particularly commercial varieties. The stress-resistant plants include seeds, callus tissue, intact plants, and cells.
[0059] Furthermore, the method described above, which involves upregulating or enhancing or increasing the expression level of the gene encoding the aforementioned protein in the target soybean or the content of the protein, involves introducing the gene encoding the aforementioned protein into the target soybean.
[0060] The aforementioned proteins also fall within the scope of protection of this invention.
[0061] The present invention also provides a biomaterial, wherein the biomaterial is any one of the following:
[0062] B1) Nucleic acid molecules that encode the aforementioned proteins;
[0063] B2), an expression cassette containing the nucleic acid molecule described in B1);
[0064] B3), a recombinant vector containing the nucleic acid molecule described in B1), or a recombinant vector containing the expression cassette described in B2);
[0065] B4) Recombinant microorganisms containing the nucleic acid molecules described in B1), or recombinant microorganisms containing the expression cassette described in B2), or recombinant microorganisms containing the recombinant vector described in B3);
[0066] B5) Nucleic acid molecules that inhibit, reduce, or downregulate the expression of genes encoding the aforementioned proteins, or nucleic acid molecules that inhibit, reduce, or downregulate the activity or content of the aforementioned proteins.
[0067] B6) The gene encoding the nucleic acid molecule described in B5);
[0068] B7), an expression cassette containing the gene described in B6);
[0069] B8), a recombinant vector containing the gene described in B6), or a recombinant vector containing the expression cassette described in B7;
[0070] B9) recombinant microorganisms containing the gene described in B6), or recombinant microorganisms containing the expression cassette described in B7), or recombinant microorganisms containing the recombinant vector described in B4).
[0071] The gene encoding the protein of the nucleic acid molecule described in B1) of the above-mentioned biological material as shown in C1), C2), or C3):
[0072] C1) The coding sequence is a cDNA molecule or DNA molecule of SEQ ID No. 1 or SEQ ID No. 3;
[0073] A cDNA or DNA molecule that hybridizes with a cDNA or DNA molecule defined by C2 and C1 and encodes a protein with the same function.
[0074] This invention obtains transgenic soybean plants with GmCOL2a or / and GmCOL2b gene knockout by knocking out the genes encoding the GmCOL2a or / and GmCOL2b proteins. The transgenic soybeans exhibit reduced plant height and increased number of seeds per pod. The GmCOL2a or / and GmCOL2b proteins and their encoding genes provided by this invention have significant theoretical and practical value in the study of soybean plant height and the number of seeds per pod. Attached Figure Description
[0075] Figure 1 The gene structures of GmCOL2a and GmCOL2b and the CRISPR / Cas9 target sequences were determined.
[0076] Figure 2 This refers to site-directed mutations at the target sites of GmCOL2a and GmCOL2b mediated by CRISPR / Cas9.
[0077] Figure 3 Phenotypes of Jack, Gmcol2a, and Gmcol2b mutants in fields in Langfang. Detailed Implementation
[0078] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.
[0079] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.
[0080] Agrobacterium tumefaciens EHA105 is described in the following literature: Cai Y, Chen L, Liu X, Guo C, Sun S, Wu C, Jiang B, Han T and Hou W (2018), CRISPR / Cas9-mediated targeted mutagenesis of GmFT2a delays flowering time in soya bean. Plant Biotechnol J 16, 176-185. It is available to the public from the Institute of Crop Science, Chinese Academy of Agricultural Sciences.
[0081] Cultivated soybean Jack is documented in the following literature: Chen L, Cai Y, Liu X, Yao W, Guo C, Sun S, Wu C, Jiang B, Han T, Hou W (2018), Improvement of soybean Agrobacterium-mediated transformation efficiency by adding glutamine and asparagine into the culturemedia. International Journal of Molecular Sciences 19, 3039. It is available to the public from the Institute of Crop Science, Chinese Academy of Agricultural Sciences.
[0082] The high-fidelity enzyme used for PCR amplification was PhantaMax Super-Fidelity DNA Polymerase, purchased from Nanjing Novizan Biotechnology Co., Ltd., catalog number P505-d3. MS salt: PhytoTech, catalog number M524. MS organic: PhytoTech, catalog number M533. B5 organic: PhytoTech, catalog number G219. B5 salt: PhytoTech, catalog number G768.
[0083] LB solid medium components: NaCl 10g / L, yeast extract 5g / L, tryptone 10g / L, agar 15g / L, solvent is water.
[0084] YEP solid medium consists of a solvent and a solute; the solutes and their concentrations in YEP solid medium are as follows: NaCl 5 g / L, yeast extract 5 g / L, tryptone 10 g / L, and agar 15 g / L; the solvent is water. The pH of YEP solid medium is 7.0.
[0085] Germination medium (pH 5.8): 3.12 g / L B5 salt, 1 mL / L B5 organic, 20 g / L sucrose, 7.5 g / L agar, balance water.
[0086] Liquid culture medium (pH 5.4): 0.43 g / L MS salt, 1 mL / LB5 organic, 40 mg / L acetylsalicylic acid, 150 mg / L dithiothreitol, 100 mg / L L-cysteine, 30 g / L sucrose, 3.9 mg / L 2-morpholinoethanesulfonic acid, balance water.
[0087] Co-culture medium (pH 5.4): 0.43 g / L MS salt, 1 mL / L B5 organic, 40 mg / L acetylsuccinone, 150 mg / L dithiothreitol, 100 mg / L L-cysteine, 30 g / L sucrose, 7.5 g / L agar, 3.9 mg / L 2-morpholinoethanesulfonic acid, balance water.
[0088] Recovery medium (pH 5.4): 3.1 g / L B5 salt, 1 mL / L B5 organic, 30 g / L sucrose, 150 mg / L cephalosporin, 150 mg / L termethin, 1 mg / L 6-BA, 0.98 g / L 2-morpholinoethanesulfonic acid, 7.5 g / L agar, 4 mL / L Fe salt (200×), 50 mg / L L-asparagine, 50 mg / L L-glutamine, balance water.
[0089] Screening medium (pH 5.4): 3.1 g / L LB5 salt, 1 mL / L LB5 organic, 0.98 g / L 2-morpholinoethanesulfonic acid, 30 g / L sucrose, 150 mg / L cephalosporin, 150 mg / L termethin, 1 mg / L 6-BA, 6 mg / L glufosinate, 7.5 g / L agar, 4 mL / L Fe salt (200×), 50 mg / L L-asparagine, 50 mg / L L-glutamine, balance water.
[0090] Elongation medium (pH 5.6): 4.0 g / L MS salt, 1 mL / L B5 organic, 0.6 g / L 2-morpholinoethanesulfonic acid, 30 g / L sucrose, 150 mg / L cephalosporin, 150 mg / L termethin, 0.1 mg / L IAA, 0.5 mg / L GA, 1 mg / L 6-BA, 6 mg / L glufosinate, 7.5 g / L agar, 4 mL / L Fe salt (200×), 50 mg / L L-asparagine, 50 mg / L L-glutamine, balance water.
[0091] Rooting medium (pH 5.7): 2.165 g / L MS salt, 1 mL / LB5 organic, 0.6 g / L 2-morpholinoethanesulfonic acid, 20 g / L sucrose, 7.5 g / L agar, 50 mg / L L-asparagine, 50 mg / L L-glutamine, with the remainder being water.
[0092] Example 1: Site-specific knockout of soybean genes GmCOL2a and GmCOL2b using CRISPR / Cas9
[0093] 1.1 sgRNA target sequence selection
[0094] The genome sequence, CDS sequence (SEQ ID No. 1), and protein sequence (SEQ ID No. 2) of soybean GmCOL2a (Glyma.13G050300) and the genome sequence, CDS sequence (SEQ ID No. 3), and protein sequence (SEQ ID No. 4) of GmCOL2b (Glyma.19G039000) were obtained from the Phytozome database. Figure 1 As shown, a target GmCOL2a-sgRNA was selected on the first exon of GmCOL2a, with the target sequence 5′-TTGGTGGCAGCACCGGCACC-3′, targeting positions 1026-1045 of the GmCOL2a genome sequence, which is positions 26 to 45 of SEQ ID No. 1; a target GmCOL2b-sgRNA was selected on the first exon of GmCOL2b, with the target sequence 5′-GCAGCAACACTGGCACCACC-3′, targeting positions 1032-1051 of the GmCOL26 genome sequence, which is positions 32 to 51 of SEQ ID No. 3.
[0095] 1.2 Construction of CRISPR / Cas9 vectors
[0096] (1) The vector pUC57-sgRNA-SpCas was double-digested with restriction endonucleases NHeI and BbsI in a 50 μL system at 37 °C for 93 hours. The target band (approximately 3201 bp in size) was detected by 1% agarose gel electrophoresis. The target band contained sgRNA expression cassette elements driven by the AtU6 promoter.
[0097] (2) Synthesize primers according to the target sequence of sgRNA. The sequence is as follows (lowercase letters represent the target sites):
[0098] GmCOL2a-Cas9-F:5′-TCGAAGTAGTGATTGttggtggcagcaccggcaccGTTTTAGAGCTAGAA-3′
[0099] GmCOL2a-Cas9-R:5′-TTCTAGCTCTAAAACggtgccggtgctgccaccaacAATCACTACTTCGA-3′
[0100] GmCOL2b-Cas9-F:5′-TCGAAGTAGTGATTGgcagcaacactggcaccaccGTTTTAGAGCTAGAA-3′
[0101] GmCOL2b-Cas9-R:5′-TTCTAGCTCTAAAACggtggtgccagtgttgctgcCAATCACTACTTCGA-3′
[0102] (3) Add 5 μL of each of the corresponding F / R primers (10 μM concentration) to a centrifuge tube, and add 15 μL of ddH2O to bring the reaction volume to 25 μL. Anneal at 95℃ for 3 min, then anneal at 0.1℃ / s to 16℃, and hold at 16℃ for 10 min to complete the annealing process and form oligo dimers. The oligo dimer formed by GmCOL2a-Cas9-F and GmCOL2a-Cas9-R encodes an sgRNA targeting GmCOL2a (named sgRNA-GmCOL2a). The oligo dimer formed by GmCOL2b-Cas9-F and GmCOL2b-Cas9-R encodes an sgRNA targeting GmCOL2b (named sgRNA-GmCOL2b).
[0103] (4) Take 3.5 μL of the annealing product and 1.5 μL of the digested pUC57-sgRNA-SpCas9 carrier gel-recovered product, and use... The Ultra One Step Cloning Kit (Nanjing Novizan Biotechnology Co., Ltd., catalog number C115-02) was used for ligation. *E. coli* DH5α was transformed using the freeze-thaw method, plated on LB+Amp solid medium, and incubated overnight at 37°C. Single colonies were picked, shaken, and sequenced. The sequencing primers were pSgRNA-CX: 5′-CGCCAGGGTTTTCCCAGTCACGAC-3′. Bacterial cultures with correct sequencing results were used for plasmid extraction. The extracted plasmids were named pUC57-sgRNA-SpCas9-GmCOL2a and pUC57-sgRNA-SpCas9-GmCOL2b. The pUC57-sgRNA-SpCas9-GmCOL2a plasmid was ligated with the sgRNA sequence targeting GmCOL2a. The pUC57-sgRNA-SpCas9-GmCOL2b plasmid was ligated with the target sgRNA sequence targeting GmCOL2b.
[0104] (5) Double digestion was performed using restriction endonucleases PacI and PmeI in a 50 μL system at 37 °C for 3 hours. After double digestion of the pUC57-sgRNA-SpCas9-GmCOL2a plasmid, the target fragment sgRNA-SpCas9-GmCOL2a (approximately 586 bp) was recovered. After double digestion of the pUC57-sgRNA-SpCas9-GmCOL2b plasmid, the target fragment sgRNA-SpCas9-GmCOL2b (approximately 586 bp) was recovered. After double digestion of the empty vector PTF101-SpCas9 plasmid, the target fragment sgRNA-SpCas9-GmCOL2b (approximately 14232 bp) was recovered. Subsequently, sgRNA-SpCas9-GmCOL2a and sgRNA-SpCas9-GmCOL2b were ligated to the target fragment of the empty vector PTF101-SpCas9 plasmid (double digestion) using T4 DNA ligase (NEB, catalog number M0202V). The ligation was then performed on E. coli DH5α using the freeze-thaw method, plated on LB+Spe solid medium, and incubated overnight at 37°C. Single clones were selected, cultured, and sequenced to verify successful ligation. The sequencing primers were sgRNA-TYJC: 5′-TGGGAATCTGAAAGAAGAGAAGCA-3′. The verified plasmids were named pGmCOL2a-sgRNA and pGmCOL2b-sgRNA. The recombinant plasmid pGmCOL2a-sgRNA is a recombinant expression vector in which the expression cassette element of sgRNA containing the target sequence of GmCOL2a replaces the nucleotide sequence between the PacI and Pme restriction enzyme recognition sites on the PTF101-SpCas9 plasmid, while maintaining other nucleotide sequences. The recombinant plasmid pGmCOL2b-sgRNA is a recombinant expression vector in which the expression cassette element of sgRNA containing the target sequence of GmCOL2b replaces the nucleotide sequence between the PacI and Pme restriction enzyme recognition sites on the PTF101-SpCas9 plasmid, while maintaining other nucleotide sequences.
[0105] (6) Extract plasmids and transform the plasmids pGmCOL2a-sgRNA and pGmCOL2b-sgRNA into competent cells of Agrobacterium tumefaciens EHA105 by freeze-thaw method to obtain recombinant Agrobacterium, named EHA105-GmCOL2a-sgRNA and EHA105-GmCOL2b-sgRNA.
[0106] 1.3 Agrobacterium tumefaciens-mediated genetic transformation of soybean
[0107] The recombinant bacteria EHA105-GmCOL2a-sgRNA and EHA105-GmCOL2b-sgRNA were transformed into the soybean cultivar Jack. The specific steps are as follows:
[0108] 1. Seed sterilization
[0109] 1) Take healthy, plump, uniform, and dry Jack soybean seeds that are free from pests, diseases, and spots, spread them evenly in a petri dish, and then place the petri dish in a desiccator.
[0110] 2) After completing step 1), place a 100mL beaker in the desiccator, pour 90mL of 12M sodium hypochlorite aqueous solution into the beaker, then slowly add 5mL of concentrated hydrochloric acid, and then quickly cover the desiccator, seal it with petroleum jelly, and place it for 16 hours for chlorine sterilization.
[0111] 2. Preparation of infecting bacterial solution
[0112] 1) Incubate the obtained EHA105-GmCOL2a-sgRNA and EHA105-GmCOL2b-sgRNA bacterial cultures at 28℃, resuspend in liquid culture medium, and obtain OD. 600nm =0.6-0.8 in the infecting bacterial solution.
[0113] 2) Place the Jack soybean seeds treated in step 1 into a clean bench. Under a microscope, peel off the seed coat, separate the two cotyledons along the long axis, and keep the cotyledon with the complete hypocotyl. Make scratches at the junction of the hypocotyl and cotyledon, usually 3-5 scratches per cotyledon. Then, immerse the seeds in a 28°C incubator for 2 hours.
[0114] 3) Place the cotyledons with the inner (smooth) side up on a co-culture medium lined with sterile filter paper, and incubate in the dark at 22°C for 5 days.
[0115] 4) After 5 days of co-culture, the hypocotyl of the explants elongated to 2 cm. Part of the hypocotyl was cut off, leaving 0.5 cm. The treated explants were then placed in recovery medium and cultured at 28°C under 16 hours of light / 8 hours of darkness for 7 days.
[0116] 5) Remove the explants from the recovery medium, remove the new shoots, cut off part of the hypocotyl, leaving 0.5 cm of the hypocotyl, and then transfer the trimmed explants into the selection medium and culture them at 28°C for 21 days under 16 hours of light / 8 hours of darkness.
[0117] 6) After 21 days of selection and induction, the explants produced a large number of adventitious buds. The cotyledons and brown leaves were removed, and the remaining parts were transferred to elongation medium for culture at 28°C under 16 hours of light / 8 hours of darkness.
[0118] 7) In the elongation medium, when the clustered buds produce 5-8 cm young stems, cut them off from the base of the adventitious buds; dip the stem base in 1 mg / LIBA solution for 1 minute, and then transfer it to the rooting medium for culture. Culture at 28°C under 16 hours of light / 8 hours of darkness for one week. After a large number of roots are produced at the base of the stem, transplant them into pots. The resulting plants are T0 generation transgenic soybean plants.
[0119] 1.4 Detection of Editing Type in Gene-Edited Plants
[0120] (1) DNA was extracted from the leaves of T0 generation transgenic soybean plants using a plant genomic DNA extraction kit (Kangwei Century, CW0553).
[0121] (2) PCR amplification was performed using DNA as a template. The primers for amplifying the target GmCOL2a-sgRNA were GmCOL2a-SpCas9-F: 5′-AGGGATAACATGAGATTTTGACTGG-3′ and GmCOL2a-SpCas9-R: 5′-AGGGATAACATGAGATTTTGACTGG-3′, with an amplified fragment length of 788 bp. The sequencing primer was COL-CX: 5′-CACGATGGCACCCTTGATTCTCCC-3′. The primers for amplifying the target GmCOL2b-sgRNA were GmCOL2b-SpCas9-F: 5′-ACACGTGTCTCCAAGTTGTGT-3′ and GmCOL2b-SpCas9-R: 5′-ACGCGTTTGTGATTGTGCTC-3′, with an amplified fragment length of 756 bp. The sequencing primer was GmCOL2b-SpCas9-F. PCR reaction system: 2× Max Buffer 25μL, dNTP Mix (10mM each) 1μL, DNA (200ng / μL) 2μL, F (10pmol / μL) 2μL, R (10pmol / μL) 2μL, Max Super-Fidelity DNA Polymerase 1 μL, ddH2O 17 μL, total volume 50 μL. Amplification reaction system: 95℃ 3 min; 95℃ 15 sec, 56℃ 15 sec, 72℃ 45 sec, 35 cycles; 72℃ 5 min. Sequencing analysis was performed on the amplified products. There are three possible sequencing results: 1) Overlapping peaks appear from near the CRISPR / Cas9 cleavage site onwards, but no overlapping peaks upstream of the target site, indicating a heterozygous editing type; 2) No overlapping peaks appear from beginning to end, and the sequence is compared with the wild-type reference sequence. If they are completely identical, no site-directed mutation has occurred; 3) No overlapping peaks appear from beginning to end, and the sequence is compared with the wild-type reference sequence. If they are not completely identical, the specific mutation type is determined, which generally includes base insertion, deletion, and mismatch.
[0122] Upon testing, 22 out of 41 T0 generation GmCOL2a-SpCas9 transgenic soybean plants showed gene editing, and 16 out of 36 T0 generation GmCOL2b-SpCas9 transgenic soybean plants showed gene editing. The T0 generation gene-edited plants were self-crossed once to obtain the T1 generation gene-edited plants. The mutation type in the T1 generation gene-edited plants was then analyzed using the method described above.
[0123] The results showed that two homozygous mutation types were generated at the GmCOL2a-sgRNA target site. Figure 2 The two subtypes (A) were named Gmcol2a-51 and Gmcol2a-78, respectively. Compared to the genomic DNA of wild-type soybean:
[0124] The GmCOL2a protein was deleted on both homologous chromosomes of the Gmcol2a-51 gene, specifically from positions 35 to 42 of SEQ ID No. 1, a total of 8 nucleotides. This caused a frameshift in the translation of the amino acids after the editing site, prematurely terminating the protein sequence and thus knocking out the GmCOL2a gene.
[0125] The GmCOL2a protein encoding the gene in Gmcol2a-78 underwent the same deletion on both homologous chromosomes, resulting in the deletion of 398 nucleotides from positions 682 to 1079 of the genome (including bases 1 to 79 of SEQ ID No. 1). This deletion caused a frameshift in the amino acid translation after the edited site, prematurely terminating the protein sequence and thus knocking out the GmCOL2a gene. The mutated genome of the Gmcol2a-78 mutant is shown in the mutated genome sequence of the Gmcol2a-78 mutant described below, as are its coding sequence (CDS) and amino acid sequence. The GmCOL2a protein is functionally lost.
[0126] The results showed that two homozygous mutation types were generated at the GmCOL2b-sgRNA target site. Figure 2 The two subtypes (B) were named Gmcol2b-35 and Gmcol2b-46, respectively. Compared to the genomic DNA of wild-type soybean:
[0127] The GmCOL2b protein encoding the gene in Gmcol2b-35 underwent the same deletion on both homologous chromosomes, specifically the deletion of position 48 of SEQ ID No. 3, totaling one nucleotide. This resulted in a frameshift of the amino acid translation after the editing site, premature termination of the protein sequence, and thus knockout of the GmCOL2b gene. The mutated genome of the Gmcol2b-35 mutant is shown in the mutated genome sequence of the Gmcol2b-35 mutant described below, as are its coding sequence (CDS) and amino acid sequence. The GmCOL2b protein is functionally lost.
[0128] The same deletion occurred on both homologous chromosomes of the gene encoding the GmCOL2b protein in Gmcol2b-46, with a total deletion of 3 nucleotides from positions 46 to 48 of SEQ ID No. 3, resulting in the deletion of one amino acid at the editing site.
[0129] Subsequently, the T1 generation Gmcol2a-78 and Gmcol2b-35 mutants were self-crossed for one generation to obtain T2 generation homozygous Gmcol2a-78 and Gmcol2b-35 mutant seeds, which were named Gmcol2a mutant and Gmcol2b mutant for subsequent field experiments.
[0130] Example 2: Phenotypic observation of wild-type, Gmcol2a, and Gmcol2b mutants in Langfang field.
[0131] Wild-type soybean varieties Jack, Gmcol2a, and Gmcol2b mutants were summer-sown in a field in Langfang (39.5°N, 116.6°E). A randomized block design with three replicates was used. Each treatment plot was 15 square meters (6.0 m × 2.5 m), with 6 rows sown per plot, each row 6 m long; row spacing was 50 cm, and plant spacing was 10 cm. Five plants were randomly selected from the middle four rows of each plot, for a total of 15 plants selected from all three plots for each material. Agronomic traits such as plant height (cm), number of main stem nodes, number of pods per plant, and number of seeds per plant were observed and statistically analyzed. GraphPad Prism 8 software was used for data processing. Experimental results are expressed as mean ± standard deviation. One-way ANOVA was used, with P < 0.05 (*) indicating a significant difference compared to the wild-type soybean variety Jack, and P < 0.01 (**) indicating a highly significant difference compared to the wild-type soybean variety Jack. The phenotypes of wild-type Jack, Gmcol2a, and Gmcol2b mutants in the Langfang field are as follows: Figure 3 As shown in Table 1. Statistical results indicate that, compared with the wild-type Jack, the Gmcol2a and Gmcol2b mutants have reduced plant height, significantly increased number of pods and seeds per plant, but no significant difference in the number of nodes on the main stem.
[0132] Table 1. Field phenotypes of wild-type, Gmcol2a, and Gmcol2b mutants in Langfang.
[0133] plant Plant height (cm) Number of pods per plant Number of seeds per plant (seeds) Number of nodes on the main stem (nodes) Jack 103.4±3.9 85.6±17.7 204.5±40.1 21.6±0.7 Gmcol2a 74.7±5.8** 101.9±13.8** 235.3±39.3** 20.9±1.3 Gmcol2b 71.2±8.0** 101.1±12.6** 242.4±34.1** 21.3±1.0
[0134] Example 3: Phenotypic observation under three planting density conditions in a Beijing field.
[0135] Wild-type soybean varieties Jack, Gmcol2a, and Gmcol2b mutants were summer-sown in a field in Beijing (40.2°N, 116.6°E). Each treatment plot was 5 square meters (2.0 m × 2.5 m), with 6 rows sown per plot, each row 2 m long and 50 cm apart. Three planting densities were established: density 1 (286,000 plants / ha), density 2 (200,000 plants / ha), and density 3 (154,000 plants / ha). A randomized block design with three replicates was used. Fifteen plants were randomly selected from the middle four rows of each plot, for a total of 45 plants selected from all three plots for each material and density condition. Agronomic traits such as plant height, number of nodes on the main stem, number of pods per plant, and number of seeds per plant were observed and statistically analyzed. Data were processed using GraphPad Prism 8 software. Experimental results are expressed as mean ± standard deviation. One-way ANOVA was used, with P < 0.05 (*) indicating a significant difference compared to the wild-type soybean variety Jack at the same planting density, and P < 0.01 (**) indicating a highly significant difference. Statistical results showed that under the three planting density conditions, compared to wild-type Jack, both Gmcol2a and Gmcol2b mutants exhibited reduced plant height and significantly increased number of pods and grains per plant, but no significant difference in the number of nodes on the main stem (see Table 2 for details). This indicates that loss of function of GmCOL2a or GmCOL2b proteins leads to a significant decrease in soybean plant height, a significant increase in the number of pods per plant, and a significant increase in the number of grains per plant. Subsequently, the actual yield of each plot was determined. For each plot, six rows of material were harvested, two edge rows were removed, and the remaining four rows were mixed and weighed. The weight was converted to kg / ha. The results showed that the wild-type Jack had harvest yields of 3405.5±311.9 kg / ha at densities of 1, 2, and 3, respectively; the Gmcol2a mutant had harvest yields of 4366.7±260.3 kg / ha at densities of 1, 2, and 3, respectively; and the Gmcol2a mutant had harvest yields of 4366.7±260.3 kg / ha at densities of 1, 2, and 3, respectively, representing increases of 28.2%, 27.9%, and 21.2% compared to the control Jack. The Gmcol2b mutant yielded 3850.0±115.5 kg / ha at densities of 1, 2, and 3, respectively, representing increases of 13.1%, 19.3%, and 38.7% compared to the control Jack.
[0136] Table 2. Field phenotypes of wild-type, Gmcol2a, and Gmcol2b mutants at different planting densities in Beijing.
[0137] plant Plant height (cm) Number of pods per plant Number of seeds per plant (seeds) Number of nodes on the main stem (nodes) Jack - Density 1 151.0±8.1 58.7±7.4 152.6±20.6 23.2±1.5 Gmcol2a-Density 1 129.7±7.9** 72.0±8.2** 184.8±21.9** 22.2±1.5 Gmcol2b-Density 1 128.4±6.9** 66.8±10.6** 170.4±28.0** 21.8±1.5 Jack - Density 2 136.2±6.3 69.3±12.8 179.0±34.0 23.4±1.7 Gmcol2a-density2 124.2±6.6** 82.7±8.3** 214.2±21.0** 22.8±1.3 Gmcol2b-density2 122.4±6.8** 79.9±10.9** 207.8±31.4** 22.9±1.6 Jack - Density 3 130.7±8.8 74.2±14.7 196.5±39.0 23.5±1.9 Gmcol2a-Density3 119.2±8.1** 90.0±17.8** 235.9±47.6** 23.3±1.9 Gmcol2b-density3 115.07±10.6** 101.1±24.7** 256.9±63.7** 23.6±1.3
[0138] GmCOL2a genomic sequence
[0139] 5'-AGGATCGTCCAAGTCAATATAGCTTATGCTGGTTCAAAGTGCAAATCCGATCCAA GAAGAAGCTCGACCTTATCTGAGTTTGATAACATTAATTGCTTGTATGTCACATGGCTTGCGATTTACAACAACCATCAACAACTTTTCCTATGCATTCTACATACATATAAGGCATTCTTATAACTTCAAGTTCATATCTAATCTTTCATTTCTTGCCTATTTCTTACAAAAGCTAACTTGAACGTTGATGCCTTTACAAGTAGCGACCATCAAACCCCTCCTTTATAACCACCATGGTGCTACTATGTTATTCCTAAGTACACCAAACAAAATTTATTATGAACAATAATATTTGAACACCATGTTATTTGAAACATTCCACCTTGACAGTTCTTCTAACTCCACCAACAAAAAATTACTATTTATTTTCATAATCACAAAGACTCCACATTCTATCTGCAAAATCAGTAATCATCTAAACACGTGCTTTAATCCTTTTACGTACAGCCAATAAGATTACATGAAAGGAAAGCATATATATTAAGGGATAACATGAGATTTTGACTGGTGCATGTATATAGAATCTACATGGACCCAAAATTACACGTGTCTCCAAGTTGTGCCTCCATAGCAATACGAACAAAGCCTCAACATCTCTTTAATGATT
[0140] CCAAAAGGCAAAGGAAGCAACATGCACAGGTCCAACAAAAAAAAAAAAAAAAAACCT
[0141] CTCTCCTATTCAACCCCAAAACACAACCAACCACAAGATAAGCAAGAACTTGTGCACGC
[0142] CACTAACTTGCTGCCACGATGGCACCCTTGATTCTCCCCCAACACTACTTGGTTCTCCTC
[0143] ACTGAACTCAACTTCTCACTCACACTCACTTCTCTTCCTCCAAATTAGTTCCAACTCCAA
[0144] GTTAGTTCCAACACAAACACAAACACACACACATACATAAACCAAGCAAAGAAAGACT
[0145] CTACAAACACTTCACTACTCATACAATATTTTCAGACACAACATGTTGAAGGAAGGCACC
[0146] AACAACGTTGGTGGCAGCACCGGCACCTGGTCACATGTCTGCGACACGTGCCGGTCAG
[0147] CGCCATGCGTACTGTACTGCCATGCAGACTCAGCATACCTTTGCTCATCCTGCGATGCTC
[0148] GTGTCCATGCGGCTAACCGTGTGGCCTCAAGACACGAGCGTGTGTGGGTGTGCGAAGC
[0149] GTGTGAACGTGCTCCCGCAGCGTTTCTATGCAAAGCCGACGCAGCTTCTCTTTGTTCTTC
[0150] CTGTGATGCTGACATTCACTCAGCAAACCCTCTCGCTAGCCGCCACCACCGCGTGCCCA
[0151] TTCTCCCGATCTCTGGTTCCCTCTTCGGGGAACCAGAGCATGAACGCGTGTACGCGTTC
[0152] GTGAATGAAGTGGAAGCGGAGGAGGAAGAGGAAGAGGTTTTTGATGAGTATGATGAGG
[0153] TTGAAGCAGCTTCGTGGTTGTTGCCACATCCTATGAAAAATGATAAAATTGACGAGAATG
[0154] GTGGTGATAAGGGTTTTTTGTTTGGTGATGAGTATTTCGACAACCTTGTTGATTGTAACTC
[0155] ATGTGGTCACAATAACAACCAGTTTAGCAACGTTTATGATCAGCACCAGCAGAATTACA
[0156] GCAACACTGTCCCTCAGAACTATGCAGTGGTTCCAGTTCAGGTGCCGCAGCATTTTCAA
[0157] CCGGGTTTGGACTTTGACTCATCAAAAGCTGGGTTCAGTTACGATGGTTCTCTTAGTCAA
[0158] AGTGTAAGTTACCTTCTTTTTACTTCTATGTGTTTCTTTTTTGGTTTGTTGTTTATATCAATG
[0159] GGGTTTTCCCTTCTTTTTTTTTTTTTTTTTCTTTTTGCAAACTCTTTGGATCATTCATCCAT
[0160] TGGATTATACTACTTTATTAGATTTTGCCTTGAAACCTGTGCTAACTCTTCCTTCTAATTGA
[0161] CAGTTTCTTTTTCTCCATTCGTATGAAACCAAGGTTTTAAACTGCAGTCATAGTCATAGTC
[0162] AAGATTATAGTTTATTCAAAAACATTAGCATTGTAATACAATTGTACTGTAATTGTCGTCG
[0163] TGAAAAAGTTAAAAAAAAGTTTAAAAAATCTCAATCGCATACCTTCTATAAAAATCTTGA
[0164] TTGAAACCAAAGTGGGTTTCATAAGTAGTTTCTTTTACAAGAAAAACTAGCAATAAAGT
[0165] CAACTACTCTAGATCCATTGCCATACATGTGACCTTGATCCATTCCACAACGTTTCGTATT
[0166] TCTTTGCCTACCCTTTTCCCATTGACTACTTTTCTGTCTTGAGAGACATTTATTTTATTTTT
[0167] TTCGTATGCAATAGCTTTCCCTTTTTGATGTCTTGTATTTGCTAGTTATAGTTGTGATTCCT
[0168] CATAAAACAGATACTTACTTTTATAGGGTAACTTAGCTTTTCACCAATTTTCGCCCATACAT
[0169] AGTTAGACTTCCCTAAACTTTGTTTCCTCATTTCTTCAACTATAATTTGATCAATTGTATAT
[0170] TAATTTCCCATAATTTTCAGCTTCAATTAAAATTTGAGCCTAAGTGCCTGTTTTATTAA
[0171] ACCAAACAATGGGGGGTTTAATATTTTAAAAACTTTCTTCTATATTTCTTTTACATTTCCT
[0172] CTGTTTTCTAATATATCTTTTCTTTTTTCTATTACATCATCTTTGATATATTTTCACCTTTTTC
[0173] TTCTTACTTTGGGGACAAACATTATTGTTTCACAAAATTGGTTAGCTACTAATAGTTTATT
[0174] GCCTTACCTGTAATCTTCCAACTTCATACACGTTTTTGTGTTAATCTTTTATATTTCATTT
[0175] TTAATTTAGGTTTCGGTTTCATCGATGGATGTTGGTGTTGTACTCGAATCAACAATAAGTG
[0176] ACATCTCAATGTCCCACTCAAAGTCGCCAATAGGGACAACTGACCTATTTCCTCCCCTTC
[0177] CCATGCCTTCACATCTCACACCAATGGACAGAGAGGCAAGAGTCCTAAGATACAGGGAG
[0178] AAAAAGAAGACAAGAAAATTTGAGAAGAAAATAAGGTATGCCTCAAGGAAGGCCTATG
[0179] CAGAGACTAGACCCCGCATAAAGGGTCGTTTTGCAAAGAGAACCGATGTAGAAGCTGA
[0180] AGTGGACCAGATGTTCTCCACAACACTATTCACTGAAGTTGGAGGTAGCATTTTTCCCACTTTCTAGAATGGAAAGAAAGTAATGGCCAGGCCATAAAAGAGAAGGTT-3'
[0181] GmCOL2b genomic sequence
[0182]
[0183] pUC57-sgRNA-SpCas9 vector sequence
[0184]
[0185] The sequence of the PTF101-SpCas9 vector is as follows:
[0186] 5'-AGTACTTTAAAGTACTTTAAAGTACTTTAAAGTACTTTGATCCAACCCCTCCGCTG CTATAGTGCAGTCGGCTTCTGACGTTCAGTGCAGCCGTCTTCTGAAAACGACATGTCGC
[0187] ACAAGTCCTAAGTTACGCGACAGGCTGCCGCCCTGCCCTTTTCCTGGCGTTTTCTTGTCG
[0188] CGTGTTTTAGTCGCATAAAGTAGAATACTTGCGACTAGAACCGGAGACATTACGCCATGA
[0189] ACAAGAGCGCCGCCGCTGGCCTGCTGGGCTATGCCCGCGTCAGCACCGACGACCAGGA
[0190] CTTGACCAACCAACGGGCCGAACTGCACGCGGCCGGCTGCACCAAGCTGTTTTCCGAG
[0191] AAGATCACCGGCACCAGGCGCGACCGCCCGGAGCTGGCCAGGATGCTTGACCACCTAC
[0192] GCCCTGGCGACGTTGTGACAGTGACCAGGCTAGACCGCCTGGCCCGCAGCACCCGCGA
[0193] CCTACTGGACATTGCCGAGCGCATCCAGGAGGCCGGCGCGGGCCTGCGTAGCCTGGCA
[0194] GAGCCGTGGGCCGACACCACCACGCCGGCCGGCCGCATGGTGTTGACCGTGTTCGCCG
[0195] GCATTGCCGAGTTCGAGCGTTCCCTAATCATCGACCGCACCCGGAGCGGGCGCGAGGCC
[0196] GCCAAGGCCCGAGGCGTGAAGTTTGGCCCCCGCCCTACCCTCACCCCGGCACAGATCG
[0197] CGCACGCCCGCGAGCTGATCGACCAGGAAGGCCGCACCGTGAAAGAGGCGGCTGCACT
[0198] GCTTGGCGTGCATCGCTCGACCCTGTACCGCGCACTTGAGCGCAGCGAGGAAGTGACG
[0199] CCCACCGAGGCCAGGCGGCGCGGTGCCTTCCGTGAGGACGCATTGACCGAGGCCGACG
[0200] CCCTGGCGGCCGCCGAGAATGAACGCCAAGAGGAACAAGCATGAAACCGCACCAGGA
[0201] CGGCCAGGACGAACCGTTTTTCATTACCGAAGAGATCGAGGCGGAGATGATCGCGGCCG
[0202] GGTACGTGTTCGAGCCGCCCGCGCACGTCTCAACCGTGCGGCTGCATGAAATCCTGGCC
[0203] GGTTTGTCTGATGCCAAGCTGGCGGCCTGGCCGGCCAGCTTGGCCGCTGAAGAAACCG
[0204] AGCGCCGCCGTCTAAAAAGGTGATGTGTATTTGAGTAAAACAGCTTGCGTCATGCGGTC
[0205] GCTGCGTATATGATGCGATGAGTAAATAAACAAATACGCAAGGGGAACGCATGAAGGTT
[0206] ATCGCTGTACTTAACCAGAAAGGCGGGTCAGGCAAGACGACCATCGCAACCCATCTAGC
[0207] CCGCGCCCTGCAACTCGCCGGGGCCGATGTTCTGTTAGTCGATTCCGATCCCCAGGGCA
[0208] GTGCCCGCGATTGGGCGGCCGTGCGGGAAGATCAACCGCTAACCGTTGTCGGCATCGAC
[0209] CGCCCGACGATTGACCGCGACGTGAAGGCCATCGGCCGGCGCGACTTCGTAGTGATCGA
[0210] CGGAGCGCCCCAGGCGGCGGACTTGGCTGTGTCCGCGATCAAGGCAGCCGACTTCGTG
[0211] CTGATTCCGGTGCAGCCAAGCCCTTACGACATATGGGCCACCGCCGACCTGGTGGAGCT
[0212] GGTTAAGCAGCGCATTGAGGTCACGGATGGAAGGCTACAAGCGGCCTTTGTCGTGTCGC
[0213] GGGCGATCAAAGGCACGCGCATCGGCGGTGAGGTTGCCGAGGCGCTGGCCGGGTACGA
[0214] GCTGCCCATTCTTGAGTCCCGTATCACGCAGCGCGTGAGCTACCCAGGCACTGCCGCCG
[0215] CCGGCACAACCGTTCTTGAATCAGAACCCGAGGGCGACGCTGCCCGCGAGGTCCAGGC
[0216] GCTGGCCGCTGAAATTAAATCAAAACTCATTTGAGTTAATGAGGTAAAGAGAAAATGAG
[0217] CAAAAGCACAAACACGCTAAGTGCCGGCCGTCCGAGCGCACGCAGCAGCAAGGCTGC
[0218] AACGTTGGCCAGCCTGGCAGACACGCCAGCCATGAAGCGGGTCAACTTTCAGTTGCCG
[0219] GCGGAGGATCACACCAAGCTGAAGATGTACGCGGTACGCCAAGGCAAGACCATTACCG
[0220] AGCTGCTATCTGAATACATCGCGCAGCTACCAGAGTAAATGAGCAAATGAATAAATGAGT
[0221] AGATGAATTTTAGCGGCTAAAGGAGGCGGCATGGAAAATCAAGAACAACCAGGCACCG
[0222] ACGCCGTGGAATGCCCCATGTGTGGAGGAACGGGCGGTTGGCCAGGCGTAAGCGGCTG
[0223] GGTTGTCTGCCGGCCCTGCAATGGCACTGGAACCCCCAAGCCCGAGGAATCGGCGTGA
[0224] GCGGTCGCAAACCATCCGGCCCGGTACAAATCGGCGCGGCGCTGGGTGATGACCTGGT
[0225] GGAGAAGTTGAAGGCCGCGCAGGCCGCCCAGCGGCAACGCATCGAGGCAGAAGCACG
[0226] CCCCGGTGAATCGTGGCAAGCGGCCGCTGATCGAATCCGCAAAGAATCCCGGCAACCG
[0227] CCGGCAGCCGGTGCGCCGTCGATTAGGAAGCCGCCCAAGGGCGACGAGCAACCAGATT
[0228] TTTTCGTTCCGATGCTCTATGACGTGGGCACCCGCGATAGTCGCAGCATCATGGACGTGG
[0229] CCGTTTTCCGTCTGTCGAAGCGTGACCGACGAGCTGGCGAGGTGATCCGCTACGAGCTT
[0230] CCAGACGGGCACGTAGAGGTTTCCGCAGGGCCGGCCGGCATGGCCAGTGTGTGGGATT
[0231] ACGACCTGGTACTGATGGCGGTTTCCCATCTAACCGAATCCATGAACCGATACCGGGAA
[0232] GGGAAGGGAGACAAGCCCGGCCGCGTGTTCCGTCCACACGTTGCGGACGTACTCAAGT
[0233] TCTGCCGGCGAGCCGATGGCGGAAAGCAGAAAGACGACCTGGTAGAAACCTGCATTCG
[0234] GTTAAACACCACGCACGTTGCCATGCAGCGTACGAAGAAGGCCAAGAACGGCCGCCTG
[0235] GTGACGGTATCCGAGGGTGAAGCCTTGATTAGCCGCTACAAGATCGTAAAGAGCGAAAC
[0236] CGGGCGGCCGGAGTACATCGAGATCGAGCTAGCTGATTGGATGTACCGCGAGATCACAG
[0237] AAGGCAAGAACCCGGACGTGCTGACGGTTCACCCCGATTACTTTTTGATCGATCCCGGC
[0238] ATCGGCCGTTTTCTCTACCGCCTGGCACGCCGCGCCGCAGGCAAGGCAGAAGCCAGATG
[0239] GTTGTTCAAGACGATCTACGAACGCAGTGGCAGCGCCGGAGAGTTCAAGAAGTTCTGT
[0240] TTCACCGTGCGCAAGCTGATCGGGTCAAATGACCTGCCGGAGTACGATTTGAAGGAGGA
[0241] GGCGGGGCAGGCTGGCCCGATCCTAGTCATGCGCTACCGCAACCTGATCGAGGGCGAA
[0242] GCATCCGCCGGTTCCTAATGTACGGAGCAGATGCTAGGGCAAATTGCCCTAGCAGGGGA
[0243] AAAAGGTCGAAAAGGTCTCTTTCCTGTGGATAGCACGTACATTGGGAACCCAAAGCCGT
[0244] ACATTGGGAACCGGAACCCGTACATTGGGAACCCAAAGCCGTACATTGGGAACCGGTC
[0245] ACACATGTAAGTGACTGATATAAAAGAGAAAAAAGGCGATTTTTCCGCCTAAAACTCTTT
[0246] AAAACTTATTAAAACTCTTAAAACCCGCCTGGCCTGTGCATAACTGTCTGGCCAGCGCA
[0247] CAGCCGAAGAGCTGCAAAAAGCGCCTACCCTTCGGTCGCTGCGCTCCCTACGCCCCGCC
[0248] GCTTCGCGTCGGCCTATCGCGGCCGCTGGCCGCTCAAAAATGGCTGGCCTACGGCCAGG
[0249] CAATCTACCAGGGCGCGGACAAGCCGCGCCGTCGCCACTCGACCGCCGGCGCCCACAT
[0250] CAAGGCACCCTGCCTCGCGCGTTTCGGTGATGACGGTGAAAACCTCTGACACATGCAGC
[0251] TCCCGGAGACGGTCACAGCTTGTCTGTAAGCGGATGCCGGGAGCAGACAAGCCCGTCA
[0252] GGGCGCGTCAGCGGGTGTTGGCGGGTGTCGGGGCGCAGCCATGACCCAGTCACGTAGC
[0253] GATAGCGGAGTGTATACTGGCTTAACTATGCGGCATCAGAGCAGATTGTACTGAGAGTGC
[0254] ACCATATGCGGTTGGAAATACCGCACAGATGCGTAAGGAGAAAATACCGCATCAGGCGC
[0255] TCTTCCGCTTCCTCGCTCACTGACTCGCTGCGCTCGGTCGTTCGGCTGCGGCGAGCGGTA
[0256] TCAGCTCACTCAAAGGCGGTAATACGGTTATCCACAGAATCAGGGGATAACGCAGGAAA
[0257] GAACATGTGAGCAAAAGGCCAGCAAAAGGCCAGGAACCGTAAAAAGGCCGCGTTGCT
[0258] GGCGTTTTTCCATAGGCTCCGCCCCCCTGACGAGCATCACAAAAATCGACGCTCAAGTC
[0259] AGAGGTGGCGAAACCCGACAGGACTATAAAGATACCAGGCGTTTCCCCCTGGAAGCTC
[0260] CCTCGTGCGCTCTCCTGTTCCGACCCTGCCGCTTACCGGATACCTGTCCGCCTTTCTCCC
[0261] TTCGGGAAGCGTGGCGCTTTCTCATAGCTCACGCTGTAGGTATCTCAGTTCGGTGTAGGT
[0262] CGTTCGCTCCAAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGACCGCTGCGCCT
[0263] TATCCGGTAACTATCGTCTTGAGTCCAACCCGGTAAGACACGACTTATCGCCACTGGCAG
[0264] CAGCCACTGGTAACAGGATTAGCAGAGCGAGGTATGTAGGCGGTGCTACAGAGTTCTTG
[0265] AATGGTGGCCTAACTACGGCTACACTAGAAGGACAGTATTTGGTATCTGCGCTCTGCTG
[0266] AAGCCAGTTACCTTCGGAAAAAGAGTTGGTAGCTCTTGATCCGGCAAACAAACCACCG
[0267] CTGGTAGCGGTGGGTTTTTTTGTTTGCAAGCAGCAGATTACGCGCAGAAAAAAAGGATCT
[0268] CAAGAAGATCCTTTGATCTTTTCTACGGGGTCTGACGCTCAGTGGAACGAAAACTCACG
[0269] TTAAGGGATTTTGGTCATGCATGATATATCTCCCAATTTGTGTAGGGCTTATTATGCACGCT
[0270] TAAAAATAATAAAAGCAGACTTGACCTGATAGTTTGGCTGTGAGCAATTATGTGCTTAGT
[0271] GCATCTAACGCTTGAGTTAAGCCGCGCCGCGAAGCGGCGTCGGCTTGAACGAATTTCTA
[0272] GCTAGACATTATTTGCCGACTACCTTGGTGATCTCGCCTTTCACGTAGTGGACAAATTCTT
[0273] CCAACTGATCTGCGCGCGAGGCCAAGCGATCTTCTTCTTGTCCAAGATAAGCCTGTCTA
[0274] GCTTCAAGTATGACGGGCTGATACTGGGCCGGCAGGCGCTCCATTGCCCAGTCGGCAGC
[0275] GACATCCTTCGGCGCGATTTTGCCGGTTACTGCGCTGTACCAAATGCGGGACAACGTAA
[0276] GCACTACATTTCGCTCATCGCCAGCCCAGTCGGGCGGCGAGTTCCATAGCGTTAAGGTTT
[0277] CATTTAGCGCCTCAAATAGATCCTGTTCAGGAACCGGATCAAAGAGTTCCTCCGCCGCT
[0278] GGACCTACCAAGGCAACGCTATGTTCTCTTGCTTTTGTCAGCAAGATAGCCAGATCAATG
[0279] TCGATCGTGGCTGGCTCGAAGATACCTGCAAGAATGTCATTGCGCTGCCATTCTCCAAAT
[0280] TGCAGTTCGCGCTTAGCTGGATAACGCCACGGAATGATGTCGTCGTGCACAACAATGGT
[0281] GACTTCTACAGCGCGGAGAATCTCGCTCTCTCCAGGGGAAGCCGAAGTTTCCAAAAGG
[0282] TCGTTGATCAAAGCTCGCCGCGTTGTTTCATCAAGCCTTACGGTCACCGTAACCAGCAA
[0283] ATCAATATCACTGTGTGGCTTCAGGCCGCCATCCACTGCGGAGCCGTACAAATGTACGGC
[0284] CAGCAACGTCGGTTCGAGATGGCGCTCGATGACGCCAACTACCTCTGATAGTTGAGTCG
[0285] ATACTTCGGCGATCACCGCTTCCCCCATGATGTTTAACTTTGTTTTAGGGCGACTGCCCTG
[0286] CTGCGTAACATCGTTGCTGCTCCATAACATCAAACATCGACCCACGGCGTAACGCGCTTG
[0287] CTGCTTGGATGCCCGAGGCATAGACTGTACCCCAAAAAAACAGTCATAACAAGCCATGA
[0288] AAACCGCCACTGCGCCGTTACCACCGCTGCGTTCGGTCAAGGTTCTGGACCAGTTGCGT
[0289] GACGGCAGTTACGCTACTTGCATTACAGCTTACGAACCGAACGAGGCTTATGTCCACTG
[0290] GGTTCGTGCCCGAATTGATCACAGGCAGCAACGCTCTGTCATCGTTACAATCAACATGCT
[0291] ACCCTCCGCGAGATCATCCGTGTTTCAAACCCGGCAGCTTAGTTGCCGTTCTTCCGAATA
[0292] GCATCGGTAACATGAGCAAAGTCTGCCGCCTTACAACGGCTCTCCCGCTGACGCCGTCC
[0293] CGGACTGATGGGCTGCCTGTATCGAGTGGTGATTTTGTGCCGAGCTGCCGGTCGGGGAG
[0294] CTGTTGGCTGGCTGGTGGCAGGATATATTGTGGTGTAAACAAATTGACGCTTAGACAACT
[0295] TAATAACACATTGCGGACGTTTTTAATGTACTGAATTAACGCCGAATTGCTCTAGCATTCG
[0296] CCATTCAGGCTGCGCAACTGTTGGGAAGGGCGATCGGTGCGGGCCTCTTCGCTATTACG
[0297] CCAGCTGGCGAAAGGGGGATGTGCTGCAAGGCGATTAAGTTGGGTAACGCCAGGGTTT
[0298] TCCCAGTCACGACGTTGTAAAACGACGGCCAGTGCCAAGCTAATTCGCTTCAAGACGTG
[0299] CTCAAATCACTATTTCCACACCCCTATATTTCTATTGCACTCCCTTTTAACTGTTTTTTATT
[0300] ACAAAAATGCCCTGGAAAATGCACTCCCTTTTTGTGTTTGTTTTTTTGTGAAACGATGTT
[0301] GTCAGGTAATTTATTTGTCAGTCTACTATGGTGGCCCATTATATTAATAGCAACTGTCGGT
[0302] CCAATAGACGACGTCGATTTTCTGCATTTGTTTAACCACGTGGATTTTATGACATTTTATAT
[0303] TAGTTAATTTGTAAAACCTACCCAATTAAAGACCTCATATGTTCTAAAGACTAATACTTAA
[0304] TGATAACAATTTTCTTTTAGTGAAGAAAGGGATAATTAGTAAATATGGAACAAGGGCAGA
[0305] AGATTTATTAAAGCCGCGGTAAGAGACAACAAGTAGGTACGTGGAGTGTCTTAGGTGAC
[0306] TTACCCACATAACATAAAGTGACATTAACAAACATAGCTAATGCTCCTATTTGAATAGTGC
[0307] ATATCAGCATACCTTATTACATATAGATAGGAGCAAACTCTAGCTAGATTGTTGAGCAGAT
[0308] CTCGGTGACGGGCAGGACCGGACGGGGCGGTACCGGCAGGCTGAAGTCCAGCTGCCA
[0309] GAAACCCACGTCATGCCAGTTCCCGTGCTTGAAGCCGGCCGCCCGCAGCATGCCGCGG
[0310] GGGGCATATCCGAGCGCCTCGTGCATGCGCACGCTCGGGTCGTTGGGCAGCCCGATGAC
[0311] AGCGACCACGCTCTTGAAGCCCTGTGCCTCCAGGGACTTCAGCAGGTGGGTGTAGAGC
[0312] GTGGAGCCCAGTCCCGTCCGCTGGTGGCGGGGGGAGACGTACACGGTCGACTCGGCCG
[0313] TCCAGTCGTAGGCGTTGCGTGCCTTCCAGGGGCCCGCGTAGGCGATGCCGGCGACCTCG
[0314] CCGTCCACCTCGGCGACGAGCCAGGGATAGCGCTCCCGCAGACGGACGAGGTCGTCCG
[0315] TCCACTCCTGCGGTTCCTGCGGCTCGGTACGGAAGTTGACCGTGCTTGTCTCGATGTAGT
[0316] GGTTGACGATGGTGCAGACCGCCGGCATGTCCGCCTCGGTGGCACGGCGGATGTCGGC
[0317] CGGGCGTCGTTCTGGGCTCATGGTAGATCCCCCGTTCGTAAATGGTGAAAATTTTCAGAA
[0318] AATTGCTTTTGCTTTAAAAGAAATGATTTAAATTGCTGCAATAGAAGTAGAATGCTTGATT
[0319] GCTTGAGATTCGTTTGTTTTGTATATGTTGTGTTGAGAATTAATTCTCGAGGTCCTCTCCA
[0320] AATGAAATGAACTTCCTTATATAGAGGAAGGGTCTTGCGAAGGATAGTGGGATTGTGCGT
[0321] CATCCCTTACGTCAGTGGAGATATCACATCAATCCACTTGCTTTGAAGACGTGGTTGGAA
[0322] CGTCTTCTTTTTCCACGATGCTCCTCGTGGGTGGGGGTCCATCTTTGGGACCACTGTCGG
[0323] TAGAGGCATCTTGAACGATAGCCTTTCCTTTATCGCAATGATGGCATTTGTAGGAGCCAC
[0324] CTTCCTTTTCCACTATCTTCACAATAAAGTGACAGATAGCTGGGCAATGGAATCCGAGGA
[0325] GGTTTCCGGATATTACCCTTTGTTGAAAAGTCTCAATTGCCCTTTGGTCTTCTGAGACTGT
[0326] ATCTTTGATATTTTTGGAGTAGACAAGTGTGTCGTGCTCCACCATGTTATCACATCAATCC
[0327] ACTTGCTTTGAAGACGTGGTTGGAACGTCTTCTTTTTCCACGATGCTCCTCGTGGGTGG
[0328] GGGTCCATCTTTGGGACCACTGTCGGCAGAGGCATCTTCAACGATGGCCTTTCCTTTATC
[0329] GCAATGATGGCATTTGTAGGAGCCACCTTCCTTTTCCACTATCTTCACAATAAAGTGACA
[0330] GATAGCTGGGCAATGGAATCCGAGGAGGTTTCCGGATATTACCCTTTGTTGAAAAGTCTC
[0331] AATTGCCCTTTGGTCTTCTGAGACTGTATCTTTGATATTTTTGGAGTAGACAAGTGTGTCG
[0332] TGCTCCACCATGTTGACCTGCAGGCATGCAAGCTTGCATGCCTGCAGGTCCCCAGATTA
[0333] GCCTTTTCAATTTCAGAAAGAATGCTAACCCACAGATGGTTAGAGAGGCTTACGCAGCA
[0334] GGTCTCATCAAGACGATCTACCCGAGCAATAATCTCCAGGAAATCAAATACCTTCCCAAG
[0335] AAGGTTAAAGATGCAGTCAAAAGATTCAGGACTAACTGCATCAAGAACACAGAGAAAG
[0336] ATATATTTCTCAAGATCAGAAGTACTATTCCAGTATGGACGATTCAAGGCTTGCTTCACAA
[0337] ACCAAGGCAAGTAATAGAGATTGGAGTCTCTAAAAAGGTAGTTCCCACTGAATCAAAGG
[0338] CCATGGAGTCAAAGATTCAAATAGAGGACCTAACAGAACTCGCCGTAAAGACTGGCGA
[0339] ACAGTTCATACAGAGTCTCTTACGACTCAATGACAAGAAGAAAATCTTCGTCAACATGG
[0340] TGGAGCACGACACACTTGTCTACTCCAAAAATATCAAAGATACAGTCTCAGAAGACCAA
[0341] AGGGCAATTGAGACTTTTCAACAAAGGGTAATATCCGGAAACCTCCTCGGATTCCATTG
[0342] CCCAGCTATCTGTCACTTTATTGTGAAGATAGTGGAAAAGGAAGGTGGCTCCTACAAAT
[0343] GCCATCATTGCGATAAAGGAAAGGCCATCGTTGAAGATGCCTCTGCCGACAGTGGTCCC
[0344] AAAGATGGACCCCCACCCACGAGGAGCATCGTGGAAAAAGAAGACGTTCCAACCACGT
[0345] CTTCAAAGCAAGTGGATTGATGTGATATCTCCACTGACGTAAGGGATGACGCACAATCC
[0346] CACTATCCTTCGCAAGACCCTTCCTCTATATAAGGAAGTTCATTTCATTTGGAGAGAACA
[0347] CGGGGGACTCTAGAATGGCCCCTAAGAAGAAGAGAAAGGTCGGTATTCACGGCGTTCC
[0348] TGCGGCGATGGACAAGAAGTATAGTATTGGTCTGGACATTGGGACGAATTCCGTTGGCT
[0349] GGGCCGTGATCACCGATGAGTACAAGGTCCCTTCCAAGAAGTTTAAGGTTCTGGGGAAC
[0350] ACCGATCGGCACAGCATCAAGAAGAATCTCATTGGAGCCCTCCTGTTCGACTCAGGCGA
[0351] GACCGCCGAAGCAACAAGGCTCAAGAGAACCGCAAGGAGACGGTATACAAGAAGGAA
[0352] GAATAGGATCTGCTACCTGCAGGAGATTTTCAGCAACGAAATGGCGAAGGTGGACGATT
[0353] CGTTCTTTCATAGATTGGAGGAGAGTTTCCTCGTCGAGGAAGATAAGAAGCACGAGAGG
[0354] CATCCTATCTTTGGCAACATTGTCGACGAGGTTGCCTATCACGAAAAGTACCCCACAATC
[0355] TATCATCTGCGGAAGAAGCTTGTGGACTCGACTGATAAGGCGGACCTTAGATTGATCTAC
[0356] CTCGCTCTGGCACACATGATTAAGTTCAGGGGCCATTTTCTGATCGAGGGGGATCTTAAC
[0357] CCGGACAATAGCGATGTGGACAAGTTGTTCATCCAGCTCGTCCAAACCTACAATCAGCT
[0358] CTTTGAGGAAACCCAAATTAATGCTTCAGGCGTCGACGCCAAGGCGATCCTGTCTGCAC
[0359] GCCTTTCAAAGTCTCGCCGGCTTGAGAACTTGATCGCTCAACTCCCGGGCGAAAAGAA
[0360] GAACGGCTTGTTCGGGAATCTCATTGCACTTTCGTTGGGGCTCACACCAAACTTCAAGA
[0361] GTAATTTTGATCTCGCTGAGGACGCAAAGCTGCAGCTTTCCAAGGACACTTATGACGAT
[0362] GACCTGGATAACCTTTTGGCCCAAATCGGCGATCAGTACGCGGACTTGTTCCTCGCGCGC
[0363] GAAGAATTTGTCGGACGCGATCCTCCTGAGTGATATTCTCCGCGTGAACACCGAGATTAC
[0364] AAAGGCCCCGCTCTCGGCGAGTATGATCAAGCGCTATGACGAGCACCATCAGGATCTGA
[0365] CCCTTTTGAAGGCTTTGGTCCGGCAGCAACTCCCAGAGAAGTACAAGGAAATCTTCTTTT
[0366] GATCAATCCAAGAACGGCTACGCTGGTTATATTGACGGCGGGGCATCGCAGGAGGAATT
[0367] CTACAAGTTTATCAAGCCAATTCTGGAGAAGATGGATGCACAGAGGAACTCCTGGTGA
[0368] AGCTCAATAGGGAGGACCTTTTGCGGAAGCAAAGAACTTTCGATAACGGCAGCATCCCT
[0369] CACCAGATTCATCTCGGGGAGCTGCACGCCATCCTGAGAAGGCAGGAAGACTTCTACCC
[0370] CTTTCTTAAGGATAACCGGGAGAAGATCGAAAAGATTCTGACGTTCAGAATTCCGTACTA
[0371] TGTCGGACCACTCGCCCGGGGTAATTCCAGATTTGCGTGGATGACCAGAAAGAGCGAG
[0372] GAAACCATCACACCTTGGAACTTCGAGGAAGTGGTCGATAAGGGCGCTTCCGCACAGA
[0373] GCTTCATTGAGCGCATGACAAATTTTGACAAGAACCTGCCTAATGAGAAGGTCCTTCCC
[0374] AAGCATTCCCTCCTGTACGAGTATTTCACTGTTTATAACGAACTCACGAAGGTGAAGTAT
[0375] GTGACCGAGGGAATGCGCAAGCCCGCCTTCCTGAGCGGCGAGCAAAAGAAGGCGATCG
[0376] TGGACCTTTTGTTTAAGACCAATCGGAAGGTCACAGTTAAGCAGCTCAAGGAGGACTAC
[0377] TTCAAGAAGATTGAATGCTTCGATTCCGTTGAGATCAGCGGCGTGGAAGACAGGTTTAA
[0378] CGCGTCACTGGGGACTTACCACGATCTCCTGAAGATCATTAAGGATAAGGACTTCTTGG
[0379] ACAACGAGGAAAATGAGGATATCCTCGAAGACATTGTCCTGACTCTTACGTTGTTTGAG
[0380] GATAGGGAAATGATCGAGGAACGCTTGAAGACGTATGCCCATCTCTTCGATGACAAGGT
[0381] TATGAAGCAGCTCAAGAGAAGAAGATACACCGGATGGGGAAGGCTGTCCCGCAAGCTT
[0382] ATCAATGGCATTAGAGACAAGCAATCAGGGAAGACAATCCTTGACTTTTTGAAGTCTGA
[0383] TGGCTTCGCGAACAGGAATTTTATGCAGCTGATTCACGATGACTCACTTACTTTCAAGGA
[0384] GGATATCCAGAAGGCTCAAGTGTCGGGACAAGGTGACAGTCTGCACGAGCATATCGCCA
[0385] ACCTTGCGGGATCTCCTGCAATCAAGAAGGGTATTCTGCAGACAGTCAAGGTTGTGGAT
[0386] GAGCTTGTGAAGGTCATGGGACGGCATAAGCCCGAGAACATCGTTATTGAGATGGCCAG
[0387] AGAAAATCAGACCACACAAAAGGGTCAGAAGAACTCGAGGGAGCGCATGAAGCGCAT
[0388] CGAGGAAGGCATTAAGGAGCTGGGGAGTCAGATCCTTAAGGAGCACCCGGTGGAAAAC
[0389] ACGCAGTTGCAAAATGAGAAGCTCTATCTGTACTATCTGCAAAATGGCAGGGATATGTAT
[0390] GTGGACCAGGAGTTGGATATTAACCGCCTCTCGGATTACGACGTCGATCATATCGTTCCT
[0391] CAGTCCTTCCTTAAGGATGACAGCATTGACAATAAGGTTCTCACCAGGTCCGACAAGAA
[0392] CCGCGGGAAGTCCGATAATGTGCCCAGCGAGGAAGTCGTTAAGAAGATGAAGACTAC
[0393] TGGAGGCAACTTTTGAATGCCAAGTTGATCACACAGAGGAAGTTTGATAACCTCACTAA
[0394] GGCCGAGCGCGGAGGTCTCAGCGAACTGGACAAGGCGGGCTTCATTAAGCGGCAACTG
[0395] GTTGAGACTAGACAGTCACGAAGCACGTGGCGCAGATTCTCGATTCACGCATGAACAC
[0396] GAAGTACGATGAGAATGACAAGCTGATCCGGGAAGTGAAGGTCATCACCTTGAAGTCA
[0397] AAGCTCGTTTCTGACTTCAGGAAGGATTTCCAATTTTATAAGGTGCGCGAGATCAACAAT
[0398] TATCACCATGCTCATGACGCATACCTCAACGCTGTGGTCGGAACAGCATTGATTAAGAAG
[0399] TACCCGAAGCTCGAGTCCGAATTCGTGTACGGTGACTATAAGGTTTACGATGTGCGCAA
[0400] GATGATCGCCAAGTCAGAGCAGGAAATTGGCAAGGCCACTGCGAAGTATTTCTTTTACT
[0401] CTAACATTATGAATTTCTTTAAGACTGAGATCACGCTGGCTAATGGCGAAATCCGGAAGA
[0402] GACCACTTATTGAGACCAACGGCGAGACAGGGAAATCGTGTGGGACAAGGGGAGGGA
[0403] TTTCGCCACAGTCCGCAAGGTTCTCTCTATGCCTCAAGTGAATATTGTCAAGAAGACTGA
[0404] AGTCCAGACGGGCGGGTTCTCAAAGGAATCTATTCTGCCCAAGCGGAACTCGGATAAGC
[0405] TTATCGCCAGAAAGAAGGACTGGGACCCGAAGAAGTATGGAGGTTTCGACTCACCAAC
[0406] GGTGGCTTACTCTGTCCTGGTTGTGGCAAAGGTGGAGAAGGGAAAGTCAAAGAAGCTC
[0407] AAGTCTGTCAAGGAGCTCCTGGGTATCACCATTATGGAGAGGTCCAGCTTCGAAAAGAA
[0408] TCCGATCGATTTTCTCGAGGCGAAGGGATATAAGGAAGTGAAGAAGGACCTGATCATTA
[0409] AGCTTCCAAAGTACAGTCTTTTCGAGTTGGAAAACGGCAGGAAGCGCATGTTGGCTTCC
[0410] GCAGGAGAGCTCCAGAAGGGTAACGAGCTTGCTTTGCCGTCCAAGTATGTGAACTTCCT
[0411] CTATCTGGCATCCCACTACGAGAAGCTCAAGGGCAGCCCAGAGGATAACGAACAGAAG
[0412] CAACTGTTTGTGGAGCAACACAAGCATTATCTTGACGAGATCATTGAACAGATTTCGGA
[0413] GTTCAGTAAGCGCGTCATCCTCGCCGACGCGAATTTGGATAAGGTTCTCTCAGCCTACAA
[0414] CAAGCACCGGGACAAGCCTATCAGAGAGCAGGCGGAAAATATCATTCATCTCTTCACCC
[0415] TGACAAACCTTGGGGCTCCCGCTGCATTCAAGTATTTTGACACTACGATTGATCGGAAG
[0416] AGATACACTTCTACGAAGGAGGTGCTGGATGCAACCCTTATCCACCAATCGATTACTGGC
[0417] CTCTACGAGACGCGGATCGACTTGAGTCAGCTCGGGGGGGATAAGAGAGACCAGCGGCAA
[0418] CCAAGAAGGCAGGACAAGCGAAGAAGAAGAAGTAGGTTAACGAATTTCCCCGATCGTT
[0419] CAAACATTTGGCAATAAAGTTTCTTAAGATTGAATCCTGTTGCCGGTCTTGCGATGATTAT
[0420] CATATAATTTCTGTTGAATTACGTTAAGCATGTAATAATTAACATGTAATGCATGACGTTAT
[0421] TTATGAGATGGGTTTTTATGATTAGAGTCCCGCAATTATACATTTAATACGCGATAGAAAA
[0422] CAAAATATAGCGCGCAAACTAGGATAAATTATCGCGCGCGGTGTCATCTATGTTACTAGAT
[0423] CGGGAATTCTTAATTAAGGATCCGTAATCATGTCATAGCTGTTTCCTGTGTGAAATTGTTA
[0424] TCCGCTCACAATTCCACACAACATACGAGCCGGAAGCATAAAGTGTAAAGCCTGGGGTG
[0425] CCTAATGAGTGAGCTAACTCACATTAATTGCGTTGCGCTCACTGCCCGCTTTCCAGTCGG
[0426] GAAACCTGTCGTGCCAGCTGCATTAATGAATCGGCCAACGCGCGGGGAGAGGCGGTTT
[0427] GCGTATTGGAGCTTGAGCTTGGATCAGATTGTCGTTTCCCGCCTTCAGGCGCGCCGTTTA
[0428] AACTATCAGTGTTTGACAGGATATATTGGCGGGTAAACCTAAGAGAAAAGAGCGTTTATT
[0429] AGAATAATCGGATATTTAAAAGGGCGTGAAAAGGTTTATCCGTTCGTCCATTTGTATGTGCATGCCAACCACAGGGTTCCCCTCGGGATCAA-3'
[0430] The mutated genome sequence of the Gmcol2a-78 mutant
[0431]
[0432] The mutated coding sequence (CDS) in the Gmcol2a-78 mutant.
[0433] 5'-CATGCGTACTGTACTGCCATGCAGACTCAGCATACCTTTGCTCATCCTGCGATGCTCGTGTCCATGCGGCTAACCGTGTGGCCTCAAGACACGAGCGTGTGTGGGTGTGCGAAGCGTGTGAACGTGCTCCCGCAGCGTTTCTATGCAAAGCCGACGCAGCTTCTCTTTGTTCTTCCTGTGATGCTGACATTCACTCAGCAAACCCTCTCGCTAGCCGCCACCACCGCGTGCCCATTCTCCCGATCTCTGGTTCCCTCTTCGGGGAACCAGAGCATGAACGCGTGTACGCGTTCGTGAATGAAGTGGAAGCGGAGGAGGAAGAGGAAGAGGTTTTTGATGAGTATGATGAGGTTGAAGCAGCTTCGTGGTTGTTGCCACATCCTATGAAAAATGATAAAATTGACGAGAATGGTGGTGATAAGGGTTTTTTGTTTGGTGATGAGTATTTCGACAACCTTGTTGATTGTAACTCATGTGGTCACAATAACAACCAGTTTAGCAACGTTTATGATCAGCACCAGCAGAATTACAGCAACACTGTCCCTCAGAACTATGCAGTGGTTCCAGTTCAGGTGCCGCAGCATTTTCAACCGGGTTTGGACTTTGACTCATCAAAAGCTGGGTTCAGTTACGATGGTTCTCTTAGTCAAAGTGTTTCGGTTTCATCGATGGATGTTGGTGTTGTACTCGAATCAACAATAAGTGACATCTCAATGTCCCACTCAAAGTCGCCAATAGGGACAACTGACCTATTTCCTCCCCTTCCCATGCCTTCACATCTCACACCAATGGACAGAGAGGCAAGAGTCCTAAGATACAGGGAGAAAAAGAAGACAAGAAAATTTGAGAAGAAAATAAGGTATGCCTCAAGGAAGGCCTATGCAGAGACTAGACCCCGCATAAAGGGTCGTTTTGCAAAGAGAACCGATGTAGAAGCTGAAGTGGACCAGATGTTCTCCACAACACTATTCACTGAAGTTGGAGGTAGCATTTTTCCCACTTTCTAG-3'
[0434] Mutated amino acid sequence in the Gmcol2a-78 mutant
[0435] MRTVLPCRLSIPLLILRCSCPCG
[0436] Mutated genomic sequence in the Gmcol2b-35 mutant
[0437] 5'-TAAAAAAAAAGTATCATATTTAGAGAGACTAAAAGATTATTTAAGTCAAAAAATA ATTTTTGTTTAGCATTTGTTCTCACCGAGTTCTAGGCAAGTGATCCGTCTTCGTCTTCCCTTTAGTAGTTGAACAGGTCAATTGGTTAAACCACAGATTGATGGAAGTGATATAATGATTAAATATGCGATATATATTACTTGTGTCACATATTAATGAAATTGTGCGGATGGTCTAGTGGTAAATATACTAGAAAGATTTTCCTAGACAAAGATCGACTCTATGGTTGGTAATTTACTTTTACAATTTCCTTAAAATAGTCAAGTAACTAGCTGTAAACCTCAGCAAATTTAGTACTTGGAGCAACTATGAACAATTAATGATATTTATTTTAAACATTCCACCTTAATTAACATCTTAGTTCTC
[0438] CAAACTCCACCAACCAAAAATTTCTATTTATTTTCATAATCACAAGCACTCCACATTCTAT
[0439] CTGCAAAATCAGTAATTAACCAAACACTTGCTTTAATCCTTTTAGAGCCAATAAGATTAC
[0440] TGAGAGGAAAGGATATATTAAGGGGTAACAAGAGATTCTAACTGGTGCAAGTATAGAAT
[0441] CTACATGGACCCAAATGACACGTGTCTCCAAGTTGTGTCACCATAGCAATACGAACAAA
[0442] GCCTCAGCATCTCTTTAATGATTCCAAAAGGCAAAGAAGCAACATGCATGGACAGGTCC
[0443] AACACCAAAACAAAAAACCCTCTCTCCTAATCAACCCTGAAACACAACCAACCACAAG
[0444] ATAAGCAAGAACTTGTGCACGCCACTAACTTGCTGCCACGTTGGCACTCTTGATTCTCCC
[0445] CCCAACACTACTTGGTTCACCTCACTAAACTCAATTTCTCACTCACACTCACTTCAGTTC
[0446] AATTCACTTCCCTTCCTCCAAATTAGTTCCAACACAAAAACCAAGCAAAGAAAAAGACT
[0447] CTACCGAAACACTTCACTACTCATACAATATTTTCAGACACAACATGTTGAAGGAAGGC
[0448] ACCAACAACGTTGGTGGCAGCAACACTGGCACACCTGGTCACGTGTCTGTGACACGTG
[0449] CCTGTCTGCGCCATGCGTGCTGTACTGCCATGCAGACTCAGCATATCTTTGCTCTTCCTGT
[0450] GATGCTCGTGTGCATGCAGCTAACCGTGTGGCCTCGAGACACAAGCGTGTGTGGGTGTG
[0451] CGAAGCGTGTGAGCGTGCTCCGGCGGCGTTTCTATGCAAAGCTGACGCAGCTTCTCTTT
[0452] GTTCTTCCTGTGATGCTGACATTCACTCAGCAAACCCTCTCGCTAGCCGCCACAACCGTG
[0453] TTCCCATTCTCCCGATCTCAGGTTCCCTCTTCAGGGAACCAGAGCACAATCACAAACGC
[0454] GTGGAACACGCGTTCGTGAATGAGGTTGAGGAGGAAGAGGAAGGGGTTTTTGATGAGT
[0455] ATGAAGACGAGGTTGAAGCTGCTTCGTGGTTGTTGCCACATCCTATGAAAAATAATGATG
[0456] AAATTGAGGAGAATGATTGTGGTGACGAGGGTTTTTTGTTTGTTGATGAGTATTTGGACA
[0457] ATCTTGTTGATTGTTGTAACTCATGTGGTCACAATGACAACCAGTTTAGCAACGTTTATC
[0458] AGCACCAGCAGAATTACAACACTGTCCCTCAGAACTATGTAGTGGTTCCAGTTCAGGTG
[0459] CCGCAGCATTTTCAACCCGGTTTGGACTTTGACTCATCAAAAGCTGGGTTCAGTTACGAT
[0460] GGTTCTCTTAGTCAAAGTGTAAGACTCTTCTTTTCAATTACTTCTGTGTTTCTTTTTCGGT
[0461] TTGTTTTTTATATGAATAGTGTTTTCCCTTTTTCTCTTTCTTTTTCTTTTCCCAGACTCTTTG
[0462] GATCATTCATCCATTGGATTATACTTTATCAGATTTTGATTTTGCCTTGACACCTTTGCTAA
[0463] CTCTTCCTTCCAATTGGCAGTTTCTTTCCTTTCTCCGTTCATATGAAACTAGGGTTTTAAA
[0464] TTATACTCATGATCAATAATTTGGTTGCAATTTTTGATATTGTAGTACAATTGCTATTGTTG
[0465] TTGTGAAAAGTTAAAAAAATCTTAATGTTGTAACCGAAATCTCAATCACATACTGTTTGT
[0466] AAAATCCTTGGTTGAAATCAAAGTGGGTTTCCCAAGTAGCTTCTTTTACTAGAAACGCTA
[0467] GCAAAGTCAACTACTCTAGTTCAATTGCCATATATGTGACTAATTTGATTCATTCCACAAC
[0468] TCTTCATATTTCTTTGACTACCCTTTACCGTTCACTTTGCGGTCTTGAGAGACTTTTATTT
[0469] ATTTCTTTTTTGTCGTATATAAGATCAGCTTTGCCTTTATGGTGTCTTTTATTTGCTAGTTA
[0470] TAGTTGTGATTCCTCTTAAACTTTGTTTCCTCATTTCATCAACTGTATATTAATTTCCCATA
[0471] ATTTTCAGCTTCAATTAAAAAAATTGAGCCTAAGTGCCTGATTTATTAAAGCCACAAAA
[0472] TGGGTGGTTTAATTTTTTTCGACACCTTTCTTTTATATTTCTTTTACTCTTCTTCTCTTTTCTT
[0473] TTTCTATTACATCATCTATAATATATCTTCACCTTTTTCTTCTTACTTTGGGGACAAACATTT
[0474] TTGTTTCACAAAGCTGGTTAGCTACTAGTTTATTGCTTAGAAAATTTATAGCCTTACCTGT
[0475] GATCTCTCAAACTTCATACACGATTTTGTGTTAATCTTTTATATTTCATTTTTAATTTAGGTT
[0476] TCGGTTTCATCAATGGATGTTGGAGTTGTACCCGAATCAACAGTAAGTGGCATCTCAATG
[0477] TCCCACTCAAAGTCACCAATAGGGACAAATGACCTATTTCCTCCCCTTCTCATGCCTTCA
[0478] CATCTCACACCAATGGACAGAGAGGCAAGAGTCCTAAGATATAGGGAGAAAAAGAAGA
[0479] CAAGAAAGTTCGAGAAGAAAATAAGATATGCCTCAAGGAAGGCCTATGCAGAGACTAG
[0480] ACCCCGCATAAAGGGTCGTTTTGCAAAGAGAACTGATGTAGAAGCTGAAGTGGATCAGA
[0481] TGTTCTCCACAAAACTATTCAATGAAGTTGGAGGTAGCATTTTTCCCACTTTCTAGAATG
[0482] GAAAGAAAGTAATTGCCAAGCCATAGAAGAGAAGGTTGGGAACCACACATCATTGGAA
[0483] GAACTTCATGTAATTTTTATGTAGTCTAATTAGTTTGGTATTATATTTACTTCTATCTTGTGCTTTATAAATATAAAGTAT-3'
[0484] The mutated coding sequence (CDS) in the Gmcol2b-35 mutant
[0485]
[0486] The mutated amino acid sequence in the Gmcol2b-35 mutant
[0487] MLKEGTNNVGGSNTGTPGHVSVTRACLRHACCTAMQTQHIFALPVMLVCMQLTVWPRDTSVCGCAKRVSVLRRRFYAKLTQLLFVLPVMLTFTQQTLSLAATTVFPFSRSQVPSSGNQSTITNAWNTRS
[0488] The present invention has been described in detail above. Those skilled in the art will recognize that the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. While specific embodiments have been provided, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein.
Claims
1. The use of a substance that inhibits, reduces, or downregulates the expression of the gene encoding the protein GmCOL2a or GmCOL2b, or a substance that inhibits, reduces, or downregulates the content of said protein GmCOL2a or GmCOL2b, in any of the following: D1) Increase soybean yield; D2) Reduces soybean plant height; D3) Prepare products that reduce soybean plant height; D4) Cultivating semi-dwarf soybeans; D5) Preparation of products from semi-dwarf soybeans; D6) Increase the number of pods per soybean plant; D7) Prepare products that increase the number of pods per soybean plant; D8) Cultivate soybeans with an increased number of pods per plant; D9) Prepare soybean products with an increased number of pods per plant; D10) increases the number of soybean grains per plant; D11) Prepare products that increase the number of soybean grains per plant; D12) Cultivating soybeans with increased number of seeds per plant; D13) Prepare soybean products with increased number of grains per plant; D14) Cultivate soybeans with reduced plant height and increased number of pods and seeds per plant. The protein is any one of the following: A1) The amino acid sequence is GmCOL2a as shown in SEQ ID No. 2; A2) The amino acid sequence is GmCOL2b as shown in SEQ ID No. 4; A3) A fusion protein with the same function obtained by attaching a tag to the N-terminus and / or C-terminus of A1) or A2); The substance is any one of the following: B1) Nucleic acid molecules that inhibit, reduce, or downregulate the expression of the gene encoding the protein GmCOL2a or GmCOL2b, or nucleic acid molecules that inhibit, reduce, or downregulate the content of the protein GmCOL2a or GmCOL2b. B2) Genes that express the nucleic acid molecules described in B1); B3), an expression cassette containing the gene described in B2); B4), a recombinant vector containing the gene described in B2), or a recombinant vector containing the expression cassette described in B3); B5) Recombinant microorganisms containing the gene described in B2), or recombinant microorganisms containing the expression cassette described in B3), or recombinant microorganisms containing the recombinant vector described in B4); B6) A transgenic plant cell line containing the gene described in B2), or a transgenic plant cell line containing the expression cassette described in B3), or a transgenic plant cell line containing the recombinant vector described in B4); B7) Transgenic plant tissue containing the gene described in B2), or transgenic plant tissue containing the expression cassette described in B3), or transgenic plant tissue containing the recombinant vector described in B4); B8) A transgenic plant organ containing the gene described in B2), or a transgenic plant organ containing the expression cassette described in B3), or a transgenic plant organ containing the recombinant vector described in B4).
2. A method for increasing soybean yield, reducing soybean plant height, increasing the number of pods per soybean plant and / or increasing the number of soybean grains, comprising downregulating or reducing or weakening the expression level of the gene encoding the protein GmCOL2a or GmCOL2b described in claim 1 or the content of the protein GmCOL2a or GmCOL2b in the target soybean to obtain transgenic soybeans, wherein the transgenic soybeans have a lower plant height than the target soybeans, and the transgenic soybeans have a higher number of pods per plant and a higher number of soybean grains per plant than the target soybeans; wherein downregulating or reducing or weakening the expression level of the gene encoding the protein GmCOL2a or GmCOL2b described in claim 1 or the content of the protein GmCOL2a or GmCOL2b in the target soybeans is achieved by introducing the gene described in claim 1 (B2) into the target soybeans.