Plant genome repetitive sequence deletion technology

By designing specific sgRNA and using CAS9/CAS12 enzymes to accurately shear redundant repeats in the genome of wheat and other plants, the problem of redundant genome sequencing data is solved, the sequencing efficiency and analysis efficiency are improved, and the cost is reduced.

CN120060261APending Publication Date: 2025-05-30CHENGDU TIANCHENG SMART AGRICULTURAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510319459.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

There are a large number of redundant repeat sequences in the genomes of complex genomic crops such as wheat, resulting in redundant genomic sequencing data, increasing the difficulty of data storage and analysis, and it is difficult for the prior art to efficiently edit these redundant sequences in vitro.

Method used

Design specific sgRNAs through bioinformatics methods, combine CAS9/CAS12 enzymes, accurately cut repeat sequences in the genomes of wheat and other plants, remove redundant sequences, reduce redundancy in sequencing data, and improve sequencing efficiency and analysis efficiency.

Benefits of technology

Effectively remove redundant repeat sequences in the genome, reduce redundancy in sequencing data, reduce data storage and analysis costs, and improve the efficiency of genomic information analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention relates to the technical field of plant genomics and gene editing, in particular to an in-vitro genome editing technology aiming at wheat and other complex genome crops, and provides a plant genome repetitive sequence deletion technology and application thereof. According to the method, specific sgRNA is designed, and CAS9 / CAS12 enzyme is guided to precisely shear repetitive sequences in a target genome, so that redundant repetitive sequences in the genome are effectively removed, sequencing data redundancy is reduced, and the efficiency and precision of genome information analysis are improved. The method comprises the following steps: firstly, identifying and annotating a repetitive sequence region by analyzing genome data of crops such as wheat; then, designing an sgRNA sequence aiming at the repetitive sequence, and optimizing the specificity and amplification efficiency of the sgRNA sequence; then, synthesizing an sgRNA library by utilizing a high-throughput DNA synthesis technology, and carrying out an in-vitro shearing reaction together with CAS9 / CAS12 enzyme; and finally, sequencing the sheared genome through a high-throughput sequencing platform, and analyzing the shearing efficiency and accuracy. The technology not only improves the efficiency of genome simplification and redundant sequence removal, but also can provide powerful data support for genome research and application of crops such as wheat.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of plant genomics and gene editing technology, especially an in vitro genome editing technology for crops with complex genomes such as wheat. Specifically, it is a technology for deleting repetitive sequences in plant genomes, aiming to reduce the demand for genomic sequencing data, improve the efficiency of genomic information analysis, and greatly reduce the detection cost and analysis cost by removing redundant repetitive sequences in the genome. Background Art

[0002] There are a large number of repetitive sequences in the genomes of crops such as wheat. These repetitive sequences not only occupy most of the genome space but also cause data redundancy during whole-genome sequencing, increasing the difficulty of data storage and analysis.

[0003] Although existing technologies have conducted some research on genomic repetitive sequences, most of them focus on in vivo editing and operate on single genes or small ranges of sequences. There is still a lack of in vitro editing technology for free DNA. Efficiently deleting redundant repetitive sequences in vitro and ensuring the integrity of functional genomic information are still a difficult point in current genomics research. Summary of the Invention

[0004] In view of the above-mentioned disadvantages of the prior art, the present invention provides an operation method for deleting repetitive sequences in plant genomes. Specific sgRNAs are designed by bioinformatics methods to guide the CAS9 / CAS12 enzyme to precisely cut repetitive sequences in the genomes of plants such as wheat. This technology can effectively remove repetitive sequences, thereby reducing data redundancy during sequencing, improving sequencing efficiency, and reducing the costs of data storage and analysis.

[0005] Technical Solution

[0006] Step 1: Use the publicly published plant genomic data that has been completed to perform the annotation work of repetitive sequences. Identify and annotate the repetitive sequences in the genomes of plants such as wheat through methods such as sequence alignment (such as using BLAST, Bowtie, or BWA for genome alignment), repetitive sequence identification (such as using RepeatMasker, Tandem Repeats Finder, MISA to detect microsatellite sequences), and genomic feature analysis (such as calculating GC content and evaluating sequence coverage).

[0007] Step 2: According to the repetitive sequence information obtained in Step 1, use bioinformatics methods (such as RepeatMasker and TRF for repetitive sequence analysis, CRISPR-P and CHOPCHOP for sgRNA design, Cas-OFFinder for off-target effect prediction, and RNAfold for secondary structure analysis) to design sgRNA sequences targeting the common repetitive sequences of wheat crops. Also, comprehensively consider factors such as the specificity, GC content, and sequence length of the sgRNA to ensure its efficient and accurate recognition of the target sequence.

[0008] Step 3: Optimize and group the designed sgRNA sequences according to the characteristics of the sgRNA sequences to ensure efficient amplification in the same reaction system and avoid interference between different sequences.

[0009] Step 4: Use high-throughput DNA synthesis technologies (such as microarray synthesis, silicon wafer synthesis, liquid-phase synthesis, and enzymatic synthesis) to efficiently synthesize the grouped sgRNA sequences and assemble them to finally prepare a high-density sgRNA library. These technologies use methods such as solid-phase or liquid-phase chemical synthesis and enzymatic synthesis to parallelly synthesize a large number of DNA sequences in a short time, ensuring that each sgRNA sequence can cover different regions of the target repetitive sequence, thereby enhancing the comprehensiveness and efficiency of in vitro cleavage.

[0010] Step 5: Combine the high-density sgRNA library prepared in Step 4 with the CAS9 / CAS12 enzyme and perform an in vitro cleavage reaction using the in vitro CRISPR-CAS9 / CAS12 enzyme cleavage system. This reaction is usually carried out in a suitable buffer system, including the CAS9 / CAS12 enzyme, sgRNA, target DNA template, and necessary ionic conditions (such as Mg²⁺). Under the guidance of the sgRNA, the CAS9 / CAS12 enzyme precisely recognizes and cleaves the repetitive sequences in the DNA library, thus efficiently completing the DNA cleavage of the target region and providing accurate fragmented DNA products for subsequent analysis.

[0011] Step 6: Sequence the DNA library after the in vitro cleavage reaction through a high-throughput sequencing platform to obtain data on non-cleaved fragments. Further analyze the sequencing data to confirm the efficiency and accuracy of DNA cleavage and identify the variations generated after cleavage. Description of the Drawings

[0012] Figure 1 It shows the flow chart of the plant genome repetitive sequence deletion technology in Example 1 of the present invention.

[0013] Figure 2 It shows the key flow chart of the repetitive sequence deletion technology in Example 1 of the present invention.

[0014] Figure 3Shown is the result comparison of genome sequencing using this technical invention in Embodiment 1 of the present invention. Detailed implementation manners

[0015] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in this technology can easily understand other advantages and effects of the present invention from the content disclosed in this specification.

[0016] Embodiment 1:

[0017] A plant genome repetitive sequence deletion technology includes the following steps:

[0018] Step 1: Collect the genome information of target crops (such as Chinese Spring wheat, Aikang 58, Jimai 22, etc.). Analyze the repetitive sequence regions according to the genome file, and identify and mark the repetitive sequence regions. The genome data of the target crops is obtained from public databases such as NCBI, Ensembl, etc., and software such as IGV (Integrative Genomics Viewer), BEDTools, etc. are used for genome visualization and analysis. According to the characteristics of the genome sequence and repetitive sequences, use alignment algorithms to identify the repetitive regions and obtain the target regions to be edited.

[0019] Step 2: Design a group of sgRNA sequences targeting these repetitive regions according to the characteristics of the repetitive sequences. Use CRISPR design tools (such as CRISPOR, Benchling, etc.) to screen these repetitive sequences to ensure that the sgRNA can specifically target the repetitive sequences without affecting other non-repetitive regions. When designing, ensure that the length of the sgRNA is between 30bp and 80bp, and the number of copies in the genome is not less than 100. The designed sgRNA sequences are synthesized by a synthesis company.

[0020] Step 3: Synthesize the designed sgRNA sequences using synthetic technology and construct a library. The T7 RNA synthesis kit (Thermo Fisher, product number: AM1334) is used to synthesize the sgRNA. The synthesis reaction system includes: T7 RNA polymerase (5μL, 5U / μL), RNA synthesis buffer (5μL), NTPs (5μL, 10mM for each NTP), and template DNA (10μL, 0.5μg / μL). The NEBNext Ultra II DNA Library Prep Kit (product number: E7645) is used to construct the library, efficiently construct the library, and perform library purification, ligation, and enrichment. Finally, a high-density library containing multiple sgRNAs is obtained.

[0021] Step 4: Clone the synthesized sgRNA library into a plasmid vector suitable for expression to ensure that each sgRNA can be effectively expressed in in vitro experiments. Select plasmid vectors such as pUC57, pX330, etc. Use T4 DNA ligase (NEB, catalog number: M0202) to ligate the sgRNA library to the vector. The amount of T4 DNA ligase used in the reaction system is 1 μL (5 U / μL). Transform it into Escherichia coli (such as DH5α) cells, and use the heat shock transformation method for transformation. The transformed cells are cultured in LB medium for screening and amplification.

[0022] Step 5: Synthesize or purchase Cas9 / Cas12 proteins to ensure that they have sufficient activity for genome cleavage. The Cas9 protein is purchased from Addgene (catalog number: 62921) or customized through a synthesis company. The working concentration of the Cas9 protein is 500 ng / μL, and it is stored and diluted using PBS buffer (containing 1 mM DTT). To improve the editing efficiency, variant Cas9 such as nCas9 or eCas9 is selected, which has higher cleavage specificity and efficiency.

[0023] Step 6: Prepare the CRISPR / Cas9 in vitro editing reaction system according to the designed sgRNA library. The reaction system includes: Cas9 enzyme (500 ng / μL), sgRNA (100 nM), target DNA template (1 μg), NEBuffer™ 3.1 for Cas9 (10 μL), MgCl 2 2 (10 mM, 1 μL), and the total volume of the reaction system is 50 μL. The reaction system is incubated at 37 °C for 1 - 2 hours to allow the Cas9 enzyme to bind to the sgRNA and perform DNA cleavage.

[0024] Step 7: Incubate the in vitro editing system under suitable conditions. The Cas9 enzyme targets and cleaves the repetitive sequences in the DNA according to the guidance of the sgRNA. Optimize the reaction temperature, enzyme concentration, and reaction time. The reaction conditions are incubation at 37 °C for 1 - 2 hours, the concentration of the Cas9 enzyme ranges from 100 ng / μL to 1 μg / μL, and the reaction time is from 30 minutes to 2 hours to ensure that the reaction proceeds fully.

[0025] Step 8: Repair the sheared DNA fragments by non-homologous end joining (NHEJ) repair or homologous recombination repair (HDR). For NHEJ repair, use NEB T4 DNA Ligase (Catalog No.: M0202), with a dosage of 1 μL (5 U / μL) of T4 DNA ligase, 2 μL of T4 DNA ligation buffer, and 500 ng of the repaired DNA fragments in a total reaction volume of 20 μL. For homologous recombination repair, use NEB TrueCut™ Cas9 HDR Kit (Catalog No.: E1321S), and incubate at 37°C for 1 hour. After the repair is completed, the obtained DNA fragments proceed to the next step of analysis and screening.

[0026] Step 9: Remove the redundant fragments generated during the in vitro editing process by restriction enzyme digestion or PCR technology, and only retain the DNA fragments that have been effectively sheared and repaired. Digest the repaired DNA fragments with appropriate restriction enzymes (such as EcoRI, BamHI), with a dosage of 1 μL (10 U / μL) of the enzyme in the reaction system, use NEBuffer™ 2.1 for the reaction, and the reaction time is 1 hour. Alternatively, use PCR technology to amplify the target region. In the PCR reaction system, use Taq DNA Polymerase (Thermo Fisher, Catalog No.: EP0401), the primer concentration is 0.5 μM, and the dNTPs (Thermo Fisher, Catalog No.: R0192) concentration is 0.2 mM, with a total reaction volume of 50 μL.

[0027] Step 10: Perform high-throughput sequencing on the in vitro edited DNA fragments using BGIseq of BGI to obtain detailed information on the genome after the repetitive sequences are sheared. Use the BGISeq-500 platform for sequencing to ensure that the coverage and depth meet the analysis requirements. The library construction and sequencing process includes the library construction of DNA fragments using NEBNext Ultra II DNA Library Prep Kit (Catalog No.: E7645), and then perform high-throughput sequencing. The sequencing data is subjected to preliminary quality control and filtering through the online platform or data analysis software provided by BGI.

[0028] Step 11: Use bioinformatics methods to align and analyze the sequencing data to confirm whether the repetitive sequences have been successfully sheared and removed, and at the same time evaluate the editing efficiency and accuracy. Use software such as BLAST, Bowtie, and BWA to align the sequencing data, and use software such as IGV to view the editing results. The alignment results are used to confirm whether the repetitive sequences have been successfully removed, and to evaluate the editing accuracy and off-target effects.

[0029] Step 12: Use techniques such as PCR and qPCR to verify the deletion effect of the repetitive sequences and evaluate the integrity of the genome. In the PCR reaction system, Taq DNA Polymerase (product number: EP0401) is used, the primer concentration is 0.5 μM, the dNTPs concentration is 0.2 mM, and the total volume of the reaction system is 50 μL. For the integrity of the genome, qPCR method is used for quantitative detection. Power SYBR® Green PCR Master Mix (product number: 4368706) is used, appropriate primers and probes are set, and quantitative analysis is carried out to ensure that there are no unexpected changes in the structure of the edited genome.

[0030] Example 2:

[0031] A plant genome repetitive sequence deletion technology includes the following steps:

[0032] Step 1: Collect the complete genome sequences of target crops (such as maize B73, Mo17, etc.), and annotate and analyze the repetitive sequences in the genome. First, download the complete genome data of maize B73 and Mo17 from the NCBI or Ensembl database, ensure that the data file format is in FASTA format, and verify the sequence integrity. Subsequently, use RepeatMasker (version 4.1.2, embedded with Repbase database) to scan the genome, set the species parameter as "Zea mays", the threshold parameter as the minimum match length of 100 bp, identify and annotate all repetitive sequence regions, and at the same time count the types and distribution of various repetitive sequences (such as dispersed repeats, tandem repeats, transposons, etc.).

[0033] Step 2: Based on the characteristics of the repetitive sequences, use CRISPR design software to design appropriate sgRNA sequences for each repetitive sequence region. Use CHOPCHOP (website: http: / / chopchop.cbu.uib.no / ) or CRISPR-ERA for design, input the target repetitive region sequence, and set the sgRNA length to be 30 bp to 80 bp. The screening criteria require that the copy number of each sgRNA in the whole genome is not less than 100, and the GC content is controlled at 40% - 60% to ensure the targeting specificity and effectiveness. After the design results are exported, number and record each sgRNA sequence for subsequent synthesis and experimental tracking.

[0034] Step 3. According to the designed sgRNA sequences, synthesize the required sgRNA by chemical synthesis or in vitro transcription. Use the Thermo Fisher T7 RNA Synthesis Kit (Catalog No. AM1334) to prepare the in vitro transcription system. Add 5 μL of T7 RNA polymerase (5 U / μL), 5 μL of RNA synthesis buffer, 5 μL of each NTP (10 mM), and 10 μL of template DNA (0.5 μg / μL) to the reaction system, and adjust the total reaction volume to 50 μL. React according to the kit instructions and incubate at 37°C for 2 hours. Group the synthesized sgRNAs according to the GC content and sequence length to optimize the subsequent in vitro cleavage efficiency.

[0035] Step 4. Clone all the designed sgRNA sequences into a suitable vector to construct a high-density library containing multiple sgRNAs targeting repetitive sequences in the target genome. Select the pX330 vector (Addgene, Catalog No. 42230) as the cloning vector, use PCR amplification to obtain double-stranded DNA containing the sgRNA insert fragment, and perform a ligation reaction in a 20 μL reaction system using T4 DNA ligase (NEB, Catalog No. M0202, 1 μL, 5 U / μL). The ligation reaction is maintained at room temperature for 30 minutes. Subsequently, heat shock transform the ligation product into DH5α Escherichia coli and culture it in LB medium in a shaker at 37°C for 16 hours to obtain a cloned strain containing a high-density sgRNA library.

[0036] Step 5. Select a highly efficient CAS9 / CAS12 enzyme or its variant, and obtain an active and stable Cas9 protein through in vitro expression or purification. Order Streptococcus pyogenes Cas9 from Addgene (Catalog No. 62934), adjust the concentration to 500 ng / μL after dissolution, formulate it in PBS buffer (containing 1 mM DTT), and store it at -80°C. Operate strictly according to the manufacturer's recommendations to ensure that the Cas9 protein has optimal activity and stability in subsequent in vitro editing.

[0037] Step 6. Mix the sgRNA library with the Cas9 enzyme to form a CRISPR-Cas9 complex, and add the target genomic DNA or free DNA to the in vitro reaction system to perform the CRISPR-Cas9 in vitro cleavage reaction. Take 1 μL of Cas9 enzyme (500 ng / μL), 1 μL of sgRNA (100 nM), 1 μg of target genomic DNA, add 10 μL of NEBuffer™ 3.1 for Cas9 and MgCl 21 μL (10 mM), and finally adjust the reaction system to 50 μL. After mixing evenly, incubate at 37°C for 1 to 2 hours to allow the Cas9 enzyme to introduce double-strand breaks in the repetitive sequence region and precisely excise the redundant repetitive sequences.

[0038] Step 7. Screen and remove the genomic fragments with redundant sequences using PCR or restriction enzyme digestion methods. Use the restriction enzyme EcoRI (NEB, 1 μL, 10 U / μL) to digest the in vitro reaction products in the NEBuffer™ 2.1 reaction system, and set the reaction conditions as 37°C for 1 hour; or use PCR amplification, use Thermo Fisher Taq DNA Polymerase (product number EP0401), set the primer concentration at 0.5 μM, the dNTPs concentration at 0.2 mM, the total reaction volume at 50 μL, and cycle 35 times to amplify the target region, and then detect the specificity and integrity of the amplified fragments by gel electrophoresis.

[0039] Step 8. Repair the sheared genomic DNA through the homologous recombination or non-homologous end joining (NHEJ) mechanism. If NHEJ repair is used, add NEB T4 DNA Ligase (1 μL, 5 U / μL, product number M0202) and T4 DNA ligation buffer (2 μL), and react at room temperature for 1 hour in a 20 μL reaction system; if homologous recombination repair is used, use NEBTrueCut™ Cas9 HDR Kit (product number E1321S), and react at 37°C for 1 hour to ensure correct repair of the broken ends and generate continuous genomic fragments.

[0040] Step 9. Recover the genomic fragments with redundant sequences removed using gel recovery or column purification methods, and thoroughly remove the residual enzymes, sgRNA and other impurities in the in vitro reaction. After separating the PCR products by 1% agarose gel electrophoresis, use the Qiagen Gel Extraction Kit (product number 28704) to extract the target fragments, or use the Qiagen PCRPurification Kit (product number 28104) to perform column purification on the mixture, and finally adjust the purified product to 30 μL to ensure the purity meets the requirements for subsequent sequencing.

[0041] Step 10. Perform high-throughput sequencing on the genome with redundant sequences removed. Use the BGIseq platform of BGI to perform sequencing, and use the NEBNext Ultra II DNA Library Prep Kit (product number E7645) for library construction to ensure that the library concentration reaches 2 nM and control the sequencing depth to reach 30X coverage. During the sequencing process, strictly follow the operating procedures of the platform, and perform preliminary quality control and data filtering by the BGI data analysis platform.

[0042] Step 11: Use bioinformatics tools to align and analyze the sequencing data, verify the deletion effect of redundant sequences and the degree of genome simplification, and at the same time check the integrity and function of the genome after removing redundant sequences. Use Bowtie 2 or BWA for alignment, set the alignment parameter that the maximum allowed number of mismatches does not exceed 2, sort and count the alignment results through SAMtools, then use IGV for visual inspection, and analyze the editing results in combination with functional annotation software such as ANNOVAR to finally determine whether the redundant sequences are completely deleted, the genome simplification effect, and potential off-target effects.

Claims

1. A plant genome repetitive sequence deletion technology, characterized in that: The following steps are involved: Step 1: Design specific sgRNA to accurately identify repetitive sequences in the plant genome; Step 2: The sgRNA and CAS9 / CAS12 protein are co-introduced into an in vitro double-end sequencing library system, and the CRISPR- CAS9 / CAS12 system is used to perform precise shearing at the target repetitive sequence position to destroy the library structure with repetitive sequence fragments; Step 3: Process the sheared DNA and use universal primers to amplify, purify and control the quality of the fragments. Fragments with repetitive sequences cannot be amplified because they are sheared. Step 4: Use magnetic beads to screen the amplified DNA library samples, and perform high-throughput sequencing on the screened target fragments.

2. The plant genome repetitive sequence deletion technology according to claim 1, characterized in that: The design of the sgRNA is based on dozens of existing wheat genomes and repetitive sequence annotations to form a universal wheat repetitive sgRNA sequence.

3. The plant genome repetitive sequence deletion technology according to claim 1, characterized in that: The sgRNA sequences are optimized and grouped according to their characteristics, and combined to prepare a high-density sgRNA library.

4. The plant genome repetitive sequence deletion technology according to claim 1, characterized in that: The plant genome repetitive sequence deletion technology system is used for precise shearing of target repetitive sequences in vitro.

5. The plant genome repetitive sequence deletion technology according to claim 1, characterized in that: The sheared DNA library samples are amplified and purified, and then sequenced using a high-throughput sequencing platform.

6. The plant genome repetitive sequence deletion technology according to claim 1, characterized in that: The technology can remove 60%-80% of the repetitive sequences in the wheat genome.

7. The plant genome repetitive sequence deletion technology according to claim 1, characterized in that: The technology can be widely used in accurate identification of wheat genome, genetic diversity research and molecular breeding.