Method for assembling myxobacteria T2T genome
By combining multiple sequencing technologies and software tools, along with telomere sequence screening and local sequence clustering screening, a high-quality slime mold T2T genome was successfully assembled, solving the problem of incomplete slime mold genome assembly and realizing a chromosome-level reference genome, providing an important research foundation for slime mold biological genetics and evolution research.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JILIN AGRICULTURAL UNIV
- Filing Date
- 2026-01-16
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies struggle to assemble complete and high-quality slime mold genomes, especially chromosome-level reference genomes, due to limitations imposed by high-density short tandem repeat sequences and contamination by symbiotic bacteria, resulting in poor continuity and integrity of the assembly results.
Second-generation Illumina short-read sequencing, Oxford Nanopore long-read sequencing, and PacBio HiFi long-read sequencing were used, combined with HiFiasm, NextDenovo, and NextPolish software for assembly and error correction. Contamination was removed by telomere sequence screening and local sequence clustering screening, and non-slime mold sequences were removed using blastn and Tiara software. Finally, high-quality slime mold genomes were obtained through NextPolish2 error correction.
A high-quality slime mold T2T genome was successfully assembled, realizing a chromosome-level reference genome from telomere to telomere. This improved the integrity and continuity of the genome, reduced experimental costs, and overcame the challenges of slime mold genome assembly, laying the foundation for slime mold biological genetics and evolutionary research.
Smart Images

Figure CN121938464A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of genomics and microbial science, and in particular to a method for assembling the T2T genome of slime molds. Background Technology
[0002] Myxomycetes are a unique and early-branching group of eukaryotes, classified in the Protist kingdom. They are among the most diverse amoeba groups in soil and an important component of soil native biodiversity. Myxomycetes are frequently used in cell biology and developmental biology research, involving multiple areas such as cell cycle regulation, cell differentiation, cell fusion, DNA replication and gene expression during development, and histone modification. Besides their value in basic research, myxomycetes also have significant applications in agriculture, food, and cosmetics. Given the potential importance of myxomycetes in ecosystems, basic research, and applications, in-depth research on this unique group of eukaryotes will provide significant scientific value and practical implications for understanding eukaryotic evolution, exploring new biological resources, and promoting the development of related industries.
[0003] Currently, slime mold research mainly focuses on taxonomy, resource diversity, and ecological functions. However, in genomics, a complete chromosome-level reference genome has remained unassembled for nearly 20 years. Assembling a slime mold genome is extremely challenging due to its fragmented nature and the unexpected complexity of the symbiotic bacteria it carries. This remains a critical bottleneck, and the lack of a reference genome significantly hinders in-depth exploration of key aspects of slime mold cell biology, origin, genetic evolution, and ecology. A complete and high-quality genome provides the foundation for revealing the genetic basis and other biological characteristics of slime molds.
[0004] Currently, commercial sequencing companies primarily employ a three-step assembly process. First, preliminary assembly is performed based on obtained third-generation sequencing data. Then, second-generation data is used to correct errors based on overlapping base sequences. Finally, Hi-C technology or optical mapping techniques are used to optimize the draft, resulting in a chromosome-level genome. This process has significant limitations: it relies on a reference genome, highly continuous base overlap, and advanced sequencing technology, making it expensive. Slime mold genomes contain high-density short tandem repeats and introns, making them extremely complex genomes. Base sequences are easily broken, leading to poor continuity and integrity in the assembly results. Therefore, existing assembly methods are difficult to use. Furthermore, slime molds often harbor numerous symbiotic bacteria, easily causing genome contamination. These factors all increase the difficulty of genome assembly, and no complete, gap-free reference genome assembly for any slime mold species has yet been published. Therefore, existing assembly methods are unsuitable for slime molds, and there is an urgent need to develop an optimized process for slime mold genome assembly. Summary of the Invention
[0005] The purpose of this invention is to provide a method for assembling the T2T genome of slime molds, which solves the problems of incomplete and low continuity in the assembly of slime mold genomes in the prior art. At the same time, it also solves the problem that no chromosome-level reference genome has been successfully assembled in slime molds. The first successfully assembled slime mold T2T genome lays an important research foundation for the evolution of eukaryotes and the exploration of new biological resources.
[0006] To achieve the above objectives, the present invention provides the following solution:
[0007] This invention provides a method for assembling the T2T genome of slime molds, comprising the following steps:
[0008] Slime molds were subjected to second-generation Illumina short-read sequencing, nanopore Oxford Nanopore long-read sequencing, and PacBio HiFi long-read sequencing to obtain second-generation Illumina sequencing data, nanopore Oxford Nanopore sequencing data, and PacBio HiFi sequencing data.
[0009] The PacBio HiFi sequencing data was assembled and contigs were extracted using HiFiasm software to obtain PacBio HiFi de novo assembled contigs; the Oxford Nanopore sequencing data was assembled using NextDenovo software to obtain ONT assembly results; and then the ONT assembly results were corrected using NextPolish software to obtain the corrected ONT assembly results.
[0010] The assembled contigs of PacBio HiFi de novo and the assembled results of ONT after error correction were compared and the aligned sequences were extracted to obtain a preliminary assembled slime mold genome draft.
[0011] The slime mold genome draft was subjected to telomere sequence screening and local sequence clustering screening to obtain a slime mold genome with contamination removed;
[0012] The decontamination-removed slime mold genome is corrected to obtain the slime mold T2T genome.
[0013] Preferably, the telomere sequence screening includes the step of extracting telomere repeat sequences from the slime mold genome draft and using the telomere repeat sequences as search targets.
[0014] Preferably, the local sequence clustering screening includes the steps of using blastn software to separate the mitochondrial genome sequence in the slime mold genome draft, comparing and removing the sequence; and then using Tiara software to monitor and remove the sequence.
[0015] Preferably, the telomere repeat sequence is (TTAGGG). n .
[0016] Preferably, the blastn software comparison includes the steps of comparing the mitochondrial genome sequence with the mitochondrial genome of *Hylocereus polycephalomycin* using blastn software and removing the sequence.
[0017] Preferably, the parameter for comparison is "-evalue 1e-10 -outfmt '6 qseqid sseqid pidentlength evalue bitscore'".
[0018] Preferably, the criterion for removing sequences by the blastn software is sequences with a confidence level greater than 50%.
[0019] Preferably, the criterion for removing sequences by the Tiara software is sequences that are shown as non-slime molds.
[0020] More preferably, the sequence of the non-slipid bacterium is a bacterial or plastid sequence, or a sequence with a length and confidence level both greater than 90.
[0021] Preferably, the slime mold includes *Sclerotium calcareum*.
[0022] This invention provides the application of the slime mold T2T genome obtained using the above-described slime mold T2T genome assembly method in the study of slime mold biological genetic and evolutionary mechanisms.
[0023] The present invention discloses the following technical effects:
[0024] Existing three-step assembly processes involve first using assembly software to roughly construct the genome sequence, then performing multiple rounds of correction and error correction using next-generation sequencing data to obtain a draft genome sequence, and finally optimizing the draft using Hi-C technology or optical mapping technology to obtain a chromosome-level genome. However, this approach is not suitable for slime molds. This invention aims to provide a more cost-effective and high-quality assembly method applicable to slime mold genomes. This method mainly includes four steps: global assembly, telomere sequence screening, local sequence clustering screening, and multiple rounds of correction using next-generation sequencing data. Compared to existing processes, the method provided in this invention adds a contamination removal strategy of telomere sequence screening and local sequence clustering screening. Using the written telo.tiqu.pl script, telomere sequence screening and local sequence clustering analysis are performed in a global sequence context. The conservation of telomere sequences is utilized to screen out the main genome sequences containing slime mold telomere repetitive sequences, effectively overcoming the problem of complex and short tandem repetitive sequences in slime mold genomes. This makes the analysis of complex regions possible, eliminating the need for fingerprinting, Hi-C technology, or optical sequencing, thus reducing experimental costs. Furthermore, it ensures that all sequences belonging to slime molds are extracted, effectively removing interference from repetitive sequences and deep prokaryotic contamination. Highly similar repetitive sequence fragments are restored and located to their accurate genomic positions, significantly improving genome integrity and continuity. This invention successfully solves the long-standing problem of a lack of high-quality reference genomes for slime molds, completing the first chromosome-level reference genome of slime molds from telomere to telomere, laying an important research foundation for the analysis of protist genetic and evolutionary mechanisms.
[0025] Furthermore, a high-quality genome of slime mold was obtained based on the assembly method provided in this application, solving the problem that slime mold genomes have been difficult to assemble successfully to date. A genome assembly process suitable for slime molds was proposed, and a high-quality reference genome was assembled.
[0026] In summary, this invention is the first to achieve chromosome-level reference-level assembly of the telomere-to-telomere T2T level genome of *Slime scabra*, becoming the first high-quality T2T reference genome of a slime mold. This lays an important research foundation for the analysis of the genetic and evolutionary mechanisms of protists and has significant scientific value. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 A schematic diagram of the improved slime mold genome assembly process;
[0029] Figure 2 The image shows protoplasmic masses of six slime molds; where a is *Squamae*, b is *Squamae*, c is *Squamae*, d is *Squamae*, e is *Squamae*, and f is *Squamae*; the scale bar is 1 cm.
[0030] Figure 3 The flow cytometry histogram (A) and distribution map (B) of the genome of *Calvatia scabra*.
[0031] Figure 4 This is a comparison of GC content and sequencing depth before and after decontamination of the *Calvatia scabra* genome; where A represents before decontamination and B represents after decontamination. Detailed Implementation
[0032] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as a limitation of the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.
[0033] It should be understood that the terminology used in this invention is merely for describing particular embodiments and is not intended to limit the invention. Furthermore, with respect to numerical ranges in this invention, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Any stated value or intermediate value within a stated range, as well as each smaller range between any other stated value or intermediate value within said range, is also included in this invention. The upper and lower limits of these smaller ranges may be independently included or excluded from the range.
[0034] Unless otherwise stated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. While only preferred methods and materials have been described herein, any methods and materials similar or equivalent to those described herein may be used in the implementation or testing of this invention. All references to this specification are incorporated by way of citation to disclose and describe methods and / or materials associated with those references. In the event of any conflict with any incorporated reference, the content of this specification shall prevail.
[0035] Various modifications and variations can be made to the specific embodiments described in this specification without departing from the scope or spirit of the invention, as will be apparent to those skilled in the art. Other embodiments derived from this specification will also be obvious to those skilled in the art. This specification and embodiments are merely exemplary.
[0036] The terms “include,” “including,” “have,” “contain,” etc., used in this article are all open-ended terms, meaning that they include but are not limited to.
[0037] Example 1
[0038] 1) The experimental material samples were collected from Xinglong Mountain, Gansu Province, specimen number 20150701001. Protoplasmic masses were obtained in the laboratory by spore germination. This strain has been published in the literature "Species Composition and Distribution Pattern of Slime Fungi in Qinba Mountains of Southern Shaanxi" (Dou Wenjun, Li Shu, Peng Xueyan, et al. Species Composition and Distribution Pattern of Slime Fungi in Qinba Mountains of Southern Shaanxi [J]. Acta Mycologica Sinica, 2023, 42(01):196-216. DOI:10.13346 / j.mycosystema.220405.) and is deposited in the Herbarium of the Engineering Research Center for Edible and Medicinal Fungi of the Ministry of Education, Jilin Agricultural University.
[0039] 2) Experimental Instruments: ICOC-ZSA302 stereomicroscope, DM1000 optical microscope, clean bench, intelligent artificial climate chamber, benchtop pure water system, vertical pressure steam sterilizer, CytoFLEX S flow cytometer, tissue homogenizer, water bath, Centrifuge 5425 R benchtop high-speed refrigerated centrifuge, basic vortex mixer, Nanodrop ONEc, ribonuclease (RNAse 10 mg / mL), plant genomic DNA extraction kit, TransZol Up, solid-phase surface RNA scavenger, mercaptoethanol, isopropanol, anhydrous ethanol, RNase-free water, citric acid, anhydrous disodium hydrogen phosphate, potassium hydroxide, anhydrous ethanol, propidium iodide (PI 10 mg / mL), calcium- and magnesium-free PBS, sodium hypochlorite, Tween 20, cedar oil, 2 mL EP tubes, forceps, petri dishes, glass slides, hemocytometer, thermometer, oat flour, water agar medium, 10μL pipettes, 1000μL pipettes, incubator, spectrophotometer.
[0040] 3) Experimental Procedure
[0041] Step 1: Sample screening, collection, and sequencing
[0042] 1. The genome size of *Agaricus bisporus* was assessed by flow cytometry. Spores were stained with PI fluorescence to obtain a stained spore suspension. The stained spore suspension was shaken and mixed before being analyzed by flow cytometry. Fluorescence from 610 / 20 channels was collected using a 561 nm laser, and the PI emission fluorescence intensity was measured. At least 10,000 fluorescence signals from the target cell population were collected and recorded, with the coefficient of variation (CV) controlled below 5%. Each sample was performed in triplicate. The data were analyzed using the instrument's built-in CytExpert software, with *Agaricus bisporus* (JELange) Imbach as a reference, to estimate the spore cell nuclear genome content.
[0043] 2. Under a dissecting microscope, pick the fruiting bodies of *Calvatia scabra* into a 2 mL EP tube, add 1000-1700 μL of RO water, and gently grind with forceps to release spores, forming a spore suspension with a concentration of 2 × 10⁻⁶. 5 -5×10 6 Cells / mL. Incubate in the dark until motile cells or myomorphosomes form, then place them on water agar medium containing oat flour and incubate in the dark until protoplasts form. Continuously purify the protoplasts to obtain purified protoplasts.
[0044] The purified protoplasts were cultured under the following conditions: 23.5℃ with oat feed, 23.5℃ without feed, 23.5℃ without feed + protoplast fragmentation, 23.5℃ without feed + protoplast aggregation, 23.5℃ without feed + mid-stage fruiting body, and 23.5℃ without feed + mature fruiting body.
[0045] 23.5℃ Oat-fed group: The protoplasm of *Calvatia scabra* was cultured at 23.5℃ in the dark for 3 days. The protoplasm tissue was collected to obtain *Calvatia scabra* protoplasm samples fed with nutrients. The culture medium used was water agar medium and the nutrient was oat grains.
[0046] 23.5℃ No-feed group: The protoplasts of *Calvatia scabra* were cultured at 23.5℃ in the dark for 3 days. The protoplast tissues were collected to obtain *Calvatia scabra* protoplast samples without feeding nutrients. The culture medium used was water agar medium.
[0047] 23.5℃ feedless + protoplast fragmentation group: The protoplasts of *Calvatia scabra* were cultured at 23.5℃ in the dark for 5-6 days. Protoplast tissues during the fragmentation period were collected to obtain *Calvatia scabra* protoplast fragmentation period samples. The culture medium used was water agar medium.
[0048] 23.5℃ feedless + protoplast aggregation group: The protoplasts of *Calvatia scabra* were cultured at 23.5℃ in the dark for 6-7 days. Protoplast tissues during the aggregation period were collected to obtain *Calvatia scabra* protoplast aggregation period samples. The culture medium used was water agar medium.
[0049] 23.5℃ No Feed + Mid-stage Fruiting Body Group: The protoplasm of *Calvatia scabra* was cultured at 23.5℃ in the dark for 8-9 days. Mid-stage fruiting body tissues were collected to obtain mid-stage fruiting body samples of *Calvatia scabra*. The culture medium used was water agar medium.
[0050] 23.5℃ No feed + mature fruiting body group: The protoplasm of *Calvatia scabra* was cultured at 23.5℃ in the dark for 11-13 days. Mature fruiting body tissues were collected to obtain mature *Calvatia scabra* fruiting body samples. The culture medium used was water agar medium.
[0051] 3. DNA was extracted from the *Calvatia scabra* tissue samples collected above at five different growth stages: early protoplast stage, protoplast fragmentation stage, protoplast aggregation stage, mid-fruiting body formation stage, and fruiting body maturity stage. Total RNA was extracted according to the instructions of the Plant Genomic DNA Extraction Kit (TianGen), and total RNA was extracted according to the instructions of the TransZol Up Kit (Transgen). The samples were stored at -80℃. All instruments and consumables were treated with a solid-phase surface RNase remover by soaking and wiping to ensure a clean operating environment. The concentrations of the obtained DNA and RNA samples were estimated using Nanodrop, and the integrity of the samples was verified using agarose gel electrophoresis.
[0052] 4. Construct high-quality genome libraries. Prepare sequencing libraries according to the standardized procedures provided by each sequencing platform. Perform second-generation Illumina short-read sequencing (sequencing on the Illumina NovaSeq sequencing platform, sequencing depth at least ~100×, sequencing read length 150bp) on *Calvatia scabra* DNA, single-molecule sequencing (sequencing on an Oxford Nanopore Technology sequencer, sequencing depth at least ~200×, sequencing read length 19kb) and PacBio HiFi long-read sequencing (SMRT sequencing on PacBio Revio and Sequel II sequencing platforms, sequencing depth ~200×, sequencing read length 18kb) on *Calvatia scabra* tissue RNA at different growth stages for library construction and sequencing. Obtain second-generation Illumina sequencing data, nanopore Oxford Nanopore sequencing data, PacBio HiFi sequencing data, and second-generation RNA-seq sequencing data (for genome annotation, i.e., for step 5).
[0053] Step 2: Improved slime mold genome assembly process and genome map construction
[0054] 1. Slime molds are extremely small in size. Due to their predatory habits and the presence of symbiotic bacteria, their genomes are often contaminated by bacteria and fungi, resulting in a large number of contaminated sequences in the sequencing data, which increases the difficulty of genome assembly. To address this issue, this embodiment improves the genome assembly and decontamination process for slime mold species after multiple experiments. This optimized process mainly includes four steps: global assembly, telomere sequence screening, local sequence clustering screening, and multiple rounds of second-generation data correction. The specific process is as follows: Figure 1 As shown.
[0055] 2. Preliminary assembly of the *Calvatia scabra* genome: A hybrid assembly was performed using PacBio HiFi sequencing data, Oxford Nanopore sequencing data, and Illumina sequencing data. HiFiasm (v0.19.9-r616) was used to assemble the PacBio HiFi sequencing data, yielding the PacBio HiFi assembly results and extracting contigs. NextDenovo (v2.5.0) was used to assemble the Oxford Nanopore sequencing data, yielding the ONT assembly results. NextPolish (v1.4.1) was used to correct errors in the ONT assembly results based on second-generation Illumina sequencing data, yielding the corrected ONT assembly results. Then, blastn (v2.16.0) software was used to align the two assembly results (PacBio HiFi assembly results and the corrected ONT assembly results) and extract the aligned sequences, resulting in a preliminary draft of the *Calvatia scabra* genome, denoted as Contig V1.
[0056] 3. Decontamination of the *Calvatia scabra* genome Contig V1: Telomere repeat sequences were extracted from the initially assembled *Calvatia scabra* genome Contig V1 using telo.tiqu.pl, confirming the telomere repeat sequence as (TTAGGG). n Using this telomere repeat sequence as the search target, it was annotated in all sequences of *Calvatia scabra* (telomere sequence screening). The mitochondrial genome sequence in the initially assembled Contig V1 was isolated using blastn (v2.16.0) software. Alignment was performed with the mitochondrial genome of *Phyllostachys multicephalomys* (genome version AB027295.1) using the parameters "-evalue 1e-10 -outfmt '6 qseqid sseqid pident length evalue bitscore'". Sequences with a confidence score greater than 50 were removed. The *Calvatia scabra* genome after removing the mitochondrial sequence was designated Contig V2. Sequence classification was monitored using Tiara (v1.0.3), and sequences indicating bacteria or plastids were removed. The alignment results were manually reviewed, and sequences with both alignment length and confidence score greater than 90 were removed. The *Calvatia scabra* genome after removing non-myxophyte sequences was designated Contig V3 (local sequence clustering screening). Integrating the above contamination-free sequences from Contig V2 and Contig V3, the resulting contamination-free *Calvatia scabra* genome is denoted as Contig V4.
[0057] The script code (telo.tiqu.pl) is as follows:
[0058] #! / usr / bin / perl
[0059] use strict;
[0060] use warnings;
[0061] use Bio::SeqIO;
[0062] # Check the number of parameters
[0063] if ($#ARGV < 3) {
[0064] die "Usage: perl $0 sequence tel_length repeat_time species_name\n";
[0065] }
[0066] my $tel_len = $ARGV[1];
[0067] my $rep_unit = $ARGV[2];
[0068] my $species_name = $ARGV[3];
[0069] my $fh;
[0070] if ($ARGV[0] =~ / \.gz$ / ) {
[0071] open $fh, "gunzip -c $ARGV[0] |" or die "Cannot open file: $!";
[0072] } else {
[0073] open $fh, '<', $ARGV[0] or die "Cannot open file: $!";
[0074] }
[0075] my $seqio = Bio::SeqIO->new(-fh => $fh, -format => 'Fasta');
[0076] while (my $seqBio = $seqio->next_seq) {
[0077] my $string = $seqBio->seq;
[0078] my $idhead = $seqBio->id;
[0079] my $tot_len = length($string);
[0080] my $mid_len = $tot_len - 2 * $tel_len;
[0081] if ($mid_len > 0) {
[0082] my $left_seq = substr($string, 0, $tel_len);
[0083] my $right_seq = substr($string, $tot_len - $tel_len, $tel_len);
[0084] my $mid_seq = substr($string, $tel_len, $mid_len);
[0085] my $left = 0;
[0086] my $right = 0;
[0087] my ($left_pos, $mid_pos, $right_pos) = (0, 0, 0);
[0088] my ($left_rep1, $mid_rep1, $right_rep1) = ('NA', 'NA', 'NA');
[0089] my $markleft = "NAcla";
[0090] my $markright = "NAcla";
[0091] my $markmid = "NAcla";
[0092] # Check the left sequence
[0093] if ($left_seq =~ / ((CCCTAA){$rep_unit,}) / ) {
[0094] $left_rep1 = $1;
[0095] $left_pos = pos($left_seq) - length($1) + 1;
[0096] $left = 1;
[0097] $markleft = "SLIMEMOLDS";
[0098] }
[0099] # Check right-handed sequence
[0100] if ($right_seq =~ / ((TTAGGG){$rep_unit,}) / ) {
[0101] $right_rep1 = $1;
[0102] $right_pos = pos($right_seq) - length($1) + 1 + $tel_len + length($mid_seq);
[0103] $right = 1;
[0104] $markright = "SLIMEMOLDS";
[0105] }
[0106] my $tolrepeats1 = 0;
[0107] my $tolrepeats2 = 0;
[0108] my $tolrepeats3 = 0;
[0109] my $tolrepeats4 = 0;
[0110] # Check for CCCTAA repeats in intermediate sequences
[0111] while ($mid_seq =~ / (CCCTAA){$rep_unit,} / g) {
[0112] $tolrepeats1 += length($&) / 6;
[0113] $markmid = "SLIMEMOLDS";
[0114] }
[0115] # Check for TTAGGG repeats in intermediate sequences
[0116] while ($mid_seq =~ / (TTAGGG){$rep_unit,} / g) {
[0117] $tolrepeats1 += length($&) / 6;
[0118] $markmid = "SLIMEMOLDS";
[0119] }
[0120] print "Sequence: $idhead\n";
[0121] print "Total length: $tot_len\n";
[0122] print "Left telomere: $markleft, Position: $left_pos, Repeat: $left_rep1\n";
[0123] print "Right telomere: $markright, Position: $right_pos, Repeat: $right_rep1\n";
[0124] print "Middle repeats: $tolrepeats1\n";
[0125] print "---\n";
[0126] }
[0127] }
[0128] close $fh;
[0129] #perl telo.tiqu.pl sequence.fasta 1000 3 species_name
[0130] Script text parsing:
[0131] I. Script Description
[0132] This script is used to analyze the distribution of telomere-related repeat sequences in DNA sequences, mainly detecting two six-base repeat patterns: CCCTAA and TTAGGG.
[0133] II. Input Parameter Description
[0134] Four parameters are required:
[0135] perl telo.tiqu.pl <sequence file> <telomere detection length> <minimum number of repetitions> <output file>.
[0136] III. Textual Description
[0137] 1. Input:
[0138] A sequence file in FASTA format
[0139] tel_len: An integer representing the length of the subsequence to be checked from both ends of the sequence when checking for telomere duplication.
[0140] rep_unit: An integer representing the minimum number of repetitions. For example, if rep_unit=3, then the repeating pattern must appear at least 3 times consecutively before it is counted.
[0141] 2. For each sequence:
[0142] a. Read the sequence ID and sequence string.
[0143] b. Calculate the total length of the sequence tot_len.
[0144] c. Calculate the length of the intermediate region: mid_len = tot_len - 2 * tel_len. If mid_len <= 0, skip the sequence (do not process).
[0145] d. Extract the left-hand sequence: take tel_len bases from the beginning.
[0146] e. Extract the right-hand sequence: Take tel_len bases from the end.
[0147] f. Extract intermediate sequence: Starting from position tel_len, extract mid_len bases.
[0148] 3. Telomere repeatability detection:
[0149] a. Left telomeres: Search the left sequence for a “CCCTAA” pattern that is repeated consecutively at least `rep_unit` times. If found, record:
[0150] - Repeating sequence string (left_rep1)
[0151] - Repeat starting position (left_pos, note that the position here is counted from 1, relative to the beginning of the entire sequence)
[0152] - Markleft is set to "SLIMEMOLDS" otherwise it is set to "NAcla".
[0153] b. Right-hand telomeres: Search the right-hand sequence for a “TTAGGG” pattern that is repeated consecutively at least `rep_unit` times. If found, record:
[0154] - Repeating sequence string (right_rep1)
[0155] - Repeat the start position (right_pos, note that this position is relative to the end of the entire sequence, calculated by adding the lengths of the left and middle sequences to the position in the right subsequence, thus converting it to the position in the entire sequence).
[0156] - Mark the right as "SLIMEMOLDS" or "NAcla".
[0157] 4. Repeated counting in the middle region:
[0158] a. Initialize the total number of intermediate repetitions tolrepeats1 to 0.
[0159] b. In the intermediate sequence, search for all consecutive repeating patterns of “CCCTAA” at least rep_unit times. For each one found, divide the length of the repeating sequence by 6 (because a repeating unit is 6 bases) and then add it to tolrepeats1.
[0160] c. Similarly, search for all consecutive "TTAGGG" patterns that repeat at least rep_unit times and add them to tolrepeats1.
[0161] d. If at least one such repetition exists in the middle region, markmid is set to "SLIMEMOLDS"; otherwise, it is set to "NAcla". (Note: In the program, markmid is assigned a value in the while loop, but its initial value is "NAcla". It is only overwritten as "SLIMEMOLDS" when at least one repetition is found. However, in the program, if the while loop is not entered even once, markmid remains "NAcla".)
[0162] 5. Output:
[0163] For each sequence, output:
[0164] Sequence: [Sequence ID]
[0165] Total length: [Total length]
[0166] Left telomere: [marker], Position: [left telomere position], Repeat: [left repeat sequence]
[0167] Right telomere: [marker], Position: [right telomere position], Repeat: [right-side repeat sequence]
[0168] Middle repeats: [Total number of repeats in the middle area]
[0169] 4. To improve the genome draft Contig V4 to T2T assembly quality and reduce the presence of splicing error sites, firstly, the genome from step 3 was back-aligned to PacBio HiFi sequencing data using minimap2 (v2.28). Secondly, based on the k-mer dataset generated from the HiFi back-alignment data (obtained by combining minimap2 with PacBio HiFi sequencing data) and Illumina sequencing data, NextPolish2 (v0.2.1) was used to correct errors in Contig V4, generating the final T2T-level chromosome genome, denoted as Contig V5.
[0170] 5. Second-generation RNA-seq transcriptome assembly and gene prediction and functional annotation: The *Calvatia scabra* transcriptome was assembled de novo using the *rnaSPAdes* tool in SPAdes (v4.0.0). The transcriptome was then compared with the UNIVEC contamination database and contaminated data was filtered using the PASAseqclean tool. Finally, the decontaminated transcriptome was compared back to the *Calvatia scabra* genome using the blat and gmap tools in PASA (v2.0.2) to obtain the open reading frame (ORF) sequences and coding gene annotation files for training the gene model in Augustus (v3.5.0). The transcriptome data input to GeneMarkS-T (v4.33) was obtained by combining Hisat2 (v2.2.1) and Stringtie (v2.1.4). The results of homology annotation, de novo annotation, and transcriptome-based annotation were integrated according to certain weights using EvidenceModeler (v2.1.0).
[0171] 6. *Calvatia scabra* Genome Assessment: The *Calvatia scabra* genome contig N50 was calculated using Quast (v5.3.0) to assess assembly continuity. A higher contig N50 value indicates greater genome assembly continuity and better assembly quality. The *Calvatia scabra* genome consistency QV value and K-mer integrity were calculated using Merqury (v1.4.1) based on K-mer. The *Calvatia scabra* genome integrity was assessed using BUSCO (v5.7.1) based on the eukaryota_odb10 eukaryotic database, combined with manually calculated GQ values, to jointly assess the completeness of annotated functional genes.
[0172] 7. Repetitive sequences in the *Calvatia scabra* genome were annotated using EDTA (v2.2.1), Tandem Repeat Finder (v4.09.1), RepeatMasker (v1.88), and the latest Repbase database.
[0173] 8. Non-coding RNAs of the *Calvatia scabra* genome were annotated using tRNAScan-SE (v2.0.12), RNAmmer (v1.2), Infernal (v1.1.5), and the Pfam domain database.
[0174] 9. Potential centromere sequence localization analysis was performed using quarTeT (v1.2.5).
[0175] 4) Experimental Results
[0176] (1) Flow cytometry estimation results of genome
[0177] The original specimen (sample number LGP) of Didymium squamulosum (Alb. & Schwein.) Fr. & Palmquist, collected from Xinglong Mountain, Gansu Province, specimen number 20150701001, was obtained from the protoplasm by spore germination in the laboratory. This strain has been published in the literature "Species Composition and Distribution Pattern of Slime Fungi in the Qinling-Bashan Mountains of Southern Shaanxi" (Dou Wenjun, Li Shu, Peng Xueyan, et al. Species Composition and Distribution Pattern of Slime Fungi in the Qinling-Bashan Mountains of Southern Shaanxi [J]. Acta Mycosystema Sinica, 2023, 42(01):196-216. DOI:10.13346 / j.mycosystema.220405.). The genome size of Didymium squamulosum was estimated to be approximately 89.64 Mb by flow cytometry, as shown in the results. Figure 3As shown in A. Subsequently, following a modified traditional wet-chamber culture method (water is added to 2.0 mL EP tubes, and myxomycete spores are inoculated for fed culture; myxomycete culture is added, first culturing to obtain myxomycete zoocytes and myxomorphs; zoocytes or myxomorphs combine and undergo meiosis to form zygotes, subsequently forming protoplasmic masses; primordia are weak, and later the primordia gradually mature, the veins and their leading edges thicken, yielding the myxomycete protoplasm), protoplasmic masses of the screened species were cultured under laboratory conditions, successfully obtaining protoplasmic masses of six myxomycetes: *Calvatia scabra*, *Calvatia xanthoides*, *Calvatia scabra*, *Calvatia scabra*, *Calvatia scabra*, *Calvatia scabra*, and *Calvatia scabra*. Figure 2 The specific process of improving the traditional wet chamber culture method (a culture method for artificially culturing myxomycete spores to form protoplasts) is disclosed in Chinese patent "CN115305205B A culture method for artificially culturing myxomycete spores to form protoplasts".
[0178] (2) Statistical results of different versions of the genome of *Calvatia scabra* obtained by the general assembly process and the method provided in this invention
[0179] In this embodiment, during the experiment, a sequencing company was commissioned to sequence and assemble the genome of *Calvatia scabra* using existing assembly procedures. The results showed that the *Calvatia scabra* genome size was 98.58 Mb, containing 535 contigs, with a Contig N50 of 0.57 Mb, and the longest sequence was 6.09 Mb (Table 1). Compared to the company's assembled version, the ContigV5 version in this embodiment showed superior overall quality, containing 47 sequences, reducing the number of fragmented sequences, with a genome size of 83.55 Mb, a Contig N50 increased to 1.75 Mb, and a significantly improved N90 value to 1.44 Mb. The L50 and L90 of Contig V5 were 20 and 41, respectively, both superior to the company's version and Contig V1. Furthermore, as... Figure 4 As shown, the genome data after the assembly process ( Figure 4 B) compared to Contig V1 ( Figure 4 A) is more concentrated and cleaner, further demonstrating its higher genome assembly integrity and continuity. Therefore, the assembly process proposed in this embodiment has strong decontamination ability, can obtain a purer version, and can further effectively improve the quality and reliability of slime mold genomes.
[0180] Table 1. Statistics of different versions of *Calvatia scabra* genomes obtained using the general assembly process and the method provided in this embodiment.
[0181]
[0182] Note: The higher the N50 value, the higher the genome continuity and the higher the quality of genome assembly.
[0183] (3) Statistical results of telomere and centromere characteristics of the *Calmineraria scabra* genome
[0184] The method provided in this invention successfully obtained the first high-quality slime mold reference genome. The genome assembly size of *Scutellaria calcei* was 83.55 Mb, and 47 chromosomes were successfully assembled. Each chromosome was annotated with telomere sequences, and the size of the centromere region ranged from 51.17 Kb to 273.87 Kb. All chromosomes were assembled in a gap-free, continuous manner (Tables 2 and 3). The contig N50 of this genome was 1.73 Mb (Table 3), indicating high continuity, and the average GC content was 34.78%. The final assembled size of the *Scutellaria calcei* genome was highly consistent with the genome size previously estimated using flow cytometry and Illumina sequencing data. Figure 3 The accuracy and reliability of the assembly results were further verified by A and B in the figure. The integrity of the annotated functional genes was assessed by BUSCO, with a result of 91% and a GQ of 95.5%, further verifying the high quality of the *Calvatia scabra* genome assembly obtained using the method provided in this invention.
[0185] Table 2. Telomere and centromere characteristics of the *Calvatia scabra* genome.
[0186]
[0187] Table 3. Telomere and centromere characteristics of the *Calvatia scabra* genome.
[0188]
[0189] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A method for assembling the T2T genome of slime molds, characterized in that, Includes the following steps: Slime mold was subjected to second-generation Illumina short-read sequencing, nanopore Oxford Nanopore long-read sequencing, and PacBioHiFi long-read sequencing to obtain second-generation Illumina sequencing data, nanopore Oxford Nanopore sequencing data, and PacBioHiFi sequencing data. The PacBio HiFi sequencing data was assembled and contigs were extracted using HiFiasm software to obtain PacBioHiFi de novo assembled contigs; the Oxford Nanopore sequencing data was assembled using NextDenovo software to obtain ONT assembly results; and then the ONT assembly results were corrected using NextPolish software to obtain the corrected ONT assembly results. The assembled contigs of PacBio HiFi de novo and the assembled results of ONT after error correction were compared and the aligned sequences were extracted to obtain a preliminary assembled slime mold genome draft. The slime mold genome draft was subjected to telomere sequence screening and local sequence clustering screening to obtain a slime mold genome with contamination removed; The decontamination-removed slime mold genome is corrected to obtain the slime mold T2T genome.
2. The method for assembling the T2T genome of slime molds according to claim 1, characterized in that, The telomere sequence screening includes the step of extracting telomere repeat sequences from the slime mold genome draft and using these telomere repeat sequences as search targets.
3. The method for assembling the T2T genome of slime molds according to claim 1, characterized in that, The local sequence clustering screening includes the steps of using blastn software to separate the mitochondrial genome sequence from the slime mold genome draft, comparing and removing the sequence; and then using Tiara software to monitor and remove the sequence.
4. The method for assembling the T2T genome of slime molds according to claim 2, characterized in that, The telomere repeat sequence is (TTAGGG). n .
5. The method for assembling the T2T genome of slime molds according to claim 3, characterized in that, The blastn software comparison includes the steps of comparing the mitochondrial genome sequence with the mitochondrial genome of *Hylocereus polycephalomycin* using blastn software and removing the sequence.
6. The method for assembling the T2T genome of slime molds according to claim 5, characterized in that, The parameters for comparison are "-evalue 1e-10 -outfmt '6 qseqid sseqid pident length evalue bitscore'".
7. The method for assembling the T2T genome of slime molds according to claim 3 or 5, characterized in that, The criterion for removing sequences using the blastn software is sequences with a confidence level greater than 50%.
8. The method for assembling the T2T genome of slime molds according to claim 3, characterized in that, The criterion for removing sequences by the Tiara software is to identify sequences that are shown as non-slime mold.
9. The method for assembling the T2T genome of slime molds according to claim 1, characterized in that, The slime molds include scaly calcareous dermatitis.
10. The application of the slime mold T2T genome obtained by the slime mold T2T genome assembly method according to any one of claims 1-9 in the study of slime mold biological genetic and evolutionary mechanisms.
Citation Information
Patent Citations
A method for artificially culturing slime mold spores to form protoplasm
CN115305205B