A method for assembling ganoderma bicomb genome based on haplotype resolution

By employing an assembly strategy based on haplotype segregation, the Ganoderma lucidum binuclear genome was split into haplotypes for sequencing. By combining multiple sequencing technologies and data processing steps, the accuracy and continuity issues of Ganoderma lucidum genome assembly in existing technologies were resolved, and high-quality genome assembly was achieved.

CN121565256BActive Publication Date: 2026-05-19JILIN AGRICULTURAL UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JILIN AGRICULTURAL UNIV
Filing Date
2026-01-21
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies suffer from high error rates, poor continuity, and insufficient assembly integrity when processing the Ganoderma lucidum heterophyllum-coordinating binucleosome genome.

Method used

An assembly strategy based on haplotype segregation was adopted to separate heterozygous diploids into homozygous haploids for sequencing. Enzymatic digestion was performed using a compound enzyme digestion solution. The assembly results of Nanopore and PacBio HiFi were combined, and a mitochondrial de-genome reads step was introduced. Chromosome mounting and gap filling were performed using Hi-C data, and correction was performed using Illumina NovaSeq data.

Benefits of technology

It significantly improved the accuracy and integrity of genome assembly, reduced the false fusion rate of high repetitive regions, enhanced the accuracy and continuity of assembly, and generated a seamless, non-redundant chromosome backbone, providing a unique, non-redundant reference coordinate system for subsequent telomere-centromere filling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121565256B_ABST
    Figure CN121565256B_ABST
Patent Text Reader

Abstract

The application discloses a Ganoderma lucidum double-nucleus genome assembly method based on haplotype analysis and belongs to the technical field of bioinformatics. The application aims to overcome the precision and accuracy of high-hybrid genome assembly. The application provides a Ganoderma lucidum double-nucleus genome assembly method based on haplotype analysis, obtains single-nucleus Ganoderma lucidum cells, respectively performs whole genome sequencing on the cells by using Illumina NovaSeq, Nanopore, Hi-C and PacBio HiFi, obtains precise genome information from Illumina NovaSeq sequencing data and Nanopore sequencing data, and realizes a haplotype separation assembly strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bioinformatics technology, specifically relating to a method for assembling the Ganoderma lucidum binucleate genome based on haplotype analysis. Background Technology

[0002] Ganoderma lucidum mycelia are heterothallic binucleate genomes, with each cell containing two sets of unfused haploid nuclei, exhibiting highly heterozygous characteristics. Existing assembly techniques based on binucleate hybrid sequencing generally suffer from high error rates, poor continuity, and insufficient assembly integrity when processing such genomes. We innovatively propose an assembly strategy based on haplotype segregation, separating the heterozygous diploid into homozygous haploids for separate sequencing. Currently, most researchers use a single enzymatic digestion solution; we have uniquely prepared a composite enzymatic digestion solution using 20 mg / mL each of lysozyme, snailase, cellulase, and lysozyme, which allows for more thorough digestion and improves single-nucleus efficiency. Currently, most researchers directly assemble the Ganoderma lucidum genome; however, the Ganoderma lucidum genome is relatively small, and the mitochondrial genome can significantly interfere with its nuclear genome assembly. Therefore, we innovatively introduce a step of removing mitochondrial genome reads to improve the accuracy and integrity of genome assembly. Currently, most people do not use Nanopore assembly results as a framework for scaffolding PacBio HiFi assembly results. In fact, Nanopore assembly has high continuity and an error rate of about 2%-5%; while PacBio HiFi assembly accuracy is >99.9%, but the contigs are shorter. Summary of the Invention

[0003] The purpose of this invention is to overcome the limitations of precision and accuracy in assembling highly heterozygous genomes. To this end, an innovative assembly strategy based on haplotype segregation is proposed, which separates heterozygous diploids into homozygous haploids for sequencing.

[0004] This invention provides a method for assembling the Ganoderma lucidum binuclear genome based on haplotype analysis. The steps of the method are as follows:

[0005] Step 1: Processing Ganoderma lucidum to obtain two monokaryotic strains of Ganoderma lucidum. After extracting the DNA of each monokaryotic strain, whole genome sequencing was performed using Illumina NovaSeq, Nanopore, Hi-C and PacBio HiFi respectively. Quality control was performed on the Illumina NovaSeq sequencing data and the Nanopore sequencing data to obtain high-quality read length data.

[0006] Step 2: The mitochondrial data of Ganoderma species downloaded from the database are compared with the Hi-C, PacBio HiFi, Illumina NovaSeq and Nanopore data from Step 1, and the mitochondrial data are removed.

[0007] Step 3: Based on the Illumina NovaSeq sequencing data, determine the genome size of the PacBio HiFi data and Nanopore data obtained in Step 2, and perform de novo assembly to obtain the PacBio HiFi contig and Nanopore contig.

[0008] Step 4: Align the PacBio HiFi data obtained in Step 2 with the PacBio HiFi contiguous group obtained in Step 3 to obtain repeating sequence information. Then, perform self-alignment on the PacBio HiFi contiguous group obtained in Step 3 to obtain repeating sequence information. Then, based on the above repeating sequences as redundant information, use Purge_Dups (v1.2.5) to process the redundant information to obtain a PacBio HiFi contiguous group without redundancy.

[0009] Step 5: Align the Nanopore contigs obtained in Step 3 with the non-redundant PacBio HiFi contigs obtained in Step 4. The sequences with 95% homology are considered collinear segments. Replace the collinear segments of the Nanopore contigs with the collinear segments of the PacBio HiFi contigs to obtain the candidate genome.

[0010] Step 6: Prepare a chromosome interaction matrix from the Hi-C data obtained in Step 1. Use the 3D-DNA algorithm to process the candidate genomes and chromosome interaction matrix obtained in Step 5 to obtain chromosome clusters. Then, use the chromosome clusters to prepare an interaction heatmap. Adjust the candidate genomes according to the chromosome map to obtain the genome sequence.

[0011] Step 7: Use LR_gapclosed software to analyze the PacBio HiFi sequencing data and Nanopore sequencing data obtained in Step 1 to obtain gap data, and then fill the gap data into the gene sequence obtained in Step 6 to obtain new gene sequences;

[0012] Step 8: Using the Illumina NovaSeq data obtained in Step 1 as a benchmark, perform whole-genome correction on the new gene sequences obtained in Step 8 to obtain a high-quality genome, and then integrate the genomes of two mononuclear Ganoderma lucidum.

[0013] To further define the method for removing mitochondrial sequences in step 2, the e-value threshold is set to 1E-5, and then reads with a length greater than 500bp or a length-to-read length ratio greater than 0.5 are matched and removed from the PacBio Hi-C data, PacBio HiFi data obtained in step 1, and the Illumina NovaSeq data and Nanopore data obtained in step 2.

[0014] Further specifying, in step 3, the PacBio HiFi data is processed using the Canu software package (v2.2), the minimum coverage (corMinCoverage) in the calibration step is set to 4, the output coverage (corOutCoverage) is set to 40×, the minimum read length (minReadLength) and the minimum overlap length (minOverlapLength) for assembly are set to 1000bp and 500bp respectively, and whole genome assembly is performed to obtain the PacBio HiFi contig.

[0015] Further, in step 3, the Nanopore data is processed using the nextdenovo software package. In the calibration step, the minimum coverage (corMinCoverage) is set to 4, the output coverage (corOutCoverage) is set to 40×, and the minimum read length (minReadLength) and minimum overlap length (minOverlapLength) for assembly are set to 1000bp and 500bp, respectively. Whole genome assembly is then performed to obtain the Nanopore contigs.

[0016] To further define the criteria, of the four overlap groups obtained in step 4, the second group is retained as the PacBioHiFi overlap group without redundancy.

[0017] To further specify, in step 6, the chromosome interaction matrix is ​​obtained using Juicer software.

[0018] Further specifying, the indicators analyzed in step 7 are: minimum alignment length of 1000, maximum gap length of 100000, minimum read length support of 2, and minimum alignment identifier of 80.

[0019] To further specify, in step 7, gap information obtained from PacBio HiFi sequencing data is used to fill the gaps first, and then gap information obtained from Nanopore sequencing data is used to fill the gaps again.

[0020] This invention provides a method for assembling a Ganoderma lucidum binuclear genome based on haplotype analysis, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for assembling a Ganoderma lucidum binuclear genome based on haplotype analysis as described above.

[0021] The present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described method for assembling a Ganoderma lucidum binucleate genome based on haplotype analysis.

[0022] Beneficial effects: 1. This invention is based on a haplotype separation assembly strategy, which separates heterozygous diploids into homozygous haploids for sequencing. We innovatively use a compound enzyme digestion solution. Compared with the general method of using only a single enzyme digestion solution, we have created a compound enzyme digestion solution with 20 mg / mL each of lysin, snail enzyme, cellulase and lysozyme, which can make the enzyme digestion more complete and improve the single nucleus efficiency.

[0023] II. This invention is based on a haplotype segregation assembly strategy, which separates heterozygous diploids into homozygous haploids for sequencing. Because the Ganoderma lucidum genome is small, the mitochondrial genome can cause significant interference to its nuclear genome assembly. Therefore, we innovatively introduce the step of removing mitochondrial genome reads to improve the accuracy and integrity of genome assembly.

[0024] Third, this invention is based on a haplotype-separation assembly strategy, which separates heterozygous diploids into homozygous haploids for separate sequencing. Nanopore assembly offers high continuity with an error rate of approximately 2%-5%, while PacBio HiFi assembly boasts accuracy >99.9%, but produces shorter contigs. Therefore, we innovatively use Nanopore assembly results as a framework to perform scaffolding on PacBio HiFi assembly results to obtain preliminary assembled sequences.

[0025] IV. This invention is based on a haplotype-separated assembly strategy, which separates heterozygous diploids into homozygous haploids for sequencing. At the data combination level, it innovatively integrates the three-piece set of "HiFi error correction-assembly → Purge redundancy removal → Hi-C mounting" to form a single-base accurate (QV ≥ 45) and chromosome-complete "T2T-ready" backbone, providing a unique and non-redundant reference coordinate system for subsequent telomere-centromere filling. At the quality control threshold level, it significantly reduces the false fusion rate of high repetitive regions, improving the mounting accuracy from the traditional 85% to ≥95%. At the output specification level, it simultaneously generates two versions, AGP and "gap-unmasked", which can be directly connected to the downstream T2T gap-filling process to achieve a closed-loop strategy of "one-time mounting and iterative gap filling" and avoid rearrangement.

[0026] V. This invention is based on a haplotype segregation assembly strategy, which separates heterozygous diploids into homozygous haploids for sequencing. It innovatively links Hi-C heatmap, HiFi self-correction, and local assembly, and fills ≥80% of the "Hi-C verification gaps" to HiFi quality in one go without increasing data costs, providing a seamless, error-free, and redundancy-free chromosome backbone for subsequent T2T gap filling.

[0027] VI. This invention is based on a haplotype separation assembly strategy, which separates heterozygous diploids into homozygous haploids for sequencing. It innovatively introduces a three-level assessment of "size-sequence-function" using "seqkit + Merqury + BUSCO", providing quantitative conclusions on whether the genome is complete, clean, and usable. This is currently the most recognized assembly quality control combination in international journals. Attached Figure Description

[0028] Figure 1 This is a structural diagram of a binucleate bacterium;

[0029] Figure 2 This is a structural diagram of a mononuclear bacterium;

[0030] Figure 3 The graph shows the effect of assembling the Ganoderma lucidum binucleate genome obtained by the methods of Example 1 and Comparative Example 1 (haploid I); Note: The horizontal axis represents the genome assembled by the original method, and the vertical axis represents the genome assembled by Example 1. The lines (dots) on the graph represent the same regions of the genome assembled by the two methods.

[0031] Figure 4 The graph shows the effect of assembling the Ganoderma lucidum binucleate genome obtained by the methods of Example 1 and Comparative Example 1 (haploid I); Note: The horizontal axis represents the genome assembled by the original method, and the vertical axis represents the genome assembled by Example 1. The lines (dots) on the graph represent the same regions of the genome assembled by the two methods.

[0032] Figure 5 The graph shows the effect of assembling the Ganoderma lucidum binucleate genome obtained by the methods of Example 1 and Comparative Example 1 (haploid I); Note: The horizontal axis represents the genome assembled by the original method, and the vertical axis represents the genome assembled by Example 1. The lines (dots) on the graph represent the same regions of the genome assembled by the two methods.

[0033] Figure 6 The graph shows the effect of assembling the Ganoderma lucidum binucleate genome obtained by the methods of Example 1 and Comparative Example 1 (haploid II); Note: The horizontal axis represents the genome assembled by the original method, and the vertical axis represents the genome assembled by Example 1. The lines (dots) on the graph represent the same regions of the genome assembled by the two methods.

[0034] Figure 7 The graph shows the effect of assembling the Ganoderma lucidum binucleate genome obtained by the methods of Example 1 and Comparative Example 1 (haploid II); Note: The horizontal axis represents the genome assembled by the original method, and the vertical axis represents the genome assembled by Example 1. The lines (dots) on the graph represent the same regions of the genome assembled by the two methods.

[0035] Figure 8The graph shows the effect of assembling the Ganoderma lucidum binucleate genome obtained by the methods of Example 1 and Comparative Example 1 (haploid II); Note: The horizontal axis represents the genome assembled by the original method, and the vertical axis represents the genome assembled by Example 1. The lines (dots) on the graph represent the same regions of the genome assembled by the two methods.

[0036] Figure 9 The graph shows the effect of assembling the Ganoderma lucidum binucleate genome obtained by the methods of Example 1 and Comparative Example 1 (haploid II); Note: The horizontal axis represents the genome assembled by the original method, and the vertical axis represents the genome assembled by Example 1. The lines (dots) on the graph represent the same regions of the genome assembled by the two methods. Detailed Implementation

[0037] MYG medium: 10 g maltose, 5 g glucose, 5 g yeast extract, (15 g agar powder) and 1000 mL distilled water. Mix well and dispense into 250 mL Erlenmeyer flasks. Add 100 mL of liquid medium to each Erlenmeyer flask. Sterilize at 121 °C for 30 min and cool for later use.

[0038] Example 1. Method for obtaining a single-kine strain of Ganoderma lucidum

[0039] The *Ganoderma lucidum* strain was inoculated into MYG culture dishes and incubated statically at 25°C until the dish was full. 0.1g of mycelium was scraped off and immediately transferred to a MYG liquid shake flask for further culture to obtain the desired protoplasts. The dried mycelium was then hydrolyzed using a compound enzymatic hydrolysate (prepared with 20mg / mL each of lysis enzyme, snail enzyme, cellulase, and lysozyme, which ensures more complete hydrolysis and improves mononuclear efficiency) at 29°C and 60 r / min for 3.5 h (1mL of hydrolysate was added for every 0.5g of dried mycelium). After lysis was complete...

[0040] 1. Pour the liquid inoculum into a funnel and filter out the culture medium;

[0041] 2. Rinse three times with sterile water to remove excess culture medium, transfer the hyphae to a centrifuge tube, add osmotic stabilizer (mannitol, concentration 0.6mol / L) and rinse three times, stir evenly with tweezers, and remove the osmotic stabilizer from the centrifuge tube;

[0042] 3. Place the mycelium on filter paper to absorb moisture and the permeation stabilizer;

[0043] 4. Weigh and mark the empty centrifuge tubes, put in the mycelium, weigh them again, and calculate the dry weight of the mycelium;

[0044] 5. Add 3 times the dry weight volume of the mycelium to the compound enzymatic hydrolysate through a bacterial filter, stir and mix well to ensure that the mycelium and enzymatic hydrolysate are thoroughly mixed.

[0045] 6. Place in a metal bath and enzymatically hydrolyze at around 28°C for 3 hours. Filter out the mycelium using absorbent cotton and retain the solution.

[0046] 7. Centrifuge at 3000 rpm for 5-8 minutes, discard the supernatant and retain the precipitate;

[0047] 8. Slowly add osmotic stabilizer to wash the centrifuge tube wall twice to remove residual enzymatic hydrolysate, then discard the supernatant and retain the precipitate. If there is little precipitate, the number of washing cycles can be reduced or washing can be omitted.

[0048] 9. Add 1 ml of osmotic stabilizer, mix thoroughly by suction and blow, and prepare a protoplast solution;

[0049] 1) Hemocytometer concentration determination: Cover a clean hemocytometer with a coverslip, gently add 40 μL of protoplast solution to the edge of the coverslip, being careful not to generate air bubbles, absorb excess bacterial suspension with filter paper, let stand for a moment, and count under an optical microscope. Repeat three times and record the concentration as the original solution concentration.

[0050] 2) Dilution to target concentration: Take 100 μL of protoplast solution, add 900 μL of sterile water, mix well, and dilute again until the target concentration is obtained;

[0051] 10. Plate spreading: Take 200 μL of protoplast solution and evenly drop it into the regeneration medium. Use a sterile spreader to spread it evenly clockwise until the protoplast solution is completely absorbed by the medium. Incubate in the dark at a constant temperature. Pick the regenerated colonies into a 5ml tube of MYG to obtain the protoplast suspension.

[0052] Dilute to 10 3 / mL, 10 4 200 μL of a protoplast suspension of / mL was spread onto MYGS regeneration plates and incubated at 25℃ for 4–6 days. Colonies from the regeneration plates were stained, and those without clamp connections were selected and stained with Congo red. Microscopic examination was performed to observe the presence of clamp connections. The presence of clamp connections indicated a binucleate strain, while the absence of clamp connections indicated a monokaryotic strain. Results are as follows: Figure 1-2 As shown, Figure 1 It has a dual-core structure. Figure 2 It has a single-core structure.

[0053] Example 2. A method for assembling the Ganoderma lucidum binucleate genome based on haplotype analysis

[0054] S1. Genome sequencing:

[0055] The DNA of the isolated mononuclear strain was sequenced using Illumina NovaSeq, Nanopore, Hi-C, and PacBioHiFi platforms.

[0056] S2. Automated genome assembly:

[0057] First, use Fastp software to perform quality control on Illumina NovaSeq sequencing data, with a quality threshold set to 20, retaining reads longer than 150 bp that exist in pairs, in order to obtain high-quality reads;

[0058] The Nanopore sequencing data was quality controlled using Fastplong software with a quality threshold of 20. Then, NanoFilt software was used to remove the first and last 50 bp of the Nanopore sequencing reads, retaining only the higher-precision sequences in the middle. The PacBio HiFi sequencing data was converted from BAM format to FastQ using samtools software. Since PacBio HiFi data has very high base precision, no quality control is required.

[0059] II. Identify and remove reads belonging to the mitochondrial genome from Ganoderma haplotype genome sequencing data;

[0060] Because the Ganoderma lucidum genome is small, the mitochondrial genome can cause significant interference with the assembly of its nuclear genome. Existing methods do not include a mitochondrial removal step, resulting in extremely low accuracy and integrity of the assembled genome. Therefore, we innovatively introduce a step of removing mitochondrial genome reads to improve the accuracy and integrity of genome assembly.

[0061] Mitochondrial data from all Ganoderma species were downloaded to construct a Ganoderma mitochondrial sequence database. BLAST+ was used to align all sequencing data obtained in step S1 of step one with the mitochondrial sequence database. After setting the e-value threshold to 1E-5, reads with a matching length greater than 500 bp or a matching length to read length ratio greater than 0.5 were identified as mitochondrial reads and removed from the sequencing data obtained in step S2 of step one (Illumina NovaSeq, Nanopore, PacBio Hi-C, and PacBio HiFi sequencing data).

[0062] 3. Obtain the overlap group by de novo assembly of PacBio HiFi data (data obtained in step 2) and Nanopore data (data obtained in step 2);

[0063] Using KmerGenie software, the genome size of the Illumina NovaSeq sequencing data (data obtained in step two) was estimated. Then, using the Canu software package (v2.2), the minimum coverage (corMinCoverage) in the correction step was set to 4, the output coverage (corOutCoverage) was set to 40×, and the minimum read length (minReadLength) and minimum overlap length (minOverlapLength) for assembly were set to 1000bp and 500bp, respectively. The genome size was set according to the prediction results of KmerGenie software, and whole-genome assembly was performed to obtain the PacBio HiFi contig.

[0064] Nanopore data (data obtained in step 2) were obtained using the nextdenovo software package. In the calibration step, the minimum coverage (corMinCoverage) was set to 4, the output coverage (corOutCoverage) was set to 40×, the minimum read length (minReadLength) and the minimum overlap length (minOverlapLength) for assembly were set to 1000bp and 500bp, respectively. The genome size was set according to the prediction results of the KmerGenie software, and whole genome assembly was performed to obtain the Nanopore contigs.

[0065] Fourth, the assembled genome is deredundant by combining it with PacBio HiFi data (data obtained in step two) and genome size.

[0066] The PacBio HiFi data (obtained in step two) was aligned with the PacBio HiFi contigs using the map-pb mode of Minimap2 (v2.17) to obtain repetitive sequence information 1. Then, the PacBio HiFi contigs themselves were aligned using the xassm5 mode to obtain repetitive sequence information 2. Redundant genomic sequences were identified and eliminated using Purge_Dups (v1.2.5) with default settings. The results were divided into four categories based on depth, and only the second category was retained to obtain a PacBio HiFi contigs without redundancy.

[0067] 5. Using the Nanopore assembly results (data obtained in step 3) as the framework, scaffolding is performed on the PacBio HiFi assembly results (data obtained in step 4) to obtain a preliminary assembly sequence;

[0068] Nanopore assembly exhibits high continuity with an error rate of approximately 2%-5%; PacBio HiFi assembly accuracy is >99.9%, but contigs are shorter. By comparing Nanopore and PacBio HiFi contigs, sequences with 95% homology (considered as collinear regions) from the Nanopore contigs were replaced with sequences from the PacBio HiFi contigs. This yielded both longer and more accurate candidate contigs, improving single-base accuracy and obtaining candidate genomes.

[0069] VI. Using Hi-C data to perform chromosome-level mounting on candidate genomes;

[0070] Step 1: Based on the obtained Hi-C data (data obtained in Step 2), firstly, Juicer software is used to filter, align, and standardize the effective interactive reads. [1. Remove PCR duplications: Only retain unique read pairs at the same position in the genome (same aligned chromosome, start site, and orientation), and remove duplicate data introduced by PCR amplification;

[0071] 2. Remove low-quality alignments: By default, MAPQ>= 1 is used as the filtering threshold, and only read pairs with minimum mapping quality are retained;

[0072] 3. Remove chimeric reads: Remove reads whose ends align to different chromosomes (inter-chromosomal) or are too far apart on the same chromosome (the default distance threshold is usually >25 kb). These signals may be due to experimental artifacts or incorrect connections.

[0073] 4. Remove "self-circulating" read pairs: Filter out read pairs whose ends align to the exact same genomic location and are in the same orientation (invalid ligation products) and generate a chromosome interaction matrix;

[0074] Step 2: Subsequently, using the 3D-DNA algorithm, based on "intra-chromosomal contact frequency"... The principle of "inter-chromosomal contact frequency" (the interaction frequency within chromosomes is much higher than the interaction frequency between chromosomes; in the Hi-C interaction matrix, this manifests as strongly interacting squares distributed along the diagonal of the matrix (corresponding to a single chromosome), while the signal in regions outside the squares (interactions between different chromosomes) is much weaker) is used to cluster, sort, and orient the candidate genomes and chromosome interaction matrices obtained in step five. [Scaffolds are clustered into chromosome clusters using graph theory methods, and the traveling salesman problem is solved within each cluster using the gradient decay property of chromatin contact frequency to determine the optimal linear order and direction of scaffolds, ultimately obtaining a chromosome-level genome draft], achieving initial chromosome cluster construction, completing automatic chromosome mounting, and outputting a chromosome-level assembled sequence file (FASTA) and a binary interaction matrix file (.hic format) that can be directly used for visualization.

[0075] Step 3: Obtain the interaction heatmap: Import this directly visualizeable binary interaction matrix file (.hic format) along with the assembled sequence file (FASTA) into Juicebox Assembly Tools (JBAT) to generate a Hi-C interaction heatmap. In this heatmap, a correct chromosome assembly will appear as clear blocks distributed diagonally with gradient-decreasing internal interaction signals. According to the Hi-C interaction heatmap, in a correct assembly, the interaction heatmap should show chromosomal blocks distributed diagonally with gradient-decreasing internal signals. Incorrect assembly (such as misalignment or inversion) will result in abnormally strong interaction signals perpendicular or parallel to the diagonal within the blocks, ultimately yielding a genome with a consistent topology supported by Hi-C data.

[0076] Based on this, the final chromosome-level assembled sequence (FASTA) and its structure description file (AGP) are output, ensuring that ≥90% of the contig sequences are accurately mounted onto the chromosome, and finally obtaining a high-quality chromosome-level genome validated by Hi-C data. The final chromosome-level genome sequence and the corresponding AGP file are output, ensuring that ≥90% of the candidate genomes are accurately mounted onto the chromosome, and ensuring single-base accuracy and structural integrity, thus obtaining the chromosome-level genome sequence.

[0077] 7. Fill gaps in the genome sequence after Hi-C correction (data obtained in step 6);

[0078] Using 5kb intervals to the left and right of the gap as anchor points, PacBio HiFi sequencing data (data obtained in step two) and Nanopore sequencing data (data obtained in step two) were extracted. The PacBio HiFi sequencing data was obtained using LR_gapclosed software with default parameters (minimum alignment length (bp) defaults to 1000, meaning long reads must be aligned to at least 1000 bp to be considered; maximum gap length (bp) defaults to 100000, meaning only gaps ≤ 100 kb are closed; minimum read support number defaults to 2, meaning at least 2 long reads are needed for closure; minimum alignment identifier (%) defaults to 80, meaning sequence similarity ≥ 80%). This data was first filled into the genome sequence to obtain the filled genome sequence. Then, the Nanopore sequencing data was processed using LR_gapclosed software with default parameters (minimum alignment length (bp) defaults to 1000, meaning long reads must be aligned to at least 1000 bp). Only when the maximum gap length (bp) is considered is it considered. The default is 100,000, which means that only gaps with a length ≤ 100 kb are closed. The default is 2, which means that at least 2 long reads are required to support closure. The default is 80, which means that the alignment similarity is ≥ 80%. The data obtained is then filled into the above-mentioned filled genome sequence to obtain the genome with gaps filled, which maintains QV≥45 while significantly improving chromosome continuity.

[0079] 8. Use bwa software to compare the Illumina NovaSeq data (the data obtained in step 2) with the genome (the genome with gaps filled in step 7). Use the Illumina NovaSeq data as the standard to find the assembly errors. Then use freebayes software to correct the genome based on the alignment results to finally obtain a high-quality genome assembly result.

[0080] 9. Integrate data from two mononuclear strains.

[0081] Evaluation of genome assembly results:

[0082] Basic information about the assembled genome was statistically analyzed using seqkit (v2.10.0), including genome size, number of chromatids, GC content, and N50 information. Next, the assembly results were evaluated using Illumina resequencing data and Merqury (v1.3) software to determine k-mer integrity and the proportion of false duplications. Finally, the integrity of the assembled gene elements was assessed using BUSCO (v5.8.2) software. Merqury assessed k-mer integrity and false duplications; BUSCO assessed genome integrity; other indicators were statistically derived using seqkit software.

[0083] (Merqury: reference-free quality, completeness, and phasingassessment for genome assemblies.

[0084] BUSCO: Assessing Genome Assembly and Annotation Completeness.)

[0085] This allows for a three-tiered assessment of the genome size, sequence, and function using "seqkit + Merqury + BUSCO," providing quantitative conclusions on genome integrity, cleanliness, and usability, along with charts. It is currently the most widely recognized assembly quality control combination in international journals. (Comparative Example 1)

[0086] The performance of our proposed method in assembling the Ganoderma lucidum binucleate genome was compared with that of the traditional assembly method (publication number CN202510884408.8). Evaluation results using BUSCO and Merqury software showed that the Ganoderma lucidum binucleate genome assembled by our method exhibited significantly improved integrity and accuracy. Furthermore, we compared the results of the Ganoderma lucidum binucleate genome assembled using our method with those obtained by previous methods, such as... Figure 3-9 As shown in the figure, inconsistent genomic regions represent areas where previous methods resulted in assembly errors. Genomic collinearity analysis was used to compare the effectiveness of our method and the previous method in assembling the binucleate genome of Ganoderma lucidum.

[0087] Results: The Busco genome integrity score was 78% using the traditional assembly method, while the Busco genome integrity score using the method in Example 2 was 99.2%. The N50 index of the method in Example 2 was twice that of the traditional assembly method. The Kmer integrity of the method in Example 2 was 99.1%, compared to 79% for the traditional assembly method.

Claims

1. A method for assembling the Ganoderma lucidum binuclear genome based on haplotype analysis, characterized in that, The steps of the method are as follows: Step 1: Processing Ganoderma lucidum to obtain two monokaryotic strains of Ganoderma lucidum. After extracting the DNA of each monokaryotic strain, whole genome sequencing was performed using Illumina NovaSeq, Nanopore, Hi-C and PacBio HiFi respectively. Quality control was performed on the Illumina NovaSeq sequencing data and the Nanopore sequencing data to obtain high-quality read length data. Step 2: The mitochondrial data of Ganoderma species downloaded from the database are compared with the Hi-C, PacBio HiFi, Illumina NovaSeq and Nanopore data from Step 1, and the mitochondrial data are removed. Step 3: Based on the Illumina NovaSeq sequencing data, determine the genome size of the PacBio HiFi data and Nanopore data obtained in Step 2, and perform de novo assembly to obtain the PacBio HiFi contig and Nanopore contig. Step 4: Align the PacBio HiFi data obtained in Step 2 with the PacBio HiFi overlap group obtained in Step 3 to obtain repeating sequence information. Then, perform self-alignment on the PacBio HiFi overlap group obtained in Step 3 to obtain repeating sequence information. Then, use the repeating sequence as redundant information and use software to process the redundant information to obtain a PacBio HiFi overlap group without redundancy. Step 5: Align the Nanopore contigs obtained in Step 3 with the non-redundant PacBio HiFi contigs obtained in Step 4. The sequences with 95% homology are considered collinear segments. Replace the collinear segments of the Nanopore contigs with the collinear segments of the PacBio HiFi contigs to obtain the candidate genome. Step 6: Prepare a chromosome interaction matrix from the Hi-C data obtained in Step 1. Use the 3D-DNA algorithm to process the candidate genomes and chromosome interaction matrix obtained in Step 5 to obtain chromosome clusters. Then, use the chromosome clusters to prepare an interaction heatmap. Adjust the candidate genomes according to the chromosome map to obtain the genome sequence. Step 7: Use software to analyze the PacBio HiFi sequencing data and Nanopore sequencing data obtained in Step 1 to obtain gap data, and then fill the gap data into the gene sequence obtained in Step 6 to obtain a new gene sequence; Step 8: Using the Illumina NovaSeq data obtained in Step 1 as a benchmark, perform whole-genome correction on the new gene sequences obtained in Step 8 to obtain a high-quality genome, and then integrate the genomes of two mononuclear Ganoderma lucidum.

2. The method according to claim 1, characterized in that, The method for removing mitochondrial sequences in step 2 is as follows: the e-value threshold is set to 1E-5, and then reads with a length greater than 500bp or a length-to-read length ratio greater than 0.5 are matched and removed from the PacBio Hi-C data, PacBio HiFi data obtained in step 1 and the Illumina NovaSeq data and Nanopore data obtained in step 2.

3. The method according to claim 1, characterized in that, Of the four overlap groups obtained in step 4, the second group is retained as the PacBio HiFi overlap group without redundancy.

4. The method according to claim 1, characterized in that, The metrics analyzed in step 7 are: minimum alignment length of 1000, maximum gap length of 100000, minimum read length support of 2, and minimum alignment identifier of 80.

5. The method according to claim 1, characterized in that, In step 7, gap information obtained from PacBio HiFi sequencing data is first used for filling, and then gap information obtained from Nanopore sequencing data is used for filling.

6. A method for assembling the Ganoderma lucidum binuclear genome based on haplotype analysis, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a method for assembling a Ganoderma lucidum binucleate genome based on haplotype analysis as described in any one of claims 1-5.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements a method for assembling a Ganoderma lucidum binucleate genome based on haplotype analysis as described in any one of claims 1-5.