Genome assembling method
By combining data fusion assembly from the PacBio HiFi and Oxford Nanopore platforms, performing multiple rounds of correction and redundant sequence removal, and using Hi-C data to optimize structures and fill gaps, the assembly challenges of complex genomes were resolved, enabling the acquisition of high-quality genomic data.
Patent Information
- Application Number
- CN202510884408.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-30
AI Technical Summary
Existing genome assembly technologies face problems with read length, accuracy, and coverage uniformity when faced with structurally complex or highly repetitive genomes, which challenges the integrity and accuracy of the assembly results.
PacBio HiFi technology was used for preliminary assembly, combined with ultra-long read data from the Oxford Nanopore platform for independent assembly, and the assembly results of the two platforms were fused through QuickMerge. Subsequently, multiple rounds of high-precision correction and redundant sequence removal were implemented, and the chromosome framework was constructed using Hi-C data for structural optimization. Finally, MaSuRCA and TGS-GapCloser were used to fill the gaps.
It significantly improves the continuity and integrity of the genome, increases base accuracy, solves the problem of assembling complex genomes, and provides higher quality genome data.
Smart Images

Figure SMS_1
Abstract
Description
Technical Field
[0001] The present invention relates to a genome assembly method. Background Art
[0002] With the rapid development of biotechnology, human beings have been conducting more and more in-depth research on genomes. As an important technology in genome research, genome assembly is of great significance for understanding the structure and function of biological genomes. With the development of sequencing technology and computing methods, genome assembly technology has gradually evolved from the early stage of low efficiency, high cost, and high manual intervention to the current high-throughput and continuously improved automation process system. However, in the face of genomes with complex structures or high repetitiveness, existing sequencing methods still have limitations in read length, accuracy, and coverage uniformity. In addition, due to the bottleneck of computing resources, the integrity and accuracy of the assembly results are still challenged. In the future, it is necessary to further improve the accuracy and integrity of the assembly of complex genomes to promote the development of precision medicine and functional genomics. Summary of the Invention
[0003] The present invention provides a genome assembly method that can effectively improve the accuracy and completeness of the genome.
[0004] The genome assembly method of the present invention is carried out according to the following steps:
[0005] First, we used high-precision long-read raw reads obtained using PacBio HiFi technology to directly perform preliminary genome assembly using hifiasmv0.19.8-r603 with default parameters. Simultaneously, based on ultra-long read data from the Oxford Nanopore platform, we screened out sequences with read lengths greater than 20 kb and independently assembled them using NextDenovo v2.5.2 with default parameters.
[0006] Second, using the assembly results of HiFi data as a reference, use nucmer to filter out fragments longer than 100 bp. Compare the two assembly results in step 1, and then use delta-filter to filter out fragments less than 10 kb.
[0007] 3. Use QuickMerge v0.3 to merge the assemblies of the PacBio HiFi and Oxford Nanopore platforms;
[0008] 4. Implement multiple rounds of high-precision polishing to maximize base accuracy;
[0009] Fifth, using HiFi data as a reference, minimap2 v2.17 was used to generate a paf file, and Purge-Dups v1.2.6 was used to remove redundant sequences based on the file;
[0010] 6. Chromosome frameworks were constructed using Hi-C data, and chromosome interaction maps were generated using Juicer v1.6. Three rounds of structure optimization and manual correction were performed using 3D-DNA and JuiceBox v2.13.07, using default parameters.
[0011] 7. Use MaSuRCA v4.1.0 in combination with TGS-GapCloser v1.1.1, using the ont data as the raw data, to complete the gap filling in the remaining regions, that is, to achieve genome assembly.
[0012] The assembly result parameters in step 3 are quickmerge -d out.rq.delta -qont.assembly.fa -r hifi.assembly.fa -hco 5.0 -c 1.5 -l 0 -ml 5000.
[0013] The method for improving base accuracy in step 4 is as follows: Illumina second-generation data is polished once, PacBio HiFi data is polished five times, and Nanopore data is polished three times; the tool is NextPolish v1.4.1.
[0014] In step 5, the redundant sequence removal parameters are purge_dups -2 -f.95 -l 100 -T cutoffs -cPB.base.cov asm.split.self.paf.gz > dups.bed, and other parameters are default parameters.
[0015] In step 7, the parameters of MaSuRCA v4.1.0 combined with TGS-GapCloser v1.1.1 are tgsgapcloser --scaff assembly.chr.fasta --reads ont.reads.fa --output un_gap.chr.fa --minmap_arg '-x splice / splice:hq' --tgstype pb --thread 180.
[0016] The method presented in this paper utilizes high-precision long-read data obtained with PacBio HiFi technology to directly perform a preliminary chromosome-level genome assembly using hifiasmv0.19.8-r603. Simultaneously, ultra-long read data from the Oxford Nanopore platform were independently assembled using NextDenovo v2.5.2. Subsequently, the assembly results from the two platforms were merged using QuickMerge v0.3, using the HiFi data as a benchmark to enhance sequence integrity and continuity. Furthermore, multiple rounds of high-precision polishing were performed to maximize base accuracy. To further enhance genome continuity and non-redundancy, redundant sequences were removed using Purge-Dups v1.2.6 according to the documentation. Subsequently, chromosome frameworks were constructed using Hi-C data, and chromosome interaction maps were generated using Juicer v1.6. Three rounds of structure optimization and manual proofreading were performed using 3D-DNA and JuiceBox v2.13.07 to ensure the accuracy and structural rationality of the chromosome assembly. To address possible residual genomic gaps, the present invention further used MaSuRCA v4.1.0 to construct an ultra-long scaffold and combined it with TGS-GapCloser v1.1.1 to complete the gap filling in the remaining regions, thereby comprehensively improving the continuity and integrity of the genome.
[0017] Compared with previous genome assembly methods, the assembly method of the present invention has achieved the following significant breakthroughs: First, predecessors often only used Oxford Nanopore data for gap filling, while the present invention uses it as an independent assembly data source and participates in data fusion, significantly improving the ability to resolve ultra-long repetitive regions; Second, predecessors mostly adopted a one-time polishing strategy, while the present invention combines multi-round, multi-platform data correction methods, significantly improving the accuracy at the base level and the overall quality of the genome. DETAILED DESCRIPTION
[0018] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0019] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0020] Example 1: Assembling the whitebait genome using the genome assembly method of the present invention
[0021] The steps for genome assembly are:
[0022] First, we used high-precision long-read raw reads obtained using PacBio HiFi technology to directly perform preliminary genome assembly using hifiasmv0.19.8-r603 with default parameters. Simultaneously, based on ultra-long read data from the Oxford Nanopore platform, we screened out sequences with read lengths greater than 20 kb and independently assembled them using NextDenovo v2.5.2 with default parameters.
[0023] Second, using the assembly results of HiFi data as a reference, use nucmer to filter out fragments longer than 100 bp. Compare the two assembly results in step 1, and then use delta-filter to filter out fragments less than 10 kb.
[0024] 3. Use QuickMerge v0.3 to merge the assembly results of the two platforms. The parameters for the assembly result in step 3 are quickmerge -d out.rq.delta -q ont.assembly.fa -r hifi.assembly.fa -hco 5.0 -c 1.5 -l 0 -ml 5000 to improve the integrity and continuity of the sequence;
[0025] Fourth, multiple rounds of high-precision polishing were performed to maximize base accuracy, including one polishing run on Illumina second-generation data, five polishing runs on PacBio HiFi data, and three polishing runs on Nanopore data, using NextPolish v1.4.1.
[0026] 5. To further improve genome continuity and non-redundancy, minimap2v2.17 was used to generate a paf file based on HiFi data. Purge-Dups v1.2.6 was used to remove redundant sequences based on this file. In step 5, the redundant sequence removal parameters were purge_dups -2 -f.95 -l 100 -T cutoffs -c PB.base.covasm.split.self.paf.gz > dups.bed, and other parameters were default parameters.
[0027] 6. Chromosome frameworks were constructed using Hi-C data, and chromosome interaction maps were generated using Juicer v1.6. Three rounds of structure optimization and manual correction were performed using 3D-DNA and JuiceBox v2.13.07, using default parameters.
[0028] VII. To resolve possible remaining genomic gaps, MaSuRCA v4.1.0 was further used in combination with TGS-GapCloser v1.1.1 (the parameters of MaSuRCA v4.1.0 and TGS-GapCloser v1.1.1 are tgsgapcloser --scaff assembly.chr.fasta --reads ont.reads.fa --output un_gap.chr.fa --minmap_arg '-x splice / splice:hq' --tgstype pb --thread 180). The ont data were used as the raw data to complete the gap filling in the remaining regions and obtain the giant whitebait genome.
[0029] This example uses high-precision long-read data obtained by PacBio HiFi technology to directly perform preliminary chromosome-level genome assembly using hifiasmv0.19.8-r603. At the same time, NextDenovo v2.5.2 was used to independently assemble the ultra-long read data based on the Oxford Nanopore platform. Subsequently, QuickMerge v0.3 was used to fuse the assembly results of the two platforms based on the assembly results of the HiFi data to improve the integrity and continuity of the sequence. On this basis, multiple rounds of high-precision correction (polishing) were implemented to maximize the accuracy of the bases. Specifically, it includes: one polishing using Illumina second-generation data, five polishings using PacBio HiFi data, and three polishings using Nanopore data, using NextPolish v1.4.1 as the tool. In order to further improve the continuity and non-redundancy of the genome, Purge-Dups v1.2.6 was used to remove redundant sequences. Subsequently, the chromosome framework was constructed using Hi-C data, and a chromosome interaction map was generated using Juicer v1.6. Three rounds of structure optimization and manual correction were performed using 3D-DNA and JuiceBox v2.13.07 to ensure the accuracy and structural rationality of the chromosome assembly. To address possible remaining genomic gaps, MaSuRCA v4.1.0 was used to construct an extra-long scaffold, and TGS-GapCloser v1.1.1 was used to fill gaps in the remaining regions, resulting in the generation of the giant whitebait genome.
[0030] Compared with previous genome assemblies of the whitebait, the method of this embodiment achieved significant breakthroughs in the following aspects: First, while previous researchers often only used Oxford Nanopore data for gap filling, this embodiment used it as an independent assembly data source and participated in data fusion, significantly improving the parsing ability of very long repeat regions; second, previous researchers often used a one-time polishing strategy, while this embodiment combined multi-round, multi-platform data correction methods to significantly improve base-level accuracy and overall genome quality. This embodiment comprehensively improved the continuity and integrity of the whitebait genome.
[0031] The high-completeness of the whitebait genome data obtained using the method in this example can be widely applied in multiple fields. In biomedicine, it can be used for genome alignment and functional analysis related to human diseases, assisting in drug target screening and gene therapy development. In aquaculture breeding, it can leverage validated molecular markers (such as the mutant mc1r) to enable targeted breeding for specific body color or skeletal traits. In evolutionary biology research, this genomic resource can be used for comparative genomic analyses of the phylogeny and trait evolution of whitebait and its closely related populations. The results of this invention have broad scientific research value and potential for industrial transformation.
[0032] Example 2 Comparison of the method of Example 1 with the existing method
[0033] The experimental groups were set up as follows: 2025 (method of the present invention), 2025 (HiFi and Hi-C assembly, gap filling), 2023 (HiFi reads and Hi-C assembly), 2020 (PacBio Clr reads assembly), and 2017 (Illumina reads assembly). The experimental results are shown in Table 1.
[0034] Table 1 Assembly results of the whitebait genome
[0035]
[0036] As shown in Table 1, the high-quality genome of Hypomesus nipponensis assembled in this paper has a total length of 430.5 Mb, significantly larger than the existing reference genome of 383.8 Mb. BUSCO assessment shows that the genome completeness of this assembly reaches 94.4%, also significantly higher than the 89.6% of the previous version. This improved assembly quality lays a solid foundation for the accurate identification and annotation of key functional genes.
Claims
1. A genome assembly method, characterized in that The genome assembly method follows these steps: First, we used high-precision long-read raw reads obtained using PacBio HiFi technology to directly perform preliminary genome assembly using hifiasm v0.19.8-r603 with default parameters. Simultaneously, based on ultra-long read data from the Oxford Nanopore platform, we screened out sequences with read lengths greater than 20 kb and independently assembled them using NextDenovo v2.5.2 with default parameters. Second, using the assembly results of HiFi data as a reference, use nucmer to filter out fragments longer than 100 bp. Compare the two assembly results in step 1, and then use delta-filter to filter out fragments less than 10 kb.
3. Use QuickMerge v0.3 to merge the assemblies of the PacBio HiFi and Oxford Nanopore platforms; 4. Implement multiple rounds of high-precision polishing to maximize base accuracy; Fifth, using HiFi data as a reference, minimap2 v2.17 was used to generate a paf file, and Purge-Dups v1.2.6 was used to remove redundant sequences based on the file; 6. Chromosome frameworks were constructed using Hi-C data, and chromosome interaction maps were generated using Juicer v1.
6. Three rounds of structure optimization and manual correction were performed using 3D-DNA and JuiceBox v2.13.07, using default parameters.
7. Use MaSuRCA v4.1.0 in combination with TGS-GapCloser v1.1.1, using the ont data as the raw data, to complete the gap filling in the remaining regions, that is, to achieve genome assembly.
2. A genome assembly method according to claim 1, characterized in that, The assembly result parameters in step 3 are quickmerge -d out.rq.delta -q ont.assembly.fa -r hifi.assembly.fa -hco5.0 -c 1.5 -l 0 -ml 5000.
3. A genome assembly method according to claim 1, characterized in that, Methods for improving base accuracy in step 4: polish once for Illumina second-generation data, five times for PacBio HiFi data, and three times for Nanopore data; The tool is NextPolish v1.4.
1.
4. A genome assembly method according to claim 1, characterized in that, In step 5, the redundant sequence removal parameter purge_dups -2 -f.95 -l 100 -T cutoffs -c PB.base.covasm.split.self.paf.gz > dups.bed is used, and other parameters are default parameters.
5. A genome assembly method according to claim 1, characterized in that, In step 7, the parameters of MaSuRCA v4.1.0 combined with TGS-GapCloser v1.1.1 are tgsgapcloser --scaffassembly.chr.fasta --reads ont.reads.fa --output un_gap.chr.fa --minmap_arg'-x splice / splice:hq' --tgstype pb --thread 180.
Citation Information
Patent Citations
Preparation method for circular single-stranded DNA integrated with aptamer and applications of circular single-stranded DNA integrated with aptamer in DNA origami
CN113278607A
Polyploidy genome assembling method and device based on third-generation sequencing
CN113496760A
Telomere-to-telomere genome assembly method
CN115691673A
Method and system for multiple dot plot analysis
KR101832834B1
Apparatus and method constructing consensus reference genome map
KR1020180083706A
Cited By
Lucid ganoderma binuclear genome assembly method based on haplotype analysis
CN121565256A
A method for assembling ganoderma bicomb genome based on haplotype resolution
CN121565256B