Reference-level high-quality genome assembling method
By combining second-generation Illumina short read length, third-generation HiFi sequencing and Hi-C chromatin conformation capture sequencing, the high cost of genome assembly of cypress can be solved by using compHiFi tools, and efficient T2T-level genome assembly is achieved, reducing sequencing costs and improving assembly quality.
Patent Information
- Application Number
- CN202510410811.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art has high cost problems when performing reference-level high-quality genome assembly, especially for the genome of phyton, and insufficient assembly technology, making it difficult to achieve efficient and economical T2T level genome assembly.
The second-generation Illumina short-read, third-generation HiFi sequencing technology and Hi-C chromatin conformation capture sequencing were used, and genome assembly was performed in combination with compHiFi tools, and initial assembly was performed through Hifiasm software. Clustering and orientation were performed using Juicer and 3D-DNA processes, error correction was used using Juicebox, and gaps were filled through verkko, canu, flye and nextdenovo tools to finally generate high-quality genomes.
It realizes the sequencing cost significantly reduces the reference-level high quality and even T2T level genome assembly without high cost ONT sequencing, providing cost-effective and effective technical support for the research of cypress.
Smart Images

Figure CN120260685A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of genomics and bioinformatics, and particularly to a method for assembling a reference-quality genome and its application in genome assembly of Beckmannia syzigachne at the telomere-to-telomere (T2T) level. Background Art
[0002] The emergence and application of second-generation and third-generation sequencing technologies have rapidly promoted the development of high-throughput genomics research. Among them, de novo genome assembly based on data from these two sequencing technologies is the basis of genomics research. The second-generation sequencing technology is widely used due to its high accuracy and low cost advantages. However, due to its short read lengths, for example, the paired-end sequencing of the Hiseq platform of Illumina company is only about 150bp, making it difficult to span repetitive sequences and complex regions rich in GC content in genome assembly. In contrast, third-generation sequencing technologies, such as the single-molecule real-time sequencing (SMRT) platform of Pacific Bioscience company, have significantly increased the read lengths. The longest read length of the PacBio RS II P5-C3 platform can reach 30kb, with an average read length of about 8kb. The P6-C4 platform has 50% of the read lengths exceeding 20kb, and the accuracy can be significantly improved by increasing the sequencing depth (one cell produces 90G of data, about 15,000 yuan). The MinION sequencing platform of Oxford Nanopore Technologies (ONT) company can generate ultra-long read lengths, up to the Mb level, but its accuracy is relatively low and the cost is also higher (one cell produces 15G of data, about 12,000 yuan). The third-generation sequencing technologies have greatly promoted the de novo genome assembly work, being able to cross high-repeat regions and obtain high-quality and highly continuous chromosome-level or even T2T-level genomes. Usually, to improve the assembly quality, HiFi data is generally used to correct the ONT ultra-long sequencing data to make up for the low accuracy problem, resulting in extremely high sequencing costs. Therefore, considering the balance between effect and cost, the present invention can achieve the splicing of a reference-quality genome and effectively improve the analysis ability of complex regions by formulating reasonable sequencing standards and optimized assembly strategies, combining second-generation and third-generation HiFi sequencing technologies, providing a more efficient and economical assembly technical solution for ordinary laboratories and biological research institutions.
[0003] Beckmannia syzigachne (Steud.) Fernald. (2n = 14) is an annual or perennial diploid self-pollinating plant of the genus Beckmannia in the Poaceae family, widely distributed globally (https: / / powo.science.kew.org). Beckmannia syzigachne has strong spreading ability and is a dominant and pernicious weed in wheat fields and rapeseed fields in the main grain-producing areas of China (especially in the middle and lower reaches of the Yangtze River), where rice-wheat and rice-rape double cropping systems are practiced. Beckmannia syzigachne has well-developed roots and sufficient field seed sources. Generally, it can reduce wheat yields by 10 - 30%, and in severely affected fields, the reduction can even exceed 50%. Under the long-term selection pressure of herbicides, Beckmannia syzigachne has developed resistance to major herbicides. Limited by the lack of a reference genome, the research on related functions and mechanisms has been hindered. A complete, high-quality, and well-annotated genome is the basis for the study of the adaptive evolution of Beckmannia syzigachne and the identification of resistance genes.
[0004] The above references are as follows:
[0005] Pop, M. and S. L. Salzberg, Bioinformatics challenges of new sequencing technology. Trends Genet, 2008. 24(3): p. 142 - 9.
[0006] Tilak, M. K., et al., Illumina Library Preparation for Sequencing the GC-Rich Fraction of Heterogeneous Genomic DNA. Genome Biol Evol, 2018. 10(2): p. 616 - 622.
[0007] Koren, S., et al., Hybrid error correction and de novo assembly of single-molecule sequencing reads. Nat Biotechnol, 2012. 30(7): p. 693 - 700.
[0008] Eisenstein, M., Oxford Nanopore announcement sets sequencing sector abuzz. Nat Biotechnol, 2012. 30(4): p. 295 - 6.
[0009] Jain,M.,et al.,The Oxford Nanopore MinION:delivery of nanopore sequencing to the genomics community.Genome Biol,2016.17(1):p.239.
[0010] Rao,N.,et al.,Influence of environmental factors on seed germination and seedling emergence of American sloughgrass(Beckmannia syzigachne).Weed Science,2008.56(4):p.529-533.
[0011] Li,Y.,Weed Flora of China.1998,Beijing,China:China Agriculture Press.1171-1172.
[0012] Li,L.,et al.,Molecular mechanism of mesosulfuron-methyl resistance in multiply-resistant American sloughgrass(Beckmannia syzigachne).Weed Science,2015.63(4):p.781-787.
[0013] Jun,L.I.,et al.,Occurrence dynamics of Bechmannia syzigachne in winter wheat and its influence on growth and production of winter wheat.Journal of Nanjing Agricultural University,2010.33(3):p.67-70. Summary of the Invention
[0014] The object of the present invention is to provide a more cost-effective sequencing solution and an efficient tool for reference-level high-quality and even T2T-level genome assembly, aiming at the high cost problem and the deficiencies of assembly technology faced by using ONT ultra-long sequencing in reference-level high-quality genome assembly projects.
[0015] To achieve the above object, the present invention provides a reference-level high-quality genome assembly method, adopting the following technical solutions:
[0016] Step 1, construct a high-quality genomic DNA library. High-quality genomic DNA is a prerequisite for constructing a high-fidelity (HiFi) sequencing library. Use the FineOut Universal Plant DNA Kit to extract genomic DNA, and use FEMTOPulse (Agilent) to evaluate the quality of the isolated DNA. Prepare the HiFi sequencing library according to the standardized process provided by PacBio. Use the standard Hi-C method to prepare a high-throughput chromosome conformation capture (Hi-C) library from cross-linked chromatin.
[0017] Step 2, high-throughput genome sequencing. Use Illumina short-read sequencing technology to evaluate the size of the genome and determine the amount of data required for sequencing. Subsequently, sequence the HiFi sequencing library described in Step 1 using single-molecule real-time (SMRT) sequencing technology on the Revio platform to generate high-fidelity HiFi sequencing data. Sequence the Hi-C library using the DNBSEQ platform to obtain paired-end Hi-C sequencing data, and use the HiC-Pro pipeline to evaluate the quality of the Hi-C experiment.
[0018] Step 3, build and optimize the genome assembly tool. Through the compHiFi tool, the HiFi reads generated by the PacBio Revio instrument are initially assembled by the Hifiasm software (Cheng, H., et al., 2021) to obtain an initial highly contiguous genome assembly. Use the paired-end Hi-C sequencing data to cluster, orient, and connect the initial highly contiguous genome assembly into pseudochromosomes through the Juicer (Durand, N.C., et al., 2016) and 3D-DNA pipeline (Dudchenko, O., et al., 2017), and then manually correct potential assembly errors using Juicebox (Robinson, J.T., et al., 2018) to generate a high-quality genome assembly sequence.
[0019] Step 4, perform refinement processing on the assembled genome. Use the developed compHiFi tool to assemble the HiFi sequencing data with multiple genome assembly tools (verkko, canu, flye, and nextdenovo) to obtain contigs, and use quartet to fill the gaps in each chromosome (Hu, J., et al., 2024; Kolmogorov, M., et al., 2019; Koren, S., et al., 2017; Lin, Y., et al., 2023; Rautiainen, M., et al., 2023), and finally generate a reference-level high-quality genome up to the T2T level.
[0020] Step 5, evaluate the quality of genome assembly. Evaluate the integrity of genome assembly through the BUSCO pipeline ( F.A., et al., 2015), and use the LTR assembly index (LAI) of long terminal repeat retrotransposons (LTR-RTs) as an indicator independent of the reference genome to evaluate the quality of genome assembly. Further, determine the LAI score of the assembled genome through the Merqury tool (Ou, S. and N. Jiang, 2018) and its default settings. And use the Merqury tool to calculate the quality value (QV) of the genome (Rhie, A., et al., 2020) to ensure the high quality of the generated genome.
[0021] Among them, the genome assembly tool compHiFi is applicable to integrating Illumina short reads, PacBio HiFi data, and Hi-C data, and supports the assembly of reference-level high-quality genomes up to the T2T level for multiple biological individuals.
[0022] Preferably, in step 1, a single plant should be selected, and after separating the spike, leaves, stems, and roots, they should be immediately stored in dry ice to ensure the integrity of DNA.
[0023] Preferably, in step 2, the average read length of the HiFi sequencing library is at least 25 kb to span most transposable elements. The HiFi sequencing depth is at least 70×, the Illumina short-read sequencing depth is at least 100×, and the Hi-C sequencing library is sequenced at least 140× paired-end.
[0024] Preferably, the restriction endonuclease used for Hi-C in step 2 is [DpnII], and the default parameters are used for HiCUP for quality assessment.
[0025] Preferably, the default parameters are used for Hifiasm used for de novo assembly in step 3.
[0026] Preferably, the command used to mount 3D-DNA onto chromosomes in step 3:
[0027] run-asm-pipeline.sh -r 0hifiasm.nextpolish.faa aligned / merged_nodups.txt.
[0028] Preferably, in step 3, the resulting output is manually verified using juicebox, and then the final assembled genome is generated. The command is as follows: run-asm-pipeline-post-review.sh -r genome.0.review.assembly genome.fa aligned / merged_nodups.txt.
[0029] Preferably, when using the nextDenovo, Flye, and Canu assembly tools in step 4, the estimated genome size of Beckmannia syzigachne is 2.85 G. According to the computing resources, a total of 128 G of memory and 64 threads are used for the assembly calculation.
[0030] Preferably, the lineage dataset used by BUSCO in step 5 is embryophyta_odb10.
[0031] Preferably, when Merqury is used to evaluate the genome assembly quality in step 5, k is set to 21, and 128 G of memory and 64 threads are used.
[0032] The beneficial effects of the present invention are as follows: The present invention provides an economical and effective method for assembling reference-level high-quality genomes. Without the need for high-cost ONT sequencing, it can perform genome assembly of various target species using only second-generation Illumina short reads, third-generation HiFi (Revio platform, average read length of 25 kb), and Hi-C chromatin conformation capture sequencing, by using the sequencing standards and compHiFi tools of the present invention, thereby significantly reducing the sequencing cost and obtaining a reference-level high-quality genome with high accuracy and high continuity, even up to the T2T level. In addition, the present invention uses this sequencing standard and assembly strategy to achieve the genome assembly of the malignant weed Beckmannia syzigachne, providing new technical support for research in the field of genomics and data support for the rapid adaptation and evolution of weeds under environmental stress. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a flowchart of the sequencing standard and the use of the reference-level high-quality genome assembly tool compHiFi;
[0034] Figure 2 It is the genome assembly of Beckmannia syzigachne; whereinFigure 2 In which, A is the Hi-C map of the Beckmannia syzigachne genome: showing all interaction relationships across the entire genome at a resolution of 500 kb; Figure 2 In which, B is the Circos map of the Beckmannia syzigachne genome: a, gene number; b, GC content; c, Copia transposon coverage; d, Gypsy transposon coverage; e, Helitron transposon coverage; f, tandem repeats; the inner gray lines show the collinearity relationships within and between chromosomes; Figure 2 In which, C is the visualization of the GFA (Graphical Fragment Assembly) result, showing the high heterozygosity of the gap region on chromosome 6; Figure 2 In which, D is the BUSCO assessment result of genome assembly and annotation; Figure 2 In which, E is the evaluation of genome assembly quality by the LTR Assembly Index (LAI);
[0035] Figure 3 are the characteristics of the telomere and centromere regions of the Beckmannia syzigachne genome;
[0036] Figure 4 is the distribution of the telomere repeat unit (TTTAGGG / CCCTAAA) identified by the method of the present invention in melon and carrot. Among them, Figure 4 In which, A is its distribution in melon, Figure 4 In which, B is its distribution in carrot;
[0037] Figure 5 is the comparison of the chromosome assembly result consistency between the method of the present invention (HiFi+Hi-C) and the method reported in the literature (second-generation+HiFi+Hi-C+ONT). Among them, Figure 5 In which, A is the comparison of the chromosome assembly result consistency of melon, Figure 5 In which, B is the comparison of the chromosome assembly result consistency of carrot. Detailed implementation mode
[0038] The present invention will be further clarified below with reference to the accompanying drawings and specific embodiments.
[0039] Example 1
[0040] The present invention provides a reference-level high-quality genome assembly tool and can be applied to the de novo assembly of high-quality genomes of animals and plants including Beckmannia syzigachne. The specific process is as Figure 1 shown. At the same time, the present invention realizes the nearly T2T assembly of the Beckmannia syzigachne genome for the first time, including the following steps:
[0041] Step 1, construct a high-quality genomic DNA library. High-quality genomic DNA is a prerequisite for constructing a high-fidelity (HiFi) sequencing library.
[0042] First, single plants of *Beckmannia syzigachne* were selected, and the spikes, leaves, stems, and roots were quickly separated. Subsequently, the samples were immediately stored in dry ice. Genomic DNA was extracted using the FineOut Universal Plant DNA Kit, and the quality of the isolated DNA was evaluated using FEMTO Pulse (Agilent). A HiFi sequencing library was prepared according to the standardized protocol provided by PacBio. Meanwhile, a high-throughput chromosome conformation capture (Hi-C) library was prepared from cross-linked chromatin using the standard Hi-C method.
[0043] Step 2, high-throughput genome sequencing. Illumina short-read sequencing technology was used to evaluate the genome size and determine the amount of data required for sequencing. Subsequently, the HiFi sequencing library described in Step 1 was sequenced on the Revio platform using single-molecule real-time (SMRT) sequencing technology to generate high-fidelity HiFi sequencing data. The Hi-C library was sequenced using the DNBSEQ platform, and the HiC-Pro pipeline was used to evaluate the quality of the Hi-C experiment.
[0044] To develop a suitable genome assembly and sequencing protocol, the present invention estimated the genome size of *Beckmannia syzigachne* to be approximately 2,820 megabase pairs (Mb) and the heterozygosity rate to be approximately 0.38% based on k-mer frequency analysis using 373G second-generation short-read data. To construct a high-quality genome of *Beckmannia syzigachne*, the present invention generated 209.7 Gb of PacBio-HiFi data (74.9× genome coverage) for genome assembly and 401 G of high-throughput chromosome conformation capture (Hi-C) data (143.2× genome coverage) for chromosome scaffolding and validation of the assembly results (Table 1).
[0045] Table 1 Genome libraries used for assembly
[0046] Library Total base number Total number of reads Coverage PacBio HiFi 209,707,775,031 11,736,567 ~74.89x NGS short reads 373,869,030,000 2,492,460,200 ~133.21x Hi-C 401,421,248,700 2,676,141,658 ~143.21x
[0047] Step 3, setting up and optimizing the genome assembly tool. Using the compHiFi tool developed by the present invention, the HiFi reads generated by the PacBio Revio instrument were initially assembled using the Hifiasm software (Cheng, H., et al., 2021) to generate highly contiguous genome sequences. Using paired-end Hi-C sequencing data, the assembled contigs were clustered, oriented, and joined into pseudochromosomes through the Juicer (Durand, N.C., et al., 2016) and 3D-DNA pipelines (Dudchenko, O., et al., 2017), and then potential assembly errors were manually corrected using Juicebox (Robinson, J.T., et al., 2018).
[0048] Based on the data described in Step 2, the present invention uses the developed compHiFi tool to assemble the Beckmannia syzigachne genome. First, Hifiasm was used for preliminary assembly with default parameters, obtaining 733 highly continuous contigs, totaling 2,825,601,699 bp, and the N50 length reaching 404.19 Mb. Subsequently, Hi-C data was used to scaffold the genome ( Figure 2 A in
[0049] Table 2 Characteristics of the 7 chromosomes of Beckmannia syzigachne
[0050] Chromosome Length GC content Chr1 499748366 44.23 Chr2 420064200 44.12 Chr3 414362802 44.23 Chr4 404191001 44.4 Chr5 373819403 44.17 Chr6 356945089 43.98 Chr7 317446781 44.02
[0051] Step 4, refine the assembled genome. Using the developed compHiFi tool, integrate the contigs obtained from various genome assembly tools (verkko, canu, flye, and nextdenovo), and use quartet to fill the gaps in each chromosome (Hu, J., et al., 2024; Kolmogorov, M., et al., 2019; Koren, S., et al., 2017; Lin, Y., et al., 2023; Rautiainen, M., et al., 2023), finally generating a reference-level high-quality genome.
[0052] During the Hi-C scaffolding process, there were 17 remaining gaps. To fill these gaps, the present invention used different de novo assembly tools, including verkko, nextdenovo, canu, flye, to assemble the HiFi sequencing data, and used the obtained contigs to fill the gaps. When nextdenovo, flye, and canu were assembling, the estimated genome size was 2.85 G. According to the computing resources, a total of 128 G of memory and 64 threads were used for the assembly calculation. By this method, the present invention filled 16 gaps, leaving only one highly heterozygous gap on chromosome 6 ( Figure 2in C), finally produced an almost complete T2T genome consisting of 7 chromosomes with a size of 2786.6 Mb ( Figure 2 in B).
[0053] In the assembled Beckmannia syzigachne genome, a total of 2460 Mb of repetitive sequences were identified, accounting for 88.35% of the total assembly. Except for a gap on chromosome 6, the remaining chromosomes of Beckmannia syzigachne have reached the T2T gapless level. This high-quality genome provides a basis for further exploring the telomere and centromere regions with complex repetitive DNA sequences ( Figure 3 ).
[0054] Telomere region. To identify the location of telomeres, the present invention searched for tandem repeats in the region of 300 kb at both ends of each chromosome and found the telomere repeat unit (TTTAGGG / CCCTAAA), which is widely present and highly conserved in the telomeres of higher organisms. Finally, a total of 14 telomeres were identified on 7 chromosomes. Among them, the 3'-end telomere of chromosome 5 is the shortest, with only 212 copies, and the 5'-end telomere of chromosome 3 is the longest, with 1833 copies (Table 3).
[0055] Centromere region. Previous studies have found that there are a large number of tandem repeats and retrotransposons in the centromere region (Lee, H.R., et al., 2005; Nagaki, K., et al., 2003). To clarify the location of the centromere, a genome-wide scan was performed on candidate tandem repeats with a size of 30 - 500 bp, and it was found that elements such as unit64bp and unit74bp showed specific distributions in each chromosome ( Figure 3 ). Multiple TEs were also co-localized with these tandem repeats, concentrated in the centromere region of each chromosome and rarely distributed in other positions ( Figure 3 ). Finally, the present invention determined the centromere region of each chromosome, with a length between 3.1 Mb and 6.3 Mb (Table 3).
[0056] Table 3 Telomere and centromere characteristics of the Beckmannia syzigachne genome
[0057] Chromosome Telomere length (bp) Telomere copy number Centromere position Chr1 6,601;8,449 943;1207 244,763,500-250,547,417 Chr2 7,364;5,103 1052;729 244,431,819-250,565,848 Chr3 12,831;9,429 1833;1347 225,357,678-228,770,461 Chr4 1,995;9,471 285;1353 166,362,932-172,672,932 Chr5 7,511;1,484 1073;212 177,462,037-181,955,379 Chr6 6,727;6,566 961;938 208,023,749-213,401,167 Chr7 2,751;4,984 393;712 124,402,193-127,539,812
[0058] Step 5, evaluate the quality of genome assembly. Through the BUSCO process ( F.A., et al. (2015) evaluated the completeness of the genome assembly and used the LTR Assembly Index (LAI) of long terminal repeat retrotransposons (LTR-RTs) as an indicator independent of the reference genome to evaluate the quality of the genome assembly. Further, through the Merqury tool (Ou, S. and N. Jiang, 2018) and its default settings, the LAI score of the assembled genome was determined. And the quality value (QV) of the genome was calculated using the Merqury tool (Rhie, A., et al., 2020) to ensure the high quality of the generated reference-level high-quality genome.
[0059] Finally, the completeness and quality of the Beckmannia syzigachne genome assembly were evaluated using BUSCO (completeness = 99.1%) and the LTR Assembly Index (LAI, 21.2) as well as Merqury (average QV = 57.35) ( Figure 2 D and E in). The lineage dataset used by BUSCO was embryophyta_odb10. When Merqury evaluated the quality of the genome assembly, k = 21 was set, 128G of memory and 64 threads were used. The results showed that the assembly quality of the Beckmannia syzigachne genome was extremely high (Table 4).
[0060] All the original sequence data of the Beckmannia syzigachne reference genome assembly described in the present invention are stored in the Genome Sequence Archive with the accession number PRJCA022644 (https: / / ngdc.cncb.ac.cn / gsa / s / hcBcvXA9). The code of the compHiFi tool can be obtained from GitHub (https: / / github.com / eric-hy / comphifi).
[0061] Table 4 Statistics of the Beckmannia syzigachne genome assembly results
[0062] Feature Value Number of chromosomes 7 Genome length 2,784,954,642 GC content 44.16% Mean LAI 21.2 BUSCO C:99.1% Proportion of repetitive sequences 88.35% Number of telomeres 14 Number of centromeres 7 Number of gaps 1 Number of protein-coding genes 48,711 Mean gene length 3424bp Average number of exons per gene 4.7
[0063] Example 2
[0064] To verify the universality and reliability of the assembly tool and the compHiFi process described in the present invention, whole-genome sequencing data of melon (Cucumis melo) and carrot (Daucus carota) that have been publicly released were further selected as the validation dataset. This dataset covers second-generation Illumina short-read sequencing, third-generation HiFi long-read sequencing, Hi-C chromatin conformation capture sequencing, and ONT nanopore sequencing data. The test results show that based only on the third-generation HiFi and Hi-C sequencing data, the genomic assembly results obtained using the assembly tool and the compHiFi process described in the present invention can identify the distribution of telomere repeat units in melon and carrot ( Figure 4 ), and compared with the chromosomal assembly results integrating ONT and second-generation sequencing data, it shows a high degree of consistency ( Figure 5 ). The above results confirm that the technical solution described in the present invention has broad species applicability and excellent assembly performance, and can provide reliable technical support for the reference-level high-quality genomic assembly of other species.
[0065] The present invention is not limited only to what is described in the specification and embodiments. Therefore, additional advantages and modifications can be easily achieved by those skilled in the art. Thus, without departing from the spirit and scope of the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details, representative solutions, and examples described herein.
Claims
1. A reference-level high-quality genome assembly method, characterized in that It includes the following steps: Use Illumina short-read sequencing technology to evaluate the genome size and determine the amount of data required for sequencing; Prepare HiFi sequencing libraries and Hi-C libraries; Sequence the HiFi sequencing library to obtain HiFi sequencing data; Sequence the Hi-C library to obtain paired-end Hi-C sequencing data; Through the compHiFi tool, preliminarily assemble the HiFi sequencing data using the Hifiasm software to obtain an initial highly contiguous genome assembly; Use the paired-end Hi-C sequencing data, cluster, orient, and connect the initial highly contiguous genome assembly into pseudochromosomes through the Juicer and 3D-DNA pipelines, and manually correct potential assembly errors in the pseudochromosomes using Juicebox to generate a high-quality genome assembly sequence; Through the compHiFi tool, use multiple de novo assembly tools to assemble the HiFi sequencing data, and the obtained contigs are used to fill the gaps in the high-quality genome assembly sequence, finally generating a reference-level high-quality genome.
2. The reference-level high-quality genome assembly method according to claim 1, wherein The preparation of the HiFi sequencing library and Hi-C library is specifically as follows: Extract genomic DNA using the FineOut Universal Plant DNA Kit, and use FEMTOPulse to evaluate the quality of the isolated DNA; Prepare the HiFi sequencing library according to the PacBio standardization process; Prepare the Hi-C library from cross-linked chromatin using the standard Hi-C method.
3. The reference-level high-quality genome assembly method according to claim 1, characterized in that The sequencing of the HiFi sequencing library uses the Revio platform and single-molecule real-time sequencing technology, and the sequencing of the Hi-C library uses the DNBSEQ platform.
4. The reference-level high-quality genome assembly method according to claim 1, wherein The Illumina short-read sequencing depth is at least 100×, the average read length of the HiFi sequencing library is at least 25 kb to span most transposable elements, the HiFi sequencing depth is at least 70×, and the paired-end sequencing of the Hi-C library is at least 140×.
5. The reference-level high-quality genome assembly method according to claim 1, wherein It also includes: After sequencing the Hi-C library, use the HiC-Pro pipeline to evaluate the quality of the paired-end Hi-C sequencing data.
6. The reference-level high-quality genome assembly method according to claim 1, wherein The multiple de novo assembly tools include verkko, canu, flye, and nextdenovo.
7. The reference-level high-quality genome assembly method according to claim 1, characterized in that, After generating the reference-level high-quality genome, it also includes evaluating the quality of the reference-level high-quality genome, specifically: Evaluate the integrity of the genome assembly through the BUSCO pipeline, and calculate the quality value of the reference-level high-quality genome using the Merqury tool.
8. The reference-level high-quality genome assembly method according to claim 7, wherein It also includes: Use the LAI as another indicator to evaluate the quality of the genome assembly, and determine the LAI score of the reference-level high-quality genome through the Merqury tool.
9. The reference-level high-quality genome assembly method according to any one of claims 1 to 8, characterized in that, The applicable biological individuals at least include: Beckmannia syzigachne, melon, carrot.
Citation Information
Cited By
Macrobrachium rosenbergii reference genome structure error correction and quality evaluation method and system
CN121811963A
Macrobrachium rosenbergii reference genome structure correction and quality assessment method and system
CN121811963B