Haplotype genome assembly method and device, and related applications
By using short-read data for error correction and optimizing long-read data during genome assembly, the problem of requiring diverse sequencing data for high-quality assembly was solved, achieving efficient and low-cost genome assembly.
Patent Information
- Application Number
- PCT/CN2024/096827
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2025-12-04
AI Technical Summary
High-quality genome assembly in existing technologies requires the use of three types of sequencing data, resulting in large amounts of data and high costs, and failing to fully utilize the advantages of nanopore ultra-long reads and next-generation sequencing.
By acquiring short-read and long-read data, the short-read data is used to correct errors in the long-read data. The genome is then assembled based on the corrected long-read data, and the short-read data is used to optimize the initial assembled sequence, reducing the coverage by 30 times to reduce the amount of data.
It achieved high-quality haplotype genome assembly, reduced the amount of data and cost required for genome assembly, and made full use of the advantages of short-read and long-read data.
Smart Images

Figure CN2024096827_04122025_PF_FP_ABST
Abstract
Description
Methods, apparatus and related applications for haplotype genome assembly Technical Field
[0001] This disclosure relates to the field of genome assembly, and more specifically, to a method, apparatus, and related applications for assembling a haplotype genome. Background Technology
[0002] Genome assembly is the process of constructing a complete genome sequence based on sequencing data, and it is a crucial step in genomics research. In recent years, advancements in sequencing technology have led to a rapid increase in long-read sequencing data. Mainstream long-read sequencing technologies (LRS technologies) are primarily divided into High Fidelity reads generated using circular consensus sequencing (CCS) and Nanopore sequencing, which can produce DNA fragment reads of thousands to tens of thousands of bp in length. New sequencing technologies have overcome the limitations of short-read sequencing (SRS, which typically produces DNA fragment reads of 100–300 bp in length), and long-read sequencing data has significantly improved the quality of genome assembly. The rapid development of third-generation sequencing technologies has opened up possibilities for hybrid long-short read assembly algorithms. These sophisticated methods combine the accuracy of short-read second-generation sequencing data (such as MGI rolling circle amplification or Illumina sequencing-by-synthesis) with the structural information provided by long-read sequencing data to generate high-quality and complete genome assembly results.
[0003] Hybrid assembly of second- and third-generation sequencing data is a method that combines DNA sequence data generated by different sequencing platforms to obtain more comprehensive and accurate biological information. Currently, most hybrid assembly algorithms rely primarily on HiFi data. The ultra-long read advantage of error-prone nanopore sequencing technology is mainly used for filling gap regions in the assembled structure and does not participate in assembly when sufficient HiFi data coverage is available. Furthermore, nanopore data is generally not utilized in the calibration process performed before assembly, which fails to fully leverage the enormous potential of ultra-long reads.
[0004] Hybrid assembly algorithms in related technologies all require Hi-Fi data and primarily rely on it for the main assembly process and post-assembly error correction. Nanopore sequencing technology only performs optional optimization modules after assembly, so the advantages of nanopore's ultra-long reads are not fully realized. Furthermore, next-generation sequencing (NGS) is also an additional option in existing methods, and its ultra-high accuracy (>99.99%) is not fully utilized. This means that current high-quality hybrid assembly results require high-throughput third-generation Hi-Fi sequencing data, nanopore data, and second-generation data. However, since third-generation sequencing is currently much more expensive than second-generation sequencing, achieving high-quality assembly requires significant sequencing costs.
[0005] There is currently no effective solution to the above problems.
[0006] Summary of the Invention
[0007] This disclosure provides a method, apparatus, and related applications for assembling a single genome, which at least solves the technical problem in related technologies that require three types of sequencing data to obtain high-quality genome assembly, resulting in a large amount of data usage.
[0008] According to one aspect of the present disclosure, a method for assembling a haplotype genome is provided, comprising: acquiring short-read data and long-read data obtained by sequencing the same biological sample to be tested; correcting errors in the long-read data based on the short-read data to obtain corrected long-read data; assembling the genome based on the corrected long-read data to obtain a preliminary assembled sequence; and optimizing the preliminary assembled sequence based on the short-read data to obtain a target assembled sequence.
[0009] According to another aspect of the present disclosure, a method for assembling the genome of a polyploid is also provided, wherein the genome assembly method for at least one haplotype of a polyploid is implemented.
[0010] According to another aspect of the present disclosure, a method for determining an unknown species type is also provided, comprising: acquiring short-read data and long-read data obtained by sequencing the same biological sample to be tested; correcting errors in the long-read data based on the short-read data to obtain corrected long-read data; assembling a genome based on the corrected long-read data to obtain a preliminary assembled sequence; optimizing the preliminary assembled sequence based on the short-read data to obtain a target assembled sequence; comparing the target assembled sequence with a known biological genome sequence, and determining the species type of the biological sample to be tested based on the comparison results.
[0011] According to another aspect of the embodiments of this disclosure, an assembly-based variation detection method is also provided, comprising: acquiring short-read data and long-read data obtained by sequencing the same biological sample to be tested; correcting errors in the long-read data based on the short-read data to obtain corrected long-read data; assembling the genome based on the corrected long-read data to obtain a preliminary assembled sequence; optimizing the preliminary assembled sequence based on the short-read data to obtain a target assembled sequence; and determining the gene variation result of the biological sample to be tested based on the target assembled sequence and a reference genome sequence.
[0012] According to another aspect of the embodiments of this disclosure, a monomorphic genome assembly apparatus is also provided, comprising: a first acquisition module for acquiring short-read data and long-read data obtained by sequencing the same biological sample to be tested; a first error correction module for correcting errors in the long-read data based on the short-read data to obtain corrected long-read data; a first assembly module for assembling the genome based on the corrected long-read data to obtain a preliminary assembled sequence; and a first optimization module for optimizing the preliminary assembled sequence based on the short-read data to obtain a target assembled sequence.
[0013] According to another aspect of the present disclosure, a polyploid genome assembly apparatus is also provided, comprising: a plurality of haplotype genome assembly apparatuses, wherein each haplotype genome assembly apparatus comprises: a second acquisition module for acquiring short-read data and long-read data obtained by sequencing the same biological sample to be tested; a second error correction module for correcting errors in the long-read data based on the short-read data to obtain corrected long-read data; a second assembly module for assembling the genome based on the corrected long-read data to obtain a preliminary assembled sequence; and a second optimization module for optimizing the preliminary assembled sequence based on the short-read data to obtain a target assembled sequence.
[0014] According to another aspect of the embodiments of this disclosure, an apparatus for determining the type of an unknown species is also provided, comprising: a third acquisition module for acquiring short-read data and long-read data obtained by sequencing the same biological sample to be tested; a third error correction module for correcting errors in the long-read data based on the short-read data to obtain corrected long-read data; a third assembly module for assembling a genome based on the corrected long-read data to obtain a preliminary assembled sequence; a third optimization module for optimizing the preliminary assembled sequence based on the short-read data to obtain a target assembled sequence; and a first determination module for comparing the target assembled sequence with a known biological genome sequence, and determining the species type of the biological sample to be tested based on the comparison results.
[0015] According to another aspect of the embodiments of this disclosure, an assembly-based variant detection device is also provided, comprising: a fourth acquisition module for acquiring short-read data and long-read data obtained by sequencing the same biological sample to be tested; a fourth error correction module for correcting errors in the long-read data based on the short-read data to obtain corrected long-read data; a fourth assembly module for assembling a genome based on the corrected long-read data to obtain a preliminary assembled sequence; a fourth optimization module for optimizing the preliminary assembled sequence based on the short-read data to obtain a target assembled sequence; and a second determination module for determining the gene variant result of the biological sample to be tested based on the target assembled sequence and a reference genome sequence.
[0016] According to another aspect of the present disclosure, an electronic device is also provided, including: a memory and a processor; the memory for storing program instructions; and the processor, connected to the memory, for executing the above-described haploid genome assembly method, or executing the above-described polyploid genome assembly method, or executing the above-described method for determining unknown species types, or executing the above-described assembly-based variation detection method.
[0017] According to another aspect of the present disclosure, a non-volatile storage medium is also provided, the non-volatile storage medium including a stored computer program, wherein the device on which the non-volatile storage medium is located executes the above-described haploid genome assembly method, or executes the above-described polyploid genome assembly method, or executes the above-described method for determining unknown species types, or executes the above-described assembly-based variation detection method by running the computer program.
[0018] According to another aspect of the present disclosure, a computer program product is also provided, including computer instructions that, when executed by a processor, implement the above-described haploid genome assembly method, or the above-described polyploid genome assembly method, or the above-described method for determining unknown species types, or the above-described assembly-based variation detection method.
[0019] In this embodiment, short-read and long-read data obtained from sequencing the same biological sample are acquired separately; the long-read data is corrected based on the short-read data to obtain corrected long-read data; the genome is assembled based on the corrected long-read data to obtain a preliminary assembled sequence; and the preliminary assembled sequence is optimized based on the short-read data to obtain the target assembled sequence. This achieves the goal of high-quality haplotype genome assembly based on both long-read and short-read data, thereby reducing the amount of data required for genome assembly. This solves the technical problem in related technologies where obtaining high-quality genome assembly requires the use of three types of sequencing data, resulting in a large amount of data usage. Attached Figure Description
[0020] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this disclosure, illustrate exemplary embodiments of the present disclosure and are used to explain the disclosure, but do not constitute an undue limitation of the disclosure. In the drawings:
[0021] Figure 1 is a hardware structure block diagram of a computer terminal for implementing a haplotype genome assembly method according to an embodiment of the present disclosure;
[0022] Figure 2 is a flowchart of a haplotype genome assembly method according to an embodiment of the present disclosure;
[0023] Figure 3 is a flowchart of a genome assembly of a monomer based on nanopore data according to an embodiment of the present disclosure;
[0024] Figure 4 is a graph showing the accuracy detection result of a nanopore after error correction using short read length, according to an embodiment of the present disclosure.
[0025] Figure 5 is a comparison of assembly results of different assembly methods under different sequencing throughput according to an embodiment of the present disclosure;
[0026] Figure 6 is a flowchart of a method for determining an unknown species type according to an embodiment of the present disclosure;
[0027] Figure 7 is a flowchart of an assembly-based variant detection method according to an embodiment of the present disclosure;
[0028] Figure 8 is a structural diagram of a monomeric genome assembly device according to an embodiment of the present disclosure;
[0029] Figure 9 is a structural diagram of an apparatus for determining an unknown species type according to an embodiment of the present disclosure;
[0030] Figure 10 is a structural diagram of an assembly-based variation detection device according to an embodiment of the present disclosure. Detailed Implementation
[0031] To enable those skilled in the art to better understand the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present disclosure.
[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0033] First, some nouns or terms that appear in the explanation of the embodiments of this disclosure shall be interpreted as follows:
[0034] De Bruijn graphs are directional diagrams illustrating the overlapping relationships between symbolic sequences and are the core foundation of tools for assembling genomes using short-read sequencing. The k-mer units constructed during assembly are based on the reads generated by short-read sequencing. That is, these reads are cut into smaller, overlapping fragments according to the length of the k-mer. k-mers can be 2-mer, 3-mer, 4-mer, or 5-mer. For example, a read with the sequence ACTGCTGAT can be cut into the following 2-mer fragments with one overlapping base: AC, CT, TG, GC, CT, TG, GA, and AT. Using 3-mer, it can be cut into the following 3-mer fragments with one overlapping base: ACT, CTG, TGC, GCT, CTG, TGA, and GAT. Using 4-mer, it can be cut into the following 4-mer fragments with one overlapping base: ACTG, CTGC, TGCT, GCTG, CTGA, and TGAT, and so on.
[0035] Compacted de Bruijn graphs are one of the most fundamental data structures in computational genomics.
[0036] Compacted and colored de Bruijn graphs are a variant based on a set of sequences, where each k-mer is associated with the sequence in which it appears.
[0037] Shasta assembly tool: Developed to improve the efficiency of genome assembly, Shasta performs significantly faster and more powerfully than widely used long-read sequencing tools like Canu. In most Shasta assembly stages, long-read sequences are stored in a compressed, aggregated form. In this storage format, identical base sequences are compressed and saved as single bases plus a repetition count. For example, GATTTACCA is saved as (GATACA, 113121). This avoids errors caused by the aggregation of identical bases, a common problem in Nanopore sequencing. This method also improves assembly accuracy. Simultaneously, each sequence is represented using a notation system. The MinHash algorithm is used to find the number of times the notation m appears in each pair of overlapping sequences. The calculation of the notation representation during sequence alignment is highly efficient. Shasta utilizes a large amount of memory on a single computing node for computation, and all data structures are stored in memory.
[0038] In related technologies, the method of assembling mixed short-read and long-read sequencing data is often used to improve the assembly quality and continuity of genomes or transcriptomes, especially when dealing with complex or highly repetitive genomes. The general steps of this process are as follows:
[0039] 1) Data generation: First, the same biological sample is sequenced using different sequencing platforms (e.g., MGI or Illumina for short read sequencing, and PacBio or Oxford Nanopore for long read sequencing) to generate corresponding sequencing data.
[0040] 2) Data quality control: Quality control is performed on the generated sequencing data, including removing low-quality sequences, removing adapter sequences, and filtering low-quality bases.
[0041] 3) Mixed Data Assembly: Assembly is performed using mixed data. This typically involves initial assembly using short-read sequencing data to obtain a basic framework with high coverage and accuracy. Then, long-read sequencing data is used to fill in repetitive regions, resolve difficult-to-assemble areas, or improve the continuity of the assembly.
[0042] 4) Error correction: Long-read sequencing data is used to correct errors that may have been introduced in short-read sequencing. This can be achieved by comparing the consistency between short-read and long-read data.
[0043] 5) Resolving repetitive regions: Long-read sequencing data are generally better suited for resolving highly repetitive genomic regions, and can therefore be used to resolve these difficult-to-assemble parts.
[0044] 6) Assembly evaluation: Evaluate the results of the hybrid assembly, including the evaluation of indicators such as N50, integrity and accuracy.
[0045] The following are some of the commonly used software and methods in the mixed assembly of short-read and long-read sequencing data in related technologies:
[0046] MaSuRCA (Maryland Super Read Cabog Assembler): This is a tool specifically designed for hybrid assembly, capable of efficiently integrating short-read and long-read sequencing data for assembly.
[0047] SPAdes: Can handle a variety of sequencing data types, including Illumina and PacBio / Nanopore data.
[0048] Unicycler: A tool specifically designed for microbial assembly, supporting mixed assembly and improving assembly continuity by combining long and short reads.
[0049] Racon is a tool for correcting short-read sequencing data, often used in conjunction with Canu (an assembly tool for PacBio / Nanopore data) to improve assembly accuracy.
[0050] HybridSPAdes: This is a module of SPAdes specifically designed for hybrid assembly, supporting the integration of Illumina and PacBio / Nanopore data.
[0051] Flye is a tool specifically designed for assembling long reads (PacBio / Nanopore) and can be used with short read sequencing data to improve assembly quality.
[0052] While these hybrid methods are very effective, they typically require long-read sequencing data with very high coverage or depth (e.g., 50X), which incurs significant sequencing costs.
[0053] Furthermore, hybrid assembly tools in related technologies utilize both conventional short-read sequencing (PCR, PCR-Free sequencing) and Hi-C (High-throughput chromosome conformation capture) for genome genotyping, enabling the acquisition of almost complete diploid genome assemblies. However, this typically requires very high data depths, generally recommending sequencing depths exceeding 50-fold as input. This means that current genome assembly tool parameters are not suitable for low-sequencing-depth data.
[0054] Verkko is a state-of-the-art assembly software applicable to telomere-to-telomere T2T telomere assembly of diploid genomes. It mixes HiFi and nanopore data, and uses Canu's error correction module (a long read assembly and error correction tool based on overlap graphs, which uses PacBio's long reads and Illumina's short reads to generate high-quality genome sequences through assembly and error correction) to correct HiFi data and build a multiplex de Bruijn graph. Then, it aligns the nanopore data onto the graph, gradually resolving cyclic and tangled regions, and finally uses Canu's consensus module to obtain the final result.
[0055] Hifiasm is an algorithm that effectively utilizes HiFi sequencing technology to reliably represent haplotype information in a phased assembly graph. Hifiasm analysis consists of three parts: the first stage corrects potential sequencing errors by comparing all sequences; the second stage constructs a phased string graph based on the overlap between sequences; and the third stage selects one side of the bubble to construct the primary assembly and the other side to construct the alternative assembly, thus forming haplotype assemblies.
[0056] Existing high-quality hybrid assembly results require high-depth long-read Hi-Fi sequencing data, nanopore data, and short-read data. However, since long-read sequencing technology is currently extremely expensive compared to short-read sequencing, completing high-quality assembly requires a significant amount of sequencing costs.
[0057] The primary approach to reducing the ultimate cost of genome assembly is to use lower coverage or sequencing depth. However, most assembly tools are still in the research phase, focusing on optimizing runtime or improving performance across various quality metrics, and therefore require high coverage inputs, often exceeding 50-fold. Conversely, assembly methods with coverage below 30-fold are rare.
[0058] To address the aforementioned issues, this disclosure provides a haploid genome assembly method. Based on the haploid genome assembly method, it also proposes a polyploid genome assembly method, a method for determining unknown species types, and an assembly-based variation detection method. The haploid genome assembly method, the polyploid genome assembly method, the method for determining unknown species types, and the assembly-based variation detection method can all be run on the computer terminal shown in Figure 1. The following description uses the haploid genome assembly method as an example to illustrate the computer terminal.
[0059] The haplotype genome assembly method embodiments provided in this disclosure can be executed in a mobile terminal, computer terminal, or similar computing device. Figure 1 shows a hardware structure block diagram of a computer terminal for implementing the haplotype genome assembly method. As shown in Figure 1, the computer terminal 10 may include one or more processors (shown as 102a, 102b, ..., 102n in the figure) (the processor may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission module 106 for communication functions connected via wired and / or wireless networks. In addition, it may also include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. Those skilled in the art will understand that the structure shown in Figure 1 is only illustrative and does not limit the structure of the above-described electronic device. For example, the computer terminal 10 may also include more or fewer components than shown in Figure 1, or have a different configuration than shown in Figure 1.
[0060] It should be noted that the aforementioned one or more processors and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10. As per the embodiments of this disclosure, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0061] The memory 104 can be used to store software programs and modules for application software, such as the program instructions / data storage device corresponding to the haplotype genome assembly method in this embodiment of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the haplotype genome assembly method described above. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0062] The transmission module 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission module 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission module 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0063] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10.
[0064] It should be noted that, in some alternative embodiments, the computer terminal shown in FIG1 may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be pointed out that FIG1 is merely one example of a specific instance and is intended to illustrate the types of components that may exist in the aforementioned computer terminal.
[0065] In the above operating environment, this disclosure provides an embodiment of a method for assembling a single genome. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Also, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0066] Figure 2 is a flowchart of a haplotype genome assembly method according to an embodiment of the present disclosure. As shown in Figure 2, the method includes the following steps:
[0067] Step S202: Obtain short read data and long read data obtained from sequencing the same biological sample to be tested.
[0068] In step S202 above, short reads refer to the relatively short length of the nucleotide sequence obtained in a single sequencing operation. Typically, these sequences are at most 500 base pairs (bp) long, while common sequencing fragment lengths are between 100 and 300 bp. Long reads refer to the relatively long length of the nucleotide sequence obtained in a single sequencing operation. These sequences are usually longer than 1000 base pairs, and sometimes even reach hundreds of thousands of base pairs. Using both short and long read sequencing technologies simultaneously, to fully utilize the advantages of both technologies, can yield more comprehensive and accurate bioinformatics analysis results.
[0069] It should be noted that, in this embodiment, long read data refers to data greater than 1000 bp, and short read data refers to data less than 1000 bp. In an optional embodiment, long read data may be, for example, nanopore data.
[0070] Step S204: Correct the long read length data based on the short read length data to obtain the corrected long read length data.
[0071] In step S204 above, short-read data generally has high accuracy, but each sequence has a short read length, meaning it can cover multiple regions of the genome, but may not be able to read large genes or genomic regions at once. Long-read data has longer read lengths, capable of covering large segments of the genome. However, due to its longer read length, its accuracy may be lower, especially at the ends of sequences or in complex regions. Because short-read data has high accuracy and deeper sequencing depth, it can be used to verify the accuracy of long-read data.
[0072] Step S206: Based on the error-corrected long read data, the genome is assembled to obtain the preliminary assembled sequence.
[0073] In step S206 above, genome assembly is a process of splicing short fragments obtained from sequencing (here, error-corrected long read data) into a complete genome sequence. This process is typically implemented, for example, using genome assembly software to identify overlapping portions between different fragments and splice them together based on these overlaps. After processing by the genome assembly software, the error-corrected long read data will be spliced into one or more consecutive sequences, which are the preliminary assembled sequences described above.
[0074] Step S208: Optimize the preliminary assembly sequence based on the short read length data to obtain the target assembly sequence.
[0075] In step S208 above, although the preliminary assembled sequence can cover most regions of the genome, due to the inherent limitations of long-read data (such as sequencing errors and read length restrictions), some errors, gaps, or repetitive sequences may exist in the preliminary assembled sequence. Short-read data is usually generated by high-throughput sequencing technologies (such as MGI's DNBseq and Illumina sequencing), which have the characteristics of high sequencing depth, high accuracy, and relatively short read lengths. These characteristics give short-read data a unique advantage in the post-genome assembly optimization stage. Using short-read data to further optimize the preliminary assembled sequence, such as error correction, gap filling, and repetitive sequence handling, can further improve the accuracy and integrity of the genome sequence.
[0076] Through steps S202 to S208 of the haplotype genome assembly method described above, the goal of achieving high-quality haplotype genome assembly based on both long-read and short-read data is achieved. This reduces the amount of data required for genome assembly and solves the technical problem of requiring three types of sequencing data and consuming large amounts of data to obtain high-quality genome assembly in related technologies. The following explanation follows.
[0077] In the above-mentioned haplotype genome assembly methods, the coverage of short read data and long read data is less than or equal to 30-fold.
[0078] In some embodiments of this disclosure, coverage (or depth of coverage) is an important metric, representing the average number of times the sequenced reads cover the genome. Specifically, coverage refers to the average number of times all sequenced reads are superimposed at all positions on the genome. Related technologies and genome assembly tools often require coverage exceeding 50-fold, resulting in high costs. This disclosure uses short-read and long-read data with a coverage of less than or equal to 30-fold for genome assembly, thus reducing costs.
[0079] In step S204 of the haplotype genome assembly method described above, long read data is corrected based on short read data to obtain corrected long read data. This includes: determining a first length and a second length, wherein the first length is shorter than the second length; constructing a first de Bruin graph corresponding to the division of short read data according to the first length, and constructing a second de Bruin graph corresponding to the division of short read data according to the second length; anchoring the long read data to the first de Bruin graph to obtain first erroneous data in the long read data, wherein the first erroneous data is data that is inconsistent with the long read data in the first de Bruin graph; correcting the first erroneous data to obtain an intermediate result; anchoring the intermediate result to the second de Bruin graph to obtain second erroneous data in the intermediate result, wherein the second erroneous data is data that is inconsistent with the intermediate result in the second de Bruin graph; and correcting the second erroneous data to obtain corrected long read data.
[0080] In some embodiments of this disclosure, the first length and the second length are predetermined. The first sequence corresponding to the first length and the second sequence corresponding to the second length can be understood as two k-mers (sequence fragments composed of k consecutive bases) of different lengths. The de Bruin diagram is constructed using k-mers of different lengths so as to detect and correct errors in long read data at different granularities.
[0081] The first de Brouin graph (hereinafter referred to as G1) is constructed by partitioning short read data according to a first length (k1-mer). In G1, each node represents a k1-mer, and edges indicate overlap between two k1-mers. G1 can help identify frequent and consistent k1-mer sequences in short read data. The second de Brouin graph (hereinafter referred to as G2) is constructed by partitioning short read data according to a second length (k2-mer, where k2>k1). Similar to G1, G2 provides contextual information for longer sequences, which helps identify more complex error patterns.
[0082] The long-read data is compared with G1 (i.e., the anchoring process described above). This process identifies the parts of the long-read data that are inconsistent with G1, i.e., the first erroneous data. These inconsistencies may be caused by sequencing errors, assembly errors, or other reasons. Based on the information in G1, the first erroneous data is corrected to obtain intermediate results. This correction process typically involves operations such as substitution, deletion, or insertion of bases to make the long-read data more consistent with the pattern of short-read data. The corrected intermediate results (i.e., intermediate results) are compared with G2. This process can discover more complex errors that may exist in the intermediate results, i.e., second erroneous data. Based on the information in G2, the second erroneous data is corrected to obtain the final error-corrected long-read data. By constructing de Bruin diagrams of different granularities to detect and correct errors in the long-read data, the overall quality of the data is improved, resulting in improved accuracy of the error-corrected long-read data.
[0083] The anchoring process described above is explained using the example of "comparing long read data with G1":
[0084] 1) Divide the long read data into a series of k-mer segments using the same k-mer length as the first de Bruin graph (i.e., the first length);
[0085] 2) For each k-mer segment, search in G1 to find if a node with an exact match exists. For example, a hash table or similar data structure can be used for fast searching.
[0086] 3) If a k-mer segment finds a matching node in G1, then the k-mer segment is anchored to that node, meaning the position of the k-mer segment in the long read data corresponds to its position in G1; if a k-mer segment does not find a completely matching node in G1, but finds a partially matching node (i.e., there is partial overlap), then some extended search or fuzzy matching may be needed to determine the optimal anchoring position; if a k-mer segment does not find any matching nodes in G1, then it may be a new sequence segment that has not appeared in the short read data, or an incorrect sequence segment.
[0087] 4) During the anchoring process, record all k-mer segments that are inconsistent with G1, which is the first error data mentioned above.
[0088] In some embodiments of this disclosure, the process of correcting long-read data based on short-read data to obtain corrected long-read data can be implemented using the Ratatosk tool, which is based on a compact, colored de Bruijn graph for mixed correction of genomic long-read sequences using short-read data. Ratatosk is specifically designed to avoid overcorrection to prevent the use of incorrect haplotypes or homologous regions, as doing so may remove real variations or add artificial variations. Ratatosk introduces several new features not included in other mixed correction tools. First, a de Bruijn graph is constructed using short-read data, with the short-read and long-read data coloring the vertices of the de Bruijn graph to highlight existing paths for correction. Graph coloring can prune the search space when traversing the graph by removing chimeric paths. Second, long-read data is anchored to the graph using precise and inprecise k-mer (such as k1-mer and k2-mer as described above). Using inprecise k-mer (such as k2-mer) can improve the anchoring of highly erroneous regions in the long-read data. Third, candidate SNPs are labeled in the anchored de Bruijn plot to distinguish small variations between haplotypes that are difficult to capture from erroneous long-read data. Fourth, two corrections are performed using both short-read and long-read data to utilize all available data, and the k-mer size is increased to eliminate errors generated during the first correction. Finally, optional reference-guided preprocessing of the input data is recommended to improve the error rate and scale Ratatosk to a large number of computational nodes.
[0089] In the above steps, the intermediate result is obtained by correcting the first erroneous data, including: determining the first region of the first erroneous data in the short read data, and determining the second region of the first erroneous data in the long read data; replacing the data contained in the second region with the data contained in the first region to obtain the intermediate result.
[0090] In some embodiments of this disclosure, short-read data typically has high accuracy and can therefore be used as a reference for correcting errors in long-read data. Through a comparison and anchoring process, the corresponding position or region of the first erroneous data in the short-read data can be determined; this region is the aforementioned first region. Locating the portion in the long-read data corresponding to the first erroneous data yields the aforementioned second region. After determining the first and second regions, accurate data from the short-read data (first region) can be used to replace the erroneous data in the long-read data (second region), thereby improving the accuracy of the long-read data.
[0091] In the above steps, correcting the second erroneous data to obtain the corrected long read data includes: determining the third region of the second erroneous data in the short read data, and determining the fourth region of the second erroneous data in the intermediate result; replacing the data contained in the fourth region with the data contained in the third region to obtain the corrected long read data.
[0092] In some embodiments of this disclosure, through a comparison and anchoring process, errors that were not identified during the correction of the first erroneous data, or newly discovered errors in the long read data, can be detected, thereby obtaining the aforementioned second erroneous data. By determining the corresponding position or region of the second erroneous data in the short read data, this region is the aforementioned third region; and by locating the portion in the intermediate result corresponding to the second erroneous data, the aforementioned fourth region is obtained. After determining the third and fourth regions, the erroneous data in the intermediate result (fourth region) can be replaced with accurate data from the short read data (third region), thereby obtaining a more accurate long read data, namely the aforementioned corrected long read data.
[0093] In step S206 of the above-mentioned haplotype genome assembly method, genome assembly is performed based on the error-corrected long read data to obtain a preliminary assembled sequence, including: determining multiple subsequences corresponding to the error-corrected long read data, wherein the sequence length of each subsequence can be the same; determining the similarity between each subsequence and other subsequences to obtain a similarity set corresponding to each subsequence; determining the subsequence corresponding to the maximum similarity in the similarity set as a ligandable sequence; and performing genome assembly on the error-corrected long read data based on the ligandable sequence to obtain a preliminary assembled sequence.
[0094] In some embodiments of this disclosure, the MinHash algorithm can be used to estimate the similarity between two sets (i.e., any two subsequences) to determine the initial assembly sequence. Specifically, this includes the following process: dividing the error-corrected long-read data into multiple subsequences (or fragments or k-mers). For ease of subsequent alignment and assembly operations, these subsequences can have the same length. The similarity between each pair of these subsequences is calculated. Similarity calculation is typically based on sequence alignment algorithms, such as the Smith-Waterman algorithm or BLAST, which can assess the degree of matching between two sequences. For each subsequence, its similarity to all other subsequences is calculated, and these similarity values are grouped into a similarity set. Within each subsequence's similarity set, the value with the highest similarity is found. This maximum similarity value indicates that the subsequence has a high sequence match with another subsequence, and therefore they are likely adjacent parts of the genome sequence. The subsequences corresponding to these maximum similarity values are marked as ligatable sequences, and the assembly of the error-corrected long-read data begins. The assembly process involves linking the ligable sequences according to their relative positions in the genome (i.e., their order and relative distance in long-read data) to obtain a preliminary assembled sequence.
[0095] In some embodiments of this disclosure, genome assembly is performed based on the error-corrected long-read data to obtain preliminary assembled sequences, which can be achieved using the Shasta tool. The error-corrected long-read data is stored in a compressed form to reduce errors caused by the aggregation of identical bases. The MinHash algorithm is used to find the number of times the marker m appears in each pair of overlapping sequences and to achieve preliminary haplotype assembly.
[0096] In most Shasta assembly stages, read sequences are stored in a homopolymer compressed form using a length-encoded method. In this form, identical consecutive bases are merged, and base and repeat counts are stored. For example, GATTTACCA would be represented as (GATACA, 113121). This representation is insensitive to homopolymer length errors, thus addressing the main error pattern of long-read data. Therefore, higher similarity alignments can be achieved due to reduced assembly noise caused by read errors. A number of labeled representations are also used, where each read sequence is represented by the sequence of occurrences of a predetermined, fixed set of short k-mers (label representations) in its run-length representation. A modified MinHash scheme is used to find candidate overlapping sequence pairs, using the m consecutively occurring labels (default m=4) as MinHash features to compute the best alignment under the label representation for all candidate data. The alignment results are used for initial assembly of the error-corrected long-read data. A label graph can also be created for assembling sequences after a series of simplification steps, where each vertex in the graph represents a label found in a set of multiple segments.
[0097] It's worth noting that similarity calculations between subsequences can also be achieved using algorithms such as LSH (Locality Sensitive Hashing), SimHash, KShingle, KSentence, and Jaccard Similarity. LSH is a hashing method that maps data to multiple hash tables, increasing the probability that similar data items will be mapped to the same hash table. SimHash is another algorithm for detecting text similarity; it determines text similarity by converting text into feature vectors and then calculating the Hamming distance between these vectors. KShingle is also a text deduplication algorithm that compares text similarity by extracting phrases (shingles) from the text and converting them into feature vectors. KSentence is based on the assumption that the K longest sentences in two repeated texts should be identical; it deduplicates text by extracting the longest sentences and calculating their MD5 values as text fingerprints. Jaccard similarity is the basis for the MinHash algorithm to estimate set similarity; MinHash compares the similarity of two sets by estimating Jaccard similarity. These algorithms have wide applications in fields such as text data processing, image recognition, and bioinformatics. Each algorithm has its specific application scenarios and advantages and disadvantages. The choice of which algorithm to use usually depends on the requirements of the specific problem and the characteristics of the data, and no restrictions are imposed here.
[0098] In step S208 of the haplotype genome assembly method described above, the preliminary assembled sequence is optimized based on short read data to obtain the target assembled sequence. This includes: using a recurrent neural network model to determine the first variant site that is inconsistent with the short read data, wherein the recurrent neural network model is used to determine whether the preliminary assembled sequence has mutated, and the first variant site is an erroneous data point with a variant site length of less than 50 bp; comparing the short read data and the preliminary assembled sequence to obtain the second and third variant sites in the preliminary assembled sequence, wherein the second variant site is an erroneous data point with a variant site length of less than 50 bp, and the third variant site is an erroneous data point with a variant site length greater than or equal to 50 bp; and optimizing the data of the first, second, and third variant sites based on the short read data to obtain the target assembled sequence.
[0099] In the above steps, the recurrent neural network model is used to determine the first variant site that is inconsistent with the short read data in the preliminary assembled sequence, including: using the recurrent neural network model to determine the single nucleotide polymorphism (SNP) site that is inconsistent with the short read data in the preliminary assembled sequence; determining the alignment file with haplotype markers based on the SNP sites, wherein the haplotype markers are used to indicate the parental origin of the SNP sites; and determining the first variant site in the preliminary assembled sequence based on the alignment file with haplotype markers.
[0100] In the above steps, the data of the first variant site, the second variant site, and the third variant site are optimized based on the short read data to obtain the target assembly sequence. This includes: determining the first position of the first variant site in the short read data and determining the second position of the first variant site in the preliminary assembly sequence; determining the third position of the second variant site in the short read data and determining the fourth position of the second variant site in the preliminary assembly sequence; determining the fifth position of the third variant site in the short read data and determining the sixth position of the third variant site in the preliminary assembly sequence; replacing the data corresponding to the second position with the data corresponding to the first position, replacing the data corresponding to the fourth position with the data corresponding to the third position, and replacing the data corresponding to the sixth position with the data corresponding to the fifth position to obtain the target assembly sequence.
[0101] In some embodiments of this disclosure, the mutation sites can be determined in two ways. The first way is to use a recurrent neural network model to determine the first mutation site, where the mutation length of the first mutation site is less than 50 base pairs (bp), for example, it can be implemented using PEPPER. The second way is not to use a neural network model, but to directly determine the second and third mutation sites by comparing short read data and preliminary assembled sequences. The mutation length of the second mutation site is less than 50 bp, and the mutation length of the third mutation site is greater than or equal to 50 bp, for example, it can be implemented using Pilon. By optimizing the mutation sites determined by the two methods, the target assembled sequence can be obtained.
[0102] In the first approach, the steps for determining the first variant site include:
[0103] 1) Use recurrent neural networks to identify single nucleotide polymorphism (SNP) sites that are inconsistent with the short read data in the preliminary assembled sequence. These SNP sites are potential sequencing errors in the long read data or real individual mutations. This step can be achieved, for example, by using PEPPER-SNP.
[0104] 2) Determine alignment files with haplotype markers based on SNP sites. This step can be achieved, for example, by using Margin. Margin uses the SNPs reported by PEPPER-SNP and generates alignment files with haplotype markers using a Hidden Markov Model (HMM).
[0105] 3) Determine the first variant site in the preliminary assembled sequence based on the alignment file with haplotype markers. For example, this can be achieved through PEPPER-HP. PEPPER-HP uses alignment files with haplotype markers to independently evaluate each haplotype and generate haplotype-specific candidate SNPs and INDEL pattern errors present in the assembly, which are the aforementioned first variant sites. Here, SNP refers to a single base error, and INDEL refers to an insertion or deletion of less than 50 bp in length.
[0106] In the second approach, for example using Pilon, short-read data and preliminary assembled sequences are scanned to obtain alignment results. Second variant sites with a variant length less than 50 bp are identified, including type errors of the aforementioned SNPs and INDELs. Then, by checking coverage and alignment differences, potential assembly errors and larger variants (i.e., inverted positions of certain segments or unreasonable segment repetition counts in the assembly result, referred to here as Type II errors, i.e., structural variations of long segments) are identified, thus obtaining third variant sites with a variant length greater than or equal to 50 bp. By determining the first, third, and fifth positions of the aforementioned first, second, and third variant sites in the short-read data, and the second, fourth, and sixth positions of the aforementioned first, second, and third variant sites in the preliminary assembled sequence, the data corresponding to the first position is used to replace the data corresponding to the second position, the data corresponding to the third position is used to replace the data corresponding to the fourth position, and the data corresponding to the fifth position is used to replace the data corresponding to the sixth position, thereby optimizing the data and obtaining the target assembled sequence. This target assembly sequence combines the advantages of preliminary assembly sequences and short read data, removing most of the errors and more closely resembling the actual genome sequence.
[0107] In this embodiment, by utilizing the PEPPER-Pilon combination strategy and employing high-precision short-read data for single-base correction, assembly result correction, and blank region filling, a more accurate genome assembly sequence can be obtained. It should be noted that the processes for determining the first variant site and the second variant site are performed simultaneously, without any specific order.
[0108] The haplotype genome assembly method provided in this disclosure, based on short-read and long-read data, corrects long-read data (such as nanopore sequencing data) using high-precision short-read data before assembly, increasing the accuracy of long-read data from 90% to over 99%. Then, the high-accuracy, error-corrected long-read data is used for preliminary haplotype genome assembly. After preliminary assembly, the high accuracy advantage of the short-read data is leveraged to iteratively optimize the preliminary assembled sequence, including using recurrent neural networks, coverage, and alignment difference anchoring to identify discrepancies, correcting erroneous sequences in the preliminary assembled sequence, and further improving the accuracy of the assembly results. This achieves the assembly effect of traditional methods using short-read, nanopore, and Hi-Fi data. The above-mentioned haplotype genome assembly method effectively reduces the coverage requirement of sequencing data for high-quality assembly results, requiring only short-read and long-read sequencing data with a coverage of less than or equal to 30 times, thereby maximizing the application potential of high-precision short-read and long-read data and reducing the cost of assembling high-quality sequencing data.
[0109] In this embodiment of the disclosure, when the long-read data used is nanopore data, the flowchart of genome assembly for the corresponding haplotype is shown in Figure 3. In Figure 3, genome assembly is achieved through three steps, including:
[0110] Step 1: Error correction before assembly. The raw nanopore data is corrected using the Ratatosk tool and a short read tool to obtain the corrected nanopore data.
[0111] Step 2: Preliminary Assembly. The corrected nanopore data is preliminarily assembled using the Shasta tool to obtain the preliminary assembly result (i.e., the preliminary assembly sequence mentioned above);
[0112] Step 3: Post-assembly optimization. The initial assembly results are optimized using short read data and the PEPPER and Pilon strategies to obtain the final assembly result (i.e., the target assembly sequence mentioned above). The assembly result can be evaluated using the Quast and Merqury tools.
[0113] The beneficial effects of this disclosure will be further explained in detail below with reference to specific embodiments. It should be noted that the sources of the sequencing data used for testing in the following embodiments are as follows:
[0114] MGI data:
[0115] Source 1:
[0116] https: / / ftp-trace.ncbi.nlm.nih.gov / ReferenceSamples / giab / data / AshkenazimTrio / HG002_NA24385_son / MGISEQ / PCR-free / NA24385 /
[0117] Source 2:
[0118] CNSA: https: / / db.cngb.org / cnsa; accession number CNP0004858
[0119] Illumina data:
[0120] Source 2:
[0121] https: / / ftp-trace.ncbi.nlm.nih.gov / giab / ftp / data / AshkenazimTrio / HG002_NA24385_son / NIST_Illumina_2x250bps / reads /
[0122] Source 3:
[0123] https: / / trace.ncbi.nlm.nih.gov / Traces / ? view=run_browser&acc=SRR22476789&displa y=metadata
[0124] HiFi data:
[0125] Source 4:
[0126] https: / / s3-us-west-2.amazonaws.com / human-pangenomics / NHGRI_UCSC_panel / HG002 / hpp_HG002_NA24385_son_v1 / PacBio_HiFi / downsampled /
[0127] Source 5:
[0128] https: / / s3-us-west-2.amazonaws.com / human-pangenomics / NHGRI_UCSC_panel / HG002 / hpp_HG002_NA24385_son_v1 / PacBio_HiFi / 15kb / m64015_190920_185703.Q20.fastq
[0129] Source 6:
[0130] https: / / s3-us-west-2.amazonaws.com / human-pangenomics / NHGRI_UCSC_panel / HG002 / hpp_HG002_NA24385_son_v1 / PacBio_HiFi / 15kb / m64015_190922_010918.Q20.fastq
[0131] Nanopore data:
[0132] Source 7:
[0133] https: / / s3-us-west-2.amazonaws.com / human-pangenomics / NHGRI_UCSC_panel / HG002 / hpp_HG002_NA24385_son_v1 / nanopore / downsampled / greater_than_100kb / HG002_uc sc_ONT_lt100kb.fastq.gz
[0134] Example 1
[0135] Step 1: Short read length error correction for nanopore data
[0136] In hybrid assembly, one application of short-read sequencing data is error rate correction of error-prone ultra-long nanopore data. Pre-assembly correction and post-assembly correction are two common methods. Based on test results, the Ratatosk tool was selected for correction, which uses a color-compressed de Bruijn map algorithm based on accurate short-read sequences. In this embodiment, two MGI samples (sources 1 and 2 above) and two Illumina samples (sources 2 and 3 above) were used to correct nanopore sequencing data (source 7 above). After correction, the accuracy of the corrected nanopore sequencing data was calculated. Correcting the read sequences in the nanopore sequencing data using MGI or Illumina sequencing data yielded a similar accuracy distribution (see Figure 4), demonstrating an accuracy of 99%.
[0137] Step 2: Hybrid assembly of short read length and nanopore data
[0138] Different assembly tools were tested on human samples from the Ashkenazim Trio human standard sample. To reduce computational load, chromosome 20 was extracted first and tested on it. Generally, genome assembly tools recommend a coverage of 50x or better. However, these studies did not publish performance data at low coverage. Therefore, in this embodiment of the disclosure, the input sequencing data was adjusted to different average genome coverages, ranging from 10x to 50x, with intervals of 5 for testing assembly results with different coverages.
[0139] The latest hybrid assembly tools Verkko and Hifiasm were compared with Shasta and Wtdbg2 based on nanopore sequencing data. The comparison results are shown in Figure 5.
[0140] As shown in Figure 5 (AD), a higher metric (NG50) in A is better, a lower metric (number of contigs) in B is better, a higher metric (completeness based on K-mer) in C is better, and a higher metric (quality value, QV) in D is better. Shasta and Shasta+Ratatosk are both ONT-based methods (i.e., ONT+WGS) (the scheme proposed in this disclosure), while Hifiasm and Verkko are Hifi-based methods (i.e., Hifi+ONT+WGS). wtdbg2 supports two input types, so ONT or Hifi is indicated in the figure for distinction.
[0141] Because nanopore sequencing data is less accurate than HiFi or short-read sequencing data, and only the latest R10 nanopore libraries (R10 represents the latest chip number for ONT nanopore sequencing technology; compared to the older R9 version, R10 achieves a significant accuracy improvement from 90% to 99%, but its use is still limited due to its cost being three times that of R9 and the sequencing length decreasing from >100kb in R9 to ~10kb) can reach an average quality value of 30, one way to improve the accuracy of nanopore sequencing data is to perform error correction before assembly. There are generally two approaches to error correction. The first is to use the nanopore's own sequencing data for error correction. However, due to the high error rate (accuracy only 90%), nanopore sequencing data is not suitable for self-correction. The second approach is to use correction tools such as Ratatosk, which can use short-read sequencing data to correct nanopore data.
[0142] As shown in Figure 5, Shasta after Ratatosk correction exhibits a significant performance improvement (i.e., the display results of Shasta+Ratatosk compared to Shasta alone), and its performance is close to that of hybrid assembly tools (i.e., Verkko and Hifiasm). Furthermore, Figure 5 shows that some assembly tools exhibit a rapid decline in results when coverage is below 30. In particular, Verkko shows a sharp drop in genome assembly quality when coverage is below 25, while Wtdbg2 shows unexpectedly uniform performance at low coverage (as seen in Figure 5C, Shasta and Verkko show a significant decrease when coverage on the horizontal axis is less than 25 (higher is better), but Wtdbg2 and Shasta+Ratatosk results are consistent). Additionally, Figure 5A shows that the Wtdbg2 ONT results are poor, failing to reach the level of Shasta, suggesting that Wtdbg2+Ratatosk likely does not reach the level of Shasta+Ratatosk). Therefore, Shasta+Ratatosk can even outperform methods based on short read length + HiFi + nanopores at low coverage.
[0143] As mentioned above, nanopore data has high potential for assembly applications due to its long sequencing length (which can exceed 100kb). However, in existing assembly methods that only utilize nanopore sequencing data, although some methods attempt to use short reads for error correction, the assembly results after error correction still differ significantly from those achieved using PacBio Hifi sequencing data. The method in this embodiment, however, utilizes only nanopore sequencing data and short-read DNBSEQ data, maximizing the technical advantages of both throughout the entire assembly process—before, during, and after assembly—and achieving assembly results almost equivalent to those obtained with PacBio Hifi data.
[0144] Furthermore, when investigating the possibility of using short-read sequencing data + nanopore sequencing data as an alternative to hybrid short-read + HiFi + nanopore assembly, the two most commonly used short-read technologies, MGI and Illumina, have not yet been clearly compared. Therefore, to determine which short-read technology performs better in hybrid short-read + nanopore assembly, MGI and Illumina were compared as short-read data for haplotype identification, pre-correction assembly, and QualityValue (QV) score evaluation as part of the Meryl database. Two samples were tested on parental sample data (HG003, HG004), and three samples were tested on daughter sample data HG002. The results showed that MGI exhibited similar or slightly better quality control statistics (see Table 1 below).
[0145] Table 1. Comparison of results between MGI and Illumina in hybrid assembly
[0146] Quality assessment is a crucial step in genome assembly. One metric used for this purpose is the Quality Value (QV) score and k-mer-based integrity assessment. However, the results above suggest that relying solely on sequences from the same sample for QV assessment may not provide comprehensive information and can produce varying results depending on the input data.
[0147] To further investigate this, a Meryl database for QV evaluation was constructed using combinations of read sequences from different sequencing technologies. Significant differences were observed in QV evaluation using the HiFi, nanopore, or short-read Meryl databases in the HiFiasm assembly. Specifically, using HiFi as the k-mer Meryl database resulted in approximately twice the QV score as when using nanopores. In contrast, MGI and Illumina showed similar results in this case (see Table 2). Furthermore, combinations of different sequencing technologies only led to higher QV scores.
[0148] Table 2. Comparison of QV values using different data as the database
[0149] Step 3: Optimize the assembly results using short read length data.
[0150] To further submit preliminary assembly results, based on the test results, the PEPPER+Pilon correction method was selected. As shown in Table 3, the QV value of the assembly result optimized by PEPPER can be increased from 26 to 29, and the QV value of the assembly result optimized by PEPPER combined with Pilon can be further increased to 32. This proves that short read data can be used to further optimize and improve the preliminary assembly structure after assembly.
[0151] Table 3. Comparison of post-assembly error correction analysis using MGI and Illumina
[0152] As can be seen from the above description, the embodiments of this disclosure achieve the following technical effects:
[0153] 1) Using only short read length combined with nanopore long read length data for de novo assembly of monomers, high-quality assembly results are achieved with existing Hifi + nanopore + short read length.
[0154] 2) High-quality mixed assembly results with existing high-depth (greater than 50 times) long and short reads can be obtained using low-coverage sequencing data (less than or equal to 30 times).
[0155] It should be noted that the short-read data used in the embodiments of this disclosure are the most common PCR-free short-read sequencing data. Similar or better assembly results can be achieved using other short-read sequencing technologies, such as linked or Hi-C sequencing data. Furthermore, the tools used in the embodiments of this disclosure can be further adjusted with different parameters to achieve even better results for specific sequencing data.
[0156] This disclosure also provides a method for assembling a polyploid genome, comprising the steps of assembling the genomes of multiple haplotypes using short-read and long-read data from a biological sample to be tested. The steps for assembling the genome of each haplotype include: acquiring short-read and long-read data obtained by sequencing the same biological sample to be tested; correcting errors in the long-read data based on the short-read data to obtain corrected long-read data; assembling the genome based on the corrected long-read data to obtain a preliminary assembled sequence; and optimizing the preliminary assembled sequence based on the short-read data to obtain a target assembled sequence.
[0157] In the aforementioned polyploid genome assembly method, error correction is performed on long read data based on short read data to obtain corrected long read data. This includes: determining a first length and a second length, wherein the first length is shorter than the second length; constructing a first de Bruin graph corresponding to the division of short read data according to the first length, and constructing a second de Bruin graph corresponding to the division of short read data according to the second length; anchoring the long read data to the first de Bruin graph to obtain first erroneous data in the long read data, wherein the first erroneous data is data that is inconsistent with the long read data in the first de Bruin graph; correcting the first erroneous data to obtain an intermediate result; anchoring the intermediate result to the second de Bruin graph to obtain second erroneous data in the intermediate result, wherein the second erroneous data is data that is inconsistent with the intermediate result in the second de Bruin graph; and correcting the second erroneous data to obtain the corrected long read data.
[0158] In the above-mentioned polyploid genome assembly method, genome assembly is performed based on the error-corrected long read data to obtain a preliminary assembled sequence, including: determining multiple sub-sequences corresponding to the error-corrected long read data, wherein the sequence length of each sub-sequence can be the same; determining the similarity between each sub-sequence and other sub-sequences to obtain a similarity set corresponding to each sub-sequence; determining the sub-sequence corresponding to the maximum similarity in the similarity set as a ligandable sequence; and performing genome assembly on the error-corrected long read data based on the ligandable sequence to obtain a preliminary assembled sequence.
[0159] In the aforementioned polyploid genome assembly method, the preliminary assembled sequence is optimized based on short-read data to obtain the target assembled sequence. This includes: using a recurrent neural network model to identify the first variant site that is inconsistent with the short-read data, wherein the recurrent neural network model is used to determine whether the preliminary assembled sequence has mutated, and the first variant site is an erroneous data point with a variant site length of less than 50 bp; comparing the short-read data and the preliminary assembled sequence to obtain the second and third variant sites in the preliminary assembled sequence, wherein the second variant site is an erroneous data point with a variant site length of less than 50 bp, and the third variant site is an erroneous data point with a variant site length greater than or equal to 50 bp; and optimizing the data of the first, second, and third variant sites based on the short-read data to obtain the target assembled sequence.
[0160] In the above steps, the recurrent neural network model is used to determine the first variant site that is inconsistent with the short read data, including: using the recurrent neural network model to determine the single nucleotide polymorphism (SNP) site that is inconsistent with the short read data; determining the alignment file with haplotype markers based on the SNP site, wherein the haplotype markers are used to indicate the parental origin of the SNP site, and the number of haplotype markers is greater than or equal to 3; and determining the first variant site in the preliminary assembled sequence based on the alignment file with haplotype markers.
[0161] In the above steps, the data of the first variant site, the second variant site, and the third variant site are optimized based on the short read data to obtain the target assembly sequence. This includes: determining the first position of the first variant site in the short read data and determining the second position of the first variant site in the preliminary assembly sequence; determining the third position of the second variant site in the short read data and determining the fourth position of the second variant site in the preliminary assembly sequence; determining the fifth position of the third variant site in the short read data and determining the sixth position of the third variant site in the preliminary assembly sequence; replacing the data corresponding to the second position with the data corresponding to the first position, replacing the data corresponding to the fourth position with the data corresponding to the third position, and replacing the data corresponding to the sixth position with the data corresponding to the fifth position to obtain the target assembly sequence.
[0162] In the aforementioned polyploid genome assembly methods, the mixed target assembly results can also be separated into polyploid assembly results based on genotype information through phasing or genotyping. Genome phasing is a process that determines whether multiple heterozygous variant sites originate from the same sequencing fragment, i.e., whether they belong to a specific haplotype chromosome. It represents the process of locating alleles to the paternal or maternal chromosome, thereby distinguishing the genome assembly results into various haplotypes.
[0163] It should be noted that the genome assembly method for polyploids described above is based on the genome assembly method for haploids shown in Figure 2. Therefore, the relevant explanations in Figure 2 regarding the genome assembly method for haploids also apply to the genome assembly method for this polyploid, and will not be repeated here.
[0164] Figure 6 is a flowchart of a method for determining an unknown species type according to an embodiment of the present disclosure. As shown in Figure 6, the method includes:
[0165] Step S402: Obtain short read data and long read data obtained from sequencing the same biological sample to be tested;
[0166] Step S404: Correct the long read length data based on the short read length data to obtain the corrected long read length data;
[0167] Step S406: Perform genome assembly based on the error-corrected long read data to obtain the preliminary assembled sequence;
[0168] Step S408: Optimize the preliminary assembly sequence based on the short read length data to obtain the target assembly sequence;
[0169] Step S410: Compare the target assembly sequence with known biological genome sequences. Based on the comparison results, determine the species type of the biological sample to be tested. The comparison process identifies similar regions between the target assembly sequence and known biological genome sequences. The degree of similarity between these regions determines the homology between the target assembly sequence and the known biological genome sequences, thus determining which species the biological sample most likely belongs to.
[0170] It should be noted that steps S402 to S408 in Figure 6 are the same as those in Figure 2. Therefore, the relevant explanations in Figure 2 also apply to the method for determining the unknown species type, and will not be repeated here.
[0171] Figure 7 is a flowchart of an assembly-based variant detection method according to an embodiment of the present disclosure. As shown in Figure 7, the method includes:
[0172] Step S502: Obtain short read data and long read data obtained from sequencing the same biological sample to be tested;
[0173] Step S504: Correct the long read length data based on the short read length data to obtain the corrected long read length data;
[0174] Step S506: Perform genome assembly based on the error-corrected long read data to obtain the preliminary assembled sequence;
[0175] Step S508: Optimize the preliminary assembly sequence based on the short read length data to obtain the target assembly sequence;
[0176] Step S510: Based on the target assembly sequence and the reference genome sequence, determine the gene variation results of the organism to be tested.
[0177] It should be noted that steps S502 to S508 in Figure 7 are the same as those in Figure 2. Therefore, the relevant explanations in Figure 2 also apply to this assembly-based variation detection method, and will not be repeated here.
[0178] Figure 8 is a structural diagram of a haplotype genome assembly apparatus according to an embodiment of the present disclosure. As shown in Figure 8, the apparatus includes:
[0179] The first acquisition module 60 is used to acquire short-read data and long-read data obtained by sequencing the same biological sample to be tested, respectively.
[0180] The first error correction module 62 is used to correct the long read data based on the short read data to obtain the corrected long read data.
[0181] The first assembly module 64 is used to assemble the genome based on the error-corrected long read data to obtain a preliminary assembled sequence.
[0182] The first optimization module 66 is used to optimize the preliminary assembly sequence based on the short read length data to obtain the target assembly sequence.
[0183] It should be noted that the haplotype genome assembly device shown in Figure 8 is used to implement the haplotype genome assembly method shown in Figure 2. Therefore, the relevant explanations in the haplotype genome assembly method above also apply to this haplotype genome assembly device, and will not be repeated here.
[0184] This disclosure also provides a polyploid genome assembly apparatus, comprising: multiple haplotype genome assembly apparatuses, wherein each haplotype genome assembly apparatus comprises: a second acquisition module for acquiring short-read data and long-read data obtained by sequencing the same biological sample to be tested; a second error correction module for correcting errors in the long-read data based on the short-read data to obtain corrected long-read data; a second assembly module for assembling the genome based on the corrected long-read data to obtain a preliminary assembled sequence; and a second optimization module for optimizing the preliminary assembled sequence based on the short-read data to obtain a target assembled sequence.
[0185] Figure 9 is a structural diagram of an apparatus for determining an unknown species type according to an embodiment of the present disclosure. As shown in Figure 9, the apparatus includes:
[0186] The third acquisition module 70 is used to acquire short-read data and long-read data obtained by sequencing the same biological sample to be tested, respectively.
[0187] The third error correction module 72 is used to correct the long read length data based on the short read length data to obtain the corrected long read length data.
[0188] The third assembly module 74 is used to assemble the genome based on the error-corrected long read data to obtain the preliminary assembled sequence.
[0189] The third optimization module 76 is used to optimize the preliminary assembly sequence based on the short read length data to obtain the target assembly sequence;
[0190] The first determining module 78 is used to compare the target assembly sequence with known biological genome sequences, and based on the comparison results, to determine the species type of the biological sample to be tested.
[0191] It should be noted that the device for determining the unknown species type shown in Figure 9 is used to execute the method for determining the unknown species type shown in Figure 6. Therefore, the relevant explanations in the method for determining the unknown species type in Figure 6 also apply to this device for determining the unknown species type, and will not be repeated here.
[0192] Figure 10 is a structural diagram of an assembly-based variant detection device according to an embodiment of the present disclosure. As shown in Figure 10, the device includes:
[0193] The fourth acquisition module 80 is used to acquire short read data and long read data obtained by sequencing the same biological sample to be tested, respectively.
[0194] The fourth error correction module 82 is used to correct the long read length data based on the short read length data to obtain the corrected long read length data;
[0195] The fourth assembly module 84 is used to assemble the genome based on the error-corrected long read data to obtain the preliminary assembled sequence;
[0196] The fourth optimization module 86 is used to optimize the preliminary assembly sequence based on the short read length data to obtain the target assembly sequence;
[0197] The second determining module 88 is used to determine the gene variation results of the organism to be tested based on the target assembly sequence and the reference genome sequence.
[0198] It should be noted that the assembly-based mutation detection device shown in Figure 10 is used to execute the assembly-based mutation detection method shown in Figure 7. Therefore, the relevant explanations in the assembly-based mutation detection method in Figure 7 also apply to this assembly-based mutation detection device, and will not be repeated here.
[0199] This disclosure also provides an electronic device, including: a memory and a processor; the memory for storing program instructions; and the processor, connected to the memory, for executing the above-described haploid genome assembly method, or executing the above-described polyploid genome assembly method, or executing the above-described method for determining unknown species types, or executing the above-described assembly-based variation detection method.
[0200] This disclosure also provides a non-volatile storage medium including a stored computer program, wherein the device containing the non-volatile storage medium executes the above-mentioned haplotype genome assembly method, or the above-mentioned polyploid genome assembly method, or the above-mentioned method for determining unknown species types, or the above-mentioned assembly-based variation detection method by running the computer program.
[0201] This disclosure also provides a computer program product, including computer instructions that, when executed by a processor, implement the above-described haploid genome assembly method, or the above-described polyploid genome assembly method, or the above-described method for determining unknown species types, or the above-described assembly-based variation detection method.
[0202] This disclosure also provides a computer program that, when executed by a processor, implements the above-described polyploid genome assembly method, the above-described method for determining unknown species types, or the above-described assembly-based variation detection method.
[0203] The sequence numbers of the embodiments disclosed above are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0204] In the above embodiments of this disclosure, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0205] In the several embodiments provided in this disclosure, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.
[0206] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0207] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0208] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0209] The above description is only a preferred embodiment of this disclosure. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles of this disclosure, and these improvements and modifications should also be considered within the scope of protection of this disclosure. Industrial applicability
[0210] The haplotype genome assembly method provided in this disclosure relates to the field of genome assembly. It involves acquiring short-read and long-read data from sequencing the same biological sample, respectively; correcting errors in the long-read data based on the short-read data to obtain corrected long-read data; assembling the genome based on the corrected long-read data to obtain a preliminary assembled sequence; and optimizing the preliminary assembled sequence based on the short-read data to obtain the target assembled sequence. This method achieves the goal of high-quality haplotype genome assembly based on both long-read and short-read data, thereby reducing the amount of data required for genome assembly. It also solves the technical problem in related technologies where obtaining high-quality genome assembly requires the use of three types of sequencing data, resulting in a large amount of data usage.
Claims
1. A method for assembling a haplotype genome, characterized in that, The assembly method includes: Obtain short-read and long-read data respectively from sequencing the same biological sample to be tested; The long read data is corrected based on the short read data to obtain the corrected long read data. Genome assembly was performed based on the corrected long-read data to obtain a preliminary assembled sequence; The initial assembly sequence is optimized based on the short read length data to obtain the target assembly sequence.
2. The method according to claim 1, characterized in that, The coverage of the short read data and the long read data is less than or equal to 30 times.
3. The method according to claim 1, characterized in that, The long read data is corrected based on the short read data to obtain the corrected long read data, including: Construct a first de Brouin graph corresponding to the division of the short read length data according to a first length, and construct a second de Brouin graph corresponding to the division of the short read length data according to a second length, wherein the first length is less than the second length; The long read data is anchored to the first de Bruin graph to obtain the first erroneous data in the long read data, wherein the first erroneous data is data that is inconsistent between the long read data and the first de Bruin graph; The intermediate result is obtained by correcting the first erroneous data; The intermediate results are anchored to the second de Bruin graph to obtain the second erroneous data in the intermediate results, wherein the second erroneous data is data inconsistent between the intermediate results and the second de Bruin graph; The second erroneous data is corrected to obtain the corrected long read length data.
4. The method according to claim 3, characterized in that, The intermediate results obtained by correcting the first erroneous data include: Determine a first region of the first erroneous data in the short read data, and determine a second region of the first erroneous data in the long read data; The intermediate result is obtained by replacing the data contained in the second region with the data contained in the first region.
5. The method according to claim 3, characterized in that, The corrected long read length data is obtained by correcting the second erroneous data, including: The second erroneous data is determined in a third region of the short read data, and the second erroneous data is determined in a fourth region of the intermediate result; The data contained in the fourth region is replaced with the data contained in the third region to obtain the long read data after error correction.
6. The method according to claim 1, characterized in that, Genome assembly is performed based on the corrected long read data to obtain a preliminary assembled sequence, including: Determine multiple subsequences corresponding to the error-corrected long read data; Determine the similarity between each subsequence and other subsequences among the plurality of subsequences to obtain the similarity set corresponding to each subsequence; The subsequence corresponding to the maximum similarity value in the similarity set is determined as a connectable sequence; The genome is assembled from the error-corrected long-read data based on the ligandable sequence to obtain the preliminary assembled sequence.
7. The method according to claim 1, characterized in that, The initial assembly sequence is optimized based on the short read length data to obtain the target assembly sequence, including: A recurrent neural network model is used to determine the first mutation site that is inconsistent with the short read data in the preliminary assembled sequence. The recurrent neural network model is used to determine whether the preliminary assembled sequence has been mutated. The first mutation site is an erroneous data with a mutation site length of less than 50 bp. By comparing the short read data with the preliminary assembled sequence, the second and third variant sites in the preliminary assembled sequence are obtained, wherein the second variant site is an erroneous data with a variant site length of less than 50 bp, and the third variant site is an erroneous data with a variant site length of greater than or equal to 50 bp. The data of the first variant site, the second variant site, and the third variant site are optimized based on the short read data to obtain the target assembly sequence.
8. The method according to claim 7, characterized in that, The recurrent neural network model is used to determine the first variant site that is inconsistent between the preliminary assembled sequence and the short read data, including: The recurrent neural network model was used to identify single nucleotide polymorphisms (SNPs) that were inconsistent between the preliminary assembled sequence and the short read data. Based on the SNP locus, an alignment file with a haplotype marker is determined, wherein the haplotype marker is used to indicate the parental origin of the SNP locus; The first variant in the preliminary assembly sequence is determined based on the alignment file with the haplotype marker. Abnormal site.
9. The method according to claim 7, characterized in that, Based on the short read data, the data of the first variant site, the second variant site, and the third variant site are optimized to obtain the target assembly sequence, including: Determine the first position of the first variant site in the short read data, and determine the second position of the first variant site in the preliminary assembled sequence; The third position of the second variant site in the short read data was determined, and the fourth position of the second variant site in the preliminary assembled sequence was determined; The fifth position of the third variant site in the short read data was determined, and the sixth position of the third variant site in the preliminary assembled sequence was determined; The data corresponding to the second position is replaced by the data corresponding to the first position, the data corresponding to the fourth position is replaced by the data corresponding to the third position, and the data corresponding to the sixth position is replaced by the data corresponding to the fifth position to obtain the target assembly sequence.
10. A method for assembling a polyploid genome, characterized in that, For at least one haplotype of the polyploid, the genome assembly method of the haplotype as described in any one of claims 1-9 is implemented.
11. A method for determining an unknown species type, characterized in that, include: Obtain short-read and long-read data respectively from sequencing the same biological sample to be tested; The long read data is corrected based on the short read data to obtain the corrected long read data. Genome assembly was performed based on the corrected long-read data to obtain a preliminary assembled sequence; The preliminary assembly sequence is optimized based on the short read length data to obtain the target assembly sequence; By comparing the target assembled sequence with known biological genome sequences, the species type of the biological sample to be tested can be determined based on the comparison results.
12. An assembly-based variation detection method, characterized in that, include: Obtain short-read and long-read data respectively from sequencing the same biological sample to be tested; The long read data is corrected based on the short read data to obtain the corrected long read data. Genome assembly was performed based on the corrected long-read data to obtain a preliminary assembled sequence; The preliminary assembly sequence is optimized based on the short read length data to obtain the target assembly sequence; Based on the target assembly sequence and the reference genome sequence, the genetic variation results of the organism to be tested are determined.
13. A haplotype genome assembly device, characterized in that, include: The first acquisition module is used to acquire short-read data and long-read data obtained by sequencing the same biological sample to be tested, respectively. The first error correction module is used to correct the long read data based on the short read data to obtain the corrected long read data. The first assembly module is used to assemble the genome based on the error-corrected long read data to obtain a preliminary assembled sequence. The first optimization module is used to optimize the preliminary assembly sequence based on the short read length data to obtain the target assembly sequence.
14. A polyploid genome assembly device, characterized in that, include: A genome assembly apparatus for multiple haplotypes, wherein each haplotype genome assembly apparatus includes: The second acquisition module is used to acquire short-read data and long-read data obtained by sequencing the same biological sample to be tested, respectively. The second error correction module is used to correct the long read data based on the short read data to obtain the corrected long read data. The second assembly module is used to assemble the genome based on the error-corrected long read data to obtain a preliminary assembled sequence. The second optimization module is used to optimize the preliminary assembly sequence based on the short read length data to obtain the target assembly sequence.
15. A device for determining an unknown species type, characterized in that, include: The third acquisition module is used to acquire short-read data and long-read data obtained by sequencing the same biological sample to be tested, respectively. The third error correction module is used to correct the long read data based on the short read data to obtain the corrected long read data. The third assembly module is used to assemble the genome based on the error-corrected long read data to obtain a preliminary assembled sequence. The third optimization module is used to optimize the preliminary assembly sequence based on the short read length data to obtain the target assembly sequence; The first determining module is used to compare the target assembled sequence with a known biological genome sequence, and based on the comparison result, to determine the species type of the biological sample to be tested.
16. An assembly-based variation detection device, characterized in that, include: The fourth acquisition module is used to acquire short-read data and long-read data obtained from sequencing the same biological sample to be tested. The fourth error correction module is used to correct the long read data based on the short read data to obtain the corrected long read data. The fourth assembly module is used to assemble the genome based on the error-corrected long read data to obtain a preliminary assembled sequence; The fourth optimization module is used to optimize the preliminary assembly sequence based on the short read length data to obtain the target assembly sequence; The second determining module is used to determine the gene variation results of the organism to be tested based on the target assembly sequence and the reference genome sequence.
17. An electronic device, characterized in that, include: A memory and a processor, wherein the memory is used to store program instructions; The processor, connected to the memory, is used to execute the haploid genome assembly method according to any one of claims 1 to 9, or the polyploid genome assembly method according to claim 10, or the method for determining an unknown species type according to claim 11, or the assembly-based variation detection method according to claim 12.
18. A non-volatile storage medium, characterized in that, The non-volatile storage medium includes a stored computer program, wherein the device containing the non-volatile storage medium executes the haploid genome assembly method according to any one of claims 1 to 9, or the polyploid genome assembly method according to claim 10, or the method for determining an unknown species type according to claim 11, or the assembly-based variation detection method according to claim 12 by running the computer program.
19. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the genome assembly method for haploids according to any one of claims 1 to 9, or the genome assembly method for polyploids according to claim 10, or the method for determining unknown species types according to claim 11, or the variation detection method based on assembly according to claim 12.
Citation Information
Patent Citations
Method and system for assembling genomic sequence
CN104017883A
Error correction method of genomic sequence obtained by assembling PacBio sequencing data
CN107563151A
Transcriptome analysis method and system without reference genome sequence
CN112397149A
Method and device for correcting long read length assembly result by using short read length sequence
CN113012758A
Plant mitochondrial genome multi-configuration assembly method based on third-generation whole genome sequencing data
CN115449543A
Cited By
Lucid ganoderma binuclear genome assembly method based on haplotype analysis
CN121565256A
A method for assembling ganoderma bicomb genome based on haplotype resolution
CN121565256B