Assembly method of extended chromosome telomere

By combining Nanopore and Pacbio sequencing technology, the telomere information in the ONT Ultra-long reads sequence is used for comparison and screening, and the problem of telomere information loss in genome assembly is solved, achieving complete extension and high-quality assembly of chromosomal telomeres.

CN120108496APending Publication Date: 2025-06-06THE SHENNONG LABORATORY +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411993282.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing genome assembly technology is difficult to fully assemble telomeres information of chromosomes when processing low-deep regions, resulting in the loss of telomeres information of some chromosomes.

Method used

By combining Nanopore ultra-long sequencing technology and Pacbio HiFi sequencing technology, the telomere information in the ONT Ultra-long reads sequence is used for alignment and screening, and consistent sequence assembly and error correction are performed to achieve complete extension of chromosomal telomeres.

Benefits of technology

It significantly improves the integrity of genome assembly, achieves complete and high-quality extension of chromosomal telomeres, and finally obtains high-quality telomeres to telomeres (T2T) assembly results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108496A_ABST
    Figure CN120108496A_ABST
Patent Text Reader

Abstract

The invention discloses an assembly method of extended chromosome telomeres. The assembly method comprises the following steps: obtaining a chromosome sequence through an assembly technology; the method comprises the following steps: detecting a chromosome sequence to obtain a target genome chromosome sequence without telomere information, and obtaining an ONT reads sequence containing telomere information from whole genome sequencing data; comparing the ONT reads sequence containing the telomere information with a target genome chromosome sequence, and extracting an ONT reads target sequence containing the telomere information according to a comparison result; the obtained ONT Ultra-long reads target sequence containing the telomere information is subjected to consistent sequence assembly, and a chromosome extension candidate sequence is obtained; comparing the obtained residual ONT reads target sequence containing telomere information to the chromosome extension candidate sequence, and carrying out error correction to obtain a high-quality chromosome extension candidate sequence; and connecting the high-quality chromosome extension candidate sequence to a target genome chromosome sequence to prepare a chromosome sequence containing telomere information. According to the method, the sequencing length advantage of Nanopore and the sequencing precision advantage of HiFi are successfully fused, accurate extension of the chromosome sequence deletion telomere is realized, and a high-quality telomere-to-telomere assembly result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of genome assembly, and specifically to an assembly method for extending chromosome telomeres. Background Art

[0003] Genome assembly is a critical task that involves integrating DNA sequence fragments extracted from biological samples into a complete genome sequence. As a core component of modern genomics and bioinformatics, this process is essential for in-depth exploration of the genetic code of organisms, revealing evolutionary contexts, and understanding biological functions. In recent years, the rapid development of sequencing technology, especially the widespread application of high-throughput sequencing technology, has greatly improved the length and accuracy of sequencing, making genome assembly more efficient and economically feasible. These advanced sequencing technologies can produce a large number of short sequences (such as those produced by Illumina and BGI platforms) and long sequences (such as those provided by Pacbio and Nanopore platforms) at a low cost. However, due to the inherent limitations of DNA extraction and sequencing technology, it is currently impossible to directly obtain a complete genome sequence through a sequencer. In addition, the presence of sequencing errors, repeated sequences, and heterozygosity makes it a very challenging task to accurately splice many long sequences into a complete genome. Although telomere-to-telomere assembly of some chromosomes has been achieved for small genomes, some telomeres still need other means to be correctly added to the genome.

[0004] In recent years, with the release of the human T2T reference genome (T2T-CHM13), this field has made breakthrough progress, filling in the blank areas in the genome, such as telomeres and centromeres, and other highly complex structures, thus opening a new era of telomere-to-telomere assembly of the genome. At present, telomere-to-telomere assembly technology mainly relies on the single-molecule real-time sequencing (SMRT) technology of Pacific Biotechnology (PacBio) and the single-molecule sequencing technology of Oxford Nanopore Technologies (Nanopore). These two technologies have their own characteristics: Pacbio's HiFi technology can generate long reads with high accuracy (90% of the bases meet the Q30 standard), while Nanopore's Ultra-long sequencing technology can produce long reads data with N50 up to 100kb. These two technologies are undoubtedly indispensable tools for achieving telomere-to-telomere genome assembly.

[0005] Through the comprehensive use of HiFi, Hi-C and Ultra-long sequencing data, many genomes close to the T2T standard have been successfully published. However, it is worth noting that only a very small number of species have achieved complete telomere-to-telomere assembly of all chromosomes. Because the telomeric region is rich in repetitive sequences and the sequencing coverage is relatively low, the current genome assembly software shows high accuracy and continuity when processing regions with high sequencing depth, but for low-depth regions of some chromosomes, the continuity may be reduced. This results in the telomere information of some chromosomes being missing after the genome assembly is completed.

[0006] Therefore, there is an urgent need to develop new bioinformatics algorithms to perform extended assembly of chromosomes lacking telomere sequences to generate complete chromosome sequences. Summary of the invention

[0007] The purpose of this application is to provide an assembly method for extending chromosome telomeres, which successfully combines the sequencing length advantage of Nanopore and the sequencing accuracy advantage of Pacbio, not only improving the integrity of genome assembly, but also achieving the precise extension of chromosome sequence-missing telomeres, and obtaining high-quality telomere-to-telomere assembly results.

[0008] To achieve the above objectives, this application provides the following technical solutions: In a first aspect, the present application proposes an assembly method for extending chromosome telomeres, the assembly method comprising the following steps: Obtain high-quality chromosome sequences through various assembly technologies; Obtaining the target genome chromosome sequence lacking telomere information by detecting the chromosome sequence, and obtaining the ONT Ultra-long reads sequence containing telomere information from the whole genome sequencing data; The ONT Ultra-long reads sequence containing telomere information is aligned with the target genome chromosome sequence, and the ONT Ultra-long reads target sequence containing telomere information is extracted and aligned to the target genome chromosome sequence according to the alignment results; The ONT Ultra-long reads target sequences containing telomere information are assembled with consensus sequences to obtain chromosome extension candidate sequences, or the longest reads are extracted as chromosome extension candidate sequences; The remaining ONT Ultra-long reads containing telomere information are aligned to the chromosome extension candidate sequence, and error correction is performed to obtain high-quality chromosome extension candidate sequences; The high-quality chromosome extension candidate sequence is connected to the target genome chromosome sequence to obtain a chromosome sequence containing telomere information.

[0009] As a specific solution in the technical solution of the present application, high-quality chromosome sequences are obtained through various assembly technologies, including: assembling high-quality chromosome sequences by adopting various sequencing technologies, and the sequencing technologies include but are not limited to high-fidelity technology, high-throughput chromosome conformation capture technology, Oxford nanopore sequencing technology and Pacbio single molecule, long read sequencing technology.

[0010] As a specific solution in the technical solution of the present application, the target genome chromosome sequence lacking telomere information is obtained by detecting the chromosome sequence, and the ONTUltra-long reads sequence containing telomere information is obtained from the whole genome sequencing data, including: According to the fuzzy matching algorithm, the number of consecutive repetitions of the telomeric motif in the specified length before and after the chromosome sequence and the whole genome sequencing sequence is calculated; Or when the telomeric motif is not repeated continuously, a window of a certain length is set to continuously count the number of telomeric motif repeats; Count the telomere motif position information and the number of telomere motif repetitions within the specified length at both ends of the chromosome sequence and the whole genome sequencing sequence; The target genome chromosome sequence lacking telomere information and the ONT Ultra-long reads sequence containing telomere information are obtained according to the telomere motif position information and the number of repetitions of the telomere motif.

[0011] As a specific solution in the technical solution of the present application, the ONT Ultra-long reads sequence containing telomere information is compared with the target genome chromosome sequence, and the ONT Ultra-long reads target sequence containing telomere information is extracted and compared to the target genome chromosome sequence according to the comparison result, including: Select ONT Ultra-long reads sequences containing telomere information; Based on the specified parameters of minimap2 software --MD -x map-ont, the ONT Ultra-long reads sequence containing telomere information was aligned with the target genome chromosome sequence, and the ONT Ultra-long reads sequence matching the target genome chromosome sequence was extracted, and the ONT Ultra-long reads target sequence that was aligned to the specified region of the target genome chromosome and contained telomere information was obtained.

[0012] As a specific solution in the technical solution of the present application, the method of assembling the consensus sequence of the ONT Ultra-longreads target sequence containing telomere information to obtain the chromosome extension candidate sequence, or extracting the longest ONTreads target sequence as the chromosome extension candidate sequence includes: First, an assembly software is used to assemble the ONT Ultra-long reads target sequence containing telomere information to obtain a spliced ​​long, continuous genome sequence as the first chromosome extension candidate sequence, or the longest ONT Ultra-long reads target sequence containing telomere information is selected as the first chromosome extension candidate sequence; The remaining ONT Ultra-long reads target sequence containing telomere information was aligned and corrected with the first chromosome extension candidate sequence using minimap2 software to obtain a more accurate ONT Ultra-long reads target sequence as the second chromosome extension candidate sequence.

[0013] As a specific solution in the technical solution of the present application, medaka_consensus software is used to correct the first chromosome extension candidate sequence.

[0014] As a specific solution in the technical solution of the present application, the chromosome extension candidate sequence is connected to the target genome chromosome sequence to obtain a chromosome sequence containing telomere information, including: Minimap2 software was used to align the chromosome extension candidate sequence with the target genome chromosome sequence, and the optimal alignment position was selected. The chromosome extension candidate sequence was connected to the optimal alignment position of the target genome chromosome sequence to obtain a chromosome sequence containing telomere information.

[0015] As a specific solution in the technical solution of this application, the assembly method also includes: Align the obtained high-quality sequencing data to the reference genome, and identify the sequence information of single nucleotide variations, insertion or deletion variations, and structural variations based on the alignment results; Select the variant sequence information with high support from high-fidelity sequencing technology, confirm it as the real sequence, replace the original incorrectly assembled sequence, and obtain a high-quality chromosome genome; High-quality sequencing data were obtained using, but not limited to, Oxford Nanopore ultra-long sequencing technology and Pacbio single-molecule, long-read HiFi sequencing technology; The reference genome is the chromosome sequence genome after the telomere sequence is finally added.

[0016] As a specific solution in the technical solution of this application, the assembly method also includes verifying the chromosome sequence containing telomere information after error correction by obtaining ONT Ultra-long reads sequencing data using, but not limited to, Oxford Nanopore ultra-long sequencing technology and Pabio single molecule and long read sequencing technology.

[0017] As a specific scheme in the technical scheme of the present application, the verification method includes: Winnowmap software compares the obtained ONT Ultra-long reads sequencing data to the chromosome sequence containing telomere information after error correction based on the -ax map-ont and -ax map-hifi parameters, and uses IGV software to view the sequencing depth distribution of the telomere region. The sequencing depth of the telomere region of the obtained chromosome sequence containing telomere information does not show a significant decrease.

[0018] Compared with the prior art, the beneficial effects of the present application are as follows: the present invention proposes a genome assembly method for telomere extension, and by combining the advantages of Nanopore ultra-long sequencing technology and HiFi sequencing technology, a complete and accurate extension of chromosome sequence-deficient telomeres is achieved. Nanopore sequencing technology relies on its excellent ability to cross high-repetitive regions, and its sequencing results often contain abundant telomere sequence information. However, limited by the performance of current assembly software, the telomere information of some low-depth regions is difficult to be fully assembled, resulting in the lack of telomere information in the chromosome assembly results. Although HiFi sequencing technology is limited by the length of the library and cannot cover telomere regions exceeding its size, its sequencing accuracy is extremely high, and therefore it is often used to construct an accurate skeleton of the genome. In view of this, the present invention creatively combines the advantages of Nanopore and HiFi sequencing technologies, firstly uses HiFi sequencing data to construct an accurate basic framework of the genome, combines other technologies to construct a high-quality genome at the chromosome level, and then uses the ultra-long reads characteristics of Nanopore technology to accurately extend the missing telomere region. Not only did it significantly improve the integrity of the genome assembly, it also achieved complete and high-quality extension of chromosome telomeres, and ultimately obtained T2T (telomere to telomere) high-quality assembly results. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A flowchart of an assembly method for extending chromosome telomeres is provided for an embodiment of the present application, and black represents telomere signals; Figure 2 The telomere information loss of the assembled high-quality chromosomes provided in the embodiments of the present application; Figure 3 Provide a telomere signal of a genome obtained after extending the telomeres for the embodiment of the present application; Figure 4Verification of the chromosome heat map after telomere extension using Hi-C data provided in the embodiments of this application; Figure 5 This is the verification of the alignment of the HiFi and ONT sequencing data provided in the embodiments of the present application to the chromosomes after telomere extension. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be further clearly and completely described below. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0021] In order to improve the assembly method for extending chromosome telomeres proposed in the background technology, this assembly method uses high-fidelity sequencing technology (HiFi), ultra-long read sequencing technology (Ultra-long) and high-throughput chromosome conformation capture technology (Hi-C) to complete the chromosome genome assembly. The main purpose is to extend the telomeres of the assembled chromosome genome to ensure that all missing telomere sequences are supplemented, thereby achieving the complete assembly of each chromosome from telomere to telomere (T2T).

[0022] The assembly method provided in the specific embodiment of the present application includes: (1) High-quality genome assembly: Use a variety of sequencing technologies to assemble a high-quality genome at the chromosome level, especially Ultra-long data, with a coverage depth of at least 50X and a continuous sequence N50 of at least 100kb for reads to improve the accuracy and continuity of the assembly. Some of the resulting chromosome genomes may contain sequences that are partially missing from chromosome telomeres, and telomere sequences may be missing from the resulting chromosome genome. Sequencing technologies include but are not limited to high-fidelity technology, high-throughput chromosome conformation capture technology, Oxford nanopore sequencing technology, and ultra-long read sequencing technology.

[0023] (2) Telomere detection and extraction: The method comprises: obtaining a target genome chromosome sequence lacking telomere information by detecting a chromosome sequence, and obtaining an ONT Ultra-long reads sequence containing telomere information by sequencing a whole genome, including: calculating the number of consecutive repetitions of a telomere motif in a specified length before and after the chromosome sequence and the whole genome sequence according to a fuzzy matching algorithm; or when the telomere motif is not repeated continuously, setting a window of a certain length, and continuously calculating the number of telomere motif repetitions; counting the telomere motif position information and the number of repetitions of the telomere motif within a specified length at both ends of the chromosome sequence and the whole genome sequence; and obtaining a target genome chromosome sequence lacking telomere information and an ONT Ultra-long reads sequence containing telomere information according to the telomere motif position information and the number of repetitions of the telomere motif.

[0024] (3) ONT Ultra-long reads comparison and screening: The ONT Ultra-long reads with telomere signals were aligned to the target genome chromosome sequence using minimap2 (https: / / github.com / lh3 / minimap2, -x map-ont), and the ONT Ultra-long reads target sequence that was aligned to the target genome chromosome sequence and contained telomere information was extracted based on the alignment results. Specifically: first, the ONT Ultra-long reads sequence containing telomere information was selected; secondly, the ONT Ultra-long reads sequence containing telomere information was aligned with the target genome chromosome sequence based on the specified parameters of the minimap2 software --MD -x map-ont, and the ONT Ultra-long reads sequence that matched the target genome chromosome sequence was extracted to obtain the ONT Ultra-long reads target sequence containing telomere information.

[0025] (4) Consensus sequence assembly: Consensus sequence assembly and alignment error correction strategy were used for the extracted ONT Ultra-long reads target sequence containing telomere information to obtain chromosome extension candidate sequences; This step includes two sub-steps: assembly and obtaining consensus sequences: first, assembly software is used to assemble the ONT Ultra-long reads target sequence containing telomere information to obtain a spliced, continuous genome sequence as the first chromosome extension candidate sequence; or the longest ONT Ultra-long reads target sequence containing telomere information is selected as the first chromosome extension candidate sequence; minimap2 software is used to align and correct the remaining ONT Ultra-long reads target sequence containing telomere information with the first chromosome extension candidate sequence to obtain a more accurate chromosome extension candidate sequence as the second chromosome extension candidate sequence. Assembly software includes but is not limited to Next-denovo, wtdbg2 and Flye, and error correction includes but is not limited to medaka_consensus.

[0026] (5) Telomere extension: The chromosome extension candidate sequence is connected to the target genome chromosome sequence to obtain a chromosome sequence containing telomere information, including: using minimap2 software to compare the chromosome extension candidate sequence with the target genome chromosome sequence, selecting the optimal comparison position, and connecting the chromosome extension candidate sequence to the optimal comparison position of the target genome chromosome sequence to obtain a chromosome sequence containing telomere information.

[0027] (6) Chromosome genome error correction: After obtaining the complete genome sequence, the chromosome sequence containing telomere information is corrected based on the default parameters of the Inspector software and the single-molecule, long-length sequencing reads obtained by Pacbio sequencing. Specifically, the obtained high-quality sequencing data is aligned to the reference genome, and the sequence information of single nucleotide variations, insertion or deletion variations, and structural variations are identified based on the alignment results; the variant sequence information with high support by high-fidelity sequencing technology is selected and confirmed as the true sequence, and the original assembly error sequence is replaced to obtain a high-quality chromosome genome; high-quality sequencing data is obtained using, but not limited to, Oxford Nanopore ultra-long sequencing technology and Pacbio single-molecule, long-read HiFi sequencing technology; the reference genome is the chromosome sequence genome after the telomere sequence is finally added.

[0028] (7) Chromosome genome verification: The corrected genome is fully verified by combining the high-throughput chromosome conformation capture technology (Hi-C) obtained by sequencing, Nanopore ultra-long read sequencing technology (Ultra-long) and Pacbio's single-molecule, long-length sequencing technology (HiFi). Ensure the accuracy and integrity of telomere extension and genome assembly. The verification method provided in this embodiment is that Winnowmap software uses the -ax map-ont or -ax map-hifi parameters to align the long-length sequencing data to the chromosome sequence containing telomere information after error correction, and uses IGV software to view the sequencing depth distribution of the telomere region. The sequencing depth of the telomere region of the chromosome sequence containing telomere information does not show a significant decrease. Example

[0030] This embodiment provides a method for assembling extended chromosome telomeres using sesame as an example, and the specific implementation process is as follows: Figures 1 to 5 As shown, the following steps are included: (1) Telomere information detection was performed on the assembled 13 high-quality target genomes and the ultra-long reads obtained by Oxford Nanopore sequencing technology (ONT) sequencing ( Figure 1 ); (2) According to the telomere information detection results of the genome, the chromosomes with missing telomeres in the target genome (chr2, chr4, chr7, chr12) are selected, and the positions of the missing telomeres are counted. Chr12 is selected as an example for telomere extension ( Figure 1 ); (3) Based on the telomere information detection results of ONT sequencing reads, select the sequences with telomere information in ONT Ultra-long reads, and count the positions of the telomeres ( Figure 1 ); (4) Use minimap2 software or Winnowmap to align the ONT Ultra-long reads with telomere information to chromosome chr12, screen the ONT Ultra-long reads sequences that match chromosome chr12, and obtain the ONT Ultra-long reads target sequences that contain telomere information and align to the specified region ( Figure 1 ); (5) For the extracted ONT Ultra-longreads that are aligned to the specified chromosome region and contain telomere information, use flye, next-denovo, wtdbg2 and other software to assemble the consensus sequence and obtain the spliced ​​long, continuous genome sequence as the first chromosome extension candidate sequence; when the extracted sequence is relatively small, select the longest sequence as the first chromosome extension candidate sequence ( Figure 1 ); (6) Using medaka_consensus software, align and correct all the spliced ​​long, continuous ONTUltra-long reads sequences obtained in step (5) with the first chromosome extension candidate sequence to obtain a more accurate chromosome extension candidate sequence as the second chromosome extension candidate sequence ( Figure 1 ); (7) The corrected second chromosome extension candidate sequence obtained above is aligned with the target genome chromosome sequence by using minimap2 software, the optimal alignment position is selected, and the chromosome extension candidate sequence is connected to the optimal alignment position of the target genome chromosome sequence to obtain a chromosome sequence containing telomere information, such as Figure 2 and Figure 3 As shown; (8) Use chromap software to align Hi-C sequencing data to the reference genome, and use heat map display to verify the accuracy of chromosomes based on the interaction strength, such as Figure 4 shown.

[0031] (9) Winnowmap software was used to align the ONT and HiFi sequencing data to the chromosome sequences containing telomere information after error correction based on the -ax map-ont and -ax map-hifi parameters. The sequencing depth distribution of the telomere region was checked using IGV software. The sequencing depth of the telomere region of the chromosome sequences containing telomere information did not show a significant decrease, as shown in Figure 2. Figure 5 .

[0032] Although the embodiments of the present application have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present application, and that the scope of the present application is defined by the appended claims and their equivalents.

Claims

1. A method for assembling a chromosome telomere extension, characterized in that: The assembly method comprises the following steps: Obtain high-quality chromosome sequences through assembly technology; Obtaining the target genome chromosome sequence lacking telomere information by detecting the chromosome sequence, and obtaining the ONT Ultra-long reads sequence containing telomere information from the whole genome sequencing data; The ONT Ultra-long reads sequence containing telomere information is aligned with the target genome chromosome sequence, and the ONT Ultra-long reads target sequence containing telomere information is extracted and aligned to the target genome chromosome sequence according to the alignment result; The obtained ONT Ultra-long reads target sequence containing telomere information is assembled with a consensus sequence to obtain a chromosome extension candidate sequence, or the longest ONT Ultra-long reads target sequence is selected as a chromosome extension candidate sequence; Align the remaining ONT Ultra-long reads target sequences containing telomere information to the chromosome extension candidate sequences, perform error correction, and obtain high-quality chromosome extension candidate sequences; The high-quality chromosome extension candidate sequence is connected to the target genome chromosome sequence to obtain a chromosome sequence containing telomere information.

2. The assembly method according to claim 1, characterized in that: Obtain high-quality chromosome sequences through assembly technology, including: High-quality chromosome sequences are assembled by adopting various sequencing technologies, including but not limited to high-fidelity technology, high-throughput chromosome conformation capture technology, Oxford Nanopore ultra-long sequencing technology and Pabio single-molecule, long-read sequencing technology.

3. The assembly method according to claim 1, characterized in that: The method of obtaining a target genome chromosome sequence lacking telomere information by detecting the chromosome sequence, and obtaining an ONT Ultra-long read sequence containing telomere information from whole genome sequencing data, comprises: According to the fuzzy matching algorithm, the number of consecutive repetitions of the telomeric motif in the specified length before and after the chromosome sequence and the whole genome sequencing sequence is calculated; Or when the telomeric motif is not repeated continuously, a window of a certain length is set to continuously count the number of telomeric motif repeats; Count the telomere motif position information and the number of telomere motif repetitions within the specified length at both ends of the chromosome sequence and the whole genome sequencing sequence; The target genome chromosome sequence lacking telomere information and the ONT Ultra-long reads sequence containing telomere information are obtained according to the telomere motif position information and the number of repetitions of the telomere motif.

4. The assembly method according to claim 1, characterized in that: The ONT Ultra-long reads sequence containing telomere information is compared with the target genome chromosome sequence, and the ONT Ultra-long reads target sequence containing telomere information is extracted and compared to the target genome chromosome sequence according to the comparison result, including: Select ONT Ultra-long reads sequences containing telomere information; Based on the specified parameters of minimap2 software --MD -x map-ont, the ONT Ultra-longreads sequence containing telomere information was aligned with the target genome chromosome sequence, and the ONT Ultra-long reads target sequence that was aligned to the specified region of the target genome chromosome and contained telomere information was extracted.

5. The assembly method according to claim 1, characterized in that: The method of assembling a consensus sequence of an ONT Ultra-long read target sequence containing telomere information to obtain a chromosome extension candidate sequence, or selecting the longest ONT Ultra-long read target sequence as a chromosome extension candidate sequence, comprises: First, an assembly software is used to assemble the ONT Ultra-long reads target sequence containing telomere information to obtain a spliced ​​long, continuous genome sequence as the first chromosome extension candidate sequence, or the longest ONT Ultra-long reads target sequence containing telomere information is selected as the first chromosome extension candidate sequence; Using minimap2 software, the remaining ONT Ultra-long reads target sequence containing telomere information was aligned and corrected with the first chromosome extension candidate sequence to obtain a more accurate extension candidate sequence as the second chromosome extension candidate sequence.

6. The assembly method according to claim 5, characterized in that: The medaka_consensus software was used to correct the errors of the candidate sequences of the first chromosome extension.

7. The assembly method according to claim 1, characterized in that: Connecting the chromosome extension candidate sequence to the target genome chromosome sequence to obtain a chromosome sequence containing telomere information, including: Minimap2 software was used to align the chromosome extension candidate sequence with the target genome chromosome sequence, and the optimal alignment position was selected. The chromosome extension candidate sequence was connected to the optimal alignment position of the target genome chromosome sequence to obtain a chromosome sequence containing telomere information.

8. The assembly method according to claim 1, characterized in that: The assembly method further comprises: Align the obtained high-quality sequencing data to the reference genome, and identify the sequence information of single nucleotide variations, insertion or deletion variations, and structural variations based on the alignment results; Select the variant sequence information with high support from high-fidelity sequencing technology, confirm it as the real sequence, replace the original incorrectly assembled sequence, and obtain a high-quality chromosome genome; High-quality sequencing data were obtained using, but not limited to, Oxford Nanopore ultra-long sequencing technology and Pacbio single-molecule, long-read HiFi sequencing technology; The reference genome is the chromosome sequence genome after the telomere sequence is finally added.

9. The assembly method according to claim 8, characterized in that: The assembly method further comprises verifying the chromosome sequence containing telomere information after error correction by obtaining ONT Ultra-long reads and HiFi sequencing data including but not limited to Oxford Nanopore ultra-long sequencing technology and Pacbio single molecule and long read sequencing technology.

10. The assembly method according to claim 9, characterized in that: The verification method comprises: using Winnowmap software to compare the obtained ONT Ultra-long reads sequencing data to the chromosome sequence containing telomere information after error correction based on -ax map-ont and -ax map-hifi parameters, using IGV software to check the sequencing depth distribution of the telomere region, and the sequencing depth of the telomere region of the obtained chromosome sequence containing telomere information does not show a significant decrease.

Citation Information

Cited By

  • Method and apparatus for whole genome sequencing assembly of plasmid-containing bacteria

    CN122619111A