Tool kit and method for supplementing chromosome telomere in T2T genome
Through the alignment, splicing and polishing module of the toolkit, the telomere sequence is accurately identified and supplemented, and the problem of incomplete assembly of telomere regions in the existing technology is solved, and efficient and precise assembly of the genome is achieved.
Patent Information
- Application Number
- CN202510595750.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-15
AI Technical Summary
Existing telomere recognition tools cannot fully utilize existing sequencing data, making it difficult to achieve complete assembly of telomere regions, and lack flexibility in handling ultra-long read lengths and multi-source data integration in complex genomes.
Provide a toolkit, including a alignment and screening module, a genome assembly and grinding module, a new genome generation module and a visual output module, through alignment, splicing, grinding and visualization methods, accurately identify and supplement telomerular sequences to form a complete new genome.
Extend the identified telomeres sequences at the end of chromosomes, improve the degree of refinement and overall integrity of genome assembly, and significantly improve the efficiency and accuracy of telomere region assembly and analysis.
Smart Images

Figure CN120496632A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of genome assembly, and in particular to a tool kit and method for supplementing chromosome telomeres in a T2T genome. Background Art
[0002] The completeness and accuracy of genome assemblies are fundamentally important in genomic research and are crucial for ensuring the reliability of downstream functional studies, evolutionary analyses, and various biological explorations. The achievement of telomere-to-telomere (T2T) genome assembly not only meets the requirements for complete genome assembly but also provides a solid foundation for comprehensive analysis of gene function and the laws of genome evolution.
[0003] Telomeres, as special structures at the ends of chromosomes, are composed of nucleoprotein complexes formed by tandem repeat sequences and bound protective proteins. They play a role in protecting the ends of chromosomes and maintaining genomic stability. The integrity of telomeres is crucial to the function of the genome, and their key roles in cell division, aging, maintaining genomic stability, and disease development have been widely recognized. However, during the genome assembly process, because telomeres are located at the ends of chromosomes, the high complexity of their repeat sequences and the significant variation in length make it difficult to accurately resolve or completely assemble the telomere regions. This assembly defect often manifests as the loss or incorrect assembly of telomere sequences, making it difficult to achieve telomere-to-telomere (T2T) levels of the genome.
[0004] Currently, although third-generation sequencing technologies such as Oxford Nanopore Technology (ONT)'s ultra-long read sequencing and Pacific Biosciences (PacBio)'s high-fidelity (HiFi) sequencing technology have shown significant advantages in genome assembly, the use of a single technology is still difficult to completely overcome the assembly challenges of telomere regions. Although some tools dedicated to telomere identification can effectively identify telomere sequences to a certain extent, facilitating the analysis of telomere regions in the genome, their main functions are still focused on the simple positioning and identification of telomere sequences. They lack the ability to further integrate and supplement unidentified telomere regions, and are unable to fully utilize existing sequencing data to improve the assembly integrity of telomere regions. In addition, when faced with complex genomes, current telomere identification tools often have difficulty handling the integration of ultra-long reads and multi-source data, and are unable to provide flexible solutions based on the genomic characteristics of different species.
[0005] Therefore, it is necessary to propose a toolkit and method to supplement the chromosome telomeres in the T2T genome, which can accurately identify the telomere sequences in the genome, make full use of the existing sequencing data to improve the assembly integrity of the telomere region, and achieve complete assembly of the telomere region. Summary of the Invention
[0006] In view of this, the present invention provides a tool kit and method for supplementing chromosome telomeres in the T2T genome, so as to solve the technical problem that existing telomere identification tools cannot fully utilize existing sequencing data to improve the assembly integrity of telomere regions due to the lack of further integration and supplementation capabilities of unidentified telomere regions.
[0007] In order to achieve the above technical objectives, the present invention adopts the following technical solutions:
[0008] In a first aspect, the present invention provides a kit for supplementing chromosome telomeres in a T2T genome, comprising:
[0009] An alignment and screening module is used to align the long-read sequencing sequence of the genome to be tested with the reference genome to obtain a primary alignment result, screen the primary alignment result to obtain the original genome in which the genome to be tested matches the reference genome, and a first out-of-range sequence that exceeds the end of the chromosome of the reference genome, and trim the first out-of-range sequence based on a preset coverage;
[0010] The genome assembly and polishing module is used to determine the relevant fragments of read matching based on the overlapping relationship between the fragments in the trimmed first overrange sequence, splice the relevant fragments, construct the draft genome, and polish the draft genome to improve the accuracy of the assembled sequence;
[0011] A new genome generation module is used to align the polished draft genome with the original genome to obtain a secondary alignment result. Based on the secondary alignment result, the draft genome is distinguished from the sequences belonging to the original genome and the second out-of-range sequence that extends beyond the end of the chromosome. The start and end position information of the second out-of-range sequence is determined, the start and end position information is matched with the original genome, the telomere sequence is extracted, and the telomere sequence is pasted back to the original genome to obtain a new genome.
[0012] The visualization output module is used to obtain a telomere sequence analysis report based on the original genome, the new genome, and the specific telomere motifs of the species corresponding to the genome to be tested.
[0013] On the other hand, the present invention also provides a method for supplementing chromosome telomeres in a T2T genome, which is implemented using the tool kit for supplementing chromosome telomeres in a T2T genome described in the above technical solution, comprising:
[0014] Aligning the long-read sequencing sequence of the genome to be tested with the reference genome to obtain a primary alignment result, screening the primary alignment result to obtain an original genome in which the genome to be tested matches the reference genome, and a first out-of-range sequence that extends beyond the end of the chromosome of the reference genome, and trimming the first out-of-range sequence based on a preset coverage;
[0015] According to the overlapping relationship between the fragments in the trimmed first overrange sequence, the relevant fragments of the read matching are determined, the relevant fragments are spliced, the draft genome is constructed, and the draft genome is polished to improve the accuracy of the assembled sequence;
[0016] The polished draft genome is aligned with the original genome to obtain a secondary alignment result. Based on the secondary alignment result, the sequences in the draft genome that belong to the original genome and the second out-of-range sequence that extends beyond the end of the chromosome are distinguished. The start and end position information of the second out-of-range sequence is determined, and the start and end position information is matched with the original genome. The telomere sequence is extracted and pasted back into the original genome to obtain a new genome.
[0017] A telomere sequence analysis report is obtained based on the specific telomere motifs of the original genome, the new genome, and the species corresponding to the genome to be tested.
[0018] Furthermore, the screening of the initial alignment results to obtain a first out-of-range sequence beyond the end of the reference genome chromosome includes:
[0019] The chromosome ends containing at least one telomere marker sequence corresponding to the species to be detected and whose coverage reaches a preset threshold are screened, and the filtered readings are output according to the left and right ends of the chromosomes to obtain the first out-of-range sequence.
[0020] Furthermore, the splicing of related fragments to construct a draft genome includes:
[0021] Calculate the overlap relationship between all read fragments, generate an overlap graph, and determine the connection order and direction of the fragments;
[0022] Optimize the paths of overlapping graphs, eliminate redundant connections and erroneous nodes, integrate fragments into continuous genome sequences, and mark duplicate and uncertain regions;
[0023] Repetitive regions are removed or isolated through preset algorithms to reduce redundancy and errors caused by repetitive sequences during the assembly process;
[0024] Outputs a draft genome containing uncertain regions.
[0025] Furthermore, polishing the draft genome sequence includes:
[0026] Aligning short-read data from Qualcomm sequencing to the draft genome improves local sequence accuracy by statistically analyzing alignment depth and variation frequency.
[0027] Correction of long-range structural errors using long-read sequencing sequences;
[0028] Filter low coverage areas and artifact sequences;
[0029] The draft genome was iteratively optimized based on the alignment results of Qualcomm sequencing short-read data and long-read sequencing sequences to improve the consistency of the genome sequence and complete the polishing of the draft genome sequence.
[0030] Furthermore, the step of matching the start and end position information with the position information of the left and right ends of the original genome to extract the telomere sequence includes:
[0031] If the position information matches and the directions are consistent, the matching sequence is selected as the telomere sequence;
[0032] If the position information matches but the directions are inconsistent, the reverse complementary matching sequence is selected as the telomere sequence.
[0033] Furthermore, pasting the telomere sequence back into the original genome to obtain a new genome comprises:
[0034] Before pasting the telomere sequence back to the original genome, determine whether the ratio of the telomere sequence length to the supplemented sequence length reaches a preset ratio threshold;
[0035] If the preset ratio threshold is reached, the telomere sequence is pasted back to the original genome to form a new genome with supplementary telomeres.
[0036] Furthermore, if a preset ratio threshold is reached, the telomere sequence is pasted back to the original genome, including:
[0037] According to the alignment direction of the draft genome sequence and the left and right end information of the original chromosome, the telomere sequence is spliced to the corresponding position.
[0038] Furthermore, the telomere sequence analysis report is obtained based on the specific telomere motifs of the original genome, the new genome, and the species corresponding to the genome to be tested, including:
[0039] Cut sequence reads of a preset length at the ends of the original genome;
[0040] Cut a region of the same preset length from the end of the new genome and add the extracted telomere sequence;
[0041] The processed original genome and the new genome are aligned, the telomere density is calculated based on the species-specific telomere motif, a collinearity comparison diagram of the original genome and the new genome is obtained, and telomere sequence statistics are output.
[0042] Furthermore, the telomere sequence statistical information includes the telomere information of the original genome and the telomere information of the new genome after the telomeres are supplemented;
[0043] The telomere information of the original genome at least includes the length and distribution position of the telomere sequence;
[0044] The telomere information of the new genome after the telomeres are supplemented includes at least the length, position, type and number of the newly added telomere sequences.
[0045] Compared with the prior art, the advantages of the tool kit and method for supplementing chromosome telomeres in T2T genomes provided by the present invention are:
[0046] 1. This toolkit can effectively extend the identified telomere sequences at the ends of chromosomes, making the assembly of telomere regions more complete and improving the refinement of genome assembly.
[0047] 2. It can capture telomere sequences missed in the traditional assembly process, improve the overall integrity of the genome by completing the genome ends, and provide technical support for T2T level genome assembly.
[0048] 3. Providing full-process and visual operations, users can intuitively compare the differences in the genome before and after supplementation, significantly improving the efficiency and accuracy of telomere region assembly and analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 A schematic diagram of the structure of a tool kit for identifying and replenishing chromosome telomeres in the T2T genome provided by the present invention;
[0050] Figure 2 The present invention provides a schematic flow chart of a method for supplementing chromosome telomeres in a T2T genome;
[0051] Figure 3 A flowchart of the TeloComp process provided by the present invention;
[0052] Figure 4 This is a schematic diagram of the visual analysis report provided by the present invention. DETAILED DESCRIPTION
[0053] The preferred embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, and are not used to limit the scope of the present invention.
[0054] See Figure 1 This embodiment provides a tool kit 100 for supplementing chromosome telomeres in a T2T genome, comprising:
[0055] An alignment and screening module 101 is configured to align the long-read sequencing sequence of the genome to be tested with the reference genome to obtain a primary alignment result, screen the primary alignment result to obtain an original genome in which the genome to be tested matches the reference genome, and a first out-of-range sequence that extends beyond the end of the chromosome of the reference genome, and trim the first out-of-range sequence based on a preset coverage;
[0056] The genome assembly and polishing module 102 is used to determine the relevant fragments with read matching based on the overlapping relationship between the fragments in the trimmed first out-of-range sequence, splice the relevant fragments, construct a draft genome, and polish the draft genome to improve the accuracy of the assembled sequence;
[0057] A new genome generation module 103 is configured to align the polished draft genome with the original genome to obtain a secondary alignment result, distinguish, based on the secondary alignment result, sequences in the draft genome that belong to the original genome and second out-of-range sequences that extend beyond the ends of the chromosomes, determine the start and end position information of the second out-of-range sequences, match the start and end position information with the original genome, extract telomere sequences, and paste the telomere sequences back into the original genome to obtain a new genome;
[0058] The visualization output module 104 is used to obtain a telomere sequence analysis report based on the original genome, the new genome, and the specific telomere motifs of the species corresponding to the genome to be detected.
[0059] The toolkit provided in this embodiment supplements and improves telomere regions to obtain a more complete genome, especially at the chromosome ends, avoiding parts that may be missed by traditional methods. Through precise trimming and alignment of out-of-range sequences, it is possible to extract complete telomere sequences and paste them back into the original genome, ensuring the integrity and accuracy of the telomeres. The original genome and the new genome are aligned with species-specific telomere motifs to provide clear and intuitive results and generate a telomere sequence analysis report. The toolkit provided in this embodiment systematically solves the problem of supplementing telomere regions and, combined with efficient alignment, assembly, polishing, visualization and other functions, provides an efficient analysis tool for genomic research.
[0060] As a specific example, the toolkit was designed and executed using Python version 3.11.6, with Linux as the primary platform. TeloComp also relies on samtools, minimap2, bwa, Flye, racon, NextPolish, teloclip, and GenomeSyn.
[0061] It should be noted that the four modules in this toolkit can be run separately for different genomes or as an overall process, and TeloComp is compatible with multi-core parallel operations.
[0062] Based on the above-mentioned tool kit, this embodiment further provides a method for supplementing chromosome telomeres in a T2T genome using the tool kit, comprising:
[0063] Step S101: Align the long-read sequencing sequence of the genome to be tested with the reference genome to obtain a primary alignment result, screen the primary alignment result to obtain the original genome in which the genome to be tested matches the reference genome, and the first out-of-range sequence that extends beyond the end of the chromosome of the reference genome, and trim the first out-of-range sequence based on a preset coverage;
[0064] Step S102: determining relevant fragments with read matches based on the overlapping relationships between the fragments in the trimmed first out-of-range sequence, splicing the relevant fragments to construct a draft genome, and polishing the draft genome to improve the accuracy of the assembled sequence;
[0065] Step S103: Align the polished draft genome to the original genome to obtain a secondary alignment result. Based on the secondary alignment result, distinguish the sequences in the draft genome that belong to the original genome and the second out-of-range sequences that extend beyond the chromosome ends, determine the start and end position information of the second out-of-range sequences, match the start and end position information with the original genome, extract the telomere sequences, and paste the telomere sequences back into the original genome to obtain a new genome.
[0066] Step S104: obtaining a telomere sequence analysis report based on the original genome, the new genome, and the specific telomere motifs of the species corresponding to the genome to be detected.
[0067] To provide a more detailed description of the workflow for the T2T genome telomere complementation method, the following uses the telomere complementation toolkit (referred to as TeloComp for ease of description) as an example.
[0068] Step 1: Sequencing data alignment and telomere read screening
[0069] First, you need to obtain the long-read sequencing data of the genome to be tested and the reference genome sequence. There are two data input methods to choose from.
[0070] The first data input method is to directly input the original genome file (usually in FASTA format) and third-generation data such as ONT or HiFi, that is, long-read sequencing data files (usually in FASTQ format). The genome file contains the reference genome sequence of the target species for subsequent alignment analysis; the ONT or HiFi file contains the long-read data obtained by high-throughput sequencing, which is the key input for identifying and supplementing telomere sequences.
[0071] In this way, minimap2 is used to align the sequencing data with the reference genome to generate the corresponding alignment file (SAM format). Although this process requires more computing resources and time, it is suitable for processing genome alignment tasks from scratch.
[0072] The second data input method is to directly input a preprocessed BAM file. BAM files are binary versions of SAM files, typically generated by alignment tools (such as minimap2) and subsequent tools (such as samtools). They contain information about the alignment of sequence reads to the reference genome. Using BAM files as input can skip the alignment step, significantly reducing run time and making it suitable for tasks with existing alignment results. Furthermore, BAM files contain information such as alignment coverage and alignment quality, facilitating subsequent read screening and processing steps.
[0073] Regardless of the method used, the final output is reads filtered for specific telomeric motifs. These reads are filtered by the teloclip tool to meet the requirement of containing at least one telomeric motif and are trimmed according to the user-selected coverage (e.g., 0-100%) to remove low-quality or redundant sequences to improve accuracy.
[0074] Finally, the screening results are output to corresponding folders according to the left and right ends of the chromosome, providing high-quality input data for subsequent assembly and telomere supplementation.
[0075] Step 2: Assembly and error correction of telomere sequences
[0076] The first step is to input a directory file containing trimmed left-end or right-end reads, assemble using Flye's overlapping graph method, and construct a draft genome. Subsequently, combined with whole-genome sequencing (WGS) data, the draft genome sequence is polished using NextPolish software to improve the accuracy of the assembled sequence.
[0077] Flye constructs a draft genome by analyzing the overlapping relationships between DNA sequence fragments and piecing together the matching read fragments. The specific assembly process includes the following steps:
[0078] Step S21: Constructing an overlap graph: Calculating the overlap relationship between all read fragments, generating an overlap graph, and determining the connection order and direction of the fragments.
[0079] Step S22: Path optimization and assembly: By optimizing the paths of the overlapping graph, eliminating redundant connections and erroneous nodes, the fragments are integrated into a continuous genome sequence, and the repeated regions and uncertain regions are marked.
[0080] Step S23: Repeated region processing: Special processing is performed on the repeated regions to eliminate redundancy and errors in assembly as much as possible, ensuring the accuracy and consistency of the assembly results.
[0081] After the above steps, Flye outputs a preliminary constructed draft genome.
[0082] The specific processing steps for polishing the draft genome sequence using NextPolish in combination with whole genome sequencing (WGS) data include:
[0083] Step S31: Error correction based on second-generation sequencing data (Qualcomm sequencing short-read data): Align the second-generation sequencing data such as Illumina to the draft genome, and correct single nucleotide variations (SNPs), insertions and deletions (Indels) and structural errors by statistically analyzing the alignment depth and variation frequency to improve the local accuracy of the sequence.
[0084] Step S32: Structural optimization based on third-generation sequencing data: Utilize the high coverage of long reads such as ONT or HiFi to correct long-range structural errors that are difficult to identify in second-generation data (such as incorrect connection, deletion or inversion of repeated sequences).
[0085] Step S33: Eliminate artifact regions: Detect and filter low coverage regions and possible artifact sequences, and remove mismatches and unreliable fragments to prevent these regions from affecting downstream analysis.
[0086] Step S34: Generate consensus sequence: Comprehensively analyze the alignment results of the second-generation and third-generation data, iteratively optimize, and finally generate a highly consistent genome sequence to ensure the overall accuracy and integrity of the draft genome.
[0087] Through Flye assembly and NextPolish polishing, the optimized draft genome is finally output to the same directory file, providing reliable basic data for subsequent telomere supplementation and genome improvement.
[0088] Step 3: Integration of telomere sequences and genome complementation
[0089] After polishing the draft genome, the directory file containing the polished draft genome is used as input. Minimap2 is used to align the draft genome with the original genome, generating a SAM file for further analysis of the alignment information. The pysam tool is then used to parse the alignment results in the SAM file, distinguishing between sequences in the draft genome that belong to the original genome and sequences that were assembled as excess, determining the start and end positions of these sequences, and then matching these start and end positions with the left and right ends of the original genome. Particular attention should be paid to the separation point between the telomere sequence and the original genome sequence.
[0090] Based on this alignment, the results are used to determine whether each telomere sequence correctly corresponds to the left or right end of the original chromosome. It should be noted that if the orientation of the telomere sequence in the draft genome is inconsistent with that of the original chromosome, the draft genome sequence must be reverse-complemented to ensure the correct sequence orientation.
[0091] Telomere sequences are cut according to the alignment demarcation points, and the draft genomic sequence beyond the original chromosome end is extracted. The proportion of telomere sequence is calculated. If the proportion of the cut telomere sequence to the total length of DNA beyond the original chromosome end reaches a preset ratio (usually set to 20% in practice) or above, the telomere sequence is considered a valid sequence and included in subsequent processing.
[0092] For telomere sequences that meet the requirements, they are attached back to the original chromosome ends to form a new genome that complements the telomeres.
[0093] Specifically, when replying, the telomere sequence is accurately spliced to the corresponding position based on the alignment direction of the draft genome sequence and the left and right ends of the original chromosome. For example, for the left end of the telomere sequence, the excised telomere sequence is spliced to the left end of the original chromosome according to its direction; for the right end of the telomere sequence, the excised telomere sequence is spliced to the right end of the original chromosome according to its direction.
[0094] The above steps generate a new genome output with detailed information, including the length and location of the supplemented telomere sequences, telomere type (such as species-specific telomere motifs), and the number of supplemented telomeres. Furthermore, to visually demonstrate the telomere supplementation status of the new genome, a density plot of the supplemented telomere sequences at the chromosome ends is generated, showing the distribution and density of telomere sequences at chromosome ends, providing a comprehensive assessment of the telomere supplementation results. This output provides rich data support for further analysis of genome integrity and telomere characteristics.
[0095] Step 4: Result Verification and Visualization
[0096] From the original genome, 100 kb of sequence was extracted from each chromosome's ends, and these sequences served as reference fragments for the original genome. For the new genome, 100 kb of sequence corresponding to the original chromosome ends was extracted from the corresponding positions on each chromosome, and supplemented with telomere sequences to form comparison fragments for the new genome. These fragments were used for subsequent collinearity analysis.
[0097] The GenomeSyn tool (https: / / github.com / jmsong2 / GenomeSyn) was used to compare the collinearity of the original and new genomes. The core of GenomeSyn is to align chromosome end sequences and generate collinearity plots, showing similarities and differences between the two. To focus on telomere complementation, we only analyze the collinearity of chromosome ends that can complement telomeres, generating a set of collinearity plots and outputting the following information:
[0098] 1. Telomere information of the original genome: including the length and distribution position of the telomere sequence.
[0099] 2. Telomere information of the new genome after telomere completion: length, position, type (telomere motif), number, etc. of the newly added telomere sequences.
[0100] like Figure 3 As shown, Figure 3 A diagram showing the workflow of TeloComp is shown. Figure 3 Part A is the process of inputting genomic data, performing read alignment, filtering, and extracting reads with corresponding coverage; Figure 3 Part B is the process of assembling, reading and polishing. Figure 3 Part C is the process of extracting telomere sequences and integrating them into the original genome to output a telomere-complemented genome; Figure 3 Part D in the figure outputs telomere position, length, and number information and visualizes the results, including the process of telomere density and collinearity analysis.
[0101] The following combination Figure 4 Visualization of mulberry genome data.
[0102] Figure 4 Part A shows a collinear alignment of the telomere region of mulberry. Sequences are identical over a 100 kb region on the chromosome, and the extended region represents the complementary telomeric sequence. Telomeric repeat counts are calculated using a 100 bp window. Figure 4 The orange and cyan colors in part A represent the right and left ends of the chromosomes, respectively (only the chromosome ends with complementary telomere sequences are shown in the figure).
[0103] Figure 4 Part B in the figure shows the alignment of the sequences, showing an IGV screenshot of the ONT reads supporting the right end of mulberry chromosome 2. The ONT reads supporting the 150kb region of the right end of mulberry chromosome 2 are separated from the original genome and the supplementary telomeric genome sequences by a red dotted line (86,298,228bp).
[0104] Figure 4 Part C is the telomere density map, i.e., the telomere density statistics of the 30kb complemented telomere region of mulberry chromosome 2. The original genomic region (left) and the complemented telomere region are distinguished by a red vertical line. The telomere density is calculated using a 100bp window.
[0105] from Figure 4As can be seen, TeloComp significantly complemented telomere sequences at multiple chromosome ends in the mulberry genome. In particular, at the right end of chromosome 2, where no telomere sequence had been assembled in the original genome, TeloComp successfully completed a 20kb stretch of DNA, including approximately 18kb of telomeric repeats. This improvement not only improved genome integrity but also provided high-quality genome sequences for subsequent analysis. Collinearity plots and telomere length comparison plots clearly demonstrate the number, distribution, and length of complemented telomeres, providing important insights for subsequent genome structure adjustments and evolutionary studies.
[0106] The overall process can be summarized as follows:
[0107] 1. Data preprocessing: Screen candidate reads covering telomeric regions.
[0108] 2. Sequence construction: Assemble and polish the telomere draft sequence.
[0109] 3. Telomere integration: Accurately add effective telomere sequences to the original genome.
[0110] 4. Result output: Verify completeness through collinearity analysis and density statistics.
[0111] This process integrates multi-source sequencing data to make up for the shortcomings of traditional tools in supplementing telomere regions, and ultimately achieves high-integrity telomere-to-telomere (T2T) assembly of chromosomes.
[0112] The present invention provides a toolkit and method for supplementing chromosome telomeres in the T2T genome, which can not only accurately identify telomere sequences in the genome, but also use multiple sequencing data (such as ONT ultra-long reads, PacBioHiFi reads, and high-throughput short reads) to supplement unidentified regions, thereby achieving complete assembly of the telomeric region of the genome.
[0113] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed by the present invention should be covered by the scope of protection of the present invention.
Claims
1. A kit for replenishing chromosome telomeres in a T2T genome, characterized in that: include: An alignment and screening module is used to align the long-read sequencing sequence of the genome to be tested with the reference genome to obtain a primary alignment result, screen the primary alignment result to obtain the original genome in which the genome to be tested matches the reference genome, and a first out-of-range sequence that exceeds the end of the chromosome of the reference genome, and trim the first out-of-range sequence based on a preset coverage; The genome assembly and polishing module is used to determine the relevant fragments of read matching based on the overlapping relationship between the fragments in the trimmed first overrange sequence, splice the relevant fragments, construct the draft genome, and polish the draft genome to improve the accuracy of the assembled sequence; A new genome generation module is used to align the polished draft genome with the original genome to obtain a secondary alignment result. Based on the secondary alignment result, the draft genome is distinguished from the sequences belonging to the original genome and the second out-of-range sequence that extends beyond the end of the chromosome. The start and end position information of the second out-of-range sequence is determined, the start and end position information is matched with the original genome, the telomere sequence is extracted, and the telomere sequence is pasted back to the original genome to obtain a new genome. The visualization output module is used to obtain a telomere sequence analysis report based on the original genome, the new genome, and the specific telomere motifs of the species corresponding to the genome to be tested.
2. A method for supplementing chromosome telomeres in a T2T genome, characterized in that: The method is implemented using the tool kit for supplementing chromosome telomeres in a T2T genome according to claim 1, comprising: Aligning the long-read sequencing sequence of the genome to be tested with the reference genome to obtain a primary alignment result, screening the primary alignment result to obtain an original genome in which the genome to be tested matches the reference genome, and a first out-of-range sequence that extends beyond the end of the chromosome of the reference genome, and trimming the first out-of-range sequence based on a preset coverage; According to the overlapping relationship between the fragments in the trimmed first overrange sequence, the relevant fragments of the read matching are determined, the relevant fragments are spliced, the draft genome is constructed, and the draft genome is polished to improve the accuracy of the assembled sequence; The polished draft genome is aligned with the original genome to obtain a secondary alignment result. Based on the secondary alignment result, the sequences in the draft genome that belong to the original genome and the second out-of-range sequence that extends beyond the end of the chromosome are distinguished. The start and end position information of the second out-of-range sequence is determined, and the start and end position information is matched with the original genome. The telomere sequence is extracted and pasted back into the original genome to obtain a new genome. A telomere sequence analysis report is obtained based on the specific telomere motifs of the original genome, the new genome, and the species corresponding to the genome to be tested.
3. The method for supplementing chromosome telomeres in a T2T genome according to claim 2, characterized in that: The screening of the initial alignment results to obtain a first out-of-range sequence beyond the end of the reference genome chromosome includes: The chromosome ends containing at least one telomere marker sequence corresponding to the species to be detected and whose coverage reaches a preset threshold are screened, and the filtered readings are output according to the left and right ends of the chromosomes to obtain the first out-of-range sequence.
4. The method for supplementing chromosome telomeres in a T2T genome according to claim 2, characterized in that: The method of splicing the relevant fragments to construct a draft genome comprises: Calculate the overlap relationship between all read fragments, generate an overlap graph, and determine the connection order and direction of the fragments; Optimize the paths of overlapping graphs, eliminate redundant connections and erroneous nodes, integrate fragments into continuous genome sequences, and mark duplicate and uncertain regions; Repetitive regions are removed or isolated through preset algorithms to reduce redundancy and errors caused by repetitive sequences during the assembly process; Outputs a draft genome containing uncertain regions.
5. The method for supplementing chromosome telomeres in a T2T genome according to claim 2, characterized in that: The step of polishing the draft genome sequence comprises: Aligning short-read data from Qualcomm sequencing to the draft genome improves local sequence accuracy by statistically analyzing alignment depth and variation frequency. Correction of long-range structural errors using long-read sequencing sequences; Filter low coverage areas and artifact sequences; The draft genome was iteratively optimized based on the alignment results of Qualcomm sequencing short-read data and long-read sequencing sequences to improve the consistency of the genome sequence and complete the polishing of the draft genome sequence.
6. The method for supplementing chromosome telomeres in a T2T genome according to claim 2, characterized in that: The step of matching the start and end position information with the position information of the left and right ends of the original genome to extract the telomere sequence includes: If the position information matches and the directions are consistent, the matching sequence is selected as the telomere sequence; If the position information matches but the directions are inconsistent, the reverse complementary matching sequence is selected as the telomere sequence.
7. The method for supplementing chromosome telomeres in a T2T genome according to claim 2, characterized in that: Pasting the telomere sequence back into the original genome to obtain a new genome comprises: Before pasting the telomere sequence back to the original genome, determine whether the ratio of the telomere sequence length to the supplemented sequence length reaches a preset ratio threshold; If the preset ratio threshold is reached, the telomere sequence is pasted back to the original genome to form a new genome with supplementary telomeres.
8. The method for supplementing chromosome telomeres in a T2T genome according to claim 7, characterized in that: If the preset ratio threshold is reached, the telomere sequence is pasted back to the original genome, including: According to the alignment direction of the draft genome sequence and the left and right end information of the original chromosome, the telomere sequence is spliced to the corresponding position.
9. The method for supplementing chromosome telomeres in a T2T genome according to claim 2, characterized in that: The telomere sequence analysis report is obtained based on the specific telomere motifs of the original genome, the new genome, and the species corresponding to the genome to be tested, including: Cut sequence reads of a preset length at the ends of the original genome; Cut a region of the same preset length from the end of the new genome and add the extracted telomere sequence; The processed original genome and the new genome are aligned, the telomere density is calculated based on the species-specific telomere motif, a collinearity comparison diagram of the original genome and the new genome is obtained, and telomere sequence statistics are output.
10. The method for supplementing chromosome telomeres in a T2T genome according to claim 9, characterized in that: The telomere sequence statistical information includes the telomere information of the original genome and the telomere information of the new genome after the telomeres are supplemented; The telomere information of the original genome at least includes the length and distribution position of the telomere sequence; The telomere information of the new genome after the telomeres are supplemented includes at least the length, position, type and number of the newly added telomere sequences.