Method for calculating length of autosomal telomere of crowd

Through the third-generation sequencing technology, the accuracy and efficiency of telomere length measurement in the existing technology are solved, and efficient and accurate telomere length measurement is achieved, which is suitable for large-scale sample analysis.

CN120544679APending Publication Date: 2025-08-26FIFTH AFFILIATED HOSPITAL OF ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510623012.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Existing telomere length measurement methods such as Southern blot hybridization technology are cumbersome, have low throughput and are difficult to meet the needs of large-scale clinical applications, while the read length of the second-generation sequencing technology leads to limited accuracy of measurement results.

Method used

Using the third-generation sequencing technology, the telomere sequence is identified and compared by obtaining the long fragment data of the third-generation sequencing gene of the population, grouping it according to the chromosome and sequence directions, and the minimum starting and maximum termination positions of the telomere sequence are calculated, and the advantages of long reading and length spanning the telomere region is used, and accurate calculations are carried out in combination with custom scripts.

Benefits of technology

It realizes efficient and accurate telomere length measurement, simplifies operational processes, reduces costs, is suitable for large-scale sample analysis, and improves the reliability of measurement results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544679A_ABST
    Figure CN120544679A_ABST
Patent Text Reader

Abstract

The invention discloses a method for calculating the length of autosomal telomeres of a crowd. The method comprises the following steps: S1, acquiring three-generation sequencing gene long fragment data of different age groups of the crowd; s2, identifying a long fragment containing a telomere sequence; s3, comparing the long fragment data of the third-generation sequencing gene in the step S1 to a reference genome; s4, grouping in chromosome and sequence directions according to a comparison result; s5, identifying and extracting positions of telomere sequences in the same chromosome and direction; sequencing according to the front and back positions of the telomere sequence on the chromosome, and determining a minimum starting position and a maximum ending position; according to the method, the advantage of long reading length of third-generation sequencing is utilized, the whole telomere area can be directly crossed, deviation caused by reading length of second-generation sequencing is avoided, the measurement result is more accurate and reliable, the efficiency is high, and the cost is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of telomere (molecule) technology, and in particular to a method for calculating the length of autosomal telomeres in a population. Background Art

[0002] Telomeres are specialized DNA-protein complexes at the ends of linear chromosomes in eukaryotes, composed of hundreds to thousands of repeated DNA sequences (for example, the telomere repeat sequence in humans is TTAGGG) and a series of binding proteins. Telomeres protect chromosome ends from degradation, fusion, and recombination, maintaining genomic integrity and stability, and playing important roles in processes such as cellular aging, apoptosis, and tumorigenesis. Telomere length is an important indicator of cell replication history and aging status. Normally, telomeres shorten with each cell division. When telomeres shorten to a certain extent, cells cease dividing and enter a state of senescence or apoptosis. Abnormal changes in telomere length are closely associated with the development and progression of various diseases, such as cancer, cardiovascular disease, and neurodegenerative diseases. Therefore, accurate measurement of telomere length is important for early diagnosis, prognosis, and personalized treatment of diseases. Telomere length can vary significantly among different populations due to genetic and environmental factors.

[0003] The current telomere hypothesis posits that telomeres are relatively long at birth due to active cell division during embryonic and infant stages. As children grow, telomeres shorten more rapidly due to the high cell division rate. However, telomere length during this period remains sufficient to maintain normal cellular function. Telomere length gradually decreases in early adulthood. Although the rate of cell division slows, the cumulative number of cell divisions causes telomeres to continue shortening. During this period, lifestyle changes (such as diet, exercise, and stress management) may begin to have a significant impact on telomere length. Telomere shortening continues in middle age, but at a slower pace. This is due to a decrease in the frequency of cell division and the activation of cellular repair mechanisms. However, long-term lifestyle habits and external environmental factors, such as smoking, pollution, and chronic stress, may accelerate telomere shortening. In older age, telomere length decreases significantly, approaching the critical threshold for cellular aging. Telomere shortening during this period is associated with a variety of age-related diseases, such as cardiovascular disease, diabetes, and certain types of cancer. Extreme telomere shortening may lead to cellular dysfunction and age-related health problems. Globally, research on telomere length has made considerable progress. These studies have shown that telomere length varies significantly across different populations due to genetic and environmental factors. Studying the distribution and standardization of telomere length across populations is important for understanding genetic traits and their role in health and disease. Understanding the average standardization of telomere length across the 21 autosomes in different populations will facilitate applications in genetic research, disease prediction, and personalized medicine.

[0004] Traditional methods for measuring telomere length rely primarily on Southern blotting. This method requires digesting genomic DNA with restriction endonucleases, separating DNA fragments of varying lengths by gel electrophoresis, and then hybridizing with probes complementary to telomeric repeat sequences to measure telomere length. However, Southern blotting is cumbersome, has low throughput, is time-consuming and labor-intensive, and requires large amounts of DNA samples, making it difficult to meet the demands of large-scale clinical applications.

[0005] In recent years, the development of next-generation sequencing (NGS) technology has provided new insights into telomere length measurement. NGS can perform high-throughput sequencing of millions to billions of DNA fragments, thereby obtaining a large amount of genomic information. By analyzing the number and distribution of reads of telomeric repeat sequences in the sequencing data, telomere length can be indirectly inferred. However, the read length of NGS is relatively short (usually only a few hundred bases), making it difficult to span the entire telomeric region, which limits the accuracy of the measurement results.

[0006] Compared with second-generation sequencing technology, third-generation sequencing technology (TGS), also known as long-read sequencing technology, can produce ultra-long read sequences (usually thousands to hundreds of thousands of bases), making it possible to span the entire telomere region and providing a new technical means for more accurate measurement of telomere length. Currently, commonly used third-generation sequencing technologies mainly include nanopore sequencing technology from Oxford Nanopore Technologies (ONT) and single-molecule real-time sequencing technology from Pacific Biosciences (PacBio). Summary of the Invention

[0007] Therefore, based on the above background, the present invention provides a method for calculating the length of autosomal telomeres in a population.

[0008] The technical solution provided by the present invention is

[0009] The method for calculating the length of autosomal telomeres in the human population includes the following steps:

[0010] S1: Obtain long fragment data of third-generation sequencing genes of different age groups;

[0011] S2: Identify long fragments containing telomere sequences from the long fragment data of the third-generation sequencing gene in step S1;

[0012] S3: Align the long fragment data of the third-generation sequencing gene in step S1 to the reference genome;

[0013] S4: grouping by chromosome and sequence direction based on the alignment results;

[0014] S5: Identify the location where telomere sequences on the same chromosome and orientation are extracted;

[0015] Sort the telomere sequences according to their front and back positions on the chromosome to determine the minimum starting position and the maximum ending position;

[0016] The telomere length of each chromosome was calculated based on the difference between the maximum and minimum positions.

[0017] Furthermore, the long fragment data of the third-generation sequencing gene in step S1 can be obtained from the platform PacBio or Oxford Nanopore, or the database NCBI;

[0018] Alternatively, DNA can be extracted from peripheral blood samples of the test population and then subjected to third-generation whole genome sequencing.

[0019] Furthermore, in step S1, Filtlong software is used to filter the long fragment data of the third-generation sequencing gene to remove low-quality reads and adapter sequences.

[0020] Furthermore, in step S2, the recognition results are filtered. Specifically, a specific threshold is set to filter out possible false positive results.

[0021] Furthermore, the threshold is the number of repetitions or the sequence matching ratio.

[0022] Furthermore, in step S3, Minimap2 or BWA-MEM is used to align the long fragment data of the third-generation sequencing gene in step S1 to the reference genome.

[0023] Furthermore, in step S5, a custom script gatherpos.pl was used to identify and extract the positions of telomere sequences on the same chromosome and direction;

[0024] The custom script telo.pl was used to sort the telomere sequences according to their front and back positions on the chromosome to determine the minimum starting position and the maximum ending position;

[0025] The telomere length of each chromosome was calculated based on the difference between the maximum and minimum positions using a custom script tique.pl.

[0026] Furthermore, the results obtained in step S4 according to the following steps are grouped according to chromosome and sequence direction:

[0027] 1) For the long fragments containing telomeric sequences identified in step S2, clustering the subtelomeric regions of the telomeric reads to generate a unique consensus anchor sequence;

[0028] 2) Mapping subtelomeric anchor sequences to chromosome ends based on T2T-CHM13 assembly results;

[0029] 3) Output telomere sequence information;

[0030] 4) Mapping the telomere sequence identified in step 3) to the chromosome position aligned in step S3, and converting the starting position of the telomere sequence obtained in step 3) on the third-generation sequencing long fragment sequence to the reference genome position in step S3.

[0031] Furthermore, the specific formula for calculating the telomere length of each chromosome according to the difference between the maximum position and the minimum position in step S5 is as follows:

[0032] Telomere length = MAX_END - MIN_START

[0033] In the formula, MAX_END represents the maximum end position of all telomere sequence alignments on the chromosome in each group, and MIN_START represents the minimum start position of all telomere sequence alignments on the chromosome in each group.

[0034] The above technical solution has the following beneficial effects:

[0035] High accuracy: Taking advantage of the long read length of third-generation sequencing, it can directly span the entire telomere region, avoiding the deviation caused by the short read length of second-generation sequencing, and the measurement results are more accurate and reliable.

[0036] High efficiency: It has a high degree of automation, is easy to operate, can quickly process large amounts of data, and is suitable for large-scale sample analysis.

[0037] Low cost: No need to perform tedious experiments such as Southern blot hybridization, which can effectively reduce experimental costs.

[0038] Telomere length is closely related to the occurrence and development of various diseases. The present invention calculates the telomere length of a population and can prevent or diagnose diseases that may frequently occur in this population based on the calculation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Attachment Figure 1 This is a result diagram of determining the minimum starting position and maximum ending position of the telomere sequence corresponding to step S5 in Example 1.

[0040] Attachment Figure 2 This is the relationship between age and telomere length of the population in Example 1.

[0041] Attachment Figure 3 The figure is a scatter plot of the average telomere length of males and females at different ages in the population of Example 1. DETAILED DESCRIPTION

[0042] To make the objectives, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0043] A method for calculating the length of autosomal telomeres in a population, characterized by comprising the following steps:

[0044] S1: Obtain long fragment data of third-generation sequencing genes of different age groups;

[0045] The long fragment data of the third-generation sequencing gene in this step can be obtained from the platform PacBio or Oxford Nanopore, or the database NCBI;

[0046] Alternatively, DNA can be extracted from peripheral blood samples of the test population and then subjected to third-generation whole genome sequencing.

[0047] In this step, Filtlong software is used to filter the long fragment data of the third-generation sequencing gene to remove low-quality reads and adapter sequences.

[0048] S2: Identify long fragments containing telomere sequences within the long fragment data from the third-generation sequencing (NGS) sequence in step S1. Telomere sequences have specific repeat sequences in many species. For example, in humans, telomeres are composed of TTAGGG repeats. Use a custom script to scan long fragments and identify these repeat sequences.

[0049] In this step, the recognition results are filtered. Specifically, a specific threshold is set to filter out possible false positive results. The threshold is the number of repetitions or the sequence matching ratio.

[0050] S3: Align the long fragment data from the third-generation sequencing (NGS) sequence in step S1 to the reference genome, generating an alignment result file (e.g., in SAM / BAM format). Minimap2 or BWA-MEM is used in this step to align the long fragment data from the third-generation sequencing (NGS) sequence in step S1 to the reference genome. The alignment results are then checked for quality indicators, such as alignment rate and mismatch rate, to ensure alignment accuracy. If necessary, adjust alignment parameters or use a different reference genome for alignment.

[0051] S4: grouping by chromosome and sequence direction based on the alignment results;

[0052] In this step, the results obtained from the following steps are grouped by chromosome and sequence direction:

[0053] 1) For the long fragments containing telomeric sequences identified in step S2, clustering the subtelomeric regions of the telomeric reads to generate a unique consensus anchor sequence;

[0054] 2) Mapping subtelomeric anchor sequences to chromosome ends based on T2T-CHM13 assembly results;

[0055] 3) Output telomere sequence information;

[0056] 4) Map the telomere sequences identified in step 3) to the chromosome positions aligned in step S3. The starting positions of the telomere sequences obtained in step 3) on the long fragments of the third-generation sequencing are converted to the reference genome positions in step S3. This is equivalent to annotating these genes to the reference genome to determine their positions.

[0057] In this step: Read the direction information: In the alignment results, the direction information of each long fragment is usually represented by "+" or "-", indicating its alignment direction with the reference sequence.

[0058] Classification and organization: Based on the direction information, long fragments are divided into two categories: sense strands and antisense strands.

[0059] S5: Identify the location where telomere sequences on the same chromosome and orientation are extracted;

[0060] Sort the telomere sequences according to their front and back positions on the chromosome to determine the minimum starting position and the maximum ending position;

[0061] The telomere length of each chromosome was calculated based on the difference between the maximum and minimum positions. The specific calculation formula is as follows: Telomere length = MAX_END - MIN_START; where MAX_END represents the maximum end position of all telomere sequence alignments on the chromosome in each group, and MIN_START represents the minimum start position of all telomere sequence alignments on the chromosome in each group.

[0062] One implementation method is to use a custom script gatherpos.pl to identify and extract the positions of telomere sequences on the same chromosome and direction;

[0063] The specific example of customizing the script gatherpos.pl is as follows:

[0064]

[0065] The custom script telo.pl was used to sort the telomere sequences according to their front and back positions on the chromosome to determine the minimum starting position and the maximum ending position;

[0066] The specific custom script telo.pl is demonstrated below:

[0067]

[0068]

[0069]

[0070] The telomere length of each chromosome was calculated based on the difference between the maximum and minimum positions using a custom script tique.pl.

[0071] The code example of defining the script tique.pl is as follows:

[0072]

[0073]

[0074] In the exemplary embodiment, the specific calculation operations of steps S2 to S5 can be seen below:

[0075] (1) The long-fragment whole-genome data from the third-generation sequencing were aligned to the reference genome. Minimap2 software was used for alignment, using the command "-z 600200 -x map-ont".

[0076]

[0077] (2) Use the script telocal.Py to identify and analyze telomere sequences in the long-fragment whole-genome data of third-generation sequencing.

[0078] Use default parameters.

[0079]

[0080] The specific demonstration of the script telocal.Py is as follows:

[0081]

[0082]

[0083]

[0084]

[0085]

[0086]

[0087]

[0088]

[0089]

[0090]

[0091]

[0092]

[0093]

[0094]

[0095]

[0096]

[0097]

[0098]

[0099] If the script telocal.Py is not used for this step, the specific implementation steps include:

[0100] 2.1 Identify long fragments containing telomere sequences in third-generation sequencing.

[0101] 2.2 Clustering of subtelomeric regions of telomeric reads to generate unique consensus anchor sequences.

[0102] 2.3 Mapping subtelomeric anchor sequences to chromosome ends based on T2T-CHM13 assembly results

[0103] 2.4 Output telomere sequence information.

[0104] (3) The telomere sequence identified in the second step is mapped to the chromosome position aligned in the first step. The result of the second step contains the starting position of the telomere sequence on the long fragment sequence of the third-generation sequencing. Its position on the long fragment sequence of the third-generation sequencing is converted to the reference genome position in the first step.

[0105] (4) Group the results of the third step according to chromosome and sequence direction.

[0106] (5) Identify the position of each telomere sequence on the chromosome in each group and record them as START1, END1; START2, END2; START3, END3.....

[0107] (6) Sort by the front and back position of the telomere sequence on the chromosome

[0108] Calculate according to the following formula:

[0109] Telomere length = MAX_END - MIN_START

[0110] In the formula, MAX_END represents the maximum end position of all telomere sequence alignments on the chromosome in each group.

[0111] MIN_START represents the minimum starting position of all telomere sequence alignments on the chromosome in each group.

[0112] Example 1: The present invention was used to calculate the autosomal telomere length of the Central Plains population. In this example, DNA extracted from peripheral blood samples of 37 healthy local people in Henan, Central Plains, was subjected to third-generation whole-genome sequencing to obtain long-fragment whole-genome sequencing data (BAM format).

[0113] Specifically, library construction and third-generation sequencing: DNA library was constructed using ONT's Ligation Sequencing Kit, and sequencing was performed on the PromethION sequencing platform to obtain sequencing data in Bam format.

[0114] Sequencing data quality control and filtering: Filtlong software was used to filter the sequencing data to remove low-quality reads and adapter sequences.

[0115] Sequence alignment and telomere sequence identification: Minimap2 software was used to align the filtered sequencing data to the human reference genome GRCh38, and TRF software was used to identify the telomeric repeat sequence TTAGGG in the sequencing data.

[0116] Telomere sequence positioning and grouping: Based on the alignment results, the identified telomere repeat sequences are grouped according to chromosome and direction.

[0117] Identify the location where telomere sequences on the same chromosome and orientation were extracted;

[0118] Sort the telomere sequences according to their front and back positions on the chromosome to determine the minimum starting position and the maximum ending position;

[0119] The telomere length of each chromosome was calculated based on the difference between the maximum position and the minimum position. Figure 1 To the attached Figure 3 , attached Figure 1 To the attached Figure 3 The telomere length is the average telomere length of each chromosome.

[0120] The present invention and its embodiments are described above. This description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs structures and embodiments similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.

Claims

1. A method for calculating the length of autosomal telomeres in a population, characterized in that: The steps include: S1: Obtain long fragment data of third-generation sequencing genes of different age groups; S2: Identify long fragments containing telomere sequences from the long fragment data of the third-generation sequencing gene in step S1; S3: Align the long fragment data of the third-generation sequencing gene in step S1 to the reference genome; S4: grouping by chromosome and sequence direction based on the alignment results; S5: Identify the location where telomere sequences on the same chromosome and orientation are extracted; Sort the telomere sequences according to their front and back positions on the chromosome to determine the minimum starting position and the maximum ending position; The telomere length of each chromosome was calculated based on the difference between the maximum and minimum positions.

2. The method for calculating the telomeres of the human autosome according to claim 1, wherein: The long fragment data of the third-generation sequencing gene in step S1 can be obtained from the platform PacBio or Oxford Nanopore, or the database NCBI; Alternatively, DNA can be extracted from peripheral blood samples of the test population and then subjected to third-generation whole genome sequencing.

3. The method for calculating the telomere length of autosomes in a population according to claim 1, wherein: In step S1, Filtlong software is used to filter the long fragment data of the third-generation sequencing gene to remove low-quality reads and adapter sequences.

4. The method for calculating the telomere length of autosomes in a population according to claim 1, wherein: In step S2, the recognition results are filtered. Specifically, a specific threshold is set to filter out possible false positive results.

5. The method for calculating the average standard of autosomal telomere length of a population according to claim 4, characterized in that: The threshold is the number of repetitions or the sequence matching ratio.

6. The method for calculating the telomere length of an autosome in a population according to claim 1, wherein: In step S3, Minimap2 or BWA-MEM is used to align the long fragment data of the third-generation sequencing gene in step S1 to the reference genome.

7. The method for calculating the length of autosomal telomeres in a population according to claim 1, wherein: The results obtained in step S4 according to the following steps are grouped according to chromosome and sequence direction: 1) For the long fragments containing telomeric sequences identified in step S2, clustering the subtelomeric regions of the telomeric reads to generate a unique consensus anchor sequence; 2) Mapping subtelomeric anchor sequences to chromosome ends based on T2T-CHM13 assembly results; 3) Output telomere sequence information; 4) Mapping the telomere sequence identified in step 3) to the chromosome position aligned in step S3, and converting the starting position of the telomere sequence obtained in step 3) on the third-generation sequencing long fragment sequence to the reference genome position in step S3.

8. The method for calculating the length of autosomal telomeres in a population according to claim 1, wherein: The specific formula for calculating the telomere length of each chromosome based on the difference between the maximum position and the minimum position in step S5 is as follows: Telomere length = MAX_END - MIN_START In the formula, MAX_END represents the maximum end position of all telomere sequence alignments on the chromosome in each group, and MIN_START represents the minimum start position of all telomere sequence alignments on the chromosome in each group.