Method, device and storage medium for short tandem repeat typing based on third generation sequencing data

By precisely processing third-generation sequencing data, including read repetition alignment and SNV analysis, the problem of inaccurate estimation of the number of repetitions in long read data has been solved, achieving higher genotyping accuracy and greater depth of downstream analysis.

CN122117018APending Publication Date: 2026-05-29SHENZHEN ANJI KANGER MEDICAL LAB

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN ANJI KANGER MEDICAL LAB
Filing Date
2026-02-14
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies are not accurate enough in estimating the number of repetitions when processing long read data, and fail to fully consider the impact of single nucleotide variants on genotyping, resulting in large deviations in genotyping results and limiting the depth and breadth of downstream analysis.

Method used

A method based on third-generation sequencing data was adopted, which improved data quality and accurately estimated the number of repetitions by steps such as target read extraction, read overlap alignment and filtering, copy number calculation of repeat units, and short tandem repeat sequence genotyping confirmation, combined with SNV mutation site analysis.

Benefits of technology

It achieves more accurate genotyping results, reduces error rates, and is applicable to third-generation sequencing platform data with different read lengths, especially high-fidelity long-read data, maximizing the advantages of third-generation sequencing read length.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122117018A_ABST
    Figure CN122117018A_ABST
Patent Text Reader

Abstract

The application discloses a method and device for short tandem repeat sequence typing based on third-generation sequencing data and a storage medium. The method for short tandem repeat sequence typing based on third-generation sequencing data provided by the application first extracts all read segments completely covering the short tandem repeat sequence region according to a short tandem repeat sequence site directory; and filters the read segments by average alignment quality to remove low-quality read segments; then calculates the copy number of the short tandem repeat sequence in each read segment, and combines mutation site information of the short tandem repeat sequence to confirm short tandem repeat sequence typing. The method provided by the application can accurately calculate the repetition number of the short tandem repeat sequence, and combines the mutation site to determine the short tandem repeat sequence typing, so that more accurate genotyping results can be provided. Moreover, the method provided by the application is suitable for third-generation sequencing data of different sequencing platforms, has strong applicability, and can maximize the advantages of the third-generation sequencing read length.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of short tandem repeat genotyping technology, and in particular to a method, apparatus and storage medium for short tandem repeat genotyping based on third-generation sequencing data. Background Technology

[0002] Short tandem repeats (STRs), comprising 3% of the human genome and exhibiting copy number variations in their 1-13 bp motifs, as well as non-classical motif insertions (such as ataxia-inducing variants in the DAB1 gene), are closely associated with over 60 diseases, including Huntington's disease and Fragile X syndrome, as well as autism and cancer. They also participate in gene expression regulation, such as transcription factor binding sites. Furthermore, the copy number and genotype of repeat sequences are correlated with disease severity.

[0003] Traditional short-read technologies cannot fully cover STR regions, while the long-read advantage of third-generation sequencing can overcome this limitation. However, it requires specialized software to process its data characteristics, such as accurately identifying repeat units, quantifying copy numbers, and performing genotyping, in order to meet the needs of disease mechanism research, genetic analysis of complex traits, and clinical STR variant detection. Therefore, third-generation sequencing STR analysis software has become a core tool for bridging technological advantages and research applications.

[0004] Currently, several techniques exist for STR genotyping of long-read sequencing data. However, these existing techniques generally have shortcomings. Some tools, such as LongTR and STRdust, do not accurately estimate repeat counts when processing long-read data, leading to significant biases in genotyping results. Furthermore, most existing techniques fail to adequately consider the impact of single nucleotide variants (SNVs) on genotyping, limiting the depth and breadth of downstream analysis.

[0005] Therefore, how to accurately estimate the number of repetitions from high-fidelity long-read data and improve the accuracy of STR genotyping remains a key research focus and challenge in the field of STR genotyping technology. Summary of the Invention

[0006] The purpose of this application is to provide a new method, apparatus, and storage medium for typing short tandem repeat sequences based on third-generation sequencing data.

[0007] To achieve the above objectives, this application adopts the following technical solution:

[0008] The first aspect of this application discloses a method for genotyping short tandem repeat sequences based on third-generation sequencing data, comprising the following steps:

[0009] The target read extraction step includes traversing the alignment file based on the short tandem repeat sequence site directory to extract all reads that completely cover the short tandem repeat sequence regions in the third-generation sequencing data to be analyzed; wherein, the short tandem repeat sequence site directory is a directory composed of existing known short tandem repeat sequence sites or short tandem repeat sequence sites of interest obtained by downloading; the alignment file is the alignment file between the third-generation sequencing data to be analyzed and the reference genome;

[0010] The read re-alignment and filtering steps include obtaining the average alignment quality of the reads extracted in the target read extraction step; for reads with an average alignment quality lower than the threshold, the short tandem repeat sequence region corresponding to the read and the nearby sequence in the reference genome are selected for re-alignment with the read, and the average alignment quality of the re-alignment is obtained; if the average alignment quality of the re-alignment is still lower than the threshold, the read is filtered out and removed, and the retained reads are used for subsequent analysis.

[0011] The steps for calculating the copy number of repeating units include processing the reads retained in the read re-alignment and filtering steps as follows: 1) Determine whether the read is an inverted complementary read. If so, invert the read based on the base complementary pairing principle. The inverted complementary sequence of the read is used for subsequent analysis; 2) At the starting point of the short tandem repeat sequence region where the read matches, move the length of the repeating unit of the short tandem repeat sequence for each alignment and matching until no matching repeating unit can be found; 3) Record the number of times the current read can continuously match repeating units, which is taken as the copy number of repeating units in the read.

[0012] The short tandem repeat sequence genotyping confirmation steps include the following processing of the reads retained in the read re-alignment and filtering steps: 1) Obtaining the base information of all mutation sites on the reads; 2) Filtering out reads with only one mutation site; 3) For all reads matching the same short tandem repeat sequence site region, taking the intersection of their mutation sites, and filtering out reads with only one mutation site in the intersection; 4) Genotyping the short tandem repeat sequence based on the haplotype of the mutation site in the read;

[0013] The results output steps include the location information, repeat unit, and reference repeat number of the short tandem repeat sequence corresponding to the output read, as well as the copy number result of the repeat unit calculation step and the short tandem repeat genotyping result of the short tandem repeat sequence genotyping confirmation step; wherein, the reference repeat number refers to the number of times the short tandem repeat sequence is repeated on the reference genome.

[0014] It should be noted that the method for short tandem repeat sequence genotyping based on third-generation sequencing data in this application can provide more accurate genotyping results through precise copy number calculation of repeat units and haplotype analysis of SNV mutation sites, resulting in higher accuracy and reduced error rate. Furthermore, the method in this application is applicable to data processing of third-generation sequencing platforms with different read lengths, especially suitable for high-fidelity long-read data, such as PacBio HiFi data and ONT R10 data, which can maximize the advantages of third-generation sequencing read length.

[0015] In one implementation of this application, the threshold for the average alignment quality of the segment re-matching and filtering steps is 20.

[0016] It should be noted that the average alignment quality threshold reflects the probability of an alignment error. The threshold is 20, i.e., MQ20, which corresponds to a 1% error rate, meaning there is a 1% chance that the read segment is misaligned. This application improves the accuracy of subsequent analysis by filtering out read segments with an average alignment quality lower than MQ20.

[0017] In one implementation of this application, the sequence near the short tandem repeat sequence region in the read segment re-alignment and filtering steps is the sequence 1kb upstream and 1kb downstream of the short tandem repeat sequence region.

[0018] It should be noted that this application improves the average alignment quality of read segments by realigning them to a defined short tandem repeat sequence region and its upstream and downstream sequences by 1kb. If the threshold requirement is still not met, the read segment is filtered out.

[0019] In one implementation of this application, during the process of calculating the copy number of the repeating unit by moving the length of the repeating unit of the short tandem repeating sequence for comparison and matching, a maximum of one mismatched base is allowed.

[0020] The second aspect of this application discloses an apparatus for genotyping short tandem repeat sequences based on third-generation sequencing data, including a target read extraction module, a read overlap alignment and filtering module, a repeat unit copy number calculation module, a short tandem repeat sequence genotyping confirmation module, and a result output module;

[0021] The target read extraction module includes a function to extract all reads that completely cover short tandem repeat regions in the third-generation sequencing data to be analyzed, based on a directory of short tandem repeat sites and by traversing the alignment files. The directory of short tandem repeat sites consists of a directory of existing known short tandem repeat sites or short tandem repeat sites of interest obtained from downloads. The alignment files are the alignment files between the third-generation sequencing data to be analyzed and the reference genome.

[0022] The read re-alignment and filtering module includes a module for obtaining the average alignment quality of reads extracted by the target read extraction module. For reads with an average alignment quality lower than a threshold, the module selects the short tandem repeat sequence region corresponding to the read in the reference genome and the sequence in its vicinity to re-align with the read and obtain the average alignment quality of the re-alignment. If the average alignment quality of the re-alignment is still lower than the threshold, the read is filtered out and removed. The remaining reads are used for subsequent analysis.

[0023] The repeat unit copy number calculation module includes the following processing of read segments retained by the read segment re-alignment and filtering module: 1) Determine whether the read segment is an inverted complementary read segment. If so, invert the read segment based on the base complementary pairing principle. The inverted complementary sequence of the read segment is used for subsequent analysis; 2) When the read segment matches the starting point of the short tandem repeat sequence region, move the length of the repeat unit of the short tandem repeat sequence for each alignment and matching until no matching repeat unit can be found; 3) Record the number of times the current read segment can continuously match repeat units, which is taken as the repeat unit copy number of the read segment.

[0024] The short tandem repeat sequence genotyping confirmation module includes the following processing for reads retained by the read re-alignment and filtering module: 1) obtaining the base information of all mutation sites on the read; 2) filtering out reads with only one mutation site; 3) for all reads matching the same short tandem repeat sequence site region, taking the intersection of their mutation sites, and filtering out reads with only one mutation site in the intersection; 4) genotyping the short tandem repeat sequence according to the haplotype of the mutation site of the read;

[0025] The results output module includes the position information, repeat unit, and reference repeat number of the short tandem repeat sequence corresponding to the read segment, as well as the repeat unit copy number result of the repeat unit copy number calculation module and the short tandem repeat sequence typing result of the short tandem repeat sequence typing confirmation module; wherein, the reference repeat number refers to the number of times the short tandem repeat sequence is repeated on the reference genome.

[0026] It should be noted that the apparatus for short tandem repeat genotyping based on third-generation sequencing data in this application is actually implemented by each module separately as described in this application. Therefore, the specific limitations of each module can be referenced in the method for short tandem repeat genotyping based on third-generation sequencing data in this application. For example, the average alignment quality threshold of 20, the use of the short tandem repeat sequence region and its upstream and downstream 1kb sequences for re-alignment, and the allowance of a maximum of one mismatched base during the alignment and matching of repeat units can all refer to the method for short tandem repeat genotyping based on third-generation sequencing data in this application.

[0027] A third aspect of this application discloses an apparatus for genotyping short tandem repeat sequences based on third-generation sequencing data. The apparatus includes a memory and a processor. The memory includes a program for storing data. The processor includes a program for executing the program stored in the memory to implement the method of genotyping short tandem repeat sequences based on third-generation sequencing data as described in this application.

[0028] The fourth aspect of this application discloses a computer-readable storage medium storing a program that can be executed by a processor to implement the method of short tandem repeat genotyping based on third-generation sequencing data.

[0029] Due to the adoption of the above technical solutions, the beneficial effects of this application are as follows:

[0030] This application presents a method for short tandem repeat (STDL) genotyping based on third-generation sequencing data. This method can accurately calculate the number of repetitions of STDL sequences and determine STDL genotyping by combining mutation sites. It can provide more accurate genotyping results. Furthermore, this method is applicable to third-generation sequencing data from different sequencing platforms, has strong applicability, and can maximize the advantages of third-generation sequencing read length. Attached Figure Description

[0031] Figure 1 This is a flowchart of the method for typing short tandem repeat sequences based on third-generation sequencing data in the embodiments of this application;

[0032] Figure 2 This is a structural block diagram of the device for short tandem repeat genotyping based on third-generation sequencing data in the embodiments of this application;

[0033] Figure 3 This is a schematic diagram illustrating the STR typing based on haplotype in an embodiment of this application. Detailed Implementation

[0034] Existing technologies for processing long-read data suffer from insufficient accuracy in estimating the number of repetitions, leading to significant biases in genotyping results. Furthermore, most existing technologies fail to adequately consider the impact of single nucleotide variants (SNVs) on genotyping, limiting the depth and breadth of downstream analysis. This application addresses these issues by optimizing the estimation of repetition counts from high-fidelity long-read data. Through data re-alignment and filtering, data quality is improved, and SNVs are incorporated into the analysis, enhancing the accuracy of genotyping and expanding the possibilities for downstream analysis.

[0035] The core technology of this application is based on the precise identification and analysis of STR sites in long-read sequencing data. Data quality is improved by re-aligning and filtering low-quality data. Repeat counts are estimated by using a sliding window in the STR region and allowing for matching errors. Furthermore, this application integrates single nucleotide variant information into the genotyping process, achieving more accurate genotyping and richer downstream analysis through comprehensive analysis of multiple genetic information.

[0036] Based on the above research and understanding, this application develops a method for short tandem repeat genotyping based on third-generation sequencing data, such as... Figure 1 As shown, the steps include target segment extraction step 11, segment overlap comparison and filtering step 12, repeating unit copy number calculation step 13, short tandem repeat sequence classification confirmation step 14, and result output step 15.

[0037] The target read extraction step 11 includes traversing the comparison file according to the short tandem repeat sequence site directory to extract all reads that completely cover the short tandem repeat sequence region in the third-generation sequencing data to be analyzed; wherein, the short tandem repeat sequence site directory is a directory composed of existing known short tandem repeat sequence sites or short tandem repeat sequence sites of interest obtained by downloading; the comparison file is the alignment file between the third-generation sequencing data to be analyzed and the reference genome.

[0038] In one implementation of this application, the following data are first prepared: third-generation sequencing alignment file, STR locus directory file, human reference genome sequence file, and sample VCF file. Specifically, the hg38 version of the genome is downloaded from the UCSC website (https: / / hgdownload.soe.ucsc.edu / goldenPath / hg38 / bigZips / analysisSet / hg38.analysisSet.fa.gz), or other versions of the human genome are downloaded from other websites. The STR locus directory is downloaded from UCSC (https: / / hgdownload.soe.ucsc.edu / goldenPath / hg38 / bigZips / latest / hg38.trf.bed.gz), or other STR locus directories of interest are selected. The STR locus directory should include at least the following information: the first to third columns are the chromosome number, genome start coordinates, and genome end coordinates, respectively; the last two columns are the repetitive sequence units and copy number, respectively. An example is shown in Table 1 below.

[0039] Table 1

[0040]

[0041] The read re-alignment and filtering step 12 includes obtaining the average alignment quality of the reads extracted in the target read extraction step. For reads with an average alignment quality lower than a threshold, the short tandem repeat sequence region corresponding to the read and the nearby sequence in the reference genome are selected for re-alignment with the read, and the average alignment quality of the re-alignment is obtained. If the average alignment quality of the re-alignment is still lower than the threshold, the read is filtered out and removed, and the retained reads are used for subsequent analysis.

[0042] In one implementation of this application, the alignment file is traversed according to the given STR locus directory, all read segments that completely cover the STR region are extracted, that is, the read segment contains the complete STR region, and the average alignment quality of the read segment is recorded.

[0043] All reads with an average alignment quality below a threshold (typically 20) are re-aligned. The reference genome selected for re-alignment is limited to the vicinity of the STR site (typically 1 kb upstream and downstream). Reads with an average alignment quality above the threshold are retained for further analysis; otherwise, they are filtered out.

[0044] Step 13, which calculates the copy number of repeating units, includes the following processing of the reads retained in the read re-alignment and filtering steps: 1) Determine whether the read is an inverted complementary read. If so, invert the read based on the base complementary pairing principle. The inverted complementary sequence of the read is used for subsequent analysis; 2) At the starting point of the short tandem repeat sequence region where the read matches, move the length of the repeating unit of the short tandem repeat sequence for each alignment and matching until no matching repeating unit can be found; 3) Record the number of times the current read can be continuously aligned and matched with repeating units, which is taken as the copy number of repeating units of the read.

[0045] The short tandem repeat sequence typing confirmation step 14 includes the following processing of the reads retained in the read re-alignment and filtering steps: 1) Obtain the base information of all mutation sites on the reads; 2) Filter out reads with only one mutation site; 3) For all reads that match the same short tandem repeat sequence site region, take the intersection of their mutation sites and filter out reads with only one mutation site in the intersection; 4) Type the short tandem repeat sequence according to the haplotype of the mutation site of the read.

[0046] In one implementation of this application, for example, the base information of SNV mutation sites in all VCF files on the read segment is recorded. A read segment must have at least two or more mutation sites; otherwise, the read segment will be filtered out and will not be processed in the following steps. For all read segments that match the same STR site, the intersection of their VCF sites is taken as the haplotype for determining the STR type. The final intersection must have at least two remaining sites; otherwise, these read segments will be filtered out and will not be processed in the following steps.

[0047] For each STR locus, the STR is genotyped based on haplotype. The genotyping logic is as follows: first, determine the haplotype combination on the reads. For example, some reads contain the haplotype AGT, while others contain ACA. Reads containing the AGT haplotype all have a copy number of 17, while reads containing the ACA haplotype have copy numbers of 24 and 28 in some reads, respectively. The former reads have a higher copy number than the latter, so the final STR genotype is determined to be 17|24. Figure 3 As shown.

[0048] The output step 15 includes the position information, repeat unit, and reference repeat number of the short tandem repeat sequence corresponding to the output read, as well as the copy number result of the repeat unit calculation step and the short tandem repeat sequence typing result of the short tandem repeat sequence typing confirmation step; wherein, the reference repeat number refers to the number of times the short tandem repeat sequence is repeated on the reference genome.

[0049] In one implementation of this application, the output includes the chromosome where the STR is located, the start point, the end point, the repeat unit of the STR reference genome, the reference repeat number, the repeat unit copy number result of the repeat unit copy number calculation step, and the short tandem repeat sequence typing result of the short tandem repeat sequence typing confirmation step, etc.

[0050] Those skilled in the art will understand that all or part of the functions of the methods described above can be implemented in hardware or by computer programs. When all or part of the functions in the above embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to achieve the above functions. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be achieved. Alternatively, when all or part of the functions in the above embodiments are implemented by computer programs, the program can also be stored in a server, another computer, disk, optical disk, flash drive, or portable hard drive, etc., and can be downloaded or copied to the memory of a local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be achieved.

[0051] Therefore, based on the method for short tandem repeat genotyping using third-generation sequencing data in this application, this application proposes a device for short tandem repeat genotyping using third-generation sequencing data, such as... Figure 2 As shown, it includes a target read segment extraction module 21, a read segment re-comparison and filtering module 22, a repeating unit copy number calculation module 23, a short tandem repeat sequence classification confirmation module 24, and a result output module 25.

[0052] The target read extraction module 21 includes a method for extracting all reads that completely cover short tandem repeat (STDL) regions from the third-generation sequencing data to be analyzed, based on a directory of STDL sites and by traversing the alignment file. The STDL site directory is a directory of existing known STDL sites or STDL sites of interest obtained through downloading. The alignment file is an alignment file of the third-generation sequencing data to be analyzed with a reference genome. The read re-alignment and filtering module 22 includes a method for obtaining the average alignment quality of the reads extracted by the target read extraction module and filtering the average alignment quality. For reads with copy number below a threshold, the corresponding short tandem repeat sequence region and nearby sequences in the reference genome are selected for re-alignment with the read, and the average alignment quality of the re-alignment is obtained. If the average alignment quality of the re-alignment is still less than the threshold, the read is filtered out, and the retained reads are used for subsequent analysis. The repeat unit copy number calculation module 23 includes the following processing for read re-alignment and the reads retained by the filtering module: 1) Determine whether the read is an inverted complementary read. If so, the read is inverted based on the base pairing principle. The inverted complementary sequence of the inverted read is used for... For subsequent analysis; 2) When the read segment matches the starting point of the short tandem repeat sequence region, the length of the repeat unit of the short tandem repeat sequence is moved for comparison and matching each time until no repeat unit can be matched; 3) Record the number of times the current read segment can be continuously matched with repeat units, as the copy number of repeat units in the read segment; The short tandem repeat sequence genotyping confirmation module 24 includes the following processing for the read segments retained by the read segment re-alignment and filtering module: 1) Obtain the base information of all mutation sites on the read segment; 2) Filter out read segments with only one mutation site; 3) For read segments matched with the same short tandem repeat sequence, the length of the repeat unit of the short tandem repeat sequence is moved for comparison and matching each time until no repeat unit can be matched; 4) For all reads in the repetitive sequence site region, take the intersection of their mutation sites, and filter out reads with only one mutation site in the intersection; 5) Genotype the short tandem repeat sequence according to the haplotype of the mutation site of the read; The result output module 25 includes the position information, repeat unit, and reference repeat number of the short tandem repeat sequence corresponding to the read, as well as the repeat unit copy number result of the repeat unit copy number calculation module and the short tandem repeat sequence genotyping result of the short tandem repeat sequence confirmation module; where the reference repeat number refers to the number of times the short tandem repeat sequence is repeated on the reference genome.

[0053] Another implementation of this application provides an apparatus for short tandem repeat (STDL) genotyping based on third-generation sequencing data. The apparatus includes a memory and a processor. The memory includes a program for storing data. The processor includes a method for executing the program stored in the memory to implement the following steps: a target read extraction step, comprising traversing a comparison file according to a short STDL site directory to extract all reads from the third-generation sequencing data to be analyzed that completely cover the short STDL region; wherein the short STDL site directory is a directory composed of existing known short STDL sites or short STDL sites of interest obtained through download; the comparison file... The file is the alignment file between the third-generation sequencing data to be analyzed and the reference genome; the read re-alignment and filtering steps include obtaining the average alignment quality of the reads extracted in the target read extraction step; for reads with an average alignment quality lower than the threshold, selecting the corresponding short tandem repeat sequence region in the reference genome and the nearby sequence to re-align with the read, and obtaining the average alignment quality of the re-alignment; if the average alignment quality of the re-alignment is still lower than the threshold, the read is filtered out, and the retained reads are used for subsequent analysis; the repeat unit copy number calculation step includes the following processing of the reads retained in the read re-alignment and filtering steps: 1) Determining whether the read is a copy number of a repeat unit. If it is not a reverse complementary read, then the read is reversed based on the base pairing principle. The reverse complementary sequence of the reversed read is used for subsequent analysis; 2) When the read matches the start of the short tandem repeat sequence region, the length of the repeat unit of the short tandem repeat sequence is moved each time for comparison and matching until no repeat unit can be matched; 3) Record the number of times the current read can be continuously matched with repeat units as the copy number of the repeat unit of the read; The short tandem repeat sequence genotyping confirmation step includes the following processing of the reads retained in the read re-alignment and filtering steps: 1) Obtain the base information of all mutation sites on the read; 2) Filter out reads with only 3) For all reads that match the same short tandem repeat sequence region, take the intersection of their mutation sites and filter out reads with only one mutation site in the intersection; 4) Genotype the short tandem repeat sequence according to the haplotype of the mutation site of the read; The result output steps include the position information, repeat unit, and reference repeat number of the short tandem repeat sequence corresponding to the read, as well as the repeat unit copy number result of the repeat unit copy number calculation step and the short tandem repeat sequence genotyping result of the short tandem repeat sequence genotyping confirmation step; where the reference repeat number refers to the number of times the short tandem repeat sequence is repeated on the reference genome.

[0054] Another implementation of this application also provides a computer-readable storage medium, which includes a program executable by a processor to implement the following method: a target read extraction step, comprising, based on a directory of short tandem repeat sequence sites, traversing a comparison file to extract all reads from the third-generation sequencing data to be analyzed that completely cover short tandem repeat sequence regions; wherein, the directory of short tandem repeat sequence sites is a directory composed of existing known short tandem repeat sequence sites or short tandem repeat sequence sites of interest obtained through download; the comparison file is a comparison file of the third-generation sequencing data to be analyzed and a reference genome; read re-alignment and... The filtering step includes obtaining the average alignment quality of the reads extracted in the target read extraction step. For reads with an average alignment quality lower than a threshold, the corresponding short tandem repeat sequence region in the reference genome and its nearby sequences are selected for re-alignment with the read, and the average alignment quality of the re-alignment is obtained. If the average alignment quality of the re-alignment is still lower than the threshold, the read is filtered out, and the retained reads are used for subsequent analysis. The repeat unit copy number calculation step includes the following processing of the reads retained in the read re-alignment and filtering steps: 1) Determining whether the read is a reverse complementary read; if so, based on base complementarity... The following steps are performed: 1) Reverse the read segment according to the principle, and use the inverse complementary sequence of the reversed read segment for subsequent analysis; 2) When the read segment matches the starting point of the short tandem repeat sequence region, move the length of the repeat unit of the short tandem repeat sequence for each comparison and matching until no repeat unit can be matched; 3) Record the number of times the current read segment can be continuously matched with repeat units, as the copy number of the repeat unit of the read segment; The short tandem repeat sequence genotyping confirmation steps include the following processing of the read segments retained in the read segment re-alignment and filtering steps: 1) Obtain the base information of all mutation sites on the read segment; 2) Filter out read segments with only one mutation site; 3) For all reads matching the same short tandem repeat sequence region, take the intersection of their mutation sites and filter out reads with only one mutation site in the intersection; 4) Genotype the short tandem repeat sequence according to the haplotype of the mutation site of the read; The result output steps include outputting the position information, repeat unit, and reference repeat number of the short tandem repeat sequence corresponding to the read, as well as the repeat unit copy number result of the repeat unit copy number calculation step and the short tandem repeat sequence genotyping result of the short tandem repeat sequence confirmation step; where the reference repeat number refers to the number of times the short tandem repeat sequence is repeated on the reference genome.

[0055] The method and apparatus of this application have the following advantages compared with the prior art:

[0056] 1. Higher accuracy: By using precise duplicate count estimation and SNV haplotype analysis, this invention can provide more accurate genotyping results and reduce the error rate compared to existing technologies.

[0057] 2. Strong data adaptability: It is suitable for data processing of third-generation sequencing platforms with different read lengths, especially for high-fidelity long read data (such as PacBio HiFi data and ONT R10 data), maximizing the advantages of third-generation sequencing read length.

[0058] It is understood that, without changing the overall functionality, the methods and apparatus of this application can use other high-performance string matching and statistical analysis algorithms to replace the existing repetition estimation algorithm, in order to further optimize performance or adapt to specific data characteristics. Furthermore, other classification methods, such as GMM model classification, can be added to the existing STR classification method to further improve the accuracy of the classification.

[0059] The present application will be further described in detail below through specific embodiments. The following embodiments are only for further illustration of the present application and should not be construed as limiting the present application.

[0060] Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art.

[0061] Example

[0062] This example uses a specific third-generation sequencing dataset for short tandem repeat genotyping, specifically including:

[0063] 1) Download the third-generation sequencing data of individual HG002 from the ONT platform (https: / / human-pangenomics.s3.amazonaws.com / index.html?prefix=submissions / 0CB931D5-AE0C-4187-8BD8-B3A9C9BFDADE--UCSC_HG002_R1041_Duplex_Dorado / Dorado_v0.1.1 / sereo_duplex / ) from the website provided by the Human Pangenomics project, and download the STR baseline data of HG002 from GIAB (https: / / ftp-trace.ncbi.nlm.nih.gov / ReferenceSamples / giab / release / AshkenazimTrio / HG002_NA24385_son / TandemRepeats_v1.0 / GRCh38 / ), and obtain the STR locus catalog from UCSC.

[0064] 2) Perform the SNV call step on the downloaded BAM file. The tool used is the paid software Sentieon. An example of how to use it is as follows:

[0065] sentieon-cli dnascope-longread -r GRCh38_no_alt_analysis_set_GCA_000001405.15.fasta -i HG002_ONT_GRCh38.bam -m DNAscopeONT2.1 -g -t 12 --techHiFi --align -d dbSNP151.hg38.vcf HG002_ONT_GRCh38_3.vcf.gz

[0066] 3) Perform STR typing according to the above method. The input files include the third-generation sequencing bam of HG002, the STR directory bed downloaded from UCSC, and the VCF file generated in the second step. The final output STR typing results are shown in Table 2 below.

[0067] Table 2

[0068]

[0069] 4) The detection results were statistically analyzed using the gold standard STR variant VCF file of HG002 downloaded from GIAB.

[0070] To be considered consistent with the STR loci in the gold standard, the detection of STR loci in VCF must meet the following two rules:

[0071] Site matching rule: STR sites (S_det) in VCF are compared with STR sites (S_ref) in the gold standard VCF. If the chromosomes are consistent and the coordinates overlap ≥90% (or within ±2bp, since there may be a 1bp difference in VCF coordinates), they are determined to be the same STR site.

[0072] Genotype matching rules: In the same STR locus, the number of allele repeats in the test results (e.g., 20 / 22) must be completely consistent with the gold standard. For diploid samples, both alleles must match. Regardless of the order, the genotype is considered correct.

[0073] First, calculate the following indicators:

[0074] True positive (TP): The STR site (S_ref) present in the gold standard has a consistent site (S_det) in the test result.

[0075] False positive (FP): The STR locus (S_det) present in the test result is not consistent with the gold standard.

[0076] False negative (FN): The STR locus (S_ref) present in the gold standard is not consistent with the locus in the test result.

[0077] The calculated TP is 88009, FP is 2572, and FN is 4127.

[0078] Next, statistical indicators for measuring the performance of the invention are calculated:

[0079] Accuracy: The proportion of "true positive sites" to "all positive detection sites" in the test results reflects the "accuracy" of the test results. The calculation method is TP / (TP+FP), which yields 0.9716.

[0080] Recall rate: The proportion of "correctly detected sites" in the gold standard to the "total sites in the gold standard", reflecting the "sensitivity" of the detection tool. It is calculated as TP / (TP+FN) and yields 0.9552.

[0081] F1 score: The harmonic mean of precision and recall, balancing "accuracy" and "sensitivity" to avoid bias from a single metric. It is calculated as 2*(precision*recall) / (precision+recall), which gives 0.9633.

[0082] The results show that the above-end tandem repeat sequence typing method is accurate and reliable, and highly consistent with the results of the gold standard. The results are shown in Table 3.

[0083] Table 3

[0084] Accuracy Recall rate F1 score 0.9716 0.9552 0.9633

[0085] The third-generation sequencing bam of HG002 was analyzed using conventional methods. After comparison and statistical analysis with the gold standard, the statistical results are shown in Table 4 below.

[0086] Table 4

[0087] Accuracy Recall rate F1 score 0.7908 0.8792 0.8327

[0088] The results above show that the method for short tandem repeat genotyping based on third-generation sequencing data in this example can obtain accurate genotyping results, and the results are highly consistent with the existing gold standard results.

[0089] The above description, in conjunction with specific embodiments, provides a further detailed explanation of this application and should not be construed as limiting the specific implementation of this application to these descriptions. Those skilled in the art to which this application pertains can make several simple deductions or substitutions without departing from the concept of this application.

Claims

1. A method for genotyping short tandem repeat sequences based on third-generation sequencing data, characterized in that: Includes the following steps, The target read extraction step includes traversing the alignment file according to the short tandem repeat sequence site directory to extract all reads that completely cover the short tandem repeat sequence region in the third-generation sequencing data to be analyzed; the short tandem repeat sequence site directory is a directory composed of existing known short tandem repeat sequence sites or short tandem repeat sequence sites of interest obtained by downloading; the alignment file is the alignment file between the third-generation sequencing data to be analyzed and the reference genome; The read re-alignment and filtering steps include obtaining the average alignment quality of the reads extracted in the target read extraction step; for reads with an average alignment quality lower than a threshold, selecting the short tandem repeat sequence region corresponding to the read in the reference genome and the sequence in its vicinity for re-alignment with the read, and obtaining the average alignment quality of the re-alignment; if the average alignment quality of the re-alignment is still lower than the threshold, the read is filtered out and removed, and the retained reads are used for subsequent analysis. The copy number calculation steps include processing the reads retained in the read re-alignment and filtering steps as follows: 1) Determine whether the read is an inverted complementary read. If so, invert the read based on the base pairing principle. The inverted complementary sequence of the read is used for subsequent analysis; 2) At the starting point of the short tandem repeat sequence region where the read matches, move the length of the repeat unit of the short tandem repeat sequence for each alignment and matching until no matching repeat unit can be found; 3) Record the number of times the current read can continuously match repeat units, which is taken as the copy number of the repeat unit of the read. The short tandem repeat sequence genotyping confirmation step includes the following processing of the reads retained in the read re-alignment and filtering steps: 1) obtaining the base information of all mutation sites on the reads; 2) filtering out reads with only one mutation site; 3) for all reads matching the same short tandem repeat sequence site region, taking the intersection of their mutation sites, and filtering out reads with only one mutation site in the intersection; 4) genotyping the short tandem repeat sequence according to the haplotype of the mutation site of the read; The results output step includes the location information, repeat unit, and reference repeat number of the short tandem repeat sequence corresponding to the output read, as well as the copy number result of the repeat unit calculation step and the short tandem repeat sequence typing result of the short tandem repeat sequence typing confirmation step; the reference repeat number refers to the number of times the short tandem repeat sequence is repeated on the reference genome.

2. The method according to claim 1, characterized in that: In the segment re-comparison and filtering steps, the threshold for average comparison quality is 20.

3. The method according to claim 1, characterized in that: In the read segment re-alignment and filtering steps, the sequences near the short tandem repeat sequence region are the sequences 1kb upstream and 1kb downstream of the short tandem repeat sequence region.

4. The method according to any one of claims 1-3, characterized in that: In the step of calculating the copy number of the repeating unit, at most one mismatched base is allowed during each comparison and matching process of moving the length of the repeating unit of the short tandem repeating sequence.

5. A device for genotyping short tandem repeat sequences based on third-generation sequencing data, characterized in that: It includes a target read segment extraction module, a read segment overlap comparison and filtering module, a repeating unit copy number calculation module, a short tandem repeat sequence typing confirmation module, and a result output module; The target read extraction module includes a function to extract all reads that completely cover short tandem repeat regions in the third-generation sequencing data to be analyzed by traversing the alignment file according to the short tandem repeat site directory; the short tandem repeat site directory is a directory composed of existing known short tandem repeat sites or short tandem repeat sites of interest obtained by downloading; the alignment file is an alignment file between the third-generation sequencing data to be analyzed and the reference genome; The read re-alignment and filtering module includes a method for obtaining the average alignment quality of reads extracted by the target read extraction module. For reads with an average alignment quality lower than a threshold, the module selects the short tandem repeat sequence region corresponding to the read in the reference genome and the sequence in its vicinity to re-align with the read and obtain the average alignment quality of the re-alignment. If the average alignment quality of the re-alignment is still lower than the threshold, the read is filtered out and removed, and the retained reads are used for subsequent analysis. The repeat unit copy number calculation module includes the following processing for the read segments retained by the read segment re-alignment and filtering module: 1) Determine whether the read segment is an inverse complementary read segment. If so, reverse the read segment based on the base complementary pairing principle. The inverse complementary sequence of the reversed read segment is used for subsequent analysis; 2) When the read segment matches the starting point of the short tandem repeat sequence region, move the length of the repeat unit of the short tandem repeat sequence for each comparison and matching until no repeat unit can be matched; 3) Record the number of times the current read segment can be continuously matched with repeat units, as the repeat unit copy number of the read segment; The short tandem repeat sequence genotyping confirmation module includes the following processing for the reads retained by the read re-alignment and filtering module: 1) obtaining the base information of all mutation sites on the read; 2) filtering out reads with only one mutation site; 3) for all reads matching the same short tandem repeat sequence site region, taking the intersection of their mutation sites, and filtering out reads with only one mutation site in the intersection; 4) genotyping the short tandem repeat sequence according to the haplotype of the mutation site of the read; The results output module includes the position information, repeat unit, and reference repeat number of the short tandem repeat sequence corresponding to the read segment, as well as the repeat unit copy number result of the repeat unit copy number calculation module and the short tandem repeat sequence typing result of the short tandem repeat sequence typing confirmation module; the reference repeat number refers to the number of times the short tandem repeat sequence is repeated on the reference genome.

6. The apparatus according to claim 5, characterized in that: In the segment re-comparison and filtering module, the threshold for average comparison quality is 20.

7. The apparatus according to claim 5, characterized in that: In the read segment re-alignment and filtering module, the sequence near the short tandem repeat sequence region is the sequence 1kb upstream and 1kb downstream of the short tandem repeat sequence region.

8. The apparatus according to any one of claims 5-7, characterized in that: In the repeating unit copy number calculation module, during each comparison and matching process of moving the length of the repeating unit of the short tandem repeating sequence, a maximum of one mismatched base is allowed.

9. A device for genotyping short tandem repeat sequences based on third-generation sequencing data, characterized in that: The device includes a memory and a processor; The memory is used to store programs; The processor is configured to implement the method for short tandem repeat genotyping based on third-generation sequencing data as described in any one of claims 1-4 by executing a program stored in the memory.

10. A computer-readable storage medium, characterized in that: The method includes a program that can be executed by a processor to implement the method for short tandem repeat genotyping based on third-generation sequencing data as described in any one of claims 1-4.