A Genotyping Method for Short Tandem Repeats Based on Next-Generation Sequencing
By designing capture probes and likelihood estimation models, the accuracy and throughput issues of next-generation sequencing in STR genotyping were solved, achieving rapid and accurate STR genotyping, which is suitable for forensic identification and disease screening, thus enhancing the application value of next-generation sequencing in STR testing.
Patent Information
- Application Number
- CN202311055099.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-22
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2043-08-22
AI Technical Summary
In existing technologies, second-generation sequencing suffers from problems such as inaccurate genotyping, limited throughput, and sequencing errors affecting alignment accuracy during the genotyping of short tandem repeat sequences, especially in STR detection where reliable and optimized algorithms are lacking.
By designing capture probes, establishing a reference genome library, performing sequence quality filtering and alignment, classifying read types, combining likelihood estimation models, calculating the genotyping results that best match the observations, considering the impact of sequencing errors and mutations, performing weighted scoring and normal distribution analysis, and integrating read information to improve genotyping accuracy.
It has improved the accuracy and practicality of second-generation sequencing data in STR typing, enabling rapid and accurate STR region typing, suitable for forensic identification and disease screening, and improving detection throughput and adaptability, especially the ability to detect degraded DNA samples.
Smart Images

Figure CN117037906B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics analysis, specifically relating to a genotyping method for short tandem repeat sequences based on next-generation sequencing. Background Technology
[0002] Short tandem repeats (STRs), also known as microsatellite DNA, are simple repeat sequences due to their relatively simple repeat units. They exhibit significant diversity across the biological world and even among different populations, the underlying material basis being the polymorphism of genomic DNA. Human genome STRs are DNA sequences formed by tandem repeats of a few base pairs as core units. Due to differences in repeat units and the number of repeats, their distribution varies greatly among different races and populations, constituting the genetic diversity of STRs. On average, there is one STR site every 15-20 kb in the genome, accounting for 10% of the genome. They are mostly found in non-coding regions and introns, with repeat units of 2-6 bp, repeat numbers of 10-60, and fragment sizes ranging from 100-500 bp. They exhibit co-dominant inheritance according to Mendelian laws.
[0003] Due to variations in repeat units and repeat numbers, STRs exhibit significant differences in distribution across different ethnicities and populations, constituting STR genetic polymorphism. The number of repeats at a homologous STR locus varies among individuals. By identifying specific sequence repeats at specific loci in the genome, it is possible to create an individual's genetic profile. Currently, tens of thousands of STR loci have been publicly disclosed. STR analysis methods have become crucial for individual identification and kinship testing in the field of forensic medicine, and can be used for genetic fingerprinting.
[0004] Genotyping of short tandem repeats (STRs) is typically performed using capillary electrophoresis combined with fluorescent labeling. Specific primers are designed for STR sites to amplify amplification products of varying lengths and fluorescent labels, which are then distinguished by capillary electrophoresis. The number of fluorescent channels used is generally limited to no more than six colors; otherwise, leakage at different wavelengths can occur, limiting the throughput of this method. However, by adding different sequence tags to different samples, next-generation sequencing (NGS) can simultaneously detect hundreds or even thousands of samples, significantly increasing throughput. Capillary electrophoresis-based STR detection only provides the number of tandem repeats, not the DNA sequence information of the repeat regions. NGS, in addition to providing information on the length of the repeat regions, also provides the DNA sequence information of the repeat regions, enhancing the individual STR identification capability. When using capillary electrophoresis, base insertions or deletions in the repeat regions of a sample's STR may result in the sample being grouped into the same bin, leading to STR genotyping errors. Capillary electrophoresis requires longer DNA fragments and cannot detect some degraded DNA. NGS requires smaller amplicon fragments than first-generation sequencing, making it more suitable for STR detection of degraded DNA. Current methods for STR detection using next-generation sequencing (NGS) data still have some shortcomings. For example, NGS reads are generally short, leading to various alignment issues with long repeat regions and resulting in inaccurate STR genotyping. The sample's snv (snives) and other inherent sequencing errors during NGS can also affect alignment accuracy. Currently, reliable and optimized algorithms for STR genotyping using NGS data are still lacking. Summary of the Invention
[0005] To address the problems existing in the prior art, this invention proposes a genotyping method for short tandem repeat sequences based on next-generation sequencing.
[0006] The technical solution adopted by the present invention to achieve the above objectives is as follows:
[0007] This invention provides a genotyping method for short tandem repeat sequences based on next-generation sequencing, comprising the following steps:
[0008] S01) Organize common DNA tandem repeat regions in the genome, design and synthesize capture probes based on sequence coordinates, establish a reference genome library, and create an index;
[0009] S02) Randomly select samples, use the same experimental protocol to obtain probe-captured second-generation sequencing raw FASTQ data, filter the sequence quality of the obtained raw data, and then align it to the reference genome library established in step S01). Sort the alignment results and remove redundant sequences in the samples.
[0010] S03) Classify the types of reads based on the data comparison results;
[0011] S04) Estimate the STR classification length for each type of reads;
[0012] S05) Integrate read information to perform likelihood estimation and calculate the classification result that best matches the observations.
[0013] Furthermore, step S03) specifically involves:
[0014] S31), Set the left boundary coordinates of the reference gene sequence STR region as... The coordinates of the right boundary are Store the coordinate results as a tuple A pair of reads aligned to the reference gene sequence STR region includes the left read1 and the right read2. The left and right coordinates of read1 are stored as a tuple. and The left and right coordinates of read2 on the right are stored as tuples. ;
[0015] S32), Set the left boundary coordinates of the reference gene sequence STR region as... The coordinates of the right boundary are Store the coordinate results as a tuple A pair of reads aligned to the reference gene sequence STR region includes the left read1 and the right read2. The left and right coordinates of read1 are stored as a tuple. and The left and right coordinates of read2 on the right are stored as tuples. ;
[0016] S33), if or or , but In this context, class(F) represents a region in a pair of reads where only one read falls within the reference gene sequence STR region.
[0017] S34), if or , but In this context, class(S) represents a pair of reads in which one read completely crosses the STR region of the reference gene sequence.
[0018] S35) Determine the type of reads based on the coordinates of the read and the coordinates of the STR region of the reference gene sequence. If or ,but That is, in a pair of reads, at least one read is completely contained within the STR region of the reference gene sequence.
[0019] Furthermore, considering the impact of sequencing errors or mutations, the alignment results of reads are weighted and scored, with +1 point awarded for each matching base. This is based on the Phred quality score.
[0020]
[0021] It can be used to assess sequencing quality. Typically, it will... The bases are considered high-quality bases. Bases with a low quality mismatch are considered low-quality bases. Based on this standard, this method scores each low-quality mismatch as follows: Each high-quality mismatch is scored as -1 point. The total score for the length is m, and the weighted score after normalization is n, where n ranges from m to n. ,if Determine the STR repetition interval of the sequence, and then proceed to step S04.
[0022] Furthermore, when reads belong to class(I), step S04) is:
[0023] a) If the STR repeat length is longer than the read length, and the paired sequence is not in the repeat region, the repeat number is estimated based on the length completely contained within the repeat region. That is, the repeat number can be estimated by calculating the length difference of the read within the repeat region.
[0024] b) If the paired sequence is also anchored to the repeating region, then the number of repeats within the repeating region can be estimated based on the coordinates of the two reads. That is, assuming the left and right coordinates of the left read1 of the paired sequence are... and The left and right coordinates of read2 on the right are and The number of repetitions can be estimated by calculating the distance difference between two reads.
[0025] c) If the paired sequence portion is in a repeating region, the repeat count can be estimated based on the repeat count of the mate sequence, the spacing, and the repeat count of the sequence itself. Specifically, it is necessary to estimate the repeat count of the mate sequence within the repeating region, the spacing between the mate sequence and the current read, and the repeat count of the current read itself, and then estimate the repeat count within the repeating region of the STR based on these three quantities.
[0026] Furthermore, when reads belong to class(F), step S04) is: Let the number of STR repetitions observed for each pair of reads be... All observations follow a normal distribution, and the number of observed repetitions is... It equals the ratio of the length of the read repetitions to the number of repetition units; the left and right boundaries of STR are... The position of each pair of reads compared to the repeated region is: ,if , Store in list L1, if , Store the data in list L2; the start and end positions in L1 also conform to a normal distribution, thus determining the left boundary of the STR maximum likelihood estimate. Similarly, the start and end positions in L2 also conform to a normal distribution, thus determining the right boundary of the STR maximum likelihood estimate. Specifically, let the number of STR repetitions be... The length of the read STR is The length of the repeating unit is ,but The length of the repeating unit here. This can be obtained from a reference sequence; our task is to estimate the number of STR repetitions. According to the law of large numbers, the number of repetitions... It approximately follows a normal distribution There is a set of observations here. The corresponding likelihood function is:
[0027]
[0028] in It is a repeating number variance Let be the number of repetitions obtained for each observation. Due to the additivity of the normal distribution, the start and end positions of the list also follow a normal distribution, thus allowing us to calculate the left and right boundaries of the maximum likelihood estimate of STR.
[0029] Furthermore, when reads belong to class(S), step S04) is: evaluate all data aligned to the STR reference sequence range, determine whether there are repeating units in these sequence regions, and align the flanking sequences of reads that cross the STR with the flanking sequences of the STR reference sequence. If the alignment weight score of reads in the STR region and its flanking sequences is greater than 0.9, then determine the number of repeating units based on the alignment region.
[0030] Further, step S05 is as follows: the observed values of the sequence alignment results and a set of reference short tandem repeat values are used as inputs and the estimated diploid repeat length is output to construct a likelihood estimation model. The information from the alignment of second-generation sequencing paired-end reads is integrated into this model. The model searches for all possible alleles and returns the maximum likelihood diploid genotype. Any STR repeat number a that supports one or more read classes is added to the STR genotyping list. The STR repeat number a is used to perform one-dimensional optimization of the likelihood function to find allele b, and the genotyping result (a, b) that maximizes the likelihood is obtained. Next, multiple rounds of loops are executed to find the genotyping result (S1, S2) with the maximum global likelihood.
[0031] Furthermore, the probe-captured next-generation sequencing data is a hybridization capture probe reagent, which needs to cover the designed STR region.
[0032] Furthermore, quality control is performed on the sequencing data, samples that fail quality control are removed, and the average sequencing depth of the STR region of the tested samples is required to be greater than 60X.
[0033] Furthermore, the samples include samples within the normal range and samples with extended abnormal results.
[0034] The beneficial effects of this invention are as follows:
[0035] (1) The tandem repeat number typing results obtained by the present invention through second-generation sequencing data analysis are consistent with the capillary electrophoresis results, which can quickly and accurately perform correct typing of the STR region of ATXN3.
[0036] (2) The algorithm provided by this invention proves the accuracy and practicality of second-generation sequencing data in genotyping of short tandem repeat sequences in the genome. It can be used for STR detection and applied to forensic identification, disease screening and other fields. Attached Figure Description
[0037] Figure 1 This is a flowchart illustrating the method.
[0038] Figure 2 This is a schematic diagram illustrating the classification of reads types in Example 1;
[0039] Figure 3 This is a schematic diagram of the reference gene sequence in Example 1;
[0040] Figure 4 This is a schematic diagram of the comparison data results in Example 1;
[0041] Figure 5 This is a schematic diagram showing the results of capillary electrophoresis detection of the test samples in the example. Detailed Implementation
[0042] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0043] Example 1
[0044] This example uses the CAG repeat sequence in the ATXN3 gene as an example.
[0045] CAG polymorphic repeats in the ATXN3 gene are unstable and amplify to an abnormal range in all individuals with SCA3. Normal-length alleles have 12 to 44 CAG repeat sequences. Intermediate-length alleles have CAG repeat lengths between clearly normal and fully penetrating alleles, some of which are not associated with the classic clinical features of SCA3. Fully penetrating alleles represent affected individuals with the classic SCA3 phenotype, with allele lengths ranging from 56 to 87. Second-generation sequencing reads are generally no longer than 150 bp, and the aligned reads may exhibit various positional relationships with the STR region. Classification processing can more accurately estimate STR genotyping results.
[0046] like Figure 1 As shown, the genotyping method for CAG repeat sequences in the ATXN3 gene in this embodiment includes the following steps:
[0047] S01) Organize the STR repeat region sequence of the ATXN3 gene on chromosome 14, extend it 500bp upstream and downstream as reference sequences, and create an index, such as... Figure 3 As shown.
[0048] The reference sequence is as follows:
[0049] .
[0050] S02) Take 26 samples with known amplification abnormalities and 30 samples with normal ranges together to perform hybridization capture experiments and next-generation sequencing. The designed capture probe covers the ATXN3 STR region.
[0051] (S03) The raw BCL format files output by the next-generation sequencer are stored in the directory path. Use bcl2fastq to split the BCL files into FastQ format files and store them in the specified path. Use the raw data quality control software FastP to filter the raw data and perform basic quality control such as Q30.
[0052] S04) Use BWA to align the fastq file and the reference sequence to obtain the bam file, and use sambamba to sort and mark redundancy in the alignment file, and remove redundant sequences.
[0053] S05) Perform quality control on the sequencing data. If the average sequencing depth of the STR region is greater than 60X, subsequent data analysis can proceed; otherwise, additional sequencing or re-sequencing is required until the data meets the requirements. Filter reads aligned to the STR region, retaining only reads with MAPQ > 20. For each pair of reads meeting the criteria, obtain their specific alignment information based on the cigar value and other parameters in the BAM file. The alignment data results are as follows: Figure 4 As shown.
[0054] S06), Set the left boundary coordinates of the reference gene sequence STR region as... The coordinates of the right boundary are Store the coordinate results as a tuple A pair of reads aligned to the reference gene sequence STR region includes the left read1 and the right read2. The left and right coordinates of read1 are stored as a tuple. and The left and right coordinates of read2 on the right are stored as tuples. .
[0055] Define three types of double-ended read lengths:
[0056] class(I) represents a pair of reads in which at least one read is completely contained within the repeated region. If or ,but .
[0057] class(F) represents a pair of reads in which only one read's portion falls into the duplicate region. If
[0058] or or , but .
[0059] class(S) represents a pair of reads in which one read completely crosses the concatenated repeating region. If or , but ,
[0060] S07) Estimate the STR typing length for each type of read.
[0061] Consider comparing reads with fully repeat sequences in cases of shift or reverse complementation (for SCA3 CAG mode, forward alignment yields fully repeat sequences of CAG, AGC, and GCA, or reverse fully repeat sequences of CTG, TGC, and GCT). Taking into account the impact of sequencing errors or mutations, the alignment results of reads are weighted and scored. The scoring criteria are +1 point for each matching base, -0.5 points for each low-quality mismatch, and -1 point for each high-quality mismatch. The total length score is calculated as m, and the normalized weighted score is n, where n ranges from (-1, 1). If n is greater than 0.9, the STR repeat interval of the sequence is determined.
[0062] When reads belong to class(I), the STR repeat length is longer than the read length. If the paired sequence is not in the repeat region, the repeat count is estimated based on the length completely contained within the repeat region. If the paired sequence is also anchored to the repeat region, the repeat count within the repeat region is estimated based on the coordinates of the two reads. If the paired sequence is partially in the repeat region, the repeat count is estimated based on the repeat count of the mate sequence, the spacing, and the repeat count of the sequence itself.
[0063] When reads belong to class(F), assume that the number of STR repetitions observed for each pair of reads is... All observations follow a normal distribution. The number of observed repetitions. This is equal to the ratio of the length of the read repetitions to the number of repetition units. The left and right boundaries of STR are... The position of each pair of reads compared to the repeated region is: ,if Store in list L1, if Store in list L2. The start and end positions in L1 also conform to a normal distribution, thus determining the left boundary of the STR maximum likelihood estimate. The start and end positions in L2 also conform to a normal distribution, thus determining the right boundary of the STR maximum likelihood estimate.
[0064] Here, variance is introduced to assess the number of observed repetitions. The difference between the actual number of repetitions s and the actual number of repetitions s
[0065]
[0066] A smaller variance indicates a more concentrated distribution of observations, suggesting a more accurate estimated number of replicates. Conversely, a larger variance indicates a more dispersed distribution of observations, making the estimated number of replicates unreliable.
[0067] When reads belong to class(S), they are evaluated across all data aligned to the SCA3 STR reference sequence. The presence of repeating units in these sequence regions is determined, and the flanking sequences of reads spanning the STR are compared with the flanking sequences of the STR reference sequence. If the alignment weight score of the reads in the STR region and its flanking sequences is greater than 0.9, the number of repeats is determined based on the alignment region.
[0068] S08) Integrate read information, assuming the number of repetitions here. Obtain the parameter as For a normal distribution, likelihood estimation can be performed using the following function:
[0069]
[0070]
[0071] in Let be the joint probability density function of the normal distribution. and is the parameter of the normal distribution. Let be the log-likelihood function. The total number of repetitions
[0072] Calculate the classification result that best matches the observed values.
[0073] This step takes the sequence alignment observations and a set of reference short tandem repeat values as input and outputs the estimated diploid repeat length. A likelihood estimation model is constructed, integrating various information from second-generation sequencing paired-end read alignments. This model is applied to the detection of SCA3 tandem repeat regions. By default, the algorithm searches for all possible alleles and returns the maximum likelihood diploid genotype. Any STR repeat number 'a' supporting one or more read classes is added to the STR genotyping list. STR repeat number 'a' is used to perform one-dimensional optimization of the likelihood function to find allele 'b', resulting in the genotyping result (a, b) that maximizes the current likelihood. Next, multiple iterations are performed to find the genotyping result (S1, S2) with the highest global likelihood. In the implementation, the search process is optimized using a non-linear search, which allows for rapid estimation of the number of all possible allele repeats.
[0074] The genotyping results for all test samples were calculated using this method, and capillary electrophoresis was performed on the test samples. The results are as follows: Figure 5 As shown.
[0075] The GATK haplotypecaller algorithm, used for variant detection in second-generation sequencing data, can only report one repeat number and cannot perform genotyping. The results are shown in Table (A) below. The genotyping results using the method of this invention (B) and the genotyping results using capillary electrophoresis (C) can be compared and judged. See Table 1 for details.
[0076] Table 1
[0077]
[0078] As the above comparison shows, using the technical solution of this invention, the tandem repeat number genotyping results obtained through next-generation sequencing data analysis are consistent with the capillary electrophoresis results, enabling rapid and accurate genotyping of the ATXN3 STR region. However, traditional capillary electrophoresis-based genotyping methods perform poorly when alleles have potentially different sequences, and require significant time and manpower investment in development and genotyping, while also exhibiting low levels of automation and data standardization. Compared to capillary electrophoresis, next-generation sequencing offers higher throughput, capable of processing hundreds or even thousands of samples simultaneously; in addition to obtaining information on the length of short tandem repeat regions, it can also acquire DNA sequence information from the repeat regions, enhancing the ability to identify individual STRs; it is suitable for degraded DNA samples because the required amplicon fragments are shorter, making it more suitable for STR detection of degraded DNA than capillary electrophoresis.
[0079] Current second-generation sequencing datasets typically detect SNVs, small indels, and SVs for individual genome analysis. However, STRs are rarely analyzed due to a lack of reliable and optimized algorithms. This algorithm demonstrates the accuracy and practicality of second-generation sequencing data in genotyping of short tandem repeat sequences in the genome and can be used for STR detection, with applications in forensic identification, disease screening, and other fields.
[0080] The above description is merely the basic principle and preferred embodiment of the present invention. Improvements and substitutions made by those skilled in the art based on the present invention are within the scope of protection of the present invention.
Claims
1. A method for typing short tandem repeat sequences based on second-generation sequencing, characterized by: The method comprises the following steps: S01) arranging the sequence of the common DNA tandem repeat region in the genome, designing and synthesizing the capture probe according to the sequence coordinates, establishing the reference genome library, and establishing the index; S02) randomly selecting samples, using the same experimental scheme to obtain the probe capture next-generation sequencing raw fastq data, filtering the measured raw data according to the sequence quality, then aligning to the reference genome library established in step S01), sorting the alignment results, and removing the redundant sequences in the sample; S03) classifying reads according to the alignment results of the data; S04) estimating the STR typing length of each type of reads; S05) integrating the read information to estimate the likelihood, and calculating the typing result that best fits the observed value; In step S03), the specific process is: S31) Set the left boundary coordinate of the reference gene sequence STR region as , the right boundary coordinate as , and store the coordinate result as a tuple ; a pair of reads aligned to the reference gene sequence STR region includes a left read 1 and a right read 2, the left and right coordinates of the read 1 are stored as a tuple and , and the left and right coordinates of the right read 2 are stored as a tuple ; S32) Set the left boundary coordinate of the reference gene sequence STR region as , the right boundary coordinate as , and store the coordinate result as a tuple ; a pair of reads aligned to the reference gene sequence STR region includes a left read 1 and a right read 2, the left and right coordinates of the read 1 are stored as a tuple and , and the left and right coordinates of the right read 2 are stored as a tuple ; S33) if or or , then , wherein class(F) represents that only a partial region of one read in the pair falls within the STR region of the reference sequence. S34) if or , then , wherein class(S) represents that one of the paired reads completely crosses the STR region of the reference genomic sequence STR. S35) judging the type of reads according to the coordinates of reads and the coordinates of STR region of reference gene sequence; if or then that is, at least one read in the pair of reads is completely contained in the STR region of the reference gene sequence; In step S04), the specific process is: a) if the STR repeat length is longer than the read length, and if the paired sequence is not in the repeat region, the repeat number is estimated according to the length completely contained in the repeat region, that is, the repeat number can be calculated by the length difference of the read in the repeat region; b) If the paired sequence is also anchored to the repeat region, estimate the number of repeats within the repeat region according to the coordinates of the two reads: i.e. when the left read 1 of the paired sequence has left and right coordinates and , and the right read 2 has left and right coordinates and , the estimate of the number of repeats can be made by calculating the difference in distance between the two reads; c) if the paired sequence is partially in the repeat region, the repeat number of the mate sequence in the repeat region, the interval between the mate sequence and the current read, and the repeat number of the current read itself need to be estimated, and then the repeat number in the STR repeat region is estimated according to the three quantities; In step S05), the specific process is: The observed value of the sequence alignment result and a set of reference short tandem repeat values are taken as inputs and the estimated diploid repeat length is taken as output, a likelihood estimation model is constructed, the information from the paired-end read alignment of the next-generation sequencing is integrated into the model, the model searches all possible alleles, and returns the maximum likelihood diploid genotype, any STR repeat number a that supports one or more read classes is added to the STR typing list, the STR repeat number a is used to perform one-dimensional optimization of the likelihood function to find the allele b, and the current typing result (a, b) that maximizes the likelihood rate is obtained, and then a plurality of cycles are performed to find the typing result (S1, S2) that maximizes the global likelihood rate.
2. The method for typing short tandem repeat sequences based on second-generation sequencing according to claim 1, characterized in that: In step S02), the probe capture next-generation sequencing raw fastq data is a hybrid capture probe reagent, and the average sequencing depth of the STR region of the sample to be detected is required to be greater than 60X.
3. The method of claim 1, wherein the method is a next generation sequencing based short tandem repeat typing method. In S03), the alignment results of the reads are weighted and scored, and the scoring standard is +1 point for each matching base, and the sequencing quality is evaluated according to the Phred quality score: Will The bases are considered high-quality bases. Bases that are considered low-quality bases are then scored as follows: Each high-quality mismatch is scored out of 100. The total length score is calculated as m, and the weighted score after normalization is n, where n ranges from 1 to 1. ,if Determine the STR repetition interval of the sequence, and then proceed to step S04.
4. The method for typing short tandem repeat sequences based on second-generation sequencing according to claim 1 or 3, characterized in that: When the reads belong to class (I), step S04) is: the STR repeat length is longer than the read length, and if the paired sequence is not in the repeat region, the repeat number is estimated according to the length completely contained in the repeat region; if the paired sequence is also anchored to the repeat region, the repeat number in the repeat region is estimated according to the coordinates of the two reads; if the paired sequence is partially in the repeat region, the repeat number is estimated according to the repeat number of the mate sequence, the interval and the repeat number of the sequence itself.
5. The method for short tandem repeat sequence typing based on second-generation sequencing according to claim 4, characterized in that: When , step S04) is: set the STR repeat number observed by each pair of reads as , all observations follow a normal distribution, the observed repeat number is equal to the ratio of the length of the read repeat to the repeat unit; the left and right boundaries of the STR are , the position of each pair of reads aligned to the repeat region is , if , , store in list L1, if , , store in list L2; the start and end positions in L1 also conform to a normal distribution, determine the left boundary of the maximum likelihood estimate of the STR, and the start and end positions in L2 also conform to a normal distribution, determine the right boundary of the maximum likelihood estimate of the STR: set the STR repeat number as , the length of the read STR is , and the length of the repeat unit is , then , the r is the length of the repeat unit , the repeat number approximately obeys a normal distribution . A set of observations , and the corresponding likelihood function is: wherein is the variance of the number of repetitions is the number of repetitions obtained for each observation. 6. The method for short tandem repeat sequence typing based on second-generation sequencing according to claim 5, characterized in that: When the reads belong to class (S), step S04) is: among all the data aligned to the STR reference sequence range, it is judged whether there is a repeat unit in these sequence regions, and the flanking sequence of the reads crossing the STR is aligned with the flanking sequence of the STR reference sequence, if the alignment weight score of the reads in the STR region and its flanking sequence is greater than 0.9, then the number of repeats is judged according to the alignment region.
Citation Information
Patent Citations
Whole-exome sequencing data analysis system
CN106021984A
Method for detecting and typing repeat number of short tandem repeat sequence
CN113362892A