Method, device, equipment and storage medium for analyzing short tandem repeats
Patent Information
- Application Number
- CN202311397900.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-26
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-10-26
AI Technical Summary
[0003]但是,由于测序技术不同、测序偏好等因素,使得不同STR分析预测结果存在偏差和不稳定性
[0049]第四方面,本申请实施例提供一种计算机可读存储介质,所述计算机可读存储介质上存储有计算机可执行指令,所述计算机可执行指令被处理器运行时执行上述的短串联重复的分析方法。
Smart Images

Figure CN117373531B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of gene sequencing technology, and in particular to an analytical method, apparatus, device, and storage medium for short tandem repeats. Background Technology
[0002] Short tandem repeats (STRs) are sequence structures consisting of tandemly linked nucleic acid sequence modules ranging in length from 1 to 6 bp. Due to the high heterogeneity and abundance of the number of core repeat units among individuals, STR loci exhibit genetic polymorphism. Currently, STRs are frequently used in research fields such as genetic mapping, forensic identification, kinship analysis, disease gene localization, and species polymorphism. In addition, STR repeat amplification is associated with severe neurological or neuromuscular diseases such as Huntington's disease, various ataxias, amyotrophic lateral sclerosis (ALS), frontotemporal dementia, and Fragile X syndrome. Furthermore, the number of STR repeats is closely related to disease severity, age of onset, and clinical symptoms; therefore, accurate detection of STR repeat counts can provide important information for disease diagnosis and management.
[0003] However, due to differences in sequencing technologies and sequencing preferences, the prediction results of different STR analyses exhibit biases and instabilities. Summary of the Invention
[0004] In view of the above, embodiments of this application provide a method, apparatus, device, and computer-readable storage medium for analyzing short serial repetitions in order to solve at least one problem existing in the prior art.
[0005] In a first aspect, embodiments of this application provide a method for analyzing short tandem repeats, comprising:
[0006] Based on the analysis of the exome sequencing data of the test sample, the short tandem repeat region of the test sample is obtained, the confidence level of the repeat count data of the preset related genes in the short tandem repeat region is determined, and a first preset score corresponding to the confidence level is obtained.
[0007] Based on the repeat count data of preset related genes within the short tandem repeat region, determine the range of abnormal amplification counts related to preset diseases in the preset related genes and obtain a second preset score corresponding to the range of abnormal amplification counts;
[0008] A third preset score corresponding to the coverage degree of the two alleles in the short tandem repeat region is obtained.
[0009] Based on the enrichment degree of the preset related genes related to the preset disease obtained from the transcriptome sequencing data analysis of the sample to be tested, a fourth preset score corresponding to the enrichment degree is determined;
[0010] The type of short tandem repeat is determined based on the first preset score, the second preset score, the third preset score, and the fourth preset score; the type includes polymorphic short tandem repeat, pathogenic short tandem repeat, and undetermined short tandem repeat.
[0011] In conjunction with the first aspect, in an optional implementation, the first preset score includes a first score or a first second score;
[0012] The step of determining the confidence level of the repeat count data of preset related genes within the short tandem repeat region of the test sample obtained from the analysis of exome sequencing data of the test sample, and obtaining a first preset score corresponding to the confidence level, includes:
[0013] Based on the analysis of the exome sequencing data of the sample to be tested, the short tandem repeat regions of the sample are determined to determine whether the number of deoxyribonucleic acid fragments in the short tandem repeat regions meets the preset number condition; the preset number condition is determined based on the preset number of deoxyribonucleic acid fragments.
[0014] If the preset quantity condition is met, the confidence level of the repeat count data of the preset related genes in the short tandem repeat region is determined to be high confidence level and the first score corresponding to the high confidence level is obtained;
[0015] If the preset quantity condition is not met, the confidence level of the repeat count data of the preset related genes in the short tandem repeat region is determined to be low confidence level, and the first and second scores corresponding to the low confidence level are obtained.
[0016] In conjunction with the first aspect, in an optional implementation, the second preset score includes a second first score, a second second score, or a second third score;
[0017] The step of determining the range of abnormal amplification counts related to a preset disease within the preset related genes based on the repeat count data of preset related genes in the short tandem repeat region and obtaining a second preset score corresponding to the range of abnormal amplification counts includes:
[0018] Based on the repeat count data of preset related genes in the short tandem repeat region, determine whether the maximum value of abnormal amplification count related to preset disease in the preset related genes is less than or equal to the reference maximum value of normal amplification count;
[0019] If the score is less than or equal to the maximum reference value for the number of normal amplifications, a second score corresponding to the maximum value is obtained.
[0020] If the number of amplifications is greater than the maximum reference value for normal amplifications, determine whether the maximum value is greater than or equal to the minimum reference value for abnormal amplifications.
[0021] If the number of abnormal amplifications is greater than or equal to the minimum reference value, a second score corresponding to the maximum value is obtained;
[0022] If the score is less than the minimum reference value for the number of abnormal amplifications, the second or third score corresponding to the maximum value is obtained.
[0023] In conjunction with the first aspect, in an optional implementation, the third preset score includes a third first score or a third second score;
[0024] The step of obtaining a third preset score corresponding to the coverage degree of the two alleles in the short tandem repeat region includes:
[0025] If the coverage of the two alleles in the short tandem repeat region meets the first condition, a third score corresponding to the coverage is obtained.
[0026] If the coverage of the two alleles in the short tandem repeat region meets the second condition, a third score corresponding to the coverage is obtained.
[0027] The distinction between the first and second types of cases is determined based on the number of consecutive coverages of the short tandem repeats occurring in a single exon region or multiple exon regions.
[0028] In conjunction with the first aspect, in an optional implementation, the fourth preset score includes a fourth first score or a fourth second score;
[0029] The step of determining a fourth preset score corresponding to the enrichment degree of preset related genes related to preset diseases based on the transcriptome sequencing data analysis of the test sample includes:
[0030] Based on the average enrichment level of the preset related genes related to the preset disease obtained from the transcriptome sequencing data analysis of the sample to be tested, determine whether the absolute value of the difference between the average value and the enrichment reference value is greater than or equal to a preset threshold.
[0031] If it is greater than or equal to a preset threshold, obtain the fourth score corresponding to the average value;
[0032] If it is less than the preset threshold, the fourth score corresponding to the average value is obtained.
[0033] In conjunction with the first aspect, in an optional implementation, determining the type of the short serial repetition based on the first preset score, the second preset score, the third preset score, and the fourth preset score includes:
[0034] The type of short tandem repeat is determined based on the total score determined by the sum of the first preset score, the second preset score, the third preset score, and the fourth preset score, and the repeat count data of preset related genes in the short tandem repeat region.
[0035] In conjunction with the first aspect, in an optional implementation, determining the type of the short tandem repeat based on the total score determined by the sum of the first preset score, the second preset score, the third preset score, and the fourth preset score, and the repeat count data of preset related genes within the short tandem repeat region, includes:
[0036] If the maximum value of the number of abnormal amplifications related to the preset disease is less than the maximum value of the number of normal amplifications, and the total score is greater than the first score threshold, the short tandem repeat is determined to be a short tandem repeat to be undetermined.
[0037] If the maximum value of the number of abnormal amplifications related to the preset disease is less than the maximum value of the number of normal amplifications, and the total score is less than or equal to the first score threshold, the short tandem repeat is determined to be a polymorphic short tandem repeat.
[0038] If the maximum value of the number of abnormal amplifications related to the preset disease is greater than or equal to the maximum value of the number of normal amplifications and less than or equal to the minimum value of the number of abnormal amplifications, and the total score is greater than the first score threshold, the short tandem repeat is determined to be a short tandem repeat to be undetermined.
[0039] If the maximum value of the number of abnormal amplifications related to the preset disease is greater than or equal to the maximum value of the number of normal amplifications and less than or equal to the minimum value of the number of abnormal amplifications, and the total score is less than or equal to the first score threshold, the short tandem repeat is determined to be a polymorphic short tandem repeat.
[0040] If the maximum value of the number of abnormal amplifications related to the preset disease is greater than the minimum reference value of the number of abnormal amplifications, and the total score is less than or equal to the second score threshold, the short tandem repeat is determined to be a short tandem repeat to be undetermined.
[0041] If the maximum value of the number of abnormal amplifications related to a preset disease is greater than the minimum reference value for the number of abnormal amplifications, and the total score is greater than the second score threshold, the short tandem repeat is determined to be a pathogenic short tandem repeat.
[0042] Secondly, embodiments of this application provide an analysis apparatus for short tandem repetitions, comprising:
[0043] The first determining unit is used to determine the confidence level of the repeat count data of a preset related gene in the short tandem repeat region of the test sample obtained by analyzing the exome sequencing data of the test sample and to obtain a first preset score corresponding to the confidence level.
[0044] The second determining unit is used to determine the range of abnormal amplification times related to a preset disease in the preset related genes based on the repeat count data of preset related genes in the short tandem repeat region, and to obtain a second preset score corresponding to the range of abnormal amplification times.
[0045] The third determining unit is used to obtain a third preset score corresponding to the coverage degree of the two alleles in the short tandem repeat region.
[0046] The fourth determining unit is used to determine a fourth preset score corresponding to the enrichment degree of the preset related genes related to the preset disease obtained from the transcriptome sequencing data analysis of the sample to be tested.
[0047] The fifth determining unit is used to determine the type of the short tandem repeat based on the first preset score, the second preset score, the third preset score, and the fourth preset score; the type includes polymorphic short tandem repeat, pathogenic short tandem repeat, and undetermined short tandem repeat.
[0048] Thirdly, embodiments of this application provide an analysis device for short serial repetitions, including a processor and a memory, wherein the memory stores computer-executable instructions, and the computer-executable instructions are executed by the processor to perform the above-described analysis method for short serial repetitions.
[0049] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, perform the aforementioned short serial repetition analysis method.
[0050] The beneficial effects of the technical solution provided in this application include: by determining the type of short tandem repeat based on the first preset score, the second preset score, the third preset score, and the fourth preset score, a quality control standard is introduced to control the analysis and ensure the accuracy of STR analysis. Furthermore, analysis combined with transcriptome data improves the accuracy of STR analysis. In addition, multiple short tandem repeat types, including polymorphic short tandem repeats, pathogenic short tandem repeats, and undetermined short tandem repeats, are identified, providing analysis functions for non-pathogenic STR sites and providing a basis for other analyses such as kinship and loss of heterozygosity.
[0051] Additional aspects and advantages of the embodiments of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the embodiments of this application. Attached Figure Description
[0052] The accompanying drawings, incorporated in and forming part of this specification, illustrate embodiments conforming to this application and, together with the description, serve to explain the principles of this application. To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort. These drawings and textual descriptions are not intended to limit the scope of the concept of this application in any way, but rather to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. In the drawings:
[0053] Figure 1 This is a schematic diagram illustrating a specific example of the method for analyzing short serial repetitions in this application.
[0054] Figure 2 This is a schematic diagram illustrating the first type of coverage of two alleles in the embodiments of this application;
[0055] Figure 3 This is a schematic diagram illustrating the second type of coverage of two alleles in the embodiments of this application;
[0056] Figure 4 This is a schematic diagram illustrating the baseline of key genes in disease-related pathways for normal populations in this application embodiment.
[0057] Figure 5 This is a schematic diagram illustrating the expression of gene sets related to pathways in patients with amyotrophic lateral sclerosis (ALS) in an embodiment of this application.
[0058] Figure 6 This is a schematic diagram illustrating a specific example of score allocation in an embodiment of this application.
[0059] Figure 7 This is a schematic block diagram of a specific example of an analysis device for short serial repetition in the embodiments of this application. Detailed Implementation
[0060] To make the technical solutions and beneficial effects of the embodiments of this application more apparent and understandable, a detailed description is provided below by listing specific embodiments. The accompanying drawings are not necessarily drawn to scale, and local features may be enlarged or reduced to more clearly show the details of the local features; unless otherwise defined, the technical and scientific terms used herein have the same meanings as those in the technical field to which the embodiments of this application pertain.
[0061] It should be noted that the terms "first," "second," etc., may be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish one element from another. When "first" is described, it does not imply the existence of a "second"; and when "second" is discussed, it does not imply the existence of a "first." The singular forms "a," "an," and "the" may also be intended to include the plural forms unless the context clearly indicates otherwise. The term "comprising" is used to identify the presence of included features, but does not exclude the presence or addition of one or more other features. The term "and / or" includes any and all combinations of the related listed items.
[0062] STR identification often uses techniques such as polymerase chain reaction (PCR) combined with capillary electrophoresis or DNA blotting. However, most techniques for identifying disease-related STRs target only a single gene. Therefore, before testing, clinicians need to make a preliminary judgment based on the clinical phenotype and select appropriate tests. Because STR-related diseases are characterized by diverse clinical phenotypes, incomplete penetrance, and onset at all ages, and because other mutations such as single nucleotide variants (SNPs) and small insertion / deletion variants (Indels) also account for a large proportion of affected individuals, choosing suitable testing methods is crucial when dealing with diseases exhibiting high genetic phenotypic heterogeneity.
[0063] Next-generation sequencing-based whole-exome sequencing (WES) and whole-genome sequencing (WGS) have made significant contributions to numerous fields, including molecular diagnostics of hereditary diseases, and are highly efficient in identifying SNPs and indels. Various analytical methods are already available for STR detection based on next-generation sequencing data. Currently, both WGS and WES have accumulated substantial datasets and are widely used in human genetic disease research, forensic identification, and molecular diagnostics. Therefore, accurate STR analysis based on high-throughput sequencing data is particularly important.
[0064] Therefore, embodiments of this application provide a method for analyzing short tandem repeats (STRs), applicable to STR analysis based on high-throughput sequencing data, which may include exome sequencing data and transcriptome sequencing data. Figure 1 As shown, the analysis method for short tandem repeats includes:
[0065] Step S10: Based on the analysis of the exome sequencing data of the test sample, the short tandem repeat region of the test sample is determined, the confidence level of the repeat count data of the preset related genes in the short tandem repeat region is determined, and the first preset score corresponding to the confidence level is obtained.
[0066] Step S20: Based on the repeat count data of preset related genes in the short tandem repeat region, determine the range of abnormal amplification counts related to preset diseases in the preset related genes and obtain the second preset score corresponding to the range of abnormal amplification counts.
[0067] Step S30: Based on the coverage of the two alleles in the short tandem repeat region, obtain the third preset score corresponding to the coverage.
[0068] Step S40: Based on the enrichment degree of preset related genes related to preset diseases obtained from the transcriptome sequencing data analysis of the sample to be tested, determine the fourth preset score corresponding to the enrichment degree.
[0069] Step S50: Determine the type of short tandem repeat based on the first preset score, the second preset score, the third preset score, and the fourth preset score; the types include polymorphic short tandem repeat, pathogenic short tandem repeat, and undetermined short tandem repeat.
[0070] In this embodiment, the step of obtaining short tandem repeat regions of the test sample based on exome sequencing data analysis can be achieved by first assessing the repeat size in the target region using spanning reads, flanking reads, and in-repeat reads patterns in a BAM (Bio-Organic Alignment Model) file. Then, a graph-based classification model is used to rearrange the locus structure and calculate the repeat count, yielding the short tandem repeat results. Furthermore, comprehensive annotation and visualization of the short tandem repeat analysis results can further graphically distinguish STR fragment read coverage patterns, allowing for verification and calibration of STR repeat counts.
[0071] Preset related genes can be set according to actual needs, such as the 37 genes in the table below. Reference genomic regions and amplification number reference ranges related to the occurrence of abnormal amplification associated with the preset disease can be obtained from these preset related genes (e.g., they can be obtained from databases such as OMIM and PubMed). The amplification number reference range includes the minimum reference value of normal amplification number norm_low, the maximum reference value of normal amplification number norm_up, the minimum reference value of abnormal amplification number aff_low, and the maximum reference value of abnormal amplification number aff_up.
[0072]
[0073]
[0074]
[0075]
[0076] In this embodiment, the accuracy of short tandem repeat analysis is ensured by introducing quality control standards based on confidence level, number of aberrations, and coverage of two alleles in the analysis of short tandem repeats using exome sequencing data. Furthermore, combining this analysis with transcriptome sequencing data improves the accuracy and stability of short tandem repeat analysis.
[0077] In one optional implementation, the first preset score includes either a first score or a first second score;
[0078] Step S10, which involves analyzing the short tandem repeat regions of the test sample based on the exome sequencing data, determining the confidence level of the repeat count data of preset related genes within the short tandem repeat regions, and obtaining a first preset score corresponding to the confidence level, includes:
[0079] Step S101: Based on the analysis of the exome sequencing data of the sample to be tested, determine whether the number of deoxyribonucleic acid fragments in the short tandem repeat region of the sample meets the preset number condition; the preset number condition is determined according to the preset number of deoxyribonucleic acid fragments.
[0080] Step S102: If the preset quantity condition is met, determine the confidence level of the repeat count data of the preset related genes in the short tandem repeat region as high confidence level and obtain the first score corresponding to the high confidence level.
[0081] Step S103: If the preset quantity condition is not met, determine the confidence level of the repeat count data of the preset related genes in the short tandem repeat region as low confidence and obtain the first and second scores corresponding to the low confidence level.
[0082] In this embodiment, the preset quantity condition can be set according to actual needs. For example, it can be that the depth within the STR region is greater than or equal to 10X (i.e., the number of reads covered per base in the target region is less than 10) or the number of reads crossing single-end / double-end breakpoints is greater than or equal to 5. The number of reads refers to the number of deoxyribonucleic acid (DNA) fragments read by the sequencing instrument during gene sequencing. That is, if the depth within the STR region is less than 10X or the number of reads crossing single-end / double-end breakpoints is less than 5, then the repeat count data of the STR region cannot be accurately determined or has low confidence, and a first score is obtained. Conversely, the repeat count data of the STR region has high confidence, and a first and second score are obtained. The first and second scores can be set according to actual needs; for example, the first and second scores can be greater than the first score. By determining the first preset score corresponding to the confidence level based on the confidence level of the repeat count data of the preset related genes within the short tandem repeat region, the accuracy of short tandem repeat analysis is improved.
[0083] In one optional implementation, the second preset score includes a second first score, a second second score, or a second third score;
[0084] Step S20, which involves determining the range of abnormal amplification counts related to a preset disease within the preset related genes based on the repeat count data of preset related genes within the short tandem repeat region and obtaining a second preset score corresponding to the range of abnormal amplification counts, includes:
[0085] Step S201: Based on the repeat count data of preset related genes in the short tandem repeat region, determine whether the maximum value of the abnormal amplification count related to the preset disease in the preset related genes is less than or equal to the reference maximum value of the normal amplification count.
[0086] Step S202: If the score is less than or equal to the maximum reference value for the number of normal amplifications, obtain the second score corresponding to the maximum value;
[0087] Step S203: If the number of amplifications is greater than the maximum reference value for normal amplifications, determine whether the maximum value is greater than or equal to the minimum reference value for abnormal amplifications.
[0088] Step S204: If the number of abnormal amplifications is greater than or equal to the minimum reference value, obtain the second score corresponding to the maximum value;
[0089] Step S205: If the score is less than the minimum reference value for the number of abnormal amplifications, obtain the second and third scores corresponding to the maximum value.
[0090] In this embodiment, the second first score, the second second score, and the second third score can be set according to actual needs. For example, the second second score can be greater than the second third score, and the second third score can be greater than the second first score. By determining the second preset score corresponding to the range of abnormal amplification times related to a preset disease in preset related genes, the accuracy of short tandem repeat analysis is improved.
[0091] In one optional implementation, the third preset score includes either the third first score or the third second score;
[0092] Step S30, which involves obtaining a third preset score corresponding to the coverage degree of the two alleles in the short tandem repeat region, includes:
[0093] Step S301: If the coverage of the two alleles in the short tandem repeat region satisfies the first case, obtain the third score corresponding to the coverage.
[0094] Step S302: If the coverage of the two alleles in the short tandem repeat region meets the second case, obtain the third score corresponding to the coverage.
[0095] The distinction between the first and second cases is determined by the number of consecutive coverages of short tandem repeats occurring in a single exon region or multiple exon regions.
[0096] In this embodiment of the application, the two alleles within the STR region are observed. The coverage of the two alleles in a high-confidence STR is relatively uniform. For example... Figure 2 As shown, the first category can include three cases: A, B, and C. Case A is where the STR is located on a single exon (i.e., spanning reads); case B is where the STR is continuous and spans multiple exon regions (i.e., in-repeat reads); and case C is where one allele is a spanning read and the other is an in-repeat read. Figure 3 As shown, the second category can include three cases: D, E, and F. Case D is where the coverage within the STR repeating region is lower than the coverage of the surrounding areas, such as... Figure 3 As indicated by the middle arrow, E indicates that the presence of multiple indels in in-repeat reads strongly suggests alignment or sequencing errors leading to an excessively high STR repeat count. Figure 3 As indicated by the middle arrow, F represents a short allele location with only one or very few spanning reads, or a situation in the STR where some flanking reads are overestimated, resulting in higher coverage than surrounding alleles. Figure 3 As indicated by the middle arrow.
[0097] The third and third-second scores can be set according to actual needs; for example, the third-first score can be greater than the third-second score. By determining the third preset score corresponding to the coverage degree of the two alleles in the short tandem repeat region, the accuracy of short tandem repeat analysis is improved.
[0098] In one optional implementation, the fourth preset score includes either the fourth first score or the fourth second score;
[0099] Step S40 involves analyzing the transcriptome sequencing data of the test sample to determine the enrichment level of preset related genes associated with preset diseases, and then determining the fourth preset score corresponding to the enrichment level, including:
[0100] Step S401: Based on the transcriptome sequencing data of the sample to be tested, analyze the average enrichment level of the preset related genes related to the preset disease, and determine whether the absolute value of the difference between the average value and the enrichment reference value is greater than or equal to the preset threshold.
[0101] Step S402: If it is greater than or equal to the preset threshold, obtain the fourth score corresponding to the average value;
[0102] Step S403: If it is less than the preset threshold, obtain the fourth score corresponding to the average value.
[0103] In this embodiment, the enrichment reference value can be based on gene expression data from different tissue sites (blood, muscle, adipose tissue, etc.) of healthy individuals collected from the GTEx database as a background value (i.e., enrichment reference value). The step of analyzing the enrichment level of preset related genes associated with preset diseases based on the transcriptome sequencing data of the test sample can involve using transcriptome expression data to sort the selected STR genes and their gene sets (GeneSET) according to their expression levels, scoring each gene in the gene set, and finally calculating an enrichment score (ES) for the gene set based on the cumulative distribution function (this process is called the ssGSEA analysis method) to calculate the association with the disease. Key gene sets can be identified based on the corresponding disease through literature or pathway analysis, as shown in the table below, using GTEx normal population samples to construct the expression baseline of STR-related genes.
[0104]
[0105]
[0106]
[0107] In this embodiment, the preset threshold can be set according to actual needs. A threshold greater than or equal to the preset threshold indicates a significant difference between the enrichment score of the preset related gene associated with the preset disease and the enrichment reference value (background value). A threshold less than the preset threshold indicates a difference between the enrichment score of the preset related gene associated with the preset disease and the enrichment reference value (background value), but without significant difference. The fourth score (number 41) and the fourth score (number 42) can be set according to actual needs; for example, the fourth score (number 41) can be greater than the fourth score (number 42).
[0108] As a concrete example, expression values from three tissues—fat, blood, and muscle—in the GTEx database can be used to calculate and visualize scores in the WiKiPathways (WP34, WP2857) and KEGG (hsa05017, has05022) pathways in the normal population using the ssGSEA analysis method. This data can then serve as a baseline for STR-related gene transcriptome analysis. Figure 4As shown. For example, for muscle tissue, the expression values of muscle tissue in the GTEx database can be used first, and the scores of normal individuals in the WiKiPathways (WP34, WP2857) and KEGG (hsa05017, has05022) pathways can be calculated using the ssGSEA analysis method, serving as a reference baseline for STR-related gene transcriptome analysis. Subsequently, the expression values of muscle transcriptome data from 23 patients with amyotrophic lateral sclerosis (ALS) (12 FUS-type ALS samples and 11 sporadic ALS samples) were used, and the scores calculated using the ssGSEA analysis method were compared with the aforementioned reference baseline. The comparison results are shown below. Figure 5 As shown. By Figure 5 It can be seen that the muscle transcriptome score (mean) of patients with amyotrophic lateral sclerosis (ALS) is significantly higher than the baseline data, that is, the ALS-related gene set in the patient sample is significantly different from that in the normal population.
[0109] Relying solely on sequencing results of short DNA reads to predict the possible number and length of STR repeats based on the statistical distribution of sequencing fragments only allows for analysis of samples under ideal sequencing conditions. In real-world scenarios, accurate analysis is often impossible due to sample quality, sequencing biases, and other issues, particularly regarding STR mutations on the boundary between pathogenic and non-pathogenic variants, which are especially difficult to identify. Furthermore, current analyses only target diseases associated with amplified repeats. However, besides being associated with certain diseases, non-pathogenic STR sites exist that can provide evidence for kinship analysis and loss of heterozygosity, so analyses in this area are currently lacking.
[0110] Therefore, in an optional implementation, determining the type of short serial repetition based on the first preset score, the second preset score, the third preset score, and the fourth preset score in step S50 includes:
[0111] Step S501: Determine the type of short tandem repeat based on the total score determined by the sum of the first preset score, the second preset score, the third preset score, and the fourth preset score, and the repeat count data of preset related genes in the short tandem repeat region.
[0112] In an optional implementation, step S501, which determines the type of short tandem repeat based on the total score determined by the sum of the first preset score, the second preset score, the third preset score, and the fourth preset score, and the repeat count data of preset related genes within the short tandem repeat region, includes:
[0113] Step S5011: If the maximum value of the number of abnormal amplifications related to the preset disease is less than the maximum value of the number of normal amplifications, and the total score is greater than the first score threshold, then the short tandem repeat is determined to be a short tandem repeat to be undetermined.
[0114] Step S5012: If the maximum value of the number of abnormal amplifications related to the preset disease is less than the maximum value of the number of normal amplifications, and the total score is less than or equal to the first score threshold, the short tandem repeat is determined to be a polymorphic short tandem repeat.
[0115] Step S5013: If the maximum value of the number of abnormal amplifications related to the preset disease is greater than or equal to the maximum value of the number of normal amplifications and less than or equal to the minimum value of the number of abnormal amplifications, and the total score is greater than the first score threshold, then the short tandem repeat is determined to be a short tandem repeat to be undetermined.
[0116] Step S5014: If the maximum value of the number of abnormal amplifications related to the preset disease is greater than or equal to the maximum value of the number of normal amplifications and less than or equal to the minimum value of the number of abnormal amplifications, and the total score is less than or equal to the first score threshold, then the short tandem repeat is determined to be a polymorphic short tandem repeat.
[0117] Step S5015: If the maximum value of the number of abnormal amplifications related to the preset disease is greater than the minimum reference value of the number of abnormal amplifications, and the total score is less than or equal to the second score threshold, the short tandem repeat is determined to be a short tandem repeat to be undetermined.
[0118] Step S5016: If the maximum value of the number of abnormal amplifications related to the preset disease is greater than the minimum reference value of the number of abnormal amplifications, and the total score is greater than the second score threshold, then the short tandem repeat is determined to be a pathogenic short tandem repeat.
[0119] In this embodiment, the type of short tandem repeat is determined based on a first preset score, a second preset score, a third preset score, and a fourth preset score. A quality control standard is introduced to control the analysis and ensure the accuracy of STR analysis. Furthermore, analysis is performed in conjunction with transcriptome data, improving the accuracy of STR analysis. In addition, various short tandem repeat types, including polymorphic short tandem repeats, pathogenic short tandem repeats, and undetermined short tandem repeats, have been identified. This provides analysis capabilities for non-pathogenic STR sites and can provide a basis for other analyses such as kinship and loss of heterozygosity.
[0120] As a specific example, such as Figure 6As shown, the determination process of the total score is: 1. Perform comparison and analysis on the maximum value of abnormal amplification times related to a preset disease in preset related genes within the STR region of the sample to be tested and the reference maximum value of normal amplification times (norm_up). If the maximum value is less than or equal to norm_up, the second preset score is a second-first score, and the STR can be marked as a candidate polymorphic STR. If the maximum value is greater than norm_up, proceed to the next judgment, and perform comparison and analysis between the maximum value and the reference minimum value of abnormal amplification times (aff_low). If the maximum value is greater than or equal to aff_low, the second preset score is a second-second score, and the STR can be marked as a candidate pathogenic STR. If the maximum value is less than aff_low, the second preset score is a second-third score.
[0121] As a specific example, the process of determining the type of short tandem repeat based on the total score and the repeat times data of preset related genes in the short tandem repeat region is as follows: 1) If the number of DNA fragments meets the preset quantity condition, 2 points are assigned; if the preset quantity condition is not met, 1 point is assigned. 2) If the repeat times < norm_up, no point is assigned; if norm_up ≤ repeat times ≤ aff_low, 1 point is assigned; if repeat times > aff_low, 2 points are assigned. 3) If the coverage degree meets A-C, 2 points are assigned; if it meets D-F, 2 points are deducted. 4) If the enrichment degree is significantly correlated with the pathogenic mechanism of STR-related diseases, 2 points are assigned; if it does not meet the requirement, 2 points are deducted. 5) Record the sum of all the above scores as the total score; 6) Determine the STR classification based on the total score and the repeat times data of preset related genes in the short tandem repeat region, as shown in the following table:
[0122]
[0123] It can be seen from the above table that when the number of repeats < norm_up and the total score > 2, it indicates that there may be a false negative result caused by the limitation of high-throughput sequencing, and it is necessary to supplement other experiments such as capillary electrophoresis or third-generation sequencing for supplementary analysis; when the number of repeats < norm_up and the total score ≤ 2, it indicates that this STR is a polymorphic STR with high confidence. When norm_up ≤ the number of repeats ≤ aff_low and the total score > 2, it indicates that there may be a false negative result caused by the limitation of high-throughput sequencing, and it is necessary to supplement other experiments such as capillary electrophoresis or third-generation sequencing for supplementary analysis; when norm_up ≤ the number of repeats ≤ aff_low and the total score ≤ 2, it indicates that this STR is a polymorphic STR with high confidence. When the number of repeats > aff_low and the total score ≤ 4, it indicates that this STR is a pathogenic STR but with low confidence, and it is necessary to supplement other experiments such as capillary electrophoresis or third-generation sequencing for supplementary analysis; when the number of repeats > aff_low and the total score > 4, it indicates that this STR is a pathogenic STR with high confidence.
[0124] As a specific example, 15 samples containing disease-related abnormally amplified genes detected by dynamic mutation products were selected for exome sequencing, and analyzed according to the short tandem repeat analysis method of the examples of the present application. Finally, 11 WES data positive for disease-related abnormal amplification were obtained, with a total of 180 records. After combined transcriptome analysis, all pathogenic STRs were identified. The results are shown in the following table. Among them, 8 were identified as pathogenic STRs and 169 were polymorphic STRs. Partial detection results are shown in the following table:
[0125]
[0126] Among them, for the sample to be tested CN-2321543, the preset related gene is FMR1. The identification result of pathogenic STR can be obtained by the short tandem repeat analysis method of the example of the present application, while by Figure 6 the analysis method (that is, analysis only based on exome sequencing data) that obtains candidate polymorphic STRs and candidate pathogenic STRs after separately comparing the maximum value of the preset disease-related abnormal amplification times in the preset related genes in the STR region with norm_up and aff_low shown in , did not identify the CN-2321543 sample to be tested. It can be seen that the analysis performed by combining transcriptome data improves the accuracy of STR analysis.
[0127] As shown in the table above, the sensitivity of analysis based solely on exome sequencing data is TPR (true positive rate) = TP / (TP+FN) = 8 / (8+1) = 88.89%, where TP represents the number of true positives (i.e., the number of pathogenic STRs identified) and FN represents the number of false negatives. The specificity is TNR (true negative rate) = TN / (FP+TN) = (180-8) / (0+172) = 100%, where TN represents the number of true negatives and FP represents the number of false positives.
[0128] The sensitivity TPR of the short tandem repeat analysis method in this application embodiment is 9 / 9 = 100%, which shows that the sensitivity of the short tandem repeat analysis method in this application embodiment can be improved to 100%.
[0129] For the combined transcriptome analysis in this application embodiment, the reverse transcription PCR method can be used to detect the target, and the expression and sequence information of the target gene can also be detected.
[0130] This application also provides an analysis apparatus for short series repetitions, corresponding to the above-described analysis method for short series repetitions. For example... Figure 7 As shown, the short tandem repeating analysis device 100 includes:
[0131] The first determining unit 10 is used to determine the confidence level of the repeat count data of preset related genes in the short tandem repeat region of the test sample obtained by analyzing the exome sequencing data of the test sample and obtain the first preset score corresponding to the confidence level.
[0132] The second determining unit 20 is used to determine the range of abnormal amplification times related to a preset disease in the preset related genes based on the repeat count data of preset related genes in the short tandem repeat region and to obtain a second preset score corresponding to the range of abnormal amplification times.
[0133] The third determining unit 30 is used to obtain a third preset score corresponding to the coverage degree of the two alleles in the short tandem repeat region.
[0134] The fourth determining unit 40 is used to determine the fourth preset score corresponding to the enrichment degree of preset related genes related to preset diseases based on the enrichment degree of preset related genes obtained from the transcriptome sequencing data analysis of the sample to be tested.
[0135] The fifth determining unit 50 is used to determine the type of short tandem repeat based on the first preset score, the second preset score, the third preset score, and the fourth preset score; the types include polymorphic short tandem repeat, pathogenic short tandem repeat, and undetermined short tandem repeat.
[0136] In one optional implementation, the first preset score includes either a first score or a first second score;
[0137] The first determining unit 10 includes:
[0138] The first judgment unit is used to determine whether the number of deoxyribonucleic acid fragments in the short tandem repeat region of the test sample obtained by analyzing the exome sequencing data of the test sample meets the preset number condition; the preset number condition is determined according to the preset number of deoxyribonucleic acid fragments.
[0139] The first obtaining unit is used to determine the confidence level of the repeat count data of the preset related genes in the short tandem repeat region as high confidence level and obtain the first score corresponding to the high confidence level if the preset quantity condition is met.
[0140] The second obtaining unit is used to determine the confidence level of the repeat count data of the preset related genes in the short tandem repeat region as low confidence level and obtain the first and second scores corresponding to the low confidence level if the preset quantity condition is not met.
[0141] In one optional implementation, the second preset score includes a second first score, a second second score, or a second third score;
[0142] The second determining unit 20 includes:
[0143] The second judgment unit is used to determine whether the maximum value of the abnormal amplification number related to the preset disease in the preset related genes is less than or equal to the reference maximum value of the normal amplification number, based on the repeat count data of the preset related genes in the short tandem repeat region.
[0144] The third obtaining unit is used to obtain the second score corresponding to the maximum value if the score is less than or equal to the reference maximum value of the normal amplification number.
[0145] The third judgment unit is used to determine whether the maximum value is greater than or equal to the minimum value of the abnormal amplification number if the number of amplifications is greater than the maximum reference value of the normal amplification number.
[0146] The fourth obtaining unit is used to obtain the second score corresponding to the maximum value if the number of aberrations is greater than or equal to the reference minimum value.
[0147] The fifth obtaining unit is used to obtain the second and third scores corresponding to the maximum value if the score is less than the minimum reference value for the number of abnormal amplifications.
[0148] In one optional implementation, the third preset score includes either the third first score or the third second score;
[0149] The third determining unit 30 includes:
[0150] The sixth obtaining unit is used to obtain the third score corresponding to the coverage degree if the coverage degree of the two alleles in the short tandem repeat region meets the first case.
[0151] The seventh unit is used to obtain the third score corresponding to the coverage degree if the coverage of two alleles in the short tandem repeat region meets the second case.
[0152] The distinction between the first and second cases is determined by the number of consecutive coverages of short tandem repeats occurring in a single exon region or multiple exon regions.
[0153] In one optional implementation, the fourth preset score includes either the fourth first score or the fourth second score;
[0154] The fourth determining unit 40 includes:
[0155] The fourth judgment unit is used to determine whether the absolute value of the difference between the average enrichment level of the preset related genes and the preset disease obtained from the transcriptome sequencing data analysis of the sample to be tested is greater than or equal to the preset threshold.
[0156] The eighth obtaining unit is used to obtain the fourth score corresponding to the average value if it is greater than or equal to a preset threshold.
[0157] The ninth obtaining unit is used to obtain the fourth score corresponding to the average value if the score is less than a preset threshold.
[0158] In an optional embodiment, the fifth determining unit 50 includes:
[0159] The sixth determining unit is used to determine the type of short tandem repeat based on the total score determined by the sum of the first preset score, the second preset score, the third preset score, and the fourth preset score, and the repeat count data of preset related genes in the short tandem repeat region.
[0160] In an optional implementation, the sixth determining unit includes:
[0161] The tenth obtaining unit is used to determine short tandem repeats as undetermined short tandem repeats if the maximum value of the number of abnormal amplifications related to a preset disease is less than the maximum value of the number of normal amplifications and the total score is greater than the first score threshold.
[0162] The eleventh obtaining unit is used to determine that the short tandem repeat is a polymorphic short tandem repeat if the maximum value of the number of abnormal amplifications related to the preset disease is less than the maximum value of the normal amplifications reference, and the total score is less than or equal to the first score threshold.
[0163] The twelfth obtaining unit is used to determine short tandem repeats as undetermined short tandem repeats if the maximum value of the number of abnormal amplifications related to the preset disease is greater than or equal to the maximum value of the number of normal amplifications and less than or equal to the minimum value of the number of abnormal amplifications, and the total score is greater than the first score threshold.
[0164] The thirteenth obtaining unit is used to determine that the short tandem repeat is a polymorphic short tandem repeat if the maximum value of the abnormal amplification number related to the preset disease is greater than or equal to the reference maximum value of the normal amplification number and less than or equal to the reference minimum value of the abnormal amplification number, and the total score is less than or equal to the first score threshold.
[0165] The fourteenth obtaining unit is used to determine short tandem repeats as undetermined short tandem repeats if the maximum value of the number of abnormal amplifications related to a preset disease is greater than the minimum reference value of the number of abnormal amplifications and the total score is less than or equal to the second score threshold.
[0166] The fifteenth obtaining unit is used to determine that short tandem repeats are pathogenic short tandem repeats if the maximum value of the number of abnormal amplifications related to a preset disease is greater than the minimum reference value of the number of abnormal amplifications and the total score is greater than the second score threshold.
[0167] This application also provides an analysis device for short serial repetitions, including a processor and a memory. The memory stores computer-executable instructions, which are executed by the processor to perform the above-described analysis method for short serial repetitions.
[0168] The processor may be a microcontroller unit (MCU) or other form of processing unit with data processing and / or instruction execution capabilities, and may control other components in a short series of repetitive analysis devices to perform the desired functions.
[0169] The memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor may execute the program instructions to implement the steps in the above-described short serial repetition analysis method and / or other desired functions.
[0170] In an alternative embodiment, the short-serialized repetitive analysis device may further include input devices and output devices, which are interconnected via a bus system and / or other forms of connection mechanisms. Furthermore, the input devices may include, for example, a keyboard, mouse, microphone, etc. The output devices can output various information externally, and may include, for example, a display, speaker, printer, and communication networks and their connected remote output devices, etc.
[0171] This application also provides a computer-readable storage medium storing computer-executable instructions, which are executed by a processor to perform the above-described short serial repetition analysis method.
[0172] The embodiments of this application may be systems, methods, and / or computer program products. A computer program product may include a computer-readable storage medium loaded with computer-readable program instructions for causing a processor to implement various aspects of the embodiments of this application. The computer program product may be written in any combination of one or more programming languages to perform operations of the embodiments of this application. Programming languages include object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on a user's computing device, partially on a user's device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computers, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing state information of computer-readable program instructions. These electronic circuits can execute computer-readable program instructions to implement various aspects of the embodiments of this application.
[0173] Computer-readable storage media can take the form of any combination of one or more readable media. A readable medium can be a readable signal medium or a readable storage medium. A computer-readable storage medium is a tangible device capable of holding and storing instructions for use by an instruction execution device. A readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combinations thereof. The computer-readable storage medium as used herein is not to be construed as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0174] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0175] Various aspects of embodiments of this application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0176] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0177] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0178] It should be noted that the short series repetition analysis method embodiments, short series repetition analysis device embodiments, storage medium embodiments, and short series repetition analysis equipment embodiments provided in this application belong to the same concept; the technical features in the technical solutions described in each embodiment can be arbitrarily combined without conflict.
[0179] It should be understood that the above embodiments are exemplary and are not intended to encompass all possible implementations included in the claims. Various modifications and changes can be made to the above embodiments without departing from the scope of this disclosure. Similarly, the various technical features of the above embodiments can be arbitrarily combined to form other embodiments of this application that may not be explicitly described. Therefore, the above embodiments only illustrate several implementations of this application and do not limit the scope of protection of this patent application.
Claims
1. A method for analyzing short tandem repeats, characterized in that, include: Based on the analysis of the exome sequencing data of the test sample, the short tandem repeat region of the test sample is obtained, the confidence level of the repeat count data of the preset related genes in the short tandem repeat region is determined, and a first preset score corresponding to the confidence level is obtained. Based on the repeat count data of preset related genes within the short tandem repeat region, determine the range of abnormal amplification counts related to preset diseases in the preset related genes and obtain a second preset score corresponding to the range of abnormal amplification counts; A third preset score corresponding to the coverage degree of the two alleles in the short tandem repeat region is obtained. Based on the enrichment degree of the preset related genes related to the preset disease obtained from the transcriptome sequencing data analysis of the sample to be tested, a fourth preset score corresponding to the enrichment degree is determined; The type of short tandem repeat is determined based on the first preset score, the second preset score, the third preset score, and the fourth preset score; the type includes polymorphic short tandem repeat, pathogenic short tandem repeat, and undetermined short tandem repeat.
2. The method for analyzing short tandem repeats according to claim 1, characterized in that, The first preset score includes either a first score or a first second score; The step of determining the confidence level of the repeat count data of preset related genes within the short tandem repeat region of the test sample obtained from the analysis of exome sequencing data of the test sample, and obtaining a first preset score corresponding to the confidence level, includes: Based on the analysis of the exome sequencing data of the sample to be tested, the short tandem repeat regions of the sample are obtained, and it is determined whether the number of deoxyribonucleic acid fragments in the short tandem repeat regions meets the preset number condition; the preset number condition is determined based on the preset number of deoxyribonucleic acid fragments. If the preset quantity condition is met, the confidence level of the repeat count data of the preset related genes in the short tandem repeat region is determined to be high confidence level and the first score corresponding to the high confidence level is obtained; If the preset quantity condition is not met, the confidence level of the repeat count data of the preset related genes in the short tandem repeat region is determined to be low confidence level, and the first and second scores corresponding to the low confidence level are obtained.
3. The method for analyzing short tandem repeats according to claim 2, characterized in that, The second preset score includes a second first score, a second second score, or a second third score; The step of determining the range of abnormal amplification counts related to a preset disease in the preset related genes based on the repeat count data of preset related genes within the short tandem repeat region and obtaining a second preset score corresponding to the range of abnormal amplification counts includes: Based on the repeat count data of preset related genes in the short tandem repeat region, determine whether the maximum value of abnormal amplification count related to preset disease in the preset related genes is less than or equal to the reference maximum value of normal amplification count; If the score is less than or equal to the maximum reference value for the number of normal amplifications, a second score corresponding to the maximum value is obtained. If the number of amplifications is greater than the maximum reference value for normal amplifications, determine whether the maximum value is greater than or equal to the minimum reference value for abnormal amplifications. If the number of abnormal amplifications is greater than or equal to the minimum reference value, a second score corresponding to the maximum value is obtained; If the score is less than the minimum reference value for the number of abnormal amplifications, the second or third score corresponding to the maximum value is obtained.
4. The method for analyzing short tandem repeats according to claim 3, characterized in that, The third preset score includes either the third first score or the third second score; The step of obtaining a third preset score corresponding to the coverage degree of the two alleles in the short tandem repeat region includes: If the coverage of the two alleles in the short tandem repeat region meets the first condition, a third score corresponding to the coverage is obtained. If the coverage of the two alleles in the short tandem repeat region meets the second condition, a third score corresponding to the coverage is obtained. The distinction between the first and second types of cases is determined based on the number of consecutive coverages of the short tandem repeats occurring in a single exon region or multiple exon regions.
5. The method for analyzing short tandem repeats according to claim 4, characterized in that, The fourth preset score includes either the fourth first score or the fourth second score; The step of determining a fourth preset score corresponding to the enrichment degree of preset related genes related to preset diseases based on the transcriptome sequencing data analysis of the test sample includes: Based on the average enrichment level of the preset related genes related to the preset disease obtained from the transcriptome sequencing data analysis of the sample to be tested, determine whether the absolute value of the difference between the average value and the enrichment reference value is greater than or equal to a preset threshold. If it is greater than or equal to a preset threshold, obtain the fourth score corresponding to the average value; If it is less than the preset threshold, the fourth score corresponding to the average value is obtained.
6. The method for analyzing short tandem repeats according to claim 5, characterized in that, Determining the type of the short serial repetition based on the first preset score, the second preset score, the third preset score, and the fourth preset score includes: The type of short tandem repeat is determined based on the total score determined by the sum of the first preset score, the second preset score, the third preset score, and the fourth preset score, and the repeat count data of preset related genes in the short tandem repeat region.
7. The method for analyzing short tandem repeats according to claim 6, characterized in that, The determination of the type of short tandem repeat based on the total score determined by the sum of the first preset score, the second preset score, the third preset score, and the fourth preset score, and the repeat count data of preset related genes within the short tandem repeat region, includes: If the maximum value of the number of abnormal amplifications related to the preset disease is less than the maximum value of the number of normal amplifications, and the total score is greater than the first score threshold, the short tandem repeat is determined to be a short tandem repeat to be undetermined. If the maximum value of the number of abnormal amplifications related to the preset disease is less than the maximum value of the number of normal amplifications, and the total score is less than or equal to the first score threshold, the short tandem repeat is determined to be a polymorphic short tandem repeat. If the maximum value of the number of abnormal amplifications related to the preset disease is greater than or equal to the maximum value of the number of normal amplifications and less than or equal to the minimum value of the number of abnormal amplifications, and the total score is greater than the first score threshold, the short tandem repeat is determined to be a short tandem repeat to be undetermined. If the maximum value of the number of abnormal amplifications related to the preset disease is greater than or equal to the maximum value of the number of normal amplifications and less than or equal to the minimum value of the number of abnormal amplifications, and the total score is less than or equal to the first score threshold, the short tandem repeat is determined to be a polymorphic short tandem repeat. If the maximum value of the number of abnormal amplifications related to the preset disease is greater than the minimum reference value of the number of abnormal amplifications, and the total score is less than or equal to the second score threshold, the short tandem repeat is determined to be a short tandem repeat to be undetermined. If the maximum value of the number of abnormal amplifications related to a preset disease is greater than the minimum reference value for the number of abnormal amplifications, and the total score is greater than the second score threshold, the short tandem repeat is determined to be a pathogenic short tandem repeat.
8. An analytical apparatus for short series repetition, characterized in that, include: The first determining unit is used to determine the confidence level of the repeat count data of a preset related gene in the short tandem repeat region of the test sample obtained by analyzing the exome sequencing data of the test sample and to obtain a first preset score corresponding to the confidence level. The second determining unit is used to determine the range of abnormal amplification times related to a preset disease in the preset related genes based on the repeat count data of preset related genes in the short tandem repeat region, and to obtain a second preset score corresponding to the range of abnormal amplification times. The third determining unit is used to obtain a third preset score corresponding to the coverage degree of the two alleles in the short tandem repeat region. The fourth determining unit is used to determine a fourth preset score corresponding to the enrichment degree of the preset related genes related to the preset disease obtained from the transcriptome sequencing data analysis of the sample to be tested. The fifth determining unit is used to determine the type of the short tandem repeat based on the first preset score, the second preset score, the third preset score, and the fourth preset score; the type includes polymorphic short tandem repeat, pathogenic short tandem repeat, and undetermined short tandem repeat.
9. An analytical device for short tandem repetition, characterized in that, It includes a processor and a memory, the memory storing computer-executable instructions, which are executed by the processor at runtime as described in any one of claims 1-7, for the analysis method of short serial repetitions.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, perform the short serial repetition analysis method as described in any one of claims 1-7.
Citation Information
Patent Citations
Pathogenicity analysis method and device for short tandem repeat sequence, and server
CN113990392A
Method for detecting short tandem repeat expansion and genotyping, electronic equipment and storage medium
CN115240770A