A method and system for processing whole exome sequencing data, and a system for detecting disease-associated abnormal expansion of short tandem repeats

By using ExpansionHunter and exSTRa software in whole-exome sequencing data, and combining actual coverage and amplification times for comparison, the problem of STR analysis results being affected by algorithms, platforms, and software was solved, enabling more accurate diagnosis of short tandem repeat diseases.

CN115312120BActive Publication Date: 2026-05-29CIPHERGENE BEIJING TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CIPHERGENE BEIJING TECH CO LTD
Filing Date
2022-08-08
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In current technologies, STR analysis results in whole-exome sequencing data are greatly affected by algorithms, sequencing platforms, and alignment software, leading to inaccurate STR typing and difficulty in accurately diagnosing short tandem duplication diseases.

Method used

By acquiring reference data and test sample data, and using ExpansionHunter and exSTRa software for analysis, the software prediction results are corrected by comparing the actual coverage and amplification times, thus determining the coverage and amplification times within the target regions of STR-related disease genes and reducing the impact of algorithms, platforms, and software.

Benefits of technology

It improved the accuracy of STR analysis of whole-exome sequencing data, reduced false positive and false negative results, and improved the diagnostic accuracy of short tandem repeat diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003787160190000091
    Figure BDA0003787160190000091
  • Figure BDA0003787160190000111
    Figure BDA0003787160190000111
  • Figure BDA0003787160190000141
    Figure BDA0003787160190000141
Patent Text Reader

Abstract

The application provides a processing method and processing system of whole exome sequencing data and a system for detecting short tandem repeat disease-related abnormal amplification. The application defines the STR-related genes that can be detected in the sample by the actual sample true coverage in the WES sequencing data, which is more accurate than the evaluation of whether the bed region of the WES probe and the bed+flanking region overlap. The processing method of whole exome sequencing data provided by the application is less affected by different algorithms, different sequencing platforms, different probes and different alignment software, and the data results obtained are more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical technology, and in particular to a method and system for processing whole-exome sequencing data, and a system for detecting disease-related amplifications of short tandem repeats. Background Technology

[0002] Exons comprise only about 1-2% of the total genome sequence, yet they contain up to 85% of disease-related pathogenic variants. Whole-exome sequencing (WES) is a high-throughput sequencing method that enriches DNA from the exon regions of the entire genome using sequence capture or targeting technologies. Due to its comprehensiveness, effectiveness, and extremely high cost-effectiveness, WES has become the preferred molecular diagnostic approach for most clinically heterogeneous diseases. It can simultaneously detect SNPs, Indels, and CNVs. With the addition of mitochondrial loop capture, it can also simultaneously detect mitochondrial loop gene variants.

[0003] Short tandem repeats (STRs) are DNA sequences consisting of 1 to 6 motifs. The number of repeats varies greatly among individuals and is numerous, exhibiting rich genetic polymorphism. It is estimated that there are over one million STR sites in the human genome, accounting for approximately 3% of the human genome. Amplification of short tandem repeats can lead to a range of diseases, including Huntington's disease, various ataxias, amyotrophic lateral sclerosis (ALS), frontotemporal dementia, Fragile X syndrome, and other neurological disorders. Numerous studies have also shown that tandem repeat polymorphism (TRP) plays an important role in gene expression regulation in polygenic diseases. Tandem repeat-related diseases (TRDs) are not phenotypically defined as simply having or not having them (comparing affected and healthy individuals). Due to their specific characteristics, variations in the number of tandem repeats typically lead to continuous magnitude changes (such as age of onset, disease severity, etc.).

[0004] Currently, routine molecular diagnosis for these diseases relies on precise PCR amplification or Southern blotting analysis. This requires laboratories to accurately amplify each different repetitive sequence, and clinicians need to make accurate diagnoses, determine which diseases are most likely associated with the patient, and submit appropriate tests. However, STR-related diseases overlap in clinical symptoms, penetrance variation, and onset time, mainly depending on allele size and the role of modifying genes. In up to 50% of ataxia patients, other mutations such as SNPs and indels may be the cause. Therefore, when performing molecular diagnosis for these diseases, routine sequencing of candidate genes, such as panel sequencing and WES, is usually required. Some genetic diseases have extremely high clinical phenotypic heterogeneity, which may lead to misdiagnosis and the selection of inappropriate testing methods, resulting in patients not receiving a correct molecular diagnosis. For example, dentate globus pallidoluysianatrophy (DRPLA) is a progressive autosomal dominant genetic disorder characterized by myoclonic epilepsy, ataxia, choreoathetosis / dystonia, cognitive impairment, dementia, and mental disorders. It is caused by tandem repeats of the CAG trinucleotide in the ATN1 gene; the normal number of repeats is 7-23, while those affected typically have 49-88 repeats. DRPLA has an age of onset ranging from 0 to 70 years, with an average age of 30 years. Clinical manifestations vary depending on the age of onset: children are characterized by ataxia, intellectual disability, behavioral changes, myoclonus, and epilepsy; adults are characterized by ataxia, choreoathetosis, and dementia. Patients with onset before age 20 typically present with a progressive myoclonic epilepsy (PME) phenotype, characterized by myoclonus, seizures, ataxia, and progressive intellectual decline. Various forms of generalized seizures (including tonic, atonic, clonic, or tonic-clonic seizures) have also been observed. Routine WES testing is often recommended for early-onset patients due to their prediagnosis of seizures and intellectual disability.

[0005] While next-generation sequencing (NGS) technology has enabled the detection of millions of STRs across the entire genome, genotyping in bioinformatics analysis remains challenging: high GC content, inability to cover short-read sequences with complete repetitions, mapping to large deletions / insertions of STR variants that differ from the reference genome, poor or unmapped repetitive characteristics of the repetitive sequences themselves, and noise from stutter products (shadow bands or DNA polymerase slip products) caused by PCR amplification. Although Illumina has developed an amplification-free (PCR-) library preparation method that eliminates STR stuttering errors during PCR amplification in sample preparation (PCR+), improving the accuracy of STR genotyping, the PCR+ method has already generated a large amount of sequencing data, and the PCR- method still faces limitations in terms of cost and complexity. While PCR-based whole-genome sequencing (WGS) has many advantages, whole-genome sequencing (WES) plays a crucial role in human genetic disease research and diagnosis due to its low cost and high coverage. In domestic genetic diagnosis, WES is the preferred detection method for diseases with high clinical heterogeneity; therefore, accurate STR genotyping from PCR+ sequencing data is essential. For WES data, numerous tools have been developed for STR genotyping, but most are limited to detecting STRs within a specific read length range. Furthermore, due to differences in their algorithms and principles, they all have limitations in identifying disease-related STRs. For example, exSTRa, an algorithm primarily used to detect user-specified STR sequences in sequencing cohort samples, is an outlier detection method that assumes that most (>85%) individuals have normal alleles at specific STR loci. Another example is ExpansionHunter, mainly used for STR analysis of WGS data, which tends to rely on PCR-based library preparation and uses predetermined thresholds to determine the presence of STR amplification in individuals.

[0006] Current research on STR analysis of NGS short-read sequencing data focuses more on analysis algorithms. However, different algorithms, sequencing platforms, probes, and alignment software all have a significant impact on the final STR analysis results, resulting in most molecular diagnostic samples having different degrees of outliers indicated by the analysis software. Summary of the Invention

[0007] In view of this, the present invention provides a method and system for processing whole exome sequencing data, and a system for detecting disease-related abnormal amplifications of short tandem repeats. The whole exome sequencing data processing method provided by the present invention is less affected by different algorithms, sequencing platforms, probes, and alignment software, and the obtained data results are more accurate.

[0008] This invention provides a method for processing whole-exome sequencing data, comprising the following steps:

[0009] Step S1: Obtain first reference data, which includes base coverage data of the target region of STR-related disease genes in the reference sample and sample proportion data under the predetermined base coverage.

[0010] Obtain second reference data, which includes data on the number of amplifications of the negative reference sample;

[0011] Step S2: Obtain test sample data, which includes base coverage data of the target region of STR-related disease genes in the test sample, sample proportion data under the predetermined base coverage, and test sample amplification number data.

[0012] The base coverage data and sample proportion data under the predetermined base coverage of the STR-related disease gene target region of the test sample are compared with the base coverage data and sample proportion data under the predetermined base coverage of the reference sample to obtain the first comparison result data.

[0013] If the first comparison result data is inconsistent, the amplification count data of the detected sample is compared with the amplification count data of the negative reference sample to obtain the second comparison result data.

[0014] In some embodiments, step S1 specifically includes:

[0015] Acquire STR-related disease gene data and identify target regions for abnormal amplification in STR-related diseases;

[0016] Obtain WES sequencing data of a reference sample, compare the WES sequencing data of the reference sample with the target region data of abnormal amplification of STR-related diseases, and obtain the base coverage data of the target region of STR-related disease genes in the reference sample and the sample proportion data under the predetermined base coverage.

[0017] The WES sequencing data of negative samples in the reference sample were analyzed using the ExpansionHunter software to obtain the amplification count data of the negative reference samples.

[0018] In some embodiments, step S2 specifically includes:

[0019] Obtain WES sequencing data of the test sample, compare the WES sequencing data of the test sample with the target region data of abnormal amplification of STR-related diseases, and obtain the base coverage data of the target region of STR-related disease genes in the test sample and the sample proportion data under the predetermined base coverage.

[0020] The WES sequencing data of the test samples were analyzed using the ExpansionHunter software to obtain the amplification count data of the test samples.

[0021] In some embodiments, step S1 further includes: correcting the amplification count data of the negative reference sample to obtain corrected amplification count data of the negative reference sample.

[0022] In some embodiments, the correction of the amplification count data of the negative reference sample specifically involves:

[0023] Obtain data on the actual number of amplifications and WES sequencing data for positive samples;

[0024] The WES sequencing data of positive samples were analyzed using ExpansionHunter software to obtain the predicted number of amplifications for the positive samples.

[0025] The amplification count data of the negative reference sample is corrected based on the actual amplification count data and the predicted amplification count data of the positive sample.

[0026] In some embodiments, it also includes:

[0027] The WES sequencing data of the reference sample were analyzed using exSTRa software to obtain the STR scores of the positive samples in the reference sample.

[0028] The WES sequencing data of the test samples were analyzed using exSTRa software to obtain the STR score of the test samples.

[0029] The present invention also provides a whole exome sequencing data processing system, including a first reference data acquisition unit, the first reference data unit being used to acquire first reference data, the first reference data including base coverage data of STR-related disease gene target regions of reference samples and sample proportion data under predetermined base coverage;

[0030] The second reference data acquisition unit is used to acquire second reference data, which includes negative reference sample amplification count data.

[0031] The test sample data acquisition unit is used to acquire test sample data, which includes base coverage data of the target region of STR-related disease genes of the test sample, sample proportion data under the predetermined base coverage, and test sample amplification number data.

[0032] The first comparison unit is used to compare the base coverage data and the sample proportion data under the predetermined base coverage in the target region of the STR-related disease gene of the test sample with the base coverage data and the sample proportion data under the predetermined base coverage in the target region of the STR-related disease gene of the reference sample to obtain the first comparison result data.

[0033] The second comparison unit is used to compare the amplification count data of the detection sample with the amplification count data of the negative reference sample to obtain the second comparison result data.

[0034] In some embodiments, the first reference data acquisition unit includes an STR-related disease gene data acquisition unit, which is used to acquire STR-related disease gene data and determine target region data for abnormal amplification of STR-related diseases.

[0035] A reference sample WES sequencing data acquisition unit is used to acquire WES sequencing data of a reference sample.

[0036] The third comparison unit is used to compare the WES sequencing data of the reference sample with the target region data of the abnormal amplification of the STR-related disease, and to obtain the base coverage data of the STR-related disease gene target region of the reference sample and the sample proportion data under the predetermined base coverage.

[0037] In some embodiments, the test sample data acquisition unit includes a test sample WES sequencing data acquisition unit, which is used to acquire the WES sequencing data of the test sample.

[0038] The fourth comparison unit is used to compare the WES sequencing data of the test sample with the target region data of the abnormal amplification of the STR-related disease to obtain the base coverage data of the STR-related disease gene target region of the test sample and the sample proportion data under the predetermined base coverage.

[0039] The sample amplification count data processing unit is used to analyze the WES sequencing data of the test samples using ExpansionHunter software to obtain the sample amplification count data.

[0040] The present invention also provides a system for detecting disease-related amplification of short tandem repeats, including a first reference data acquisition unit, the first reference data unit being used to acquire first reference data, the first reference data including base coverage data of the target region of STR-related disease genes of the reference sample and sample proportion data under a predetermined base coverage.

[0041] The second reference data acquisition unit is used to acquire second reference data, which includes negative reference sample amplification count data.

[0042] The test sample data acquisition unit is used to acquire test sample data, which includes base coverage data of the target region of STR-related disease genes of the test sample, sample proportion data under the predetermined base coverage, and test sample amplification number data.

[0043] The first comparison unit is used to compare the base coverage data and the sample proportion data under the predetermined base coverage in the target region of the STR-related disease gene of the test sample with the base coverage data and the sample proportion data under the predetermined base coverage in the target region of the STR-related disease gene of the reference sample to obtain the first comparison result data.

[0044] The second comparison unit is used to compare the amplification count data of the detection sample with the amplification count data of the negative reference sample to obtain the second comparison result data;

[0045] A prediction system is used to obtain prediction results of short tandem repeat disease-related amplifications based on first and second alignment result data.

[0046] This invention defines detectable STR-related genes in samples based on the actual coverage of the samples in WES sequencing data, which is more accurate than using the overlap of the bed region / bed+flanking region of the WES probe for evaluation. The whole-exome sequencing data processing method provided by this invention is less affected by different algorithms, sequencing platforms, probes, and alignment software, resulting in more accurate data results. Attached Figure Description

[0047] Figure 1 This is a schematic diagram of the process for detecting disease-related abnormal amplifications of short tandem repeats provided in an embodiment of the present invention;

[0048] Figure 2 This is a flowchart illustrating the process of obtaining the coverage of STR-related disease gene target regions according to an embodiment of the present invention;

[0049] Figure 3 These are actual observed values ​​and software-predicted values ​​within the normal amplification range of STR disease-related genes.

[0050] Figure 4 This is a schematic diagram of the filtering indicators and evaluation criteria for pathogenic STR variants. Detailed Implementation

[0051] This invention provides a method and system for processing whole-exome sequencing data, and a system for detecting disease-related amplifications of short tandem repeats. Those skilled in the art can refer to the content of this document and appropriately modify the process parameters to achieve the desired results. It should be particularly noted that all similar substitutions and modifications are obvious to those skilled in the art and are considered to be included in this invention. The methods and applications of this invention have been described through preferred embodiments, and those skilled in the art can obviously make modifications or appropriate alterations and combinations to the methods and applications described herein without departing from the content, spirit, and scope of this invention to implement and apply the technology of this invention.

[0052] See Figure 1 , Figure 1 This is a schematic diagram of the process for detecting disease-related amplifications with short tandem repeats provided in an embodiment of the present invention. Specifically, the method for detecting disease-related amplifications with short tandem repeats provided in an embodiment of the present invention includes the following steps:

[0053] 1) Obtaining the list of STR-related diseases and the coverage of WES products within their target regions:

[0054] First, we collect gene variation information of rare genetic diseases related to STR and WES sequencing data of actual production data samples (i.e. WES bam files after comparison of production data samples). We analyze the two types of data to obtain target region data of abnormal amplification of STR-related diseases (STR variant and flanking sequence related regions bam files) and obtain data such as the coverage of STR variant and flanking sequence related regions.

[0055] 2) Setting up standard analysis software workflows and building local libraries:

[0056] For example, the software ExpansionHunter and exSTRa were used to analyze and process WES sequencing data from different samples, and local negative sample (i.e., normal people) STR sample set and local specific gene STR abnormality reference set were constructed respectively.

[0057] 3) Evaluation of the difference between software prediction and experimental amplification times

[0058] 4) Determination of pathogenic STR variant filtering indicators and assessment criteria

[0059] The above analysis results were evaluated based on STR disease-related abnormality filtering indicators and assessment criteria, and were included in the Y-label of key concern or the L-label of subsequent review.

[0060] 6) Sensitivity and specificity assessment of the criteria for identifying disease-related abnormal amplification STRs.

[0061] This invention first obtains the coverage of target regions for STR-related disease genes, see [link to relevant documentation]. Figure 2 , Figure 2 This is a flowchart illustrating the process of obtaining the coverage of STR-related disease gene target regions according to an embodiment of the present invention.

[0062] This invention first identifies publicly available STR-related disease genes in existing literature through a search. For example, using keywords such as STR, short tandem repeat, or genetic disorder, 38 STR-related disease genes were retrieved from databases such as OMIM and PubMed. The genomic regions and ranges of disease-related amplification were then determined, including the minimum and maximum amplification values ​​(min_abnorm and Max_abnorm). For STR intervals with a target region length of less than 10 bp, the upstream and downstream regions were extended to 25 bp to identify the target region data for STR-related disease amplification, generating an STR-gene bed file.

[0063] This application obtains WES sequencing data (saved as BAM files) from production data, i.e., actual samples, for example, 200 samples. It statistically analyzes the coverage of all bases within the target region of abnormal amplification in STR-related diseases, the percentage of bases achieving sequencing coverage of 1X, 10X, 20X, etc., and the proportion of actual samples meeting the predetermined coverage conditions, such as greater than 95%. Using the coverage from actual WES probe products under real-world detection conditions, rather than the overlap values ​​from the product BAM files, better determines whether there are sufficient reads covering the target region to assess abnormal amplification. The results are shown in Table 1, which shows the percentage of samples meeting different coverage conditions within the target region of the WES sequencing data for STR-related disease genes.

[0064] Table 1. Percentage of samples meeting different coverage levels within the target region of WES sequencing data for STR-related disease genes.

[0065]

[0066] Among them: existing area coverage a, 1: complete; 2: coverage fluctuates; 3: poor coverage.

[0067] This invention uses ExpansionHunter and exSTRa as standard analysis software. Based on the 38 disease-related genes obtained above, relevant locus files were prepared according to the software requirements. The WES sequencing data of the samples were analyzed using the software's default parameters, for example:

[0068] ExpansionHunter:

[0069] To analyze each positive sample using the software's default parameters, the command line is as follows:

[0070] ExpansionHunter --reads smp.bam --reference reference.fa --variant-catalog / path / hg38 / variant_catalog.json --output-prefix smp

[0071] Parameter description:

[0072] --reads the BAM files to be read

[0073] --reference reference genome FASTA file

[0074] --variant-catalog STR locus files with known variant information

[0075] --output-prefix Output file prefix

[0076] Each test sample can obtain the predicted number of amplifications for candidate sites.

[0077] exSTRa:

[0078] Based on the STR-gene bed, for each STR region, a new STR-gene flanking500 bed file was obtained by extending 500bp to both sides. The original file was then used to extract the region bam from the STR-gene flanking500 bed to obtain the control target bam file. Perl scripts and modules from https: / / github.com / bahlolab / Bio-STR-exSTRa were used to read the target bam files of both control and test samples and generate STR counts. The RexSTRa package was used to calculate candidate STR scores, including P-value, t-value, and differential visualization.

[0079] The ExpansionHunter results of approximately 317 normal individuals (i.e., negative samples) were compiled to count the number of STR amplifications in local normal individuals, thus constructing a local normal individual database. The amplification counts of negative samples were extracted, sorted, and the minimum (low) and second-largest (up) values ​​were selected. These two values ​​were used as the range of amplification counts for normal individuals (Norm_min_CG; Norm_max_CG). Because software predictions may produce some false positives, the second-largest value was chosen to avoid false negatives caused by some false positive data increasing the upper limit of the normal amplification range.

[0080] Table 2. Software-predicted amplification counts of different STR genes in local healthy individuals.

[0081]

[0082] Sample data indicates the total number of samples in the test that can indicate the number of times the gene has been amplified. For regions with poor or no coverage, ExpansionHunter cannot predict the number of amplifications.

[0083] This application, after establishing a local library, evaluates the difference between the software prediction results and the actual number of amplifications in experiments, thereby correcting the amplification count data. Commercially available dynamic mutation detection products often use PCR + capillary electrophoresis for detection. Commonly detected genes or diseases include dentate nucleus-globus pallidus Lewy body atrophy (DRPLA), Friedrich syndrome (FRDA), Kennedy syndrome (SBMA), Fragile X syndrome (FX), myotonic dystrophy (DM), Huntington's disease (HD), spinocerebellar ataxia types 8 (1-3, 6-8, 12, 17), and spinocerebellar ataxia types 10 (1-3, 6-8, 12, 17, FRDA, DRPLA). This study summarizes experimental amplification data related to dynamic mutation products submitted clinically, and obtains WES data for relevant samples to perform ExpansionHunter amplification count prediction. A total of 80 samples were analyzed, involving 13 genes, all of which have high clinical diagnostic rates. The differences between actual observed values ​​and software predicted values ​​within the normal amplification range of STR disease-related genes are compared. See the results below. Figure 3 , Figure 3 These are actual observed values ​​and software predicted values ​​within the normal amplification range of STR disease-related genes. Figure 3 In the diagram, the X-axis represents genes related to some dynamic mutations, and the Y-axis represents the difference between observed and predicted values. Blue represents the minimum difference between two alleles, and yellow represents the maximum difference between two alleles. The results showed that observed values ​​were generally smaller than predicted values, while a few genes exhibited larger fluctuations in the difference between observed and predicted values, with the largest difference reaching 20 replicates. For example, SCA17 had an average of 20 fewer replicates, and SCA3 had an average of 5 fewer replicates. SCA2 showed a more average difference, with only 2 replicates. The five best-performing genes were ATN1, ATXN1, PPP2P2B, ATXN2, HTT, and TBP, followed by ATXN3, CACNA1A, ATXN7, ATXN8, and DMPK. The average difference between the actual observation and the software prediction within the normal amplification range for each STR was defined as Diff. The Min_abnorm corresponding to genes with Diff was revised as Min_abnorm_revised = Min_abnorm - Diff.

[0084] In some embodiments of the present invention, candidate pathogenic STR variant filtering indicators and evaluation criteria are used to classify the data obtained according to the method described above. See [link to relevant documentation]. Figure 4 , Figure 4 The flowchart illustrates the filtering indicators and evaluation criteria for candidate pathogenic STR variants. The specific process is as follows:

[0085] The Y-standard filtering criteria are as follows: These records require close attention based on the patient's clinical symptoms.

[0086] 1) Obtain relevant data according to the method described above, select genes whose target region >20X coverage accounts for more than 90% of the sample, and use the conditional genes as the starting genes for group L.

[0087] 2) Select disease-related minimum amplification values ​​within the NGS sequencing read length range (i.e., Min_abnorm* amplified base units < 150 bp, and use the conditional genes as Y standard inclusion genes).

[0088] 3) Select data whose maximum gene amplification count for the corresponding sample is greater than Norm_max_CG;

[0089] 4) Select data whose maximum gene amplification count for the test sample is greater than Min_abnorm;

[0090] 5) The exSTRa analysis was used to determine the significant difference in amplification between the tested sample and the control library (P value < 0.05).

[0091] The L-standard filtration conditions are as follows:

[0092] Under the condition that Y criterion 1) is met, STR amplification data that meet any of the criteria 3), 4), or 5) are selected as L criteria for inclusion in subsequent large-scale retrospective analysis.

[0093] This application processed 38 STR gene data using the method described above, and the results are as follows:

[0094] The standard inclusion genes for Y are: ATXN1, ATXN2, ATXN3, CACNA1A, ATXN7, PPP2R2B, TBP, DMPK, HTT, and ATN1.

[0095] The L-standard inclusion genes are: PPP2R2B, TBP, ATXN1, ATXN2, NOP56, ATXN3, CACNA1A, ATXN7, ATXN8, JPH3, HTT, DMPK, AR, ATN1, LRP12, TCF4, GLS, NOTCH2NLC, and NUTM2BA_S1.

[0096] This application further evaluates the sensitivity and specificity of the disease-related amplified STR identification criteria, as follows: The evaluation was conducted only on the Y-criteria included genes (10 genes). Eighteen samples of disease-related amplified genes detected by dynamic mutation products were selected for WES sequencing and analyzed according to the above procedure. A total of 180 WES records with positive results for disease-related amplified amplification were obtained (see Table 3). Among them, 17 met the Y-criteria, and 48+17 met the L-criteria.

[0097] Under the Y standard:

[0098] Sensitivity (TPR): True positive rate, describes the proportion of all positive cases identified out of all positive cases;

[0099] The calculation formula is: TPR = TP / (TP + FN). TP: true positive, FN: false negative.

[0100] TPR = 17 / 17 + 1 = 94.4%;

[0101] Specificity (TNR): true negative rate, describes the proportion of identified negative examples out of all negative examples;

[0102] The calculation formula is: TNR = TN / (FP + TN). TN: True Negative, FP: False Positive.

[0103] TNR: (180-18) / (0+162)=100%

[0104] In the included positive samples, under the Y standard, the sensitivity was 94.4% and the specificity was 100%. This ensures that records classified as Y by this method provide sufficient evidence of a true amplification. ATXN3 (SCA3) appearing in false negative samples was defined as L standard because its exSTRa pvalue did not meet the requirements. Based on the experimental and predictive difference analysis described above, SCA3 showed slightly larger fluctuations, leading to bias in the statistical significance of the difference. The ATXN3 gene under the L standard requires special attention.

[0105] Table 3. WES data and STR detection results of experimental samples with positive dynamic mutations.

[0106]

[0107] Based on this, the present invention provides a method for processing whole exome sequencing data, comprising the following steps:

[0108] Step S1: Obtain first reference data, which includes base coverage data of the target region of STR-related disease genes in the reference sample and sample proportion data under the predetermined base coverage.

[0109] Obtain second reference data, which includes data on the number of amplifications of the negative reference sample;

[0110] Step S2: Obtain test sample data, which includes base coverage data of the target region of STR-related disease genes in the test sample, sample proportion data under the predetermined base coverage, and test sample amplification number data.

[0111] The base coverage data and sample proportion data under the predetermined base coverage of the STR-related disease gene target region of the test sample are compared with the base coverage data and sample proportion data under the predetermined base coverage of the reference sample to obtain the first comparison result data.

[0112] If the first comparison result data is inconsistent, the amplification count data of the detected sample is compared with the amplification count data of the negative reference sample to obtain the second comparison result data.

[0113] The present invention first obtains first reference data and second reference data, the specific process of which is as follows:

[0114] Acquire STR-related disease gene data and identify target regions for abnormal amplification in STR-related diseases;

[0115] Obtain WES sequencing data of a reference sample, compare the WES sequencing data of the reference sample with the target region data of abnormal amplification of STR-related diseases, and obtain the base coverage data of the target region of STR-related disease genes in the reference sample and the sample proportion data under the predetermined base coverage.

[0116] The WES sequencing data of negative samples in the reference sample were analyzed using the ExpansionHunter software to obtain the amplification count data of the negative reference samples.

[0117] This application first acquires STR-related disease data and identifies target regions for abnormal amplification in STR-related diseases. In one specific implementation, this application acquires STR-related disease data by searching existing literature, such as retrieving STR-related disease genes from databases like OMIM and PubMed, and identifies the gene regions and ranges where abnormal amplification occurs in STR-related diseases, i.e., the target regions. Specifically, the target region data for abnormal amplification in STR-related diseases includes the minimum gene fragment length (Min_abnorm) and the maximum gene fragment length (Max_abnorm). To improve the accuracy of subsequent data processing, for target regions with gene fragment lengths less than 10 bp, the upstream and downstream sequences are extended, for example, to 25 bp, to obtain the target region data for abnormal amplification in STR-related diseases.

[0118] The processing method provided in this application includes the step of obtaining WES sequencing data of a reference sample. In some embodiments, the WES sequencing data of the reference sample can be WES sequencing results related to STR-related diseases obtained by a testing institution. The method involves statistically analyzing the coverage of all bases within the target region of abnormal amplification in STR-related diseases, determining the target region of abnormal amplification in STR-related diseases in the WES sequencing data of the reference sample, and determining the percentage of bases in the target region of abnormal amplification in STR-related diseases with coverage reaching 1X, 10X, 20X, 30X, etc., based on the coverage of all bases within the target region of abnormal amplification in STR-related diseases. Simultaneously, the method also statistically analyzes the proportion of reference samples whose base percentage meets a predetermined coverage condition, such as greater than 95%.

[0119] The processing method provided by this invention includes the step of obtaining second reference data, which includes amplification count data of negative reference samples. In some possible implementations, this application uses ExpansionHunter software to analyze the WES sequencing data of negative samples in the reference samples to obtain amplification count data of negative reference samples. Specifically, the amplification count data of negative reference samples is an amplification count range, with the minimum amplification count (Norm_min_CG) and the second maximum amplification count (Norm_max_CG) serving as the upper and lower limits of the amplification count range, respectively.

[0120] Furthermore, the second reference data also includes positive reference sample amplification count data, the processing method of which is similar to the method for obtaining negative reference sample amplification count data, and will not be described in detail here.

[0121] In some possible implementations, the second reference data also includes STR calculation score data of positive samples in the reference sample. Specifically, this application uses exSTRa software to analyze the WES sequencing data of the reference sample to obtain STR calculation score data of positive samples in the reference sample, such as P value, t value, differential visualization, etc.

[0122] In some possible implementations, this application may further include a process for correcting the amplification count data of the negative reference sample, specifically as follows:

[0123] Step 1: Obtain the actual number of amplifications and WES sequencing data for positive samples;

[0124] Step 2: Use ExpansionHunter software to analyze the WES sequencing data of positive samples to obtain the predicted number of amplifications of positive samples.

[0125] Step 3: Correct the amplification count data of the negative reference sample based on the actual amplification count data and the predicted amplification count data of the positive sample.

[0126] Specifically, steps 1 and 2 are similar to the acquisition process described above, and will not be repeated here. After obtaining the actual amplification count and the predicted amplification count of the positive sample, the difference between the two is calculated, and the amplification count data of the negative reference sample is corrected based on this difference.

[0127] The processing method provided by the present invention further includes step S2:

[0128] Acquire test sample data, which includes base coverage data of the target region of STR-related disease genes in the test sample, sample proportion data under the predetermined base coverage, and test sample amplification number data;

[0129] The base coverage data and sample proportion data under the predetermined base coverage of the STR-related disease gene target region of the test sample are compared with the base coverage data and sample proportion data under the predetermined base coverage of the reference sample to obtain the first comparison result data.

[0130] If the first comparison result data is inconsistent, the amplification count data of the detected sample is compared with the amplification count data of the negative reference sample to obtain the second comparison result data.

[0131] The process of obtaining the detection sample data in step S2 is similar to the process of obtaining the first reference data and the second reference data in step S1, specifically as follows:

[0132] Obtain WES sequencing data of the test sample, compare the WES sequencing data of the test sample with the target region data of abnormal amplification of STR-related diseases, and obtain the base coverage data of the target region of STR-related disease genes in the test sample and the sample proportion data under the predetermined base coverage.

[0133] The WES sequencing data of the test samples were analyzed using the ExpansionHunter software to obtain the amplification count data of the test samples.

[0134] After obtaining the test sample data, the base coverage data and sample percentage data under the predetermined base coverage of the STR-related disease gene target region of the test sample are compared with the base coverage data and sample percentage data under the predetermined base coverage of the STR-related disease gene target region of the reference sample to obtain the first comparison data. For example, if the base coverage of the predetermined target region reaches 20X or more and the sample percentage data of the bases is 90% or more, if the corresponding data of the test sample meets the above predetermined conditions, the data of the test sample is included in the L standard inclusion gene.

[0135] If the data of the test sample does not meet the above predetermined conditions, a further judgment is made. In a specific implementation, the test sample data also includes the genomic region where the abnormal amplification of the test sample occurs and the range of disease-related abnormal amplification, including a minimum value Min_abnorm and a maximum value Max_abnorm. If the data of the test sample does not meet the above predetermined conditions, the minimum value Min_abnorm of the genomic region where the abnormal amplification of the test sample occurs is compared to obtain second alignment result data. Specifically, Min_abnorm is compared with the NGS sequencing read length. If Min_abnorm is within the NGS sequencing read length range, the test sample data is included in the Y standard inclusion gene; if Min_abnorm is not within the NGS sequencing read length range, further alignment analysis is performed on the test sample data.

[0136] Specifically, further comparative analysis includes: comparing the maximum number of amplifications in the test sample data with the number of amplifications in the negative sample data, including at least one of the following comparative analyses:

[0137] (1) Compare the maximum number of amplifications of the test sample data with the second largest number of amplifications of the negative sample (Norm-max-CG). If the maximum number of amplifications is not greater than Norm-max-CG, the data is listed as L standard inclusion gene. If the maximum number of amplifications is greater than Norm-max-CG, it is listed as Y standard and used as key data for clinical interpretation.

[0138] (2) Compare the maximum number of amplifications of the test sample data with the minimum number of amplifications of the corrected negative sample (Min_abnorm correction). If the maximum number of amplifications is not greater than the minimum number of amplifications of the corrected negative sample, the data is listed as L standard inclusion gene. If the maximum number of amplifications is greater than the minimum number of amplifications of the corrected negative sample, it is listed as Y standard and used as key data for clinical interpretation.

[0139] (3) Obtain the P value calculated using exSTRa from the test sample data. If the P value is ≥0.05, the data is listed as L standard inclusion genes. If the P value is <0.05, the data is listed as Y standard and is used as key data for clinical interpretation.

[0140] This invention provides a whole-exome sequencing data processing system, including a first reference data acquisition unit. The first reference data acquisition unit is used to acquire first reference data, which includes base coverage data of the target region of STR-related disease genes of the reference sample and sample proportion data under a predetermined base coverage.

[0141] The second reference data acquisition unit is used to acquire second reference data, which includes negative reference sample amplification count data.

[0142] The test sample data acquisition unit is used to acquire test sample data, which includes base coverage data of the target region of STR-related disease genes of the test sample, sample proportion data under the predetermined base coverage, and test sample amplification number data.

[0143] The first comparison unit is used to compare the base coverage data and the sample proportion data under the predetermined base coverage in the target region of the STR-related disease gene of the test sample with the base coverage data and the sample proportion data under the predetermined base coverage in the target region of the STR-related disease gene of the reference sample to obtain the first comparison result data.

[0144] The second comparison unit is used to compare the amplification count data of the detection sample with the amplification count data of the negative reference sample to obtain the second comparison result data.

[0145] In one specific implementation, the first reference data acquisition unit includes an STR-related disease gene data acquisition unit, which is used to acquire STR-related disease gene data and determine the target region data of abnormal amplification of STR-related diseases.

[0146] A reference sample WES sequencing data acquisition unit is used to acquire WES sequencing data of a reference sample.

[0147] The third comparison unit is used to compare the WES sequencing data of the reference sample with the target region data of the abnormal amplification of the STR-related disease, and to obtain the base coverage data of the STR-related disease gene target region of the reference sample and the sample proportion data under the predetermined base coverage.

[0148] In one specific implementation, the test sample data acquisition unit includes a test sample WES sequencing data acquisition unit, which is used to acquire the WES sequencing data of the test sample.

[0149] The fourth comparison unit is used to compare the WES sequencing data of the test sample with the target region data of the abnormal amplification of the STR-related disease to obtain the base coverage data of the STR-related disease gene target region of the test sample and the sample proportion data under the predetermined base coverage.

[0150] The sample amplification count data processing unit is used to analyze the WES sequencing data of the test samples using Min_abnorm software to obtain the sample amplification count data.

[0151] The data processing system provided by this invention is used to implement the above-mentioned data processing method. The function of each unit is to implement each step of the data processing method, which will not be described in detail here.

[0152] The present invention also provides a system for detecting disease-related amplification of short tandem repeats, including a first reference data acquisition unit, the first reference data unit being used to acquire first reference data, the first reference data including base coverage data of the target region of STR-related disease genes of the reference sample and sample proportion data under a predetermined base coverage.

[0153] The second reference data acquisition unit is used to acquire second reference data, which includes negative reference sample amplification count data.

[0154] The test sample data acquisition unit is used to acquire test sample data, which includes base coverage data of the target region of STR-related disease genes of the test sample, sample proportion data under the predetermined base coverage, and test sample amplification number data.

[0155] The first comparison unit is used to compare the base coverage data and the sample proportion data under the predetermined base coverage in the target region of the STR-related disease gene of the test sample with the base coverage data and the sample proportion data under the predetermined base coverage in the target region of the STR-related disease gene of the reference sample to obtain the first comparison result data.

[0156] The second comparison unit is used to compare the amplification count data of the detection sample with the amplification count data of the negative reference sample to obtain the second comparison result data;

[0157] A prediction system is used to obtain prediction results of short tandem repeat disease-related amplifications based on first and second alignment result data.

[0158] Specifically, the prediction system analyzes the first and second alignment results to obtain prediction results for abnormal amplifications related to short tandem repeat diseases. After obtaining the test sample data, the system compares the base coverage data and sample percentage data under a predetermined base coverage of the target region of the STR-related disease gene in the test sample with the base coverage data and sample percentage data under a predetermined base coverage of the target region of the STR-related disease gene in the reference sample to obtain first comparison data. For example, if the base coverage of the predetermined target region reaches 20X or more and the sample percentage data of the bases is 90% or more, the prediction system determines whether to include the test sample in the L inclusion standard gene or proceed to the next step based on the first comparison data. If the corresponding data of the test sample meets the above predetermined conditions, the data of the test sample is included in the L standard inclusion gene.

[0159] If the data of the test sample does not meet the above predetermined conditions, a further judgment is made. In a specific implementation, the test sample data also includes the genomic region where the abnormal amplification of the test sample occurs and the range of disease-related abnormal amplification, including a minimum value Min_abnorm and a maximum value Max_abnorm. If the data of the test sample does not meet the above predetermined conditions, the minimum value Min_abnorm of the genomic region where the abnormal amplification of the test sample occurs is compared to obtain second alignment result data. Specifically, Min_abnorm is compared with the NGS sequencing read length. If Min_abnorm is within the NGS sequencing read length range, the prediction system includes the test sample data in the Y standard inclusion gene group; if Min_abnorm is not within the NGS sequencing read length range, further alignment analysis is performed on the test sample data.

[0160] Specifically, further comparative analysis includes: comparing the maximum number of amplifications in the test sample data with the number of amplifications in the negative sample data, including at least one of the following comparative analyses:

[0161] (1) Compare the maximum number of amplifications of the test sample data with the second largest number of amplifications of the negative sample (Norm_max_CG). If the maximum number of amplifications is not greater than Norm_max_CG, the data is listed as L standard inclusion gene. If the maximum number of amplifications is greater than Norm_max_CG, it is listed as Y standard and used as key data for clinical interpretation.

[0162] (2) Compare the maximum number of amplifications of the test sample data with the minimum number of amplifications of the corrected negative sample (Min_abnorm correction). If the maximum number of amplifications is not greater than the minimum number of amplifications of the corrected negative sample, the data is listed as L standard inclusion gene. If the maximum number of amplifications is greater than the minimum number of amplifications of the corrected negative sample, it is listed as Y standard and used as key data for clinical interpretation.

[0163] (3) Obtain the P value calculated using exSTRa from the test sample data. If the P value is ≥0.05, the data is listed as L standard inclusion genes. If the P value is <0.05, the data is listed as Y standard and is used as key data for clinical interpretation.

[0164] This invention defines detectable STR-related genes in samples based on the actual coverage of the samples in WES sequencing data, which is more accurate than using the overlap of the bed region / bed+flanking region of the WES probe for evaluation. The whole-exome sequencing data processing method provided by this invention is less affected by different algorithms, sequencing platforms, probes, and alignment software, resulting in more accurate data results.

[0165] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for processing whole-exome sequencing data, characterized in that, Includes the following steps: Step S1: Obtain first reference data, which includes base coverage data of the target region of STR-related disease genes in the reference sample and sample proportion data under the predetermined base coverage. Obtain second reference data, which includes data on the number of amplifications of the negative reference sample; Step S2: Obtain test sample data, which includes base coverage data of the target region of STR-related disease genes in the test sample, sample proportion data under the predetermined base coverage, and test sample amplification number data. The base coverage data and sample proportion data under the predetermined base coverage of the STR-related disease gene target region of the test sample are compared with the base coverage data and sample proportion data under the predetermined base coverage of the reference sample to obtain the first comparison result data. If the first comparison result data is inconsistent, the amplification count data of the test sample is compared with the amplification count data of the negative reference sample to obtain the second comparison result data; Step S1 specifically includes: Acquire gene data related to STR-related diseases and identify target regions for abnormal amplification in STR-related diseases; Obtain WES sequencing data of a reference sample, compare the WES sequencing data of the reference sample with the target region data of abnormal amplification of STR-related diseases, and obtain the base coverage data of the target region of STR-related disease genes in the reference sample and the sample proportion data under the predetermined base coverage. The WES sequencing data of negative samples in the reference sample were analyzed using the ExpansionHunter software to obtain the amplification count data of the negative reference samples. Step S1 further includes: correcting the amplification count data of the negative reference sample to obtain corrected amplification count data of the negative reference sample; Step S2 specifically involves: Obtain WES sequencing data of the test sample, compare the WES sequencing data of the test sample with the target region data of abnormal amplification of STR-related diseases, and obtain the base coverage data of the target region of STR-related disease genes in the test sample and the sample proportion data under the predetermined base coverage. The WES sequencing data of the test samples were analyzed using the ExpansionHunter software to obtain the amplification count data of the test samples.

2. The processing method according to claim 1, characterized in that, The correction to the amplification count data of the negative reference sample is specifically as follows: Obtain data on the actual number of amplifications and WES sequencing data for positive samples; The WES sequencing data of positive samples were analyzed using ExpansionHunter software to obtain the predicted number of amplifications for the positive samples. The amplification count data of the negative reference sample is corrected based on the actual amplification count data and the predicted amplification count data of the positive sample.

3. The processing method according to claim 1, characterized in that, Also includes: The WES sequencing data of the reference sample were analyzed using exSTRa software to obtain the STR scores of the positive samples in the reference sample. The WES sequencing data of the test samples were analyzed using exSTRa software to obtain the STR score of the test samples.

4. A whole-exome sequencing data processing system for implementing the method as described in any one of claims 1 to 3, characterized in that, It includes a first reference data acquisition unit, which is used to acquire first reference data, which includes base coverage data of the target region of STR-related disease genes of the reference sample and sample proportion data under a predetermined base coverage. The second reference data acquisition unit is used to acquire second reference data, which includes negative reference sample amplification count data. The test sample data acquisition unit is used to acquire test sample data, which includes base coverage data of the target region of STR-related disease genes of the test sample, sample proportion data under the predetermined base coverage, and test sample amplification number data. The first comparison unit is used to compare the base coverage data and the sample proportion data under the predetermined base coverage in the target region of the STR-related disease gene of the test sample with the base coverage data and the sample proportion data under the predetermined base coverage in the target region of the STR-related disease gene of the reference sample to obtain the first comparison result data. The second comparison unit is used to compare the amplification count data of the detection sample with the amplification count data of the negative reference sample to obtain the second comparison result data.

5. The processing system according to claim 4, characterized in that, The first reference data acquisition unit includes an STR-related disease gene data acquisition unit, which is used to acquire STR-related disease gene data and determine the target region data of abnormal amplification of STR-related diseases. A reference sample WES sequencing data acquisition unit is used to acquire WES sequencing data of a reference sample. The third comparison unit is used to compare the WES sequencing data of the reference sample with the target region data of the abnormal amplification of the STR-related disease, and to obtain the base coverage data of the STR-related disease gene target region of the reference sample and the sample proportion data under the predetermined base coverage.

6. The processing system according to claim 5, characterized in that, The sample data acquisition unit includes: A sample WES sequencing data acquisition unit is used to acquire the WES sequencing data of the sample. The fourth comparison unit is used to compare the WES sequencing data of the test sample with the target region data of the abnormal amplification of the STR-related disease to obtain the base coverage data of the STR-related disease gene target region of the test sample and the sample proportion data under the predetermined base coverage. The sample amplification count data processing unit is used to analyze the WES sequencing data of the test samples using ExpansionHunter software to obtain the sample amplification count data.

7. A system for detecting disease-related aberrant amplifications of short tandem repeats, used to implement the method as described in any one of claims 1 to 3, characterized in that, It includes a first reference data acquisition unit, which is used to acquire first reference data, which includes base coverage data of the target region of STR-related disease genes of the reference sample and sample proportion data under a predetermined base coverage. The second reference data acquisition unit is used to acquire second reference data, which includes negative reference sample amplification count data. The test sample data acquisition unit is used to acquire test sample data, which includes base coverage data of the target region of STR-related disease genes of the test sample, sample proportion data under the predetermined base coverage, and test sample amplification number data. The first comparison unit is used to compare the base coverage data and the sample proportion data under the predetermined base coverage in the target region of the STR-related disease gene of the test sample with the base coverage data and the sample proportion data under the predetermined base coverage in the target region of the STR-related disease gene of the reference sample to obtain the first comparison result data. The second comparison unit is used to compare the amplification count data of the detection sample with the amplification count data of the negative reference sample to obtain the second comparison result data; A prediction system is used to obtain prediction results of short tandem repeat disease-related amplifications based on first and second alignment result data.