Tumor gene mutation interpretation method based on NGS platform
By using preset lists and dynamic adaptation interpretation criteria in the tumor gene mutation interpretation method, the problem of frequent false positives in the existing technology is solved, and high-accuracy and high-efficiency tumor gene mutation detection is achieved.
Patent Information
- Application Number
- CN202511099556.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-07
AI Technical Summary
Existing methods for interpreting tumor gene mutations cannot effectively distinguish low-abundance mutations, background noise, or misjudged systematic false positives, resulting in frequent false positive results, especially when faced with complex gene mutations, where the accuracy of the judgment is low.
By obtaining the mutation detection results of the sample to be analyzed, it is determined whether the candidate mutation exists in the preset list, and the interpretation result is generated according to the list type. It is further interpreted in combination with the hotspot mutation type, sample quality type and mutation data, and the interpretation standard is dynamically adapted.
It has improved the standardization level and clinical applicability of tumor gene mutation interpretation, reduced the false positive and missed detection rates, and improved the accuracy and efficiency of interpretation.
Smart Images

Figure CN120600111A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of tumor gene mutation detection, and in particular to a method for interpreting tumor gene mutations based on an NGS platform. Background Art
[0002] With the advancement of genomics technology, tumor gene mutation detection methods based on NGS platforms have been widely used in clinical cancer diagnosis and targeted treatment decision-making. By deeply sequencing gene mutations in tumor tissue samples, NGS technology can accurately identify cancer-related gene mutations and provide support for personalized treatment.
[0003] Existing methods for interpreting tumor gene mutations typically rely on specific mutation detection algorithms and threshold criteria to classify and analyze mutation data. These methods are unable to effectively distinguish low-abundance mutations, background noise, or systematic false positives due to misidentification. Furthermore, when faced with complex gene mutations, such as low-frequency mutations, mutations with low dispersion, or rare mutation sites, the accuracy of their interpretation is low, leading to frequent false-positive results. Although false-positive mutations can be filtered out using various mechanisms, the variety of false-positive mutation types and their mechanisms of occurrence are not yet clearly understood. Therefore, the current false-positive filtering effect is limited, and there is room for improvement. Summary of the Invention
[0004] This application provides a method for interpreting tumor gene mutations based on an NGS platform, which can improve the standardization level and clinical applicability of tumor gene mutation interpretation, while effectively reducing the false positive misjudgment rate and missed detection rate.
[0005] The above-mentioned invention objectives of this application are achieved through the following technical solutions:
[0006] A method for interpreting tumor gene mutations based on an NGS platform, comprising:
[0007] Obtaining mutation detection results for the sample to be analyzed, the mutation detection results including at least one candidate mutation and mutation data corresponding to the candidate mutation, the mutation data including mutation type, mutation abundance, number of supporting reads, site sequencing depth, and manual interpretation labels;
[0008] According to the mutation detection result, determining whether the candidate mutation exists in a preset list, wherein the preset list includes at least a blacklist and a graylist;
[0009] If the candidate mutation exists in the preset list, generating a judgment result corresponding to the candidate mutation according to the list type of the preset list;
[0010] If the candidate mutation does not exist in the preset list, determining the hotspot mutation type and sample quality type of the candidate mutation of the sample to be analyzed according to the mutation detection result;
[0011] Based on the hotspot mutation type, the sample quality type, and the mutation data, a judgment result corresponding to the candidate mutation is generated.
[0012] By adopting the above technical solution, by obtaining the mutation detection results of the samples to be analyzed, the basic data characteristics of each candidate mutation can be accurately grasped, thereby avoiding interpretation bias due to incomplete mutation information. By determining whether the candidate mutation exists in the preset list, mutations that are known to be unreliable or have interpretation differences can be eliminated in the early stage of the interpretation process, thereby improving the overall screening efficiency and accuracy. By generating interpretation results according to the list type, preliminary judgments with high credibility or those requiring review can be made quickly, thereby shortening the necessary time for manual review intervention and improving the overall interpretation speed. After the candidate mutation does not exist in the preset list, further interpretation is performed based on the hotspot mutation type, sample quality type and mutation data, which can achieve dynamic adaptation and personalized segmentation of the interpretation standard, thereby ensuring interpretation accuracy and adaptability under different sample conditions, thereby improving the standardization level and clinical applicability of tumor gene mutation interpretation.
[0013] In a preferred example, the present application may be further configured as follows: generating the interpretation result corresponding to the candidate mutation according to the list type of the preset list, specifically including:
[0014] If the candidate mutation exists in the blacklist, generating and marking the interpretation result of the candidate mutation as unqualified;
[0015] If the candidate mutation exists in the gray list, a reading result of the candidate mutation is generated and marked as requiring manual review.
[0016] By adopting the above technical solution, by generating the interpretation results of candidate mutations according to the list type of the preset list, the interpretation conclusion can be accurately output directly according to the type characteristics of the blacklist and gray list, thereby avoiding the ordinary interpretation logic from processing abnormal or controversial mutations and improving the overall interpretation accuracy.
[0017] In a preferred example, the present application may be further configured as follows: determining the hotspot mutation type and sample quality type of the candidate mutation of the sample to be analyzed based on the mutation detection result, specifically including:
[0018] Determine whether the candidate mutation of the sample to be analyzed exists in a preset whitelist, where the whitelist is constructed based on known hotspot mutations in public clinical databases;
[0019] If the candidate mutation exists in the whitelist, determining the hotspot mutation type of the candidate mutation as a hotspot mutation;
[0020] If the candidate mutation does not exist in the whitelist, determining the hotspot mutation type as a non-hotspot mutation;
[0021] Obtaining mutation spectrum characteristics and background noise levels of the sample to be analyzed from the mutation detection results, and determining a sample degradation type of the sample to be analyzed based on the mutation spectrum characteristics;
[0022] determining a sample noise type of the sample to be analyzed according to the background noise level;
[0023] A sample quality type of the sample to be analyzed is determined based on the sample degradation type and the sample noise type.
[0024] By adopting the above technical solution, by determining the hotspot mutation type and sample quality type of the candidate mutation based on the mutation detection results, it is possible to set differentiated interpretation standards for samples of different qualities based on the clinical value of the mutation and the differences in sequencing quality of the samples, thereby improving the targetedness of mutation screening and the scientific rationality of the interpretation standards, and avoiding misjudgment caused by one-size-fits-all treatment of samples with different characteristics under the same standard.
[0025] In a preferred example, the present application may be further configured as follows: determining the sample degradation type of the sample to be analyzed based on the mutation spectrum characteristics, specifically including:
[0026] Counting the total number of mutations with SNV type in the mutation data of the sample to be analyzed;
[0027] From the total number of mutations, the number of mutations whose mutation form is C>T and whose mutation abundance in the mutation data is less than a preset abundance threshold is obtained;
[0028] Based on the total number of mutations and the number of mutations, the degradation ratio of the sample is calculated;
[0029] If the sample degradation ratio is greater than or equal to the degradation judgment threshold, the sample degradation type of the sample to be analyzed is determined to be a degraded sample;
[0030] If the sample degradation ratio is less than the degradation judgment threshold, the sample degradation type of the sample to be analyzed is determined to be a non-degraded sample.
[0031] By adopting the above technical solution, by determining the sample degradation type based on the mutation spectrum characteristics, it is possible to identify potentially degraded samples based on the proportion characteristics of the overall low-frequency C>T mutations in the sample, thereby providing early warning of the risk of sequencing data distortion and improving the sample quality recognition rate before interpretation. By setting the degradation judgment threshold and calculating the degradation ratio, the degree of sample degradation can be quantitatively analyzed, thereby objectively supporting the selection of different subsequent interpretation standards and avoiding reliance on subjective experience.
[0032] In a preferred example, the present application may be further configured as follows: determining the sample noise type of the sample to be analyzed according to the background noise level specifically includes:
[0033] Based on the background noise level, obtaining a hotspot mutation noise level of at least one candidate mutation whose hotspot mutation type is a hotspot mutation in the sample to be analyzed;
[0034] Calculating the single-sample background noise level statistics of all hotspot mutations in each sample to be analyzed based on the hotspot mutation noise level;
[0035] If the single sample background noise level statistic is greater than or equal to a preset threshold, determining that the sample noise type of the sample to be analyzed is a high noise sample;
[0036] If the single-sample background noise level statistic is less than the preset threshold, it is determined that the sample noise type of the sample to be analyzed is a non-high noise sample.
[0037] By adopting the above technical solution and determining the sample noise type based on the background noise level, the background noise level changes of hot spot mutations in the sample can be used to objectively evaluate the degree of data interference, thereby timely identifying samples with low data quality. By determining the noise type based on the background noise level statistics and the preset threshold, the influence of noise anomalies can be effectively eliminated, ensuring the accuracy and consistency of mutation interpretation.
[0038] In a preferred example, the present application may be further configured as follows: determining the sample quality type of the sample to be analyzed based on the sample degradation type and the sample noise type, specifically including:
[0039] If the sample degradation type of the sample to be analyzed is a degraded sample and / or the sample noise type of the sample to be analyzed is a high-noise sample, determining that the sample quality type of the sample to be analyzed is a low-quality sample;
[0040] If the sample degradation type of the sample to be analyzed is a non-degraded sample and / or the sample noise type of the sample to be analyzed is a non-high noise sample, it is determined that the sample quality type of the sample to be analyzed is a normal sample.
[0041] By adopting the above technical solution, by determining the sample quality type based on the sample degradation type and the sample noise type, it is possible to scientifically divide the quality of samples based on different data quality dimensions, and objectively distinguish high-noise samples from non-high-noise samples, thereby effectively identifying potential abnormal samples that are severely affected by background interference, improving the adaptability and robustness of the method under complex sample conditions, and avoiding misjudgment or missed judgment caused by poor-quality samples.
[0042] In a preferred example, the present application may be further configured as follows: generating a judgment result corresponding to the candidate mutation based on the hotspot mutation type, the sample quality type, and the mutation data, specifically including:
[0043] Based on the hotspot mutation type and the sample quality type of the sample to be analyzed, obtaining a corresponding preset dynamic threshold;
[0044] According to whether the mutation data of the candidate mutation meets the preset dynamic threshold standard, if it does, the interpretation result of the candidate mutation is generated and marked as qualified; if it does not meet the standard, the interpretation result of the candidate mutation is generated and marked as requiring manual review.
[0045] By adopting the above technical solution, by generating candidate mutation interpretation results based on hotspot mutation type, sample quality type and mutation data, the corresponding preset interpretation threshold standards can be dynamically called according to the actual sample and mutation characteristics, thereby achieving priority protection of high-risk hotspot mutations and isolation of low-quality sample risks, reducing the risks of false positives and false negatives, and improving the credibility and clinical applicability of mutation interpretation results.
[0046] In a preferred example, the present application may be further configured as follows: before the step of obtaining the corresponding preset dynamic threshold based on the hotspot mutation type and the sample quality type of the sample to be analyzed, the tumor gene mutation interpretation method further includes:
[0047] Obtain a historical sample data set, and divide the historical sample data set into a test set and a validation set according to a preset ratio;
[0048] Grouping the test set and the validation set according to the hotspot mutation type and the sample quality type;
[0049] For each group of the test set, all mutation sites with the manual interpretation label as unqualified are selected, and an initial threshold boundary is set according to the mutation abundance, the number of mutation supporting reads and the sequencing depth of the site of all mutation sites;
[0050] Based on the initial threshold boundary, a multidimensional grid space is constructed and discretely divided into threshold combination candidate areas, and a grid search and optimization algorithm is used to determine the threshold combination with the maximum weighted objective function value for the corresponding group in the test set based on the threshold combination candidate areas and using a weighted objective function;
[0051] Utilizing the validation set for validation, recording mutations in each sample in the validation set that fall within the threshold combination of the corresponding group as qualified mutations, and calculating the consistency index of the qualified mutations in the corresponding group in the validation set;
[0052] If the consistency index is greater than or equal to the preset consistency threshold, the preset dynamic threshold of the corresponding group is determined according to the threshold combination.
[0053] By adopting the above technical solution, by obtaining historical sample data sets and dividing them into test sets and validation sets, a reasonable training and verification mechanism can be established based on real historical judgment data, thereby ensuring that dynamic threshold optimization has sufficient data support and generalization capabilities. By conducting group training based on hot spot mutation types and sample quality types, the judgment standards of different categories of samples can be optimized in a targeted manner, thereby improving the judgment accuracy and stability. By constructing a multi-dimensional grid space and a weighted objective function to search for the optimal threshold combination, an optimal balance can be achieved between specificity and coverage, thereby maximizing the risk of missed judgments and misjudgments. By verifying the consistency index and screening the dynamic threshold in the validation set, it can be ensured that the final dynamic judgment standard is also highly stable and reliable on independent sample sets.
[0054] In a preferred example, the present application may be further configured as follows: before the step of determining whether the candidate mutation exists in the preset list, the tumor gene mutation interpretation method further includes:
[0055] Screening the mutation sites in the historical sample data set that are manually interpreted as unqualified, and calculating the screening frequency and mutation abundance standard deviation of the screened mutation sites;
[0056] According to the screening frequency, the mutation sites whose screening frequency is greater than or equal to a preset frequency threshold are classified as a high-frequency unqualified mutation site set;
[0057] In the set of high-frequency unqualified mutation sites, a first site whose mutation abundance standard deviation is less than a preset stability threshold is selected, and the selected first site is manually reviewed and confirmed with known clinical significance;
[0058] The first site that is still unqualified and has no known clinical significance after manual review is included in the blacklist.
[0059] By adopting the above technical solution, by screening the mutation sites that are manually interpreted as unqualified in historical sample data and calculating the screening frequency and abundance standard deviation, it is possible to accurately identify false positive mutations with poor stability and no clinical significance that have repeatedly appeared in history, thereby achieving scientific collection of blacklist candidate sites, and adding sites that meet the standards to the blacklist after manual review and confirmation, which can ensure the rigor and authority of the list, thereby quickly and accurately eliminating high-risk false positive mutations in the interpretation process, and improving the overall interpretation efficiency and accuracy.
[0060] In a preferred example, the present application may be further configured as follows: before the step of determining whether the candidate mutation exists in the preset list, the tumor gene mutation interpretation method further includes:
[0061] Screening out mutation sites that do not belong to the whitelist and the blacklist from the mutation sites, and calculating the abundance dispersion index of the screened mutation sites;
[0062] The mutation sites whose abundance dispersion index is greater than the preset dispersion threshold are included in the gray list.
[0063] By adopting the above technical solution, by screening mutation sites whose abundance dispersion index is greater than the preset dispersion threshold from the remaining mutation sites that do not belong to the whitelist and blacklist, difficult mutation sites with large fluctuations in interpretation parameters and large review differences can be identified, thereby constructing a gray list, which is convenient for giving priority to manual review in the subsequent interpretation process, avoiding misjudgment or missed judgment due to fluctuations in parameter boundaries, thereby improving the reliability of the overall interpretation process and the review efficiency of human-computer collaboration.
[0064] In summary, this application includes at least one of the following beneficial technical effects:
[0065] 1. By obtaining mutation detection results for samples to be analyzed, the basic data characteristics of each candidate mutation can be accurately grasped, thereby avoiding interpretation bias caused by incomplete mutation information. By determining whether the candidate mutation exists in the preset list, mutations known to be unreliable or with conflicting interpretations can be eliminated early in the interpretation process, improving the efficiency and accuracy of the overall screening. By generating interpretation results based on the list type, preliminary judgments of high confidence or those requiring review can be quickly made, thereby shortening the time required for manual review and intervention, and improving the overall interpretation speed. If the candidate mutation is not on the preset list, further interpretation can be carried out based on hotspot mutation type, sample quality type, and mutation data, enabling dynamic adaptation and personalized segmentation of interpretation standards, thereby ensuring interpretation accuracy and adaptability under different sample conditions, thereby improving the standardization and clinical applicability of tumor gene mutation interpretation;
[0066] 2. By screening mutation sites that have been manually interpreted as unqualified in historical sample data and calculating the screening frequency and abundance standard deviation, we can accurately identify historically recurring false-positive mutations with poor stability and no clinical significance, thereby achieving a scientific collection of blacklist candidate sites. After manual review and confirmation, sites that meet the standards are added to the blacklist, ensuring the rigor and authority of the list. This allows for the rapid and accurate elimination of high-risk false-positive mutations during the interpretation process, improving overall interpretation efficiency and accuracy.
[0067] 3. By screening mutation sites whose abundance dispersion index is greater than the preset dispersion threshold from the remaining mutation sites that are not on the whitelist and blacklist, difficult mutation sites with large fluctuations in interpretation parameters and large review differences can be identified, thereby constructing a gray list, which is convenient for prioritizing manual review in the subsequent interpretation process, avoiding misjudgments or missed judgments due to fluctuations in parameter boundaries, thereby improving the reliability of the overall interpretation process and the efficiency of human-machine collaborative review. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 This is a flowchart of a method for interpreting tumor gene mutations based on an NGS platform in one embodiment of the present application;
[0069] Figure 2 This is another implementation flowchart of a method for interpreting tumor gene mutations based on an NGS platform in one embodiment of the present application;
[0070] Figure 3 This is another implementation flowchart of a method for interpreting tumor gene mutations based on an NGS platform in one embodiment of the present application;
[0071] Figure 4 This is another implementation flowchart of a method for interpreting tumor gene mutations based on an NGS platform in one embodiment of the present application. DETAILED DESCRIPTION
[0072] The present application is further described in detail below with reference to the accompanying drawings.
[0073] In one embodiment, if Figure 1 As shown, the present application discloses a method for interpreting tumor gene mutations based on an NGS platform, which specifically includes the following steps:
[0074] S10: Obtain the mutation detection results of the sample to be analyzed. The mutation detection results include at least one candidate mutation and mutation data corresponding to the candidate mutation. The mutation data includes mutation type, mutation abundance, number of supporting reads, site sequencing depth, and manual interpretation labels.
[0075] Specifically, the mutation detection results of the samples to be analyzed are obtained through the NGS platform. For each detected candidate mutation, detailed information is recorded, including the mutation type such as SNP, InDel, mutation abundance AF, that is, the proportion of mutation reads to total reads, support reads AD, that is, the number of reads supporting the mutation, site sequencing depth DP, that is, the total number of reads covering the site, and manual review labels such as qualified Pass and unqualified Fail. Among them, the mutation type refers to whether the mutation is a single nucleotide variation SNV, insertion / deletion variation InDel, etc., the mutation abundance AF refers to the proportion of the mutation in all detected genomes, the support reads AD refers to the number of sequencing reads supporting the mutation, and the site sequencing depth DP refers to the sequencing depth of the sequencing coverage of the site. The manual interpretation label is manually determined to mark whether the mutation meets the clinical standards and ensure the reliability of the data quality.
[0076] S20: According to the mutation detection result, determine whether the candidate mutation exists in a preset list, where the preset list includes at least a blacklist and a graylist.
[0077] Specifically, the preset lists include blacklists and graylists. The blacklist contains known high-frequency false-positive mutation sites. These sites are easily mistakenly identified as mutation sites in NGS sequencing due to various technical reasons such as sequencing errors and alignment bias. The graylist contains complex mutation sites that are difficult to automatically interpret and require manual review. These sites may have low mutation abundance, complex variation patterns, or be located in complex regions of the genome. For each candidate mutation, based on the mutation data of the candidate mutation, namely mutation type, mutation abundance AF, supporting read number AD and other information, a quick match is made to see whether the mutation meets the conditions of the blacklist or graylist. That is, whether the mutation data of the candidate mutation meets the data conditions for matching the blacklist or whether the mutation data of the candidate mutation meets the data conditions for matching the blacklist, or whether the coordinates of the candidate mutation match any site in the blacklist, then the mutation is considered to be likely to be a false positive. If the coordinates of the candidate mutation match any site in the graylist, then the mutation is considered to require further manual review and no clear interpretation result can be given.
[0078] S30: If the candidate mutation exists in the preset list, a judgment result corresponding to the candidate mutation is generated according to the list type of the preset list.
[0079] Specifically, if it is determined that the candidate mutation exists in the preset list, a preliminary interpretation result is generated based on the type of list in which the candidate mutation is located. If the candidate mutation exists in the blacklist, the interpretation result of the candidate mutation is generated and marked as Fail, indicating that the mutation is considered to be a false positive and does not require further analysis. If the candidate mutation exists in the graylist, the interpretation result of the candidate mutation is generated and marked as requiring manual review, indicating that the mutation requires further judgment by human experts and no clear pass or fail conclusion can be given.
[0080] S40: If the candidate mutation does not exist in the preset list, the hotspot mutation type and sample quality type of the candidate mutation of the sample to be analyzed are determined according to the mutation detection result.
[0081] Specifically, if the candidate mutation is not in the blacklist and graylist, its nature and sample quality need to be further analyzed to determine how to interpret it. First, query the preset hotspot mutation whitelist to determine whether the candidate mutation is a hotspot mutation. Hotspot mutations refer to mutations that frequently appear in specific genes and have important clinical significance, such as the common L858R mutation in the EGFR gene, or the common G12C mutation in the KRAS gene. Then, evaluate the quality of the sample to be analyzed, including determining whether the sample has degraded, the background noise level of the sample, and whether the sample is a high-noise sample. Sample quality will significantly affect the accuracy of mutation detection, so it needs to be considered to obtain the hotspot mutation type and sample quality type of the candidate mutation of the sample to be analyzed.
[0082] S50: Based on the hotspot mutation type, sample quality type and mutation data, generate the interpretation result corresponding to the candidate mutation.
[0083] Specifically, after determining the hotspot mutation type and sample quality type, the corresponding dynamic threshold is selected based on the two. Each type of sample and mutation combination has a corresponding dynamic threshold, and the specific threshold is set according to the historical data of the training set and the results of the validation set. For example, for hotspot mutations in high-quality samples, a looser threshold may be used, while for non-hotspot mutations in low-quality samples, a stricter threshold is used. The mutation data of the candidate mutation, such as mutation abundance AF, number of supporting reads AD, and site sequencing depth DP, are compared with the selected dynamic threshold. If the threshold requirements are met, the interpretation result of the candidate mutation is generated and marked as qualified Pass. Otherwise, the interpretation result of the candidate mutation is generated and marked as requiring manual review, thereby obtaining the interpretation result corresponding to the candidate mutation. The interpretation result can also be output to the mutation detection result of the sample, and the uninterpreted results are further reviewed and confirmed manually.
[0084] In one embodiment, in step S30, based on the list type of the preset list, generating the interpretation result corresponding to the candidate mutation specifically includes:
[0085] S31: If the candidate mutation exists in the blacklist, a judgment result of the candidate mutation is generated and marked as unqualified.
[0086] Specifically, if a candidate mutation is determined to be in the blacklist, it means that the candidate mutation is considered to be a high-probability false-positive result. Since the mutations in the blacklist are not clinically significant, the interpretation result of the candidate mutation is generated and marked as Fail. This can avoid unnecessary subsequent analysis of these known false-positive mutations, improve efficiency and reduce false positives.
[0087] S32: If the candidate mutation exists in the gray list, the interpretation result of the candidate mutation is generated and marked as requiring manual review.
[0088] Specifically, if a candidate mutation is determined to be present in the gray list, it indicates that the mutation is complex or difficult to automatically judge. These mutations may have low accuracy in automatic interpretation due to various reasons such as low frequency, complex genomic regions, etc. Therefore, the interpretation results of the candidate mutation are generated and marked as requiring manual review. Then, these candidate mutations marked as requiring manual review are submitted to professional clinical geneticists or pathologists for manual review to ensure the accuracy and reliability of the interpretation.
[0089] In one embodiment, in step S40, based on the mutation detection results, determining the hotspot mutation type and sample quality type of the candidate mutations of the sample to be analyzed specifically includes:
[0090] S41: Determine whether the candidate mutation of the sample to be analyzed exists in a preset whitelist. The whitelist is constructed based on known hotspot mutations in public clinical databases.
[0091] Specifically, by querying a pre-established whitelist L whitelist , the whitelist L whitelist It contains known hotspot mutation information compiled from public clinical databases such as COSMIC and OncoKB. Hotspot mutations are mutations that occur frequently in specific genes and are closely related to the occurrence, development or treatment response of tumors. For example, the p.L858R mutation of the EGFR gene and the p.G12D mutation of the KRAS gene are included in the whitelist. If a candidate mutation of the sample to be analyzed is found in the whitelist, it indicates that the mutation is a known and highly credible hotspot mutation, and it is determined to hit the whitelist; otherwise, it is considered a miss.
[0092] S42: If the candidate mutation exists in the whitelist, determining the hotspot mutation type of the candidate mutation as a hotspot mutation.
[0093] Specifically, if a candidate mutation of the sample to be analyzed is found in the whitelist, the hotspot mutation type of the candidate mutation is marked as a hotspot mutation. A hotspot mutation refers to a genetic variation that plays a key role in the occurrence and development of tumors and has been clinically verified. For example, the TP53 gene p.R273H mutation is widely present in a variety of solid tumors and has important guiding significance for treatment strategies. Therefore, when a mutation such as TP53 p.R273H appears in a sample and matches the whitelist, it is immediately identified as a hotspot mutation. There is no need to further rely on mutation abundance AF, supporting reads AD, etc. to determine whether it is a hotspot, ensuring priority protection and identification of clinically important mutations.
[0094] S43: If the candidate mutation does not exist in the whitelist, the hotspot mutation type is determined to be a non-hotspot mutation.
[0095] Specifically, if a candidate mutation fails to match the whitelist, the hotspot mutation type of the candidate mutation will be marked as a non-hotspot mutation. Non-hotspot mutations refer to mutations for which there is currently insufficient clinical evidence to support their relevance to tumor occurrence, development, or treatment. For example, some newly discovered rare mutations, such as some non-pathogenic mutations in the BRCA2 gene, are classified as non-hotspot mutations even if they are detected due to the lack of a large number of clinical research verifications. This allows more stringent or prudent screening criteria to be applied to non-hotspot mutations in the subsequent interpretation process to reduce the misjudgment rate.
[0096] S44: Obtaining mutation spectrum characteristics and background noise levels of the sample to be analyzed from the mutation detection results, and determining a sample degradation type of the sample to be analyzed based on the mutation spectrum characteristics.
[0097] Specifically, in addition to the hotspot mutation type, the quality of the sample itself will also affect the accuracy of mutation detection. Two key sample quality indicators are extracted from the mutation detection results: mutation spectrum characteristics and background noise level. Mutation spectrum characteristics refer to the distribution of various mutation types in the sample. For example, degraded samples usually have a higher proportion of C>T (after correction of the positive and negative chains of the reference genome) mutations. Background noise level refers to low-frequency error signals that also appear in normal tissues. High background noise levels may lead to more false positive results. By analyzing these two indicators, the sample quality can be more accurately assessed and the samples can be divided into different quality types so that the corresponding thresholds or standards can be applied in subsequent interpretations. Then, based on the extracted mutation spectrum characteristics, a preset algorithm or model is used to determine whether the sample has been degraded. For example, if the proportion of C>T mutations exceeds a certain threshold, the sample is considered to have been degraded, thereby obtaining the sample degradation type of the sample to be analyzed.
[0098] S45: Determine the sample noise type of the sample to be analyzed according to the background noise level.
[0099] Specifically, the extracted background noise level data is used to further determine the noise type of the sample. The noise level of the sample is usually compared with a preset threshold. If the noise level of the sample is higher than the threshold, the sample is considered to be a high-noise sample, otherwise it is a non-high-noise sample. High-noise samples may require stricter filtering or a higher mutation abundance threshold to avoid false positives, thereby obtaining the sample noise type of the sample to be analyzed.
[0100] S46: Determine a sample quality type of the sample to be analyzed based on the sample degradation type and the sample noise type.
[0101] Specifically, after comprehensively considering the degradation type and noise type of the sample, the sample will be classified into an overall quality type. For example, if the degradation type of the sample to be analyzed is a degraded sample or the noise type is a high noise sample, and either one meets the low-quality type judgment criteria, then the sample to be analyzed is classified as a low-quality sample; otherwise, it is classified as a normal sample, thereby obtaining the sample quality type of the sample to be analyzed.
[0102] In one embodiment, in step S44, the sample degradation type of the sample to be analyzed is determined based on the mutation spectrum characteristics, specifically including:
[0103] S441: Count the total number of mutations of SNV type in the mutation data of the sample to be analyzed.
[0104] Specifically, the total number of mutations N whose mutation type is single nucleotide variation (SNV) in the sample to be analyzed is counted. SNV SNV refers to the change of a single nucleotide in the DNA sequence, for example, from A to G, or from C to T. By analyzing the mutation detection results of the sample, all candidate mutations marked as SNV in the mutation type field are screened out, and these mutations are counted. The total number obtained is used for subsequent calculation of degradation characteristics. For example, in an NGS test report, a total of 420 mutation records were detected, of which 360 SNV mutations were screened and confirmed. The total number of SNV mutations in the sample is recorded as 360, which is used for subsequent calculation of sample degradation ratio. For example, N SNV It can be expressed as , where N total is the total number of all mutations in the sample, 1{condition} is the indicator function, which takes 1 when the condition is met, otherwise it takes 0, N SNV is the total number of SNV mutations in the sample.
[0105] S442: From the total number of mutations, the number of mutations whose mutation form is C>T and whose mutation abundance in the mutation data is less than a preset abundance threshold is obtained.
[0106] Specifically, within the total number of SNV mutations, a specific subset of SNVs is further screened. The screening criteria are two: the mutation type is C>T, meaning the base C in the DNA sequence changes to T; and the mutation abundance AF is less than a preset abundance threshold, such as 3%. Mutation abundance refers to the proportion of sequencing reads carrying the mutated base to the total number of reads. Here, a low threshold, such as 3%, is used to focus on low-frequency mutations, thereby counting the number of C>T mutations N that meet these two conditions. C→T,AF<3% , among which, low-frequency C>T mutations are enriched in degraded DNA samples, so counting their number helps to judge the sample quality. For example, among the 360 SNV mutations mentioned above, there are 90 C>T mutations, of which 72 have mutation abundances less than 3%. The number of C>T low-abundance mutations is recorded as 72 for subsequent degradation ratio calculation. For example, N C→T,AF<3% It can be expressed as , where N total is the total number of all mutations in the sample, 1{condition} is the indicator function, which takes 1 when the condition is met, otherwise it takes 0, N C→T,AF<3% is the total number of C>T low-frequency mutations that meet the criteria.
[0107] S443: Based on the total number of mutations and the number of mutations, the degradation ratio of the sample is calculated.
[0108] Specifically, after obtaining the total number of SNV mutations N SNV and the number of C>T low-abundance mutations N C→T,AF<3% Then, the number of low-frequency C>T mutations N C→T,AF<3% Divide by the total number of mutations N SNV , and the sample degradation ratio P is obtained C→T,AF<3% This ratio represents the proportion of low-frequency C>T mutations in all SNV mutations, which can reflect the degree of sample degradation. For example, if the number of C>T low-abundance mutations is 72 and the total number of SNV mutations is 360, the degradation ratio is 72 / 360=0.2, or 20%, which is used to determine whether the sample is a degraded sample. For example, P C→T,AF<3% It can be expressed as , where P C→T,AF<3% is the proportion of C>T low-frequency mutations that meet the conditions.
[0109] S444: If the sample degradation ratio is greater than or equal to the degradation judgment threshold, the sample degradation type of the sample to be analyzed is determined to be a degraded sample.
[0110] Specifically, after calculating the sample degradation ratio, it is compared with a preset degradation judgment threshold, such as 65%. If the sample degradation ratio is greater than or equal to the degradation judgment threshold, it is considered that the sample has undergone obvious degradation, and the sample degradation type is determined to be a degraded sample. For example, when the degradation ratio of a sample is 72%, which is greater than the threshold of 65%, the sample is determined to be a degraded sample, and the dynamic threshold standard set for the degraded sample is used in the subsequent mutation interpretation process.
[0111] S445: If the sample degradation ratio is less than the degradation judgment threshold, the sample degradation type of the sample to be analyzed is determined to be a non-degraded sample.
[0112] Specifically, if the degradation ratio of a sample is less than the preset degradation judgment threshold, the degree of degradation of the sample is considered to be relatively light, and the degradation type of the sample is determined to be a non-degraded sample. For example, when the degradation ratio of a sample is 17%, which is lower than the judgment threshold of 65%, the sample is determined to be a non-degraded sample.
[0113] In one embodiment, in step S45, the sample noise type of the sample to be analyzed is determined based on the background noise level, specifically including:
[0114] S451: Based on the background noise level, obtain the hotspot mutation noise level of at least one candidate mutation whose hotspot mutation type is a hotspot mutation in the sample to be analyzed.
[0115] Specifically, background noise refers to low-frequency false mutation signals in normal tissues devoid of tumor cells, caused by sequencing errors or other technical reasons. For the sample to be analyzed, the first step is to determine which candidate mutations are hotspot mutations. Hotspot mutations are typically identified based on a preset whitelist. Then, the noise levels of these hotspot mutations are extracted from the sequencing data. If a sample contains multiple hotspot mutation sites, the background noise value for each site is extracted for subsequent averaging. For example, if a sample contains two hotspot mutation sites, EGFR p.L858R and KRAS p.G12D, both on the whitelist, the background noise levels of these two mutation sites need to be calculated. Their background noise levels are 0.4% and 0.5%, respectively. Therefore, the sample to be analyzed is matched against the whitelist to identify the mutation site, and the background noise level for each mutation site is then extracted.
[0116] S452: Calculate the single-sample background noise level statistics of all hotspot mutations in each sample to be analyzed based on the hotspot mutation noise level.
[0117] Specifically, after extracting the noise level of each hotspot mutation site, the background noise values of all hotspot sites in the same sample are sorted, and the 95% quantile is determined based on the sorting results as the statistical value of the background noise level of a single sample. , 95% quantile means that 95% of the data is less than or equal to this value, and 5% of the data is greater than this value. It is used to exclude the interference of abnormal small values and more accurately reflect the actual background noise of the sample. For example, the hotspot mutation noise levels in a certain sample are 0.2%, 0.3%, 0.4%, 0.6%, and 0.7% respectively. Then, after arranging them in order of size, the 95% quantile is taken and the single sample background noise level of the sample is calculated to be 0.7%. For example, the statistical value of the single sample background noise level is It can be expressed as , where M is the total number of mutations in the hotspot mutation whitelist, x i,j is the hotspot mutation noise level of the i-th hotspot mutation in the j-th sample, is the 95% quantile value of the noise level of all hotspot mutations in the jth sample.
[0118] S453: If the single-sample background noise level statistic is greater than or equal to a preset threshold, it is determined that the sample noise type of the sample to be analyzed is a high-noise sample.
[0119] Specifically, in order to distinguish whether the noise level of a sample is normal or abnormally high, a preset threshold needs to be set. This preset threshold is usually based on the noise level data of a large number of normal clinical samples, and is obtained by using the 95% quantile method. Specifically, the calculated single-sample background noise level statistical value is compared with this preset threshold. If the noise level statistical value of the sample to be analyzed is greater than or equal to this preset threshold, it is considered that the noise level of the sample is significantly higher than the normal level and is a high-noise sample. A high-noise sample means that there may be more false-positive mutations in its sequencing results, requiring more stringent quality control and subsequent manual review. For example, the preset threshold T can be expressed as , where N is the number of clinical samples, N ≥ 1000. For example, when the background noise level of a single sample reaches 0.6% and the preset threshold is 0.5%, the sample is marked as a high-noise sample, and subsequent interpretation uses more stringent parameter screening or prompts manual review.
[0120] S454: If the single-sample background noise level statistic is less than a preset threshold, the sample noise type of the sample to be analyzed is determined to be a non-high noise sample.
[0121] Specifically, if the single-sample background noise level of the sample to be analyzed is less than the preset threshold, the noise level of the sample is considered to be within the normal range, and the sample noise type is determined to be a non-high-noise sample. A non-high-noise sample means that the background signal interference is within an acceptable range, and the conventional dynamic threshold can be used for mutation interpretation without additional processing. For example, if the single-sample background noise level of a sample is only 0.3%, which is significantly lower than the statistical value of the sample background noise level of 0.5%, it is classified as a non-high-noise sample and directly participates in the subsequent conventional interpretation process.
[0122] In one embodiment, in step S46, the sample quality type of the sample to be analyzed is determined based on the sample degradation type and the sample noise type, specifically including:
[0123] S461: If the sample degradation type of the sample to be analyzed is a degraded sample and / or the sample noise type of the sample to be analyzed is a high-noise sample, determine that the sample quality type of the sample to be analyzed is a low-quality sample.
[0124] Specifically, after determining the degradation type and noise type of the sample to be analyzed, if the degradation type is determined to be a degraded sample or the noise type is determined to be a high-noise sample, the sample to be analyzed is directly marked as a low-quality sample. Degraded samples usually have a decreased sequencing accuracy due to severe DNA fragmentation, and high-noise samples indicate that the background interference level is too high, which is not conducive to the true identification of mutations. For example, if the degradation ratio of a sample to be analyzed is 75%, which is greater than the threshold of 65%, and the 95% quantile of the hotspot mutation background noise is 0.55%, which is greater than the noise threshold of 0.5%, then the sample exceeds the standard in any indicator and is therefore determined to be a low-quality sample. In the subsequent mutation interpretation process, a more relaxed interpretation standard needs to be used or resampling is recommended.
[0125] S462: If the sample degradation type of the sample to be analyzed is a non-degraded sample and / or the sample noise type of the sample to be analyzed is a non-high noise sample, determine that the sample quality type of the sample to be analyzed is a normal sample.
[0126] Specifically, after determining the degradation type and noise type of the sample to be analyzed, if the degradation type is determined to be a non-degraded sample and the noise type is determined to be a non-high-noise sample, the sample to be analyzed will be marked as a normal sample. Normal samples have good DNA integrity and low background noise levels, and are suitable for applying standard dynamic thresholds for high-confidence mutation interpretation. For example, if the degradation ratio of a sample to be analyzed is only 15%, which is lower than the set standard of 65%, and the 95% quantile of the hotspot mutation background noise is 0.3%, which is far lower than the noise threshold of 0.5%, the system determines that the sample is a normal sample. In subsequent steps, mutations can be strictly screened according to standard parameters to ensure the reliability of the results.
[0127] In one embodiment, in step S50, based on the hotspot mutation type, sample quality type, and mutation data, a call result corresponding to the candidate mutation is generated, specifically including:
[0128] S51: Based on the hotspot mutation type and sample quality type of the sample to be analyzed, a corresponding preset dynamic threshold is obtained.
[0129] Specifically, when obtaining dynamic thresholds based on the hotspot mutation type and sample quality type of the sample to be analyzed, first, based on whether the mutation is a hotspot mutation and whether the sample is a normal sample, a corresponding group of dynamic judgment threshold combinations that have been trained and optimized in advance is selected. The dynamic threshold combination includes the minimum standard values of mutation abundance AF, supporting reads number AD, and site sequencing depth DP. Each group of thresholds is determined based on historical sample training data through weighted optimization of specificity and coverage. For example, if the sample is a normal sample and the mutation is a hotspot mutation, the threshold group of AF≥1%, AD≥30, and DP≥300 corresponding to the normal sample-hotspot mutation group is called to ensure that accurate judgment standards are set for different types of samples and mutation characteristics.
[0130] S52: Based on whether the mutation data of the candidate mutation meets the preset dynamic threshold standard, if so, a judgment result of the candidate mutation is generated and marked as qualified; if not, a judgment result of the candidate mutation is generated and marked as requiring manual review.
[0131] Specifically, after obtaining the corresponding dynamic threshold, the candidate mutation data of the sample to be analyzed are compared with the dynamic threshold standard item by item in turn. If the mutation abundance AF, supporting reads number AD and site sequencing depth DP all meet the corresponding standards, the judgment result of the candidate mutation is generated and marked as qualified Pass. Conversely, if any indicator does not meet the threshold requirements, the judgment result of the candidate mutation is generated and marked as requiring manual review. For example, for a mutation, if its mutation abundance AF is 1.2%, the supporting reads number AD is 35, and the site depth DP is 420, all of which meet the threshold standards corresponding to the normal sample-hotspot mutation category, the candidate mutation is judged as qualified Pass. Otherwise, if the supporting reads AD is only 25 and below the standard of 30, it is marked as requiring manual review and pushed to the review process.
[0132] In one embodiment, if Figure 2 As shown, before step S50, a method for interpreting tumor gene mutations based on an NGS platform further includes:
[0133] S501: Obtain a historical sample data set, and divide the historical sample data set into a test set and a validation set according to a preset ratio.
[0134] Specifically, when obtaining historical sample data sets, large-scale NGS detection data are collected, where each historical sample data contains mutation abundance AF, supporting reads number AD, site sequencing depth DP and corresponding manual interpretation labels. To ensure the objectivity and generalization ability of the parameter optimization process, the historical data are divided into a test set and a validation set according to a fixed ratio of 4:1. The test set is used for training and optimization of dynamic threshold combinations, and the validation set is used to finally evaluate the consistency index of the obtained threshold combination. For example, after collecting 10,000 historical mutation data, 8,000 are divided as a test set and 2,000 as a validation set, providing a data basis for the subsequent search for optimal interpretation parameters.
[0135] S502: Group the test set and validation set according to the hotspot mutation type and sample quality type.
[0136] Specifically, after completing the data set division, the test set and validation set are grouped separately according to the hotspot mutation type corresponding to each mutation, i.e. hotspot mutation or non-hotspot mutation, and the sample quality type, i.e. normal sample or low-quality sample, to form four combinations: normal sample-hotspot mutation, normal sample-non-hotspot mutation, low-quality sample-hotspot mutation, and low-quality sample-non-hotspot mutation. The threshold is optimized independently for the internal data of each combination to avoid interpretation bias caused by mixed processing of mutations of different qualities or different clinical significances. For example, if a test set mutation comes from a low-quality sample and is a non-hotspot mutation, it will be classified into the low-quality sample-non-hotspot mutation group to ensure the targetedness and accuracy of parameter training.
[0137] S503: For each group of the test set, select all mutation sites that are manually judged as unqualified, and set the initial threshold boundary based on the mutation abundance, supporting read number and site sequencing depth of all mutation sites.
[0138] Specifically, all mutation sites that are manually interpreted as unqualified Fail are screened out in each group, and the 95% quantiles of these unqualified Fail mutations in three dimensions, namely, mutation abundance AF, number of supporting reads AD, and site sequencing depth DP, are counted as the preliminary search boundaries. The initial boundaries are used to limit the subsequent dynamic threshold search range to avoid inefficient training due to excessively large search space. For example, in the normal sample-hotspot mutation group, the unqualified Fail mutation data are counted, and the 95% quantile of AF is 1%, the 95% quantile of AD is 30, and the 95% quantile of DP is 300. These values are then used as the starting boundaries of the dynamic threshold search to provide an effective range for subsequent optimization. For example, the initial threshold boundary T init It can be expressed as , where T init is the initial threshold boundary, R fail A collection of unqualified Fail mutations.
[0139] S504: Based on the initial threshold boundary, by constructing a multidimensional grid space and discretely dividing the threshold combination candidate area, through grid search and optimization algorithm, based on the threshold combination candidate area, using a weighted objective function, determine the threshold combination with the largest weighted objective function value for the corresponding group in the test set.
[0140] Specifically, after determining the initial boundary, a three-dimensional grid space G is constructed based on the three parameter axes of mutation abundance AF, supporting reads AD and site sequencing depth DP. i,j,k , the boundaries of each dimension are discretized into several small intervals to form a large number of candidate threshold combination areas. For each set of candidate threshold combinations, the specificity and coverage are calculated in the test set respectively. The specificity is the proportion of mutations that are actually manually judged as qualified passes among the mutations judged as qualified passes, and the coverage is defined as the proportion of mutations that are correctly judged as qualified passes among the mutations that are manually judged as qualified passes. A comprehensive evaluation is performed according to the preset weighted objective function F=λ×Specificity+(1-λ)×Coverage, where λ is a weight factor, such as 0.7 or 0.8. Finally, the threshold combination with the highest objective function F value is selected as the dynamic threshold T of the corresponding group. final For example, in the normal sample-hotspot mutation group, the best combination is selected after grid search as AF≥1%, AD≥35, DP≥320, ensuring the best comprehensive interpretation performance. final It can be expressed as ,in and , , where For the grid G i,j,k The number of mutations that were manually interpreted as qualified passes, is the grid G i,j,k The number of mutations that were manually judged as unqualified Fail, Specificity is the proportion of mutations that were actually manually judged as qualified Pass among the mutations that were judged as qualified Pass, indicating how many proportions of unqualified Fail mutations were successfully excluded from the credible region, Coverage is the proportion of manually judged qualified Pass mutations that were correctly judged as qualified Pass, indicating how many proportions of qualified Pass mutations were included in the credible region, (α, β, γ) is a combination variable that represents the threshold value in three-dimensional space, α is the lower limit of mutation abundance AF, β is the lower limit of the number of supported reads AD, and γ is the lower limit of the site sequencing depth DP. Therefore, through various (α, β, γ) combinations, the combination that maximizes the weighted index value is selected as the final threshold T final .
[0141] S505: Validate using the validation set, record the mutations of each sample in the validation set that fall within the threshold combination of the corresponding group as qualified mutations, and calculate the consistency index of the qualified mutations in the corresponding group in the validation set.
[0142] Specifically, in determining the T of each group in the test set final After combination, use the validation set to test T final Combine and verify, and substitute each candidate mutation in the verification set into the T of the corresponding group final Threshold combination is used for interpretation. If the mutation abundance AF, supporting read number AD and site sequencing depth DP of the candidate mutation are all greater than or equal to T final If the mutation corresponds to the standard, it is judged as a qualified mutation, otherwise it is judged to require manual review or fail. The number of mutations judged as qualified Pass that are consistent with the manual review label is counted, and the consistency index P is calculated. The consistency index P is defined as the number of mutations judged as qualified Pass and manually reviewed as qualified Pass in the verification set divided by the total number of mutations manually reviewed as qualified Pass in the verification set. For example, if there are 1000 mutations manually reviewed as qualified Pass in the verification set, use T final After combined interpretation, 950 of them are correctly identified as qualified Pass mutations, and the consistency index is 950 / 1000=95%. For example, the consistency index P can be expressed as .
[0143] S506: If the consistency index is greater than or equal to the preset consistency threshold, determine the preset dynamic threshold of the corresponding group according to the threshold combination.
[0144] Specifically, after calculating the consistency index P of each group in the validation set, the consistency index is compared with the preset consistency threshold. If the consistency index P is greater than or equal to the set consistency standard value, such as 99%, the T selected in the test set is confirmed. final The combination is used as the formal dynamic threshold for the group. Otherwise, the threshold search range needs to be readjusted or the optimization strategy needs to be retrained. For example, if the normal sample-hotspot mutation group is applied to the validation set, T final The combined verification consistency is 99.5%, which is higher than the set standard of 99%. The dynamic threshold combination AF≥1%, AD≥35, and DP≥320 determined in the test set is used as the official dynamic interpretation standard for this category of mutations to ensure the stability and high confidence of the interpretation of subsequent new samples. This dynamic threshold standard is directly applied in the automatic interpretation of new samples. For example, the preset dynamic threshold T for each group is final It can be expressed as .
[0145] In one embodiment, if Figure 3As shown, before step S20, a method for interpreting tumor gene mutations based on an NGS platform further includes:
[0146] S2010: Screen the mutation sites in the historical sample data set that are manually interpreted as unqualified, and calculate the screening frequency and mutation abundance standard deviation of the selected mutation sites.
[0147] Specifically, all mutation sites that were manually interpreted as Fail were screened from the historical sample dataset. For each Fail mutation site that was screened out, its frequency of occurrence in all samples was counted as the screening frequency f(l), and the standard deviation σ of the mutation abundance AF of the site in different samples was calculated. AF (l), the screening frequency indicates the proportion of samples that are judged as unqualified Fail at this site, and the standard deviation of mutation abundance indicates the degree of fluctuation of the mutation abundance of this site in different samples. For example, in historical data, if a site has 120 unqualified Fail mutations in 5000 samples, and its abundance standard deviation is 0.08, then its screening frequency is recorded as 2.4% and its abundance standard deviation is 0.08, providing basic data for subsequent list screening. For example, the screening frequency f(l) can be expressed as , S is the clinical sample set, S≥1000, standard deviation σ AF (l) can be expressed as , AF i is the mutation abundance AF of the i-th sample, so the screening frequency of the mutation site and the standard deviation of the mutation abundance can be obtained.
[0148] S2011: Based on the screening frequency, the mutation sites with a screening frequency greater than or equal to a preset frequency threshold are classified as a high-frequency unqualified mutation site set.
[0149] Specifically, a preset frequency threshold is set, such as 50%, and the calculated screening frequency f(l) is compared with the threshold. Mutation sites with a screening frequency f(l) greater than or equal to the threshold are classified as a high-frequency unqualified Fail mutation site set L. high-fail-freq , which means that these sites are judged as Fail in more than half of the samples. They are likely to be systematic false positives, which can further narrow the scope of analysis. For example, if the screening frequency of a mutation site is 55%, which is higher than the preset frequency threshold of 50%, it will be included in the high-frequency Fail mutation site set for further processing in the subsequent blacklist screening process. For example, the high-frequency Fail mutation site set L high-fail-freq It can be expressed as L high-fail-freq ={l∈L | f(l)≥0.5}, where L is the set of all mutation sites.
[0150] S2012: In the set of high-frequency unqualified mutation sites, screen the first site whose mutation abundance standard deviation is less than the preset stability threshold, and manually review the first site screened and confirm its known clinical significance.
[0151] Specifically, in the obtained high-frequency unqualified Fail mutation site set L high-fail-freq In the experiment, we further screened the mutation abundance standard deviation σ AF (l) The sites that are less than the preset stability threshold α, for example, 2%, constitute the site set L similar These sites not only have a high frequency of occurrence, but also have relatively stable mutation abundance in different samples, which further increases the possibility that they are false positives. These sites are taken as the first sites and manually reviewed. That is, professional pathologists or genomics experts use tools such as IGV to carefully check the sequencing data of these sites to confirm whether they are really false positives. At the same time, relevant databases and literature are consulted to confirm whether these sites have known clinical significance. If these sites are real mutations, they may play an important role in certain tumor types, avoiding the mistaken addition of real, clinically significant mutations to the blacklist. For example, if the standard deviation of the abundance of a mutation site is 0.06 and there is no known clinical pathogenic annotation, the standard is met and the next step of blacklist screening is entered. For example, the site set L similar It can be expressed as L similar ={l∈L high-fail-freq |σ AF (l)<α}.
[0152] S2013: The first site that is still unqualified after manual review and has no known clinical significance will be included in the blacklist.
[0153] Specifically, for the first sites that have been manually reviewed, if the manual review confirms that they are still unqualified Fail, i.e., false positives, and after consulting the literature and databases, it is found that they have no known clinical significance, these sites will be included in the blacklist. blacklist The sites in the blacklist will be directly interpreted as unqualified Fail in the subsequent automatic interpretation without further analysis. For example, if a high-frequency, low-dispersion mutation site has no clinical significance after review and the manual interpretation keeps the unqualified Fail label for a long time, the site will be entered into the blacklist and marked as unqualified Fail in the subsequent interpretation, thus improving the accuracy and efficiency of the overall interpretation. blacklist It can be expressed as L blacklist ={l∈L similar |Manual review (l) = fail ∧ ClinSig (l) = 0}, ClinSig (l) is the clinical significance of site l, expressed as a Boolean value, 0 = meaningless, 1 = meaningful.
[0154] More specifically, after the initial construction of the blacklist set, the performance of the mutation sites hitting the blacklist in new samples is continuously monitored during the subsequent interpretation process. If a blacklist site is manually reviewed and confirmed as qualified Pass in the newly collected large sample data many times, and its mutation abundance AF is stable and the background noise is low, indicating that the site may have certain clinical significance due to sample size expansion or biological background changes, then the site will be automatically included in the list for review. After further manual confirmation and literature database retrieval, if it is indeed clinically relevant or passes the review, the site will be removed from the blacklist. Conversely, if a site continues to maintain a high-frequency Fail characteristic in the new samples and there is no significant improvement in the abundance stability, its weight index in the blacklist will be updated or an additional annotation will be added. At the same time, if a new mutation site that has newly emerged and meets the high-frequency, low-dispersion, and no clinical significance standards is found in the new samples, the screening frequency and standard deviation will be recalculated according to the initial screening logic. When the conditions are met, the blacklist will be automatically expanded to achieve dynamic maintenance and intelligent iteration of the blacklist.
[0155] In one embodiment, if Figure 4 As shown, before step S20, a method for interpreting tumor gene mutations based on an NGS platform further includes:
[0156] S2020: Screen out mutation sites that are not on the whitelist and blacklist from the mutation sites, and calculate the abundance dispersion index of the screened mutation sites.
[0157] Specifically, from all detected mutation sites, those that belong to the white list, i.e., known, clinically significant hotspot mutations, and the black list, i.e., known false-positive mutations, are excluded. For the remaining mutation sites L candidate , calculate their abundance dispersion index CV(l), which is used to measure the degree of variation of the mutation abundance of these sites in different samples. In the embodiment, the coefficient of variation CV, that is, the standard deviation divided by the mean, is used as the abundance dispersion index. The larger the CV value, the greater the variation of the mutation abundance in different samples. These sites may be difficult to interpret automatically and require manual review. For example, if the abundance value of a mutation site varies dramatically in different samples and the standard deviation reaches 0.45, it means that the dispersion is large and meets the gray list screening conditions. Among them, L candidate It can be expressed as L candidate ={ l∈L |l∉L blacklist ∧l∉L whitelist ∧f(l)>0}.
[0158] S2021: Mutation sites whose abundance dispersion index is greater than the preset dispersion threshold will be included in the gray list.
[0159] Specifically, a preset dispersion threshold value, such as 0.3, is set, and the calculated abundance dispersion index is compared with the threshold value CV(l), and the mutation sites whose abundance dispersion index is greater than the threshold value are included in the gray list L graylist The sites in the gray list will be marked as requiring manual review in the subsequent automatic interpretation, and manual review will be mandatory. For example, if the abundance standard deviation of a mutation site in multiple batches of samples is 0.36, which is significantly greater than the set threshold of 0.3, it will be included in the gray list. When such sites are hit during interpretation, manual review will be prompted to avoid automatic missed or misjudgment. graylist It can be expressed as L graylist ={ l∈L graylist |CV(l)>0.3}.
[0160] More specifically, after the initial construction of the gray list set, the latest interpretation results of the gray list sites are collected during the interpretation process. If a gray list site is manually reviewed and confirmed as qualified in a large number of subsequent samples, and the mutation abundance AF, supporting reads number AD and site sequencing depth DP indicators are stable, it can be removed from the gray list after manual review and database verification. If a gray list site has multiple interpretation differences in new samples, and the mutation data fluctuations intensify, or it is re-confirmed as a false positive after manual review, the site will be transferred to the blacklist for strict interception management. At the same time, in the process of continuously introducing new sample data, based on the comparison of the new sample mutation parameters with the dynamic threshold, when a large number of mutation sites close to the judgment boundary and inconsistent with the manual review results are found, they are screened according to the abundance dispersion index CV(l) and the preset dispersion threshold. The newly screened mutation sites can be expanded to the gray list, thereby maintaining the synchronous update and adaptive optimization of the gray list and the distribution characteristics of the sample data.
[0161] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0162] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A method for interpreting tumor gene mutations based on an NGS platform, characterized in that: The tumor gene mutation interpretation method comprises: Obtaining mutation detection results for the sample to be analyzed, the mutation detection results including at least one candidate mutation and mutation data corresponding to the candidate mutation, the mutation data including mutation type, mutation abundance, number of supporting reads, site sequencing depth, and manual interpretation labels; According to the mutation detection result, determining whether the candidate mutation exists in a preset list, wherein the preset list includes at least a blacklist and a graylist; If the candidate mutation exists in the preset list, generating a judgment result corresponding to the candidate mutation according to the list type of the preset list; If the candidate mutation does not exist in the preset list, determining the hotspot mutation type and sample quality type of the candidate mutation of the sample to be analyzed according to the mutation detection result; Based on the hotspot mutation type, the sample quality type, and the mutation data, a judgment result corresponding to the candidate mutation is generated.
2. The method for interpreting tumor gene mutations according to claim 1, wherein: Generating the interpretation result corresponding to the candidate mutation according to the list type of the preset list specifically includes: If the candidate mutation exists in the blacklist, generating and marking the interpretation result of the candidate mutation as unqualified; If the candidate mutation exists in the gray list, a reading result of the candidate mutation is generated and marked as requiring manual review.
3. The method for interpreting tumor gene mutations according to claim 2, wherein: Determining the hotspot mutation type and sample quality type of the candidate mutation of the sample to be analyzed based on the mutation detection result specifically includes: Determine whether the candidate mutation of the sample to be analyzed exists in a preset whitelist, where the whitelist is constructed based on known hotspot mutations in public clinical databases; If the candidate mutation exists in the whitelist, determining the hotspot mutation type of the candidate mutation as a hotspot mutation; If the candidate mutation does not exist in the whitelist, determining the hotspot mutation type as a non-hotspot mutation; Obtaining mutation spectrum characteristics and background noise levels of the sample to be analyzed from the mutation detection results, and determining a sample degradation type of the sample to be analyzed based on the mutation spectrum characteristics; determining a sample noise type of the sample to be analyzed according to the background noise level; A sample quality type of the sample to be analyzed is determined based on the sample degradation type and the sample noise type.
4. The method for interpreting tumor gene mutations according to claim 3, wherein: Determining the sample degradation type of the sample to be analyzed based on the mutation spectrum characteristics specifically includes: Counting the total number of mutations with SNV type in the mutation data of the sample to be analyzed; From the total number of mutations, the number of mutations whose mutation form is C>T and whose mutation abundance in the mutation data is less than a preset abundance threshold is obtained; Based on the total number of mutations and the number of mutations, the degradation ratio of the sample is calculated; If the sample degradation ratio is greater than or equal to the degradation judgment threshold, the sample degradation type of the sample to be analyzed is determined to be a degraded sample; If the sample degradation ratio is less than the degradation judgment threshold, the sample degradation type of the sample to be analyzed is determined to be a non-degraded sample.
5. The method for interpreting tumor gene mutations according to claim 3, wherein: Determining the sample noise type of the sample to be analyzed according to the background noise level specifically includes: Based on the background noise level, obtaining a hotspot mutation noise level of at least one candidate mutation whose hotspot mutation type is a hotspot mutation in the sample to be analyzed; Calculating the single-sample background noise level statistics of all hotspot mutations in each sample to be analyzed based on the hotspot mutation noise level; If the single sample background noise level statistic is greater than or equal to a preset threshold, determining that the sample noise type of the sample to be analyzed is a high noise sample; If the single-sample background noise level statistic is less than the preset threshold, it is determined that the sample noise type of the sample to be analyzed is a non-high noise sample.
6. The method for interpreting tumor gene mutations according to any one of claims 3 to 5, wherein: The determining the sample quality type of the sample to be analyzed based on the sample degradation type and the sample noise type specifically includes: If the sample degradation type of the sample to be analyzed is a degraded sample and / or the sample noise type of the sample to be analyzed is a high-noise sample, determining that the sample quality type of the sample to be analyzed is a low-quality sample; If the sample degradation type of the sample to be analyzed is a non-degraded sample and / or the sample noise type of the sample to be analyzed is a non-high noise sample, it is determined that the sample quality type of the sample to be analyzed is a normal sample.
7. The method for interpreting tumor gene mutations according to claim 6, characterized in that: Generating a judgment result corresponding to the candidate mutation based on the hotspot mutation type, the sample quality type, and the mutation data specifically includes: Based on the hotspot mutation type and the sample quality type of the sample to be analyzed, obtaining a corresponding preset dynamic threshold; According to whether the mutation data of the candidate mutation meets the preset dynamic threshold standard, if it does, the interpretation result of the candidate mutation is generated and marked as qualified; if it does not meet the standard, the interpretation result of the candidate mutation is generated and marked as requiring manual review.
8. The method for interpreting tumor gene mutations according to claim 7, characterized in that: Before the step of obtaining a corresponding preset dynamic threshold based on the hotspot mutation type and the sample quality type of the sample to be analyzed, the tumor gene mutation interpretation method further includes: Obtain a historical sample data set, and divide the historical sample data set into a test set and a validation set according to a preset ratio; Grouping the test set and the validation set according to the hotspot mutation type and the sample quality type; For each group of the test set, all mutation sites with the manual interpretation label as unqualified are selected, and an initial threshold boundary is set according to the mutation abundance, the number of mutation supporting reads and the sequencing depth of the site of all mutation sites; Based on the initial threshold boundary, a multidimensional grid space is constructed and discretely divided into threshold combination candidate areas, and a grid search and optimization algorithm is used to determine the threshold combination with the maximum weighted objective function value for the corresponding group in the test set based on the threshold combination candidate areas and using a weighted objective function; Utilizing the validation set for validation, recording mutations in each sample in the validation set that fall within the threshold combination of the corresponding group as qualified mutations, and calculating the consistency index of the qualified mutations in the corresponding group in the validation set; If the consistency index is greater than or equal to the preset consistency threshold, the preset dynamic threshold of the corresponding group is determined according to the threshold combination.
9. The method for interpreting tumor gene mutations according to claim 8, characterized in that: Before the step of determining whether the candidate mutation exists in the preset list, the tumor gene mutation interpretation method further includes: Screening the mutation sites in the historical sample data set that are manually interpreted as unqualified, and calculating the screening frequency and mutation abundance standard deviation of the screened mutation sites; According to the screening frequency, the mutation sites whose screening frequency is greater than or equal to a preset frequency threshold are classified as a high-frequency unqualified mutation site set; In the set of high-frequency unqualified mutation sites, a first site whose mutation abundance standard deviation is less than a preset stability threshold is selected, and the selected first site is manually reviewed and confirmed with known clinical significance; The first site that is still unqualified and has no known clinical significance after manual review is included in the blacklist.
10. The method for interpreting tumor gene mutations according to claim 9, characterized in that: Before the step of determining whether the candidate mutation exists in the preset list, the tumor gene mutation interpretation method further includes: Screening out mutation sites that do not belong to the whitelist and the blacklist from the mutation sites, and calculating the abundance dispersion index of the screened mutation sites; The mutation sites whose abundance dispersion index is greater than the preset dispersion threshold are included in the gray list.
Citation Information
Patent Citations
NGS-based brain tumor molecular diagnosis analysis method
CN112102944A
Method and system for rapid genetic analysis
US20190325988A1