A tumor gene mutation interpretation method based on NGS platform
By using a pre-defined list and dynamic matching technology in the tumor gene mutation interpretation method, the problem of frequent false positives in existing methods has been solved, achieving efficient and accurate tumor gene mutation detection and improving the standardization and clinical applicability of interpretation.
Patent Information
- Application Number
- CN202511099556.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-07
AI Technical Summary
Existing methods for interpreting tumor gene mutations cannot effectively distinguish low-abundance mutations, background noise, or misjudged systematic false positives, resulting in frequent false positive results, especially when faced with complex gene mutations, where the accuracy of the judgment is low.
By obtaining the mutation detection results of the sample to be analyzed, it is determined whether the candidate mutation exists in the preset list, and the corresponding interpretation results are generated. The interpretation is dynamically adapted according to the hot mutation type, sample quality type and mutation data, including the use of blacklist, gray list and whitelist, and personalized subdivision interpretation is performed in combination with sample degradation type and noise type.
It improves the standardization and clinical applicability of tumor gene mutation interpretation, reduces false positive and false negative rates, enhances the accuracy and efficiency of interpretation, and ensures adaptability and reliability under different sample conditions.
Smart Images

Figure CN120600111B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of tumor gene mutation detection technology, and in particular to a tumor gene mutation interpretation method based on an NGS platform. Background Technology
[0002] Currently, with the advancement of genomics technology, tumor gene mutation detection methods based on NGS platforms have been widely applied in clinical cancer diagnosis and targeted therapy decision-making. NGS technology, through deep sequencing of gene mutations in tumor tissue samples, can accurately identify cancer-related gene mutations, providing support for personalized treatment.
[0003] Existing methods for interpreting tumor gene mutations typically rely on specific mutation detection algorithms and threshold standards to classify and analyze mutation data. These methods are unable to effectively distinguish between low-abundance mutations, background noise, or systemic false positives. Furthermore, they suffer from low accuracy when dealing with complex gene mutations, such as low-frequency mutations, low-dispersion mutations, or rare mutation sites, leading to frequent false positives. While various mechanisms can be used to filter false positive mutations, the types of false positive mutations are numerous and their mechanisms are not yet fully understood. Therefore, current false positive filtering methods have limited effectiveness and room for improvement. Summary of the Invention
[0004] This application provides a tumor gene mutation interpretation method based on the NGS platform, which can improve the standardization level and clinical applicability of tumor gene mutation interpretation, while effectively reducing the false positive rate and the missed detection rate.
[0005] The above-mentioned inventive objective of this application is achieved through the following technical solutions:
[0006] A method for identifying tumor gene mutations based on an NGS platform, the method comprising:
[0007] Obtain the mutation detection results of the sample to be analyzed. The mutation detection results include at least one candidate mutation and the mutation data corresponding to the candidate mutation. The mutation data includes mutation type, mutation abundance, number of supporting reads, site sequencing depth, and manually interpreted tags.
[0008] Based on the mutation detection results, it is determined whether the candidate mutation exists in a preset list, which includes at least a blacklist and a graylist;
[0009] If the candidate mutation exists in the preset list, then according to the list type of the preset list, the interpretation result corresponding to the candidate mutation is generated;
[0010] If the candidate mutation does not exist in the preset list, then based on the mutation detection results, the hot spot mutation type and sample quality type of the candidate mutation of the sample to be analyzed are determined;
[0011] Based on the hotspot mutation type, the sample quality type, and the mutation data, the interpretation result corresponding to the candidate mutation is generated.
[0012] By adopting the above technical solutions, the basic data characteristics of each candidate mutation can be accurately grasped by obtaining the mutation detection results of the sample to be analyzed, thereby avoiding interpretation bias caused by incomplete mutation information. By determining whether the candidate mutation exists in the preset list, known unreliable or discrepancies in interpretation can be eliminated in the early stage of the interpretation process, improving the efficiency and accuracy of the overall screening. By generating interpretation results based on the list type, a preliminary judgment with high credibility or requiring verification can be made quickly, thereby shortening the necessary time for manual review and improving the overall interpretation speed. After the candidate mutation does not exist in the preset list, further interpretation can be carried out by hot mutation type, sample quality type and mutation data, which can achieve dynamic adaptation and personalized subdivision of interpretation standards, thereby ensuring the accuracy and adaptability of interpretation under different sample conditions, and thus improving the standardization level and clinical applicability of tumor gene mutation interpretation.
[0013] In a preferred embodiment, this application can be further configured such that: generating the interpretation result corresponding to the candidate mutation based on the list type of the preset list specifically includes:
[0014] If the candidate mutation exists in the blacklist, the interpretation result of the candidate mutation is generated and marked as unqualified;
[0015] If the candidate mutation exists in the gray list, the interpretation result of the candidate mutation is generated and marked as requiring manual review.
[0016] By adopting the above technical solution, and generating the interpretation results of candidate mutations based on the list type of the preset list, the interpretation conclusion can be accurately output directly based on the type characteristics of the blacklist and graylist, thereby avoiding abnormal or controversial mutations in ordinary interpretation logic processing and improving the overall interpretation accuracy.
[0017] In a preferred embodiment, this application can be further configured such that: determining the hotspot mutation type and sample quality type of the candidate mutations in the sample to be analyzed based on the mutation detection results specifically includes:
[0018] Determine whether the candidate mutation of the sample to be analyzed exists in a preset whitelist, which is constructed based on known hotspot mutations in a public clinical database;
[0019] If the candidate mutation exists in the whitelist, then the hotspot mutation type of the candidate mutation is determined to be a hotspot mutation;
[0020] If the candidate mutation does not exist in the whitelist, then the hotspot mutation type is determined to be a non-hotspot mutation;
[0021] The mutation spectrum characteristics and background noise level of the sample to be analyzed are obtained from the mutation detection results, and the sample degradation type of the sample to be analyzed is determined based on the mutation spectrum characteristics.
[0022] Based on the background noise level, determine the sample noise type of the sample to be analyzed;
[0023] Based on the sample degradation type and the sample noise type, the sample quality type of the sample to be analyzed is determined.
[0024] By adopting the above technical solution, and by determining the hotspot mutation type and sample quality type of candidate mutations based on mutation detection results, it is possible to set differentiated interpretation criteria for samples of different quality based on the clinical value of mutations and the differences in sequencing quality of samples. This improves the targeting of mutation screening and the scientific rationality of interpretation criteria, and avoids misjudgment caused by treating samples with different characteristics in a one-size-fits-all manner under the same standard.
[0025] In a preferred embodiment, this application can be further configured such that: determining the sample degradation type of the sample to be analyzed based on the mutation spectrum characteristics specifically includes:
[0026] Count the total number of mutations of type SNV in the mutation data of the sample to be analyzed;
[0027] From the total number of mutations, the number of mutations with the mutation form C>T and whose mutation abundance in the mutation data is less than a preset abundance threshold is obtained;
[0028] Based on the total number of mutations and the number of mutations, the sample degradation percentage is calculated;
[0029] If the percentage of sample degradation is greater than or equal to the degradation judgment threshold, then the sample degradation type of the sample to be analyzed is determined to be a degraded sample.
[0030] If the percentage of sample degradation is less than the degradation judgment threshold, then the sample to be analyzed is determined to be a non-degraded sample.
[0031] By adopting the above technical solution, the sample degradation type can be determined based on the mutation spectrum characteristics. Potentially degraded samples can be identified based on the proportion of low-frequency C>T mutations in the overall sample, thereby providing early warning of sequencing data distortion risks and improving the sample quality identification rate before interpretation. By setting a degradation judgment threshold and calculating the degradation ratio, the degree of sample degradation can be quantitatively analyzed, thereby objectively supporting the selection of different interpretation standards and avoiding reliance on subjective experience judgment.
[0032] In a preferred embodiment, this application can be further configured such that: determining the sample noise type of the sample to be analyzed based on the background noise level specifically includes:
[0033] Based on the background noise level, obtain the hotspot mutation noise level of at least one candidate mutation whose hotspot mutation type is hotspot mutation in the sample to be analyzed;
[0034] Based on the hotspot mutation noise level, the single-sample background noise level statistics of all hotspot mutations in each sample of the sample to be analyzed are calculated.
[0035] If the statistical value of the background noise level of a single sample is greater than or equal to a preset threshold, then the sample noise type of the sample to be analyzed is determined to be a high-noise sample.
[0036] If the statistical value of the background noise level of a single sample is less than the preset threshold, then the sample noise type of the sample to be analyzed is determined to be a non-high noise sample.
[0037] By adopting the above technical solution, the sample noise type can be determined based on the background noise level. The change in background noise level of hot spot mutations in the sample can be used to objectively assess the degree of data interference, thereby timely identifying samples with low data quality. By determining the noise type based on the statistical value of the background noise level and the preset threshold, the influence of abnormal noise can be effectively eliminated, ensuring the accuracy and consistency of mutation interpretation.
[0038] In a preferred embodiment, this application can be further configured such that: determining the sample quality type of the sample to be analyzed based on the sample degradation type and the sample noise type specifically includes:
[0039] If the sample degradation type of the sample to be analyzed is a degraded sample and / or the sample noise type of the sample to be analyzed is a high-noise sample, then the sample quality type of the sample to be analyzed is determined to be a low-quality sample.
[0040] If the sample degradation type of the sample to be analyzed is non-degraded sample and / or the sample noise type of the sample to be analyzed is non-high noise sample, then the sample quality type of the sample to be analyzed is determined to be normal sample.
[0041] By adopting the above technical solution, the sample quality type is determined based on the sample degradation type and sample noise type. This allows for a comprehensive and scientific classification of sample quality across different data quality dimensions. It also enables an objective distinction between high-noise samples and non-high-noise samples, thereby effectively identifying potentially abnormal samples that are severely affected by background interference. This improves the adaptability and robustness of the method under complex sample conditions and avoids misjudgment or missed judgment caused by inferior samples.
[0042] In a preferred embodiment, this application can be further configured such that: generating the interpretation result corresponding to the candidate mutation based on the hotspot mutation type, the sample quality type, and the mutation data specifically includes:
[0043] Based on the hotspot mutation type and the sample quality type of the sample to be analyzed, obtain the corresponding preset dynamic threshold;
[0044] If the mutation data of the candidate mutation meets the preset dynamic threshold standard, the candidate mutation is generated and marked as qualified; otherwise, the candidate mutation is generated and marked as requiring manual review.
[0045] By adopting the above technical solution, and generating candidate mutation interpretation results based on hotspot mutation type, sample quality type, and mutation data, the corresponding preset interpretation threshold standard can be dynamically invoked according to the actual sample and mutation characteristics. This enables priority protection of high-risk hotspot mutations and risk isolation of low-quality samples, reduces the risk of false positives and false negatives, and improves the reliability and clinical applicability of mutation interpretation results.
[0046] In a preferred embodiment, this application can be further configured such that, before the step of obtaining the corresponding preset dynamic threshold based on the hotspot mutation type and the sample quality type of the sample to be analyzed, the tumor gene mutation interpretation method further includes:
[0047] Obtain a historical sample dataset and divide the historical sample dataset into a test set and a validation set according to a preset ratio;
[0048] The test set and the validation set are grouped according to the hotspot mutation type and the sample quality type;
[0049] For each group in the test set, all mutation sites whose manually identified labels are unqualified are selected, and an initial threshold boundary is set based on the mutation abundance, the number of mutation-supporting reads, and the sequencing depth of the site for all mutation sites.
[0050] Based on the initial threshold boundary, a multi-dimensional grid space is constructed and threshold combination candidate regions are discretized. Through gridded search and optimization algorithms, based on the threshold combination candidate regions, a weighted objective function is used to determine the threshold combination with the largest weighted objective function value for the corresponding group in the test set.
[0051] The validation set is used for validation. Mutations in each sample in the validation set that fall within the threshold combination of the corresponding group are recorded as qualified mutations. The consistency index of the qualified mutations in the corresponding group of the validation set is calculated.
[0052] If the consistency index is greater than or equal to the preset consistency threshold, then the preset dynamic threshold of the corresponding group is determined according to the combination of the thresholds.
[0053] By adopting the above technical solutions, and by acquiring historical sample datasets and dividing them into test and validation sets, a reasonable training and validation mechanism can be established based on real historical interpretation data. This ensures that dynamic threshold optimization has sufficient data support and generalization ability. By grouping training based on hotspot mutation types and sample quality types, the interpretation criteria for different categories of samples can be optimized in a targeted manner, thereby improving the interpretation accuracy and stability. By constructing a multi-dimensional grid space and searching for the optimal threshold combination using a weighted objective function, an optimal balance can be achieved between specificity and coverage, thereby maximizing the protection against the risks of missed and false judgments. By validating consistency indicators and screening dynamic thresholds in the validation set, it can be ensured that the finally determined dynamic interpretation criteria also have high stability and high reliability on independent sample sets.
[0054] In a preferred embodiment, this application can be further configured such that, prior to the step of determining whether the candidate mutation exists in the preset list, the tumor gene mutation interpretation method further includes:
[0055] Filter out mutation sites in the historical sample dataset whose manually interpreted labels are unqualified, and calculate the screening frequency and standard deviation of mutation abundance of the selected mutation sites;
[0056] Based on the screening frequency, mutation sites with a screening frequency greater than or equal to a preset frequency threshold are classified into a set of high-frequency unqualified mutation sites.
[0057] In the set of high-frequency unqualified mutation sites, the first site whose mutation abundance standard deviation is less than a preset stability threshold is selected, and the first site selected is manually reviewed and its known clinical significance is confirmed.
[0058] The first locus that is still deemed unqualified by manual review and has no known clinical significance will be included in the blacklist.
[0059] By adopting the above technical solution, and by screening mutation sites in historical sample data that are manually identified as unqualified, and calculating the screening frequency and abundance standard deviation, it is possible to accurately identify false positive mutations that have repeatedly occurred in history and have poor stability and no clinical significance. This enables the scientific collection of candidate sites for the blacklist. After manual review and confirmation, sites that meet the standards are included in the blacklist, which ensures the rigor and authority of the list. This allows for the rapid and accurate removal of high-risk false positive mutations in the interpretation process, thereby improving the overall interpretation efficiency and accuracy.
[0060] In a preferred embodiment, this application can be further configured such that, prior to the step of determining whether the candidate mutation exists in the preset list, the tumor gene mutation interpretation method further includes:
[0061] Mutation sites that do not belong to the whitelist and the blacklist are screened from the mutation sites, and the abundance dispersion index of the screened mutation sites is calculated.
[0062] Mutation sites whose abundance dispersion index is greater than a preset dispersion threshold are included in the gray list.
[0063] By adopting the above technical solution, mutation sites with abundance dispersion indices greater than a preset dispersion threshold can be screened from the remaining mutation sites that are not on the whitelist or blacklist. This allows for the identification of difficult mutation sites with large fluctuations in interpretation parameters and significant discrepancies in review, thereby constructing a gray list. This facilitates prioritizing manual review in subsequent interpretation processes, avoiding misjudgments or omissions caused by fluctuations in parameter boundaries, and thus improving the reliability of the overall interpretation process and the efficiency of human-machine collaborative review.
[0064] In summary, this application includes at least one of the following beneficial technical effects:
[0065] 1. By obtaining the mutation detection results of the sample to be analyzed, the basic data characteristics of each candidate mutation can be accurately grasped, thereby avoiding interpretation bias caused by incomplete mutation information. By determining whether the candidate mutation exists in the preset list, known unreliable or controversial mutations can be eliminated in the early stage of the interpretation process, improving the efficiency and accuracy of the overall screening. By generating interpretation results according to the list type, a preliminary judgment with high credibility or requiring verification can be made quickly, thereby shortening the necessary time for manual review and intervention and improving the overall interpretation speed. After the candidate mutation does not exist in the preset list, further interpretation can be carried out by hot mutation type, sample quality type and mutation data, which can realize dynamic adaptation and personalized subdivision of interpretation standards, thereby ensuring the accuracy and adaptability of interpretation under different sample conditions, and thus improving the standardization level and clinical applicability of tumor gene mutation interpretation.
[0066] 2. By screening mutation sites in historical sample data that are manually identified as unqualified and calculating the screening frequency and abundance standard deviation, we can accurately identify false positive mutations that have repeatedly occurred in history and have poor stability and no clinical significance. This enables the scientific collection of candidate sites for the blacklist. After manual review and confirmation, sites that meet the standards are included in the blacklist, which ensures the rigor and authority of the list. This allows for the rapid and accurate removal of high-risk false positive mutations in the interpretation process, improving the overall interpretation efficiency and accuracy.
[0067] 3. By screening mutation sites whose abundance dispersion index is greater than the preset dispersion threshold from the remaining mutation sites that are not on the whitelist or blacklist, it is possible to identify difficult mutation sites with large fluctuations in interpretation parameters and large discrepancies in review, thereby constructing a gray list. This facilitates prioritizing manual review in subsequent interpretation processes, avoiding misjudgments or omissions caused by fluctuations in parameter boundaries, and thus improving the reliability of the overall interpretation process and the efficiency of human-machine collaborative review. Attached Figure Description
[0068] Figure 1 This is a flowchart illustrating the implementation of a tumor gene mutation interpretation method based on an NGS platform in one embodiment of this application.
[0069] Figure 2 This is another implementation flowchart of a tumor gene mutation interpretation method based on an NGS platform in one embodiment of this application;
[0070] Figure 3 This is another implementation flowchart of a tumor gene mutation interpretation method based on an NGS platform in one embodiment of this application;
[0071] Figure 4 This is a flowchart illustrating another implementation of a tumor gene mutation interpretation method based on an NGS platform in one embodiment of this application. Detailed Implementation
[0072] The present application will be further described in detail below with reference to the accompanying drawings.
[0073] In one embodiment, such as Figure 1 As shown, this application discloses a method for identifying tumor gene mutations based on an NGS platform, which specifically includes the following steps:
[0074] S10: Obtain the mutation detection results of the sample to be analyzed. The mutation detection results include at least one candidate mutation and the corresponding mutation data. The mutation data includes mutation type, mutation abundance, number of supporting reads, site sequencing depth, and manually interpreted tags.
[0075] Specifically, the mutation detection results of the samples to be analyzed are obtained through an NGS platform. For each detected candidate mutation, its detailed information is recorded, including mutation type (e.g., SNP, InDel), mutation abundance (AF), the proportion of mutant reads to total reads, number of supporting reads (AD), the number of sequencing reads supporting the mutation, site sequencing depth (DP), and manual review labels (e.g., Pass, Fail). Mutation type refers to whether the mutation is a single nucleotide variant (SNV), insertion / deletion variant (InDel), etc.; mutation abundance (AF) refers to the proportion of the mutation in all detected genomes; number of supporting reads (AD) refers to the number of sequencing reads supporting the mutation; site sequencing depth (DP) refers to the sequencing depth covering the site; and manual review labels are determined manually to indicate whether the mutation meets clinical standards, ensuring the reliability of data quality.
[0076] S20: Based on the mutation detection results, determine whether the candidate mutation exists in the preset list, which includes at least the blacklist and the graylist.
[0077] Specifically, the pre-defined lists include a blacklist and a graylist. The blacklist contains known high-frequency false-positive mutation sites, which are easily misidentified as mutation sites in NGS sequencing due to various technical reasons such as sequencing errors and alignment biases. The graylist contains complex mutation sites that are difficult to automatically interpret and require manual review. These sites may have low mutation abundance, complex variation patterns, or be located in complex regions of the genome. For each candidate mutation, based on the mutation data of the candidate mutation, such as mutation type, mutation abundance (AF), and supporting read count (AD), the mutation is quickly matched to see if it meets the conditions of the blacklist or graylist. That is, whether the mutation data of the candidate mutation meets the data conditions for matching the blacklist or graylist, or whether the coordinates of the candidate mutation match any site in the blacklist, then the mutation is considered likely to be a false positive. If the coordinates of the candidate mutation match any site in the graylist, then the mutation is considered to require further manual review, and a definitive interpretation cannot be given.
[0078] S30: If the candidate mutation exists in the preset list, then generate the interpretation result corresponding to the candidate mutation according to the list type of the preset list.
[0079] Specifically, if a candidate mutation is identified as existing in a preset list, a preliminary interpretation result is generated based on the list type in which the candidate mutation exists. If the candidate mutation exists in a blacklist, an interpretation result for that candidate mutation is generated and marked as "Fail," indicating that the mutation is considered a false positive and does not require further analysis. If the candidate mutation exists in a gray list, an interpretation result for that candidate mutation is generated and marked as "requires manual review," indicating that the mutation requires further judgment by human experts and a clear "Pass" or "Fail" conclusion cannot be given.
[0080] S40: If the candidate mutation does not exist in the preset list, then based on the mutation detection results, determine the hot spot mutation type and sample quality type of the candidate mutations in the sample to be analyzed.
[0081] Specifically, if a candidate mutation is not on the blacklist or graylist, its nature and sample quality need further analysis to determine how to interpret it. First, the pre-defined hotspot mutation whitelist is consulted to determine whether the candidate mutation is a hotspot mutation. Hotspot mutations refer to mutations that occur frequently in specific genes and have significant clinical importance, such as the L858R mutation commonly found in the EGFR gene or the G12C mutation commonly found in the KRAS gene. Then, the quality of the sample to be analyzed is evaluated, including whether the sample has been degraded and the background noise level of the sample. Is the sample a high-noise sample? Sample quality significantly affects the accuracy of mutation detection, so it needs to be considered to obtain the hotspot mutation type and sample quality type of the candidate mutation in the sample to be analyzed.
[0082] S50: Based on hotspot mutation types, sample quality types, and mutation data, generate interpretation results corresponding to candidate mutations.
[0083] Specifically, after identifying the hotspot mutation types and sample quality types, a corresponding dynamic threshold is selected based on these two factors. Each type of sample and mutation combination has a corresponding dynamic threshold, which is set based on historical data from the training set and results from the validation set. For example, a more lenient threshold might be used for hotspot mutations in high-quality samples, while a stricter threshold is used for non-hotspot mutations in low-quality samples. The mutation data of candidate mutations, such as mutation abundance (AF), supporting read count (AD), and site sequencing depth (DP), are compared with the selected dynamic threshold. If the threshold requirements are met, an interpretation result for the candidate mutation is generated and marked as "Pass." Otherwise, an interpretation result for the candidate mutation is generated and marked as requiring manual review, thus obtaining the interpretation result corresponding to the candidate mutation. The interpretation result can also be output to the mutation detection results of the sample, and uninterpreted results are further reviewed and confirmed manually.
[0084] In one embodiment, step S30, namely generating the interpretation result corresponding to the candidate mutation according to the list type of the preset list, specifically includes:
[0085] S31: If a candidate mutation exists in the blacklist, the interpretation result of the candidate mutation is generated and marked as unqualified.
[0086] Specifically, if a candidate mutation is identified as existing in the blacklist, it means that the candidate mutation is considered a high-probability false positive result. Since mutations in the blacklist are not clinically significant, an interpretation result for the candidate mutation is generated and marked as unqualified (Fail). This avoids unnecessary follow-up analysis of these known false positive mutations, improves efficiency, and reduces false alarms.
[0087] S32: If a candidate mutation exists in the gray list, the interpretation result of the candidate mutation is generated and marked as requiring manual review.
[0088] Specifically, if a candidate mutation is identified as existing in the gray list, it indicates that the mutation is complex or difficult to automatically identify. These mutations may have low accuracy in automatic interpretation due to various reasons such as low frequency or complex genomic regions. Therefore, the interpretation result of the candidate mutation is generated and marked as requiring manual review. Then, these candidate mutations marked as requiring manual review are submitted to professional clinical geneticists or pathologists for manual review to ensure the accuracy and reliability of the interpretation.
[0089] In one embodiment, step S40, namely, determining the hotspot mutation type and sample quality type of the candidate mutations in the sample to be analyzed based on the mutation detection results, specifically includes:
[0090] S41: Determine whether the candidate mutation of the sample to be analyzed exists in the preset whitelist, which is constructed based on known hotspot mutations in public clinical databases.
[0091] Specifically, by querying a pre-established whitelist L whitelist The whitelist L whitelist It includes information on known hotspot mutations compiled from public clinical databases such as COSMIC and OncoKB. Hotspot mutations are mutations that occur frequently in specific genes and are closely related to the occurrence, development, or treatment response of tumors. For example, the p.L858R variant of the EGFR gene and the p.G12D variant of the KRAS gene are included in the whitelist. If a candidate mutation of the sample to be analyzed is found in the whitelist, it indicates that the mutation is a known and highly reliable hotspot mutation and is thus identified as a match in the whitelist; otherwise, it is considered a miss.
[0092] S42: If the candidate mutation exists in the whitelist, then the hotspot mutation type of the candidate mutation is determined to be a hotspot mutation.
[0093] Specifically, if a candidate mutation is found in the whitelist of samples to be analyzed, the hotspot mutation type of the candidate mutation is marked as a hotspot mutation. Hotspot mutations refer to gene variations that play a key role in the occurrence and development of tumors and have been validated by multiple clinical trials. For example, the TP53 gene p.R273H mutation is widely present in a variety of solid tumors and has important guiding significance for treatment strategies. Therefore, when a mutation such as TP53 p.R273H appears in the sample and matches the whitelist, it is immediately identified as a hotspot mutation. There is no need to further rely on mutation abundance (AF), supporting read count (AD), etc. to determine whether it is a hotspot, thus ensuring the priority protection and identification of clinically important mutations.
[0094] S43: If the candidate mutation does not exist in the whitelist, then the hotspot mutation type is determined to be a non-hotspot mutation.
[0095] Specifically, if a candidate mutation fails to match the whitelist, the hotspot mutation type of the candidate mutation is marked as a non-hotspot mutation. Non-hotspot mutations refer to variants for which there is currently insufficient clinical evidence to support their association with tumor occurrence, development, or treatment. For example, some newly discovered rare mutations, such as some non-pathogenic mutations of the BRCA2 gene, are classified as non-hotspot mutations even if they are detected due to a lack of extensive clinical research validation. This is to allow for the application of more stringent or cautious screening criteria for non-hotspot mutations in the subsequent interpretation process, thereby reducing the misjudgment rate.
[0096] S44: Obtain the mutation spectrum characteristics and background noise level of the sample to be analyzed from the mutation detection results, and determine the sample degradation type of the sample to be analyzed based on the mutation spectrum characteristics.
[0097] Specifically, in addition to hotspot mutation types, the quality of the sample itself also affects the accuracy of mutation detection. Two key sample quality indicators are extracted from the mutation detection results: mutation spectrum characteristics and background noise level. Mutation spectrum characteristics refer to the distribution of various mutation types in the sample. For example, degraded samples usually have a high proportion of C>T (after correction of positive and negative strands in the reference genome) mutations. Background noise level refers to low-frequency error signals that also appear in normal tissues. High background noise levels may lead to more false positive results. By analyzing these two indicators, sample quality can be assessed more accurately, and samples can be classified into different quality types so that corresponding thresholds or standards can be applied in subsequent interpretations. Then, based on the extracted mutation spectrum characteristics, a preset algorithm or model is used to determine whether the sample has degraded. For example, if the proportion of C>T mutations exceeds a certain threshold, the sample is considered to have degraded, thus obtaining the sample degradation type of the sample to be analyzed.
[0098] S45: Determine the sample noise type of the sample to be analyzed based on the background noise level.
[0099] Specifically, the extracted background noise level data is used to further determine the noise type of the sample. Typically, the noise level of the sample is compared with a preset threshold. If the noise level of the sample is higher than the threshold, the sample is considered to be a high-noise sample, and vice versa. High-noise samples may require stricter filtering or a higher mutation abundance threshold to avoid false positives, thereby obtaining the sample noise type of the sample to be analyzed.
[0100] S46: Determine the sample quality type of the sample to be analyzed based on the sample degradation type and sample noise type.
[0101] Specifically, after comprehensively considering the degradation type and noise type of the sample, the sample will be classified into an overall quality type. For example, if the degradation type of the sample to be analyzed is a degraded sample or the noise type is a high-noise sample, and either of these meets the low quality type judgment criteria, then the sample to be analyzed will be classified as a low quality sample; otherwise, it will be classified as a normal sample, thus obtaining the sample quality type of the sample to be analyzed.
[0102] In one embodiment, step S44, namely determining the sample degradation type of the sample to be analyzed based on the mutation spectrum characteristics, specifically includes:
[0103] S441: Count the total number of mutations of type SNV in the mutation data of the sample to be analyzed.
[0104] Specifically, the total number N of mutations of the single nucleotide variant (SNV) type in the sample to be analyzed is counted. SNV SNVs refer to single nucleotide changes in a DNA sequence, such as changing from A to G or from C to T. By analyzing the mutation detection results of a sample, all candidate mutations marked as SNVs in the mutation type field are screened, and these mutations are counted. The total number is used to calculate degradation characteristics. For example, in an NGS test report, a total of 420 mutation records were detected, of which 360 were confirmed as SNVs. Therefore, the total number of SNV mutations in that sample is 360, which is used to calculate the degradation percentage of the sample. For example, N... SNV It can be represented as In the formula, N total Let N be the total number of mutations in the sample, and {condition} be an indicator function that takes the value 1 if the condition is met and 0 otherwise. SNV This represents the total number of SNV-type mutations in the sample.
[0105] S442: From the total number of mutations, obtain the number of mutations with the mutation form C>T and whose mutation abundance in the mutation data is less than a preset abundance threshold.
[0106] Specifically, within the total number of SNV mutations, a specific subset of SNVs is further screened. Two screening criteria are used: the mutation form is C>T (i.e., the base C in the DNA sequence is changed to T); and the mutation abundance (AF) is less than a preset abundance threshold, such as 3%. Mutation abundance refers to the proportion of sequencing reads carrying mutated bases to the total reads; here, a low threshold, such as 3%, is used to focus on low-frequency mutations, thereby counting the number N of C>T mutations that meet these two conditions. C→T,AF<3% Low-frequency C>T mutations tend to accumulate in degraded DNA samples, so counting their number helps determine sample quality. For example, among the 360 SNV mutations mentioned above, 90 are C>T mutations, of which 72 have an abundance of less than 3%. Therefore, the number of low-abundance C>T mutations is recorded as 72 for subsequent degradation percentage calculations. For instance, N C→T,AF<3% It can be represented as In the formula, N total Let N be the total number of mutations in the sample, and {condition} be an indicator function that takes the value 1 if the condition is met and 0 otherwise. C→T,AF<3% The total number of C>T low-frequency mutations that meet the criteria.
[0107] S443: Based on the total number of mutations and the number of mutations, the percentage of sample degradation is calculated.
[0108] Specifically, after obtaining the total number of SNV mutations N SNV The number N of low abundance mutations of C>T C→T,AF<3% Then, the number N of low-frequency C>T mutations obtained by statistics will be... C→T,AF<3% Divide by the total number of mutations N obtained statistically. SNV The percentage of sample degradation, P, was obtained. C→T,AF<3% This proportion represents the percentage of low-frequency C>T mutations among all SNV mutations, reflecting the degree of sample degradation. For example, if the number of low-abundance C>T mutations is 72 and the total number of SNV mutations is 360, then the degradation percentage is 72 / 360 = 0.2, or 20%, used to subsequently determine whether a sample is a degraded sample. For instance, P... C→T,AF<3% It can be represented as In the formula, P C→T,AF<3% The percentage of low-frequency C>T mutations that meet the criteria.
[0109] S444: If the proportion of sample degradation is greater than or equal to the degradation judgment threshold, then the sample to be analyzed is determined to be a degraded sample.
[0110] Specifically, after calculating the percentage of sample degradation, it is compared with a preset degradation judgment threshold, such as 65%. If the percentage of sample degradation is greater than or equal to the degradation judgment threshold, the sample is considered to have undergone significant degradation, and the sample degradation type is determined to be a degraded sample. For example, when the degradation percentage of a sample is 72%, which is greater than the threshold of 65%, the sample is determined to be a degraded sample, and the dynamic threshold standard set for degraded samples is used in the subsequent mutation interpretation process.
[0111] S445: If the percentage of sample degradation is less than the degradation judgment threshold, then the sample to be analyzed is determined to be a non-degraded sample.
[0112] Specifically, if the degradation percentage of a sample is less than the preset degradation judgment threshold, the sample is considered to have a relatively mild degradation degree, and the sample degradation type is determined to be a non-degraded sample. For example, if the degradation percentage of a sample is 17%, which is lower than the judgment threshold of 65%, then the sample is determined to be a non-degraded sample.
[0113] In one embodiment, step S45, namely determining the sample noise type of the sample to be analyzed based on the background noise level, specifically includes:
[0114] S451: Based on the background noise level, obtain the hotspot mutation noise level of at least one candidate mutation in the sample to be analyzed that is a hotspot mutation.
[0115] Specifically, background noise level refers to the low-frequency spurious mutation signals in normal tissues without tumor cells, caused by sequencing errors or other technical reasons. For the sample to be analyzed, it is first necessary to determine which candidate mutations are hotspot mutations. Hotspot mutations are usually identified based on a pre-defined whitelist. Then, the noise level of these hotspot mutations is extracted from the sequencing data. If multiple hotspot mutation sites exist in the sample, the background noise value of each site is extracted separately for subsequent averaging. For example, if the sample contains two hotspot mutation sites, EGFR p.L858R and KRAS p.G12D, which are on the whitelist, the background noise levels of these two mutation sites need to be calculated, with background noise levels of 0.4% and 0.5%, respectively. Therefore, by matching the sample to be analyzed with the whitelist, identifying the mutation sites in the sample, and then extracting the background noise level of each mutation site, the analysis can be performed.
[0116] S452: Based on the hotspot mutation noise level, calculate the statistical value of the single-sample background noise level of all hotspot mutations in each sample of the sample to be analyzed.
[0117] Specifically, after extracting the noise level of each hotspot mutation site, the background noise values of all hotspot sites within the same sample are sorted, and the 95th percentile is determined based on the sorting results as the statistical value of the background noise level of a single sample. The 95th percentile refers to the value in which 95% of the data are less than or equal to the value, and 5% of the data are greater than the value. It is used to exclude the interference of outliers and more accurately reflect the actual background noise of the sample. For example, if the hotspot abrupt noise levels in a sample are 0.2%, 0.3%, 0.4%, 0.6%, and 0.7%, then after arranging them in order of magnitude and taking the 95th percentile, the calculated single-sample background noise level is 0.7%. For example, the statistical value of the single-sample background noise level... It can be represented as Where M is the total number of mutations in the hotspot mutation whitelist, and x i,j Let represent the hotspot mutation noise level of the i-th hotspot mutation in the j-th sample. is the 95th percentile value of the noise level of all hotspot mutations in the j-th sample.
[0118] S453: If the statistical value of the background noise level of a single sample is greater than or equal to the preset threshold, then the sample noise type of the sample to be analyzed is determined to be a high-noise sample.
[0119] Specifically, to distinguish whether the noise level of a sample is normal or abnormally high, a preset threshold needs to be set. This preset threshold is usually calculated based on noise level data from a large number of normal clinical samples using the 95th percentile method. Specifically, the calculated statistical value of the background noise level of a single sample is compared with this preset threshold. If the statistical value of the noise level of the sample to be analyzed is greater than or equal to this preset threshold, the sample is considered to have a significantly higher noise level than normal, belonging to a high-noise sample. High-noise samples mean that their sequencing results may contain more false-positive mutations, requiring stricter quality control and subsequent manual review. For example, the preset threshold T can be expressed as... Where N is the number of clinical samples, N≥1000. For example, if the background noise level of a sample reaches 0.6% and the preset threshold is 0.5%, the sample is marked as a high-noise sample, and subsequent interpretations will use stricter parameter screening or prompt manual review.
[0120] S454: If the statistical value of the background noise level of a single sample is less than the preset threshold, then the sample noise type of the sample to be analyzed is determined to be a non-high noise sample.
[0121] Specifically, if the background noise level of a single sample of the sample to be analyzed is less than a preset threshold, the noise level of the sample is considered to be within the normal range, and the sample noise type is determined to be a non-high noise sample. A non-high noise sample means that the background signal interference is within an acceptable range, and abrupt changes can be judged using conventional dynamic thresholds without additional processing. For example, if the background noise level of a single sample is only 0.3%, which is significantly lower than the statistical value of the sample background noise level of 0.5%, it is classified as a non-high noise sample and directly participates in the subsequent conventional judgment process.
[0122] In one embodiment, step S46, which determines the sample quality type of the sample to be analyzed based on the sample degradation type and sample noise type, specifically includes:
[0123] S461: If the sample degradation type of the sample to be analyzed is a degraded sample and / or the sample noise type of the sample to be analyzed is a high-noise sample, then the sample quality type of the sample to be analyzed is determined to be a low-quality sample.
[0124] Specifically, after determining the degradation type and noise type of the sample to be analyzed, if the degradation type is determined to be a degradation sample or the noise type is determined to be a high-noise sample, the sample to be analyzed is directly marked as a low-quality sample. Degradation samples usually lead to a decrease in sequencing accuracy due to severe DNA fragmentation, while high-noise samples indicate that the background interference level is too high, which is not conducive to the true identification of mutations. For example, if the degradation rate of a sample to be analyzed is 75%, which is greater than the 65% threshold, and the 95th percentile of the background noise of hotspot mutations is 0.55%, which is greater than the 0.5% noise threshold, then the sample exceeds the standard in both indicators and is therefore judged as a low-quality sample. In the subsequent mutation interpretation process, a more lenient interpretation standard should be used or resampling should be recommended.
[0125] S462: If the sample degradation type of the sample to be analyzed is non-degraded sample and / or the sample noise type of the sample to be analyzed is non-high noise sample, then the sample quality type of the sample to be analyzed is determined to be normal sample.
[0126] Specifically, after determining the degradation type and noise type of the sample to be analyzed, if the degradation type is determined to be a non-degradation sample and the noise type is determined to be a non-high noise sample, the sample to be analyzed is marked as a normal sample. Normal samples have good DNA integrity and low background noise levels, making them suitable for applying standard dynamic thresholds for high-confidence mutation interpretation. For example, if the degradation rate of a sample to be analyzed is only 15%, which is lower than the set standard of 65%, and the 95th percentile of the hot spot mutation background noise is 0.3%, which is far lower than the noise threshold of 0.5%, then the system determines that the sample is a normal sample. Mutations can be strictly screened according to standard parameters in subsequent steps to ensure the reliability of the results.
[0127] In one embodiment, step S50, which generates the interpretation result corresponding to the candidate mutation based on the hotspot mutation type, sample quality type, and mutation data, specifically includes:
[0128] S51: Based on the hotspot mutation type and sample quality type of the sample to be analyzed, obtain the corresponding preset dynamic threshold.
[0129] Specifically, when dynamically acquiring thresholds based on the hotspot mutation type and sample quality type of the sample to be analyzed, the first step is to select a corresponding set from four pre-trained and optimized dynamic interpretation threshold combinations, depending on whether the mutation belongs to a hotspot mutation and whether the sample is a normal sample. The dynamic threshold combination includes the minimum standard values of mutation abundance (AF), supported read count (AD), and site sequencing depth (DP). Each set of thresholds is determined by weighted optimization based on specificity and coverage according to historical sample training data. For example, if the sample is a normal sample and the mutation is a hotspot mutation, the threshold set corresponding to the normal sample-hotspot mutation group with AF≥1%, AD≥30, and DP≥300 is called to ensure that accurate interpretation criteria are set for different types of samples and mutation characteristics.
[0130] S52: Based on whether the mutation data of the candidate mutation meets the preset dynamic threshold standard, if it does, the interpretation result of the candidate mutation is generated and marked as qualified; if it does not meet the standard, the interpretation result of the candidate mutation is generated and marked as requiring manual review.
[0131] Specifically, after obtaining the corresponding dynamic threshold, the candidate mutation data of the sample to be analyzed are compared with the dynamic threshold standard item by item. If the mutation abundance (AF), the number of supported reads (AD), and the site sequencing depth (DP) all meet the corresponding standard, the interpretation result of the candidate mutation is generated and marked as qualified (Pass). Conversely, if any indicator does not meet the threshold requirement, the interpretation result of the candidate mutation is generated and marked as requiring manual review. For example, for a mutation, if its mutation abundance (AF) is 1.2%, the number of supported reads (AD) is 35, and the site sequencing depth (DP) is 420, all of which meet the threshold standard corresponding to the normal sample-hotspot mutation category, the candidate mutation is interpreted as qualified (Pass). Otherwise, if the number of supported reads (AD) is only 25, which is lower than the standard of 30, it is marked as requiring manual review and pushed to the review process.
[0132] In one embodiment, such as Figure 2 As shown, prior to step S50, a tumor gene mutation interpretation method based on an NGS platform further includes:
[0133] S501: Obtain the historical sample dataset and divide it into a test set and a validation set according to a preset ratio.
[0134] Specifically, when acquiring historical sample datasets, large-scale NGS detection data is collected. Each historical sample data includes mutation abundance (AF), number of supported reads (AD), sequencing depth (DP), and corresponding manual interpretation labels. To ensure the objectivity and generalization ability of the parameter optimization process, the historical data is divided into a test set and a validation set at a fixed ratio of 4:1. The test set is used for training and optimization of dynamic threshold combinations, while the validation set is used to evaluate the consistency index of the obtained threshold combinations. For example, after collecting 10,000 historical mutation data, 8,000 are divided into a test set and 2,000 into a validation set to provide a data foundation for subsequent search for optimal interpretation parameters.
[0135] S502: Group the test set and validation set according to hotspot mutation type and sample quality type.
[0136] Specifically, after the dataset is divided, the test set and validation set are grouped according to the hotspot mutation type (hotspot mutation or non-hotspot mutation) and the sample quality type (normal sample or low-quality sample) for each mutation, forming four combinations: normal sample-hotspot mutation, normal sample-non-hotspot mutation, low-quality sample-hotspot mutation, and low-quality sample-non-hotspot mutation. The thresholds within each combination are optimized independently to avoid interpretation bias caused by mixing mutations of different quality or clinical significance. For example, if a mutation in a test set comes from a low-quality sample and is a non-hotspot mutation, it is classified into the low-quality sample-non-hotspot mutation group to ensure the targeting and accuracy of parameter training.
[0137] S503: For each group in the test set, select all mutation sites whose manually interpreted labels are unqualified, and set an initial threshold boundary based on the mutation abundance, number of supported reads, and sequencing depth of all mutation sites.
[0138] Specifically, within each group, all mutation sites manually labeled as "Fail" are selected. The 95th percentiles of these Fail mutations are calculated across three dimensions: mutation abundance (AF), supported reads number (AD), and sequencing depth (DP). These initial boundaries serve as preliminary search boundaries, limiting the subsequent dynamic threshold search range and preventing an excessively large search space that leads to inefficient training. For example, in the normal sample-hotspot mutation group, if the 95th percentile of Fail mutations is 1%, the 95th percentile of AD is 30, and the 95th percentile of DP is 300, these values are used as the starting boundaries for the dynamic threshold search, providing an effective range for subsequent optimization. For instance, the initial threshold boundary T... init It can be represented as In the formula, T init As the initial threshold boundary, R fail This is a set of unqualified Fail mutations.
[0139] S504: Based on the initial threshold boundary, a multi-dimensional grid space is constructed and threshold combination candidate regions are discretized. Through gridded search and optimization algorithms, a weighted objective function is used to determine the threshold combination with the largest weighted objective function value for the corresponding group in the test set.
[0140] Specifically, after determining the initial boundary, a three-dimensional grid space G is constructed based on three parameter axes: mutation abundance (AF), number of supporting reads (AD), and site sequencing depth (DP). i,j,k The boundaries of each dimension are discretized into several small intervals, forming a large number of candidate threshold combinations. For each candidate threshold combination, specificity and coverage are calculated in the test set. Specificity is the proportion of mutations that are actually manually interpreted as qualified passes, and coverage is the proportion of mutations that are correctly interpreted as qualified passes. A comprehensive evaluation is then performed according to a preset weighted objective function F = λ × Specificity + (1 - λ) × Coverage, where λ is a weighting factor, for example, 0.7 or 0.8. Finally, the threshold combination with the highest objective function F value is selected as the dynamic threshold T for the corresponding group. final For example, in the normal sample-hotspot mutation grouping, after grid search, the optimal combination is selected as AF≥1%, AD≥35, DP≥320, ensuring optimal overall interpretation performance. For instance, T... final It can be represented as ,in and , In the formula, For grid G i,j,k The number of mutations that are manually identified as passing passes. For grid G i,j,k The number of mutations manually identified as "Fail" is denoted as T. Specificity represents the proportion of mutations manually identified as "Pass" that were actually manually identified as "Pass," indicating how many "Fail" mutations were successfully excluded from the confidence region. Coverage represents the proportion of "Pass" mutations correctly identified as "Pass," indicating how many "Pass" mutations are contained within the confidence region. (α, β, γ) are combination variables representing the threshold values in three-dimensional space. α is the lower limit of mutation abundance (AF), β is the lower limit of support read count (AD), and γ is the lower limit of site sequencing depth (DP). Therefore, through various combinations of (α, β, γ), the combination that maximizes this weighted index value is selected as the final threshold T. final .
[0141] S505: Use the validation set for validation. Record the mutations in each sample in the validation set that fall within the threshold combination of the corresponding group as qualified mutations, and calculate the consistency index of qualified mutations in the corresponding group of the validation set.
[0142] Specifically, in determining T for each group in the test set final After combining, use the validation set on T final The validation was performed by combining the results, substituting each candidate mutation in the validation set into the corresponding group's T. final Threshold combinations are used for interpretation. If the mutation abundance (AF), supporting reads (AD), and sequencing depth (DP) of the candidate mutation are all greater than or equal to T, then... final If the mutation matches the corresponding standard, it is considered a pass; otherwise, it is considered a fail requiring manual review. The number of mutations judged as pass that also meet the manual review label is counted, and a consistency index P is calculated. The consistency index P is defined as the number of mutations in the validation set that are both judged as pass and manually reviewed as pass, divided by the total number of mutations in the validation set that are manually reviewed as pass. For example, if there are 1000 mutations in the validation set that are manually reviewed as pass, T... final After combined interpretation, if 950 of these are correctly identified as qualified Pass mutations, then the consistency index is 950 / 1000 = 95%. For example, the consistency index P can be expressed as... .
[0143] S506: If the consistency index is greater than or equal to the preset consistency threshold, then the preset dynamic threshold of the corresponding group is determined according to the threshold combination.
[0144] Specifically, after calculating the consistency index P for each group in the validation set, the consistency index is compared with a preset consistency threshold. If the consistency index P is greater than or equal to the set consistency standard value, such as 99%, then the selected T in the test set is confirmed. final The combination serves as the formal dynamic threshold for this grouping; otherwise, the threshold search range needs to be readjusted or the optimization strategy retrained. For example, if the normal sample-hotspot mutation grouping is applied to the validation set using T... final The combined validation consistency is 99.5%, higher than the set standard of 99%. Therefore, the dynamic threshold combination AF≥1%, AD≥35, and DP≥320 determined in the test set will be used as the formal dynamic interpretation standard for this category of mutations to ensure the interpretation stability and high reliability of subsequent new samples. This dynamic threshold standard will be directly applied when automatically interpreting new samples. For example, the preset dynamic threshold T for each group... final It can be represented as .
[0145] In one embodiment, such as Figure 3As shown, prior to step S20, a tumor gene mutation interpretation method based on the NGS platform further includes:
[0146] S2010: Filter mutation sites in the historical sample dataset that are manually identified as unqualified, and calculate the screening frequency and standard deviation of mutation abundance of the selected mutation sites.
[0147] Specifically, from the historical sample dataset, all mutation sites that were manually identified as "Fail" were selected. For each selected "Fail" mutation site, its frequency of occurrence in all samples was calculated as the selection frequency f(l), and the standard deviation σ of the mutation abundance AF of that site in different samples was also calculated. AF (l) represents the proportion of samples in which a particular locus is judged as a "Fail," and the mutation abundance standard deviation represents the degree of fluctuation in the mutation abundance of that locus across different samples. For example, in historical data, if a locus exhibits 120 "Fail" mutations in 5000 samples, and its abundance standard deviation is 0.08, then its screening frequency is recorded as 2.4%, and its abundance standard deviation is 0.08, providing basic data for subsequent list selection. For example, the screening frequency f(l) can be expressed as... S is the clinical sample set, S≥1000, and the standard deviation σ AF (l) can be expressed as AF i Let AF be the mutation abundance of the i-th sample, so that the screening frequency of mutation sites and the standard deviation of mutation abundance can be obtained.
[0148] S2011: Based on the screening frequency, mutation sites with a screening frequency greater than or equal to a preset frequency threshold are classified into a set of high-frequency unqualified mutation sites.
[0149] Specifically, a preset frequency threshold is set, for example, 50%. The calculated screening frequency f(l) is compared with this threshold. Mutation sites with a screening frequency f(l) greater than or equal to this threshold are classified into a high-frequency unqualified Fail mutation site set L. high-fail-freq This means that these sites were judged as "Fail" in more than half of the samples, and they are likely to be systematic false positives, thus further narrowing down the scope of analysis. For example, if a mutation site has a screening frequency of 55%, which is higher than the preset frequency threshold of 50%, it is included in the high-frequency "Fail" mutation site set for further processing in the subsequent blacklist screening process. For example, the high-frequency "Fail" mutation site set L... high-fail-freq It can be represented as L high-fail-freq ={l∈L | f(l)≥0.5}, where L is the set of all mutation sites.
[0150] S2012: In the set of high-frequency unqualified mutation sites, the first site with a mutation abundance standard deviation less than the preset stability threshold is selected, and the selected first site is manually reviewed and its known clinical significance is confirmed.
[0151] Specifically, in the obtained set L of high-frequency unqualified fail mutation sites high-fail-freq In the process, the standard deviation of mutation abundance σ was further screened. AF (l) Sites that are less than a preset stability threshold α, for example, 2%, constitute the site set L. similar These sites not only occur frequently, but their mutation abundance is also relatively stable across different samples, further increasing the likelihood of false positives. These sites are used as primary sites for manual verification. This involves professional pathologists or genomics experts carefully examining the sequencing data of these sites using tools such as IGV to confirm whether they are truly false positives. Simultaneously, relevant databases and literature are consulted to confirm whether these sites have known clinical significance. If these sites are genuine mutations, they may play an important role in certain tumor types. This avoids mistakenly adding genuine, clinically significant mutations to the blacklist. For example, if the standard deviation of a mutation site's abundance is 0.06 and there are no known clinically pathogenic annotations, then it meets the criteria and proceeds to the next step of blacklist screening. For instance, the site set L... similar It can be represented as L similar ={l∈L high-fail-freq |σ AF (l)<α}.
[0152] S2013: First loci that are still deemed unqualified by manual review and have no known clinical significance will be added to the blacklist.
[0153] Specifically, for the first locus that has undergone manual review, if the manual review confirms that it is still a false positive (i.e., a failure), and after reviewing the literature and databases, it is found that it has no known clinical significance, then these loci will be added to the blacklist. blacklist Sites in the blacklist will be directly identified as "Fail" in subsequent automatic interpretations without further analysis. For example, if a high-frequency, low-dispersion mutation site has no clinical significance after review and consistently remains labeled as "Fail" by manual interpretation, then that site will be added to the blacklist. This will allow for immediate marking as "Fail" upon detection during subsequent interpretations, improving the overall accuracy and efficiency of the interpretation process. The blacklist L... blacklist It can be represented as L blacklist ={l∈L similar | Manual review (l) = fail ∧ ClinSig (l) = 0}, ClinSig (l) is the clinical significance of site l, represented by a Boolean value, 0 = no significance, 1 = significant.
[0154] More specifically, after the initial blacklist set is constructed, the performance of mutation sites that hit the blacklist in new samples is continuously monitored during subsequent interpretation. If a blacklist site is repeatedly manually reviewed and confirmed as qualified (Pass) in newly collected large sample data, and its mutation abundance (AF) is stable with low background noise, indicating that the site may have certain clinical significance due to sample size expansion or changes in biological background, then the site is automatically added to the list to be reviewed. After manual confirmation and literature database search, if it is indeed clinically relevant or passes the review, the site is removed from the blacklist. Conversely, if a site continues to maintain high-frequency unqualified (Fail) characteristics in new samples and its abundance stability does not improve significantly, its weight index or additional annotation in the blacklist is updated. At the same time, if a new mutation site that meets the criteria of high frequency, low dispersion, and no clinical significance is found in new samples, the screening frequency and standard deviation are recalculated according to the initial screening logic. After meeting the conditions, the blacklist is automatically expanded to achieve dynamic maintenance and intelligent iteration of the blacklist.
[0155] In one embodiment, such as Figure 4 As shown, prior to step S20, a tumor gene mutation interpretation method based on the NGS platform further includes:
[0156] S2020: Screen out mutation sites that do not belong to the whitelist and blacklist from the mutation sites, and calculate the abundance dispersion index of the screened mutation sites.
[0157] Specifically, from all detected mutation sites, those belonging to the whitelist (known, clinically significant hotspot mutations) and the blacklist (known false-positive mutations) are excluded. For the remaining mutation sites L... candidate The abundance dispersion index CV(l) is calculated for these sites. The abundance dispersion index measures the degree of variation in the mutation abundance of these sites across different samples. In this example, the coefficient of variation CV, i.e., the standard deviation divided by the mean, is used as the abundance dispersion index. The larger the CV value, the greater the variation in mutation abundance across different samples. These sites may be difficult to interpret automatically and require manual review. For example, if the abundance value of a mutation site varies drastically across different samples, with a standard deviation reaching 0.45, it indicates high dispersion and meets the gray list screening criteria. Where L... candidate It can be represented as L candidate ={ l∈L |l∉L blacklist ∧l∉L whitelist ∧ f(l)>0}.
[0158] S2021: Mutation sites with abundance dispersion indices greater than a preset dispersion threshold will be included in the gray list.
[0159] Specifically, a preset dispersion threshold is set, for example, 0.3. The calculated abundance dispersion index is compared with this threshold CV(l). Mutation sites with abundance dispersion indices greater than this threshold are included in the gray list L. graylist Sites in the gray list will be marked for manual review in subsequent automatic interpretations, forcing manual review. For example, if the abundance standard deviation of a mutation site in multiple batches of samples is 0.36, significantly greater than the set threshold of 0.3, it will be included in the gray list. When such sites are detected during interpretation, a manual review will be prompted to avoid automatic omissions or misjudgments. The gray list L... graylist It can be represented as L graylist ={ l∈L graylist | CV(l)>0.3}.
[0160] More specifically, after the initial graylist set is constructed, the latest interpretation results of graylist sites are collected during the interpretation process. If a graylist site is consistently confirmed as qualified (Pass) by manual review in a large number of subsequent samples, and the mutation abundance (AF), supported reads (AD), and site sequencing depth (DP) indicators are stable, it can be removed from the graylist after manual review and database verification. If a graylist site shows multiple interpretation discrepancies in new samples, and the mutation data fluctuations intensify, or it is reconfirmed as a false positive by manual review, the site is transferred to the blacklist for strict interception and management. At the same time, during the continuous introduction of new sample data, based on the comparison of new sample mutation parameters and dynamic thresholds, when a large number of mutation sites close to the judgment boundary and inconsistent with manual review results are found, they are screened according to the abundance dispersion index CV(l) and the preset dispersion threshold. The newly screened mutation sites can be expanded into the graylist, thereby maintaining the synchronous update and adaptive optimization of the graylist and sample data distribution characteristics.
[0161] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0162] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for identifying tumor gene mutations based on an NGS platform, characterized in that, The tumor gene mutation interpretation method includes: Obtain the mutation detection results of the sample to be analyzed. The mutation detection results include at least one candidate mutation and the mutation data corresponding to the candidate mutation. The mutation data includes mutation type, mutation abundance, number of supporting reads, site sequencing depth, and manually interpreted tags. Based on the mutation detection results, it is determined whether the candidate mutation exists in a preset list, which includes at least a blacklist and a graylist; If the candidate mutation exists in the preset list, then according to the list type of the preset list, the interpretation result corresponding to the candidate mutation is generated; If the candidate mutation does not exist in the preset list, then based on the mutation detection results, the hot spot mutation type and sample quality type of the candidate mutation of the sample to be analyzed are determined; Based on the hotspot mutation type, the sample quality type, and the mutation data, the interpretation result corresponding to the candidate mutation is generated; Specifically, generating the interpretation result corresponding to the candidate mutation based on the hotspot mutation type, the sample quality type, and the mutation data includes: Based on the hotspot mutation type and the sample quality type of the sample to be analyzed, obtain the corresponding preset dynamic threshold; If the mutation data of the candidate mutation meets the preset dynamic threshold standard, the candidate mutation is generated and marked as qualified; otherwise, the candidate mutation is generated and marked as requiring manual review.
2. The tumor gene mutation interpretation method according to claim 1, characterized in that, The step of generating the interpretation result corresponding to the candidate mutation based on the list type of the preset list specifically includes: If the candidate mutation exists in the blacklist, the interpretation result of the candidate mutation is generated and marked as unqualified; If the candidate mutation exists in the gray list, the interpretation result of the candidate mutation is generated and marked as requiring manual review.
3. The tumor gene mutation interpretation method according to claim 2, characterized in that, The step of determining the hotspot mutation type and sample quality type of the candidate mutations in the sample to be analyzed based on the mutation detection results specifically includes: Determine whether the candidate mutation of the sample to be analyzed exists in a preset whitelist, which is constructed based on known hotspot mutations in a public clinical database; If the candidate mutation exists in the whitelist, then the hotspot mutation type of the candidate mutation is determined to be a hotspot mutation; If the candidate mutation does not exist in the whitelist, then the hotspot mutation type is determined to be a non-hotspot mutation; The mutation spectrum characteristics and background noise level of the sample to be analyzed are obtained from the mutation detection results, and the sample degradation type of the sample to be analyzed is determined based on the mutation spectrum characteristics. Based on the background noise level, determine the sample noise type of the sample to be analyzed; Based on the sample degradation type and the sample noise type, the sample quality type of the sample to be analyzed is determined.
4. The tumor gene mutation interpretation method according to claim 3, characterized in that, The step of determining the sample degradation type of the sample to be analyzed based on the mutation spectrum characteristics specifically includes: Count the total number of mutations of type SNV in the mutation data of the sample to be analyzed; From the total number of mutations, the number of mutations with the mutation form C>T and whose mutation abundance in the mutation data is less than a preset abundance threshold is obtained; Based on the total number of mutations and the number of mutations, the sample degradation percentage is calculated; If the percentage of sample degradation is greater than or equal to the degradation judgment threshold, then the sample degradation type of the sample to be analyzed is determined to be a degraded sample. If the percentage of sample degradation is less than the degradation judgment threshold, then the sample to be analyzed is determined to be a non-degraded sample.
5. The tumor gene mutation interpretation method according to claim 3, characterized in that, The step of determining the sample noise type of the sample to be analyzed based on the background noise level specifically includes: Based on the background noise level, obtain the hotspot mutation noise level of at least one candidate mutation whose hotspot mutation type is hotspot mutation in the sample to be analyzed; Based on the hotspot mutation noise level, the single-sample background noise level statistics of all hotspot mutations in each sample of the sample to be analyzed are calculated. If the statistical value of the background noise level of a single sample is greater than or equal to a preset threshold, then the sample noise type of the sample to be analyzed is determined to be a high-noise sample. If the statistical value of the background noise level of a single sample is less than the preset threshold, then the sample noise type of the sample to be analyzed is determined to be a non-high noise sample.
6. The method for interpreting tumor gene mutations according to any one of claims 3-5, characterized in that, The determination of the sample quality type of the sample to be analyzed based on the sample degradation type and the sample noise type specifically includes: If the sample degradation type of the sample to be analyzed is a degraded sample and / or the sample noise type of the sample to be analyzed is a high-noise sample, then the sample quality type of the sample to be analyzed is determined to be a low-quality sample. If the sample degradation type of the sample to be analyzed is non-degraded sample and / or the sample noise type of the sample to be analyzed is non-high noise sample, then the sample quality type of the sample to be analyzed is determined to be normal sample.
7. The method for determining tumor gene mutations according to claim 6, characterized in that, Before the step of obtaining the corresponding preset dynamic threshold based on the hotspot mutation type and the sample quality type of the sample to be analyzed, the tumor gene mutation interpretation method further includes: Obtain a historical sample dataset and divide the historical sample dataset into a test set and a validation set according to a preset ratio; The test set and the validation set are grouped according to the hotspot mutation type and the sample quality type; For each group in the test set, all mutation sites whose manually identified labels are unqualified are selected, and an initial threshold boundary is set based on the mutation abundance, the number of mutation-supporting reads, and the sequencing depth of the site for all mutation sites. Based on the initial threshold boundary, a multi-dimensional grid space is constructed and threshold combination candidate regions are discretized. Through gridded search and optimization algorithms, based on the threshold combination candidate regions, a weighted objective function is used to determine the threshold combination with the largest weighted objective function value for the corresponding group in the test set. The validation set is used for validation. Mutations in each sample in the validation set that fall within the threshold combination of the corresponding group are recorded as qualified mutations. The consistency index of the qualified mutations in the corresponding group of the validation set is calculated. If the consistency index is greater than or equal to the preset consistency threshold, then the preset dynamic threshold of the corresponding group is determined according to the combination of the thresholds.
8. The method for determining tumor gene mutations according to claim 7, characterized in that, Before the step of determining whether the candidate mutation exists in the preset list, the tumor gene mutation interpretation method further includes: Filter out mutation sites in the historical sample dataset whose manually interpreted labels are unqualified, and calculate the screening frequency and standard deviation of mutation abundance of the selected mutation sites; Based on the screening frequency, mutation sites with a screening frequency greater than or equal to a preset frequency threshold are classified into a set of high-frequency unqualified mutation sites. In the set of high-frequency unqualified mutation sites, the first site whose mutation abundance standard deviation is less than a preset stability threshold is selected, and the first site selected is manually reviewed and its known clinical significance is confirmed. The first locus that is still deemed unqualified by manual review and has no known clinical significance will be included in the blacklist.
9. The method for interpreting tumor gene mutations according to claim 8, characterized in that, Before the step of determining whether the candidate mutation exists in the preset list, the tumor gene mutation interpretation method further includes: Mutation sites that do not belong to the whitelist and the blacklist are screened from the mutation sites, and the abundance dispersion index of the screened mutation sites is calculated. Mutation sites whose abundance dispersion index is greater than a preset dispersion threshold are included in the gray list.
Citation Information
Patent Citations
NGS-based brain tumor molecular diagnosis analysis method
CN112102944A
Method and system for rapid genetic analysis
US20190325988A1