Evidence optimization method, device, computer equipment and storage medium based on Bayesian model

Through the evidence optimization method based on the Bayesian model, the evidence strength level of the ACMG/AMP guidelines was quantified and the threshold was determined, which solved the problem of insufficient accuracy and consistency in the interpretation of genetic variations, improved the accuracy and classification consistency of clinical genetic testing, and supported early intervention and personalized treatment.

CN119719955BActive Publication Date: 2025-09-05WEST CHINA HOSPITAL SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411855983.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-09-05
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

Existing technologies lack software and tools specifically designed to quantify the strength of evidence and establish appropriate thresholds, resulting in low accuracy and poor consistency in the interpretation of genetic variants in clinical genetic testing.

Method used

A Bayesian model-based evidence optimization method was used to determine the theoretical pathogenicity probability corresponding to each evidence strength level by calculating the posterior probability corresponding to the 17 evidence combinations given in the ACMG/AMP guidelines under given prior probabilities. The threshold was determined by the positive likelihood ratio and local positive likelihood ratio, and the evidence labels of the variants were relabeled to expand the evidence strength level to improve classification accuracy and consistency.

Benefits of technology

It improves the accuracy and consistency of genetic variation interpretation, improves the classification of variants of unknown clinical significance, and supports early intervention and personalized treatment of Mendelian genetic diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719955B_ABST
    Figure CN119719955B_ABST
Patent Text Reader

Abstract

The present invention discloses a Bayesian model-based evidence optimization method, apparatus, computer equipment, and storage medium, relating to the technical field of healthcare informatics. The method comprises the following steps: obtaining a text file containing variant classification results; calculating the posterior probabilities corresponding to the 17 evidence combinations given in the ACMG / AMP guidelines; and obtaining the theoretical pathogenicity probability corresponding to each evidence strength level when the selected evidence combination rules satisfy the posterior probability classification; subclassifying variants of unknown clinical significance (VUS) into six levels; determining the thresholds corresponding to different evidence strength levels by comparing the positive likelihood ratio or local positive likelihood ratio with the theoretical pathogenicity probability corresponding to each evidence strength level; and relabeling the variants with ACMG / AMP evidence labels to provide the variant pathogenicity classification results. The present invention can reduce the difficulty of evidence optimization and improve the accuracy and consistency of genetic variation interpretation in clinical genetic testing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical care informatics, and in particular to a Bayesian model-based evidence optimization method, device, computer equipment, and storage medium. Background Art

[0002] Early genetic molecular diagnosis is crucial for early intervention of patient symptoms and improving their quality of life. For example, Mendelian genetic diseases are inherited diseases caused by single gene mutations. These diseases often pose serious health risks to patients and may even lead to birth defects, disability, or death. In clinical practice, targeted sequencing, whole-exome sequencing, or whole-genome sequencing, combined with bioinformatics analysis and interpretation of variant pathogenicity, can identify genetic molecular causes for patients. However, accurately determining the pathogenicity of a large number of variants of unknown significance is one of the main challenges facing clinical genetic diagnosis.

[0003] To assist in interpreting the clinical significance of variants, the American College of Medical Genetics and Genomics (ACMG) and the Association for Molecular Pathology (AMP) released the Standards and Guidelines for the Interpretation of Sequence Variants in 2015. Based on eight aspects including population frequency, phenotypic co-segregation, and functional prediction scores, a total of 28 evidence items (involving 5 evidence strength levels: independent, very strong, strong, moderate, and supportive) and 18 evidence combination rules, variants are classified into five categories: pathogenic (P), likely pathogenic (LP), benign (B), likely benign (LB), and uncertain significance (US). After the release of the ACMG / AMP guidelines, they have been widely recognized by international peers, greatly enhancing the global clinical gene testing and data interpretation capabilities. However, the evidence combination rules of the ACMG / AMP guidelines only qualitatively analyze the pathogenicity of variants based on a simple evidence strength level accumulation rule, lacking a mathematical model to prove its effectiveness. In 2018, ClinGen proposed that the qualitative framework of the ACMG / AMP guidelines could be transformed into a quantitative framework based on the Naive Bayes model. This quantitative framework calculates the posterior probability (Post_P) of variant pathogenicity according to the prior probability of variant pathogenicity and the pathogenicity odds corresponding to the evidence strength level, and then realizes variant classification. Specifically, when Post_P > 0.99, 0.90 < Post_P ≤ 0.99, 0.001 ≤ Post_P < 0.10, Post_P < 0.001, and 0.10 ≤ Post_P ≤ 0.90, the corresponding pathogenicity classifications of the variants are P, LP, LB, B, and US, respectively. Based on this quantitative framework, ClinGen SVI demonstrated that among the 18 evidence combination rules proposed by the ACMG / AMP Standards and Guidelines for the Interpretation of Sequence Variants, 16 were consistent with the classification results calculated by the Bayesian model, proving the rationality of the evidence combination rules of the ACMG / AMP guidelines. In 2020, ClinGen further transformed the Bayesian framework into a scoring framework, where different evidence strength levels correspond to different scores, and variants can be classified by adding up the scores. These quantitative models provide a mathematical basis for further quantifying the evidence strength level and expanding the combination rules. Researchers quantified and expanded the strength levels of multiple evidence items, such as PS1 / PM5, PS3 / BS3, and PP3 / BP4, according to the Bayesian framework, improving the recognition rate of pathogenic variants, and proving that the standardization and stratification of evidence are crucial for establishing objective, unbiased, and consistent variant classification.

[0004] However, there is currently a lack of software and tools specifically designed to quantify the strength of evidence and establish appropriate thresholds to improve the ACMG / AMP guidelines, resulting in low accuracy and poor consistency in the interpretation of genetic variants in clinical genetic testing. Summary of the Invention

[0005] In response to the above-mentioned deficiencies in the prior art, the present invention provides a Bayesian model-based evidence optimization method, apparatus, computer device, and storage medium, which aims to reduce the difficulty of evidence optimization and improve the accuracy and consistency of genetic variation interpretation in clinical genetic testing by expanding the evidence strength level of the ACMG / AMP guidelines.

[0006] In a first aspect, the present invention provides an evidence optimization method based on a Bayesian model, comprising the following steps:

[0007] S10: Obtain a text file containing variant classification results, wherein the variant classification results include variants, ACMG / AMP evidence, and variant pathogenicity classification information;

[0008] S20: When there is an exponential relationship between the pathogenicity probabilities of different evidence strength levels, the posterior probabilities (Post_P) corresponding to the 17 evidence combinations given in the ACMG / AMP guidelines are calculated according to a preset number of cycles under a given prior probability (Prior_P). When the selected evidence combination rule satisfies the posterior probability (Post_P) classification, the minimum pathogenicity probability value corresponding to each evidence strength level is obtained, and the minimum pathogenicity probability value is used as the theoretical pathogenicity probability corresponding to each evidence strength level.

[0009] S30: Variants of uncertain clinical significance (VUS) are subclassified into first-level hot, second-level warm, third-level tepid, fourth-level cool, fifth-level cold, and sixth-level ice-cold. Variants of uncertain clinical significance (VUS) classified as first-level hot, second-level warm, and third-level tepid are defined as PL-VUS with a tendency to cause disease. Variants of uncertain clinical significance (VUS) classified as fourth-level cool, fifth-level cold, and sixth-level ice-cold are defined as BL-VUS with a tendency to cause benign disease. PL-VUS with a tendency to cause disease are filtered out during evidence optimization.

[0010] S40: For categorical variables, calculate the positive likelihood ratio LR + For continuous variables, the local positive likelihood ratio lr+ is calculated by comparing the positive likelihood ratio LR + or local positive likelihood ratio lr + and the theoretical pathogenicity probability corresponding to each level of evidence strength, and determine the thresholds corresponding to different levels of evidence strength;

[0011] S50: Based on the optimized evidence strength corresponding level and corresponding threshold, the ACMG / AMP evidence label is re-labeled for the variant, and the pathogenicity classification results of the variant are given according to the evidence combination rules, Bayesian framework and scoring system given in the ACMG / AMP guidelines.

[0012] Furthermore, in step S20, the posterior probability Post_P classification is specifically as follows: the posterior probability Post_P of the pathogenic combination satisfies: posterior probability Post_P>0.99; the posterior probability Post_P of the possible pathogenic combination satisfies: 0.90< posterior probability Post_P ≤ 0.99; the posterior probability Post_P of the possible benign combination satisfies: 0.001 ≤ posterior probability Post_P<0.10; the posterior probability Post_P of the benign combination satisfies: posterior probability Post_P<0.001;

[0013] The calculation formula for the theoretical pathogenicity probability corresponding to each level of evidence strength is:

[0014] ;

[0015] OP: Odds Path, the pathogenicity probability obtained by combining ACMG / AMP evidence; O PVSt : The probability of pathogenicity corresponding to the Very Strong pathogenicity evidence strength level; N PVSt : The number of evidences with a pathogenicity strength level of very strong (Very Strong); N PSt : the number of evidences with strong pathogenicity evidence level; N PM : The number of evidences with a moderate level of pathogenicity evidence; N PSu : The number of evidences with the pathogenicity strength level as supporting (Supporting); N BSt : The number of evidences with strong level of benign evidence; N BSu : The number of evidences with the benign evidence strength level as supporting (Supporting);

[0016] The calculation formula for the posterior probability Post_P is:

[0017] ;

[0018] Among them, Post_P: posterior probability of variant pathogenicity; Prior_P: prior probability of variant pathogenicity.

[0019] Furthermore, in step S40, for categorical variables, various statistical indicators for each test condition are calculated by counting the number of P / LP and BL-VUS / B / LB variants that meet or do not meet the test conditions, including true positives, false positives, true negatives, accuracy, true positive rate, true negative rate, positive predictive value, negative predictive value, F1 score, and positive likelihood ratio LR. + , the positive likelihood ratio LR is obtained by the bootstrapping algorithm + The 95% confidence interval estimate of the positive likelihood ratio LR is then compared + The evidence strength level and threshold are determined by combining the lower limit of the 95% confidence interval of the obtained result with the theoretical pathogenicity probability corresponding to each evidence strength level obtained in step S20.

[0020] Furthermore, the positive likelihood ratio LR + The calculation formula is:

[0021] ;

[0022] Wherein, TP: True positive, the number of P / LP variants that meet the test conditions; FP: False positive, the number of BL-VUS / B / LB variants that meet the test conditions; TN: True negative, the number of BL-VUS / B / LB variants that do not meet the test conditions; FN: False negative, the number of P / LP variants that do not meet the test conditions;

[0023] The local positive likelihood ratio lr + The calculation formula is:

[0024] ;

[0025] Among them, p: probability; s: test value; v: variation; : represents the density of the benign variation score under a certain test value; : represents the density of the distribution of pathogenic variant scores under a certain test value.

[0026] Furthermore, in step S40, for continuous variables, all observations are sorted, and with each observation as the center, for a given sliding window, the local posterior probability of pathogenicity of each test value in the interval is calculated, and the 95% confidence interval of the local posterior probability of each test value is estimated through a preset number of iterations to evaluate the threshold corresponding to the evidence strength level.

[0027] Furthermore, the calculation formula of the local posterior probability of pathogenicity is:

[0028] ;

[0029] Where, Local posterior probability: local posterior probability; #P / LP variants in interval: the number of P / LP variants in a given sliding window; #non-pathogenic variants in interval: the number of BL-VUS / B / LB variants in a given sliding window; Weight: the probability of BL-VUS / B / LB variants in the dataset.

[0030] Furthermore, the test value of the categorical variable is a binary variable; and the test value of the continuous variable is a continuous variable.

[0031] In a second aspect, the present invention provides an evidence optimization device based on a Bayesian model, comprising:

[0032] an acquisition module, configured to obtain a text file containing variant classification results, wherein the variant classification results include variants, ACMG / AMP evidence, and variant pathogenicity classification information;

[0033] A calculation and selection module is used to calculate the posterior probability (Post_P) corresponding to the 17 evidence combinations given in the ACMG / AMP guidelines under a given prior probability (Prior_P) according to a preset number of cycles when there is an exponential relationship between the pathogenicity probabilities of each evidence strength level. When the selected evidence combination rule satisfies the posterior probability (Post_P) classification, the minimum pathogenicity probability value corresponding to each evidence strength level is obtained, and the minimum pathogenicity probability value is used as the theoretical pathogenicity probability corresponding to each evidence strength level.

[0034] A subclassification module is used to subclassify VUSs into first-level hot, second-level warm, third-level tepid, fourth-level cool, fifth-level cold, and sixth-level ice-cold. VUSs classified as first-level hot, second-level warm, and third-level tepid are considered pathogenic variants (PL-VUS). VUSs classified as fourth-level cool, fifth-level cold, and sixth-level ice-cold are considered benign variants (BL-VUS). PL-VUSs with pathogenic tendencies are filtered out during evidence optimization.

[0035] Threshold determination module, used to calculate the positive likelihood ratio LR for categorical variables + ; For continuous variables, calculate the local positive likelihood ratio lr+ , by comparing the positive likelihood ratio LR + or local positive likelihood ratio lr + and the theoretical pathogenicity probability corresponding to each level of evidence strength, and determine the thresholds corresponding to different levels of evidence strength;

[0036] The result output module is used to re-label the variants with ACMG / AMP evidence labels based on the optimized evidence strength corresponding levels and corresponding thresholds, and to provide the pathogenicity classification results of the variants according to the evidence combination rules, Bayesian framework and scoring system given in the ACMG / AMP guidelines.

[0037] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.

[0038] In a fourth aspect, the present invention provides a storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.

[0039] The beneficial effects of the present invention are:

[0040] The present invention first obtains a text file containing the variant classification results, and calculates the posterior probability corresponding to the 17 evidence combinations given by the ACMG / AMP guidelines under a given prior probability. When the selected evidence combination rule meets the posterior probability Post_P classification, the theoretical pathogenicity probability corresponding to each evidence strength level is obtained. For categorical variables, the positive likelihood ratio LR is calculated. + ; For continuous variables, calculate the local positive likelihood ratio lr + By comparing the positive likelihood ratio LR + or local positive likelihood ratio lr + The theoretical pathogenicity probability corresponding to each evidence strength level can be determined, and the thresholds corresponding to different evidence strength levels can be determined, thereby improving the flexibility and convenience of quantifying ACMG / AMP evidence strength levels and establishing corresponding thresholds, which can promote the transition from subjective assessment to quantitative and objective assessment.

[0041] Then, based on the optimized evidence strength corresponding level and corresponding threshold, the ACMG / AMP evidence label is re-labeled for the variant, and the pathogenicity classification results of the variant are given according to the evidence combination rules, Bayesian framework and scoring system given in the ACMG / AMP guidelines. Based on the evidence optimization results, the pathogenicity of the variant is reclassified, and the classification is not limited to using the ACMG / AMP combination rules, but can also be classified according to the posterior probability calculated according to the Bayesian framework. By expanding the evidence strength level of the ACMG / AMP guidelines, the present invention is expected to further improve the accuracy and consistency of variant interpretation in clinical genetic testing and improve the classification of variants of unknown clinical significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 1 is a flow chart of an evidence optimization method based on a Bayesian model provided by an embodiment of the present invention;

[0043] Figure 2 is a corresponding relationship diagram of variant pathogenicity classification, posterior probability, and evidence combination provided by an embodiment of the present invention;

[0044] Figure 3 This is a functional module block diagram of an evidence optimization device based on a Bayesian model provided by an embodiment of the present invention.

[0045] icon:

[0046] Acquisition module 100; Calculation and selection module 200; Sub-classification module 300;

[0047] Threshold determination module 400; Result output module 500.

[0048] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. It should be understood that the specific embodiments described herein are only used to illustrate the present invention and are not intended to limit the present invention. Example

[0049] like Figure 1 As shown, an evidence optimization method based on a Bayesian model provided by an embodiment of the present invention may include the following steps S10 to S40:

[0050] Step S10: Obtain a text file containing the variant classification results.

[0051] The variant classification results include at least the variant, ACMG / AMP evidence, and variant pathogenicity classification information. The variant, ACMG / AMP evidence, and variant pathogenicity classification information are separated by a tab.

[0052] Step S20: When there is an exponential relationship between the pathogenicity probabilities of each evidence strength level, the posterior probability Post_P corresponding to the 17 evidence combinations given in the ACMG / AMP guidelines is calculated according to a preset number of cycles under a given prior probability Prior_P. When the selected evidence combination rule meets the posterior probability Post_P classification, the minimum pathogenicity probability value corresponding to each evidence strength level is obtained, and the minimum pathogenicity probability value is used as the theoretical pathogenicity probability corresponding to each evidence strength level.

[0053] In this example, an exponential relationship between the pathogenicity probabilities of each evidence strength level is first assumed. During implementation, the preset number of cycles is preferably 10,000. Within this preset number of cycles, the posterior probabilities (Post_P) corresponding to the 17 evidence combinations provided by the ACMG / AMP guidelines are calculated under a given prior probability (Prior_P). When the selected evidence combination rule satisfies the posterior probability (Post_P) classification, the minimum pathogenicity probability value corresponding to each evidence strength level is obtained, and this minimum pathogenicity probability value is used as the theoretical pathogenicity probability corresponding to each evidence strength level.

[0054] In one embodiment, the posterior probability Post_P classification is specifically divided into the following cases: the posterior probability Post_P of the pathogenic combination satisfies: posterior probability Post_P>0.99; the posterior probability Post_P of the possible pathogenic combination satisfies: 0.90<posterior probability Post_P ≤ 0.99; the posterior probability Post_P of the possible benign combination satisfies: 0.001 ≤ posterior probability Post_P<0.10; the posterior probability Post_P of the benign combination satisfies: posterior probability Post_P<0.001.

[0055] Furthermore, the calculation formula for the theoretical pathogenicity probability corresponding to each level of evidence strength is:

[0056] ;

[0057] OP: Odds Path, the pathogenicity probability obtained by combining ACMG / AMP evidence; O PVSt : The probability of pathogenicity corresponding to the Very Strong pathogenicity evidence strength level; N PVSt : The number of evidences with a pathogenicity strength level of very strong (Very Strong); N PSt : the number of evidences with strong pathogenicity evidence level; N PM : The number of evidences with a moderate level of pathogenicity evidence; N PSu : The number of evidences with the pathogenicity strength level as supporting (Supporting); N BSt: The number of evidences with strong level of benign evidence; N BSu : The number of pieces of evidence with the strength of benign evidence level as supporting.

[0058] Furthermore, the calculation formula of the posterior probability Post_P is:

[0059] ;

[0060] Among them, Post_P: posterior probability of variant pathogenicity; Prior_P: prior probability of variant pathogenicity.

[0061] In this embodiment, the corresponding relationship between the pathogenicity classification of the variant, the posterior probability Post_P and the evidence combination is as follows: Figure 2 As shown, there are 17 combinations of evidence. Figure 2 In, O PVSt : Very Strong pathogenicity evidence strength level corresponding to the probability of pathogenicity; O PSt : Strong pathogenicity evidence strength level corresponding to the probability of pathogenicity; O PM : The probability of pathogenicity corresponding to the moderate pathogenicity evidence strength level; O PSu : Supporting pathogenicity evidence strength level corresponding to the probability of pathogenicity; O BSt : Strong (Strong) benign evidence strength level corresponding to the probability of pathogenicity; O BSu : The probability of pathogenicity corresponding to the strength level of evidence supporting benignity.

[0062] It should be noted that the pathogenicity classification of variants includes: P (pathogenic), LP (likely pathogenic), B (benign), LB (likely benign) and US (uncertain significance).

[0063] Step S30: Variants of unknown clinical significance (VUS) are subclassified into first-level hot, second-level warm, third-level tepid, fourth-level cool, fifth-level cold, and sixth-level ice-cold. Variants of unknown clinical significance (VUS) classified as first-level hot, second-level warm, and third-level tepid are defined as PL-VUS with a tendency to cause disease. Variants of unknown clinical significance (VUS) classified as fourth-level cool, fifth-level cold, and sixth-level ice-cold are defined as BL-VUS with a tendency to cause benign disease. During evidence optimization, PL-VUS with a tendency to cause disease are filtered out.

[0064] Step S40: For categorical variables, calculate the positive likelihood ratio LR + ; For continuous variables, calculate the local positive likelihood ratio lr + , by comparing the positive likelihood ratio LR + or local positive likelihood ratio lr + The theoretical pathogenicity probability corresponding to each level of evidence strength is used to determine the threshold values ​​corresponding to different levels of evidence strength.

[0065] In one embodiment, for categorical variables, various statistical indicators including true positive, false positive, true negative, accuracy, true positive rate, true negative rate, positive predictive value, negative predictive value, F1 score and positive likelihood ratio LR are calculated for each test value by counting the number of P / LP and BL-VUS / B / LB variants that meet or do not meet the test conditions. + , the positive likelihood ratio LR is obtained by the bootstrapping algorithm + The 95% confidence interval estimate of the positive likelihood ratio LR is then compared + The evidence strength level and threshold are determined by combining the lower limit of the 95% confidence interval of the obtained result with the theoretical pathogenicity probability corresponding to each evidence strength level obtained in step S20.

[0066] It should be noted that the bootstrapping process is formally described as follows: for a given natural language processing task, a specific guided method for training a classification model is selected. Two datasets are then required: a small labeled dataset L and an unlabeled dataset U. Finally, the labeled dataset is gradually expanded using the unlabeled dataset U, thereby training the final classifier to achieve the specific natural language processing task. Bootstrapping is an iterative process that can obtain a large amount of annotated data with high confidence from a small amount of annotated data. Bootstrapping is implemented through two main steps: first, providing a heuristic rule or other classification method that can effectively classify a small amount of data. Second, evaluating the newly annotated data generated by the classifier. This evaluation process generates annotated data with high confidence, which can then be iteratively obtained to obtain a larger annotated dataset. The iterations are terminated by a threshold number of iterations or when the amount of newly annotated data generated is too low.

[0067] Here, the categorical variable refers to a binary variable with a test value, for example, the test is whether the variant is marked with a certain evidence label (such as PM2, PS1, etc.), such as whether the variant is significantly enriched in the population (PS4 evidence, the corresponding test value is OR>5 and p<0.05). For categorical variables, a four-cell table will be obtained, and then the positive likelihood ratio LR will be calculated. + value.

[0068] Furthermore, the positive likelihood ratio LR + The calculation formula is:

[0069] ;

[0070] TP: True positive, the number of P / LP variants above the test value; FP: False positive, the number of BL-VUS / B / LB variants above the test value; TN: True negative, the number of BL-VUS / B / LB variants below the test value; FN: False negative, the number of P / LP variants above the test value.

[0071] The local positive likelihood ratio lr + The calculation formula is:

[0072] ;

[0073] Among them, p: probability; s: test value; v: variation; : represents the density of the benign variation score under a certain test value; : represents the density of the distribution of pathogenic variant scores under a certain test value.

[0074] In one embodiment, for continuous variables, all observations are sorted, and with each observation as the center, for a given sliding window, the local posterior probability of pathogenicity of each test value in the interval is calculated, and the 95% confidence interval of the local posterior probability of each test value is estimated through a preset number of iterations to evaluate the threshold corresponding to the strength of evidence level.

[0075] A continuous variable refers to a variable whose test value is continuous. For example, PM2 / BS1 evidence corresponds to a threshold for the frequency of a variant in the population, and PP3 / BP4 corresponds to a threshold for the pathogenicity score of a variant predicted by the software. When determining the thresholds corresponding to different levels of evidence strength, various frequency values ​​or scores are tested, and these values ​​are continuous and have a range. During implementation, the preset number of tests is set by the user, preferably 100, 1000, and 5000.

[0076] Furthermore, the calculation formula of the local posterior probability of pathogenicity is:

[0077] ;

[0078] Where, Local posterior probability: local posterior probability; #P / LP variants in interval: the number of P / LP variants in a given sliding window; #non-pathogenic variants in interval: the number of BL-VUS / B / LB variants in a given sliding window; Weight: the probability of BL-VUS / B / LB variants in the dataset.

[0079] This embodiment improves the flexibility and convenience of quantifying ACMG / AMP evidence strength levels and establishing corresponding thresholds by determining thresholds corresponding to evidence strength levels, thereby promoting the transition from subjective assessment to quantitative and objective assessment.

[0080] Step S50: Re-label the variant with the ACMG / AMP evidence label based on the optimized evidence strength corresponding level and corresponding threshold, and give the pathogenicity classification results of the variant according to the evidence combination rules, Bayesian framework and scoring system given in the ACMG / AMP guidelines.

[0081] This example reclassifies variant pathogenicity based on the evidence optimization results. Furthermore, the classification is not limited to using the ACMG / AMP combined rules but can also be performed based on the calculation of posterior probabilities within a Bayesian framework. By expanding the evidence strength level of the ACMG / AMP guidelines, the present invention is expected to further enhance the accuracy and consistency of variant interpretation in clinical genetic testing and improve the classification of variants of unknown clinical significance. This may, for example, provide strong support for early intervention and personalized treatment of Mendelian genetic diseases.

[0082] Example 2

[0083] like Figure 3 As shown, an embodiment of the present invention further provides an evidence optimization device based on a Bayesian model, which may include an acquisition module 100, a calculation and selection module 200, a sub-classification module 300, a threshold determination module 400 and a result output module 500.

[0084] The acquisition module 100 is used to acquire a text file containing a variant classification result, wherein the variant classification result includes variant, ACMG / AMP evidence, and variant pathogenicity classification information.

[0085] The calculation and selection module 200 is used to calculate the posterior probability Post_P corresponding to the 17 evidence combinations given in the ACMG / AMP guidelines under a given prior probability Prior_P according to a preset number of cycles when there is an exponential relationship between the pathogenicity probabilities of each evidence strength level. When the selected evidence combination rule meets the posterior probability Post_P classification, the minimum pathogenicity probability value corresponding to each evidence strength level is obtained, and the minimum pathogenicity probability value is used as the theoretical pathogenicity probability corresponding to each evidence strength level.

[0086] The subclassification module 300 is used to subclassify variants of clinical significance (VUS) into first-level hot, second-level warm, third-level tepid, fourth-level cool, fifth-level cold, and sixth-level ice-cold, and to classify variants of clinical significance (VUS) classified as first-level hot, second-level warm, and third-level tepid as variants of clinical significance (PL-VUS) with a tendency to be pathogenic, and to classify variants of clinical significance (VUS) classified as fourth-level cool, fifth-level cold, and sixth-level ice-cold as variants of clinical significance (BL-VUS) with a tendency to be benign. During evidence optimization, variants of clinical significance (PL-VUS) with a tendency to be pathogenic are filtered out.

[0087] The threshold determination module 400 is used to calculate the positive likelihood ratio LR for the categorical variable. + ; For continuous variables, calculate the local positive likelihood ratio lr + , by comparing the positive likelihood ratio LR+ or local positive likelihood ratio lr + The theoretical pathogenicity probability corresponding to each level of evidence strength is used to determine the threshold values ​​corresponding to different levels of evidence strength.

[0088] The result output module 500 is used to re-label the ACMG / AMP evidence label for the variant according to the optimized evidence strength corresponding level and corresponding threshold, and to provide the pathogenicity classification results of the variant according to the evidence combination rules, Bayesian framework and scoring system provided by the ACMG / AMP guidelines.

[0089] In detail, each module in the Bayesian model-based evidence optimization device provided in the embodiment of the present invention adopts the same technical means as the Bayesian model-based evidence optimization method provided in the above-mentioned Example 1 when used, and can produce the same technical effects, which will not be repeated here.

[0090] Example 3

[0091] An embodiment of the present invention further provides a computer device, comprising a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method described in embodiment 1.

[0092] In some embodiments, the processor may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor is the control core (Control Unit) of the electronic device, connecting the various components of the entire electronic device using various interfaces and lines, and executing programs or modules stored in the memory (such as an evidence optimization program based on a Bayesian model, etc.), as well as calling data stored in the memory, to perform various functions and process data.

[0093] In some embodiments, the memory may be an internal storage unit of an electronic device, such as a mobile hard disk of the electronic device. In other embodiments, the memory may also be an external storage device of an electronic device, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device. Furthermore, the memory may also include both an internal storage unit and an external storage device of the electronic device. The memory can be used not only to store application software and various types of data installed in the electronic device, such as the code of an evidence optimization program based on a Bayesian model, but also to temporarily store data that has been output or is to be output.

[0094] Example 4

[0095] An embodiment of the present invention further provides a storage medium having a computer program stored thereon. When the computer program is executed by a processor, the processor executes the steps of the method described in Example 1.

[0096] The storage medium is a readable storage medium, which may include a flash memory, a mobile hard disk, a multimedia card, a card-type memory (eg, an SD or DX memory), a magnetic memory, a magnetic disk, an optical disk, and the like.

[0097] In summary, the embodiment of the present invention first obtains a text file containing the variant classification results, and calculates the posterior probability corresponding to the 17 evidence combinations given by the ACMG / AMP guidelines under a given prior probability. When the selected evidence combination rule meets the posterior probability Post_P classification, the theoretical pathogenicity probability corresponding to each evidence strength level is obtained. For categorical variables, the positive likelihood ratio LR is calculated. + ; For continuous variables, calculate the local positive likelihood ratio lr + By comparing the positive likelihood ratio LR + or local positive likelihood ratio lr + The theoretical pathogenicity probability corresponding to each evidence strength level can be determined, and the thresholds corresponding to different evidence strength levels can be determined, thereby improving the flexibility and convenience of quantifying ACMG / AMP evidence strength levels and establishing corresponding thresholds, which can promote the transition from subjective assessment to quantitative and objective assessment.

[0098] Then, based on the optimized evidence strength corresponding level and corresponding threshold, the ACMG / AMP evidence label is re-labeled for the variant, and the pathogenicity classification results of the variant are given according to the evidence combination rules, Bayesian framework and scoring system given in the ACMG / AMP guidelines. Based on the evidence optimization results, the pathogenicity of the variant is reclassified, and the classification is not limited to using the ACMG / AMP combination rules, but can also be classified according to the posterior probability calculated according to the Bayesian framework. By expanding the evidence strength level of the ACMG / AMP guidelines, the present invention is expected to further improve the accuracy and consistency of variant interpretation in clinical genetic testing and improve the classification of variants of unknown clinical significance.

[0099] In the embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic. For example, the division of the modules is merely a logical functional division, and there may be other division methods in actual implementation. The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, the functional modules in the various embodiments of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional modules.

[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An evidence optimization method based on a Bayesian model, characterized in that: The following steps are involved: S10: Obtain a text file containing variant classification results, wherein the variant classification results include variants, ACMG / AMP evidence, and variant pathogenicity classification information; S20: When there is an exponential relationship between the pathogenicity probabilities of different evidence strength levels, the posterior probabilities (Post_P) corresponding to the 17 evidence combinations given in the ACMG / AMP guidelines are calculated according to a preset number of cycles under a given prior probability (Prior_P). When the selected evidence combination rule satisfies the posterior probability (Post_P) classification, the minimum pathogenicity probability value corresponding to each evidence strength level is obtained, and the minimum pathogenicity probability value is used as the theoretical pathogenicity probability corresponding to each evidence strength level. The posterior probability Post_P classification is specifically as follows: the posterior probability Post_P of the pathogenic combination satisfies: Post_P>0.99; the posterior probability Post_P of the possible pathogenic combination satisfies: 0.90<Post_P≤0.99; the posterior probability Post_P of the possible benign combination satisfies: 0.001≤Post_P<0.10; the posterior probability Post_P of the benign combination satisfies: Post_P<0.001; The calculation formula for the theoretical pathogenicity probability corresponding to each level of evidence strength is: OP: Odds Path, the pathogenicity probability obtained by combining ACMG / AMP evidence; O PVSt : The probability of pathogenicity corresponding to the very strong pathogenicity evidence strength level; N PVSt : The number of evidences with a very strong pathogenicity evidence level; N PSt : The number of strong evidences for pathogenicity; N PM : The number of evidences with medium pathogenicity strength level; N PSu : The number of pieces of evidence supporting the pathogenicity; N BSt : The number of strong evidences; N BSu : The strength of evidence level for benign is the number of supporting evidence; The calculation formula for the posterior probability Post_P is: Wherein, Post_P: posterior probability of variant pathogenicity; Prior_P: prior probability of variant pathogenicity; S30: Variants of uncertain clinical significance (VUS) are subclassified into first-level hot, second-level warm, third-level tepid, fourth-level cool, fifth-level cold, and sixth-level ice-cold. Variants of uncertain clinical significance (VUS) classified as first-level hot, second-level warm, and third-level tepid are defined as PL-VUS with a tendency to cause disease. Variants of uncertain clinical significance (VUS) classified as fourth-level cool, fifth-level cold, and sixth-level ice-cold are defined as BL-VUS with a tendency to cause benign disease. PL-VUS with a tendency to cause disease are filtered out during evidence optimization. S40: For categorical variables, calculate the positive likelihood ratio LR + ; For continuous variables, calculate the local positive likelihood ratio lr + , by comparing the positive likelihood ratio LR + or local positive likelihood ratio lr + and the theoretical pathogenicity probability corresponding to each level of evidence strength, and determine the thresholds corresponding to different levels of evidence strength; The positive likelihood ratio LR + The calculation formula is: Wherein, TP: True positive, the number of P / LP variants that meet the test conditions; FP: False positive, the number of BL-VUS / B / LB variants that meet the test conditions; TN: True negative, the number of BL-VUS / B / LB variants that do not meet the test conditions; FN: False negative, the number of P / LP variants that do not meet the test conditions; P, LP, B, LB, and BL-VUS represent different variant types; P: pathogenic; LP: likely pathogenic; B: benign; LB: likely benign; BL-VUS: benign-leaning VUS; VUS: variants of uncertain significance. The local positive likelihood ratio lr + The calculation formula is: Where p is probability; s is test value; v is variant; p(s|v is benign): represents the density of the fraction of benign variants under a certain test value; p(s|v is pathogenic): represents the density of the fraction of pathogenic variants under a certain test value; S50: Based on the optimized evidence strength corresponding level and corresponding threshold, the ACMG / AMP evidence label is re-labeled for the variant, and the pathogenicity classification results of the variant are given according to the evidence combination rules, Bayesian framework and scoring system given in the ACMG / AMP guidelines.

2. The Bayesian model-based evidence optimization method according to claim 1, characterized in that: In step S40, for categorical variables, various statistical indicators for each test condition are calculated by counting the number of P / LP and BL-VUS / B / LB variants that meet or do not meet the test conditions, including true positive, false positive, true negative, accuracy, true positive rate, true negative rate, positive predictive value, negative predictive value, F1 score and positive likelihood ratio LR + , the positive likelihood ratio LR is obtained by the bootstrapping algorithm + The 95% confidence interval estimate is then compared by comparing the positive likelihood ratio LR + The 95% confidence interval lower limit value of the obtained result and the theoretical pathogenicity probability corresponding to each evidence strength level obtained in step S20 are used to determine the evidence strength level and the threshold value.

3. The Bayesian model-based evidence optimization method according to claim 1, characterized in that: In step S40, for continuous variables, all observations are sorted, and with each observation as the center, for a given sliding window, the local posterior probability of pathogenicity of each test value in the interval is calculated, and the 95% confidence interval of the local posterior probability of each test value is estimated through a preset number of iterations to evaluate the threshold corresponding to the evidence strength level.

4. The Bayesian model-based evidence optimization method according to claim 3, characterized in that: The calculation formula of the local posterior probability of pathogenicity is: Where, Local posterior probability: local posterior probability; #P / LP variants in interval: the number of P / LP variants in a given sliding window; #non-pathogenic variants in interval: the number of BL-VUS / B / LB variants in a given sliding window; Weight: the probability of BL-VUS / B / LB variants in the dataset.

5. The Bayesian model-based evidence optimization method according to claim 1, characterized in that: The test value of the categorical variable is a binary variable; the test value of the continuous variable is a continuous variable.

6. An evidence optimization device based on a Bayesian model, characterized in that: include: an acquisition module, configured to obtain a text file containing variant classification results, wherein the variant classification results include variants, ACMG / AMP evidence, and variant pathogenicity classification information; A calculation and selection module is used to calculate the posterior probability (Post_P) corresponding to the 17 evidence combinations given in the ACMG / AMP guidelines under a given prior probability (Prior_P) according to a preset number of cycles when there is an exponential relationship between the pathogenicity probabilities of each evidence strength level. When the selected evidence combination rule satisfies the posterior probability (Post_P) classification, the minimum pathogenicity probability value corresponding to each evidence strength level is obtained, and the minimum pathogenicity probability value is used as the theoretical pathogenicity probability corresponding to each evidence strength level. The posterior probability Post_P classification is specifically as follows: the posterior probability Post_P of the pathogenic combination satisfies: Post_P>0.99; the posterior probability Post_P of the possible pathogenic combination satisfies: 0.90<Post_P≤0.99; the posterior probability Post_P of the possible benign combination satisfies: 0.001≤Post_P<0.10; the posterior probability Post_P of the benign combination satisfies: Post_P<0.001; The calculation formula for the theoretical pathogenicity probability corresponding to each level of evidence strength is: OP: Odds Path, the pathogenicity probability obtained by combining ACMG / AMP evidence; O PVSt : The probability of pathogenicity corresponding to the very strong pathogenicity evidence strength level; N PVSt : The number of evidences with a very strong pathogenicity evidence level; N PSt : The number of strong evidences for pathogenicity; N PM : The number of evidences with medium pathogenicity strength level; N PSu : The number of pieces of evidence supporting the pathogenicity; N BSt : The number of strong evidences; N BSu : The strength of evidence level for benign is the number of supporting evidence; The calculation formula for the posterior probability Post_P is: Wherein, Post_P: posterior probability of variant pathogenicity; Prior_P: prior probability of variant pathogenicity; A subclassification module is used to subclassify VUSs into first-level hot, second-level warm, third-level tepid, fourth-level cool, fifth-level cold, and sixth-level ice-cold. VUSs classified as first-level hot, second-level warm, and third-level tepid are considered pathogenic variants (PL-VUS). VUSs classified as fourth-level cool, fifth-level cold, and sixth-level ice-cold are considered benign variants (BL-VUS). PL-VUSs with pathogenic tendencies are filtered out during evidence optimization. Threshold determination module, used to calculate the positive likelihood ratio LR for categorical variables + ; For continuous variables, calculate the local positive likelihood ratio lr + , by comparing the positive likelihood ratio LR + or local positive likelihood ratio lr + and the theoretical pathogenicity probability corresponding to each level of evidence strength, and determine the thresholds corresponding to different levels of evidence strength; The positive likelihood ratio LR + The calculation formula is: Wherein, TP: True positive, the number of P / LP variants that meet the test conditions; FP: False positive, the number of BL-VUS / B / LB variants that meet the test conditions; TN: True negative, the number of BL-VUS / B / LB variants that do not meet the test conditions; FN: False negative, the number of P / LP variants that do not meet the test conditions; The local positive likelihood ratio lr + The calculation formula is: Where p is probability; s is test value; v is variant; p(s|v is benign): represents the density of the fraction of benign variants under a certain test value; p(s|v is pathogenic): represents the density of the fraction of pathogenic variants under a certain test value; The result output module is used to re-label the variants with ACMG / AMP evidence labels based on the optimized evidence strength corresponding levels and corresponding thresholds, and to provide the pathogenicity classification results of the variants according to the evidence combination rules, Bayesian framework and scoring system given in the ACMG / AMP guidelines.

7. A computer device, characterized in that: The invention comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 5.

8. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Systems and methods for classifying, prioritizing and interpreting genetic variants and therapies using a deep neural network

    CA2894317A1

  • Evidence judgment method, system and device for genetic variation and medium

    CN115798579A