Automatic classification method, device and equipment for genetic variation of embryonic line and storage medium
By acquiring a set of variant data to generate an initial set of evidence, using a machine learning model to optimize conflicting evidence, and combining it with a Bayesian framework for classification decisions, the problem of incomplete evidence coverage in ACMG was solved, and efficient and standardized germline gene variant classification was achieved.
Patent Information
- Application Number
- CN202510967608.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-11-21
AI Technical Summary
Existing automated assessment tools for ACMG evidence are incomplete, subject to subjective bias, and difficult to achieve efficient and standardized classification of germline genetic variations.
By acquiring a set of variant data, an initial set of evidence is generated. Conflicting evidence is optimized using a machine learning model. A classification decision is made using a Bayesian framework, and the classification result and confidence level are output.
It achieves efficient and standardized classification of germline gene variations, solves the problems of incomplete evidence coverage and insufficient gene specificity of ACMG, and improves the consistency and dynamic updating capability of classification.
Smart Images

Figure CN120998295A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of bioinformatics and medical genetics, and particularly relates to a germline gene variation automatic classification method, device, equipment and storage medium. BACKGROUND
[0002] At present, the ACMG / AMP 2015 guidelines provide a standardized framework for clinical interpretation of germline variations, defining 28 evidence criteria, which are divided into pathogenic evidence (PVS1, PS1-PS5, PM1-PM7, PP1-PP6) and benign evidence (BA1, BS1-BS5, BP1-BP8). These evidences involve multiple aspects such as variation type, population frequency, functional experiment, computational prediction, and family cosegregation. However, manual application of ACMG guidelines is time-consuming and subjective, especially in high-throughput sequencing (such as whole exome sequencing WES or whole genome sequencing WGS), which may produce tens of thousands to hundreds of thousands of variations each time, and the manual classification efficiency is low and the consistency is poor.
[0003] Although there are tools that attempt to achieve automatic evaluation of ACMG evidence and variation classification, including InterVar, VarCards2, MAGI-ACMG, vaRHC and UMD Variant Interpretation Tool. These tools have improved classification efficiency to some extent, but still have limitations: incomplete evidence coverage, some data still need to rely on manual input, or manual intervention, introducing subjective bias, poor consistency, and therefore difficult to dynamically adjust the classification results.
[0004] The above content is only used to assist in understanding the technical solutions of the present application and does not represent the acknowledgement of the above content as prior art. SUMMARY
[0005] The main purpose of the present application is to provide a germline gene variation automatic classification method, device, equipment and storage medium, which aims to solve the technical problem that the ACMG evidence cannot be fully automatically evaluated in the prior art.
[0006] To achieve the above purpose, the present application provides a germline gene variation automatic classification method, which comprises:
[0007] Obtaining an input variation data set, the variation data set containing genes, transcripts, nucleotide changes, amino acid changes, variation types, chromosome positions, reference alleles, variation alleles, population frequencies, functional domain information, ClinVar records, functional experiment data, family pedigree data, phenotype data, computational prediction scores, ClinGen gene-specific guidelines and variation context;
[0008] parsing the variant data set to generate an initial evidence set, the initial evidence set containing status values of 28 ACMG evidence criteria;
[0009] optimizing conflicting evidence in the initial evidence set by a machine learning model to generate an optimized evidence set;
[0010] classifying the variant according to the optimized evidence set in combination with a Bayesian framework to output a classification result and a confidence level.
[0011] In an embodiment, the step of parsing the variant data set to generate an initial evidence set, the initial evidence set containing status values of 28 ACMG evidence criteria comprises:
[0012] extracting a variant type from the variant data set, determining whether the variant is nonsense, frameshift or large fragment deletion, if it meets and is located in ClinGen annotated loss-of-function sensitive gene, setting PVS1 status value to True;
[0013] comparing known pathogenic variants from the ClinVar record, if the variant causes the same amino acid change as the known pathogenic variant, setting PS1 status value to True;
[0014] analyzing the family pedigree data, if the parents do not carry the variant but the patient carries it, setting PS2 status value to True;
[0015] retrieving functional experiment data from PubMed and ClinVar, if there is clear pathogenicity experiment support, setting PS3 status value to True;
[0016] based on case-control database, calculating the frequency difference of the variant in case group and control group, if the difference is significant, setting PS4 status value to True;
[0017] determining pathogenic variants from ClinVar, if the pathogenic variants are different nucleotides but the same amino acids, setting PS5 status value to True;
[0018] obtaining functional domain information from UniProt database, if the variant is located in the key functional domain, setting PM1 status value to True;
[0019] querying gnomAD database, if the population frequency is less than 0.0001, setting PM2 status value to True;
[0020] analyzing the family pedigree data, if there is a trans variation in the recessive disease, setting PM3 status value to True;
[0021] Check for indels or stop loss in non-repetitive regions, if consistent, set PM4 status value to True;
[0022] Compare with ClinVar records, if different pathogenic variants exist at the same amino acid position, set PM5 status value to True;
[0023] Analyze the family pedigree data, if the patient carries it, set PM6 status value to True;
[0024] Use CADD and DeepSEA models to predict the impact of non-coding region variants, if the prediction result supports pathogenicity, set PM7 status value to True;
[0025] Calculate the LOD score based on the family pedigree data, if the score is greater than the preset threshold, set PP1 status value to True;
[0026] Check ClinGen records, if the gene is mainly dominated by missense variants, set PP2 status value to True; integrate SIFT, PolyPhen-2 and CADD prediction scores, if multiple tools consistently support pathogenicity, set PP3 status value to True;
[0027] Compare patient phenotypes with disease-related phenotypes in the HPO database, if the matching degree is higher than the preset threshold, set PP4 status value to True;
[0028] Query ClinVar records, if there are reliable pathogenic records, set PP5 status value to True;
[0029] Query gnomAD database, if the population frequency is greater than 0.05, set BA1 status value to True;
[0030] Based on population frequency, functional experiment and cosegregation data, reverse evaluate benign evidence, generate BS1 to BS5 status values;
[0031] Based on multi-tool prediction results and ClinGen records, generate BP1 to BP8 status values.
[0032] In an embodiment, the optimizing processing of the conflicting evidence in the initial evidence set by the machine learning model includes:
[0033] Construct a gradient boosting model, and the training data comes from the pathogenic and benign variants annotated in the ClinVar database;
[0034] Input the initial evidence set as input into the gradient boosting model, and the model outputs the priority of the conflicting evidence;
[0035] According to the priority adjustment of the state value of the evidence, an optimized evidence set is generated.
[0036] In an embodiment, the step of making a classification decision on the variant according to the optimized evidence set in combination with the Bayesian framework, and outputting a classification result and a confidence level comprises:
[0037] A pathogenicity score is calculated, which is a PVS1 weight multiplied by a state value plus a PS series weight multiplied by a sum of state values, and a benignity score is calculated, which is a BA1 weight multiplied by a state value plus a sum of BS series and BP series weights multiplied by a sum of state values.
[0038] A pathogenic probability is calculated based on the Bayesian framework, which is the pathogenicity score divided by a sum of the pathogenicity score and the benignity score.
[0039] According to the score and the pathogenic probability, a classification result and a confidence level value are determined, wherein the classification result is pathogenic, possibly pathogenic, of unknown significance, possibly benign, or benign.
[0040] In an embodiment, before the step of obtaining the input set of variant data, the method further comprises:
[0041] A minimum set of variant fields is determined, which includes gene, transcript, nucleotide change, amino acid change, variant type, chromosome location, reference allele, variant allele, population frequency, functional domain information, ClinVar record, functional experiment data, family pedigree data, phenotype data, calculated prediction score, ClinGen gene-specific guidelines, and variant context.
[0042] It is verified whether the input data meets the minimum set requirement, and if not, the user is prompted to supplement the missing fields.
[0043] In an embodiment, before the step of parsing the set of variant data to generate an initial evidence set, the method further comprises:
[0044] The state values of the 28 ACMG evidence criteria in the initial evidence set are initialized, and the state values of the 28 ACMG evidence criteria are set to False.
[0045] In an embodiment, the method for automatic classification of germline gene variants comprises:
[0046] The ClinVar database is connected through an API interface to obtain the latest pathogenicity and benignity records.
[0047] The gnomAD database is connected through the API interface to obtain the latest population frequency data.
[0048] Connect the PubMed database through the API interface to obtain the latest functional experimental data;
[0049] The updated data is re-input into the rule engine and machine learning model to generate new classification results.
[0050] In addition, to achieve the above-mentioned purpose, the application also provides an embryonic gene variation automatic classification device, which comprises:
[0051] A data input module is configured to obtain an input variation data set, wherein the variation data set comprises genes, transcripts, nucleotide changes, amino acid changes, variation types, chromosome positions, reference alleles, variation alleles, population frequencies, functional domain information, ClinVar records, functional experimental data, family pedigree data, phenotype data, computational prediction scores, ClinGen gene-specific guidelines, and variation contexts.
[0052] A data analysis module is configured to analyze the variation data set to generate an initial evidence set, wherein the initial evidence set comprises state values of 28 ACMG evidence standards.
[0053] An evidence evaluation module is configured to optimize conflicting evidence in the initial evidence set by a machine learning model to generate an optimized evidence set.
[0054] A classification decision module is configured to make a classification decision on the variation based on the optimized evidence set and a Bayesian framework, and output a classification result and a confidence level.
[0055] In addition, to achieve the above-mentioned purpose, the application also provides an embryonic gene variation automatic classification device, which comprises a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the embryonic gene variation automatic classification method as described above.
[0056] In addition, to achieve the above-mentioned purpose, the application also provides a storage medium, which is a computer-readable storage medium, and the storage medium stores a computer program, wherein the computer program is executed by a processor to implement the steps of the embryonic gene variation automatic classification method as described above.
[0057] In addition, to achieve the above-mentioned purpose, the application also provides a computer program product, which comprises a computer program, wherein the computer program is executed by a processor to implement the steps of the embryonic gene variation automatic classification method as described above.
[0058] The one or more technical solutions provided in the application have at least the following technical effects: by acquiring an inputted variant data set, analyzing the variant data set, generating an initial evidence set containing state values of 28 ACMG evidence standards, optimizing conflict evidence in the initial evidence set by a machine learning model to generate an optimized evidence set, classifying the variant according to the optimized evidence set in combination with a Bayesian framework, and outputting a classification result and a confidence level. The problems of incomplete ACMG evidence coverage, insufficient gene specificity, and limited dynamic updating capability in the prior art are solved, and efficient and standardized germline gene variant classification is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0059] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the application and serve to explain the principles of the application together with the specification.
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the accompanying drawings required to be used in the embodiments or prior art description will be briefly introduced. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative labor.
[0061] Figure 1 A flowchart is provided for the germline gene variant automatic classification method embodiment one of the application;
[0062] Figure 2 An algorithm flowchart is provided for the germline gene variant automatic classification method one embodiment of the application;
[0063] Figure 3 A module structure diagram of the germline gene variant automatic classification device of the embodiment of the application is provided;
[0064] Figure 4 A device structure diagram of the hardware running environment involved in the germline gene variant automatic classification method in the embodiment of the application is provided.
[0065] The purpose implementation, functional characteristics and advantages of the application will be further explained with reference to the accompanying drawings in combination with the embodiments. DETAILED DESCRIPTION
[0066] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the application and not to limit the application.
[0067] In order to better understand the technical solutions of the application, the following will be described in detail in combination with the drawings in the specification and specific embodiments.
[0068] The main solution of the embodiment of the application is:
[0069] In this embodiment, for ease of description, the following is described with the identification germline gene variation automatic classification device as the execution subject.
[0070] Since there are tools in the prior art that attempt to achieve automatic evaluation and variation classification of ACMG evidence, including InterVar, VarCards2, MAGI-ACMG, vaRHC and UMD Variant Interpretation Tool. These tools improve the classification efficiency to a certain extent, but still have limitations: incomplete evidence coverage, some data still need to rely on manual input, or manual intervention, introducing subjective bias, poor consistency, so it is difficult to dynamically adjust the classification results.
[0071] The present application provides a solution, by acquiring the input variation data set, parsing the variation data set, generating an initial evidence set, the initial evidence set contains the state value of 28 ACMG evidence standards, through the machine learning model, the conflicting evidence in the initial evidence set is optimized, and an optimized evidence set is generated, according to the optimized evidence set, combined with the Bayesian framework, the variation is classified and decided, and the classification result and the confidence are output. Solve the problem of incomplete ACMG evidence coverage, insufficient gene specificity and limited dynamic updating ability in the prior art, and realize efficient and standardized germline gene variation classification.
[0072] From the above embodiment, the present application acquires the input variation data set, parses the variation data set, generates an initial evidence set, the initial evidence set contains the state value of 28 ACMG evidence standards, through the machine learning model, the conflicting evidence in the initial evidence set is optimized, and an optimized evidence set is generated, according to the optimized evidence set, combined with the Bayesian framework, the variation is classified and decided, and the classification result and the confidence are output. Solve the problem of incomplete ACMG evidence coverage, insufficient gene specificity and limited dynamic updating ability in the prior art, and realize efficient and standardized germline gene variation classification.
[0073] It should be noted that the execution subject of the present embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, a germline gene variation automatic classification device, etc. The following will take the germline gene variation automatic classification device as an example to illustrate the present embodiment and the following embodiments.
[0074] Based on this, the present application embodiment provides a germline gene variation automatic classification method, referring to Figure 1 , Figure 1A flowchart of a first embodiment of the germline gene variant automatic classification method of the present application.
[0075] In this embodiment, the germline gene variant automatic classification method includes steps S10-S40:
[0076] Step S10, obtain the inputted variant data set, the variant data set contains genes, transcripts, nucleotide changes, amino acid changes, variant types, chromosome positions, reference alleles, variant alleles, population frequencies, domain information, ClinVar records, functional experiment data, family pedigree data, phenotype data, calculated prediction scores, ClinGen gene-specific guidelines, and variant context;
[0077] It should be noted that, with reference to Figure 2 , Figure 2For algorithm flowchart. The variant data set includes fields for supporting AMCG evidence evaluation, germline variants. Gene: the gene where the variant is located (e.g. BRCA1, CFTR). Transcript: the reference transcript ID (e.g. NM_000492.4). Nucleotide Change: the change at cDNA level (e.g. c.5165G>A). Amino Acid Change: the change at protein level (e.g. p.Arg1722Gln). Variant Type: missense, nonsense, frameshift, indel, splice, non-coding, etc. Chromosome Position: e.g. chr1:123456. Reference Allele: e.g. G. Alternate Allele: e.g. A. Allele Frequency (AF): population frequency (e.g. gnomAD AF=0.0001). Functional Domain: whether the variant is located in a key functional domain (e.g. UniProt annotated enzyme active site). ClinVar Annotation: known pathogenicity / benignity record (e.g. Pathogenic). Functional Evidence: e.g. in vitro experimental results (ClinVar, PubMed). Pedigree Data: PED file, containing parent and patient variant status. Phenotype Data: patient phenotype (HPO term, e.g. HP:0001428). In Silico Scores: e.g. SIFT, PolyPhen-2, CADD scores. ClinGen Gene-Specific Guidelines: e.g. PVS1 adjustment rule for BRCA1. Variant Context: e.g. whether it triggers nonsense-mediated mRNA decay (NMD). When inputting the variant data set, VCF file, PED file, phenotype data can be inputted. Pseudocode is shown as follows:
[0078] def AutoX(vcf_file, ped_file, phenotype_data).
[0079] In one possible implementation, before the acquiring the input variant data set, the method further comprises:
[0080] Determine the minimum set of variant fields, which includes genes, transcripts, nucleotide changes, amino acid changes, variant types, chromosomal locations, reference alleles, variant alleles, population frequencies, functional domain information, ClinVar records, functional experimental data, pedigree data, phenotypic data, calculated prediction scores, ClinGen gene-specific guidelines, and variant context.
[0081] Verify whether the input data meets the minimum set requirement. If not, prompt the user to add the missing fields.
[0082] In practical implementation, the completeness of the data needs to be checked before the user inputs the variant dataset. This involves determining the minimum set of variant fields, which represents the most basic and essential information fields that must be included when describing or analyzing gene variants. The minimum set includes gene, transcript, nucleotide changes, amino acid changes, variant type, chromosomal location, reference allele, variant allele, population frequency, functional domain information, ClinVar records, functional experimental data, pedigree data, phenotypic data, calculated prediction scores, ClinGen gene-specific guidelines, and variant context. During validation, the system checks whether the input data contains all of the above fields. If a field is missing, it is recorded. If a missing field is found, the system should prompt the user to supplement the missing information, clearly indicating the name of the missing field, for example: "Please supplement the amino acid change information of the variant." By validating the minimum set, the completeness of the input data can be ensured, avoiding misjudgments due to missing information. Complete data provides a more reliable basis for subsequent variant pathogenicity assessment and functional prediction.
[0083] Step S20: parse the mutation data set to generate an initial evidence set, which contains the state values of 28 ACMG evidence criteria.
[0084] It should be noted that the initial evidence set contains 28 ACMG evidence criteria status values, which can be divided into pathogenic evidence (PVS1, PS1-PS5, PM1-PM7, PP1-PP6) and benign evidence (BA1, BS1-BS5, BP1-BP8).
[0085] In the specific implementation, before parsing the variant data set and generating the initial evidence set, it is necessary to initialize the state values of the 28 ACMG evidence criteria in the initial evidence set by setting the state values of these 28 ACMG evidence criteria to False. The corresponding pseudocode can be represented as follows:
[0086] variants = preprocess_vcf(vcf_file) # use ANNOVAR / VEP to annotate;
[0087] pedigree = load_pedigree(ped_file) ;
[0088] phenotypes = parse_phenotypes(phenotype_data).
[0089] Input: VCF file, PED file, and phenotype data (HPO terms).
[0090] Extract fields using ANNOVAR and VEP: gene, transcript, nucleotide change, amino acid change, variant type, etc.
[0091] Query gnomAD for allele frequency, UniProt for domain information, ClinVar for known pathogenic records.
[0092] Upon initialization, evidence and classification results can be initialized. The pseudocode for the initialization process can be represented as:
[0093]
[0094]
[0095] In one possible implementation, the step of parsing the variant data set, generating an initial evidence set containing status values for 28 ACMG evidence criteria includes:
[0096] Extracting the variant type from the variant data set, determining whether the variant is nonsense, frameshift, or large fragment deletion, and if so and if it is located in a ClinGen-annotated loss-of-function intolerance gene, setting the PVS1 status value to True;
[0097] Comparing known pathogenic variants from the ClinVar records, and if the variant causes the same amino acid change as a known pathogenic variant, setting the PS1 status value to True;
[0098] Analyzing the family pedigree data, and if the parents do not carry the variant but the patient does, setting the PS2 status value to True;
[0099] Retrieving functional experiment data from PubMed and ClinVar, and if there is clear pathogenic experimental support, setting the PS3 status value to True;
[0100] Based on the case-control database, calculate the frequency difference of the variation in the case group and the control group, and if the difference is significant, set the PS4 status value to True;
[0101] Determine pathogenic variations from ClinVar, and if the pathogenic variations are different nucleotides but the same amino acid, set the PS5 status value to True;
[0102] Obtain functional domain information from the UniProt database, and if the variation is located in a key functional domain, set the PM1 status value to True;
[0103] Query the gnomAD database, and if the population frequency is less than 0.0001, set the PM2 status value to True;
[0104] Analyze the family pedigree data, and if there is a trans variation in a recessive disease, set the PM3 status value to True;
[0105] Check the insertion / deletion or loss of termination in non-repetitive regions, and if it is consistent, set the PM4 status value to True;
[0106] Compare ClinVar records, and if there are different pathogenic variations at the same amino acid position, set the PM5 status value to True;
[0107] Analyze the family pedigree data, and if the patient carries it, set the PM6 status value to True;
[0108] Use CADD and DeepSEA models to predict the impact of non-coding region variations, and if the prediction result supports pathogenicity, set the PM7 status value to True;
[0109] Calculate the LOD score based on the family pedigree data, and if the score is greater than the preset threshold, set the PP1 status value to True;
[0110] Check ClinGen records, and if the gene is mainly dominated by missense variations, set the PP2 status value to True; integrate SIFT, PolyPhen-2 and CADD prediction scores, and if multiple tools consistently support pathogenicity, set the PP3 status value to True;
[0111] Compare the patient's phenotype with the disease-related phenotypes in the HPO database, and if the matching degree is higher than the preset threshold, set the PP4 status value to True;
[0112] Query ClinVar records, and if there are reliable pathogenic records, set the PP5 status value to True;
[0113] Query gnomAD database, if population frequency is greater than 0.05, set BA1 status value to True;
[0114] Based on population frequency, functional experiment and co-segregation data, reverse evaluate benign evidence, generate BS1 to BS5 status value;
[0115] Based on multi-tool prediction results and ClinGen records, generate BP1 to BP8 status value.
[0116] In a specific implementation, the classification decision process for parsing the variant data set to generate the initial evidence set can be described as follows:
[0117] PVS1: Check if it is a loss-of-function variant (nonsense, frameshift, large deletion), combined with ClinGen gene-specific rules and NMD prediction.
[0118] PS1: Compare ClinVar to confirm if it causes the same amino acid change as the known pathogenic variant.
[0119] PS2: Analyze PED files to verify de novo mutations (parents have no mutations, patients carry).
[0120] PS3: Query ClinVar and PubMed to confirm functional experimental evidence.
[0121] PS4: Use case-control database to calculate the difference in mutation frequency (Fisher's exact test).
[0122] PS5: Check if it is a pathogenic variant in ClinVar with different nucleotides but the same amino acid.
[0123] PM1: Based on UniProt, determine if the variant is located in a key functional domain.
[0124] PM2: If AF < 0.0001 in gnomAD, assign PM2.
[0125] PM3: Analyze PED files to confirm anti-trans variants in recessive diseases.
[0126] PM4: Check for insertions / deletions or loss of termination in non-repetitive regions.
[0127] PM5: Compare ClinVar to confirm different pathogenic variants at the same amino acid position.
[0128] PM6: Similar to PS2, but without parental confirmation.
[0129] PM7: Use CADD, DeepSEA to predict the impact of non-coding region variants.
[0130] PP1: Calculate family-based LOD score based on PED file.
[0131] PP2: Check ClinGen record to confirm if gene is predominantly missense.
[0132] PP3: Integrate SIFT, PolyPhen-2, CADD predictions.
[0133] PP4: Compare patient phenotype to gene-associated diseases (HPO database).
[0134] PP5: Query ClinVar to confirm trusted pathogenic record.
[0135] BA1: Assign BA1 if AF > 0.05 (gnomAD).
[0136] BS1-BS5: Reverse evaluation based on population frequency, functional experiments, cosegregation, etc.
[0137] BP1-BP8: Evaluate benign evidence, such as BP4 (multiple tool predictions of no impact).
[0138] The corresponding pseudocode can be represented as:
[0139] evidence['PVS1'] = evaluate_pvs1(variant, clinvar, clingen_rules)
[0140] evidence['PS1'] = evaluate_ps1(variant, clinvar)
[0141] evidence['PS2'] = evaluate_ps2(variant, pedigree)
[0142] evidence['PS3'] = evaluate_ps3(variant, clinvar, pubmed)
[0143] evidence['PS4'] = evaluate_ps4(variant, case_control_db)
[0144] evidence['PS5'] = evaluate_ps5(variant, clinvar)
[0145] evidence['PM1'] = evaluate_pm1(variant, uniprot)
[0146] evidence['PM2'] = evaluate_pm2(variant, gnomad)
[0147] evidence['PM3'] = evaluate_pm3(variant, pedigree, clinvar)
[0148] evidence['PM4'] = evaluate_pm4(variant)
[0149] evidence['PM5'] = evaluate_pm5(variant, clinvar)
[0150] evidence['PM6'] = evaluate_pm6(variant, pedigree)
[0151] evidence['PM7'] = evaluate_pm7(variant, cadd, deepsea)
[0152] evidence['PP1'] = evaluate_pp1(variant, pedigree)
[0153] evidence['PP2'] = evaluate_pp2(variant, clingen)
[0154] evidence['PP3'] = evaluate_pp3(variant, sift, polyphen, cadd)
[0155] evidence['PP4'] = evaluate_pp4(variant, phenotypes, hpo_db)
[0156] evidence['PP5'] = evaluate_pp5(variant, clinvar)
[0157] evidence['BA1'] = evaluate_ba1(variant, gnomad)
[0158] evidence['BS1-BS5'] = evaluate_bs(variant, gnomad, clinvar, pedigree) evidence['BP1-BP8'] = evaluate_bp(variant, clingen, sift, polyphen, cadd)
[0159] Step S30, the conflicting evidence in the initial evidence set is optimized by a machine learning model to generate an optimized evidence set;
[0160] It should be noted that the machine learning model can learn the relationship between different evidences by learning a large number of known variant evidence sets, and optimize the conflicting evidence. The initial evidence set containing conflicting evidence is input into the trained machine learning model, and the model weighs and optimizes the conflicting evidence according to the learned pattern to generate a more reasonable classification result. The model not only outputs the final classification, but also adjusts the weight of the evidence or reinterprets the conflicting evidence to generate an optimized evidence set.
[0161] In a specific implementation, the conflicting evidence in the initial evidence set is optimized by a machine learning model to generate an optimized evidence set, including: constructing a gradient boosting model, and the training data is derived from the pathogenic and benign variants annotated in the ClinVar database; input the initial evidence set as input into the gradient boosting model, and the model outputs the priority of the conflicting evidence; adjust the state value of the conflicting evidence according to the priority to generate an optimized evidence set. Specifically, the gradient boosting model is a powerful machine learning algorithm that gradually constructs multiple weak learners and combines them to form a strong learner. The model uses the pathogenic and benign variants annotated in the ClinVar database as training data, extracts relevant features (such as variant type, population frequency, and functional prediction score), and trains the gradient boosting model. The model can output the priority of each evidence. Input the initial evidence set containing conflicting evidence into the model, adjust the state value of the conflicting evidence according to the priority output by the model, and finally generate an optimized evidence set. If there is evidence conflict (such as PM1 and BP4), use the gradient boosting model (trained on ClinVar data) to predict the final classification.
[0162] The process of solving evidence conflicts can be implemented with reference to the pseudo code:
[0163] conflict_resolved = resolve_conflicts(evidence, gradient_boosting_model)
[0164] In evidence evaluation, the evaluation function is:
[0165] def evaluate_pvs1(variant,clinvar,clingen_rules):
[0166] if variant['type'] in ['nonsense', 'frameshift', 'large_deletion']:
[0167] if variant['gene'] in clingen_rules['lof_sensitive']:
[0168] if triggers_nmd(variant['transcript'], variant['position']):
[0169] return True
[0170] return False
[0171] def evaluate_pm2(variant,gnomad):
[0172] if variant['allele_frequency'] < 0.0001:
[0173] return True
[0174] return False
[0175] Step S40, according to the optimized evidence set, combined with the Bayesian framework, the classification decision of the variation, output classification results and confidence.
[0176] It should be noted that the output classification results include pathogenic, possibly pathogenic, uncertain significance (VUS), possibly benign or benign. Confidence is the probability value of the output result.
[0177] In a specific implementation, the Bayesian framework uses prior probability and likelihood function, combined with each evidence in the optimized evidence set, to calculate the posterior probability of the variation belonging to different classifications (such as pathogenicity, benignity, etc.). By comparing these posterior probabilities, the final classification of the variation can be determined, and the corresponding confidence is output.
[0178] In one possible implementation, the classification decision of the variation according to the optimized evidence set combined with the Bayesian framework, the output classification results and the confidence include:
[0179] A pathogenicity score is calculated as PVS1 weight multiplied by state value plus PS series weight multiplied by state value sum, and a benignity score is calculated as BA1 weight multiplied by state value plus BS series and BP series weight multiplied by state value sum;
[0180] A pathogenic probability is calculated based on a Bayesian framework as the pathogenicity score divided by the sum of the pathogenicity score and the benignity score;
[0181] According to the score and the pathogenic probability, an output classification result and a confidence value are determined, and the output classification result is pathogenic, possibly pathogenic, of unknown significance, possibly benign or benign.
[0182] In a specific implementation, when classification is performed, pathogenicity and benignity scores are calculated according to ACMG combination rules (such as 1 PVS1+1 PS, or 2 PM), a pathogenic probability is calculated using a Bayesian framework, and VUS classification is solved. The pseudo code of the classification decision Bayesian framework is represented as:
[0183] classification,confidence=classify_variant(conflict_resolved,acmg_rules)
[0184] The pseudo code of the classification decision process is:
[0185]
[0186]
[0187] The Bayesian probability calculation process is:
[0188]
[0189] The pseudo code of the final output result can be represented as:
[0190]
[0191] In a feasible implementation, the germline gene variant automatic classification method comprises:
[0192] The ClinVar database is connected through an API interface to obtain the latest pathogenicity and benignity records;
[0193] The gnomAD database is connected through the API interface to obtain the latest population frequency data;
[0194] The PubMed database is connected through the API interface to obtain the latest functional experiment data;
[0195] The updated data is re-input into the rule engine and machine learning model to generate new classification results.
[0196] In a specific implementation, the latest evidence is queried in real time through an API, and the ClinGen rule base can be updated monthly to ensure gene specificity. For non-coding regions, CADD and DeepSEA models can be used to predict the impact of non-coding region variants on transcription and splicing, and assign PM7 or BP4. To ensure the accuracy and timeliness of gene variant analysis, ClinVar, gnomAD, and PubMed databases are connected through API interfaces to obtain the latest pathogenic and benign records, population frequency data, and functional experiment data. These updated data are integrated into the initial evidence set and re-input into the rule engine and machine learning model to generate new classification results. The rule engine classifies the variants according to predefined rules (such as the ACMG guidelines), while the machine learning model (such as a gradient boosting model) further optimizes the evidence set, resolves conflicting evidence, and improves classification accuracy. Finally, the posterior probability of the variant belonging to different classifications is calculated based on the Bayesian framework, and the new classification results and confidence are output. The final performance indicators are:
[0197] Accuracy: On the ClinVar validation set, the classification consistency is > 95%.
[0198] Coverage: 100% coverage of 28 ACMG evidence.
[0199] The pseudo-code for dynamic updating is represented as:
[0200]
[0201] The embodiment provides an automatic classification method for germline gene variants, which comprises the following steps: obtaining an input variant data set, parsing the variant data set to generate an initial evidence set, the initial evidence set containing state values of 28 ACMG evidence standards, optimizing the conflicting evidence in the initial evidence set through a machine learning model to generate an optimized evidence set, and classifying the variants based on the optimized evidence set and a Bayesian framework to output classification results and confidence. The problems of incomplete ACMG evidence coverage, insufficient gene specificity, and limited dynamic updating capability in the prior art are solved, and efficient and standardized classification of germline gene variants is achieved.
[0202] It should be noted that the above examples are only used for understanding the present application and do not constitute a limitation on the automatic classification method for germline gene variants of the present application. Further simple transformations based on this technical concept are within the scope of protection of the present application.
[0203] The present application also provides an automatic classification device for germline gene variants, which is described in detail as follows:Figure 3 The embryonic gene variation automatic classification device comprises:
[0204] A data input module 10 is configured to obtain an inputted variation data set, wherein the variation data set comprises genes, transcripts, nucleotide changes, amino acid changes, variation types, chromosome positions, reference alleles, variation alleles, population frequencies, domain information, ClinVar records, functional experiment data, family pedigree data, phenotype data, computational prediction scores, ClinGen gene-specific guidelines, and variation contexts.
[0205] A data analysis module 20 is configured to analyze the variation data set to generate an initial evidence set, wherein the initial evidence set comprises state values of 28 ACMG evidence criteria.
[0206] An evidence evaluation module 30 is configured to optimize conflicting evidence in the initial evidence set by a machine learning model to generate an optimized evidence set.
[0207] A classification decision module 40 is configured to make a classification decision on the variation based on the optimized evidence set in combination with a Bayesian framework to output a classification result and a confidence level.
[0208] In an embodiment, the data analysis module 20 is further configured to extract a variation type from the variation data set, determine whether the variation is nonsense, frameshift, or large fragment deletion, and if so and if the variation is located in a ClinGen-labeled loss-of-function sensitive gene, set a PVS1 state value to True.
[0209] Compare a known pathogenic variation from the ClinVar record, and if the variation causes the same amino acid change as the known pathogenic variation, set a PS1 state value to True.
[0210] Analyze the family pedigree data, and if the parents do not carry the variation but the patient does, set a PS2 state value to True.
[0211] Retrieve functional experiment data from PubMed and ClinVar, and if there is clear pathogenic experiment support, set a PS3 state value to True.
[0212] Based on a case-control database, calculate a frequency difference of the variation between a case group and a control group, and if the difference is significant, set a PS4 state value to True.
[0213] Determine a pathogenic variation from ClinVar, and if the pathogenic variation is different nucleotides but the same amino acid, set a PS5 state value to True.
[0214] Obtain functional domain information from UniProt database, if the variation is located within a key functional domain, set PM1 status value as True;
[0215] Query gnomAD database, if population frequency is less than 0.0001, set PM2 status value as True;
[0216] Analyze the family pedigree data, if there is a trans variation in the recessive disease, set PM3 status value as True;
[0217] Check the insertion / deletion or loss of termination in non-repetitive regions, if it is consistent, set PM4 status value as True;
[0218] Compare ClinVar records, if there are different pathogenic variations at the same amino acid position, set PM5 status value as True;
[0219] Analyze the family pedigree data, if the patient carries it, set PM6 status value as True;
[0220] Use CADD and DeepSEA models to predict the impact of non-coding region variations, if the prediction result supports pathogenicity, set PM7 status value as True;
[0221] Calculate the LOD score based on the family pedigree data, if the score is greater than the preset threshold, set PP1 status value as True;
[0222] Check ClinGen records, if the gene is mainly dominated by missense variations, set PP2 status value as True; integrate SIFT, PolyPhen-2 and CADD prediction scores, if multiple tools consistently support pathogenicity, set PP3 status value as True;
[0223] Compare patient phenotypes with disease-related phenotypes in the HPO database, if the matching degree is higher than the preset threshold, set PP4 status value as True;
[0224] Query ClinVar records, if there are reliable pathogenic records, set PP5 status value as True;
[0225] Query gnomAD database, if the population frequency is greater than 0.05, set BA1 status value as True;
[0226] Based on population frequency, functional experiment and cosegregation data, reverse evaluate benign evidence, generate BS1 to BS5 status values;
[0227] Based on multi-tool prediction results and ClinGen records, generate BP1 to BP8 status values.
[0228] In an embodiment, the evidence evaluation module 30 is further configured to build a gradient boosting model, and the training data is derived from ClinVar database annotated pathogenic and benign variants;
[0229] The initial evidence set is input into the gradient boosting model, and the model outputs the priority of conflicting evidence;
[0230] According to the priority, the state value of the conflicting evidence is adjusted to generate an optimized evidence set.
[0231] In an embodiment, the classification decision module 40 is further configured to calculate a pathogenic score and a benign score, wherein the pathogenic score is a PVS1 weight multiplied by a state value plus a PS series weight multiplied by a state value sum, and the benign score is a BA1 weight multiplied by a state value plus a BS series and BP series weight multiplied by a state value sum;
[0232] Based on the Bayesian framework, a pathogenic probability is calculated, which is the pathogenic score divided by the sum of the pathogenic score and the benign score;
[0233] According to the score and the pathogenic probability, an output classification result and a confidence value are determined, wherein the output classification result is pathogenic, possibly pathogenic, unknown significance, possibly benign, or benign.
[0234] In an embodiment, the data input module 10 is further configured to determine a minimum set of variant fields, which includes gene, transcript, nucleotide change, amino acid change, variant type, chromosome location, reference allele, variant allele, population frequency, functional domain information, ClinVar record, functional experiment data, family pedigree data, phenotype data, computational prediction score, ClinGen gene-specific guidelines, and variant context;
[0235] Verify whether the input data meets the minimum set requirement, and if not, prompt the user to supplement the missing fields.
[0236] In an embodiment, the data parsing module 20 is further configured to initialize the state values of the 28 ACMG evidence criteria in the initial evidence set, and set the state values of the 28 ACMG evidence criteria to False.
[0237] In an embodiment, the dynamic updating module 50 is further configured to connect the ClinVar database through an API interface to obtain the latest pathogenic and benign records;
[0238] Connect the gnomAD database through the API interface to obtain the latest population frequency data;
[0239] Connect the PubMed database through the API interface to obtain the latest functional experimental data.
[0240] The updated data is re-input into the rule engine and the machine learning model to generate a new classification result.
[0241] The embryonic gene variation automatic classification device provided in the present application adopts the embryonic gene variation automatic classification method in the above embodiments, and can solve the technical problem that ACMG evidence cannot be fully automatically evaluated in the prior art. Compared with the prior art, the embryonic gene variation automatic classification device provided in the present application has the same beneficial effects as the embryonic gene variation automatic classification method provided in the above embodiments, and other technical features in the embryonic gene variation automatic classification device are the same as the features disclosed in the above embodiment method, which will not be repeated here.
[0242] The present application provides an embryonic gene variation automatic classification device, which comprises at least one processor and a memory in communication connection with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the embryonic gene variation automatic classification method in the above embodiment one.
[0243] Reference will be made to the following Figure 4 which shows a structural schematic diagram of an embryonic gene variation automatic classification device suitable for being used to implement the embodiments of the present application. The embryonic gene variation automatic classification device in the embodiments of the present application can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description), PMPs (Portable Media Player), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. Figure 4 The shown embryonic gene variation automatic classification device is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present application.
[0244] As Figure 4As shown, the germline gene variation automatic classification device can include a processing device 1001 (for example, a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. In the RAM 1004, various programs and data required for the operation of the germline gene variation automatic classification device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; the storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the germline gene variation automatic classification device to communicate with other devices wirelessly or by wire to exchange data. Although the germline gene variation automatic classification device with various systems is shown in the figure, it should be understood that all the systems shown are not required to be implemented or possessed. More or fewer systems can be alternatively implemented or possessed.
[0245] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program codes for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through a communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are performed.
[0246] The germline gene variation automatic classification device provided by the present application adopts the germline gene variation automatic classification method in the above-mentioned embodiments, and can solve the technical problem that the ACMG evidence cannot be comprehensively and automatically evaluated in the prior art. Compared with the prior art, the germline gene variation automatic classification device provided by the present application has the same beneficial effects as the germline gene variation automatic classification method provided by the above-mentioned embodiments, and other technical features in the germline gene variation automatic classification device are the same as the features disclosed in the previous embodiment method, which will not be repeated here.
[0247] It should be understood that portions of the application disclosed can be implemented in hardware, software, firmware, or combinations thereof. In the description of the embodiments above, specific features, structures, materials or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0248] The above description is merely illustrative of the application and is not intended to limit the scope of the application. Any changes or modifications that can be made to the application in light of the teachings described herein would, however, be encompassed by the application. Consequently, the legal scope of the application is to be defined by the appended claims.
[0249] The application provides a computer readable storage medium having stored thereon computer readable program instructions (i.e., a computer program) for performing the germline genetic variation automatic classification method in the above-described embodiments.
[0250] The computer readable storage medium provided by the application may, for example, be a U disk, but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, system, or device, or any combination thereof. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more conductive wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present embodiment, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer readable storage medium can be transmitted in any suitable medium, including but not limited to electrical wire, optical cable, RF (Radio Frequency), etc., or any suitable combination of the above.
[0251] The above computer readable storage medium can be included in the germline genetic variation automatic classification device; or can exist separately and not be assembled into the germline genetic variation automatic classification device.
[0252] The above computer readable storage medium carries one or more programs, which, when executed by the germline genetic variation automatic classification device, cause the germline genetic variation automatic classification device to:
[0253] obtaining an inputted variant data set, the variant data set comprising genes, transcripts, nucleotide changes, amino acid changes, variant types, chromosome locations, reference alleles, variant alleles, population frequencies, domain information, ClinVar records, functional experimental data, family pedigree data, phenotype data, computational prediction scores, ClinGen gene-specific guidelines, and variant context;
[0254] parsing the variant data set to generate an initial evidence set, the initial evidence set comprising status values for 28 ACMG evidence criteria;
[0255] optimizing conflicting evidence in the initial evidence set by a machine learning model to generate an optimized evidence set;
[0256] making a classification decision on the variant based on the optimized evidence set in combination with a Bayesian framework, and outputting a classification result and a confidence score.
[0257] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0258] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or combinations of hardware and software.
[0259] The modules involved in the embodiments of the present application can be implemented in the form of software or in the form of hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.
[0260] The readable storage medium provided by the present application is a computer readable storage medium, which stores computer readable program instructions (i.e. computer program) for executing the above-mentioned germline gene variation automatic classification method, and can solve the technical problem that the ACMG evidence cannot be fully automatically evaluated in the prior art. Compared with the prior art, the computer readable storage medium provided by the present application has the same beneficial effects as the germline gene variation automatic classification method provided by the above-mentioned embodiments, which will not be repeated here.
[0261] The present application also provides a computer program product comprising a computer program, which, when executed by a processor, implements the steps of the above-mentioned germline gene variation automatic classification method.
[0262] The computer program product provided by the present application can solve the technical problem that the ACMG evidence cannot be fully automatically evaluated in the prior art. Compared with the prior art, the computer program product provided by the present application has the same beneficial effects as the germline gene variation automatic classification method provided by the above-mentioned embodiments, which will not be repeated here.
[0263] The above only describes some embodiments of the present application, and does not limit the patent scope of the present application, and any equivalent structural transformation made by using the content of the present application specification and drawings, or direct / indirect application in other related technical fields is included in the patent protection scope of the present application.
Claims
1. An automatic classification method for germline gene variations, characterized in that, The automatic classification method for germline gene variations includes: Obtain the input set of variant data, which includes genes, transcripts, nucleotide changes, amino acid changes, variant types, chromosome positions, reference alleles, variant alleles, population frequencies, functional domain information, ClinVar records, functional experimental data, pedigree data, phenotypic data, calculated prediction scores, ClinGen gene-specific guidelines, and variant context. The mutation data set is parsed to generate an initial evidence set, which contains the state values of 28 ACMG evidence criteria. The conflicting evidence in the initial evidence set is optimized using a machine learning model to generate an optimized evidence set. Based on the optimized evidence set, a classification decision is made on the variant using a Bayesian framework, and the classification result and confidence level are output.
2. The method as described in claim 1, characterized in that, The step of parsing the mutation data set to generate an initial evidence set, the initial evidence set containing state values of 28 ACMG evidence criteria, includes: Extract the mutation type from the mutation dataset, determine whether the mutation is nonsense, frame shift, or large fragment deletion, and if it matches and is located in the loss-of-function sensitive genes marked by ClinGen, then set the PVS1 status value to True. If a known pathogenic variant is compared with the ClinVar record, and the variant causes the same amino acid change as the known pathogenic variant, then the PS1 state value is set to True. Analyze the family pedigree data; if the parents do not carry the variant but the patient does, then set the PS2 status value to True. Retrieve functional experimental data from PubMed and ClinVar; if there is clear pathogenicity experimental support, set the PS3 status value to True. Based on the case-control database, the frequency difference of the variant in the case group and the control group is calculated. If the difference is significant, the PS4 status value is set to True. Pathogenic variants are identified from ClinVar. If the pathogenic variant is a different nucleotide but the same amino acid, the PS5 state value is set to True. Retrieve functional domain information from the UniProt database. If the mutation is located within a critical functional domain, set the PM1 state value to True. Query the gnomAD database; if the population frequency is less than 0.0001, set the PM2 state value to True. If the family pedigree data is analyzed and a trans variant is found in the recessive disease, the PM3 status value is set to True. Check for insertion / deletion or termination loss in non-repeating regions; if the condition is met, set the PM4 status value to True. By comparing ClinVar records, if different pathogenic variations exist at the same amino acid position, the PM5 status value is set to True; Analyze the family pedigree data; if the patient carries the virus, set the PM6 status value to True. The CADD and DeepSEA models are used to predict the impact of non-coding region variations. If the prediction results support pathogenicity, the PM7 state value is set to True. The LOD score is calculated based on family genealogy data. If the score is greater than a preset threshold, the PP1 status value is set to True. Examine the ClinGen records. If the gene is predominantly a missense variant, set the PP2 status value to True. Integrate the SIFT, PolyPhen-2, and CADD prediction scores. If multiple tools consistently support pathogenicity, set the PP3 status value to True. Compare the patient's phenotype with disease-related phenotypes in the HPO database. If the match is higher than a preset threshold, set the PP4 status value to True. Query ClinVar records; if a reliable pathogenic record exists, set the PP5 status value to True. Query the gnomAD database; if the population frequency is greater than 0.05, set the BA1 state value to True. Based on population frequency, functional experiments and co-segregation data, the benign evidence is evaluated in reverse to generate BS1 to BS5 state values. Based on the prediction results of multiple tools and ClinGen records, BP1 to BP8 state values are generated.
3. The method as described in claim 1, characterized in that, The step of optimizing conflicting evidence in the initial evidence set using a machine learning model to generate an optimized evidence set includes: A gradient boosting model was constructed, with training data derived from pathogenic and benign variants labeled in the ClinVar database; The initial set of evidence is used as input to the gradient boosting model, and the model outputs the priority of conflicting evidence. Adjust the state values of conflicting evidence according to the priority to generate an optimized evidence set.
4. The method as described in claim 1, characterized in that, The step of classifying variants based on the optimized evidence set and using a Bayesian framework, and outputting classification results and confidence scores, includes: Calculate the pathogenicity score and the benignity score. The pathogenicity score is the sum of the PVS1 weight multiplied by the state value and the PS series weight multiplied by the state value. The benignity score is the sum of the BA1 weight multiplied by the state value and the BS series and BP series weight multiplied by the state value. The pathogenicity probability is calculated based on a Bayesian framework, whereby the pathogenicity score is divided by the sum of the pathogenicity score and the benignity score. Based on the score and the pathogenicity probability, the output classification result and confidence value are determined, wherein the output classification result is pathogenic, possibly pathogenic, of unknown significance, possibly benign, or benign.
5. The method as described in claim 1, characterized in that, Before obtaining the input set of variant data, the process also includes: Determine the minimum set of variant fields, which includes genes, transcripts, nucleotide changes, amino acid changes, variant types, chromosomal locations, reference alleles, variant alleles, population frequencies, functional domain information, ClinVar records, functional experimental data, pedigree data, phenotypic data, calculated prediction scores, ClinGen gene-specific guidelines, and variant context. Verify whether the input data meets the minimum set requirement. If not, prompt the user to add the missing fields.
6. The method as described in claim 1, characterized in that, Before the step of parsing the mutation data set to generate an initial evidence set, the method further includes: The state values of the 28 ACMG evidence criteria in the initial evidence set are initialized by setting the state values of the 28 ACMG evidence criteria to False.
7. The method as described in claim 1, characterized in that, The automatic classification method for germline gene variations includes: Connect to the ClinVar database via the API interface to obtain the latest pathogenic and benign records; Connect to the gnomAD database through the API interface to obtain the latest population frequency data; Connect to the PubMed database through the API interface to obtain the latest functional experimental data; The updated data is then re-input into the rules engine and machine learning model to generate new classification results.
8. An automatic classification device for germline gene variations, characterized in that, The device includes: The data input module is used to acquire the input set of variant data, which includes genes, transcripts, nucleotide changes, amino acid changes, variant types, chromosome positions, reference alleles, variant alleles, population frequencies, functional domain information, ClinVar records, functional experimental data, pedigree data, phenotypic data, calculated prediction scores, ClinGen gene-specific guidelines, and variant context. The data parsing module is used to parse the variant data set and generate an initial evidence set, which contains the state values of 28 ACMG evidence criteria. The evidence evaluation module is used to optimize conflicting evidence in the initial evidence set using a machine learning model, and generate an optimized evidence set. The classification decision module is used to make classification decisions on the variants based on the optimized evidence set and in conjunction with the Bayesian framework, and output the classification results and confidence scores.
9. An automatic classification device for germline gene variations, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the automatic classification method for germline genetic variations as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the automatic classification method for germline gene variation as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Tumor-related gene variation pathogenicity classification method and device and storage medium
CN114882946A
Variant gene pathogenicity evaluation method and device based on Bayesian algorithm
CN116665773A
Automatic interpretation method and device for gene variation, equipment and storage medium
CN118762746A
Tumor whole course intelligent management platform and method
CN119314693A
Sequence variation analysis method and system, and storage medium
WO2023087277A1