A gene mutation grading method, system and storage medium

Through the comprehensive scoring module, the gene mutations are scored and pathogenicity judgments are solved, and the problems of high subjectivity and high cost in the grading of gene mutations in the prior art are achieved, and more objective and accurate gene mutation grading is achieved, especially in the grading of tumor somatic mutations.

CN118969077BActive Publication Date: 2025-05-23WUHAN KAIDEWEI MEDICAL LAB CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410869142.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-01
Publication Date
2025-05-23
Estimated Expiration
2044-07-01

AI Technical Summary

Technical Problem

The prior art has problems such as high subjectivity, high cost and grading differences between different units and personnel in the grading process of gene mutations. Especially in the grading of tumor somatic mutations, existing software cannot comprehensively evaluate the pathogenicity of SNV and Indel.

Method used

A gene mutation rating method is provided. The parameters of the preset population frequency calculation module, the list of class I mutations, the list of hot spots of gene mutations, the software scoring module, the database number of cases and OR value module, and the Cl inVar database module are comprehensively scored, and the pathogenicity level is determined based on the pathogenicity rating rules.

Benefits of technology

It improves the objectivity and accuracy of gene mutation classification, especially in the determination of Class I/II variants, and can be automated, reducing the subjectivity and cost of manual evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118969077B_ABST
    Figure CN118969077B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of gene detection technology, and in particular, to a gene mutation grading method, system and storage medium. The method comprises the steps of obtaining a gene mutation to be graded; scoring the gene mutation to be graded respectively according to a preset population frequency calculation module, a preset list of class I mutations, a preset list of hot spots of gene mutations, a preset software scoring module, a preset number of cases included in the database and an OR value module, and parameters of a Cl inVar database module; obtaining a total score of the gene mutation to be graded according to multiple scores; comparing the total score of the gene mutation to be graded with a preset pathogenicity grading rule to determine the pathogenicity grading result of the gene mutation to be graded. The grading result of the method is more objective and accurate, has a high determination accuracy for class I / class II mutations, and can also interpret frameshift mutations; and has a good auxiliary effect for the grading of gene mutations in clinical and scientific research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of gene detection technology, and in particular to a gene mutation grading method, system and storage medium. Background Art

[0002] High-throughput sequencing (NGS, also known as second-generation sequencing) has developed into one of the important technical means in the field of molecular diagnosis. NGS can obtain massive gene variation information in samples by sequencing dozens, hundreds or even whole genomes at one time. Through the analysis and research of this information, the disease gene characteristics of different patients can be grasped, thereby assisting clinical diagnosis and treatment decisions. However, separating useful information from massive variation information is still a difficult problem.

[0003] Genetic variation includes many types of variation, such as single nucleotide substitution (SNV), insertion and deletion (Indel), copy number variation (Copy Number Variation), gene rearrangement, etc. Among them, SNV and Indel are the most common and largest types of variation. In tumors, a one-time NGS test can obtain several or even dozens of SNVs and Indels (hereinafter referred to as mutations). By combining various knowledge base information (such as COSMIC, ClinVar, 1000G, etc.), literature reports, and software predictions (such as SIFT, PolyPhen2, etc.), and based on the tumor somatic mutation grading guidelines issued by ACMG / AMP (PMID: 27993330), such mutations can be divided into four levels: clear clinical significance (or pathogenic mutations / class I mutations), potential clinical significance (or possible pathogenic mutations / class II mutations), unknown clinical significance (or unknown significance mutations / class III mutations), and benign / likely benign (class IV mutations). In clinical practice, the first two levels of mutations are usually the focus.

[0004] However, the above guidelines only serve as an overall guide in the actual grading process. The guidelines mention that the basis for the classification of each level includes population frequency, mutation type, system / germline variation classification, database inclusion, software prediction, literature reporting, etc. In the process of grading a specific mutation, the grading personnel need to make judgments based on the above indicators, especially the literature reporting section, which requires the grading personnel to read the literature to evaluate the pathogenicity of the mutation, which is highly subjective. This process will lead to differences in the grading of the same mutation by different units and different personnel, and the need to read English literature makes the labor cost high.

[0005] There are some auxiliary grading software on the market, such as InterVar, but this software is only for germline mutation grading, and its pathogenicity determination needs to be included in the family analysis link, which is not suitable for tumor somatic mutation grading. In addition, there are similar patents in existing patents, including CN111063392A, CN117672382A, etc., but the former also targets germline mutations rather than somatic mutations, and the latter is only evaluated through software harmfulness prediction, and can only evaluate SNVs, but not Indel pathogenicity, which has certain limitations in clinical application. Summary of the invention

[0006] The technical problem to be solved by the present invention is to provide a gene mutation grading method, system and storage medium. The technical solution of the present invention to solve the above technical problem is as follows:

[0007] The present invention provides a method for grading gene mutations, comprising the following steps:

[0008] Obtaining gene mutations to be classified;

[0009] According to the parameters of the preset population frequency calculation module, the preset Class I mutation list, the preset gene mutation hotspot area list, the preset software scoring module, the preset database collection case number and OR value module, and the ClinVar database module, the gene mutation to be classified is scored respectively; and the total score of the gene mutation to be classified is obtained according to the multiple scores;

[0010] The total score of the gene mutation to be graded is compared with the preset pathogenicity grading rules to determine the pathogenicity grading result of the gene mutation to be graded.

[0011] Based on the above technical solution, the present invention can also be improved as follows.

[0012] Further, the gene mutation to be graded is scored according to a preset population frequency calculation module, including:

[0013] Acquire the preset population frequency calculation module according to multiple search libraries, wherein the preset population frequency calculation module includes first mutation information corresponding to all mutations of genes in the preset gene list recorded in the population frequency database and the maximum population frequency corresponding to each mutation;

[0014] According to the first mutation information of the plurality of genes, a first target gene mutation matching the mutation information of the gene mutation to be classified is obtained from the plurality of the genes, and according to the maximum population frequency of the first target gene mutation, a first score of the gene mutation to be classified is obtained.

[0015] Furthermore, the gene mutation to be classified is scored according to the preset Class I mutation list, including:

[0016] Obtaining the preset Class I mutation list according to authoritative guidelines, wherein the preset Class I mutation list includes mutation types, mutation sites and / or mutation intervals of multiple gene mutations;

[0017] The gene mutation to be classified is matched with the preset Class I mutation list, and a second score of the gene mutation to be classified is obtained according to the matching result.

[0018] Furthermore, the gene mutations to be graded are scored according to a preset list of gene mutation hotspots, including:

[0019] Acquire the preset gene mutation hotspot region list according to the cancer gene mutation database, wherein the preset gene mutation hotspot region list includes a mutation hotspot region collection;

[0020] The gene mutation to be classified is compared with the collection of mutation hotspot regions. When the gene mutation to be classified is located in any mutation hotspot region in the collection of mutation hotspot regions, the mutation type of the gene mutation to be classified is determined, and a third score of the gene mutation to be classified is obtained according to the mutation type.

[0021] Furthermore, the preset software scoring module includes a plurality of bioinformatics software, and the scoring of the gene mutation to be graded according to the preset software scoring module includes:

[0022] Using a plurality of the bioinformatics software to predict the harmfulness of the mutation site in the gene mutation to be classified, and obtaining a fourth score of the gene mutation to be classified according to the predicted mutation type and harmfulness of the mutation site;

[0023] Wherein, the mutation type is a missense mutation, a nonsense mutation or a splice mutation.

[0024] Furthermore, the gene mutation to be classified is scored according to the preset number of cases included in the database and the OR value module, including:

[0025] Obtaining the preset database collection case number and OR value module according to the gene mutation database, wherein the preset database collection case number and OR value module includes the second mutation information of multiple gene mutations, the database collection case number and the OR value score;

[0026] According to the second mutation information of the plurality of gene mutations, the gene mutation to be classified is matched with the preset database collection case number and OR value module to obtain a matched second target gene mutation;

[0027] The maximum value of the case number score of the second target gene mutation and the OR value score is used as the fifth score of the gene mutation to be classified.

[0028] Further, the gene mutation to be graded is scored according to the ClinVar database, including:

[0029] A sixth score of the gene mutation to be classified is obtained according to the collection result of the gene mutation to be classified in the ClinVar database.

[0030] Furthermore, the preset pathogenicity grading rule is to divide the disease into four score ranges, which are divided into Class I, Class II, Class III and Class IV according to the score ranges from high to low; the Class I variation indicates a variation with clear clinical significance or a pathogenic variation, the Class II variation indicates a variation with potential clinical significance or a possible pathogenic variation, the Class III variation indicates a variation with unknown clinical significance, and the Class IV variation indicates a benign or possibly benign variation.

[0031] The present invention also provides a gene mutation grading system, comprising:

[0032] An input unit, used to input a gene set to be graded;

[0033] The scoring acquisition unit scores the gene mutation to be classified according to the preset population frequency calculation module, the preset Class I mutation list, the preset gene mutation hotspot area list, the preset software scoring module, the preset database case number and OR value module, and the parameters of the ClinVar database module; and adds up a plurality of the scores to obtain a total score of the gene mutation to be classified;

[0034] The pathogenicity grading unit compares the total score of the gene mutation to be graded with a preset pathogenicity grading rule to determine the pathogenicity grading result of the gene mutation to be graded.

[0035] The present invention also provides a computer-readable storage medium, on which is stored a computer program programmed or configured to execute the above-mentioned gene mutation grading method.

[0036] The beneficial effects of the present invention are:

[0037] (1) The gene mutation grading method of the present invention makes the grading result of gene mutation more objective and accurate by comprehensively scoring the mutation population frequency, mutation type and mutation region determination, software scoring, number of cases included in the database, OR value module and ClinVar database inclusion;

[0038] (2) The gene mutation grading method of the present invention has a high accuracy in determining class I / II mutations and can also determine frameshift mutations (Inde l);

[0039] (3) After the preset parameter configuration is completed, the gene mutation grading method of the present invention has a simple and objective process when performing gene mutation grading, and does not require additional literature reading to conduct subjective evaluation of the harmfulness of the mutation;

[0040] (4) The gene mutation grading method of the present invention effectively realizes the automation of gene mutation grading and is an objective, accurate and efficient grading method;

[0041] (5) The gene mutation grading method of the present invention has a good auxiliary effect on gene mutation grading in clinical and scientific research. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 The present invention is a method for grading gene mutations, a schematic diagram of a process of grading gene mutations in one embodiment;

[0043] Figure 2 The method for grading gene mutations of the present invention is a schematic flow chart of another method for grading gene mutations provided in one embodiment. DETAILED DESCRIPTION

[0044] The principles and features of the present invention are described below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.

[0045] See also Figure 1 and Figure 2 The gene mutation grading method provided by the present invention comprises the following steps:

[0046] S10, obtaining the gene mutation to be graded;

[0047] S20, scoring the gene mutations to be classified respectively according to the preset population frequency calculation module, the preset Class I mutation list, the preset gene mutation hotspot area list, the preset software scoring module, the preset database case number and OR value module, and the parameters of the ClinVar database module; obtaining the total score of the gene mutation to be classified according to the multiple scores;

[0048] S30, comparing the total score of the gene mutation to be classified with the preset pathogenicity classification rule to determine the pathogenicity classification result of the gene mutation to be classified.

[0049] The gene mutation grading method of the present invention is an automated, multi-dimensional method for automatic classification and grading of tumor somatic gene mutations. This method specifically performs comprehensive scoring on the frequency of mutation populations, mutation types and mutation regions, software scoring, database collection cases and OR value modules, and ClinVar database collection. The device can grade the pathogenicity of tumor somatic mutations by only importing the public database mutation detection results of related tumors. After completing the preset parameter configuration, when grading the specific gene mutation to be graded, the process is simple and objective, and there is no need to perform additional literature reading to evaluate the harmfulness of the mutation. This method can well assist clinical and scientific researchers in determining the level of mutations, especially for Class I / Class II mutations with high determination accuracy. This method can interpret frameshift mutations (Inde l), and the grading results are more accurate and objective.

[0050] The gene mutation grading method of the present invention can score subsequent multiple related gene mutations to be graded after the initial preset configuration of each module is completed.

[0051] Preferably, the method of the present invention is applicable to mutation grading of gene mutations obtained after actual sequencing of a limited gene set (commonly referred to as "Panel" in the industry). When performing the test, the relevant gene list is a determined gene list.

[0052] Preferably, the preset pathogenicity grading rule is to divide the four score ranges into Class I, Class II, Class III and Class IV according to the score range from high to low; Class I variation indicates variation with clear clinical significance or pathogenic variation, Class II variation indicates variation with potential clinical significance or possible pathogenic variation, Class III variation indicates variation with unknown clinical significance, and Class IV variation indicates benign or possible benign variation. Then the specific determination method of the pathogenicity grading result of the above-mentioned gene mutation to be graded can be: determine the target score range in which the total score of the gene mutation to be graded is located, and determine the pathogenicity grading result corresponding to the target score range as the pathogenicity grading result of the gene mutation to be graded.

[0053] The data sources of each preset module used in the method of the present invention are specifically:

[0054] The crowd frequency calculation module is preset based on the crowd database. The crowd database used in the present invention can theoretically be various public databases; preferably, the crowd database used is ESP, 1000 Genomes (1000G), and gnomAD.

[0055] Class I mutation list: The data in this module are based on the gene mutations, gene mutation ranges or gene mutation types related to the tumors that are clearly indicated in the authoritative guidelines at home and abroad corresponding to the tumors with gene mutations to be classified. For example, when the gene mutations to be classified correspond to lung cancer, the data in this module can be derived from the mutations mentioned in the NCCN guidelines for lung cancer. In addition, the data in this module can be manually collected and entered by mutations in the guidelines.

[0056] List of hotspots of gene mutations. The data in this module is obtained based on the mutation data of corresponding cancer types in public databases. For example, when the gene mutation to be classified belongs to breast cancer, the data in this module can be selected from all mutation data of breast cancer in the COSMIC database, or the mutation data provided in the literature attachment / TCGA corresponding cancer mutation data and other public data.

[0057] Software scoring module, which uses a variety of bioinformatics software, including ClinPred, BayesDel, MCAP, MetaRNN, reveal, AlphaMissense, CADD database, etc.

[0058] The module for the number of cases and OR value included in the database is based on the public database data to calculate the number of cases and the OR value of the mutation position of the relevant gene recorded in the database. Specifically, the database used by this module can be the COSMIC database or other databases such as the mutation appendix of the literature or the TCGA database data.

[0059] ClinVar database. In this method, the gene mutation to be classified is determined using the information included in the ClinVar database.

[0060] The preset methods and specific scoring methods for each module are as follows:

[0061] (1) Scoring the gene mutation to be graded according to a preset population frequency calculation module, including:

[0062] A preset population frequency calculation module is obtained according to multiple search libraries, wherein the preset population frequency calculation module includes first mutation information corresponding to all mutations of genes in a preset gene list recorded in a population frequency database and a maximum population frequency corresponding to each gene mutation.

[0063] Preferably, the specific preset process is to download the three population databases, ESP, 1000 Genomes (1000G), and gnomAD, locally.

[0064] The maximum value of the population frequency of all mutations in the ESP database, the 1000 Genomes Project database (the version published in August 2015), and the gnomAD database, as well as the first mutation information, are used to form a population frequency search library. The first mutation information specifically includes chromosomes, gene mutation group coordinates, reference bases, and variant bases.

[0065] Taking the JAK2 gene chr9:5073770C>T mutation as an example, the maximum population frequency of this mutation in the three databases is 0.0007. The storage format of the information of this gene mutation is shown in Table 1.

[0066] Table 1 Information storage format of JAK2 gene chr9:5073770C>T mutation

[0067] Gene chromosome coordinate Ref Alt AF_max JAK2 9 5073770 C T 0.0007

[0068] After completing the above presets, based on the first mutation information of multiple genes, a first target gene mutation that matches the mutation information of the gene mutation to be classified is obtained from the multiple genes, and based on the maximum population frequency of the first target gene mutation, a first score of the gene mutation to be classified is obtained.

[0069] Preferably, the specific scoring process is that when the first mutation information of the first 5 columns in Table 1 are all matched, the value of the AF_max column in Table 1 is extracted, and it is compared with the corresponding tumor population incidence rate (referring to the population incidence rate of the detected tumor corresponding to the preset gene list) and the value of 1% (the tumor population incidence rate is less than 1%), and a score is given according to the comparison result.

[0070] Preferably, if the AF_max value is greater than 1%, the subsequent module scoring is stopped and the grading result is directly output as a Class IV variation, that is, the gene mutation is determined to be a possible benign / benign variation; if the AF_max value is between the incidence rate of the tumor population and 1%, the score is -1 point; if the AF_max value is less than the above-mentioned population incidence rate, the score is 1 point.

[0071] (2) Score the gene mutations to be classified according to the preset Class I mutation list, including:

[0072] The preset Class I mutation list is obtained according to authoritative guidelines. The preset Class I mutation list includes mutation types, mutation sites and / or mutation ranges of multiple gene mutations.

[0073] Among them, the Class I mutation list is the "white list" determined by this method. Mutations in this list will be assigned high scores to ensure that the corresponding gene mutations to be classified can tend to be judged as higher-level mutation types.

[0074] Preferably, the specific preset process is that this list is derived from the gene mutation, gene mutation range or gene mutation type related to the tumor clearly pointed out in the authoritative domestic and foreign guidelines corresponding to the preset tumor. According to the above database configuration, the class mutation list in the information record format shown in Table 2 is completed to preset the class I mutation list.

[0075] Table 2 Information recording format of Class I mutation list

[0076]

[0077] According to the mutation patterns of multiple gene mutations, the gene mutation to be classified is matched with a preset Class I mutation list, and a second score of the gene mutation to be classified is obtained according to the matching result.

[0078] Preferably, when judging a mutation to be classified, the information is matched with the list in Table 2 in the following manner: the name of the gene and the mutation pattern (i.e., the 1st and 2nd columns of Table 2) are completely matched; if the mutation column is not empty, the mutation column is also required to be completely matched; if the mutation column does not specify the mutated protein, only the protein position is matched. For example, in Table 2, when T315I is written as T315, the protein position is configured to be 315; if the base start and base end have values, the base change to be mutated must be within the range of the start position and the end position; if the protein (amino acid) start and end have values, the amino acid change to be mutated must be within the range of the start position and the end position.

[0079] Preferably, according to the above matching rules, when the gene mutation to be classified successfully matches all rows in the list, the gene mutation to be classified is scored 5 points; no points are scored in other cases. For example, for the p.Ile1236Leu mutation of the TET2 gene, when compared with Table 2, the gene name to which it belongs matches the first column, the mutation type is a missense mutation (i.e., Missense) matches the second column, and the protein mutation position is between 1134 and 1444 in the above table, therefore, the mutation is scored 5 points.

[0080] (3) Score the gene mutations to be graded based on a preset list of gene mutation hotspots, including:

[0081] A preset gene mutation hotspot region list is obtained according to a cancer gene mutation database, wherein the preset gene mutation hotspot region list includes a mutation hotspot region collection.

[0082] Preferably, the specific preset process is to select the mutation data of the corresponding cancer in the public database, download the data to the computer, and the data must include all the included mutations of the selected cancer, as well as the chromosome number where the mutation is located, the starting genome coordinates of the mutation, the ending genome coordinates of the mutation, the reference base of the mutation, the mutant base of the mutation, and the sample number corresponding to the mutation. Calculate and count each gene and obtain a collection of mutation hotspot areas. The calculation and statistical process is:

[0083] A. When presetting this module, you first need to define a variable value RMC (Region-Mutat ion count), which is defined as the number of mutations within a certain interval.

[0084] B. Segment the exon regions of all genes in the gene list one by one with 100 bp as an interval length (denoted as R 1 , R 2 …R N ), and count the RMC values ​​(denoted as RMC) in each interval of the corresponding gene in the mutation data set obtained from the database i ). After that, the following calculations are performed for each gene one by one:

[0085] B1. RMC to be obtained i After sorting in descending order, add them one by one from high to low. When the sum exceeds 80% of the total number of mutations in the gene, stop and output the corresponding segmentation region number at this time.

[0086] B2. Merge continuous segmented regions. For example, the result is R 3 , R 4 , R 6 , then R 3 , R 4 The region is merged (denoted as R 3-4 ). At this time, R 3-4 and R 6 Candidate regions for mutation hot spots of genes.

[0087] B3. Convergence of hotspot areas: For the areas obtained above, each area is converged according to the following process:

[0088] The regions were segmented at 10 bp intervals and the RMC values ​​were calculated;

[0089] Sort the RMC values ​​in descending order and add them up. When the sum exceeds 90% of the total number of mutations in the region, stop and merge the adjacent segmented regions.

[0090] The final area obtained by counting the above is the mutation hotspot area.

[0091] C. For the regions obtained by the above process, if the sum of the mutation hotspot regions of a gene exceeds 80% of the total coding region length of the gene, the gene is judged to have no hotspot region, otherwise it is considered to have a mutation hotspot region. The collection of mutation hotspot regions of all genes is the list of gene mutation hotspot regions corresponding to this module.

[0092] The gene mutation to be classified is compared with the collection of mutation hotspot regions. When the gene mutation to be classified is located in any mutation hotspot region in the collection of mutation hotspot regions, the mutation type of the gene mutation to be classified is determined, and a third score of the gene mutation to be classified is obtained according to the mutation type.

[0093] Preferably, when determining whether it is located in a mutation hotspot region, the specific method is as follows:

[0094] If the position of the gene mutation to be classified is within the list of mutation hotspots calculated for the corresponding gene, the mutation type will be determined. If the mutation type is a missense mutation, 1 point will be scored; if the mutation type is an inactivating mutation (including frameshift mutations, nonsense mutations, splice mutations, and start codon deletion mutations), 2 points will be scored; no points will be scored for other cases.

[0095] (4) Scoring the gene mutation to be graded according to a preset software scoring module, including:

[0096] The preset software scoring module uses a variety of bioinformatics software to score the harmfulness of a gene mutation to be classified, and performs scoring statistics to output the mutation bioinformatics harmfulness prediction results.

[0097] For missense mutations and nonsense mutations, six softwares were used for analysis: ClinPred, BayesDe l, MCAP, MetaRNN, revel, and AlphaMissense. Each software predicted the harmfulness of the mutation site separately. If the software judged it to be harmful, 1 point was recorded, and if it was judged to be harmless, -1 point was recorded. Then the scores of the six softwares were added together. If the result was greater than or equal to 2, 1 point was recorded; if the result was less than or equal to -2, -1 point was recorded; if it was less than 2 and greater than -2, no point was recorded.

[0098] For splice mutations, the CADD database was used for scoring. If the CADD database score was greater than or equal to 20, 1 point was given; if the result was less than 20, no point was given.

[0099] (5) Scoring the gene mutations to be classified according to the preset number of cases included in the database and the OR value module, including:

[0100] The module for obtaining the preset database inclusion cases and OR values according to the gene mutation database. The preset database inclusion cases and OR values module includes the second mutation information, case count scores, and OR value count scores of multiple gene mutations.

[0101] Preferably, the data basis of this module comes from the COSMIC database or other data such as literature mutation appendices or TCGA database data. In actual application, it is best to use two databases simultaneously. However, if one of the databases is lacking, it is also possible to use the other database for calculation and statistics.

[0102] Preferably, the preset process of this module is divided into COSMIC database hot spot calculation and OR value data preset.

[0103] A. COSMIC database hot spot calculation:

[0104] Extract all mutations of genes belonging to the detected gene list from the COSMIC database to form a mutation data set.

[0105] Extract the number of mutation cases of the corresponding cancer type in the mutation inclusion. For example, if the corresponding cancer type is hematological tumor and the example mutation is the c.G2503T(p.D835Y) mutation of the FLT3 gene, and its COSMIC record is {ID = COSV54042116; OCCURENCE = 384(haematopoiet ic_and_lymphoid_t i ssue)}, at this time, the extracted value should be the value 384 before the string “(haematopoiet ic_and_lymphoid_t i ssue)”. This data is the case count, and its meaning is that the COSMIC database has 384 reports of this mutation in hematological tumors.

[0106] Score all mutations to obtain the case count score. If the case count of this mutation is greater than 10, count 2 points; if the count is greater than or equal to 3 and less than or equal to 10, count 1 point; if the count is less than 3, count 0 points.

[0107] B. Literature appendix or TCGA database data statistics:

[0108] Use the literature mutation appendix or the TCGA corresponding tumor database mutation set as the case group and the gnomad2 database mutation set as the control group to calculate the OR value and 95% confidence interval (95% CI) for each mutation position of the gene. The calculation method is as follows:

[0109] For the case group (the case group refers to individuals suffering from the research disease), calculate the detection rate ratio of each mutation site (a / (a + b)) and the ratio of the corresponding mutation position in gnomad2 (c / (c + d)), where:

[0110] a: the number of cases with a specific mutation detected in the case group; b: the number of cases without the mutation in the case group;

[0111] c: The number of cases with the same mutation in a in the gnomad2 group; d: The number of cases without the same mutation in a in the gnomad2 group;

[0112] OR value = (a / (a+b)) / (c / (c+d)); standard error (SE) = sqrt((1 / a)+(1 / b)+(1 / c)+(1 / d));

[0113] Lower limit of 95% CI = logOR-1.96*SE; Upper limit of 95% CI = logOR+1.96*SE;

[0114] If the OR value is greater than 10 and 1 is not within the 95% CI interval, 2 points are scored; if the OR value is greater than or equal to 5 and less than 10 and 1 is not within the 95% CI interval, 1 point is scored; no points are scored in other cases. The OR value scores of each gene mutation are obtained by statistics.

[0115] C. Compare the proportion score and OR value score, take the maximum value, and store each gene mutation, genomic position, mutation base change, amino acid change, and maximum value in the format shown in Table 3.

[0116] Table 3. Number of cases included in the database and storage format of mutation information of OR value module

[0117] Gene chromosome coordinate Ref Alt Base changes Amino acid changes Number of cases Score (maximum value) JAK2 9 5073770 C T c.1849G>T p.Val617Phe 42380 2

[0118] According to the second mutation information of multiple gene mutations, the gene mutation to be classified is matched with the preset database collection case number and OR value module to obtain the matched second target gene mutation, and the fifth score of the gene mutation to be classified in this module is obtained.

[0119] Preferably, when matching, it needs to be consistent with the first 5 columns of information in Table 3. When the match is consistent, the corresponding score column value is output, which is the score of the gene mutation to be graded in this module. If no match is achieved, the default score result is 0 points.

[0120] (6) Scoring the gene mutation to be classified according to the ClinVar database, including:

[0121] The sixth score of the gene mutation to be classified is obtained according to the inclusion result of the gene mutation to be classified in the ClinVar database.

[0122] Preferably, the ClinVar database is downloaded. For a gene mutation in a certain gene list, if the ClinVar included result is Pathogen ic, Like ly_pathogen ic, Pathogen ic / Like ly_pathogen ic, 1 point is calculated; if the result is Begin, Like ly_ben ign, Ben ign / Like ly_ben ign, -1 point is calculated; if it is other results, no point is calculated.

[0123] After the modules of the present invention are preset, the gene mutations to be graded can be scored. Preferably, in the specific implementation process, for the data obtained from a specific NGS sequencing sample, the mutation information of the gene mutation is obtained using a conventional NGS analysis and annotation process, including chromosome number, mutation start genomic position, mutation end genomic position, mutation reference base, mutation variant base, mutation base change (such as: c.112C>T), mutation protein change (such as: p.T84M)

[0124] According to the aforementioned scoring method, each module is imported and run for scoring, and the total score after adding up the scores is used to determine the pathogenicity according to the following rules. Preferably, in the grading method of the present invention, the scoring ranges of each level and the corresponding grading are shown in Table 4.

[0125] Table 4 Scoring ranges and corresponding gradings at each level

[0126] Total score after addition determination ≥ 4 points Class I variants (clinically significant variants / pathogenic variants) 2 points ≤ Total score < 4 points Class II variants (potentially clinically significant variants / likely pathogenic variants) 0 points ≤ total score < 2 points Class III variants (clinical significance unknown) Total score < 0 points Class IV mutations (benign / likely benign mutations)

[0127] Based on the same principle as the above gene mutation grading method, the present invention also provides a gene mutation grading system, comprising:

[0128] An input unit, used for inputting a set of gene mutations to be graded;

[0129] The scoring acquisition unit scores the gene mutation to be classified according to the parameters of the preset population frequency calculation module, the preset Class I mutation list, the preset gene mutation hotspot area list, the preset software scoring module, the preset database case number and OR value module, and the ClinVar database module; and adds up a plurality of the scores to obtain a total score of the gene mutation to be classified;

[0130] The pathogenicity grading unit compares the total score of the gene mutation to be graded with a preset pathogenicity grading rule to determine the pathogenicity grading result of the gene mutation to be graded.

[0131] The computer-readable storage medium of the present invention stores a computer program programmed or configured to execute the gene mutation grading method of the present invention.

[0132] The gene mutation grading system of the embodiment of the present invention can execute the gene mutation grading method provided by the embodiment of the present invention, and its implementation principle is similar. The actions performed by each module and unit in the gene mutation grading system or device in each embodiment of the present invention correspond to the steps in the gene mutation grading method in each embodiment of the present invention. For the detailed functional description of each unit of the gene mutation grading system, please refer to the description in the corresponding gene mutation grading method shown in the previous text, which will not be repeated here.

[0133] The gene mutation grading system may be a computer program (including program code) running in a computer device, for example, the gene mutation grading system is an application software; the device may be used to execute the corresponding steps in the method provided in the embodiment of the present invention.

[0134] In some embodiments, the gene mutation grading system provided by the embodiment of the present invention can be implemented in a combination of software and hardware. As an example, the gene mutation grading system provided by the embodiment of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the gene mutation grading method provided by the embodiment of the present invention. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASIC, Application Specific Integrated Circuit), DSP, programmable logic device (PLD, Programmable Logic Device), complex programmable logic device (CPLD, Complex Programmable Logic Device), field programmable gate array (FPGA, Field-Programmable Gate Array) or other electronic components. In other embodiments, the gene mutation grading system provided by the embodiment of the present invention can be implemented in a software manner. The gene mutation grading system stored in the memory can be software in the form of programs and plug-ins, and includes a series of units for implementing the gene mutation grading method provided by the embodiment of the present invention.

[0135] The technical solution of the present invention is illustrated below by means of specific embodiments.

[0136] Example 1

[0137] This example takes myelodysplastic syndrome (MDS), a common subtype of blood tumors, as an example to verify the grading accuracy of the method of the present invention.

[0138] Specifically, this embodiment uses the results of mutation determination in the appendix of the MDS cluster typing document (PMID: 38319256) published in the journal NEJM Evidence as a reference standard. The document determines mutations as "Pathogen ic" and "VUS", Pathogen ic corresponds to the mutations determined as Class I and Class II in the present invention, and VUS corresponds to the mutations determined as Class III and Class IV in the present invention.

[0139] The grading process of this embodiment is as follows: Before grading, the gene list is first confirmed. The 31 genes involved in the clustering algorithm in the document are used as the initial gene list of the present invention, and the genes are shown in Table 5.

[0140] Table 5 List of genes in myelodysplastic syndrome (MDS)

[0141] ASXL1 BCOR BCORL1 CBL CEBPA DNMT3A ETNK1 ETV6 EZH2 FLT3 GATA2 GNB1 IDH1 IDH2 KRAS NF1 NPM1 NRAS PHF6 PPM1D PRPF8 PTPN11 RUNX1 SETBP1 SF3B1 SRSF2 STAG2 TP53 U2AF1 WT1

[0142] The gene mutation grading method of the present invention is used for scoring:

[0143] Each module is preset, wherein: in the population frequency calculation module preset in this embodiment, the tumor population incidence rate is set to the blood tumor incidence rate of 0.0062%. When scoring, the maximum population frequency of each gene mutation to be graded is compared with the incidence rate and 1% value for scoring.

[0144] The list of Class I mutations preset in this embodiment is based on the disease-related gene mutations and mutation types mentioned in the 2023 edition of the NCCN MDS Guidelines of the United States. The list of hotspot areas of gene mutations preset in this embodiment is constructed using data from the COSMIC database.

[0145] The data of the database collection case number and OR value module preset in this embodiment are derived from the COSMIC database data and the TCGA database data.

[0146] After the above preset configuration was completed, the pathogenicity of 7802 gene mutations in the gene list detected from 2797 MDS samples in the appendix of the NEJM Evidence magazine article (PMID: 38319256) was determined. The results are shown in Table 6, and the determination accuracy is shown in Table 7.

[0147] Table 6 Judgment results and comparison

[0148] Literature judgmentPathogenic Literature determination VUS Total number of mutations: The present invention is determined to be Class I / II 6748 261 7009 The present invention determines the class III / IV 357 436 793 Total number of mutations 7105 697 7802

[0149] Table 7 Determination accuracy

[0150] PPV 94.98% NPV 62.55% Sensitivity 96.28% Specificity 54.98%

[0151] As can be seen from the above, the present invention has good PPV and sensitivity, that is, it has strong accuracy in determining class I / II variant sites, but there may be misjudgments for class III / IV sites. In clinical or actual scientific research activities, the focus is on the association between class I / II sites and diseases, so the present invention can well assist in filtering the pathogenicity of mutation sites in clinical and scientific research.

[0152] Example 2

[0153] This embodiment uses acute myeloid leukemia (AML) as an example disease, and uses the system device of the present invention to grade related gene mutations, such as Figure 2 As shown, the specific process is as follows:

[0154] 1. Confirm the list of genes to be tested. Select common gene mutations in acute myeloid leukemia, as shown in Table 8.

[0155] Table 8 Common gene mutations in acute myeloid leukemia

[0156] ABL1 ASXL1 BCOR BCORL1 CALR CBL CEBPA CSF3R DNMT3A ETV6 EZH2 FLT3 GNAS GNB1 IDH1 IDH2 JAK2 KIT KRAS MPL NF1 NPM1 NRAS PIGA PPM1D PRPF40B RUNX1 SF3B1 SRSF2 TET2 TP53 U2AF1 WT1 ZRSR2

[0157] 2. Preset each module in the system.

[0158] (1) Preset population frequency calculation module: Download the three population databases, ESP, 1000 Genomes (1000G), and gnomAD, to the local computer. Construct the population frequency database according to the invention scheme. The incidence of AML is about 0.004%, so in this module, the incidence of tumor disease population is set to 0.004%.

[0159] (2) Preset Class I mutation list: According to the steps of the present invention, the mutations mentioned in the NCCN guidelines are collected and organized into a list, and the organized list is shown in Table 9.

[0160] Table 9 List of Class I mutations

[0161]

[0162]

[0163]

[0164] (3) Preset gene mutation hotspot area list: Using COSMIC database as input data, the hotspot area is determined according to the gene mutation hotspot area calculation step in the method of the present invention. The hotspot area list after calculation is shown in Table 10.

[0165] Table 10 List of gene mutation hotspots

[0166]

[0167]

[0168]

[0169]

[0170]

[0171]

[0172]

[0173]

[0174]

[0175]

[0176]

[0177]

[0178]

[0179]

[0180]

[0181] (4) Preset software scoring module: construct the software scoring module according to the method of the present invention.

[0182] (5) Module for presetting the number of cases included in the database and the OR value: COSMIC database data and TCGA public database data are selected as input data, and this module is constructed according to the preset method in the method of the present invention.

[0183] (6) ClinVar database module: The module is constructed according to point 6 of the invention scheme.

[0184] 3. Perform mutation grading

[0185] After the preset construction of the above modules is completed, the mutation data of 8 AML specimens NGS are used as the data to be judged and graded. SNPEFF or other similar conventional bioinformatics software is used to annotate the VCF files of the 8 sample data for gene mutations, and the base changes and amino acid changes of the gene mutations are obtained. The mutation annotations are imported into the above modules for scoring and grading, and the results are compared by manual grading. The mutation annotation results, the automatic grading results of the present invention, and the manual grading results are shown in Table 11.

[0186] Table 11AML mutation annotation results, automatic grading results and manual grading results

[0187]

[0188] There are a total of 14 mutations mentioned above, 11 mutations whose automatic grading results of the present invention are consistent with the manual results, and 3 mutations that are inconsistent, with a consistency of 78%. The mutations that do not conform to Class I are mainly mutations with less evidence. Such mutations are often difficult to grade with sufficient evidence, and are usually not mutations of focus. In general, the automatic determination results using the method of the present invention have a strong guiding role in assisting the determination of pathogenic mutations in clinical or scientific research.

[0189] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A gene mutation grading system, characterized in that: The gene mutation grading method is used for grading, and the method comprises the following steps: Obtaining gene mutations to be classified; According to the preset population frequency calculation module, the preset Class I mutation list, the preset gene mutation hotspot area list, the preset software scoring module, the database case number and OR value module, and the parameters of the ClinVar database module, the gene mutation to be classified is scored respectively; the total score of the gene mutation to be classified is obtained by adding up the scores; The step of scoring the gene mutation to be graded according to a preset population frequency calculation module includes: Acquire the preset population frequency calculation module according to multiple search libraries, wherein the preset population frequency calculation module includes first mutation information of multiple genes and the maximum population frequency corresponding to each gene mutation; According to the first mutation information of the plurality of genes, a first target gene mutation matching the mutation information of the gene mutation to be determined is obtained from the plurality of the genes, and according to the maximum population frequency of the first target gene mutation, a first score of the gene mutation to be determined is obtained; if the maximum population frequency is between the tumor population incidence rate and 1%, a score of -1 point is assigned; if the maximum population frequency is less than the above population incidence rate, a score of 1 point is assigned; The step of scoring the gene mutation to be classified according to the preset Class I mutation list includes: Obtaining the preset Class I mutation list according to the tumor gene mutation guide, wherein the preset Class I mutation list includes mutation types, mutation sites and / or mutation intervals of multiple gene mutations; According to the mutation types, mutation sites and / or mutation intervals of the multiple gene mutations, the gene mutation to be classified is matched with the preset Class I mutation list, and a second score of the gene mutation to be classified is obtained according to the matching result; when the gene mutation to be classified successfully matches a row in the preset Class I mutation list, 5 points are given to the gene mutation to be classified; no points are given for other cases; The step of scoring the gene mutation to be graded according to a preset list of gene mutation hot spots includes: Acquiring the preset gene mutation hotspot region list according to the cancer gene mutation database, wherein the preset gene mutation hotspot region list includes a collection of mutation hotspot regions; Comparing the gene mutation to be classified with the collection of mutation hotspot regions, when the gene mutation to be classified is located in any mutation hotspot region in the collection of mutation hotspot regions, determining the mutation type of the gene mutation to be classified, and obtaining a third score of the gene mutation to be classified according to the mutation type; The specific preset process of the preset gene mutation hotspot area list is to select the mutation data of the corresponding cancer type in the public database, download the mutation data to the computer, and the mutation data includes all the included mutations of the selected cancer type, the chromosome number where the mutation is located, the starting genome coordinates of the mutation, the ending genome coordinates of the mutation, the reference base of the mutation, the mutant base of the mutation, and the sample number corresponding to the mutation; Calculation and statistics are performed on each gene to obtain the mutation hotspot region collection, and the calculation and statistics process is as follows: A. When presetting this module, you first need to define a variable value RMC, which is defined as the number of mutations within a certain interval; B. Segment the exon regions of all genes in the gene list one by one with an interval length of 100 bp, denoted as R1, R2…R N , the RMC values ​​of the corresponding genes in each interval in the mutation data set obtained from the database are counted and recorded as RMC i ; Afterwards, the following calculations are performed for each gene one by one: B1. RMC to be obtained i After sorting in descending order, add them one by one from high to low. When the sum exceeds 80% of the total number of mutations in the gene, stop and output the corresponding segmentation region number at this time; B2. Merge the continuous segmented regions, merge the obtained continuous regions, and use the merged continuous and discontinuous regions as candidate regions of the mutation hotspot regions of the gene; B3. Convergence of hotspot areas: For the areas obtained above, each area is converged according to the following process: The regions were segmented at 10 bp intervals and the RMC values ​​were calculated; Sort the RMC values ​​in descending order and add them up. When the sum exceeds 90% of the total number of mutations in the region, stop and merge the adjacent segmented regions. The final area obtained by counting the above is the mutation hotspot area; C. For the regions obtained by the above process, if the sum of the mutation hotspot regions of a gene exceeds 80% of the total coding region length of the gene, the gene is judged to have no hotspot region, otherwise it is considered to have a mutation hotspot region. The collection of mutation hotspot regions of all genes is the list of gene mutation hotspot regions corresponding to this module; When the position of the gene mutation to be classified is within the mutation hotspot region list calculated for the corresponding gene, the mutation type is determined. If the mutation type is a missense mutation, 1 point is scored, if the mutation type is an inactivating mutation, 2 points are scored, and no points are scored for other cases. The preset software scoring module includes a plurality of bioinformatics software, and the scoring of the gene mutation to be graded according to the preset software scoring module includes: Using a plurality of the bioinformatics software to predict the harmfulness of the gene mutation to be classified, and obtaining a fourth score of the gene mutation to be classified according to the predicted mutation type and harmfulness of the mutation site; Wherein, the mutation type is a missense mutation, a nonsense mutation or a splice mutation; For missense mutations and nonsense mutations, six softwares, ClinPred, BayesDel, MCAP, MetaRNN, revel, and AlphaMissense, were used for analysis, and each software independently predicted the harmfulness of the mutation site; For splice subunit mutations, the CADD database was used for scoring; Each software predicts the harmfulness of the mutation site separately. If the software judges it to be harmful, 1 point is given, and if it is judged to be harmless, -1 point is given. Then the scores of the six software are added together. If the result is greater than or equal to 2, 1 point is given; if the result is less than or equal to -2, -1 point is given; if it is less than 2 and greater than -2, no point is given. For splice subunit mutations, the CADD database was used for scoring. If the CADD database score was greater than or equal to 20, 1 point was given; if the result was less than 20, no point was given; The step of scoring the gene mutation to be classified according to the preset number of cases collected in the database and the OR value module includes: Obtaining the preset database collection case number and OR value module according to the gene mutation database, wherein the preset database collection case number and OR value module includes second mutation information, case number score values ​​and OR value score values ​​of multiple gene mutations; According to the second mutation information of the plurality of gene mutations, the gene mutation to be classified is matched with the preset database collection case number and OR value module to obtain a matched second target gene mutation; The maximum value of the case number score of the second target gene mutation and the OR value score is used as the fifth score of the gene mutation to be classified; The gene mutation database includes COSMIC database and TCGA database; For COSMIC database hot spots, if the number of cases of the mutation is greater than 10, 2 points are given; if the count is greater than or equal to 3 and less than or equal to 10, 1 point is given; For the TCGA database data statistics, if the count is less than 3, it is scored as 0; The scoring of the comparison ratio and OR value is as follows: if the OR value is greater than 10 and 1 is not within the 95% CI interval, it is scored as 2 points; if the OR value is greater than or equal to 5 and less than 10 and 1 is not within the 95% CI interval, it is scored as 1 point; no points are given in other cases; The step of scoring the gene mutation to be graded according to the ClinVar database comprises: Obtaining a sixth score for the gene mutation to be classified according to the inclusion result of the gene mutation to be classified in the ClinVar database; If the ClinVar result is Pathogenic, Likely_pathogenic, or Pathogenic / Likely_pathogenic, 1 point will be given; if the result is Begin, Likely_benign, or Benign / Likely_benign, -1 point will be given; if it is other results, no point will be given; Comparing the total score of the gene mutation to be classified with the preset pathogenicity classification rules to determine the pathogenicity classification result of the gene mutation to be classified; The preset pathogenicity grading rule is to divide the mutation into four score ranges, which are divided into Class I, Class II, Class III and Class IV according to the score ranges from high to low; Class I mutations are represented as clinically significant mutations or pathogenic mutations, Class II mutations are represented as potential clinically significant mutations or possible pathogenic mutations, Class III mutations are represented as mutations of unknown clinical significance, and Class IV mutations are represented as benign or possibly benign mutations; The scoring ranges and corresponding gradings for each level are as follows: The system comprises: An input unit, used for inputting a set of gene mutations to be graded; The scoring acquisition unit scores the gene mutation to be classified according to the parameters of the preset population frequency calculation module, the preset Class I mutation list, the preset gene mutation hotspot area list, the preset software scoring module, the preset database case number and OR value module, and the ClinVar database module; and adds up a plurality of the scores to obtain a total score of the gene mutation to be classified; The pathogenicity grading unit compares the total score of the gene mutation to be graded with a preset pathogenicity grading rule to determine the pathogenicity grading result of the gene mutation to be graded.

2. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program programmed or configured to execute the gene mutation grading method as claimed in claim 1.

Citation Information

Patent Citations

  • Gene missense mutation pathogenicity prediction system based on deep learning

    CN117672382A

  • Genetic variation determination method and system and storage medium

    CN109243530A

  • Gene mutation pathogenicity detection method and system based on neural network and medium

    CN111063392A