A method, system and apparatus for classifying a MODY type of diabetes

By combining clinical physiological and biochemical indicators with DNA sequencing data, the problem of misdiagnosis of MODY type diabetes has been solved, enabling accurate diabetes classification and the provision of treatment plans.

CN113870947BActive Publication Date: 2026-02-10石家庄市第二医院 +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111299859.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-04
Publication Date
2026-02-10
Estimated Expiration
2041-11-04

AI Technical Summary

Technical Problem

Existing methods for diagnosing and classifying MODY type diabetes have a high misdiagnosis rate and cannot accurately determine the type of diabetes through gene sequencing data, thus making it impossible to provide appropriate treatment.

Method used

By combining clinical physiological and biochemical indicators and DNA sequencing data, and utilizing multiple variant databases and variant site hazard detection tools, the probability of MODY type diabetes is calculated and statistically tested by comprehensively analyzing and scoring the patient's gene sequencing data annotation information and clinical physiological and biochemical data, thus achieving accurate subtyping.

Benefits of technology

It improves the diagnostic accuracy of MODY type diabetes, reduces the misdiagnosis rate, enables more precise treatment plans, and improves diagnostic efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113870947B_ABST
    Figure CN113870947B_ABST
Patent Text Reader

Abstract

The application discloses a MODY type diabetes typing method, system and device, and the method comprises the following steps: acquiring a gene variation data file, analyzing and screening pathogenic genes, scoring clinical indexes, data testing and probability calculation. The system comprises a variation data file acquisition subsystem, a pathogenic gene analysis and screening subsystem, a diabetes physiological and biochemical index scoring subsystem, a MODY type diabetes typing result calculation subsystem and the like. The device comprises a data acquisition device, a processor and a memory. Based on clinical physiological and biochemical indexes and DNA sequencing data, by using a plurality of variation databases and variation site hazard detection tools, the gene sequencing data annotation information and the clinical physiological and biochemical data of a patient are comprehensively analyzed and scored, the probability of MODY is calculated, and the significance of the pathogenic genes related to MODY type diabetes is statistically tested by using a database, so that the accurate typing of MODY type diabetes is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of diabetes classification, in particular to a MODY type diabetes classification method, system and device. BACKGROUND

[0002] Diabetes mellitus is increasing globally and has become a serious public health problem. According to the conventional classification, most of them are Type 1 Diabetes Mellitus (T1DM) and Type 2 Diabetes Mellitus (T2DM). In addition, there is a kind of diabetes caused by single gene mutation. According to the age of onset of clinical symptoms, single gene diabetes can be divided into two categories: Neonatal Diabetes Mellitus (NDM) and Maturity Onset Diabetes of Young (MODY).

[0003] Single gene diabetes is a special type of diabetes caused by mutation of a single gene that plays a key role in the development, function or insulin signaling pathway of pancreatic beta cells, accounting for about 1%-5% of all types of diabetes, which brings great burden to families and society. MODY type diabetes is an autosomal dominant single gene hereditary disease, that is, a disease or pathological trait controlled by a pair of alleles. A class of hereditary diseases that can cause disease as long as a single gene mutation occurs in the human body, which further leads to beta cell dysfunction, and is commonly seen in children and adolescents. Clinically, only according to the physiological and biochemical indicators of patients, MODY type diabetes is often misdiagnosed as T1DM or T2DM type diabetes, which leads to the failure to obtain the best symptomatic treatment.

[0004] MODY is the most common type of single gene diabetes, and its diagnosis and classification depend on gene detection. Based on gene detection technology, 14 genes have been confirmed to be associated with MODY, including HNF4α-MODY1, GCK-MODY2, HNF1α-MODY3, PDX1-MODY4, HNF1β-MODY5, NEUROD1-MODY6, KLF11-MODY7, CEL-MODY8, PAX4-MODY9, INS-MODY10, BLK-MODY11, ABCC8-MODY12, KCNJ11-MODY13, APPL1-MODY14, etc.

[0005] Although the gene detection technology can be used to determine which genes have been mutated, the degree of mutation and the importance of the mutation site are different, resulting in different results. In the field of life sciences, genetic variation is divided into pathogenic, likely pathogenic, uncertain significance, likely benign, and benign. Therefore, only according to the gene sequencing related genetic variation information, the specific type of MODY type diabetes cannot be accurately determined. Accurate diagnosis of pathogenic genes of monogenic genetic diseases can bring precise treatment plan in some cases, greatly relieving or curing the disease of the patient to some extent. Therefore, accurate diagnosis of pathogenic genes of monogenic diabetes mellitus has great significance and clinical value in clinical practice

[0006] Since the age of onset, clinical manifestations and treatment response to drugs of patients with different types of MODY are not the same, it is particularly important to distinguish MODY from T1DM and T2DM, correctly diagnose and timely give appropriate treatment, evaluate the complications, prognosis and analyze and genetic counseling for family members of patients, and carry out individualized diagnosis and treatment.

[0007] In summary, the existing MODY type diabetes based on physiological and biochemical indicators of patients has the defect of misdiagnosis, and the existing diagnosis and further typing of MODY type diabetes based on gene sequencing data have the problem of inaccuracy, therefore, it is necessary to develop a MODY type diabetes typing method, system and device, which provides a high-efficiency, practical and fast MODY type diabetes typing tool for relevant medical staff, researchers and other related personnel. SUMMARY

[0008] Therefore, the present application provides a MODY type diabetes typing method, system and device, which is based on clinical physiological and biochemical indicators and DNA sequencing data, uses a variety of variation databases and variation site hazard detection tools, and comprehensively analyzes and scores the gene sequencing data annotation information and clinical physiological and biochemical data of patients to calculate the probability of MODY, and uses the database to statistically test the significance of the pathogenic genes related to MODY type diabetes, thereby realizing accurate typing of MODY type diabetes.

[0009] In order to achieve the above purpose, the present application provides the following technical solutions:

[0010] A MODY type diabetes typing method, the method comprising:

[0011] Obtaining the variation data file of the patient;

[0012] filtering the mutation quality, mutation frequency and mutation type of each mutation site in the mutation data file to obtain high-quality mutation sites, aligning and analyzing the high-quality mutation sites by using a genetic disease diagnosis database and a mutation hazard prediction software, annotating the mutation sites to obtain mutation gene information, and classifying and scoring the mutation gene information to obtain scoring data of the mutation gene information;

[0013] obtaining clinical physiological and biochemical indexes of a patient, scoring the clinical physiological and biochemical indexes to obtain scoring data of the clinical physiological and biochemical indexes;

[0014] processing data in a diabetes epidemiology data set to obtain a score distribution of the data in the diabetes epidemiology data set, testing the scoring data to determine whether the scoring data is consistent with the bias of the MODY type diabetes distribution, and if so, calculating a MODY type diabetes probability and combining the mutation gene information to obtain a MODY type.

[0015] Preferably, the obtaining of the mutation data file of the patient specifically comprises: extracting MODY related gene sequence information by using human reference genome information, constructing MODY gene reference sequence information and an index file thereof from the MODY related gene sequence information, reading gene sequencing data of the patient, performing quality control on the gene sequencing data, and aligning the gene sequencing data with the MODY gene reference sequence to obtain the mutation data file of the patient.

[0016] Preferably, the genetic disease diagnosis database comprises InterVar, ClinVar, HGMD and OMIM databases; and the mutation hazard prediction software comprises SIFT, PolyPhen2 and gerp++ software.

[0017] Preferably, the clinical physiological and biochemical indexes of the patient comprise BMI, age, family information, diabetes onset age, disease duration, glycosylated hemoglobin, fasting blood glucose, C peptide and urea nitrogen creatinine ratio, the scoring of the clinical physiological and biochemical indexes is weighted scoring according to the contribution of the clinical physiological and biochemical indexes, and the scoring data of the clinical physiological and biochemical indexes and the scoring data of the mutation gene information are summarized.

[0018] Preferably, the testing of the scoring data comprises one-sided testing of Kolmogorov-Smirnov testing; the scoring data comprises scoring data of clinical physiological and biochemical indexes, or scoring data of clinical physiological and biochemical indexes and scoring data of mutation gene information; and the calculation of the MODY type diabetes probability adopts a Bayesian formula.

[0019] The application further provides a MODY type diabetes typing system, which comprises:

[0020] a variant data file acquisition subsystem configured to acquire a variant data file of a patient;

[0021] a pathogenic gene analysis and screening subsystem configured to filter a variant quality, a mutation frequency and a variant type of each variant site in the variant data file, to obtain high-quality variant sites, to perform alignment and analysis on the high-quality variant sites by using a genetic disease diagnosis database and a variant hazard prediction software, to obtain variant gene information by annotating the variant sites, to obtain scoring data of the variant gene information by classifying and scoring the variant gene information, and to obtain a MODY type of diabetes by calculating a probability of MODY type diabetes.

[0022] a diabetes physiological and biochemical index scoring subsystem configured to acquire clinical physiological and biochemical indexes of a patient, to score the clinical physiological and biochemical indexes, and to obtain scoring data of the clinical physiological and biochemical indexes;

[0023] a MODY type diabetes classification result calculation subsystem configured to process physiological and biochemical indexes in a diabetes epidemiology data set, to obtain a score distribution of the physiological and biochemical indexes in the diabetes epidemiology data set, to perform a test on the scoring data, to determine whether the scoring data is consistent with a bias of a MODY type diabetes distribution, to calculate a probability of MODY type diabetes if the scoring data is consistent with the bias of the MODY type diabetes distribution, and to obtain a MODY type in combination with the variant gene information.

[0024] Preferably, the function of the variant data file acquisition subsystem specifically comprises the following steps: extracting a MODY related gene sequence by using human reference genome information, constructing a MODY gene reference sequence and an index file thereof by using the MODY related gene sequence, reading gene sequencing data of a patient, performing quality control on the gene sequencing data, and performing alignment on the gene sequencing data and the MODY gene reference sequence to obtain a variant data file of the patient.

[0025] Preferably, the function of the diabetes physiological and biochemical index scoring subsystem further comprises the following steps: performing weighted scoring according to a contribution degree of the clinical physiological and biochemical indexes, and performing aggregation on the scoring data of the clinical physiological and biochemical indexes and the scoring data of the variant gene information.

[0026] The application further provides an application of a MODY type diabetes classification system.

[0027] The application further provides a MODY type diabetes classification device, which comprises a data acquisition device, a memory and a processor.

[0028] The data collection device is used to collect data required for the MODY type diabetes classification; the memory is used to store one or more program instructions for executing the method as described above; and the processor is used to execute one or more program instructions for executing the method as described above.

[0029] The present application also provides a computer readable storage medium comprising one or more program instructions, which, when executed on a computer, cause the computer to execute the method as described above.

[0030] In one or more specific embodiments, the MODY type diabetes classification method, system and device provided by the present application have the following technical effects: ① Based on clinical physiological and biochemical indicators and DNA sequencing data, a variety of variant databases and variant site hazard detection tools are used to comprehensively analyze and score the gene sequencing data annotation information and clinical physiological and biochemical data of the patient, calculate the probability of MODY, and statistically test the significance of the pathogenic gene related to MODY type diabetes, so as to realize the accurate classification of MODY type diabetes, especially the accurate classification of early-onset diabetes; ② The method, system and device of the present application can not only use whole exome sequencing data, but also use whole genome or target region sequencing and other different ways of second-generation sequencing data; ③ Compared with the prior art, the present application has higher efficiency and more accurate detection results, and can achieve better use effect. BRIEF DESCRIPTION OF DRAWINGS

[0031] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings needed in the following embodiment or prior art description will be briefly introduced. Obviously, the drawings in the following description are only exemplary, and for those skilled in the art, other implementation drawings can be obtained without creative labor on the basis of the provided drawings.

[0032] The structures, proportions, sizes, etc. shown in the specification are only used to cooperate with the content disclosed in the specification, to be understood and read by those skilled in the art, and do not define the limiting conditions for the implementation of the present application, so they do not have technical significance. Any modification of structure, change of proportion relationship or adjustment of size, without affecting the effect and purpose that can be achieved by the present application, should still fall within the scope of the technical content disclosed by the present application.

[0033] Figure 1 Flowchart of the MODY type diabetes classification method provided by the present application;

[0034] Figure 2 Flowchart of one embodiment of the MODY type diabetes classification method provided by the present application;

[0035] Figure 3 A schematic diagram of the MODY type diabetes classification system provided by the present application. DETAILED DESCRIPTION

[0036] The embodiments of the present application will be described in detail by specific embodiments, and those skilled in the art can easily understand other advantages and effects of the present application from the disclosure. Obviously, the described embodiments are part of the embodiments of the present application, not all. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0037] The MODY type diabetes classification method, system and device provided by the present application are based on clinical physiological and biochemical indicators and DNA sequencing data, utilize various variant databases and variant site hazard detection tools, comprehensively analyze and score the gene sequencing data annotation information and clinical physiological and biochemical data of the patient, calculate the probability of MODY, and statistically test the significance of the pathogenic gene related to MODY type diabetes, so as to realize the accurate classification of MODY type diabetes.

[0038] In one embodiment, the present application provides the following technical solutions:

[0039] A MODY type diabetes classification method, as shown in the figure, the method comprises: Figure 1

[0040] Obtaining the variant data file (VCF) of the patient;

[0041] Filtering the mutation quality, mutation frequency and mutation type of each variant site in the variant data file to obtain high-quality variant sites, using genetic disease diagnosis database and variant hazard prediction software to align and analyze the high-quality variant sites, annotating the variant sites to obtain variant gene information, classifying and scoring the variant gene information to obtain scoring data of the variant gene information;

[0042] Obtaining the clinical physiological and biochemical indicators of the patient, scoring the clinical physiological and biochemical indicators to obtain scoring data of the clinical physiological and biochemical indicators;

[0043] Processing the data in the diabetes epidemiology data set to obtain the score distribution of the data in the diabetes epidemiology data set, testing the scoring data to determine whether the scoring data conforms to the bias of the MODY type diabetes distribution, if it conforms to the distribution of the MODY type diabetes, calculating the probability of the MODY type diabetes, and combining the variant gene information to obtain the MODY type.

[0044] ​In an embodiment, the variant quality, mutation frequency and variant type in the variant data file are filtered to retain high-quality variant site information, genetic disease diagnosis databases such as InterVar, ClinVar, HGMD and OMIM are integrated, and variant sites are annotated by combining variant harmfulness prediction software such as SIFT, PolyPhen2 and gerp++, to obtain variant gene information, and the variant gene information is classified and scored to obtain scoring data of the variant gene information.

[0045] In an embodiment, physiological and biochemical indicators in a diabetes epidemiology data set of a certain region are processed to obtain a score distribution of the physiological and biochemical indicators in the diabetes epidemiology data set, form a database, and store the database in a local computing device as a local database.

[0046] In an embodiment, various clinical data in a MODY type diabetes epidemiology data set of a certain region are processed to obtain a score distribution of the clinical data in the MODY type diabetes epidemiology data set, form a database, and store the database in a local computing device as a local database.

[0047] In an embodiment, the database formed by the physiological and biochemical indicators is used to calculate the prevalence of MODY type diabetes; or the database formed by the physiological and biochemical indicators and the database formed by the MODY type diabetes epidemiology data are used to calculate the prevalence of MODY type diabetes.

[0048] In an embodiment, as shown in Figure 2 In an embodiment, the obtaining of the variant data file of the patient specifically includes: extracting MODY related gene sequence information by using human reference genome information, constructing MODY gene reference sequence information and an index file thereof from the MODY related gene sequence information; reading gene sequencing data of the patient, performing quality control on the gene sequencing data, and aligning the gene sequencing data with the MODY gene reference sequence to obtain the variant data file of the patient.

[0049] In an embodiment, the human reference genome information is human reference genome (hg38) information.

[0050] In an embodiment, the gene sequencing data of the patient includes whole genome sequencing data, target region sequencing data, and also includes whole exon or other related second-generation sequencing data.

[0051] In an embodiment, the genetic disease diagnosis database includes InterVar, ClinVar, HGMD and OMIM databases; and the variant harmfulness prediction software includes SIFT, PolyPhen2 and gerp++ software.

[0052] In an embodiment, the clinical physiological and biochemical indicators of the patient include BMI, age, family information, diabetes onset age, duration of illness, glycosylated hemoglobin, fasting blood glucose, C-peptide, and urea nitrogen creatinine ratio. The scoring of the clinical physiological and biochemical indicators is weighted scoring according to the contribution of the clinical physiological and biochemical indicators. The scoring data of the clinical physiological and biochemical indicators and the scoring data of the variant gene information are summarized.

[0053] In an embodiment, the indicators of the patient are scored according to the variant information and physiological and biochemical indicator scoring rules listed in Table 1, and then the scoring data is summarized. In an embodiment, the scoring data is summarized as follows: the scoring data is calculated by weighted summation. The weights of interVar, ClinVar, HGMD, onset age, and family history of diabetes (the items marked with "*" in the table) are 10%, 10%, 15%, 10%, and 15%, respectively, and the weights of the remaining items are all 4%. The final score is obtained by multiplying the scores of the items by the corresponding weights and summing them up.

[0054] In an embodiment, the verification of the scoring data includes one-sided Kolmogorov-Smirnov test. The scoring data includes the scoring data of the clinical physiological and biochemical indicators, or the scoring data of the clinical physiological and biochemical indicators and the scoring data of the variant gene information. The calculation of the MODY diabetes probability uses the Bayes formula.

[0055] In addition to the above method, the present application also provides a MODY diabetes typing system, as shown in Figure 3 The system comprises:

[0056] The variant data file acquisition subsystem is configured to acquire the variant data file of the patient.

[0057] The pathogenic gene analysis and screening subsystem is configured to filter the variant quality, mutation frequency, and variant type of each variant site in the variant data file to obtain high-quality variant sites, compare and analyze the high-quality variant sites using a genetic disease diagnosis database and variant hazard prediction software, annotate the variant sites to obtain variant gene information, classify and score the variant gene information to obtain scoring data of the variant gene information.

[0058] The diabetes physiological and biochemical indicator scoring subsystem is configured to acquire the clinical physiological and biochemical indicators of the patient, score the clinical physiological and biochemical indicators, and obtain scoring data of the clinical physiological and biochemical indicators.

[0059] The MODY type diabetes typing result calculation subsystem is used for processing physiological and biochemical indexes in a diabetes epidemiology data set, obtaining a score distribution of the physiological and biochemical indexes in the diabetes epidemiology data set, testing the scoring data, judging whether the scoring data conforms to the bias of the MODY type diabetes distribution, calculating a MODY type diabetes probability if the scoring data conforms to the bias of the MODY type diabetes distribution, and obtaining a MODY type in combination with variant gene information.

[0060] In a specific embodiment, the typing system of the present application has a general computer structure, has general CPU, memory, display and other conventional computer hardware devices, and can run a general operating system to access network resources such as the Internet, can run the above-mentioned typing methods, and achieves the purpose of the present application.

[0061] In a specific embodiment, the function of the variant data file acquisition subsystem specifically comprises: extracting a MODY related gene sequence by using human reference genome information, constructing a MODY gene reference sequence and an index file thereof from the MODY related gene sequence, reading gene sequencing data of a patient, performing quality control on the gene sequencing data, aligning the gene sequencing data with the MODY gene reference sequence, and obtaining a variant data file of the patient.

[0062] In a specific embodiment, the function of the diabetes physiological and biochemical index scoring subsystem further comprises weighting scoring according to the contribution degree of the clinical physiological and biochemical indexes, and aggregating the scoring data of the clinical physiological and biochemical indexes and the scoring data of the variant gene information.

[0063] Based on the same technical concept, the present application further provides an application of a MODY type diabetes typing system, which is used for the development and application of bioinformatics software.

[0064] In a specific embodiment, the steps of implementing gene variant detection filtering, clinical data scoring, distribution testing and probability calculation in the typing system of the present application are constructed by using a programming language such as Python to build a pre-loadable module, which can be directly called and maintained in analysis, can be used for calling by other bioinformatics software, and is used for the development and application of these software.

[0065] Based on the same technical concept, the present application further provides a MODY type diabetes typing device, which comprises a data acquisition device, a memory and a processor.

[0066] The data acquisition device is used for acquiring data required for MODY type diabetes typing; the memory is used for storing one or more program instructions for executing the method as described above; and the processor is used for executing the one or more program instructions for executing the method as described above.

[0067] In a specific embodiment, the data required for the typing of MODY type diabetes includes patient gene sequencing data, patient physiological and biochemical indicators and other clinical data, and also includes data of a certain region diabetes epidemiology data set, data of a certain region MODY type diabetes epidemiology data set, and the like.

[0068] Based on the same technical concept, the present application also provides a computer readable storage medium, which comprises one or more program instructions, when the program instructions are run on a computer, the computer executes the method as described above.

[0069] The present application is further described in detail in the following specific embodiment, so that those skilled in the art can better understand the present application and implement it. In the present embodiment, the methods and default parameters of the prior art are used unless otherwise specified.

[0070] The present embodiment uses the typing system of the present application to run the typing method of the present application, and achieves the corresponding technical effects. The typing system uses a Linux-based computing server, the server is Intel Xeon CPU E5-2650, the memory is 512GB, the hard disk is 20TB, and the general display and input / output devices are used.

[0071] The specific process is as follows:

[0072] 1. Obtain the patient's variant data file by reference genome alignment of whole exome sequencing data and variant information annotation and filtering;

[0073] 1a) Extract the sequence of 14 MODY genes (as shown in Table 2) in the human genome (hg38) sequence and merge it into a reference sequence, and use BWA (Burrows-Wheeler-Alignment) to construct the merged reference sequence into an index file;

[0074] 1b) The input file is whole exome sequencing data of second-generation sequencing, and Trimmomatic software is used to remove the adapter sequence and low-quality data in the original sequencing data to obtain clean reads;

[0075] 1c) Use BWA to align the data to the above-mentioned MODY reference sequence to generate an alignment file; preferably, in order to ensure better hard disk space utilization, the alignment file can be compressed into bam format;

[0076] 1d) Use the GATK software package to detect variant information on the MODY gene. First, use the program's DuplicatesMarking to mark or delete duplicate sequences during GATK analysis. Then, use RealignerTargetCreator to locally realign reads near the indel to reduce the probability of alignment errors. Finally, use HaplotypeCaller to detect variants and output a variant result file (VCF file).

[0077] It should be noted that in step 1a, the index file only needs to be built the first time it is used, and subsequent analyses directly use BWA to call the built index file; in step 1d, preferably, BaseRecalibrator can be used to recalibrate the base quality values ​​of the reads in the bam file to ensure that the base quality values ​​of the reads in the final output bam file are closer to the actual probability of mismatch with the reference genome.

[0078] 2. Filter, compare, analyze, annotate, and score the contents of the variant data file to obtain scored data of variant gene information.

[0079] This step includes a series of subroutines for variant annotation and screening, used to automatically screen variant sites in pathogenic genes from VCF files. The filtering is based on the quality of variants, mutation frequency, and variant type in the variant data files, retaining high-quality variant site information. It integrates genetic disease diagnostic databases such as InterVar, ClinVar, HGMD, and OMIM, and combines variant hazard prediction software such as SIFT, PolyPhen2, and gerp++ to annotate variant sites, thereby quickly classifying and scoring patient variant information.

[0080] 2a) Quality control: The analysis module filters out low-quality variant sites based on the score output by the GATK software for each variant site, sequencing depth, location and type of the variant site;

[0081] 2b) Mutation frequency screening: Mutation frequency annotation is performed using internal databases of public database sets such as dbSNP, 1000Genome and ExAC. The distribution frequency of each site in the population is calculated, and mutation sites with a minimum mutation frequency (MAF) of less than 0.05 are automatically screened.

[0082] 2c) Pathogenicity screening: using HGMD, ClinVar, OMIM, interVar and other public or commercial pathogenic literature databases for comparison, screening the remaining literature reported pathogenic and potentially pathogenic variation sites, at the same time, this step also uses SIFT, PolyPhen2 and gerp++ and other variation hazard prediction software to annotate the variation sites that have significant regulation on protein function;

[0083] 2d) Variation type screening: according to the variation information, screening dominant mutations, if there is variation data of the sample parents, the variation sites that do not meet the genetic mode will be automatically filtered.

[0084] 3. According to the variation information data and clinical physiological and biochemical indicators of the patient, scoring is carried out, and scoring data of mutation sites and clinical physiological and biochemical indicators are obtained

[0085] The variation data and physiological and biochemical indicators of the patient will be scored according to the following table, and the score is divided into different categories of variation, physiology and biochemistry, which will affect the subsequent weighted score. The scoring rules are as follows:

[0086] Table 1 Scoring rules of variation information and physiological and biochemical indicators

[0087]

[0088]

[0089] The projects marked with "*" in the table will be weighted in the subsequent combination of variation data scoring to process the total score.

[0090] 4. Process the epidemiological data set to obtain the scoring distribution of the related indicators, test the scoring data, calculate the probability, and obtain the MODY type

[0091] Using 1000Genome or internal database, the pathogenic variation sites in each healthy person's genome are screened. For each healthy person's pathogenic variation site, the typing system of the present application calculates its frequency in the healthy population, on this basis, the cumulative probability distribution of the frequency of pathogenic variation sites is calculated, using the cumulative probability distribution, the cumulative probability corresponding to the frequency of the pathogenic variation site screened in the patient is inferred, and recorded as the p value of the variation site.

[0092] 4a) Clinical data is entered using text documents by corresponding items, and the typing system automatically calls VCF annotation results and clinical data for automatic scoring; if it contains a three-generation direct relative diabetes history and an onset age < 25 years, the corresponding index will be scored with a weighting of 0.8, and the scores are added; the variant data is scored according to the disease database annotation, with +3 points for each pathogenic site, +1 point for a potential pathogenic site, and +0 points for others; the site hazard annotation is scored according to each harmful site +3 points, potential harmful site +1 point, and other +0 points;

[0093] 4b) Using 1000Genome, ExAC and other databases, for each healthy person's pathogenic variant site, the system of the present application will calculate its frequency in the healthy population, and on this basis, calculate the cumulative probability distribution of the frequency of pathogenic variant sites, use K-S test probability distribution to infer the cumulative probability corresponding to the frequency of pathogenic variant sites screened in patients, and record it as the p value of the variant site;

[0094] 4c) Using the local diabetes epidemiology database, the probability of the patient having MODY type diabetes k when the patient has clinical phenotype a is calculated by the Bayesian model. The specific calculation is as follows: define a certain phenotype of diabetes as Diabetes a , then the probability can be written as conditional probability P(MDOY k |Diabetes a ), which is calculated by the Bayes formula:

[0095]

[0096] The joint probability P(MDOY k ,Diabetes a ) and P(Diabetes a ) are calculated, wherein P(Diabetes a ) is the probability of the patient showing a certain defined phenotype of Diabetes a , and P(MDOY k |Diabetes a ) represents the conditional probability of MDOY type diabetes MDOY a when the patient shows a certain defined phenotype of Diabetes k .

[0097] 4d) Finally, according to the probability, it is judged whether the input data is a MODY case, if so, output all the screened site corresponding genes as candidate pathogenic genes, and define the p value as the p value of the variant site on it, and sort according to the p value, the smaller the p value, the greater the possibility of pathogenic gene.

[0098] In summary, the typing system of the present application can automatically integrate the weighted scoring from the patient clinical data and variation data, calculate the probability of suffering from MODY type diabetes using the Bayesian formula, and use the K-S test to statistically determine the p value of the pathogenic gene, output the ranking of the candidate pathogenic gene based on the patient clinical data and whole exome sequencing data, the smaller the p value, the greater the possibility of the gene being a pathogenic gene, and finally accurately determine the MODY type. The system of the present application has higher efficiency and more accurate detection results compared with the prior art, thereby achieving better clinical use effect.

[0099] Table 2 14 type pathogenic genes related to MODY type diabetes

[0100]

[0101]

[0102] The typing system is different from the clinical diagnosis and simple gene detection method. Compared with the diagnosis based on clinical phenotype, the system can reduce the misdiagnosis probability of MODY type diabetes by using the clinical data scoring system and distribution model. Compared with the simple gene detection method, the system can help to determine the pathogenic gene site and probability of MODY by using clinical data and statistical test, and can more accurately provide a treatment plan.

[0103] Each embodiment of the above method in the specification is described in a progressive manner, and the same and similar parts between each embodiment can be referred to each other. Each embodiment mainly describes the difference from other embodiments. The related parts can refer to the part of the method embodiment.

[0104] It should be noted that although the operations of the method of the present application are described in a specific order in the accompanying drawings, this does not require or imply that the operations must be performed in this specific order, or that all of the shown operations must be performed to achieve the desired result. Additionally or alternatively, certain steps can be omitted, multiple steps can be combined into one step, and / or one step can be divided into multiple steps.

[0105] Although the present application provides method operational steps in the form of embodiments or flowcharts, more or fewer operational steps can be included based on conventional or non-creative means. The order in which the steps are listed in the embodiments is only one of the many possible execution orders of the steps, and does not represent the only execution order. In actual device or client product execution, the method order shown in the embodiments can be executed in sequence or in parallel (for example, in a parallel processor or multi-threaded processing environment, or even in a distributed data processing environment). The terms "comprise", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, product or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, product or device. Without more limitations, it does not exclude the presence of other same or equivalent elements in the process, method, product or device including the elements.

[0106] The units, devices or modules and the like illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. For the convenience of description, the above devices are described as various modules respectively described in function. Of course, in the implementation of the present application, the functions of each module can be implemented in the same or more software and / or hardware, or the modules implementing the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The above described device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the coupling or direct coupling or communication connection between the displayed or discussed each other can be through some interface, indirect coupling or communication connection between devices or units, which can be electrical, mechanical or other forms.

[0107] Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer readable program code, the same function can also be implemented by logically programming the method steps in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers. Therefore, such a controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the devices for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0108] The application can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, classes, etc., that perform particular tasks or implement particular abstract data types. The application can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including memory storage devices.

[0109] From the above description of the embodiments of the application, those skilled in the art can clearly understand that the application can be implemented by means of software plus necessary universal hardware platforms. Based on such an understanding, the technical solutions of the application can be embodied in a software product form, and the computer software product can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, or an optical disk, and includes a plurality of instructions to cause a computer device (which can be a personal computer, a mobile terminal, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the application.

[0110] The various embodiments in the specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. The application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable electronic devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, etc.

[0111] The above specific embodiments further illustrate the purpose, technical solutions and beneficial effects of the application. It should be understood that the above description is only for specific embodiments of the application and is not used to limit the protection scope of the application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the application should be included in the protection scope of the application.

Claims

1. A method for classifying MODY type diabetes, characterized in that, The method includes: Obtain the patient's variant data file; The mutation quality, mutation frequency, and mutation type of each mutation site in the mutation data file are filtered to obtain high-quality mutation sites. The high-quality mutation sites are compared and analyzed using a genetic disease diagnosis database and mutation hazard prediction software. The mutation sites are annotated to obtain mutation gene information. The mutation gene information is classified and scored to obtain the scoring data of the mutation gene information. Obtain the patient's clinical physiological and biochemical indicators, score the clinical physiological and biochemical indicators, and obtain the scoring data of the clinical physiological and biochemical indicators; The data in the diabetes epidemiology dataset is processed to obtain the score distribution of the data in the diabetes epidemiology dataset. The score data including clinical physiological and biochemical indicators and the score data of variant gene information are examined to determine whether they conform to the bias of the distribution of MODY type diabetes. If they conform to the distribution of MODY type diabetes, the probability of MODY type diabetes is calculated, and the MODY type is obtained by combining the variant gene information. The patient's clinical physiological and biochemical indicators include BMI, age, family information, age of onset of diabetes, duration of disease, glycated hemoglobin, fasting blood glucose, and C-peptide to urea nitrogen-creatinine ratio. The clinical physiological and biochemical indicators are scored by weighting the scores based on their contribution. The scores of the clinical physiological and biochemical indicators are then combined with the scores of the mutated gene information.

2. The MODY type diabetes classification method as described in claim 1, characterized in that, The specific files for obtaining patient variation data include: MODY-related gene sequence information was extracted using human reference genome information, and the MODY-related gene sequence information was used to construct MODY gene reference sequence information and its index file. The patient's gene sequencing data is read, the gene sequencing data is quality controlled, and compared with the MODY gene reference sequence to obtain the patient's variant data file.

3. The MODY type diabetes classification method as described in claim 1, characterized in that, The genetic disease diagnostic database includes the InterVar, ClinVar, HGMD, and OMIM databases; The software for predicting the hazard of mutations includes SIFT, PolyPhen2, and gerp++.

4. The MODY type diabetes classification method as described in claim 1, characterized in that, The scoring data were tested, including the Kolmogorov-Smirnov test and one-sided test. The probability of MODY type diabetes is calculated using Bayes' theorem.

5. A MODY type diabetes classification system, characterized in that, The system includes: The variant data file acquisition subsystem is used to acquire the patient's variant data files; The pathogenic gene analysis and screening subsystem is used to filter the mutation quality, mutation frequency and mutation type of each mutation site in the mutation data file to obtain high-quality mutation sites. The high-quality mutation sites are compared and analyzed using a genetic disease diagnosis database and mutation hazard prediction software. The mutation sites are annotated to obtain mutation gene information. The mutation gene information is classified and scored to obtain the scoring data of the mutation gene information. A diabetes physiological and biochemical index scoring subsystem is used to acquire the patient's clinical physiological and biochemical indexes, score the clinical physiological and biochemical indexes, and obtain the scoring data of the clinical physiological and biochemical indexes. The MODY type diabetes classification result calculation subsystem is used to process the physiological and biochemical indicators in the diabetes epidemiology dataset to obtain the score distribution of the physiological and biochemical indicators in the diabetes epidemiology dataset. It examines the scoring data including clinical physiological and biochemical indicators and the scoring data of variant gene information to determine whether it conforms to the bias of the MODY type diabetes distribution. If it does, it calculates the probability of MODY type diabetes and obtains the MODY type by combining the variant gene information. The patient's clinical physiological and biochemical indicators include BMI, age, family information, age of onset of diabetes, duration of disease, glycated hemoglobin, fasting blood glucose, and C-peptide to urea nitrogen-creatinine ratio. The clinical physiological and biochemical indicators are scored by weighting the scores based on their contribution. The scores of the clinical physiological and biochemical indicators are then combined with the scores of the mutated gene information.

6. The MODY type diabetes classification system as described in claim 5, characterized in that, The specific functions of the subsystem for obtaining variant data files include: extracting MODY-related gene sequences using human reference genome information, constructing a MODY gene reference sequence and its index file from the MODY-related gene sequences, reading the patient's gene sequencing data, performing quality control on the gene sequencing data, comparing it with the MODY gene reference sequence, and obtaining the patient's variant data file.

7. An application of a MODY type diabetes classification system, characterized in that, The MODY type diabetes classification system as described in claim 5 or 6 is used for the development and application of bioinformatics software.

8. A MODY type diabetes classification device, characterized in that, The device includes: a data acquisition device, a memory, and a processor; The data acquisition device is used to acquire the data required for the MODY type diabetes classification; the memory is used to store one or more program instructions for executing the method as described in any one of claims 1-4; the processor is used to execute one or more program instructions for executing the method as described in any one of claims 1-4.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains one or more program instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Diagnosis model for distinguishing MODY from T1D and T2D

    CN111584080A