Application of protein marker in preparation of reagent, kit or device for predicting attack of type 2 diabetes mellitus
By constructing 23 plasma protein scoring models based on machine learning, combined with multiple factors, the problem of failure to cover the risk prediction of type 2 diabetes in the existing technology is solved, and precise prevention and control of type 2 diabetes and early screening is achieved.
Patent Information
- Application Number
- CN202510943014.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-07-09
AI Technical Summary
The prior art has failed to provide a risk prediction of type 2 diabetes covering multiple time points, and lacks a proteomic model for the Chinese population, making it difficult to achieve precise prevention and control of type 2 diabetes.
Based on machine learning algorithms, a protein scoring model of 23 plasma proteins was constructed, and combined with factors such as age, gender, education level, smoking status and waist circumference, a cross-period prediction model was established, and a type 2 diabetes incidence prediction model was verified using Olink proteome PEA technology and discovery-internal verification strategy.
It realizes multi-time risk prediction for type 2 diabetes, improves prediction efficiency, supports early screening and precise prevention and control, and is suitable for the full-cycle management of type 2 diabetes.
Smart Images

Figure CN120446499A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical technology, and in particular to a proteomics-based prediction model for the onset of type 2 diabetes and its application. Background Art
[0002] Type 2 diabetes (T2D) has posed a significant global disease burden. In 2021, the global prevalence of diabetes reached 6.1% (95% confidence interval: 5.8%-6.5%), of which type 2 diabetes accounted for 96.0% (Lancet. 2023;402(10397):203-234). T2D has an insidious onset, a long course, and is difficult to cure. It can cause serious complications across multiple organs, and once diagnosed, lifelong glucose control is required. Therefore, early prediction and assessment of future T2D risk, screening of high-risk individuals, and early intervention are among the most effective and cost-effective strategies for reducing the burden of T2D.
[0003] T2D is a chronic and complex disease that is influenced by both genetic and environmental factors. Proteins, as downstream products of gene expression, can also respond to external environmental factors and are ideal biomarkers for reflecting the body's glucose metabolism function and measuring the risk of T2D. Previous studies have found that plasma protein levels such as IGFBP1, IGFBP2, and GHR are associated with the risk of T2D. Furthermore, using Mendelian randomization methods, potential causal associations between plasma GCKR, RAB1A, SHBG, ATP1B2, and GSTA1 and the risk of diabetes have been identified (Cell Rep Med. 2023;4(9):101174.; Diabetes Care. 2023;46(4):733-741). Some studies have also evaluated the predictive power of T2D-related plasma proteins for the risk of T2D development in the next 10 or 20 years. The results showed that incorporating proteins into the model can significantly improve the predictive power of the basic model (Diabetes Care. 2023;46(4):733-741; Diabetes Care. 2025:dc242478), showing the great potential of applying plasma proteins to the long-term risk prediction of T2D.
[0004] However, current proteomic studies of T2D have primarily focused on European and American populations, while evidence for this in the Chinese population remains lacking. Furthermore, given the progressive and long-term course of T2D, the protein profiles associated with T2D onset in different years of the future may differ. For example, short- to medium-term predictions over the next 3-5 years can capture the metabolic imbalances currently developing, while long-term predictions over 10 years or more can reveal more structural changes related to genetic factors. Protein models constructed in previous studies only assess T2D risk at a single point in the future and fail to provide a comprehensive and integrated assessment across all time periods. Screening for protein combinations that accurately predict across multiple timeframes could comprehensively reflect the core biological pathways underlying the continuous evolution of type 2 diabetes, providing a more cost-effective and efficient solution for its full-cycle prevention and control. Therefore, it is still necessary to map the protein profiles associated with T2D risk in the Chinese population and establish a T2D plasma protein prediction model with robust performance across multiple timeframes to achieve precise T2D prevention and control. Summary of the Invention
[0005] In light of this, the present invention provides a proteomics-based prediction model for type 2 diabetes and its uses. Based on a machine learning algorithm, the present invention proposes a protein score constructed from 23 plasma proteins. Based on this score, a type 2 diabetes prediction model was established and validated, facilitating precise prevention and control of T2D across time periods.
[0006] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions:
[0007] The present invention provides protein markers, which are FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3.
[0008] The present invention provides the application of the above protein markers in predicting the onset of type 2 diabetes.
[0009] The present invention also provides the use of the above protein marker in preparing a reagent, a kit or a device for predicting the onset of type 2 diabetes.
[0010] In some specific embodiments of the present invention, the above-mentioned prediction of the onset of type 2 diabetes includes prediction based on PRS;
[0011] The PRS = NPX1 × 0.355114 + NPX2 × 0.313233 - NPX3 × 0.30225 +NPX4 × 0.284929 + NPX5 × 0.265607 - NPX6 × 0.24851 - NPX7 × 0.24182 -NPX8 × 0.21008 + NPX9 × 0.198402 - NPX10 × 0.14131 + NPX11 × 0.125386 -NPX12 × 0.0958 - NPX13 × 0.08129 + NPX14 × 0.054934 + NPX15 × 0.052256 +NPX16 × 0.042321 - NPX17 × 0.03355 + NPX18 × 0.030498 - NPX19 × 0.02234 +NPX20 × 0.011776 - NPX21 × 0.00255 + NPX22 × 0.000875 - NPX23 × 0.00019;
[0012] Among them, NPX1 to NPX23 are the NPX values of FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3 of the subjects, respectively.
[0013] In some specific embodiments of the present invention, the above-mentioned prediction of the onset of type 2 diabetes includes prediction based on the subject's age, gender, education level, waist circumference and the PRS.
[0014] In some specific embodiments of the present invention, the above-mentioned prediction of the onset of type 2 diabetes is predicted by risk score;
[0015] The risk score = the PRS × 1.730733072 + age factor + gender factor + education level factor + smoking status factor + waist circumference factor - 2.941020761;
[0016] The age factor is the subject's age × 0.033150918;
[0017] The value of the gender factor is as follows: if the subject is male, the value is 0; if the subject is female, the value is 0.499399374;
[0018] The value of the education level factor is as follows: if the subject's education level is primary school or below, the value is 0; if the subject's education level is junior high school or high school, the value is -0.549186868; if the subject's education level is college or above, the value is -0.636837765;
[0019] The value of the smoking status factor is as follows: if the subject is not a current smoker, the value is 0; if the subject is a current smoker, the value is 1.037613865;
[0020] The waist circumference factor is the subject's waist circumference value × 0.022096412, and the unit of the waist circumference value is centimeters.
[0021] In some specific embodiments of the present invention, the UniProt accession numbers of the FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3 used above are as follows: The next ones are Q96MK3, Q6QNK2, Q6P1J6, O15031, Q13478, Q15517, Q13316, Q9NQ79, P05107, P06858, Q8N4F0, Q9NQ30, Q9UI42, Q9Y5Q6, Q9UNE0, P23141, Q9UJA9, Q96C92, Q9P0G3, P01178, P19021, P22079 and Q9UM47.
[0022] The present invention also provides reagents, and detection targets include FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3.
[0023] In some specific embodiments of the invention, the reagents comprise specific antibodies against FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO, and NOTCH3.
[0024] The present invention also provides a kit comprising the above reagents.
[0025] The present invention also provides a device comprising the above reagent.
[0026] In some specific embodiments of the present invention, the above device further comprises:
[0027] An acquisition module is used to obtain the NPX values of FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO, and NOTCH3 of the subject;
[0028] an analysis module, configured to calculate a risk score and predict the risk of a subject developing type 2 diabetes based on the risk score;
[0029] The risk score = PRS × 1.730733072 + age factor + gender factor + education level factor + smoking status factor + waist circumference factor - 2.941020761;
[0030] The age factor is the subject's age × 0.033150918;
[0031] The value of the gender factor is as follows: if the subject is male, the value is 0; if the subject is female, the value is 0.499399374;
[0032] The value of the education level factor is as follows: if the subject's education level is primary school or below, the value is 0; if the subject's education level is junior high school or high school, the value is -0.549186868; if the subject's education level is college or above, the value is -0.636837765;
[0033] The value of the smoking status factor is as follows: if the subject is not a current smoker, the value is 0; if the subject is a current smoker, the value is 1.037613865;
[0034] The waist circumference factor is the subject's waist circumference value × 0.022096412, and the unit of the waist circumference value is centimeters;
[0035] The PRS = NPX1 × 0.355114 + NPX2 × 0.313233 - NPX3 × 0.30225 +NPX4 × 0.284929 + NPX5 × 0.265607 - NPX6 × 0.24851 - NPX7 × 0.24182 -NPX8 × 0.21008 + NPX9 × 0.198402 - NPX10 × 0.14131 + NPX11 × 0.125386 -NPX12 × 0.0958 - NPX13 × 0.08129 + NPX14 × 0.054934 + NPX15 × 0.052256 +NPX16 × 0.042321 - NPX17 × 0.03355 + NPX18 × 0.030498 - NPX19 × 0.02234 +NPX20 × 0.011776 - NPX21 × 0.00255 + NPX22 × 0.000875 - NPX23 × 0.00019;
[0036] Among them, NPX1 to NPX23 are the NPX values of FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3 of the subjects, respectively.
[0037] Based on Olink proteomic PEA technology, machine learning technology and discovery-internal validation strategy, the present invention obtained a prediction model including 23 plasma proteins: FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO, and NOTCH3, which provides new ideas for the full-cycle prediction and prevention and control of type 2 diabetes, and is of great significance for the early screening and precise prevention and control of type 2 diabetes. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for describing the embodiments or the prior art.
[0039] Figure 1Shows the research design flow chart of the present invention;
[0040] Figure 2 The Cox proportional hazards regression model is used to preliminarily screen T2D-related plasma protein association diagrams in the examples;
[0041] Figure 3 Graph showing the relationship between the regularization parameter λ and the partial likelihood estimation deviation of the LASSO-Cox regression model in the embodiment;
[0042] Figure 4 The time-dependent ROC curves of the prediction model in the training set are shown in the embodiment (3 years, 5 years, 10 years, and 15 years);
[0043] Figure 5 The time-dependent ROC curves of the prediction model in the test set (3 years, 5 years, 10 years, and 15 years) are shown in the embodiment.
[0044] Figure 6 The forest plot of the area under the ROC curve of the prediction model in the training and test sets in the embodiment is shown;
[0045] Figure 7 The survival curves of the risk groups are shown in the embodiment using a cutoff value of 10% for the 10-year T2D incidence risk. DETAILED DESCRIPTION
[0046] The present invention discloses a proteomics-based prediction model for the onset of type 2 diabetes and its uses. Those skilled in the art can refer to the contents of this article and appropriately improve the process parameters to achieve the desired results. It should be noted that all similar substitutions and modifications are obvious to those skilled in the art and are considered to be included in the present invention. The methods and applications of the present invention have been described through preferred embodiments, and relevant personnel can obviously modify or appropriately change and combine the methods and applications described herein without departing from the content, spirit and scope of the present invention to implement and apply the technology of the present invention.
[0047] It should be understood that the expression "one or more of" includes individually each of the items recited after the expression and various combinations of two or more of the recited items, unless otherwise apparent from the context and usage. The expression "and / or" in conjunction with three or more recited items should be understood to have the same meaning, unless otherwise apparent from the context.
[0048] The terms "comprising", "having" or "containing", including their grammatical synonyms, should generally be understood as open and non-restrictive, e.g., not excluding other unrecited elements or steps, unless otherwise specifically stated or understood from the context.
[0049] It should be understood that the order of steps or the order in which certain actions are performed are not important as long as the present invention remains operable. Additionally, two or more steps or actions may be performed simultaneously.
[0050] The use of any and all examples or exemplary language such as "for example" or "including" herein is intended only to better illustrate the present invention and does not limit the scope of the present invention. No language in this specification should be construed as indicating any non-claimed element is essential to the practice of the present invention.
[0051] In addition, the numerical ranges and parameters used to define the present invention are approximate values. The relevant numerical values in the specific examples have been presented as accurately as possible. However, any numerical value inherently inevitably contains standard deviations due to individual testing methods. Therefore, unless otherwise expressly stated, it should be understood that all ranges, amounts, values, and percentages used in this disclosure are modified by the word "about." As used herein, "about" generally means that the actual value is within plus or minus 10%, 5%, 1%, or 0.5% of a particular value or range.
[0052] The technical solution of the present invention is obtained according to the following steps:
[0053] Step 1: From the baseline survey population of the China Chronic Diseases Cohort (CKB), a sub-cohort without diabetes was screened and matched with questionnaire survey, physical examination, long-term follow-up and genotyping data of this sub-cohort;
[0054] Step 2: Integrate clinical knowledge, previous literature, etc. to select traditional risk factors and common predictors to be used in subsequent model construction;
[0055] Step 3: Use the Olink Explore 3072 detection platform to test the plasma protein levels of the sub-cohort population to obtain the plasma protein levels of the population;
[0056] Step 4: In the sub-cohort population, the Cox proportional hazards regression model was used to preliminarily screen traditional risk factors and plasma proteins associated with T2D;
[0057] Step 5: The sub-cohort population was randomly divided into a training set and a test set. The plasma proteins were screened again in the training set using the LASSO-Cox regression model to obtain the regression coefficients of the plasma proteins and traditional predictors included in the model;
[0058] Step 6: Based on the LASSO-Cox model established in the training set, the final predicted protein value combination is obtained;
[0059] Step 7: The model prediction performance was evaluated using the area under the time-dependent curve and the net reclassification index in the training set, and verified in the test set, ultimately obtaining the protein score for T2D onset prediction of the present invention.
[0060] Traditional risk factors include sex, age, education, smoking status, alcohol consumption, physical activity level, body mass index (BMI), waist circumference, and family history of diabetes. After screening, the risk factors used to construct the basic T2D model were sex, age, education, smoking status, and waist circumference.
[0061] The common predictors are random glucose (RPG) and genetic risk score (GRS).
[0062] The protein score PRS for T2D onset prediction is as follows: FAM20A * 0.355114 + ADGRD1 * 0.313233 - PLB1 * 0.30225 + PLXNB2 * 0.284929 + IL18R1 * 0.265607 - CDSN * 0.24851 -DMP1 * 0.24182 -CRTAC1 * 0.21008 + ITGB2 * 0.198402 - LPL * 0.14131 + BPIFB2 *0.125386 - ESM1 * 0.0958 - CPA4 * 0.08129 + INSL5 * 0.054934 + EDAR *0.052256 + CES1 * 0.042321 - ENPP5 * 0.03355 + ENTR1 * 0.030498 - KLK14 * 0.02234 + OXT * 0.011776 - PAM * 0.00255 + LPO * 0.000875 - NOTCH3 * 0.00019.
[0063] The risk score of the T2D onset prediction protein model is: risk score = 0.033150918 × age + 0.499399374 (sex == 1) - 0.549186868 (education level == 2) - 0.636837765 (education level == 3) + 1.037613865 (current smoker == 1) + 0.022096412 × waist circumference + 1.730733072 × protein score - 2.941020761.
[0064] The protein is detected in a blood sample, and the blood sample may be whole blood, plasma, or serum.
[0065] The protein can be detected by enzyme-linked immunosorbent assay (ELISA), mass spectrometry, nucleic acid amplification, immunoassay, rapid test kit and the like.
[0066] Unless otherwise specified, the raw materials, reagents, consumables and instruments involved in the present invention are all common commercial products and can be purchased from the market.
[0067] The present invention will be further described below with reference to the embodiments.
[0068] Example: Construction and validation of a multi-time point diagnostic model for type 2 diabetes based on plasma proteome
[0069] This example is divided into seven parts: (1) selection of research subjects and data collection; (2) evaluation of risk factors and assessment of common predictive factors; (3) detection of plasma proteins using the Olink Explore 3072 platform; (4) preliminary screening of traditional predictive factors and plasma proteins; (5) establishment of a LASSO-Cox regression model; (6) evaluation and validation of predictive efficacy; (7) relevant parameters of the T2D protein model; and (8) a device for detecting protein levels of T2D predictive factors. The process of this example is as follows: Figure 1 shown.
[0070] (1) Selection of research subjects and data collection
[0071] Selection of research subjects: The research population was from the China Kadoorie Biobank (CKB). The baseline survey of the CKB project was conducted from June 2004 to July 2008, and 512,714 adults were recruited in 10 project areas across the country for questionnaire surveys, physical examinations, blood sample collection, etc. Among all the research subjects who completed the baseline survey, the CKB project selected approximately 30,000 new cases of cardiovascular disease and chronic obstructive pulmonary disease during the follow-up period and matched controls, as well as approximately 70,000 randomly selected research subjects, using a chip custom-designed for the Han Chinese population (Affymetrix Axiom ® The CKB array performed whole-genome genotyping on blood samples collected at baseline. After completing genotyping determination, quality control, and imputation, available genotyping data were obtained for a total of 100,706 study subjects. Based on this, the CKB project randomly selected 2,026 study subjects to form a sub-cohort and performed plasma protein testing based on the following criteria: ① having genotyping data and no kinship; ② self-reported no cardiovascular disease and no statin use at baseline. Based on this sub-cohort, the present invention further eliminated those with missing associated variables and ultimately identified 1,889 study subjects (see Figure 1 ), and their baseline characteristics are shown in Table 1.
[0072] Table 1: Baseline characteristics of all study subjects
[0073]
[0074] Baseline data collection: The baseline survey of the CKB project was conducted using a predetermined standardized survey protocol. Uniformly trained investigators used electronic questionnaires to collect basic demographic information, socioeconomic status, behavioral lifestyle, medical history, and medication history of the subjects. Physical examination data, such as height and weight, were collected using a standardized tool. Venous blood samples were also collected, of which 10 μL was used for immediate random blood glucose testing, and the remaining blood samples were frozen for subsequent whole-genome and other omics testing.
[0075] Disease outcome follow-up: The CKB project has initiated long-term follow-up of cohort members since the baseline survey, comprehensively collecting information on cohort members' deaths, major chronic disease events (including diabetes), hospitalization events, and migration loss. Among them, the channels for obtaining morbidity and mortality information include the death surveillance system of each project area, the population registration system, the routine disease surveillance system, the national health insurance database, and the active targeted monitoring of project staff. All morbidity and mortality classifications are based on the International Classification of Diseases, 10th edition (ICD10). th The CKB project is coded using the International Clinical Diagnosis and Treatment (ICD)-10 (revised version, ICD-10). As of December 31, 2022, the loss to follow-up rate for the CKB project was less than 0.5%. In this invention, the ICD-10 code corresponding to the predicted disease outcome type 2 diabetes is E11.
[0076] All studies were approved by the Biomedical Ethics Committee of Peking University. Before data collection, each subject signed an informed consent form.
[0077] (2) Evaluation of risk factors and assessment of common predictors
[0078] This study incorporates nine traditional risk factors for T2D into the prediction model, combining currently recognized T2D risk factors, previous literature reports (Pang Yao, Diabetes Care, 2024), and CKB project data. These factors include gender, age, education level, smoking status, alcohol consumption, physical activity level, body mass index (BMI), waist circumference, and family history of diabetes. Furthermore, because random plasma glucose (RPG) has been shown to have strong predictive efficacy for T2D, this study incorporates RPG as an independent predictor, distinct from the above risk factors, in a subsequent separate prediction model for comparison.
[0079] The coding methods of the above-mentioned risk factors and predictors are shown in Table 2.
[0080] Table 2: Categories and codes of risk factors and predictors
[0081]
[0082] Genetic factors are a key factor in the development of T2D. To assess the ability of individual genes to improve T2D prediction through traditional factors and compare this with subsequent protein-based prediction, a weighted genetic risk score (GRS) was constructed using 46 single nucleotide polymorphisms (SNPs) previously validated by Wen Gan et al. in a CKB population. The GRS is shown in Formula I.
[0083] Formula I
[0084] in, is 46, For the The weight corresponding to each SNP, For the Research subjects The number of alleles for each SNP. Specific SNP information is shown in Table 3.
[0085] Table 3: Locus information used in T2D genetic risk score
[0086]
[0087]
[0088] (3) Olink Explore 3072 platform for plasma protein detection
[0089] Plasma levels of 2,944 targeted proteins were measured using the Olink Explore 3072 assay platform. Baseline plasma samples, frozen at -80°C, were thawed and aliquoted into 96-well plates (with 8 wells reserved for quality control). Each plate contained both case and cohort samples, arranged in the order in which they were retrieved from the Wolfson Laboratory in Oxford. The assayed plates were shipped cold chain on dry ice to the Olink laboratory in Uppsala, Sweden, where they were analyzed using the Olink Explore 3072 assay platform, which includes panels with similar numbers of detectable proteins in four categories: cardiovascular metabolic, inflammatory, neurological, and oncological. Plasma protein levels were normalized using intra- and inter-plate controls, and all proteins were converted using predefined correction factors. Protein limits of detection (LODs) were determined using a negative control sample (buffered buffer without antigen). Plasma protein levels after pretreatment were expressed as the logarithm of normalized protein expression (NPX) in arbitrary units.
[0090] (4) Preliminary screening of traditional predictors and plasma proteins
[0091] First, the present invention screened the aforementioned traditional risk factors for T2D and determined the predictive factors for constructing a traditional prediction model. A multivariate Cox proportional hazard regression model was used, with the incidence of T2D as the dependent variable, and gender, age, education level, smoking status, drinking status, physical activity level, BMI, waist circumference, and family history of diabetes were included. In addition to gender and age being retained as variables in the basic model, the remaining variables were screened for T2D-related variables with a significance threshold of P < 0.05 and included in the basic model. The results of the multivariate Cox regression analysis are shown in Table 4. Based on this result, the basic model variables selected included gender, age, education level, smoking status, and waist circumference.
[0092] Table 4: Results of multivariate Cox regression analysis of traditional risk factors and T2D incidence
[0093]
[0094] Subsequently, the types of plasma proteins were initially screened. A Cox proportional hazard regression model was used, with the standardized plasma protein NPX level as the independent variable and the incidence of T2D as the dependent variable. Gender, age, project area, education level, smoking status, drinking status, physical activity level, BMI, fasting time, and protein test batch were adjusted. With a false discovery rate (FDR) of <0.05, 258 T2D-related proteins were screened (see Figure 2 ).
[0095] (5) Establishing LASSO-Cox regression model
[0096] In the entire construction population, 1889 subjects were randomly divided into training and test sets (see Figure 1 A LASSO-Cox regression model was constructed for the 944 subjects in the training set. This model introduces a penalty term based on the Cox proportional hazards regression model. Its advantages include variable screening to obtain better performance parameters and complexity adjustment to avoid model overfitting.
[0097] The general Cox proportional hazards regression model formula is shown in Formula II.
[0098] Formula II
[0099] in, For the Samples at time risk, is the baseline hazard function, is the regression coefficient vector, For the The characteristic vector of a sample can be estimated using the partial likelihood function, as shown in Formula III.
[0100] Formula III
[0101] in, is the sample size, For time The set of samples that are still at risk when For the Whether the sample has T2D. Take the logarithm of the above partial likelihood function and record it as Subsequently, the L1 regularization of LASSO is introduced as shown in Equation IV.
[0102] Formula IV
[0103] in, is the total number of proteins to be screened, For the The coefficient of the protein to be screened, λ is the LASSO regularization parameter. This formula is equivalent to formula V.
[0104] Formula V
[0105] Using the above method, the 258 proteins and 5 traditional predictors that were initially screened were included in the example. The model was set to penalize only proteins through parameter setting, and the optimal parameter λ was selected using 5-fold cross validation. The mean square error (MSE) changes corresponding to different λ values are shown in Figure 3 The two vertical lines represent λ.min and λ.1se, respectively. The former is the λ value that minimizes the MSE, while the latter is the λ value that keeps the MSE within 1 standard error of the minimum MSE while also reducing model complexity. In this example, λ.min = 0.02073961 was selected. Based on this value, 23 proteins were screened for predictive model construction, and their corresponding coefficients were obtained (see Table 5).
[0106] Table 5: Proteins and their coefficients selected by LASSO-Cox model
[0107]
[0108] Note: Cardiometabolic, cardiovascular metabolism; Inflammation, inflammation; Oncology, tumor; Neurology, nerve
[0109] (6) Prediction performance evaluation and verification
[0110] Based on the above traditional risk factors, RPG, GRS and proteins, four T2D prediction models were constructed using the Cox proportional hazards regression model in the training set population and predicted in the test set:
[0111] ① Base model, which includes five traditional risk factors;
[0112] ② Blood glucose model (Base+RPG), which incorporates RPG on top of the basic model;
[0113] ③Genetic model (Base+GRS), which incorporates the GRS of T2D on top of the basic model;
[0114] ④ Protein model (Base+PRS): A protein risk score (PRS) consisting of 23 additional proteins is incorporated into the basic model. The score weight uses the regression coefficient obtained from the LASSO-Cox model in Table 5, and the calculation formula is shown in Formula VI.
[0115] Formula VI
[0116] in, For the The weight corresponding to each protein, For the Research subjects NPX value of each protein.
[0117] In the training set and test set, the time-dependent area under the curve (AUC) and net reclassification index (NRI) of the four prediction models were calculated to evaluate the actual prediction effect of the prediction models. The time-dependent AUC results showed that the blood glucose model and the protein model could significantly improve the prediction efficiency of the basic model for the onset of T2D in the next 3, 5, 10 and 15 years, and the improvement of the protein model was higher than that of the blood glucose model ( Figure 4-Figure 6 The genetic model did not significantly improve the predictive performance of the basic model. Regarding the NRI, the protein model significantly improved the basic model in predicting T2D onset 3, 5, 10, and 15 years in the future, while the blood glucose model and the genetic model showed no significant improvement in the training or test sets for some years (Table 6). These results consistently demonstrate that the method and constructed model of the present invention can accurately predict the risk of T2D onset at multiple time points.
[0118] Table 6: NRI results of T2D prediction models in training and test sets
[0119]
[0120] (7) Related parameters of T2D prediction model
[0121] The specific calculation formula for the risk score obtained by the aforementioned protein model is as follows:
[0122] Protein model risk score = 0.033150918 × age + 0.499399374 (sex == 1) - 0.549186868 (education level == 2) - 0.636837765 (education level == 3) + 1.037613865 (current smoker == 1) + 0.022096412 × waist circumference + 1.730733072 × PRS - 2.941020761.
[0123] To facilitate model application, taking the risk of 10% in the next 3, 5, 10, and 15 years as an example, the cutoff values of the protein model risk score were 3.6139, 3.0976, 2.1760, and 1.3391. Based on this series of cutoff values, the population was divided into high and low risk groups. There were significant differences in the risk of T2D between the populations ( Figure 7 The risk score calculation formulas and corresponding cutoff values of the other models are shown in Table 7.
[0124] Table 7: Prediction model formulas and cutoff values
[0125]
[0126] (8) Protein level detection device for T2D predictors
[0127] This embodiment provides a protein level detection device for T2D prediction. Depending on actual needs, the protein detection device can be prepared into a variety of optional detection kits, with no limitation on specific formats and detection methods, such as ELISA, immunofluorescence kits, or detection using mass spectrometry, nucleic acid amplification, and other methods.
[0128] In the aforementioned kits, taking an ELISA kit as an example, the kit should include a solid phase carrier, preferably an enzyme-labeled plate (such as a 96-well or other polystyrene microplate), a membrane carrier (such as a nitrocellulose membrane, glass cellulose membrane, or nylon membrane), or microspheres, pre-coated with an antibody specific for the protein to be tested. In some embodiments, the membrane carrier may also include a positive control to facilitate detection and result interpretation. The kit should further include a primary antibody that binds to the target and an enzyme-labeled secondary antibody, the latter of which may be labeled with horseradish peroxidase (HRP) or alkaline phosphatase.
[0129] Kits of all formats are capable of quantitative protein detection and are particularly suitable for the quantitative analysis of multiple protein markers in blood specimens. Taking an ELISA kit as an example, when detecting a protein marker such as LPL in serum, it first binds to an LPL-specific antibody pre-coated on a solid support. Subsequently, an enzyme-labeled secondary antibody is added, which also binds to the complex and becomes immobilized on the solid surface. The amount of enzyme bound to the solid surface is positively correlated with the concentration of LPL protein in the blood specimen. Upon addition of a specific enzyme substrate, the enzyme catalyzes the formation of a colored product, the color of which reflects the content of the protein being tested. Due to the extremely high catalytic efficiency of the enzyme reaction, the signal is significantly amplified, improving the sensitivity and accuracy of the test and enabling qualitative and quantitative analysis of the target protein.
[0130] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. Use of a protein marker in the preparation of a reagent, kit or device for predicting the onset of type 2 diabetes, characterized in that: The protein markers are FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3.
2. The use according to claim 1, characterized in that The prediction of the onset of type 2 diabetes includes prediction based on PRS; The PRS = NPX1 × 0.355114 + NPX2 × 0.313233 - NPX3 × 0.30225 + NPX4 × 0.284929 + NPX5 × 0.265607 - NPX6 × 0.24851 - NPX7 × 0.24182 - NPX8 × 0.21008 + NPX9 × 0.198402 - NPX10 × 0.14131 + NPX11 × 0.125386 - NPX12 × 0.0958 - NPX13 × 0.08129 + NPX14 × 0.054934 + NPX15 × 0.052256 + NPX16 × 0.042321 - NPX17 × 0.03355 + NPX18 × 0.030498 - NPX19 × 0.02234 + NPX20 × 0.011776 - NPX21 × 0.00255 + NPX22 × 0.000875 - NPX23 × 0.00019; Among them, NPX1 to NPX23 are the NPX values of FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3 of the subjects, respectively.
3. The use according to claim 2, characterized in that The prediction of the onset of type 2 diabetes includes prediction based on the subject's age, gender, education level, waist circumference and the PRS.
4. The use according to claim 3, characterized in that The prediction of the onset of type 2 diabetes is predicted by a risk score; The risk score = the PRS × 1.730733072 + age factor + gender factor + education level factor + smoking status factor + waist circumference factor - 2.941020761; The age factor is the subject's age × 0.033150918; The value of the gender factor is as follows: if the subject is male, the value is 0; if the subject is female, the value is 0.499399374; The value of the education level factor is as follows: if the subject's education level is primary school or below, the value is 0; if the subject's education level is junior high school or high school, the value is -0.549186868; if the subject's education level is college or above, the value is -0.636837765; The value of the smoking status factor is as follows: if the subject is not a current smoker, the value is 0; if the subject is a current smoker, the value is 1.037613865; The waist circumference factor is the subject's waist circumference value × 0.022096412, and the unit of the waist circumference value is centimeters.
5. The use according to claim 1, characterized in that The UniProt accession numbers of FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3 are Q96MK3, Q 6QNK2, Q6P1J6, O15031, Q13478, Q15517, Q13316, Q9NQ79, P05107, P06858, Q8N4F0, Q9NQ30, Q9UI42, Q9Y5Q6, Q9UNE0, P23141, Q9UJA9, Q96C92, Q9P0G3, P01178, P19021, P22079, and Q9UM47.
6. Reagent, characterized in that Detection targets include FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO, and NOTCH3.
7. The reagent according to claim 6, wherein Includes specific antibodies for FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO, and NOTCH3.
8. A kit, characterized in that Contains the reagent according to claim 6 or 7.
9. The device, characterized in that Contains the reagent according to claim 6 or 7.
10. The device according to claim 9, wherein Also includes: An acquisition module is used to obtain the NPX values of FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO, and NOTCH3 of the subject; an analysis module, configured to calculate a risk score and predict the risk of a subject developing type 2 diabetes based on the risk score; The risk score = PRS × 1.730733072 + age factor + gender factor + education level factor + smoking status factor + waist circumference factor - 2.941020761; The age factor is the subject's age × 0.033150918; The value of the gender factor is as follows: if the subject is male, the value is 0; if the subject is female, the value is 0.499399374; The value of the education level factor is as follows: if the subject's education level is primary school or below, the value is 0; if the subject's education level is junior high school or high school, the value is -0.549186868; if the subject's education level is college or above, the value is -0.636837765; The value of the smoking status factor is as follows: if the subject is not a current smoker, the value is 0; if the subject is a current smoker, the value is 1.037613865; The waist circumference factor is the subject's waist circumference value × 0.022096412, and the unit of the waist circumference value is centimeters; The PRS = NPX1 × 0.355114 + NPX2 × 0.313233 - NPX3 × 0.30225 + NPX4 × 0.284929 + NPX5 × 0.265607 - NPX6 × 0.24851 - NPX7 × 0.24182 - NPX8 × 0.21008 + NPX9 × 0.198402 - NPX10 × 0.14131 + NPX11 × 0.125386 - NPX12 × 0.0958 - NPX13 × 0.08129 + NPX14 × 0.054934 + NPX15 × 0.052256 + NPX16 × 0.042321 - NPX17 × 0.03355 + NPX18 × 0.030498 - NPX19 × 0.02234 + NPX20 × 0.011776 - NPX21 × 0.00255 + NPX22 × 0.000875 - NPX23 × 0.00019; Among them, NPX1 to NPX23 are the NPX values of FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3 of the subjects, respectively.
Citation Information
Patent Citations
Biomarker group, kit and system for predicting poor prognosis of cerebral arterial thrombosis patient
CN115011687A
Application of protein marker in preparation of product for predicting future coronary heart disease onset risk of subject
CN119959557A
Biomarkers for predicting type 2 diabetes status
US20250208147A1