Use of protein markers for the preparation of a reagent, kit or device for the prediction of the onset of type 2 diabetes
By constructing 23 plasma protein scoring models based on machine learning and combining them with traditional factors, the problem of the inability of existing technologies to comprehensively assess the risk of type 2 diabetes at multiple time points has been solved, thus achieving precise prevention and early screening of type 2 diabetes.
Patent Information
- Application Number
- CN202510943014.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-07-09
AI Technical Summary
Existing technologies fail to provide a comprehensive assessment of the risk of developing type 2 diabetes across multiple time points and lack proteomics models specific to the Chinese population, making it difficult to achieve precise prevention and control of type 2 diabetes.
Based on machine learning algorithms, a protein scoring model containing 23 plasma proteins was constructed. Combined with factors such as age, gender, education level, smoking status, and waist circumference, a predictive model for the incidence of type 2 diabetes was established. Plasma proteins were screened using Olink proteomics PEA technology and LASSO-Cox regression model to achieve multi-time-point prediction.
It improves the predictive efficacy of type 2 diabetes at various future time periods, provides a full-cycle risk assessment of type 2 diabetes, supports early screening and precise prevention and control, and is applicable to the Chinese population.
Smart Images

Figure CN120446499B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of biomedical technology, in particular to a type 2 diabetes onset prediction model based on proteomics and use thereof. BACKGROUND
[0002] Type 2 diabetes (T2D) has caused a serious disease burden worldwide. The global prevalence of diabetes in 2021 was 6.1% (95% confidence interval: 5.8%-6.5%), of which type 2 diabetes accounted for 96.0% (Lancet. 2023;402(10397):203-234). T2D is insidious in onset, long in course, and difficult to cure, which can cause serious complications in multiple organs of the body. Once diagnosed, it needs to be controlled for a lifetime. Therefore, early prediction, assessment of future T2D risk, screening of high-risk groups and early intervention are one of the most effective and cost-effective strategies to reduce the disease burden of T2D.
[0003] T2D is a chronic and complex disease, which is affected by both genetic and environmental factors. Proteins, as downstream products of gene expression, can also respond to external environmental factors, and are ideal biomarkers for reflecting the function of glucose metabolism and measuring the risk of T2D. Previous studies have found that plasma protein levels such as IGFBP1, IGFBP2, GHR are associated with the risk of T2D, and further through Mendelian randomization method, determined the potential causal relationship between plasma GCKR, RAB1A, SHBG, ATP1B2 and GSTA1 and the risk of diabetes (Cell Rep Med. 2023;4(9):101174.; Diabetes Care. 2023;46(4):733-741). Some studies have also evaluated the predictive performance of T2D-related plasma proteins on the risk of T2D in the next 10 or 20 years, and the results showed that the inclusion of proteins in the model can significantly improve the predictive performance of the basic model (Diabetes Care. 2023;46(4):733-741; Diabetes Care. 2025:dc242478), showing the great potential of plasma proteins in long-term risk prediction of T2D.
[0004] However, current T2D-related proteomic studies are mostly based on European and American populations, and there is still a lack of evidence in this regard for Chinese populations. In addition, due to the progressive and long-term development of T2D, the protein spectrum of T2D onset at different time points in the future may have certain differences. For example, short-term and medium-term prediction in the next 3-5 years can capture the current gradually formed metabolic imbalance, while long-term prediction of more than 10 years can reveal more structural changes related to genetic factors. The protein model constructed in previous studies only assesses the risk of T2D occurrence at a single time point in the future, and cannot provide a complete and comprehensive assessment of each time period in the future; while screening a combination of proteins that can accurately predict multiple time spans can comprehensively reflect the core biological pathways of the continuous evolution of type 2 diabetes mellitus, and provide a more economical and efficient solution for the prevention and control of the whole cycle. Therefore, it is still necessary to draw the protein map of T2D onset risk in Chinese population, and to establish a T2D plasma protein prediction model covering multiple time points with good prediction performance, in order to realize the precise prevention and control of T2D. SUMMARY
[0005] Therefore, the present application provides a type 2 diabetes onset prediction model based on proteomics and use. Based on machine learning algorithm, the present application proposes a protein score constructed by 23 plasma proteins, and establishes and verifies a type 2 diabetes onset prediction model according to the score, which is convenient for the precise prevention and control of T2D across time periods.
[0006] In order to achieve the above-mentioned application purposes, the present application provides the following technical solutions:
[0007] The present application provides protein markers, which are FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDAR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3.
[0008] The present application provides the application of the above-mentioned protein markers in type 2 diabetes onset prediction.
[0009] The present application also provides the application of the above-mentioned protein markers in the preparation of reagents, kits or devices for type 2 diabetes onset prediction.
[0010] In some specific embodiments of the present application, the type 2 diabetes onset prediction of the above-mentioned application comprises prediction based on PRS;
[0011] the PRS = NPX1 x 0.355114 + NPX2 x 0.313233 - NPX3 x 0.30225 + NPX4 x 0.284929 + NPX5 x 0.265607 - NPX6 x 0.24851 - NPX7 x 0.24182 - NPX8 x 0.21008 + NPX9 x 0.198402 - NPX10 x 0.14131 + NPX11 x 0.125386 - NPX12 x 0.0958 - NPX13 x 0.08129 + NPX14 x 0.054934 + NPX15 x 0.052256 + NPX16 x 0.042321 - NPX17 x 0.03355 + NPX18 x 0.030498 - NPX19 x 0.02234 + NPX20 x 0.011776 - NPX21 x 0.00255 + NPX22 x 0.000875 - NPX23 x 0.00019;
[0012] wherein NPX1 to NPX23 are NPX values of FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDAR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3 of the subject, respectively.
[0013] In some embodiments of the present application, the type 2 diabetes incidence prediction of the above application comprises prediction based on the age, gender, education level, waist circumference and the PRS of the subject.
[0014] In some embodiments of the present application, the type 2 diabetes incidence prediction of the above application is predicted by a risk score;
[0015] the risk score = the PRS x 1.730733072 + age factor + gender factor + education level factor + smoking status factor + waist circumference factor - 2.941020761;
[0016] the age factor takes a value of the age value of the subject x 0.033150918;
[0017] the value of the gender factor is 0 if the gender of the subject is male, and is 0.499399374 if the gender of the subject is female.
[0018] The value rule of the education level factor is: if the education level of the subject is primary school or below, the value is 0; if the education level of the subject is junior high school or high school, the value is -0.549186868; if the education level of the subject is university or above, the value is -0.636837765.
[0019] The value rule of the smoking status factor is: if the subject is a non-current smoker, the value is 0; if the subject is a current smoker, the value is 1.037613865.
[0020] The waist circumference factor is the waist circumference value of the subject multiplied by 0.022096412, and the unit of the waist circumference value is centimeter.
[0021] In some embodiments of the present application, the UniProt accession numbers of the FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDAR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3 in the above application are Q96MK3, Q6QNK2, Q6P1J6, O15031, Q13478, Q15517, Q13316, Q9NQ79, P05107, P06858, Q8N4F0, Q9NQ30, Q9UI42, Q9Y5Q6, Q9UNE0, P23141, Q9UJA9, Q96C92, Q9P0G3, P01178, P19021, P22079 and Q9UM47, respectively.
[0022] The present application also provides a reagent for detecting the target including FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDAR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3.
[0023] In some embodiments of the present application, the reagent comprises a specific antibody of FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDAR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3.
[0024] The present application also provides a kit comprising the above-mentioned reagent.
[0025] The present application also provides a device comprising the above-mentioned reagent.
[0026] In some embodiments of the present application, the above-mentioned device further comprises:
[0027] an acquisition module configured to acquire the NPX values of FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDAR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3 of the subject;
[0028] an analysis module configured to calculate a risk score and predict the risk of the subject of developing type 2 diabetes based on the risk score;
[0029] the risk score = PRS x 1.730733072 + age factor + gender factor + education level factor + smoking status factor + waist circumference factor - 2.941020761;
[0030] the age factor is equal to the age value of the subject x 0.033150918;
[0031] the value of the gender factor is 0 if the subject is male, and is 0.499399374 if the subject is female;
[0032] the value of the education level factor is 0 if the subject has a primary school education or below, is -0.549186868 if the subject has a junior high school or high school education, and is -0.636837765 if the subject has a university education or above;
[0033] the value of the smoking status factor is 0 if the subject is a non-current smoker, and is 1.037613865 if the subject is a current smoker;
[0034] the waist circumference factor is equal to the waist circumference value of the subject x 0.022096412, and the unit of the waist circumference value is centimeter;
[0035] the PRS = NPX1 x 0.355114 + NPX2 x 0.313233 - NPX3 x 0.30225 + NPX4 x 0.284929 + NPX5 x 0.265607 - NPX6 x 0.24851 - NPX7 x 0.24182 - NPX8 x 0.21008 + NPX9 x 0.198402 - NPX10 x 0.14131 + NPX11 x 0.125386 - NPX12 x 0.0958 - NPX13 x 0.08129 + NPX14 x 0.054934 + NPX15 x 0.052256 + NPX16 x 0.042321 - NPX17 x 0.03355 + NPX18 x 0.030498 - NPX19 x 0.02234 + NPX20 x 0.011776 - NPX21 x 0.00255 + NPX22 x 0.000875 - NPX23 x 0.00019;
[0036] wherein NPX1 to NPX23 are NPX values of FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDAR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO, and NOTCH3 of the subject, respectively.
[0037] The present application is based on Olink proteome PEA technology, machine learning technology and discovery-internal validation strategy, and obtains a prediction model including 23 plasma proteins: FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDAR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO, and NOTCH3, which provides a new idea for whole cycle prediction and prevention and control of type 2 diabetes, and has important significance for early screening and precise prevention and control of type 2 diabetes. BRIEF DESCRIPTION OF DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, brief introductions will be given below to the drawings needed to be used in the embodiments or prior art descriptions.
[0039] Figure 1A research design flowchart of the present application is shown in Figure 1.
[0040] Figure 2 A Cox proportional hazards regression model in an embodiment is used to preliminarily screen T2D-related plasma protein association graphs;
[0041] Figure 3 A LASSO-Cox regression model in an embodiment is used to regularize parameter λ and partial likelihood estimation bias relationship graphs;
[0042] Figure 4 A time-dependent ROC curve graph (3 years, 5 years, 10 years, 15 years) of a prediction model in an embodiment in the training set is shown in Figure 4;
[0043] Figure 5 A time-dependent ROC curve graph (3 years, 5 years, 10 years, 15 years) of a prediction model in an embodiment in the test set is shown in Figure 5;
[0044] Figure 6 A forest plot of the area under the ROC curve of a prediction model in an embodiment in the training and test sets is shown in Figure 6;
[0045] Figure 7 A survival curve graph of a risk population divided by a 10% cutoff value of 10-year T2D incidence risk in an embodiment is shown in Figure 7. DETAILED DESCRIPTION
[0046] The present application discloses a type 2 diabetes mellitus onset prediction model based on proteomics and use, and those skilled in the art can refer to the content herein to appropriately improve process parameters for implementation. It is particularly pointed out that all similar substitutions and changes are obvious to those skilled in the art, and they are all considered to be included in the present application. The methods and applications of the present application have been described by preferred embodiments, and relevant personnel can obviously modify or appropriately change and combine the methods and applications described herein without departing from the content, spirit and scope of the present application, to realize and apply the present application technology.
[0047] It should be understood that the expression "one or more of" includes each object recited after the expression and various different combinations of two or more of the recited objects, unless otherwise understood from the context and usage. The expression "and / or" in combination with three or more recited objects should be understood to have the same meaning, unless otherwise understood from the context.
[0048] The terms "comprising", "having", or "including", including the use of their grammatical synonyms, should generally be understood to be open-ended and non-limiting, for example, not excluding other unrecited elements or steps, unless otherwise specifically stated or understood from the context.
[0049] It should be understood that the order of steps or order for performing certain actions is immaterial so long as the application remains operable. Moreover, two or more steps or actions can be conducted simultaneously.
[0050] The use of any and all examples, or exemplary language herein, e.g., "such as" or "including", is intended merely to better illustrate the application and does not indicate that any non-claimed element is essential to the practice of the application. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the application.
[0051] Further, the numerical ranges and parameters setting forth the broadest scope of the application are approximations, and are only used to convey generally understood precision. Numerical parameters are only approximations of numerical values in certain embodiments. Therefore, every numerical value is approximated for the purpose of conveying general understanding of the application. Generally, a numerical value can be indicated within a range by using a designation that such a range will be understood to include every number between the upper and lower limits of that range, as interpreted according to the doctrine of equivalents. In this context, a statement that a parameter is "at least" a certain value is understood as a statement that the parameter is greater than or equal to that value. Similarly, a statement that a parameter is "at most" a certain value is understood as a statement that the parameter is less than or equal to that value. Thus, every numerical value is approximated for the purpose of conveying general understanding of the application. Generally, a numerical value can be indicated within a range by using a designation that such a range will be understood to include every number between the upper and lower limits of that range, as interpreted according to the doctrine of equivalents. In this context, a statement that a parameter is "at least" a certain value is understood as a statement that the parameter is greater than or equal to that value. Similarly, a statement that a parameter is "at most" a certain value is understood as a statement that the parameter is less than or equal to that value. Thus, every numerical value is approximated for the purpose of conveying general understanding of the application.
[0052] The technical solutions of the application are obtained according to the following steps:
[0053] Step one: In the baseline survey population of the Chinese Chronic Disease Prospective Study (CKB) cohort, a sub-cohort of people without diabetes at baseline was screened out, and the questionnaire survey, physical examination, long-term follow-up and genotyping data of this part of the population were matched;
[0054] Step two: Based on clinical knowledge, previous literature, etc., traditional risk factors and common predictors to be used for subsequent model construction were selected;
[0055] Step three: The plasma protein levels of the sub-cohort population were detected by using the Olink Explore 3072 detection platform, and the plasma protein levels of the population were obtained;
[0056] Step four: In the sub-cohort population, Cox proportional hazards regression model was used to preliminarily screen traditional risk factors and plasma proteins related to T2D;
[0057] Step five: The sub-cohort population was randomly divided into training set and test set, and LASSO-Cox regression model was used in the training set to screen plasma proteins again, and the regression coefficients of the plasma proteins and traditional predictors included in the model were obtained;
[0058] Step six: Based on the LASSO-Cox model established in the training set, the final prediction protein value combination was obtained;
[0059] Step seven: the area under the time-dependent curve, net reclassification index are used to evaluate the prediction performance of the model in the training set, and are verified in the test set. Finally, the protein score for predicting the onset of T2D of the application is obtained.
[0060] The traditional risk factors include gender, age, education level, smoking status, drinking status, physical activity level, body mass index (BMI), waist circumference, and family history of diabetes. After screening, the risk factors used to construct the T2D basic model are gender, age, education level, smoking status, and waist circumference.
[0061] The common predictors are random plasma glucose (RPG) and genetic risk score (GRS).
[0062] The protein score for predicting the onset of T2D PRS = FAM20A * 0.355114 + ADGRD1 * 0.313233- PLB1 * 0.30225 + PLXNB2 * 0.284929 + IL18R1 * 0.265607 - CDSN * 0.24851 - DMP1 * 0.24182 - CRTAC1 * 0.21008 + ITGB2 * 0.198402 - LPL * 0.14131 + BPIFB2 * 0.125386 - ESM1 * 0.0958 - CPA4 * 0.08129 + INSL5 * 0.054934 + EDAR * 0.052256 + CES1 * 0.042321 - ENPP5 * 0.03355 + ENTR1 * 0.030498 - KLK14 * 0.02234 + OXT * 0.011776 - PAM * 0.00255 + LPO * 0.000875 - NOTCH3 * 0.00019.
[0063] The risk score of the protein model for predicting the onset of T2D is: risk score = 0.033150918 x age + 0.499399374 (gender == 1) - 0.549186868 (education level == 2) - 0.636837765 (education level == 3) + 1.037613865 (current smoker == 1) + 0.022096412 x waist circumference + 1.730733072 x protein score - 2.941020761.
[0064] The protein is detected in a blood sample, and the blood sample can be whole blood, plasma, or serum.
[0065] The protein can be detected by methods such as enzyme-linked immunosorbent assay (ELISA), mass spectrometry, nucleic acid amplification, immunoassay, and rapid test kits.
[0066] Unless otherwise specified, the raw materials, reagents, consumables and instruments involved in this invention are all commercially available products and can be purchased from the market.
[0067] The present invention will be further illustrated below with reference to the embodiments.
[0068] Example: Construction and validation of a multi-timepoint diagnostic model for type 2 diabetes based on plasma proteomics
[0069] This embodiment consists of seven parts: (1) selection of research subjects and data collection; (2) risk factor assessment and evaluation of common predictive factors; (3) detection of plasma proteins using the Olink Explore 3072 platform; (4) preliminary screening of traditional predictive factors and plasma proteins; (5) establishment of a LASSO-Cox regression model; (6) predictive efficacy evaluation and validation; (7) relevant parameters of the T2D protein model; and (8) protein level detection device for T2D predictive factors. The procedure of this embodiment is as follows: Figure 1 As shown.
[0070] (1) Selection of research subjects and data collection
[0071] Study Participant Selection: The study population was drawn from the China Kadoorie Biobank (CBB). The CKB project's baseline survey was conducted from June 2004 to July 2008, recruiting 512,714 adults across 10 project areas nationwide for questionnaires, physical examinations, and blood sample collection. Of all participants who completed the baseline survey, the CKB project extracted approximately 30,000 newly diagnosed cardiovascular disease and chronic obstructive pulmonary disease cases and matched controls during the follow-up period, as well as approximately 70,000 randomly selected participants, using an Affymetrix Axiomtek chip specifically designed for the Han Chinese population. ® The CKB array was used to perform whole-genome genotyping on baseline blood samples. After genotyping, quality control, and imputation, usable genotyping data were obtained for 100,706 participants. Based on this, the CKB project randomly selected 2,026 participants to form a subcohort and performed plasma protein testing according to the following criteria: ① having genotyping data and no kinship among participants; ② reporting no cardiovascular disease and not taking statins at baseline. Based on this subcohort, after further excluding participants with missing correlation variables, 1,889 participants were finally identified (see [link to CKB array]). Figure 1 The baseline characteristics are shown in Table 1.
[0072] Table 1: Baseline characteristics of all study subjects
[0073]
[0074] Baseline data collection: The baseline survey of the CKB project was conducted using a pre-determined standardized survey protocol, and basic demographic information, socioeconomic status, behavioral lifestyle, disease history and medication history of the study subjects were collected by trained investigators using an electronic questionnaire. Physical examination data such as height and weight of the study subjects were collected using a uniform tool, and venous blood samples were collected, of which 10 μL was used for immediate random blood glucose detection, and the remaining blood samples were frozen for subsequent whole genome and other omics detection.
[0075] Disease outcome follow-up: The CKB project started long-term follow-up of cohort members from the beginning of the baseline survey, and collected information on death events, major chronic disease incidence events (including diabetes), hospitalization events and migration loss of cohort members. Among them, the ways to obtain the incidence and death information include the death monitoring system, population registration system, regular disease monitoring system, national medical insurance database and project staff active targeted monitoring in the project area. All incidence or cause of death classification is coded using the International Classification of Diseases, 10th revision (ICD-10). As of December 31, 2022, the loss to follow-up rate of the CKB project was less than 0.5%. In the present invention, the ICD-10 code corresponding to the predicted disease outcome type diabetes is E11. th revision, ICD-10). As of December 31, 2022, the loss to follow-up rate of the CKB project was less than 0.5%. In the present invention, the ICD-10 code corresponding to the predicted disease outcome type diabetes is E11.
[0076] All studies were approved by the Biomedical Ethics Committee of Peking University. Before data collection, each study subject signed an informed consent form.
[0077] (2) Risk factor evaluation and common predictor assessment
[0078] In combination with the currently recognized T2D risk factors, previous literature reports (Pang Yao, Diabetes Care, 2024) and CKB project collection, the present invention comprehensively considers the inclusion of 9 traditional risk factors for constructing a prediction model, including gender, age, education level, smoking status, drinking status, physical activity level, body mass index (BMI), waist circumference and family history of diabetes. In addition, since random blood glucose (RPG) has been proven to have good prediction performance for T2D, the present invention additionally includes RPG as an independent predictor different from the above risk factors, and constructs a separate prediction model for comparison in the subsequent.
[0079] The coding methods of the above-mentioned risk factors and predictors are shown in Table 2.
[0080] Table 2: Categories and codes of risk factors and predictors
[0081]
[0082] Genetic factors are one of the important factors in the occurrence of T2D. In order to evaluate the ability of individual genes to improve T2D prediction by traditional factors and compare it with the subsequent protein prediction efficacy, this invention uses 46 single nucleotide polymorphism (SNP) sites that have been validated in the CKB population by Wen Gan et al. to construct a weighted gene risk score (GRS), which is shown in Formula I.
[0083] Formula I
[0084] in, It is 46. For the first The weights corresponding to each SNP For the first No. 1 research subject The number of alleles for each SNP. See Table 3 for detailed SNP information.
[0085] Table 3: Locus Information Used in T2D Gene Risk Score
[0086]
[0087]
[0088] (3) Detection of plasma proteins using the Olink Explore 3072 platform
[0089] The 2944 plasma protein levels of the study subjects were targetedly detected by using Olink Explore 3072 detection platform. The baseline plasma samples stored at -80°C were thawed and aliquoted into 96-well plates (8 wells were reserved for quality control in each plate), ensuring that each plate contained case and sub-cohort samples and arranged in the order of retrieval from the Oxford Wolfson Laboratory. The detection plate to be detected was transported to the Olink laboratory in Uppsala, Sweden, using dry ice, and analyzed using the Olink Explore 3072 detection platform, which includes four types of panels that can detect phase similar proteins, namely cardiovascular metabolism, inflammation, nerve and tumor. The measured plasma protein levels were normalized using intra- and inter-plate controls, and all proteins were converted using a pre-determined correction factor. The protein detection limit (LOD) was determined using negative control samples, i.e. buffer without antigen. The pre-processed plasma protein levels were represented by the logarithmic value of the normalized protein expression (NPX) in arbitrary units.
[0090] (4) Preliminary screening of traditional predictors and plasma proteins
[0091] Firstly, the present application screens the traditional risk factors of the aforementioned T2D to determine the predictors for constructing the traditional prediction model. A multivariate Cox proportional hazards regression model is used, with T2D onset as the dependent variable, and gender, age, education level, smoking status, drinking status, physical activity level, BMI, waist circumference and family history of diabetes are included. In addition to gender and age as reserved variables for the base model, the remaining variables are screened for T2D-related variables with a significance threshold of P<0.05. The results of the multivariate Cox regression analysis are shown in Table 4. According to the results, the base model variables include gender, age, education level, smoking status, and waist circumference.
[0092] Table 4: Multivariate Cox regression analysis results of traditional risk factors and T2D onset
[0093]
[0094] Subsequently, the plasma protein species were screened. A Cox proportional hazards regression model was used, with the normalized plasma protein NPX level as the independent variable and T2D onset as the dependent variable, adjusting for gender, age, project area, education level, smoking status, drinking status, physical activity level, BMI, fasting time and protein test batch, and selecting 258 T2D-related proteins with a false discovery rate (FDR) <0.05 (see Table 5). Figure 2
[0095] (5) Establish the LASSO-Cox regression model
[0096] In the entire population, 1889 participants were randomly divided into training and test sets (see [link to training set]). Figure 1 A LASSO-Cox regression model was constructed on 944 subjects in the training set. This model introduces a penalty term into the Cox proportional hazards regression model, which has the advantages of allowing variable selection to obtain better performance parameters and complexity adjustment to avoid model overfitting.
[0097] The general formula for the Cox proportional hazards regression model is shown in Equation II.
[0098] Formula II
[0099] in, For the first Each sample in time The risk, For the baseline risk function, For the regression coefficient vector, For the first The feature vector of a sample can be estimated using a partial likelihood function, as shown in Equation III.
[0100] Formula III
[0101] in, For sample size, For time The sample set is still at risk. For the first Does each sample experience T2D? Take the logarithm of the partial likelihood function above, denoted as . Subsequently, L1 regularization of LASSO is introduced, as shown in Equation IV.
[0102] Formula IV
[0103] in, This represents the total number of proteins to be screened. For the first The coefficients of the proteins to be screened, where λ is the LASSO regularization parameter. This equation is equivalent to equation V.
[0104] Formula V
[0105] Using the above method, the examples included 258 proteins initially screened and 5 traditional predictive factors. Parameter settings were used to penalize only proteins, and 5-fold cross-validation was employed to select the optimal parameter λ. The mean squared error (MSE) variations corresponding to different λ values are shown in [the table / link]. Figure 3 The two vertical lines represent λ.min and λ.1se, respectively. The former is the value of λ that minimizes the MSE, and the latter is the value of λ that minimizes the MSE to within one standard error of the minimum MSE while reducing model complexity. In this embodiment, λ.min = 0.02073961 is selected. Based on this value, 23 proteins suitable for predictive model construction are selected, and their corresponding coefficients are obtained (see Table 5).
[0106] Table 5: Proteins screened by the LASSO-Cox model and their coefficients
[0107]
[0108] Note: Cardiometabolic; Inflammation; Oncology; Neurology
[0109] (6) Predictive performance evaluation and verification
[0110] Based on the above traditional risk factors, RPG, GRS, and proteins, four T2D prediction models were constructed using the Cox proportional hazards regression model in the training set population, and predictions were performed on the test set:
[0111] ① The base model incorporates five traditional risk factors;
[0112] ② Blood glucose model (Base+RPG): RPG is added on top of the basic model;
[0113] ③ Genetic model (Base+GRS): The T2D GRS is added on top of the basic model;
[0114] ④ Protein model (Base+PRS): In addition to the basic model, a protein risk score (PRS) consisting of 23 proteins is added. The score weights are the regression coefficients obtained from the LASSO-Cox model in Table 5, and the calculation formula is shown in Equation VI.
[0115] Style VI
[0116] in, For the first The weights corresponding to each protein For the first No. 1 research subject The NPX value of each protein.
[0117] The area under the curve (AUC) and net reclassification index (NRI) of the four prediction models were calculated in the training set and the test set, respectively, to evaluate the actual prediction effect of the prediction models. The time-dependent AUC results showed that the blood glucose model and the protein model could significantly improve the prediction performance of the basic model for the incidence of T2D in the future 3 years, 5 years, 10 years and 15 years, and the improvement degree of the protein model was higher than that of the blood glucose model Figures 4-6 ); the genetic model did not significantly improve the prediction performance of the basic model. For NRI, it can be seen that the protein model has a greater improvement on the basic model in predicting the incidence of T2D in the future 3 years, 5 years, 10 years and 15 years, while the blood glucose model and the genetic model have no significant improvement in some years of the training set or the test set (Table 6). The above results consistently show that the method and the constructed model of the present application can accurately predict the multi-time point incidence risk of T2D.
[0118] Table 6: NRI results of T2D prediction models in training and test sets
[0119]
[0120] (7) Related parameters of T2D prediction model
[0121] The risk score obtained by the protein model is calculated according to the following formula:
[0122] The risk score of the protein model = 0.033150918 x age + 0.499399374 (gender == 1) - 0.549186868 (education level == 2) - 0.636837765 (education level == 3) + 1.037613865 (current smoker == 1) + 0.022096412 x waist circumference + 1.730733072 x PRS - 2.941020761.
[0123] In order to facilitate the application of the model, taking the incidence risk of 10% in the future 3 years, 5 years, 10 years and 15 years as an example, the cut-off values of the risk score of the protein model are 3.6139, 3.0976, 2.1760 and 1.3391, respectively. According to the series of cut-off values, the population is divided into high and low risk groups, and there is a significant difference in the incidence risk of T2D between the population groups Figure 7 ). The risk score calculation formula and the corresponding cut-off values of the other models are shown in Table 7.
[0124] Table 7: Formula and cut-off value of each prediction model
[0125]
[0126] (8) Protein level detection device for T2D predictor
[0127] The present embodiment provides a protein level detection device for the aforementioned T2D prediction. According to actual needs, the protein detection device can be prepared into various optional detection kits, and the specific form and detection method are not limited, such as ELISA, immunofluorescence kit, or detection by mass spectrometry, nucleic acid amplification, etc.
[0128] In the above kit, taking the ELISA kit as an example, the kit should contain a solid phase carrier, preferably an enzyme-labeled plate (such as a 96-well or other size polystyrene microwell plate), a membrane carrier (such as a nitrocellulose membrane, a glass fiber membrane, or a nylon membrane), or a microsphere, and the solid phase carrier is pre-coated with specific antibodies against the protein to be detected. In the embodiment, the membrane carrier can also be provided with a positive control to facilitate detection and result judgment. The kit should further include an anti-target primary antibody and an enzyme-labeled secondary antibody, and the latter can be a horseradish peroxidase (HRP) or an alkaline phosphatase-labeled secondary antibody.
[0129] Various forms of kits all have the ability of protein quantitative detection, and are especially suitable for quantitative analysis of various protein markers in blood samples. Taking the ELISA kit as an example, when detecting the protein marker LPL in serum, it first binds to the LPL-specific antibody pre-coated on the solid phase carrier. Then, an enzyme-labeled secondary antibody is added, which also binds to the complex and is fixed on the solid phase surface. At this time, the amount of enzyme bound to the solid phase surface is positively correlated with the concentration of LPL protein in the blood sample. After adding a specific enzyme substrate, a colored product is generated under the catalysis of the enzyme, and the color depth reflects the content of the protein to be detected. Since the enzyme reaction has extremely high catalytic efficiency, the signal is significantly amplified, improving the detection sensitivity and accuracy, and realizing the qualitative and quantitative analysis of the target protein.
[0130] The above only describes the preferred embodiments of the present application, and it should be noted that for those skilled in the art, without departing from the principles of the present application, several improvements and refinements can be made, and these improvements and refinements should also be considered within the protection scope of the present application.
Claims
1. Use of a protein marker for the manufacture of a reagent, kit or device for the prediction of the onset of type 2 diabetes, characterized in that, the protein markers are FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDAR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3; the UniProt accession numbers of the FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDAR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3 are Q96MK3, Q6QNK2, Q6P1J6, O15031, Q13478, Q15517, Q13316, Q9NQ79, P05107, P06858, Q8N4F0, Q9NQ30, Q9UI42, Q9Y5Q6, Q9UNE0, P23141, Q9UJA9, Q96C92, Q9P0G3, P01178, P19021, P22079 and Q9UM47, respectively.
2. Use according to claim 1, wherein the type 2 diabetes onset prediction comprises PRS-based prediction; the PRS = NPX1 × 0.355114 + NPX2 × 0.313233 - NPX3 × 0.30225 + NPX4 × 0.284929 + NPX5 × 0.265607 - NPX6 × 0.24851 - NPX7 × 0.24182 - NPX8 × 0.21008 + NPX9 × 0.198402 - NPX10 × 0.14131 + NPX11 × 0.125386 - NPX12 × 0.0958 - NPX13 × 0.08129 + NPX14 × 0.054934 + NPX15 × 0.052256 + NPX16 × 0.042321 - NPX17 × 0.03355 + NPX18 × 0.030498 - NPX19 × 0.02234 + NPX20 × 0.011776 - NPX21 × 0.00255 + NPX22 × 0.000875 - NPX23 × 0.00019; Wherein, NPX1 to NPX23 are NPX values of FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDAR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3 of the subject in sequence.
3. Use according to claim 2, wherein the compound is ###0002### The type 2 diabetes onset prediction comprises prediction based on the age, gender, education level, waist circumference and the PRS of the subject.
4. The use according to claim 3, wherein the compound is ###0002### The type 2 diabetes onset prediction is predicted by a risk score; The risk score = the PRS × 1.730733072 + age factor + gender factor + education level factor + smoking status factor + waist circumference factor - 2.941020761; The age factor takes the value of the age value of the subject × 0.033150918; The value of the gender factor is 0 if the gender of the subject is male, and is 0.499399374 if the gender of the subject is female; The value of the education level factor is 0 if the education level of the subject is primary school or below, is -0.549186868 if the education level of the subject is junior high school or high school, and is -0.636837765 if the education level of the subject is university or above; The value of the smoking status factor is 0 if the subject is a non-current smoker, and is 1.037613865 if the subject is a current smoker; The waist circumference factor is the waist circumference value of the subject × 0.022096412, and the unit of the waist circumference value is centimeter.
5. A reagent for the prediction of the onset of type 2 diabetes, characterized in that, Specific antibodies of FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDAR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3; The UniProt accession numbers of the FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDAR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3 are Q96MK3, Q6QNK2, Q6P1J6, O15031, Q13478, Q15517, Q13316, Q9NQ79, P05107, P06858, Q8N4F0, Q9NQ30, Q9UI42, Q9Y5Q6, Q9UNE0, P23141, Q9UJA9, Q96C92, Q9P0G3, P01178, P19021, P22079 and Q9UM47, respectively.
6. A kit for the prediction of the onset of type 2 diabetes, characterized in that, The reagent of claim 5.
7. Device for the prediction of the onset of type 2 diabetes, characterized in that, The reagent of claim 5.
8. The apparatus of claim 7, wherein, Also included are: an acquisition module, configured to acquire the NPX values of FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDAR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO and NOTCH3 of the subject; an analysis module, configured to calculate a risk score and predict the risk of the subject of developing type 2 diabetes based on the risk score; the risk score = PRS × 1.730733072 + age factor + gender factor + education level factor + smoking status factor + waist circumference factor - 2.941020761; the age factor is the age value of the subject × 0.033150918; the value of the gender factor is 0 if the subject is male, and is 0.499399374 if the subject is female; the value of the education level factor is 0 if the subject has a primary school education or below, is -0.549186868 if the subject has a junior high school or high school education, and is -0.636837765 if the subject has a university education or above; the value of the smoking status factor is 0 if the subject is a non-current smoker, and is 1.037613865 if the subject is a current smoker; the waist circumference factor is the waist circumference value of the subject × 0.022096412, and the unit of the waist circumference value is centimeter; the PRS = NPX1 x 0.355114 + NPX2 x 0.313233 - NPX3 x 0.30225 + NPX4 x 0.284929 + NPX5 x 0.265607 - NPX6 x 0.24851 - NPX7 x 0.24182 - NPX8 x 0.21008 + NPX9 x 0.198402 - NPX10 x 0.14131 + NPX11 x 0.125386 - NPX12 x 0.0958 - NPX13 x 0.08129 + NPX14 x 0.054934 + NPX15 x 0.052256 + NPX16 x 0.042321 - NPX17 x 0.03355 + NPX18 x 0.030498 - NPX19 x 0.02234 + NPX20 x 0.011776 - NPX21 x 0.00255 + NPX22 x 0.000875 - NPX23 x 0.00019; wherein NPX1 to NPX23 are NPX values of FAM20A, ADGRD1, PLB1, PLXNB2, IL18R1, CDSN, DMP1, CRTAC1, ITGB2, LPL, BPIFB2, ESM1, CPA4, INSL5, EDAR, CES1, ENPP5, ENTR1, KLK14, OXT, PAM, LPO, and NOTCH3, respectively, of the subject.
Citation Information
Patent Citations
Application of protein marker in preparation of product for predicting future coronary heart disease onset risk of subject
CN119959557A
Biomarkers for predicting type 2 diabetes status
US20250208147A1