Protein marker and clinical indicator combination, product and application thereof
By combining protein markers and clinical indicators, a binary logistic regression model was constructed, which solved the problems of high cost of MRI examinations and limited predictive ability of traditional indicators, and achieved rapid and economical assessment of white matter lesions with high predictive accuracy.
Patent Information
- Application Number
- CN202510825808.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-06-19
AI Technical Summary
Existing technologies make it difficult to accurately assess and predict the extent of white matter lesions. MRI examinations are expensive and equipment accessibility is limited. Traditional hematological indicators have limited predictive capabilities, making large-scale screening and early identification difficult to achieve.
A combination of protein markers (P05198, P00966, Q8IZP0, and P48061) and related clinical indicators (MMSE, CTTtrail1, and SDMT) was provided. A prediction model was constructed using a binary logistic regression model, and the optimal combination was screened using machine learning algorithms (XGBoost and LASSO) for the rapid and convenient assessment of white matter lesions.
A convenient and economical assessment of white matter lesions without relying on MRI has been achieved. The prediction model performs well, has high accuracy and feasibility, and is suitable for large-scale screening and early identification.
Smart Images

Figure CN120801716A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of biological medicine combined with machine learning, and particularly relates to a combination of protein markers and clinical indicators related to white matter lesions, a product and an application thereof. BACKGROUND
[0002] White matter lesions (WML) is a brain pathological change closely related to vascular risk factors and aging, which is characterized by demyelination, axonal degeneration and cerebral small vessel disease in the white matter region. WML is an important risk factor for vascular cognitive impairment, stroke, Alzheimer's disease, Parkinson's disease and other neurodegenerative diseases. WML is mainly detected by head magnetic resonance imaging (MRI) T2 weighted imaging or fluid attenuated inversion recovery (FLAIR) sequence, which shows point-like, patchy or fused lesions with high signal. The detection of white matter lesions is often graded by Fazekas score for the degree of damage. White matter lesions are common in the elderly population. In the population of 60-70 years old, more than 87% of them have WML; and in the population of 80-90 years old, the detection rate of WML is as high as 95%-100%. A large number of studies have shown that vascular risk factors such as hypertension, diabetes, hyperlipidemia, coronary heart disease and smoking can promote the occurrence and accelerate the progression of WML. However, in individuals exposed to these vascular risk factors for a long time, there is still a large difference in the severity of white matter damage, suggesting that the susceptibility of individuals to vascular risk factors may be affected by other biological factors.
[0003] Although white matter damage is common in the elderly exposed to vascular risk factors for a long time, there is a significant individual difference in the severity of the damage. This heterogeneity may be related to genetic factors, lifestyle, environmental exposure and other factors. Therefore, it is difficult to accurately assess and predict the severity of white matter damage by relying only on traditional clinical indicators. Currently, the assessment of white matter damage mainly relies on head magnetic resonance imaging (MRI) technology, and Fazekas score is a commonly used semi-quantitative assessment method. Fazekas score divides white matter damage into mild (score < 2) and severe (score ≥ 2).
[0004] However, the high cost of MRI examination, long scanning time and limited equipment accessibility limit the feasibility of large-scale screening and early identification, so a more convenient and economical biomarker is needed to assist in the assessment of white matter lesions. On the other hand, although traditional hematological indicators (such as blood lipids, blood glucose, inflammatory markers, etc.) are related to WML, the predictive ability of a single indicator is limited, and it is difficult to accurately reflect the pathological changes of individual WML. Therefore, exploring new biomarkers to establish a simple, economical and generalizable blood test model has important clinical significance for the early screening and disease management of WML. SUMMARY
[0005] In view of the above technical problems, the present application provides a combination of protein markers and clinical indicators related to white matter lesions, products and applications thereof.
[0006] The technical solutions provided by the present application are as follows: In a first aspect, the present application provides a protein marker combination, comprising: P05198, P00966, Q8IZP0 and P48061.
[0007] In a second aspect, the present application provides an application of a reagent for detecting a protein marker combination in the preparation of a product for diagnosing white matter lesions, wherein the protein marker combination comprises P05198, P00966, Q8IZP0 and P48061.
[0008] In a third aspect, the present application provides a combination of protein markers and clinical indicators, wherein the protein markers comprise: P05198, P00966, Q8IZP0 and P48061; and the clinical indicators comprise MMSE, CTTtrail1 and SDMT.
[0009] In a fourth aspect, the present application provides an application of a reagent for detecting a combination of protein markers and clinical indicators in the preparation of a product for diagnosing white matter lesions, wherein the protein marker combination comprises P05198, P00966, Q8IZP0 and P48061; and the clinical indicators comprise MMSE, CTTtrail1 and SDMT.
[0010] In a fifth aspect, the present application provides a kit comprising detection reagents for detecting the combination of protein markers and clinical indicators of the third aspect.
[0011] In a sixth aspect, the present application provides an application of the kit of the fifth aspect in the preparation of a product for detecting white matter lesions, and an application of the detection reagents in the kit of the fifth aspect in the preparation of a kit for diagnosing white matter lesions.
[0012] In a seventh aspect, the present application provides a program product related to white matter lesion, which is used to diagnose the risk of a subject suffering from white matter lesion, comprising the following steps: obtaining the expression amount of each protein marker in the plasma of the subject; the protein markers include P05198, P00966, Q8IZP0 and P48061; substituting the expression amount of each protein marker and the value of the clinical index into a binary logistic regression equation to calculate the log of the odds of the subject y; calculating the probability P of the subject being a healthy person according to y, P = exp(y) / {1 + exp(y)}, and exp(y) represents a natural exponential function; diagnosing or predicting whether the subject suffers from white matter lesion or has the risk of suffering from white matter lesion according to the comparison result of the probability P and the reference value.
[0013] In a possible implementation manner, the formula of the binary logistic regression equation is as follows:
[0014] wherein A is an intercept term, B1-B4 are regression coefficients of independent variables; x1, x2, x3 and x4 are the expression detection values of the proteins P05198, P00966, Q8IZP0 and P48061 respectively.
[0015] Further, the method for obtaining the binary logistic regression equation comprises the following steps: obtaining the expression amount of the protein markers of the healthy person and the patient with white matter lesion; performing binary logistic regression training by taking the expression amount of the protein markers as input and whether suffering from the disease as output to obtain the binary logistic regression equation.
[0016] In an eighth aspect, the present application provides a program product related to white matter lesion, which is used to diagnose the risk of a subject suffering from white matter lesion, comprising the following steps: obtaining the expression amount of each protein marker and the value of each clinical index in the plasma of the subject; the protein markers include P05198, P00966, Q8IZP0 and P48061; and the clinical indexes include MMSE, CTTtrail1 and SDMT; substituting the expression amount of each protein marker and the value of the clinical index into a binary logistic regression equation to calculate the log of the odds of the subject y; calculating the probability P of the subject being a healthy person according to y, P = exp(y) / {1 + exp(y)}, and exp(y) represents a natural exponential function; According to the comparison result of the probability P and the reference value, whether the to-be-tested object has leukoaraiosis or has the risk of having leukoaraiosis is diagnosed or predicted.
[0017] In a possible implementation manner, a formula of the binary logistic regression equation is as follows:
[0018] wherein, A is an intercept term, B1-B7 are regression coefficients of independent variables; x1, x2, x3, x4 are detection values of expression amounts of proteins P05198, P00966, Q8IZP0 and P48061 respectively, and x5, x6, x7 are detection values of clinical indexes SDMT, MMSE and CTTtrail1 respectively.
[0019] Further, the method for obtaining the binary logistic regression equation comprises the following steps. obtaining expression amounts of protein markers and numerical values of clinical indexes of healthy people and patients with leukoaraiosis; using the expression amounts of protein markers and the numerical values of clinical indexes as inputs and whether having the disease as output, performing binary logistic regression training to obtain a binary logistic regression equation.
[0020] In a ninth aspect, the application provides a method for constructing a combination model of protein markers and clinical indexes, comprising the following steps. obtaining clinical index information and protein expression amount data in plasma of patients with leukoaraiosis; screening first differentially expressed protein data by comparing with protein expression amount data of healthy people; calculating importance scores and rankings of the differentially expressed proteins in differentiating different groups of samples using an XGBoost machine learning model and a LASSO regression model, and screening second differentially expressed protein data; randomly combining the second differentially expressed protein data and the clinical index information as inputs and whether having the disease as output, and constructing multiple combination models using logistic regression; screening an optimal combination model based on AUC values.
[0021] The application has the following beneficial effects: (1) The application provides a combination, product and application of protein markers and clinical indexes for evaluating the degree of leukoaraiosis, which does not depend on clinical scales and neuroimaging, and can be used for quickly, conveniently and economically identifying leukoaraiosis.
[0022] (2) The application also provides a computer program product, which can use the combination of protein markers and clinical indicators to construct a prediction model to effectively evaluate the degree of white matter lesions. The prediction model has good performance and shows relatively accurate prediction ability, and has clinical use and promotion value. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 Statistical number of protein quantification in Example 1; Figure 2 Screening of differential proteins in Example 1; Figure 3 Prediction model parameters of protein marker combination in Example 1; Figure 4 Prediction model parameters of protein marker and clinical indicator combination in Example 1; Figure 5 Nomogram of protein marker combination model; Figure 6 ROC analysis diagram of prediction model of protein marker combination; Figure 7 Confusion matrix diagram of prediction model of protein marker combination; Figure 8 Score box plot of prediction model of protein marker combination; Figure 9 Score box plot of prediction model of protein marker and clinical indicator combination; Figure 10 ROC analysis diagram of prediction model of protein marker and clinical indicator combination; Figure 11 Confusion matrix analysis diagram of prediction model of protein marker and clinical indicator combination. DETAILED DESCRIPTION
[0024] The content of the application will be further described below in combination with specific embodiments, and the content of the application is not limited thereto.
[0025] At present, the high cost of MRI examination, long scanning time and device accessibility limit the feasibility of large-scale screening and early identification.
[0026] Therefore, the application provides a combination of protein markers and clinical indicators related to white matter lesions, products and applications thereof.
[0027] For ease of understanding, the relevant professional terms appearing in the examples are uniformly explained.
[0028] Protein marker: P05198: Gene name EIF2S1, Protein name Eukaryotic translation initiation factor 2 subunit 1; P00966: Gene name ASS1, Protein name Argininosuccinate synthase; Q8IZP0: Gene name ABI1, Protein name Abl interactor 1; P48061: Gene name CXCL12, Protein name Stromal cell-derived factor 1.
[0029] Clinical indicators: CI: Cognitive impairment SDMT: Symbol Digit Modalities Test MMSE: Mini-Mental State Examination HLP: Hyperlipoproteinemia education_year: Years of education MoCA: Montreal Cognitive Assessment CTT_Trail1: Color Trail Test, Number 1 smoke_status: Smoking status CTT_Trail2: Color Trail Test, Number 2 sex: Sex CDT3: Clock Drawing Test (3 minutes) TMT_B: Trail Making Test B TMT_A: Trail Making Test A BNT_15: Boston Naming Test (15 items) CHD: Coronary heart disease AF: Atrial fibrillation VHD: Valvular heart disease age: Age HAMA: Hamilton Anxiety Scale HAMD: Hamilton Depression Scale diabetes: Diabetes mellitus drink_status: Drinking status BMI: Body mass index, weight ÷ height squared MI: Myocardial infarction hypertension: Hypertension TIA: Transient ischemic attack PVD: peripheral vascular disease HCY: hyperhomocysteinemia Example 1: Screening of white matter lesion markers and construction of diagnostic model 1. Sample information The samples used in this example were derived from 181 patients in Zhongnan Hospital of Wuhan University. The patient information is shown in Table 1.
[0030] Table 1. Patient clinical information table
[0031] Note: Group 0: healthy control group, Group 1: white matter lesion group; sex: 0 represents male, 1 represents female; smoke_status: 0 represents current non-smoker, 1 represents smoker; drink_status: 0 represents current non-drinker, 1 represents drinker; hypertension, diabetes, HLP, MI, CHD, AF, VHD, PVD, TIA, CI, HCY: 0 represents not sick, 1 represents sick; CDT3: 0, 1, 2, 3 respectively indicate the score; categorical variable information, such as the sex (%) column, the numerical value represents the number of patients in each category, such as 70, which means that there are 70 males among all patients; continuous variable information, such as the age (median [IQR]) column, the numerical value represents the median, such as 65.00, which means that the median age of all patients is 65; other numerical data are the same.
[0032] 2. Sample pretreatment Take 20 microliters of plasma sample, dilute with loading buffer (10 mM Tris-Cl, 1 mM EDTA, 150 mM KCl, 0.05% CHAPS), mix with 1 mg of superparamagnetic iron oxide nanobead suspension, and incubate at 37°C for 1 hour. Wash the magnetic beads twice with loading buffer, and then wash once with loading buffer without CHAPS [3-[(3-cholamidopropyl)dimethylammonio]-1-propanesulfonic acid inner salt] (10 mM Tris-Cl, 1 mM EDTA, 150 mM KCl). Absorb the magnetic beads on a magnetic stand, and discard the supernatant to obtain the protein-enriched nanobeads. Add lysis solution (1% SDC / 100 mM Tris-HCl, pH=8.5 / 10 mM TCEP / 40 mM CAA) to the sample, and incubate at 60°C for 30 min to perform the reduction alkylation reaction. Add an equal volume of ddH2O to dilute the SDC (sodium deoxycholate) concentration to below 0.5%, add 1 microgram of trypsin, and incubate at 37°C overnight with shaking to perform the enzymatic digestion. The next day, add TFA to terminate the enzymatic digestion, take the supernatant to perform SDB-RPS desalting column desalting, and then vacuum dry and store at -20°C for later use.
[0033] 3. LC-MS / MS detection and analysis The mass spectrometric detection of the sample used an UltiMate 3000 RSLCnano nanoflow liquid chromatograph (Thermo) coupled to a timsTOF Pro mass spectrometer (Bruker). The peptide sample was injected by an autosampler, which was coupled to a C18 trapping column (75 µm x 2 cm, 3 µm particle size, 100 Å pore size, Thermo), and then entered an analytical column (75 µm x 15 cm, 1.7 µm particle size, 100 Å pore size, IonOpticks) for separation. The analysis gradient was established using mobile phase A (0.1% formic acid) and mobile phase B (0.1% formic acid in ACN). The mass spectrometer was used for data acquisition in diaPASEF mode. The capillary voltage was set to 1500 V. The scan range of MS1 and MS2 spectra was set to 100-1700 m / z. The ion mobility range was set to 0.6-1.6 Vs / cm 2 . The accumulation time and ramp time were set to 50 ms. According to the distribution of mass-to-charge ratio-ion mobility, the diaPASEF acquisition window was set using the timsControl software. The collision energy was set according to the ion mobility from 1 / K0=1.6 Vs / cm 2linear decrease from 59 eV to 1 / K0=0.6 Vs / cm^2 at 20 eV. DIA raw data files were further obtained.
[0034] 4. Data preprocessing DIA raw data files were analyzed by DIA-NN software (v 1.8). The database used for search was the proteome reference database of Human in Uniprot (2022-02-09, containing 20375 protein sequences). A spectral library was predicted by the deep learning algorithm in DIA-NN, and the spectral library was extracted from the DIA raw data using the predicted spectral library and the spectral library obtained by MBR function. Protein expression quantification values were calculated based on mass spectrometry detection signals. The final results were screened at the parent ion and protein level with 1% FDR. The proteome quantification information after screening was used for subsequent analysis. As shown in Figure 1 , 957-2912 different proteins were quantified in each sample, respectively.
[0035] 5. Differential protein screening The differential expression proteins between the diseased group and the non-diseased group were screened with the criteria of greater than 1.2-fold change in protein expression abundance and less than 0.05 in T-test P value. As shown in Figure 2 , the expression of 329 proteins was significantly different between the population with higher degree of white matter lesions and the population with lower degree of white matter lesions, 190 proteins were up-regulated in the high lesion group, and 139 proteins were down-regulated in the high lesion group.
[0036] 6. XGBoost and LASSO analysis The importance scores and rankings of the above differential proteins in distinguishing different groups of samples were calculated using the XGBoost machine learning model and the LASSO regression model, and 41 most useful diagnostic protein features were further screened out as plasma protein markers that can be used to distinguish the diseased group from the non-diseased group. The information of each protein marker is shown in Table 2.
[0037] Table 2 Marker protein information table
[0038] 7. Logistic regression analysis The above protein markers and clinical indicators were randomly combined, and a combination model was constructed using binary logistic regression. The optimal combination of protein markers and clinical indicators was selected based on the AUC value. The optimal combination was a combination of four proteins and three clinical indicators (P05198 + P00966 + Q8IZP0 + P48061 + SDMT + MMSE + CTT_Trail1). A clinical diagnostic model was constructed based on the above protein markers and clinical indicators. The parameters of the protein combination model and the protein clinical indicator combination model are as follows: Figure 3 、 4 As shown. Based on the combined model, a nomogram is established to demonstrate the application of the model in diagnosis, as shown Figures 5-8 shown.
[0039] The protein combination model is shown in formula (1): (1) The model of protein and clinical index combination is shown in formula (2): (2) Example 2: Testing and evaluation of a white matter lesion assessment model This example uses 78 subjects from Zhongnan Hospital of Wuhan University as a test set, and uses a combination model of 4 proteins and 3 clinical indicators (Formula (2) in Example 1) to verify the effect of evaluating the degree of white matter lesions.
[0040] The clinical information of the subjects is shown in Table 3.
[0041] Table 3 Clinical information of test subjects
[0042] The sample source, pretreatment, LC-MS / MS detection and data preprocessing in Example 2 are the same as those in Example 1.
[0043] The test results are as follows Figures 9-11 As shown in Figure 2, the AUC was 0.776 (0.671-0.881). The results showed that the model had good discrimination and could effectively distinguish between the high-risk and low-risk groups of white matter lesions. The sensitivity of the model reached 65.7% and the specificity was 74.4%, indicating that the model had good predictive ability and accuracy.
[0044] Example 3: Kit The present embodiment provides a kit comprising reagents for detecting protein markers (P05198 + P00966 + Q8IZP0 + P48061) and clinical indicators (SDMT + MMSE + CTT_Trail1_time).
[0045] 3.1. Reagents for detecting protein markers Loading buffer 1 (10 mM Tris-Cl, 1 mM EDTA, 150 mM KCl, 0.05% CHAPS); Loading buffer 2 (10 mM Tris-Cl, 1 mM EDTA, 150 mM KCl); Lysis solution (1% SDC / 100 mM Tris-HCl, pH=8.5 / 10 mM TCEP / 40 mM CAA); Trypsin; SDB-RPS desalting column.
[0046] 3.2. Reagents for detecting clinical indicators Reagents used for detecting relevant clinical indicators.
[0047] Embodiment 4: Computer program A program product related to white matter lesions, the computer program product is used for executing diagnosis of the risk of the subject suffering from white matter lesions, comprising the following steps: (1) obtaining the expression amount of each protein marker and the value of each clinical indicator in the plasma of the subject to be tested; the protein markers include P05198, P00966, Q8IZP0 and P48061; the clinical indicators include MMSE, CTTtrail1 and SDMT; (2) substituting the expression amount of each protein marker and the value of the clinical indicator into a binary logistic regression equation to calculate the log of the advantage of the subject y; The binary logistic regression equation is as follows:
[0048] Wherein, A is the intercept term, B1-B6 are the regression coefficients of the independent variables; x1, x2, x3, x4 are the detection values of the expression amounts of proteins P05198, P00966, Q8IZP0 and P48061 respectively, x5, x6, x7 are the detection values of the clinical indicators SDMT, MMSE and CTTtrail1 respectively.
[0049] A=1.51, B1=0.41, B2=-0.3, B3=-0.22, B4=0.12, B5=-0.08, B6=-0.04, B7=0.02.
[0050] (3) calculating the probability P that the to-be-tested subject is a healthy person according to y, P=exp(y) / {1+exp(y)}, exp(y) represents a natural exponential function; (4) diagnosing or predicting whether the to-be-tested subject has leukoaraiosis or has a risk of having leukoaraiosis according to the comparison result of the probability P and the reference value.
[0051] It can be understood that the probability P generally takes a value of 0.5, and can also be adjusted according to actual conditions.
[0052] The above description is only the preferred specific implementation of the present application, but the scope of protection of the present application is not limited to this. Any modification, equivalent replacement and improvement made by any person skilled in the art within the technical range disclosed by the present application shall be included in the protection scope of the present application.
Claims
1. A protein marker combination, characterized in that: include: P05198, P00966, Q8IZP0, and P48061.
2. Use of a reagent for detecting a protein marker combination in the preparation of a product for diagnosing white matter lesions, wherein the protein marker combination includes P05198, P00966, Q8IZP0, and P48061.
3. A combination of a protein marker and a clinical indicator, characterized in that: The protein markers include: P05198, P00966, Q8IZP0 and P48061; the clinical indicators include Mini-Mental State Examination (MMSE), Color-Digit Trail Test (CTTtrail1) and Symbol-Digit Translation Test (SDMT).
4. Use of a reagent for detecting a combination of protein markers and clinical indicators in the preparation of a product for diagnosing white matter lesions, wherein the protein marker combination includes P05198, P00966, Q8IZP0, and P48061; and the clinical indicators include MMSE, CTTtrail1, and SDMT.
5. A kit, characterized in that The kit comprises a detection reagent for detecting the combination of the protein marker and clinical indicator according to claim 3.
6. Use of the kit according to claim 5 in the preparation of a product for detecting white matter lesions, and use of the detection reagent in the kit according to claim 5 in the preparation of a kit for diagnosing white matter lesions.
7. A program product related to white matter lesions, characterized in that: The computer program product is used to diagnose the risk of a subject suffering from white matter lesions, comprising the following steps: Obtaining the expression level of each protein marker and the numerical value of each clinical indicator in the plasma of the test subject; the protein markers include P05198, P00966, Q8IZP0 and P48061; the clinical indicators include MMSE, CTTtrail1 and SDMT; Substituting the expression level of each protein marker and the numerical value of the clinical index into a binary logistic regression equation to calculate the logarithm y of the odds of the subject to be tested; Calculate the probability P of the subject being healthy based on y, P = exp(y) / {1+exp(y)}, where exp(y) represents the natural exponential function; Based on the comparison result of the probability P with the reference value, it is diagnosed or predicted whether the subject to be tested suffers from white matter lesions or has the risk of suffering from white matter lesions.
8. The program product according to claim 7, wherein The formula of the binary logistic regression equation is: Among them, A is the intercept term, B1-B7 are the regression coefficients of the independent variables; x1, x2, x3, and x4 are the expression levels of proteins P05198, P00966, Q8IZP0, and P48061, respectively; x5, x6, and x7 are the detection values of clinical indicators SDMT, MMSE, and CTTtrail1, respectively.
9. The program product according to claim 7 or 8, characterized in that The method for obtaining the binary logistic regression equation includes: Obtain the expression levels of protein markers and clinical indicators of healthy controls and patients with white matter lesions; The expression levels of protein markers and the values of clinical indicators were used as inputs, and whether the patient was ill was used as output. Binary logistic regression training was performed to obtain a binary logistic regression equation.
10. A method for constructing a combined model of protein markers and clinical indicators, characterized in that: include: Obtain clinical indicator information and plasma protein expression data of patients with white matter lesions; By comparing the protein expression data with that of healthy people, the first differentially expressed protein data was screened out; The XGBoost machine learning model and LASSO regression model were used to calculate the importance scores and rankings of differentially expressed proteins in distinguishing different group samples, and the second differentially expressed protein data were screened out; The second differentially expressed protein data and clinical indicator information are randomly combined as input, and whether the disease is serious is used as output, and multiple combination models are constructed using logistic regression; Based on the AUC value, the optimal combination model was screened out.
Citation Information
Patent Citations
Method for genome editing
CN102858985A
Probe set and kit for detecting whole exons of extended genetic diseases and application of probe set
CN110499364A
Biomarker combination for white matter lesion and application of biomarker combination
CN114264757A
Vascular depression recognition model construction method and system
CN118016271A
Genomic editing of neurodevelopmental genes in animals
US20110023143A1