Protein marker for diagnosing cognitive impairment of heavy metal-related middle-aged and elderly people, and application, kit and risk assessment model thereof
By identifying protein markers related to cadmium-induced cognitive impairment, providing a combination of TSKU, SOD3, CES2, PSMB10 and C1QA as diagnostic markers, it solves the problem of difficulty in early diagnosis of heavy metal-related cognitive impairment in the prior art, and achieves efficient diagnosis and risk assessment.
Patent Information
- Application Number
- CN202510255415.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to diagnose cognitive impairment in heavy metal-related middle-aged and elderly people in early stages, and the lack of accurate and efficient protein markers has led to the failure to meet clinical needs.
By identifying and analyzing protein markers associated with cadmium-induced cognitive impairment response, a combination of TSKU, SOD3, CES2, PSMB10 and C1QA is provided as diagnostic markers, and early diagnostic kits and risk assessment models are developed.
It has achieved early diagnosis of cognitive impairment in heavy metal-related elderly people, improved the sensitivity and efficiency of detection, and can delay the progression of cognitive impairment and reduce the disability mortality rate.
Smart Images

Figure CN120044253A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical technologies, and particularly to a protein biomarker for diagnosing cognitive impairment in middle-aged and elderly people related to heavy metals, its application, kit, and risk assessment model. Background Art
[0002] Cognitive impairment is a disease state, manifested as the cognitive functions (such as memory, attention, thinking, judgment, understanding, calculation, learning ability, spatial orientation ability, and language ability, etc.) of patients being significantly lower than their normal levels, and being sufficient to affect their daily life and social functions. The etiology of cognitive impairment is unknown and may involve multiple factors such as genetics, environment, and lifestyle. Previous studies have shown that cadmium can enter the body through inhalation, oral ingestion, and skin contact. Then, a small amount of cadmium can enter the brain tissue through the blood-brain barrier, inducing neurodegenerative lesions and memory defects in the brain, etc., causing the occurrence and progression of cognitive impairment in the body.
[0003] Considering the large number of middle-aged and elderly people and their susceptibility to cognitive impairment, as well as the rapid growth trend of the latter in China, finding protein biomarkers for early diagnosis of cognitive impairment in middle-aged and elderly people related to heavy metals has become an important challenge. In the international and domestic markets, there are no protein biomarkers available for such early diagnosis, which severely restricts the satisfaction of clinical needs. Therefore, researchers urgently need to strive to find more accurate and efficient protein biomarkers to discover and improve the current situation of cognitive impairment in middle-aged and elderly people related to heavy metals in a timely manner. Summary of the Invention
[0004] This study mainly explored the role of proteins in the process of cognitive impairment response in middle-aged and elderly groups after cadmium exposure, with particular attention to protein markers related to cadmium-induced cognitive impairment response. By identifying and analyzing these markers, this study aims to provide a new method and tool for early diagnosis of the possibility of cognitive impairment in middle-aged and elderly groups due to cadmium exposure.
[0005] On the one hand, the present invention provides a protein biomarker for diagnosing cognitive impairment in middle-aged and elderly people related to heavy metals, and the protein biomarker is a combination of TSKU, SOD3, CES2, PSMB10, and C1QA.
[0006] Among them, compared with healthy controls, the expression levels of TSKU and CES2 are down-regulated in patients with cognitive impairment in middle-aged and elderly people related to heavy metals, and the expression levels of SOD3, PSMB10, and C1QA are up-regulated in patients with cognitive impairment in middle-aged and elderly people related to heavy metals.
[0007] Among them, the source of the protein biomarker is tissue or blood, etc. Specifically, the test sample is from the plasma, serum or blood extract, brushings, biopsy, or surgically resected tissue or fluid sample of the subject, etc. Preferably, the source of the protein biomarker is blood.
[0008] On the other hand, the present invention discloses the application of the aforementioned protein biomarker in a kit for early diagnosis of cognitive impairment in middle-aged and elderly people related to heavy metals.
[0009] Among them, the kit for early diagnosis of cognitive impairment in middle-aged and elderly people related to heavy metals includes detection reagents capable of detecting the expression levels of TSKU, SOD3, CES2, PSMB10, and C1QA. Among them, the detection reagents include reagents for detecting protein levels by methods such as immunoblotting, enzyme-linked immunosorbent assay, mass spectrometry, radioimmunoassay, radioimmunodiffusion, immunoelectrophoresis, tissue immunostaining, immunoprecipitation analysis, complement fixation analysis, fluorescence-activated cell sorting, mass analysis, or protein microarray.
[0010] Specifically, the detection reagent is a binder that specifically binds to the proteins encoded by TSKU, SOD3, CES2, PSMB10, and C1QA. More specifically, the binder for the protein includes peptides, peptidomimetics, aptamers, spiegelmers, darpins, ankyrin repeat proteins, Kunitz-type domains, antibodies, single-domain antibodies, and / or monovalent antibody fragments, etc.
[0011] On the other hand, the present invention discloses a risk assessment model for cognitive impairment in middle-aged and elderly people related to heavy metals. The risk assessment model uses the expression levels of the aforementioned protein biomarkers (a combination of TSKU, SOD3, CES2, PSMB10, and C1QA) as variables (corresponding to corresponding scores), and uses a nomogram to visualize the model results. The sum of the individual scores corresponding to different values of each protein biomarker is obtained to get the total score, and then it corresponds to the probability of cognitive impairment events occurring in middle-aged and elderly people related to heavy metals.
[0012] Those skilled in the art can implement and realize the steps of associating the biomarker level with a certain possibility or risk in different ways. Preferably, the measured concentrations of one or more protein biomarkers are combined mathematically, and the combined value is associated with the early diagnosis problem. The measured combinations of biomarker values can be combined by any suitable existing mathematical method.
[0013] A function for associating a biomarker combination with a disease, preferably using an algorithm developed and obtained by applying statistical methods. For example, suitable statistical methods are discriminant analysis (DA) (i.e., linear, quadratic, regular DA), Kernel methods (i.e., SVM), non-parametric methods (i.e., k-nearest neighbor classifier), PLS (partial least squares), tree-based methods (i.e., logistic regression, CART, random forest method, boosting / bagging method), generalized linear models (i.e., log regression), principal component-based methods (i.e., SIMCA), generalized additive models, fuzzy logic-based methods, neural network- and genetic algorithm-based methods, etc. A person skilled in the art will have no problem in selecting a suitable statistical method to evaluate the biomarker combination of the present invention and thereby obtaining a suitable mathematical algorithm.
[0014] Specifically, in the nomogram, the score range of TSKU is 20 - 15.5, corresponding to 0 - 39 in the score scale; the score range of SOD3 is 16.2 - 18.4, corresponding to 0 - 63 in the score scale; the score range of CES2 is 11.8 - 8.8, corresponding to 0 - 93 in the score scale; the score range of PSMB10 is 11.2 - 13.2, corresponding to 0 - 67 in the score scale; the score range of C1QA is 13.5 - 18, corresponding to 0 - 100 in the score scale; the probabilities of cognitive impairment events occurring in the elderly population related to heavy metals take values of 0.01, 0.1, 0.3, 0.5, 0.7, and 0.9, corresponding to 89 - 91, 116 - 118, 132 - 135, 142 - 144, 150 - 152, and 166 - 168 in the total score scale respectively.
[0015] The beneficial effects of the present invention are as follows: By detecting the expression levels of TSKU, SOD3, CES2, PSMB10, and C1QA, the present invention can achieve early diagnosis of cognitive impairment in the elderly population related to heavy metals, increase the sensitivity of detection, improve the detection ability and efficiency, and actively take intervention measures, which can not only delay the progression of cognitive impairment, reduce the disability and fatality rates, but may even achieve the effect of early cure. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for describing the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0017] Figure 1 is a flowchart for model construction; Figure 2 is the differentially expressed proteins screened from the cognitive impairment group and the healthy control group; Figure 3It is the principal component analysis diagram; Figure 4 It is the relationship diagram between the fitting error and log(λ); Figure 5 It is the LASSO coefficient distribution diagram of differentially expressed proteins; Figure 6 It is the box plot of the protein expression level of the protein marker TSKU; Figure 7 It is the box plot of the protein expression level of the protein marker SOD3; Figure 8 It is the box plot of the protein expression level of the protein marker CES2; Figure 9 It is the box plot of the protein expression level of the protein marker PSMB10; Figure 10 It is the box plot of the protein expression level of the protein marker C1QA; Figure 11 It is the receiver operating characteristic curve ROC diagram of the early protein diagnosis model; Figure 12 It is the confusion matrix diagram of the early protein diagnosis model; Figure 13 It is the calibration curve diagram of the early protein diagnosis model; Figure 14 It is the decision curve DCA diagram of the early protein diagnosis model; Figure 15 It is the protein modeling nomogram for predicting the risk of cognitive impairment in middle-aged and elderly people related to heavy metals; Figure 16 It is the receiver operating characteristic curve ROC diagram of the early clinical diagnosis model; Figure 17 It is the confusion matrix diagram of the early clinical diagnosis model; Figure 18 It is the calibration curve diagram of the early clinical diagnosis model; Figure 19 It is the decision curve DCA diagram of the early clinical diagnosis model; Figure 20 It is the clinical modeling nomogram for predicting the risk of cognitive impairment in middle-aged and elderly people related to heavy metals; Figure 21 It is the analysis diagram of the proteomics of rat brain tissue predicted by heavy metal exposure. Detailed implementation manners
[0018] In the present invention, the protein markers include TSKU, SOD3, CES2, PSMB10, and C1QA. Protein markers such as TSKU (UniprotID: Q8WUA8), SOD3 (UniprotID: P08294), CES2 (UniprotID: O00748), PSMB10 (UniprotID: P40306), C1QA (UniprotID: P02745); including genes and their encoded proteins and their homologs, mutations, and isoforms. UniprotID can be obtained at https: / / www.uniprot.org / .
[0019] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. The following embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. For the experimental methods without specific conditions noted in the embodiments, they are usually carried out under conventional conditions or according to the conditions recommended by the manufacturers.
[0020] 1. Sample collection and storage Collect serum samples from 25 middle-aged and elderly patients with cognitive impairment (cognitive impairment group) and 25 middle-aged and elderly healthy control subjects (healthy control group). When collecting samples, record in detail the basic information of the patients and the healthy control group, as well as relevant laboratory test indicators, including age, gender, weight, BMI, triglyceride, total cholesterol, hypertension, low-density lipoprotein cholesterol, high-density lipoprotein cholesterol, and fasting blood glucose, etc.
[0021] Inclusion criteria for the case group: 1. Age: The age of the participants is over 55 years old; 2. There is a situation of cognitive decline. According to the score of the Mini-Mental State Examination (MMSE) used to evaluate cognitive function and dementia screening (Li et al., 2016): illiterate < 17, MMSE < 20 for those with 1 - 6 years of education; MMSE < 24 for individuals with 7 years or more of education; 3. Agree to participate: The participant or their legal guardian must agree to participate in this study and sign an informed consent form; 4. Have not used antidepressants or antipsychotics for at least 3 months; 5. Language comprehension ability: The participant has basic language comprehension and expression abilities and can conduct tests and communicate; 6. Living environment: The participant is a permanent resident in the cadmium-polluted area - Gangkou Town, Jiujiang City, Jiangxi Province.
[0022] Inclusion criteria for the healthy group: 1. Age: The age of the participants is over 55 years old; 2. The participants have no symptoms of cognitive function decline. According to the MMSE cognitive scale, the score of illiterate ≥ 17; MMSE ≥ 20 for those with 1 - 6 years of education; MMSE ≥ 24 for individuals with 7 years or more of education; 3. Agree to participate: The participant or their legal guardian must agree to participate in this study and sign an informed consent form; 4. Language comprehension ability: The participant has basic language comprehension and expression abilities and can conduct tests and communicate; 5. Living environment: The participant is also a permanent resident in the cadmium-polluted area - Gangkou Town, Jiujiang City, Jiangxi Province.
[0023] Exclusion criteria: 1. Age: People under 55 years old will be excluded; 2. Health status: Participants with severe heart disease, lung disease, kidney disease, or other severe physical diseases, or those with mental diseases such as schizophrenia, severe depression, etc. will be excluded; 3. Without obtaining the consent of the participant or their legal guardian, the participant will be excluded; 4. Language comprehension ability: Participants with severely impaired language comprehension and expression abilities and unable to communicate effectively will be excluded.
[0024] 2. Experimental methods This patent uses the next-generation label-free quantitative proteomics technology to complete the analysis. In the data-independent acquisition (DIA) mode, the latest high-resolution mass spectrometry is used to simultaneously collect the peptide ion characteristics in terms of mass number and retention time. The detection and analysis process steps are as follows: 1) Sample preparation Enrich the sample using superparamagnetic iron oxide nanoparticles. Take 20 μL of serum sample, dilute it with the loading buffer (10 mM Tris-HCl, 1 mM EDTA, 150 mM KCl, 0.05% CHAPS), mix it evenly with 1 mg of nanoparticle suspension, and incubate at 37 °C for 1 hour. Wash the nanoparticles twice with the loading buffer and then once with the loading buffer without CHAPS (10 mM Tris-HCl, 1 mM EDTA, 150 mM KCl). Adsorb the nanoparticles on a magnetic stand, discard the supernatant, and obtain the nanoparticles enriched with proteins. Add the lysis buffer (1% SDC / 100 mM Tris-HCl, pH = 8.5 / 10 mM TCEP / 40 mM CAA) to the sample, and incubate at 60 °C for 30 min for reduction alkylation reaction. Add an equal volume of ddH2O to dilute the SDC concentration to less than 0.5%, add 1 μg of trypsin, and incubate and shake overnight at 37 °C for digestion. The next day, add TFA to terminate the digestion, take the supernatant and desalt it using an SDB-RPS desalting column, vacuum dry it, and store it at -20 °C for later use.
[0025] 2) Mass spectrometry detection The mass spectrometry detection of the sample was performed using an UltiMate 3000 RSLCnano nanoliter liquid phase (Thermo) tandem timsTOFPro mass spectrometer (Bruker). The peptide sample was injected through an autosampler, bound to a C18 trapping column (75 µm * 2 cm, 3 µm particle size, 100 Å pore size, Thermo), and then separated in an analytical column (75 µm * 25 cm, 1.6 µm particle size, 100 Å pore size, IonOpticks). An analytical gradient was established using mobile phase A (0.1% formic acid) and mobile phase B (0.1% formic acid in ACN). The flow rate for the analysis was set at 300 nL / min. Mass spectrometry data was collected in the diaPASEF mode. The capillary voltage was set at 1400 V. The scanning ranges for MS1 and MS2 spectra were set at 100 - 1700 m / z. The ion mobility range was set at 0.6 - 1.6 Vs / cm2. The accumulation time and ramp time were set at 100 ms. According to the distribution law of mass-to-charge ratio - ion mobility, the diaPASEF acquisition window was set using the timsControl software. The collision energy was set to linearly increase from 59 eV at 1 / K0 = 1.6 Vs / cm2 to 20 eV at 1 / K0 = 0.6 Vs / cm2 according to the ion mobility.
[0026] 3) Data analysis 1 Database search The DIA raw data was analyzed using the library-free mode of the DIA-NN (1.8.1) software. The database used for the search was the HUMAN protein sequence database downloaded from Uniprot. The search parameters mainly adopted the default settings, and the key parameters are described as follows: The Precursor ion generation option was enabled for the prediction of the theoretical spectral library; Trypsin / P was used, allowing a maximum of 1 missed cleavage site; Carbamidomethyl (C) modification was set as a fixed modification; Oxidaton (M) and Acetylation (protein N-terminal) were set as variable modifications; The MS1 and MS2 mass tolerances were set at 15 ppm; MBR was enabled; Heuristic protein inference was enabled; The FDR was set at 1%. The protein quantification information was normalized using the MaxLFQ algorithm.
[0027] 2 Data cleaning, filtering, transformation, and filling Remove contaminating proteins from the quantitative protein matrix, filter proteins with a sample missing ratio greater than 70%, and perform logarithmic transformation. Then, fill in the missing values with values representing a normal distribution near the detection limit of the mass spectrometer. To this end, we determined the mean and standard deviation of the actual intensity distribution, and then created a new distribution shifted down by 1.8 standard deviations and with a width of 0.25 standard deviations. These values were used to estimate the missing values in the protein matrix. Finally, proteins with a missing ratio exceeding 50% in both groups were removed, and the remaining proteins were used for subsequent analysis.
[0028] 3 Differential expression analysis Compare the protein quantification data of the cognitive impairment group vs the healthy control group in the middle-aged and elderly population, and screen for differentially expressed proteins in the comparison between the two groups using the fold change (FC) and differential test (two-sided P<0.05) as criteria. The screening criteria were |log2FC|≥0.26, p<0.05. The results are as Figure 3 shown.
[0029] 4 Key protein screening Randomly divide the total sample into a training set and a test set at a ratio of 4:1. Use the Lasso regression model for five-fold cross-validation in the training set to screen out N differential protein features; then, perform three-fold cross-validation screening on these N proteins, and select the optimal diagnostic model through the discrimination index AUC and calibration index Brier score calculated by three-fold cross-validation to select the optimal combination of protein markers for constructing a logistic regression model. The differentially expressed proteins in the optimal combination are the disease-related protein markers.
[0030] 5 Model construction and evaluation Apply Logistic regression and cross-validation methods in the training set to randomly select protein combinations from the protein markers to construct multiple prediction models, and select the best according to the area under the receiver operating characteristic curve (AUC) and Brier score. Subsequently, evaluate the selected optimal prediction model. Evaluate the discrimination, calibration, and clinical utility performance of the model in the training set and test set through the receiver operating characteristic (ROC) curve, calibration curve, and decision curve in the test set and training set respectively. Generally, it is considered that 0.90≤AUC<1.00, the model has excellent discrimination ability; 0.75≤AUC<0.90, the model has good discrimination ability; 0.60≤AUC<0.75, the model has certain discrimination ability, but it is not recommended to use; AUC<0.60, the model has poor discrimination ability. All statistical analyses were completed using R (version 4.3.2) and Python (version 3.10.12).
[0031] 3. Results and analysis 1) From the middle-aged and elderly population in the questionnaire, 25 pairs of information-matched middle-aged and elderly patients with cognitive impairment and healthy control groups of middle-aged and elderly people were selected for proteomics research, and an early diagnosis model was established. Figure 1 The flow chart of the construction of the early diagnosis model for cognitive impairment in the middle-aged and elderly population was described.
[0032] 2) Analyze the clinical data of middle-aged and elderly people with cognitive impairment related to heavy metals in this study (description of the study population, univariate analysis, interaction test of grouping, multi-model comparison, and sensitivity analysis) (1) The demographic and clinical characteristics of the middle-aged and elderly people participating in this study were summarized, and the results are shown in Table 1.
[0033] Table 1 Demographic and clinical characteristics of the middle-aged and elderly population
[0034]
[0035] Results showed: The baseline characteristics of the 50 included subjects were presented. There were no differences in age (β = 0.16, 95%CI: -0.40, 0.71, P = 0.581) and gender composition (β = 0, 95%CI: -0.55, 0.55, P = 1.000) between the cognitively impaired group and the non-cognitively impaired group. The HAMA scores for anxiety in the cognitively impaired group were significantly higher than those in the non-cognitively impaired group (β = 1.44, 95%CI: 0.82, 2.06, P < 0.001), and the HAMD scores for depression in the cognitively impaired group were also significantly higher than those in the non-cognitively impaired group (β = 1.68, 95%CI: 1.03, 2.32, P < 0.001). Similarly, 80% of the individuals in the cognitively impaired group had anxiety (β = 1.50, 95%CI: 0.87, 2.13, P < 0.001), and the proportion of depression also reached 88% (β = 2.08, 95%CI: 1.39, 2.77, P < 0.001). Regarding heavy metal exposure, the mercury (β = 0.90, 95%CI: 0.32, 1.48, P = 0.003) and cadmium (β = 0.72, 95%CI: 0.15, 1.29, P = 0.014) levels in the plasma of the cognitively impaired group were significantly higher than those in the non-cognitively impaired group. There was a statistically significant difference in the detection of the ApoE3 gene between the two groups (β = 0.77, 95%CI: 0.20, 1.35, P = 0.011). Interestingly, the HDL-C in the cognitively impaired group was significantly higher than that in the non-cognitively impaired group (β = 0.77, 95%CI: 0.19, 1.34, P = 0.009). The average MMSE score of the cognitively impaired group was 21.00 ± 2.74, and the average MMSE score of the non-cognitively impaired group was 26.72 ± 1.97.
[0036] (2) Univariate analysis was performed on the demographic and clinical characteristics of the middle-aged and elderly population participating in this study, and the results are shown in Table 2.
[0037] Table 2 Univariate analysis of demographic and clinical characteristics of cognitive impairment in the middle-aged and elderly
[0038] Data in the table: β (95%CI) P value / OR (95%CI) P value; Outcome variable: MMSE / Cognitive impairment; Adjusted variable: None Taking the continuous variable MMSE as the dependent variable, regression analysis was performed with other relevant variables, and the data in the table show the effect values (β), 95% confidence intervals, and P values of the univariate analysis.
[0039] VFT, Lead is associated with a higher MMSE score. For every 1 μg / dL increase in blood lead level, the MMSE score increases by 0.43 (P = 0.0227). Among continuous variables, the effect sizes (β) of HAMA, HAMD, IADL, TMT-A, TMT-B, Stroop color word, and VST are significant, but the absolute values of the effect sizes (|β|) are small. Hg and Cd are negatively correlated with a higher MMSE score. For every 1 μg / L increase in blood mercury level, the MMSE score decreases by 0.67 (P = 0.0178); for every 1 μg / L increase in blood cadmium level, the MMSE score decreases by 0.51 (P = 0.0021). The correlations between the remaining continuous variables and the MMSE score are not statistically significant.
[0040] According to the heavy metal exposure standards recommended by the World Health Organization, the relationships between the exceeding of three heavy metals and the MMSE score are as follows: The MMSE score of the population with excessive blood mercury is 2.81 lower than that of the non-excessive population (P = 0.0394); the MMSE score of the population with excessive blood cadmium is 3.16 lower than that of the non-excessive population (P = 0.0021); the MMSE score of the population with excessive blood lead is 2.08 higher than that of the non-excessive population (P = 0.0488).
[0041] The MMSE score of anxious patients is 4.20 lower than that of non-anxious patients (P < 0.0001), and the MMSE score of depressive patients is 5.32 lower than that of non-depressive patients (P < 0.0001). The MMSE score of the population carrying the ApoE3 gene is 2.24 higher than that of non-carriers (P = 0.0331).
[0042] Taking the binary variable Cognitive impairment as the dependent variable and conducting a regression analysis with other relevant variables, the OR, 95% confidence interval, and P value of the univariate analysis are shown in the table.
[0043] An increase in VFT and Lead will reduce the risk of cognitive impairment. For every 1 μg / dL increase in blood lead level, the risk of cognitive impairment is reduced by 21% (P = 0.0383). Among continuous variables, an increase in TMT-A, TMT-B, Stroop color, Stroop color word, and VST is associated with an increase in the risk of cognitive impairment, but the magnitude is not large. An increase in Hg and Cd will increase the risk of cognitive impairment. For every 1 μg / L increase in blood mercury level, the risk of cognitive impairment increases by 67% (P = 0.0178), and for every 1 μg / L increase in blood cadmium level, the risk of cognitive impairment increases by 27% (P = 0.0187). It is worth noting that for every one-unit increase in HDL-C, the hazard ratio of cognitive impairment is 64.0 (P = 0.0059). The correlations between the remaining continuous variables and cognitive impairment are not statistically significant.
[0044] Exceedance of both Hg and Cd increases the risk of cognitive impairment. Compared with the non-exceedance population, the hazard ratio of cognitive impairment in the population with blood mercury exceedance is 11.29 (P = 0.0285), and the hazard ratio of cognitive impairment in the population with blood cadmium exceedance is 3.86 (P = 0.0255). Compared with the non-exceedance population, the risk of cognitive impairment in the population with blood lead exceedance is reduced by 78% (P = 0.0127).
[0045] The risk of cognitive impairment in anxious patients is 16.0 times that of non-anxious people (P < 0.0001), and the risk of cognitive impairment in depressive patients is 38.5 times that of non-depressive people (P < 0.0001). The risk of cognitive impairment in the population carrying the ApoE3 gene is reduced by 78% compared with non-carriers (P = 0.0127).
[0046] (3) Multiple model validations were performed on the cognitive impairment of the middle-aged and elderly population with heavy metals participating in this study, and the results are shown in Table 3.
[0047] Table 3 Relationship between Hg, Pb, Cd and MMSE score in multiple adjusted models
[0048] Non-adjusted model: No adjustment variables were included; Model 1: Adjusted for age, gender, and years of education; Model 2: Adjusted for age, gender, years of education, BMI, and ApoE3.
[0049] To better understand the relationship between heavy metals (Hg, Pb, and Cd) and the risk of cognitive impairment, we conducted a detailed multivariate regression analysis with variable adjustment in multiple models. The regression analysis after adjusting for age, gender, and years of education is shown in Model 1. The results showed that for every 1 μg / L increase in Hg, the MMSE score decreased by 0.71 points (β = -0.71, 95% CI: -1.27, -0.16, P = 0.0156). For every 1 μg / L increase in Cd, the MMSE score decreased by 0.54 points (β = -0.54, 95% CI: -0.86, -0.21, P = 0.0024). Similarly, when divided by whether it exceeded the standard, compared with the group with non-exceeding Hg, the MMSE score of the group with exceeding Hg decreased by 3.52 points (β = -3.52, 95% CI: -6.21, -0.83, P = 0.0136); compared with the group with non-exceeding Cd, the MMSE score of the group with exceeding Cd decreased by 3.25 points (β = -3.25, 95% CI: -5.33, -1.16, P = 0.0038). Secondly, through logistic regression analysis, it was found that for every 1 μg / L increase in Hg, the risk of developing cognitive impairment increased to 1.76 times the original (OR = 1.76, 95% CI: 1.19, 2.62, P = 0.0051); for every 1 μg / L increase in Cd, the risk of developing cognitive impairment increased to 1.29 times the original (OR = 1.29, 95% CI: 1.04, 1.59, P = 0.0217). When divided by whether it exceeded the standard, the risk of developing cognitive impairment in the group with exceeding Hg was 18.04 times that of the group with non-exceeding Hg (OR = 18.04, 95% CI: 1.71, 189.77, P = 0.0160); the risk of developing cognitive impairment in the group with exceeding Cd was 4.13 times that of the group with non-exceeding Cd (OR = 4.13, 95% CI: 1.14, 15.00, P = 0.0309). Similarly, the results of Model 2 after adjusting for age, gender, years of education, BMI, and ApoE3 variables showed that for every 1 μg / L increase in Hg, the MMSE score decreased by 0.66 points (β = -0.66, 95% CI: -1.21, -0.11, P = 0.0223), and the risk of developing cognitive impairment increased to 1.85 times the original (OR = 1.85, 95% CI: 1.20, 2.86, P = 0.0057). For every 1 μg / L increase in Cd, the MMSE score decreased by 0.49 points (β = -0.49, 95% CI: -0.84, -0.13, P = 0.0108).
[0050] In summary, Hg and Cd are independent risk factors for cognitive impairment. However, it is worth noting that in Model 1, the role of Pb appears to be a protective factor. The risk of cognitive impairment in people with elevated Pb levels is 0.18 times that of those without elevated Pb levels (OR = 0.18, 95% CI: 0.05, 0.66, P = 0.0093). However, after adjusting for additional variables such as BMI and ApoE3 in Model 1, the results of Model 2 showed that the effect of Pb on cognitive impairment was not significant, suggesting that BMI and APOE3 may be confounding factors affecting the relationship between Pb and cognitive impairment.
[0051] (4) Subgroup analysis was performed to evaluate the consistency of the association between the three heavy metals and cognitive impairment in different populations. The results are shown in Tables 4 - 6.
[0052] Table 4 Subgroup analysis of Hg (μg / L) vs MMSE score and cognitive impairment
[0053] Table 5 Subgroup analysis of Pb (μg / dL) vs MMSE score and cognitive impairment
[0054] Table 6 Subgroup analysis of Cd (μg / L) vs MMSE score and cognitive impairment
[0055] The results showed that hyperlipidemia (P = 0.0483) and depression (P = 0.0217) were interaction factors affecting the association between Hg and cognitive ability; Hg could interact with blood cadmium levels and affect the MMSE score. The only variable with a significant interaction test result when Pb was used as the independent variable was age. In the analysis with MMSE as the dependent variable, there were no interactions between Hg, Pb, Cd and other stratification variables, nor between them pairwise. In other subgroups, the associations between the three heavy metals, lead, cadmium, and mercury, and cognitive impairment were stable, and gender, education, BMI, smoking, hypertension, diabetes, and high - density lipoprotein cholesterol had no significant effect on this association.
[0056] (5)Sensitivity analysis of the model in different subgroups (demographic and clinical characteristics) Table 7 Sensitivity analysis of clinical modeling in different subgroups
[0057] Shows the AUC and 95% CI of the protein prediction model and the clinical prediction model in different subgroups.
[0058] (3) Mass spectrometry quality control information (1) Using data-independent acquisition (DIA) technology, serum proteomic profiles of 25 middle-aged and elderly patients with cognitive impairment and 25 healthy controls in the middle-aged and elderly population were obtained. A total of 2,044 proteins were quantified in the entire cohort using DIA technology. (2) The expression of 73 proteins was significantly different between middle-aged and elderly patients with cognitive impairment and healthy controls in the middle-aged and elderly population. Compared with the healthy control group, the expression of 26 proteins was up-regulated and 47 proteins were down-regulated in the anxiety group. The results are as Figure 2 shown. (3) The results of principal component analysis (PCA) showed the intensity of 73 differentially expressed proteins, showing a significant difference between middle-aged and elderly patients with cognitive impairment and healthy controls in the middle-aged and elderly population. The results are as Figure 3 shown.
[0059] (4) Protein screening The total samples were randomly divided into a training set and a test set at a ratio of 4:1. The Lasso regression model was used for five-fold cross-validation in the training set to screen out 18 important proteins from 73 differentially expressed proteins. Logistic regression analysis was performed on these 18 important proteins by random combination to construct a diagnostic model, and the optimal diagnostic model was screened by the discrimination index AUC and the calibration index Brier score calculated by three-fold cross-validation. The screening results are as Figure 4 and Figure 5 shown. Finally, TSKU, SOD3, CES2, PSMB10, and C1QA were selected as the best protein biomarker combination, and a clinical risk model was established. The model results are shown in Table 8.
[0060] Table 8 Logistic regression model Unnamed: 0 Estimate Std. Error z value Pr(>|z|) (Intercept) -65.313 32.083 -2.036 0.042 TSKU -0.774 0.542 -1.428 0.153 SOD3 2.547 1.351 1.885 0.059 CES2 -2.728 1.389 -1.965 0.049 PSMB10 2.927 1.337 2.189 0.029 C1QA 1.956 0.747 2.619 0.009 "Estimate" is the regression coefficient, "Std.Error" is the standard error of the regression coefficient, "z value" is the test statistic for the Z test of the regression coefficient, and "Pr(>|z)" is the P value for the Z test of the regression coefficient.
[0061] Box plots were made of the expression levels of the five protein biomarkers TSKU, SOD3, CES2, PSMB10, and C1QA screened in the two groups of people to visually transform the expression differences of these five protein biomarkers in the sera of middle-aged and elderly patients with cognitive impairment and healthy controls in the middle-aged and elderly population. The box plots showed that compared with the healthy control group, the expression levels of TSKU and CES2 were down-regulated in middle-aged and elderly patients with heavy metal-related cognitive impairment, while the expression levels of SOD3, PSMB10, and C1QA were up-regulated in middle-aged and elderly patients with heavy metal-related cognitive impairment. The results are as Figures 6 - 10 shown.
[0062] Protein model evaluation When evaluating the early diagnosis model composed of protein markers TSKU, SOD3, CES2, PSMB10, and C1QA, 50 samples were divided into two sets according to a certain proportion. Among them, 80% of the data was used as the training set (n = 40), and 20% of the data was used as the test set (n = 10). The training set and the test set were used for model evaluation.
[0063] 1) Receiver operating characteristic (ROC) curve: The results are as Figure 11 shown. The ROC curve shows that the early diagnosis model for cognitive impairment in the middle-aged and elderly population obtained this time has good discrimination ability.
[0064] 2) Confusion matrix: The results are as Figure 12 shown. The results of the confusion matrix diagram show that the model sensitivity is 84% (21 / 25), and the specificity is 84% (21 / 25), indicating that the prediction ability of this early diagnosis model is relatively good.
[0065] 3) Calibration curve: The results are as Figure 13 shown. The results of the calibration curve show that the consistency between the predicted probability and the observed probability of the early diagnosis model obtained this time is relatively good.
[0066] 4) Decision curve: The results are as Figure 14 shown. The decision curve diagram shows that within a certain range, the net return rate of the early diagnosis model is relatively high and higher than the two options of all treatments and no treatments at all, proving that the clinical utility of the early diagnosis model obtained this time is relatively large.
[0067] Protein model visualization According to Figure 15The presented results visualized the model results through a nomogram, which helped predict the disease risk. The sum of the individual scores (Points) of each protein marker at different values was used to obtain the total score (TotalPoints), which was then mapped to the Risk score, representing the probability of cognitive impairment events in the middle-aged and elderly population. Specifically, in the nomogram, the score range of TSKU was 20 - 15.5, corresponding to 0 - 39 on the score scale; the score range of SOD3 was 16.2 - 18.4, corresponding to 0 - 63 on the score scale; the score range of CES2 was 11.8 - 8.8, corresponding to 0 - 93 on the score scale; the score range of PSMB10 was 11.2 - 13.2, corresponding to 0 - 67 on the score scale; the score range of C1QA was 13.5 - 18, corresponding to 0 - 100 on the score scale; the probability values of cognitive impairment events in the middle-aged and elderly population related to heavy metals were 0.01, 0.1, 0.3, 0.5, 0.7, and 0.9, corresponding to 89 - 91, 116 - 118, 132 - 135, 142 - 144, 150 - 152, and 166 - 168 on the total score scale, respectively. It can be seen from the visualization results that the combination of these five markers is of great significance for risk prediction, and the upper limit of the predicted probability can reach 0.9.
[0068] Clinical Protein Model Evaluation When evaluating the early diagnosis model composed of baseline characteristics such as Cd, TMT - B, Stroop - color, and protein markers CES2 and C1QA, 50 samples were divided into two sets according to a certain proportion. Among them, 80% of the data was used as the training set (n = 40), and 20% of the data was used as the test set (n = 10). The training set and the test set were used for model evaluation.
[0069] 1) Receiver Operating Characteristic (ROC) Curve: The results are as Figure 16 shown. The ROC curve indicates that the early diagnosis model for cognitive impairment in the middle-aged and elderly population obtained this time has good discrimination ability.
[0070] 2) Confusion Matrix: The results are as Figure 17 shown. The results of the confusion matrix diagram show that the sensitivity of the model is 76% (19 / 25), and the specificity is 84% (21 / 25), indicating that the early diagnosis model has good predictive ability.
[0071] 3) Calibration Curve: The results are as Figure 18 shown. The results of the calibration curve show that the consistency between the predicted probability and the observed probability of the early diagnosis model obtained this time is good.
[0072] 4) Decision Curve: The results are as Figure 19As shown, the decision curve shows that within a certain range, the net benefit rate of the early diagnosis model is relatively high and higher than the two options of all treatments and no treatments at all, proving that the clinical utility of the early diagnosis model obtained this time is relatively large.
[0073] 7. Visualization of the clinical protein model According to Figure 20 the results presented, in the nomogram, the score range of Cd is 0 - 10, corresponding to 0 - 32.5 in the corresponding score scale; the score range of TMT-B is 20 - 260, corresponding to 0 - 19 in the score scale; the score range of Stroop-color is 10 - 100, and the corresponding score scale is 0 - 49; the score range of CES2 is 11.8 - 8.8, and the corresponding score range is 0 - 52; the score range of C1QA is 13.5 - 18, and the corresponding score range is 0 - 100; the probabilities of cognitive impairment events occurring in the elderly population related to heavy metals take values of 0.01, 0.1, 0.3, 0.5, 0.7, and 0.9, corresponding to 52 - 54, 66 - 68, 75 - 77, 81 - 83, 86 - 88, and 94 - 96 in the total score scale respectively. It can be seen from the visualization results that the combination of these five biomarkers is of great significance for risk prediction, and the upper limit of the prediction probability can reach 0.9.
[0074] 8. Proteomic analysis of the brain tissues of cadmium-exposed rats According to Figure 21 the results presented, we established a cadmium exposure model in 17 rats to explore the expression of the biomarkers found in the above clinical cohort at the tissue level. The rats were randomly divided into a control group, a low cadmium exposure group, and a high cadmium exposure group. After eight weeks of intervention, the prefrontal cortex tissues of the rats were taken for DIA proteomic detection and differential protein screening, and one-way ANOVA was used for comparison between groups ( Figure 21 A).
[0075] The quantitative heat map of the differential proteins shows that there are significant differences in the expression of DEPs among the three groups of rats ( Figure 21 B). Similarly, PCA analysis shows that the differential proteins can clearly distinguish the control group, the low cadmium exposure group, and the high cadmium exposure group ( Figure 21 D). The GO enrichment analysis of the differential proteins shows that the most enriched BP, CC, and MF pathways are cellular process, extracellular space, and identical protein binding respectively ( Figure 21 C). Salmonella infection, Protein processing in endoplasmic reticulum, and Carbon metabolism are the top three most enriched pathways of the differential proteins in the KEGG database.Figure 21 E).
[0076] To determine the differentially expressed proteins in the population cohort and whether the proteins screened for constructing the prediction model show differences in the brain tissues of the rat experiment, the intersection of the differentially expressed proteins in the prefrontal cortex tissues of cadmium-exposed rats and the differentially expressed proteins in the population cohort was taken. The results showed that there were 9 intersection proteins in total, namely PGM1, WDR81, IDE, FDPS, COPS7A, CCDC126, SERPINC1, PFDN5, and NID2. Among them, PFDN5 is an important protein in the XGBoost machine learning model of the population cohort and is also the hub protein at the intersection of the four algorithms of the protein-protein interaction network described above. NID2 was included in the screening results of the LASSO analysis ( Figure 21 F). The results of the rat experiment verified the reliability of the results of the population cohort study.
[0077] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A protein marker for diagnosing heavy metal-related cognitive impairment in the elderly, characterized in that: The protein marker is a combination of the following proteins: TSKU, SOD3, CES2, PSMB10 and C1QA.
2. The protein marker for diagnosing heavy metal-related cognitive impairment in the elderly according to claim 1, characterized in that: The expression levels of TSKU and CES2 are downregulated in patients with heavy metal-related cognitive impairment in the elderly and middle-aged population, and the expression levels of SOD3, PSMB10 and C1QA are upregulated in patients with heavy metal-related cognitive impairment in the elderly and middle-aged population.
3. The protein marker for diagnosing heavy metal-related cognitive impairment in the elderly according to claim 1, characterized in that: The protein markers are extracted from tissue or blood.
4. Use of the protein marker for diagnosing heavy metal-related cognitive impairment in the elderly as claimed in claim 1 in a kit for early diagnosis of symptoms of heavy metal-related cognitive impairment in the elderly.
5. An early diagnosis kit for detecting heavy metal-related cognitive impairment in the elderly, characterized in that: The detection reagent in the kit has the ability to detect the expression levels of TSKU, SOD3, CES2, PSMB10 and C1QA.
6. The early diagnosis kit according to claim 5, characterized in that: The detection reagent can specifically bind to proteins encoding TSKU, SOD3, CES2, PSMB10 and C1QA.
7. The early diagnosis kit according to claim 6, characterized in that: The detection reagents comprise peptides, peptide mimetics, aptamers, spiegelmers, darpins, ankyrin repeat proteins, Kunitz-type domains, antibodies, single domain antibodies and / or monovalent antibody fragments.
8. A risk assessment model for diagnosing heavy metal-related cognitive impairment in the elderly, characterized in that: The expression level of the protein marker described in claim 1 is used as a variable, and the model results are visualized by using a nomogram. The individual scores corresponding to each protein marker at different values are summed up to obtain a total score, and then the total score is corresponded to the probability of heavy metal-related cognitive impairment events in the elderly population.
9. The risk assessment model for diagnosing heavy metal-related cognitive impairment in the elderly according to claim 8, characterized in that: In the nomogram, the score range for each gene and the corresponding score scale. For example, the score range of TSKU is 20-15.5, corresponding to 0-39 in the score scale; the score range of SOD3 is 16.2-18.4, corresponding to 0-63 in the score scale; the score range of CES2 is 11.8-8.8, corresponding to 0-93 in the score scale; the score range of PSMB10 is 11.2-13.2, corresponding to 0-67 in the score scale; the score range of C1QA is 13.5-18, corresponding to 0-100 in the score scale; the probability of heavy metal-related cognitive impairment events in the middle-aged and elderly population is 0.01, 0.1, 0.3, 0.5, 0.7 and 0.9, corresponding to 89-91, 116-118, 132-135, 142-144, 150-152 and 166-168 in the total score scale, respectively.
10. The risk assessment model for diagnosing heavy metal-related cognitive impairment in the elderly according to claim 8, characterized in that: Clinical modeling was performed after incorporating demographic characteristics and clinical factors. After optimizing the model, a clinical protein model constructed by cadmium, TMT-B, Stroop-color, CES2, and C1QA was obtained. In the nomogram, the score range of each gene and the corresponding score scale were shown. For example, the score range of Cd is 0-10, corresponding to 0-32.5 in the score scale; the score range of TMT-B is 20-260, corresponding to 0-19 in the score scale; the score range of Stroop-color is 10-100, corresponding to 0-49 in the score scale; the score range of CES2 is 11.8-8.8, corresponding to 0-52; the score range of C1QA is 13.5-18, corresponding to 0-100; the probability of cognitive impairment events in the middle-aged and elderly population related to heavy metals is 0.01, 0.1, 0.3, 0.5, 0.7 and 0.9, corresponding to 52-54, 66-68, 75-77, 81-83, 86-88 and 94-96 in the total score scale respectively.