Alzheimer's disease onset risk prediction method and device based on multiple factors
By obtaining the immunoglobulin IgG level values and personal information that specifically react with Toxoplasma gondii CST1 antigen in serum samples, and combining the XGboost algorithm to predict Alzheimer's disease risk, the problem of integrating infection factors in the existing technology is solved, and early personalized evaluation and low-cost detection are achieved.
Patent Information
- Application Number
- CN202510623817.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-26
AI Technical Summary
The prior art is difficult to integrate infection factors with traditional risk factors for Alzheimer's risk prediction, and biomarker detection is highly invasive, costly and difficult to predict early.
By obtaining the immunoglobulin IgG level values and personal basic information that specifically respond to Toxoplasma gondii CST1 antigen in the serum samples of the tester, the training-completed risk prediction model was used to predict using the XGboost algorithm, and personalized risk assessment was conducted based on gender and age.
Early personalized Alzheimer's risk assessment was achieved, improving the accuracy and feasibility of predictions, and reducing the invasiveness and cost of detection.
Smart Images

Figure CN120544883A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of disease-assisted diagnosis, and more specifically, to a method and device for predicting the risk of Alzheimer's disease based on multiple factors. Background Art
[0002] Alzheimer's disease (AD) is a complex neurodegenerative disease whose risk is affected by multiple factors, including age, gender, genetic factors, environment and lifestyle, as well as infections with bacteria, fungi, viruses and parasites (such as herpes virus, Porphyromonas gingivalis, Helicobacter pylori, Toxoplasma gondii, etc.) proposed by emerging research.
[0003] Existing methods are primarily based on clinical assessment, biomarker testing, and imaging techniques. Biomarker testing often relies on cerebrospinal fluid markers (such as Aβ and Tau), which have high specificity and sensitivity and are an important basis for diagnosing AD. However, these methods are highly invasive, costly, and difficult to predict early. Recent studies have suggested that bacterial, fungal, viral, and parasitic infections may contribute to the disease process through inflammation or other mechanisms, but there is currently no method that integrates infection with traditional risk factors for Alzheimer's disease risk prediction. Summary of the Invention
[0004] In view of this, the present invention provides a multi-factor based method and device for predicting the risk of Alzheimer's disease. By using specific infection indicators and combining the patient's gender, age and other personal information, a trained prediction model is used to predict the risk prediction probability of the test person, which can perform personalized risk assessment and early intervention.
[0005] To achieve the above objectives, the following solutions are proposed:
[0006] A multi-factor-based method for predicting the risk of Alzheimer's disease, including:
[0007] Obtaining the level of immunoglobulin IgG that specifically reacts with Toxoplasma gondii CST1 antigen in the serum sample of the test person;
[0008] Obtain the tester's basic personal information, including gender and age;
[0009] The test person's serum anti-CST1 IgG level and personal basic information are input into a pre-trained risk prediction model for prediction to obtain the test person's Alzheimer's disease risk.
[0010] Preferably, the risk prediction model is trained using the XGboost algorithm, with the training user's disease label, serum anti-CST1 IgG level and personal basic information as training samples and the training user's disease risk probability as sample label.
[0011] Preferably, the process of predicting the risk of developing Alzheimer's disease in a test person includes:
[0012] The risk prediction model obtains the risk prediction probability of the test person based on the test person's serum anti-CST1 IgG level and personal basic information;
[0013] Based on the risk prediction probability, the test person's Alzheimer's disease risk is determined.
[0014] Preferably, if batch prediction is performed, a data frame is created based on the tester information and standardized;
[0015] Input each data frame into the risk prediction model in batches to obtain the Alzheimer's disease risk of each test person.
[0016] Preferably, the IgG level of immunoglobulin specifically reactive with the CST1 antigen in the serum sample is detected by ELISA.
[0017] A device for predicting the risk of Alzheimer's disease based on multiple factors, comprising:
[0018] An IgG level acquisition unit, used for acquiring the level value of immunoglobulin IgG that specifically reacts with Toxoplasma gondii CST1 antigen in the serum sample of the test person;
[0019] An information collection unit is used to obtain the basic personal information of the tester, wherein the basic personal information includes gender and age;
[0020] The risk prediction unit is used to input the test person's serum anti-CST1 IgG level value and personal basic information into a pre-trained risk prediction model to predict the test person's Alzheimer's disease risk.
[0021] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0022] The multi-factorial Alzheimer's disease risk prediction method provided by the present invention first obtains the level of immunoglobulin IgG (antigen IgG) specifically reactive with the Toxoplasma gondii CST1 antigen in a serum sample of a test subject; then obtains the test subject's basic personal information, including gender and age; then inputs the test subject's serum anti-CST1 IgG level and basic personal information into a pre-trained risk prediction model to determine the test subject's Alzheimer's disease risk. This method uses a trained prediction model to predict the test subject's risk probability by combining specific infection indicators with the patient's gender, age, and other personal information, enabling personalized risk assessment and early intervention. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0024] Figure 1 A flowchart of a method for predicting the risk of Alzheimer's disease based on multiple factors provided by an embodiment of the present invention;
[0025] Figure 2 a-2b is a schematic diagram of a batch test result code provided by an embodiment of the present invention;
[0026] Figure 3 A graph showing test results for different items of the risk prediction model provided by an embodiment of the present invention;
[0027] Figure 4 A schematic diagram of the structure of a device for predicting the risk of Alzheimer's disease based on multiple factors provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0029] First, combine Figure 1 The present invention provides a method for predicting the risk of Alzheimer's disease based on multiple factors. Figure 1 As shown, the method includes:
[0030] Step S01: obtaining the level of immunoglobulin IgG that specifically reacts with Toxoplasma gondii CST1 antigen in the serum sample of the test person.
[0031] Specifically, the system can collect the level of immunoglobulin IgG that specifically reacts with the Toxoplasma gondii CST1 antigen in the test person's serum sample input by the staff, collect the patient's serum sample and test for indicators of specific pathogen infection.
[0032] The level of immunoglobulin IgG that specifically reacts with the Toxoplasma CST1 antigen in serum samples can be measured by ELISA. The process of ELISA testing is as follows:
[0033] 1) Add 100 μl of coating buffer containing 0.1 μg of purified recombinant CST1 protein to each well of a Highbinding 96-well microtiter plate. Place in a humidified chamber and coat overnight at 4°C. The amplified sequence corresponding to the recombinant CST1 protein is the 1122-bp N-terminal portion (328-1449) of the CST1 protein-encoding gene (XM_002368554.2). The gene sequence is shown in Sequence Listing 1.
[0034] 2) Wash the plate with PBST at the following volume ratios: 150 μl, 200 μl, 250 μl, 300 μl, and 350 μl, washing each well 6 times.
[0035] 3) Add 400 μl of 3% skim milk to each well and block the wells in a humidified chamber at room temperature for one hour.
[0036] 4) After aspirating or tapping off the blocking solution, add 100 μl of 1:400 diluted serum (Alzheimer's disease and normal human serum are diluted with PBS).
[0037] 5) Wash the plate with PBST at the following volumes: 150 μl, 200 μl, 250 μl, 300 μl, 350 μl, and 400 μl, washing each well 6 times.
[0038] 6) Dilute the HRP-labeled anti-human IgG full-length fragment with 3% skim milk at a dilution ratio of 1:1000, then add 100 μl to each well, place in a humidified container at room temperature, and incubate for 1 hour.
[0039] 7) Wash six times with PBST at 150 μl, 200 μl, 250 μl, 300 μl, 350 μl, and 400 μl per well.
[0040] 8) After discarding all liquid from the wells, add 200 μl of pre-prepared, light-proof OPD colorimetric solution to each well. Allow to develop in the dark for 30 minutes at room temperature. After development, read the absorbance at 450 nm using a microplate reader and record it for comparison.
[0041] 9) After color development is complete, retain the liquid and directly add 50 μl of 1M H2SO4 to each well to terminate the color development reaction. Use a microplate reader to measure the absorbance at a wavelength of 490 nm.
[0042] Step S02: Obtain the basic personal information of the tester, wherein the basic personal information includes gender and age.
[0043] Specifically, information such as the gender, age, and family genetic disease history of the input tester is obtained.
[0044] Step S03: Inputting the serum anti-CST1 IgG level value and personal basic information of the test person into a pre-trained risk prediction model for prediction.
[0045] Specifically, the serum anti-CST1 IgG level value and personal basic information of the above-mentioned test person are input into the risk prediction model to generate the risk prediction probability of Alzheimer's disease for the test person; the risk of Alzheimer's disease of the test person is judged based on the predicted risk prediction probability. Testers whose risk prediction probability exceeds the preset probability threshold can be identified as high risk. For example: the risk prediction probability threshold is set to X (for example, 0.6). When the score exceeds X, it indicates high risk. An 87-year-old woman with a CST1 ELISA test value (IgG level value) of 1.3 and a risk prediction probability of 0.699 is classified as high risk.
[0046] The multi-factorial Alzheimer's disease risk prediction method provided by the present invention first obtains the level of immunoglobulin IgG (antigen IgG) specifically reactive with the Toxoplasma gondii CST1 antigen in a serum sample of a test subject; then obtains the test subject's basic personal information, including gender and age; then inputs the test subject's serum anti-CST1 IgG level and basic personal information into a pre-trained risk prediction model to determine the test subject's Alzheimer's disease risk. This method uses specific infection indicators combined with the patient's gender, age, and other personal information to predict the test subject's risk probability using the trained prediction model, enabling personalized risk assessment and early intervention.
[0047] Furthermore, the embodiment of the present invention takes into account the need to perform risk prediction on multiple test persons in batches. A data frame can be created first based on the test person's information, and the data in the data frame can be standardized. The data frame contains the serum anti-CST1 IgG level value and personal basic information corresponding to the test person. The data frames are input into the risk prediction model in batches to obtain the Alzheimer's disease risk of each test person. If batch prediction is performed, the corresponding data frame can be created and then standardized for batch prediction. For example, if three individuals are predicted at the same time, No. 1 is a 55-year-old female with an ELISA test value of 2.1 for CST1, No. 2 is a 68-year-old male with an ELISA test value of 1.0 for CST1, and No. 3 is a 77-year-old female with an ELISA test value of 1.0 for CST1. Create a data frame such as Figure 2 As shown in a, the data frame is input into the risk prediction model. The risk prediction model predicts that the probability of risk No. 1 is 0.517, the probability of risk No. 2 is 0.639, and the probability of risk No. 3 is 0.699. Figure 2 b shows the corresponding disease risks for the three individuals.
[0048] Next, the present invention describes the training process of the risk prediction model. The model is trained using the XGboost algorithm, using the training user's disease label (whether they have Alzheimer's disease), serum anti-CST1 IgG level (ELISA test value for recombinant protein CST1), and basic personal information (gender, age) as training samples, and the training user's disease risk probability as the sample label.
[0049] 1. Data Collection
[0050] The following data were collected from 235 subjects (114 healthy controls and 121 patients with Alzheimer's disease):
[0051] Serum samples were collected and the level of immunoglobulin G (IgG) in the serum that specifically reacts with the Toxoplasma CST1 antigen was detected using ELISA.
[0052] Gender (male / female), age (range 55 years and above).
[0053] The results showed that the level of specific IgG antibodies against CST1 antigen in the serum of the patient group was significantly higher than that of the control group (P<0.01), and the risk was higher in elderly individuals.
[0054] Table 1 shows the ELISA test results of CST1 and crude antigen against different sera
[0055] Serum class CST1 Alzheimer's patients 0.559±0.044 Healthy people 0.147±0.008 Statistic U 1297 P-value <0.001
[0056] (OD490 ±SEM)
[0057] 2. Model construction and training
[0058] 1) XGboost is a gradient boosting tree-based algorithm that optimizes the objective function by iteratively adding decision trees. Its objective function is:
[0059]
[0060] Represents the loss function, measuring the true value y i and predicted values For binary classification tasks, the loss function is cross entropy loss, i.e. Log Loss. Ω(f k ) is a regularization term used to control model complexity. It can be mathematically expressed as:
[0061]
[0062] f k represents the prediction function of the kth tree, γ is the gamma parameter used to control the split, T is the number of leaf nodes, w is the leaf weight, and λ is the L2 regularization parameter.
[0063] 2) XGBoost predictions are obtained by accumulating the results of multiple trees. The final output is converted into a prediction probability through the sigmoid function, which is mathematically expressed as:
[0064]
[0065] 3) XGBoost optimizes the objective function using gradient descent, a gradient boosting method based on a second-order Taylor expansion. At each step, the first-order derivative (gradient) and second-order derivative (Hessian) of the loss function are calculated and used to fit a new tree. The learning rate (step size) eta controls the contribution of each new tree.
[0066] Therefore, in R language, the training process of XGBoost model is completed by controlling parameters, including task type, evaluation method, tree complexity, regularization, learning rate and randomness. The detailed parameters are:
[0067] (1) objective = "binary:logistic", which specifies the task type and objective function of the model. The task type is binary classification and the objective function is logistic regression.
[0068] (2) eval_metric = "auc", specifies the model evaluation metric during training.
[0069] (3) max_depth = 3, controls the depth of the tree.
[0070] (4) Learning rate (step size reduction factor) eta = 0.1, used to control the magnitude of the model update in each iteration. The smaller the value, the slower the model learns but the more stable it is, usually requiring more iterations. eta = 0.1 is a common default value.
[0071] (5) gamma = 0.1, which is the mathematical expression of the regularization term and determines whether to split the tree. If the loss reduction caused by the split is less than gamma, the split is not performed. 0.1 means that the requirement for splitting is slightly stricter, which helps to prevent overfitting.
[0072] (6) The sample subsampling ratio is set to subsample = 0.8. In each tree construction process, 80% of the training data is randomly sampled to train the model to prevent overfitting and increase the generalization ability of the model.
[0073] (7) The feature subsampling ratio is set to colsample_bytree = 0.8. Each time a tree is built, 80% of the features are randomly sampled to train the model. This is similar to subsample, but acts on the feature dimension to increase the diversity and robustness of the model.
[0074] 3. Model verification and analysis
[0075] The test results of different items of the risk prediction model are as follows Figure 3 First, according to the feature dependency of the model based on SHAP value Figure 3 The analysis shows that when gender changes from 0 (female) to 1 (male), the SHAP value decreases, indicating that females are a risk factor relative to males. However, the change in SHAP value is very small, indicating that gender has little impact on the prediction results.
[0076] Secondly, according to Figure 3 b analysis shows that as age increases, the effect of age on the predicted outcome increases, which is nonlinear and has different effects in different age groups. The color coding is CST1 (0 to 5), the purple points (high CST1) are concentrated in the area where the SHAP value corresponding to age is greater than 1, and the yellow points (low CST1) are widely distributed, suggesting that there may be an interactive effect between CST1 and age. Figure 3Analysis of CST1 c reveals that as CST1 levels increase, the SHAP value initially increases rapidly before leveling off, indicating that after reaching a certain threshold, CST1 reflects similar infection states and has a similar impact on prediction results. Furthermore, purple points (high age) are concentrated in areas with higher CST1 levels, while yellow points (low age) are distributed in areas with lower CST1 levels, indicating an interactive effect between age and CST1. Higher CST1 levels (especially when combined with higher age) increase the probability of developing the disease.
[0077] Finally, through Figure 3 d analysis revealed a weak interaction between sex and CST1, possibly due to gender-specific lifestyle differences, such as women being more likely to keep cats. A stronger interaction between CST1 and age could be attributed to a decline in immune health with aging. Figure 3 e shows that there is a strong interactive effect between age and CST1 level. The combination of high age and high CST1 level has the greatest positive contribution to the model prediction. The elderly group (75 years old) with high CST1 level is a high-risk group and should be given priority attention. In addition, Figure 3 As shown in Figure 5, on the test set, the overall prediction accuracy of the model is 85.7% and the AUC value is 0.92.
[0078] The following describes an Alzheimer's disease risk prediction device based on multiple factors provided in an embodiment of the present invention. The Alzheimer's disease risk prediction device based on multiple factors described below and the Alzheimer's disease risk prediction method based on multiple factors described above can be referenced to each other.
[0079] First combine Figure 4 , introduces the Alzheimer's disease risk prediction device based on multiple factors, such as Figure 4 As shown, the multi-factor based Alzheimer's disease risk prediction device may include:
[0080] IgG level acquisition unit 100, which acquires the level of immunoglobulin IgG that specifically reacts with Toxoplasma gondii CST1 antigen in the serum sample of the test person;
[0081] The information collection unit 200 is used to obtain the basic personal information of the tester, wherein the basic personal information includes gender and age;
[0082] The risk prediction unit 300 is used to input the serum anti-CST1 IgG level value and personal basic information of the test person into a pre-trained risk prediction model to predict the test person's Alzheimer's disease risk.
[0083] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0084] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0085] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for predicting the risk of Alzheimer's disease based on multiple factors, characterized in that: include: Obtaining the level of immunoglobulin IgG that specifically reacts with Toxoplasma gondii CST1 antigen in the serum sample of the test person; Obtain the tester's basic personal information, including gender and age; The test person's serum anti-CST1 IgG level and personal basic information are input into a pre-trained risk prediction model for prediction to obtain the test person's Alzheimer's disease risk.
2. The method for predicting the risk of Alzheimer's disease based on multiple factors according to claim 1, characterized in that: The risk prediction model is trained using the XGboost algorithm, with the training user's disease label, serum anti-CST1 IgG level and personal basic information as training samples and the training user's disease risk probability as sample label.
3. The method for predicting the risk of Alzheimer's disease based on multiple factors according to claim 1, characterized in that: The process of predicting a person's risk of developing Alzheimer's disease includes: The risk prediction model obtains the risk prediction probability of the test person based on the test person's serum anti-CST1 IgG level and personal basic information; Based on the risk prediction probability, the test person's Alzheimer's disease risk is determined.
4. The method for predicting the risk of Alzheimer's disease based on multiple factors according to claim 1, characterized in that: If batch prediction is performed, a data frame is created based on the tester information and standardized; Input each data frame into the risk prediction model in batches to obtain the Alzheimer's disease risk of each test person.
5. The method for predicting the risk of Alzheimer's disease based on multiple factors according to claim 1, characterized in that: The levels of immunoglobulin IgG that specifically reacts with CST1 antigen in serum samples were detected by ELISA.
6. A device for predicting the risk of Alzheimer's disease based on multiple factors, characterized in that: include: An IgG level acquisition unit, used for acquiring the level value of immunoglobulin IgG that specifically reacts with Toxoplasma gondii CST1 antigen in the serum sample of the test person; An information collection unit is used to obtain the basic personal information of the tester, wherein the basic personal information includes gender and age; The risk prediction unit is used to input the test person's serum anti-CST1 IgG level value and personal basic information into a pre-trained risk prediction model to predict the test person's Alzheimer's disease risk.