Multi-system atrophy risk prediction model

By collecting and processing data of patients with multi-system atrophy, analyzing and controlling bias, establishing Cox regression and machine learning models, and building risk prediction models, it solves the problem of difficult to deal with multi-system atrophy patient data in the existing technology, and achieves accurate prediction of disease progress and improvement of clinical management.

CN119943391APending Publication Date: 2025-05-06FOURTH MILITARY MEDICAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510023032.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

It is difficult to collect and process data from patients with multi-system atrophy, analyze the bias of control data, identify risk factors for assisted walking, establish effective predictive models, and conduct prognosis evaluation.

Method used

By collecting the basic characteristics and clinical characteristics of patients with multi-system atrophy, establishing an information database, sorting and processing data, analyzing data bias, identifying risk factors for assisted walking, establishing Cox regression prediction models and machine learning models, building risk prediction models, and conducting prognosis assessments.

Benefits of technology

Help clinicians understand the natural history of the disease, predict the disease progression, improve data quality, accurately predict the impact of the disease on quality of life, identify risk factors for assisted walking, and improve clinical management and treatment strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005231950440000042
    Figure BDA0005231950440000042
  • Figure BDA0005231950440000122
    Figure BDA0005231950440000122
  • Figure FDA0005231950430000021
    Figure FDA0005231950430000021
Patent Text Reader

Abstract

The invention discloses a multi-system atrophy risk prediction model, relates to the technical field of medical treatment, and solves the problem that data of a multi-system atrophy patient is difficult to collect and process. Secondly, data bias of multi-system atrophy patients is difficult to analyze and control; then, multi-system atrophy auxiliary walking risk factors are difficult to analyze, and the auxiliary walking risk factors are identified; according to analysis and recognition results, it is difficult to establish a multi-system atrophy assisted walking Cox regression prediction model, a multi-system atrophy assisted walking machine learning model and a multi-system atrophy risk prediction model; and finally, prognosis evaluation is difficult to carry out according to a multi-system atrophy risk prediction model. According to the method, clinical characteristics of a multi-system atrophy patient are collected, risk factors of walking assistance are analyzed, a prediction model is established, and the stability of the prediction model is verified, so that clinical doctors are helped to understand natural history of diseases and predict disease progresses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of medical technology, and specifically is a multiple system atrophy risk prediction model. Background Art

[0002] Multiple-system atrophy (MSA) is a rare, adult-onset progressive neurodegenerative disease with a variable combination of autonomic failure, cerebellar ataxia, pyramidal tract damage, and parkinsonism as clinical manifestations. Compared with other neurodegenerative diseases, MSA progresses rapidly and has a poor prognosis. At present, the diagnosis of MSA mainly refers to the third edition of the diagnostic criteria proposed by MDS in 2022, which divides MSA into four levels: "neuropathologically confirmed MSA", "clinically confirmed", "clinically likely MSA", and "possible prodromal MSA", as well as two phenotypes: parkinsonian (MSA-P) and cerebellar (MSA-C). The onset is usually in the sixth decade of life, with no gender difference. The prognosis of MSA varies considerably, with an average survival of 8 to 10 years after the onset of symptoms, and a few patients survive for more than 15 years. Given the aggressiveness of the disease and the lack of effective treatment to date, in-depth understanding of the prognostic factors of MSA, establishment of an effective prognostic model, and accurate judgment of patient prognosis can guide clinical trial design and therapeutic intervention, prolong the survival of MSA patients, and improve the quality of life of patients, which is crucial to providing the best treatment and care support for MSA patients.

[0003] The existing problems are as follows: first, it is difficult to collect and process data of patients with multiple system atrophy; second, it is difficult to analyze and control the bias of data of patients with multiple system atrophy; then, it is difficult to analyze the risk factors for assisted walking in multiple system atrophy and identify the risk factors for assisted walking; based on the analysis and identification results, it is difficult to establish a Cox regression prediction model for assisted walking in multiple system atrophy, a machine learning model for assisted walking in multiple system atrophy, and a risk prediction model for multiple system atrophy; finally, it is difficult to conduct prognostic assessment based on the risk prediction model for multiple system atrophy. Summary of the invention

[0004] The present invention aims to solve at least one of the technical problems existing in the prior art; to this end, the present invention proposes a multiple system atrophy risk prediction model to solve the following technical problems:

[0005] First, it is difficult to collect and process data from patients with multiple system atrophy; second, it is difficult to analyze and control the bias of data from patients with multiple system atrophy; then, it is difficult to analyze the risk factors for assisted walking in multiple system atrophy and identify the risk factors for assisted walking; based on the analysis and identification results, it is difficult to establish a Cox regression prediction model for assisted walking in multiple system atrophy, a machine learning model for assisted walking in multiple system atrophy, and a risk prediction model for multiple system atrophy; finally, it is difficult to conduct prognostic assessment based on the risk prediction model for multiple system atrophy.

[0006] To solve the above problems, the first aspect of the present invention provides a multiple system atrophy risk prediction model, comprising the following steps:

[0007] S1: Collect the basic characteristics and clinical characteristics of patients with multiple system atrophy, form an information collection form and follow-up form, and establish an information database;

[0008] S2: Data collation and entry processing based on information database;

[0009] S3: Analyze the bias of the data of patients with multiple system atrophy and control the data quality of patients with multiple system atrophy;

[0010] S4: Analyze the risk factors of assisted walking in multiple system atrophy based on the information database, and identify the risk factors of assisted walking; establish a Cox regression prediction model for assisted walking in multiple system atrophy based on the analysis and identification results; establish a machine learning model for assisted walking in multiple system atrophy based on the information database; predict the course of the disease by constructing a risk prediction model for multiple system atrophy, and verify the predictive efficacy and stability of the risk prediction model for multiple system atrophy;

[0011] S5: Prognostic assessment based on the multiple system atrophy risk prediction model.

[0012] As a further solution of the present invention: Step S1 comprises the following steps:

[0013] According to the collected basic characteristics and clinical features of patients with multiple system atrophy, an information collection form and a follow-up form were formed;

[0014] The information collection form is divided into five parts: basic characteristics, clinical manifestations, course characteristics, auxiliary examinations, and baseline scale scores; among them, the basic characteristics include: gender, age, height, weight, education level, onset age, and diagnostic classification; clinical manifestations include: first symptoms and red flag signs; course characteristics include: time from symptom onset to medical treatment, time from the onset of motor symptoms and non-motor symptoms, and time from onset to assisted walking; auxiliary examinations include: magnetic resonance imaging (MRI) and functional magnetic resonance imaging (fMRI), anal sphincter electromyography, video polysomnography (VPSG), and urinary system ultrasound; all patients with multiple system atrophy use a unified multiple system atrophy scoring scale to collect the baseline scale scores of patients with multiple system atrophy;

[0015] The follow-up table includes: follow-up time, disease changes and prognosis assessment endpoint events; by evaluating the patient's disease changes according to the follow-up time, to assist walking, including: using a cane, a walker or a companion's arm to provide support at any time as a prognosis assessment endpoint event;

[0016] Based on the information collection form and follow-up form, patients with multiple system atrophy were followed up every 6 months through face-to-face or telephone follow-up, and an information database was established based on the follow-up content.

[0017] As a further solution of the present invention: the step S2 comprises the following steps:

[0018] According to the established information database, double entry is performed by using EpiData software, and the two entered databases are checked by setting data intervals;

[0019] Delete samples and factors with more than 30% missing data; interpolate missing data by using the nearest neighbor method;

[0020] Let X = {x1, x2, …, x n} is a set of measurement data, and the normality test of the measurement data is carried out through the analysis formula:

[0021]

[0022] Where W is the test statistic, n is the sample size, and a i is the coefficient based on sample sorting, x (i) is the sorted data, x is the sample mean; when it conforms to the normal distribution, it is expressed as mean x±standard deviation s, and the t test is used; when it does not conform to the normal distribution, it is expressed as median, quartile P25 and P75, and the Mann-Whitney U test is used;

[0023] The count data were expressed as frequency n and percentage %, and Pearson χ 2 test.

[0024] As a further solution of the present invention: Step S3 comprises the following steps:

[0025] During the follow-up collection and data analysis stage, the quality of the data of patients with multiple system atrophy was controlled by adjusting the bias; the bias included: loss to follow-up bias, zero point bias, aggregation bias, information bias and confounding bias; among them, the information bias included: investigation bias, recall bias and blind judgment of outcomes.

[0026] As a further solution of the present invention: the step S4 analyzes the risk factors of assisted walking for multiple system atrophy according to the information database and identifies the risk factors of assisted walking, including the following steps:

[0027] Univariate Cox proportional hazard regression model and multivariate Cox proportional hazard regression model were used to analyze the risk factors of assisted walking for multiple system atrophy.

[0028] Collect clinical data of patients with multiple system atrophy according to the information database, and clean the collected clinical data, wherein the clinical data include: motor symptoms, autonomic dysfunction, cognitive function, genetic and environmental factors;

[0029] A univariate Cox regression model was established for each independent variable in the collected and processed clinical data; each independent variable in the clinical data was input into the univariate Cox regression model for training, and the relationship between each independent variable and the risk of assisted walking was evaluated; based on the trained univariate Cox regression model, the significant p value was evaluated by looking at the HR risk ratio and 95% confidence interval of each independent variable, and significantly related variables were screened out according to the significant p value or AIC / BIC index;

[0030] According to the significantly correlated variables screened out by the univariate Cox regression model, the significantly correlated variables screened out were used as independent variables, and the start time or status of assisted walking was used as the dependent variable to be input into the multivariate Cox regression model for training; during the training process of the multivariate Cox regression model, it was checked whether there was an interaction between the variables. When there was an interaction, the corresponding interaction term was added to the multivariate Cox regression model. When there was no interaction, the corresponding interaction term was not added to the multivariate Cox regression model; according to the trained multivariate Cox regression model, the influence of each variable after controlling other variables was evaluated by checking the adjusted HR risk ratio and 95% confidence interval; the multivariate Cox regression model was evaluated and verified by checking the goodness of fit and residual analysis of the multivariate Cox regression model;

[0031] Statistical methods, including survival analysis and multivariate regression analysis, were used to identify factors associated with the rate of disease progression and survival. Based on the analysis results of the univariate Cox hazard proportional regression model and the multivariate Cox hazard proportional regression model, the risk factors for assisted walking were identified by using the HR risk ratio, significant p value and 95% confidence interval in the analysis results.

[0032] As a further solution of the present invention: in step S4, a Cox regression prediction model for assisted walking of multiple system atrophy is established according to the analysis and recognition results, comprising the following steps:

[0033] Based on the results of the analysis to identify risk factors for assisted walking, the data of all patients with multiple system atrophy in the information database were randomly divided into a training set and a validation set in a ratio of 8:2;

[0034] In the training set, the univariate Cox proportional hazard model was used to screen the risk factors for assisted walking with a P value less than 0.01. The screened risk factors for assisted walking were input into the multivariate Cox proportional hazard model and stepwise regression was performed in the training set. Based on the results of stepwise regression, the Cox regression prediction model for assisted walking of multiple system atrophy was established.

[0035] Use forest plots and nomograms to visualize the Cox regression prediction model for assisted walking in multiple system atrophy; draw forest plots to display the HR value and confidence interval of each risk factor, and draw nomograms to predict the risk of assisted walking in patients;

[0036] On the training set, the prediction performance of the Cox regression prediction model for assisted walking of multiple system atrophy was evaluated by calculating the AUC value, C index and credible interval of the Cox regression prediction model for assisted walking of multiple system atrophy. The evaluation steps of the training set were repeated on the validation set to test the generalization ability of the Cox regression prediction model for assisted walking of multiple system atrophy.

[0037] As a further solution of the present invention: in step S4, a multiple system atrophy assisted walking machine learning model is established according to the information database, comprising the following steps:

[0038] In the information database, the data of all patients with multiple system atrophy were randomly divided into a training set and a validation set in a ratio of 8:2, and the factor variables were onehot encoded;

[0039] After using the k-fold validation method to screen features, Xgboost and lightgbm were used to establish multiple systemic atrophy assisted walking machine learning models respectively; the screened features were input into the multiple systemic atrophy assisted walking machine learning model for training;

[0040] According to the trained multiple system atrophy assisted walking machine learning model, the AUC value, accuracy, balance accuracy, weighted F1 score and C index of the multiple system atrophy assisted walking machine learning model were analyzed and calculated.

[0041] As a further solution of the present invention: in step S4, predicting the disease process by constructing a multiple system atrophy risk prediction model and verifying the predictive efficacy and stability of the multiple system atrophy risk prediction model comprises the following steps:

[0042] According to the screened features, a statistical or machine learning model is selected, and the screened features and corresponding labels are input into a multiple system atrophy risk prediction model for training, wherein the labels include: whether the patient suffers from multiple system atrophy or the severity of the disease process; during the training process, the parameters of the multiple system atrophy risk prediction model are adjusted by using parameter adjustment technology and integrated learning methods including: random forest or gradient boosting machine;

[0043] Based on the trained multiple system atrophy risk prediction model, the generalization ability of the multiple system atrophy risk prediction model was evaluated by using cross-validation, and the multiple system atrophy risk prediction model was applied to external independent data sets; new data of multiple system atrophy patients collected and processed in real time were input into the trained multiple system atrophy risk prediction model for prediction; the predictive effectiveness and stability of the multiple system atrophy risk prediction model were evaluated by using accuracy, recall rate, and area under the ROC curve AUC.

[0044] As a further solution of the present invention: the step S5 comprises the following steps:

[0045] By developing and using quality of life assessment tools, we monitor the treatment effects and changes in quality of life of patients with multiple system atrophy; based on the analysis results of the multiple system atrophy risk prediction model, we convert it into a prognostic scoring system; by building a WebApp, we develop a predictive model scoring system to monitor and evaluate patients with multiple system atrophy, and promote its application through web pages.

[0046] As a further scheme of the present invention: the following steps are included: the data of all patients with multiple system atrophy are analyzed based on IBMSPSS 25 and R software version 4.4.1, and the rms package, pROC package, caret package, timeROC package, ggplot2 package, dcurves package, survminer package, glmnet package, xgboost package, and lightgbm package are used to perform data segmentation, model construction, and drawing of ROC curves and correction curves.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] The present invention collects clinical characteristics of patients with multiple system atrophy, analyzes risk factors for assisted walking, establishes a prediction model, and verifies its stability, thereby helping clinicians understand the natural history of the disease and predict the course of the disease;

[0049] The present invention controls the data quality of multiple system atrophy patients by analyzing the bias of the data of multiple system atrophy patients, and can more accurately predict and reduce the impact of the disease on the quality of life of patients;

[0050] The present invention identifies the risk factors for assisted walking by analyzing the risk factors for assisted walking in multiple system atrophy; by analyzing the identification results, a Cox regression prediction model for assisted walking in multiple system atrophy and a machine learning model for assisted walking in multiple system atrophy are established, which can provide a deeper understanding of the disease mechanism of multiple system atrophy, and predict the disease course by constructing a risk prediction model for multiple system atrophy, verify the predictive efficacy and stability of the risk prediction model for multiple system atrophy, provide patients with more accurate risk assessment and prognosis prediction, thereby improving clinical management and treatment strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0052] Figure 1 The present invention is a flow chart of the method. DETAILED DESCRIPTION

[0053] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0054] See also Figure 1 As shown, the first embodiment of the present invention provides a multiple system atrophy risk prediction model, comprising the following steps:

[0055] S1: Collect the basic characteristics and clinical characteristics of patients with multiple system atrophy, form an information collection form and follow-up form, and establish an information database;

[0056] S2: Data collation and entry processing based on information database;

[0057] S3: Analyze the bias of the data of patients with multiple system atrophy and control the data quality of patients with multiple system atrophy;

[0058] S4: Analyze the risk factors of assisted walking in multiple system atrophy based on the information database, and identify the risk factors of assisted walking; establish a Cox regression prediction model for assisted walking in multiple system atrophy based on the analysis and identification results; establish a machine learning model for assisted walking in multiple system atrophy based on the information database; predict the course of the disease by constructing a risk prediction model for multiple system atrophy, and verify the predictive efficacy and stability of the risk prediction model for multiple system atrophy;

[0059] S5: Prognostic assessment based on the multiple system atrophy risk prediction model.

[0060] Specifically, data were collected through electronic health records, patient questionnaires, laboratory tests and other methods. Information collection forms and follow-up plans were designed based on the basic characteristics and clinical characteristics of patients with multiple system atrophy. The information database established was used to store and manage the collected patient information and follow-up data, and to perform preliminary data cleaning and verification. The collected data were sorted to remove duplicate, erroneous or missing data. The sorted data were entered into the information database to ensure the accuracy and consistency of the data. The collected data were analyzed to identify possible biases to control the data quality of patients with multiple system atrophy. Statistical methods were used to analyze the data and identify risk factors associated with assisted walking in multiple system atrophy. Based on the identification results, a Cox regression prediction model was established to predict the risk of assisted walking in patients. A suitable machine learning algorithm was selected to establish a machine learning model for assisted walking in multiple system atrophy. The model was trained using the training data set, and the optimal model parameters were found through parameter adjustment technology. The generalization ability of the model was evaluated using cross-validation, and the model was applied to an external independent data set for testing. The established risk prediction model for multiple system atrophy was used to evaluate the prognosis of patients and predict their future risk of assisted walking and the rate of disease progression. Based on the evaluation results, a long-term follow-up study will be designed and implemented to collect detailed information on disease progression and patient survival to verify the accuracy and stability of the prediction model.

[0061] In one embodiment of the present invention, the step S1 comprises the following steps:

[0062] According to the collected basic characteristics and clinical features of patients with multiple system atrophy, an information collection form and a follow-up form were formed;

[0063] The information collection form is divided into five parts: basic characteristics, clinical manifestations, course characteristics, auxiliary examinations, and baseline scale scores; among them, the basic characteristics include: gender, age, height, weight, education level, onset age, and diagnostic classification; clinical manifestations include: first symptoms and red flag signs; course characteristics include: time from symptom onset to medical treatment, time from the onset of motor symptoms and non-motor symptoms, and time from onset to assisted walking; auxiliary examinations include: magnetic resonance imaging (MRI) and functional magnetic resonance imaging (fMRI), anal sphincter electromyography, video polysomnography (VPSG), and urinary system ultrasound; all patients with multiple system atrophy use a unified multiple system atrophy scoring scale to collect the baseline scale scores of patients with multiple system atrophy;

[0064] The follow-up table includes: follow-up time, disease changes and prognosis assessment endpoint events; by evaluating the patient's disease changes according to the follow-up time, to assist walking, including: using a cane, a walker or a companion's arm to provide support at any time as a prognosis assessment endpoint event;

[0065] Based on the information collection form and follow-up form, patients with multiple system atrophy were followed up every 6 months through face-to-face or telephone follow-up, and an information database was established based on the follow-up content.

[0066] Specifically, patients with multiple system atrophy in the Department of Neurology of the First Affiliated Hospital of Air Force Medical University, the Department of Neurology of Shaanxi Provincial People's Hospital, and the Third Hospital of Xi'an were collected continuously. Inclusion criteria include: meeting the MSA diagnostic criteria proposed by MDS in 2022, cooperating with follow-up and signing informed consent, and detailed and reliable medical history recall. Exclusion criteria include: suffering from other neurodegenerative diseases, other reasons affecting limb movement and condition assessors, and combined with other underlying diseases or serious complications. According to the information collection form, record the patient's gender information, current age, height, weight, highest education level, age when symptoms of multiple system atrophy first appeared, and record the patient's diagnostic level according to the diagnostic criteria of multiple system atrophy; record the first symptoms of the patient and record whether the patient has key symptoms or signs of diagnostic significance; record the time interval from the first onset of symptoms to medical treatment and the time interval from the onset of motor symptoms to the simultaneous onset of non-motor symptoms; record the time interval from the onset of the disease to the need for auxiliary walking tools, such as canes, walkers, etc.; record the patient's nuclear magnetic resonance MRI and functional magnetic resonance fMRI examination results, anal sphincter electromyography examination results, video polysomnography VPSG examination results, and urinary system ultrasound examination results; use a unified multiple system atrophy scoring scale to quantify the severity of patients' symptoms and quality of life. Record the date and time of each follow-up according to the follow-up form; record the changes in the patient's condition since the last follow-up, including improvement or worsening of symptoms, the emergence of new symptoms, etc.; use assisted walking as the endpoint event of prognostic evaluation, and specifically record whether the patient starts using a cane, walker, or companion's arm for support, as well as the starting time and frequency of use. Follow-up is conducted face-to-face or by phone every 6 months. In order to increase the follow-up rate, each patient leaves 3 contact numbers and a dedicated person is assigned to follow up. According to the contents of the information collection form and the follow-up form, enter the patient's information into the information database. The database should have good data management and storage functions to ensure the accuracy and completeness of the data. As the follow-up progresses, the patient information in the database should be updated in a timely manner to ensure the timeliness and accuracy of the data.

[0067] In one embodiment of the present invention, the step S2 comprises the following steps:

[0068] According to the established information database, double entry is performed by using EpiData software, and the two entered databases are checked by setting data intervals;

[0069] Delete samples and factors with more than 30% missing data; interpolate missing data by using the nearest neighbor method;

[0070] Let X = {x1, x2, …, x n} is a set of measurement data, and the normality test of the measurement data is carried out through the analysis formula:

[0071]

[0072] Where W is the test statistic, n is the sample size, and a i is the coefficient based on sample sorting, x (i) is the sorted data, x is the sample mean; when it conforms to the normal distribution, it is expressed as mean x±standard deviation s, and the t test is used; when it does not conform to the normal distribution, it is expressed as median, quartile P25 and P75, and the Mann-Whitney U test is used;

[0073] The count data were expressed as frequency n and percentage %, and Pearson χ 2 test.

[0074] Specifically, a dedicated person is responsible for checking the integrity of the information collection form and the follow-up form, establishing a database for the follow-up content, and trained personnel use EpiData software for double entry. According to the admission rate of outpatient and inpatient multiple system atrophy patients in each hospital and the machine learning experience estimate, and considering the 10% loss rate, at least 330 cases should be collected. In EpiData, double entry is a commonly used data entry quality control method. It requires two people to enter the same data twice to reduce entry errors. First, enter the first data and save it as a database file. Then copy the database file for the second entry. Close the original database file and open the copied file for the second entry. In double entry mode, EpiData can compare the data entered twice in real time, or enter the same data in different databases and then compare them to find entry errors. Before entering data, it is necessary to set a reasonable data interval to ensure that the entered data is within a reasonable range. This can be achieved through the verification function of EpiData. The verification function can check the rationality and consistency of the data, such as whether there is data beyond the set range, whether there is a logical error, etc. If an error is found, EpiData will prompt the data entry to modify it. After the modification is completed, it is necessary to verify again until the data is correct. For samples and factors with more than 30% missing data, you can consider deleting them directly. Because a large amount of missing data will affect the accuracy and reliability of data analysis. When deleting, you need to pay attention to deleting all data related to the sample or factor at the same time to maintain data consistency. For samples or factors with less missing data, interpolation methods can be used to fill them. The neighbor method is a commonly used interpolation method, which infers the value of the missing data based on the data values ​​before and after the missing data. After the interpolation is completed, the interpolated data needs to be verified to ensure the accuracy and rationality of the interpolated data. Normality test is an important step before data analysis. It is used to determine whether the data conforms to the normal distribution. Because normal distribution is the premise assumption of many statistical methods. For quantitative data, normality tests can be performed using methods such as Shapiro-Wilk test or Kolmogorov-Smirnov test. If the data conform to normal distribution, the mean ± standard deviation can be used to express the data, and the t-test method can be used for data analysis. If the data do not conform to normal distribution, the median, quartile P25 and P75 need to be used to express the data, and non-parametric test methods such as Mann-Whitney U test can be used for data analysis. For count data, n (%) can be used to express it, and Pearson χ 2 Test methods for data analysis.

[0075] In one embodiment of the present invention, step S3 comprises the following steps:

[0076] During the follow-up collection and data analysis stage, the quality of the data of patients with multiple system atrophy was controlled by adjusting the bias; the bias included: loss to follow-up bias, zero point bias, aggregation bias, information bias and confounding bias; among them, the information bias included: investigation bias, recall bias and blind judgment of outcomes.

[0077] Specifically, due to the long duration of the study, loss to follow-up bias is more common, mostly caused by patient migration, going out, unwillingness to cooperate or death from non-endpoint events. In order to improve the follow-up rate of patients, eligible and compliant research subjects can be selected, and a follow-up time of every 6 months can be set. Various forms such as outpatient, inpatient or telephone follow-up can be adopted. Each patient leaves 3 contact numbers and a dedicated person is responsible for follow-up. Zero-point bias is a common bias in prognostic studies, which is caused by the inability to select the same moment of observation of the disease. In order to control zero-point bias, the medical history should be understood in detail, and the main clinical events should be recorded strictly according to the patient's course of disease. The research subjects should be kept at the same stage of the disease as much as possible, and the first admission time should be used as the starting point for follow-up. Multicenter research and a three-level doctor diagnosis system are adopted to avoid misdiagnosis and reduce the bias caused by examination and diagnosis. Formulate inclusion criteria, strictly follow the relevant diagnostic criteria for inclusion, and reduce collection bias. Information bias includes investigation bias, recall bias and blind judgment of outcomes. In order to control quality, investigators undergo unified training and use unified guidance terms to conduct truthful and objective investigations on the subjects to reduce investigation bias. The clinical characteristics are collected by first-, second- and third-line physicians respectively. The discrepant clinical characteristics are collected repeatedly many times to finally form a unified opinion and reduce information bias. Most patients are middle-aged and elderly, and there may be errors between oral recollections and the actual situation due to memory distortion. The spouse or children of the respondents are required to be present during the investigation, and the method of verifying information with the spouse or children of the respondents is adopted to reduce recall bias. In order to avoid suspicion bias and expectation bias, blind judgment of outcomes is adopted. Factors such as age and gender are usually the most common confounding biases, which can be controlled in the data analysis stage.

[0078] In one embodiment of the present invention, the step S4 analyzes the risk factors for assisted walking of multiple system atrophy according to the information database and identifies the risk factors for assisted walking, including the following steps:

[0079] Univariate Cox proportional hazard regression model and multivariate Cox proportional hazard regression model were used to analyze the risk factors of assisted walking for multiple system atrophy.

[0080] Collect clinical data of patients with multiple system atrophy according to the information database, and clean the collected clinical data, wherein the clinical data include: motor symptoms, autonomic dysfunction, cognitive function, genetic and environmental factors;

[0081] A univariate Cox regression model was established for each independent variable in the collected and processed clinical data; each independent variable in the clinical data was input into the univariate Cox regression model for training, and the relationship between each independent variable and the risk of assisted walking was evaluated; based on the trained univariate Cox regression model, the significant p value was evaluated by looking at the HR risk ratio and 95% confidence interval of each independent variable, and significantly related variables were screened out according to the significant p value or AIC / BIC index;

[0082] According to the significantly correlated variables screened out by the univariate Cox regression model, the significantly correlated variables screened out were used as independent variables, and the start time or status of assisted walking was used as the dependent variable to be input into the multivariate Cox regression model for training; during the training process of the multivariate Cox regression model, it was checked whether there was an interaction between the variables. When there was an interaction, the corresponding interaction term was added to the multivariate Cox regression model. When there was no interaction, the corresponding interaction term was not added to the multivariate Cox regression model; according to the trained multivariate Cox regression model, the influence of each variable after controlling other variables was evaluated by checking the adjusted HR risk ratio and 95% confidence interval; the multivariate Cox regression model was evaluated and verified by checking the goodness of fit and residual analysis of the multivariate Cox regression model;

[0083] Statistical methods, including survival analysis and multivariate regression analysis, were used to identify factors associated with the rate of disease progression and survival. Based on the analysis results of the univariate Cox hazard proportional regression model and the multivariate Cox hazard proportional regression model, the risk factors for assisted walking were identified by using the HR risk ratio, significant p value and 95% confidence interval in the analysis results.

[0084] Specifically, clinical data of patients with multiple system atrophy, including but not limited to motor symptoms, autonomic dysfunction, cognitive function, genetic and environmental factors, are collected from the information database, and the collected clinical data are cleaned. For each independent variable in the clinical data, a univariate Cox regression model is established to evaluate the relationship between each independent variable and the risk of assisted walking. Each independent variable is input into the corresponding univariate Cox regression model for training. After the training is completed, the degree of association between the variable and the risk of assisted walking is evaluated by looking at the HR hazard ratio and 95% confidence interval of each independent variable and evaluating the significant p value. The significantly related variables are screened out according to the significant p value or AIC / BIC index, and these variables will be used as candidate independent variables for subsequent multivariate Cox regression model analysis. The significantly related variables screened out by the univariate Cox regression model are used as independent variables, and the start time or status of assisted walking is used as the dependent variable, and are input into the multivariate Cox regression model for training. During the training process, it is necessary to check whether there is an interaction between the variables. If there is an interaction, the corresponding interaction term needs to be added to the model; if there is no interaction, there is no need to add the interaction term. After the training was completed, the effect of each variable after controlling for other variables was evaluated by looking at the adjusted HR hazard ratio and 95% confidence interval. Goodness of fit and residual analysis were performed on the multivariate Cox regression model to verify the reliability and accuracy of the model. This included checking whether the residuals of the model met the assumptions of normality and independence. Statistical methods such as survival analysis and multivariate regression analysis were used to further identify factors associated with the rate of progression and survival of multiple system atrophy. Based on the analysis results of the univariate Cox hazard proportional regression model and the multivariate Cox hazard proportional regression model, the risk factors for assisted walking were identified using information such as HR hazard ratio, significant p value and 95% confidence interval. These risk factors may include age of onset, disease stage, severity of motor symptoms, degree of autonomic dysfunction, etc.

[0085] In one embodiment of the present invention, in step S4, a Cox regression prediction model for multiple system atrophy assisted walking is established based on the analysis and recognition results, including the following steps:

[0086] Based on the results of the analysis to identify risk factors for assisted walking, the data of all patients with multiple system atrophy in the information database were randomly divided into a training set and a validation set in a ratio of 8:2;

[0087] In the training set, the univariate Cox proportional hazard model was used to screen the risk factors for assisted walking with a P value less than 0.01. The screened risk factors for assisted walking were input into the multivariate Cox proportional hazard model and stepwise regression was performed in the training set. Based on the results of stepwise regression, the Cox regression prediction model for assisted walking of multiple system atrophy was established.

[0088] Use forest plots and nomograms to visualize the Cox regression prediction model for assisted walking in multiple system atrophy; draw forest plots to display the HR value and confidence interval of each risk factor, and draw nomograms to predict the risk of assisted walking in patients;

[0089] On the training set, the prediction performance of the Cox regression prediction model for assisted walking of multiple system atrophy was evaluated by calculating the AUC value, C index and credible interval of the Cox regression prediction model for assisted walking of multiple system atrophy. The evaluation steps of the training set were repeated on the validation set to test the generalization ability of the Cox regression prediction model for assisted walking of multiple system atrophy.

[0090] Specifically, according to the results of analyzing and identifying the risk factors for assisted walking, the data of all patients with multiple system atrophy were extracted from the information database, and these data were randomly divided into a training set and a validation set at a ratio of 8:2. The training set was used to establish the model, and the validation set was used to test the generalization ability of the model. In the training set, each risk factor was screened using a univariate Cox hazard proportional model. The significance level was set to P<0.01, and the risk factors significantly associated with the risk of assisted walking were screened out. The screened risk factors were used as independent variables, and the start time or status of assisted walking was used as the dependent variable, and entered into the multivariate Cox hazard proportional model for stepwise regression. During the stepwise regression process, independent variables were gradually added or deleted according to the model fit and the significance of the variables until the optimal model structure was obtained. According to the results of stepwise regression, a Cox regression prediction model for assisted walking in multiple system atrophy was established to predict the risk of assisted walking in patients. The use of a forest plot to display the HR risk ratio and confidence interval of each risk factor helps to intuitively understand the degree of influence of each risk factor on the risk of assisted walking. The risk of assisted walking was predicted according to the specific situation of the patient, and a nomogram was drawn. The nomogram quantitatively associates risk factors with predicted risks, making it easier for clinicians and patients to understand the risk situation. On the training set, the predictive performance of the model was evaluated by calculating the area under the curve, C index, and credible interval of the Cox regression prediction model for assisted walking in multiple system atrophy. These indicators will reflect the fit and prediction accuracy of the model on the training set. The evaluation steps of the training set were repeated on the validation set to test the generalization ability of the Cox regression prediction model for assisted walking in multiple system atrophy. By comparing the evaluation results on the training set and the validation set, the stability and reliability of the model on different data sets can be understood.

[0091] In one embodiment of the present invention, the step S4 establishes a multiple system atrophy assisted walking machine learning model according to the information database, including the following steps:

[0092] In the information database, the data of all patients with multiple system atrophy were randomly divided into a training set and a validation set in a ratio of 8:2, and the factor variables were onehot encoded;

[0093] After using the k-fold validation method to screen features, Xgboost and lightgbm were used to establish multiple systemic atrophy assisted walking machine learning models respectively; the screened features were input into the multiple systemic atrophy assisted walking machine learning model for training;

[0094] According to the trained multiple system atrophy assisted walking machine learning model, the AUC value, accuracy, balance accuracy, weighted F1 score and C index of the multiple system atrophy assisted walking machine learning model were analyzed and calculated.

[0095] Specifically, the data of all patients with multiple system atrophy were extracted from the information database, and the data were randomly divided into a training set and a validation set in a ratio of 8:2, and the factor variables in the data, i.e., categorical variables, were one-hot encoded. Among them, the training set was used for model training and optimization, and the validation set was used to evaluate the performance of the model. The k-fold validation method was used to screen the data for features. The k-fold validation method is a commonly used cross-validation method that divides the data set into k non-overlapping subsets, and then takes one of the subsets as the validation set in turn, and the remaining subsets as the training set for training and validation. In each fold, the importance of the features is evaluated based on the performance of the model on the validation set, and the features that have a significant impact on the model performance are selected. Using the screened features, Xgboost and LightGBM machine learning models are established respectively. Xgboost is an efficient and powerful machine learning technology that is widely used in classification, regression, and ranking problems. LightGBM is a gradient boosting framework based on decision trees, which has the advantages of high efficiency, flexibility, and scalability. The screened features are input into the two models and trained using the training set. According to the trained machine learning model of multiple system atrophy assisted walking, the AUC value, accuracy, balanced accuracy, weighted F1 score and C index of the machine learning model of multiple system atrophy assisted walking were analyzed and calculated. Among them, the AUC value is the area under the curve, which is used to evaluate the classification performance of the model; the accuracy is the proportion of the number of correctly predicted samples to the total number of samples; the balanced accuracy is to calculate the accuracy of each category for the unbalanced data set and take the average; the weighted F1 score is the harmonic mean of the precision and recall rate, and the weighted F1 score takes into account the imbalance of the number of samples in different categories; the C index is a consistency index, which is used to evaluate the consistency between the model prediction results and the actual results. According to the results of the evaluation indicators, the model is optimized. The parameters of the model can be adjusted, and a more appropriate feature combination can be selected. Repeat the process of feature screening, model building and evaluation optimization until satisfactory model performance is obtained.

[0096] In one embodiment of the present invention, the step S4 predicts the disease process by constructing a multiple system atrophy risk prediction model and verifies the predictive efficacy and stability of the multiple system atrophy risk prediction model, including the following steps:

[0097] According to the screened features, a statistical or machine learning model is selected, and the screened features and corresponding labels are input into a multiple system atrophy risk prediction model for training, wherein the labels include: whether the patient suffers from multiple system atrophy or the severity of the disease process; during the training process, the parameters of the multiple system atrophy risk prediction model are adjusted by using parameter adjustment technology and integrated learning methods including: random forest or gradient boosting machine;

[0098] Based on the trained multiple system atrophy risk prediction model, the generalization ability of the multiple system atrophy risk prediction model was evaluated by using cross-validation, and the multiple system atrophy risk prediction model was applied to external independent data sets; new data of multiple system atrophy patients collected and processed in real time were input into the trained multiple system atrophy risk prediction model for prediction; the predictive effectiveness and stability of the multiple system atrophy risk prediction model were evaluated by using accuracy, recall rate, and area under the ROC curve AUC.

[0099] Specifically, according to the features related to the risk of multiple system atrophy, including the patient's age, gender, genetic information, clinical manifestations, imaging indicators, etc., the feature screening can be performed by statistical methods including correlation analysis, chi-square test, etc. or machine learning algorithms including recursive feature elimination, model-based feature selection, etc. According to the screened features, a suitable statistical or machine learning model is selected as the risk prediction model for multiple system atrophy. Commonly used models include logistic regression, support vector machine, random forest, gradient boosting machine, etc. When selecting a model, it is necessary to consider the complexity, interpretability, training efficiency and feasibility of the model in practical applications. The screened features and the corresponding labels, i.e., whether the patient has multiple system atrophy or the severity of the disease process, are input into the selected model for training. During the training process, parameter adjustment techniques such as grid search, random search, etc. are used to optimize the parameters of the model to improve the prediction performance of the model. In order to further improve the prediction performance of the model, an integrated learning method such as random forest or gradient boosting machine can be used. Random forest improves the stability and accuracy of the model by constructing multiple decision trees and integrating their prediction results. The gradient boosting machine constructs a strong classifier by iteratively training multiple weak classifiers and weightedly combining their results. When using ensemble learning methods, relevant parameters need to be adjusted, such as the number of trees and maximum depth in random forests; learning rate, number of iterations, regularization parameters, etc. in gradient boosting machines. Parameter adjustment can be performed by methods such as cross-validation to find the optimal parameter combination. The cross-validation method is used to evaluate the generalization ability of the risk prediction model for multiple system atrophy. Cross-validation divides the data set into multiple subsets, and one of the subsets is used as the validation set in turn, and the remaining subsets are used as the training set for training and validation. The stability and accuracy of the model are evaluated by calculating the average performance on multiple validation sets. The trained risk prediction model for multiple system atrophy is applied to an external independent data set for validation, which helps to evaluate the generalization ability and stability of the model on different data sets. The external independent data set should have similar feature distribution and label distribution as the training data set to ensure the reliability of the validation results. New data of patients with multiple system atrophy, including clinical manifestations, imaging indicators, etc., are collected and processed in real time. New data are preprocessed, such as missing value filling and outlier processing, to ensure the quality and consistency of the data. The processed new data is input into the trained multiple system atrophy risk prediction model for prediction. The model will output the predicted results of the patient's risk of multiple system atrophy or the severity of the disease process. Evaluation indicators such as accuracy, recall, and area under the ROC curve AUC are used to evaluate the predictive effectiveness and stability of the multiple system atrophy risk prediction model. The accuracy rate indicates the proportion of correct predictions made by the model; the recall rate indicates the proportion of patients who are truly ill that the model identifies; and the AUC value comprehensively reflects the performance of the model at different thresholds. As new data continues to accumulate, it is necessary to continuously monitor the performance of the model and make necessary updates and optimizations.By regularly evaluating the predictive effectiveness and stability of the model, the accuracy and reliability of the model in practical applications can be ensured.

[0100] In one embodiment of the present invention, the step S5 comprises the following steps:

[0101] By developing and using quality of life assessment tools, we monitor the treatment effects and changes in quality of life of patients with multiple system atrophy; based on the analysis results of the multiple system atrophy risk prediction model, we convert it into a prognostic scoring system; by building a WebApp, we develop a predictive model scoring system to monitor and evaluate patients with multiple system atrophy, and promote its application through web pages.

[0102] Specifically, the results of the prognostic model are converted into clinically applicable tools, such as prognostic scoring systems, and applied to the optimization of treatment strategies, including drug therapy, physical therapy, lifestyle intervention, etc., to improve treatment efficacy and patient satisfaction. As new data continue to accumulate, the prognostic model is regularly updated to ensure the timeliness and accuracy of its predictive ability. Participate in updating clinical guidelines and incorporate new research results into the diagnosis and treatment of patients with multiple system atrophy to guide clinical practice.

[0103] In one embodiment of the present invention, the following steps are included: the data of all patients with multiple system atrophy are analyzed based on IBMSPSS 25 and R software version 4.4.1, and the rms package, pROC package, caret package, timeROC package, ggplot2 package, dcurves package, survminer package, glmnet package, xgboost package, and lightgbm package are used for data segmentation, model construction, and drawing of ROC curves and correction curves.

[0104] Specifically, the data of all patients with multiple system atrophy were cleaned using IBM SPSS 25 and R software version 4.4.1 to remove duplicates, missing or outliers; the data set was divided into training set, validation set and test set using tools such as rms package and caret package, where the training set was used to build the model, the validation set was used to adjust the model parameters, and the test set was used to evaluate the model performance. The glmnet package was used for feature selection and regularization to improve the generalization ability of the model. The machine learning model was constructed using the xgboost package and lightgbm package. The pROC package was used to draw the ROC curve to evaluate the classification performance of the model. The optimal model parameters were selected based on the AUC value, sensitivity, specificity and other indicators of the ROC curve. The test set was used to evaluate the performance of the constructed model, and indicators such as accuracy, recall and F1 score were calculated. The timeROC package was used to draw the time-related ROC curve to evaluate the performance of the model at different time points. The ggplot2 package was used to draw scatter plots, box plots, heat maps, etc. to intuitively display the data distribution and feature relationships. The dcurves package was used to draw survival curves and cumulative risk curves to evaluate the prognosis of patients. The survminer package was used for survival analysis. The constructed model was applied to new data of patients with multiple system atrophy for prognosis prediction and risk assessment; based on the prediction results, personalized treatment strategies and lifestyle interventions could be formulated for patients. In practical applications, new patient data were continuously collected to verify and update the model. The accuracy and reliability of the model were evaluated by comparing the model prediction results with the actual clinical results.

[0105] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A risk prediction model for multiple system atrophy, characterized in that: The following steps are involved: S1: Collect the basic characteristics and clinical characteristics of patients with multiple system atrophy, form an information collection form and follow-up form, and establish an information database; S2: Data collation and entry processing based on information database; S3: Analyze the bias of the data of patients with multiple system atrophy and control the data quality of patients with multiple system atrophy; S4: Analyze the risk factors of assisted walking in multiple system atrophy based on the information database, and identify the risk factors of assisted walking; establish a Cox regression prediction model for assisted walking in multiple system atrophy based on the analysis and identification results; establish a machine learning model for assisted walking in multiple system atrophy based on the information database; predict the course of the disease by constructing a risk prediction model for multiple system atrophy, and verify the predictive efficacy and stability of the risk prediction model for multiple system atrophy; S5: Prognostic assessment based on the multiple system atrophy risk prediction model.

2. A multiple system atrophy risk prediction model according to claim 1, characterized in that: The step S1 comprises the following steps: According to the collected basic characteristics and clinical features of patients with multiple system atrophy, an information collection form and a follow-up form were formed; The information collection form is divided into five parts: basic characteristics, clinical manifestations, course characteristics, auxiliary examinations, and baseline scale scores; among them, the basic characteristics include: gender, age, height, weight, education level, onset age, and diagnostic classification; clinical manifestations include: first symptoms and red flag signs; course characteristics include: time from symptom onset to medical treatment, time from the onset of motor symptoms and non-motor symptoms, and time from onset to assisted walking; auxiliary examinations include: magnetic resonance imaging (MRI) and functional magnetic resonance imaging (fMRI), anal sphincter electromyography, video polysomnography (VPSG), and urinary system ultrasound; all patients with multiple system atrophy use a unified multiple system atrophy scoring scale to collect the baseline scale scores of patients with multiple system atrophy; The follow-up table includes: follow-up time, disease changes and prognosis assessment endpoint events; by evaluating the patient's disease changes according to the follow-up time, to assist walking, including: using a cane, a walker or a companion's arm to provide support at any time as a prognosis assessment endpoint event; Based on the information collection form and follow-up form, patients with multiple system atrophy were followed up every 6 months through face-to-face or telephone follow-up, and an information database was established based on the follow-up content.

3. A multiple system atrophy risk prediction model according to claim 1, characterized in that: The step S2 comprises the following steps: According to the established information database, double entry is performed by using EpiData software, and the two entered databases are checked by setting data intervals; Delete samples and factors with more than 30% missing data; interpolate missing data by using the nearest neighbor method; Let X = {x1, x2, …, x n } is a set of measurement data, and the normality test of the measurement data is carried out through the analysis formula: Where W is the test statistic, n is the sample size, and a i is the coefficient based on sample sorting, x (i) is the sorted data, x is the sample mean; when it conforms to the normal distribution, it is expressed as mean x±standard deviation s, and the t test is used; when it does not conform to the normal distribution, it is expressed as median, quartile P25 and P75, and the Mann-Whitney U test is used; The count data were expressed as frequency n and percentage %, and Pearson χ 2 test.

4. A multiple system atrophy risk prediction model according to claim 1, characterized in that: The step S3 comprises the following steps: During the follow-up collection and data analysis stage, the quality of the data of patients with multiple system atrophy was controlled by adjusting the bias; the bias included: loss to follow-up bias, zero point bias, aggregation bias, information bias and confounding bias; among them, the information bias included: investigation bias, recall bias and blind judgment of outcomes.

5. The multiple system atrophy risk prediction model according to claim 1, characterized in that: The step S4 analyzes the risk factors for assisted walking of multiple system atrophy according to the information database and identifies the risk factors for assisted walking, including the following steps: Univariate Cox proportional hazard regression model and multivariate Cox proportional hazard regression model were used to analyze the risk factors of assisted walking for multiple system atrophy. Collect clinical data of patients with multiple system atrophy according to the information database, and clean the collected clinical data, wherein the clinical data include: motor symptoms, autonomic dysfunction, cognitive function, genetic and environmental factors; A univariate Cox regression model was established for each independent variable in the collected and processed clinical data; each independent variable in the clinical data was input into the univariate Cox regression model for training, and the relationship between each independent variable and the risk of assisted walking was evaluated; based on the trained univariate Cox regression model, the significant p value was evaluated by looking at the HR risk ratio and 95% confidence interval of each independent variable, and significantly related variables were screened out according to the significant p value or AIC / BIC index; According to the significantly correlated variables screened out by the univariate Cox regression model, the significantly correlated variables screened out were used as independent variables, and the start time or status of assisted walking was used as the dependent variable to be input into the multivariate Cox regression model for training; during the training process of the multivariate Cox regression model, it was checked whether there was an interaction between the variables. When there was an interaction, the corresponding interaction term was added to the multivariate Cox regression model. When there was no interaction, the corresponding interaction term was not added to the multivariate Cox regression model; according to the trained multivariate Cox regression model, the influence of each variable after controlling other variables was evaluated by checking the adjusted HR risk ratio and 95% confidence interval; the multivariate Cox regression model was evaluated and verified by checking the goodness of fit and residual analysis of the multivariate Cox regression model; Statistical methods, including survival analysis and multivariate regression analysis, were used to identify factors associated with the rate of disease progression and survival. Based on the analysis results of the univariate Cox hazard proportional regression model and the multivariate Cox hazard proportional regression model, the risk factors for assisted walking were identified by using the HR risk ratio, significant p value and 95% confidence interval in the analysis results.

6. A risk prediction model for multiple system atrophy according to claim 5, characterized in that: In step S4, a Cox regression prediction model for assisted walking for multiple system atrophy is established based on the analysis and recognition results, including the following steps: Based on the results of the analysis to identify risk factors for assisted walking, the data of all patients with multiple system atrophy in the information database were randomly divided into a training set and a validation set in a ratio of 8:2; In the training set, the univariate Cox proportional hazard model was used to screen the risk factors for assisted walking with a P value less than 0.

01. The screened risk factors for assisted walking were input into the multivariate Cox proportional hazard model and stepwise regression was performed in the training set. Based on the results of stepwise regression, the Cox regression prediction model for assisted walking of multiple system atrophy was established. Use forest plots and nomograms to visualize the Cox regression prediction model for assisted walking in multiple system atrophy; draw forest plots to display the HR value and confidence interval of each risk factor, and draw nomograms to predict the risk of assisted walking in patients; On the training set, the prediction performance of the Cox regression prediction model for assisted walking of multiple system atrophy was evaluated by calculating the AUC value, C index and credible interval of the Cox regression prediction model for assisted walking of multiple system atrophy. The evaluation steps of the training set were repeated on the validation set to test the generalization ability of the Cox regression prediction model for assisted walking of multiple system atrophy.

7. The multiple system atrophy risk prediction model according to claim 1, characterized in that: In step S4, a multiple system atrophy assisted walking machine learning model is established according to the information database, including the following steps: In the information database, the data of all patients with multiple system atrophy were randomly divided into training set and validation set in a ratio of 8:2, and the factor variables were onehot encoded; After using the k-fold validation method to screen features, Xgboost and lightgbm were used to establish the multiple systemic atrophy assisted walking machine learning model respectively; the screened features were input into the multiple systemic atrophy assisted walking machine learning model for training; According to the trained multiple system atrophy assisted walking machine learning model, the AUC value, accuracy, balance accuracy, weighted F1 score and C index of the multiple system atrophy assisted walking machine learning model were analyzed and calculated.

8. The multiple system atrophy risk prediction model according to claim 1, characterized in that: In step S4, the disease process is predicted by constructing a multiple system atrophy risk prediction model, and the prediction efficacy and stability of the multiple system atrophy risk prediction model are verified, including the following steps: According to the screened features, a statistical or machine learning model is selected, and the screened features and corresponding labels are input into a multiple system atrophy risk prediction model for training, wherein the labels include: whether the patient suffers from multiple system atrophy or the severity of the disease process; during the training process, the parameters of the multiple system atrophy risk prediction model are adjusted by using parameter adjustment technology and integrated learning methods including: random forest or gradient boosting machine; Based on the trained multiple system atrophy risk prediction model, the generalization ability of the multiple system atrophy risk prediction model was evaluated by using cross-validation, and the multiple system atrophy risk prediction model was applied to external independent data sets; new data of multiple system atrophy patients collected and processed in real time were input into the trained multiple system atrophy risk prediction model for prediction; the predictive effectiveness and stability of the multiple system atrophy risk prediction model were evaluated by using accuracy, recall rate, and area under the ROC curve AUC.

9. The multiple system atrophy risk prediction model according to claim 1, characterized in that: The step S5 comprises the following steps: By developing and using quality of life assessment tools, we monitor the treatment effects and changes in quality of life of patients with multiple system atrophy; based on the analysis results of the multiple system atrophy risk prediction model, we convert it into a prognostic scoring system; by building a Web App, we develop a predictive model scoring system to monitor and evaluate patients with multiple system atrophy, and promote its application through web pages.

10. A risk prediction model for multiple system atrophy according to claim 1, characterized in that: The following steps are involved: The data of all patients with multiple system atrophy were analyzed based on IBM SPSS25 and R software version 4.4.

1. The rms package, pROC package, caret package, timeROC package, ggplot2 package, dcurves package, survminer package, glmnet package, xgboost package, and lightgbm package were used for data segmentation, model construction, and drawing of ROC curves and correction curves.

Citation Information

Cited By

  • Cognitive risk layering method and system based on double-task gaits

    CN122158149A