Construction method of circulating microorganism abundance prognostic scoring model
Through the construction method of the prognostic scoring model of circulating microbial abundance, the shortcomings in the construction of microbial feature selection and prognostic model in the prior art are solved, and more accurate and robust tumor prognosis evaluation is achieved, which is suitable for different types of cancer data.
Patent Information
- Application Number
- CN202510260498.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art has shortcomings in the selection of microbial characteristics and construction of prognostic models, making it difficult to fully understand the complex interactions between microorganisms and tumors, and the existing models have generalization capabilities and robustness problems in verification and application.
The construction method of the circulating microbial abundance prognostic scoring model was adopted, and the microbial characteristics that independently affect the overall survival were screened out through multi-level screening processes such as univariate COX regression, LASSOCOX regression and multivariate COX regression, and microbial microbial characteristics that independently affect the overall survival were constructed, and the MAPS scoring model was ensured through cross-validation and other evaluation methods.
This method can take into account microbial characteristics and clinical factors more comprehensively, improve the accuracy and stability of prognostic evaluation, enhance the predictive ability and clinical application value of the model, and is suitable for different types of cancer data.
Smart Images

Figure CN120199339A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedicine, and particularly to a method for constructing a prognostic scoring model of circulating microbial abundance. Background Art
[0002] In the current fields of cancer research and clinical treatment, exploring accurate and effective prognostic assessment methods has always been one of the core goals. With the in-depth research, people have gradually realized that there are intricate connections between microorganisms and tumors, and microbial characteristics play a crucial role in the occurrence, development, and patient prognosis of tumors.
[0003] More and more studies have shown that the human microbial community is closely related to the occurrence and development of tumors. The microbial community is widely distributed in various parts of the human body, such as the gut, oral cavity, respiratory tract, etc. They interact with host cells and affect the physiological and pathological states of the host. During tumorigenesis, microorganisms can play roles through various pathways. For example, certain microorganisms can produce carcinogenic metabolites, which may damage the DNA of host cells, leading to gene mutations and thus promoting tumorigenesis. In addition, microorganisms can also regulate the host immune system, affecting the activity and function of immune cells, and thus influencing tumor immune escape and growth.
[0004] In terms of cancer treatment, microorganisms also show important effects. Some studies have found that the composition of the gut microbial community can affect the response of cancer patients to chemotherapy, radiotherapy, and immunotherapy. For example, specific gut microorganisms can enhance the effect of immunotherapy and improve the survival rate of patients; while some other microorganisms may cause patients to develop drug resistance to treatment and reduce the treatment effect. Therefore, in-depth study of the interaction between microorganisms and tumors is of great significance for improving the levels of cancer diagnosis, treatment, and prognostic assessment.
[0005] However, currently in microbial feature selection, traditional methods have obvious deficiencies. Traditional microbial feature selection methods often overly focus on single biomarkers. Although this approach can, to a certain extent, reveal the associations between certain microorganisms and tumors, it ignores the complex interactions between the microbial community as a whole and tumors. The occurrence and development of tumors is a multi-factor and multi-step process, involving the synergistic effects of multiple microorganisms and their metabolites. Merely focusing on single biomarkers may miss a lot of important information, resulting in an inability to comprehensively understand the relationship between microorganisms and tumors, and thus affecting the accuracy of prognostic assessment.
[0006] With the rapid development of big data and multi-omics technologies, we are able to obtain massive amounts of microbiome data and related clinical data. These data contain rich information, but traditional microbial feature selection methods are difficult to fully tap their value. Traditional methods are unable to cope with high-dimensional and complex data and cannot effectively identify microbial feature combinations that are closely related to tumor prognosis. They often fail to take into account complex factors such as interactions between microorganisms, interactions between microorganisms and host genes, and associations between microorganisms and environmental factors, thus limiting the in-depth understanding of microbial-tumor interactions.
[0007] In addition to the limitations of microbial feature selection, many existing microbial-based prognostic models also have serious problems in validation. Many studies only build prognostic models based on specific data sets and evaluate and validate them on these data sets. However, different data sets may have very different characteristics and distributions of data due to differences in sample sources, detection methods, experimental conditions, and other factors. If a model only performs well on a specific data set and cannot be validated in other data sets, its generalization ability will be seriously questioned.
[0008] The lack of extensive external validation makes these models face great risks in actual clinical applications. When these models are applied to new and different data sets, their predictive accuracy may drop significantly and fail to provide reliable support for clinical decision-making. This has led to many research results that, although they have certain value in theory, are difficult to translate into actual clinical applications, limiting the role of microorganisms in tumor prognosis assessment.
[0009] In summary, there are many urgent problems to be solved in the selection of microbial features and the construction of prognostic models. There is an urgent need for a prognostic scoring model construction method that can comprehensively consider multiple microbial features and clinical factors, has good generalization ability and clinical application value, so as to improve the prognostic evaluation level of tumor patients and provide a more accurate and reliable basis for clinical treatment decisions. Summary of the invention
[0010] In order to solve the above problems, especially to address the deficiencies in the prior art, the present invention provides a method for constructing a prognostic scoring model for circulating microbial abundance that can solve the above problems.
[0011] To achieve the above purpose, the present invention adopts the following technical means:
[0012] A method for constructing a circulating microbial abundance prognostic scoring model comprises the following steps:
[0013] Step 1: Univariate COX regression analysis:
[0014] Using univariate COX regression analysis, circulating microbial features significantly associated with the overall survival of patients were screened out, where the criterion for significant association was a p-value < 0.05;
[0015] Step 2: LASSO COX regression analysis and cross-validation:
[0016] Perform LASSO COX regression analysis on the screened microbial features, and further narrow down the range of microbiome features through 10-fold cross-validation to ensure the selection of the most predictive features;
[0017] Step 3: Multivariate COX regression analysis:
[0018] Based on the features selected by LASSO, perform multivariate COX regression analysis to further confirm which microbial features are independent factors affecting overall survival. The criterion for significant association is a p-value < 0.05. At the same time, calculate the regression coefficients of these microbial features;
[0019] Step 4: Construction of the MAPS model:
[0020] Multiply the abundance of each microbial feature significantly associated with overall survival by its risk coefficient obtained in the multivariate COX regression analysis, and sum these products to obtain the MAPS score for each patient;
[0021] Step 5: Determination of the optimal cut-off value:
[0022] Use the R language package "maxstat" to determine the optimal cut-off value, and divide all patients into high-risk and low-risk groups;
[0023] Step 6: Performance evaluation:
[0024] Use the Kaplan-Meier survival curve to evaluate the predictive ability of the MAPS score for overall survival, and use the ROC curve to evaluate the predictive performance of the MAPS model, including the AUC values for 1-year, 3-year, and 5-year survival;
[0025] Step 7: Clinical correlation analysis:
[0026] Combined with clinical features, perform correlation analysis between the MAPS score and clinical features.
[0027] A further aspect of the present invention is that in step 1, the univariate COX regression analysis is performed using the statistical software R language, and missing value processing and standardization preprocessing are performed on the collected circulating microbial feature data before analysis.
[0028] A further aspect of the present invention is that in step 2, data normalization processing is performed on the screened microbial features.
[0029] In a further aspect of the present invention, in step three, the stepwise regression method is used to determine the microbial features that finally enter the model.
[0030] In a further aspect of the present invention, in step four, logarithmic transformation is performed on the abundance data of the microbial features.
[0031] In a further aspect of the present invention, in step five, the maximum log-rank statistic method in the "maxstat" package is used to find the optimal cut-off point.
[0032] In a further aspect of the present invention, in step six, the Bootstrap resampling method is used to verify the results of the Kaplan-Meier survival curve and the ROC curve.
[0033] In a further aspect of the present invention, in step seven, Spearman rank correlation analysis is used to evaluate the correlation between the MAPS score and clinical features.
[0034] In a further aspect of the present invention, in step seven, the clinical features include age, M stage, N stage, and clinical stage.
[0035] Advantages of the present invention:
[0036] 1. The present invention comprehensively considers microbial features and clinical factors, providing more comprehensive information for the prognostic model. It not only focuses on microbial features but also combines clinical factors such as age and tumor stage, obtaining information related to tumor prognosis from multiple dimensions, overcoming the limitation of traditional methods that only focus on a single biomarker, and more comprehensively considering all microbial features related to tumors, thus significantly enhancing the predictive ability of the model.
[0037] 2. The present invention adopts a multi-level screening process, effectively reducing the risk of overfitting. Through multi-level screening methods such as univariate COX regression, LASSO regression, and multivariate COX regression, unimportant features are gradually removed, ensuring that the finally selected features have strong independent prognostic ability and improving the stability and reliability of the model.
[0038] 3. The present invention realizes dynamic evaluation and personalized grouping. By using the R language package "maxstat" to determine the optimal cut-off value and dividing patients into different risk groups, it can accurately judge the risk level of each patient according to their specific situation, helping doctors formulate personalized treatment plans for patients in different risk groups and improving the pertinence and effectiveness of treatment.
[0039] 4. The present invention uses efficient evaluation and verification methods to ensure the robustness of the model. Through 10-fold cross-validation, the performance of the model on different data subsets is comprehensively evaluated to avoid inaccurate evaluation caused by the contingency of data partitioning. At the same time, the Kaplan-Meier survival curve is used to intuitively display the survival status of patients in different risk groups, and the ROC curve and the AUC values of 1-year, 3-year, and 5-year survival are used to quantify the prediction accuracy of the model, scientifically evaluating the model from multiple perspectives.
[0040] 5. The present invention improves the clinical application value of the model. This model can not only accurately predict the overall survival of patients, but also combine the prediction results with clinical information, providing more valuable references for clinicians in formulating treatment plans and judging patient prognosis, etc., helping doctors make more scientific and reasonable clinical decisions, and thus improving the survival rate and quality of life of patients.
[0041] 6. The present invention has good adaptability and scalability. The structure of the solution is flexible and can be applied to other types of cancer data or different microbiome data, having a wide application prospect in tumor research and clinical practice, and being able to provide effective prognostic evaluation services for more types of tumor patients.
[0042] 7. The present invention is more in-depth in mining microbial features. Through a multi-level analysis method, complex microbial-tumor interactions can be mined, and compared with traditional methods, it reveals the relationship between microorganisms and tumor prognosis more comprehensively, providing a new perspective and method for tumor research.
[0043] 8. The personalized evaluation of the present invention is more in line with the actual situation of patients. Considering the individual differences among different patients, the model can perform personalized risk assessment based on the microbial features and clinical information of patients, providing a prognostic prediction more in line with the actual situation of each patient, reflecting the concept of precision medicine.
[0044] 9. The present invention helps to improve the rational allocation of medical resources. Through accurate risk assessment and personalized grouping, doctors can give more attention and resource investment to high-risk patients and adopt relatively conservative treatment plans for low-risk patients, thus realizing the optimal allocation of medical resources.
[0045] 10. The present invention promotes the cross-integration of multiple disciplines. The construction of this model involves multiple disciplinary fields such as microbiology, statistics, and clinical medicine, promoting the communication and cooperation between different disciplines and providing interdisciplinary ideas and methods for solving complex medical problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a flowchart of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0047] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0048] Example 1
[0049] As Figure 1 shown, a method for constructing a prognostic scoring model of circulating microbial abundance includes the following steps:
[0050] Step 1. Univariate COX regression analysis:
[0051] Using univariate COX regression analysis, circulating microbial features significantly correlated with the overall survival (OS) of patients are screened out, where the determination criterion for significant correlation is a p-value < 0.05;
[0052] Step 2. LASSO COX regression analysis and cross-validation:
[0053] Perform LASSO COX regression analysis on the screened microbial features, and further narrow the range of microbiome features through 10-fold cross-validation to ensure the selection of the most predictive features;
[0054] Step 3. Multivariate COX regression analysis:
[0055] Based on the features screened by LASSO, perform multivariate COX regression analysis to further confirm which microbial features are independent factors affecting overall survival. The determination criterion for significant correlation is a p-value < 0.05. At the same time, calculate the regression coefficients of these microbial features;
[0056] Step 4. Construction of the MAPS model:
[0057] Multiply the abundance of each microbial feature significantly correlated with the overall survival (OS) by its risk coefficient obtained in the multivariate COX regression analysis, and sum these products to obtain the MAPS score for each patient;
[0058] Step 5. Determination of the optimal cut-off value:
[0059] Use the R language package "maxstat" to determine the optimal cut-off value, and divide all patients into high-risk groups and low-risk groups;
[0060] Step 6. Performance evaluation:
[0061] The Kaplan-Meier survival curve was used to evaluate the predictive ability of the MAPS score for overall survival (OS), and the ROC curve was used to evaluate the predictive performance of the MAPS model, including the AUC values for 1-year, 3-year, and 5-year survival;
[0062] Step 7. Clinical correlation analysis:
[0063] Combined with clinical characteristics, the correlation between the MAPS score and clinical characteristics was analyzed.
[0064] In univariate COX regression analysis, univariate COX regression analysis was performed using the statistical software R language. Before analysis, missing value processing and standardization preprocessing were performed on the collected circulating microbial feature data.
[0065] The benefits of the above settings are as follows:
[0066] 1. Benefits of using R language for univariate COX regression analysis
[0067] Powerful statistical functions
[0068] Rich function library: R language has many packages specifically for statistical analysis, such as the survival package, which contains the function coxph() required for COX regression analysis. It can conveniently and quickly implement univariate COX regression analysis and accurately calculate the relationship between each microbial feature and the overall survival of patients, including important statistical indicators such as hazard ratio (HR) and p-value.
[0069] Flexible data analysis: R language allows users to flexibly adjust the analysis process according to specific needs. For example, it is possible to easily analyze different subsets of microbial features or perform various transformations and screenings on the data to meet the requirements of different research questions.
[0070] Open source and free
[0071] Cost reduction: R language is open-source software, and users can use it for free without paying high software licensing fees. This greatly reduces the research cost for scientific research institutions and researchers, especially when large-scale data analysis is required, and the advantage is more obvious.
[0072] Extensive community support: Due to its open-source nature, R language has a large user community. Researchers can seek help in the community when encountering problems during use, and can also share their code and experience to promote academic communication and cooperation.
[0073] Good visualization effect
[0074] Intuitive display of results: R provides rich visualization tools, such as the ggplot2 package, which can display the results of univariate COX regression analysis in an intuitive chart form, such as a forest plot. These visual charts help researchers better understand the analysis results and discover patterns and trends in the data.
[0075] 2. Benefits of handling missing values
[0076] Improve data quality
[0077] Avoid data bias: The existence of missing values may lead to data bias and affect the accuracy of analysis results. Through reasonable missing value handling methods, such as deleting samples with fewer missing values, filling missing values with the mean or median, predicting missing values based on models, etc., the data can be made more complete and accurate, reducing errors caused by missing values.
[0078] Ensure the validity of analysis: Some statistical analysis methods have high requirements for data integrity. If missing values are not handled, the analysis may not proceed normally or incorrect results may be obtained. Handling missing values can ensure the validity and reliability of univariate COX regression analysis.
[0079] 3. Benefits of performing standard preprocessing
[0080] Eliminate the influence of dimensions
[0081] Fairly compare features: Different circulating microbial features may have different dimensions and value ranges. For example, the abundance of some microorganisms may be between 0 and 1, while the abundance of others may be in the dozens or even hundreds. Standardization can unify the value ranges of all features to a similar scale, making each feature have the same weight in the analysis, avoiding some features having too much influence on the analysis results due to different dimensions, and thus more fairly comparing the relationships between various microbial features and the overall survival of patients.
[0082] Accelerate the model convergence speed
[0083] Optimize computational efficiency: When performing univariate COX regression analysis, the standardized data can make the model converge faster. This is because the standardized data distribution is more concentrated and uniform, and the model is more likely to find the optimal solution during the iteration process, thus improving computational efficiency and reducing the time required for analysis.
[0084] In LASSO COX regression analysis and cross-validation, perform data normalization on the selected microbial features.
[0085] The benefits of the above settings are:
[0086] 1. Improve model performance
[0087] Enhance model stability
[0088] Different microbial characteristics may have different value ranges and dimensions. For example, the abundances of certain microorganisms may fluctuate within a very small numerical range, while the abundance values of other microorganisms are relatively large. Without normalization, the LASSOCOX regression model may overly focus on features with large value ranges during the learning process and ignore features with small value ranges, resulting in reduced stability and reliability of the model. Data normalization unifies the value ranges of all features to a similar scale, enabling the model to treat each feature more fairly, thereby enhancing the model's stability and reducing the impact caused by differences in feature dimensions.
[0089] Reduce the risk of overfitting
[0090] Normalization helps reduce the numerical differences between features, preventing the model from overly relying on certain features during training. In LASSOCOX regression, it already has a feature selection function that screens out important features by penalizing the coefficients. Data normalization can further optimize this process, enabling the model to select features more accurately and reasonably, avoiding the selection of irrelevant features that seem important only because of their large values, thus reducing the risk of overfitting and improving the model's generalization ability.
[0091] 2. Improve computational efficiency
[0092] Accelerate the convergence rate
[0093] When performing LASSOCOX regression analysis, the model usually needs to find the optimal coefficient solution through iteration. When the data is not normalized, the gradient magnitudes of different features may vary greatly, which can slow down the model's convergence rate during iteration and may even lead to getting stuck in a local optimal solution. Data normalization makes the gradients of all features fall within a similar magnitude, which can accelerate the model's convergence rate, reduce the number of iterations, thereby improving computational efficiency and saving computational time and resources.
[0094] 3. Facilitate feature comparison and interpretation
[0095] Fairly compare the importance of features
[0096] The normalized data makes each feature numerically comparable, enabling a more intuitive comparison of the impact of different microbial characteristics on the overall survival of patients. In the results obtained from LASSOCOX regression, the magnitude of the feature coefficients can to some extent reflect their importance. Since the data is normalized, the magnitude of the coefficients can more truly reflect the relative importance of the features without being interfered by the original dimensions of the features, facilitating researchers to accurately determine which microbial characteristics play a key role in prognostic assessment.
[0097] Enhanced result interpretability
[0098] For model users and clinicians, the normalized data and model results are easier to understand and interpret. Based on the standardized features and coefficients, they can more clearly understand the relationship between each microbial feature and the patient's prognosis, thus providing more targeted and operable suggestions for clinical decision-making.
[0099] In the multivariate COX regression analysis, the stepwise regression method was used to determine the microbial features finally included in the model.
[0100] The benefits of the above settings are as follows:
[0101] 1. Optimize the model structure
[0102] Screen important features
[0103] The stepwise regression method can screen out the features that have a significant impact on the patient's overall survival (OS) from numerous microbial features. In actual research, a large amount of microbial feature data may be collected, but not all features are closely related to the patient's prognosis. Through stepwise regression, features can be gradually introduced or excluded according to certain statistical criteria (such as p-value), and finally those features that contribute the most to the model and can best explain the patient's survival situation are retained, making the model structure more concise and reasonable.
[0104] Avoid redundant information
[0105] There may be a certain correlation between microbial features, and some features may contain part of the information of other features. The stepwise regression method can identify and exclude these redundant features, avoiding too much repeated or similar information in the model. This can reduce the complexity of the model, reduce the variance of the model, and improve the stability and prediction accuracy of the model.
[0106] 2. Improve the model performance
[0107] Enhance the prediction ability
[0108] By only including the features closely related to the patient's prognosis, the model constructed by the stepwise regression method can more accurately capture the relationship between microbial features and the patient's overall survival. Compared with the model including all features, the screened model can focus more on key factors, thus improving the model's prediction ability for the patient's prognosis and providing a more reliable basis for clinical decision-making.
[0109] Reduce the risk of overfitting
[0110] Too many features may cause the model to perform well on the training data but poorly on new data, i.e., the phenomenon of overfitting occurs. Stepwise regression reduces unnecessary parameters in the model by screening features, reduces the complexity of the model, and thus effectively reduces the risk of overfitting, enabling the model to have better generalization ability and maintain good performance on different data sets.
[0111] 3. Improve the interpretability of results
[0112] Identify key factors
[0113] The features included in the final model determined by stepwise regression are strictly screened, and the relationship between these features and the patient's prognosis is clearer. Clinicians and researchers can more easily understand and interpret the impact of these features on the patient's survival, and thus better apply the model results to clinical practice and develop personalized treatment plans for patients.
[0114] Simplify model interpretation
[0115] A simple model is easier to interpret and analyze than a complex one. The model obtained by stepwise regression only contains a few important microbial features, reducing the interference of irrelevant factors, enabling researchers to more clearly observe and analyze the relationship between each feature and the patient's prognosis, and providing convenience for further research and exploration.
[0116] 4. Meet the actual research needs
[0117] Effective utilization of resources
[0118] In actual research, the resources for data collection and analysis are limited. Stepwise regression can help researchers focus on the most important features with limited resources, avoid unnecessary analysis of a large number of irrelevant features, thereby improving research efficiency and saving time and costs.
[0119] Adapt to data characteristics
[0120] Microbial data usually has characteristics such as high dimensionality and complex correlations. Stepwise regression can dynamically select appropriate features to enter the model according to the actual situation of the data, adapt to the characteristics of microbial data, and better mine useful information in the data.
[0121] In the construction of the MAPS model, logarithmic transformation is performed on the abundance data of microbial features.
[0122] The benefits of the above settings are as follows:
[0123] 1. Improve data distribution
[0124] Approach normal distribution
[0125] The abundance data of microbial characteristics often exhibit a skewed distribution, that is, the data are concentrated in a certain interval while being less distributed in other intervals. This skewed distribution may affect the performance of some statistical methods and models because many statistical models assume that the data follow a normal distribution. Logarithmic transformation can convert the skewed distribution data into a form closer to the normal distribution, making the data distribution more symmetric and uniform. For example, the abundance of some microorganisms may have a few extremely high values. Logarithmic transformation can compress these high values while stretching the low values, making the data distribution more reasonable, thus meeting the requirements of the model for data distribution and improving the accuracy and reliability of the model.
[0126] Reduce variance heterogeneity
[0127] In the original microbial abundance data, the variances of different characteristics may vary greatly. This variance heterogeneity will affect the fitting effect of the model. Logarithmic transformation can, to a certain extent, reduce the variance heterogeneity, making the variances of different characteristics more stable and similar. This helps the model to be more fair when dealing with each characteristic, avoiding excessive influence on the model due to the too large variance of some characteristics, thereby improving the stability and generalization ability of the model.
[0128] 2. Enhance model stability
[0129] Reduce the influence of extreme values
[0130] There may be some extreme values in the microbial abundance data, which may be caused by experimental errors, sample contamination or other special reasons. These extreme values will have a greater interference on the training and prediction of the model, making the parameter estimation of the model inaccurate and even leading to overfitting of the model. Logarithmic transformation can effectively compress the influence of extreme values, convert extreme values into relatively small values, thereby reducing their interference on the model and making the model more robust.
[0131] Optimize parameter estimation
[0132] When the data are logarithmically transformed, the model will be more stable and accurate in parameter estimation. Logarithmic transformation can make the scale of the data more reasonable, avoiding the instability of model parameter estimation caused by too large differences in data scales. In the MAPS model, more accurate parameter estimation can improve the model's ability to predict the prognosis of patients and provide a more reliable basis for clinical decision-making.
[0133] 3. Facilitate feature relationship analysis
[0134] Linearize feature relationships
[0135] In some cases, there may be a non-linear relationship between the abundance of microbial features and the prognosis of patients. Logarithmic transformation can convert this non-linear relationship into an approximately linear relationship, making it easier for the model to capture the association between features and prognosis. Linear relationships are easier to understand and interpret, allowing researchers to more intuitively analyze the degree of influence of each microbial feature on the overall survival of patients, facilitating further research and clinical applications.
[0136] Enhance feature comparability
[0137] Logarithmic transformation can make the abundance data of different microbial features numerically more comparable. After transformation, the value ranges of each feature are closer, enabling the model to treat each feature more fairly and thus more accurately evaluate their importance in prognosis assessment.
[0138] 4. Outlier handling
[0139] Weaken the weight of outliers
[0140] Logarithmic transformation can reduce the weight of outliers in the model. Outliers may have large numerical values in the original data and have a significant impact on the model. Through logarithmic transformation, the numerical values of outliers are compressed, and their influence on the model is correspondingly reduced, enabling the model to focus more on the overall trend and features of the data rather than being dominated by individual outliers.
[0141] In determining the optimal cut-off value, the maximum log-rank statistic method in the "maxstat" package is used to find the optimal cut-off point.
[0142] The benefits of the above settings are as follows:
[0143] 1. Statistical validity
[0144] Maximize the difference between groups
[0145] The core objective of the maximum log-rank statistic method is to find a cut-off point such that the two groups (such as the high-risk group and the low-risk group) divided according to this cut-off point have the maximum difference in survival outcomes. The log-rank statistic is used to compare the differences between the survival curves of the two groups. The larger its value, the more obvious the separation of the two survival curves, that is, the more significant the difference in the survival conditions of the two groups of patients. By using this method, it can be ensured that the two groups divided have the maximum statistical discrimination, thus more accurately identifying patient groups with different prognoses and improving the model's predictive ability for patient prognosis.
[0146] Objective and scientific selection
[0147] This method is based on strict statistical principles. By evaluating and comparing all possible cut-off values and using the log-rank statistic as the criterion, the optimal cut-off point is objectively determined. This data-driven method avoids the interference of subjective factors, making the selection of the cut-off point more scientific and reliable. It can truly reflect the information contained in the data and provide a solid statistical basis for subsequent analysis and decision-making.
[0148] 2. Clinical practicability
[0149] Guiding clinical decision-making
[0150] In clinical practice, accurately differentiating high-risk and low-risk patients is crucial for formulating personalized treatment plans. The optimal cut-off point determined by the maximum log-rank statistic method can clearly divide patients into different risk groups, and doctors can formulate corresponding treatment strategies according to the risk groups to which patients belong. For example, for high-risk group patients, more aggressive treatment measures such as intensive chemotherapy and surgical intervention can be taken; while for low-risk group patients, relatively conservative treatment plans can be selected to avoid unnecessary risks and burdens brought by over-treatment.
[0151] Improving treatment effect
[0152] By reasonably dividing risk groups, treatment can be more precisely targeted at the actual situation of patients, improving the pertinence and effectiveness of treatment. Concentrating limited medical resources on patients who truly need them helps to improve the survival outcome of patients, enhance the overall treatment effect, and at the same time improve the utilization efficiency of medical resources.
[0153] 3. Intuitiveness of results
[0154] Visual display
[0155] The optimal cut-off point determined by the maximum log-rank statistic method can be intuitively displayed through survival curves. The Kaplan-Meier survival curve can clearly present the survival situations of high-risk and low-risk group patients. Doctors and researchers can intuitively observe the survival differences between the two groups, thus better understanding the significance and role of the cut-off point. This visual result helps clinical doctors quickly understand and apply the prediction results of the model, providing an intuitive basis for clinical decision-making.
[0156] Easy to interpret
[0157] The cut-off point determined by this method has a clear clinical significance and is easy to explain to patients and their families. Doctors can clearly explain to patients their risk status and corresponding treatment suggestions based on the survival curve and the cut-off point, improving patients' understanding and compliance with treatment plans, and promoting effective communication and cooperation between doctors and patients.
[0158] 4. Generalizability of the method
[0159] Applicable to multiple data types
[0160] The maximum log-rank statistic method in the "maxstat" package has strong versatility and is applicable to various types of continuous or ordinal categorical data. Whether it is microbial abundance data, biomarker concentration data, or other clinical indicator data, this method can be used to determine the optimal cut-off point, providing convenience for research and clinical applications in different fields.
[0161] Compatible with other analysis methods
[0162] This method can be combined with other statistical analysis methods and models to further expand its application scope. For example, the determined risk groups can be combined with multivariate analysis to explore the influence of other factors on the prognosis of patients; it can also be applied to validation studies to evaluate the performance of the model on different datasets, enhancing the reliability and universality of research results.
[0163] In performance evaluation, the Bootstrap resampling method is used to verify the results of Kaplan-Meier survival curves and ROC curves.
[0164] The benefits of the above settings are as follows:
[0165] 1. Enhance the reliability and stability of the results
[0166] Reduce the influence of sampling error
[0167] In actual research, the samples used are only a subset of the population, and the randomness of the samples may lead to certain biases in statistical results. The Bootstrap resampling method draws a large number of new samples (usually hundreds to thousands of times) from the original sample with replacement, and the size of each new sample is the same as the original sample. Based on these resampled samples, relevant indicators of Kaplan-Meier survival curves and ROC curves (such as survival probability, AUC value, etc.) are calculated respectively, and finally these results are integrated to evaluate the model performance. This can simulate the situation of multiple samplings, thereby reducing the error caused by single sampling and making the results more stable and reliable.
[0168] Provide more accurate confidence intervals
[0169] The distribution of various evaluation metrics (such as AUC value) can be obtained through the Bootstrap method, and then a more accurate confidence interval can be calculated. The confidence intervals calculated by traditional methods may be based on some assumptions (such as data following a normal distribution), but these assumptions do not necessarily hold in practical applications. The Bootstrap method does not rely on specific distribution assumptions and can estimate the confidence interval according to the characteristics of the data itself, more accurately reflecting the true fluctuation range of the evaluation metrics and providing a more precise basis for judging the model performance.
[0170] 2. Adapt to different data distributions and sample characteristics
[0171] Do not rely on specific distribution assumptions
[0172] Microbial data and clinical data often have complex distribution characteristics and may not meet the normal distribution or other specific distributions required by traditional statistical methods. The Bootstrap resampling method is a non-parametric method that does not require any assumptions about the distribution of the data and is applicable to various types of data distributions, including skewed distributions, multimodal distributions, etc. This makes the method highly adaptable when dealing with the complex and diverse data in practical research and can more accurately evaluate the performance of the model under different data conditions.
[0173] Handle small sample data
[0174] In some cases, due to difficulties in obtaining samples, etc., the research may only obtain a small sample size. Small sample data is prone to unstable and biased statistical results. The Bootstrap method increases the utilization efficiency of the samples through resampling, making up for the deficiencies of small sample data to a certain extent. It can mine more information from limited samples, enabling the reliable evaluation of the model performance even in the case of small samples and providing more valuable references for the research.
[0175] 3. Comprehensively evaluate the model performance
[0176] Evaluate the robustness of the model
[0177] Through multiple resamplings and calculations, the Bootstrap method can show the performance of the model under different sampling situations. If the model can maintain good performance in most resampled samples, it indicates that the model has strong robustness and can adapt to different data changes. On the contrary, if the model performance fluctuates greatly during the resampling process, it suggests that the model may have certain instability and needs further optimization or improvement.
[0178] Explore the variability of the results
[0179] This method can reveal the variability of the evaluation results, enabling researchers to understand the fluctuations of model performance metrics (such as AUC values) in different samples. This exploration of result variability helps to gain a deeper understanding of the model's performance and provides more comprehensive information for model improvement and application. For example, if a large fluctuation range of AUC values is found, it may be necessary to consider increasing the sample size or adjusting the model parameters to improve the model's stability.
[0180] 4. Facilitates combination with other methods
[0181] Integrates multiple evaluation metrics
[0182] The Bootstrap resampling method can be combined with other evaluation methods and metrics to conduct a more comprehensive evaluation of the model. For example, metrics such as sensitivity, specificity, and accuracy can be combined simultaneously, and the distributions of these metrics are calculated respectively during the resampling process, thereby evaluating the model's performance from multiple perspectives. This integrated evaluation method can provide richer information and helps to more accurately judge the advantages and disadvantages of the model.
[0183] Assists in model selection and optimization
[0184] When comparing different models or optimizing a model, the Bootstrap method can provide strong support for decision-making. By comparing the performance of different models on resampled samples, the optimal model can be selected more objectively; at the same time, based on the analysis of the resampling results, the weak links of the model can be identified, and the model can be optimized and improved targetedly to enhance the model's prediction ability and clinical application value.
[0185] In the clinical relevance analysis, Spearman rank correlation analysis is used to evaluate the correlation between the MAPS score and clinical characteristics.
[0186] The benefits of the above settings are as follows:
[0187] 1. Low requirements for data distribution
[0188] Does not rely on normal distribution
[0189] Many clinical data, such as clinical characteristics like age and tumor stage, as well as the MAPS score, may not follow a normal distribution. Spearman rank correlation analysis is a non-parametric statistical method that does not require the data to satisfy a specific distribution form and only considers the ranks of the data (i.e., the order after sorting the data from smallest to largest). This makes it highly adaptable in processing various types of clinical data. Whether the data is skewed, multimodal, or other irregular distributions, it can accurately evaluate the correlation between variables.
[0190] Robustly handles outliers
[0191] There may be outliers in clinical data, which may be caused by measurement errors, special cases, etc. Traditional parametric correlation analysis methods (such as Pearson correlation analysis) are sensitive to outliers, and outliers may seriously affect the calculation results of the correlation coefficient, leading to incorrect judgments about the relationship between variables. Spearman rank correlation analysis is calculated based on ranks, and the impact of outliers on ranks is relatively small. Therefore, it can handle data containing outliers more robustly and provide a more reliable correlation assessment.
[0192] 2. Capturing monotonic relationships
[0193] Identifying monotonic trends
[0194] Spearman rank correlation analysis mainly measures the monotonic relationship between two variables, that is, the value of one variable shows a consistent increasing (or decreasing) trend as the value of the other variable increases (or decreases). In clinical research, there may be such a monotonic association between the MAPS score and certain clinical characteristics. For example, as the tumor stage increases, the MAPS score may also show an upward trend. Through Spearman rank correlation analysis, this monotonic relationship can be effectively identified, providing valuable information for clinicians to help them understand the internal connection between the MAPS score and clinical characteristics.
[0195] Not restricted by linear relationships
[0196] Different from Pearson correlation analysis which can only detect linear relationships, Spearman rank correlation analysis can detect non-linear monotonic relationships between variables. In actual clinical situations, the relationship between the MAPS score and clinical characteristics may not be a strictly linear relationship, but there may be a monotonic change trend. Spearman rank correlation analysis can capture this broader relationship, providing a more comprehensive perspective for the study and helping to discover some potential clinical associations.
[0197] 3. Intuitive result interpretation
[0198] Simple and easy-to-understand correlation coefficient
[0199] The value range of the Spearman rank correlation coefficient is between -1 and 1. The closer its absolute value is to 1, the stronger the monotonic relationship between the two variables; a positive correlation coefficient indicates a positive correlation, that is, when the value of one variable increases, the value of the other variable also tends to increase; a negative correlation coefficient indicates a negative correlation, that is, when the value of one variable increases, the value of the other variable tends to decrease. This simple and intuitive result representation method enables clinicians and researchers to easily understand the strength and direction of the correlation between the MAPS score and clinical characteristics, facilitating the application of research results to clinical practice and decision-making.
[0200] Auxiliary clinical decision-making
[0201] The correlation results obtained through Spearman rank correlation analysis can help clinicians better understand the relationship between the MAPS score and clinical characteristics, thus making more scientific and reasonable decisions in formulating treatment plans, judging the prognosis of patients, etc. For example, if it is found that the MAPS score is highly positively correlated with the tumor stage, clinicians can more accurately evaluate the progress of the tumor based on the patient's MAPS score and then select a more appropriate treatment method.
[0202] 4. Wide range of applications
[0203] Processing ordinal categorical data
[0204] Ordinal categorical data are often included in clinical characteristics, such as the grade of tumors (well-differentiated, moderately-differentiated, poorly-differentiated), the performance status score of patients (ECOG score), etc. Spearman rank correlation analysis is very suitable for processing such ordinal categorical data. It can transform the ordinal categorical data into ranks and then calculate the correlation with the MAPS score. This method can make full use of the information contained in the ordinal categorical data and provide more accurate analysis results for clinical research.
[0205] Compatibility for multi-field applications
[0206] Spearman rank correlation analysis is not only applicable to the medical field but also has extensive applications in many other fields. Using this method in clinical correlation analysis facilitates comparison and communication with research results in other fields, promotes cooperation and knowledge sharing among different disciplines, and drives the development of medical research.
[0207] In clinical correlation analysis, clinical characteristics include age, M stage, N stage, and clinical stage.
[0208] The benefits of the above settings are as follows:
[0209] 1. Comprehensively evaluate the patient's condition
[0210] Reflect information at different levels
[0211] Age: Age is a basic clinical characteristic. There are differences in the physical functions, immunity, and disease tolerance of patients in different age groups. For example, elderly patients may have multiple underlying diseases and poor physical recovery ability, and may need to be more cautious in choosing treatment plans for tumor treatment; while young patients have relatively better physical functions and may have stronger tolerance to treatment.
[0212] M stage: distant metastasis, reflecting whether the tumor has spread to other organs in the body far away from the primary site, is an important indicator for judging the severity and prognosis of the tumor. If there is distant metastasis, it usually means that the disease is in the late stage, difficult to treat, and has a poor prognosis.
[0213] N stage: stands for regional lymph node metastasis. Knowing whether the tumor has invaded the local lymph nodes is crucial for determining the spread of the tumor and formulating treatment strategies. The extent of regional lymph node metastasis is closely related to the local progression and recurrence risk of the tumor.
[0214] Clinical staging: It comprehensively considers the size of the primary tumor (T staging), regional lymph node metastasis (N staging), and distant metastasis (M staging). It can comprehensively and systematically evaluate the overall progression of the tumor and provide an important reference for clinical treatment.
[0215] Understanding the condition from multiple dimensions
[0216] By analyzing these clinical features together, we can fully understand the patient's condition from multiple dimensions. A single feature may only reflect one aspect of the condition, while the combination of multiple features can provide richer and more comprehensive information, helping doctors to more accurately grasp the patient's disease status and develop a more reasonable treatment plan.
[0217] 2. Accurately determine the patient's prognosis
[0218] Predicting survival
[0219] These clinical characteristics are closely related to the patient's prognosis. Different combinations of age, M stage, N stage and clinical stage can predict the patient's survival time and risk of recurrence. For example, patients who are older, have distant metastasis (later M stage), extensive regional lymph node metastasis (later N stage) and advanced clinical stage often have a poor prognosis; while young patients with no distant metastasis, no regional lymph node metastasis and early clinical stage have a relatively good prognosis. By analyzing the correlation between the MAPS score and these clinical characteristics, the patient's prognosis can be judged more accurately, providing patients and their families with more accurate disease information.
[0220] Evaluating treatment effectiveness
[0221] Combining clinical characteristics and MAPS scores, the effect of treatment can also be evaluated. If the patient's clinical stage improves during treatment and the MAPS score decreases accordingly, it means that the treatment plan may be effective; conversely, if the clinical stage does not improve or even worsens, and the MAPS score increases, it indicates that the treatment effect is poor and the treatment plan needs to be adjusted in time.
[0222] 3. Guide the formulation of personalized treatment plans
[0223] Develop targeted strategies
[0224] Different combinations of clinical features correspond to different treatment needs and strategies. For example, for patients with older age, poor physical condition and distant metastasis, palliative treatment may be more suitable, mainly aiming to relieve symptoms and improve quality of life; while for young patients with good physical condition and early clinical stage, active surgical treatment or comprehensive treatment may be more suitable. By analyzing the correlation between the MAPS score and clinical features, a personalized treatment plan can be developed for each patient, improving the pertinence and effectiveness of treatment.
[0225] Optimize treatment decisions
[0226] Clinicians can, based on information such as the patient's age, M stage, N stage and clinical stage, combined with the MAPS score, weigh the pros and cons of various treatment methods and make more informed treatment decisions. For example, when deciding whether to perform surgery, radiotherapy, chemotherapy or targeted therapy, consider the patient's overall condition and prognostic risk, and select the most suitable treatment method for the patient to avoid over-treatment or under-treatment.
[0227] 4. Verify the effectiveness and practicality of the model
[0228] Evaluate model performance
[0229] Performing a correlation analysis between the MAPS score and these important clinical features can verify the effectiveness and practicality of the MAPS model. If the MAPS score has a good correlation with clinical features such as clinical stage, M stage, N stage, etc., it indicates that the model can reflect the severity of the patient's condition and prognosis, and has high clinical application value; conversely, if the correlation is poor, the model needs to be further optimized and improved.
[0230] Promote model application
[0231] Through clinical correlation analysis, demonstrating the close connection between the MAPS score and clinical features can enhance clinicians' trust and recognition of the model, and promote the wide application of the model in clinical practice. Clinicians can, based on the comprehensive results of the MAPS score and clinical features, better manage the treatment and follow-up of patients, improving the quality of medical care and the survival rate of patients.
[0232] Example 2
[0233] A method for constructing a prognostic scoring model for the circulating microbial abundance of cervical cancer patients, comprising the following steps:
[0234] Data collection
[0235] Three hundred cervical cancer patients admitted to a well-known tumor specialized hospital in recent years were selected as the research subjects. Detailed clinical data of each patient were collected, covering information such as the patient's age, reproductive history, HPV infection status, tumor size, FIGO stage (M stage, N stage, clinical stage), and overall survival (OS). At the same time, peripheral blood samples of the patients were collected, and advanced metagenomic sequencing technology was used to obtain circulating microbiome data, which can comprehensively and accurately detect the types and abundances of microorganisms in the samples.
[0236] Model construction steps
[0237] 1. Univariate COX regression analysis
[0238] Univariate COX regression analysis was performed on the obtained circulating microbial features using the survival package in R language. The significance level was set as p-value < 0.05 to screen out the microbial features significantly related to the overall survival of the patients. After careful analysis, it was found that 22 microbial features were significantly related to OS.
[0239] The specific R language is as follows:
[0240] library(survival)
[0241] # Assume data is a data frame containing microbial features and survival information
[0242] # time is the survival time, status is the survival status, and microbe_features is the column of microbial features
[0243] univ_cox<-lapply(microbe_features,function(x)
[0244] {coxph(Surv(time,status)~data[[x]],data=data)})
[0245] p_values<-sapply(univ_cox,function(x)
[0246] summary(x)$coefficients[5])
[0247] selected_features<-microbe_features[p_values<0.05]
[0248] 2. LASSO COX regression analysis and cross-validation
[0249] Perform LASSO COX regression analysis on the 22 selected microbial features and use the 10-fold cross-validation method to determine the optimal penalty parameter. In this process, the dataset is randomly divided into 10 subsets, and 9 of them are used as the training set and 1 as the test set in turn. The model is trained and validated multiple times to ensure that the selected features have good generalization ability. Finally, 12 most predictive microbial features are selected.
[0250] The specific R code is as follows:
[0251] library(glmnet)
[0252] # Extract the screened microbial feature data
[0253] X<-as.matrix(data[,selected_features])
[0254] y<-Surv(data$time,data$status)
[0255] # Perform 10-fold cross-validation LASSO COX regression
[0256] cv_fit<-cv.glmnet(X,y,family="cox",nfolds=10)
[0257] # Select the best lambda value
[0258] best_lambda<-cv_fit$lambda.min
[0259] # Perform LASSO regression based on the best lambda value
[0260] lasso_fit<-glmnet(X,y,family="cox",lambda=best_lambda)
[0261] # Extract the features with non-zero coefficients
[0262] final_features<-selected_features[which(coef(lasso_fit)!=0)]
[0263] 3. Multivariate COX regression analysis
[0264] Perform a multivariable COX regression analysis on the 12 features selected based on LASSO to further determine which microbial features are independent factors affecting overall survival, and calculate the regression coefficients of these microbial features. This step can accurately evaluate the independent impact of each microbial feature on the overall survival of patients while controlling other factors.
[0265] The specific R code is as follows:
[0266] # Extract the data of the finally selected features
[0267] X_final<-as.matrix(data[,final_features])
[0268] multi_cox<-coxph(Surv(time,status)~X_final,data=data)
[0269] # Extract the regression coefficients
[0270] coefficients<-coef(multi_cox)
[0271] 4. Construction of the MAPS model
[0272] Based on the regression coefficients obtained from the multivariable COX regression analysis, calculate the MAPS score for each patient. The specific method is to multiply the abundance of each microbial feature significantly related to OS by its corresponding risk coefficient, and then sum these products to obtain the MAPS score for each patient.
[0273] The specific R code is as follows:
[0274] # Calculate the MAPS score for each patient
[0275] maps_scores<-apply(data[,final_features],1,function(x){sum(x*coefficients)})
[0276] 5. Determination of the optimal cut-off value
[0277] Use the R package "maxstat" to determine the optimal cut-off value and divide all patients into high-risk and low-risk groups. The "maxstat" package tries all possible cut-off values and calculates the log-rank statistic corresponding to each cut-off value, and selects the cut-off value that maximizes the log-rank statistic as the optimal cut-off point.
[0278] The specific R code is as follows:
[0279] library(maxstat)
[0280] cutoff <- maxstat.test(Surv(data$time, data$status) ~ maps_scores, data = data)$estimate
[0281] risk_groups <- ifelse(maps_scores > cutoff, "High Risk", "Low Risk")
[0282] 6. Performance Evaluation
[0283] 1). Use the Kaplan-Meier survival curve to evaluate the predictive ability of the MAPS score for overall survival. The plotted survival curve intuitively shows the difference in survival between high-risk and low-risk groups of patients. The results show that the survival curves of the two groups are significantly separated, indicating that the MAPS score can effectively distinguish patients with different prognoses.
[0284] The specific R code is as follows:
[0285] library(survminer)
[0286] surv_obj <- Surv(data$time, data$status)
[0287] km_fit <- survfit(surv_obj ~ risk_groups, data = data)
[0288] ggsurvplot(km_fit, pval = TRUE)
[0289] 2). Use the ROC curve to evaluate the predictive performance of the MAPS model, and calculate the AUC values for 1-year, 3-year, and 5-year survival respectively. The closer the AUC value is to 1, the better the predictive performance of the model. The calculation results show that the AUC values for 1-year, 3-year, and 5-year survival are 0.81, 0.79, and 0.77 respectively, indicating that the model has good predictive ability.
[0290] The specific R code is as follows:
[0291] library(survivalROC)
[0292] # Calculate the AUC values for 1-year, 3-year, and 5-year
[0293] roc_1yr <- survivalROC(Stime = data$time, status = data$status, marker = maps_scores, predict.time = 1, method = "KM")
[0294] roc_3yr <- survivalROC(Stime = data$time, status = data$status, marker = maps_scores, predict.time = 3, method = "KM")
[0295] roc_5yr <- survivalROC(Stime = data$time, status = data$status, marker = maps_scores, predict.time = 5, method = "KM")
[0296] auc_1yr <- roc_1yr$AUC
[0297] auc_3yr <- roc_3yr$AUC
[0298] auc_5yr <- roc_5yr$AUC
[0299] 7. Clinical correlation analysis
[0300] Combined with the clinical characteristics of the patients, such as age, HPV infection status, FIGO stage, etc., the correlation analysis between MAPS score and clinical characteristics was carried out. The results showed that the MAPS score was positively correlated with the FIGO stage of the patients and also significantly correlated with the high-risk HPV infection status. This indicates that the MAPS score can not only reflect the relationship between microbial characteristics and prognosis, but also be closely related to commonly used clinical indicators, further verifying the clinical rationality and practicality of the model.
[0301] Specifically in R language:
[0302] # Assume clinical_features is the column of clinical characteristics
[0303] correlation_matrix <- cor(data[, c("maps_scores", clinical_features)])
[0304] The circulating microbial abundance prognostic scoring model constructed through the above steps can provide strong support for the prognostic evaluation of cervical cancer patients and help clinicians formulate more personalized and precise treatment plans.
[0305] Example 3
[0306] A method for constructing a prognostic scoring model of circulating microbial abundance in colorectal cancer patients, comprising the following steps:
[0307] Data collection
[0308] Collect the clinical information and circulating microbiome data of 150 colorectal cancer patients in a large general hospital. The clinical information covers the basic situation, tumor stage, treatment history, and overall survival of the patients; the circulating microbiome data is obtained through metagenomic sequencing technology.
[0309] Model construction steps
[0310] The model construction steps are basically the same as those in Example 2, specifically as follows:
[0311] Univariate COX regression analysis: Screen out 20 microbial features significantly related to OS.
[0312] LASSO COX regression analysis and cross-validation: After 10-fold cross-validation, screen out 10 microbial features with the most predictive ability.
[0313] Multivariate COX regression analysis: Determine the factors independently affecting the overall survival and calculate the regression coefficients.
[0314] MAPS model construction: Calculate the MAPS score for each patient.
[0315] Determination of the optimal cut-off value: Use the "maxstat" package to determine the optimal cut-off value and divide the high-risk group and the low-risk group.
[0316] Performance evaluation
[0317] The Kaplan-Meier survival curve shows a significant difference in the survival of patients in the high-risk group and the low-risk group.
[0318] The AUC values for 1-year, 3-year, and 5-year survival are 0.78, 0.76, and 0.74, respectively.
[0319] Clinical correlation analysis: It is found that the MAPS score is significantly correlated with the tumor differentiation degree and lymph node metastasis of the patients.
[0320] The examples given in the present invention are not intended to limit the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here, and the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. A method for constructing a prognostic scoring model for circulating microbial abundance, characterized in that: The following steps are involved: Step 1: Univariate COX regression analysis: Univariate COX regression analysis was used to screen out circulating microbial features that were significantly associated with the patient's overall survival, with a significant correlation determined by a p value of < 0.
05. Step 2: LASSOCOX regression analysis and cross validation: LASSOCOX regression analysis was performed on the selected microbial features, and 10-fold cross-validation was used to further narrow the range of microbiome features to ensure that the most predictive features were selected; Step 3: Multivariate COX regression analysis: Multivariate COX regression analysis was performed based on the characteristics screened by LASSO to further confirm which microbial characteristics were independent factors affecting overall survival. The criterion for significant correlation was p value < 0.
05. At the same time, the regression coefficients of these microbial characteristics were calculated. Step 4: MAPS model construction: The abundance of each microbial signature significantly associated with overall survival was multiplied by its risk coefficient obtained in multivariate COX regression analysis, and these products were summed to obtain the MAPS score for each patient; Step 5: Determination of the optimal cutoff value: The R language package "maxstat" was used to determine the optimal cutoff value and divide all patients into high-risk and low-risk groups; Step 6: Performance evaluation: The Kaplan-Meier survival curve was used to evaluate the predictive ability of the MAPS score for overall survival, and the ROC curve was used to evaluate the predictive performance of the MAPS model, including the AUC values for 1-, 3-, and 5-year survival; Step 7: Clinical relevance analysis: Combined with clinical characteristics, the correlation analysis between MAPS score and clinical characteristics was performed.
2. The method for constructing a circulating microbial abundance prognostic scoring model according to claim 1, characterized in that: In the step 1, the univariate COX regression analysis is performed using the statistical software R language, and the collected circulating microbial characteristic data are processed for missing values and standardized before analysis.
3. The method for constructing a circulating microbial abundance prognostic scoring model according to claim 1, characterized in that: In the step 2, data normalization is performed on the screened microbial features.
4. The method for constructing a circulating microbial abundance prognostic scoring model according to claim 1, characterized in that: In step three, a stepwise regression method is used to determine the microbial characteristics that ultimately enter the model.
5. The method for constructing a circulating microbial abundance prognostic scoring model according to claim 1, characterized in that: In the step 4, the abundance data of the microbial characteristics are logarithmically transformed.
6. The method for constructing a circulating microbial abundance prognostic scoring model according to claim 1, characterized in that: In step 5, the maximum log-rank statistics method in the "maxstat" package is used to find the optimal cutoff point.
7. The method for constructing a circulating microbial abundance prognostic scoring model according to claim 1, characterized in that: In step six, the results of the Kaplan-Meier survival curve and the ROC curve are verified using the Bootstrap resampling method.
8. The method for constructing a circulating microbial abundance prognostic scoring model according to claim 1, characterized in that: In step seven, Spearman rank correlation analysis was used to evaluate the correlation between MAPS scores and clinical characteristics.
9. The method for constructing a circulating microbial abundance prognostic scoring model according to claim 8, characterized in that: In step seven, the clinical characteristics include age, M stage, N stage, and clinical stage.