Biomarker combination for prognosis evaluation of acute respiratory distress syndrome and application of biomarker combination
By combining chromosome structure maintenance protein 4, indole derivatives and other biomarkers, a prognostic assessment model was constructed, which solved the problem of inaccurate ARDS prediction in existing technologies, achieved efficient risk assessment, and improved the ability to provide individualized treatment for ARDS patients.
Patent Information
- Application Number
- CN202511885710.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-12-15
AI Technical Summary
Existing technologies are insufficient to accurately predict the clinical outcomes of patients with acute respiratory distress syndrome (ARDS). Existing diagnostic criteria and scoring systems cannot meet the needs of individualized treatment, and existing biomarker prediction models have failed to effectively integrate multidimensional biological information.
Using chromosome structure maintenance protein 4 (SMC4) or indole derivatives such as indoleacetic acid (IAA) and 3-formylindole (3-FI), Acinetobacter baumannii, pulmonary surfactant protein A (SFTPA1 and SFTPA2), tryptophan metabolism pathway-related genes, and other biomarkers, a prognostic assessment model was constructed. Combined with machine learning algorithms such as random forest, logistic regression, and neural networks, the 28-day mortality risk of ARDS patients after intervention was predicted by detecting the expression level or abundance of biomarkers.
It improved the sensitivity and specificity of ARDS prediction, increasing the AUC value from 0.7226 to 0.984, providing more accurate risk assessment and having significant clinical translational value.
Smart Images

Figure CN121577903A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of biomedical technology and precision medicine technology, specifically to a combination of biomarkers for prognostic assessment of acute respiratory distress syndrome and their application. Background Technology
[0002] Acute respiratory distress syndrome (ARDS) is a critical illness in intensive care medicine, characterized by high heterogeneity in clinical presentation and biological features, resulting in a persistently high mortality rate. Current clinical diagnostic criteria and severity scoring systems, such as APACHE II and SOFA scores, are insufficient to accurately predict individual patient outcomes, failing to meet the clinical need for precise risk stratification and individualized treatment. In recent years, molecularly based subtyping studies (such as "high inflammation" and "low inflammation" subtypes) have shown potential in predicting prognosis and treatment response, highlighting the urgent need to develop novel molecular diagnostic tools.
[0003] Even though existing technologies already use biomarkers for prognostic assessment of ARDS, such as patent document (CN113252896A) which discloses the use of ND1 content in bronchoalveolar lavage fluid to assess ARDS prognosis, the AUC value predicted solely by the content of this protein is only 0.7226. Moreover, existing technologies do not have analytical and predictive models that can integrate multi-dimensional biological information (including microbial function, metabolites, and host response). Summary of the Invention
[0004] In a first aspect, the present invention provides a biomarker for prognostic assessment of acute respiratory distress syndrome (ARDS), said biomarker comprising one or both of chromosome structure maintenance protein 4 or indole derivatives.
[0005] Preferably, the indole derivative includes indole-3-aceticacid (IAA) and / or 3-formylindole (3-FI).
[0006] Preferably, the biomarkers also include Acinetobacter baumannii and / or pulmonary surfactant protein A.
[0007] Preferably, the pulmonary surfactant protein A includes SFTPA1 and / or SFTPA2.
[0008] Preferably, the biomarkers also include genes related to the tryptophan metabolism pathway.
[0009] In one specific embodiment of the present invention, the tryptophan metabolism pathway-related genes include one or more of the following KEGG database numbers: K00128, K00382, K01426, K01556, K01825, K03781, K03782, or K21801.
[0010] in, K00128 is aldehyde dehydrogenase (NAD + ), indicating aldehyde dehydrogenase (NAD) + ) or its encoding gene.
[0011] K00382 stands for dihydrolipoyl dehydrogenase, which represents dihydrolipoamide dehydrogenase or its encoding gene.
[0012] K01426 stands for amidase, indicating an amidase or its encoding gene.
[0013] K01556 stands for kynureninase, representing kynureninase or its encoding gene.
[0014] K01825 stands for 3-hydroxyacyl-CoA dehydrogenase, representing 3-hydroxyacyl-CoA dehydrogenase or its encoding gene.
[0015] K03781 stands for catalase, representing catalase or its encoding gene.
[0016] K03782 stands for catalase-peroxidase, indicating that it is a hydrogen peroxide-peroxidase or its encoding gene.
[0017] K21801 stands for indoleacetamide hydrolase, representing indoleacetamide hydrolase or its encoding gene.
[0018] Preferably, the biomarkers further include one or more of C-reactive protein, monocyte percentage, or neutrophil percentage.
[0019] In one specific embodiment of the present invention, the biomarker is chromosome structure maintenance protein 4.
[0020] In one specific embodiment of the present invention, the biomarker is an indole derivative, which is indole-3-aceticacid (IAA) and 3-formylindole (3-FI).
[0021] In one specific embodiment of the present invention, the biomarkers are Acinetobacter baumannii, indoleacetic acid, 3-formylindole, SFTPA1, SFTPA2, chromosome structure maintenance protein 4, and enzymes or their encoding genes with KEGG database numbers K00128, K00382, K01426, K01556, K01825, K03781, K03782, and K21801.
[0022] In one specific embodiment of the present invention, the biomarkers are Acinetobacter baumannii, indoleacetic acid, 3-formylindole, SFTPA1, SFTPA2, chromosome structure maintenance protein 4, enzymes or their encoding genes with KEGG database numbers K00128, K00382, K01426, K01556, K01825, K03781, K03782 and K21801, C-reactive protein, monocyte percentage and neutrophil percentage.
[0023] Preferably, the biomarker is a biomarker in cells, tissues, organs, or body fluids.
[0024] Preferably, the body fluid includes blood, plasma, or bronchoalveolar lavage fluid.
[0025] In a second aspect, the present invention provides a chip, reagent kit, test strip, membrane strip, or device for assessing the prognosis of ARDS, wherein the chip, reagent kit, test strip, membrane strip, or device includes reagents for detecting the aforementioned biomarkers.
[0026] Preferably, the chip, reagent kit, test strip, membrane strip, or device further includes reagents for processing samples.
[0027] Preferably, the sample includes cells, tissues, organs, or body fluids.
[0028] A third aspect of the present invention provides a method for constructing a prognostic assessment model for ARDS, the method comprising: 1) Collect samples from the ARDS death group and the ARDS survival group, and detect the expression level or abundance of the biomarkers mentioned in the first aspect; 2) Construct a prognostic assessment model based on the information collected in step 1); Alternatively, the construction method may include: I) Collect ARDS patient samples and detect the expression levels or abundance of the biomarkers described in the first aspect; II) Based on the patients' clinical outcomes, ARDS patients were divided into an ARDS death group and an ARDS survival group; III) Construct a prognostic assessment model based on the test results of I) and the clinical diagnosis results of II).
[0029] Preferably, the biomarker is a biomarker in cells, tissues, organs, or body fluids.
[0030] Preferably, the body fluid includes blood, plasma, or bronchoalveolar lavage fluid.
[0031] Preferably, the machine learning model used to construct the prognostic assessment model includes one or more of Random Forest, Logistic Regression, Neural Network, or Naïve Bayes.
[0032] Preferably, the prognostic assessment model is used to predict the risk of death 28 days after a patient receives interventional treatment.
[0033] Preferably, the intervention includes mechanical ventilation.
[0034] Preferably, the prognostic assessment includes obtaining a prognostic risk score.
[0035] The prognostic risk score is obtained using the following formula: Risk-score = (1.88 × X_Ab) + (1.65 × X_BALF_IAA) + (0.95 × X_Plasma_IAA) + (0.82 × X_SFTPA1)- (1.50 × X_SFTPA2) +(0.92 × X_Neu) - (0.85 × X_Mono) - (0.42 × Here, BALF represents bronchoalveolar lavage fluid, and Plasma represents blood plasma. For example, "X_BALF_3FI" represents the standardized value (X) of 3-formylindole in bronchoalveolar lavage fluid.
[0036] X_Ab represents the standardized value (X) of Acinetobacter baumannii derived from bronchoalveolar lavage fluid.
[0037] X_SMC4 represents the standardized value (X) of SMC4 derived from plasma.
[0038] X_SFTPA1 and X_SFTPA2 represent the standardized values (X) of SFTPA1 and SFTPA2 derived from bronchoalveolar lavage fluid.
[0039] X_CRP represents the standardized value (X) of CRP derived from plasma.
[0040] X_Neu represents the standardized value (X) of the proportion of neutrophils derived from plasma.
[0041] X_Mono represents the standardized value (X) of a monocyte derived from plasma.
[0042] Score_Trp represents the score of KO genes in the tryptophan metabolism pathway derived from bronchoalveolar lavage fluid, calculated as follows: Score_Trp = 0.15 × (K21801 + K03782) + 0.12 × (K01825 + K01426+K00382+K00128) + 0.10 × (- K01556 + K03781).
[0043] In a fourth aspect, the present invention provides a prognostic evaluation model obtained by the above-described construction method.
[0044] A fifth aspect of the present invention provides a system for assessing the prognosis of ARDS, the system comprising: A data detection device, unit, or module for detecting the abundance or expression level of the biomarkers described in the first aspect in a sample; A data input device, unit, or module for inputting the abundance or expression level data of the biomarker; Data analysis devices, units, or modules that calculate the 28-day mortality risk of ARDS patients after interventional treatment based on biomarker abundance or expression level data; A data output device, unit, or module for outputting analysis results of the 28-day mortality risk of ARDS patients after interventional treatment.
[0045] The implementation of the system includes the manual and / or automatic execution or completion of selected tasks.
[0046] The system can perform tasks via software and / or hardware.
[0047] This application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0048] The data analysis device, unit, or module includes the prognostic evaluation model obtained by the above construction method.
[0049] A sixth aspect of the present invention provides an apparatus comprising the system.
[0050] In a seventh aspect, the present invention provides the use of a biomarker, the chip, a reagent kit, a test strip, a membrane strip or device, the prognostic assessment model, the system or apparatus in the preparation of products for prognostic assessment of ARDS. Preferably, the biomarker is a biomarker in cells, tissues, organs, or body fluids.
[0051] In one specific embodiment of the present invention, the body fluid includes blood, plasma, or bronchoalveolar lavage fluid.
[0052] Preferably, among the biomarkers, 1) Acinetobacter baumannii is Acinetobacter baumannii derived from bronchoalveolar lavage fluid; and / or, 2) 3-Formylindole and / or indoleacetic acid are 3-formylindole and / or indoleacetic acid derived from bronchoalveolar lavage fluid and / or plasma; and / or, 3) SFTPA1 and / or SFTPA2 are derived from bronchoalveolar lavage fluid; 4) Chromosomal structure maintenance protein 4 is derived from plasma; and / or, 5) Genes related to the tryptophan metabolism pathway are derived from bronchoalveolar lavage fluid; and / or, 6) The proportion of C-reactive protein, monocytes, or neutrophils in plasma refers to the proportion of C-reactive protein, monocytes, or neutrophils in plasma.
[0053] Preferably, ARDS prognostic assessment includes predicting the risk of death 28 days after the patient receives intervention.
[0054] Preferably, the intervention includes mechanical ventilation.
[0055] Preferably, prognostic assessment of ARDS includes detecting the expression level or abundance of biomarkers.
[0056] Preferably, the methods for detecting the expression level or abundance of biomarkers include one or more of whole genome sequencing, chemiluminescence, immunofluorescence, colorimetric assay, ELISA, protein chip, liquid chromatography or mass spectrometry.
[0057] Preferably, the ARDS prognostic assessment further includes comparing the expression level or abundance of detected biomarkers with a threshold.
[0058] The threshold mentioned above was obtained through previous experiments, that is, the threshold was determined by the difference in the expression level or abundance of biomarkers between patients who died after having ARDS and patients who survived after having ARDS through experiments and data analysis.
[0059] When the expression level or abundance of a detected biomarker differs or is significantly different from a threshold, for example, 1) When the abundance of Acinetobacter baumannii is higher than or significantly higher than the threshold, it is predicted to be a high risk of death; 2) When the abundance of indoleacetic acid is higher than or significantly higher than the threshold, it is predicted to be a high risk of death; 3) When the abundance of 3-formylindole is higher than or significantly higher than the threshold, it is predicted to indicate a high risk of death; 4) When the expression level of SFTPA1 is below or significantly below the threshold, it is predicted to indicate a high risk of death; 5) When the expression level of SFTPA2 is below or significantly below the threshold, it is predicted to indicate a high risk of death; 6) When the expression level of chromosome structure maintenance protein 4 is higher than or significantly higher than the threshold, it is predicted to be a high risk of death; 7) When the expression levels of one or more of the genes related to the tryptophan metabolism pathway (K00128, K00382, K01426, K01825, K03781, K03782 or K21801) are higher than or significantly higher than the threshold, and / or the expression level of K01556 is lower than or significantly lower than the threshold, a high risk of death is predicted. 8) When the expression level of C-reactive protein is below or significantly below the threshold, it is predicted to be a high risk of death; 9) When the proportion (abundance) of monocytes is below or significantly below the threshold, it is predicted to be a high risk of death; 10) When the proportion (abundance) of neutrophils is higher than or significantly higher than the threshold, it is predicted to be a high risk of death.
[0060] Preferably, the product includes one or more of the following: chip, reagent kit, test strip, membrane strip, or device.
[0061] An eighth aspect of the present invention provides a method for prognostic assessment of ARDS, the method comprising detecting the expression level or abundance of the biomarker described in the first aspect, or the method comprising performing an assessment using the aforementioned system, apparatus or prognostic assessment model.
[0062] Preferably, the method includes the following steps: a) Obtain patient biological samples, including bronchoalveolar lavage fluid (BALF) and / or plasma; b) Detect the expression level or abundance of at least one of the biomarkers described in the first aspect of the present invention in the biological sample; c) Based on the test results of step b), assess the risk of the patient having an adverse outcome (such as death within 28 days).
[0063] The "patient" described in this invention can be a human or a non-human mammal. The non-human mammal can be a wild animal, a zoo animal, an economic animal, a pet, a laboratory animal, etc. Preferably, the non-human mammal includes, but is not limited to, pigs, cattle, sheep, horses, donkeys, foxes, jackals, minks, camels, dogs, cats, rabbits, mice (e.g., rats, mice, guinea pigs, hamsters, gerbils, chinchillas, squirrels) or monkeys, etc.
[0064] The term "and / or" as used in this invention includes all combinations of connected items, and should be considered as each combination having been individually listed in this application. For example, "A and / or B" includes "A", "B", and "A and B", and "A, B and / or C" includes "A", "B", "C", "A and B", "A and C", "B and C", and "A and B and C".
[0065] The use of "comprising" or "including" in this invention is an open-ended description, encompassing the specified components or steps that are connected, as well as other specified components or steps that do not substantially affect the meaning.
[0066] The "method" or "application" described in this invention may be for diagnostic purposes or for non-diagnostic purposes.
[0067] The beneficial effects of this invention are: 1. This invention identified multiple differential biomarkers in ARDS death and survival groups and verified that SMC4 or indole derivatives alone have high sensitivity, specificity, and AUC values when used to predict ARDS mortality risk. For example, the AUC value for SMC4 as a predictor can reach up to 0.752, and when indoleacetic acid and 3-formylindole are used as biomarkers, the predicted AUC value can reach 0.885. 2. Innovative biological axis: This invention reveals and verifies for the first time the crucial role of a core pathogenic axis—"microorganism (Acinetobacter baumannii) - metabolite (indole derivative) - host protein (SFTPA1, SFTPA2, SMC4)"—in ARDS prognosis, providing a new perspective for understanding the disease mechanism. Using the biological axis of this application as a biomarker to predict ARDS mortality risk, the AUC value can reach 0.921. Further combining it with clinical indicators (C-reactive protein, monocyte percentage, and neutrophil percentage) further improves the AUC value to 0.984, which greatly enhances predictive performance and has significant clinical translational value. Attached Figure Description
[0068] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings, wherein: Figure 1 : Figure 1Figure a shows a comparison of the relative abundance of Acinetobacter baumannii in BALF from patients in the death and survival groups of ARDS. Figure 1 Figure b shows the genus-level statistical results of microorganisms in the Death and Survival groups of ARDS patients.
[0069] Figure 2 Abundance levels of indoleacetic acid (IAA) and 3-formylindole (3-FI) in samples (BALF or plasma samples) from ARDS patients in the death and survival groups.
[0070] Figure 3 Expression levels of SFTPA1, SFTPA2, and SMC4 in ARDS patients in the death and survival groups.
[0071] Figure 4 Comparison of C-reactive protein, monocyte ratio (%), and neutrophil ratio (%) in ARDS patients from the death and survival groups.
[0072] Figure 5 : Figure 5 Figure a shows the results of KEGG pathway enrichment analysis; Figure 5 Figure b shows the heatmap analysis results of genes related to the tryptophan metabolism pathway. In the figure, * indicates p < 0.05, ** indicates p < 0.01, *** indicates p < 0.001, and **** indicates p < 0.0001.
[0073] Figure 6 The diagram shows the "microbe-indole-host" interaction network based on multi-omics data. Pink lines represent positive correlations, and light blue lines represent negative correlations. Only data with a correlation coefficient R greater than 0.2 from Pearson correlation analysis were used for plotting.
[0074] Figure 7 : ROC curves of the predictive performance of biomarker-based machine learning models on the validation set. Here, RF represents Random Forest, LR represents Logistic Regression, NN represents Neural Network, NB represents Naive Bayes, CV represents Hierarchical Five-Fold Cross-Validation, and LOOCV represents Leave-One-Out-of-One Count (LOOCV). Taking RF_CV as an example, it represents the ROC curve plotted using the Random Forest machine learning method with hierarchical five-fold cross-validation. Figure 7Figure a shows the ROC curve of the predictive performance of the biomarker group (4). Figure 7 Figure b shows the ROC curve of the predictive performance of the biomarker group (3). Figure 7 Figure c shows the ROC curve of the predictive performance of the biomarker group (2). Figure 7 Figure d shows the ROC curve of the predictive performance of the biomarker group (1). Specific implementation manners
[0075] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0076] The samples, experimental methods and data analysis involved in the embodiments are as follows: 1. Research cohort and sample collection This study is a multi - center, prospective observational study, including 47 ARDS patients receiving mechanical ventilation from six medical centers. The research protocol of this study has been approved by the ethics committees of each participating center. The participating centers of this study involving human participants include the Ethics Committee of the Eighth Medical Center of the Chinese PLA General Hospital and the Ethics Committee of the First Medical Center of the Chinese PLA General Hospital (Ethical Approval Number: 2022113030901836), Beijing Anzhen Hospital Affiliated to Capital Medical University (Number: KS2024031), the First Affiliated Hospital of Zhengzhou University (Number: 2024 - KY - 0156 - 01), Zhengzhou Central Hospital (formerly: Zhengzhou Fourth People's Hospital) (Number: ZXYY202428), and Henan Provincial People's Hospital (Number: 2024 Ethics Review No. 71).
[0077] The inclusion criteria for ARDS patients include: meeting the ARDS diagnostic criteria defined in Berlin in 2012, aged ≥ 18 years, and being moderate - to - severe ARDS patients receiving endotracheal intubation and mechanical ventilation. The exclusion criteria include: patients who did not consent to participate; patients with intracranial hypertension, acute coronary syndrome, pleural fistula or pneumothorax, subcutaneous emphysema; pregnant and lactating patients; and patients with severe hemodynamic instability (vasopressor increase > 30% in the first 6 hours, or norepinephrine > 0.5 μg / kg / min).
[0078] Based on the patients' survival outcomes at 28 days post-enrollment, they were divided into a death group and a survival group. On the day of enrollment, bronchoalveolar lavage fluid (BALF) and EDTA-anticoagulated whole blood were collected simultaneously. The whole blood was centrifuged to separate plasma. The BALF concentrate was used for metagenomic sequencing, while the BALF supernatant and plasma samples were stored at -80°C for subsequent metabolomics and proteomics analysis. There were no significant differences between the two groups in key baseline characteristics such as age, sex, and disease severity scores (APACHE II, GCS, SOFA) (see Table 1), making them comparable.
[0079] Table 1 Baseline characteristics of ARDS patients included in the study
[0080] Note: Normally distributed measurement items are expressed as mean ± standard deviation; non-normally distributed items are expressed as median (25th–75th percentile); count items are expressed as number of cases (%). Comparisons between groups were performed using the Chi-square test (categorical items), independent samples t-test (normally distributed), or Mann-Whitney U test (non-normally distributed).
[0081] The two groups of patients were randomly divided into two cohorts (Table 2).
[0082] Table 2
[0083] 2. Detection methods for biomarkers (1) Detection of relative abundance of Acinetobacter baumannii (metagenomics) On the day of enrollment, BALF stock solution was collected, and the library was constructed according to the manufacturer's procedures and sequenced on the Illumina PE150 platform. After quality control of the raw sequences, adapters and low-quality sequences were removed, and host sequences were removed by alignment with the human reference genome (GRCh38 / hg38) using Bowtie2. Non-host reads were assembled using MEGAHIT, and ORFs were predicted using Meta-Gene Mark. Redundancy was removed across samples to construct a non-redundant gene set. Reads were back-linked for quantification and their length and library size were normalized (relative abundance / TPM). Constitutive data were transformed using CLR. DIAMOND was aligned with MicroNR and relative abundance of species was obtained using LCA, and the data were simultaneously annotated to KEGG. Orthology (KO) was used to summarize KO abundance; to reduce noise, low abundance features with relative abundance <0.1% were removed before downstream analysis, and only samples with grouping information were retained; α diversity (Chao1, Shannon, Simpson) was calculated from the species table, β diversity was based on Bray-Curtis and presented using PCoA, PERMANOVA (999 permutations) was used to test differences between groups and betadisper was used to assess dispersion; differential species / KO were corrected using two-sided Wilcoxon and Benjamini-Hochberg methods, and effect sizes and q values were reported.
[0084] (2) Detection of indoleacetic acid (IAA) and 3-formylindole (3-FI) levels (metabolic genomics) Plasma and bronchoalveolar lavage fluid samples were analyzed by the CalOmics non-targeted metabolomics platform at the Hangzhou Calibra laboratory. Samples were extracted with methanol (1:4), centrifuged, dried under nitrogen, and then reconstituted. Metabolite analysis was performed using an ACQUITY2DUPLC ultra-high performance liquid chromatography system (Waters, Milford, MA, USA) coupled with a QExactive (QE) high-resolution mass spectrometer (Thermo Fisher Scientific, San Jose, USA). Four chromatographic methods were used, combined with C18 and HILIC columns, and data were acquired in both positive and negative ion modes. The mobile phase contained water, methanol, and acetonitrile, with buffers such as formic acid, PFPA, ammonium formate, or ammonium bicarbonate added according to the mode to achieve broad metabolite coverage. Metabolomics data processing: The abundance of each metabolite in each sample was expressed as its area under the curve (AUC), with undetected samples recorded as blank values. The raw peak areas were used to plot box plots and calculate inter-group mean ratios. The standardization process includes: performing a log2 transformation on each metabolite, then dividing it by the log2 median of all samples in the same batch to normalize the median to 1; missing values are filled with the minimum value for that metabolite. Subsequent statistical analysis is based on the standardized data.
[0085] (3) Detection of expression levels of SFTPA1, SFTPA2 and SMC4 proteins (proteomics) BALF supernatant and plasma samples were collected and mass spectrometry analysis was performed in data-independent acquisition (DIA) mode. Serum proteins were enriched using magnetic nanomaterials, and raw data were acquired by mass spectrometry in DIA mode. Identification and quantification were performed using Spectronaut Pulsar 17.5 (Biognosys) based on the UniProt-Homosapiens (9606, 2023.2.1) database. Data quality control and preprocessing were performed according to the existing workflow: only proteins with unique peptide ≥1 were retained; entries with a missing proportion >50% in all three groups were removed; when the effective value in a group was ≥50%, the missing value was filled with the mean of that group, and the remaining missing values were filled with half of the minimum value of the sample; then log2 transformation and z-normalization were performed for downstream analysis. Differential abundance proteins (DAPs) were screened with a threshold of |log2FC|≥1.0 and p<0.05 (P values were controlled for FDR using the Benjamini-Hochberg method), and intergroup separation was assessed using principal component analysis (PCA). KEGG enrichment was performed at the pathway level using the R package Cluster Profiler.
[0086] 3. Correlation analysis of multi-omics characteristics A correlation coefficient greater than 0.2 was considered to be relevant, and a correlation network was drawn using Cytoscape software.
[0087] 4. Establishment and evaluation of multi-omics machine learning models To predict the 28-day survival outcome of ARDS patients after intervention, this application trained four classifiers—Random Forest (RF), Logistic Regression (LR), Neural Network (NN, Multilayer Perceptron), and Naive Bayes (NB)—based on multi-omics features, all implemented in Python / scikit-learn. A hierarchical data partitioning strategy was used to maintain consistent outcome proportions. All preprocessing (z-score normalization, missing data handling, and necessary feature selection) was fitted within the training layer and applied only to the corresponding validation layer to avoid information leakage. Hyperparameters were determined solely within the training data using grid search: RF adjusted n_estimators, max_depth, and max_features; LR employed L2 regularization and optimized C; NN adjusted the hidden layer size and regularization terms (dropout, early stopping); and NB optimized the smoothing term.
[0088] 5. Statistical Analysis Statistical analysis of clinical data was performed using Python. First, the Shapiro–Wilk test was used to assess the normality of quantitative variables; normally distributed variables were expressed as mean ± standard deviation (mean ± SD), and non-normally distributed variables as median (25th–75th percentile); categorical variables were expressed as frequency and percentage (n). For intergroup comparisons: chi-square test was used for categorical variables; independent samples t-test was used for normally distributed quantitative variables; and Mann-Whitney U test was used for non-normally distributed quantitative variables. Unless otherwise stated, all tests were two-tailed, and the statistical significance threshold was set at p < 0.05.
[0089] Dimensionality reduction and correlation analysis of multi-omics data were performed in R (stats::prcomp, vegan::adonis2, tidyverse). Non-numerical variables were removed before PCA, and numerical variables were centered and standardized to unit variance. PCA was calculated using prcomp, and PC1–PC2 were extracted and the variance explained was reported. Based on the same standardized matrix, PERMANOVA (vegan::adonis2) was used to assess differences between groups, and the R results were reported. 2 The p-value was compared with the permutation test. Pearson correlation (stats::cor.test) was used to determine the pairwise associations between variables such as metabolites and microbiota, and the correlation coefficient r and p-value were reported; p < 0.05 was considered significant.
[0090] Example 1: Multi-omics analysis and biomarker discovery Metagenomics, metabolomics, proteomics, clinical indicators, and metabolic pathway detection were performed on the training set samples. Recursive Feature Elimination (IFE) was used for feature selection: within a hierarchical 5-fold cross-validation framework, a base learner (random forest, RF) was fitted only on each training fold. Variables were ranked according to feature importance (or permutation importance), and the features with the lowest contribution were iteratively removed with a fixed step size. The results are as follows: 1. Metagenomic analysis: Metagenomic sequencing of BALF samples revealed that, at the species level, the relative abundance of *Acinetobacter baumannii* was significantly higher in the death group than in the survival group (see [link to study]). Figure 1 (Figure a). Functional analysis revealed significant differences in the KEGGOrthology of the two groups. At the genus level, the abundance of *Acinetobacter* genus was significantly higher in the deceased group than in the surviving group (see Figure a). Figure 1 (Figure b in the middle)
[0091] 2. Metabolomics Analysis: Untargeted metabolomics analysis of BALF and plasma samples revealed significantly higher abundance levels of indole derivatives, particularly indoleacetic acid (IAA) and 3-formylindole (3-FI), in the BALF and plasma of deceased patients compared to surviving patients (see [link to study]). Figure 2 ).
[0092] 3. Proteomics Analysis: Proteomics analysis of BALF and plasma samples revealed that, compared with the survival group, the expression levels of SFTPA1 and SFTPA2 in the BALF of deceased patients were significantly downregulated, while the expression level of chromosome structure maintenance protein 4 (SMC4) in plasma was significantly upregulated (see [link to study]). Figure 3 ).
[0093] 4. Clinical Indicator Analysis: Analysis of clinical indicators in the survival and death groups revealed that, compared with the survival group, the death group showed a significant decrease in C-reactive protein expression and monocyte proportion, while a significant increase in neutrophil proportion (see [link to clinical data]). Figure 4 ).
[0094] 5. Metabolic pathway analysis: Differential metabolic pathways between the survivor and non-survivor groups were analyzed (p<0.05), including the tryptophan metabolism pathway. Figure 5 Figure a) shows an analysis of 10 related genes in this metabolic pathway (see Figure a). Figure 5 Figure b shows that the expression levels (abundance) of these 10 genes were positively correlated with the abundance of *Acinetobacter baumannii*, and also highly correlated with the abundance of indole metabolites such as indoleacetic acid and 3-formylindole. Further analysis of the expression differences of these 10 genes between the survival and death groups revealed significant differences in the expression of genes with KEGG database numbers K00128, K00382, K01426, K01556, K01825, K03781, K03782, and K21801 between the survival and death groups.
[0095] Example 2: Construction of the "microorganism-indole-host" axis This application constructed a core biological axis driving ARDS prognosis through cross-omics correlation network analysis. This correlation network analysis ( Figure 6 )show: The abundance of Acinetobacter baumannii was significantly positively correlated with the abundance of IAA and 3-FI (r=0.59, P<0.0001).
[0096] Activation of the IAA / 3-FI axis is negatively correlated with the levels of SFTPA1 and SFTPA2.
[0097] Activation of the IAA / 3-FI axis is positively correlated with the level of SMC4.
[0098] This result reveals a pathogenic pathway driven by Acinetobacter baumannii in the lungs, which affects local alveolar barrier proteins and systemic inflammatory proteins in the host through its metabolite indole derivatives.
[0099] In addition, this application also found correlations between the abundance of Acinetobacter baumannii, the abundance of IAA and 3-FI, the expression levels of SFTPA1 and SFTPA2, the expression level of SMC4, tryptophan metabolism pathway-related genes, and clinical indicators (including C-reactive protein levels, monocyte percentage, and neutrophil percentage). Figure 6 ).
[0100] Example 3: Validation of biomarkers or combinations thereof The above biomarkers were detected in the validation set samples, and it was found that each biomarker showed the same trend as the training set.
[0101] This embodiment uses four models: random forest, logistic regression, neural network, and Naive Bayes model. Different models are used to verify the predictive efficacy of different biomarkers or combinations of biomarkers for assessing the risk of death in ARDS patients 28 days after intervention.
[0102] Group 1) Using the expression level of SMC4 in plasma as a biomarker, the ROC curves of various models are shown in the figure. Figure 7 As shown in the middle d-figure, the AUC value can reach as high as 0.752.
[0103] Group 2) Using the abundance of indoleacetic acid and 3-formylindole in plasma and indoleacetic acid and 3-formylindole in BALF as biomarkers, the ROC curves of various models are shown below. Figure 7 As shown in Figure c, the AUC value can reach as high as 0.885.
[0104] Group 3) Using Acinetobacter baumannii abundance, indoleacetic acid abundance, 3-formylindole abundance, SFTPA1 expression level, SFTPA2 expression level, SMC4 expression level, and the expression levels of genes K00128, K00382, K01426, K01556, K01825, K03781, K03782, and K21801 as biomarker combinations, the ROC curves of various models are shown below. Figure 7 As shown in Figure b, the AUC value can reach as high as 0.921.
[0105] Group 4) Using Acinetobacter baumannii abundance, indoleacetic acid abundance, 3-formylindole abundance, SFTPA1 expression level, SFTPA2 expression level, SMC4 expression level, C-reactive protein expression level, monocyte proportion, neutrophil proportion, and K00128, K00382, K01426, K01556, K01825, K03781, K03782, and K21801 gene expression levels as biomarker combinations, the ROC curves of various models are shown below. Figure 7 As shown in Figure a, the AUC value can reach as high as 0.984.
[0106] Taking the logistic regression model as an example, the predictive power of the four biomarkers is shown in Table 3.
[0107] Table 3
[0108] This demonstrates the core value and clinical applicability of the biomarker combination proposed in this invention in the prognostic assessment of ARDS.
[0109] Furthermore, A linear prognostic scoring formula (Risk-score) based on the principle of logistic regression: Data preprocessing (standardization) To eliminate the significant differences in order of magnitude and units among different biomarkers (such as metabolite peak area, protein expression level, and cell proportion) and ensure the accuracy of the scoring model, this invention employs Z-score standardization to process the raw detection data. Before being substituted into the formula for calculation, the detection values (Raw Values) of each biomarker need to be converted into standardized values (X).
[0110] The conversion rules are as follows: (1) For metabolite abundance (IAA, 3-FI) and protein expression levels (SFTPA1, SFTPA2, SMC4, CRP): Due to the large range of their original detection values (e.g., mass spectrometry peak areas are usually on the order of 10^5 to 10^8), a log2 transformation is first performed, followed by Z-score normalization. The calculation formula is: X = (log2(detection value) - μ) / σ; (2) For cell proportions (neutrophils (Neu), monocytes (Mono)) and microbial abundance (Acinetobacter baumannii (Ab): no logarithmic transformation is required; Z-score standardization is performed directly. The calculation formula is as follows: X = (detection value - μ) / σ; In the formula, μ (mean) and σ (standard deviation) are statistical parameters of the training set. In practical applications, these statistical parameters will be treated as fixed constants.
[0111] Calculate the fraction of KO genes in the tryptophan metabolism pathway derived from bronchoalveolar lavage fluid: Score_Trp = 0.15 × (K21801 + K03782) + 0.12 × (K01825 + K01426+K00382+K00128) + 0.10 × (- K01556 + K03781) Calculate the overall prognostic risk score: Risk-score = (1.88 × X_Ab) + (1.65 × X_BALF_IAA) + (0.95 × X_Plasma_IAA) + (0.82 × X_SFTPA1)- (1.50 × X_SFTPA2) +(0.92 × X_Neu) - (0.85 × X_Mono) - (0.42 × In the formula, BALF represents bronchoalveolar lavage fluid, and Plasma represents blood plasma. For example, "X_BALF_3FI" represents the standardized value (X) of 3-formylindole in bronchoalveolar lavage fluid.
[0112] X_Ab represents the standardized value (X) of Acinetobacter baumannii derived from bronchoalveolar lavage fluid.
[0113] X_SMC4 represents the standardized value (X) of SMC4 derived from plasma.
[0114] X_SFTPA1 and X_SFTPA2 represent the standardized values (X) of SFTPA1 and SFTPA2 derived from bronchoalveolar lavage fluid.
[0115] X_CRP represents the standardized value (X) of CRP derived from plasma.
[0116] X_Neu represents the standardized value (X) of the proportion of neutrophils derived from plasma.
[0117] X_Mono represents the standardized value (X) of a monocyte derived from plasma.
[0118] When the Risk-score is ≥ 0.82, it is considered a high risk of death.
[0119] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solution of the present invention, and these simple modifications all fall within the protection scope of the present invention.
[0120] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the present invention will not describe the various possible combinations separately.
Claims
1. A biomarker for ARDS prognosis assessment, characterized in that, The biomarkers include one or both of chromosome structure maintenance protein 4 or indole derivatives; the indole derivatives preferably include indoleacetic acid and / or 3-formylindole.
2. The biomarker of claim 1, wherein, The biomarkers also include Acinetobacter baumannii and / or pulmonary surfactant protein A; the pulmonary surfactant protein A preferably includes SFTPA1 and / or SFTPA2.
3. The biomarker according to claim 1 or 2, characterized in that, The biomarkers also include genes related to the tryptophan metabolism pathway; Preferably, the tryptophan metabolism pathway-related genes include one or more of the following KEGG database numbers: K00128, K00382, K01426, K01556, K01825, K03781, K03782 or K21801. Preferably, the biomarkers further include one or more of C-reactive protein, monocyte percentage, or neutrophil percentage.
4. The biomarker of any one of claims 1-3, wherein, The biomarkers mentioned are derived from cells, tissues, organs, or body fluids; Preferably, the body fluid includes blood, plasma, or bronchoalveolar lavage fluid.
5. A chip, kit, test paper, membrane strip or device for ARDS prognosis evaluation, characterized in that, The chip, reagent kit, test strip, membrane strip, or device includes reagents for detecting any of the biomarkers described in claims 1-4. 6.A method for constructing a prognosis evaluation model of ARDS, characterized in that, The construction method includes: 1) Collect samples from the ARDS death group and the ARDS survival group, and detect the expression level or abundance of any of the biomarkers described in claims 1-4; 2) Construct a prognostic assessment model based on the information collected in step 1); Alternatively, the construction method may include: I) Collect ARDS patient samples and detect the expression level or abundance of any of the biomarkers described in claims 1-4; II) Based on the patients' clinical outcomes, ARDS patients were divided into an ARDS death group and an ARDS survival group; III) Construct a prognostic assessment model based on the test results of I) and the clinical diagnosis results of II); Preferably, the machine learning model used to construct the prognostic assessment model includes one or more of the following: random forest, logistic regression, neural network, or Naive Bayes model.
7. A system for assessing the prognosis of ARDS, characterized in that, The system includes: A data detection device, unit, or module for detecting the abundance or expression level of any of the biomarkers of claims 1-4 in a sample; A data input device, unit, or module for inputting the abundance or expression level data of the biomarker; Data analysis devices, units, or modules that calculate the 28-day mortality risk of ARDS patients after interventional treatment based on biomarker abundance or expression level data; A data output device, unit, or module for outputting analysis results of the 28-day mortality risk of ARDS patients after interventional treatment.
8. An apparatus, comprising: The device includes the system of claim 7.
9. The use of any of the biomarkers of claims 1-4, the chip, reagent kit, test strip, membrane strip or device of claim 5, the prognostic assessment model obtained by the construction method of claim 6, the system of claim 7 or the device of claim 8 in the preparation of products for prognostic assessment of ARDS.
10. The application according to claim 9, characterized in that, ARDS prognostic assessment includes predicting the risk of death 28 days after a patient receives intervention; Preferably, the intervention includes mechanical ventilation; Preferably, prognostic assessment of ARDS includes detecting the expression level or abundance of biomarkers; Preferably, the product includes one or more of the following: chip, reagent kit, test strip, membrane strip, or device.
Citation Information
Patent Citations
Kit for predicting prognosis of acute respiratory distress syndrome by using ND1 in alveolar lavage fluid
CN113252896A
Apparatus for harvesting crop
KR102024031B1
Application of reagent for detecting expression level of 40 biomarkers in sample in preparation of kit for evaluating colorectal cancer risk
CN115074446A
Lactobacillus plantarum composition for weight loss and blood sugar reduction regulation and application of lactobacillus plantarum composition
CN117264828A
Application of indole-3-propionic acid as biomarker in early warning of cognitive impairment after stroke
CN120369864A