Serum biomarker-based typing method for severe covid-19 patients

By screening for biomarkers SI and S-II, and combining feature selection algorithms and clustering methods, machine learning models were used to classify critically ill COVID-19 patients, solving the problem of lack of molecular subtyping in existing technologies and enabling precise treatment of critically ill patients.

CN116660541BActive Publication Date: 2026-01-16GUANGZHOU NAT LAB +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310374289.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-07
Publication Date
2026-01-16
Estimated Expiration
2043-04-07

AI Technical Summary

Technical Problem

Current technologies lack molecular subtyping methods for patients with severe COVID-19, resulting in significant differences in treatment strategy responses and making it difficult to achieve precision treatment.

Method used

Based on serum proteomics and metabolomics data, biomarkers SI and S-II were screened out. Combining feature selection algorithms and clustering methods, a machine learning model was used to classify critically ill patients. A device with acquisition and judgment modules was used for classification.

Benefits of technology

It enabled accurate classification of critically ill COVID-19 patients, provided precise treatment strategies, and improved the treatment outcomes for critically ill patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116660541B_ABST
    Figure CN116660541B_ABST
Patent Text Reader

Abstract

The application provides a typing diagnosis method for severe COVID-19 patients based on serum characteristic molecules, and specifically provides a group of biomarkers, including at least one of the following molecular subtypes: S-I and S-II, wherein S-I contains at least one of 5 highly expressed proteins, and S-II contains at least one of 8 highly expressed proteins. By determining the two molecular subtypes of S-I and S-II, accurate molecular typing of COVID-19 severe patients can be performed, accurate treatment strategies are determined, and the treatment effect of severe patients is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biotechnology, in particular, the present application relates to a typing diagnosis method for severe COVID-19 patients based on serum characteristic molecules. BACKGROUND

[0002] Coronavirus disease 2019 (COVID-19) is a new respiratory and systemic disease caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). Fever, cough, and fatigue are the main symptoms of COVID-19 patients, and few have typical symptoms such as taste and smell disorders. Bilateral lung ground-glass opacities are the most common imaging findings. People usually develop symptoms within 1-2 weeks after being infected with SARS-CoV-2, and most are diagnosed as mild or ordinary cases. A small number of severe cases develop respiratory distress and / or hypoxemia, and if not treated in time, they will quickly progress to acute respiratory distress syndrome, septic shock, and even uncorrectable metabolic acidosis and coagulation dysfunction. According to reports, the mortality rate of COVID-19 hospitalized patients is 3.28%. At present, the treatment of COVID-19 severe patients is mainly based on experience, and clinical data shows that patients have a large difference in response to treatment strategies, indicating that there is a large molecular heterogeneity among severe patients.

[0003] Therefore, based on serum proteomics and metabolomics technology, the molecular typing of severe patients will help the implementation of precise treatment strategies, thereby improving the treatment effect of severe patients. SUMMARY

[0004] The present application aims to at least partially solve at least one of the technical problems existing in the prior art.

[0005] Since the COVID-19 pandemic, multiple studies have integrated multi-omics analysis to systematically study the molecular pathological characteristics of COVID-19 patients, but there is a lack of in-depth exploration of molecular typing of severe patients. Based on human serum proteomics and metabolomics data, the inventors provide serum characteristic molecules (proteins and metabolites) that can be used for molecular typing of severe COVD-19 patients.

[0006] Therefore, in one aspect of the present application, the present application proposes a set of biomarkers. According to an embodiment of the present application, the biomarkers comprise at least one of the following molecular subtypes: S-I and S-II, wherein the S-I comprises at least one of Peptidoglycan Recognition Protein 1 (PGRP1), Von Willebrand Factor (VWF), Myeloperoxidase (PERM), Neutrophil Gelatinase-Associated Lipocalin (NGAL) and V-Set and Immunoglobulin Domain Containing 4 Protein (VSIG4); and the S-II comprises at least one of Apolipoprotein L1 (APOL1), Apolipoprotein C1 (APOC1), Apolipoprotein C2 (APOC2), Apolipoprotein C3 (APOC3), Transthyretin (TTHY), Insulin-Like Growth Factor Binding Protein 3 (IBP3), Superoxide Dismutase 1 (ALS) and Proteoglycan 4 (PRG4). By determining that the two molecular subtypes S-I and S-II can accurately classify COVID-19 severe patients, accurate treatment strategies can be determined, thereby improving the treatment effect of severe patients.

[0007] According to an embodiment of the present application, the above-mentioned biomarkers can further comprise at least one of the following additional technical features:

[0008] According to an embodiment of the present application, the biomarkers are screened by a feature selection algorithm and a clustering method.

[0009] According to an embodiment of the present application, the feature selection algorithm comprises at least one selected from the group consisting of standard deviation, median absolute deviation and coefficient of variation, or other methods having similar functions to standard deviation, median absolute deviation or coefficient of variation.

[0010] According to an embodiment of the present application, the feature selection algorithm is selected from median absolute deviation.

[0011] According to an embodiment of the present application, the clustering method is selected from at least one of hclust, kmeans, skmeans, pam and mclust. According to an embodiment of the present application, the clustering method is selected from kmeans. The inventors have found that when the algorithm is median absolute deviation and the clustering method is kmeans, the biomarkers screened by integrating the two methods are optimal in terms of precision and accuracy in classifying severe patients.

[0012] According to an embodiment of the present application, the feature selection algorithm is median absolute deviation, the clustering method is kmeans, and the kmeans parameter is set as follows: number of clusters (k) = 2, sampling rate = 0.8, and sampling times = 1000. Under this condition, effective classification of severe patients can be achieved.

[0013] In yet another aspect of the present application, the present application provides a method for training a machine learning model. According to an embodiment of the present application, the method comprises: obtaining training samples, the training samples being composed of a plurality of samples with known classification results; obtaining expression information of the biomarkers in front of the training samples; training a machine learning model by taking the expression information of the biomarkers in front as input features and the known classification results as training labels, so as to obtain a trained machine learning model.

[0014] According to an embodiment of the present application, the machine learning model is selected from, but not limited to, at least one of support vector machine, decision tree, random forest, neural network or logistic regression algorithm.

[0015] In yet another aspect of the present application, the present application provides a device for classifying patients with severe COVID-19 infection. According to an embodiment of the present application, the device comprises: an obtaining module for obtaining expression information of the biomarkers in front of serum samples of patients with severe COVID-19 infection; and a judging module for inputting the expression information of the biomarkers in front into a trained machine learning model in the method in front, so as to obtain classification results of the patients with severe COVID-19 infection. This method has good prediction performance. According to the method of the embodiment of the present application, the serum samples of the patients with severe COVID-19 infection can be effectively molecularly typed, which lays a foundation for subsequent researchers to continue analyzing the to-be-tested samples, or the severity of COVID-19 patients can be diagnosed or evaluated according to this method, reflecting the recovery of the patients, which is conducive to doctors to accurately predict the progress of the disease and timely perform clinical intervention.

[0016] According to an embodiment of the present application, the device for classifying patients with severe COVID-19 infection can further comprise at least one of the following additional technical features:

[0017] According to an embodiment of the present application, the obtaining module obtains the expression information of the biomarkers in front of serum samples of patients with severe COVID-19 infection by a clustering method.

[0018] According to an embodiment of the present application, the clustering method is selected from at least one of hclust, kmeans, skmeans, pam and mclust.

[0019] According to an embodiment of the present application, the clustering method is kmeans.

[0020] According to an embodiment of the present application, the parameters of the clustering method kmeans are set as follows: the number of clusters (k) = 2, the sampling rate = 0.8, and the sampling times = 1000.

[0021] According to an embodiment of the present application, when the expression level of the S-I biomarker is higher than the expression level of the S-II biomarker, the classification result of the severe COVID-19 infection patient is S-I type severe COVID-19 infection.

[0022] According to an embodiment of the present application, when the expression level of the S-II biomarker is higher than the expression level of the S-I biomarker, the classification result of the severe COVID-19 infection patient is S-II type severe COVID-19 infection.

[0023] According to an embodiment of the present application, the device can further comprise a probe, an antibody, a receptor or a ligand for detecting the expression level of the biomarker.

[0024] It should be noted that the "S-I type" and "S-II type" described in the present application are two molecular subtypes of COVID-19 severe patients, wherein the proteins highly expressed in the S-I molecular subtype include PGRP1, VWF, PERM, NGAL and VSIG4, and the patients with the S-I molecular subtype are clinically diagnosed as high inflammatory reactivity (high level of CRP molecules is detected by blood routine test), including excessive activation of neutrophil immune response (high level of neutrophil percentage is analyzed by blood test), and the condition is more serious; wherein the proteins highly expressed in the S-II molecular subtype include APOL1, APOC1, APOC2, APOC3, TTHY, IBP3, ALS and PRG4, and the patients with the S-II molecular subtype are clinically diagnosed as abnormally active lipid metabolism, with more serious liver damage, and the condition is less serious.

[0025] In another aspect of the present application, a system for determining the source of a to-be-tested sample is provided. According to an embodiment of the present application, the system comprises: a determination device for determining the expression information of the aforementioned biomarkers in a to-be-tested sample; and a determination device connected to the determination device, for determining the source of the to-be-tested sample based on the expression information of the biomarkers obtained by the determination device, wherein the to-be-tested sample is a serum sample. The system according to the embodiment of the present application can execute the method for determining the source of the to-be-tested sample, and can effectively perform molecular typing on the sample, thereby laying a foundation for subsequent researchers to continue analyzing the to-be-tested sample, or the system can be used for clinically diagnosing or evaluating the severity of COVID-19 patients, reflecting the recovery of the patients, and being beneficial to doctors for accurately predicting the progress of the disease and timely performing clinical intervention.

[0026] According to an embodiment of the present application, when the expression level of the S-I biomarker of the to-be-tested sample is higher than the expression level of the S-II biomarker, it is an indication that the serum sample is derived from a S-I type COVID-19 infection patient.

[0027] According to an embodiment of the present application, when the expression level of the to-be-tested sample S-II biomarker is higher than that of the S-I biomarker, it is an indication that the serum sample is derived from a patient infected with S-II coronavirus.

[0028] According to an embodiment of the present application, the system can further comprise a probe, an antibody, a receptor or a ligand for detecting the expression level of the biomarker.

[0029] Advantages:

[0030] By using the biomarker and the device of the present application, severe coronavirus infection patients can be quickly and effectively classified, and doctors can formulate precise treatment strategies for patients, thereby improving the treatment effect of severe patients.

[0031] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS

[0032] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description, taken in conjunction with the accompanying drawings, in which:

[0033] Figure 1 is a system diagram for determining the source of a to-be-tested sample according to an embodiment of the present application;

[0034] Figure 2is a graph of the construction of the subtyping matrix for different feature parameter selection algorithm and clustering algorithm combinations using the run_all_consensus_partition_methods function in the R package cola according to an embodiment of the present application, to select the best clustering parameter (the best number of clusters (k) for each combination is listed in the table). The feature selection algorithms used include: standard deviation (SD), median absolute deviation (MAD), coefficient of variance (CV), and ability to correlate to other rows (ATC); the clustering algorithms used include hclust, kmeans, skmeans, pam, and mclust. The unsupervised classification is based on 80% sampling rate and 1000 times of repeated partitioning. The PAC in the table represents the proportion of the ambiguous subgrouping (PAC). The Silhouette score represents the similarity of samples within the group. The Concordance score represents the average proportion of samples that have the same subtype label when all clustering partitions are run;

[0035] Figure 3 is a line chart for the MAD-kmeans clustering parameter combination according to an embodiment of the present application, to select the best number of clusters (k) (all subgraphs from left to right and from top to bottom are: the empirical cumulative distribution function curve; the PAC score curve; the average Silhouette score curve; the Concordance score curve; the area under the curve of the empirical cumulative distribution function curve growth curve (evaluating the increase in area compared to the previous k value); the rank curve (evaluating the percentage of sample pairs that belong to or do not belong to the same subtype for k and k-1 values); the jaccard score graph (evaluating the ratio of sample pairs belonging to the same subtype for k and k-1 values));

[0036] Figure 4 is a sample subtype clustering heatmap presenting the results of 5 random samplings and 1000 times of repeated partitioning according to an embodiment of the present application (the rows represent the variables used for clustering, and the top x parameter is presented according to random sampling clustering; the columns represent samples, and the annotation heatmap at the top represents the probability of the sample belonging to a specific subtype and the subtyping label during the entire running process);

[0037] Figure 5is a consistency clustering matrix heat map according to an embodiment of the present application (the heat map is from 5000 runs of sufficient partitioning for the MAD-kmeans algorithm combined with binary classification. Darker parts in the figure indicate that the samples are clustered into the same subtype more times. The right annotation heat map annotates the probability of sample subtyping, silhouette score and subtyping label);

[0038] Figure 6 is a uniform manifold approximation and projection (UMAP) plot representing sample classification according to an embodiment of the present application (the points in the figure represent samples, and the colors represent different classification variables sample subtype, gender, age, sample processing batch, mass spectrometry detection batch and SARS-CoV-2-specific polymerase chain reaction positive days (PPDs));

[0039] Figure 7 is a volcano plot representing differentially expressed molecules between the two subtypes of severe patients according to an embodiment of the present application (the biological class characteristics of serum molecules are identified by color);

[0040] Figure 8A is a molecular subtyping heat map of severe COVID-19 patients according to an embodiment of the present application. The upper annotation heat map from top to bottom represents the probability of patients belonging to different subtypes, subtype labels and sample classification Silhouette scores. The middle heat map represents the expression levels (z-value normalized) of differentially expressed molecules between the two subtypes. The identification criteria for differentially expressed molecules are P < 0.05 (independent sample t-test). The bottom represents the level difference of clinical characteristics of severe patients between the two subtypes (Wilcoxon test). ALT, alanine transaminase; IBIL, indirect bilirubin; IgG, immunoglobulin G; IgM, immunoglobulin M; PPD, SARS-CoV-2-specific polymerase chain reaction positive day; TT, thrombin time;

[0041] Figure 8B is the functional annotation of the up-regulated expression molecules in S-I subtype according to an embodiment of the present application. The left side is a radar chart representing the distribution of metabolite species; the middle is the top 10 GO biological processes (GOBP) to which the proteins are enriched; and the right side is the tissue or cell to which the proteins are enriched. AAP, amino acid and peptide; FAC, fatty acid and conjugate; FAL, fatty alcohol; FE, fatty ester; GPL, glycerophospholipid; LA, lineolic acid and derivative; OOC, organooxygen compound; PG, phosphatidylglycerol; PL, prenol lipid; PS, glycerophosphoserine; SD, steroid and steroid derivative; TAG, triradylglycerol;

[0042] Figure 8CFigure 9 is a radar plot showing the distribution of metabolite classes for the S-II subtype according to embodiments of the present application. AAP, amino acid and peptide; FAC, fatty acid and conjugate; FAL, fatty alcohol; FE, fatty ester; GPL, glycerophospholipid; LA, lineolic acid and derivative; OOC, organooxygen compound; PG, phosphatidylglycerol; PL, prenol lipid; PS, glycerophosphoserine; SD, steroid and steroid derivative; TAG, triradylglycerol.

[0043] Figure 8D-1 Figure 10 is a heatmap showing the validation of the molecular subtypes of severe COVID-19 patients by external data according to embodiments of the present application. The heatmap annotation is the same as (A). The validation data is from Sun et al. (Front Immunol. 2022 Jul 26; 13:893943.), BMI, body mass index; CRP, C-reaction protein; FiO2, fraction of inspired oxygen; TBIL, total bilirubin; WBC, white blood cell counts.

[0044] Figure 8D-2Validation of molecular subtypes of severe COVID-19 patients by external data according to embodiments of the present application, the heatmap is annotated as (A). The validation data comes from the study of Demichev et al. (Cell Syst. 2021 Aug 18; 12(8): 780-794.e7.), BMI, body mass index; CRP, C-reaction protein; FiO2, fraction of inspired oxygen; TBIL, total bilirubin; WBC, white blood cell counts;

[0045] Figure 8D-3 Validation of molecular subtypes of severe COVID-19 patients by external data according to embodiments of the present application, the heatmap is annotated as (A). The validation data comes from the study of Shen et al. (Cell. 2020 Jul 9; 182(1): 59-72.e15.), BMI, body mass index; CRP, C-reaction protein; FiO2, fraction of inspired oxygen; TBIL, total bilirubin; WBC, white blood cell counts;

[0046] Figure 8D-4 Validation of molecular subtypes of severe COVID-19 patients by external data according to embodiments of the present application, the heatmap is annotated as (A). The validation data comes from the study of Wu et al. (Nat Commun. 2021 Jul 27; 12(1): 4543.), BMI, body mass index; CRP, C-reaction protein; FiO2, fraction of inspired oxygen; TBIL, total bilirubin; WBC, white blood cell counts;

[0047] Figure 8E Evaluation of the prediction performance of the molecular subtype characteristics of severe COVID-19 patients on the clinical outcomes of severe COVID-19 patients according to embodiments of the present application. The left panel is the protein expression heatmap of molecular typing, and the right panel is the ROC curve for evaluating the performance of predicting the clinical outcomes of severe patients based on the molecular typing characteristics;

[0048] Figure 9is a box plot of the differences in clinical characteristics between the two subtypes of severe COVID-19 patients, S-I and S-II, according to embodiments of the present application. In the box plot, the upper and lower edges of the box are the upper and lower quartiles, respectively, and the upper and lower ends of the line are the maximum and minimum values, respectively, while the position of the median is also marked. The comparison between the two groups used a two-sided Wilcoxon test. A.G. rate (Age), albumin to Globulin rate; ALP, alkaline phosphatase; ALT, alanine transaminase; APTT, activated partial thromboplastin time; AST, aspartate aminotransferase; BNP, B-type natriuretic peptide; BUN, blood urea nitrogen; CK, creatine kinase activity; CKMB, creatine kinase-MB; CRP, C reactive protein; DBIL, direct bilirubin; GGT, gamma-glutamyl transferase; HBDH, a-Hydroxybutyrate dehydrogenase; IBIL, indirect bilirubin; IgG, immunoglobulin G; IgM, immunoglobulin M; IL6, interleukin 6; INR, international normalized ratio; LDH, lactate dehydrogenase; MB, creatine kinase-MB mass; MCH, mean corpuscular hemoglobin; MCHC, mean corpuscular hemoglobin concentration; MCV, mean corpuscular volume; MPV, mean platelet volume;NRBC, nucleated red blood cell count; PPD, SARS-CoV-2-specific polymerase chain reaction positive day; PT, prothrombin time; RDW, red cell distribution width; TBA, total bile acid; TBIL, total bilirubin; TNI, troponin I; TP, total protein; TT, thrombin time; UA, uric acid. Detailed Implementation

[0049] This invention proposes a device for classifying severely ill COVID-19 patients. For example... Figure 1 As shown, the device includes an acquisition module 100 and a judgment module 200. When the sample to be tested, i.e., the serum sample, enters the acquisition module 100, the clustering method in the acquisition module determines the expression information of biomarkers in the serum sample. Subsequently, the machine learning model in the judgment module 200 can obtain the classification result of the critically ill COVID-19 infected patient based on the expression information of the biomarkers obtained in the acquisition module 100.

[0050] The embodiments of the present invention are described in detail below. These embodiments are exemplary and are only used to explain the present invention, and should not be construed as limiting the invention. Where specific techniques or conditions are not specified in the embodiments, they are performed according to the techniques or conditions described in the literature in the art or according to the product instructions. Reagents or instruments used, unless otherwise specified, are all commercially available conventional products.

[0051] Example 1: Classification of Severe COVID-19 Patients Based on Serum Characteristic Molecules

[0052] 1. Collecting samples from severely ill COVID-19 patients:

[0053] Specifically, according to the National Health Commission COVID-19 patient diagnosis and management guidelines (7th edition), patients were divided into mild group (M) and severe group (S): mild patients showed mild clinical symptoms and no viral infection; severe patients showed dyspnea, respiratory rate ≥ 30 / min, oxygen saturation ≤ 93%, arterial oxygen partial pressure to inhaled oxygen concentration ratio (PaO2 / FiO2) < 300, and / or 24 to 48 hours lung imaging showed lung infiltration greater than 50%. At the same time, healthy physical examination subjects were recruited to form a healthy control group (H).

[0054] 2. Serum samples of the subjects were analyzed for proteomics and metabolomics data using high-performance liquid chromatography-mass spectrometry technology, and the specific process was as follows:

[0055] (1) Proteomics analysis:

[0056] 1) The ProteoMinerTM kits (Bio-Rad, USA) kit was used to remove high abundance proteins from serum and enrich low abundance proteins;

[0057] 2) The samples were sequentially treated with reduction (10 mM dithiothreitol, 37°C, 60 min), alkylation (40 mM iodoacetamide, room temperature, 30 min), and enzymolysis (trypsin: sample = 1:50, 37°C, 12 h);

[0058] 3) The enzymolyzed peptides were desalted by C18 (The Nest Group, USA) and freeze-dried;

[0059] 4) Library preparation: ① Equal amounts of treated sample peptide solutions were mixed and divided into 6 fractions and freeze-dried under vacuum; ② Equal amounts of sample serum were mixed and after removing high abundance proteins (High Select™ Top14 Abundant Protein Depletion Mini Spin Columns, Thermo Fisher, USA), the peptide solutions were prepared according to the previous method and divided into 18 fractions using a liquid chromatography system and freeze-dried under vacuum;

[0060] 5) Mass spectrometry analysis: ① Library construction: The polypeptide samples after enzymatic digestion were placed in formic acid solution and formic acid acetonitrile solution with iRT (Biognosys, Schlieren, Switzerland) standard peptide at a concentration of 0.1%, 5 μL was loaded on the EASY-nLC1000 system (Thermo Fisher Scientific, USA) nanoliter liquid chromatograph, the flow rate was set to 300 nL / min, and separation was performed on the analytical column. It was detected for 65 min using positive ion scanning. Data-dependent analysis (DDA) was completed using Q Exactive™HF-X mass spectrometer (Thermo Fisher Scientific, USA). The primary mass spectrometry scan range: 300-1 800 m / z, mass spectrometry resolution: 60 000 (m / z 200), AGC target: 3e6, Maximum IT: 50 ms, DDA data was directly imported into Spectronaut (Version 14.10; Biognosys AG, USA) software to construct the library. ② Non-data-dependent mode analysis (DIA): 2 μg of peptide segments (with appropriate iRT standard peptides) were taken from each sample, and the same detection platform as the library construction step was used for DIA analysis. DIA mode contains 60 variable scan windows. The mass spectrometry setting parameters are: 120,000 (m / z 350-1250); NCE = 27%; AGC target = 3e6; max IT = 60 ms. The cycle time is 3 seconds. DIA data is imported into Spectronaut software for analysis.

[0061] (2) Metabolomics analysis:

[0062] 1) Metabolite extraction: ① Hydrophobic metabolite extraction. Take 40 μL of serum sample and add 300 μL of methanol solution containing internal standard (PC[12:0-13:0], PE [12:0-13:0], Cer / Sph mixture I and FFA [19:0]), shake for 2 min, then add 1000 μL of methyl tert-butyl ether and 250 μL of deionized water in turn to extract hydrophobic metabolites, freeze-dried in liquid nitrogen, and store at -80℃. ② Hydrophilic metabolite extraction. Take 100 μL of serum sample and add 300 μL of methanol containing internal standard (TMAO-D9), centrifuge and take the supernatant, freeze-dried in liquid nitrogen, and store at -80℃;

[0063] 2) Mass spectrometry analysis: A combined detection platform was used, employing a Dionex™ MltiMate™ 3000 Rapid Separation LC (RSLC) system (Thermo Scientific, USA) and a Q Exactive™ hybrid quadrupole Orbitrap mass spectrometer (Thermo Scientific, USA). Hydrophobic metabolites were detected using an ACQUITY UPLC BEH C8 column (1.7 μm, 2.1 mm × 100 mm; Waters, USA), with mass spectrometry analysis performed in both anion and cation modes. Hydrophilic metabolites were detected using an UPLC BEH Amide column (2.1 mm × 100 mm, 1.7 µm; Waters, USA), with mass spectrometry analysis performed in cation mode. The mass spectrometry analysis parameters were set as follows: scan range = m / z 150-1500; mass spectrometry resolution = 70,000; AGC = 3e6; max IT = 50 ms; secondary mass spectrometry resolution was 17,500. Data processing was performed using Xcalibur 2.2SP1.48 software (Thermo Fisher Scientific, USA).

[0064] 3. Data processing and analysis: The R package 'cola' was used to perform molecular subtyping of critically ill patients, select the optimal subtyping, analyze the molecular characteristics of different subtypes, and validate the subtype analysis using public datasets.

[0065] Molecular subtyping of severely ill COVID-19 patients was performed based on serum proteomics and metabolomics data. The specific process is as follows:

[0066] (1) The proteomics data and metabolomics data were transformed by log2, normalized by median and filled with minimum values, and used as the dataset for building machine learning models;

[0067] (2) Cluster analysis was performed using the R package cola. The run_all_consensus_partition_methods function was used to construct a fractal matrix by combining different feature parameter selection algorithms and clustering algorithms, in order to select the optimal clustering parameters. Figure 2). The feature selection algorithms mentioned above include: standard deviation (SD), median absolute deviation (MAD), coefficient of variance (CV), and ability to correlate to other rows (ATC); the clustering methods include: hclust, kmeans, skmeans, pam, mclust. The setting parameters of clustering functions are as follows: partition repeat = 1,000; p_sampling = 0.8 (resampling 80%); the maximum cluster number is set to 6. The selection criteria of the best clustering strategy are as follows: (i) Jaccard index < 0.95; (ii) when 1 - PAC ≥ 0.90, the maximum cluster number (k) is selected; otherwise, the k value with the maximum 1-PAC, the maximum average silhouette, or the maximum concordance is selected Figure 3

[0068] (3) The distribution of serum metabolites differentially expressed among different subtypes was analyzed by manual annotation; based on the gene ontology biological processes (GOBP) database and the CellMarker (Zhang et al., 2019) database, the differentially expressed proteins were subjected to biological pathway enrichment analysis and tissue / cell type enrichment analysis by hypergeometric test Figure 7 and Figure 8B ).

[0069] 4. Verification of molecular typing data

[0070] (1) The repeatability of molecular typing was verified by using four reported data sets Figure 8D-1~Figure 8D-4 ​), respectively: (i) Demichev et al. study, including 44 severe patients (Cell Syst. 2021 Aug 18; 12(8): 780-794.e7.); (ii) Wu et al. study, including 27 severe patients (Nat Commun. 2021 Jul 27; 12(1): 4543.); (iii) Shen et al. study, including 13 severe patients (Cell. 2020 Jul 9; 182(1): 59-72.e15.); (iv) Sun et al. study, including 15 severe patients (Front Immunol. 2022 Jul 26; 13: 893943.). Supervised classification was performed on the feature molecules of the four datasets, respectively, and K-means (k = 2) was used to divide the patients into two subtypes; among them, the group with high expression of S-I feature molecules identified in the discovery cohort was considered as S-I subtype, and the group with high expression of S-II feature molecules was considered as S-II subtype; at the same time, combined with the different inflammatory, immune or clinical characteristics of patients in the four different datasets, the statistical differences between different subtypes were analyzed, whether it was consistent with the results in the discovery cohort population, and embodied the clinical significance of this molecular typing. (2) Using a reported dataset to verify the clinical prognosis difference of molecular typing Figure 8E ): Kimura et al. study including 10 severe patients (Sci Rep. 2021; 11: 20638.), of which 5 patients had favorable outcomes, and 5 patients had adverse outcomes (death or need for extracorporeal membrane oxygenation therapy). The differentially expressed molecules between the two groups of patients were analyzed, with P value less than 0.05 as statistically significant, and ROC curve was used to evaluate the prediction accuracy of molecular typing for the clinical outcome of severe patients;

[0071] (3) In step (2), the feature molecules of the molecular subtypes were screened Figure 8E The screening criteria are as follows: (i) There is a significant difference in expression abundance between the two subtypes (t-test, BH-adjusted P < 0.01); (ii) These differential molecules are also differentially expressed in Kimura's study cohort.

[0072] Example 2 Screening of feature molecules

[0073] 1. Recruitment of the study cohort. The study cohort includes healthy controls (H, n = 30), mild group (M, n = 42) and severe group (S, n = 82).

[0074] 2. Serum sample omics detection results. Using the method described in step 2 of Example 1, 717 serum metabolites and 628 serum proteins were identified (stable in 50% of patients in at least one study group).

[0075] 3. Two subtypes were identified in critically ill patients through consistent cluster analysis and named SI (n = 22) and S-II (n = 60). Figure 4 , Figure 5 , Figure 6 and Figure 8A The specific operating parameters are: feature selection algorithm (top_value_method) = MAD, clustering algorithm (partition_method) = kmeans, number of clusters (k) = 2, sampling rate (p_sampling) = 0.8, number of samples (partition_repeat) = 1000. SI has a higher total granulocyte count and clotting time than S-II; while S-II has higher bilirubin, eosinophil ratio, and lymphocyte count than SI. Figure 9 Since a decrease in eosinophil percentage is associated with worsening of COVID-19 disease and mortality (Liu et al., 2020; Tong et al., 2021; Zhang et al., 2020a), it is believed that patients with SI (severe eosinophilic syndrome) have more severe disease and a more intense in vivo response. Functional enrichment analysis of differentially expressed proteins between the two subtypes showed that SI patients exhibited high inflammatory reactivity, including excessive activation of neutrophils and the immune response. Figure 8B S-II patients, on the other hand, exhibit abnormally active lipid metabolism. Figure 8C Corresponding tissue enrichment analysis showed that the SI characteristic protein was mainly enriched in the bone marrow, lungs, and kidneys; while S-II was mainly enriched in the intestines, liver, and blood vessels.

[0076] Validation results from external data show that the COVID-19 severe subtype of this invention was repeatedly detected. Figure 8A~Figure 8E). In Demichev's data, S-I patients had higher neutrophil ratio (P < 0.001) and C-reactive protein level (P < 0.001) than S-II, suggesting a higher inflammatory reactivity, and S-I patients had slightly higher mortality than S-II (P = 0.10); while S-II patients had slightly higher alanine aminotransferase, suggesting more severe liver damage. This feature of S-II was also confirmed in Shen's data: S-II patients had higher alanine aminotransferase (P = 0.08) and bilirubin (P = 0.03) than S-I. Although Wu's study provided limited clinical features, it was found that S-I patients expressed lower levels of anti-inflammatory factor IFIT1 (P = 0.09), immune regulation related protein CD3E (P = 0.05), and interferon stimulated antiviral factor IFIT1 (P = 0.05). These results further confirmed that S-I is a high inflammatory reactivity patient with more severe condition, while S-II is mainly a patient with abnormal lipid metabolism.

[0077] 4. The original data set of Kimura's study was re-analyzed, and 29 proteins were identified to be differentially expressed between patients with favorable outcome and patients with adverse outcome, of which 13 proteins were also identified in step 3 to be differentially expressed between the two severe patient subtypes (Table 1). Cluster analysis was performed on the Kimura study data set with the 13 proteins, using the same running parameters in step 3, and the results are shown in Figure 8E . It can be seen that the coincidence rate of S-I and S-II obtained by the typing method of the present application with the clinical outcome of the research subjects is 90% (Table 2); Fisher's exact test P = 0.047; the area under the ROC curve (AUC) is 0.9167.

[0078] 5、Finally, 13 subtype signature molecules were identified, which were differentially expressed in patients with severe COVID-19 with different clinical outcomes, including 5 proteins highly expressed in S-I subtype (adverse clinical outcome), including Peptidoglycan recognition protein 1 (PGRP1), Von Willebrand factor (VWF), Myeloperoxidase (PERM), Neutrophil gelatinase-associated lipocalin (NGAL) and V-Set and immunoglobulin domain-containing 4 protein (VSIG4); and 8 proteins highly expressed in S-II subtype (favorable clinical outcome), including Apolipoprotein L1 (APOL1), Apolipoprotein C1 (APOC1), Apolipoprotein C2 (APOC2), Apolipoprotein C3 (APOC3), Transthyretin (TTHY), Insulin-like growth factor-binding protein 3 (IBP3), Superoxide dismutase 1 (ALS), Proteoglycan 4 (PRG4) (Table 1). In summary, the identified subtype signature molecules of patients with severe COVID-19 will help individualized treatment of COVID-19 patients. Figure 8E , Table 1). In summary, the identified subtype signature molecules of patients with severe COVID-19 will help individualized treatment of COVID-19 patients.

[0079] Table 1 Subtype signature molecules of patients with severe COVID-19

[0080]

[0081] Note: log2FC, log2 of the fold change (FC) of protein expression between two subtypes; S-I and S-II are subtypes of patients with severe COVID-19; if the log2FC (S-I / S-II) of a molecule is positive, it means that the expression of this molecule in S-I subtype is higher than that in S-II subtype.

[0082] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the skilled person in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.

[0083] Although the embodiments of the present application have been shown and described above, it is understood that the above-described embodiments are exemplary and are not to be construed as limiting the present application, and that changes, modifications, substitutions and variations can be made by those skilled in the art without departing from the scope of the present application.

Claims

1. A panel of biomarkers characterized in that, consisting of S-I combination and S-II combination, wherein, the S-I combination consists of Peptidoglycan recognition protein 1, Von Willebrand factor, Myeloperoxidase, Neutrophil gelatinase-associated lipocalin and V-set and immunoglobulin domain-containing protein 4; the S-II combination consists of Apolipoprotein L1, Apolipoprotein C1, Apolipoprotein C2, Apolipoprotein C3, Transthyretin, Insulin-like growth factor-binding protein 3, Superoxide dismutase 1 and Proteoglycan 4.

2. The biomarker of claim 1, wherein, The biomarker is screened by a feature selection algorithm and a clustering method.

3. The biomarker of claim 2, wherein, The feature selection algorithm is selected from at least one of standard deviation, median absolute deviation and coefficient of variation.

4. The biomarker of claim 3, wherein, The feature selection algorithm is selected from median absolute deviation.

5. The biomarker of claim 2, wherein, The clustering method is selected from at least one of hclust, kmeans, skmeans, pam and mclust.

6. The biomarker of claim 5, wherein, The clustering method is selected from kmeans.

7. The biomarker of claim 5, wherein, The clustering method is set as follows: number of clusters (k) = 2, sampling rate = 0.8, and sampling times = 1000.

8. A method of training a machine learning model, the method comprising: Comprising: obtaining training samples, the training samples consisting of a plurality of samples with known classification results; obtaining expression information of the biomarker of any one of claims 1-7 in the training samples; using the expression information of the biomarker of any one of claims 1-7 as input features and the known classification results as training labels to train a machine learning model, so as to obtain a trained machine learning model.

9. The method of claim 8, wherein, The machine learning model is selected from at least one of support vector machine, decision tree, random forest, neural network or logistic regression algorithm.

10. A device for classifying critically ill COVID-19 patients, characterized in that, Comprising: an obtaining module for obtaining expression information of the biomarker of any one of claims 1-7 in serum samples of severe COVID-19 patients; and a judging module for inputting the expression information of the biomarker of any one of claims 1-7 into the trained machine learning model in the method of any one of claims 8-9, so as to obtain a classification result of the severe COVID-19 patient.

11. The apparatus of claim 10, wherein, The obtaining module obtains the expression information of the biomarker of any one of claims 1-7 in serum samples of severe COVID-19 patients by a clustering method.

12. The apparatus of claim 11, wherein, The clustering method is selected from at least one of hclust, kmeans, skmeans, pam and mclust.

13. The apparatus of claim 12, wherein, The clustering method is selected from kmeans.

14. The apparatus of claim 12, wherein, The clustering method is set as follows: number of clusters (k) = 2, sampling rate = 0.8, and sampling times = 1000.

15. The apparatus of any of claims 10-14, wherein, When the expression level of the S-I biomarker is higher than that of the S-II biomarker, the classification result of the severe COVID-19 patient is S-I type severe COVID-19; When the expression level of the S-II biomarker is higher than that of the S-I biomarker, the classification result of the severe COVID-19 patient is S-II type severe COVID-19.

16. A system for determining the origin of a sample under test, characterized by Comprising: a measuring device for determining expression information of the biomarker of any one of claims 1-7 in a to-be-tested sample; A determination device connected with the determination device, used for determining the source of the sample to be tested based on the expression information of the biomarker obtained in the determination device, wherein the sample to be tested is a serum sample.

17. The system of claim 16, wherein, When the expression level of the biomarker in the sample to be tested S-I is higher than that of S-II, it is an indication that the serum sample is derived from a patient with S-I type of COVID-19; When the expression level of the biomarker in the sample to be tested S-II is higher than that of S-I, it is an indication that the serum sample is derived from a patient with S-II type of COVID-19.

18. The apparatus of any of claims 10-15, or the system of any of claims 16-17, wherein, Also included are probes, antibodies, receptors or ligands for detecting the expression level of the biomarker.

Citation Information

Patent Citations

  • Use of MAIT cells as biomarkers and biotargets in covid-19

    WO2022043496A2