Biomarker combination for predicting severe progress of novel coronavirus infected patient caused by Ommike variant and application of biomarker combination for predicting severe progress of novel coronavirus infected patient caused by Ommike variant

By using 10-episozolin oxidized cholaride, propofol glucuronide and diethyl 2-hydroxyglutarate as biomarkers, and combining machine learning algorithms to construct a prediction model, the prediction problem of the severe course of patients with novel coronavirus infection in the Omickron variant strain was solved, and highly accurate disease evaluation and treatment plan selection were achieved.

CN120233012APending Publication Date: 2025-07-01XIANGYANG CENT HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510166262.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art cannot effectively predict whether patients with novel coronavirus infection caused by the Omickron mutation will develop into severe illness, resulting in improper allocation of medical resources and inability to intervene in time when the condition worsens.

Method used

Three metabolites, 10-episozolin oxidized cholaride, propofol glucuronide and diethyl 2-hydroxyglutarate, were used as biomarkers, and were detected in combination with high-performance liquid chromatography, gas chromatography, enzyme-linked immunosorbent assays or rapid liquid chromatography, and predictive models were constructed through machine learning algorithms such as logistic regression, neural networks, extreme gradient enhancement, decision trees and random forests.

Benefits of technology

Accurate prediction of the severity of patients with novel coronavirus infection has been achieved, with high specificity and sensitivity, and an AUC value of 1, which can guide early clinical management and resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120233012A_ABST
    Figure CN120233012A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a biomarker combination for predicting the severe progress of a novel coronavirus infected patient caused by an Ommitake variant and application of the biomarker combination, and the biomarker combination for predicting the severe progress of the novel coronavirus infected patient caused by the Ommitake variant is composed of 10-epi-eupatoroxin, 10-epi-eupatoroxin, 10-epi-eupatoroxin, 10-epi-eupatoroxin, 10-epi-eupatoroxin, 10-epi-eupatoroxin and 10-epi-eupatoroxin. The compound is composed of three metabolites, i.e., propofol glauronide and 2-hydroxyglutarate, and the compound is prepared from the following three metabolites: propofol glauronide, propofol glauronide and 2-hydroxyglutarate. The marker combination is used as a detection marker, is high in accuracy, can predict the illness state degree of novel coronavirus infection (novel coronavirus) caused by Omicron infection, and possibly guides the selection of a treatment scheme of the novel coronavirus infection (novel coronavirus).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of biotechnology, and particularly to a biomarker combination for predicting the severe disease progression of patients infected with the novel coronavirus caused by the Omicron variant and its application. Background Art

[0002] The novel coronavirus infection is an acute respiratory infectious disease. The research on the etiology, vaccines, and specific drugs of the novel coronavirus has been attracting much attention. Although different types of novel coronavirus vaccines and therapeutic drugs have been successively approved for use, there is still no specific drug or vaccine that can eliminate the epidemic of the novel coronavirus. The main reason for the difficulty in eradicating the novel coronavirus is that its genome has strong plasticity and can continue to prevail and spread in the population through frequent mutations and recombinations. Among them, the Omicron variant was discovered in the infected population in November 2021. Due to its significantly enhanced transmissibility and immune escape ability, it quickly replaced the Delta variant at the beginning of 2022 and has now become the globally dominant epidemic strain. As of the beginning of September 2023, 5 subtypes (BA.1, BA.2, BA.3, BA.4, BA.5) of the Omicron strain have evolved into more than 750 sub-branches in a series of generations. Currently, the number of confirmed and death cases of the novel coronavirus caused by the Omicron epidemic strain is still increasing. Compared with the initial strain, although its pathogenicity and mortality are significantly reduced, and it mainly presents as asymptomatic and mild upper respiratory symptoms including cough, expectoration, nasal congestion, and runny nose in clinical practice, immunocompromised, elderly, or people with underlying diseases and other immunocompromised populations are considered high-risk groups for the novel coronavirus, and they are more likely to develop severe infections and even die. Currently, the molecular dynamic changes during different disease stages and severities related to Omicron infection are not clear. Once the condition of Omicron-infected patients worsens, they need to be immediately transferred to a professional intensive care unit for treatment, otherwise they will face life-threatening risks. Given that existing drugs and vaccines cannot completely eradicate the Omicron variant.

[0003] Therefore, there is an urgent clinical need to determine which novel coronavirus patients will develop severe diseases, so as to take effective clinical management measures in the early stage of the disease course and reasonably allocate medical resources for rapid treatment, thereby saving lives and reducing mortality. Summary of the Invention

[0004] The object of the present invention is to provide a biomarker combination for predicting the severe disease progression of patients infected with the novel coronavirus caused by the Omicron variant and its application, which can predict the severity of the novel coronavirus infection (novel coronavirus) caused by Omicron infection and may guide the selection of its treatment plan.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] In the first aspect of the embodiments of the present invention, a biomarker combination for predicting the severe disease progression of patients infected with the novel coronavirus caused by the Omicron variant is provided. The biomarker combination consists of three metabolites: 10-epi-eupatoroxin, Propofol glucuronide, and 2-hydroxyglutarate.

[0007] In the second aspect of the embodiments of the present invention, an application of the detection reagent of the biomarker combination for predicting the severe disease progression of patients infected with the novel coronavirus caused by the Omicron variant in the preparation of a product for predicting the severe disease progression of patients infected with the novel coronavirus caused by Omicron is provided.

[0008] Further, the detection reagent contains reagents for one or more detection methods selected from the following: high performance liquid chromatography, gas chromatography, enzyme-linked immunosorbent assay, mass spectrometry, or rapid liquid chromatography.

[0009] Further, the product includes at least one of a reagent, a kit, a test strip, a chip, and a system.

[0010] Further, the test sample of the product is selected from at least one of the tissue, cells, and secretions of the object to be tested.

[0011] In the third aspect of the embodiments of the present invention, a method for constructing a model for predicting the severe disease progression of patients infected with the novel coronavirus caused by Omicron is provided. The method includes using the biomarker combination for model construction to obtain a computational model for predicting the severe disease progression of patients infected with the novel coronavirus caused by the Omicron variant.

[0012] Further, the construction method includes at least one of logistic regression, neural network, extreme gradient boosting, decision tree, and random forest.

[0013] As a specific implementation manner, the method for constructing the model specifically includes:

[0014] Obtaining the laboratory test indexes of patients infected with the novel coronavirus caused by Omicron;

[0015] Performing baseline analysis and correlation analysis on the laboratory test indexes for preliminary screening to obtain the preliminarily screened variables, and randomly dividing them into a training set and a test set;

[0016] In the training set, a machine learning method is used to rank the importance of the preliminarily screened variables, and Boruta feature selection is adopted to score the variable set to obtain the finally screened variables; the finally screened variables are input, and multiple initial prediction models are used to respectively construct multiple trained prediction models;

[0017] The multiple trained prediction models are evaluated using model evaluation metrics, and the optimal trained prediction model after evaluation is determined as the model for predicting the severe disease progression of patients infected with the novel coronavirus caused by Omicron.

[0018] Further, the multiple initial prediction models include: extreme gradient boosting tree, logistic regression algorithm, random forest algorithm, decision tree, and neural network algorithm model in machine learning;

[0019] The preset model evaluation metrics include: plotting the receiver operating characteristic curves of the multiple prediction models, and evaluating the efficacy of the prediction models through the area under the ROC curve (AUC) and accuracy.

[0020] In the fourth aspect of the embodiments of the present invention, a model for predicting the severe disease progression of patients infected with the novel coronavirus caused by the Omicron variant constructed by using the above method is provided.

[0021] In the fifth aspect of the embodiments of the present invention, a system for predicting the severe disease progression of patients infected with the novel coronavirus caused by Omicron is provided. The system includes:

[0022] A processor and a memory. The memory is coupled to the processor. The memory stores instructions. When the instructions are executed by the processor, the above model is used to calculate the detection results of the biomarker combination to obtain the prediction of the severe disease progression of patients infected with the novel coronavirus caused by the Omicron variant.

[0023] In the sixth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are used: inputting the detection results of the biomarker combination for predicting the severe disease progression of patients infected with the novel coronavirus caused by the Omicron variant into the system, and obtaining the severe disease progression of patients infected with the novel coronavirus caused by the Omicron variant through calculation.

[0024] In the seventh aspect of the embodiments of the present invention, a product for predicting the severe disease progression of patients infected with the novel coronavirus caused by the Omicron variant is provided. The product includes the above system for predicting the severe disease progression of patients infected with the novel coronavirus caused by Omicron or the above computer-readable storage medium.

[0025] Furthermore, the product further includes a detection reagent for the biomarker combination for predicting the severe disease progression of patients infected with the novel coronavirus caused by the Omicron variant.

[0026] One or more technical solutions in the embodiments of the present invention have at least the following technical effects or advantages:

[0027] Among the biomarker combinations for predicting the severe disease progression of patients infected with the novel coronavirus caused by the Omicron variant provided in the embodiments of the present invention, the present invention first discovers that 10-epi-eupatoroxin, Propofol glucuronide, and 2-hydroxyglutarate can predict the severity of the novel coronavirus infection (SARS-CoV-2) caused by Omicron infection, with an AUC value reaching 1, where the optimal cut-off value is 0.437, the corresponding specificity is 1, and the sensitivity is 1. Using this biomarker as a detection marker, it has strong specificity, high sensitivity, and high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0029] Figure 1 It is a box plot of the laboratory parameter concentrations between the non-severe and severe novel coronavirus infection groups.

[0030] Figure 2 It is the principal component analysis (PCA) and orthogonal partial least squares discriminant analysis (OPLS-DA) of the potential structure for patients infected with the novel coronavirus. Among them, A is the principal component analysis diagram of metabolites in the healthy, non-severe, and severe novel coronavirus infection groups, B is the OPLS-DA diagram of metabolites between non-severe COVID-19 patients and healthy people, C is the OPLS-DA diagram of metabolites between severe COVID-19 patients and healthy people, D is the OPLS-DA diagram of metabolites between severe and non-severe COVID-19 patients, E is the permutation test diagram of OPLS-DA of metabolites between non-severe COVID-19 patients and healthy people, F is the permutation test diagram of OPLS-DA of metabolites between severe COVID-19 patients and healthy people, and G is the permutation test diagram of OPLS-DA of metabolites between severe and non-severe COVID-19 patients.

[0031] Figure 3Screening of differential metabolites in patients with novel coronavirus infection. Among them, A is the unsupervised hierarchical clustering of serum metabolomics analysis of the healthy group (n = 30), non-severe novel coronavirus group (n = 30) and severe novel coronavirus group (n = 31). The clustering uses the Euclidean distance measurement method and the Ward clustering algorithm. B is the Venn diagram analysis of the overlapping distribution of metabolites in the three groups. C is the heat map visualization of the significantly different differential metabolites (DEMs) in the severe group, non-severe novel coronavirus group and healthy control group. The color bar represents the relative intensity of the identified metabolites from -2 to 2. D is the volcano plot of differential metabolites between non-severe COVID-19 patients and healthy people. E is the volcano plot of differential metabolites between severe COVID-19 patients and healthy people. F is the volcano plot of differential metabolites between severe COVID-19 patients and non-severe COVID-19 patients. Each point in the volcano plot represents a metabolite. The names of some important DEMs are shown in the figure. The number of significantly down-regulated (green) and up-regulated (red) metabolites is shown at the top.

[0032] Figure 4 Heat map of overlapping DEMs in healthy, non-severe, and severely novel coronavirus-infected populations.

[0033] Figure 5 Identification of specific metabolite clusters in patients with novel coronavirus infection. Among them, A is the cluster with continuously increasing metabolite content from healthy control to non-severe COVID-19 patients to severe COVID-19 patients. B is the cluster with continuously decreasing metabolite content from healthy control to non-severe COVID-19 patients to severe COVID-19 patients. C is the set of other clusters except A and B.

[0034] Figure 6 KEGG enrichment analysis of DEMs in patients with novel coronavirus infection. Among them, A is the bar chart of enrichment analysis of differential metabolites between healthy control and non-severe COVID-19 patients. B is the bubble chart of enrichment analysis of differential metabolites between healthy control and non-severe COVID-19 patients. C is the bar chart of enrichment analysis of differential metabolites between healthy control and severe COVID-19 patients. D is the bubble chart of enrichment analysis of differential metabolites between healthy control and severe COVID-19 patients. E is the bar chart of enrichment analysis of differential metabolites between non-severe COVID-19 patients and severe COVID-19 patients. F is the bubble chart of enrichment analysis of differential metabolites between non-severe COVID-19 patients and severe COVID-19 patients.

[0035] Figure 7 Spearman correlation between laboratory test indicators and DEMs in severe and non-severe novel coronavirus-infected patients. Among them, A is the heat map of 455 DEMs highly correlated with laboratory tests. B shows the representative DEMs highly correlated with laboratory test parameters in the second and third groups.

[0036] Figure 8The number of key metabolic biomarkers for identifying and screening the progression of novel coronavirus infection based on a machine learning model. Among them, A is a block diagram of important variables selected for severe and non-severe novel coronavirus patients using the Boruta algorithm. B is the prediction accuracy of the multivariate random forest model with different numbers of variables in the first dataset. C shows the importance plot of variables to predict recovery based on the contribution of variables to the RF model of three metabolites based on the first dataset. D is the prediction accuracy of the multivariate random forest model with different numbers of variables in the second dataset. E shows the importance plot of variables to predict recovery based on the contribution of variables to the RF model of three metabolites based on the second dataset.

[0037] Figure 9 To identify key metabolic biomarkers that affect the progression of novel coronavirus infection based on a machine learning model. Among them, A is the flow chart of MCCV (Monte Carlo cross-validation) and boruta feature selection based on the random forest model. B is the prediction performance of decision tree (DT), random forest (RF), extreme gradient boosting (XGBoost), logistic regression (LR), and neural network (NNET) models in predicting the risk of severe novel coronavirus using AUC (area under the curve) in the training cohort (the first dataset). C is the ROC curve of the first dataset generated by the multivariate random forest model based on MCCV. The AUC-ROC value and its 95% CI are shown in the figure. D is the ROC curve of the second dataset generated by the multivariate random forest model based on MCCV. The AUC-ROC value and its 95% CI are shown in the figure. E is to predict the top 3 metabolites using the random forest model and draw the ROC curve to distinguish between non-severe and severe novel coronavirus groups.

[0038] Figure 10 The block diagram shows the change in the average correlation ranking of the top three differentially abundant metabolites in non-severe and severe novel coronavirus infection groups according to the random forest model. Detailed implementation manners

[0039] The following will specifically elaborate on the embodiments of the present invention in combination with the detailed implementation manners and examples. The advantages and various effects of the embodiments of the present invention will be presented more clearly therefrom. Those skilled in the art should understand that these detailed implementation manners and examples are used to illustrate the embodiments of the present invention, rather than limiting the embodiments of the present invention.

[0040] Throughout the specification, unless otherwise specifically stated, the terms used herein should be understood as having the meanings commonly used in the art. Therefore, unless otherwise defined, all technical and scientific terms used herein have the same meaning as those in the field to which the embodiments of the present invention belong

[0041] The biomarkers of the present application and their applications will be described in detail below in combination with experimental data.

[0042] Example 1: Biomarker combination for predicting the severe progression of patients infected with the Omicron variant of the novel coronavirus and its screening method

[0043] I. Sample collection

[0044] A total of 91 volunteers were recruited. Among them, 30 healthy individuals served as the healthy control group (Healthy group), 30 patients with mild / moderate novel coronavirus (Non-severe group), and 31 patients with severe / critical novel coronavirus (Severe group) served as the experimental group. The determination of clinical classification was carried out in accordance with the "Diagnosis and Treatment Protocol for Novel Coronavirus Infection (Trial Version 10)". The ages and genders of the above-mentioned volunteers were collected, and each participant underwent nucleic acid testing and antibody testing for the novel coronavirus. The clinical signs and symptoms of novel coronavirus patients were recorded, and routine blood tests, biochemical tests, and coagulation tests were completed simultaneously.

[0045] 2 ml of cubital venous blood was drawn from each subject into a dry tube. The serum samples were centrifuged at 12,000 g for 10 min, and the supernatant was transferred to a new centrifuge tube and stored frozen at -80 °C for subsequent metabolite detection.

[0046] II. Metabolomics processing and research on 91 samples

[0047] 1. Serum sample processing: (1) After the samples were thawed, they were vortexed for 10 s to mix evenly, and 50 μL of the samples was transferred to centrifuge tubes with corresponding numbers; (2) 300 μL of 20% acetonitrile methanol internal standard extraction solution was added, and the mixture was vortexed for 3 min and centrifuged at 12,000 r / min for 10 min at 4 °C; (3) After centrifugation, 200 μL of the supernatant was transferred to another centrifuge tube with a corresponding number and left to stand in a -20 °C refrigerator for 30 min; (4) Centrifuged again at 12,000 r / min for 3 min at 4 °C, and 180 μL of the supernatant was transferred to the liner of the corresponding injection vial for on-machine analysis.

[0048] 2. TM Broad Target Detection: TM Broad Target combines non-target (high resolution, wide coverage) with broad target (high sensitivity, precise quantification). First, the sample is subjected to a second-order spectrum scan using a high-resolution quadrupole time-of-flight mass spectrometer to extract multiple reaction monitoring (MRM) ion pair information; then, the broad target library is integrated, and the AI prediction library of the Metware integrated public database (including databases such as Metlin (Metabolite Link), HMDB (The Human Metabolome Database), and KEGG) is used to qualitatively analyze the substances in the high-resolution mass spectrum; subsequently, the ion pairs of the qualitatively analyzed substances are transferred to the mass spectrometer for retention time (RT) data acquisition to obtain peak areas and construct a specific database for the sample; finally, a triple quadrupole linear ion trap mass spectrometer is used to perform precise MRM detection on the substances in the library.

[0049] 3. Statistical Analysis of Data Using Excel: The R language (version 4.2.2) is used to analyze the differences in age and gender among the Healthy, Non-severe, and Severe groups, and to analyze the differences in clinical signs, symptoms, and laboratory parameters among the Non-severe and Severe groups. Statistical significance is achieved when P < 0.05. The t-test is used for data with a normal distribution among continuous variables, the wilcox test is used for data with a non-normal distribution among continuous variables, and the chi-square test or Fisher's exact test is used for all categorical variables. The ggplot2 (version 3.4.0) is used to plot the box plots of the changes in laboratory parameters of the Non-severe and Severe groups and to label whether the differences are significant. *P ≤ 0.05, **P ≤ 0.01, ***P ≤ 0.001.

[0050] III. Analysis and Screening of Biomarkers

[0051] 1. Data Preprocessing: The K-Nearest Neighbor (KNN) algorithm is used to fill in the missing data. PCA: The prcomp function in the R language is used to perform PCA principal component analysis on the Healthy, Non-severe, and Severe groups. OPLS-DA: The Metware metabolomics analysis platform is used to perform OPLS-DA analysis on any two of the Healthy, Non-severe, and Severe groups respectively, and the VIP value is obtained. To evaluate whether the OPLS-DA model is reliable, 200 random permutation and combination experiments are performed on the data. R2X, R2Y, and Q2 are used as evaluations. Among them, R2X and R2Y respectively represent the interpretation rates of the established model for the X and Y matrices, and Q2 represents the prediction ability of the model.

[0052] 2. Sample clustering: Unsupervised average-linkage hierarchical clustering based on Euclidean distance measurement was used to perform clustering analysis on the samples. Differential analysis: After adjusting for age and gender, differential analysis was performed between any two of the Healthy group, Non-severe group, and Severe group. Wilcox test was used for hypothesis testing. Screening was performed using p-value, variable importance in projection VIP value, and fold change (FC). Upregulation was defined as p-value < 0.05, VIP ≥ 1, and FC > 1.5. Downregulation was defined as p-value < 0.05, VIP ≥ 1, and FC < 0.7. A volcano plot was drawn using the ggplo2 package (version 3.4.0) in R language. A heatmap of the concentration changes of the total differential metabolites between any two groups in the Healthy group, Non-severe group, and Severe group was drawn using the Heatmap function of the ComplexHeatmap package (version 2.14.0).

[0053] 3. Using the Mfuzz package (version 2.60.0) in R language, metabolite concentration changes were analyzed for all metabolites after differential analysis between any two of the Healthy group, Non-severe group, and Severe group based on the fuzzy c-means algorithm.

[0054] 4. Using MetaboAnalyst 5.0, KEGG pathway enrichment was performed on the differential metabolites obtained from differential analysis between any two of the Healthy group, Non-severe group, and Severe group. The enrichment analysis used the hypergeometric test, and the formula is as follows:

[0055]

[0056] Among them, N is the number of metabolites participating in the KEGG metabolic pathway among all metabolites, n is the number of differential metabolites in N, y is the number of metabolites annotated to a certain KEGG pathway, and x is the number of differential metabolites enriched in a certain KEGG pathway. If the ratio condition x / n > y / N is satisfied, then this pathway is a KEGG enrichment pathway. Using the hypergeometric test method, the p-value of pathway enrichment was obtained and corrected using the BH method. The corrected p-value < 0.05 was used as the threshold, and the KEGG pathways that met this condition were defined as the KEGG pathways significantly enriched in the differential metabolites. These pathways were classified and annotated using KEGG, and a bar chart was drawn using ggplot2 in R language. At the same time, the pathways enriched with differential metabolites were drawn as a bubble chart using ggplot2 to visualize the KEGG enrichment results.

[0057] 5. Calculate the correlation between the differential metabolites and laboratory test parameters in the Non-Severe group and the Severe group using the corr.test function in the psych package (version 2.2.9). Select the spearman method, and use the t-test to screen the laboratory parameters significantly correlated with the differential metabolites. Adjust the P(FDR) value to less than 0.05 as the screening criterion. Use the Heatmap function in the ComplexHeatmap package (version 2.14.0) to draw the correlation heatmap. And perform hierarchical clustering on the laboratory parameters.

[0058] IV. Validation of Biomarker Combinations for Predicting the Severe Progression of Patients Infected with the Omicron Variant of SARS-CoV-2

[0059] To verify whether the identified metabolites can be used as potential predictors for differentiating between the Non-severe and Severe groups, this study selected 450 differential metabolites from 61 subjects in the Non-severe and Severe groups as the candidate dataset. 60% of the samples were used as the first dataset, and 40% of the samples were used as the second dataset. Use Boruta feature selection to identify key classification variables, and use classic models such as DT, RF, XGBoost, LR, and NNET in the tidymodels (version 1.1.1) in R language to train the training set (the first dataset), and select the best-performing random forest as the model selection.

[0060] The MetaboAnalyst 5.0 was used to generate a multivariate random forest model for each of the two datasets through MCCV to select the best biomarker combination. The variables in the two datasets were sorted according to importance and the intersection was taken, and then prediction was performed through the random forest model. In the Severe group, it was divided into the Recovery group and the Death group according to whether the outcome was death. The OPLS-DA plot was drawn using the metabolomics analysis platform of Wuhan Metware Biotechnology Co., Ltd. The differential metabolites between the Recovery group and the Death group were screened, and the screening conditions were satisfied: 1) VIP≥1; 2) FC>1.5 or <0.7; 3) p-value<0.05. The screened differential metabolites were used to draw a volcano plot with the ggplo2 package in R language, and a heatmap of the concentration changes of the differential metabolites in the Recovery group and the Death group was drawn using the Heatmap function of the ComplexHeatmap package (version 2.14.0). Based on the differential metabolite data from 11 Death patients and 8 Recovery patients as the third dataset, and the differential metabolite data from 7 Death patients and 5 Recovery patients as the fourth dataset, the same strategy was adopted to generate a multivariate random forest model for each of the two datasets through Monte Carlo cross-validation (MCCV) to select the best biomarker combination. The variables in the two datasets were sorted according to importance and the intersection was taken, and then prediction was performed through the random forest model.

[0061] 91 volunteers, including 39 females and 52 males. Eventually, 18 patients died in the Severe group. All patients with novel coronavirus were tested by real-time quantitative polymerase chain reaction (qRT-PCR) and determined to be nucleic acid positive, and the virus sequencing results showed that all were infected with the Omicron variant. Clinical laboratory examinations were performed on the first day of admission of patients with novel coronavirus and serum was isolated. At the same time, the clinical signs (Table 1) and laboratory test parameters (Table 2 and Figure 1 ) of the Non-Severe group and the Severe group of patients with novel coronavirus were collected and analyzed.

[0062] Table 1 Demographic and baseline characteristics of patients with novel coronavirus infection

[0063]

[0064]

[0065] Table 2 Laboratory parameters of patients in the novel coronavirus infection cohort determined by disease severity

[0066]

[0067]

[0068]

[0069] We first evaluated the associations among basic clinical factors, including patient age, gender, length of hospital stay, clinical examinations, underlying disease status, etc.

[0070] As shown in Table 1, we found significant differences in age, length of hospital stay, and mortality between the Severe group and the Non-Severe group (P≤0.05), where elderly patients were more likely to have more severe symptoms and even death compared to younger patients. In terms of clinical signs and symptoms, compared with the Non-Severe group, the number of patients with sore throat symptoms decreased in the Severe group, but the number of patients with increased body temperature and shortness of breath increased significantly (P≤0.05).

[0071] The results of the analysis of the differences in laboratory parameters between the Non-Severe group and the Severe group showed that, compared with the Non-Severe group, the levels of white blood cell count (WBC), neutrophil count (NEUT), lymphocyte count (LYMPH), basophil count (BASO), and eosinophil count (EO) in the Severe group were significantly decreased (P≤0.05), indicating that the immune system of the patients in the Severe group had been disordered; the levels of inflammatory indicators such as erythrocyte sedimentation rate (ESR), C-reactive protein (CRP), procalcitonin (PCT), and interleukin-6 (IL-6) were significantly increased (P≤0.05), and both the mean and median values exceeded the normal range, indicating that the inflammatory response in the Severe group was significantly higher than that in the Non-Severe group; the levels of indicators such as creatinine (Cr), urea, and human cystatin C (CysC) were significantly increased (P≤0.05), and both the mean and median values also exceeded the normal range, indicating that renal dysfunction had occurred in the Severe group; the levels of indicators such as serum total protein (T-PROT), serum albumin (ALB), and albumin / globulin ratio (A / G) were decreased (P≤0.05), while the levels of indicators such as gamma-glutamyltransferase (GGT), alanine transaminase (ALT), and glutamic oxalacetic transaminase (AST) were increased, indicating that liver dysfunction had occurred in the patients in the Severe group; the thrombin time (TT) indicator was significantly decreased (P≤0.05), and the prothrombin time (PT) and D-dimer indicators were significantly increased (P<0.001), indicating that coagulation dysfunction had occurred in the patients in the Severe group. In addition, we also found that the fasting blood glucose indicator was significantly increased (P<0.001), and the Ca ion concentration was significantly decreased (P≤0.05), etc.

[0072] These changes all indicate significant differences between the Non-Severe group and the Severe group in terms of inflammation, immunity, biochemistry, coagulation, and liver and kidney functions (Table 2 and Figure 1 ).

[0073] V. Principal Component Analysis

[0074] A total of 1,864 metabolites were detected by TM broad-targeted metabolomics in 91 serum samples included in this project. To test whether metabolic analysis could distinguish Non-Severe patients, Severe patients, and Healthy individuals infected with the novel coronavirus caused by Omicron, we used Principal Component Analysis (PCA).

[0075] The results are as Figure 2 shown in Figure A: The metabolites in the Healthy group partially overlap with those in the Non-Severe group, and there is also partial overlap between the Severe group and the Non-Severe group, while there are obvious differences between the Severe group and the Healthy group. The results of PCA analysis indicate that there are relatively obvious boundaries between different groups.

[0076] VI. Orthogonal Partial Least Squares Discriminant Analysis

[0077] To further explore the differences between groups, we used supervised Orthogonal Projections to Latent Structures Discriminant Analysis (OPLS-DA) for further research, and compared the Healthy group, Non-Severe group, and Severe group pairwise

[0078] The results are as Figure 2 shown. We found that the samples of each group were evenly distributed along the Y-axis, and the distinction between groups was obvious, confirming that there were obvious differences in the test results between any two groups ( Figure 2 Figures B - D), and the evaluation of the OPLS-DA model further confirmed that the groups could be significantly distinguished from each other ( Figure 2 Figures E - G).

[0079] The results of unsupervised hierarchical clustering based on Euclidean distance measurement showed that among the three groups of samples, except for a few samples in the Non-Severe group and the Severe group clustering together, the groups were clearly stratified and well-separated ( Figure 3 Figure A). After further quantifying the metabolites among the three groups while correcting for age and gender, we performed differential analysis between any two groups. The results showed that there were a total of 710 differential metabolites, among which 28 metabolites had significant differences between any two groupsFigure 3 Band Figure 4 )。 Figure 3 Heatmap C shows the changes in the concentrations of differential metabolites (DEMs) across different groups. More than one-third of the metabolite concentrations show a gradual increase from the Healthy group to the Non-Severe group and then to the Severe group. Conversely, nearly two-thirds of the metabolite concentrations show a gradual decrease from the Healthy group to the Non-Severe group and then to the Severe group. The volcano plot results show that compared with the Healthy group, 110 metabolites are upregulated and 83 metabolites are downregulated in the Non-Severe group ( Figure 3 D), and the upregulated metabolites include 20-Hydroxy Prostaglandin F2α, etc., and the downregulated metabolites include Cortisol, etc. When performing differential analysis between the Severe group and the Healthy group, the results show that compared with the Healthy group, 164 metabolites are upregulated and 289 metabolites are downregulated in the Severe group ( Figure 3 E), and the upregulated metabolites include Butyrylcarnitine, etc., and the downregulated metabolites include Hyodeoxycholic acid, etc. When performing differential analysis between the Severe group and the Non-Severe group, the results show that compared with the Non-Severe group, 187 metabolites are upregulated and 263 metabolites are downregulated in the Severe group ( Figure 3 F). It is worth noting that compared with the Non-Severe group, many of the upregulated metabolites in the Severe group are related to the inflammatory response and can trigger a cytokine storm, such as Cholesterol and Carnitine C4:0, etc., and the downregulated metabolites have anti-inflammatory effects, such as Hyodeoxycholic acid, Docosahexaenoic acid (DHA), etc.

[0080] Using the Mfuzz software package to process the metabolite abundance matrix, these 710 differential metabolites can be divided into 8 clusters according to their concentrations in different groups ( Figure 5 ). According to whether there are turning changes in the metabolites, these 8 clusters can be further divided into 3 groups. Figure 5 The metabolite concentrations in 3 clusters in A increase with the aggravation of disease severity (61 metabolites in Cluster 1 group, 99 metabolites in Cluster 2 group, 118 metabolites in Cluster 3 group); Figure 5The concentrations of 3 clusters of metabolites in B decreased with the aggravation of disease severity (92 metabolites in Cluster 4 group, 110 metabolites in Cluster 5 group, 106 metabolites in Cluster 6 group); while Figure 5 the metabolite concentrations in C first increased and then decreased with the aggravation of disease severity (61 metabolites in Cluster 7 group, 63 metabolites in Cluster 8 group). The above results dynamically revealed the dynamic changes of these 710 differential metabolites in different infection processes.

[0081] The differential metabolites were subjected to Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis, showing that compared with the Healthy group, the metabolic pathways such as butyric acid metabolism, fructose and mannose metabolism, pentose phosphate pathway, galactose metabolism, steroid hormone biosynthesis, and D-glutamine and D-glutamate metabolism were disordered in the Non-Severe group patients (such as Figure 6 A and 6B). Compared with the Healthy group, the metabolic pathways such as pentose phosphate pathway, tyrosine metabolism, tryptophan metabolism, pyrimidine metabolism, and phenylalanine metabolism were disordered in the Severe group (such as Figure 6 C and 6D). Compared with the Non-Severe group, the metabolic pathways such as nicotinic acid and nicotinamide metabolism, alanine-aspartate-glutamate metabolism, pentose phosphate pathway, fructose and mannose metabolism, glycerolipid metabolism, and butyric acid metabolism were disordered in the Severe group (such as Figure 6 E and 6F). These disordered metabolic pathways may be important reasons for the aggravation of the condition of patients with novel coronavirus.

[0082] We performed a correlation analysis of 450 differential metabolites (i.e., 187 up-regulated metabolites and 263 down-regulated metabolites) between the Severe group and the Non-Severe group with clinical laboratory test parameters, and found that a large number of laboratory test indicators were significantly correlated with the differential metabolites, and P<0.05. The results of the laboratory parameter cluster analysis showed that the second-layer clustering classified these test results into 4 categories, and the test results of the second and third categories had the highest correlation with the differential metabolites ( Figure 7 ). These test indicators were mainly concentrated in coagulation (D-dimer), inflammatory response (PCT, IL-6, etc.), immune function (WBC, NEUT, etc.), and liver and kidney function (Cr, Urea, T-PORT, etc.), which were consistent with the Figure 1 shown differential laboratory test indicators, further confirming the accuracy of the differential metabolites we obtained, and laying a foundation for further screening potential markers that can predict the progression of severe novel coronavirus infection caused by Omicron infection.

[0083] VII. Further Screening

[0084] We attempted to classify the Severe group and the Non-Severe group using the above 450 differential metabolites. Therefore, we used the differential metabolite data from 19 Severe patients and 18 Non-Severe patients as the first dataset and training set, and the differential metabolite data from 12 Severe patients and 12 Non-Severe patients as the second dataset and test set. First, the Boruta algorithm was used for feature selection. Among them, 59 metabolites were considered to play an important role in grouping, and 26 metabolites were considered to possibly play an important role ( Figure 8 A). After removing the metabolites that might be clinical therapeutic drugs, there were 49 metabolites left. The machine learning models we used included Decision Tree (DT), Random Forest (RF), eXtreme Gradient Boosting (XGBoost), Logistic Regression (LR), and Neural Network (NNET) models, and it was found that the prediction effect of the random forest model was the best ( Figure 9 B). Multivariate random forest models were generated through Monte Carlo cross-validation (MCCV) in the two datasets respectively, and the performance of different multivariate random forest models was evaluated using the values of area under the receiver-operating characteristic curve (AUC-ROC) and Accuracy (ACC), and variables that played an important role in model classification were screened. We found that using 3 metabolites (10-epi-Eupatoroxin, Azelaoyl PAF, and Gly-Leu-Arg-Val-Phe) for prediction in the first dataset could show strong prediction performance, with AUC value = 0.996 and ACC = 95.7% ( Figure 8 B and Figure 9 C); similarly, when using 3 metabolites (10-epi-Eupatoroxin, LPC(O-18:1), 2-Hydroxyglutarate) for prediction in the second dataset, AUC value = 0.976 and ACC = 95.8% ( Figure 8 D and Figure 9 D). Under the condition of using 3 metabolites for prediction, the variables in the two datasets were ranked according to the average importance ( Figure 8 C and Figure 8E) The top 3 metabolites in the sum and sorting are 10-epi-eupatoroxin, Propofol glucuronide, and 2-hydroxyglutarate. Among them, 2-Hydroxyglutarate and 10-epi-Eupatoroxin were significantly down-regulated in the Severe group, and Popofol Glucuronide was significantly up-regulated in the Severe group( Figure 10 ). A random forest model was constructed using the selected 3 metabolites in the training set and predicted in the test cohort. It was found that these 3 metabolites could accurately classify the Severe group and the Non-Severe group, with AUC = 1( Figure 9 E), where the optimal cut-off value was 0.437, the corresponding specificity was 1, and the sensitivity was 1. Therefore, we believe that 10-epi-eupatoroxin, Propofol glucuronide, and 2-hydroxyglutarate can be used as key biomarkers to distinguish the Non-Severe group and the Severe group in patients with novel coronavirus.

[0085] Example 2. A model for predicting the severe progression of patients infected with novel coronavirus caused by Omicron and its construction method

[0086] A method for constructing a model for predicting the severe progression of patients infected with novel coronavirus caused by Omicron includes:

[0087] S1. Perform TM wide-target detection on serum samples to construct a specific database of the samples. Randomly select 60% of the data as the training set to train the model, and the remaining 40% of the abundance data as the validation set;

[0088] S2. In the training set, use machine learning methods to sort and identify key classification variables for differential metabolites;

[0089] S3. Input the finally selected markers and use multiple initial prediction models to construct multiple post-training prediction models respectively; among them, the multiple initial prediction models include: extreme gradient boosting tree, logistic regression algorithm, random forest algorithm, decision tree, and neural network algorithm in machine learning;

[0090] S4. Evaluate the multiple post-training prediction models using the validation set, and determine the optimal post-training prediction model after evaluation as the prediction model for predicting the severe progression of patients infected with novel coronavirus caused by Omicron variant. Among them, the preset model evaluation indicators include: draw the receiver operating characteristic curve of the prediction model, and evaluate the performance of the prediction model through the area under the ROC curve AUC and accuracy.

[0091] The receiver operating characteristic curve (ROC curve) test was performed for analysis to obtain the cutoff value (optimal cut-off value). From the above results, it can be seen that 10-epi-eupatoroxin, Propofolglucuronide, and 2-hydroxyglutarate can predict the severity of coronavirus disease 2019 (COVID-19) caused by Omicron infection. The AUC value reached 1, with the optimal cut-off value being 0.437, the corresponding specificity being 1, and the sensitivity being 1.

[0092] Example 3: System for predicting the severe progression of patients with coronavirus disease 2019 caused by Omicron

[0093] An embodiment of the present invention provides a system for predicting the severe progression of patients with coronavirus disease 2019 caused by Omicron. The system includes:

[0094] A processor and a memory. The memory is coupled to the processor, and the memory stores instructions that, when executed by the processor, use the following steps:

[0095] Judge the severe progression of patients with coronavirus disease 2019 caused by Omicron according to the detection results of the input biomarker combination.

[0096] In the above technical solution, the input is the detection result of the peak area of the biomarker, and the model / system outputs a value of either 1 or 0. If the output value is 1, it indicates that the patient with coronavirus disease 2019 caused by Omicron has reached the severe level. If the output value is 0, it indicates that the patient with coronavirus disease 2019 caused by Omicron has not reached the severe level.

[0097] Example 4: Computer-readable storage medium

[0098] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the method in Example 2 and / or the method in Example 3.

[0099] Of course, for a storage medium containing computer-executable instructions provided by an embodiment of the present invention, the computer-executable instructions are not limited to the method operations as described above, and can also execute related operations in the methods provided by any embodiment of the present invention.

[0100] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disc of a computer, etc., including several instructions for causing an electronic device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0101] It should be noted that in the above embodiments, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.

[0102] Application Example 1: Predicting the severe disease progression of patients infected with the novel coronavirus caused by Omicron

[0103] 1. Experimental method

[0104] S1: Extract samples;

[0105] S2: Use mass spectrometry (ordinary mass spectrometry instruments) to detect the peak areas of the contents of three metabolites, 10-epi-eupatoroxin, Propofol glucuronide, and 2-hydroxyglutarate, in the samples.

[0106] S3: According to the quantitative detection results of step S2, output the situation of the severe disease progression of patients infected with the novel coronavirus caused by Omicron.

[0107] 2. Experimental results

[0108] The peak area of 10-epi-eupatoroxin is 560011, the peak area of Propofol glucuronide is 8700080, and the peak area of 2-hydroxyglutarate is 880773. Finally, the value 1 in the output model indicates that the patients infected with the novel coronavirus caused by Omicron have reached the severe degree.

[0109] Finally, it should also be noted that the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or apparatus.

[0110] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.

[0111] Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if these modifications and variations of the embodiments of the present invention fall within the scope of the claims of the embodiments of the present invention and their equivalent technologies, the embodiments of the present invention are also intended to include these modifications and variations.

Claims

1. A biomarker combination for predicting the progression of severe illness in patients infected with the novel coronavirus caused by an Omicron variant, characterized in that: The biomarker combination consists of three metabolites: 10-epi-eupatoroxin, propofol glucuronide and 2-hydroxyglutarate.

2. Use of the detection reagent of the biomarker combination according to claim 1 in the preparation of a product for predicting the progression of severe illness in patients infected with the new coronavirus caused by the Omicron variant.

3. The use according to claim 2, characterized in that: The detection reagent comprises reagents of one or more detection methods selected from the group consisting of: high performance liquid chromatography, gas chromatography, enzyme-linked immunosorbent assay, mass spectrometry or rapid liquid chromatography.

4. A method for constructing a model for predicting the progression of severe illness in patients infected with the new coronavirus caused by the Omicron variant, the method comprising constructing a model using the biomarker combination described in claim 1 to obtain a computational model for predicting the progression of severe illness in patients infected with the new coronavirus caused by the Omicron variant.

5. The construction method according to claim 4, characterized in that: The construction method includes at least one of logistic regression, neural network, extreme gradient boosting, decision tree, and random forest.

6. A model constructed according to the method according to any one of claims 4-5 for predicting the progression of severe illness in patients infected with the new coronavirus caused by the Omicron variant.

7. A system for predicting the progression of severe illness in patients infected with the novel coronavirus caused by a variant of Omicron, characterized in that: The system comprises: A processor and a memory, wherein the memory is coupled to the processor, and the memory stores instructions. When the instructions are executed by the processor, the model described in claim 6 is used to calculate the detection results of the biomarker combination described in claim 1 to obtain a prediction of the progression of severe illness in patients infected with the new coronavirus caused by the Omicron variant.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, the following steps are used: the detection results of the biomarker combination used to predict the progression of severe illness in patients infected with the new coronavirus caused by the Omicron variant are input into the system as described in claim 7, and the progression of severe illness in patients infected with the new coronavirus caused by the Omicron variant is obtained through calculation.

9. A product for predicting the progression of severe illness in patients infected with the novel coronavirus caused by a variant of Omicron, characterized in that: The product includes: the system for predicting the progression of severe illness in patients infected with the new coronavirus caused by Omicron as described in claim 7 or the computer-readable storage medium as described in claim 8.

10. A product for predicting the progression of severe illness in patients infected with the novel coronavirus caused by the variant strain of Omicron according to claim 1 according to claim 9, characterized in that: The products also include: detection reagents for a combination of biomarkers for predicting the progression of severe illness in patients infected with the new coronavirus caused by the Omicron variant.