A marker for predicting the risk of severe COVID-19 and applications thereof

By detecting biomarkers such as VSIG4 and KLRK1 in COVID-19 patients, a mathematical model was established to solve the problem of lagging COVID-19 severe disease risk assessment in existing technologies, and to achieve early and accurate severe disease risk assessment.

CN116004810BActive Publication Date: 2026-02-10SHENZHEN LUWEI BIOTECHNOLOGY (BIOMANIFOLD TECH CO) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310096917.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-18
Publication Date
2026-02-10
Estimated Expiration
2043-01-18

AI Technical Summary

Technical Problem

The lack of existing molecular biological methods to accurately assess the risk of severe illness in COVID-19 patients in advance leads to delays in clinical judgment.

Method used

Using a combination of biomarkers including VSIG4, KLRK1, TREML1, SAMD14, GIMAP5, SAMSN1, CD160, DGKH, PXYLP1, NUDT16, CMTM5, CSGALNACT2, TMED8, RASGEF1A, UHRF1, GIMAP7, SIRT5, EGFL7, JDP2, IRF2BPL, SIPA1L2, TRABD2A, CKAP4, and TEX2, a mathematical model was established to assess the risk of severe illness in COVID-19 patients through gene and protein level detection.

Benefits of technology

It enables a more accurate assessment of the risk of severe illness in COVID-19 patients, provides an early prediction method, and reduces the lag in clinical judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116004810B_ABST
    Figure CN116004810B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of molecular diagnosis, in particular to a marker for predicting the risk of COVID-19 severe cases and application thereof. The marker comprises at least N markers selected from VSIG4, KLRK1, TREML1, SAMD14, GIMAP5, SAMSN1, CD160, DGKH, PXYLP1, NUDT16, CMTM5, CSGALNACT2, TMED8, RASGEF1A, UHRF1, GIMAP7, SIRT5, EGFL7, JDP2, IRF2BPL, SIPA1L2, TRABD2A, CKAP4 and TEX2, and N is an optional positive integer from 1 to 24. The above marker can be used to more accurately evaluate the risk of COVID-19 severe cases, and is of great significance for rapidly evaluating the risk of COVID-19 severe cases and seeking a suitable treatment scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of molecular diagnostics, and in particular to a biomarker for predicting the risk of severe COVID-19 and its application. Background Technology

[0002] Although the proportion of severe and critically ill (intensive care unit, ICU) patients among those diagnosed with COVID-19 (SARS-CoV-2 infection) has decreased from 2% at the beginning of the epidemic to 0.2% currently, the absolute number of critically ill patients remains high due to the still large number of infected individuals. Therefore, predicting the risk of critically ill patients in advance is an important research project.

[0003] In related technologies, the National Institutes of Health (NIH) defines severe illness as one of the following: ① respiratory rate >30 breaths per minute; ② resting inspiratory oxygen saturation <94%; ③ arterial oxygen partial pressure / inspired oxygen concentration <300 mmHg; ④ lung infiltration >50%. However, these clinical indicators are relatively lagging in the rescue of critically ill patients. Therefore, it is necessary to propose a new method based on molecular biology-based molecular diagnostic technology to accurately assess and predict the risk of critical illness in advance. Summary of the Invention

[0004] This invention aims to address at least one of the technical problems existing in the prior art. To this end, this invention proposes a biomarker for predicting the risk of severe COVID-19 and its application, which can more accurately assess the risk of severe COVID-19 in patients.

[0005] In a first aspect, the present invention provides the use of a substance for detecting biomarkers in the preparation of reagents for assessing the risk or prognosis of severe COVID-19, said biomarkers comprising at least N of the following: VSIG4, KLRK1, TREML1, SAMD14, GIMAP5, SAMSN1, CD160, DGKH, PXYLP1, NUDT16, CMTM5, CSGALNACT2, TMED8, RASGEF1A, UHRF1, GIMAP7, SIRT5, EGFL7, JDP2, IRF2BPL, SIPA1L2, TRABD2A, CKAP4, and TEX2, wherein N is selected as a positive integer from 1 to 24.

[0006] According to the application of the embodiments of the present invention, at least the following beneficial effects are achieved: the above-mentioned biomarkers have good diagnostic characteristics, and by using one or more of the above-mentioned biomarkers as targets to detect their levels, the risk of severe illness in COVID-19 patients can be evaluated more accurately, thereby seeking appropriate treatment options.

[0007] In some embodiments of the present invention, the prognostic assessment includes a predictive assessment of long COVID-19 sequelae.

[0008] VSIG4 (V-set and immunoglobulin domain containing 4) is a gene encoding a protein containing the V-set and immunoglobulin domain. The protein encoded by this gene is structurally associated with the B7 family of immunomodulatory proteins. It is a phagocytic cell receptor and a potent negative regulator of T cell proliferation and IL-2 production, as well as an effective inhibitor of alternative complement pathway converting enzymes. It strongly inhibits T cell activation, but VSIG4 gene expression is limited to tissue-resident macrophages, such as liver Kupffer cells, peritoneal macrophages, pancreatic macrophages, synovial lining macrophages, and cardiac interstitial macrophages. In critically ill COVID-19 patients, the normalized expression level of the VSIG4 gene is 10 times higher than that in healthy or asymptomatic individuals.

[0009] KLRK1 (Killer Cell Lectin Like Receptor K1) is a gene that encodes the NKG2D protein, a member of the immune-activated receptor natural killer group. In critically ill COVID-19 patients, the normalized expression level of KLRK1 was 4.69 times that of healthy or asymptomatic individuals.

[0010] TREML1 (Triggering Receptor Expressed on Myeloid Cells-like-1) is a triggering receptor (TREM)-like transcript-1 gene expressed on myeloid cells. The protein it encodes is a platelet-specific receptor, a typical type 1 monoimmunoglobulin domain membrane protein receptor.

[0011] SAMD14 (Sterile Alpha Motif Domain Containing 14) is a gene that encodes a sterile α motif domain containing 14. The encoded protein can activate actin filament binding activity and is predicted to be involved in actin filament organization, calcium-mediated signal transduction, and neuronal projection development.

[0012] The GIMAP5 (GTPase immune-associated protein 5) gene encodes a protein belonging to the GTP-binding superfamily and the immune-associated nucleotide (IAN) subfamily of nucleotide-binding proteins. In humans, IAN subfamily genes are located in the 7q36.1 cluster. This gene encodes an anti-apoptotic protein that plays a role in T cell survival. There is readthrough transcription between this gene and its neighboring upstream GIMAP1 (GTPase, IMAP family member 1) gene.

[0013] SAMSN1 (SAM Domain, SH3 Domain and Nuclear Localization Signals 1) is a gene that encodes the SAM domain, SH3 domain, and nuclear localization signal 1 protein. The protein encoded by this gene is a negative regulator of B cell activation, which can downregulate cell proliferation, promote the formation of RAC1-dependent membrane folds and the reorganization of the actin cytoskeleton, regulate cell diffusion and cell polarization, stimulate HDAC1 activity, and regulate LYN activity by modulating its tyrosine phosphorylation.

[0014] CD160 (CD160 Molecule) is the gene encoding the glycoprotein CD160, which is a receptor on immune cells that transmits stimulatory or inhibitory signals that regulate cell activation and differentiation. As a receptor or ligand for TNFRSF14, a member of the TNF superfamily, CD160 participates in bidirectional cell-cell signaling between antigen-presenting cells and lymphocytes. After binding to TNFRSF14, it provides stimulatory signals to NK cells, enhancing IFNG production and anti-tumor immune responses.

[0015] The protein encoded by the DGKH (Diacylglycerol kinase eta) gene belongs to the diacylglycerol kinase (DGK) family. It contains a conserved C'-terminal catalytic domain and two cysteine-rich Zn2+ finger motifs, and has different regulatory domains.

[0016] PXYLP1 (2-Phosphoxylose Phosphatase 1) is a gene encoding xylose-2-phosphate phosphatase 1. The protein it encodes is responsible for the 2-O dephosphorylation of xylose in the glycosaminoglycan-protein linker of proteoglycans, thereby regulating the amount of mature glycosaminoglycan (GAG) chains. Furthermore, the protein encoded by the PXYLP1 gene participates in the biosynthesis of chondroitin sulfate proteoglycans, glycosaminoglycans, and heparan sulfate proteoglycans, and is located within the Golgi apparatus of the cell.

[0017] NUDT16 (Nudix Hydrolase 16) is a gene encoding Nudix hydrolase 16. The encoded protein is an RNA-binding and decapping enzyme that mainly catalyzes the breaking of the cap structure of snoRNA and mRNA in a metal-dependent manner. It may participate in the degradation of snoRNA and mRNA, the hydrolysis of inositol diphosphate (IDP) and deoxy-inositol diphosphate (dIDP), and may exclude non-canonical purines from the RNA and DNA precursor libraries, thereby preventing their incorporation into RNA and DNA and avoiding chromosome damage. Therefore, NUDT16 affects chromosome stability and is very important for cell growth.

[0018] The CMTM5 (CKLF Like MARVEL Transmembrane Domain Containing 5) gene encodes a CKLF-like MARVEL containing a 5-cell transmembrane domain. It belongs to the chemokine-like factor superfamily and is a multichannel membrane protein similar to chemokines and transmembrane 4 superfamily signaling molecules, which may exhibit tumor suppressor activity.

[0019] CSGALNACT2 (Chondroitin Sulfate N-Acetylgalactosaminyltransferase 2) is a gene encoding chondroitin sulfate N-acetylgalactosyltransferase 2 protein, primarily expressed in the Golgi apparatus of cells. It transfers 1,4-N-acetylgalactosyltransferase (GalNAc) from UDP GalNAc to the non-reducing terminus of glucuronic acid (GlcUA), adding the first GalNAc to the core tetrasaccharide linker and elongating the chondroitin chain. In critically ill COVID-19 patients, CSGALNACT21 expression is upregulated, with a normalized expression level 2.77 times higher than that in healthy individuals and asymptomatic individuals.

[0020] TMED8 (Transmembrane P24 Trafficking Protein Family Member 8) is a gene encoding member 8 of the transmembrane P24 trafficking protein family and is a paralog of ACBD3. Proteins encoded by ACBD3 play a crucial role in the sorting and modification of proteins exported from the Golgi complex and endoplasmic reticulum, and participate in the maintenance of Golgi apparatus structure and function through interaction with the global membrane protein giantin. Furthermore, it may also be involved in the hormonal regulation of steroid formation.

[0021] RASGEF1A (RasGEF Domain Family Member 1A) is a member of the RasGEF domain family. It encodes a guanine nucleotide exchange factor (GEF) that is specific (in vitro) for RAP2A, KRAS, HRAS, and NRAS. It participates in the positive regulation of cell migration and Ras protein signal transduction and may be located in the cytoplasm.

[0022] UHRF1 (Ubiquitin Like With PHD And Ring Finger Domains 1) is a ubiquitin-1 ligase with both PHD and ring finger domains, and also a member of the E3 type ubiquitin ligase family with ring finger domains. UHRF1 is a key epigenetic regulator linking DNA methylation and chromatin modification, playing a crucial role in stabilizing DNA methylation and modifying chromatin. In critically ill COVID-19 patients, UHRF1 gene expression is upregulated, with a normalized expression level 2.71 times higher than that of healthy individuals and asymptomatic individuals.

[0023] GIMAP7 (GTPase, IMAP Family Member 7) is a gene that encodes immune-associated nucleotide-binding protein 7 (GIMAP7), a GTPase belonging to the immune-associated nucleotide-binding protein GTPase family. Its expression level is downregulated in peripheral blood cells of critically ill COVID-19 patients.

[0024] SIRT5 (Sirtuin 5) is a gene encoding Sirtuin-5, an NAD-dependent protein deacetylase. The encoded protein is primarily located in mitochondria. SIRT5 has an affinity for negatively charged acyl groups and mainly catalyzes the dearylation, desuccinification, and demethylation of lysine residues, while also exhibiting weak deacetylase activity. In COVID-19, SIRT5 expression is upregulated in critically ill patients, with a normalized expression level 2.64 times higher than that in healthy individuals and asymptomatic individuals. In a 9-gene linear model of critical illness risk, SIRT5 has the largest positive weight, indicating that upregulated SIRT5 expression is a significant contributing factor to the progression to critical illness. Therefore, SIRT5 inhibitors may reduce the rapid metabolism caused by infection, thereby reducing the risk of critical illness.

[0025] EGFL7 (EGF-like Domain Multiple 7) is a gene encoding epidermal growth factor (EGF)-like domain protein 7. The protein encoded by this gene regulates tubule formation in vivo, inhibits platelet-derived growth factor (PDGF)-BB-induced smooth muscle cell migration, and promotes endothelial cell adhesion to the extracellular matrix and angiogenesis. Recent mouse experiments have shown that EGFL7 is a potential inhibitor of macrophage adhesion to mouse aortic endothelial cells. Recombinant EGFL7 protein can alleviate stress overload-induced cardiac remodeling by blocking PI3Kγ / AKT / NFκB signaling in macrophages.

[0026] JDP2 (Jun Dimerization Protein 2) encodes JUN dimerization protein 2, a component of the AP-1 transcription factor. This protein inhibits transcriptional activation mediated by JUN family proteins and participates in various AP-1-related transcriptional responses, such as UV-induced apoptosis, cell differentiation, tumorigenesis, and anti-tumor activity. Simultaneously, the protein encoded by this gene can also act as a repressor by recruiting histone deacetylase 3 / HDAC3 to the promoter region of JUN, controlling transcription through direct regulation of histone modification and chromatin assembly. The normalized expression level of JDP2 in severely ill COVID-19 patients was 2.60 times higher than that in healthy individuals and asymptomatic individuals. Since high JDP2 expression may lead to cardiac dysfunction, this is one of the causes of severe illness.

[0027] IRF2BPL (Interferon Regulatory Factor 2 Binding Protein Like) encodes an interferon regulatory factor 2 binding protein-like protein that enables E3 ubiquitin ligase to participate in proteasome-mediated ubiquitin-dependent target protein degradation, playing a role in central nervous system development and neuronal maintenance.

[0028] SIPA1L2 (Signal Induced Proliferation Associated 1-Like 2) is a member of the signal-induced proliferation-associated 1-like family. This family includes a GTPase activation domain, a PDZ domain, and a C-terminal helical coil domain with a leucine zipper. Amphibiosomes are organelles in the autophagy pathway, formed by the fusion of autophagosomes and late endosomes. The biogenesis of autophagosomes and late endosomes continues at the axonal terminals. SIPA1L2 is involved in the retrograde transport of neuronal BDNF / TrkB cells to human cells during presynaptic clonus.

[0029] TRABD2A (TraB Domain Containing 2A) encodes a TraB domain containing 2A and is a metalloproteinase that acts as a negative regulator of the Wnt signaling pathway by mediating the cleavage of eight N-terminal residues of the Wnt protein subset. Following cleavage, the Wnt protein is oxidized and forms large disulfide oligomers, leading to its inactivation. Furthermore, the protein encoded by this gene can cleave WNT3A and WNT5, but not WNT11.

[0030] CKAP4 (Cytoskeleton Associated Protein 4) encodes cytoskeleton-associated protein 4, which mediates the anchoring of microtubules in the endoplasmic reticulum. Furthermore, the protein encoded by the CKAP4 gene regulates the proliferation and apoptosis of primary vascular smooth muscle cells, and is a key component in the pathogenesis of abdominal aortic aneurysms.

[0031] TEX2 (Testis Expressed 2) encodes the testis-expressed protein 2. During endoplasmic reticulum stress or when cellular ceramide levels increase, the protein encoded by this gene induces contact between the endoplasmic reticulum and the medial Golgi complex to promote non-cystic transport of ceramide from the endoplasmic reticulum to the Golgi complex and its conversion into complex sphingolipids, preventing the accumulation of toxic ceramides.

[0032] In some embodiments of the present invention, the markers include at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen, at least sixteen, at least seventeen, at least eighteen, at least nineteen, at least twenty, at least twenty-one, at least twenty-two, at least twenty-three, and all twenty-four of the following: VSIG4, KLRK1, TREML1, SAMD14, GIMAP5, SAMSN1, CD160, DGKH, PXYLP1, NUDT16, CMTM5, CSGALNACT2, TMED8, RASGEF1A, UHRF1, GIMAP7, SIRT5, EGFL7, JDP2, IRF2BPL, SIPA1L2, TRABD2A, CKAP4, and TEX2.

[0033] In some embodiments of the present invention, the markers include at least N of the following: VSIG4, TREML1, SAMD14, DGKH, UHRF1, SIRT5, TEX2, KLRK1, and GIMAP7, where N is any positive integer from 1 to 9.

[0034] In some embodiments of the present invention, the markers include at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, and all nine of VSIG4, TREML1, SAMD14, DGKH, UHRF1, SIRT5, TEX2, KLRK1, and GIMAP7.

[0035] In some embodiments of the present invention, the markers include VSIG4, KLRK1, TREML1, SAMD14, GIMAP5, SAMSN1, CD160, DGKH, PXYLP1, NUDT16, CMTM5, CSGALNACT2, TMED8, RASGEF1A, UHRF1, GIMAP7, SIRT5, EGFL7, JDP2, IRF2BPL, SIPA1L2, TRABD2A, CKAP4, and TEX2.

[0036] In some embodiments of the present invention, the markers include VSIG4, TREML1, SAMD14, DGKH, UHRF1, SIRT5, TEX2, KLRK1, and GIMAP7.

[0037] In some embodiments of the present invention, the markers include KLRK1 and DGKH.

[0038] In some embodiments of the present invention, the markers include KLRK1 and SIRT5.

[0039] In some embodiments of the present invention, the markers include VSIG4 and KLRK1.

[0040] In some embodiments of the present invention, the markers include KLRK1 and PXYLP1.

[0041] In some embodiments of the present invention, the markers include UHRF1 and SIRT5.

[0042] In some embodiments of the present invention, the markers include KLRK1, DGKH, and UHRF1.

[0043] In some embodiments of the present invention, the markers include KLRK1, DGKH, and GIMAP7.

[0044] In some embodiments of the present invention, the markers include KLRK1, PXYLP1, and UHRF1.

[0045] In some embodiments of the present invention, the markers include KLRK1, UHRF1, and SIRT5.

[0046] In some embodiments of the present invention, the markers include KLRK1, SAMD14, and DGKH.

[0047] In some embodiments of the present invention, the markers include KLRK1, DGKH, UHRF1, and GIMAP7.

[0048] In some embodiments of the present invention, the markers include KLRK1, SAMD14, DGKH, and GIMAP7.

[0049] In some embodiments of the present invention, the markers include KLRK1, SAMSN1, DGKH, and GIMAP7.

[0050] In some embodiments of the present invention, the markers include KLRK1, TREML1, DGKH, and GIMAP7.

[0051] In some embodiments of the present invention, the markers include KLRK1, DGKH, UHRF1, and TEX2.

[0052] In some embodiments of the present invention, the markers include VSIG4, KLRK1, SIRT5, UHRF1, and TRABD2A.

[0053] In some embodiments of the present invention, the markers include VSIG4, KLRK1, SIRT5, UHRF1, and GIMAP7.

[0054] In some embodiments of the present invention, the markers include VSIG4, KLRK1, SIRT5, UHRF1, and CSGALNACT2.

[0055] In some embodiments of the present invention, the markers include VSIG4, KLRK1, SIRT5, UHRF1, and SAMSN1.

[0056] In some embodiments of the present invention, the markers include VSIG4, KLRK1, SIRT5, UHRF1, and TREML1.

[0057] In some embodiments of the present invention, the markers include VSIG4, TREML1, SAMD14, DGKH, UHRF1, SIRT5, TEX2, KLRK1, and GIMAP7.

[0058] In some embodiments of the present invention, the biomarkers include TEX2, DGKH, UHRF1, CSGALNACT2, and KLRK1. Using these biomarkers as targets, the risk of severe COVID-19 in adult patients aged 20-54 can be predicted relatively accurately.

[0059] In some embodiments of the present invention, the markers include SIRT5, TEX2, UHRF1, DGKH, VSIG4, CSGALNACT2, and KLRK1.

[0060] In some embodiments of the present invention, the markers include TEX2, DGKH, UHRF1, VSIG4, CSGALNACT2, and KLRK1.

[0061] In some embodiments of the present invention, the markers include SIRT5, TMED8, TEX2, UHRF1, DGKH, GIMAP7, KLRK1, and CSGALNACT2.

[0062] In some embodiments of the present invention, the markers include SIRT5, TMED8, UHRF1, DGKH, GIMAP7, CSGALNACT2, and KLRK1.

[0063] In some embodiments of the present invention, the markers include TEX2, DGKH, CSGALNACT2, and KLRK1.

[0064] In some embodiments of the present invention, the biomarkers include UHRF1, SIRT5, DGKH, TREML1, IRF2BPL, EGFL7, KLRK1, CMTM5, and RASGEF1A. Using these biomarkers as targets, the risk of severe COVID-19 in elderly patients aged 55-90 years can be predicted relatively accurately.

[0065] In some embodiments of the present invention, the markers include SIRT5, UHRF1, IRF2BPL, DGKH, KLRK1, and RASGEF1A.

[0066] In some embodiments of the present invention, the markers include SIRT5, UHRF1, IRF2BPL, and RASGEF1A.

[0067] In some embodiments of the present invention, the markers include SIRT5, DGKH, RASGEF1A, and KLRK1.

[0068] In some embodiments of the present invention, the markers include SIRT5, UHRF1, and RASGEF1A.

[0069] In some embodiments of the present invention, the samples for the COVID-19 severe illness risk assessment or prognosis assessment are selected from tissues, blood, bronchoalveolar lavage fluid, or urine.

[0070] A second aspect of the present invention provides a product comprising a substance for detecting the aforementioned markers.

[0071] In some embodiments of the present invention, the substance for detecting biomarkers includes a substance for detecting biomarkers at the gene level and / or protein level;

[0072] Preferably, the substance comprises a substance for use in one or more detection techniques or methods selected from the group consisting of: immunohistochemistry, Western blotting, Northern blotting, PCR, and microarray.

[0073] Preferably, the immunohistochemical method is selected from at least one of the following: immunofluorescence analysis, reverse enzyme-linked immunosorbent assay (ELISA), and immunogold assay.

[0074] In some embodiments of the present invention, the product comprises reagents, kits, test strips, chips, or systems.

[0075] A third aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions for causing a computer to perform the following operations:

[0076] Step 1: Obtain information on the expression levels of the above biomarkers in samples from COVID-19 patients;

[0077] Step 2: Perform mathematical correlation on the expression levels to obtain a score; the score is used to indicate the risk of severe illness in COVID-19 patients.

[0078] In some embodiments of the present invention, the COVID-19 patient's sample is derived from at least one of the patient's blood, tissue, or cell samples.

[0079] In some embodiments of the present invention, the scoring Where ai represents the expression level of the marker, bi represents the assigned weight of the marker, and n represents the number of markers.

[0080] In some embodiments of the present invention, when the score is higher than a set value, it indicates that the COVID-19 patient has a higher risk of severe illness.

[0081] A fourth aspect of the present invention provides a computer device including a processor and a memory, wherein the memory stores a computer program executable on the processor, and the processor performs the following operations when running the computer program:

[0082] Step 1: Obtain information on the expression levels of the above biomarkers in samples from COVID-19 patients;

[0083] Step 2: Perform mathematical correlation on the expression levels to obtain a score; the score is used to indicate the risk of severe illness in COVID-19 patients.

[0084] In some embodiments of the present invention, the COVID-19 patient's sample is derived from at least one of the patient's blood, tissue, or cell samples.

[0085] In some embodiments of the present invention, the scoring is described in the formula. Where ai is the expression level of the marker, bi is the assigned weight of the marker, n is the number of markers, and n≤N.

[0086] In some embodiments of the present invention, when the score is higher than a set value, it indicates that the COVID-19 patient has a higher risk of severe illness.

[0087] In some embodiments of the present invention, the set value is the Cutoff value, i.e., the criticality score threshold value.

[0088] The criticality score threshold can be determined in one of three ways: 1. Statistically analyze the scores of non-critical individuals, including healthy people, asymptomatic COVID-19 patients, and those with mild symptoms, and then take the q% percentile (e.g., 95%) of the scores for this group, which can control the false positive rate to no more than 5%; 2. Alternatively, statistically analyze the scores of known critical individuals, and then take the minimum value or p% percentile (e.g., p=0, or 2) of the scores for critical individuals to ensure that the false negative rate in the training dataset is 0% or 2%; 3. Alternatively, use an ROC curve to treat false positives and false negatives as equally important optimal points.

[0089] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. Attached Figure Description

[0090] The present invention will be further described below with reference to the accompanying drawings and embodiments, wherein:

[0091] Figure 1 The ROC plot shows the top 12 genes with the largest AUC for predicting critical illness based on gene expression levels, as presented in this invention.

[0092] Figure 2 This is a Venn diagram showing the t-test results of an embodiment of the present invention.

[0093] Figure 3 This is the ROC curve used to validate the biomarker regression model in all samples in Example 24 of this invention.

[0094] Figure 4 This is the ROC curve used to validate the biomarker regression model in all samples in Example 9 of this invention.

[0095] Figure 5 This is a histogram showing the expression levels of seven positively weighted genes in different populations in the nine biomarker regression model of this invention.

[0096] Figure 6 This is a histogram showing the expression levels of two negatively weighted genes in different populations in the 9 biomarker regression model of this invention.

[0097] Figure 7 The ROC curve of the GSE157103 proteome dataset was used to validate the regression model of the nine biomarkers of this invention.

[0098] Figure 8 This is a graph showing the percentage of severe cases at different ages in the 217 samples of this invention.

[0099] Figure 9 This is the univariate ROC curve for predicting criticality indicators using age in this invention.

[0100] Figure 10 This is a Venn diagram showing the genes associated with the risk of severe COVID-19 in the adult and elderly groups of this invention. Detailed Implementation

[0101] The following will describe the concept and technical effects of the present invention clearly and completely with reference to embodiments, so as to fully understand the purpose, features and effects of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the scope of protection of the present invention.

[0102] In the description of this invention, the terms "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0103] Unless otherwise specified in the examples, the procedures should be performed under standard conditions or conditions recommended by the manufacturer. Reagents or instruments whose manufacturers are not specified are all commercially available products.

[0104] Example 1

[0105] This embodiment utilizes mRNA gene expression data to screen for biomarkers used to predict the risk of severe COVID-19, as follows:

[0106] I. Dataset Preparation

[0107] 1. Training Dataset Collection: The training dataset for this embodiment is selected from the public datasets GSE152418, GSE157103 and GSE162562. This training dataset consists of transcriptome NGS data from 251 peripheral blood samples, including 88 uninfected healthy human samples (denoted as HC) and 163 COVID-19 infected samples, including 43 asymptomatic (denoted as Asymptom), 58 mild cases (denoted as Mild or Moderate), and 62 severe cases (denoted as Severe) and ICU cases.

[0108] The Severity Flag was defined as follows: all severe cases and ICU patients were marked as 1, totaling 62 cases; others were marked as 0, including those in HC, asymptomatic and mild cases, totaling 189 cases.

[0109] In addition, age information was missing in 34 of the 251 patients. Among the 217 patients with age information, the mean age was 50 years, the median age was 53 years, the minimum age was 1 year, the maximum age was 90 years, the Q1 (25th percentile) age was 32 years, and the Q3 (75th percentile) age was 64 years.

[0110] 2. Data Normalization: Since the three sets of data came from different laboratories' next-generation sequencing (NGS) and data analysis workflows, although all used the Illumina NovaSeq 6000 platform, differences in laboratory procedures and bioinformatics pipelines necessitated data normalization. The steps are as follows:

[0111] A. For each set of data, first perform a logarithmic (log2) transformation on the expression level readings.

[0112] B. Perform quartile normalization on each sample. Given a sample, calculate the 25th and 75th percentiles of the gene expression vector. Percentiles can be calculated by first sorting the expression values ​​from smallest to largest, then looking at the values ​​corresponding to the 25th and 75th percentiles, denoted as [Q1, Q3]. Using the linear mapping of [Q1, Q3] to the unit interval [0, 1], calculate: Normalized expression level = (Original expression level - Q1) / (Q3 - Q1). Therefore, the expression vectors of all genes are aligned. Sample normalization eliminates variations in the input sample size before sequencing.

[0113] C. Perform the same quartile normalization on each gene. At this point, [Q1, Q3] correspond to the normalized readings of that gene in different samples calculated in B. Gene normalization removes variations introduced during sequencing, such as shifts in the absolute values ​​of gene expression levels caused by different library preparation methods or sequencing depths.

[0114] D. Take the intersection subset of the genomes from the three datasets, resulting in 6136 (denoted by n) identical genes. Assemble the normalized expression levels of the three sets of samples (number of samples m1, m2, and m3) corresponding to the n genes into the final dataset. Specifically, assuming M1, M2, and M3 are the normalized expression level matrices of the three datasets, with dimensions m1×n, m2×n, and m3×n respectively, then after assembly (keeping the matrix columns unchanged and stacking the rows, for example, using the R function rbind), the final expression matrix is ​​(m1+m2+m3)×n.

[0115] II. Marker Selection and Modeling

[0116] 1. For the 6136 genes in the above training dataset, calculate the univariate ROC curve and corresponding AUC of the gene expression level to predict critical illness indicators, and sort them in descending order of AUC.

[0117] For example, Figure 1 The ROC plots for the 12 genes with the largest AUCs are shown. Among them, the following genes are downregulated in critically ill patients: KLRK1 (AUC=0.83), CRTAM (AUC=0.82), TRAF3IP3 (AUC=0.80), TMEM229B (AUC=0.80), GIMAP5 (AUC=0.78), CD160 (AUC=0.78), GIMAP7 (AUC=0.78), and CCDC65 (AUC=0.78); while the following genes are upregulated in critically ill patients: VSIG4 (AUC=0.81), UHRF1 (AUC=0.80), CHPT1 (AUC=0.78), and TMED8 (AUC=0.78).

[0118] 2. From 6136 genes, genomes specific to critical illness were screened, and t-tests were used to compare the following populations:

[0119] A: HC vs COVID-19, denoted as COVID_INDX;

[0120] B: Asymptomatic vs. ICU, denoted as Asymptom_Severe;

[0121] C: Mild cases vs. ICU, denoted as Mild_Severe;

[0122] D: ICU=0 vs ICU=1, denoted as SeverityFlag;

[0123] E: HC vs. asymptomatic, denoted as HC_Asymptom.

[0124] Genes with t-test p-values ​​less than 0.05 in AD and genes with t-test p-values ​​not less than 0.05 in E are selected as the intersection of the two groups. Figure 2 The Venn diagram shown above reveals 390 overlapping genes across the five groups, indicating that the expression levels of these genes were not significantly different between the HC and asymptomatic individuals, but showed statistically significant differences in the AD group.

[0125] 3. Using the SeverityFlag t-test to compare the mean expression levels Mu0 and Mu1 of the SeverityFlag (critical group) (SeverityFlag = 0, i.e., all critically ill and ICU patients) and (SeverityFlag = 1, i.e., HC, asymptomatic, and mildly ill individuals), the fold change of each gene was calculated as: 2^(|Mu1 - Mu0|). For the fold change of 390 genes, the 95th percentile Q95 (approximately 2.5) was used as the lower limit. Finally, candidate genes meeting the following requirements were screened:

[0126] ① Corresponding to AUC > 0.65;

[0127] ② The multiple change is greater than 2.5.

[0128] Using the above method, a total of 24 genes associated with the risk of severe illness were screened, namely: VSIG4, KLRK1, TREML1, SAMD14, GIMAP5, SAMSN1, CD160, DGKH, PXYLP1, NUDT16, CMTM5, CSGALNACT2, TMED8, RASGEF1A, UHRF1, GIMAP7, SIRT5, EGFL7, JDP2, IRF2BPL, SIPA1L2, TRABD2A, CKAP4, and TEX2.

[0129] 4. Using the above 24 genes, linear regression statistics were performed on all samples to obtain a 24-marker regression model for predicting severe illness. The parameters of each gene in this model are shown in Table 1:

[0130] Table 1.24 Relevant parameters of each gene in the biomarker regression model

[0131]

[0132]

[0133] The risk of severe COVID-19 is calculated based on the weights of the biomarkers described above. The formula for the severe risk score (critical score) is: 0.0194×VSIG4 - 0.0604×KLRK1 + 0.0342×TREML1 + 0.0196×SAMD14 + 0.0018×GIMAP5 - 0.0095×SAMSN1 + 0.0111×CD160 + 0.0387×DGKH + 0.0231×PXYLP1 - 0.0022×NUDT16 - 0.0409×CMTM5 - 0.0435×C SGALNACT2+0.0523×TMED8-0.0053×RASGEF1A+0.0605×UHRF1-0.0322×GIMAP7+0.0641×SIRT5+0.0166×EGFL7+0.007×JDP2+0.0112×IRF2BPL-0.0155×SIPA1L2+0.0284×TRABD2A+0.0137×CKAP4+0.0565×TEX2, where the abbreviations of the markers in the formula represent the standardized values ​​of the expression levels of the corresponding markers.

[0134] ROC curves were plotted on all the above samples using the established 24-marker regression model to test the model's ability to assess the risk of severe illness in patients. The results are as follows: Figure 3 As shown, the AUC is 0.949, and the specificity (1-false positive rate) corresponding to the optimal decision point on the ROC curve (as shown by the dashed line) is 93%, and the sensitivity is 87%.

[0135] In practical applications, given the urgency of predicting the progression of COVID-19 infection to severe illness, using qPCR to detect the expression of 24 genes also involves a significant amount of laboratory work. This embodiment further optimizes the regression model for the aforementioned 24 biomarkers, primarily to reduce the number of genes in the model without significantly compromising accuracy. The specific optimization method is as follows:

[0136] The method employs iterative linear regression, where an upper limit p-value (pMax) is selected. In each iteration, linear regression is used to model the model, and genes with p-values ​​greater than pMax are removed. This process is repeated until every p-value in the model is less than pMax. In this embodiment, pMax is set to 0.01, 0.02, 0.03, ..., 0.1, and 10 models are built iteratively. Finally, all genes in the models are combined to build the final model, and genes that seem to conflict are removed, resulting in an optimized model.

[0137] The final optimized 9-biomarker regression model includes 9 genes associated with the risk of severe illness, namely VSIG4, TREML1, SAMD14, DGKH, UHRF1, SIRT5, TEX2, KLRK1, and GIMAP7. Among them, the genes highly expressed in critically ill patients are VSIG4, TREML1, SAMD14, DGKH, UHRF1, SIRT5, and TEX2; the genes with low expression are KLRK1 and GIMAP7.

[0138] The parameters of each gene in the above 9 biomarker prediction model are shown in Table 2 below:

[0139] Table 2.9 Relevant parameters of each gene in the biomarker prediction model

[0140]

[0141] The model accuracy indices, such as the modeling AUC, corresponding to the regression model based on the above 9 biomarkers are shown in Table 3. The modeling AUC refers to the weighted sum calculated for each sample using the 251 samples in the training set, based on the expression levels of the 9 genes and the weights listed in Table 2: VSIG4×0.0149-KLRK1×0.0557+TREML1×0.0139+SAMD14×0.0215+DGKH×0.0413+UHRF1×0.0656-GIMAP7×0.0273+SIRT5×0.078+TEX2×0.0327+0.0925, to obtain the severe illness risk score (critical illness score) for that sample.

[0142] ROC curves were plotted on all 251 samples using the established 9-biomarker regression model to test the model's ability to assess the risk of severe illness in COVID-19 patients. The results are as follows: Figure 4 As shown, the AUC is 0.923, and the specificity (1 - false positive rate) corresponding to the optimal decision point on the ROC curve (as shown by the dashed line) is 94%, and the sensitivity is 87%.

[0143] The ROC curve is plotted by iterating through all possible severity score thresholds. Starting from the lowest severity score and progressing to the highest, the jump step size is increased each time (typically 1% of the interval length). The corresponding False Positive Rate (FPR) and True Positive Rate (TPR) are calculated, and the corresponding point on the ROC curve is plotted at coordinates (FPR, TPR). If the jump step size is 1% of the severity score interval length, the ROC curve will have 101 points. The optimal severity score threshold is taken as the point with the closest Euclidean distance to the top-left corner of the ROC unit frame at (FPR, TPR) = (0, 1). (0, 1) represents 0% false positives and 100% true positives, representing the most perfect prediction result. Figure 4 The value at (0.06, 0.87) is shown. Taking the 13th percentile (1-TPR) of the critical score of all samples with a criticality index of 1, i.e., the percentile of (1-0.87), we get 0.3227, which is the critical score threshold (CUTOFF) in Table 3. Similarly, we can take the 94th percentile (1-FPR) of the critical score of samples with a criticality index of 0 to get the same critical score threshold. Here, it is assumed that false positives and false negatives are equally important in the application.

[0144] Table 3. Accuracy Indicators

[0145] AUC FPR TPR CUTOFF ACC N0 N1 0.923 0.06 0.87 0.3227 0.92 189 62

[0146] Combined with Table 3 and Figure 4 It can be seen that the criticality score threshold of the 9-marker model is 0.3227. In 251 patients, the model accuracy rate is 92.3%, with 6% false positives in 189 non-critical cases and 13% false negatives in 62 critical cases.

[0147] For the 9-marker regression model constructed above, the histogram of single-gene expression levels is as follows: Figure 5 and Figure 6 As shown, where Figure 5 Histograms showing the expression levels of seven positively weighted genes (VSIG4, TREML1, SAMD14, DGKH, UHRF1, SIRT5, TEX2, KLRK1, and GIMAP7) in different populations in a 9-marker regression model. Figure 6 Histograms showing the expression levels of two negatively weighted genes (KLRK1 and GIMAP7) in different populations in a 9-marker regression model.

[0148] Example 2

[0149] All gene subsets of the aforementioned 24 genes can also be used to build models for predicting severe COVID-19. In this embodiment, K (2, 3, 4, and 5) genes were randomly selected from the aforementioned 24 biomarkers, and the model was reconstructed and validated according to the aforementioned method. The results are shown in Table 4.

[0150] Table 4. Accuracy verification of models constructed from different gene subsets of the 24 genes.

[0151]

[0152]

[0153] As shown in the table above, the AUC corresponding to combinations of 2, 3, 4 or 5 genes randomly selected from the 24 genes is between 0.89 and 0.93, indicating that the models constructed from other random gene subsets of the 24 genes also have good predictive accuracy. These gene subsets can be used clinically to predict severe cases of COVID-19.

[0154] Example 3

[0155] The 9-marker regression model of Example 1 was used to validate the protein level risk assessment of severe COVID-19 in peripheral blood samples from 100 COVID-19 patients in the GSE157103 proteome dataset, including 50 non-critical cases and 50 ICU severe (critical) cases.

[0156] The ROC curve of the validation dataset using the 9-marker regression model of this embodiment is shown below. Figure 7 As shown, the results indicate that its AUC is 0.902, specificity is 92%, sensitivity is 86%, and accuracy is 89%.

[0157] Example 4

[0158] This embodiment uses a random subset of the aforementioned 24 genes to construct a regression model to assess the risk of severe illness in COVID-19 patients of different age groups. The dataset comes from the aforementioned 251 samples (of which 217 samples have age information). The specific method is as follows:

[0159] (1) First, the 217 samples were grouped, with each group consisting of 5-year increments. Each age point A corresponds to a group of 23 people with ages A, A+1, A+2, A+3, and A+4. For example, Age=55 represents the 55-59 age group. Each point displays the total number of people in that group. For instance, Age=55 (age 55-59) has 23 people, of whom 8 are severe cases (severe case rate ~35%). For detailed grouping information, please refer to [link / reference]. Figure 8 As shown, the incidence of severe illness has two peaks in age, one between 55 and 80 years old and the other between 35 and 45 years old.

[0160] (2) Plot the univariate ROC curve for predicting criticality indicators using age, specifically as follows: Figure 9 As shown, the AUC is 0.739, and the optimal separation point is located at (FPR, TPR) = (0.37, 0.8). According to the method of selecting the optimal threshold value on ROC as described in Example 1, the corresponding optimal threshold value is 55 years old. Therefore, the above samples are divided into two groups according to age: the middle-aged group (age 20-54 years old) and the elderly group (age 55-90 years old).

[0161] (3) Regression models were constructed using different gene subsets to verify the accuracy of critical illness risk assessment for the middle-aged group (age 20-54 years) and the elderly group (age 55-90 years).

[0162] The accuracy verification results for different gene subsets in the adult group are shown in Table 5.

[0163] Table 5. Accuracy verification of different gene subsets in the middle-aged group (20-54 years old)

[0164]

[0165] As shown in Table 5, the regression model constructed using a random subset of the 24 genes in this application has good evaluation accuracy for the adult group. The best model is the 5-genome TEX2, DGKH, UHRF1, CSGALNACT2 and KLRK1, with an AUC value of 0.977, a specificity of 89% and a sensitivity of 100%.

[0166] The results of the accuracy verification for different gene subsets in the elderly group (55-90 years old) are shown in Table 6.

[0167] Table 6. Validation of the accuracy of different gene subsets in the elderly group (55-90 years old)

[0168]

[0169]

[0170] As shown in Table 6, the optimal model consists of 9 genomes: UHRF1, SIRT5, DGKH, TREML1, IRF2BPL, EGFL7, KLRK1, CMTM5, and RASGEF1A, with an AUC value of 0.962, a specificity of 92%, and a sensitivity of 85%.

[0171] The Venn diagrams corresponding to all genes in Tables 5 and 6 are shown below. Figure 10As shown, DGKH, KLRK1, SIRT5, TREML1, and UHRF1 are severe illness risk-related genes common to both the young adult and elderly groups; CSGALNACT2, GIMAP7, PXYLP1, TEX2, TMED8, and VSIG4 are severe illness risk-related genes unique to the young adult group; and CMTM5, EGFL7, IRF2BPL, and RASGEF1A are severe illness risk-related genes unique to the elderly group.

[0172] Example 5

[0173] This embodiment provides a kit for assessing the risk of severe COVID-19. The kit includes reagents capable of quantitatively detecting the mRNA levels of the following 24 genes: VSIG4, KLRK1, TREML1, SAMD14, GIMAP5, SAMSN1, CD160, DGKH, PXYLP1, NUDT16, CMTM5, CSGALNACT2, TMED8, RASGEF1A, UHRF1, GIMAP7, SIRT5, EGFL7, JDP2, IRF2BPL, SIPA1L2, TRABD2A, CKAP4, and TEX2. The reagents include reverse transcriptase, primers, Taq enzyme, fluorescent dyes, etc.

[0174] Example 6

[0175] This embodiment provides a device for assessing the risk of severe illness in COVID-19 patients. The device includes a processor and a memory, with the memory storing a computer program executable by the processor. The method for assessing the risk of severe illness in COVID-19 patients using this device is as follows:

[0176] 1. Select peripheral blood samples from COVID-19 patients to extract mRNA.

[0177] 2. The extracted mRNA was fed into the detection device to obtain information on the quantitative expression levels of VSIG4, KLRK1, TREML1, SAMD14, GIMAP5, SAMSN1, CD160, DGKH, PXYLP1, NUDT16, CMTM5, CSGALNACT2, TMED8, RASGEF1A, UHRF1, GIMAP7, SIRT5, EGFL7, JDP2, IRF2BPL, SIPA1L2, TRABD2A, CKAP4, and TEX2.

[0178] 3. Score according to the scoring formula. Where a i b represents the level of expression of the marker. iWeights are assigned to the biomarkers, where n is the number of biomarkers and n≤24. The expression levels of the 24 genes are substituted to calculate the risk assessment scores. Then, according to one or more pre-set threshold values, the risk scores of the subjects are divided into different risk groups, and different treatment plans are considered for different risk groups.

[0179] The embodiments of the present invention have been described in detail above. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention. Furthermore, the embodiments of the present invention and the features thereof can be combined with each other unless otherwise specified.

Claims

1. Application of substances for detecting mRNA expression levels of biomarkers in the preparation of COVID-19 severe illness risk assessment reagents, wherein the biomarkers are composed of VSIG4, KLRK1, TREML1, SAMD14, GIMAP5, SAMSN1, CD160, DGKH, PXYLP1, NUDT16, CMTM5, CSGALNACT2, TMED8, RASGEF1A, UHRF1, GIMAP7, SIRT5, EGFL7, JDP2, IRF2BPL, SIPA1L2, TRABD2A, CKAP4, and TEX2; The sample used for the COVID-19 severe illness risk assessment was a peripheral blood sample.

2. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing the computer to perform the following operations: Step 1: Obtain mRNA expression level information of the biomarker as described in claim 1 from a sample from a COVID-19 patient, wherein the sample is a peripheral blood sample; Step 2: Perform mathematical correlation on the mRNA expression levels of the biomarkers to obtain a score; wherein the formula for the score is: 0.0194×VSIG4 - 0.0604×KLRK1 + 0.0342×TREML1 + 0.0196×SAMD14 + 0.0018×GIMAP5 - 0.0095×SAMSN1 + 0.0111×CD160 + 0.0387×DGKH + 0.0231×PXYLP1 - 0.0022×NUDT16 - 0.0409×CMTM5 - 0.0435×CSGALNACT2 + 0.0523×TMED8 -0.0053×RASGEF1A+0.0605×UHRF1-0.0322×GIMAP7+0.0641×SIRT5+0.0166×EGFL7+0.007×JDP2+0.0112×IRF2BPL-0.0155×SIPA1L2+0.0284×TRABD2A+0.0137×CKAP4+0.0565×TEX2, where the names of the biomarkers represent the normalized values ​​of the corresponding mRNA expression levels; The score is used to indicate the risk of severe illness in COVID-19 patients; when the score is higher than a set value, it indicates a higher risk of severe illness in COVID-19 patients.

3. A computer device, characterized in that, It includes a processor and a memory, the memory storing a computer program that can run on the processor, the processor performing the following operations when running the computer program: Step 1: Obtain mRNA expression level information of the biomarker as described in claim 1 from a sample from a COVID-19 patient, wherein the sample is a peripheral blood sample; Step 2: Perform mathematical correlation on the mRNA expression levels of the biomarkers to obtain a score; wherein the formula for the score is: 0.0194×VSIG4 - 0.0604×KLRK1 + 0.0342×TREML1 + 0.0196×SAMD14 + 0.0018×GIMAP5 - 0.0095×SAMSN1 + 0.0111×CD160 + 0.0387×DGKH + 0.0231×PXYLP1 - 0.0022×NUDT16 - 0.0409×CMTM5 - 0.0435×CSGALNACT2 + 0.0523×TMED8 -0.0053×RASGEF1A+0.0605×UHRF1-0.0322×GIMAP7+0.0641×SIRT5+0.0166×EGFL7+0.007×JDP2+0.0112×IRF2BPL-0.0155×SIPA1L2+0.0284×TRABD2A+0.0137×CKAP4+0.0565×TEX2, where the names of the biomarkers represent the normalized values ​​of the corresponding mRNA expression levels; The score is used to indicate the risk of severe illness in COVID-19 patients. When the score is higher than a set value, it indicates that the risk of severe illness in COVID-19 patients is higher.

Citation Information

Patent Citations

  • Prognostic marker related to radiotherapy sensitivity of head and neck squamous cell carcinoma and application of prognosis and radiotherapy risk assessment method

    CN115198019A

  • Spatial detection of SARS-COV-2 using templated ligation

    WO2022271820A1