Methods for diagnosing bacterial and viral infections
Biomarker-based gene expression analysis using TSPO, EMR1, NINJ2, ACPP, and others addresses the challenge of distinguishing bacterial from viral infections, enhancing diagnostic accuracy and reducing antibiotic misuse.
Patent Information
- Application Number
- JP2023061560
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-06-07
- Filing Date
- 2023-04-05
- Publication Date
- 2025-10-07
- Estimated Expiration
- 2037-06-05
AI Technical Summary
Current diagnostic methods are inadequate for accurately distinguishing between bacterial and viral infections, leading to inappropriate antibiotic use and increased antibiotic resistance, with a need for highly sensitive and specific tests.
Utilizing biomarkers such as TSPO, EMR1, NINJ2, ACPP, and others to determine the presence of bacterial or viral infections through gene expression analysis, combined with clinical parameters for prognosis and treatment monitoring.
Provides a highly sensitive and specific diagnostic method for distinguishing between bacterial and viral infections, reducing inappropriate antibiotic use and improving patient outcomes.
Smart Images

Figure 0007750895000066 
Figure 0007750895000067 
Figure 0007750895000068
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefit of U.S. Provisional Application No. 62 / 346,962, filed June 7, 2016, the application of which is incorporated herein by reference.
[0002] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT This invention was made with government support under Contract Nos. AI109662 and AI057229 awarded by the National Institutes of Health. The U.S. Federal Government has certain rights in this invention.
[0003] Technical Field The present invention relates generally to methods for diagnosing bacterial and viral infections. In particular, the present invention relates to the use of biomarkers that can distinguish whether a patient with acute inflammation has a bacterial or viral infection. [Background technology]
[0004] Early and accurate diagnosis of infection is key to improving patient outcomes and reducing antibiotic resistance. Mortality from bacterial sepsis increases by 8% for every hour that antibiotics are delayed. 1 Prescribing antibiotics to patients without bacterial infection increases morbidity and antimicrobial resistance. The rate of inappropriate antibiotic prescribing in the hospital setting is estimated to be 30-50%, but this is expected to improve with improved diagnostics. 2,3 An astonishing 95% of patients treated with antibiotics for suspected typhoid fever have negative cultures. 4 Currently, there are no optimal point-of-care diagnostics that can universally determine the presence and type of infection. Therefore, the U.S. government established the National Action Plan for Combating Antibiotic-Resistant Bacteria, which calls for "point-of-need diagnostic tests that rapidly distinguish between bacterial and viral infections." 5.
[0005] New PCR-based molecular diagnostic methods can profile pathogens directly from blood cultures. 6 However, such methods rely on the presence of sufficient numbers of pathogens in the blood. Furthermore, PCR-based molecular diagnostic methods are limited to detecting a discrete range of pathogens. As a result, there is growing interest in molecular diagnostic methods that profile the host's genetic response. These molecular diagnostic methods are diagnostic methods that can distinguish the presence of infection compared to inflamed but non-infected patients, and among them, our 11-gene "sepsis metascore" 7 (Sepsis MetaScore:SMS) (validated across multiple cohorts) 8 ) and other diagnostic methods 9,10 Other research groups are focusing on gene sets that can distinguish between types of infection, such as bacterial versus viral infections. 11~13 Tsalik et al. described a model that distinguished between all three classes (i.e., non-infected patients and patients with bacterial or viral disease), but this model required the measurement of 122 probes. 14 We have previously described a "metaviral signature" that describes a general response to viral infection but contains too many genes (396) for clinical application. 15 Overall, the field shows great promise, but diagnostic methods for infection by host gene expression have not yet been translated into clinical practice.
[0006] Data from these biomarker studies, as well as numerous other genome-wide expression studies in sepsis and acute infection, have been published and deposited for further study in public databases such as the NIH Gene Expression Omnibus (GEO) and EBI ArrayExpress. These data represent a largely untapped resource that can be used for both biomarker discovery and validation. We demonstrate that our integrated multi-cohort analysis of gene expression has the potential to significantly improve the efficacy and safety of sepsis biomarkers in acute infections. 7 , specific types of viral infections 15 , and active tuberculosis 16 Furthermore, these data are also useful as benchmark and validation tools for novel host gene expression diagnostic methods. 17 However, because technical differences between studies preclude direct comparison of diagnostic scores between cohorts, such validation in published data has so far been limited to cohorts containing at least two classes of interest (i.e., classes that allow direct comparison between classes). Summary of the Invention [Problem to be solved by the invention]
[0007] There remains a need for a highly sensitive and specific diagnostic test that can distinguish between bacterial and viral infections. [Means for solving the problem]
[0008] The present invention relates to the use of biomarkers that can determine whether a patient with acute inflammation has a bacterial or viral infection. These biomarkers can be used alone or in combination with one or more additional biomarkers or relevant clinical parameters in the prognosis, diagnosis, or treatment monitoring of the infection.
[0009] In one embodiment, the invention provides a method of devising a classification for use in diagnosing an infection in a patient, comprising the step of: (a) measuring the expression levels of at least two biomarkers in a biological sample from the patient, wherein the at least two biomarkers are selected from one or both of a first set of biomarkers where high expression levels are indicative of a bacterial infection and a second set of biomarkers where high expression levels are indicative of a viral infection, wherein the first set of biomarkers includes at least one of TSPO, EMR1, NINJ2, ACPP, TBXAS1, PGD, S100A12, SORT1, TNIP1, RAB31, SLC12A9, PLP2, IMPA2, GPAA1, LTA4H, RTN3, CETP, TALD01, HK3, ACAA1, CAT, DOK3, SORL1, PYGL, DYSF, TWF2, TKT, CTSB, FLII, PROS1, NRD1, STAT5B, CYBRD1, PTAFR, and LAPTM5. The second set of biomarkers was OAS1, IFIT1, SAMD9, ISG15, HERC5, DDX60, HESX1, IFI6, MX1, OASL, LAX1, IFIT5, IFIT3, KCTD14, OAS2, RTP4, PARP12, LY6E, ADA, IFI44L, IFI27, RSAD2, IFI44, OAS3, IFIH1, SIGLEC1, JUP, STAT1, CUL1, DNMT1, IFIT2, CHST12, ISG20, DHX5 (b) devising a classification or generation algorithm that can use the expression levels of the biomarkers to determine the presence or probability of a bacterial or viral infection in a patient; and (c) applying the algorithm to diagnose the patient as having, or likely to have, a bacterial or viral infection.
[0010] In one embodiment, the invention is directed to a method for diagnosing an infection in a patient, the method comprising analyzing expression levels of at least two genes, where the at least two genes are predictive of viral or bacterial infection, and the expression levels of the at least two genes provide an area under the curve for predicting viral or bacterial infection of at least 0.80; and diagnosing the patient as having a bacterial or viral infection.
[0011] In one embodiment, the invention is directed to a method for diagnosing and treating an infection in a patient, the method comprising the steps of: (a) obtaining a biological sample from the patient; (b) measuring the expression levels of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, and CTSB in the biological sample; (c) analyzing the expression level of each biomarker along with the biomarker's respective reference range, wherein an increase in the expression level of the biomarkers IFI27, JUP, LAX1 compared to the reference range for the biomarker in a control subject is indicative of the patient having a viral infection, and an increase in the expression level of the biomarkers HK3, TNIP1, GPAA1, CTSB compared to the reference range for the biomarker in a control subject is indicative of the patient having a bacterial infection; and (d) administering to the patient an effective amount of an antiviral agent if the patient is diagnosed with a viral infection, or administering to the patient an effective amount of an antibiotic agent if the patient is diagnosed with a bacterial infection.
[0012] In any embodiment, the biological sample may comprise whole blood or peripheral blood mononuclear cells (PBMCs).
[0013] In any embodiment, the levels of the biomarkers can be compared to time-matched reference values for infected or non-infected subjects.
[0014] In any embodiment, the method may include calculating a bacterial / viral metascore for the patient based on the levels of the biomarkers, where a positive bacterial / viral metascore for the patient indicates that the patient has a viral infection and a negative bacterial / viral metascore for the patient indicates that the patient has a bacterial infection.
[0015] In any embodiment, the method may include normalizing the data using COCONUT normalization.
[0016] In any embodiment, the patient can be a human.
[0017] In any embodiment, the step of measuring the levels of the plurality of biomarkers may include performing microarray analysis, polymerase chain reaction (PCR), reverse transcriptase polymerase chain reaction (RT-PCR), Northern blot, or serial analysis of gene expression (SAGE).
[0018] In one embodiment, the present invention provides a method of diagnosing and treating a patient with inflammation, comprising the steps of: (a) obtaining a biological sample from the patient; (b) measuring the expression levels of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, CTSB, CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 in the biological sample; and (c) first determining the expression levels of each biomarker. and analyzing the expression levels of the biomarkers, along with their respective reference ranges, wherein an increase in the expression levels of the biomarkers CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, and C3AR1, and a decrease in the expression levels of the biomarkers KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1, compared to the reference ranges of the biomarkers for uninfected control subjects, indicates that the patient has an infection ... (d) if the patient is diagnosed with an infection, further analyzing the expression levels of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, and CTSB, wherein the absence of differential expression of the biomarkers CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 indicates that the patient does not have an infection; and (e) if the patient is diagnosed with an infection, further analyzing the expression levels of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, and CTSB, wherein the absence of differential expression of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, and CTSB indicates that the patient does not have an infection. wherein an increase in the expression levels of JUP, LAX1 compared to the reference range of the biomarkers for the control subject indicates that the patient has a viral infection, and an increase in the expression levels of the biomarkers HK3, TNIP1, GPAA1, CTSB compared to the reference range of the biomarkers for the control subject indicates that the patient has a bacterial infection; and (e) administering, or if the patient is diagnosed with a bacterial infection, administering an effective amount of an antibiotic to the patient.
[0019] In any embodiment, the method may include calculating a sepsis metascore for the patient, where a sepsis metascore above the reference range for non-infected control subjects indicates that the patient has an infection, and a sepsis metascore within the reference range for non-infected control subjects indicates that the patient has a non-infectious inflammatory condition.
[0020] In any embodiment, the method may include calculating a bacterial / viral metascore for the patient if the patient is diagnosed with an infection, where a positive bacterial / viral metascore for the patient indicates that the patient has a viral infection and a negative bacterial / viral metascore for the patient indicates that the patient has a bacterial infection.
[0021] In any embodiment, the levels of the biomarkers can be compared to time-matched reference values for infected or non-infected subjects.
[0022] In any embodiment, the non-infectious inflammatory condition may be selected from the group of systemic inflammatory response syndrome (SIRS), autoimmune disorders, traumatic injury, and surgery.
[0023] In any embodiment, the patient can be a human.
[0024] In any embodiment, the step of measuring the level of the biomarker may include performing microarray analysis, polymerase chain reaction (PCR), reverse transcriptase polymerase chain reaction (RT-PCR), Northern blot, or serial analysis of gene expression (SAGE).
[0025] In one embodiment, the present invention is directed to a kit comprising agents for measuring the levels of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, and CTSB.
[0026] In any embodiment, the kit may include agents for measuring the levels of the biomarkers CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1.
[0027] In any embodiment, the kit may include a microarray.
[0028] In any embodiment, the microarray may comprise oligonucleotides that hybridize to IFI27 polynucleotides, oligonucleotides that hybridize to JUP polynucleotides, oligonucleotides that hybridize to LAX1 polynucleotides, oligonucleotides that hybridize to HK3 polynucleotides, oligonucleotides that hybridize to TNIP1 polynucleotides, oligonucleotides that hybridize to GPAA1 polynucleotides, and oligonucleotides that hybridize to CTSB polynucleotides.
[0029] In any embodiment, the microarray may comprise oligonucleotides that hybridize to CEACAM1 polynucleotides, oligonucleotides that hybridize to ZDHHC19 polynucleotides, oligonucleotides that hybridize to C9orf95 polynucleotides, oligonucleotides that hybridize to GNA15 polynucleotides, oligonucleotides that hybridize to BATF polynucleotides, oligonucleotides that hybridize to C3AR1 polynucleotides, oligonucleotides that hybridize to KIAA1370 polynucleotides, oligonucleotides that hybridize to TGFBI polynucleotides, oligonucleotides that hybridize to MTCH1 polynucleotides, oligonucleotides that hybridize to RPGRIP1 polynucleotides, and oligonucleotides that hybridize to HLA-DPB1 polynucleotides.
[0030] In any embodiment, the kit may include information in electronic or paper form with instructions for correlating the detected level of each biomarker with sepsis.
[0031] In one embodiment, the method is directed to a computer-implemented method for diagnosing a patient suspected of having an infection, wherein the computer performs the steps of: (a) receiving input patient data including values for levels of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, and CTSB in a biological sample from the patient; b) analyzing the level of each of the biomarkers and comparing it to the biomarker's respective reference value range; c) calculating a bacterial / viral metascore for the patient based on the levels of the biomarkers, wherein a positive bacterial / viral metascore for the patient indicates that the patient has a viral infection, and a negative bacterial / viral metascore for the patient indicates that the patient has a bacterial infection; and (d) displaying information regarding the patient's diagnosis.
[0032] In any embodiment, the biological sample may comprise whole blood or peripheral blood mononuclear cells (PBMCs).
[0033] In one embodiment, the present invention is directed to a diagnostic system for performing a computer-implemented method, the diagnostic system including: (a) a storage component for storing data, the storage component having instructions stored therein for determining a patient's diagnosis; (b) a computer processor for processing the data, the computer processor coupled to the storage component and configured to receive patient data and execute the instructions stored in the storage component to analyze the patient data according to one or more algorithms; and (c) a display component for displaying information regarding the patient's diagnosis.
[0034] In any embodiment, the storage component may include instructions for calculating a bacterial / viral metascore.
[0035] In one embodiment, the invention provides a computer-implemented method for diagnosing a patient with inflammation, the method comprising the steps of: a) receiving input patient data, the input patient data including values for the levels of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, CTSB, CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 in a biological sample from the patient; b) analyzing the level of each of the biomarkers and comparing it to the biomarker's respective reference value range; and c) calculating a sepsis metascore for the patient. a) calculating a bacterial / viral metascore for the patient if the sepsis score indicates that the patient has an infection, wherein a sepsis metascore above the reference range for non-infected control subjects indicates that the patient has an infection, and a sepsis metascore within the reference range for non-infected control subjects indicates that the patient has a non-infectious inflammatory condition; b) calculating a bacterial / viral metascore for the patient if the sepsis score indicates that the patient has an infection, wherein a positive bacterial / viral metascore for the patient indicates that the patient has a viral infection, and a negative bacterial / viral metascore for the patient indicates that the patient has a bacterial infection; and c) displaying information regarding the patient's diagnosis.
[0036] In any embodiment, the biological sample may comprise whole blood or peripheral blood mononuclear cells (PBMCs).
[0037] In one embodiment, the present invention is directed to a diagnostic system for performing a computer-implemented method, the diagnostic system including: a) a storage component for storing data, the storage component having instructions stored therein for determining a patient's diagnosis; b) a computer processor for processing the data, the computer processor coupled to the storage component and configured to receive patient data and execute the instructions stored in the storage component to analyze the patient data according to one or more algorithms; and c) a display component for displaying information regarding the patient's diagnosis.
[0038] In any embodiment, the storage component may include instructions for calculating a sepsis metascore and a bacterial / viral metascore.
[0039] In one embodiment, the present invention provides a method for diagnosing and treating an infection in a patient, comprising the steps of: a) obtaining a biological sample from the patient; and b) measuring the expression levels of a set of viral response genes and a set of bacterial response genes in the biological sample, wherein the set of viral response genes is selected from the group consisting of OAS2, CUL1, ISG15, CHST12, IFIT1, SIGLEC1, ADA, MX1, RSAD2, IFI44L, GZMB, KCTD14, LY6E, IFI44, HESX1, OASL, OAS1, OAS3, EIF2AK2, DDX60, DNMT1, HERC5, IFIH1, SAMD9, IFI6, IFIT3, IFIT5, XAF1, ISG20, PARP12, IFIT2, DHX58, and STAT1. group, wherein the set of bacterial response genes comprises one or more genes selected from the group: SLC12A9, ACPP, STAT5B, EMR1, FLII, PTAFR, NRD1, PLP2, DYSF, TWF2, SORT1, TSPO, TBXAS1, ACAA1, S100A12, PGD, LAPTM5, NINJ2, DOK3, SORL1, RAB31, IMPA2, LTA4H, TALDO1, TKT, PYGL, CETP, PROS1, RTN3, CAT, CYBRD1; and c) analyzing the expression level of each biomarker along with their respective reference ranges in uninfected control subjects, wherein the differential expression of the viral response genes is compared to the reference values.
[0040] In any embodiment, the set of viral response genes and the set of bacterial response genes are selected from the group consisting of: a) a set of viral response genes comprising OAS2 and CUL1, and a set of bacterial response genes comprising SLC12A9, ACPP, STAT5B; b) a set of viral response genes comprising ISG15 and CHST12, and a set of bacterial response genes comprising EMR1 and FLII; c) a set of viral response genes comprising IFIT1, SIGLEC1, and ADA, and a set of bacterial response genes comprising PTAFR, NRD1, PLP2; d) a set of viral response genes comprising MX1, and a set of bacterial response genes comprising DYSF, TWF2; e) a set of viral response genes comprising RSAD2, and and a set of bacterial response genes comprising SORT1 and TSPO; f) a set of viral response genes comprising IFI44L, GZMB, and KCTD14, and a set of bacterial response genes comprising TBXAS1, ACAA1, and S100A12; 16g) a set of viral response genes comprising LY6E, and a set of bacterial response genes comprising PGD and LAPTM5; h) a set of viral response genes comprising IFI44, HESX1, and OASL, and a set of bacterial response genes comprising NINJ2, DOK3, SORL1, and RAB31; and i) a set of viral response genes comprising OAS1, and a set of bacterial response genes comprising IMPA2 and LTA4H.
[0041] In any embodiment, the biological sample may comprise whole blood or peripheral blood mononuclear cells (PBMCs).
[0042] In any embodiment, the levels of the biomarkers can be compared to time-matched reference values for infected or non-infected subjects.
[0043] In any embodiment, the method may include calculating a bacterial / viral metascore for the patient based on the levels of the biomarkers, where a positive bacterial / viral metascore for the patient indicates that the patient has a viral infection and a negative bacterial / viral metascore for the patient indicates that the patient has a bacterial infection.
[0044] In any embodiment, the method comprises the steps of measuring the expression levels of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, CTSB, CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 in the biological sample; and analyzing the expression level of each biomarker along with the respective reference ranges for the biomarkers, wherein the expression levels of the biomarkers CEACAM1, CEACAM2, CEACAM3, CEACAM4, CEACAM5, CEACAM6, CEACAM7, CEACAM8, CEACAM9, CEACAM10, CEACAM11, CEACAM12, CEACAM13, CEACAM14, CEACAM15, CEACAM16, CEACAM17, CEACAM18, CEACAM19, CEACAM19, CEACAM16, CEACAM18, CEACAM19, CEACAM19, CEACAM19, CEACAM11, CEACAM15, CEACAM16, CEACAM18, CEACAM19, CEACAM19, CEACAM11, CEACAM12, CEACAM13, CEACAM14, CEACAM15, CEACAM15, CEACAM16, CEACAM18, CEACAM19, CEACAM15, CEACAM19, CEACAM15, CEACAM15, CEACAM16, CEACAM18, CEACAM19, CEACAM19, CEACAM11, CEACAM15 ...5, CEACAM15, C The method may include steps in which an increase in the expression levels of ZDHHC19, C9orf95, GNA15, BATF, and C3AR1 and a decrease in the expression levels of the biomarkers KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 indicates that the patient has an infection, and an absence of differential expression of the biomarkers CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 compared to uninfected control subjects indicates that the patient does not have an infection.
[0045] In one embodiment, the present invention provides: (a) a set of viral response genes comprising OAS2 and CUL1, and a set of bacterial response genes comprising SLC12A9, ACPP, STAT5B; (b) a set of viral response genes comprising ISG15 and CHST12, and a set of bacterial response genes comprising EMR1 and FLII; b) a set of viral response genes comprising IFIT1, SIGLEC1, and ADA, and a set of bacterial response genes comprising PTAFR, NRD1, PLP2; c) a set of viral response genes comprising MX1, and a set of bacterial response genes comprising DYSF and TWF2; d) a set of viral response genes comprising RSAD2, and a set of bacterial response genes comprising SORT1 and TSPO; e) IFI44L The present invention relates to a kit comprising an agent for measuring the expression levels of a set of viral response genes and a set of bacterial response genes selected from the group consisting of: f) a set of viral response genes comprising LY6E and a set of bacterial response genes comprising PGD and LAPTM5; g) a set of viral response genes comprising IFI44, HESX1, and OASL and a set of bacterial response genes comprising NINJ2, DOK3, SORL1, and RAB31; and h) a set of viral response genes comprising OAS1 and a set of bacterial response genes comprising IMPA2 and LTA4H.
[0046] In any embodiment, the kit may include a microarray.
[0047] In one embodiment, the invention provides a computer-implemented method for diagnosing a patient suspected of having an infection, the method comprising the steps of: a) receiving input patient data comprising values for expression levels in the biological sample of a set of viral response genes and a set of bacterial response genes, wherein the set of viral response genes comprises one or more genes selected from the group consisting of OAS2, CUL1, ISG15, CHST12, IFIT1, SIGLEC1, ADA, MX1, RSAD2, IFI44L, GZMB, KCTD14, LY6E, IFI44, HESX1, OASL, OAS1, OAS3, EIF2AK2, DDX60, DNMT1, HERC5, IFIH1, SAMD9, IFI6, IFIT3, IFIT5, XAF1, ISG20, PARP12, IFIT2, DHX58, STAT1; and the patient comprises one or more genes selected from the group consisting of SLC12A9, ACPP, STAT5B, EMR1, FLII, PTAFR, NRD1, PLP2, DYSF, TWF2, SORT1, TSPO, TBXAS1, ACAA1, S100A12, PGD, LAPTM5, NINJ2, DOK3, SORL1, RAB31, IMPA2, LTA4H, TALDO1, TKT, PYGL, CETP, PROS1, RTN3, CAT, and CYBRD1; (b) analyzing the expression levels of the set of viral response genes and the set of bacterial response genes and comparing them with respective reference ranges for uninfected control subjects; (c) calculating a bacterial / viral metascore for the patient based on the expression levels of the set of viral response genes and the set of bacterial response genes; and (d) displaying information regarding the diagnosis of the patient.
[0048] In one embodiment, the present invention is directed to a diagnostic system for performing a computer-implemented method, the diagnostic system including: a) a storage component for storing data, the storage component having instructions stored therein for determining a patient's diagnosis; b) a computer processor for processing the data, the computer processor coupled to the storage component and configured to receive patient data and execute the instructions stored in the storage component to analyze the patient data according to one or more algorithms; and c) a display component for displaying information regarding the patient's diagnosis.
[0049] In one embodiment, the present invention provides a method for diagnosing an infection in a patient, comprising the step of: (a) measuring the expression levels of at least two biomarkers in a biological sample from the patient, wherein the at least two biomarkers are selected from one or both of a first set of biomarkers, high expression levels of which are indicative of a bacterial infection, and a second set of biomarkers, high expression levels of which are indicative of a viral infection, wherein the first set of biomarkers includes TSPO, EMR1, NINJ2, ACPP, TBXAS1, PGD, S100A12, SORT1, TNIP1, RAB31, SLC12A9, PLP2, IMPA2, GPAA1, LTA4H, RTN3, CETP, TALD01, HK3, ACAA1, CAT, DOK3, SORL1, PYGL, DYSF, TWF2, TKT, CTSB, FLII, PROS1, NRD1, STAT5B, CYBRD and a second set of biomarkers comprising at least one of OAS1, IFIT1, SAMD9, ISG15, HERC5, DDX60, HESX1, IFI6, MX1, OASL, LAX1, IFIT5, IFIT3, KCTD14, OAS2, RTP4, PARP12, LY6E, ADA, IFI44L, IFI27, RSAD2, IFI44, OAS3, IFIH1, SIGLEC1, JUP, STAT1, CUL1, DNMT1, IFIT2, CHST12, ISG20, DHX58, EIF2AK2, XAF1, and GZMB; and (b) analyzing the expression level of each biomarker along with the respective reference ranges for the biomarkers to determine viral or bacterial infection.
[0050] In any embodiment, the method may include administering to the patient an effective amount of an antiviral agent if the patient is diagnosed with a viral infection, or an effective amount of an antibiotic agent if the patient is diagnosed with a bacterial infection.
[0051] In any embodiment, the expression levels of at least two biomarkers may result in an area under the curve of at least 0.80.
[0052] In any embodiment, the first set of biomarkers may include at least one of HK3, TNIP1, GPAA1, and CTSB, and the second set of biomarkers may include at least one of IFI27, JUP, and LAX1.
[0053] In any embodiment, the biological sample may comprise whole blood or peripheral blood mononuclear cells (PBMCs).
[0054] In any embodiment, the levels of the biomarkers can be compared to time-matched reference values for infected or non-infected subjects.
[0055] In any embodiment, the method may include calculating a bacterial / viral metascore for the patient based on the levels of the biomarkers, where a positive bacterial / viral metascore for the patient indicates that the patient has a viral infection and a negative bacterial / viral metascore for the patient indicates that the patient has a bacterial infection.
[0056] In any embodiment, the method may include normalizing the data using COCONUT normalization, which includes (a) separating data from multiple cohorts into healthy and diseased components; (b) co-normalizing the healthy component using ComBat co-normalization without using covariates; (c) obtaining ComBat estimated parameters for each dataset for the healthy component; and (d) applying the ComBat estimated parameters to the diseased component.
[0057] In any embodiment, the patient can be a human.
[0058] In any embodiment, the step of measuring the levels of the plurality of biomarkers may include performing microarray analysis, polymerase chain reaction (PCR), reverse transcriptase polymerase chain reaction (RT-PCR), Northern blot, or serial analysis of gene expression (SAGE).
[0059] In one embodiment, the present invention provides a method of diagnosing and treating a patient with inflammation, comprising the steps of: (a) measuring the expression levels of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, CTSB, CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 in a biological sample from the patient; and (b) first analyzing the expression level of each biomarker, along with the respective reference ranges for the biomarkers, and comparing the expression levels to the reference ranges for the biomarkers for uninfected control subjects. wherein increased expression levels of the biomarkers CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, and C3AR1 and decreased expression levels of the biomarkers KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 indicate that the patient has an infection, and wherein the absence of differential expression of the biomarkers CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 compared to uninfected control subjects indicates that the patient does not have an infection;(c) further analyzing the expression levels of at least two biomarkers in the patient's biological sample to determine bacterial or viral infection, wherein the at least two biomarkers are selected from one or both of a first set of biomarkers, high expression levels of which are indicative of bacterial infection, and a second set of biomarkers, high expression levels of which are indicative of viral infection, wherein the first set of biomarkers is selected from TSPO, EMR1, NINJ2, ACPP, TBXAS1, PGD, S100A12, SORT1, TNIP1, RAB31, SLC12A9, PLP2, IMPA2, GPAA1, LTA4H, RTN3, CETP, TALD01, HK3, ACAA1, CAT, DOK3, SORL1, PYGL, DYSF, TWF2, and a second set of biomarkers including at least one of OAS1, IFIT1, SAMD9, ISG15, HERC5, DDX60, HESX1, IFI6, MX1, OASL, LAX1, IFIT5, IFIT3, KCTD14, OAS2, RTP4, PARP12, LY6E, ADA, IFI44L, IFI27, RSAD2, IFI44, OAS3, IFIH1, SIGLEC1, JUP, STAT1, CUL1, DNMT1, IFIT2, CHST12, ISG20, DHX58, EIF2AK2, XAF1, and GZMB;
[0060] In any embodiment, the method may include calculating a sepsis metascore for the patient, where a sepsis metascore above the reference range for non-infected control subjects indicates that the patient has an infection, and a sepsis metascore within the reference range for non-infected control subjects indicates that the patient has a non-infectious inflammatory condition.
[0061] In any embodiment, the method may include calculating a bacterial / viral metascore for the patient if the patient is diagnosed with an infection, where a positive bacterial / viral metascore for the patient indicates that the patient has a viral infection and a negative bacterial / viral metascore for the patient indicates that the patient has a bacterial infection.
[0062] In any embodiment, the levels of the biomarkers can be compared to time-matched reference values for infected or non-infected subjects.
[0063] In any embodiment, the non-infectious inflammatory condition may be selected from the group of systemic inflammatory response syndrome (SIRS), autoimmune disorders, traumatic injury, and surgery.
[0064] In any embodiment, the patient can be a human.
[0065] In any embodiment, the step of measuring the level of the biomarker may include performing microarray analysis, polymerase chain reaction (PCR), reverse transcriptase polymerase chain reaction (RT-PCR), Northern blot, or serial analysis of gene expression (SAGE).
[0066] In one embodiment, the method is a kit comprising agents for measuring the levels of at least two biomarkers in a biological sample from a patient, wherein the at least two biomarkers are selected from one or both of a first set of biomarkers where high expression levels are indicative of a bacterial infection and a second set of biomarkers where high expression levels are indicative of a viral infection, wherein the first set of biomarkers includes TSPO, EMR1, NINJ2, ACPP, TBXAS1, PGD, S100A12, SORT1, TNIP1, RAB31, SLC12A9, PLP2, IMPA2, GPAA1, LTA4H, RTN3, CETP, TALD01, HK3, ACAA1, CAT, DOK3, SORL1, PYGL, DYSF, TWF2, and a second set of biomarkers comprising at least one of OAS1, IFIT1, SAMD9, ISG15, HERC5, DDX60, HESX1, IFI6, MX1, OASL, LAX1, IFIT5, IFIT3, KCTD14, OAS2, RTP4, PARP12, LY6E, ADA, IFI44L, IFI27, RSAD2, IFI44, OAS3, IFIH1, SIGLEC1, JUP, STAT1, CUL1, DNMT1, IFIT2, CHST12, ISG20, DHX58, EIF2AK2, XAF1, and GZMB.
[0067] In any embodiment, the kit may include agents for measuring the levels of the biomarkers CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1.
[0068] In any embodiment, the kit may include a microarray.
[0069] In any embodiment, the microarray may comprise oligonucleotides that hybridize to IFI27 polynucleotides, oligonucleotides that hybridize to JUP polynucleotides, oligonucleotides that hybridize to LAX1 polynucleotides, oligonucleotides that hybridize to HK3 polynucleotides, oligonucleotides that hybridize to TNIP1 polynucleotides, oligonucleotides that hybridize to GPAA1 polynucleotides, and oligonucleotides that hybridize to CTSB polynucleotides.
[0070] In any embodiment, the microarray may comprise oligonucleotides that hybridize to CEACAM1 polynucleotides, oligonucleotides that hybridize to ZDHHC19 polynucleotides, oligonucleotides that hybridize to C9orf95 polynucleotides, oligonucleotides that hybridize to GNA15 polynucleotides, oligonucleotides that hybridize to BATF polynucleotides, oligonucleotides that hybridize to C3AR1 polynucleotides, oligonucleotides that hybridize to KIAA1370 polynucleotides, oligonucleotides that hybridize to TGFBI polynucleotides, oligonucleotides that hybridize to MTCH1 polynucleotides, oligonucleotides that hybridize to RPGRIP1 polynucleotides, and oligonucleotides that hybridize to HLA-DPB1 polynucleotides.
[0071] In any embodiment, the kit may include information in electronic or paper form with instructions for correlating the detected level of each biomarker with sepsis.
[0072] In one embodiment, the invention provides a computer-implemented method for diagnosing a patient suspected of having an infection, the method comprising the steps of: (a) receiving input patient data including values for the levels of at least two biomarkers in a biological sample of the patient, the at least two biomarkers being selected from one or both of a first set of biomarkers, high expression levels of which are indicative of a bacterial infection, and a second set of biomarkers, high expression levels of which are indicative of a viral infection, the first set of biomarkers being selected from T and a second set of biomarkers comprising at least one of SPO, EMR1, NINJ2, ACPP, TBXAS1, PGD, S100A12, SORT1, TNIP1, RAB31, SLC12A9, PLP2, IMPA2, GPAA1, LTA4H, RTN3, CETP, TALD01, HK3, ACAA1, CAT, DOK3, SORL1, PYGL, DYSF, TWF2, TKT, CTSB, FLII, PROS1, NRD1, STAT5B, CYBRD1, PTAFR, and LAPTM5, and and at least one of the following biomarkers in the sample: OAS1, IFIT1, SAMD9, ISG15, HERC5, DDX60, HESX1, IFI6, MX1, OASL, LAX1, IFIT5, IFIT3, KCTD14, OAS2, RTP4, PARP12, LY6E, ADA, IFI44L, IFI27, RSAD2, IFI44, OAS3, IFIH1, SIGLEC1, JUP, STAT1, CUL1, DNMT1, IFIT2, CHST12, ISG20, DHX58, EIF2AK2, XAF1, and GZMB. (b) analyzing the level of each of the biomarkers and comparing it to a respective reference value range for the biomarker; (c) calculating a bacterial / viral metascore for the patient based on the levels of the biomarkers, wherein a positive bacterial / viral metascore for the patient indicates that the patient has a viral infection and a negative bacterial / viral metascore for the patient indicates that the patient has a bacterial infection; and (d) displaying information regarding the patient's diagnosis.
[0073] In any embodiment, the biological sample may comprise whole blood or peripheral blood mononuclear cells (PBMCs).
[0074] In one embodiment, the present invention is directed to a diagnostic system for performing a computer-implemented method, the diagnostic system including: (a) a storage component for storing data, the storage component having instructions stored therein for determining a patient's diagnosis; (b) a computer processor for processing the data, the computer processor coupled to the storage component and configured to receive patient data and execute the instructions stored in the storage component to analyze the patient data according to one or more algorithms; and (c) a display component for displaying information regarding the patient's diagnosis.
[0075] In any embodiment, the storage component may include instructions for calculating a bacterial / viral metascore.
[0076] In one embodiment, the invention provides a computer-implemented method for diagnosing a patient with inflammation, the method comprising the steps of: (a) receiving input patient data having values for levels of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, CTSB, CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 in a biological sample from the patient; (b) analyzing the level of each of the biomarkers and comparing it to the biomarker's respective reference value range; and (c) calculating a sepsis metascore for the patient. (d) if the sepsis score indicates that the patient has an infection, calculating a bacterial / viral metascore for the patient, wherein a positive bacterial / viral metascore for the patient indicates that the patient has a viral infection, and a negative bacterial / viral metascore for the patient indicates that the patient has a bacterial infection; and (e) displaying information regarding the patient's diagnosis.
[0077] In any embodiment, the biological sample may comprise whole blood or peripheral blood mononuclear cells (PBMCs).
[0078] In one embodiment, the present invention is directed to a diagnostic system for performing a computer-implemented method, the diagnostic system including: (a) a storage component for storing data, the storage component having instructions stored therein for determining a patient's diagnosis; (b) a computer processor for processing the data, the computer processor coupled to the storage component and configured to receive patient data and execute the instructions stored in the storage component to analyze the patient data according to one or more algorithms; and (c) a display component for displaying information regarding the patient's diagnosis.
[0079] In any embodiment, the storage component may include instructions for calculating a sepsis metascore and a bacterial / viral metascore.
[0080] In one embodiment, the present invention provides a method for diagnosing and treating an infection in a patient, comprising the steps of: (a) obtaining a biological sample from the patient; and (b) measuring the expression level of an arbitrary set of at least two biomarkers in the patient's biological sample, wherein the at least two biomarkers are selected from one or both of a first set of biomarkers, high expression levels of which are indicative of a bacterial infection, and a second set of biomarkers, high expression levels of which are indicative of a viral infection. the first set of markers is selected from the group consisting of at least one of TSPO, EMR1, NINJ2, ACPP, TBXAS1, PGD, S100A12, SORT1, TNIP1, RAB31, SLC12A9, PLP2, IMPA2, GPAA1, LTA4H, RTN3, CETP, TALD01, HK3, ACAA1, CAT, DOK3, SORL1, PYGL, DYSF, TWF2, TKT, CTSB, FLII, PROS1, NRD1, STAT5B, CYBRD1, PTAFR, and LAPTM5; and the second set of biomarkers included OAS1, IFIT1, SAMD9, ISG15, HERC5, DDX60, HESX1, IFI6, MX1, OASL, LAX1, IFIT5, IFIT3, KCTD14, OAS2, RTP4, PARP12, LY6E, ADA, IFI44L, IFI27, RSAD2, IFI44, OAS3, IFIH1, SIGLEC1, JUP, STAT1, CUL1, DNMT1, IFIT2, CHST12, ISG20, DHX58, EIF2 (c) analyzing the expression level of each biomarker along with their respective reference ranges for non-infected control subjects, wherein differential expression of viral response genes compared to the reference ranges for non-infected control subjects indicates that the patient has a viral infection, and differential expression of bacterial response genes compared to the reference ranges for non-infected control subjects indicates that the patient has a bacterial infection.
[0081] In any embodiment, the set of viral response genes and bacterial response genes is selected from the group consisting of: (a) a set of viral response genes comprising OAS2 and CUL1, and a set of bacterial response genes comprising SLC12A9, ACPP, STAT5B; (b) a set of viral response genes comprising ISG15 and CHST12, and a set of bacterial response genes comprising EMR1 and FLII; (c) a set of viral response genes comprising IFIT1, SIGLEC1, and ADA, and a set of bacterial response genes comprising PTAFR, NRD1, PLP2; (d) a set of viral response genes comprising MX1, and a set of bacterial response genes comprising DYSF, TWF2; (e) a set of viral response genes comprising RSAD2, and (f) a set of viral response genes comprising IFI44L, GZMB, and KCTD14, and a set of bacterial response genes comprising TBXAS1, ACAA1, and S100A12; (g) a set of viral response genes comprising LY6E, and a set of bacterial response genes comprising PGD and LAPTM5; (h) a set of viral response genes comprising IFI44, HESX1, and OASL, and a set of bacterial response genes comprising NINJ2, DOK3, SORL1, and RAB31; and (i) a set of viral response genes comprising OAS1, and a set of bacterial response genes comprising IMPA2 and LTA4H.
[0082] In any embodiment, the biological sample may comprise whole blood or peripheral blood mononuclear cells (PBMCs).
[0083] In any embodiment, the levels of the biomarkers can be compared to time-matched reference values for infected or non-infected subjects.
[0084] In any embodiment, the method may include calculating a bacterial / viral metascore for the patient based on the levels of the biomarkers, where a positive bacterial / viral metascore for the patient indicates that the patient has a viral infection and a negative bacterial / viral metascore for the patient indicates that the patient has a bacterial infection.
[0085] In any embodiment, the method comprises the steps of measuring the expression levels of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, CTSB, CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 in the biological sample; and analyzing the expression level of each biomarker along with the respective reference ranges for the biomarkers, wherein the expression levels of the biomarkers CEACAM1, CEACAM2, CEACAM3, CEACAM4, CEACAM5, CEACAM6, CEACAM7, CEACAM8, CEACAM9, CEACAM10, CEACAM11, CEACAM12, CEACAM13, CEACAM14, CEACAM15, CEACAM16, CEACAM17, CEACAM18, CEACAM19, CEACAM19, CEACAM16, CEACAM18, CEACAM19, CEACAM19, CEACAM19, CEACAM11, CEACAM15, CEACAM16, CEACAM18, CEACAM19, CEACAM19, CEACAM11, CEACAM12, CEACAM13, CEACAM14, CEACAM15, CEACAM15, CEACAM16, CEACAM18, CEACAM19, CEACAM15, CEACAM19, CEACAM15, CEACAM15, CEACAM16, CEACAM18, CEACAM19, CEACAM19, CEACAM11, CEACAM15 ...5, CEACAM15, C The method may include steps in which an increase in the expression levels of ZDHHC19, C9orf95, GNA15, BATF, and C3AR1 and a decrease in the expression levels of the biomarkers KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 indicates that the patient has an infection, and an absence of differential expression of the biomarkers CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 compared to uninfected control subjects indicates that the patient does not have an infection.
[0086] In one embodiment, the method comprises: (a) a set of viral response genes comprising OAS2 and CUL1, and a set of bacterial response genes comprising SLC12A9, ACPP, STAT5B; (b) a set of viral response genes comprising ISG15 and CHST12, and a set of bacterial response genes comprising EMR1 and FLII; (c) a set of viral response genes comprising IFIT1, SIGLEC1, and ADA, and a set of bacterial response genes comprising PTAFR, NRD1, PLP2; (d) a set of viral response genes comprising MX1, and a set of bacterial response genes comprising DYSF, TWF2; (e) a set of viral response genes comprising RSAD2, and SORT1 and T The present invention relates to a kit comprising an agent for measuring the expression levels of a set of viral response genes and a set of bacterial response genes selected from: (f) a set of bacterial response genes comprising SPO; (f) a set of viral response genes comprising IFI44L, GZMB, and KCTD14, and a set of bacterial response genes comprising TBXAS1, ACAA1, and S100A12; (h) a set of viral response genes comprising IFI44, HESX1, and OASL, and a set of bacterial response genes comprising NINJ2, DOK3, SORL1, and RAB31; and (i) a set of viral response genes comprising OAS1, and a set of bacterial response genes comprising IMPA2 and LTA4H.
[0087] In any embodiment, the kit may include a microarray.
[0088] In one embodiment, the invention provides a computer-implemented method for diagnosing a patient suspected of having an infection, the method comprising the steps of: (a) receiving input patient data including values for expression levels of at least two biomarkers in a biological sample of the patient, wherein the at least two biomarkers are selected from one or both of a first set of biomarkers where high expression levels are indicative of a bacterial infection and a second set of biomarkers where high expression levels are indicative of a viral infection; and wherein the set of viral response genes includes OAS2, CUL1, ISG15, CHST12, IFIT1, SIGLEC1, ADA, MX1, RSAD2, IFI44L, GZMB, KCTD14, LY6E, IFI44, HESX1, OASL, OAS1, OAS3, EIF2AK2, DDX60, DNMT1, HERC5, IFIH1, SAMD9, IFI6, IFIT3, IFIT5, XAF1, ISG20, PARP12, IFIT2, DHX58, S and TAT1, and the set of bacterial response genes comprises one or more genes selected from the group of SLC12A9, ACPP, STAT5B, EMR1, FLII, PTAFR, NRD1, PLP2, DYSF, TWF2, SORT1, TSPO, TBXAS1, ACAA1, S100A12, PGD, LAPTM5, NINJ2, DOK3, SORL1, RAB31, IMPA2, LTA4H, TALDO1, TKT, PYGL, CETP, PROS1, RTN3, CAT, CYBRD1; (b) analyzing the expression levels of the set of viral response genes and the set of bacterial response genes and comparing them with respective reference ranges for uninfected control subjects; (c) calculating a bacterial / viral metascore for the patient based on the expression levels of the set of viral response genes and the set of bacterial response genes; and (d) displaying information regarding the diagnosis of the patient.
[0089] In one embodiment, the present invention is directed to a diagnostic system for performing a computer-implemented method, the diagnostic system including: (a) a storage component for storing data, the storage component having instructions stored therein for determining a patient's diagnosis; (b) a computer processor for processing the data, the computer processor coupled to the storage component and configured to receive patient data and execute the instructions stored in the storage component to analyze the patient data according to one or more algorithms; and (c) a display component for displaying information related to the patient's diagnosis.
[0090] These and other embodiments of the present invention will readily occur to those of ordinary skill in the art given the disclosure herein. [Brief explanation of the drawings]
[0091] [Figure 1A] Figures 1A and 1B show summary receiver operating characteristic (ROC) curves for the discovery dataset (Figure 1A) and the direct validation dataset (Figure 1B) for the bacterial / viral metascore. The summary ROC curve is shown in black, with the 95% confidence intervals in dark gray. [Figure 1B] Figures 1A and 1B show summary receiver operating characteristic (ROC) curves for the discovery dataset (Figure 1A) and the direct validation dataset (Figure 1B) for the bacterial / viral metascore. The summary ROC curve is shown in black, with the 95% confidence intervals in dark gray. [Figure 2]Figure 2 shows bacterial / viral scores for the COCONUT co-normalized whole blood discovery dataset. The PBMC dataset is excluded from Figure 2 because it is expected to have different gene levels than whole blood. The global AUC across all whole blood discovery datasets is 0.92. Score distribution by dataset (dark gray = bacterial, light gray = viral), individual gene level, and housekeeping genes (grayscale) is shown. The dotted line indicates possible global thresholds. The width of each violin corresponds to the distribution of scores within a given dataset. Within each violin, the vertical bars span the 25th to 75th percentiles, and the central open dash indicates the mean score. Housekeeping genes (POLG, ATP6V1B1, and PEG10) show predicted invariance across datasets after COCONUT normalization. [Figure 3A] Figures 3A-3C show the Integrated Antibiotic Decision Model (IADM) across published gene expression data co-normalized by COCONUT that met the inclusion criteria. Figure 3A shows a schematic of the IADM. [Figure 3B] Figures 3A-3C show the Integrated Antibiotic Decision Model (IADM) across published gene expression data co-normalized by COCONUT that met the inclusion criteria, and Figure 3B shows the distribution of scores and cutoffs for IADM on COCONUT co-normalized data. [Figure 3C] Figures 3A-3C show the integrated antibiotic decision model (IADM) across published gene expression data co-normalized by COCONUT that met the inclusion criteria. Figure 3C shows the confusion matrix for diagnoses: sensitivity for bacterial infection: 94.0%; specificity for bacterial infection: 59.8%; sensitivity for viral infection: 53.0%; specificity for viral infection: 90.6%. [Figure 4A]Figures 4A-4E show gene expression data from children with SIRS / sepsis from the GPSSSI cohort (total N=96; 36 SIRS, 49 bacterial sepsis, 11 viral sepsis) that were not studied by microarray using targeting NanoStrings. Figure 4A shows an overview of infected patients by organism type. [Figure 4B] Figures 4A-4E show gene expression data from children with SIRS / sepsis from the GPSSSI cohort (total N=96; 36 SIRS, 49 bacterial sepsis, 11 viral sepsis) that were not studied by microarray using targeting NanoStrings. Figures 4B and 4C show ROC curves for SMS and bacterial / viral metascores. [Figure 4C] Figures 4A-4E show gene expression data from children with SIRS / sepsis from the GPSSSI cohort (total N=96; 36 SIRS, 49 bacterial sepsis, 11 viral sepsis) that were not studied by microarray using targeting NanoStrings. Figures 4B and 4C show ROC curves for SMS and bacterial / viral metascores. [Figure 4D] Figures 4A-4E show gene expression data from children with SIRS / sepsis from the GPSSSI cohort (total N=96; 36 SIRS, 49 bacterial sepsis, 11 viral sepsis) that were not studied by microarray. Figure 4D shows the score distribution and cutoff for IADM. [Figure 4E]Figures 4A-4E show gene expression data from children with SIRS / sepsis from the GPSSSI cohort (total N = 96; 36 SIRS, 49 bacterial sepsis, 11 viral sepsis) that were not examined by microarray. Figure 4E shows the confusion matrix for IADM, with sensitivity for bacterial infection: 89.7%; specificity for bacterial infection: 70.0%; sensitivity for viral infection: 54.5%; and specificity for viral infection: 96.5%. [Figure 5A] 5A and 5B show that the Sepsis Metascore (SMS) cannot determine the type of pathogen by itself. The schematic diagram in Fig. 5A shows how to build a decision model. [Figure 5B] Figures 5A and 5B show that the sepsis metascore (SMS) cannot independently determine the type of pathogen. Figure 5B shows the distribution of SMS in patients with bacterial infections compared with patients with viral infections. Of the 11 datasets, only three showed significant differences in SMS distribution between bacterial and viral infections. [Figure 6] FIG. 1 is a schematic diagram of the workflow for multi-cohort analysis and bacterial-viral meta-signature discovery. [Figure 7A] Figure 1 shows a forest plot for genes in bacterial / viral metascores across the discovery dataset. The x-axis represents the standardized mean difference between bacterial and viral infected samples, calculated as Hedges' g-value, on a log2 scale. The size of the black box is inversely proportional to the standard error of the mean across studies. Boxes and whiskers represent the 95% confidence interval. The light gray diamonds represent the overall combinatorial mean difference for a given gene. The width of the diamond represents the 95% confidence interval for the overall combinatorial mean difference. [Figure 7B]Figure 1 shows a forest plot for genes in bacterial / viral metascores across the discovery dataset. The x-axis represents the standardized mean difference between bacterial and viral infected samples, calculated as Hedges' g-value, on a log2 scale. The size of the black box is inversely proportional to the standard error of the mean across studies. Boxes and whiskers represent the 95% confidence interval. The light gray diamonds represent the overall combinatorial mean difference for a given gene. The width of the diamond represents the 95% confidence interval for the overall combinatorial mean difference. [Figure 8] Figure 1 shows a forest plot for a random-effects meta-analysis of the summary ROC parameters alpha and beta for the discovery dataset. Alpha roughly controls the distance from the line of direct proportion (high alpha = high AUC), and beta controls the skewness of the actual ROC curve (beta = 0 means no skewness). [Figure 9] Figure 1 shows a forest plot for a random-effects meta-analysis of the summary ROC parameters alpha and beta for the validation dataset. Alpha roughly controls the distance from the line of direct proportion (high alpha = high AUC), and beta controls the skewness of the actual ROC curve (beta = 0 means no skewness). [Figure 10] This figure shows the ROC of the bacterial / viral metascore for GSE53166, a monocyte-derived dendritic cell model stimulated in vitro with LPS or influenza virus, for a total of N=75 cases (39 cases: LPS, 36 cases: influenza virus). [Figure 11] Schematic diagram of COCONUT co-normalization. Light grey indicates healthy ("H"), medium grey means viral ("V"), and dark grey means bacterial ("B"). Different cross-hatching is intended to indicate different batch effects. For formal mathematical details, see Methods. [Figure 12A]Figures 12A and 12B show data from the whole blood discovery dataset. The PBMC dataset is excluded from Figures 12A and 12B because it is expected to have different gene levels than whole blood. Figure 12A shows the raw data. [Figure 12B] Figures 12A and 12B show data from the whole blood discovery dataset. The PBMC dataset is excluded from Figures 12A and 12B because it is expected to have different gene levels than whole blood. Figure 12B shows COCONUT co-normalized data. COCONUT co-normalization was performed to reset each gene to the same position and scale as control patients. The distribution of genes within the datasets is unchanged (the median difference in T-statistics for healthy versus diseased within a dataset is 0, with a range of (-1 x 10-13, 1 x 10-13) across all genes and all datasets). The housekeeping gene ATP6V1B1 exhibits the expected invariance with disease and is invariant across datasets after normalization. Genes predicted to be induced by disease, such as CEACAM1, exhibit invariance across healthy controls but may vary between datasets in disease states. The upper colored bar indicates the data set and the lower colored bar indicates the disease class. [Figure 13] Figure 1 shows the bacteria / virus scores in the global ROC for COCONUT co-normalization of the whole blood validation dataset. The global AUC across all whole blood validation datasets is 0.93. The score distribution by dataset (dark gray violins = bacteria, light gray violins = viruses) and housekeeping genes (grayscale) is shown. The width of each violin corresponds to the distribution of scores within a given dataset. Within each violin, the vertical bars span the 25th to 75th percentiles, and the central open dash indicates the mean score. The dotted line indicates a possible global threshold. The housekeeping genes (POLG, ATP6V1B1, and PEG10) show predicted invariance across datasets after COCONUT normalization. [Figure 14]Figure 14 shows the bacterial / viral scores in the global ROC for the non-conormalized whole blood discovery dataset. The PBMC dataset is excluded from Figure 14 because it is expected to have different gene levels than whole blood. The global AUC across all whole blood discovery datasets is 0.93. The distribution of scores by dataset (dark gray violins = bacteria, light gray violins = viruses) and housekeeping genes (grayscale) is shown. The width of each violin corresponds to the distribution of scores within a given dataset. Within each violin, the vertical bars span the 25th to 75th percentiles, and the central open dash indicates the mean score. Note the highly variable position and scale of the housekeeping genes POLG, ATP6V1B1, and PEG10. [Figure 15] Figure 15 shows bacterial / viral scores in the global ROC for the non-conormalized whole blood validation dataset. The PBMC dataset is excluded from Figure 15 because it is expected to have different gene levels than whole blood. Score distribution by dataset (dark gray violins = bacteria, light gray violins = viruses) and housekeeping genes (grayscale) are shown. The width of each violin corresponds to the distribution of scores within a given dataset. Within each violin, the vertical bars span the 25th to 75th percentiles, and the central open dash indicates the mean score. Note the highly variable position and scale of the housekeeping genes POLG, ATP6V1B1, and PEG10. [Figure 16]Figure 1 shows the bacterial / viral scores in the global ROC for COCONUT co-normalization of the PBMC validation dataset. The PBMC dataset is considered separately because it is expected to have different gene levels than whole blood. The global AUC across all PBMC validation datasets is 0.92. The distribution of scores by dataset (dark gray violins = bacteria, light gray violins = viruses) and housekeeping genes (grayscale) is shown. The dotted line indicates possible global thresholds. The width of each violin corresponds to the distribution of scores within a given dataset. Within each violin, the vertical bars span the 25th to 75th percentiles, and the central open dash indicates the mean score. Housekeeping genes (POLG, ATP6V1B1) show predicted invariance across datasets after COCONUT normalization. [Figure 17] Figure 1 shows bacterial / viral scores in the global ROC for the non-conormalized PBMC validation dataset. The PBMC dataset is considered separately because it is expected to have different gene levels than whole blood. Score distributions are shown by dataset (dark gray violins = bacteria, light gray violins = viruses), individual gene levels, and housekeeping genes (grayscale). The width of each violin corresponds to the distribution of scores within a given dataset. Within each violin, the vertical bars span the 25th to 75th percentiles, and the central open dash indicates the mean score. Note the highly variable position and scale of the housekeeping genes POLG and ATP6V1B1. [Figure 18] FIG. 1 shows the distribution of mean AUC across all discovery datasets for 10,000 randomly chosen pairs of two genes. [Figure 19A]Figures 19A-19D show the effect of age on sepsis metascores in the COCONUT co-normalized data. Figure 19A shows age versus SMS for each pathogen type to assess whether pathogen type drives age differences in SMS. Figure 19B shows log10(age) versus SMS for each pathogen type, indicating that at the oldest ages, SMS may also have different maximum values that can be achieved. Figure 19C shows log10(age) versus SMS for each dataset, confirming that the relationship between age and SMS is dataset-independent. Figures 19A-19C incorporate only infected patient samples, while Figure 19D shows both healthy and non-infectious SIRS samples, in addition to showing baseline across ages. In all cases, the age data in GSE25504 are randomly distributed according to the mean age, given in their manuscripts approximately in 2-week ± 1-week increments, to demonstrate data density. All age=0 were reset as age=1 / 365. [Figure 19B] Figures 19A-19D show the effect of age on sepsis metascores in the COCONUT co-normalized data. Figure 19A shows age versus SMS for each pathogen type to assess whether pathogen type drives age differences in SMS. Figure 19B shows log10(age) versus SMS for each pathogen type, indicating that at the oldest ages, SMS may also have different maximum values that can be achieved. Figure 19C shows log10(age) versus SMS for each dataset, confirming that the relationship between age and SMS is dataset-independent. Figures 19A-19C incorporate only infected patient samples, while Figure 19D shows both healthy and non-infectious SIRS samples, in addition to showing baseline across ages. In all cases, the age data in GSE25504 are randomly distributed according to the mean age, given in their manuscripts approximately in 2-week ± 1-week increments, to demonstrate data density. All age=0 were reset as age=1 / 365. [Figure 19C]Figures 19A-19D show the effect of age on sepsis metascores in the COCONUT co-normalized data. Figure 19A shows age versus SMS for each pathogen type to assess whether pathogen type drives age differences in SMS. Figure 19B shows log10(age) versus SMS for each pathogen type, indicating that at the oldest ages, SMS may also have different maximum values that can be achieved. Figure 19C shows log10(age) versus SMS for each dataset, confirming that the relationship between age and SMS is dataset-independent. Figures 19A-19C incorporate only infected patient samples, while Figure 19D shows both healthy and non-infectious SIRS samples, in addition to showing baseline across ages. In all cases, the age data in GSE25504 are randomly distributed according to the mean age, given in their manuscripts approximately in 2-week ± 1-week increments, to demonstrate data density. All age=0 were reset as age=1 / 365. [Figure 19D] Figures 19A-19D show the effect of age on sepsis metascores in the COCONUT co-normalized data. Figure 19A shows age versus SMS for each pathogen type to assess whether pathogen type drives age differences in SMS. Figure 19B shows log10(age) versus SMS for each pathogen type, indicating that at the oldest ages, SMS may also have different maximum values that can be achieved. Figure 19C shows log10(age) versus SMS for each dataset, confirming that the relationship between age and SMS is dataset-independent. Figures 19A-19C incorporate only infected patient samples, while Figure 19D shows both healthy and non-infectious SIRS samples, in addition to showing baseline across ages. In all cases, the age data in GSE25504 are randomly distributed according to the mean age, given in their manuscripts approximately in 2-week ± 1-week increments, to demonstrate data density. All age=0 were reset as age=1 / 365. [Figure 20A]Figures 20A and 20B show the sepsis metascore across all whole blood data (both discovery and validation whole blood data) before (Figure 20B) and after (Figure 20A) COCONUT co-normalization. The global AUC is 0.86 (95% CI: 0.84-0.89) after COCONUT co-normalization. The score distribution by dataset (light gray violins = non-infectious inflammation, dark gray violins = infection / sepsis) and housekeeping genes (grayscale) is shown. The dotted line indicates possible global thresholds. The width of each violin corresponds to the distribution of scores within a given dataset. Within each violin, the vertical bars span the 25th to 75th percentiles, and the central open dash indicates the mean score. Note the invariance of the housekeeping genes POLG, ATP6V1B1, and PEG10 after COCONUT normalization in Figure 20A, and also note the highly variable positions and scales of the housekeeping genes before normalization in Figure 20B. [Figure 20B] Figures 20A and 20B show the sepsis metascore across all whole blood data (both discovery and validation whole blood data) before (Figure 20B) and after (Figure 20A) COCONUT co-normalization. The global AUC is 0.86 (95% CI: 0.84-0.89) after COCONUT co-normalization. The score distribution by dataset (light gray violins = non-infectious inflammation, dark gray violins = infection / sepsis) and housekeeping genes (grayscale) is shown. The dotted line indicates possible global thresholds. The width of each violin corresponds to the distribution of scores within a given dataset. Within each violin, the vertical bars span the 25th to 75th percentiles, and the central open dash indicates the mean score. Note the invariance of the housekeeping genes POLG, ATP6V1B1, and PEG10 after COCONUT normalization in Figure 20A, and also note the highly variable positions and scales of the housekeeping genes before normalization in Figure 20B. [Figure 21A]Figures 21A and 21B show IADM across published gene expression data co-normalized by COCONUT, including healthy controls. The included datasets (and score cutoffs used) are the same as those in Figures 3A-3C. Figure 21A shows the score distribution and cutoffs for IADM in the COCONUT co-normalized data. [Figure 21B] Figures 21A and 21B show IADM across published gene expression data co-normalized by COCONUT, including healthy controls. The included datasets (and score cutoffs used) are the same as those in Figures 3A-3C. Figure 21B shows the confusion matrix for diagnosis. Sensitivity for bacterial infection: 94.2%; specificity for bacterial infection: 68.5%; sensitivity for viral infection: 53.0%; specificity for viral infection: 94.1%. "SIRS" refers to non-infectious inflammation. [Figure 22] Figure 1 shows NPV and PPV versus prevalence for a diagnostic test with a sensitivity of 94.0% and a specificity of 59.8%. The red line indicates an NPV of 98.3% at a prevalence of 15% as a rough estimate of the actual incidence of infection. [Figure 23A] Figures 23A-23D show results for the GSE63990 dataset (adults with acute respiratory infections). Figures 23A and 23B show ROC curves for the sepsis metascore and the bacterial / viral metascore. [Figure 23B] Figures 23A-23D show results for the GSE63990 dataset (adults with acute respiratory infections). Figures 23A and 23B show ROC curves for the sepsis metascore and the bacterial / viral metascore. [Figure 23C] Figures 23A-23D show results for the GSE63990 dataset (adults with acute respiratory infections), and Figure 23C shows the distribution of scores and cutoffs for IADM. [Figure 23D]Figures 23A-23D show results for the GSE63990 dataset (adults with acute respiratory infections). Figure 23D shows the confusion matrix for IADM, with sensitivity for bacterial infection: 94.3%; specificity for bacterial infection: 52.2%; sensitivity for viral infection: 52.2%; specificity for viral infection: 94.3%. DETAILED DESCRIPTION OF THE INVENTION
[0092] Unless otherwise indicated, the practice of the present invention will employ conventional methods of pharmacology, chemistry, biochemistry, recombinant DNA technology, and immunology, within the skill of the art. Such techniques are explained fully in the literature, see, e.g., J.E.Bennett, R. Dolin, and M.J.Blaser, "Mandell, Douglas, and Bennett's Principles and Practice of Infectious Diseases" (Saunders, 8th ed., 2014); J.R.Brown, "Sepsis: Symptoms, Diagnosis and Treatment" (Public Health in the 21st Century, 2014). stCentury Series, Nova Science Publishers, Inc., 2013); "Sepsis and Non-infectious Systemic Inflammation: From Biology to Critical Care" (J. Cavaillon, C. Adrie, eds., Wiley-Blackwell, 2008); "Sepsis: Diagnosis, Management and Health Outcomes" (Allergies and Infectious Diseases, N. Khardori, ed., Nova Science Pub Inc., 2014); "Handbook of Experimental Immunology," Volumes I-IV (D.M. Weir and C.C. Blackwell, eds., Blackwell Scientific Publications); A.L. Lehninger, Biochemistry (Worth Publishers, Inc., latest expanded edition); Sambrook et al., "Molecular Cloning: A Laboratory Manual" (3rd ed., 2001); "Methods In Enzymology" (S. Colowick and N. Kaplan, eds., Academic Press, Inc.).
[0093] All publications, patents, and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety. I. Definition
[0094] In describing the present invention, the following terms will be utilized and are intended to be defined as indicated below.
[0095] As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a biomarker" includes a mixture of two or more biomarkers, and the like.
[0096] In particular, the term "about" when referring to a given quantity is intended to encompass a deviation of plus or minus 5 percent.
[0097] The term area under the curve (AUC) as used herein is understood to refer to the area under the receiver operating characteristic curve (ROC curve).
[0098] A "biomarker" in the context of the present invention refers to a biological compound, such as a polynucleotide, that is differentially expressed in a sample taken from a patient with an infection compared to an equivalent sample taken from a control subject (e.g., a person with a negative diagnosis, a normal or healthy subject, or an uninfected subject). A biomarker can be a nucleic acid, a fragment of a nucleic acid, a polynucleotide, or an oligonucleotide that can be detected and / or quantified. Biomarkers are polynucleotides comprising a nucleotide sequence derived from a gene or an RNA transcript of a gene, and include IFI27, JUP, LAX1, OAS2, CUL1, ISG15, CHST12, IFIT1, SIGLEC1, ADA, MX1, RSAD2, IFI44L, GZMB, KCTD14, LY6E, IFI44, HESX1, OASL, OAS1, OAS3, EIF2AK2, DDX60, DNMT1, HERC5, IFIH1, SAMD9, IFI6, IFIT3, IFIT5, XAF1, ISG20, PARP12, IFIT2, DHX58, STAT1, HK3, TNIP1, GPAA1, CTSB, These polynucleotides include, but are not limited to, SLC12A9, ACPP, STAT5B, EMR1, FLII, PTAFR, NRD1, PLP2, DYSF, TWF2, SORT1, TSPO, TBXAS1, ACAA1, S100A12, PGD, LAPTM5, NINJ2, DOK3, SORL1, RAB31, IMPA2, LTA4H, TALDO1, TKT, PYGL, CETP, PROS1, RTN3, CAT, CYBRD1, CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1.
[0099] "Viral response genes" refer to genes that are differentially expressed in samples taken from patients with viral infections compared to comparable samples taken from control subjects (e.g., individuals with negative diagnoses, normal or healthy subjects, or uninfected subjects). Viral response genes include, but are not limited to, IFI27, JUP, LAX1, OAS2, CUL1, ISG15, CHST12, IFIT1, SIGLEC1, ADA, MX1, RSAD2, IFI44L, GZMB, KCTD14, LY6E, IFI44, HESX1, OASL, OAS1, OAS3, EIF2AK2, DDX60, DNMT1, HERC5, IFIH1, SAMD9, IFI6, IFIT3, IFIT5, XAF1, ISG20, PARP12, IFIT2, DHX58, and STAT1.
[0100] "Bacterial response gene" refers to a gene that is differentially expressed in a sample taken from a patient with a bacterial infection compared to a comparable sample taken from a control subject (e.g., a person with a negative diagnosis, a normal or healthy subject, or an uninfected subject). Bacterial response genes include, but are not limited to, HK3, TNIP1, GPAA1, CTSB, SLC12A9, ACPP, STAT5B, EMR1, FLII, PTAFR, NRD1, PLP2, DYSF, TWF2, SORT1, TSPO, TBXAS1, ACAA1, S100A12, PGD, LAPTM5, NINJ2, DOK3, SORL1, RAB31, IMPA2, LTA4H, TALDO1, TKT, PYGL, CETP, PROS1, RTN3, CAT, and CYBRD1.
[0101] "Sepsis response gene" refers to a gene that is differentially expressed in a sample taken from a patient with sepsis or infection compared to a comparable sample taken from a control subject (e.g., a person with a negative diagnosis, a normal or healthy subject, or a non-infected subject). Sepsis response genes include, but are not limited to, CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1.
[0102] The terms "polypeptide" and "protein" refer to a polymer of amino acid residues and are not limited to a minimum length. Thus, peptides, oligopeptides, dimers, multimers, and the like, are included within the definition. Both full-length proteins and fragments thereof are encompassed by the definition. The terms also include post-expression modifications of the polypeptide, such as glycosylation, acetylation, phosphorylation, hydroxylation, oxidation, and the like.
[0103] As used herein, the terms "polynucleotide," "oligonucleotide," "nucleic acid," and "nucleic acid molecule" are used to include polymeric forms of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. The terms refer only to the primary structure of the molecule. Thus, the terms include triple-stranded, double-stranded, and single-stranded DNA, as well as triple-stranded, double-stranded, and single-stranded RNA. The terms may also include unmodified forms of polynucleotides, as well as modifications such as methylation and / or capping. More specifically, the terms "polynucleotide," "oligonucleotide," "nucleic acid," and "nucleic acid molecule" include polydeoxyribonucleotides (containing 2-deoxy-D-ribose), polyribonucleotides (containing D-ribose), and any other type of polynucleotide that is an N- or C-glycoside of a purine or pyrimidine base. No distinction in length is intended between the terms "polynucleotide," "oligonucleotide," "nucleic acid," and "nucleic acid molecule," and these terms are used interchangeably.
[0104] The phrase "differentially expressed" refers to, for example, a difference in the quantity and / or frequency of a biomarker present in a sample taken from a patient with an infection (e.g., a viral or bacterial infection) compared to a control subject or an uninfected subject. For example, a biomarker can be a polynucleotide that is present at a higher or lower level in samples from patients with an infection (e.g., a viral or bacterial infection) compared to samples from control subjects. Alternatively, a biomarker can be a polynucleotide that is detected at a higher or lower frequency in samples from patients with an infection (e.g., a viral or bacterial infection) compared to samples from control subjects. A biomarker can be differentially present in terms of quantity, frequency, or both.
[0105] A polynucleotide is differentially expressed between two samples if the amount of the polynucleotide in one sample is statistically significantly different from the amount of the polynucleotide in the other sample. For example, a polynucleotide is differentially expressed in two samples if it is present at least about 120%, at least about 130%, at least about 150%, at least about 180%, at least about 200%, at least about 300%, at least about 500%, at least about 700%, at least about 900%, or at least about 1000% more than it is present in the other sample, or if it is detectable in one sample but not detectable in the other sample.
[0106] Alternatively, or in addition, a polynucleotide is differentially expressed in two sample sets if it is detected at a statistically significant greater or less frequent frequency in samples from patients with sepsis than in control samples. For example, a polynucleotide is differentially expressed in two sample sets if it is detected at least about 120%, at least about 130%, at least about 150%, at least about 180%, at least about 200%, at least about 300%, at least about 500%, at least about 700%, at least about 900%, or at least about 1000% more or less frequently in one set of samples than in the other set of samples.
[0107] A "similarity value" is a number that represents the degree of similarity between two things being compared. For example, a similarity value can be a number that indicates the overall similarity of a patient's expression profile using a specific phenotype-associated biomarker to a reference value range for the biomarker in one or more control samples, or to a reference expression profile (e.g., similarity to a "viral infection" expression profile or a "bacterial infection" expression profile). A similarity value can be expressed as a similarity metric, such as a correlation coefficient, or simply as the difference in expression level, or the sum of the difference in expression level, between the level of the biomarker in the patient sample and the level of the biomarker in the control sample or the reference expression profile.
[0108] As used herein, the terms "subject," "individual," and "patient" are used interchangeably to refer to any mammalian subject, particularly humans, for whom diagnosis, prognosis, treatment, or therapy is desired. Other subjects may include cows, dogs, cats, guinea pigs, rabbits, rats, mice, horses, etc. Optionally, the methods of the present invention are used in the development of laboratory animals, veterinary applications, and animal models for disease, including, but not limited to, rodents, including mice, rats, and hamsters; and primates.
[0109] As used herein, a "biological sample" refers to a sample of tissue, cells, or bodily fluid isolated from a subject, including, but not limited to, blood, buffy coat, plasma, serum, blood cells (e.g., peripheral blood mononuclear cells (PBMCs)), feces, urine, bone marrow, bile, spinal fluid, lymphatic fluid, skin samples, external secretions of the skin, respiratory, intestinal, and genitourinary tract, tears, saliva, breast milk, organs, biopsies, and also in vitro cell culture components, including, but not limited to, cells and tissues in culture medium, e.g., conditioned medium obtained from the growth of recombinant cells, and samples of cellular components.
[0110] A "test amount" of a biomarker refers to the amount of biomarker present in a test sample. The test amount can be an absolute amount (e.g., μg / ml) or a relative amount (e.g., relative intensity of signals).
[0111] A "diagnostic amount" of a biomarker refers to the amount of the biomarker in a subject's sample that is consistent with a diagnosis of infection (e.g., viral or bacterial infection). A diagnostic amount can be an absolute amount (e.g., μg / ml) or a relative amount (e.g., relative intensity of signals).
[0112] A "control amount" of a biomarker can be any amount or range of amounts compared to a test amount of a biomarker. For example, a control amount of a biomarker can be the amount of the biomarker in a person without an infection (e.g., a viral or bacterial infection). A control amount can be an absolute amount (e.g., μg / ml) or a relative amount (e.g., relative intensity of signals).
[0113] The term "antibody" refers to polyclonal and monoclonal antibody preparations, as well as hybrid, modified, chimeric, and humanized antibody molecules, as well as hybrid (chimeric) antibody molecules (see, e.g., Winter et al. (1991) Nature 349:293-299; and U.S. Pat. No. 4,816,567); F(ab')2 fragments and F(ab)2 fragments; Fv molecules (noncovalent heterodimers, see, e.g., Inbar et al. (1972) Proc Natl Acad Sci USA 69:2659-2662; and Ehrlich et al. (1980) Biochem 19:4091-4096); single-chain Fv molecules (sFv) (see, e.g., Huston et al. (1988) Proc Natl Acad Sci USA USA 85:5879-5883); dimeric and trimeric antibody fragment constructs; minibodies (see, e.g., Pack et al. (1992) Biochem 31:1579-1584; Cumber et al. (1992) J Immunology 149B:120-126); humanized antibody molecules (see, e.g., Riechmann et al. (1988) Nature 332:323-327; Verhoeyan et al. (1988) Science 239:1534-1536; and GB Patent Publication No. 2,276,169, published 21 September 1994); and preparations comprising any functional fragments derived from such molecules, which fragments retain the specific binding properties of the parent antibody molecule.
[0114] "Detectable moieties" or "detectable labels" contemplated for use in the present invention include radioisotopes, fluorescent dyes such as fluorescein, phycoerythrin, Cy-3, Cy-5, allophycocyanin, DAPI, Texas Red, rhodamine, Oregon Green, Lucifer Yellow, and the like, green fluorescent protein (GFP), red fluorescent protein (DsRed), cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), Cerianthus orange fluorescent protein (cOFP), alkaline phosphatase (AP), beta-lactamase, chloramphenicol acetyltransferase (CAT), adenosine deaminase (ADA), aminoglycoside phosphotransferase (neo), and the like. r , G418 rEnzyme tags, including but not limited to, dihydrofolate reductase (DHFR), hygromycin B-phosphotransferase (HPH), thymidine kinase (TK), lacZ (encoding β-galactosidase), and xanthine guanine phosphoribosyltransferase (XGPRT), beta-glucuronidase (gus), placental alkaline phosphatase (PLAP), secreted fetal alkaline phosphatase (SEAP), or firefly luciferase or bacterial luciferase (LUC), are used in conjunction with their cognate substrates. The term also includes color-coded microspheres with known fluorescence intensities (see, e.g., microspheres with xMAP technology from Luminex, Austin, TX); microspheres containing quantum dot nanocrystals, e.g., containing different ratios and combinations of quantum dot colors (see, e.g., Qdot nanocrystals from Life Technologies, Carlsbad, CA); glass-coated metal nanoparticles (see, e.g., SERS nanotags from Nanoplex Technologies, Inc., Mountain View, CA); barcode materials (see, e.g., submicron-sized striped metal rods, such as Nanobarcodes from Nanoplex Technologies, Inc.); coded microparticles with colored barcodes (see, e.g., CellCard from Vitra Bioscience, vitrabio.com); and glass microparticles with digital holographic code images (see, e.g., CyVera microbeads from Illumina, San Diego, CA). As with many of the standard procedures associated with the practice of the present invention, one of skill in the art will be aware of additional labels that may be used.
[0115] "Devout a classifier" refers to using input variables to generate an algorithm or classifier that is capable of distinguishing between two or more states.
[0116] As used herein, "diagnosis" generally includes the determination of whether a subject is likely to be affected by a given disease, disorder, or dysfunction. Those skilled in the art often make a diagnosis based on one or more diagnostic indicators, i.e., biomarkers, the presence, absence, or amount of which indicates the presence or absence of a disease, disorder, or dysfunction.
[0117] As used herein, "prognosis" generally refers to a prediction of the likely course and outcome of a clinical condition or disease. A prognosis for a patient is typically made by assessing disease factors or symptoms that indicate a favorable or unfavorable course or outcome of the disease. It is understood that the term "prognosis" does not necessarily refer to the ability to predict the course or outcome of a condition with 100% accuracy. Indeed, those skilled in the art will understand that the term "prognosis" refers to an increased probability of a particular course or outcome occurring; i.e., the likelihood of a course or outcome occurring in patients who exhibit a given condition compared to individuals who do not exhibit the condition.
[0118] "Substantially purified" refers to nucleic acid molecules or proteins that have been removed from their natural environment, isolated or separated, and are at least about 60% free, preferably about 75% free, and most preferably about 90% free from other components with which they are naturally associated. II. MODES OF CARRYING OUT THE INVENTION
[0119] Before describing the present invention in detail, it is to be understood that this invention is not limited to particular formulations or process parameters, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments of the invention only and is not intended to be limiting.
[0120] Although a number of methods and materials similar or equivalent to those described herein can be used in the practice of the present invention, the preferred materials and methods are described herein.
[0121] The present invention is based on the discovery of biomarkers (see Example 1) that can be used to diagnose infection. In particular, the present invention relates to the use of biomarkers that can be used to determine whether a patient with acute inflammation has a bacterial infection that would benefit from treatment with an antibiotic or a viral infection that would benefit from treatment with an antiviral. To further understand the present invention, a more detailed discussion of the identified biomarkers and methods of using them in the diagnosis and treatment of infection is presented below. A. Biomarkers
[0122] Biomarkers that can be used in the practice of the present invention are "viral response genes," which are polynucleotides comprising a nucleotide sequence derived from a gene or an RNA transcript of a gene, that are differentially expressed in patients with a viral infection compared to control subjects (e.g., individuals with a negative diagnosis, normal or healthy subjects, or uninfected subjects without a viral infection), and include IFI27, JUP, LAX1, OAS2, CUL1, ISG15, CHST12, IFIT1, SIGLEC1, ADA, MX1, RSAD2, IGF-1, IGF-2, IGF-3, IGF-4, IGF-5, IGF-6, IGF-7, IGF-8, IGF-9, IGF-11, IGF-12, IGF-13, IGF-14, IGF-15, IGF-16, IGF-17, IGF-18, IGF-19, IGF-20, IGF-21, IGF-22, IGF-23, IGF-24, IGF-25, IGF-26, IGF-27, IGF-28, IGF-29, IGF-30, IGF-31, IGF-32, IGF-33, IGF-34, IGF-35, IGF-36, IGF-37, IGF-38, IGF-39, IGF-40, IGF-41, IGF-42, IGF-43, IGF-44, IGF-45, IGF-46, IGF-47, IGF-48, IGF-49, IGF-50, IGF-51, IGF-52, IGF-53, IGF-54, IGF-55, IGF-56, IGF-57, IGF-58, IGF-59, IGF-60, IGF-61, IGF-62, IGF-63, IGF-64, IGF "viral response genes" such as, but not limited to, FI44L, GZMB, KCTD14, LY6E, IFI44, HESX1, OASL, OAS1, OAS3, EIF2AK2, DDX60, DNMT1, HERC5, IFIH1, SAMD9, IFI6, IFIT3, IFIT5, XAF1, ISG20, PARP12, IFIT2, DHX58, and STAT1; in patients with bacterial infection, control subjects (e.g., those with a negative diagnosis, normal or healthy subjects, or uninfected subjects without bacterial infection). ), including HK3, TNIP1, GPAA1, CTSB, SLC12A9, ACPP, STAT5B, EMR1, FLII, PTAFR, NRD1, PLP2, DYSF, TWF2, SORT1, TSPO, TBXAS1, ACAA1, S100A12, PGD, LAPTM5, NINJ2, DOK3, SORL1, RAB31, IMPA2, LTA4H, TALDO1, TKT, PYGL, CETP, PROS1, RTN3, CAT, and CYBRD1. "Bacterial response genes" include, but are not limited to, "bacterial response genes"; and "sepsis response genes" that are differentially expressed in patients with sepsis or infection compared to control subjects (e.g., individuals with a negative diagnosis, normal or healthy subjects, or uninfected subjects without sepsis), including polynucleotides containing "sepsis response genes" such as, but not limited to, CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1.
[0123] In one aspect, the invention includes a method of diagnosing an infection in a patient, the method comprising the steps of: a) obtaining a biological sample from the patient; b) measuring, in the biological sample, the expression levels of a set of viral response genes that exhibit differential expression associated with viral infection and a set of bacterial response genes that exhibit differential expression associated with bacterial infection; and c) analyzing the expression levels of the viral response genes and the bacterial response genes along with their respective reference ranges.
[0124] When analyzing biomarker levels in biological samples, the reference range may represent the levels of one or more biomarkers found in one or more samples for one or more subjects without infection (e.g., healthy or uninfected subjects). Alternatively, the reference range may represent the levels of one or more biomarkers found in one or more samples for one or more subjects with a viral or bacterial infection. In certain embodiments, the levels of the biomarkers are compared to time-matched reference ranges for uninfected or infected subjects.
[0125] In certain embodiments, the set of viral response genes and the set of bacterial response genes are: a) a set of viral response genes comprising IFI27, JUP, and LAX1, and a set of bacterial response genes comprising HK3, TNIP1, GPAA1, and CTSB; b) a set of viral response genes comprising OAS2 and CUL1, and a set of bacterial response genes comprising SLC12A9, ACPP, STAT5B; c) a set of viral response genes comprising ISG15 and CHST12, and a set of bacterial response genes comprising EMR1 and FLII; d) a set of viral response genes comprising IFIT1, SIGLEC1, and ADA, and a set of bacterial response genes comprising PTAFR, NRD1, PLP2; e) a set of viral response genes comprising MX1, and a set of bacterial response genes comprising DYSF, TWF f) a set of viral response genes comprising RSAD2, and a set of bacterial response genes comprising SORT1 and TSPO; g) a set of viral response genes comprising IFI44L, GZMB, and KCTD14, and a set of bacterial response genes comprising TBXAS1, ACAA1, and S100A12; h) a set of viral response genes comprising LY6E, and a set of bacterial response genes comprising PGD and LAPTM5; i) a set of viral response genes comprising IFI44, HESX1, and OASL, and a set of bacterial response genes comprising NINJ2, DOK3, SORL1, and RAB31; and j) a set of viral response genes comprising OAS1, and a set of bacterial response genes comprising IMPA2 and LTA4H.
[0126] The biological sample obtained from the patient to be diagnosed is typically whole blood or blood cells (e.g., PBMCs), but can be any sample containing bodily fluids, tissues, or cells containing expressed biomarkers. As used herein, a "control" sample refers to a biological sample, such as a non-diseased bodily fluid, tissue, or cell. That is, a control sample is obtained from a normal or non-infected subject (e.g., an individual known to be free of viral infection, bacterial infection, sepsis, or inflammation). Biological samples can be obtained from patients by conventional techniques. For example, blood can be obtained by venipuncture, and solid tissue samples can be obtained by surgical procedures according to methods well known in the art.
[0127] In certain embodiments, a panel of biomarkers is used for diagnosing infection. Biomarker panels of any size can be used in practicing the invention. Biomarker panels for diagnosing infection typically include at least three biomarkers and up to 30 biomarkers, including any number therebetween, such as 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 biomarkers. In certain embodiments, the invention includes biomarker panels that include at least three, or at least four, or at least five, or at least six, or at least seven, or at least eight, or at least nine, or at least ten, or at least 11, or more biomarkers. Although small biomarker panels are typically less expensive, large biomarker panels (i.e., greater than 30 biomarkers) have the advantage of providing detailed information and can also be used to practice the present invention.
[0128] In certain embodiments, the invention comprises a panel of biomarkers for diagnosing infection, the panel comprising one or more polynucleotides comprising a nucleotide sequence derived from a gene or RNA transcript of a gene selected from the group consisting of IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, and CTSB. In another embodiment, the panel of biomarkers further comprises one or more polynucleotides comprising a nucleotide sequence derived from a gene or RNA transcript of a gene selected from the group consisting of CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1.
[0129] In certain embodiments, the biomarkers described herein for distinguishing between viral and bacterial infections are combined with additional biomarkers capable of distinguishing whether inflammation in a subject is caused by an infection or a non-infectious source of inflammation (e.g., traumatic injury, surgery, autoimmune disease, thrombosis, or systemic inflammatory response syndrome (SIRS)). A first diagnostic test is used to determine whether the acute inflammation is caused by an infectious or non-infectious source, and if the source of the inflammation is infection, a second diagnostic test is used to determine whether the infection is a viral or bacterial infection that would benefit from treatment with an antiviral or antibiotic, respectively.
[0130] In one embodiment, the present invention provides a method of diagnosing and treating a patient with inflammation, comprising the steps of: a) obtaining a biological sample from the patient; b) measuring the expression levels of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, CTSB, CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 in the biological sample; and c) first determining the expression level of each biomarker by: and analyzing the biomarkers together with their respective reference ranges, wherein an increase in the expression levels of the biomarkers CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, and C3AR1 and a decrease in the expression levels of the biomarkers KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 compared to the reference ranges of the biomarkers for uninfected control subjects indicates that the patient has an infection, and wherein an increase in the expression levels of the biomarkers CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, and C3AR1 and a decrease in the expression levels of the biomarkers KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 compared to the reference ranges of the biomarkers for uninfected control subjects indicates that the patient has an infection. wherein the absence of differential expression of HHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 indicates that the patient does not have the infection; and d) second, if the patient is diagnosed with the infection, analyzing the expression levels of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, and CTSB, wherein the expression levels of the biomarkers IFI27, JUP, LAX1 are compared with those of the control subjects. wherein an elevation of the expression levels of the biomarkers HK3, TNIP1, GPAA1, CTSB compared to the reference range of the biomarkers for a control subject indicates that the patient has a bacterial infection; and e) administering to the patient an effective amount of an antiviral agent if the patient is diagnosed with a viral infection, or an effective amount of an antibiotic agent if the patient is diagnosed with a bacterial infection.
[0131] In another embodiment, the method further includes calculating a sepsis metascore for the patient, wherein a sepsis metascore above the reference range for non-infected control subjects indicates that the patient has an infection, and a sepsis metascore within the reference range for non-infected control subjects indicates that the patient has a non-infectious inflammatory condition.
[0132] In another embodiment, the method further comprises calculating a bacterial / viral metascore for the patient if the patient is diagnosed with an infection, wherein a positive bacterial / viral metascore for the patient indicates that the patient has a viral infection and a negative bacterial / viral metascore for the patient indicates that the patient has a bacterial infection.
[0133] In another embodiment, the invention includes a method of treating a patient suspected of having an infection, the method comprising the steps of: a) receiving information regarding the patient's diagnosis according to the methods described herein; and b) administering a therapeutically effective amount of an antiviral agent if the patient is diagnosed with a viral infection, or administering an effective amount of an antibiotic agent if the patient is diagnosed with a bacterial infection.
[0134] In certain embodiments, a patient diagnosed with a viral infection by the methods described herein is administered a therapeutically effective amount of an antiviral agent, such as a broad-spectrum antiviral agent, an antiviral vaccine, a neuraminidase inhibitor (e.g., zanamivir (Relenza) and oseltamivir (Tamiflu)), a nucleoside analog (e.g., acyclovir, zidovudine (AZT), and lamivudine), an antisense antiviral agent (e.g., phosphorothioate antisense antivirals (e.g., fomivirsen (Vitravene) for cytomegalovirus retinitis), morpholino antisense antivirals), a viral threshing inhibitor (e.g., amantadine and rimantadine for influenza, pleconaril for rhinovirus), a viral entry inhibitor (e.g., Fuzeon for HIV), a viral assembly inhibitor (e.g., rifampicin), or an antiviral agent that stimulates the immune system (e.g., interferon).Exemplary antiviral agents include abacavir, acyclovir, adefovir, amantadine, amprenavir, ampligen, arbidol, atazanavir, atripla (fixed dose), baravir, cidofovir, combivir (fixed dose), dolutegravir, darunavir, delavirdine, didanosine, docosanol, edoxudine, efavirenz, emtricitabine, enfuvirtide, entecavir, ecolieber, famciclovir, fixed dose combination (antiretroviral), fomivirsen, fosamprenavir, foscarnet, fosfonet, fusion inhibitors, ganciclovir, ibacitabine, immunovir, idoxuridine, imiquimod, indinavir, inosine, integrase inhibitors, type III interferon, type II interferon, type I interferon, interferon, la Mivudine, lopinavir, loviride, maraviroc, moroxydine, methisazone, nelfinavir, nevirapine, nexavir, nitazoxanide, nucleoside analogues, novir, oseltamivir (Tamiflu), pegylated interferon alfa-2a, penciclovir, peramivir, pleconaril, podophyllotoxin, protease inhibitors, raltegravir, reverse transcriptase inhibitors, ribavirin, rimantadine , ritonavir, pyramidine, saquinavir, sofosbuvir, stavudine, synergistic enhancers (antiretrovirals), telaprevir, tenofovir, tenofovir disoproxil, tipranavir, trifluridine, trizivir, tromantadine, Truvada, valacyclovir (Valtrex), valganciclovir, vicriviroc, vidarabine, viramidine, zalcitabine, zanamivir (Relenza), or zidovudine.
[0135] In certain embodiments, a patient diagnosed with a bacterial infection by the methods described herein is administered a therapeutically effective amount of an antibiotic, which may include a broad-spectrum antibiotic, a bactericidal antibiotic, or a bacteriostatic antibiotic. Exemplary antibiotics include aminoglycosides such as amikacin, Amikin, gentamicin, Garamycin, kanamycin, Kantrex, neomycin, Neo-Fradin, netilmicin, Netromycin, tobramycin, Nebcin, paromomycin, Humatin, streptomycin, spectinomycin (Bs), and Trobicin; ansamycins such as geldanamycin, herbimycin, rifaximin, and Xifaxan; carbacephems such as loracarbef and Lorabid; carbapenems such as ertapenem, Invanz, doripenem, Doribax, imipenem / cilastatin, Primaxin, meropenem, Merrem; cefadroxil, Duricef, cefazolin, Ancef, cephalothin, Cefalotin or Cephalothin, Keflin, Cephalexin, Cephalosporins such as Syn, Keflex, Cefaclor, Distaclor, Cefamandole, Mandol, Cefoxitin, Mefoxin, Cefprozil, Cefzil, Cefuroxime, Ceftin, Cefixime, Cefdinir, Cefditoren, Cefoperazone, Cefotaxime, Cefpodoxime, Ceftazidime, Ceftibuten, Ceftizoxime, Ceftriaxone, Cefepime, Maxipime, Ceftaroline Fosamil, Teflaro, Ceftobiprole, and Zeftera; glycopeptides such as Teicoplanin, Targocid, Vancomycin, Vancocin, Telavancin, Vibativ, Dalbavancin, Dalvance, Oritavancin, and Orbactiv; lincosamides such as Clindamycin, Cleocin, Lincomycin, and Lincocin; lipopeptides such as Daptomycin and Cubicin;Macrolides such as Azithromycin, Zithromax, Sumamed, Xithrone, Clarithromycin, Biaxin, Dirithromycin, Dynabac, Erythromycin, Erythocin, Erythroped, Roxithromycin, Troleandomycin, Tao, Telithromycin, Ketek, Spiramycin, Rovamycine; Monobactams such as Aztreonam, Azactam; Furazolidone, Fu Nitrofurans such as roxone, nitrofurantoin, Macrodantin and Macrobid; oxazolidinones such as linezolid, Zyvox, VRSA, pocizolid, radezolid and torezolid; penicillins such as penicillin V, Veetids (Pen-Vee-K), piperacillin, Pipracil, penicillin G, Pfizerpen, temocillin, Negaban, ticarcillin and Ticar; amoxicillin / clavulanic acid, Penicillin combinations such as Augmentin, ampicillin / sulbactam, Unasyn, piperacillin / tazobactam, Zosyn, ticarcillin / clavulanic acid, and Timentin; polypeptides such as bacitracin, colistin, Coly-Mycin-S, and polymyxin B; ciprofloxacin, Cipro, Ciproxin, Ciprobay, enoxacin, Penetrex, gatifloxacin, Tequin, and gemifloxacin Quinolones / fluoroquinolones such as Syn, Factive, Levofloxacin, Levaquin, Lomefloxacin, Maxaquin, Moxifloxacin, Avelox, Nalidixic Acid, NegGram, Norfloxacin, Noroxin, Ofloxacin, Floxin, Ocuflox, Trovafloxacin, Trovan, Grepafloxacin, Raxar, Sparfloxacin, Zagam, Temafloxacin, Omniflox;Amoxicillin, Novamox, Amoxil, Ampicillin, Principen, Azlocillin, Carbenicillin, Geocillin, Cloxacillin, Tegopen, Dicloxacillin, Dynapen, Flucloxacillin, Floxapen, Mezlocillin, Mezlin, Methicillin, Staphcillin, Nafcillin, Unipen, Oxacillin, Prostaphlin, Penicillin G, Pentids, etc., Mafenide, Sulfamylon, Sulfacetamide, Sulamyd, Bleph-10, Sulfadiazine, Micro-Sulfon, Silver Sulfadiazine, Silvadene, Sulfadimethoxine, Di-Methox, Albon, Sulfamethizole, Thiosulfil Sulfonamides such as Forte, sulfamethoxazole, Gantanol, sulfanilimide, sulfasalazine, Azulfidine, sulfisoxazole, Gantrisin, trimethoprim-sulfamethoxazole, cotrimoxazole, TMP-SMX, Bactrim, Septra, sulfonamide chrysoidine, and Prontosil; Demeclocycline, Declomycin, Doxycycline, Vibramycin, Minocycline, Minocin, Oxytetracycline, Terramycin, Tetracycline, Sumycin, and Achromycin Tetracyclines such as V and Steclin; drugs against mycobacteria such as clofazimine, Lamprene, dapsone, Avlosulfon, capreomycin, Capastat, cycloserine, seromycin, ethambutol, myambutol, ethionamide, trector, isoniazid, INH, pyrazinamide, aldinamide, rifampicin, rifadin, rimactane, rifabutin, mycobutin, rifapentine, priftin, and streptomycin;Other antibiotics include: Arsphenamine, Salvarsan, Chloramphenicol, Chloromycetin, Fosfomycin, Monurol, Monuril, Fusidic Acid, Fucidin, Metronidazole, Flagyl, Mupirocin, Bactroban, Platensimycin, Quinupristin / Dalfopristin, Synercid, Thiamphenicol, Tigecycline, Tigacyl, Tinidazole, Tindamax, Fasigyn, Trimethoprim, Proloprim, and Trimpex; B. Biomarker Detection and Measurement
[0136] It is understood that biomarkers in a sample can be measured by any suitable method known in the art. Measuring the expression level of a biomarker can be direct or indirect. For example, the abundance level of RNA or protein can be directly quantified. Alternatively, the amount of a biomarker can be determined indirectly by measuring the abundance level of cDNA, amplified RNA, or DNA, or by measuring the quantity or activity of RNA, protein, or other molecules (e.g., metabolites) that indicate the expression level of the biomarker. Methods for measuring biomarkers in a sample have many applications. For example, measuring one or more biomarkers can aid in the diagnosis of an infection, determine an appropriate treatment for a subject, monitor a subject's response to a treatment, or identify therapeutic compounds that modulate the expression of a biomarker in vivo or in vitro. Detection of biomarker polynucleotides
[0137] In one embodiment, the expression level of a biomarker is determined by measuring the polynucleotide level of the biomarker. The transcript level of a specific biomarker gene can be determined from the amount of mRNA, or polynucleotides derived therefrom, present in a biological sample. Polynucleotides can be detected and quantified by a variety of methods, including, but not limited to, microarray analysis, polymerase chain reaction (PCR), reverse transcriptase polymerase chain reaction (RT-PCR), Northern blot, serial analysis of gene expression (SAGE), RNA switches, and solid nanovesicle detection. See, e.g., Draghici, "Data Analysis Tools for DNA Microarrays," Chapman and Hall / CRC, 2003; Simon et al., "Design and Analysis of DNA Microarray Investigations," Springer, 2004; "Real-Time PCR: Current Technology and Applications," Logan, Edwards, and Saunders (eds.), Castor Academic Press, 2009; "Bustin AZ of Quantitative PCR" (IUL Biotechnology, No. 5), International University Press, 2004; Velculescu et al. (1995) Science, 270:484-487; Matsumura et al. (2005) Cell. Microbiol., 7:11-18; "Serial Analysis of Gene Expression (SAGE): Methods and Protocols" (Methods in Molecular Biology), Humana Press, 2008, which are incorporated herein by reference in their entireties.
[0138] In one embodiment, a microarray is used to measure the levels of the biomarkers. The advantage of microarray analysis is that the expression of each of the biomarkers can be measured simultaneously, and the microarray can be specifically designed to provide a diagnostic expression profile for a particular disease or condition (e.g., sepsis).
[0139] Microarrays are prepared by selecting probes containing polynucleotide sequences and then immobilizing such probes on a solid support or surface. For example, the probes may contain DNA sequences, RNA sequences, or DNA and RNA copolymer sequences. The polynucleotide sequences of the probes may also contain DNA and / or RNA analogs, or combinations thereof. For example, the polynucleotide sequences of the probes may be full-length genomic DNA or partial fragments of genomic DNA. The polynucleotide sequences of the probes may also be synthetic nucleotide sequences, such as synthetic oligonucleotide sequences. The probe sequences may be enzymatically synthesized in vivo, enzymatically synthesized in vitro (e.g., by PCR), or non-enzymatically synthesized in vitro.
[0140] The probes used in the methods of the present invention are preferably immobilized on a solid support, which may be porous or non-porous. For example, the probe may be a polynucleotide sequence covalently attached at its 3' or 5' end to a nitrocellulose or nylon membrane or filter. Such hybridization probes are well known in the art (see, e.g., Sambrook et al., "Molecular Cloning: A Laboratory Manual" (3rd ed., 2001). Alternatively, the solid support or surface may be a glass, silicon, or plastic surface. In one embodiment, the level of hybridization is measured with a probe microarray consisting of a solid phase having immobilized thereon a population of polynucleotides, such as DNA or DNA mimics, or alternatively, RNA or RNA mimics. The solid phase may be a non-porous material, or optionally a porous material such as a gel, or a porous wafer such as a TipChip (Axela, Ontario, Canada).
[0141] In one embodiment, a microarray comprises a support or surface with an ordered array of binding (e.g., hybridization) sites or "probes," each representing one of the biomarkers described herein. The microarray is preferably an addressable array, and more preferably a positionally addressable array. More particularly, each probe of the array is preferably located at a known, predetermined location on the solid support, such that the identity (i.e., sequence) of each probe can be determined from its location in the array (i.e., on the support or surface). Each probe is preferably covalently attached to the solid support at a single site.
[0142] Microarrays can be made in a number of ways, some of which are described below. However, once made, microarrays share certain characteristics. Arrays are replicable, allowing multiple copies of a given array to be made and easily compared with each other. Microarrays are preferably made from materials that are stable under binding (e.g., nucleic acid hybridization) conditions. Microarrays are generally small, e.g., 0.1 cm 2 ~25cm 2 Although larger arrays can also be used, for example, in screening arrays, a given binding site or unique set of binding sites in a microarray will preferably specifically bind to (e.g., hybridize with) the product of a single gene in a cell (e.g., with a specific mRNA, or with a specific cDNA derived from the mRNA). However, generally, other related or similar sequences will also cross-hybridize with a given binding site.
[0143] As mentioned above, a "probe" to which a particular polynucleotide molecule specifically hybridizes contains a complementary polynucleotide sequence. Microarray probes typically consist of a nucleotide sequence no longer than 1,000 nucleotides. In some embodiments, array probes consist of a nucleotide sequence of 10 to 1,000 nucleotides. In one embodiment, the nucleotide sequence of the probes ranges from 10 to 200 nucleotides in length and is the genomic sequence of a single species of organism, such that multiple different probes are present with sequences that are complementary to, and therefore hybridize with, the genome of such species of organism, and can be ordered across all or part of the genome. In other embodiments, the probes range in length from 10 to 30 nucleotides, 10 to 40 nucleotides, 20 to 50 nucleotides, 40 to 80 nucleotides, 50 to 150 nucleotides, 80 to 120 nucleotides, or are 60 nucleotides in length.
[0144] The probes may comprise DNA or DNA "mimics" (e.g., derivatives and analogs) that correspond to portions of an organism's genome. In another embodiment, the microarray probes are complementary RNA or RNA mimics. DNA mimics are polymers composed of subunits capable of specific Watson-Crick-like hybridization with DNA or specific hybridization with RNA. Nucleic acids can be modified at the base moiety, sugar moiety, or phosphate backbone (e.g., phosphorothioate).
[0145] DNA can be obtained, for example, by polymerase chain reaction (PCR) amplification of genomic DNA or cloned sequences. PCR primers are preferably selected based on known genomic sequences that result in the amplification of specific fragments of genomic DNA. Computer programs well known in the art, such as Oligo version 5.0 (National Biosciences), are useful for designing primers with the desired specificity and optimal amplification properties. Each probe in the microarray is typically between 10 and 50,000 bases in length, usually between 20 and 200 bases in length. PCR methods are well known in the art and are described, for example, in Innis et al. (eds.), "PCR Protocols: A Guide To Methods And Applications," Academic Press Inc., San Diego, Calif. (1990), incorporated herein by reference in its entirety. Those skilled in the art will appreciate that robotic control systems are useful for isolating and amplifying nucleic acids.
[0146] Alternatively, a preferred means for generating polynucleotide probes is by synthesis of synthetic polynucleotides or oligonucleotides, for example, using N-phosphonate or phosphoramidite chemistry (Froehler et al., Nucleic Acids Res., 14:5399-5407 (1986); McBride et al., Tetrahedron Lett., 24:246-248 (1983)). Synthetic sequences are typically between about 10 and about 500 bases in length, more typically between about 20 and about 100 bases, and most preferably between about 40 and about 70 bases in length. In some embodiments, synthetic nucleic acids include unnatural bases, such as, but not limited to, inosine. As mentioned above, nucleic acid analogs can be used as binding sites for hybridization. An example of a suitable nucleic acid analog is peptide nucleic acid (see, eg, Egholm et al., Nature, 363:566-568 (1993); US Pat. No. 5,539,083).
[0147] Probes are preferably selected using algorithms that take into account binding energy, base composition, sequence complexity, cross-hybridization binding energy, and secondary structure. See Friend et al., WO 01 / 05935, published January 25, 2001; Hughes et al., Nat. Biotech., 19:342-7 (2001).
[0148] Those skilled in the art will also recognize that positive control probes, e.g., probes known to be complementary to and hybridizable with sequences in the target polynucleotide molecules, and negative control probes, e.g., probes known not to be complementary to and not hybridizable with sequences in the target polynucleotide molecules, should be incorporated into the array. In one embodiment, the positive controls are synthesized along the perimeter of the array. In another embodiment, the positive controls are synthesized diagonally across the array. In yet another embodiment, the reverse complement of each probe is synthesized next to the probe's position and used as a negative control. In yet another embodiment, sequences from other species of organisms are used as negative or "spike-in" controls.
[0149] The probes are attached to a solid support or surface, which can be made of, for example, glass, plastic (e.g., polypropylene, nylon), polyacrylamide, nitrocellulose, gel, silicon, or other porous or non-porous materials. One method for attaching nucleic acids to a surface is by printing onto a glass plate, as generally described by Schena et al., Science, 270:467-470 (1995). The COCONUT method is particularly useful for preparing cDNA microarrays (see also DeRisi et al., Nature Genetics, 14:457-460 (1996); Shalon et al., Genome Res., 6:639-645 (1996); and Schena et al., Proc. Natl. Acad. Sci. USA, 93:10539-11286 (1995), which are incorporated herein by reference in their entireties).
[0150] A second method for making microarrays produces high density oligonucleotide arrays. Techniques are known for producing arrays containing thousands of oligonucleotides complementary to defined sequences at defined locations on a surface using photolithographic techniques for in situ synthesis (see Fodor et al., 1991, Science, 251:767-773; Pease et al., 1994, Proc. Natl. Acad. Sci. USA, 91:5022-5026; Lockhart et al., 1996, Nature Biotechnology, 14:1675; U.S. Patent Nos. 5,578,832; 5,556,752; and 5,510,270, which are incorporated herein by reference in their entireties), or other methods for the rapid synthesis and deposition of defined oligonucleotides (Blanchard et al., Biosensors & Bioelectronics, 11:687-690, which are incorporated herein by reference in their entireties). Using these methods, oligonucleotides (e.g., 60-mers) of known sequence are synthesized directly on a surface such as a derivatized glass slide. Typically, the arrays produced are redundant, with multiple oligonucleotide molecules per RNA.
[0151] Other methods for making microarrays, such as masking (Maskos and Southern, 1992, Nuc. Acids. Res. 20:1679-1684, incorporated herein by reference in its entirety), can also be used. In principle, any type of array could be used, such as a dot blot on a nylon hybridization membrane (see Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., 2001). However, as will be recognized by those skilled in the art, very small arrays will often be preferred due to the reduced hybridization volume.
[0152] Microarrays can also be fabricated by means of inkjet printing devices for oligonucleotide synthesis, using the methods and systems described, for example, by Blanchard, U.S. Pat. No. 6,028,189; Blanchard et al., 1996, Biosensors and Bioelectronics, 11:687-690; and Blanchard, 1998, "Synthetic DNA Arrays," Genetic Engineering, Vol. 20, J.K. Setlow (ed.), Plenum Press, New York, pp. 111-123, which are incorporated herein by reference in their entireties. Specifically, the oligonucleotide probes in such microarrays are synthesized in the array, e.g., on a glass slide, by sequentially depositing individual nucleotide bases in "microdroplets" of a high-surface tension solvent, such as propylene carbonate. The microdroplets have small volumes (e.g., 100 pL or less, more preferably 50 pL or less) and are separated from one another (e.g., by hydrophobic domains) in the microarray to form spherical surface tension wells that define the locations of the array elements (i.e., different probes). Microarrays fabricated by this inkjet method are typically high density, preferably over 1 cm. 2 The polynucleotide probes have a density of at least about 2,500 different probes per support. The polynucleotide probes are covalently attached to the support at either the 3' or 5' end of the polynucleotide.
[0153] Biomarker polynucleotides that may be measured by microarray analysis may be expressed RNA or nucleic acids derived from RNA (e.g., cDNA, or amplified RNA derived from cDNA incorporating an RNA polymerase promoter), including naturally occurring nucleic acid molecules as well as synthetic nucleic acid molecules. In one embodiment, the target polynucleotide molecule is expressed as total cellular RNA, poly(A) +RNA includes, but is not limited to, messenger RNA (mRNA) or a fraction thereof, cytoplasmic mRNA, or RNA transcribed from cDNA (i.e., cRNA; see, e.g., Linsley and Schelter, U.S. Patent Application Serial No. 09 / 411,074, filed October 4, 1999, or U.S. Patent Nos. 5,545,522, 5,891,636, or 5,716,785). In the art, total RNA and poly(A) + Methods for preparing RNA are well known and are generally described, for example, in Sambrook et al., "Molecular Cloning: A Laboratory Manual" (3rd ed., 2001). RNA can also be extracted from cells of interest using guanidinium thiocyanate followed by CsCl centrifugation (Chirgwin et al., 1979, Biochemistry, 18:5294-5299), silica gel-based columns (e.g., RNeasy (Qiagen, Valencia, Calif.) or StrataPrep (Stratagene, LaJolla, Calif.)), or using phenol and chloroform as described in Ausubel et al., eds., 1989, "Current Protocols in Molecular Biology," Vol. III, Green Publishing Associates, Inc., John Wiley & Sons, Inc., New York, pp. 13.12.1-13.12.5. Poly(A) + RNA can be selected, for example, by selection with oligo-dT cellulose, or alternatively, by oligo-dT primed reverse transcription of total cellular RNA. RNA can be fragmented to generate RNA fragments by methods known in the art, for example, incubation with ZnCl.
[0154] In one embodiment, total RNA, mRNA, or nucleic acids derived therefrom are isolated from a sample taken from a patient with infection or inflammation. Biomarker polynucleotides that are lowly expressed in specific cells can be enriched using normalization methods (Bonaldo et al., 1996, Genome Res., 6:791-806).
[0155] As described above, biomarker polynucleotides can be detectably labeled at one or more nucleotides. Any method known in the art can be used to label target polynucleotides. This labeling preferably incorporates the label uniformly along the entire length of the RNA, and more preferably is performed with high efficiency. For example, polynucleotides can be labeled by oligo-dT-primed reverse transcription. Random primers (e.g., 9-mers) can be used in reverse transcription to incorporate labeled nucleotides uniformly throughout the entire length of the polynucleotide. Alternatively, random primers can be used in conjunction with PCR or T7 promoter-based in vitro transcription to amplify polynucleotides.
[0156] The detectable label can be a luminescent label. For example, fluorescent labels, bioluminescent labels, chemiluminescent labels, and colorimetric labels can be used to practice the invention. Fluorescent labels that can be used include, but are not limited to, derivatives of fluorescein, phosphoryl, rhodamine, or polymethine dyes. Chemiluminescent labels that can be used include, but are not limited to, luminol. In addition, commercially available fluorescent labels can be used, including, but not limited to, FluorePrime (Amersham Pharmacia, Piscataway, NJ), Fluoredite (Miilipore, Bedford, Mass.), FAM (ABI, Foster City, Calif.), and fluorescent phosphoramidites such as Cy3 or Cy5 (Amersham Pharmacia, Piscataway, NJ). Alternatively, the detectable label can be a radiolabeled nucleotide.
[0157] In one embodiment, biomarker polynucleotide molecules from a patient sample are differentially labeled from corresponding polynucleotide molecules from a reference sample. The reference may include polynucleotide molecules from a normal biological sample (i.e., a control sample, e.g., blood or PBMCs from a subject without infection or inflammation) or a reference biological sample (e.g., blood or PBMCs from a subject with a viral or bacterial infection).
[0158] Nucleic acid hybridization and washing conditions are selected so that target polynucleotide molecules specifically bind to or hybridize with specific array sites where complementary polynucleotide sequences, preferably complementary DNA, are arranged on the array. Arrays containing arranged double-stranded probe DNA are preferably subjected to denaturing conditions to convert the DNA into single strands before contacting with target polynucleotide molecules. Arrays containing single-stranded probe DNA (e.g., synthetic oligodeoxyribonucleic acid) may need to be denatured before contacting with target polynucleotide molecules, for example, to remove hairpins or dimers formed due to self-complementary sequences.
[0159] Optimal hybridization conditions will depend on the length (e.g., oligomers versus polynucleotides greater than 200 bases) and type (e.g., RNA or DNA) of the probe and target nucleic acid. Those skilled in the art will recognize that as oligonucleotides become shorter, their lengths may need to be adjusted to achieve a relatively uniform melting temperature for desirable hybridization results. General parameters for nucleic acid-specific (i.e., stringent) hybridization conditions are described in Sambrook et al., "Molecular Cloning: A Laboratory Manual" (3rd ed., 2001), and Ausubel et al., "Current Protocols in Molecular Biology," Vol. 2, Current Protocols Publishing, New York (1994). Typical hybridization conditions for cDNA microarrays according to Schena et al. are hybridization in 5x SSC + 0.2% SDS at 65°C for 4 hours, followed by a wash in low stringency wash buffer (1x SSC + 0.2% SDS) at 25°C, followed by a wash in high stringency wash buffer (0.1x SSC + 0.2% SDS) for 10 minutes at 25°C (Schena et al., Proc. Natl. Acad. Sci. USA, 93:10614 (1993)). Useful hybridization conditions are also presented, for example, in Tijessen, 1993, "Hypridization With Nucleic Acid Probes," Elsevier Science Publishers BV; and Kricka, 1992, "Nonisotopic DNA Probe Techniques," Academic Press, San Diego, Calif. Particularly preferred hybridization conditions include hybridization in 1 M NaCl, 50 mM MES buffer (pH 6.5), 0.5% sodium sarcosine, and 30% formamide at a temperature that is at or near the mean melting temperature of the probes (e.g., 51°C or less, more preferably 21°C or less).
[0160] When fluorescently labeled gene products are used, the fluorescence emission at each site on the microarray is preferably detected by scanning confocal laser microscopy. In one embodiment, a separate scan is performed using the appropriate excitation line for each of the two fluorophores used. Alternatively, a laser capable of simultaneously irradiating the sample with wavelengths specific to the two fluorophores can be used, and the emissions from the two fluorophores can be analyzed simultaneously (see Shalon et al., 1996, "A DNA microarray system for analyzing complex DNA samples using two-color fluorescent probe hybridization," Genome Research, 6:639-645, incorporated by reference in its entirety for all purposes). The array can be scanned with a laser-based fluorescence scanner equipped with a computer-controlled XY stage and a micron-range objective. Sequential excitation of the two fluorophores is achieved with a multi-line, mixed-gas laser, and the emitted light is separated by wavelength and detected with two photomultiplier tubes. Fluorescence laser scanning devices are described in Schena et al., Genome Res., 6:639-645 (1996), and other references cited herein. Alternatively, the fiber optic bundle described by Ferguson et al., Nature Biotech., 14:1681-1684 (1996) can be used to simultaneously monitor mRNA abundance at multiple sites. Alternatively, the probe can be labeled with a fluorophore and the target measured with a quencher, such that amplification is followed by measuring a decrease in signal intensity.
[0161] In certain embodiments, the invention comprises a microarray comprising a plurality of probes for detecting the expression of genes in a set of viral response genes, a set of bacterial response genes, and / or a set of sepsis response genes.
[0162] In one embodiment, the microarray comprises oligonucleotides that hybridize to IFI27 polynucleotides, oligonucleotides that hybridize to JUP polynucleotides, oligonucleotides that hybridize to LAX1 polynucleotides, oligonucleotides that hybridize to HK3 polynucleotides, oligonucleotides that hybridize to TNIP1 polynucleotides, oligonucleotides that hybridize to GPAA1 polynucleotides, and oligonucleotides that hybridize to CTSB polynucleotides.
[0163] In another embodiment, the microarray further comprises oligonucleotides that hybridize to CEACAM1 polynucleotides, oligonucleotides that hybridize to ZDHHC19 polynucleotides, oligonucleotides that hybridize to C9orf95 polynucleotides, oligonucleotides that hybridize to GNA15 polynucleotides, oligonucleotides that hybridize to BATF polynucleotides, oligonucleotides that hybridize to C3AR1 polynucleotides, oligonucleotides that hybridize to KIAA1370 polynucleotides, oligonucleotides that hybridize to TGFBI polynucleotides, oligonucleotides that hybridize to MTCH1 polynucleotides, oligonucleotides that hybridize to RPGRIP1 polynucleotides, and oligonucleotides that hybridize to HLA-DPB1 polynucleotides.
[0164] Polynucleotides can also be analyzed by other methods, including, but not limited to, Northern blotting, nuclease protection assays, RNA fingerprinting, polymerase chain reaction, ligase chain reaction, Qbeta replicase, isothermal amplification, strand displacement amplification, transcription-based amplification systems, nuclease protection (S1 nuclease protection assay or RNase protection assay), SAGE, as well as the methods disclosed in WO 88 / 10315 and WO 89 / 06700, and International Applications PCT / US87 / 00880 and PCT / US89 / 01025, which are incorporated by reference in their entireties.
[0165] Standard Northern blot assays can be used to confirm the size of RNA transcripts or, alternatively, to identify the relative amounts of spliced RNA transcripts and mRNA in a sample, according to conventional Northern hybridization techniques known to those skilled in the art. In Northern blots, RNA samples are first separated by size by electrophoresis in agarose gels under denaturing conditions. The RNA is then transferred to a membrane, crosslinked, and hybridized with a labeled probe. Nonisotopic radiolabeled probes or highly radiospecific radiolabeled probes can be used, including random-primed, nick-translated, or PCR-generated DNA probes, in vitro-transcribed RNA probes, and oligonucleotides. In addition, sequences with only partial homology (e.g., cDNAs from different species or genomic DNA fragments that may contain exons) can be used as probes. Labeled probes, such as full-length DNA, single-stranded DNA, or radiolabeled cDNA containing fragments of these DNA sequences, can be at least 20, at least 30, at least 50, or at least 100 contiguous nucleotides in length. Probes can be labeled by any of many different methods known to those skilled in the art. The most commonly employed labels in these studies are radioactive elements, enzymes, chemicals that fluoresce when exposed to ultraviolet light, and other labels. Numerous fluorescent materials are known and can be utilized as labels. These include, but are not limited to, fluorescein, rhodamine, auramine, Texas Red, AMCA Blue, and Lucifer Yellow. A particular detection material is anti-rabbit antibody prepared in goats and conjugated to fluorescein via isothiocyanate. Proteins can also be labeled with radioactive elements or enzymes. Radioactive labels can be detected by any of the currently available counting procedures. Isotopes that can be used include: 3 H, 14 C. 32 P, 35 S, 36 Cl, 35 Cr, 57 Co, 58Co, 59 Fe, 90 Y, 125 I, 131 I, and 186 Examples of suitable labeling materials include, but are not limited to, Re. Enzyme labels are also useful and can be detected by any of the currently available colorimetric, spectrophotometric, fluorospectrophotometric, amperometric, or gasometric methods. The enzyme is conjugated to the selected particle by reaction with a bridging molecule such as carbodiimide, diisocyanate, or glutaraldehyde. Any enzyme known to those skilled in the art can be utilized. Examples of such enzymes include, but are not limited to, peroxidase, beta-D-galactosidase, urease, glucose oxidase plus peroxidase, and alkaline phosphatase. See, for example, U.S. Patent Nos. 3,654,090, 3,850,752, and 4,016,043 for disclosure of alternative labeling materials and methods.
[0166] Nuclease protection assays (including both ribonuclease protection assays and S1 nuclease assays) can be used to detect and quantify specific mRNAs. In a nuclease protection assay, an antisense probe (labeled, e.g., radioactively or nonisotopically labeled) is hybridized to an RNA sample in solution. After hybridization, the single-stranded, unhybridized probe and RNA are degraded by nucleases. Acrylamide gels are used to separate the remaining protected fragments. Solution hybridization is more efficient than membrane-based hybridization and can typically accommodate up to 100 μg of sample RNA, compared to a maximum of 20–30 μg for blot hybridization.
[0167] The most common type of nuclease protection assay, the ribonuclease protection assay, requires the use of RNA probes. Assays containing S1 nuclease can only use oligonucleotides and other single-stranded DNA probes. The single-stranded antisense probe is typically completely homologous to the target RNA to prevent cleavage of the probe:target hybrid by the nuclease.
[0168] Serial analysis of gene expression (SAGE) can also be used to determine the abundance of RNA in a cell sample. See, for example, Velculescu et al., 1995, Science, 270:484-7; Carulli et al., 1998, Journal of Cellular Biochemistry Supplements, 30 / 31:286-96, which are incorporated herein by reference in their entirety. SAGE analysis does not require specialized devices for detection and is one of the preferred analytical methods for simultaneously detecting the expression of multiple transcripts. First, polyA +RNA is extracted from cells. Next, using a biotinylated oligo(dT) primer, the RNA is converted into cDNA and treated with a four-base-recognizing restriction enzyme (anchoring enzyme: AE), resulting in AE-treated fragments containing a biotin group at the 3' end. The AE-treated fragments are then incubated with streptavidin for binding. The bound cDNA is divided into two fractions, and each fraction is then ligated to a different double-stranded oligonucleotide adapter (linker) A or B. These linkers consist of (1) a single-stranded overhanging portion with a sequence complementary to the overhanging portion formed by the action of the anchoring enzyme; (2) a 5'-nucleotide that recognizes the sequence of a type IIS restriction enzyme (which cuts at a predetermined position no more than 20 bp away from the recognition site) used as a tagging enzyme (TE); and (3) an additional sequence of sufficient length to construct a PCR-specific primer. The linker-ligated cDNA is cleaved using a tagging enzyme, leaving only the linker-ligated cDNA sequence portion present in the form of a short sequence tag. Next, a pool of short sequence tags derived from two different types of linkers is ligated to each other, followed by PCR amplification using primers specific to linkers A and B. As a result, an amplified product is obtained as a mixture containing an infinite number of sequences with two adjacent sequence tags (double tags) attached to linkers A and B. The amplified product is treated with an anchoring enzyme, and the free double tag portions are ligated into a chain in a standard ligation reaction. The amplified product is then cloned. The nucleotide sequence of the clone can be used to obtain a readout of a certain length of continuous double tags. The presence of mRNA corresponding to each tag can then be identified from the nucleotide sequence of the clone and information about the sequence tags.
[0169] Quantitative reverse transcriptase PCR (qRT-PCR) can also be used to determine the expression profile of biomarkers (see, e.g., U.S. Patent Application Publication No. 2005 / 0048542A1, incorporated herein by reference in its entirety). The first step in gene expression profiling by RT-PCR is reverse transcription of the RNA template into cDNA, followed by its exponential amplification in a PCR reaction. The two most commonly used reverse transcriptases are avian myeloblastosis virus reverse transcriptase (AMV-RT) and Moloney murine leukemia virus reverse transcriptase (MLV-RT). The reverse transcription step is typically primed using specific primers, random hexamers, or oligo-dT primers, depending on the context and goal of expression profiling. For example, extracted RNA can be reverse transcribed using a GeneAmp RNA PCR kit (Perkin Elmer, Calif., USA) according to the manufacturer's instructions. The derived cDNA can then be used as a template in the subsequent PCR reaction.
[0170] The PCR step can use a variety of thermostable, DNA-dependent DNA polymerases, but typically employs Taq DNA polymerase, which possesses 5'-3' nuclease activity but lacks 3'-5' proofreading endonuclease activity. Thus, TAQMAN PCR typically utilizes the 5' nuclease activity of Taq or Tth polymerase to hydrolyze hybridization probes bound to its target amplicon, although any enzyme with equivalent 5' nuclease activity can be used. Two oligonucleotide primers are used to generate an amplicon typical of a PCR reaction. A third oligonucleotide, or probe, is designed to detect the nucleotide sequence located between the two PCR primers. The probe is non-extendible by Taq DNA polymerase enzyme and is labeled with a reporter fluorescent dye and a quencher fluorescent dye. When the two dyes are placed close together, as in the probe, any laser-induced emission from the reporter dye is quenched by the quencher dye. During the amplification reaction, the Taq DNA polymerase enzyme cleaves the probe in a template-dependent manner. The resulting probe fragments dissociate in solution, and the signal from the released reporter dye is liberated from the quenching effect of the second fluorophore. For each new molecule synthesized, one molecule of reporter dye is released, and detection of the unquenched reporter dye provides the basis for quantitative interpretation of the data.
[0171] TAQMAN RT-PCR can be performed using commercially available equipment, such as the ABI PRISM 7700 Sequence Detection System (Perkin-Elmer-Applied Biosystems, Foster City, Calif., USA) or the Lightcycler (Roche Molecular Biochemicals, Mannheim, Germany). Alternatives include, but are not limited to, sample-to-answer (STA) point-of-care devices such as the cobas Liat (Roche Molecular Diagnostics, Pleasanton, Calif., USA) or the GeneXpert system (Cepheid, Sunnyvale, Calif., USA). Those skilled in the art will appreciate that the present invention is not limited to the listed devices, and that other devices may also be used for TAQMAN-PCR. In a preferred embodiment, the 5' nuclease step is performed in a real-time quantitative PCR device, such as the ABI PRISM 7700 Sequence Detection System. The system consists of a thermocycler, laser, charge-coupled device (CCD), camera, and computer. The system includes software for running the instrument and analyzing the data. 5' nuclease assay data are initially expressed as Ct, or cycle threshold. Fluorescence values are recorded at every cycle and represent the amount of product amplified up to this point in the amplification reaction. The cycle threshold (Ct) is the point at which the fluorescent signal is first recorded as statistically significant. Alternatives to standard thermal cycling include, but are not limited to, continuous thermal gradient amplification or isothermal amplification with endpoint detection, and other devices known to those skilled in the art. To minimize errors and the effects of sample-to-sample variability, RT-PCR is typically performed using an internal standard. An ideal internal standard is expressed at consistent levels across different tissues and is unaffected by experimental treatments. The RNAs most frequently used to normalize gene expression patterns are the mRNAs of the housekeeping genes glyceraldehyde-3-phosphate dehydrogenase (GAPDH) and beta-actin.
[0172] A more recent variation of the RT-PCR method is real-time quantitative PCR, which uses a dual-labeled fluorogenic probe (i.e., TAQMAN probe) to measure the accumulation of PCR products. Real-time PCR is compatible with both quantitative competitive PCR, which uses an internal competitor for each target sequence for normalization, and quantitative comparative PCR, which uses a normalization gene contained in the sample or a housekeeping gene for RT-PCR. For further details, see, for example, Held et al., Genome Research, 6:986-994 (1996).
[0173] An alternative method is to detect PCR products using digital counting methods. These alternative methods include, but are not limited to, digital droplet PCR and solid nanovesicle PCR product detection. In these methods, the count of the product of interest can be normalized to the count of a housekeeping gene. Other methods of PCR detection known to those skilled in the art can also be used, and the present invention is not limited to the listed methods. Biomarker data analysis
[0174] To diagnose the type of infection, including assessing whether a patient has inflammation arising from non-infectious sources, such as traumatic injury, surgery, autoimmune disease, thrombosis, or systemic inflammatory response syndrome (SIRS), or an infection, and, if the patient is diagnosed with an infection, determining whether the patient has a viral or bacterial infection, the biomarker data can be analyzed by a variety of methods to identify biomarkers and determine the statistical significance of differences in biomarker levels observed between test and reference expression profiles. In certain embodiments, patient data is analyzed by one or more methods, including, but not limited to, multivariate linear discriminant analysis (LDA), receiver operating characteristic (ROC) analysis, principal component analysis (PCA), ensemble data mining methods, microarray significance analysis (SAM), cell-specific microarray significance analysis (csSAM), spanning-tree progression analysis of density-normalized events (SPADE), and multidimensional protein identification technology (MUDPIT) analysis.(See, e.g., Hilbe (2009), "Logistic Regression Models," Chapman & Hall / CRC Press; McLachlan (2004), "Discriminant Analysis and Statistical Pattern Recognition," Wiley Interscience; Zweig et al. (1993), Clin. Chem., 39:561-577; Pepe (2003), "The statistical evaluation of medical tests for classification and prediction," New York, NY: Oxford; Sing et al. (2005), Bioinformatics, 21:3940-3941; Tusher et al. (2001), Proc. Natl. Acad. Sci. USA, 98:5116-5121; Oza (2006), "Ensemble data mining," NASA Ames Research Center, Moffett, which are incorporated herein by reference in their entireties. Field, CA, USA; English et al. (2009), J. Biomed. Inform., 42(2):287-295; Zhang (2007), Bioinformatics, 8:230; Shen-Orr et al. (2010), Journal of Immunology, 184:144-130; Qiu et al. (2011), Nat. Biotechnol, 29(10):886-891; Ru et al. (2006), J. Chromatogr. A., 1111(2):166-174; Jolliffe, "Principal Component Analysis", (Springer Series in Statistics, 2nd ed., Springer, NY, 2002); Koren et al. (2004), IEEE Trans Vis Comput Graph, 10:459-470. C. Kit
[0175] In yet another aspect, the present invention provides a kit for diagnosing an infection in a subject, which kit can be used to detect a biomarker of the present invention. For example, the kit can be used to detect any one or more of the biomarkers described herein that are differentially expressed in samples from patients with viral or bacterial infections and healthy or uninfected subjects. The kit can include one or more agents for measuring the expression levels of a set of viral response genes and a set of bacterial response genes; a container for holding a biological sample isolated from a human subject suspected of having an infection; and instructions for reacting the agents with the biological sample or a portion of the biological sample to measure the expression levels of the set of viral response genes and the set of bacterial response genes in the biological sample. The agents can be packaged in separate containers. The kit can further include one or more control samples and reagents for performing immunoassays, PCR, or microarray analysis.
[0176] In one embodiment, the kit comprises agents for measuring the levels of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, and CTSB to distinguish viral from bacterial infections.
[0177] In another embodiment, the kit further comprises agents for measuring the levels of CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1, which are biomarkers for distinguishing whether inflammation is caused by an infectious or non-infectious source.
[0178] In certain embodiments, the kit further comprises a microarray for analyzing a plurality of biomarker polynucleotides, hi one embodiment, the microarray comprises oligonucleotides that hybridize to an IFI27 polynucleotide, an oligonucleotide that hybridizes to a JUP polynucleotide, an oligonucleotide that hybridizes to an LAX1 polynucleotide, an oligonucleotide that hybridizes to an HK3 polynucleotide, an oligonucleotide that hybridizes to a TNIP1 polynucleotide, an oligonucleotide that hybridizes to a GPAA1 polynucleotide, and an oligonucleotide that hybridizes to a CTSB polynucleotide.
[0179] In another embodiment, the kit further comprises a microarray comprising oligonucleotides that hybridize to CEACAM1 polynucleotides, oligonucleotides that hybridize to ZDHHC19 polynucleotides, oligonucleotides that hybridize to C9orf95 polynucleotides, oligonucleotides that hybridize to GNA15 polynucleotides, oligonucleotides that hybridize to BATF polynucleotides, oligonucleotides that hybridize to C3AR1 polynucleotides, oligonucleotides that hybridize to KIAA1370 polynucleotides, oligonucleotides that hybridize to TGFBI polynucleotides, oligonucleotides that hybridize to MTCH1 polynucleotides, oligonucleotides that hybridize to RPGRIP1 polynucleotides, and oligonucleotides that hybridize to HLA-DPB1 polynucleotides.
[0180] The kit may include one or more containers for the compositions contained in the kit. The compositions may be in liquid form or may be lyophilized. Suitable containers for the compositions include, for example, bottles, vials, syringes, and test tubes. The containers may be made of a variety of materials, including glass or plastic. The kit may also include a package insert containing instructions for how to diagnose the infection.
[0181] The kits of the present invention have many applications. For example, the kits can be used to determine whether a subject has an infection or some other inflammatory condition arising from a non-infectious source, such as traumatic injury, surgery, autoimmune disease, thrombosis, or systemic inflammatory response syndrome (SIRS). If a patient is diagnosed with an infection, the kit can be used to further determine the type of infection (i.e., viral or bacterial infection). In another example, the kit can be used to determine whether a patient with acute inflammation should be treated with, for example, a broad-spectrum antibiotic or an antiviral agent. In another example, the kit can be used to monitor the effectiveness of treatment of a patient with an infection. In a further example, the kit can be used to identify compounds that modulate the expression of one or more of the biomarkers in an in vitro or in vivo animal model to determine the effectiveness of treatment. D. Diagnostic Systems and Computerized Methods for Diagnosis of Infection
[0182] In a further aspect, the invention includes a computer-implemented method for diagnosing a patient suspected of having an infection, wherein the computer performs the steps of receiving input patient data including values for the expression levels of one or both of a set of viral response genes and a set of bacterial response genes in a biological sample from the patient, analyzing the expression levels of the set of genes, calculating a bacterial / viral metascore for the patient based on the expression levels of the set of genes, the value of the bacterial / viral metascore indicating whether the patient has a viral or bacterial infection, and displaying information regarding the patient's diagnosis.
[0183] In certain embodiments, the input patient data includes: a) a set of viral response genes including IFI27, JUP, and LAX1, and a set of bacterial response genes including HK3, TNIP1, GPAA1, and CTSB; b) a set of viral response genes including OAS2 and CUL1, and a set of bacterial response genes including SLC12A9, ACPP, STAT5B; c) a set of viral response genes including ISG15 and CHST12, and a set of bacterial response genes including EMR1 and FLII; d) a set of viral response genes including IFIT1, SIGLEC1, and ADA, and a set of bacterial response genes including PTAFR, NRD1, PLP2; e) a set of viral response genes including MX1, and a set of bacterial response genes including DYSF, TWF2; f) a set of viral response genes including RSAD2 g) a set of viral response genes including IFI44L, GZMB, and KCTD14, and a set of bacterial response genes including TBXAS1, ACAA1, and S100A12; h) a set of viral response genes including LY6E, and a set of bacterial response genes including PGD and LAPTM5; i) a set of viral response genes including IFI44, HESX1, and OASL, and a set of bacterial response genes including NINJ2, DOK3, SORL1, and RAB31; j) a set of viral response genes including OAS1, and a set of bacterial response genes including IMPA2 and LTA4H.
[0184] In another embodiment, the invention comprises a computer-implemented method for diagnosing a patient suspected of having an infection, wherein the computer performs the steps of: a) receiving input patient data including values for the levels of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, and CTSB in a biological sample from the patient; b) analyzing the level of each of the biomarkers and comparing it to the biomarker's respective reference value range; c) calculating a bacterial / viral metascore for the patient based on the expression levels of the biomarkers, wherein a positive bacterial / viral metascore for the patient indicates that the patient has a viral infection, and a negative bacterial / viral metascore for the patient indicates that the patient has a bacterial infection; and d) displaying information regarding the patient's diagnosis.
[0185] In certain embodiments, the input patient data further includes values for the expression levels of a set of sepsis response genes including CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1, wherein the computer-implemented method further includes calculating a sepsis metascore for the patient, wherein a sepsis metascore above the reference range for non-infected control subjects indicates that the patient has an infection, and wherein a sepsis metascore within the reference range for non-infected control subjects indicates that the patient has a non-infectious inflammatory condition.
[0186] In another embodiment, the invention provides a computer-implemented method for diagnosing a patient with inflammation, the method comprising the steps of: a) receiving input patient data, the input patient data including values for the levels of the biomarkers IFI27, JUP, LAX1, HK3, TNIP1, GPAA1, CTSB, CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1 in a biological sample from the patient; b) analyzing the level of each of the biomarkers and comparing it to the biomarker's respective reference value range; and c) calculating a sepsis metascore for the patient. a) calculating a bacterial / viral metascore for the patient if the sepsis score indicates that the patient has an infection, wherein a positive bacterial / viral metascore for the patient indicates that the patient has a viral infection, and a negative bacterial / viral metascore for the patient indicates that the patient has a bacterial infection; and b) displaying information regarding the patient's diagnosis.
[0187] In a further aspect, the present invention includes a diagnostic system for performing the computer-implemented methods described. The diagnostic system includes a computer containing a processor, a storage component (i.e., memory), a display component, and other components typically found in a general-purpose computer. The storage component stores information accessible by the processor, including instructions that may be executed by the processor and data that may be retrieved, manipulated, or stored by the processor.
[0188] The storage component includes instructions for determining a patient's diagnosis. For example, the storage component includes instructions for calculating a bacterial / viral metascore and / or a sepsis metascore, as described herein (see Example 1). In addition, the storage component may further include instructions for performing multivariate linear discriminant analysis (LDA), receiver operating characteristic (ROC) analysis, principal component analysis (PCA), ensemble data mining, cell-specific microarray significance analysis (csSAM), or multidimensional protein identification technology (MUDPIT) analysis. The computer processor is coupled to the storage component and configured to receive the patient data and execute the instructions stored in the storage component to analyze the patient data according to one or more algorithms. The display component displays information regarding the patient's diagnosis.
[0189] The storage component may be any type of storage component capable of storing information accessible by the processor, such as a hard drive, memory card, ROM, RAM, DVD, CD-ROM, USB flash drive, writable memory, and read-only memory. The processor may be any known processor, such as a processor manufactured by Intel Corporation. Alternatively, the processor may be a dedicated controller, such as an ASIC.
[0190] Instructions may be any set of instructions that are executed directly by a processor (such as machine code) or indirectly by a processor (such as a script). In this regard, the terms "instructions," "steps," and "program" may be used interchangeably herein. Instructions may be stored in object code form for direct processing by a processor, or in any other computer language, including a script or a collection of independent source code modules that are interpreted on demand or pre-compiled.
[0191] Data can be retrieved, stored, and modified by a processor in accordance with instructions. For example, the diagnostic system is not limited to any particular data structure, but data can be stored in a computer register, a relational database as a table with multiple distinct fields and records, an XML document, or a flat file. Data can also be formatted in any computer-readable format, such as, but not limited to, binary values, ASCII, or Unicode. Furthermore, data can include any information sufficient to identify relevant information, such as numbers, descriptive text, proprietary codes, pointers, data stored in other memory (including other network locations), or information used by a function to calculate relevant data.
[0192] In certain embodiments, a processor and storage component may include multiple processors and storage components, which may or may not be stored within the same physical enclosure. For example, some instructions and data may be stored on a removable CD-ROM, while other instructions and data may be stored within a read-only computer chip. Some or all of the instructions and data may be stored in a location physically remote from the processor, but still accessible by the processor. Similarly, a processor may in fact include a collection of processors, which may or may not operate in parallel.
[0193] In one aspect, the computer is a server that communicates with one or more client computers. Like a server, each client computer can be configured with a processor, storage components, and instructions. Each client computer can be a personal computer intended for personal use and having all the internal components typically found in a personal computer, such as a central processing unit (CPU), a display (e.g., a monitor that displays information processed by the processor), a CD-ROM, a hard drive, user input devices (e.g., a mouse, keyboard, touchscreen, or microphone), speakers, a modem, and / or a network interface device (telephone, cable, or other form), as well as all of the components used to connect these elements and enable them to communicate (directly or indirectly) with each other. Furthermore, computers in accordance with the systems and methods described herein can include any device capable of processing instructions and transmitting data to and from humans and other computers, including network computers lacking local storage capabilities.
[0194] While the client computer may include a full-sized personal computer, many aspects of the system and method are particularly advantageous when used in conjunction with a mobile device capable of wirelessly exchanging data with a server over a network such as the Internet. For example, the client computer may be a wireless-enabled PDA, such as a Blackberry phone, an Apple iPhone, an Android, or other Internet-enabled cell phone. In such a regard, a user may input information using a small keyboard, keypad, touchscreen, or any other user input means. The computer may have an antenna for receiving wireless signals.
[0195] The server and client computers can communicate directly and indirectly, such as through a network. While only a few computers may be used, it should be appreciated that a typical system may include many connected computers, with each different computer residing at a different node of the network. The network and intervening nodes may include a variety of combinations of devices and communication protocols, including the Internet, the World Wide Web, intranets, virtual private networks, wide area networks, local networks, cell phone networks, private networks using communication protocols proprietary to one or more companies, Ethernet, WiFi, and HTTP. Such communication may be facilitated by any device capable of transmitting data to and from other computers, such as modems (e.g., dial-up or cable), networks, and wireless interfaces. The server may be a web server.
[0196] As noted above, certain advantages are obtained when transmitting or receiving information, but other aspects of the systems and methods are not limited to any particular manner of transmitting information. For example, in some aspects, information can be sent via media such as disk, tape, flash drive, DVD, or CD-ROM. In other aspects, information can be sent in a non-electronic format and manually entered into the system. Still further, while certain functions are indicated to occur at a server and other functions at a client, various aspects of the systems and methods can be implemented by a single computer having a single processor. [Example]
[0197] III. Experiment Below are examples of specific embodiments for carrying out the present invention. The examples are presented for illustrative purposes only and are not intended to limit the scope of the invention in any way.
[0198] Efforts have been made to ensure accuracy with respect to numbers used (eg, amounts, temperature, etc.), but some experimental error and deviation should, of course, be allowed for.
[0199] Example 1 Robust classification of bacterial and viral infections via integrated host gene expression diagnostics Introduction Here, we sought to improve the diagnostic power of the Sepsis Metascore (SMS) by adding the ability to distinguish bacterial from viral infections. Therefore, to derive new biomarkers for differentiating infection types, we applied our multicohort analysis framework to clinical microarray cohorts to compare host responses to bacterial and viral infections. We further developed a novel method to conormalize gene expression data across multiple cohorts, allowing for direct comparison of diagnostic scores across cohorts. Finally, we combined the Sepsis Metascore and the novel bacterial / viral diagnostic method with an integrated antibiotic decision model (IADM), which can determine whether patients with acute inflammation from any source have an underlying bacterial infection.
[0200] result Derivation of bacterial / viral metascores for seven genes Our previously published 11-gene SMS failed to reliably distinguish between bacterial and viral infections, and most showed insignificant differences in score distribution between patients with bacterial and viral infections (Figures 5A and 5B). 15 We hypothesized that a classifier for viral versus bacterial infections would allow for improved diagnostic models. Therefore, we performed a systematic search for gene expression microarray cohorts studying patients with viral and / or bacterial infections. We included eight cohorts containing N>5 patients with both viral and bacterial infections.11、18~26 We identified 426 patient samples (142 viral infections and 284 bacterial infections) from both whole blood and PBMCs (Table 1A). The eight cohorts consisted of 426 patient samples (142 viral infections and 284 bacterial infections), including pediatric and adult patients, medical and surgical patients, and patients with multiple infection sites. We performed a multicohort analysis of the eight cohorts as previously described (Figure 6). 7、15、16、27 We set the significance thresholds for the leave-one-dataset-out round-robin analysis at an effect size of >2-fold and an FDR of <1%. However, to ensure that neither tissue type biased the results, we further selected only genes with an effect size >1.5-fold in both the PBMC and whole blood cohorts separately. This step resulted in 72 significantly differentially expressed genes (Supplementary Table 1). We then performed a forward search using the greedy method. 7 Using this method, we found a diagnostically optimized gene set, resulting in seven genes (high viral infection: IFI27, JUP, LAX1; high bacterial infection: HK3, TNIP1, GPAA1, CTSB; Figures 7A-7B). As expected, a "bacterial / viral metascore" based on these seven genes robustly distinguished viral from bacterial infection in all eight discovery cohorts (Summary ROC AUC = 0.97, 95% CI = 0.89-0.99, Figures 1A and 8).
[0201] We next compared the six remaining independent clinical cohorts directly between bacterial and viral infections (a total of 341 samples, 138 bacterial infections and 203 viral infections). 13、14、28~30 We investigated a set of seven genes and found a summary ROC AUC of 0.91 (95% CI = 0.82-0.96) (Table 1B, Figure 1B, Figure 9). As a test of the generalizability of the signature, we also investigated whether cells stimulated with LPS or influenza virus in vitro could be separated by the bacterial / viral metascore (GSE53166).31 , N=75, AUC=0.99) (Figure 10).
[0202] Global Validation via COCONUT Co-Normalization While many microarray cohorts have been published that have studied either bacterial or viral infections, none have studied both, making direct (within dataset) estimation of diagnostic power for separating bacterial and viral diseases impossible. To apply and compare gene scores across these cohorts, a new method was needed that could remove batch effects between datasets while maintaining unbiased diagnosis for affected patients. Here, we used ComBat to obtain unbiased correction of disease samples relative to healthy controls. 32 We designed and implemented a novel type of array normalization using an empirical Bayes normalization method (which we refer to as the COMbat CO-Normalization Using Control method, or "COCONUT" method, in the following section and Figure 11). Importantly, while each gene retains the same distribution between disease and control within each dataset, housekeeping genes remain invariant across both disease and cohorts after COCONUT co-normalization (Figures 12A and 12B). Because the method assumes that all healthy samples are derived from the same distribution, and different immune cell types have significantly different baseline gene expression distributions, we separate whole blood samples from PBMC samples. Using COCONUT co-normalization, we were able to show that the bacterial / viral metascore had a global AUC of 0.92 (95% CI: 0.89-0.96) in the discovery cohort (Figure 2; pre-normalized data: Figure 14). We then applied this method to all published microarray cohorts that met the inclusion criteria and used whole blood (20 cohorts that assessed bacterial or viral infection, but not both, in addition to the four direct validation cohorts that included control patients). 33~49The bacterial / viral metascore was tested across a total of 1,040 cohorts (including 143 + 897 = 1,040), demonstrating an overall ROC AUC of 0.93 (95% CI: 0.91-0.94) across these data (Table 2, Figure 13; pre-normalized data: Figure 15). The wide clinical diversity of the data, including a wide range of infection types (Gram-positive, Gram-negative, atypical bacteria, common respiratory viruses, and dengue virus) and severities (mild infection to septic shock), is particularly striking. Thus, we were able to establish a single cutoff (shown as a horizontal dotted line) across all cohorts. Finally, we performed the same procedure separately on the available PBMC validation cohorts (six cohorts). 50~54 N=259, global AUC=0.92 (95% CI: 0.87-0.97, Figure 16; pre-normalized data: Figure 17). Using COCONUT co-normalization, the AUCs of the three global ROCs (discovery whole blood = 0.92, validation whole blood = 0.93, validation PBMC = 0.92) were all nearly comparable to the summary AUC of the direct validation cohort (0.91), providing a high degree of confidence in this level of diagnostic power.
[0203] Supplementary Table 4 shows the bacterial / viral metascores for all combinations of two genes selected from 71 gene sets obtained by iterating the greedy forward algorithm in the discovery dataset. All combinations of two genes from the 71 gene sets show a mean AUC greater than or equal to 0.80 (≥ 0.80). In comparison, Figure 18 shows the distribution of mean AUCs in the discovery dataset for 10,000 randomly selected two-gene pairs, demonstrating that an AUC greater than or equal to 0.80 is not achievable by chance alone. As illustrated in Figure 18, randomly selected two-gene pairs result in a normal distribution of mean AUCs bounded by greater than 0.2 (> 0.20) and less than 0.80 (< 0.80). Two-gene combinations presented in Supplementary Table 4 with an AUC equal to or greater than 0.80 (≥ 0.80) provide clinically useful decisions about whether an infection is viral or bacterial.
[0204] Integrated antibiotic decision model A key clinical need is to diagnose whether patients with signs and symptoms of inflammation have an underlying bacterial infection, as prompt and appropriate administration of antibiotics is key to improving patient outcomes. Neither the SMS nor the bacterial / viral metascore alone can robustly distinguish between all three classes: (1) non-infectious inflammation, (2) bacterial disease, and (3) viral disease. Therefore, to increase clinical relevance, we first developed a novel SMS that we have previously described. 7We investigated an "integrated antibiotic decision model" (IADM), which applies a criterion to examine the presence of infection, and then examine samples that test positive for infection using a bacterial / viral metascore (Figure 3A). As mentioned above, the only way to simultaneously establish test features for IADM across cohorts is to use COCONUT co-normalization. However, we found that SMS in COCONUT-co-normalized data was strongly affected by age differences between healthy and infected patients, or both (Figures 19A and 19B). Therefore, we excluded cohorts focused on infants (children <1 year old) from IADM, resulting in a total of 20 cohorts (N = 1,057). The resulting global AUC for SMS across the available data was 0.86 (95% CI: 0.84-0.89) (Supplementary Table 2; Figures 20A and 20B). We set a global threshold of 95% for SMS sensitivity for infection and a global threshold of 95% for bacterial / viral metascore sensitivity for bacterial infection. This resulted in overall sensitivities and specificities of 94.0% and 59.8%, respectively, for bacterial infection and 53.0% and 90.6%, respectively, for viral infection (Figures 3A-3C). The overall sensitivity and specificity remained largely unchanged when healthy patients were included in the non-infected class (Figures 21A and 21B). Thus, the overall positive and negative likelihood ratios for bacterial infection in IADM were 2.34 (LR+) and 0.10 (LR-), whereas a recent meta-analysis of procalcitonin showed a negative LR of 0.29 (95% CI: 0.22-0.38). 55 We plotted the NPV and PPV for these test characteristics against prevalence, and found that the NPV and PPV for bacterial infection at a prevalence of 15% are 98.3% and 29.2% (Figure 22).
[0205] A dataset (GSE63990) that includes non-infectious SIRS patients and patients with both bacterial and viral illnesses, but does not include healthy controls, precluding their inclusion in the global calculations.14 ), but only one was present. Therefore, we investigated IADM with a locally derived test threshold. We found a sensitivity and specificity for all bacterial infections of 94.3% and 52.2%, respectively (Figures 21A and 21B). Validation with NanoString
[0206] Finally, we present a novel method for targeting NanoString nCounter 56 We used gene expression assays to validate these results in individual whole blood samples from children with sepsis in the Genomics of Pediatric SIRS and Septic Shock Investigators (GPSSSI) cohort (total N = 96: 36 SIRS, 49 bacterial sepsis, and 11 viral sepsis patients; Figures 4A–4E). The GPSSSI cohort was also leveraged by the GSE66099 dataset, but the children profiled here were not part of the discovery dataset because they were not profiled via microarray. In the NanoString validation cohort, the AUC of SMS was 0.81 (AUC in GSE66099: 0.80). Similarly, the AUC of the bacterial / viral metascore was 0.84 (AUC in GSE66099: 0.83). Thus, the AUC of the microarray was preserved when new patients were examined with targeted gene expression assays. Applying the same IADM, the sensitivity and specificity were 89.7% and 70.0% for bacterial infections and 54.5% and 96.5% for viral infections, respectively.
[0207] Consideration Good diagnostics for acute infections are required in both inpatient and outpatient settings. In low-risk outpatient settings, even a simple diagnostic that can distinguish bacterial from viral infections may be sufficient to guide appropriate antibiotic use. In high-risk settings, excluding non-infectious inflammatory causes becomes increasingly important, so decision models for antibiotic prescribing must incorporate non-infectious (non-healthy) cases. Therefore, reliable diagnosis requires distinguishing between all three cases (non-infectious inflammation, bacterial infection, and viral infection). Here, using 426 samples from eight cohorts, we derived a set of only seven genes that can accurately distinguish bacterial from viral infections across a wide range of clinical conditions in independent cohorts (a total of 30 cohorts consisting of 1,299 patients). By coupling our previous sepsis metascore (which identifies the presence or absence of infection) with this new bacterial / viral metascore (which determines the type of infection) into a single integrated antibiotic decision model, we further demonstrate that we can determine with high accuracy which patients will benefit from antibiotics. Finally, we confirmed the diagnostic power of both the seven-gene set and IADM using a targeted NanoString assay in independent samples, demonstrating that the signature retains diagnostic power even without relying on microarrays.
[0208] IADM has a small negative likelihood ratio (0.10) and a large estimated NPV, implying its potential usefulness as a rule-out test. A meta-analysis of procalcitonin, including 3,244 patients from 30 studies, resulted in an overall estimated negative likelihood ratio of 0.29 (95% CI: 0.22-0.38). 55It is noteworthy that the negative likelihood ratios for IADM are therefore significantly smaller than those for procalcitonin. Furthermore, these test characteristics do not presuppose knowledge about the patient and are only estimates of the actual clinical utility of such tests. Medical history and physical examination, vital signs, and laboratory values may all also aid in diagnosis. Even with these caveats, recent efficient decision models for screening ICU patients for hospital-acquired infections suggest that tests such as IADM can accurately diagnose bacterial and viral infections and may be cost-effective. 57 Ultimately, only interventional trials will be possible to establish the cost-effectiveness and clinical utility of new diagnostic methods.
[0209] We used the NanoString assay to validate our diagnosis in pediatric sepsis patients from the GPSSSI cohort. NanoString is a highly accurate and useful tool for measuring the expression levels of multiple genes simultaneously, but it is likely too slow for clinical application (4–6 hours per assay). Thus, while the assay confirms that our gene set is robust for targeted measurement, further work will be required to improve test turnaround time. There are several possibilities for a final commercial product based on rapid, multiplexed qPCR. However, this technical challenge is one that all gene expression-based infection diagnostics must overcome to achieve clinical relevance.
[0210] Several research groups have published models for diagnosing infections based on host gene expression, but none have yet achieved clinical utility. Most previous classifiers either were not tested in multiple independent cohorts, or included too many genes to allow for the rapid profiling required for useful diagnosis, or both. For example, Suarez et al. created a 10-gene, k-nearest neighbor classifier but did not test this classifier outside of their published dataset (GSE60244). 13 Tsalik et al. created a classifier of 122 probes (120 genes) based on multiple regression models, but when they examined this classifier outside the GEO cohort, they retrained their regression coefficients on each new dataset. 14 Such model retraining would either introduce a strong upward bias to these validation rounds (assuming the final model is not locally retrained) or suggest that each new clinical site must recruit a large prospective cohort to train the model before implementation. Other research groups have created gene expression classifiers for sepsis but have not incorporated models to discriminate against viral infections. 7、9、10 Our novel IADM is robust across a wide range of disease types and severities, but has relatively little sensitivity to viral infection. Biomarkers other than gene expression have also been used to diagnose infection. Procalcitonin has been extensively studied in the context of sepsis diagnosis, but it cannot distinguish between uninfected and those with viral infection. 58 Protein panel assays have been shown to distinguish bacterial from viral infections, but they cannot distinguish patients with non-infectious inflammation. 59、60 Therefore, all of these classifiers have certain advantages and disadvantages that will become more apparent with further prospective testing and head-to-head comparisons.
[0211] While our goal in this study was to identify novel biomarkers, not necessarily novel biology, it is still important that the biomarker set has biological plausibility. Of the seven genes in the bacterial / viral metascore, six have already been linked to infection or leukocyte activation. While both IFI27 and JUP have been shown to be induced in response to viral infection in single-cohort genome-wide expression studies, 52,61 , TNIP1, and CTSB have been shown to be important in modulating the NF-kB-mediated and necroptotic responses to bacterial infection. 62,63 Finally, LAX1, which is upregulated during viral infection, is involved in T and B cell activation. 64 whereas HK3 plays an instrumental role in the neutrophil differentiation pathway. 65 Thus, the role of these transcripts as biomarkers for types of infection is novel, but not unprecedented.
[0212] Here, we rely on a new method, COCONUT, to directly compare our model across a large pool of single-class cohorts that would otherwise be unusable for benchmarking new diagnostic methods. COCONUT assumes that all controls come from the same distribution; that is, genes in each control group are reconfigured to have the same mean and variance with batch parameters empirically learned from the gene group. This method corrects for differences in microarray processing and batch processing between cohorts, allowing the creation of a global ROC curve with a single threshold. This is a more "realistic" measure of diagnostic power than simply reporting multiple validation ROC curves, since different cohorts may not achieve the same test characteristics with a single cutoff. 16The most important takeaway from the COCONUT co-normalized data is that both the bacterial / viral metascore and the IADM retain diagnostic power across a wide range of infection types and severities, with overall AUCs similar to the summary AUCs from head-to-head comparisons within the cohort.
[0213] Overall, we have derived a highly robust model for improving infection diagnosis using our demonstrated multi-cohort analysis pipeline. Using novel methods, we were able to validate this model in multiple independent microarray cohorts. We also validated the use of targeted NanoString assays in pediatric sepsis patients. While IADM still needs to be optimized for rapid testing times and prospective intervention trials, it seems clear that molecular profiling of the host genome will become part of the clinical toolkit of the future.
[0214] Those skilled in the art will understand that alternatives to the bacterial / viral metascore can be used to develop classifiers capable of distinguishing between bacterial and viral infections. Any machine learning method known in the art can be used to develop a classifier. Methods for developing classifiers can include ensemble algorithms created from multiple algorithms, such as logistic regression, support vector machines, and decision trees such as random forests and gradient-boosted trees. Classifications can be developed using neural networks containing multiple nodes arranged in layers, with the output from a node in a first layer used as input for a node in the next layer. Alternatively, classifications can be developed using support vector machine models, which represent examples as points in space, with examples of distinct types mapped to separate them by as large a clear gap as possible. New examples are then mapped into the same space, and the type assignment of the new example is predicted based on which side of the gap the new example lies on. Those skilled in the art will understand that any number of machine learning algorithms can be used to develop classifiers capable of distinguishing between bacterial and viral infections.
[0215] method Systematic search and multi-cohort analysis We performed a systematic search in NIH GEO and EBI ArrayExpress for published human microarray genome-wide expression studies using the search terms: bact[wildcard], vir[wildcard], infection, sepsis, SIRS, ICU, hospital-acquired, fever, and pneumonia. Abstracts were screened to exclude all studies: (1) nonclinical studies, (2) studies performed using tissues other than whole blood or PBMCs, or (3) studies comparing patients not matched for clinical time.
[0216] All microarray data were renormalized from raw data (when available) using standardized methods. Affymetrix arrays were renormalized using gcRMA (on arrays with perfect match probes) or RMA. Illumina, Agilent, GE, and other commercial arrays were renormalized via normal-exponential background correction followed by quartile normalization. Custom arrays were not renormalized. Data were log2 transformed, and probes for genes within each study were aggregated using a fixed-effects model. Within each study, cohorts assayed with different types of microarrays were treated as independent cohorts.
[0217] The inventors have already described 7、15、16、27 A multi-cohort meta-analysis was performed as described in
[10] . Briefly, genes were aggregated using Hedges' g dose, and a random effects model by DerSimonian-Leard was used for the meta-analysis, followed by Benjamini-Hochberg's multiple hypothesis correction. 66 Patients with bacterial infections within a study were compared to patients with viral infections, such that a positive effect size indicated that the gene was more highly expressed in patients with viral infections, and a negative effect size indicated that the gene was more highly expressed in patients with bacterial infections.
[0218] To find a set of genes highly conserved in differential expression between bacterial and viral infections, we selected all cohorts that directly compared patients with bacterial and viral infections. Patients with documented co-infections (i.e., both bacterial and viral infections) were excluded. Cohorts were required to have >5 patients in each group for inclusion in the meta-analysis. Both PBMC and whole blood cohorts were included. Significant genes were those with an effect size >2-fold and an FDR <1% in round-robin analyses using the LOO method. However, to ensure that both tissue types were represented in the final gene set, we also performed separate meta-analyses for the PBMC and whole blood cohorts and separately excluded all genes with an effect size <1.5-fold in either tissue type. The remaining genes were considered significant.
[0219] Derivation of a set of 7 genes To find a set of highly diagnostic genes, significant genes from meta-analysis were compared with previously described 7 , and subjected to a greedy forward search. Briefly, the algorithm starts with zero genes and adds one gene at a time in each cycle that most significantly improves the AUC for diagnosis in the discovery cohort until no new genes can improve the discovery AUC beyond some threshold. The resulting genes are used to calculate a single "bacteria / virus metascore," calculated as the geometric mean of the "virus" response genes minus the geometric mean of the "bacteria" response genes multiplied by the ratio of the number of genes in each set. The resulting continuous score can then be examined for diagnostic power using an ROC curve.
[0220] Derivation of further gene sets To identify additional diagnostic gene sets, we performed a recursive, greedy forward search, excluding the resulting diagnostic gene set from the set of possible significant genes at the end of the algorithm and running the algorithm again. The first gene set was taken for further validation, but other gene sets were noted to perform similarly in the discovery cohort (Supplementary Table 3).
[0221] Direct validation of the 7-gene set The resulting gene sets were first validated in the remaining published gene expression cohorts that directly compare bacterial to viral infections but are too small to be used in meta-analyses. After completing our meta-analysis, we compared the two cohorts (GSE60244 13 and GSE63990 14 ) was published and used for validation. To demonstrate generalizability, we also considered one large in vitro data set comparing exposure to LPS with exposure to influenza in monocyte-derived dendritic cells, but did not incorporate it into the summary AUC because it is not expected to come from the same distribution as the clinical study.
[0222] Summary ROC Curve For both the discovery and validation cohorts, the method by Kester and Van Chincks 67 and methods already described 16 Summary ROC curves were constructed according to the following: Briefly, for each ROC curve, a linear exponential model was created, and the parameters of these individual curves were aggregated using a random effects model to estimate the overall summary ROC curve parameters. The alpha parameter controls the AUC (specifically, the distance of the line from the line of direct proportion), and the beta parameter controls the skewness of the ROC curve. Confidence intervals for the summary AUC are estimated from the standard errors of alpha and beta in the meta-analysis.
[0223] COCONUT co-normalization There are numerous published microarray cohorts that profile patients with bacterial or viral infections, but not both. Comparing gene scores across these cohorts would be advantageous, but this has not yet been possible because different microarrays have widely varying background measurements for each gene, and large batch effects exist even among studies using the same type of microarray. To use these data, we needed to compare these cohorts in a way that (1) did not introduce bias that could affect the final classification (i.e., the normalization protocol was blinded to the diagnosis); (2) did not vary the distribution of genes within a study; and (3) after normalization, conormalized genes to ensure the same distribution across studies. A method with these features would allow our gene scores to be calculated and compared across multiple studies, thus broadly examining their generalizability.
[0224] Empirical Bayes Regularization with ComBat 32 While frequently used for cross-platform normalization, it assumes equal distribution across disease states and therefore falls critically short of our desired criteria. Therefore, we devised a modification of the ComBat method that co-normalizes control samples from different cohorts, allowing direct comparison of diseased samples from these same cohorts. We call this method COMBAT Co-Normalization Using control, or "COCONUT." COCONUT makes one explicit assumption: it forces control / healthy patients from different cohorts to represent the same distribution. Briefly, all cohorts are divided into healthy and diseased components. The healthy component is subjected to ComBat co-normalization without covariates. For the healthy component of each dataset, the ComBat estimated parameters [ka] is obtained and then applied to the diseased component (Figure 10). This forces the diseased components of all cohorts to come from the same background distribution but preserve their relative distance from the healthy component (due to floating-point arithmetic, only the T-statistics within the dataset are different after COCONUT). COCONUT also does not require a priori knowledge of the disease classification (i.e., bacterial or viral infection), so it is important that our pre-specified criteria are met. The COCONUT method has an important requirement: healthy / control patients are present in the dataset for pooling with other available data. The COCONUT method also places healthy / control patients in the same distribution, so it is used only when such assumptions are valid (i.e., within the same tissue type, between the same species, etc.).
[0225] ComBat model and COCONUT method As described by Johnson et al., the ComBat model first solves a least-squares model for gene expression, then iteratively uses the solved empirical Bayes estimator to correct for the location and scale of each gene by reducing the resulting parameters. 32 Formally, the expression level of each gene, Y ijg (of gene g for sample j in batch i) as the global gene expression α g , the regression coefficient is β g The design matrix for sample condition X, the additive batch effect, and the synergistic batch effect, γ ig and δ ig , and the error term ε ijg It consists of: Y ijg =α g +Xβ g +γ ig +δ ig ε ijg Let's say.
[0226] By estimating parameters using least squares regression, Y ijg into a new term Zijg (In the formula, [ka] is ε ijg (where is the standard deviation of ):
number
[0227] Thus, the standardized data [ka] [In the formula, [ka] ] It is distributed according to
[0228] The inverse of gamma is taken as the standard uninformative prior. The remaining hyperparameters are estimated empirically, with derivations and solutions found in the original references. 32 Then, the estimated batch effect is [ka] The standardized data is then batch adjusted by empirical Bayes using the [ka] :
number
[0229] In our modified version of this method (COCONUT), all of the above is performed according to the original method without any modifications. However, the COCONUT modified method is applied only to healthy / control patients in each dataset (i.e., Y is a matrix for only healthy patient samples). Estimation parameters [ka] and calculate the matrix D consisting of only affected patient samples (which must be ordered in the same way as Y):
number
number
[0230] Thus, the inventors have identified the affected sample D * , which corrects for differences between healthy controls, but i For each sub-matrix D i A batch-corrected form can be obtained that does not change
[0231] Global ROC We used COCONUT co-normalization to examine (1) all discovery cohorts and (2) all validation cohorts, even those containing only bacterial or only viral diseases. We did this separately for PBMC and whole blood data for the reasons described above. After co-normalization, distributions for the individual cohorts were plotted together to allow direct comparison. For each plot, we show (1) the distribution of scores for each dataset, (2) the normalized gene expression level for each gene in the diagnostic test, and (3) housekeeping genes that are predicted to show no differences between classes based on meta-analysis. Healthy patients were excluded from these plots. However, to demonstrate that the distribution of genes does not change between healthy and diseased patients within the cohort after COCONUT co-normalization, we also show plots with both patient types, including both target and housekeeping genes (Figure 11). Genes with the smallest effect size and smallest variance in the meta-analysis were selected as housekeeping genes.
[0232] For each comparison, a single global ROC AUC was calculated and a single threshold was set to allow estimation of the actual test diagnostic performance. Because false negative results for bacterial infections (i.e., recommendations not to administer antibiotics when they are needed) can be significant, the threshold for the cutoff for bacterial versus viral infections was set to approximate a bacterial sensitivity of 90%.
[0233] Integrated antibiotic decision-making model Although the SMS can distinguish patients with severe acute infections from patients with inflammation from other sources, it cannot distinguish between types of infection (Figures 5A and 5B). Therefore, we investigated an integrated antibiotic decision model (IADM), which applies an 11-gene SMS followed by a 7-gene bacterial / viral metascore. Thus, the IADM model (1) identifies whether a patient has an infection and (2) if so, what type of infection (bacterial or viral) is present. Because we were unable to identify a sufficient validation cohort with patients with noninfectious inflammation that also included healthy controls, we used both the discovery and validation cohorts in constructing the global ROC. A global threshold was established across all included cohorts using COCONUT co-normalization, and the global threshold was applied to each individual dataset to examine the ability of IADM to properly distinguish between patients with non-infectious inflammation, bacterial infection, and viral infection. Healthy patients were not included as a diagnostic class because they were used in the co-normalization procedure. IADM was also applied separately to all cohorts that did not have healthy controls but included both (1) non-infectious SIRS patients and (2) patients with both bacterial and viral infections.
[0234] Positive predictive value and negative predictive value (PPV and NPV) depend on prevalence. However, because the prevalence data used here do not correspond to the prevalence of infections in hospital settings, we calculated PPV and NPV curves based on the sensitivity and specificity for bacterial infection achieved by the integrated antibiotic decision model. Formally, NPV = specificity × (1 − prevalence) / ((1 − sensitivity) × prevalence + specificity × (1 − prevalence)); PPV = sensitivity × prevalence / (sensitivity × prevalence + (1 − specificity) × (1 − prevalence)).
[0235] Validation with NanoString Finally, Targeting NanoString 56Genomics of Pediatric SIRS and Septic Shock Investigators Study Using Digital Multiplex Gene Quantification Assays 18~22 Ninety-six samples from individual patients (i.e., patients not profiled via microarray) were examined. The 18 genes were not renormalized to any housekeeping genes. Both SMS and bacterial / viral metascore genes were assayed, and the diagnostic performance of IADM was calculated.
[0236] All analyses were performed in the R statistical computing language (version 3.1.1). Code for rerunning the multi-cohort meta-analysis has been deposited and is available at khatrilab.stanford.edu / sepsis.
[0237] [Table 1] TIFF0007750895000013.tif57149
[0238] [Table 2] TIFF0007750895000015.tif117149
[0239] [Table 3] TIFF0007750895000017.tif129149 TIFF0007750895000018.tif85149
[0240] [Table 4]
[0241] [Table 5]
[0242]
Table 6
[0243] References 1. Ferrer, R., et al. Empiric antibiotictreatment reduces mortality in severe sepsis and septic shock from the firsthour: results from a guideline-based performance improvement program*. CritCare Med 42, 1749-1755 (2014). 2. Fridkin, S., et al. Vital signs:improving antibiotic use among hospitalized patients. MMWR Morb Mortal Wkly Rep63, 194-200 (2014). 3. Grijalva, C.G., Nuorti, J.P. & Griffin,M.R. Antibiotic prescription rates for acute respiratory tract infections in USambulatory settings. JAMA 302, 758-766 (2009). 4. Andrews, J. High rates of enteric feverdiagnosis and low burden of disease in rural Nepal. in 9th International Conferenceon Typhoid and Invasive NTS Disease (Bali, Indonesia, 2015). 5. National strategy and action plan forcombating antibiotic resistant bacteria, (New York : Nova Publishers, 2015). 6. Liesenfeld, O., Lehman, L., Hunfeld,K.P. & Kost, G. Molecular diagnosis of sepsis: New aspects and recentdevelopments. Eur J Microbiol Immunol (Bp) 4, 1-25 (2014). 7. Sweeney, T.E., Shidham, A., Wong, H.R.& Khatri, P. A comprehensive time-course-based multicohortanalysis of sepsis and sterile inflammation reveals a robust diagnostic geneset. Sci Transl Med 7, 287ra271 (2015). 8. Sweeney, T.E. & Khatri, P.Comprehensive Validation of the FAIM3:PLAC8 Ratio in Time-matched PublicGene Expression Data. Am J Respir Crit Care Med 192, 1260-1261 (2015). 9. McHugh, L., et al. A Molecular HostResponse Assay to Discriminate Between Sepsis and Infection-Negative SystemicInflammation in Critically Ill Patients: Discovery and Validation inIndependent Cohorts. PLoS Med 12, e1001916 (2015). 10. Scicluna, B.P., et al. A MolecularBiomarker to Diagnose Community-acquired Pneumonia on Intensive Care UnitAdmission. Am J Respir Crit Care Med (2015). 11. Hu, X., Yu, J., Crosby, S.D. &Storch, G.A. Gene expression profiles in febrile children with defined viraland bacterial infection. Proc Natl Acad Sci U S A 110, 12792-12797 (2013). 12. Zaas, A.K., et al. A host-based rt-PCRgene expression signature to identify acute respiratory viral infection. SciTransl Med 5, 203ra126 (2013). 13. Suarez, N.M., et al. Superiority ofTranscriptional Profiling Over Procalcitonin for Distinguishing Bacterial FromViral Lower Respiratory Tract Infections in Hospitalized Adults. J Infect Dis(2015). 14. Tsalik, E.L., et al. Host geneexpression classifiers diagnose acute respiratory illness etiology. Sci TranslMed 8, 322ra311 (2016). 15. Andres-Terre, M., et al. Integrated,Multi-cohort Analysis Identifies Conserved Transcriptional Signatures acrossMultiple Respiratory Viruses. Immunity 43, 1199-1211 (2015). 16. Sweeney, T.E., Braviak, L., Tato, C.M.& Khatri, P. Multi-Cohort Analysis of Genome-Wide Expression for Diagnosisof Pulmonary Tuberculosis. Lancet Resp Med (Accepted,Jan 2016). 17. Sweeney, T.E. & Khatri, P.Benchmarking sepsis gene expression diagnostics using public data. UnderReview, Jan 2016. 18. Shanley, T.P., et al. Genome-levellongitudinal expression of signaling pathways and gene networks in pediatricseptic shock. Mol Med 13, 495-508 (2007). 19. Wong, H.R., et al. Genome-levelexpression profiles in pediatric septic shock indicate a role for altered zinchomeostasis in poor outcome. Physiol Genomics 30, 146-155 (2007). 20. Cvijanovich, N., et al. Validating thegenomic signature of pediatric septic shock. Physiol Genomics 34, 127-134(2008). 21. Wong, H.R., et al. Genomic expressionprofiling across the pediatric systemic inflammatory response syndrome, sepsis,and septic shock spectrum. Crit Care Med 37, 1558-1566 (2009). 22. Wong, H.R., et al. Interleukin-27 is anovel candidate diagnostic biomarker for bacterial infection in critically illchildren. Crit Care 16, R213 (2012). 23. Ramilo, O., et al. Gene expressionpatterns in blood leukocytes discriminate patients with acute infections. Blood109, 2066-2077 (2007). 24. Parnell, G., et al. Aberrant cell cycleand apoptotic changes characterise severe influenza A infection--ameta-analysis of genomic signatures in circulating leukocytes. PLoS One 6,e17186 (2011). 25. Parnell, G.P., et al. A distinctinfluenza infection signature in the blood transcriptome of patients withsevere community-acquired pneumonia. Crit Care 16, R157 (2012). 26. Herberg, J.A., et al. Transcriptomicprofiling in childhood H1N1 / 09 influenza reveals reduced expression of proteinsynthesis genes. J Infect Dis 208, 1664-1668 (2013). 27. Khatri, P., et al. A common rejectionmodule (CRM) for acute rejection across multiple organs identifies noveltherapeutics for organ transplantation. J Exp Med 210, 2205-2221(2013). 28. Popper, S.J., et al. Gene transcriptabundance profiles distinguish Kawasaki disease from adenovirus infection. JInfect Dis 200, 657-666 (2009). 29. Smith, C.L., et al. Identification of ahuman neonatal immune-metabolic network associated with bacterial infection.Nat Commun 5, 4649 (2014). 30. Almansa, R., et al. Critical COPDrespiratory illness is linked to increased transcriptomic activity ofneutrophil proteases genes. BMC Res Notes 5, 401 (2012). 31. Lee, M.N., et al. Common geneticvariants modulate pathogen-sensing responses in human dendritic cells. Science343, 1246980 (2014). 32. Johnson, W.E., Li, C. & Rabinovic,A. Adjusting batch effects in microarray expression data using empirical Bayesmethods. Biostatistics 8, 118-127 (2007). 33. Zhai, Y., et al. Host TranscriptionalResponse to Influenza and Other Acute Respiratory Viral Infections--AProspective Cohort Study. PLoS Pathog 11, e1004869 (2015). 34. Kwissa, M., et al. Dengue virusinfection induces expansion of a CD14(+)CD16(+) monocyte population thatstimulates plasmablast differentiation. Cell Host Microbe 16, 115-127 (2014). 35. Mejias, A., et al. Whole blood geneexpression profiles to assess pathogenesis and disease severity in infants withrespiratory syncytial virus infection. PLoS Med 10, e1001549 (2013). 36. Berdal, J.E., et al. Excessive innateimmune response and mutant D222G / N in severe A (H1N1) pandemic influenza. JInfect 63, 308-316 (2011). 37. Bermejo-Martin, J.F., et al. Hostadaptive immunity deficiency in severe pandemic influenza. Crit Care 14, R167(2010). 38. Zaas, A.K., et al. Gene expressionsignatures diagnose influenza and other symptomatic respiratory viralinfections in humans. Cell Host Microbe 6, 207-217 (2009). 39. van de Weg, C.A., et al. Time sinceonset of disease and individual clinical markers associate with transcriptionalchanges in uncomplicated dengue. PLoS Negl Trop Dis 9, e0003522 (2015). 40. Conejero, L., et al. The BloodTranscriptome of Experimental Melioidosis Reflects Disease Severity and ShowsConsiderable Similarity with the Human Disease. J Immunol 195, 3248-3261(2015). 41. Cazalis, M.A., et al. Early and dynamicchanges in gene expression in septic shock patients: a genome-wide approach.Intensive Care Med Exp 2, 20 (2014). 42. Lill, M., et al. Peripheral blood RNAgene expression profiling in patients with bacterial meningitis. Front Neurosci7, 33 (2013). 43. Ahn, S.H., et al. Gene expression-basedclassifiers identify Staphylococcus aureus infection in mice andhumans. PLoS One 8, e48979 (2013). 44. Thuny, F., et al. The gene expressionanalysis of blood reveals S100A11 and AQP9 as potential biomarkers of infectiveendocarditis. PLoS One 7, e31490 (2012). 45. Sutherland, A., et al. Development andvalidation of a novel molecular biomarker diagnostic test for the earlydetection of sepsis. Crit Care 15, R149 (2011). 46. Berry, M.P., et al. Aninterferon-inducible neutrophil-driven blood transcriptional signature in humantuberculosis. Nature 466, 973-977 (2010). 47. Pankla, R., et al. Genomictranscriptional profiling identifies a candidate blood biomarker signature forthe diagnosis of septicemic melioidosis. Genome Biol 10, R127 (2009). 48. Irwin, A.D., et al. Novel biomarkercombination improves the diagnosis of serious bacterial infections in Malawianchildren. BMC Med Genomics 5, 13 (2012). 49. Bloom, C.I., et al. Transcriptionalblood signatures distinguish pulmonary tuberculosis, pulmonary sarcoidosis,pneumonias and lung cancers. PLoS One 8, e70630 (2013). 50. Ardura, M.I., et al. Enhanced monocyteresponse and decreased central memory T cells in children with invasiveStaphylococcus aureus infections. PLoS One 4, e5446 (2009). 51. Liu, K., Chen, L., Kaur, R. &Pichichero, M. Transcriptome signature in young children with acute otitismedia due to Streptococcus pneumoniae. Microbes Infect 14, 600-609 (2012). 52. Ioannidis, I., et al. Plasticity andvirus specificity of the airway epithelial cell immune response duringrespiratory virus infection. J Virol 86, 5422-5436 (2012). 53. Popper, S.J., et al. Temporal dynamicsof the transcriptional response to dengue virus infection in Nicaraguanchildren. PLoS Negl Trop Dis 6, e1966 (2012). 54. Brand, H.K., et al. Olfactomedin 4Serves as a Marker for Disease Severity in Pediatric Respiratory SyncytialVirus (RSV) Infection. PLoS One 10, e0131927 (2015). 55. Wacker, C., Prkno, A., Brunkhorst, F.M.& Schlattmann, P. Procalcitonin as a diagnostic marker forsepsis: a systematic review and meta-analysis. Lancet Infect Dis 13, 426-435(2013). 56. Kulkarni, M.M. Digital multiplexed geneexpression analysis using the NanoString nCounter system. Curr Protoc Mol BiolChapter 25, Unit25B.10 (2011). 57. Tsalik, E.L., et al. PotentialCost-effectiveness of Early Identification of Hospital-acquired Infection inCritically Ill Patients. Ann Am Thorac Soc (2015). 58. Gilbert, D.N. Procalcitonin as abiomarker in respiratory tract infection. Clin Infect Dis 52 Suppl 4, S346-350(2011). 59. Oved, K., et al. A novel host-proteomesignature for distinguishing between acute bacterial and viral infections. PLoSOne 10, e0120012 (2015). 60. Valim, C., et al. Responses toBacteria, Virus, and Malaria Distinguish the Etiology of Pediatric ClinicalPneumonia. Am J Respir Crit Care Med 193, 448-459 (2016). 61. Tolfvenstam, T., et al.Characterization of early host responses in adults with dengue disease. BMCInfect Dis 11, 209 (2011). 62. Vanden Berghe, T., Linkermann, A.,Jouan-Lanhouet, S., Walczak, H. & Vandenabeele, P. Regulated necrosis: theexpanding network of non-apoptotic cell death pathways. Nat Rev Mol Cell Biol15, 135-147 (2014). 63. Ashida, H., et al. A bacterial E3ubiquitin ligase IpaH9.8 targets NEMO / IKKgamma to dampen the hostNF-kappaB-mediated inflammatory response. Nat Cell Biol 12, 66-73; sup pp 61-69(2010). 64. Zhu, M., et al. Negative regulation oflymphocyte activation by the adaptor protein LAX. J Immunol 174, 5612-5619(2005). 65. Federzoni, E.A., et al. PU.1 is linkingthe glycolytic enzyme HK3 in neutrophil differentiation and survival of APLcells. Blood 119, 4963-4970 (2012). 66. Benjamini, Y. & Hochberg, Y.Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society, Series B 57, 289-300(1995). 67. Kester, AD & Buntinx, F. Meta-analysis of ROC curves. Med Decis Making 20, 430-439 (2000).
[0244] While the preferred embodiment of the invention has been illustrated and described, it will be appreciated that various changes can be made therein without departing from the spirit and scope of the invention.
Claims
1. 1. A method for aiding in the diagnosis of an infection in a patient, comprising: a) measuring in vitro the expression level of biomarker polynucleotides in a patient biological sample, wherein said biomarker polynucleotides comprise transcripts of at least one pair of adjacent gene pairs 1 and 2 listed in Table 1 below; b) analyzing the expression level of each biomarker polynucleotide, along with a respective reference range for each biomarker polynucleotide, to determine viral or bacterial infection; a method comprising: 【Table 1】 。
2. The at least one gene pair is selected from the group consisting of SIGLEC1 and SLC12A9, IFI27 and HK3, IFI27 and S100A12, SIGLEC1 and IMPA2, SIGLEC1 and TBXAS1, IFI27 and DYSF, IFI27 and TNIP1, SIGLEC1 and ACAA1, SIGLEC1 and DYSF, IFI27 and TSPO, OAS2 and SLC12A9, IFI27 and EMR1, SIGLEC1 and HK3, IFI27 and SLC12A9, IFI27 and SORT1, OAS3 and HK3, SIGLEC1 and ST 2. The method of claim 1, wherein the target gene is selected from the group consisting of AT5B, IFIT1 and HK3, SIGLEC1 and EMR1, IFI27 and PGD, CUL1 and IFI27, IFI27 and JUP, IFI27 and ACAA1, IFI27 and GPAA1, IFI27 and NRD1, IFI27 and STAT5B, IFIT1 and DYSF, OAS1 and HK3, OAS1 and SLC12A9, OAS2 and PTAFR, OAS3 and SLC12A9, SIGLEC1 and FLII, SIGLEC1 and TSPO, and CHST12 and IFI27.
3. 3. The method of claim 1 or 2, wherein the expression levels of the two biomarker polynucleotides result in an area under the receiver operating characteristic curve of at least 0.
80.
4. 4. The method of claim 3, wherein the expression levels of the two biomarker polynucleotides result in an area under the receiver operating characteristic curve of at least 0.
84.
5. The method of any one of claims 1 to 4, wherein the biological sample comprises whole blood or peripheral blood mononuclear cells (PBMCs).
6. 6. The method of any one of claims 1 to 5, wherein the levels of the biomarker polynucleotides are compared to time-matched reference values for infected or non-infected subjects, wherein the time-matched reference values are matched for clinical time.
7. The method of any one of claims 1 to 6, wherein the biomarker polynucleotide is mRNA or a polynucleotide derived therefrom.
8. 8. The method of any one of claims 1 to 7, wherein measuring the level of the biomarker polynucleotide comprises performing one or more methods comprising microarray analysis via detection of fluorescent, chemiluminescent, or electrical signals, polymerase chain reaction (PCR), reverse transcriptase polymerase chain reaction (RT-PCR), digital droplet PCR (ddPCR), solid-state nanovesicle detection, isothermal amplification, RNA switch activation, Northern blot, or serial analysis of gene expression (SAGE).
9. 1. A kit for detecting whether a patient has a viral or bacterial infection, comprising: an agent for measuring the level of a biomarker polynucleotide in a biological sample of a patient; and Instructions for correlating the detection level of each biomarker with viral or bacterial infection Including, a kit, wherein the biomarker polynucleotides comprise transcripts of at least one pair of adjacent genes 1 and 2 listed in Table 2 below; 【Table 2】 。
10. The at least one gene pair is selected from the group consisting of SIGLEC1 and SLC12A9, IFI27 and HK3, IFI27 and S100A12, SIGLEC1 and IMPA2, SIGLEC1 and TBXAS1, IFI27 and DYSF, IFI27 and TNIP1, SIGLEC1 and ACAA1, SIGLEC1 and DYSF, IFI27 and TSPO, OAS2 and SLC12A9, IFI27 and EMR1, SIGLEC1 and HK3, IFI27 and SLC12A9, IFI27 and SORT1, OAS3 and HK3, SIGLEC1 and ST The kit of claim 9, wherein the target protein is selected from the group consisting of AT5B, IFIT1 and HK3, SIGLEC1 and EMR1, IFI27 and PGD, CUL1 and IFI27, IFI27 and JUP, IFI27 and ACAA1, IFI27 and GPAA1, IFI27 and NRD1, IFI27 and STAT5B, IFIT1 and DYSF, OAS1 and HK3, OAS1 and SLC12A9, OAS2 and PTAFR, OAS3 and SLC12A9, SIGLEC1 and FLII, SIGLEC1 and TSPO, and CHST12 and IFI27.
11. The kit of claim 9 or 10, further comprising an agent for measuring the levels of the biomarker polynucleotides CEACAM1, ZDHHC19, C9orf95, GNA15, BATF, C3AR1, KIAA1370, TGFBI, MTCH1, RPGRIP1, and HLA-DPB1.
Citation Information
Patent Citations
Characteristics and determinants for diagnosing infection, and methods of using them.
JP2015510122A
Signatures and determinants for distinguishing between a bacterial and viral infection and methods of use thereof
US20110275542A1