Methods, kits and systems for risk prediction or detection of sepsis and / or infection

Specific nucleic acid biomarker signatures in biological samples, combined with mathematical methods, enable early and accurate detection of sepsis and infection, addressing the limitations of current diagnostic methods and reducing antimicrobial resistance.

GB2700474APending Publication Date: 2026-02-11THE SEC OF STATE FOR DEFENCE IN HER BRITANNIC MAJESTYS GOVERNMENT OF THE UK OF GREAT BRITAIN & NORTHERN IRELAND
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
GB2025003978
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-20
Filing Date
2025-03-19
Publication Date
2026-02-11

AI Technical Summary

Technical Problem

Current diagnostic methods for sepsis and infection lack the ability to accurately predict or detect these conditions before symptoms arise, particularly in vulnerable populations, and existing biomarkers fail to provide high predictive accuracy, leading to inappropriate use of antimicrobials and antimicrobial resistance.

Method used

Utilization of specific nucleic acid biomarker signatures, such as DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, SNORA53, BEND2, CR1, EIF4G3, LARP1, LGALS2, MIAT, QSOX1, RPL13A, RRBP1, SGSH, SLC36A1, TDRD9, TEX2, and TRIP13, to analyze expression levels in biological samples, supported by mathematical methods and neural networks, for early detection and differentiation of infection and sepsis.

Benefits of technology

Achieves high predictive accuracy (AUC >0.9) in distinguishing between uninfected individuals, patients with infection, and those at risk of sepsis, enabling timely treatment and reducing antimicrobial misuse.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Method for risk prediction or detection / diagnosis, or absence, of sepsis and / or infection in a subject comprising the steps of: a) providing at least one biological sample from said subject; b) detect
Need to check novelty before this filing date? Find Prior Art

Description

The present invention is concerned with methods, computer implemented methods, systems and kits for the risk prediction or detection / diagnosis of sepsis and / or infection in a subject through measuring the presence or level of expression of specific nucleic acid markers in a biomarker signature. Following exposure to a microbial pathogen there is often a lag phase before symptoms of infection, which could further result in symptoms of organ dysfunction, and development of sepsis. After the onset of clinical symptoms, the effectiveness of treatment often decreases as the disease progresses, so the time taken to make any diagnosis is critical. It is likely that a detection or diagnostic assay will be the first confirmed indicator of infection or sepsis. The availability, rapidity and predictive accuracy of such an assay, especially at any point on the trajectory of disease onset, will therefore be crucial in determining the outcome. Any time saved will speed up the implementation of medical countermeasures and will have a significant impact on recovery. The development of technologies to facilitate rapid detection of infection and sepsis is a key concern for all at risk, which includes the ability to differentiate between a systemic inflammatory response to infection, from a systemic inflammatory response to an alternative stressor such as trauma, surgery, acute inflammation, ischemia or malignancy, cellulitis, autoimmune disease, vasculitis or infarction. During the initial stages of infection many biological agents are either absent from, or present at very low concentrations in, typical clinical samples (e.g. blood). It is therefore likely that agent-specific assays would have limited utility in detecting infection before clinical symptoms arise. Previous studies have shown that infection elicits a pattern of immune response involving changes in the expression of a variety of biomarkers that is indicative of the type of agent. Such patterns of biomarker expression have proven to be diagnostic for a variety of infectious agents. It is now possible to distinguish patterns of gene expression in blood leukocytes from symptomatic patients with acute infections caused by four common human pathogens (Influenza A, Staphylococcus aureus, Streptococcus pneumoniae and Escherichia coli) using whole transcriptome analysis. More recently, researchers have been able to reduce the number of host biomarkers required to make a diagnosis through use of appropriate bioinformatic analysis techniques to select key biomarkers for the diagnosis of infectious disease. While host biomarker signatures represent an attractive solution for the prediction of microbial infection, their discovery relies on the exploitation of laboratory models of infection whose fidelity to the pathogenesis of disease in humans varies. An alternative approach for biomarker discovery in humans is to exploit a common sequela of biological agent infection; such as the life-threatening condition sepsis, which now requires organ dysfunction for a positive diagnosis. Sepsis has traditionally been defined as a systemic inflammatory response syndrome (SIRS) in response to infection which, when associated with acute organ dysfunction, may ultimately cause severe lifethreatening complications. However, sepsis is now defined as life-threatening organ dysfunction caused by a dysregulated host response to infection, wherein organ dysfunction can be identified as an increase in the total sequential organ failure assessment (SOFA; originally the Sepsis-related Organ Failure Assessment) score of 2 or more points from one day to the next following infection. A higher SOFA score is associated with an increased probability of mortality. The score grades abnormality by organ system and accounts for clinical interventions. The baseline SOFA score can be assumed to be zero in patients not known to have pre-existing organ dysfunction. A SOFA score >2 reflects an overall mortality risk of approximately 10% in a general hospital population with suspected infection. Please note, in the UK, the National Early Warning Score 2 (NEWS2) scoring system is now more often used to detect and characterise acute illness severity in patients. Sepsis is a major cause of morbidity and mortality in intensive care units (ICU). In the UK, sepsis is believed to be responsible for about 27% of all ICU admissions. Across Europe the average incidence of sepsis in the ICU is about 30%, with a mortality rate of 27%. In the USA, hospital-associated mortality from sepsis ranges between 18 to 30%; an estimated 9.3% of all deaths occurred in patients with sepsis. Clearly there is a very accessible patient population that could be used to study predictive markers for the onset of sepsis. Despite greatly improved diagnosis, treatment and support, serious infection and sepsis remain significant causes of death and often result in chronic ill-health or disability in those who survive acute episodes. Although sudden, overwhelming infection is comparatively rare amongst otherwise healthy adults, it constitutes an increased risk in immunocompromised individuals, seriously ill patients in intensive care, burns patients and young children. In a proportion of cases, an apparently treatable infection leads to the development of sepsis. The ability to detect potentially serious infections and sepsis as early as possible and, especially, to predict the onset of sepsis in susceptible individuals is clearly advantageous. Alternatively, the ability to identify or predict individuals (such as patients) who are unlikely to develop an infection, or indeed sepsis, is also of great value, thus a method for ruling out infection or sepsis is also desired. Indeed a key challenge in this field is the need to identify patients who are unlikely to develop an infection (thus rule out methodology), in order for the use of antimicrobial agents (such an antibiotics) to be much better managed and controlled. Antimicrobial resistance (AMR) is one of the top global public health threats, which has at least in part come about from the misuse and overuse of antimicrobial agents, with them being prescribed as a precaution, just in case a patient has an infection. This behaviour has been a key contributor in the growth of drug resistant pathogens, and in infections becoming a lot more difficult to treat. Priorities to address AMR in human health include avoidance of inappropriate use of antimicrobial agents, and ensuring the availability of appropriate tests to promptly rule out the presence of an infection, and for such testing to be universally available, to inform appropriate treatment. Furthermore, kits and methods for indicating the severity of infection are also desired, in which the severity may include development of sepsis, or the presence of an antimicrobial resistant pathogen, where the result of a test could inform an appropriate treatment regime. Although a number of biomarkers (markers), such as nucleic acid markers or protein markers, have been shown to correlate with infection and sepsis and some give an indication of the seriousness of the condition, no single marker or combination of markers has yet been shown to be a reliable diagnostic test. Further, many of the markers currently used in tests have in fact been discovered in samples from seriously ill patients, who are often already septic, whereas it would be more beneficial to consider patients, and expression of biomarkers, prior to formal clinical diagnosis or the development of recognisable symptoms (pre-symptomatic), by at least a few days, in order to consider the prediction, rather than simply diagnosis, of whether infection and / or sepsis is likely to occur, or alternatively is unlikely (rule-out methodology), and thereby identification of a suitable treatment regime. Moreover, clinicians also struggle to confirm the presence of infection, a key component of the current sepsis diagnostic criteria, and so a rigorous protocol, using multiple clinicians, is advisable in any study. Extracting reliable diagnostic patterns and robust prognostic indications from changes over time in complex sets of variables including traditional clinical observations, clinical chemistry, biochemical, immunological and cytometric data requires sophisticated methods of analysis. The use of expert systems and artificial intelligence, including neural networks, for medical diagnostic applications has been being developed for some time. The ability to detect the earliest signs of infection and / or sepsis has clear benefits in terms of allowing treatment as soon as possible. Indications of the severity of the condition and likely outcome if untreated inform decisions about treatment options. This is relevant both in vulnerable hospital populations, such as those in intensive care or in the post-operative recovery period, or who are burned or immunocompromised, and in other groups in which there is an increased risk of serious infection and subsequent sepsis. It is also relevant in patients presenting to Emergency Departments, for which clinicians must decide whether to admit or discharge. The use or suspected use of biological weapons in both battlefield and civilian settings is an example where a rapid and reliable means of testing for the earliest signs of infection or organ dysfunction (i.e. sepsis) in individuals exposed would also be advantageous. In addition, rapid prognostication of patients with overt symptoms of infection who may deteriorate and develop sepsis is also a key concern for those managing patient treatment. However, until now the majority of investigations focused on developing a group of biomarkers and / or a test for sepsis were based on the previous definition of sepsis of SIRS in response to infection, and / or in samples from patients who were already seriously ill, and thus have generally focussed on identifying a group of biomarkers and / or a test to predict SIRS in response to infection. Sepsis is, however, now defined as life-threatening organ dysfunction caused by a dysregulated host response to infection, and thus groups of biomarkers and / or tests are required which are capable of confirming the presence of infection and identifying individuals at risk of developing life-threatening organ dysfunction caused by a dysregulated host response to infection. Until now neither a test nor a list of biomarkers has been identified / produced which can detect or predict sepsis at any stage of the trajectory of disease with a good / high predictive accuracy (for example with an area under the curve (AUC) >0.90, and more preferably >0.95). It is notable that clinicians often disagree on whether patients have sepsis based on clinical criteria. Further, whilst several discrete molecular tests are now available, confirmation of the presence of bacterial infection is particularly challenging when the standard of care, microbial culture, fails to yield a result in up to 50% of cases. The present invention thus aims to provide biomarker signatures (groups of biomarkers), and methods for classifying biological samples using the biomarker signatures, to confirm the presence or absence (rule-out) of infection, predict / detect the development of sepsis, and preferably both infection and sepsis, with a high predictive accuracy. The present invention further aims to provide methods for differentiating uninfected patients (controls) from patients with sepsis, and for differentiating patients presenting with SIRS (with no infection) from patients having an infection. Thus in a first aspect the present invention provides a method for the risk prediction or detection / diagnosis, or absence, of sepsis and / or infection in a subject, comprising the steps of: a) providing at least one biological sample from said subject; b) detecting the presence or level of expression of at least two nucleic acid markers selected from a biomarker signature, in said sample, compared to a control (optionally the expression level of at least one control nucleic acid or housekeeping nucleic acid), and / or a baseline level of expression, and / or expression levels with respect to other nucleic acid markers in said group, and / or through use of a mathematical method, algorithm or neural network; and c) concluding, from said presence or level of expression of said at least two nucleic acid markers, the risk of the development, or detection / diagnosis, or absence, of sepsis and / or infection, and classifying the subject accordingly; wherein the biomarker signature is either a first group consisting of DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, and SNORA53, or a second group consisting of BEND2, CR1, EIF4G3, LARP1, LGALS2, MIAT, QSOX1, RPL13A, RRBP1, SGSH, SLC36A1, TDRD9, TEX2, and TRIP13, or a third group consisting of a combination of the first group and the second group. Thus, in one embodiment, is a method for the risk prediction or detection / diagnosis, or absence, of sepsis and / or infection in a subject, comprising the steps of: a) providing at least one biological sample from said subject; b) detecting the presence or level of expression of at least two nucleic acid markers selected from a biomarker signature consisting of DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, and SNORA53, in said sample, compared to a control, or a baseline level of expression, or expression levels with respect to other nucleic acid markers in said biomarker signature; and c) concluding, from said presence or level of expression of said at least two nucleic acid markers, the risk of the development, or detection / diagnosis, or absence of sepsis and / or infection, and classifying the subject accordingly. The method may comprise detecting the presence or level of expression of 2, 3, 4, 5, 6 or 7 of the nucleic acid markers in the group consisting of DDX11L9 (DEAD / H-Box Helicase 11 Like 9 (Pseudogene)), TIFA (TRAF Interacting Protein With Forkhead Associated Domain), B4GALT5 (Beta-1,4-Galactosyltransferase 5), PTGES3 (Prostaglandin E Synthase 3), TPM3 (Tropomyosin 3), CDC42 (Cell Division Cycle 42), and SNORA53 (Small Nucleolar RNA, H / ACA Box 53). In one embodiment, the method comprises detecting the presence or level of expression of at least the markers B4GALT5 and SNORA53. In another embodiment, the method comprises detecting the presence or level of expression of all of the nucleic acid markers in the group consisting of DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, and SNORA53. The Applicant has shown that the combination of all seven nucleic acid markers, using a method that considers the expression of each of the markers, achieves a high performance. However, embodiments with subsets of these seven markers also achieve high performance, and especially with the four gene signature B4GALT5, PTGES3, TPM3, and SNORA53, with the highest performance, although surprisingly smaller subsets such as two markers also provides for high differentiation, with AUCs in excess of 0.9. Single markers also provide for some differentiation, especially as concerns negative predictive value, and so ability to rule out infection and / or sepsis. The expression of the markers may be increased or decreased with reference to a control, a baseline level of expression, and / or indeed with respect to each other, but it is the combined analysis (and overall expression) of markers (be that all seven or indeed subsets) that provides for the high accuracy of the method. In another embodiment, is a method for the risk prediction or detection / diagnosis, or absence, of sepsis and / or infection in a subject, comprising the steps of: a) providing at least one biological sample from said subject; b) detecting the presence or level of expression of at least two nucleic acid markers selected from a biomarker signature consisting of BEND2, CR1, EIF4G3, LARP1, LGALS2, MIAT, QSOX1, RPL13A, RRBP1, SGSH, SLC36A1, TDRD9, TEX2, and TRIP13, in said sample, compared to a control, or a baseline level of expression, or expression levels with respect to other nucleic acid markers in said biomarker signature; and c) concluding, from said presence or level of expression of said at least two nucleic acid markers, the risk of the development, or detection / diagnosis, or absence of sepsis and / or infection, and classifying the subject accordingly. The method may comprise detecting the presence or level of expression of 2, 3, 4, 5, 6, 7, 8, 9,10, 11, 12, or 13 of the nucleic acid markers in the group consisting of BEND2, CR1, EIF4G3, LARP1, LGALS2, MIAT, QSOX1, RPL13A, RRBP1, SGSH, SLC36A1, TDRD9, TEX2, and TRIP13. In a preferred embodiment, the method comprises detecting the presence or level of expression of all of the nucleic acid markers in the group consisting of BEND2 (BEN Domain-Containing Protein 2), CR1 (Complement C3b / C4b Receptor 1 (Knops Blood Group)), EIF4G3 (Eukaryotic Translation Initiation Factor 4 Gamma 3), LARP1 (La Ribonucleoprotein 1, Translational Regulator), LGALS2 (Galectin 2), MIAT (Myocardial Infarction Associated Transcript), QSOX1 (Quiescin Sulfhydryl Oxidase 1), RPL13A (Ribosomal Protein L13a), RRBP1 (Ribosome Binding Protein 1), SGSH (N-Sulfoglucosamine Sulfohydrolase), SLC36A1 (Solute Carrier Family 36 Member 1), TDRD9 (Tudor Domain Containing 9), TEX2 (Testis Expressed 2), and TRIP13 (Thyroid Hormone Receptor Interactor 13). The Applicant has shown that the combination of all fourteen markers, using a method that considers the expression of each of the markers, achieves an especially high performance. The expression of the markers may be increased or decreased with reference to a control, or indeed with respect to each other, but it is the combined analysis (and overall expression) of all fourteen markers that provides for the high accuracy of the method. In another embodiment, is a method for the risk prediction or detection / diagnosis, or absence, of sepsis and / or infection in a subject, comprising the steps of: a) providing at least one biological sample from said subject; b) detecting the presence or level of expression of at least two nucleic acid markers selected from a biomarker signature consisting of DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, SNORA53 BEND2, CR1, EIF4G3, LARP1, LGALS2, MIAT, QSOX1, RPL13A, RRBP1, SGSH, SLC36A1, TDRD9, TEX2, and TRIP13 ('group 3' - a combination of group 1 and group 2), in said sample, compared to a control, or a baseline level of expression, or expression levels with respect to other nucleic acid markers in said biomarker signature; and c) concluding, from said presence or level of expression of said at least two nucleic acid markers, the risk of the development, or detection / diagnosis, or absence of sepsis and / or infection, and classifying the subject accordingly. The method may comprise detecting the presence or level of expression of 2, 3, 4, 5, 6, 7, 8, 9,10, 11, 12, 13, 14, 15, 16,17,18, 19, or 20 of the nucleic acid markers in the group, or indeed all 21 of the nucleic acid markers. The Applicant has shown the three marker combination of QSOX1, RRBP1 and SNORA53 to be especially effective in considering risk prediction of infection in a subject (and especially differentiating from individuals displaying SIRS) with an AUC of 0.98, especially when utilising the single housekeeping gene RPLPO for normalisation. However, signatures having any two of QSOX1, RRBP1 and SNORA53, are also capable of differentiating infection from SIRS with AUCs of between 0.93 and 0.95. Moreover, the individual markers are also capable of differentiating infection from SIRS, especially SNORA53 alone achieving an impressive AUC of 0.93, and RRBP1 alone an AUC of 0.88. Thus the method of the first aspect (and other aspects throughout the application) could alternatively be undertaken through just the detection or expression of a single nucleic acid marker, especially SNORA53 or RRBP1, compared to a control, or a baseline level of expression, to provide a determination. The method for considering the expression of each of the nucleic acid markers, especially with respect to each other, a baseline level of expression, and / or a control, may be achieved through use of a mathematical method, a trained neural network, or an appropriate algorithm, and may be computer-implemented. Through such an approach, the expression of the markers may synergistically combine to provide the high predictive accuracy of the method. An algorithm may be based on Linear Discriminant Analysis (LDA) or Shrinkage Discriminant Analysis (SDA), and may especially utilise the weightings detailed throughout this application, such as those in Tables 1 to 4, or Table 8, to generate a high level of predictivity for either of these biomarker signatures, or groups. The control may be a nucleic acid which has an expression that remains largely constant in samples from a subject / patient irrespective of infection or development of sepsis, and thus enables normalisation within the method. The control may for example be one or more housekeeping genes / nucleic acids whose expression is predictable or relatively static irrespective of whether infection and / or sepsis may develop. For example, the reference gene / nucleic acid may be OTULIN (also known as FAM105B) and / or RANBP3, and / or RPLPO. The Applicant has shown that the housekeeping gene RPLPO, as a control, provides for the best overall performance of the method, with RANBP3 also being especially effective. The control may also be a combination of housekeeping genes. The level of expression of the at least two nucleic acid markers may be determined using a targeted assay to specifically measure the level of expression, which targeted assay may be a nucleic acid amplification assay, which may be a qPCR-based assay. The targeted assay may alternatively, or indeed additionally, utilise a microarray based assay or platform. The method may further comprise a step of identifying a treatment for the subject based on the classification of the subject. The Applicant has determined two groups (or signatures) of nucleic acid markers (biomarkers) which can be used, separately or in combination, in whole or in part, to predict the development of, or absence of, and / or detect / diagnose, infection and / or sepsis, potentially at any point on the trajectory of disease onset, with the possibility of predicting prior to the onset of symptoms (pre-symptomatic). The Applicant has identified through a comprehensive analysis of the host transcriptome, sourced from blood samples from human patients, a first panel of 7 (or indeed 8 with the inclusion of XLOC_003212) nucleic acids, and a second panel of 14 nucleic acids, wherein the expression, and / or combined expression pattern of each marker, within a panel (or combination of first and second panel) has been shown to be highly significant to predicting infection and sepsis, or indeed rule out infection or sepsis, or the likelihood of developing an infection or sepsis. The Applicant has identified two new candidate biomarker signatures through collecting and comparing samples collected at two days prior to the onset of sepsis. The dataset consisted of three classes of sample: sepsis, (non-septic) infection, and comparator (or control), with the aim of developing a model to separate or differentiate sepsis from the other two (i.e. infection; control). Given the challenges in identifying clinical truth in the world of infection, the study undertaken by the Applicant utilised a four clinician expert panel to adjudicate on whether and when a patient was deemed to have an infection (or not), or indeed sepsis, and where possible the severity. The applicant used marker set pre-filter steps (variance and intensity, and / or feature list), and feature selection methodology (recursive feature elimination and / or forward feature selection), followed by the statistical methods partial least squares regression (PLS) and shrinkage discriminant analysis (SDA) to perform a comprehensive analysis of the host transcriptome, sourced from blood samples from human patients. The two biomarker signatures derived from this analysis were then tested with several independent samples sets, to discriminate sepsis patients from uninfected controls, patients with an infection, and patients displaying SIRS (without infection), and also patients with an infection from uninfected controls, and patients displaying SIRS (without infection). The two biomarker signatures were tested with an independent sample set, consisting of 31 sepsis patients and 31 uninfected normal patients, undertaken through qPCR analysis, followed by an analysis of the performance of each signature, at only one time point. The data were optimally normalised, with reference to the expression of control nucleic acids, before calculating the predictive accuracy of the biomarker signatures. For the seven biomarker signature, the AUC for normalised data was 0.994 for differentiating samples from sepsis patients versus samples from control patients, through analysis of the expression (and expression pattern), supported by a mathematical method, of all seven biomarkers. The AUC for non-normalised data was 0.941. This biomarker signature is thus capable of a predictive accuracy in excess of an AUC of 0.9, and in excess of an AUC of 0.94, and indeed in excess of an AUC of 0.99. These accuracies are a phenomenal improvement over previously identified biomarker signatures. The 14 biomarker signature also displayed an AUC of 0.993 for normalised data for differentiating samples from sepsis patients versus samples from control patients, through analysis of the expression (and expression pattern), supported by a mathematical method, of all 14 biomarkers. The AUC for non-normalised data was 0.88. This biomarker signature is thus also capable of a predictive accuracy in excess of an AUC of 0.9, and in excess of an AUC of 0.94, and indeed in excess of an AUC of 0.99. These accuracies are also a phenomenal improvement over previously identified biomarker signatures. The biomarker signatures of DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, and SNORA53 (seven markers), DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, and SNORA53 (six markers), DDX11L9, B4GALT5, PTGES3, TPM3, and SNORA53 (five markers), B4GALT5, PTGES3, TPM3, and SNORA53 (four markers), B4GALT5, TPM3, and SNORA53 (three markers), B4GALT5, and SNORA53 (two markers), and B4GALT5 (one marker) were also tested against the following sample sets, using a comparable approach: Patients with an infection (80 samples) from patients presenting with SIRS (with no infection) (87 samples), providing an AUC 0.974 (0.953-0.994), with a positive predictive value (PPV) 0.915 and a negative predictive value (NPV) 0.952 (using the prevalence of the cohorts used for the analysis), and a PPV 0.674 and NPV 0.990 (assuming a prevalence [for infection in a SIRS population] of 0.15), for the seven markers; AUC 0.972 (0.952-0.993), with a PPV 0.895 and a NPV 0.962 (using the prevalence of the cohorts used for the analysis), and a PPV 0.646 and NPV 0.990 (assuming a prevalence of 0.15), for the six markers; AUC 0.977 (0.958-0.997), with a PPV 0.906 and NPV 0.975 (using the prevalence of the cohorts used for the analysis), and a PPV 0.677 and NPV 0.933 (assuming a prevalence of 0.15), for the five markers; AUC 0.991 (0.981-1.0), with a PPV 0.962 and NPV 0.965 (using the prevalence of the cohorts used for the analysis), and a PPV 0.829 and NPV 0.993 (assuming a prevalence of 0.15), for the four markers; AUC 0.978 (0.961-0.995), with a PPV 0.926 and NPV 0.952 (using the prevalence of the cohorts used for the analysis), and a PPV 0.708 and NPV 0.991 (assuming a prevalence of 0.15), for the three markers; AUC 0.938 (0.905 -0.972), with a PPV 0.822 and NPV 0.922 (using the prevalence of the cohorts used for the analysis), and a PPV 0.371 and NPV 0.991 (assuming a prevalence of 0.15), for the two markers; and AUC 0.545 (0.457-0.634), with a PPV 0.842 and NPV 0.567 (using the prevalence of the cohorts used for the analysis), and a PPV 0.154 and NPV 0.901 (assuming a prevalence of 0.15), for the one marker. Thus evidencing the ability of all of these biomarker signatures to especially accurately identify patients without an infection, or unlikely to develop an infection, as a result of the high NPV achieved with these biomarkers. Patients with a viral infection (43 samples) from patients presenting with SIRS (87 samples), providing an AUC 0.971 (0.941-1.), with a PPV 0.872 and a NPV 0.975 (using the prevalence of the cohorts used for the analysis), and a PPV 0.777 and NPV 0.983 (assuming a prevalence of 0.15), for the seven markers; AUC 0.969 (0.940-0.999), with a PPV 0.836 and a NPV 0.975 (using the prevalence of the cohorts used for the analysis), and a PPV 0.734 and NPV 0.983 (assuming a prevalence of 0.15), for the six markers; AUC 0.972 (0.942-1.0), with a PPV 0.854 and NPV 0.975 (using the prevalence of the cohorts used for the analysis), and a PPV 0.675 and NPV 0.991 (assuming a prevalence of 0.15), for the five markers; AUC 0.986 (0.97-1.0), with a PPV 0.93 and NPV 0.965 (using the prevalence of the cohorts used for the analysis), and a PPV 0.930 and NPV 0.983 (assuming a prevalence of 0.15), for the four markers; AUC 0.978 (0.958-0.997), with a PPV 0.872 and NPV 0.975 (using the prevalence of the cohorts used for the analysis), and a PPV 0.777 and NPV 0.983 (assuming a prevalence of 0.15), for the three markers; AUC 0.940 (0.902-0.978), with a PPV 0.714 and NPV 0.959 (using the prevalence of the cohorts used for the analysis), and a PPV 0.471 and NPV 0.985 (assuming a prevalence of 0.15); AUC 0.565 (0.456-0.673), with a PPV 0.394 and NPV 0.745 (using the prevalence of the cohorts used for the analysis), and a PPV 0.171 and NPV 0.893 (assuming a prevalence of 0.15), for the one marker. Patients with a bacterial infection (19 samples) from patients presenting with SIRS (87 samples), with an AUC 0.976 (0.952-0.999) with a PPV 0.655 and a NPV 1.0 (using the prevalence of the cohorts used for the analysis), and a PPV 0.752 and NPV 0.962 (assuming a prevalence of 0.15) for the seven markers; AUC 0.975 (0.951-0.999), with a PPV 0.633 and a NPV 1.0 (using the prevalence of the cohorts used for the analysis), and a PPV 0.752 and NPV 0.962 (assuming a prevalence of 0.15), for the six markers; AUC 0.981 (0.961-1.0), with a PPV 0.703 and NPV 1.0 (using the prevalence of the cohorts used for the analysis), and a PPV 0.858 and NPV 0.963 (assuming a prevalence of 0.15), for the five markers; AUC 0.994 (0.986-1.0), with a PPV 0.863 and NPV 1.0 (using the prevalence of the cohorts used for the analysis), and a PPV 0.929 and NPV 0.981 (assuming a prevalence of 0.15), for the four markers; AUC 0.972 (0.944-1.0), with a PPV 0.809 and NPV 0.976 (using the prevalence of the cohorts used for the analysis), and a PPV 0.774 and NPV 0.981 (assuming a prevalence of 0.15), for the three markers; AUC 0.926 (0.868-0.984), with a PPV 0.515 and NPV 0.972 (using the prevalence of the cohorts used for the analysis), and a PPV 0.462 and NPV 0.967 (assuming a prevalence of 0.15); AUC 0.580 (0.421-0.739), with a PPV 0.666 and NPV 0.865 (using the prevalence of the cohorts used for the analysis), and a PPV 0.159 and NPV 0.924 (assuming a prevalence of 0.15), for the one marker. Patients with infection (80 samples) plus patients with sepsis (20 samples) from patients presenting with SIRS (87 samples), with an AUC 0.973 (0.953-0.993), with a PPV 0.931 and a NPV 0.941 (using the prevalence of the cohorts used for the analysis), and a PPV 0.598 and NPV 0.994 (assuming a prevalence of 0.15) for the seven markers; AUC 0.970 (0.949-0.991), with a PPV 0.914 and a NPV 0.951 (using the prevalence of the cohorts used for the analysis), and a PPV 0.620 and NPV 0.992 (assuming a prevalence of 0.15), for the six markers; AUC 0.978 (0.960-0.996), with a PPV 0.932 and NPV 0.963 (using the prevalence of the cohorts used for the analysis), and a PPV 0.679 and NPV 0.995 (assuming a prevalence of 0.15), for the five markers; AUC 0.992 (0.983-1.0), with a PPV 0.97 and NPV 0.965 (using the prevalence of the cohorts used for the analysis), and a PPV 0.830 and NPV 0.995 (assuming a prevalence of 0.15), for the four markers; AUC 0.978 (0.961-0.996), with a PPV 0.941 and NPV 0.952 (using the prevalence of the cohorts used for the analysis), and a PPV 0.711 and NPV 0.992 (assuming a prevalence of 0.15), for the three markers; AUC 0.937 (0.904-0.970), with a PPV 0.853 and NPV 0.910 (using the prevalence of the cohorts used for the analysis), and a PPV 0.373 and NPV 0.993 (assuming a prevalence of 0.15), for the two markers; and AUC 0.505 (0.422-0.589), with a PPV 0.857 and NPV 0.506 (using the prevalence of the cohorts used for the analysis), and a PPV 0.154 and NPV 0.906 (assuming a prevalence of 0.15), for the one marker. Patients with sepsis only (20 samples) from patients presenting with SIRS (87 samples), with an AUC 0.968 (0.94-0.995), with a PPV 0.679 and NPV 0.981 (assuming a prevalence [of sepsis in a SIRS population] of 0.15) for the seven markers. Thus also evidencing the ability to especially identify patients without sepsis, or unlikely to develop sepsis (and thus a rule-out methodology), as a result of the high NPV achieved. These results are especially highly compelling for signatures comprising between two and seven markers that can be used to differentiate not only patients with sepsis from controls, but also patients with infection or sepsis from patients displaying SIRS (with no infection), with AUCs between 0.926 and 0.994. However, perhaps even more compelling are the negative predictive values obtained, of between 0.910 and 1.0, thus indicating the value of this biomarker signature in a 'rule-out test', for either ruling out the possibility of a patient having an infection or sepsis. The biomarker signature of QSOX1, RRBP1 and SNORA53 and an SDA algorithm, with respective coefficients, was also established and tested using a comparable approach: Patients with an infection (84 samples) from patients presenting with SIRS (with no infection) (88 samples). The performance was ascertained through application of a 4-fold cross validation, undertaken 1000 times, where in each fold 75% of the samples were randomly used to develop a model, and the remaining 25% of samples used as the test samples in each fold. The data, including accuracy values, are thus derived from the test samples only, to provide performance estimates from the cross validation. This accordingly provided an AUC 0.98 (0.93 - 1.0), with a positive predictive value (PPV) 0.79 (0.56 - 1.0) and a negative predictive value (NPV) 0.99 (0.96 - 1.0) (assuming a prevalence [for infection in a SIRS population] of 0.20). This result is again especially highly compelling for a signature consisting of only three markers that can be used to differentiate patients with infection or sepsis from patients displaying SIRS (with no infection). However, perhaps even more compelling is the negative predictive value obtained, of 0.99 thus indicating the value of this three biomarker signature in a 'rule-out test', for either ruling out the possibility of a patient having an infection or sepsis. Indeed, the combination of any two of QSOX1, RRBP1 and SNORA53, also allowed for high predictivity, with AUCs between 0.93 and 0.95, with even single biomarkers also providing surprisingly high predictivity, with SNORA53 an AUC of 0.93, and RRBP1 and AUC of 0.88. Details of the particular nucleic acids, from their annotated names (as provided throughout this application), can be found from human gene databases such as GeneCards®: The Human Gene Database (genecards.org), or the Human Gene Resources at the National Center for Biotechnology Information (NCBI) (ncbi.nlm.nih.gov / genome / guide / human / ). The Applicant has down-selected lists of nucleic acid markers, as detailed above, critical to predicting or detecting / diagnosing sepsis or infection with a high level of confidence. The term 'biological sample' includes, but not exclusively, blood, serum, plasma, urine, saliva, cerebrospinal fluid or any other form of material, preferably fluid-based or capable of being converted into a fluid-like state (e.g. tissue which can be broken down or separated in a solution, such as a buffered solution), which can be extracted or collected from a patient. In a second aspect, the present invention provides a kit for diagnosing and / or predicting development, or absence, of infection and / or sepsis in a subject, said kit comprising reagents and / or systems for determining levels of at least two nucleic acid markers from a biomarker signature, in a biological sample from a subject, wherein the biomarker signature is either a first group consisting of DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, and SNORA53, or a second group consisting of BEND2, CR1, EIF4G3, LARP1, LGALS2, MIAT, QSOX1, RPL13A, RRBP1, SGSH, SLC36A1, TDRD9, TEX2, and TRIP13, or a third group (combination of first group and second group) and optionally reagents and / or systems for determining levels of at least one control nucleic acid or housekeeping nucleic acid. The reagents and / or systems for determining levels of at least two nucleic acid markers from a biomarker signature, are preferably reagents specific to the selected nucleic acids, so enabling specificity to and determination of each particular nucleic acid, preferably with minimal or no crossreactivity with other nucleic acids. The kit may comprise reagents and / or systems for determining levels of 2, 3, 4, 5, 6 or 7 of the nucleic acid markers in the group consisting of DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, and SNORA53. In one embodiment, the reagents and / or systems may be for determining levels of at least B4GALT5 and SNORA53, which combination of two markers provides for an AUC in excess of 0.9. In a preferred embodiment, the kit comprises reagents and / or systems for determining levels of each of the nucleic acid markers in the group consisting of DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, and SNORA53. The Applicant has shown that the combination of all seven nucleic acid markers achieves an especially high performance, as do subsets of at least two markers, with single markers also providing for differentiation. The kit may comprise reagents and / or systems for determining levels of 2, 3, 4, 5, 6, 7, 8, 9,10,11,12, or 13 of the nucleic acid markers in the group consisting of BEND2, CR1, EIF4G3, LARP1, LGALS2, MIAT, QSOX1, RPL13A, RRBP1, SGSH, SLC36A1, TDRD9, TEX2, and TRIP13. In a preferred embodiment, the kit comprises reagents and / or systems for determining levels of each of the nucleic acid markers in the group consisting of BEND2, CR1, EIF4G3, LARP1, LGALS2, MIAT, QSOX1, RPL13A, RRBP1, SGSH, SLC36A1, TDRD9, TEX2, and TRIP13. The Applicant has shown that the combination of all fourteen nucleic acid markers achieves an especially high performance. The kit may comprise reagents and / or systems for determining levels of the nucleic acid markers QSOX1, RRBP1 and SNORA53, or subsets thereof. In a particular embodiment, the reagents or systems comprises means for detecting a nucleic acid, for example RNA, such as mRNA. The kit of the second aspect may comprise reagents and / or systems for determining levels of at least one control nucleic acid or housekeeping nucleic acid, wherein the at least one control nucleic acid or housekeeping nucleic acid may be OTULIN and / or RANBP3 and / or RPLPO. The reagents or systems may include use of recognition elements, or microarray-based and / or Next Generation Sequencing (NGS) methods. Thus in a particular embodiment, the kit of the invention comprises a microarray on which are immobilised probes suitable for binding to RNA expressed by each nucleic acid of a biomarker signature. In an alternative embodiment, the kit comprises at least some of the reagents suitable for carrying out amplification of nucleic acids of the biomarker signature, or regions thereof. In one embodiment the reagents or systems are for or use real-time (RT) polymerase chain reaction (PCR). In such cases, the reagents may comprise primers for amplification of said nucleic acids or regions thereof (and preferably primers specific to each nucleic acid marker). The kits may further comprise labels, in particular fluorescent labels, and / or oligonucleotide probes to allow the PCR to be monitored in real-time using any of the known assays, such as TaqMan, LUX, etc. The kits may also contain reagents such as buffers, enzymes, salts such as MgCI etc. required for carrying out a nucleic acid amplification reaction. The reagents, especially for nucleic acid amplification, may comprise for example one or more of fluorescently-labelled oligonucleotide probes or fluorescently-labelled primers, wherein the fluorescently-labelled oligonucleotide probes or fluorescently-labelled primers may consist of probes and primers each capable of specific binding and detection of nucleic acid products. Primers and / or probes to the specific nucleic acid sequences of each biomarker may be custom synthesised, or may be sourced / purchased from a number of suppliers, such as ThermoFisher Scientific, Integrated DNA Technologies (IDT), or OriGene. For example, the ThermoFisher catalogue includes gene expressions assays, including primers and / or probes, for TIFA (ID: Hs01848307_sl), B4GALT5 (ID: Hs00941041_ml), PTGES3 (ID: Hs00832847_gH), TPM3 (ID: Hs00383595_ml), CDC42 (ID: Hs00741586_mH), SNORA53 (ID: Hs03298722_sl), FAM015B (ID: Hs00385644_ml), and RANBP3 (ID: Hs00999818_ml). The method of the first aspect may advantageously be computer-implemented to handle the complexity in monitoring and analysis of the numerous biomarkers, and their respective relationships to each other. Such a computer-implemented invention could enable a yes / no answer as to whether infection and / or sepsis is present, or absent, and / or likely to develop, or at least provide an indication of how likely the development is. The method preferably uses mathematical tools and / or algorithms to monitor and assess the (expression) levels of the biomarkers (the nucleic acid markers, or products thereof) both qualitatively and quantitatively. The tools could in particular include support vector machine (SVM) algorithms, decision trees, random forests, artificial neural networks, quadratic discriminant analysis, and Bayes classifiers. In one embodiment the data from monitoring all biomarkers in the biomarker signature is assessed by means of an artificial neural network, or algorithm, which may be derived from a partial least squares regression analysis of training data, or a linear or shrinkage discriminant analysis of training data. In one embodiment, the computer implemented method utilises an algorithm based on Shrinkage Discriminant Analysis (SDA), which may comprise use of the weightings detailed throughout this application, such as detailed in Table 1, Table 2, Table 3, Table 4 or Table 8 to provide a risk prediction with high predictivity, which may be at least 90%, at least 95%, or at least 99% accurate. A computer implemented method may comprise measuring the level of expression of biomarkers within a specific biomarker signature to generate a prediction as to whether a subject has or may develop sepsis, or indeed is infected, and whether they may become septic. In one embodiment of the first aspect the method is a computer-implemented method wherein the monitoring, measuring and / or detecting comprises producing quantitative, and optionally qualitative, data for all nucleic acid markers, inputting said data into an analytical process on the computer, using at least one mathematical method, that may compare the data with reference (control) data, and producing an output from the analytical process which provides a prediction for the likelihood of developing infection and / or sepsis, or absence of infection and / or sepsis, or positive detection or diagnosis of infection or sepsis. Reference data may include data from healthy subjects, subjects diagnosed with sepsis, subjects with infection, and subjects with SIRS, but no infection. Reference data may include, or comprise, expression data for control or housekeeping nucleic acids. The output from the analytical process may enable the time to onset of symptoms to be predicted, such as 1, 2, or 3 days prior to onset of symptoms, and consequently may be particularly valuable and useful to a medical practitioner in suggesting a course of treatment, especially when the choice of course of treatment is dependent on the progression of the disease. The method may also enable monitoring of the success of any treatment, assessing whether the likelihood of onset of symptoms decreases over the course of treatment. In a third aspect, the present invention provides a computer implemented method for the risk prediction or detection / diagnosis, or absence, of sepsis and / or infection in a subject, comprising the steps of: a) inputting the level of expression of at least two nucleic acid markers selected from a biomarker signature, from a sample obtained from the subject; and b) concluding from the level of expression of said at least two nucleic acid markers, the risk of the development, or detection / diagnosis, or absence, of sepsis and / or infection, and classifying the subject accordingly; wherein the biomarker signature is either a first group consisting of DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, and SNORA53, or a second group consisting of BEND2, CR1, EIF4G3, LARP1, LGALS2, MIAT, QSOX1, RPL13A, RRBP1, SGSH, SLC36A1, TDRD9, TEX2, and TRIP13, or a third group consisting a combination of the first group and the second group. The computer-implemented method may utilise an algorithm based on Shrinkage Discriminant Analysis or partial Least Squares regression to evaluate the level of expression of the at least two nucleic acid markers, to conclude the risk of the development, or detection / diagnosis, or absence, of sepsis and / or infection, and classify the subject accordingly. In a fourth aspect, the present invention provides an apparatus or system comprising a processor and one or more computer readable media instructions that, when executed by the processor, cause the processor to implement the method of the first or third aspects. The apparatus or system may be, or be part of, a nucleic acid amplification system. The system or apparatus of the fourth aspect may comprise the kit of the second aspect, and / or a computer programme product, to implement the method of the first or third aspects and / or a computer readable medium. In a fifth aspect, the present invention provides a computer readable medium or media storing instructions that, when executed by a processor, cause the processor to implement the method of the first or third aspects. In a sixth aspect the present invention provides a computer programme product comprising instructions that, when executed by a processor, cause the processor to implement the method of the first or third aspects. The biomarker signatures of the different aspects of the present invention may be used to monitor and / or predict the response of a subject to a particular therapeutic agent, such as a sepsis targeting drug or an antibiotic. For example, the expression of particular nucleic acid markers, which may be elevated or reduced in a subject on a course to develop sepsis, or having already developed sepsis, could be monitored to establish whether the levels are returning to the levels expected for a subject without, or unlikely to develop, sepsis, which could be an indication of the therapeutic agent successfully treating the subject. A therapeutic agent may be one targeted to particular subsets of markers in an attempt to treat the subject, and indeed the choice of therapeutic agent to be used in a subject may be determined by the expression of specific nucleic acid markers which are most affected, or differ most from a control or from a patient not predicted to develop sepsis, as a result of developing sepsis. For example, the elevation of certain markers may suggest the use of one therapeutic agent, whereas elevation of a different subset of markers, may suggest use of another therapeutic agent. Any feature in one aspect of the invention may be applied to any other aspects of the invention, in any appropriate combination. In particular, method aspects may be applied to kit aspects and vice versa. The invention extends to methods, uses, systems, and kits substantially as herein described, with reference to the Example(s). In all aspects, the invention may comprise, consist essentially of, or consist of any feature or combination of features. Examples Shrinkage Discriminant Analysis (SDA) Model resulting from a Comprehensive Analysis of the Host Transcriptome, Sourced from Blood Samples from Human Patients identified as Septic Class, or Non-Septic Class The method was based on linear discriminant analysis (LDA) with shrinkage of the covariance matrix between nucleic acids in classifier training, and feature selection using correlation-adjusted t scores. The Applicant has identified through a comprehensive analysis of the host transcriptome, sourced from blood samples from human patients, a first panel of 7 nucleic acids, and a second panel of 14 nucleic acids, wherein the expression and / or combined expression pattern of each marker, within a panel has been shown to be highly significant to predicting infection and sepsis, or indeed rule out infection or sepsis, or the likelihood of developing an infection or sepsis. The Applicant has identified the two new candidate biomarker signatures through collecting and comparing samples collected at two days prior to the onset of sepsis. The dataset consisted of three classes of sample: sepsis, (non-septic) infection, and comparator (or control), with the aim of developing a model to separate or differentiate sepsis from the other two (i.e. infection; control). The applicant used marker set pre-filter steps (variance and intensity, and / or feature list), and feature selection methodology (recursive feature elimination and / or forward feature selection), followed by shrinkage discriminant analysis (SDA) to perform a comprehensive analysis of the host transcriptome, sourced from blood samples from human patients, to identify the panels of markers, and thereby derive parameters and coefficients (pw and ref) to enable the analysis and classification of new samples from a subject, to conclude from the level of expression of at least two nucleic acid markers (or even at least one nucleic acid marker), the risk of the development, or detection / diagnosis, or absence, of sepsis and / or infection, and classify the subject accordingly In order to score a new sample, discriminant score for each class was calculated from the nucleic acid expression data as follows (where A is the non-septic class and B is the septic class, or alternatively A is the non-infected (SIRS) class, and B is the infected class, as applicable):- (7 or 14 \ (Class A pwi x (xL — Class A ref^ j (7 or 14 \ Z (Class B pwt Class B reQ) \ i=l / Where xt is the expression of nucleic acid i; this is combined with class-specific prediction coefficients pw and ref for each nucleic acid. Nucleic acid and class specific prediction coefficients (pw and ref) are provided in Table 1, Table 2, Table 3 and Table 4; ref is the nucleic acid expression centroid for each nucleic acid within each class, and pw is feature weight. Note that for model training both priors were set to 0.5 based on training endpoint distribution. A and B are then transformed using a scaling factor defined by the following equations: A_Transform = exp(A — max (A, B)) B_Transform = exp(B — max(A., B)) The transformed values for A and B are then used to calculate a probability of class membership for group A and B where: A Transform pr ryY) =__________-____________________ A_Transform + B_Transform B Transform Pr(B) =--------=---------------- A_Transform + B_Transform The probability scores are subject to a final transformation to the logit scale. If Pr(B)<0.5 / Pr(B) \ Final Signature Score = In J If Pr(B) >=0.5 / Pr (A) \ Final Signature Score = — In J Following the final transformation of the signature scores the discrimination threshold is expected to be 0 (equivalent to probability of 0.5). Table 1. Nucleic acid and class specific prediction coefficients (pw and ref) for an SDA based machine learning classification algorithm utilising the seven nucleic acid signature of DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, and SNORA53. Prediction Coefficients Nucleic Acid ( / j Class A pw Class B pw Class A ref Class B ref DDX11L9 0.7302 -0.7162 8.8985 8.5383 TIFA -0.6897 0.6765 3.5041 3.8555 B4GALT5 -0.921 0.9033 9.9029 10.397 PTGES3 -1.1716 1.149 4.9484 5.226 TPM3 -1.683 1.6506 10.0707 10.3842 CDC42 -0.3251 0.3189 9.2296 9.5256 SNORA53 1.3384 -1.3126 4.9808 4.7061 The nucleic acid expression centroid for each gene within each class, ref, can also be considered as the average normalized expression per class. The closer the actual normalized expression is to the ref value for a class, the more likely it is to be in that class. The feature weights pw indicate how much to weight the deviation from the average normalized expression per class. Note pw for a class 5 is positive if that class generally has a higher average expression, and negative if that class generally has a lower average expression. Thus, if class B (septic class) has a higher re / than class A (non-septic class), and indeed pw for Class B is a positive integer, then we can generally infer that over-expression is associated with class B, and under-expression is associated with class A. Over-expression and under-expression are of course 10 relative to the specific nucleic acid, and the ref value of each nucleic acid, given the ref values range from 3.8555 to 10.397 (the value of ref is on the Iog2 scale, and thus a nucleic acid with an expression level of 10 is 128 times more highly expressed than a nucleic acid with an expression level of 3). So to interpret the data and parameters in Table 1, TIFA, B4GALT5, PTGES3, TPM3, and CDC42 can be regarded as over-expressed in Class B, the septic class, and DDX11L9 and SNORA53 as underexpressed in Class B. Table 2. Nucleic acid and class specific prediction coefficients (pw) for an SDA based machine 5 learning classification algorithm utilising down-selected nucleic acid signatures from the seven nucleic acid signature in Table 1; the six nucleic acid signature of DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, and SNORA53, five nucleic acid signature of DDX11L9, B4GALT5, PTGES3, TPM3, and SNORA53, four nucleic acid signature of B4GALT5, PTGES3, TPM3, and SNORA53, and three nucleic acid signature of B4GALT5, TPM3, and SNORA53. Note that the ref values remain constant for all 10 signatures, it is just the pw coefficients that change when the signature is re-trained after feature removal from 7 genes down to 6, 5, 4 and 3, and thus Class A re / and Class B ref values are as for Table 1 for each nucleic acid, respectively. A designates Class A; B designates Class B. Nucleic Acid 6 nucleic acid signature (pw) 5 nucleic acid signature (pw) 4 nucleic acid signature (pw) 3 nucleic acid signature (pw) A / B A / B A / B A / B DDX11L9 0.72846 / -0.71445 0.79133 / -0.77611 TIFA -0.73595 / 0.72179 B4GALT5 -1.0027 / 0.98345 -0.83775 / 0.82164 -1.1158 / 1.0944 -1.2292 / 1.2056 PTGES3 -1.2224 / 1.1989 -1.3511 / 1.3251 -1.2378 / 1.214 TPM3 -1.6796 / 1.6473 -2.1049 / 2.0644 -1.2034 / 1.1803 -0.92764 / 0.9098 SNORA53 1.3091 / -1.284 1.285 / -1.2603 1.3998 / -1.3729 1.3201 / -1.2948 Table 3. Nucleic acid and class specific prediction coefficients (pw) for an SDA based machine learning classification algorithm utilising down-selected nucleic acid signatures from the seven gene signature in Table 1; the two nucleic acid signature of B4GALT5 and SNORA53, and one nucleic acid signature of B4GALT5. Note that the ref values remain constant for all signatures, it is just the pw 5 coefficients that change when the signature is re-trained after feature removal from 7 nucleic acids down to 2 and 1, and thus Class A ref and Class B ref values are as for Table 1 for each nucleic acid, respectively. A designates Class A; B designates Class B. Nucleic Acid 2 nucleic acid signature (pw) 1 nucleic acid signature (pw) A / B A / B B4GALT5 -1.5337 / 1.5042 -1.3932 / 1.3664 SNORA53 1.0207 / -1.0011 Table 4. Nucleic acid and class specific prediction coefficients (pw and ref) for an SDA based machine 10 learning classification algorithm utilising the fourteen nucleic acid signature of BEND2, CR1, EIF4G3, LARP1, LGALS2, MIAT, QSOX1, RPL13A, RRBP1, SGSH, SLC36A1, TDRD9, TEX2, and TRIP13. Prediction Coefficients Nucleic Acid ( / j Class A pw Class B pw Class A ref Class B ref BEND2 0.1328 -0.2946 5.2914 5.3207 CR1 -0.2126 0.4718 7.8738 8.2461 EIF4G3 -0.2279 0.5058 7.0118 7.1881 LAR Pl -0.0225 0.0499 9.6557 9.8405 LGALS2 0.0271 -0.0602 9.7915 9.1742 MIAT -0.5754 1.2767 6.7212 6.9290 QSOX1 0.0629 -0.1395 8.6734 8.8306 RPL13A -0.0156 0.0347 13.3132 13.2167 RRBP1 -0.0901 0.1999 7.2705 7.5056 SGSH 0.2321 -0.5149 9.7869 9.9580 SLC36A1 0.1409 -0.3126 8.9118 9.1601 TDRD9 -0.0929 0.2061 8.5812 9.1308 TEX2 -0.6683 1.4828 6.3044 6.6212 TRIP13 0.1576 -0.3498 5.0975 5.1238 Partial Least Squares (PLS) Model resulting from a Comprehensive Analysis of the Host Transcriptome, Sourced from Blood Samples from Human Patients identified as Septic Class or Non- Septic Class: Fourteen Nucleic Acid Signature 5 A partial Least Squares regression approach, or model, was also created, used and trained with expression data from patient samples denoted as septic, infection, or control, and the resulting algorithm used to characterise samples from patients. The equation for the PLS model is: Signature Score = my + x (xi - mxi) 10 Where my (mean y), b and mx (mean x) are calculated model parameters and x is the gene level expression data for nucleic acids i = 1-14 Table 5. Nucleic acid and class specific prediction coefficients (b and mx) for a PLS based machine learning classification algorithm utilising the fourteen nucleic acid signature of BEND2, CR1, EIF4G3, LARP1, LGALS2, MIAT, QSOX1, RPL13A, RRBP1, SGSH, SLC36A1, TDRD9, TEX2, and TRIP13. Mean y (my) 15 is 0.3107. Nucleic Acid b Mean x BEND2 -0.0142 5.3005 CR1 0.036 7.9895 EIF4G3 0.054 7.0666 LAR Pl 0.1142 9.7131 LGALS2 -0.024 9.5997 Ml AT 0.0868 6.7857 QSOX1 -0.0147 8.7222 RPL13A 0.0117 13.2832 RRBP1 0.0501 7.3436 SGSH -0.0167 9.8401 SLC36A1 -0.0362 8.9889 TDRD9 0.0525 8.752 TEX2 0.2311 6.4028 TRIP13 -0.029 5.1057 Through the analysis and model development, it has been shown that subsets of markers from either signature (seven, or fourteen) were able to differentiate categories of patients, such as sepsis patients from controls, and infected patients from non-infected SIRS patients. Ranking in importance, for 5 differentiating between categories (such as sepsis versus control), of the seven nucleic acid signature were B4GALT5 (rank 1), SNORA53, TPM3, PTGES3, DDX11L9, TIFA, CDC42 (rank 7), with subsets of markers, for example subsets of two or three markers, especially B4GALT5 and SNORA53 (two markers), and B4GALT5, SNORA53, and TPM3 (three markers), evidenced as being capable of differentiating between the different categories of sample, such as sepsis and controls. It is however believed that a single biomarker, from either model 1 or model 2, is capable of differentiating between the different categories of sample, especially in a rule-out test. Evaluation of the two predictive panel of biomarkers for infection and sepsis, with utility at any part of the trajectory of disease onset. This study tested the performance of the two candidate signatures in a trial cohort through the use of qPCR data. The two signatures consist of (1) DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, and SNORA53 (Model 1), and (2) BEND2, CR1, EIF4G3, LARP1, LGALS2, MIAT, QSOX1, RPL13A, RRBP1, SGSH, SLC36A1, TDRD9, TEX2, and TRIP13 (Model 2), which were derived from an analysis of expression data from a large cohort of patients and control patients, through marker set pre-filter steps (variance and intensity, and / or feature list), feature selection methodology (recursive feature elimination and / or forward feature selection), and shrinkage discriminant analysis (SDA) and / or partial least squares regression (PLS). The evaluation utilised the SDA Model, and the parameters detailed in Tables 1 to 4, as applicable. The number of samples in this study is 62, with samples from patients with sepsis (n = 31) and controls (n = 31). The event rate is thus 50%. It was decided to perform a screening of different data transformations and try to understand which transformations would result in the best transformation from the qPCR platform to a microarray platform. A dataset was created for each combination of the following transformations: Housekeeping nucleic acid normalisation, and bridging to microarray data. All data sets were scale inverted. The order of the transformations was: 1) scale inversion 2) housekeeping nucleic acid normalisation 3) bridging of data to microarray data. Scale Inversion Microarray data is Iog2 transformed counts. The larger the value the higher the amount of mRNA. qPCR data is Ct values. The larger the Ct value the lower the amount of mRNA. The scale of qPCR data must be inverted, such that a higher value means a higher amount of mRNA, to better reflect the format of microarray data. This was done by subtracting the measurements for each gene in each sample from the respective LLOQ. (lower limit of quantification) value. Housekeeping nucleic acid (gene) normalisation Housekeeping nucleic acid normalisation was used to account for the differences in mRNA content across all samples. The housekeeping nucleic acids used were OTULIN and RANBP3. Bridging to microarray data By comparing the measurements of the same samples run on both the qPCR and a microarray platform (Agilent), it was possible to correct eventual nucleic acid-specific median shifts in the measurements caused by the technical differences between the platforms. The bridging is performed after the scale inversion. Inspect the effect of the data transformations on the order of the scores Both a matrix of scatter plots comparing all combinations of data transformations and individual scatterplots were created. The matrix of scatter plots with spearman correlation coefficients was created using the base R function pairs (R Core Team, 2022). The individual scatter plots were created using the R-package ggplot2, and the spearman regression coefficient was calculated using stat_cor from the R-package ggpubr. AUC analysis The performance of the two signatures were reported as are under the ROC curve (AUC), sensitivity, specificity, positive predictive value (PPV) and negative predictive value (NPV) investigated with bootstrapped 95 % confidence intervals using 2000 stratified bootstrapped replicates. Youdens index also referred to as the best point, was used to identify the cut-off with the highest sensitivity and specificity. The performance metrics were calculated using the function roc from the R-Package pRoc. The formulas used for NPV and PPV are: PPV = True positives / (True positives + False positives); NPV = True negatives / (False negatives + True negatives). Sixteen rounds of AUC analysis were performed. Eight rounds for each of the two signatures, Software information All analyses described in this report were made in R version 4.2.2 (2022-10-31). Missing data points were imputed using impute.knn from the R-package impute version 1.58.0. Scatter plots were created using the R-package ggplot2 version 3.4.1. Spearman correlation coefficients were calculated using the stat_cor function from the R-package ggpubr version 0.6.0. The performance metrics of the signatures were calculated using the function roc from the R-Package pRoc version 1.18.0. The matrix of scatter plots with spearman correlation coefficients was created using the base R function pairs. Results AUC analysis The results from the rounds of AUC analysis are shown in Table 6. All rounds showed great performance. The highest AUC value for both signatures is achieved with the housekeeping nucleic acid normalised data (0.99 (0.98-1) for both). Table 6: Results from the AUC analysis showing performance metrics for the two different signatures. 5 Abbreviations: PPV = positive predictive value, NPV = negative predictive value, BPt = best point "Youdens Index" i.e. the threshold, HKG = house keeping nucleic acid (gene) ID AUC Sensitivity Specificity PPV NPV BPt Model 1 Original 0.94 (0.88-1) 0.87 (0.8-1) 0.92 (0.72-1) 0.93 (0.81-1) 0.85 (0.78-1) -24.97 Bridged 0.94 (0.88-1) 0.87 (0.8-1) 0.92 (0.76-1) 0.93 (0.81-1) 0.85 (0.79-1) -2.5 HKG norm 0.99 (0.98-1) 0.93 (0.87-1) 1 (0.88-1) 1 (0.91-1) 0.92 (0.86-1) -24.15 HKG norm & Bridged 0.99 (0.98-1) 0.93 (0.87-1) 1 (0.88-1) 1 (0.91-1) 0.92 (0.86-1) -0.91 Model 2 - Original 0.87 (0.78- 0.96) 0.9 (0.67-1) 0.77 (0.67- 1) 0.8 (0.74-1) 0.88 (0.74- 1) 0 Bridged 0.87 (0.78- 0.96) 0.9 (0.67-1) 0.77 (0.67- 1) 0.8 (0.74-1) 0.88 (0.74- 1) 1.98 HKG norm 0.99 (0.98- 1) 1 (0.87-1) 0.93 (0.87- 1) 0.93 (0.88- 1) 1 (0.88-1) -0.01 HKG norm & Bridged 0.99 (0.98- 1) 1 (0.9-1) 0.93 (0.87- 1) 0.93 (0.88- 1) 1 (0.91-1) 1.98 Both signatures perform very well in this cohort, especially when the data has been housekeeping nucleic acid normalised before analysis. The procedure of bridging the data to Agilent microarray data does not affect the performance of the signatures. The qPCR data has successfully been transformed to be comparable to Agilent microarray data. The best performance for both signatures is obtained using scale inversion and housekeeping nucleic acid normalisation. Bridging the data to Agilent microarray data does not affect performance. The best metrics for model 1 are AUC = 0.99 (0.98-1), sensitivity = 0.93 (0.87-1), specificity = 1 (0.88-1), PPV = 1 (0.91-1), and NPV = 0.92 (0.86-1). The best metrics for model 2 are AUC = 0.99 (0.98-1), Sensitivity = 1 (0.87-1), specificity = 0.93 (0.87-1), PPV = 0.93 (0.88-1), NPV = 1 (0.88-1). The Applicant has further shown, using a comparable approach with these 62 samples, that an eight biomarker signature comprising the seven marker signature of Model 1 together with the nucleic acid XLOC_003212 also performs as well as the seven marker signature, with an AUC of 0.99 (0.98-1.0) for normalised data for differentiating samples from sepsis patients versus samples from control patients, through analysis of the expression (and expression pattern) of all eight biomarkers. The AUC for nonnormalised data was 0.94 (0.88-1.0). This eight biomarker signature is thus capable of a predictive accuracy in excess of an AUC of 0.9, and in excess of an AUC of 0.94, and indeed in excess of an AUC of 0.99. However, XLOC_003212 is a non-coding nucleic acid, and also has very low expression in many samples, and so the investigation has generally focussed on the seven marker signature. The aspects of the invention may therefore utilise this eight biomarker signature in place of the seven biomarker signature, and subsets thereof. Further studies using a comparable approach to that detailed above (scale inversion, housekeeping nucleic acid normalisation (OTULIN and RANBP3), bridging to microarray data, SDA Model, and ROC analysis), with the seven marker signature DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, and SNORA53, and subsets thereof of DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, and SNORA53 (six markers), DDX11L9, B4GALT5, PTGES3, TPM3, and SNORA53 (five markers), B4GALT5, PTGES3, TPM3, and SNORA53 (four markers), B4GALT5, TPM3, and SNORA53 (three markers), B4GALT5 and SNORA53 (two markers), and B4GALT5 (one marker) were also undertaken to consider the ability of the signature to differentiate between (1) patients with an infection (80 samples) from patients presenting with SIRS (87 samples); (2) patients with a viral infection (43 samples) from patients presenting with SIRS (87 samples); (3) patients with a bacterial infection (19 samples) from patients presenting with SIRS (87 samples); and (4) patients with infection (80 samples) or sepsis (20 samples) from patients presenting with SIRS (87 samples). The nomenclature (1) to (4) is used in the following sections to refer to these four comparative analyses. The seven marker signature was also evaluated for discriminating patients with sepsis (20 samples) from patients presenting with SIRS (87 samples). The seven marker signature achieved (1) AUC 0.974 (0.953-0.994), with a PPV 0.915 and a NPV 0.952 (using the prevalence of the cohorts used for the analysis), and a PPV 0.674 and NPV 0.990 (assuming a prevalence [for infection in a SIRS population] of 0.15), and a PPV 0.834 and NPV 0.977 (assuming a prevalence of 0.30); (2) AUC 0.971 (0.941-1.), with a PPV 0.872 and a NPV 0.975 (using the prevalence of the cohorts used for the analysis), and a PPV 0.777 and NPV 0.983 (assuming a prevalence of 0.15), and a PPV 0.894 and NPV 0.959 (assuming a prevalence of 0.30); (3) AUC 0.976 (0.952-0.999) with a PPV 0.655 and a NPV 1.0 (using the prevalence of the cohorts used for the analysis), and a PPV 0.752 and NPV 0.962 (assuming a prevalence of 0.15), and a PPV 0.880 and NPV 0.913 (assuming a prevalence of 0.30); (4) AUC 0.973 (0.953-0.993), with a PPV 0.931 and a NPV 0.941 (using the prevalence of the cohorts used for the analysis), and a PPV 0.598 and NPV 0.994 (assuming a prevalence of 0.15), and a PPV 0.783 and NPV 0.986 (assuming a prevalence of 0.30), respectively. The six marker signature achieved (1) AUC 0.972 (0.952-0.993), with a PPV 0.895 and a NPV 0.962 (using the prevalence of the cohorts used for the analysis), and a PPV 0.646 and NPV 0.990 (assuming a prevalence of 0.15), and a PPV 0.816 and NPV 0.977 (assuming a prevalence of 0.30); (2) AUC 0.969 (0.940-0.999), with a PPV 0.836 and a NPV 0.975 (using the prevalence of the cohorts used for the analysis), and a PPV 0.734 and NPV 0.983 (assuming a prevalence of 0.15), and PPV 0.870 and NPV 0.959 (assuming a prevalence of 0.30); (3) AUC 0.975 (0.951-0.999), with a PPV 0.633 and a NPV 1.0 (using the prevalence of the cohorts used for the analysis), and a PPV 0.752 and NPV 0.962 (assuming a prevalence of 0.15), and PPV 0.880 and NPV 0.913 (assuming a prevalence of 0.30); (4) AUC 0.970 (0.949-0.991), with a PPV 0.914 and a NPV 0.951 (using the prevalence of the cohorts used for the analysis), and a PPV 0.620 and NPV 0.992 (assuming a prevalence of 0.15), and PPV 0.798 and NPV 0.981 (assuming a prevalence of 0.30), respectively. The five marker signature achieved (1) AUC 0.977 (0.958-0.997), with a PPV 0.906 and NPV 0.975 (using the prevalence of the cohorts used for the analysis), and a PPV 0.677 and NPV 0.933 (assuming a prevalence of 0.15), and a PPV 0.836 and NPV 0.983 (assuming a prevalence of 0.30); (2) AUC 0.972 (0.942-1.0), with a PPV 0.854 and NPV 0.975 (using the prevalence of the cohorts used for the analysis), and a PPV 0.675 and NPV 0.991 (assuming a prevalence of 0.15), and PPV 0.835 and NPV 0.979 (assuming a prevalence of 0.30); (3) AUC 0.981 (0.961-1.0), with a PPV 0.703 and NPV 1.0 (using the prevalence of the cohorts used for the analysis), and a PPV 0.858 and NPV 0.963 (assuming a prevalence of 0.15), and a PPV 0.936 and NPV 0.915 (assuming a prevalence of 0.30); (4) AUC 0.978 (0.960-0.996), with a PPV 0.932 and NPV 0.963 (using the prevalence of the cohorts used for the analysis), and a PPV 0.679 and NPV 0.995 (assuming a prevalence of 0.15), and PPV 0.837 and NPV 0.986 (assuming a prevalence of 0.30), respectively. The four marker signature achieved (1) AUC 0.991 (0.981-1.0), with a PPV 0.962 and NPV 0.965 (using the prevalence of the cohorts used for the analysis), and a PPV 0.829 and NPV 0.993 (assuming a prevalence of 0.15), and a PPV 0.922 and NPV 0.983 (assuming a prevalence of 0.30); (2) AUC 0.986 (0.97-1.0), with a PPV 0.93 and NPV 0.965 (using the prevalence of the cohorts used for the analysis), and a PPV 0.930 and NPV 0.983 (assuming a prevalence of 0.15), and a PPV 0.970 and NPV 0.961 (assuming a prevalence of 0.30); (3) AUC 0.994 (0.986-1.0), with a PPV 0.863 and NPV 1.0 (using the prevalence of the cohorts used for the analysis), and a PPV 0.929 and NPV 0.981 (assuming a prevalence of 0.15), and a PPV 0.970 and NPV 0.956 (assuming a prevalence of 0.30); (4) AUC 0.992 (0.983-1.0), with a PPV 0.97 and NPV 0.965 (using the prevalence of the cohorts used for the analysis), and a PPV 0.830 and NPV 0.995 (assuming a prevalence of 0.15), and PPV 0.922 and NPV 0.987 (assuming a prevalence of 0.30), for the four markers, respectively. The three marker signature achieved (1) AUC 0.978 (0.961-0.995), with a PPV 0.926 and NPV 0.952 (using the prevalence of the cohorts used for the analysis), and a PPV 0.708 and NPV 0.991 (assuming a prevalence of 0.15), and a PPV 0.855 and NPV 0.978 (assuming a prevalence of 0.30); (2) AUC 0.978 (0.958-0.997), with a PPV 0.872 and NPV 0.975 (using the prevalence of the cohorts used for the analysis), and a PPV 0.777 and NPV 0.983 (assuming a prevalence of 0.15), and a PPV 0.894 and NPV 0.959 (assuming a prevalence of 0.30); (3) AUC 0.972 (0.944-1.0), with a PPV 0.809 and NPV 0.976 (using the prevalence of the cohorts used for the analysis), and a PPV 0.774 and NPV 0.981 (assuming a prevalence of 0.15); (4) AUC 0.978 (0.961-0.996), with a PPV 0.941 and NPV 0.952 (using the prevalence of the cohorts used for the analysis), and a PPV 0.711 and NPV 0.992 (assuming a prevalence of 0.15), and a PPV 0.856 and NPV 0.982 (assuming a prevalence of 0.30), respectively. The two marker signature achieved (1) AUC 0.938 (0.905 -0.972), with a PPV 0.822 and NPV 0.922 (using the prevalence of the cohorts used for the analysis), and a PPV 0.371 and NPV 0.991 (assuming a prevalence of 0.15), and a PPV 0.589 and NPV 0.978 (assuming a prevalence of 0.30); (2) AUC 0.940 (0.902-0.978), with a PPV 0.714 and NPV 0.959 (using the prevalence of the cohorts used for the analysis), and a PPV 0.471 and NPV 0.985 (assuming a prevalence of 0.15), and PPV 0.684 and NPV 0.965 (assuming a prevalence of 0.30); (3) AUC 0.926 (0.868-0.984), with a PPV 0.515 and NPV 0.972 (using the prevalence of the cohorts used for the analysis), and a PPV 0.462 and NPV 0.967 (assuming a prevalence of 0.15), and a PPV 0.676 and NPV 0.924 (assuming a prevalence of 0.30); (4) AUC 0.937 (0.904-0.970), with a PPV 0.853 and NPV 0.910 (using the prevalence of the cohorts used for the analysis), and a PPV 0.373 and NPV 0.993 (assuming a prevalence of 0.15), and a PPV 0.591 and NPV 0.982 (assuming a prevalence of 0.30), respectively. The one marker signature achieved (1) AUC 0.545 (0.457-0.634), with a PPV 0.842 and NPV 0.567 (using the prevalence of the cohorts used for the analysis), and a PPV 0.154 and NPV 0.901 (assuming a prevalence of 0.15), and a PPV 0.307 and NPV 0.789 (assuming a prevalence of 0.30); (2) AUC 0.565 (0.456-0.673), with a PPV 0.394 and NPV 0.745 (using the prevalence of the cohorts used for the analysis), and a PPV 0.171 and NPV 0.893 (assuming a prevalence of 0.15), and a PPV 0.334 and NPV 0.775 (assuming a prevalence of 0.30); (3) AUC 0.580 (0.421-0.739), with a PPV 0.666 and NPV 0.865 (using the prevalence of the cohorts used for the analysis), and a PPV 0.159 and NPV 0.924 (assuming a prevalence of 0.15), and a PPV 0.314 and NPV 0.834 (assuming a prevalence of 0.30); (4) AUC 0.505 (0.422-0.589), with a PPV 0.857 and NPV 0.506 (using the prevalence of the cohorts used for the analysis), and a PPV 0.154 and NPV 0.906 (assuming a prevalence of 0.15), and a PPV 0.306 and a NPV 0.799 (assuming a prevalence of 0.30), respectively. The seven marker signature also discriminated patients with sepsis from patients presenting with SIRS with an AUC 0.968 (0.94-0.995), with a PPV 0.679 and NPV 0.981 (assuming a prevalence (for sepsis in a SIRS population] of 0.15). Thus also evidencing the ability to especially identify patients without sepsis, or unlikely to develop sepsis (and thus a rule-out methodology), as a result of the high NPV. An investigation was also undertaken to establish which housekeeping nucleic acid (OTULIN and RANBP3) provided for the best results, AUC, PPV and NPV, through housekeeping nucleic acid normalisation of the data, and whether one or a combination of two housekeeping nucleic acids provided for a better outcome. The conclusion was that in most circumstances RANBP3 alone provided for the best performance, with a slight increase in performance over the combination of RANBP3 and OTULIN, with OTULIN alone providing a slightly decreased performance compared to both RANBP3 and OTULIN. For example, the seven marker signature for discriminating patients with an infection (80 samples) from patients presenting with SIRS (87 samples) using both housekeeping nucleic acids for normalisation, provided an AUC of 0.974, whereas with RANBP3 alone an AUC of 0.984 was achieved, and with OTULIN alone an AUC of 0.960. Further studies using a comparable approach to that detailed above (scale inversion, housekeeping nucleic acid normalisation (RPLPO), SDA Model, and ROC analysis), with the three biomarker signature consisting of QSOX1, RRBP1 and SNORA53 were also undertaken to consider the ability of the signature to differentiate between patients with an infection from patients presenting with SIRS [172 samples in total: 88 SIRS; 84 Infection], The sample data set was chosen to have approximately equal numbers of SIRS and Infection samples with similar distributions of C-reactive protein (CRP) in each set. In the clinical setting it is believed that the prevalence of infection is lower, possibly 20-30% of samples. The prevalence does not affect the calculated sensitivity and specificity, but it does affect the positive and negative predictive values (PPV and NPV) of the test. The PPV and NPV were calculated at various prevalences, to examine the effect. Random sampling was used to create a training subset of 75% of the data and a testing set of 25% of the data. The model was fitted with the training data and evaluated with the test data. This was repeated 1000 times with different subsets and the results aggregated. Algorithm Generation The three signature gene CT values were first normalised individually by subtracting the housekeeping gene CT value. The labels SIRS and INF (infection) were assigned so that INF was the positive label. Hence, a positive score from the multiplex model indicates a high risk of infection. Then, a Linear Discriminant Model (LDA) was fitted to the training data, with shrinkage applied. Shrinkage is a mathematical procedure that can improve the quality of the fit in certain circumstances. The coefficients mo, mR, ms and the offset b are model parameter values determined from the fitted LDA algorithm. The algorithm is a linear combination of the normalised CT values for each gene and an offset, yielding a continuous model score (s). For each sample, the continuous model score is determined as: s - mQ(CTQ - CTn) + mR(CTR - CTn) + ms(CTs - CTn) + b The test score = 1 if S >0; and 0 otherwise. Where s is the continuous algorithm score, parameters mQ, mR, mS are the coefficients associated with the QSOX1, RRBP1 and SNORA53 genes respectively, b is the offset, CTq, CTr, CTs and CTN are the CT values measured for the QSOX1, RRBP1 and SNORA53 and RPLPO genes for each sample respectively. A test score of 1 indicates a high risk of infection and 0 indicates a low risk of infection. Algorithm robustness Further calculations were made to show that the model was robust and not subject to overfitting. Overfitting means that the model becomes very accurate at predicting its own training data but is not able to generalise to new unseen data. It is a particular concern when the number of features (genes) is high relative to the number of observations (samples). In this study, sufficient samples (n=172) were taken relative to the number of signature genes (p=3) to make this less of a risk. To check this and generate model performance estimates, 1000 times 4-fold cross validation was performed as follows. Random sampling was used to create a training subset of 75% of the data and a testing set of 25% of the data. The model was fitted with the training data and evaluated with the test data. This was repeated 1000 times with different subsets and the results aggregated. A robust model should show good performance when predicting the (previously unseen) test data and furthermore the generated parameters should be consistent between repeats. A further test was to determine a null distribution for learning in the data, by randomly shuffling the labels assigned to each sample (label permutation preserved proportion). The model should not be able to learn from this dataset to the same extent as with the true clinical labels. Label permutation will retain a degree of correlation with the true clinical labels, so some degree of learning may be possible where correlation is high by chance. The distribution of learning, as measured by AUC, should be normally distributed about AUC = 0.50. Any deviation from this mean value suggests some bias in the learning process or in the data itself. If it were still able to predict the random labels with high performance, this would indicate some hidden factor or unobserved structure in the data, not infection status, was driving the results. Results The model developed using SDA considers delta Cq (cycle threshold) for each target nucleic acid 5 (biomarker Cq minus RPLPO Cq), and has the following weights (mq; mR; ms) and offset (b): QSOX1: 2.04; RRBP1: -3.36; SNORA53: -1.58; Offset: -1.11. Testing performance is as detailed in Table 7, assuming a prevalence [for infection in a SIRS population] of either 20% (0.2), 30% (0.3), 40% (0.4), or 50% (0.5) Table 7. Results from the AUC analysis showing performance metrics for the three marker signature. 10 Abbreviations: PPV = positive predictive value, NPV = negative predictive value. Prevalence AUC Sensitivity Specificity PPV NPV 20% 0.98 (0.93-1) 0.97 (0.85-1) 0.93 (0.82-1) 0.79 (0.56-1) 0.99 (0.96-1) 30% 0.98 (0.93-1) 0.97 (0.85-1) 0.93 (0.82-1) 0.86 (0.70-1) 0.98 (0.94-1) 40% 0.98 (0.93-1) 0.97 (0.85-1) 0.93 (0.82-1) 0.91 (0.78-1) 0.98 (0.90-1) 50% 0.98 (0.93-1) 0.97 (0.85-1) 0.93 (0.82-1) 0.93 (0.84-1) 0.96 (0.87-1) An analysis was also undertaken with reduced signature variants of the three biomarker signature, for QSOX1 / RRBP1, QSOX1 / SNORA53, RRBP1 / SNORA53, and each biomarker alone. The SDA models built using these reduced signatures had the coefficients and performances as detailed in Table 8. 15 Table 8. Coefficients and performances for the reduced signatures. QSOX1 weight RRBP1 weight SNORA53 weight Offset / constant AUC Sensitivity Specificity NPV PPV QR 2.16 -3.95 7.89 0.95 0.94 0.89 0.94 0.90 QS 0.17 -1.92 -9.47 0.93 0.88 0.89 0.89 0.89 RS -1.52 -1.60 -4.21 0.95 0.91 0.91 0.91 0.91 Q -0.23 0.10 0.56 0.55 0.68 0.65 0.69 R -1.99 4.63 0.88 0.80 0.87 0.83 0.88 S -1.88 -9.15 0.93 0.88 0.89 0.89 0.90 As can be seen, SNORA53 alone achieved an impressive AUC of 0.93, and RRBP1 alone an AUC of 0.88, whereas the three different combinations of 2 markers all achieved AUCs of between 0.93 to 0.95, further supporting the overall performance of these three biomarkers either alone or in combination 5 to differentiate patients with an infection from patients having SIRS. These results are highly compelling for a signature comprising between two and seven markers that can be used to differentiate not only patients with sepsis from controls, but also patients with infection or sepsis from patients displaying SIRS, with AUCs between 0.926 and 0.994. However, perhaps even more compelling are the negative predictive values obtained, of between 0.910 and 1.0 (using the 10 prevalence of the cohorts used for the analysis), thus indicating the value of these biomarker signatures in a 'rule-out test', for either ruling out the possibility of a patient having an infection or sepsis. The aim of this program of work was to develop at least one predictive panel of biomarkers for infection and sepsis, and indeed absence of infection and / or sepsis, through comprehensive analysis of the host transcriptome, sourced from blood samples from human patients, and to develop biomarker signatures that may indicate whether and when clinical symptoms will arise. In so doing it would yield a suitably powered bioinformatic model for identifying and differentiating patients developing infection and / or sepsis based on transcriptomic biomarker signatures. In turn, this will assist in the development of (RT-PCR) methods for infection and sepsis diagnosis and / or prediction, where this capability should provide timely diagnosis and treatment when medical countermeasures are most effective. This work has demonstrated that the biomarker signatures of the present invention can be used at any stage of the trajectory of disease. The work reported here is unique since it has down-selected nucleic acids capable of discriminating between infection, sepsis and other patient cohorts, and proved that it is possible to identify clinically useful host biomarker signatures (with low numbers of markers) to diagnose and / or predict infection and sepsis, including at an early stage of disease. Successful down-selection of nucleic acids to manageable numbers is vital for transition to a platform. Consequently, a machine learning algorithm approach has been used to select appropriate targets and classify patients based on host nucleic acid expression with output compared to clinical diagnosis to determine predictive accuracy. The success of this approach is evidenced by high AUC and NPV values and small biomarker signatures, when comparing sepsis, infection and comparator patients.

Claims

1. A method for the risk prediction or detection / diagnosis, or absence, of sepsis and / or infection in a subject, comprising the steps of:a) providing at least one biological sample from said subject;b) detecting the presence or level of expression of at least two nucleic acid markers selected from a biomarker signature consisting of DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, SNORA53, BEND2, CR1, EIF4G3, LARP1, LGALS2, MIAT, QSOX1, RPL13A, RRBP1, SGSH, SLC36A1, TDRD9, TEX2, and TRIP13, in said sample, compared to a control, or a baseline level of expression, or expression levels with respect to other nucleic acid markers in said biomarker signature; andc) concluding, from said presence or level of expression of said at least two nucleic acid markers, the risk of the development, or detection / diagnosis, or absence of sepsis and / or infection, and classifying the subject accordingly;wherein the at least two nucleic acid markers comprises RRBP1 and SNORA53, RRBP1 and QSOX1, or SNORA53 and QSOX1.

2. A method according to Claim 1, wherein the at least two nucleic acid markers comprises RRBP1, QSOX1, and SNORA53.

3. A method according to Claim 1 or Claim 2, wherein the level of expression of the at least two nucleic acids is determined using a targeted assay to specifically measure the level of expression.

4. A method according to Claims 1 to 3, wherein the level of expression of each marker is compared to the level of expression of at least one control nucleic acid, wherein the at least one control nucleic acid is RPLPO.

5. A method according to Claim 3, wherein the targeted assay is a qPCR-based assay.

6. A method according to Claims 1 to 5, wherein the method further comprises identifying a treatment for the subject based on the classification of the subject.

7. A kit for diagnosing and / or predicting development, or absence, of infection and / or sepsis in a subject, said kit comprising reagents and / or systems for determining levels of at least two nucleic acid markers in a biological sample from a subject, wherein the at least two nucleic acid markers are selected from a biomarker signature consisting of DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, SNORA53 BEND2, CR1, EIF4G3, LARP1, LGALS2, MIAT, QSOX1, RPL13A, RRBP1, SGSH, SLC36A1, TDRD9, TEX2, and TRIP13, and optionally reagents and / or systems for determining levels of at least one control nucleic acid or housekeeping nucleic acid; wherein the at least two nucleic acid markers comprises RRBP1 and SNORA53, RRBP1 and QSOX1, or SNORA53 and QSOX1, and the reagents and / or systems for determining levels of at least two nucleic acid markers are reagents and / or systems specific to determining levels of at least RRBP1 and SNORA53, RRBP1 and QSOX1, or SNORA53 and QSOX18. A kit according to Claim 7, comprising reagents and / or systems specific to determining levels of RRBP1, QSOX1, and SNORA53 in a biological sample from a subject.

9. A kit according to Claims 7 to 8, comprising reagents and / or systems for determining levels of at least one control nucleic acid or housekeeping nucleic acid, wherein the at least one control nucleic acid or housekeeping nucleic acid is RPLPO, and the reagents and / or systems for determining levels of at least one control nucleic acid or housekeeping nucleic acid are reagents and / or system specific to determining levels of RPLPO.

10. Use of nucleic acid markers selected from a list consisting of:DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, SNORA53, BEND2, CR1, EIF4G3, LARP1, LGALS2, MIAT, QSOX1, RPL13A, RRBP1, SGSH, SLC36A1, TDRD9, TEX2, and TRIP13; for predicting presence,development or absence of infection and / or sepsis in a subject, wherein the nucleic acid markers at least comprise RRBP1 and SNORA53, RRBP1 and QSOX1, or SNORA53 and QSOX111. Use according to Claim 16, wherein the nucleic markers consist of RRBP1, QSOX1, and SNORA53.

12. A computer implemented method for the risk prediction or detection / diagnosis, or absence, of sepsis and / or infection in a subject, comprising the steps of:a) inputting the level of expression of at least two nucleic acid markers selected from a biomarker signature consisting of DDX11L9, TIFA, B4GALT5, PTGES3, TPM3, CDC42, SNORA53, BEND2, CR1, EIF4G3, LARP1, LGALS2, MIAT, QSOX1, RPL13A, RRBP1, SGSH, SLC36A1, TDRD9, TEX2, and TRIP13, from a sample obtained from the subject; andb) concluding from the level of expression of said at least two nucleic acid markers, the risk of the development, or detection / diagnosis, or absence, of sepsis and / or infection, and classifying the subject accordingly;wherein the at least two nucleic acid markers comprises RRBP1 and SNORA53, RRBP1 and QSOX1, or SNORA53 and QSOX113. A computer implemented method according to Claim 12, wherein the at least two nucleic acid markers consists of RRBP1, QSOX1, and SNORA53.

14. A computer implemented method according to Claims 12 or Claim 13, wherein the computer-implemented method utilises an algorithm based on Shrinkage Discriminant Analysis to evaluate the level of expression of the at least two nucleic acid markers, to conclude the risk of the development, or detection / diagnosis, or absence, of sepsis and / or infection, and classify the subject accordingly.

15. A system comprising a processor and one or more computer readable media storing instructions that, when executed by the processor, cause the processor to implement the method of any one of claims 12 to 14.

16. A computer readable medium or media storing instructions that, when executed by a processor, cause the processor to implement the method of any of Claims 12 to 14.

17. A computer programme product comprising instructions that, when executed by a processor, cause the processor to implement the method of any of Claims 12 to 14.

Citation Information

Patent Citations

  • Apparatus, kits and methods for predicting the development of sepsis

    GB2601222A

  • Sepsis biomarker panels and methods of use

    US20210388443A1

  • Systems and methods for targeting covid-19 therapies

    WO2023091587A1

  • Methods for determining menstrual cycle time point

    WO2023245243A1