Disease prognosis using RNA velocity

RNA velocity analysis through the VeloCD algorithm addresses the limitations of static transcriptomic studies by predicting future disease states and treatment efficacy using unspliced to spliced mRNA transcript ratios, enhancing personalized medicine and resource allocation.

WO2026038025A1PCT designated stage Publication Date: 2026-02-19IMPERIAL COLLEGE INNVOATIONS LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2025/051778
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-12
Filing Date
2025-08-12
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Existing transcriptomic studies provide a limited, static view of disease states at the time of sampling, failing to predict future disease progression or resolution, which is crucial for personalized medicine and resource allocation.

Method used

The method utilizes RNA velocity analysis by calculating the ratio of unspliced to spliced mRNA transcripts to predict future gene expression levels, employing the VeloCD algorithm for prognostic predictions based on whole-blood RNA-Seq data, incorporating dimensionality reduction and transition probability calculations.

Benefits of technology

Enables accurate prediction of disease outcomes and treatment efficacy from a single patient sample, providing insights into future disease states and clinical trajectories without the need for sequential sampling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025051778_19022026_PF_FP_ABST
    Figure GB2025051778_19022026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to methods for predicting a subject's condition and / or providing a prognosis of a subject's condition, and methods of determining the efficacy of treating a subject suffering from a condition. The invention also relates to methods of selecting a gene or genes for prognosing a condition in a subject.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Disease Prognosis

[0002] The present invention relates to methods for predicting a subject's condition and / or providing a prognosis of a subject's condition, and methods of determining the efficacy of treating a subject suffering from a condition. The invention also relates to methods of selecting a gene or genes for prognosing a condition in a subject.

[0003] Transcriptomics is the study of all the RIMA in a sample; the link between the genotype of an organism and the main functional component (proteome) responsible for its phenotypic characteristics. Therefore, transcriptomics unlocks an abundance of information about a manifold of biological processes. Whole-blood transcriptomics has been particularly useful in the study of infectious disease, and it has widely been observed in the literature that blood RNA expression is highly representative of a patient's current disease state, specifically disease cause (diagnosis) and severity levels. Consequently, many of these studies have identified new biomarkers for diagnostics, prognostics, and therapeutic targets, by looking at genes and biological pathways involved in pathophysiological processes implicated in disease.

[0004] Traditionally, these studies use the spliced transcript component of the RNA pool (PolyA selected transcriptomic datasets), or only examine expression at the gene-level across transcripts. Spliced transcripts are RNA transcripts that have been fully processed by the RNA splicing machinery and do not contain introns; they represent the current transcriptomic state of the organism at the time of sampling, because they are the component translated to generate proteins.

[0005] More recently, numerous studies have revealed the potential utility of separating out and independently examining the other component of the RNA pool, namely unspliced transcripts, which are pre-processed, immature RNA that have not yet been spliced, i.e., they contain introns. Single-exon genes do not contain intronic sequences, whereas multi-exon genes do contain intronic sequences and their transcripts undergo splicing. Numerous studies have illustrated that unspliced transcripts, quantified using intronmapping reads, follow biological patterns that represent future yet-to-be-spliced transcript expression.

[0006] These observations acted as the precursor for the concept of RNA velocity, where the relative abundance of spliced and unspliced transcripts are used to predict the future transcriptomic states of single cells undergoing differentiation. Since the publication of this method, many alterations and extensions have been published illustrating much interest in its concept, which has elucidated future temporal dynamics of gene expression for numerous biological systems.

[0007] RIMA velocity overcomes the main limitation of many transcriptomic studies, which is that they only examine a fixed view of the biological process of the time of sampling, whilst RNA velocity allows for extrapolation of future transcriptomic states. Indeed, most prognostic studies only provide a limited, static view of the blood transcriptome at the time of sampling. This is highly representative of current disease states, but often reveals very little about future disease progression or resolution. The ability to predict which patients and / or subjects are likely to recover or deteriorate, within clinically relevant timescales, and without the need for time-consuming sequential sampling, could be an invaluable tool in clinical settings for personalised medicine, patient triage, and resource allocation.

[0008] There is, therefore, a need for a novel platform that may be used as a prognostic tool, which, utilising RNA velocity data, can predict future disease states of individual patients from just a single patient sample.

[0009] Thus, according to a first aspect of the invention, there is provided a method for predicting a subject's condition and / or providing a prognosis of a subject's condition, the method comprising: i) performing RNA analysis on a subject; ii) calculating an RNA velocity value based on the ratio of unspliced mRNA transcripts to spliced mRNA transcripts, for at least one gene in the sample; and iii) using the RNA velocity value to predict future expression levels of the at least one gene, to thereby provide a prediction and / or prognosis of the subject's condition.

[0010] In a second aspect of the invention, there is provided a method for determining the efficacy of treating a subject suffering from a condition with a therapeutic agent or a specialised diet, the method comprising: i) performing RNA analysis on a subject; ii) calculating an RNA velocity value based on the ratio of unspliced mRNA transcripts to spliced mRNA transcripts, for at least one gene in the sample; and iii) using the RNA velocity value to predict future expression levels of the at least one gene, to thereby determine whether treatment with the therapeutic agent or the specialised diet is effective or ineffective.

[0011] As described in the Examples, the inventors have adapted the original RNA velocity algorithm, called velocyto, into a new bioinformatic approach to overcome its previous limitations. The inventors have surprisingly demonstrated that the RNA velocity algorithm can be used as a prognostic tool, here called VeloCD, "RNA Velocity-based transition predictions of future disease states of subjects in Clinical Datasets", to predict disease outcomes from host RNA-Seq data.

[0012] VeloCD performs prognostic predictions under several core assumptions, and these core assumptions are as follows: (i) current transcriptomic, i.e., spliced transcript, expression represents current disease states; and (ii) RNA velocities can capture future gene expression which corresponds to future disease states.

[0013] The inventors of the present disclosure have demonstrated how these assumptions are met in the whole-blood RNA-Seq datasets analysed to date and have illustrated a proof of concept of the prognostic capabilities of this tool across a range of diseases.

[0014] The inventors also believe that the methods of the invention can be used to treat a subject suffering from a condition, by calculating the RNA velocity of at least one gene in the subject's sample and prognosing their condition. For example, once a subject has been prognosed with a condition, a therapeutic agent may be administered, depending on the disease in question.

[0015] According to a third aspect of the invention, there is provided a method of treating a subject suffering from a condition, the method comprising: i) performing RNA analysis on a subject; ii) calculating an RNA velocity value based on the ratio of unspliced mRNA transcripts to spliced mRNA transcripts, for at least one gene in the sample; iii) using the RNA velocity value to predict future expression levels of the at least one gene, to thereby provide a prediction and / or prognosis of the subject's condition; and i) administering, or having administered, to the subject, a therapeutic agent, which reduces, delays or prevents progression of the condition. The RNA analysis may comprise the use of an RIMA array or RNA sequencing, for example RNA-seq analysis generated in short-read (e.g. Illumina) or long-read (e.g. Oxford Nanopore, Pacific Biosciences) sequencing technologies. The RNA analysis may be performed using RNA-Seq, microarray, qPCR, LAMP, or technologies using CRISPR-Cas. It will be appreciated that RNA analysis may be a sequencing technique that typically uses next-generation sequencing (NGS) to reveal the presence and quantity of RNA in a biological sample, representing an aggregated snapshot of the cells' dynamic pool of RNAs, i.e. the transcriptome. The RNA analysis may also be based on the detection of specific sequences by using complementary probes, which can be done at scale by multiplexing. In some embodiments, the RNA analysis is configured to analyse spliced and unspliced transcripts.

[0016] In some embodiments, the RNA analysis may be performed using RNA-seq. In some embodiments, single-cell RNA-seq may be used.

[0017] It is known to the skilled person that an RNA velocity calculation is based on the relative abundance of unspliced RNA and spliced RNA, to estimate the future rate of RNA splicing and degradation, i.e. infer the directionality of transcription events. RNA velocity calculations are described in more detail here: La Manno, G., Soldatov, R., Zeisel, A., Braun, E., Hochgerner, H., Petukhov, V., Lidschreiber, K., Kastriti, M. E., Lbnnerberg, P.

[0018] & Furlan, A. (2018) RNA velocity of single cells. Nature. 560 (7719), 494-498, the contents of which are incorporated herein by reference. Accordingly, in some embodiments, calculating the RNA velocity will be understood to mean predicting the future expression of the at least one gene.

[0019] The RNA sequencing data may be pre-processed to remove low-quality cells and genes, normalise gene expression values and handle any technical variations. Pre-processing may also comprise removing genes that do not have spliced and unspliced transcripts.

[0020] Thus, in some embodiments, the method comprises the step of pre-processing the RNA sequencing data. Pre-processing the RNA sequencing data may comprise removing low quality cells and / or genes, normalising gene expression values, and / or handling any technical variations. In some embodiments, pre-processing the RNA sequencing data may comprise removing genes that do not have spliced and unspliced transcripts.

[0021] In order to calculate the ratio of spliced to unspliced mRNA transcripts, the expression of spliced mRNA transcripts and unspliced mRNA transcripts is calculated. Thus, in some embodiments, the method comprises the step of calculating the expression of spliced mRNA transcripts and unspliced mRNA transcripts for at least one gene in the sample.

[0022] The RNA velocity tool (i.e. VeloCD) constructs low dimensional representations of the spliced transcript expression, i.e., the current transcriptomic state, of genes in the samples in a dataset. These genes are pre-selected as those with both spliced and unspliced transcript expression, i.e., multi-exon genes. Principal Component Analysis (PCA), t-distributed stochastic neighbor embedding (tSNE), and Uniform Manifold Approximation and Projection (UMAP) are the three implemented options for performing this step.

[0023] Accordingly, in one embodiment, the method comprises the step of transforming the spliced transcript expression into low-dimensional transcriptomic space. In one embodiment, Principal Component Analysis (PCA), t-distributed stochastic neighbor embedding (tSNE), or Uniform Manifold Approximation and Projection (UMAP), are used to transform the spliced transcript expression into low-dimensional transcriptomic space.

[0024] The VeloCD algorithm, and its predecessor (velocyto), work under the null hypothesis that no gene expression change is occurring, i.e., gene expression is at a steady state. RNA velocities are then calculated as a deviation from this assumption. Accordingly, in one embodiment, calculating the RNA velocity value comprises subtracting the unspliced transcript expression estimated under a steady state from the actual unspliced transcript expression.

[0025] After performing dimensionality reduction, VeloCD calculates gene and sample specific RNA velocity values. First, a weighted gamma model, also called a weighted least square fit equation, is fitted to each gene in the dataset: y = w(mx + c)2

[0026] Estimated steady state unspliced transcript expression = weights slope x spliced + c)2

[0027] Accordingly, in one embodiment, the method comprises fitting a weighted gamma model to the at least one gene. The weighted gamma model may comprise the following equation: y = w(mx + c)2

[0028] Estimated steady state unspliced transcript expression = weights slope x spliced + c)2 The offset is represented by c in these equations. Accordingly, in one embodiment, the method may comprise calculating the offset. This is a rearrangement of a deterministic continuous-time rate equation that states that the change in the spliced transcript abundance is the rate of splicing multiplied by the unspliced abundance minus the degradation rate multiplied by the spliced transcript expression. This is based on the central dogma of molecular biology - that spliced transcript expression proceeds unspliced.

[0029] In this model, P(t) is assumed to be 1: ds / dt = / ?(t)iz(t) — y(t)s(t)

[0030] P(t) = splicing rate u(t) = unspliced transcript abundance s(t) = spliced transcript abundance y(t) = degradation rate of spliced mRNA molecules

[0031] These equations were taken from the first supplementary file of the original RIMA velocity study (La Manno et al., 2018). The parameters of this model, slope (degradation rate) and offset, are estimated using each gene's unspliced and spliced transcript expression value in the samples assumed to be under a steady state. These samples' spliced transcript expression values are in the top and bottom 99thpercentile of expression values.

[0032] The offset is an optional gene-specific parameter used to adjust for potential alignment errors and intronic sequences that originate from unannotated transcripts. A steady state is defined as the point where the behaviour of a system does not change even when a process is still occurring. In the case of gene expression, this state is reached when the ratio of spliced to unspliced transcript expression is constant and the gene expression is not changing. Under this steady state, y(t) is estimated as the ratio of unspliced to spliced transcript expression of the samples with the top and bottom 99th percentile of expression values: y(t) = unspliced / spliced

[0033] Accordingly, in one embodiment, the method comprises estimating the degradation rate of spliced mRNA molecules, using the following equation: y(t) = unspliced / spliced.

[0034] Using this parameter, the algorithm then predicts the unspliced transcript expression using the following equation: predicted unspliced = (degradation rate x spliced transcript abundance) + offset

[0035] Accordingly, in one embodiment, the method comprises predicting the unspliced transcript expression using the following equation: predicted unspliced = (degradation rate x spliced transcript abundance) + offset

[0036] The predicted unspliced transcript abundance is then used to calculate the RNA velocity values:

[0037] RNA velocity = unspliced transcript abundance - predicted unspliced transcript abundance

[0038] Accordingly, in one embodiment, the method comprises calculating the RNA velocity value using the following equation:

[0039] RNA velocity = unspliced transcript abundance - predicted unspliced transcript abundance

[0040] Hence non-zero RNA velocity values are a deviation from the gene being in a steady state because they indicate that the actual unspliced transcript abundance is different from that predicted under this state. Positive values indicate that the observed unspliced transcript abundance is higher than the predicted and there is gene expression upregulation. Negative values indicates that the observed unspliced transcript abundance is lower than the predicted and there is gene expression downregulation.

[0041] Accordingly, in one embodiment, a positive RNA velocity value indicates an increase in future expression levels for the at least one gene in the sample. Alternatively, a negative velocity value indicates a decrease in future expression levels for the at least one gene in the sample. In some embodiments, a positive RNA velocity value indicates an increase in future expression levels compared to the modelled gene expression, for the at least one gene in the sample. Alternatively, a negative velocity value indicates a decrease in future expression levels compared to the modelled gene expression, for the at least one gene in the sample.

[0042] VeloCD then projects the RNA velocity values onto low-dimensional transcriptomic space (so-called RNA velocity fate maps), generated using the tSNE, 2-dimensional UMAP and 3-dimensional UMAP embedding methods separately. Accordingly, in one embodiment, the method comprises embedding the RNA velocity value on the low-dimensional transcriptomic space. For each sample, VeloCD calculates the correlation between the RIMA velocity values of each gene and the difference between it and every other sample's spliced transcript expression. Accordingly, in one embodiment, the method comprises calculating the correlation between the RNA velocity values of the at least one gene and the difference between it and every other sample's spliced transcript expression. These values are firstly calculated for each gene (gene-level correlation values) and then summed across genes.

[0043] The correlation values are then used to calculate the probability of each sample transitioning to each other sample, which are the sample-to-sample transition probability values. This is performed for each sample row to give the probability of transition of each row to each sample column. Accordingly, in one embodiment, the method comprises calculating the sample-to-sample transition probability.

[0044] These are calculated using the following equation:

[0045] Sample to sample transition probabilities = Z e correlation coefficient scaling factor ) - i / n

[0046] Firstly, the correlation values are scaled:

[0047] To control for the density of samples, the algorithm then completes a nearest neighbour search of the low-dimensional transcriptomic space, to identify the nearest neighbours of each sample. Neighbour samples are selected based on their Euclidean distance from the sample of interest, with closer samples selected over those that are further away. These relationships are then returned as a matrix consisting of Os and Is. Samples whose columns were assigned 1 indicates that they were selected as one of that sample's nearest neighbours. Samples whose columns were assigned 0 were not selected. The number of samples assigned as 1 for each row is defined by n.

[0048] For each row, the correlation values of the columns assigned 0 are converted to 0. The correlation values for the samples whose columns were assigned 1 are then adjusted ensuring that they sum to 1 for each row, splitting 1 unequally across n number of samples (columns) based on these scaled values. These adjusted values are the sample- to-sample transition probabilities. The percentage probability of the transition of each sample to each of its nearest neighbours is its transition probability multiplied by 100. The probability of transitioning from self to self is always zero.

[0049] The value of n used to control for the density of points is referred to as the "transition probability number of neighbours value" (TPNN) hyperparameter to avoid confusion with the term used to specify the embedding number of neighbours value for dimensionality reduction. The sample-to-sample transition probability values embed the RIMA velocities onto low dimensional transcriptomic space.

[0050] In short, the x, y and z coordinates of each sample are extracted from the tSNE and UMAP embeddings. The differences between the embedding coordinates of each sample and each other sample are used to form each axis-specific vector of the unitary matrix. The transition probabilities are then multiplied by each vector of this unitary matrix. The resulting values (u, v, w) are then added to the low dimensional embedding coordinates of each sample (x, y, z), giving the coordinates of the heads of the RNA velocity arrows. These are then visualised on RNA velocity fate maps, which are low-dimensional representations of transcriptomic space.

[0051] One of the most important distinctions between VeloCD and velocyto is that VeloCD also uses the sample-to-sample transition probability values to calculate sample-to-group transition probabilities. Accordingly, in one embodiment, the method comprises calculating the sample-to-group transition probabilities.

[0052] The sample-to-sample transition probability values are extracted, and prespecified group classifications are used divide these values by group, which are then summed. This can be illustrated with a theoretical example. For example, the transition probability of sample A to sample B is 0.6, from A to C is 0.3, A to D is 0.05 and A to E is 0.05. B and C belong to group X and D and E to group Y. The probability of sample A transitioning to group X is: 0.6 (A to B) + 0.3 (A to C) = 0.9. A belongs to group Y. Therefore, probability of sample A remaining in group Y is: 0.05 (A to D) + 0.05 (A to E) + 0 (A to A) = 0.1.

[0053] As currently implemented, A needs to have some sort of group classification for RNA velocity analysis. This can be its actual group or sample A can be specified as its own group of n=l for convenience. The group with the highest probability or whose value is above a certain threshold is then chosen as the one to which the sample is most likely to transition. These values are then used for downstream analysis to calculate the number of samples predicted to transition to group X or Y. These groups can represent the different "future states" of the subjects, including disease outcomes: death or survival.

[0054] Once the RNA velocity value has been calculated, it can be used to provide a prognosis of the subject's condition, by predicting the clinical trajectory with reference to known patterns of gene activity in mild or severe disease. The RIMA velocity value may also be used in the process of calculating the probability of a certain disease outcome.

[0055] In the context of a prognostic tool, VeloCD is run using a "drop-one in" approach; it requires sets of samples belonging to two or more "reference groups", with clearly defined current disease states, i.e., at the time of sampling, representative of the possible disease outcomes the user would like to predict i.e., death or survival. For optimal predictions, this algorithm also requires a pre-selected set of highly predictive genes, but it can also run using all of the expressed multi-exon genes in a dataset.

[0056] For prediction, the user then drops one "test" sample into the algorithm, and its RNA velocity values are used to calculate the probability of transition to each reference group outcome. This concept is visualised by the cartoon fate map shown in Figure 1. Results are then examined across a range of TPNN hyperparameter values, i.e., analysis runs, where transition probabilities are collated across runs for each test subject. In some embodiments, the TPNN is set to a single value for each dataset, or is an average (median) across all possible TPNN values.

[0057] Accordingly, in one embodiment, the method comprises comparing the predicted future expression levels of the at least one gene, with expression levels of the gene in at least one reference sample. In another embodiment, the method comprises comparing the predicted future expression levels of the at least one gene, with expression levels of the gene in at least two reference samples.

[0058] The reference sample(s) may be a sample taken from a positive control and / or a negative control. For example, a negative control sample may be a sample taken from a subject who does not have the condition that the subject is suffering from. A positive control sample may be a sample taken from a subject who does have the condition that the subject is suffering from.

[0059] VeloCD uses the velocity values in the context of the samples in the low-dimensional transcriptomic space to calculate probabilities of transition for each sample to each predefined sample group. These give the probability that a subject will transition into each of the at least one reference groups (together summing to 1). For example, the reference groups may be health deterioration (going to PICU) or recovery (going home).

[0060] VeloCD has two main hyperparameters; the first is called the transition probability number of neighbours (TPNN). The TPNN is the number of close samples the algorithm takes into consideration when calculating transition probabilities, and the second is the number of neighbours (NN, also called perplexity in the literature) used by tSNE or UMAP to generate the low-dimensional embeddings. PCA does not use this hyperparameter.

[0061] NN: This hyperparameter is used to generate the tSNE and UMAP embeddings (not required by PCA). It is the number of each sample's neighbouring samples that are considered in high-dimensional space to recreate the relationships between samples in low-dimensional space. I.e. if NN = 5 then each sample's five closest samples will be used to help recreate these relationships in low-dimensional space. The higher the number the more global structure of the data is preserved, i.e., every sample's spatial relationship to each other is considered.

[0062] TPNN: Using the position of the samples in low-dimensional transcriptomic space and to control for the density of samples, the algorithm completes a nearest neighbour search of the low-dimensional transcriptomic space to identify the n nearest neighbours of each sample. Neighbour samples are selected based on their Euclidean distance from the sample of interest, with closer samples selected over those that are further away. These relationships are then returned as a matrix consisting of Os and Is. Samples whose columns were assigned 1 indicates that they were selected as one of that samples nearest neighbours. Samples whose columns were assigned 0 were not selected. The number of samples assigned as 1 for each row is defined by n - defined by the user using the TPNN hyperparameter. Therefore, TPNN defines the number of close neighbouring samples taken into consideration when calculating the transition probability number of neighbour values.

[0063] Prognosis may relate to determining the therapeutic outcome in a subject that has been diagnosed with a condition. Prognosis may relate to predicting the rate of progression or improvement and / or the duration of the condition in a subject, the probability of survival, and / or the efficacy of various treatment regimes. Thus, a poor prognosis may be indicative of disease progression, low probability of survival and reduced efficacy of a treatment regime. A favourable prognosis may be indicative of disease resolution, high probability of survival and increased efficacy of a treatment regime.

[0064] Providing a prognosis of the subject's condition may comprise predicting the probability of the subject transitioning to a user-defined group, such as binary disease outcomes (for example, death or survival) or different states (for example, asymptomatic, uncomplicated or severe disease). Accordingly, in one embodiment, providing a prognosis of the subject's condition comprises predicting their future disease state. For example, their disease state may be asymptomatic, uncomplicated or severe disease. Alternatively, in another embodiment, the disease state may be developing or not developing a particular condition (e.g. an infection), or a particular complication of an infection. Alternatively, in another embodiment, providing a prognosis of the subject's condition comprises predicting trajectories of illness severity. The illness severity may be represented by the requirement for different levels of clinical care. For example, mild severity groups (sent home), moderate severity groups (admitted to a hospital ward), or severe / very severe groups (sent to Intensive Care Unit). Alternatively, in another embodiment, providing a prognosis of the subject's condition may comprise providing a binary disease outcome, i.e. death or survival.

[0065] As used herein, "predicting a subject's condition" may comprise predicting whether a disease will occur at all, following an exposure that puts someone at risk of the disease. In other words, "predicting a subject's condition" may comprise predicting whether a subject will test positive for a disease before they become positive.

[0066] As used herein, the phrase "future expression levels" will be understood to mean the expression levels of the at least one gene at a time point after the RIMA analysis has been performed. In one embodiment, "future expression levels" may be expression levels of the at least one gene up to one, two, three, four or five hours after the RNA analysis has been performed. In another embodiment, "future expression levels" may be expression levels of the at least one gene up to six, seven, eight, nine or ten hours after the RNA analysis has been performed. Alternatively, in another embodiment, "future expression levels" may be expression levels of the at least one gene up to 12, 14, 16, 18, 20, 22 or 24 hours after the RNA analysis has been performed. Alternatively, in another embodiment, "future expression levels" may be expression levels of the at least one gene up to 36, 48 or 72 hours after the RNA analysis has been performed.

[0067] Alternatively, in another embodiment, "future expression levels" may be expression levels of the at least one gene up to one week, two weeks, three weeks or four weeks after the RNA analysis has been performed. Alternatively, in another embodiment, "future expression levels" may be expression levels of the at least one gene up to two, three, four or five months after the RNA analysis has been performed.

[0068] The at least one gene may be a gene whose future transcriptomic states are representative of future disease states. Accordingly, in some embodiments, the at least one gene is a gene whose expression levels can provide a prognosis of a subject's condition. The at least one gene may be a gene whose expression across time is representative of diverging disease trajectories. For example, the at least one gene may be a gene which can provide a positive or negative prognosis for a condition.

[0069] In some embodiments, the at least one gene may be one, two, three, four or five genes. In another embodiment, the at least one gene may be six, seven, eight, nine or ten genes. In some embodiments, the at least one gene may be at least 100 genes, at least 1000 genes, at least 10,000 genes, or at least 50,000 genes.

[0070] The at least one gene may be selected from the group consisting of: IFI44L (Interferon Induced Protein 44 Like); LY6E (Lymphocyte antigen 6E); APOL3; and BAK1.

[0071] The at least one gene may be selected from the list of genes in Table 1, Table 2, Table 3, and / or Table 4.

[0072] The inventors have discovered a framework for selecting biologically informative genes that can be used to prognose a condition in a subject using RNA velocity values (see Figure 4). Accordingly, in some embodiments, the method further comprises selecting at least one gene whose future transcriptomic states are representative of future disease states.

[0073] Accordingly, in a fourth aspect of the invention, there is provided a method of selecting at least one gene for prognosing a condition in a subject using RNA velocity, the method comprising: i) conducting differential expression analysis (DEA) to identify at least one gene that can distinguish between a subject who is positive for the condition and a subject who is negative for the condition; ii) performing discordance-concordance (DISCO) analysis of the RNA velocity value to identify at least one gene whose RNA velocity value is concordant in a threshold number of samples; and iii) selecting the at least one gene that meets the criteria of i) and ii).

[0074] In some embodiments, the method further comprises removing at least one gene that is not suitable for downstream translation use (e.g. transfer onto clinical devices). For example, the method may comprise removing at least one gene with expression below a specific threshold, removing at least one non-protein coding gene, or removing a gene with a certain biotype including rRNA. This step may be performed before step i) of the method according to the fourth aspect.

[0075] In some embodiments, step i) of the method comprises multiple testing correction and / or logzFC (fold change) filtering. In some embodiments, step i) of the method comprises selecting at least one gene above the FDR p-value threshold.

[0076] In some embodiments, performing the discordance-concordance (DISCO) analysis (i.e. step ii)), comprises subtracting spliced transcript expression from unspliced transcript expression, for the at least one gene. In some embodiments, performing the discordance-concordance (DISCO) analysis comprises subtracting the spliced transcript expression of the previous day from the spiced transcript expression of the next day. Genes whose direction of either metric, the above and RNA velocity values (positive or negative) is concordant in at least a "threshold" (chosen by user) number of samples is taken forward for further analysis. The intersection of these two analyses is used as the initial signature for VeloCD.

[0077] In some embodiments, the method comprises using feature selection algorithms, such as Least Absolute Shrinkage and Selection Operator (LASSO) regression.

[0078] In some embodiments, the method comprises selecting the at least one gene whose percentage contribution to the top three Principal Components (PCs) of the PCA is greater than or equal to a threshold value.

[0079] In a fifth aspect, there is provided computer-readable instructions which, when executed by a computing apparatus, cause the computing apparatus to perform the method of the fourth aspect.

[0080] In a sixth aspect, there is provided a computer-readable medium (such as a non- transitory computer-readable medium) comprising program instructions stored thereon for performing the method of the fourth aspect.

[0081] In a seventh aspect, there is provided an apparatus comprising: at least one processor; and at least one memory including computer program code which, when executed by the at least one processor, causes the apparatus to perform the method of the fourth aspect. In an eighth aspect, there is provided an apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, causes the apparatus to perform the method of the fourth aspect.

[0082] The condition may be an infectious disease, an inflammatory disease, a metabolic disease, an endocrine disease, cancer, a degenerative disease, drug toxicity, a cardiovascular disease, or trauma (injury). The infectious disease may be a bacterial infection, a viral infection, a fungal infection, a parasitic infection, or a combination thereof.

[0083] The bacterial infection may be a Gram positive bacterial infection selected from the group consisting of: Corynebacterium diphtheriae, Clostridium botulinum, Clostridium difficile, Clostridium perfringens, Clostridium tetani, Enterococcus faecalis, Enterococcus faecium, Listeria monocytogenes, Staphylococcus aureus, Staphylococcus epidermidis, Staphylococcus saprophyticus, Group A Streptococcus, Group B streptococcus, Streptococcus agalactiae, Streptococcus pneumoniae, Streptococcus pyogenes, or acid fast bacteria such as Mycobacterium leprae, Mycobacterium tuberculosis, Mycobacterium ulcerans and Mycobacterium avium intercellularae, or a Gram negative bacterial infection selected from the group consisting of: Bordetella pertussis, Borrelia burgdorferi, Brucella abortus, Brucella canis, Brucella melitensis, Brucella suis, Campylobacter jejuni , Chlamydia pneumoniae, Chlamydia trachomatis, Chlamydophila psittaci, Escherichia coli, Francisella tularensis, Haemophilus influenzae, Helicobacter pylori, Legionella pneumophila, Leptospira interrogans, Mycoplasma pneumonia, Neisseria gonorrhoeae, Neisseria meningitidis, Pseudomonas aeruginosa, Pseudomonas spp, Rickettsia rickettsii, Salmonella typhi, Salmonella typhimurium, Shigella sonnei, Treponema pallidum, Vibrio cholerae, Yersinia pestis, Kingel I a kingae, Stenotrophomonas and Klebsiella.

[0084] The viral infection may be selected from the group consisting of: influenza such as Influenza A, including but not limited to: H1N1, H2N2, H3N2, H5N1, H7N7, H1N2, H9N2, H7N2, H7N3, H10N7, Influenza B and Influenza C, Respiratory Syncytial Virus (RSV), rhinovirus, enterovirus, bocavirus, parainfluenza, adenovirus, metapneumovirus, herpes simplex virus, Chickenpox virus, Human papillomavirus, Hepatitis, Epstein-Barr virus, Varicella-zoster virus, Human cytomegalovirus, Human herpesvirus, human herpesvirus 6 (HHV6), type 8 BK virus, JC virus, Human coronaviruses, Smallpox, Parvovirus B19, Human astrovirus, Norwalk virus, coxsackievirus, poliovirus, Severe acute respiratory syndrome coronaviruses (e.g. SARS-CoV-2), yellow fever virus, dengue virus, West Nile virus, Rubella virus, Human immunodeficiency virus, Guanarito virus, Junin virus, Lassa virus, Machupo virus, Sabia virus, Crimean-Congo haemorrhagic fever virus, Ebola virus, Marburg virus, Measles virus, Mumps virus, Rabies virus and Rotavirus.

[0085] The parasitic infection may be selected from the group including but not limited to: Malaria, African trypanosomiasis, babesiosis, Chagas disease, leishmaniasis, and toxoplasmosis.

[0086] The fungal infections may be selected from the group including but not limited to: Candida species, Aspergillus species, Histoplasmosis, Mucormycosis, and Pneumocystis ji roveci i.

[0087] The inflammatory disease may be selected from the group including but not limited to: systemic lupus erythematosus (SLE), juvenile idiopathic arthritis (JIA), Henoch-Schbnlein purpura (HSP), Kawasaki disease, rheumatoid arthritis, ulcerative colitis, Crohn's disease, chronic active hepatitis, celiac disease and vasculitis.

[0088] In one embodiment, the condition may be Tuberculosis (TB)-associated immune reconstitution inflammatory syndrome (TB-IRIS). In another embodiment, the condition may be sepsis. In another embodiment, the condition may be an acute febrile illness. In another embodiment, the condition may be an acute illness.

[0089] In an embodiment in which the condition is a bacterial infection, the method according to the third aspect may comprise administering an anti-bacterial agent, such as an antibiotic. Suitable anti-bacterial agents may be selected from the group including but not limited to: ceftobiprole, ceflaroline, clindamycin, dalbavancin, daptomycin, linezolid, oritavancin, tedizolid, telavancin, tigecycline, vancomycin, aminoglycosides, carbapenems, ceftazidime, ceftobiprole, fluoroquinolines, piperacillin / tazobactam, ticarcillin / clavulanic acid, streptogramins, such as amikacin, gentamicin, kanamycin, netilmicin, tobramycin, paromomycin, streptomycin, geldanamycin, herbimycin, rifaximin, loracarbef, ertapenem, doripenem, imipenem / cilastatin, meropenem, cefadroxil, cefazolin, cefalotin / cefalothin, cefalexin, cefaclor, k cefamandole, cefoxitin, cefprozil, cefuroxime, cefixime, cefdinir, cefditoren, cefoperazone, cefotaxime, cefpodoxime, ceftazidime, ceftibuten, ceftizoxime, ceftriaxone, cefepime, ceftaroline fosamil, ceftobiprole, teicoplanin, vancomycin, telavancin, dalbavancin, oritavancin, dalbavancin, oritavancin, clindamycin, linomycin, daptomycin, azithromycin, clarithromycin, dirithromycin, erythromycin, roxithromycin, troleandomycin, telithromycin, spiramycin, aztreonam, furazilidone, linezolid, posizolid, radezolid, torezolid, amoxicillin, ampicillin, azlocillin, carbenicillin, cioxacillin, dicloxacillin, flucloxacillin, mezlocillin, nafcillin, oxacillin, penicillin G, penicillin V, piperacillin, temocillin, ticarcillin, amoxicillin / clavulanate, ampicillin / sulbactam, piperacillin / tazobactam, bacitracin, colistin, polymyxin B, ciprofloxacin, enoxacin, gatifloxacin, gemifloxacin, levofloxacin, lomefloxacin, moxifloxacin, nalidixic acid, norfloxacin, ofloxacin, trovafloxacin, grepafloxacin, sparfloxacin, temafloxacin, mafenide, sulfacetamide, sulfadiazine, silver sulfadiazine, sulfadimethoxine, sulfasalazine, sulfisoxazole, trimethoprim-sulfamethoxazole, sulfonamidochrysoidine, demeclocycline, doxycycline, minocycline, oxytetracycline, tetracycline, clofazimine, dapsone, capremycin, cycloserine, ethambutol, ethionamide, isoniazid, pyrazinamide, rifampicin, rifapentine, streptomycin, chloramphenicol, fosfomycin, fusidic acid, metronidazole, mupirocin, platensimycin, quinupristin / dalfopristin, thiamphenicol, tigecycline, tinidazole and trimethoprim.

[0090] In an embodiment in which the condition is a viral infection, the method according to the third aspect may comprise administering an anti-viral agent, an immunomodulatory agent, and / or supportive care where the viral infection is suspected to be self-limiting. Suitable anti-viral agents may be selected from the group including but not limited to: Aciclovir, valaciclovir, ganciclovir, valganciclovir, oseltamivir, zanamivir, amantadine, rimantadine, cidofovir, brincidofovir, famciclovir, foscarnet, molnupiravir, nirmatrelvir, remdesivir, sotrovimab, entecavir, interferon alfa, tenofovir, adefovir, dipivoxil, lamivudine, sofosbuvir, ribavirin, and immunoglobulin.

[0091] In an embodiment in which the condition is a parasitic infection, the method according to the third aspect may comprise administering an anti-parasitic agent. Suitable anti- parasitic agents may be selected from the group including but not limited to: Praziquantel, Mebendazole, Albendazole, Metronidazole, Ivermectin, Nitazoxanide, Suramin, Amphotericin B, Pentamidine, Melarsoprol, Eflornithine, Nifurtimox, Atovaquone, Azithromycin, Clindamycin, Quinine, Doxycycline, Benznidazole, Miltefosine, Paromomycin, Sodium stibogluconate, meglumine antimoniate, Chloroquine, Artesunate, Artemether, Mefloquine, Proguanil, Primaquine, Lumefantrine, Piperaquine, Pyronaridine, Sulfadiazine, Pyrimethamine, and Spiramycin.

[0092] In an embodiment in which the condition is an inflammatory disease, the method according to the third aspect may comprise administering an agent including but not limited to: Ibuprofen, naproxen, aspirin, hydroxychloroquine, colchicine, corticosteroids, methotrexate, sulfasalazine, adalimumab, infliximab, etanercept, anakinra, canakinumab, tocilizumab, sarilumab, baricitinib, tofacitinib, imatinib, upadacitinib, filgotinib, mesalamine, rituximab, belimumab, mycophenolate mofetil, cyclophosphamide, abatacept, ustekinumab, vedolizumab, natalizumab, plasma exchange, and immunoglobulin.

[0093] The methods of the first, second or third aspect may comprise obtaining a sample from the subject. Alternatively, the methods of the first, second or third aspect may comprise performing RIMA analysis on a sample obtained from the subject.

[0094] In some embodiments, the sample comprises a biological sample. The sample may be any material that is obtainable from a subject. The biological sample may be tissue or a biological fluid. Furthermore, the sample may be blood, plasma, serum, spinal fluid, urine, sweat, saliva, tears, breast aspirate, breast milk, prostate fluid, seminal fluid, vaginal fluid, stool, cervical scraping, cytes, amniotic fluid, intraocular fluid, mucous, moisture in breath, animal tissue, cell lysates, tumour tissue, hair, skin, buccal scrapings, lymph, interstitial fluid, nails, bone marrow, cartilage, prions, bone powder, ear wax, lymph, granuloma, cerebrospinal fluid, or combinations thereof.

[0095] The sample may comprise blood, urine, tissue etc. In one embodiment, the sample comprises a blood sample. In one embodiment, the sample comprises a whole-blood sample. In one embodiment, the whole-blood sample comprises RNA from multiple cell types.

[0096] Advantageously, whole blood samples allow broad screening for a range of different disorders. The inventors believe that they are the first to have performed RNA velocity analysis on whole blood samples, for prognosing a condition in a subject.

[0097] Accordingly, in a ninth aspect of the invention, there is provided a method of predicting a subject's condition and / or providing a subject's prognosis, the method comprising performing RNA velocity analysis on whole blood of a subject.

[0098] In one embodiment, the method according to the ninth aspect comprises performing RNA velocity on a whole blood sample from the subject. In one embodiment, the method according to the ninth aspect comprises performing RNA velocity on a single whole blood sample from the subject.

[0099] The blood may be venous or arterial blood. Blood samples may be assayed immediately. Alternatively, the blood sample may be stored at low temperatures, for example in a fridge or even frozen before the method is conducted. Alternatively, the blood sample may be stored at room temperature, for example between 18 to 22 degrees Celsius, before the method is conducted. The blood sample may comprise blood serum. The blood sample may comprise blood plasma. Alternatively, the detection is carried out on whole blood or the blood sample is peripheral blood. The blood may be further processed before the diagnostic method is performed. For instance, an anticoagulant, such as citrate (such as sodium citrate), hirudin, heparin, PPACK, or sodium fluoride may be added.

[0100] The subject may be a vertebrate, mammal, or domestic animal. Most preferably, however, the subject is a human being. The subject may be a male or female. The subject may be a child or adult. The subject may be a febrile child.

[0101] The method according to the first, second, third, or ninth aspect, may be performed in vivo. However, in some embodiments, the method is performed in-vitro or ex-vivo. In some embodiments, the method is performed in-vitro. In some embodiments, the method is performed ex-vivo.

[0102] All of the features described herein (including any accompanying claims, abstracts and drawings), and / or all of the steps of any method or process so disclosed, may be combined with any of the above aspects in any combination, except combinations where at least some features and / or steps are mutually exclusive.

[0103] For a better understanding of the invention, and to show how embodiments of the same may be carried into effect, reference will now be made, by way of example, to the accompanying Figures, in which:-

[0104] Figure 1 shows one embodiment of the "drop-one-in" approach to predict future disease states and trajectories of individual subjects using the VeloCD algorithm. Each point represents a single sample. The low-dimensional transcriptomic space can be generated using PCA, tSNE or UMAP and can be visualised in 2 or 3 dimensions. The former is shown here for simplicity.

[0105] Figure 2 shows that the first core assumption of VeloCD is met in the Influenza A (A, C, E, F, H, J and N) and SARS-CoV-2 (B, D, G, I, L, M, O) RNA-Seq datasets. A and B - RNA was extracted from the blood of infected volunteers on day 0 and the days shown postinoculation. C (n=138, 21,647 genes) and D (n=267, 23,943 genes) - PCA plots of the spliced transcriptome across time. These plots show the segregation of those that test positive or remain negative of their respective viral infections at day 2 and 3 in C and day 3 and 4 D respectively. E, F and H - volcano plots showing the genes DE between those who tested PCR positive or remined negative for flu at: day 1 (E, n=6 positives versus n=6 negatives, day 2 samples used, 2,988 genes) and day 2 (F, n=9 positives versus n=6 negatives, day 3 samples used, 892 genes). G - the equivalent analysis for the SARS-CoV-2 cohort (day 3 samples used, 1,010 DEGs, n = 18 positives versus n=9 PCR negative subjects). H (n=46) and I (n=27) - PCA plots of the unique multi-exon DEGs identified in E and / or F (H, 2,115 genes) and G (I, day 3 samples shown, 715 genes) respectively. J and K (n = 138), L and M (n=267) - the spliced transcript expression of two of the top genes identified by maSigPro as changing across time and distinct between the PCR positive and negative subjects. N (1,563-genes, n=138) and O (2,726 genes, n=267) - all genes identified by maSigPro transformed onto a PCA, the large line was superimposed to illustrate how the points line up by time (day postinoculation, day 0.5 is the post-inoculation sample taken on the day of inoculation). The large circle highlights the cluster of points representing samples from those who remained PCR negative. These were generated using TPNN values of 45 (N) and 25 (O) respectively.

[0106] Figure 3 shows that the Influenza A (n = 138; A, B, E, G, I, K) and SARS-CoV-2 (n=267; C, D, F, H, J, L) human challenge RNA-Seq datasets meet the second core assumption of VeloCD. The unspliced transcript expression of IFI44L (A, C) and LY6E (B, D) respectively. E-H the temporal spliced and unspliced transcript expression of APOL3 (E, n=6 and F, n=9) and BAK1 (G, n=6 and H, n=9) in the blood samples of specific study participants. I (1,563-genes) and J (2,726 genes) -predicted future spliced expression of all genes (calculated at day 2) versus actual spliced transcript expression at day 3. K (n=138, 1,563-genes) and L (n=267, 2,726 genes) -RNA velocity fate maps. The RNA velocity arrows attached to the centre of each point represent predicted future transcript trajectories. These were generates using TPNN values of 45 (K) and 25 (L) respectively.

[0107] Figure 4 shows one embodiment of a framework for selecting genes suitable for RNA velocity analysis prognosis prediction using VeloCD. The first three steps are non- optional, the latter two are optional if a reduced gene signature is required and can include the use of feature selection tools such as LASSO regression and selecting genes contributing the most variation to the first X (chosen by user) PCs (from PCA). Step 1 shows a volcano plot from differential expression analysis of "group 1" versus "group 2", where each point represents the log2FC (x-axis) and -log 10 false discovery rate corrected p-value (y-axis) of a single gene. Those above the FDR p-value threshold are considered statistically significant and taken forward for analysis. Step 2 shows the DISCO analysis of each gene's RNA velocity value (y-axis) with its other information (x- axis) in each sample (each point). Genes whose direction of both metrics (positive or negative) is concordant in at least "threshold" (chosen by user) number of samples is taken forward for further analysis. The intersection of these two analyses is used as the initial signature for VeloCD.

[0108] Figure 5 shows the prognostic potential of VeloCD in the Influenza A and SARS-CoV-2 challenge RNA-Seq datasets. The performance of the 3-gene Influenza A signature at predicting future PCR status from day 1 post-inoculation in the Influenza A (A, 23 participants, generated using a TPNN of 16 and PCA-based embeddings) and SARS-CoV- 2 cohorts (C, 36 participants, TPNN value: 8 and PCA-based embeddings). These ROC curves were generated for each TPNN hyperparameter run, the PCR positive subjects were used to calculate sensitivity. The PCR negative subjects were used to calculate specificity, B - the performance of the 125-gene signature at predicting future PCR Status from day 1 post-inoculation on the SARS-CoV-2 cohort (TPNN: 5, PCA-based fate maps).

[0109] Figure 6 shows RNA Velocity analysis of the placebo arm of a TB-IRIS clinical drug trial. A - A PCA plot of the spliced transcript expression of all multi-exon genes (n=95, 23,041 genes), colored and shaped by if / when the trial participants developed TB-IRIS (circles: participants that developed TB-IRIS two weeks after the beginning of the trial i.e. postsampling, n=8; squares: trial participants that never developed TB-IRIS, n = 51; triangles: participants that developed TB-IRIS by the time of sampling, up to two weeks since the beginning of the trial, n=36). B - the differential expression analysis of the participants in this trial that developed TB-IRIS within two weeks of treatment versus those that never developed this syndrome. This identified 5,108 DEGs (all points above the horizontal line). C - A PCA plot of the genes from B with spliced and unspliced transcript expression (3,448 genes) D - A UMAP-based RNA velocity fate map of all the samples in this cohort constructed using a 1,980-gene signature (TPNN : 45, NN: 25, UMAP with three-dimensions). Arrows represent the future transcriptomic trajectories of each participant sample point.

[0110] Figure 7 shows RNA Velocity analysis of the PERFORM cohort (n=421). A - PCA plot of the spliced transcript expression of all genes (24,165) with samples (mild, n = 118; moderate, n=191; severe, n=37; very severe, n=75) shaped by diagnosis (bacterial, n=221 or viral, n=200). B - volcano plots showing the results of the DEA between the mild study participants recruited at ED vs. severe and very severe patients recruited in the PICU for the bacterial (6,589 DEGs) and viral subjects separately (4,991 DEGs). C - PCA plot (n=203) of the spliced transcript expression of the 1,591 multi-exon genes DE in both analyses in B with concordant log2FC values. This is shown for the mild participants recruited at ED (n=98) and severe and very severe patients recruited in PICU (n = 105). D - UMAP-based RIMA Velocity fate map (57 genes), n=222 in total, n = 19 healthy controls, n=98 mild participants, n=105 severe (n=32) and very severe (n=73) patients recruited in the PICU, showing two alternative disease trajectories (deterioration and recovery). Arrows represent the future transcriptomic trajectories of each participant sample point. This was generated with an NN of 20 and TPNN of 70. E - the median transition probabilities (to the PICU) of 20 mild (n = 10 bacterial, n = 10 viral), 5 severe, 2 very severe and 184 moderate (ward) patients (n=90 viral, n=94 bacterial). These were calculated across TPNN values 10 to 70 and projected into low-dimensional space using UMAP-based fate maps (NN:20).

[0111] Figure 8 shows the RNA velocity analysis of the INSTINCT cohort, (a) Summary table descripting the characteristics of the INSTINCT cohort, (b) Scatterplots showing the spliced and unspliced transcript expression of 1 FI44L and Ly6E in this dataset (n=138 samples), (c) A UMAP-based RNA velocity fate map of 231-genes from the 232 SARS- CoV-2 signature that predicts future PCR status in the SARS-CoV-2 CHIM dataset. Individual points represent individual samples coloured by timepoint and shaped by the PCR status of the original participant at the beginning of the study. Sample-specific arrows represent RNA velocities (future transcriptomic trajectories). The superimposed circle indicates the clustering of the PCR negative (squares and circles) participants and latter timepoints of the PCR positives (triangles). The superimposed arrows demonstrate the direction of the day 0 and 7 samples of the PCR positive participants.

[0112] Examples

[0113] Whole-blood transcriptomics has revolutionised infectious disease research by providing an abundance of information about the immunopathology of infection. These studies often examine the transcriptome at the gene or spliced transcript level to identify genes and biological pathways representative of disease severity at the time of sampling. More recently, RNA velocity analysis has unlocked the potential of the unspliced transcriptome to model and predict future transcriptomic states.

[0114] Accordingly, the inventors hypothesised that if a patient's current disease state is well- represented by their current (spliced transcript) transcriptome at the time of sampling, then their future disease states can also be captured by the component of the transcriptome that represents future gene expression (unspliced transcripts). Under these assumptions, the inventors set out to expand and adapt RNA velocity into a new prognostic tool, called VeloCD, which can predict future disease states of individual patients from host whole-blood RNA-Seq. Materials and Methods

[0115] This section provides details of the analysis of the 4 whole-blood RNA-Seq datasets discussed in this study.

[0116] RNA-Seq Pre-processing

[0117] The reads from all RNA-Seq datasets were first quality control checked using FastQC (Andrews, 2010), mapped using STAR (Dobin et al., 2013) and quantified using featurecounts (Liao, Smyth & Shi, 2014). The strandedness of the data was confirmed using the infer_experiment.py programme of the RSeQC package (Wang, Wang & Li, 2012).

[0118] FeatureCounts was run using the meta feature option meaning reads were collated from all exons (per gene exon level) or introns (per gene intron level) or features (gene) that overlapped a single gene into three separate gene specific counts. This was performed using the human GRCh38 genome assembly fasta and GTF files from Ensembl. Prior to this pre-processing, the GTF file was edited to reduce the transcripts of the same gene down into the minimal non-overlapping set, with introns introduced between the new set of non-overlapping exons. This ensured a 1 : 1 ratio between the number of genes with one non-overlapping transcript and the number of "transcript" level features. This method also ensures that intronic sequences do not overlap any exons of each genes annotated alternative transcript.

[0119] The exon level reads were used as an approximation for the spliced transcript expression. The intron level reads were used as an approximation for the unspliced transcript level expression. Genes whose introns overlap any exons of another gene on the same strand were removed from downstream analysis. Similarly, genes with any introns labelled with the biotype "retained intron" were also removed. Only genes with both spliced and unspliced transcript expression were used for this analysis. Spliced and unspliced transcript expression is then separately normalised for library size differences using the edgeR trimmed mean of m-values method and returned as Iog2 transformed counts per million values.

[0120] PCA

[0121] PCA was performed using the R packages: PCAtools (Blighe, Kevin & Lun, 2020), ggplot2 (Wickham, 2016), ggbiplot (Vincent, 2011), factoMineR (Le, Josse & Husson, 2008) and factoExtra (Kassambara, A. et al., 2017; Kassambara, Alboukadel & Mundt, 2020). The latter two packages were also used to extract the percentage contribution of each gene to each of the PCs of PCA. The scale=FALSE argument was used for all PCA plots presented in this study.

[0122] DEA

[0123] The DEA of the CHIM datasets (Influenza A and SARS-CoV-2) was performed using the glmFit and glmLRT functions from the edgeR R package. The DESeq function from the DeSeq2 package was used to perform the DEA of the TB-IRIS and PERFORM datasets (Love, Anders & Huber, 2014). Gene-level raw expression values were used for this analysis with pre-expression filtering of at least 5 in at least 3 samples used across all datasets. All volcano plots were generated using the EnhancedVolcano package in R (Blighe, K. et al., 2021).

[0124] Statistical Tests

[0125] For the analysis of the PERFORM cohort, Fisher's exact tests were used to compare median TPs (to the PICU group, grouped into >= 0.75 and <0.75) and categorical clinical features (fisher. test() function in R). Pearson correlations were performed using the cor.test() function with the "pearson" method argument in R. For all statistical tests, a p-value cut-off of 0.05 was used to determine significance. For the DEA, the Benjamini-Hochberg method was used to adjust for multiple testing and a threshold of 0.05 was used to determine statistical significance.

[0126] Influenza A

[0127] The generation of the human controlled infection model (HCIM) Influenza A whole-blood paired-end RNA-Seq data is described in Rosenheim et al., 2023 (Rosenheim et al., 2023). This cohort consisted of 23 individuals with 7 or 8 samples each taken at day 0 pre-inoculation and then day 1, 2, 3, 7, 10, 14 and 28 post-inoculation (n = 178). 17 of the patients went on to test positive, 6 remained negative by PCR throughout the study.

[0128] There was an average of 24,340,283 (92.5%) uniquely mapped reads per sample (range: 18,566,465-29,205,267, 87.5-95.6%). Day 14 and 28 samples were excluded from downstream analysis. Spliced and unspliced transcript expression values were then calculated from the quantified reads and PCA performed for exploratory analysis (21,647 genes, n=138).

[0129] DEA then identified genes DE between those who tested positive vs those who remained negative (by PCR, n=6) using the day 2 sample of those who tested positive on day 1 (n=6, 2,988 DEGs). For those who tested positive on day 2 (n=9), the day 3 samples were used for this analysis (892 DEGs). Sex was included as a covariate in the design matrices of these analyses.

[0130] The R package maSigPro was used to perform backward feature selection to identify genes that were changing across time and whose expression values diverged between PCR status (n = 138). Genes with low spliced transcript expression were filtered prior to this analysis (>= 3 in 3 samples for both spliced and unspliced transcript expression, 7,726 genes), which selected 1,563 genes with a linear relationship with time and were DE between PCR groups. The participant IDs were included as the "Replicate" design matrix covariate in this analysis. The 1,563-gene set was put into VeloCD to generate fate maps for exploratory analysis, using the hyperparameter values NN: 25 (tSNE and UMAP only) and TPNN: 45.

[0131] DISCO analysis was used to identify genes whose future change in spliced transcript expression were well-modelled by the RNA velocity algorithm. This analysis used the day 1 and 2 samples of all subjects (n=46), day 3 samples of the: negative (n=6), day 2 positive people (n=9) and single day 3 positive person (n=l, n=62 in total). The RNA velocity values of all genes (21,647 genes) were compared to future observed change in spliced transcript expression calculates as: spliced transcript expression day on X + 1 - spliced transcript expression on day X. For example, these were calculated for the day 1 samples as day 2 - day 1, and the day 2 samples as day 3 - day 2. There were 1,646 genes concordant in direction (value sign) in at least 30 samples. The intersection between these genes and those from the DEA of this dataset (3,389 unique genes) was 190 genes.

[0132] The spliced transcript expression of this 190-gene set was then input into LASSO regression to reduce this signature down into a minimal more highly predictive set able to classify the "reference" samples by PCR status. The reference samples were picked as the day after the subject tested positive (PCR positive reference). The day 3 samples of the PCR negative subjects were also used. The day 3 sample of the single person who tested positive on day 3 was used due to a lack of day 4 sampling in this cohort. No samples from the single day 4 positive subject were used in the reference set (n=22 reference samples). This method identified a 5-gene signature, which was then run through PCA. The top three genes which contributed most to the top three PCs were then selected as the final gene signature for RNA velocity analysis, where each participant's day 1 test sample day was dropped into the algorithm with the reference (with self-removed reference sample removed). Its probability of transition to each outcome (PCR positive PCR negative) was then recorded and collated across participant runs, which used the same hyperparameter values. The range of TPNN values run was 2- 19. PCA was used to embed the spliced transcript expression into low-dimensional transcriptomic space.

[0133] Sensitivity and specificity values were then calculated across each equivalent run and ROC curves, Area Under the Curve (AUC) and 95% confidence intervals, were constructed and calculated using the roc(), rocit(), auc() and ci.auc() functions from the ROCit (Khan, Md Riaz Ahmed Ahmed, 2019) and pROC (Robin et al., 2011) R packages. The hyperparameter run with the highest AUC was then selected as the "best performance" of the signature. This was a TPNN value of 16 for this analysis.

[0134] Table 1. Influenza A 3-aene signature

[0135] SARS-CoV-2

[0136] The generation of the HCIM SARS-CoV-2 whole-blood paired-end RNA-Seq data is described in Rosenheim et al., 2023 (Rosenheim et al., 2023). In short, peripheral blood was collected from 36 participants on day 0 pre- and post-inoculation and on days 1, 2, 3, 4, 5, 7, 10, 14 and 28 post-inoculation (n=385). 9 of these subjects remained negative throughout this trial, 18 had sustained replicative infections.

[0137] There were an average 57,790,882 (82.3%) uniquely mapped reads per sample (range: 23,776,992-128,131,106, 66.1-90.3%). The day 28 samples were excluded from downstream analysis as were the samples belonging to the two participants that sero- converted during quarantine. The 7 subjects with ambiguous PCR results were also removed (n=267, Participant IDs excluded: 439679, 636163, 645438, 666475, 672533, 677306, 677696). DEA identified 1,010 DEGs that distinguished between the PCR positive and negative subjects on day 3. Sex and age were included as covariates in this analysis.

[0138] Using a backward selection approach, maSigPro identified genes whose spliced transcript expression was both DE by PCR status and correlated with time. In this model the postinoculation day 0 samples were given with the value 0.5. Low expression pre-filtering was performed prior to this method (spliced and unspliced transcript expression >= 3 in >= 3 samples, 7,837 genes), which selected 2,726 genes input into VeloCD for exploratory analysis (TPNN: 25). DISCO analysis was performed using the day 0 (post-inoculation), day 1 and day 3 samples. Genes whose RIMA velocity values were concordant in direction with an approximation of the future gene expression change: unspliced transcript expression minus the spliced transcript expression, in at least 70 samples were then selected (2,643 genes). There were 125 genes intersecting between this set and the DEA of this dataset. The 125-gene signature was then input into VeloCD, where the day 1 sample of each participant was dropped into the analysis with the PCR positive and negative reference groups (with self- reference sample removed). The probability of transition of each day 1 sample to each of the latter-day groups was then collated between participant runs, which used the same hyperparameter values (TPNN: 2-26). Sensitivity and specificity values were then calculated using these values and ROC curves, AUC and 95% confidence intervals calculated as described in the previous section. PCA was used to embed the spliced transcript expression in low-dimensional space. The optimal result presented for this analysis used a TPNN value of 5.

[0139] The 3-gene signature from the Influenza A cohort analysis was also used to predict future PCR status in the SARS-CoV-2 cohort. These genes were extracted from the spliced and unspliced transcript files and run using the drop-one-in approach. Similarly, the 125-gene signature from the SARS-CoV-2 analysis was also used to predict future PCR status in the Influenza A cohort.

[0140] Table 1. SARS-CoV-2 125 gene signature

[0141]

[0142] TB-IRIS

[0143] The week 2 samples of the placebo arm of the clinical trial described in Meintjes et al., 2018 (Meintjes et al., 2018) were used for this study. One subject with sepsis was excluded from downstream analysis. Another whose blood sample had a low uniquely mapping read percentage (< 50%) was also removed.

[0144] Of the remaining subjects, 51 never developed TB-IRIS, 36 had developed this syndrome by the time of sampling and 8 developed it at least two weeks after the beginning of the trial (post-sampling, n=95). The blood RNA-Seq samples had a uniquely mapping average of 33,645,255 (79.2%) reads per sample (range: 22,624,044- 55,714,568, 57.9-89.1%). Quality control revealed evidence of rRNA contamination.

[0145] Consequently, genes whose biotype was labelled as rRNA, Mt_rRNA, rRNA_pseudogene, and ribozyme, were removed.

[0146] There were 5,108 DEGs between the subjects who never developed TB-IRIS (n=51) and those that developed it within two weeks of the start of the trial (n=36). Sex was included as a covariate in this analysis. Xist was then removed from the set of significant genes. DISCO analysis identified 9,670 genes with concordance between the direction of the RNA velocity values and unspliced minus the spliced transcript expression. There were 1,980 shared genes between these DEA and DISCO analyses.

[0147] This 1,980-gene signature was input into VeloCD using the "drop-one-in" approach to predict the future development of TB-IRIS in the subjects (n=8) that had not developed the syndrome at the time of sampling. Each of these "test" subjects were put into this algorithm with all reference subjects (n=87 total, n = 51 non-TB-IRIS subjects, n=36 TB- IRIS subjects). PCA was used to embed the spliced transcript expression into low dimensional transcriptomic space.

[0148] The probabilities of transition of each test subject to the TB-IRIS group were then collated between runs which used the same hyperparameter values (TPNN range: 5-50). Sensitivity values were then calculated as number of subjects whose probability of transition to the TB-IRIS was greater than or equal to the probability threshold divided by the total number of test subjects (n=8) multiplied by 100.

[0149] Table 3. TB-IRIS gene signature

[0150]

[0151]

[0152]

[0153]

[0154]

[0155] PERFORM

[0156] The PERFORM whole-blood RNA-Seq dataset consists of 3,052 paired-end samples representing different lane-files per clinical event (1-4 lanes per event sample). A clinical event is defined as each unique time a study participant presented to hospital. The RNA- Seq samples from this study had a range of uniquely mapped reads between 842,305 and 105,918,580 (12.3-95.8%) and mean of 18,545,806 (94.4%).

[0157] For the analysis presented in this study, samples from clinical events without severity group classifications and those with non-viral or bacterial phenotypic diagnosis were removed. Bacterial groups were defined as anyone belonging to the bacterial syndrome, probable bacterial or definite bacterial groups. The viral group was defined as anyone belonging to the viral syndrome, probable viral or definite viral group. This reduced the sample size down in 421 clinical events, each represented by a single time-point sample (time-point 1 only, read counts summed between sequencing lanes). 105 of these study participants were recruited in the PICU and 316 were recruited in the ED. The severity groups consisted of mild (n = 118: n=79 bacterial, n=39 viral), moderate (n=191), severe (n=37) and very severe (n=75).

[0158] The data was divided into two sets. The first (n=203, reference set) consisted of the mild patients recruited in the ED (with 20 randomly removed, n=98) and the severe and very severe patients recruited in the PICU (n = 105). This set was used for proceeding DEA and LASSO regression. The second (n=218, test set) consists of the 20 randomly selected mild patients, 7 severe or very severe patients (recruited in the ED) and 191 moderate subjects also recruited in the ED. This set was used to test the signature derived by the first set. The first set was also used as reference samples for the second set.

[0159] DEA was performed between the mild patients and those recruited in PICU for the viral and bacterial groups separately. Genes whose biotype was labelled as rRNA, Mt_rRNA, rRNA_pseudogene and ribozyme were removed prior to this analysis and no covariates were included in the design matrix. The multi-exon genes DE in both analyses with concordant log2FC values were then identified (1,371 genes), and their spliced transcript expression transformed using PCA. The participants used to construct the spliced-unspliced transcript plots in Figure 3 E-H were SPFC018 (E), 634105 (F), SFC001 (G) and 680539 (H). The arrows were added between each point (ordered by time in ascending order) using the arrows() function in R.

[0160] DISCO analysis was performed using all 24,165 multi-exon expressed genes and all 421 samples. Genes whose RNA velocity values was concordant in direction with the unspliced-spliced transcript expression in at least 200 samples were selected (13,689 genes). There were 1,371 genes in common between the above DEA (1,371 genes) and this DISCO analysis. These genes and the samples from the first set were put into LASSO regression to distinguish mild from PICU samples. This selected 57 genes, which were then input into VeloCD (NN:20, TPNN: 10-70) with the mild and PICU samples (n=203) combined with 19 healthy controls (n=222 in total).

[0161] The inventors then used this 57-gene signature to predict future disease states for the test samples. These disease states were defined as two opposing clinical trajectories: the recovery (towards the mild patients that went home) and deterioration direction (towards the PICU patients). The inventors used the "drop-one-in" approach for this analysis, dropping one test sample with the reference set (n=203 + 1 test sample). This analysis was run across TPNN values of 10 to 70 for UMAP (NN:20). Results were collected for each sample across hyperparameter values and median TP values (to the PICU group) were calculated. These were plotted on a raincloud plot.

[0162] The inventors then examined the relationships between the following clinical features and these TP values for the moderate group only (n=184): gender (male or female), triage code (immediate, very urgent and urgent combined group, non-urgent and standard combined group), life-saving intervention (all interventions combined into a yes group), ill appearance (yes or no), full recovery (yes or no), level of disability at baseline and discharge (good performance, and a combined mild, moderate or severe "disability" group), sepsis (yes or no), febrile illness resolved (yes or no), new symptoms developed (yes or no), re-presentation healthcare (yes or no), and final phenotype diagnosis (combined into viral and bacterial groups, "diagnosis"). For "febrile illness resolved" and "new symptoms developed", if either follow-up call labelled this as yes, then the participant was put into the yes group. Participants with follow-up calls when one is labelled as no, and the other "null" were not used for analysis because it can't be definitively determined if the unknown null call was a clinical event or not. Fisher's exact tests were used to determine the significance of the relationship between each feature and the median TP values (to the PICU) grouped by a threshold value (> = or < 0.75). A cut-off of 0.05 was used to determine significance. These tests were repeated for the bacterial and viral patient groups separately.

[0163] The 7 clinical event subjects in the moderate severity group were removed prior to the clinical feature analysis because they had ambiguous clinical metadata (it was unclear if they were admitted to the ward). They had the following IDs: BIV-1408-1517-E1, BIV- 1301-1008-E1, BIV-1301-1188-E1, BIV-1701-1101-E1, BIV-1801-1011-E1, BIV-2001- 1076-E1, BIV-1701-1185-E1.

[0164] Table 4. PERFORM 57-qene signature Example 1 - Current gene expression captures current disease states

[0165] The inventors first illustrated how the two core assumptions of VeloCD are met using two Controlled Human Infection Model (CHIM) whole-blood RNA-Seq datasets. CHIMs (or "human challenge" studies) involve the intentional infection of volunteers with a well- characterised strain of a pathogen of interest, followed by monitoring over several weeks, all under highly controlled conditions (Zhou et al., 2023).

[0166] The inventors used two viral CHIM datasets, where subjects were infected with Influenza A (Influenza A / Belgium / 4217 / 2015, 23 participants) or SARS-CoV-2 (Asp614Gly, 36 participants) and sequential blood samples were taken for RNA-Seq (Rosenheim et al., 2023; Zhou et al., 2023). Blood samples were taken from the Influenza A patients on day 0 (pre-inoculation) and then the following days post inoculation: 1, 2, 3 ,7, 10, 14 and 28 (n = 178 samples). Blood samples were taken from the SARS-CoV-2 patients on day 0 (pre-inoculation) and the following days post-inoculation: day 0 (post-inoculation), 1, 2, 3, 4, 5, 7, 10, 13, 14 and 28 (n=385 samples). This experimental set-up is summarised in Figure 2 A and B.

[0167] In each CHIM cohort, some volunteers became infected with the challenge virus, whereas others remained negative throughout by PCR-testing. Amongst infected volunteers there was variation in time from inoculation to first positive PCR test, and variation in symptom severity, but none became severely unwell (requiring medical intervention).

[0168] In the Influenza A study, 6 of the participants tested positive on day 1 post inoculation, 9 tested positive on day 2, one on day 3, and one on day 4. In total 17 people tested positive throughout the study and 6 remained negative. In the SARS-CoV-2 study, 2 participants seroconverted between pre-challenge screening and inoculation (Zhou et al., 2023) and were removed from analysis. Of the remaining 34 participants, 7 had non- consecutive PCR positive, meaning the course of their infection was unclear and they were therefore also removed. Of the remaining 27 subjects, 18 tested positive and 9 remained negative for SARS-CoV-2. 6 subjects tested positive on day one (n=4 during the evening post-blood sampling PCR sample) and 12 on day 2.

[0169] These time-course datasets were used to test VeloCD because they allowed us to observe how well RNA velocities appeared to capture gene expression change across time in whole-blood RNA-Seq and note how these diverge between the infected (became PCR positive at some point during the trial) and PCR-negative groups. The exact timing of PCR positivity was also known allowing us to test whether RIMA velocity could predict this as a future disease using only earlier time-point samples.

[0170] Before predicting these outcomes, the inventors first tested whether the assumptions of RNA velocity hold true for whole blood samples. The inventors first confirmed that spliced transcript expression ("current" gene expression) is representative of the current disease states of the subjects in these cohorts. The day 14 and 28 samples were excluded from the Influenza A cohort analysis, leaving 138 samples with 21,647 detectable genes, and the day 28 samples were removed from the SARS-CoV-2 dataset, leaving 267 samples with 23,943 detectable genes. These samples were excluded because gene expression had returned to baseline by these timepoints. The spliced transcript expression of all genes with spliced and unspliced transcript expression, was transformed into low-dimensional transcriptomic space (Figure 2 C and D) using PCA (Blighe, Kevin & Lun, 2020). There was clear segregation of samples in low-dimensional transcriptomic space by both time since inoculation and infection status (PCR positive or negative), with the segregation of those who tested positive from negative most prominent from day 2 for Influenza A and day 3 for SARS-CoV-2, respectively.

[0171] The inventors then illustrated how Differential Expression Analysis (DEA) can identify the genes that distinguish between those who tested positive from those who remained negative at specific time-points. Gene-level expression values were required for this analysis as edgeR requires raw non-normalised expression (Robinson, McCarthy & Smyth, 2010).

[0172] The Influenza A cohort infected subjects were segregated by when they first tested positive and the day after first positivity used for DEA. On day 2 there were 2,988 Differentially Expressed Genes (DEGs) between those who tested positive on day 1 (n=6) and those who remained negative (n=6). On day 3, there were 892 DEGs between those who tested positive on day 2 (n=9) and the negatives (n=6). Between these two comparisons there were 3,389 unique genes. The day after PCR-positivity was used for this analysis because these time-points represented the maximum differences in gene expression, as evident in the PCA plot (Figure 2 C-D).

[0173] There were 1,010 genes Differentially Expressed (or with differential expression, DE) between the SARS-CoV-2 infected subjects (n = 18) and those who remained PCR- negative (n=9). Day 3 samples were used for this analysis (Figure 2 E, F and G). The spliced expression values from the multi-exon subset of these DEGs (suitable for later RNA velocity analysis) were transformed into low-dimensional transcriptomic space (PCA), illustrating how well they segregate the infected and PCR-negative subjects at these time-points (Figure 2 H and I). In these Figures, PCR status refers to whether they will test positive at any point during the challenge study (PCR positive or "infected") or remain negative (PCR negative).

[0174] Concurrent with this analysis, the maSigPro backward selection algorithm was used to identify genes with DE across time and between the infected subjects those that remained PCR negative for each virus (Conesa et al., 2006). The spliced transcript expression of two of the top genes (ordered by multiple testing corrected p-values) identified by both analyses, IFI44L (Interferon Induced Protein 44 Like) and LY6E (Lymphocyte antigen 6E), are shown in Figure 2 J, K, L and M. They show clear patterns of temporal expression change for the PCR positive subjects and a lack of change in samples derived from the PCR negatives subjects. This pattern is representative of many of the genes identified by this method, 1,563 and 2,726 for Influenza A and SARS-CoV-2 respectively, as illustrated by the PCA plots in Figure 2 M and N, which show clear temporal ordering (superimposed red line) of the PCR positive samples that are distinct from the tight cluster of samples from those who remained PCR negative (blue circle). Together, these analyses illustrate that the spliced transcript expression appears to represent different disease states (and sample timing) supporting the first fundamental assumption of VeloCD.

[0175] Example 2 - RNA velocities capture future disease states The inventors then investigated the second assumption of VeloCD: that RNA velocities can capture future gene expression, which corresponds to future disease states. The unspliced transcript expression of IFI44L and LY6E shown in Figure 3 A-D illustrates how the unspliced transcript expression follows expected biologically informative patterns across time. This would not be observed if these reads represented noise from genomic- DNA contamination, unannotated exons or retained introns - ideas posed by earlier studies (DeLuca et al., 2012; Mortazavi et al., 2008; Sultan et al., 2008), that the more recent literature (Lee, S. et al., 2020) and this study opposes.

[0176] The relationship between spliced and unspliced expression is shown for two genes (APOL3 and BAK1) significant in the maSigPro analysis of both datasets (Figure 3 E-H). These illustrate expected patterns of expression upregulation between day 0 and 2-3 of the trial followed repression from day 3 onwards where the expression appears to decrease. Future spliced transcript expression values are predicted during the calculation of the RIMA Velocity values. For both CHIM datasets and using all the significant genes identified in the maSigPro analysis, the inventors took the predicted future spliced transcript expression values calculated from the day 2 samples and correlated them with measured expression values on day 3 (Figure 3 I and J). The range of Pearson correlation coefficients across subjects was 0.86-0.99 (mean: 0.96, all p-values < 2.2e-16) and 0.76-0.99 (mean: 0.93, all p-values < 2.2e-16) in the Influenza A and SARS-CoV-2 cohorts respectively, indicating that this method is accurately modelling the temporal expression change of these genes.

[0177] These same genes were then input into VeloCD, and fate maps generated (Figure 3 K and L). Each point on these maps represents a single time-point sample from each participant in these studies and the RNA velocity arrows represent predicted future transcriptomic states. The superimposed large grey arrow highlights how the RNA velocity arrows represent clear transitions across time in the infected subjects. The circled points are the segregated samples belonging to the PCR negative study participants (Figure 3 K and L), thus illustrating how these fate maps are also capturing diverging trajectories: gene expression and disease state change across time for the infected individuals, whilst expression varies very little in the PCR-negative individuals, consistent with the gene-level plots shown in Figure 3 A-D.

[0178] Together these analyses illustrate RNA velocities calculated with VeloCD correspond to both future transcriptomic and current disease states.

[0179] Example 3 - VeloCD has prognostic potential

[0180] Having demonstrated that the assumptions of VeloCD hold true in both CHIM datasets, the inventors then assessed whether this algorithm could be used on samples from day 1 post-inoculation to predict future infection status. To perform this analysis VeloCD requires reference samples to represent each outcome group (Figure 1). For the Influenza A cohort, samples from one day after each infected subject first tested positive were used as the infected reference. The day 3 sample of the single person who tested positive on day 3 was added to this, due to a lack of sampling on day 4 in this cohort. The day 3 samples from the PCR-negative group were used as the negative reference. The single subject testing positive on day 4 subject was excluded from the reference, leaving a total of n=22 reference samples. The day 3 samples were used as the reference groups for the SARS-CoV-2 cohort analysis (n=27). VeloCD works more optimally with the pre-selection of highly predictive genes whose expression across time is representative of diverging disease trajectories: infected or PCR-negative. Theoretically all genes could be used for prediction, however many will have velocity values close to 0, and whose expression is not associated with disease progression. These genes may mask and dilute the signal from the more informative genes and consequently should be removed prior to the RIMA velocity analysis.

[0181] The framework for selecting biologically informative genes was developed using the Influenza A dataset and then partially replicated using the SARS-CoV-2 dataset.

[0182] This framework has several key steps (Figure 4):

[0183] 1. DEA of the reference subjects. Take the DEGs from this step after multiple testing correction and log2FC (fold change) filtering (optional). Prior to this analysis, users also have the option to remove any genes that are not suitable for downstream translation use (including transfer onto clinical devices) such as genes with expression below specific thresholds and non-protein coding genes.

[0184] 2. Perform Discordance-Concordance (DISCO) analysis between the RNA velocity values of all genes (in the reference and test samples) and other available information that can be used to identify how well the direction (positive or negative) of these values are being estimated. The latter includes the unspliced minus the spliced transcript expression and the spliced transcript expression of the next day minus the previous day. A threshold number of samples in which these values are concordant is then used to select a set of highly concordant genes.

[0185] 3. Take the intersection of 1. and 2. as your gene selection for input into VeloCD.

[0186] 4. Optional: Reduce the genes from 3. into a smaller set using feature selection algorithms such as Least Absolute Shrinkage and Selection Operator (LASSO) regression.

[0187] 5. Optional: Perform further gene signature reduction by selecting the genes (from 4.) whose percentage contribution to the top 3 Principal Components (PCs) of the PCA is greater than or equal to a threshold value.

[0188] For the Influenza A cohort, the 3,389 unique genes from the analysis presented in Figure 2 also correspond to the output of step 1. The DISCO analysis then compared the direction of each gene's RNA velocity values to the spliced transcript expression of the earlier time points subtracted from the future of the respective sample: day 2 minus day 1 expression, day 3 minus day 2 expression. This was performed for the day 1 to 3 samples (n=63) only because they have consecutive timepoints. Genes concordant in at least 30 samples were then selected: 1,646 genes. There were 190 common genes between these two steps (step 3). These were then used as the input for LASSO regression, where their spliced transcript expression was used to classify the reference samples by PCR status (step 4). This selected 5 genes. These were then reduced to the top three genes that contributed most to the top 3 PCs of the PCA of the same samples (step 5).

[0189] VeloCD was then run using the "drop-one-in" approach (Figure 1) with each participant's day 1 sample (23 participants) dropped in with the reference samples. The reference sample from the same study participant was removed prior to each run to prevent bias due to sample identity. The probability of transition for each test sample to the infected group were then collated between hyperparameter (TPNN) runs. Sensitivity and specificity values were then used to measure this predictor's performance for the subjects who went on to test positive and remain negative respectively. The runs which used a TPNN of 16 were optimal for these predictions, with an area under an Area Under the Receiver Operating Characteristic (AUROC) Curve value of 0.89 (95% Confidence Interval: 0.76-1; Figure 5 A) to predict infection status between a few hours to 3 days into the future (most subjects tested positive 24 hours after sampling, n=9, reference samples were extracted 24-48 hours after day 1).

[0190] This analysis was then repeated for the SARS-CoV-2 dataset. The 1,010 DEGs shown in the volcano plot in Figure 2 were the output of the first step of this gene selection framework. DISCO analysis (step 2) then compared the direction of the RNA velocities against the spliced transcript expression subtracted from the unspliced transcript expression for each of the reference and test samples. The day 0 post-inoculation samples were also used in this analysis (n=81 in total). Genes concordant in at least 70 samples were selected (2,643 genes). 125 genes were shared between the DEA and DISCO analysis (step 3). These genes were then used as the input for VeloCD to predict future PCR status in this cohort. Across hyperparameter runs (for all 36 participants) the optimal analysis had an AUROC of 0.76 (95% Confidence Interval: 0.57-0.94, Figure 5 B, TPNN: 5) to predict infection status from several to 24 hours into the future (reference samples were extracted 48 hours after day 1).

[0191] Influenza A and SARS-CoV-2 are both viruses with single stranded RNA genomes (McCauley & Mahy, 1983; Naqvi et al., 2020). The literature also suggests common immune pathways that are stimulated in response to viral infections (Tsalik et al., 2021). Consequently, the inventors hypothesised that the genes selected from each cohort could be predictive in the other and follow similar patterns of expression change across time. The optimal runs of the 3-gene influenza A signature to predict (on average) infection status 24 hours into the future had an AUROC of 0.73 (95% Confidence Interval: 0.53-0.93, TPNN: 8) for predicting SARS-CoV-2 status at 24 hours into the future (Figure 5 C). 2 of the 3 genes in this signature showed clear patterns of expression change across time in both datasets with spliced transcript expression change in the SARS-CoV-2 infected subject lagging slightly behind those with the Influenza A infection, consistent with the divergence between the infected and PCR-negative blood transcriptomes appearing to occur from day 3 and day 2 respectively (Figure 2 C-D). There are no intersecting genes between the 3 the 125-gene signatures.

[0192] The 125-gene SARS-CoV-2 signature did not perform well at predicting future PCR status from day 1 in the Influenza A cohort (optimal TPNN: 2, AUROC: 0.57, 95% Confidence Interval: 0.30-0.84; Figure 5), possibly because it was optimised for later timepoint prediction (reference samples were extracted 48 hours after day 1 in the SARS-CoV-2 cohort). Despite this, many of the genes in this signature had similar patterns of expression change across time.

[0193] Overall, this analysis has demonstrated the prognostic potential of this method in the context of predicting future disease state in whole-blood RNA-Seq datasets with multiple time-points.

[0194] Example 4 - VeloCD can indicate the future development of TB-IRIS

[0195] The inventors then investigated whether VeloCD can predict the onset of a different infectious disease manifestation, outside of the highly controlled situation of CHIM. Tuberculosis (TB)-associated Immune reconstitution inflammatory syndrome (TB-IRIS) is a paradoxical immunopathological reaction that develops in some patients with TB and HIV-coinfection shortly after they begin Anti-Retroviral Treatment (ART). It is characterized by new or recurrent symptoms at sites of TB infection (Lai, Meintjes & Wilkinson, 2016). Recent studies have demonstrated that TB-IRIS has a distinct host transcriptomic signature in blood which may have prognostic potential (Mbandi et al., 2022).

[0196] The inventors analysed whole blood transcriptomes of subjects (n=95) from the placebo arm of a clinical trial in which subjects who had been on anti-TB treatment for less than 30 days were started on ART and randomised to receive either prednisone or placebo (Meintjes et al., 2018). Using samples collected two weeks after initiation of ART, the inventors assembled two reference groups: 1) those who never developed TB-IRIS (n=51) during at least 12 weeks of follow-up, and 2) those who developed TB-IRIS within the first two weeks of the trial (n=36). To test whether RNA-velocity could indicate the future occurrence of TB-IRIS, the inventors used samples from participants who had not yet developed the syndrome at the time of sampling ("test" samples, n=8), but went on to develop TB-IRIS between 15 and 37 days after starting ART.

[0197] The inventors first illustrated that the core assumptions of VeloCD were met in this dataset. The first is clearly demonstrated by the segregation of the TB-IRIS and non-TB- IRIS patients in low-dimensional transcriptomic space of the spliced transcript expression of all genes (Figure 6 A 23,041 genes). Interestingly, the test subjects (n=8) are located between the two reference groups in this space.

[0198] DEA identified 5,108 DEGs between the TB-IRIS and no TB-IRIS groups (Figure 6 B), the subset of these with spliced transcript expression was then input back into PCA showing high levels of sample segregation by TB-IRIS status (Figure 6 C). These genes were the output of the first step of the gene selection framework introduced in the previous section (Figure 4). Following this framework, DISCO analysis identified high levels of concordance between the direction of the RNA velocity values and the spliced subtracted from the unspliced transcript expression with 9,670 concordant in at least 50 samples (n=95 in total). There were 1,980 genes in common between the DEA and DISCO analysis.

[0199] This 1,980-gene signature was used as an input for VeloCD, where each test participant (n=8) was individually dropped into the algorithm ("drop-one-in" approach, Figure 1) with the reference samples (n = 51 non-TB-IRIS, n=36 TB-IRIS). The probability of transition of each test subject to the TB-IRIS group was collated across hyperparameter runs (TPNN range: 5-50). Predictions were stable across these runs. Using the PCA embedding method to generate the fate maps and a TPNN of 30, 7 of 8 (87.5%) of these subjects were correctly predicted to develop this syndrome at a probability of at least 30%.

[0200] The UMAP-based RNA velocity fate map in Figure 6 D, consisting of all test and reference subjects, illustrates how the two reference groups form a continuum with velocity arrows pointing away from each other with all but one test subject pointing towards the TB-IRIS group. Despite the small size of the test group (n=8), this analysis illustrates the prognostic potential of this method in an RNA-Seq dataset consisting of data collected at a single time-point.

[0201] Example 5 - VeloCD indicates changes in severity of acute febrile illness

[0202] When clinicians assess any acutely unwell patient, they must try to identify both the cause and severity of illness. Many clinical decisions, such as admission to hospital, investigations, and medication choices, depend on perception of the current severity of illness and the likelihood that a patient's condition will improve or deteriorate. This is particularly the case for febrile children, who can often appear unwell, but most will have self-limiting viral illnesses, whilst some will have more severe viral and bacterial infections.

[0203] Having shown that VeloCD is able to predict future disease states (TB-IRIS, infections status), the inventors next wanted to assess whether it can indicate trajectories of illness severity, represented by requirement for different levels of clinical care. For this analysis, the inventors used whole-blood RNA-Seq data generated from the "Personalised Risk assessment in Febrile illness to Optimise Real-life Management across the European Union" (PERFORM, IRAS ID: 209035; https: / / www.hra.nhs.uk / planning-and-improving- research / application-summaries / research-summaries / perform / ) study. PERFORM was a multi-country prospective observational study of children with suspected infection presenting to hospitals in Europe. One of the core goals of PERFORM was to identify new host response-based diagnostic signatures to distinguish bacterial from viral infections. The inventors chose this dataset for analysis because of its rich clinical metadata including the grouping of the participants' clinical events into severity groups: mild (child sent home), moderate (child admitted to the ward), and severe / very severe (child sent to the Paediatric Intensive Care Unit, PICU). Participants in this study were either recruited in the Emergency Department (ED) or PICU. The participants were phenotyped into non-infectious causes of fever, or definite, probable and "syndrome" viral and bacterial groups (Nijman et al., 2021). For this analysis only the bacterial (combined definitive bacterial, probable bacterial and bacterial syndrome group) and viral (combined definitive viral, probable viral and viral syndrome group) participant events were used (n=421 in total: n=200 viral, n=221 bacterial). The inventors also used 19 healthy control subjects for one component of the exploratory analysis (for aim 3 below). For each subject a single, presentation time-point whole-blood RNA-Seq sample was analysed.

[0204] The analysis of the PERFORM dataset had three main aims:

[0205] 1. To illustrate that the first assumption of VeloCD is also met in this dataset: that current disease severity is captured by whole-blood gene expression.

[0206] 2. To illustrate the diverging trajectories of recovery and deterioration as RNA velocity fate maps. The inventors hypothesised that mild patients (n=98: n=69 viral, n=29 bacterial) recruited in the ED would form one branch of a trajectory, whilst deteriorating patients would follow a trajectory towards the severe and very severe patients recruited in PICU (called the "PICU" group, n = 105: n=29 viral, n=76 bacterial). To perform this analysis, genes associated with disease severity were identified by the framework outlined in Figure 4. This analysis also tested the second assumption of VeloCD: that RNA velocities capture future disease trajectories.

[0207] 3. To assess the predictive abilities of the gene signature identified in the above analysis. Using the mild and PICU subjects as "recovery" and "deterioration" reference groups "drop-one-in" analysis (Figure 1) was performed on 20 mild patients not used in previous analysis, 7 patients recruited in the emergency department but requiring immediate admission to the PICU and the 184 participants that were sent to the ward (moderate severity group). These groups were examined in isolation and the relationship between transition probabilities for the deterioration direction towards the severe group of the moderate group and their clinical features including diagnostic phenotype (viral or bacterial) was inspected.

[0208] To achieve the first aim, the spliced transcript expression of all genes was first reduced into low-dimensional transcriptomic space using PCA. This showed that the samples clearly segregate by severity with the mild and PICU participants most distinct from each other with the moderate severity group positioned between them (Figure 7 A). "Diagnostic group" (bacterial or viral) is also clearly intercorrelated with disease severity on these plots. This is unsurprising given that the bacterial group are over- re presented in the PICU subset (76 / 105- 72.4%).

[0209] DEA identified 4,991 and 6,589 genes DE between the mild and PICU participants in the viral and bacterial groups respectively (Figure 7 B). 2,366 of these genes were DE in both analyses and had concordant log2FC values. The subset of these intersecting genes with spliced and unspliced transcript expression (1,591 genes) were used to generate the PCA plot in Figure 7 C, which shows very clear segregation by severity. Together these analyses show how this dataset meets the first assumption of VeloCD - that disease severity at the time of sampling is captured by spliced transcript (current) gene expression.

[0210] To capture the two diverging trajectories, recovery, and determination, the inventors first used the gene selection framework outlined in Figure 4, to identify genes that both capture severity and are suitable for RNA velocity analysis. The 2,366 intersecting genes from the above DEA were used as the output of the first step of this analysis, which consisted of the mild and PICU samples only (n=203). The intersecting genes were used to attempt to capture the shared aspects of the viral and bacterial immune response. DISCO analysis of all samples (n=421) identified 13,689 genes with concordant RNA Velocity and unspliced-spliced transcript expression in at least 200 samples. There were 1,371 intersecting genes between the DEA and DISCO analysis. These were input into LASSO regression to identify a smaller set of genes with spliced transcript expression that distinguish the mild from PICU subjects. This identified a 57-gene signature, which was input into VeloCD.

[0211] To more definitively illustrate that the direction of the RIMA velocity arrows of the mild patients was truly representative of "recovery", the inventors added the blood samples of 19 "healthy" patients to this analysis. These participants were also recruited as a part of the PERFORM study in hospital as non-infectious controls who had non-infectious issues that were unlikely to affect the blood transcriptome including pre-surgical samples for polydactyly and tonsillectomy.

[0212] The UMAP-based fate map shown in Figure 7D illustrates how the 57-gene signature clearly captures the recovery and deterioration pathways of these study participants. Deterioration is shown in the right top corner direction (away from the healthy participants), recovery is represented by the left-bottom corner direction with the RNA velocity arrows of the mild patients pointing towards the healthy subjects. Interestingly, the bacterial mild patients appear segregated form the viral mild patients on this fate map.

[0213] Having achieved the first two aims of this analysis, the inventors lastly wanted to examine the prognostic potential of the 57-gene signature. The inventors first looked at how well it performs for 20 mild patients (10 viral, 20 bacterial) randomly selected and excluded from all prior analysis and 7 participants who were recruited in the ED but went straight to the PICU. VeloCD was run for each participant across a range of TPNN values (10-70, NN: 20, UMAP-based fate maps), results were collated and the median transition probabilities to PICU examined for each participant sample. Median values were used because of the high levels of consistency between hyperparameter runs (excluding any outlier runs; Figure 7 E). As expected, the mild viral participants have very low TPs (to the PICU), the bacterial mild participants are a lot more split with 30% (n=3) having median transition probabilities over 0.75. Unsurprisingly, the severe (n = 5) and very severe patients (n=2) have very high median transition probabilities (to the PICU) - all over 0.5.

[0214] Lastly, the inventors examined the relationship between the median TP values of the moderate (ward) group and their clinical features (Table 5). 7 participants were removed from the original moderate group (n = 191 to n = 184) due to conflicting information in the clinical metadata regarding their transfer to the ward. Median TPs were grouped by a threshold value of below or above / equal to 0.75 and examined across each categorical clinical feature using Fisher's exact tests. Clinical features are separated into those present at presentation ("on-presentation") and those present at discharge, or only determined later (i.e. disease phenotype, "future severity").

[0215] As expected, gender has no significant relationship to these values and has an odds ratio close to 1, indicating that the 57-gene signature is not capturing sex-specific severity effects. For the other features, except for disability (at baseline and discharge), all odds ratios indicate that the more severe group ("exposed group", i.e. the presence of sepsis) have higher odds of having TPs (to the PICU) >= 0.75, indicating that these values are capturing some aspects of both present and future disease severity. More specifically, both triage code and re-presentation to healthcare are approaching but do not quite reach significance in this analysis (Table 5).

[0216] The median TPs (to the PICU) of the bacterial and viral ward patients have slightly different distributions (Figure 7 D), with the viral subjects having overall lower values. Fisher's exact test found a significant association between these values and diagnosis (bacterial, n=94, viral, n=90; p-value: 0.02, Table 5). Due to the potential confounding effect of this factor on the other clinical features, these tests were repeated, when possible, for the bacterial (Table 6) and viral subjects separately (Table 7).

[0217] Interestingly, for the bacterial subjects, patients with sepsis had 2.8-fold higher odds of being in the high (>= 0.75) TP group. This value is borderline not significant (p-value: 0.051, Table 6). No features were significant in analysis of the viral subjects, but almost all odds ratios from both analyses indicate that the more severe group ("exposed group", i.e. the presence of sepsis) have higher odds of having TP (to the PICU) values >= 0.75, they just do not reach significance, due to small group sizes. Together these results illustrate how these predictions are representative of some facets of the biology and clinical features, both present and future, of these infections.

[0218] In summary, the analysis of the PERFORM dataset has further illustrated the prognostic potential of VeloCD.

[0219] Table 5. The association between clinical features and median transition probabilities in the PERFORM subjects who were sent to the ward (n = 184).

[0220] Odds ratios were calculated using Fisher's exact test for the female, bacterial and more severe phenotypic groups ("yes", triage code: "very urgent", baseline and discharge: "disability"), meaning that they give the odds of i.e. presence of sepsis being in the high transition probability group (p-value <0.05). TP=Transition Probability. Percentages are given to one decimal place. All other values are given to two decimal places.

[0221] Table 6. The association between clinical features and median transition probabilities in the PERFORM subjects who were sent to the ward and had a bacterial infection (n=94). Odds ratios were calculated using Fisher's exact test for the female and more severe phenotypic groups ("yes", triage code: "very urgent", baseline and discharge: "disability"), meaning that they give the odds of i.e. presence of sepsis being in the high transition probability group (p-value <0.05). TP=Transition Probability. Percentages are given to one decimal place. All other values are given to two decimal places.

[0222]

[0223] Table 7 ■ The association between clinical features and median transition probabilities in the PERFORM subjects who were sent to the ward and had a viral infection (n=90). Odds ratios were calculated using Fisher's exact test for the female and more severe phenotypic groups ("yes", triage code: "very urgent", baseline and discharge: "disability"), meaning that they give the odds of i.e. presence of sepsis being in the high transition probability group (p-value <0.05). TP=Transition Probability. Percentages are given to one decimal place. All other values are given to two decimal places.

[0224] Discussion VeloCD is a new prognostic tool that predicts future disease states from whole-blood RNA-Seq data. It adapts the concept of RNA velocity into a new context overcoming the main limitation of most blood transcriptomic studies - that they only capture a limited view of what is happening in an organism at the time of sampling and provide little information about what will happen in the future.

[0225] The core assumptions of VeloCD were based on the observations in the literature that whole-blood transcriptomic data provides much information about a patient's current disease state at the time of sampling and that RNA velocity values can predict future transcriptomic state. As well as demonstrating how these assumptions hold true in the datasets analysed in this study, the inventors show that RNA velocity can be applied to predict future disease states in whole-blood RNA-Seq.

[0226] This study also demonstrated that, using the gene selection framework described in Figure 4, genes can be identified whose future transcriptomic states are also representative of future disease states. This was illustrated across diseases (Influenza A, SARS-CoV-2 and TB-IRIS), study designs (CHIM, clinical trial, perspective observational study) and outcome variables: future PCR status, future development of TB-IRIS and change in clinical care level.

[0227] Example 6 - INSTINCT study - additional dataset to support modelling in viral infections To provide additional support to the claimed invention, i.e., that RNA velocity can model and capture future disease trajectories in natural viral infections, the inventors modelled future disease trajectories using RNA-Seq data from the SARS-CoV-2 (IRAS ID: 282820) household contact surveillance INSTINCT study.

[0228] INSTINCT consisted of 56 (Figure 8a) individuals who were: i) PCR positive for SARS- CoV-2 (n=24) at the start of the study, ii) were a close contact of someone who tested positive but remained negative (n=24), or iii) non-contact PCR negative controls (n=8), throughout the testing period. Samples were collected at day 0, 7, 8, 14 and 28 post-the beginning of the study (n=138 samples).

[0229] Because study participants were already exposed to the pathogen and PCR positive or negative at the start of the study, the inventors could not test the ability of the 232- gene signature, discovered in the SARS-CoV-2 Controlled Human Infection Model (CHIM) dataset, to predict future PCR status in this cohort. Instead, the inventors examined the expression dynamics of some of these genes and determined how well the RNA velocity fate map could capture the PCR positive participants returning to the "PCR negative state" as the study progressed and their infections resolved.

[0230] The inventors first examined the expression dynamics of two genes previously identified as significantly associated with PCR status and change over time (using maSigPro), IFI44L and Ly6E respectively, in both the flu and previous SARS-Cov-2 CHIM datasets. Both the spliced and unspliced transcripts show biologically relevant and similar patterns of change over time with the unspliced transcripts having expression more intermingled between PCR groups earlier than the spliced transcripts (Figure 8b and 8c).

[0231] Unsurprisingly because individuals were much further into their illness than in the CHIM cohorts, the spliced and unspliced transcript expression is already separated by PCR status at day 0 but then become much closer and completely intermingled by day 28. This resembles the latter timepoints in both CHIM datasets, replicating the inventors' previous findings.

[0232] They next took the 232 genes in the predictive signature from the SARS-CoV-2 CHIM- analysis, 231 of which were expressed in this dataset, and examined how well it captures disease trajectories in a fate map constructed from all timepoints in INSTINCT (Figure 8d). As expected, the timepoint 0 samples from the PCR positives were well separated from the negatives, appear to be going away and then coming back towards the next timepoint (day 7) whose RNA velocity arrows are pointing towards the day 8 and latter timepoint cluster of samples (blue circle). This is also where the PCR negative samples are situated, which as expected, have very small RNA velocity arrows, indicative of small amounts of gene expression change. These observations are consistent with the dynamics of IFI44L and Ly6E, where people start separated by PCR status at the start of the study and then come back around to having similar transcriptomes to the PCR negatives as their infections resolve.

[0233] Conclusion

[0234] This study has illustrated the prognostic potential of VeloCD across a range of infections including specific pathogens (Influenza A, SARS-CoV-2), syndromes (TB-IRIS), and broad diagnostic groups (bacterial and viral), as well as for a variety of outcome variables (future PCR status, syndrome onset, and level of clinical care). This potential has also been illustrated across a range of whole-blood RNA-Seq datasets from those with sequential sampling to single time-points. In summary, the proof of concept of this new prognostic method has been illustrated by this disclosure.

Claims

Claims1. A method for predicting a subject's condition and / or providing a prognosis of a subject's condition, the method comprising: i) performing RIMA analysis on a subject; ii) calculating an RNA velocity value based on the ratio of unspliced mRNA transcripts to spliced mRNA transcripts, for at least one gene in the sample; and iii) using the RNA velocity value to predict future expression levels of the at least one gene, to thereby provide a prediction and / or prognosis of the subject's condition.

2. A method for determining the efficacy of treating a subject suffering from a condition with a therapeutic agent or a specialised diet, the method comprising: i) performing RNA analysis on a subject; ii) calculating an RNA velocity value based on the ratio of unspliced mRNA transcripts to spliced mRNA transcripts, for at least one gene in the sample; and iii) using the RNA velocity value to predict future expression levels of the at least one gene, to thereby determine whether treatment with the therapeutic agent or the specialised diet is effective or ineffective.

3. The method according to either claim 1 or claim 2, wherein the RNA analysis is performed using RNA-Seq, microarray, qPCR, LAMP, or technologies using CRISPR-Cas.

4. The method according to any one of the preceding claims, wherein the method comprises the step of pre-processing the RNA sequencing data, optionally wherein pre-processing the RNA sequencing data comprises removing low quality cells and / or genes, normalising gene expression values, handling any technical variations, and / or removing genes that do not have spliced and unspliced transcripts.

5. The method according to any one of the preceding claims, wherein the method comprises the step of calculating the expression of spliced mRNA transcripts and unspliced mRNA transcripts for the at least one gene in the sample.

6. The method according to claim 5, wherein the method comprises the step of transforming the spliced transcript expression into low-dimensional transcriptomic space, optionally wherein Principal Component Analysis (PCA), t-distributed stochastic neighbor embedding (tSNE), or Uniform Manifold Approximation and Projection (UMAP), is used to transform the spliced transcript expression into lowdimensional transcriptomic space.

7. The method according to either claim 5 or claim 6, wherein calculating the RIMA velocity value comprises subtracting the unspliced transcript expression calculated under a steady state from the actual unspliced transcript expression.

8. The method according to any one of the preceding claims, wherein a positive RNA velocity value indicates an increase in future expression levels for the at least one gene in the sample, and a negative velocity value indicates a decrease in future expression levels for the at least one gene in the sample.

9. The method according to any one of the preceding claims, wherein the method comprises embedding the RNA velocity value on a low-dimensional transcriptomic space.

10. The method according to any one of the preceding claims, wherein the method comprises comparing the predicted future expression levels of the at least one gene, with expression levels of the gene in at least one reference sample or at least two reference samples.

11. The method according to any one of the preceding claims, wherein providing a prognosis of the subject's condition comprises predicting their future disease state, optionally wherein their disease state is asymptomatic, uncomplicated or severe disease.

12. The method according to any one of claims 1 to 10, wherein providing a prognosis of the subject's condition comprises predicting trajectories of illness severity, optionally wherein the illness severity is represented by mild severity groups (sent home), moderate severity groups (admitted to a hospital ward), or severe / very severe groups (sent to Intensive Care Unit).

13. The method according to any one of the preceding claims, wherein future expression levels is expression levels of the at least one gene up to one, two, three, four or five hours after the RNA analysis has been performed, or up to six, seven, eight, nine or ten hours after the RNA analysis has been performed.

14. The method according to any one of the preceding claims, wherein future expression levels is expression levels of the at least one gene up to 12, 14, 16, 18, 20, 22 or 24 hours after the RIMA analysis has been performed, or up to 36, 48 or 72 hours after the RNA analysis has been performed, or up to one week, two weeks, three weeks or four weeks after the RNA analysis has been performed.

15. The method according to any one of the preceding claims, wherein the at least one gene is one, two, three, four or five genes, or wherein the at least one gene is six, seven, eight, nine or ten genes, or wherein the at least one gene is at least 100 genes, at least 1000 genes, at least 10,000 genes, or at least 50,000 genes.

16. The method according to any one of the preceding claims, wherein the method further comprises selecting at least one gene whose future transcriptomic states are representative of future disease states.

17. The method according to any one of the preceding claims, wherein the condition is an infectious disease, an inflammatory disease, a metabolic disease, an endocrine disease, cancer, a degenerative disease, drug toxicity, a cardiovascular disease, trauma (injury), Tuberculosis (TB)-associated immune reconstitution inflammatory syndrome (TB-IRIS), sepsis, acute febrile illness, or an acute illness.

18. The method according to claim 17, wherein the infectious disease is a bacterial infection, a viral infection, a fungal infection, a parasitic infection, or a combination thereof.

19. The method according to any one of the preceding claims, wherein the sample comprises a blood sample, optionally wherein the sample comprises a wholeblood sample.

20. A method of predicting a subject's condition and / or providing a subject's prognosis, the method comprising performing RNA velocity analysis on whole blood of a subject.

21. The method according to claim 20, wherein the method comprises performing RNA velocity on a single whole blood sample from the subject.

22. A method of selecting at least one gene for prognosing a condition in a subject using RNA velocity, the method comprising:i) conducting differential expression analysis (DEA) to identify at least one gene that can distinguish between a subject who is positive for the condition and a subject who is negative for the condition; ii) performing discordance-concordance (DISCO) analysis of the RIMA velocity value to identify at least one gene whose RNA velocity value is concordant in a threshold number of samples; and iii) selecting the at least one gene that meets the criteria of i) and ii).

23. The method according to claim 22, wherein the method further comprises removing at least one gene that is not suitable for downstream translation use, optionally wherein the method comprises removing at least one gene with expression below a specific threshold, removing at least one non-protein coding gene, or removing a gene with a certain biotype.

24. The method according to claim 22 or claim 23, wherein step i) of the method comprises multiple testing correction and / or logzFC (fold change) filtering.

25. The method according to any one of claims 22 to 24, wherein performing the discordance-concordance (DISCO) analysis comprises subtracting spliced transcript expression from unspliced transcript expression, for the at least one gene.

26. Computer-readable instructions which, when executed by a computing apparatus, cause the computing apparatus to perform the method of any one of claims 22 to 25.

27. A computer-readable medium comprising program instructions stored thereon for performing the method according to any one of claims 22 to 25.

28. An apparatus comprising: at least one processor; and at least one memory including computer program code which, when executed by the at least one processor, causes the apparatus to perform the method according to any one of claims 22 to 25.

29. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, causes the apparatus to perform the method according to any one of claims 22 to 25.

Citation Information

Patent Citations

  • Combined biomarker detection method for body health status assessment

    CN115678973A

  • Liver cancer prognosis model construction method and application

    CN118262916A

  • Computational models to analyze RNA velocity

    US20230268024A1

  • Use of cancer cell expression of cadherin 12 and cadherin 18 to treat muscle invasive and metastatic bladder cancers

    US20240115699A1

  • Methods characterizing multiple analytes from individual cells or cell populations

    WO2019157529A1