Methods and compositions for monitoring and diagnosing health and disease conditions
A method for generating a transcriptome-wide expression profile of platelets using RNA sequencing addresses the lack of intra-individual variability in current diagnostics, enabling precise health and disease monitoring and diagnosis.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- UNIV OF UTAH RES FOUND
- Filing Date
- 2026-04-13
- Publication Date
- 2026-07-24
AI Technical Summary
Current gene expression diagnostics lack accurate criteria for intra-individual variability in platelet gene expression, relying on large cross-sectional cohorts to indirectly account for noise, which does not correct for genetic variants affecting gene expression.
A method for preparing a transcriptome-wide expression profile of isolated platelets using RNA sequencing to generate a dataset that accounts for intra-individual variability, providing stable gene expression signatures for health and disease monitoring.
The method provides a reliable and accurate basis for monitoring and diagnosing health and disease states by accounting for intra-individual variability, enabling precise screening, diagnosis, and prognosis using platelet gene expression signatures.
Smart Images

Figure 2026121367000002 
Figure 2026121367000003 
Figure 2026121367000004
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims the benefit as of the filing date of U.S. Provisional Application No. 62 / 951,004, filed on December 20, 2019. The contents of this prior application are incorporated herein by reference in their entirety.
[0002] Statement regarding federally funded research This invention was made with government support under grant numbers AG048022 and HL144957, awarded by the National Institutes of Health. The government has certain rights in this invention.
[0003] Incorporation of sequence lists This application includes a sequence listing submitted simultaneously with the filing of this application via EFS-Web, filed with the filename "21101_0410P1_SL.txt", which is 4,096 bytes in size, created on December 6, 2020, and is incorporated herein by reference in its entirety.
[0004] This disclosure relates to compositions and methods relating to healthy gene expression signature criteria for platelets that can be used for monitoring and diagnosing health and disease states in subjects. These platelet healthy gene expression signature criteria can be used to screen for disease states, diagnose them, monitor their onset, monitor their progression, identify and diagnose them, or as prognostic indicators. Treatment plans can also be established and evaluated using these platelet healthy gene expression signature criteria. [Background technology]
[0005] Current gene expression diagnostics rely on the use of large cross-sectional cohorts, including healthy controls, to account for noise from inter-individual and intra-individual variability. There are no available criteria for expected variability in platelets from healthy individuals over time. Cross-sectional studies require a large number of healthy controls to indirectly account for intra-individual variability. These cross-sectional studies do not directly correct for intra-individual variability or genetic variants that may significantly affect gene expression. Diagnostic tests need more accurate "healthy" criteria. [Overview of the Initiative]
[0006] A method for preparing a transcriptome-wide expression profile of a biological sample is disclosed herein, comprising: a) measuring the expression level of one or more genes present in a first biological sample, wherein the measurement includes RNA sequencing; b) determining the expression level of a gene from the measured expression level obtained in step a); c) combining the results of step a) and b) to generate a transcriptome-wide expression profile; and d) providing the transcriptome-wide expression profile as a dataset, wherein the first biological sample consists of isolated platelets. [Brief explanation of the drawing]
[0007] [Figure 1A] This shows the intra-individual and inter-individual stability of platelet RNA expression over a period of 4 months (Cohort 1) to 4 years (Cohort 2). Figure 1A shows unsupervised clustering and heatmaps of total RNA expression in platelets derived from Cohort 1 samples. The histogram on the left of each heatmap shows the distribution of distances between sample pairs, and the intensity of the blue color indicates the similarity between sample pairs. Samples clustered as neighbors in the heatmap dendrogram reflect the transcriptome with the highest similarity. The nearest neighbor self-pairs are highlighted in yellow and gray, and the nearest neighbor non-self-pairs are highlighted in orange. [Figure 1B]This shows the intra-individual and inter-individual stability of platelet RNA expression over 4 months (Cohort 1) to 4 years (Cohort 2). Figure 1B shows an example of an individual correlation plot of transcripts from Cohort 1. Each data point represents the normalized logarithmic expression level (RLD) of a single transcript from a specified donor at time 0 (x-axis) versus time 0, 2 weeks, 4 months, or 4 years (y-axis), either within the same individual (upper panel) or between different individuals (lower panel). Points are heat-colored according to density. P-values are derived from Pearson correlations. [Figure 1C] Figures 1C and 1F show the intra-individual and inter-individual stability of platelet RNA expression over a period of 4 months (Cohort 1) to 4 years (Cohort 2). Figures 1C and 1F show box plots summarizing the intra-individual and inter-individual RNA expression Pearson correlations for all tested time points (left) or individually specified time points (right). Note that for the specified time points in Figure 1F, the mean inter-individual correlation did not significantly decrease when comparing further analyzed samples. For example, there was no significant difference when comparing the mean intra-individual correlation at time point 0 versus 2 weeks to the mean intra-individual correlation at time point 0 versus 4 years. Box plots for Cohort 1 (Figure 1C) before and after adjustment for age, sex, and race are shown; these are not adjusted for Cohort 2 (Figure 1F) due to the smaller sample size. P-values are from the Wilcoxon test and are adjusted. [Figure 1D] This shows the intra-individual and inter-individual stability of platelet RNA expression over a period of 4 months (Cohort 1) to 4 years (Cohort 2). Figure 1D shows unsupervised clustering and heatmaps of total RNA expression in platelets derived from Cohort 2 samples. The histogram on the left of each heatmap shows the distribution of distances between sample pairs, and the intensity of the blue color indicates the similarity between sample pairs. Samples clustered as neighbors in the heatmap dendrogram reflect the transcriptome with the highest similarity. The nearest neighbor self-pairs are highlighted in yellow and gray, and the nearest neighbor non-self-pairs are highlighted in orange. [Figure 1E]This shows the intra-individual and inter-individual stability of platelet RNA expression over 4 months (Cohort 1) to 4 years (Cohort 2). Figure 1E shows an example of an inter-individual correlation plot of transcripts from Cohort 2. Each data point represents the normalized log-transformed expression level (RLD) of a single transcript from a specified donor at time 0 (x-axis) versus time 0, 2 weeks, 4 months, or 4 years (y-axis), either within the same individual (upper panel) or between different individuals (lower panel). Points are heat-colored according to density. P-values are derived from Pearson correlations. [Figure 1F] Figures 1C and 1F show the intra-individual and inter-individual stability of platelet RNA expression over a period of 4 months (Cohort 1) to 4 years (Cohort 2). Figures 1C and 1F show box plots summarizing the intra-individual and inter-individual RNA expression Pearson correlations for all tested time points (left) or individually specified time points (right). Note that for the specified time points in Figure 1F, the mean inter-individual correlation did not significantly decrease when comparing further analyzed samples. For example, there was no significant difference when comparing the mean intra-individual correlation at time point 0 versus 2 weeks to the mean intra-individual correlation at time point 0 versus 4 years. Box plots for Cohort 1 (Figure 1C) before and after adjustment for age, sex, and race are shown; these are not adjusted for Cohort 2 (Figure 1F) due to the smaller sample size. P-values are from the Wilcoxon test and are adjusted. [Figure 2A] This section shows a comparison of intra-individual and total variability of each transcript in platelets across cohorts. Mean intra-individual variability and total variability (standard deviation, SD) were calculated from the normalized log-transformed expression (RLD) of each transcript. Figure 2A shows the intra-individual variability, plotting each transcript in Cohort 1 (x-axis) against the variability of each transcript in Cohort 2 (y-axis). The horizontal and vertical lines at 0.5 represent arbitrary thresholds for variability used in the Venn diagrams in Figures 2B and 2D. [Figure 2B]This section shows a comparison of intra-individual and total variability of each transcript in platelets across cohorts. Mean intra-individual and total individual variability (standard deviation, SD) were calculated from the normalized log-transformed expression (RLD) of each transcript. The horizontal and vertical lines at 0.5 indicate arbitrary thresholds of variability used in the Venn diagrams in Figures 2B and 2D. Figure 2B shows the Venn diagram of overlapping transcripts with the highest total individual variability (D). Below each Venn, the significantly enriched GO terms for transcripts overlapping between the two cohorts are listed. FDR = Benjamini false discovery rate calculated by David (Huang DW, et al. Nat Protoc. 2009;4:44-57). [Figure 2C] This section shows a comparison of intra-individual and total variability of each transcript in platelets across cohorts. Mean intra-individual and total individual variability (standard deviation, SD) were calculated from the normalized log-transformed expression (RLD) of each transcript. Figure 2C shows the total individual variability, plotting each transcript in Cohort 1 (x-axis) against the respective variability of each transcript in Cohort 2 (y-axis). The horizontal and vertical lines at 0.5 represent arbitrary thresholds for variability used in the Venn diagrams in Figures 2B and 2D. [Figure 2D] This section shows a comparison of within-individual and total variability of each transcript in platelets across cohorts. Mean within-individual and total individual variability (standard deviation, SD) were calculated from the normalized log-transformed expression (RLD) of each transcript. The horizontal and vertical lines at 0.5 indicate arbitrary thresholds of variability used in the Venn diagrams in Figures 2B and 2D. Figure 2D shows the Venn diagram of overlapping transcripts with the highest within-individual variability (B) and total individual variability (D). Below each Venn, the significantly enriched GO terms for transcripts overlapping between the two cohorts are listed. FDR = Benjamini false discovery rate calculated by David (Huang DW, et al. Nat Protoc. 2009;4:44-57). [Figure 3A]This shows the enrichment of heritable traits and eQTLs in transcripts ranked by repetition. Figure 3A shows a table of the transcripts with the highest repetition in the RNA-seq data of Cohort 1, and their associations with race, sex, or eQTLs reported in the PRAX1 (Simon LM, et al. Am J Hum Genet. 2016;98:883-97, and Simon LM, et al. Blood. 2014;123(16):e37-45) microarray data. Associations with FDR < 1e-4 are highlighted in pink. NS = not significant. [Figure 3B] This shows that the heritable traits and eQTLs of transcripts ranked by repeatability are enriched. Figure 3B shows a correlation plot of MFN2 RNA expression (log-normalized) at time 0 (x-axis) and time 4 months (y-axis). Points are colored according to the rs1474868 genotype (ND = not determined). The above is a density histogram showing a bimodal distribution by genotype. Bimodal p-values from Hartigan's dip test for multimodality. [Figure 3C] This shows the enrichment of heritable traits and eQTLs in transcripts ranked by repetition. Figure 3C (top) shows enrichment plots for the presence of eQTLs ranked according to different measures of intra-individual variability, mean expression abundance, total variability, or repetition. The axes below the plots show the gene rank according to each measure and the value of the repetition measure (values for other measures are not shown on the axes). Genes with known eQTLs are shown in red, and genes without are shown in blue. Thus, genes with the highest repetition are nearly 100% eQTL genes, while genes with the lowest repetition are nearly 0%. Figure 3C (bottom) shows plots of cumulative enrichment scores for each metric. [Figure 3D] This shows that the heritable traits and eQTLs of transcripts ranked by repetition are enriched. Figure 3D shows the odds ratio of the likelihood of identifying eQTLs of a gene at a specified repetition threshold, compared to the same number of genes ranked by total variation. [Figure 3E]This shows enrichment of heritable traits and eQTLs in transcripts ranked by repetition. Figure 3E shows box plots of LINC01089 expression by rs1168863 genotype in cohorts 1 and NL at time 0 and 4 months (Best MG, et al. Cancer Cell. 2017;32:238-252). *P values adjusted for age, sex, and race (cohort 1) or population structure (cohort 2) (inferred genetic ancestry (Purcell S, et al. Am J Hum Genet. 2007;81:559-75, and Chang CC, et al. Gigascience. 2015;4:7)). [Figure 3F] This shows enrichment of heritable traits and eQTLs in transcripts ranked by repetition. Figure 3F shows box plots illustrating the allelic imbalance of rs1168863 in heterozygotes in cohorts 1 and NL. The ratio of RNA-seq reads containing A nucleotides to RNA-seq reads containing T nucleotides was calculated and plotted for each heterozygote. [Figure 4A] This shows the intra-individual and inter-individual stability of exon skipping in platelets. Figure 4A shows a schematic diagram of how the exon skipping event is defined. Exon splice-in (PSI) percentages are calculated using splice junction leads, which is the ratio of exon-containing junction leads to total junction leads. [Figure 4B] This shows the intra-individual and inter-individual stability of exon skipping in platelets. Figure 4B shows the correlation plot of PSI for exon skipping events within an individual (left panel) and between individuals (right panel). Each point represents a single exon skipping event from a specified donor at time 0 (x axis) versus time 0 or 4 months (y axis). [Figure 4C]This shows the intra-individual and inter-individual stability of exon skipping in platelets. Figure 4C shows box plots summarizing the intra-individual versus inter-individual Pearson correlations, adjusted for age, sex, and race, when analyzing the PSI of all exon skipping events at time 0 and 4 months. *Wilcoxon test, adjusted. [Figure 5A] This shows the recurrence rate of exon 14 skipping in SELP and its association with race. Figure 5A shows a table of the most recurring exon skipping events in platelets. [Figure 5B] This shows the repetition and racial association of exon 14 skipping in SELP. Figure 5B shows representative IGV plots of sequencing reads from two different individuals at time 0 and 4 months, illustrating the differential distribution of reads between individuals that align with or skip exon 14 in SELP. This histogram shows the cumulative abundance of reads aligned with each exon. Subsets of individual reads are shown below each histogram, with thin lines (not present in the read) connecting to thick lines (mapped portion of the read) indicating splice junction reads. Red and blue reads are splice junction reads that align with or skip exon 14, respectively. [Figure 5C] This shows the repetition of exon 14 skipping in SELP and its association with race. Figure 5C shows the correlation plot of SELP exon 14 PSI. Each point represents the PSI of an individual donor at time 0 (x axis) and time 4 months (y axis). Donors represented in the IGV plot in B are labeled in red. [Figure 5D] This shows the repetition of exon 14 skipping in SELP and its association with race. Figure 5D shows box plots of mean SELP exon 14 PSI by race at time 0 and 4 months. [Figure 6A]rs6128 is shown to be a platelet SELP exon 14 splicing QTL. Figure 6A shows an enlarged IGV plot depicting the read distribution across SELP exon 14 for individuals with rs6128 A / A and relatively high levels of exon skipping reads (top) or individuals with the rs6128 (T / T) variant (bottom). The change from C to T does not alter the amino acid sequence but changes exon splicing silencer and enhancer sites as predicted by Ex-Skip (Raponi M, et al. Hum Mutat. 2011;32:436-444). [Figure 6B] rs6128 is shown to be a platelet SELP exon 14 splicing QTL. Figure 6B shows a box-and-whisker plot of the mean PSI of SELP exon 14 by rs6128 genotype inferred from RNA-seq in cohort 1 at time point 0 and 4 months. [Figure 6C] rs6128 is shown to be a platelet SELP exon 14 splicing QTL. Figure 6C shows a box-and-whisker plot of the mean PSI of SELP exon 14 by rs6128 genotype inferred from publicly available RNA-seq data in the NL cohort. *P-values adjusted for age, sex, and ethnicity (Figure 6B) or population structure (Figure 6C) (inferred genetic ancestry (Purcell S, et al. Am J Hum Genet. 2007;81:559-75, and Chang CC, et al. Gigascience. 2015;4:7)). [Figure 7A]rs6128 is shown to directly control exon 14 skipping in SELP and alter the ratio of surface to soluble P-selectin protein expression. Figure 7A is a schematic diagram of the SELP minigene construct containing the ORF of SELP and the intron adjacent to exon 14. The C / C construct and the T / T construct differ by only one nucleotide at rs6128. The constructs were cloned into vectors with two different promoters (CMV or MSCV). After transfection into HEK293 cells, the intron is spliced and exon 14 is alternatively spliced (skipped). The degree of exon 14 skipping is measured by PCR using exon 14 flanking primers that generate two PCR products of different sizes. [Figure 7B] rs6128 is shown to directly control exon 14 skipping in SELP and alter the ratio of surface to soluble P-selectin protein expression. Figure 7B shows the RT-PCR analysis of SELP exon 14 skipping after transfection of HEK293 cells with the rs6128 C / C vector or the T / T vector. Representative results from five independent experiments are shown. Below are the bar graph and summary of standard errors of PSI calculated according to the densitometric analysis of the exon 14 inclusion band (upper band) divided by the sum (total) of the upper and lower bands. *Paired t-test, n = five independent experiments. [Figure 7C] rs6128 is shown to directly control exon 14 skipping in SELP and alter the ratio of surface to soluble P-selectin protein expression. Figure 7C shows the flow cytometry analysis of P-selectin surface expression after transfection of HEK293 cells with the rs6128 C / C vector or the T / T vector. Above is a representative histogram overlay of P-selectin surface expression 24 hours after transfection with the CMV promoter-empty vector, rs6128 C / C, or T / T. Below are the bar graph and summary of standard errors of the fold change in surface P-selectin MFI (normalized to transfection). *Paired t-test, n = 5 - 6 pairs per group. [Figure 7D] This demonstrates that rs6128 directly regulates exon 14 skipping in SELP and alters the surface ratio to soluble P-selectin protein expression. Figure 7D shows ELISA analysis of soluble P-selectin in the supernatant of HEK293 cells after transfection with rs6128 C / C vector or T / T vector. *Paired t-test, n=12-14 pairs per group. [Figure 8A] This shows the intra-individual and inter-individual stability of platelet non-coding RNA expression over 4 months (Cohort 1) and 4 years (Cohort 2). Figure 8A shows unsupervised clustering and heatmaps of non-coding RNA expression in platelets from Cohort 1 samples. The histogram on the left of each heatmap shows the distribution of distances between sample pairs, and the intensity of the blue color indicates the similarity between sample pairs. Samples clustered as neighbors in the heatmap dendrogram reflect the non-coding transcriptome with the highest similarity. The nearest neighbor self-pairs are highlighted in yellow and gray, and the nearest neighbor non-self-pairs are highlighted in orange. [Figure 8B] Figure 8B shows the intra-individual and inter-individual stability of platelet non-coding RNA expression over 4 months (Cohort 1) and 4 years (Cohort 2). An example of individual correlation plots for non-coding transcripts in Cohort 1 is shown. Each data point represents the normalized logarithmic expression level (RLD) of a single non-coding transcript from a specified donor at time 0 (x-axis) versus time 0, 2 weeks, 4 months, or 4 years (y-axis), either within the same individual (upper panel) or between different individuals (lower panel). Points are color-coded according to their density. [Figure 8C]Figures 8C and 8F show the intra-individual and inter-individual stability of platelet non-coding RNA expression over 4 months (Cohort 1) and 4 years (Cohort 2). Figures 8C and 8F show box plots summarizing the total intra-individual and inter-individual non-coding RNA expression Pearson correlations at time 0 and 4 months (Figure 8C), or (Figure 8F) time (left) or individually specified time (right). Figure 8C shows box plots for Cohort 1 before and after adjustment for age, sex, and race, while Cohort 2 is unadjusted (due to smaller sample size). P-values from Wilcoxon test, adjusted. [Figure 8D] This shows the intra-individual and inter-individual stability of platelet non-coding RNA expression over 4 months (Cohort 1) and 4 years (Cohort 2). Figure 8D shows unsupervised clustering and heatmaps of non-coding RNA expression in platelets derived from samples from Cohort 2 (Figure 8D). The histogram on the left of each heatmap shows the distribution of distances between sample pairs, and the intensity of the blue indicates the similarity between sample pairs. Samples clustered as neighbors in the heatmap dendrogram reflect the non-coding transcriptome with the highest similarity. The nearest neighbor self-pairs are highlighted in yellow and gray, and the nearest neighbor non-self-pairs are highlighted in orange. [Figure 8E] This shows the intra-individual and inter-individual stability of platelet non-coding RNA expression over 4 months (Cohort 1) and 4 years (Cohort 2). Figure 8E shows an example of individual correlation plots for non-coding transcripts in Cohort 2 (Figure 8E). Each data point represents the normalized logarithmic expression level (RLD) of a single non-coding transcript from a specified donor at time 0 (x-axis) versus time 0, 2 weeks, 4 months, or 4 years (y-axis) within the same individual (upper panel) or between different individuals (lower panel). Points are heat-colored according to their density. [Figure 8F]Figures 8C and 8F show the intra-individual and inter-individual stability of platelet non-coding RNA expression over 4 months (Cohort 1) and 4 years (Cohort 2). Figures 8C and 8F show box plots summarizing the total intra-individual and inter-individual non-coding RNA expression Pearson correlations at time 0 and 4 months (Figure 8C), or (Figure 8F) time (left) or individually specified time (right). Figure 8C shows box plots for Cohort 1 before and after adjustment for age, sex, and race, while Cohort 2 is unadjusted (due to smaller sample size). P-values from Wilcoxon test, adjusted. [Figure 9A] This section compares the intra-individual variability and total variability of each transcript in platelets. Mean intra-individual variability and total individual variability (standard deviation, SD) were calculated from the normalized log-transformed expression (RLD) of each transcript. Figure 9A shows the normalized expression (x-axis) plotted against the intra-individual variability (y-axis) for each transcript in Cohort 1. Labeled points represent representative transcripts with low intra-individual variability and high total individual variability. [Figure 9B] This section compares the intra-individual variability and total variability of each transcript in platelets. Mean intra-individual variability and total variability (standard deviation, SD) were calculated from the normalized log-transformed expression (RLD) of each transcript. Figure 9B shows the normalized expression (x-axis) plotted against the intra-individual variability of each transcript (y-axis) in Cohort 2 (Figure 9B). Labeled points represent representative transcripts with low intra-individual variability and high total variability. [Figure 9C] This section compares the intra-individual variability and total variability of each transcript in platelets. Mean intra-individual variability and total individual variability (standard deviation, SD) were calculated from the normalized log-transformed expression (RLD) of each transcript. Figure 9C shows the normalized expression (x-axis) plotted against total variability for each transcript (y-axis) in Cohort 1. Labeled points represent representative transcripts with low intra-individual variability and high total individual variability. [Figure 9D]This section compares the intra-individual variability and total variability of each transcript in platelets. Mean intra-individual variability and total variability (standard deviation, SD) were calculated from the normalized log-transformed expression (RLD) of each transcript. Figure 9D shows the normalized expression (x-axis) plotted against total variability for each transcript (y-axis) in Cohort 2 (Figure 9D). Labeled points represent representative transcripts with low intra-individual variability and high total variability. [Figure 9E] This section compares the intra-individual variability and total variability of each transcript in platelets. Mean intra-individual variability and total individual variability (standard deviation, SD) were calculated from the normalized log-transformed expression (RLD) of each transcript. Figure 9E shows the total individual variability of each transcript (x-axis) plotted against the intra-individual variability (y-axis) of each transcript for Cohort 1. Labeled points represent representative transcripts with low intra-individual variability and high total individual variability. [Figure 9F] This section compares the intra-individual variability and total variability of each transcript in platelets. Mean intra-individual variability and total individual variability (standard deviation, SD) were calculated from the normalized log-transformed expression (RLD) of each transcript. Figure 9F shows the total individual variability of each transcript (x-axis) plotted against the intra-individual variability (y-axis) of each transcript for Cohort 2 (Figure 9F). Labeled points represent representative transcripts with low intra-individual variability and high total individual variability. [Figure 10] This shows a partition analysis of platelet gene expression. The violin plot shows the distribution of the percentage of variance for each transcript (y axis) attributable to the indicated covariates (x axis). The width of the violins indicates the probability density of the transcript at each y value. The box plot shows the median and interquartile range, and outliers are plotted as individual points. For example, sex explains less than 50% of the variance for most transcripts, with the exception of the Y-chromosome genes EIF1AY, TMSB4Y, and UTY, which vary almost exclusively with sex. The plot on the far right shows that for most transcripts (more than 50%), inter-individual differences explain the majority of the variance. [Figure 11A]This table shows the transcripts with the highest expression in RNA-seq data from Cohort 1 (Figure 11A), and their associations with race, sex, or eQTLs reported in PRAX1 microarray data. Associations with FDR < 1e-4 are highlighted in pink. NS = not significant. [Figure 11B] This table shows the transcripts with the lowest intra-individual variability in RNA-seq data from Cohort 1 (Figure 11B), and their associations with race, sex, or eQTLs reported in PRAX1 microarray data. Associations with FDR < 1e-4 are highlighted in pink. NS = not significant. [Figure 11C] This table shows the transcripts with the highest total variation (Figure 11C) in the RNA-seq data of Cohort 1, and their associations with race, sex, or eQTLs reported in the PRAX1 microarray data. Associations with FDR < 1e-4 are highlighted in pink. NS = not significant. [Figure 11D] Figure 11D shows the transcripts with the highest repeatability (low intra-individual variability, high inter-individual variability) in RNA-seq data from Cohort 1, and their associations with race, sex, or eQTLs reported in PRAX1 microarray data. Associations with FDR < 1e-4 are highlighted in pink. NS = not significant. [Figure 12] This study shows that FDR is associated with repetition for transcripts with eQTLs reported in PRAX1. The x-axis represents the lowest reported log FDR (i.e., -125 = 10-125) for each transcript associated with the eQTL. The y-axis represents the 1-repetition rate (cohort 1) for each transcript. [Figure 13]This shows unsupervised clustering and heatmaps based on exon PSI for all 245 identified exon skipping events in platelets from 31 individuals in Cohort 1 at time 0 and 4 months. The histogram on the left shows the distribution of distances between all sample pairs, and the intensity of the blue indicates the similarity between sample pairs. Samples clustered as neighbors in the heatmap dendrogram reflect the samples with the highest similarity at the exon skipping level. The nearest neighbor self-pair is highlighted in yellow. The bars on the left are colored according to the sequencing batch or lane. [Figure 14] This report demonstrates PCR confirmation of SELP exon 14 skipping in platelets. Platelet RNA from five different individuals was reverse transcribed, and the cDNA was amplified with primers adjacent to SELP exon 14. The bands were extracted and sequenced by Sanger sequencing to confirm the sequences. [Figure 15A] This shows that SELP exon 14 skipping remains associated with rs6128 in disease. Box plots of SELP exon 14 mean PSI by rs6128 genotype inferred from RNA-seq in healthy and diseased samples reported in the NL cohort when analyzed total (Figure 15A). *p values adjusted for age, sex, smoking, hospital, and storage time. [Figure 15B] This shows that SELP exon 14 skipping remains associated with rs6128 in disease. Box plots of SELP exon 14 mean PSI by rs6128 genotype inferred from RNA-seq in healthy and diseased samples reported in the NL cohort when analyzed by disease (Figure 15B). Figure 15B shows diseases with multiple samples containing at least two different genotypes. *p-values adjusted for age, sex, smoking, hospital, and storage time. [Figure 16]This demonstrates that rs6128 directly controls exon 14 skipping of the P-selectin protein. Western blot analysis of P-selectin (antibody against c-terminal DYK tag) in HEK293 cells after transfection with rs6128 C / C or T / T SELP (CMV promoter) constructs. Representative blots from four independent experiments are shown. Below are bar graphs and standard error summaries of PSI calculated according to density metric analysis of the exon 14-containing band (upper band) divided by the sum of the upper and lower bands (aggregate). *Paired t-test, n=4 independent experiments. Note that while the anti-DYK antibody detected two distinct bands differing by approximately 19 kDa, exon 14 encodes 40 amino acids (<5 kDa). This difference is due to heavy glycosylation of exon 14, as deglycosylation of the solubilizer by PNGase made the band sizes indistinguishable. [Figure 17A] This section compares RNA-seq with genomic variant calling and population stratification. Figure 17A shows the allele frequencies (x-axis) of 641 variants tested for the presence of eQTLs called by RNA-seq in the NL cohort, compared to the allele frequencies reported in the Dutch genome database (GoNL (Genome of the Netherlands Consortium, Francioli LC, Menelaou A, et al. Nat Genet. 2014;46:818-825)). Pearson correlation = 0.93 (p<2.2e-16). Allele frequencies in Cohort 1 were also evaluated compared to allele frequencies reported in the 1000 Genome Database (Gibbs RA, et al. Nature. 2015;526:68-74), but are not shown here. For RNA-seq calls from Caucasian individuals in Cohort 1 compared to the European superpopulation genome, the correlation was 0.87 (p<2.2e-16). For RNA-seq calls from Black / African American individuals in Cohort 1 compared to the African superpopulation genome, the correlation was 0.89 (p<2.2e-16). [Figure 17B]This shows a comparison of RNA-seq with genomic variant calling and population stratification. Figure 17B shows a PCA analysis comparing the allele frequencies of 641 variants called by RNA-seq in Black / African American (AA), White, or NL cohorts in Cohort 1 with allele frequencies reported in GoNL and 1000 Genome Database hyper-subpopulations: East Asian (EAS), South Asian (SAS), Mixed American (AMR), European (EUR), or African (AFR). [Figure 17C] This section compares RNA-seq with genome variant calling and population stratification. Figure 17C shows MDS analysis of population structure for each individual in Cohort 1 and the NL cohort fixed to individuals from 1000 genomes (Purcell S, et al. Am J Hum Genet. 2007;81:559-75, and Chang CC, et al. Gigascience. 2015;4:7). 1994 variants were simultaneously identified in the RNA-seq cohort, and 1000 genomes were used for this analysis. The results show that Caucasian individuals in Cohort 1 and RNA-seq individuals in the NL cohort cluster almost exclusively with the EUR genome supersubpopulation, while Black / African American individuals in Cohort 1 and other / unknown RNA-seq individuals cluster with the AFR supersubpopulation, suggesting mixing. The top three MDS components are plotted. If the density of Caucasian individuals in Cohort 1, individuals in the NL cohort, and European individuals is high, include a zoom box for clarity. [Modes for carrying out the invention]
[0008] This disclosure can be more readily understood by referring to embodiments for carrying out the following inventions, the drawings and examples contained herein.
[0009] Before the methods and gene expression panels of the present invention are disclosed and described, it should be understood that they are not limited to specific synthesis methods or specific reagents unless otherwise specified, and are therefore subject to change. It should also be understood that the terms used herein are intended solely to describe specific embodiments and are not intended to limit them. Any methods and materials similar to or equivalent to those described herein may be used in carrying out or testing the present invention, examples of methods and materials are described herein.
[0010] Furthermore, unless otherwise expressly stated, none of the methods described herein are ever intended to be construed as requiring their steps to be performed in a particular order. Therefore, if a claim for a method does not actually enumerate the order in which its steps should be followed, or if it is not otherwise specifically stated in the claims or description that the steps are limited to a particular order, no order is ever intended to be inferred in any way. This also applies to any possible grounds for interpretation that could be unclear, including logical issues relating to the sequence or flow of steps, plain meanings arising from grammatical structure or punctuation, or the number or type of embodiments described in the specification.
[0011] All publications referenced herein are incorporated herein by reference to disclose and describe methods and / or materials in connection with the citation of such publications. Publications described herein are provided solely for the purpose of such disclosure prior to the filing date of this application. Nothing herein should be construed as admitting that the present invention does not have prior rights to such publication by prior art. Furthermore, publication dates provided herein may differ from actual publication dates, which may require independent verification.
[0012] definition As used herein and in the appended claims, the singular forms "a," "an," and "the" refer to multiple subjects unless the context clearly indicates otherwise.
[0013] As used herein, the word "or" means any one member of a given list, and also includes any combination of members of that list.
[0014] In this specification, a range may be expressed as “about” or “approximately” from one particular value and / or “about” or “approximately” to another particular value. Where such a range is expressed, further aspects include from one particular value and / or to another particular value. Similarly, where a value is expressed as an approximation, the preceding use of “about” or “approximately” will be understood to mean that the particular value forms further aspects. It will also be understood that each endpoint of a range is important in relation to the other endpoints and independently of the other endpoints. It will also be understood that there are several values disclosed herein, and each value is disclosed herein “about” that particular value, in addition to the value itself. For example, if the value “10” is disclosed, “about 10” is also disclosed. It will also be understood that each unit between two particular units is also disclosed. For example, if 10 and 15 are disclosed, 11, 12, 13, and 14 are also disclosed.
[0015] As used herein, the terms “optional” or “optionally” mean that the events or circumstances described below may or may not occur, and that the descriptions include both instances in which such events or circumstances occur and instances in which they do not.
[0016] As used herein, the term “sample” means a solution containing one or more molecules derived from tissue or organs from a subject, cells (either intracellular, directly collected from the subject, or maintained from culture or a cultured cell line), cell solubilizes (or solubilize fractions) or cell extracts, or cells or cellular material (e.g., polypeptides or nucleic acids) assayed as described herein. A sample may also be any bodily fluid or excretion containing cells or cellular components (e.g., blood, urine, feces, saliva, tears, bile).
[0017] As used herein, the term “subject” refers to the target of administration, e.g., human. Therefore, the subjects of the methods of this disclosure may be vertebrates such as mammals, fish, birds, reptiles, or amphibians. The term “subject” also includes domesticated animals (e.g., cats, dogs), livestock (e.g., cattle, horses, pigs, sheep, goats), and laboratory animals (e.g., mice, rabbits, rats, guinea pigs, fruit flies). In one embodiment, the subject is a mammal. In another embodiment, the subject is a human. This term is not intended to indicate a specific age or sex. Therefore, it is intended to encompass adult, child, adolescent, and neonatal subjects, as well as fetuses, regardless of whether they are male or female.
[0018] As used herein, the term “patient” refers to an object suffering from a disease or disorder. The term “patient” includes human and veterinary subjects. In some aspects of the methods of this disclosure, the “patient” is diagnosed as needing treatment for a disease, for example, prior to the administration step.
[0019] In some embodiments, terms such as “patient,” “subject,” and “individual” are used interchangeably herein and refer to any animal or its cells capable of accepting the methods described herein, whether in vitro or in sight. In some embodiments, patient, subject, or individual is human.
[0020] As used herein, the term “contains” may include “consisting of” and “essentially consisting of.”
[0021] As used herein, the terms “normal” or “healthy” refer to an individual, sample, or subject that is free from disease or impairment, or has no increased susceptibility to developing disease or impairment.
[0022] As used herein, the term “susceptibility” refers to the likelihood that an object will be clinically diagnosed as having the disease. For example, a human object with increased susceptibility to the disease may refer to a human object with increased likelihood that it will be clinically diagnosed as having the disease.
[0023] As used herein, the term “polypeptide” refers to any peptide, oligopeptide, polypeptide, gene product, expression product, or protein. A polypeptide consists of a sequence of amino acids. The term “polypeptide” encompasses naturally occurring or synthetic molecules. As used herein, the term “amino acid sequence” refers to a list of abbreviations, letters, characters, or words representing amino acid residues.
[0024] As used herein, the terms “peptide,” “polypeptide,” and “protein” are interchangeable and refer to compounds containing amino acid residues covalently linked by peptide bonds. A protein or peptide must contain at least two amino acids, and there is no limit to the maximum number of amino acids that can constitute a protein or peptide sequence. A polypeptide includes any peptide or protein containing two or more amino acids linked to one another by peptide bonds. As used herein, this term refers to both short chains, also commonly referred to in the art as peptides, oligopeptides, and oligomers, and longer chains, also commonly referred to in the art as proteins, of which many types exist. A “polypeptide” includes, for example, biologically active fragments, substantially homologous polypeptides, oligopeptides, homodimers, heterodimers, polypeptide variants, modified polypeptides, derivatives, analogs, and fusion proteins. Polypeptides include native peptides, recombinant peptides, synthetic peptides, or combinations thereof.
[0025] As used herein, the term “gene” refers to a region of DNA that codes for functional RNA or protein. “Functional RNA” refers to an RNA molecule that is not translated into protein. Generally, gene symbols are shown in italics, and protein symbols are shown in non-italics.
[0026] As used herein, the term “nucleic acid” refers to naturally occurring or synthetic oligonucleotides or polynucleotides that can hybridize with complementary nucleic acids by Watson-Crick base pairing, whether they are DNA, RNA, or DNA-RNA hybrids, single-stranded or double-stranded, sense or antisense. The nucleic acids of the present invention may also include nucleotide analogs (e.g., BrdU) and non-phosphodiester nucleoside bonds (e.g., peptide nucleic acids (PNA) or thiodiester bonds). Specifically, nucleic acids may include, but are not limited to, DNA, RNA, cDNA, gDNA, ssDNA, dsDNA, or any combination thereof.
[0027] Nucleic acids may each include pyrimidine bases and purine bases, preferably cytosine, thymine, and uracil, as well as any polymers or oligomers of adenine and guanine. In fact, the present invention intends any deoxyribonucleotide, ribonucleotide, or peptide nucleic acid component, as well as any chemical variants thereof, such as methylated, hydroxymethylated, or glucosylated forms of their bases. The polymers or oligomers may be heterogeneous or homogeneous in the composition and may be isolated from naturally occurring sources or produced artificially or synthetically. In addition, nucleic acids may be DNA or RNA, or mixtures thereof, and may exist permanently or transiently in single-stranded or double-stranded forms, including homo-double-stranded, hetero-double-stranded, and hybrid states.
[0028] An “oligonucleotide” or “polynucleotide” is a nucleic acid, or a compound that specifically hybridizes to a polynucleotide, having a length of at least 2, preferably at least 8, 15, or 25 nucleotides, but which may be up to 50, 100, 1000, or 5000 nucleotides. A polynucleotide contains a sequence of deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) or a mimic thereof, which may be isolated from natural sources, recombinantly produced, or artificially synthesized. A further example of a polynucleotide in the present invention may be a peptide nucleic acid (PNA). (See U.S. Patent No. 6,156,501, which is incorporated herein by reference in its entirety.) The present invention also encompasses situations in which non-traditional base pairings exist, such as Hoogsteen-type base pairings, which are specified in certain tRNA molecules and are assumed to be present in a triple helix. In this disclosure, “polynucleotide” and “oligonucleotide” are used interchangeably. Where a nucleotide sequence is represented herein by a DNA sequence (e.g., A, T, G, and C), it will be understood that this also includes the corresponding RNA sequence (e.g., A, U, G, C) in which "U" is replaced with "T".
[0029] As used herein, “polynucleotide” includes cDNA, RNA, DNA / RNA hybrids, antisense RNA, ribozymes, genomic DNA, synthetic forms, and mixed polymers, both sense and antisense strands, and may be chemically or biochemically modified to include non-natural, derivatized, synthetic, or semi-synthetic nucleotide bases. Modifications of wild-type or synthetic genes are also intended, including, but not limited to, deletions, insertions, substitutions, or fusions of one or more nucleotides into other polynucleotide sequences.
[0030] "Isolated polypeptide" or "purified polypeptide" means a polypeptide (or fragment thereof) that substantially does not contain the material with which the polypeptide normally associates in nature. The polypeptides or fragments thereof of the present invention can be obtained, for example, by extraction from natural sources (e.g., mammalian cells), by expression of recombinant nucleic acids encoding the polypeptide (e.g., intracellular or in a cell-free translation system), or by chemical synthesis of the polypeptide. In addition, polypeptide fragments can be obtained by any of these methods or by cleaving a full-length polypeptide.
[0031] "Isolated nucleic acid" or "purified nucleic acid" means DNA that does not contain genes adjacent to the gene in the naturally occurring genome of the organism in which the DNA of the present invention is induced. Therefore, this term includes, for example, recombinant DNA incorporated into a vector such as a self-replicating plasmid or virus, or recombinant DNA incorporated into the genomic DNA of a prokaryotic or eukaryotic organism (e.g., a transgene), or recombinant DNA existing as a separate molecule (e.g., cDNA or genome or cDNA fragment produced by PCR, restriction endonuclease digestion, or chemical or in vitro synthesis). This also includes recombinant DNA that is part of a hybrid gene encoding an additional polypeptide sequence. The term "isolated nucleic acid" also refers to RNA, for example, mRNA molecules encoded by an isolated DNA molecule, or chemically synthesized mRNA molecules, or mRNA molecules isolated from or substantially containing at least some cellular components, for example, other types of RNA molecules or polypeptide molecules.
[0032] "Specifically binding" means that an antibody recognizes its own antigen, physically interacts with it, and does not significantly recognize or interact with other antigens. Such an antibody may be a polyclonal or monoclonal antibody produced by techniques well known in the art.
[0033] "Probe," "primer," or oligonucleotide means a single-stranded DNA or RNA molecule with a defined sequence that can base-pair with a second DNA or RNA molecule ("target") containing a complementary sequence. The stability of the resulting hybrid depends on the degree of base pairing that occurs. The degree of base pairing is influenced by parameters such as the degree of complementarity between the probe and the target molecule, and the degree of stringency of the hybridization conditions. The degree of hybridization stringency is influenced by parameters such as temperature, salt concentration, and concentration of organic molecules such as formamide, and is determined by methods known to those skilled in the art. Probes or primers specific to a particular nucleic acid (e.g., genes and / or mRNA) have at least 80% to 90% sequence complementarity, preferably at least 91% to 95%, more preferably at least 96% to 99%, and most preferably 100% sequence complementarity with respect to the region of nucleic acid they hybridize. Probes, primers, and oligonucleotides can be labeled detectably either radioactively or non-radioactively by methods well known to those skilled in the art. Probes, primers, and oligonucleotides are used in methods including nucleic acid hybridization, such as nucleic acid sequencing, reverse transcription and / or nucleic acid amplification by polymerase chain reaction, single-strand conformational polymorphism (SSCP) analysis, restriction fragment polymorphism (RFLP) analysis, Southern hybridization, Northern hybridization, Insights hybridization, and electrophoretic mobility shift assays (EMSA).
[0034] As used herein, the term “probe” refers to an oligonucleotide (i.e., a sequence of nucleotides) that can hybridize to another oligonucleotide of interest, whether naturally occurring as found in purified restriction digests, or produced synthetically, recombinantly, or by PCR amplification. Probes may be single-stranded or double-stranded. Probes are useful for detecting, identifying, and isolating specific gene sequences.
[0035] The term "primer" refers to an oligonucleotide that can act as a synthetic starting point along the complementary strand when the conditions are favorable for the synthesis of the primer extension product. Synthetic conditions include the presence of four different deoxyribonucleotide triphosphates and at least one polymerization inducer, such as reverse transcriptase or DNA polymerase. These are present in a suitable buffer, which may contain cofactors or components that influence conditions such as pH at various suitable temperatures. Primers are preferably single-stranded sequences to optimize amplification efficiency, but double-stranded sequences can also be utilized.
[0036] "Specific hybridization" means that, under high stringency conditions, a probe, primer, or oligonucleotide recognizes and physically interacts (i.e., base-pairs) with substantially complementary nucleic acids, but does not substantially base-pair with other nucleic acids.
[0037] "High-stringency conditions" means conditions that enable hybridization equivalent to those achieved by using a DNA probe of at least 40 nucleotides in a buffer containing 0.5 M NaHPO4, pH 7.2, 7% SDS, 1 mM EDTA, and 1% BSA (fraction V) at a temperature of 65°C, or in a buffer containing 48% formamide, 4.8x SSC, 0.2 M Tris-Cl, pH 7.6, 1x Denhardt's solution, 10% dextran sulfate, and 0.1% SDS at a temperature of 42°C. Other conditions for high-stringency hybridization, such as PCR, Northern, Southern, or Insights hybridization, and DNA sequencing, are well known to those skilled in the field of molecular biology. (See, for example, F. Ausubel et al., Current Protocols in Molecular Biology, John Wiley & Sons, New York, NY, 1998).
[0038] The term “abnormal” is used to refer to an organism, tissue, cell, or component thereof that differs from an organism, tissue, cell, or component thereof that exhibits “normal” (expected) characteristics in at least one observable or detectable characteristic (e.g., age, treatment, time of day). A characteristic that is normal or expected in one cell type or tissue type may be abnormal in a different cell type or tissue type.
[0039] The term "amplification" refers to the operation of multiplying the copy number of a target nucleotide sequence present in a sample.
[0040] As used herein, the term “antibody” refers to an immunoglobulin molecule capable of specifically binding to a particular epitope on an antigen. Antibodies can be intact immunoglobulins derived from natural or recombinant sources, or they can be the immunoreactive portion of intact immunoglobulins. The antibodies of the present invention may exist in various forms, including, for example, polyclonal antibodies, monoclonal antibodies, intracellular antibodies ("intrabodies"), Fv, Fab, and F(ab)2, as well as single-chain antibodies (scFv), heavy-chain antibodies, such as camelid antibodies, synthetic antibodies, chimeric antibodies, and humanized antibodies (Harlow et al., 1999, Using Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory Press, NY; Harlow et al., 1989, Antibodies: A Laboratory Manual, Cold Spring Harbor, NY; Houston et al., 1988, Proc. Natl. Acad. Sci. USA 85:5879-5883; Bird et al., 1988, Science 242:423-426).
[0041] As used herein, “immunoassay” refers to any binding assay that detects and quantifies a target molecule using an antibody capable of specifically binding to that target molecule.
[0042] As used herein, the term “coding sequence” refers to a sequence or portion thereof of a nucleic acid or its complement that can be transcribed and / or translated to produce mRNA and / or polypeptides or fragments thereof. The coding sequence includes exons in genomic DNA or immature primary RNA transcripts, which are joined together by the cell’s biochemical mechanisms to provide mature mRNA. The antisense strand is the complement of such nucleic acid, from which the coding sequence can be inferred. In contrast, as used herein, the term “non-coding sequence” refers to a sequence or portion thereof of a nucleic acid or its complement that is not translated into amino acids in vivo, or where tRNA does not interact to arrange, or attempts to arrange, amino acids. Non-coding sequences include both intron sequences in genomic DNA or immature primary RNA transcripts and gene-associated sequences such as promoters, enhancers, and silencers.
[0043] As used herein, the terms “complementary” or “complementarity” are used in relation to polynucleotides (i.e., sequences of nucleotides) linked by base pairing rules. For example, the sequence “AGT” is complementary to the sequence “TCA”. Complementarity may be “partial,” meaning only some of the bases of the nucleic acid match according to base pairing rules. Alternatively, “complete” or “total” complementarity may exist between nucleic acids. The degree of complementarity between nucleic acid strands has a significant impact on the efficiency and strength of hybridization between nucleic acid strands. This is particularly important in amplification reactions, as well as in detection methods that rely on binding between nucleic acids.
[0044] As used herein, the term “diagnosis” refers to the determination of the presence of a disease or disorder. In some embodiments, methods are provided for making a diagnosis that enables the determination of the presence of a particular disease or disorder.
[0045] A "disease" is a state of animal health in which the animal is unable to maintain homeostasis, and if the disease does not improve, the animal's health continues to deteriorate. In contrast, a "disorder" in animals is a state of health in which the animal is able to maintain homeostasis, but the animal's health is less desirable than its health in the absence of the disorder. If left untreated, a disorder does not necessarily lead to a further deterioration of the animal's health.
[0046] As used herein, the term “code” refers to the inherent properties of a particular nucleotide sequence in a polynucleotide, such as a gene, cDNA, or mRNA, to serve as a template for the synthesis of other polymers and macromolecules in biological processes having either a defined nucleotide sequence (i.e., rRNA, tRNA, and mRNA) or a defined amino acid sequence, and the biological properties arising therefrom. Thus, a gene codes for a protein if the transcription and translation of the mRNA corresponding to that gene produces a protein in a cell or other biological system. Both the coding strand (whose nucleotide sequence is identical to the mRNA sequence and is typically presented in a sequence listing) and the non-coding strand (used as a template for the transcription of a gene or cDNA) can be said to code for a protein or other product of that gene or cDNA.
[0047] As used herein, the term “hybridization” is used in relation to the pairing of complementary nucleic acids. Hybridization and the intensity of hybridization (i.e., the intensity of association between nucleic acids) are influenced by factors such as the degree of complementarity between nucleic acids, the stringency of the conditions involved, the Tm of the hybrid formed, and the G:C ratio within the nucleic acid. A single molecule containing a pair of complementary nucleic acids within its structure is said to be “self-hybridized.” A single DNA molecule with internal complementarity can take on a variety of secondary structures, including loops, kinks, or coils in the case of long base pair extensions.
[0048] The term “Educational Materials,” as used herein, includes publications, records, figures, or any other medium of representation that can be used to communicate the usefulness of the nucleic acids, peptides, and / or compounds of the present invention within a kit for identifying, diagnosing, or mitigating or treating the various diseases or disorders listed herein. Optionally or alternatively, the educational materials may describe one or more methods for identifying, diagnosing, or mitigating a disease or disorder in the cells or tissues of interest. The educational materials of a kit may, for example, be attached to a container containing one or more components of the present invention, or shipped together with a container containing one or more components of the present invention. Alternatively, the educational materials may be shipped separately from the container with the recipient's intention to use the educational materials and components collaboratively.
[0049] The term "isolated" means that something has been altered or removed from its natural state. For example, a nucleic acid or peptide that is naturally present in a living animal is not "isolated," but the same nucleic acid or peptide that has been partially or completely separated from its natural coexisting substances is "isolated." Isolated nucleic acids or proteins can exist in a substantially purified form or in a non-natural environment, such as a host cell.
[0050] As used herein, the term “labeling” refers to a detectable compound or composition that is directly or indirectly conjugated to a probe to produce a “labeled” probe. The labeling may be detectable on its own (e.g., radioisotope labeling or fluorescent labeling), or, in the case of enzymatic labeling, may catalyze a chemical change in a detectable substrate compound or composition (e.g., avidin-biotin). In some cases, primers can be labeled to detect PCR products.
[0051] The terms “microarray” and “array” broadly refer to “DNA microarrays,” “DNA chips,” “protein microarrays,” and “protein chips,” and encompass all solid supports recognized in the art, as well as all methods recognized in the art for immobilizing nucleic acid, peptide, and polypeptide molecules. Preferred arrays typically include multiple different nucleic acid or peptide probes bound to the surface of a substrate at different known positions. These arrays, also known as “microarrays” or colloquially “chips,” are commonly described in the art, for example, in U.S. Patents No. 5,143,854, 5,445,934, 5,744,305, 5,677,195, 5,800,992, 6,040,193, 5,424,186, and Fodor et al., 1991, Science, 251:767-777, each of which is incorporated by reference as a whole for all purposes. Arrays can generally be produced using a variety of techniques, such as mechanical synthesis or optically oriented synthesis, which incorporate a combination of photolithography and solid-phase synthesis. Techniques for synthesizing these arrays using mechanical synthesis methods are described, for example, in U.S. Patents No. 5,384,261 and No. 6,040,193, which are incorporated herein by reference in their entirety for all purposes. Planar array surfaces are preferred, but arrays may be fabricated on surfaces of substantially any shape or even a variety of surfaces. Arrays may be nucleic acids on beads, gels, polymer surfaces, fibers such as optical fibers, glass, or any other suitable substrate. (See U.S. Patents 5,770,358, 5,789,162, 5,708,153, 6,040,193, and 5,800,992, which are incorporated herein by reference in their entirety for all purposes.) The array may be packaged in a manner that enables diagnostic use, or may be made into an integrated device, see, for example, U.S. Patents 5,856,174 and 5,922,591, which are incorporated herein by reference in their entirety for all purposes.Arrays are commercially available from companies such as Affymetrix (Santa Clara, Calif.) and Applied Biosystems (Foster City, Calif.) and are intended for a variety of purposes, including genotyping, diagnosis, mutation analysis, marker expression, and gene expression monitoring of various eukaryotes and prokaryotes. The number of probes on a solid support can be varied by changing the size of the individual features. In some embodiments, the feature size is 20 × 25 square microns, while in other embodiments, the features may be, for example, 8 × 8, 5 × 5, or 3 × 3 square microns, resulting in approximately 2,600,000, 6,600,000, or 18,000,000 individual probe features.
[0052] Assays for amplifying known sequences are also disclosed. For example, primers for PCR may be designed to amplify a region of a sequence. In the case of RNA, a first reverse transcriptase step may be used to generate double-stranded DNA from single-stranded RNA. Arrays may be designed to detect sequences from an entire genome, or from one or more regions of a genome, such as a selected region of a genome, e.g., a region encoding a protein or RNA of interest, or from conserved regions from multiple genomes, or from multiple genomes, arrays, and methods for performing genetic analysis using arrays are described in Cutler, et al., 2001, Genome Res. 11(11):1913-1925 and Warrington, et al., 2002, Hum Mutat 19:402-409, and U.S. Patent Publication No. 20030124539, respectively, which are incorporated herein by reference in their entirety.
[0053] As used herein, the term “polymerase chain reaction” (PCR) refers to the KBMullis methods (U.S. Patents 4,683,195, 4,683,202, and 4,965,188, incorporated herein by reference), which describe methods for increasing the concentration of a segment of a target sequence in a genomic DNA mixture without cloning or purification. This process for amplifying a target sequence consists of introducing a large excess of two oligonucleotide primers into a DNA mixture containing the desired target sequence, and then introducing the precise sequence of thermal cycling in the presence of DNA polymerase. These two primers are complementary to the respective strands of the double-stranded target sequence. To bring about amplification, the mixture is denatured, and then the primers are annealed to their complementary sequences within the target molecule. After annealing, the primers are extended with polymerase to form a new complementary strand pair. By repeating the denaturation step, primer annealing step, and polymerase extension step multiple times (i.e., denaturation, annealing, and extension constitute one “cycle,” and there may be many “cycles”), a highly concentrated amplified segment of the desired target sequence can be obtained. The length of the amplified segment of the desired target sequence is determined by the relative positions of the primers to each other, and therefore this length is a controllable parameter. Based on the repeated mode of this process, this method is referred to as “polymerase chain reaction” (hereinafter, PCR). The desired amplified segment of the target sequence becomes the dominant sequence (in terms of concentration) in the mixture, and therefore they are said to be “PCR amplified.” As used herein, the terms “PCR product,” “PCR fragment,” “amplified product,” or “amplicon” refer to a mixture of compounds obtained after two or more cycles of the PCR steps of denaturation, annealing, and extension have been completed. These terms encompass the case where amplification of one or more segments of one or more target sequences is present.
[0054] As used in relation to organisms, tissues, cells, or components thereof, the term “abnormal” refers to an organism, tissue, cell, or component thereof that differs from the organism, tissue, cell, or component thereof that exhibits the respective (expected) characteristics in at least one observable or detectable characteristic (e.g., age, treatment, time of day, etc.). A characteristic that is normal or expected in one cell type or tissue type may be abnormal in a different cell type or tissue type.
[0055] The term "amplification" refers to the operation of multiplying the copy number of a target nucleotide sequence present in a sample.
[0056] Platelets are readily available and abundant blood cells that are increasingly being used in gene expression studies. Like nucleated cells, platelets have a diverse portfolio of RNA, including coding mRNA, small non-coding RNA, and lncRNA (Rowley JW, et al. Blood. 2011;118:e101-e111, Bray PF, et al. BMC Genomics. 2013;14:1, and Gnatenko DV, et al. Blood. 2003;101:2285-931-3). Furthermore, their anucleated nature is advantageous for gene expression studies compared to nucleated cells. For example, ex vivo processing of nucleated cells (cell isolation method, processing time, buffers, etc.) can immediately affect the expression of thousands of transcripts (Beliakova-Bethell N, et al. Cytometry A. 2014;85:94-104, Bhattacharjee J, et al. F1000Research. 2017;6:2045, and Baechler EC, et al. Genes Immun. 2004;5:347-534-6). On the other hand, platelets are transcriptionally unaffected by isolation (Angenieux C, et al. PLoS One. 2016;11:e0148064, and Best MG, et al. Cancer Cell. 2015;28:666-676), allowing for the capture of their innate in vivo gene expression signatures. These attractive characteristics make them an excellent choice for RNA diagnostics and gene expression studies.
[0057] Platelets are used in GWAS studies focusing on RNA abundance, gene phenotyping studies (Kondkar AA, et al. J Thromb Haemost. 2010;8:369-78, and Edelstein LC, et al. Nat Med. 2013;19:1609-16), diagnostic studies (Best MG, et al. Cancer Res. 2018;78:3407-3412), and differential expression studies (Schubert S, et al. Blood. 2014;124:493-502) to elucidate the mediators of platelet reactivity in healthy and diseased states. Genetic modifiers of RNA abundance in platelets, known as quantitative phenotypic loci (eQTLs), have also been described (Simon LM, et al. Am J Hum Genet. 2016;98:883-97, and Kong X, et al. Thromb Haemost. 2017;117:962-970). eQTLs are DNA sequence variants associated with gene expression that affect nearby (cis-) genes or distant (trans-) genes in a cell-type-specific manner. eQTLs are particularly important in genetic research because they provide intermediate and mechanistic links between phenotype and gene association.
[0058] Beyond RNA abundance, platelets and megakaryocytes (Schubert S, et al. Blood. 2014;124:493-502) are known to possess alternative structural features of RNA, such as alternative start and stop sites, and alternative splicing, which diversify the transcriptome and proteome (Nassa G, et al. Sci Rep. 2018;8:498, and Schwertz H, et al. J Exp Med. 2006;203:2433-40) and alter cellular function. In platelets, RNA splicing is induced upon activation, thereby regulating the expression of functional proteins (Nassa G, et al. Sci Rep. 2018;8:498, and Denis MM, et al. Cell. 2005;122:379-91). Genetic variants called splicing QTLs (sQTLs) may also affect the fundamental activation-dependent RNA splicing levels in platelets. However, sQTLs in platelets remain unexplained.
[0059] Other major knowledge gaps exist regarding RNA abundance and structure in platelets. Most platelet studies are cross-sectional, examining gene expression at a single time point. Moreover, gene expression can change over time, both between and within individuals. Hormonal changes, circadian rhythms, inflammation, diet, and aging are examples of environmental factors that can alter gene expression in healthy individuals (Bryois J, et al. Genome Res. 2017;27:545-552, Waaseth M, et al. BMC Med Genomics. 2011;4:29, and Arnardottir ES, et al. Sleep. 2014;37:1589-600). Such normal changes in gene expression can mask the ability of differential gene expression studies, diagnostic studies, and genetic studies to detect signals and can complicate their analysis. Therefore, understanding the comparison between intra-individual and inter-individual variability in gene expression is crucial for designing and interpreting gene expression studies and can be used to prioritize candidates in gene research.
[0060] Regarding genetic studies, several reports suggest using repetition to identify eQTL genes (Barendse W. BMC Genomics. 2011;12:232, Carlborg O, et al. Bioinformatics. 2004;21:2383-93, and Hoffman GE, Schadt EE. BMC Bioinformatics. 2016;17:48321-23). In vivo repetition can only be calculated from multiple samples from the same individual and refers to the ratio of variation resulting from a comparison of inter-individual variability and intra-individual variability (Lessells CM, Boag PT. Auk. 1987;104:116-121). Frankly speaking, repetition imposes an upper limit on heritability in a broad sense (Dohm MR.Funct Ecol.2002;16:273-280), and is insufficient to detect heritable gene signals when inter-individual differences are not repetitive, due to low inter-individual variability and / or high intra-individual variability. For this reason, it is recommended to measure the repetition of traits before performing GWAS21. Carlborg et al. (Carlborg O, et al. Bioinformatics.2004;21:2383-93) found that censoring mouse eQTL data for repetition is an effective way to prioritize transcripts with a high a priori probability of successful eQTL identification. Hoffman et al. (Hoffman GE, Schadt EE. BMC Bioinformatics.2016;17:483) also demonstrated the potential use of intra-individual technical variability to narrow down candidates and facilitate eQTL prediction, although this study used single-time point replication. In summary, these studies suggest that longitudinal analysis of gene expression can facilitate the discovery of promising eQTL (and sQTL) genes. Surprisingly, longitudinal analysis of gene expression is scarce in primary cells from healthy individuals and absent in platelets. The repeatability of temporal alternative splicing has not been established for any primary cell type. A method for constructing transcriptome-wide expression profiles of biological samples, which has not previously existed in the art, is disclosed herein.For example, a method for preparing a transcriptome-wide expression profile of a biological sample is disclosed herein, comprising: a) measuring the expression level of one or more genes present in a first biological sample, wherein the measurement includes RNA sequencing; b) determining the expression level of a gene from the measured expression level obtained in step a); c) combining the results of step a) and b) to generate a transcriptome-wide expression profile; and d) providing the transcriptome-wide expression profile as a dataset, wherein the first biological sample consists of isolated platelets.
[0061] This specification discloses findings that platelet transcriptomes remain largely stable over four years in healthy individuals, providing longitudinal criteria for disease diagnosis. Platelet gene expression signatures have recently been used to accurately classify and diagnose cancer. Changes in platelet transcriptomes compared to healthy individuals have also been associated with numerous other diseases (sepsis, influenza infection, lupus, and acute myocardial infarction), increasing the potential of platelet gene expression signatures for the diagnosis and classification of various diseases.
[0062] Excessive noise is known to limit the use of gene expression signatures in diagnostics, including platelet gene expression-based diagnostics. Most gene expression-based diagnostics compare gene expression in diseased individuals to healthy controls. Another approach to gene expression-based diagnostics that minimizes environmental noise and internally corrects for inter-individual (i.e., genetic) variability is longitudinal sampling, which compares the gene signature of individuals with disease to the gene signature of the same individuals at a time when they were healthy. This approach relies on the assumption that gene signatures are relatively stable over time in healthy individuals, or that changes over time in healthy individuals are known and consistent. However, the degree of variability in gene expression over time is unknown for most cell types, including platelets. Reference gene signatures for longitudinal studies in platelets are currently unavailable. To address this issue, we evaluated gene expression in platelets of healthy individuals over up to four years and identified transcriptome-wide expression profiles in healthy individuals. This profile includes the most stable and least stable genes and splicing events over time in platelets from healthy individuals. The transcriptome-wide expression profile can be used as a criterion for platelet differential gene expression analysis and diagnosis by providing expected variability within healthy individuals over time. Deviations from this criterion may indicate disease.
[0063] As described herein, transcriptome-wide expression profiles can be used as a criterion for platelet differential gene expression analysis and diagnosis by providing expected variations over time within a single healthy individual or multiple healthy individuals. In some embodiments, the transcriptome-wide expression profiles disclosed herein can be used as a criterion for platelet differential gene expression analysis and diagnosis by providing expected variations over time within a single individual or multiple individuals with disease. In some embodiments, the transcriptome-wide expression profiles disclosed herein can be used as a criterion for platelet differential gene expression analysis and diagnosis by providing expected variations within a single individual or multiple individuals, where a single individual or multiple individuals are healthy, have disease, or are a combination of healthy and diseased. Deviations from this criterion may be useful as a periodic screening for distinguishing between healthy and diseased states. For example, intra-individual and inter-individual variations of the least and most variable genes expressed in platelets, as well as repeatable measurement results, are provided herein. These can be used as a benchmark for future disease studies, particularly for longitudinal evaluation of gene expression, i.e., when disease is suspected in a healthy individual, comparing the gene expression of a sample derived from a healthy individual with the gene expression of a sample longitudinally isolated from the same individual. The repetition of gene expression in platelets for predicting the presence of eQTLs is demonstrated herein. This list can be used to narrow down the number of genetic variants that may be useful for inclusion in disease diagnostic signatures.
[0064] The advantages of the methods disclosed herein include, but are not limited to, providing more accurate “healthy” criteria for diagnostic testing and enabling individual patients to be tested over time and serve as patient-specific criteria. The methods disclosed herein can also be used as standards for other diagnostic tests, such as genetic testing. Furthermore, the use of longitudinal criteria can reduce the costs of developing disease diagnostic signatures that would otherwise require longitudinal sampling of gene expression in additional healthy individuals.
[0065] Transcriptome-wide expression profiles and methods for generating them are disclosed herein. A transcriptome-wide expression profile obtained from a platelet sample is disclosed herein. A transcriptome-wide expression profile of platelets is disclosed herein. For example, a transcriptome-wide expression profile obtained from a healthy individual is disclosed herein. This transcriptome-wide expression profile is obtained by sampling the same individual over a period of time (4 years), thereby allowing the transcriptome-wide expression profile to explain intra-individual variation over time. In some embodiments, the transcriptome-wide expression profile can be obtained from individuals having a specific disease, condition, disorder, or injury, or a specific set of diseases, conditions, disorders, or injuries. Other reference gene signatures rely on sampling many different people at a single point in time to find the mean signature. Longitudinal methods allow for the investigation of natural variations in the expression of individual genes in a single individual, thereby enabling the identification of the most stable and least stable genes over time.
[0066] A transcriptome-wide expression profile can be a gene signature or gene expression signature, which is a group of single or combined genes within a cell that has a distinctive gene expression pattern resulting from a modified or non-modified biological process or pathogenic disease condition. The clinical applications of transcriptome-wide expression profiles are categorized into prognostic signatures, diagnostic signatures, and predictive signatures. The phenotypes that can theoretically be defined by transcriptome-wide expression profiles range from those that predict the survival or prognosis of individuals with a disease, to those used to distinguish different subtypes of the disease, to those that predict the activation of specific pathways. In summary, the disclosed method addresses the need for transcriptome-wide expression profiles for longitudinal studies of platelet gene signatures. The method disclosed herein encompasses platelet RNA gene signatures or maps for disease diagnosis and can therefore be used, for example, as a highly sensitive analytical method for analyzing tumor components in bodily fluids such as blood in liquid biopsies.
[0067] This specification discloses a method for using transcriptome-wide expression profiles (intra-individual and inter-individual variability) as a criterion for platelet differential gene expression analysis and diagnosis by providing the expected amount of variation over time within healthy individuals.
[0068] This specification discloses a method for comparing gene expression in a sample derived from a healthy individual with that of a sample longitudinally isolated from the same individual, when a disease is suspected in a healthy individual, using transcriptome-wide expression profiles (intra-individual and inter-individual variability).
[0069] A method for predicting the presence of quantitatively expressible trait loci (eQTLs) and splice-quantitative trait loci (sQTLs) over time is disclosed herein, using transcriptome-wide expression profiles (intra-individual and inter-individual variability).
[0070] A method for determining eQTL gene expression using transcriptome-wide expression profiles (intra-individual and inter-individual variability) is disclosed herein.
[0071] Method for constructing transcriptome-wide expression profiles A method for preparing a transcriptome-wide expression profile of a biological sample is disclosed herein, comprising: a) measuring the expression level of one or more genes present in a first biological sample, wherein the measurement includes RNA sequencing; b) determining the expression level of a gene from the measured expression level obtained in step a); c) combining the results of step a) and b) to generate a transcriptome-wide expression profile; and d) providing the transcriptome-wide expression profile as a dataset, wherein the first biological sample consists of isolated platelets. In some embodiments, the method can be repeated at least once. In some embodiments, the method can be repeated with a second biological sample, the second biological sample consisting of isolated platelets. In some embodiments, the first and second biological samples can be obtained from the same or different subjects. In some embodiments, steps a) to d) can be repeated with each biological sample. In some embodiments, the transcriptome-wide expression profiles of subjects can be compared. In some embodiments, the first and second biological samples may be obtained from the same subject at the first and second time points. In some embodiments, the first and second biological samples may be obtained from different subjects, and further include obtaining additional biological samples from the subject, which may be obtained at different time points. In some embodiments, the first and second time points may be different time points. In some embodiments, a transcriptome-wide expression profile from the first time point may be compared with a transcriptome-wide expression profile from the second time point. In some embodiments, the method further includes repeating these steps until a valid transcriptome-wide expression profile can be identified. In some embodiments, the expression level may be measured using a device selected from the group consisting of microarrays, bead arrays, liquid arrays, and nucleic acid sequences.
[0072] Methods for preparing transcriptome-wide expression profiles of biological samples are disclosed herein. In some embodiments, the method may include measuring the expression levels of one or more genes present in a first biological sample. In some embodiments, the measurement may include RNA sequencing. In some embodiments, the method may include determining the expression levels of genes from the measured expression levels obtained in the first biological sample, and such measurement may include RNA sequencing. In some embodiments, the method may include generating a transcriptome-wide expression profile by combining the results of measuring the expression levels of genes present in the first biological sample with the results of determining the expression levels of genes from the measured expression levels. In some embodiments, the method may further provide the transcriptome-wide expression profiles as a dataset. In some embodiments, the first biological sample may consist of isolated platelets. In some embodiments, the method may be repeated at least once. In some embodiments, the method may be repeated with a second biological sample, the second biological sample consisting of isolated platelets. In some embodiments, the first and second biological samples may be obtained from the same or different subjects. In some embodiments, steps a) to d) can be repeated for each biological sample. In some embodiments, the transcriptome-wide expression profiles of the subjects can be compared. In some embodiments, the first and second biological samples can be obtained from the same subject at the first and second time points. In some embodiments, the first and second biological samples can be obtained from different subjects, and further include obtaining additional biological samples from the subject, which are obtained at different time points. In some embodiments, the first and second time points can be different time points. In some embodiments, the transcriptome-wide expression profile from the first time point can be compared with the transcriptome-wide expression profile from the second time point.In some embodiments, the method further includes repeating those steps until an effective transcriptome-wide expression profile can be identified. In some embodiments, the expression level can be measured using a device selected from the group consisting of microarrays, bead arrays, liquid arrays, and nucleic acid sequences.
[0073] A method for preparing a transcriptome-wide expression profile of a biological sample is disclosed herein. The method for preparing a transcriptome-wide expression profile of a biological sample is disclosed herein, comprising: a) measuring the expression level of one or more genes present in a first biological sample, wherein the measurement includes RNA sequencing; b) determining the expression level of a gene from the measured expression level obtained in step a); c) combining the results of step a) and b) to generate a transcriptome-wide expression profile; and d) providing the transcriptome-wide expression profile as a dataset, wherein the first biological sample consists of isolated platelets. In some embodiments, steps a) to d) can be repeated for each biological sample. In some embodiments, the transcriptome-wide expression profiles of subjects can be compared. In some embodiments, the first and second biological samples can be obtained from the same subject at a first and second time point. In some embodiments, the first and second biological samples may be obtained from different subjects, and further include obtaining additional biological samples from the subjects, the additional biological samples being obtained at different time points. In some embodiments, the first and second time points may be different time points. In some embodiments, the transcriptome-wide expression profile from the first time point can be compared with the transcriptome-wide expression profile from the second time point. In some embodiments, the method further includes repeating those steps until a valid transcriptome-wide expression profile can be identified. In some embodiments, the expression levels may be measured using a device selected from the group consisting of microarrays, bead arrays, liquid arrays, and nucleic acid sequences.
[0074] A method for identifying gene expression differences between two transcriptome-wide expression profiles is also disclosed herein, comprising determining one or more variations in the platelet transcriptome of a subject using a method for constructing a transcriptome-wide expression profile of a biological sample, wherein the method comprises: a) measuring the expression level of one or more genes present in a first biological sample, wherein the measurement includes RNA sequencing; b) determining the expression level of a gene from the measured expression level obtained in step a); c) combining the results of step a) and step b) to generate a transcriptome-wide expression profile; and d) providing the transcriptome-wide expression profile as a dataset, wherein the first biological sample consists of isolated platelets, and the method is performed at different time points, thereby identifying and providing gene expression differences and comparing the time-dependent changes from the subject to a reference, wherein the method is performed at different time points, thereby identifying and comparing gene expression differences. In some embodiments, the method can be repeated at least once. In some embodiments, the method can be repeated with a second biological sample, the second biological sample consisting of isolated platelets. In some embodiments, the first and second biological samples may be obtained from the same or different subjects. In some embodiments, steps a) to d) may be repeated for each biological sample. In some embodiments, the transcriptome-wide expression profiles of the subjects may be compared. In some embodiments, the first and second biological samples may be obtained from the same subject at a first and second time point. In some embodiments, the first and second biological samples may be obtained from different subjects, further including obtaining additional biological samples from the subjects, the additional biological samples being obtained at different time points. In some embodiments, the first and second time points may be different time points. In some embodiments, the transcriptome-wide expression profile from the first time point may be compared with the transcriptome-wide expression profile from the second time point.In some embodiments, the method further includes repeating those steps until an effective transcriptome-wide expression profile can be identified. In some embodiments, the expression level can be measured using a device selected from the group consisting of microarrays, bead arrays, liquid arrays, and nucleic acid sequences.
[0075] A method for measuring time-dependent gene expression differences in platelets derived from a subject is further disclosed herein, comprising determining one or more variations in the platelet transcriptome of the subject using a method for producing a transcriptome-wide expression profile of a biological sample, wherein the method comprises: a) measuring the expression level of one or more genes present in a first biological sample, wherein the measurement includes RNA sequencing; b) determining the gene expression level from the measured expression level obtained in step a); c) generating a transcriptome-wide expression profile by combining the results of step a) and step b); and d) providing the transcriptome-wide expression profile as a dataset, wherein the first biological sample consists of isolated platelets, the method is performed at different time points, thereby identifying, providing, and comparing the time-dependent changes from the subject to a reference. In some embodiments, the method can be repeated at least once. In some embodiments, the method can be repeated with a second biological sample, the second biological sample consisting of isolated platelets. In some embodiments, the first and second biological samples may be obtained from the same or different subjects. In some embodiments, steps a) to d) may be repeated for each biological sample. In some embodiments, the transcriptome-wide expression profiles of the subjects may be compared. In some embodiments, the first and second biological samples may be obtained from the same subject at a first and second time point. In some embodiments, the first and second biological samples may be obtained from different subjects, further including obtaining additional biological samples from the subjects, the additional biological samples being obtained at different time points. In some embodiments, the first and second time points may be different time points. In some embodiments, the transcriptome-wide expression profile from the first time point may be compared with the transcriptome-wide expression profile from the second time point.In some embodiments, the method further includes repeating those steps until an effective transcriptome-wide expression profile can be identified. In some embodiments, the expression level can be measured using a device selected from the group consisting of microarrays, bead arrays, liquid arrays, and nucleic acid sequences.
[0076] A biological sample can be any biological tissue or bodily fluid. In many cases, the sample is a “clinical sample,” which is a patient-derived sample. A biological sample can contain any biological material suitable for detecting a desired biomarker and may include cellular and non-cellular material obtained from an individual. Biological samples can be obtained by appropriate methods, such as blood collection, fluid collection, or biopsy. Examples of such samples include, but are not limited to, blood, lymph, urine, gynecological fluids, biopsies, amniotic fluid, and smears. Samples that are essentially liquid are referred to herein as “bodily fluids.” Body samples can be obtained from a patient by a variety of techniques, including, for example, scraping or swabbing an area or using a needle to aspirate bodily fluids. Methods for collecting various body samples are well known in the art. In many cases, the sample is a “clinical sample,” i.e., a patient-derived sample. Such samples include, but are not limited to, body fluids that may or may not contain cells, such as blood (e.g., whole blood, serum, or plasma), urine, saliva, tissue, or microneedle biopsy samples, and record-based samples with a known history of diagnosis, treatment, and / or outcome. In some embodiments, the biological sample may include blood cells. In some embodiments, the sample (or biological sample) may be tissue, blood, serum, or plasma. In some embodiments, the biological sample may include platelets. In some embodiments, the sample may be isolated platelets. In some embodiments, the biological sample may be derived from a healthy subject. In some embodiments, the biological sample may be derived from a subject having one or more diseases, conditions, disorders, or injuries.
[0077] In some embodiments, the method may further include extracting RNA from isolated platelets.
[0078] In some embodiments, the transcriptome-wide expression profile from a first time point can be compared with the transcriptome-wide expression profile from a second time point.
[0079] In some embodiments, the method may further include repeating those steps until an effective transcriptome-wide expression profile is identified.
[0080] In some embodiments, expression levels can be measured using a device selected from the group consisting of microarrays, bead arrays, liquid arrays, and nucleic acid sequences.
[0081] Identification of markers or biomarkers Methods for identifying disease-related biomarkers are disclosed herein. In some embodiments, the method may include: a) obtaining or having obtained a sample from a subject having a disease, wherein the sample contains isolated platelets; b) sequencing the isolated platelets in the sample; c) determining the gene expression of the sequence in step b); d) repeating steps a), b), and c) at different time points; and e) comparing the gene expression at different time points to identify, compare, a biomarker associated with the gene of the subject if the change in gene expression is at least two standard deviations. In some embodiments, a biomarker associated with the gene of the subject can be identified if the change in gene expression is at least three standard deviations. In some embodiments, the method for identifying disease-related biomarkers may be used for diagnosing disease, assessing disease severity, and assessing recovery from disease by detecting differentially expressed biomarkers in a biological sample obtained from a subject compared to a control or reference sample.
[0082] A method for producing a transcriptome-wide expression profile of a biological sample is disclosed herein, comprising: a) measuring the expression level of one or more genes present in a first biological sample, wherein the measurement includes RNA sequencing; b) determining the expression level of a gene from the measured expression level obtained in step a); c) combining the results of step a) and b) to generate a transcriptome-wide expression profile; and d) providing the transcriptome-wide expression profile as a dataset, wherein the first biological sample consists of isolated platelets, wherein the method can be used for the diagnosis of disease, assessment of disease severity, and assessment of recovery from disease by detecting differentially expressed biomarkers in a biological sample obtained from a subject compared to a control or reference sample.
[0083] This specification discloses transcriptome-wide expression profiles that can be used for disease diagnosis, disease severity assessment, and disease recovery assessment by detecting differentially expressed biomarkers in biological samples obtained from subjects compared to control or reference samples.
[0084] Transcriptome-wide expression profiles generated by methods disclosed herein can be used to diagnose diseases, assess disease severity, and evaluate recovery from diseases by detecting differentially expressed biomarkers in biological samples obtained from subjects compared to control or reference samples.
[0085] Methods that can be used to identify differences between transcriptome-wide expression profiles are disclosed herein. In some embodiments, these methods may include identifying differences between healthy transcriptome-wide expression profiles and non-healthy (e.g., diseased) transcriptome-wide expression profiles. In some embodiments, a method for identifying gene expression differences between two transcriptome-wide expression profiles includes determining one or more variations in the target platelet transcriptome using a method for producing a transcriptome-wide expression profile of a biological sample, wherein the method includes a) measuring the expression levels of one or more genes present in a first biological sample, wherein the measurement includes RNA sequencing; b) determining the gene expression levels from the measured expression levels obtained in step a); c) generating a transcriptome-wide expression profile by combining the results of step a) and step b); and d) providing the transcriptome-wide expression profile as a dataset, wherein the first biological sample consists of isolated platelets, the method is performed at different time points, thereby identifying and providing gene expression differences, and comparing such time-dependent changes from the target to a reference, the method is performed at different time points, thereby identifying and comparing gene expression differences. In some embodiments, the method can be repeated at least once. In some embodiments, the method can be repeated with a second biological sample, the second biological sample consists of isolated platelets. In some embodiments, the first and second biological samples may be obtained from the same or different subjects. In some embodiments, steps a) to d) may be repeated for each biological sample. In some embodiments, the transcriptome-wide expression profiles of the subjects may be compared. In some embodiments, the first and second biological samples may be obtained from the same subject at the first and second time points.In some embodiments, the first and second biological samples may be obtained from different subjects, and further include obtaining additional biological samples from the subjects, the additional biological samples being obtained at different time points. In some embodiments, the first and second time points may be different time points. In some embodiments, the transcriptome-wide expression profile from the first time point can be compared with the transcriptome-wide expression profile from the second time point. In some embodiments, the method further includes repeating those steps until a valid transcriptome-wide expression profile can be identified. In some embodiments, the expression levels may be measured using a device selected from the group consisting of microarrays, bead arrays, liquid arrays, and nucleic acid sequences.
[0086] A method for detecting differentially expressed markers by nucleic acid microarrays is disclosed herein. In some embodiments, the method can be used to detect and measure the levels of differentially expressed marker expression products, such as RNA and proteins, to measure the levels of one or more differentially expressed marker expression products, as is known to those skilled in the art.
[0087] Methods for detecting or measuring gene expression using methods focused on cellular components (cytoscopy) or methods focused on extracellular components (fluidography) are disclosed herein. Because gene expression involves the regular production of several different molecules, cytoscopy or fluidography can be used to detect or measure a variety of molecules, including RNA, proteins, and several molecules that may be modified as a result of protein function. Typical nucleic acid-focused diagnostic methods include amplification techniques, e.g., PCR and RT-PCR (including quantitatively modified versions), as well as hybridization techniques, e.g., insight hybridization, microarrays, and blotting. Typical protein-focused diagnostic methods include binding techniques, e.g., ELISA, immunohistochemistry, microarrays, and functional techniques, e.g., enzyme assays.
[0088] Methods for identifying transcriptome variations in a subject are disclosed herein. In some embodiments, the method may include: a) obtaining or having obtained a sample from a subject, wherein the sample contains isolated platelets; b) sequencing the transcriptome of the isolated platelets in the sample; c) determining the gene expression of the transcriptome in step b; d) repeating steps a), b), and c) at least once at different time points; and e) comparing the gene expression determined in steps c) and d) to identify transcriptome variations in the subject if the change in gene expression is at least two standard deviations. In some embodiments, transcriptome variations in a subject can be identified if the change in gene expression is at least three standard deviations. In some embodiments, the variations may be biomarkers. In some embodiments, biomarkers may be used for disease diagnosis, assessment of disease severity, and assessment of recovery from disease by detecting differentially expressed biomarkers in a biological sample obtained from a subject compared to a control or reference sample. In some embodiments, the presence of sequence variations can be detected in platelet quantitative trait loci (eQTL) genes. In some embodiments, the presence of sequence variations may be in platelet splice quantitative trait loci (sQTL) genes.
[0089] Methods for preparing datasets are disclosed herein. In some embodiments, the method may include: a) obtaining or having obtained samples from two or more subjects, wherein the samples contain isolated platelets; b) sequencing the transcriptome in the samples from a); c) determining gene expression in the transcriptome from the sequences from b); and d) identifying variable genes (where the change in gene expression of the variable gene is at least two standard deviations) or identifying recurrent genes (where the change in gene expression of the recurrent gene is less than two standard deviations). In some embodiments, variable genes or recurrent genes can be identified if the change in gene expression is at least three standard deviations. In some embodiments, variable genes or recurrent genes may be biomarkers. In some embodiments, biomarkers may be used for the diagnosis of disease, assessment of disease severity, and assessment of recovery from disease by detecting differentially expressed biomarkers in biological samples obtained from subjects compared to control or reference samples.
[0090] In some embodiments, the dataset can be the output generated by the analysis of RNA-Seq reads. The dataset can include all or substantially all known features of the reference (e.g., exons, introns, etc.). The dataset can be used to identify biomarkers.
[0091] In some embodiments, the methods disclosed herein may include the step of aligning sequence reads that are substantially comprehensively represented in an annotated reference. Alignment algorithms can be used to rapidly map reads, even when there are many reads associated with RNA-Seq results and substantially comprehensive references.
[0092] In some embodiments, the method may include analyzing a transcriptome by obtaining multiple sequence reads from it, finding alignments that each have an alignment score that meets a predetermined criterion, and identifying transcripts within the transcriptome. This method is suitable for analyzing reads obtained by RNA-Seq. The predetermined criterion may be, for example, the alignment with the highest score.
[0093] In some embodiments, the methods and systems disclosed herein may include transcriptome analysis in which a dataset can be used as a reference.
[0094] In some embodiments, the dataset can represent substantially all known exons of at least one chromosome. In some embodiments, alignments can be found by comparing each sequence read with at least a large portion of the possible paths through the dataset. The method may include assembling multiple sequence reads into a contig based on the found alignments.
[0095] In some embodiments, transcript identification includes identifying known and novel biomarkers.
[0096] The methods disclosed herein may include determining the expression level of a transcript.
[0097] In some embodiments, the methods disclosed herein may further include monitoring disease progression, monitoring residual disease, monitoring therapy, diagnosing a condition, prognosing a condition, or selecting a therapy based on discovered variants or biomarkers.
[0098] In some embodiments, the methods disclosed herein may further include the identification of variants or biomarkers that can be tracked by imaging tests (e.g., CT, PET-CT, MRI, X-ray, ultrasound) for localization of tissue abnormalities suspected to cause identified variants or biomarkers.
[0099] In some embodiments, the method can be used to identify any gene. In some embodiments, the gene can be a heritable gene. In some embodiments, any of the methods disclosed herein can be repeated multiple times. In some embodiments, any of the methods disclosed herein can be repeated one, two, three, four, five, six, seven, eight, nine, ten, or more times.
[0100] Genes identified as differentially expressed can be evaluated using various nucleic acid detection assays to detect or quantify the expression levels of one or more genes in a given sample. For example, gene expression levels can be detected using traditional Northern blotting, nuclease protection, RT-PCR, microarrays, and differential display methods. Methods for assaying mRNA include Northern blotting, slot blotting, dot blotting, and hybridization of oligonucleotides into regular arrays. Any method can be used to specifically and quantitatively measure a particular protein or mRNA or DNA product. However, methods and assays are most efficiently designed using array or chip hybridization-based methods for detecting the expression of multiple genes. Any hybridization assay format can be used, including solution-based and solid support-based assay formats.
[0101] The protein products of the genes identified herein can also be assayed to determine their expression levels. Methods for assaying proteins include Western blotting, immunoprecipitation, and radioimmunoassays. The proteins to be analyzed can be localized intracellularly (most commonly by immunohistochemistry) or extracellularly (most commonly by immunoassays such as ELISA).
[0102] In some embodiments, biological samples can be obtained from a subject before signs or symptoms of disease or injury appear. For example, in some embodiments, biological samples can be obtained approximately 1 minute, 5 minutes, 10 minutes, 30 minutes, 1 hour, 2 hours, 4 hours, 6 hours, 12 hours, 18 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 10 days, 2 weeks, 3 weeks, 4 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 9 months, 1 year, or 2 years before signs or symptoms of disease or injury appear. In some embodiments, samples can be obtained more than 2, 3, 4, 5, 6, 7, 8, 9, or 10 years before signs or symptoms of disease or injury appear. In some embodiments, samples can be obtained less than 1 minute before signs or symptoms of disease or injury appear. In some embodiments, multiple biological samples can be obtained at one or more different points in time.
[0103] In some embodiments, biological samples can be obtained from a subject after the onset of signs or symptoms of disease or injury. For example, in some embodiments, biological samples can be obtained approximately 1 minute, 5 minutes, 10 minutes, 30 minutes, 1 hour, 2 hours, 4 hours, 6 hours, 12 hours, 18 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 10 days, 2 weeks, 3 weeks, 4 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 9 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, or 10 years after the onset of signs or symptoms of disease or injury. In certain embodiments, samples can be obtained more than 2, 3, 4, 5, 6, 7, 8, 9 years, or 10 years after the onset of signs or symptoms of disease or injury. In some embodiments, samples can be obtained less than 1 minute after the onset of signs or symptoms of disease or injury. In a particular embodiment, multiple biological samples can be obtained at one or more different time points in time.
[0104] In some embodiments, the method may further include obtaining a second sample from a subject. In some embodiments, the subject may be the same subject. In some embodiments, the second sample derived from the same subject may be obtained between 4 months and 4 years. In some embodiments, the subject has one or more signs or symptoms of a disease.
[0105] The control group sample may be a sample derived from a normal or healthy subject, or a sample derived from a subject with a known disease or injury. In some embodiments, the control group sample may be a sample derived from a subject that has recovered from a known disease or injury, or a sample derived from a subject that has not recovered. As described herein, comparison of the expression pattern of the sample under test with the expression pattern of a control can be used to diagnose a disease or injury, assess the severity of a disease or injury, or evaluate recovery from a disease or injury. In some embodiments, the control group is for the purpose of establishing an initial cutoff or threshold for the assay described herein. Therefore, in some embodiments, the systems and methods disclosed herein can be used to diagnose a disease or injury, assess the severity of a disease or injury, or evaluate recovery from a disease or injury without requiring comparison with a control group.
[0106] Diagnostic methods Disclosed herein are methods for diagnosing disease or injury in subjects who have experienced signs or symptoms of disease or injury, or who have not experienced them, for assessing the severity of disease or injury, and for assessing recovery from disease or injury, by generating transcriptome-wide expression profiles.
[0107] In some embodiments, a method for producing a transcriptome-wide expression profile of a biological sample can be used for diagnosing disease or injury, assessing the severity of disease or injury, and evaluating recovery from disease or injury in subjects who have experienced signs or symptoms of disease or injury, or who have not. In some embodiments, a method for producing a transcriptome-wide expression profile of a biological sample may include: a) measuring the expression level of one or more genes present in a first biological sample, wherein the measurement includes RNA sequencing; b) determining the expression level of a gene from the measured expression level obtained in step a); c) combining the results of steps a) and b) to generate a transcriptome-wide expression profile; and d) providing the transcriptome-wide expression profile as a dataset, wherein the first biological sample consists of isolated platelets. In some embodiments, the first and second biological samples may be obtained from the same or different subjects. In some embodiments, steps a) to d) may be repeated for each biological sample. In some embodiments, the transcriptome-wide expression profiles of subjects can be compared. In some embodiments, the first and second biological samples may be obtained from the same subject at the first and second time points. In some embodiments, the first and second biological samples may be obtained from different subjects, and further include obtaining additional biological samples from the subject, which may be obtained at different time points. In some embodiments, the first and second time points may be different time points. In some embodiments, a transcriptome-wide expression profile from the first time point may be compared with a transcriptome-wide expression profile from the second time point. In some embodiments, the method further includes repeating these steps until a valid transcriptome-wide expression profile can be identified. In some embodiments, the expression level may be measured using a device selected from the group consisting of microarrays, bead arrays, liquid arrays, and nucleic acid sequences.
[0108] A method for producing a transcriptome-wide expression profile of a biological sample is disclosed herein, comprising: a) measuring the expression level of one or more genes present in a first biological sample, wherein the measurement includes RNA sequencing; b) determining the expression level of a gene from the measured expression level obtained in step a); c) combining the results of step a) and b) to generate a transcriptome-wide expression profile; and d) providing the transcriptome-wide expression profile as a dataset, wherein the first biological sample consists of isolated platelets. The method for producing a transcriptome-wide expression profile can be used for diagnosing disease or injury, assessing the severity of disease or injury, and assessing recovery from disease or injury in subjects who have experienced signs or symptoms of disease or injury or who have not.
[0109] This specification discloses transcriptome-wide expression profiles that can be used for disease diagnosis, disease severity assessment, and disease recovery assessment by detecting differentially expressed biomarkers in biological samples obtained from subjects compared to control or reference samples.
[0110] Transcriptome-wide expression profiles generated by methods disclosed herein can be used to diagnose diseases, assess disease severity, and evaluate recovery from diseases by detecting differentially expressed biomarkers in biological samples obtained from subjects compared to control or reference samples.
[0111] Methods that can be used to identify differences between transcriptome-wide expression profiles are disclosed herein. In some embodiments, these methods may include identifying differences between healthy transcriptome-wide expression profiles and non-healthy (e.g., diseased) transcriptome-wide expression profiles. In some embodiments, a method for identifying gene expression differences between two transcriptome-wide expression profiles includes determining one or more variations in the target platelet transcriptome using a method for producing a transcriptome-wide expression profile of a biological sample, wherein the method includes a) measuring the expression levels of one or more genes present in a first biological sample, wherein the measurement includes RNA sequencing; b) determining the gene expression levels from the measured expression levels obtained in step a); c) generating a transcriptome-wide expression profile by combining the results of step a) and step b); and d) providing the transcriptome-wide expression profile as a dataset, wherein the first biological sample consists of isolated platelets, the method is performed at different time points, thereby identifying and providing gene expression differences, and comparing such time-dependent changes from the target to a reference, the method is performed at different time points, thereby identifying and comparing gene expression differences. In some embodiments, the method can be repeated at least once. In some embodiments, the method can be repeated with a second biological sample, the second biological sample consists of isolated platelets. In some embodiments, the first and second biological samples may be obtained from the same or different subjects. In some embodiments, steps a) to d) may be repeated for each biological sample. In some embodiments, the transcriptome-wide expression profiles of the subjects may be compared. In some embodiments, the first and second biological samples may be obtained from the same subject at the first and second time points.In some embodiments, the first and second biological samples may be obtained from different subjects, and further include obtaining additional biological samples from the subjects, the additional biological samples being obtained at different time points. In some embodiments, the first and second time points may be different time points. In some embodiments, the transcriptome-wide expression profile from the first time point can be compared with the transcriptome-wide expression profile from the second time point. In some embodiments, the method further includes repeating those steps until a valid transcriptome-wide expression profile can be identified. In some embodiments, the expression levels may be measured using a device selected from the group consisting of microarrays, bead arrays, liquid arrays, and nucleic acid sequences.
[0112] In some embodiments, a method for identifying subjects with disease or injury, including asymptomatic subjects or subjects exhibiting nonspecific indicators of disease or injury, by generating a transcriptome-wide expression profile and detecting one or more biomarkers. These biomarkers are also useful for monitoring subjects receiving treatment and therapy for disease or injury and / or disease or injury-related conditions, as well as for selecting or modifying therapies and treatments that would be effective for subjects with disease or injury (here, selection and use of such therapies and treatments). Such therapies can treat the disease or injury by delaying or preventing symptoms associated with the disease or injury.
[0113] Improved methods for diagnosing and prognosing diseases or injuries are disclosed herein. Diagnosis or prognosis of a disease or injury can be assessed by measuring one or more of the biomarkers described herein and comparing the measured values to comparison, reference, or index values. Such comparisons can be performed using mathematical algorithms or formulas to combine information from the results of multiple individual biomarkers (e.g., signature, variant genes) with other parameters to obtain a single measurement or index. Subjects identified as having a disease or injury may be selectively selected to receive treatment regimens, such as the administration of therapeutic compounds to prevent, treat, or delay symptoms associated with the disease or injury.
[0114] Identifying subjects who have a disease or injury within several hours before or after the onset of signs or symptoms of the disease or injury may enable the selection and initiation of various therapeutic interventions or treatment regimens to delay, alleviate, or prevent symptoms associated with the disease or injury and to improve recovery. Monitoring the treatment process is also possible by monitoring the levels of at least one biomarker. For example, samples may be provided from subjects receiving a treatment regimen or therapeutic intervention. Such treatment regimens or therapeutic interventions may include, but are not limited to, the administration of pharmaceuticals and treatment with therapeutic or prophylactic drugs used for subjects diagnosed or identified as having a disease or injury. Samples may be obtained from subjects at various time points before, during, or after treatment.
[0115] Accordingly, biomarkers (or signatures) identified by the methods disclosed herein can be used to generate biomarker profiles or signatures for (i) subjects without disease or injury, (ii) subjects with disease or injury, or with symptoms or signs of disease or injury, and / or (iii) subjects who are recovering from or have recovered from disease or injury. By comparing the subject's biomarker profile to a predetermined or comparative biomarker profile or reference biomarker profile, disease or injury can be diagnosed, the progression or rate of progression of symptoms or pathologies associated with disease or injury can be monitored, and the effectiveness of disease or injury treatment can be monitored. Data relating to biomarkers identified by the methods disclosed herein may be combined with or correlated with other data or test results, including, but not limited to, clinical parameters or measurements of other algorithms relating to disease or injury. Other data may include age, sex, ethnicity, body mass index (BMI), neurological examinations, EEG recording data, EKG recording data, imaging results (e.g., CT scan, MRI, angiography), etc. Data may also include subject information such as medical history and any relevant family history.
[0116] Methods are disclosed herein that can be used to identify agents for treating a disease or injury that are appropriate for a specific subject or otherwise customized. In this regard, test samples can be taken from subjects exposed to a therapeutic agent or drug, and the levels of one or more biomarkers can be determined. The levels of one or more biomarkers can be compared to samples obtained from subjects before and after treatment, or to samples obtained from one or more subjects that showed improvement in risk factors as a result of such treatment or exposure.
[0117] In some embodiments, the methods described herein can utilize a biological sample (e.g., platelets) for the detection of one or more biomarkers in the sample. In some embodiments, the method includes the detection of one or more biomarkers in the platelets of interest.
[0118] In some embodiments, transcriptome-wide expression profiles generated as disclosed herein can be used to diagnose a disease, condition, disorder, or injury of interest. In some embodiments, the disease or injury includes, but is not limited to, infectious diseases including, bacterial, fungal, viral, and parasitic infections, as well as diseases resulting from infection, including, but not limited to, sepsis, severe sepsis, and septic shock, COVID-19, multisystem inflammatory syndrome (MIS-C) in children, and systemic inflammation. In some embodiments, the disease or injury includes, but is not limited to, cancer, autoimmune diseases, skin diseases, eye diseases, endocrine diseases, neurological disorders, and cardiovascular diseases.
[0119] In some embodiments, cancer can be a solid tumor cancer that includes, but is not limited to, colon cancer, pancreatic cancer, brain cancer, bladder cancer, breast cancer, prostate cancer, lung cancer, ovarian cancer, uterine cancer, liver cancer, kidney cancer, spleen cancer, thymic cancer, thyroid cancer, nerve tissue cancer, epithelial tissue cancer, lymph node cancer, bone cancer, muscle cancer, and skin cancer.
[0120] In some aspects, autoimmune diseases include achlorhydric autoimmune active chronic hepatitis, acute disseminated encephalomyelitis, acute hemorrhagic leukoencephalitis, Addison's disease, agammaglobulinemia, alopecia areata, amyotrophic lateral sclerosis, ankylosing spondylitis, anti-GBM / TBM nephritis, antiphospholipid syndrome, anti-synthetase syndrome, polyarthritis, atopic allergy, atopic dermatitis, autoimmune aplastic anemia, autoimmune cardiomyopathy, autoimmune intestinal disease, autoimmune hemolytic anemia, autoimmune hepatitis, autoimmune inner ear disease, autoimmune lymphoproliferative syndrome, autoimmune peripheral neuropathy, autoimmune pancreatitis, and autoimmune diseases. Immune polyendocrine syndrome, autoimmune progesterone dermatitis, autoimmune thrombocytopenic purpura, autoimmune uveitis, Baro disease / Baro concentric sclerosis, Behçet's syndrome, Berger disease, Vickerstaff encephalitis, Blau syndrome, bullous pemphigoid, Castleman disease, celiac disease, Chagas disease, chronic fatigue immune deficiency syndrome, chronic inflammatory demyelinating polyneuropathy, chronic relapsing multifocal osteomyelitis, chronic Lyme disease, chronic obstructive pulmonary disease, Churg-Strauss syndrome, scarring pemphigoid, celiac disease, Cogan syndrome, cold agglutinin disease, complement component 2 deficiency, cranial arteritis, Crest syndrome, Crohn's disease, Cushing's syndrome, cutaneous leukocytoclastic vasculitis, Degos disease, Darkham's disease, herpetiform dermatitis, dermatomyositis, type 1 diabetes, diffuse systemic cutaneous sclerosis, Dressler's syndrome, lupus discoid, eczema, endometriosis, enthesitis-associated arthritis, eosinophilic fasciitis, eosinophilic gastroenteritis, acquired epidermolysis bullosa, erythema nodosum, essential mixed cryoglobulinemia, Evans syndrome, progressive ossifying fibrosis, fibromyalgia / fibromyositis, fibrotic alveolitis, gastritis, gastroenteritis-like pemphigoid, giant cell arteritis, glomerulonephritis, Goodpasture syndrome, Graves' disease, Guillain-Barré syndrome - Syndrome, Hashimoto's encephalitis, Hashimoto's thyroiditis, hemolytic anemia, Henoch-Schönlein purpura, herpes zoster of pregnancy, sweat gland abscess, Hughes' syndrome, hypogammaglobulinemia, idiopathic inflammatory demyelinating disease, idiopathic pulmonary fibrosis, idiopathic thrombocytopenic purpura, IgA nephropathy, inclusion body myositis, inflammatory demyelinating polyneuropathy, interstitial cystitis, irritable bowel syndrome (IBS), juvenile idiopathic arthritis, juvenile rheumatoid arthritis, Kawasaki disease, Lambert-Eaton myasthenia gravis, leukocytoclastic vasculitis, lichen planus, lichen sclerosing, linear IgA disease, Lou Gehrig's disease, lupoid hepatitis, lupus erythematosus, Magid's syndrome,Meniere's disease, microscopic polyangiitis, Miller-Fischer syndrome, mixed connective tissue disease, focal scleroderma, Mucha-Habermann disease, Mackle-Wells syndrome, multiple myeloma, multiple sclerosis, myasthenia gravis, myositis, narcolepsy, neuromyelitis optica, neurogenic myotonica, ocular scarring pemphigoid, opsoclonus-myoclonus syndrome, old thyroiditis, relapsing rheumatoid arthritis, PA NDAS, paraneoplastic cerebellar degeneration, paroxysmal nocturnal hemoglobinuria, Parry-Romberg syndrome, Personage-Turner syndrome, ciliary body squamous cellulitis, pemphigus, pemphigus vulgaris, pernicious anemia, perivenosis encephalomyelitis, POEMS syndrome, polyarteritis nodosa, polymyalgia rheumatica, polymyositis, primary biliary cirrhosis, primary sclerosing cholangitis, progressive inflammatory neuropathy, psoriasis, psoriatic arthritis, necrosis This list may include, but is not limited to, pyoderma angina, pure red cell aplasia, Rasmussen syndrome, Raynaud's phenomenon, relapsing polychondritis, Reiter's syndrome, restless legs syndrome, retroperitoneal fibrosis, rheumatoid arthritis, rheumatic fever, sarcoidosis, schizophrenia, Schmidt syndrome, Schnitzler syndrome, scleritis, scleroderma, Sjögren's syndrome, spondyloarthritis, adhesive blood syndrome, Still's disease, Stiff Person syndrome, subacute bacterial endocarditis (SBE), Suzak syndrome, Sweet's syndrome, Sydenham chorea, sympathetic ophthalmitis, Takayasu's arteritis, temporal arteritis, Trosa Hunt syndrome, transverse myelitis, ulcerative colitis, undifferentiated connective tissue disease, undifferentiated spondyloarthritis, vasculitis, vitiligo, Wegener's granulomatosis, Wilson's syndrome, and Wiscott-Aldrich syndrome.
[0121] In some aspects, skin diseases include acneiform rash, autoinflammatory syndrome, chronic blister formation, mucosal conditions, skin appendage conditions, subcutaneous fat conditions, congenital abnormalities, connective tissue diseases (such as abnormalities of dermal fibrous tissue and elastic tissue), skin and subcutaneous growth, dermatitis (such as atopic dermatitis, contact dermatitis, eczema, pustular dermatitis, and seborrheic dermatitis), hyperpigmentation abnormalities, drug eruptions, endocrine-related skin diseases, eosinophilia, epidermal nevi, neoplasms, cysts, erythema, hereditary skin diseases, infection-related skin diseases, lichenoid rash, and phosphorus. This may include, but is not limited to, vasoconstrictive skin conditions, phenotypic nevus and neoplasms (such as melanoma), monocyte and macrophage-associated skin conditions, mucinosis, neurocutaneous conditions, noninfectious immunodeficiency-associated skin conditions, nutrition-associated skin conditions, papulosquamous hyperkeratosis (such as palmoplantar keratoderma), pregnancy-associated skin conditions, pruritus, psoriasis, reactive neutrophils, refractory palmoplantar eruption, conditions resulting from metabolic errors, conditions resulting from physical factors (such as ionizing radiation induction), urticaria and angioedema, and vascular-associated skin conditions.
[0122] In some embodiments, endocrine disorders may include, but are not limited to, adrenal disorders, glucose homeostasis disorders, thyroid disorders, calcium homeostasis disorders and metabolic bone disorders, pituitary disorders, and sex hormone disorders.
[0123] In some aspects, eye diseases include H00-H06 disorders of the eyelids, lacrimal system, and orbit; H10-H13 disorders of the conjunctiva; H15-H22 disorders of the sclera, cornea, iris, and ciliary body; H25-H28 disorders of the lens; H30-H36 disorders of the choroid and retina (H30 chorioretinal inflammation, H31 other choroidal disorders, H32 chorioretinal disorders in diseases classified elsewhere, H33 retinal detachment and tears, H34 retinal vascular occlusion). This may include, but is not limited to, H35 other retinal disorders (including retinal disorders in diseases classified elsewhere in H36), H40-H42 glaucoma, H43-H45 disorders of the vitreous humor and eyeball, H46-H48 disorders of the optic nerve and visual pathway, H49-H52 disorders of the extraocular muscles, binocular movement, accommodation, and reflexes, H53-H54.9 visual impairment and blindness, and H55-H59 other disorders of the eye and adnexa.
[0124] In some forms, neurological disorders include: loss of gravity, acquired epileptic aphasia, acute disseminated encephalomyelitis, adrenoleukodystrophy, corpus callosum agenesis, agnosia, Aicardi syndrome, Alexander disease, alien hand syndrome, sensori-spheric inversion, Alpers disease, alternating hemiplegia, Alzheimer's disease, amyotrophic lateral sclerosis (see motor neuron disease), anencephaly, Angelman syndrome, hemangioma, oxygen deficiency, aphasia, apraxia, arachnoid cyst, arachnoiditis, Arnold-Chiari malformation, arteriovenous malformation, telangiectasia ataxia, attention deficit hyperactivity disorder, and auditory processing disorder. Damage, autonomic nervous system dysfunction, back pain, Batten's disease, Behçet's disease, Bell's palsy, benign idiopathic blepharospasm, benign intracranial hypertension, bilateral polymicrogyria of the frontofront, Binswanger's disease, blepharospasm, Bloch-Salzberger syndrome, brachial plexus injury, brain abscess, brain injury, brain tumor, Brown-Séquard syndrome, Canavan disease, carpal tunnel syndrome, burning pain, central pain syndrome, central pontine myelin breakdown, central nucleus myopathy, head injury, cerebral aneurysm, cerebral arteriosclerosis, cerebral atrophy, cerebral gigantism, cerebral palsy, cerebrovascular disease, cervical spinal stenosis, Charcot-Marie-Tooth disease, Chiari malformation, chorea, chronic fatigue syndrome Syndrome, chronic inflammatory demyelinating polyneuropathy (CIDP), chronic pain, Coffin-Lowley syndrome, coma, complex regional pain syndrome, compressive neuropathy, congenital bilateral facial nerve palsy, corticobasal degeneration, cranial arteritis, craniosynostosis, Creutzfeldt-Jakob disease, cumulative traumatic disease, Cushing's syndrome, giant cell inclusion body disease (CIBD), cytomegalovirus infection, Dandy-Walker syndrome, Dawson's disease, Demorsia syndrome, Degerin-Klumpke palsy, Degerin-Sottas disease, delayed sleep syndrome, dementia, dermatomyositis, developmental behavioral disorders, diabetic neuropathy, diffuse Sclerosis, Dravet syndrome, autonomic nervous system disorder, dyscalculia, dysgraphia, dyslexia, dystonia, empty cell syndrome, encephalitis, brain herniation, trigeminal nerve hemangioma, encopresis, epilepsy, Erb's palsy, erythromelalgia, essential tremor, Fabry disease, Fahl syndrome, syncope, familial spastic paralysis, febrile seizures, Fisher syndrome, Friedreich's ataxia, fibromyalgia, Gaucher disease, Gerstmann syndrome, giant cell arteritis, giant cell inclusion disease, globoid cell leukoatrophy, gray matter ectopic formation, Guillain-Barré syndrome, HTLV-1 associated myelopathy, Harrelforden-Spats disease, head injury,Headache, hemifacial spasm, hereditary spastic paraplegia, hereditary polyneurotic ataxia, herpes zoster, shingles, Hirayama syndrome, holoprosencephalopathy, Huntington's disease, anencephaly, hydrocephalus, hyperadrenocorticism, hypoxia, immune-mediated encephalomyelitis, inclusion body myositis, incontinentia pigmenti, infantile phytanate storage, infantile Refsum disease, infantile seizures, inflammatory muscle disease, intracranial cysts, increased intracranial pressure, Joubert syndrome, Carac syndrome, Keens-Sayer syndrome, Kennedy disease, Kinsborne syndrome, Klippel-Feil syndrome, Krabbe disease, Kugelberg-Wellander disease, Kuru disease, Lafora disease, Lambert-Eaton myasthenic syndrome, Landau-Kleffner syndrome, Wallenberg syndrome, learning disabilities, Leigh disease, Lennox-Gastaut syndrome, Lesch-Nyhan syndrome, cerebral white matter atrophy, Lewy body dementia, gyral defects, locked-in syndrome, Lou Gehrig's disease (see motor neuron disease), intervertebral disc disease, lumbar stenosis, Lyme disease - neurological sequelae, Machado-Joseph disease ( Spinocerebellar ataxia type 3, cerebral encephalopathy, macropsia, megacephaly, Melkerson-Rosenthal syndrome, Meniere's disease, meningitis, Menkes disease, metachromatic leukodystrophy, cephalopia, micropsia, migraine, Miller-Fischer syndrome, petit mal seizures (transient ischemic attacks), mitochondrial myopathy, Moebius syndrome, monoliary muscular atrophy, motor neuron disease, motor neuropathy, Moyamoya disease, mucopolysaccharidosis, multiple stroke dementia, multifocal motor neuropathy, multiple sclerosis, multiple system atrophy, muscular dystrophy, myalgic encephalomyelitis Myasthenia gravis, myelin-disintegrating diffuse sclerosis, infant myoclonic encephalopathy, myoclonus, myopathy, myotubomyopathy, congenital myotonia, narcolepsy, neurofibromatosis, neuroleptic malignant syndrome, neurological symptoms of AIDS, neurological sequelae of lupus, neurogenic myotonia, neuronal ceroid lipofuscinosis, neuronal migration disorder, Niemann-Pick disease, non-24-hour sleep-wake syndrome, nonverbal learning disorder, O'Sullivan-MacLeod syndrome, occult neuralgia, secondary to occult spinal cord nonunion Spinal dysraphism sequence, Ohtahara syndrome, olivopontocerebellar atrophy, opsoclonus-myoclonus syndrome, optic neuritis, orthostatic hypotension, overuse syndrome, recurrent vision, paresthesia, Parkinson's disease, congenital paramyotonia, paraneoplastic disorders, paroxysmal seizures, Parry-Romberg syndrome, Pelizaeus-Merzbacher disease,Periodic paralysis, peripheral neuropathy, persistent vegetative state, pervasive developmental disorder, photic sneeze reflex, phytanic acid storage, Pick's disease, nerve compression, pituitary tumor, PMG, polio, polymicrogyria, polymyositis, porencephaly, post-polio syndrome, postherpetic neuralgia (PHN), post-infectious encephalomyelitis, post-orthostatic hypotension, Prader-Willi syndrome, primary lateral sclerosis, prion disease, progressive hemifacial atrophy, progressive multifocal white matter brain damage, progressive supranuclear palsy, pseudotumor, rabies, Ramsay Hunt syndrome (type i) (and type II), Rasmussen's encephalitis, reflex neurovascular dystrophy, Refsum's disease, repetitive movement disorder, repetitive stress injury, restless leg syndrome, retrovirus-associated myelopathy, Rett syndrome, Reye's syndrome, rhythmic movement disorder, Romberg syndrome, chorea, Sandhoff's disease, schizophrenia, Schilder's disease, schizencephaly, sensory integration disorder, septal-optic nerve dysplasia, shaken baby syndrome, herpes zoster, Shy-Drager syndrome, Sjögren's syndrome, sleep apnea, sleeping sickness, SNATIATIO N, Sotos syndrome, spasticity, spinal aryspinal fissure, spinal cord injury, spinal cord tumor, spinal muscular atrophy, spinocerebellar ataxia, Steele-Richardson-Olsewski syndrome, stiff person syndrome, seizures, Sturge-Weber syndrome, subacute sclerosing panencephalitis, subcortical arteriosclerotic encephalopathy, surface iron deposition disease, Sydenham's chorea, syncope, synesthesia, syringomyelia, tarsal tunnel syndrome, tardive dyskinesia, Tahlob's cyst, Tay-Sachs disease, temporal arteritis, scleralization, tethered spinal cord syndrome, Thomsen's disease, thoracic outlet syndrome This may include, but is not limited to, painful tics, Todd's palsy, Tourette's syndrome, toxic encephalopathy, transient ischemic attack, infectious spongiform encephalopathy, transverse myelitis, traumatic brain injury, tremor, trigeminal neuralgia, tropical spastic paraplegia, trypanosomiasis, tuberous sclerosis, von Hippel-Lindau disease, Biliwisk encephalomyelitis, Wallenberg syndrome, Werdnig-Hoffmann disease, West syndrome, whiplash, Williams syndrome, Wilson's disease, and Zellweger syndrome.
[0125] In some embodiments, cardiovascular disease may include, but is not limited to, aneurysms, angina pectoris, atherosclerosis, cerebrovascular disorders (attacks), cerebrovascular diseases, congestive heart failure, coronary artery disease, myocardial infarction (heart attack), and peripheral vascular disease.
[0126] In some embodiments, transcriptome-wide expression profiles can be used to determine whether a subject has recovered from disease or injury.
[0127] In some embodiments, the method includes detecting that one or more biomarkers in a sample obtained before or after the onset of signs or symptoms of a disease or before or after the diagnosis of a disease are upregulated compared to a control sample. In some embodiments, the method includes detecting that one or more biomarkers in a sample obtained before or after the onset of signs or symptoms of a disease or before or after diagnosis of a disease are downregulated compared to a control sample. In some embodiments, the control sample is the baseline level of one or more biomarkers measured in a sample obtained before the onset of signs or symptoms of a disease or injury. For example, in some embodiments, the method may include detecting upcontrol by determining that a change in the expression of one or more biomarkers in a target sample obtained within 1 minute, 5 minutes, 10 minutes, 30 minutes, 1 hour, 2 hours, 4 hours, 6 hours, 12 hours, 18 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 10 days, 2 weeks, 3 weeks, 4 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 9 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, or 10 years exceeds the change in the expression of one or more biomarkers in a control subject or population that does not have signs or symptoms of the disease or has not been diagnosed with the disease, compared to baseline.In some embodiments, the method may include detecting that one or more biomarkers in a target sample obtained before or after the onset of signs or symptoms of the disease, or before or after the diagnosis of the disease, are downregulated compared to a control. For example, in some embodiments, the method may include detecting downregulation by determining that a change in the expression of one or more biomarkers in a subject obtained from a sample within 1 minute, 5 minutes, 10 minutes, 30 minutes, 1 hour, 2 hours, 4 hours, 6 hours, 12 hours, 18 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 10 days, 2 weeks, 3 weeks, 4 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 9 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, or 10 years is less than the change in the expression of one or more biomarkers in a control subject or population that has not experienced or has not had signs or symptoms of the disease or injury, or has not been diagnosed with the disease, compared to baseline.
[0128] In some embodiments, the method may include detecting one or more markers in a biological sample of interest. In some embodiments, the level of one or more markers in a biological test sample of interest is compared to the level of a biomarker in a control. Non-limiting examples of a control include, but are not limited to, a negative control, a positive control, a standard control, a standard value, an expected normal background value for the subject, an estimated historical normal background value for the subject, a reference criterion, a reference level, an estimated normal background value for the population to which the subject is a member, or an estimated historical normal background value for the population to which the subject is a member. In some embodiments, the control may be the level of one or more biomarkers in a sample obtained from the subject before the onset of signs or symptoms of disease or injury. In some embodiments, the control may be the level of one or more biomarkers in a sample previously obtained from the subject after the onset of signs or symptoms of disease or injury, but before the collection of the test sample.
[0129] In some embodiments, the method may include monitoring the progression of disease or injury in a subject by evaluating the levels of one or more markers in the biological sample of the subject.
[0130] In some embodiments, the subjects may be human subjects of any race, sex, and age. In some embodiments, the subjects may be healthy or may have a disease or injury.
[0131] In some embodiments, the information obtained from the methods described herein can be used alone or in combination with other information from or obtained from a biological sample of a subject (e.g., disease status, disease history, vital signs, blood chemistry, neurological scores, etc.).
[0132] In some embodiments of the methods disclosed herein, it can be determined that the level of one or more markers has increased if the level of one or more markers has increased by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 125%, at least 150%, at least 175%, at least 200%, at least 250%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, at least 1000%, at least 1500%, at least 2000%, at least 2500%, at least 3000%, at least 4000%, or at least 5000% compared to a control.
[0133] In some embodiments of the methods disclosed herein, a decrease in the level of one or more markers can be determined if the level of one or more markers decreases by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 125%, at least 150%, at least 175%, at least 200%, at least 250%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, at least 1000%, at least 1500%, at least 2000%, at least 2500%, at least 3000%, at least 4000%, or at least 5000% compared to a control.
[0134] In some embodiments, a biological sample derived from a subject can be evaluated for one or more levels of markers in a biological sample obtained from a patient. The level of one or more markers in the biological sample can be determined by evaluating the amount of one or more polypeptides of the biomarkers in the biological sample, the amount of one or more mRNAs of the biomarkers in the biological sample, the amount of enzyme activity of one or more enzymes of the biomarkers in the biological sample, or a combination thereof.
[0135] This specification discloses compositions and methods relating to biomarkers that can be used to create transcriptome-wide expression profiles or to diagnose diseases in subjects. These biomarkers can be used for disease screening, diagnosis, onset monitoring, progression monitoring, and assessment of recovery. These biomarkers can be used to establish and evaluate treatment plans. In some embodiments, biomarkers can be identified using transcriptome-wide expression profiles. In some embodiments, once biomarkers are identified, gene signatures can be created using these biomarkers, which can then be used for disease screening, diagnosis, onset monitoring, progression monitoring, and assessment of recovery.
[0136] Compositions and methods for producing transcriptome-wide expression profiles are disclosed herein. Compositions and methods that can be used for disease diagnosis and prognosis are disclosed herein. In some embodiments, the method may include examining relevant biomarkers and their expression. In some embodiments, biomarker expression may include transcription to messenger RNA (mRNA) and translation to protein. In several embodiments, the method may include determining whether the expression level of the relevant biomarker is differentially expressed compared to a control. In some embodiments, the control may be the level of the relevant biomarker in a control sample of a subject without the disease, a population without the disease, a subject who has not recovered from the disease, a population who has not recovered from the disease, a subject who has recovered from the disease, a population who has recovered from the disease, and a subject being diagnosed (this control sample is obtained before the onset of signs or symptoms of the disease). In some embodiments, the method may include determining whether the expression level of the relevant biomarker in a sample obtained from a subject is differentially expressed compared to the expression level of the relevant biomarker in a sample previously obtained from the same subject, where the previously obtained sample is obtained at an early point in time after the onset of risk factors or signs or symptoms of the disease. In some embodiments, the method may include detecting the expression levels of relevant biomarkers across multiple samples obtained from the same subject over time, thereby providing a time-dependent view of biomarker expression.
[0137] Accordingly, in some embodiments, a method for diagnosing a disease is provided. The method includes a) providing a biological sample from a subject; b) analyzing the biological sample using an assay that specifically detects at least one biomarker of the present invention in the biological sample; and c) comparing the level of at least one biomarker in the sample with the level in a control sample or a previously obtained biological sample, wherein a statistically significant difference between the level of at least one biomarker in the sample and the level in a control sample or a previously obtained biological sample indicates brain damage. In some embodiments, the method further includes d) implementing a treatment regimen based thereon. In some embodiments, the method includes analyzing changes in gene expression in a subject over a defined time interval and comparing the detected changes in gene expression with changes in gene expression observed in a control subject.
[0138] In some embodiments, methods are provided for determining the prognosis or treatment regimen of a disease. These methods may include: a) providing a biological sample from a subject; b) analyzing the biological sample using an assay that specifically detects at least one biomarker in the biological sample; and c) comparing the level of at least one biomarker in the sample with the level in a control sample or a previously obtained biological sample, wherein a statistically significant difference between the level of at least one biomarker in the sample and the level in the control sample or a previously obtained biological sample indicates the prognosis or treatment regimen of a disease. In some embodiments, these methods may include analyzing changes in gene expression in a subject over a defined time interval and comparing the detected changes in gene expression with changes in gene expression observed in a control subject.
[0139] In some embodiments, the biomarker type may include mRNA biomarkers. In some embodiments, mRNA can be detected by at least one of the following: mass spectrometry, PCR microarray, thermal sequencing, capillary array sequencing, solid-phase sequencing, etc.
[0140] In some embodiments, the biomarker type may include polypeptide biomarkers. In some embodiments, the polypeptide can be detected by at least one of the following: ELISA, Western blotting, flow cytometry, immunofluorescence, immunohistochemistry, mass spectrometry, etc.
[0141] Biomarker detection Methods for detecting one or more mRNA biomarkers, polypeptide biomarkers, or combinations thereof in a biological sample are disclosed herein. Biomarkers can generally be measured and detected by a variety of assays, methods, and detection systems known to those skilled in the art.
[0142] Examples of methods include, but are not limited to, immunoassays, microarrays, PCR, RT-PCR, refractive index spectroscopy (RI), ultraviolet spectroscopy (UV), fluorescence analysis, electrochemical analysis, radiochemical analysis, near-infrared spectroscopy (NIR), infrared (IR) spectroscopy, nuclear magnetic resonance (NMR), light scattering analysis (LS), mass spectrometry, pyrolysis mass spectrometry, turbidimetric analysis, dispersion Raman spectroscopy, gas chromatography, liquid chromatography, gas chromatography combined with mass spectrometry, liquid chromatography combined with mass spectrometry, matrix-assisted laser desorption / ionization-time-of-flight (MALDI-TOF) combined with mass spectrometry, ion spray spectroscopy combined with mass spectrometry, capillary electrophoresis, colorimetric methods, and surface plasmon resonance (such as those provided by systems offered by Biacore Life Sciences). In some embodiments, biomarkers can be measured using the detection methods described above or other methods known to those skilled in the art. Other biomarkers can likewise be detected using reagents specifically designed or formulated to detect them.
[0143] Different types of biomarkers and their measurement results can be combined using the compositions and methods described herein. In some embodiments, the protein morphology of the biomarker can be measured. In some embodiments, the nucleic acid morphology of the biomarker can be measured. In some embodiments, the nucleic acid morphology can be mRNA. In some embodiments, the measurement results of protein biomarkers can be used in conjunction with the measurement results of nucleic acid biomarkers.
[0144] Methods for measuring polypeptide levels in biological samples obtained from a subject are disclosed herein, including but not limited to immunochromatography assays, immunodot assays, Luminex assays, ELISA assays, ELISPOT assays, protein microarray assays, ligand-receptor binding assays, ligand substitution from receptor assays, ligand substitution from covalent receptor assays, immunostaining assays, Western blot assays, mass spectrometry assays, radioimmunoassays (RIA), radioimmunodiffusion assays, liquid chromatography-tandem mass spectrometry assays, Octalony immunodiffusion assays, reverse-phase protein microarrays, rocket immunoelectrophoresis assays, immunohistochemistry assays, immunoprecipitation assays, complement fixation assays, FACS, enzyme-substrate binding assays, enzyme assays, enzyme assays using detectable molecules, such as chromophores, fluorophores, or radioactive substrates, substrate binding assays using such substrates, substrate substitution assays using such substrates, and protein chip assays.
[0145] Methods for detecting nucleic acids (e.g., mRNA), such as RT-PCR, real-time PCR, microarrays, branched DNA, and NASBA, are well known in the art. Using sequence information provided by database entries for biomarker sequences, the expression of biomarker sequences (if any) can be detected and measured using techniques known to those skilled in the art. For example, using sequences from sequence database entries or sequences disclosed herein, probes for detecting biomarker RNA sequences can be constructed, for example, by Northern blot hybridization analysis or by methods for specifically and quantitatively amplifying specific nucleic acid sequences. Alternatively, sequences can be used to construct primers for specifically amplifying biomarker sequences using amplification-based detection methods, such as reverse transcription-based polymerase chain reaction (RT-PCR). If changes in gene expression are associated with gene amplification, deletion, polymorphism, and mutation, sequence comparisons in test and reference populations can be performed by comparing the relative amounts of the tested DNA sequences in the test cell population and the reference cell population. In addition to Northern blotting and RT-PCR, RNA can be measured using other targeted amplification methods (e.g., TMA, SDA, NASBA), signal amplification methods (e.g., bDNA), nuclease protection assays, and insight hybridization.
[0146] In some embodiments, quantitative hybridization methods such as Southern spectroscopy, Northern spectroscopy, or Insights hybridization may be used. As used herein, “nucleic acid probe” may be a DNA probe or an RNA probe. A probe may be, for example, a gene, a gene fragment (e.g., one or more exons), a gene-containing vector, a probe, or a primer. A nucleic acid probe may be, for example, a full-length nucleic acid molecule or a portion thereof, e.g., an oligonucleotide of at least 15, 30, 50, 100, 250, or 500 nucleotides in length, and sufficient to specifically hybridize to a suitable target mRNA or cDNA under stringent conditions. The hybridization sample may be maintained under conditions sufficient to allow specific hybridization of the nucleic acid probe to mRNA or cDNA. Specific hybridization may be performed under high-stringency or moderate-stringency conditions, as needed. In some embodiments, the hybridization conditions for specific hybridization may be high-stringency. The specific hybridization, if present, is then detected using standard methods. If specific hybridization occurs between nucleic acid probes containing mRNA or cDNA in a test sample, the level of mRNA or cDNA in the sample can be evaluated. Two or more nucleic acid probes can also be used simultaneously in this manner. Specific hybridization of any one of the nucleic acid probes may indicate the presence of the target mRNA or cDNA, as described herein.
[0147] Alternatively, peptide nucleic acid (PNA) probes can be used instead of nucleic acid probes in the quantitative hybridization methods described herein. PNA is a DNA mimetic having a peptide-like inorganic backbone, such as an N-(2-aminoethyl)glycine unit, in which an organic base (A, G, C, T, or U) is bound to a glycine nitrogen via a methylene carbonyl linker. PNA probes can be designed to hybridize specifically to a target nucleic acid sequence. Hybridization of PNA probes to nucleic acid sequences can be used to determine the level of the target nucleic acid in a biological sample.
[0148] In some embodiments, the levels of one or more biomarkers in a biological sample obtained from a subject can be determined using an array of oligonucleotide probes complementary to a target nucleic acid sequence in the biological sample. Using an array of oligonucleotide probes, the levels of one or more biomarkers can be determined either individually or in relation to the levels of one or more other nucleic acids in the biological sample. Oligonucleotide arrays typically comprise a plurality of different oligonucleotide probes bound to the surface of a substrate at different known positions. These oligonucleotide arrays, also known as "Genechips," are commonly described in the art, for example, in U.S. Patent No. 5,143,854, and PCT Patent Publications WO90 / 15070 and WO92 / 10092. These arrays can generally be produced using mechanical synthesis or photo-directed synthesis methods, incorporating a combination of photolithography and solid-phase oligonucleotide synthesis.
[0149] After the oligonucleotide array is prepared, the target nucleic acid can be hybridized with the array and its level can be quantified. Hybridization and quantification are generally carried out by the methods described herein. Briefly, the target nucleic acid sequence can be amplified by a well-known amplification technique, e.g., PCR. Typically, this involves the use of a primer sequence complementary to the target nucleic acid. Asymmetric PCR techniques can also be used. The amplified target, which generally incorporates a label, can then be hybridized with the array under appropriate conditions. Once the hybridization and washing of the array are complete, the array can be scanned to determine the amount of hybridized nucleic acid. The hybridization data obtained from the scan can typically be in the form of fluorescence intensity as a function of the amount or relative amount of the target nucleic acid in the biological sample. The target nucleic acid can be hybridized to the array in combination with one or more comparison controls (e.g., positive control, negative control, quantity control, etc.) to improve the quantification of the target nucleic acid in the sample.
[0150] The probes and primers described herein may be directly or indirectly labeled with radioactive or non-radioactive compounds by methods known to those skilled in the art to obtain a detectable and / or quantifiable signal, and the labeling of the primers or probes may be carried out with radioactive elements or non-radioactive molecules. Of the radioactive isotopes used, 32 P, 33 P, 35 S, or 3H may be mentioned. Non-radioactive entities can be selected from ligands, e.g., biotin, avidin, streptavidin, or digoxigenin, haptens, dyes, and luminescent agents, e.g., radioluminescent agents, chemiluminescent agents, bioluminescent agents, fluorescent agents, or phosphorescent agents.
[0151] Nucleic acids can be obtained from cells using known techniques. In this specification, nucleic acids refer to RNA, including mRNA, and DNA, including cDNA. Nucleic acids can be double-stranded or single-stranded (i.e., sense or antisense single-stranded) and can be complementary to nucleic acids encoding polypeptides. Nucleic acid content can also be determined by RNA or DNA extraction performed on biological samples, including biological fluids and fresh or fixed tissue samples.
[0152] Many methods for detecting and quantifying specific nucleic acid sequences are known in the art, and new methods are continuously being reported. Most known specific nucleic acid detection and quantification methods utilize nucleic acid probes in specific hybridization reactions. In some embodiments, detection of hybridization to double-stranded forms can be performed using the Southern blotting technique. In the Southern blotting technique, nucleic acid samples can be separated in an agarose gel based on size (molecular weight), immobilized on a membrane, denatured, and exposed to (mixed with) a labeled nucleic acid probe under hybridization conditions. If the labeled nucleic acid probe forms a hybrid with the nucleic acid on the blot, the label can be bound to the membrane.
[0153] In Southern blotting, nucleic acid probes can be labeled with tags. In some embodiments, the tags can be radioisotopes, fluorescent dyes, or other well-known materials. Another type of process for the specific detection of nucleic acids in biological samples is hybridization. In some embodiments, nucleic acid probes consisting of at least 10 nucleotides, at least 15 nucleotides, or at least 25 nucleotides having a sequence complementary to the nucleic acid of interest can be hybridized in a sample subjected to depolymerization conditions, and this sample can be treated with an ATP / luciferase system, which will fluoresce if the nucleic acid sequence is present. In quantitative Southern blotting, the level of the nucleic acid of interest can be compared to the level of a second nucleic acid of interest and / or one or more control nucleic acids (e.g., positive control, negative control, quantitative control, etc.).
[0154] In some embodiments, a useful method for detecting and quantifying nucleic acids can be polymerase chain reaction (PCR). The PCR process is well known in the art. To briefly summarize PCR, a nucleic acid primer complementary to the opposite strand of the nucleic acid amplification target sequence is allowed to be annealed to a denatured sample. DNA polymerase (typically thermally stable) extends the DNA double helix from the hybridized primer. This process can be repeated to amplify the nucleic acid target. If the nucleic acid primer does not hybridize to the sample, there is no corresponding amplified PCR product. In this case, the PCR primer acts as a hybridization probe.
[0155] In PCR, nucleic acid probes can be labeled with tags as described herein. In some embodiments, double-strand detection can be performed using at least one primer directed to the nucleic acid of interest. In some embodiments of PCR, hybridized double-strand detection may include electrophoretic gel separation, followed by visualization based on dyes.
[0156] Typical hybridization and washing stringency conditions depend in part on the size of the oligonucleotide probe (i.e., the number of nucleotides is the length), the base composition, and the concentrations of monovalent and divalent cations.
[0157] In some embodiments, the process for determining the quantitative and qualitative profiles of the target nucleic acid described herein may be characterized in that amplification is real-time amplification performed using a labeled probe or a labeled hydrolysis probe that can specifically hybridize with the target nucleic acid segment under stringent conditions. The labeled probe can emit a detectable signal after each amplification cycle, thereby making it possible to measure the signal obtained for each cycle.
[0158] Real-time amplification, such as real-time PCR, is well known in the art, and various known techniques can be used in the best way to carry out this process. These techniques can be performed using probes from various categories, such as hydrolysis probes, hybridization neighbor probes, or molecular beacons. Techniques using hydrolysis probes or molecular beacons are based on the use of a fluorescent quencher / reporter system, while hybridization neighbor probes are based on the use of a fluorescent receptor / donor molecule.
[0159] Hydrolysis probes with fluorescent quenchers / reporter systems are commercially available. Many fluorescent dyes can be used, such as FAM dyes (6-carboxy-fluorescein) or any other dye phosphoramidite reagents.
[0160] One of the stringent conditions applicable to any one of the hydrolysis probes described herein is a Tm in the range of about 65°C to 75°C. In some embodiments, the Tm of any one of the hydrolysis probes described herein may be in the range of about 67°C to about 70°C. In some embodiments, the Tm applicable to any one of the hydrolysis probes described herein may be about 67°C.
[0161] In some embodiments, the methods described herein may include a primer complementary to the nucleic acid of interest, more specifically, the primer may include 12 or more consecutive nucleotides substantially complementary to the nucleic acid of interest. In some embodiments, the primers that can be used in the methods described herein may include nucleotide sequences sufficiently complementary to hybridize to a nucleic acid sequence of about 12 to 25 nucleotides.
[0162] In some embodiments, the primer differs from the target adjacent nucleotide sequence by only 1, 2, or 3 nucleotides or less. In some embodiments, the length of the primer may vary, for example, from about 15 to 28 nucleotides (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 nucleotides).
[0163] The concentration of biomarkers in a sample can be determined by any suitable assay. Suitable assays may include one or more of the following methods, enzyme assays, immunoassays, mass spectrometry, chromatography, electrophoresis, or antibody microarrays, or any combination thereof. Thus, as will be understood by those skilled in the art, the systems and methods disclosed herein may include any methods known in the art for detecting biomarkers in a sample.
[0164] Methods for a multiplexing analysis platform are also disclosed herein. In some embodiments, the method includes an analytical method for multiplexing the analytical measurement of markers.
[0165] As used herein, the term “expression” may mean, when used in relation to the expression or determination or detection of the expression level of one or more biomarkers, the determination or detection of the transcription of a biomarker (e.g., determination at the gene, i.e., mRNA level), and / or the determination or detection of the translation of a biomarker (e.g., determination or detection of the protein produced). Determining the expression level of a biomarker means determining whether or not the biomarker is expressed, and if so, to what relative degree it is expressed.
[0166] The expression levels of one or more biomarkers disclosed herein can be determined directly (e.g., by immunoassay, mass spectrometry) or indirectly (e.g., by determining the mRNA expression of a protein or peptide). Examples of mass spectrometry include ionization sources such as EI, CI, MALDI, and ESI, and analysis, spectroscopy, isotope ratio mass spectrometry (IRMS), thermal ionization mass spectrometry (TIMS), spark light source mass spectrometry, multiple reaction monitoring (MRM), or SRM. Any of these techniques can be performed in combination with pre-fractionation or concentration methods. Examples of immunoassays include immunoblotting, Western blotting, enzyme-linked immunosorbent assay (ELISA), enzyme immunoassay (EIA), and radioimmunoassay.
[0167] Immunoassays use antibodies for detection, and the determination of antigen levels is well known in the art. Antibodies can be immobilized on solid supports such as sticks, plates, beads, microbeads, or arrays.
[0168] The expression levels of one or more of the biomarkers described herein can also be determined indirectly by determining the mRNA expression of one or more biomarkers in a biological sample. RNA expression methods include, but are not limited to, cell mRNA extraction and Northern blotting using labeled probes that hybridize to transcripts encoding all or part of a gene, mRNA amplification using gene-specific primers, polymerase chain reaction (PCR), and reverse transcription polymerase chain reaction (RT-PCR), followed by quantitative detection of gene products by various methods, RNA extraction from cells followed by labeling, then use for probes to cDNA or oligonucleotides encoding a gene, insight hybridization, and detection of reporter genes.
[0169] Methods for measuring protein expression levels include, but are not limited to, Western blotting, immunoblotting, ELISA, radioimmunoassay, immunoprecipitation, surface plasmon resonance, chemiluminescence, fluorescence polarization, phosphorescence, immunohistochemical analysis, microcytometry, microarrays, microscopy, fluorescence-activated cell classification (FACS), and flow cytometry. These methods may also include assays based on specific protein properties, including, but are not limited to, enzyme activity or enzymatic interactions with other protein partners. Binding assays can also be used and are well known in the art. For example, the binding constant of a complex between two proteins can be determined using a BIAcore instrument. Other suitable assays for determining or detecting the binding of one protein to another include immunoassays such as ELISA and radioimmunoassay. Binding can be determined by monitoring changes in spectroscopy, or the optical properties of the protein can be determined by fluorescence, ultraviolet absorption, circular dichroism, or nuclear magnetic resonance (NMR).
[0170] As used herein, the term "comparative control" may be used interchangeably with "reference," "reference expression," "reference sample," "reference value," "control," and "control sample," and when used in relation to a sample or expression level of one or more biomarkers, one or more genes or proteins refer to a reference standard, which is expressed at a certain level in the sample, unaffected by experimental conditions, and indicates a level in the sample of a given disease state. The reference value may be a predetermined standard value or a range of predetermined standard values, representing disease or absence, or disease or a predetermined type or severity of disease.
[0171] The reference expression may be the level of one or more biomarkers or genes described herein in a reference sample derived from a subject or pool of subjects that is free from the disease or free from a given severity or type of the disease. In some embodiments, the reference value may be the level of one or more biomarkers or genes described herein in a sample derived from one or more subjects, where one or more subjects are considered healthy and free from the specific disease.
[0172] In some embodiments, the expression levels of one or more biomarkers or genes can be compared. By comparing the expression levels of one or more biomarkers obtained from a subject with reference expression levels, the subject's susceptibility to a disease can be determined.
[0173] Determining the expression levels of one or more biomarkers or genes may include determining whether the biomarkers or genes are upregulated or increased, downregulated or decreased, or unchanged compared to a control or reference sample. As used herein, the terms “upregulated” and “increased expression level” or “increased expression level” refer to sequences corresponding to one or more biomarkers or genes that are expressed, and the measure of the amount of the sequence is an increased expression level compared to a reference sample (e.g., from a “disease control” or “normal” control). For example, the terms “upregulated” and “increased expression level” or “increased expression level” refer to sequences corresponding to one or more biomarkers or genes that are expressed, and the measure of the amount of the sequence is an increased expression level of one or more biomarkers (e.g., proteins and / or mRNAs) compared to the expression of the same mRNA from a reference sample (e.g., from a “disease control” or “normal” control). "Increased expression level" means an increase of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or more, for example, 20%, 30%, 40%, or 50%, 60%, 70%, 80%, 90%, or more, or more than 1x, up to 2x, 3x, 4x, 5x, 10x, 50x, 100x, or more in expression. As used herein, the terms "downregulated" and "reduced expression level" or "reduced expression level" refer to sequences corresponding to one or more biomarkers or genes that are expressed, and the measure of the amount of the sequence is that it exhibits a reduced expression level compared to a reference sample (e.g., from a "disease control" or "normal" control). For example, the terms “downregulated” and “reduced expression levels” or “reduced expression levels” refer to sequences corresponding to one or more biomarkers or genes that are expressed, and the measure of the amount of the sequence is that it exhibits reduced expression levels of one or more biomarkers (e.g., proteins and / or mRNAs) compared to the expression of the same mRNA from a reference sample (e.g., from a “disease control” or “normal” control)."Reduced expression level" refers to a decrease of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or more, for example, 20%, 30%, 40%, or 50%, 60%, 70%, 80%, 90%, or more, or a decrease of more than 1x, up to 2x, 3x, 4x, 5x, 10x, 50x, 100x, or more.
[0174] In some embodiments, the method may include determining whether a subject has increased susceptibility to a disease. As described herein, the expression ratio of a sample derived from a subject may be determined by comparing it to a reference sample to determine whether the subject has increased susceptibility to a disease. The reference sample may be derived from a subject having one or more biomarkers or genes at “normal” levels. Suitable statistical and other analyses may be performed to confirm changes in one or more biomarkers compared to the reference sample (e.g., increased expression or higher levels), and the ratio of the sample expression level of one or more biomarkers to the reference expression level of one or more biomarkers indicates a higher expression level of one or more biomarkers in the sample. In some embodiments, the ratio of the sample expression levels of two or more, three or more, four or more, five or more, or six or more biomarkers to the reference expression levels of two or more, three or more, four or more, five or more, or six or more biomarkers indicates a higher expression level of two or more, three or more, four or more, five or more, or six or more biomarkers in the sample, indicating that the subject has increased susceptibility to a disease.
[0175] A higher or increased expression level of one or more biomarkers compared to a reference expression level of one or more biomarkers may indicate increased susceptibility to a disease. Observing characteristic patterns of increased (higher) or decreased (lower) sample expression levels of one or more biomarkers compared to a reference expression level can indicate disease susceptibility (e.g., higher or lower) in a subject.
[0176] The expression levels of one or more biomarkers or genes described herein may, for example, be a measure of one or more biomarkers or genes per unit weight or unit volume. In some embodiments, the expression level may be a ratio (for example, the amount of one or more biomarkers or genes in a sample relative to the amount of one or more biomarkers or markers in a reference value).
[0177] In some embodiments, a percentage change can be determined by comparing a sample derived from a subject with a reference sample to determine whether the subject has increased susceptibility to a disease. In other words, expression levels can be expressed as a percentage. For example, a percentage change in the expression level of one or more biomarkers or genes, where the expression level of one (or two, three, four, five, or six) or more biomarkers increases (or is higher than) 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% compared to the reference expression level of the biomarker, indicates increased susceptibility to a disease. Alternatively, the percentage change in the expression level of one or more biomarkers or genes may be 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% lower compared to the reference expression level.
[0178] In some embodiments, an increase or decrease in the expression level of a biomarker, gene, or protein, or a combination thereof, may indicate increased susceptibility to a disease or a diagnosis of a disease in a subject. In some embodiments, a characteristic pattern of increased or decreased expression levels of one or more of the biomarkers, genes, or proteins disclosed herein is shown.
[0179] In some embodiments, the methods disclosed herein may further include methods for preventing disease morbidity and / or death. For example, the method includes providing a subject with further examinations (which may include examinations for the disease), e.g., a routine physical examination in which increased susceptibility to the disease is diagnosed. The method further includes administering a therapy to prevent the onset or spread of the disease, thereby reducing disease morbidity and / or mortality.
[0180] kit Kits useful for the methods described herein are disclosed herein. In some embodiments, the kits may include various combinations of components useful for any of the methods described herein, for example, materials for quantitatively analyzing biomarkers (e.g., polypeptides and / or nucleic acids), materials for evaluating the activity of biomarkers (e.g., polypeptides and / or nucleic acids), and teaching materials. For example, in some embodiments, the kits may include components useful for quantifying a desired nucleic acid in a biological sample. In some embodiments, the kits may include components useful for quantifying a desired polypeptide in a biological sample. In some embodiments, the kits may include components useful for evaluating the activity (e.g., enzyme activity, substrate binding activity, etc.) of a desired polypeptide in a biological sample.
[0181] In some embodiments, the kit may include components of an assay for monitoring the effectiveness of a treatment administered to a subject in need of treatment, including components for determining whether the levels of a biomarker in a biological sample obtained from the subject are modulated during or after treatment. In some embodiments, to determine whether the levels of a biomarker in a biological sample obtained from a subject are modulated, the levels of the biomarker can be compared to at least one control included in the kit, e.g., the level of a positive control, negative control, historical control, historical baseline, or the level of another reference molecule in the biological sample. In some embodiments, the ratio of the biomarker to the reference molecule can be determined to assist in monitoring the treatment.
[0182] In one embodiment, a kit is provided for measuring the RNA (e.g., RNA products) of one or more biomarkers disclosed herein. The kit may include materials and reagents that can be used to measure the RNA expression of one or more biomarkers. Examples of preferred kits include RT-PCR or microarrays. These kits may include the reagents necessary to perform the measurement of RNA expression levels. Alternatively, these kits may further include additional materials and reagents. For example, these kits may include materials and reagents necessary to measure the RNA expression levels of any number of genes, up to 1, 2, 3, 4, 5, 10, or more, that are not biomarkers disclosed herein.
[0183] treatment method Disclosed herein are methods for diagnosing disease or injury in subjects who have experienced signs or symptoms of disease or injury, methods for assessing the severity of disease or injury, and methods for assessing recovery from disease or injury, by generating transcriptome-wide expression profiles. In some embodiments, methods for producing transcriptome-wide expression profiles of biological samples can be used for diagnosing disease or injury, assessing the severity of disease or injury, and assessing recovery from disease or injury in subjects who have experienced signs or symptoms of disease or injury, or subjects who have not. In some embodiments, methods for producing transcriptome-wide expression profiles of biological samples may include: a) measuring the expression level of one or more genes present in a first biological sample, wherein the measurement includes RNA sequencing; b) determining the expression level of a gene from the measured expression level obtained in step a); c) combining the results of steps a) and b) to generate a transcriptome-wide expression profile; and d) providing the transcriptome-wide expression profile as a dataset, wherein the first biological sample consists of isolated platelets. In some embodiments, the first and second biological samples may be obtained from the same or different subjects. In some embodiments, steps a) to d) may be repeated for each biological sample. In some embodiments, the transcriptome-wide expression profiles of the subjects may be compared. In some embodiments, the first and second biological samples may be obtained from the same subject at a first and second time point. In some embodiments, the first and second biological samples may be obtained from different subjects, further including obtaining additional biological samples from the subjects, the additional biological samples being obtained at different time points. In some embodiments, the first and second time points may be different time points. In some embodiments, the transcriptome-wide expression profile from the first time point may be compared with the transcriptome-wide expression profile from the second time point.In some embodiments, the method further includes repeating those steps until an effective transcriptome-wide expression profile can be identified. In some embodiments, the expression level can be measured using a device selected from the group consisting of microarrays, bead arrays, liquid arrays, and nucleic acid sequences. In some embodiments, the method may further include treating a subject diagnosed with a disease or injury or one or more symptoms of a disease or injury.
[0184] Therapies comprising administering disease modifiers to subjects are disclosed herein. Disease modifiers may be therapeutic or prophylactic agents used in subjects diagnosed or identified with the disease, or at risk of developing the disease. In some embodiments, modification of therapy means changing the duration, frequency, or intensity of therapy, for example, changing the dosage level.
[0185] In some embodiments, administering treatment may include getting the subject to receive treatment or communicating to the subject the need for treatment. In some embodiments, the therapy may be surgery.
[0186] In some embodiments, measuring biomarker levels allows for monitoring the course of disease treatment. The effectiveness of a disease treatment regimen can be monitored by detecting effective amounts of one or more biomarkers in samples obtained from the subject over time and comparing the amounts of the detected biomarkers. For example, the first sample can be obtained before the subject receives treatment, and one or more subsequent samples can be collected after or during treatment. Changes in biomarker levels across samples can provide an indicator of the effectiveness of the therapy.
[0187] In some embodiments, test samples derived from a subject may be exposed to a therapeutic agent or drug to identify an appropriate therapeutic agent or drug for a particular subject, and the levels of one or more biomarkers can be determined. The biomarker levels can be compared to samples obtained from the subject before and after treatment or before and after exposure to the therapeutic agent or drug, or to samples obtained from one or more subjects who showed improvement in the disease as a result of such treatment or exposure. Accordingly, a method for evaluating the effectiveness of a therapy is disclosed herein, comprising: performing a first measurement of a biomarker panel in a first sample derived from a subject; administering the therapy to the subject; performing a second measurement of a biomarker panel in a second sample derived from the subject; and evaluating the effectiveness of the therapy by comparing the first measurement results with the second measurement results.
[0188] In addition, therapeutic agents suitable for administration to a particular subject can be identified by detecting one or more effective amounts of biomarkers in a sample obtained from the subject, and by exposing the sample to a test compound that determines the amount of biomarkers in the sample. Therefore, a treatment or therapeutic regimen for use in a subject with a disease (or one or more signs or symptoms of the disease) can be selected based on the amount of biomarkers in a sample obtained from the subject and compared to a reference value. Two or more treatments or therapeutic regimens can be evaluated in parallel to determine which treatment or therapeutic regimen is most effective for the subject in delaying the onset or progression of the disease. In some embodiments, suggestions can be made regarding initiating or continuing treatment for the disease.
[0189] In some embodiments, the administration of therapy may include administering disease-modulating agents to the subject. The subject may be treated with one or more drugs until the altered levels of the measured biomarkers return to baseline levels measured in the disease-free population, the population with relatively milder disease, or the population showing improvement in disease biomarkers as a result of drug treatment. In some embodiments, the subject may be treated with one or more drugs until the altered levels of the measured biomarkers return to baseline levels measured in a pre-symptomatic sample obtained from the subject. In addition, improvements related to altered levels of biomarkers or clinical parameters may be a result of treatment with disease-modulating agents.
[0190] Any drug or combination of drugs disclosed herein may be administered to a subject to treat a disease. The drugs herein may often be formulated in any number of ways, according to various formulations known in the art, or as disclosed or referenced herein.
[0191] In some embodiments, any drug or combination of drugs disclosed herein will not be administered to a subject for the treatment of a disease. In some embodiments, a healthcare professional may refrain from administering a drug or combination of drugs, recommend that a subject not be administered a drug or combination of drugs, or prevent a subject from being administered a drug or combination of drugs.
[0192] In some embodiments, one or more additional drugs may be administered optionally in addition to the recommended or administered drugs. These additional drugs are typically not unrecommended or drugs to be avoided. [Examples]
[0193] Example 1: Repeat longitudinal RNA-seq analysis of gene expression and splicing in human platelets identifies SELP splice QTLs. Longitudinal studies are necessary to distinguish between intra-individual and inter-individual variability and to study the repetition of gene expression. These studies have progressed to the point of decoding gene signals from environmental noise, and are expected to have applications in gene variant and expression research. However, longitudinal analysis of gene expression in healthy individuals lacks data for most primary cell types, including platelets, particularly regarding alternative splicing.
[0194] We evaluated the repetition of gene expression and splicing in platelets and used this repetition to identify novel platelet eQTLs and sQTLs.
[0195] Transcriptomes of platelets repeatedly isolated from healthy individuals for up to four years were sequenced. Easily measurable alternative splicing events, such as intra-individual and inter-individual variability and repetition of platelet RNA expression, as well as exon skipping, were examined. The results indicate that platelet gene expression is generally stable inter-individual and intra-individual over time, with the exception of a subset of genes enriched with inflammatory gene ontology. The results also show enrichment between repetitive genes in relation to heritable traits, including known and novel platelet eQTLs. Several exon skipping events were also highly repetitive, suggesting a heritable pattern of splicing in platelets. One of the most repetitive was exon 14 skipping of SELP. Therefore, rs6128 was identified as a platelet sQTL, and an rs6128-dependent association between SELP exon 14 skipping and race was defined. In vitro experiments demonstrate that this single-nucleotide variant directly affects exon 14 skipping and alters the ratio of transmembrane P-selectin protein production to soluble P-selectin protein production.
[0196] The platelet transcriptome has remained largely stable for four years. The findings demonstrate the use of gene expression and splicing repetitions to identify novel platelet eQTLs and sQTLs. rs6128 is a platelet sQTL that alters the ratio of SELP exon 14 skipping and soluble P-selectin protein production to transmembrane P-selectin protein production.
[0197] In this example, longitudinal RNA-seq analysis was used to investigate intra-individual and inter-individual variability in the human platelet transcriptome. Repetition and exon skipping, readily measurable alternative splicing events, were examined. Repetition retrospectively demonstrated its use in decoding heritable signals from environmental noise and identifying eQTL genes. Furthermore, repetition was prospectively used to prioritize candidate eQTL and sQTL genes for the discovery of novel platelet eQTLs and sQTLs.
[0198] Materials and methods. Human subjects and platelet isolation. [Table 1]
[0199] The subjects were healthy and had no active medical conditions. None of the subjects had undergone surgery in the past four months. Any disease required symptom resolution for at least seven days prior to sampling. Cohort 2 subjects were prospectively recruited. Cohort 1 subjects were part of a clinical study in which subjects had been previously exposed to aspirin but had not received aspirin for at least four weeks prior to the collection of each blood sample. Blood was collected by venous puncture into a citrate tube (Cohort 1) or acid-citrate-dextrose (Cohort 2), and sample processing was initiated within 30 minutes of venotomy. Platelets were isolated by magnetic leukocyte depletion using CD45 microbeads (Miltenyi) (Rowley JW, et al. Blood. 2011;118:e101-e111, Bray PF, et al. BMC Genomics. 2013;14:1, and Voora D, Cyr D, et al. J Am Coll Cardiol. 2013;62:1267-76).
[0200] RNA isolation and sequencing. RNA was isolated using phenol-chloroform extraction (Cohort 1) (Rowley JW, et al. Blood. 2011;118:e101-e111) or the DirectZol kit (Cohort 2, Zymogen). Sequenced libraries from Cohort 1 and Cohort 2 were barcoded and prepared using the KAPA Stranded mRNA-Seq Kit (Roche, no. KK8421) and the poly(A)-selective TruSeq unstranded v2 (Illumina, no. RS-122), respectively. The libraries were sequenced to a mapping read depth of approximately 20 to 40 million per sample on an Illumina HiSeq 4000 (Cohort 1) or Illumina HiSeq 2000 (Cohort 2) with 50 cycles, single-ended. The Fastq files are deposited in the NIH Sequence Read Archive PRJNA531691.
[0201] For RNA-seq analysis and expression differential analysis, reads were aligned to GRCh38 / hg38 using Novoalign (Novocraft) (Rowley JW, et al. Blood. 2011;118:e101-e111). Reads were assigned to flattened Ensembl gene annotations using the USeq analysis package (Nix DA, et al. BMC Bioinformatics. 2010;11:455). Read counts were normalized separately for each cohort using the DESeq2 analysis package (Love MI, et al. Genome Biology. 2014;15:550). Non-coding RNAs were selected according to Ensembl transcription biotypes. Heatmaps, clustering (fully linked), density plots, box plots, and scatter plots were generated using R (R Development Core Team. R: A Language and Environment for Statistical Computing. 2018, available from http: / / www.r-project.org). Read distribution plots were generated using the Integrated Genomics Viewer (IGV) (Thorvaldsdottir H, et al. Brief Bioinform. 2013;14:178-92). Gene ontology was analyzed using DAVID (Huang DW, et al. Nat Protoc. 2009;4:44-57).
[0202] Isolation and Sequencing of RNA. After isolation, fresh platelet pellets were suspended in 1 mL of Trizol and frozen at -80 °C until RNA isolation. RNA from cohort 1 was isolated using phenol-chloroform extraction, isopropanol precipitation in the presence of glycogen, and 75% ethanol washes. Samples were treated with DNAse (Invitrogen, number AM1907), and RNA was re-purified using ammonium acetate / isopropanol precipitation (Rowley JW, et al. Blood. 2011;118:e101-e111). RNA from cohort 2 was isolated and treated with DNAse using the DirectZol kit and column (Zymogen).
[0203] Analysis of Individual Transcript Variation. RNA-seq data have a strong mean-variance relationship. DESeq2 regularized log transformation (RLD) (Love MI, et al. Genome Biology. 2014;15:550) was applied to the counts to preferentially reduce the overall variance among low-abundance transcripts (while retaining outliers), thus enabling easier comparison of transcript variation across expression levels. The lowest-abundance transcripts were also optionally excluded. Intra-individual variation was calculated as the standard deviation of samples from the same individual. Total variation was calculated as the standard deviation of samples across individuals in each cohort, and thus intra-individual variation is a sub-component of total individual variation. Sources of variance were quantified in a linear mixed model using the R package variancePartition (Hoffman GE, et al. BMC Bioinformatics. 2016;17:483). Reproducibility was calculated using the formula σ b 2 / (σ w 2 +σ b 2 ) in the R package "heritability" (Kruijer W, et al. Genetics. 2015;199:379-98), including sex correction in the reproducibility calculation for eQTL enrichment analysis and gene prioritization for eQTL and sQTL discovery.
[0204] Exon skipping analysis. Exon skipping events were identified from triplicate structures of exon / exon and exon / intron junctions with more than 5 reads per junction. The calculated splice-in percentage (PSI, see Figure 5D) was greater than 0.05 and less than 0.95 in more than 30% of the samples. Junctions where adjacent exon / intron junction pairs changed by more than 10-fold in more than 70% of the samples were excluded. This strategy identified 245 exons (derived from 194 different transcripts), which were skipped in a significant proportion of the transcripts in some samples.
[0205] The complete open reading frame of the SELP transcript ENST00000263686.10, which has a c-terminal DYK tag, as well as introns 13 and 14 adjacent to exon 14, were cloned into PCDNA-CMV vectors and pCDH-MSCV-GFP vectors. A single nucleotide change from C to T was performed at rs6128. 293T cells (HEK293T / 17, ATCC CRL-11268) were maintained according to ATCC recommendations, and passages 10-20 were used. 293T cells were transfected with lipofectamine 2000. 24 hours after transfection, SELP RNA splicing was analyzed by PCR for 25 and 30 cycles using primers adjacent to exon 14 (5'-gtcaactaccgtgccaacct (SEQ ID NO: 1), 5'-taaggactcgggtcaaatgc (SEQ ID NO: 2)). For flow cytometry experiments, cells were either co-transfected with a GFP plasmid (PCDNA-CMV) or GFP was incorporated within the same scaffold as SELP (PCDH-MSCV). Surface expression of P-selectin was evaluated by staining with Psel.KO2.3 APC antibody (ThermoFisher, no. 17-0626-82) and analyzed using a CytoFLEX analyzer (Beckman). P-selectin was analyzed by flow cytometry (Coulter). MFI of P-selectin was normalized to GFP expression, which was assessed in live / transfected cells gated according to forward / side scattering (live) and FL-1 (GFP+) intensity. Soluble P-selectin was measured using the Quantikine ELISA kit (R&D Systems, no. DPSE00).For Western blotting, cells were lysed with RIPA, the lysate was denatured and reduced, proteins were separated using 10% SDS-PAGE, transferred to a PVDF membrane, and blotted for tagged P-selectin with anti-DYK antibody (Cell Signaling Tech., no. 2368S), followed by blotting for P-selectin labeled with anti-rabbit HRP secondary antibody (Rockland, no. 18-8816-33), and detection was performed by chemiluminescence (ThermoFisher, no. 34580).
[0206] Novel platelet eQTL and sQTL analysis. RNA-seq fastq files from 234 previously published samples (Best et.al. (Best MG, et.al. Cancer Cell. 2015;28:666-676, and Best MG, et.al. Cancer Cell. 2017;32:238-252), Netherlands cohort, hereinafter referred to as the NL cohort) were searched from the NCBI short-read archive PRJNA353588 (Best MG, et.al. Cancer Cell. 2017;32:238-252). These and the fastq files from Cohort 1 were aligned to a human reference genome (HG38 constructed) using splice recognition with STAR (Dobin A, et al. Bioinformatics. 2013;29:15-21). Variants were then called and filtered using a workflow constructed from the Genome Analysis Toolkit (GATK) (McKenna A, et al. Genome Res. 2010;20:1297-1303) best practices for variants that call RNA-seq. Variants tested for eQTLs were limited to within 2kb of genes not identified as eQTLs by the PRAX1 (Simon LM, et al. Am J Hum Genet. 2016;98:883-97) dataset, which had a repeatability of greater than 0.9 (238 genes). For comparison, a similar number of genes with the lowest repeatability were also included. The combined filtering process yielded 641 variants across 181 genes whose association with gene abundance and variants was tested.The RNA-seq allele frequencies of these variants were comparable to those reported by the Genome of the Netherlands project (GoNL (Genome of the Netherlands Consortium, Francioli LC, Menelaou A, et al. Whole-genome sequence variation, population structure and demographic history of the Dutch population. Nat Genet. 2014;46:818-825)) and 1000 Genomes (Gibbs RA, et al. Nature. 2015;526:68-74) (Figure 17A), and were clustered according to the subpopulations predicted by PCA analysis of allele frequencies (Figure 17B). RNA-seq and predicted allele frequencies of significant variants. Transcriptome-wide variants were invoked to assess population stratification. Multidimensional scaling (MDS) was performed using Plink 1.9 (Purcell S, et al. Am J Hum Genet. 2007;81:559-75, and Chang CC, et al. Gigascience. 2015;4:7) on 1994 variants (filtered and pruned: FS>30.0, QD<2.0, clusterSize=3, clusterWindowSize=35, --geno 0.2, --hwe 10e-6, --maf 0.01, --indep-pairwise 50 5 0.2) from RNA-seq in Cohort 1 and NL cohort, and simultaneously identified from 1000 genomes. Visual inspection of the MDS plots showed that individual RNA-seq samples were clustered according to their expected genetic ancestry, and their population structure was captured within the first four MDS components (Figure 17C). Gene expression was normalized using variance-stabilized transformation (VST) within the DESeq2 package (Love MI, et al. Genome Biology. 2014;15:550).The association between genes and variants was investigated using the SNPassoc R package (Gonzalez JR, et al. Bioinformatics. 2007; 1, 2) with an additive variant-allelic dosage model (0, 1, 2), while controlling for covariates sex, age, and population structure (Purcell S, et al. Am J Hum Genet. 2007; 81: 559-75, and Chang CC, et al. Gigascience. 2015; 4: 7) (see Figure 17 and Methods for details). Benjamini-Hochberg FDR corrections for multiple tests (641 gene-variant tests) have been reported. However, novel eQTLs were filtered as if a genome-wide analysis had been performed using a conservative significance threshold p < 1e-6 (Simon LM, et al. Am J Hum Genet. 2016; 98: 883-97). Allelic imbalances for each significant eQTL were assessed using the Wilcoxon rank-sum test for the ratio of reference variant reads to total reads in each heterozygous individual. After examining significance, eQTLs were limited to those reported in dbSNPs to minimize the possibility of RNA-specific (i.e., RNA editing) calls. At this threshold, 27 variants were identified across 11 novel platelet eQTL genes. These remained significant even after further control for the first five latent variables estimated from proxy variable analysis (SVA) (Leek JT, Storey JD. PLoS Genet. 2007;3:e161), with explicit adjustments being for sex and age. Subsequently, the 27 significant gene-variant associations were tested in Cohort 1 using the strategy described herein for the NL cohort, including sex, age, and race as covariates. However, due to the small sample size, a codominant model was observed when one of the three genotypes required for additive model testing was missing. The results of the eQTL analysis in Cohort 1 included a relatively small sample size, and therefore, significance in Cohort 1 was not expected and was not considered a criterion for selecting novel platelet eQTLs.
[0207] A generalized linear model was used in logistic regression analysis in R to investigate the association between SELP exon 14 splicing and self-reported race or rs6128 variant-allele dosage. For this purpose, SELP exon 14 inclusion counts plus exclusion counts were used as the binary response variable against total inclusion counts. Where specified, the model was controlled for latent covariates including race or population structure, rs6128 genotype, sex, and age.
[0208] There are recognized strengths and limitations to using RNA-seq to infer genetic variants (Piskol R, Ramaswaami G, Li JB. Am J Hum Genet. 2013;93:641-51). RNA variant calling is of a different quality than genome calling. They are limited to the regions in which they are expressed, making fine-tuning of the causative eQTLs impossible. Genetic variants may also be missed in the presence of extreme allele imbalances. As with other eQTL analyses, false positives are possible, such as from confounding LDs and genetic substructures that are not explained by large-scale population stratification. Detailed structural analysis is impossible due to the low density of variants in RNA-seq data. Due to these limitations, additional observations would be useful in clarifying these results. The allele frequencies of significant variant RNA-seq calls were consistent with expected allele frequencies (with the exception of rs879095052 in HBG1), 22 of the 27 eQTLs (7 of the 11 eGenes) had been reported in other tissues (The Genotype-Tissue Expression (GTEx) Project), and 8 of the 11 eGenes showed allele imbalances that were unidirectionally consistent with the influence of eQTLs on expression.
[0209] Statistical significance for multimodality was calculated according to the Hartigan dip test statistic. Similarity in the distribution of inter-sample and intra-sample correlations was tested using the Wilcoxon test adjusted for multiple comparisons. The Kolmogorov-Smirnov test was used to test the enrichment of ranks of genes associated with heritable traits / heritability. Pre-ranked gene set enrichment analysis (GSEA) (Subramanian A, et al. Proc Natl Acad Sci US A. 2005; 102: 15545-50) was used to assess the enrichment between genes tested for the presence of significant eQTLs identified in the PRAX1 cohort. For this purpose, a subset of genes tested by PRAX1 was included. The association between eQTL presence and different repeatability thresholds was estimated using odds ratios (odds at each threshold compared to no threshold) and significance assessed by independent chi-square tests. A two-tailed t-test and correlation test (alpha = 0.05) for significance were performed using the functions cor.test and t.test in R (R Development Core Team. R: A Language and Environment for Statistical Computing. 2018, available from www.r-project.org), setting the lower limit of the p-value to 2.2e-16.
[0210] Results. To evaluate intra-individual and inter-individual variability in platelet transcriptomes, RNA-seq analysis of leukocyte-depleted platelets longitudinally isolated from two independent cohorts of healthy individuals was used. In Cohort 1, platelets from 31 individuals were analyzed at initial examination (time 0) and 4 months later. In Cohort 2, platelets from 7 individuals obtained longitudinally at time 0 and then over a 4-year period were analyzed. The characteristics of these two cohorts are detailed in Table 1.
[0211] Platelets contain stable intra-individual gene expression signatures. Unsupervised clustering analysis of pairwise distances within and between individual transcriptomes in Cohort 1 revealed robust intra-individual (self) RNA expression signatures (Figure 1A), with most self pairs clustering as the nearest neighbor pair. As shown in Figures 1B–1C, the mean intra-individual correlation of platelet transcriptomes isolated at 4-month intervals was very high (r mean ± standard deviation = 0.987 ± 0.012). Inter-individual correlation of the entire platelet transcriptome was also high (0.947 ± 0.024), but significantly lower than intra-individual correlation (p < 2.2e-16). Differences were partially corrected by grouping samples by race, age, or sex (Figure 1C). Intra-individual clustering of non-coding RNAs was also robust (Figure 8A), with a mean intra-individual correlation of 0.984 ± 0.013. The mean inter-individual correlation for non-protein-coding transcripts (0.905 ± 0.031, Figures 8B–8C) was significantly weaker compared to protein-coding transcripts (p < 2.2e-16), which is consistent with previous cross-sectional studies. Raw and normalized counts for each transcript were determined in Cohort 1.
[0212] Unsupervised clustering of the total RNA transcriptome in Cohort 2 yielded robust self-clustering, suggesting minimal transcriptional drift over four years (Figure 1D). Intra-individual correlations of samples isolated at four-year intervals remained comparable to those of samples isolated at two-week intervals and were significantly higher than inter-individual correlations at all time points (Figures 1E-1F). As shown in Figure 8D, intra-individual non-coding RNA signatures were also robust, with individuals uniquely identified at every time point over the four-year period, reflecting significantly higher intra-individual correlations compared to inter-individual correlations in non-coding RNA expression (Figures 8E-8F). Raw and normalized counts of each transcript in Cohort 2 were determined.
[0213] Intra-individual and inter-individual transcript variability in platelets is repetitive across cohorts. In summary, the data in Figure 1 show similar patterns of gene expression variability between cohort 1 and cohort 2, with similar mean intra-individual correlations (0.983±0.13 vs. 0.987±0.12) and similar mean inter-individual correlations (0.958±0.021 vs. 0.947±0.024). Specific transcripts exhibiting minimum and maximum overall variability in each cohort were further defined. As shown in Figures 2A and 9A–9B, intra-individual variability was limited to a small number of moderately expressed transcripts that consistently showed variability in both cohort 1 and cohort 2. The transcripts with the most variability within individuals in both cohorts were enriched against those within the inflammatory and protective response gene ontology (GO, Figure 2B) (Bryois J, et al. Genome Res. 2017;27:545-552, and Whitney AR, et al. Proc Natl Acad Sci US A. 2003;100:1896-901). As shown in Figures 2C and 9C–9D, transcripts with high total variability spanned a wider range of expression levels, but the degree of variability remained consistent between cohort 1 and cohort 2. The transcripts with the highest total variability mainly overlapped with those with the highest intra-individual variability (Figures 9E–9F), resulting in enrichment of the inflammatory and protective response GO (Figure 2D). In addition to inflammatory transcripts, several genes were found to have high overall variability previously associated with sex, race, or platelet eQTLs (Edelstein LC, et al. Nat Med. 2013;19:1609-16, Simon LM, et al. Am J Hum Genet. 2016;98:883-97, and Simon LM, et al. Blood. 2014;123(16):e37-45). However, unlike inflammatory transcripts, most of these did not show high intra-individual variability (see the lower right quadrant in Figures 9E-9F).
[0214] Using variance partitioning analysis (Hoffman GE, Schadt EE. BMC Bioinformatics. 2016;17:483), the amount of variation attributable to sex, race, and other covariates was further partitioned and quantified for each gene, separating intra-individual variation from total variation. As shown in Figure 10, for more than half of the genes, inter-individual variation accounted for the majority (over 50%) of the gene expression variation, followed by the remainder of intra-individual variation. On the other hand, sex and race affected only a few genes. Other known covariates, including age and sample treatment, contributed only a small fraction of the variation for each gene.
[0215] Repetition defines hereditary platelet gene expression and predicts eQTL genes.
[0216] As discussed herein, we tested whether genes could be prioritized for eQTL analysis using repeatability, which captures intra-individual and inter-individual variability in a single indicator (see Methods). Therefore, we calculated the repeatability of each gene and retrospectively tested whether repeatability is associated with the genetic and heritable regulation of gene expression in platelets. To do this, we utilized the publicly available PRAX1 dataset (Simon LM, et al. Am J Hum Genet. 2016;98:883-97) as an independent (non-overlapping with the current cohort) and cross-platform (microarray) validation dataset. PRAX1 has previously associated platelet transcripts with sex and race and identified 612 platelet eQTL genes. Cohort 1 was used for this analysis because its size was larger than Cohort 2 and it better matched PRAX1 diversity and demographics. Genes in the PRAX1 microarray were ranked according to the repeatability calculated by RNA-seq in Cohort 1. Ranking by repetition resulted in significant enrichment of genes associated with sex, race, or eQTLs (p<2e-16), which was significantly greater than when genes were ranked by abundance, intraspecific variability, or total variability (p<2e-6; see also Figure 11). As shown in Figure 3A, 100% of the top 15 repetition-ranked genes were significantly associated with sex, race, or cis-eQTLs in terms of expression. For example, MFN2 is an established platelet eQTL gene among the most repetition-ranked genes (repetition = 0.98) (Simon LM, et al. Am J Hum Genet. 2016;98:883-97). This is because MFN2 expression varies depending on the eQTL genotype (Simon LM, et al. Am J Hum Genet. 2016;98:883-97), and while this variation is more than eightfold between individuals, it remains relatively constant over time within individuals (Figure 3B).
[0217] In particular, when evaluating eQTL gene enrichment, GSEA showed significant enrichment of known eQTLs as repetition increased (p=0), with repetition greater than 0.68 (tip) explaining the enrichment (Figure 3C). Significant enrichment was also observed in differences ranked by mean expression abundance or total variability, but the enrichment scores for these measures were lower than those for repetition. According to bin analysis of odds ratios, the odds of identifying eQTLs of genes with repetition less than 0.5 were 3.2 times lower than randomly testing genes (p=5e-15) (Figure 3D). The odds of identifying eQTLs of genes with repetition greater than 0.9 were 6.2 times higher than random (p=6e-45) and 1.8 times higher than ranking by total variability (p=0.006, adjusted). Furthermore, in the case of known platelet cis-eQTLs, there was a significant correlation between the repetition of the eQTL FDR and its associated transcript (Figure 12). Therefore, repetition indicates enrichment and intensity of cis-eQTL signals in platelets and can be a useful filtering and prioritization strategy for identifying genes with eQTL signals.
[0218] Microarrays differ from RNA-seq in terms of accuracy, sensitivity, and comprehensiveness, and some eQTL genes identifiable by RNA-seq may have been missed by PRAX1. Therefore, we re-examined 238 genes with high repeatability (≥0.9) using RNA-seq data, but no previously found cis-eQTLs were identified. These were then tested for cis-eQTLs in platelet RNA-seq from 234 healthy individuals (NL cohort) using publicly available datasets collected in the Netherlands (Simon LM, et al. Am J Hum Genet. 2016;98:883-97, and Simon LM, et al. Blood. 2014;123(16):e37-45). Genetic variants were identified from RNA-seq reads across each gene and tested for their association with RNA-seq abundance. Despite known limitations of RNA-seq in calling genetic variants (e.g., most variants are located within promoters and introns), 11 novel potential platelet eQTL genes were identified. In contrast, no additional eQTLs were identified when analyzing the same number of genes with the lowest repeatability, and this difference was statistically significant (11 / 238 vs. 0 / 238, p=0.0009, Fisher's exact test). Allele-specific expression (ASE) analysis of allele imbalance, measuring the ratio of read counts from each allele in heterozygotes, confirmed a significant, unidirectional, and consistent in-sample eQTL effect on allele imbalance for 8 of the 11 genes. One example of a novel platelet eQTL gene is the long non-coding RNA LINC01089, which is one of the transcripts with the highest repeatability (0.95) and abundance (top 10% by RNA-seq) in platelets. As shown in Figure 3E, there is a strong additive allele dose effect of rs1168663 on LINC01089 expression between individuals in Cohort 1 and between individuals in the NL cohort at both time points. As shown in Figure 3F, LINC01089 expression shows a significant allele imbalance between Cohort 1 and the NL cohort.In summary, these data define several novel platelet cis-eQTLs, demonstrate the repetitive utility of cross-sectional expression data, predict heritable gene expression variations, and prioritize targets for prospective identification of cis-eQTL genes.
[0219] Exon skipping in platelets is maintained within the organism over time. To evaluate the intra-individual and inter-individual stability of alternative splicing, we focused on exon skipping, an alternative splicing event readily measurable in RNA-seq data. Exon skipping events in Cohort 1 (see Methods) were strictly identified, and the percentage of exon splice-in (PSI, Figure 4A) for each was calculated. As shown in Figure 4B, a wide range of exon skipping levels existed across different exon skipping events, which remained fairly consistent between inter-individual and intra-individual clustering over time. Unsupervised clustering analysis using PSI yielded a preference for intra-individual clustering compared to inter-individual clustering (Figure 13), although this was less robust than expression-based clustering. Nevertheless, intra-individual correlations of PSI were significantly higher than inter-individual correlations, independent of age, race, or sex (Figures 4B-4C), suggesting a heritable component of exon skipping levels in platelets.
[0220] Identification of race-related platelet splice QTLs affecting exon 14 skipping in SELP. Unlike eQTLs, platelet sQTLs had not been previously identified. Therefore, the goal was to prioritize exon skipping events using repetition and identify novel, robust, physiologically significant platelet cis-sQTLs. To achieve this goal, intra-individual / inter-individual variability of PSI was assessed for each exon skipping event, and each was ranked by repetition. As shown in Figure 5A, exon 14 of SELP, encoding the leukocyte adhesion and platelet activation marker P-selectin, was ranked second among the repetitive exon skipping events. While there are differences in exon 14 exon skipping between donors, intra-individual stability is shown by the alignment plot in Figure 5B and the correlation plot in Figure 5C.
[0221] Exon 14 skipping predicts in-frame deletion of the transmembrane domain of P-selectin. Exon 14-deficient isoforms of P-selectin have been previously detected in endothelial cells and platelets (Johnston GI, et al. J Biol Chem. 1990;265:21381-5, McEver RP. Blood Cells. 1990;16:73-80, and Ishiwata N, et al. J Biol Chem. 1994;269:23708-15) and have been detected at significant levels in the human circulatory system (McEver RP. Blood Cells. 1990;16:73-80, Ishiwata N, et al. J Biol Chem. 1994;269:23708-15, and Semenov AV, et al. Biochem Biokhimiia. 1999;64:1326-35). PCR (Figure 14), cloning, and Sanger sequencing verified that RNA isoforms predicted by RNA-seq in the cohort matched the aforementioned soluble protein isoforms in plasma. In summary, this suggests that exon 14 skipping is the source of variability in inter-individual P-selectin protein cell surface and soluble plasma levels.
[0222] Previous studies have associated soluble P-selectin in plasma with various clinical and genetic factors, including race and single nucleotide polymorphisms (SNPs) (Ataga KI, et al. N Engl J Med. 2017;376:429-439, Lee DS, et al. J Thromb Haemost. 2007;6:20-31, Burger PC, Wagner DD. Blood. 2003;101:2661-2666, and Penman A, Hoadley S, et al. Am J Ophthalmol. 2015;159:1152-1160.e2). As shown in Figure 5D, a significant increase in SELP exon 14 inclusion was observed among Black / African Americans compared to White Americans at both time points. Searching for the most likely causative genetic variant identified rs6128, a SNP within exon 14 of SELP, which has a significantly higher homozygous MAF(T / T) ratio in Africans (0.29) compared to Europeans (0.04) (Gibbs RA, et al. Nature. 2015;526:68-74). Interestingly, rs6128 has been associated with plasma P-selectin levels (Sun BB, et al. Nature. 2018;558:73-79) and diabetic retinopathy, particularly among African Americans (Penman A, et al. Am J Ophthalmol. 2015;159:1152-1160.e2). However, the direct relationship between rs6128 and soluble P-selectin remains unclear because the C-to-T transition does not alter the protein sequence or modify the standard splice site. Bioinformatics analysis of the sequence surrounding rs6128 (Raponi M, et al. Hum Mutat. 2011;32:436-444) predicted the net loss of the exon splicing silencer (ESS) motif and the net gain of the two exon splicing enhancer (ESE) motifs (Figure 6A).To determine whether rs6128 is associated with SELP exon 14 splicing in platelets, rs6128 genotypes were inferred from RNA-seq reads of cohorts 1 and NL (Best MG, et al. Cancer Cell. 2017;32:238-252). As shown in Figures 6B–6C, rs6128 SNPs were significantly associated with SELP exon 14 skipping in platelets in both cohorts. The association between rs6128 and SELP exon 14 skipping was independent of age, sex, race, or population structure.
[0223] Subsequently, we tested whether the difference in rs6128 MAF between Africans and Europeans could explain the association between exon 14 skipping and race, as shown in Figure 5D. Consistent with this, after adjusting for rs6128, there was no difference in exon 14 skipping levels between Black / African Americans and Whites (p=0.4 and 0.3 at time 0 and 4 months, respectively).
[0224] Given the importance of P-selectin in disease, we extended the SELP exon 14 splicing analysis to additional diseases (also available from Best et al. (Best MG, et al. Cancer Cell. 2015;28:666-676, and Best MG, et al. Cancer Cell. 2017;32:238-252)). As shown in Figure 15, none of the diseases examined (non-small cell lung cancer, multiple sclerosis, or pulmonary hypertension) were significantly associated with SELP exon 14 splicing or the effect of rs6128 on splicing. This suggests that the level of exon 14 skipping is stable within the organism, even with respect to environmental stressors that have been shown to cause changes in platelet transcript abundance (Best MG, et al. Cancer Res. 2018;78:3407-3412, and Best MG, et al. Cancer Cell. 2017;32:238-252).
[0225] Rs6128 directly affects SELP exon 14 skipping and the ratio of soluble P-selectin to surface P-selectin in vitro. Non-causal markers are generally falsely identified in genetic association studies due to their association with other unobserved variables (Platt A, et al. Genetics. 2010;186:1045-52). To specifically test the causal effect of rs6128 on SELP exon 14 splicing, we generated minigene constructs of SELP ORFs with rs6128 C / C or T / T, including an intron adjacent to exon 14 (Figure 7A). Since promoter differences are known to affect splicing (Cramer P, et al. Proc Natl Acad Sci US A. 1997;94:11456-60), two different promoters (CMV or MSCV) were tested for each minigene construct. The construct was expressed in HEK293 cells lacking endogenous P-selectin. After transfection, RT-PCR analysis confirmed that the single nucleotide change from C / C to T / T resulted in a significant shift in the ratio of SELP RNA isoforms, consistent with RNA-seq results (Figure 7B). This occurred in both promoters, but a more pronounced shift was observed in the CMV promoter. Western blot analysis showed that the T / T variant resulted in a shift to exon 14 inclusion in the P-selectin protein (Figure 16). As shown in Figure 7C, the single nucleotide change from C / C to T / T significantly increased the amount of surface P-selectin on HEK293 cells (2-fold), consistent with transmembrane domain inclusion. In contrast, the T / T variant significantly decreased the amount of soluble P-selectin in the supernatant, as measured by ELISA (Figure 7D). In summary, this data demonstrates a causal relationship between the rs6128 genotype, the amount of exon 14 inclusion in SELP RNA, and the ratio of soluble P-selectin expression to surface P-selectin expression.
[0226] Discussion. An analysis comparing intra-individual and inter-individual variability in two independent cohorts showed that the human platelet transcriptome is highly stable for up to four years in healthy individuals. There are few available longitudinal studies in primary nucleated cells for comparison. One study with a similar design by Radich et al. (Radich JP, et al. Genomics. 2004;83:980-988) observed a 30% intra-individual misclassification rate for the leukocyte transcriptome, even after selecting gene signatures that maximized inter-individual variability. Although there are differences between this published study and the results described herein, it was identified that platelets had a lower intra-individual misclassification rate (10–12%) without signature selection. Their anucleated nature and lifespan of 7–10 days are thought to mitigate in vitro and in vivo RNA changes in platelets, promoting stable, defined in vivo healthy gene expression signatures.
[0227] Platelet gene expression profiling by RNA sequencing is emerging as a relevant tool for platelet function research, defining causal relationships and causes of diseases, and disease diagnosis (Best MG, et al. Cancer Cell. 2015;28:666-676, Kondkar AA, et al. J Thromb Haemost. 2010;8:369-78, Edelstein LC, et al. Nat Med. 2013;19:1609-16, Schubert S, et al. Blood. 2014;124:493-502, Kong X, et al. Thromb Haemost. 2017;117:962-970, Best MG, et al. Cancer Cell. 2017;32:238-252, and Campbell (RA, et al. Blood. 2019; blood-2018-09-873984). However, most published studies to date rely on single-time point-in-time comparisons of platelet transcriptomes between disease cohorts and healthy subjects. Since healthy subjects are often used as the “baseline” or “control” condition in these studies, it is important to understand whether the platelet transcriptome is durable in healthy individuals in order to understand the robustness of these comparisons. The finding that the platelet transcriptome is generally stable over four years in healthy individuals validates these comparisons. Thorough analysis of individual transcripts that differ within and between individuals also suggests some limitations and caveats.
[0228] Most transcripts were stable, but a small number varied significantly within the organism. Most of the intracellularly variable transcripts were associated with inflammation. These could provide information for studies evaluating the effects of inflammation on platelet gene expression and may be relevant to the clinical findings that inflammatory stress is associated with platelet count and function. 58 When evaluating the impact of inflammatory gene changes on disease, it may be worthwhile to investigate the range of inflammatory transcript fluctuations in healthy individuals compared to those in overtly inflammatory diseases.
[0229] Significant differences in platelet expression were observed between individuals. Sex and race were the causes of some of the major differences. Other sources of individual variation contributed more. This data suggests a prominent role for cis-eQTLs. Regardless of the source of variation, trends in genes that vary between (or within) healthy individuals may be taken into consideration when interpreting differential disease gene studies and designing validation experiments. Genes with high intrinsic variability require larger sample sizes to achieve statistical confidence. With small sample sizes, the likelihood of false positives is higher for genes with high intrinsic variability. Correction for known covariates may be useful in this regard. While corrections for sex, race, and age are often considered in differential gene expression analyses, eQTLs are usually unavailable or ignored.
[0230] Stability information (along with the Minimum Information for Publication of Quantitative Real-Time PCR Experiments (MIQE (Bustin SA, et al. Clin Chem. 2009;55:611-622) guidelines) can sometimes guide the selection of reference genes for normalization controls. Intra-individual / inter-individual stable genes such as SYK, AKT1 / 2, GP1BA, and ACTB may be good choices. On the other hand, genes with high intra-individual variability, such as the TUBB1 gene, should generally be avoided as reference genes.
[0231] As an application of longitudinal datasets, we identified novel platelet eQTLs using repetition as a filtering strategy. Due to the limitations of multiple testing (Altshuler D, et al. Science. 2008;322:881-888), filtering and prioritization strategies are also used in larger gene expression association studies to avoid testing thousands of genes (and further alternative splice events) against millions of loci. Common filtering strategies include hard filtering against abundance (Kumar V, et al. PLoS Genet. 2013;9:e1003201) or total variance. However, filtering against total variance alone can enrich transcripts that are overly influenced by technical or environmental noise. To avoid this, several studies (Barendse W. BMC Genomics. 2011;12:232, Carlborg O, et al. Bioinformatics. 2004;21:2383-93, and Hoffman GE, Schadt EE. BMC Bioinformatics. 2016;17:483) have suggested using repetition instead. Here, we experimentally tested this idea using platelet longitudinal data. Significant enrichment of cis-eQTLs was observed among the most repetitive genes, significantly improving the ability to identify eQTL genes compared to when abundance or total variation was used.
[0232] eQTLs were significantly enriched, but the association with repeatability was not complete. Cohort 1 was assayed on a different platform (microarray versus RNA-seq) and was smaller than PRAX1 (31 vs. 154) and had similar but not identical demographics (race (Black / African American: 45% vs. 42% (no significant difference), sex (male): 49% vs. 32% (no significant difference), age: 42+ / -11 vs. 29+ / -7, p<0.05)). Repeatability measured in the same samples used for eQTL analysis would probably provide the best predictive value. However, large-scale longitudinal studies are often impractical due to the cost and challenges of repeated sampling. This is not accurate. Furthermore, repetition can deepen the confidence in eQTL results when measured in independent cohorts. For this reason, a larger sample size that reflects the demographics and environment of the study cohort would probably be better. Cohorts that are too small to capture genetic variation (i.e., eQTLs with low MAF) or cohorts subject to systematic environmental variability suffer from lower repetition overall and lack the sensitivity to predict eQTLs. Additional research is needed to determine how repetition applies more broadly to the analysis of additional datasets, cell types, and gene-environment interactions.
[0233] Using an additional RNA-seq cohort (NL cohort), we tested unreported platelet eQTL genes among the most repetitive genes in Cohort 1. Of the 27 eQTLs identified, 22 (7 out of 11 eQTL genes) have been reported in other tissues (The Genotype-Tissue Expression (GTEx) Project). Of particular note among the candidate eQTL genes is TECPR2, which had previously been associated with platelet count by GWAS (Astle WJ, et.al. Cell. 2016;167:1415-1429.e19). Long non-coding RNA eQTL genes: LINC01089, MAGI2-AS3, and KANSL1-AS1 were also identified. Similar to these three genes, non-coding RNAs were found to be generally stable within individuals but more variable between individuals compared to protein-coding RNAs, suggesting greater genetic diversity among non-coding RNAs. Long non-coding RNAs have attracted attention for their multifaceted ability to regulate gene expression, but they are still being studied in platelets.
[0234] By further applying repetition to prioritize exon skipping events, we identified those most likely to identify biologically manageable sQTL signals. A robust association between rs6128 and SELP exon 14 skipping was identified. A significant association between rs6128 splicing and exon 14 skipping was also identified in whole blood samples (Zhernakova DV, et.al. Nat Genet. 2017;49:139-145), further reinforcing this finding.
[0235] To establish a causal relationship, we performed transfection experiments in HEK293 cells that do not favorably express endogenous SELP. The results strongly suggest that rs6128 is responsible for differential SELP exon 14 splicing in platelets, but do not rule out the potential contribution of the endogenous promoter or additional related or associated variants. Confirmation of platelet observations in unrelated cell lines suggests that rs6128 may influence SELP splicing in multiple tissues, including endothelial cells, another major P-selectin producing cell line.
[0236] Surface P-selectins mediate leukocyte interactions and inflammation, are involved in atherogenesis, and contribute to tumor metastasis. Soluble P-selectins are functionally and clinically significant platelet proteins in the cardiovascular system associated with various diseases (Ataga KI, et al. N Engl J Med. 2017;376:429-439, and Ludwig RJ, et al. Expert Opin Ther Targets. 2007;11:1103-1117). While the main source of soluble P-selectins is associated with activation-induced shedding, a significant amount is hereditary. An association between rs6128 in plasma and soluble P-selectin protein levels has been reported (Penman A, et al. Am J Ophthalmol. 2015;159:1152-1160.e2, and Sun BB, Maranville JC, Peters JE, et al. Genomic atlas of the human plasma proteome. Nature. 2018;558:73-79). The observed effects of rs6128 on surface P-selectin localization compared to exon 14 SELP splicing and soluble P-selectin localization establish an association between these observations. They may explain previous clinical studies that linked rs6128 to plasma P-selectin and diabetic retinopathy in African Americans (Penman A, et al. Am J Ophthalmol. 2015;159:1152-1160.e2). Finally, the finding that SELP is differentially spliced by race may have therapeutic significance in light of promising clinical trials that effectively used P-selectin blockade to treat painful episodes in sickle cell anemia, a disease that primarily affects individuals of African descent (Ataga KI, et.al. N Engl J Med. 2017;376:429-439).
[0237] While human platelets possess a rich RNA repertoire, platelet RNA expression differs between healthy and diseased individuals and is associated with platelet function. Single-point studies of the platelet transcriptome are increasingly being used for biological discoveries in human health and disease, and the results disclosed herein provide new information.
[0238] The disclosed results show that platelet RNA expression remains stable and repeatable over time in healthy human donors for up to four years, and that the integrated use of longitudinal repeatability indicators significantly facilitates the discovery of genetic variants affecting gene expression and splicing, demonstrating that genetic variants of the SELP gene induce the removal of the P-selectin transmembrane domain.
[0239] The disclosed results demonstrate that platelets, though anucleated, possess a rich and dynamic transcriptome. The use of platelet transcriptomics is increasing, with applications ranging from cancer diagnosis to gene discovery. The results disclosed herein add to the art by establishing the stability or repeatability of platelet gene expression and splicing in healthy donors repeatedly evaluated over up to four years. This type of longitudinal evaluation is lacking not only in platelets but in any primary human cell. These results demonstrate that the platelet transcriptome is remarkably stable in healthy conditions, which can aid in comparisons across disease states and enhance diagnostic and prognostic diagnoses using platelet RNA. Furthermore, the data demonstrate that integrating measures of repeatability (e.g., comparison of inter-individual and intra-individual variability in platelet RNA expression and splicing) leads to improved detection of genes influenced by nearby genetic variants. This technique can be applied to the discovery of platelet SELP sQTLs. Functionally, these splice QTLs induce the removal of the transmembrane domain of P-selectin. The findings disclosed herein indicate that this transmembrane domain deletion is reduced in Black individuals compared to Caucasians. In vitro, this leads to an increase in surface P-selectin but a decrease in soluble P-selectin, suggesting that this may affect diseases more common in Black individuals for which P-selectin is a therapeutic target (e.g., sickle cell disease).
Claims
1. A method for preparing transcriptome-wide expression profiles of biological samples, a) Measuring the expression level of one or more genes present in a first biological sample, wherein the measurement includes RNA sequencing. b) Determining the expression level of the gene from the measured expression level obtained in step a), c) Combining the results of step a) and step b) to generate a transcriptome-wide expression profile, d) Providing the transcriptome-wide expression profiles as a dataset, A method wherein the first biological sample consists of isolated platelets.
2. The method according to claim 1, wherein the above method is repeated at least once.
3. The method according to claim 1, wherein the above method is repeated with a second biological sample, and the second biological sample consists of isolated platelets.
4. The method according to claim 3, wherein the first biological sample and the second biological sample are obtained from the same or different subjects.
5. The method according to claim 4, wherein the first biological sample and the second biological sample are obtained from the same subject at a first time point and a second time point.
6. The method according to claim 5, wherein the first time point and the second time point are different time points.
7. The method according to claim 6, wherein the transcriptome-wide expression profile from the first time point is compared with the transcriptome-wide expression profile from the second time point.
8. The method according to any one of the prior claims, further comprising repeating those steps until an effective transcriptome-wide expression profile is identified.
9. The method according to claim 1, wherein the expression level is measured by a device selected from the group consisting of microarrays, bead arrays, liquid arrays, and nucleic acid sequences.
10. The method according to claim 4, wherein steps a) to d) are repeated for each biological sample.
11. The method according to claim 10, wherein the transcriptome-wide expression profiles of the subject are compared.
12. The method according to claim 5, further comprising obtaining the first biological sample and the second biological sample from different subjects, and obtaining an additional biological sample from the subject, wherein the additional biological sample is obtained at different points in time.
13. The method according to claim 12, wherein the transcriptome-wide expression profile from the first time point is compared with the transcriptome-wide expression profile from the second time point.
14. A method for identifying gene expression differences between two transcriptome-wide expression profiles, comprising determining one or more variations of a target platelet transcriptome using the method of claim 0, wherein the method is performed at different time points, thereby identifying and determining gene expression differences.
15. A method for measuring time-dependent gene expression differences in platelets derived from a subject, comprising determining one or more changes in the platelet transcriptome of the subject using the method according to claim 0, wherein the method is performed at different points in time, and thereby includes identifying and determining gene expression differences and comparing the time-dependent changes from the subject to a reference.