Exploring the host response in infected lung organoids by processing gene expression data
The method computes magnitude-altitude scores from gene expression data in lung organoids to address the challenge of identifying infections in conventional lung organoid analysis techniques, enabling effective clinical diagnostics and treatment decisions.
Patent Information
- Application Number
- PCT/US2024/061732
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-23
- Publication Date
- 2025-06-26
AI Technical Summary
Conventional lung organoid analysis techniques lack effective methods for quantifying gene expression activity in various pathogenic conditions and across time, making it difficult to quickly and accurately identify infections suitable for clinical diagnostics.
A computer-implemented method and system for computing magnitude-altitude scores corresponding to genes expressing in various ways during one or more time periods, using statistical results from differential gene expression analysis in organ tissue equivalents under varying infection conditions.
Enables the quick and accurate identification of viral infections and tracks the severity of respiratory viral infections, providing timely and appropriate medical treatment decisions.
Smart Images

Figure US2024061732_26062025_PF_FP_ABST
Abstract
Description
Attorney Docket No.32103 / 59579 / PC EXPLORING THE HOST RESPONSE IN INFECTED LUNG ORGANOIDS BY PROCESSING GENE EXPRESSION DATA CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Application No.63 / 614,473, filed December 22, 2023, and entitled "EXPLORING THE HOST RESPONSE IN INFECTED LUNG ORGANOIDS BY PROCESSING GENE EXPRESSION DATA," which is incorporated herein by reference in its entirety. STATEMENT OF GOVERNMENT SUPPORT
[0002] This invention was made with U.S. Government support under Base Agreement No. W15QKN- 16-9-1002, awarded by the ACC-NJ to the MCDC. The Government has certain rights in the invention. FIELD OF THE DISCLOSURE
[0003] The present disclosure is generally directed to methods and systems for exploring the host response in infected lung organoids (e.g., using NanoString technology, RNASeq technology, etc.) to process gene expression data, and more particularly for computing magnitude-altitude scores corresponding to genes expressing in various ways during one or more time periods during which we compare at least one organ tissue equivalent in different infection conditions to conduct differential expression analysis. BACKGROUND
[0004] Human microphysiological systems are three-dimensional, bioengineered constructs containing multiple cell types that are organized into tissue-like structures. The improved morphology, functional maturation, and physiological fidelity of these constructs renders them a valuable tool for studying organ function under healthy and diseased conditions.
[0005] Organoids are three-dimensional ex vivo cultures that reproduce the biological structures, interactions and arrangements found in living tissues. Organoids may be produced using tissue-specific adult stem cells sourced from biopsy samples or cadavers, or from pluripotent stem cell populations, such embryonic stem cells or induced pluripotent stem cells. While many organoid protocols rely on the expansion, differentiation and self-assembly of pluripotent stem cells into emergent structures, recent work has shown that combining primary adult stem cells with extracellular matrices can create micro- engineered organoids which recreate the architecture of a specific tissue. Indeed, scaffold-based “organ-tissue equivalents” (OTEs) exhibit several advantages over scaffold-free organoids or two- dimensional cell cultures, including accelerated functional maturation; defined phenotypic composition; microscale tissue architecture; and recapitulation of functionally important, organ-specific cell-cell, cell- extracellular matrix (ECM) and mechanical interactions.Attorney Docket No.32103 / 59579 / PC
[0006] Generally, organoids are intricate in vitro models formed through tissue engineering, that faithfully reproduce the complex structure and functionality of corresponding in vivo tissues, and find wide-ranging applications in human tissue development, repair, diagnostics, disease modeling, drug discovery, and personalized medicine. Organoids may be derived from pluripotent or tissue-resident stem cells, and emulate diverse tissues, including tumors, and offer a comprehensive platform for investigating tissue biology. While conventional methods involve pluripotent stem cell self-assembly, recent innovations have introduced novel approaches, including exploration of patient-derived organoid cultures to study colorectal cancer (CRC) mechanisms, revealing how culture methods influence gene expression patterns. Other studies have dissected human cornea organoids and donor corneas, elucidating early developmental states and their relevance for corneal diseases. Still further studies have generated cerebral organoids from individuals with bipolar disorder, uncovering gene expression differences and synaptic impairments. Other studies have employed lung organoids to study SARS- CoV-2 infection, revealing virus-cell interactions and testing therapeutic interventions. Integrating adult stem cells and matrices, micro-scale organ-tissue equivalents replicate specific tissue characteristics, accelerating functional maturation and enabling interactions unique to each organ. Commercial sources and tissue-specific media continue to expand OTE possibilities.
[0007] Biomedical professionals often use the PCR method for mRNA gene expression quantification. However, many challenges exist involving PCR because it requires the reverse transcription and amplification of sample RNA. Thus, biomedical researchers began to look for new gene expression quantification methods. The recent growth in commercial sources offering primary human cells and technical advances in tissue-specific growth media have made it possible to create and evaluate OTEs for a large variety of human organs.
[0008] One emerging technology is the NanoString nCounter which does not require reverse transcription and amplification like PCR. Some studies have shown that the NanoString nCounter had far better gene expression quantification results compared to PCR, as discussed for example in Reis, Patricia P et al. “mRNA transcript quantification in archival samples using multiplexed, color-coded probes.” BMC biotechnology vol.1146.9 May.2011, doi:10.1186 / 1472-6750-11-46, incorporated by reference herein in its entirety for all purposes. It should be appreciated that the present techniques may use, in addition to, or alternative to NanoString, any platform capable of multiplexed, quantitative gene expression analysis that uses either hybridization-based detection, amplification-based quantification and / or direct molecular counting.
[0009] As noted, organoids accurately replicate internal processes and reactions. These complex structures are increasingly being used to study human physiology and disease. One method to compare the effects of perturbations on organoid gene expression is by direct quantitation of mRNA moleculesAttorney Docket No.32103 / 59579 / PC using target-specific, color-coded probe pairs. This method has the advantage of being quick and relatively low-cost compared to full transcriptomic sequencing. For example, the NanoString platform allows for highly multiplexed quantitation of mRNA targets using labeled probes, enabling the simultaneous processing of gene expression for hundreds of transcripts. By comparing the gene expression patterns of a focused panel of genes between different categories of organoids, researchers can infer functional differences in the responsiveness of OTEs to perturbation. The utilization of NanoString technology in organoid research highlights the power of modern molecular biology techniques to advance knowledge of complex biological systems.
[0010] One specific area utilizing organoids is the gastrointestinal field, as discussed in Dieterich, Walburga et al. “Intestinal ex vivo organoid culture reveals altered programmed crypt stem cells in patients with celiac disease.” Scientific reports vol.10,13535.26 Feb.2020, doi:10.1038 / s41598-020- 60521-5, incorporated by reference herein in its entirety for all purposes. This study determined that celiac disease patients and healthy controls had vastly different gene expression patterns. Numerous ECM proteins, such as COL1A2, COL3A1, FN1, and TIMP3, showed limited expression patterns in CD patients compared to healthy patients.
[0011] Another study utilized NanoString technology to highlight potential immunotherapeutic antigen targets to treat melanoma, as discussed in Beard, Rachel E et al. “Gene expression profiling using nanostring digital RNA counting to identify potential target antigens for melanoma immunotherapy.” Clinical cancer research : an official journal of the American Association for Cancer Research vol.19,18 (2013): 4941-50. doi:10.1158 / 1078-0432.CCR-13-1253, incorporated by reference herein in its entirety for all purposes. Using NanoString technology, the authors identified seven potential genes that could help treat metastatic cancer. These genes were overexpressed in the melanoma cell lines and metastatic melanoma tumors and underexpressed in normal tissues. They found that normal tissues and tumor tissues have different gene expression profiles, and that a number of genes were differentially expressed in normal and tumor tissues. From those target genes, the authors identified seven genes as potential immunotherapeutic approaches to treat melanoma: CSAG2, MAGEA3, MAGEC2, IL13RA2, PRAME, CSPG4, and SOX10.
[0012] Further, cerebral organoids and NanoString technology were used to compare gene expression levels in wild-type and DISC1-disrupted organoids, as discussed in Daniel K. Crawford, Jasper Mullenders, Johanna Pott, Sylvia F. Boj, Shira Landskroner-Eiger, Matthew M. Goddeeris, Targeting G542X CFTR nonsense alleles with ELX-02 restores CFTR function in human-derived intestinal organoids, Journal of Cystic Fibrosis, Volume 20, Issue 3, 2021,Pages 436-442, ISSN 1569- 1993, https: / / doi.org / 10.1016 / j.jcf.2021.01.009, incorporated by reference herein in its entirety for all purposes. Among other things, this study showed that many of the gene expression changes found inAttorney Docket No.32103 / 59579 / PC DISC1 mutation could be replicated by WNT agonism, and that BRN2, CALB2, and EAAT2 decreased expression, and OLFM1, CALB1, FEZF2, and NRG1 increased expression in DISC1-mutated organoids compared to the wild-type organoids.
[0013] Finally, another study involving perfused human kidney organoids transplanted into rats investigated the therapeutic action of GFB-887, a new kidney drug, in phase 2 clinical trials, as discussed in Westerling-Bui, Amy D et al. “Transplanted organoids empower human preclinical assessment of drug candidate for the clinic.” Science advances vol.8,27 (2022): eabj5633. doi:10.1126 / sciadv.abj5633, incorporated by reference herein in its entirety for all purposes.
[0014] Molecular classification of tumors is an effective approach for determining the suitable treatment for bladder cancer patients. However, the technology needed for molecular classification is expensive and rarely available in the medical field. Thus, biomedical professionals are searching for new techniques that easily categorize the molecular subtypes of bladder cancer tumors, as discussed in Lopez-Beltran, Antonio et al. “Molecular Classification of Bladder Urothelial Carcinoma Using NanoString-Based Gene Expression Analysis.” Cancers vol.13,215500.1 Nov.2021, doi:10.3390 / cancers13215500, incorporated by reference herein in its entirety for all purposes. This research shows that NanoString technology has the potential to be a cheap, easy, and largely reproducible technique to distinguish the different molecular subtypes of bladder cancer.
[0015] In sum, conventional lung organoid analysis techniques are lacking techniques for quantifying the expression activity of genes in various pathogenic conditions and across time. Thus, the conventional techniques are lacking in approaches (e.g., diagnostic tools) for quick and accurate identification of infections that are suitable for clinical diagnostics (e.g., tracking severity of respiratory viral infections).
[0016] Therefore, there is an opportunity for improved tests, tools and techniques for assessing the health status of individuals, identifying diseases / conditions, monitoring disease progression and determining severity of illness by assessing infection impact on patient respiratory system and health. This information is critical for guiding appropriate medical treatment and intervention decisions. BRIEF SUMMARY
[0017] In one aspect, a computer-implemented method for computing magnitude-altitude scores corresponding to genes expressing in various ways during one or more time periods during which at least one organ tissue equivalent transits from one or more conditions includes (1) receiving, via one or more processors, statistical results including one or more raw p-values that (a) are indicative of the differential gene expression in the organ tissue equivalent under varying infection conditions, from an initial condition to a subsequent condition, over a period of time, and (b) describe significance of geneAttorney Docket No.32103 / 59579 / PC expression differences from the initial condition to the subsequent condition; (2) generating, via one or more processors, a respective relaxed magnitude-altitude score for one or more genes identified as significant in the statistical results, by computing pairwise comparisons of general linear model quasi likelihood relaxed magnitude altitude scores (GLMQL-RMAS) across diverse conditions; and (3) displaying, via one or more processors, at least some of the GLMQL-RMAS scores for the one or more genes on a display.
[0018] In another aspect, a computer system for computing magnitude-altitude scores corresponding to genes expressing in various ways during one or more time periods during which at least one organ tissue equivalent transits from one or more conditions includes one or more processors, and one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computer system to: (1) receive statistical results including one or more raw p-values that (a) are indicative of the differential gene expression in the organ tissue equivalent under varying infection conditions, from an initial condition to a subsequent condition, over a period of time, and (b) describe significance of gene expression differences from the initial condition to the subsequent condition; (2) generate a respective relaxed magnitude-altitude score for one or more genes identified as significant in the statistical results, by computing pairwise comparisons of general linear model quasi likelihood relaxed magnitude altitude scores (GLMQL-RMAS) across diverse conditions; and (3) displaying, via one or more processors, at least some of the GLMQL-RMAS scores for the one or more genes on a display.
[0019] In yet another aspect, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause a computer to: (1) receive statistical results including one or more raw p-values that (a) are indicative of the differential gene expression in the organ tissue equivalent under varying infection conditions, from an initial condition to a subsequent condition, over a period of time, and (b) describe significance of gene expression differences from the initial condition to the subsequent condition; (2) generate a respective relaxed magnitude-altitude score for one or more genes identified as significant in the statistical results, by computing pairwise comparisons of general linear model quasi likelihood relaxed magnitude altitude scores (GLMQL-RMAS) across diverse conditions; and (3) display, via one or more processors, at least some of the GLMQL-RMAS scores for the one or more genes on a display.
[0020] In one aspect, a computer-implemented method for computing magnitude-altitude scores includes: (1) receiving, via one or more processors, statistical results including raw p-values indicative of differential gene expression during various conditions; (2) generating, via one or more processors, a relaxed magnitude-altitude score for identified significant genes by computing pairwise comparisons ofAttorney Docket No.32103 / 59579 / PC scores across conditions; and (3) displaying, via one or more processors, some of the scores for the genes on a display.
[0021] In another aspect, a computer system for computing magnitude-altitude scores includes one or more processors and memories that includes: (1) receive statistical results including raw p-values indicative of differential gene expression during various conditions; (2) generate a relaxed magnitude- altitude score for identified significant genes by computing pairwise comparisons of scores across conditions; and (3) display some of the scores for the genes on a display.
[0022] In yet another aspect, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause a computer to: (1) receive statistical results including raw p-values indicative of differential gene expression during various conditions; (2) generate a relaxed magnitude-altitude score for identified significant genes by computing pairwise comparisons of scores across conditions; and (3) display some of the scores for the genes on a display.
[0023] In still another aspect, a computer-implemented method for training a machine learning model using multi-scale gene expression data includes: collecting molecular biology data over various time points for baseline and treated conditions; applying a training-validation step to the collected data; utilizing cross-validation techniques for model training; applying an algorithm to identify a set of genes related to infection and time specific for each fold in the cross-validation; and selecting an optimal machine learning classifier based on the identified genes and its performance evaluated on a hold-out validation set.
[0024] In a further aspect, a computer-implemented method for testing a trained machine learning model on multi-scale gene expression data includes: holding out data to form a test set; applying the selected machine learning classifiers to an RNA-Seq dataset; and testing the performance of the machine learning classifiers using a validation technique, considering different scales in the dataset.
[0025] In another further aspect, a computer-implemented method for identifying differentially expressed genes under infection conditions using generalized linear models and quasi-likelihood methods includes: collecting RNA-Seq data from organ tissue equivalents infected with a virus; employing models with tests to analyze the RNA-Seq data; introducing algorithms to navigate the RNA- Seq data; and identifying genes differentially expressed under various infection conditions as potential biomarkers.
[0026] In an additional aspect, a system for identifying gene signatures and pathways associated with virus-induced gastrointestinal complications includes: a data processing unit configured to receive RNA- Seq data from colon organoid models exposed to the virus; a machine learning module connected to the data processing unit, trained to analyze the RNA-Seq data to identify gene signatures and pathwaysAttorney Docket No.32103 / 59579 / PC indicative of gastrointestinal complications caused by the virus; and an output interface configured to display the identified gene signatures and pathways.
[0027] In a further additional aspect, a computer-implemented method for comparing gene expression analysis techniques includes: obtaining a first set of gene expression data using a NanoString analysis technique; obtaining a second set of gene expression data using an RNA-Seq analysis technique; processing both sets of gene expression data to normalize the datasets; comparing the normalized gene expression data obtained from both techniques; and determining differences in gene expression measurements between the two techniques.
[0028] In yet another additional aspect, a computer-implemented method for analyzing RNA-seq data for breast cancer includes: receiving RNA-seq data derived from breast cancer tissue samples; applying an algorithm to the RNA-seq data to identify patterns and correlations specific to breast cancer; utilizing a technique to refine the analysis by selecting the most predictive models of breast cancer based on the identified patterns and correlations; and generating a report that includes the results of the analysis, identifying key genetic markers and potential therapeutic targets for breast cancer.
[0029] Advantages will become more apparent to those of ordinary skill in the art from the following description of the preferred embodiments which have been shown and described by way of illustration. As will be realized, the present embodiments may be capable of other and different embodiments, and their details are capable of modification in various respects. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] FIG.1 depicts a block diagram of an exemplary computing environment for implementing the present techniques, according to some aspects.
[0031] FIG.2A depicts a block diagram of an exemplary computer-implemented method for computing statistically significant genes for viruses as organ-tissue equivalents move between conditions, according to some aspects.
[0032] FIG.2B depicts a block diagram of an exemplary computer-implemented method for applying a Benjamini-Hochberg adjustment procedure to the method of FIG.2A, according to some aspects.
[0033] FIG.2C depicts a block diagram of an exemplary computer-implemented method for computing magnitude-altitude scores corresponding to genes expressing in various ways during one or more time periods during which at least one organ tissue equivalent is compared under varying conditions, according to some aspects.Attorney Docket No.32103 / 59579 / PC
[0034] FIG.3A depicts an exemplary three-dimensional visualization of OTEs over time, according to some aspects.
[0035] FIG.3B depicts an exemplary three-dimensional visualization of OTEs over time, according to some aspects.
[0036] FIG.3C depicts average points of all replicates pinned on the same 3-D PCs coordinate system in which replicates are projected, according to some aspects.
[0037] FIG.3D depicts a three-dimensional visualization of sample classification using IFIT1, IFIT2, and ELOVL4 as axes, underscoring the distinct clustering of samples based on infection and time- dependent gene expression patterns, facilitated by the GLMQL-RMAS methodology.
[0038] FIG.3E depicts a confusion matrix for an 8-class classification using IFIT1, IFIT2, and ELOVL4 as predictors, wherein the matrix highlights the model’s accuracy in distinguishing between different viral infections and time points, reflecting a mean accuracy of 92% as achieved through a 6-fold stratified cross-validation process.
[0039] FIG.3F depicts a confusion matrix for the 8-class classification after replacing RMAS ranking with rankings based on the smallest p-value (EdgeR ranking method).
[0040] FIG.3G depicts a confusion matrix for the 8-class classification when genes are ranked based on logFC (EdgeR ranking method).
[0041] FIG.4A depicts an exemplary line chart illustrating a BH-cutoff line and a number of BH-based significant genes of the first two pair conditions for IAV, across a 24-hour period.
[0042] FIG.4B depicts an exemplary line chart illustrating a BH-cutoff line and a number of BH-based significant genes of the first two pair conditions for IAV, across a 72-hour period.
[0043] FIG.5 depicts an exemplary bar chart illustrating positive false discovery rates of the BH adjustment procedure for a period of time, according to some aspects.
[0044] FIG.6A depicts an exemplary volcano plot of hyperparameter tuning when (M,A) is set to (0,1), according to some aspects.
[0045] FIG.6B depicts an exemplary volcano plot of hyperparameter tuning when (M,A) is set to (1,0), according to some aspects.
[0046] FIG.6C depicts an exemplary volcano plot of hyperparameter tuning when (M,A) is set to (1,1), according to some aspects.Attorney Docket No.32103 / 59579 / PC
[0047] FIG.7A depicts an exemplary bar chart illustrating IAV affected genes as OTEs in negative and experimental control groups are contrasted, according to some aspects.
[0048] FIG.7B depicts an exemplary bar chart illustrating IAV affected genes as OTEs in negative and experimental control groups are contrasted, according to some aspects.
[0049] FIG.7C depicts an exemplary bar chart illustrating IAV affected genes as OTEs in negative and experimental control groups are contrasted, according to some aspects.
[0050] FIG.7D depicts an exemplary bar chart illustrating PIV3 affected genes as OTEs in negative and experimental control groups are contrasted over time, according to some aspects.
[0051] FIG.7E depicts an exemplary bar chart illustrating PIV3 affected genes as OTEs in negative and experimental control groups are contrasted over time, according to some aspects.
[0052] FIG.7F depicts an exemplary bar chart illustrating PIV3 affected genes as OTEs in negative and experimental control groups are contrasted over time, according to some aspects.
[0053] FIG.7G depicts an exemplary bar chart illustrating MPV affected genes as OTEs in negative and experimental control groups are contrasted over time, according to some aspects.
[0054] FIG.8A depicts an exemplary plot of common BH-significant genes among all viruses as OTEs transit UV-treated conditions over time, according to some aspects.
[0055] FIG.8B depicts an exemplary plot of common BH-significant genes among all viruses as OTEs in UV-treated and experimental conditions are compared over time, according to some aspects.
[0056] FIG.8C depicts an exemplary three-dimensional visualization of OTEs showing separation of replicates of particular genes transit UV-treated conditions over time, according to some aspects.
[0057] FIG.8D depicts an exemplary three-dimensional visualization of OTEs showing separation of replicates of particular genes transit UV-treated conditions over time, according to some aspects.
[0058] FIG.8E depicts an exemplary three-dimensional visualization of OTEs showing separation of replicates of particular genes transit UV-treated conditions over time, according to some aspects.
[0059] FIG.8F depicts an exemplary three-dimensional visualization of OTEs showing separation of replicates of particular genes transit UV-treated conditions over time, according to some aspects.
[0060] FIG.8G depicts an exemplary three-dimensional visualization of OTEs showing separation of replicates of particular genes transit UV-treated conditions over time, according to some aspects.
[0061] FIG.9A depicts an exemplary computer-implemented exploratory analysis visualization plot, according to some aspects.Attorney Docket No.32103 / 59579 / PC
[0062] FIG.9B depicts an exemplary computer-implemented differential analysis visualization plot, according to some aspects.
[0063] FIG.9C depicts an exemplary plot showing a comparison of the activity of IAV over time, relative to its Mock-infected counterpart, according to some aspects.
[0064] FIG.9D depicts an exemplary plot showing a comparison of the activity of MPV over time, relative to its Mock-infected counterpart, according to some aspects.
[0065] FIG.9E depicts an exemplary plot showing a comparison of the activity of PIV3 over time, relative to its Mock-infected counterpart, according to some aspects.
[0066] FIG.9F depicts an exemplary plot showing a comparison over time of virus activity over time, according to some aspects.
[0067] FIG.10A depicts a PCA-based visualization of 96 OTEs studied, according to some aspects.
[0068] FIG.10B depicts UV versus None-UV and 24-hours versus 72-hours identified pattern for IAV- infected OTEs, according to some aspects.
[0069] FIG.10C depicts a heatmap showing quasi-likelihood F-test of selected genes, according to some aspects.
[0070] FIG. 10D depicts a distribution of upregulated and downregulated significant genes in IAV, MPV, and PIV3 infected samples compared to Mock samples at 24 h post-infection, using raw p-values with a significance level of α = 0.05, according to some aspects.
[0071] FIG. 10E depicts a distribution of upregulated and downregulated significant genes in IAV, MPV, and PIV3 infected samples compared to Mock samples at 72 h post-infection, using raw p-values with a significance level of α = 0.05, according to some aspects.
[0072] FIG.10F depicts Venn diagrams showing the overlap of significant genes associated with IAV, MPV, and PIV3 infections compared to Mock samples at 24 / 72 h post-infection, wherein genes are ranked based on RMAS scores and aggregated RMAS rankings to highlight their significance across the groups ((A) At 24 h post-infection; (B) At 72 h post-infection).
[0073] FIG.10G depicts Venn diagram of significant upregulated genes with logFC >1 at 72 h post- infection; consistent with the previous figures, FIG 10G demonstrates the dependability of RMAS ranking in maintaining the identification of the top gene despite varying logFC thresholds, wherein (A) is a Venn diagram of significant upregulated genes with logFC >0 at 24 h post-infection; (B) is a Venn diagram of significant upregulated genes with logFC > 0 at 72 h post-infection; (C) is a Venn diagram of significantAttorney Docket No.32103 / 59579 / PC upregulated genes with logFC > 1 at 24 h post-infection; and (D) is a Venn diagram of significant upregulated genes with logFC > 1 at 72 h post-infection, all according to various aspects.
[0074] FIG.11A depicts a GLMQL-MAS-based Volcano Plot: Mock-24 to IAV-None-24, according to some aspects.
[0075] FIG.11B depicts a GLMQL-MAS-based Volcano Plot: Mock-72 to IAV-None-72, according to some aspects.
[0076] FIG.11C depicts a GLMQL-MAS-based Volcano Plot: IAV-None-24 to IAV-None-72, according to some aspects.
[0077] FIG. 11D depicts Volcano Plots and Analysis for IAV-Infected OTEs compared to Mock infection at 24 hours post-infection.
[0078] FIG.11E depicts Volcano Plots and Analysis for MPV-Infected OTEs compared to Mock infection at 24 hours post-infection.
[0079] FIG.11F depicts Volcano Plots and Analysis for PIV3-Infected OTEs compared to Mock infection at 24 hours post-infection.
[0080] FIG.11G depicts Volcano Plots and Analysis for IAV-Infected OTEs compared to Mock infection at 72 hours post-infection.
[0081] FIG.11H depicts Volcano Plots and Analysis for MPV-Infected OTEs compared to Mock infection at 72 hours post-infection.
[0082] FIG.11I depicts Volcano Plots and Analysis for PIV3-Infected OTEs compared to Mock infection at 72 hours post-infection.
[0083] FIG.11J depicts a Venn Diagram of Differentially Expressed Genes at 24 Hours Post-Infection, illustrating the overlap and unique DEGs among samples infected with IAV, MPV, and PIV3 compared to UV-treated controls at 24 hours post-infection, wherein each circle represents a set of DEGs associated with a specific viral infection, with overlapping areas indicating common DEGs across different viruses, and the numbers indicate the count of DEGs unique to each condition or shared between conditions, highlighting the gene expression patterns that are specific to each viral treatment and those that are commonly induced by different viruses at the early stage of infection
[0084] FIG.11K depicts a Venn Diagram of Differentially Expressed Genes at 72 Hours Post- Infection, showing the intersection and unique DEGs identified in OTE models infected with IAV, MPV, and PIV3 relative to UV-treated controls at 72 hours post-infection; similar to the 24-hour time point of FIG.11J, the circles correspond to DEGs for each virus, with the intersections revealing DEGs that areAttorney Docket No.32103 / 59579 / PC shared; and the numbers reflect the DEGs that are exclusive to a given viral treatment or common between treatments, providing insights into the transcriptional responses that persist or emerge at the later stage of infection.
[0085] FIG.11L depicts a distribution of upregulated and downregulated significant genes in IAV, MPV, and PIV3 infected samples compared to Mock samples at 24 hours post-infection, adjusted using the Benjamini-Hochberg (BH) method for controlling the false discovery rate.
[0086] FIG.11M depicts a distribution of upregulated and downregulated significant genes in IAV, MPV, and PIV3 infected samples compared to Mock samples at 72 hours post-infection, adjusted using the Benjamini-Hochberg (BH) method for controlling the false discovery rate.
[0087] FIG.11N depicts a distribution of upregulated and downregulated significant genes in IAV, MPV, and PIV3 infected samples compared to Mock samples at 24 hours post-infection, using the Bonferroni correction to adjust for multiple comparisons.
[0088] FIG.11O depicts a distribution of upregulated and downregulated significant genes in IAV, MPV, and PIV3 infected samples compared to Mock samples at 72 hours post-infection, using the Bonferroni correction to adjust for multiple comparisons.
[0089] FIG.11P depicts a Venn diagram illustrating the overlap of BH significant genes in response to IAV, MPV, and PIV3 infection compared to Mock samples at 24 hours post-infection; and a Venn diagram illustrating the overlap of BH significant genes in response to IAV, MPV, and PIV3 infection compared to Mock samples at 72 hours post-infection.
[0090] FIG.11Q depicts a Venn diagram illustrating the overlap of Bonferroni significant genes in response to IAV, MPV, and PIV3 infection compared to Mock samples at 24 hours post-infection; and a Venn diagram illustrating the overlap of Bonferroni significant genes in response to IAV, MPV, and PIV3 infection compared to Mock samples at 72 hours post-infection.
[0091] FIG.12 depicts a block diagram of a computer-implemented method 1200 for computing magnitude-altitude scores corresponding to genes expressing in various ways during one or more time periods, according to some aspects.
[0092] FIG.13 depicts a block diagram of a cross-modal predictive modeling framework employing a Magnitude-Altitude Score Analysis for Tracking Infection and Time-Dependent Genes (MASIT) algorithm, according to some aspects.
[0093] FIG.14A depicts a visual representation illustrating the main objectives of the comparative analysis between RNA-Seq and NanoString platforms, highlighting five key objectives: 1) CorrelationAttorney Docket No.32103 / 59579 / PC Analysis, 2) Bland-Altman Analysis, 3) Regression and Residual Analysis, 4) Concordance Analysis and 5) Gene Ontology Analysis.
[0094] FIG.14B depicts a comparison of spearman and distance correlation coefficients across infection conditions, according to some aspects.
[0095] FIG.14C depicts Bland-Altman plots for RNA-Seq and NanoString gene expression comparisons, according to some aspects.
[0096] FIG.14D depicts Bland-Altman plots comparing gene expression measurements from RNA - Seq and NanoString platforms for each condition, according to some aspects.
[0097] FIG.14E depicts a comparison of significant genes identified using the MAS algorithm in both RNA-Seq and NanoString datasets across various infection condition.
[0098] FIG.14F depicts volcano plots illustrating gene expression contrasts between multiple control and experimental groups within the RNA-Seq and NanoString datasets.
[0099] FIG.14G displays the top 50 differentially expressed genes identified using the MAS algorithm, highlighting the common genes between the RNA-Seq and NanoString platforms within this top-50 list.
[0100] FIG.14H depicts examples of the top 50 differentially expressed genes identified using the MAS algorithm, according some aspects.
[0101] FIG.14I presents the top 10 out of 567 significant Gene Ontology (GO) processes based on q- values for BH-significantly upregulated genes in the IAV-None-24 versus Mock-24 comparison across both platforms.
[0102] FIG.14J shows the top 10 of 92 significant GO processes for BH-significantly upregulated genes in the PIV3-None-24 versus Mock-24 comparison across both platforms.
[0103] FIG.14K displays the top 10 of 573 significant GO processes, based on q-values for BH- significantly upregulated genes in the IAV-None-72 versus Mock-72 comparison across both platforms.
[0104] FIG.14L illustrates the top 10 of 464 significant GO processes for BH-significantly upregulated genes in the PIV3-None-72 versus Mock-72 comparison across both platforms.
[0105] FIG.15A illustrates the projection of all subjects’ TMM-normalized gene expression onto a two- dimensional plane, defined by t-SNE coordinates Tsne1 and Tsne2.
[0106] FIG.15B presents a volcano plot of genes identified as BH-significant through the GLMQL- MAS system.
[0107] FIG.15C displays the projection of all subjects on TSNE1 and TSNE2 axes.Attorney Docket No.32103 / 59579 / PC
[0108] FIG.15D illustrates the number of upregulated and downregulated BH-significant genes identified by the GLMQL-MAS.
[0109] FIG.15E illustrates the effectiveness of the GLMQL-MAS methodology in enhancing the discrimination capabilities of a logistic regression model used for analyzing gene expression data.
[0110] FIG.15F depicts a Distribution of the top 20 Genes by BH-Significance Occurrence Using GLMQL-MAS. This figure illustrates the top-20 genes that achieved the highest frequency of BH- significance in an analysis of 500 iterations, with |LogFC| upper thresholds set at 1. It highlights the genes that consistently demonstrate significant differential expression in lymph node metastasis of breast cancer.
[0111] FIG.15G depicts top 10 GO Processes Related to Lymph Node Metastasis. This figure details the GO processes most intimately connected with the pathology of lymph node metastasis, providing insights into the molecular functions and cellular components affected.
[0112] FIG.15H illustrates a selection of the significant pathways identified from the GSEA using Hallmark gene sets, highlighting the predominant biological mechanisms influenced by the GLMQLMAS selected genes between 42 ALNM positive (ALNM+) and 65 ALNM negative (ALNM−) breast cancer samples.
[0113] FIG.15I depicts top 100 consistent genes in random selections meeting GLMQL-MAS criteria displayed in one of the 50 GSEA Hallmark sets, specifically highlighting the key genes in cancer progression and metastasis in red.
[0114] FIG.15J depicts common gene ontology (GO) processes and their associated upregulated BH- significant genes from analyses of 42 ALNM+ samples compared to ALNM− samples in two scenarios: Case 1, comparing 65 ALNM− versus 42 ALNM+; and Case 2, involving a random sampling of 42 ALNM− to compare against the same 42 ALNM+.
[0115] FIG.15K depicts a network that maps the most relevant Hallmark pathways to breast cancer, highlighting the common GLMQL-MAS selected genes from both Case 1 and Case 2 of FIG.15J.
[0116] FIG.15L depicts a flowchart of a computer-implemented method for refining the initial dataset of 1078 patients down to 104 untreated individuals based on the availability of ALNM information and absence of prior treatment. The final cohort is categorized into two groups: 65 without ALNM (ALNM−) and 42 with ALNM (ALNM+), facilitating the study of genetic and molecular markers associated with lymph node metastasis.Attorney Docket No.32103 / 59579 / PC
[0117] FIG.16A depicts representative hematoxylin and eosin (H&E) stained tissue images from ALNM+ (top row) and ALNM- (bottom row) of patients from the experimental cohort, according to some aspects.
[0118] FIG.16B depicts a computer-implemented method of an SMAS process flow, from data division and normalization through to gene selection and validation, according to some aspects. Key steps illustrate the methodical approach to minimizing overfitting risks while maximizing predictive accuracy, showcasing the statistical and methodological strengths of the GLMQL-MAS strategy in analyzing gene expression data for lymph node metastasis in breast cancer.
[0119] FIG.16C depicts a principal component analysis showing the distribution of samples in a two- dimensional space defined by the first two principal components before TMM and GLMQL-MAS are applied.
[0120] FIG.16D depicts a Receiver Operating Characteristic Area Under the Curve (ROCAUC) for various predictive models: Supervised GLMQL-MAS (SMAS), Logistic Regression, SVM, Decision Tree, and Random Forest with 3 trees and a maximum depth of 8.
[0121] FIG.16E depicts a projection of samples onto the first two principal components (PC1 and PC2) after log transformation and Z-score normalization, based on genes selected by the GLMQL-MAS.
[0122] FIG.16F depicts the top 100 genes identified as most important by the SMAS model for ALNM prediction, as analyzed across all five folds using cumulative normalized scores (details in Algorithm 2). The red bars specifically highlight those transcripts that were significantly upregulated, as determined by the Benjamini-Hochberg procedure across all five folds.
[0123] FIG.16G depicts top 25 Gene Ontology processes related to the genes chosen by SMAS with positive cumulative normalized scores are illustrated in this figure.
[0124] FIG.17A depicts a timeline of responses for different nonhuman primates (NHPs) to EBOV exposure, according to some aspects.
[0125] FIG.17B encapsulates the structured approach of Objective 1, according to some aspects.
[0126] FIG.17C depicts analysis and visualization of gene expression and predictive modeling in Objective 1.1, wherein (A) depicts a volcano plot that contrasts gene expressions of NHPs with strong positive RT-qPCR results against their status on day 0, when RT-qPCR results were strongly negative, highlighting top MAS-selected genes with a logFC >1; (B) depicts a three-dimensional visualization is provided, showing all strongly positive and negative NHPs using the top three selected MAS genes: IFI27, IFI6, and HP; (C) depicts the Receiver Operating Characteristic (ROC) curves for each of the five folds of the five-fold stratified cross-validation, along with their mean, using logistic regression on the topAttorney Docket No.32103 / 59579 / PC selected gene, IFI27. This panel also includes AUC and accuracy metrics for each fold, offering a detailed assessment of the predictive model’s performance. The focus on a single gene and linear model emphasizes the model’s robust generalizability, demonstrating the effectiveness of IFI27 as a predictive marker.
[0127] FIG.17D depicts an overview of gene expression analyses in Objective 1.2, wherein (A) shows a volcano plot comparing gene expressions of NHPs on day 6 vs. day 0, highlighting significant changes; (B) shows a similar plot for day 6 vs. day 3, pinpointing early infection response genes; (c) illustrates the lack of BH significant genes when comparing day 3 to day 0; (D) reveals FLT1 as a top significant gene with relaxed BH criteria; and (E) depicts a comparison of the expression patterns of IFI27 (from Objectives 1.1 and 1.2), IFI6 (from Objective 1.1), and FLT1, emphasizing their roles across different infection stages.
[0128] FIG.17E depicts a gene expression analysis and classification in Objective 1.3, wherein (A) depicts Venn Diagram of BH Significant Genes—this diagram compares NHPs on day 10 vs. days 0 and 3, showcasing BH significant genes with logFC >1, highlighting highlights the top genes (ISG15, IFI6, IFI44) expressed significantly in both contrasts and the lack of unique significant genes for the day 10 vs. day 3 comparison; (B) depicts unique BH significant genes across objectives—this figure reveals IFI27 as the only BH significant gene with logFC >1 when comparing ND-0-10, N-0-6, and NDL-0-strong, underscoring its unique presence across different comparisons; (C) depicts common BH significant genes across objectives—displaying IFI6, HP, and IFI27 as the top three BH significant genes with logFC >1, this figure shows the commonality of these genes among ND-0-10, ND-3-10, and NDL-0- strong comparisons; (D) depicts ROC AUC for Classifying NHPs Using SMAS Predictors—this graph illustrates the ROC AUC for classifying NHPs as negative (RT-qPCR on day 0) or positive (day 10) using IFI6, IFI27, or ISG15 (only one gene) as the sole predictor in SMAS; and (E) depicts ROC AUC for classifying NHPs using SMAS predictors—this graph illustrates the ROC AUC for classifying NHPs as negative (RT-qPCR on day 0) or positive (day 10) using IFI6, IFI27, or ISG15 (only one gene) as the sole predictor in SMAS.
[0129] FIG.17F depicts analysis and visualization of gene expression for objective 1.4, wherein (A) shows a volcano plot contrasting all positive vs. negative NHPs in RT-qPCR, highlighting BH significant genes with logFC >1 using MAS ranking; (B) illustrates the ROC AUC from a five-fold stratified cross- validation using the top gene, OAS1, in the Simplified MAS (SMAS) approach; (C) depicts the visualization of all samples using the first three principal components (PC1, PC2, PC3); and (D) presents a 3D visualization of all samples using the top three MAS-selected genes, offering insights into their spatial distribution and expression patterns.Attorney Docket No.32103 / 59579 / PC
[0130] FIG.17G depicts gene expression analysis and predictive modeling in Objective 2, wherein (A) depicts the selected MAS genes for All-N-P, NDL-0-strong, ND-0-10, and ND-3-10, and due to the complexity of presenting all common and unique genes across these groups, this figure specifically illustrates only those for the first three groups; however, it highlights IFI6, IFI27, and MX1 as the top three common genes shared among all four groups, underscoring their significance across various analysis categories; (B) presents a 3D visualization of gene expression, focusing on the top three selected genes: IFI6, IFI27, and MX1; and (C) depicts a heat map with hierarchical clustering of the top three selected genes: IFI6, IFI27, MX1, and HP.
[0131] FIG.17H depicts a Comparative Analysis of Gene Selection between Traditional Ranking and MAS Methods for Objective 1.1, wherein the figure displays the top 20 genes selected by the traditional ranking method (left) and the MAS method (right) for objective 1.1, and it highlights differences in gene prioritization between the two methods, revealing potential insights into gene selection for distinguishing negative and positive samples in RT-qPCR NHPs.
[0132] FIG.17I depicts five most significant GO terms associated with upregulated genes in nonhuman primates following exposure to the Ebola virus, based on the negative logarithm of the q- values, as determined using the Supervised Magnitude-Altitude Scoring (SMAS) method and adjusted by the Benjamini-Hochberg procedure. The titles within each subplot indicate the total number of significant GO terms found for the respective gene, showcasing their broad impact on cellular processes.
[0133] FIG.17J depicts distribution of Top 10 Significant GO Terms for IFI27, IFI6, MX1, and HP. Each panel represents one of the genes studied, illustrating the GO terms with the highest statistical significance in their association with gene upregulation. The number indicated in each title reflects the total significant GO terms identified for that gene, emphasizing the biological impact of each gene’s expression changes.
[0134] FIG.18A depicts a Schematic Overview of the Research Approach for Dengue Host-Pathogen Interaction Analysis, including the two-fold objective framework of the study: Objective 1 details the comprehensive differential gene expression analysis across three stages of Dengue infection contrasted with the control group, while Objective 2 depicts the development of a binary classification model using the most significant gene identified through Objective 1 for robust Dengue infection classification.
[0135] FIG.18B illustrates the top 20 MAS-selected genes for Objectives 1.1 to 1.3.
[0136] FIG.18C depicts Identification of Key Genes Meeting Benjamini-Hochberg Significance and logFC > 2 Criteria Across Objectives 1.1 to 1.3.Attorney Docket No.32103 / 59579 / PC
[0137] FIG.18D depicts ROC Curves and Performance Metrics for Logistic Regression and SVM with Linear Kernel.
[0138] FIG.19 depicts (a) Identification of Key Genes Meeting Benjamini-Hochberg Significance and logFC > 2 Criteria Across Objectives 1.1 to 1.3; and (b) ROC Curves and Performance Metrics for Logistic Regression and SVM with Linear Kernel.
[0139] FIG.20A depicts enriched ontologies using Enrichr-KG and GO Biological Processes 2021 based on the 112 upregulated DEGs overlapping between all dengue groups (DI, DWS, DS) with logFC > 0.
[0140] FIG.20B depicts STRING version 12.0 Protein-Protein-Interaction network based on the BH- significant DEGs across all three dengue groups (DI, DWS, DS) with logFC > 0.
[0141] FIG.20C depicts identification of key DEGs with BH-adjusted p-values and logFC > 2 unique to each dengue group based on disease progression (DI, DWS, and DS).
[0142] FIG.21A depicts volcano plots to demonstrate the efficacy of different gene ranking systems following GLMQL application on MPXV I infected samples versus Mock samples: (a) RMAS ranking, which integrates both log fold change (LogFC) and p-values to prioritize genes; (b) Ranking based solely on p-values, highlighting genes with the most statistical significance irrespective of effect size; and (c) Ranking based solely on LogFC, emphasizing genes with the greatest expression changes without considering statistical significance.
[0143] FIG.21B depicts a comparison of significant gene selection across different statistical correction methods and LogFC thresholds after applying GLMQL-RMAS / MAS to contrast MPXV I infected samples against Mock (baseline) samples. Panels (a), (b), and (c) illustrate the number and identity of genes determined to be significant (a) using raw p-values, (b) Benjamini-Hochberg correction, and (c) Bonferroni correction, respectively. Each subpanel within (a), (b), and (c) represents varying LogFC thresholds, from 0 to 3 for upregulated genes and 0 to -3 for downregulated genes, highlighting the influence of statistical methodology and threshold settings on the identification of significant genes.
[0144] FIG.21C depicts gene expression analysis results using GLMQL-RMAS, identifying significant upregulated and downregulated genes across various MPXV clade comparisons with Mock, based on raw p-values. The genes are listed in descending order of their RMAS, where the top gene displays the maximum log fold change (logFC) and the minimum p-value, indicating the most significant expression difference in this analysis.
[0145] FIG.21D depicts differential gene expression analysis using GLMQL-MAS, with results adjusted using the Benjamini-Hochberg method. This figure displays significant upregulated andAttorney Docket No.32103 / 59579 / PC downregulated genes across different comparisons between Mock and various MPXV clades. Each gene in the list is ranked according to its MAS score, with the first gene showing the highest differential expression based on the combined criteria of maximum logFC and minimum BH adjusted p-value.
[0146] FIG.21E depicts top 20 GO biological processes associated with genes significantly upregulated in MPXV I compared to Mock, based on raw p-values. This representation includes genes deemed significant before the application of the Benjamini-Hochberg adjustment, allowing for a broader inclusion of differentially expressed genes.
[0147] FIG.21F depicts a Venn diagram illustrating the categorization of significantly upregulated GLMQL-RMAS selected genes (LogFC > 0) based on their uniqueness or overlap among different contrasts (Mock vs. MPXV I, Mock vs. MPXV IIa, and Mock vs. MPXV IIb). The diagram displays the seven potential groups of genes, ranked using the Cross-RMAS method. In this figure, genes are organized by their Cross-RMAS scores; the first gene mentioned has the highest score, reflecting the strongest statistical and biological relevance among the displayed genes in each group.
[0148] FIG.21G depicts utilization of Gene Ontology (GO) terms derived from GLMQL-RMAS upregulated significant genes from the Mock vs. MPXV I contrast. This figure details the selected GO terms for each of the top Cross-RMAS identified genes, highlighting the pathways and biological processes involved.
[0149] FIG.21H depicts confusion matrices depicting the performance of Logistic Regression and SVM with a linear kernel using the first three principal components of top genes identified by Cross- RMAS. These models use One-Versus-One (OVO) and One-Versus-Rest (OVR) strategies to ifferentiate between mock and various MPXV clades.
[0150] FIG.21I depicts classification accuracy using PC1 derived from TFF1, FOS, GSTA5, and HSPA6 gene expressions, demonstrating 100% differentiation between different Mpox clades and mock samples.
[0151] FIG.21J depicts a heatmap showing hierarchical clustering of samples using the top Cross- RMAS identified genes (TFF1, FOS, GSTA5, HSPA6) within TMM normalized RNA-seq data. Clustering employed Euclidean distance and Ward’s linkage method, with data log2 transformed adding a pseudocount of 1.
[0152] FIG.21K depicts a 3D visualization using TFF1, EGR1, and GSTA5 as axes, based on log2 transformed data with a pseudocount of 1. This figure illustrates the distinct separation of the samples, providing a spatial representation of gene expression differences among the different conditions.Attorney Docket No.32103 / 59579 / PC
[0153] FIG.22A depicts an overview of the cross-modal predictive modeling framework employing MASIT. This figure illustrates the structured methodology of our study, detailing the two-phase process of training and validating of on NanoString data followed by testing on RNA-Seq data. It highlights the categorization of data by post-infection times, the application of MASIT for gene selection during training, and the use of optimal linear classifiers for cross-validation and final testing to evaluate the translational efficacy of predictive insights across different genomic technologies.
[0154] FIG.22B depicts Hierarchical clustering of infected OTEs using selected MASIT genes. This figure illustrates the hierarchical clustering of all infected OTEs within the RNA -Seq dataset, using the eight genes identified through the MASIT algorithm applied exclusively to the NanoString dataset. The genes include six infection-dependent genes (IFIT1, IFIT2, IFIT3, OASL, OAS3 and IFI44) marked with blue in the annotation, and two time-dependent genes (IL33, CCL20) highlighted in red.
[0155] FIG.22C depicts the performance accuracies of various models using MASIT-selected genes (IFIT1, CXCL10, and IL33) on both NanoString (left panel) and RNA-Seq (right panel) datasets. The models evaluated include SVM, LR, Random Forest with 5 trees and a maximum depth of MASIT-RF-5- 4 (or another algorithm), XGBoost with 3 trees and a maximum depth of 4 (MASIT-XGBoost-3-4), and AdaBoost with 3 trees and a maximum depth of 4 (MASIT-AdaBoost-3-4).
[0156] FIG.22D depicts the validation accuracies of various models trained using MASIT-selected genes (IFIT1, CXCL10, and IL33) and the entire gene set (19,671 genes) within RNA -Seq data. The models evaluated include SVM with OVO and OVR configurations, Logistic Regression with OVO and OVR configurations, Random Forest with 5 trees and a maximum depth of 4 (RF-5-4), XGBoost with 3 trees and a maximum depth of 4 (XGBoost-3-4), and AdaBoost with 3 trees and a maximum depth of 4 (AdaBoost-3-4)
[0157] FIG.22E depicts a 3D visualization comparing the log2 -transformed expression of IFIT1, CXCL10, and IL33 between NanoString and RNA -Seq datasets. The left panels represent the NanoString data, while the right panels represent the RNA-Seq data. The top panels show all samples, and the bottom panels focus specifically on IAV-infected samples and Mock samples as examples. Each axis corresponds to the log2 -transformed expression levels of one of the three genes, allowing for a visual comparison of their expression levels across different conditions and platforms.
[0158] FIG.23A depicts a computer-implemented method for expanding the identification of virus- responsive genes selected by the MASIT algorithm to include Alternative genes within the RNA-Seq dataset, utilizing Spearman correlation analysis. This analysis focuses on discovering genes not included in the NanoString dataset but showing high correlation with established virus-responsive genes fingerprints, thereby identifying potential novel biomarkers for viral infections.Attorney Docket No.32103 / 59579 / PC
[0159] FIG.23B depicts a correlation analysis between infection-dependent MASIT-selected genes and genes that are both common and Not-in-Common with NanoString, and in particular top-10 correlated genes associated with IFIT1, encompassing both the Common and Not-in-Common sets, highlighting genes that may serve as gene fingerprints similar to IFIT1; and top-10 correlated genes associated with CXCL10, encompassing both the Common and Not-in-Common sets, highlighting genes that may serve as gene fingerprints similar to CXCL10.
[0160] FIG.23C depicts the hierarchical clustering of the top-10 selected genes from the Common and Not-in-Common gene sets associated with IFIT1. The heatmap color codes log2-transformed expression levels to reveal patterns and similarities in gene expression across different infection conditions.
[0161] FIG.23D a 3-dimensional visualization of all Organ Tissue Equivalents (OTEs), where the x-, y-, and z-axes represent the log2-transformed expression levels of selected genes from the Not-in- Common set. The x-axis displays the expression of DDX60, identified as the top Alternative gene corresponding to IFIT1, emphasizing its role in infection analysis. The y-axis showcases USP18, identified as the top Alternative gene to CXCL10, underlining its importance for detecting infection. The z-axis depicts the expression of IL33, which, although a time-dependent gene and not directly related to infection, serves as a reference for temporal variations in gene expression.
[0162] FIG.23E exclusively focuses on the top 10 genes that are not in common (Not-in-Common set), which are correlated with IFIT1. Additionally, it preserves the RMAS ranking from the analysis of the entire set of 19,671 genes, providing a comprehensive and relative overview of each gene’s significance across various infected conditions compared to Mock samples. It showcases volcano plots annotated in the following format: (original RMAS rank) Gene Symbol.
[0163] FIG.23F depicts Comprehensive Gene Ontology Analysis of top-10 Not-in-Common genes correlated to IFIT1.
[0164] FIG.23G Comparative Performance of Predictive Models Using MASIT-Selected and Alternative Genes. The left panel displays the classification accuracies for models using the traditional MASIT-selected genes (IFIT1, CXCL10, and IL33), while the right panel shows results for models utilizing the alternative virus response genes (DDX60, USP18) not included in the NanoString dataset. This comparison highlights the effectiveness of both sets of genes in predicting viral infection states and time points across different machine learning frameworks under a 6-fold cross-validation scheme. DETAILED DESCRIPTIONAttorney Docket No.32103 / 59579 / PC
[0165] The present techniques provide methods and systems for, inter alia, exploring the host response in infected lung organoids (e.g., using NanoString technology) to analyze gene expression data, and more particularly for identifying patterns of gene expression enabling the development of diagnostic tools for quick and accurate identification of viral infections and track the severity of respiratory viral infections.
[0166] The present techniques may include quantifying changes in cellular gene expression of upper airway organ tissue equivalents (OTEs) infected by Influenza A virus (IAV), Human metapneumovirus (MPV), and / or Parainfluenza virus type 3 (PIV3). At a high level, the present techniques may utilize two complementary techniques; specifically, (1) differential expression analysis and (2) exploratory analysis.
[0167] For example, in some aspects, organoids such as lung organoids and / or OTEs may be constructed using techniques described in U.S. Patent No.11,001,811, “Multi-Layer Airway Organoids and Methods of Making and Using the Same,” to Murphy et al, filed May 11, 2021, and hereby incorporated by reference in its entirety for all purposes.
[0168] In particular, the present techniques include methods and systems for assessing influence of IAV, MPV and / or PIV3 on gene expression. In some aspects, the present techniques may utilize a Benjamini-Hochberg (BH) adjustment procedure as part of the assessment. For example, the present techniques may include measuring changes in gene expression IAV, MPV, and PIV3, at different times after infection. The present techniques may include identifying genes whose levels were significantly changed by infection with IAV and PIV3 at specific times after infection, potentially serving as fingerprints for infection by these viruses. The present techniques may further include identifying common patterns of gene expression or fingerprints that may identify infections caused by IAV, MPV, and PIV3, as well as fingerprints that inform on the progression of the viral infection and possibly viral disease. Advantageously, these findings may be used to develop diagnostic tools for quick and accurate identification of viral infections and track the severity of respiratory viral infections, ultimately reducing the impact of these infections on human health.
[0169] As noted, the present techniques may include techniques for analyzing gene expression patterns of respiratory viruses (e.g., IAV, MPV, and PIV3) in infected lung-tissue-equivalents using a statistical techniques. And as noted, in some aspects, the present statistical techniques may be based or utilize the Benjamini-Hochberg adjustment procedure. The present techniques may use principal component analysis (PCA) visualization and a significance gene score to quantify a magnitude and a significance of gene expression changes.
[0170] Empirical results illustrate that IAV had the most significant impact on gene expression compared to the other two viruses. PIV3 significantly affected genes between 24-72 hours post-Attorney Docket No.32103 / 59579 / PC infection, while MPV was the least intense virus during the first 72 hours of infection. Advantageously, the empirical findings enabled by the present techniques and included in this disclosure provide insights into the mechanisms of viral infections and may aid in the development of effective treatments for respiratory diseases caused by these viruses. The present techniques may be applicable to other virus / disease pathogenesis, and may advantageously aid in development of effective treatments for those processes and mechanisms, and specific clinical symptoms or disease manifestations.
[0171] As discussed herein, the present techniques have been used to identify several genes that are significantly affected by infections with IAV, MPV, and PIV3 during different post-infection time frames. These results show that IFNL1 and CCL20 are the most affected genes during the first 24 hours of IAV post-infection. For MPV, IFIT1, IFIT2, MX1, OAS2, and IFI44 are the most affected genes during the first 24 hours of post-infection. Finally, for PIV3, IFIT1, MX1, OAS2, and OAS1 are the most affected genes during the first 24 hours after infection. These results also demonstrate that several genes such as CCL5, CXCL10, CXCL11, and TNFSF13B show significant changes during post-infection time points 24- and 72-hours for IAV. Similarly, IL33, CCL20, CXCL6, TNFSF18 and CXCL8 show significant changes for MPV during post-infection time points 24- and 72-hours. For PIV3 the genes IFIT3, IFIT2, IFIT1, HERC5, IFNL1, OAS1, OAS3 and ZBP1 significantly change within the post-infection timeframe between 24- and 72-hours. These findings advantageously contribute to a better understanding of the host-virus interaction and aid in the development of antiviral therapies.
[0172] Through differential expression analysis, the present techniques have been used to identify several unique genes that are significantly activated or changed by IAV and PIV3 infections during specific time points. These genes can serve as potential biomarkers or fingerprints for detecting and characterizing these infections. Specifically, IFNB1, OASL, IFNL2 / 3 and ISG15 are uniquely and significantly activated by IAV during the first 24 hours of infection, while TNF, PECAM1, and VWF are unique genes that change significantly between 24- and 72-hours after IAV infection. Additionally, IL31 is a unique gene that changes significantly during the first 24 hours of PIV3 post-infection, and IRF9 is a gene that significantly changes between PIV3 post-infection time points 24- and 72-hours. These findings provide advantageous and valuable insights into the molecular mechanisms of viral infections and may lead to the development of more effective diagnostic and therapeutic approaches.
[0173] The present techniques were used to identify five genes, namely OAS1, EIF2AK2, IFIT1, OAS2, and MX1 as common fingerprints for identifying respiratory viral infections caused by IAV, MPV and PIV3. Further, three genes (CCL20, IL6, and IL33) were identified as biomarkers for monitoring the progression of the disease between post-infection time points 24- and 72-hours. Advantageously, these findings may help develop diagnostic tools for quick and accurate identification of viral infections andAttorney Docket No.32103 / 59579 / PC track the severity of respiratory viral infections to provide timely and appropriate treatment. The present techniques may have significant implications for reducing the impact of viral infections on human health.
[0174] In some aspects, the present techniques may be applied to less active viruses with a lower multiplicity of infection, which may improve accuracy and generalizability. Furthermore, collecting data at more frequent post-infection time points may help reduce bias and improve the validity of the results. Moreover, the use of additional modalities may help address potential confounding factors and enhance robustness.
[0175] Differential expression analysis enables identification of top genes that are significantly affected by different viruses in terms of biological or statistical effect. The present techniques may also include exploratory analysis, which has a broader perspective or point-of-view of genes. In particular, the global gene expression landscape may be considered.
[0176] In the present techniques, differentially expressed genes from an organ tissue equivalent (OTE) of the human lung infected with three respiratory viruses may be processed using two novel algorithms. Levels of expression for genes (e.g., 19,671 genes obtained through RNA-Seq or genes obtained through NanoString) may be examined at distinct time points following infection with both intact and UV-inactivated virus.
[0177] The first algorithm, termed GLMQL-MAS, may include gauging the expression levels of targeted genes across specific viral infections. In particular, the parameters for a generalized linear model (GLM), employed for the identification of differentially expressed genes, may be estimated using robust quasi-likelihood (QL) methods. As discussed below, this GLMQL-MAS algorithm, which harmoniously combines Generalized Linear Models and Quasi-Likelihood techniques, may evaluate the extent of gene expression variation (magnitude) and its statistical significance (altitude). The synthesis of these two components empowers the MAS score to effectively rank genes according to their distinctive differential expression patterns under specific viral infection conditions.
[0178] The second algorithm, termed GLMQL-RMAS, relaxes the Benjamini-Hochberg adjustment (or other adjustment algorithm) for multiple comparisons. This algorithm improves the detection of potentially significant changes in the temporal pattern of gene expression.
[0179] These analyses were applied to the lung OTEs infected with influenza A virus (IAV), human metapneumovirus (MPV), and parainfluenza virus type 3 (PIV3). Key genes showing differential expression included many well-known interferon-responsive genes such as EIF5, IFNL1, IFIT1, IFIT2, CXCL10, OASL, IFIT3, IFNL2, IFNL3, CXCL11, RSAD2, IFNB1 and IFITM1 as well as some genes not associated with the interferon response such as GJB2 (gap junction beta 2), CA12 (carbonic anhydrase 12), and NEFL (neurofilament light polypeptide).Attorney Docket No.32103 / 59579 / PC
[0180] By 72 hours after infection, virus-specific gene expression for all viruses showed a substantial increase compared to non-infected samples or samples infected with UV-inactivated virus. OTEs infected with PIV3 showed the greatest change in differentially expressed genes compared to IAV- and MPV-infected OTEs. By contrast, IAV-infected OTEs showed the greatest change compared to PIV3- and MPV-infected OTEs. These results highlight the differential and potentially dynamic nature of virus- induced gene expression, with IAV showing early and sustained changes, PIV3 demonstrating increased changes over time, and MPV exhibiting the smallest changes among the three viruses analyzed here. This work and these methods of analysis may provide greater insight into the impact of viruses and host gene expression.
[0181] The present techniques include methods and systems for studying the intricate interplay between upper airway lung Organoid-Tissue Equivalents (OTEs) and three respiratory viruses: Influenza A virus (IAV), Human metapneumovirus (MPV), and Parainfluenza virus type 3 (PIV3). In some aspects, leveraging RNA-Seq data encompassing a comprehensive set of genes (e.g., 19,671 genes or more) may include analyzing RNA-Seq data using the Generalized Linear Models (GLM), Quasi-Likelihood (QL), and Magnitude-Altitude Score (MAS) algorithm. The present techniques may reveal genes with distinctive differential expression patterns across various experimental groups. To ensure the precise assessment of relative gene expression levels, the present techniques may employ the Trimmed Mean of M-values (TMM) normalization
[0182] The present techniques comprehensively investigate the interplay between OTEs and three respiratory viruses (IAV, MPV, and PIV3). The present techniques include applying the Generalized Linear Model (GLM) and Quasi-Likelihood (QL) F-test, to conduct a comprehensive analysis of gene expression dynamics across 16 infection conditions. By designating the untreated OTEs at the 24-hour time point (Naïve-24) as the control, this approach provides an overarching perspective of gene expression shifts across diverse experimental setups.
[0183] Further, the present techniques include performing pairwise comparison analysis, harnessing Generalized Linear Models with Quasi-Likelihood (GLMQL) to discern the impact of infection conditions on gene expression. Guided by the MAS algorithm, the present techniques include identifying differentially expressed genes between specific condition pairs, unraveling the molecular underpinnings of these disparities. The MAS algorithm, anchored in quantifying log-fold changes in gene expression levels and evaluating the statistical significance of these alterations using the Benjamini-Hochberg procedure, efficiently ranks genes based on their distinct differential expression patterns under specific viral infection conditions. This enables the present techniques to pinpoint genes that undergo significant changes in response to varying conditions, shedding light on the intricate mechanisms driving these differences. This enables us to pinpoint genes that undergo significant changes in response to varyingAttorney Docket No.32103 / 59579 / PC conditions, shedding light on the intricate mechanisms driving these differences. In general, this is referred to as “differential expression analysis” herein.
[0184] Lastly, the present techniques include a comprehensive exploration of gene expression patterns in viral infections, harnessing the power of General Linear Models (GLMs) in conjunction with Quasi-Likelihood (QL). This analytical approach not only enables dissection of the intricate interplay between IAV, MPV, and PIV3 infections and gene expression dynamics, but also uncovers the temporal intricacies of gene activity. Within this analytical framework, the present techniques introduce the Relaxed Magnitude-Altitude Score (RMAS) algorithm, which incorporates a relaxation of the Benjamini- Hochberg adjustment for multiple comparisons. By integrating this algorithm with GLMs and QL, the present techniques uncover significant trends, correlations, and variations underlying gene expression levels during viral infections. This analysis, generally referred to herein as “exploratory analysis of expression,” or merely “exploratory analysis,” advantageously enhances understanding of the associations between genes and viruses and their dynamic changes over time, shedding light on critical regulatory mechanisms that govern responses to viral challenges.
[0185] Thus, the present disclosure relates to a computer-implemented method for identifying differentially expressed genes under infection conditions using generalized linear models and quasi- likelihood methods. This method may include the collection of RNA-Seq data from organ tissue equivalents (OTEs) that have been infected with a virus. Specifically, the RNA-Seq data may be collected from OTEs infected with Influenza A virus (IAV), Human metapneumovirus (MPV), or Parainfluenza virus type 3 (PIV3). The analysis of this data may employ Generalized Linear Models (GLMs) with Quasi-Likelihood (QL) F-tests, which are particularly suited for modeling count data that follow distributions from the exponential family, such as Poisson or negative binomial distributions. This approach allows for the accurate and reliable inference of differentially expressed genes in RNA-Seq data, capturing the complex distributional characteristics of such data and managing gene-specific variability. Furthermore, the method may include the introduction of Magnitude-Altitude Score (MAS) and Relaxed Magnitude-Altitude Score (RMAS) algorithms. These algorithms are designed to navigate the RNA-Seq data effectively, integrating both the magnitude of expression change and its statistical significance. This integration facilitates the focused identification of genes significantly impacted by the infection, earmarking them as potential biomarkers. The MAS and RMAS algorithms prioritize genes for their potential as infection-specific biomarkers, offering a nuanced approach to gene prioritization by capturing a gene's overall impact on the study's biological context. Additionally, the method may include conducting Gene Ontology (GO) enrichment analysis to interpret the biological significance of the identified differentially expressed genes. This analysis facilitates the understanding of the biological processes, cellular components, and molecular functions enriched among the list of differentiallyAttorney Docket No.32103 / 59579 / PC expressed genes. This analysis also provides insights into the host's defense mechanisms and viral strategies exploiting host cellular functions. The GO enrichment analysis identifies key biological processes activated in response to viral infections, including the activation of interferon-stimulated genes and innate immune mechanisms, highlighting the host's reliance on innate immune mechanisms to combat viral threats. Moreover, the method may further include employing a multinomial logistic regression model to classify samples into distinct classes based on the identified differentially expressed genes. Empirically, this classification achieved a mean accuracy of 92% in distinguishing between different viral infections using a stratified k-fold cross-validation approach. The employment of multinomial logistic regression is useful in translating differential expression analysis into actionable insights, accommodating the complexities of biological data to ensure that identified biomarkers are not only statistically significant but also biologically relevant and capable of predicting specific viral infections accurately. Exemplary Computing Environment
[0186] FIG.1 depicts a block diagram of an exemplary computing environment 100 for implementing the present techniques, according to some aspects. The environment 100 may include a server computing device 102 and a client computing device communicatively coupled via an electronic network 106. Some aspects may include a plurality of server computing devices 102 and / or a plurality of client computing devices 104. In some aspects, some or all of the components of the environment 100 may be implemented using public / private / hybrid cloud computing resources (e.g., virtualized computing instances, virtualized databases, etc.).
[0187] The server computing device 102 includes a processor 110 and a network interface controller (NIC) 112. The server computing device may include an electronic database 108. The server 102 may include a library of client bindings for accessing the database 108. In some embodiments, the database 108 is located remote from the server 102. The database 108 may be a structured query language (SQL) database (e.g., a MySQL database, an Oracle database, etc.) or another type of database (e.g., a not only SQL (NoSQL) database). The database 108 may include data sets processed by the present techniques, as discussed herein.
[0188] The processor 110 may include any suitable number of processors and / or processor types, such as CPUs and one or more graphics processing units (GPUs). Generally, the processor 110 is configured to execute software instructions stored in a memory 114. The memory 114 may include one or more persistent memories (e.g., a hard drive / solid state memory) and stores one or more set of computer executable instructions / modules 150, including a statistical results module 152, a pair conditions module 154, an adjustment module 156, a magnitude-altitude module 158, a hyperparameterAttorney Docket No.32103 / 59579 / PC tuning module 160, a visualization module 162 and a diagnostics module 164. Each of the modules 160 implements specific functionality related to the present techniques. The modules 160 may execute concurrently / in parallel (e.g., on the CPUs 110) as respective computing processes, in some embodiments.
[0189] The client computing device 104 may be an individual server, a group (e.g., cluster) of multiple servers, or another suitable type of computing device or system (e.g., a collection of computing resources). For example, the client computing device 104 may be any suitable computing device (e.g., a server, a mobile computing device, a smart phone, a tablet, a laptop, a wearable device, etc.). Generally, a user (e.g., a researcher, a scientist, a diagnostician, an administrator, etc.) may access the data, processing software / modules and results generated by the server computing device 102 via the client computing device 102.
[0190] The network 106 may be a single communication network, or may include multiple communication networks of one or more types (e.g., one or more wired and / or wireless local area networks (LANs), and / or one or more wired and / or wireless wide area networks (WANs) such as the Internet). The network 106 may enable bidirectional communication between the server computing device 102 and the client computing device 104, and / or between multiple server computing devices 102, for example.
[0191] The NIC 112 may include any suitable network interface controller(s), such as wired / wireless controllers (e.g., Ethernet controllers), and facilitate bidirectional / multiplexed networking over the network 106 between the server computing device 102, the client computing device 104 and other components of the environment 100 (e.g., another client computing device 104, the electronic database 108, etc.).
[0192] The input device 120 may include one or more systems and / or devices for generating data for processing. For example, the input device 120 may include cell culture equipment (e.g., a CO2 incubator, a microscope, etc.), RNA extraction / purification equipment, quantification and quality assessment equipment, reverse transcription and / or PCR equipment, sequencing equipment (including next-generation sequencing equipment / platforms) such as high-throughput RNA equipment (e.g., RNA- seq) and / or data processing software. In some aspects, the input device 140 may include a suitable device or devices for receiving input, such as one or more microphones, one or more cameras, a hardware keyboard, a hardware mouse, a capacitive touch screen, etc.
[0193] The output device 130 may include any suitable device for conveying output, such as a hardware speaker, a computer monitor, a touch screen, etc. In some cases, the input device 120 and the output device 130 may be integrated into a single device, such as a touch screen device that acceptsAttorney Docket No.32103 / 59579 / PC user input and displays output. For example, the output device 130 may display one or more visualizations. In some aspects, the output device may be a virtual device.
[0194] Each of the modules 160 implements specific respective functionality related to the present techniques. For example, in some aspects, the statistical results module 152 includes instructions for receiving / retrieving statistical results from an experimental data sources (e.g., via the network 106 and / or the database 108). For example, the statistical results module 152 may receive / retrieve raw data summaries of gene expression data (e.g., measures such as mean, median, variance, range, etc.), NanoString data, RNASeq data, other data distributions (e.g., histograms), quality control results, normalized / transformed data, etc. The statistical results module 152 may store the received statistical results data in the database 108, for example.
[0195] The pair conditions module 154 may load the data retrieved / received by the statistical data module 152 and generate pair conditions for one or more pathogens, as described herein, for example with respect to FIG.2A. For example, the pair conditions module 154 may include instructions for generating combination tables using one or more viruses, one or more treatment conditions, one or more control groups, and / or one or more time periods, as discussed further herein. The pair conditions module 154 may make the generated combination data available for further processing by other modules in the modules 150 and / or store the combined information in the database 108, for example, or the memory 114.
[0196] The adjustment module 156 may include instructions for performing one or more data processing steps, such as computing of t-value, p-value and other statistical measures as discussed herein, for example, with respect to FIG.2B. The adjustment module 156 may include instructions for performing one or more adjustment procedures, such as a Benjamini-Hochberg algorithm, a Bonferroni correction algorithm, and / or other suitable techniques. The adjustment module 156 may correct p-values and other values, and store the adjusted values in a memory or database for further processing, and / or make such values available to other modules for further processing. The adjustment module 156 may include further instructions for sorting values prior to performing the adjustment procedures, in some aspects. The adjust module 156 may include instructions for computing a positive false discovery rate.
[0197] The magnitude-altitude module 158 may include instructions for computing scoring operations on the data adjusted by the adjustment module 156, as discussed in greater detail below, for example with respect to FIG.2C. For example, the magnitude-altitude module 158 may include instructions for computing magnitude-altitude scores based on the data generated by the pair conditions module 154 and the adjustment module 156. The magnitude-altitude module 158 may include instructions for sorting genes in terms of the computed scoring operations, and ranking genes according to min-max scores.Attorney Docket No.32103 / 59579 / PC The magnitude-altitude module 158 may store the scores, gene rankings, etc. and / or provide that information to other of the modules 150.
[0198] The hyperparameter tuning module 160 may enable hyperparameter values used for computations (e.g., by the magnitude-altitude module 158) to be adjusted. For example, the hyperparameter tuning module 160 may enable a user to view results of analysis performed using different hyperparameters, and to adjust hyperparameters (e.g., via the client computing device 104).
[0199] The visualization module 162 may include sets of computer-executable instructions for generating charts, graphs, plots, and other visual representations of the data processing of modules 152- 160, in some aspects. For example, visualizations that may be shown to users (e.g., via the output device 130 and / or the client computing device 104) may include those of FIGs.3A-3B, 4A-4B, 5, 6A-6C, 7A-7G and 8A-8G. It should be appreciated that as discussed below, visualizations may enable researchers and other users to quickly spot patterns in data, and as such, many other visualization techniques are envisioned, beyond those examples discussed herein such as heatmaps and volcano plots. For example, other visualization types that may be used include pathology images, flow cytometry, biomedical signal analysis, clinical data, etc.
[0200] The diagnostics module 164 may include instructions enabling diagnostic uses of the data processed by the modules 150. For example, the diagnostic module 164 may include a software application to facilitate review of the data generated by the modules 150. Examples of diagnostic usage are discussed in further detail below.
[0201] The interaction and operation of the modules 150 will now be described with respect to computer-implemented methods. Exemplary Computer-Implemented OTE Processing
[0202] The present techniques include methods and systems for comparing how upper airway lung OTEs respond to infection by respiratory viruses, including IAV, MPV and / or PIV3. Changes in cellular gene expression may be measured using the nCounter Host Response Panel of 785 cellular genes (or NanoString data). In some aspects, the present techniques may analyze RNASeq data, comprising many more cellular genes (e.g., 20,000 genes or more). The present techniques may include evaluating changes in gene expression by analyzing patterns of gene expression after infection with IAV, MPV and / or PIV3, over different time schedules after infection, to identify statistically significant changes most affected by each virus by comparing infected OTEs to control conditions (e.g., mock-infected or infected with UV-inactivated virus), as discussed below. For example, the NanoString data may be sourced from the input device 120 and / or the database 108 of FIG.1.Attorney Docket No.32103 / 59579 / PC
[0203] The present techniques may include identifying genes that are uniquely affected by each virus (IAV, MPV, and PIV3) as OTEs from an experimental group (e.g., infected with active virus) and a negative experimental group (e.g., infected with UV-inactivated virus) are contrasted. This processing may be performed (e.g., by the statistical results module 152 and pair conditions module 154) from two times after infection (e.g., 24 and 72 hours) to gain a better understanding of the temporal trajectory of gene expression in the infected OTEs, as discussed below.
[0204] The present techniques may include identifying genes affected by all three viruses (IAV, MPV, and PIV3) as OTEs move from the experimental condition (e.g., infected with active virus) to the negative control condition (e.g., infected with UV-inactivated virus). The processing may be repeated for data obtained 24 and 72 hours after infection to further evaluate changes in the pattern of gene expression over time. This processing may include identifying statistically significant genes whose expression varied in in all three virus infections, as discussed below.
[0205] In an aspect, a NanoString data set may include quantified transcript levels for a number (e.g., 773) of host response genes associated with viral infections. Target genes may be compared in a number (e.g., 16) of different infection conditions, where each infection condition involves a number (e.g., six) of OTE replicates. For example, 16 different infection conditions classified based on virus, treat, and post-infection time are shown below.
[0206] In another aspect, an RNA-Seq dataset may span many (e.g., 19,671 genes). Gene expression levels across 16 distinct infection conditions may be analyzed, each involving six OTE replicates. These conditions are categorized based on virus type, treatment, and post-infection time. The study may involve three viruses: IAV, MPV, and PIV3. The treatment category may include UV and None-UV (active), and two post-infection time points may be considered: 24 hours and 72 hours. "Treat" may refer to virus handling: UV treatment renders the virus non-infectious, while None-UV uses active viruses.
[0207] When a virus is exposed to inactivating doses of ultraviolet (UV) light, it is no longer able to productively infect cells. Thus, the level of UV exposure may be adjusted to ensure that the virus is inactivated without causing any other effects. The None-UV treatment refers to active viruses that can infect the organoids. The use of UV-treated infected OTEs as a negative control group enables the present techniques to compare changes that occur in None-UV infected OTEs with the changes that would be expected if there were no virus present at all.
[0208] To monitor temporal changes in UV or None-UV infected OTEs over time, additional groups of OTEs (e.g., naïve (i.e., untreated) and Mock) may be included as control groups. Naïve OTEs may be untreated samples, meaning that no virus or other treatment is applied to them, and samples are collected from these OTEs at specified time points. Mock-infected samples may be those that areAttorney Docket No.32103 / 59579 / PC treated in the same way as the virus-infected samples, except that no actual virus is introduced to the tissue. Mock samples serve as controls, ensuring gene expression changes in virus-infected samples result from the virus presence. Mock-infected OTEs compare gene expression between UV-treated and None-UV infected OTEs, while Naïve OTEs control for infection process effects. Herein, Condition- Treatment-Time" indicates OTE infection state. For example, IAV-None-24 implies active IAV infection at 24 hours. Mock-None-24 (or Mock-24) represents Mock-infected OTEs at 24 hours. Data distribution details, including time points, groups, and controls.
[0209] These may be referred to collectively as “mock-infected” samples. In this example, the Mock- infected OTEs may be used as a control group to compare the changes in gene expression between the UV-inactivated-treated and None-UV infected OTEs. Additionally, Naïve OTEs may be used as a control group for the Mock-infected OTEs to ensure that any changes in gene expression observed in the Mock- infected OTEs are not due to the process of the infection itself but rather are specifically related to the presence or absence of the virus. The pair conditions module 154 may include instructions for creating the control groups and experimental groups.
[0210] Table 1 depicts the distribution of data points between post-infection time points, experimental groups, negative control groups, and control groups. CONDITI VIRUS TREATME POST- CONDITI VIRUS TREATME POST- ON NT INFECTI ON NT INFECTI ON TIME ON TIME IAV-UV-24 IAV UV 24- IAV-UV-72 IAV UV 72- HOURS HOURS IAV- IAV NONE-UV 24- IAV- IAV NONE-UV 72- NONE-24 HOURS NONE-72 HOURS MPV-UV- MPV UV 24- MPV-UV- MPV UV 72- 24 HOURS 72 HOURS MPV- MPV NONE-UV 24- MPV- MPV NONE-UV 72- NONE-24 HOURS NONE-72 HOURS PIV3-UV- PIV3 UV 24- PIV3-UV- PIV3 UV 72- 24 HOURS 72 HOURS PIV3- PIV3 NONE-UV 24- PIV3- PIV3 NONE-UV 72- NONE-24 HOURS NONE-72 HOURS MOCK-24 MOCK - 24- MOCK-72 MOCK - 72- HOURS HOURS NAÏVE-24 UNTREAT - 24- NAÏVE-72 UNTREAT - 72- ED HOURS ED HOURS Table 1 - Sixteen infection conditions classified based on (1) Virus, (2) Treatment, and (3) post-infection time.
[0211] Table 2 depicts a distribution of replicates between post-infection timepoints, experimental groups, negative and regular control groups:Attorney Docket No.32103 / 59579 / PC AT POST- 6 REPLICATES FOR EACH OF THE FOLLOWINGS: INFECTION EXPERIMENTAL GROUP: NONE-UV TREATED IAV, NONE-UV TREATED MPV, TIME POINT NONE-UV TREATED PIV3 24-HOURS NEGATIVE CONTROL GROUP: UV TREATED IAV, UV TREATED MPV, UV TREATED PIV3 CONTROL GROUP: MOCK, NAÏVE AT POST- 6 REPLICATES FOR EACH OF THE FOLLOWINGS: INFECTION EXPERIMENTAL GROUP: NONE-UV TREATED IAV, NONE-UV TREATED MPV, TIME POINT NONE-UV TREATED PIV3 72-HOURS NEGATIVE CONTROL GROUP: UV TREATED IAV, UV TREATED MPV, UV TREATED PIV3 CONTROL GROUP: MOCK, NAÏVE Table 2 – Distribution of Replicates
[0212] The present techniques may include comparing a number of pair conditions for the pathogens under study, as shown for example in the following table: CONDITION A CONDITION B CONDITION A CONDITION B VIRUS-NONE-24 ↔ VIRUS-UV-24 VIRUS-UV-24 ↔ MOCK-24 VIRUS-NONE-72 ↔ VIRUS-UV-72 VIRUS-UV-24 ↔ NAÏVE-24 VIRUS-NONE-24 ↔ VIRUS-NONE-72 VIRUS-NONE-72 ↔ MOCK-72 VIRUS-UV-24 ↔ VIRUS-UV-72 VIRUS-NONE-72 ↔ NAÏVE-72 VIRUS-NONE-24 ↔ MOCK-24 VIRUS-UV-72 ↔ MOCK-72 VIRUS-NONE-24 ↔ NAÏVE-24 VIRUS-UV-72 ↔ NAÏVE-72 Table 3 – 12 different pair conditions for each virus (IAV, MPV and PIV3)
[0213] Table 3, above, depicts 12 different pair condition comparisons for viruses analyzed to determine the most statistically significant genes. To find the most affected genes as OTE infection conditions change for each virus, the present techniques may consider each of the 12 different pairs of comparisons. In particular, the present techniques may identify genes that significantly increase or decrease as six replicates of OTEs move from a condition A to a condition B, wherein this transition represents the infection condition for six replicates. Mechanically, this identification involves performing a two-sample independent t-test for genes, and computation of p-values with a given significance level (e.g., α=0.05), as illustrated in FIG.2. Exemplary Computer-Implemented Methods
[0214] FIG.2A depicts a block diagram of an exemplary computer-implemented method 200 for computing statistically significant genes for viruses as OTEs move between conditions, according to some aspects. In general, the method 200 may be performed by one or more modules of the computing modules 150 of FIG.1 (e.g., the statistical results module 152, the pair conditions module 154, the adjustment module 156, the magnitude-altitude module 158, the hyperparameters module 160, the visualization module 162, and / or the diagnostics module 164). The method 200 may include performing t-test and computing p-values for all genes as OTEs move from a condition A to a condition B, where r isAttorney Docket No.32103 / 59579 / PC the number of OTE replicates in each condition (e.g., r=6) and G is the number of genes (e.g., G=773). In some aspects, condition A may be referred to as an “initial condition” and condition B may be referred to as a “subsequent condition.”
[0215] A proportion of significant results in any statistical test includes false positives. To control the false discovery rate, the present techniques may include applying a Benjamini-Hochberg (BH) procedure that adjusts the significance level first and then computes the BH-adjusted p-values, as discussed in Benjamini, Yoav and Hochberg, Yosef, "Controlling the false discovery rate: a practical and powerful approach to multiple testing," Journal of the Royal statistical society: series B (Methodological), vol.57, no.1, pp.289-300, 1995; and in Benjamini, Yoav and Heller, Ruth and Yekutieli, Daniel, "Selective inference in complex research," Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, vol.367, pp.4255-4271, 2009; both of which are incorporated by reference in their entireties for all purposes.
[0216] FIG.2B depicts applying the BH procedure to the method 200 of FIG.2A, according to some aspects. The method 200 may further include computing BH-adjusted p-values for multiple p-values obtained from t-tests (block 220). The method 200 may further include sorting all p-values of mhypotheses in ascending order, ^(^), ^(^), ... , ^(^), to enable BH to control the false discovery rate byrejecting all hypotheses whose ordered p-are ^(^) ≤ ^^ ^, where ^, ^, and ^ are the rank of genebased on sorted p-values, the total number of hypotheses, and significance level (block 222). To measure how effective the BH procedure is, the present techniques may utilize a positive false discovery rate (pFDR), defined as the expected number of false positives out of all significant tests, as discussed in J. D. Storey, "A direct approach to false discovery rates," Journal of the Royal Statistical Society: Series B (Statistical Methodology), vol.64, no.3, pp.479-498, 2002., incorporated by reference herein in its entirety for all purposes. Exemplary Magnitude-Altitude Score (MAS) Aspects
[0217] In some aspects, the present techniques may be used to determine one or more most affected genes that are deemed significant by the BH adjustment procedure. Specifically, FIG.3 depicts a block diagram of an exemplary computer-implemented method 200 for computing magnitude-altitude score (MAS) for all BH-based significant genes, according to some aspects. The method 200 may advantageously enable identification of genes that are not only BH-based significant, but also genes whose expression increases or decreases considerably as infection conditions change.
[0218] For example, a number (e.g., six) of OTE replicates may transition from condition A to condition B, like discussed above. The above-described BH adjusted method 200 may be applied to find all BH-Attorney Docket No.32103 / 59579 / PC significant genes, also like discussed above (block 242). In that case, the method 200 may include computing Let ^^^^^^ ( )^^ ^^ ^ ^^(FC)^ = ^^^^(^^^^^^^ (^) ),^for ^ = 1, 2, … , ^, where ^ is the rejected based BH adjusted method,^^^^^^^(^)^and ^^^^^^^(^)^are average of gene ^ of replicates in condition A and B, respectively(block 244). that if ^^^^(FC)^ < 0, then it indicates that the average value of gene l decreases asOTEs move from condition A to B. Rejected null hypotheses are those for which adjusted p-values are considered significant. This means that the data provides enough evidence to conclude that the null hypothesis is unlikely to be true, and there is a significant effect.Thus, if log2FC^ > 0, then it indicates that the average value of gene ^ increases as OTEs move fromcondition A to B. If ^^^^^^^ (^)^is ! times as large as ^^^^^^^ (^)^, i.e., ^^^^^^^ (^)^= !(^^^^^^^ (^)^), then^^^^(FC)^ < −1, '( 0 < ! < 0.5,ì≤ ^^^^ ^ < 0, '( 0.5 ≤ ! < 1,'( ! = 1,1, '( 1 < ! ≤ 2,1, '( ! > 2.The method 200 may then identify genes having a maximum value of | ^^^^(FC)^| and |^^^^+(^^,-)|, simultaneously, by computing the following MAS algorithm: .^ / ^ = | ^^^^(FC) 0 ,- 1^| |^^^^+(^^ )| ,for ^ = 1, 2, … , ^, where ^ is thebased BH adjusted method,M and A are parameters (block 246).
[0219] The present techniques may interpret | ^^^^(FC)^| and |^^^^+(^^,-)| as magnitude and altitude control variables based on visualization of volcano plots, according to some aspects.
[0220] The method 200 may control the balance between magnitude and altitude variables by twohyper-parameters . and ^, where ., ^ ≥ 0. Without loss of generality, suppose | ^^^^(FC)056|>>| ^^^^(FC)^78|, where | ^^^^(FC)056| and | ^^^^(FC)^78| are respectively absolute values ofAttorney Docket No.32103 / 59579 / PC log2FC corresponding to two genes with maximum rate of increase and decrease as OTEs move from condition A to B. In that case, hyper-parameters . and ^ must be selected in a way that a gene with R=1 (MAS-based rank) is the one with minimum distance from the point(| ^^^ (FC) |, max {|^^^ (^,-^ 056 ^ ^+ ^ )|}). On the other hand, if | ^^^^(FC)056| <<| ^^^^(FC)^78|, then hyper-parameters . and ^ must be selected in a way that a gene with R=1 (MAS-based rank) is the one withminimum distance from the point (| ^^^^(>?)^78|, max^{|^^^^+(^,-^ )|}). It should be appreciated thatMAS-based rank is a symmetric metric that rank to whether OTEs move from condition A to B or vice versa.
[0221] The method 200 may include sorting the genes in terms of .^ / ^in a descending order and ranking the gene(s) with the largest MAS as rank 1 (R=1), and the gene(s) with the smallest MAS as rank s (R=s) (block 248). The method 200 may include returning the genes’ respective ranks (block 250). Determining Unique Statistically Significant Genes Affected by Respective Viruses During OTE Condition Transitions
[0222] The unique statistically significant genes that are largely affected by each respective virus as OTEs move between conditions is of interest. The method 200 may be extended to compute this information. For example, in some aspects, the method 200 may include finding all BH-based significant genes that are only affected by a particular virus. Continuing the example of having three viruses IAV, MPV and PIV3, this means that n=3. Thus, the method 200 may include computing a unique gene selection for a particular virus by computing the following: Input: Given infection conditions A and B.Step 1: for ' = 1,2, … , @ do:• By means of method 200 (as depicted in FIGs.2A-2C), as OTEs infected by virus i transmit between conditions A and B, determine BH-based significant genes and their MAS-based ranks and store them into a set, namely BH-MAS-virus-i. End (for)Step 2: for A = 1,2, … , @ do:• Construct a set Unique-BH-MAS-virus-j whose elements are differentially expressed genes of BH-MAS-virus-j obtained in Step 1 that do not belong to BH-MAS-virus-k for ^ = 1,2, … , @ and^ ≠ A .End (for)Output: Unique-BH-MAS-virus-j for A = 1, 2, ... , @.
[0223] Further, the present techniques may include visualizing the unique statistically significant genes that are affected by the respective viruses, as discussed below with respect to FIGs.7A-7G. Determining Statistically Significant Genes Affected by All Viruses During OTE Condition TransitionsAttorney Docket No.32103 / 59579 / PC
[0224] Furthermore, the intersection set of BH-based significant genes among all viruses as the OTE replicates move between conditions is also of interest. The method 200 may be extended to compute this information, alternatively or in addition to finding unique statistically significant genes: Input: Given infection conditions A and B.Step 1: for ' = 1,2, … , @ do:• By means of method 200 (as depicted in FIGs.2A-2C), as OTEs infected by virus i transmit between conditions A and B, determine BH-based significant genes and their MAS-based ranks and store them into a set namely BH-MAS-virus-i. End (for)Step 2: for A = 1,2, … , @ do:• Construct a set Common-BH-MAS whose elements are differentially expressed genes of BH- MAS-virus-j obtained in Step 1 that do belong to all BH-MAS-virus-k for ^ = 1,2, … , @.End (for) Output: Common-BH-MAS
[0225] Further, the present techniques may include visualizing the common statistically significant genes during OTE condition transitions, as discussed below with respect to FIGs.8A-8B. Exemplary Computer-Implemented Visualization Aspects
[0226] FIG.3A depicts an exemplary three-dimensional visualization 300 of OTEs over time, according to some aspects. The visualization 300 visualizes all OTEs by means of PCA, for example, as discussed in Abdi, Hervé, and Lynne J. Williams, "Principal component analysis," Wiley interdisciplinary reviews: computational statistics, vol.2, no.4, pp.433-459, 2010, herein incorporated by reference in its entirety for all purposes. Since the NanoString data contains different genes of different scales, for the sake of accurate visualization, the present techniques may include standardizing the data by converting the scales into a common scale (through a min-max normalization strategy, for example, as discussed in Patro, S and Sahu, Kishore Kumar, "Normalization: A preprocessing stage," arXiv preprint arXiv:1503.06462, 2015, herein incorporated by reference in its entirety for all purposes). In particular, the visualization 300 depicts the projection of all OTEs (e.g., 96 in all infection conditions) onto a three- dimensional subspace that is generated by the first three principal components.
[0227] The present visualization techniques are advantageous as they enable patterns to be immediately identified as infected OTEs move between conditions (e.g., UV versus None-UV) and over time (e.g., 24-hour versus 72-hour conditions).
[0228] For example, FIG.3B depicts an example of identified patterns based on the PCA analysis of FIG.3A, according to some aspects. FIG.3B enables a viewer to single out the identified pattern for IAV-infected OTEs as they move between conditions and over time (e.g., UV versus None-UV and 24- hours versus 72-hours), and identified patterns for MPV and PIV3 infected OTEs. In some aspects, anAttorney Docket No.32103 / 59579 / PC average of OTE replicates is used for visualization. For example, the present techniques may take an average of projected replicates and pin them on the same three-dimensional coordinate system in which replicates are projected, as shown in FIG.3C. By viewing FIG.3B, the viewer can immediately see the relationship of UV and None-UV treated samples, as well as the pattern of progression of those samples over different time scales. It should be appreciated that the present techniques use UV / None-UV conditions for exemplary purposes, and that many other conditions may be added, in some aspects. Further, it should be appreciated that other time scales may be added, which may be shorter or longer. In some aspects, the relevant time scales may be much shorter (e.g., one second or less) or much longer (e.g., one week or longer).
[0229] The present techniques identified three genes: IFIT1 (infection-dependent), IFIT2 (infection- dependent), and ELOVL4 (time-dependent), using GLMQL-RMAS as predictors for multinomial logistic regression. After conducting a 6-fold stratified cross-validation, the present techniques obtained a mean accuracy of 92% with a 95% confidence interval of [85, 98%]. FIG.3D displays a 3D visualization of all samples using only these three genes as the coordinate axes. FIG.3E shows the confusion matrix for the 8-class classification using these genes. FIG.3F presents the confusion matrix for the 8-class classification when replacing RMAS with the common ranking method, ranking genes based on the smallest p-value, and FIG.3G illustrates the classification when genes are ranked based on logFC, with all other aspects remaining identical.
[0230] FIG.4A depicts an exemplary line chart 402 illustrating a BH-cutoff line and a number of BH- based significant genes of the first two pair conditions for IAV, across a 24-hour period. FIG.4B depicts an exemplary line chart 404 illustrating a BH-cutoff line and a number of BH-based significant genes of the first two pair conditions for IAV, across a 72-hour period. The chart 402 and the chart 404 reflect the status of the BH and p-value products after execution of the block 222 of the method 200 in FIG.2B.
[0231] FIG.5 depicts an exemplary bar chart 500 illustrating positive false discovery rates of the BH adjustment procedure for a period of time, according to some aspects. Exemplary Computer-Implemented Hyperparameter Tuning
[0232] The present techniques may include tuning one or more hyperparameter values (e.g., the values of (Magnitude, Altitude)). For example, the hyperparameters may be chosen as follows: (0,1), (1,0), and (1,1). Assuming OTEs move from infection condition IAV-None-24 to IAV-UV-24, by means of method 200, blocks 242-250, the present techniques may compute MAS-based rankings of genes according to the changed (M, A) values. FIGs.6A-6C depict the result plotting these rankings using volcano plots. Specifically, FIG.6A depicts an exemplary volcano plot of hyperparameter tuning when (M,A) is set to (0,1), according to some aspects. FIG.6B depicts an exemplary volcano plot ofAttorney Docket No.32103 / 59579 / PC hyperparameter tuning when (M,A) is set to (1,0), according to some aspects. FIG.6C depicts an exemplary volcano plot of hyperparameter tuning when (M,A) is set to (1,1), according to some aspects.
[0233] FIGs.6A-6C display the respective volcano plots and MAS-based ranking of genes corresponding to cases when (M, A) is equal to (0,1), (1,0), and (1,1), respectively. These plots indicate that the balance between the magnitude and altitude variables occurs if M= A. And further, that if (M, A) = (1,0) and (M, A) = (0,1), then CXCL10 and OAS2 are selected as MAS-based rank #1 (R=1), respectively, but neither of them satisfies the optimal balance between magnitude and altitude. FIGs. 6A-6C further indicate that when the values of (M, A) are set to (1, 0), the method’s ranking weights shift towards larger magnitudes, suggesting a bias towards higher values. Conversely, when (M, A) is set to (0, 1), the ranking weights shift towards larger altitudes, indicating a bias towards higher altitude values. However, if (M, A) are set to (1, 1), there is a balance between the magnitude and altitude variables. Inthis case, the gene with the lowest distance from the point (| ^^^ ,-^(FC)056|, max^{C(^^^^+(^^ D}) and anappropriate balance between the magnitude and altitude #1(R=1), which is IFNL1.
[0234] Additional volcano plots may be generated to depict MAS-based ranking of genes corresponding as OTEs move between different infection conditions when M=A=1. For example, such plots may display how top-20 MAS-based selected genes change within OTE replicates, illustrate the averages of top-20 genes for identical infection condition change, demonstrate 3-D visualizations of OTE replicates using the top-3 MAS-based selected genes, and provide the values of | ^^^^(FC)^| and |^^^^+(^^,-)| for the top-20 selected genes after all MAS scores are normalized (through the min-max normalization strategy
[0020] ) as OTEs move between 36 infection conditions i.e.12 different pair infection conditions comparisons for each virus, as shown in Table 3, supra.
[0235] Table 4, Table 5 and Table 6 (infra) may display the top-10 MAS-based selected genes as infected OTEs by IAV, MPV, and PIV3, respectively, move between 12 infection pair conditions. Moreover, to provide a better understanding of selected genes, Table 7 (infra) may display the MAS- based selected genes as OTEs move between control / negative-control group conditions. As shown, MAS is an appropriate global (within all pair infection conditions) metric to compare the intensity of infection conditions. Genes with larger MAS scores are more affected as OTEs move between two infection conditions. CONDITION CONDITION THE TOP-10 MAS-BASED SELECTED GENES A B IAV-NONE-24 ↔ IAV-UV-24 IFNL1, IFIT2, IFNB1, RSAD2, IFIT3, OASL, IFNL2 / 3,Attorney Docket No.32103 / 59579 / PC IFIT1, ISG15, HERC5 IAV-NONE-72 ↔ IAV-UV-72 IFNL1, CCL5, IFNL2 / 3, CXCL10, CXCL11, IFIT2, TNFSF13B, IFIT1, RSAD2, OASL IAV-NONE-24 ↔ IAV-NONE- CCL20, IFNB1, PECAM1, VWF, IL1B, IL33, CXCL1, 72 CCL28, CXCL17, BNIP3 IAV-UV-24 ↔ IAV-UV-72 CCL20, IL33, RSAD2, IL6, SOD2, MX1, CXCL17, CXCL8, VEGFA, IL32 IAV-NONE-24 ↔ MOCK-24 IFIT2, RSAD2, IFIT3, IFNL1, IFIT1, IFNB1, ISG15, OAS2, OASL, MX1 IAV-NONE-24 ↔ NAÏVE-24 RSAD2, IFNL1, IFIT2, IFNB1, IFIT3, OASL, IFNL2 / 3, IFIT1, HERC5, MX1 IAV-UV-24 ↔ MOCK-24 IL27, CCL1, IL23R, MS4A2, IL12B, LILRB2, SH2D1A, PTPRC, IL9, CD79A IAV-UV-24 ↔ NAÏVE-24 SLC2A3, BNIP3, VEGFA, PRDM1, TGFB1, PLAUR, MS4A2, CD68, IKBKE, IL1RN IAV-NONE-72 ↔ MOCK-72 IFIT1, CXCL10, RSAD2, IFITM1, IFIT3, IFI44, OAS2, IFIT2, OAS1, CCL5 IAV-NONE-72 ↔ NAÏVE-72 IFIT1, RSAD2, IFIT3, OAS1, IFIT2, IFI44, IFITM1, CCL5, OASL, OAS3 IAV-UV-72 ↔ MOCK-72 OAS2, RSAD2, IFI44, IFIT1, MX1, IFI6, IFITM1, IL12B, OAS1, IFIT3 IAV-UV-72 ↔ NAÏVE-72 RSAD2, OAS1, MX1, IFI44, IFIT1, OAS2, IFIT3, CXCL3, IL36RN, CXCL6 Table 4 – Top-10 MAS-based selected genes for IAV-infected OTEs as they transit between 12 pair conditions CONDITION CONDITION THE TOP-10 MAS-BASED SELECTED GENES A B MPV-NONE- ↔ MPV-UV-24 IFIT1, IFIT2, IFIT3, RSAD2, MX1, OAS1, OAS2, IFI44, 24 IFI35, STAT1 MPV-NONE- ↔ MPV-UV-72 IRF7, CSF1 72 MPV-NONE- ↔ MPV-NONE- IL33, CCL20, CXCL5, CXCL6, IL36G, TNFSF18, CXCL8, 24 72 IL1B, SOD2, IL6 MPV-UV-24 ↔ MPV-UV-72 IL33, CCL20, TNFSF18, CXCL17, CXCL1, IFI27, IL6, CXCL6, LRG1, CXCL8 MPV-NONE- ↔ MOCK-24 IFIT1, IFIT2, CXCL10, IFIT3, RSAD2, TXK, MX1, 24 CD79A, OAS2, IFI44 MPV-NONE- ↔ NAÏVE-24 SLC2A3, BNIP3, PLAUR, VEGFA, MYC, CD68, PARP1, 24 MAP3K8, PFKFB3, IKBKE MPV-UV-24 ↔ MOCK-24 OSM, FCAR, IL27, IFNL4, NOX1, NOD2, IL24, DEFA4, CSF2RA, IL37 MPV-UV-24 ↔ NAÏVE-24 CXCL10, VEGFA, SLC2A3, MYC, BNIP3, DDIT3, PARP1, CXCL11, MAP3K8, CD68Attorney Docket No.32103 / 59579 / PC MPV-NONE- ↔ MOCK-72 IL17A, MAP3K7, MAP3K5, GBA, PAK1, CTSL, PELI1, 72 VAMP3, EVL, CARD11 MPV-NONE- ↔ NAÏVE-72 CXCL5, CXCL3, CXCL6, IL17A, EVL, IL1B, LIMK2, 72 NFKB1, CXCL1, LDHB MPV-UV-72 ↔ MOCK-72 IL17A, CXCL12, GBA, MAP3K7, TOLLIP, MAP3K5, RAB31, EVL, IL9, CARD11 MPV-UV-72 ↔ NAÏVE-72 IL36RN, CXCL3, CXCL10, CXCL5, CXCL6, PRDM1, TGFB1, IL1RN, CXCL1, SLC2A3 Table 5 – Top-10 MAS-based selected genes for MPV-infected OTEs as they transit between 12 pair conditions CONDITION CONDITION THE TOP-10 MAS-BASED SELECTED GENES A B PIV3-NONE- ↔ PIV3-UV-24 IFIT1, MX1, OAS2, OAS1, DTX3L, IFITM1, LAG3, 24 TRIM21, UBE2L6, EIF2AK2 PIV3-NONE- ↔ PIV3-UV-72 IFIT3, IFIT1, RSAD2, MX1, IFIT2, HERC5, OAS3, 72 DDX58, IFI44, OAS1 PIV3-NONE- ↔ PIV3-NONE- RSAD2, IFIT3, IFIT2, IL33, IFIT1, HERC5, IFNL1, OAS1, 24 72 OAS3, ZBP1 PIV3-UV-24 ↔ PIV3-UV-72 IL33, CXCL17, CYSTM1, LTF, TNFSF18, PFKFB3, CXCL6, RPS6KA1, IL6, BST2 PIV3-NONE- ↔ MOCK-24 IFIT1, CXCL10, IFIT3, IFIT2, MX1, OSM, RSAD2, 24 HERC5, OAS2, OAS1 PIV3-NONE- ↔ NAÏVE-24 SLC2A3, BNIP3, PLAUR, VEGFA, IL1RN, CD68, PRDM1, 24 TGFB1, PFKFB3, ULK1 PIV3-UV-24 ↔ MOCK-24 LILRB2, MS4A2, IL23R, IL27, IL2RB, IFNL4, IFNA5, SH2D1A, CCL19, KLRC1 PIV3-UV-24 ↔ NAÏVE-24 CXCL10, SLC2A3, PLAUR, VEGFA, PFKFB3, PRDM1, ULK1, BNIP3, CD68, TGFB1 PIV3-NONE- ↔ MOCK-72 IFIT3, RSAD2, IFIT1, IFIT2, MX1, IFI44, OAS3, OAS2, 72 HERC5, DDX58 PIV3-NONE- ↔ NAÏVE-72 IFIT3, RSAD2, IFIT1, MX1, IFIT2, OAS3, IFI44, HERC5, 72 OAS1, OAS2 PIV3-UV-72 ↔ MOCK-72 IL12B, CD244, KLRC1, DEFA4, IL9, CXCL12, SH2D1A, TRAT1, CD3G, FAM30A PIV3-UV-72 ↔ NAÏVE-72 CXCL5, GBP1, CXCL6, KLRC1, CXCL10, CXCL13, TRIM22, LTF, CXCL3, IFI35 Table 6 – Top-10 MAS-based selected genes for PIV3-infected OTEs as they transit between 12 pair conditions CONDITION CONDITION THE TOP-10 MAS-BASED SELECTED GENES A B MOCK-24 ↔ NAÏVE-24 BNIP3, CXCL10, PLAUR, VEGFA, IL1RN, CD68, PARP1, MAP3K8, SLC2A3, MYC MOCK-72 ↔ NAÏVE-72 IL36RN, TGFB1, SLC2A3, MME, PRDM1, CXCL6, IL36B, CXCL1, CXCL10, BNIP3 MOCK-24 ↔ MOCK-72 IL36G, IL33, CXCL17, IL6, VEGFA, CXCL1, JUN, BNIP3, CXCL6, SERPINA1 NAÏVE-24 ↔ NAÏVE-72 CXCL10, IFNA6, S100A12, OAS2, IFIT1, IFIT3, ISG15, IFI16, PTGS2, IKBKEAttorney Docket No.32103 / 59579 / PC Table 7 – Top-10 MAS-based selected genes for OTEs as they transit between control group conditions Visualizing Unique Statistically Significant Genes Affected by Each Virus During OTE Condition Transitions
[0236] As discussed above, the present techniques may include quantifying unique statistically significant genes that are largely affected by each virus as OTEs move between conditions. After computing the unique gene selection for the particular virus, the present techniques may generate one or more visualizations depicting p-values and other computed values. For example, bar charts may be plotted after computing p-values corresponding to pair infection conditions to find the unique genes corresponding to each virus.
[0237] FIG.7A depicts an exemplary bar chart illustrating IAV affected genes as OTEs in negative and experimental control groups are contrasted, according to some aspects.
[0238] FIG.7B depicts an exemplary bar chart illustrating IAV affected genes as OTEs in negative and experimental control groups are contrasted, according to some aspects.
[0239] FIG.7C depicts an exemplary bar chart illustrating IAV affected genes as OTEs in negative and experimental control groups are contrasted, according to some aspects.
[0240] FIG.7D depicts an exemplary bar chart illustrating PIV3 affected genes as OTEs in negative and experimental control groups are contrasted over time, according to some aspects.
[0241] FIG.7E depicts an exemplary bar chart illustrating PIV3 affected genes as OTEs in negative and experimental control groups are contrasted over time, according to some aspects.
[0242] FIG.7F depicts an exemplary bar chart illustrating PIV3 affected genes as OTEs in negative and experimental control groups are contrasted over time, according to some aspects.
[0243] FIG.7G depicts an exemplary bar chart illustrating MPV affected genes as OTEs in negative and experimental control groups are contrasted over time, according to some aspects. Visualizing Common Statistically Significant Genes Largely Affected by All Viruses During OTE Condition Transitions
[0244] As discussed above, the present techniques may include quantifying common statistically significant genes during OTE condition transitions. In particular, the present techniques may include generating one or more visualizations that depict common statistically significant genes that are largely affected to all viruses as OTEs move between two conditions. The visualizations may be generated after constructing the Common-BH-MAS set, as described above. There may be five BH-significant genesAttorney Docket No.32103 / 59579 / PC that are common among all viruses at post-infection timepoint 24-hours, in some aspects (e.g., OAS1, EIF2AK2, IFIT1, OAS2 and MX1).
[0245] For example, FIG.8A depicts a log2-transformed heatmap of averages of these five genes, according to some aspects. FIG.8B depicts a log2-transformed heatmap of five BH-significant genes at time 72-hours, according to some aspects. Advantageously, by visually comparing FIG.8A and FIG.8B, the viewer is able to quickly and intuitively compare how a gene (e.g., IFIT1) is affected as OTEs move between UV and None-UV infection conditions.
[0246] FIG.8C depicts an exemplary heatmap plot showing separation of replicates of particular genes transit UV-treated conditions over time, according to some aspects. FIG.8D depicts an exemplary heatmap plot showing separation of replicates of particular genes transit UV-treated conditions over time, according to some aspects. In particular, FIG.8C and 8D depict that IFIT1 can separate all OTE replicates in UV-treated conditions from None-UV infected conditions at time 24-hours. FIG.8E depicts an exemplary heatmap plot showing separation of replicates of particular genes transit UV-treated conditions over time, according to some aspects. FIG.8F depicts an exemplary heatmap plot showing separation of replicates of particular genes transit UV-treated conditions over time, according to some aspects. FIG.8E and 8F depict that IFIT1 can separate all replicates in UV-treated conditions from None-UV infected conditions at time 72-hours.
[0247] FIG.8G depicts an exemplary heatmap plot showing separation of replicates of particular genes transit UV-treated conditions over time, according to some aspects. In particular, the method 200 may be applied as OTEs move from 24 hours to 72 hours for all conditions and all viruses to determine that three common genes separate OTEs during the transition from 24 hour from 72 hour conditions. In particular, FIG.8G depicts that genes CCL20, IL6 and IL33 separate the OTEs in 24 hour to 72 hour conditions. Advantages and Improvements
[0248] The present MAS scoring and visualization capabilities represent significant improvements over conventional techniques. Specifically, the ability to rapidly compare how organ tissue equivalent genes are affected as they move between disease conditions has profound implications for disease understanding, diagnosis, treatment, and research. This empowers scientists, clinicians, and healthcare providers with actionable insights that may lead to improved patient care, enhanced treatment strategies, and accelerated advancements in the field of biomedicine, biomedical research and healthcare.
[0249] Further specific improvements flowing from the present techniques include early disease detection and diagnosis, wherein rapid comparison of gene expression patterns across different disease conditions can aid in the early detection and diagnosis of diseases. Identifying unique gene signatures / Attorney Docket No.32103 / 59579 / PC fingerprints associated with specific diseases allows for more accurate and timely diagnosis, leading to early intervention and improved patient outcomes.
[0250] The present techniques also improve precision medicine capabilities, by helping clinicians and researchers to understand how OTE genes respond differently to various disease conditions. This may enable the development of personalized treatment approaches. By tailoring treatments based on an individual's genetic makeup, healthcare providers may increase treatment efficacy and reduce potential side effects.
[0251] The present techniques may improve drug discovery and development methods and systems. For example, rapid gene expression comparisons between disease conditions may contribute to drug discovery and development by helping researchers to identify potential drug targets and pathways that are specific to certain diseases, facilitating the design of targeted therapies.
[0252] Still further, comparing gene expression profiles may improve biomarker identification (specific genes or gene patterns that are indicative of certain diseases), an essential task for disease prognosis, monitoring of treatment response, and evaluation of disease progression.
[0253] Furthermore, many additional advantages are envisioned. For example, examining how OTE genes behave across different disease conditions may provide advantageous insights into the underlying mechanisms of diseases that is fundamental for unraveling disease pathogenesis and developing effective interventions. The ability to quickly compare gene expressions using the present techniques may also streamline research processes. For example, researchers may rapidly explore large datasets and identify significant gene changes associated with different disease states, accelerating the pace of scientific discovery. The ability to swiftly compare gene expression in different disease conditions may generate data-driven insights that can guide research hypotheses, inform experimental designs, and direct further investigations. Monitoring changes in gene expression during the course of treatment may advantageously help assess treatment effectiveness. It enables clinicians to make informed decisions about treatment adjustments based on real-time molecular responses. Rapid gene expression comparisons may contribute to understanding the spread and impact of diseases at a population level. Identifying gene expression patterns associated with disease outbreaks can aid in epidemiological studies and public health interventions. Bridging the gap between basic research and clinical applications is facilitated by quick gene expression comparisons, such that insights gained from laboratory studies can be more efficiently translated into clinical practice. Healthcare providers can use gene expression information to develop individualized treatment plans, monitor disease progression, and optimize patient management strategies.Attorney Docket No.32103 / 59579 / PC
[0254] A concrete example of developing learnings to further diagnostic techniques follows. As discussed, the present techniques may be used to analyze gene expression patterns respiratory viruses (e.g., IAV, MPV, and PIV3) in infected lung-tissue-equivalents using a statistical technique. PCA visualization and scoring were used to quantify the magnitude and significance of gene expression changes. In empirical testing, IAV had the most significant impact on gene expression compared to the other two viruses. PIV3 had significantly affected genes between 24-72 hours post-infection, while MPV was the least intense virus during the first 72 hours of infection. The findings provide insights into the mechanisms of viral infections and could aid in the development of effective treatments for respiratory diseases caused by these viruses.
[0255] In this example, several genes that are significantly affected by infections of IAV, MPV and PIV3 may be identified during different post-infection time frames. Analyzing these results may show that IFNL1 and CCL20 are the most affected genes during the first 24 hours of IAV post-infection. For MPV, IFIT1, IFIT2, MX1, OAS2, and IFI44 are the most affected genes during the first 24 hours of post- infection. Finally, for PIV3, IFIT1, MX1, OAS2, and OAS1 are the most affected genes during the first 24 hours after infection.
[0256] Using the present techniques, empirical testing also found that that several genes such as CCL5, CXCL10, CXCL11, and TNFSF13B show significant changes during post-infection time points 24- and 72-hours for IAV. Similarly, IL33, CCL20, CXCL6, TNFSF18, and CXCL8 show significant changes for MPV during post-infection time points 24- and 72-hours. For PIV3, IFIT3, IFIT2, IFIT1, HERC5, IFNL1, OAS1, OAS3, and ZBP1 are genes that significantly change within the post-infection timeframe between 24- and 72-hours. These findings may advantageously contribute to a better understanding of the host-virus interaction and aid in the development of antiviral therapies.
[0257] Further, processing that identified several unique genes that are significantly activated or changed by IAV and PIV3 infections during specific time points may serve as potential biomarkers or fingerprints for detecting and characterizing these infections. Specifically, IFNB1, OASL, IFNL2 / 3, and ISG15 are uniquely and significantly activated by IAV during the first 24 hours of infection, while TNF, PECAM1, and VWF are unique genes that change significantly between 24- and 72-hours after IAV infection. Additionally, IL31 is a unique gene that changes significantly during the first 24 hours of PIV3 post-infection, and IRF9 is a gene that significantly changes between PIV3 post-infection time points 24- and 72-hours. These findings provide valuable insights into the molecular mechanisms of viral infections and can potentially lead to the development of more effective diagnostic and therapeutic approaches.
[0258] Using the present techniques, five genes were identified, namely OAS1, EIF2AK2, IFIT1, OAS2, and MX1, as common fingerprints for identifying respiratory viral infections caused by IAV, MPV,Attorney Docket No.32103 / 59579 / PC and PIV3. Additionally, three genes, CCL20, IL6, and IL33, were identified as biomarkers for monitoring the progression of the disease between post-infection time points 24- and 72-hours. These findings advantageously may help develop diagnostic tools for quick and accurate identification of viral infections and track the severity of respiratory viral infections to provide timely and appropriate treatment. These specific findings may have significant implications for reducing the impact of viral infections on human health. Exemplary Computer-Implemented Exploratory Analysis
[0259] As discussed above, the present techniques may include exploratory analysis, which adopts a wider view of which genes are in scope. Instead of looking for genes with highest expression for a given infection, the present techniques may analyze all genes regardless of their biological significance. Differential analysis may be used to determine whether given OTEs are affected by viruses, and it is indeed very important to identify genes that are affected. At the same time, looking at the OTE as a whole and seeing the genes as a whole is important too, to determine if interplay between less significant genes.
[0260] It should be appreciated that exploratory analysis may still consider only genes that are statistically significant, but with a more expansive view. In other words, differential expression analysis may consider significant genes with a tightly restricted view, using techniques like Benjamini-Hochberg to select genes with highest expression significance. Thus, statistical significance is still important, but on a slightly more relaxed basis in exploratory analysis. This advantageously enables a good view of the infection process. Table 8 depicts a comparison between the two techniques.Attorney Docket No.32103 / 59579 / PC Table 8 – Comparing Exploratory Analysis in Gene Expression and Differential Expression Analysis
[0261] In general, RNASeq data is less precise than NanoString. Thus, it is necessary to relax the criteria of statistical significance when working with RNASeq to enable the methodology. Exemplary Computer-Implemented Exploratory Analysis Visualizations
[0262] FIG.9A depicts an exemplary computer-implemented exploratory analysis visualization plot 900 of RMAS scores over time, using the relaxed exploratory analysis approach. This figure enables patterns of expression over different OTEs to be analyzed, for different viral infections. The X axis of plot 900 depicts OTEs infected with respective viruses (IAV, MPV and PIV3) with a baseline of mock infected, at time 24. The plot shows that gene expression landscape for IAV over the first 24 hours is very active, while for MPV and PIV3, it is less so.
[0263] Next, the X axis of plot 900 depicts OTEs infected with respective viruses (IAV, MPV and PIV3) with a baseline of mock infected, at time 72. The plot shows that gene expression landscape for IAV and PIV3 at time 72 hours is very active, while for MPV, it is less so. Comparing the two time periods indicates similar intensity of expression of IAV, but a significant jump for PIV3 from time 24 to 72.
[0264] Specifically, plot 900 illustrates gene GLMQL-RMAS scores for different pairs of conditions. The scatter plots depict the expression patterns of genes under varying infection conditions, including (Mock-24, Virus-None-24) and (Mock-72, Virus-None-72) for viruses IAV, MPV, and PIV3. The x-axisAttorney Docket No.32103 / 59579 / PC represents the different pairs of conditions, while the y-axis represents the gene GLMQL-Relaxed-MAS scores. The scatter plots provide a visual representation of gene expression patterns across different conditions and time points, facilitating the identification of trends, variations, and potential fingerprints.
[0265] The ability to plot the exploratory analysis results advantageously allows the viewer (e.g., a researcher, clinician, academic, end user, etc.) to quickly determine the behavior and intensity of a virus over different time periods, and to reach conclusions regarding viruses. For example, the viewer of FIG. 9A would conclude that IAV is a quickly expressing virus, with strong and durable expression, while PIV3 expression happens later. Another evident finding is that MPV does not express much at all within 72 hours.
[0266] As noted, differential analysis adopts, by design, a narrower view of what is considered statistically significant expression values for given genes. The expression landscape of MPV in FIG.9A indicates a primary advantage of exploratory analysis vis-a-vis differential expression analysis; namely, that MPV is a large virus with many genes, and when using Benjamini-Hochberg in conjunction with differential analysis, the expression change from 24 to 72 hours is not captured because it is out of scope. In other words, the cutoff for expression is too high in differential analysis to capture the slight increase in expression between the two time periods. However, exploratory analysis allows these more subtle changes to be captured and presented to the viewer.
[0267] Even more interesting, FIG.9A enables the viewer to compare expression among different viruses at different time steps for purposes of forming conclusions. For example, the expression of MPV at time 72 is similar to that of PIV3 at time 24, which suggests that the genes expressed similarly but at different times. It should be noted that the differential expression analysis (e.g., using Benjamini- Hochberg, GLMQL-MAS) does not captured this nuance (FIG.9B). Thus, exploratory analysis advantageously provides a more granular and precise understanding of viral gene expression activity.
[0268] In sum, the present exploratory analysis techniques enable understanding of how different viruses affect gene expression in a broader context rather than pinpointing specific genes. Using exploratory analysis, the present techniques enable visualization of the global gene expression landscape when cells are infected by different viruses. This is akin to viewing an entire forest’s health (exploratory analysis) rather than the health of specific trees (differential expression analysis). This allows researchers to discern patterns or trends that might be consistent across multiple viral infections. For instance, if two different viruses cause similar broad changes in gene expression, they might share some mechanism of action or target similar cellular pathways. Additional visualizations of the data are possible. For example, FIG.9C depicts an exemplary plot showing a comparison of the activity of IAV over time, relative to its Mock-infected counterpart, according to some aspects. FIG.9D depicts anAttorney Docket No.32103 / 59579 / PC exemplary plot showing a comparison of the activity of MPV over time, relative to its Mock-infected counterpart, according to some aspects. FIG.9E depicts an exemplary plot showing a comparison of the activity of PIV3 over time, relative to its Mock-infected counterpart, according to some aspects.
[0269] Further, the ability to understand the activity of expression among viruses over time enables other types of questions to be answered. For example, given an OTE exposed to an unknown virus, the present techniques may be used to determine (1) which virus the OTE is exposed to and (2) how much time has passed since exposure. Likewise, the present techniques may be used to identify an unknown virus (either as de novo or as a previously-unknown variation of a known virus) by comparing the expression profile of the unknown virus to known virus expression profiles.
[0270] The present exploratory analysis techniques may advantageously assist researchers to unearth hidden gene relationships. For example, as noted, relationships that might not be evident through differential expression analysis alone may be discovered, by using exploratory analysis to construct gene expression networks, identifying genes that tend to be co-expressed or those that show opposite expression patterns. Here, the goal may shift from finding the most changed genes to understanding how the entire network of genes shifts and adapts in response to an external factor, like a viral infection. This could unveil genes that, while not significantly differentially expressed themselves, play key roles in regulating or interacting with those that are. For example, a gene might not show major changes itself but could be crucial in regulating another set of genes that do show significant changes.
[0271] Exploratory analysis may also be used to conduct time-series analyses of gene expressions. For example, temporal dynamics may be explored using exploratory analysis by monitoring the progression of gene expression changes over time, especially in response to a viral infection. Instead of just comparing two points in time (pre-infection and post-infection), EA can analyze gene expression at multiple intervals post-infection, creating a dynamic picture of how cells respond over time, as shown in FIG.9A. By using a time series, transient changes may be captured, in addition to stages of response (early, mid, late), and future changes may be predicted. Differential analysis might show that a gene is upregulated post-infection, but exploratory analysis in a time-series context can show when that upregulation peaked, if the upregulation has decreased or increased, and how it corresponds with other genes’ expression.
[0272] For example, the following tables include results from analysis of differentially expressed genes (DEGs) in organ tissue equivalents (OTEs) infected with Influenza A virus (IAV), Human metapneumovirus (MPV), and Parainfluenza virus type 3 (PIV3) compared to mock-infected controls. This analysis is conducted at both 24 and 72 hours post-infection, utilizing the Benjamini-Hochberg (BH) method and the Bonferroni correction to adjust for multiple comparisons. The data presented in TablesAttorney Docket No.32103 / 59579 / PC S1 through S14 offer a detailed view of the transcriptional responses elicited by these viral infections, highlighting the specific genes and biological processes that are significantly affected. Table S1. Top 10 Significant P-Values and Q-Values for GO Biological Process 2023 in Upregulated Significant Genes When Comparing IAV-None-24 to Mock-24. Term p-value q-value Overlap genes Defense Response 1.744462e- 5.899769e-41 IFITM3, RTP4, IFITM1, IFITM2, IFIT5, CGAS, DDX60L, IFIT1, TNF, IFI44L, To Virus 44 IFIT3, IFIT2, OASL, IFIH1, TRIM5, DHX58, CASP1, TRIM25, TRIM26, (GO:0051607) TRIM21, IKBKE, GBP5, GBP7, RSAD2, RIPK3, NT5C3A, PLSCR1, AIM2, IFI27, OAS1, OAS2, IRF1, OAS3, IFIT1B, IRF2, IRF7, TLR3, TLR2, IFI6, DDX60, SAMHD1, USP18, IFNL2, IFNL1, IFI16, PMAIP1, GBP2, IFNL3, APOBEC3F, APOBEC3G, MLKL, IFNB1, STAT1, MX2, MX1, EIF2AK2, IFNLR1, ISG15, PML, BST2, ISG20, CXCL10, ZNFX1, TRIM31, SHFL, APOBEC3A, MYD88, APOBEC3B Defense Response 6.415411e- 1.084846e-39 IFITM3, RTP4, IFITM1, IFITM2, IFIT5, CGAS, DDX60L, IFIT1, IFI44L, IFIT3, To Symbiont 43 IFIT2, OASL, IFIH1, TRIM5, CASP1, IKBKE, GBP5, GBP7, RSAD2, RIPK3, (GO:0140546) NT5C3A, PLSCR1, AIM2, IFI27, OAS1, OAS2, IRF1, OAS3, IFIT1B, IRF2, IRF7, TLR3, TLR2, IFI6, DDX60, SAMHD1, IFNL2, IFNL1, IFI16, PMAIP1, GBP2, IFNL3, APOBEC3F, APOBEC3G, MLKL, IFNB1, STAT1, MX2, MX1, EIF2AK2, IFNLR1, ISG15, BST2, ISG20, ZNFX1, TRIM31, SHFL, APOBEC3A, MYD88, APOBEC3B Negative 9.861638e- 1.111735e-30 IFITM3, IFITM1, IFITM2, IFIT5, ZC3HAV1, IFIT1, TNF, OASL, IFIH1, ZFP36, Regulation Of Viral 34 IFI16, CCL5, TRIM21, IFNL3, N4BP1, APOBEC3F, RSAD2, APOBEC3G, Process IFNB1, STAT1, MX1, EIF2AK2, ISG15, BST2, ISG20, FAM111A, PLSCR1, (GO:0048525) ZNFX1, OAS1, TNIP1, OAS2, OAS3, TRIM14, TRIM31, SHFL, APOBEC3A Negative 1.429776e- 1.208876e-27 IFITM3, IFITM1, IFITM2, IFIT5, ZC3HAV1, IFIT1, TNF, OASL, IFIH1, IFI16, Regulation Of 30 CCL5, IFNL3, N4BP1, APOBEC3F, RSAD2, APOBEC3G, IFNB1, MX1, Viral Genome EIF2AK2, ISG15, BST2, ISG20, FAM111A, PLSCR1, ZNFX1, OAS1, TNIP1, Replication OAS2, OAS3, SHFL, APOBEC3A, APOBEC3B (GO:0045071) Regulation Of 6.718836e- 4.544621e-26 IFITM3, IFITM1, CXCL8, IFITM2, STAU1, IFIT5, ZC3HAV1, IFIT1, TNF, Viral Genome 29 OASL, IFIH1, IFI16, CCL5, IFNL3, N4BP1, GBP7, APOBEC3F, RSAD2, Replication APOBEC3G, IFNB1, MX1, EIF2AK2, ISG15, BST2, ISG20, FAM111A, (GO:0045069) PLSCR1, ZNFX1, OAS1, TNIP1, OAS2, OAS3, SHFL, APOBEC3A Response To 2.492655e- 1.405027e-17 IFITM3, CD274, CSF3, IFITM1, SP100, IFITM2, CALCOCO2, ADAR, Cytokine 20 CXCL16, RELB, NUB1, IRAK2, LAMP3, KYNU, RIPK1, LGALS9, JAK2, (GO:0034097) TRIM21, IKBKE, MCL1, GCH1, STAT1, MX2, MX1, EIF2AK2, LIFR, IFNLR1, ISG15, PML, BST2, CH25H, PLSCR1, IL1B, XAF1, SHFL, MYD88 Response To Type 9.547845e- 4.612973e-17 IFITM3, IFITM1, SP100, IFITM2, CALCOCO2, CX3CL1, CXCL16, NUB1, II Interferon 20 KYNU, CCL5, CCL4, CASP1, CCL3, LGALS9, GBP2, TRIM21, GBP1, GBP4, (GO:0034341) HLA-DPA1, GBP6, GBP5, CCL22, GCH1, CCL20, STAT1, BST2, CD47, SHFL, TLR2 Positive 1.923014e- 8.129543e-16 RET, CSF3, BMPR2, PLEKHF1, NCF1, RNF13, TRAF3IP2, RTKN2, HTR2B, Regulation Of 18 IFIT5, SECTM1, IFI35, TNF, CX3CL1, BBC3, CASP10, TRIM5, NUP62, Intracellular TNFSF10, CASP1, TRIM25, LGALS9, SOX9, JAK2, IKBKE, TRIM21, Signal TRIM22, NKX3-1, RIPK2, ARRDC3, IRAK4, TICAM1, IL23A, IL1B, PELI1, Transduction ADAM8, RBCK1, NUPR1, TMEM106A, S100A8, TLR3, HBEGF, BIRC3, (GO:1902533) TLR2, NOD2, MST1R, CLEC7A, CCL5, CCL4, MIER1, CCL3, PMAIP1, RICTOR, RIPK1, APOL2, LYN, TBX1, TCF7L2, NDFIP1, DAB2IP, LIF, EIF2AK2, SEPTIN4, CFLAR, BST2, BMP2, TNIP2, HPSE, TRIM38, PIK3AP1, MYD88, TRIM34 Regulation Of I- 1.211121e- 4.551124e-15 TRAF3IP2, IFIT5, HTR2B, SECTM1, TNFAIP3, NOD2, TNF, CX3CL1, kappaB 17 CASP10, TRIM5, CLEC7A, ZC3H12A, MIER1, NUP62, TNFSF10, CASP1, kinase / NF-kappaB TRIM25, RIPK1, LGALS9, TRIM21, IKBKE, TRIM22, APOL2, GBP7, NDFIP1, Signaling TLE1, RIPK2, STAT1, DAB2IP, NR1D1, IRAK4, CFLAR, TICAM1, BST2, (GO:0043122) TNIP1, TNIP2, IL1B, PELI1, RBCK1, TRIM38, OPTN, TMEM106A, MYD88, BIRC3, TRIM34Attorney Docket No.32103 / 59579 / PC Cellular Response 4.127660e- 1.395974e-13 RIPOR2, IFNA7, CSF3, IFITM2, CXCL8, IFNA1, IFIT1, CX3CL1, ZFP36, ZC3H12A, To Cytokine 16 CASP1, LGALS9, SOX9, JAK2, HLA-DPA1, NKX3-1, GBP6, GBP5, LIFR, IRAK4, Stimulus EREG, IL1A, AIM2, OAS1, IL1B, IRF1, CD47, TLR2, IL2RG, ZFP36L2, SOCS1, IFI16, (GO:0071345) CCL5, CCL4, CCL3, RIPK1, GBP2, GBP1, DUOX2, GBP4, GBP3, CCL22, PNPT1, CCL20, IFNB1, STAT1, DAB2IP, IL36G, NR1D1, NFKBIAAttorney Docket No.32103 / 59579 / PC Table S2. Top 10 Significant P-Values and Q-Values for GO Biological Process 2023 in Downregulated Significant Genes When Comparing IAV-None-24 to Mock-24. Term p-value q-value Overlap genes Cilium Assembly 3.878250e-10 0.000001 UNC119B, GALNT11, LRGUK, TRAF3IP1, ARL3, DNAH5, FNBP1L, SPATA6, (GO:0060271) KIF3B, TCTN3, TEKT1, CFAP43, CFAP20, DYNC2H1, BBS2, RSPH4A, GBF1, RFX2, CSNK1D, IFT80, RPGRIP1L, TMEM231, EHD2, AHI1, IFT88, UBXN10, TMEM216, TTC30A, CEP83, WDPCP, CCDC65, B9D2, CEP41, BBOF1, ATP6V0D1, IFT46, GAS8, SPACA9 Plasma Membrane 4.085312e-09 0.000006 UNC119B, GALNT11, SPEF1, TRAF3IP1, ARL3, DNAH5, FNBP1L, Bounded Cell SLC9A3R1, KIF3B, TCTN3, TEKT1, ARFIP2, INPP5K, SRGAP3, CFAP43, Projection Assembly CFAP20, DYNC2H1, BBS2, SPAG6, GBF1, RFX2, IFT80, EMP2, PARVB, (GO:0120031) TMEM231, EHD2, AHI1, IFT88, UBXN10, TMEM216, TTC30A, CEP83, WDPCP, CCDC65, B9D2, CEP41, ATP6V0D1, ARHGEF7, IFT46 Cilium Organization1.120370e-08 0.000011UNC119B, GALNT11, TRAF3IP1, ARL3, DNAH5, FNBP1L, TTC29, (GO:0044782) SLC9A3R1, KIF3B, TCTN3, TEKT1, CFAP43, CFAP20, DYNC2H1, BBS2, GBF1, RFX2, IFT80, TMEM231, ARMC2, EHD2, AHI1, IFT88, UBXN10, TMEM216, TTC30A, CEP83, WDPCP, CCDC65, B9D2, CEP41, MNS1, ATP6V0D1, IFT46 Organelle Assembly 1.101206e-07 0.000084 UNC119B, GALNT11, TRAF3IP1, ARL3, DNAH5, FNBP1L, FXR1, AP4M1, (GO:0070925) KIF3B, TCTN3, CHMP1B, RB1CC1, TP53INP2, TEKT1, TP53INP1, CFAP43, CFAP20, DYNC2H1, BBS2, YTHDF2, GBF1, RFX2, IFT80, TMEM231, EHD2, AHI1, IFT88, ATG16L1, UBXN10, TMEM216, TTC30A, CEP83, WDPCP, CCDC65, B9D2, CEP41, ATP6V0D1, IFT46, VPS25, ATG2B Mitochondrial 1.135464e-06 0.000694 DHX30, MTERF3, FASTKD2, NOA1, MRPS2, MRM2 Ribosome Assembly (GO:0061668) Cilium Movement 5.351788e-06 0.002609 DNAI4, RSPH4A, SPEF1, DNAH7, DNAAF2, SPAG6, DNAH5, DNAAF5, (GO:0003341) TTC29, CABYR, TEKT1, HYDIN, GAS8 Regulation Of 5.972299e-06 0.002609 GSK3B, CROCC, ENTR1, GAS8, LZTFL1 Protein Localization To Cilium (GO:1903564) Positive Regulation 2.983244e-05 0.010213 GSK3B, CROCC, ENTR1, GAS8 Of Protein Localization To Cilium (GO:1903566) Organelle 3.148164e-05 0.010213 DYRK3, STOML2, PEX11A, FNBP1L, TTC29, PHB2, SYNE1, SLC9A3R1, FAM174B, Organization CHMP1B, TMED3, TRAPPC12, PACSIN2, MLST8, CLASP2, DYNC2H1, COG8, (GO:0006996) JAGN1, YTHDF2, AGTPBP1, VPS13D, GBF1, COG2, LIG4, CSNK1D, PEX2, MTOR, ARMC2, SEC23IP, ABI2, VAPA, NOA1, KAT6A, RAB34, GOLGB1, TRIP11, MTFR1, WDPCP, MNS1, ARHGEF7, BRWD1 Cilium-Dependent 3.339782e-05 0.010213 DNAH3, DNAH2, DNAH7, DNAAF2, TEKT1, CCDC65, GAS8 Cell Motility (GO:0060285)Attorney Docket No.32103 / 59579 / PC Table S3. Top 10 Significant P-Values and Q-Values for GO Biological Process 2023 in Upregulated Significant Genes When Comparing MPV-None-24 to Mock-24. Term p-value q-value Overlap genes Defense Response To Symbiont 7.474605e- 5.060307e- IFITM1, RSAD2, MX2, MX1, ISG15, IFIT1, TANK, (GO:0140546) 14 11 IFIT3, IFIT2, PLSCR1, OAS1, OAS2, OAS3 Negative Regulation Of Viral Genome 1.381954e- 3.966404e- IFITM1, PLSCR1, RSAD2, OAS1, OAS2, OAS3, MX1, Replication (GO:0045071) 12 10 ISG15, IFIT1 Defense Response To Virus (GO:0051607) 1.757639e- 3.966404e- IFITM1, RSAD2, MX2, MX1, ISG15, IFIT1, TANK, 12 10 IFIT3, IFIT2, PLSCR1, OAS1, OAS2, OAS3Negative Regulation Of Viral Process 5.257900e- 8.898996e- IFITM1, PLSCR1, RSAD2, OAS1, OAS2, OAS3, MX1, (GO:0048525) 12 10 ISG15, IFIT1 Regulation Of Viral Genome Replication 1.269322e- 1.718661e- IFITM1, PLSCR1, RSAD2, OAS1, OAS2, OAS3, MX1, (GO:0045069) 11 09 ISG15, IFIT1 Positive Regulation Of Type I Interferon 3.483968e- 3.931077e- OAS1, OAS2, OAS3, ISG15, PTPN11, TANK Production (GO:0032481) 07 05Positive Regulation Of Interferon-Beta 4.700003e- 4.545574e- OAS1, OAS2, OAS3, ISG15, PTPN11 Production (GO:0032728) 07 05Interleukin-27-Mediated Signaling Pathway 7.900057e- 5.942599e- OAS1, OAS2, MX1 (GO:0070106) 07 05Regulation Of Ribonuclease Activity 7.900057e- 5.942599e- OAS1, OAS2, OAS3 (GO:0060700) 07 05Regulation Of Nuclease Activity 1.575049e- 1.016462e- OAS1, OAS2, OAS3 (GO:0032069) 06 04Table S4. Top 10 Significant P-Values and Q-Values for GO Biological Process 2023 in Downregulated Significant Genes When Comparing MPV-None-24 to Mock-24. Term p-value q-value Overlap genes Lysosomal Lumen Acidification 0.000003 0.003232 ATP6V0B, RNASEK, CCDC115, ATP6V1F (GO:0007042) Ribosome Assembly (GO:0042255) 0.000011 0.004007 EFNA1, RPS28, EIF6, MRPS2, RPS27L Regulation Of Lysosomal Lumen pH 0.000013 0.004007 ATP6V0B, RNASEK, CCDC115, ATP6V1F (GO:0035751) Golgi Lumen Acidification (GO:0061795) 0.000016 0.004007 ATP6V0B, RNASEK, ATP6V1F Vacuolar Acidification (GO:0007035) 0.000024 0.004772 ATP6V0B, RNASEK, CCDC115, ATP6V1F Endosomal Lumen Acidification 0.000032 0.005202 ATP6V0B, RNASEK, ATP6V1F (GO:0048388) Response To Cadmium Ion (GO:0046686) 0.000246 0.030167 NCF1, MT1F, MT1G Endosome Organization (GO:0007032) 0.000270 0.030167 VPS18, ATP6V0B, RNASEK, ATP6V1F Cellular Response To Cadmium Ion 0.000284 0.030167 NCF1, MT1F, MT1G (GO:0071276) Regulation Of Neutrophil Extravasation 0.000341 0.030167 MDK, CD99 (GO:2000389)Attorney Docket No.32103 / 59579 / PC Table S5. Top 10 Significant P-Values and Q-Values for GO Biological Process 2023 in Upregulated Significant Genes When Comparing PIV3-None-24 to Mock-24. Term p-value q-value Overlap genes Defense Response To Virus 3.544457e- 4.462471e- RTP4, IFITM1, IFIT5, DDX60L, IFIT1, DDX60, TANK, USP18, IFI44L, (GO:0051607) 33 30 IFIT3, IFIT2, OASL, IFIH1, IFNL2, IFNL1, PMAIP1, TRIM25, GBP5, RSAD2, STAT1, MX2, MX1, EIF2AK2, ISG15, NT5C3A, CXCL10,PLSCR1, ZNFX1, OAS1, OAS2, OAS3, IRF2, F2RL1 Defense Response To 2.897708e- 1.824107e- RTP4, IFITM1, IFIT5, DDX60L, IFIT1, DDX60, TANK, IFI44L, IFIT3, Symbiont (GO:0140546) 32 29 IFIT2, OASL, IFIH1, IFNL2, IFNL1, PMAIP1, GBP5, RSAD2, STAT1, MX2, MX1, EIF2AK2, ISG15, NT5C3A, PLSCR1, ZNFX1, OAS1,OAS2, OAS3, IRF2, F2RL1 Negative Regulation Of Viral 1.044136e- 4.381890e- IFITM1, RSAD2, MX1, IFIT5, EIF2AK2, ISG15, ZC3HAV1, IFIT1, Genome Replication 20 18 OASL, IFIH1, PLSCR1, ZNFX1, OAS1, OAS2, OAS3, TASOR (GO:0045071)Negative Regulation Of Viral 1.340752e- 4.220016e- IFITM1, RSAD2, STAT1, MX1, IFIT5, EIF2AK2, ISG15, ZC3HAV1, Process (GO:0048525) 19 17 IFIT1, OASL, IFIH1, PLSCR1, ZNFX1, OAS1, OAS2, OAS3 Regulation Of Viral Genome 7.091943e- 1.785751e- IFITM1, CXCL8, RSAD2, MX1, IFIT5, EIF2AK2, ISG15, ZC3HAV1, Replication (GO:0045069) 19 16 IFIT1, OASL, IFIH1, PLSCR1, ZNFX1, OAS1, OAS2, OAS3 Antiviral Innate Immune 8.350296e- 1.752171e- IFIH1, CXCL10, OAS1, MX1, TRIM25, EIF2AK2, IFIT1, USP18, IFIT3, Response (GO:0140374) 12 09 IFIT2 Interleukin-27-Mediated 5.584486e- 1.004410e- OAS1, OAS2, STAT1, MX1, OASL Signaling Pathway 11 08(GO:0070106)Cytokine-Mediated Signaling 1.227468e- 1.931728e- LYN, IL11, CXCL6, CXCL8, STAT1, MX1, PTPN11, CXCL5, OASL, Pathway (GO:0019221) 08 06 CXCL10, CXCL11, OAS1, OAS2, TRAF5, IL13RA2 Regulation Of Ribonuclease 3.150419e- 4.407087e- OAS1, OAS2, OAS3, OASL Activity (GO:0060700) 08 06Regulation Of Nuclease 9.384887e- 1.181557e- OAS1, OAS2, OAS3, OASL Activity (GO:0032069) 08 05Table S6. Top 10 Significant P-Values and Q-Values for GO Biological Process 2023 in Downregulated Significant Genes When Comparing PIV3-None-24 to Mock-24. Term p-value q-value Overlap genes Carboxylic Acid Catabolic Process (GO:0046395) 0.000039 0.027633 BCKDK, TST, ETFB Negative Regulation Of Cellular Response To Hypoxia (GO:1900038) 0.000153 0.046339 ENO1, NOL3 Glycolytic Process (GO:0006096) 0.000201 0.046339 LDHA, PGAM4, ENO1 Response To Hydroperoxide (GO:0033194) 0.000319 0.046339 GPX3, PRKD1 Cellular Response To Vascular Endothelial Growth Factor Stimulus 0.000325 0.046339 MT1G, PRKD1, PGF (GO:0035924) Amino Acid Catabolic Process (GO:0009063) 0.000453 0.053824 BCKDK, HAL, ETFB Carbohydrate Catabolic Process (GO:0016052) 0.000567 0.057803 LDHA, PGAM4, ENO1 Positive Regulation Of B Cell Differentiation (GO:0045579) 0.000679 0.059032 XBP1, PPP2R3C Positive Regulation Of Vasculature Development (GO:1904018) 0.000769 0.059032 XBP1, GRN, PRKD1, PGF Pentose-Phosphate Shunt (GO:0006098) 0.000828 0.059032 TALDO1, PGLSAttorney Docket No.32103 / 59579 / PC Table S7. Top 10 Significant P-Values and Q-Values for GO Biological Process 2023 in Upregulated Significant Genes When Comparing IAV-None-72 to Mock-72. Term p-value q-value Overlap genes Defense Response To Virus 4.106471e- 1.530482e- IFITM3, RTP4, IFITM1, CD40, IFITM2, IFIT5, DDX60L, IFIT1, TNF, (GO:0051607) 37 33 IFI44L, IFIT3, IFIT2, OASL, IFIH1, TRIM5, DHX58, CASP1, TRIM25, TRIM21, NCK1, GBP5, GBP7, RSAD2, RNASE1, NT5C3A, PLSCR1, AIM2, IFI27,OAS1, OAS2, IRF1, OAS3, IFIT1B, IRF7, TLR3, TLR2, IFI6, DDX60, SAMHD1, TANK, USP18, IFNL2, IFNL1, IFI16, PMAIP1, IFNL3, APOBEC3F, APOBEC3G, MLKL, IFNB1, STAT1, MX2, STAT2, MX1, EIF2AK2, ISG15, PML, BST2, ISG20, CXCL10, ZNFX1, SHFL, APOBEC3A, MYD88, APOBEC3B Defense Response To Symbiont 8.983444e- 1.674065e- IFITM3, RTP4, IFITM1, CD40, IFITM2, IFIT5, DDX60L, IFIT1, IFI44L, (GO:0140546) 36 32 IFIT3, IFIT2, OASL, IFIH1, TRIM5, CASP1, GBP5, GBP7, RSAD2, RNASE1, NT5C3A, PLSCR1, AIM2, IFI27, OAS1, OAS2, IRF1, OAS3,IFIT1B, IRF7, TLR3, TLR2, IFI6, DDX60, SAMHD1, TANK, IFNL2, IFNL1, IFI16, PMAIP1, IFNL3, APOBEC3F, APOBEC3G, MLKL, IFNB1, STAT1, MX2, STAT2, MX1, EIF2AK2, ISG15, BST2, ISG20, ZNFX1, SHFL, APOBEC3A, MYD88, APOBEC3B Negative Regulation Of Viral 1.457869e- 1.555334e- IFITM3, IFITM1, IFITM2, IFIT5, ZC3HAV1, IFIT1, TNF, OASL, IFIH1, Process (GO:0048525) 28 25 ZFP36, IFI16, CCL5, TRIM21, IFNL3, N4BP1, APOBEC3F, RSAD2, APOBEC3G, IFNB1, STAT1, MX1, EIF2AK2, ISG15, BST2, ISG20,PLSCR1, ZNFX1, OAS1, TNIP1, OAS2, OAS3, TRIM14, SHFL, APOBEC3A Negative Regulation Of Viral 1.669261e- 1.555334e- IFITM3, IFITM1, IFITM2, IFIT5, ZC3HAV1, IFIT1, TNF, OASL, Genome Replication 28 25 RESF1, IFIH1, IFI16, CCL5, IFNL3, N4BP1, APOBEC3F, RSAD2, (GO:0045071) APOBEC3G, IFNB1, MX1, EIF2AK2, ISG15, BST2, ISG20, PLSCR1, ZNFX1,OAS1, TNIP1, OAS2, OAS3, SHFL, APOBEC3A, APOBEC3B Regulation Of Viral Genome 3.590358e- 2.676253e- IFITM3, IFITM1, CXCL8, IFITM2, IFIT5, ZC3HAV1, IFIT1, TNF, Replication (GO:0045069) 24 21 OASL, IFIH1, IFI16, CCL5, IFNL3, N4BP1, GBP7, APOBEC3F, RSAD2, APOBEC3G, IFNB1, MX1, EIF2AK2, ISG15, BST2, ISG20, PLSCR1,ZNFX1, OAS1, TNIP1, OAS2, OAS3, SHFL, APOBEC3A Response To Cytokine 3.350446e- 2.081186e- IFITM3, CD274, IFITM1, CD40, SP100, IFITM2, CALCOCO2, ADAR, (GO:0034097) 18 15 CXCL16, RELB, NUB1, IRAK2, LAMP3, UBD, SMPD1, TIMP2, LGALS9, JAK2, TRIM21, MCL1, GCH1, STAT1, SPHK1, MX2, MX1,EIF2AK2, LIFR, ISG15, OSMR, PML, BST2, PLSCR1, COL3A1, XAF1, SHFL, MYD88 Regulation Of Inflammatory 1.098193e- 5.847091e- PTGER4, SEMA7A, TNFAIP6, SERPINE1, CXCL17, TNFAIP3, IFI35, Response (GO:0050727) 16 14 NOD2, METRNL, ETS1, TNF, USP18, CX3CL1, CAMK2N1, MDK, TMSB4X, CCL5, CASP1, CCL3, CCN4, CCN3, SLC39A8, JAK2, SNX6,IL15, SPHK1, WNT5A, MMP3, NR1D1, KLF4, IL22RA1, NFKBIA, ACE2, CYLD, PSMA6, FNDC4, AIM2, SELENOS, TNIP1, NINJ1, CD47, TEK, PIK3AP1, TLR3, MYD88, BIRC2, TLR2, BIRC3 Cellular Response To 1.123715e- 5.235105e- CD274, CXCL9, CXCL8, SERPINE1, TNFAIP3, NOD2, TNF, ZFP36, Lipopolysaccharide 15 13 CASP7, CCL5, ZC3H12A, ANKRD1, CASP1, CCL3, CCL2, IL36RN, (GO:0071222) GBP3, LYN, WNT5A, DAB2IP, IL36G, NR1D1, PDCD1LG2, TICAM1,CXCL10, CXCL11, SELENOS, TNIP1, AXL, TNIP2, CD68, MYD88, CHMP5 Response To 2.337121e- 9.678278e- CD274, CXCL9, CXCL8, SERPINE1, TNFAIP3, NOD2, ZFP36, CASP7, Lipopolysaccharide 15 13 ZC3H12A, ANKRD1, CASP1, CCL2, IL12A, IL36RN, LGALS9, JAK2, (GO:0032496) GBP3, GCH1, WNT5A, DAB2IP, ERBIN, IL36G, NR1D1, PDCD1LG2,TICAM1, CXCL10, CXCL11, SELENOS, TNIP1, AXL, TNIP2, PELI1, TAB2, TRIB1, CD68, MYD88, CHMP5 Regulation Of I-kappaB 3.316584e- 1.236091e- CD40, IFIT5, HTR2B, SECTM1, TNFAIP3, NOD2, TANK, TNF, kinase / NF-kappaB Signaling 15 12 CX3CL1, CASP10, TRIM5, UBD, ZC3H12A, MIER1, NUP62, TNFSF10, (GO:0043122) CASP1, TRIM25, SLC39A8, LGALS9, TRIM21, TRIM22, APOL2,GBP7, NDFIP1, STAT1, WNT5A, DAB2IP, SHISA5, NR1D1, TRAF1, CFLAR, TICAM1, BST2, TNIP1, TNIP2, PELI1, TAB2, TRIM38, OPTN, TMEM9B, MYD88, BIRC2, BIRC3, TRIM34Attorney Docket No.32103 / 59579 / PC Table S8. Top 10 Significant P-Values and Q-Values for GO Biological Process 2023 in Downregulated Significant Genes When Comparing IAV-None-72 to Mock-72. Term p-value q-value Overlap genes Translation (GO:0006412) 1.716745e- 5.337361e- RPL4, RPL5, RPL3, RPL31, RPLP1, RPLP0, MRPL37, RPL8, APEH, 34 31 RPL10A, RPL9, RPL6, EEF1B2, RPS4X, RPS14, RPL7A, RPS17, RPS16,RPL18A, NARS1, RPS18, RACK1, RPS10, RARS2, RPS9, MRPS27, RARS1,RPL21, MRPS24, RPS8, PABPC4, RPS5, RPL22, RPS6, MRPS2, RPL13A, RPSA, RPS3A, MRPL45, DARS1, EEF1A1, EEF1G, SARS1, EEF1D, LARS1, EEF1A2, RPL37A, RPL29, UBA52, RPL10, MRPS34, RPL11, RPS4Y1, RPS3, RPL13, IARS2, RPL15, RPS2, RPL18, RPS27A, RPL17, RPL19, EEF2, WARS2, GARS1, EPRS1, RPS21, RPS24, RPS23 Cytoplasmic Translation 8.489110e- 1.319632e- RPL4, RPL5, RPL3, RPL10, RPL31, RPLP1, RPLP0, RPL11, RPL10A, RPL8, (GO:0002181) 34 30 RPL9, RPL6, RPS4X, RPL7A, RPS14, RPS17, RPS16, RPL18A, RPS18, RACK1, RPS3, RPL13, RPL15, RPS2, RPL18, RPS27A, RPS10, RPL17,RPL19, RPS9, RPL21, RPS8, RPS5, RPL22, RPS6, RPL13A, RPS3A, RPSA, SARS1, RPL37A, RPL29, UBA52, RPS21, RPS24, RPS23 Macromolecule 3.075262e- 3.186997e- RPL4, RPL5, RPL3, RPL31, RPLP1, RPLP0, MRPL37, RPL10A, RPL8, RPL9, Biosynthetic Process 27 24 RPL6, EEF1B2, RPS4X, RPS14, RPL7A, RPS17, RPS16, RPL18A, RPS18, (GO:0009059) RPS10, RPS9, RPL21, RPS8, RPS5, PABPC4, RPL22, RPS6, RPL13A, RPSA,RPS3A, DARS1, EEF1A1, EEF1G, SARS1, EEF1D, EEF1A2, RPL37A, RPL29, RPL10, RPL11, RPS4Y1, RPS3, RPL13, APOE, RPL15, RPS2, RPL18, RPS27A, RPL17, RPL19, EEF2, RPS21, RPS24, RPS23 Peptide Biosynthetic 6.476539e- 5.033890e- RPL4, RPL5, RPL3, RPL31, RPLP1, RPLP0, MRPL37, RPL10A, RPL8, RPL9, Process (GO:0043043) 26 23 RPL6, RPS4X, RPS14, RPL7A, RPS17, RPS16, RPL18A, RPS18, RPS10, RPS9, RPL21, RPS8, RPS5, PABPC4, RPL22, RPS6, RPL13A, RPSA, RPS3A, DARS1, EEF1A1,SARS1, EEF1A2, RPL37A, RPL29, RPL10, RPL11, RPS4Y1, RPS3, RPL13, RPL15, RPS2, RPL18, RPS27A, RPL17, RPL19, RPS21, RPS24, RPS23 Cilium Movement 3.222724e- 2.003890e- SPAG17, SPEF2, SPEF1, DNAH7, DNAH5, DNAH9, RSPH9, CFAP206, (GO:0003341) 21 18 TTC29, CFAP251, TEKT1, HYDIN, ROPN1L, SPA17, DNAI1, DNAI4, CCDC39, DNAH11, RSPH4A, DNAI3, CFAP73, SPAG6, CFAP91, DNAAF6,NME5, CFAP100, CFAP53, GAS8 Gene Expression 7.381584e- 3.824891e- RPL4, RPL5, RPL3, RPL31, RPLP1, RPLP0, MRPL37, RPL8, RPL10A, RPL9, (GO:0010467) 15 12 RPL6, RBM3, RPS4X, RPS14, RPL7A, RPS17, RPS16, RPL18A, RPS18, RPS0, RPS9, RPL21, RPS8, PABPC4, RPS5, RPL22, RPS6, RPL13A, RPSA, RPS3A,DARS1, EEF1A1, SARS1, EEF1A2, RPL37A, RPL29, RPL10, RPL11, RPS4Y1, RPS3, RPL13, RPL15, RPS2, RPL18, RPS27A, RPL17, RPL19, SORL1, RPS21, RPS24, RPS23 Axoneme Assembly 1.386094e- 6.156236e- DNAI4, SPAG17, TTC26, CCDC39, RSPH4A, SPEF1, DNAAF3, CFAP91, (GO:0035082) 14 12 DNAAF6, DNAJB13, RSPH9, CFAP206, RP1, RSPH1, HYDIN, CCDC65, HOATZ, GAS8, SPACA9Cilium Assembly 6.669096e- 2.591777e- UNC119B, SPAG17, CEP126, TTC26, TRAF3IP1, ARL3, DNAH5, RSPH9, (GO:0060271) 12 09 CFAP206, SPATA6, HSPB11, CDC14B, DYNC2LI1, ABLIM1, KIF3A, TEKT1, CCDC96, CFAP47, CFAP43, CC2D2A, DYNC2H1, CCDC39,RSPH4A, GSN, DNAAF3, CCDC113, TTC39C, IFT88, RP1, RSPH1, CFAP161, CEP83, IFT27, CCDC65, HOATZ, B9D2, BBOF1, ZMYND10, CIBAR2, GAS8, SPACA9 Axonemal Dynein 1.157914e- 3.999950e- DNAI4, CCDC39, DNAH2, CFAP73, DNAI3, DNAH7, DNAAF3, DNAH5, Complex Assembly 09 07 DNAAF6, CFAP100, CCDC65, ZMYND10, DNAI1, DNAL1 (GO:0070286)Cilium Organization 3.851706e- 1.197495e- UNC119B, CEP126, TTC26, TRAF3IP1, ARL3, DNAH5, TTC29, HSPB11, (GO:0044782) 08 05 CDC14B, SLC9A3R1, ABLIM1, KIF3A, TEKT1, CCDC96, CFAP47, MARK4, CFAP43, CC2D2A, DYNC2H1, GSN, CCDC113, TTC39C, IQCG, ARMC2,IFT88, CFAP161, CEP83, IFT27, CCDC65, HOATZ, B9D2, MNS1, CIBAR2Attorney Docket No.32103 / 59579 / PC Table S9. Top 10 Significant P-Values and Q-Values for GO Biological Process 2023 in Upregulated Significant Genes When Comparing MPV-None-72 to Mock-72. Term p-value q-value Overlap genes Defense Response To Virus 3.318522e- 6.258732e- IFITM3, RTP4, IFITM1, IFIT5, IFI6, DDX60L, IFIT1, DDX60, SAMHD1, (GO:0051607) 36 33 USP18, IFI44L, IFIT3, IFIT2, OASL, IFIH1, IFI16, DHX58, TRIM25, TRIM21, RSAD2, STAT1, MX2, MX1, EIF2AK2, ISG15, PML, BST2,ISG20, NT5C3A, CXCL10, PLSCR1, ZNFX1, IFI27, OAS1, OAS2, IRF1, OAS3, IFIT1B, IRF7, SHFL, APOBEC3A, TLR3 Defense Response To 1.218238e- 1.148799e- IFITM3, RTP4, IFITM1, IFIT5, IFI6, DDX60L, IFIT1, DDX60, SAMHD1, Symbiont (GO:0140546) 32 29 IFI44L, IFIT3, IFIT2, OASL, IFIH1, IFI16, RSAD2, STAT1, MX2, MX1, EIF2AK2, ISG15, BST2, ISG20, NT5C3A, PLSCR1, ZNFX1, IFI27, OAS1,OAS2, IRF1, OAS3, IFIT1B, IRF7, SHFL, APOBEC3A, TLR3 Negative Regulation Of Viral 3.360369e- 2.112552e- IFITM3, IFITM1, RSAD2, STAT1, MX1, IFIT5, EIF2AK2, ISG15, Process (GO:0048525) 26 23 ZC3HAV1, IFIT1, OASL, IFIH1, BST2, ISG20, PLSCR1, ZNFX1, OAS1, IFI16, OAS2, OAS3, TRIM21, SHFL, APOBEC3ANegative Regulation Of Viral 1.434042e- 6.761510e- IFITM3, IFITM1, RSAD2, MX1, IFIT5, EIF2AK2, ISG15, ZC3HAV1, Genome Replication 24 22 IFIT1, OASL, IFIH1, BST2, ISG20, PLSCR1, ZNFX1, OAS1, IFI16, OAS2, (GO:0045071) OAS3, SHFL, APOBEC3ARegulation Of Viral Genome 4.800034e- 1.810573e- IFITM3, IFITM1, RSAD2, MX1, IFIT5, EIF2AK2, ISG15, ZC3HAV1, Replication (GO:0045069) 22 19 IFIT1, OASL, IFIH1, BST2, ISG20, PLSCR1, ZNFX1, OAS1, IFI16, OAS2, OAS3, SHFL, APOBEC3AResponse To Cytokine 5.825177e- 1.831047e- IFITM3, IFITM1, SP100, IL1R1, STAT1, MX2, MX1, EIF2AK2, ISG15, (GO:0034097) 16 13 ADAR, PML, BST2, PLSCR1, COL3A1, LAMP3, SMPD1, TIMP3, LGALS9, XAF1, TRIM21, SHFLResponse To Interferon-Beta 1.251634e- 3.372260e- IFITM3, BST2, IFITM1, PLSCR1, IFI16, OAS1, STAT1, IRF1, XAF1, (GO:0035456) 11 09 SHFL Antiviral Innate Immune 1.015002e- 2.392867e- IFIH1, CXCL10, OAS1, DHX58, MX1, TRIM25, EIF2AK2, IFIT1, USP18, Response (GO:0140374) 10 08 IFIT3, IFIT2 Interleukin-27-Mediated 9.390252e- 1.967779e- OAS1, OAS2, STAT1, MX1, OASL Signaling Pathway 10 07(GO:0070106)Response To Type I 1.174973e- 2.216000e- SP100, SMPD1, MX1, ISG15, IFIT1, SHFL Interferon (GO:0034340) 09 07Table S10. Top 10 Significant P-Values and Q-Values for GO Biological Process 2023 in Downregulated Significant Genes When Comparing MPV-None-72 to Mock-72. Term p-value q-value Overlap genes Regulation Of Neutrophil Migration (GO:1902622) 0.000262 0.159959 RAC2, OLFM4 Regulation Of Cytoplasmic Transport (GO:1903649) 0.000511 0.159959 MAP2K2, SRC Formation Of Cytoplasmic Translation Initiation Complex 0.000840 0.159959 EIF3M, EIF3H (GO:0001732) G Protein-Coupled Acetylcholine Receptor Signaling Pathway 0.001248 0.159959 GNA15, GRK2 (GO:0007213) Regulation Of Early Endosome To Late Endosome Transport 0.001563 0.159959 MAP2K2, SRC (GO:2000641) Regulation Of Centrosome Cycle (GO:0046605) 0.002294 0.159959 NPM1, CCNL2 Negative Regulation Of Protein Kinase Activity (GO:0006469) 0.002581 0.159959 DBNDD1, NPM1, CHP1 Cytoplasmic Translation (GO:0002181) 0.003019 0.159959 RPL4, EIF3M, RPL17 Intracellular Protein Transport (GO:0006886) 0.003335 0.159959 AKIRIN2, RAB1A, NPM1, CHP1, STX10 Cytoplasmic Translational Initiation (GO:0002183) 0.003889 0.159959 EIF3M, EIF3HAttorney Docket No.32103 / 59579 / PC Table S11. Top 10 Significant P-Values and Q-Values for GO Biological Process 2023 in Upregulated Significant Genes When Comparing PIV3-None-72 to Mock-72. Term p-value q-value Overlap genes Defense Response To 1.205983e- 4.276417e- IFITM3, RTP4, IFITM1, IFITM2, IFIT5, CGAS, DDX60L, IFIT1, IFI44L, Symbiont (GO:0140546) 48 45 IFIT3, IFIT2, OASL, PYCARD, IFIH1, TBK1, TRIM5, CASP1, DHX15, AZI2, IKBKE, GBP5, GBP7, RSAD2, NT5C3A, PLSCR1, AIM2, IFI27, OAS1, OAS2,IRF1, OAS3, IFIT1B, IRF2, IRF7, TLR3, TLR2, TRIM56, IFI6, DDX60, SAMHD1, TANK, IFNL2, IFNL1, IFI16, PMAIP1, IFNL3, APOBEC3C, APOBEC3F, APOBEC3G, MLKL, IFNB1, STAT1, MX2, STAT2, MX1, EIF2AK2, ISG15, BST2, ISG20, MOV10, ZNFX1, STING1, F2RL1, TRIM31, SHFL, APOBEC3A, MYD88, APOBEC3B Defense Response To Virus 6.670557e- 1.182690e- IFITM3, RTP4, IFITM1, IFITM2, IFIT5, CGAS, DDX60L, IFIT1, IFI44L, (GO:0051607) 48 44 IFIT3, IFIT2, OASL, PYCARD, IFIH1, TBK1, TRIM5, DHX58, CASP1, DHX15, TRIM25, AZI2, TRIM21, IKBKE, NCK1, GBP5, GBP7, RSAD2,NT5C3A, PLSCR1, AIM2, IFI27, OAS1, OAS2, IRF1, OAS3, IFIT1B, IRF2, IRF7, TLR3, TLR2, TRIM56, IFI6, DDX60, SAMHD1, TANK, USP18, IFNL2, IFNL1, IFI16, PMAIP1, IFNL3, APOBEC3C, APOBEC3F, APOBEC3G, MLKL, IFNB1, STAT1, MX2, STAT2, MX1, EIF2AK2, ISG15, PML, BST2, ISG20, CXCL10, MOV10, ZNFX1, STING1, F2RL1, TRIM31, SHFL, APOBEC3A, MYD88, APOBEC3B Negative Regulation Of Viral 5.709537e- 6.748673e- IFITM3, IFITM1, IFITM2, IFIT5, ZC3HAV1, IFIT1, OASL, IFIH1, ZFP36, Process (GO:0048525) 30 27 IFI16, CCL5, TRIM21, IFNL3, N4BP1, APOBEC3C, APOBEC3F, RSAD2, APOBEC3G, IFNB1, STAT1, MX1, EIF2AK2, ISG15, BST2, ISG20,FAM111A, PLSCR1, ZNFX1, OAS1, OAS2, OAS3, TRIM14, TRIM31, SHFL, APOBEC3A Negative Regulation Of Viral 1.669261e- 1.479800e- IFITM3, IFITM1, IFITM2, IFIT5, ZC3HAV1, IFIT1, OASL, RESF1, IFIH1, Genome Replication 28 25 IFI16, CCL5, IFNL3, N4BP1, APOBEC3C, APOBEC3F, RSAD2, (GO:0045071) APOBEC3G, IFNB1, MX1, EIF2AK2, ISG15, BST2, ISG20, FAM111A, PLSCR1,ZNFX1, OAS1, OAS2, OAS3, SHFL, APOBEC3A, APOBEC3B Regulation Of Viral Genome 3.590358e- 2.546282e- IFITM3, IFITM1, IFITM2, STAU1, IFIT5, ZC3HAV1, IFIT1, OASL, IFIH1, Replication (GO:0045069) 24 21 IFI16, CCL5, IFNL3, N4BP1, APOBEC3C, GBP7, APOBEC3F, RSAD2, APOBEC3G, IFNB1, MX1, EIF2AK2, ISG15, BST2, ISG20, FAM111A,PLSCR1, ZNFX1, OAS1, OAS2, OAS3, SHFL, APOBEC3A Response To Cytokine 3.350446e- 1.980114e- IFITM3, CD274, IFITM1, SP100, IFITM2, CALCOCO2, ADAR, CXCL16, (GO:0034097) 18 15 NUB1, IRAK2, LAMP3, UBD, CASP3, LGALS9, JAK2, TRIM21, IKBKE, MCL1, GCH1, IL1R1, STAT1, MX2, MX1, EIF2AK2, LIFR, ISG15, IRAK3,OSMR, PML, BST2, CH25H, PLSCR1, XAF1, SHFL, MYD88, TRIM56 Regulation Of Innate Immune 3.068497e- 1.554413e- CFH, CGAS, TNFAIP3, XIAP, IFI35, SAMHD1, IFI16, CCL5, DHX58, Response (GO:0045088) 16 13 TRIM21, N4BP1, GBP5, AKIRIN2, IFNB1, TREX1, ERAP1, TRAFD1, IRAK3, PARP9, EREG, HLA-E, PLSCR1, IRF1, ADAM8, BIRC2, BIRC3,TRIM56 Positive Regulation Of 5.735796e- 2.542392e- RET, BCAR3, BMPR2, PLEKHF1, NCF1, RNF13, RTKN2, HTR2B, IFIT5, Intracellular Signal 14 11 SECTM1, IFI35, FGF2, CX3CL1, BBC3, PYCARD, TBK1, CASP10,Transduction (GO:1902533) TRIM5, FAM110C, TNFSF10, DHX15, CASP1, TRIM25, LGALS9, SOX9,JAK2, IKBKE, TRIM21, TRIM22, ARRDC3, SHISA5, TICAM1, DDIT3, PELI1, ADAM8, RBCK1, NUPR1, BIRC2, TLR3, HBEGF, BIRC3, TRIM56, AVPI1, TLR2, NOD2, ADRB2, CCL5, UBD, DHX36, MIER1, PMAIP1, RICTOR, IGFBP6, APOL2, LYN, TCF7L2, NDFIP1, TGFB1, DAB2IP, LIF, EIF2AK2, SEPTIN4, CFLAR, BST2, AXL, F2RL1, TRIM38, PIK3AP1, MYD88, TRIM34 Regulation Of I-kappaB 7.258738e- 2.859943e- IFIT5, HTR2B, SECTM1, TNFAIP3, RORA, NOD2, TANK, CX3CL1, kinase / NF-kappaB Signaling 14 11 PYCARD, TBK1, CASP10, TRIM5, UBD, DHX36, MIER1, TNFSF10, (GO:0043122) DHX15, CASP1, TRIM25, LGALS9, TRIM21, IKBKE, TRIM22, APOL2,GBP7, NDFIP1, TLE1, STAT1, DAB2IP, SHISA5, NR1D1, CFLAR, TICAM1, BST2, PELI1, F2RL1, RBCK1, TRIM38, OPTN, MYD88, BIRC2, BIRC3, TRIM34 Response To Type II 5.049963e- 1.790717e- IFITM3, GBP6, GBP5, IFITM1, SP100, IFITM2, GCH1, CALCOCO2, Interferon (GO:0034341) 13 10 STAT1, CX3CL1, CXCL16, BST2, NUB1, CCL5, UBD, CASP1, LGALS9, CD47, TRIM21, GBP1, SHFL, GBP4, HLA-DPA1, TLR2Attorney Docket No.32103 / 59579 / PC Table S12. Top 10 Significant P-Values and Q-Values for GO Biological Process 2023 in Downregulated Significant Genes When Comparing PIV3-None-72 to Mock-72. Term p-value q-value Overlap genes Translation 3.64138 1.185634e- RPL4, RPL5, RPL30, RPL3, RPL32, RPL31, RPL34, RPL8, APEH, RPL10A, RPL9, RPL6, (GO:0006412) 3e-65 61 RPL7, EEF1B2, RPS14, RPS17, RPS16, RPL18A, RPS18, RPL38, RPL37, RPS11, RPS10, RARS2, RARS1, RPL21, RPS7, RPS8, RPS5, RPL22, RPS6, RPSA, TUFM, EEF1A1,LARS1, RPL27, RPL26, AURKAIP1, RPL29, UBA52, DAP3, MRPL12, RPS4Y1, EIF4EBP2, VARS1, MRPL23, EEF2, HARS2, RPS26, EPRS1, RPL27A, RPS20, RPS21, RPS24, FARSB, MRPS15, RPLP1, RPLP0, MRPL36, MRPL37, MRPL34, RPS4X, MRPL3, RPL7A, RACK1, RPLP2, MRPS28, MRPS24, PABPC4, MRPS2, RPL13A, RPS3A, MRPL45, MRPS6, EEF1G, SARS1, MRPL51, EEF1D, RPL37A, RPL10, RPL12, RPL11, MRPL55, RPS15A, RPS3, RPL13, IARS2, RPL15, RPS2, IARS1, RPL18, RPS27A, RPL17, RPL19, GADD45GIP1, RWDD1, GARS1, ABCE1 Cytoplasmic 1.38069 2.247765e- RPL4, RPL5, RPL30, RPL3, RPL32, RPL31, RPL34, RPLP1, RPLP0, RPL10A, RPL8, RPL9, Translation 1e-56 53 RPL6, RPL7, RPS4X, RPL7A, RPS14, RPS17, RPS16, RPL18A, RPS18, RACK1, RPLP2, (GO:0002181) RPL38, RPL37, RPS11, RPS10, RPL21, RPS7, RPS8, RPS5, RPL22, RPS6, RPL13A, RPS3A,RPSA, SARS1, RPL37A, RPL27, RPL26, RPL29, UBA52, RPL10, RPL12, RPL11, RPS15A, RPS3, RPL13, RPL15, RPS2, RPL18, RPS27A, RPL17, RPL19, RWDD1, RPS26, EIF3M, RPL27A, RPS20, RPS21, RPS24 Macromolecule 3.18695 3.458910e- RPL4, RPL5, RPL30, MRPS15, RPL3, RPL32, RPL31, RPL34, RPLP1, RPLP0, MRPL36, Biosynthetic 7e-56 53 MRPL37, RPL10A, RPL8, MRPL34, RPL9, RPL6, RPL7, EEF1B2, RPS4X, RPS14, MRPL3, ProcessRPL7A, RPS17, RPS16, RPL18A, RPS18, RPLP2, RPL38, RPL37, RPS11, RPS10, RPL21, RPS7,(GO:0009059) RPS8, RPS5, PABPC4, RPL22, RPS6, RPL13A, RPSA, RPS3A, MRPS6, TUFM, EEF1A1, EEF1G, SARS1, MRPL51, EEF1D, RPL37A, RPL27, RPL26, RPL29, RPL10, RPL12, RPL11, MRPL12, RPS4Y1, MRPL55, RPS15A, POLD2, RPS3, EIF4EBP2, RPL13, APOE, RPL15, RPS2, RPL18, RPS27A, RPL17, RPL19, EEF2, MRPL23, HARS2, RPS26, RPL27A, RPS20, RPS21, EIF4G2, RPS24, FARSB Peptide 2.18612 1.779509e- RPL4, RPL5, RPL30, MRPS15, RPL3, RPL32, RPL31, RPL34, RPLP1, RPLP0, MRPL36, Biosynthetic 9e-52 49 MRPL37, RPL10A, RPL8, MRPL34, RPL9, RPL6, RPL7, RPS4X, RPS14, MRPL3, RPL7A, ProcessRPS17, RPS16, RPL18A, RPS18, RPLP2, RPL38, RPL37, RPS11, RPS10, RPL21, RPS7,(GO:0043043) RPS8, RPS5, PABPC4, RPL22, RPS6, RPL13A, RPSA, RPS3A, MRPS6, EEF1A1, SARS1, MRPL51, RPL37A, RPL27, RPL26, RPL29, RPL10, RPL12, RPL11, MRPL12, RPS4Y1, MRPL55, RPS15A, RPS3, EIF4EBP2, RPL13, RPL15, RPS2, RPL18, RPS27A, RPL17, RPL19, MRPL23, HARS2, RPS26, RPL27A, RPS20, RPS21, RPS24, FARSB Gene Expression 4.48887 2.923152e- RPL4, RPL5, RPL30, MRPS15, RPL3, RPL32, RPL31, RPL34, RPLP1, RPLP0, MRPL36, (GO:0010467) 0e-35 32 MRPL37, RPL8, MRPL34, RPL10A, RPL9, RPL6, RPL7, RBM3, RPS4X, RPS14, MRPL3, RBM4, RPL7A, RPS17, RPS16, RPL18A, RPS18, RPLP2, RPL38, RPL37, RPS11, RPS10, RPL21,RPS7, RPS8, PABPC4, RPS5, RPL22, RPS6, RPL13A, RPSA, RPS3A, MRPS6, EEF1A1, SARS1, MRPL51, RPL37A, RPL27, RPL26, RPL29, RPL10, RPL12, RPL11, MRPL12, RPS4Y1, MRPL55, RPS15A, RPS3, RPL13, EIF4EBP2, RPL15, RPS2, RPL18, RPS27A, RPL17, RPL19, PRSS37, MRPL23, HARS2, SORL1, RPS26, RPL27A, RPS20, STUB1, RPS21, RPS24, FARSB Regulation Of 2.56402 1.391412e- RPL5, DDX6, RPL10, PRKDC, HSPB1, MSI2, YBX1, BZW2, FXR1, RBM3, RPS4X, Translation 7e-11 08 RBM4, ENC1, RACK1, RPS3, METTL5, EIF4B, NANOS1, EIF2B4, NPM1, PRMT1, ELP2, (GO:0006417) RPL13A, EEF2, LRPPRC, MTOR, EIF6, EPRS1, SERBP1, EIF3H, EIF3E, RPL26, VIM,GAPDH, EIF3D, EIF4G2 Aerobic 3.03285 1.410711e- NDUFA13, NDUFB7, NDUFA6, NDUFB10, UQCRB, NDUFB11, MDH2, NDUFA10, Respiration 4e-10 07 NDUFA2, NDUFA1, CHCHD10, UQCRH, OXA1L, NDUFS5, NDUFAB1, UQCRC1, (GO:0009060) MTFR1, NDUFV1protein-RNA 6.23927 2.539385e- RPL5, HSP90AB1, RPL10, PRKDC, RPLP0, RPL11, RPL6, RPS14, SNRPD2, PUF60, Complex 5e-09 06 RPL38, TXNL4A, XRCC5, RPS5, RPSA, NUDT21, EIF3M, EIF2S3, EIF6, EIF3K, EIF3L, AssemblyAGO1, STRAP, EIF3H, EIF3E, EIF3D, NOP53(GO:0022618) Cellular 2.97230 1.075312e- NDUFA13, NDUFB7, NDUFA6, NDUFB10, UQCRB, NDUFB11, MDH2, COX4I1, Respiration 1e-08 05 NDUFA10, NDUFA2, NDUFA1, ETFB, UQCRH, OXA1L, NDUFS5, NDUFAB1, UQCRC1, (GO:0045333) MTFR1, NDUFV1Positive 4.45028 1.449012e- RPL5, NDUFA13, HDAC2, PRKDC, FAF1, RAB1B, FXR1, RBM3, RPS4X, METTL5, Regulation Of 2e-08 05 APOE, MAP2K5, NPM1, PRMT1, WFS1, IL18, WNT7A, EEF2, SORL1, MTOR, DDB1,Attorney Docket No.32103 / 59579 / PC Protein MetabolicEIF6, EIF3E, STUB1, RPL26, VIM, AURKAIP1, SEC22B, EIF3D, VPS28Process (GO:0051247)Attorney Docket No.32103 / 59579 / PC Table S13. Top 10 Significant P-values and Q-values for GO Biological Process 2023 in Upregulated Significant Genes Common to IAV / MPV / PIV3 vs. Mock at 24 Hours Post-infection. Term p-value q-value Overlap genes Defense Response To Symbiont (GO:0140546) 3.055391e- 5.102503e- IFITM1, PLSCR1, RSAD2, OAS1, OAS2, MX2, OAS3, 22 20 MX1, ISG15, IFIT1, IFIT3, IFIT2Defense Response To Virus (GO:0051607) 6.289549e- 5.251774e- IFITM1, PLSCR1, RSAD2, OAS1, OAS2, MX2, OAS3, 21 19 MX1, ISG15, IFIT1, IFIT3, IFIT2Negative Regulation Of Viral Genome 1.501452e- 8.358081e- IFITM1, PLSCR1, RSAD2, OAS1, OAS2, OAS3, MX1, Replication (GO:0045071) 19 18 ISG15, IFIT1 Negative Regulation Of Viral Process 5.856740e- 2.445189e- IFITM1, PLSCR1, RSAD2, OAS1, OAS2, OAS3, MX1, (GO:0048525) 19 17 ISG15, IFIT1 Regulation Of Viral Genome Replication 1.440569e- 4.811501e- IFITM1, PLSCR1, RSAD2, OAS1, OAS2, OAS3, MX1, (GO:0045069) 18 17 ISG15, IFIT1 Antiviral Innate Immune Response 4.308594e- 1.199225e- OAS1, MX1, IFIT1, IFIT3, IFIT2 (GO:0140374) 10 08Response To Cytokine (GO:0034097) 9.218174e- 2.199193e- IFITM1, PLSCR1, MX2, MX1, ISG15, XAF1 10 08Regulation Of Ribonuclease Activity 6.112990e- 1.134299e- OAS1, OAS2, OAS3 (GO:0060700) 09 07Interleukin-27-Mediated Signaling Pathway 6.112990e- 1.134299e- OAS1, OAS2, MX1 (GO:0070106) 09 07Response To Interferon-Beta (GO:0035456) 1.075154e- 1.795507e- IFITM1, PLSCR1, OAS1, XAF1 08 07Table S14. Top 10 Significant P-values and Q-values for GO Biological Process 2023 in Upregulated Significant Genes Common to IAV / MPV / PIV3 vs. Mock at 72 Hours Post-infection. Term p-value q-value Overlap genes Defense Response To Virus 1.207965e- 1.211589e- IFITM3, RTP4, IFITM1, IFI6, IFIT5, DDX60L, IFIT1, DDX60, SAMHD1, (GO:0051607) 52 49 USP18, IFI44L, IFIT3, IFIT2, OASL, IFIH1, IFI16, DHX58, TRIM25, TRIM21, RSAD2, STAT1, MX2, MX1, EIF2AK2, ISG15, PML, ISG20,NT5C3A, BST2, CXCL10, PLSCR1, ZNFX1, IFI27, OAS1, OAS2, IRF1, OAS3, IFIT1B, IRF7, SHFL, APOBEC3A, TLR3 Defense Response To Symbiont 1.588775e- 7.967708e- IFITM3, RTP4, IFITM1, IFI6, IFIT5, DDX60L, IFIT1, DDX60, SAMHD1, (GO:0140546) 46 44 IFI44L, IFIT3, IFIT2, OASL, IFIH1, IFI16, RSAD2, STAT1, MX2, MX1, EIF2AK2, ISG15, ISG20, NT5C3A, BST2, PLSCR1, ZNFX1, IFI27, OAS1,OAS2, IRF1, OAS3, IFIT1B, IRF7, SHFL, APOBEC3A, TLR3 Negative Regulation Of Viral 7.229291e- 2.416993e- IFITM3, IFITM1, RSAD2, STAT1, MX1, IFIT5, EIF2AK2, ISG15, Process (GO:0048525) 35 32 ZC3HAV1, IFIT1, OASL, IFIH1, BST2, ISG20, PLSCR1, ZNFX1, OAS1, IFI16, OAS2, OAS3, TRIM21, SHFL, APOBEC3ANegative Regulation Of Viral 1.889924e- 4.738984e- IFITM3, IFITM1, RSAD2, MX1, IFIT5, EIF2AK2, ISG15, ZC3HAV1, Genome Replication 32 30 IFIT1, OASL, IFIH1, BST2, ISG20, PLSCR1, ZNFX1, OAS1, IFI16, OAS2, (GO:0045071) OAS3, SHFL, APOBEC3ARegulation Of Viral Genome 7.135177e- 1.431317e- IFITM3, IFITM1, RSAD2, MX1, IFIT5, EIF2AK2, ISG15, ZC3HAV1, Replication (GO:0045069) 30 27 IFIT1, OASL, IFIH1, BST2, ISG20, PLSCR1, ZNFX1, OAS1, IFI16, OAS2, OAS3, SHFL, APOBEC3AResponse To Cytokine 4.515305e- 7.548085e- IFITM3, IFITM1, SP100, IL1R1, STAT1, MX2, MX1, EIF2AK2, ISG15, (GO:0034097) 19 17 ADAR, PML, BST2, PLSCR1, LAMP3, LGALS9, XAF1, TRIM21, SHFL Response To Interferon-Beta 2.920577e- 4.184770e- IFITM3, BST2, IFITM1, PLSCR1, IFI16, OAS1, STAT1, IRF1, XAF1,Attorney Docket No.32103 / 59579 / PC (GO:0035456) 15 13 SHFL Antiviral Innate Immune 1.126368e- 1.412183e- IFIH1, CXCL10, OAS1, DHX58, MX1, TRIM25, EIF2AK2, IFIT1, USP18, Response (GO:0140374) 14 12 IFIT3, IFIT2 Response To Type II Interferon 7.265284e- 8.096756e- IFITM3, BST2, IFITM1, SP100, STAT1, LGALS9, CD47, TRIM21, GBP1, (GO:0034341) 12 10 SHFL, GBP4 Interleukin-27-Mediated 1.401159e- 1.405363e- OAS1, OAS2, STAT1, MX1, OASL Signaling Pathway 11 09(GO:0070106)
[0273] The GO analysis aimed to decipher the complex interactions between viral infections and host cellular mechanisms, focusing on the biological processes impacted by the upregulation and downregulation of significant genes.
[0274] At 24 h post-infection, GO analysis in tables S1, S3 and S5 above revealed an upregulation of genes associated with innate immune responses and antiviral defense mechanisms across all three viral infections, indicating a robust host defense strategy that involves interferon-stimulated genes and cytokine signaling pathways. Concurrently, the downregulation of genes (Tables S2, S4, S6) highlighted the viruses’ impact on host cellular structures, metabolism, and organelle functions, suggesting viral strategies to disrupt normal cell processes and evade immune surveillance. By 72 h post-infection, the host response showed a sustained activation of antiviral defense mechanisms (Tables S7, S9, S11), with a continued focus on cytokine-mediated immune responses and the inhibition of viral replication. This period also exhibited a significant downregulation in genes related to translation, cellular respiration, and gene expression (Tables S8, S10, S12), reflecting a deeper viral influence on host metabolic and biosynthetic pathways. Tables S13, S14 provided a comparative analysis of the common host responses to IAV, MPV, and PIV3 infections at both time points. These analyses underscored a conserved set of defense mechanisms activated by the host, including the upregulation of genes crucial for blocking viral entry, replication, and the modulation of immune signaling. This shared response highlights the fundamental aspects of the host’s antiviral defense and points to potential targets for broad-spectrum therapeutic interventions.
[0275] Figures 11A-11Q offer a detailed view of the transcriptional responses elicited by these viral infections, highlighting the specific genes and biological processes that are significantly affected.
[0276] It should be appreciated that in some aspects, both differential analysis and exploratory analysis may be complementary techniques that are used together, to understandAttorney Docket No.32103 / 59579 / PC the landscape of gene expression in addition to top genes. For example, in some aspects, a comprehensive analysis may include determining a gene expression of OTEs that have undergone viral infection. Differential analysis may use NanoString, which is expensive but includes cleaner expression data. RNASeq is more cost effective, but many genes are included that are unexpressed. It is also anticipated that exploratory analysis may be used with NanoString data. Time fingerprint identifier: Comprehensive Analysis of Gene Expression Dynamics Across 16 Infection Conditions Using Generalized Linear Models (GLMs) and Quasi-Likelihood (QL) Approach for Differential Expression Analysis
[0277] The present techniques may include analyzing RNA-Seq data and identifying genes that exhibit significant differential expression across 16 infection condition groups while demonstrating maximal expression with minimal variation compared to Naive-24. This process, akin to a one-way ANOVA test, is designed to unveil genes with consistent expression patterns across diverse conditions. To achieve this, the present techniques may employ Generalized Linear Models (GLM) in conjunction with the Quasi-Likelihood (QL) F-test. GLMs extend traditional linear models to accommodate non-normally distributed response data, capturing the intricate interplay between mean and variance. Within RNA-Seq analysis, GLMs serve to model gene expression levels relative to experimental conditions. The QL F-test is specifically chosen within the GLMs framework due to its handling of gene-specific dispersion uncertainty, ensuring robust control over error rates, particularly with our 6 replicates per condition. To implement the GLM, the present techniques construct a design matrix that integrates factors representing experimental conditions like treatments, time points, and infection conditions. This matrix encapsulates the connection between gene expression levels and conditions, facilitating subsequent statistical comparisons. Once the matrix is established, the present techniques fit RNA-Seq data to the GLM using the glmQLFit function within EdgeR. Specifically, the matrix and fitting operations may be performed by the statistical results module 152 of FIG.1, in some aspects.
[0278] In the present techniques, the reference level may be set as the minimum control group Naïve-24, establishing a baseline for comparing other conditions. This procedureAttorney Docket No.32103 / 59579 / PC assesses whether contrasts (Condition-Treatment-Time, Naïve-24) are non-zero. Condition- Treatment-Time can be any of the 16 conditions listed in Table 1, except Naïve-24. Comparing the GLM QL F-test with Naïve-24 as the reference to 15 infection conditions identifies a top gene with significant differential expression across conditions. It examines if any of the 15 conditions differ significantly from Naïve-24. This process of identifying genes exhibiting consistent expression patterns across conditions has the potential to serve as time biomarkers, allowing us to pinpoint the infection time point for OTEs regardless of the specific infection condition.
[0279] The culmination of our statistical methodology involves identifying potential biomarkers that can inform on the specific viral infection and its impact on gene expression. Post-application of the Generalized Linear Models with Quasi-Likelihood F-test (GLMQL), the present analysis proceeds to pinpoint genes exhibiting significant differential expression between the conditions compared, using both raw p-values and adjusted significance levels. The present techniques may utilize the GLMQL framework for precise identification of DE genes across experimental conditions, emphasizing those with significant p-values and notable log fold changes (logFC). This meticulous approach facilitates the discovery of genes whose expression profiles are distinctively altered in response to viral infections, earmarking them as potential biomarkers. Recognizing the importance of minimizing false discoveries in high-throughput data analysis, the present techniques may apply rigorous multiple testing corrections. This may includes the Benjamini-Hochberg (BH) method to control the false discovery rate and / or the Bonferroni correction for a stringent significance threshold. Such adjustments ensure that identified DE genes are not artifacts of multiple comparisons but reflect genuine biological differences. In the present analysis, the Magnitude-Altitude Score (MAS) and Relaxed Magnitude-Altitude Score (RMAS) algorithms may be used to prioritize genes for their potential as infection-specific biomarkers. These methodologies transcend conventional prioritization based solely on p- values or log fold changes (logFC), as commonly seen in tools like EdgeR (Robinson et al., 2010) or DESeq2 (Love et al., 2014).Attorney Docket No.32103 / 59579 / PC
[0280] The use of MAS and RMAS offers a nuanced approach to gene prioritization by capturing a gene’s overall impact on the study’s biological context. This is particularly crucial in RNA-Seq data analysis for several reasons: Holistic Gene Evaluation: Traditional prioritization based on p-values or logFC alone might overlook genes that, despite having moderate changes in expression or borderline statistical significance, play crucial roles in biological processes or pathways. MAS and RMAS integrate both dimensions, offering a more rounded assessment of each gene‘s potential impact. Enhanced Biological Relevance: By balancing the magnitude of change with statistical significance, MAS and RMAS help identify genes that are not only statistically significant but also biologically meaningful. This is vital for understanding complex biological responses, such as those seen in viral infections, where the interaction between host and pathogen can affect gene expression in nuanced ways. Adaptability to Data Variability: RNA-Seq datasets are characterized by inherent variability and complexity. The flexibility of MAS and RMAS in considering both expression change and significance level makes them particularly suited for such data, enabling the identification of relevant genes. Exploratory Insight: RMAS, with its use of raw p-values, allows for the exploration of data beyond the constraints of traditional statistical thresholds. This exploratory nature is invaluable for uncovering potential biomarkers or therapeutic targets that might be dismissed by more conservative methods.
[0281] To interpret the biological implications of the significant changes in gene expression observed upon viral infection, the present techniques may employ Gene Ontology (GO) enrichment analysis. This analysis facilitates the understanding of the biological processes, cellular components, and molecular functions enriched among the list of differentially expressed genes, thereby offering insights into the host’s defense mechanisms and the potential strategies employed by viruses to evade these defenses.
[0282] For the GO enrichment analysis, the present techniques may utilize the Enrichr platform (Chen et al., 2013), a comprehensive web-based tool designed to analyze gene lists forAttorney Docket No.32103 / 59579 / PC enrichment of specific GO terms. Enrichr incorporates multiple gene set libraries and employs robust statistical methods to identify significantly enriched terms, providing a deeper understanding of the biological themes associated with the gene lists.
[0283] Each list of significant upregulated and downregulated genes (raw p-values) identified for IAV, MPV, and PIV3 at 24- and 72-h post-infection was separately analyzed using Enrichr. The analysis process entailed the submission of gene symbols to the Enrichr platform, where the enrichment of GO terms across three main categories, Biological Process, Cellular Component, and Molecular Function, was assessed. The output from Enrichr included lists of enriched GO terms associated with each gene list, along with statistical metrics such as p- values and combined scores, which facilitated the prioritization and interpretation of the most relevant biological processes impacted by viral infection. Through this systematic GO enrichment analysis, the present techniques aimed to identify and compare the host cellular processes modulated in response to infection by each virus, thereby shedding light on the complexity of host-pathogen interactions and the dynamic nature of the host defense mechanisms. This methodological approach not only allowed for the exploration of the specific effects of individual viruses on the host cellular landscape but also provided a framework for identifying commonalities and differences in the host responses to these respiratory viral pathogens.
[0284] The present classification approach commenced with the application of the GLMQL- RMAS method to identify top genes by comparing IAV-infected OTEs against Mock samples at 24 h post-infection. Similar comparative analyses were conducted for MPV and PIV3-infected samples against Mock samples at the same time point. It is important to note that the top gene with the highest RMAS score is assigned a rank of 1, thereby assigning an RMAS rank to each gene. This initial phase enabled identification of significant genes uniquely expressed in response to each viral infection. The present techniques further refined the analysis by identifying common significant genes across these contrasts. For these common significant genes, an aggregated RMAS rank was defined as the summation of RMAS ranks for each comparison. Utilizing the RMAS scores, the gene with the smallest aggregated RMAS rank was selected as the fingerprint marker, capable of distinguishing all infected samples from the mockAttorney Docket No.32103 / 59579 / PC samples at the 24-h post-infection mark. This procedure was replicated for samples collected at 72 h post-infection, yielding two pivotal genes indicative of infection at 24- and 72-h post- infection intervals.
[0285] In instances where the top-selected gene at 24- and 72-h post-infection are the same, the present techniques opted for the second-highest ranked gene at 24 h as the primary selection. This adjustment was made because, within the first 24 h post-infection, the samples are less separable due to the nature of the viruses, and selecting an alternate top gene enhances the separation of samples during this critical early phase. This strategy ensures a more robust marker is utilized for distinguishing samples in the initial post-infection period.
[0286] To elucidate the dynamic changes between the two time points, the present techniques applied the GLMQL-RMAS method again, comparing the gene expressions of Mock, IAV, MPV, and PIV3 infected samples at 72 h against their respective 24-h expressions. An aggregated RMAS rank for all common significant genes was then defined as the summation of ranks from all four contrasts. Subsequently, the gene with the smallest aggregated RMAS rank, capable of separating samples at times 24 and 72, was chosen as the third pivotal gene.
[0287] Following this, a log-transformation was applied to the expression data of the three selected genes to normalize the distribution of expression levels, thereby enhancing the comparability and interpretability of gene expression across samples. To classify the samples into eight distinct classes (three viruses at two time points, along with two mock conditions), the present techniques employed a multinomial logistic regression model. This model was selected for its ability to handle multiple classes and provide probabilistic insights into class membership.
[0288] The present techniques adopted a Stratified K-Fold cross-validation strategy (Rodriguez et al., 2009) with six splits to evaluate the model’s performance. This approach ensured that each fold accurately represented the entire dataset by maintaining the proportion of samples for each class, thus allowing the present techniques to gauge the model’s generalizability and robustness across different data subsets. The model’s classification performance was quantitatively assessed using metrics such as accuracy, precision, recall, and the F1-score (Theodoridis and Koutroumbas, 2006). Furthermore, an aggregated confusionAttorney Docket No.32103 / 59579 / PC matrix provided a comprehensive visual summary of the model’s performance across all folds, detailing true positives, false positives, and misclassifications.
[0289] To evaluate the efficacy of RMAS relative to ranking methods commonly utilized in EdgeR (Robinson et al., 2010) and DESeq2 (Love et al., 2014), the present techniques replicated the aforementioned process, substituting RMAS with two alternative ranking criteria: once based on the smallest p-value and once based on the largest log2-fold change (log2FC). This comparative analysis allowed us to assess the performance and discriminative power of RMAS against traditional approaches in identifying pivotal genes indicative of viral infection. Infection fingerprint identifier: Identification of Differentially Expressed Genes for Pairwise Comparisons using Generalized Linear Models with Quasi-Likelihood (GLMQL) and the MAS Algorithm
[0290] In some aspects, the present techniques may employ GLMQL pairwise comparisons to investigate the impact of infection conditions on gene expression, and specifically, to identify genes exhibiting differential expression between specific pairs of conditions, such as Condition A and B (see Tables 1 for all 16 conditions). The present techniques may apply GLM to the data using a design matrix that captures the relationship between gene expression and conditions. Quasi-Likelihood F-tests evaluate expression disparities, revealing genes influenced by infection conditions. GLMQL's insights deepen understanding of how conditions affect genes, shedding light on biological processes.
[0291] Within the GLMQL approach, the present techniques may assess differential expression through p-values in condition comparisons. Log-fold changes generally provide insight into the direction and magnitude of expression differences. To control the false discovery rate, the present techniques may employ adjusted p-values, BH adjusted p-values. By analyzing these metrics, the present techniques may identify genes with significant differential expression, enhancing comprehension of condition distinctions and potential biomarker discovery. The present techniques may employ the MAS algorithm to prioritize genes, considering both adjusted p-values and log-fold changes. MAS highlights highly affected genes by striking a balance between these measures.Attorney Docket No.32103 / 59579 / PC
[0292] As discussed above, MAS computes using log-fold change (log2FC) and the logarithm (base 10) of BH-adjusted p-value, considering the magnitude of expression changes and their statistical significance. To investigate genes significantly affected by changes in infection conditions (IAV, MPV, and PIV3), the present techniques conduct twelve pairwise comparisons between different conditions as shown above in Table 3, encompassing virus-treatment-time triples. Exploratory Analysis of Gene Expression Patterns in Viral Infections using General Linear Models (GLMs) for Unveiling Gene-Virus Associations and Temporal Dynamics
[0293] This section performs a comprehensive analysis through pairwise comparisons of gene GLMQL-RMAS scores across diverse conditions associated with viruses IAV, MPV, and PIV3. The introduced RMAS scores (Equation 2) are derived from GLMQL-MAS scores (Equation 1) by substituting Benjamini-Hochberg adjusted p-values (pGFH) with raw p-values (pF), MAS = | lo K GH PJ g2FCJ| |LOG^+(PJ )| ,
[0294] Utilizing theinto gene expression patterns, enriching understanding of gene activity dynamics. GLMQL-RMAS scores primarily serve for exploration, contrasting the prioritization of significant differential expression by GLMQL-MAS. They provide a broader perspective on relationships and patterns, extending beyond rigid analysis. By considering these scores, the present techniques develop a comprehensive understanding of gene expression dynamics within various infection conditions associated with viruses. The present techniques perform a series of analyses to delve into the intricate dynamics of gene expression under viral infection conditions. Each analysis sheds light on distinct aspects of the interplay between viruses and gene activity. Comprehensive Virus-Condition Comparison
[0295] The present techniques begin our analysis by computing GLMQL-RMAS scores for six distinct infection condition changes: (Mock-24, IAV-None-24), (Mock-24, MPV-None-24), (Mock- 24, PIV3-None-24), (Mock-72, IAV-None-72), (Mock-72, MPV-None-72), and (Mock-72, PIV3-Attorney Docket No.32103 / 59579 / PC None-72). It's important to emphasize that, in all these scenarios, the present techniques utilize the Mock-infected control group as a consistent reference. This approach establishes a unified baseline for comparing the influence of different viruses on gene expression over time. The obtained scores, referred to as Virus-RMAS-Time (or Virus-MAS-Time) scores, offer insights into gene expression dynamics across various time points and different viruses following infection. Temporal Dynamics within Individual Viruses
[0296] Extending the analysis, the present techniques investigate temporal variations unique to each virus. Utilizing time-based pairs of GLMQL-RMAS scores for each virus i.e. (IAV-RMAS- 24, IAV-RMAS-72), (MPV-RMAS-24, MPV-RMAS-72), and (PIV3-RMAS-24, PIV3-RMAS-72), the present techniques explore how gene expression levels evolve over time within specific viral infections. This exploration reveals temporal trends and variations, offering insights into the evolving impact of viruses on gene expression patterns. Further, the present techniques delve into the interplay of gene expression profiles within specific virus infections. Employing General Linear Models (GLMs), the present techniques probe pairwise correlations among GLMQL- RMAS scores for different viruses at 24- and 72- hours. The application of GLMs enables quantification of the linear relationship's slope between gene expression scores, shedding light on gene expression levels at distinct time points. This approach illuminates how genes evolve from 24 to 72 hours after infection by a specific virus. Exploring Relationships via GLMs
[0297] Continuing the investigation, the present techniques shift focus to the interplay of gene expression profiles within distinct infection conditions. Utilizing General Linear Models (GLMs), the present techniques delve into the pairwise correlations among GLMQL-RMAS scores for different viruses at both 24 and 72 hours. Through GLM analysis, the present techniques derive the slope of the linear relationship between gene expression scores, thereby enhancing our comprehension of gene interactions and their contributions to the overall expression patterns under varying infection conditions. Additionally, the present techniques conduct a comprehensive analysis of the linear correlation among the intensity levels of different viruses using GLMs. Our study encompasses six pairwise comparisons: (IAV-RMAS-24, MPV-RMAS-Attorney Docket No.32103 / 59579 / PC 24), (IAV-RMAS-24, PIV3-RMAS-24), (MPV-RMAS-24, PIV3-RMAS-24), (IAV-RMAS-72, MPV- RMAS-72), (IAV-RMAS-72, PIV3-RMAS-72), and (MPV-RMAS-72, PIV3-RMAS-72). These comparisons provide valuable insights into the relationship and correlation between GLMQL- RMAS scores for different viruses at both 24 and 72-hour time points. Employing GLMs to analyze the data, the present techniques extract the slope of the linear relationship between GLMQL-RMAS scores for each pair of viruses. This slope grants deeper insights into the direction and magnitude of the correlation between virus-related intensities, bolstering comprehension of the interplay between different viruses at distinct time points. Identifying Correlated Genes
[0298] Drawing upon the insights gleaned from the previous analyses, the present techniques move to identify genes that exhibit consistent correlations with both viral infection and time. By amalgamating information from GLMQL-MAS and GLMQL-RMAS analyses, linear correlation assessments, and relationship explorations, the present techniques pinpoint genes that stand out due to significant trends, relationships, or changes in expression levels. These genes hold potential as valuable biomarkers, essential regulators, or indicators of the intricate host response to viral infections.
[0299] For the purpose of data visualization in the PCA, the present techniques applied a log2-transformation to the RNA-Seq data post-TMM normalization, facilitating the equalization of variance across samples. FIG.10A resents the 3D PCA plot, which provides a clear spatial segregation of the samples according to the defined experimental groups identified herein, demonstrating the distinct transcriptional profiles induced by each viral infection condition. It is important to note that the log2-transformation was specifically for PCA visualization; all other analyses reported in this paper are based on non-log-transformed data to preserve the original distribution and scale of expression values.
[0300] Throughout this study, in the GLMQL-RMAS / MAS analysis, the present techniques set both the M and A parameters to 1, with a significance level of α = 0.05. FIGs.10D and 10E, infra, display the total counts of upregulated and downregulated significant genes in IAV, MPV, and PIV3 infected samples, contrasted against Mock samples at 24- and 72-h post-infection, respectively. Subsequent FIGs.11 depict volcano plots for six contrasts at both 24- and 72-hAttorney Docket No.32103 / 59579 / PC post-infection using raw p-values. UV-treated samples served primarily as in-house controls to ascertain the baseline effects of viral component presence without replication. For a thorough analysis, the present techniques also conducted multiple hypothesis testing across all actively infected samples compared to their UV-treated counterparts, as shown in FIGs.11.
[0301] FIGs.11 also display the total counts of upregulated and downregulated genes, significant according to both BH and Bonferroni corrections, in samples infected with IAV, MPV, and PIV3, contrasted against Mock samples at 24- and 72-h post-infection, respectively. After applying the Benjamini-Hochberg adjustment and Bonferroni correction, all MPV-associated genes were rendered insignificant, attributed to the nature of the MPV virus. Consequently, for the remainder of this paper, we will focus our discussion on significant genes identified using raw p-values, as shown in FIGs.10.
[0302] FIGs.10F and 10G depict Venn diagrams illustrating the common significant genes associated with IAV, MPV, and PIV3, contrasted against Mock samples at 24- and 72-h post- infection, respectively. FIGs.11P and 11Q depict Venn diagrams showcasing the common BH and Bonferroni significant genes related to IAV, MPV, and PIV3 when compared against Mock samples at the same post-infection intervals. The genes are ranked according to their RMAS scores and aggregated RMAS rankings, reflecting their overall significance across two or three groups. These figures show the Venn diagrams for significant upregulated genes when logFC >0 and logFC >1. This aims to demonstrate the stability of RMAS ranking, showing that it remains consistent regardless of the logFC lower threshold (with the top gene being identically selected), unlike rankings based on only p-values.
[0303] FIG.10A depicts a PCA-based visualization of 96 OTEs studied, according to some aspects. In particular, outcomes derived from the methodologies outlined above are analyzed using PCA to deepen understanding, by applying PCA to visually represent all OTEs. FIG.10A depicts the projection of these OTEs (96 in total across all infection conditions) onto a three- dimensional subspace. This subspace is defined by the first three principal components extracted from the log2-transformed data. FIG.10A reveals a distinct pattern as OTEs undergo infection treatment changes between UV and None-UV treatments and across the 24-hour and 72-hour time points. For a closer examination, FIG.10B focuses specifically on OTEs infectedAttorney Docket No.32103 / 59579 / PC with the IAV virus. Notably, this pattern holds true for other viruses as well, albeit with an impact on gene expression that may be comparatively subdued. Time fingerprint identifier: Comprehensive Analysis of Gene Expression Dynamics Across 16 Infection Conditions Using Generalized Linear Models (GLMs) and Quasi-Likelihood (QL) Approach for Differential Expression Analysis
[0304] FIG.10C illustrates a heatmap of the outcomes of the analysis discussed above in regards to “Time fingerprint identifier: Comprehensive Analysis of Gene Expression Dynamics Across 16 Infection Conditions Using Generalized Linear Models (GLMs) and Quasi-Likelihood (QL) Approach for Differential Expression Analysis.” This heatmap clustering visualization showcases the top 20 selected genes: EIF5 (p-value=7.34E-31), GYS1 (p-value=2.29E-26), GJB6 (p-value=2.46E-25), GJB2 (p-value=8.76E-24), CA12 (p-value=8.03E-23), LARP1B (p- value=1.04E-22), RNF152 (p-value=1.15E-22), WNT5B (p-value=2.51E-22), PRPSAP1 (p- value=3.45E-22), NEFL (p-value=3.82E-22), PGAM1 (p-value=4.64E-22), GDE1 (p- value=4.84E-22), LURAP1L (p-value=6.24E-22), TMEM70 (p-value=7.43E-22), CPSF6 (p- value=8.47E-22), SSR3 (p-value=1.97E-21), KLF9 (p-value=2.34E-21), SKIL (p-value=2.53E- 21), and ENO1 (p-value=1.22E-20), while utilizing Naive-24 as the reference group. Notably, the heatmap exclusively portrays the log-fold change of genes from Naïve-24 to the other 15 infection conditions.
[0305] This selection emphasizes the most substantial gene expression differences observed across the experimental conditions. By designating Naive-24 as the reference group, the present techniques discerned and contrasted genes exhibiting significant changes in expression relative to this baseline condition, while also showcasing maximal expression with minimal variation compared to Naive-24. Each heatmap row corresponds to a specific gene, while each column represents an experimental condition. The intensity of the color gradient reflects the magnitude of the Log2 Fold Change, signifying the extent of differential expression as Naïve-24 is considered the control group. This process of identifying genes exhibiting consistent expression patterns across conditions serves as time biomarkers, allowing the present techniques to pinpoint the infection time point for OTEs regardless of the specific infection condition.Attorney Docket No.32103 / 59579 / PC Infection fingerprint identifier: Identification of Differentially Expressed Genes for Pairwise Comparisons using Generalized Linear Models with Quasi-Likelihood (GLMQL) and the MAS Algorithm
[0306] The present techniques may include applying the methodology described above in regards to “Infection fingerprint identifier: Identification of Differentially Expressed Genes for Pairwise Comparisons using Generalized Linear Models with Quasi-Likelihood (GLMQL) and the MAS Algorithm” to explore the impact of infection conditions on gene expression. To visualize the results, the present techniques may utilize volcano plots, which are effective tools for illustrating gene expression changes between conditions. Volcano plots provide a comprehensive depiction of gene expression variations by plotting the statistical significance (represented by -log10 BH adjusted p-values) against the magnitude of change (represented by log2 fold change) for each gene. For the specific pair conditions corresponding to IAV (Mock-24 to IAV-None-24, Mock-72 to IAV-None-72 and IAV-None-24 to IAV-None-72), the present techniques may include instructions for generating volcano plots as shown in FIGs.11A-11C, with genes ranked using the MAS algorithm to prioritize their significance.
[0307] In particular, FIG.11A depicts a GLMQL-MAS-based Volcano Plot: Mock-24 to IAV- None-24, according to some aspects. This volcano plot displays the differential gene expression analysis between the Mock-24 and IAV-None-24 conditions. FIG.11B depicts a GLMQL-MAS- based Volcano Plot: Mock-72 to IAV-None-72, according to some aspects. This volcano plot displays the differential gene expression analysis between the Mock-72 and IAV-None-72 condition. FIG.11C depicts a GLMQL-MAS-based Volcano Plot: IAV-None-24 to IAV-None-72, according to some aspects.. This volcano plot displays the differential gene expression analysis between the IAV-None-72 and IAV-None-72 conditions. Exploratory Analysis of Gene Expression Patterns in Viral Infections using General Linear Models (GLMs) for Unveiling Gene-Virus Associations and Temporal Dynamics Comprehensive Virus-Condition Comparison
[0308] As shown above, FIG.9A illustrates the gene GLMQL-MAS scores for different pairs of conditions. These scatter plots enable the relationship between gene expression patterns andAttorney Docket No.32103 / 59579 / PC the presence of specific viruses across various infection conditions and time points to be understood. Identifying Correlated Genes
[0309] In the results presented above regarding “Time fingerprint identifier: Comprehensive Analysis of Gene Expression Dynamics Across 16 Infection Conditions Using Generalized Linear Models (GLMs) and Quasi-Likelihood (QL) Approach for Differential Expression Analysis,” EIF5 stands out as the top gene as a time biomarker, allowing the infection time point for OTEs to be pinpointed regardless of the specific infection condition. The section above entitled “Infection fingerprint identifier: Identification of Differentially Expressed Genes for Pairwise Comparisons using Generalized Linear Models with Quasi-Likelihood (GLMQL) and the MAS Algorithm” demonstrates the consistent presence of IFIT1 across various conditions indicates its active involvement in the immune response against the investigated viruses. Its significant expression levels suggest a potential role in the defense mechanism against IAV and PIV3 infections. Furthermore, the findings shown in FIG.9A provide suggestive evidence that MX1 is influenced by the presence of MPV.
[0310] FIG.12A depicts the OTEs in a three-dimensional space, where the axes represent the logarithm with base 2 of IFIT1, MX1, and EIF5. This visualization provides valuable insights into the relationship and interactions among the infected OTEs. FIGs.9C-9E present a comparison of gene expression levels in OTEs infected with IAV, MPV, and PIV3 over time, relative to their respective Mock-infected counterparts. This comparison is visualized in a three- dimensional space, with the axes corresponding to the logarithm with base 2 of IFIT1 (infection fingerprint), MX1 (infection fingerprint), and EIF5 (time fingerprint) gene expressions. Discussion
[0311] FIG.10A and FIG.10B reveal noticeable patterns observed in all OTEs between the 24-hour and 72-hour post-infection time points. Irrespective of the virus type, all infected OTEs tend to shift towards the positive axis of the first principal component. However, during the time frame from 24 to 72 hours after infection, it becomes apparent that PIV3-infected OTEs move a greater distance compared to IAV- and MPV-infected OTEs. Furthermore, IAV-infected OTEsAttorney Docket No.32103 / 59579 / PC exhibit a greater shift compared to MPV-infected OTEs. Regarding the time frame between 0 and 24 hours after infection, IAV-infected OTEs at the 24-hour post-infection time point (IAV-24) exhibit a larger movement towards the positive axis of the first principal component compared to PIV3- and MPV-infected OTEs (PIV3-24 and MPV-24). IAV-24 OTEs are located further away from Mock-24 OTEs compared to PIV3-24 and MPV-24 OTEs. Moreover, PIV3-24 OTEs are positioned farther from Mock-24 OTEs compared to MPV-24 OTEs. This is because we consider Mock-infected OTEs at the 24- and 72-hour post-infection time points (Mock-24 and Mock-72) as the control groups. Time fingerprint identifier: Comprehensive Analysis of Gene Expression Dynamics Across 16 Infection Conditions Using Generalized Linear Models (GLMs) and Quasi-Likelihood (QL) Approach for Differential Expression Analysis
[0312] The outcomes of the analysis outlined above are depicted in 10C, featuring a heatmap clustering visualization of the top 20 selected genes. These genes- EIF5, GYS1, GJB6, GJB2, CA12, LARP1B, RNF152, WNT5B, PRPSAP1, NEFL, PGAM1, GDE1, LURAP1L, TMEM70, CPSF6, SSR3, KLF9, SKIL, and ENO1- exhibit maximal expression with minimal variation in comparison to Naive-24, which serves as the control group. As shown in 10C, these genes serve as effective "time-fingerprints," allowing us to determine the infection time point of an OTE (active, inactive, mock, or untreated). This identification is made by pinpointing genes that consistently maintain high expression levels over time, regardless of OTE infection condition. Additionally, FIG.10C demonstrates a clear distinction between OTEs at 24 hours and those at 72 hours post-infection. This distinction highlights the potential of these genes to act as distinguishing markers for determining the time of infection in OTEs, essentially serving as unique fingerprints for different experiment time points. Further, FIG.10D depicts a distribution of upregulated and downregulated significant genes in IAV, MPV, and PIV3 infected samples compared to Mock samples at 24 h post-infection, using raw p-values with a significance level of α = 0.05. FIG.10E depicts a distribution of upregulated and downregulated significant genes in IAV, MPV, and PIV3 infected samples compared to Mock samples at 72 h post-infection, using raw p-values with a significance level of α = 0.05, according to some aspects. FIG.10F depicts Venn diagrams showing the overlap of significant genes associated with IAV, MPV, and PIV3Attorney Docket No.32103 / 59579 / PC infections compared to Mock samples at 24 / 72 h post-infection, wherein genes are ranked based on RMAS scores and aggregated RMAS rankings to highlight their significance across the groups. (A) At 24 h post-infection. (B) At 72 h post-infection. FIG.10G depicts Venn diagram of significant upregulated genes with logFC >1 at 72 h post-infection; consistent with the previous figures, FIG 10G demonstrates the dependability of RMAS ranking in maintaining the identification of the top gene despite varying logFC thresholds, wherein (A) is a Venn diagram of significant upregulated genes with logFC >0 at 24 h post-infection; (B) is a Venn diagram of significant upregulated genes with logFC > 0 at 72 h post-infection; (C) is a Venn diagram of significant upregulated genes with logFC > 1 at 24 h post-infection; and (D) is a Venn diagram of significant upregulated genes with logFC > 1 at 72 h post-infection, all according to various aspects. Infection fingerprint identifier: Identification of Differentially Expressed Genes for Pairwise Comparisons using Generalized Linear Models with Quasi-Likelihood (GLMQL) and the MAS Algorithm
[0313] In this subsection the results of the section above entitled “Infection fingerprint identifier: Identification of Differentially Expressed Genes for Pairwise Comparisons using Generalized Linear Models with Quasi-Likelihood (GLMQL) and the MAS Algorithm,” are discussed, with specific focus on the differentially expressed analysis. This analysis requires to solely concentrate on the results of GLMQL-MAS based selected genes. The GLMQL-MAS method revealed distinct expression profiles of genes under various experimental conditions, shedding light on their potential roles in the cellular response against viral infections.
[0314] Regarding IFNL1, the present techniques indicate that UV treatment influences IFNL1 expression during both IAV and PIV3 infections. The present techniques observed differential expression between UV-treated and None-UV-treated IAV-infected OTEs at 24 and 72 hours, as well as between UV-treated and None-UV-treated PIV3-infected OTEs at 72 hours. These results suggest the involvement of UV-induced signaling pathways in modulating IFNL1 expression during viral infections. Additionally, we observed differential expression between mock-treated OTEs and None-UV-treated IAV- and PIV3-infected OTEs, highlighting the specific impact of viral infection on IFNL1 expression. These findings underscore the importanceAttorney Docket No.32103 / 59579 / PC of IFNL1 in the immune response against both IAV and PIV3 infections, and further research is needed to elucidate the underlying mechanisms and functional significance of IFNL1 in antiviral defense.
[0315] The expression patterns of IFIT1 reveal its potential role against viral infections. The present techniques observed differential expression between PIV3-infected OTEs with and without UV treatment at 72 hours, indicating the influence of UV-induced signaling pathways on IFIT1 expression during PIV3 infection. Regulation of IFIT1 was also observed during PIV3 infection, with differential expression between PIV3-infected OTEs at 24 and 72 hours without UV treatment. In the context of IAV infection, differential expression was observed between mock-treated OTEs and IAV-infected OTEs with or without UV treatment at 24 hours, emphasizing the specific impact of IAV infection and UV treatment on IFIT1 expression. Furthermore, differential expression of IFIT1 was observed between mock-treated OTEs and None-UV-treated IAV- or PIV3-infected OTEs at 72 hours, indicating the influence of viral infection and infection duration on IFIT1 expression. These findings suggest that IFIT1 is involved in the immune response against both IAV and PIV3 infections, and further exploration of its regulation could offer potential therapeutic avenues for combating viral infections.
[0316] IFIT2 expression patterns reveal its potential involvement in the cellular response against IAV and PIV3 infections. The present techniques observed differential expression between IAV-infected OTEs with and without UV treatment at 24 and 72 hours. Similarly, differential expression was observed between PIV3-infected OTEs with and without UV treatment at 72 hours. Moreover, differential expression of IFIT2 was observed between mock- treated OTEs and None-UV-treated IAV-infected OTEs at both 24 and 72 hours, highlighting the influence of IAV infection on IFIT2 expression. These findings suggest that IFIT2 plays a role in the immune response against IAV and PIV3 infections, and further investigation is warranted to uncover the underlying mechanisms and explore its therapeutic potential.
[0317] CXCL10, an important immune-related gene, exhibited differential expression patterns in response to IAV and PIV3 infections. The present techniques observed differential expression between IAV-infected OTEs with and without UV treatment at 24 and 72 hours, indicating the influence of UV treatment on CXCL10 expression during IAV infection. Similarly, differentialAttorney Docket No.32103 / 59579 / PC expression was observed between PIV3-infected OTEs with and without UV treatment at 72 hours. Moreover, differential expression of CXCL10 was observed between mock-treated OTEs and None-UV-treated IAV-infected OTEs at both 24 and 72 hours. These findings suggest the involvement of CXCL10 in the immune response against IAV and PIV3 infections. Further research is needed to unravel the specific mechanisms and functional significance of CXCL10 in antiviral defense.
[0318] OASL, another immune-related gene, exhibited differential expression patterns in the context of IAV and PIV3 infections. The present techniques observed differential expression between IAV-infected OTEs with and without UV treatment at 24 hours, as well as between PIV3-infected OTEs at 72 hours with and without UV treatment. Additionally, differential expression of OASL was observed between mock-treated OTEs and IAV-infected OTEs without UV treatment at both 24 and 72 hours. These findings suggest that OASL may play a role in the cellular response against IAV and PIV3 infections. Further investigation is necessary to uncover the underlying mechanisms and functional significance of OASL in antiviral defense.
[0319] The expression patterns of IFIT3 reveal its potential involvement in the immune response against IAV and PIV3 infections. The present techniques observed differential expression between PIV3-infected OTEs with and without UV treatment at both 24 and 72 hours. Differential expression of IFIT3 was observed between mock-treated OTEs and IAV- or PIV3-infected OTEs without treatment at 72 hours. These findings suggest that IFIT3 may contribute to the cellular response against IAV and PIV3 infections. Further research is necessary to elucidate the specific mechanisms and functional role of IFIT3 in antiviral defense.
[0320] The expression patterns of IFNL2 exhibited interesting dynamics during IAV and PIV3 infections. The present techniques observed differential expression between IAV-infected OTEs with and without UV treatment at both 24 and 72 hours. Similarly, differential expression was observed between PIV3-infected OTEs with and without UV treatment at 72 hours. Additionally, differential expression of IFNL2 was observed between untreated OTEs and IAV-infected OTEs without treatment at both 24 and 72 hours. These findings suggest the involvement of IFNL2 in the immune response against IAV and PIV3 infections. Further investigation is required toAttorney Docket No.32103 / 59579 / PC uncover the specific mechanisms and functional significance of IFNL2 and IFNL3 in antiviral defense.
[0321] The gene IFNL3 exhibits intriguing expression patterns under various experimental conditions, shedding light on its potential role in the cellular response against influenza A virus (IAV). Differential expression patterns were observed between IAV-infected OTEs treated with UV at 24 and 72 hours and IAV-infected OTEs without UV treatment, indicating the potential impact of UV treatment on IFNL3 expression during IAV infection. Furthermore, differential expression patterns were observed between Mock-treated OTEs and IAV-infected OTEs without UV treatment at 24 hours, emphasizing the specific influence of IAV infection on IFNL3 expression. These findings suggest the involvement of IFNL3 in the immune response against IAV, and further research is needed to uncover the underlying mechanisms and explore its potential as a therapeutic target.
[0322] The gene CXCL11 shows differential expression patterns under different experimental conditions, providing insights into its potential role in the immune response against IAV and PIV3 infections. The comparison between IAV-infected OTEs treated with UV and IAV-infected OTEs without UV treatment at 24 and 72 hours suggests that UV treatment may influence CXCL11 expression during IAV infection. Similarly, the comparison between mock-treated OTEs and IAV-infected OTEs without UV treatment at 24 hours highlights the specific impact of IAV infection on CXCL11 expression during the early stage of infection. Additionally, the comparison between PIV3-infected OTEs at 24 and 72 hours without any treatment suggests that PIV3 infection may also influence CXCL11 expression. These findings indicate the potential involvement of CXCL11 in the immune response against IAV and PIV3 infections, warranting further investigation to uncover the underlying mechanisms and functional significance of CXCL11 in antiviral defense.
[0323] RSAD2 exhibited differential expression patterns in the context of IAV and PIV3 infections. The present techniques observed differential expression between PIV3-infected OTEs with and without UV treatment at both 24 and 72 hours. Additionally, differential expression of RSAD2 was observed between Mock-treated OTEs and IAV-infected OTEs without treatment at both 24 and 72 hours. These findings suggest that RSAD2 may play a roleAttorney Docket No.32103 / 59579 / PC in the cellular response against IAV and PIV3 infections. Further investigation is necessary to uncover the underlying mechanisms and functional significance of RSAD2 in antiviral defense.
[0324] Differential expression patterns were observed for the gene IFNB1, an IAV-related gene, at the 24-hour time point. These patterns were observed between IAV-infected OTEs treated with UV and those without UV treatment, indicating the potential impact of UV treatment on IFNB1 expression during IAV infection. Furthermore, there were differential expression patterns between mock-treated OTEs and IAV-infected OTEs, suggesting the influence of IAV infection on IFNB1 expression. Additionally, differential expression patterns were observed between untreated OTEs and IAV-infected OTEs, highlighting the role of viral infection in modulating IFNB1 expression. These findings provide insights into the regulation of IFNB1 in response to IAV infection at the 24-hour time point, shedding light on its involvement in the immune response against IAV. Further research is needed to elucidate the underlying mechanisms and functional significance of IFNB1 in antiviral defense.
[0325] Similarly, at the 72-hour time point, the gene IFITM1, another IAV-related gene, exhibited differential expression patterns. These patterns were observed between mock-treated OTEs and IAV-infected OTEs treated with UV, suggesting the influence of UV treatment on IFITM1 expression during IAV infection. Likewise, differential expression patterns were observed between mock-treated OTEs and IAV-infected OTEs without any treatment, indicating the specific impact of IAV infection on IFITM1 expression. Moreover, there were differential expression patterns between untreated OTEs and IAV-infected OTEs, emphasizing the influence of IAV infection on IFITM1 expression. These findings provide insights into the regulation of IFITM1 in response to IAV infection at the 72-hour time point, suggesting its involvement in the host's immune response against IAV. Further exploration of its regulation could offer potential therapeutic avenues for combating IAV infections. Statistical methodology and gene ontology analysis
[0326] The PCA visualization in FIG.10A reveals that samples progressively shift towards the positive direction of PC1 from 24 to 72 h post-infection. FIGs.10E, 10F and 11L-11O suggest the stability of the MAS / RMAS ranking method, as it maintains consistent rankings even when applying Benjamini-Hochberg or Bonferroni corrections, provided the genes stillAttorney Docket No.32103 / 59579 / PC meet the significance criteria. Notably, IFNB1 and IFIT1 remain the top-ranked upregulated significant genes when contrasting IAV samples against Mock samples at both 24- and 72-h post-infection. This stability in ranking is further highlighted in Figure 5, where using RMAS for ranking shows that even when adjusting the lower logFC threshold from 0 to 1, the ranking of upregulated significant genes remains unchanged.
[0327] Analyzing the upregulated genes at 24 h post-infection (Figure 2), a distinct immunological response pattern for each virus was observed:
[0328] For IAV, genes such as IFNB1 (interferon-beta 1) and IFIT2 (interferon-induced protein with tetratricopeptide repeats 2) are central to the innate antiviral response. IFNB1 is pivotal for initiating a broad-spectrum antiviral state, while IFIT2 is known for its role in inhibiting viral protein synthesis. OASL (2′-5′-oligoadenylate synthetase-like) and RSAD2 (radical S- adenosyl methionine domain-containing 2) enhance viral RNA degradation, indicating a robust cellular mechanism to thwart viral replication.
[0329] For MPV, OAS2 and MX1 (myxovirus resistance 1) suggest activation of similar antiviral pathways. MX1 is involved in the inhibition of viral replication in the cytoplasm, adding an additional layer of cellular defense. ATG14 and ZDHHC16 (zinc finger DHHC-type containing 16) may indicate the engagement of autophagy-related pathways and membrane trafficking adjustments as cellular strategies against viral infection.
[0330] In the case of PIV3, IFIT3 (interferon-induced protein with tetratricopeptide repeats 3) upregulation is significant for the interception of viral replication, complementing the activities of IFIT1 and IFIT2. The increase in OAS2 and MX1 once again underscores a shared antiviral strategy across different viral infections, and ISG15 (interferon-stimulated gene 15) points to a broader modulation of the immune response, given its role in protein ubiquitination related to antiviral defense.
[0331] At the 72-h mark (Figure 3), the trend of gene expression changes suggests an adaptive shift in the host response. For instance, the upregulation of IFITM1 and OAS1 in IAV infections might signify a sustained defense mechanism, possibly transitioning from an immediate to a more regulated long-term response. The continued downregulation of structuralAttorney Docket No.32103 / 59579 / PC genes like ITGB4 and metabolic genes like EEF2 across different viruses could reflect a sustained reprogramming of cellular processes in the extended phase of viral infection. At 72 h post-infection, there is a clear pattern of upregulation for genes involved in the OTE’s response to viral infection: • For IAV, a sustained upregulation of IFIT1 and IFITM1 is observed, which are critical for the ongoing defense against viral replication and signaling to other immune system components. OAS3, OASL, and OAS1 continue to illustrate the activation of the oligoadenylate synthetase pathway, crucial for degrading viral RNA. RSAD2 and DDX58 (RIG-I) remain central for recognizing RNA viruses and initiating immune responses, suggesting a prolonged active defense mechanism. • For MPV, the persistence of MX1 and OAS1 upregulation is noted, suggesting a long- lasting immune response activation. Proteins like ZC3HAV1 (zinc-finger antiviral protein) play roles in viral mRNA sensing and decay, indicating a continued cellular effort to suppress viral gene expression. SRP68 and PLSCR1 implicate ongoing cellular adjustments in response to infection stress. • The PIV3 response highlights a similar trend with OAS3 and OAS1 again pointing to the significance of the antiviral state at this later time point. IFIT3, ISG15, and IFI44L emphasize a maintained immune response, with ISG15 indicating a broader immune system communication and potential modulation of the inflammatory response.
[0332] In the present extensive investigation of the host response to IAV, MPV, and PIV3 infections in OTEs at 24- and 72-h post-infection, Gene Ontology (GO) analysis has been instrumental in elucidating the complex interactions between viral invasion and host defense mechanisms. By analyzing both upregulated and downregulated genes at these time points, we have painted a detailed portrait of how host cells activate various biological processes to counteract viral threats. Concurrently, these cells undergo substantial changes in cellular functions to support the viral lifecycle. Given the absence of BH or Bonferroni significant genes for the MPV virus, the present techniques conducted the GO analysis using significant genes determined by raw p-values. This approach ensured a comprehensive understanding of the host’s response mechanisms across all studied viruses.Attorney Docket No.32103 / 59579 / PC
[0333] Gene ontology analysis of the host (OTE) response to influenza A virus
[0334] The present examination of OTEs infected with IAV at 24 h post-infection through Gene Ontology (GO) analysis (Table S1) has provided deep insights into the host’s defense mechanisms. This analysis reveals a multifaceted response to viral invasion, highlighting the significant roles played by a diverse array of genes. Key among these are interferon-stimulated genes such as IFITM3, IFITM1, IFITM2, and OAS1, which spearhead the defense against viral entry, replication, and spread. Additionally, genes integral to the innate immune response, such as CGAS and DDX60, have been identified, emphasizing the host’s rapid mobilization against the virus. The breadth of the host’s defense is further illustrated through the “Defense response to symbiont” category, pointing to a strategy effective against a range of pathogens. This broad- based approach is complemented by more targeted responses, such as those in the “Negative regulation of viral process” and “Negative regulation of viral genome replication” categories, where genes like ZC3HAV1 and APOBEC3G play pivotal roles in thwarting viral RNA degradation and replication. The analysis also uncovers the nuanced balance the host maintains through the “Regulation of viral genome replication,” revealing the intricate dance between host and virus. The importance of cytokine signaling in coordinating the immune response is underscored by findings in “Response to cytokine” and “Response to type II interferon” categories, with genes like CD40 and GCH1 highlighting the orchestration of inflammation and antiviral states.
[0335] Moreover, the host’s adaptive response is evident in the modulation of immune and inflammatory pathways, as shown by the involvement of genes in “Positive regulation of intracellular signal transduction” and “Regulation of I-kappaB kinase / NF-kappaB signaling.” This adaptability is crucial for mounting an effective defense and is further highlighted by the host’s responsiveness to cytokine signals, as seen in the “Cellular response to cytokine stimulus” category.
[0336] On the flip side, the analysis of downregulated genes (Table S2) reveals a significant reconfiguration of cellular priorities in the wake of IAV invasion. This includes alterations in cilium assembly and mitochondrial ribosome assembly, indicative of a strategic shift towards supporting viral replication at the expense of certain host functions.Attorney Docket No.32103 / 59579 / PC
[0337] At 72 h post-infection (Tables S7, S8), the present GO analysis traces the evolution of the host’s defense mechanisms, marking both continuity and adaptation. The sustained antiviral response, enriched in “Defense response to virus” and broadened by “Defense response to symbiont,” highlights enduring strategies against viral challenges. The ongoing efforts to curb viral RNA degradation and replication, as seen in “Negative regulation of viral process” and its counterpart for viral genome replication, reflect a persistent molecular defense.
[0338] This period also showcases the host’s regulatory finesse in “Regulation of viral genome replication,” ensuring viral control without compromising cellular integrity. The pivotal role of cytokine signaling in this extended defense phase is evident, with genes implicated in cytokine and inflammatory response modulation playing key roles in maintaining the systemic defense against IAV.
[0339] Furthermore, the activation typically associated with bacterial infections, as seen in “Cellular response to lipopolysaccharide” and “Response to lipopolysaccharide,” suggests a heightened state of immune readiness, potentially enhancing the host’s capability to manage co-infections.
[0340] The critical role of NF-kappaB signaling in mediating these responses, as highlighted in “Regulation of I-kappaB kinase / NF-kappaB signaling,” underscores the complexity and efficacy of the host’s defense mechanisms, offering insights into potential therapeutic targets to bolster resistance against IAV.
[0341] The examination of downregulated genes at this juncture sheds light on the long-term impacts of IAV on the host’s cellular machinery, revealing strategies likely aimed at optimizing viral replication and survival. This includes a notable suppression of host cell translation machinery and biosynthetic processes, suggesting a comprehensive viral strategy to reshape host cellular architecture and function.
[0342] Collectively, these insights from the GO analysis at both 24- and 72-h post-infection provide a comprehensive view of the host’s dynamic and evolving response to IAV, highlighting the complexity of antiviral defense mechanisms and identifying avenues for therapeutic intervention to support the host’s immune defense and recovery.Attorney Docket No.32103 / 59579 / PC
[0343] Gene ontology analysis of host (OTE) response to human metapneumovirus
[0344] The present investigation into the response of OTEs to MPV infection, at both the early and later stages post-infection (Tables S3, S4, S9, S10) through Gene Ontology (GO) analysis, sheds light on the intricate defense mechanisms mobilized by the host. This analysis identifies key biological processes and genes that are pivotal in the host’s strategy against MPV, revealing a nuanced and adaptive response to the viral threat.
[0345] Central to the host’s defense are the “Defense response to symbiont” and “Defense response to virus” processes, which underscore a comprehensive antiviral response encompassing a wide spectrum of strategies. This is exemplified by the activation of genes such as IFITM1, RSAD2, and the OAS gene family, alongside interferon-stimulated genes (ISGs) like MX1, MX2, and ISG15. These genes are instrumental in halting viral replication and spread, highlighting the host’s rapid and targeted defense mechanisms designed to counteract MPV invasion efficiently.
[0346] The host’s efforts to thwart viral proliferation are further evidenced by the “Negative regulation of viral genome replication” and “Negative regulation of viral process,” which focus on the molecular inhibition of viral replication. This strategic suppression, involving genes like IFIT1 and PLSCR1, signifies the host’s calculated approach to limit viral dissemination. Moreover, the “Positive regulation of type I interferon production” and related categories emphasize the host’s endeavor to boost the production of critical antiviral cytokines, a vital step in fortifying the host’s antiviral state and alerting adjacent cells to the viral presence.
[0347] Additionally, the GO analysis brings to light the modulation of the immune response through specific cytokines and enzymatic activities, as seen in the “Interleukin-27-mediated signaling pathway.” This not only demonstrates the host’s readiness and sophisticated response to viral encounters but also points to potential therapeutic targets that could enhance host defenses against MPV.
[0348] Conversely, the examination of downregulated genes during MPV infection unveils viral strategies aimed at manipulating host cellular functions and immune responses. This includes interference with “Lysosomal Lumen Acidification” and “Ribosome Assembly,”Attorney Docket No.32103 / 59579 / PC potentially affecting protein degradation and synthesis. Such viral tactics may aim to divert host mechanisms to favor viral replication or impede host defenses.
[0349] Insights into the “Regulation of Neutrophil Migration” and “Regulation of Cytoplasmic Transport” reveal viral influences on immune cell dynamics and intracellular movement, possibly altering the inflammatory landscape and host protein production priorities. Furthermore, changes in “G Protein-Coupled Acetylcholine Receptor Signaling Pathway” and “Regulation of Early Endosome to Late Endosome Transport” suggest viral modifications to cellular signaling and trafficking, which could impact viral entry and immune responses.
[0350] Through this comprehensive analysis, the present techniques delve into the dynamic interactions between MPV and the host, uncovering the adaptive and multifaceted nature of the host’s defense mechanisms. These findings not only deepen our understanding of the biological processes at play during MPV infection but also offer a groundwork for developing strategies to restore host functions and bolster the immune defense, providing a clear direction for future therapeutic interventions.
[0351] Gene ontology analysis of host (OTE) response to parainfluenza virus type 3
[0352] The present investigation into PIV3 infection in OTEs at 24- and 72-h post-infection, documented in Tables S5, S6, S11, S12 has yielded insightful revelations about the host’s defense strategy through Gene Ontology (GO) analysis. This detailed examination highlights a nuanced host response meticulously orchestrated to counteract PIV3 invasion, spotlighting specific biological processes and genes pivotal in mounting an effective defense.
[0353] At the forefront of this defense is the “Defense response to virus” category, which brings to light the host’s reliance on interferon-stimulated genes (ISGs) like IFITM1, IFIT5, and the OAS gene family. These genes are instrumental in blocking viral entry and replication, epitomizing the host’s immediate and robust countermeasures against the viral threat.
[0354] Moreover, the “Defense response to symbiont” term underscores the host’s broad immune capability, showcasing its adaptability to combat not just viruses but a spectrum of pathogens. This is complemented by efforts detailed under “Negative regulation of viral genomeAttorney Docket No.32103 / 59579 / PC replication,” where the host employs a targeted molecular assault on the viral lifecycle, as evidenced by the actions of RSAD2 and MX1.
[0355] The analysis also shines a light on the “Antiviral innate immune response,” underscoring the innate mechanisms activated to detect and neutralize viral components, thereby initiating a comprehensive immune response. The significance of cytokine signaling in orchestrating these defenses is elaborated through the “Interleukin-27-mediated signaling pathway” and “Cytokine-mediated signaling pathway,” highlighting the crucial roles of genes that facilitate antiviral states and mobilize immune cells to the infection site.
[0356] On the flip side, the examination of downregulated genes presents a clear picture of PIV3’s strategic impact on host metabolism and cellular processes. This includes notable disruptions in carboxylic acid catabolism and the cellular response to hypoxia, indicating a viral- induced shift in host energy metabolism and adaptations to the metabolic requirements imposed by the infection.
[0357] Processes like “Carboxylic Acid Catabolic Process” and “Negative Regulation of Cellular Response to Hypoxia” point to a deliberate alteration in energy management and hypoxic responses, suggesting the virus’s influence on host cellular conditions to favor its replication. Additionally, “Response to Hydroperoxide” and related changes in the oxidative stress response highlight the host’s adjustments to manage cellular damage and foster tissue repair under viral attack.
[0358] This comprehensive analysis not only deepens the understanding of the sophisticated interplay between PIV3 and its host but also illuminates potential targets for therapeutic intervention. By highlighting the intricate defense mechanisms activated by the host and the strategic viral maneuvers to subvert these responses, the present techniques identify avenues to strengthen host defenses, aiming to mitigate the impact of PIV3 infection and support the host’s recovery and resilience against viral challenges.
[0359] Comparative analysis of host (OTE) response to IAV, MPV and PIV3 infections
[0360] The present extensive Gene Ontology (GO) analysis across OTEs infected with Influenza A virus (IAV), Human Metapneumovirus (MPV), and Parainfluenza virus type 3 (PIV3)Attorney Docket No.32103 / 59579 / PC at both 24 and 72 h post-infection unveils a dynamic and multifaceted host defense against these respiratory viruses. Each virus triggers a distinctive host response, leveraging a blend of common and unique strategies to counter viral threats effectively. This response encompasses a broad activation of interferon-stimulated genes, innate immune mechanisms, and a strategic modulation of cellular processes to optimize defense while navigating viral evasion tactics.
[0361] IAV elicits a potent antiviral defense, characterized by the mobilization of key interferon-stimulated genes (IFITM3, IFITM1, IFITM2, OAS1) and critical innate immune response genes (CGAS, DDX60). This robust response is augmented by a broad “Defense response to symbiont,” signifying the host’s versatile defense strategy against various pathogens. The host’s approach is further defined by targeted molecular interventions (ZC3HAV1, APOBEC3G) to curb viral proliferation, underpinned by a critical emphasis on cytokine signaling to orchestrate an integrated immune and inflammatory response, highlighting the role of genes such as CD40 and GCH1.
[0362] MPV infection showcases a similar reliance on a comprehensive antiviral response, with genes like MX1, MX2, ISG15, IFITM1, and RSAD2 playing pivotal roles in halting viral replication. This response is coupled with strategic molecular efforts (IFIT1, PLSCR1) to suppress viral replication and amplify type I interferon production, underscoring the host’s adaptive immune response modulation.
[0363] PIV3 prompts a refined host defense, focusing on obstructing viral entry and replication through the activation of IFITM1, IFIT5, and the OAS gene family. This targeted defense is complemented by a “Defense response to symbiont,” illustrating a broad immune capability. Moreover, PIV3 infection underlines the host’s direct efforts (RSAD2, MX1) to dampen viral processes and genome replication, spotlighting the role of innate immune mechanisms and cytokine signaling in mounting a coordinated defense.
[0364] Upregulated genes and host defense mechanisms
[0365] The present extensive Gene Ontology (GO) analysis across OTEs infected with Influenza A virus (IAV), Human Metapneumovirus (MPV), and Parainfluenza virus type 3 (PIV3) at both 24 and 72 h post-infection unveils a dynamic and multifaceted host defense againstAttorney Docket No.32103 / 59579 / PC these respiratory viruses. Each virus triggers a distinctive host response, leveraging a blend of common and unique strategies to counter viral threats effectively. This response encompasses a broad activation of interferon-stimulated genes, innate immune mechanisms, and a strategic modulation of cellular processes to optimize defense while navigating viral evasion tactics.
[0366] IAV elicits a potent antiviral defense, characterized by the mobilization of key interferon-stimulated genes (IFITM3, IFITM1, IFITM2, OAS1) and critical innate immune response genes (CGAS, DDX60). This robust response is augmented by a broad “Defense response to symbiont,” signifying the host’s versatile defense strategy against various pathogens. The host’s approach is further defined by targeted molecular interventions (ZC3HAV1, APOBEC3G) to curb viral proliferation, underpinned by a critical emphasis on cytokine signaling to orchestrate an integrated immune and inflammatory response, highlighting the role of genes such as CD40 and GCH1.
[0367] MPV infection showcases a similar reliance on a comprehensive antiviral response, with genes like MX1, MX2, ISG15, IFITM1, and RSAD2 playing pivotal roles in halting viral replication. This response is coupled with strategic molecular efforts (IFIT1, PLSCR1) to suppress viral replication and amplify type I interferon production, underscoring the host’s adaptive immune response modulation.
[0368] PIV3 prompts a refined host defense, focusing on obstructing viral entry and replication through the activation of IFITM1, IFIT5, and the OAS gene family. This targeted defense is complemented by a “Defense response to symbiont,” illustrating a broad immune capability. Moreover, PIV3 infection underlines the host’s direct efforts (RSAD2, MX1) to dampen viral processes and genome replication, spotlighting the role of innate immune mechanisms and cytokine signaling in mounting a coordinated defense.
[0369] Downregulated genes and cellular alterations
[0370] Notably, the downregulation of genes associated with “Cilium Assembly” (GO:0060271) and “Organelle Assembly” (GO:0070925) at 24 h post-IAV infection (Table S2), and similar trends in MPV (Table S4) and PIV3 infections (Table S6), suggests a strategic viral interference with cellular structures and organelle functions. This might facilitate viral evasion orAttorney Docket No.32103 / 59579 / PC replication by altering cellular priorities and homeostasis. Furthermore, the downregulation observed in “Translation”(GO:0006412) and “Cytoplasmic Translation” (GO:0002181) at 72 h post-IAV and PIV3 infection (Tables S8, S12) points towards a viral strategy to dominate the host cell’s protein synthesis machinery.
[0371] Integrated host response across viral infections
[0372] The GO analysis of genes common to IAV, MPV, and PIV3 infections at both 24 (Table S13) and 72 h (Table S14) post-infection reveals a core set of biological processes activated in response to these viral infections. This integrated host defense mechanism, involving both upregulation of antiviral genes and downregulation of genes related to cellular maintenance and metabolism, suggests a strategic host response that prioritizes defense mechanisms while modulating cellular functions to mitigate viral invasion.
[0373] Comparative analysis and insights for therapeutic intervention
[0374] The present comparative analysis underscores both unique and shared host response pathways across IAV, MPV, and PIV3 infections. While core antiviral mechanisms, such as interferon signaling and viral genome replication inhibition, are universally activated, variations in specific GO terms and associated genes highlight the unique interactions between the host and each virus. This understanding not only deepens our insights into the molecular underpinnings of host-virus interactions but also illuminates potential targets for broad-spectrum and virus-specific therapeutic interventions aimed at enhancing host defenses and mitigating the adverse effects of viral infections.
[0375] In conclusion, this extensive GO analysis across early and later stages post-infection offers a valuable framework for future research into the mechanisms of host resistance and virus pathogenesis. It highlights the complexity of the host’s defense strategies, the adaptability of viral mechanisms to circumvent these defenses, and the potential for identifying novel therapeutic targets. These findings contribute significantly to our understanding of viral infections in OTEs, providing a solid foundation for the development of effective antiviral strategies and enhancing our capacity to combat respiratory viral pathogens.Attorney Docket No.32103 / 59579 / PC Exploratory Analysis of Gene Expression Patterns in Viral Infections using General Linear Models (GLMs) for Unveiling Gene-Virus Associations and Temporal Dynamics
[0376] FIG.9A demonstrates the gene activity patterns of viruses IAV, MPV, and PIV3 during different time frames after infection. Within the first 24 hours after infection, IAV exhibits the highest level of activity among the three viruses. A majority of genes are impacted by IAV during this time period, indicating a significant influence on gene expression. On the other hand, both MPV and PIV3 show relatively lower levels of activity compared to IAV. However, PIV3 exhibits a higher intensity of activity compared to MPV during the first 24 hours.
[0377] Moving to the post-infection time frame between 24 and 72 hours, a notable shift in gene activity is observed. PIV3 emerges as the most influential virus during this time period, significantly impacting a large number of genes. The intensity of PIV3's influence on gene expression increases substantially, surpassing that of IAV. Despite this, IAV still maintains its activity during this time frame, albeit to a lesser extent compared to PIV3. MPV, on the other hand, remains relatively less active even during the 24- to 72-hour post-infection period.
[0378] These findings highlight the dynamic nature of virus-induced gene expression patterns. IAV shows early and sustained activity, while PIV3 demonstrates an increase in activity over time, becoming highly intense during the later stage of infection. MPV exhibits comparatively lower activity throughout the analyzed time frames. Note that these observations are based on the analysis of GLMQL-RMAS scores, which provide a broader exploration of virus activity patterns and relationships across conditions. These scores offer valuable insights into the gene expression dynamics beyond differential expression alone and contribute to a comprehensive understanding of the interplay between viruses and gene expression.
[0379] We extended our analysis by fitting General Linear Models (GLMs) to explore the relationship between the GLMQL-RMAS scores within individual viruses at different time points. Specifically, we examined the pairs (IAV-RMAS-24, IAV-RMAS-72), (MPV-RMAS-24, MPV- RMAS-72), and (PIV3-RMAS-24, PIV3-RMAS-72). By fitting GLMs to these pairs, we estimated the slope and associated statistical measures, providing insights into the changes in virus intensity over time.Attorney Docket No.32103 / 59579 / PC
[0380] The analysis of linear correlation using General Linear Models (GLMs) and GLMQL- RMAS scores provides valuable insights into the relationship and interplay between different viruses at various time points. This approach allows us to explore how the intensity levels of these viruses are related to each other and how they influence gene expression patterns.
[0381] The GLM intercept and slope values for each pair comparison provide further insights into the relationship between virus intensities. For instance, a slope of 0.01 indicates a weak positive correlation between the intensity of IAV and MPV at the 24-hour time point. This suggests that as the intensity of IAV increases, there is a slight tendency for the intensity of MPV to also increase. Similarly, the slope of 0.06 between IAV and PIV3 at the same time point indicates a slightly stronger positive correlation between these two viruses. As we move to the 72-hour time point, the relationships between virus intensities become more pronounced. The slope of 0.08 between IAV and MPV suggests a slightly stronger positive correlation, indicating that both viruses exhibit increased activity. However, the slope of 0.96 between IAV and PIV3 at this time point indicates a much stronger positive correlation, suggesting that PIV3 becomes significantly more active and influences gene expression patterns more intensely compared to IAV. Furthermore, the slope of 6.40 between MPV and PIV3 indicates a substantial positive correlation, indicating a significant influence of PIV3 on gene expression patterns at the 72-hour time point.
[0382] In summary, at the 24-hour time point, IAV exhibits the highest activity, significantly impacting gene expression patterns. MPV shows relatively lower activity compared to IAV, while PIV3 demonstrates intermediate activity. Between the 24- and 72-hour time points, IAV maintains its activity level, MPV remains relatively less active, and PIV3 becomes significantly more active, exerting a stronger influence on gene expression patterns. These findings highlight the dynamic nature of viral activity over time, with IAV consistently being the most active virus and PIV3 displaying an intensified activity during the later time frame.
[0383] FIG 9C-9F provide further insights into the temporal activity of the three viruses, IAV, MPV, and PIV3, by utilizing the genes IFIT1, MX1, and EIF5. These figures allow the observation of the dynamics of gene expression in response to viral infections. Notably, a consistent pattern is observed where higher viral intensity is associated with increasedAttorney Docket No.32103 / 59579 / PC expression of IFIT1. This suggests the active involvement of IFIT1 in the immune response against the investigated viruses. Moreover, EIF5 serves as a reliable indicator of time, as its expression levels exhibit distinct patterns over different time points.
[0384] By examining these figures, a deeper understanding of the dynamic regulation of IFIT1, MX1, and EIF5 in the context of viral infections is gained. Overall, these findings highlight the significance of IFIT1 as a key player in the immune response against IAV and PIV3 infections, while also suggesting the influence of MPV on MX1 expression. Additionally, EIF5 serves as a valuable marker for tracking the progression of viral infections over time. By examining the interplay between these genes and viruses, we gain valuable insights into the dynamics and regulation of gene expression during viral infections.
[0385] The comprehensive analysis of gene expression dynamics using Generalized Linear Models (GLMs) and Quasi-Likelihood (QL) methods provides valuable insights into the interplay between viral infections and gene expression. The study investigates RNA-Seq data encompassing many (e.g., 19,671 or more) genes across various infection conditions, focusing on Influenza A virus (IAV), Human metapneumovirus (MPV), and Parainfluenza virus type 3 (PIV3). By employing the GLM-QL framework, differentially expressed genes are identified, shedding light on the underlying biological processes and molecular mechanisms associated with infections. To further enhance the analysis, two algorithms, namely GLMQL-MAS and GLMQL-RMAS, are employed. GLMQL-MAS enables the investigation of gene-virus associations, quantifying the strength of association between individual genes and specific viruses. This information provides valuable insights into potential markers for determining the time of infection and the cellular response against viral infections. On the other hand, GLMQL- RMAS relaxes the Benjamini-Hochberg adjustment for explanatory analysis, facilitating the measurement of virus intensity over time.
[0386] The results reveal distinct patterns in the gene expression of infected OTEs between 24- and 72-hours post-infection. All infected OTEs shift towards the positive axis of the first principal component, with PIV3-infected OTEs exhibiting the most significant movement compared to IAV- and MPV-infected OTEs. Within the 24-hour timeframe, IAV-infected OTEs show a larger shift compared to PIV3- and MPV-infected OTEs. These findings highlight theAttorney Docket No.32103 / 59579 / PC dynamic nature of virus-induced gene expression patterns. Furthermore, specific genes, including IFNL1, IFIT1, IFIT2, CXCL10, OASL, IFIT3, IFNL2, IFNL3, CXCL11, RSAD2, IFNB1, and IFITM1, are identified as exhibiting significant differential expression across different infection conditions and time points. These genes serve as potential markers for determining the time of infection and play a crucial role in the cellular response against viral infections.
[0387] Overall, the present techniques contribute to a comprehensive understanding of viral infections by elucidating the complex relationship between infection conditions and gene expression. The findings emphasize the importance of considering gene-virus associations and virus intensity over time. By uncovering the dynamic nature of virus-induced gene expression patterns, this research provides valuable insights into the underlying mechanisms involved in viral infections. These insights can potentially guide future research and aid in the development of effective strategies for combating viral diseases.
[0388] The present techniques showcase the advanced methodology of biomarker identification through differential expression analysis, underscored by the application of Generalized Linear Models with Quasi-Likelihood F-tests (GLMQL-RMAS). This rigorous approach allowed us to identify genes with significant differential expression due to viral infection, utilizing both raw p-values and adjusted significance levels for an accurate representation of biological differences.
[0389] The differential expression analysis, anchored in the GLMQL framework, pinpoints genes with notable differences in expression across experimental conditions. By emphasizing significant p-values and log fold changes (logFC), we unearth potential biomarkers indicative of specific viral infections. The RMAS algorithm refines gene prioritization by integrating the magnitude of expression changes with statistical significance, surpassing traditional methods that rely solely on p-values or logFC. The RMAS and MAS calculations, with hyperparameters M and A set to 1, emphasize the equal importance of expression magnitude and statistical robustness, promoting a balanced evaluation of each gene’s potential as a biomarker.
[0390] The selection of the top three genes through RMAS ranking, as illustrated in Figures 6, 7, demonstrates a remarkable 92% mean accuracy in distinguishing between eight different groups using a stratified k-fold cross-validation approach. This precision starkly contrasts withAttorney Docket No.32103 / 59579 / PC the lesser efficiency of prioritizing genes based solely on p-value or logFC, which achieved a mean accuracy of 85% each, as depicted in Figures 8, 9. This comparison underscores the superiority of integrating both p-value and logFC information in biomarker identification.
[0391] A significant aspect of the present techniques focuses on the requirement for an adequate sample size to ensure the reliability of biomarkers in prediction models. Adhering to the principle that at least 10 samples per variable (gene) are necessary, our approach is geared towards maximizing the predictive power and generalizability of the model. The utilization of multinomial logistic regression further enhances the model’s robustness, enabling it to generalize across different biological contexts and experimental conditions.
[0392] The employment of multinomial logistic regression is crucial in translating differential expression analysis into actionable insights. This statistical model accommodates the complexities of biological data, ensuring that identified biomarkers are not only statistically significant but also biologically relevant and capable of predicting specific viral infections accurately.
[0393] Our comprehensive approach, from differential expression analysis through to biomarker identification using MAS and RMAS, followed by validation via multinomial logistic regression, sets a new standard in the field. It highlights the importance of integrating statistical rigor with biological relevance, paving the way for the identification of robust biomarkers that can inform on viral infections with high precision. This methodology not only contributes to our understanding of viral pathogenesis but also opens avenues for the development of targeted diagnostic and therapeutic strategies, underscoring the potential of advanced statistical models in biomedical research.
[0394] The present techniques provide a comprehensive examination of gene expression dynamics in response to viral infections, focusing specifically on Influenza A virus (IAV), Human metapneumovirus (MPV), and Parainfluenza virus type 3 (PIV3). Utilizing Generalized Linear Models (GLMs) with Quasi-Likelihood (QL) F-tests (GLMQL) to analyze RNA-Seq data from 19,671 genes, our aim was to identify genes differentially expressed under various infection conditions. The application of GLMQL, along with the innovative Magnitude-Altitude Score (MAS) and Relaxed Magnitude-Altitude Score (RMAS) algorithms, facilitated the preciseAttorney Docket No.32103 / 59579 / PC identification of potential biomarkers indicative of specific viral infections. Our Gene Ontology (GO) analysis further enriched our understanding of the host’s defense mechanisms, identifying key biological processes activated in response to these viral infections. This analysis consistently highlighted the activation of interferon-stimulated genes across all three viruses, underlining the host’s reliance on innate immune mechanisms to combat viral threats. Moreover, the analysis of downregulated genes unveiled viral strategies aimed at manipulating host cellular functions, emphasizing the need for targeted therapeutic interventions.
[0395] The GO analysis offered deep insights into the host’s comprehensive response to viral infection, pinpointing critical biological processes, cellular components, and molecular functions impacted by IAV, MPV, and PIV3. Key findings included the activation of a broad range of interferon-stimulated genes (e.g., IFIT1, IFIT2, IFIT3, OAS1) and innate immune response genes (e.g., CGAS, DDX60), showcasing the host’s robust defense mechanisms against viral entry, replication, and spread. Notably, the analysis also revealed significant changes in cellular functions and structures, such as cilium assembly and mitochondrial ribosome assembly, indicating a strategic shift in cellular priorities to support viral replication. The stability of the MAS / RMAS ranking method, even under stringent statistical corrections, underscores the reliability of our approach in identifying key biomarkers.
[0396] The present implementation of the GLMQL-RMAS methodology enabled the precise identification of genes with significant differential expression, setting the stage for biomarker identification. This rigorous statistical framework, combined with the strategic use of the MAS and RMAS algorithms, allowed for the nuanced prioritization of genes based on both the magnitude of expression changes and their statistical significance. The classification of respiratory virus infections in OTEs, leveraging a multinomial logistic regression model validated through stratified k-fold cross-validation, achieved remarkable accuracy. This method not only demonstrated a 92% mean accuracy in distinguishing between different viral infections but also underscored the superior efficacy of integrating both p-value and logFC information over traditional methods. Crucially, our study highlights the importance of having an adequate sample size for each variable (gene) used as a predictor in the model, ensuring the reliability and generalizability of our findings.Attorney Docket No.32103 / 59579 / PC Further Exemplary Computer-Implemented Method Aspects
[0397] FIG.12 depicts a block diagram of a computer-implemented method 1200 for computing magnitude-altitude scores corresponding to genes expressing in various ways during one or more time periods, according to some aspects.
[0398] The method 1200 may include receiving, via one or more processors, statistical results including one or more raw p-values that (a) are indicative of a differential gene expression in the organ tissue equivalent under varying infection conditions, from an initial condition to a subsequent condition, over a period of time, and (b) describe significance of gene expression differences from the initial condition to the subsequent condition (block 1202). The method 1200 may include generating, via one or more processors, a respective relaxed magnitude-altitude score for one or more genes identified as significant in the statistical results, by computing pairwise comparisons of general linear model quasi likelihood relaxed magnitude altitude scores (GLMQL-RMAS) across diverse conditions (block 1204). The method 1200 may include block displaying, via one or more processors, at least some of the GLMQL-RMAS scores for the one or more genes on a display (block 1206).
[0399] In some aspects, the organ tissue equivalent is a lung organoid. In some aspects, the initial condition and / or the subsequent condition corresponds to infected, ultraviolet treated, non- ultraviolet treated, mock, or naive. In some aspects, the period of time corresponds to 24 hours or 72 hours. In some aspects, the organ tissue equivalent is infected by one or more pathogens. In some aspects, the one or more pathogens are respective viruses. In some aspects, the respective viruses include at least one of influenza A virus, human metapneumovirus, or parainfluenza virus type 3. In some aspects, receiving the statistical results includes receiving RNA-Seq data.
[0400] In some aspects, the method 1200 may further include determining, via one or more processors, unique statistically significant genes affected by respective viruses during organ tissue equivalent condition transitions, by: for each of the respective viruses, repeating step 2 of claim 1 and step 3 of claim 1 to generate a significance score set that includes respective significant genes and respective ranked relaxed magnitude-altitude scores; and for each of theAttorney Docket No.32103 / 59579 / PC respective viruses, generating a set of unique significant genes including respective differentially expressed genes.
[0401] In some aspects, the method 1200 may include generating, via one or more processors, a visualization of the set of unique significant genes including respective differentially expressed genes. In some aspects, the method 1200 may include determining, via one or more processors, common statistically significant genes affected by all viruses during organ tissue equivalent condition transitions, by: for each of the respective viruses, repeating step 2 of claim 1 and step 3 of claim 1 to generate a significance score set that includes respective significant genes and respective ranked relaxed magnitude-altitude scores; and for each respective virus, generating a set of common significant genes including respective differentially expressed genes. In some aspects, the method 1200 may include generating, via one or more processors, a visualization of the set of common significant genes including respective differentially expressed genes. In some aspects, the method 1200 may include generating, via one or more processors, a visualization of adjusted statistical results.
[0402] In some aspects, the method 1200 may include generating, via one or more processors, a visualization of the respective magnitude-altitude score for the one or more genes; and based on the visualization, tuning hyperparameters M and A.
[0403] The MAS (Magnitude Altitude Score) algorithm can be extended and applied in various ways. For example, as discussed MAS has been utilized not only in analyzing RNA-seq data but also in NanoString data, demonstrating its versatility across different data modalities. Furthermore, the algorithm has been adapted to explore biomarkers in cancer data, leveraging the TCGA cancer data set for comprehensive analysis. This highlights its potential in oncological research and personalized medicine. Additionally, the development of the MASIT algorithm (discussed infra) represents a significant advancement, incorporating time domain analysis and multi-omics data to enhance predictive power and applicability. This extension allows for a more nuanced understanding of biological processes over time and across various levels of biological information.
[0404] The algorithm’s validation across diverse datasets, including those related to infectious diseases such as Ebola, Dengue fever, and Mpox or other Orthopoxviruses,Attorney Docket No.32103 / 59579 / PC underscores its broad applicability and potential in addressing global health challenges. Through these extensions and applications, the MAS algorithm and its derivatives offer powerful tools for biomedical research, with implications for disease diagnosis, prognosis, and treatment strategies.
[0405] The MAS algorithm is generally an unsupervised algorithm. The present techniques may include further techniques (e.g., MASIT), that can include supervised training on one dataset (e.g., NanoString) and inference on data of a different modality (e.g., RNASeq data). Another potential distinction between MAS and MASIT is that MASIT may include temporal dependencies. For example, FIG.13 depicts a block-flow diagram of a computer-implemented method 1300 for training, validating, and testing a machine learning (ML) model using multi- scale gene expression data, according to some aspects. The method 1300 may include a Magnitude-Altitude Score Analysis for Tracking Infection and Time-Dependent Genes algorithm (MASIT) Training-Validation step (block 1302). At block 1302, the method 1300 may include collecting molecular biology data (e.g., NanoString data) over various time points (Time T1, Time T2, ..., Time Tn), both for baseline and treated conditions, with the subscript 'n' indicating different samples or conditions. The method 1300 may further include the application of cross- validation techniques (block 1304), for example, using a validation K-Fold Stratified Cross Validation approach for model training, where the dataset is split into K subsets. In each cycle k (k=1, ..., K), a different subset is used as the validation set, and the remaining subsets are combined to form the training set. During each cycle, MASIT may be applied to identify a set of genes, which consists of genes related to infection and time specific for that fold (block 1306). The method 1300 may include selecting the optimal ML classifier based on the identified genes and its performance evaluated on the hold-out validation set for that cycle (block 1308).
[0406] The method 1300 may further include holding out data to form a test set (block 1310). In this case, RNA-Seq data may be similarly collected over various time points. The method 1300 may include applying the ML classifiers that were previously selected using the selected genes from the training-validation phase at block 1302 to the RNA-Seq dataset (block 1312). The method 1300 may include testing the performance of the ML classifiers using the same validation technique (e.g., K-Fold Stratified Cross Validation), considering different scales in theAttorney Docket No.32103 / 59579 / PC dataset (block 1312). The method 1300 may include constructing a set of genes that includes genes identified by a majority within the k-fold cross-validation as being either infection or time signature genes (block 1314). After training and validation, the method 1300 may include testing the model using an independent RNA-Seq dataset to evaluate its generalizability and performance. In general, the method 1300 may be used in bioinformatics to create predictive models for biological data, such as gene expression profiles, and to assess the impact of treatments over time. The goal of this methodology may be to identify signature genes that can be used as biomarkers and / or to understand the underlying mechanisms of response to treatment.
[0407] The present techniques may include transfer learning techniques, in some aspects, wherein one or more models are trained on one modality and applied to another modality. By training on NanoString to generate trained models, researchers can leverage the high specificity and sensitivity of NanoString data to inform and refine models that can then be applied to the more widely available and less expensive RNASeq data. This approach allows for the identification of biomarkers and the understanding of gene expression dynamics across different conditions and diseases, including infectious diseases like Ebola and Dengue fever, as well as in cancer research. The methodology, known as MAS (Magnitude Altitude Score), serves as the foundation for this cross-modality application, enabling the prediction and analysis of gene expression with high accuracy. This innovative approach not only maximizes the utility of high- quality NanoString data but also enhances the applicability and relevance of RNASeq data in biomedical research. The method 1300 may be performed by the components depicted in FIG. 1; for example, the server computing device 102 may include instructions in the modules 150 (e.g., the magnitude-altitude module 158) may include instructions for training, validating, operating and testing the ML models and classifiers described herein.
[0408] RNASeq techniques are widely available, albeit of lesser accuracy compared to NanoString or similar technology. RNASeq play a crucial role in the equation by offering a cost- effective and accessible means to apply the insights gained from the more expensive NanoString data across a broader range of samples and conditions. By training models on NanoString data, researchers can identify key biomarkers and gene expression patterns withAttorney Docket No.32103 / 59579 / PC high specificity and sensitivity. These models can then be applied to RNASeq data, which is more widely available and less costly, thereby extending the reach and applicability of the findings. This cross-modality application leverages the strengths of both technologies: the high- quality data from NanoString for model training and the broad accessibility of RNASeq for widespread application. This approach not only enhances the value of RNASeq data by providing a method to interpret it with greater accuracy but also democratizes access to high- quality genomic insights, making it possible to conduct extensive and inclusive research across various biological and clinical settings.
[0409] In the context of the invention, the difference in the number of genes that can be analyzed by NanoString and RNASeq technologies is significant. NanoString technology is limited to analyzing a specific set of genes, typically up to around 950 genes, due to its design and the nature of the probes used in the technology. This limitation means that when models are trained using NanoString data, they focus on a relatively small, predefined set of genes. However, these genes are analyzed with high specificity and sensitivity, making NanoString data particularly valuable for identifying key biomarkers and understanding intricate gene expression dynamics within those constraints. On the other hand, RNASeq technology does not have the same limitations regarding the number of genes it can analyze. It is capable of analyzing the expression of tens of thousands of genes across the entire genome. This wide coverage makes RNASeq a powerful tool for exploratory research and for studies where the full breadth of genomic expression needs to be considered. However, RNASeq is generally considered to be less accurate per gene compared to NanoString, particularly in terms of quantification and detection of low-abundance transcripts. The present techniques strengths of both technologies by using the high-quality, focused analysis provided by NanoString to inform models that can then be applied to the broader, more comprehensive datasets generated by RNASeq. This approach allows for the identification of biomarkers and gene expression patterns with the precision and reliability of NanoString data, while also taking advantage of the wide coverage and accessibility of RNASeq data. Essentially, the invention bridges the gap between the depth and accuracy of NanoString's analysis of a limited number of genes and the breadth and scalability of RNASeq's genome-wide analysis. This synergy enhances the predictive power and applicability of the models developed, making it possible to extend high-Attorney Docket No.32103 / 59579 / PC quality genomic insights across a much wider array of genes and applications than would be possible using either technology alone.
[0410] K-fold cross-validation plays a role in the method 1300, by enhancing the robustness and generalizability of the models developed from NanoString data when applied to RNASeq data. The method 1300 may divide the dataset into 'k' number of subsets (or folds). In the context of this invention, the NanoString dataset may be divided in such a manner. For each iteration of the model training process, one of these subsets is held out as the test set, while the remaining k-1 subsets are used as the training set. This process is repeated 'k' times, with each subset serving as the test set exactly once. This approach ensures that the model is trained and validated on different segments of the data, enhancing its ability to generalize well to new, unseen data. By training the model on different portions of the NanoString data and validating it on the held-out portions, the method 1300 can assess how well the model is likely to perform on new data, which ensures that the insights and patterns the model learns are not specific to the particularities of the training data but are generalizable across different datasets, including RNASeq data. K-fold cross-validation allows for the fine-tuning of model parameters to optimize performance. By evaluating the model's performance across different folds, researchers can identify the best set of parameters that lead to the most reliable and accurate predictions. This optimization facilitates applying the model to RNASeq data, as it ensures that the model's predictions are as accurate and reliable as possible. Training a model on a limited set of high- quality NanoString data poses a risk of overfitting, where the model learns the noise in the training data instead of the underlying patterns. K-fold cross-validation mitigates this risk by ensuring that the model's performance is tested on unseen data throughout the training process. This helps in developing a model that is robust and performs well not just on the NanoString data it was trained on but also on broader RNASeq datasets.
[0411] The number of genes that NanoString and RNASeq technologies can analyze, respectively, also plays an important role in the decision to train models on NanoString data for application to RNASeq data, instead of the inverse. NanoString technology is designed to analyze a specific, predefined set of genes, typically up to around 950 genes. This limitation means that the models trained using NanoString data are highly focused on a relatively smallAttorney Docket No.32103 / 59579 / PC set of genes, allowing for deep insights into the specific gene expression dynamics within that set. The high specificity and sensitivity of NanoString technology ensure that the data for these genes is of high quality, making it ideal for training robust and accurate models. On the other hand, RNASeq technology is capable of analyzing the expression of tens of thousands of genes across the entire genome. While this broad coverage is a significant advantage for exploratory research and comprehensive analysis, it also introduces challenges in terms of data volume, variability, and the potential for noise. Training models directly on RNASeq data would require navigating these challenges and could potentially dilute the focus on specific genes of interest due to the vast amount of additional data. By training models on the focused and high-quality NanoString data, the present techniques aim to capture detailed insights into the expression dynamics of a targeted set of genes. These models can then be applied to the broader RNASeq data, extending the precision and accuracy of NanoString to the comprehensive analysis possible with RNASeq. This strategy effectively leverages the strengths of NanoString technology to enhance the utility and interpretability of RNASeq data, making it possible to conduct high-quality genomic analysis across a much wider array of genes than would be possible using NanoString data alone. Comparing NanoString and RNA-Seq
[0412] The present techniques include a comparative analysis of gene expression data across RNA-Seq and NanoString platforms. While RNA-Seq covered 19,671 genes and NanoString targeted 773 genes associated with immune responses to viruses, the present techniques’ primary focus was on the 754 genes found in both platforms. An experiment involved 16 different infection conditions, with samples derived from 3D airway organ-tissue equivalents subjected to three virus types, influenza A virus (IAV), human metapneumovirus (MPV), and parainfluenza virus 3 (PIV3). Post-infection measurements, after UV (inactive virus) and Non-UV (active virus) treatments, were recorded at 24-h and 72-h intervals. Including untreated and Mock-infected OTEs as control groups enabled differentiating changes induced by the virus from those arising due to procedural elements. Through a series of methodological approaches (including Spearman correlation, Distance correlation, Bland-Altman analysis, Generalized Linear Models Huber regression, the Magnitude-Altitude Score (MAS) algorithmAttorney Docket No.32103 / 59579 / PC and Gene Ontology analysis) a study meticulously contrasted RNA-Seq and NanoString datasets. The Magnitude-Altitude Score algorithm, which integrates both the amplitude of gene expression changes (magnitude) and their statistical relevance (altitude), offers a comprehensive tool for prioritizing genes based on their differential expression profiles in specific viral infection conditions. Empirically, a strong congruence between the platforms was observed, especially in identifying key antiviral defense genes. Both platforms consistently highlighted genes including ISG15, MX1, RSAD2, and members of the OAS family (OAS1, OAS2, OAS3). The IFIT proteins (IFIT1, IFIT2, IFIT3) were emphasized for their crucial role in counteracting viral replication by both platforms. Additionally, CXCL10 and CXCL11 were pinpointed, shedding light on the organ tissue equivalent’s innate immune response to viral infections. While both platforms provided invaluable insights into the genetic landscape of organoids under viral infection, the NanoString platform often presented a more detailed picture in situations where RNA-Seq signals were more subtle. The combined data from both platforms emphasize their joint value in advancing our understanding of viral impacts on lung organoids.
[0413] The human airways are a nexus of intricate relationships, intricately binding cellular interactions, extracellular matrix (ECM) proteins, and the biomechanical milieu. At the forefront of replicating these complexities, the present techniques include a 3D airway organ tissue equivalent (OTE) model functioning at an air-liquid interface (ALI). Incorporating native pulmonary fibroblasts, solubilized lung ECM, and a tunable hydrogel substrate, this model stands as a revolutionary contribution to airway biology research. By evaluating the influence of our model on the phenotype of human bronchial epithelial (HBE) cells over a 28-day ALI culture duration, the present techniques noted its pronounced ability in nurturing well-differentiated ALI cultures. These cultures notably manifest barrier functionality and mature epithelial marker expression. A unique feature of the present model is the adjustable stiffness of the hydrogel, offering potential avenues for further phenotype modulation research. The present techniques build on foundational methodologies, emphasizing the versatility and precision of the 3D airway OTE model in simulating the multifarious dimensions of the human airway’s 3D microenvironment (Leach et al., 2023).Attorney Docket No.32103 / 59579 / PC
[0414] To enhance understanding of the 3D airway OTE model’s response to viral infections, the present techniques employed RNA-Seq (Wang et al., 2009) and NanoString (Geiss et al., 2008) technologies to dissect the complex virus-host interactions at the molecular level. The present techniques focused on samples collected at two critical time points, 24- and 72-h post- infection, to capture both the immediate and prolonged cellular responses to viral invasion. This strategy aims to provide a comprehensive view of the dynamic interactions between host cells and infecting viruses during these pivotal infection phases.
[0415] As above, RNA-Seq data encompassing many (e.g., 19,671) genes may be analyzed to explore gene expression dynamics following infection with both active and UV-inactivated viruses: Influenza A virus (IAV), Human metapneumovirus (MPV), and Parainfluenza virus type 3 (PIV3). We employed two algorithms, GLMQL-MAS and GLMQL-Relaxed-MAS, which integrate Generalized Linear Models (GLM), Quasi-Likelihood (QL) F-tests, and the Magnitude- Altitude Score (MAS). These methodologies robustly identified key differentially expressed genes, particularly those involved in interferon signaling pathways such as IFIT1, IFIT2, IFIT3, and OAS1, which play crucial roles in the innate immune response.
[0416] The present techniques may include a comparative analysis to demonstrate the consistency of gene selection between the RNA-Seq and NanoString platforms, with the aim of validating the reproducibility of gene expression data across these technologies. Objectives may be as follows: Correlation analysis: Assess the consistency between RNA-Seq and NanoString data across multiple infection conditions using Spearman and Distance correlation metrics, providing a comprehensive evaluation of the agreement between these two platforms. Bland-Altman analysis: This analysis aids in visualizing the level of concordance between RNA- Seq and NanoString measurements, highlighting any systematic biases or discrepancies to enhance our understanding of each platform’s reliability. Generalized linear model and huber regression analysis: By employing these robust statistical tools, aim to evaluate the relationship between RNA-Seq and NanoString data across diverseAttorney Docket No.32103 / 59579 / PC infection conditions, ensuring that our findings are resilient against potential data outliers and deviations. Concordance analysis: Utilizing expression analysis and the GLMQL-MAS algorithm, identify biologically meaningful changes in gene expression, ensuring that the significant genes detected remain consistent across both technological platforms. Gene ontology (GO) analysis for common BH-significant transcripts across two platforms: This analysis confirms the biological relevance of the significant changes detected by both platforms, enriching understanding of the molecular mechanisms underlying the response to viral infections.
[0417] FIG.14A displays a schematic overview of the main objectives of the present comparative analysis between RNA-Seq and NanoString platforms. By addressing these objectives, the present techniques not only aims to validate the agreement between RNA-Seq and NanoString technologies but also to enhance the biological insights derived from the 3D airway OTE model.
[0418] The virus infection medium (iDMEM) was concocted using Dulbecco-modified Eagle’s minimal essential medium, supplemented with 0.1% heat-inactivated fetal bovine serum, 0.3% purified bovine serum albumin, 20 mM HEPES [pH 7.5], and 0.2 mM Glutamax. The concoction was subsequently filter-sterilized through a 0.2 µm filter and preserved at 4CC. Titers of the Influenza virus were ascertained using Madin-Darby Canine Kidney (MDCK) cells (ATCC, #CCL-34), while the titers for all other viruses were determined using the LLC-MK2 rhesus monkey kidney cell line (ATCC, #CCL-185). Each lung OTE was calculated to encompass approximately 1.6 × 105 epithelial cells. The nominal multiplicity of infection was calculated based on the premise that solely the epithelial cells were vulnerable to the viral infections. A 24- well culture plate filled with modified PneumaCult ALI Medium (Stemcell Technologies) was permitted to reach equilibrium at 37CC in a humidified incubator with 5% CO2 for 30 min. Inserts holding the OTEs were sterily relocated to the balanced culture plate. The surface of each OTE slated for infection or Mock infection was rinsed by delicately adding 0.3 mL of warm Hank’s balanced saline solution, followed by cautious aspiration. The virus, diluted in iDMEM and brought to ambient temperature, was then dispensed in a 40 µL volume onto the apicalAttorney Docket No.32103 / 59579 / PC surface of the OTE. Subsequently, the plate harboring the infected OTEs was situated on a rocking table in the 37CC CO2 incubator for 1 hour before being retrieved from the rocking table and restored to standard growth conditions for the designated durations.
[0419] The RNA extraction from the OTEs was meticulously performed using the Direct-zol RNA Miniprep Plus Kit. This protocol included sample preparation, cell lysis, RNA purification, and DNase I treatment to eliminate potential DNA contaminants. We stored the purified RNA at −80CC to preserve it for subsequent in-depth sequencing and gene expression analyses. These steps were critical for accurately dissecting the complex interplay between the OTE model and the introduced viral pathogens, ensuring the integrity of the samples for further molecular analysis.
[0420] Following RNA extraction, we conducted RNA sequencing to delve deeper into the molecular underpinnings of virus-host interactions. We crafted cDNA libraries from 50 ng of the extracted RNA using the NEXTFLEX® Combo-SeqTM mRNA / miRNA Kit. The processing was performed on a Sciclone® G3 NGSx Workstation, with libraries quantified using a KAPA Library Quantification Kit and evaluated for average fragment size using a 4,200 TapeStation System. After library normalization, we performed high-throughput sequencing on an Illumina® NovaSeq 6000 System, producing 76-bp single-end reads. This setup laid the groundwork for comprehensive sequence analysis.
[0421] For data analysis, we utilized the Partek® Flow® software. The raw sequence data underwent meticulous processing that included adaptor trimming and quality base filtering with Cutadapt (Martin, 2011), alignment against the hg38 GENCODE reference database using the STAR algorithm, and transcript quantification using an expectation / maximization (E / M) algorithm (Xing et al., 2006). We normalized transcript-level counts to gene-level data employing the median-of-ratios method from DESeq2 (Love et al., 2014), with results log2-transformed for enhanced clarity.
[0422] Additionally, we leveraged the NanoString nCounter® Analysis System to perform highly multiplexed detection of mRNA targets relevant to our study of viral infections. This technology is particularly suited for samples like those derived from our OTE model where RNA integrity is variable, as it does not rely on amplification or fluorescence intensity for targetAttorney Docket No.32103 / 59579 / PC detection. Instead, detection is based on barcoded sample processing, which yields precise and reproducible gene expression data. We ensured data quality at every analysis stage, beginning with general assay performance and followed by background correction, data normalization, and comprehensive quality control checks on the resultant expression metrics.
[0423] Two distinct normalization methods were applied to the RNA-Seq and NanoString data. The RNA-Seq data were normalized using the Trimmed Mean of M-values (TMM) method. This method scales the library sizes by a normalization factor, thus allowing for more accurate comparisons by mitigating the influence of highly expressed genes. For the NanoString data, a two-step normalization process was used. Firstly, a Positive Control Normalization factor was calculated using the positive controls added to each sample. This step helps to adjust for technical variations across samples, lanes, cartridges, and different days of experimentation. Secondly, a CodeSet Content Normalization factor was computed using housekeeping genes.
[0424] In our study, special emphasis is placed on Non-UV (active) samples, where active viral infections are facilitated, allowing the viruses to replicate and dynamically interact with the host cells within the OTEs. This condition is pivotal as it most accurately simulates the natural infection environment, providing critical insights into the host’s cellular and molecular responses under active viral attack. To ensure the accuracy of our findings and clearly delineate the effects of viral infections, our study utilized a robust set of control conditions, including Mock, UV- treated, and naïve (untreated) samples. UV-treated samples, exposed to ultraviolet light to inactivate the viruses, served as crucial controls for examining the impact of viral components without active replication. Naïve samples, which are untreated OTEs, acted as internal controls to set a baseline for gene expression across our experiments. Mock-infected samples, treated with a vehicle or sham procedure, offered comparative data to underscore the specific gene expression responses triggered by active viral infections in the Non-UV (active) samples, where viruses capable of replication were used.
[0425] Our analyses covered 773 immune response genes identified in the NanoString dataset and a broader spectrum of 19,671 genes covered by RNA-Seq. These genes were evaluated across 16 distinct infection conditions, each involving six replicates of OTEs, as detailed in Table T1. The conditions were categorized by virus type (IAV, MPV, or PIV3),Attorney Docket No.32103 / 59579 / PC treatment type (UV or Non-UV / active / None), and post-infection times (24-h and 72-h post infection), which is denoted as Virus-Treatment-Time. This comprehensive setup allowed us to rigorously test and validate the biological significance and reproducibility of our data across different experimental and control conditions, as detailed in Table T2, which illustrates the distribution of data across these groups. From the available data, we identified 754 genes that were common to both the RNA-Seq and NanoString platforms, facilitating a consistent and comparative analysis of gene expression dynamics in response to viral challenges, particularly focusing on the critical role of Non-UV (active) samples in our study.
[0426] Table T1 displays 16 different infection conditions that are classified based on (1) Virus, (2) Treat, and (3) post-infection time. In this study, OTEs are infected by three viruses: IAV, MPV, and PIV3; there are two kinds of treatments: UV and Non-UV (or active); and there are two post-infection time points, 24-hours, and 72-hours. Table T2 describes the distribution of data points between post-infection time points, experimental groups, negative control groups, and control groups.
[0427] Table T1. Sixteen infection conditions classified based on (1) Virus, (2) Treat, and (3) post-infection time. Condition Virus Treat Post- Condition Virus Treat Post- infection infection Time Time IAV-UV-24 IAV UV 24-hours IAV-UV-72 IAV UV 72-hours IAV-None-24 IAV Non-UV 24-hours IAV-None-72 IAV Non-UV 72-hours MPV-UV-24 MPV UV 24-hours MPV-UV-72 MPV UV 72-hours MPV-None-24 MPV Non-UV 24-hours MPV-None-72 MPV Non-UV 72-hours PIV3-UV-24PIV3UV 24-hours PIV3-UV-72PIV3UV 72-hoursPIV3-None-24 PIV3 Non-UV 24-hours PIV3-None-72 PIV3 Non-UV 72-hours Mock-24 Mock - 24-hours Mock-72 Mock - 72-hours Naïve-24Untreated -24-hours Naïve-72Untreated -72-hoursAttorney Docket No.32103 / 59579 / PC
[0428] Table T2. Distribution of replicates between post-infection timepoints, experimental groups, negative and regular control groups. At post- 6 replicates for each of the followings: infection Experimental Group: Non-UV treated IAV, Non-UV treated MPV, Non-UV treated PIV3 time pointNegative Control Group: UV treated IAV, UV treated MPV, UV treated PIV324-hoursControl Group: Mock, NaïveAt post- 6 replicates for each of the followings: infection Experimental Group: Non-UV treated IAV, Non-UV treated MPV, Non-UV treated PIV3 time point Negative Control Group: UV treated IAV, UV treated MPV, UV treated PIV3 72-hoursControl Group: Mock, Naïve
[0429] Table T3. Top 20 Significant GO Biological Processes of 2023 for Upregulated, Significant Common Genes Between RNA-Seq and NanoString When Comparing IAV-None-24 against Mock-24. ID Description qvalue Count Gene Symbols GO:0009615 response to virus 8.16E-53 45 IFIT2, RSAD2, IFIT3, IFNL1, IFIT1, IFNB1, ISG15, OAS2, OASL, MX1, HERC5, OAS1, OAS3, IFI44, DDX58, IFIH1, CXCL10, CCL5, IFI6, IFITM1, EIF2AK2, IFI16, STAT1, IRF1, ZBP1, GBP1, TNF, DTX3L, IFITM2, SAMHD1,GO:0051607 defense response to3.47E-52 41 IFIT2, RSAD2, IFIT3, IFNLA1P,OIFBITE1C,3IFGN, ABI1M,2IS,G15, OAS2, OASL, MX1, virus HERC5, OAS1, OAS3, DDX58, IFIH1, CXCL10, IFI6, IFITM1, EIF2AK2, IFI16, STAT1, IRF1,GO:0140546 defense response to 3.47E-52 41 IFZITB2P,1R,SGABDP21,,IDFITTX33,LIF,NIFLIT1,MIF2,ITS1A,MIFHNDB1,,AISPGO1B5E,CO3AGS,2A,IOMA2S,LT,LMRX3,1,symbiont HERC5, OAS1, OAS3, DDX58, IFIH1, CXCL10, IFI6, IFITM1, EIF2AK2, IFI16, STAT1, IRF1,GO:0048525 negative regulation of 9.99E-33 23ZBRPS1A,DG2B,PIF1IT, D1,TIXF3NLB,1IF, IISTMG125,,SOAAMSH2D,1O, ASPLO,BMEXC13,GO,AASI1M,2O,ATSLR3,3,viral process IFIH1, CCL5, TRIM21, IFITM1, EIF2AK2, IFI16, STAT1, TNF, IFITM2, APOBEC3G, GO:0050792 regulation of viral 4.87E-30 25 RSAD2, IFIT1, IFNB1, ISG15T,ROIAMS52,, OASL, MX1, OAS1, OAS3, process IFIH1, CCL5, LAMP3, TRIM21, IFITM1, EIF2AK2, IFI16, STAT1, TNF, IFITM2, GO:0045071 negative regulation of 1.00E-29 19 RSAD2, IFIT1, IFNB1, ISGA1P5O, OBEACS23,GO,ASL, MX1, OAS1, OAS3, viral genome IFIH1, CCL5, replication IFITM1, EIF2AK2, IFI16, TNF, IFITM2, APOBEC3G, IFITM3, BST2 GO:1903900 regulation of viral 1.47E-29 24 RSAD2, IFIT1, IFNB1, ISG15, OAS2, OASL, MX1, OAS1, OAS3, life cycle IFIH1, CCL5, LAMP3, TRIM21, IFITM1, EIF2AK2, IFI16, TNF, IFITM2, APOBEC3G, GO:0019221 cytokine-mediated 1.67E-29 33 IFNB1, ISG15, OAS2, OASL,TMRXIM1,5O,AS1, OAS3, CXCL10, CCL5, signaling pathway IFITM1, CXCL11, STAT1, TNFSF13B, IRF1, ZBP1, TNF, SOCS1, SP100, IFITM2, SAMHD1,GO:0002831 regulation of1.55E-28 29 IFAITIM1,2I,FPNABR1P,9IS,GIF1IT5,MO3A, ISFLI2,7H,ECRACS5P,1O,AJASK1,2O, MASY3D,8D8D, FXA5S8,,CILC15L,5, response to biotic CD274, TRIM21, stimulus IFI35, IFI16, STAT1, IRF1, ZBP1, DTX3L, SOCS1, LAG3, SAMHD1,GO:0045069regulation of viral5.51E-26 19RSAD2, IFIT1, IFNB1, ISGA1P5O, OBEACS23,GO,ASL, MX1, OAS1, OAS3, genome replication IFIH1, CCL5,GO:0019079viral genome6.72E-24 20IFRISTMAD1,2E, IIFI2TA1K,2IF,NIFBI16,,ISTGN1F5,,IFOIATMS2,,OAPASOLB,EMCX31G,,OIFAIST1M,3O,ABS3T,2 replication IFIH1, CCL5, IFITM1, EIF2AK2, IFI16, TNF, IFITM2, APOBEC3G, IFITM3, IFI27,Attorney Docket No.32103 / 59579 / PC GO:0034340 response to type I 2.31E-23 16 IFIT1, IFNB1, ISG15, OAS2, MX1, OAS1, OAS3, IFITM1, STAT1, interferon ZBP1, SP100, GO:0019viral life cycle 4.4 2 RS IAFDIT2M,2 IF,I STA1,M IHFNDB1,1, IF ISITGM135,, I OFIA2S7,2 M, OYADS8L8, MX1, OAS1, 058 6E- 5 OAS3, IFIH1, CCL5, 23 LAMP3, TRIM21, IFITM1, EIF2AK2, IFI16, TNF, IFITM2, APOBEC3G, TRIM5,Attorney Docket No.32103 / 59579 / PC GO:001 viral process 6.5 2 RSAD2, IFIT1, IFNB1, ISG15, OAS2, OASL, MX1, 6032 6E 7 OAS1, OAS3, IFIH1, CCL5, -23 LAMP3, TRIM21, IFITM1, EIF2AK2, IFI16, GO:007 cellular 2.9 1 IFIT1, IFNB1, ISG15, OAS2, OAS1, OAS3, IFITM1, 1357 response to 8E 5 STAT1, ZBP1, SP100, IFITM2, GO:006type I - 12.62 1 IFNB S1A, IMSGH1D51,, O IAFIST2M, O3,A ISF1I2,7 O,A MSY3,D I8F8ITM1, 0337 interferon 4E 4 STAT1, ZBP1, SP100, IFITM2, GO:003re ssigpnoanlsieng to - 42.50 CCL51, GBP4, TRIM21, I SFAITMMH1D, S1T, IAFTIT1M, I3R,F I1F,I2 G7B,P M1Y, SDO8C8S1, 4341 interferon- 9E 8 SP100, IFITM2, GBP5, TLR3, PARP9, IFITM3, GO:003 response to - 22.70 1 IFNB1, OAS1, IFITM1, IFI16, STAT1, IRF1, XAF1, 5456 interferon- 6E 2 IFITM2, AIM2, TLR3, IFITM3, GO:000positive - 41.39 2 RSAD2, IFNL1, ISGB1S5T, O2AS2, OAS1, OAS3, 1819 regulation of 4E 5 DDX58, IFIH1, CD274, EIF2AK2, cytokine -19 IFI16, STAT1, IRF1, TNF, GBP5, AIM2, TLR3, GO:004regulation of 4.2 1 IFNB1, ISG I1L51,2 OAA,S C1G,A OSA,S C3A, CSPC1L,5 J,A TKR2I,M21, IFI35, 5088 innate 0E 9 IFI16, IRF1, ZBP1, SOCS1, immune -18LAG3, SAMHD1, GBP5, AIM2, TRIM5, PARP9,
[0430] Table T4. Top 20 Significant GO Biological Processes of 2023 for Upregulated, Significant Common Genes Between RNA-Seq and NanoString When Comparing PIV3-None-24 against Mock-24. ID Description qvalue Count Gene Symbols GO:0051607 defense response to virus 1.63E-23 14 IFIT1, CXCL10, IFIT3, IFIT2, MX1, RSAD2, HERC5, OAS2, ISG15, GO:0140546 defense response to1.63E-23 14 IFIT1, CXOCAL1S03,,IDFIDTX35,8IF,IETI2F,2MAXK12,,ROSAASDL,2P, HAERRPC95, OAS2, symbiont ISG15, GO:0009615 response to virus 1.14E-21 14 IFIT1, CXOCAL1S03,,IDFIDTX35,8IF,IETI2F,2MAXK12,,ROSAASDL,2P, HAERRPC95, OAS2,ISG15, GO:0045071 negative regulation of viral 1.44E-15 8 IFIT1, MXO1,ARSS3,ADD2,XO58A,SE2I,FI2SAGK125,,OOAASSL3,,PEAIFR2PA9K2, OASLgenome replication GO:0045069 regulation of viral genome 3.87E-14 8 IFIT1, MX1, RSAD2, OAS2, ISG15, OAS3, EIF2AK2, OASL replication GO:0048525 negative regulation of viral 6.23E-14 8 IFIT1, MX1, RSAD2, OAS2, ISG15, OAS3, EIF2AK2, OASL process GO:0019079 viral genome replication 9.80E-13 8 IFIT1, MX1, RSAD2, OAS2, ISG15, OAS3, EIF2AK2, OASL GO:1903900 regulation of viral life cycle 2.32E-12 8 IFIT1, MX1, RSAD2, OAS2, ISG15, OAS3, EIF2AK2, OASL GO:0050792 regulation of viral process 4.75E-12 8 IFIT1, MX1, RSAD2, OAS2, ISG15, OAS3, EIF2AK2, OASL GO:0016032 viral process 1.27E-10 9 IFIT1, MX1, RSAD2, OAS2, ISG15, OAS3, EIF2AK2, OASL, GO:0019058 viral life cycle 7.83E-10 8 IFIT1, MX1, RSAD2, OAS2 P,A ISRGP915, OAS3, EIF2AK2, OASL GO:0140374 antiviral innate immune 7.24E-09 ...
Claims
Attorney Docket No.32103 / 59579 / PC WHAT IS CLAIMED:
1. A computer-implemented method for computing magnitude-altitude scores corresponding to genes expressing in various ways during one or more time periods during which at least one organ tissue equivalent transits from one or more conditions, the method comprising: (1) receiving, via one or more processors, statistical results including one or more raw p- values that (a) are indicative of a differential gene expression in the organ tissue equivalent under varying infection conditions, from an initial condition to a subsequent condition, over a period of time, and (b) describe significance of gene expression differences from the initial condition to the subsequent condition; (2) generating, via one or more processors, a respective relaxed magnitude-altitude score for one or more genes identified as significant in the statistical results, by computing pairwise comparisons of general linear model quasi likelihood relaxed magnitude altitude scores (GLMQL-RMAS) across diverse conditions; and (3) displaying, via one or more processors, at least some of the GLMQL-RMAS scores for the one or more genes on a display.
2. The computer-implemented method of claim 1, wherein the organ tissue equivalent is a lung organoid.
3. The computer-implemented method of claim 1, wherein the initial condition and / or the subsequent condition corresponds to infected, ultraviolet treated, non-ultraviolet treated, mock, or naive.
4. The computer-implemented method of claim 1, wherein the period of time corresponds to 24 hours or 72 hours.
5. The computer-implemented method of claim 1, wherein the organ tissue equivalent is infected by one or more pathogens.Attorney Docket No.32103 / 59579 / PC 6. The computer-implemented method of claim 5, wherein the one or more pathogens are respective viruses.
7. The computer-implemented method of claim 6, wherein the respective viruses include at least one of influenza A virus, human metapneumovirus, or parainfluenza virus type 3.
8. The computer-implemented method of claim 1, wherein receiving the statistical results includes receiving RNA-Seq data.
9. The computer-implemented method of claim 8, further comprising: determining, via one or more processors, unique statistically significant genes affected by respective viruses during organ tissue equivalent condition transitions, by: for each of the respective viruses, repeating step 2 of claim 1 and step 3 of claim 1 to generate a significance score set that includes respective significant genes and respective ranked relaxed magnitude-altitude scores; and for each of the respective viruses, generating a set of unique significant genes including respective differentially expressed genes.
10. The computer-implemented method of claim 9, further comprising: generating, via one or more processors, a visualization of the set of unique significant genes including respective differentially expressed genes.
11. The computer-implemented method of claim 8, further comprising: determining, via one or more processors, common statistically significant genes affected by all viruses during organ tissue equivalent condition transitions, by:Attorney Docket No.32103 / 59579 / PC for each of the respective viruses, repeating step 2 of claim 1 and step 3 of claim 1 to generate a significance score set that includes respective significant genes and respective ranked relaxed magnitude-altitude scores; and for each respective virus, generating a set of common significant genes including respective differentially expressed genes.
12. The computer-implemented method of claim 11, further comprising: generating, via one or more processors, a visualization of the set of common significant genes including respective differentially expressed genes.
13. The computer-implemented method of claim 1, further comprising: generating, via one or more processors, a visualization of adjusted statistical results.
14. The computer-implemented method of claim 1, further comprising: generating, via one or more processors, a visualization of the respective magnitude- altitude score for the one or more genes; and based on the visualization, tuning hyperparameters M and A.
15. A computer system for computing magnitude-altitude scores corresponding to genes expressing in various ways during one or more time periods during which at least one organ tissue equivalent transits from one or more conditions, comprising: one or more processors, and one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computer system to: (1) receive statistical results including one or more raw p-values that (a) are indicative of a differential gene expression in the organ tissue equivalent under varying infection conditions, from an initial condition to a subsequent condition, over a period of time, and (b) describeAttorney Docket No.32103 / 59579 / PC significance of gene expression differences from the initial condition to the subsequent condition; (2) generate a respective relaxed magnitude-altitude score for one or more genes identified as significant in the statistical results, by computing pairwise comparisons of general linear model quasi likelihood relaxed magnitude altitude scores (GLMQL-RMAS) across diverse conditions; and (3) display, via one or more processors, at least some of the GLMQL-RMAS scores for the one or more genes on a display.
16. The computer system of claim 15, wherein the organ tissue equivalent is infected by one or more viruses; and the memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computer system to: determine, via one or more processors, unique statistically significant genes affected by respective viruses during organ tissue equivalent condition transitions, by: for each of the respective viruses, repeating step 2 of claim 15 and step 3 of claim 15 to generate a significance score set that includes respective significant genes and respective ranked relaxed magnitude-altitude scores; and for each of the respective viruses, generating a set of unique significant genes including respective differentially expressed genes.
17. The computer system of claim 15, wherein the organ tissue equivalent is infected by one or more viruses; and the memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computer system to: determine, via one or more processors, common statistically significant genes affected by all viruses during organ tissue equivalent condition transitions, by:Attorney Docket No.32103 / 59579 / PC for each of the respective viruses, repeating step 2 of claim 15 and step 3 of claim 15 to generate a significance score set that includes respective significant genes and respective ranked relaxed magnitude-altitude scores; and for each respective virus, generating a set of common significant genes including respective differentially expressed genes.
18. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause a computer to: (1) receive statistical results including one or more raw p-values that (a) are indicative of a differential gene expression in an organ tissue equivalent under varying infection conditions, from an initial condition to a subsequent condition, over a period of time, and (b) describe significance of gene expression differences from the initial condition to the subsequent condition; (2) generate a respective relaxed magnitude-altitude score for one or more genes identified as significant in the statistical results, by computing pairwise comparisons of general linear model quasi likelihood relaxed magnitude altitude scores (GLMQL-RMAS) across diverse conditions; and (3) display, via one or more processors, at least some of the GLMQL-RMAS scores for the one or more genes on a display.
19. The non-transitory computer-readable medium of claim 18, wherein the organ tissue equivalent is infected by one or more viruses, and having stored thereon instructions that, when executed by the one or more processors, cause a computer to: determine, via one or more processors, unique statistically significant genes affected by respective viruses during organ tissue equivalent condition transitions, by: for each of the respective viruses, repeating step 2 of claim 15 and step 3 of claim 15 to generate a significance score set that includes respective significant genes and respective ranked relaxed magnitude-altitude scores; andAttorney Docket No.32103 / 59579 / PC for each of the respective viruses, generating a set of unique significant genes including respective differentially expressed genes.
20. The non-transitory computer-readable medium of claim 18, wherein the organ tissue equivalent is infected by one or more viruses, and having stored thereon instructions that, when executed by the one or more processors, cause a computer to: determine, via one or more processors, common statistically significant genes affected by all viruses during organ tissue equivalent condition transitions, by: for each of the respective viruses, repeating step 2 of claim 15 and step 3 of claim 15 to generate a significance score set that includes respective significant genes and respective ranked relaxed magnitude-altitude scores; and for each respective virus, generating a set of common significant genes including respective differentially expressed genes.
21. A computer-implemented method for training a machine learning model using multi-scale gene expression data, the method comprising: collecting molecular biology data over various time points for baseline and treated conditions; applying a Magnitude-Altitude Score Analysis for Tracking Infection and Time- Dependent Genes (MASIT) Training-Validation step to the collected data; utilizing cross-validation techniques for model training; applying the MASIT algorithm to identify a set of genes related to infection and time specific for each fold in the cross-validation; and selecting an optimal machine learning classifier based on the identified genes and its performance evaluated on a hold-out validation set.
22. A computer-implemented method for testing a trained machine learning model on multi-scale gene expression data, the method comprising:Attorney Docket No.32103 / 59579 / PC holding out data to form a test set; applying the selected machine learning classifiers to an RNA-Seq dataset; and testing the performance of the machine learning classifiers using a validation technique, considering different scales in the dataset.
23. The method of claim 21, wherein the molecular biology data comprises NanoString data or similar data collected over various time points for both baseline and treated conditions.
24. The method of claim 21, wherein the cross-validation technique used for model training is a K-Fold Stratified Cross Validation approach, wherein the dataset is split into K subsets and, in each cycle, a different subset is used as the validation set while the remaining subsets are combined to form the training set.
25. The method of claim 21, further comprising constructing a set of genes that includes genes identified by a majority within the k-fold cross-validation as being either infection or time signature genes.
26. The method of claim 21, further comprising testing the model using an independent RNA-Seq dataset to evaluate its generalizability and performance.
27. The method of claim 21, wherein the method includes transfer learning techniques, wherein one or more models are trained on one modality and applied to another modality, specifically training on NanoString data to generate trained models that are then applied to RNASeq data.
28. The method of claim 21, wherein the method leverages the strengths of both NanoString and RNASeq technologies by using the high-quality, focused analysis provided byAttorney Docket No.32103 / 59579 / PC NanoString to inform models that can then be applied to the broader, more comprehensive datasets generated by RNASeq.
29. A computer-implemented method for training a machine learning model using multi-scale gene expression data, the method comprising: receiving molecular biology data over various time points for baseline and treated conditions; applying a Magnitude-Altitude Score Analysis for Tracking Infection and Time- Dependent Genes (MASIT) Training-Validation step to the collected data; utilizing cross-validation techniques for model training; applying the MASIT algorithm to identify a set of genes related to infection and time specific for each fold in the cross-validation; selecting an optimal machine learning classifier based on the identified genes and its performance evaluated on a hold-out validation set; holding out data to form a test set; applying the selected machine learning classifiers to an RNA-Seq dataset; and testing the performance of the machine learning classifiers using the same validation technique, considering different scales in the dataset.
30. The method of claim 29, wherein the molecular biology data comprises NanoString data collected over various time points for both baseline and treated conditions.
31. The method of claim 29, wherein the cross-validation technique used for model training is a K-Fold Stratified Cross Validation approach, wherein the dataset is split into K subsets and, in each cycle, a different subset is used as the validation set while the remaining subsets are combined to form the training set.
32. The method of claim 29, further comprising constructing a set of genes that includes genes identified by a majority within the k-fold cross-validation as being either infection or time signature genes.Attorney Docket No.32103 / 59579 / PC 33. The method of claim 29, further comprising testing the model using an independent RNA-Seq dataset to evaluate its generalizability and performance.
34. The method of claim 29, wherein the method includes transfer learning techniques, wherein one or more models are trained on one modality and applied to another modality, specifically training on NanoString data to generate trained models that are then applied to RNASeq data.
35. The method of claim 29, wherein the method leverages the strengths of both NanoString and RNASeq technologies by using the high-quality, focused analysis provided by NanoString to inform models that can then be applied to the broader, more comprehensive datasets generated by RNASeq.
36. The method of claim 29, wherein the method includes dividing the NanoString dataset into 'k' number of subsets (or folds) for K-fold cross-validation, enhancing the robustness and generalizability of the models developed from NanoString data when applied to RNASeq data.
37. A computer-implemented method for identifying differentially expressed genes under infection conditions using generalized linear models and quasi-likelihood methods, the method comprising: collecting RNA-Seq data from organ tissue equivalents infected with a virus; employing Generalized Linear Models (GLMs) with Quasi-Likelihood (QL) F-tests to analyze the RNA-Seq data; introducing Magnitude-Altitude Score (MAS) and Relaxed Magnitude-Altitude Score (RMAS) algorithms to navigate the RNA-Seq data; andAttorney Docket No.32103 / 59579 / PC identifying genes differentially expressed under various infection conditions as potential biomarkers.
38. The method of claim 37, wherein the RNA-Seq data is collected from organ tissue equivalents infected with Influenza A virus (IAV), Human metapneumovirus (MPV), or Parainfluenza virus type 3 (PIV3).
39. The method of claim 37, wherein the GLMs with QL F-tests are utilized to model the RNA-Seq count data, allowing for response variables to have error distribution models other than a normal distribution.
40. The method of claim 37, wherein the MAS algorithm integrates both the magnitude of expression change and its statistical significance to prioritize genes for their potential as infection-specific biomarkers.
41. The method of claim 37, further comprising conducting Gene Ontology (GO) enrichment analysis to interpret the biological significance of the identified differentially expressed genes.
42. The method of claim 41, wherein the GO enrichment analysis identifies key biological processes activated in response to viral infections, including activation of interferon- stimulated genes and innate immune mechanisms.
43. The method of claim 37, further comprising employing a multinomial logistic regression model to classify samples into distinct classes based on the identified differentially expressed genes.Attorney Docket No.32103 / 59579 / PC 44. The method of claim 37, wherein the classification of samples achieves a mean accuracy of 92% in distinguishing between different viral infections using a stratified k-fold cross-validation approach.
45. A system for identifying gene signatures and pathways associated with mpox virus-induced gastrointestinal complications, comprising: a data processing unit configured to receive RNA-Seq data derived from colon organoid models exposed to the mpox virus; a machine learning module operatively connected to the data processing unit, the machine learning module trained to: analyze the RNA-Seq data to identify gene signatures and pathways indicative of gastrointestinal complications caused by the mpox virus; and an output interface configured to display the identified gene signatures and pathways, wherein the system facilitates the understanding of mpox virus-induced gastrointestinal complications through the analysis of colon organoid model data.
46. The system of claim 45, wherein the machine learning module is further trained to differentiate between gene signatures associated with mpox virus-induced gastrointestinal complications and those associated with other gastrointestinal diseases.
47. The system of claim 45, wherein the RNA-Seq data is obtained from colon organoid models that have been exposed to varying concentrations of the mpox virus to simulate different stages of infection.
48. The system of claim 45, further comprising a data storage unit for storing the RNA- Seq data, the identified gene signatures, and pathways for subsequent analysis.Attorney Docket No.32103 / 59579 / PC 49. The system of claim 45, wherein the output interface is further configured to provide recommendations for potential therapeutic interventions based on the identified gene signatures and pathways.
50. The system of claim 45, wherein the machine learning module includes a neural network trained on a dataset comprising RNA-Seq data from colon organoid models both with and without exposure to the mpox virus.
51. The system of claim 45, further comprising a sample preparation module for preparing the colon organoid models for RNA-Seq analysis, including isolating RNA from the colon organoid models.
52. A computer-implemented method for comparing gene expression analysis techniques, comprising: obtaining a first set of gene expression data from a biological sample using a NanoString analysis technique; obtaining a second set of gene expression data from the same biological sample using an RNA-Seq analysis technique; processing the first and second sets of gene expression data to normalize the data sets; comparing the normalized gene expression data obtained from the NanoString analysis technique with the normalized gene expression data obtained from the RNA-Seq analysis technique; and determining differences in gene expression measurements between the NanoString and RNA-Seq analysis techniques, wherein the method facilitates the evaluation of the comparative effectiveness of NanoString and RNA-Seq in gene expression analysis.Attorney Docket No.32103 / 59579 / PC 53. The method of claim 52, wherein the biological sample is derived from a subject diagnosed with a viral infection, and the gene expression data relates to genes implicated in the viral infection response.
54. The method of claim 52, further comprising applying statistical analysis to the differences in gene expression measurements to identify statistically significant differences between the NanoString and RNA-Seq analysis techniques.
55. The method of claim 52, wherein the normalization of the data sets includes the use of control genes known to have stable expression across different analysis techniques.
56. The method of claim 52, further comprising visualizing the compared gene expression data using a graphical representation to highlight the differences in gene expression measurements between the NanoString and RNA-Seq analysis techniques.
57. The method of claim 52, wherein the determining step includes quantifying the extent of discrepancy in gene expression measurements between the NanoString and RNA-Seq analysis techniques for specific genes of interest.
58. The method of claim 52, further including correlating the differences in gene expression measurements with the sensitivity and specificity of the NanoString and RNA-Seq analysis techniques in detecting specific gene expression changes.
59. A computer-implemented method for analyzing RNA-seq data for breast cancer, comprising: receiving RNA-seq data derived from breast cancer tissue samples; applying a Generalized Linear Model Quasi-Likelihood (GLMQL) algorithm to the RNA- seq data to identify patterns and correlations specific to breast cancer;Attorney Docket No.32103 / 59579 / PC utilizing a Model Averaging and Selection (MAS) technique to refine the analysis by selecting the most predictive models of breast cancer based on the identified patterns and correlations; and generating a report that includes the results of the GLMQL and MAS analysis, wherein the report identifies key genetic markers and potential therapeutic targets for breast cancer.
60. The method of claim 59, wherein the RNA-seq data is obtained from a cohort of subjects diagnosed with various stages of breast cancer, allowing for the identification of stage- specific genetic markers and therapeutic targets.
61. The method of claim 59, further comprising comparing the identified genetic markers and potential therapeutic targets with a database of known genetic markers and therapies for breast cancer to validate the findings of the GLMQL and MAS analysis.
62. The method of claim 59, wherein the GLMQL algorithm is specifically configured to identify gene expression patterns that are predictive of treatment response in breast cancer patients.
63. The method of claim 59, further comprising applying a cross-validation technique to assess the predictive accuracy of the models selected by the MAS technique.
64. The method of claim 59, wherein the report includes a risk assessment for each identified genetic marker, indicating the likelihood of the marker being associated with aggressive forms of breast cancer.Attorney Docket No.32103 / 59579 / PC 65. The method of claim 59, further comprising utilizing the identified genetic markers to design personalized medicine approaches for the treatment of breast cancer, including the selection of targeted therapies.
66. A computer-implemented method for analyzing the impact of the Ebola virus on gene expression in nonhuman primates, comprising: collecting RNA-seq data from tissue samples of nonhuman primates infected with the Ebola virus; applying a machine learning algorithm to the RNA-seq data to identify changes in gene expression patterns caused by the Ebola virus infection; comparing the identified changes in gene expression patterns with those from uninfected control nonhuman primates to determine the specific impact of the Ebola virus on gene expression; and generating a report detailing the differences in gene expression between infected and uninfected nonhuman primates, wherein the report aids in understanding the molecular mechanisms of Ebola virus pathogenesis in nonhuman primates.
67. The method of claim 66, wherein the machine learning algorithm is trained using a dataset comprising RNA-seq data from both Ebola virus-infected and uninfected nonhuman primates across multiple time points post-infection to capture the dynamic changes in gene expression over the course of the infection.
68. The method of claim 66, further comprising identifying gene expression biomarkers that are predictive of disease severity or outcome in nonhuman primates infected with the Ebola virus.Attorney Docket No.32103 / 59579 / PC 69. The method of claim 66, wherein the comparison includes the use of statistical analysis techniques to identify statistically significant differences in gene expression patterns between infected and uninfected nonhuman primates.
70. The method of claim 66, further comprising correlating the identified changes in gene expression with specific pathological features observed in the nonhuman primates, such as hemorrhagic manifestations or immune response alterations.
71. The method of claim 66, wherein the machine learning algorithm includes a deep learning model that is capable of identifying complex, nonlinear relationships in the RNA-seq data.
72. The method of claim 66, further comprising utilizing the findings from the report to develop therapeutic strategies aimed at mitigating the impact of the Ebola virus on gene expression in nonhuman primates, potentially translating these strategies to human treatments.
73. A computer-implemented method for analyzing the gene expression profile of dengue-infected peripheral blood mononuclear cells (PBMCs), comprising: collecting RNA-seq data from PBMCs infected with the dengue virus; applying machine learning techniques to the RNA-seq data to identify specific gene expression profiles associated with dengue virus infection; comparing the identified gene expression profiles with those from uninfected PBMCs to determine the unique impact of the dengue virus on gene expression; and generating a report that includes the analysis results, wherein the report provides insights into the molecular response of PBMCs to dengue virus infection and identifies potential biomarkers for dengue virus infection detection and progression monitoring.Attorney Docket No.32103 / 59579 / PC 74. The method of claim 73, wherein the machine learning techniques include supervised learning algorithms trained on a dataset comprising RNA-seq data from dengue virus-infected and uninfected PBMCs to accurately classify the samples based on their gene expression profiles.
75. The method of claim 73, further comprising using the identified gene expression profiles to investigate the immune response mechanisms of PBMCs to dengue virus infection, including the activation of specific immune pathways.
76. The method of claim 73, wherein the comparison of gene expression profiles involves the use of differential expression analysis to identify genes that are significantly upregulated or downregulated in dengue virus-infected PBMCs compared to uninfected PBMCs.
77. The method of claim 73, further comprising validating the identified gene expression profiles and potential biomarkers using quantitative PCR (qPCR) on a separate cohort of dengue virus-infected and uninfected PBMC samples.
78. The method of claim 73, wherein the machine learning techniques include the use of feature selection algorithms to identify a subset of genes whose expression levels most effectively distinguish between dengue virus-infected and uninfected PBMCs.
79. The method of claim 73, further comprising integrating the gene expression data with clinical data from the subjects from whom the PBMC samples were collected to explore correlations between gene expression profiles and clinical manifestations of dengue virus infection.
80. A computer-implemented method for identifying gene signatures and pathways in mpox virus-induced gastrointestinal complications using colon organoid models, comprising:Attorney Docket No.32103 / 59579 / PC collecting RNA-seq data from colon organoid models infected with the mpox virus; applying machine learning analysis to the RNA-seq data to identify gene signatures and pathways associated with gastrointestinal complications induced by the mpox virus; comparing the identified gene signatures and pathways with those from uninfected colon organoid models to determine the specific impact of the mpox virus on gene expression; and generating a report that includes the analysis results, wherein the report provides insights into the molecular mechanisms underlying mpox virus-induced gastrointestinal complications and identifies potential targets for therapeutic intervention.
81. The method of claim 80, wherein the machine learning analysis includes the use of a supervised learning algorithm trained on a dataset comprising RNA-seq data from both mpox virus-infected and uninfected colon organoid models to enhance the accuracy of identifying gene signatures and pathways.
82. The method of claim 80, further comprising utilizing the identified gene signatures and pathways to screen for compounds that can modulate these pathways, aiming to mitigate the gastrointestinal complications induced by the mpox virus.
83. The method of claim 80, wherein the comparison of gene signatures and pathways involves the application of differential expression analysis to pinpoint genes that are significantly upregulated or downregulated in response to mpox virus infection.
84. The method of claim 80, further comprising validating the identified gene signatures and pathways through additional experimental approaches, such as quantitative PCR or Western blotting, on separate sets of colon organoid models.Attorney Docket No.32103 / 59579 / PC 85. The method of claim 80, wherein the machine learning analysis incorporates a feature selection process to identify a subset of genes and pathways that most significantly contribute to the gastrointestinal complications observed in mpox virus-infected colon organoid models.
86. The method of claim 80, further comprising correlating the identified gene signatures and pathways with clinical data from subjects infected with the mpox virus to assess the relevance of these molecular findings to human disease.
87. A computer-implemented method for analyzing RNA-seq data from human lung organoids in response to viral infections, comprising: applying a Generalized Linear Model Quasi-Likelihood- Relaxed Magnitude-Altitude Score (GLMQL-RMAS) algorithm to RNA-seq data obtained from human lung organoids infected with at least one of Influenza A virus (IAV), Human metapneumovirus (MPV), and Parainfluenza virus type 3 (PIV3); identifying genes capable of differentiating between mock and infected samples at two post-infection time points using the GLMQL-RMAS algorithm; and generating a report that includes the identified genes, wherein the GLMQL-RMAS algorithm demonstrates superior effectiveness in gene ranking compared to traditional methods based on p-values or log fold change (logFC).
88. A computer-implemented method for identifying genes associated with axillary lymph node metastasis in cancer studies, comprising: analyzing RNA-seq data from subjects with and without positive axillary lymph node metastasis using a Generalized Linear Model Quasi-Likelihood- Relaxed Magnitude-Altitude Score (GLMQL-RMAS) algorithm; identifying genes capable of differentiating between subjects with positive axillary lymph node metastasis and those without;Attorney Docket No.32103 / 59579 / PC validating the identified genes through Gene Ontology (GO) and Gene Set Enrichment Analysis (GSEA) Hallmark pathway analysis; and generating a report that includes the analysis results, wherein the GLMQL-RMAS algorithm is utilized for effective gene identification and ranking.
89. A computer-implemented method for evaluating the efficacy of the GLMQL-RMAS algorithm in viral infection studies, comprising: applying the GLMQL-RMAS algorithm to RNA-seq data from Ebola-infected nonhuman primates; identifying top genes capable of differentiating between positive and negative samples in a held-test set; comparing the accuracy of the GLMQL-RMAS algorithm in identifying the top genes with the accuracy of traditional methods based on p-values or logFC, such as EdgeR or DESeq2; and generating a report that demonstrates the superior performance of the GLMQL-RMAS algorithm in gene ranking with 100% accuracy compared to the best performance of traditional methods.
90. A computer-implemented method for predictive modeling of multi-omic data in 3D airway organ tissue equivalents during viral infection, comprising: collecting multi-omic data, including at least two of genomic, transcriptomic, proteomic, and metabolomic data, from 3D airway organ tissue equivalents exposed to a viral infection; applying a cross-modal predictive modeling algorithm to the collected multi-omic data to identify patterns and correlations across the different omic datasets that are indicative of the response of the 3D airway organ tissue equivalents to the viral infection; generating predictive models based on the identified patterns and correlations that forecast the outcomes of viral infection on the 3D airway organ tissue equivalents; andAttorney Docket No.32103 / 59579 / PC producing a report that includes the predictive models and their potential implications for understanding the mechanisms of viral infection in 3D airway organ tissue equivalents, wherein the method enables the integration of multi-omic data to enhance the prediction and understanding of viral infection impacts.
91. The method of claim 90, wherein the cross-modal predictive modeling algorithm incorporates machine learning techniques to enhance the accuracy and reliability of the predictive models by learning from complex patterns in the multi-omic data.
92. The method of claim 90, further comprising the step of validating the predictive models using additional 3D airway organ tissue equivalents not included in the initial dataset to assess the generalizability of the models across different viral infections.
93. The method of claim 90, wherein the multi-omic data includes transcriptomic data obtained through RNA sequencing and proteomic data obtained through mass spectrometry, providing a comprehensive overview of the molecular response of the 3D airway organ tissue equivalents to viral infection.
94. The method of claim 90, further comprising the step of applying the predictive models to design targeted therapeutic interventions aimed at mitigating the impact of viral infections on 3D airway organ tissue equivalents, based on the identified patterns and correlations.
95. The method of claim 90, wherein the viral infection is at least one of Influenza A virus, Human metapneumovirus, and Parainfluenza virus type 3, allowing for the assessment of the method's effectiveness across different types of viral infections.
96. The method of claim 90, further comprising the step of integrating environmental data, such as temperature and humidity conditions, with the multi-omic data in the cross-modalAttorney Docket No.32103 / 59579 / PC predictive modeling algorithm to assess their impact on the outcomes of viral infection in 3D airway organ tissue equivalents.
97. A computer-implemented method for analyzing biological data to identify molecular responses and predictive markers in disease conditions, comprising: collecting biological data from samples exposed to a disease condition, wherein the biological data includes at least one of RNA-seq data, multi-omic data, and gene expression data; applying a computational analysis algorithm, including at least one of Generalized Linear Model Quasi-Likelihood- Relaxed Magnitude-Altitude Score (GLMQL-RMAS), machine learning techniques, and cross-modal predictive modeling, to the collected biological data to identify patterns, correlations, gene signatures, and pathways associated with the disease condition; comparing the identified molecular responses and predictive markers with those from control samples to determine the specific impact of the disease condition; validating the identified molecular responses and predictive markers through additional experimental or computational methods; and generating a report that includes the analysis results, wherein the report provides insights into the molecular mechanisms underlying the disease condition and identifies potential targets for therapeutic intervention or diagnostic markers, thereby enabling the integration and analysis of diverse biological data types to enhance the understanding and management of disease conditions.
98. A computer-implemented method for training a computational analysis algorithm for disease condition analysis, comprising: collecting a training dataset that includes biological data from samples exposed to various disease conditions and control samples not exposed to disease conditions, wherein the biological data comprises RNA-seq data, multi-omic data, and gene expression data; preprocessing the collected biological data to normalize and standardize the data for analysis;Attorney Docket No.32103 / 59579 / PC applying machine learning techniques to the preprocessed biological data to train a computational analysis algorithm, wherein the machine learning techniques include at least one of supervised learning, unsupervised learning, and deep learning; evaluating the performance of the trained computational analysis algorithm using a validation dataset distinct from the training dataset to ensure accuracy and reliability in identifying patterns, correlations, gene signatures, and pathways associated with disease conditions; and refining the computational analysis algorithm based on the evaluation results to optimize its performance for analyzing biological data related to disease conditions.
99. A computer-implemented method for developing a predictive model for disease condition outcomes, comprising: assembling a dataset of multi-omic data from samples under various disease conditions and corresponding outcomes; employing a cross-modal predictive modeling approach to analyze the multi-omic data, wherein the approach integrates data across genomic, transcriptomic, proteomic, and metabolomic profiles; training a predictive model using the cross-modal analysis to forecast disease condition outcomes based on identified patterns and correlations in the multi-omic data; validating the predictive model's accuracy and generalizability across different datasets and disease conditions; and updating the predictive model iteratively based on feedback from ongoing validation efforts to enhance its predictive accuracy for disease condition outcomes.
100. A computer-implemented method for training a gene ranking algorithm for use in disease condition analysis, comprising:Attorney Docket No.32103 / 59579 / PC compiling a comprehensive dataset that includes gene expression data from both diseased and healthy samples, with the diseased samples representing a variety of disease conditions; implementing a Generalized Linear Model Quasi-Likelihood- Relaxed Magnitude-Altitude Score (GLMQL-RMAS) framework to analyze the gene expression data; training the GLMQL-RMAS algorithm to rank genes based on their ability to differentiate between diseased and healthy samples, taking into account factors such as p-values and log fold change (logFC); conducting cross-validation tests to assess the gene ranking algorithm's effectiveness in accurately distinguishing between different disease conditions; and fine-tuning the GLMQL- RMAS algorithm parameters based on cross-validation results to improve its gene ranking performance for disease condition analysis.
101. A method for training a Magnitude-Altitude Score Analysis for Tracking Infection and Time-Dependent Genes (MASIT) algorithm, comprising: collecting a dataset that includes time-series gene expression data from biological samples subjected to infection over multiple time points; preprocessing the collected gene expression data to correct for variations and normalize the data across different time points and samples; implementing the MASIT algorithm to analyze the preprocessed time-series gene expression data, wherein the algorithm is designed to identify and score genes based on changes in expression magnitude and altitude over the course of the infection; training the MASIT algorithm using a subset of the dataset where the infection status and time points are known, enabling the algorithm to learn patterns of gene expression changes that are indicative of infection progression; validating the trained MASIT algorithm on a separate subset of the dataset not used in training to assess its effectiveness in accurately tracking infection and identifying time- dependent gene expression changes; andAttorney Docket No.32103 / 59579 / PC refining the training process based on validation results, including adjusting algorithm parameters and selection criteria, to enhance the MASIT algorithm's ability to provide insights into the dynamics of gene expression in response to infection over time.
102. A method for training a Magnitude-Altitude Score Analysis for Tracking Infection and Time-Dependent Genes (MASIT) algorithm, comprising: collecting a dataset that includes time-series gene expression data from biological samples subjected to infection over multiple time points; preprocessing the collected gene expression data to correct for variations and normalize the data across different time points and samples; implementing the MASIT algorithm to analyze the preprocessed time-series gene expression data, wherein the algorithm is designed to identify and score genes based on changes in expression magnitude and altitude over the course of the infection; and training the MASIT algorithm using a subset of the dataset where the infection status and time points are known, enabling the algorithm to learn patterns of gene expression changes that are indicative of infection progression.
103. A method for utilizing a trained Magnitude-Altitude Score Analysis for Tracking Infection and Time-Dependent Genes (MASIT) algorithm, comprising: inputting time-series gene expression data from biological samples subjected to infection into the trained MASIT algorithm; analyzing the inputted gene expression data with the trained MASIT algorithm to identify and score genes based on changes in expression magnitude and altitude over the course of the infection; and outputting results from the MASIT algorithm that highlight time-dependent gene expression changes indicative of infection progression, wherein the outputted results provide insights into the dynamics of gene expression in response to infection over time.Attorney Docket No.32103 / 59579 / PC 104. A system for training a Magnitude-Altitude Score Analysis for Tracking Infection and Time-Dependent Genes (MASIT) algorithm, comprising: a data collection module configured to collect a dataset including time-series gene expression data from biological samples subjected to infection over multiple time points; a preprocessing module configured to correct for variations and normalize the collected gene expression data across different time points and samples; a processing unit comprising a processor and memory, wherein the memory stores instructions that, when executed by the processor, implement the MASIT algorithm to: analyze the preprocessed time-series gene expression data by identifying and scoring genes based on changes in expression magnitude and altitude over the course of the infection; and a training module configured to train the MASIT algorithm using a subset of the dataset where the infection status and time points are known, enabling the algorithm to learn patterns of gene expression changes indicative of infection progression.
105. A computing system for utilizing a trained Magnitude-Altitude Score Analysis for Tracking Infection and Time-Dependent Genes (MASIT) algorithm, comprising: an input module configured to receive time-series gene expression data from biological samples subjected to infection; a processing unit comprising a processor and memory, wherein the memory stores instructions that, when executed by the processor, utilize the trained MASIT algorithm to analyze the received gene expression data by identifying and scoring genes based on changes in expression magnitude and altitude over the course of the infection; and an output module configured to provide results from the MASIT algorithm that highlight time-dependent gene expression changes indicative of infection progression, offering insights into the dynamics of gene expression in response to infection over time.Attorney Docket No.32103 / 59579 / PC 106. A computer-readable medium having stored thereon instructions that, when executed by a processor, perform a method for training a Magnitude-Altitude Score Analysis for Tracking Infection and Time-Dependent Genes (MASIT) algorithm, the method comprising: collecting a dataset including time-series gene expression data from biological samples subjected to infection over multiple time points; preprocessing the collected gene expression data to correct for variations and normalize the data across different time points and samples; implementing the MASIT algorithm to analyze the preprocessed time-series gene expression data by identifying and scoring genes based on changes in expression magnitude and altitude over the course of the infection; and training the MASIT algorithm using a subset of the dataset where the infection status and time points are known, enabling the algorithm to learn patterns of gene expression changes indicative of infection progression.
107. A computer-readable medium having stored thereon instructions that, when executed by a processor, perform a method for: utilizing a trained Magnitude-Altitude Score Analysis for Tracking Infection and Time- Dependent Genes (MASIT) algorithm, the method comprising: receiving time-series gene expression data from biological samples subjected to infection; analyzing the received gene expression data with the trained MASIT algorithm to identify and score genes based on changes in expression magnitude and altitude over the course of the infection; and outputting results from the MASIT algorithm that highlight time-dependent gene expression changes indicative of infection progression, providing insights into the dynamics of gene expression in response to infection over time.
Citation Information
Patent Citations
Method of toxicological evaluation, method of toxicological screening and associated system
US20130060550A1
Active water phantom for three-dimensional ion beam therapy quality assurance
US20160135765A1
Methods for determining spatial and temporal gene expression dynamics in single cells
US20190218276A1
Gene expression-based biomarker for the detection and monitoring of bronchial premalignant lesions
US20210254171A1
Detection and prediction of infectious disease
US20210403986A1
Cited By
Gene abnormality regulation and control detection method based on regulation and control pathway disturbance analysis
CN120913644A