Method for developing a cancer organoid digital twin for combinatorial therapeutic efficacy prediction and system and applications associated therewith
A digital twin system using patient-derived organoids with 3D imaging and transcriptomics predicts cancer therapy efficacy, addressing the limitations of existing methods by providing personalized and scalable treatment solutions.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMUELSSON ERIK REINHOLD
- Filing Date
- 2025-11-26
- Publication Date
- 2026-06-04
Smart Images

Figure SE2025010038_04062026_PF_FP_ABST
Abstract
Description
[0001] Method for developing a cancer organoid digital twin for combinatorial therapeutic efficacy prediction and system and applications associated therewith
[0002] Technical field
[0003] The present invention refers to the fields of cancer treatment, digital twins and combinatorial therapeutic efficacy prediction.
[0004] Technical background
[0005] Efforts utilizing cerebral organoids to model various cancers, viral infections, developmental and neurological disorders are in the early stages of application and represent a new frontier for understanding and developing treatments. In particular, the development and usage of organoids to model glioblastoma models (GB) has seen a spike in recent years, with applications ranging from assessing oncogene expression and mutation, tumor infiltration, formation of microtubule networks consistent with GB pathophysiology, and intra- and inter-tumor heterogeneity.
[0006] An ex vivo glioblastoma organoid model has been disclosed in US2023 / 0324371A1, wherein the similarity of organoid phenotypes to the current known 4 phenotypes (MES-like, OPC-like, NPC-like and AC-like) which were identified from a set of 29 glioblastoma patients is disclosed.
[0007] US20210202098A1 describes methods involving organoid models but appears to not combine all the specific steps of this method in the same manner. It emphasizes the integration of multiple data types for therapy optimization. W02023004149A1 focuses on personalized cancer treatments but appears not to emphasize the use of spatial transcriptomics or tissue clearing data. W02024050133A1 involves organoids and cancer therapies but lacks detailed procedural steps outlined, particularly the specific data types collected and their integration into a digital model. U US20220341914A1 discloses "large-scale organoid analysis" with classifier training over sequencing-derived features and organoid sensitivities, but it neither teaches nor enables (i) tri-modal data capture that includes tissue-cleared 3-D imaging and spatial-omics, (ii) longitudinal (time-series) fusion of those modalities, nor (iii) a closed-loop controller that updates combination-therapy recommendations by comparing a digital twin to new patient and / or organoid measurements.
[0008] US20220341914A1 describes software and systems for predicting treatment outcomes using organoid data and sequencing-derived features. It does not address multimodal data integration or iterative control aspects in detail. Thus, new therapies precl i nical ly tested on 2D cell culture models and animal models almost uniformly fail. The main reason is that the models do not represent the true states or full range of states present in actual patient tumors. Existing methods are used for preclinical assessment of therapies, but do not take into account personalized medicine, thereby defining therapeutic regimens that meet individual patient needs.
[0009] Hence, there is a need in the art for improved methods for overcoming the drawbacks of the existing solutions.
[0010] Summary of the invention The present invention aims at solving the problems and drawbacks of the art, thereby providing an improved method, a system for performing the method, and related applications and uses.
[0011] In a first aspect, the invention relates to a method for developing a digital twin of an organoid model for optimising cancer therapy for at least one patient diagnosed with cancer as defined in the appended claims.
[0012] Hereby, by using organoids cultivated with patient-derived cancer cells, one can compare organoids treated with therapeutics at different concentrations and durations. Further, by selecting patient organoids at different time points during treatment, one can then extract them and process them using e.g. (i) tissue clearing & 3D imaging, (ii) scRNAseq and (iii) spatial transcriptomics. The extract data can be used to develop a digital twin of each individual patient's cancer and based on the differing responses of the organoid model to different therapies, make predictions on the combinations of therapies that can either halt or eliminate cancer. The predictions can then be validated in a new experiment, and the results used to validate or adjust the prediction.
[0013] In some embodiments, the cancer cells are derived from glioblastoma. In some embodiments, the cancer cells are derived from liver, lung, kidney, intestines, pancreas, bone marrow, skin, heart, adipose tissue, bladder, breast, colon, or brain. Hence, the method of the present disclosure could be used for cancer cells of any kind, and these are examples of tissues for which the method could be used.
[0014] In some embodiments, the cancer cells are derived specifically from glioblastoma patients, and the digital twin incorporates glioblastoma-specific biomarkers identified through advanced transcriptomic analysis. The present method has been developed primarily for use for glioblastoma patients. By using incorporating specific biomarkers in the digital twin, a higher success rate is expected. Thus, the examples provided in this disclosure (iCNV and the clinical definitions) are used to define a cancer cell, but typically the metaprogram analysis will be used to profile the identified cancer cells, i.e. the "advanced transcriptomic analysis" referred to.
[0015] In some embodiments, the organoids are derived from tissues such as liver, lung, kidney, intestines, pancreas, bone marrow, skin, heart, adipose tissue, bladder, breast, colon, or brain, such as cerebral organoids, for use in studying primary cancers or metastases chosen from those targeting the brain or originating in non-brain tissues.
[0016] In some embodiments, the organoids are cerebral organoids, liver organoids, lung organoids, kidney organoids, intestinal organoids, pancreatic organoids, bone marrow organoids, skin organoids, cardiac organoids, adipose tissue organoids, or bladder organoids, or organoids from breast or colon, and the method includes organoid-specific data such as neural layer patterns obtained through immunofluorescent staining. Hereby, cerebral organoids and organoids from other tissues, as models for developing treatments are used.
[0017] Thus, the present invention includes both scenarios (1) modelling metastasis of non-glioblastoma cancers to the brain and (2) using organoids derived from non-brain tissues to study metastasis or cancer-specific interactions in those organs. For example, cerebral organoids can be used for brain- targeted metastases while accommodating organoids from other tissues for non-brain cancer studies, i.e. including glioblastoma and other cancer types. In some embodiments, the digital copies are developed for a large cohort of patients, and the aggregated data is used to identify population-wide trends in therapy responses and to refine individual predictions. Hereby, the method of this disclosure is used for developing understanding of suitable treatments on a population level.
[0018] In some embodiments, the patient and the model are added and compared to an integrated patient atlas of glioblastoma. Hereby, the patient and organoid data extracted is added to a global glioblastoma cell atlas, thereby constantly improving the global glioblastoma reference and providing context for patient-specific vs population-specific features.
[0019] The number of metaprograms identified from the inventor's cohort of 140 GB (glioblastoma) patients scRNAseq data was 30. As a potential advantage, the metaprograms can likely change as the number of patients and samples increases.
[0020] For example, if one increases the number of patient-derived scRNAseq datasets from 140 to 240, one would likely end up with an increased number of metaprograms that better defines the range of transcriptomic states present in glioblastoma.
[0021] Basically, the more samples added to the reference atlas (both from patients and organoids) the more high-resolution the picture of GB is, and the better one is able to predict and assess patient outcome / drug response, etc. Thus, this defines a dynamic approach that is meant to change (and improve) with the samples added to it.
[0022] In some embodiments, the prediction of combinations of therapies that can halt or eliminate the cancer development in the patient(s) is based on a comparison of the digital twin data with real-time transcriptomics / sequencing data from the patient during therapy, enabling dynamic adjustments to the treatment regimen.
[0023] In some embodiments, the method further comprises validating the adjusted regimen by re-dosing a subsequent batch of organoids and confirming improved efficacy, thereby completing a closed-loop iteration. Hence, the closed-loop controller may trigger re-dosing of organoids to validate updated therapy recommendations.
[0024] In some embodiments, integration of two modalities (e.g., RNA-seq and 3D imaging) is sufficient to generate a functional digital twin, with optional compensation using CNV or biomarker data.
[0025] In a second aspect, the invention relates to a system as defined in the appended claims. In some embodiments, the system is designed for performing the method of the first aspect.
[0026] In some embodiments, the system comprises a storage unit for securely storing patient data, organoid data, digital copies, and therapy predictions, ensuring compliance with data protection regulations.
[0027] In some embodiments, the system comprises a user interface configured to allow researchers or clinicians to: - (i) input patient-specific data, - (ii) view and analyze the digital copies and predicted therapy outcomes, and - (iii) adjust treatment plans based on real-time feedback from the prediction engine. In some embodiments, the system comprises a communication interface for integrating the patient and organoid data into a global glioblastoma cell atlas, facilitating the identification of populationwide trends and the refinement of individual predictions.
[0028] In a third aspect, the invention relates to a non-transitory computer-readable medium as defined in the appended claims.
[0029] In a fourth aspect, the invention relates to the use of the method of the first aspect, the system of the second aspect, or the non-transitory computer-readable medium of the third aspect as defined in the appended claims, in any of the following applications or products:
[0030] (i) a personalized cancer therapy platform; (ii) a drug development and testing platform;
[0031] (iii) integrated patient care systems;
[0032] (iv) a glioblastoma-specific clinical decision support tool;
[0033] (v) a global cancer research database subscription;
[0034] (vi) predictive analytics for clinical trials; or
[0035] (vii) personalized oncology services.
[0036] Hence, the method, system, non-transitory computer-readable medium and use provided in this disclosure are used to assess and predict combinations of therapies to produce intended effects. A common limitation of testing combination therapies is the number of models and patients needed: as the number of combined therapies increases (along with the range of dosages), the number of test conditions increases beyond available patients and models. Using a system of paired wet-lab to digital twin experiments can remove this limitation and enable the testing and prediction of combination therapies for either individual patients or large cohorts of patients.
[0037] Further, implementation of the present method, system, computer-readable medium and use could eventually see the dropping of traditional drug testing methods for this new approach.
[0038] Moreover, by relating the range of states mapped in the models to patient data, one could predict the course of a patient's disease by taking a biopsy for analysis (transcriptomic) and comparing it to the projected development and ranges in the model.
[0039] Also, the data extracted and disclosed herein can be used for the invention to assess the state of the cancer cells in the model over time, assess deviations from controls and compare it to reference patient data. By treating the organoid models with multiple types of therapeutics, one can map the range of states (transcriptomic, spatial, number of cells, etc) and predict working combinations on a population (if one is testing a cohort of patient-derived cancer cells) or on a personal (precisionmedicine) basis. The resulting data can be used to then predict and train a digital twin so that eventually, the model will have a high-enough fidelity to the actual cancer development that a single patient biopsy sample can be extracted, sequenced, and input into the model to predict a working therapeutic combination for that patient. The digital twin could be used in a clinical or research setting to predict therapeutic response and disease progression. ummarise some of the advantages and unexpected effects of the present invention:
[0040] 1. Enhanced Precision Medicine: The invention enables highly personalized treatment regimens by incorporating patient-specific data into the digital twin, ensuring that therapeutic strategies are tailored to the unique characteristics of each patient's cancer. This approach significantly increases the likelihood of therapeutic success and reduces the risk of adverse effects.
[0041] 2. Dynamic Therapy Adjustment: By integrating real-time transcriptomic and sequencing data from patients during therapy, the digital twin allows for continuous monitoring and dynamic adjustment of treatment regimens. This adaptability ensures that the therapy remains effective even as the cancer evolves, preventing resistance and relapse.
[0042] 3. Reduction of Preclinical Attrition Rates: The method overcomes the limitations of traditional 2D cell cultures and animal models by providing a more accurate representation of human cancer biology. This leads to better predictive power for therapeutic efficacy, thereby reducing the high attrition rates seen in preclinical drug development.
[0043] 4. Scalable Testing for Combination Therapies: The paired wet-lab and digital twin approach allows for the scalable testing of a vast number of combinatorial therapies, including varying dosages and sequences. This scalability enables the exploration of complex therapeutic strategies that would otherwise be impractical due to the sheer number of required experimental conditions.
[0044] 5. Population-Wide Insights and Individual Predictions: By developing digital twins for a large cohort of patients and aggregating the data, the invention provides insights into populationwide trends in therapy responses. This dual benefit allows for the refinement of both individual predictions and broader therapeutic strategies, leading to more effective treatments across diverse patient populations.
[0045] 6. Integration with Global Cancer Atlases: The invention's ability to compare patient and organoid data with a global cancer atlas enhances the understanding of disease heterogeneity and evolution. This integration provides a powerful tool for researchers and clinicians to identify commonalities and differences in cancer progression and response to therapy across different populations.
[0046] 7. Early-Stage Disease Prediction: The method enables early detection and prediction of disease progression by comparing patient biopsy data with the digital twin's projected development. This early intervention capability can lead to more timely and effective treatments, potentially improving patient outcomes and survival rates.
[0047] 8. Cost and Time Efficiency: By leveraging digital twins, the invention reduces the need for extensive and time-consuming physical experiments. This efficiency accelerates the drug development process, lowers costs, and expedites the translation of new therapies from the lab to the clinic.
[0048] 9. Potential to Replace Traditional Drug Testing: As the fidelity of the digital twin improves, the method could eventually replace traditional drug testing methods, offering a more accurate and efficient means of developing and optimizing cancer therapies. This shift could revolutionize the preclinical testing landscape, making it more predictive of clinical success.
[0049] 10. Cross-Disease Applicability: While the method is particularly suited for glioblastoma and other cancers, its principles could be adapted for other complex diseases that benefit from organoid modeling. This cross-disease applicability broadens the scope and impact of the invention, making it a versatile tool in biomedical research and personalized medicine.
[0050] Brief description of the drawings
[0051] Figure 1 (A-E) refers to a Metaprogram Definition and Expression Signatures Among Malignant GB, Cells, wherein:
[0052] A, Heatmap of Jaccard Similarity Indices with Cluster Annotations. This heatmap displays the Jaccard similarity indices among 1,090 robustcNMF programs, clustered into 30 metaprograms. The color intensity represents the degree of similarity between samples, with darker colors indicating higher similarity. The clusters are labeled and color-coded along the axes, providing a visual representation of the internal similarity within clusters and the distinct separation between them.
[0053] B, Cumulative Expression Heatmap of Genes within Metaprograms in Malignant Cells. This heatmap displays the cumulative expression levels of genes within each of the 30 metaprograms (MP 0 to MP 29) across malignant cells. The relative expression levels are color-coded, with red indicating higher expression, blue indicating lower expression, and white representing mid-level expression. Cells are hierarchically clustered based on their expression profiles to reveal potential subpopulations and patterns of co-regulation among metaprograms.
[0054] C, Proportions of Cells Assigned to Each Metaprogram (MP) and Corresponding Cell Counts. The pie chart illustrates the distribution of cells assigned to various MPs, highlighting that MP 21 and MP 2 represent the largest portions at 45.8% and 32.0%, respectively. The accompanying table lists the exact number of cells assigned to each MP, providing a detailed view of the data distribution.
[0055] Figure D-E refers to ITH and Evolution:
[0056] D, MPs per Subclone. X axis, patient donor subclones. Y axis, proportion of malignant cells assigned to MPs. The heatmap illustrates the distribution of second-highest MP (marker profile) percentages for cells, excluding self-assignments and specific MPs. The x-axis represents primary MP assignments, while the y-axis indicates the second-highest MP scores. The color intensity in the heatmap, ranging from blue / purple (lower percentages) to yellow (higher percentages), signifies the percentage of cells for which a particular MP is the second-highest score.
[0057] E, Heatmap of chi-squared test residuals for metaprogram assignments within patient subclones. Each row displays the residual values for a specific patient's subclones, highlighting differences between observed and expected MP assignments. Positive residuals signify MPs that are over- represented in a subclone, while negative residuals denote under-represented MPs. The heatmap illustrates the correlation between MPs across all malignant cells in the Extended GBMap.
[0058] Figure 2 refers to Developmental Brain MPs and their Similarity to Malignant MPs: Heatmap of Jaccard Similarity Indices: Metaprograms vs Cortical2023, Fetal, and Hallmark Metaprograms. The heatmap represents the Jaccard similarity indices between gene sets of various metaprograms and clusters from Cortical2023, Fetal, and Hallmark datasets. The color gradient ranges from white (low similarity) to dark purple (high similarity), highlighting areas of significant overlap between the gene sets.
[0059] Figure 3 refers to Similarities Between GB Metaprograms: Network diagram of overlap (edges) between MPs (nodes), as determined via the top genes of each MP. A connection was drawn between a MP if there were 5 or more overlapping genes. Bold lines indicate 10+ genes, normal lines indicate 5-9 genes. Overlapping genes among malignant metaprograms. The figure illustrates significant overlaps (10% or more overlap, >5 genes) between various metaprograms. The overlaps are represented with connecting lines, where the thickness of each line is proportional to the number of overlapping genes. Different colors are used to distinguish the comparisons between malignant metaprograms.
[0060] Figure 4 refers to Organoid Assessment:
[0061] A. UMAP of Malignant Cells by Condition, Sample, and Timepoint. The first plot shows UMAP of malignant cells color-coded based on their condition: 314 (blue), 520 (orange), Cancer Spheroid (green), and Cancer Spheroid Matrigel Embedded (red). The second plot displays the UMAP of malignant cells by sample, with each sample represented by a different color, as indicated in the detailed legend. The third plot illustrates the UMAP of malignant cells by timepoint, with timepoints 1.0, 2.0, 3.0, and 4.0 represented in blue, orange, green, and red respectively.
[0062] B. Heatmap of Jaccard similarity indices for organoid metaprograms with cluster annotations. This panel shows the pairwise Jaccard similarity indices between metaprograms inferred from individual organoid samples. Each row and column corresponds to a metaprogram derived from nonnegative matrix factorization (NMF) of organoid single-cell RNA-seq data. Metaprograms are grouped into clusters based on their pattern of mutual overlap, and both rows and columns are colored according to their assigned cluster (here, Clusters 1-6). The color scale encodes the Jaccard index between metaprogram gene sets, with darker shades indicating higher overlap. The colored side bars and accompanying legend indicate cluster membership; detailed mappings of metaprograms to organoid samples and cluster IDs are provided in the corresponding table rather than as axis labels in the drawing.
[0063] C. Heatmap of Jaccard Similarity Indices: Organoid Metaprograms vs. GB Patient Metaprograms and Hallmarks. This heatmap visualizes the Jaccard similarity indices between various clusters of organoid metaprograms and the metaprograms derived from GB patient samples as well as hallmark gene sets. The intensity of the color represents the degree of similarity, with darker shades indicating higher similarity. The rows represent different metaprograms and hallmarks, while the columns represent different organoid clusters.
[0064] D. UMAP of Malignant Cells Colored by Assigned Gene Metaprogram (GB MP). This plot visualizes the clustering of malignant cells based on their gene expression profiles. Each point represents a single cell, colored according to its assigned GB MP. The distinct clusters indicate the differentiation of malignant cells into specific metaprograms, highlighting potential functional differences.
[0065] E. UMAP of Malignant Cells Colored by Assigned Organoid Metaprogram (MP). This plot displays the clustering of malignant cells based on their gene expression profiles, with each cell colored according to its assigned organoid MP. The clusters indicate the potential differentiation and functional states of the cells within organoid structures.
[0066] F. Comparison of Correlation of Metaprogram Scores with Timepoint between Conditions 314, 520, and Cancer Spheroid. The bar plot shows the Spearman correlation coefficients of metaprogram (MP) scores with timepoint for three different conditions: 314 (blue), 520 (green), and Cancer Spheroid (yellow). Each bar represents the correlation for a specific MP, with positive correlations indicating an increase in MP score over time and negative correlations indicating a decrease. Significant correlations (p < 0.05) are marked accordingly.
[0067] Figure 5 refers to Assessment of Metaprogram Dynamics across Time and Condition. A, Correlation between Metaprograms across Malignant cells in the Organoid Model. (A) All cells. The correlation matrix displays the relationships between the expression patterns of the thirty metaprograms (MP 0 to MP 29) across malignant cells. The correlation values range from -1 to 1, where positive values (closer to 1) indicate a positive correlation, and negative values (closer to -1) indicate a negative correlation. B, 520 T1 Assessment of Metaprogram Dynamics across Time for Condition 520 Tl. The correlation matrix displays the relationships between the expression patterns of the thirty metaprograms (MP 0 to MP 29) across malignant cells. The correlation values range from -1 to 1, where positive values (closer to 1) indicate a positive correlation, and negative values (closer to -1) indicate a negative correlation,
[0068] C, 314 T2 Assessment of Metaprogram Dynamics across Time and Condition 314 T2. The correlation matrix displays the relationships between the expression patterns of the thirty metaprograms (MP 0 to MP 29) across malignant cells. The correlation values range from -1 to 1, where positive values (closer to 1) indicate a positive correlation, and negative values (closer to -1) indicate a negative correlation. D, Cancer Spheroid Tl: Assessment of Metaprogram Dynamics across Time and Condition Cancer Spheroid Tl. The correlation matrix displays the relationships between the expression patterns of the thirty metaprograms (MP 0 to MP 29) across malignant cells. The correlation values range from -1 to 1, where positive values (closer to 1) indicate a positive correlation, and negative values (closer to -1) indicate a negative correlation.
[0069] Figure 6 refers to Quantitative Analysis of Segmented 3D Organoid Data.
[0070] A-D, Time series of separate 520F GB organoids extracted. A, At 10 days (COLEC1_D11), B, 17 days (COLEC2_A2), C, 24 days (COLEC3_W3), and D, 38 days (COLEC4_W5). In all cases, the GBSC line was U3013 from the Human Glioblastoma Cell Culture resource. The U3013 cells are lentivirus transduced cells expressing ZsGreen driven by an EFla promoter. DAPI stained nuclei are shown in blue, ZsGreen expression is shown in yellow, SOX2 IF staining in pink and TBR1 staining in green.
[0071] E, Median number of cells per time point for two conditions (Condition 314 and Condition 520) with standard deviation. The x-axis represents sequential time points, and the y-axis displays the median number of cells. Error bars indicate standard deviation, capturing variability within each condition at each time point. Condition 314 is represented in blue, and Condition 520 is represented in orange.
[0072] F, Mean distance and standard deviation of distances measured over four time points (Tl, T2, T3, and T4). The blue line represents the mean distance, while the orange line represents the standard deviation of distances. Both metrics are plotted on the same y-axis scale to illustrate relative changes across the time points., G, Temporal trend of density across four time points. The plot indicates a sequential decrease in density values, with data points connected by a line to highlight the trend over time,
[0073] H, Line plot showing the temporal trend of nearest neighbour distance across four time points,
[0074] I, Line plot showing the temporal trend of compactness across four time points (1 to 4).
[0075] Figure 7 (Supplementary Fig. SI) refers to Metaprogram Identification and Validation of Approach Against Neftel et al 2019 (An Integrative Model of Cellular States, Plasticity, and Genetics for Glioblastoma. Cell. 2019 Aug 8;178(4):835-849.e21. doi: 10.1016 / j. cell.2019.06.024. Epub 2019 Jul 18. PMID: 31327527; PMCID: PMC6703186.):
[0076] A, Overlap of original Neftel MPs with ours. Network diagram of overlap (edges) between MPs (nodes), as determined via the top genes of each MP. A connection was drawn between a MP if there were 5 or more overlapping genes. Bold lines indicate 10+ genes, normal lines indicate 5-9 genes.
[0077] B, Sankey Plot of GBMap / Neftel Annotations to ours.
[0078] C, Circle plots showing the distribution of the top five metaprograms (MPs) in each donor, based on the highest-scoring MP per malignant cell. Only donors from the extended GBMap with >50 malignant cells are shown. Before assignment, MPs 0, 4, 5 (immune inflammation), 14 (post- transcriptional regulation), 16 (cilia), 19 (DNA damage response), 22 (stress response), 23, and 27 (senescence) were excluded; all remaining MPs are colour-coded consistently across donors. Each mini-circle represents one donor, with wedge size proportional to the fraction of malignant cells assigned to each of that donor's five most frequent MPs. The underlying donor x MP top-5 distributions (both counts and fractions) are provided in Figure7C_donor_top5_MP_counts.csv and Figure7C_donor_top5_MP_fractions.csv, respectively.
[0079] Figure 8 (Supplementary Fig. S2) refers to Similarities Between GB Metaprograms, marked by overlap with Neftel, Hallmarks, Fetal and Cortical MPs. Network diagram of overlap (edges) between MPs (nodes), as determined via the top genes of each MP. A connection was drawn between a MP if there were 5 or more overlapping genes. Bold lines indicate 10+ genes, normal lines indicate 5-9 genes. A, Neftel. B, Hallmarks. C, Fetal D, Cortical.
[0080] Figure 9 (Supplementary Fig. S3) refers to Quality Control of scRNAseq Data. A, Violin plots of quality control metrics before filtering. Each dot represents a cell. From left: Number of detected genes, total number of counts, percent mitochondrial gene counts, percent ribosomal counts, percent spike-in counts, percent hemoglobin counts. B, Predicted doublet detection results. Scatterplot of number of detected genes vs total counts per cell. The color indicates whether the cell is a predicted doublet or not, using the Scrublet algorithm in Scanpy.
[0081] Figure 10 (Fig. S4) refers to time series of separate 520F GB organoids extracted. A, At 10 days (COLEC1_D11), B, 17 days (COLEC2_A2), C, 24 days (COLEC3_W3), and D, 38 days (COLEC4_W5). In all cases, the GBSC line was U3013 from the Human Glioblastoma Cell Culture resource. The U3013 cells are lentivirus transduced cells expressing ZsGreen driven by an EFla promoter. DAPI stained nuclei are shown in blue, ZsGreen expression is shown in yellow, SOX2 IF staining in pink and TBR1 staining in green. Figure 11 (Supplementary Fig. 8) refers to 3D image analysis pipeline for GB Organoid COLEC4_W2, where A, the image is normalized with background correction to enhance contrast, B, cell segmentation performed with the Voronoi-Otsu-Labeling workflow (a combination of Gaussian blur, spot detection, thresholding and binary watershed), C, cell coordinates (red dots) of malignant cells and the yellow star is site of spheroid placement based on density of cancer cells, D, Network diagram built with Delaunay Triangulation from all cells.
[0082] Figure 12: MPs per Subclone. X axis, patient donor subclones. Y axis, proportion of malignant cells assigned to MPs.
[0083] Figure 13. Residuals ofchi-squared test for metaprogram assignments across genetic subclones. The heatmap shows the residuals from a chi-squared test of independence between assigned metaprograms (MPs) and genetic subclones. Columns correspond to metaprograms (MP 0-MP 29, excluding MP 4), and rows correspond to individual subclones. For readability, subclone labels are not shown on the y-axis; instead, they are listed in the accompanying data file. Each cell represents the difference between the observed and expected number of cells for that metaprogram-subclone combination (residual - observed - expected), with warm colors indicating over-representation and cool colors indicating under-representation relative to the global null model. The underlying numerical values are provided in Figurel3_MP_subclone_residuals_matrix.csv, where each row is a subclone (e.g. P3017_subclonel, C3N-03188_subclone4) in the same order as the rows of the heatmap, and each column is an assigned metaprogram (matching the x-axis labels). The entries are the residual values used to generate the heatmap.
[0084] Figure 14: (A-D) F, G, H, I, Kaplan-Meier survival plots for glioblastoma patients, categorized by high and low expression levels of malignant metaprograms. The survival analysis reveals significant differences between expression groups, with p-values below 0.05. Low expression levels are represented by the cyan line, while high expression levels are shown in red.
[0085] Figure 15: Mean expression and prevalence of Synapse, Invasion, Developmental, and Wound Healing metaprograms across cells assigned to Glioblastoma malignant metaprograms cellular groups. Each row represents a specific MP, while the size of each dot indicates the fraction of cells in each group expressing that MP, as shown by the key on the right. The color intensity reflects the mean expression level within each group. Figure 16: Correlation between Metaprograms across all Patient-derived malignant cells in the Extended GBMap. The correlation matrix displays the relationships between the expression patterns of the thirty metaprograms (MP 0 to MP 29) across malignant cells. The correlation values range from -1 to 1, where positive values (closer to 1) indicate a positive correlation, and negative values (closer to -1) indicate a negative correlation.
[0086] Figure 17: MPs per GB Cell Line. X axis, patient donor subclones. Y axis, proportion of malignant cells assigned to MPs.
[0087] Figure 18 (Flowchart — Pipeline). Figure 18 illustrates the pipeline corresponding to steps (a)-(f) of Claim 1. Steps: organoid cultivation -> treatment / control split tri-modal data capture (timepoints tO...tn) -> fusion digital twin -> prediction validation loop. Figure 19 (Flowchart — Controller loop). Figure 19 illustrates the closed-loop controller referenced in step (f). New data in divergence metric computed if threshold exceeded, adjust combination + update twin -> re-measure repeat.
[0088] Figure 20 (Schematic — Data fusion). Figure 20 illustrates the data fusion process referenced in step (d). Blocks for 3D imaging / RNA-seq / spatial-omics feature extraction -> alignment joint embedding twin parameters.
[0089] Definitions:
[0090] Divergence metric. As used herein, a "divergence metric" is a scalar or vector function that quantifies a difference between predicted behaviour of the digital twin and newly observed data from the patient and / or organoid. Examples include, without limitation: Ku II back— Leibler (KL) divergence between predicted and observed metaprogram distributions, mean-square error (MSE) between predicted and measured viability over time, earth-mover distance between predicted and observed spatial expression patterns, and cosine distance between state vectors representing tumor cell states.
[0091] Threshold criterion. As used herein, a "threshold" or "threshold criterion" denotes a predefined quantitative condition applied to a divergence metric that, when satisfied, triggers an action by the closed-loop controller. The threshold may be defined, for example, as a fixed numeric value, a range, or a percentile-derived cut-off that corresponds to a statistically or clinically significant deviation of the divergence metric from a baseline, training cohort, or atlas reference.
[0092] Digital twin. As used herein, a "digital twin" means a computational representation constructed by integrating at least (i) tissue-cleared 3-D volumetric imaging, (ii) RNA-sequencing, and (iii) spatial- omics readouts from a patient-derived organoid corresponding to a tumor of a subject, the representation parameterized to simulate disease trajectory and therapy response and configured for longitudinal updating based on new organoid and / or patient measurements.
[0093] Longitudinal fusion. "Longitudinal (time-series) fusion" means integration of two or more temporally separated measurement sets from the same patient-derived organoid and / or the patient, producing change-aware parameters (e.g., Ametaprogram scores, viability trend vectors) consumed by the comparator and controller.
[0094] Closed-loop controller. A "closed-loop controller" is a decision module that ingests divergence metrics between predicted and observed tumor behavior, applies prespecified policies or learned models, outputs a therapy recommendation that triggers a concrete action (e.g., organoid dosing change or updated regimen proposal), and updates model parameters for the next cycle.
[0095] Divergence metric. A "divergence metric" is a scalar or vector function (e.g., KL-divergence of predicted vs observed metaprogram distributions; mean-square error on viability / ti me-to-ki 11; CNV- anchored state distance) computed between digital-twin predictions and new measurements.
[0096] Quality assurance. Prior to twin generation, data are filtered by modality-specific thresholds, including but not limited to imaging signal-to-background and segmentation completeness; RNA-seq read depth, mitochondrial read fraction, and cell count; spatial-omics spot / UMI count and coverage. Organoids failing viability or growth criteria are excluded. In some embodiments, imaging quality thresholds are defined to exclude low-quality organoid images. For example, tissue-cleared 3-D images may be required to exhibit a signal-to-background ratio above a predefined value (e.g., >3:1) and a segmentation completeness above a predefined proportion of the organoid volume (e.g., at least 70-80% of the organoid volume successfully segmented). Images that fail these criteria can be discarded or flagged for manual review prior to twin generation. For RNA sequencing data, quality thresholds may include minimum total reads per sample and basic per-cell or per-spot metrics. For example, bulk RNA-seq libraries may be required to have at least about 1-5 x 10A6 mapped reads. Single-cell RNA-seq data may be filtered such that individual cells are retained only if they exhibit, for example, at least about 1,000 detected transcripts (UMIs) and / or at least about 500 detected genes, and a mitochondrial read fraction below a predefined percentage (e.g., less than about 20%). Cells or libraries failing such criteria can be excluded from downstream analysis. For spatial-omics measurements, quality thresholds may include minimum counts per spot or feature and coverage across the organoid. For example, spatial transcriptomics spots may be retained only if they contain at least about 100-200 detected transcripts and at least about 50 detected genes, and if the total number of retained spots per organoid exceeds a minimum coverage threshold selected by the operator. Spots or regions that do not meet these criteria may be masked or removed prior to integration into the digital twin. Viability and growth criteria can also be applied at the organoid level. For example, organoids may be required to reach a minimum size or area (e.g., exceeding a predefined diameter or volume) and / or a minimum viability (e.g., greater than about 60-70% viable cells by a suitable assay) before being included. Organoids that fail to meet these viability or growth thresholds can be excluded from the analysis to ensure that the digital twin is constructed from biologically representative samples. The specific numerical values and ranges mentioned above are merely illustrative examples of thresholds that can be used to obtain data of sufficient quality for constructing the digital twin. A person skilled in the art will appreciate that appropriate threshold values can be adapted to the particular sequencing platform, imaging modality, tumor type, or organoid protocol employed, and may be tuned empirically. More stringent or more permissive thresholds may be used without departing from the scope of the invention.
[0097] Combination-therapy efficacy. A candidate combination is prioritized when organoid readouts demonstrate (i) reduction or stabilization of malignant metaprogram scores and / or (ii) suppression of emergent therapy-resistant subclones across measured time points.
[0098] Atlas alignment. The digital twin can be embedded into a reference atlas to identify nearest-neighbor patients or tumor subtypes. The divergence metric quantifies deviation from atlas baselines and informs when bespoke combinations are required.
[0099] Software embodiment. The software instructions implement the steps of organoid data ingestion, tri- modal or bi-modal integration, twin synthesis, longitudinal comparison, divergence computation, and closed-loop recommendation updates.
[0100] As used herein, the term "non-transitory computer-readable medium" refers to any tangible storage device or medium that can store instructions for execution by one or more processors. The medium is non-transitory in that it does not consist of a signal per se and excludes carrier waves or other ephemeral forms of transmission. Examples of such media include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks), optical storage devices (e.g., CDs, DVDs), solid-state storage devices (e.g., flash memory, SSDs), and semiconductor memory devices (e.g., RAM, ROM).
[0101] The non-transitory computer-readable medium may store instructions that, when executed by one or more processors, cause the processor(s) to perform operations comprising:
[0102] • Receiving patient-specific tumor data, including genomic, transcriptomic, proteomic, and / or imaging data; • Generating a patient-derived organoid model representing the biological characteristics of the patient's tumor;
[0103] • Constructing a digital twin of the patient's tumor based on the organoid model and patientspecific data;
[0104] • Simulating tumor progression and therapy response within the digital twin under different treatment regimens;
[0105] • Analyzing simulation outputs to identify an optimized cancer therapy tailored to the patient.
[0106] The instructions may be implemented in any suitable programming language and may be stored in any format that allows execution by a computing system. The computing system may include one or more processors, memory, and interfaces for data input and output.
[0107] As used herein, the term "spatial omics" refers to any analytical approach that combines molecular profiling with spatial context within a biological sample, enabling the mapping of biomolecules (e.g., RNA, DNA, proteins, metabolites) to their precise locations in tissue or organoid structures. Spatial omics techniques provide both quantitative molecular data and positional information, thereby preserving the architecture and microenvironment of the sample.
[0108] Examples of spatial omics include, but are not limited to:
[0109] • Spatial Transcriptomics, in which gene expression is measured across tissue sections or organoids while retaining spatial coordinates, allowing visualization of transcriptomic patterns within the three-dimensional structure. In the present embodiments, spatial transcriptomics (optionally together with immunofluorescent marker imaging) is a primary example of spatial omics used in the organoid-derived digital twin.
[0110] • Other spatially resolved omics may be used analogously, for example spatial proteomics (mapping protein abundance and localization using multiplexed imaging or imaging mass spectrometry) or spatial epigenomics (profiling chromatin accessibility or DNA methylation patterns in situ) where such assays are available for the sample type. The term "synergistic effect" refers to a combined therapeutic effect greater than the sum of individual effects, as measured by viability reduction or metaprogram suppression.
[0111] Detailed description of the invention
[0112] Glioblastoma multiforme (GB) is the most common type of primary malignant brain tumor, in addition to being the deadliest. Overall median survival remains at 11-15 months with standard treatment. The conventional treatment pathway of surgery, radiotherapy and chemotherapy is drastically limited by the high rates of recurrence, resistance to therapy and neurological deterioration seen in patients. As of 2018, there are an estimated 297,000 new cases of brain and nervous system tumors worldwide per year, with glioblastoma multiforme (GB) accounting for approximately 50% of them, leading to approximately 150,000 cases of GB worldwide in 2018. Of all cancer types, GB has seen the least improvement in mortality rates over the past several decades. In recent years, new technologies for treating glioblastomas such as gene therapy, focused radiation therapy, immunotherapy, as well as new drugs are being developed for eventual use and are being tested in experimental trials. However, these early-stage technologies have boosted patient's survival rates by a median of three months, seeing little overall change in patient outcome. Clearly, improvements are badly needed and merit a new approach to validate potential treatments for this type of brain cancer.
[0113] With incidence in the EU and USA at 2-3 per 100,000 people, GB is an orphan disease according to regulatory governing bodies and eligible for the accelerated approval pathways assigned to orphan diseases. Despite this, only 6.2% of orphan drug development projects reach the market. This is due to the low approval rate of orphan drugs in oncology. Most of the oncologic orphan drugs fail in efficacy (Phase II) trials, with only 2.8% of orphan oncology drugs making it past Phase II trials. With that being said, the same study found that oncology trials using biomarkers had higher success rates (i.e. oncology drugs paired with biomarkers had a probability of success of 10.7% of making it through clinical trials (Phase l-lll), while oncology drugs that had no biomarker had a probability of success of 1.6%).
[0114] Standard efforts in the area of GB research center on assessing new mechanisms in cancer proliferation and developing novel compounds. But these efforts are hamstrung by the limitations of the tools used. Current clinical assessment tools are limited because they are either one-dimensional (cell culture) and lack realistic complexity or do not recapitulate human tumoral biology (mouse models / xenografts), which leads to therapies that are not developed in the right pathophysiological context.
[0115] A paradigm of cancer research is of intratumoral heterogeneity (ITH); that cancers are assemblies of heterogeneous populations of various distinct cell types, genetic compositions and phenotypes. ITH further compounds the difficulties in establishing efficacious treatments, as the variability between individual patients' tumor profiles (intra- and intertumoral heterogeneity) confounds current treatment and efforts to establish common clinical biomarkers. Awareness of this factor is critical in diagnosing patients, devising treatments, and developing new therapies.
[0116] Efforts to classify the nature and features of brain tumors are historical and ongoing. Starting in 1858, Virchow postulated that some cancers could originate from developmental-stage tissues. The first attempt of brain tumor classification was published by Bailey and Cushing in 1926, based on the histological appearance of different types of immature neural cells within the tumors. In the transcriptomics age, a subtyping system for glioblastoma based on bulk RNA-seq data from the TCGA glioblastoma dataset by Verhaak et al 2010 (Cancer Genome Atlas Research Network. Integrated genomic analysis identifies clinically relevant subtypes of glioblastoma characterized by abnormalities in PDGFRA, IDH1, EGFR, and NF1. Cancer Cell. 2010 Jan 19;17(l):98-110. doi:
[0117] 10.1016 / j.ccr.2009.12.020. PMID: 20129251; PMCID: PMC2818769.) revealed four subtypes (Proneural, Neural, Classical and Mesenchymal). Further assessment via single-cell RNA sequencing (scRNAseq) classified malignant glioblastoma cells phenotypes based on their expression patterns similar to astrocytes, mesenchymal, oligodendrocyte-, and neural precursor cells (AC-like, MES-like, OPC-like, and NPC-like, respectively). Furthermore, the study established the plasticity of said glioblastoma stem-like cells (GBSCs) suggesting these tumors reflect a reversion to an embryonic cell state. The work of the inventors utilizing spatial transcriptomics to profile glioblastoma architecture in patient sections further contextualizes the localization of single-cell malignant phenotypes within the tumor microenvironment (TME), and the niches that can form within patients.
[0118] With advances in imaging and next-generation sequencing (NGS) techniques, it is now possible to better characterize new models by combining methods developed within the past decade but burdened with their own limitations. 3D imaging can accurately depict the spatial distribution of cancer cells, but lack a means of subtyping the cancer cells. scRNAseq, which examines the transcriptional information within individual cells, is limited in that samples need first to be dissociated to form single cell suspensions, thereby losing all spatial distribution and potential interaction between cancer cells and the tumor microenvironment inside the tumor. Moreover, little of the mutational burden of each cell is recovered. Spatial transcriptomics provides the spatial mapping of genetic expression guided and targeted by specifically designed probes; however, it is necessary to know beforehand what genes and mutations are going to be examined before running the experiment. The intention of the inventors is to combine these three approaches.
[0119] Used together, scRNAseq provides the targeting system for spatial transcriptomics while 3D imaging can guide the selection of zones for spatial transcriptomics. By combining the approaches, a new picture of the GB tumor emerges; that of phenotype contextualized with the TME, and the distribution and phenotype can be determined.
[0120] Efforts utilizing cerebral organoids to model various cancers, viral infections, developmental and neurological disorders are in the early stages of application and represent a new frontier for understanding and developing treatments. In particular, the development and usage of organoids to model GB have seen a spike in recent years, with applications ranging from assessing oncogene expression and mutation, tumor infiltration, and intra- and inter- tumor heterogeneity. Current cerebral organoid models still lack some features of the GB ecosystem, such as vasculature, immune cells or blood-brain barrier function. What we get in return as opposed to an animal environment, is a 3D human model that is easily accessible for permutations, and the retention of enough biologically relevant features of GB to be relevant in terms of the applicability of organoid models for drug testing, they are already being adopted and used to achieve real clinical improvements.
[0121] The organoids themselves generate hypoxic niches and necrotic centers (both clinical features of GB) due to the limit of diffusion of oxygen and nutrients as the organoids grow. The GB organoid models retain phenotypic heterogeneity among GBSCs, reactively upregulate genes seen in clinical GB subtypes, and retain the differential sensitivities to chemotherapy and radiation seen in GB patients. In terms of their applicability for developing personalized therapies and treatments, GB organoids are increasingly made using patient-derived GBsc's. GB organoid models are ready to be used for a variety of purposes that were previously unsatisfied by 2D cell culture and animal models.
[0122] What remains is to link a patient's clinical features to the models and assess the extent to which the model recapitulates; this project aims to harness and validate the current state of GB organoid models in relation to real tumors, for the purpose of gaining ground in the efforts to decrypt patient status and better map out the progression of this disease and connect them to clinically relevant treatments.
[0123] The current disclosure leverages both large-scale scRNA-seq data from GB patients and GB organoid (GO) models to characterize GB's transcriptional diversity and model its environmental interactions. By integrating patient-derived transcriptional profiles with organoid-based insights, this research aims to deepen the understanding of GB's adaptive landscape and assess the potential of organoids to serve as patient-relevant platforms for studying GB biology. This dual approach seeks to uncover new avenues for therapeutic exploration by addressing both the intrinsic complexity of GB and the critical role of its environment in influencing tumor behavior.
[0124] The method Therefore, in a first aspect the invention refers to a method for developing a digital twin of an organoid model for optimising cancer therapy for at least one patient diagnosed with cancer, comprising the steps of:
[0125] (a) providing organoids cultivated with patient-derived cancer cells;
[0126] (b) treating a first group of organoids of step (a) with a plurality of different therapeutics at different concentrations and durations;
[0127] (c) maintaining a second group of organoids of step (a) that are untreated to act as controls;
[0128] (d) selecting organoids of step (b) and (c) at a plurality of different time points to obtain comprehensive data, said data comprising tissue clearing data, 3D imaging data, RNA sequencing data and / or spatial transcriptomics data;
[0129] (e) integrating the comprehensive data obtained for the selected organoids to develop a digital twin representing each patient's cancer development; and
[0130] (f) predicting combinations of therapies that can halt or eliminate the cancer development in the patient(s), based on the differing responses of the organoid model to different therapies.
[0131] Hereby, with the inherent flexibility to integrate new data types, predict more personalized therapies, and continuously refine the global glioblastoma atlas the invention's potential in personalized cancer treatment is realized.
[0132] The following variations of features and items of the method of the invention are included, as well as the embodiments listed:
[0133] 1. Variations in the Methodology:
[0134] Organoid Type:
[0135] • Variation in Tissue Sources: The method can be adapted for different organoid models beyond cerebral organoids, such as co-culture models combining glioblastoma cells with brain microvascular endothelial cells to simulate the blood-brain barrier or integrating immune cells to study tumor-immune interactions.
[0136] • Additional Cancer Types: Although the primary focus is on glioblastoma, the method may also be applied to other cancer types, such as breast cancer or lung cancer, by modifying the organoid model to include relevant tumor microenvironments.
[0137] Therapeutic Agents:
[0138] • Broadening of Therapeutics: In addition to small molecules and biologies, the method incorporates other therapeutic modalities such as gene editing technologies (e.g., CRISPR), immunotherapies (e.g., CAR-T cells), or oncolytic viruses. This would involve specific protocols for delivering these agents into the organoids.
[0139] • Multi-modal Treatment Regimens: A variation in the method involves combining multiple types of therapies (e.g., a small molecule inhibitor with immunotherapy) and analyzing their synergistic effects on the organoids.
[0140] Data Collection Approaches:
[0141] • Integration of Epigenomic Data: Beyond RNA sequencing and spatial transcriptomics, the method incorporates epigenomic data such as DNA methylation profiles or chromatin accessibility (ATAC-seq) to provide further insights into the regulatory mechanisms driving glioblastoma progression and therapeutic response.
[0142] • Real-Time Monitoring: Real-time, non-invasive monitoring of organoid health and drug response using live-cell imaging techniques or metabolic assays may be introduced to refine and adjust treatments during the experiment.
[0143] 2. Lists of Genes for Glioblastoma Metaprograms:
[0144] • Metaprograms Identified from 140 Glioblastoma Patients: o The method may be enhanced by the specific lists of genes associated with distinct metaprograms (groups of co-expressed genes) that were identified from a cohort of 140 glioblastoma patients.
[0145] • Normal Reference Metaprograms for Comparison: o To differentiate between malignant and non-malignant metaprograms, a list of reference genes representing normal brain functions or developmental processes are included. For example, the Cortical and Fetal sheets I uploaded include the metaprograms derived from normal cells in fetal developing brain and young cortical brain. o The normal reference programs are dynamically derived and can be extracted as necessary from relevant datasets, ensuring adaptability to new data while avoiding a rigid framework.
[0146] • Patient-Specific vs. Population-Wide Metaprograms: o Specific examples of patient-specific metaprograms that could diverge from population-wide trends are provided. For instance, some patients may show amplification of growth factor pathways not commonly upregulated in other patients. A comparison of the gene expression signatures can help in identifying unique features of an individual's cancer.
[0147] 3. Embodiments of Data Integration for Digital Copies:
[0148] Multi-Omics Data Integration:
[0149] • Embodiment 1: Data from RNA sequencing, 3D imaging, and spatial transcriptomics is combined using specialized bioinformatics pipelines. For example, RNA-seq data identifying gene expression profiles can be spatially mapped onto the 3D organoid structure to create a detailed digital model of both gene activity and physical organization.
[0150] • Embodiment 2: The method incorporates additional proteomic data obtained via mass spectrometry, enabling deeper insights into protein expression levels and post-translational modifications in the organoids.
[0151] Advanced Data Analytics and Al Integration:
[0152] • Embodiment 3: Machine learning algorithms may be applied to the integrated dataset to predict which combination of therapies would be most effective for a given patient. For example, algorithms could learn from past therapy responses and use the organoid data to predict patient outcomes based on similar molecular profiles.
[0153] • Embodiment 4: An Al-driven system continually updates and refines predictions based on real-time data from patient-derived organoids and global cancer data, incorporating population-wide trends as well as individual-specific information.
[0154] 4. Additional Variations on Predictive Therapeutic Testing: Personalized Immunotherapy Development:
[0155] • Embodiment 5: The digital twin may be used to simulate immune responses, such as incorporating data from immune checkpoint proteins (e.g., PD-L1) and immune cell interactions with the organoid. This allows personalized testing of immunotherapies like checkpoint inhibitors or CAR-T cell therapies.
[0156] Dynamic Drug Testing and Adaptation:
[0157] • Embodiment 6: In addition to predicting combinations of therapies that halt cancer progression, the system may monitor resistance development in real-time and recommend adaptive therapy adjustments (e.g., adding a secondary drug if resistance signatures are detected).
[0158] 5. Embodiment of Global Glioblastoma Atlas Integration:
[0159] • Embodiment 7: Patient-derived data from the method may be added to a global glioblastoma atlas that tracks treatment responses and gene expression profiles across a large cohort. For example, when a new patient's organoid is analyzed, its data is compared to thousands of previous cases to identify patterns that align with known successful therapies.
[0160] • Embodiment 8: The global atlas could be further refined to create subtype-specific atlases, allowing the identification of commonalities among patients with specific glioblastoma subtypes (e.g., EGFR-amplified tumors vs. IDHl-mutated tumors), aiding in the personalization of treatment strategies based on molecular subtypes.
[0161] Metaprograms
[0162] In the context of this disclosure, "metaprogram" refers to groups of co-expressed genes that are identified in the patients studied. See Example section for more details of co-expressed genes for metaprograms.
[0163] Score genes
[0164] In score_genes (see e.g. Fig. 1) the score is calculated in relation to a randomly selected baseline. In this case (cumulative expression), it's just adding up the expression of each gene per metaprogram and seeing what it is in each cell. In this case, it's better to use cumulative to see if the metaprograms are expressed at all, and if they clearly mark cell groups. It also helps distinguish MPs that are more 'functional', e.g see moderate expression in most cells with high in others, versus 'lineage' that see negative expression (by taking the log of the cumulative scores) relative to others. Meaning, certain lineages are unlikely to be active simultaneously.
[0165] Therapeutics
[0166] By therapeutics is, in the context of this disclosure, referred to any agent that interferes or kills cancer cells. Thus, the present invention may be used as a setup for testing any therapeutic and therapeutic combination, whether it is a traditional one (Stupp regimen, chemotherapy + radiotherapy or Temazolomide) or not-so-traditional. For example, temazolomide as a DNA alkylating agent is a commonly used therapeutic in GB and is included in the scope of the invention.
[0167] Thus, the present invention has a broad adaptability for diverse therapeutic interventions.
[0168] With regard to concentrations and durations used for the therapeutics used in the present invention, any concentration in the range of 1 nM to IM, such as 1 nM, 10 nM, 100 nM, 1 uM, lOuM, lOOuM, ImM, 10 mM, 100 mM, IM and any duration from a few seconds to months, such as Is, 1 minute, 1 hour, 1 day, 1 week, 1 month would be applicable, depending on the therapeutic and other varying conditions. Typically, a reason for the usefulness of the method and system of the present invention is to assess longer-term systemic effects of treatments on the model system, which is a broader range than other models are capable of sustaining. Thus, time may need to be accounted for making dynamic adjustments. More specifically, the following would typically apply:
[0169] 1. Concentration Limits:
[0170] Lower Limit:
[0171] Biologically Active Dose: The minimum concentration should be sufficient to elicit a detectable biological effect in the organoid model, such as inhibition of cancer cell growth or induction of cell death. Concentrations below this threshold may not produce meaningful results.
[0172] Example Range: For small molecule therapeutics, the lower limit is typically in the range of 1 nM to 10 nM, depending on the potency and mechanism of the drug.
[0173] Upper Limit:
[0174] Toxic Dose for Organoids: The upper limit should be below the concentration where the drug induces widespread toxicity or causes non-specific cell death across both malignant and healthy cells. Concentrations beyond this can obscure the specific therapeutic effects.
[0175] Example Range: For small molecules, concentrations may reach up to 1 pM to 10 pM, but highly potent compounds may exhibit toxicity at lower concentrations. For biologies like monoclonal antibodies, upper limits can be between 100 ng / mL and 10 pg / mL.
[0176] Consideration: Larger biomolecules may require higher concentrations due to their molecular size and mechanism of action, often necessitating higher dosing ranges compared to small molecules.
[0177] 2. Duration Limits:
[0178] Lower Limit:
[0179] Minimal Therapeutic Effect Window: The minimum duration should allow sufficient time for the drug to exert its intended biological effect, such as inhibiting key signaling pathways or initiating cell death mechanisms. Shorter durations might not capture the full cellular response.
[0180] Example Duration: For fast-acting therapies like chemotherapeutic agents, a short duration of 24 to 48 hours can often produce observable effects. However, for slower-acting drugs or biologies, a duration of 72 hours might be required to capture significant changes.
[0181] Upper Limit:
[0182] Saturation of Response or Non-specific Effects: The upper duration limit should account for the point at which the therapeutic effect plateaus, or where prolonged exposure leads to non-specific toxicities or a breakdown in organoid structure. Extended treatment should avoid cumulative toxicity or degradation of the model.
[0183] Example Duration:
[0184] For most in vitro drug testing on organoids, 7 to 14 days is often a reasonable duration for observing sustained drug effects and assessing how the cells adapt or resist treatment.
[0185] For chronic studies, where long-term drug effects are desired, the upper duration can be extended up to 4-8 weeks (28 to 56 days) or even longer, depending on the stability of the organoids and the need to simulate prolonged therapeutic exposure. During these extended studies, intermittent dosing (e.g., weekly) can be applied to mimic clinical treatment cycles and reduce non-specific toxicity. 3. Dynamic Adjustments:
[0186] Real-Time Monitoring: Incorporating real-time assessments like live-cell imaging, metabolic assays, or transcriptomic analyses can help adjust duration and concentration dynamically based on how the organoids are responding to treatment. This enables a tailored approach to each experimental system, maximizing the effectiveness of the drug while minimizing unwanted side effects.
[0187] Specific Considerations for Organoids:
[0188] Organoid Size and Maturity: Larger, more complex organoids, or those representing tissues like the brain or liver, may require longer treatment durations and / or higher drug concentrations to penetrate deeply into the tissue and elicit responses across all cell types.
[0189] Drug Solubility and Stability: Some drugs may degrade or lose potency over time in culture, necessitating more frequent replenishment or dosing adjustments during extended treatments.
[0190] Tissue-Specific Drug Response: Different tissue types may require varying concentrations and durations due to differences in drug absorption, metabolism, and interaction with the extracellular matrix in the 3D organoid structure.
[0191] Conclusion:
[0192] Concentration Limits: Typically, 1 nM to 10 pM for small molecules, and 100 ng / mL to 10 pg / mL for biologies. Adjustments may be needed based on the drug's potency and the specific biological context.
[0193] Duration Limits: 24 hours to 7-14 days for short-term studies, but for chronic studies, durations can extend up to 4-8 weeks (28 to 56 days) or longer, depending on the stability of the organoid model and the therapeutic goals.
[0194] This extended range for chronic studies allows for the modelling of long-term drug effects, including potential resistance mechanisms, adaptive cellular responses, and the prolonged impact on organoid viability and structure.
[0195] With regard to selecting organoids of step (b) and (c) at a plurality of different time points to obtain comprehensive data,, the different "time points" may for example be at least two and up to several hundred, such as 2, 5, 10, 25, 50, 100, 250, 500.
[0196] In general, therapeutics are tested across biologically relevant ranges of concentrations and durations, tailored to individual compounds.
[0197] In the context of this disclosure, the method is applicable to a broad range of interventions (e.g., temozolomide for chemotherapy, radiotherapy for radiation, CAR-T cells for biologica Is, and electrical fields for procedures).
[0198] Interventions
[0199] The present invention is applicable for a large variety of interventions, such as drugs, radiation therapies, biologicals, procedures, diets, training programs, surgeries, evaluations etc, of different kinds and for different purposes, as well as combinations thereof.
[0200] Comprehensive data
[0201] "comprehensive data" would typically be referred to as a multidimensional dataset that captures various aspects of the organoid's biological state and response to treatments. Specifically, it includes the following types of data: 1. Tissue Clearing Data: Information obtained through tissue clearing techniques that allow for the visualization and analysis of the organoid's internal structures in three dimensions. This data provides insight into the spatial organization of cells, the integrity of tissue architecture, and the presence of specific cellular or structural markers throughout the organoid.
[0202] 2. 3D Imaging Data: High-resolution imaging data, such as that obtained from confocal or lightsheet microscopy, which provides detailed three-dimensional reconstructions of the organoid. This data helps in assessing the morphology, cell-cell interactions, and the formation of microenvironments within the organoid model.
[0203] 3. RNA Sequencing Data (scRNA-seq or Bulk RNA-seq): Transcriptomic data that captures gene expression profiles at either the single-cell level (scRNA-seq) or the bulk tissue level (bulk
[0204] RNA-seq). This data reveals the activity of genes and pathways within the organoid, highlighting how different cell types respond to treatment at a molecular level.
[0205] 4. Spatial Transcriptomics Data: Data that combines gene expression information with spatial context, showing where specific genes are being expressed within the tissue. This data provides insights into how the spatial arrangement of cells and their microenvironment influences gene expression patterns and treatment responses.
[0206] 5. Cellular and Molecular Profiling: Additional profiling data such as protein expression (via immunohistochemistry or mass spectrometry), metabolomics, or epigenetic modifications. This data adds another layer of understanding to the cellular and molecular changes occurring within the organoid.
[0207] 6. Time-Resolved Data: Data collected at multiple time points during the experiment, allowing for the observation of dynamic changes in the organoid over time. This includes tracking the progression of treatment responses, changes in cell populations, and the evolution of the cancerous state within the organoid.
[0208] 7. Quantitative Analysis of Treatment Effects: Measurements of treatment efficacey, including metrics such as cell viability, proliferation rates, apoptosis, and other phenotypic changes. These quantitative assessments provide a direct comparison between treated and control organoids across different time points.
[0209] By integrating these diverse types of data, "comprehensive data" provides a holistic view of the organoid's biological state and its response to various treatments. In general, the invention supports integration of transcriptomic, proteomic, and spatial data for holistic analysis of organoid responses.
[0210] With regard to the comprehensive data, it would be possible in accordance with the invention to apply a limited (minimum) approach as well as an enhanced or even full approach, i.e. including less or more data. The following example approaches may be provided:
[0211] 1. Minimum Viable Data Set:
[0212] • RNA Sequencing Data (scRNA-seq or Bulk RNA-seq):
[0213] • This remains crucial for understanding gene expression profiles and how different cell types within the organoid respond to treatment.
[0214] 3D Imaging Data (Including Tissue Clearing): • Essential for visualizing the spatial organization, morphology, and internal structures of the organoid. This provides important context for how treatments affect the architecture and cell interactions within the organoid.
[0215] 2. Enhanced Data Set:
[0216] • Spatial Transcriptomics Data:
[0217] • This adds the spatial dimension to the gene expression data, helping to link molecular changes with specific locations within the organoid. It's particularly valuable if the focus is on understanding microenvironments or heterogeneity within the tumor.
[0218] 3. Comprehensive Data Set:
[0219] • All of the Above:
[0220] • Combining RNA sequencing, 3D imaging (including tissue clearing), and spatial transcriptomics will provide a thorough and nuanced view of both the molecular and structural aspects of the organoid. This approach is ideal for creating a highly detailed and accurate digital twin.
[0221] Digital twin
[0222] By "developing a digital twin" is referred to the creation of a computational model or simulation that accurately represents the biological state and behavior of the patient-derived organoid based on the comprehensive data collected.
[0223] The following are required: integration of multimodal data, creation of a virtual 3D model with molecular simulation, dynamic simulation of treatment responses, validation and iterative calibration against experimental data, and user interface tools for visualization and analysis.
[0224] 1. Data Integration:
[0225] • Multimodal Data Integration:
[0226] • The various types of data collected (e.g., RNA sequencing, 3D imaging, spatial transcriptomics) must be integrated into a unified model. This involves aligning and correlating the different datasets so that spatial, molecular, and structural information is cohesively represented.
[0227] • Normalization and Preprocessing:
[0228] • Raw data from different sources may need to be normalized to account for differences in scale, noise, or measurement techniques. Preprocessing steps might include data cleaning, dimensionality reduction, and alignment of spatial and molecular data.
[0229] 2. Computational Modeling:
[0230] • Creation of a Virtual 3D Model:
[0231] • Using the 3D imaging data (including tissue clearing), a virtual three-dimensional model of the organoid is constructed. This model represents the spatial organization and architecture of the organoid, including cellular structures and tissue morphology.
[0232] • Molecular Simulation: • Incorporating RNA sequencing and spatial transcriptomics data to simulate gene expression and molecular pathways within the virtual 3D model. This allows the digital twin to reflect the dynamic molecular state of the organoid.
[0233] • Dynamic Simulation of Treatment Response:
[0234] • The model should simulate how the organoid would respond to different treatments over time, based on the collected data. This might involve modeling cell proliferation, apoptosis, and other phenotypic changes in response to the therapies applied in the wet-lab experiments.
[0235] 3. Validation and Calibration:
[0236] • Model Validation:
[0237] • The digital twin must be validated against the actual experimental data to ensure its accuracy. This could involve comparing predicted treatment outcomes with observed results from the physical organoid experiments.
[0238] • Iterative Calibration:
[0239] • The digital model may require iterative adjustments and refinements to accurately reflect the biological reality of the organoid. This might involve tuning parameters or incorporating additional data until the model's predictions closely match experimental observations.
[0240] 4. User Interface and Data Accessibility:
[0241] • Visualization Tools:
[0242] • The digital twin should include tools for visualizing the 3D structure, molecular data, and treatment responses. This could involve interactive 3D models or dashboards that allow researchers to explore the digital organoid and its predicted behaviors.
[0243] • Data Accessibility:
[0244] • The model should be accessible for further analysis, including the ability to run in silico experiments (e.g., testing new treatment combinations or dosages) and compare results across different patient-derived organoid models.
[0245] 5. Scalability and Generalizability:
[0246] • Adaptability to New Data:
[0247] • The digital twin should be designed to incorporate new data as it becomes available, allowing for ongoing refinement and adaptation to new therapeutic insights.
[0248] • Generalization to Other Organoids:
[0249] • Ideally, the methods used to develop the digital twin should be applicable to other patient-derived organoids, enabling broader use across different cancer types or patient populations.
[0250] Combination therapy
[0251] To successfully "halt" cancer development, a combination therapy must stabilize or reverse malignant metaprograms, reduce or stabilize cancer cell populations in the 3D model as evidenced by fluorescent reporters, suppress oncogenic pathways without inducing resistance, identify and promote metaprograms associated with long-term survival, and maintain structural and spatial stability within the organoid. The therapy must also prevent cancer cells from adopting pseudonormal states that could allow them to evade treatment. These effects must be sustained over time, ensuring that the cancer does not resume its progression or adopt more aggressive traits. Thus, the present invention has a broad adaptability for diverse therapeutic interventions.
[0252] Minimum Requirements for a Successful Combination:
[0253] Stabilization or Reversal of Malignant Metaprograms:
[0254] • Inactivation or Downregulation of Malignant Metaprograms: The therapy successfully downregulates or inactivates key metaprograms associated with cancer progression, invasion, or resistance. This is evidenced by a significant reduction in the expression of gene sets within these metaprograms, as observed through RNA sequencing and spatial transcriptomics.
[0255] • Monitoring for Pseudo-Normal Metaprograms: The therapy is assessed for its potential to induce metaprograms that mimic normal cellular functions, which may indicate that cancer cells are adopting a "hiding" strategy. Detection of such metaprograms is crucial, as it may signal that the cancer cells are evading therapeutic pressure by shifting to a less detectable state.
[0256] Quantitative Reduction or Stabilization of Cancer Cell Populations:
[0257] • Fluorescent Reporter Analysis in 3D: Image analysis of the 3D data, utilizing fluorescent reporters, shows a significant reduction in the number of cancer cells over time. This reduction indicates effective inhibition of cell proliferation and possible induction of cell death pathways.
[0258] • Stabilization of Cancer Cell Numbers: Alternatively, the number of cancer cells stabilizes, with no further increase, reflecting a halt in proliferation even if cell death is not predominant.
[0259] Suppression of Metaprograms Linked to Oncogenic Pathways:
[0260] • Down regulation of Metaprograms Driving Oncogenesis: The therapy results in the downregulation of metaprograms that are linked to key oncogenic pathways, such as those involved in cell cycle progression, survival signaling, and metastasis. This downregulation should be consistent across multiple time points, indicating a sustained therapeutic effect.
[0261] • Metaprogram Association with Survival Since certain metaprograms correlate with better or worse outcomes, this means that within a patient there can cancer cells that, if targeted, could lead to better outcomes for the patient and decreased likelihood of recurrence. o We have a process for assessing bulk-RNA samples (validated via the GB-TCGA) for metaprogram expression. o Can see that high vs low expressors of individual metaprograms can predict survival, for metaprograms.
[0262] • No Emergence of Resistance-Related Metaprograms: The therapy does not trigger the activation of metaprograms associated with drug resistance or other aggressive cancer phenotypes, such as stem cell-like behavior or epithelial-mesenchymal transition (EMT).
[0263] Identification of Metaprograms Associated with Long-Term Survival: • Characterization of Beneficial Metaprograms: The therapy should ideally induce or maintain metaprograms associated with long-term survival and disease stabilization, such as those linked to immune response activation, cellular senescence, or controlled differentiation that does not facilitate cancer cell evasion.
[0264] Spatial and Functional Stability in the 3D Model:
[0265] • Maintenance of Structural Integrity: The 3D imaging data should show that the overall structure of the organoid remains stable, with no signs of invasive growth or disruption of normal tissue architecture. This is reflected in the spatial organization of cells and the distribution of fluorescent reporters.
[0266] • Spatial Metaprogram Analysis: Spatial transcriptomics should reveal that the malignant metaprograms do not spread or intensify in specific regions of the organoid, indicating that the therapy has effectively contained the cancerous activity within the model.
[0267] Consistency Across Time Points:
[0268] • Time-Resolved Metaprogram Stability: The criteria above must be met consistently across multiple time points, showing that the malignant metaprograms remain suppressed and that the number of cancer cells does not increase over the duration of the experiment.
[0269] • No Recurrence of Malignant Metaprograms: Even after the cessation of therapy, the malignant metaprograms do not re-emerge, and the 3D imaging analysis continues to show stable or reduced cancer cell populations.
[0270] Biomarkers
[0271] With regard to "biomarkers" to be used in the present disclosure, the 2021 World Health Organization (WHO) Classification of Tumors of the Central Nervous System (CNS 5) includes several molecular biomarkers for gliomas, including:
[0272] • I DH 1 / 2 mutation status: A primary genetic marker for gliomas
[0273] • lp / 19q co-deletion: A primary genetic marker for gliomas
[0274] • H3F3A alterations: A primary genetic marker for gliomas
[0275] • ATRX gene mutations: A primary genetic marker for gliomas
[0276] • MGMT promoter methylation status: A powerful predictive and prognostic marker
[0277] • Loss of CDKN2A: A primary genetic marker for gliomas
[0278] • EGFR amplification: A primary genetic marker for gliomas
[0279] • Combined gain of chromosome 7 and loss of chromosome 10: A primary genetic marker for gliomas
[0280] • TERT promoter pathogenic variants: A primary genetic marker for glioma
[0281] A glioblastoma (IDH-wildtype, H3-wildtype, TERTp mutant, EGFR amplified or chr +7 / 10-) is the primary focus in this disclosure, but there are other clinical definitions for other cancers.
[0282] More specifically, for the approach of the present inventio, inferred copy number variations (CNVs) are typically the ideal method for defining tumor-specific biomarkers due to their significant role in driving the disease's aggressive behavior. CNVs involve changes in the number of copies of specific genomic regions, leading to either amplification or deletion of genes that are critical to tumor growth, survival, and therapy resistance. These CNVs are central to the pathogenesis of glioblastoma and are therefore essential for developing accurate digital copies of organoid models that can predict therapeutic responses.
[0283] Key Glioblastoma-Specific CNVs
[0284] 7p Gain (Chromosome 7p Amplification):
[0285] • Description: One of the hallmark genetic alterations in glioblastoma is the gain of chromosome 7p. This region includes the EGFR gene, which is frequently amplified in GB, leading to overexpression and hyperactivation of the EGFR signaling pathway.
[0286] • Significance: The 7p gain is a defining feature of glioblastoma and a critical biomarker. Amplification of EGFR due to this CNV is associated with enhanced tumor cell proliferation and survival. Targeted therapies against EGFR are often considered in GB cases with 7p amplification. lOq Loss (Chromosome lOq Deletion):
[0287] • Description: The loss of chromosome lOq is another defining CNV in glioblastoma. This region harbors the PTEN gene, a tumor suppressor that regulates the PI3K / AKT signaling pathway. Loss of PTEN due to lOq deletion results in unchecked pathway activation, promoting tumor growth and resistance to apoptosis.
[0288] • Significance: The lOq loss is critical in glioblastoma biology. The deletion of PTEN is associated with poor prognosis and resistance to certain therapies. This CNV serves as a key biomarker for therapeutic decisions, especially in targeting the PI3K / AKT / mTOR pathway.
[0289] Advanced analysis
[0290] With regard to "advanced transcriptomics analysis", the following examples can be provided:
[0291] Implementing Inferred CNVs in the Digital twin for Glioblastoma
[0292] Incorporating inferred CNVs, particularly the 7p gain and lOq loss, into the digital twin of glioblastoma organoid models enhances the precision and predictive power of these models. The process involves several key steps:
[0293] Detection and Inference of CNVs:
[0294] • Sequencing and Analysis: Advanced genomic techniques, such as whole-genome or whole-exome sequencing, are employed to detect CNVs. Bioinformatics tools then infer the presence of CNVs like the 7p gain and lOq loss, identifying these critical alterations in the tumor genome.
[0295] • Specificity to GB: The focus is on identifying CNVs that are characteristic of glioblastoma, particularly those that define the disease, such as the 7p gain (EGFR amplification) and lOq loss (PTEN deletion).
[0296] • Identification of Heterogenous Genetic Profiles in Patients: In a CNV assessment, the individual patients in a cohort may exhibit different CNV patterns across subgroups of patients. The implications of this is that 1) GB may actually be several genomically distinct brain cancers, 2) our method is an approach for identifying them, and 3) any new genomically-defined GB subtypes could have different clinical features and benefit from different treatments. • Metaprogram Correlation with Subclones: We have the ability to identify Subclones from the scRNAseq CNV process (via SCEVAN) and can see that the metaprograms correlate to specific subclones within patients. Since certain metaprograms correlate with better or worse outcomes, this means that within a patient there can exist subclonal lineages that, if targeted, could lead to better outcomes for the patient and decreased likelihood of recurrence.
[0297] Integration into the Digital Model:
[0298] • Data Integration: The inferred CNVs are integrated with other molecular data (e.g., transcriptomics, proteomics) to construct a comprehensive and accurate digital twin of the glioblastoma organoid model.
[0299] • Functional Simulation: The digital model simulates the effects of these CNVs on cellular behavior, including gene expression, signaling pathways, and response to therapies, ensuring precise prediction of therapeutic outcomes.
[0300] Model Calibration and Validation:
[0301] • Validation Against Experimental Data: The digital twin's predictions, especially those related to the impact of CNVs like 7p gain and lOq loss, are validated by comparing them with experimental results from the organoid models.
[0302] • Adjustment for CNV Impact: The model is fine-tuned to accurately reflect how these CNVs influence tumor growth, invasion, and resistance, ensuring it can be effectively used in personalized treatment planning.
[0303] Clinical Relevance and Application:
[0304] • Therapeutic Targeting: The inferred CNVs are used to identify potential therapeutic targets. For instance, EGFR-targeted therapies may be prioritized in cases with 7p gain, while PTEN loss (lOq deletion) may guide the use of therapies targeting the PI3K / AKT / mTOR pathway.
[0305] • Predictive Modeling: The digital twin, incorporating these specific CNVs, allows for the prediction of how glioblastoma tumors will respond to different treatment regimens, aiding in the optimization of personalized therapeutic strategies.
[0306] • Therapeutic Targeting: Inferred CNVs are used to identify therapeutic targets and potential resistance mechanisms, allowing for personalized treatment strategies tailored to the genetic landscape of the patient's tumor.
[0307] Organoid-specific data
[0308] With regard to "organoid-specific data" the examples below can be provided. Of particular note is the potential for modifying and growing other types of organoids to more closely mimic the tumor microenvironment in glioblastoma. For example, adding vascular cells to assess vascularization, or immune cells to assess immune response.
[0309] In addition to neural layer patterns obtained through immunofluorescent staining, other examples of organoid-specific data that can be integrated into the method for developing and using cerebral organoids as models for treatment development include:
[0310] 1. Cellular Composition and Lineage Tracing: • Single-Cell RNA Sequencing (scRNA-seq): This technique can be used to map the diverse cell types present in the organoid, providing detailed information on the cellular composition and the presence of specific cell lineages. It helps in understanding how different cell types contribute to the organoid's overall structure and function.
[0311] • Lineage Tracing Studies: Using genetic markers or reporter systems to trace the development of specific cell populations within the organoid over time, which can provide insights into the differentiation pathways and the stability of cell lineages. Structural Organization:
[0312] • 3D Imaging and Reconstruction: Techniques such as light-sheet fluorescence microscopy or confocal microscopy can be used to visualize and reconstruct the 3D architecture of the organoid, including the spatial arrangement of different cell types, extracellular matrix components, and vasculature.
[0313] • Morphometric Analysis: Quantitative analysis of the organoid's size, shape, and structural complexity, including measurements of cortical folding, lumen formation, and overall organoid morphology. nctional Activity and Electrophysiology:
[0314] • Calcium Imaging: This technique measures calcium influx in neurons to assess functional activity, such as synaptic activity, neural network formation, and responsiveness to stimuli.
[0315] • Electrophysiological Recordings: Techniques like patch-clamp or multi-electrode arrays (MEAs) can be used to record electrical activity in neurons, providing data on the functional maturation and connectivity of neural networks within the organoid. ene Expression Patterns and Spatial Transcriptomics:
[0316] • Spatial Transcriptomics: This technique allows for the mapping of gene expression across different regions of the organoid, providing spatial context to molecular data and revealing patterns of regional specialization or dysfunction.
[0317] • In Situ Hybridization (ISH): Used to visualize the expression of specific genes within the tissue, giving insights into the localization and distribution of key markers associated with neural development or disease states. igenetic Profiling: • Chromatin Accessibility (ATAC-seq): This technique measures the openness of chromatin regions to infer regulatory elements that are active within different cell types or regions of the organoid.
[0318] • DNA Methylation Patterns: Analysis of methylation status at specific gene loci to understand the epigenetic regulation of gene expression within the organoid, particularly in relation to neural differentiation or disease modeling. otein Expression and Signaling Pathways:
[0319] Proteomic Analysis: Mass spectrometry-based proteomics can be used to identify and quantify proteins expressed in the organoid, providing insights into active signaling pathways and cellular processes. Phosphoproteomics: Focuses on the detection of phosphorylated proteins, which are often involved in signaling cascades, to understand the dynamic regulation of cellular functions within the organoid.
[0320] 7. Metabolic Profiling:
[0321] • Metabolomics: Profiling of metabolites within the organoid to assess the metabolic state, including energy production, neurotransmitter synthesis, and lipid metabolism.
[0322] • Live-Cell Metabolic Imaging: Techniques like fluorescence lifetime imaging microscopy (FLIM) can be used to measure metabolic activity in live organoids, providing real-time data on how metabolism is altered under different conditions. 8. Immune Cell Infiltration and Microenvironment Analysis:
[0323] • Immune Cell Profiling: Flow cytometry or single-cell RNA sequencing can be used to identify and characterize immune cells within the organoid, providing insights into immune responses or inflammation.
[0324] • Extracellular Matrix (ECM) Composition: Analysis of ECM components, such as collagen or laminin, within the organoid, which play a critical role in tissue architecture and cell signaling.
[0325] 9. Synaptic Connectivity and Plasticity:
[0326] • Synaptic Protein Staining: Immunostaining for synaptic markers such as synapsin, PSD-95, or synaptophysin to assess the development and density of synaptic connections.
[0327] • Synaptic Plasticity Assays: Functional assays to measure synaptic plasticity, such as long-term potentiation (LTP) or long-term depression (LTD), which are indicative of learning and memory processes in the organoid.
[0328] 10. Vascularization and Blood-Brain Barrier (BBB) Modeling:
[0329] • Vascular Network Formation: Imaging and analysis of blood vessel formation within the organoid, often through staining for endothelial markers like CD31 or VE-cadherin.
[0330] • Blood-Brain Barrier Integrity: Assessment of BBB properties by evaluating tight junction protein expression (e.g., claudin-5, occludin) and permeability assays to determine the functional integrity of the barrier.
[0331] Organoids
[0332] To study the invasion and potentially model metastatic behavior in different microenvironments, several types of organoids could be utilized:
[0333] 1. Brain Organoids (Cerebral Organoids):
[0334] • Relevance: Cerebral organoids are the most directly relevant model for studying glioblastoma invasion within the brain. These organoids can be designed to mimic the architecture and cellular composition of the human brain, including the formation of neural layers and the development of a blood-brain barrier.
[0335] • Application: By using cerebral organoids, researchers can observe how glioblastoma cells invade different regions of the brain, interact with neural cells, and disrupt normal brain function. The inclusion of neural layer patterns and blood-brain barrier models provides insights into how glioblastoma cells penetrate and spread within the CNS. 2. Choroid Plexus Organoids:
[0336] • Relevance: The choroid plexus is a structure within the brain that produces cerebrospinal fluid (CSF) and forms part of the blood-CSF barrier.
[0337] • Application: Choroid plexus organoids could be used to study how glioblastoma cells interact with and potentially breach the blood-CSF barrier, providing insights into how cancer cells might migrate within the CSF and contribute to disease spread within the CNS.
[0338] 3. Blood-Brain Barrier (BBB) Organoids:
[0339] • Relevance: BBB organoids specifically model the selective permeability of the blood-brain barrier, which is a significant factor in glioblastoma's ability to invade brain tissue and in drug delivery challenges.
[0340] • Application: These organoids allow researchers to study how glioblastoma cells penetrate the BBB and spread within the brain. They also provide a platform to test how different treatments can cross the BBB to reach glioblastoma cells.
[0341] 4. Tumor-Stroma Organoids:
[0342] • Relevance: Tumor-stroma organoids include both cancer cells and the surrounding stroma, which consists of extracellular matrix components, immune cells, and fibroblasts.
[0343] • Application: These organoids can be used to study the interactions between glioblastoma cells and the tumor microenvironment, particularly how the stroma influences tumor invasion and resistance to therapy. This model can provide insights into the metastatic potential of glioblastoma-like behavior in other environments, such as through manipulation of the organoid's composition.
[0344] 5. Vascularized Organoids:
[0345] • Relevance: Vascularized organoids contain functional blood vessels, which are crucial for studying how glioblastoma cells interact with and invade vasculature, mimicking the process of angiogenesis and vessel co-option seen in tumors.
[0346] • Application: These organoids allow for the observation of how glioblastoma cells might invade and utilize blood vessels within the brain, contributing to their invasive spread and the formation of new, abnormal vasculature within the tumor.
[0347] 6. Meningeal Organoids: • Relevance: The meninges are the protective layers surrounding the brain and spinal cord.
[0348] Glioblastoma cells often invade these layers, contributing to their spread within the CNS.
[0349] • Application: Meningeal organoids can be used to study how glioblastoma cells invade these protective layers and migrate along the meningeal surfaces, which is particularly relevant in understanding the local spread of the disease within the CNS.
[0350] 7. Spinal Cord Organoids: Relevance: Although glioblastoma primarily affects the brain, the spinal cord can sometimes be involved, particularly in diffuse midline gliomas.
[0351] • Application: Spinal cord organoids can be used to study how glioblastoma cells might invade or spread to spinal cord tissue, providing insights into potential pathways of disease progression in the CNS.
[0352] 8. Co-Culture Organoids (Glioblastoma with Other Brain Structures):
[0353] • Relevance: Co-culture organoids involve the combination of glioblastoma cells with other brain cell types or structures, such as neurons, astrocytes, or endothelial cells.
[0354] • Application: These models can be tailored to study specific interactions between glioblastoma cells and other components of the brain, such as how cancer cells influence neural signaling, disrupt astrocytic function, or manipulate the vasculature to facilitate invasion.
[0355] Also, the present invention could be used for studying metastasis of glioblastoma cancer cells to other organs. Thus, in addition to brain-related organoids, using organoids derived from other nonbrain organs could provide valuable insights into glioblastoma's potential behavior in different tissue environments, especially when considering the theoretical possibility of metastasis or to study the mechanisms of local invasion, cell migration, or treatment responses. Here are some examples of non-brain organoids that may be relevant:
[0356] 1. Liver Organoids:
[0357] • Relevance: The liver is a common site for metastasis in many cancers. While glioblastoma typically does not metastasize outside the CNS, liver organoids can be used to study how glioblastoma cells might behave in a different microenvironment.
[0358] • Application: Liver organoids allow researchers to observe how glioblastoma cells might interact with hepatocytes and the liver's unique extracellular matrix. This can provide insights into the adaptability of glioblastoma cells and their responses to treatments in non-brain environments.
[0359] 2. Lung Organoids:
[0360] • Relevance: The lung is another common site for metastasis in many cancers. Studying glioblastoma cells in lung organoids could provide insights into how these cells would interact with the lung's microenvironment.
[0361] • Application: Lung organoids can be used to model the behavior of glioblastoma cells in the lung tissue, including their ability to survive, proliferate, and interact with lung-specific cells. This could also help in understanding the potential for distant metastasis, even if theoretical, and how therapies might differ in this context.
[0362] 3. Kidney Organoids:
[0363] • Relevance: Kidney organoids represent a complex tissue environment that could be used to study the adaptability and invasiveness of glioblastoma cells.
[0364] • Application: Researchers can use kidney organoids to investigate how glioblastoma cells might invade and proliferate within renal tissue. The unique vascular and extracellular matrix composition of the kidney can offer insights into the mechanisms by which glioblastoma cells might exploit non-CNS tissues.
[0365] 4. Intestinal Organoids:
[0366] • Relevance: The intestine is a highly vascularized organ with a distinct immune environment. Intestinal organoids could model how glioblastoma cells might interact with a different immune and vascular milieu.
[0367] • Application: Intestinal organoids allow the study of glioblastoma cell behavior in a gut-like environment, particularly how these cells might interact with immune cells and the gut epithelium. This model could help in understanding the potential immune evasion strategies of glioblastoma cells outside the brain.
[0368] 5. Pancreatic Organoids:
[0369] • Relevance: The pancreas is another organ with a dense stromal environment. Pancreatic organoids can be used to explore how glioblastoma cells might adapt to or invade such a microenvironment.
[0370] • Application: Pancreatic organoids provide a platform to study the interaction between glioblastoma cells and pancreatic cells, including how the cancer cells might respond to the pancreatic microenvironment's unique signalling and metabolic conditions.
[0371] 6. Bone Marrow Organoids:
[0372] • Relevance: The bone marrow is a primary site for hematopoiesis and contains a rich immune cell population. Bone marrow organoids can model interactions between glioblastoma cells and immune cells.
[0373] • Application: These organoids can be used to study how glioblastoma cells might interact with bone marrow-derived cells, such as immune cells or hematopoietic stem cells. This can provide insights into immune evasion mechanisms or how glioblastoma cells might influence the bone marrow environment.
[0374] 7. Skin Organoids:
[0375] • Relevance: Skin organoids are useful for studying interactions between glioblastoma cells and epithelial cells in a highly stratified tissue environment.
[0376] • Application: Skin organoids can be used to explore how glioblastoma cells might behave in a barrier tissue, including their ability to invade, induce angiogenesis, or alter the local microenvironment.
[0377] 8. Cardiac Organoids:
[0378] • Relevance: Cardiac organoids model heart tissue, which has a highly specialized cellular environment. Studying glioblastoma cells in this context could reveal how they interact with different cell types such as cardiomyocytes. • Application: Cardiac organoids allow for the exploration of how glioblastoma cells might impact or adapt to heart tissue, particularly in terms of how these cells might respond to the mechanical and electrical signalling unique to the heart.
[0379] 9. Adipose Tissue Organoids:
[0380] • Relevance: Adipose tissue organoids represent fat tissue, which is involved in energy storage and has a distinct metabolic profile. Studying glioblastoma in this context could provide insights into metabolic interactions.
[0381] • Application: Adipose tissue organoids can be used to study the interactions between glioblastoma cells and fat cells, particularly how these cancer cells might exploit or alter the lipid- rich environment for their growth.
[0382] 10. Bladder Organoids:
[0383] • Relevance: The bladder, with its unique urothelial lining and exposure to different chemical environments, could provide insights into how glioblastoma cells adapt to epithelial tissues.
[0384] • Application: Bladder organoids can model how glioblastoma cells interact with urothelial cells and the bladder microenvironment, providing insights into potential pathways of invasion in epithelial tissues.
[0385] System
[0386] In a second aspect, the invention refers to a system for performing the method of the invention.
[0387] Explanation of system related features: • Processor and Modules: The processor is central to executing the various steps of the method, connected to modules responsible for cultivation, treatment, and control of the organoids, as well as data collection and analysis.
[0388] • Data Acquisition System: This system ensures the collection of comprehensive data from the organoids, which is critical for creating the digital twin and making accurate therapy predictions.
[0389] • Data Integration and Analysis Module: This component is key to integrating the data and generating a digital twin of the organoid model that reflects the patient's cancer progression.
[0390] • Prediction Engine: This software component is responsible for predicting therapy outcomes and making real-time adjustments to treatment plans based on patient data.
[0391] • Storage Unit: Secure storage is essential for handling sensitive patient and research data, ensuring that it is kept safe and accessible for analysis.
[0392] • User Interface: The Ul provides an interactive platform for researchers and clinicians to engage with the system, input data, and make informed decisions based on the system's outputs.
[0393] • Communication Interface: This interface allows the system to contribute data to a global atlas, enhancing the broader understanding of glioblastoma and refining the system's predictions.
[0394] While all the components listed contribute to the system's functionality, their necessity may vary depending on the specific application or use case: Components (a) processor, (b) a data acquisition system, (c) a data integration and analysis module, and (d) a prediction engine are fundamental for performing the core functions of the method. These include processing patient-derived cancer cells, acquiring comprehensive data, integrating the data into a digital twin, and using predictive models to refine therapy.
[0395] The following components are potentially optional:
[0396] (e) Storage Unit: While secure data storage is crucial in most clinical and research settings, its implementation might be optional if the system is integrated with external cloud-based storage solutions managed by the user or institution.
[0397] (f) User Interface: The user interface could be considered optional in certain automated systems designed purely for data integration and prediction without direct clinician interaction. However, it is critical for applications where human oversight or intervention is needed.
[0398] (g) Communication Interface: Integration with a global atlas is invaluable for research and broader predictive modeling, but in standalone systems focused solely on individual patient care, this component might be optional.
[0399] Applications Context:
[0400] The inclusion or exclusion of components like (e), (f), and (g) may depend on the intended application (e.g., a personalized therapy platform vs. a global research database).
[0401] Other products related to the method and system of the present disclosure
[0402] In other aspects, the invention relates to products, services and applications wherein the method and / or the system of the present disclosure is / are used.
[0403] The primary product (a platform for creating and using digital twins of patient-derived organoids) could be applied in several products and services, ranging from clinical decision support tools to drug development platforms and personalized cancer treatment services. The technology's ability to reduce costs, improve treatment outcomes, and provide actionable insights at both the individual and population levels make it a compelling opportunity in the rapidly growing field of precision oncology.
[0404] Some key potentials and products:
[0405] 1. Personalized Cancer Therapy Platform
[0406] • Product: A comprehensive software and hardware platform that generates digital twins of patient-derived organoids to predict the most effective cancer therapies. This platform could be sold to hospitals, cancer research centers, and pharmaceutical companies.
[0407] • Commercial Potential: High. With the increasing emphasis on personalized medicine, such a platform could become a cornerstone in the treatment planning process, allowing oncologists to tailor therapies to individual patients' tumor profiles, improving outcomes and reducing the trial- and-error approach often seen in cancer treatment.
[0408] 2. Drug Development and Testing Platform
[0409] Product: A service or platform that pharmaceutical companies can use to test the efficacy of new drugs on digital copies of cancer models derived from organoids. This would allow drug developers to predict how different patient subpopulations might respond to new therapies before clinical trials.
[0410] • Commercial Potential: High. This product would shorten the drug development timeline, reduce costs, and increase the likelihood of successful clinical trials by enabling more precise targeting of therapies.
[0411] 3. Integrated Patient Care Systems
[0412] • Product: An integrated system that includes patient monitoring, data collection, and real-time therapy adjustment based on ongoing data from both the patient and their organoid model. This system could be licensed to hospitals and clinics. • Commercial Potential: Moderate to High. As healthcare moves towards more data-driven and adaptive approaches, this system could become integral to managing complex cancer cases, ensuring that treatment regimens are continually optimized for the best outcomes.
[0413] 4. Glioblastoma-Specific Clinical Decision Support Tool
[0414] • Product: A specialized software tool designed for glioblastoma treatment, leveraging the global glioblastoma cell atlas to help clinicians make informed decisions based on the latest research and patient-specific data.
[0415] • Commercial Potential: High within the niche market of neuro-oncology. Given the complexity and poor prognosis of glioblastoma, a tool that offers precise, data-driven guidance could be highly valuable to specialists in this field. 5. Global Cancer Research Database Subscription
[0416] • Product: A subscription-based service that provides access to a continually updated global cancer cell atlas, including patient-derived organoid data. Researchers, clinicians, and pharmaceutical companies could use this database to identify new treatment targets and understand cancer biology at a deeper level.
[0417] • Commercial Potential: Moderate to High. While niche, the continuous influx of new data and insights would be invaluable to both academic and industry researchers, providing a steady revenue stream through subscriptions.
[0418] 6. Predictive Analytics for Clinical Trials
[0419] • Product: A predictive analytics tool that uses organoid-derived digital twins to identify ideal patient cohorts for clinical trials, predict outcomes, and optimize trial design.
[0420] • Commercial Potential: High. This tool would significantly enhance the efficiency and success rates of clinical trials, making it highly attractive to pharmaceutical companies looking to de-risk their investment in new therapies.
[0421] 7. Personalized Oncology Services
[0422] • Product: A personalized oncology service where patients can have organoids created from their tumor samples, and the most effective treatment regimens are identified using the digital twin technology. • Commercial Potential: Moderate to High. This could be offered as a premium service through specialized cancer treatment centers or as a partnership with insurance providers, offering a new level of personalized care.
[0423] The invention will now be further described by way of the following non-limiting examples.
[0424] EXAMPLES
[0425] Results
[0426] To build upon the foundational insights of this disclosure, the inventors applied consensus Nonnegative Matrix Factorization (cNMF) to a large GB cohort, revealing additional metaprograms that illuminate GB's transcriptional diversity. By analyzing scRNA-seq data from 140 GB patients from the merged and harmonized GBMap resource and applying cNMF to capture a comprehensive range of GB-specific metaprograms, the inventors ultimately identify 30 unique transcriptional states. These include not only well-known lineage-like states but also novel signatures, such as those associated with hypoxia, neuroepithelial traits, and immune modulation. This expanded perspective reveals the intricate interplay of transcriptional states within GB, highlighting both commonalities with normal brain development and GB-specific adaptations.
[0427] To better understand how these states manifest in a controlled environment, the inventors developed GB organoid models by implanting GB cell lines into cerebral organoids, creating a brain-like tumorhost system. The inventors hypothesized that cancer cells would exhibit distinct behaviors in this 3D environment compared to traditional 2D cultures, offering insights into how organoids might serve as robust models for studying GB plasticity. Using organoid scRNA-seq data across time points and conditions, we assessed gene expression dynamics, stability, and condition-specific behaviors to evaluate how well these models recapitulate patient-derived states.
[0428] Through a comparative analysis of gene expression networks, we aimed to validate our organoid models as representative systems for GB biology. By tracking metaprogram (MP) correlations over time, the inventors captured temporal changes in gene expression, revealing the progression of transcriptional states across various conditions. Additionally, we examined condition-specific behaviors across different experimental conditions, identifying how environmental factors shape cancer cell states. The inventor's analysis also identified stable and variable MP correlations over time, which likely reflect core regulatory processes and adaptive responses, respectively. Finally, the inventors related these findings to tumor structure through 3D imaging, leveraging tissue clearing techniques to map the spatial organization of MPs within organoids.
[0429] This study provides a refined perspective on GB transcriptional states by combining a large-scale patientderived dataset with a 3D organoid model system. Our work emphasizes the need for comprehensive, multi-scale approaches in studying GB, capturing both the consistency of key processes and the adaptability of tumor cells to varying microenvironmental conditions across time.
[0430] Example 1 - Identification of patient-derived GB metaprograms
[0431] To better structure, and to facilitate interpretation of transcriptional ITH, the inventors decided to conduct a metaprogram analysis using a large collection of GB scRNAseq data. A metaprogram refers to a higher-level collection of gene expression patterns that capture more complex, often broader biological states or processes. Metaprograms represent overarching themes or states of cellular activity that can encompass multiple pathways and broader regulatory networks. Each metaprogram represents a group of genes that are co-expressed and potentially co-regulated, reflecting specific biological processes or cellular states. The identification of metaprograms has been used effectively to characterize transcriptomic states in tumors before, either in the case of Neftel et al 2019 which identified 6 metaprograms in GB, or Gavish et al 2023 (Hallmarks of transcriptional intratumour heterogeneity across a thousand tumours. Nature. 2023 Jun;618(7965):598-606. doi: 10.1038 / s41586-023-06130-4. Epub 2023 May 31. PMID: 37258682.) which identified hallmark metaprograms present across many types of cancer. To reassess the range of states present in GB, we selected 140 patients accessed via the GBMap (Harmonized single-cell landscape, intercellular crosstalk and tumor architecture of glioblastoma Cristian Ruiz-Moreno, Sergio Marco Salas, Erik Samuelsson, Sebastian Brandner, Mariette E.G. Kranendonk, Mats Nilsson, Hendrik G. Stunnenberg bioRxiv 2022.08.27.505439; doi: https: / / doi.org / 10.1101 / 2022.08.27.505439), a merged and harmonized dataset of publically available GB scRNAseq data. The donors included in the metaprogram assessment were chosen from the extended GBMap based on two criterias; that they had more than 10% malignant cells, and a sample size of less than 10,000 cells.
[0432] To identify the heterogeneous transcriptomic programs present across the samples, the inventors have applied an approach similar to the one pursued by Gavish et al 2023, but with the aim of characterizing GB tumor-type specific programs rather than pan-tumor programs. We performed consensus Non-negative Matrix factorization (cNMF) on the data of each individual patient to identify the programs of each sample. Since application of cNMF requires a "K" parameter that influences the results, the inventors run cNMF using different values (l<=5,6,7,8,9,10) and generate 45 programs for each tumor. Each cNMF program is summarized by the top 100 genes representing that program based on cNMF coefficients. Then, the inventors identified the most robust cNMF programs across the patient cohort as those that recur within the tumor (gene lists have at least 70% overlap), recur across tumor (have at least 20% overlap with any other cNMF program in other patients analyzed), and are non-redundant within the tumor (rank programs by similarity with programs from other tumors, then remove programs that have at least a 20% overlap with other programs within the same patient). From the robust cNMF programs, the inventors clustered them together based on their Jaccard similarity, and identified the most consistent genes present across the grouped cNMF programs.
[0433] The inventors identified 30 metaprograms from the GB patient cohort. To validate the clusters of cNMFs that define each metaprogram, the inventors assessed the Jaccard similarity index for all robust cNMFs, grouped by metaprogram (Fig. 2A). Overall, the heatmap reveals clear patterns of similarity within each cluster, with well-defined blocks along the diagonal indicating strong intracluster cohesion. The distinct separation between clusters highlights the differences in feature sets among them, providing a comprehensive view of the relationships and similarities within the dataset. Furthermore, to assess the approach for identifying metaprograms in GB, the inventors ran it on the adult GB samples from the original Neftel et al 2019 publication. The approach generated similar results to the original Neftel metaprograms (see Fig. 7).
[0434] To assess the expression of the metaprograms in GB patient-derived malignant cells, the inventors examined the relative expression of the genes in the extracted malignant cells from the core GBMap. The cumulative gene expression heatmap reveals distinct patterns of expression across the 30 metaprograms in malignant cells, providing insights into potential subpopulations and co-regulatory mechanisms. The relative expression scores provide a quantitative measure of gene activity. Hierarchical clustering of the cells based on their cumulative expression profiles revealed potential subpopulations within the malignant cells. These subpopulations are characterized by consistent expression patterns across multiple metaprograms, suggesting shared regulatory mechanisms or functional states (Fig. 2B). Certain metaprograms, such as MP 0, MP 1, MP 4, and MP 23, exhibit distinct clusters of high expression, suggesting their genes are highly active in specific subpopulations of cells. Metaprograms MP 2, MP 3, MP 7, MP 8, MP 12, MP 13, MP 17, MP 18, MP 20, MP 21, MP 24, MP 28, and MP 29 display varied expression levels, with notable high expression in certain cell clusters, indicating their involvement in specific biological processes or functional states. MP 5, MP 6, MP 9, MP 10, MP 11, MP 14, MP 15, MP 16, MP 19, MP 22, MP 25, MP 26, and MP 27 show relatively uniform low to moderate expression across cells, suggesting these metaprograms may play fundamental or supportive roles in cellular processes. Overall, the heatmap offers a comprehensive overview of the cumulative gene expression profiles of metaprograms in malignant cells, facilitating the identification of significant expression patterns and subpopulations for further biological exploration.
[0435] To assess the expression of the metaprograms within the malignant cells, the inventors then annotated each malignant cell based on the highest scoring metaprogram. The analysis of cell assignments to Metaprograms (MPs) revealed a highly skewed distribution. MP 21 and MP 2 were the most dominant, encompassing 184,896 (45.8%) and 129,061 (32.0%) of the cells, respectively (Fig. 1C). These two MPs alone account for nearly 78% of the total cell population, indicating significant roles in the dataset. Other MPs, such as MP 9 and MP 1, also contributed a notable number of cells, with 35,575 (8.8%) and 13,597 (3.4%) cells, respectively. However, the majority of the MPs represented relatively small proportions, with MPs like MP 20 (2.0%), MP 8 (1.9%), and MP 11 (1.2%) each accounting for less than 2% of the total cells. To provide a clearer view of the data, the inventors also presented the number of cells assigned to each MP in a table format (Fig. 1C). The table highlights the exact cell counts for each MP, emphasizing the disparities in cell distribution. For instance, MPs such as MP 6, MP 7, and MP 29 contribute marginally, with counts of 2,842 (0.7%), 2,249 (0.6%), and 2,164 (0.5%), respectively. The least represented MPs, such as MP 18 with only 20 cells (0.0%) and MP 28 with 172 cells (0.0%), underscore the presence of rare cell types or potential outliers within the dataset. This skewed distribution underscores the heterogeneity within the malignant cell population and suggests that MP 21 and MP 2 might define crucial biological roles.
[0436] Due to the plastic nature of malignant cells in GB, the inventors know that they have the ability to change state. For each cell assigned to an MP, the inventors determined the most likely second- highest scoring MP to assess hybrid states indicative of plastic transitions. The inventors excluded specific MPs ('MP O’, ’MP 4', 'MP 5’, ’MP 14', 'MP 16’, ’MP 19’, 'MP 22', 'MP 23’, and 'MP 27’) from the analysis because they represented cycling (MP 0 and MP 23), noise (MP 4), or were more indicative of specific functions rather than interconnected plastic states in malignant cells. These MPs did not share significant overlaps with other MPs as seen in Fig. 2-3. The heatmap (Fig. ID) reveals significant insights into the relationships between different marker profiles (MPs), excluding the specified MPs due to their roles in cycling, noise, or specific functions. Notably, MP 2 and MP 21 frequently appear as the second-highest scores across various primary MP assignments. For instance, MP 2 is prominently the second-highest MP for primary MPs 1, 8, 20, 21, 25, 26, and 29, indicating a widespread relevance or hybrid tendency with these profiles. Similarly, MP 21 is a recurrent second- highest MP for primary MPs 1, 7, 9, 10, 11, 12, and 21. This suggests a strong secondary association or hybrid characteristic involving MP 21. Additionally, some primary MPs, like MP 6 and MP 17, show a high concentration of second-highest scores in specific MPs (MP 9 and MP 21, respectively), highlighting potential strong hybrid or secondary marker relationships. These findings underline the hybrid nature of certain cells and suggest specific MP overlaps that might be crucial for understanding cell differentiation and identity within this dataset.
[0437] To identify metaprograms with similar correlation patterns, the inventors performed hierarchical clustering on a correlation matrix of the metaprogram scores in the malignant cells of the Extended GBMap (Fig. IB). This clustering method helped to identify groups of metaprograms with similar correlation patterns, facilitating the identification of functional modules and potential regulatory relationships.
[0438] The printed correlation matrix provides detailed numerical values of the correlation coefficients between each pair of metaprograms (MPs), revealing several key observations. Clusters of correlations can be identified within the matrix. Group 1, comprising MPs like MP_1, MP_2, MP_20, and MP_27, shows strong positive correlations among its members, suggesting that these MPs might be part of a common functional module. Group 2, including MPs like MP_8, MP_25, MP_28, and MP_29, also exhibits high positive correlations, indicating that these MPs share related functions or are co-regulated. On the other hand, MP_15 demonstrates strong negative correlations with many MPs in Groups 1 and 2, implying a potential antagonistic relationship. This could indicate that MP_15 might have an opposite functional role, such as being involved in differentiation or inhibitory pathways. This analysis highlights common association of states across many patients, representing the core consistencies of the disease.
[0439] Characterization of Patient-Derived GB Metaprograms
[0440] MPs were annotated based on gene set enrichment analysis (GSEA), previously identified GB metaprograms from Neftel et al 2019, and hallmark metaprograms from the Gavish et al. 2023 study on biological Hallmarks across many varieties of cancer (Fig 2, Supplementary Note 1). We did identify MPs that, based on their similarity to the original Neftel and Gavish publications, most resembled the original GB MPs, namely MES-like (MP 11), AC-like (MP 1), NPC-like (MP 29) and OPC- like (MP 3), along with Cycling (MP 0 and MP23). A further 5 were similar to those from the cancer hallmarks paper. 19 were new (Fig. 2). They were further grouped into families based on their related functions and overlap with each other (Fig. 1C). The 30 MPs are described briefly below and in detail in Supplementary Note 1.
[0441] Several MPs were identified with roles in general processes such as cell cycling, stress responses, ciliation, damage response and senescence. Interestingly, we identified a MP that consisted of genes associated with post-transcriptional modifications (MP 14), indicating a potential signature of transcriptomic regulation in GB. Besides the general MPs, the main families included MPs with oligodendroglial, neuronal, radial glia, astrocytic, mesenchymal, stem and neuroepithelial traits. The oligodendroglial lineage MPs included MP 3 (OPC-like 1), MP 6 (OC-like), and MP 17 (OPC-like 2). These metaprograms reflect characteristics associated with OPC differentiation, process that has been shown to occur within GB as tumor cells interact with white matter as they migrate to demyelinated regions. Neuronal-associated MPs comprised MP 12 (IN-like), MP 15 (N-like), and MP 29 (NPC-like). These MPs demonstrate features related to neuronal differentiation, highlighting the adaptability of GB cells in adopting neuronal-like functions that enable synaptic network integration45 and regulation via calcium channel oscillations4 Astrocytic MPs were represented by MP 1 (AC-like 1), MP 7 (AC-like 2), and MP 18 (AC-like 3) and exhibited markers and pathways aligned with astrocytic differentiation. MP 11 (MES-like) and MP 26 (MES-like Inflammation), showed features characteristic of mesenchymal cells. The radial glial-like MPS included MPs 9 (RG-like 1) and 21 (RG- like 2), and likely represent intermediate states on the path of neuroglial-like development of GB cells as they differentiate towards other lineages. MP 8 (NSC-like) reflected stem cell-like characteristics, crucial for maintaining the tumor's ability to self-renew and adapt under varying conditions.
[0442] Most intriguing are the set MPs that resemble adult olfactory and developmental euroepithelium, which have not, to knowledge, been previously described. MP 2 (NE-like Hypoxia 1), MP 10 (NE-like Hypoxia 2), and MP 24 (NE-like Hypoxia 4 / Flagella) were associated with neuroepithelial traits and responses to hypoxia. Hypoxia-related MPs play a significant role in tumor survival and adaptation under low-oxygen conditions, driving transcriptomic changes in GB. A second group, including MP 28 (NE-like Inflammation 1), MP 13 (NE-like Inflammation 2), MP 25 (NE-like Inflammation 3), displayed traits characteristic of neuroepithelial cells responding to inflammation and immune regulation, characteristics of GB progression. These MPs suggest a dual role, where the neuroepithelial properties can exist in hypoxic and inflammatory conditions, and the inflammation-linked traits could modulate immune activity and potentially aid in immune evasion.
[0443] Example 2 - Results: Comparison of Malignant Metaprograms to Normal References
[0444] GB shares features of normal neurodevelopmental hierarchy, and further recapitulates mature brain features such as synapse formation and integration into neuronal circuitry. Furthermore, GB shares hallmarks seen in a variety of other types of cancers, such as hypoxia and vascularization.
[0445] Characterization & Comparison of Malignant Metaprograms to Normal References
[0446] GB exhibits features aligning with the neurodevelopmental hierarchy observed in normal brain development. Additionally, it recapitulates mature brain characteristics, such as synaptic formation and the capacity to integrate into neural circuits. GB also shares hallmark traits common across diverse cancers, including mechanisms related to hypoxia and vascularization.
[0447] Building on these insights, the inventors utilized gene set enrichment analysis (GSEA) alongside previously characterized metaprograms from key studies (Neftel et al., 2019; Gavish et al., 2023) to annotate the identified malignant metaprograms (MPs). This approach highlighted similarities between our MPs and well-defined GB-associated states, such as mesenchymal-like (MES-like, MP 11), astrocytic- 1 ike (AC-like, MP 1), neuronal precursor-like (NPC-like, MP 29), and oligodendrocyte precursor-like (OPC-like, MP 3) programs, as well as cycling MPs (MP 0, MP 23). Five MPs showed parallels with those in the cancer hallmark framework described by Gavish et aL, while 19 MPs were newly identified in this study. These 30 MPs were categorized into functional families based on their overlapping characteristics and related summary of these MPs is provided below, with comprehensive descriptions.
[0448] Several MPs were linked to fundamental cellular functions, including cell division, stress response, ciliation, DNA damage repair, and senescence. Notably, MP 14, enriched in genes linked to post- transcriptional modifications, suggests a signature related to transcriptomic regulation in GB. Beyond these general functions, other MPs were categorized into lineage-specific families, reflecting diverse differentiation potentials.
[0449] • Oligodendroglial Lineage MPs: MP 3 (OPC-like 1), MP 6 (OC-like), and MP 17 (OPC-like 2) highlighted features tied to oligodendrocyte precursor differentiation, an essential aspect of GB cells' interaction with white matter and demyelinated regions during migration.
[0450] • Neuronal-associated MPs: MP 12 (IN-like), MP 15 (N-like), and MP 29 (NPC-like) showcased characteristics of neuronal differentiation, suggesting GB cells can adopt neuronal-like behaviors, aiding their integration into neural networks and enabling calcium-channel- mediated signaling.
[0451] • Astrocytic MPs: MP 1 (AC-like 1), MP 7 (AC-like 2), and MP 18 (AC-like 3) reflected astrocytic traits and differentiation pathways.
[0452] • Mesenchymal MPs: MP 11 (MES-like) and MP 26 (MES-like Inflammation) exhibited mesenchymal properties often associated with GB's invasive and adaptive capacities.
[0453] • Radial Glial-like MPs: MPs 9 (RG-like 1) and 21 (RG-like 2) represented intermediate states in the neuroglial differentiation trajectory of GB cells.
[0454] • Stem-like MPs: MP 8 (NSC-like) displayed stem cell-associated traits, essential for the tumor's plasticity and self-renewal.
[0455] • Neuroepithelial-like MPs: Several MPs resembled adult olfactory and developmental neuroepithelial traits. These include MPs associated with hypoxia-driven adaptations. (MP 2, MP 10, MP 24) and those linked to inflammation and immune modulation (MP 28, MP 13, MP 25). These findings underscore GB's ability to exploit neuroepithelial properties for survival and progression, with potential implications for hypoxia-driven and inflammation- mediated mechanisms in immune evasion and tumor adaptability.
[0456] To investigate the range of shared features overlaps between our derived GB MPs, the inventors conducted a comprehensive analysis using gene expression data from various studies. Specifically, the inventors examined the overlap among MPs 0-29 derived from malignant cells in the Extended GBMap, MPs from the Velmeshev et al. 2023 (Single-cell analysis of prenatal and postnatal human cortical development. Science. 2023 Oct 13;382(6667): eadf0834. doi: 10.1126 / science.adf0834. Epub 2023 Oct 13. PMID: 37824647; PMCID: PMC11005279.) Cortical brain study, MPs from the Braun et al. 2023 (Comprehensive cell atlas of the first-trimester developing human brain. Science. 2023 Oct 13;382(6667):eadfl226. doi: 10.1126 / science.adfl226. Epub 2023 Oct 13. PMID: 37824650.) study of Fetal brain, and MPs from the Gavish et al. 2023 study on biological Hallmarks across many varieties of cancer. The key differences between the datasets were that the fetal brain represented cells extracted from the XXX regions in the first trimester of development, while the cortical represented XXX regions from month X to Month X post-birth, representing early and later stages of brain development.
[0457] The heatmap (Fig. 2) of the Jaccard similarity indices between gene sets of various metaprograms and clusters from the Cortical2023, Fetal, and Hallmark datasets reveals shared features between brain developmental programs, general cancer programs and our GB metaprograms. Several key observations can be made from the heatmap. See Fig. 8 (Example 20) for a full list of overlapping genes with the GB Metaprograms. MPs with the highest overlap included MPO and MP23 (both cycling) to Cortical_MP7 and F4, MP8 (NSC-like) to F21, MP16 (Cilia) to F7, MP19 (DNA Damage Response) to F8, MP21 (RG-like 2) to Cortical_MP2, and MP29 (NPC-like) to Fl. MPs with moderate overlap included MP1 (AC-like) to F21, MP3 (OPC-like 1) to F10, MP6 (OC-like) to Cortical_MP6 and F10, MP7 (AC-like 2) to Cortical_MP2, MP9 (RG-like 1) to F10, MP12 (IN-like) to Fl, and MP15 (N-like) to Cortical_MPl. And finally, those that were unique to GB included MP2 (NE-like Hypoxia 1), MP5 (Immune / lnflation), MP10 (NE-like Hypoxia 2), MP11 (MES-like), MP13 (NE-like Inflammation 2), MP14 (Post-transcriptional regulation), MP17 (OPC-like 2), MP18 (AC-like 3), MP20 (NE-like Hypoxia 3), MP22 (Stress), MP24 (NE-like Hypoxia 4), MP25 (NE-like Inflammation 3), MP26 (MES-like Inflammation), MP27 (NE-like Adhesion / Senescence), MP28 (NE-like Inflammation 1).
[0458] In summary, certain MPs show higher similarity indices with specific clusters, indicating potential shared functional roles, particularly in processes like cell cycle regulation, damage response, and normal development. Conversely, many MPs exhibit minimal overlap, indicating distinct gene sets with functionalities unique to GB. Most notable are the fact that the NE-like and MES-like MPs saw little overlap with normal brain development, and that the MPs that were similar to normal brain development varied between early (fetal) and late (cortical) stage development. This indicates that while GB has similarities to a wide temporal range of developmental states, there are those that do not.
[0459] Minimal Overlap MPs: Several metaprograms, including MPO, MP4, MP5, MP8, MP9, MP11, MP12, MP15, MP17, MP18, MP20, MP22, MP24, and MP26, exhibit minimal overlap across all clusters. This suggests that these MPs contain unique gene sets with specific functional attributes not widely shared with other metaprograms. They likely represent distinct biological processes or conditions in GB.
[0460] Low Similarity with Slight Overlap: MP2, MP3, MP6, MP10, MP16, MP21, MP25, and MP28 show low similarity indices across most clusters but have slight overlaps with specific clusters. For instance, MP2 has a noticeable overlap with Cortical_MP17 and Hallmark Cell Cycle - G2 / M, indicating the shared feature of cell cycle processes. MP3 shows a slightly higher overlap with Cortical_MP7 and Hallmark Hypoxia, indicating potential roles in stress or hypoxia responses. MP6 exhibits slight overlap with Hallmark EMT-IV, suggesting involvement in epithelial-mesenchymal transition processes, while MP16 overlaps somewhat with Fetal F8 and Hallmark Hypoxia, suggesting involvement in stress responses.
[0461] Moderate Overlap with Specific Clusters: MP1, MP13, MP14, MP19, MP23, and MP29 display low to moderate overlap with several clusters. MP1 shows overlap with Fetal clusters F4 and F21, suggesting a partial functional similarity with these fetal-specific gene sets. MP13 and MP19 exhibit higher similarity with Cortical_MP10 and Hallmark Cell Cycle - G2 / M, indicating potential involvement in cell cycle regulation. MP14 and MP23 show higher similarity with Cortical_MP7, suggesting significant overlap in gene sets involved in cortical development or neurogenesis. MP29 displays higher similarity with Hallmark EMT-III, indicating significant overlap in genes involved in epithelial- mesenchymal transition processes.
[0462] High Similarity with Specific Clusters: MP7, MP19, MP23, and MP27 exhibit higher similarity with specific clusters, indicating significant overlap in their gene sets. MP7 and MP23 show notable overlap with Cortical_MP17 and Cortical_MP7, respectively, suggesting significant involvement in cortical development or function. MP19 and MP27 exhibit higher similarity with Hallmark Cell Cycle - G2 / M, indicating significant overlap in cell cycle regulation genes. MP23 also overlaps significantly with Fetal F7.
[0463] In summary, the detailed analysis of the 30 metaprograms (MPs) highlights specific areas of significant overlap with the Hallmark, Fetal, and Cortical metaprograms. Certain MPs, such as MP13, MP14, MP19, MP23, MP27, and MP29, show higher similarity indices with specific clusters, suggesting potential shared functional roles, particularly in processes like cell cycle regulation, stress response, and development. Conversely, many MPs exhibit minimal overlap, indicating distinct gene sets with unique functionalities. This detailed assessment provides insights into the functional relationships and distinctiveness of the metaprograms in the context of the Cortical2023, Fetal, and Hallmark datasets.
[0464] Based on the overlapping genes shared with the metaprograms derived from other studies, the inventors have assessed the likely states that are specific to each GB MP.
[0465] Example 3: Results - Internal Comparison of GB Metaprograms
[0466] The analysis revealed significant overlaps (>5 genes) between various metaprograms, indicating shared gene expression patterns among different malignant and nonmalignant states. The investigation into the overlaps within the MPs derived from malignant cells in the Extended GBMap revealed several significant overlaps.
[0467] Example 4: Results: ITH and Evolution in GB - A Reassessment
[0468] To assess if MPs align with genetic subclones in malignant cells, the inventors used SCEVAN59 to identify distinct subclones within patient-derived malignant cells based on CNV profiles. In a cohort of 112 patients who passed quality control, we detected a total of 527 subclones, with 1 to 19 subclones per tumor. Each malignant cell was assigned to a metaprogram according to its highest MP score, revealing that subclones often contained cells spanning multiple MPs (Fig. 12). The median number of MPs per donor was 9, and no donors had cells represented in all MPs.
[0469] To explore whether MPs are linked to specific genetic subclones, the inventors analyzed residuals, which quantify differences between observed and expected distributions of MP-assigned cells within each subclone. The expected distribution assumes random assignment of MP-associated cells across subclones, based on each patient's overall MP proportions (Fig. 13). Positive residuals indicate overrepresentation, while negative residuals indicate under-representation of specific MPs within subclones. Residuals were calculated for each patient individually. Among the 112 patients, 73 had p- values below 0.05, rejecting the null hypothesis of random distribution, while 36 had p-values above, showing no significant deviation from random distribution. This suggests two groups: patients where MP and subclone associations exist, and others where MP assignment is random. Patients showing no MP-subclone association may represent more "plastic" tumors, potentially lacking niche restrictions or selective bottlenecks that enforce specific MPs. Interestingly, some patients exhibit residual patterns across specific MP pairs, often inversely correlated (e.g., high residuals in MP2 (NE-like Hypoxia 1) paired with low residuals in MP21 (RG-like-2)). Excluding MP14 (post-translational modifications), the highest scoring MPs also show these patterns: MP6 (OC-like) and MP9 (RG-like 1), as well as MP8 (NSC-like) and MP11 (Mesenchymal-like). Positive relationships include high MP2 (NE- like Hypoxia 1) with high MP1 (AC-like) and MP8 (NSC-like). These findings suggest that specific subclones tend to adopt distinct transcriptomic states, likely optimized for unique microenvironmental conditions, and may exert selective pressures influencing further tumor evolution. Conversely, subclones without strong MP associations appear more plastic, lacking restrictive niche pressures and thus retaining a broader adaptability within the tumor ecosystem.
[0470] Example 5: MP Clinical Associations
[0471] The inventors explored the relationship between GB MPs and survival outcomes by analyzing the bulk IDH-wildtype GB specimens from TCGA-GB, identified through reclassification in an earlier study. Bulk RNA profiles, which capture a mix of malignant and nonmalignant cell types and states in the TME, provide an approximate measure of the abundance of cellular states. Prior to generating survival curves for TCGA samples based on the cumulative high or low expression of metaprogram genes, the inventors conducted univariate Cox proportional hazard regression on each gene to determine its survival association, selecting the most significant genes (based on p-value) for inclusion in the metaprogram survival analysis. Each bulk sample was then scored for expression of the survival- associated genes within the 30 malignant GB MPs, and we assessed the correlation between expression levels and survival (Fig. 14 A-D). I ntriguingly, high expression of MP 2 (NE-like Hypoxia 1, p- value - 0.0049), MP 5 (Immune / lnflammation, p-value - 0.0018), MP 15 (N-like, p-value = 0.0034), and MP 29 (NPC-like, pvalue = 0.013) was associated with poorer survival.
[0472] These findings highlight the prognostic significance of specific GB metaprograms, underscoring how distinct cellular states within the tumor microenvironment relate to survival outcomes. The association of high expression in NE-like Hypoxia 1, Immune Inflammation, N-like, and NPC-like MPs with poorer survival suggests that hypoxic adaptation, immune signaling, synaptic-like activity, and neurodevelopmental features contribute to aggressive tumor behavior. These insights emphasize the potential value of targeting these pathways to improve therapeutic outcomes in GB.
[0473] Example 6: Comparison of Patient-Derived GB Metaprograms to Programs for Synapse Formation, Invasion, Development, and Wound Healing
[0474] In this analysis, the inventors compared patient-derived GB MPs to established transcriptional programs associated with synapse formation (22), invasion (23), development, and wound healing (24), as described in previous studies. By scoring each cell's expression of genes within these functional MPs, the inventors assessed whether MP-assigned cells in GB tumors exhibit overlap with these functional signatures, highlighting the versatility and complexity of their transcriptomic states.
[0475] The results revealed that MP-assigned cells in GB often align with multiple functional signatures (Fig. 15). MP 17, for instance, shows high invasion and developmental scores, indicating a strongly invasive and developmental gene profile. This pattern suggests an aggressive GB phenotype that leverages both invasion and growth mechanisms. MP 13, with slightly lower scores in these areas, may represent a similar, though less intense, transcriptional state. For wound healing, MP 17 and MP 28 score highly, implying engagement in repair mechanisms. MP 16 also shows notable wound-healing activity, possibly reflecting a reactive or regenerative profile alongside invasive features. In contrast, synapse-related scores are generally low, except for MP 6, which has a positive score, hinting at neuronal interactions— a trait occasionally co-opted by GB.
[0476] Differing invasion scores further distinguish MPs: MP 16 and MP 23 have high invasion scores, whereas MP 28 may focus more on wound healing than invasion. Collectively, these patterns suggest MP 17 as a highly aggressive, multi-functional state, while MP 13 shows similar traits at a reduced intensity. MP 28 and MP 16 are more aligned with wound healing, and MP 6 with potential neuronal mimicry. This analysis provides a foundation for further exploration of GB's varied transcriptional states and associated functions.
[0477] Example 7: Plasticity in GB
[0478] GB cells exhibit substantial plasticity, adapting their states as the tumor progresses. This plasticity allows variations in MP expression based on intrinsic cellular factors, local microenvironmental conditions, and tumor progression stages. While MPs define discrete transcriptomic states, GB operates on a continuum, with multiple MPs often expressed within single cells. To identify correlated states, the inventors performed hierarchical clustering on malignant cells based on cumulative MP expression, revealing distinct expression patterns (Fig. IB). These subpopulations, marked by consistent multi-MP expression, suggest shared regulatory mechanisms or functional states. MPs such as MP 0, 1, 2, 3, 5, 6, 12, 20, 21, 23, 28, and 29 showed distinct clusters of high expression, highlighting active states within certain cell subpopulations. In contrast, MPs like MP 7, 8, 9, 13, 17, 19, 25, and 26 displayed variable expression, suggesting specialized roles in specific processes or states. MPs 10, 11, 14, 15, 16, 18, 22, 24, and 27 exhibited more uniform low to moderate expression across cells, hinting at supportive or fundamental roles in cellular processes. These results enable us to identify subpopulations with distinct expression patterns, providing a basis for further biological investigation into GB's transcriptomic diversity.
[0479] In the context of GB plasticity, identifying correlations among MPs is essential to understanding interactions and transitions between transcriptomic states. By examining correlated MPs, we capture common associations across patients. To uncover metaprograms with similar patterns, we conducted hierarchical clustering on a correlation matrix of MP scores across malignant cells (Fig. IE), which highlights clusters of co-expressed MPs, revealing dynamic associations.
[0480] The inventors identified two main groups of highly correlated MPs, suggesting shared regulatory mechanisms. Group 1 includes MP 5 (Immune Inflammation), MP 15 (N-like), MP 16 (Cilia), MP 18 (AC-like 3), MP 14 (Post-transcriptional Regulation), MP 24 (NE-like Hypoxia 4), MP 12 (I N-like), and MP 19 (DNA-Damage Response). This group likely reflects a coordinated response to tumor- associated stressors like inflammation and hypoxia. Group 2 contains three sub-groups: Sub-group A (MP 11 and MP 26, MES-like 1 and 2) shows the strongest internal correlation; Sub-group B (MP 6, OC-like; MP 9, RG-like 1; MP 7, AC-like 2; and MP 21, RG-like 2) is linked to glial lineage traits; and Sub-group C (MP 3, OPC-like 1; MP 29, NPC-like; MP 10, NE-like Hypoxia 2; MP 22, Stress; MP 8, NSC- like; MP 27, Senescence; MP 25, NE-like Inflammation 3; MP 28, NE-like Inflammation 2; MP 2, NE-like Hypoxia 2; MP 20, NE-like Hypoxia 3; MP 1, AC-like 1; and MP 13, NE-like Inflammation 2) likely represents a spectrum of neural lineage, inflammatory, hypoxic, and stress-related states. These clusters suggest an adaptive continuum, where GB cells shift between regenerative, inflammatory, and stress-responsive states in response to environmental pressures.
[0481] Patient-specific analyses (Example 22) further highlight stable and variable relationships among GB MPs across individuals. Stable correlations, such as MP 1 (AC-like 1) with MP 7 (AC-like 2) and MP 3 (OPC-like 1) with MP 17 (OPC-like 2) and MP 6 (OC-like), are observed across most samples, indicating consistent astrocytic and oligodendroglial differentiation trajectories in GB tumors. Variable correlations, especially in MPs associated with hypoxia and inflammation (e.g., MP 2 and MP 10), suggest context-dependent shifts likely driven by unique tumor microenvironments. These samplespecific variations underscore the tumor's adaptability and the importance of individualized therapeutic approaches in GB.
[0482] In plastic cancers, a metaprogram correlation matrix reveals adaptability and complex state transitions. High correlations indicate co-activated programs supporting survival and adaptation, while negative correlations suggest mutually exclusive states. These patterns offer insight into cellular heterogeneity and plasticity, revealing potential therapeutic targets to disrupt adaptive processes and address treatment resistance. The correlations represent general associations across patients, but GB's plastic nature implies these relationships may shift with changing conditions or treatment stages. Repeated analyses over time could further clarify how MP interactions evolve throughout disease progression, enhancing understanding of GB's dynamic transcriptomic landscape.
[0483] Example 8: Characterizing Transcriptomic States in a Patient-Derived Organoid Model
[0484] To determine how accurately GB organoids model native GB, the inventors compared scRNA-seq data from organoid-derived and native GB cells (see Methods for details on organoid cultivation and sampling, Example 23 for Quality Control). To compare the transcriptomic states in our model to patient data, we performed the same MP analysis as on the patient-derived scRNAseq data to identify the organoid model-specific metaprograms (Fig. 4B) and compare them to both the newly identified GB metaprograms and the cancer hallmark metaprograms identified by Gavish et al 2023 (Fig. 4C). Like our results with patient samples (Fig. 1A), the Jaccard similarity index for all robust cNMFs in the organoid dataset revealed well-defined metaprograms (Fig. 4B). When assessing similarity of the organoid-derived metaprograms to the patient-derived and cancer hallmark ones, we observed that out of the 6 identified MPs in the organoid model, 4 are similar to patient-derived ones and 2 are unique to the model (Fig. 4C). Organoid_Clusterl is similar to MP 2 (NE-like Hypoxia 1, Organoid_Cluster2 is similar to the cycling associated metaprograms (MP 0 and 23 from our data, and Cell Cycle - G2 / M), Organoid_Cluster4 is similar to both our MP 2 (NE-like Hypoxia 1) and Gavish et al. 2023’s EMT-I, and Organoid_Cluster6 is similar to both our MP 19 (DNA Damage Response) and Gavish et al. 2023's Cell Cycle - Gl / S. Organoid_Cluster3 and Organoid_Cluster5 have low similarity with most GB patient metaprograms and hallmark gene sets. This variation in similarity emphasizes the heterogeneity within organoid clusters and their unique alignment relative to patient-derived metaprograms and established hallmarks. These clusters likely represent distinct biological states specific to the organoids, with minimal overlap with the GB or Cancer Hallmark profiles.
[0485] Example 9: Mapping Transcriptomic States in a Patient-Derived Organoid Models from another Study
[0486] In a recent study on GB organoids, scRNA-seq analysis was conducted to map transcriptomic states, with several differences from the inventor's approach. Specifically, they used a single time point (two weeks post-attachment), nine patient-derived GB cell lines rather than a single line, and older organoids (4-6 months old compared to our 1-month-old models). The 7,658 cancer cells extracted from their organoids for sequencing were categorized based on the older Neftel et al. 2019 metaprograms, with the conclusion that their organoid models effectively represented the transcriptomic states of GB. To validate this, the inventors downloaded their dataset and applied our new metaprogram assignments, assigning each cancer cell based on the highest-scoring metaprogram (Fig. 17).
[0487] Our analysis reveals key limitations in the ability of organoid models to fully recapitulate the complexity of GB. Out of the 30 metaprograms identified from our cohort of 140 GB patient biopsy scRNA-seq datasets, many were either underrepresented or absent in the organoid models. Notably, MP 21 (RG-like 2), a metaprogram that constitutes a substantial proportion of cells in patient-derived data, was largely absent in their organoid models. However, the inventors observed alignment in some areas: MP 2 (NE-like) and MP20 (NE-like 2) were dominant across organoid conditions, mirroring trends the inventors identified. In the organoid model, more than half of the patient-derived GB MPs have a negligible presence in the organoid models. Several MPs are either entirely absent (no cells assigned at all) or largely absent (defined as having fewer than 30 cells), highlighting limitations in capturing the full spectrum of GB transcriptomic diversity. The MPs completely missing from the dataset include MP 5, MP 6, MP 15, MP 16, MP 17, MP 18, MP 21, and MP 26. Additionally, several MPs show very low cell counts, indicating limited representation within the model. These include MP 25, MP 28, MP 24, MP 13, MP 7, MP 19, and MP 11. This analysis suggests that a substantial portion of the transcriptomic diversity observed in GB patient-derived samples is not fully captured in the organoid model. Overall, our findings underscore the utility of multiple cell lines for a more comprehensive representation of GB heterogeneity and highlight that current organoid models, while valuable, only partially capture the complete spectrum of cancer cell states inherent to GB.
[0488] Example 10: Comparison of GB Metaprogram Dynamics between Model System and GB Patient Cohort
[0489] Comparing organoid cells (from our single patient-derived line) (Fig. 5A) with patient-derived samples (140 patients) (Fig. IE) highlights key differences. Organoid cells show high, consistent correlations among metaprograms, indicating a more homogeneous expression pattern, while patient samples exhibit diverse, variable correlations due to genetic, environmental, and treatment differences. Certain metaprograms, like MP 1, MP 2, MP 6, MP 7, and MP 9, display stable correlations in organoids but more variable ones in patient data, emphasizing the controlled nature of organoids for fundamental cancer studies. In contrast, the patient data reveal both strong positive and negative correlations, capturing the complex heterogeneity crucial for personalized approaches in cancer treatment.
[0490] Example 11: Mapping Transcriptomic States in a Patient-Derived Organoid Model across Time and Condition
[0491] The inventors started by evaluating the prevalence of organoid model-specific and GB patient-derived metaprograms and assigned cells to metaprograms based on their highest scores. To start, Fig. 4A shows UMAP projections of malignant cells categorized by timepoint and condition. The first plot, organized by timepoints, shows a temporal progression, with early timepoints (1.0) clustering together and transitioning through intermediate stages (2.0 and 3.0) to a distinct cluster for the latest timepoint (4.0). This progression captures dynamic gene expression changes in the malignant cells over time and condition. In the second plot, cells are grouped by condition (314, 520, Cancer Spheroid, and Cancer Spheroid Matrigel Embedded), forming distinct clusters that reflect gene expression differences driven by each condition.
[0492] Analysis of GB patient-derived metaprogram scores over time across three conditions— 314, 520, and Cancer Spheroid— revealed distinct correlation patterns (Fig. 4F). MPs 1, 11, 13, 25, and 28 showed strong negative correlations with time across all conditions, especially in condition 314, where MPs like 1, 13 and 28 were most pronounced. Positive correlations for MPs 3 and 6 indicated an increase over time, with MP 3's strongest correlation in Cancer Spheroid versus 314 and 520. Conditionspecific trends emerged: Cancer Spheroid uniquely displayed positive correlations for MPs 12, 19, and 21, while MPs 7 and 8 had higher positive and negative correlations, respectively. Condition 314 consistently had the strongest negative correlations overall. MPs 2 and 17 showed weak or insignificant correlations across all conditions, suggesting stable scores over time. These findings underscore the nuanced metaprogram dynamics in different experimental settings.
[0493] Example 12: Assessment of GB Metaprogram Dynamics across Time and Conditions in our Model System The inventor's dataset reveals varied correlations between MPs across time and experimental conditions, with dynamic changes indicating that some MPs adapt over time, reflecting shifts in cellular states due to environmental factors such as organoid development and cancer cell interactions. The inventor's analysis captures these evolving MP relationships, showing how gene expression dynamics unfold across different conditions. The inventors began by examining correlation matrices across conditions (314- and 520-iPSC-derived organoids, and Cancer Spheroids) and timepoints (1.0: 17 days post-attachment; 2.0: 24 days; 3.0: 31 days; 4.0: 38 days).
[0494] In the "520" condition (Fig. 4F), most MPs were dynamic, but MPs such as MP 13, MP 15, MP 17, and MP 19 showed significant changes only in specific intervals, while MP 16 remained stable early on but became dynamic later. In "314" (Fig. 5B), nearly all MPs showed dynamic changes between timepoints 2.0 and 4.0, with only MP 12 and MP 15 remaining stable. In "Cancer Spheroids" (Fig. 4D), many MPs displayed significant changes over time, although MP 22 remained stable throughout. MPs like MP 2, MP 9, MP 10, MP 13, MP 16, MP 18, MP 21, MP 24, MP 26, and MP 29 were stable between timepoints 2.0 and 3.0.
[0495] Comparing the overall correlation matrix (Fig. 4A) with time- and condition-specific matrices (Fig. B- D) provided insights into MP dynamics across conditions. In the "Cancer Spheroid" condition, MPs like MP 6, MP 7, and MP 11 showed high variability at timepoint 1.0, which stabilized at later timepoints. The "314" condition exhibited early variability, especially at timepoint 2.0, with stabilization by 4.0. The "520" condition showed high initial variability, with MPs like MP 11 and MP 24 still dynamic at timepoint 4.0.
[0496] Dynamic behavior was generally higher at early timepoints (1.0 and 2.0) across conditions, indicating that MPs are more adaptable in initial stages. By later stages (3.0 and 4.0), correlations stabilize, reflecting a more settled state. Each condition displays unique trends in MP correlation evolution, highlighting the context-dependent nature of gene interaction dynamics.
[0497] Example 13: Assessment of Invasion and Migration of Cancer Cells in our GB Organoid Model
[0498] To assess the inventor's models for changes as labeled cancer stem cells are introduced into organoids, the inventors measured tumor size, invasion patterns, and cancer cell distribution within each organoid. The inventors generated a stable model conducive to GBSC cultivation, allowing extended studies. The GB organoids were fixed, CUBIC-cleared, IF-stained for SOX2 and TBR1, and imaged to assess GBSC infiltration and migration (Fig. 6 A, B, C, D, Fig 10). A workflow enabled extraction of quantitative 3D information (Fig. 11 A, B), providing spatial coordinates for each segmented cell group and allowing network diagram creation for each organoid (Fig. 11D).
[0499] Qualitative analysis revealed cancer cell distribution patterns over time (Fig. 11). Spatial metrics analysis over four timepoints showed an increase in cell count per organoid, though individual counts varied (Fig. 7A). The mean nearest neighbor distance indicated initial clustering, a slight dispersion at timepoint 3, and closer clustering by timepoint 4, reflecting dynamic changes in spatial distribution (Fig. 6F). Cell density, assessed by Kernel Density Estimation (KDE), progressively decreased, indicating greater dispersal over time (Fig. 6G). Pairwise distance analysis revealed samples with greater mean distances at the earliest time point, suggesting dispersed cells, while later on they were more clustered (Fig. 6H). Compactness, reflecting cell shape and organization, fluctuated, increasing markedly by timepoint 4, suggesting more irregular, elongated structures (Fig. 51).
[0500] Overall, spatial distribution dynamics showed cells starting clustered, dispersing, and forming more complex, less dense arrangements by timepoint 4, as cells proliferated and migrated from the spheroid, leading to increasingly dispersed and irregular patterns. Example 14: Results: Mapping State Transitions in a Patient-Derived Organoid Model
[0501] The inventor's strategy for generating GB organoid models was to implant cells from the GB cell lines into cerebral organoid pre-cultured to a certain age and size. The organoid thus represents the host 'brain', in which the patient-derived GB cells will form our 'tumor'. The cerebral organoids were generated from publicly available hiPSC lines. The inventors have access to patient-derived GB cells lines from the human glioblastoma cell culture (www.hgcc.se, Uppsala) biobank with full genome sequencing of approx 40 patients available. For this study, we utilized the U3013 line.
[0502] The inventors collected GB organoids from four discrete time points: 7, 14, 21 and 28 days postimplantation, to generate two conditions that can be compared with respect to cellular composition and state changes. In all, the tumors are growing in cerebral organoids that from a developmental point resemble a very young brain, but at the later time point, the tumors have grown to a size where hypoxic niches and necrotic zones are likely to have formed on the basis of their size (more than 1 mm in diameter) and the limits of diffusion (200-400 micron). The time series allows for changes in the microenvironment to take place, and as a consequence of the tumor cells migrating and invading into the cerebral organoid, form structures that cause a shift in the state preference of tumor cells.
[0503] Using our in-house protocol for cultivation as a basis, the inventors examined different approaches to induce tumor formation in the generated cerebral organoids. The approach was to co-culture the developing cerebral organoids with labeled cancer cells or tumor spheroids, based on a method allowing rapid and efficient assaying of GB cell invasion into cerebral organoids. The inventors investigated several modes of co-culture but settled on the attachment of tumor cell spheroids with the host cerebral organoid.
[0504] Some cells in the HGCC biobank have been used to systematically profile the pharmacological and genomic profiles of 100 patient-derived 2D GB cell culture models and some mouse xenograft models to a drug library. However, thorough analysis of the genomic profiling of several of these GB lines has shown the cells undergo significant genomic and transcriptional change with each passage when cultured in 2D (Baskaran, et al. Neuro-Oncol. 20, 1080-1091 (2018)). The inventors hypothesized that the cells will behave differently when cultured in 3D, which is one of the objectives to characterize. To validate the developed model, the inventors compare the scRNAseq-data of the GBSC's from the organoids, to native GBSC from GB samples. For the 3D culturing, the inventors aim to identify the differentiated progeny of the founder GBSCs in the organoid model, so that after the GBSCs have been co-cultured in the organoid for a specific span of time, the inventors have taken the organoid models, dissociated them, sorted the cells and separated out the fluorescent GBSC daughter cells via flow cytometry, and ran scRNAseq to examine the genetic signature of said cells and compared them to the expression patterns seen in native tumor GB cells. The aim was to examine to what extent the organoid-GBCs recapitulate the expression patterns in GB.
[0505] After quality control (see Fig. 9 (Example 21), our dataset comprised 2470 cells and 37321 genes, including metadata such as original identity, RNA counts, feature counts, sample information, condition, timepoint, date, labels, barcodes, and various quality control metrics. Cells with more than 15% mitochondrial and 40% ribosomal content were also filtered out (totally 7 cells) (Fig. 9A). 45 cells were identified as possible doublets but since these were evenly distributed in the gene vs count scatterplot (Fig. 9B), there seems to be no doublets present in the data. As expected, after dimensional reduction and clustering, these predicted doubles are still scattered. Fig. 4A presents UMAP projections of malignant cells categorized by condition, sample, and timepoint. In the first plot, malignant cells are segregated based on their conditions (314, 520, Cancer Spheroid, and Cancer Spheroid Matrigel Embedded), with each condition forming distinct clusters, indicating differences in gene expression profiles influenced by these conditions. The second plot, which color-codes cells by sample, reveals a finer segregation, with each sample contributing to distinct regions of the UMAP, suggesting unique transcriptomic signatures for individual samples. The third plot, color-coded by timepoints, displays a clear temporal progression. Early timepoints (1.0) cluster together and transition smoothly through intermediate timepoints (2.0 and 3.0) to the latest timepoint (4.0), which forms its own distinct cluster. This temporal progression highlights dynamic changes in gene expression as the malignancy advances over time.
[0506] The inventors performed the same analysis as on the patient-derived scRNAseq data to identify the organoid model-specific metaprograms (Fig. 4B) and compare them to both the newly identified GB metaprograms and the cancer hallmark metaprograms identified by Gavish et al 2023 (Fig. 4C). Fig. 4B illustrates the Jaccard similarity indices among the cNMFs generated from the organoid samples, highlighting the clustering patterns within the data. Distinct clusters of MPs are evident, particularly noticeable in the blocks along the diagonal of the heatmap. These blocks indicate groups of samples with high within-group similarity and low between-group similarity.
[0507] In Fig. 4C, several noteworthy patterns emerge from the heatmap. For instance, Organoid_Clusterl displays a high similarity with MP 0 and MP 23, as evidenced by the dark-colored cells in these intersections. This indicates that the gene expressions and pathways in Organoid_Clusterl share considerable overlap with those in MP 0 and MP 23 from GB patient samples, suggesting common underlying biological processes or cellular states. Similarly, Organoid_Cluster5 demonstrates a notable similarity with the cell cycle G2 / M hallmark. This suggests that Organoid_Cluster5 shares significant genetic signatures related to cell cycle regulation, which is a critical aspect of tumor proliferation and growth.
[0508] Furthermore, the heatmap reveals that certain organoid MPs, such as Organoid_Cluster4 and Organoid_Cluster6, have moderate to low similarity with most GB patient metaprograms and hallmark gene sets. This variation in similarity levels highlights the heterogeneity within organoid clusters and their differential alignment with patient-derived metaprograms and established hallmarks. Specifically, Organoid_Cluster4 and Organoid_Cluster6 have lower levels of overlap with GB patient metaprograms and hallmarks. This suggests that these clusters may represent distinct genetic profiles or biological states that are not as prevalent in the patient samples or hallmark sets analyzed.
[0509] To assess the prevalence of the identified organoid model-specific and GB patient-derived metaprograms in the sequencing data, the inventors scored the metaprograms and assigned identities to the cells based on the highest scoring metaprogram. The UMAP plots for both GB MP (Fig. 4D) and organoid MP (Fig. 4E) reveal distinct clusters of malignant cells, each colored according to their assigned metaprograms. In 4D, the clusters indicate the presence of specific gene expression profiles that correspond to different functional states or subtypes of malignant cells based on gene metaprogram scoring. Similarly, in 4E, the clusters represent different gene expression profiles associated with the organoid metaprograms, reflecting the functional diversity and differentiation states of the cells within the organoid environment. The clear separation between clusters in both plots suggests robust differentiation based on metaprogram scoring, which could provide insights into the heterogeneity and potential therapeutic targets within malignant cell populations. The consistency of clustering patterns across both GB MP and organoid MP underscores the biological relevance of these metaprograms in defining malignant cell behavior and their response to different microenvironmental contexts.
[0510] The analysis of metaprogram scores over time across three conditions— 314, 520, and Cancer Spheroid— revealed distinct patterns in the Spearman correlation coefficients.
[0511] Metaprograms MP 1, MP 11, MP 13, MP 25, and MP 28 showed strong negative correlations with time points across all conditions, indicating a consistent decrease over time. This trend was particularly pronounced in condition 314, where these MPs had the strongest negative correlations, such as MP 1 (-0.759), MP 13 (-0.811), and MP 28 (-0.765).
[0512] Positive correlations were observed for MP 3 and MP 6 across all conditions, suggesting an increase over time. Notably, the correlation for MP 3 was stronger in the Cancer Spheroid condition (0.361) compared to 314 (0.079) and 520 (0.230).
[0513] Certain metaprograms exhibited condition-specific patterns. The Cancer Spheroid condition displayed unique positive correlations for MPs such as MP 12 (0.369), MP 19 (0.135), and MP 21 (0.215), which were not as prominent in the other conditions. Additionally, MPs like MP 7 and MP 8 showed higher positive and negative correlations, respectively, in the Cancer Spheroid condition compared to 314 and 520. Condition 314 consistently demonstrated the strongest negative correlations overall, particularly for MPs 1, 11, 13, 25, and 28. In contrast, MPs such as MP 2 and MP 17 exhibited weak or insignificant correlations across all conditions, indicating that their scores did not significantly change over time.
[0514] In summary, while some metaprograms exhibited consistent trends across all conditions, others displayed distinct patterns depending on the specific condition. These findings highlight the nuanced differences in metaprogram dynamics over time in different experimental settings.
[0515] Example 15: Results - Assessment of GB Metaprogram Dynamics across Time and Condition in our Model System (extension of Example 12)
[0516] The primary aim of these comparisons is to understand the dynamics and stability of gene expression networks in different experimental conditions and to evaluate how well organoid models replicate the complexity of human cancer as observed in patient-derived samples. This includes:
[0517] The inventors performed a detailed comparative analysis of gene expression dynamics and interactions using correlation matrices derived from the organoid model scRNA-seq data across different experimental conditions and timepoints. The conditions studied include various states of cancer spheroids and organoid models, which were compared to a large dataset of patient-derived samples. The primary objectives of this analysis were to:
[0518] 1) Assess Temporal Dynamics: By examining changes in metaprogram (MP) correlations over various timepoints, the study aims to capture the temporal progression of gene expression and interaction networks. This helps in understanding how gene regulatory mechanisms evolve over time in response to different experimental conditions,
[0519] 2) Evaluate Condition-Specific Behavior: Comparing the MP correlations across different experimental conditions (Cancer Spheroid, Condition 314, Condition 520) helps in identifying condition-specific trends and behaviors. This can provide insights into how different environments or treatments impact gene expression and interaction networks,
[0520] 3) Understanding Stability and Variability: Analyzing the stability of MP correlations over time and across conditions allows researchers to identify which gene interactions are consistently stable and which are more variable. Stable interactions might represent fundamental biological processes, while variable interactions could indicate adaptive responses to changing conditions,
[0521] 4) Model Validation: By comparing the organoid data (from a single patient-derived cancer cell line) with patient-derived data (140+ patients), the study aims to validate the organoid model as a representative system for studying cancer biology. This comparison helps in identifying the strengths and limitations of the organoid model in replicating the complexity and heterogeneity of human cancer, and
[0522] 5) Identifying Key Gene Networks: The study seeks to identify key gene networks and interactions that are preserved across different conditions and timepoints. Understanding these core networks can provide valuable targets for therapeutic interventions and further research into the molecular mechanisms of cancer. The inventor's analysis captures the dynamic changes in correlations between metaprograms, providing insights into how gene expression relationships evolve over time under different experimental conditions. To start, the inventors examined correlation matrices across condition (314- or 520-iPSC derived organoids, as well as Cancer Spheroids that were cultured without being attached to organoids) and timepoint (1.0, which corresponds to 17 days post-attachment of the cancer spheroid to the organoids, 2.0 to 24 days, 3.0 to 31 days, and 4.0 to 38 days.)
[0523] The "520" condition (Fig. 5B) exhibited a high degree of dynamism across all intervals from 1.0 to 4.0. Dynamic MPs included MP 1, MP 2, MP 3, MP 5, MP 6, MP 7, MP 8, MP 9, MP 10, MP 11, MP 12, MP 14, MP 18, MP 20, MP 21, MP 22, MP 24, MP 25, MP 26, MP 27, MP 28, and MP 29. MPs such as MP 13, MP 15, MP 17, and MP 19 showed significant changes primarily in specific intervals. MP 16 demonstrated stability in the initial interval (1.0 to 2.0) but was dynamic in later intervals.
[0524] In the "314" condition (Fig. 5C), dynamic behavior was observed for almost all metaprograms between the 2.0 and 4.0 timepoints. MPs including MP 1, MP 2, MP 3, MP 5, MP 6, MP 7, MP 8, MP 9, MP 10, MP 11, MP 13, MP 14, MP 16, MP 17, MP 18, MP 19, MP 20, MP 21, MP 22, MP 24, MP 25, MP 26, MP 27, MP 28, and MP 29 all showed significant changes. Only MP 12 and MP 15 were found to be relatively stable during this period.
[0525] In the "Cancer Spheroid" condition (Fig. 5D), a significant number of metaprograms (MPs) demonstrated dynamic behavior, indicating substantial changes in their correlations over time. Specifically, MPs such as MP 1, MP 3, MP 5, MP 6, MP 7, MP 8, MP 11, MP 12, MP 14, MP 15, MP 17, MP 19, MP 20, MP 25, MP 27, and MP 28 showed changes between both the 1.0 to 2.0 and 2.0 to 3.0 timepoints. Conversely, MPs like MP 22 remained stable across both intervals, indicating consistent correlations over time. MPs such as MP 2, MP 9, MP 10, MP 13, MP 16, MP 18, MP 21, MP 24, MP 26, and MP 29 exhibited stability in the latter interval (2.0 to 3.0).
[0526] Comparing the correlation matrix examining all cells regardless of time or condition (Fig. 5A) with the matrices sorted by time and condition (Fig. 5B, C, D), several insights into the dynamics of metaprogram (MP) correlations under different experimental contexts were gained. In the "Cancer Spheroid" condition, there is considerable variability in MP behavior when comparing timepoint 1.0 to the aggregated data, with MPs like MP 6, MP 7, and MP 11 showing notable differences. These differences decrease at timepoints 2.0 and 3.0, indicating increasing stabilization of MP correlations over time. In the "314" condition, significant changes are observed in almost all MPs at timepoint 2, indicating a highly dynamic state early in the condition, with some stabilization occurring by timepoint 4.0. The "520" condition shows high variability at timepoint 1.0, with differences decreasing at later timepoints, although MPs like MP 11 and MP 24 still show significant changes even at timepoint 4.0.
[0527] Dynamic behavior is generally higher at early timepoints (1.0 and 2.0) across conditions, indicating that MPs are highly dynamic during the initial stages of the experiment. As conditions mature (moving to 3.0 and 4.0), MP correlations tend to stabilize, reflecting a more settled state in the interaction networks. Each condition displays unique trends in how MP correlations evolve over time, highlighting the context-dependent nature of gene interaction dynamics.
[0528] Comparing the organoid cells (from a single patient-derived cancer cell line) (Fig. 5A) and the patient- derived samples (140+ patients) (Fig. IE) reveals several observations. The organoid cells exhibit high correlations among many metaprograms, suggesting a more consistent or homogeneous interaction pattern within the organoid model. The patient-derived samples show a more diverse correlation pattern, reflecting the biological variability and heterogeneity present across the 140+ patients. Certain metaprograms such as MP 1, MP 2, MP 6, MP 7, and MP 9 show strong and consistent correlations in the organoid cells, indicating stable interactions within this controlled model. However, the patient data reveal more variable correlations, likely due to differences in genetic, environmental, and treatment backgrounds among the patients.
[0529] Specific metaprograms like MP 5 and MP 13 exhibit contrasting behaviors between the two datasets. In the organoid cells, these MPs might have specific and strong correlations, while in patient samples, their interactions can be significantly weaker or more variable. The patient-derived data matrix includes both strong positive and negative correlations, indicating complex interactions that may not be as prevalent in the organoid model. The uniformity in the organoid model suggests it could be a useful tool for studying fundamental aspects of cancer biology and testing therapeutic interventions in a controlled environment. However, the heterogeneity in patient-derived samples underscores the importance of personalized approaches in cancer treatment, as the diverse correlation patterns reflect the unique interaction networks in different patients.
[0530] Example 16: Methods - Wet Lab Methods
[0531] GB Organoid Model Cultivation:
[0532] The inventor's strategy for creating GB organoid models involved implanting GB cell lines into cerebral organoids pre-cultured to a specific age and size, representing the ’brain’ host, while the patient- derived GB cells formed the 'tumor.' The organoids were made from publicly available hiPSC lines, and GB cell lines were sourced from the Uppsala HGCC biobank, which includes genome-sequenced cell lines from approximately 40 patients. The HGCC biobank provides pharmacological and genomic profiling of GB cell lines in 2D cultures and some mouse xenografts. Previous studies showed significant genomic and transcriptional changes in GB cells at higher passages in 2D culture.
[0533] To assess the extent to which GB organoid models recapitulates actual GB, we compared scRNA-seq data from GBSCs in organoids to native malignant GB cells, focusing on identifying the differentiated progeny of founder GBSCs. For this study, we used the U3013 cell line. Building on our cultivation protocol, we tested various methods to establish tumors in cerebral organoids, settling on coculturing organoids with labeled cancer cell spheroids. Attaching tumor cell spheroids proved most effective for studying GB invasion. GB organoids were collected at 17, 24, 31 and 38 days postimplantation to observe changes in cellular composition and state. The initial time points represented tumors within organoids mimicking a young brain, allowing the microenvironment to transform and tumor cells to migrate and invade, affecting their state preferences. After specified co-culture periods, we dissociated organoids, sorted fluorescent GBSCs, and performed scRNA-seq. This enabled us to analyze their genetic signatures and compare them to native GB expression patterns, assessing how well organoid-GBCs replicated native GB cells' expression.
[0534] For the organoid cultivation, the inventors cultivated them according to an in-house protocol based on the original cerebral organoid cultivation publication. Briefly, two iPSC (314M and 520F, FUJI) lines were plated and expanded in standard cultivation media in 6-well plates. Then, the iPSCs were detached from the plates and dissociated in media to create a single-cell suspension in media per cell line. The single cell suspensions were then cell-counted to determine the cells / mL for each preparation and the appropriate concentration for 10000 cells per well were placed into each well in a curve-bottomed 384-well plate. The cells were then cultured in neural induction media for 2 days to induce the formation of spheroids. Then, each spheroid was embedded in a Matrigel droplet and transferred to a 6-well plate on a shaker (approximately 5-6 embedded spheroids per well) and maintained for 30 days in maturation media.
[0535] For the cancer cell spheroid cultivation, the inventors obtained from the Uppsala HGCC biobank the ZS-green transduced U3013 patient-derived GB cancer cell line, from which the inventors cultivated spheroids in a similar manner as described above for the normal iPSC spheroid cultivation. The spheroids were made by seeding the GB cancer cells in a ULA 96 well plate, and after incubation for 2 days, were attached to the cerebral organoids
[0536] Upon maturation of the organoids, single cancer cell spheroids were attached to the organoids in a Matrigel droplet. The now-formed GB organoids were then transferred to 96-well plates at one GB- organoid per well and placed back on the shaker. After 10 days they were then transferred to a 6-well plate on a shaker (5-6 per well) and maintained for the duration of the experiment.
[0537] Tissue Clearing and Staining:
[0538] The inventors utilized the CUBIC clearing approach to prepare the organoids for 3D imaging. Briefly, the extracted organoids were fixed overnight in 4% PFA, washed in PBS and stored in methanol until the inventors were ready to start the staining. The organoids were then transferred to PBS to clean out the methanol at +4C. Next, the deli pidation process was started by transferring the organoids to 50% CUBIC-L solution overnight. The organoids were then transferred to 100% CUBIC-L solution O / N to complete the delipidation process. The inventors then started the immunofluorescent staining (IF) process. Please note that all organoids in the 3D imaging portion of this study were handled in individual Eppendorfs but using the same master mix for all steps. After washing the organoids in PBS for 2+ hours at +4C, the organoids were placed in blocking solution overnight. Next, the organoids were placed in primary Ab solution for 5 days at +4C on a rocker. The primary Ab solution was replenished daily. Primary Abs included SOX2 (SC-365964, 1:500) and TBR1 (Abeam ab31940, 1:2000). The organoids were then washed 5x times for 1 hour with 0.1% Triton X-100 in PBS. Then, the organoids were placed in a secondary Ab solution for 4 days, with the solution replaced for each sample every day. The samples were again washed 5x times for 1 hour with 0.1% Triton X-100 in PBS. Upon completion of the IF staining, the organoids were washed in PBS for 2 hours and fixed in 1% PFA overnight.
[0539] To complete the CUBIC clearing and prepare the organoids for light sheet imaging, the organoids were washed in PBS for 2 hours before being embedded in 2% agarose gel and trimmed to size. For refractive index-matching, the embedded organoids were then placed in 50% CUBIC-R solution until they sank, upon which they were transferred to 100% CUBIC-R solution overnight.
[0540] Light Sheet Imaging:
[0541] The agarose-embedded organoids were then each placed into an imaging chamber containing imaging solution. Images were acquired on a Zeiss Lightsheet Z.l using a 5x detection objective (EC- Plan-Neuofluar-0.16NA) and 5x illumination objectives (light-sheet fluorescence microscopy, 0.2NA). Samples were illuminated from both directions and Z-stacks were acquired for the entire organoid using a lx optical zoom, 1.31 pm light-sheet thickness, and 2.92 pm Z-intervaL Stitched images were generated as necessary for larger samples.
[0542] Dissociation & Fluorescence-activated cell sorting (FACS) of GB Organoids:
[0543] At the relevant time points, GB organoids were extracted from culture, dissociated had FACS performed to extract the ZsGreen-expressing cancer cells. The cancer cells were then sorted into lysis buffer plates and frozen in preparation for SmartSeq3 sequencing. FACs was performed on an Arial 11 (Beckton Dickinsson) cell sorter, with the following settings. Sort setup 100 micron, frequency 30.0, amplitude 12.5, phase 0.00, Drop Delay 26.18, purity mask 32, phase mask 16, plates voltage 2500, voltage centering -137).
[0544] SmartSeq3:
[0545] Individual cell sequencing libraries were generated using the protocol outlined in Hagemann-Jensen et al. 2020 (Single-cell RNA counting at allele and isoform resolution using Smart-seq3. Nat Biotechnol. 2020 Jun;38(6):708-714. doi: 10.1038 / s41587-020-0497-0. Epub 2020 May 4. PMID: 32518404.), with minor improvements as detailed in Smart-seq3 V.3 on protocols. io (link). The lysis buffer contains ERCC spike in molecules (Ambion, 4456740) diluted 1:80000000 of the original concentrations in the RT reaction volume. Samples were sequenced on NovaSeq6000 (NovaSeq Control Software 1.8.0 / RTA v3.4.4) with a 85nt(Readl)-10nt(lndexl)-10nt(lndex2)-133nt(Read2) setup using 'NovaSeqStandard' workflow in 'SP' mode flowcell. The Bel to FastQ. conversion was performed using bcl2fastq_v2.20.0.422 from the CASAVA software suite. The quality scale used is Sanger / phred33 / Illumina 1.8+. Each flowcell was demultiplexed plate-wise. The output data was then further demultiplexed cell-wise and processed with the zUMIs pipeline.
[0546] Example 17: Methods - Computational Methods Metaprogram Generation:
[0547] To identify the heterogeneous transcriptomic programs present across the samples, the inventors have applied an approach similar to the one pursued by Gavish et al 2023, but with the aim of characterizing GB tumor-type specific programs rather than pan-tumor programs. We performed consensus Non-negative Matrix factorization (cNMF) on the data of each individual patient to identify the programs of each sample. Since application of cNMF requires a "K" parameter that influences the results, the inventors run cNMF using different values (l<=5,6,7,8,9,10) and generate 45 programs for each tumor. Each cNMF program is summarized by the top 100 genes representing that program based on cNMF coefficients.
[0548] Then, the inventors identified the most robust cNMF programs across the patient cohort as those that recur within the tumor (gene lists have at least 70% overlap), recur across tumor (have at least 20% overlap with any other cNMF program in other patients analyzed), and are non-redundant within the tumor (rank programs by similarity with programs from other tumors, then remove programs that have at least a 20% overlap with other programs within the same patient).
[0549] From the robust cNMF programs, the inventors clustered them together based on their Jaccard similarity and identified the most consistent genes present across the grouped cNMF programs. This is how the inventors end up with our meta-programs. The clustering process resulted in 30 metaprograms (clusters), each containing samples with high internal similarity.
[0550] Assessment of Clustered cNMF's Comprising Each Metaprogram:
[0551] To validate the resulting metaprograms, the inventors assessed the overall similarity via Jaccard indices of eachcNMF assigned to a cluster and their tendency to be more similar to cNMFs within said cluster than outside of it. The Jaccard index, defined as the size of the intersection divided by the size of the union of two gene sets, was computed using the formula Jaccard lndex=|AnB||AUB|\text{Jaccard Index} = \frac{|A \cap B | }{ | A \cup B | }Jaccard lndex=|AUB||AnB|, where A and B are two gene sets. First, the inventors calculated the Jaccard similarity indices for all pairs of cNMFs within a dataset, resulting in a similarity matrix. This matrix represents the proportion of shared features between pairs of samples. Next, the inventors grouped the cNMFs together based on the metaprogram clustering which grouped the samples into distinct clusters based on their similarity indices. The similarity matrix could then be visualized as a heatmap to provide a clear visual representation of the internal similarity of the cNMFs within clusters and the separation between different clusters.
[0552] Gene Set Enrichment Analysis (GSEA) of Glioblastoma Metaprograms
[0553] The inventors conducted Gene Set Enrichment Analysis (GSEA) on predefined glioblastoma metaprograms using the clusterProfiler package in R to identify enriched pathways associated with each metaprogram. For each metaprogram, the top 8000 variable genes in the GB dataset were ranked for malignant cells assigned to that metaprogram based on their metaprogram scores. Specifically, cells were assigned to a metaprogram if it held the highest score among all metaprograms for that cell, ensuring that the gene list for each metaprogram reflected the most relevant transcriptomic features. To facilitate this analysis, gene sets were loaded from the Molecular Signatures Database (MSigDB) using the msigdbr package, specifically focusing on the C2 category to capture curated gene sets with biological relevance to cancer. Each metaprogram's ranked gene list was loaded from CSV files, containing Ensembl gene identifiers and corresponding log fold change values. Gene sets from MSigDB were filtered to include only those relevant to glioblastoma and neural biology, using keyword matching on terms such as "glioma," "brain," "oncogenesis," and "neurogenesis" to capture pathways pertinent to the biological context. For each metaprogram, genes were ranked by log fold change values, and GSEA was performed with a p-value cutoff of 0.05 to identify significantly enriched pathways. Pathways with significant normalized enrichment scores (NES) were identified, and only those with positive NES values were selected for visualization
[0554] Cumulative Expression of Metaprograms:
[0555] Malignant cells from the Core GBMap were identified based on the iCNV attribute, specifically selecting those labeled as 'aneuploid'. Metaprograms were defined by the gene sets from each of the 30 metaprograms (Table 2 / GB_MP_FigureVersion.csv). The gene expression data for the malignant cells was normalized using StandardScaler and log-transformed to emphasize relative differences in expression levels. For each metaprogram, cumulative expression was calculated by summing the expression levels of its constituent genes. To quantify the relative expression score for each cell, the cumulative expression values were centered around zero, with positive values indicating higher expression and negative values indicating lower expression relative to the mean. Hierarchical clustering was performed on the cells using the average linkage method, and the resulting order was applied to the cumulative expression data.
[0556] Scoring of Metaprograms in Cells:
[0557] To analyze the distribution of cells across different metaprograms, the inventors used scanpy's tl.score_genes function to calculate gene scores for each metaprogram, using the list of genes present within the data. The inventors excluded specific MPs ('MP O', 'MP 4', 'MP 5', 'MP 14', 'MP 16', 'MP 19', 'MP 22', 'MP 23', and 'MP 27') from the analysis because they represented cycling (MP 0 and MP 23), noise (MP 4), or were more indicative of specific functions rather than interconnected plastic states in malignant cells. Each cell was then assigned to the MP with the highest score.
[0558] Assessment of Metaprogram Residuals Across Patient and Subclones
[0559] The inventors first used SCEVAN32 to identify subclones within malignant cells based on their CNV profiles. To examine the relationship between subclones and MP assignments within individual patients, the inventors performed residual calculations using a chi-squared test, analyzing each patient separately. The inventors identified unique patients in our dataset by the patientjd field in scRNA-seq annotations. For each patient, data was grouped by subclone and MP assignment, creating a contingency table to represent the distribution of MP assignments across subclones. To avoid issues related to zero values, we added a small constant (0.5) to each cell. A chi-squared test of independence was then conducted on each patient's contingency table to assess the significance of the association between subclone and MP assignment. This test generated a chisquared statistic, degrees of freedom, and p-value, indicating whether the observed MP distribution within subclones deviated significantly from the expected distribution. Subsequently, the inventors calculated residuals by subtracting the expected counts from observed counts in each cell. Positive residuals indicate overrepresentation of an MP within a subclone, while negative values indicate under-representation.
[0560] TCGA BRCA Analysis This study utilized publicly available gene expression and clinical data from The Cancer Genome Atlas (TCGA) for glioblastoma (GB) to conduct a survival analysis. The methodology encompassed multiple steps, including data acquisition, processing, and statistical analysis. Specifically, RNA-Seq gene expression data was retrieved, processed, and combined with clinical data to generate survival curves for metaprograms identified in GB. The GB IDH-WT samples were extracted from the TCGA using the information from this study that recategorized the TCGA-GB samples by current standards.
[0561] Cox Proportional Hazard Regression Analysis
[0562] Gene expression data for the entire TCGA-GB cohort was obtained using the TCGAbiolinks package. The query targeted the "Transcriptome Profiling" data category, specifically RNA-Seq data processed with the STAR workflow, focusing on "Gene Expression Quantification." Data was downloaded and processed using the GDCprepare function to create the FPKM-UQ (Fragments Per Kilobase of transcript per Million mapped reads - Upper Quartile normalized) gene expression matrix. Clinical data for the TCGA-GB cohort was retrieved using the GDCquery_clinic function, and relevant survival information (patient barcode, days to last follow-up, and vital status) was selected for further analysis. Univariate Cox proportional hazards regression was performed for each gene in the matched expression data using the coxph function from the survival package, assessing associations between gene expression levels and patient survival (time to last follow-up and vital status). The hazard ratio (HR), 95% confidence intervals, and p-values for each gene were calculated and saved in a results table. Significant results (p < 0.05) were filtered and merged with metaprogram gene lists derived from previous studies, categorizing genes into specific metaprograms associated with malignant, nonmalignant, and reference normal brain tissue samples. The genes identified as significant in the Cox regression analysis were used to refine the metaprogram gene lists, excluding non-significant genes, and saved in CSV format for continued exploration.
[0563] Kaplan-Meier Survival Curve Analysis
[0564] The Cox-filtered gene lists were loaded from the previous step. Clinical data for the TCGA-GB cohort was retrieved using the GDCquery_clinic function, providing essential survival information such as vital status and survival time for each patient. An external Excel file with reclassification data was used to identify and extract GB-specific patient barcodes, which were matched with the clinical data. Necessary columns from the clinical data were selected and formatted, including converting time- related variables to numeric values and categorizing patient vital status as "Alive" or "Dead." RNA-Seq gene expression data for the TCGA-GB cohort was queried and downloaded via the TCGAbiolinks package. The data was normalized using variance-stabilizing transformation (VST) from the DESeq2 package to ensure consistency in scaling for downstream analyses. For each metaprogram, gene expression values within the metaprogram were extracted from the normalized gene expression matrix. These values were averaged to compute a "module expression" score for each patient, which was then stratified into "LOW" or "HIGH" groups based on median expression. The clinical and gene expression data were merged to enable survival analysis, incorporating both the calculated module expression scores and relevant clinical variables (e.g., overall survival time and vital status). A Cox proportional hazards model was fitted to assess the effect of each metaprogram's expression on patient survival. Survival curves were generated using the KaplanMeier method, with significance assessed through log-rank tests. The survfit function from the survival package was used to fit the survival curves, and the ggsurvplot function from the survminer package was employed for visualization.
[0565] Identification of Second-highest scoring Meta prog ram in Cells:
[0566] To generate the heatmap (Figure ID), the inventors first calculated the second-highest MP score for each cell by identifying the second-largest value in each row of the scores dataframe. The inventors added the information about the second-highest MP to the observation dataframe under the column second_highest_MP and created a Boolean column is_hybrid to flag cells that have a second-highest MP score. The inventors then identified the most likely hybrid for each primary MP by grouping the data by the primary MP assignment and determining the most frequent second-highest MP for each group. Next, the inventors constructed a matrix to count the occurrences of each MP being the second-highest score for each primary MP assignment, excluding self-assignments. This matrix was then normalized to show percentages by dividing the counts by the total number of cells assigned to each primary MP and multiplying by 100.
[0567] Correlation of Metaprograms Across Cells:
[0568] The generation of the correlation matrix involved the dataset utilized for this analysis comprising metaprogram (MP) scores for various cells. The correlation coefficients, specifically Pearson correlation coefficients, were calculated using the corr() method in pandas.
[0569] Comparison of GB Metaprograms to Reference Metaprograms:
[0570] The analysis involved three datasets: the Cortical2023 dataset, which contains the top 50 genes per cluster; the Fetal dataset, which includes genes per metaprogram; and the Hallmark dataset, representing genes per hallmark. To determine the similarities between these gene sets, the Jaccard similarity index was calculated for each pair of columns (representing gene sets) from the metaprograms dataset and each of the Cortical2023, Fetal, and Hallmark datasets.
[0571] To extract the overlapping genes between the metaprograms, the inventors then developed a custom function to calculate the overlap between genes in each pair of MPs across different datasets. This function iteratively compared each MP from one dataset with every MP from another, identifying common genes and recording the number and identities of overlapping genes. To ensure clarity, selfcomparisons within the same dataset were excluded, and only unique comparisons were considered to avoid redundancy (e.g., avoiding both MP0 vs. MP29 and MP29 vs. MP0). For each comparison, the inventors generated two dataframes: one to capture the count of overlapping genes and another to store the actual overlapping gene names. These dataframes were then used to filter overlaps with five or more genes, which were considered significant for further analysis. The results were saved into CSV files for documentation and further review. Finally, the significant overlaps were formatted and printed to provide a comprehensive view of the gene overlaps across different metaprogram datasets.
[0572] CNV Assessment in GB Patient and Organoid Data To generate the CNV heatmap, the inventors first filtered the malignant cell data to include only specific patient IDs. The inventors then created a DataFrame from the CNV data, ensuring all values were numeric, and included donor IDs for grouping purposes. To generate the CNV heatmap of malignant cells grouped by metaprogram (MP), the inventors used the following steps: loaded the metaprogram and scored the genes for each metaprogram using the sc.tl.score_genes function from the Scanpy library. Cells were sorted by their assigned MP, and the CNV data was reordered accordingly. CNV frequencies were then computed for each cell by applying a lambda function across the CNV DataFrame, determining the presence of CNVs as a frequency value. To identify CNVs present in a significant proportion of cells, a threshold was set at 50%. For each metaprogram, CNVs with high frequency within the group but low frequency in others were identified using specific thresholds (60% and 40%, respectively).
[0573] Assessment of Temporal Correlation of Metaprograms in the Organoid Model:
[0574] In this study, the inventors analyzed the correlation between metaprogram (MP) scores and timepoints for three different conditions: 314, 520, and Cancer Spheroid. The inventors used the scanpy library to preprocess the data and compute the necessary scores. The inventors scored the genes for each metaprogram using the scanpy.tl.score_genes function. The inventors excluded specific metaprograms, namely MP 0, MP 23, and MP 4, from further analysis. To facilitate correlation analysis, the inventors factorized the timepoint variable to a numeric representation (timepoint_numeric). The inventors then focused our analysis on three specific conditions: 314, 520, and Cancer Spheroid. For each condition, the inventors calculated the Spearman correlation between each MP score and the numeric timepoint.
[0575] Assessment of Metaprogram Correlation Across Time and Conditions:
[0576] To analyze changes in gene expression correlations over time, correlation matrices for different conditions and timepoints were generated from a malignant dataset.Conditions and their respective timepoints included Cancer Spheroid (timepoints 1.0, 2.0, 3.0), 314 (timepoints 2.0, 4.0), 520 (timepoints 1.0, 2.0, 3.0, 4.0), and Cancer Spheroid Matrigel Embedded (timepoint 3.0). To quantify changes over time, a custom function was used to compute the difference between successive timepoints' correlation matrices for each condition. This function iteratively subtracted the correlation matrix of the earlier timepoint from that of the later timepoint, storing the resulting difference matrices in a dictionary keyed by the compared timepoints (e.g., 1.0_vs_2.0).
[0577] Image Segmentation & Cancer Cell Coordinate Assessment:
[0578] The images were normalized with background correction to enhance contrast of cell vs background. Cell segmentation was then performed with the Voronoi-Otsu-Labeling workflow (a combination of Gaussian blur, spot detection, thresholding and binary watershed) in 3D. Cell coordinates were obtained of the segmented malignant cells. Network diagrams were then built via Delaunay Triangulation from all cells. The inventors then estimated migration from distance to the spheroid center, in the number of steps in the diagram.
[0579] 3D Cell Mapping: In this study, the inventors analyzed spatial point cloud data from cancer organoid models across four distinct timepoints. The data, consisting of spatial coordinates (pX, pY, pZ) for each cell, was processed to derive several spatial metrics to understand changes in cell distribution and organization over time. First, nearest neighbor distances were calculated for each sample using a k-dimensional tree (KDTree) to determine the distance to the closest neighboring cell. Kernel Density Estimation (KDE) was then applied to estimate the density of cells within the organoid at each timepoint.
[0580] Clustering analysis was performed using the DBSCAN algorithm to identify clusters of closely packed cells. Voronoi tessellation was conducted for a subset of samples to examine the spatial partitioning of cells, although this was limited to 2D (pX, pY) data. Spatial autocorrelation was assessed using Moran's I statistic to measure the degree of spatial clustering, though the results were inconclusive due to limitations in the dataset. Principal Component Analysis (PCA) was employed to reduce the dimensionality of the spatial data and identify major patterns in cell distribution. Finally, compactness, a shape descriptor, was calculated based on the Convex Hull of the cell distributions to assess changes in the spatial structure of the organoids over time. All analyses were conducted on data subsets corresponding to individual samples, and the results were aggregated to identify temporal trends.
[0581] Example 18: Discussion
[0582] Metaprogram Identification and Phenotypic Complexity
[0583] The cNMF approach utilized in this study has revealed a nuanced understanding of the phenotypic complexity in glioblastoma (GB) that surpasses previous classification systems, such as the four discrete states identified by Neftel and better capture the plasticity of GB cells. Rather than viewing cells with high expression of multiple metaprograms (MPs) as hybrids, the inventors propose they are plastic, dynamically navigating the landscape of these MPs. Some MPs coexpress more frequently, reflecting a complex network of phenotypic expressions that provide a more detailed understanding of GB cell behavior.
[0584] Functional Roles and States Within Malignant Cell Populations
[0585] The expression patterns observed in Figure 2B suggest various functional roles or states within the malignant cell population, such as differences in metabolic activity, immune response, or cell cycle status. Investigating these patterns further could uncover the molecular mechanisms driving malignancy and potentially identify new therapeutic targets. The skewed distribution of certain MPs, as noted in Figure 2C, emphasizes the heterogeneity within the malignant population and hints that MPs like MP 21 and MP 2 might play critical biological roles.
[0586] Overlap with Cortical, Fetal, and Cancer Hallmark Metaprograms
[0587] The significant overlap between GB-derived MPs and those from fetal, cortical, and cancer hallmark metaprograms underscores the interconnectedness of these gene expression patterns across different developmental and pathological contexts (Figure 3).
[0588] 1. Cortical Metaprograms: The overlap, such as that between MP 0 and Cortical_MP7 and Cortical_MP19, suggests GB cells share proliferative characteristics with cortical cells, potentially adopting neurogenic states. Similarly, the overlap of MP 29 with Cortical_MP5, characterized by neuronal differentiation and synaptic function genes, indicates that GB may retain or hijack neurodevelopmental programs.
[0589] 2. Fetal Metaprograms: The alignment of MP 0 with fetal program F4, involving cell proliferation and mitosis, suggests GB cells may recapitulate mechanisms from early developmental stages, contributing to their aggressive proliferation. MP 29's overlap with fetal program Fl, which includes genes involved in neuronal differentiation, further supports this notion.
[0590] 3. Cancer Hallmark Metaprograms: The overlaps with cancer hallmark metaprograms, such as MP 0 with the Cell Cycle - G2 / M program, highlight the critical role of cell division in GB. The overlap of MP 1 with the Mesenchymal Transition hallmark suggests GB cells may undergo epithelial-to-mesenchymal transition (EMT), enhancing their invasive potential.
[0591] Shared Gene Expression Patterns
[0592] The analysis of shared gene expression patterns among GB MPs has revealed profound insights into the tumor's biology (Figure 4). The significant overlaps, particularly between MP 0 and MP 23, emphasize common pathways driving cell proliferation and survival. Other overlaps, such as those involving mesenchymal transition, immune response, and astrocytic markers, reveal critical processes underlying tumor invasiveness and immune evasion.
[0593] 1. Metabolic Adaptation and Hypoxia: The overlaps between MP 2 and MP 24, involving hypoxia response genes, illustrate GB's metabolic adaptation to hypoxic conditions, contributing to its aggressive behavior and therapeutic resistance.
[0594] 2. Neural Differentiation and Developmental Pathways: Relationships between MPs 3, 12, and 29 suggest a link to neural progenitor cell biology and differentiation processes, highlighting the tumor's exploitation of developmental pathways for proliferation and survival.
[0595] 3. Myelination and Oligodendrocyte Markers: Shared genes among MP 6 and MP 8, linked to myelination, indicate interactions with neural tissues and potential therapeutic targets.
[0596] 4. Immune Response and Epithelial Pathways: Overlaps between MP 13 and MP 25 highlight immune evasion strategies, while MP 28's role in protein maturation and immune response underscores mechanisms supporting tumor growth.
[0597] ITH and Evolution Diversity and Applications of Organoid Models
[0598] The diversity within organoid clusters, as shown in Figure 4B and 4C, captures the biological variability of tumor biology. Different clusters reflect various aspects of tumor heterogeneity, providing a broader understanding of tumor behavior and treatment response. This variability is crucial for personalized medicine, allowing researchers to select relevant models for studying specific tumor characteristics or testing therapies.
[0599] Metaprogram Dynamics in Different Conditions
[0600] Figure 5 A-D reveals the dynamic nature of gene interaction networks across different conditions. In cancer spheroids and other conditions, early dynamism followed by stabilization suggests periods of adjustment and reorganization in gene interactions. Organoid models, while consistent and reproducible, lack the complexity of patient-derived samples, highlighting the importance of personalized approaches in cancer treatment.
[0601] 3D Tumor Structure
[0602] The analysis of spatial metrics within the cancer organoid models provides critical insights into the dynamic nature of tumor architecture over time. As tumors evolve, the spatial distribution and organization of cells play a significant role in shaping their biological behavior and response to treatment. The observed trends in cell clustering, density, and compactness across the four time points suggest an intricate interplay between cellular migration, structural reorganization, and possibly, the differentiation of cell subpopulations within the organoids. These findings highlight the dynamic and evolving nature of tumor structure within the organoid models. The observed trends in spatial metrics suggest that the tumor microenvironment undergoes significant reorganization over time, potentially driven by underlying biological processes such as cell migration, differentiation, or response to microenvironmental cues. The increasing dispersion and irregularity in cell distribution may reflect the heterogeneity of the tumor, with different cell populations occupying distinct niches or engaging in various functional roles within the organoid.
[0603] Conclusion
[0604] The comprehensive analysis of shared gene expression patterns among GB MPs has provided significant insights into the tumor's biology. Identifying overlapping genes highlights potential therapeutic targets, revealing common pathways driving malignancy. Future research should integrate results from model systems and patient data to enhance the translational relevance of preclinical studies, ultimately improving personalized treatment strategies for GB.
[0605] Supplementary Results
[0606] Example 19 (SI): Validation of the Metaprogram Assessment
[0607] For a reference glioblastoma dataset consisting of 23 adult IDH-wildtype GB patients from Neftel et al. (2019), accessed via the GBMap merged and harmonized atlas, the inventors identified eight noncycling metaprograms (N1-N8) in addition to a cycling program (Fig. 7A). These metaprograms showed substantial overlap with the original Neftel programs: for example, N5 aligned with NPC2, N2 with AC, and N3 with MES2. Some metaprograms, such as N4 and N8, showed mixed expression of NPC1 and OPC markers, reflecting hybrid progenitor-like states. In addition, the inventors identified a metaprogram N7 characterized by mature oligodendrocyte / myelination genes (e.g. MOBP, MAL, PLP1, MAG, MBP), corresponding to an oligodendrocyte program later reported by Gavish et al. (2023). Per-patient metaprogram compositions (Fig. 7C) demonstrated that some tumors were dominated by specific MPs (e.g. NPC-like, AC-like, MES-like), whereas others exhibited mixed MP distributions, confirming that the metaprogram framework captures known GB cell states and their patient-specific combinations.
[0608] Example 20 (S2) Overlaps Between Metaprograms and Cortical2023, Fetal, and Hallmarks MPs:
[0609] The inventors next compared metaprograms (MPs 0-29) with external reference programs derived from cortical development ("Cortical2023"), fetal brain atlases, and curated hallmark signatures. This analysis demonstrated extensive and systematic overlap between the GB metaprograms and known biological programs, confirming that the metaprograms capture proliferative, glial, neuronal, hypoxia, stress, and immune-related states.
[0610] Table 1.
[0611] Cycling MPs (MP 0 and MP 23) showed strong overlap with cell-cycle and G2 / M-associated signatures and proliferative cortical programs, including shared genes such as PTTG1, AURKB, UBE2C, MKI67, TPX2, CENPF, PRC1, CDK1, CCNB1 and TOP2A. Astrocyte-like and glial MPs (e.g. MP 1, MP 7, MP 18, MP 21) overlapped astrocyte and immature glia reference sets and included genes such as APOE, SLC1A3, AQP4, GFAP, CST3, and CLU. OPC-like and oligodendroglial MPs (MP 3, MP 6, MP 17) overlapped oligodendrocyte progenitor and mature oligodendrocyte programs, sharing genes such as OLIG1, OLIG2, BCAN, CLDN11, MBP, PLP1, and CNP.
[0612] Hypoxia-associated MPs (MP 2, MP 10, MP 20, MP 24) overlapped hypoxia and metabolic stress signatures, with representative genes including PGK1, ADM, ALDOA, CA12, BNIP3, BNIP3L, VEGFA, PDK1, LDHA and NDRG1. Stress and DNA-damage-associated MPs (e.g. MP 10, MP 19, MP 22) showed enrichment for chaperones and DNA-repair genes (for example HSPA1A, HSPA1B, HSPA6, DNAJB1, PCNA, DNMT1, RAD51AP1, ATF4, DDIT3, GADD45A). Interferon and immune-related MPs (MP 5, MP 13, MP 25, MP 26, MP 28) overlapped interferon / MHC-ll, secreted-factor and inflammation programs, including genes such as HLA-A, HLA-C, B2M, ISG15, MX1, CXCL10, CXCL11, BST2, NAM PT, BIRC3, TNFAIP3 and CCL2.
[0613] Neuronal and progenitor-like MPs (MP 3, MP 8, MP 9, MP 12, MP 15, MP 29) overlapped multiple neuronal and progenitor signatures across cortical, fetal, and NPC-glioma references, with representative overlapping genes such as DCX, STMN2, UCHL1, MAP1B, SOX4, ELAVL4, NRXN1, GRIA2, and GNG3. Together, these overlaps demonstrate that each GB metaprogram corresponds to a coherent and externally anchored biological state rather than an arbitrary cluster. Detailed gene-by- gene overlap counts for selected MP-reference combinations is provided in table 1.
[0614] Example 21 (S3) Sample Specific MP correlations
[0615] The inventors assessed pairwise correlations among GB metaprograms across patient samples to characterize relationships between different malignant states. Certain metaprogram pairs showed consistently positive correlations across tumors. For example, astrocyte-like MPs (MP 1 and MP 7) and OPC-like MPs (MP 3 and MP 17) were positively correlated in most samples, indicating stable lineage-related co-expression modules. These associations suggest conserved astrocytic and oligodendroglial differentiation trajectories within malignant GB cells.
[0616] Other pairs showed consistently negative correlations, such as AC-like MP 1 versus MES-like MP 11, or NE-like hypoxia MP 2 versus interneuron-like MP 12, consistent with opposing differentiation or stress-response programs. Many MP pairs, particularly those involving hypoxia- and inflammation- associated MPs (e.g. MP 2, MP 10, MP 13, MP 28), exhibited mixed correlation signs across samples, reflecting context-dependent relationships driven by local hypoxia, immune activity, or metabolic status.
[0617] Overall, the analysis indicates that some MP relationships (e.g. AC-like AC-like, OPC-like OPC- like) are stable and recurrent across patients, while others vary in a manner consistent with phenotypic plasticity and microenvironmental adaptation. This supports the use of metaprogram correlations as a flexible representation of GB state architecture that can be integrated into digitaltwin and longitudinal modelling.
[0618] Example 22 (S4) : Quality Control of scRNAseq Data
[0619] In one embodiment, single-cell RNA-sequencing (scRNA-seq) data are processed using a standard quality control and normalization workflow. Briefly, the raw output matrix is constructed using unique molecular identifiers (UMIs) to correct for PCR amplification, and both exonic and intronic reads may be included to capture a broad range of transcriptional states. Empty wells are removed, and genes are annotated (e.g. mitochondrial, ribosomal, hemoglobin and spike-in classes).
[0620] Per-cell quality metrics, such as detected gene count, total counts, and the percentage of mitochondrial and spike-in reads, are computed (Fig. 9A). Low-quality cells with low gene detection and / or high spike-in content are filtered out using modality-appropriate thresholds; for example, a cell may be required to exhibit at least about 2,000 detected genes and a spike-in fraction below about 0.2%. Genes detected in only a small number of cells (for example fewer than three cells) may be removed. Putative doublets can be evaluated with a standard tool such as Scrublet (Fig. 9B); in one dataset, the predicted doublets did not form a coherent cluster and appeared evenly distributed, consistent with an absence of problematic doublets.
[0621] After QC, library size normalization is performed so that each cell is scaled to a common target count, followed by log-transformation (e.g. loglp). Highly variable genes (HVGs) are identified (for example several thousand HVGs), and optional regression of cell-cycle scores may be applied to emphasize other sources of variation. The data are then scaled, principal component analysis (PCA) is performed on HVGs, and a low-dimensional embedding such as UMAP is generated using the leading principal components. In one implementation, this procedure yields a dataset with approximately 2,000-3,000 high-quality cells and on the order of 30,000-40,000 expressed genes, accompanied by metadata such as sample identity, condition, timepoint, labels, barcodes and QC metrics.
[0622] Example 23 (S4): Time series Visualization of Cleared, IF Stained and Imaged GB Organoids
[0623] In one embodiment, patient-derived GB organoids are immunostained, tissue-cleared, and imaged at multiple timepoints to generate three-dimensional (3D) volumetric datasets. These time-series images, exemplified in Figures 10 and 11, illustrate changes in organoid size, morphology, and marker expression over time and under different treatment conditions. Such volumetric data are integrated into the digital twin as one modality contributing to state estimation and longitudinal divergence metrics.
[0624] Example 24 - Limited Relationship between Genetic Subclones and Intra-tumoral State Diversity
[0625] In some embodiments, large-scale copy-number alterations (CNAs), such as chromosome-arm gains and losses, are inferred per cell from the scRNA-seq data using established expression-based methods. This enables identification of multiple CNV-defined genetic subclones within a tumor. The distribution of metaprogram scores can then be evaluated for each subclone to assess whether particular subclones are confined to specific MP-defined states. In one analysis, each of several dozen CNV-defined subclones contained cells spanning multiple metaprogram states, and only a minority of subclones showed a strong bias toward a single state. This suggests that much of the intra-tumoral diversity in metaprogram states is not strictly determined by CNV subclones, but reflects non-genetic plasticity and microenvironmental influences, consistent with prior observations in other glioma subtypes.
[0626] Example 25: Supplementary Notes - Description of Metaprograms
[0627] The following notes summarize exemplary glioblastoma metaprograms (MPs 0-29) identified in malignant cells. Each MP corresponds to a coherent gene-expression module linked to specific biological functions. The invention is not limited to any individual gene or exact gene list; a skilled person can define equivalent metaprogram signatures based on these principles. MP 0: Cycling / Proliferation 1
[0628] • Brief description: MP 0 is a proliferative program enriched for G2 / M cell-cycle genes and mitotic regulators.
[0629] • Representative genes: PTTG1, AURKB, UBE2C, CDCA3, BIRC5, NUSAP1, MKI67, TPX2, CENPF, PRC1, DLGAP5, SMC4, CENPE, KIF2C, CDK1, CCNB1, TOP2A.
[0630] • Key overlaps: strongly overlaps external proliferative and G2 / M programs from cortical and fetal datasets and Hallmark cell-cycle signatures.
[0631] MP 1: Astrocyte-like 1 (AC-like 1)
[0632] • Brief description: MP 1 is enriched for astrocytic markers and extracellular matrix / immune- modulating genes, consistent with astrocyte-like malignant states.
[0633] • Representative genes: SPARC, CIS, C1R, IGFBP7, VIM, SPARCL1, ID3, SERPING1, APOE, SLC1A3, AQ.P4, PON2, ALDOC, ID4, AGT, TTYH1, GJA1, ATP1B2, NTRK2, CST3.
[0634] • Key overlaps: overlaps astrocyte reference programs and mesenchymal / interferon signatures in external datasets.
[0635] MP 2: Neuroepithelial-like Hypoxia 1 (NE-like Hypoxia 1)
[0636] • Brief description: MP 2 is a neuroepithelial-like hypoxia program enriched for metabolic and hypoxia-response genes.
[0637] • Representative genes: VIM, CAV1, LGALS3, PGK1, ADM, ALDOA, HILPDA, CA12, BNIP3L, IGFBP3, ENO2, GBE1, PLOD2, SLC2A1, ERO1A, VEGFA, BNIP3, EGLN3, LDHA, NDRG1. • Key overlaps: overlaps hypoxia / metabolic signatures and neuroepithelial programs.
[0638] MP 3: Oligodendrocyte Progenitor-like 1 (OPC-like 1)
[0639] • Brief description: MP 3 corresponds to oligodendrocyte progenitor-like and neural progenitor-like states, and overlaps multiple progenitor programs.
[0640] • Representative genes: OLIG1, SCD5, BCAN, SMOC1, FGF12, DLL3, NKAIN4, SOX8, NEU4, GPR17, SCRG1, TNR, OLIG2, CA10, SOX4, DCX, STMN4, CD24, MLLT11.
[0641] • Key overlaps: overlaps OPC and NPC glioma programs and external oligodendrocyte progenitor signatures.
[0642] MP 4: Ribosomal
[0643] • Brief description: MP 4 consists predominantly of ribosomal genes and is associated with general translational activity rather than a specific malignant state.
[0644] • Representative genes: enriched in ribosomal structural components (e.g. RPL and RPS family genes).
[0645] MP 5: Immune / Inflammation
[0646] • Brief description: MP 5 is associated with antigen presentation and immune activation, particularly interferon / MHC-ll-related functions.
[0647] • Representative genes: CD74, HLA-DQB1, HLA-DRB1, HLA-DMA, HLA-DRA.
[0648] • Key overlaps: overlaps interferon / MHC-ll programs. MP 6: Oligodendrocyte-like (OC-like)
[0649] • Brief description: MP 6 reflects mature oligodendrocyte and myelination programs and is enriched for proneural-like signatures.
[0650] • Representative genes: CLDN11, APOD, UGT8, TF, MBP, MAG, SLC44A1, SIRT2, PLP1, TUBB4A, BINI, CNP, PTGDS.
[0651] • Key overlaps: overlaps oligodendrocyte progenitor and mature oligodendrocyte programs and proneural signatures.
[0652] MP 7: Astrocyte-like 2 (AC-like 2)
[0653] • Brief description: MP 7 is a glial program related to AC-like 1, with strong enrichment for glial markers.
[0654] • Representative genes: RFX4, LRIG1, ANKFN1, NTRK2, GLIS3, GFAP, NFIA, PACRG, DCLK2, AQP4, GPM6A.
[0655] • Key overlaps: overlaps astrocyte reference programs and immature glia signatures.
[0656] MP 8: Neural Stem Cell-like (NSC-like)
[0657] • Brief description: MP 8 is associated with neural stem / progenitor markers and neuroepithelial functions.
[0658] • Representative genes: NES, VIM, HOPX, RAMP1, BCAN, SLC1A3, ATP1B2, GATM, PON2, CST3, TTYH1, HEPN1, SCRG1, FABP7, SIOOB.
[0659] • Key overlaps: overlaps neuroepithelial, astrocyte, and oligodendrocyte progenitor programs. MP 9: Radial Glia-like 2 (RG-like 2 / intermediate)
[0660] • Brief description: MP 9 reflects an intermediate glial-progenitor state with features of NPC and OPC as well as neuronal connectivity.
[0661] • Representative genes: MAP2, MMP16, CADM2, GRIA2, TANC2, BRINP3, NPAS3, GRID2, PTPRZ1, PTN, SEMA5A.
[0662] • Key overlaps: overlaps NPC / OPC and radial glia-related programs.
[0663] MP 10: Neuroepithelial-like Hypoxia 2 (NE-like Hypoxia 2)
[0664] • Brief description: MP 10 is a stress and hypoxia program with neuroepithelial enrichment.
[0665] • Representative genes: HSPA6, DNAJA1, HSPH1, DNAJB1, PPP1R15A, HSPA1A, HSPA1B, HILPDA, NDRG1, PDK1, PGK1, INSIG2, PLOD2. • Key overlaps: overlaps stress and hypoxia signatures.
[0666] MP 11: Mesenchymal-like (MES-like)
[0667] • Brief description: MP 11 is a mesenchymal program in glioma, enriched for Verhaak mesenchymal signatures and associated with invasion and ECM remodelling.
[0668] • Representative genes: FRMD5, CLIC4, CD44, NAMPT, GAP43, AKAP12, ANXA2, VCL, TNC, SAMD4A, EMP1, WWTR1, VMP1, CAMK2D.
[0669] • Key overlaps: overlaps mesenchymal glioma signatures.
[0670] MP 12: Interneuron-like (IN-like) • Brief description: MP 12 corresponds to interneuron-like and neural differentiation programs.
[0671] • Representative genes: STMN2, SOX4, DCX, PKIA, STMN4, KIF5C, SYT1, BASP1, MIAT, MEG3, PAK3, SOX11, DLX6-AS1, NREP, ELAVL4.
[0672] • Key overlaps: overlaps NPC and interneuron-like glioma programs and cortical neuronal sets.
[0673] MP 13: Neuroepithelial-like Inflammation 2 (NE-like Inflammation 2)
[0674] • Brief description: MP 13 combines neuroepithelial features with interferon and immune- response genes.
[0675] • Representative genes: CXCL10, CXCL11, HLA-A, HLA-C, HLA-E, MX1, IFI35, ISG15, OAS1, OAS2, OAS3, IFIT3, GBP1, STAT1. • Key overlaps: overlaps interferon / MHC-ll programs and inflammatory signatures.
[0676] MP 14: Post-Transcriptional Regulation
[0677] • Brief description: MP 14 is associated with chromatin organization, RNA processing and post- transcriptional regulation.
[0678] • Representative genes: ARGLU1, KCNQ1OT1, MACF1, GOLGB1, DST, NEAT1, ZBTB20, MALAT1, SON, AKAP9, LUC7L3, DDX5, KMT2C, PNISR.
[0679] MP 15: Neuron-like (N-like)
[0680] • Brief description: MP 15 reflects canonical neuronal functions, including synaptic transmission and neural differentiation.
[0681] • Representative genes: CELF4, ANKS1B, SYT1, MYT1L, FRMD4A, MEG3, ASIC2, FRMPD4, KHDRBS2, RYR2, ZNF385B, RBFOX1, SLC35F3, PRKCB.
[0682] • Key overlaps: overlaps cortical neuronal programs and NPC glioma signatures.
[0683] MP 16: Cilia
[0684] • Brief description: MP 16 is a cilium-associated program involved in ciliary structure and motility.
[0685] • Representative genes: CAPS, IFT57, PIFO, CETN2, RSPH1, DNAAF1, KIF9, DNALI1, LRRIQ1, DYNLRB2, F0XJ1, FAM183A, TPPP3, EFHC1, ZMYND10, SPA17.
[0686] • Key overlaps: overlaps cilia-related signatures in cortical and fetal datasets.
[0687] MP 17: Oligodendrocyte Progenitor-like 2 (OPC-like 2)
[0688] • Brief description: MP 17 is a second OPC-like program overlapping MP 3 and enriched for oligodendroglial differentiation and stress-associated genes.
[0689] • Representative genes: MTRNR2L family members, PTPRZ1, BCAN, TCF12, POU5F1, SOX2-OT.
[0690] MP 18: Astrocyte-like 3 (AC-like 3)
[0691] • Brief description: MP 18 is enriched for glial and neural markers and represents a further astrocyte-like / immature glia program.
[0692] • Representative genes: ADGRV1, SLC1A3, SLC1A2, TRPM3, DTNA, NCAN, LRRC3B, GPC5, SFXN5, RGS20, PITPNC1. MP 19: DNA Damage / Replication Stress
[0693] • Brief description: MP 19 is associated with DNA replication, repair and replication stress responses.
[0694] • Representative genes: DNMT1, CENPK, CLSPN, HELLS, RFC4, CENPU, CDC6, RAD51AP1, UHRF1, PCNA, ASF1B, ATAD2, FEN1.
[0695] • Key overlaps: overlaps DNA-damage and replication stress signatures.
[0696] MP 20: Neuroepithelial-like Hypoxia 3 (NE-like Hypoxia 3)
[0697] • Brief description: MP 20 combines hypoxia-related and mesenchymal traits with neuroepithelial enrichment, linking cell motility, ECM remodelling and stress response. • Representative genes: RHOC, SH3BGRL3, TAGLN2, S100A10, S100A6, ANXA2, VIM, EMP1,
[0698] TNFRSF12A, LGALS3, S100A16, LGALS1, ALDOA, IGFBP3.
[0699] • Key overlaps: overlaps hypoxia, EMT and mesenchymal glioma signatures.
[0700] MP 21: Radial Glia-like 2 (RG-like 2)
[0701] • Brief description: MP 21 reflects radial glia-like and immature astrocyte / neuron traits.
[0702] • Representative genes: SPON1, NPAS3, QKI, PTPRZ1, NFIA, CST3, NCAN, SLC4A4, SLC1A3, ATP13A4, SLC1A2, SHROOM3, DTNA, RORA, ATP1A2.
[0703] MP 22: Stress (in vitro)
[0704] • Brief description: MP 22 is a stress-response program associated with integrated stressresponse and amino-acid metabolism genes, particularly in vitro. • Representative genes: CEBPG, CEBPB, SHMT2, NUPR1, SNHG15, XBP1, SLC3A2, PHGDH,
[0705] SQSTM1, EIF1B, ATF4, ASNS, TXNIP, ATF3, HERPUD1, TRIB3, DDIT3, GADD45A, PSAT1, HSPA9.
[0706] MP 23: Cycling / Proliferation 2
[0707] • Brief description: MP 23 is a second proliferative program, also enriched for G2 / M cell-cycle genes and mitotic regulators, closely related to MP 0.
[0708] • Representative genes: PTTG1, UBE2C, BIRC5, UBE2S, CDKN3, NUSAP1, KNSTRN, MKI67, TACC3, TPX2, KPNA2, CENPF, PRC1, CKAP2, SMC4, CENPE, KIF2C, CDC20, GTSE1, PLK1, KIF20B, CDK1, AURKA, CCNB1, TOP2A, CCNB2.
[0709] MP 24: Neuroepithelial-like Hypoxia 4 (NE-like Hypoxia 4)
[0710] • Brief description: MP 24 is a hypoxia-associated program overlapping NE-like Hypoxia 1 and emphasizing glycolytic and oxygen-response genes.
[0711] • Representative genes: GBE1, PFKP, VEGFA, BNIP3L, ERO1A, PDK1, NDRG1.
[0712] MP 25: Neuroepithelial-like Inflammation 3 (NE-like Inflammation 3)
[0713] • Brief description: MP 25 combines neuroepithelial features with antigen presentation and immune regulation.
[0714] • Representative genes: BST2, LGALS3BP, HLA-A, HLA-C, B2M.
[0715] MP 26: Mesenchymal-like Inflammation (MES-like Inflammation) • Brief description: MP 26 is a mesenchymal-inflammation program enriched for secreted inflammatory mediators and overlaps with MES-like MP 11.
[0716] • Representative genes: NAMPT, BIRC3, CXCL2, LUCAT1, ICAM1, TNFAIP3, CCL2.
[0717] MP 27: Neuroepithelial-like Senescence (NE-like Senescence)
[0718] • Brief description: MP 27 is associated with epithelial-like senescence and secreted factors in a neuroepithelial context.
[0719] • Representative genes: PSCA, ELF3, CLDN7, MUC4, AQP3, MAL2, WFDC2, SLPI, TACSTD2, LCN2, TNFSF10, KRT19, CD24.
[0720] MP 28: Neuroepithelial-like Inflammation 1 (NE-like Inflammation 1) • Brief description: MP 28 is enriched for protein processing and immune-response genes, linking neuroepithelial and glioblastoma-associated inflammatory states.
[0721] • Representative genes: APLP2, LGALS3BP, HLA-A, APP, PDIA6, HLA-C.
[0722] MP 29: Neural Progenitor-like (NPC-like)
[0723] • Brief description: MP 29 corresponds to neural progenitor-like and neuronal differentiation states, with strong overlap to NPC and neuronal programs.
[0724] • Representative genes: DCX, TERF2IP, RBFOX2, NREP, RND3, STMN4, STMN1, DLX5, NSG1, ARL4D, MLLT11, STMN2, ENO2, SOX4, TUBB3, TUBB2A, CD24, UCHL1, MAP1B, RBP1, PAK3, GNG3, ELAVL4.
Claims
1. Claims:
1. A method for developing a digital twin of a patient's tumor using a patient-derived organoid model, for optimizing cancer therapy for said patient, comprising:(a) cultivating an organoid from cancer cells derived from the patient;(b) exposing a first set of said organoids to a plurality of different therapeutic agents at one or more concentrations and treatment durations, and maintaining a second set untreated as controls;(c) at a plurality of time points during and after step (b), collecting multi-modal data from treated and control organoids, the data comprising at least one of: (i) tissue- cleared three-dimensional (3D) imaging data, (ii) RNA sequencing data, and (iii) spatial-omics data mapping gene expression within the organoid structure;(d) integrating the data collected in step (c) to generate a digital twin representing the patient's cancer state and trajectory, the twin informed by the organoid's 3D microstructure, cellular expression programs, and / or spatial organization;(e) using the digital twin to predict one or more combination therapies capable of halting or eliminating the patient's cancer progression based on differential responses observed in the organoid model; and(f) longitudinally updating the digital twin during patient therapy by comparing new patient and / or organoid measurements to predictions of the digital twin to compute one or more divergence metrics, and operating a closed-loop controller configured to adjust the recommended therapy when a divergence metric satisfies a predefined threshold criterion; wherein the method is configured to adjust the recommended therapy when divergence metrics exceed a threshold.
2. The method of claim 1, wherein the cancer is glioblastoma and the digital twin incorporates glioblastoma-specific molecular features.
3. The method of claim 2, wherein the features include metaprogram gene-expression states identified from single-cell RNA-seq of the organoid.
4. The method of any of the preceding claims, wherein the tumor tissue is selected from brain, liver, lung, kidney, intestines, pancreas, bone marrow, skin, heart, adipose tissue, bladder, breast, or colon.
5. The method of claim 4, wherein the organoid is a cerebral organoid and the integrated data includes organoid-specific phenotypic data including neural layer differentiation patterns obtained by immunofluorescence.
6. The method of any of the preceding claims, further comprising copy-number variation (CNV) inference integrated into the twin to represent tumor genomic instability and to guide combination selection.
7. The method of any of the preceding claims, further comprising classifying the patient as a responder or non-responder to a therapy or combination based on the twin's predicted response profile relative to reference profiles.
8. The method of any of the preceding claims, wherein the divergence metric quantifies deviation of new patient and / or organoid data from the twin's predicted trajectory and / or a population atlasbaseline, and the controller adjusts the recommended therapy when the divergence metric satisfies a predefined threshold criterion.
9. The method of any of the preceding claims, further comprising aligning the patient's twin within an integrated patient atlas to identify nearest-neighbor patients or tumor subtypes to inform therapy selection.
10. The method of any of the preceding claims, further comprising quality-assurance (Q.A) filtering that excludes organoids or data failing viability or data-quality thresholds prior to twin generation.
11. The method of any of the preceding claims, wherein a recommended combination is selected if it(i) reduces or stabilizes malignant metaprogram signatures and / or (ii) suppresses therapy-resistant subclones across time points in the organoid.
12. The method of any of the preceding claims, wherein the therapeutic panel comprises at least three agents including at least two distinct mechanism classes, and the recommended regimen comprises at least two agents demonstrating synergistic or complementary effects in the organoid.
13. The method of any of the preceding claims, further comprising validating the adjusted regimen by re-dosing a subsequent batch of organoids and confirming improved efficacy, thereby completing a closed-loop iteration.
14. The method of claims 1, 5 or 8-13, wherein step (c) comprises acquiring at least two modalities selected from: (i) 3D cleared imaging, (ii) RNA-seq, and (iii) spatial-omics, the acquired modalities being integrated to generate the digital twin, and steps (e)-(f) being performed as recited in claim 1.
15. The method of claim 14, wherein the modalities are (i)+(ii); or (ii)+(iii); or (i)+(iii).
16. The method of claim 14 or 15, wherein the twin is further enhanced with CNV or clinical biomarker data to compensate for the absent modality.
17. The method of any of claims 1 or 8, wherein the threshold criterion corresponds to a predefined quantitative change in the divergence metric that is selected based on statistical or clinically relevant parameters derived from organoid and / or patient data.
18. A system comprising:(a) a cultivation module for growing patient-derived organoids and a treatment module for administering therapeutic agents with programmable concentrations and durations;(b) a control module maintaining untreated organoids; (c) a data acquisition subsystem configured to collect, at multiple time points, (i) 3D cleared imaging data, (ii) RNA-seq data, and (iii) spatial-omics data from organoids;(d) a data-integration and analysis module configured to synthesize the digital twin from multi-modal time-series data; and(e) a prediction engine (closed-loop controller) configured to: (i) identify combination therapies predicted to halt or eliminate cancer progression based on organoid responses, and (ii) compare predictions to incoming patient and / or organoid data to update the twin and modify therapy recommendations iteratively.
19. The system of claim 18, for performing the method of any of claims 1-17.
20. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause performance of the method of any of claims 1-17.
21. Use of the method according to any of claims 1-17, the system according to claims 18-19, or the non-transitory computer-readable medium according to claim 20 in any of the following applications or products:(i) a personalized cancer therapy platform;(ii) a drug development and testing platform;(iii) integrated patient care systems;(iv) a glioblastoma-specific clinical decision support tool; (v) a global cancer research database subscription;(vi) predictive analytics for clinical trials; or(vii) personalized oncology services.