Artificial intelligence architecture for gene network analysis and related disease-modeling systems

An AI-driven method using patient-derived brain organoids and gene embeddings addresses the limitations of current models by providing a precise and scalable platform for disease modeling and therapeutic target identification in neurological disorders.

WO2026085448A1PCT designated stage Publication Date: 2026-04-23BRAINSTORM THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BRAINSTORM THERAPEUTICS INC
Filing Date
2025-10-17
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Current methods for discovering therapeutic targets for neurological disorders like Parkinson's Disease face challenges due to the lack of reliable animal models that mimic human brain biology and the complexity of genetic, environmental, and lifestyle factors, leading to limited understanding of disease mechanisms and ineffective treatments.

Method used

An AI-driven approach using patient-derived brain organoids integrated with foundation model-based gene embeddings and single-cell RNA sequencing to construct dynamic gene networks, enabling precise disease modeling and target identification.

Benefits of technology

This method provides a scalable and context-aware platform for understanding disease mechanisms and identifying therapeutic targets with enhanced precision, accelerating the development of effective treatments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025051486_23042026_PF_FP_ABST
    Figure US2025051486_23042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are artificial intelligence architecture for gene network analysis and related organoid-based systems and methods for in vitro and in silico disease modeling and identification of disease-modifying target gene and therapeutic candidates.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: 14876-002-228 ARTIFICIAL INTELLIGENCE ARCHITECTURE FOR GENE NETWORK ANALYSIS AND RELATED DISEASE-MODELING SYSTEMS 1. CROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 709,480, filed on October 20, 2024, and U.S. Provisional Application No.63 / 880,668, filed on September 12, 2025, the disclosure of each of which is incorporated by reference herein in its entirety. 2. FIELD

[0002] The present disclosure is in the field of artificial intelligence (AI) models, analytical methods and systems for discovering therapeutic gene targets and agents for human diseases. 3. BACKGROUND

[0003] The discovery and development of effective disease-modifying therapies for neurological disorders, such as Parkinson’s Disease (PD), have faced significant challenges due to the lack of reliable animal models that can mimic the complexity of human brain biology and accurately predict human efficacy and toxicity. Additionally, understanding of the diverse biological factors underlying brain diseases remains limited. For example, PD is a multifaceted syndrome influenced by a range of genetic, environmental, and lifestyle factors that contribute to the degeneration of dopamine neurons and the progression of the disease. Parkinson’s disease affects approximately 10 million people worldwide, yet despite substantial investment, no disease-modifying therapies are currently available, with existing treatments providing only symptomatic relief.

[0004] Recent advances in patient-derived brain organoid models and AI-driven methodologies have opened new avenues for disease modeling and drug discovery. Innovative approaches are needed for modeling disease progression, identifying drug targets, and discovering biomarkers to facilitate the development of novel therapies. Additionally, prediction of therapeutic gene targets through transcriptomics relies on extensive data sets of transcriptional perturbations. For instance, high-throughput single-cell RNA- sequencing (scRNA-seq) allows profiling of genome-wide expression of thousands of individual cells subjected to such perturbations with single-cell precision. See Chen et al., 2022, “Recent - 1 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 advance in high-throughput single-cell transcriptomics and spatial transcriptomics,” Lab Chip 22, p.4774, the content of which is herein incorporated by reference in its entirety. However, the vast amount of data sets generated by single cell sequencing proposed huge challenges in data analysis. Traditional approaches to molecular interpretation, such as pathway enrichment analysis, often rely on static gene sets that lack cell-type specificity and may obscure subtle but critical disease-relevant signals. Conventional single-cell gene co-expression approaches, such as hdWGCNA, are constrained by the limited number of genes detected per cell and by stringent inclusion thresholds, e.g., requiring genes to be expressed in at least 5% of all cells. As a result, these methods typically capture only 3,000–5,000 genes in the co- expression network, leaving much of the transcriptome unexplored. There exists a need for new methods and systems that offer scalable, context-aware representations that enable the construction of dynamic gene networks directly from experimental data. The present disclosure meets this need. 4. SUMMARY

[0005] Given the above background, what is needed in the art are systems and methods for identification of therapeutic gene targets for drug discovery using a reliable and reproducible translational model. The present disclosure addresses the above-identified shortcomings, at least in part, by using machine learning methods that take advantage of both transcriptional information and functional information about candidate genes to identify potential clinical targets using AI inferred gene network.

[0006] AI foundation model–based gene embeddings offer scalable, context-aware representations that enable the construction of dynamic gene networks directly from experimental data. Trained on datasets comprising 10–100 million single cells, these models generate embeddings that span nearly the entire transcriptome, covering more than 16,000 genes and thus greatly expanding the scope of network-level analysis. The gene network combined with the in vitro organoid system provided an unprecedented approach to establish a preclinical translational model to understand the disease mechanisms, identify therapeutic targets, and evaluate the drug mechanisms of action. Here, provided is a robust and translational midbrain organoid platform integrated with foundation model-based gene network analysis for Parkinson’s disease. By coupling organoid-derived transcriptomics with AI-powered network analysis, a unified framework is established to interrogate disease mechanisms and target discovery with - 2 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 enhanced precision and scalability. This integrative strategy addresses key limitations in current CNS drug discovery workflows and represents a critical step toward building clinically relevant, hypothesis-free platforms for neurodegenerative research.

[0007] The present disclosure makes use of a novel machine learning model that learns transcriptional information to predict genes that impact disease states of interest. The model is trained on a perturbational dataset, consisting of transcriptional information of a plurality of genes across different cell types under the impact of a target disease, to generate a disease- specific gene embedding. The disease-specific gene embedding is obtained by fine-tuning a pre- trained zero-shot deep learning model that has been trained on single-cell transcriptional information of millions of cells to constructs disease-specific gene networks that accurately reflect disease-specific biological processes, including disrupted cellular pathways, gene functions, cell fates, and biomarkers that are associated with the target diseases.

[0008] By integrating patient-derived brain organoids with high-throughput single-cell RNA sequencing (scRNA-seq), the present disclosure enables detailed characterization of cell-type- specific transcriptional responses to genetic and chemical perturbations. These technologies provide a human-relevant platform for modeling disease progression, identifying therapeutic targets, and discovering biomarkers with greater accuracy than traditional animal models. The incorporation of AI-driven analytical methods further enhances the ability to interpret high- dimensional transcriptomic data, accelerating the identification of candidate interventions. Together, these strategies addresses critical gaps in understanding of disease mechanisms and holds promise for enabling more effective, personalized treatments.

[0009] In one aspect, provided herein is a method for generating a disease-specific gene network for a target disease, comprising: (a) providing a training dataset comprising gene transcriptional data collected from a diseased biological sample, wherein the diseased biological sample comprises an cellular sphere comprising cells descending from a cell isolated from a diseased subject suffering from the target disease; (b) feeding the training data set to a foundation deep learning model pre-trained with a sufficient quantity of gene transcriptional data, thereby obtaining a fine-tuned model specific for the target disease; (c) extracting gene embeddings as output from the fine-tuned model; and (d) generating the disease-specific gene - 3 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 network based on the gene embeddings; wherein the gene network comprises at least one cluster of related genes.

[0010] In a related aspect, provided herein is a system for modeling disease-associated mechanism of actions (MOA), comprising: (I) a disease-modeling organoid for a target disease, and (II) a computer system comprising one or more processors and a non-transitory computer readable storage medium including software stored thereon, wherein the software comprises executable instructions that, as a result of execution, causes the one or more processors of the computer system to: (a) load a foundation deep learning model pre-trained with a sufficient quantity of gene transcriptional data; (b) receive a training dataset comprising gene transcriptional data collected from the disease-modeling organoid; (c) process the training dataset using the deep learning model to generate a disease-specific gene network as output; and (d) infer a disease-associated MOA for the target disease, or a disease-modifying solution for the target disease, in either case, based on the generated output in (c).

[0011] In a related aspect, provided herein is a method for modeling a disease-associated mechanism of action (MOA) for a target disease, comprising: (a) providing a disease-modeling organoid for the target disease; (b) collecting gene transcriptional data from the disease-modeling organoid into a training dataset; (c) feeding the training data set to a foundation deep learning model pre-trained with a sufficient quantity of gene transcriptional data, thereby obtaining a fine- tuned model specific for the target disease; (d) extracting gene embeddings as output from the fine-tuned model; and (e) inferring the disease-associated MOA for the target disease, or a disease-modifying solution for the target disease, in either case, based on the generated output in (d).

[0012] In other related aspects, provided herein are also disease-modeling organoid and disease-specific gene network generated by the present method and systems. 5. BRIEF DESCRIPTION OF THE FIGURES

[0013] FIG.1 illustrates an AI-driven workflow of disease modeling and drug discovery using single-cell sequencing (SCS) data integrated with a large language model (LLM)-based AI engine. The workflow begins with input data from preclinical models, such as in vitro cells and organoids, in vivo models, and CRISPR-engineered cells or animal models, as well as human tissue samples from normal donors and patients. The AI engine contains a foundation model - 4 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 trained on large quantity (e.g., over 30 million) of single-cell sequencing data points, which is fine-tuned using the input SCS data to generate gene embeddings. In some embodiments, the gene embeddings can be used to build a gene co-expression network. In some embodiments, the model further undergoes in silico perturbation to simulate gene deletions and downstream effects on the gene network. The outputs of this system include an optimized AI engine tailored for disease-specific applications, validated disease-specific preclinical models, identification of therapeutic targets based on the gene network biology and / or perturbation analysis, and identification of drug candidate based on understanding of drug mechanisms of action through the gene network biology and / or perturbation analysis.

[0014] FIG.2 is a schematic illustration of midbrain organoid generation from human iPSCs.

[0015] FIG.3 illustrates a schematic workflow on constructing a Foundation Model-based gene network. The workflow employs large language models (LLM) trained using single cell sequencing data (such as scGPT and Geneformer) to learn and represent complex biological interactions based on single-cell RNA-seq and bulk RNA-seq data from normal and patient- derived brain organoids. Disease-specific gene clusters were identified using the gene embeddings for dimension reduction, community detection, and cluster annotation matching known biological information. Finally, the network with gene expression data from patient- derived organoids and patient brain tissue were overlaid, enabling disease modeling and target discovery.

[0016] FIG.4 shows an analysis workflow utilizing pre-trained single-cell foundation models, including the scGPT and scFoundation models implemented using NVIDIA® CUDA® and the Geneformer® model implemented using NVIDIA® BioNeMo® and Hugging Face®. The models were fine-tuned using scRNA-seq data from midbrain organoids derived from patients with mutations in GBA1 or LRRK2. Disease-specific gene clusters were identified using the gene embeddings for dimension reduction, community detection, and cluster annotation matching known biological information. Finally, the network with gene expression data from patient-derived organoids and patient brain tissue were overlaid, enabling disease modeling and target discovery.

[0017] FIG.5 shows results from a study demonstrating the patient-derived midbrain organoid models captured the major pathogenic pathways in Parkinson’s Disease (PD) patient’s - 5 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 dopamine neurons. The left graph represents DEGs identified using bulk RNA-Seq for GBA1 PD derived organoid compared with wild-type (WT) control organoids. In bulk RNA-Seq, RNAs from the whole organoids were studied. The right graph represents DEGs identified using single- cell sequencing on dopamine neurons from idiopathic PD brain tissues. In single-cell sequencing, each single cell was sequenced separately. The DEGs were identified by comparing dopamine neuron single cells from PD versus WT control. In each panel of the figure, each node represents a gene, with clusters of closely positioned genes indicating similar expression profiles and shared biological functions, as detected by the fined-tuned model. Red nodes indicate genes up- regulated in PD versus wild-type dopamine neurons, while blue for down-regulated genes.

[0018] FIG.6 shows results from a study demonstrating the patient-derived midbrain organoid models captured the major pathogenic pathways in Parkinson’s Disease (PD) patient’s dopamine neurons and revealed patient genotype- specific biological processes using single-cell sequencing. In each panel of the figure, each node represents a gene, with clusters of closely positioned genes indicating similar expression profiles and shared biological functions, as detected by the fined-tuned model. Red nodes indicate genes up-regulated in PD versus wild- type dopamine neurons, while blue for down-regulated genes.

[0019] FIGS.7A-7E shows gene module identification and benchmarking of the foundation model-derived network. FIG.7A shows gene module assignments visualized on a t-SNE plot. FIG.7B shows Gene Ontology (GO) biological process enrichment analysis for the top five gene modules ranked by best enrichment score. FIG.7C shows functional annotation of gene modules. FIG.7D Reactome pathway enrichment analysis of the top five gene modules. FIG. 7E shows Benchmarking of gene module enrichment performance using GO terms across networks built from pre-trained and fine-tuned foundation models, and non foundation model based method, hdWGCNA. Gene modules associated with neurological functions are highlighted in red. The top five gene modules ranked by best GO enrichment scores for each network are shown using bar plots.

[0020] FIGS.8A-8G show bulk RNA-Seq network analysis of PD GBA1 and wildtype midbrain organoids. FIG.8A shows Pearson correlation of bulk RNA-Seq samples across six experimental batches. FIG.8B shows Principal Variance Component Analysis (PVCA) of factors contributing to the variance of gene expressions in RNA-Seq. FIG.8C shows GO enrichment of differentially expressed genes (DEGs) between PD and control midbrain - 6 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 organoids. Positive Z-scores indicate more up-regulated genes in the term, and negative Z-scores as more down-regulated genes in the term. FIG.8D shows gene network visualization of PD vs Ctrl DEGs, with upregulated genes in red and downregulated genes in blue. FIG.8E shows enrichment of DEGs in gene network modules. FIG.8F shows DEGs in GO:0022008 Neurogenesis term that fall into M7 neurogenesis gene module. FIG.8G shows DEGs in M10 neurogenesis module.

[0021] FIGS.9A-9F show single-nucleus RNA-Seq network analysis of PD GBA1 and wildtype midbrain organoids. FIG.9A shows cell clusters in snRNA-Seq data in UMAP. FIG. 9B shows expression profiles of cell-type markers. FIG.9C shows visualization of marker genes on the gene network, only modules significantly enriched with cell markers are labeled (BH adjusted p<0.05) FIG.9D shows gene module enrichment of cell markers and comparison to bulk RNA-Seq PD vs Ctrl enrichment scores. FIG.9E shows cell-type composition by cell deconvolution in bulk RNA-Seq. FIG.9F shows proportions of dopaminergic neurons and radial glial cells in PD and control organoids.

[0022] FIGS.10A-10E show cross-model gene network analysis of snRNA-Seq data from PD GBA1, PD LRRK2 organoids, and idiopathic PD brain tissue. FIG.10A shows Differentially expressed genes between PD vs Ctrl in dopaminergic neurons in PD GBA1 organoids, PD LRRK2 organoids, and idiopathic PD brain samples. Red indicates upregulation; blue indicates downregulation. FIG.10B shows network of DEGs in dopaminergic neurons from any of the three above comparisons. Dark red (2), up-regulated in two of the three above comparisons, light red (1), up-regulated in one of the three above comparisons, light blue (-1), down-regulated in one of the three above comparisons, dark blue (-2), down-regulated in two of the three above comparisons, gray (0) , opposite direction in two of the three above comparisons. FIGS.10C and 10D show Count of connections by cosine similarity >0.25 and average cosine similarity of the genes in the above categories. FIG.10E shows Spearman correlation of DEGs from the neurogenesis module (m10) in GBA1 organoids with DEGs in dopaminergic neurons from LRRK2 PD and idiopathic PD brain.

[0023] FIG.11A shows module enrichment scores overlapping with Parkinson’s Disease GWAS hits (P<1e-5) from different networks. FIG.11B shows module enrichment scores overlapping with Parkinson’s Disease GWAS hits (P<5e-8) from different networks. FIG.11C shows the enrichment scores of the top five modules by DEGs from bulk RNA-Seq study of PD - 7 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 vs WT organoids. Scgpt_br_noft, zero-shot scGPT brain model; scgpt_all_noft, zero-shot scGPT human model; geneformer_all_noft, zero-shot geneformer model; scfound_all_noft, zero- shot scFoundation model; Scgpt_br_ft, fine-tuned scGPT brain model; scgpt_all_ft, fine-tuned scGPT human model; geneformer_all_ft, fine-tuned geneformer model; scfound_all_ft, fine- tuned scFoundation model.

[0024] FIG.12 shows gene network built with pre-trained versus fine-tuned foundation models. DEGs from PD vs. Ctrl midbrain bulk RNA-Seq are plotted with up-regulation in red and down-regulation in blue.

[0025] FIG.13A shows small worldness calculated by cosine similarity cutoff of 0.3 for gene modules from different networks. FIG.13B shows Gini coefficient by cosine similarity for gene modules from different networks.

[0026] FIGS.14A-14D show gene set enrichment scores from Reactome, GO MF, PanglaoDB and WikiPathways for PD vs Ctrl in bulk RNA-Seq for GBA1 organoid.

[0027] FIGS.15A-15D show DEGs in GO terms plotted on the network. These DEGs in the GO terms fall into different gene modules in the network. The number of genes in each module and Z-scores are shown using the bar plots.

[0028] FIG.16 shows cell markers expression in PD and Ctrl midbrain organoid snRNA- Seq UMAP.

[0029] FIG.17A shows cell trajectory analysis of PD and Ctrl midbrain organoid snRNA- Seq using Monocle 3. FIG.17B shows hierarchical clustering tree of cell clusters.

[0030] FIG.18A shows cell markers from different cell clusters in PD. FIG.18B shows Ctrl midbrain organoid snRNA-Seq plotted on the gene network.

[0031] FIG.19 shows DEGs from different cell clusters between PD vs Ctrl in midbrain organoid snRNA-Seq

[0032] FIG.20 shows Spearman correlation of DEGs from the neurogenesis module (m10) in GBA1 organoids with DEGs in dopaminergic neurons from idiopathic PD brain in m10, m3 and m8, and random genes. 6. DETAILED DESCRIPTION

[0033] Provided herein are artificial intelligence (AI) architecture for gene network analysis and related organoid-based systems and methods for in vitro and in silico disease modeling and identification of disease-modifying target gene and therapeutic candidates. Additional features - 8 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 of the present disclosure will become apparent to those skilled in the art upon consideration of the following detailed description of particular embodiments. 6.1 General Techniques

[0034] The foregoing description, for purpose of explanation, has been described with reference to specific implementations. However, the illustrative discussions below are not intended to be exhaustive or to limit the implementations to the precise forms disclosed. Many modifications and variations are possible in view of the above teachings. The implementations are chosen and described in order to best explain the principles and their practical applications, to thereby enable others skilled in the art to best utilize the implementations and various implementations with various modifications as are suited to the particular use contemplated.

[0035] In the interest of clarity, not all of the routine features of the implementations described herein are shown and described. It will be appreciated that, in the development of any such actual implementation, numerous implementation-specific decisions are made in order to achieve the designer’s specific goals, such as compliance with use case- and business-related constraints, and that these specific goals will vary from one implementation to another and from one designer to another. Moreover, it will be appreciated that such a design effort might be complex and time-consuming, but nevertheless be a routine undertaking of engineering for those of ordering skill in the art having the benefit of the present disclosure.

[0036] Some portions of this description describe the embodiments of the invention in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the data processing arts to convey the substance of their work effectively to others skilled in the art. These operations, while described functionally, computationally, or logically, are understood to be implemented by computer programs or equivalent electrical circuits, microcode, or the like.

[0037] The language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the invention be limited not by this detailed description, but rather by any claims that issue on an application based hereon. Accordingly, the disclosure of the embodiments of the invention is intended to be illustrative, but not limiting, of the scope of the invention. - 9 - NAI-5004565027v1Attorney Docket No.: 14876-002-228

[0038] In general, terms used in the claims and the specification are intended to be construed as having the plain meaning understood by a person of ordinary skill in the art. Certain terms are defined below to provide additional clarity. In case of conflict between the plain meaning and the provided definitions, the provided definitions are to be used.

[0039] Any terms not directly defined herein shall be understood to have the meanings commonly associated with them as understood within the art of the invention. Certain terms are discussed herein to provide additional guidance to the practitioner in describing the compositions, devices, methods and the like of aspects of the invention, and how to make or use them. It will be appreciated that the same thing may be said in more than one way. Consequently, alternative language and synonyms may be used for any one or more of the terms discussed herein. No significance is to be placed upon whether or not a term is elaborated or discussed herein. Some synonyms or substitutable methods, materials and the like are provided. Recital of one or a few synonyms or equivalents does not exclude use of other synonyms or equivalents, unless it is explicitly stated. Use of examples„ including examples of terms, is for illustrative purposes only and does not limit the scope and meaning of the aspects of the invention herein. 6.2 Terminology

[0040] Unless described otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art. For purposes of interpreting this specification, the following description of terms will apply and whenever appropriate, terms used in the singular will also include the plural and vice versa. All patents, applications, published applications, and other publications are incorporated by reference in their entirety. In the event that any description of terms set forth conflicts with any document incorporated herein by reference, the description of term set forth below shall control.

[0041] As used herein, the terms “abundance,” “abundance level,” or “expression level” refers to an amount of a cellular constituent (e.g., a gene product such as an RNA species, e.g., mRNA or miRNA, or a protein molecule) present in one or more cells, or an average amount of a cellular constituent present across multiple cells. When referring to mRNA or protein expression, the term generally refers to the amount of any RNA or protein species corresponding to a particular genomic locus, e.g., a particular gene. However, in some embodiments, an abundance can refer to the amount of a particular isoform of an mRNA or protein corresponding to a - 10 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 particular gene that gives rise to multiple mRNA or protein isoforms. The genomic locus can be identified using a gene name, a chromosomal location, or any other genetic mapping metric.

[0042] As used interchangeably herein, a “cell state” or “biological state” refers to a state or phenotype of a cell or a population of cells. For example, a cell state can be healthy or diseased. A cell state can be one of a plurality of diseases. A cell state can be a response to a compound treatment and / or a differentiated cell lineage. A cell state can be characterized by a measure (e.g., an activation, expression, and / or measure of abundance) of one or more cellular constituents, including but not limited to one or more genes, one or more proteins, and / or one or more biological pathways.

[0043] As used herein, a “cell state transition” or “cellular transition” refers to a transition in a cell’s state from a first cell state to a second cell state. In some embodiments, the second cell state is an altered cell state (e.g., a healthy cell state to a diseased cell state). In some embodiments, one of the respective first cell state and second cell state is an unperturbed state and the other of the respective first cell state and second cell state is a perturbed state caused by an exposure of the cell to a condition. The perturbed state can be caused by exposure of the cell to a compound A cell state transition can be marked by a change in cellular constituent abundance in the cell, and thus by the identity and quantity of cellular constituents (e.g., mRNA, transcription factors) produced by the cell (e.g., a perturbation signature).

[0044] As used herein, the term “dataset” in reference to cellular constituent abundance measurements for a cell or a plurality of cells can refer to a high-dimensional set of data collected from a single cell (e.g., a single-cell cellular constituent abundance dataset) in some contexts. In other contexts, the term “dataset” can refer to a plurality of high-dimensional sets of data collected from single cells (e.g., a plurality of single-cell cellular constituent abundance datasets), each set of data of the plurality collected from one cell of a plurality of cells.

[0045] A “disease-modifying” gene as used herein, refers to a gene, for which an alteration, such as through mutation, regulation, and / or targeted therapeutic intervention, leads to a measurable change in the progression, manifestation, or severity of a disease. In some embodiments, disease modification can result in delaying, halting, or reversing the underlying pathological or pathophysiological processes of the disease, thereby improving clinical signs, symptoms, or patient-reported outcomes. According to the present disclosure, such genes may - 11 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 serve as direct therapeutic targets or be utilized in gene therapy approaches to achieve disease modification, including treatment or cure of the disease.

[0046] Multiple genes can be functionally related to one another in a “disease-modifying biological pathway” where a series of interconnected molecular or cellular events associated with those genes within a biological system leads to a measurable change in the progression, manifestation, or severity of a disease.

[0047] As used herein, the term “differentially expressed gene” or “DEG” refers to a gene that has differential expression as defined herein. The term “upregulation” or a grammatical variant, when used in connection with a gene, refers to the process or event by which the expression level of the gene is increased relative to a reference state, resulting in a higher abundance of the gene’s RNA transcript, and often, consequently, increased production of the encoded nucleic acid or protein. Without being limited to any theory, gene upregulation may be due to factors such as activation of transcription factors, signaling pathways, or epigenetic modifications that enhance gene transcription, including effects related to disease pathology. The term “downregulation” or a grammatical variant, when used in connection with a gene, refers to the process or event by which the expression level of a gene is decreased relative to a reference state, resulting in a reduced abundance of the gene’s RNA transcript, and often decreased production of the gene’s encoded nucleic acid or protein. Without being limited to any theory, gene downregulation may result from inhibitory regulatory mechanisms, such as transcriptional repression, RNA degradation, or epigenetic changes that suppress gene expression, including effects related to disease pathology.

[0048] As used herein, the term “differential abundance” or “differential expression” refers to differences in the quantity and / or the frequency of a cellular constituent present in a first entity (e.g., a first cell, plurality of cells, and / or sample) as compared to a second entity (e.g., a second cell, plurality of cells, and / or sample). In some embodiments, a first entity is a sample characterized by a first cell state (e.g., a diseased phenotype) and a second entity is a sample characterized by a second cell state (e.g., a normal or healthy phenotype). For example, a cellular constituent can be a polynucleotide (e.g., an mRNA transcript) which is present at an elevated level or at a decreased level in entities characterized by a first cell state compared to entities characterized by a second cell state. In some embodiments, a cellular constituent can be a polynucleotide which is detected at a higher frequency or at a lower frequency in entities - 12 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 characterized by a first cell state compared to entities characterized by a second cell state. A cellular constituent can be differentially abundant in terms of quantity, frequency or both. In some instances, a cellular constituent is differentially abundant between two entities if the amount of the cellular constituent in one entity is statistically significantly different from the amount of the cellular constituent in the other entity. For example, a cellular constituent is differentially abundant in two entities if it is present at least about 120%, at least about 130%, at least about 150%, at least about 180%, at least about 200%, at least about 300%, at least about 500%,, at least about 700%,, at least about 900%, or at least about 1000%, or greater in one entity than it is present in the other entity, or if it is detectable in one entity and not detectable in the other. In some instances, a cellular constituent is differentially expressed in two sets of entities if the frequency of detecting the cellular constituent in a first subset of entities (e.g., cells representing a first subset of annotated cell states) is statistically significantly higher or lower than in a second subset of entities (e.g., cells representing a second subset of annotated cell states). For example, a cellular constituent is differentially expressed in two sets of entities if it is detected at least about 120%, at least about 130%, at least about 150%, at least about 180%, at least about 200%, at least about 300%, at least about 500%, at least about 700%, at least about 900%, or at least about 1000% more frequently or less frequently observed in one set of entities than the other set of entities.

[0049] As used herein, the term “healthy” refers to a sample characterized by a healthy state (e.g., obtained from a subject possessing good health). A healthy subject can demonstrate an absence of any malignant or non-malignant disease. A “healthy” individual can have other diseases or conditions, unrelated to the condition being assayed, which can normally not be considered “healthy.”

[0050] As used herein, the term “perturbation” in reference to a cell (e.g., a perturbation of a cell or a cellular perturbation) refers to any exposure of the cell to one or more conditions, such as a treatment by one or more compounds. These compounds can be referred to as “perturbagens.” In some embodiments, the perturbagen can include, e.g., a small molecule, a biologic, a therapeutic, a protein, a protein combined with a small molecule, an ADC, a nucleic acid, such as an siRNA or interfering RNA, a cDNA overexpressing wild-type and / or mutant shRNA, a cDNA over-expressing wild-type and / or mutant guide RNA (e.g., Cas9 system or other gene editing system), or any combination of any of the foregoing. A perturbation can - 13 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 induce or be characterized by a change in the phenotype of the cell and / or a change in the expression or abundance level of one or more cellular constituents in the cell (e.g., a perturbation signature). For instance, a perturbation can be characterized by a change in the transcriptional profile of the cell.

[0051] As used herein, the term “sample,” “biological sample,” or “patient sample” refers to any sample isolated from a subject or derived from a subject, which can reflect a biological state associated with the subject. Examples of samples include, but are not limited to, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of the subject. A sample can include any tissue or material derived from a living or dead subject. A sample can be a cell-free sample. A sample can comprise one or more cellular constituents. For instance, a sample can comprise a nucleic acid (e.g., DNA or RNA) or a fragment thereof, or a protein.

[0052] A biological sample that is “isolated from” a subject refers to any material obtained from a human or non-human organism. The sample is typically collected under controlled conditions and may be used to assess physiological or pathological states, perform molecular profiling, or derive biological models such as cell cultures or organoids.

[0053] A biological sample that is “derived from” a subject, such as a “patient-derived” sample, refers to biological materials that are produced using a sample originally collected from a subject. Examples of such derived biological samples include, but are not limited to, (a) induced pluripotent stem cells (iPSC cells) produced from a cell (e.g., skin cells) collected from a subject, and (b) in vitro cultured cell lines, cell cultures, or cell populations descending from stem cells (including naturally-existing iPSC cells) that are collected from a subject. In some embodiments, a biological sample that is derived from an iPSC cell line include descending cells that have different genotype, phenotype, cell type, and / or cell population’s architectures or functions as compared to the original iPSC cells isolated from the subject, such as (a) mutated cell lines obtained by introducing genetic mutations to the original iPSC cells isolated or derived from a subject, and (b) cell populations, including any 2-dimensional or 3-dimensional cell cultures, spheres, and organoids, that have different constituent cell types, intercellular networks, architectures, and / or functions compared to the original iPSC cells isolated or derived from a subject, including those obtained by inducing differentiation or trans-differentiation of the original iPSC cells. - 14 - NAI-5004565027v1Attorney Docket No.: 14876-002-228

[0054] A “cellular sphere” as used herein refers to a three-dimensional aggregate of cells formed through self-assembly or culture in non-adherent or scaffold-free conditions. In some embodiments, a cellular sphere can arise from a single cell type or multiple cell types and are commonly used as in vitro models to mimic aspects of tissue architecture, cell–cell interactions, and microenvironmental conditions. In some embodiments, cellular spheres are used in applications such as cancer biology (e.g., tumor spheroids), stem cell research (e.g., embryoid bodies or neurospheres). In some embodiments, cellular sphere is used in drug screening due to their ability to better recapitulate physiological conditions compared to two-dimensional cell cultures. In some embodiments, a cellular sphere is a functional organoid.

[0055] A “organoid” as used herein refers to a three-dimensional, spheroid-shaped structure composed of self-organizing cells that mimics key structural and functional aspects of a specific organ. In some embodiments, an organoid is derived from stem cells or primary tissue. In some embodiments, an organoid represents a miniaturized and simplified version of an organ in vitro and is capable of performing one or more of its physiological functions, such as secretion, absorption, or signal transduction. In some embodiments, a cellular sphere may constitute an organoid when it exhibits tissue-specific architecture and function reflective of its organ of origin.

[0056] A “subject-derived organoid” as used herein refers to an organoid that is grown in vitro from cells obtained directly from a biological sample isolated from a donor subject. In some embodiments, the subject is a patient suffering a target disease. In some embodiments, the subject is at an elevated risk of suffering from a target disease. In some embodiments, the subject is suspected to have acquired a target disease. In some embodiments, the subject is a healthy subject. In some embodiments, a subject-derived organoid retains the genetic, histological, and phenotypic features of the donor subject. In some embodiments, a patient-derived organoids mimic key structural and functional characteristics of a diseased tissue or organ isolated from a patient. In some embodiments, a patient-derived organoid is used to model the disease of the donor subject. In some embodiments, a patient-derived organoid is used to identify or test drug molecules to which the donor subject is likely to respond.

[0057] The terms “subject” and “patient” may be used interchangeably. As used herein, in certain embodiments, a subject is a mammal, such as a non-primate (e.g., cow, pig, horse, cat, dog, rat, etc.) or a primate (e.g., monkey and human). In specific embodiments, the subject is a - 15 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 human. In one embodiment, the subject is a mammal (e.g., a human) having a neurological disease. In another embodiment, the subject is a mammal (e.g., a human) at risk of developing a neurological disease.

[0058] As used herein, the term “wild-type” refers to organisms, cells, genes, proteins, oligonucleotides, and the like that are not found in the target disease, and are found in Nature and are unchanged relative to these components found in Nature (native or in the wild).

[0059] As used herein, the term “phenotype” refers to well-known detectable characteristics of the cells referred to herein. The neuronal phenotype can be, but is not limited to, one or more of: neuronal morphology, expression of one or more neuronal markers, electrophysiological characteristics of neurons, synapse formation and release of neurotransmitter. For example, neuronal phenotype encompasses but is not limited to: characteristic morphological aspects of a neuron such as presence of dendrites, an axon and dendritic spines; characteristic neuronal protein expression and distribution, such as presence of synaptic proteins in synaptic puncta, presence of MAP2 in dendrites; and characteristic electrophysiological signs such as spontaneous and evoked synaptic events. Phenotypes that distinguish a neuron from a non- neuron cell (e.g., a glial cell) as well as method for detecting and measuring such phenotypes are known to those of ordinary skill in the art.

[0060] As used herein, the terms “manage,” “managing,” and “management” refer to the beneficial effects that a subject derives from a therapy (e.g., a prophylactic or therapeutic agent), which does not result in a cure of the disease. In certain embodiments, a subject is administered one or more therapies (e.g., prophylactic or therapeutic agents to “manage” a neuronal disorder, one or more symptoms thereof, so as to prevent the progression or worsening of the disease.

[0061] As used herein, the terms “prevent,” “preventing,” and “prevention” refer to reducing the likelihood of the onset (or recurrence) of a disease, disorder, condition, or associated symptom(s) (e.g., Parkinson’s disease).

[0062] As used herein, the terms “treat,” “treating,” and “treatment” refer to (i) the management, (ii) prevention, and (iii) the substantial or complete elimination (e.g., a cure) of the disease, disorder, condition, or associated symptom(s) in question.

[0063] As used interchangeably herein in the context of computer based technology, the term “model,” “algorithm,” “regressor,” and / or “classifier” refers to a machine learning model or algorithm. In some embodiments, a model is an unsupervised learning algorithm. In some - 16 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 embodiments, a model is supervised machine learning. Nonlimiting examples of supervised learning algorithms include, but are not limited to, logistic regression, neural networks, support vector machines, Naive Bayes algorithms, nearest neighbor algorithms, random forest algorithms, decision tree algorithms, boosted trees algorithms, multinomial logistic regression algorithms, linear models, linear regression, GradientBoosting, mixture models„hidden Markov models, Gaussian NB algorithms, linear discriminant analysis, or any combinations thereof.

[0064] The term “foundation model” is a term of art in the field of artificial intelligence, which depending on the context, may refer to a large-scale machine learning model trained on extensive and diverse datasets, enabling it to perform a wide range of tasks with fine-tuning. These models serve as a base upon which various AI applications are built, offering a versatile starting point for developing specialized systems. The term was introduced in a 2021 paper by researchers, defining it as any model that is trained on broad data and adaptable to a wide range of downstream tasks. See Bommasani et al., “On the Opportunities and Risks of Foundation Models” arXiv:2108.07258 (doi.org / 10.48550 / arXiv.2108.07258) page 1.

[0065] As used herein, the term “vector” is an enumerated list of elements, such as an array of elements, where each element has an assigned meaning. As an example, if a vector comprises the abundance counts, in a plurality of cells, for a respective cellular constituent, there exists a predetermined element in the vector for each one of the plurality of cells. For ease of presentation, in some instances a vector may be described as being one-dimensional. However, the present disclosure is not so limited. A vector of any dimension may be used in the present disclosure provided that a description of what each element in the vector represents is defined (e.g., that element 1 represents abundance count of cell 1 of a plurality of cells, etc.).

[0066] The term “embedding” or grammatical variants thereof, when used as a verb in the context of data processing, refers to the process of transforming raw data (e.g., information relating to genes) into a numerical representation, referred to as a vector, that captures the data’s essential meaning and relationships. The term “embedding” or “embeddings” can also be used as a noun to refer to the vectors. In some embodiments, these vectors (embeddings) are lower- dimensional than the original data, making them more efficient for machine learning models to process and analyze. In some embodiments, gene embedding comprises rendering representations of genes as numerical vectors in a multi-dimensional network, capturing meaningful biological properties of a gene and relationships between multiple genes. In some - 17 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 embodiments, gene embeddings contain information such as gene names, gene’s known functions, gene expression status, gene regulation status, temporal and / or spatial (e.g., cell-type specific) patterns of gene expression or regulation, functional interactions or similarities between multiple genes. In some embodiments, gene embeddings are generated from data sets containing single-cell RNA sequencing (scRNA-Seq) data, single-cell DNA sequencing (scDNA-Seq) data, single-nucleus RNA sequencing (snRNA-Seq) data, bulk RNA-Seq data, transcriptomic data, epigenomic data, and / or proteomic data.

[0067] In some embodiments, gene embeddings undergo a dimensionality reduction process, in which data with many features (high dimensionality) are transformed into a dataset with fewer features (lower dimensionality) while preserving the most important characteristics of the original data. In some embodiments, dimension reduction can simplify the data, reducing computational costs, and improve performance of the AI model.

[0068] As used herein, and unless otherwise indicated, the term “about” or “approximately” means an acceptable error for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined. In certain embodiments, the term “about” or “approximately” means within 1, 2, 3, or 4 standard deviations. In certain embodiments, the term “about” or “approximately” means within 20%, 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.05%, or less of a given value or range. As used herein, when “about” is used in connection with a numerical range, the term “about” is meant to apply to both ends of such modified range (e.g., “about 5 to 10” means “about 5 to about 10”).

[0069] The singular terms “a,” “an,” and “the” as used herein include the plural reference unless the context clearly indicates otherwise.

[0070] All publications, patent applications, accession numbers, and other references cited in this specification are herein incorporated by reference in their entirety as if each individual publication or patent application were specifically and individually indicated to be incorporated by reference. The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided can be different from the actual publication dates which can need to be independently confirmed. - 18 - NAI-5004565027v1Attorney Docket No.: 14876-002-228

[0071] A number of embodiments of the invention have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the invention. Accordingly, the descriptions in the Experimental section and examples are intended to illustrate but not limit the scope of invention described in the claims. 6.3 Methods and Systems

[0072] According to the present disclosure, provided herein are artificial intelligence (AI) architecture for gene network analysis and related disease-modeling systems.

[0073] In one aspect of the present disclosure, provided herein is a method for utilizing pre- trained foundation models, fine-tuned with organoid-specific data, to construct a disease-specific gene interaction network. This method employs models such as scGPT and Geneformer to learn and represent complex biological interactions based on single-cell RNA-seq and bulk RNA-seq data from normal and patient-derived brain organoids (FIG.4). This approach enhances the accuracy of disease models by capturing pathogenic pathways, such as dopamine neuron loss in Parkinson’s disease, which are typically overlooked by traditional enrichment methods.

[0074] In one aspect of the present disclosure, provided herein is a human brain organoid- based network for neurological disease analysis. In some embodiments, provided herein is a system for constructing disease-specific biological networks from human brain organoid data using fine-tuned AI foundation models. The network is optimized for the study of neurological disorders, enabling the identification of key gene clusters, critical pathways, and fundamental biological processes such as neurogenesis, axon guidance, and neurotransmitter secretion. This approach offers unprecedented insights into disease mechanisms at both the cellular and molecular levels, significantly advancing disease modeling for complex conditions like Parkinson’s disease. Moreover, the network-based analysis provides a novel capability to distinguish and quantify the distinct impacts of various disease genotypes on patient-specific phenotypes, enabling more precise stratification and understanding of neurological disorders.

[0075] In one aspect of the present disclosure, provided herein are applications of the present systems and methods for drug discovery and in silico perturbation. In some embodiments, provided herein is a system that leverages the AI-generated network to identify novel drug targets by analyzing gene clusters and their roles in disease progression. The network supports in silico perturbation screening by simulating genetic or pharmacological interventions, enabling - 19 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 the prediction of drug efficacy and the identification of potential therapeutic compounds. This method provides a cost-effective and scalable approach for high-throughput drug screening.

[0076] In one aspect of the present disclosure, provided herein are applications of the present systems and methods for Mechanism of Action (MOA) studies and biomarker discovery for a target disease. In some embodiments, provided herein is a method for utilizing the AI-generated network to investigate the mechanisms of action (MOA) of identified drug candidates. By analyzing the impact of drugs on this network, researchers can gain deeper insights into drug efficacy, off-target effects, and potential side effects. The network's ability to map interactions between multiple targets enhances the understanding of combination therapies and drug synergies. Furthermore, the network facilitates biomarker discovery by tracking shifts in gene expression and network dynamics in response to perturbations, enabling more precise patient stratification and the development of diagnostic tools.

[0077] According to the present disclosure, in one aspect, provided herein is a method for generating a disease-specific gene network for a target disease. In some embodiments, the gene network is a systematic representation of interconnected genes that interact with each other through regulatory relationships, including transcriptional, translational, and signaling pathways, to coordinate and control biological processes within a cell or organism. These interactions collectively influence gene expression patters, cellular functions, and phenotypic outcomes.

[0078] In some embodiments, the gene network is generated by or though the use of an artificial intelligence (AI) model. In some embodiments, the AI model is pre-trained with sufficient quantity of gene transcriptional data for a target organism or species. In some embodiments, the gene transcriptional data have a resolution at single cells. In some embodiments, the target organism or species is human. In some embodiments, the target organism or species is a non-human primate. In some embodiments, the target organism or species is a non-human mammal.

[0079] In some embodiments, the AI model is pretrained with gene transcriptional data of more than 1 million single cells. In some embodiments, the AI model is pretrained with gene transcriptional data of more than 5 million single cells. In some embodiments, the AI model is pretrained with gene transcriptional data of more than 10 million single cells. In some embodiments, the AI model is pretrained with gene transcriptional data of more than 20 million single cells. In some embodiments, the AI model is pretrained with gene transcriptional data of - 20 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 more than 30 million single cells. In some embodiments, the AI model is pretrained with gene transcriptional data of more than 50 million single cells.

[0080] In some embodiments, the AI model is a multinomial classifier algorithm. In some embodiments, the AI model is a transformer model. In some embodiments, the AI model is a bidirectional transformer (BERT) model. In some embodiments, the AI model is a deep neural network (e.g., a deep-and-wide sample-level model). In some embodiments, a classifier or model of the present disclosure has 25 or more, 100 or more, 1000 or more 10,000 or more, 100,000 or more or 106or more parameters and thus the calculations of the model cannot be mentally performed.

[0081] As used herein, the term “parameter” refers to any coefficient or, similarly, any value of an internal or external element (e.g., a weight and / or a hyperparameter) in an algorithm, model, regressor, and / or classifier that can affect (e,g, modify, tailor, and / or adjust) one or more inputs, outputs, and / or functions in the algorithm, model, regressor and / or classitier. For example, in some embodiments, a parameter refers to any coefficient, weight, and / or hyperparameter that can be used to control, modify, tailor, and / or adjust the behavior, learning, and / or performance of an algorithm, model, regressor, and / or classifier. In some instances, a parameter is used to increase or decrease the influence of an input (e.g., a feature) to an algorithm, model, regressor, and / or classifier. As a nonlimiting example, in some embodiments, a parameter is used to increase or decrease the influence of a node (e.g, of a neural network), where the node includes one or more activation functions. Assignment of parameters to specific inputs, outputs, and / or functions is not limited to any one paradigm for a given algorithm, model, regressor, and / or classifier but can be used in any suitable algorithm, model, regressor, and / or classifier architecture for a desired performance. In some embodiments, a parameter has a fixed value. In some embodiments, a value of a parameter is manually and / or automatically adjustable. In some embodiments, a value of a parameter is modified by a validation and / or training process for an algorithm, model, regressor, and / or classifier (e.g., by error minimization and / or backpropagation methods). In some embodiments, an algorithm, model, regressor, and / or classifier of the present disclosure includes a plurality of parameters In some embodiments, the plurality of parameters is n parameters, where n≥2, n≥5, n≥10, n≥25, n≥40, n≥50, n≥75, n≥100, n≥125, n≥150, n≥200, n≥225, n≥250, n≥350, n≥500, n≥600, n≥750, n≥1,000, n≥2,000, n≥4,000, n≥5,000, n≥7,500, n≥10,000, n≥20,000, n≥40,000, n≥75,000, n≥100,000, n≥200,000, n≥500,000, - 21 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 n≥1 × 106, n≥5 × 106, or n≥1 × 107. In some embodiments, n is between 10,000 and 1 x 107, between 100,000 and 5 x 106, or between 500,000 and 1 x 106. As such, the algorithms, models, regressors, and / or classifiers of the present disclosure cannot be mentally performed. In some embodiments, the algorithms, models, regressors, and / or classifier of the present disclosure operate in a k-dimensional space, where k is a positive integer of 5 or greater (e.g., 5, 6, 7, 8, 9, 10, etc.). As such, the algorithms, models, regressors, and / or classifiers of the present disclosure cannot be mentally performed.

[0082] In some embodiments, the AI model is a transformer-based model (e.g., a Transformer encoder, decoder, or encoder-decoder architecture, such as a Bidirectional Encoder Representations from Transformers (BERT) model). Transformer models, also known as attention-based deep learning models, include encoder-only, decoder-only, and encoder-decoder architectures. Transformer models can be machine learning models that may be trained to map an input data set to an output data set, where the model comprises a sequence of interconnected layers utilizing self-attention and feed-forward neural networks. For example, in some embodiments, the transformer architecture may comprise at least an input embedding layer, one or more transformer blocks (each block including multi-head self-attention and feed-forward sublayers), and an output layer. The transformer may comprise any total number of layers, and any number of attention heads, where the self-attention mechanism functions as a trainable feature extractor that allows mapping of a set of input tokens to an output value or set of output values.

[0083] In some embodiments, the AI model is a neural network (e.g., a convolutional neural network and / or a residual neural network). Neural network models, also known as artificial neural networks (ANNs), include convolutional and / or residual neural network models (deep learning models). Neural networks can be machine learning models that may be trained to map an input data set to an output data set, where the neural network comprises an interconnected group of nodes organized into multiple layers of nodes. For example, in some embodiments, the neural network architecture may comprise at least an input layer, one or more hidden layers, and an output layer. The neural network may comprise any total number of layers, and any number of hidden layers, where the hidden layers function as trainable feature extractors that allow mapping of a set of input data to an output value or set of output values. - 22 - NAI-5004565027v1Attorney Docket No.: 14876-002-228

[0084] In some embodiments, the AI model is a deep learning model. In some embodiments, the deep learning model can be a neural network comprising a plurality of hidden layers, e.g., two or more hidden layers. In some embodiments, each layer of the neural network can comprise a number of nodes (or “neurons”). In some embodiments, a node can receive input that comes either directly from the input data or the output of nodes in previous layers, and perform a specific operation, e.g., a summation operation. In some embodiments, a connection from an input to a node is associated with a parameter (e.g., a weight and / or weighting factor). In some embodiments, the node may sum up the products of all pairs of inputs, xi, and their associated parameters. In some embodiments, the weighted sum is offset with a bias, b. In some embodiments, the output of a node or neuron may be gated using a threshold or activation function, f, which may be a linear or non-linear function. In some embodiments, the activation function may be, for example, a rectified linear unit (ReLU) activation function, a Leaky ReLU activation function, or other function such as a saturating hyperbolic tangent, identity, binary step, logistic, arcTan, softsign, parametric rectified linear unit, exponential linear unit, softPlus, bent identity, softExponential, Sinusoid, Sine, Gaussian, or sigmoid function, or any combination thereof.

[0085] In some embodiments, the weighting factors, bias values, and threshold values, or other computational parameters of the neural network, may be “taught” or “learned” in a training phase using one or more sets of training data. For example, in some embodiments, the parameters may be trained using the input data from a training dataset and a gradient descent or backward propagation method so that the output value(s) that the ANN computes are consistent with the examples included in the training dataset. The parameters may be obtained from a back propagation neural network training process. In some embodiments, the machine learning makes use of a pre-trained and / or transfer-learned ANN or deep learning architecture.

[0086] As used interchangeably herein, the term “neuron,” “node,” “unit,” “hidden neuron,” “hidden unit,” or the like, refers to a unit of a neural network that accepts input and provides an output via an activation function and one or more parameters (e.g., coefficients and / or weights). For example, in some embodiments, a hidden neuron can accept one or more inputs from a prior layer and provide an output that serves as an input for a subsequent layer. In some embodiments, a neural network comprises only one output neuron. In some embodiments, a neural network comprises a plurality of output neurons. In some embodiments, the output is a prediction value, - 23 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 such as a probability or likelihood, a binary determination (e.g., a presence or absence, a positive or negative result), and / or a label (e.g., a classification and / or a correlation coefficient) of a condition of interest such as a covariate, a cell state annotation, or a cellular process of interest. In some embodiments, for single-class classification models, the output can be a likelihood (e.g., a correlation coefficient and / or a weight) of an input feature (e.g., one or more cellular constituent modules) having a condition (e.g., a covariate, a cell state annotation, and / or a cellular process of interest). For multi-class classification models, multiple prediction values can be generated, with each prediction value indicating the likelihood of an input feature for each condition of interest.

[0087] In some embodiments, the AI model is a foundation model. The term “foundation model” is a term of art in the field of artificial intelligence and was introduced in a 2021 paper by researchers, defining it as any model that is trained on broad data and adaptable to a wide range of downstream tasks. See Bommasani et al., “On the Opportunities and Risks of Foundation Models” arXiv:2108.07258 (doi.org / 10.48550 / arXiv.2108.07258) page 1. In some embodiments of the present disclosure, the foundation model is pre-trained on extensive and diverse datasets, enabling the model to perform a wide range of tasks with fine-tuning.

[0088] In some embodiments, the foundation model is pre-trained with a sufficiently large quantity of gene transcriptional data. In some embodiments, the foundation model is a pre- trained deep learning model. In some embodiments, the foundation model is a pre-trained large language model. In some embodiments, the foundation model is a pre-trained generative pre- trained transformer (GPT).

[0089] In some embodiments, the foundational model is pre-trained with gene transcriptional data having the single-cell resolution collected from more than 1 million cells, more than 5 million cells, more than 10 million cells, more than 20 million cells, more than 30 million cells or more than 50 million cells. In some embodiments, the more than 1 million cells, more than 5 million cells, more than 10 million cells, more than 20 million cells, more than 30 million cells or more than 50 million cells have diversified cell types and / or cell states.

[0090] In some embodiments, the gene transcriptional data used to pre-train the foundation model comprises various types of data. In some embodiments, the gene transcriptional data comprises single-cell RNA sequencing (scRNA-seq) data. In some embodiments, the scRNA- seq data capture the transcriptome of individual cells to reveal cell-specific gene expression - 24 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 profiles. In some embodiments, the scRNA- seq data offer full-length transcript coverage. In some embodiments, the gene transcriptional data comprises single-nucleus RNA sequencing (snRNA-seq) data that profile RNA from isolated nuclei. In other embodiments, the gene transcriptional data comprises total RNA sequencing data at a single-cell resolution. In some embodiments, the gene transcriptional data capture both coding and non-coding RNAs for a more comprehensive view of the transcriptome.

[0091] In some embodiments, the gene transcriptional data can be provided by methods and systems known in the art. For example, in some embodiments, the gene transcriptional data are provided by the Smart-seq or Smart-seq2 method or system. In other embodiments, the gene transcriptional data are provided using high-throughput methods that utilize droplet microfluidics to analyze gene expression at the single-cell level. In some embodiments, the gene transcriptional data are provided using the 10x Genomics Chromium method or system. In some embodiments, the gene transcriptional data are provided using the Drop-seq method or system. In some embodiments, the gene transcriptional data are provided using or the inDrop method or system. In some embodiments, the gene transcriptional data are provided using the Seq-Well method or system.

[0092] In some embodiments, the gene transcriptional data comprise spatial transcriptomic data. In some embodiments, the gene transcriptional data retain spatial context of gene expression in tissue sections. In some embodiments, the gene transcriptional data are provided by the 10x Genomics Visium method or system. In some embodiments, the gene transcriptional data are provided by the MERFISH method or system. In some embodiments, the gene transcriptional data are provided by the seqFISH method or system. In some embodiments, the gene transcriptional data are provided by the Slide-seq method or system.

[0093] In some embodiments, the gene transcriptional data comprises data derived from CRISPR-based perturbation assays that combine genetic perturbations with single-cell transcriptomics to study gene function and regulatory interactions. In some embodiments, the gene transcriptional data is provided by the Perturb-seq method or system. In some embodiments, the gene transcriptional data includes multimodal measurements such as data provided by the CITE-seq or REAP-seq method or system, where both mRNA and surface protein expression are simultaneously profiled in individual cells using antibody-derived tags. In other embodiments, the gene transcriptional data is analyzed in conjunction with chromatin - 25 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 accessibility data, such as using the scATAC-seq method or system, to link transcriptional activity with epigenomic regulation at the single-cell level.

[0094] In specific embodiments, the pre-trained foundation model is scGPT (Cui H, Wang C, Maan H, Pang K, Luo F, Duan N, et al. scGPT: toward building a foundation model for single- cell multi-omics using generative AI. Nat Methods.2024. Epub 2024 / 02 / 27. doi: 10.1038 / s41592-024-02201-0. PubMed PMID: 38409223; the content of which is incorporated herein by reference in its entirety). In specific embodiments, the pre-trained foundation model is Geneformer (Theodoris CV, Xiao L, Chopra A, Chaffin MD, Al Sayed ZR, Hill MC, et al. Transfer learning enables predictions in network biology. Nature.2023;618(7965):616-24. Epub 2023 / 06 / 01. doi: 10.1038 / s41586-023-06139-9. PubMed PMID: 37258680; PubMed Central PMCID: PMCPMC10949956; the content of which is incorporated herein by reference in its entirety). In specific embodiments, the pre-trained foundation model is scFoundation (Hao M, Gong J, Zeng X, Liu C, Guo Y, Cheng X, et al. Large-scale foundation model on single-cell transcriptomics. Nat Methods.2024;21(8):1481-91. Epub 2024 / 06 / 07. doi: 10.1038 / s41592-024- 02305-7. PubMed PMID: 38844628; the content of which is incorporated herein by reference in its entirety).

[0095] In some embodiments, the foundation model is implemented using a computing platform. In some embodiments, the computing platform comprises one or more selected from NVIDIA® CUDA®, NVIDIA® BioNeMo®, and Hugging Face®. In specific embodiments, the foundation model is scGPT implemented using NVIDIA® CUDA®. In specific embodiments, the foundation model is scFoundation implemented using NVIDIA® CUDA®, Geneformer implemented using NVIDIA® BioNeMo®. In some embodiment, the foundation model is Geneformer implemented using PyTorch from Hugging Face®.

[0096] In some embodiments, the present method involves providing a training dataset for fine-tuning a pre-trained foundation model for specific tasks. In some embodiments, the training dataset comprises a gene transcriptional data collected from a biological sample. In some embodiments, the biological sample comprises biological materials isolated from a subject. In some embodiments, the biological sample comprises an in vitro culture of biological materials isolated from a subject. In some embodiments, the isolated biological materials comprises a cell, a tissue, an / or an organ. In some embodiments, the biological sample comprises an in vitro culture of biological materials derived from a subject. In specific embodiments, the biological - 26 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 sample comprises an in vitro culture of cells descending from an ancestor cell isolated from a subject. In some embodiments, cells descending from an ancestor cell is the product of cell proliferation and expansion. In some embodiments, cells descending from an ancestor cell can carry one or more mutations that is not present in the ancestor cell.

[0097] In some embodiments, the biological sample comprises an in vitro culture of cells derived from an ancestor cell isolated from a subject and the derived cells are different from the ancestor cell in one or more features selected from genotype, phenotype, cell state, cell type, cell function, cell fate, architectures of cell populations, and function of cell populations. In some embodiments, the ancestor cell is a stem cell, and cells derived from the ancestor cell is the product of differentiation of the stem cell. In specific embodiments, the ancestor cell is a stem cell isolated from a subject. In some embodiments, the ancestor cell is an embryonic stem cells. In some embodiments, the ancestor cell is an adult stem cell. In some embodiments, the ancestor cell is an undifferentiated or partially differentiated progenitor cell. In specific embodiments, the ancestor cell is a stem cell that is produced from a differentiated or partially differentiated cell isolated from a subject. In specific embodiments, the ancestor cell is an induced pluripotent stem cell (iPSC) derived from a differentiated or partially differentiated cell isolated from a subject. In specific embodiments, the ancestor cell is an iPSC derived from a skin tissue isolated from a subject.

[0098] In some embodiments, the subject providing the biological sample or the biological materials from which a biological sample is derived suffers from a target disease, is suspected of suffering from a target disease, or is at an elevated risk of developing a target disease. In some embodiments, the subject providing the biological sample or the biological materials from which a biological sample is derived is a subject who does not have a target disease, or is not suspected of having a target disease, or is not at an elevated risk of having a target disease. In some embodiments, the subject providing the biological sample or the biological materials from which a biological sample is derived is a healthy subject.

[0099] In some embodiments, the target disease is a neurological disease, and the disease- modeling organoid is a midbrain organoid. In some embodiments, the target disease is a neurological disease, and the disease-modeling organoid is a cerebral organoid. In some embodiments, the disease-modeling organoid carries at least one gene mutation associated with the target neurological disease. In some embodiments, the target disease is selected from - 27 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 Alzheimer’s Disease, Rett Syndrome, Parkinson’s Disease, Huntington’s Disease, Epilepsy, Addiction, Neuropsychiatry, Rett Syndrome, and CDKL5 Deficiency Disorder.

[0100] In specific embodiments, the target disease is Parkinson’s Disease, and the disease- modeling organoid carries one or more gene mutations associated with the Parkinson’s Disease. In specific embodiments, the target disease is Parkinson’s Disease, the disease-modifying organoid is a patient-derived midbrain organoid derived from a Parkinson’s Disease patient. In specific embodiments, the target disease is Parkinson’s Disease, and the disease-modifying organoid carries a mutation in the GBA1 (Glucocerebrosidase 1) gene. In specific embodiments, the target disease is Parkinson’s Disease, and the disease-modifying organoid carries a plurality of gene mutations associated with Parkinson’s Disease, wherein the gene mutations comprises a mutation in the GBA1 gene. In specific embodiments, the midbrain organoid is homozygote for the GBA1 gene mutation. In specific embodiments, the midbrain organoid is heterozygote for the GBA1 gene mutation. In specific embodiments, the midbrain organoid comprises one or more GBA1 gene mutations selected from N370S, L444P, D409H and the RecNcil mutation. In specific embodiments, the midbrain organoid comprises GBA1 gene mutations N370S, L444P, D409H and the RecNcil mutation. In specific embodiments, the midbrain organoid has a genotype of GBA1 N370S / N370S, GBA1 L444P / L444P, or GBA1 N370S / L444P, wherein the mutation comprises a missense substitution in exon 9 (e.g., c.1226A>G (N370S)) or exon 10 (e.g., c.1448T>C (L444P). Additional pathogenic GBA1 gene mutations implicated with Parkinson’s Disease can be found in Zhou, Y., Wang, Y., Wan, J. et al. Mutational spectrum and clinical features of GBA1 variants in a Chinese cohort with Parkinson’s disease. npj Parkinsons Dis.9, 129 (2023). doi.org / 10.1038 / s41531-023-00571-4.

[0101] In specific embodiments, the target disease is Parkinson’s Disease, and the disease- modifying organoid carries a mutation in the LRRK2 (Leucine-rich repeat kinase 2) gene. In specific embodiments, the target disease is Parkinson’s Disease, and the disease-modifying organoid carries a plurality of gene mutations associated with Parkinson’s Disease, wherein the gene mutations comprises a mutation in the LRRK2 gene. In specific embodiments, the disease- modifying organoid is homozygote for the LRRK2 gene mutation. In specific embodiments, the disease-modifying organoid is heterozygote for the LRRK2 gene mutation. In specific embodiments, the midbrain organoid comprises one or more LRRK2 gene mutations selected from G2019S, R1441C, R1441G, R1441H, R1441S, I12020T, and Y1699C. In specific - 28 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 embodiments, the midbrain organoid comprises LRRK2 gene mutations G2019S, R1441C / G / H / S, I12020T, and Y1699C. In specific embodiments, the disease-modifying organoid has a genotype of LRRK2 G2019S / G2019S or LRRK2 G2019S / wild-type, wherein the mutation comprises a missense substitution c.6055G>A (p.Gly2019Ser) located in the kinase domain of the LRRK2 protein.

[0102] In specific embodiments, the target disease is Parkinson’s Disease, and the disease- modifying organoid carries one or more mutations in genes selected from GBA1, LRRK2, SNCA (α-synuclein), PINK1 (PTEN-induced kinase 1), PARK2 (Parkin), and DJ-1. Additional

[0103] In specific embodiments, the target disease is Alzheimer’s Disease, and the disease- modifying organoid carries one or more mutations in genes selected from APP (Amyloid precursor protein), PSEN1 (Presenilin 1), PSEN2 (Presenilin 2), and APOE ε4.

[0104] In specific embodiments, the target disease is Amyotrophic Lateral Sclerosis (ALS), and the disease-modifying organoid carries one or more mutations in genes selected from SOD1 (Superoxide dismutase 1), C9orf72 (hexanucleotide repeat expansion), TARDBP (TDP-43), FUS (Fused in sarcoma), ANG, and OPTN.

[0105] In specific embodiments, the target disease is Huntington’s Disease, and the disease- modifying organoid carries one or more mutations in the HTT (Huntingtin) gene.

[0106] In specific embodiments, the target disease is Frontotemporal Dementia (FTD), and the disease-modifying organoid carries one or more mutations in the genes selected from MAPT (Microtubule-associated protein tau), GRN (Progranulin), C9orf72, VCP, and CHMP2B.

[0107] In specific embodiments, the target disease is Spinal Muscular Atrophy (SMA), and the disease-modifying organoid carries one or more mutations in the genes selected from SMN1 (Survival motor neuron 1) and SMN2.

[0108] In specific embodiments, the target disease is Charcot-Marie-Tooth Disease (CMT), and the disease-modifying organoid carries one or more mutations in the genes selected from PMP22 (Peripheral myelin protein 22), GJB1, MPZ, MFN2, and EGR2.

[0109] In specific embodiments, the target disease is Epilepsy, and the disease-modifying organoid carries one or more mutations in the genes selected from SCN1A, SCN2A, KCNQ2, DEPDC5, GABRA1, and CDKL5.

[0110] In specific embodiments, the target disease is Rett Syndrome, and the disease- modifying organoid carries a mutation in the MECP2 (Methyl CpG binding protein 2) gene. - 29 - NAI-5004565027v1Attorney Docket No.: 14876-002-228

[0111] In specific embodiments, the target disease is Fragile X Syndrome, and the disease- modifying organoid carries a mutation in the FMR1 gene.

[0112] In specific embodiments, the target disease is ataxia, and the disease-modifying organoid carries one or more mutations in the genes selected from ATXN1, ATXN2, ATXN3, and FXN.

[0113] In specific embodiments, the target disease is Multiple Sclerosis (MS), and the disease-modifying organoid carries one or more mutations in genes selected from HLA-DRB1, IL2RA, IL7R, and TNFRSF1A.

[0114] In specific embodiments, the target disease is Dystonia, and the disease-modifying organoid carries one or more mutations in the genes selected from TOR1A, THAP1, GNAL, and ANO3.

[0115] In specific embodiments, the target disease is a Leukodystrophy, and the disease- modifying organoid carries one or more mutations in the genes selected from ABCD1, ARSA, GALC, and EIF2B1–EIF2B5.

[0116] In specific embodiments, the target disease is Neurofibromatosis, and the disease- modifying organoid carries one or more mutations in the genes selected from NF1, NF2, SMARCB1, and LZTR1.

[0117] In specific embodiments, the target disease is Tay-Sachs Disease, and the disease- modifying organoid carries a mutation in the HEXA gene, which encodes beta-hexosaminidase A.

[0118] In specific embodiments, the target disease is Canavan Disease, and the disease- modifying organoid carries a mutation in the ASPA gene, which encodes the enzyme aspartoacylase.

[0119] In specific embodiments, the target disease is Pelizaeus-Merzbacher Disease, and the disease-modifying organoid carries a mutation in the PLP1 gene, encoding proteolipid protein 1 essential for myelin formation.

[0120] In specific embodiments, the target disease is Alexander Disease, and the disease- modifying organoid carries a mutation in the GFAP gene, encoding glial fibrillary acidic protein.

[0121] In specific embodiments, the target disease is Wilson’s Disease, and the disease- modifying organoid carries a mutation in the ATP7B gene, which encodes a copper-transporting ATPase involved in hepatic copper regulation. - 30 - NAI-5004565027v1Attorney Docket No.: 14876-002-228

[0122] In some embodiments, the in vitro culture comprises a two-dimensional (2D) monolayer culture of cells. In some embodiments, cells are grown as a monolayer on flat surfaces, such as tissue culture-treated plastic substrates (e.g., polystyrene) or coated glass. In some embodiments, the 2D culture system allows for straightforward observation, manipulation, and high-throughput analysis of cellular responses.

[0123] Methods and systems for generating 2D cell cultures are well known in the art. In some embodiments, cultured cells are plating onto a 2D surface at defined densities in a culture media. Commonly used surface coatings include extracellular matrix proteins such as collagen, fibronectin, or laminin, which enhance cell attachment and proliferation. Culture media are supplemented with essential nutrients, growth factors, and serum or defined additives specific to the cell type. In some embodiments, cells are plating on a 2D surface at defined densities.

[0124] While 2D cultures offer simplicity and reproducibility, they can be limited, under certain scenarios, in their ability to mimic the structural and functional complexity of native tissues due to the lack of three-dimensional cell-cell and cell-matrix interactions. Accordingly, in some embodiments, the in vitro culture comprises a three-dimensional (3D) culture system, where constituent cells are allowed to mimic the structural and functional interactions existing in native tissues. In some embodiments, a 3D culture system supports the spatial organization of cells and their interaction with extracellular matrix components.

[0125] In some embodiments, the in vitro 3D culture comprises cellular spheres comprising multicellular aggregates of cells that assemble in suspension or within scaffold-free or scaffold- based environments. In some embodiments, in vitro cultured cellular spheres recapitulate key features of tissue architecture, such as but are not limited to nutrient and oxygen gradients, temporal and special differential proliferation or differentiation, and cell-cell communication.

[0126] In some embodiments, the in vitro 3D culture comprises a functional organoid. In some embodiments, the organoid comprises a self-organizing populations of cells derived from stem cells (e.g., embryonic stem cells, induced pluripotent stem cells, or adult stem / progenitor cells) that recapitulate the cellular heterogeneity, structural organization, and some functional aspects of a native organ. In various embodiments, organoids can be derived from various tissues, including but not limited to intestine, brain, liver, kidney, heart and lung. In various embodiments, organoids can mimic various organs, including but not limited to intestine, brain, liver, kidney, heart, and lung. In some embodiments, organoids are cultured in specialized - 31 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 matrices (e.g., Matrigel or synthetic hydrogels) that provide necessary biochemical and biomechanical cues.

[0127] Several well-established methods can be used to generate cellular spheres. One approach is the hanging drop method, where droplets containing suspended cells are placed on the inverted lid of a culture dish, allowing gravity to facilitate cell aggregation at the lowest point of each droplet. Another technique involves the use of low-attachment plates, in which cells are seeded into ultra-low adhesion multi-well plates that inhibit cell attachment to the surface, thereby promoting spontaneous aggregation and spheroid formation in suspension. Additionally, rotating bioreactors or spinner flasks can be employed to create dynamic culture conditions, which improve spheroid uniformity and enhance the diffusion of oxygen and nutrients throughout the spheroid structure.

[0128] In some embodiments, cellular sphere or organoids can be derived from stem cells (e.g., embryonic stem cells (ESCs), induced pluripotent stem cells (iPSCs), or adult tissue- resident stem or progenitor cells). In some embodiments, the stem cells or progeny cells are subjected to stage-specific differentiation cues in a stepwise manner to mimic embryonic development and guide lineage specification. In some embodiments, following induction, the cells are embedded in a biomimetic extracellular matrix (e.g., such as Matrigel, collagen, or synthetic hydrogels), which supports 3D tissue organization and morphogenesis. In some embodiments, the embedded cells are cultured in organ-specific media under either static or dynamic conditions to promote maturation into organized, tissue-like structures. In some embodiments, the cellular sphere or organoids are isogenic having the same genotype as the ancestor stem cell.

[0129] In some embodiments, the in vitro 3D culture comprises a functional organoid mimicking an organ affected by a target disease. For example, in some embodiments, the target disease is a neurological disorder, and the disease-modeling organoid is a midbrain organoid. In some embodiments, the target disease is a neurological disorder, and the disease modeling organoid is a cerebral organoid.

[0130] Various methods and systems are available to examine and characterize the structural and functional features of organoids. For example, in some embodiments, morphological assessment of organoids can be performed using phase-contrast or confocal microscopy to evaluate cell shape, polarity, and spatial organization. Immunofluorescence staining may - 32 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 additionally be used to detect tissue-specific markers and their localization within the 3D structure. In some embodiments, molecular characterization may include the use of quantitative PCR or RNA sequencing to assess gene expression profiles, while protein-level analyses may be conducted using techniques such as Western blotting or enzyme-linked immunosorbent assay (ELISA) to detect specific signaling pathways or marker proteins. In some embodiments, functional assays may be utilized to evaluate physiological performance. In some embodiments, electrophysiological methods, such as calcium imaging or patch-clamp recordings, can be used to characterize neural or cardiac organoids. In some embodiments, barrier integrity for epithelial models can be assessed using transepithelial electrical resistance (TEER) measurements. In some embodiments, metabolic activity of an organoid can be evaluated using assays that measure oxygen consumption or glucose uptake. In some embodiments, histological analysis can be used to examining tissue architecture, such as distribution of specific cell types or extracellular matrix components, of an organoid. Non-limiting examples of such methods include but are not limited to paraffin embedding, sectioning, and hematoxylin and eosin (H&E) staining, or immunohistochemistry. In some embodiments, live-cell imaging and time-lapse microscopy can be used to monitor dynamic biological processes such as organoid growth, differentiation, and cell migration over time. In some embodiments, single-cell analysis techniques—such as flow cytometry or single-cell RNA sequencing—can be applied to assess cellular heterogeneity and population dynamics within the organoid. These methods, while not exhaustive, provide non- limiting examples of tools that can be used to ensure that organoid cultures exhibit appropriate phenotypic and functional characteristics for use in disease modeling, regenerative medicine, and therapeutic screening.

[0131] In some embodiments, the training dataset used for fine-tuning a pre-trained foundation model for specific tasks comprises gene transcriptional data collected from an in vitro organoid model. In some embodiments, the training dataset comprises single-cell transcriptional data collected from at least 1000 cells in an in vitro cell culture, at least 5000 cells in an in vitro cell culture, at least 10,000 cells in an in vitro cell culture, at least 50,000 cells in an in vitro cell culture, at least 100,000 cells in an in vitro cell culture, at least 500,000 cells in an in vitro cell culture. In some embodiments, the in vitro cell culture is any type of 2D or 3D cell cultures described herein. In specific embodiments, the in vitro cell culture comprises at least one cellular sphere. In specific embodiments, the in vitro cell culture comprises at least one - 33 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 organoid. In specific embodiments, the organoid is a disease-modeling organoid comprising at least one gene mutation associated with a target disease. In specific embodiments, the organoid is a disease-modeling organoid derived from a subject suffering from the target disease. In specific embodiments, the organoid is a disease-modeling organoid derived from a subject suspected of having a target disease. In specific embodiments, the organoid is a disease-modeling organoid derived from a subject at an elevated risk of developing a target disease. In specific embodiments, the organoid is a control organoid derived from healthy cells.

[0132] In some embodiments, the gene transcriptional data in the training dataset used for fine-tuning a pre-trained foundation model comprises various types of data. In some embodiments, the gene transcriptional data comprises single-cell RNA sequencing (scRNA-seq) data. In some embodiments, the scRNA-seq data capture the transcriptome of individual cells to reveal cell-specific gene expression profiles. In some embodiments, the scRNA- seq data offer full-length transcript coverage. In some embodiments, the gene transcriptional data comprises single-nucleus RNA sequencing (snRNA-seq) data that profile RNA from isolated nuclei. In other embodiments, the gene transcriptional data comprises total RNA sequencing data at a single-cell resolution. In some embodiments, the gene transcriptional data capture both coding and non-coding RNAs for a more comprehensive view of the transcriptome.

[0133] In some embodiments, the gene transcriptional data can be provided by methods and systems known in the art. For example, in some embodiments, the gene transcriptional data are provided by the Smart-seq or Smart-seq2 method or system. In other embodiments, the gene transcriptional data are provided using high-throughput methods that utilize droplet microfluidics to analyze gene expression at the single-cell level. In some embodiments, the gene transcriptional data are provided using the 10x Genomics Chromium method or system. In some embodiments, the gene transcriptional data are provided using the Drop-seq method or system. In some embodiments, the gene transcriptional data are provided using or the inDrop method or system. In some embodiments, the gene transcriptional data are provided using the Seq-Well method or system.

[0134] In some embodiments, the gene transcriptional data comprise spatial transcriptomic data. In some embodiments, the gene transcriptional data retain spatial context of gene expression in tissue sections. In some embodiments, the gene transcriptional data are provided by the 10x Genomics Visium method or system. In some embodiments, the gene transcriptional - 34 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 data are provided by the MERFISH method or system. In some embodiments, the gene transcriptional data are provided by the seqFISH method or system. In some embodiments, the gene transcriptional data are provided by the Slide-seq method or system.

[0135] In some embodiments, the gene transcriptional data comprises data derived from CRISPR-based perturbation assays that combine genetic perturbations with single-cell transcriptomics to study gene function and regulatory interactions. In some embodiments, the gene transcriptional data is provided by the Perturb-seq method or system. In some embodiments, the gene transcriptional data includes multimodal measurements such as data provided by the CITE-seq or REAP-seq method or system, where both mRNA and surface protein expression are simultaneously profiled in individual cells using antibody-derived tags. In other embodiments, the gene transcriptional data is analyzed in conjunction with chromatin accessibility data, such as using the scATAC-seq method or system, to link transcriptional activity with epigenomic regulation at the single-cell level.

[0136] In some embodiments, the present method involves extracting gene embeddings as output from the fine-tuned model. The gene embeddings extracted from the fine-tuned model capture biologically meaningful representations of genes derived from large-scale single-cell RNA sequencing (scRNA-seq) data. In some embodiments, these gene embeddings encode functional similarity and co-expression relationships by grouping together genes involved in the same biological pathways or functional modules. In some embodiments, the fine-tuned model is able to organize genes into biologically coherent clusters.

[0137] In some embodiments, the gene embeddings further encode gene regulatory network (GRN) structure. The model learns gene-gene relationships such that similarity networks built from embeddings correspond to known signaling pathways, such as those found in the Reactome database. In some embodiments, cosine similarity between gene embeddings is used to infer gene to gene co-expression relationship, supporting the idea that the embeddings reflect regulatory and functional connectivity.

[0138] In some embodiments, the gene embeddings also capture gene expression changes in response to perturbations. In some embodiments, by leveraging self-attention mechanisms, the fine-tuned model learns how specific gene knockouts affect the expression of other genes, allowing accurate prediction of gene expression profiles in untested perturbation conditions. In some embodiments, the model has the ability to represent gene interactions in a dynamic context. - 35 - NAI-5004565027v1Attorney Docket No.: 14876-002-228

[0139] In some embodiments, gene embeddings are capable of integrating multi-modal data, including gene expression, chromatin accessibility, and protein abundance. In some embodiments, the fine-tuned model supports the inclusion of condition tokens (such as modality, batch, or perturbation status), enabling the embeddings to represent context-specific biological variation.

[0140] In some embodiments, the fine-tuned model is capable of generating a gene network based on the gene embeddings. In some embodiments, the gene network comprises one or more clusters of related genes. In specific embodiments, the related genes are related in one or more features selected from co-expression pattern, function similarity, involvement in biological pathways or functional modules.

[0141] In some embodiments, the gene network is rendered in a multi-dimensional representation. In specific embodiments, the gene network is rendered in a 3D representation. In specific embodiments, the gene network is rendered in a 2D representation. In specific embodiments, the gene network is rendered in a 1D representation.

[0142] In some embodiments, the fine-tuned model is able to identify a common function shared by multiple genes in the gene cluster. In some embodiments, a common function shared by a pre-determined number of genes in one cluster is identified as the primary common function for the cluster. In some embodiments, the fine-tuned model identifies other genes in the cluster as having the primary common function. In some embodiments, the fined-tuned model is further able to annotate the generated gene network by associating the cluster with its primary common function. In some embodiments, the pre-determined number of genes is at least 2 genes, at least 3 genes, at least 4 genes, at least 5 genes, at least 10 genes, at least 20 genes.

[0143] In some embodiments, the fine-tuned model is able to identify multiple common functions shared by a pre-determined number of related genes in a gene cluster. In some embodiments, the multiple common functions relate to each other in a biological pathway. In some embodiments, the fine-tuned model identifies other genes in the cluster as having functions involved in the biological pathway. In some embodiments, the fine-tuned model is further able to annotate the generated gene network by associating the cluster with the biological pathway. In some embodiments, the pre-determined number of genes is at least 2 genes, at least 3 genes, at least 4 genes, at least 5 genes, at least 10 genes, at least 20 genes. - 36 - NAI-5004565027v1Attorney Docket No.: 14876-002-228

[0144] In some embodiments, the common function identified for a gene cluster is metabolism, cell differentiation, cell cycle, cell population proliferation, embryo development, lipd metabolism, localization, oxidative phosphorylation, signal transduction, apoptosis, immune response, DNA repair, transcriptional regulation, protein synthesis, cell migration, angiogenesis, autophagy, stress response, synaptic signaling, neurogenesis, hematopoiesis, lipid metabolism, cytokine production, mitochondrial function, epigenetic regulation, and cell-cell communication.

[0145] In some embodiments, the common function identified for a gene cluster is a neuronal function. In specific embodiments, the neuronal function is selected from neurogenesis, synaptic signaling, axon guidance, dendritic growth, synaptic plasticity, neurotransmitter release, neuronal migration, action potential propagation, ion channel regulation, myelination, neuronal differentiation, neurite outgrowth, long-term potentiation, long-term depression, neuronal survival, glial-neuronal interaction, synapse formation, neural circuit remodeling, and neurotransmitter reuptake.

[0146] In some embodiments, the common function identified for a gene cluster is being markers for a certain cell type. In specific embodiments, the genes in a cluster are markers for neurons. In specific embodiments, the genes in a cluster are markers for GABAergic neurons. In specific embodiments, the genes in a cluster are markers for dopaminergic neurons. In specific embodiments, the genes in a cluster are markers for glutamatergic neurons. In specific embodiments, the genes in a cluster are markers for serotonergic neurons. In specific embodiments, the genes in a cluster are markers for cholinergic neurons. In specific embodiments, the genes in a cluster are markers for noradrenergic neurons. In specific embodiments, the genes in a cluster are markers for histaminergic neurons. In specific embodiments, the genes in a cluster are markers for oligodendrocytes.

[0147] In specific embodiments, the genes in a cluster are markers for glial cells. In specific embodiments, the genes in a cluster are markers for astrocytes. In specific embodiments, the genes in a cluster are markers for oligodendrocytes. In specific embodiments, the genes in a cluster are markers for microglia. In specific embodiments, the genes in a cluster are markers for ependymal cells. In specific embodiments, the genes in a cluster are markers for NG2 glia cells. In specific embodiments, the genes in a cluster are markers for Schwann cells. In specific embodiments, the genes in a cluster are markers for satellite glia cells. In specific embodiments, the genes in a cluster are markers for Enteric glia cells. - 37 - NAI-5004565027v1Attorney Docket No.: 14876-002-228

[0148] In specific embodiments, the genes in a cluster are markers for neural progenitor cells. In specific embodiments, the genes in a cluster are markers for neuroepithelial cells. In specific embodiments, the genes in a cluster are markers for radial glial cells. In specific embodiments, the genes in a cluster are markers for apical radial glial cells. In specific embodiments, the genes in a cluster are markers for basal / outer radial glial cells. In specific embodiments, the genes in a cluster are markers for intermediate progenitor cells. In specific embodiments, the genes in a cluster are markers for outer subventricular zone progenitors. In specific embodiments, the genes in a cluster are markers for gliogenic progenitor cells.

[0149] In some embodiments, the fine-tuned model is able to identify the presence of one or more cell types in the biological sample by analyzing transcriptional data for cell type marker genes. In some embodiments, the fine-tune model is further able to annotate the generated gene network by associating gene clusters with one or more cell types.

[0150] In some embodiments, the fine-tuned model is able to identify one or more hub genes in a given gene cluster. In some embodiments, the one or more genes having the highest intra- cluster connectivity are ranked as having a great influence on the function of the gene cluster.

[0151] In some embodiments, intracluster connectivity is measured by distance and similarity metrics. In the analysis of gene networks, distance and similarity metrics play an important role in quantifying relationships between genes, gene expression profiles, or regulatory pathways. These metrics are used to assess how closely related two genes are based on their expression patterns, functional similarities, or positions within a biological network. By computing the degree of similarity or dissimilarity between gene vectors, researchers can identify co-expressed genes, infer gene function, detect regulatory modules, and construct or refine interaction networks. For example, Cosine Similarity is commonly used to compare gene expression profiles by focusing on the direction of expression changes rather than their magnitude, making it robust to differences in expression scale. Correlation Coefficient (e.g., Pearson or Spearman correlation) is often applied to measure the strength and direction of linear or monotonic relationships between gene expression profiles, making it particularly useful for identifying co-expressed genes across varying conditions. Euclidean Distance is effective for clustering genes based on raw expression differences, while Manhattan Distance is better suited for datasets with high-dimensional, sparse features such as single-cell RNA-seq data. Minkowski Distance provides flexibility by generalizing other distance measures depending on the - 38 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 parameter chosen. Jaccard Similarity can be used to compare binary gene activation states across different conditions, and Hamming Distance is useful for detecting mutations or variations in gene sequences. These metrics, applied appropriately, help decode the complex structure and function of gene regulatory networks.

[0152] Accordingly in some embodiments, the fine-tuned model is able to measure intracluster connectivity for any given number of genes in at least one gene cluster by measuring a distance and similarity metrics selected from Cosine Similarity, Correlation Coefficient, Euclidean Distance, Manhattan Distance, Minkowski Distance, Jaccard Similarity and Hamming Distance.

[0153] In some embodiments, the fine-tuned model is able to measure intracluster connectivity for any given number of genes in at least one gene cluster by measuring the connection number, which is the number of connections between said gene and other genes in the at least one gene cluster.

[0154] In some embodiments, the fine-tuned model is able to rank measured intracluster connectivity for a plurality of genes. In some embodiments, the fine-tuned model selects one or more genes having the top about 5% to about 10% intracluster connectivity among other genes in the cluster as the hub genes of the cluster. In some embodiments, the fine-tuned model is able to prioritize selection of hub gene based on the measured intracluster connectivity and identify one or more hub genes as a disease-modifying gene for the target disease.

[0155] In some embodiments, the training dataset comprises transcriptional data collected from both a diseased sample and a control sample that is free of the target disease, and the fine- tuned model is able to identify one or more differentially expressed genes (DEGs) that have different expression profile in diseased samples as compared to the control. In some embodiments, the fine-tuned model is able to identify one or more DEGs as a disease-modifying gene for the target disease.

[0156] In some embodiments, the fine-tuned model is able to determine whether a cluster of related genes is enriched with DEGs or other pre-selected genes that are known to associated with the target disease. In some embodiments, the primary common function of the gene clusters enriched with DEGs or other pre-selected genes is identified as a disease-modifying mechanism of the target disease. - 39 - NAI-5004565027v1Attorney Docket No.: 14876-002-228

[0157] In some embodiments, the fine-tuned model is able to compute an enrichment score for each gene cluster based on the number of DEG and / or pre-selected genes in the gene cluster and the total number of genes in the cluster. In some embodiments, one or more gene clusters having statistically significantly higher enrichment scores compared to other clusters are identified as clusters enriched with the DEGs and / or pre-selected genes. In some embodiments, the statistical test is Fisher’s Exact Test. In some embodiments, the statistical test comprises adjusting the resulting p-values for multiple hypothesis testing using one or more correction methods. In some embodiments, the correction methods can be selected from Bonferroni correction or the Benjamini–Hochberg (BH) procedure. In some embodiments, the threshold for determining statistically significance is set to be 0.05 or 0.1.

[0158] In some embodiments, the fine-tuned model is able to visualize DEGs in the disease- specific gene network with different colors representing upregulated and downregulated genes, respectively.

[0159] In some embodiments, the gene modules are identified as being enriched for one or more differentially expressed genes (DEGs), wherein the DEGs exhibit altered expression profiles in diseased samples relative to control samples. Enrichment may be assessed against a background gene set using a statistical test, such as Fisher’s Exact Test.

[0160] In some embodiments, a disease-specific training dataset used to fine-tune the pre- trained foundation model comprises data specific for a target disease, and the fine-tuned model is able to identify changes in one or more biological mechanism of actions (MOA) that are associated with the target disease. In some embodiments, the disease-specific training dataset comprises data collected from a diseased biological sample as described herein.

[0161] In some embodiments, the foundation model has been pretrained and fine-tuned with a control training dataset, where the control training dataset comprises data comparable in the amount and type to the disease-specific training dataset, which is collected from a control biological sample as described herein, where the control biological sample is free of the target disease. In specific embodiments, the diseased biological sample is isolated or derived from a subject having the target disease, and the control biological sample is isolated or derived from a subject free of the target disease. In specific embodiments, the control biological sample is isolated or derived from a healthy subject. In specific embodiments, the diseased biological - 40 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 sample comprises a disease-modeling organoid and the control biological sample comprises a healthy organoid of similar type, development stage and size.

[0162] In some embodiments, the disease-associated MOA comprises a change in the expression of one or more genes in the disease. In specific embodiments, the disease-associated MOA comprises upregulation or downregulation of one or more genes in the disease. In some embodiments, the disease-associated MOA comprises a change in the function of one or more genes or its encoded product in the disease. In some embodiments, the disease-associated MOA comprises a change in the temporal or spatial (e.g., tissue-specific) pattern of expression of one or more genes in the disease. In some embodiments, the disease-associated MOA comprises a change in one or more signaling pathways in the disease. In some embodiments, the disease- associated MOA comprises a change in cell fate of one or more types of cells in the disease. In some embodiments, the disease-associated MOA comprises a change in the differentiation, development, and / or proliferation of one or more types of cells in the disease. In some embodiments, the disease-associated MOA is causative for the target disease. In some embodiments, the disease-associated MOA is caused by the target disease. In some embodiments, the disease-associated MOA is a symptom for a target disease.

[0163] In some embodiments, the fine-tuned model trained with training dataset specific for a target disease is further able to identify a gene that is involved in one or more of the disease- associated MOA as a disease-modifying gene.

[0164] In some embodiments, the disease-associated MOA comprises one or more DEGs for which the expression is upregulated or downregulated in the disease-modeling organoid as compared to a healthy control. In some embodiments, the disease-modifying solution comprises inhibiting a DEG that is upregulated in the disease-modeling organoid for preventing, treating or managing the target disease in a subject. In some embodiments, the disease-modifying solution comprises activating a DEG that is downregulated in the disease-modeling organoid for preventing, treating or managing the target disease in a subject. In some embodiments, the disease-modifying solution comprises supplementing copies of a DEG that is downregulated in the disease-modeling organoid for preventing, treating or managing the target disease in a subject. - 41 - NAI-5004565027v1Attorney Docket No.: 14876-002-228

[0165] In some embodiments, the fine-tuned model trained with training dataset specific for a target disease is further able to infer a disease-modifying solution that is informative on the treatment, prevention and management of a target disease in a subject.

[0166] In some embodiments, the disease-associated MOA comprises one or more cellular signaling pathways that is enhanced or inhibited in the disease-modeling organoid as compared to a healthy control. In some embodiments, the disease-modifying solution comprises enhancing the cellular signaling pathway that is inhibited in the disease-modeling organoid as compared to a healthy control for treating, preventing or managing the target disease in a subject. In some embodiments, the disease-modifying solution comprises inhibiting the cellular signaling pathway that is enhanced in the disease-modeling organoid as compared to a healthy control for treating, preventing or managing the target disease in a subject.

[0167] In some embodiments, the disease-associated MOA comprises one or more cell types for which one or more features of the cell cycle are affected by the target disease. In some embodiments, differentiation of the cells is changed in the disease-modeling organoid as compared to a healthy control. In some embodiments, development of the cells is changed in the disease-modeling organoid as compared to a healthy control. In some embodiments, proliferation of the cells is changed in the disease-modeling organoid as compared to a healthy control. In some embodiments, apoptosis of the cells is changed in the disease-modeling organoid as compared to a healthy control. In some embodiments, the relative abundance of the cells to other cell types is changed in the disease-modeling organoid as compared to a healthy control.

[0168] In some embodiments, the disease-modifying solution comprises detecting one or more biomarkers for the affected cell type for diagnosis of the target disease in a subject. In some embodiments, the disease-modifying solution comprises detecting one or more biomarkers of the affected cell type in a subject who is receiving a treatment of the target disease to evaluate the subject’s response to the treatment. In some embodiments, the disease modifying solution comprises detecting one or more biomarkers of the affected cell type in a subject for evaluating the likelihood of the subject to react to a treatment of the target disease.

[0169] In specific embodiments, the training dataset comprises gene transcriptional data collected from a diseased biological sample isolated from or derived from one or more subjects - 42 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 having Parkinson’s disease. In specific embodiments, the training dataset comprises gene transcriptional data collected from a Parkinson’s Disease modeling midbrain organoid.

[0170] In one aspect, provided herein are methods for producing disease-modeling organoids carrying one or more mutations in at least one disease-modifying gene identified by the present methods and systems. In some embodiments, the method for producing a disease-modeling organoid comprises providing an iPSC cell line carrying one or more mutation in at least one disease-modifying gene, and culturing the iPSC cell line under a suitable condition to induce differentiation of cells into the disease-modifying organoid.

[0171] In one aspect, provided herein are systems and related methods for analyzing or modeling disease-associate MOA. In some embodiments, the system comprises a diseased biological sample and a computer system comprising one or more processors and a non- transitory computer readable storage medium including software stored thereon, wherein the software comprises executable instructions that, as a result of execution, causes the one or more processors of the computer system to: (a) load a pre-trained foundation model as described herein; (b) receiving a training dataset comprising gene transcriptional data collected from the diseased biological sample as described herein; (c) process the training dataset using the pre-trained foundation model to generate a disease-specific gene network as output as described herein; and (d) infer a mechanism of action associated with the target disease, or a disease- modifying solution, in either case, based on the generated output in (c).

[0172] In some embodiments, the diseased biological sample comprises a sample isolated from a subject having, or is at an elevated risk of developing, the target disease. In some embodiments, the diseased biological sample comprises a disease-modeling organoid for the target disease. In some embodiments, the disease-modeling organoid comprises cells carrying a mutation in one or more disease-modifying genes for the target disease. In some embodiments, the disease-modeling organoid comprises cells carrying a mutation in one or more genes involved in a disease-associated MOA of the target disease. In some embodiments, the disease- modifying gene is identified by the present methods or systems. In some embodiments, the disease-modifying gene is known to associate with the target disease. In some embodiments, the - 43 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 disease-associated MOA is identified by the present methods or systems. In some embodiments, the disease-associated MOA is known to associate with the target disease.

[0173] In some embodiments, provided herein is a method for analyzing or modeling disease-associate MOA via in silico perturbation. In some embodiments, the method comprises providing a disease-modeling organoid. In some embodiments, the method further comprises collecting gene transcriptional data from the disease-modifying organoid into a training dataset for a pre-trained foundation model as described herein. In some embodiments, the method further comprises feeding the training dataset to the pre-trained foundation model and extract gene embeddings as an output. In some embodiments, the method further comprises comparing the gene embeddings with control gene embeddings to infer biological perturbation induced by the gene mutation.

[0174] In some embodiments, the control gene embeddings is extracted from the pre-trained foundation model after the model is trained with a training dataset containing comparative data collected from a control biological sample. In some embodiments, the control biological sample is isolated from of derived from a healthy subject. In some embodiments, the disease-modeling organoid carries a mutation in one or more disease-modifying gene for the target disease. In some embodiments, the disease-modeling organoid carries a mutation in one or more genes that is involved with the disease-associated MOA. In some embodiments, the control biological sample comprises a disease-modeling organoid that does not contain the same gene mutation. In some embodiments, the disease-modeling organoid is exposed to a candidate therapeutic agent under a suitable condition for the disease-modeling organoid to react to the candidate therapeutic agent. In some embodiments, the control biological sample comprises a disease-modeling organoid that has not been exposed to the therapeutic agent.

[0175] In some embodiments, the mechanism of action comprises one or more DEGs for which the expression is upregulated or downregulated in the disease-modeling organoid as compared to a healthy control. In some embodiments, the disease-modifying solution comprises inhibiting a DEG that is upregulated in the disease-modeling organoid for preventing, treating or managing the target disease in a subject. In some embodiments, the disease-modifying solution comprises activating a DEG that is downregulated in the disease-modeling organoid for preventing, treating or managing the target disease in a subject. In some embodiments, the disease-modifying solution comprises supplementing copies of a DEG that is downregulated in - 44 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 the disease-modeling organoid for preventing, treating or managing the target disease in a subject.

[0176] In some embodiments, the mechanism of action comprises one or more cellular signaling pathways that is enhanced or inhibited in the disease-modeling organoid as compared to a healthy control. In some embodiments, the disease-modifying solution comprises enhancing the cellular signaling pathway that is inhibited in the disease-modeling organoid as compared to a healthy control for treating, preventing or managing the target disease in a subject. In some embodiments, the disease-modifying solution comprises inhibiting the cellular signaling pathway that is enhanced in the disease-modeling organoid as compared to a healthy control for treating, preventing or managing the target disease in a subject.

[0177] In some embodiments, the mechanism of action comprises one or more cell types for which one or more features of the cell cycle are affected by the target disease. In some embodiments, differentiation of the cells is changed in the disease-modeling organoid as compared to a healthy control. In some embodiments, development of the cells is changed in the disease-modeling organoid as compared to a healthy control. In some embodiments, proliferation of the cells is changed in the disease-modeling organoid as compared to a healthy control. In some embodiments, apoptosis of the cells is changed in the disease-modeling organoid as compared to a healthy control. In some embodiments, the relative abundance of the cells to other cell types is changed in the disease-modeling organoid as compared to a healthy control.

[0178] In some embodiments, the disease-modifying solution comprises detecting one or more biomarkers for the affected cell type for diagnosis of the target disease in a subject. In some embodiments, the disease-modifying solution comprises detecting one or more biomarkers of the affected cell type in a subject who is receiving a treatment of the target disease to evaluate the subject’s response to the treatment. In some embodiments, the disease modifying solution comprises detecting one or more biomarkers of the affected cell type in a subject for evaluating the likelihood of the subject to react to a treatment of the target disease.

[0179] In some embodiments, the disease-associated MOA inferred by the computer system is used to improve the disease-modeling organoid in the system. In specific embodiments, the disease-associated MOA involves one or more disease-modifying gene, and disease-modeling organoid is improved by mutating one or more disease-modifying gene in the organoid. - 45 - NAI-5004565027v1Attorney Docket No.: 14876-002-228

[0180] In some embodiments, the disease-modeling organoid is exposed to a candidate therapeutic agent for treating the target disease. In some embodiments, the disease-associated MOA inferred by the computer system comprises changes in expression of one or more genes in reaction to the therapeutic agent. In some embodiments, wherein such change in gene expression is shown or has been known to associate with amelioration or elimination of a symptom of the target disease, the computer system further infers a disease-modifying solution comprising the use of the candidate therapeutic agent for treating, preventing or managing the target disease.

[0181] In some embodiments, the disease-associated MOA inferred by the computer system comprises changes in function of one or more genes in reaction to the therapeutic agent. In some embodiments, wherein such change in gene function is shown or has been known to associate with amelioration or elimination of a symptom of the target disease, the computer system further infers a disease-modifying solution comprising the use of the candidate therapeutic agent for treating, preventing or managing the target disease.

[0182] In some embodiments, the disease-associated MOA inferred by the computer system comprises changes in abundance of a cell type in reaction to the therapeutic agent. In some embodiments, wherein such change in cell abundance is shown or has been known to associate with amelioration or elimination of a symptom of the target disease, the computer system further infers a disease-modifying solution comprising the use of the candidate therapeutic agent for treating, preventing or managing the target disease.

[0183] In some embodiments, the disease-associated MOA inferred by the computer system comprises changes in activity of one or more cell types in reaction to the therapeutic agent. In some embodiments, wherein such change in cell activity is shown or has been known to associate with amelioration or elimination of a symptom of the target disease, the computer system further infers a disease-modifying solution comprising the use of the candidate therapeutic agent for treating, preventing or managing the target disease.

[0184] In some embodiments, the disease-associated MOA inferred by the computer system comprises changes in organ function modeled by the disease-modeling organoid in reaction to the therapeutic agent. In some embodiments, wherein such change in organ function is shown or has been known to associate with amelioration or elimination of a symptom of the target disease, - 46 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 the computer system further infers a disease-modifying solution comprising the use of the candidate therapeutic agent for treating, preventing or managing the target disease.

[0185] In some embodiments, the disease-associated MOA inferred by the computer system comprises changes in one or more symptoms of the target disease in reaction to the therapeutic agent. In some embodiments, wherein such change in disease symptoms is shown or has been known to associate with the treatment, prevention, or management of the target disease, the computer system further infers a disease-modifying solution comprising the use of the candidate therapeutic agent for treating, preventing or managing the target disease. 7. EXAMPLES

[0186] The examples in this section (i.e., Section 7) are offered by way of illustration, and not by way of limitation.

[0187] To interrogate transcriptional programs in Parkinson’s disease (PD) midbrain organoids, a computational framework integrating foundation models with midbrain organoid transcriptomics was developed. Human iPSC-derived midbrain organoids were generated from control and familial PD lines harboring GBA1 or LRRK2 mutations (FIG.2). Then a pipeline that extracts gene embeddings from fine-tuned single-cell foundation models to construct high- dimensional gene interaction networks was implemented. The network revealed gene modules with discrete biological functions and can be used to interpret disease modeling and advance drug discovery research (FIG.4). 7.1 Example 1: Methods and Materials

[0188] Establishment of Midbrain Organoid Model

[0189] Human iPSC-derived midbrain organoids were generated from control and familial Parkinson disease cell lines harboring GBA1 or LRRK2 mutations (FIG.2).

[0190] Single cell data processing

[0191] Raw count data derived from midbrain organoid single-cell sequencing were normalized and processed utilizing the R Seurat package (Slovin S, et al. Single-Cell RNA Sequencing Analysis: A Step-by-Step Overview. Methods Mol Biol.2021;2284:343-65). Integration of various single-cell sequencing datasets was achieved through Seurat’s Canonical Correlation Analysis (CCA) method via CCAIntegration. Cell clustering was conducted employing the Leiden algorithm to identify distinct cellular populations. Marker genes for each - 47 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 cluster were identified using the Wilcoxon rank-sum test to ascertain statistically significant differences in gene expression. The annotation of cell clusters was refined by comparing the identified markers against previously published markers specific to midbrain cells (Smajic S, et al. Single-cell sequencing of human midbrain reveals glial activation and a Parkinson-specific neuronal state. Brain.2022;145(3):964-78).

[0192] Differentially expressed genes (DEGs) between PD GBA1 versus wildtype samples, PD LRRK2 versus wildtype samples, or idiopathic PD samples versus wildtype samples from single-cell sequencing were performed using Wilcox rank sum test in R Seurat. DEGs from PD GBA1 or LRRK2 were selected by log2 transformed fold change >=0.5 or <=- 0.5 and Benjamini-Hochberg (BH) adjusted P value <0.05.

[0193] Bulk RNA-Seq data processing

[0194] Raw count data from bulk RNA-Seq were analyzed using R DESeq2 Wald test. DEGs were selected by log2 transformed fold change >=1 or <=-1 and BH p value <0.05. Single cell foundation model fine tuning

[0195] Fine-tuning of single-cell foundation models, including scGPT (Cui H, Wang C, Maan H, Pang K, Luo F, Duan N, et al. scGPT: toward building a foundation model for single- cell multi-omics using generative AI. Nat Methods.2024. Epub 2024 / 02 / 27. doi: 10.1038 / s41592-024-02201-0. PubMed PMID: 38409223), Geneformer (Theodoris CV, Xiao L, Chopra A, Chaffin MD, Al Sayed ZR, Hill MC, et al. Transfer learning enables predictions in network biology. Nature.2023;618(7965):616-24. Epub 2023 / 06 / 01. doi: 10.1038 / s41586-023- 06139-9. PubMed PMID: 37258680; PubMed Central PMCID: PMCPMC10949956), and scFoundation (Hao M, Gong J, Zeng X, Liu C, Guo Y, Cheng X, et al. Large-scale foundation model on single-cell transcriptomics. Nat Methods.2024;21(8):1481-91. Epub 2024 / 06 / 07. doi: 10.1038 / s41592-024-02305-7. PubMed PMID: 38844628), was performed using single-cell sequencing data from midbrain organoids containing over 100,000 cells in total from seven familial PD patients with GBA1 mutation, two samples from a familial PD patient with LRRK2 mutations, and four wildtype controls. The full scGPT model and scGPT brain model underwent fine-tuning with a batch size of 16 over 5 epochs. Similarly, Geneformer was fine-tuned using a batch size of 16 but extended to 10 epochs. Owing to its substantial memory requirements, the scFoundation model was fine-tuned on a reduced dataset of 30,000 cells and limited to the last two layers, with a batch size of 16 over 10 epochs. - 48 - NAI-5004565027v1Attorney Docket No.: 14876-002-228

[0196] Foundation Model-based Gene Network

[0197] Gene embeddings were extracted from the fine-tuned foundation models. Relationships among genes were inferred through the application of the K-nearest neighbor algorithm, followed by the use of Louvain clustering to identify distinct gene modules. These modules were then annotated based on enrichment analyses performed using Fisher’s exact test, with comparisons against established public databases including Gene Ontology (Gene Ontology Consortium: going forward. Nucleic Acids Res.2015;43(Database issue):D1049-56), Reactome (Rothfels K, et al. Using the Reactome Database. Curr Protoc.2023;3(4):e722), and PangloDB (Franzen O, et al.. PanglaoDB: a web server for exploration of mouse and human single-cell RNA sequencing data. Database (Oxford).2019;2019). To facilitate visualization, t-SNE was employed to reduce the dimensionality of gene embeddings to two dimensions, allowing for the plotting of the network. Additionally, an interactive website was developed using Python Dash to enable dynamic visualization and exploration of the gene network.

[0198] hdWGCNA Co-expression Network Analysis

[0199] The Parkinson’s disease and control single-cell RNA sequencing dataset used for foundation model training was further analyzed with the hdWGCNA package (Morabito, S., Reese, F., Rahimzadeh, N., Miyoshi, E. & Swarup, V. hdWGCNA identifies co-expression networks in high-dimensional transcriptomics data. Cell Rep Methods 3, 100498, doi:10.1016 / j.crmeth.2023.100498 (2023)). Following the recommended workflow, genes expressed in at least 5% of all cells were retained for network construction. To reduce noise and enhance stability, metacells were generated by aggregating small groups of transcriptionally similar cells using a k-nearest neighbor (kNN) approach. These metacells were then used to build a weighted co-expression network with an approximately scale-free topology.

[0200] Network Information Content Comparison

[0201] To assess the informational content of networks constructed using various foundation models, the concordance of gene modules were analyzed with established public knowledge databases (Gene Ontology Consortium; Reactome Database; PanglaoDB; Supra). A network ideal for functional analysis of midbrain organoids should produce gene modules that are significantly enriched in neuron-related activities. Therefore, the gene modules from each network were ranked based on the p-value of the most significantly over-represented GO biological process for each module. The p-values and relevance to known neuron-related - 49 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 processes were compared across different networks. The fine-tuned scGPT human brain foundation model demonstrated the most favorable p-value scores and the highest enrichment in neuron-related processes.

[0202] Additional concordance tests involved comparing the p-values from enrichment tests for different modules with differentially expressed genes (DEGs) from GBA1 Parkinson’s Disease (PD) patient-derived midbrain organoids versus control organoids. An effective network should identify modules that are highly enriched with PD-associated DEGs. The fine-tuned scGPT human brain foundation model excelled once more, showing the best enrichment for modules associated with PD DEGs.

[0203] Network analysis of Parkinson’s Disease (PD) associated genes

[0204] PD-associated genes are defined as genes showing differentially expression between PD and wildtype samples. These genes were analyzed by the network constructed using the fine-tuned scGPT human brain foundation model. Within the network, genes were color- coded to reflect their expression patterns: red indicated genes up-regulated in PD compared to wild type, and blue denoted those down-regulated (see FIG.5). Enrichment of these DEGs within the gene modules was evaluated using Fisher’s exact test to ascertain significant associations. To compare DEG enrichment scores of different PD subtypes, the gene modules were ranked by Fisher’s exact test P values. The top modules with P values <0.05 were compared among different PD subtypes.

[0205] Data source

[0206] PD midbrain organoid single cell sequencing data are deposited in the NCBI GEO with accession IDs GSE268784 with seven familial PD patients with GBA1 mutation and three wild-type controls, and GSE133894with one familial PD patient with LRRK2 mutation and one wild-type control. PD midbrain organoid bulk RNA-Seq is BrainStorm’s private data with three familial PD patients with GBA1 mutation and three wild-type controls . Single cell sequencing data for idiopathic PD midbrain was from GSE157783 with five idiopathic PD patients and five controls.

[0207] Computational Resources

[0208] Model fine-tuning was performed using AWS EC2 instances g4dn.2xlarge (1 x NVIDIA T4 Tensor Core), g4dn.12xlarge (4 x NVIDIA T4 Tensor Core) and g6.xlarge (1 x NVIDIA L4). - 50 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 7.2 Example 2: Establishment of the AI Model

[0209] To compare performances of foundation models as applied to identifying disease- modifying genes in Parkinson's disease, gene networks constructed from pre-trained and fine- tuned versions of three foundation models, scGPT, Geneformer, and scFoundation, were compared with each other. Non foundation model based method, hdWGCNA, was also included and compared. Each foundation models were originally trained on 10 to 50 million single-cell profiles. For each model, gene embeddings were extracted and reduced to two dimensions for visualization (FIGS.7A-7D). Louvain clustering was used to identify gene modules, which were then functionally annotated by comparing to public knowledge databases including Gene Ontology (GO) and Reactome. The gene networks were evaluated for their concordance with the public knowledge databases by enrichment scores.

[0210] Genome-wide association studies (GWAS) have identified numerous genetic loci associated with Parkinson's disease. These studies have revealed a complex genetic architecture for Parkinson's disease, implicating genetic associations with disease risk. Accordingly, additional benchmarking was performed by assessing module enrichment for Parkinson’s disease (PD)-relevant signals, including PD GWAS loci and differentially expressed genes (DEGs) from patient-derived organoids..

[0211] The scGPT brain model fine-tuned on midbrain organoid data consistently outperformed all other configurations. It produced gene modules with comparable enrichment score to scGPT all cells model and significantly higher enrichment scores for neuron-specific biological processes after fine tuning (FIG.7E). Furthermore, the scGPT-derived network showed the strongest enrichment for PD GWAS signals (P < 1 × 10⁻¹⁵), indicating its ability to capture genetically relevant disease modules (FIG.11A). Bulk RNA-Seq profiles were generated from midbrain organoids derived from seven Parkinson's disease patients with GBA1 mutations and three control donors. scGPT-based modules exhibited the greatest concordance with DEGs from this dataset (FIG.11B) and visually appealing dense gene clustering (FIG.12), further validating their disease relevance.

[0212] The conventional hdWGCNA method requires genes to be expressed in more than 5% of all cells, which substantially limits the number of genes incorporated into the network. As a result, only 3,951 genes were clustered into five modules, in contrast to the >16,000 genes and >20 modules consistently captured by foundation model–based networks. Moreover, the - 51 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 enrichment scores generated by hdWGCNA were generally lower than those from the scGPT brain fine-tuned model across GO biological processes, PD GWAS (at a 1e-5cutoff), and DEGs, with the exception of PD GWAS at the more stringent 5e-8cutoff. Given the markedly smaller number of genes and modules detected, the utility of hdWGCNA for interpreting single-cell sequencing datasets was limited.

[0213] Finally, network topology metrics were computed to assess structural properties. In the context of AI models representing complex systems or data structures, the concepts of small-worldness and the Gini coefficient offer valuable insights. Small-worldness is a characteristic of complex networks that possess both high local clustering and short average path lengths between nodes. When applied to AI models, small-worldness can be used to describe the network’s structure and how information flows within it. A network with a high small-worldness implies high relevance between nodes, as connections are both localized (tightly connected clusters) and span longer distances (short paths between any two nodes). Gini coefficient can be used as a metric to evaluate the performance of classification models, particularly in credit risk assessment, by quantifying their ability to discriminate between positive and negative instances.

[0214] The fine-tuned scGPT brain model network demonstrated the highest small- worldness and Gini coefficient compared to other and fine-tuned foundation models and non foundation model method hdWGCNA, suggesting a modular and hierarchically organized architecture (FIGS.13A and 13B). Based on superior biological relevance, disease concordance, and network organization, the fine-tuned scGPT brain model was selected as the basis for all subsequent analyses and refer to it hereafter as the gene network. 7.3 Example 3: Gene network analysis identifies dysregulated neurogenesis in PD midbrain organoids

[0215] To evaluate the robustness and disease relevance of midbrain organoids derived from Parkinson Disease patients and healthy controls, bulk RNA-Seq data across six independent experimental batches were generated. Transcriptomic profiles demonstrated high reproducibility, with pairwise correlation coefficients exceeding 0.92 across samples (FIG.8A). The reproducibility of gene expression was established across organoids from different patients and different batches by analyzing bulk RNA-seq data. Principal variance component analysis (PVCA) revealed that gene expression variance was primarily driven by donor genotype, disease - 52 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 status, and gender, with minimal contribution from batch effects (FIG.8B), underscoring the reproducibility and biological fidelity of the organoid model.

[0216] To interpret gene expression changes at the pathway level, standard enrichment analysis of DEGs between PD and control organoids were performed using multiple databases, including Gene Ontology (GO), Reactome, WikiPathways, and PanglaoDB (FIG.8C, FIGS. 14A-14D). Consistent enrichment of neurogenesis, cell differentiation, and cell cycle pathways was observed. However, these predefined annotations were often overlapping, redundant, and agnostic to gene co-expression or network context.

[0217] In contrast, projecting DEGs onto the foundation model–derived gene network resolved multiple discrete modules with coherent biological themes, including cell cycle, cell differentiation, neurogenesis, synaptic signaling, and metabolism (FIGS.8D and 8E). Notably, while GO enrichment broadly identified “neurogenesis” (GO:0022008), the network distinguished two neurogenesis-associated modules with opposing regulatory trends: M7 and M10. Further inspection revealed that DEGs annotated under the GO neurogenesis term were distributed across at least three distinct modules, M3 (cell differentiation), M7, and M10, each with unique expression signatures (FIGS.15A-15D).

[0218] M7 appears to reflect a transitional progenitor or glial lineage state. It includes progenitor- and glia-associated genes such as LRP2, ST18, NKX6-2, as well as oligodendrocyte- related myelin genes MBP and PLP1, possibly representing an early neurogenic or glial diversification program (FIG.8F). Additionally, M7 includes matrix and adhesion genes (COL4A5, ITGB8, LTBP1) suggestive of structural remodeling. By contrast, M10 likely corresponds to post-mitotic, maturing neurons undergoing synaptogenesis and circuit integration. It is enriched for genes involved in neurotransmission (GRIA1, GRIK1 / 2, SCN1A / 2A, KCNJ3, KCNQ5), synaptic structure (SNCA, SV2B, SYNPR), and neurodevelopmental regulators (BCL11B, MEF2C, SOX2-OT, POU6F2) (FIG.8G).

[0219] The cell cycle module (M11) is dense gene module away from the other gene modules. M11 predominantly consisted of upregulated genes in PD organoids, including canonical mitotic regulators (CDK1, CDK2, CCNB1, CCNA2), the G2 / M transcription factor FOXM1, replication and repair genes (MCM3, MCM5, BRCA1, BRCA2, FANCI), and proliferation markers (MKI67, RRM2, TOP2A). Comparison of this module to public databases showed strong alignment with the Reactome “Cell Cycle” pathway. However, GO-based cell - 53 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 cycle terms assigned a broader gene set spanning additional modules, including M3 (cell differentiation), demonstrating the higher specificity and resolution achieved by the network- based approach (FIGS.15A-15D).

[0220] Together, these findings highlight the advantage of foundation model–based networks in resolving functionally distinct gene modules that would otherwise be conflated in traditional enrichment analyses, enabling finer dissection of disease-relevant transcriptional programs. 7.4 Example 4: snRNA-Seq reveals cell-type shifts underlying transcriptional changes in PD organoids

[0221] To dissect the cellular heterogeneity underlying the transcriptional alterations observed in Parkinson Disease midbrain organoids single-nucleus RNA sequencing (snRNA- Seq) were performed on Parkinson Disease and control iPSC-derived midbrain organoids. Dimensionality reduction using Uniform Manifold Approximation and Projection (UMAP), followed by marker gene expression analysis, revealed the presence of major neuronal and glial subtypes, including distinct clusters of dopaminergic neurons, GABAergic neurons, radial glia, and neural progenitors (FIGS.9A and 9B, FIG.16). Among the dopaminergic populations, three key subsets were identified: DA-1 (expressing TH, VGLUT2, and SNCA), DA-2 (primarily TH+), and two hybrid DA&GABA clusters, likely representing intermediate or transitional states.

[0222] To infer developmental trajectories, Monocle 3 pseudotime analysis was performed, which confirmed lineage progression from neural progenitors toward differentiated dopaminergic and GABAergic neurons, as well as a distinct branch leading to radial glial cells (FIGS.17A and 17B). This analysis revealed coordinated transcriptional programs underlying cell fate decisions and highlighted divergence between neuronal and glial lineages.

[0223] To contextualize these cell types within the gene network, cell-type–specific markers were projected onto the foundation model–derived modules. Cell markers for Dopaminergic neurons, GABAergic neurons, and progenitor cells were enriched in the downregulated M10 neurogenesis module, consistent with impaired neuronal maturation and reduced neurogenic output in Parkinson Disease organoids (FIGS.9C and 9D, FIGS.18A and 18B). In contrast, radial glial markers mapped predominantly to the upregulated M7 neurogenesis and M11 cell cycle modules, indicating an expansion of proliferative glial-like - 54 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 populations in Parkinson Disease organoids. These opposing trends between neurogenic and gliogenic programs were consistent with bulk RNA-Seq findings and suggested that the observed transcriptomic differences may be driven in part by shifts in cell-type composition.

[0224] To quantitatively validate this hypothesis, FARDEEP-based cell-type deconvolution was applied to bulk RNA-Seq profiles. The analysis confirmed a significant reduction in dopaminergic neuron proportions and a corresponding increase in radial glia in PD- derived organoids compared to controls (FIGS.9E and 9F). These findings support a model in which neurodegenerative or developmental impairments in PD organoids lead to a decline in neuronal populations and compensatory expansion of glial precursors, recapitulating features of Parkinson Disease pathophysiology in a human in vitro system. 7.5 Example 5: Gene network reveals conserved neurogenic dysfunction across PD models

[0225] To examine disease convergence, DEG profiles across GBA1 mutant Parkinson Disease PD GBA1, LRRK2 mutant Parkinson Disease PD LRRK2, and idiopathic Parkinson Disease (idiopathic PD) samples were compared using the gene network. Shared downregulation was observed for genes related to neurogenesis, synaptic signaling and cell differentiation in dopaminergic related neurons in GBA1 and LRRK2 mutant organoids and idiopathic patient brain (FIG.10A, FIG.19).

[0226] To understand the shared gene expression profiles in the three different experiments, genes that were significantly differentiated expressed across experiments were evaluated (FIG.10B). While there were no genes that were significantly differentiated in all the three experiments, the genes down-regulated in two out of the three experiments showed significantly higher network connections (cosine similarity >0.25) and average cosine similarity, highlighting their importance of regulating the neurogenesis process (FIGS.10C and 10D). These include TENM2, SEMA3E, and CHL1 that are involved in axon guidance and synapse formation. LMO3 and SAMD5 that regulate neuronal differentiation, while PDE4D and CADPS that contribute to intracellular signaling and neurotransmitter release. Their downregulation in Parkinson Disease organoids suggests impaired neuronal circuit development.

[0227] To test the consistency of expression patterns in the gene modules across different experiments, a correlation analysis was performed using DEGs in M10 from PD GBA1 versus control compared to DEGs from M10 from PD LRRK2 versus control and idiopathic PD versus - 55 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 control. The DEGs in M10 were also compared to DEGs from M3, M10 and randomly selected genes as negative controls. DEGs from the same module M10 showed significant correlation in gene embeddings (FIG.10E), and gene expression correlation from single cell sequencing and bulk RNA-Seq (FIG.20). This confirms that foundation model-derived modules capture conserved disease-relevant transcriptional architecture.

[0228] The above studies demonstrated establishment of a Parkinson Disease 3D midbrain organoid model having expected functionality and cell compositions for a Parkinson Disease midbrain. These human iPSC-derived midbrain organoids robustly recapitulate key features of Parkinson Disease-relevant brain regions, including dopaminergic and glial lineages. Bulk RNA-Seq across six independent batches demonstrated high transcriptomic reproducibility with minimal batch effects, supporting the robustness of the model. Single-nucleus RNA-Seq resolved major neuronal and glial subtypes, including TH+ / VGLUT2+ dopaminergic neurons, GABAergic neurons, and radial glia. Cell-type deconvolution revealed disease-associated shifts in cell composition, most notably, reduced dopaminergic neurons and increased radial glia in PD organoids, consistent with neurodegenerative phenotypes.

[0229] The network captured Parkinson Disease -relevant gene modules enriched for GWAS loci, synaptic signaling, and neurodevelopmental regulators, prioritizing candidate targets with genetic and functional support. Integration of single-cell and bulk transcriptomics enables screening for compounds that normalize disease-associated module expression patterns. Cell-type resolved expression within modules supports identification of cell-type specific targets (e.g., for dopaminergic neuron rescue or glial modulation). The organoid-network platform provides a scalable system for in silico perturbation, gene prioritization, and phenotypic screening of small molecules in human-relevant 3D tissue.

[0230] Conclusion and Discussion. The discovery of genetic forms of PD, especially mutations in the GBA1 and LRRK2 genes—the most common genetic mutations associated with the disease—has opened new avenues for developing disease-modifying therapies. Applicant’s research has shown that midbrain organoids derived from patient iPSCs with pathogenic GBA1 and LRRK2 mutations accurately replicate dopamine neuron loss observed in PD. Particularly, Applicant established organoid reproducibility across different patients and batches by analyzing bulk RNA-seq data, which revealed highly consistent gene expression across six independent differentiation batches (Correlation R>0.92). To further enhance disease modeling and - 56 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 therapeutic discovery, the Foundation Model-based gene network tailored to PD drug discovery was established as described in the study described above. While traditional gene set enrichment analysis provided a high-level view of key dysregulated processes, it lacked the granularity needed for insights into specific pathways and direct comparisons to human PD brain tissue. These findings confirm the stability and reliability of Applicant’s midbrain organoid platform and underscore the need for advanced data analysis to explore disease mechanisms in depth.

[0231] An analysis workflow utilizing pre-trained single-cell foundation models was developed, including the scGPT and scFoundation models implemented using NVIDIA® CUDA® and the Geneformer® model implemented using NVIDIA® BioNeMo® and Hugging Face®. The models were fine-tuned using scRNA-seq data from midbrain organoids derived from patients with mutations in GBA1 or LRRK2. Disease-specific gene clusters were identified using the gene embeddings for dimension reduction, community detection, and cluster annotation matching known biological information. Finally, the network with gene expression data from patient-derived organoids and patient brain tissue were overlaid, enabling disease modeling and target discovery. The approach was further validated by using rigorous benchmarking to compare different foundation models, including scGPT and Geneformer, as well as traditional gene clustering algorithms, such as hdWGCNA, to assess their performance in capturing PD-relevant pathways. The fine-tuned scGPT human brain model outperformed the others in detecting processes critical to PD pathology, such as neurogenesis and cell cycle dysregulation.

[0232] The study and analysis demonstrated that patient-derived midbrain organoid models effectively capture major pathogenic pathways in PD, especially in dopamine neurons. Overlaying the network with differentially-expressed genes from GBA1-PD organoids revealed consistent downregulation of genes involved in neurogenesis and cell differentiation, aligning with findings in actual PD patient brain tissue. Also discovered was cell cycle dysfunction, likely resulting from disease-related defects in neural precursor differentiation. These findings support the concept that early neurodevelopmental defects can set the stage for cellular vulnerabilities that manifest later as adult-onset neurodegenerative disorders.

[0233] The network’s sensitivity in analyzing organoids derived from PD patients with distinct mutations was further assessed. Consistent changes in neurogenesis and cell differentiation across GBA1-PD and LRKK2-PD midbrain organoids were observed. Processes - 57 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 that differ between these two genotypes were also observed, such as lipid metabolism dysfunction in the GBA1-PD organoids and oxidative phosphorylation in the LRRK2-PD organoids. This highlights the platform’s potential as a “clinical trial in a dish,” offering a high- fidelity model to dissect Parkinson’s disease mechanisms and de-risk the clinical translation of potential therapies. Moreover, the Foundation Model-based network supports in silico perturbation screening, enabling drug response simulations to pinpoint therapeutic candidates and biomarkers.

[0234] In conclusion, a robust patient-derived organoid platform that accurately mirrors PD-specific molecular biomarkers and disease phenotypes has been established. The Foundation Model-based network analysis has further strengthened this iPSC-derived midbrain organoid platform, capturing essential biological processes driving PD progression and uncovering both shared and genotype-specific disease mechanisms. This approach paves the way for precision therapeutics tailored to individual genetic and molecular profiles. - 58 - NAI-5004565027v1

Claims

Attorney Docket No.: 14876-002-228 WHAT IS CLAIMED:

1. A method for generating a disease-specific gene network for a target disease, comprising: (a) providing a training dataset comprising gene transcriptional data collected from a diseased biological sample, wherein the diseased biological sample comprises a cellular sphere comprising cells descending from a cell isolated from a diseased subject suffering from the target disease; (b) feeding the training data set to a foundation deep learning model pre-trained with a sufficient quantity of gene transcriptional data, thereby obtaining a fine-tuned model specific for the target disease; (c) extracting gene embeddings as output from the fine-tuned model; and (d) generating the disease-specific gene network based on the gene embeddings; wherein the gene network comprises at least one cluster of related genes.

2. The method of claim 1, further comprising (e) identifying a primary common function shared by a pre-determined number of related genes in at least one cluster; and annotating the disease-specific gene network by associating the cluster with the primary common function.

3. The method of claim 1, further comprising (e) identifying multiple common functions shared by a pre-determined number of related genes in at least one cluster; wherein the multiple common functions relate to each other in a biological pathway; and annotating the disease-specific gene network by associating the cluster with the biological pathway.

4. The method of any one of claims 1 to 3, further comprising (f) for at least one cluster of related genes, determining intracluster connectivity for each gene, and identifying genes having top about 5% - about 10% intracluster connectivity as the hub genes for the gene cluster; optionally wherein the method further comprises prioritizing hub genes by their respective intracluster connectivity. - 59 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 5. The method of claim 4, wherein intracluster connectivity of the gene is measured by a distance metrics based on gene embeddings; optionally wherein the distance metrics is selected from Cosine Similarity, Correlation Coefficient, Jaccard Similarity, Minkowski Distance, Manhattan Distance, and Euclidean Distance.

6. The method of claim 4 or 5, wherein intracluster connectivity of the gene is measured by the number of connections between the gene and other genes in the cluster.

7. The method of any one of claims 1 to 6, wherein the training dataset further comprises gene transcriptional data collected from a control biological sample free of the target disease; and wherein the method further comprises: (g) identifying one or more differentially expressed genes (DEGs) in at least one cluster of related genes; wherein for each DEG, expression profiles in the diseased biological sample and the control biological sample are different.

8. The method of any one of claims 4 to 7, further comprising (h) identifying the hub gene or the DEG as a disease-modifying genes for the target disease.

9. The method of claim 8, further comprising: (i) identifying one or more disease modifying genes identified in (h) as markers for a cell type and identifying the cell type as a disease-modifying cell type for the target disease; and / or identifying one or more disease modifying genes identified in (h) as having a common function or involved in a common cellular pathway, and identifying the common function or the common cellular pathway as a disease-modifying mechanism for the target disease.

10. The method of any one of claims 7 to 9, further comprising (j) identifying one or more cluster of related genes that is enriched with DEGs and / or pre-selected genes known to associate with the target disease; optionally the - 60 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 identifying is performed by (j-1) computing an enrichment score for each cluster based on the number of DEG and / or pre-selected genes and the total number of genes in the cluster and (j-2) identifying the clusters having statistically significantly higher enrichment scores as clusters enriched with the DEGs and / or pre-selected genes.

11. The method of claim 10, further comprising identifying the primary common function of the clusters identified in (j) as a disease-modifying mechanism of the target disease.

12. The method of any one of claims 7 to 11, further comprising visualizing DEGs in the disease-specific gene network with different colors representing upregulated and downregulated genes, respectively.

13. The method of claim 8 or 9 further comprising: (k) culturing in vitro a disease-modeling organoid carrying one or more mutations in at least one disease-modifying gene identified in (h).

14. The method of claim 13, wherein the mutations in the disease-modeling organoid is configured to mimic a change the disease-modifying gene’s function associated with the target disease.

15. The method of claim 13 or 14, wherein the culturing in (k) is performed by (k-1) providing an iPSC cell line carrying the mutation in at least one disease- modifying gene identified in (h); and (k-2) culturing the iPSC cell line under a suitable condition to induce differentiation of cells into the disease-modeling organoid.

16. The method of any one of claims 13 to 15, further comprising: (l) identifying a disease-modifying therapeutics for the target disease.

17. The method of claim 16, wherein the identifying in (l) is performed by: - 61 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 (l-1) contacting the disease-modeling organoid with a therapeutic candidate under a suitable condition for the disease-modeling midbrain organoid to react to the therapeutic candidate; (l-2) detecting a change in the disease-modeling midbrain organoid relevant to an improvement of a disease symptom; and (l-3) identifying the therapeutic candidate as a disease-modifying therapeutic for the target disease upon detecting the change in (l-2).

18. The method of any one of claims 1 to 17, wherein the disease-specific gene network is rendered in a three-dimensional (3D) representation.

19. The method of any one of claims 1 to 18, wherein the disease-specific gene network is rendered in a two-dimensional (2D) representation.

20. The method of any one of claims 1 to 19, wherein the diseased biological sample further comprises a diseased tissue isolated from the diseased subject.

21. The method of any one of claims 1 to 20, wherein the gene transcriptional data in the training set comprises single-cell sequencing data.

22. The method of claim 21, wherein the single-cell sequencing data in the training data set comprises bulk RNA-seq data, single cell RNA-seq data (scRNA-seq) or single nucleus RNA-seq data (snRNA-seq).

23. The method of any one of claims 1 to 22, wherein the foundation deep learning model is pre-trained with gene transcriptional data comprising single-cell sequencing data collected from a plurality of subjects of the same species as the diseased subject.

24. The method of claim 23, wherein the plurality of subjects do not suffer from the target disease. - 62 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 25. The method of claim 22, wherein the plurality of subjects are healthy subjects.

26. The method of claim 22, wherein the plurality of subjects is suffering or suspected as suffering from the target disease.

27. The method of any one of claims 23 to 26, wherein the foundation deep learning model is pre-trained with single-cell sequence data of more than 10 million, more than 30 million, more than 50 million or more than 100 million cells.

28. The method of any one of claim 1 to 27, wherein the diseased subject is a human.

29. The method of any one of claims 1 to 28, wherein the foundation deep learning model is scGPT, scFoundation, or Geneformer.

30. The method of any one of claims 1 to 28, wherein the foundation deep learning model is scGPT 31. The method of any one of claims 1 to 30, wherein the foundation deep learning model is implemented using computing platform selected from NVIDIA® CUDA®, NVIDIA® BioNeMo®, or Hugging Face®.

32. The method of any one of claims 1 to 31, wherein the foundation deep learning model is scGPT implemented using NVIDIA® CUDA®, scFoundation implemented using NVIDIA® CUDA®, Geneformer implemented using NVIDIA® BioNeMo®, or Geneformer implemented using PyTorch in Hugging Face®.

33. The method of any one of claims 1 to 32, wherein the foundation deep learning model is scGPT implemented using NVIDIA® CUDA®. - 63 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 34. The method of any one of claims 1 to 33, wherein the target disease affects an organ in the diseased subject, and wherein the cellular sphere in the diseased biological sample is an organoid capable of mimicking one or more in vivo functions of the organ.

35. The method of claim 34, wherein the organoid is derived in vitro from iPSC cells collected from the diseased subject.

36. The method of claim 35, wherein the organoid is isogenic.

37. The method of claim 35 or 36, wherein the iPSC cells or the organoid derived therefrom contain one or more gene mutations known to associate with the target disease.

38. The method of any one of claims 34 to 37, wherein the target disease is a neurological disease, and wherein the organoid is a midbrain organoid.

39. The method of claim 38, wherein the midbrain organoid comprises dopaminergic neurons, GABAergic neurons, glial cells, oligodendrocytes, glutaminergic neurons, astrocytes, radial glial cells, and neural progenitor cells.

40. The method of claim 39, wherein the dopaminergic neuron composition in the organoid is at least 15-40% of all neurons, and 5-15% of all cells.

41. The method of any one of claims 38 to 40, wherein the neurons in the midbrain organoid forms axons and dendrites.

42. The method of any one of claims 38 to 41, wherein the neurons in the midbrain organoid form synapses capable of transmitting neural signals.

43. The method of any one of claims 1 to 42, wherein the target disease is selected from Alzheimer’s Disease, Parkinson’s Disease, Huntington’s Disease, Epilepsy, Addiction, Neuropsychiatry, Rett Syndrome, and CDKL5 Deficiency Disorder. - 64 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 44. The method of any one of claims 1 to 43, wherein the target disease is known to associate with one or more gene mutations, and wherein the cellular sphere in the diseased biological sample is an organoid carrying the one or more gene mutations.

45. The method of any one of claims 1 to 43, wherein the target disease is Parkinson’s Disease, and wherein the cellular sphere in the diseased biological sample is a midbrain organoid carries a mutation in the GBA1 gene or the LRRK2 gene.

46. A disease-specific gene network generated using any one of the methods of claims 1 to 45.

47. The disease-specific gene network of claim 46, wherein the disease-specific gene network is digital.

48. A disease-modeling organoid, comprising a mutation in at least one disease-modifying gene identified by the method of claim 8.

49. A disease-modeling organoid generated by the method of any one of claims 13 to 15.

50. A system for modeling disease-associated mechanism of actions (MOA), comprising: (I) a disease-modeling organoid for a target disease, and (II) a computer system comprising one or more processors and a non-transitory computer readable storage medium including software stored thereon, wherein the software comprises executable instructions that, as a result of execution, causes the one or more processors of the computer system to: (a) load a foundation deep learning model pre-trained with a sufficient quantity of gene transcriptional data; (b) receive a training dataset comprising gene transcriptional data collected from the disease-modeling organoid; - 65 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 (c) process the training dataset using the deep learning model to generate a disease-specific gene network as output; and (d) infer a disease-associated MOA for the target disease, or a disease- modifying solution for the target disease, in either case, based on the generated output in (c).

51. The system of claim 50, wherein the disease-modeling organoid comprises cells carrying a gene mutation associated with the target disease.

52. The system of any one of claim 50 or 51, wherein the disease-modeling organoid is exposed to a candidate therapeutic agent under a suitable condition for the disease- modeling organoid to react to the candidate therapeutic agent.

53. A method for modeling a disease-associated mechanism of action (MOA) for a target disease, comprising: (a) providing a disease-modeling organoid for the target disease; (b) collecting gene transcriptional data from the disease-modeling organoid into a training dataset; (c) feeding the training data set to a foundation deep learning model pre-trained with a sufficient quantity of gene transcriptional data, thereby obtaining a fine-tuned model specific for the target disease; (d) extracting gene embeddings as output from the fine-tuned model; and (e) inferring the disease-associated MOA for the target disease, or a disease- modifying solution for the target disease, in either case, based on the generated output in (d).

54. The method of claim 53, further comprising (f) identifying a gene involved in the mechanism of action as a disease-modifying gene for the target disease.

55. The method of claim 54, further comprising - 66 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 (g) modifying the disease-modeling organoid by introducing a mutation in the disease-modifying gene into the disease-modeling organoid.

56. The method of claim 55, further comprising repeating steps (b) to (e) for one or more cycles.

57. The system or method of any one of claims 50 to 56, wherein the disease-associated MOA comprises (a) one or more DEGs that is upregulated or downregulated in the disease-modeling organoid as compared to a healthy control; (b) one or more cellular signaling pathways that is enhanced or inhibited in the disease-modeling organoid as compared to a healthy control; or (c) one or more cell types for which one or more features of the cell cycle are affected by the target disease.

58. The system or method of claim 57, wherein the one or more features of the cell cycle is selected from differentiation, development, proliferation and apoptosis of the cells.

59. The system or method of any one of claims 50 to 58, wherein the disease-modifying solution comprises (a) inhibiting the DEG that is upregulated in the disease-modeling organoid for treating, preventing or managing the target disease; (b) activating or supplementing the DEG that is downregulated in the disease- modeling organoid for treating, preventing or managing the target disease; (c) enhancing the cellular signaling pathway that is inhibited in the disease-modeling organoid for treating, preventing or managing the target disease; (d) inhibiting the cellular signaling pathway that is enhanced in the disease-modeling organoid for treating, preventing or managing the target disease; (e) detecting one or more biomarkers for the affected cell type for diagnosis of the target disease in the subject; - 67 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 (f) detecting one or more biomarkers of the affected cell type in a subject who is receiving a treatment of the target disease to evaluate the subject’s response to the treatment; and / or (g) detecting one or more biomarkers of the affected cell type in a subject for evaluating the likelihood of the subject to react to a treatment of the target disease.

60. The system or method of any one of claims 50 to 59, wherein the disease-modeling organoid is exposed to a candidate therapeutic agent under a suitable condition for the disease-modeling organoid to react to the candidate therapeutic agent.

61. The system or method of claim 60, wherein the disease-associated MOA comprises a change in (a) expression of one or more genes; (b) function of one or more genes; (c) abundance of one or more cell types; (d) activity of one or more cell types; (e) cell state transition of one or more cell types; (f) an organ function modeled by the disease-modeling organoid; or (g) one or more symptoms of the target disease; in reaction to the therapeutic agent.

62. The system or method of claim 61, wherein the change is associated with amelioration or elimination of the target disease, and wherein the method further infers the disease- modifying solution comprising using the candidate therapeutic agent for treating, preventing or managing the target disease.

63. The system or method of any one of claims 50 to 62, wherein the gene transcriptional data in the training set comprises single-cell sequencing data. - 68 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 64. The system or method of claim 63, wherein the single-cell sequencing data in the training data set comprises bulk RNA-seq data, single cell RNA-seq data (scRNA-seq) or single nucleus RNA-seq data (snRNA-seq).

65. The system or method of any one of claims 50 to 64, wherein the foundation deep learning model is pre-trained with gene transcriptional data comprising single-cell sequencing data collected from a plurality of subjects of the same species.

66. The system or method of claim 65, wherein the plurality of subjects do not suffer from the target disease.

67. The system or method of claim 65, wherein the plurality of subjects are healthy subjects.

68. The system or method of claim 65, wherein the plurality of subjects is suffering or suspected as suffering from the target disease.

69. The system or method of any one of claims 65 to 68, wherein the foundation deep learning model is pre-trained with single-cell sequence data of more than 10 million, more than 30 million, more than 50 million or more than 100 million cells.

70. The system or method of any one of claims 50 to 69, wherein the target disease is a disease of human.

71. The system or method of any one of claims 50 to 70, wherein the foundation deep learning model is scGPT, scFoundation, or Geneformer.

72. The system or method of any one of claims 50 to 70, wherein the foundation deep learning model is scGPT - 69 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 73. The system or method of any one of claims 50 to 72, wherein the foundation deep learning model is implemented using computing platform selected from NVIDIA® CUDA®, NVIDIA® BioNeMo®, or Hugging Face®.

74. The system or method of any one of claims 50 to 72, wherein the foundation deep learning model is scGPT implemented using NVIDIA® CUDA®, scFoundation implemented using NVIDIA® CUDA®, Geneformer implemented using NVIDIA® BioNeMo®, or Geneformer implemented using PyTorch in Hugging Face®.

75. The system or method of any one of claims 50 to 70, wherein the foundation deep learning model is scGPT implemented using NVIDIA® CUDA®.

76. The system or method of any one of claims 50 to 75, wherein the target disease affects an organ in a diseased subject, and wherein the diseased biological sample comprises an organoid capable of mimicking one or more in vivo functions of the organ.

77. The system or method of claim 76, wherein the organoid is derived in vitro from iPSC cells collected from the diseased subject.

78. The system or method of claim 77, wherein the organoid is isogenic.

79. The system or method of claim 77 or 78, wherein the iPSC cells or the organoid derived therefrom contain one or more gene mutations known to associate with the target disease.

80. The system or method of any one of claims 50 to 79, wherein the target disease is a neurological disease, and wherein the organoid is a midbrain organoid.

81. The system or method of claim 80, wherein the midbrain organoid comprises dopaminergic neurons, GABAergic neurons, glial cells, and neural progenitor cells. - 70 - NAI-5004565027v1Attorney Docket No.: 14876-002-228 82. The system or method of claim 81, wherein the dopaminergic neuron composition in the organoid is at least 15-40% of all neurons, and 5-15% of all cells.

83. The system or method of any one of claims 80 to 82, wherein the neurons in the midbrain organoid forms axons.

84. The system or method of any one of claims 80 to 83, wherein the neurons in the midbrain organoid form synapses capable of transmitting neural signals.

85. The system or method of any one of claims 50 to 84, wherein the target disease is selected from Alzheimer’s Disease, Rett Syndrome, Parkinson’s Disease, Huntington’s Disease, Epilepsy, Addiction, Neuropsychiatry, Rett Syndrome, and CDKL5 Deficiency Disorder.

86. The system or method of any one of claims 50 to 85, wherein the target disease is known to associate with one or more gene mutations, and wherein the diseased biological sample comprises an organoid carrying the one or more gene mutations.

87. The system or method of claim 86, wherein the target disease is Parkinson’s Disease, and wherein the cellular sphere in the diseased biological sample is a midbrain organoid carries a mutation in the GBA1 gene or the LRRK2 gene. - 71 - NAI-5004565027v1