Target, biomarker, and patient selection discovery methods using cell-type specific spatial proteomics and machine learning
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2026-03-12
AI Technical Summary
Current methods for biomarker and drug target discovery in central nervous system disorders face challenges due to limited cell-type specificity in neuronal cell generation, poor spatial resolution in protein labeling, and inefficiencies in analyzing sparse biological datasets, leading to high failure rates in clinical trials.
A platform technology using induced pluripotent stem cell differentiation, antibody-enzyme labeling, statistical data augmentation, and gradient boosting machine learning classifiers for spatial proteome profiling to identify ranked biomarkers and drug targets for neurodevelopmental and neurodegenerative disorders.
Enables the discovery of clinically relevant biomarkers and drug targets by generating disease-relevant neural cells, capturing cell-type specific protein signatures, and analyzing sparse datasets effectively, facilitating precision medicine for CNS disorders.
Smart Images

Figure US2025040175_12032026_PF_FP_ABST
Abstract
Description
TARGET, BIOMARKER, AND PATIENT SELECTION DISCOVERY METHODS USING CELL-TYPE SPECIFIC SPATIAL PROTEOMICS AND MACHINE LEARNINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 678,041 filed July 31, 2024. The entire contents of the above-identified application is hereby fully incorporated herein by reference.TECHNICAL FIELD
[0002] The subject matter disclosed herein is generally directed to the field of biomarker discovery and drug target identification, patient stratification, and patient selection for central nervous system disorders. The invention provides a platform technology combining induced pluripotent stem cell differentiation into specific neuronal subtypes, antibody-enzyme labeling of cell surface proteins, statistical data augmentation for sparse datasets, and gradient boosting machine learning classifiers. The platform enables identification of ranked biomarkers, drug targets, and patient selection for neurodevelopmental and neurodegenerative disorders including autism, Alzheimer's disease, Parkinson's disease, and schizophrenia through spatial proteome profiling and computational analysis.BACKGROUND
[0003] Central nervous system disorders (CNS), including neurodevelopmental conditions such as autism spectrum disorder, Rett syndrome, and Fragile X syndrome, as well as neurodegenerative diseases like Alzheimer's disease, Parkinson's disease, and amyotrophic lateral sclerosis, represent some of the most challenging areas in modern medicine. Drug development for CNS disorders has been notoriously difficult, with exceptionally high failure rates in clinical trials. This challenge is partially hampered by the lack of clear understanding of disease drivers, the poor translatability of molecular insights from animal models to humans, the absence of effective methods to interpret complex omics data from neural tissue systems, and to select the right patients to participate in clinical trials (patient selection).
[0004] Current approaches to biomarker and drug target discovery face several significant limitations. Existing methods for neuronal cell generation from pluripotent stem cells often relyon complex protocols with variable outcomes and limited cell-type specificity. Available protein labeling techniques, while capable of detecting target proteins in biological samples, typically lack the spatial resolution and cell-type specificity required for precise molecular profiling of neural circuits. Furthermore, conventional proteomic mapping approaches often provide broad protein detection without the ability to capture spatially localized protein networks that are critical for understanding neuronal function and dysfunction.
[0005] Machine learning applications in biological research have shown promise for analyzing complex datasets, but existing approaches often struggle with the sparse datasets typical of rare disease research and clinical studies with limited patient samples. Current methods for analyzing proteomic data frequently require large sample sizes and may not adequately address the inherent variability and missing data patterns common in biological studies. Additionally, conventional feature selection and classification algorithms may not be optimized for the specific challenges of identifying disease-relevant biomarkers from high-dimensional proteomic datasets.
[0006] There remains a need for integrated platform technologies that can effectively combine human-relevant cellular models, spatially-resolved molecular profiling, and advanced computational methods specifically designed for sparse biological datasets to accelerate the identification of novel drug targets and biomarkers for CNS disorders.
[0007] Citation or identification of any document in this application is not an admission that such a document is available as prior art to the present invention.SUMMARY
[0008] The present invention relates to methods for identifying ranked biomarkers for central nervous system disorders, spatial proteome profiling techniques, data augmentation methods for sparse biological datasets, neural differentiation protocols, patient stratification methods, drug screening applications, and related diagnostic and therapeutic applications.
[0009] In an example embodiment, there is provided a method for identifying ranked biomarkers for central nervous system disorders comprising: (a) generating neural cells from induced pluripotent stem cells (iPSCs) derived from a subject with a central nervous system disorder or healthy control subject by: (i) contacting the iPSCs with inhibitors of SMAD signaling and culturing the cells in 2D on transwell membrane (polyester) plates; (ii) culturing the cells toform neurospheres; and (iii) dissociating the neurospheres and replating the cells in 2D to further differentiate the cells into region-specific neural cells; (b) binding cell surface antigens by contacting the neural cells with an antibody-peroxidase conjugate and conducting a spatial proteome labeling reaction of proteins proximate to the antibody-peroxidase conjugate; (c) isolating the labeled proteins using affinity purification; (d) analyzing the isolated proteins by mass spectrometry to generate protein expression data; (e) computationally expanding the protein expression dataset by applying statistical data augmentation; and (f) identifying ranked biomarkers using a trained machine learning classifier.
[0010] In another embodiment, the statistical data augmentation comprises: (i) calculating standard deviations for each protein within patient and control groups; (ii) generating synthetic data points by applying multiplicative perturbations within ±4 standard deviations to existing values while preserving null values; and (iii) balancing synthetic sample numbers between patient and control groups.
[0011] In another embodiment, the central nervous system disorder is selected from the group consisting of autism spectrum disorder, schizophrenia, bipolar disease, epilepsy, rare neurodevelopmental disorders such as Rett Syndrome, CDKL5 deficiency disorder, Fragile X syndrome, SYNGAP1 disorder, CACNAlA-related disorders, CAMK2-related neurodevelopmental disorders and neurodegenerative disorders such as Alzheimer's disease, Parkinson's disease, amyotrophic lateral sclerosis, and frontotemporal dementia.
[0012] In another embodiment, the region-specific neurons are selected from the group consisting of forebrain excitatory neurons.
[0013] In another embodiment, the antibody -enzyme conjugate comprises horseradish peroxidase conjugated to an antibody that binds to a cell surface antigen expressed on layerspecific cortical neurons.
[0014] In another embodiment, the biotinylating substrate is a biotin-tyramide derivative and the catalyzing step is performed in the presence of hydrogen peroxide.
[0015] In another embodiment, the affinity purification is performed using avidin-conjugated magnetic beads.
[0016] In another embodiment, the mass spectrometry analysis is performed using nano liquid chromatography tandem mass spectrometry in data independent mode with a false discovery rate of 1%.
[0017] In another embodiment, the trained machine learning classifier is a gradient boosting classifier.
[0018] In another embodiment, the trained machine learning classifier identifies the ranked biomarkers using feature importance analysis.
[0019] In another embodiment, the feature importance analysis uses SHAP (Shapley Additive Explanations) values to rank protein contributions to disease classification.
[0020] In another embodiment, the method further comprises validating identified biomarkers by confirming their differential expression between patient and control samples.
[0021] In another embodiment, there is provided a method for cell-type specific spatial proteome profiling comprising: (a) contacting a neural cell sample with antibody-peroxidase conjugate that binds to a cell surface antigen; (b) performing spatial proteome labeling by contacting the cells with a biotinylating substrate comprising biotin-tyramide or biotin-tyramide derivatives and hydrogen peroxide to biotinylate proteins proximated to the antibody-enzyme conjugate; (c) quenching the labeling reaction; (d) isolating the biotinylated proteins; and (e) identifying the isolated proteins by liquid chromatography tandem mass spectrometry.
[0022] In another embodiment, the cell surface antigen is selected from the group consisting of GRHC3, GRIN3A, LYPD1, EPHA5, RXFP1, SLIT3, FGFR1, NRG1, CDH22, SEMA3E, and SEMA3D and the cells are layer V cortical neurons.
[0023] In another embodiment, the method further comprises comparing protein expression profiles between neurons derived from patients and control subjects.
[0024] In another embodiment, there is provided a statistical data augmentation method for sparse biological datasets comprising: (a) receiving a biological dataset containing protein expression measurements from patient and control groups; (b) segmenting the data into discrete patient and control cohorts; (c) calculating feature- wise standard deviations for each protein across non-null observations within each cohort; (d) generating synthetic samples by introducing multiplicative stochastic perturbations drawn from a uniform distribution bounded within ±4 standard deviations of original observed values; (e) preserving missing values in their originalform to maintain sparsity patterns; and (f) creating equal numbers of synthetic samples per class to address class imbalance.
[0025] In another embodiment, the biological dataset contains fewer than 1000 samples.
[0026] In another embodiment, the biological dataset comprises normalized mass spectrometry peak intensity values.
[0027] In another embodiment, the synthetic data generation maintains distributional properties of the original dataset.
[0028] In another embodiment, there is provided a method for generating cortical neurons from induced pluripotent stem cells comprising: (a) plating iPSCs as single cells at sparse density on Matrigel-coated transwell membranes (polyester membranes) in conditionally defined medium containing ROCK inhibitor; (b) removing the ROCK inhibitor and treating the cells with dual SMAD inhibitors in conditionally defined medium to induce neural fate; (c) culturing cells in conditionally defined medium without SMAD inhibitors; (d) treating cells with Trypsin-DNase I followed by DNase-CMF-PBS (Deoxyribonuclease in Calcium and Magnesium-Free Phosphate Buffered Saline) to form neurospheres; (e) culturing the neurospheres as floating spheres in conditionally defined medium; and (f) dissociating and replating the neurospheres on poly-D- Lysine and Laminin coated plates to generate adherent cortical neurons, with and without glial cells added.
[0029] In another embodiment, the dual SMAD inhibitors are LDN193189 and SB431542.
[0030] In another embodiment: (a) the dual SMAD inhibitors are applied for days 0-7; (b) the culturing in conditionally defined medium without SMAD inhibitors is performed from days 7- 20; (c) the cells are treated to form neurospheres at day 20; (d) the neurospheres are cultured until day 27 in conditionally defined medium without SMAD inhibitors; and (e) adherent cortical neurons are generated between days 37 and 70, with and without glial cells added.
[0031] In another embodiment, the conditionally defined medium comprises DMEM / F 12 with sodium bicarbonate, 0.5% BSA, 0.1 mM P-mercaptoethanol, 2 mM glutamate, 10 pM NEAA, lx N2 supplement, 1xB27 without retinoic acid, and Primocin.
[0032] In another embodiment, the cortical neurons express forebrain-specific markers and do not express hindbrain-specific markers at day 37.
[0033] In another embodiment, there is provided a method of patient stratification for central nervous system disorders comprising: (a) obtaining induced pluripotent stem cells (iPSCs) from multiple subjects with central nervous system disorders and control subjects; (b) generating neural cells from each subject's iPSCs by differentiating the iPSCs into forebrain excitatory neurons; (c) performing cell-type specific cell surface protein labeling on the neural cells from each subject using antibody-enzyme conjugates; (d) analyzing the labeled proteins by mass spectrometry to generate subject-specific spatial protein expression profiles; (e) clustering the subjects based on their spatial protein expression profiles using Principal Component Analysis and Hierarchical Clustering; (f) linking the protein expression profiles to clinical symptom severity (mild, moderate, severe); and (g) identifying patient subgroups with distinct proteomic profiles for targeted therapeutic intervention.
[0034] In another embodiment, the clustering correlates with symptom severity levels among the subjects.
[0035] In another embodiment, subjects with mild symptom severity cluster more closely to control subjects than subjects with moderate symptom severity.
[0036] In another embodiment, the method further comprises correlating the patient subgroups with genetic backgrounds of the subjects.
[0037] In another embodiment, there is provided a method for patient stratification for central nervous system disorders comprising: (a) generating neural cells from induced pluripotent stem cells derived from a patient sample; (b) performing cell-type specific cell surface protein labeling on the neural cells using antibody-enzyme conjugates; (c) analyzing the labeled proteins by mass spectrometry to generate a patient-specific protein expression profile; (d) comparing the patientspecific protein expression profile to reference protein expression profiles using computational similarity analysis, wherein each reference protein expression profile is from a previously established patient subgroup, wherein each subgroup is characterized by distinct proteomic signatures and clinical phenotypes; and (e) assigning the patient to a patient subgroup based on the highest similarity match for targeted therapeutic intervention.
[0038] In another embodiment, there is provided a method of screening drug candidates comprising: (a) identifying ranked biomarkers using the method of claim 1; (b) contacting neural cells with one or more test compounds; (c) measuring expression levels of the ranked biomarkersin the presence of the test compounds; and (d) identifying drug candidates based on modulation of the biomarker expression levels compared to control conditions.
[0039] In another embodiment, the test compounds are screened against a panel of the topranked biomarkers.
[0040] In another embodiment, the drug candidates are identified based on normalization of biomarker expression levels toward control levels.
[0041] In another embodiment, the method further comprises: (e) measuring a functional phenotype of the neural cells in the presence of the test compounds, wherein the functional phenotype is selected from the group consisting of neuronal electrical activity, synaptic function, neurite outgrowth, calcium signaling, and cell viability; and (f) identifying drug candidates based on improvement of the functional phenotype in addition to modulation of biomarker expression levels.
[0042] In another embodiment, the method further comprises comparing spatial protein expression of one region of a neural cell from a first diseased subject with that of a second diseased subject to identify variability in protein expression between diseased subjects.
[0043] In another embodiment, the method further comprises comparing spatial protein expression of one region of a neural cell from a diseased subject with that of a healthy subject to identify variability in protein expression between diseased and healthy subjects.
[0044] In another embodiment, the method further comprises comparing spatial protein expression of one region of a neural cell from a first healthy subject with that of a second healthy subject to identify baseline variability in protein expression within the healthy population.
[0045] In another embodiment, the drug candidates are identified based on normalizing protein expression variability identified through multiple subject comparisons toward healthy baseline levels.
[0046] In another embodiment, the drug candidates target at least one protein selected from the proteins in Table 1.
[0047] In another embodiment, there is provided a kit for detecting biomarkers for central nervous system disorders comprising reagents for detecting one or more biomarkers identified by the method of claim 1.
[0048] In another embodiment, the reagents comprise antibodies specific for cell surface proteins selected from the group consisting of NCAM1, SEMA7A, GRIK3, NRP1, SLITRK1, SEMA3E, SEMA3D, and EPHA5.
[0049] In another embodiment, the kit comprises reagents for detecting a panel of at least 3 biomarkers.
[0050] In another embodiment, there is provided a method of treating a subject with a central nervous system disorder comprising administering to the subject a therapeutically effective amount of a compound that modulates a biomarker identified by the method of claim 1.
[0051] In another embodiment, the central nervous system disorder is autism spectrum disorder and the compound modulates GRIK3 activity.
[0052] In another embodiment, there is provided a method of diagnosing a central nervous system disorder comprising: (a) obtaining a biological sample from a subject; (b) measuring expression levels of biomarkers identified by the method of claim 1; and (c) comparing the expression levels to reference profiles to determine disease status.
[0053] In another embodiment, the biomarkers are cell surface proteins and the measuring is performed using immunoassay techniques.
[0054] These and other aspects, objects, features, and advantages of the example embodiments will become apparent to those having ordinary skill in the art upon consideration of the following detailed description of example embodiments.BRIEF DESCRIPTION OF THE DRAWINGS
[0055] An understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention may be utilized, and the accompanying drawings of which:
[0056] FIG. 1 - is a schematic overview of the six-step platform technology for identifying drug targets and biomarkers, and perform patient stratification in central nervous system disorders. Step 1 shows the generation of neural cells from patient-derived iPSCs, which are differentiated into forebrain excitatory neurons using optimized protocols involving sparse density plating (3,000 cells / well), dual SMAD inhibition with LDN193189 and SB431542 in 2D, and novel neurosphere formation using Trypsin-DNase I treatment followed by DNase-CMF-PBS processing. Steps 2and 3 illustrate cell-type specific spatial molecular profiling, where neurons are labeled with antibody-enzyme conjugates (such as antibodies against GRIK3, GRIN3A, LYPD1, EPHA5, RXFP1, SLIT3, FGFR1, NRG1, CDH22, SEMA3E, SEMA3D) that bind to specific cell surface antigens, followed by enzymatic biotinylation using cell-impermeable biotin-tyramide-derivatives and hydrogen peroxide to label proteins within spatial vicinity of the antibody -enzyme conjugate. Step 4 involves affinity purification of biotinylated proteins using avidin-coated magnetic beads and mass spectrometry analysis using nano liquid chromatography tandem mass spectrometry in data independent mode to generate protein expression data. Step 5 demonstrates statistical data augmentation to computationally expand sparse biological datasets by applying multiplicative stochastic perturbations within ±4 standard deviations while preserving distributional properties and addressing class imbalance. Step 6 shows machine learning classification using gradient boosting classifiers with SHAP -based feature importance analysis to identify and rank biomarkers, as candidate targets. The platform integrates patient symptom severity data with molecular signatures to enable biomarker discovery, patient stratification, and drug target identification for neurodevel opmental and neurodegenerative disorders.
[0057] FIG. 2A-2B - are bar charts showing signal intensity measurements for brain region protein markers detected in iPSC-derived neuronal cultures at day 37 by mass spectrometry. Panel 2A demonstrates successful forebrain excitatory neuron differentiation by showing high expression levels of forebrain-specific markers PAX6, FOXG1, and layer-specific markers CTIP2 and CUX1, confirming proper forebrain identity, while showing no detection of hindbrain markers GBX2, H0XA2, and ATOH1, indicating successful exclusion of hindbrain cell types. Panel 2B shows validation of expression of proteins encoded by genes that have been identified as high-risk autism and neurodevelopmental disorder genes. Together, these panels validate the region-specific neural differentiation and robust expression of proteins implicated in central nervous system disorders, demonstrating the platform's capability to generate patterned neurons with appropriate molecular identities for subsequent target and biomarker discovery applications.
[0058] FIG. 3 - is a heat map visualization showing comparative protein expression levels between affinity purified samples and input control samples from an iPSC-derived forebrain culture at day 55. The grayscale intensity represents normalized protein abundance levels, with darker regions indicating higher protein concentrations. The hierarchical clustering patterndemonstrates the effectiveness of the affinity purification method in selectively enriching specific proteins compared to input samples, validating the cell-type specific protein labeling approach.
[0059] FIG. 4A-4C - are scatter plots demonstrating the reproducibility of the spatial proteome profiling method through protein expression correlations. Panel 4A shows correlation between technical replicates (R2= 0.97), demonstrating technical precision. Panel 4B displays correlation between biological replicates from different iPSC clones of the same individual (R2= 0.77), validating reproducibility within subjects. Panel 4C presents correlation between samples from different individuals (R2= 0.82), indicating consistent protein signatures across different patients / individuals and supporting the platform's utility for biomarker identification.
[0060] FIG. 5 - is a scatter plot demonstrating the specificity of the spatial proteome labeling method by showing the absence of correlation between a sample processed for cell surface labeling and a negative control sample without biotinylating substrate. The lack of correlation validates that protein labeling is specifically dependent on the enzymatic reaction and biotinylating substrate, confirming that detected proteins result from genuine spatial proteome labeling rather than nonspecific binding interactions.
[0061] FIG. 6 - is a volcano plot visualization of differentially expressed proteins identified between patient and control samples. The plot displays statistical significance (p-values) on the y- axis against magnitude of change (Log2 fold change) on the x-axis for each detected protein. Individual points represent proteins, with position indicating both the degree of differential expression and statistical confidence. This visualization identifies candidate biomarkers by highlighting proteins with substantial fold changes and high statistical significance.
[0062] FIG. 7 - is a hierarchical clustering analysis of protein expression profiles demonstrating patient stratification capability. The dendrogram reveals distinct groupings based on the distinct protein signatures detected for patient and control subjects, with patients exhibiting mild symptoms clustering most closely to control subjects and separately from patients with moderate symptom severity. The clustering pattern demonstrates that spatial proteome profiling can reliably distinguish patient subgroups based on molecular signatures that correlate with clinical phenotype severity.
[0063] FIG. 8 - is a Principal Component Analysis plot demonstrating the discriminatory power of the spatial proteome profiling method for patient stratification. The visualization showsdistinct clustering patterns where negative controls (no surface labeling) cluster separately from other samples, patients with mild symptoms cluster closest to controls, and two genetically defined neurodevelopmental conditions can be distinguished based on surface proteomics profiles, validating the method's ability to capture disease-specific molecular signatures.
[0064] FIG. 9 - illustrates the statistical data augmentation method for expanding sparse biological datasets. The figure demonstrates controlled expansion while preserving distributional properties of original protein expression data. The visualization shows how synthetic data points are generated through multiplicative stochastic perturbations within ±4 standard deviations of original values, enabling effective machine learning analysis of limited biological samples typical in rare disease research.
[0065] FIG. 10 - demonstrates validation of the machine learning-based biomarker identification method using published cerebrospinal fluid proteomics data from Rett syndrome patients and controls (Zlatic et al. 2022 iScience, 6 individuals total). The dataset containing 192 differentially expressed proteins was synthetically expanded 10-fold using the statistical data augmentation method described in Step 5, increasing sample size from 6 to 60. Missing values were imputed using a constant value (-1) and the dataset was normalized using StandardScaler. A Gradient Boosting Classifier was trained on the augmented dataset, and feature importance was assessed using SHAP values computed with TreeExplainer. The SHAP values analysis successfully identified MECP2, the known driver gene for Rett syndrome, among the top 35 proteins out of the 192 differentially expressed proteins, validating the effectiveness of the statistical data augmentation method combined with gradient boosting classification and SHAP- based feature importance analysis for discovering biologically relevant biomarkers from sparse biological datasets without prior knowledge of disease mechanisms.
[0066] FIG. 11 - demonstrates the successful identification and validation of key biomarkers through the disclosed machine learning-driven approach. The data shows statistically significant differential expression patterns between control and patient samples across six top-ranked biomarkers: SEMA7A, GRIK3, NRP1, SLITRK1, EPHA5, and NCAM1. These proteins, identified through SHAP-based feature importance analysis, exhibit distinct expression profiles that enable robust patient-control classification. NCAM1 shows the most pronounced differential expression, consistent with its #1 SHAP ranking and known role in synaptic plasticity and autism-related pathways. The consistent directional changes across multiple biomarkers validate the method's ability to capture disease-specific molecular signatures. This differential expression pattern forms the foundation for diagnostic applications, drug screening assays, and patient stratification methods claimed in the patent. The robust signal-to-noise ratio observed across these biomarkers supports their utility as reliable diagnostic targets and demonstrates the commercial viability of the disclosed biomarker detection kits and therapeutic screening platforms for central nervous system disorders.
[0067] The figures herein are for illustrative purposes only and are not necessarily drawn to scale.DETAILED DESCRIPTION OF THE EXAMPLE EMBODIMENTSOVERVIEW
[0068] Embodiments disclosed herein collectively provide a method for identifying ranked biomarkers for central nervous system disorders that addresses critical limitations in current biomarker discovery approaches by combining patient-derived cellular models, spatial proteomics, and machine learning designed for sparse biological datasets. Central nervous system disorders, including neurodevelopmental conditions such as autism spectrum disorder and rare genetic syndromes, as well as neurodegenerative diseases like Alzheimer's and Parkinson's disease, present significant challenges for biomarker identification due to the cellular heterogeneity of neural tissue, limited patient sample availability, and the inherent complexity of neural protein networks. Existing methods typically rely on bulk tissue analysis that obscures cell-type specific molecular changes, utilize animal models with questionable translational relevance to human pathophysiology, or employ computational approaches not optimized for the high-dimensional, sparse datasets characteristic of rare disease research. The disclosed method overcomes these limitations through an integrated platform that generates disease-relevant neural cells from patient- derived induced pluripotent stem cells, performs spatially-resolved biotinylation labeling to capture cell-type specific protein signatures, and applies novel statistical data augmentation techniques coupled with interpretable machine learning to identify and rank biomarkers even from limited patient cohorts. This approach enables the discovery of clinically relevant biomarkers thatreflect actual disease for clinical translation, ultimately facilitating precision medicine approaches for CNS disorder diagnosis, patient stratification, and therapeutic development.A METHOD FOR IDENTIFYING RANKED BIOMARKERS FOR CNS DISORDERS
[0069] In an embodiment, a method for identifying ranked biomarkers for CNS disorders comprises generating neural cells from induced pluripotent stem cells (iPSCs) derived from a subject with a central nervous system disorder or healthy control subject, binding cell surface antigens of the neural cells with an antibody-peroxidase conjugate and conducting a spatial proteome labeling reaction that labels proteins proximate to the antibody-peroxidase conjugate, isolating the labeled proteins, identifying the labeled proteins by mass spectrometry to generate protein expression data, computationally expanding the protein expression data set, and identifying ranked biomarkers using a trained machine learning classifier.Generating Neural Cells
[0070] The method comprises generating neural cells from induced pluripotent stem cells (iPSCs) derived from subjects with central nervous system disorders or healthy control subjects. As used herein, "neural cells" encompasses any cells of neural origin, including but not limited to neurons, astrocytes, oligodendrocytes, and neural progenitor cells. The neural cells may be derived from primary neural tissue, which consists of freshly isolated or cultured neural cells from brain tissue. Alternatively, iPSC-derived neural cells can be used, consisting of neurons generated from induced pluripotent stem cells using the methods described above or alternative differentiation protocols. A "neural cell sample" is a biological sample comprising neural cells as defined above.
[0071] As used herein, "induced pluripotent stem cells" or "iPSCs" refer to pluripotent stem cells that can be generated directly from adult somatic cells through the introduction of specific transcription factors, as first described by Takahashi and Yamanaka. iPSCs possess the ability to differentiate into any cell type of the three primary germ layers: ectoderm, mesoderm, and endoderm, including neural cell types.
[0072] In an embodiment, the specific combination of dual SMAD inhibition first in 2D on transwell membranes followed by neurosphere formation provides improved yield and regional specificity compared to conventional approaches.
[0073] In an embodiment, the neural differentiation process comprises the steps of initial neural induction, neurosphere formation, and final differentiation into regionally specific neural cells. This method may also be used independently of the method for identifying ranked biomarkers. Neural fate induction refers to the process by which pluripotent stem cells commit to neural lineage specification and begin expressing neural progenitor markers while losing pluripotency markers. This process involves the activation of neural transcription factors such as PAX6, SOX1, and NESTIN, while downregulating pluripotency factors including OCT4, NANOG, and SOX2. The dual SMAD inhibition treatment is applied for days 0-7 to achieve efficient neural induction.Initial Neural Induction
[0074] The iPSCs are treated with dual inhibitors of SMAD signaling while cultured in 2D on transwell plates. As used herein, "SMAD signaling" refers to the intracellular signaling pathway mediated by Small Mothers Against Decapentaplegic (SMAD) proteins, which are key mediators of transforming growth factor-0 (TGF-p) and bone morphogenetic protein (BMP) signaling pathways. SMAD signaling inhibition is crucial for neural induction as it blocks the inhibitory effects of BMP and TGF-0 pathways on neural fate specification.
[0075] The dual SMAD inhibitors suitable for use in the present method include: LDN193189 is a selective BMP signaling inhibitor that targets ALK2, ALK3, and ALK6 receptors with IC50 values in the nanomolar range. LDN193189 is typically used at concentrations ranging from 100 nM to 500 nM, preferably 100 nM. SB431542 is a selective inhibitor of the TGF-0 type I receptors ALK4, ALK5, and ALK7, with an IC50 of 94 nM for ALK5. SB431542 is typically used at concentrations ranging from 5 pM to 20 pM, preferably 10 pM.
[0076] In an embodiment, dual SMAD inhibition is employed using both LDN193189 at concentrations ranging from 50 nM to 200 nM, preferably 100 nM, and SB431542 at concentrations ranging from 5 pM to 15 pM, preferably 10 pM.
[0077] Transwell plates refer to cell culture devices containing permeable membrane inserts that allow for controlled culture conditions while enabling molecular exchange between compartments. The use of transwell membranes for neural differentiation offers severaladvantages, including reduced fluctuations in gene expression and improved cell health during differentiation.
[0078] The iPSCs are plated as single cells at a sparse density. Sparse density as used herein refers to plating densities significantly lower than conventional neural differentiation protocols, typically ranging from 1,000 to 10,000 cells per well of a 6-well plate, preferably 3,000 cells per well. This sparse plating density contrasts with typical high-density protocols that use 50,000- 100,000 cells per well.
[0079] As used herein, "conditionally defined medium" refers to a chemically defined cell culture medium containing known components without undefined additives such as serum. In an embodiment the conditionally defined medium comprises: DMEM / F12 with sodium bicarbonate, 0.5% BSA, 0.1 mM P-mercaptoethanol, 2 mM glutamate, 10 pM NEAA, l x N2 supplement, lxB27 without retinoic acid, and Primocin.
[0080] The cells may be initially treated with a ROCK inhibitor, such as Y-27632 at 10 pM concentration, to prevent cell death associated with single cell plating.
[0081] In an embodiment, SMAD inhibition treatment is typically maintained for 1, 2, 3, 4, 5, 6, or 7 days to induce neural fate commitment and establish forebrain identity.Neurosphere Formation
[0082] Following the initial neural induction period, the cells are cultured to form neurospheres. As used herein, "neurospheres" refer to three-dimensional cellular aggregates of neural stem cells and neural progenitor cells that form when cultured in suspension without adhesive substrates. In an embodiment, the specific combination of neurosphere formation with prior SMAD inhibition and subsequent Trypsin-DNase I treatment provides unique advantages in eliminating non-neuronal cells and generating homogeneously sized spheres.
[0083] In an embodiment, the neurosphere formation process involves briefly (4 min) treating the cells between days 20-27 of differentiation with Trypsin-DNase I, followed by incubation in DNase-CMF-PBS. Cells are then triturated in DNase-CMF-PBS with a glass pipette until spheres form. The Trypsin-DNase I treatment results in the formation of homogeneously sized neurospheres, eliminating non-neuronal cells.Final Differentiation to Region-Specific Neural Cells
[0084] The neurospheres are dissociated and replated in 2D culture to further differentiate into region and subtype-specific neural cells. As used herein, "region-specific neural cells" refer to neurons that exhibit molecular markers and functional properties characteristic of specific anatomical regions of the central nervous system, such as forebrain.
[0085] In an embodiment, cells may be dissociated and replated on poly-D-lysine and laminin- coated surfaces. The combination of poly-D-lysine and laminin coating provides optimal substrate conditions for final neural differentiation and the formation of adherent neuronal cultures with extensive neurite networks.
[0086] In another embodiment, cells may be dissociated and replated on glial feeder layers to promote neural differentiation and synaptic maturation further. The co-culture with glial cells, including astrocytes and / or oligodendrocytes, provides essential neurotrophic support, facilitates proper synaptic development, and enhances the functional maturation of the differentiated neurons through paracrine signaling and direct cell-cell interactions.Regional Identity Validation And Marker Expression
[0087] Forebrain-specific markers include transcription factors specifically expressed in forebrain-derived neural cells, including but not limited to: FOXG1 (Forkhead Box Gl), PAX6 (Paired Box 6), and specific layers of the cortex, including but not limited to: CUX1 (Cut Like Homeobox 1), and TBR1 (T-Box Brain Transcription Factor 1). Hindbrain-specific markers include transcription factors specifically expressed in hindbrain-derived neural cells, including but not limited to: GBX2 (Gastrulation Brain Homeobox 2), H0XA2 (Homeobox A2), and ATOH1 (Atonal Homolog 1). The expression of forebrain-specific markers and absence of hindbrainspecific markers at day 37 and beyond confirms the regional specificity and purity of the cortical neuron cultures generated by the disclosed method.
[0088] Region-specific neurons that can be generated using this method include forebrain excitatory neurons, which are glutamatergic projection neurons typically found in cortical layers Will and V / VI, characterized by expression of markers such as TBR1, CTIP2, and CUX1.
[0089] The disclosed forebrain differentiation method offers significant advantages over conventional published protocols through its innovative combination of transwell membraneculture and neurosphere formation. This approach achieves enhanced regional specificity without requiring WNT signaling inhibition, as the sequential culture of cells on transwell membranes followed by the neurosphere step creates homogeneous populations with improved cell viability. The result is highly pure forebrain neuronal cultures with minimal contamination from other brain regions or cell types.
[0090] The method's standardized substrate conditions and timing parameters contribute to its reproducibility, ensuring consistent outcomes across different experimental conditions and laboratories. This reliability is particularly valuable for research applications where experimental variability can compromise data interpretation and cross-study comparisons.
[0091] Furthermore, the forebrain neurons generated through this protocol demonstrate functional competence characterized by appropriate marker expression, including proteins directly implicated in central nervous system disorders. This functional integrity makes the neurons particularly well-suited for downstream applications, including the biomarker discovery and drug target identification methods described throughout this disclosure, thereby establishing a seamless integration between cell generation and therapeutic discovery platforms.Spatial Proteome Profiling
[0092] In an embodiment, cell type-specific spatial proteome labeling comprises binding cellsurface antigens of the neural cells with a labeled antibody, such as an antibody-horseradish peroxidase conjugate, that binds to a cell surface antigen, performing spatial biotinylation by contacting the cells with a biotinylating substrate comprising biotin-tyramide or biotin-tyramide derivatives and hydrogen peroxide to biotinylate proteins proximate to the antibody-enzyme conjugate.
[0093] In an embodiment, the neural cells are iPSC-derived neurons generated using the region-specific differentiation protocol described above, as these preserve patient-specific genetic backgrounds while providing sufficient cell numbers and experimental control. This method of spatial proteome profiling may also be used independently of the method of identifying ranked biomarkers.
[0094] The antibody-peroxidase conjugate may be a horseradish peroxidase (HRP) conjugate, though the method may also encompass other peroxidases and oxidative enzymes. HRP is a heme-containing enzyme with a molecular weight of approximately 40 kDa that catalyzes the oxidation of various organic substrates in the presence of hydrogen peroxide.
[0095] Cell surface antigens suitable for targeting include any protein expressed on the external surface of neural cell membranes that provides cell-type specificity within heterogeneous neural cultures. In an embodiment, cell surface antigens are selected from the group consisting of GRIK3, GRIN3A, LYPD1, EPHA5, RXFP1, SLIT3, FGFR1, NRG1, CDH22, SEMA3E, and SEMA3D, which are specifically expressed on layer V cortical neurons.Spatial Proteome Profiling Reaction
[0096] In an embodiment, spatial proteome profiling may be performed by contacting the cells with a biotinylating substrate comprising biotin-tyramide or biotin-tyramide derivatives and hydrogen peroxide to biotinylate proteins proximate to the antibody-enzyme conjugate. In an embodiment, proximate means within 10-300 nm. In an embodiment, a protein is proximate to the antibody-enzyme conjugate if it is within 200 nm.
[0097] Spatial proteome profiling represents a significant advance over traditional biochemical fractionation methods by providing spatial resolution at the subcellular level.
[0098] Biotinylating substrates encompass biotin-tyramide and biotin-tyramide derivatives. Biotin-tyramide is the prototypical substrate, consisting of biotin linked to a tyramide moiety through an amide bond. When activated by HRP in the presence of H2O2, biotin-tyramide generates biotin-phenoxyl radicals that covalently modify electron-rich amino acid residues.
[0099] The spatial radius of approximately 200 nanometers reflects the limited diffusion distance of biotin-phenoxyl radicals before they react with nearby proteins or decay.
[0100] Hydrogen peroxide serves as the terminal electron acceptor in the peroxidase reaction and is typically used at concentrations of 0.5-2 mM.
[0101] The labeling reaction is typically performed for 1-10 minutes at room temperature or 37°C, with shorter times providing more spatial restriction and longer times increasing labeling efficiency.Reaction Quenching
[0102] Following the spatial biotinylation reaction, the labeling reaction is quenched. Quenching refers to the rapid termination of the enzymatic reaction to prevent continued substrate activation and non-specific labeling.
[0103] Effective quenching strategies include antioxidant cocktails consisting of solutions containing sodium azide (10 mM), sodium ascorbate (10 mM), and Trolox (5 mM) that scavenge reactive radicals and inhibit peroxidase activity.Protein Isolation and Purification
[0104] Following the spatial biotinylation reaction, the biotinylated protein may be isolated. The isolation step may comprise cell lysis followed by binding of the biotinylated proteins using an avidin-coated substrate, such as avidin-coated magnetic beads.
[0105] Cell lysis may involve disrupting cellular membranes to release intracellular proteins while preserving protein-biotin conjugates. Suitable lysis buffers include for example RIPA buffer, which is a radio-immunoprecipitation assay buffer containing non-ionic and ionic detergents for complete protein solubilization.
[0106] Avidin-coated substrates provide the affinity matrix for capturing biotinylated proteins. Avidin is a tetrameric glycoprotein from egg whites with exceptionally high affinity for biotin (Kd ~ 1015M). The avidin-biotin interaction is one of the strongest non-covalent interactions known in biology, making it ideal for affinity purification applications.
[0107] Magnetic beads offer rapid separation through magnetic fields that enable quick isolation without centrifugation.Mass Spectroscopy Identification
[0108] The isolated proteins are then identified by liquid chromatography tandem mass spectrometry. Liquid chromatography tandem mass spectrometry (LC-MS / MS) combines the separation power of liquid chromatography with the identification capabilities of tandem mass spectrometry. This approach is suitable for proteomics analysis due to its sensitivity, specificity, and throughput capabilities.
[0109] In an embodiment, nano-scale liquid chromatography is employed to enhance sensitivity through reduced flow rates and smaller column diameters. Data-independent acquisition (DIA) modes provide comprehensive and reproducible protein quantification compared to traditional data-dependent approaches.Comparative Protein Expression Analysis
[0110] The method may further comprise comparing protein expression profiles between neurons derived from patients and control subjects. This comparative analysis enables the identification of disease-associated protein changes and the validation of biomarker candidates.[0U1] Protein expression profiles refer to the comprehensive set of proteins identified and quantified in each sample, typically represented as intensity values or spectral counts for each detected protein. Comparative analysis involves statistical methods to identify proteins that show significant differences between experimental groups.DATA AUGMENTATION OF SPARSE DATA SETS
[0112] While data augmentation techniques are known in machine learning applications, their application to biological datasets presents unique challenges that are not addressed by conventional approaches. Mass spectrometry proteomics data exhibits distinct characteristics, including high dimensionality, sparse representation, missing values, and class imbalance that require specialized augmentation strategies.
[0113] In an embodiment statistical augmentation of a protein expression data set may comprise segmenting the data into discrete patient and control cohorts, calculating feature-wise standard deviations for each protein across non-null observations within each cohort, generating synthetic samples by introducing multiplicative stochastic perturbations, preserving missing values in their original form to maintain sparsity patterns, and creating equal numbers of synthetic samples per class to address any class imbalance. This method maintains both statistical validity and biological relevance while addressing the specific characteristics of sparse proteomics datasets.
[0114] The method of statistical data augmentation may also be used independently of the process for identifying ranked biomarkers for CNS disorders. In that regard, the biological datasetmay be the protein expression data from the mass spectroscopy analysis described above or may also be any collection of molecular measurements derived from biological samples, including mass spectrometry proteomics data consisting of protein intensity values, spectral counts, or label-free quantification measurements.
[0115] Sparse datasets are common in rare disease studies involving limited patient populations for uncommon conditions. Clinical proteomics studies exhibit sparsity due to the high cost and complexity of mass spectrometry analysis.
[0116] In an embodiment, a sparse data set may be a data set having 1,000 or fewer samples. The 1000-sample threshold reflects a practical boundary where conventional machine learning approaches begin to show reduced performance due to insufficient training data, particularly for high-dimensional datasets typical in proteomics.Data Segmentation
[0117] The method comprises segmenting the data into discrete patient and control cohorts. Data segmentation involves partitioning the dataset based on predefined group assignments to enable group-specific statistical calculations.Feature-Wise Standard Deviation Calculation
[0118] The method continues with calculating feature-wise standard deviations for each protein across non-null observations within each cohort. As used herein, "feature-wise" refers to calculations performed independently for each measured variable (protein) in the dataset.
[0119] Non-null observations exclude missing values that are common in mass spectrometry data due to detection limits, ion suppression, or stochastic sampling effects.
[0120] The method proceeds with generating synthetic samples by introducing multiplicative stochastic perturbations drawn from a uniform distribution bounded within ±4 standard deviations of original observed values. Multiplicative perturbations are applied rather than additive perturbations to preserve the proportional relationships inherent in biological data. The use of multiplicative rather than additive perturbations is particularly important for mass spectrometry data, which typically exhibits log-normal distributions and heteroscedastic variance (variance that scales with mean intensity).
[0121] The ±4 standard deviation boundary may be used to ensure that synthetic values remain within a biologically plausible range while providing sufficient diversity for machine learning applications.Missing Value Preservation
[0122] The method includes preserving missing values in their original form to maintain sparsity patterns as recited in the claims. This preservation strategy represents a fundamental departure from conventional data augmentation approaches that typically require complete datasets or employ imputation methods prior to synthetic data generation. Traditional approaches often mask or eliminate the inherent sparsity structure of biological data, thereby losing critical information embedded within the pattern of missing observations.
[0123] Sparsity patterns in biological data, particularly mass spectrometry proteomics datasets, carry essential information that reflects multiple biological and technical phenomena. These patterns provide insights into detection limits where proteins fall below instrumental detection thresholds due to low abundance, ion suppression effects, or stochastic sampling limitations inherent to data-dependent acquisition methods. The missing data structure also indicates biological absence, representing proteins that are genuinely not expressed in specific cell types, developmental stages, or disease conditions, thereby providing meaningful negative information about cellular states.
[0124] Additionally, sparsity patterns reflect technical sampling effects characteristic of mass spectrometry analysis, where the stochastic nature of peptide selection for fragmentation results in variable protein detection across samples. The preservation of these patterns maintains the natural heterogeneity observed in biological systems, where protein expression exhibits inherent variability both within and between individuals, particularly in disease contexts where cellular dysfunction may lead to altered protein expression profiles.
[0125] The disclosed preservation strategy involves several key technical components. Null value flagging explicitly marks missing measurements to distinguish them from true zero values, which represent detected but quantifiably absent proteins. Pattern maintenance ensures that proteins exhibiting high missingness rates retain their sparse representation, preserving the natural detection frequency distributions observed in the original dataset. Group-specific patternpreservation maintains different missingness patterns between patient and control groups when such differences exist naturally, as these differential sparsity patterns may themselves constitute disease-relevant biomarkers.
[0126] This approach provides significant advantages over alternative strategies. Imputation before augmentation can introduce systematic biases by artificially creating protein expression values where none were detected, potentially leading to false positive biomarker identifications. Zero-filling strategies similarly distort the data structure by converting missing observations to artificial zero values, which may not accurately represent the biological state. Complete case analysis, which excludes proteins with any missing values, results in substantial loss of potentially informative features and reduced statistical power for biomarker discovery.
[0127] The preservation of missing values maintains the discriminatory power of machine learning models by ensuring that the absence of detection contributes appropriately to classification decisions. In many biological contexts, the pattern of what is not detected can be as informative as what is detected, particularly in disease states where protein loss or cellular dysfunction may result in characteristic absence patterns. This approach ensures that synthetic data generation does not artificially inflate apparent protein coverage or create unrealistic datasets that poorly represent the constraints and characteristics of actual biological measurements.
[0128] Furthermore, the method's handling of missing values enables more accurate assessment of biomarker reliability and clinical utility. Biomarkers that are consistently detected across samples are inherently more suitable for clinical implementation than those with high missingness rates, as the latter may present practical challenges for diagnostic assay development. By preserving the natural detection patterns, the disclosed method enables the identification of robust biomarkers that are likely to translate successfully to clinical applications.Class Balance Creation
[0129] Creating equal numbers of synthetic samples per class may be used to address class imbalance. Class imbalance is a common challenge in clinical studies where patient and control group sizes may differ due to recruitment constraints, disease prevalence, or study design considerations.
[0130] The equal sampling strategy generates identical numbers of synthetic samples for each class regardless of the original group sizes.Distributional Property Preservation
[0131] The synthetic data generation may be used to maintain the distributional properties of the original dataset. This represents a critical advantage over conventional augmentation methods that may alter the fundamental statistical characteristics of biological data.
[0132] Distributional properties that are preserved include mean and variance relationships through the maintenance of heteroscedastic variance patterns. Correlation structures are preserved to maintain protein-protein expression correlations. Sparsity patterns retain missing value distributions.MACHINE LEARNING-BASED BIOMARKER IDENTIFICATION AND RANKING
[0133] The method employs a trained machine learning classifier to identify and rank biomarkers from the protein expression dataset. Following statistical data augmentation, the expanded dataset undergoes preprocessing steps including missing value imputation using a constant value (-1) and normalization using StandardScaler transformation to ensure comparable protein expression scales across all features.
[0134] The initial protein dataset may be subjected to supervised feature selection to reduce computational complexity while preserving biologically relevant information. For example, given an initial dataset containing approximately 5,800 protein measurements, dimensionality may be reduced to approximately 800 proteins using feature importance scores derived from baseline Random Forest classifiers that identify proteins with the highest discriminatory power between patient and control samples.
[0135] Following the initial feature selection using Random Forest classifiers, the method employs a Gradient Boosting Classifier as the primary machine learning algorithm for final biomarker identification and ranking. The preferred machine learning classifier is implemented using scikit-leam, which provides superior performance for high-dimensional biological datasets compared to traditional classification approaches.
[0136] Biomarker ranking may be achieved by applying SHAP (Shapley Additive Explanations) analysis to the trained Gradient Boosting Classifier to extract interpretable feature importance scores that quantify each protein's marginal contribution to the classification decision. SHAP values may be computed using tree-optimized algorithms. Proteins are ranked based on their mean absolute SHAP values across all samples, providing a quantitative measure of each biomarker's importance in disease classification.
[0137] In an example embodiment, the machine learning classifier successfully identified highly ranked biomarkers that demonstrate both statistical significance and biological relevance, with exemplary implementations identifying NCAM1, SEMA7A, GRIK3, NRP1, SLITRK1, and EPHA5 as top-ranked biomarkers, each showing substantial fold-change differences between patient and control samples coupled with high feature importance scores.PATIENT STRATIFICATION METHOD FOR CENTRAL NERVOUS SYSTEM DISORDERS
[0138] The present invention provides a method of patient stratification for central nervous system disorders, which leverages the molecular profiling capabilities described herein to identify patient subgroups with distinct proteomic profiles, enabling personalized therapeutic approaches and precision medicine strategies for CNS disorders.Multi-Subject Sample Collection and iPSC Generation
[0139] In one embodiment, the method begins with obtaining induced pluripotent stem cells from multiple subjects with central nervous system disorders and control subjects. As used herein, "multiple subjects" refers to a cohort of individuals representing the heterogeneity typical of CNS disorders, typically including 10-100 subjects, preferably 20-50 subjects, to capture sufficient molecular and clinical diversity for meaningful stratification analysis.Computational Clustering and Pattern Recognition
[0140] Following mass spectrometry analysis to generate subject-specific protein expression profiles as described in earlier claims, one embodiment of the method proceeds with clustering thesubjects based on their protein expression profiles using Principal Component Analysis and Hierarchical Clustering.Clinical Correlation and Symptom Severity Assessment
[0141] In another embodiment, the method continues with linking the protein expression profiles to clinical symptom severity, establishing the clinical relevance of molecular clustering patterns. As used herein, "clinical symptom severity" refers to standardized assessments of disease presentation and functional impact using validated clinical rating scales appropriate for the specific CNS disorder under investigation.Patient Subgroup Identification and Therapeutic Stratification
[0142] The final step in one embodiment involves identifying patient subgroups with distinct proteomic profiles for targeted therapeutic intervention. As used herein, "patient subgroups" refer to molecularly distinct categories within the broader disease population, each characterized by specific protein expression patterns that may indicate different underlying disease mechanisms, progression rates, or therapeutic responsiveness profiles.CLINICAL PATIENT CLASSIFICATION METHOD
[0143] In another embodiment, the present invention provides a method for patient stratification for central nervous system disorders that enables real-time classification of individual patients based on their molecular profiles compared to previously established reference subgroups.DRUG SCREENING METHOD
[0144] The present invention provides a comprehensive drug screening method that leverages the ranked biomarkers identified through the biomarker discovery method. This screening approach represents a significant advancement over conventional drug discovery methods by utilizing disease-specific biomarkers derived from patient-specific neural models to identify therapeutically relevant compounds with enhanced precision and reduced false positive rates.GENERAL DEFINITIONS
[0145] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Definitions of common terms and techniques in molecular biology may be found in Molecular Cloning: A Laboratory Manual, 2ndedition (1989) (Sambrook, Fritsch, and Maniatis); Molecular Cloning: A Laboratory Manual, 4thedition (2012) (Green and Sambrook); Current Protocols in Molecular Biology (1987) (F.M. Ausubel et al. eds.); the series Methods in Enzymology (Academic Press, Inc.): PCR2: A Practical Approach (1995) (M.J. MacPherson, B.D. Hames, and G.R. Taylor eds ): Antibodies, A Laboratory Manual (1988) (Harlow and Lane, eds.): Antibodies A Laboratory Manual, 2ndedition 2013 (E.A. Greenfield ed.); Animal Cell Culture (1987) (R.I. Freshney, ed.); Benjamin Lewin, Genes IX, published by Jones and Bartlet, 2008 (ISBN 0763752223); Kendrew et al . (eds.), The Encyclopedia of Molecular Biology, published by Blackwell Science Ltd., 1994 (ISBN 0632021829); Robert A. Meyers (ed.), Molecular Biology and Biotechnology: a Comprehensive Desk Reference, published by VCH Publishers, Inc., 1995 (ISBN 9780471185710); Singleton etal., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, N.Y. 1994), March, Advanced Organic Chemistry Reactions, Mechanisms and Structure 4th ed., John Wiley & Sons (New York, N.Y. 1992); and Marten H. Hofker and Jan van Deursen, Transgenic Mouse Methods and Protocols, 2ndedition (2011).
[0146] As used herein, the singular forms “a”, “an”, and “the” include both singular and plural referents unless the context clearly dictates otherwise.
[0147] The term “optional” or “optionally” means that the subsequent described event, circumstance or substituent may or may not occur, and that the description includes instances where the event or circumstance occurs and instances where it does not.
[0148] The recitation of numerical ranges by endpoints includes all numbers and fractions subsumed within the respective ranges, as well as the recited endpoints.
[0149] The terms “about” or “approximately” as used herein when referring to a measurable value such as a parameter, an amount, a temporal duration, and the like, are meant to encompass variations of and from the specified value, such as variations of + / -10% or less, + / -5% or less, + / - 1% or less, and + / -0.1% or less of and from the specified value, insofar such variations areappropriate to perform in the disclosed invention. It is to be understood that the value to which the modifier “about” or “approximately” refers is itself also specifically, and preferably, disclosed.
[0150] As used herein, a “biological sample” may contain whole cells and / or live cells and / or cell debris. The biological sample may contain (or be derived from) a “bodily fluid”. The present invention encompasses embodiments wherein the bodily fluid is selected from amniotic fluid, aqueous humour, vitreous humour, bile, blood serum, breast milk, cerebrospinal fluid, cerumen (earwax), chyle, chyme, endolymph, perilymph, exudates, feces, female ejaculate, gastric acid, gastric juice, lymph, mucus (including nasal drainage and phlegm), pericardial fluid, peritoneal fluid, pleural fluid, pus, rheum, saliva, sebum (skin oil), semen, sputum, synovial fluid, sweat, tears, urine, vaginal secretion, vomit and mixtures of one or more thereof. Biological samples include cell cultures, bodily fluids, cell cultures from bodily fluids. Bodily fluids may be obtained from a mammal organism, for example by puncture, or other collecting or sampling procedures.
[0151] The terms “subject,” “individual,” and “patient” are used interchangeably herein to refer to a vertebrate, preferably a mammal, more preferably a human. Mammals include, but are not limited to, murines, simians, humans, farm animals, sport animals, and pets. Tissues, cells and their progeny of a biological entity obtained in vivo or cultured in vitro are also encompassed.
[0152] Various embodiments are described hereinafter. It should be noted that the specific embodiments are not intended as an exhaustive description or as a limitation to the broader aspects discussed herein. One aspect described in conjunction with a particular embodiment is not necessarily limited to that embodiment and can be practiced with any other embodiment(s). Reference throughout this specification to “one embodiment”, “an embodiment,” “an example embodiment,” means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment,” “in an embodiment,” or “an example embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment, but may. Furthermore, the particular features, structures or characteristics may be combined in any suitable manner, as would be apparent to a person skilled in the art from this disclosure, in one or more embodiments. Furthermore, while some embodiments described herein include some but not other features included in other embodiments, combinations of features ofdifferent embodiments are meant to be within the scope of the invention. For example, in the appended claims, any of the claimed embodiments can be used in any combination.
[0153] All publications, published patent documents, and patent applications cited herein are hereby incorporated by reference to the same extent as though each individual publication, published patent document, or patent application was specifically and individually indicated as being incorporated by reference.
[0154] Further embodiments are illustrated in the following Examples which are given for illustrative purposes only and are not intended to limit the scope of the invention.EXAMPLESEXAMPLE 1 - GENERATION OF iPSC-DERIVED FOREBRAIN NEURONS FROM NEURODEVELOPMENTAL DISORDER PATIENTS
[0155] Eight iPSC clones were obtained from patients with specific mutations and chromosomal deletions associated with neurodevelopmental abnormalities and a range of neurodevel opmental abnormalities including an autism diagnosis in up to 50% of the patients. Control iPSC lines were similarly obtained from healthy individuals.
[0156] The iPSCs were differentiated into forebrain neurons using the method described above. Briefly, iPSCs were plated as single cells at a density of 3,000 cells per well of a 6-well plate on Matrigel-coated transwell membranes in mTESR+, along with 10 pM ROCK-inhibitor Y- 27632.
[0157] The next day, cells were switched to conditionally defined comprising DMEM / F12 with sodium bicarbonate, 0.5% BSA, 0.1 mM |3-mercaptoethanol, 2 mM glutamate, 10 pM NEAA, l x N2 supplement, l x B27 without retinoic acid, and Primocin (0.1 mg / ml), containing dual SMAD inhibitors: 100 nM LDN193189 and 10 pM SB431542 for days 0-7 to induce neural induction and forebrain identity. Between days 7-20, cells were maintained in conditionally defined medium without SMAD inhibitors with the media exchanged every other day.
[0158] At DIV20, cells were scraped off the trans wells using a cell scraper and centrifuged at 1000 rpm for 4 minutes in PBS. The pellet was treated with Trypsin-Dnase I for 3.5-4 minutes, then treated with 1 ml Dnase-CMF-PBS, triturated with a fine bore pipette, and the volume was brought up to 10 ml with DMEM / F12. Cells were centrifuged at 1000 rpm for 4 minutes and thenewly formed neurospheres were plated in conditionally defined medium in non-coated 6-well plates as floating spheres with media exchanged every other day until day 27.
[0159] The cells were then dissociated and replated on poly-D-Lysine and Laminin coated 6- well plates. At day 37, forebrain and layer-specific markers PAX6, FOXG1, CUX1, CTIP2, TBRlwere detected by mass spectrometry at high signal intensities, while hindbrain-specific markers GBX2, H0XA2, and ATOH1 showed minimal detection, confirming the proper identity of the forebrain neuronal cultures.EXAMPLE 2 - CELL-TYPE SPECIFIC SPATIAL PROTEOME PROFILING OF LAYER V CORTICAL NEURONS
[0160] The iPSC-derived forebrain neurons generated in Example 1 were subjected to celltype specific spatial proteome profiling. Layer V cortical neurons were targeted using an antibody against GRIK3, a cell surface antigen specifically expressed on layer V cortical neurons.
[0161] The antibody against GRIK3 was conjugated to horseradish peroxidase (HRP) using commercial conjugation kits according to manufacturer's instructions. The HRP-conjugated antibody was applied to the neuronal cultures in HEPES-PBS for 20 minutes at 4°C.
[0162] After washing, spatial biotinylation was performed by incubating cells with 100 pM biotin-tyramide (BxxP) and 1 mM H2O2 in Tyrode's buffer for 10 minutes at room temperature. The labeling reaction was quenched by replacing the media with a quencher solution consisting of Tyrode's buffer with 10 mM sodium azide, 10 mM sodium ascorbate, and 5 mM Trolox, exchanged three times.
[0163] Cells were harvested by scraping, pelleted, and frozen at -80°C. For protein extraction, the pellet was resuspended in 400 pl high-SDS RIPA buffer and dissolved through pipetting and sonication until cleared. The samples were incubated at 95°C for 5 minutes to denature proteins, then 1.2 ml of SDS-free RIPA was added to dilute SDS to 0.2%. Samples were rotated for 30 minutes at 4°C, then centrifuged at 20,000 g for 20 minutes.
[0164] The supernatant was collected and protein concentrations were measured. The lysates were normalized for protein concentration and added to avidin-conjugated magnetic beads according to manufacturer's recommendations and rotated overnight at 4°C. Beads were washed according to manufacturer's recommendations.
[0165] The isolated proteins were analyzed by nano liquid chromatography tandem mass spectrometry (LC-MS / MS) in data independent mode with a false discovery rate of 1%. Negative controls included samples with no antibody, and no biotin-tyramide.EXAMPLE 3 - REPRODUCIBILITY ANALYSIS OF SPATIAL PROTEOME PROFILING METHOD
[0166] To validate the reproducibility of the spatial proteome profiling method, correlation analyses were performed between different sample types:
[0167] Technical replicates from the same biological sample showed high correlation (R2= 0.97), demonstrating the technical precision and consistency of the spatial proteome labeling and mass spectrometry analysis procedures.
[0168] Biological replicates derived from different iPSC clones obtained from the same individual yielded a correlation coefficient of R2= 0.77, validating that the method can reliably detect consistent protein expression patterns across different cellular preparations from the same patient.
[0169] Samples from different individuals showed a correlation coefficient of R2= 0.82, indicating that the method reliably captures consistent protein signatures across different patients. By contrast, there was no correlation between a sample processed for cell surface labeling versus a negative control without biotin-tyramide, confirming the specificity of the method.EXAMPLE 4 - STATISTICAL DATA AUGMENTATION FOR SPARSE BIOLOGICAL DATASETS
[0170] The protein expression data generated from Examples 3 and 4 was subjected to statistical data augmentation.
[0171] The data was segmented into discrete patient and control cohorts. For each protein within each group, the standard deviation of observed values was calculated across non-null observations only, preserving the natural variance structure of each protein.
[0172] Synthetic samples were generated by applying random multiplicative perturbations within ±4 standard deviations to existing non-null values. For each protein i and each existing non-null value, synthetic values were generated using the formula: x(synthetic) = x(original) x (1 + r x o / |i), where r was a random variable drawn from a uniform distribution U(-4, 4).
[0173] Missing values were preserved in their original form to maintain sparsity patterns characteristic of mass spectrometry data. Equal numbers of synthetic samples were generated for patient and control groups to address class imbalance, increasing the dataset from the original sample size to 2,000 total samples (1,000 patient, 1,000 control).EXAMPLE 5 - MACHINE LEARNING-BASED BIOMARKER IDENTIFICATION
[0174] Following data augmentation from Example 5, the expanded dataset was processed for machine learning analysis. Missing values were imputed using a constant value (-1) and the dataset was normalized using StandardScaler transformation.
[0175] Feature dimensionality was reduced from approximately 5,800 proteins to 800 using supervised feature selection based on feature importance scores derived from Random Forest classifiers. The dataset was split into training (80%) and testing (20%) sets using stratified sampling.
[0176] A Gradient Boosting Classifier was trained using stratified 5-fold cross-validation with shuffling and a fixed random seed (42) to ensure reproducibility. The model achieved high classification accuracy in distinguishing patient samples from control samples.
[0177] Feature importance analysis was performed using SHAP (Shapley Additive Explanations) values to rank protein contributions to disease classification. SHAP values were computed using TreeExplainer and proteins were ranked based on their mean absolute SHAP values.
[0178] The method successfully identified several highly ranked candidate biomarkers including NC AMI, SEMA7A, GRIK3, NRP1, SLITRK1, and EPHA5, which showed significant differential expression between patient and control samples and high feature importance scores.Table 1: Identified Biomarker Proteins and Their Characteristics
[0179] The following table summarizes the key biomarker proteins identified through the disclosed methods, their cellular functions, and relevance to CNS disorders:EXAMPLE 6 - PATIENT STRATIFICATION BASED ON PROTEOMIC PROFILES
[0180] Hierarchical clustering analysis was performed on the protein expression profiles from multiple patient samples to demonstrate patient stratification capabilities. Patients with mild overall symptom severity (able to talk and walk unassisted) clustered most closely to control subjects and to each other, while forming a distinct cluster separate from patients with moderate overall symptom severity (non-verbal, able to walk but unsteady without assistance).
[0181] Principal Component Analysis (PCA) was performed to further analyze the clustering patterns. The PCA plot revealed that negative controls clustered separately from all other samples, confirming method specificity. Patients with mild symptoms clustered most closely to controls while remaining distinguishable from patients with moderate symptoms.
[0182] Importantly, two genetically defined neurodevelopmental conditions could be distinguished from one another based solely on their surface proteomics profiles, demonstrating the method's ability to capture disease-specific molecular signatures that reflect both symptom severity and underlying genetic etiology.EXAMPLE 7 - VALIDATION USING PUBLISHED RETT SYNDROME DATASET
[0183] To further validate the machine learning approach, the method was applied to published cerebrospinal fluid proteomics data from Rett syndrome patients. The dataset contained 192 differentially expressed proteins from 6 individuals (patients and controls).
[0184] The data was synthetically expanded 10-fold using the statistical data augmentation method of Example 5, increasing the sample size from 6 to 60. Missing values were imputed using a constant value (-1) and the dataset was normalized using StandardScaler.
[0185] Because the dataset was already balanced and reduced in features (containing only differentially expressed proteins), balancing and feature selection steps were omitted. A Gradient Boosting Classifier was trained on the augmented dataset.
[0186] Feature importance was assessed using SHAP values. Examining the SHAP values revealed which proteins were driving the model predictions. Importantly, MECP2, the known driver gene for Rett Syndrome, was identified among the top 35 proteins out of the 192 differentially expressed proteins, demonstrating the method's ability to identify biologically relevant disease drivers without prior knowledge.EXAMPLE 8 - DRUG SCREENING APPLICATION
[0187] The ranked biomarkers identified in Example 6 were used for drug screening applications. Neural cells generated using the method of Example 1 were treated with various test compounds from small molecule libraries.
[0188] Expression levels of the top-ranked biomarkers (NCAM1, SEMA7A, GRIK3, NRP1, SLITRK1, EPHA5) were measured using targeted mass spectrometry approaches after compound treatment. Test compounds were screened at concentrations of 1 pM, 10 pM, and 100 pM to establish dose-response relationships.
[0189] Drug candidates were identified based on their ability to modulate biomarker expression levels compared to vehicle control conditions. Compounds achieving normalization of biomarker expression profiles toward control levels (>50% rescue index) were considered high- priority candidates for further validation.
[0190] Functional phenotypes were assessed including neuronal electrical activity using microelectrode arrays, neurite outgrowth analysis using automated imaging, and cell viabilityusing metabolic activity assays. Lead compounds showed both biomarker modulation and functional improvement.* * *
[0191] Various modifications and variations of the described methods, pharmaceutical compositions, and kits of the invention will be apparent to those skilled in the art without departing from the scope and spirit of the invention. Although the invention has been described in connection with specific embodiments, it will be understood that it is capable of further modifications and that the invention as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the invention that are obvious to those skilled in the art are intended to be within the scope of the invention. This application is intended to cover any variations, uses, or adaptations of the invention following, in general, the principles of the invention and including such departures from the present disclosure come within known customary practice within the art to which the invention pertains and may be applied to the essential features herein before set forth.
Claims
CLAIMSWhat is claimed is:
1. A method for identifying ranked biomarkers for central nervous system disorders comprising:(a) generating neural cells from induced pluripotent stem cells (iPSCs) derived from a subject with a central nervous system disorder or healthy control subject by: (i) contacting the iPSCs with inhibitors of SMAD signaling and culturing the cells in 2D on transwell membrane (polyester) plates; (ii) culturing the cells to form neurospheres; and (iii) dissociating the neurospheres and replating the cells in 2D to further differentiate the cells into region-specific neural cells;(b) binding cell surface antigens by contacting the neural cells with an antibody- peroxidase conjugate and conducting a spatial proteome profiling reaction of proteins proximate to the antibody-peroxidase conjugate;(c) isolating the labeled proteins using affinity purification;(d) analyzing the isolated proteins by mass spectrometry to generate protein expression data;(e) computationally expanding the protein expression dataset by applying statistical data augmentation; and(f) identifying ranked biomarkers using a trained machine learning classifier.
2. The method of claim 1, wherein the statistical data augmentation comprises: (i) calculating standard deviations for each protein within patient and control groups; (ii) generating synthetic data points by applying multiplicative perturbations within ±4 standard deviations to existing values while preserving null values; and (iii) balancing synthetic sample numbers between patient and control groups.
3. The method of claim 1, wherein the central nervous system disorder is selected from the group consisting of autism spectrum disorder, schizophrenia, bipolar disease, epilepsy, rare neurodevelopmental disorders including but not limited to Rett Syndrome, CDKL5 deficiency disorder, Fragile X syndrome, and neurodegenerative disorders such as Alzheimer's disease, Parkinson's disease, amyotrophic lateral sclerosis, and frontotemporal dementia.
4. The method of claim 1, wherein the region-specific neurons are selected from the group consisting of forebrain excitatory neurons.
5. The method of claim 1, wherein the antibody -enzyme conjugate comprises horseradish peroxidase conjugated to an antibody that binds to a cell surface antigen expressed on layerspecific cortical neurons.
6. The method of claim 1, wherein the biotinylating substrate is a biotin-tyramide derivative and the catalyzing step is performed in the presence of hydrogen peroxide.
7. The method of claim 1, wherein the affinity purification is performed using avidin- conjugated magnetic beads.
8. The method of claim 1, wherein the mass spectrometry analysis is performed using nano liquid chromatography tandem mass spectrometry in data independent mode with a false discovery rate of 1%.
9. The method of claim 1, wherein the trained machine learning classifier is a gradient boosting classifier.
10. The method of claim 1, wherein the trained machine learning classifier identifies the ranked biomarkers using feature importance analysis.
11. The method of claim 10, wherein the feature importance analysis uses SHAP (Shapley Additive Explanations) values to rank protein contributions to disease classification.
12. The method of claim 1, further comprising validating identified biomarkers by confirming their differential expression between patient and control samples.
13. A method for cell-type specific spatial proteome profiling comprising:(a) contacting a neural cell sample with antibody-peroxidase conjugate that binds to a cell surface antigen;(b) performing enzyme-mediated protein tagging by contacting the cells with a biotinylating substrate comprising biotin-tyramide or biotin-tyramide derivatives and hydrogen peroxide to biotinylate proteins proximated to the antibody-enzyme conjugate;(c) quenching the labeling reaction;(d) isolating the biotinylated proteins; and(e) identifying the isolated proteins by liquid chromatography tandem mass spectrometry.
14. The method of claim 13, wherein the cell surface antigen is selected from the group consisting of GRIK3, GRIN3A, LYPD1, EPHA5, RXFP1, SLIT3, FGFR1, NRG1, CDH22, SEMA3E, and SEMA3D and the cells are layer V cortical neurons.
15. The method of claim 13, further comprising comparing protein expression profiles between neurons derived from patients and control subjects.
16. A statistical data augmentation method for sparse biological datasets comprising:(a) receiving a biological dataset containing protein expression measurements from patient and control groups;(b) segmenting the data into discrete patient and control cohorts;(c) calculating feature-wise standard deviations for each protein across non-null observations within each cohort;(d) generating synthetic samples by introducing multiplicative stochastic perturbations drawn from a uniform distribution bounded within ±4 standard deviations of original observed values;(e) preserving missing values in their original form to maintain sparsity patterns; and(f) creating equal numbers of synthetic samples per class to address class imbalance.
17. The method of claim 16, wherein the biological dataset contains fewer than 1000 samples.
18. The method of claim 16, wherein the biological dataset comprises normalized mass spectrometry peak intensity values.
19. The method of claim 16, wherein the synthetic data generation maintains distributional properties of the original dataset.
20. A method for generating cortical neurons from induced pluripotent stem cells comprising:(a) plating iPSCs as single cells at sparse density on Matrigel-coated transwell membranes (polyester membranes) in mTESR+ containing ROCK inhibitor;(b) removing the ROCK inhibitor and treating the cells with dual SMAD inhibitors in conditionally defined medium to induce neural fate;(c) culturing cells in conditionally defined medium without SMAD inhibitors;(d) treating cells with Trypsin-DNase I followed by DNase-CMF-PBS (Deoxyribonuclease in Calcium and Magnesium -Free Phosphate Buffered Saline) to form neurospheres;(e) culturing the neurospheres as floating spheres in conditionally defined medium; and(f) dissociating and replating the neurospheres on poly-D-Lysine and Laminin coated plates to generate adherent cortical neurons.
21. The method of claim 20, wherein the dual SMAD inhibitors are LDN193189 and SB431542.
22. The method of claim 20, wherein: (a) the dual SMAD inhibitors are applied for days 0-7;(b) the culturing in conditionally defined medium without SMAD inhibitors is performed from days 7-20; (c) the cells are treated to form neurospheres at day 20; (d) the neurospheres are cultured until day 27 in conditionally defined medium without SMAD inhibitors; and (e) adherent cortical neurons are generated between days 37 and 70.
23. The method of claim 20, wherein the conditionally defined medium comprises DMEM / F12 with sodium bicarbonate, 0.5% BSA, 0.1 mM P-mercaptoethanol, 2 mM glutamate, 10 pM NEAA, 1 x N2 supplement, 1 x B27 without retinoic acid, and Primocin.
24. The method of claim 20, wherein the cortical neurons express forebrain-specific markers and do not express hindbrain-specific markers at day 37.
25. A method of patient stratification for central nervous system disorders comprising:(a) obtaining induced pluripotent stem cells from multiple subjects with central nervous system disorders and control subjects;(b) generating neural cells from each subject's iPSCs by differentiating the iPSCs into forebrain excitatory neurons;(c) performing cell-type specific cell surface protein labeling on the neural cells from each subject using antibody-enzyme conjugates;(d) analyzing the labeled proteins by mass spectrometry to generate subject-specific protein expression profiles;(e) clustering the subjects based on their protein expression profiles using Principal Component Analysis and Hierarchical Clustering;(f) linking the protein expression profiles to clinical symptom severity (mild, moderate, severe); and(g) identifying patient subgroups with distinct proteomic profiles for targeted therapeutic intervention.
26. The method of claim 25, wherein the clustering correlates with symptom severity levels among the subjects.
27. The method of claim 25, wherein subjects with mild symptom severity cluster more closely to control subjects than subjects with moderate symptom severity.
28. The method of claim 25, further comprising correlating the patient subgroups with genetic backgrounds of the subjects.
29. A method for patient stratification for central nervous system disorders comprising:(a) generating neural cells from induced pluripotent stem cells derived from a patient sample;(b) performing cell-type specific cell surface protein labeling on the neural cells using antibody-enzyme conjugates;(c) analyzing the labeled proteins by mass spectrometry to generate a patient-specific protein expression profile;(d) comparing the patient-specific protein expression profile to reference protein expression profiles using computational similarity analysis, wherein each reference protein expression profile is from a previously established patient subgroup, wherein each subgroup is characterized by distinct proteomic signatures and clinical phenotypes; and(e) assigning the patient to a patient subgroup based on the highest similarity match for targeted therapeutic intervention.
30. A method of screening drug candidates comprising:(a) identifying ranked biomarkers using the method of claim 1;(b) contacting neural cells with one or more test compounds;(c) measuring expression levels of the ranked biomarkers in the presence of the test compounds; and(d) identifying drug candidates based on modulation of the biomarker expression levels compared to control conditions.
31. The method of claim 30, wherein the test compounds are screened against a panel of the top-ranked biomarkers.
32. The method of claim 30, wherein the drug candidates are identified based on normalization of biomarker expression levels toward control levels.
33. The method of claim 30, further comprising: (e) measuring a functional phenotype of the neural cells in the presence of the test compounds, wherein the functional phenotype is selected from the group consisting of neuronal electrical activity, synaptic function, neurite outgrowth, calcium signaling, and cell viability; and (f) identifying drug candidates based on improvement of the functional phenotype in addition to modulation of biomarker expression levels.
34. The method of claim 30, further comprising comparing spatial protein expression of one region of a neural cell from a first diseased subject with that of a second diseased subject to identify variability in protein expression between diseased subjects.
35. The method of claim 30, further comprising comparing spatial protein expression of one region of a neural cell from a diseased subject with that of a healthy subject to identify variability in protein expression between diseased and healthy subjects.
36. The method of claim 30, further comprising comparing spatial protein expression of one region of a neural cell from a first healthy subject with that of a second healthy subject to identify baseline variability in protein expression within the healthy population.
37. The method of claim 30, wherein the drug candidates are identified based on normalizing protein expression variability identified through multiple subject comparisons toward healthy baseline levels.
38. The method of claim 30, wherein the drug candidates target at least one protein selected from the proteins in Table 1.
39. A kit for detecting biomarkers for central nervous system disorders comprising reagents for detecting one or more biomarkers identified by the method of claim 1.
40. The kit of claim 39, wherein the reagents comprise antibodies specific for cell surface proteins selected from the group consisting of NCAM1, SEMA7A, GRIK3, NRP1, SLITRK1, SEMA3E, SEMA3D, EPHA5, and proteins listed in Table 1.
41. The kit of claim 39, wherein the kit comprises reagents for detecting a panel of at least 3 biomarkers.
42. A method of treating a subject with a central nervous system disorder comprising administering to the subject a therapeutically effective amount of a compound that modulates a biomarker identified by the method of claim 1.
43. The method of claim 42, wherein the central nervous system disorder is autism spectrum disorder and the compound modulates GRIK3 activity.
44. A method of diagnosing a central nervous system disorder comprising:(a) obtaining a biological sample from a subject;(b) measuring expression levels of biomarkers identified by the method of claim 1; and(c) comparing the expression levels to reference profiles to determine disease status.
45. The method of claim 44, wherein the biomarkers are cell surface proteins and the measuring is performed using immunoassay techniques.