Drug screening method

WO2024173560A3PCT designated stage expired Publication Date: 2025-05-08DANA FARBER CANCER INSTITUTE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/015807
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-29
Filing Date
2024-02-14
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Current methods for identifying SWI/SNF modulating compounds are inadequate, particularly in understanding the biochemical contributions of ARID1A/B subunits' intrinsically disordered regions (IDRs) to chromatin remodeling and their role in human diseases like cancer and neurodevelopmental disorders.

Method used

An in vitro screening method involving incubation of candidate compounds with ARID1A or ARID1B IDRs or fragments, detecting modulation of IDR functions such as condensate formation, chromatin localization, and gene expression, to identify SWI/SNF modulating compounds.

Benefits of technology

This method effectively identifies compounds that modulate SWI/SNF activity, providing insights into chromatin remodeling and potential therapeutic targets for cancer and neurodevelopmental disorders by elucidating the functional role of IDRs in cBAF complex targeting and genomic accessibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024015807_08052025_PF_FP_ABST
    Figure US2024015807_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention is directed to an in vitro screening method to identify a SWI / SNF modulating compound.
Need to check novelty before this filing date? Find Prior Art

Description

DRUG SCREENING METHOD

[0001] This application claims priority to U.S. Provisional Application No. 63 / 445,625, filed on February 14, 2023, U.S. Provisional Application No. 63 / 537,961, filed on September 12, 2023, and U.S. Provisional Application No. 63 / 541,503, filed September 29, 2023, the entire contents of each of which are incorporated herein by reference.

[0002] All patents, patent applications and publications cited herein are hereby incorporated by reference in their entirety. The disclosures of these publications in their entireties are hereby incorporated by reference into this application in order to more fully describe the state of the art as known to those skilled therein as of the date of the invention described and claimed herein.

[0003] This patent disclosure contains material that is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or the patent disclosure as it appears in the U.S. Patent and Trademark Office patent file or records, but otherwise reserves any and all copyright rights.FIELD OF THE INVENTION

[0004] This invention is directed to an in vitro screening method to identify a SWI / SNF modulating compound. Aspects of the invention are drawn to methods of incubating a candidate compound with a protein or protein complex. For example, the protein or protein complex can comprise an ARID1 A or ARID1B intrinsically disordered region (IDR) or fragment thereof.BACKGROUND

[0005] Intrinsically disordered regions (IDRs) comprise 37-50% of the human proteome and are especially enriched in nuclear proteins. Rather than a singular structure, IDRs comprise heterogeneous conformational ensembles, indicating that IDR-mediated interactions are less specific than those mediated by folded domains. However, associations driven by specific IDRs play important roles in forming biomolecular condensates, which are regions of high local protein concentration formed via the process of phase separation or related phase transitions. IDRs and their role in driving phase separation are implicated in various aspects of nuclear organization, but much remains unclear in the context of chromatin remodeling.

[0006] The mammalian SWI / SNF (mSWI / SNF or BAF) ATP-dependent chromatin remodeling complexes collectively represent one of the most frequently mutated cellular entities in human cancer, second only to TP539. Indeed, mutational frequencies for the 29 human genes that encode for mSWI / SNF complex subunits tally to over 20% across human cancers. mSWI / SNF subunit mutations and translocations represent cancer-initiating events in a number of rare cancers and are among the most frequently perturbed genes in neurodevelopmental disorders (NDDs).

[0007] Mammalian SWI / SNF chromatin remodeling complexes, specifically, ARID1A / B- nucleated cBAF complexes, are among the most disrupted molecular entities in human disease. However, the biochemical contributions of the ARID1A / B subunits, such as their frequently- mutated N-terminal IDRs, to cBAF function remain poorly understood.SUMMARY OF THE INVENTION

[0008] Aspects of the invention are directed to in vitro screening methods to identify a SWI- SNF modulating compound. In embodiments, the method comprises incubating a candidate compound with a protein or protein complex. For example, the protein or protein complex comprises an ARID1 A or ARID1B intrinsically disordered region (IDR) or fragment thereof. In embodiments, the method comprises detecting at least one IDR function. In embodiments, modulation of the IDR function is indicative of a SWI-SNF modulating compound.

[0009] In embodiments, the candidate compound binds to the IDR.

[0010] In embodiments, the IDR function is selected from the group consisting of (i) condensate formation, (ii) chromatin localization, (iii) IDR protein engagement, and (iv) gene expression

[0011] In embodiments, the ARID1A or ARID1B intrinsically disordered region or fragment thereof, is purified from mammalian cells.

[0012] In embodiments, the protein or protein complex comprising an ARID 1 A or ARID IB IDR or fragment thereof is selected from the group consisting of a purified polypeptide, a chromatin remodeling complex, a DNA-binding protein, an epitope-tagged peptide, a core binding region (CBR) peptide, or any combination thereof. For example, the chromatin remodeling complex comprises an ATP-dependent chromatin remodeling complex. For example, the ATP-dependent chromatin remodeling complex comprises the mammalian SWI / SNF chromatin remodeling complex.

[0013] In embodiments, the candidate compound comprises a small molecule.

[0014] In embodiments, detecting comprises a sequencing technique, an immunoassay, mass spectrometry, immunoprecipitation, or a combination thereof.

[0015] In embodiments, the in vitro screening method further comprises comparing modulation of the IDR function by the candidate compound to that of a control compound. For example, the control compound comprises a positive control, a negative control, or both.

[0016] In embodiments, the in vitro screening method further comprises selecting the candidate compound as a therapeutic agent when the candidate modulates the IDR function.

[0017] In embodiments, incubating occurs in a reaction vessel. For example, the reaction vessel comprises a dish, a tube, or a multi-well plate.

[0018] In embodiments, the method is a high-throughput method.

[0019] Aspects of the invention are further drawn to a therapeutic compound identified by the in vitro screening method described herein.

[0020] Aspects of the invention are further drawn to a plate comprising one or more wells. In embodiments, at least one well comprises a candidate compound, a protein or protein complex comprising an ARID 1 A or ARID IB intrinsically disordered region (IDR) or fragment thereof, and a medium.

[0021] In embodiments, the at least one well further comprises a plurality of cells.

[0022] In embodiments, the medium comprises a cell culture medium.

[0023] Aspects of the invention are further drawn to a kit comprising the plate described herein.

[0024] In embodiments, the plate is designed for automated drug screening.

[0025] Other objects and advantages of this invention will become readily apparent from the ensuing description.BRIEF DESCRIPTION OF THE FIGURES

[0026] FIG. 1 shows the IDRs of ARID1A / B are dispensable for cBAF assembly and in vitro nucleosome remodeling. Panel A shows human cBAF complex (PDBDEV_00000056) with putative ARIDlAN-terminal region of unassigned cryo-EM density and C-terminal CBR highlighted . Panel B shows disease-associated mutations mapped onto ARID1A / B and disorder as PONDR score. Panel C provides distribution of disease-associated missense and indel mutations in ARIDIA / B’s N-terminus. Panel D provides a schematic of HA-tagged ARID 1 A expression constructs. Panel E provides immunoblots of nuclear protein input and anti-HA IPs in AN3CA (ARIDlA / B-deficient) cells expressing HA-tagged ARID1A WT or mutant variants. Panel F provides TMT mass spectrometric signal for cBAF components fromanti -HA ARID 1 A WT or mutant immunoprecipitation. Panel G (top) shows a restriction enzyme accessibility assay (REAA) time course using 2.5 nM purified cBAF carrying ARID 1 A WT or mutant variants; (bottom) REAA using 0-5 nM cBAF (t=30min) (n=2 experimental replicates each). Panel H provides ATPase (ADP-Glo) measurements for indicated conditions and timepoints, ns, not significant by one-way ANOVA test.

[0027] FIG. 2 shows ARID 1 A IDRs dictate cBAF complex condensation in vitro and in cells, which is enhanced by DNA binding. Panel A (left) shows in vitro condensation experiments of indicated 0.66 pM eGFP-tagged cBAF complexes; (right) shows condensate area per field of view. Panel B provides percent condensate-covered area with 100 nM DNA, nucleosomes, or RNA. Panel C provides confocal imaging of eGFP-tagged cBAF complexes containing ARID1 A WT or mutant variants in live AN3CA cells. Panel D provides saturation concentration, condensate count and area ARID1A puncta in AN3CA cells. PS: Phase Separation. Panel E provides an immunoblot for ARID 1 A and other cBAF subunits in AN3CA cells - / + doxycycline alongside human and murine cell ty pes. Panel F provides immunofluorescence of AN3CA cells without or with doxycycline induction of exogenous eGFP-tagged ARID1A. PCC between eGFP-ARIDlA and anti-ARIDlA immunostaining . Bottom: Immunostain for endogenous ARID 1 A in KLE (human endometrial), C2C12 myoblast (mouse), MCF-10A (human breast cancer), and primary' rat neurons. Panel G (top) provides a schematic of Corelet system used to evaluate self-interaction propensity of IDRs; (bottom) provides a schematic of IDR-containing constructs evaluated. Panel H provides representative images of U2OS cell nuclei yvithout (-light) and yvith (-Flight) light-induced oligomerization. Panel I (top) provides phase diagram schematic; (bottom) provides phase diagrams of ARID 1 A constructs: shaded area indicates two-phase region. In Panel A and Panel B, P-values calculated by one-way ANOVA test. In Panel D, by unpaired student's t-test.

[0028] FIG. 3 shows ARID 1 A IDRs and DNA-binding functions govern cBAF occupancy, DNA accessibility and gene expression in cells. Panel A shoyvs chromatin occupancy of cBAF complexes marked by HA (ARID1A), SMARCA4, and SMARCC1, H3K27ac enhancer mark occupancy and DNA accessibility (ATAC) at cBAF-occupied sites in AN3CA cells, divided into 4 clusters using k-means clustering. Panel B provides Distance-to-TSS distribution of merged CUT&Tag and ATAC-Seq peaks for the conditions, across Clusters 1-4 from (A). Panel C provides Principal Component Analysis (PC A) of cBAF-occupied enhancer sites across conditions as assayed by SMARCA4 and SMARCC1 signals. Panel D provides representative CUT&Tag and ATAC-Seq tracks at the MAP2. NCAPH. and intergenicenhancer loci in AN3CA cells across Empty and ARID1 A WT or mutant conditions. Panel E provides overlap of accessible sites by ATAC-Seq in empty vector control (Empty) versus ARID 1 A WT or mutant conditions in AN3CA cells. Gained sites relative to empty condition are highlighted in bold. Panel F provides transcription factor motif enrichment analysis (HOMER) at Clusters 2, 3, and 4 from (A). Panel G provides a box and whisker plot for the conditions comparing expression levels of top differentially expressed genes (DEGs) upon ARID! A WT introduction versus empty control.

[0029] FIG. 4 shows ARID1A IDRs mediate local proximity of cBAF complex with cellular transcriptional machinery, enabling ARID domain-dependent TF binding. Panel A provides an immunoblot for input and anti -HA IP from AN3CA cells expressing HA- ARID 1 A fused to biotin ligase TurboID (TbID). Panel B provides distribution of biotinylated proteins fold changes. Panel C provides volcano plots comparing biotinylated protein levels. Panel D provides immunofluorescence analysis of ARID1A and p300 in AN3CA cells. Panel E provides volcano plots comparing detected protein levels following IP-Mass Spec. Panel F shows overlap of ARID 1 A WT-carrying cBAF interactomes measured using proximity labeling or IP-Mass Spec. Panel G shows protein class enrichment of detected cBAF interacting proteins via IP-MS (DNA interactors in red). Panel H shows input and selected transcription factor (cJUN, NFIA, TEAD1) reciprocal IPs using AN3CA cells expressing empty vector or WT- and mutant- ARID1 A.

[0030] FIG. 5 shows sequence-specific heterotypic interactions of ARID 1 A IDR1 are required for cBAF-mediated chromatin and gene regulation. Panel A provides a schematic of ARID1 A FUSIDRand DDX4IDRfusion mutant variants. Panel B provides representative images of eGFP-tagged constructs in live AN3CA cells. Panel C provides count and average area of condensates. Statistical test: one way ANOVA. Panel D provides FRAP curves, Immobile fraction, and half time of recovery (T1 / 2) quantification for indicated constructs. Error bars: standard deviation, n = 3 biological trials, 15 cells each. Statistical test, one way ANOVA. Panel E provides chromatin occupancy of cBAF complexes marked by HA (ARID 1 A), SMARCA4 and SMARCC1, H3K27ac enhancer mark occupancy and DNA accessibility (ATAC-Seq) at Cluster 2 and 3 sites from FIG. 3, panel A. Panel F shows fold change of differentially expressed genes (DEGs) relative to empty vector. Panel G provides volcano plots comparing detected protein levels by IP -MS. Hits meeting the cut off of log2 fold change <-l and >1 and p-value <0.25 are blue and red. respectively. Panel H provides immunofluorescence analysis of ARID 1 A and p300. Panel I (top) provides metaplots ofSMARCA4 occupancy over cBAF sites (shared SMARCA4 / SMARCC1 sites) AIDR1 (left) or CBR-only (right) cBAF complex target sites; (bottom) provides metaplots of ATAC-Seq accessibility. Panel J provides example tracks of SMARCA4 occupancy and DNA accessibility in the ARID1 A CBR-only, FUSIDR, and DDX4IDRmutant conditions at the BRD2 and CD320 genomic loci.

[0031] FIG. 6 shows sequence patterning analysis enabled separation of condensation and heterotypic interaction functions in ARID 1 A IDR1. Panel A provides clustering analysis of non-random amino acid sequence features performed across the IDRs within mSWI / SNF proteins. Z-scores for enriched / ‘blocky’ or depleted / ‘well-mixed’ sequence features are shown as a green-to-purple color scale. Red arrow: ARID1A / B IDRs. IDR sequence feature key in Panel B. Panel B (left) provides enrichment of amino acid sequence features across Clusters 1-4 of mSWI / SNF IDR patterns; (right) provides IDR sequence feature key. Panel C provides a schematic for 42YS and AQG scramble ARID1A IDR1 rationally designed mutant variants. Panel D provides NARDINI plots of ARID 1 A IDR1 WT, AQG scramble and 42YS mutant IDRs. Amino acid key on left. Panel E provides an immunoblot for input and anti-HA IP from AN3CA cells. Panel F provides live cell imaging of eGFP-tagged cBAF complexes containing WT ARID1 A and the 42YS or AQG scramble IDR1 variants. Panel G provides condensation metrics for ARID 1 A WT and mutants (3 biological trials of n=25 cells each); error bars represent SEM. **p=0.002 by unpaired t-test. Panel H provides a clustered heatmap of chromatin occupancy of cBAF complexes marked by HA (ARID 1 A), SMARCA4 and H3K27ac enhancer mark occupancy and DNA accessibility (ATAC-Seq) across empty, WT ARID1A and the 42YS or AQG scramble IDR1 ARID1A mutants. Panel I shows overlap between Cluster B lost sites from H and Clusters 2,3 lost sites from FIG. 3, panel A. Panel J provides DEGs in WT and 42YS and AQGscram conditions relative to empty control. Panel K provides TbID proximity labeling results for the AQG scramble and 42Y S mutants compared to ARID1A WT. Hits meeting the cut off log2 fold change < -1 and >1 and p-value <0.2 are labeled in blue. Panel L provides immunofluorescence of p300 and eGFP-tagged cBAF complexes containing WT ARID1 A or AQG scramble. Panel M provides nuclear protein input and anti-TF IP-immunoblot studies.

[0032] FIG. 7 shows mutations in ARID1B IDR1 sequence pattern disrupt condensation and genomic targeting of cBAF. Panel A provides mutational frequencies in ARID1A / B IDRs associated with neurodevelopmental disorders (NDD) from DECIPHER. Panel B provides NDD-associated mutations (DECIPHER) plotted across the 26 sequence blocks within IDR1of ARID1B. Panel C provides a schematic of ARID1B WT, block deletion, and NDD mutants. Panel D provides an immunoblot for nuclear input and anti-HA IP experiments in AN3CA cells expressing HA-tagged ARID IB WT or mutants. Panel E provides representative images of eGFP-tagged ARID1B in AN3CA cells. Panel F provides condensation metrics of ARID1B in AN3CA cells. Panel G provides SMARCA4 genomic localization over severely lost sites in Block 9 deletion (left) and Block 13 deletion (right). Panel H provides PCA of ATAC-Seq peaks across ARID IB WT and mutant conditions. Panel I provides example tracks of cBAF localization and ATAC acces sibil ity over NCAPH and IL1B loci in ARID1B WT and mutant conditions. Panel J provides transcription factor motif enrichment analysis (HOMER) of ClusterY sites (FIG. 14). Panel K shows the change inNFI TF family TMT-MS signal in the S320 G327del mutant condition relative to WT ARID IB. Panel L provides differential gene expression changes for top upregulated genes in WT versus Block 9 and 13 deletions and mutant conditions. Panel M provides relative gene expression changes of top differential genes across WT and mutant conditions. Panel N provides a model highlighting the role of the ARID 1 A N-terminus.

[0033] FIG. 8 shows structural and functional features of ARID1 A / B cBAF subunits. Panel A (left) provides a 3D structure of the human cBAF complex (PDB:6LTJ) with the ARID1 A C-terminal Core Binding Region (CBR) highlighted; (right) provides a structure of the yeast SWI / SNF complex (PDB:7EGP) with the ARID1A homolog Swil highlighted. Residues of complex subunits within 10A of ARID 1 A or Swil are highlighted in color. Panel B provides mutational frequencies of ARID1 A and ARID1B associated with cancer (TCGA) or neurodevelopmental disorders (NDD) (DECIPHER), respectively. Percentages of total cancer and NDD cases for each are indicated. Panel C provides the amino acid level conservation between ARID1A and ARID1B, defined by pairwise alignment using EMBOSS Needle. Panel D provides structural models of ARID 1 A and ARID IB subunits using Alphafold highlights disordered regions (IDR1 and IDR2). The N- and C-termini of IDR1 and 2 and the ARID and CBR (Arm repeat) domains are labeled. Panel E provides nuclear proteins ranked based on degree of disorder using MobiDB-lite. Panel F provides an interaction model of the ARID 1 A ARID domain and dsDNA (PDB: 1RYU, 1KQQ overlap). Highlighted residues SI086, S1087 and SI 090 were mutated to glutamic acid (E) to compromise DNA binding (DBDmutmutant). Panel G (left) provides an electrophoretic mobility shift assay (EMSA) using GST-tagged wild-type (WT) or DBDmutARID domain proteins (aa958-1375) and a IRDye800 labeled dsDNA probe; (right) provides quantification of DNA binding. Panel H provides animmunoblot performed on nuclear extracts isolated from naive HEK293T, HEK293T AARID1A / B and AN3CA cells. Panel I provides immunoblots performed on nuclear protein input and anti-HA IPs in AARID1 A / B HEK293T cells with rescue of HA-tagged ARID1 A WT or mutant variants. Panel J provides an immunoblot of MG- 132 proteasome inhibitor treated AN3CA cells expressing ARID 1 A WT and mutants. Panel K provides density sedimentation analysis using 10-30% glycerol gradients performed on nuclear extracts of AN3CA cells (top) and AN3CA cells rescued with HA-WT ARID! A (bottom). Panel L provides an immunoblot and quantitative densitometry for HA, SMARCA4 and SMARCC1 performed on purified WT and mutant cBAF complexes used for in vitro nucleosome remodeling assays. Panel M provides TapeStation analysis of REAA-based in vitro nucleosome remodeling assays shown in FIG. 1, panel G.

[0034] FIG. 9 shows ARID I A / B IDRs and ARID DNA binding promote localized condensation of cBAF in vitro and in cells. Panel A provides a schematic of eGFP-tagged ARID 1 A WT and mutant constructs. Panel B (left) shows silver stain of purified cBAF complexes containing WT or mutant eGFP-tagged ARID 1 A, purified from HEK293T AARID1A / B cells; (middle) provides an immunoblot of ATPase and core cBAF subunits using purified complexes; (right) provides quantitative densitometry of HA (ARID 1 A), SMARCA4, SMARCC1, and SMARCE1. Panel C provides representative images from in vitro droplet assays performed across a range of concentrations for eGFP-tagged complexes. Scale bar, 20 pm. Panel D (left) provides representative images from in vitro droplet assay, scale bar. 20 pm; (right) provides quantification of droplet area coverage for WT and mutant ARID1 A- containing eGFP-tagged cBAF complexes at 2, 0.66, 0.22 and 0.074 pM, alone or with addition of 100 nM DNA, nucleosomes, or RNA. Error bars represent standard deviation of 8 fields of view in each condition. Panel E provides representative images of in vitro droplet assays with 100 nM DNA added in indicated conditions. Strings of droplets from in in WT ARID 1 A + DNA condition only. Panel F (top) provides an immunoblot for input and (bottom) anti-HA IP in AN3CA rescued with HA-tagged ARID1A WT or mutant eGFP tagged variants. Panel G provides quantitative densitometry of subunit protein levels across conditions from Panel F (input and IP). Panel H provides density sedimentation analysis using nuclear extracts of AN3CA cells rescued with eGFP-tagged ARID1 A showing that the eGFP tag does not disrupt complex formation. Panel I provides representative images (left) and saturation concentration (right) of Control and MG-132 proteasome inhibitor treated AN3CA cells expressing ARID1A WT-eGFP or mutants. Panel J provides representative images, puncta count and size of eGFP-tagged ARID1B WT and DBDmutin AN3CA cells (n=25 cells each). ****p<0.0001 by unpaired t-test. Panel K provides FRAP curves, half time of recover}’ (Tl / 2) and Immobile fraction quantification for ARIDlA-eGFP (left) and ARIDlB-eGFP (right) containing cBAF complexes. Scale bars, 10 pm. Error bars represent standard deviation, n = 3 biological trials containing 15 cells each. P-values were calculated using an unpaired t-test. ns, not statistically significant. Panel L provides confocal imaging of condensates using anti-ARIDlA antibody in CRL-7250. MCF10A, MDA-MB-231, and for anti-ARIDlB in primary rat neurons, alongside Hoecht stain. Scale bars, 10 pm. Panel M provides representative images of one nucleus expressing each construct without (- light) and with (+ light) light-induced oligomerization through the Corelet system. Scale bars. 10 pm. Panel N provides Corelet system phase diagrams of indicated ARID 1 A constructs in U2OS cells. Panels O-P provide Corelet system Phase Diagrams of indicated ARID1B constructs in U2OS cells. Panel Q provides representative images for repeated light activation-deactivation cycles of indicated constructs in Corelet System in U2OS cells, and subsequent Pearson Correlation Coefficient (PCC) of droplet nuclear localization. Scale bars. 10 pm. Panel R provides quantification of PCC across three activation-deactivation cycles for indicated constructs. Error bars represent standard deviation, n = 32, 20, 20 cells. P-values were calculated using a one-way ANOVA test.

[0035] FIG. 10 shows ARID1A IDRs and ARID domain mediate cBAF occupancy, DNA accessibility and gene expression in cells. Panel A provides metaplots of HA (ARID 1 A), SMARCA4, SMARCC1 , H3K27ac occupancy and DNA accessibility (AT AC) at Clusters 2, 3, and 4 sites from FIG. 3, panel A across Empty and ARID1 A WT / mutant conditions. Panel B provides Venn diagrams indicating overlap between accessible sites gained in ARID1A WT (red) and DBDmut. AIDR1, CBR mutant conditions (green, purple, brown, respectively) in AN3CA cells. Panel C provides Principal Component Analysis (PCA) of the ATAC-Seq sites in AN3CA cells across Empty control and ARID 1 A WT / mutant conditions. Panel D provides RNA-Seq PCA in AN3CA cells across Empty control and ARID 1 A WT / mutant conditions. Panel E shows overlap of gained AT AC sites shared between WT, DBDmut, AIDR1, CBR conditions. Panel F shows overlap of WT-only gained ATAC sites not overlapping with DBDmut, AIDR1, CBR mutant gained sites. Panel G shows cBAF (ARID 1 A, SMARCA4, SMARCC1) and H3K27ac chromatin occupancy (CUT&Tag) at gained DNA-accessible sites (ATAC-Seq) in AN3CA cells expressing ARID1 A WT compared to empty control. Panel H provides transcription factor motif enrichment analysis (HOMER) of Cluster 2 sites from PanelG. Panel I provides volcano plots reflecting gene expression changes (RNA-Seq) between conditions indicated. Red and blue dots indicate genes up- and down-regulated with an adjusted p-value cutoff of 0.01 and a log2 fold change threshold of 1. Panel J show s up- and down- regulated genes (DEG count) across comparisons indicated. Panel K shows expression of genes nearest to cBAF-occupied sites in Clusters 1-4 from FIG. 3, panel A across ARID1A WT or mutant conditions relative to empty control. Panel L shows differentially expressed genes (left) with differential cBAF target sites (ARID 1A / SM ARC A4 sites) within 1 kB of genes (right) in ARID 1 A WT versus empty vector conditions. Panel L provides metascape enrichment analysis performed on genes closest to Clusters 2 / 3 sites from FIG. 3, panel A. Epithelial cell differentiation term highlighted in red.

[0036] FIG. 11 shows the ARID 1 A IDRs and DNA-binding ARID domain facilitate interactions with transcription factors and transcriptional machinery. Panel A provides a schematic of HA-tagged ARID1A WT or mutant variants fused to the biotin ligase TurboID (TbID). Panel B provides an immunoblot using nuclear extract from AN3CA cells expressing TbID fused to ARID 1 A WT or mutant variants labeled with 50 pM Biotin for 10 min. Panel C provides a schematic for TbID-TMT mass spectrometry experiments. Panel D provides metascape enrichment analysis of downregulated biotinylated hits in AIDR1 and CBR compared to WT. Panel E (left) provides a heatmap of normalized TMT peptide signal of BAF complex subunits; (right) provides an immunoblot of pan-BAF, cBAF, PBAF and ncBAF specific subunits in AN3CA cells expressing ARID 1 A WT or mutant TbID fusions. Panel F provides normalized TMT signal heatmaps of Clusters 2 / 3 motif enriched transcription factors from FIG. 3, panel A, Mediator complex and RNA Polymerase Il-associated proteins across the same conditions as in I. Panel G provides an immunoblot of indicated proteins in AN3CA cells expressing ARID 1 A WT or mutant TbID fusions. Panel H provides coimmunofluorescence studies performed on AN3CA cells rescued with ARID 1 A WT and mutant variants and visualized for cBAF (eGFP) and SMARCC1. Arrows indicate colocalization. Scale bars, 10 pm. Panel I provides an immunoblot for p300 in AN3CA cells expressing Empty control, ARID1A WT-eGFP or mutants. Panel J provides molecular function (GO) enrichment analysis of proteins overlapping between IP-Mass Spectrometry and proximity labelling experiments. Panel K provides correlation of proximity labeling (TurboID) log2 fold change and changes in DNA accessibility (log2 fold change) at Cluster 2 and Cluster 3 sites from FIG. 3, panel A across human transcription factors. Factors corresponding to highly enriched motifs (from FIG. 3, panel G) are highlighted in red. Panel L provides mRNAexpression (CPM) of key TF genes across ARID1A rescue conditions in AN3CA cells shown in FIG. 4, panel H.

[0037] FIG. 12 shows IDRs of FUS and DDX4 rescue cBAF condensation in cells but not chromatin occupancy, DNA accessibility and gene expression. Panels A-B provide an immunoblot for input and anti-HA IP experiments in AN3CA cells expressing HA-tagged or HA- and eGFP dual tagged ARID1 A WT or FUS- and DDX- fusion mutant variants along with densitometry measurements. Panel C shows the saturation concentration of WT and mutant variants of ARID1A in AN3CA cells calculated using calibrated fluorescence imaging. PS: Phase Separation. Panel D provides metaplots of HA (ARID1A), SMARCA4, SMARCC1, H3K27ac occupancy and DNA accessibility (ATAC) at Clusters 2 and 3 from FIG. 3, panel A, across Empty control, ARID 1 A WT and mutant conditions. Panel E provides PC A of the ATAC-Seq sites in AN3CA cells across Empty’ control, ARID1A WT and mutants. Panel F provides volcano plots comparing global gene expression profiles (RNA-Seq) of AN3CA cells expressing ARID1A mutants (AIDR1, FUSIDR, DDX4IDR) compared to ARID1A WT with an adjusted p- value cutoff of 0.01 and a log2 fold change threshold of 1. Panel G provides normalized TMT-MS signal heatmaps of detected cBAF subunits from IP-MS experiments in AN3CA cells. Panel H (top) provides a schematic of ARID1 A WT and mutant TbID fusions; (bottom) provides an immunoblot using nuclear extract from AN3CA cells expressing ARID1A WT-TblD or mutant variants labeled with 50 pM Biotin for 10 min. Panel I provides volcano plots comparing biotinylated protein levels of cBAF in cells expressing ARID 1 A variants versus WT TbID fusions. Panel J provides normalized TMT-MS signal heatmaps of detected Mediator complex subunits, RNA Polymerase II associated proteins and Clusters 2,3 motif enriched transcription factors from FIG. 3, panel A for indicated conditions. Panel K provides immunofluorescence analysis of ARID1A and SMARCC1 in AN3CA cells expressing ARID 1 A WT-eGFP or indicated mutants. Scale bars, 10 pm. Panel L provides an immunoblot for p300 in AN3CA cells expressing ARID1 A WT-eGFP or indicated mutants.

[0038] FIG. 13 shows ARID 1 A IDR mutants affecting phase separation or protein partner interactions result in convergent defects in cBAF complex chromatin targeting and activity. Panel A provides a schematic for the IDRs across mSWI / SNF subunits. Panel B provides IDR sequence patterning conservation of ARID 1 A orthologs across eukaryotes. Panel C provides FRAP studies for WT and AQG scramble mutant sequences; half time of recovery, and mobility measurements are shown. Panel D provides Distance- to-TSS for Clusters A-C from FIG. 6, panel H. Panel E provides HOMER TF motif analysis from Cluster B sites in FIG. 6,panel H. Panel F provides PCA analyses of SMARCA4 sites, ATAC-seq, and RNA-seq datasets across empty and ARID1A WT or mutant variant conditions in AN3CA cells. Panel G provides metaplots for HA (ARID 1 A) and ATAC-Seq for Clusters A-C from FIG. 6, panel H. Panel H provides volcano plots reflecting changes in gene expression; significantly downregulated and upregulated genes are highlighted in blue and red, respectively. Panel I provides sequence grammar schematics of FUS and DDX4 IDRs, showing their abundance of aromatic residues but lack of blocky AQG patches. Panel J provides PCC correlation between ARID1 A and p300 across WT and AQGscram conditions.

[0039] FIG. 14 shows block deletion and NDD-associated mutations impact cBAF condensation and function genome-wide. Panel A provides NDD- and cancer-associated mutations (DECIPHER) plotted across distinct blocks within IDR1 of ARID1AB. Types of mutations are indicated in legend. Panel B (left) provides FRAP curves for ARID1B-WT- eGFP, Block 9 del, Block 13 del, and NDD-associated mutant- containing cBAF complexes. Error bars represent standard deviation, n = 3 biological trials containing 15 cells each. P-values were calculated using a one-way ANOVA test, ns, not statistically significant; (right) provides immobile fraction and T1 / 2 for the variants. Panel C (left) provides a schematic of experiment to test the impact of SMARCA2 / 4 ATPase inhibiton on cBAF condensate formation; (middle) provides live cell imaging of eGFP-tagged cBAF complexes containing WT ARID1A across indicated conditions. Scale bars, 10 pm. Panel D provides clustered heatmaps reflecting chromatin occupancy of cBAF complexes marked by HA(ARIDIB) and SMARCA4 and DNA accessibility (ATAC-Seq) at cBAF occupied sites. Panel E provides Distance-to-TSS plots for Clusters defined in Panel D. Panel F provides PCA of ATAC-Seq peaks for the ARID1B WT and mutant variants. Clustered heatmaps reflecting chromatin occupancy of cBAF complexes marked by HA(ARIDIB) and SMARCA4. G. PCA of ATAC-Seq peaks for the ARID1A / B variants tested.

[0040] FIG. 15 shows time lapse live imaging of AN3CA cells expressing eGFP-tagged ARID 1 A and ARID IB WT and DBDmutconstructs.

[0041] FIG. 16 shows ARID1A / B IDR features dictate cBAF condensation, genomic targeting, and partner recruitment.DETAILED DESCRIPTION OF THE INVENTION

[0042] Detailed descriptions of one or more embodiments are provided herein. It is to be understood, however, that the present invention may be embodied in various forms. Therefore,specific details disclosed herein are not to be interpreted as limiting, but rather as a basis for the claims and as a representative basis for teaching one skilled in the art to employ the present invention in any appropriate manner.

[0043] The singular forms ’‘a”, “an” and “the” include plural reference unless the context clearly dictates otherwise. The use of the word “a” or “an” when used in conjunction with the term “comprising” in the claims and / or the specification may mean “one,” but it is also consistent with the meaning of “one or more,” “at least one.” and “one or more than one.”

[0044] Wherever any of the phrases “for example,” “such as,” “including” and the like are used herein, the phrase “and without limitation” is understood to follow unless explicitly stated otherwise. Similarly, “an example,” “exemplary” and the like are understood to be nonlimiting.

[0045] The term “substantially” allows for deviations from the descriptor that do not negatively impact the intended purpose. Descriptive terms are understood to be modified by the term “substantially” even if the word “substantially” is not explicitly recited.

[0046] The terms “comprising” and “including” and “having” and “involving” (and similarly “comprises”, “includes,” “has,” and “involves”) and the like are used interchangeably and have the same meaning. Specifically, each of the terms is defined consistent with the common United States patent law definition of “comprising” and is therefore interpreted to be an open term meaning “at least the following,” and is also interpreted not to exclude additional features, limitations, aspects, etc. Thus, for example, “a process involving steps a, b, and c” means that the process includes at least steps a, b and c. Wherever the terms “a” or “an” are used, “one or more” is understood, unless such interpretation is nonsensical in context.

[0047] The term “about” is used herein to mean approximately, roughly, around, or in the region of. When the term “about” is used in conj unction with a numerical range, it modifies that range by extending the boundaries above and below the numerical values set forth. In general, the term “about” is used herein to modify a numerical value above and below the stated value by a variance of 20 percent up or down (higher or lower).

[0048] The term “in vivo” can refer to an event that takes place in a subject's body.

[0049] The term “in vitro” can refer to an event that takes places outside of a subject's body.

[0050] The term “ex vivo” can refer to outside a living subject. Examples of ex vivo cell populations include in vitro cell cultures and biological samples such as fluid or tissue samples from humans or animals. Such samples can be obtained by methods well known in the art. Exemplary biological fluid samples include blood, cerebrospinal fluid, urine, saliva.Exemplary tissue samples include tumors and biopsies thereof. In this context, the compounds can be in numerous applications, both therapeutic and experimental.

[0051] Aspects of the invention are drawn to an in vitro screening method to identify a SWI / SNF modulating compound. In embodiments, the method comprises incubating a candidate compound with a protein or protein complex, such as a protein complex comprising an ARID 1 A and / or ARID IB intrinsically disordered region (IDR), and detecting at least one IDR function, wherein modulation of the IDR function is indicative of a SWI / SNF modulating compound.

[0052] As used herein, the term “screening method” can refer to a method used for characterizing or selecting compounds from a collection of compounds based on the activities of the compounds. For example, an embodiment can be used to identify a SWI / SNF modulating compound. The screening process can allow for a deeper understanding of cell pathways that can be disrupted and / or affected by a treatment. In embodiments, the compound identified by the screening method described herein can inhibit the binding of a transcription factor. In embodiments, the compound identified by the screening method can inhibit residues required for coordinating proper targeting of the SWI / SNF complex. In embodiments, the compound identified by the screening method can alter SWI / SNF condensation or phase separation. In embodiments, the compound identified by the screening method can break ARID 1 A or ARID1B binding to other factors, such as transcription factors.

[0053] The term "compound," such as a SWI / SNF modulating compound, can refer to a substance that is composed of two or more elements that are bonded together. Not limiting examples of “compounds” can include small organic molecules, natural product extract libraries, organic molecules, inorganic molecules, including but not limited to chemicals, metals and organometallic molecules.

[0054] In embodiments, the compound can comprise a candidate compound. The term “candidate” or “candidate compound” can refer to an experimental candidate used in a screening process to identify activity, non-activity, or other modulation of a particularized biological target or pathway, such as binding to an ARID 1 A IDR, ARID IB IDR or fragment thereof. In embodiments, a candidate compound is specific for an ARID 1 A IDR, ARID IB IDR, or fragment thereof. In embodiments, a candidate compound alters protein behavior, such as binding, configuration and / or engagement with other key factors, such as transcription factors. In embodiments, a candidate compound is cell permeable. In embodiments, a candidate compound can alter where mSWI / SNF complexes bind. In embodiments, a candidatecompound plays a role in gene regulation. In embodiments, a candidate compound can cause for the destabilization or the stabilization an IDR of ARID 1 A versus and IDR of ARID IB. For example, the candidate compound can comprise small organic molecules, natural product extract libraries, organic molecules, inorganic molecules, including but not limited to chemicals, metals and organometallic molecules.

[0055] In embodiments, the candidate compound can comprise a small molecule. The term "small molecule" can refer to molecules that are less than about 1000 molecular weight or less than about 500 molecular weight. In one embodiment, small molecules do not exclusively comprise peptide bonds. In another embodiment, small molecules are not oligomeric. Exemplary small molecule compounds which can be screened for activity include, but are not limited to, peptides, peptidomimetics. nucleic acids, carbohydrates, small organic molecules (e.g., polyketides) (Cane et al. (1998) Science 282:63), and natural product extract libraries. In another embodiment, the compounds are small, organic non-peptidic compounds. In a further embodiment, a small molecule is not biosynthetic.

[0056] Embodiments as described herein can refer to a high-throughput method for identifying a candidate compound. The term “high-throughput” or “high-throughput screening” can refer to a discovery process that allows automated testing of large numbers of chemical and / or biological compounds for a specific biological target, such as an ARID 1 A or ARID IB IDR or fragment thereof.

[0057] The term “modulating”, such as in a SWI / SNF modulating compound, can refer to regulating or adjusting the degree of activity of a process or the degree of an effect. “Modulating” can include activation, inhibition, degradation, amplification, attenuation, and suppression. For example, modulating the activity of the SWI / SNF complex can refer to altering the level or activity of the SWI / SNF complex, component thereof, or a related downstream effect. The activity' level of a SWI / SNF complex can be measured using any method known in the art.

[0058] In embodiments, the compound is a SWI / SNF modulating compound. A “modulator” or “modulating compound” can refer to an compound that agonizes (activates or enhances) or antagonizes (inhibits or reduces) the function of a biological target. In embodiments, the SWI / SNF complex modulator can comprise small organic molecules, natural product extract libraries, organic molecules, inorganic molecules, including but not limited to chemicals, metals and organometallic molecules. In some embodiments, the SWI / SNF modulating compound can alter SWI / SNF condensation or phase separation.

[0059] In embodiments, the SWI / SNF modulating compound can inhibit function. The term “inhibiting” can refer to decreasing, limiting, and / or blocking a certain action, function, or interaction. In some embodiments, the SWI / SNF modulating compound can inhibit the binding of a transcription factor. In some embodiments, the SWI / SNF modulating compound can inhibit residues required for coordinating proper targeting of the mSWI / SNF complex.

[0060] In embodiments, the protein or protein complex comprising an ARID 1 A IDR, ARID1B IDR. or fragment thereof, is selected from the group consisting of a purified polypeptide, a chromatin remodeling complex, a DNA-binding protein, an epitope-tagged peptide, a core binding pepetide, or any combination thereof.

[0061] In embodiments, the chromatin remodeling complex can comprise the mammalian SWI / SNF chromatin remodeling complex. The term "SWI / SNF complex" can refer to SWitch / Sucrose Non-Fermentable, a nucleosome remodeling complex found in both eukaryotes and prokaryotes (Neigebom Carlson (1984) Genetics 108:845-858; Stem et al. (1984) J Mai. Biol. 178:853-868). The SWI / SNF complex was discovered in the yeast, Saccharomyces cerevisiae. named after yeast mating types switching (SWI) and sucrose nonfermenting (SNF) pathways (Workman and Kingston (1998) Annu Rev Biochem. 67:545- 579; Sudarsanam and Winston (2000) Trends Genet. 16:345-351). It is a group of proteins comprising, at least, SWII, SWI2 / SNF2, SWI3, SW15, and SW16, as well as other polypeptides (Pazin and Kadonaga (1997) Cell 88:737-740). A genetic screening for suppressive mutations of the SWI / SNF phenotypes identified different histones and chromatin components, indicating that these proteins were involved in histone binding and chromatin organization (Winston and Carlson (1992) Trends Genet. 8:387-391). Biochemical purification of the SWI / SNF2p in S. cerevisiae demonstrated that this protein was part of a complex containing an additional 11 polypeptides, with a combined molecular weight over 1. 5 MDa. The SWI / SNF complex contains the ATPase Swi2 / Snf2p, two actin-related proteins (Arp7p and Arp9) and other subunits involved in DNA and protein-protein interactions. The purified SWI / SNF complex alters the nucleosome structure in an ATPIO dependent manner (Workman and Kingston (1998), supra; Vignali et al. (2000) Mai Cell Biol. 20: 1899-1910). The structures of the SWI / SNF and RSC complexes are highly conserved but not identical, reflecting an increasing complexity of chromatin (e.g., an increased genome size, the presence of DNA methylation, and more complex genetic organization) through evolution.

[0062] For this reason, the SWI / SNF complex in higher eukaryotes maintains core components, but also substitute or add on other components with more specialized or tissue-specific domains. Yeast contains two distinct and similar remodeling complexes, SWI / SNF and RSC (Remodeling the Structure of Chromatin). In Drosophila, the two complexes are called BAP (Brahma Associated Protein) and PBAP (Polybromo-associated BAP) complexes. The human analogs are BAF (Brgl Associated Factors, or SWI / SNF-A) and PBAF (Polybromo-associated BAF, or SWI / SNF-B). The BAF complex comprises, at least, BAF250A (ARID1A), BAF250B (ARID1B), BAF57 (SMARCEI), BAF190 / BRM (SMARCA2), BAF47 (SMARCBI), BAF53A (ACTL6A). BRG1 / BAF190 (SMARCA4). BAF 155 (SMARCCI), and BAFI 70 (SMARCC2). The PBAF complex comprises, at last, BAF200 (ARID2), BAF180 (PBRMI), BRD7, BAF45A (PHFIO), BRG1 / BAF190 (SMARCA4), BAF155 (SMARCCI), and BAF 170 (SMARCC2).

[0063] Combinatorial association of SWI / SNF subunits can in principle give rise to hundreds of distinct complexes, although the exact number has yet to be determined (Wu et al. (2009), supra). Genetic evidence indicates that distinct subunit configurations of SWI / SNF are equipped to perform specialized functions. As an example, SWI / SNF contains one of two ATPase subunits, BRG1 or BRM / SMARCA2. which share 75% amino acid sequence identity (Khavari et al. (1993) Nature 366: 170-174). While in certain cell types BRG 1 and BRM can compensate for loss of the other subunit, in other contexts these two ATPases perform divergent functions (Strobeck et al. (2002) J Biol Chem. 277:4782-4789; Hoffman et al. (2014) Proc Natl Acad Sci US A. I l l :3128-3133). In some cell types, BRG 1 and BRM can even functionally oppose one another to regulate differentiation (Flowers et al. (2009) J Biol Chem. 284: 10067-10075). The functional specificity of BRG 1 and BRM has been linked to sequence variations near their N-terminus, which have different interaction specificities fortranscription factors (Kadam and Emerson (2003) Mai Cell. 11 :377-389). Another example of paralogous subunits that form mutually exclusive SWI / SNF complexes are ARID1A / BAF250A, ARID1B / BAF250B, and ARID2 / BAF200. ARID1A and ARID1B share 60% sequence identity, but yet can perform opposing functions in regulating the cell cycle, with MYC being an important downstream target of each paralog (Nagi et al. (2007) EMBO J. 26:752-763). ARID2 has diverged considerably from ARID1A / ARID1B and exists in a unique SWI / SNF assembly known as PBAF (or SWI / SNF-B), which contains several unique subunits not found in ARIDlA / B-containing complexes. The composition of SWI / SNF can also be dynamically reconfigured during cell fate transitions through cell type-specific expression patterns of certain subunits. For example, BAF53A / ACTL6A is repressed and replaced by BAF53B / ACTL6B during neuronal differentiation, a switch that is essential for proper neuronal functions in vivo(Lessard et al. (2007) Neuron 55:201-215). These studies stress that SWI / SNF in fact represents a collection of multi-subunit complexes whose integrated functions control diverse cellular processes, which is also incorporated in the scope of definitions of the disclosure. Two recently published meta-analyses of cancer genome sequencing data estimate that nearly 20% of human cancers harbor mutations in one ( or more) of the genes encoding SWI / SNF (Kadoch et al. (2013) Nat Genet. 45:592-601; Shain and Pollack (2013) PLoS One. 8:e55119). Such mutations can be loss-of-function, implicating SWI / SNF as a maj or tumor suppressor in diverse cancers. Specific SWI / SNF gene mutations can be linked to a specific subset of cancer lineages: SNF5 is mutated in malignant rhabdoid tumors (MRT), PBRM1 / BAF 180 is frequently inactivated in renal carcinoma, and BRG 1 is mutated in non-small cell lung cancer (NSCLC) and several other cancers. In the disclosure, the scope of SWI / SNF complex" can cover at least one fraction or the whole complex (e.g., some or all subunit proteins / other components), in the human BAF / PBAF forms or their homologs / orthologs in other species (e.g., the yeast and drosophila forms described herein). Preferably, a "SWI / SNF complex" described herein contains at least part of the full complex bio-functionality, such as binding to other subumts / components, binding to DNA / histone, catalyzing ATP, promoting chromatin remodeling, etc.

[0064] The term “BAF complex" can refer to at least one ty pe of mammalian SWI / SNF complexes. Its nucleosome remodeling activity can be reconstituted with a set of four core subunits (BRG1 / SMARCA4, SNF5 / SMARCB1, BAF155 / SMARCC1, and BAF170 / SMARCC2), which have orthologs in the yeast complex (Phelan et al. (1999) Mol Cell. 3:247-253). However, mammalian SWI / SNF contains several subunits not found in the yeast counterpart, which can provide interaction surfaces for chromatin (e.g. acetyl-lysine recognition by bromodomains) or transcription factors and thus contribute to the genomic targeting ofthe complex (Wang et al. (1996) EMBO J. 15:5370-5382; Wang et al. (1996) Genes Dev. 10:2117-2130; Nie et al. (2000) ). A key attribute of mammalian SWI / SNF is the heterogeneity' of subunit configurations that can exist in different tissues and even in a single cell type (e.g., as BAF, PBAF, neural progenitor BAF (npBAF), neuron BAF (nBAF), embryonic stem cell BAF (esBAF), etc.). In some embodiments, the BAF complex described herein refers to one type of mammalian SWI / SNF complexes, which is different from PBAF complexes.

[0065] The SWI / SNF modulating compounds of this disclosure can be used to impact the SWI / SNF complex, and find use in the treatment of various diseases, disorders, or conditionsrelated to mutations in or malfunction of SWI / SNF. In embodiments, the SWI / SNF modulating compounds can inhibit residues required for coordinating proper targeting of the mSWI / SNF complex. In embodiments, the SWI / SNF modulating compounds can alter mSWI / SNF condensation or phase separation. Human SWI / SNF (BAF) complexes are a diverse family of ATP-dependent chromatin remodelers that exhibit combinatorial specificity to regulate specific genetic programs. Due to the many diverse functions of the BAF complex, ranging from opposition of Polycomb-repressed genes to direct interaction with Topoisomerases and mediated genome stability, inhibitors of the SWI / SNF complex are of broad utility in human disease. The SWI / SNF or BAF complex is mutated in roughly 20% of human cancers, and associated with a host of neurological diseases including Coffm-Siris syndrome, autism, Nicholaides-Baraitser Syndrome, Kleefstra Syndrome, among others. Further, the SWI / SNF complex has been implicated in HIV-1 Tat mediated transcription, indicating a potential use for reversing HIV latency. Non-limiting examples of current interest include the treatment of aneurodevelopmental disorder and the treatment of cancer.

[0066] In embodiments, the candidate compound is specific for an IDR of ARID 1 A or fragment thereof. The term "ARID1 A" or "BAF250A"can refer to AT-rich interactive domaincontaining protein IA, a subunit of the SWI / SNF complex, which can be find in BAF but not PBAF complex. In humans there are two BAF250 isoforms, BAF250A / ARID1A and BAF250B / ARID1B. They can be E3 ubiquitin ligases that target histone H2B (Li et al. (2010) Mai. Cell. Biol. 30: 1673-1688). ARID1A is highly expressed in the spleen, thymus, prostate, testes, ovaries, small intestine, colon and peripheral leukocytes. ARID 1 A is involved in transcriptional activation and repres-sion of select genes by chromatin remodeling. It is also involved in vitamin D-coupled transcription regulation by associating with the WINAC complex, a chromatin-remod-eling complex recruited by vitamin D receptor. ARID 1 A belongs to the neural progenitors-specific chromatin remod-eling (npBAF) and the neuron-specific chromatin remodel-ing (nBAF) complexes, which are involved in switching developing neurons from stem / progenitors to post-mitotic chromatin remodeling as they exit the cell cycle and become committed to their adult state. ARID1A also plays key roles in maintaining embryonic stem cell pluripotency and in cardiac development and function (Lei et al. (2012) J. Biol. Chem. 287:24255-24262; Gao et al. (2008) Proc. Natl. Acad. Sci. U.S.A. 105:6656- 6661). Loss of BAF250a expression was seen in 42% of the ovarian clear cell carcinoma samples and 21 % of the endometrioid carcinoma samples, compared with just 1 % of the highgrade serous carcinoma samples. ARID 1 A deficiency also impairs the DNA damagecheckpoint and sensitizes cells to PARP inhibitors (Shen et al. (2015) Cancer Discov. 5:752- 767). Human ARID1 A protein has 2285 amino acids and a molecular mass of 242045 Da, with at least a DNA-binding domain that can specifically bind an AT-rich DNA sequence, recognized by a SWI / SNF complex at the beta-globin locus, and a C-terminus domain for glucocorticoid receptor-depen-dent transcriptional activation. ARID IA has been shown to interact with proteins such as SMARCB1 / BAF47 (Kato et al. (2002) J. Biol. Chem. 277:5498- 505; Wang et al. (1996) EMBO J. 15:5370-5382) and SMARCA4 / BRG1 (Wang et al. (1996). supra; Zhao et al. (1998) Cell 95:625-636), etc.

[0067] In embodiments, the candidate compound is specific for an IDR of ARID1B or fragment thereof. The term "ARID1B" or "BAF250B" can refer to AT-rich interactive domaincontaining protein IB. a subunit of the SWI / SNF complex, which can be find in BAF but not PBAF complex. ARID IB and ARID I A are alternative and mutually exclusive ARID-subunits of the SWI / SNF com-pl ex. Germline mutations in ARID IB are associated with Coffin-Siris syndrome (Tsurusaki et al. (2012) Nat. Genet. 44:376-378; Santen et al. (2012) Nat. Genet. 44:379-380). Somatic mutations in ARID1B are associated with several cancer subtypes, indicating that it is a tumor suppressor gene (Shai and Pollack (2013) PLoS ONE 8:e55119; Sausen et al. (2013) Nat. Genet. 45:12-17; Shain et al. (2012) Proc. Natl. Acad. Sci. U.S.A. 109:E252-E259; Fujimoto et al. (2012) Nat. Genet. 44:760-764). Human ARID IA protein has 2236 amino acids and a molecular mass of 236123 Da, with at least a DNA-binding domain that can specifically bind an AT-rich DNA sequence, recognized by a SWI / SNF complex at the beta-globin locus, and a C-terminus domain for glucocorticoid receptor-dependent transcriptional activa-tion. ARID1B has been shown to interact with SMARCA4 / BRG1 (Hurlstone et al. (2002) Biochem. J. 364:255-264; Inoue et al. (2002) J. Biol. Chem. 277:41674-41685 and SMARCA2 / BRM (Inoue et al. (2002), supra).

[0068] The terms “intrinsically disordered region” or “IDR” can refer to a polypeptide segment that does not contain sufficient hydrophobic amino acids to mediate co-operative folding. IDRs contain a higher proportion of polar or charged amino acids, and thus, lack a unique three-dimensional structure entirely or in parts in their native state. The term “intrinsically disordered protein” can refer to any protein insofar as it has an intrinsically disordered region. When the intrinsically disordered protein is used for screening, the intrinsically disordered protein can be used in a naturally occurring state, or the disordered portion can be cut out and used. The disordered portion can take any form insofar as an interaction with a test substance can be detected, and the disordered sequence portion and aportion having a domain structure can be bound to each other. In embodiments, the IDR can be a recombinant IDR. The term "‘recombinant” can refer to a substance, such as DNA, proteins, or cells, that are made by combining genetic material from two different sources.

[0069] In embodiments, the candidate compound binds to the IDR. The term “binds” or “binding” can refer to an attractive interaction between two molecules that results in a stable association in which the molecules are in close proximity to each other. In embodiments, the candidate compound binds to an ARID ! A or ARID IB IDR or fragment thereof.

[0070] Embodiments as described herein comprise incubating a candidate compound with a protein or protein complex. For example, the protein or protein complex can comprise an ARID1A or ARID1B intrinsically disordered region (IDR) or fragment thereof. The terms “incubate,” “incubating,” and “incubation” can refer to an artificial means of maintaining the controlled environmental conditions necessary for supporting an interaction, such as an interaction between a candidate compound and an ARID 1 A or ARID IB IDR or fragment thereof.

[0071] In embodiments, the candidate compound can be incubated for a period of time. For example, the candidate compound can be incubated for about 10 minutes, 20 minutes. 30 minutes, 40 minutes, 50 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 11 hours, 12 hours, or more than 12 hours.

[0072] The term “protein” can refer to any of a class of nitrogenous organic compounds that comprise large molecules composed of one or more long chains of amino acids and are an essential part of living organisms. A protein can contain various modifications to the amino acid structure such as disulfide bond formation, phosphorylations and glycosylations. A linear chain of amino acid residues can be called a “polypeptide.” A protein contains at least one polypeptide. Short polypeptides, e.g., containing less than 20-30 residues, are sometimes referred to as “peptides.”

[0073] In embodiments, the protein complex can comprise an ARID1A IDR, an ARID1B IDR, or fragment thereof. The term "protein complex" can refer to a combination of two or more proteins formed by one or more interactions between the proteins. For example, the one or more interactions can be specific non-covalent binding interactions. However, covalent bonds can also be present between the interacting partners. For instance, the two interacting partners can be covalently crosslinked so that the protein complex becomes more stable. The protein complex can include and / or be associated with other molecules such as nucleic acid, such as RNA or DNA, or lipids or further cofactors or moieties selected from a metal ions,hormones, second messengers, phosphate, sugars. A "protein complex" encompassed by the invention can also be part of or a unit of a larger physiological protein assembly.

[0074] The term "isolated protein complex" can refer to a protein complex present in a composition or environment that is different from that found in nature, in its native or original cellular or body environment. Preferably, an "isolated protein complex" is separated from at least 50%, more preferably at least 75%, most preferably at least 90% of other naturally coexisting cellular or tissue components. Thus, an "isolated protein complex" can also be a naturally existing protein complex in an artificial preparation or a non-native host cell. An "isolated protein complex" can also be a "purified protein complex", that is, a substantially purified form in a substantially homogenous preparation substantially free of other cellular components, other polypeptides, viral materials, or culture medium, or. when the protein components in the protein complex are chemically synthesized, free of chemical precursors or by-products associated with the chemical synthesis. A "purified protein complex" can refer to a preparation containing preferably at least 75%, more preferably at least 85%, and most preferably at least 95% of a particular protein complex. A "purified protein complex" can be obtained from natural or recombinant host cells or other body samples by standard purification techniques, or by chemical synthesis.

[0075] In embodiments, the candidate compound can comprise a nucleic acid molecule. The term “nucleic acid molecule'’ can refer to DNA molecules and RNA molecules. A nucleic acid molecule can be single-stranded or double-stranded. In embodiments, the nucleic acid molecule is single stranded. In embodiments, the nucleic acid molecule is double-stranded DNA. As used herein, the term “isolated nucleic acid molecule” can refer to a nucleic acid molecule in which the nucleotide sequences are free of other nucleotide sequences, which other sequences can naturally flank the nucleic acid in human genomic DNA. Non-limiting examples of a nucleic acid molecule comprise a siRNA, miRNA, shRNA, antisense RNA, guide RNA (gRNA), single-guide RNA (sgRNA), modified forms thereof, or combination thereof. For example, the nucleic acid molecule can comprise an RNA interfering agent or an antisense oligonucleotide.

[0076] The term “peptide” can refer to a polymer of amino acid residues ranging in length from 2 to about 30, or to about 40, or to about 50, or to about 60, or to about 70 residues. In certain embodiments the peptide ranges in length from about 2, 3, 4, 5, 7, 9, 10, or 11 residues to about 60, 50, 45, 40, 45, 30, 25, 20, or 15 residues. In certain embodiments the peptide ranges in length from about 8, 9, 10, 11, or 12 residues to about 15, 20 or 25 residues. In certain embodiments the amino acid residues comprising the peptide are “L-form” amino acidresidues, however, it is recognized that in various embodiments. “D” amino acids can be incorporated into the peptide. Peptides also include amino acid polymers in which one or more amino acid residues are an artificial chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers.

[0077] In embodiments, the ARID1A IDR, ARID1B IDR, or fragment, thereof can comprise a polypeptide. The terms "polypeptides" and "proteins" are, where applicable, used interchangeably herein. They can be chemically modified, e.g. post-translationally modified. For example, they can be glycosylated or comprise modified amino acid residues. They can also be modified by the addition of a signal sequence to promote their secretion from a cell where the polypeptide does not naturally contain such a sequence. They can be tagged with a tag. They can be tagged with different labels which can assist in identification of the proteins in a protein complex. Polypeptides / proteins for use in the invention can be in a substantially isolated form. It will be understood that the polypeptide / protein can be mixed with carriers or diluents which will not interfere with the intended purpose of the polypeptide and still be regarded as substantially isolated. A polypeptide / protein for use in the invention can also be in a substantially purified form, in which case it will compnse the polypeptide in a preparation in which more than 50%, e.g. more than 80%, 90%, 95% or 99%, by weight of the polypeptide in the preparation is a polypeptide encompassed by the invention.

[0078] The terms "hybrid protein", "hybrid polypeptide," "hybrid peptide", "fusion protein", "fusion polypeptide", and "fusion peptide" are used herein interchangeably and can refer to a non-naturally occurring protein having a specified polypeptide molecule covalently linked to one or more polypeptide molecules that do not naturally link to the specified polypeptide. Thus, a "hybrid protein" can be two naturally occurring proteins or fragments thereof linked together by a covalent linkage. A "hybrid protein" can also refer to a protein formed by covalently linking two artificial polypeptides together. The two or more polypeptide molecules can be linked or fused together by a peptide bond forming a single non-branched polypeptide chain.

[0079] The term "modified polypeptide" or "modified protein complex" can refer to a polypeptide or a protein complex in a composition that is different from that found in nature in its native or original cellular or body environment. The term "modification" as used herein refers to modifications of a protein or protein complex encompassed by the invention including cleavage and addition or removal of a group. In some embodiments, the "modified polypeptide" or "modified protein complex" comprises at least one modification (e.g., fragment, mutation, and the like) or subunit that is modified, i.e., different from that found in nature, in its nativeor original cellular or body environment. The "modified subunit" can be, e.g., a derivative or fragment of the native subunit from which it derives.

[0080] The term "domain" can refer to a functional portion, segment or region of a protein, or polypeptide.

[0081] The terms "derivatives" or "analogs of subunit proteins" or "variants" can refer to molecules comprising regions that are substantially homologous to the subunit proteins, in various embodiments, by at least 30%. 40%. 50%. 60%. 70%. 80%. 90%. 95% or 99% identity over an amino acid sequence of identical size or when compared to an aligned sequence in which the alignment is done by a computer homology' program known in the art, or whose encoding nucleic acid can hybridize to a sequence encoding the component protein under stringent, moderately stringent, or nonstringent conditions. It means a protein which is the outcome of a modification of the naturally occurring protein, by amino acid substitutions, deletions and additions, respectively, which derivatives still exhibit the biological function of the naturally occurring protein although not necessarily to the same degree. The biological function of such proteins can e.g. be examined by suitable available in vitro assays as provided in the invention. The term "functionally active" as used herein refers to a polypeptide, namely a fragment or derivative, having structural, regulatory, or biochemical functions of the protein according to the embodiment of which this polypeptide, namely fragment or derivative is related to.

[0082] In embodiments, the protein or protein complex can comprise an ARID 1 A IDR, an ARID1 B IDR, or fragment thereof. The terms "polypeptide fragment" or "fragment", when used in reference to a reference polypeptide, can refer to a polypeptide in which amino acid residues are deleted as compared to the reference polypeptide itself, but where the remaining amino acid sequence is identical to the corresponding positions in the reference polypeptide. Such deletions can occur at the amino-terminus, internally, or at the carboxyl-terminus of the reference polypeptide, or alternatively both. Fragments can be at least 5, 6, 8 or 10 amino acids long, at least 14 amino acids long, at least 20, 30, 40 or 50 amino acids long, at least 75 amino acids long, or at least 100, 150, 200, 300, 500 or more amino acids long. "Homologous" as used herein, refers to nucleotide sequence similarity' between two regions of the same nucleic acid strand or between regions of two different nucleic acid strands. When a nucleotide residue position in both regions is occupied by the same nucleotide residue, then the regions are homologous at that position. A first region is homologous to a second region if at least one nucleotide residue position of each region is occupied by the same residue. Homology betweentwo regions is expressed in terms of the proportion of nucleotide residue positions of the two regions that are occupied by the same nucleotide residue. By way of example, a region having the nucleotide sequence 5'ATTGCC- 3' and a region having the nucleotide sequence 5'- TATGGC-3' share 50% homology. Preferably, the first region comprises a first portion and the second region comprises a second portion, whereby, at least about 50%, and preferably at least about 75%. at least about 90%. or at least about 95% of the nucleotide residue positions of each of the portions are occupied by the same nucleotide residue. More preferably, nucleotide residue positions of each of the portions are occupied by the same nucleotide residue.

[0083] Embodiments further comprise detecting at least one IDR function, wherein modulation of the IDR function is indicative of a SWI / SNF modulating compound. The terms ■'detection" and "‘detecting” can refer to identifying the presence of specific cellular events indicative of a SWI / SNF modulating compound, such as condensate formation, chromatin localization, IDR protein engagement, and gene expression.

[0084] Non-limiting examples of methods of detecting can comprise a sequencing technique, an immunoassay (e.g., an immunoblot), mass spectrometry, an aggregation assay, ultracentrifugation, size-exclusion chromatography, gel electrophoresis, or dynamic light scattering measurements.

[0085] In embodiments, at least one IDR function can be detected by a sequencing technique. The term “sequencing technique” can refer to a technology used to determine the order of nucleotides in small targeted genomic regions or entire genomes. Non-limiting examples of sequencing techniques can comprise assay for transposase-accessible chromatin with high-throughput sequencing (ATAC-seq), chromatin immunoprecipitation followed by sequencing (ChlP-seq), and RNA-seq.

[0086] In embodiments, at least one IDR function can be detected by an immunoassay. The term “immunoassay” can refer to any assay which detects, identifies, characterizes, quantifies, or otherwise measures an amino acid target in a sample. For example, an amino acid target can be a small peptide, a polypeptide, a protein, or proteinaceous macromolecule. Immunoassays include, for example, direct or competitive binding assays using techniques such as western blots, radioimmunoassays, ELISA (enzyme linked immunosorbent assay), “sandwich” immunoassays, immunoprecipitation assays, fluorescent immunoassays, and protein A immunoassays. Immunoassays use antibodies or antibody fragments, but can also use binding proteins or carrier proteins which bind target molecules with high specificity. In embodiments, an immunoassay can be a chromatin targeting immunoassay. For example, a chromatintargeting immunoassay can comprise CUT&RUN and CUT&TAG. The term “CUT&RUN,” also known as cleavage under targets and release using nuclease, can refer to a chromatin profiling technique that enables high-resolution chromatin mapping and probing. The term “CUT&TAG,” also known as cleavage under targets and tagmentation, can refer to a chromatin profiling technique used to investigate interactions between proteins and DNA, and to identify DNA binding sites for their protein of interest.

[0087] The term “mass spectrometry” can refer to a sensitive and accurate technique for separating and identifying molecules. For example, mass spectrometry can identify protein interactors of cB AF complexes and can determine whether these interactors are lost upon ARID mutation.

[0088] For example, the IDR function detected in embodiments described herein can be selected from the group consisting of (i) condensate formation, (ii) chromatin localization, (iii) IDR protein engagement, and (iv) gene expression.

[0089] In embodiments, IDR function can comprise condensate formation. The term “condensate” can refer to a non-membrane-encapsulated compartment formed by phase separation of one or more proteins and / or other macromolecules such as nucleic acids (including the stages of phase separation).

[0090] In embodiments, IDR function can comprise chromatin localization. The term "chromatin" can refer to the larger-scale nucleoprotein structure comprising the cellular genome. Cellular chromatin comprises nucleic acid, primarily DNA, and protein, including histones and non-histone chromosomal proteins. The majority of eukar otic cellular chromatin exists in the form of nucleosomes, wherein a "nucleosome" core comprises approximately 150 base pairs of DNA associated with an octamer comprising two each of histones H2A, H2B, H3 and H4; and linker DNA (of variable length depending on the organism) extends between nucleosome cores. A molecule of histone HI can be associated with the linker DNA For the purposes of the disclosure, the term "chromatin" is meant to encompass the types of cellular nucleoprotein, both prokaryotic and eukaryotic. Cellular chromatin includes both chromosomal and episomal chromatin.

[0091] In embodiments, IDR function can comprise IDR protein engagement. The term “IDR protein engagement” can refer to an interaction between an IDR, such as an ARID 1 A IDR or ARID IB IDR, with another protein. Non-limiting examples of proteins that interact with an IDR as described herein include transcription factors, co-activators, chaperones, and transcriptional machinery.

[0092] In embodiments, IDR function can comprise gene expression. The terms “gene expression” or “expression" can refer to the process by which information (for example, gene- encoded and / or epigenetic information) is converted into the structures and operating in the cell. For example, “expression" can refer to transcription into a polynucleotide, translation into a polypeptide, or even polynucleotide and / or polypeptide modifications (for example, posttranslational modifications of a polypeptide). Fragments of the transcribed polynucleotide, the translated polypeptide, or polynucleotide and / or polypeptide modifications (for example, posttranslational modification of a polypeptide) can also be regarded as expressed whether they originate from a transcript generated by alternative splicing or a degraded transcript, or from a post-translational processing of the polypeptide, for example, by proteolysis. “Expressed genes” can include those that are transcribed into a polynucleotide as mRNA and then translated into a polypeptide, and also those that are transcribed into RNA but not translated into a polypeptide (for example, transfer and ribosomal RNAs).

[0093] The term “expression profile” can refer to a genomic expression profile. Profiles can be generated by any convenient means for determining a level of a nucleic acid sequence, nonlimiting examples of which include quantitative hybridization of microRNA, labeled microRNA, amplified microRNA, cRNA, quantitative PCR, ELISA for quantitation, and the like. Expression profiles can allow for the analysis of differential gene expression between two samples. In embodiments, a subject or patient sample, e g., cells or collections thereof, e.g., tissues, can be assayed. Samples can be collected by any convenient method, as known in the art.

[0094] In embodiments, the candidate compound can be incubated with a protein or protein complex, such as an ARID1 A IDR, an ARID1B IDR, or fragment thereof, in a reaction vessel. The term “reaction vessel” can refer to any container in which a reaction can occur in accordance with the methods described herein. For example, the reaction vessel can comprise a dish, a tube, or a multi-well plate.

[0095] In embodiments, the in vitro screening method further comprises comparing modulation of the IDR function by the candidate compound to that of a control compound. The term “comparing” can refer to a comparison of corresponding parameters or values. For example, “comparing” can refer to comparing the modulation of the IDR function by the candidate compound to that of a control compound, a reference, or a threshold value.

[0096] A “control” or “control compound” can refer to a compound in which the subjects or reagents of the experiment are treated as in a parallel experiment except for omission of aprocedure, reagent, or variable of the experiment. In some embodiments, the control compound is used as a standard of comparison in evaluating experimental effects. In embodiments, the control compound comprises a positive control, a negative control, or both. A “positive control” can refer to a compound in an experiment that receives a treatment with a known result, and therefore can show a particular change during the experiment. A “negative control” can refer to a compound in an experiment that does not receive treatment and, therefore, does not show any change during the expenment.

[0097] A “reference” or “reference compound” can refer to a compound where a specific readout and well-known response is expected.

[0098] A “threshold value” can refer to the magnitude or intensity that must be exceeded for a certain reaction, phenomenon, result, or condition to occur or be considered relevant. The relevance can depend on context, e.g., it can refer to a positive, reactive or statistically significant relevance.

[0099] Embodiments of the invention can further comprise selecting the candidate compound as a therapeutic agent when the candidate modulates IDR function. An “agent” can refer to any small molecule chemical compound, antibody, nucleic acid molecule, or polypeptide, or fragments thereof, whereas a “therapeutic agent” can refer to an agent for treating a disease or disorder. The term “disease” can refer to an abnormal condition affecting the body of an organism. The term “disorder” can refer to a functional abnormality or disturbance. In embodiments, the disease or disorder can comprise a cancer or a neurodev el opmental disorder.

[0100] As used herein, the terms “gene” and “recombinant gene” can refer to nucleic acid molecules comprising an open reading frame encoding a polypeptide corresponding to a marker of the invention. Such natural allelic variations can result in 1-5% variance in the nucleotide sequence of a given gene. Alternative alleles can be identified by sequencing the gene of interest in a number of different individuals. This can be readily carried out by using hybridization probes to identify the same genetic locus in a variety7of individuals. Any and all such nucleotide variations and resulting amino acid polymorphisms or variations that are the result of natural allelic variation and that do not alter the functional activity are intended to be w ithin the scope of the invention.

[0101] The term "activity" when used in connection with proteins or protein complexes can refer to any physiological or biochemical activities displayed by or associated with a particular protein or protein complex including but not limited to activities exhibited inbiological processes and cellular functions, ability to interact with or bind another molecule or a moiety thereof, binding affinity or specificity to certain molecules, in vitro or in vivo stability (e.g., protein degradation rate, or in the case of protein complexes ability to maintain the form of protein complex), antigenicity and immunogenecity', enzy matic activities, etc. Such activities can be detected or assayed by any of a variety of suitable methods as will be apparent to skilled artisans.

[0102] An "isolated’’ nucleic acid molecule is free of sequences (such as protein-encoding sequences) which naturally flank the nucleic acid (i.e., sequences located at the 5’ and 3' ends of the nucleic acid) in the genomic DNA of the organism from which the nucleic acid is derived. For example, in various embodiments, the isolated nucleic acid molecule can contain less than about 5 kB, 4 kB, 3 kB, 2 kB, 1 kB, 0.5 kB or 0.1 kB of nucleotide sequences which naturally flank the nucleic acid molecule in genomic DNA of the cell from which the nucleic acid is derived. Moreover, an “isolated” nucleic acid molecule, such as a cDNA molecule, can be substantially free of other cellular material or culture medium when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized.

[0103] An "accessible region" can refer to a site in cellular chromatin in which a target site in the nucleic acid can be bound by an exogenous molecule which recognizes the target site. Without wishing to be bound by any particular theory, it is believed that an accessible region is one that is not packaged into a nucleosomal structure. The distinct structure of an accessible region can often be detected by its sensitivity to chemical and enzymatic probes, for example, nucleases.

[0104] A nucleic acid molecule of the invention can be isolated using standard molecular biology techniques and the sequence information in the database records described herein. Using all or a portion of such nucleic acid sequences, nucleic acid molecules of the invention can be isolated using standard hybridization and cloning techniques (e.g., as described in Sambrook et al., ed., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989).

[0105] A nucleic acid molecule of the invention can be amplified using cDNA, mRNA. or genomic DNA as a template and appropriate oligonucleotide primers according to standard PCR amplification techniques. The nucleic acid molecules so amplified can be cloned into an appropriate vector and characterized by DNA sequence analysis. Furthermore,oligonucleotides corresponding to all or a portion of a nucleic acid molecule of the invention can be prepared by standard synthetic techniques, e.g., using an automated DNA synthesizer.

[0106] Unless otherwise specified here within, the terms “antibody” and “antibodies” broadly encompass naturally-occurring forms of antibodies (e.g. IgG, IgA, IgM, IgE) and recombinant antibodies, such as single-chain antibodies, chimeric and humanized antibodies and multi-specific antibodies, as well as fragments and derivatives of the foregoing, which fragments and derivatives have at least an antigenic binding site. Antibody derivatives can comprise a protein or chemical moiety conjugated to an antibody.

[0107] The term “antibody” as used herein also includes an “antigen-binding portion” of an antibody (or simply “antibody portion”). The term “antigen-binding portion”, as used herein, can refer to one or more fragments of an antibody that retain the ability to specifically bind to an antigen (e.g., a biomarker polypeptide or fragment thereof). It has been shown that the antigen-binding function of an antibody can be performed by fragments of a full-length antibody. Examples of binding fragments encompassed within the term “antigen-binding portion” of an antibody include (i) a Fab fragment, a monovalent fragment consisting of the VL, VH. CL and CHI domains; (ii) a F(ab')2 fragment, a bivalent fragment comprising two Fab fragments linked by a disulfide bridge at the hinge region; (iii) a Fd fragment consisting of the VH and CHI domains; (iv) a Fv fragment consisting of the VL and VH domains of a single arm of an antibody, (v) a dAb fragment (Ward el al., (1989) Nature 341 :544-546). which consists of a VH domain; and (vi) an isolated complementarity determining region (CDR). Furthermore, although the two domains of the Fv fragment, VL and VH, are coded for by separate genes, they can be joined, using recombinant methods, by a synthetic linker that allows for them to be made as a single protein chain in which the VL and VH regions pair to form monovalent polypeptides (known as single chain Fv (scFv); see e.g.. Bird et al. (1988) Science 242:423-426; and Huston et al. (1988) Proc. Natl. Acad. Sci. USA 85:5879- 5883; and Osbourn et al. 1998, Nature Biotechnology 16: 778). Such single chain antibodies are also intended to be encompassed within the term “antigen-binding portion” of an antibody. Any VH and VL sequences of specific scFv can be linked to human immunoglobulin constant region cDNA or genomic sequences, in order to generate expression vectors encoding complete IgG polypeptides or other isotypes. VH and VL can also be used in the generation of Fab, Fv or other fragments of immunoglobulins using protein chemistry or recombinant DNA technology. Other forms of single chain antibodies, such as diabodies are also encompassed. Diabodies are bivalent, bispecific antibodies inwhich VH and VL domains are expressed on a single polypeptide chain, but using a linker that is too short to allow for pairing between the two domains on the same chain, thereby forcing the domains to pair with complementary domains of another chain and creating two antigen binding sites (see e.g., Holliger et al. (1993) Proc. Natl. Acad. Sci. U.S.A. 90:6444- 6448; Poljak et al. (1994) Structure 2: 1121-1123).

[0108] Still further, an antibody or antigen-binding portion thereof can be part of larger immunoadhesion polypeptides, formed by covalent or noncovalent association of the antibody or antibody portion with one or more other proteins or peptides. Examples of such immunoadhesion polypeptides include use of the streptavidin core region to make a tetrameric scFv polypeptide (Kipriyanov et al. (1995) Human Antibodies and Hybridomas 6:93-101) and use of a cysteine residue, biomarker peptide and a C-terminal polyhistidine tag to make bivalent and biotinylated scFv polypeptides (Kipriyanov et al. (1994) Mol. Immunol. 31 : 1047-1058). Antibody portions, such as Fab and F(ab')2 fragments, can be prepared from whole antibodies using conventional techniques, such as papain or pepsin digestion, respectively, of whole antibodies. Moreover, antibodies, antibody portions and immunoadhesion polypeptides can be obtained using standard recombinant DNA techniques, as described herein.

[0109] Antibodies can be polyclonal or monoclonal; xenogeneic, allogeneic, or syngeneic; or modified forms thereof (e g. humanized, chimeric, etc.). Antibodies can also be fully human. Antibodies of the invention bind specifically or substantially specifically to a biomarker polypeptide or fragment thereof. The terms ‘’monoclonal antibodies” and “monoclonal antibody composition”, as used herein, can refer to a population of antibody polypeptides that contain only one species of an antigen binding site that can immunoreact with a certain epitope of an antigen, whereas the term “polyclonal antibodies” and “polyclonal antibody composition” can refer to a population of antibody polypeptides that contain multiple species of antigen binding sites that can interact with a certain antigen. A monoclonal antibody composition displays a single binding affinity for a certain antigen with which it immunoreacts. In embodiments, the antibody, or antigen binding fragment thereof, is murine, chimeric, humanized, mosaic, composite, or human.

[0110] Antibodies can be “humanized,” which is intended to include antibodies made by a non-human cell having variable and constant regions which have been altered to more closely resemble antibodies that can be made by a human cell. For example, by altering the non- human antibody amino acid sequence to incorporate amino acids found in human germlineimmunoglobulin sequences. The humanized antibodies of the invention can include amino acid residues not encoded by human germline immunoglobulin sequences (e.g., mutations introduced by random or site-specific mutagenesis in vitro or by somatic mutation in vivo), for example in the CDRs. The term “humanized antibody”, as used herein, also includes antibodies in which CDR sequences derived from the germline of another mammalian species, such as a mouse, have been grafted onto human framework sequences.

[0111] The term “coding region” can refer to regions of a nucleotide sequence comprising codons which are translated into amino acid residues, whereas the term “noncoding region” refers to regions of a nucleotide sequence that are not translated into amino acids (e.g., 5' and 3' untranslated regions).

[0112] The term “complementary” can refer to the broad concept of sequence complementarity between regions of two nucleic acid strands or between two regions of the same nucleic acid strand. It is known that an adenine residue of a first nucleic acid region can form specific hydrogen bonds (“base pairing”) with a residue of a second nucleic acid region which is antiparallel to the first region if the residue is thymine or uracil. Similarly, it is known that a cytosine residue of a first nucleic acid strand can base pair with a residue of a second nucleic acid strand which is antiparallel to the first strand if the residue is guanine. A first region of a nucleic acid is complementary to a second region of the same or a different nucleic acid if, when the two regions are arranged in an antiparallel fashion, at least one nucleotide residue of the first region can base pair with a residue of the second region. The first region comprises a first portion and the second region comprises a second portion, whereby, when the first and second portions are arranged in an antiparallel fashion, at least about 50%. and at least about 75%, at least about 90%, or at least about 95% of the nucleotide residues of the first portion can base pair with nucleotide residues in the second portion. Nucleotide residues of the first portion can base pair with nucleotide residues in the second portion.

[0113] In embodiments, the ARID1 A or ARID1B IDR or fragment thereof is purified from mammalian cells, such as human cells. The terms “purified” or “isolated” can refer to material that is substantially or essentially free from components that normally accompany it in its native state. Purity and homogeneity can be determined using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high performance liquid chromatography. A protein or nucleic acid that is the predominant species in a preparation is substantially purified. In particular, in some embodiments, an isolated nucleic acidcomprising a gene is separated from open reading frames that naturally flank the gene and encode proteins other than the protein encoded by the gene. An isolated antibody is separated from other non-immunoglobulin proteins and from other immunoglobulin proteins with different antigen binding specificities. It can also mean that the nucleic acid or protein is at least 85% pure, at least 95% pure, and in some embodiments, at least 99% pure. For example, the ARID 1 A or ARID IB IDR or fragment thereof is purified from human cells.

[0114] Aspects of the invention are drawn to a plate comprising one or more wells. In embodiments, at least one well comprises a candidate compound as described herein, a protein or protein complex comprising an ARID 1 A or ARID IB IDR or fragment thereof, and medium.

[0115] In embodiments, the medium can comprise a cell culture medium. The terms “medium'’, “cell culture medium”, “culture medium” can refer to a solution containing nutrients that nourish growing cells. In certain embodiments, the culture medium is useful for growing mammalian cells. A culture medium can provide essential and non-essential amino acids, vitamins, energy' sources, lipids, and trace elements required by the cell for minimal growth and / or survival. A culture medium can also contain supplementary components that enhance growth and / or survival above the minimal rate, including, but not limited to, hormones and / or other growth factors, particular ions (such as sodium, chloride, calcium, magnesium, and phosphate), buffers, vitamins, nucleosides or nucleotides, trace elements (inorganic compounds can present at very low final concentrations), amino acids, lipids, and / or glucose or other energy source. In certain embodiments, a medium is advantageously formulated to a pH and salt concentration optimal for cell survival and proliferation.

[0116] In embodiments, the culture medium can comprise “supplementary components”, which can refer to components that enhance growth and / or survival above the minimal rate. Non-limiting examples of supplementary components include hormones and / or other growth factors, ions (such as sodium, chloride, calcium, magnesium, and phosphate), buffers, vitamins, nucleosides or nucleotides, trace elements (inorganic compounds can present at very low final concentrations), amino acids, lipids, and / or glucose or other energy source. In certain embodiments, supplementary components are added to the initial cell culture. In certain embodiments, supplementary components are added after the beginning of the cell culture.

[0117] In embodiments, at least one well of the plate can comprise a plurality of cells. The terms “cells” and “plurality of cells” can refer to a population of cells (i.e., more than one cell). In embodiments, the population can be a pure population comprising one cell type. In otherembodiments, the population can include multiple cell ty pes. Accordingly, there is no limitation on the types of cells that the population of cells can contain.

[0118] Aspects of the invention are also directed towards kits, such as kits comprising plates as described herein. For example, the kit can comprise a plate designed for automated screening described herein.

[0119] In one embodiment, the kit includes (a) a plate as described herein, and optionally (b) informational material. The plate can comprise a candidate compound as described herein, a protein or protein complex comprising an ARID1 A or ARID1B IDR or fragment thereof as described herein, and a medium. The informational material can be descriptive, instructional, marketing or other material that is drawn to the methods described herein and / or the use of the agents for therapeutic benefit. In embodiments, the plate can be designed for automated drug screening.

[0120] The informational material of the kits is not limited in its form. In one embodiment, the informational material can include information about production of the compound, molecular weight of the compound, concentration, date of expiration, batch or production site information, and so forth. In one embodiment, the informational material comprises methods of administering the therapeutic combination composition, e.g., in a suitable dose, dosage form, or mode of administration (e.g., a dose, dosage form, or mode of administration described herein), to treat a subject who has cancer). The information can be provided in a variety of formats, include printed text, computer readable material, video recording, or audio recording, or information that provides a link or address to substantive material.

[0121] Other Embodiments

[0122] While the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.

[0123] The invention will be further described in the following examples, which do not limit the scope of the invention described in the claims.EXAMPLES

[0124] Examples are provided below to facilitate a more complete understanding of the invention. The following examples illustrate the exemplary modes of making and practicing the invention. However, the scope of the invention is not limited to specific embodiments disclosed in these Examples, which are for purposes of illustration only, since alternative methods can be utilized to obtain similar results.EXAMPLE 1

[0125] Mammalian SWI / SNF chromatin remodeling complexes, specifically, ARIDlA / B-nucleated cBAF complexes, are among the most disrupted molecular entities in human disease. However, the biochemical contributions of the ARID1A / B subunits, such as their frequently-mutated N-terminal intrinsically-disordered regions (IDRs), to cBAF function remain poorly understood. Here, we demonstrate that the IDRs of ARID1A / B, coupled with the DNA-binding ARID domain, drive cBAF condensate formation and chromatin localization. We define ARID1A / B IDRs as two-part systems, conferring homot pic cBAF complex interactions (i.e. valence generated by phase separation) and heterotypic complex interactions that establish a highly specific, sequence-encoded protein interaction network within condensates. Both interactions are required for appropriate genome-wide targeting, DNA accessibility, and gene expression, and are disrupted by disease-associated mutations. Taken together, these data establish a functional role for the largest IDRs within a major chromatin remodeler and explain how condensation provides a mechanism through which both genomic localization and functional partner recruitment are achieved.

[0126] We have for the first time defined the role for intrinsically disordered domains within a chromatin remodeler / regulatory complex. Revealing the function of these domains indicates new therapeutic opportunities that can disrupt specific activities and indicate that chemical modulation / binding of specific "motifs" or regions within IDRs can provide highly specialized forms of mSWI / SNF disruption for use in cancer, neurodevelopmental disorders, etc.EXAMPLE 2

[0127] A disordered region controls cBAF activity via condensation and partner recruitment

[0128] Intrinsically disordered regions (IDRs) represent a large percentage of overall nuclear protein content. The prevailing dogma is that IDRs engage in non-specific interactions because they are poorly constrained by evolutionary selection. Here, we demonstrate that condensate formation and heterotypic interactions are distinct and separable features of an IDR within the ARID1A / B subunits of the mSWI / SNF chromatin remodeler, cBAF, and establish distinct "sequence grammars’ underlying each contribution. Condensation is driven by uniformly distributed tyrosine residues, and partner interactions are mediated by non-random blocks rich in alanine, glycine, and glutamine residues. These features concentrate a specific cBAF protein-protein interaction network and are essential for chromatin localization and activity. Importantly, human disease-associated perturbations in ARID1B IDR sequence grammars disrupt cBAF function in cells. Together, these data identify IDR contributions to chromatin remodeling and explain how phase separation provides a mechanism through which both genomic localization and functional partner recruitment are achieved.

[0129] Introduction

[0130] Intrinsically disordered regions (IDRs) comprise 37-50% of the human proteome1and are especially enriched in nuclear proteins2. Rather than a singular structure, IDRs are defined by heterogeneous conformational ensembles3,4which has led to the prevailing view that IDR-mediated interactions are less specific than those mediated by folded domains5. However, associations driven by specific IDRs are known to play important roles in forming biomolecular condensates, which are regions of high local protein concentration formed via the process of phase separation or related phase transitions8. IDRs and their role in driving phase separation are implicated in various aspects of nuclear organization, but much remains unclear in the context of chromatin remodeling.

[0131] The mammalian SWI / SNF (mSWI / SNF or BAF) ATP-dependent chromatin remodeling complexes collectively represent one of the most frequently mutated cellular entities in human cancer, second only to TP539’10. Indeed, mutational frequencies for the 29 human genes that encode for mSWI / SNF complex subunits tally to over 20% across human cancers9. mSWI / SNF subunit mutations and translocations represent cancer-initiating events in a number of rare cancers11-13and are among the most frequently perturbed genes in neurodevelopmental disorders (NDDs)14-20.

[0132] The most frequently mutated genes within the mSWI / SNF family are the ARID1 genes, ARID1A an ARIDlB, which encode 250-kDa paralog subunits (ARID1A and ARID1B) that define and assemble into cBAF subcomplexes in a mutually exclusive manner9,21. ARID 1 Ais mutated in over 8% of human cancers arising from a range of cell lineages, while in neurodevelopmental disorders, ARID IB is the most recurrently mutated chromatin regulatory gene and one of the top five genes associated with autism19’22-24. These human genetic data indicate critical functional contributions of the ARID1 subunits as well as differences between the two paralogs. Recent studies21,25-27have revealed that the large ARID1 subunits (Swil in yeast SWI / SNF) connect the cBAF core with the ATPase module via a conserved core-binding region (CBR) containing 6 tandem Armadillo (Arm) repeats21,27(FIG. 1, panel A; FIG. 8, panel A). Expression of this CBR is sufficient for cBAF complex biochemical assembly, specifically, binding of the ATPase module on to the cBAF core21. Cancer-associated missense mutations in the C-terminal region destabilize ARID1A and / or prevent its assembly into cBAF complexes21,27; disease-associated mutations in both ARID1A and ARID1B are nonsense and frameshift in nature (FIG. 8, panel B). Intriguingly, the role of the remaining two-thirds of these proteins (65.69% of ARID1 A, 1501 amino acids; 68.74% of ARID1B, 1537 amino acids) remains uncharacterized. ARID1A / B N-termini contain two IDRs bridged by a structured ARID DNA-binding domain (FIG. 1, panels A-B; FIG. 8, panels C-D). Most cancer- associated mutations in ARID1A / B genes and NDD-associated mutations in ARID IB fall within the IDRs (-58% and -83%, respectively) (FIG. 1, panel C). Further, the IDRs of the ARID1 A / B N-termini make up -33% of the IDR content of the entire cBAF complex. Disorder scores (using MobiDB-Lite 3.028) for these ARID 1 A / B regions are similar to those of prionlike domains known to phase separate, including TDP-43, DDX4, FUS, and others29(FIG. 8, panel E).

[0133] IDRs within chromatin-bound proteins have p utative functional roles including influencing dynamics of chromatin-bound proteins30and transcriptional activation31, creating reaction crucibles32-34, and heterochromatic silencing35-37. Several pathogenic mutations in human cancer and Mendelian diseases map to condensate-forming proteins39. The functions imparted to nuclear proteins by IDRs remain incompletely understood in the context of ATP- dependent chromatin remodelers.

[0134] Here, we find that the ARID1A / B IDRs and DNA-binding ARID domain direct genomic targeting of the cBAF complex and subsequent generation of DNA accessibility, enhancer activation and gene expression, through an IDR-encoded specific biomolecular interaction netw ork.

[0135] Results

[0136] The ARID1A / B N-terminus is dispensable for cBAF assembly and in vitro nucleosome remodeling

[0137] To define the role of the IDR-rich ARID1A / B N-termini with respect to complex assembly and ATPase dependent nucleosome remodeling activities, we generated HA-tagged ARID1A full-length (wild-ty pe, WT) or ARID1A mutant variants that lack IDR1 (AIDR1), contain mutations in the ARID DNA-binding domain that compromise DNA binding as assayed by electrophoretic mobility shift assay (S1086E, S1087E. S1090E) (DBDmut) (FIG. 8, panels F-G), or lack the entire N-terminus, including IDR1, the ARID domain, and IDR2 (CBR only) (FIG. 1, panel D). We introduced these into AN3CA cells derived from a dedifferentiated endometrial carcinoma lacking both ARID 1 A and ARID IB subunits (and hence, lacking functional cBAF complexes) as well as ARIDlA / B-deficient HEK293T cells generated using CRISPR / Cas9-based editing21(FIG. 1, panel E; FIG. 8, panel H). The C- terminal CBR was sufficient to enable assembly of complexes in both cell types (FIG. 1, panel E; FIG. 8, panel I) Protein levels across mutants were similar to WT and unaffected by proteasome inhibition (FIG. 8, panel J). cBAF complexes purified from cells expressing WT or mutant ARID 1 A contained similar levels of BAF core and ATPase module subunits, by immunoblot and tandem-mass-tag (TMT) mass spectrometric analyses of HA immunoprecipitations (IPs) (FIG. 1, panels E-F). Expression of ARID1A in ARIDlA / B-null cells restored cBAF assembly, demonstrated by density sedimentation analysis21(FIG. 8, panel K). Intriguingly, restriction enzyme accessibility assays (REAA) revealed both WT and mutant complexes purified via HA-IP from AARID1 A / B HEK293T cells have equivalent nucleosome remodeling activities in vitro, and ATPase catalytic activities in solution (FIG. 1, panels G-H; FIG. 8, panels L-M), indicating the ARID 1 A C-terminus is sufficient for cBAF complex assembly, nucleosome remodeling, and catalytic activities, underscoring the need to investigate alternate functional contributions of the large N-terminal region.

[0138] The ARID1 disordered regions confer phase separation potential to cBAF complexes

[0139] Coupled with the high predictions for disorder and disease relevance of the N- terminal regions, we sought to examine their potential role in cBAF phase separation. We expressed individual C-terminally eGFP-tagged ARID 1 A WT, DNA-binding mutant, or truncation variants in AARID1 A / B HEK293T cells, isolated fully assembled cBAF complexes, and performed in vitro condensation (phase separation) assays (FIG. 2, panel A; FIG. 9, panels A-B). Purified protein complexes diluted to 2, 0.66, 0.2, and 0.074 pM in physiologicalsalt buffer with no additional crowding agent (150 mM NaCl, 25 mM HEPES pH 7.5), were imaged after 30 minutes on a spinning disc confocal microscope to query the presence of condensates. Complexes incorporating WT- or DBDmutARIDlA formed condensates in solution, while loss of one or both IDRs nearly completely attenuated condensate formation (FIG. 2, panel A, left; FIG. 9, panel C) We quantified the presence of condensates using a two-dimensional proxy for volume fraction: percent of the field of view covered by eGFP- positive droplets (Condensate Area. WT 9.05%; DBDmut8.63%; AIDR1 0.59%; CBR 0.04%) (FIG. 2, panel A, right). Addition of 100 nM DNA (linear, dsDNA of random sequence), nucleosomes (mixed mono-, di-, and tri-nucleosomes) or RNA showed that condensate formation was enhanced by DNA and nucleosomes, but not by RNA (Condensate Area WT Only 9.05%; WT + DNA 14.89%. WT + nucleosomes 15.50%, WT + RNA 8.24%), implicating cBAF complex DNA- and nucleosome-binding regions in promoting phase separation (FIG. 2, panel B; FIG. 9, panel D). In addition, WT but not DBDmutsamples formed strings of condensates in reactions containing DNA, similar to observations of a pioneer transcription factor40(FIG. 9, panel E). Of note, cBAF complexes contain several other DNA- binding domains within the core module; so attenuated condensation upon inactivation of the ARID domain alone indicates a prominent role for this domain. Interestingly, while the condensation of cBAF complexes carrying the DBDmutmutant was not enhanced by the addition of DNA, it was enhanced by the addition of nucleosomes, indicating that bilateral engagement of cBAF at the acidic patch regions27,41can enhance phase separation independent of ARID domain-mediated DNA binding (FIG. 2, panel B) Addition of DNA or nucleosomes moderately enhanced condensation of AIDR1 -containing cBAF complexes presumably via ARID domain-mediated DNA and nucleosome binding, though not significantly relative to complex-only control (FIG. 2, panel B; FIG. 9, panel D). Further, cBAF complexes nucleated by the ARID 1 A CBR alone failed to form condensates in any of the conditions tested, indicating that although additional IDRs are in other cBAF subunits27, they are not sufficient to induce condensation of complexes (FIG. 2, panel B).

[0140] Next, we evaluated protein dynamics and exchange using Fluorescence Recovery After Photobleaching (FRAP) in AN3CA cells expressing eGFP-tagged ARID! A / B. Addition of eGFP to ARID1A did not disrupt cBAF complex assembly in AN3CA cells as assayed by IP-immunoblot and density sedimentation analyses (FIG. 9, panels F-H). Indeed, WT ARIDlA-carrying cBAF complexes showed a clear punctate pattern by microscopy while disruption of IDRs nearly completely attenuated the presence of nuclear puncta (FIG. 2, panelsC-D) ARID 1 A DBDmut-expressing cells had consistently fewer condensates, each with increased area, reflecting enhanced coarsening enabled by loss of targeted interaction with genomic DNA (FIG. 2, panels C-D). Concentration-calibrated fluorescence imaging of eGFP- tagged ARID1A constructs demonstrated a threshold concentration of 1.13 ± 0.11 pM for WT and 1.08 ± 0.16 gM for DBDmut, above which the punctate nuclear pattern is observed (FIG. 2, panel D). Condensation and threshold concentration of the four ARID 1 A mutants were unaffected by proteasome inhibitor treatment (FIG. 9, panel I). Similar condensation patterns were obtained for the ARID IB paralog subunit (FIG. 9, panel J). Time-lapse imaging showed that individual ARID1A / B nuclear puncta are present over tens of minutes and exhibit fusion and coalescence, which characterizes purely viscous fluids or viscoelastic materials with terminally viscous properties (FIG. 15). FRAP experiments in the ARID 1 A WT condition demonstrated rapid recovery (half time of recovery T1 / 2 ~5.7 sec and T1 / 2 ~7 sec for ARID 1 A and ARID IB, respectively) with low immobile fraction (27%), consistent with viscoelastic materials that feature mobile and immobile species (FIG. 9, panel K). The ARID1A and ARID IB DNA-binding mutants (DBDmut) demonstrated similar dynamics with slight but statistically significant increases in half time of recovery (to T1 / 2 = 9.4 sec for ARID 1 A, T1 / 2 = ~15 seconds for ARID IB) but no change in immobile fraction, indicating that loss of DNA binding activity does not drastically alter protein dynamics (FIG. 9, panel K). By immunoblot, levels of exogenous ARID1A expression in AN3CA cells were comparable to endogenous ARID 1 A levels across a range of human and murine cell types (FIG. 2, panel E) and immunofluorescence detected punctate nuclear cBAF structures in endogenous contexts (FIG. 2, panel F ; FIG. 9, panel L). These data provide the first visual evidence of cBAF condensates under endogenous expression levels.

[0141] To further characterize the self-interaction capabilities of the ARID1 A / B N-terminal IDRs, we employed the Corelet System42, which makes use of a multivalent 'Core’ particle (24-mer Ferritin) to act as a scaffold for assembly of phase-separation-prone proteins in a lightdependent manner (FIG. 2, panel G, top). We generated variants of ARID1A containing IDR1, IDR2, or the full N-terminus (IDRs and ARID domain, FL), each lacking the C-terminal BAF-binding CBR region to allow for us to study the low-complexity N-terminus in isolation (FIG. 2, panel G, bottom). Notably, the ARID 1 A N-terminus formed light-dependent condensates over a wide range of concentrations and valences, while IDR alone or DBDmutexhibited significant attenuation in phase separation potential (FIG. 2, panels H-I; FIG. 9, panel M). Again, the DNA-binding domain mutant formed fewer droplets of larger size (FIG.2, panel H; FIG. 9, panel N). Similar results were obtained for ARID1B N-terminus, except that IDR1 more closely mirrored the full-length N-term variant, perhaps indicating its stronger phase separation propensity (FIG. 9, panels O-P). Repeated on-off light cycles revealed that the specific nuclear localization of ARID1 A IDR puncta, observed as high correlation between nuclear positioning in subsequent cycles, was dependent on the ARID domain (FIG. 9, panels Q-R). Together, these data highlight the functionality of the IDRs and ARID DNA-binding domain of ARID1A / B subunits in conferring phase separation and sub-nuclear localization properties to cBAF remodeling complexes.

[0142] ARID1A IDRs and ARID domain are required for cBAF targeting, chromatin accessibility and gene expression in cells

[0143] To determine the functional contributions of the IDRs and ARID domain of ARID1 A / B, we introduced ARID 1 A WT, AIDR1, DBDmut, CBR-only mutant variants (FIG. 1, panel C) or empty vector control into AN3CA cells and performed CUT&Tag43, ATAC- Seq44-45, and RNA-Seq to evaluate chromatin localization of cBAF, DNA accessibility, and gene expression, respectively. We first examined the chromatin occupancy of cBAF complexes, the enhancer mark H3K27ac, and DNA accessibility. Global clustering analyses performed on over 40,964 merged SMARCC1 / SMARCA4 sites revealed a set over which only WT ARID 1 A restored complex occupancy and accessibility7, whereas IDR deletion (AIDR1, CBR) and ARID domain (DBDmut) mutants were unable to restore these features (Cluster 2: 5042 sites; 12.3%) (FIG. 3, panel A; FIG. 10, panel A). We also identified a cluster that exhibits a similar trend to a lesser extent, with CBR-only mutant being the most deleterious (Cluster 3: 4705 sites; 11.5%) (FIG. 3, panel A; FIG. 10, panel A). Clusters 2 and 3 sites were largely Transcriptional Start Site (TSS)-distal, consistent with an important role for cBAF complexes in enhancer accessibility41,46(FIG. 3, panel B). Cluster 1 (unaffected by ARID 1 A expression) contains promoter-proximal sites (FIG. 3, panels A-B). Principal component analysis (PC A) of cBAF-occupied enhancer sites demonstrated a distinct clustering pattern, with the DBDmutmost similar to WT, and AIDR1 and CBR mutants closest to empty vector control (FIG. 3, panel C). These findings are exemplified at intragenic enhancers at \hs MAP2 and NCAPH loci and an intergenic enhancer within chromosome 2 (FIG. 3, panel D). Consistent with in vitro remodeling data demonstrating that the N-terminus is not required for ATPase activity7(FIG. 1, panels G-H), we did not identify sites with intact mutant complex targeting but loss of accessibility7, indicating genomic targeting of the complex, not core enzymatic remodeling activity, is compromised in these mutants (FIG. 3, panel A). Thenumber of sites affected genome- wide (n= 9747 total for Clusters 2 and 3) mirror those affected by complexes containing defects in the ATPase activity itself (i. e. , K.785R of SMARCA4), or complexes lacking core components such as SMARCB1 or SMARCE112>13’41’46, indicating that disruption of the IDRs of the ARID1 proteins direct similar consequences for cBAF complex targeting as disruption of these structurally integral subunits.

[0144] At a global level, WT ARID1A expression led to a significant increase in accessibility by ATAC-Seq compared to empty vector control (39170 de novo sites) (FIG. 3, panel E). Accessibility gains were reduced upon expression of each mutant variant relative to WT (DBD""" = 20797, AIDR1= 9931, CBR= 8539 sites), with CBR mutant resulting in the lowest accessibility, followed by the AIDR1 and DBD""rfmutants (WT > DBD"!i" > AIDR1 > CBR) (FIG. 3, panel E; FIG. 10, panel B). PCA performed across the ATAC-Seq and RNA- Seq conditions similarly revealed the DBDmutclusters closer to WT than AIDR1 and CBR mutants (FIG. 10, panels C-D) These data collectively indicate that loss of the ARID1 A N- terminal IDR and / or DNA-binding regions of BAF complexes result in substantial changes in targeting and genomic accessibility in cells. Accessible sites in ARIDlA-mutant conditions represented a subset of those generated by ARID1A WT (FIG. 10, panel E), with significant overlap among one another, exemplifying the convergent deficits in differentially perturbed cBAF complexes (FIG. 10, panels E-F).

[0145] Sites most affected by disruption of the ARID N-terminus were enriched in transcription factor (TF) motifs corresponding to the AP-1, FOS / Jun, NF1, and TEAD factors, several of which have been shown to localize to enhancers via interaction with mSWI / SNF complexes (FIG. 3, panel F)47,48. 72% of accessible sites gained in cells expressing WT ARID 1 A showed a concordant increase in occupancy of H3K27ac, and were enriched for similar transcription factor motifs as Cluster 2 and 3 sites (FIG. 10, panels G-H; FIG. 3, panel A; FIG. 3, panel E). Intriguingly, sites at which cBAF complex occupancy and DNA accessibility w ere reduced upon rescue with WT ARID 1 A but not the mutant variants or empty vector control were significantly enriched for CTCF and CTCFL (BORIS) motifs (Cluster 4: 5915 sites), consistent with recent observations that lack of cBAF assembly and / or function results in increased non-canonical BAF (ncBAF) complex abundance and function at its key target sites (CTCF) (FIG. 3, panel F; FIG. 3, panel A, Cluster 4, FIG. 10, panel A)4930.

[0146] Finally, we found that relative to WT ARID1 A, expression of ARID1A N-terminal mutants resulted in attenuation of gene expression (FIG. 3, panel G; FIG. 10, panel I). Globally, we identified a greater number of upregulated and downregulated transcripts inmutant conditions, with the CBR mutant resulting in the fewest upregulated transcripts (FIG.10, panels J-K). In this AN3CA endometrial cellular context specifically, N-terminal mutants failed to rescue expression of genes involved in endometnal cell differentiation (FIG. 10, panel L), indicating that condensation of cBAF is essential to its genome-localized nucleosome remodeling function. These data demonstrate that the IDR-rich N-terminus, coupled with the ARID DNA-binding domain, are together required for the stable occupancy of cBAF complexes at distal enhancers over which they establish and maintain accessibility.

[0147] Heterotypic cBAF interactions with transcription factors require IDR sequences and the ARID DNA-binding domain of ARID1A

[0148] We sought to define the mechanistic basis underlying the necessity of the ARID1A / B N-termini for cBAF function. We reasoned that proteins localizing into ARID1 A / B-containing nuclear condensates can be identified by their proximity’, and so performed proximity labeling followed by mass-spectrometry by fusing an engineered biotin ligase TurboID (TblD)51,52to the C-terminus of ARID1A WT and mutant variants to map changes in the proximal protein repertoire of cBAF complexes (FIG. 11, panel A). TbID fusion did not disrupt nucleation and assembly of cBAF (FIG. 4, panel A). Upon confirmation of self labeling of the bait (ARID 1 A), non-self labeling with biotin (50pM for 10 minutes), and visualization with streptavidin (FIG. 11, panel B), we performed TMT mass spectrometry' to identify proximal proteins for each cBAF complex variant (FIG. 11, panel C). Notably, truncation of the full N- terminal region or IDR1 alone, but not inactivation of the ARID DNA-binding domain (DBDmut), resulted in a significantly depleted repertoire of proximal proteins (FIG. 4, panel B). This set of proteins was enriched in factors associated with chromatin organization, histone modification, and transcription (FIG. 4, panel C; FIG. 11, panel D). Losses in associated proteins were IDR-dependent with no significant changes in the DBDmutmutant (relative to WT) (FIG. 4, panel C) We found that mSWI / SNF components themselves (cBAF as well as PBAF and ncBAF) were markedly reduced near cBAF complexes lacking the IDRs of ARID1 A (FIG. 11, panel E, left), with no change in total nuclear protein level (FIG. 11, panel E, right), indicative of reduced proximity due to a loss of condensate formation. Furthermore, we measured a marked reduction in the abundance of Mediator complex components, RNA Polymerase II, the p300 acetyl transferase, and selected TFs in proximity of IDR-mutant complexes relative to WT, again, absent changes in corresponding nuclear protein levels (FIG.11, panels F-G). To validate these data, we performed immunofluorescence colocalization studies of p300 with ARID1 A-WT or -mutant cBAF complexes. Consistently, we found alterednuclear distribution of p300 and a loss of co-condensation with cBAF (FIG. 4, panel D; FIG. 11, panels H-I). These findings demonstrate the critical role of the ARIDlA N-terminal IDR1 in facilitating localized condensation of cBAF complexes and their association with the transcriptional machinery, TFs, and other factors required for functional chromatin remodeling.

[0149] Though the proximal protein repertoire of ARID 1 A DBDmutcarry ing cBAF complexes was similar to that of WT cBAF, these complexes were defective in genomic localization (FIG. 3, panel A). To identify the reason behind this observation, we used IP- Mass Spectrometry' (IP-MS) to identify high-stringency protein interactors of cBAF complexes and determine whether these interactors are lost upon ARID mutation. We identified 1076 interacting proteins that were dependent on ARID1A WT for association with cBAF. >90% of which overlapped with those identified in the TurboID-based proximity labeling experiments (FIG. 4, panels E-F; FIG. 11, panel J) cBAF -interacting factors were enriched for TFs such as FOS / Jun, TEAD1, NFIA, NFIB, RELA, GATA2, ATF3, and CUX1, consistent with the roles of TF-cBAF interactions in genomic navigation53’54. The identified TFs correspond to key cognate DNA motifs that were enriched under ARID 1 A IDR-dependent sites genome-wide (FIG. 4, panel G; FIG. 3, panel F; FIG. 11, panel K). Of the DNA-interacting IP-MS hits, 75% are transcription factors, of which, six that interact with cBAF by IP-MS also have motifs enriched in Cluster 2 / 3 sites, including cJUN. NFIA and TEAD1 (FIG. 4, panel G). Importantly, by IP-MS, ARID 1 A DBDmut-carrying complexes were equally deficient for TF tethering as IDR-mutant complexes (FIG. 4, panel E; FIG. 3, panel A), indicating that the ARID domain stabilizes a broad set of TF-cBAF interactions. Similarly, transcription initiation machinery components detected by proximity labeling were not enriched by IP-MS, indicating these factors localize near to but do not directly bind cBAF complexes (FIG. 4, panel F). Finally, reciprocal co-immunoprecipitation followed by immunoblots for selected TFs demonstrated specific binding to WT but not mutant ARIDlA-containing cBAF complexes, indicating these interactions are dependent on the N-terminus (FIG. 4, panel H; FIG. 11, panel L). These parallel proximity labeling and IP-MS experiments define the related but distinct sets of proximal and stringent interactions mediated by the ARID ! A N-terminus, and the role of the ARID domain in stabilizing functional associations with TFs mediated by disordered regions within ARID 1 A.

[0150] Genomic targeting and protein interactions of cBAF complexes requires the ARIDlA-specific IDR

[0151] Given the critical role of the ARID1A / B IDRs in driving condensation, protein interactions, and genomic localization of cBAF in cells, we sought to determine whether these functions can be performed by other phase separation-prone IDRs. To evaluate this, we generated constructs replacing IDR1 of ARID 1 A with alternate well-known self-interacting IDRs from FUS and DDX455(FIG. 5, panel A). Given the retention of the ARID1A CBR, these fusion constructs were able to nucleate cBAF assembly in AN3CA cells (FIG. 12, panel A). Live-cell microscopy revealed the presence of condensates in FUSIDR- and DDX4IDR- ARID1A mutant expressing cells, comparable in count, area, saturation concentration and FRAP dynamics to those detected in the ARID 1 A WT condition (FIG. 5, panels B-D; FIG. 12, panels B-C), indicating that these alternate IDRs are sufficient for condensation of cBAF in living cells.

[0152] To define whether FUSIDR / DDX4IDR-ARID1A can rescue cBAF chromatintargeting, we profiled complex occupancy, DNA accessibility and gene expression in AN3CA cells. We focused specifically on de novo cBAF-occupied and accessible sites that were specific to the WT ARID 1 A condition (FIG. 3, panel A, Clusters 2, 3). Importantly, cBAF complexes containing FUSIDR / DDX4IDR-ARID1A were unable to recapitulate WT targeting, indicating sequence-specific functions of ARID 1 A IDR1 (FIG. 5, panel E; FIG. 12, panels D-E). PCA of ATAC-Seq sites revealed that FUSIDRand DDX4IDRARID1 A mutants clustered more closely with AIDR1 than ARID 1 A WT, indicating that although they rescue cBAF condensation, they fail to recapitulate genomic targeting, and implicating ARID 1 A IDR1 in cBAF chromatin occupancy (FIG. 12, panel E). Importantly, FUSIDR / DDX4IDR-ARID1 A variants failed to activate gene expression relative to WT ARID 1 A (FIG. 5, panel F; FIG. 12, panel F).

[0153] Next, we investigated the underlying basis for the specificity of ARID 1 A IDR1 in mediating cBAF activity using IP-MS experiments. Following confirmation that replacement of ARID1A IDR1 with FUS- or DDX4-derived IDRs does not alter cBAF assembly (FIG. 12, panel G), we found that FUSIDR- and DDX4IDR-ARID1A mutant carry ing cBAF complexes each failed to capture TFs associating with WT cBAF complexes (FIG. 5, panel G). Proximity labeling experiments using a FUSIDR-ARIDlA-TbID fusion construct confirmed that FUSIDR- ARID1 A failed to restore proximity of cBAF to TFs and the transcriptional machinery relative to ARID1A WT (FIG. 12, panels H-I). Instead, the repertoire of proteins nearby FUSIDR- ARID1A containing cBAF complexes were more similar to the AIDR1 ARID1A mutantcarrying BAF (FIG. 12, panel J; FIG. 11, panel F). Further, immunofluorescence confirmedthat p300 does not colocalize with cBAF complexes carrying FUS / DDX4IDR-ARID1 A fusions, without affecting nuclear protein levels of p300 (FIG. 5, panel H; FIG. 12, panels K-L). Finally, FUS / DDX4IDR-ARID1A fusions had increased occupancy over AIDR1 or CBR-only ARID 1 A bound sites, absent any corresponding changes in chromatin accessibility (FIG. 5, panel I), exemplified over the BRD2 and CD320 loci (FIG. 5, panel J), indicating similar off- target binding in these mutant conditions. These data indicate that generic condensation of BAF is insufficient for genomic targeting in cells, imparting a marked specificity to the ARID! A N- terminal IDR and indicating that condensate-driving IDRs need not be functionally interoperable with one another.

[0154] Analysis of ARID1A IDR1 sequence features enables uncoupling of condensation and heterotypic protein-protein interactions

[0155] To decipher the underlying basis of the specificity of ARID 1 A IDR1, we performed IDR-specific comparative analyses of the “sequence grammar” of ARID1 A / B IDRs, including distinctive compositional biases, non-random binary7sequence patterns that influence conformational properties of IDRs, and the presence, if any, of short linear motifs56. To uncover these features, we collated disordered sequences across the entire mSWI / SNF family of protein subunits (within cBAF, PBAF, ncBAF complexes) and analyzed their amino acid compositional and sequence patterning features using the NARDINI+ algorithm6-5657, which combines the work of Zarin et al., and Cohan et al, and enunciates the findings in terms of sequence feature vectors. These vectors were then hierarchically clustered using Euclidean distance and Ward’s clustering6,56. We found that IDR1 of ARID1 A and ARID1 B cBAF- defining subunits represents a distinct evolutionary7cluster away from the other mSWI / SNF IDRs, including IDR2 of ARID 1 A / B, indicating that they harbor distinctive non-random sequence features (FIG. 6, panel A).

[0156] The ARID 1 A / B IDRls are uniquely enriched in Alanine-Glutamine-Glycine stretches or “blocks ” (FIG. 6, panel B; FIG. 13, panel A) This highly non-random blocky patterning in the ARID1A / B IDRls was found to be conserved across eukaryotes (at the phyla level) despite divergence in the amino acid sequence across homologs (FIG. 13, panel B). Additionally, we identified a pronounced compositional bias, with more than 40 aromatic residues (Tyrosine, Tryptophan, and Phenylalanine) distributed uniformly across the 1016- amino acid IDR; aromatic residues contribute to pi-pi and cation-pi interactions that have been shown to drive homotypic (self) interactions and phase separation in IDRs from other condensation-prone proteins including FUS58,59. Given these features, we next generatedARID 1 A IDR1 mutant variants that disrupt blockiness of AQG patches by scrambling the amino acid content within them (AQGscram), or disrupt aromatic character by mutating 42 Tyrosines to Serines (42YS) (FIG. 6, panels B-C). Both designs maintain the overall IDR length. Disruption of AQG blocks in the AQGscram mutant and preserv ation in the 42YS mutant were confirmed using NARDINI6(FIG. 6, panel D). Following confirmation that these mutant variants maintained expression level and complex integration when expressed in AN3CA cells, we performed live condensate imaging. ARID1A 42YS mutant-containing cBAF complexes failed to form condensates in cells, while the AQGscram mutant-containing complexes formed condensates comparable to those carrying WT ARID 1 A, with slightly increased area and attenuated FRAP recovery times (FIG. 6, panels E-G; FIG. 13, panel C). These data indicate that the 42 Tyrosine residues found in ARID 1 A IDR1 are the main “stickers”60that drive phase separation of the 1.04 MDa cBAF complex, and that the evolutionarily conserved, non-random AQG blocks found in this region are not essential for cBAF condensate formation.

[0157] Importantly, both the 42YS and AQGscram ARID 1 A mutant variants show equivalent failure to rescue cBAF localization and DNA accessibility at de novo WT cBAF- occupied sites (n=9159 sites), in a manner similar to the DBDmut, AIDR1, and CBR mutants (FIG. 6, panel H; FIG. 13, panels D-H; FIG. 3, panel A) >60% of the sites with reduced occupancy of these two convergent mutants overlapped with reduced occupancy sites in AIDR1, CBR, or DBDmutcontexts (FIG. 6, panel I; FIG. 13, panels F-G; FIG. 3, panel A, Clusters 2 and 3). At the gene expression level, expression of 42YS and AQGscram ARID1 A mutants resulted in overall downregulation of genes relative to WT ARID 1 A (FIG. 6, panel J; FIG. 13, panel H)

[0158] To further understand the mechanism of action of these two IDR disruptions, we mapped the proximal protein repertoire of complexes containing the 42YS or AQGscram mutant using TbID-based proximity labeling. Intriguingly, we find that complexes carrying the 42YS mutant have a comparable proximal protein repertoire to WT, while complexes containing the AQGscram are severely deficient in their interaction network (FIG. 6, panel K). Indeed, we found a significant reduction in p300 colocalization in the setting of the ARID1A AQG scramble variant (FIG. 6, panel L; FIG. 13, panel J). Both mutants are deficient in direct TF tethering, as assayed by coimmunoprecipitation immunoblot analysis of the TFs NFIA and TEAD1 (FIG. 6, panel M). This indicates a sequence-encoded separation of functions for the ARID 1 A N-terminal IDR1 region, namely condensate formation throughTyrosine residues, and partner protein interactions through specificity imparted by AQG blocks. Both roles together are essential for direct TF tethering and proper genomic localization of cBAF in cells. Of note, FUS and DDX4 IDRs also utilize aromatic residues for pi-pi (FUS) and cation-pi (DDX4) interactions as drivers of condensation though they lack the AQG blocks found in the ARID 1 A IDR. Therefore, while these orthogonal systems were able to rescue condensation of cBAF in cells, they were unable to recapitulate the network of functionally relevant heterotypic interactions (FIG. 13, panel H; FIG. 6, panels A-C).

[0159] NDD-associated mutations in ARID1B IDR1 sequence blocks disrupt cBAF condensate formation and chromatin localization

[0160] Finally, we sought to utilize our understanding of IDR sequence grammar to rationalize human disease-associated missense mutations that localize to the IDRs of ARID1A / B. Referencing a collated list of neurodevelopmental disorder (NDD)-associated mutations from the DECIPHER database, we find that ARID IB is enriched for this category of mutations relative to its paralog, ARID1A (FIG. 7, panel A; FIG. 14, panel A). Mapping the occurrence of NDD-associated mutations within the 26 AQG-rich blocky sequences (FIG. 6) reveals that Block 9, a large A / G-rich block, and Block 13, a shorter polyA-rich sequence, are disproportionately perturbed (FIG. 7, panel B; FIG. 7, panel C). We designed and cloned ARID1B in-frame truncation variants lacking these regions (Block 9 or Block 13 deletion) and selected causal NDD-associated mutations falling within these regions (S320_327del within Block 9, and A457 G461del within Block 13) (FIG. 7, panel C) to test for condensation, genomic localization, and accessibility generation in cells. The ARID1B mutant variants did not affect cBAF complex assembly (FIG. 7, panel D). We found that the block deletions and patient-derived ARID IB mutant-carrying complexes are still capable of condensation (FIG. 7, panel E), and their mobility by FRAP is not significantly different from wild-type (FIG. 14, panel B), in line with the result that scrambling the blocky AQG sequences did not abolish condensation propensity7or significantly affect cBAF complex diffusion (FIG. 7, panel E; FIG. 6, panel F). Interestingly, the saturation concentration of Block 13 del mutant is higher than wild type, indicating that deleting this block disrupts self-interaction, though this phenotype is not significant in the shorter patient derived A457_G461del mutant (FIG. 7, panels E-F). While we do not observe major changes in condensate count (except for the S320_G327del mutant), we notice an overall increase in condensate area (FIG. 7, panel F), similar to that observed for the ARID DNA-binding domain mutant variant (FIG. 2, panel D), indicating that these mutations can disrupt TF tethering or chromatin-bound stability. Tocontextualize these results, we measured the effect of the SMARCA2 / 4 ATPase inhibitor, Compound 14, on the condensation propensity of the complex61-62. ATPase inhibition has been demonstrated to result in destabilized cBAF complexes at distal enhancers at which they interact with key TFs, resulting in accumulation of complexes over open promoters46,63. Consistently, upon ATPase inhibition, we find formation of cBAF condensates, albeit fewer puncta and with greater area per nucleus (FIG. 14, panel C).

[0161] We mapped chromatin occupancy of cBAF carrying ARID1B WT or mutants (CUT&RUN) and measured DNA accessibility (ATAC-Seq) as well as gene expression (RNA- Seq) to test the functional impact of deletion and disease-associated IDR perturbations on cBAF function. We found that the Block 9 and Block 13 deletion mutants exhibit substantial loss of localization, while patient-derived mutants S320 G327del and A457 G461del that map to Blocks 9 and 13, respectively, result in partial but significant localization defects, consistent with their compatibility with life in individuals with NDDs (FIG. 7, panels G-H; FIG. 14, panels D-F). These findings are exemplified at the NCAPH and IL1B loci on chromosome 2 (FIG. 7, panel I). Importantly, HOMER TF motif enrichment analyses identified motifs corresponding to the AP-1. FOS / Jun. NF1. and TEAD factors to be enriched over sites at which cBAF complexes were defective in targeting and accessibility generation in the mutant conditions relative to WT ARID1B (FIG. 7, panel J). In line with this, by IP-mass spectrometry, we identified a significant reduction in association of the NFI TF family with cBAF complexes carrying the NDD-associated S320 327del ARID IB variant (FIG. 7, panel K). Finally, NDD-associated mutations and block deletions resulted in a significant attenuation in gene expression activation relative to ARID1B WT over key differentiation-associated genes (FIG. 7, panels L-M). These results underscore the impact of in-frame disruptions within the ARID1A / B IDRs on cBAF remodeler function and present a foundation for the mechanistic assignment and characterization of such mutations in human disease.

[0162] Discussion

[0163] Most studies on chromatin regulatory complexes, including mSWI / SNF complexes, have focused on highly structured domains, characterizing how their physical features dictate chromatin binding and activity. Our findings provide understanding of a unique disordered domain on a remodeler, the mSWI / SNF family cBAF complex, for which localized condensation and heterotypic interactions are both essential, and independently directed by a distinct set of non-random sequence features encoded within ARID1A / B N-terminal IDRs(FIG. 7, panel N). These features are critical in governing cBAF-mediated genome-wide targeting, accessibility generation, and gene regulatory activities.

[0164] Our results reveal that IDR1 of ARID1 A / B carries a set of unique sequence features relative to the IDR sequences within the mSWI / SNF family subunits (FIG. 6, panel A). We found that deleting IDR1 alone almost entirely prevents condensate formation of full, >1.5 MDa cBAF complexes in cells. Additionally, within IDR1, short GA / A block deletions or NDD-associated mutations within these blocks maintained condensation but attenuated TF binding and genomic targeting of WT cBAF complexes. Furthermore, our data indicate that while other cBAF complex subunits contain IDRs, they do not confer self-interaction properties sufficient for condensation. Beyond cBAF, additional subunits within the mSWI / SNF family contain IDRs, indicating by extension that IDRs of related chromatin remodelers can serve as critical components of spatial genome organization (FIG. 13, panel A). Further, the protein subunits that comprise human cBAF complexes contain increased intrinsic disorder relative to those of yeast SWI / SNF complexes27. This indicates a model in which additional IDRs evolved to confer condensation properties and highly specific protein-protein interaction networks, to facilitate gene regulation in the mammalian nucleus.

[0165] Incubation of cBAF complexes with DNA in vitro potentiates condensation. This can be reversed by inactivation of ARID DNA-binding domain, despite the fact that the core module of cBAF complexes contain several other sequence non-specific DNA-binding domains, highlighting its unique function (FIG. 2, panels A-B). Moreover, the ARID domain is required for cBAF to appropriately interact with TFs in the nucleus (FIG. 4, panel E), indicating a distinct role for the ARID1A / B ARID domain. These results begin to provide insights regarding the order of events of nucleation and assembly of cBAF complexes on chromatin, their interactions with DNA, and their association with binding partners.

[0166] One notable finding of our study is that alternate low complexity IDRs derived from unrelated proteins cannot rescue cBAF genomic targeting and protein interactions in cells (FIG. 5, panel E; FIG. 5, panel G) despite identical condensation properties (FIG. 5, panels C-D), underscoring the key roles for condensation-specific and interaction-specific sequence grammars in IDRs64'75. The integrative approach used here enabled our conclusions; indeed, quantification of condensation alone would have indicated the IDR swap mutants are functionally comparable to WT ARID1A, yet when combined with genomic and biochemical evaluation, we found that condensation alone does not confer cBAF function.

[0167] Importantly, condensate formation and heterotypic biomolecular interaction networks can be distinct, each playing critical but separable roles in biological function. We demonstrate here that condensate formation and protein-protein interactions of the ARID 1 A N-terminal IDR are independent of each other, but they are both required for chromatin targeting of cBAF in cells. Our results indicate that cells can be able to regulate and evolve these features independently to create localized, compositionally defined, and functionalized, high concentration compartments in a modular way.

[0168] Our analysis of the non-random sequence features of cBAF IDRs provides a framework upon which to mechanistically assign the number of disease-associated missense and indel mutations that fall within the ARID1A / B IDRs (FIG. 1, panel C). Our data indicate thatNDD-associated changes ofjust a few amino acids within the ARID1B IDR partially alters condensation properties, TF interactions, and chromatin-level targeting in cells (FIG. 7), though these changes are more subtle than full block deletions or complete IDR deletion (FIG. 14, panel G), in agreement with the knowledge that NDD-associated mutations are live birth compatible. Intellectual disability (Coffin-Siris syndrome)-associated mutations in the C- terminal domain of the SMARCB1 subunit result in similarly subtle live-cell phenotypes41.

[0169] STAR Methods

[0170] Key Resource Table

[0171] Experimental Models and Study Participant Details

[0172] Cell lines and culture conditions

[0173] All human and mouse cell lines were grown at 37 °C with 5% CO2. HEK293T and U2OS cell lines are female human cells. HEK293T AARID1A / B and U2OS cells were grown in DMEM media (Gibco) supplemented with 10% FBS (Gibco), IX GlutaMAX (Gibco), 100 U / mL Penicillin-Streptomycin (Gibco), 1 mM Sodium Pyruvate (Gibco). IX MEM NEAA (Gibco), and 10 mM HEPES (Gibco). AN3CA endometrial cancer cells are human female cells and were grown in EMEM media (Gibco) supplemented with 15% tetracycline-free FBS (Omega), IX GlutaMAX (Gibco), 100 U / mL Penicillin-Streptomycin (Gibco), 1 mM Sodium Pyruvate (Gibco), IX MEM NEAA (Gibco), and 10 mM HEPES (Gibco). For immunofluorescence data. U2OS and MDA-MB-231 (human female breast-cancer derived) cells were grown in DMEM media with glucose, glutamine and pyruvate (ThermoFisher) supplemented with 10% FBS (Avantor) and 100 U / mL Penicillin-Streptomycin (Gibco). KLE (ATCC CRL-1622, human female uterine cell line), C2C12 (ATCC CRL-1772, female mouse myoblast cells) and CRL-7250 male human foreskin fibroblast cells were cultured in the same but with 20% FBS. MCF10A and MCF10-CA human female breast cancer cells were culturedin DMEM:F12 (Gibco 21041025) supplemented with 5% Horse serum (Sigma), 20 ng / mL EGF. 1 pg / mL Hydrocortisone, and 10 pg / mL Insulin. U2OS cell lines were authenticated by STR profiling.

[0174] Primary rat neuron dissection and culture

[0175] The inner 60 wells of 96 well glass bottom plates were treated with 0.01 mg / mL poly-D-lysine at 37 °C overnight and washed x4 in HBSS. The outer 36 wells of the 96 well plate were filled with ultrapure water. 50 pL of neuron media (Gibco Neurobasal Plus with 2% Gibco B27 Plus, 1% penstrep, and 250 ng / mL Amphotericin B) with 2% Gibco CultureOne supplement (antimitotic) was added to each well, and the plates were stored at 37°C overnight, 5% CO2. Embryos were collected from euthanized Sprague-Dawley rats (Hilltop Lab Animals Inc.) at embryonic day 17 via caesarian section. The embryos in placentas were transferred to HBSS in 10 cm glass plates. The placenta was cut from each embryo, the heads were removed and transferred to a new glass plate with HBSS. Using a dissection microscope, the skull was removed by making a medial cut from caudal to rostral following the central sulcus using small scissors held parallel to the brain, cutting just the skull layer and not into cortex. Closed scissors were used to get under the brain from the caudal side and gently flip brain out, cutting away any remaining attachment. Brains were transferred to a new glass plate with HBSS. Meninges were carefully and thoroughly removed starting with the ventral side, flipping to dorsal side, removing caudal to rostral along the central sulcus, gently unraveling the cortex from the central sulcus. The cortex was cut away from the striatum and other structures and transferred to a 10 mL conical with HBSS.

[0176] Worthington papain dissociation kit was used to dissociate cortices into individual cells in a biosafety cabinet using sterile technique. Reagents were prepared as described by the kit. HBSS was carefully removed from the cortices and 5 mL papain solution was added (100 units papain, 1000 units DNase I, 1 mM L-cysteine, 0.5 mM EDTA in HBSS). The conical was inverted thrice and then incubated at 37 °C for 20 minutes, with no agitation or inversion after the incubation. The papain solution was removed, and 3 mL of inhibitor solution (3 mg ovomucoid inhibitor. 3 mg albumin, and 500 units DNase I in HBSS) was added to the cortices, inverted thrice, and sat upright for 5 min. Supernatant was removed and replaced with 3 mL additional inhibitor solution, inverted thrice, and sat upright for 5 min. Supernatant was removed and 1.5 mL neuron media was added. A flame-treated Pasteur pipette was used to slowly triturate up and down ten times, avoiding bubbles. Cells were allowed to settle in the upright tube for 2 min. The top 750 pL of dissociated cells were removed and added to a new10 mL conical. 750 pL neuron media was added to the original tube, triturated ten times, settled in the upright tube for 2 min, and the top 750 pL of dissociated cells were transferred to the new 10 mL conical. This process was repeated one more time, for a total of three trituration steps, adding the media with cells to the new tube after the final trituration. Cells were centrifuged for 5 min at 300 g, supernatant removed, resuspended in 1 mL neuron media, and counted using a hemocytometer. Cells were diluted in additional neuron media to achieve 25.600 cells in 50 pL per well (80.000 cells per cm2growing area). 50 pL of diluted cells were added to each well of the previously prepared plates to bring the final volume to 100 pL with 1% CultureOne supplement. CultureOne supplement was not used again after this treatment on day in vitro (DIV) 0. Cells were grown at 37 °C with 5% CO2. On DIV3 100 pL more neuron media was added. Every 3-4 days after that, 95 pL media was removed from each well and replaced with 100 pL fresh media and 5 pL ultrapure water to counter evaporation. Neurons were fixed and used for immunofluorescence on DIV11.

[0177] Quantification and statistical analysis

[0178] Statistical analyses on quantified imaging data was performed with Prism. Statistical details, exact values of n and what n represents (individual cells or biological replicates) for each experiment can be found in the figure legends. In general, p values of significance less than 0.05 are denoted with one asterisk less than 0.01 with two asterisks less than 0.001 with three ‘***‘ and less than 0.0001 with four asterisks ‘****’_ No outlier data was omitted, no samples were excluded from our analyses. To identify differentially expressed genes (DEGs) or differentially interacting proteins, t-tests were performed on RNA-sequencing and mass spec data respectively. Error bar representation is indicated in the figure legends.

[0179] Method Details

[0180] Plasmids, cloning and expression

[0181] All ARID1 A / B constructs used in this study were HA-tagged at the N-terminus and cloned into a piggybac vector downstream of a Doxycycline-inducible promoter. The vector also contains a separate Tet-On 3G gene and Blasticidin or Puromycin resistance gene cassette separated by a P2A sequence, under the human EFla promoter. Constructs were sequence verified using Sanger sequencing. Piggybac plasmids were co-transfected with a mammalian expression plasmid carrying a transposase gene cassette in AN3CA using Lipofectamine 3000 (Thermo Fisher) and selected with 10 ug / ml Blasticidin or 2 ug / ml Puromycin 24 h post transfection for 3-5 days. Expression of the transgene was induced by addition of 200 ng / ml Doxycyline for 48 hours. Plasmids used in this study are listed in the STAR Methods section.

[0182] Coimmunoprecipitation

[0183] cBAF complex coimmunoprecipitation

[0184] BAF complex immunoprecipitation was performed as described previously21. Cells were washed with cold PBS and resuspended in EB0 hypotonic buffer containing 50 mM Tris- HC1 pH7.5, 0.1% NP-40,1 mM EDTA, 1 mM MgCh supplemented with protease inhibitors. Lysates were pelleted at 5,000 rpm for 5 min at 4 °C. Supernatants were discarded, and nuclei were resuspended in EB300 high salt buffer containing 50 mM Tris pH 7.5. 300 mM NaCl. 1% NP-40, 1 mM EDTA, 1 mM MgCh supplemented with protease inhibitors. Lysates were incubated on ice for 10 min with occasional vortexing and then spun at 21000 g for 11 min at 4 °C. 0.5-1 mg of nuclear lysate was used for immunoprecipitation with rabbit anti-HA antibody (1 :200 v / v) (Cell Signaling Technology) overnight at 4 °C to bind to HA-tagged ARID1A / B (bait). Protein-G Dynabeads (ThermoFisher) were then added for 2 hours and washed five times with EB300. Protein was eluted from beads with 4X LDS buffer by boiling for 7 min and loaded onto SDS-PAGE gels for Western blotting. Antibodies are listed in the STAR Methods section.

[0185] Transcription factor-cBAF complex coimmunoprecipitation

[0186] Reciprocal immunoprecipitations to validate cBAF Immunoprecipitation-Mass Spectrometry’ results were performed as follows: Cells were washed with cold PBS and resuspended in EB0 hypotonic buffer containing 50 mM Tris-HCl pH 7.5, 0.1% NP-40, 1 mM EDTA. 1 mM MgCh supplemented with protease inhibitors. Lysates were pelleted at 5.000 rpm for 5 min at 4 °C. Supernatants were discarded, and nuclei were resuspended in EB150 salt buffer containing 50 mM Tris-HCl pH 7.5, 150 mM NaCl, 1% NP-40, 1 mM EDTA, 1 mM MgCh supplemented with protease inhibitors. Lysates were incubated on ice for 10 min with occasional vortexing and then spun at 21000 g for 11 min at 4 °C. 1-2.2 mg of nuclear lysate was used for immunoprecipitation with rabbit anti-NFIA, rabbit anti-TEADL or rabbit anti- cJUN antibodies (1 :200 v / v) (Cell Signaling Technology) overnight at 4 °C. Protein-G Dynabeads (ThermoFisher) were then added for 2 hours. The beads were then extremely gently- washed on a magnet three times with EB150 supplemented with protease inhibitors to avoid disrupting low affinity interactions, followed by boiling in 4X LDS buffer for 7-10 min and loading onto SDS-PAGE gels for Western blotting. Antibodies are listed in STAR Methods.

[0187] Western blotting

[0188] Western blot analysis was performed using a standard protocol. Nuclear extracts were separated using a 4%-12% Bis-Tris PAGE gel (Bolt 4%-12%Bis-Tris Protein Gel,Thermo Fisher) and transferred onto 0.2 un Nitrocellulose membranes (Biorad) at 400 mA for 2 hours on ice. Membranes were blocked with 5% milk in IX TBST for 30 min at room temperature and then incubated with primary antibody overnight at 4 °C (1:2000 v / v for Cell Signaling antibodies, 1 : 1000 v / v for others). They were then washed thrice with IX TBST and incubated with near-infrared fluorophore-conjugated species-specific secondary antibodies (LI-COR Biosciences) for 1 hour at room temperature (1: 10,000 v / v). Following secondary antibody incubation, membranes were washed twice with IX TBST. once with IX TBS, and imaged using a Li-Cor Odyssey CLx imaging system (LI-COR Biosciences).

[0189] ATAC-seq

[0190] Omni-ATAC protocol was used to measure DNA accessibility with slight modifications covered below91. 100,000 cells per sample were trypsinized and washed with cold PBS to remove trypsin. Cell pellets were lysed in 50 pL cold resuspension buffer (RSB) supplemented with fresh NP40 (final 0.1% v / v), Tween-20 (final 0.1% v / v), Digitonin (final 0.01% v / v) (RSB recipe: 10 mM Tris-HCl pH 7.4, 10 mMNaCl, and 3 mM MgCh). Lysis step was quenched with 1 mL of RSB supplemented with Tween-20 (final 0.1% v / v) and nuclei were pelleted at 500 g for 10 min at 4 °C after incubating on ice for 3 minutes. Nuclei were then resuspended in 50 pL transposition reaction mix containing 25 pL 2X Tagment DNA buffer (Illumina), 2.5 pL Tn5 transposase (Illumina), 16.5 pL IX PBS, 0.5 pL 1% digitonin (final 0.01% v / v), 0.5 pL 10% Tween-20 (final 0.1% v / v), and 5 pL nuclease-free water. The transposition reaction was °C for 30 min with constant shaking (1000 rpm) on a thermomixer. Tagmented DNA was purified using the MinElute Reaction Cleanup Kit (Qiagen). Standard ATAC-seq amplification protocol with 7 cycles of amplification was used to amplify tagmented libraries45. Libraries were sequenced on aNextSeq 500 (Illumina) using 37 bp pairend sequencing.

[0191] CUT&Tag

[0192] CUT&Tag was performed as described previously62using a protocol developed by Epicypher (https: / / www.epicypher.com / content / documents / protocols / cutana-cut&tag- protocol.pdf) in 8-strip PCR tubes with slight modifications as described herein. Briefly, Concanavahn A (ConA) coated magnetic beads (Polysciences) were activated with Bead Activation Buffer containing 20 mM HEPES pH 7.9, 10 mM KC1, 1 mM CaCk, 1 mM MnCh; beads were stored on ice until used. 300,000 cells / sample were trypsinized and pelleted by centrifugation at room temperature (600g for 3 min). Cells were lysed using cold Nuclear Extraction Buffer containing 20 mM HEPES-KOH pH 7.9, 10 mM KC1, 0.1% Triton X-100,20% Glycerol supplemented with fresh 0.5 mM Spermidine and IX protease inhibitor (Roche) for 2 min. Nuclei were pelleted by centrifugation (600 g for 3 min), resuspended in 100 ul / sample Resuspension buffer (20 mM HEPES pH 7.5, 150 mM NaCl supplemented with fresh 0.5 mM Spermidine and IX protease inhibitor) and incubated with activated ConA beads at room temperature for 15 min. The nuclei-ConA bead complexes were then resuspended in Antibody 150 Buffer containing 20 mM HEPES pH 7.5, 150 mM NaCl, 2 mM EDTA supplemented with fresh 0.5 mM Spermidine. IX protease inhibitor, 0.01% Digitonin, and 0.5 ug primary antibody / sample. Following overnight incubation at 4°C on a nutator, supernatant was discarded, and the ConA-nuclei complexes were then incubated with Digitonin 150 buffer (20 mM HEPES pH 7.5, 150 mM NaCl, 0.5 mM Spermidine, IX protease inhibitor, 0.01% Digitonin) supplemented with 0.5 ug / sample Secondary antibody for 1 hour at room temperature on a nutator. They were then washed with Digitonin 150 Buffer twice before resuspension in 50 pL cold Digitonin 300 Buffer containing 20 mM HEPES, pH 7.5, 300 mM NaCl, 0.5 mM Spermidine, IX protease inhibitor, and 0.01% Digitonin. 2 pL CUT ANA pAG- Tn5 (Epicypher) was added to each sample and incubated on a nutator for 1 hr at room temperature. Following incubation, beads were washed twice with cold Digitonin 300 Buffer. Targeted chromatin tagmentation and library amplification were carried out according to Epicypher’s protocol mentioned above. Size distribution was measured on a DI 000 ScreemTape run on a TapeStation 2200 (Agilent). Equimolar amounts of barcoded libraries were pooled and sequenced on aNextSeq 500 (Illumina) using 37 bp pair-end sequencing with the goal of achieving a minimum of 8-10 million reads per library.

[0193] CUT&RUN

[0194] CUT&RUN was performed based largely on Epicypher’s protocol (https: / / www. epicypher. com / content / documents / protocols / cutana-cut&run-protocol.pdl) and the CUT&Tag protocol described herein but with key modifications as described herein. Briefly, Concanavalin A (ConA) coated magnetic beads (Polysciences) were activated with Bead Activation Buffer containing 20 mM HEPES pH 7.9, 10 mM KC1, 1 mM CaCh, 1 mM MnCh; beads were stored on ice until used. 500,000 cells / sample were trypsinized and pelleted by centrifugation at room temperature (600 g for 3 min). Cells were lysed using cold Nuclear Extraction Buffer containing 20 mM HEPES-KOH pH 7.9, 10 mM KC1, 0.1% Triton X-100, 20% Glycerol supplemented with fresh 0.5 mM Spermidine and IX protease inhibitor (Roche) for 2 min. Nuclei were pelleted by centrifugation (600 g for 3 min), resuspended in 100 ul / sample Resuspension buffer (20 mM HEPES pH 7.5, 150 mM NaCl supplemented withfresh 0.5 mM Spermidine and IX protease inhibitor) and incubated with activated ConA beads at room temperature for 15 min. The nuclei-ConA bead complexes were then resuspended in Antibody 150 Buffer containing 20 mM HEPES pH 7.5, 150 mM NaCl, 2 mM EDTA supplemented with fresh 0.5 mM Spermidine, IX protease inhibitor, 0.01% Digitonin, and 0.5 ug primary antibody / sample. Following overnight incubation at 4 °C on a nutator, the supernatant was discarded, and the ConA-nuclei complexes were then washed twice with Digitonin 150 buffer (20 mM HEPES pH 7.5, 150 mM NaCl, 0.5 mM Spermidine, IX protease inhibitor, 0.01% Digitonin). They were then resuspended in Digitonin 150 buffer and 2.5 pL of CUT ANA pAG-MNase (Epicypher) was added to each sample followed by incubation on a nutator for 30-60 min. The supernatant was then discarded and the ConA-nuclei complexes were washed twice with Digitonin 150 buffer and resuspended in fresh Digitonin 150 buffer supplemented followed by addition of 1 pL of 100 mM CaCh to each sample. The samples were then incubated on a nutator for 2 hours at 4 °C, followed by addition of the Stop buffer (340 mM NaCl, 20 mM EDTA, 4 mM EGTA, 50 ug / mL RNase A, 50 ug / ml Glycogen). Samples were then incubated at 37 °C for 10 minutes to release MNase-digested DNA fragments. The supernatants were then transferred to a new tube and DNA was purified using the MinElute Reaction Cleanup Kit (Qiagen). Libraries were prepared using the CUTANA CUT&RUN Library' Prep Kit (Epicypher). Size distribution was measured on a DI 000 ScreemTape run on a TapeStation 2200 (Agilent). Equimolar amounts of barcoded libraries were pooled and sequenced on aNextSeq 500 (Illumina) using 37 bp pair-end sequencing with the goal of achieving a minimum of 8-10 million reads per library'.

[0195] NGS Data Processing

[0196] CUT&Tag, CUT&RUN, ATAC-Seq, and RNA-Seq samples were sequenced on an Illumina NextSeqS 00 instrument. RNA-Seq reads were aligned to the hg!9 genome with STAR v2.5.2b77, and tracks were generated using the deepTools v2.5.3 bamCoverage function78with the normalizeUsingRPKM parameter. Output gene count tables from STAR were used as input into the edgeR v3.12.1 R software package77-80to evaluate differential gene expression. For ATAC-Seq data, read trimming was carried out by Trimmomatic v0.3681, followed by alignment, duplicate read removal, and read quality filtering using Bowtie282, Picard v2.8.0 (http: / / broadinstitute.github.io / picard / ), and SAMtools v 0.1. 1983, respectively, and ATAC-seq peaks were called with MACS2 v2.184using the BAMPE option and a broad peak cutoff of 0.001. For ATAC-Seq track generation, output BAM files were converted into BigWig files using MACS2 and UCSC utilities92in order to display coverage throughout the genome inRPM values. For CUT&Tag and CUT&RUN libraries, the CutRunTools pipeline was leveraged to perform read trimming, quality filtering, alignment, peak calling, and track building using default parameters85. Sequencing data analyzed in this study have been deposited at NCBI’s Gene Expression Omnibus under accession number GSE209961.

[0197] CUT&Tag, CUT&RUN and ATAC-seq data analyses

[0198] Heatmaps and metaplots displaying signals aligned to peak centers were generated using ngsplot v2.6386. RPM values were quantile normalized across samples, and K-means clustering was applied to partition the data into groups. The Bedtools multilntersectBed and merge functions were used for peak merging79, and distance-to-TSS peak distributions were computed utilizing Ensembl gene coordinates provided by the UCSC genome browser. Principle Component Analysis was performed using the wt. scale and fast.svd functions from the corpcor R package on CUT&Tag / CUT&RUN quantile normalized log2-transformed RPKM values within merged peaks87,88. Transcription factor motif enrichment analyses were carried out by the HOMER v4.993software.

[0199] cBAF complex purification

[0200] mSW / SNF complex purification was performed essentially as described previously21,41. Briefly, HEK293TARID1 A / B knock-out cells stably expressing HA-tagged ARID1A WT or mutants under a doxycycline-inducible promoter created using piggybac transfection (described herein) were plated in 50-100 15-cm plates. Expression of the bait (HA- ARID1A) was induced by addition of 200 ng / ml Dox for 48 hours. Cells were then scraped from plates, washed with cold PBS, and centrifuged at 5,000 rpm for 5 min at 4 °C. Pellets were resuspended in hypotonic buffer (HB: 10 mM Tris HC1 pH 7.5, 10 mM KC1, 1.5 mM MgCh, supplemented with 1 mM DTT and 1 mM PMSF) and incubated for 5 min on ice. The suspension was centrifuged at 5.000 rpm for 5 min at 4 °C. and pellets were resuspended in 5 volumes of HB containing protease inhibitor cocktail. The suspension was then homogenized using a glass Dounce homogenizer (Kimble Kontes). Nuclei were pelleted by centrifugation at 5000 rpm for 15 min at 4 °C. Nuclear pellets were resuspended in high salt buffer (HSB: 50 mM Tris HC1 pH 7.5, 300 mM KC1, 1 mM MgCbftmM EDTA, 1% NP40 supplemented with 1 mM DTT. 1 mM PMSF, and IX protease inhibitor cocktail). The homogenate was then incubated on a rotator for 1 hr at 4 °C followed by centrifugation at 20,000 rpm for 1 h at 4 °C using a SW32Ti rotor in an ultracentrifuge. The high salt nuclear extract supernatant was filtered through a 5 pm filter (EMD Millipore) and incubated with Pierce Anti-HA Magnetic Beads (Thermo Fisher) overnight at 4 °C. HA beads were washed 6 times in HSB and elutedwith HSB containing 2 mg / mL of HA peptide (GenScript) for four elutions of 2 h each followed by one overnight elution. Eluted proteins were then subjected to dialysis (Slide- A-Lyzer MINI Dialysis Device, 10K MWCO, ThermoFisher) using Dialysis Buffer (25 mM HEPES pH 8.0, 0. 1 mM EDTA, 100 mM KC1, 1 mM MgCh, 15% glycerol, and 1 mM DTT) overnight at 4 °C, and finally concentrated using Amicon Ultracentrifugal filters (30kDa MWCO, EMD Millipore). Complexes were aliquoted, flash frozen in liquid nitrogen and stored at -80 °C.

[0201] In vitro condensation assay

[0202] Purified cBAF complexes containing C-terminally eGFP-tagged ARID 1 A WT, DBDmut, AIDR1 or CBR were stored in 25 mM HEPES pH 8.0, 0. 1 mM EDTA, 100 mM KC1,1 mM MgC12,15% glycerol, and 1 mM DTT, at -80 °C. Reaction chambers for the in vitro assay were prepared by coating the interior glass of a 96-well glass bottom plate (Cellvis. P96-1.5H- N) with 1% w / v PF-127 (Pluronic F-127, ThermoFisher, P3000MP) for 15 minutes. Unused wells were filled with distilled water to maintain humidity in the nearby reaction chambers and prevent sample evaporation. Protein complexes were thawed on ice, then diluted to four concentrations (2, 0.66, 0.22, 0.074 pM) in physiological salt buffer (150 mM NaCl, 25 mM HEPES pH 7.5) in 4 pL reaction volume. For assays containing DNA, nucleosomes, or RNA. each reaction additionally contained 100 ng / pL DNA, 100 ng / pL nucleosomes, or 100 ng / pL RNA. Source of DNA was a linearized double-stranded 10 kb plasmid of random sequence. Nucleosomes were mono- di- and tri-nucleosomes purified from HeLa cells (Epicypher). Source of RNA was in vitro transcribed 18s rRNA fromHEK293T cell cDNA (ThermoFisher). Reactions were allowed to equilibrate at room temperature for 30 min for droplets to form and settle onto the coverslip. Visualization of the reaction chambers was performed on a spinningdisk confocal microscope (Yokogawa CSU-Xl) with 100X oil immersion Apo TIRF objective (NA 1.49) and Andor DU-897 EMCCD camera on a Nikon Eclipse Ti inverted microscope body. Images were obtained in DIC (Differential Interference Contrast) and GFP (488 nm laser) channels; at least 6 fields of view per sample were gathered. For quantification, images were deidentified, segmented for droplets in the GFP channel using FIJI90, and droplet area measured. "Percent Area’ metric was calculated for each image as the area in microns squared covered by droplets over the area in microns square of the entire field of view.

[0203] Fluorescence recovery after photobleaching (FRAP)

[0204] FRAP assays were performed in AN3CA patient-derived endometrial cells with doxycycline-inducible expression of C-terminally eGFP-tagged ARID1A or ARID1B constructs. Cells were plated in 24-well glass bottom plates (Cellvis) 48 hours prior to imaging.Expression was induced 24 hours prior to imaging by exchanging for media with 200 ng / mL doxycycline (Fisher Scientific). Cells were imaged on a Nikon Ti2 microscope equipped with an AIR HD25mm scanhead, with Plan Apo X 1.4 NA oil lens, maintained at 37 °C and 5% CO2 with a Tokai Hit Stagetop incubator equipped with a Ti ZWX stage insert. Images were obtained with 0.1-0.8 % laser power 488 nm with 10-70 HU gain at 11.11X zoom, 1 AU pinhole, 256x256 pixels each 0.0625 pm. Three pre-bleach images were acquired, then bleaching was performed with the 488 laser at 10% power. Post-bleach images acquired every 0.25 seconds for the first 10 seconds, every 1 sec for the next 20 sec, then every 5 sec for the next 2 minutes. For each construct, three biological replicates were prepared, and at least 15 cells bleached per replicate. For quantification, movies were registered using StackReg plugin in FIJI94, bleached area recognized by segmentation, then intensity of the bleached area in each frame measured. Measurements were normalized by subtracting the background (nucleoplasmic) intensity, then dividing over the average pre-bleach intensity from three prebleach images.

[0205] Live time-lapse movies

[0206] AN3CA cells expressing ARIDlA-WT-eGFP. ARIDlA-DBDmut-eGFP, ARID1B- WT-eGFP or ARIDlB-DBDmut-eGFP were prepared on the Nikon Ti2 microscope with AIR scanhead and Tokai Hit Stagetop incubator as described herein. Images were obtained with 0.1-0 8% laser power 488 nm with 10-70 HU gain at 4X zoom. lAu pinhole, 512x512 pixels. Images were obtained every 20 seconds for 60 minutes to observe long-term stability of condensates, or every 5 seconds for 10 minutes to observe fusion and coalescence of nuclear puncta.

[0207] Immunofluorescence

[0208] AN3CA cells were plated in 24-well glass bottom plates at 50% confluency (Cellvis) 48 hours prior to fixation, and expression induced 24 hours prior to fixation by adding 200 ng / rnE doxycycline (Fisher Scientific). Cells were washed once with DPBS, then fixed in 4% paraformaldehyde (diluted in DPBS from 16% paraformaldehyde, Electron Micrsocopy Science # 15710) for 15 minutes at room temperature. Fixed cells were washed three times, five minutes each in room temperature DPBS, then permeabilized in 0.2% PBST (Triton X-100. ThermoFisher) for 60 minutes with rocking. Permeabilized cells were washed again three times, five minutes each in room temperature DPBS, then blocked in 0. 1 % PBST with 5% goat serum (Vector Laboratories S-1000-20) + 5% BSA for 60 minutes with rocking. Cells were stained with primary’ antibody in block overnight at room temperature with rocking (1:500rabbit mAb anti-p300; 1:500 rabbit mAb anti-SMARCCl; 1: 1000 rabbit mAb anti-ARIDlA or 1 : 1000 rabbit mAb anti -ARID IB). Cells were washed three times, five minutes each with DPBS, then stained with secondary antibody (1:5000 Goat anti-rabbit highly cross-adsorbed 568-conjugated antibody) for 3 hours with rocking at room temperature. This antibody staining protocol was developed to faithfully recognize condensates in exogenous ARID1 A WT-eGFP- expressing AN3CA cells, then applied to the additional panel of cell types (KLE, CRL-7250, MDA-MB-231, MCF10A, MCF10-CA, C2C12 mouse myoblasts and primary rat cortex neurons). Antibodies are listed in the STAR Methods.

[0209] Saturation Concentration Measurement

[0210] Microscope fluorescence intensity’ to concentration calibration

[0211] Prior to imaging, the Nikon Al scanning confocal microscope and oil immersion objective (Plan Apo 60X / 1.4, Nikon) were calibrated for fluorescence-to-concentration conversion using Fluorescent Correlation Spectroscopy for mCherry and GFP (568 nm and 488 nm lasers) as in Bracha et al 201842. Briefly, mCherry fluorescence was converted to absolute concentration using FCS, then GFP fluorescence conversion was done by an exact mCherry- to-GFP fluorescence ratio with mCherry-P2A-eGFP construct. Diffusion and concentration were measured with 30 sec FCS measurement time, then a conversion table was created for fluorescence-intensity-to-concentration at specific optical settings. Activation was performed with a 488 nm excitation channel power of 84 uW / um2, measured w ith an optical power meter (PM100D, Thorlabs), and images obtained with 1% head power on 488 nm laser, with intensity 0.1 -1 %, gain between 10-70 HU, IX zoom, 1 AU pinhole (33.2 pm), 1024x1024 pixels.

[0212] Measuring saturation concentration

[0213] AN3CA cells were plated in 24-w-ell glass bottom plates (Cellvis) 48 hours prior to imaging, and expression induced 24 hours prior to imaging by adding 200 ng / mL doxycycline (Fisher Scientific). Live cells were imaged on a Nikon Al point-scanning laser confocal with 60X oil immersion lens of NA 1.4. Cells were maintained at 37 °C and 5% CO2 with Okolab stagetop incubation. To quantitatively determine saturation concentration, images of nuclei were obtained with calibrated settings, then nuclei segmented from background in FIJI by Otsu’s method and classified as having no condensates (no PS) or as having condensates (yes PS) by the variance in pixel intensity across a 4pm x 4 pm area within the nucleus that does not overlap a skewing feature like a nucleolus; those areas with no puncta have low7variance (<10% of mean intensity), while those with condensates have high variance (>10% of mean intensity). Concentrations of ARID1A / B in each nucleus were mapped and plotted, and thethreshold at which the ‘yes PS’ and ‘no PS’ categories are most separated by a logistic regression was marked as the Saturation Concentration.

[0214] Condensate count and area measurements

[0215] To quantify the number and size of condensates per nucleus, the identified nuclei counted as ‘yes PS’ were subjected to further image analysis. These images of a single z plane within nuclei were segmented by IsoData method in FIJI to recognize the puncta, then their count per nucleus and average size per nucleus was recorded using the Analyze Particles feature in FIJI. To account for cell-to-cell variability, in general three biological replicates were performed with greater than 100 cells measured in each replicate, then the averages of three replicates plotted with standard error shown as error bars.

[0216] Light cycling experiments

[0217] U2OS cells expressing the Corelet components were subjected to repeated on-off cycles of 488 laser exposure. To do this, the cells were imaged for three ‘pre-activation’ frames, one every five seconds, in only the mCherry (561 nm laser) channel. Then, images were acquired every 5 seconds for 3 minutes in both GFP and mCherry channels, which exposes them to 488 nm light and 'activates’ the Corelet system to form condensates. Droplets were then dissipated for 5 minutes by only imaging in the mCherry channel and reactivated again for two more cycles of (3 minutes activation + 5 minutes deactivation). Nuclei were registered using HyperStackReg in FIJI (doi:10.5281 / zenodo.2252521). then Pearson Correlation Coefficient of nuclear pixel intensities in the last frame of each activation cycle was calculated using the JaCoP plugin95.

[0218] Restriction Enzyme Accessibility Assay (REAA)

[0219] Purified cBAF complexes carrying ARID1 A WT or mutants were quantified using SMARCA4 protein levels via Western Blotting using SMARCA4 standards (Epicypher). Complexes were added to a 30 pL reaction containing 3 pL REAA buffer (20 mM HEPES pH 7.5, 5 mM Tris-HCL pH 7.5, 40 mM KC1, 2 mM MgCh), 1 mM DTT, 5 nM unmodified nucleosomes (Epidyne Nucleosome Remodeling Assay Substrate ST601-GATC1, 50-N-66, Biotinylated, Epicypher), 10 U / pL DpnII restriction enzyme (New England Biolabs), 0.5 mM ATP (Ultrapure ATP, Promega), °C in a PCR thermocycler. After incubation, 15 pL of the reaction was used to measure ATPase activity using ADP-Glo Max Assay kit (Promega). The rest of the reaction was quenched with 20 mM EDTA and 12 pg Proteinase K (Ambion) and incubated at 55 °C for 1 h and 80 °C for 10 min, followed by DNA purification using IXAMPure beads (Beckman Coulter) and DS 1000 High Sensitivity DNA ScreenTape analysis (Agilent).

[0220] ATPase activity measurement

[0221] 15 pL of the REAA reaction was transferred to a 96-well white bottom plate containing 5 pL water followed by addition and mixing of 20 pL of ADP Gio reagent. The plate was covered in aluminum foil and placed on a shaker for 1 hour. 40 pL of the ADP Gio detection reagent was then added and mixed, followed by another 1 hour incubation on the shaker with the plate covered in foil. Luminescence was measured using a spectrophotometer.

[0222] ARID domain purification

[0223] The ARID1A ARID domain (amino acids 958-1375) (wild-type and the DNA binding mutant S1086E, S1087E, S1091E) was cloned in an in-house bacterial expression vector downstream of a GST tag and transformed into E. coli Rosetta (DE3) cells. Colonies were grown in Terrific broth at 37 °C in the presence of 100 pg / ml Carbenecillin and 25 pg / ml Chloramphenicol until ODeoo was 0.7. Protein expression was then induced with 1 mM IPTG and the culture was incubated at room temperature for 5 hours at 225 rpm, following which cells were pelleted by centrifugation at 5000 rpm for 10 min. Pellets were washed once with cold PBS and frozen at -80 °C. For protein purification, pellets were resuspended in 40 ml cold Lysis buffer (50 mM Tris-HCl pH 7.5, 500 mM NaCl, 1% NP-40, 0.5 mg / ml lysozyme, 1 mM DTT) supplemented with protease inhibitors. Cells were lysed by sonication on ice and the lysate was centrifuged at 20,000 rpm for 1 hour at 4 °C. The clarified lysate was then incubated with magnetic Glutathione beads (ThermoFisher) (washed twice in lysis buffer) on a rotator for 2 hours at 4 °C. The beads were washed five times with Wash buffer (50 mM Tris-HCl pH 8, 500 mM NaCl, 1 mM DTT) supplemented with protease inhibitors. Five elutions were performed using Wash buffer supplemented with 20 mM reduced Glutathione (Boston Bioproducts). 10 pL of each elution fraction w as denatured in 2X LDS buffer and subjected to SDS-PAGE. The gel was stained with Coomassie Blue and fractions containing protein were pooled. The pooled fractions were buffer exchanged in dialysis buffer (25 mM HEPES pH 7.5, 100 mM KC1, 1 mM MgCh, 0.1 mM EDTA, 10% glycerol) overnight at 4 °C. Following dialysis, GST-ARID protein levels were quantified using Protein Qubit (ThermoFisher), aliquoted, flash frozen in liquid Nitrogen and stored at -80 °C.

[0224] Electrophoretic mobility shift assay (EMSA)

[0225] GST-ARID protein (WT or DBDmut) and an IRDye800-tagged dsDNA probe (random sequence) were incubated in 10 pL EMSA buffer (20 mM Tris-HCl pH 7.5, 20 mMNaCl, 20 mM KC1, 10% glycerol, 10 pg / ml BSA, 1 mM DTT) at room temperature for 30 min. Following incubation, 2 pL of Gel loading dye lacking SDS (New England Biolabs) was added to the reactions and run on 1% TAE agarose gels at 125 V for 20 min. Gels were then imaged using a Li-COR Odyssey CLx imaging system (LI-COR Biosciences).

[0226] 10-30% glycerol gradient sedimentation

[0227] Glycerol gradient-based sedimentation was performed as previously described21. 1 mg nuclear extracts were loaded on top of linear. 11 ml 10%-30% glycerol gradients containing 25 mM HEPES pH 8.0, 0.1 mM EDTA, 12.5 mM MgCb, 100 mM KC1 supplemented with 1 mM DTT and protease inhibitors. Tubes were then loaded into a SW41 rotor and centrifuged at 40,000 rpm for 16 hours at 4 °C. 550pL fractions were manually collected from the top of the gradient, to which 10 pL of Strataclean beads (Agilent) were added and incubated on a rotator for 1 hour at °C for 10 min. The mixture was then spun at 21000 g for 1 min, and the supernatants were loaded onto SDS-PAGE gels followed by Western blot analysis.

[0228] Proximity labeling and TMT Mass Spectrometry

[0229] Proximity labelling using TurboID

[0230] Proximity labelling was performed as previously described51,52. Briefly, no ligase Control or ARIDlA-TurboID (WT or mutant) fusion expressing AN3CA cells were treated with 200 pg / ml Doxycycline to induce gene expression for 48 h, following which cells were labelled with 50 pM Biotin (Sigma Aldrich) for 10 min. Media was aspirated and cells were washed five times with sterile cold PBS on the plate. They were then scraped and resuspended in EB0 hypotonic buffer containing 50 mM Tris pH7.5, 0.1 % NP-40, 1 mM EDTA, 1 mM MgCb supplemented with IX protease inhibitors. Lysates were pelleted at 5000 rpm for 5 min at 4 °C. Supernatants were discarded, and nuclei were resuspended in EB300 high salt buffer containing 50 mM Tris pH 7.5, 300 mM NaCl, 1% NP-40, 1 mM EDTA, 1 mM MgCb supplemented with IX protease inhibitors. Lysates were incubated on ice for 10 min with occasional vortexing and then spun at 21000 g for 11 min at 4 °C. Supernatants were quantified and supplemented with 1 mM DTT. 1.3 mg nuclear lysate was then incubated with magnetic Streptavidin beads (Thermo Fisher) on a rotator at 4 °C for 2 hours to isolate biotinylated proteins. Beads were then washed twice with EB300. once with 1 M KC1, and five times with 100 mM HEPES pH 8.0, following which they were resuspended in 100 pL of 100 mM HEPES pH 8.0 and flash frozen for mass spectrometry' analysis. Each sample was run in biological triplicate.

[0231] Protein Digestion

[0232] Beads were resuspended in 200 mM HEPES pH 8.5 and digested at room temperature for 13 h with Lys-C protease at a 100: 1 protein-to-protease ratio. Trypsin was then added at a 100: 1 ratio and the reaction was incubated 6 h at 37 °C. Peptides were separated from beads, vacuum centrifuged to near-dryness and desalted via StageTip.

[0233] Tandem mass tag labelling

[0234] For labeling, a final acetonitrile concentration of -30% (v / v) in 200 mM HEPES pH 8.5 was added along with 2 pL of TMT reagent (20 ng / mL) to the peptides in 25 pL total volume. Following incubation at room temperature for 1.5 h, the reaction was quenched with hydroxylamine to a final concentration of 0.3% (v / v) for 15 min. The TMT-labeled samples were pooled at a 1: 1 ratio across samples. The combined sample was vacuum centrifuged to near dryness and subjected to Cl 8 solid-phase extraction (SPE) via Sep-Pak (Waters, Milford, MA).

[0235] Off-line basic pH reversed phase (BPRP) fractionation

[0236] The pooled TMT-labeled peptide samples were fractionated using the Pierce High pH Reversed-Phase Peptide Fractionation Kit (ThermoFisher). Twelve fractions were collected using: 7.5%. 10%, 12.5%, 15%. 17.5%. 20%, 22.5%, 25%. 27.5%. 30%, 35%, and 60% acetonitrile and every sixth samples was concatenated, resulting in a total of six fractions per experiment. Samples were subsequently acidified with 1% formic acid and vacuum centrifuged to near dryness. Each fraction was desalted via StageTip, dried again via vacuum centrifugation, and reconstituted in 5% acetonitrile, 5% formic acid for LC-MS / MS processing.

[0237] Liquid chromatography and tandem mass spectrometry

[0238] Mass spectrometry data were collected using an Orbitrap Eclipse mass spectrometer (ThermoFisher) coupled to a Proxeon EASY-nLC 1200 liquid chromatography (LC) pump (ThermoFisher). Peptides were separated on a 100 mm inner diameter microcapillary column packed with -30 cm of Accucorel50 resin (2.6 pm, 150 A, ThermoFisher). For each analysis, we loaded -2 pg onto the column and separation was achieved using a 90 min gradient of 5 to 25% acetonitrile in 0. 125% formic acid at a flow rate of -450 nL / min. For the high-resolution MS2 (hrMS2) method, the scan sequence began with an MSI spectrum (Orbitrap analysis; resolution, 60.000: mass range. 400-1600 Th; automatic gain control (AGC) target 100%; maximum injection time, auto). Data were acquired with FAIMS using three CVs (-40V, -60V, and -80V) each with a 1 sec TopSpeed method. MS2 analysis consisted of high energy collision-induced dissociation (HCD) with the following settings: resolution, 50,000; AGCtarget, 200%; isolation width, 0.7 Th; normalized collision energy (NCE), 37; maximum injection time, 86 ms.

[0239] Data analysis

[0240] Mass spectra were processed using a Comet-based software pipeline96,97. Spectra were converted to mzXML using a modified version of ReAdW.exe. Database searching included entries from the human UniProt database. This database was concatenated with one composed of the protein sequences in the reversed order. Searches were performed using a 50- ppm precursor ion tolerance for total protein level profiling. TMTpro tags on lysine residues and peptide N termini (+304.207 Da) and carbamidomethylation of cysteine residues (+304.207 Da) were set as static modifications, while oxidation of methionine residues (+15.995 Da) was set as a variable modification. Peptide-spectrum matches (PSMs) were adjusted to a 1% false discovery rate (FDR)98,99PSM filtering was performed using a linear discriminant analysis, as described previously10°, while considering the following parameters: XCorr, ACn, missed cleavages, peptide length, charge state, and precursor mass accuracy. For TMT-based reporter ion quantitation, we extracted the summed signal-to-noise (S / N) ratio for each TMT channel and found the closest matching centroid to the mass of the TMT reporter ion. PSMs were identified, quantified, and collapsed to a 1% peptide false discovery rate (FDR) and then collapsed further to a final protein-level FDR of 1%. Moreover, protein assembly was guided by principles of parsimony to produce the smallest set of proteins necessary to account for observed peptides. Proteins were quantified by summing reporter ion counts across matching PSMs, as described previously10°. PSMs with poor quality and reporter summed signal-to-noise ratio less than 100, or no MS3 spectra were excluded from quantification101. Data from the samples were normalized to Acetyl-CoA Carboxylase signal (ACACA), an endogenously biotinylated protein in Streptavidin precipitations. ACACA was also omitted from downstream analyses. TMT signal of the No Ligase control was subtracted from samples of respective replicates after ACACA signal normalization and peptide filtering. Signals of replicates were averaged between replicates for downstream analyses. Unless otherwise noted, plots were generated using matplotlib and seaborn.

[0241] Coimmunoprecipitation followed by TMT mass spectrometry (IP-Mass Spec)

[0242] Cells were scraped from 1 -cm plates, washed with cold PBS and resuspended in EB0 hypotonic buffer containing 50 rnM Tris pH7.5, 0.1% NP-40, 1 rnM EDTA, 1 mM MgCb supplemented with IX protease inhibitors. Lysates were pelleted at 5000 rpm for 5 min at 4 °C. Supernatants were discarded, and nuclei were resuspended in EB150 salt buffer containing50 mM Tris pH 7.5, 150 mMNaCl, 1%NP-4O, 1 mM EDTA, 1 mM MgCk supplemented with IX protease inhibitors. Lysates were incubated on ice for 10 min with occasional vortexing. Nuclear lysate was pelleted at 21000 g for 11 min at 4 °C. Supernatants were quantified and supplemented with 1 mM DTT. 1.5 mg of nuclear lysate was used for immunoprecipitation with rabbit anti-HA antibody (Cell Signaling Technology ) overnight at 4 °Cf on a rotator to isolate cBAF complexes with HA- ARID 1 A as bait. Protein-G Dynabeads were then added and incubated on a rotator for 2 hours, washed thrice with EB150 and thrice with 100 mM HEPES pH 8.0. They were then resuspended in 5% formic acid to elute protein (2 elutions per sample, 50pL 5% formic acid per elution, 6 minutes incubation at room temperature per elution). Elutions were then pooled per sample and frozen at -80 °C.

[0243] Protein Digestion

[0244] Eluates were dried in a vacuum centrifuge and resuspended in 200 mM HEPES pH 8.5. Proteins were digested at room temperature for 13 h with Lys-C protease at a 100:1 protein-to-protease ratio. Trypsin was then added at a 100:1 ratio and the reaction was incubated 6 h at 37 °C.

[0245] Tandem mass tag labelling

[0246] For labeling, a final acetonitrile concentration of -30% (v / v) in 200 mM HEPES pH 8.5 was added along with 3 pL of TMT reagent (20 ng / pL) to the peptides in 25 pL total volume. Following incubation at room temperature for 1.5 h, the reaction was quenched with hydroxylamine to a final concentration of 0.3% (v / v) for 15 min. The TMT-labeled samples were pooled at a 1 : 1 ratio across samples. The combined sample was subsequently acidified with 1% formic acid and vacuum centrifuged to near dryness. The sample was desalted via StageTip, dried via vacuum centrifugation, and reconstituted in 5% acetonitrile, 5% formic acid for LC-MS / MS processing.

[0247] Liquid chromatography and tandem mass spectrometry

[0248] Mass spectrometry data were collected using an Orbitrap Fusion Eclipse mass spectrometer (ThermoFisher) coupled to a Proxeon EASY-nLC 1200 liquid chromatography (LC) pump (ThermoFisher). Peptides were separated on a 100 pm inner diameter microcapillary column packed with -30 cm of Accucorel50 resin (2.6 pm, 150 A, Thermo Fisher). For each analysis, we loaded one-half of the sample onto the column and separation was achieved using a 150 min gradient of 3 to 25% acetonitrile in 0.125% formic acid at a flow rate of -450 nL / min. For this high-resolution MS2 (hrMS2) method, the scan sequence began with an MS I spectrum (Orbitrap analysis; resolution, 120,000; mass range, 400-1500 Th;automatic gain control (AGC) target, "standard"; maximum injection time, “auto”). Data were acquired with FAIMS using three CVs (-40V, -60V, and -80V) each with a 1 sec. TopSpeed method. MS2 analysis consisted of high energy collision-induced dissociation (HCD) with the following settings: resolution, 50,000; AGC target, 300%; isolation width, 0.5 Th; normalized collision energy (NCE), 36; maximum injection time, 250 ms. The second half of the sample was re-analyzed with a similar method which had a different set of CVs (-30V, -50V, and - 70V).

[0249] Data analysis

[0250] Mass spectra were processed using a Comet-based software pipeline96,97. Spectra were converted to mzXML using a modified version of ReAdW.exe. Database searching included entries from the human UniProt database. This database was concatenated with one composed of the protein sequences in the reversed order. Searches were performed using a 50- ppm precursor ion tolerance for total protein level profiling. TMTpro tags on lysine residues and peptide N termini (+304.207 Da) and carbamidomethylation of cysteine residues (+304.207 Da) were set as static modifications, while oxidation of methionine residues (+15.995 Da) was set as a variable modification. Peptide-spectrum matches (PSMs) were adjusted to a 1% false discovery rate (FDR)98,99PSM filtering was performed using a linear discriminant analysis, as described previously10°, while considering the following parameters: XCorr, ACn, missed cleavages, peptide length, charge state, and precursor mass accuracy. For TMT-based reporter ion quantitation, we extracted the summed signal-to-noise (S / N) ratio for each TMT channel and found the closest matching centroid to the mass of the TMT reporter ion. PSMs were identified, quantified, and collapsed to a 1% peptide false discovery7rate (FDR) and then collapsed further to a final protein-level FDR of 1%. Moreover, protein assembly was guided by principles of parsimony to produce the smallest set of proteins necessary to account for observed peptides. Proteins were quantified by summing reporter ion counts across matching PSMs, as described previously10°. PSMs with poor quality7and reporter summed signal-to-noise ratio less than 100, or no MS3 spectra were excluded from quantification101. Scaled TMT values were normalized to the control by subtracting the scaled values of the corresponding replicate control from the scaled values of each condition. These control- normalized values were normalized to bait (ARID 1 A) by dividing the control normalized values for each condition by the control normalized values for ARID 1 A. Log-2 fold-changes between each condition and ARID 1 A WT were calculated using the mean control-bait- normalized values for each condition (any mean control-bait-normalized values less than 0were set to 0) with a pseudocount of 0.0001. Two-sample t-tests (n=2) with equal variance were used to calculate p-values. Only protein isoforms with the greatest detected peptide counts per gene were used for downstream analysis and visualization. Heatmaps were generated using the control -bait-normalized values. Volcano plots were generated using the log2 fold changes and p-values calculated as described herein. A log2FC = + / - 1 and p-value = 0.25 were used to define gained and lost proteins. Unless otherwise noted, plots were generated using matplotlib and seaborn.

[0251] Identification of non-random amino acid sequence features in disordered regions of mSWI / SNF subunits

[0252] The Swissprot database was used to download the Homo sapiens proteome (May 2015, 20882 entries). Disordered regions were then extracted from each protein sequence using MobiDB3,102. Specifically, a residue was considered disordered if the consensus prediction labeled it as being disordered. Then, consecutive disordered stretches greater than or equal to 30 residues in length were extracted to create what we refer to as the human IDRome, consisting of 24508 IDRs. Ninety sequence features previously found to be important for IDR conformational ensembles, phase separation, and function were calculated for IDRs in the human IDRome6,57. Sequence features are split into two broad categories: patterning and composition. To extract patterning z-scores we employed the NARDINI program6which calculates the degree of blockiness of groups of residues compared to 105randomly generated sequences with the same composition. Residues are grouped into the following eight types: polar=(Q, S, H, T, C, N), hydrophobic=(I, L, M, V), positive=(K, R), negative=(D, E), aromatic=(F, Y, W), alanine=A, proline=P, and glycine=G. Considering pairs of residue types leads to 36 patterning features. Positive z-scores indicate the patterning of the two residue types is more blocky than random, whereas negative z-scores indicate the patterning is more well- mixed than random.

[0253] Fifty-four compositional features were also calculated for each human IDR. localCIDER7w as utilized to calculate most of the compositional features including amino acid fractions (20 features), fraction of polar, aliphatic, aromatic, positive, negative, charged, chain expanding, and disorder promoting residues (8 features), the ratio of numbers of Rs to Ks and Es to Ds (2 features), and general features such as the net charge per residue, isoelectric point, hydrophobicity7, and polyproline II propensity (4 features). We also calculated 20 patch features defined as the fraction of the IDR in a specific residue or RG patch. Here, W w as excluded as no W patch was found in the human IDRome. A patch was calculated as a region of thesequence that had at least four occurrences of the given residue or two occurrences of RG and was not allowed to extend past two interruptions. Then, z-scores for each of the 54 compositional features were generated using the mean and standard deviation of the entire human IDRome. Here, positive z-scores indicate the compositional feature is enriched in the IDR of interest, whereas negative z-scores indicate the compositional feature is depleted in the IDR of interest.

[0254] Ninety-one IDRs were extracted from the human IDRome from the 29 mSWI / SNF proteins. Sequence feature z-score vectors of the IDRs were hierarchically clustered using the Euclidean distance and Ward’s linkage method. Only sequence features with a standard deviation > 0.1 across the 91 IDRs are shown in FIG. 6, panel A. The sequence features analyzed were divided into six categories: (1: red) patterning of X residues with Z residues. (2: orange) fraction of X residues, (3: green) fraction of IDR in X residue or RG patch, (4: blue) fraction of X+... +Z residues, (5: purple) ratio of number of X residues to Z residues, and (6: grey) additional compositional features calculated using localCIDER (http: / / pappulab.github.io / localCIDER / ). Four clusters were identified: Cluster 1 (red) consists only of the N-terminal ARID1A and ARID1B IDRs which are enriched in blocks of polar residues, alanines, and glycines. Cluster 2 (orange) consists of IDRs enriched in patches of prolines and glutamines, Cluster 3 (green) consists of highly negatively charged IDRs, and Cluster 4 (blue) consists of IDRs enriched in blocks of positive and negative residues.

[0255] To quantitatively determine the sequence features enriched / blocky in each of the four mSWI / SNF IDRome clusters, the z-score distributions from the IDRs in each cluster were compared to the z-score distributions of the remaining human IDRome. Colored values in FIG. 6, panel B indicate that sequence feature is more enriched or blockier in that cluster compared to the rest of the human IDRome. Specifically, the Kolmogorov-Smimov test was used to determine if the two distributions were identical and extract a p-value for each of the ninety sequence features. If the / ?- value was less than 0.05, then the signed logio(p-value) was calculated. A positive / negative logio(p-value) implies the mean z-score was greater / less than the cluster distribution compared to the distribution from the remaining human IDRome. Only features with signed logio(p-value) greater than zero for at least one cluster are shown in FIG.6, panel B.

[0256] NARDINI plots

[0257] Non-random sequence patterning features of the ARID 1 A and ARID IB sequences are calculated6. The 20 canonical amino acids are grouped into eight categories: polar,hydrophobic, positively charged, negatively charged, aromatic, Ala, Pro, and Gly. Here, the polar residue categories are further broken down to Q, S, H, and TCN, as noted in the figure legends. The z-scores are calculated with respect to the null model of 103randomly scrambled sequences with fixed amino acid composition. Z-scores > 0 indicate clustering of residue category' into blocks in the linear sequence, whereas z-score < 0 indicate that the residues are evenly distributed, or well-mixed, throughout the sequence.

[0258] Amino acid sequence patterning of the N-terminal IDR of eukaryotic ARID1A orthologs

[0259] Eukaryotic ARID1 A ortholog sequences were obtained from the EggNOG database (KOG2510, N = 3O7)103. The N-terminal intrinsically disordered regions were extracted for the analysis. In the heatmap, each row corresponds to an ARID1A ortholog sequence, and each column corresponds to the z-score of a sequence feature. H. sapiens ARID1 A IDR sequence is outlined in black. Each sequence (row) is color-coded by its taxonomic ranks in phylum, class, and order. The sequence patterning features were calculated as described70. The sequence composition features were calculated from the primary sequence features and the z-scores were calculated with respect to the null model of the ortholog sequences. The dendrogram was generated using the Frobenius norm of the z-score matrices, where the norms were used as Euclidean distances, and Ward’s clustering was used.

[0260] Mapping of cancer- and neurodevelopmental disorder-associated mutations on to ARID1A / B non-random pattern blocks

[0261] Cancer- and neurodevelopmental disorder-associated ARID1 A / B mutations were obtained from Valencia, Sankar et al., Nature Genetics 2023104Duplicate amino acid mutations were eliminated from the analyses. Silent mutations were not considered. Mutations found in the list were categorized into five categories: deletion, insertion, substitution, frameshift, and complex. Complex mutations indicate occurrence of more than one type of mutation.

[0262] References Cited in this Example:

[0263] 1 Oates, M. E. et al. (2013). D(2)P(2): database of disordered protein predictions. Nucleic Acids Res 41, D508-516, doi: 10. 1093 / nar / gksl226.

[0264] 2 Frege, T. & Uversky, V. N. (2015). Intrinsically disordered proteins in the nucleus of human cells. Biochem Biophys Rep 1, 33-51, doi: 10.1016 / j.bbrep.2015.03.003.

[0265] 3 Piovesan, D. et al. (2021). MobiDB: intrinsically disordered proteins in 2021. Nucleic Acids Res 49, D361-D367, doi: 10. 1093 / nar / gkaal058.

[0266] 4 Konrat, R. (2014). NMR contributions to structural dynamics studies of intrinsically disordered proteins. J Magn Reson 241, 74-85, doi: 10. 1016 / j.jmr.2013. 11.011.

[0267] 5 Cermakova, K. & Hodges, H. C. (2023). Interaction modules that impart specificity to disordered protein. Trends Biochem Set 48, 477-490, doi: 10. 1016 / j .tibs.2023.01.004.

[0268] 6 Cohan, M. C., Shinn, M. K., Lalmansingh, J. M. & Pappu, R. V. (2022). Uncovering Non-random Binary Patterns Within Sequences of Intrinsically Disordered Proteins. J Mol Biol 434, 167373, doi: 10.1016 / j.jmb.2021. 167373.

[0269] 7 Holehouse, A. S., Das, R. K., Ahad, J. N., Richardson, M. O. & Pappu, R. V. (2017). CIDER: Resources to Analyze Sequence-Ensemble Relationships of Intrinsically Disordered Proteins. Biophys . / 1 12. 16-21, doi: 10. 1016 / j .bpj .2016. 11.3200.

[0270] 8 Holehouse, A. S. in Intrinsically Disordered Proteins (ed Nicola Salvi) 209- 255 (Academic Press, 2019).

[0271] 9 Kadoch, C. et al. (2013). Proteomic and bioinformatic analysis of mammalian SWI / SNF complexes identifies extensive roles in human malignancy. Nat Genet 45, 592-601, doi: 10.1038 / ng.2628.

[0272] 10 Shain, A. H. & Pollack, J. R. (2013). The spectrum of SWI / SNF mutations, ubiquitous in human cancers. PLoS One 8, e55119, doi:10.1371 / joumal.pone.0055119.

[0273] 11 McBride, M. J. et al. (2018). The SS18-SSX Fusion Oncoprotein Hijacks BAF Complex Targeting and Function to Drive Synovial Sarcoma. Cancer Cell 33, 1128-1141 el 127, doi: 10.1016 / j.ccell.2018.05.002.

[0274] 12 Nakayama, R. T. et al. (2017). SMARCB1 is required for widespread BAF complex-mediated activation of enhancers and bivalent promoters. Nat Genet 49, 1613-1623, doi: 10.1038 / ng.3958.

[0275] 13 St Pierre, R. et al. (2022). SMARCE1 deficiency generates a targetable mSWI / SNF dependency in clear cell meningioma. Nat Genet, doi:10.1038 / s41588-022-01077- 0.

[0276] 14 Bogershausen, N. & Wollnik, B. (2018). Mutational Landscapes and Phenotypic Spectrum of SWI / SNF-Related Intellectual Disability Disorders. Front Mol Neurosci 11, 252, doi: 10.3389 / fnmol.2018.00252.

[0277] 15 Hanly, C , Shah, H„ Au, P. Y. B. & Murias, K. (2021). Description of neurodevelopmental phenotypes associated with 10 genetic neurodevelopmental disorders: A scoping review. Clin Genet 99, 335-346, doi: 10. 1111 / cge. 13882.

[0278] 16 Santen, G. W. et al. (2012). Mutations in SWI / SNF chromatin remodeling complex gene ARID IB cause Coffin-Siris syndrome. Nat Genet 44, 379-380, doi: 10.1038 / ng.2217.

[0279] 17 Santen, G. W. et al. (2013). Coffin-Siris syndrome and the BAF complex: genotype-phenotype study in 63 patients. Hum Mutat 34, 1519-1528, doi: 10.1002 / humu.22394.

[0280] 18 Santen. G. W.. Kriek, M. & van Attikum, H. (2012). SWI / SNF complex in disorder: Switching from malignancies to intellectual disability. Epigenetics 7, 1219-1224, doi: 10.4161 / epi.22299.

[0281] 19 Satterstrom, F. K. et al. (2020). Large-Scale Exome Sequencing Study Implicates Both Developmental and Functional Changes in the Neurobiology of Autism. Cell 180, 568-584 e523, doi: 10. 1016 / j.cell.2019. 12.036.

[0282] 20 Wright, C. F. et al. (2018). Making new genetic diagnoses with old data: iterative reanalysis and reporting from genome- wide data in 1,133 families with developmental disorders. Genet Med ia, 1216-1223, doi: 10.1038 / gim.2017.246.

[0283] 21 Mashtahr, N. et al. (2018). Modular Organization and Assembly of SWI / SNF Family Chromatin Remodeling Complexes. Cell 175, 1272-1288 el220, doi: 10.1016 / j. cell.2018.09.032.

[0284] 22 Kadoch, C. & Crabtree, G. R. (2015). Mammalian SWI / SNF chromatin remodeling complexes and cancer: Mechanistic insights gained from human genomics. Set Adv 1, el 500447, doi: 10. 1126 / sciadv. 1500447.

[0285] 23 Bailey, M. H. et al. (2018). Comprehensive Characterization of Cancer Driver Genes and Mutations. Cell 174, 1034-1035, doi: 10. 1016 / j.cell.2018.07.034.

[0286] 24 Lawrence, M. S. et al. (2014). Discovery and saturation analysis of cancer genes across 21 tumour types. Nature 505, 495-501, doi: 10. 1038 / naturel2912.

[0287] 25 Han, Y„ Reyes, A. A., Malik, S. & He, Y. (2020). Cryo-EM structure of SWI / SNF complex bound to a nucleosome. Nature 579, 452-455, doi: 10. 1038 / s41586-020- 2087-1.

[0288] 26 He. S. et al. (2020). Structure of nucleosome-bound human BAF complex. Science 367, 875-881, doi:10.1126 / science.aaz9761.

[0289] 27 Mashtahr, N. et al. (2020). A Structural Model of the Endogenous Human BAF Complex Informs Disease Mechanisms. Cell 183, 802-817 e824, doi : 10.1016 / j . cell.2020.09.051.

[0290] 28 Necci, M., Piovesan, D., Clementel, D., Dosztanyi, Z. & Tosatto, S. C. E. (2020). MobiDB-lite 3.0: fast consensus annotation of intrinsic disorder flavours in proteins. Bioinformatics, doi: 10. 1093 / bioinformatics / btaal045.

[0291] 29 Iglesias, V. et al. (2019). In silico Characterization of Human Prion-Like Proteins: Beyond Neurological Diseases. Front Physiol 10, 314, doi: 10.3389 / fphys.2019.00314.

[0292] 30 Boija. A. et al. (2018). Transcription Factors Activate Genes through the Phase- Separation Capacity of Their Activation Domains. Cell 175, 1842-1855 el816, doi:10.1016 / j.cell.2018.10.042.

[0293] 31 Wei, M. T. et al. (2020). Nucleated transcriptional condensates amplify gene expression. Nat Cell Biol 22, 1187-1196, doi: 10. 1038 / s41556-020-00578-6.

[0294] 32 Koga, S., Williams, D. S., Perriman, A. W. & Mann, S. (2011). Peptidenucleotide microdroplets as a step towards a membrane-free protocell model. Nat Chem 3, 720- 724, doi: 10.1038 / nchem.H 10.

[0295] 33 Shin, Y. & Brangwynne, C. P. (2017). Liquid phase condensation in cell physiology and disease. Science 357, doi: 10.1 126 / science.aaf4382.

[0296] 34 Strulson, C. A., Molden, R. C , Keating, C. D. & Bevilacqua, P. C. (2012). RNA catalysis through compartmentalization. Nat Chem 4, 941-946, doi:10.1038 / nchem. 1466.

[0297] 35 Strom, A. R. et al. (2017). Phase separation drives heterochromatin domain formation. Nature 547, 241-245. doi: 10. 1038 / nature22989.

[0298] 36 Larson, A. G. et al. (2017). Liquid droplet formation by HP1 alpha indicates a role for phase separation in heterochromatin. Nature 547, 236-240, doi:10.1038 / nature22822.

[0299] 37 Sabari, B. R. et al. (2018). Coactivator condensation at super-enhancers links phase separation and gene control. Science 361, doi: 10. 1126 / science.aar3958.

[0300] 38 Klein, I. A. et al. (2020). Partitioning of cancer therapeutics in nuclear condensates. Science 368, 1386-1392, doi: 10.1126 / science.aaz4427.

[0301] 39 Banani, S. F. et al. (2022). Genetic variation associated with condensate dysregulation in disease. Dev Cell, doi: 10.1016 / j.devcel.2022.06.010.

[0302] 40 Morin, J. A. et al. (2022). Sequence-dependent surface condensation of a pioneer transcription factor on DNA. Nature Physics 18, 271-276, doi: 10. 1038 / s41567-021- 01462-2.

[0303] 41 Valencia, A. M. et al. (2019). Recurrent SMARCB1 Mutations Reveal a Nucleosome Acidic Patch Interaction Site That Potentiates mSWI / SNF Complex Chromatin Remodeling. Cell 179, 1342-1356 el323, doi: 10.1016 / j.cell.2019.10.044.

[0304] 42 Bracha, D. et al. (2019). Mapping Local and Global Liquid Phase Behavior in Living Cells Using Photo-Oligomerizable Seeds. Cell 176, 407, doi:10.1016 / j.cell.2018.12.026.\

[0305] 43 Kaya-Okur. H. S., Janssens, D. H.. Henikoff, J. G., Ahmad. K. & Henikoff, S. (2020). Efficient low-cost chromatin profiling with CUT&Tag. Nat Protoc 15, 3264-3283, doi: 10.1038 / s41596-020-0373-x.

[0306] 44 Buenrostro, J. D., Giresi. P. G., Zaba, L. C.. Chang, H. Y. & Greenleaf, W. J. (2013). Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin, DNA-binding proteins and nucleosome position. Nat Methods 10, 1213-1218, doi: 10. 1038 / nmeth.2688.

[0307] 45 Buenrostro, J. D„ Wu, B„ Chang, H. Y. & Greenleaf, W. J. (2015). ATAC-seq: A Method for Assaying Chromatin Accessibility Genome-Wide. Curr Protoc Mol Biol 109, 21 29 21-21 29 29, doi: 10.1002 / 0471142727.mb2129sl09.

[0308] 46 Pan, J. et al. (2019). The ATPase module of mammalian SWI / SNF family complexes mediates subcomplex identity and catalytic activity-independent genomic targeting. Nat Genet 51, 618-626, doi: 10.1038 / s41588-019-0363-5.

[0309] 47 Alver, B. H. et al. (2017). The SWI / SNF chromatin remodelling complex is required for maintenance of lineage specific enhancers. Nat Commun 8, 14648, doi: 10. 1038 / ncomms 14648.

[0310] 48 Vierbuchen, T. et al. (2017). AP-1 Transcription Factors and the BAF Complex Mediate Signal-Dependent Enhancer Selection. Mol Cell 68, 1067-1082 el012, doi:10.1016 / j.molcel.2017.1 1.026.

[0311] 49 Gatchalian, J. et al. (2018). A non-canonical BRD9-containing BAF chromatin remodeling complex regulates naive pluripotency in mouse embryonic stem cells. Nat Commun 9, 5139, doi: 10.1038 / s41467-018-07528-9.

[0312] 50 Michel. B. C. et al. (2018). A non-canonical SWI / SNF complex is a synthetic lethal target in cancers driven by BAF complex perturbation. Nat Cell Biol 20, 1410-1420, doi: 10.1038 / s41556-018-0221-l.

[0313] 51 Branon. T. C. et al. (2018). Efficient proximity labeling in living cells and organisms with TurboID. Nat Biotechnol 36, 880-887, doi: 10.1038 / nbt.420L

[0314] 52 Cho, K. F. et al. (2020). Proximity labeling in mammalian cells with TurboID and split-TurboID. Nat Protoc 15, 3971-3999. doi: 10.1038 / s41596-020-0399-0.

[0315] 53 Boulay, G. et al. (2017). Cancer-Specific Retargeting of BAF Complexes by a Prion-like Domain. Cell 171, 163-178 el l9, doi: 10.1016 / j.cell.2017.07.036.

[0316] 54 Sandoval, G. J. et al. (2018). Binding of TMPRSS2-ERG to BAF Chromatin Remodeling Complexes Mediates Prostate Oncogenesis. Mol Cell 71, 554-566 e557, doi: 10. 1016 / j.molcel.2018.06.040.

[0317] 55 Shin, Y. et al. (2017). Spatiotemporal Control of Intracellular Phase Transitions Using Light-Activated optoDroplets. Cell 168, 159-171 el l4, doi: 10.1016 / j . cell.2016.11.054.

[0318] 56 Ruff, K. M. et al. (2022). Sequence grammar underlying the unfolding and phase separation of globular proteins. Mol Cell 82, 3193-3208 e3198, doi : 10. 1016 / j .molcel.2022.06.024.

[0319] 57 Zarin, T. et al. (2021). Identifying molecular features that are associated with biological function of intrinsically disordered protein regions. Elife 10, doi: 10.7554 / eLife.60220.

[0320] 58 Lin. Y.. Currie. S. L. & Rosen, M. K. (2017). Intrinsically disordered sequences enable modulation of protein phase separation through distributed tyrosine motifs. J Biol Chem 292, 19110-19120, doi:10.1074 / jbc.M117.800466.

[0321] 59 Brangwynne, Clifford P., Tompa, P. & Pappu, Rohit V. (2015). Polymer physics of intracellular phase transitions. Nature Physics 11. 899-904. doi: 10.1038 / nphys3532.

[0322] 60 Farag, M. et al. (2022). Condensates formed by prion-like low-complexity domains have small-world network structures and interfaces defined by expanded conformations. Nat Commun 13. 7722, doi: 10.1038 / s41467-022-35370-7.

[0323] 61 Papillon, J. P. N. et al. (2018). Discovery of Orally Active Inhibitors of Brahma Homolog (BRM) / SMARCA2 ATPase Activity for the Treatment of Brahma Related Gene 1 (BRG1) / SMARCA4-Mutant Cancers. J Med Chem 61, 10155-10172, doi : 10. 1021 / acs .jmedchem.8b01318.

[0324] 62 Wei, J. et al. (2023). Pharmacological disruption of mSWI / SNF complex activity restricts SARS-CoV-2 infection. Nat Genet 55. 471-483, doi: 10. 1038 / s41588-023- 01307-z.

[0325] 63 lurlaro, M. et al. (2021). Mammalian SWI / SNF continuously restores local accessibility to chromatin. Nat Genet 53, 279-287, doi: 10. 1038 / s41588-020-00768-w.

[0326] 64 Martin, E. W. et al. (2020). Valence and paterning of aromatic residues determine the phase behavior of prion-like domains. Science 367, 694-699, doi : 10. 1126 / science. aaw8653.

[0327] 65 Not, T. J. et al. (201 ). Phase transition of a disordered nuage protein generates environmentally responsive membraneless organelles. Mol Cell 57, 936-947, doi: 10.1016 / j.molcel.2015.01.013.

[0328] 66 Wang, J. et al. (2018). A Molecular Grammar Governing the Driving Forces for Phase Separation of Prion-like RNA Binding Proteins. Cell 174, 688-699 e616, doi: 10.1016 / j. cell.2018.06.006.

[0329] 67 Bremer, A. et al. (2022). Deciphering how naturally occurring sequence features impact the phase behaviours of disordered prion-like domains. Nat Chem 14, 196-207, doi : 10. 1038 / s41557-021 -00840-w.

[0330] 68 Kar, M. et al. (2022). Phase-separating RNA-binding proteins form heterogeneous distributions of clusters in subsaturated solutions. Proc Natl Acad Sci U S A 119, e2202222119, doi: 10.1073 / pnas.2202222119.

[0331] 69 Pappu, R. V., Cohen. S. R„ Dar, F„ Farag, M. & Kar. M. (2023). Phase Transitions of Associative Biomacromolecules. Chem Rev, doi: 10. 1021 / acs.chemrev.2c00814.

[0332] 70 Shinn, M. K. et al. (2022). Connecting sequence features within the disordered C-terminal linker of Bacillus subtilis FtsZ to functions and bacterial cell division. Proc Nall Acad Sci USA 119. e2211178119, doi: 10.1073 / pnas.2211178119.

[0333] 71 Bergeron-Sandoval, L. P. et al. (2021 ). Endocytic proteins with prion-like domains form viscoelastic condensates that enable membrane remodeling. Proc Natl Acad Sci USA 118, doi: 10.1073 / pnas.2113789118.

[0334] 72 Staller, M. V. et al. (2022). Directed mutational scanning reveals a balance between acidic and hydrophobic residues in strong human activation domains. Cell Syst 13, 334-345 e335, doi:10.1016 / j.cels.2022.01.002.

[0335] 73 Zeng, X., Ruff, K. M. & Pappu, R. V. (2022). Competing interactions give rise to two-state behavior and switch-like transitions in charge-rich intrinsically disordered proteins. Proc Natl Acad Sci USA 119, e2200559119. doi: 10.1073 / pnas.2200559119.

[0336] 74 Greig, J. A. et al. (2020). Arginine-Enriched Mixed-Charge Domains Provide Cohesion for Nuclear Speckle Condensation. Mol Cell 77, 1237-1250 el234, doi : 10. 1016 / j . molcel.2020.01.025.

[0337] 75 Sherry. K. P., Das, R. K„ Pappu, R. V. & Barrick, D. (2017). Control of transcriptional activity by design of charge patterning in the intrinsically disordered RAM region of the Notch receptor. Proc Natl Acad Sci U S A 114, E9243-E9252, doi: 10. 1073 / pnas. 1706083114.

[0338] 76 Santner, S. J. et al. (2001). Malignant MCF10CA1 cell lines derived from premalignant human breast epithelial MCF10AT cells. Breast Cancer Res Treat 65, 101-110, doi: 10. 1023 / a: 1006461422273.

[0339] 77 Dobin, A. et al. (2013). STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29 , 15-21, doi: 10.1093 / bioinformatics / bts635.

[0340] 78 Ramirez, F. et al. (2016). deepTools2: a next generation web server for deepsequencing data analysis. Nucleic Acids Res 44. W160-165. doi: 10.1093 / nar / gkw257.

[0341] 79 Quinlan, A. R. & Hall, I. M. (2010). BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841-842, doi: 10.1093 / bioinformatics / btq033.

[0342] 80 Love, M. I., Huber, W. & Anders, S. (2014). Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol 15, 550, doi: 10. 1186 / S13059-014-0550-8.

[0343] 81 Bolger, A. M., Lohse, M. & Usadel, B. (2012). Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics 30, 2114-2120, doi: 10. 1093 / bioinformatics / btul70 (2014).

[0344] 82 Langmead, B. & Salzberg, S. L. (2012). Fast gapped-read alignment with Bowtie 2. Nat Methods 9, 357-359, doi: 10. 1038 / nmeth. 1923.

[0345] 83 Li, H. et al. (2009). The Sequence Alignment / Map format and SAMtools. Bioinformatics 25, 2078-2079, doi: 10.1093 / bioinformatics / btp352.

[0346] 84 Zhang, Y. et al. (2008). Model-based analysis of ChlP-Seq (MACS). Genome Biol 9, R137, doi: 10.1186 / gb-2008-9-9-rl37.

[0347] 85 Zhu, Q„ Liu, N„ Orkin, S. H. & Yuan, G. C. (2019). CUT&RUNTools: a flexible pipeline for CUT&RUN processing and footprint analysis. Genome Biol 20, 192, doi: 10.1186 / sl3059-019-1802-4.

[0348] 86 Shen. L., Shao. N.. Liu, X. & Nestler, E. (2014). ngs.plot: Quick mining and visualization of next-generation sequencing data by integrating genomic databases. BMC Genomics 15, 284, doi: 10.1186 / 1471-2164-15-284.

[0349] 87 Schafer, J. & Shimmer, K. (2005). A shrinkage approach to large-scale covariance matrix estimation and implications for functional genomics. Stat Appl Genet Mol Biol 4, Article32, doi: 10.2202 / 1544-6115. 1175.

[0350] 88 Opgen-Rhein, R. & Shimmer, K. (2007). Accurate ranking of differentially expressed genes by a distribution-free shrinkage approach. Stat Appl Genet Mol Biol 6, Article^ doi: 10.2202 / 1544-6115. 1252.

[0351] 89 Heinz, S. et al. (2010). Simple combinations of lineage-determining transcription factors prime cis-regulatory elements required for macrophage and B cell identities. Mol Cell 38, 576-589, doi: 10.1016 / j.molcel.2010.05.004.

[0352] 90 Schindelin, J. et al. (2012). Fiji: an open-source platform for biological-image analysis. Nat Methods 9. 676-682. doi: 10.1038 / nmeth.2019.

[0353] 91 Corces, M. R. et al. (2017). An improved ATAC-seq protocol reduces background and enables interrogation of frozen tissues. Nat Methods 14, 959-962, doi: 10. 1038 / nmeth.4396.

[0354] 92 Kuhn, R. M., Haussler, D. & Kent, W. J. (2013). The UCSC genome browser and associated tools. Brief Bioinform 14, 144-161, doi: 10. 1093 / bib / bbs038.

[0355] 93 Heinz, S. et al. (2010). Simple Combinations of Lineage-Determining Transcription Factors Prime cis-Regulatory Elements Required for Macrophage and B Cell Identities. Molecular cell 38, 576-589, doi: 10.1016 / j.molcel.2010.05.004.

[0356] 94 Thevenaz, P., Ruttimann, U. E. & Unser. M. (1998). A pyramid approach to subpixel registration based on intensity. IEEE Trans Image Process 7, 27-41 , doi: 10. 1109 / 83.650848.

[0357] 95 Bolte. S. & Cordelieres, F. P. (2006). A guided tour into subcellular colocalization analysis in light microscopy. J Microsc 224, 213-232. doi: 10.1111 / j. 1365- 2818.2006.01706.x.

[0358] 96 Eng, J. K. et al. (2015). A deeper look into Comet— implementation and features. J Am Soc Mass Spectrom lG, 1865-1874, doi: 10.1007 / sl3361-015-1179-x.

[0359] 97 Eng, J. K, Jahan, T. A. & Hoopmann, M. R. (2013). Comet: an open-source MS / MS sequence database search tool. Proteomics 13, 22-24, doi: 10.1002 / pmic.201200439.

[0360] 98 Elias, J. E. & Gygi, S. P. (2010). Target-decoy search strategy for mass spectrometry-based proteomics. Methods Mol Biol 604, 55-71, doi: 10. 1007 / 978-1-60761-444- 9_5.

[0361] 99 Elias, J. E. & Gygi, S. P. (2007). Target-decoy search strategy for increased confidence in large-scale protein identifications by mass spectrometry. Nat Methods 4. 207- 214, doi: 10.1038 / nmethl019.

[0362] 100 Huttlin, E. L. et al. (2010). A tissue-specific atlas of mouse protein phosphory lation and expression. Cell 143, 1174-1189, doi: 10.1016 / j .cell.2010.12.001.

[0363] 101 McAlister, G. C. et al. (2012). Increasing the multiplexing capacity of TMTs using reporter ion isotopologues with isobaric masses. Analytical chemistry 84, 7469-7478. doi:10.1021 / ac301572t.

[0364] 102 UniProt, C. (2021). UniProt: the universal protein knowledgebase in 2021.Nucleic Acids Res 49, D480-D489, doi: 10. 1093 / nar / gkaal l00.

[0365] 103 Huerta-Cepas, J. et al. (2019). eggNOG 5.0: a hierarchical, functionally and phylogenetically annotated orthology resource based on 5090 organisms and 2502 viruses. Nucleic Acids Res 47, D309-D314, doi: 10. 1093 / nar / gkyl085.

[0366] 104 Valencia, A. M. et al. (2023). Landscape of mSWI / SNF chromatin remodeling complex perturbations in neurodevelopmental disorders. Nat Genet, doi: 10.1038 / s41588-023- 01451-6.EQUIVALENTS

[0367] Those skilled in the art will recognize, or be able to ascertain, using no more than routine experimentation, numerous equivalents to the specific substances and procedures described herein. Such equivalents are considered to be within the scope of this invention and are covered by the following claims.

Claims

What is claimed:

1. An in vitro screening method to identify a SWI / SNF modulating compound, wherein the method comprises: incubating a candidate compound with a protein or protein complex, wherein the protein or protein complex comprises an ARID1 A or ARID1B intrinsically disordered region (IDR) or fragment thereof, and detecting at least one IDR function, wherein modulation of the IDR function is indicative of a SWI-SNF modulating compound.

2. The in vitro screening method of claim 1, wherein the candidate compound binds to the IDR.

3. The in vitro screening method of claim 1, wherein the IDR function is selected from the group consisting of (i) condensate formation, (ii) chromatin localization, (iii) IDR protein engagement, and (iv) gene expression.

4. The in vitro screening method of claim 1, wherein the ARID 1 A or ARID IB intrinsically disordered region or fragment thereof, is purified from mammalian cells.

5. The in vitro screening method of claim 1, wherein the protein or protein complex comprising an ARID 1 A or ARID IB IDR or fragment thereof is selected from the group consisting of a purified polypeptide, a chromatin remodeling complex, a DNA- binding protein, an epitope-tagged peptide, a core binding region (CBR) peptide, or any combination thereof.

6. The in vitro screening method of claim 5, wherein the chromatin remodeling complex comprises an ATP-dependent chromatin remodeling complex.

7. The in vitro screening method of claim 6, wherein the ATP-dependent chromatin remodeling complex comprises the mammalian SWI / SNF chromatin remodeling complex.

8. The in vitro screening method of claim 1, wherein the candidate compound comprises a small molecule.

9. The in vitro screening method of claim 1, wherein detecting comprises a sequencing technique, an immunoassay, mass spectrometry, immunoprecipitation, or a combination thereof.

10. The in vitro screening method of claim 1, further comprising comparing modulation of the IDR function by the candidate compound to that of a control compound.

11. The in vitro screening method of claim 10, wherein the control compound comprises a positive control, a negative control, or both.

12. The in vitro screening method of claim 1, further comprising selecting the candidate compound as a therapeutic agent when the candidate modulates the IDR function.

13. The in vitro screening method of claim 1, wherein the incubating occurs in a reaction vessel.

14. The in vitro screening method of claim 13, wherein the reaction vessel comprises a dish, a tube, or a multi-well plate.

15. The in vitro screening method of claim 1, wherein the method is a high-throughput method.

16. A therapeutic compound identified by the in vitro screening method of claim 1.

17. A plate comprising one or more wells, wherein at least one well comprises a candidate compound, a protein or protein complex comprising an ARID 1 A or ARID IB intrinsically disordered region (IDR) or fragment thereof, and a medium.

18. The plate of claim 17, wherein the at least one well further comprises a plurality of cells.

19. The plate of claim 17, wherein the medium comprises a cell culture medium.

20. The plate of claim 17, wherein the ARID1A or ARID1B intrinsically disordered region, or fragment thereof, is purified from mammalian cells.

21. The plate of claim 17, wherein the protein or protein complex comprising an ARID1 A or ARID IB IDR or fragment thereof is selected from the group consisting of a purified polypeptide, a chromatin remodeling complex, a DNA-binding protein, an epitope-tagged peptide, a core binding region (CBR) peptide, or any combination thereof.

22. The plate of claim 21, wherein the chromatin remodeling complex comprises an ATP- dependent chromatin remodeling complex.

23. The plate of claim 22, wherein the ATP-dependent chromatin remodeling complex comprises the mammalian SWI / SNF chromatin remodeling complex.

24. The plate of claim 17, wherein the candidate compound comprises a small molecule.

25. A kit comprising the plate of claim 17.

26. The kit of claim 25, wherein the plate is designed for automated drug screening.

Citation Information

Patent Citations

  • CHD5 encoding nucleic acids, polypeptides, antibodies and methods of use thereof

    US20100256071A1

  • Biomarkers predictive of Anti-immune checkpoint response

    US20190338370A1