Antisense RNA vector technology for the modulation of alternative polyadenylation

WO2026165273A2PCT designated stage Publication Date: 2026-08-06RGT UNIV OF CALIFORNIA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
RGT UNIV OF CALIFORNIA
Filing Date
2026-01-29
Publication Date
2026-08-06

Smart Images

  • Figure US2026013136_06082026_PF_FP_ABST
    Figure US2026013136_06082026_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are antisense RNAs comprising aptamer with high affinity to RNA binding proteins (RBPs), vectors comprising antisense RNAs RBPs, and methods of regulating poly A site (PAS) selection with the antisense RNAs and / or vectors.
Need to check novelty before this filing date? Find Prior Art

Description

Atty. Dkt. No.: 114198-3760ANTISENSE RNA VECTOR TECHNOLOGY FOR THE MODULATION OF ALTERNATIVE POLYADENYLATION CROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of and priority to U.S. Provisional Application No.63 / 751,756, filed January 30, 2025, which is incorporated herein by reference in its entirety.STATEMENT OF GOVERNMENT SUPPORT

[0002] This invention was made with government support under HG004659 awarded by the National Institutes of Health. The government has certain rights in the invention.BACKGROUND|0003] Alternative polyadenylation (APA) is a widespread RNA processing mechanism, occurring in over 60% of human genes (1). APA choices affects RNA stability, translation, and localization and contributes to protein diversification by selectively using poly(A) sites (PAS) and generating mRNA isoforms with different 3’ ends. APA exhibits tissue and cell-type-specificity, which is crucial for fundamental cellular processes such as growth, proliferation, and differentiation, and plays an important role in cancer progression.

[0004] This process is facilitated by the core cleavage and polyadenylation complex, which includes many RNA Binding Proteins (RBPs). Despite substantial evidence of RBPs’ critical role in APA, a systematic characterization of these proteins in the context of APA remains incomplete.

[0005] Published approaches to identify RBPs that regulate PAS selection have utilized shRNAs or CRISPR-Cas9 to knock out genes and measure changes in PAS selection through bulk and single-cell next-generation sequencing respectively. Alternatively, to assess the functionality of PAS site selection reporters have been utilized harboring nucleotide sequences encoding for different PAS sites and reporter activity can be correlated with strength of PAS selection. Additionally, mutating genomically PAS using CRISPR-Cas9 to assess the functional consequence of PAS selection.-1- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760

[0006] Antisense RNA technology is genetically encoded, exogenously introduced to cells, and guides an APA RBP to a PAS of interest. As a result, APA RBP localization activates or inhibits PAS selection. Thus, a need exists in the art for alternative and / or additional technology for PAS site selection. This disclosure satisfies this need and provides related advantages as well.SUMMARY OF THE DISCLOSURE

[0007] Described herein is a novel antisense RNA technology that enhances or inhibits PAS selection by recruiting RBPs to a specific location on the mRNA, typically but not limited to, at the 3’UTR region of a gene. Unlike approaches utilizing a reporter to determine PAS selection strength, this antisense RNA technology can be used to determine PAS selection on the transcriptome. Unlike approaches that mutate or destroy PAS sites using CRISPR-Cas9, this technology will not generate double stranded DNA breaks and directly targets mRNA. Unlike approaches where APA RBPs are knocked down or knocked out, this antisense RNA technology can directly measure if an RNA binding protein modulates a single PAS without affecting other PAS in the transcriptome.[0()08| Antisense RNA technology is genetically encoded and exogenously introduced to cells and guides an APA RBP to a PAS of interest. As a result, APA RBP localization activates or inhibits PAS selection.

[0009] Thus, the present disclosure provides a nuclear-expressed antisense RNA comprising, from the 5' end to the 3' end, a nucleotide sequence encoding an aptamer that has high affinity to an alternative polyadenylation (APA)-modulating RNA binding protein (RBP), and a nucleotide sequence antisense to a target region proximal to a polyA site (PAS). In certain embodiments, the antisense RNA further comprises at least one RBP or ribonucleoprotein (RNP)-recruiting motif positioned at the 5' end and / or the 3' end of the antisense portion to increase stability by recruiting RBPs or RNPs to block antisense RNA degradation. In other embodiments the target region proximal to the PAS is located between about 20 nucleotides and about 200 nucleotides upstream or downstream of the PAS cleavage site. In some embodiments the aptamer binds to the APA-modulating RBP via direct binding, and in other embodiments via a tethering interaction with an exogenous RNA-binding moiety fused to the RBP. In some embodiments the aptamer and RNA-binding moiety are derived-2- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760from a bacteriophage selected from MS2, R17, , PP7, QP or GA, or are derived from an iron responsive protein (IRP), bovine immunodeficiency virus, or human U1 small nuclear ribonucleoprotein A. In some examples the RNA hairpin structure encoded by the aptamer comprises or consists essentially of the sequence of SEQ ID NO: 1 or SEQ ID NO: 2, or an equivalent thereof, and the RNA-binding moiety comprises or consists essentially of the sequence of SEQ ID NO: 11, or an equivalent thereof. In some embodiments the aptamer is specific to the RBP to be recruited for PAS modulation. Vectors and host cells comprising the RNA molecules are provided as well.

[0010] Also provided are vectors comprising a DNA molecule encoding an antisense RNA as described above. In some embodiments the vector is selected from a plasmid, a viral vector, a cosmid, or a phage, and in certain embodiments the viral vector is selected from a baculovirus, a retrovirus, or an adenovirus. In some aspects the vector further comprises a polymerase II or polymerase III promoter such as a Ul, U6, or U7 promoter operably linked to the antisense RNA coding sequence. In some embodiments the vector further comprises a nucleotide sequence encoding an RBP operably linked to the antisense RNA, wherein the RBP is a modulator of PAS selection, for example CPSF5, CPSF6, RNPS1, or GRB2. In other embodiments the RBP can be any APA-modulating RBP identified in Table 4 or FIGS.2 A or 2E or a functional equivalent thereof.

[0011] Additional embodiments claim an isolated host cell comprising one or more of the antisense RNA, the DNA encoding said antisense RNA, or the vector described above. In some embodiments the host cell further comprises a polynucleotide encoding the target RBP or the RBP itself, optionally selected from CPSF5, CPSF6, RNPS1, GRB2, or another APA-modulating RBP. The host cell may be eukaryotic or prokaryotic, and in some embodiments is a mammalian cell such as a HEK293T cell.[00121 Some aspects are directed to a composition comprising a carrier, such as a pharmaceutically acceptable carrier, and one or more of the antisense RNA, DNA encoding said antisense RNA, vector, and / or isolated host cell as described herein. In certain embodiments the carrier is selected from phosphate-buffered saline, a stabilizer, a diluent, or another suitable excipient.-3- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[00131 Further embodiments claim a method of regulating PAS selection on mRNA in a cell, comprising contacting the cell with the antisense RNA or a vector or composition comprising the antisense RNA, and further contacting the cell with an RBP or a polynucleotide encoding an RBP, wherein the antisense RNA hybridizes to the target region proximal to the PAS and wherein the RBP activates or inhibits PAS selection of the mRNA in the cell. In other embodiments the method comprises contacting the cell with a vector comprising both the antisense RNA and a nucleotide encoding the RBP, wherein the antisense RNA binds to the target region proximal to the PAS and the recruited RBP modulates PAS selection. In some embodiments the RBP is recruited via an aptamer that binds an endogenous RNA-binding moiety naturally present in the cell. In some examples the method targets mRNAs from genes implicated in disease, including haploinsufficiency disorders such as Dravet syndrome, autosomal dominant retinitis pigmentosa, or neurofibromatosis type 1, with the aim of shifting PAS usage toward a therapeutically beneficial isoform. In particular embodiments, the PAS targeting and RBP recruitment is positional, with maximum activation observed when the aptamer is located about 60 nucleotides upstream of the AAUAAA motif.

[0014] In certain dependent embodiments of the method, the antisense RNA and RBP recruiting elements are delivered to mammalian cells, animal models, or human patients via transfection reagents, or packaged in a viral vector such as lentivirus or adeno-associated virus, using a promoter that directs nuclear expression of the RNA transcript. In other dependent embodiments, the method uses a programmable U7smOPT-MS2 recruitment system to direct RBPs to a desired location near the PAS in either reporter constructs or endogenous gene targets.

[0015] Some aspects further claim a method of predicting an RBP’s function in PAS selection by obtaining the amino acid sequence of the RBP, inputting the sequence into a fine-tuned protein language model trained on APA-modulating RBPs, classifying the RBP as an activator or non-activator based on the model’s score, and optionally generating occlusion maps indicating sequence regions important for classification. In further dependent embodiments, the model is ProteinBERT fine-tuned on tethered-function screen data and filtered to exclude sequences with more than 80% identity to training sequences. Additional dependent embodiments claim a method of mapping domain-level contributions to PAS activation by generating deletion variants or isolated domain constructs of the RBP fused to a -4- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760coat protein, testing each variant in a PAS activation assay, and correlating activity changes with the presence or absence of specific domains.10016] Other embodiments claim a method of predicting protein regulators of PAS selection by obtaining amino acid sequences of a plurality of RBPs, labeling each as activator or nonactivator based on results from a tethered-function assay, inputting the sequences and labels into a sequence-based machine learning model, training the model to classify proteins, and outputting predictions for RBPs of unknown status. In dependent embodiments, the model further highlights low complexity regions or protein domains that contribute to PAS activation or inhibition, which can be exploited to design minimal RBP effector modules for antisense RNA constructs.

[0017] Through these various claims, the invention covers antisense RNA molecules tailored for PAS-proximal recruitment of RBPs, vectors and host cells for delivering such molecules, methods for regulating PAS usage in vitro and in vivo, compositions for therapeutic use, and computational tools for identifying and optimizing RBPs and their domains for use in PAS modulation systems.BRIEF DESCRIPTION OF THE DRAWINGS

[0018] FIGS. 1A- IE: Development of dual-luciferase MS2 reporter tethering assay and overview of the large-scale tethered screen (FIG. 1A) Schematic of the upstream (top) and downstream (bottom) dual-luciferase reporters. (FIG. IB) mRNA isoform schematics depicting reporter outcomes. L3 PAS inhibition by a tethered RBP leads to both Renilla and Firefly luciferase expression via transcription readthrough (top), whereas activation results in primarily Renilla expression (bottom). (FIG. 1C) Luminescence ratios for control RBPs using the upstream and downstream reporters (n=3), normalized to FLAG-MCP. (FIG. ID) RT-PCR validation showing PAS usage following binding of control RBPs to the upstream and downstream reporters (n=3). (FIG. IE) Overview of the large-scale tethering assay and criteria for selecting RBP candidates. Significance is indicated by **** P < 0.0001, *** P < 0.001, ** P < 0.01, * P < 0.05 using a two-sided t-test.FIGS. 2A - 21: Large-scale tethered screen identifies high-confidence RBPs that modulate PAS selection. (FIG. 2A) Volcano plots showing luminescence ratios normalized to FLAG-MCP and associated -log10(P values) for RBPs (n=3) following tethering to-5- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760upstream (left) and downstream (right) reporters. Top significant (P<0.05, two-sided / -test) activators and inhibitors of PAS selection are highlighted. (FIG. 2B) Overlap of RBPs activating (top) or inhibiting (bottom) PAS selection identified using upstream and downstream reporters. (FIG. 2C) Functional enrichment analysis stratified by candidate groups. (FIG. 2D) Domain enrichment analysis (Fisher’s exact test) stratified by candidate groups. (FIG. 2E) Luminescence ratios normalized to FLAG-MCP for high-confidence activators identified after secondary screening (n=3). Known (red) and novel (black) APA factors are indicated. (FIG. 2F) Comparison of RT-qPCR (n=3) and luminescence readouts for 20 tethered RBPs (FIG. 2G) Schematic of the MS2 U7 smOPT snRNA-based recruitment system and its target region on the modified reporter without MS2 stem-loops. (FIG. 2H) Luminescence ratios normalized to non-targeting control for FLAG, CPSF5, and CPSF6 fused to MCP co-expressed with guides designed 0 nt, 20 nt, 40 nt, 60 nt, and 120 nt upstream of the AAUAAA hexamer (n=3, two-sided t-test, ** P < 0.001,* P < 0.05) (FIG. 21) Luminescence ratios normalized to non-targeting control for FLAG, known CPA factors (CPSF5 and CPSF6), and novel activators (GRB2 and RNPS1) fused to MCP co-expressed with the guide targeting 60bp upstream of the PAS (n=3, two-sided t-test, ** P < 0.001,* P < 0.05).

[0020] FIGS. 3A - 3G: Training and interpreting a protein language model to predict activators of PAS selection. (FIG. 3A) Schematic of data pre-processing, model benchmarking, training, and output. (FIG. 3B) ROC-AUC and PR-AUC curves for the SVM, CNN, ProteinBERT, and sequence-driven HydRA models. (FIG. 3C) Overlap of significant (P < 0.05) occlusion peaks with domains and LCRs. (FIG. 3D) Top informative domains contributing to activator classification. (FIG. 3E) Occlusion maps for CPSF6 and SCAF8. (FIG. 3F) Fine-tuned ProteinBERT prediction scores and classifications for the ZFP set (Activator: Score >0.15, Non- Activator: Score <= 0.15). (FIG. 3G) Occlusion map of BRCA1.

[0021] FIGS. 4A- 4G: Confirmation and characterization of high-confidence activators using knockdown PAS-seq, RNA-seq, and eCLIP analyses. (FIG. 4A) Number of 3 ' end shortening (light blue) and lengthening (dark blue) events identified from PAS-seq analyses following the knockdown of each high-confidence activator in HEK293T cells. (FIG. 4B) Top-enriched functional categories for genes undergoing APA upon each activator-6- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760knockdown. (FIG. 4C) Clustering of activator candidates based on the Jaccard Index (JI) calculated based on APA and AS events, number of CPA interactors, and number of splicing interactors. (FIG. 4D) Top ten most enriched functional categories of each activator’s interactome, grouped into larger functional groups. (FIG. 4E) Enrichment (Fisher’s exact test, P < 0.05) of eCLIP signal surrounding KD-modulated PASs relative to unaffected PASs.(FIG. 4F) A scaled binding density profiles (±200 nt) comparing constitutive and knockdown-activated PAS, stratified by direction of APA shift (knockdown PAS-seq). (FIG.4G) Scaled eCLIP binding density ± 200 nt surrounding canonical PASs located in the Last Exon (left) and Intron (right).[00221 FIGS. 5A - 5E: Analysis of the GRB2 and RNPS1 interactome using AP-MS and co-IP. (FIG. 5A) The number of annotated (dark green) and unannotated (light green) GRB2 and RNPS1 interactors as identified by AP-MS. (FIG. 5B) Scatterplot of -log10( value) versus log2(Fold Change) for all interactors identified for GRB2 (left) and RNPS1 (right). Significant hits are highlighted in dark pink. (FIG. 5C) Western blots of co-IPs with anti-GRB2 or anti-RNPSl in HEK293T cells (RNase: RNase A / Tl). (FIG. 5D) Western blots of co-IPs with anti-CPSF7 in HEK293T cells. (FIG. 5E) Western blots of co-IPs with anti-FIPl in HEK293T cells overexpressing FLAG-RNPS1.

[0023] FIGS. 6A - 6L: Domain mapping and biochemical interaction assays of GRB2 and RNPS1 in PAS selection (FIG. 6A) Western blot of proteins co-immunoprecipitated with FLAG-GRB2 domain-deletion constructs in HEK293T cells. Bottom panel: Occlusion map of GRB2 from the fine-tuned ProteinBERT model. Significant windows are highlighted (dark blue: P < 0.001, light blue: P < 0.05). (FIG. 6B) Tethering assay results using GRB2 domain-deletion constructs with the upstream reporter. Statistical significance: * P < 0.05, ** P < 0.01, *** P <0.001, **** P <0.0001. (one-way ANOVA, Tukey’s multiple comparisons test). (FIG. 6C) SDS-PAGE of purified CPA subcomplexes and associated factors. (FIG. 6D) Western blot of GST pull-down assays using purified CPA subcomplexes and GST or GST-GRB2. Additional data in FIG. 12A. (FIG. 6E) Overlap of APA events between GRB2 knockdown and CFIm knockdowns. (FIG. 6F) Western blot of proteins co-immunoprecipitated with FLAG-RNPS1 domain constructs in HEK293T cells. Bottom panel: Occlusion map of RNPS1 from the fine-tuned ProteinBERT model. Significant windows are highlighted (dark: P < 0.001, light: P < 0.05). (FIG. 6G) Tethering assay results using-7- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760RNPS1 constructs with the upstream reporter. Statistical significance: * P <0.05, ** P <0.01, *** P <0.001, **** P <0.0001. (one-way ANOVA, Tukey’s multiple comparisons test). (FIG. 6H) Western blot of GST pull-down assays using purified CPA subcomplexes and GST or GST-RNPSlNterm. Additional data in FIG. 12D. (FIG. 61) APA analysis of RNPS1-overexpressing HEK293T cells using PAS-seq, displaying only genes with significant APA changes. (FIG. 6J) A upstream eCLIP density (proximal PAS - distal PAS; -500 to 0 nt) for genes proximal-shifted by RNPS1 overexpression versus unaffected controls (Wilcoxon rank-sum, P= 9.9 x 10l2). (FIG. 6K) Spearman R correlation between mean RNPS1 expression and the number of shortening events across seven cancer types. (FIG. 6L) Schematic of the human core CPA machinery, highlighting direct interactions of GRB2 and RNPS1 with CPA subunits.

[0024] FIGS. 7A - 7B: (FIG. 7A) Renilla to Firefly Luciferase luminescence ratios for tethered and untethered RBP normalized to matched FLAG control. (FIG. 7B) Primer design for reporter validation with RT-PCR.[00251 FIGS. 8A - 8G: (FIG. 8A) Comparison of the number of known cleavage and polyadenylation (CPA) factors interacting with activator and inhibitor RBP candidates (Wilcoxon rank-sum). (FIG. 8B) Overlap of activators identified in primary and secondary screening, with rank density (lower rank=strong activator, higher rank=weaker activator) of overlapping and non-overlapping primary screen candidates (FIG. 8C) Ranking (lower rank=strong activator, higher rank = weaker activator) correlation (Spearman’s R) among recovered activators between primary and secondary screening. (FIG. 8D) Positiondependent comparison of activator candidates between primary and secondary screening to determine high-confidence activators. (FIG. 8E) High confidence activators compared with candidates identified in a tethered function assay surveying RBP effects on translation and stability15. (FIG. 8F) Tethering assay in EXOSC3-FKBPF36V-FLAG-HEK293T cells treated with or without 500nM dTAG-vl for 9 h (n=3, multiple t-test, ns: not significant). (FIG. 8G) Luminescence Ratios normalized to non-targeting control for FLAG, known CPA factors (CPSF5 and CPSF6), and novel activators (GRB2 and RNPS1) fused to MCP co-expressed with the guide targeting 60 nt upstream of the PAS (n=3, two-sided t-test, ** P < 0.001,* P < 0.05).-8- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[0026[ FIGS. 9A - 9E: (FIG. 9A) Distribution of percent similarity among protein pairs of isoforms or variants and different proteins. All proteins below 50% similarity are not shown.(FIG. 9B) ROC-AUC (left) and PR-AUC (right) for the fine-tuned ProteinBERT evaluation on the heldout set. The red point represents the Fl -optimized threshold. (C-D) Upstream (FIG. 9C) and downstream (FIG. 9D) validations for ZFP set. ( / -test, *** P < 0.005, ** P < 0.01 * P < 0.05) (FIG. 9E) Confusion matrix reporting performance of the fine-tuned ProteinBERT model in classifying the ZFPs using an Fl -optimized threshold.

[0027] FIGS. 10A - 10D: (FIG. 10A) Number of genes with changes in alternative splicing following knockdown of candidate RBPs categorized by splicing event type. (FIG. 10B) Direction-stratified overlap of genes with RBP eCLIP binding and RBP targets identified through knockdown PAS-seq. (FIG. 10C) Binding location of RBP eCLIP windows that overlap RBP targets identified through knockdown PAS-seq. (FIG. 10D) Top 5 enriched motifs in binding regions and associated P values based on binding regions found in genes having shifts in PAS usage after candidate knockdown.[0028| FIGS. 11A - 11C: (FIG. 11 A) GRB2 interactome detected by AP-MS with annotated functions. (FIG. 11B) RNPS1 interactome detected by AP-MS with annotated functions. (FIG. 11C) Western blots of co-IPs with anti-FIPl in HEK293T cells.

[0029] FIGS. 12A - 12G: (FIG. 12A) Western blot of GST pull-down assays using purified CPA subcomplexes and GST or GST-GRB2. (FIG. 12B) Overlap of APA events between GRB2 knockdown and CFIm subunit knockdowns. (FIG. 12C) Directionality analysis of overlapping of APA events between GRB2 knockdown and CFIm. (FIG. 12D) Western blot of GST pull-down assays using purified CPA subcomplexes and GST or GST-RNPSlNterm.(FIG. 12E) RNPS1 eCLIP density near proximal or distal PAS of genes that undergoes shortening upon RNPS1 overexpression (left) versus unaffected (right). (FIG. 12F) Topenriched functional categories for genes undergoing APA upon RNPS1 overexpression. (FIG. 12G) APA analysis of EIF4A3 (left) or RBM8A (right) overexpressing HEK293T cells using PAS-seq, displaying only genes with significant APA changes.

[0030] FIG. 13: Antisense RNA vector features.

[0031] FIGS. 14A-14C: illustrate flowcharts of various methods described herein, in accordance with some embodiments. (FIG. 14A) is a flow chart of an implementation of a -9- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760method for predicting RBP function in PAS selection, according to some implementations.(FIG. 14B) is a flow chart of an implementation of a method for mapping domain-level contributions to PAS activation in a candidate RBP, according to some implementations. (FIG. 14C) is a flow chart of an implementation of a method 1416 for predicting protein regulators of polyadenylation site (PAS) selection.[0032| FIGS. 15A and 15B: are block diagrams depicting embodiments of computing devices that can be used in connection with the methods and systems described herein. (FIG.15A) illustrates a computing device that may include a storage device, an installation device, a network interface, an I / O controller, display devices, a keyboard and a pointing device, such as a mouse. The storage device may include, without limitation, an operating system and / or software. (FIG. 15B) illustrates a computing device that may also include additional optional elements, such as a memory port, a bridge, one or more input / output devices, and a cache memory in communication with the central processing unit.DETAILED DESCRIPTION

[0033] Throughout this application various technical and patent publications are referenced, the disclosure of which are incorporated herein to more fully describe the state of the art to which this disclosure pertains. Technical reference may be identified by an Arabic numeral wherein the complete bibliographic details of the publication are provided in the Reference section preceding the claims. The disclosures of these technical publications also are referenced herein to more fully describe the state of the art.Definitions

[0034] As used in the specification and claims, the singular form “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a cell” includes a plurality of cells, including mixtures thereof.

[0035] As used herein, the term “comprising” is intended to mean that the compositions or methods include the recited steps or elements, but do not exclude others. “Consisting essentially of’ shall mean rendering the claims open only for the inclusion of steps or elements, which do not materially affect the basic and novel characteristics of the claimed compositions and methods. “Consisting of’ shall mean excluding any element or step not-10- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760specified in the claim. Embodiments defined by each of these transition terms are within the scope of this disclosure. For example, a composition consisting essentially of the elements as defined herein would not exclude trace contaminants from the isolation and purification method and pharmaceutically acceptable carriers, such as phosphate buffered saline, preservatives and the like. “Consisting of’ shall mean excluding more than trace elements of other ingredients and substantial method steps for administering the compositions disclosed herein. Aspects defined by each of these transition terms are within the scope of the present disclosure.

[0036] As used herein, the term “about” is used to indicate that a value includes the standard deviation of error for the device or method being employed to determine the value. The term “about” when used before a numerical designation, e.g., temperature, time, amount, and concentration, including range, indicates approximations which can vary by (+) or (-) 15%, 10%, 5%, 3%, 2%, or 1 %.

[0037] As used herein, the term “animal” refers to living multi-cellular vertebrate organisms, a category that includes, for example, mammals and birds. The term “mammal” includes both human and non-human mammals.

[0038] The term “subject,” “host,” “individual,” and “patient” are as used interchangeably herein to refer to animals, typically mammalian animals. Any suitable mammal can be treated by a method, cell or composition described herein. Non-limiting examples of mammals include humans, non-human primates (e.g., apes, gibbons, chimpanzees, orangutans, monkeys, macaques, and the like), domestic animals (e.g., dogs and cats), farm animals (e.g., horses, cows, goats, sheep, pigs) and experimental animals (e.g., mouse, rat, rabbit, guinea pig). In some embodiments a mammal is a human. A mammal can be any age or at any stage of development (e.g., an adult, teen, child, infant, or a mammal in utero). A mammal can be male or female. A mammal can be a pregnant female. In some embodiments a subject is a human.

[0039] As used herein, “host cell” or “isolated host cell” refers to a prokaryotic or eukaryotic cell modified in vitro or ex vivo to contain introduced nucleic acids or proteins of interest, substantially separated from unmodified cells of its native population. Non-limiting examples include, for example: HEK293T cells co-transfected with pPASPORT reporter and-11- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760MCP-RBP vector; HEK293T cells transfected with U7smOPT antisense RNA guides targeting PAS; and EXOSC3-FKBPAF36V-FLAG-HEK293T degron cells used to test exosome effects on reporter activity.

[0040] “Eukaryotic cells” comprise all of the life kingdoms except monera. They can be easily distinguished through a membrane-bound nucleus. Animals, plants, fungi, and protists are eukaryotes or organisms whose cells are organized into complex structures by internal membranes and a cytoskeleton. The most characteristic membrane-bound structure is the nucleus. Unless specifically recited, the term “host” includes a eukaryotic host, including, for example, yeast, higher plant, insect and mammalian cells. Non-limiting examples of eukaryotic cells or hosts include simian, bovine, porcine, murine, rat, avian, reptilian and human.

[0041] “Prokaryotic cells” usually lack a nucleus or any other membrane-bound organelles and are divided into two domains, bacteria and archaea. In addition to chromosomal DNA, these cells can also contain genetic information in a circular loop called on episome.Bacterial cells are very small, roughly the size of an animal mitochondrion (about 1-2 pm in diameter and 10 pm long). Prokaryotic cells feature three major shapes: rod shaped, spherical, and spiral. Instead of going through elaborate replication processes like eukaryotes, bacterial cells divide by binary fission. Examples include but are not limited to Bacillus bacteria, E. coli bacterium, and Salmonella bacterium.[00421 A “composition” typically intends a combination of the active agent, and a naturally-occurring or non-naturally-occurring carrier, inert (for example, a detectable agent or label) or active, such as an adjuvant, diluent, binder, stabilizer, buffers, salts, lipophilic solvents, preservative, adjuvant or the like and include pharmaceutically acceptable carriers. Carriers also include pharmaceutical excipients and additives proteins, peptides, amino acids, lipids, and carbohydrates (e.g., sugars, including monosaccharides, di-, tri, tetra-oligosaccharides, and oligosaccharides; derivatized sugars such as alditols, aldonic acids, esterified sugars and the like; and polysaccharides or sugar polymers), which can be present singly or in combination, comprising alone or in combination 1-99.99% by weight or volume.Exemplary protein excipients include serum albumin such as human serum albumin (HSA), recombinant human albumin (rHA), gelatin, casein, and the like. Representative amino acid-12- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760components, which can also function in a buffering capacity, include alanine, arginine, glycine, arginine, betaine, histidine, glutamic acid, aspartic acid, cysteine, lysine, leucine, isoleucine, valine, methionine, phenylalanine, aspartame, and the like. Carbohydrate excipients are also intended within the scope of this technology, examples of which include but are not limited to monosaccharides such as fructose, maltose, galactose, glucose, D-mannose, sorbose, and the like; disaccharides, such as lactose, sucrose, trehalose, cellobiose, and the like; polysaccharides, such as raffinose, melezitose, maltodextrins, dextrans, starches, and the like; and alditols, such as mannitol, xylitol, maltitol, lactitol, xylitol sorbitol (glucitol) and myoinositol.

[0043] The compositions used in accordance with the disclosure, including cells, treatments, therapies, agents, drugs and pharmaceutical formulations can be packaged in dosage unit form for ease of administration and uniformity of dosage. The term "unit dose" or "dosage" refers to physically discrete units suitable for use in a subject, each unit containing a predetermined quantity of the composition calculated to produce the desired responses in association with its administration, i.e., the appropriate route and regimen. The quantity to be administered, both according to number of treatments and unit dose, depends on the result and / or protection desired. Precise amounts of the composition also depend on the judgment of the practitioner and are peculiar to each individual. Factors affecting dose include physical and clinical state of the subject, route of administration, intended goal of treatment (alleviation of symptoms versus cure), and potency, stability, and toxicity of the particular composition. Upon formulation, solutions will be administered in a manner compatible with the dosage formulation and in such amount as is therapeutically or prophylactically effective. The formulations are easily administered in a variety of dosage forms, such as the type of injectable solutions described herein.

[0044] As used herein, the terms “nucleic acid sequence,” “oligonucleotide,” and “polynucleotide” are used interchangeably to refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. Thus, this term includes, but is not limited to, single-, double-, or multi-stranded DNA or RNA, genomic DNA, cDNA, circRNA, DNA-RNA hybrids, or a polymer comprising purine and pyrimidine bases or other natural, synthetic, recombinantly produced, chemically or biochemically modified, nonnatural, or derivatized nucleotide bases. A polynucleotide can comprise modified nucleotides,-13- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure can be imparted before or after assembly of the polynucleotide. The sequence of nucleotides can be interrupted by non-nucleotide components. A polynucleotide can be further modified after polymerization, such as by conjugation with a labeling component. The term also refers to both double- and single-stranded molecules. Unless otherwise specified or required, any aspect of this technology that is a polynucleotide encompasses both the double-stranded form and each of two complementary single- stranded forms known or predicted to make up the double-stranded form.

[0045] The term “encode” as it is applied to nucleic acid sequences refers to a polynucleotide which is the to “encode” a polypeptide if, in its native state or when manipulated by methods well known to those skilled in the art, can be transcribed and / or translated to produce the mRNA for the polypeptide and / or a fragment thereof. The antisense strand is the complement of such a nucleic acid, and the encoding sequence can be deduced therefrom.

[0046] As used herein, the term “isolated cell” generally refers to a cell that is substantially separated from other cells of a tissue. The term includes prokaryotic and eukaryotic cells.[0047| As used herein, the term “vector” refers to a nucleic acid construct deigned for transfer between different hosts, including but not limited to a plasmid, a virus, a cosmid, a phage, a BAC, a YAC, etc. It can also be referred to as an “expression cassette.” In one aspect, the vectors are designed to carry coding sequences, promoters, or regulatory elements into a target cell to express antisense RNAs and / or RBPs. Non-limiting examples include, for example:pPASPORT bicistronic luciferase reporters containing MS2 aptamers for PAS tethering (FIGS. 1A & 2G); pcDNA3.1 vectors expressing FLAG-tagged GRB2 or RNPS1; and lentiviral shRNA vectors used for knockdown of high-confidence RBPs.

[0048] A “viral vector” is defined as a recombinantly produced virus or viral particle that comprises a polynucleotide to be delivered into a host cell, either in vivo, ex vivo or in vitro. In some embodiments, plasmid vectors can be prepared from commercially available vectors. In other embodiments, viral vectors can be produced from baculoviruses, retroviruses, adenoviruses, AAVs, etc. according to techniques known in the art. In one embodiment, the viral vector is a lentiviral vector. Examples of viral vectors include retroviral vectors,-14- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760adenovirus vectors, adeno-associated virus vectors, alphavirus vectors and the like. Infectious tobacco mosaic virus (TMV)-based vectors can be used to manufacture proteins and have been reported to express Griffithsin in tobacco leaves (O'Keefe et al. (2009) Proc. Nat. Acad. Sci. USA 106(15):6099-6104). Alphavirus vectors, such as Semliki Forest virus-based vectors and Sindbis virus-based vectors, have also been developed for use in gene therapy and immunotherapy. See, Schlesinger & Dubensky (1999) Curr. Opin. Biotechnol. 5:434-439 and Ying et al. (1999) Nat. Med. 5(7):823-827. Further details as to modem methods of vectors for use in gene transfer can be found in, for example, Kotterman et al. (2015) Viral Vectors for Gene Therapy: Translational and Clinical Outlook Annual Review of Biomedical Engineering 17. Vectors that contain both a promoter and a cloning site into which a polynucleotide can be operatively linked are known in the art. Such vectors are capable of transcribing RNA in vitro or in vivo and are commercially available from sources such as Agilent Technologies (Santa Clara, Calif.) and Promega Biotech (Madison, Wis.).

[0049] As used herein, the term “detectable marker” or “detectable label” refers to at least one marker capable of directly or indirectly, producing a detectable signal. A non-exhaustive list of markers includes enzymes which produce a detectable signal, for example by colorimetry, fluorescence, luminescence, such as horseradish peroxidase, alkaline phosphatase, P-galactosidase, glucose-6-phosphate dehydrogenase, chromophores such as fluorescent, luminescent dyes, groups with electron density detected by electron microscopy or by their electrical property such as conductivity, amperometry, voltammetry, impedance, detectable groups, for example whose molecules are of sufficient size to induce detectable modifications in their physical and / or chemical properties, such detection can be accomplished by optical methods such as diffraction, surface plasmon resonance, surface variation , the contact angle change or physical methods such as atomic force spectroscopy, tunnel effect, or radioactive molecules such as32P,35S or125I.

[0050] As used herein, the term “purification marker” or “reporter protein” refer to at least one marker useful for purification or identification. A non-exhaustive list of purification markers includes His, lacZ, GST, maltose-binding protein, NusA, BCCP, c-myc, CaM, FLAG, GFP, YFP, cherry, thioredoxin, poly(NANP), V5, Snap, HA, chitin-binding protein, Softag 1, Softag 3, Strep, or S-protein. Suitable direct or indirect fluorescence marker comprise FLAG, GFP, YFP, RFP, dTomato, cherry, Cy3, Cy 5, Cy 5.5, Cy 7, DNP, AMCA,-15- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760Biotin, Digoxigenin, Tamra, Texas Red, rhodamine, Alexa fluors, FITC, TRITC or any other fluorescent dye or hapten.10051] As used herein, the term “expression” refers to the process by which polynucleotides are transcribed into mRNA and / or the process by which the transcribed mRNA is subsequently being translated into peptides, polypeptides, or proteins. If the polynucleotide is derived from genomic DNA, expression can include splicing of the mRNA in a eukaryotic cell. The expression level of a gene can be determined by measuring the amount of mRNA or protein in a cell or tissue sample. In one aspect, the expression level of a gene from one sample can be directly compared to the expression level of that gene from a control or reference sample. In another aspect, the expression level of a gene from one sample can be directly compared to the expression level of that gene from the same sample following administration of a compound.

[0052] As used herein, “homology” or “identical”, percent “identity” or “similarity”, when used in the context of two or more nucleic acids or polypeptide sequences, refers to two or more sequences or subsequences that are the same or have a specified percentage of nucleotides or amino acid residues that are the same, e.g., at least 60% identity, preferably at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher identity over a specified region (e.g., nucleotide sequence encoding the RBP described herein). Homology can be determined by comparing a position in each sequence which can be aligned for purposes of comparison. When a position in the compared sequence is occupied by the same base or amino acid, then the molecules are homologous at that position. A degree of homology between sequences is a function of the number of matching or homologous positions shared by the sequences. The alignment and the percent homology or sequence identity can be determined using software programs known in the art, for example those described in Current Protocols in Molecular Biology (Ausubel et al., eds. 1987) Supplement 30, section 7.7.18, Table 7.7.1. Preferably, default parameters are used for alignment. A preferred alignment program is BLAST, using default parameters. In particular, preferred programs are BLASTN and BLASTP, using the following default parameters: Genetic code = standard; filter = none; strand = both; cutoff = 60; expect = 10; Matrix = BLOSUM62; Descriptions = 50 sequences; sort by = HIGH SCORE; Databases = non-redundant, GenBank + EMBL + DDB J + PDB + GenBank CDS translations +-16- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760SwissProtein + SPupdate + PIR. Details of these programs can be found at the following Internet address: ncbi.nlm.nih.gov / cgi-bin / BLAST. The terms “homology” or “identical,” percent “identity” or “similarity” also refer to, or can be applied to, the complement of a test sequence. The terms also include sequences that have deletions and / or additions, as well as those that have substitutions. As described herein, the preferred algorithms can account for gaps and the like. Preferably, identity exists over a region that is at least about 25 amino acids or nucleotides in length, or more preferably over a region that is at least 50-100 amino acids or nucleotides in length. An “unrelated” or “non-homologous” sequence shares less than 40% identity, or alternatively less than 25% identity, with one of the sequences disclosed herein.

[0053] It is to be inferred without explicit recitation and unless otherwise intended, that when the present disclosure relates to a polypeptide, protein, polynucleotide, an equivalent or a biologically equivalent of such is intended within the scope of this disclosure. As used herein, the term “biological equivalent thereof’ is intended to be synonymous with “equivalent thereof’ when referring to a reference protein, polypeptide, or nucleic acid, intends those having minimal homology while still maintaining desired structure or functionality. Unless specifically recited herein, it is contemplated that any of the above also includes equivalents thereof. For example, an equivalent intends at least about 70% homology or identity, or at least 80% homology or identity and alternatively, or at least about 85%, or alternatively at least about 90%, or alternatively at least about 95%, or alternatively at least 98% percent homology or identity and / or exhibits substantially equivalent biological activity to the reference protein, polypeptide, or nucleic acid. Alternatively, an equivalent intends at least 70% homology or identity, or at least 80% homology or identity and alternatively, or at least 85%, or alternatively at least 90%, or alternatively at least 95%, or alternatively at least 98% percent homology or identity and / or exhibits substantially equivalent biological activity to the reference protein, polypeptide, or nucleic acid. Yet further, when referring to polynucleotides, an equivalent thereof is a polynucleotide that hybridizes under stringent conditions to the reference polynucleotide or its complement.[0054 { The phrase “equivalent polypeptide” or “equivalent peptide fragment” refers to protein, polynucleotide, or peptide fragment encoded by a polynucleotide that hybridizes to a polynucleotide encoding the exemplified polypeptide or its complement of the polynucleotide -17- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760encoding the exemplified polypeptide, under high stringency and / or which exhibit similar biological activity in vivo, e.g., approximately 100%, or alternatively, over 90% or alternatively over 85% or alternatively over 70%, as compared to the standard or control biological activity. Additional embodiments within the scope of this disclosure are identified by having more than 60%, or alternatively, more than 65%, or alternatively, more than 70%, or alternatively, more than 75%, or alternatively, more than 80%, or alternatively, more than 85%, or alternatively, more than 90%, or alternatively, more than 95%, or alternatively more than 97%, or alternatively, more than 98% or 99% sequence homology. Percentage homology can be determined by sequence comparison using programs such as BLAST run under appropriate conditions. In one aspect, the program is run under default parameters.

[0055] The term “isolated” as used herein refers to molecules or biologicals or cellular materials being substantially free from other materials. In one aspect, the term “isolated” refers to nucleic acid, such as DNA or RNA, or protein or polypeptide, or cell or cellular organelle, or tissue or organ, separated from other DNAs or RNAs, or proteins or polypeptides, or cells or cellular organelles, or tissues or organs, respectively, that are present in the natural source. The term “isolated” also refers to a nucleic acid or peptide that is substantially free of cellular material, viral material, or culture medium when produced by recombinant DNA techniques, or chemical precursors or other chemicals when chemically synthesized. Moreover, an “isolated nucleic acid” is meant to include nucleic acid fragments which are not naturally occurring as fragments and would not be found in the natural state. The term “isolated” is also used herein to refer to polypeptides which are isolated from other cellular proteins and is meant to encompass both purified and recombinant polypeptides. The term “isolated” is also used herein to refer to cells or tissues that are isolated from other cells or tissues and is meant to encompass both cultured and engineered cells or tissues.

[0056] The term “protein”, “peptide” and “polypeptide” are used interchangeably and in their broadest sense to refer to a compound of two or more subunit amino acids, amino acid analogs or peptidomimetics. The subunits can be linked by peptide bonds. In another aspect, the subunit can be linked by other bonds, e.g., ester, ether, etc. A protein or peptide must contain at least two amino acids and no limitation is placed on the maximum number of amino acids which can comprise a protein’s or peptide’s sequence. As used herein the term-18- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760“amino acid” refers to either natural and / or unnatural or synthetic amino acids, including glycine and both the D and L optical isomers, amino acid analogs and peptidomimetics.10057] As used herein, the term “purified” does not require absolute purity; rather, it is intended as a relative term. Thus, for example, a purified nucleic acid, peptide, protein, biological complexes or other active compound is one that is isolated in whole or in part from proteins or other contaminants. Generally, substantially purified peptides, proteins, biological complexes, or other active compounds for use within the disclosure comprise more than 80% of all macromolecular species present in a preparation prior to admixture or formulation of the peptide, protein, biological complex or other active compound with a pharmaceutical carrier, excipient, buffer, absorption enhancing agent, stabilizer, preservative, adjuvant or other co-ingredient in a complete pharmaceutical formulation for therapeutic administration. More typically, the peptide, protein, biological complex or other active compound is purified to represent greater than 90%, often greater than 95% of all macromolecular species present in a purified preparation prior to admixture with other formulation ingredients. In other cases, the purified preparation can be essentially homogeneous, wherein other macromolecular species are not detectable by conventional techniques.

[0058] As used herein, the term “recombinant protein” refers to a polypeptide which is produced by recombinant DNA techniques, wherein generally, DNA encoding the polypeptide is inserted into a suitable expression vector which is in turn used to transform a host cell to produce the heterologous protein. In one aspect, the term includes proteins that are chemically manufactured without the use of a host cell system.

[0059] The terms “fusion” or “chimeric” and grammatical variations thereof, when used in reference to a molecule, means that a portions or part of the molecule contains a different entity distinct (heterologous) from the molecule as they do not typically exist together in nature. That is, for example, one portion of the fusion or chimera includes or consists of a portion that does not exist together in nature and is structurally distinct.

[0060] As used herein, the term “enhancer”, denotes sequence elements that augment, improve or ameliorate transcription of a nucleic acid sequence irrespective of its location and orientation in relation to the nucleic acid sequence to be expressed. An enhancer can enhance-19- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760transcription from a single promoter or simultaneously from more than one promoter. As long as this functionality of improving transcription is retained or substantially retained (e.g., at least 70%, at least 80%, at least 90% or at least 95% of wild-type activity, that is, activity of a full-length sequence), any truncated, mutated or otherwise modified variants of a wild-type enhancer sequence are also within the above definition.[0061 [ The term “promoter” as used herein refers to any sequence that regulates the expression of a coding sequence, such as a gene. Promoters can be constitutive, inducible, repressible, or tissue-specific, for example. A “promoter” is a control sequence that is a region of a polynucleotide sequence at which initiation and rate of transcription are controlled. It can contain genetic elements at which regulatory proteins and molecules can bind such as RNA polymerase and other transcription factors. Pol II or pol III promoters are exemplary only. In one aspect, the promoter refers to the DNA sequence in the expression vector that directs cellular transcription machinery to synthesize the antisense RNA. This promoter, whether a polymerase II promoter such as CMV or EFla, or a polymerase III promoter such as Ul, U6, or U7, is distinct from the RBP / RNP recruiting region. The RBP / RNP recruiting region is an RNA-level element, such as an aptamer hairpin or other motif, embedded within the transcribed antisense RNA sequence. It functions after transcription to bind an RNA-binding protein or ribonucleoprotein. In one embodiment, this recruiting motif is not part of the DNA promoter and does not have promoter activity; it cannot initiate transcription. Its role is purely to mediate protein-RNA interactions once the antisense RNA has been produced. In another embodiment, “polymerase III promoter” refers to a DNA sequence recognized by RNA polymerase III to initiate transcription of small nuclear RNAs or other noncoding RNAs. Examples include Ul, U6, and U7 promoters. Non-limiting examples include, for example: the U7 promoter driving U7smOPT antisense RNAs with MS2 aptamer insertion (FIG. 2G); the U6 promoter in small RNA expression constructs; and the Ul promoter used in analogous snRNA targeting systems.

[0062] The term “contacting” means direct or indirect binding or interaction between two or more. A particular example of direct interaction is binding. A particular example of an indirect interaction is where one entity acts upon an intermediary molecule, which in turn acts upon the second referenced entity. Contacting as used herein includes in solution, in solid-20- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760phase, in vitro, ex vivo, in a cell and in vivo. Contacting in vivo can be referred to as administering, or administration.10063] The term “RBP tethering,” “RBP tether,” or “tethered” as provided herein refers to a system comprised of an RNA and a RBP. The RNA (e.g. an antisense RNA) contains an RNA structural element that is recognized by an exogeneous RNA-binding moiety which is fused to the RBP. The RNA-binding moiety binds to the RNA structural element, thereby recruiting the RBP to the RNA. The RNA structural element in the RBP tethering system is an aptamer. An aptamer is a sequence that bind specific molecules. In one aspect, the RNA structural element is an RNA stem-loop structure (“hairpin”). The hairpin is type of aptamer structures. The number of hairpins / aptamers can be modified, and multiple RNA-binding moieties can be recruited to a hairpin. Common tethering systems include, but are not limited to, the hairpins and coat proteins derived from bacteriophages MS2, , PP7, QP, GA, the bovine immunodeficiency virus, human U1 small nuclear ribonucleoprotein-specific protein U1A, and the iron response protein tethering system. Further details as to tethering can be found, for example, in Bos, T. J., et al. Tethered Function Assays as Tools to Elucidate the Molecular Roles of RNA-Binding Proteins. Advances in experimental medicine and biology, 907, 61-88 (2016). https: / / doi.org / 10.1007 / 978-3-319-29073-7_3 and Luo, E. C. et al. Large-scale tethered function assays identify factors that regulate mRNA stability and translation. Nat. Struct. Mol. Biol. 27, 989 (2020).

[0064] As used herein, the term “polyA site (PAS) selection” refers the selection of a polyA site on an RNAto attach a polyA tail to. Additional information about PAS selection can be found for example in Michael K Leung, Andrew Delong, Brendan J Frey, Inference of the human polyadenylation code, Bioinformatics, Volume 34, Issue 17, September 2018, Pages 2889-2898, DOI: 10.1093 / bioinformatics / bty211, the contents of which are incorporated by reference herein.

[0065] As used herein, “antisense RNA” refers to an RNA molecule comprising a nucleotide sequence complementary to all or a portion of a target RNA sequence, such that the antisense RNA is capable of specifically hybridizing to the target RNA by Watson-Crick base pairing. The antisense RNA may include additional functional modules, such as aptamers, hairpin structures, or RBP / RNP recruiting motifs, and may be chemically modified or unmodified.-21- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760Non-limiting examples of antisense RNAs include, for example: a U7 small nuclear RNA (U7 snRNA)-based transcript engineered to contain an MS2 hairpin aptamer and an antisense guide sequence positioned ~60 nt upstream of the AAUAAA motif of a polyA site;an antisense RNAs targeting the L3 adenovirus major late transcript PAS as used in the dual-luciferase reporter assays described herein; and an antisense RNAs targeting an endogenous proximal PAS in genes regulated by GRB2 or RNPS1 as identified in PAS-seq experiments.

[0066] As used herein, “aptamer” refers to a nucleic acid sequence or structure (e.g., RNA or DNA) that specifically binds to, and exhibits high affinity for, a target molecule such as an RNA-binding protein (RBP). Aptamers may be naturally occurring or engineered (including mutant sequences) and may form secondary / tertiary structures such as stem-loops, hairpins, or bulges. Aptamers may be recognized by endogenous RBPs or by exogenous RNA-binding moieties fused to RBPs in tethering systems. Non-limiting examples include, for example: the MS2 bacteriophage RNA hairpin (SEQ ID NO: 1) and its high-affinity mutant (SEQ ID NO: 2) used for MCP-RBP tethering in the luciferase PAS reporters (FIGS. 1 & 2); the PP7 bacteriophage RNA hairpin (SEQ ID NO: 5) used in analogous tethered function assays; and the iron-responsive element (IRE) aptamer (SEQ ID NO: 7) recognized by endogenous iron regulatory proteins.[0067J As used herein, “APA-modulating RBP” refers to an RNA-binding protein that regulates the selection or usage of polyadenylation sites within a transcript, thereby influencing mRNA isoform production. Such modulation can be activation or inhibition of PAS usage and may occur via direct PAS-proximal binding or indirectly via recruitment of CPA factors. Non-limiting examples include, for example: CPSF5 (NUDT21) and CPSF6, CFIm complex subunits shown to act as upstream activators when tethered near PAS elements (FIGS. 1 & 2); RNPS1, identified herein as a novel upstream activator acting via its N-terminal intrinsically disordered region (IDR) binding to mPSF and CPSF6 (FIGS. 5 & 6); and GRB2, shown to activate PAS selection through its C-terminal SH3 domain and direct interaction with CPSF6 / CPSF7.[0068J As used herein, “polyA site” or “PAS” refers to a specific location within a pre-mRNA transcript where cleavage and addition of a polyadenosine tail occur during 3' end-22- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760processing. PAS selection involves cis-elements (e.g., AAUAAA hexamer, upstream UGUA motifs, G / U-rich downstream elements), accessory proteins, and cell-specific context.Non-limiting examples include, for example the L3 PAS from the adenovirus major late transcript used in luciferase reporters with or without UGUA motifs (FIGS. 1A & 2G); endogenous PASs in CFIm-regulated targets altered upon GRB2 knockdown (FIG. 6E); and intronic and last-exon PASs quantified via PAS-seq following RNPS1 overexpression (FIG.61)

[0069] As used herein, “target region proximal to the PAS” refers to an RNA sequence located upstream or downstream of the PAS cleavage site, typically within 20-200 nucleotides, selected for hybridization by antisense RNA to enable locus-specific recruitment of RBPs. Non-limiting examples include, for example: sequences 60 nt upstream of the AAUAAA hexamer targeted by U7smOPT antisense guides (FIGS. 2H & 21); the 15 nt upstream location used for MS2 aptamer insertion in upstream luciferase reporters (FIG. 1A); and the 56 nt downstream location used in downstream reporters (FIG. 1A).[0070| As used herein, “RBP / RNP-recruiting motif’ refers to one or more sequences or structures within the antisense RNA that interact specifically with RNA-binding proteins or ribonucleoprotein complexes, conferring stability, localization, or regulatory activity.Non-limiting examples include, for example: MS2 hairpin aptamers recruiting MCP-fused CPSF5, CPSF6, GRB2, or RNPS1 in the tethering assays (FIGS. 1 & 2); stem-loop elements recruiting endogenous iron-responsive proteins via IRE aptamers; and engineered PP7 aptamers for recruiting PP7 coat protein-RBP fusions in analogous APA assays.

[0071] As used herein, “modulator of PAS selection” refers to any molecule, compound, protein, or engineered construct capable of altering cleavage and polyadenylation at a given PAS. Non-limiting examples include, for example: CPSF5 and CPSF6, activating PAS in upstream tethering assays; GRB2, modulating PAS via SH3 -mediated CFIm binding;RNPS1, modulating PAS via mPSF and CPSF6 interactions; and antisense RNA guides delivered via U7smOPT-MS2 targeting.

[0072] As used herein, “high affinity” refers to a binding interaction with a dissociation constant (K_d) of less than 100 nM, preferably less than 10 nM, under physiologically relevant conditions. Non-limiting examples include, for example: MS2 hairpin-MCP binding-23- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760pair used throughout tethered assays; PP7 hairpin-coat protein binding with nanomolar affinity; and (iii) CPSF6-GRB2 SH3 interaction with proline-rich motif binding within CFIm complexes.

[0073] As used herein, “operably linked” refers to the functional association between a promoter / regulatory sequence and a coding sequence such that the promoter controls transcription of the coding sequence. Non-limiting examples include, for example: CMV promoter operably linked to dual-luciferase reporters (FIG. 1A); U7 promoter operably linked to antisense guide sequences (FIG. 2G); and EF-la promoter operably linked to RBP-MCP fusion constructs in tethering assays.

[0074] As used herein, “regulating PAS selection” refers to altering the relative usage of one PAS versus another in a transcript, via activation or inhibition, producing measurable changes in mRNA isoforms. Non-limiting examples include, for example: increasing upstream L3 PAS usage by tethering CPSF5-MCP or CPSF6-MCP to MS2 sites (FIG. 1C, FIGS. 2H & 21); inhibiting P AS usage with HNRNPCL1-MCP in luciferase reporters; and shifting endogenous PAS usage following GRB2 knockdown or RNPS1 overexpression (FIGS. 4 & 6).

[0075] As used herein, “haploinsufficiency diseases” refer to genetic disorders wherein a single functional gene copy fails to produce sufficient gene product to maintain normal function, leading to disease. Non-limiting examples include, for example: Dravet syndrome, due to SCN1A haploinsufficiency; autosomal dominant retinitis pigmentosa, due to rhodopsin or PRPF31 haploinsufficiency; and neurofibromatosis type 1, due to NF1 haploinsufficiency, each a relevant target for PAS modulation technology disclosed herein.

[0076] In some aspects of this disclosure, the term “near” refers to a location within a specified nucleotide distance of a reference site, such as a polyadenylation site (PAS) cleavage site, as measured along the RNA sequence in either the 5' to 3' or 3' to 5' direction. Unless otherwise indicated, “near” encompasses positions located within about 0-200 nucleotides upstream or downstream of the reference site. Non-limiting examples from the experimental work include MS2 aptamer insertion sites located 15 nucleotides upstream or 56 nucleotides downstream of the AAUAAA hexamer in luciferase reporters, U7smOPT4918-7893-0315.1Atty. Dkt. No.: 114198-3760antisense guide positions located 0, 20, 40, 60, and 120 nucleotides upstream of the PAS, and PAS-proximal eCLIP binding peaks located within ±200 nucleotides of modulated PASs.10077] In some aspects of this disclosure, the term “proximal” when referring to a site or region, such as “proximal PAS,” refers to a polyadenylation site or sequence element located closer to the coding sequence (i.e., upstream in the 3 ' untranslated region or within an intron) relative to an alternative distal site in the same transcript. In a spatial context on the RNA molecule, “proximal” may also mean a location within about 1-200 nucleotides upstream or downstream of a PAS when describing binding sites, antisense targeting regions, or aptamer insertion positions. Non-limiting examples from the experimental work include U7 antisense guide targeting 60 nucleotides upstream of the AAUAAA hexamer, which produced maximal PAS activation by CPSF5 or CPSF6, “proximal” PAS usage shifts in GRB2, RNPS1, and CPSF6 knockdowns identified by PAS-seq, and RNPS1 eCLIP signal upstream of proximal PASs showing increased usage upon overexpression.

[0078] In some aspects of this disclosure, the term “distal” refers to a polyadenylation site located further downstream in the transcript, usually toward the 3' end, relative to a proximal site. Distal PAS usage results in longer 3' untranslated regions. Within positioning language, “distal” is the opposite of “proximal,” and may refer to a binding site or aptamer placement more than about 200 nucleotides away from the PAS cleavage site. Non-limiting examples from the experimental work include distal PAS usage shifts detected by PAS-seq following CPSF5 or CPSF6 knockdown and PASs in last-exon locations serving as distal relative to intronic PASs.

[0079] In some aspects of this disclosure, “upstream” refers to an RNA sequence located 5' to a reference point in the transcript, such as a PAS cleavage site or AAUAAA hexamer, according to transcriptional orientation. In nucleotide terms relative to PAS in this disclosure, “upstream” encompasses sequence positions 0-200 nucleotides before the PAS cleavage site, unless otherwise specified. Non-limiting examples include MS2 aptamer insertion 15 nucleotides upstream of PAS in luciferase reporters, antisense guide positions 60 nucleotides upstream used for CPSF5 or CPSF6 tethering to activate PAS, and CPSF5 or CPSF6 eCLIP binding peaks upstream of PASs that shift upon knockdown.-25- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[0080J In some aspects of this disclosure, “downstream” refers to an RNA sequence located 3' to a reference point such as the PAS cleavage site, measured according to transcriptional orientation. Unless otherwise indicated, “downstream” encompasses sequence positions within about 0-200 nucleotides after the PAS cleavage site. Non-limiting examples from the experimental work include MS2 aptamer insertion 56 nucleotides downstream of PAS in luciferase reporters, downstream PAS candidates activated in CSTF1 tethering assays, and eCLIP binding peaks downstream of distal PASs modulated by RNPS1.

[0081] In some aspects of this disclosure, “adjacent” refers to a position immediately next to or within 20 nucleotides of a reference element such as a PAS, AAUAAA motif, or aptamer hairpin. Non-limiting examples include aptamer placement directly adjacent to PAS in modified upstream luciferase reporters with no intervening sequence, sequences adjacent to AAUAAA motifs that were mutated to AATAAG to ablate PAS function, and binding peaks located adjacent to PASs in PAS-proximal eCLIP analyses.

[0082] In one aspect, the term “low complexity region” can be detected using computational sequence analysis methods that identify segments with reduced compositional diversity relative to the remainder of the sequence. For example, detection may be accomplished by applying algorithms such as SEG, CAST, or approaches based on Shannon entropy scoring, which quantify the frequency distribution of amino acids across a sliding window and compare it to expected random distributions. In one approach suitable for this disclosure, Shannon entropy is calculated for each window, and regions falling below a defined complexity threshold are annotated as LCRs. Boundaries are recorded as coordinate ranges within the full-length protein sequence. These analyses can be performed directly on the amino acid sequence of the RNA-binding protein, using standard bioinformatics tools. Once identified, LCRs can be intersected with other annotated features such as domains or motifs, for example by using bedtools intersect or similar coordinate overlap functions. This enables correlation of low complexity segments with functional features highlighted by the fine-tuned ProteinBERT model, such as occlusion peaks that indicate contribution to classification as a PAS activator or non-activator. The methodology provides an objective and reproducible means to annotate LCRs and to relate them to the mechanisms by which the protein modulates polyadenylation site selection.-26- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760

[0083] As used herein, the term “functional equivalent,” when referring to an alternative polyadenylation (APA)-modulating RNA-binding protein (RBP), a domain thereof, or a minimal effector module thereof, means a protein, polypeptide, or domain that retains the ability to activate or inhibit polyadenylation site (PAS) selection in the same positional context as the reference protein or domain. A functional equivalent may have a sequence identity of at least 70%, 80%, 85%, 90%, or 95% compared to the reference sequence, and may include conservative substitutions, deletions, or insertions that do not abolishPAS-modulating activity. The functional equivalent can be assayed for activity in vitro using reporter-based tethering assays, or in vivo using PAS-sequencing, in accordance with the methods described herein. Sequence identity can be determined using BLAST run under default parameters.

[0084] As used herein, the term “therapeutically beneficial isoform” refers to an mRNA transcript variant whose expression or polyadenylation site usage results in increased production of functional protein, restoration of normal protein activity, or improvement of cellular physiology with respect to a disease phenotype. In certain embodiments, the therapeutically beneficial isoform may contain a longer coding sequence that encodes a full-length protein, a 3' untranslated region of desired regulatory length, or the absence of destabilizing elements that occur in an alternative isoform. In other embodiments, the beneficial isoform produces an altered protein or transcript that is more stable or more efficiently translated.[00851 As used herein, the term “activator of PAS selection” refers to an APA-modulating RBP, RBP domain, or recruited protein complex that increases the relative usage of the targeted polyadenylation site compared to a control, resulting in increased cleavage and polyadenylation at that site. PAS activation may be measured as an increase in the Renilla-to-Firefly luciferase ratio in the dual-luciferase assay described herein, an increase in PAS-proximal read counts in PAS-sequencing data, or an increase in PAS-proximal binding enrichment in enhanced crosslinking immunoprecipitation (eCLIP) data.

[0086] As used herein, the term “inhibitor of PAS selection” refers to an APA-modulating RBP, RBP domain, or recruited protein complex that decreases the relative usage of the-27- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760targeted polyadenylation site compared to a control, resulting in decreased cleavage and polyadenylation at that site. PAS inhibition may be measured as a decrease in the Renilla-to-Firefly luciferase ratio in the dual-luciferase assay described herein, a decrease in PAS-proximal read counts in PAS-sequencing data, or a decrease in PAS-proximal binding enrichment in enhanced crosslinking immunoprecipitation (eCLIP) data.

[0087] As used herein, the term “minimal effector module” refers to a polypeptide segment, domain, or intrinsically disordered region derived from an APA-modulating RBP that, on its own and when recruited near a PAS, is sufficient to reproduce at least part of thePAS-modulating activity of the full-length protein. Non-limiting examples include the C-terminal SH3 domain of GRB2 and the N-terminal intrinsically disordered region of RNPS1. Minimal effector modules can be incorporated into PAS-targeted constructs to provide activation or inhibition of PAS usage without requiring the full-length RBP.

[0088] As used herein, the term “reporter construct” refers to a genetically encoded plasmid or vector comprising a detectable reporter cassette that includes at least one reporter gene, such as a luciferase, operably linked via transcriptional or translational elements to a polyadenylation site (PAS) of interest. The PAS can be endogenous, synthetic, or engineered, and may be modified to include or exclude cis-regulatory motifs. In certain embodiments, the reporter construct comprises PAS-proximal aptamer binding sites for tetheringAPA-modulating RBPs or domains via coat protein fusions, allowing quantification of PAS-modulating activity by measuring reporter output.[0089 J Modes for Carrying Out the Disclosure

[0090] RNA-binding proteins (RBPs) are critical regulators of gene expression at multiple levels, including transcription, pre-mRNA splicing, and 3' end processing1. Yet most of the over 1,500 annotated human RBPs remain functionally uncharacterized2. Pre-mRNA 3' end processing can occur at multiple poly(A) sites (PASs) within a single gene, a phenomenon known as alternative polyadenylation (APA). APA generates transcript isoforms that differ in 3' untranslated regions (UTRs) and / or coding sequences, thereby influencing the stability, translation, and localization of mRNA3’4. It occurs in over 70% of human genes5, and aberrant APA is implicated in a wide range of diseases, including cancer and neurological-28- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760disorders6,7. While several RBPs are known to regulate APA8, a comprehensive, systematic identification of APA regulators across the human RBP repertoire is still lacking.10091 ] Several studies have profiled proteins involved in APA using genetic perturbation of known regulators9 1'. While these approaches have provided valuable insights, they primarily assessed APA changes upon perturbation of known factors, limiting the discovery of novel regulators. A separate study conducted a genome-wide, pooled CRISPR / Cas9 screen with a dual-fluorescence reporter to identify novel cleavage and polyadenylation (CPA) factors12. However, depletion-based readouts cannot readily distinguish direct regulatory effects (via RNA or CPA factor binding) from indirect effects arising from upstream regulatory networks. In addition, the screen was based on a reporter containing the model PAS sequence and thus primarily recovered core CPA factors, likely missing RBPs that act at alternative PASs with different cis-regulatory contexts genome-wide.

[0092] Large-scale tethered function assays, which employ fusion proteins to recruit RBPs to reporters via RNA aptamers, have proven to be useful tools for uncovering novel roles of RBPs in alternative splicing (AS)13, mRNA stability, and translation14’15. They offer a simple and more robust alternative to genome-wide perturbation methods, and allow the identification of direct, causal effects via targeted recruitment. Crucially, these scalable assays enable functional characterization of RBPs without prior knowledge of their endogenous targets. Because many RBPs are modular, tethered assays can also map domainfunction relationships, facilitating the engineering of programmable modulation systems for targeted RNA regulation13.

[0093] In this study, Applicant used a dual-luciferase MS2 tethered-function assay to screen a library of 879 human RBPs to identify those influencing PAS selection. This approach revealed previously unrecognized activators of PAS usage, which allowed the fine-tuning of a protein language model that predicts PAS activators. Applicant further characterized two unexpected hits — GRB2 and RNPS1 — by biochemically defining their direct binding partners within the CPA machinery and mapping domains responsible for their activity. Together, these results advance the understanding of RBP-mediated APA regulation and provide a valuable resource for future studies of RBPs and their roles in disease.]0094| Antisense RNA, vectors, and compositions-29- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760

[0095] Described herein is a nuclear expressed antisense RNA comprising, from the 5’ end to the 3’ end: a) a nucleotide sequence encoding an aptamer, wherein the aptamer has a high affinity to an alternative polyadenylation (APA)-modulating RBP; and b) a nucleotide sequence antisense to a target region proximal to a polyA site (PAS).

[0096] In some embodiments, the nuclear expressed antisense RNA further comprises an RBP / RNP (ribonucleoprotein)-recruiting motif. The RBP / RNP (ribonucleoprotein)-recruiting motif is included on the 5’ end and / or 3’ end of the antisense portion to increase stability by recruiting RBPs / RNPs to block antisense RNA degradation. In one aspect, the antisense RNA further comprises a nucleotide sequence comprising an RBP / RNP (ribonucleoprotein)-recruiting motif.

[0097] In some embodiments, the antisense RNA has a nucleotide sequence according to the antisense RNA of as shown in Table 5.

[0098] In some embodiments, the target region proximal to the PAS is approximately 20bp, approximately 40bp, approximately 60bp, approximately 80bp, approximately lOObp, approximately 120bp, approximately 140bp, approximately 160bp, approximately 180bp, or approximately 200bp upstream of the PAS. In other embodiments, the target region proximal to the PAS is approximately 20bp, approximately 40bp, approximately 60bp, approximately 80bp, approximately lOObp, approximately 120bp, approximately 140bp, approximately 160bp, approximately 180bp, or approximately 200bp downstream of the PAS. Alternatively, the PAS is 20bp, or 40bp, or 60bp, or 80bp, or lOObp, or 120bp, or 140bp, or 160bp, or 180bp, or 200bp upstream of the PAS. In other embodiments, the target region proximal to the PAS is approximately 20bp, approximately 40bp, approximately 60bp, approximately 80bp, approximately lOObp, approximately 120bp, approximately 140bp, approximately 160bp, approximately 180bp, or approximately 200bp downstream of the PAS. Yet further, the target region proximal to the PAS is 20bp, or 40bp, or 60bp, or 80bp, or lOObp, or 120bp, or 140bp, or 160bp, or 180bp, or 200bp downstream of the PAS.

[0099] The high affinity aptamer can bind to the RBP. In some aspects, the high affinity aptamer binds directly to the RBP. In other aspects, the RBP binds to the aptamer by tethering.-30- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[01OO| In one aspect, the aptamer in the antisense RNA can be recognized by an exogeneous RNA-binding moiety which is fused to the RBP. The aptamer has a high affinity to the APA-modulating RBP due to interactions with the RNA-binding moiety (i.e. tethering system). In another aspect the aptamer in the antisense RNA can be recognized by an endogenous RNA-binding moiety present in an RBP in a cell.[01011 In one aspect, the aptamer and RNA-binding moiety are derived from a bacteriophage selected from: MS2, R17, , PP7, Qp or GA; iron responsive protein (IRP); bovine immunodeficiency virus (BIV), or human U1 small nuclear ribonucleoprotein A. In on aspect, the aptamer is derived from MS2. In some embodiments, the RNA hairpin or RNA- binding moiety comprise a sequence as shown in any one of SEQ ID NOs: 1-19, as shown in Table 3. In some aspects, the aptamer may be modified (e.g. mlpseudoU) or unmodified RNA.

[0102] In another aspect, the RNA aptamer comprises a sequence consisting ofSEQ ID NO: 1 (MS2 hairpin), SEQ ID NO: 2 (MS2 high affinity mutant), or an equivalent thereof having at least 90% sequence identity thereto; the aptamer is configured to recruit an RNA-binding protein selected from the group consisting of CPSF5 (SEQ ID NO: ) CPSF5 SEQ ID NO: ), CPSF6 (SEQ ID NO: ), RNPS1 (SEQ ID NO: ), and GRB2 (SEQ ID NO: ); the antisense RNA hybridizes to a target region proximal to a polyadenylation site, wherein “proximal” comprises a region upstream or downstream of the cleavage site sufficient to modulate PAS selection, without limitation to a fixed nucleotide distance; and optionally wherein the method is used for regulating PAS selection in an mRNA of a gene associated with a disease involving aberrant alternative polyadenylation, including cancers, haploinsufficiency disorders, and other APA-related pathological conditions. For example, the target region proximal to the PAS is approximately 20bp, approximately 40bp, approximately 60bp, approximately 80bp, approximately lOObp, approximately 120bp, approximately 140bp, approximately 160bp, approximately 180bp, or approximately 200bp upstream or downstream of the PAS.

[0103] In some embodiments the RNA-binding moiety has at least 80%, 90%, 95%, 99%, or 100% sequence identity to at least one of SEQ ID NOs: 11-19 and wherein the identity is determined by BLAST run under default parameters, or alternatively at least 80%, 90%, 95%,-31- 4918-7893-0315.1Atty. Dkt. No.: 114198-376099%, or 100% sequence identity to at least one of SEQ ID NOs: 11-19 across the complete sequence and wherein the identity is determined by BLAST run under default parameters.10104] The sequences of the RNA hairpin and / or RNA-binding moiety are shown in Table 3 are representative, rather than exhaustive. In some embodiments, the RNA hairpin can bind to more than one RNA-binding moiety. In some embodiments, the aptamer / hairpin and / or RNA-binding moiety are different than what is shown in Table 3.

[0105] In one embodiment the RNA hairpin and RNA binding moiety are derived from the bacteriophage MS2. In one embodiment, the RNA hairpin and RNA binding moiety are derived from the bacteriophage MS2, and the hairpin comprises, or consists essentially of, or yet further consists of an amino acid sequence according to SEQ ID NOs: 1 or 2, or an equivalent of each thereof, and the RNA-binding moiety comprises, or consists essentially of, or yet further consists of an amino acid sequence according to SEQ ID NO: 11 or an equivalent thereof.

[0106] DNA molecules encoding the above sequences are further provided herein.

[0107] In some embodiments, the aptamer is specific to the RBP.

[0108] Applicant also provides herein a vector comprising the antisense RNA or the DNA molecule encoding the antisense RNA. In some aspects, the vector is selected from a plasmid, a viral vector, a cosmid, or a phage, optionally wherein the viral vector is selected from a baculovirus, a retrovirus, or an adenovirus.

[0109] In some aspects, the vector further comprises a promoter for expression or replication of the RNA or DNA, e.g., a polymerase II promoter or a polymerase III promoter. In one aspect, the polymerase III promoter is selected from a Ul, U6, or U7 promoter.

[0110] In some aspects, the vector further comprises a nucleotide sequence encoding an RBP operably linked to the antisense RNA, wherein the RBP is a modulator of PAS selection. In one aspect, the RBP is mammalian. The RBP may be selected from any of the RBPs in Table 4 or FIGS. 2A or 2E. In one aspect, the RBP is selected from CPSF5, CPSF6, RNPS1, and GRB2. In some aspects, the RBP is a different RBP than those in Table 4 or FIGS. 2A or 2E. In one aspect, the RBP in the vector activates PAS selection, and wherein the RBP is selected from any of the RBPs in FIG. 2E.-32- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[01111 In one aspect of these embodiments, the RPB is mammalian.

[0112] Also described herein are isolated host cells comprising one or more of the antisense RNA, DNA encoding such, or vectors as described herein. In one aspect, when the vector does not comprise a nucleotide sequence encoding an RBP operably linked to the antisense RNA, the isolated host cell may further comprise a polynucleotide encoding the RBP or the RBP which the aptamer has a high affinity to. Optionally, the RBP may be selected from any of the RBPs in FIGS. 2A or E or Table 4. Optionally, the RBP or polynucleotide encoding the RBP in the host cell is selected from CPSF5, CPSF6, RNPS1, and GRB2. Optionally, the RBP may be a different RBP.

[0113] Also described herein are compositions comprising a carrier such as a pharmaceutically acceptable carrier and one or more of the antisense RNA, DNA encoding the RNA, vector, and / or isolated host cell as described herein.

[0114] Methods of regulating polyA site (PAS) selection on mRNA in a cel.

[0115] Provided herein are methods of regulating polyA site (PAS) selection on mRNA in a cell. According to one embodiment, the method comprises, consists of, or yet further consists essentially of contacting the cell with the antisense RNA or a vector or composition comprising the antisense RNA, and further contacting the cell with an RBP, or polynucleotide encoding an RBP, wherein the antisense RNA binds to the target region proximal to the PAS, and wherein the RBP activates or inhibits PAS selection of the mRNA in the cell. In one aspect, the RBP is selected from any one of the RBPs identified in Table 4 or FIGS. 2A or E. In another aspect, the RBP is a different APA-modulating RBP.

[0116] According to another embodiment, the method of regulating polyA site (PAS) selection on mRNA in a cell, comprises, consists of, or yet further consists essentially of contacting the cell with a vector comprising i) the antisense RNA; and ii) a nucleotide encoding an RBP, or a composition comprising the vector, wherein the antisense RNA binds to the target region proximal to the PAS, and wherein the RBP activates or inhibits PAS selection of the mRNA in the cell , and wherein the RBP activates or inhibits PAS selection of the mRNA in the cell.-33- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[0117| In one aspect, the RBP is selected from any one of the RBPs identified in Table 4 or FIGS. 2A or 2E. In another aspect, the RBP is a different APA-modulating RBP.10118] According to another embodiment, the method of regulating polyA site (PAS) selection on mRNA in a cell comprises, consists of, or yet further consists essentially of contacting the cell with the antisense RNA described herein, or vector and / or composition comprising the antisense RNA, wherein an endogenous RBP binds to the aptamer, and wherein the antisense RNA binds to a target region proximal to the PAS, and wherein the RBP activates or inhibits PAS selection of the mRNA in the cell.

[0119] In some aspects, the target mRNA is a gene implicated in disease. In one aspect, the gene is implicated in haploinsufficiency diseases. For example, the haploinsufficiency disease may be selected from Dravet syndrome, autosomal dominant retinitis pigmentosa, or neurofibromatosis type 1.

[0120] When both the first genetic vector and the second genetic vector are expressed in a cell, the nuclear expressed antisense RNA binds to the targeted PAS displaying the RBP-recognizing aptamer / sequence. The APA-modulating RBP is recruited to the targeted PAS, due to its high affinity for the RBP-recognizing aptamer / sequence on the antisense RNA and based on the effect of the RBP localized to the PAS, the RBP can either activate or inhibit APA selection.

[0121] In some embodiments, the APA-modulating RBP is encoded in the same vector as the vector expressing the nuclear expressed antisense RNA, and the single vector is expressed in the cell to activate or inhibit APA selection.[0122 [ In some embodiments, the RBP-recognizing aptamer / sequence on the antisense RNA recruits endogenous RBP / RNP complexes in the cell to affect PAS selection and activate or inhibit APA selection.

[0123] The RBP recruiting aptamer or sequence is an important component for the utility of the invention, and can recruit endogenous mammalian RBP / RNPs, including U snRNP components, hnRNP components, and any other nuclear expressed RBP / RNPs that include the tether protein. Without the recruiting aptamer or sequence, the nuclear expressed antisense RNA will lack the ability to effectively modulate APA sites.-34- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760

[0124] The genetic vector may be delivered into mammalian cells, animal models, or human patients with transfection reagents or in a viral vector such as lentivirus or adeno-associated virus (AAV). The nuclear expressed antisense RNA transcript is driven by a Polymerase III promoter — such as a Ul, U6, or U7 promoter — or any such RNA polymerase promoter that drives nuclear expression of an RNA transcript.

[0125] Screening methods

[0126] Also provided herein are methods of screening for RBPs that modulate PAS selection, comprising, or consisting essentially of, or yet further consisting of: a) providing a dual-luciferase bicistronic reporter comprising a first luciferase, a poly(A) signal (PAS) of interest, and a second luciferase; b) incorporating RNA hairpin aptamers upstream or downstream of the PAS for RBP tethering via a coat protein fusion; c) expressing in cells candidate RBP-coat protein fusions; and d) determining a luminescence ratio of the first luciferase to the second luciferase, wherein a change relative to a control indicates the RBP modulates PAS selection. The cell can be a prokaryotic or eukaryotic cell, as appropriate.

[0127] In one embodiment, the reporter is an upstream reporter with the aptamer located about 15 nt upstream of the PAS, or a downstream reporter with the aptamer located about 56 nt downstream of the PAS. In another embodiment, the PAS is an L3 adenovirus major late transcript PAS lacking UGUA motifs. In a further embodiment, wherein the coat protein is MS2 coat protein (MCP).

[0128] Applicant also provides methods of identifying high-confidence PAS activators, comprising performing the method noted above, in a first screen of a library of RBP candidates, performing the method again in a second screen of candidates positive in the first screen, and designating as high-confidence those RBPs showing consistent activation in both screens with reporter-position specificity.

[0129] Further provided is a pool of RBPs identified by these methods, e.g., wherein the pool comprises CPSF5, CPSF6, RNPS1, GRB2, MBNL1, MBNL2, YTHDF1, PCBP1, SCAF8, HNRNPF, and HNRNPH2.-35- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[013O| Also provided are methods of predicting RBP function in PAS selection, comprising: a) inputting amino acid sequences of RBPs into a fine-tuned protein language model trained on screening results of the methods described herein; b) classifying the RBPs as activators or non-activators based on a model score; and c) optionally outputting occlusion maps indicating sequence regions important for classification. In one embodiment, the protein language model is ProteinBERT fine-tuned on tethered function assay data. The methods can further comprise validating model predictions by MS2 tethering assays.1'0131] Mapping methods

[0132] Methods of mapping domain-level contributions to PAS activation in a candidate RBP are provided, the methods comprising, or consisting essentially of, or yet further consisting of: a) generating deletion variants or isolated domain constructs of the RBP fused to a coat protein; b) testing PAS activation by the method as described herein; and c) associating changes in activity with the presence or absence of specific domains.

[0133] In one aspect, the RBP is GRB2 and the domains comprise N-terminal SH3, SH2, and C-terminal SH3 domains. In another aspect, the RBP is RNPS1 and the domains comprise an N-terminal intrinsically disordered region (IDR), an RNA recognition motif (RRM), and a C-terminal RS-rich region.

[0134] Methods of determining direct interactions between an RBP and cleavage / polyadenylation (CPA) factor(s) is provided, the methods comprising, or consisting essentially of, or yet further consisting of: a) purifying CPA subcomplexes; b) performing an in-vitro binding assay with the RBP or RBP fragment; and c) detecting specific subcomplex binding. In one aspect, GRB2 binds directly to CFIm complexes containing CPSF6 or CPSF7. In another aspect, RNPS1 binds directly to mPSF and CPSF6.[0135| Methods of modulating endogenous PAS selection

[0136] Also provided are methods of modulating endogenous PAS selection, comprising, or consisting essentially of, or yet further consisting of: delivering to a cell: a) an RBP-coat protein fusion; and b) a small nuclear RNA (snRNA) engineered to include an aptamer binding site for the coat protein and an antisense sequence targeting a PAS of interest. In one -36- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760aspect, the snRNA is a U7smOPT snRNA. In another aspect, the antisense sequence targets a region about 60 nt upstream of the PAS. In a further aspect, the RBP is a high-confidence activator.

[0137] Compositions for modulating PAS selection

[0138] Composition for modulating PAS selection in a target transcript are provided, comprising, or consisting essentially of, or yet further consisting of: a) a minimal effector domain of an RBP identified as an activator by the method of identifying high-confidence PAS activators; and b) a programmable RNA targeting module configured to bind near a PAS. In one aspect, the minimal effector domain is the C-terminal SH3 domain of GRB2 or the N-terminal IDR of RNPS1.

[0139] Therapeutic methods

[0140] Further provided are methods of treating a disease associated with aberrant APA, comprising, or consisting essentially of, or yet further consisting of: administering to a subject an effective amount of the composition for modulating PAS selection in a target transcript the composition comprising: a) a minimal effector domain of an RBP identified as an activator; and b) a programmable RNA targeting module configured to bind near a PAS. In one aspect, the disease is selected from cancer, Dravet syndrome, autosomal dominant retinitis pigmentosa, and neurofibromatosis type 1.

[0141] Prediction methods

[0142] Also provided are methods of predicting protein regulators of polyadenylation site (PAS) selection, comprising, or consisting of, or yet further consisting of: a) obtaining amino acid sequences of a plurality of RNA-binding proteins (RBPs); b) labeling each of the sequences as an activator or a non-activator of PAS selection based on experimental results of a tethered function assay; c) inputting the sequences and corresponding labels into a sequence-based machine learning model; d) training the machine learning model to classify proteins as activators or non-activators of PAS selection; and e) outputting one or more predictions for an RBP of unknown PAS-activator status. In one aspect, the model contains-37- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760one or more of: a support vector machine, a convolutional neural network, or a protein language model. In one aspect, the protein language model is ProteinBERT.

[0013] In a further aspect, the method further adds fine-tuning the model with class weighting and hyperparameter optimization.1 144] In one aspect, the sequences are trimmed at each terminus prior to training.

[0145] In another embodiment, the method further adds interpreting the trained model by analyzing attention weights and / or generating occlusion maps to identify amino acid regions contributory to classification. In one embodiment, the contributory regions correspond to protein domains or low-complexity regions.

[0146] In one aspect, the methods add further filtering the amino acid sequences to remove pairs of proteins having at least 80% sequence identity before splitting into training and test datasets. For example, the machine learning model is validated on a set of zinc-finger proteins not included in the training dataset. In another embodiment, the validation achieves a recall of at least 0.9 and an accuracy of at least 70%.[0147| In a further aspect, the methods add further generating interpretive occlusion maps for a predicted activator from the zinc-finger protein set, wherein the occlusion maps identify functional domains associated with PAS regulation. In one aspect, the predicted activator is BRCA1 and the functional domains comprise a RING domain and a BRCT domain.10148] Methods and systems to predict RBP function

[0149] Referring now to FIG. 14A, illustrated is a flow chart of an implementation of a method 1400 for predicting RBP function in PAS selection, according to some implementations. The method 1400 may be performed by a data processing system, such as the computing device 1500, shown and described with reference to FIGS. 15A & 15B. The method 1400 may include any number of steps and the steps may be performed in any order. The data processing system may perform the method 1400 using the methods described herein.-38- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760

[0150] At operation 1402, the data processing system can input amino acid sequences of RBPs into a fine-tuned protein language model. The fine-tune protein language model can be trained on screening results for RBPs that modulate PAS selection, as described hereon. At operation 1404, the data processing system can classify the RBPs as activators ornon-activators based on a model score generated based on activation or execution of the finetune protein language model using the input amino acid sequence of RBPs. At operation 1406, the data processing system can optionally output occlusion maps indicating sequence regions important for classification.

[0151] Referring now to FIG. 14B, illustrated is a flow chart of an implementation of a method 1408 for mapping domain-level contributions to PAS activation in a candidate RBP, according to some implementations. The method 1408 may be performed by a data processing system, such as the computing device 1500, shown and described with reference to FIGS. 15A & 15B. The method 1408 may include any number of steps and the steps may be performed in any order. The data processing system may perform the method 1408 using the methods described herein.

[0152] At operation 1410, the data processing system can generate deletion variants or isolated domain constructs of the RBP fused to a coat protein. At operation 1412, the data processing system can test PAS activation. The data processing system can do so by generating screening results for RBPs that modulate PAS selection, as described herein. At operation 1414, the data processing system can associate changes in activity with the presence or absence of specific domains.

[0153] Referring now to FIG. 14C, illustrated is a flow chart of an implementation of a method 1416 for predicting protein regulators of polyadenylation site (PAS) selection), according to some implementations. The method 1416 may be performed by a data processing system, such as the computing device 1500, shown and described with reference to FIGS. 15A & 15B. The method 1416 may include any number of steps and the steps may be performed in any order. The data processing system may perform the method 1416 using the methods described herein.

[0154] At operation 1418, the data processing system can obtain amino acid sequences of a plurality of RNA-binding proteins (RBPs). At operation 1420, the data processing system can-39- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760label each of the sequences as an activator or a non-activator of PAS selection based on experimental results of a tethered function assay. At operation 1422, the data processing system can input the sequences and corresponding labels into a sequence-based machine learning model. At operation 1424, the data processing system can train the machine learning model to classify proteins as activators or non-activators of PAS selection. At operation 1426, the data processing system can output one or more predictions for an RBP of unknown PAS-activator status.

[0155] FIGS. 15A & 15B are block diagrams depicting embodiments of computing devices that can be used in connection with the methods and systems described herein.[0156J Having discussed specific embodiments of the present solution, it may be helpful to describe aspects of the operating environment as well as associated system components (e.g., hardware elements) in connection with the methods and systems described herein.

[0157] The systems and methods discussed herein may be deployed as and / or executed on any type and form of computing device, such as a computer, network device or appliance capable of communicating on any type and form of network and performing the operations described herein. FIGS. 15A & 15B depict block diagrams of a computing device 1500 useful for practicing an embodiment of the systems and methods described herein. As shown in FIGS. 15A & 15B, each computing device 1500 includes a central processing unit 1521, and a main memory unit 1522. As shown in FIG. 15A, a computing device 1500 may include a storage device 1528, an installation device 1516, a network interface 1518, an I / O controller 1523, display devices 1524a-1524n, a keyboard 1526 and a pointing device 1527, such as a mouse. The storage device 1528 may include, without limitation, an operating system and / or software. As shown in FIG. 15B, each computing device 1500 may also include additional optional elements, such as a memory port 1503, a bridge 1570, one or more input / output devices 1530a-1530n (generally referred to using reference numeral 1530), and a cache memory 1540 in communication with the central processing unit 1521.

[0158] The central processing unit 1521 is any logic circuitry that responds to and processes instructions fetched from the main memory unit 1522. In many embodiments, the central processing unit 1521 is provided by a microprocessor unit, such as: those manufactured by Intel Corporation of Mountain View, California; those manufactured by International-40- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760Business Machines of White Plains, New York; or those manufactured by Advanced Micro Devices of Sunnyvale, California. The computing device 1500 may be based on any of these processors, or any other processor capable of operating as described herein.

[0159] Main memory unit 1522 may be one or more memory chips capable of storing data and allowing any storage location to be directly accessed by the microprocessor 1521, such as any type or variant of Static random access memory (SRAM), Dynamic random access memory (DRAM), Ferroelectric RAM (FRAM), NAND Flash, NOR Flash and Solid State Drives (SSD). The main memory 1522 may be based on any of the above described memory chips, or any other available memory chips capable of operating as described herein. In the embodiment shown in FIG. 15A, the processor 1521 communicates with main memory 1522 via a system bus 1580 (described in more detail below). FIG. 15B depicts an embodiment of a computing device 1500 in which the processor communicates directly with main memory 1522 via a memory port 1503. For example, in FIG. 15B the main memory 1522 may be DRDRAM.[0160| FIG. 15B depicts an embodiment in which the main processor 1521 communicates directly with cache memory 1540 via a secondary bus, sometimes referred to as a backside bus. In other embodiments, the main processor 1521 communicates with cache memory 1540 using the system bus 1580. Cache memory 1540 typically has a faster response time than main memory 1522 and is provided by, for example, SRAM, BSRAM, or EDRAM. In the embodiment shown in FIG. 15B, the processor 1521 communicates with various VO devices 1530 via a local system bus 1580. Various buses may be used to connect the central processing unit 1521 to any of the I / O devices 1530, for example, a VESA VL bus, an ISA bus, an EISA bus, a MicroChannel Architecture (MCA) bus, a PCI bus, a PCI-X bus, a PCI-Express bus, or a NuBus. For embodiments in which the VO device is a video display 1524, the processor 1521 may use an Advanced Graphics Port (AGP) to communicate with the display 1524. FIG. 15B depicts an embodiment of a computer 1500 in which the main processor 1521 may communicate directly with VO device 1530b, for example via HYPERTRANSPORT, RAPIDIO, or INFINIBAND communications technology. FIG. 15B also depicts an embodiment in which local busses and direct communication are mixed: the processor 1521 communicates with I / O device 1530a using a local interconnect bus while communicating with I / O device 1530b directly.-41- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[01611 A wide variety of I / O devices 1530a-1530n may be present in the computing device 1500. Input devices include keyboards, mice, trackpads, trackballs, microphones, dials, touch pads, touch screens, and drawing tablets. Output devices include video displays, speakers, inkjet printers, laser printers, projectors and dye-sublimation printers. The I / O devices may be controlled by an I / O controller 1523 as shown in FIG. 15A. The I / O controller may control one or more I / O devices such as a keyboard 1526 and a pointing device 1527, e.g., a mouse or optical pen. Furthermore, an I / O device may also provide storage and / or an installation device 1516 for the computing device 1500. In still other embodiments, the computing device 1500 may provide USB connections (not shown) to receive handheld USB storage devices such as the USB Flash Drive line of devices manufactured by Twintech Industry, Inc., of Los Alamitos, California.

[0162] Referring again to FIG. 15A, the computing device 1500 may support any suitable installation device 1516, such as a disk drive, a CD-ROM drive, a CD-R / RW drive, a DVD-ROM drive, a flash memory drive, tape drives of various formats, USB device, hard-drive, a network interface, or any other device suitable for installing software and programs. The computing device 1500 may further include a storage device, such as one or more hard disk drives or redundant arrays of independent disks, for storing an operating system and other related software, and for storing application software programs such as any program or software 1520 for implementing (e.g., configured and / or designed for) the systems and methods described herein. Optionally, any of the installation devices 1516 could also be used as the storage device. Additionally, the operating system and the software can be run from a bootable medium.

[0163] Furthermore, the computing device 1500 may include a network interface 1518 to interface to a network through a variety of connections including, but not limited to, standard telephone lines, LAN or WAN links (e.g., 802.11, Tl, T3, 156kb, X.25, SNA, DECNET), broadband connections (e.g., ISDN, Frame Relay, ATM, Gigabit Ethernet, Ethemet-over-SONET), wireless connections, or some combination of any or all of the above. Connections can be established using a variety of communication protocols (e.g., TCP / IP, IPX, SPX, NetBIOS, Ethernet, ARCNET, SONET, SDH, Fiber Distributed Data Interface (FDDI), RS232, IEEE 802.11, IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.1 In, IEEE 802.1 lac, IEEE 802.1 lad, CDMA, GSM, WiMax and direct asynchronous connections). In -42- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760one embodiment, the computing device 1500 communicates with other computing devices 1500’ via any type and / or form of gateway or tunneling protocol such as Secure Socket Layer (SSL) or Transport Layer Security (TLS). The network interface 1518 may include a built-in network adapter, network interface card, PCMCIA network card, card bus network adapter, wireless network adapter, USB network adapter, modem or any other device suitable for interfacing the computing device 1500 to any type of network capable of communication and performing the operations described herein.

[0164] In some implementations, the computing device 1500 may include or be connected to one or more display devices 1524a-1524n. As such, any of the I / O devices 1530a-1530n and / or the I / O controller 1523 may include any type and / or form of suitable hardware, software, or combination of hardware and software to support, enable or provide for the connection and use of the display device(s) 1524a-1524n by the computing device 1500. For example, the computing device 1500 may include any type and / or form of video adapter, video card, driver, and / or library to interface, communicate, connect or otherwise use the display device(s) 1524a-1524n. In one embodiment, a video adapter may include multiple connectors to interface to the display device(s) 1524a-1524n. In other embodiments, the computing device 1500 may include multiple video adapters, with each video adapter connected to the display device(s) 1524a-1524n. In some implementations, any portion of the operating system of the computing device 1500 may be configured for using multiple displays 1524a-1524n. One ordinarily skilled in the art will recognize and appreciate the various ways and embodiments that a computing device 1500 may be configured to have one or more display devices 1524a-1524n.

[0165] In further embodiments, an I / O device 1530 may be a bridge between the system bus 1580 and an external communication bus, such as a USB bus, an Apple Desktop Bus, an RS-232 serial connection, a SCSI bus, a FireWire bus, a FireWire 500 bus, an Ethernet bus, an AppleTalk bus, a Gigabit Ethernet bus, an Asynchronous Transfer Mode bus, a FibreChannel bus, a Serial Attached small computer system interface bus, a USB connection, or a HDMI bus.

[0166] A computing device 1500 of the sort depicted in FIGS. 15A & 15B may operate under the control of an operating system, which control scheduling of tasks and access to-43- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760system resources. The computing device 1500 can be running any operating system, such as any of the versions of the MICROSOFT WINDOWS operating systems, the different releases of the Unix and Linux operating systems, any version of the MAC OS for Macintosh computers, any embedded operating system, any real-time operating system, any open source operating system, any proprietary operating system, any operating systems for mobile computing devices, or any other operating system capable of running on the computing device and performing the operations described herein. Typical operating systems include, but are not limited to, Android, produced by Google Inc.; WINDOWS 7 and 8, produced by Microsoft Corporation of Redmond, Washington; MAC OS, produced by Apple Computer of Cupertino, California; WebOS, produced by Research In Motion (RIM); OS / 2, produced by International Business Machines of Armonk, New York; and Linux, a freely-available operating system distributed by Caldera Corp, of Salt Lake City, Utah, or any type and / or form of a Unix operating system, among others.

[0167] The computer system 1500 can be any workstation, telephone, desktop computer, laptop or notebook computer, server, handheld computer, mobile telephone or other portable telecommunications device, media playing device, a gaming system, mobile computing device, or any other type and / or form of computing, telecommunications or media device that is capable of communication. The computer system 1500 has sufficient processor power and memory capacity to perform the operations described herein.

[0168] In some implementations, the computing device 1500 may have different processors, operating systems, and input devices consistent with the device. For example, in one embodiment, the computing device 1500 is a smart phone, mobile device, tablet or personal digital assistant. In still other embodiments, the computing device 1500 is an Android-based mobile device, an iPhone smart phone manufactured by Apple Computer of Cupertino, California, or a Blackberry or WebOS-based handheld device or smart phone, such as the devices manufactured by Research In Motion Limited. Moreover, the computing device 600 can be any workstation, desktop computer, laptop or notebook computer, server, handheld computer, mobile telephone, any other computer, or other form of computing or telecommunications device that is capable of communication and that has sufficient processor power and memory capacity to perform the operations described herein.-44- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[0169| Although examples of communications systems described above may include devices operating according to an 802.11 standard, it should be understood that embodiments of the systems and methods described can operate according to other standards and use wireless communications devices other than devices configured as devices and APs. For example, multiple-unit communication interfaces associated with cellular networks, satellite communications, vehicle communication networks, and other non-802.11 wireless networks can utilize the systems and methods described herein to achieve improved overall capacity and / or link quality without departing from the scope of the systems and methods described herein.[0170| It should be noted that certain passages of this disclosure may reference terms such as “first” and “second” in connection with devices, mode of operation, transmit chains, antennas, etc., for purposes of identifying or differentiating one from another or from others. These terms are not intended to merely relate entities (e.g., a first device and a second device) temporally or according to a sequence, although in some cases, these entities may include such a relationship. Nor do these terms limit the number of possible entities (e.g., devices) that may operate within a system or environment.10171] It should be understood that the systems described above may provide multiple ones of any or each of those components and these components may be provided on either a standalone machine or, in some implementations, on multiple machines in a distributed system. In addition, the systems and methods described above may be provided as one or more computer-readable programs or executable instructions embodied on or in one or more articles of manufacture. The article of manufacture may be a floppy disk, a hard disk, a CD-ROM, a flash memory card, a PROM, a RAM, a ROM, or a magnetic tape. In general, the computer-readable programs may be implemented in any programming language, such as LISP, PERL, C, C++, C#, PROLOG, or in any byte code language such as JAVA. The software programs or executable instructions may be stored on or in one or more articles of manufacture as object code.EXPERIMENTAL-45- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[01721 The following examples are provided to illustrate and not limit the scope of this disclosure.EXPERIMENTAL

[0173] Development of Dual-Luciferase MS2 Tethering Assay to Identify RBPs Regulating PAS Selection[0174J Applicant constructed two dual-luciferase, bicistronic reporters and established a ratio-based readout to assess how RBP recruitment alters PAS usage (FIG. 1A). This system utilizes MS2 stem-loops and MS2 coat protein (MCP)-based tethering to precisely recruit RBPs. Each reporter consists of a cytomegalovirus (CMV) promoter followed by Renilla luciferase open reading frame (ORF), a modified L3 poly(A) signal (137 nucleotide (nt) surrounding the cleavage site (or PAS) derived from the adenovirus major late transcript) bearing MS2 hairpins, an internal ribosomal entry site (IRES), a Firefly luciferase ORF, and a mouse P-globin poly(A) signal. In this construct, the L3 poly(A) signal was modified by removing the two UGUA motifs upstream of the PAS to prevent CFIm binding, which would otherwise strongly activate the PAS and confound the tethering effects16. All other regulatory elements were retained, including the AAUAAA hexamer and the G / U-rich downstream element. To maximize detection of RBPs that modulate APA, Applicant installed three MS2 stem-loops either 15 nt upstream (“upstream reporter”) or 56 nt downstream (“downstream reporter”) of the AAUAAA hexamer, providing two positional tethering contexts.Transcription from the CMV promoter proceeds through the Renilla ORF and then either terminates at the L3 PAS or reads through the IRES into the Firefly ORF, depending on PAS usage. Consequently, greater L3 PAS usage reduces Firefly luciferase expression and lower usage increases Firefly expression, corresponding to an increase and decrease in Renilla / Firefly ratio, respectively (FIG. IB). Applicant therefore used the luminescence ratio to classify tethered proteins as activators (increase in ratio) or inhibitors (decrease in ratio) of L3 PAS usage relative to the downstream PAS.

[0175] Applicant validated the assay in HEK293T cells using FLAG-MCP as a negative control, CPSF5 (NUDT21)-MCP and CPSF6-MCP as upstream activator controls16, and HNRNPCL1-MCP, a paralog of a known inhibitor, as an inhibitor control17. Each reporter was co-transfected with the indicated MCP fusion, and Renilla / Firefly ratios were measured.-46- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760When normalized to FLAG-MCP, CPSF5-MCP and CPSF6-MCP significantly increased ratios when tethered upstream of the L3 PAS (PAS activation), but showed no activation when tethered downstream (FIG. 1C). In contrast, HNRNPCL1-MCP decreased the ratio from both upstream and downstream positions, consistent with inhibition. These locationdependent effects are consistent with established models in which CFIm subunits act from upstream16,18’19. To exclude overexpression artifacts, Applicant compared the results against untethered CPSF5, CPSF6, and HNRNPCL1 and confirmed that the phenotype requires directed recruitment (FIG. 7A). Finally, RT-qPCR results verified that the luciferase readout mirrored RNA-level changes in PAS usage and was not significantly affected by changes in translation (FIGS. ID & 7B)

[0176] After validating the assay, Applicant performed a large-scale screen using 879 full-length RBP-MCP fusions from the previously established library (FIG. IE)13’15. Applicant co-transfected each RBP-MCP plasmid with either the upstream or downstream reporter into HEK293T cells in triplicate wells. Screens were run in an arrayed, 96-well format with plate-matched controls (FLAG-MCP, CPSF5 / 6-MCP, HNRNPCL1-MCP) on every plate to manage variability and ensure batch consistency. For each well, Applicant computed the Renilla / Firefly ratio and normalized to FLAG-MCP from the same plate. Two-sided t-tests (P < 0.05) identified candidates for secondary screening.

[0177] Identification of high-confidence RBPs that modulate PAS selection

[0178] Following the large-scale tethering screen, Applicant identified 107 activators and 380 inhibitors with the upstream reporter, and 95 activators and 386 inhibitors with the downstream reporter (FIGS. 2A & 2B. 44 activators and 232 inhibitors were common across both reporters. As expected, CPSF5 and CPSF6 ranked among the strongest upstream activators, consistent with their known functions16, whereas CSTF1 (a core CPA factor) was the strongest downstream activator, in line with its recognition of G / U-rich downstream elements20. Several known APA regulators were also recovered, including MBNL1 and RBM10, HNRNPC, and its paralog HNRNPCL117‘20 23. These results demonstrate the effectiveness of the screen in identifying RBPs that modulate PAS usage.

[0179] Functional enrichment analysis revealed that activators are strongly enriched for RNA processing pathways, whereas inhibitors showed weaker specificity for RNA biology and-47- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760were more associated with protein biosynthesis / other metabolic pathways (FIG. 2C). This aligned with observations in STRINGdb, where activators displayed greater connectivity to CPA factors than inhibitors (%2test, P = 0.006; FIG. 8A). Likewise, domain overrepresentation (using the RBP-MCP library as background and Fisher’s exact test) was observed only among activators, including canonical RNA-binding and non-RBD interaction modules. While domains were present among inhibitors, none were enriched (FIG. 2D). Together, these observations suggest activation of PAS usage more often depends on specific, domain-mediated interactions, whereas inhibition may reflect heterogeneous or indirect effects (e.g., steric hindrance). Accordingly, Applicant focused subsequent analyses on activators to maximize mechanistic insight.

[0180] Applicant next conducted a second tethering screen using all activators identified from the primary screen. Applicant recovered 52% of the upstream 64% of the downstream activators from the primary screen (FIG. 8B). Rank-density and between-screen rank correlations (Spearman’s rank, R = 0.709 upstream; 0.807 downstream) indicated that RBPs with stronger effects were reproduced more faithfully (FIGS. 8B & 8C). Applying the positional-consistency filter yielded 63 activators: 17 upstream-specific, 20 downstreamspecific, and 26 location-independent (FIGS. 2F & 8D. Applicant designated these RBPs as high-confidence activators: those significant in both rounds (two-sided t-test, P < 0.05) and retaining a consistent positional effect.

[0181] Excluding positive controls (CPSF5 and CPSF6), only seven high-confidence activators had prior links to APA (MBNL1, MBNL2, YTHDF1, PCBP1, SCAF8, HNRNPF, HNRNPH2) indicating that most of the high-confidence RBPs are novel21'24 27. Among the top-ranked novel candidates were GRB2, RNPS1, and RBM22, none of which had been characterized previously in the context of APA.[01821 To ensure the signal from the screen reflects APA rather than reporter translation or stability artifacts, Applicant conducted three validation analyses. First, RT-qPCR of reporter mRNA for 20 select tethered RBPs recapitulated the luminescence changes (Pearson R = 0.78 upstream; 0.87 downstream; FIG. 2F), indicating the luminescence readout reflects steady-state RNA abundance. Second, cross-referencing with the prior tethered-function stability and translation screenl5 revealed only seven overlapping RBPs (FIG. 8E); the-48- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760overlap included canonical APA factors (e.g., CPSF5), consistent with RBPs having multiple roles across RNA processing steps rather than stability-driven artifacts. Third, acute depletion of the RNA exosome subunit EXOSC3 using a dTAG degron system did not alter reporter signals for six high-confidence activators or for controls (FIG. 8F) suggesting that the exosome-mediated RNA decay does not drive the effects observed for RBPs that Applicant tested. Collectively, these data support that the tethering assay readout observed for the disclosed high-confidence RBPs are predominantly APA-associated.

[0183] Finally, to extend Applicant’s findings beyond fixed MS2 sites, Applicant developed a programmable U7smOPT-MS2 recruitment system to direct RBPs near a PAS. Because snRNAs are modular and have a therapeutic track record, Applicant chose U7smOPT28,29. Prior work has also shown that U1 / U7 can be reprogrammed for splicing modulation, ADAR recruitment, and pseudouridylation30’31. Applicant incorporated an MS2 stem-loop into the 3' end of the U7 snRNA guide (FIG. 2G), removed the MS2 stem-loop from the upstream reporter, and designed 40-nt guides targeting 0, 20, 40, 60, or 120 nt upstream of the AAUAAA hexamer to scan for optimal positioning (FIG. 2H). Co-expression with CPSF5-MCP or CPSF6-MCP revealed a strong positional effect, with 60 nt upstream giving maximal PAS activation, consistent with prior work16. Control constructs lacking MS2 or the PAS abolished the effect (FIG. 8G). Applicant then applied the system to two screen-nominated upstream activators, GRB2 and RNPS1, and observed congruent activation using U7smOPT-MS2 (FIG. 21). Although Applicant only demonstrate proof of concept with reporters, the system is readily adaptable for endogenous targeting.

[0184] In summary, Applicant’s large-scale tethering screen identified novel RBPs that activate or inhibit PAS usage, with activators showing greater mechanistic specificity. A two-tier strategy yielded 63 high-confidence activators for follow-up. U7smOPT-MS2 platform shows that these activators can be programmatically recruited near target PASs, highlighting a path toward endogenous, locus- and position-specific APA modulation with biomedical relevance.

[0185] Fine-tuned protein language model predicts PAS activators from sequence and highlights domain-level features-49- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[0186| To translate the screen results into a predictive tool, Applicant trained a sequencebased classifier to (i) virtually screen PAS activators for experimental follow-up and (ii) predict sequence features contributing to activation (FIG. 3A). To mitigate data leakage, Applicant removed protein pairs with >80% sequence identity (FIG. 9A) while retaining more divergent paralogs (Methods). Data were split 90% / l 0% into training / held-out sets with label stratification to address the strong class imbalance between non-activators and activators. For benchmarking and model development, Applicant adopted HydRA32as the starting point. HydRA is sequence-driven, learns functional features directly from amino acid input, and offers built-in evaluation and interpretation tools, aligning with the goals.Applicant omitted the SONAR (PPI) module due to incomplete coverage for the library. Applicant referred to this variant as sequence-driven HydRA.|0187] Using 10-fold, class-stratified cross-validation, Applicant benchmarked the sequence-driven HydRA ensemble and its components (ProteinBERT, CNN, SVM). ProteinBERT was the top performer by both ROC-AUC and PR-AUC (FIG. 3B, Table 1), with the latter more informative given class imbalance.Table 1: AUC and PR-AUC Scores of each benchmarked model using 10-fold cross validation with stratified split.

[0188] The sequence-driven HydRA ranked second. Applicant speculate that this reflects the weaker performance of the SVM and CNN constituents, whereas ProteinBERT’ s large-scale pre-training likely confers few-shot learning capability and capacity to capture local motifs and long-range dependencies33. Following hyperparameter tuning (Methods), the fine-tuned ProteinBERT achieved ROC-AUC = 0.77 (random=0.5) and PR-AUC = 0.48 (random=0.2)-50- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760(FIG. 9B) on the held-out protein set, indicating effective capture of sequence features relevant to PAS activation.10189] Next, Applicant applied sliding-window occlusion to map sequence regions driving activator classification. Most RBPs with significant occlusion peaks had at least one peak overlapping an annotated domain (FIG. 3C); notably, low-complexity regions (LCRs) were also frequently highlighted as informative. Among annotated domains, RRM1, DEAD, and Helicase C were most frequently found to be important across activators (FIG. 3D). Case studies matched known biology (FIG. 3E): in CPSF6, C-terminal RS-rich LCR was predicted to be important, consistent with its interaction with FIP1 in the core CPA machinery16. For SCAF8, the model prioritizes RRM1, mirroring results in the SCAF8-like yeast protein Sebl, where RRM1 is essential for 3' end processing34’35. Together, these results indicate that the fine-tuned ProteinBERT learns biologically meaningful features underlying PAS activation.

[0190] To test generalization, Applicant evaluated an independent set of 28 zinc-finger proteins (ZFPs) not used in training. These proteins were selected from ZFPs prioritized by Gosztyla et al.36, which profiled RNA targets and activities for >100 ZFPs. ZFPs provide a relevant test case because many are putative RBPs lacking canonical RNA-binding domains. Applying the Fl -optimized threshold (0.15, see Methods) and validating predictions independently with the MS2 tethering assay, the model recovered thirteen of fourteen activators (recall = 0.93) with 71.4% accuracy across 28 proteins (FIGS. 3F & 9C-9E), supporting its use as a virtual screening tool to shortlist candidates for experimental followup. Among predicted hits, BRCA1 showed occlusion peaks in the RING (zf-C3HC4) and BRCT domains. Notably, BRCA1 (via BARD1) interacts with CSTF1 to inhibit polyadenylation37’38, suggesting context-dependent roles in PAS regulation. Together, these results show that the model generalizes to unseen proteins, learns biologically meaningful features directly from sequence, and offers interpretable domain-level hypotheses for mechanisms of PAS activation.

[0191] High-confidence activators modulate endogenous PAS selection

[0192] Applicant asked whether high-confidence activators modulate endogenous PAS selection by performing knockdown followed by RNA-seq in HEK293T cells. For each candidate, Applicant used two complementary methods: (1) PAS-seq, a 3' end RNA-51- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760sequencing method that identifies and quantifies global APA events39,40, and (2) bulk-RNA-seq, to assess alternative splicing (AS) events. After quality control steps (knockdown efficiency, library quality), 39 candidates were analyzed (Methods). Upon knockdown, PAS-seq detected global APA shifts - both proximal (shortening) and distal (lengthening) — to varying extents across genes, confirming that the screen faithfully identified RBPs that modulate endogenous PAS usage (FIG. 4A, Table 2).

[0193] Table 2: Number of APA events, stratified by direction and the location of the PAS.-52- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[0194| As expected, knockdown of CPSF5 and CPSF6 induced strong APA shifts towards proximal PAS, involving over 2,400 and 700 events, respectively, consistent with previous studies9,10,16. Among the RBPs with over 1,000 affected events were MBNL2 and hnRNPF, both previously linked to APA21,41. Whereas MBNL depletion in mouse embryonic fibroblasts has been associated with proximal shifts, in HEK293T cells Applicant observed a predominantly distal shift, suggesting cell-type-specific regulation. For hnRNPF, prior work centered on B cells and select genes41; in contrast, the transcriptome-wide PAS-seq reveals extensive hnRNPF-dependent APA in HEK293T. Two factors primarily known for splicing regulation, TRNAU1AP and RBM513’42’43also showed over 900 and 800 APA events, respectively.-53- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760

[0195] To test whether different RBPs regulate distinct gene sets, Applicant performed functional enrichment on genes showing APA after knockdown (FIG. 4B). CPSF5 and CPSF6 displayed highly overlapping categories, including intracellular transport, consistent with their shared role in the CFIm complex. Interestingly, most activators showed RBP-specific functional signatures of APA. Notably, genes undergoing APA upon GRB2 perturbation were enriched for DNA damage response, aligning with reported nuclear roles of GRB2 in DNA damage response44,45. These results support a model in which RBPs identified from the screen preferentially modulate APA of discrete functional modules.

[0196] APA and AS are mechanistically linked53946 4S. To assess each RBP’s involvement in both processes, Applicant analyzed global AS events using bulk RNA-seq following RBP knockdown (FIG. 10A) and computed the Jaccard index (JI) between each factor’s APA- and AS-affected gene sets. Applicant then clustered RBPs using JI together with the number of known CPA and splicing interactors from the STRING physical-interaction subnetwork49as features (FIG. 4C). Six clusters emerged: Clusters 1, 5, and 6 displayed higher JI and stronger association with either splicing or CPA interactors, consistent with factors that impact both APA and AS but are more tightly connected to one machinery. Clusters 2 and 4 showed lower JI with interactors skewed toward splicing or CPA, respectively, suggesting specialized roles in one process or the other. Cluster 3 lacked substantial interactions with either machinery, potentially representing uncharacterized or multifunctional RBPs with context-specific roles in PAS selection. To further probe these observations, Applicant performed functional enrichment on each candidate’s interactors. The top ten functional categories for all clusters except Cluster 3 fell almost exclusively within “Gene expression and RNA processing,” while those in Cluster 3 spanned broader functional categories, suggesting that they are multifunctional and have context-specific roles in PAS selection (FIG. 4D)

[0017] To assess whether the candidate RBPs regulate APA by directly binding to the affected RNAs, Applicant generated new eCLIP data and integrated published profiles13,36for a total of 14 RBPs spanning 5 of 6 clusters. After stratifying by APA direction (shortening vs. lengthening), 8 of 12 RBPs showed more than 50% gene-level overlap between high-confidence eCLIP signal and KD-induced APA genes (FIG. 10B), indicating a broad association between binding and APA modulation. Within the overlapping sets, signals -54- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760mapped to different regions of the transcript, showing factor-dependent enrichment in coding sequence (CDS), introns, or 3' UTRs (FIG. IOC). Motif analysis revealed factor-specific patterns, including the expected UGUA motif for CPSF5 (FIG. 10D).

[0198] To test whether RBPs bind near the PASs that they regulate, Applicant assessed PAS-proximal enrichment by comparing eCLIP peaks within ±200 nt of modulated PASs to those near unaffected PASs (FIG. 4E). CPSF5, CPSF6, ZC3HAV1, RBM5, MBNL2, and GRB2 showed significantly enriched binding near modulated PASs, consistent with RNA binding-mediated regulation. Applicant then computed direction-stratified binding density (A = constitutive - KD-favored) within ±200 nt, which revealed factor-specific positional biases across RBPs (FIG. 4F). CPSF5 and CPSF6 exhibited a strong upstream bias among shortening events (the predominant KD class), recapitulating known mechanisms16and the screen; MBNL2 likewise showed upstream enrichment, in line with prior work21. Binding preferences also varied with the location of the constitutive PAS (intronic vs. last-exon; FIG.4G). Overall, these analyses indicate that some RBPs act via PAS-proximal RNA binding, whereas others likely influence PAS choice through protein-protein interactions with CPA components. Applicant notes that RNA-mediated regulation occurring beyond the ±200 nt window could be missed, and limitation of eCLIP (cross-linking efficiency, antibody performance) may obscure true PAS-proximal binding.

[0199] Together, these results show that high-confidence activators modulate PAS selection in endogenous contexts, reshaping APA across distinct gene programs. eCLIP and interaction analyses indicate two modes of action: RNA-binding-mediated and RNA-independent (likely via protein-protein contacts with CPA). These findings reveal a broader and more diverse set of APA regulators than previously recognized.

[0200] AP-MS reveals interactions between the candidates and the CPA machinery [02011 GRB2 and RNPS1 emerged as unexpected top hits from the screen and represent distinct groups from the functional clustering analysis (FIGS. 2 & 4). GRB2 is a well-studied adaptor protein that is ubiquitously expressed in eukaryotes and plays a critical role in signal transduction pathways, such as the RAS / mitogen-activated protein kinase (RAS / MAPK) pathway50. While GRB2 has been extensively studied in the context of its cytoplasmic role in cell signaling51,52, it has also been reported to localize to the nucleus53,54, and has been-55- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760implicated in DNA damage response44,45,55. However, its other nuclear functions, especially in mRNA processing, remain unknown. Notably, GRB2 lacks a canonical RNA-binding domain. However, it was recovered by UV-crosslinking-based RNA interactome capture in HeLa cells56, and was therefore included in the library. This study further confirms its RNA-binding activity and binding preference (FIGS. 4E-4G). RNPS1 is a splicing factor and a peripheral component of the exon junction complex (EJC)57 59, an integral part of the spliced messenger ribonucleoprotein (mRNP)60,61 and plays a crucial role in post-transcriptional gene regulation, including mRNA transport, nonsense-mediated decay (NMD), and translation62. RNPS1 has been shown to influence 3' processing in vitro63and is enriched in nuclear speckles, membraneless organelles that function as hubs for splicing and 3' end processing64,65. However, its involvement in APA or interaction with CPA factors has not been previously investigated.

[0202] To investigate how GRB2 and RNPS1 influence PAS selection, Applicant characterized their protein interactomes using affinity purification-mass spectrometry (AP-MS) in HEK293T cells. Applicant identified 220 and 236 interactors for GRB2 and RNPS1, respectively, of which 57 GRB2 and 77 RNPS1 interactors were previously documented in the STRING physical interaction subnetwork database (FIGS. 5A, HA & 11B. As expected, the GRB2 interactome included canonical cytoplasmic partners involved in RAS signaling, such as S0S1, S0S2, and GABI66(FIG. 5B). RNPS1 interactome included the EJC components, such as the core factor EIF4A3 and peripheral factors ACINI and SRR 162. Notably, both interactomes contained core CPA factors. GRB2 interactome was enriched for CFIm subunits (CPSF5, CPSF6, CPSF7) and PABPC1, while RNPS1 co-purified a broader set encompassing CFIm and additional CPA subcomplexes — mPSF (CPSF1, CPSF4, FIP1), mCF complex (CPSF2, CPSF3), CstF (CSTF1)— along with RBBP6 and PABPN1.

[0203] To confirm these findings, Applicant performed co-immunoprecipitation (co-IP) experiments for GRB2 and RNPS1 using antibodies against these proteins in RNase + / -conditions (FIG. 5C). Consistent with the AP-MS results, both proteins co-precipitated multiple CPA complex subunits in an RNA-independent manner. Notably, RNase treatment further strengthened the interactions, particularly for RNPS1, suggesting that portions of these interactions are partially RNA-shielded. This observation is consistent with a recent-56- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760study, which reported that many biologically relevant protein-protein interactions involved in pre-mRNA processing are RNA-shielded67.10204] As orthogonal validation, Applicant performed reciprocal IPs using antibodies against CPA subunits under the same + / - RNase conditions. Anti-CPSF7 and anti-FIPl IPs coprecipitated GRB2, with signal increasing after RNase, consistent with partially RNA-shielded contacts (FIGS. 5D & 5E). In contrast, endogenous RNPS1 was not detectable in these reciprocal IPs, which likely reflects low occupancy and / or low affinity due to stoichiometry of RNPS1-CPA interactions (FIGS. 5D & 11C). Supporting this interpretation, anti-FIPl IP robustly recovered FLAG-RNPS1 upon overexpression, again enhanced by RNase treatment (FIG. 5E). Together, these data indicate that GRB2 and RNPS1 associate with core CPA factors, and that RNPS1 may bind only a subset of CPA complexes at the endogenous expression level in a sub-stoichiometric, context-dependent manner.

[0205] GRB2 modulates PAS selection through a direct interaction with the CFIm via its SH3 domains

[0206] Applicant characterized the mechanism by which GRB2 and RNPS1 modulate PAS selection by interrogating the roles of their constituent domains. GRB2 consists of three structural domains: a central SH2 domain, which specifically binds to phosphorylated tyrosine residues, flanked by two SH3 domains that interact with proline-rich motifs.Applicant first investigated which protein domain(s) of GRB2 are required for its interaction with CPA factors. To do this, Applicant overexpressed FL AG-tagged full-length (FL) GRB2 or its truncated variants with specific domain deletions (AN-SH3, ASH2, and AC-SH3) followed by FLAG-IP in HEK293T cells (FIG. 6A). Applicant found that, similar to the FL protein, the ASH2 variant retained interactions with CPA factors CPSF1, CPSF6, CPSF7, and RBBP6. However, deletion of either the N-terminal or C-terminal SH3 domain completely abolished CPA factor binding, suggesting that both SH3 domains are required for GRB2's interactions with CPA factors. Applicant then compared the PAS selection activities of the FL and truncated variants fused to MCP using the same tethering assay with the upstream dual-luciferase reporter. These results showed that the AC-SH3 variant failed to activate PAS, whereas the AN-SH3 variant displayed significantly lower activity compared to the FL (FIG.6B) Consistent with these results, the fine-tuned ProteinBERT’s occlusion analysis predicted-57- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760both SH3 domains — particularly the C-terminal SH3 — as critical for PAS activation (FIG.6A). Collectively, these results demonstrate that the SH3 domains are crucial for GRB2's interactions with CPA factors and its role in activating PAS selection.

[0207] Although the AP-MS and co-IP assays detected interactions between GRB2 and the core CPA factors, it is unclear whether these interactions are direct. Therefore, Applicant sought to determine which CPA factor(s) may directly bind to GRB2. To this end, Applicant purified recombinant CPA subcomplexes (mPSF, mCF, CstF, CFIm, and CFIIm) and the N-terminal RBBP6 fragment (RBBP61-350). These factors are sufficient for reconstituting mRNA cleavage in vitro68,69. Next, Applicant mixed recombinant GST-GRB2 with these purified CPA factors and performed a GST pulldown assay. Applicant found that GST-GRB2 specifically bound to the CFIm complex, but not the mPSF, mCF, CstF, or CFIIm complexes (FIGS. 6D & 12A). The CFIm complex consists of a CPSF5 homodimer and two paralogous proteins, CPSF6 or CPSF7. The pulldown assay showed that GRB2 does not bind CPSF5 alone. In contrast, CPSF5-CPSF7 and CPSF5-CPSF6 complexes both bound GRB2, suggesting that GRB2 associates with the CFIm complex via CPSF7 or CPSF6, which contain proline-rich motifs recognized by GRB2 SH3 domains. Additionally, the N-terminal fragment of RBBP6, which also contains a proline-rich motif, bound to GRB2. However, this interaction was weaker than CPSF6 or CPSF7. These findings are consistent with a previous GRB2 AP-MS study70and with reports that GRB2 can be ubiquitinated by RBBP6 in the context of DNA damage repair45. Together, these results suggest that GRB2 directly engages CPSF6, CPSF7, and RBBP6, providing a molecular basis for its recruitment to the CPA machinery.

[0208] The CFIm complex is a position-dependent activator of PAS selection and plays an important role in APA regulation during development! 6,71. Since GRB2 directly interacts with the CFIm complex, Applicant asked if GRB2 knockdown-induced APA changes phenocopies CFIm depletion. Of the 215 significant APA changes upon GRB2 KD, 73% overlapped with those observed upon depletion of CFIm subunits (CPSF5 / 6 / 7); among these overlaps, 76% showed the same direction of change (shortening vs. lengthening) (FIGS. 6E< 12B & 12C). These results support a model in which GRB2 modulates APA, at least in part, by directly interacting with CFIm via its SH3 domains (FIG. 6L).-58- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[02O9| RNPS1 modulates PAS selection through a direct interaction with the mPSF and CPSF6 via its N-terminal IDR|0210] RNPS1 is composed of a structured RNA recognition motif (RRM) flanked by intrinsically disordered regions (IDRs) at the N- and C-termini, enriched with serine and arginine / serine repeats, respectively. To identify which regions are required for interacting with CPA factors, Applicant divided RNPS1 into three subregions: N-terminal (1-160 aa), RRM (161-240 aa), and C-terminal (241-305 aa). FLAG-GST fusion constructs for each subregion (GST included to stabilize / express IDRs) were transiently expressed in HEK293T cells and analyzed by FLAG co-IP. The N-terminal region alone was sufficient to associate with core CPA factors, such as RBBP6, CPSF1, and CPSF6, at levels comparable to full-length RNPS1 (FIG. 6F). In contrast, the RRM and C-terminal regions displayed little, if any, interaction with CPA factors, suggesting that the N-terminal IDR mediates RNPSl's interaction with CPA factors. To determine whether these contacts contribute to PAS activation, Applicant performed the tethering assay using the upstream reporter and RNPS1 subdomain variants fused to MCP. The N-terminal region alone was sufficient to activate PAS whereas neither the RRM nor the C-terminal region showed activation, mirroring the coIP results (FIG. 6G). Consistent with these results, occlusion analysis from the fine-tuned ProteinBERT predicted the N-terminal IDR, particularly its serine-rich tract, as a key contributor for PAS activation (FIG. 6F). The model also highlighted the RRM. Although the RRM alone neither bound CPA factors nor activated PAS in the tethering assay, it may still be required for PAS regulation in the endogenous (non-tethered) context.|021I] To identify which CPA subunit(s) bind RNPS1 directly, Applicant performed GST pull-down assays with purified CPA subcomplexes. Applicant used the same N-terminal fragment of RNPS1 (1-160 aa) fused to GST (GST-RNPSlNterm), as this region is sufficient for CPA association and PAS activation in the cellular assays (FIGS. 6F-6G). In contrast to GRB2, which directly interacted with the CFIm complex and RBBP6, RNPS1 binds the mPSF complex — the poly(A) signal recognition module — and CPSF6 (FIGS. 6H & 12D).These findings indicate that RNPS1 is a bona fide interactor of the CPA machinery and that GRB2 and RNPS1 engage the CPA complex through distinct mechanisms.-59- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[02121 RNPS1 was among the strongest activators in the screen, yet its knockdown led to relatively few APA changes (FIG. 4A). The co-IP data suggest that RNPS1-CPA contacts are sub-stoichiometric (FIGS. 5D & 11C), prompting us to test whether increased RNPS1 expression can modulate PAS selection. Interestingly, overexpressing full-length RNPS1 in HEK293T cells followed by PAS-seq analysis revealed 368 significant APA events, >95% of which reflected mRNA shortening (FIG. 61). eCLIP analysis indicated that proximal PASs with increased usage upon RNPS1 overexpression exhibited stronger upstream RNPS1 signal than unaffected genes (FIGS. 6J & 12E), consistent with the screening data which defined RNPS1 as an upstream activator (FIG. 2). Genes affected by RNPS1 overexpression were enriched for phosphorylation and RNA-processing functions, suggesting that RNPS1 -driven APA preferentially tunes signaling pathways and RNA-metabolism networks (FIG. 12F). Notably, overexpression of core EJC subunits (EIF4A3, RBM8A) produced few, if any, APA events, suggesting that the RNPS1 phenotype is largely EJC-independent (FIG. 12G).

[0213] mRNA shortening, via 3' UTR shortening or intronic polyadenylation, is a hallmark of cancer72 74. To investigate whether increased levels of RNPS1 are linked to cancer-associated shortening, Applicant analyzed the correlation between RNPS1 expression and the number of genes exhibiting mRNA shortening across seven cancer types (TCGA)75(FIG. 6K). RNPS1 showed a strong positive correlation (Spearman’s R=0.71), consistent with the idea that elevated RNPS1 may promote proximal PAS usage and contribute to mRNA shortening in tumors. Together with the mechanistic data, these results support a model in which elevated, upstream-positioned RNPS1 promotes proximal PAS usage, at least in part by stabilizing mPSF recognition of the poly(A) signal (FIG. 6L).

[0214] Experimental Discussion

[0215] In this work, Applicant deployed a dual-luciferase MS2 tethering screen across 879 RBPs and identify 63 high-confidence activators of PAS selection, most of which were not previously linked to APA. Applicant validated these activators with KD PAS-seq, KD RNA-seq, and eCLIP, and mechanistically dissected two unexpected hits, GRB2 and RNPS1. The fine-tuned ProteinBERT model predicts PAS activators from an unseen dataset and highlights important domains. Applicant also developed a programmable U7-MS2 system that can generalize RBP targeting beyond reporters. Together and without being bound by theory, it is-60- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760anticipated that these findings and methods will serve as a useful resource and framework for dissecting mechanisms of PAS selection and other RNA-mediated processes.10216] GRB2, a widely studied adapter protein, has been primarily studied for its cytoplasmic function as a mediator of cell signaling, such as linking RTK to the RAS / MAPK pathway in a variety of cellular contexts50. Though it can localize to the nucleus, its roles in mRNA processing were previously unknown. These results demonstrate a direct role for GRB2 in pre-mRNA 3' end processing, affecting over 200 genes. This function is mediated by its SH3 domains, at least partially through direct binding to the CFIm complex in the CPA machinery (FIG. 6). Notably, genes undergoing APA upon GRB2 perturbation were most significantly enriched for DNA damage response and related processes (FIG. 4). This suggests that GRB2-mediated APA may directly regulate this process, aligning with its known nuclear role in DNA damage response44,45’55. Whether this nuclear function of GRB2 is connected to its cell signaling function remains to be investigated. While mechanisms regulating APA have been primarily studied in the context of RBP binding, the role of cell signaling in mediating APA has rarely been explored and represents an intriguing avenue for future study.|0217] RNPS1, a peripheral EJC component and splicing factor with no established role in APA, emerges in the study as a regulator of PAS selection. Although most EJC and splicing functions map to the RRM and C-terminal region57’58’76, Applicant find that its N-terminal IDR directly interacts with the mPSF complex and CPSF6 (FIG. 6). Although the tethering screen classified RNPS1 as a strong activator, endogenous co-IPs suggest sub-stoichiometric engagement with CPA, with robust functional effects becoming evident when RNPS1 expression is elevated. This pattern mirrors SRRM1 (SRml60), the only other EJC-associated factor reported to affect 3' end processing, which likewise shows sub-stoichiometric CPA engagement and expression-dependent activity63. Applicant found that RNPS1 overexpression induces widespread APA dominated by mRNA shortening, and RNPS1 expression correlates with proximal PAS shifts across multiple cancer types.Therefore, it is plausible that elevated RNPS1 acts outside its canonical EJC / splicing roles to drive context-specific APA and contribute to dysregulated mRNA processing in pathological contexts.-61- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[0218J While Applicant did not delineate mechanisms for every RBP identified, these candidates represent promising entry points for future work on RNA processing and disease. Notably, the U7smOPT-MS2 recruitment system may be amenable to endogenous RNA targeting, suggesting that minimal effector fragments (e.g., GRB2 SH3 modules, RNPS1 N-terminal IDR) could be deployed to bias PAS choice at selected transcripts. In principle, this platform enables locus- and cell-type-specific activation or repression of pathogenic PASs. Collectively, our screen, models, and U7smOPT-MS2 platform set the stage for systematic mapping of proteins involved in PAS regulation and targeted APA intervention.

[0219] Table 6 - Key Resources Table-62- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-63- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-64- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-65- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-66- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760&4918-7893-0315.1Atty. Dkt. No.: 114198-3760

[0220] Experimental Model and Study Participant Details

[0221] HEK293T cells were purchased from ATCC (CRL-3216) and were not further authenticated. Cells were cultured with DMEM + 10% FBS (Gibco, Sigma). Cells were routinely evaluated for mycoplasma contamination using the MycoAlert mycoplasma test kit (Lonza) and were found negative for mycoplasma. EXOSC3-FKBPF36V-FLAG-HEK293T cell lines were generated following the procedure described by Soles et al., 202577.

[0222] HEK293T cells were purchased from ATCC (CRL-3216) and were not further authenticated. Cells were cultured with DMEM + 10% FBS (Gibco, Sigma). Cells were routinely evaluated for mycoplasma contamination using the MycoAlert mycoplasma test kit (Lonza) and were found negative for mycoplasma. EXOSC3-FKBPF36V-FLAG-HEK293T cell lines were generated following the procedure described by Soles et al., 202577.

[0223] Method Details

[0224] Construction of reporter constructs

[0225] pPASPORT bicistronic reporter16,88’89, containing the L3 PAS from the adenovirus major late transcript, was modified as follows: 1. Three MS2 hairpin sequences were inserted 15 bp upstream (upstream reporter) or 56 bp downstream (downstream reporter) of the AAUAAA hexamer within L3 poly(A) signal to enable tethering by MCP -fused RBPs. 2. UGUA motifs, which serve as binding sites for the CFIm complex, were deleted from the L3 PAS to weaken its strength, facilitating more accurate identification of PAS activators.

[0226] Dual-Luciferase Tethering Assay

[0227] Plate preparation, transfection, and read-out

[0228] The tethered function assay was completed using 96-well Solid Black Flat Bottom Polystyrene TC-treated Microplates (Corning, 3916). The plates were coated with 75uL of poly-D-lysine hydrobromide (PDL) (Sigma-Aldrich, P6407-5MG) dissolved in water at 1g L’1and diluted 1:5 in lx DPBS (Corning, 21-031-CV). The plates were dried overnight inside a tissue culture incubator. Prior to transfection, the plates were rinsed twice with IxDPBS. The expression plasmids of the RBP-MCP library used were obtained from previously conducted screens13,15. Expression plasmids and reporter plasmids were diluted to 50 ng / uL. A transfection master mix was prepared using a 1 : 1 mixture of upstream or downstream-68- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760reporter and RBP-MCP expression plasmid using the Lipofectamine 3000 and P3000 reagents (Invitrogen, L3000015) diluted in Opti-MEM Reduced Serum Media (Gibco, 31985062). The mixture was incubated for 15 minutes, after which it was distributed across the prepared 96-well PDL-coated plate. 75uL of HEK293T cells were then plated in each well of the same 96-well PDL-coated plate, at a concentration of 266,666 cells per milliliter and incubated in a tissue culture incubator for 48h. Each plate contained three replicates / triplicates (three separate wells assigned) for each expression plasmid being tested and three replicates of each control (FLAG-MCP, CPSF5-MCP, CPSF6-MCP, HNRNPCL1-MCP).[0229J The Dual-Glo Luciferase Assay System (Promega, E2980) and associated protocol was used to conduct Luminescence measurements. Following the 48h incubation, plates were removed from the tissue culture incubator and cooled to room temperature for 30 minutes. First, 75uL of Dual-Glo Luciferase Reagent was added to each well and mixed thoroughly using a Microplate Genie Plate Shaker (Scientific Industries) then centrifuged. The reaction was incubated at room temperature for 10 minutes. The luminescence measurement was obtained using a Teacan infinite 200Pro plate reader using automatic attenuation, 500 ms integration time, and 0ms settle time at room temperature. The process was repeated with Dual-Glo Stop & Gio Reagent.

[0230] Analysis of Luminescence Measurements[02311 The luminescence ratios of Renilla to Firefly were calculated for each expression vector and replicate. A t-test was employed using scipy vl .11 ,478to determine the statistical significance (P < 0.05) of the difference in effects seen between the RBP-MCP and FLAG-MCP control. The magnitude of the difference was calculated as the fold-change between the mean of luminescence ratios across replicates of RBP-MCP and FLAG-MCP(mean normalized). Luminescence ratios of RBP-MCPs that were less than those of plate-matched FLAG-MCP control were categorized as inhibitors and those greater than categorized as activators. Analysis of measurements for primary screen, secondary screen, and zinc-finger proteins followed the same approach. After secondary screening, recovered activator candidates (ones that showed significant ( / -test, P < 0.05) PAS activation in both rounds of screening) were identified individually for the upstream and downstream reporters.-69- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760To compare relative ranks of candidates across screens, candidates were ranked by the mean normalized value and assigned a rank. Lower rank values indicate stronger activators and higher rank values indicate weaker activators. Final high-confidence activator candidates were those displaying consistent reporter specificity across screens: upstream only, downstream only, or both upstream and downstream.[02321 RT-PCR RNA Validation

[0233] For RT-PCR validation, samples were prepared in accordance with the plate preparation and transfection methods described using standard 96-well tissue culture plates (Costar, 3596). Following 48h incubation, total RNA was isolated by lysing cells in TRIzol (Invitrogen, 15596018) and purified using the Direct-zol RNA Miniprep Kit (Zymo Research, R2052). For reverse transcription, total RNA was converted to cDNA using the All-In-One 5* RT MasterMix with gDNA Removal (abm, G592) according to the manufacturer’s instructions. The resulting cDNA was used for qPCR with PowerUp™ SYBR™ Green Master Mix (Applied Biosystems, A25742). Relative quantification between Renilla and Firefly ORF regions was determined by the 2A-AACt method using the following primers: Renilla, Forward: TAACGCGGCCTCTTCTTATTT, Renilla, Reverse:GATTTGCCTGATTTGCCCATAC; Firefly, Forward: CGGAAAGACGATGACGGAAA, Firefly, Reverse: CGGTACTTCGTCCACAAACA.

[0234] Domain and Functional Enrichment Analysis

[0235] Protein sequences of each RBP were obtained from the previously published RBP-MCP library13,15. Domains associated with each RBP ORF were called using hmmscan v3.1b290with parameters: — noali, — notextw. For each ORF, Applicant resolved overlapping or redundant HMM matches by retaining the entry with the lowest e-value. For each candidate category (upstream activator, downstream activator, upstream inhibitor, downstream inhibitor), Applicant constructed contingency tables for every domain comparing the number of RBPs ORFs within the category versus the background set (excluding proteins in that category) that contained the domain. Fisher’s exact test was then applied to assess statistical enrichment using scipy vl.11.478. Functional enrichment analysis of candidates in each category was completed using Decoupler vl.5.079, using the background set of screened RBPs.-70- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[ 02361 Interactors Analysis of Candidates

[0237] The physical interaction subnetwork vl2.0 was obtained from STRINGdb49. For each candidate, the number of interactions with members of the CPSF, CSTF, and CPEB family of proteins, in addition to RBBP6, PABPN1, and PABPC1 were counted. The statistical significance was computed with the Wilcoxon rank-sum test, using the scipy vl .11.478package.

[0238] MS2 U7 snRNA Recruitment

[0239] Cloning ofMS2 U7 snRNA Guides

[0240] Construction of MS2 variants in U7 smOPT backbone was done based on Smargon et al.31. Base U7 smOPT backbones (Table 7) were previously cloned into PUC19 backbone. MS2 hairpins were first cloned using Golden Gate Assembly91.Table 7<><>4918-7893-0315.1Atty. Dkt. No.: 114198-3760-72- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-73- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-74- 4918-7893-0315.1Atty. Dkt. No.: 114198-376010241 ] Oligos, containing MS2 variants in Bpil cloning sites for downstream guide cloning, were obtained from IDT, annealed using T4 Polynucleotide Kinase (NEB, M0201S), and -75- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760incubated at 37°C for 30 minutes then 95°C for 5 minutes. Base U7 smOPT plasmid was digested with Bpil (Thermo Fisher Scientific, FD1014) for 1 hour at 37 C and then purified using QIAquick PCR Purification kit (Qiagen, 28104). Annealed MS2 oligos were then ligated with digested U7 smOPT backbone using Quick Ligation (NEB, M2200S) and incubated at 25 °C for 10 minutes. Ligation products were then transformed, grown, isolated, sequenced, and verified as above. These were then cloned into MS2 U7 smOPT base plasmids using methods as described above.

[0242] Guides (Table 8) were then cloned into MS2 U7 smOPT base plasmids using methods as described above.[0243J Table 8-76- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760

[0244] Cloning no MS2 and no PAS APA Reporter

[0245] Reporters were cloned for snRNA-MS2 targeting to remove the MS2 hairpin and delete the PAS. The ablation of the PAS was done using an AATAAG, instead of the canonical AATAA, to ablate PAS selection92. PCR of the original upstream reporter (see Table 7, above) was performed using Q5 Hot Star High-Fidelity Master Mix (NEB, M0494S) to remove the MS2. The PCR product was isolated using agarose gel electrophoresis and purified using QIAquick Gel Extraction kit (Qiagen, 28704). Single strand Ultramer DNA oligonucleotides were obtained from IDT and used as the insert for Gibson cloning using HiFi DNA Assembly Master Mix (NEB, E2621). Gibson products were then transformed using DH5 Alpha Competent Cells (Biopioneer, GACC-96). Colonies were cultured in LB-Agar with Ampicillin antibiotic, and then miniprepped using QIAprep Spin Miniprep Kit (Qiagen, 27106). The sequence was confirmed using plasmid sequencing.

[0246] Transfection ofMS2 U7 snRNA Guides]0247| At 48h post-transfection, luminescence measurements were obtained using Dual-Glo Luciferase Assay System (Promega, E2920) according to the manufacturers protocol.-77- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760Luminescence was measured using a Teacan infinite 200Pro plate reader with the Costar 96 flat white setting, automatic attenuation, and integration times of 1000ms for Firefly and Renilla. Data was then analyzed in R31,93.

[0248] Model Training, Evaluation, and Application

[0249] Individual protein sequence files were created for each RBP using the RBP-MCP library information published previously 13, 15. Applicant used MMseqs2 vl7.b8O41*0easy-search to quantify the similarity between all protein sequences in a pairwise manner and output proteins with greater than 50% similarity using parameters: — min-seq-id 0.5. To avoid data leakage in training and evaluation, Applicant determined a similarity threshold by assessing the percent similarity distribution among protein isoforms or variants and different proteins and used the lower end of the values among the protein isoforms / variants distribution as the threshold. With this threshold, Applicant excluded protein pairs with greater than 80% similarity. After filtering, Applicant kept 90% for training and 10% for heldout / testing, ensuring that each contains similar class representations. Using the training data, Applicant conducted 10-fold cross validation on the ensemble model and each sequence component of HydRA with hydra vO.1.21.3832HydRa2_train_eval command and default parameters. Model selection was based on ROC-AUC and PR-AUC. Final fine-tuning, implementation of class weighting, and parameter tuning was completed using ProteinBERT vl .0.133and the previously described training data. After fine-tuning ProteinBERT with all training data and selected parameters, evaluation of the final model was conducted on the heldout dataset to evaluate ROC-AUC and PR-AUC of the final model and to determine the fine-tuned ProteinBERT Score threshold with the best Fl performance, for an Fl -optimized threshold to leverage for final classification. Protein sequences for the unseen ZFP dataset were obtained from UniProt IDs 94 and used as input to the final fine-tuned ProteinBERT model for classification as activators or non-activators. The Fl -optimized score threshold was used for classification.

[0250] Occlusion scores were generated for each sliding window across all proteins with HydRA vO.1.21.3832occlusion_map3 and default parameters using the final fine-tuned ProteinBERT model. To account for length effects, proteins were grouped into Fibonaccibased length bins. For each bin, Applicant calculated the mean and standard deviation of-78- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760window scores and assessed significance using the normal CDF implemented inscipy. stats. norm vl.11.478. For each sequence, low complexity regions (LCRs) were identified using Shannon Entropy Scoring, domains were identified as outlined previously with using hmmscan v3. Ib290, and RNA binding domains were assigned based on Gerstberger et al.1Significantly occluded regions were annotated by intersecting the corresponding regions with the regions harboring domains and LCRs using bedtools v2.27.186intersect with default parameters.

[0251] Knockdown RNA sample and library preparation and analysis

[0252] Preparation of knockdown RNA

[0253] Protein knockdown samples were generated using short targeting RNA (shRNA) plasmids from the MISSION lentiviral collection (Sigma-Aldrich). HEK293T cells were seeded onto 12-well plates at a density of 0. IxlO6cells per well in ImL of DMEM (LifeTech, 11995065) + 10% FBS and grown overnight. The media was replaced the following day using fresh media containing 8ug / mL polybrene (Hexadimethrine bromide - Sigma Aldrich cat# H9268) and lOOul of lentivirus specific to RBP open reading frames. The lentiviral medium was removed after 24h and replaced with fresh media containing 4ug / mL Puromycin and incubated in a standard tissue culture incubator for 48h. Total RNA extraction was completed using the Maxwell RSC simplyRNA tissue kit (Promega, AS 1340). Knockdown of each RBP was evaluated using qPCR, completed using the iScript cDNA Synthesis kit (Biorad, 1708891) and Phusion Hot Start Flex DNA polymerase (NEB, M0535).[02541 Bulk RNA-seq[0255[ RNA-seq libraries were prepared using the TruSeq Stranded mRNA library preparation kit (Illumina, 20020595) and the IDT for Illumina TruSeq RNA UD Index set (20040871) for samples having greater than 50% knockdown of the RBP. Samples were pooled in an equimolar fashion and 100 base-pair paired-end reads were generated following sequencing on the Illumina NovaSeq 6000 S4 flow cell. The reads were then aligned to hg3895. Following alignment, alternative splicing events were identified and quantified for each knockdown sample relative to non-targeting control using rMATS v4.3.087with parameters: — readLength 100. Principal component analysis (PCA) was used to ensure that samples did not exhibit batch effects.-79- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[0256| PAS-seq

[0257] 2 pg of total RNA was used for PAS-seq library preparation, following the protocol detailed in our methodological paper40. PAS-seq was aligned to hg3895and further analysis was conducted as described by Wang et al. (2021)96.

[0258] Downstream analysis of RBP knockdown data

[0259] Following analysis of bulk RNA-seq and PAS-seq for each RBP knockdown, genes having significant splicing changes and PAS changes were identified and a Jaccard index was calculated using these sets of genes. Using the STRINGdb physical interaction subnetwork49, the number of immediate CPA and splicing factor interactors were identified for each RBP. To identify clusters of RBPs, hierarchical clustering analysis with Euclidean distance metric was completed using the Jaccard index, number of CPA interactors, number of splicing factors for each RBP results. Functional enrichment analysis of targets undergoing APA following RBP knockdown was completed using Decoupler vl.5.079.|0260] Functional enrichment of direct RBP interactors

[0261] All immediate interactors, regardless of whether they were CPA or splicing factors, were identified using the STRINGdb physical interaction subnetwork49. Functional enrichment analysis was then completed on each RBP’s set of interactors using Decoupler vl.5.079, and filtered for statistically significant enrichment (P < 0.05). Applicant plotted the top ten terms for each RBP interactor set plotted using Seaborn97, coloring each point by the cluster it belongs to as previously identified and categorizing functional terms into broader functional groups for ease of interpretation.[02621 eCLIP Sample preparation and analysis

[0263] Preparation of eCLIP samples

[0264] eCLIP was performed in HEK293T cells in accordance with Yeo laboratory standard operating procedures98. Endogenous eCLIP was performed for CPSF6 with CPSF6 antibody (Bethyl, A301-356A), GRB2 with GRB2 antibody (Invitrogen, MA5-35238), CPSF5 with CPSF5 / NUDT21 antibody (Invitrogen, Catalog # 702871), EIF4B with EIF4B antibody (Bethyl, A301-766A), RBM10 with RBM10 antibody (Cell Signaling Technology, 47729S), and RBM22 with RBM22 antibody (Bethyl, A303-923A) . V5-tagged eCLIP was performed -80- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760for RNPS1, LGALS3, MBNL2, with V5 antibody (Bethyl, A190-120A) following overexpression of each protein by transfecting the protein expression plasmid in HEK293T cells as described previously. To generate cell pellets prior to eCLIP, a set of four confluent 10cm cell culture dishes (Thermo Fisher Scientific, FB012924) were prepared for each RBP. Each plate was washed with 3mL of DPBS and crosslinked using a UV cycle set to 400 mJ / cm2. Cells were harvested and centrifuged at 300xg for 3 minutes. Cells were then resuspended in DPBS, centrifuged again at 300xg for 3 minutes and flash frozen after aspirating media. Input samples were taken from a portion of the immunoprecipitation (IP) samples. eCLIP data for STAU2 and TRNAU1AP was obtained from Schmok et al.13and eCLIP data for RBM5, ZC3HAV1, and ZMAT3 was obtained from Gosztyla et al.36

[0265] Analysis of eCLIP data

[0266] eCLIP data was analyzed using Skipper81 with default parameters and sample-matched adapter sequences. Reads were mapped to hg3895. Reproducible enriched windows were filtered for those with enrichment_12or_mean > 3 and p min < 0.05. Metadensity analyses were conducted using the Metadensity package82. Motif enrichment analysis was completed using HOMER Motif enrichment analysis was completed using HOMER83fmdMotifsGenome.pl with parameters: — len 6,8 -rna. Modulated PAS sites for profiling binding were identified from the knockdown PAS-seq dataset. Enrichment of PAS binding around modulated PAS sites was determined using a Fisher’s exact test implemented with scipy vl .11 ,478after constructing contingency tables with the number of eCLIP binding sites ±200 bp around the unaffected and affected PAS sites.[02671 AP-MS data generation and analysis

[0268] HEK293T cells were transfected with expression plasmids containing either FLAG-V5-MCP, RNPS1-V5-MCP, or GRB2-V5-MCP in accordance with previously described methods. For each RBP, six samples were prepared: three replicates each for V5 antibody (Bethyl, A190-120A) and three replicates each for Rabbit IgG Isotype Control (Invitrogen, 02-6102). For each set of samples, cells were lysed in lysis buffer (150mMNaCl, 50mM Tris pH 7.5, 1% IGPAL-CA-630 Sigma #18896, 5% glycerol) and split evenly. Lysates were incubated on ice for 20 minutes, clarified, and total protein was quantified by BSA quantification. 5ug antibody (V5 antibody, Bethyl, A 190- 120 A; Rabbit IgG Isotype Control,-81- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760Invitrogen, 02-6102) was added to Img total protein lysate per replicate. The cell lysates with antibody were incubated with magnetic beads overnight at 4C. Supernatants were removed, beads were washed 2 times with wash buffer plus IGPAL (50 mM Tris pH 7.5, 150 mM NaCl, 5% glycerol, 0.05% IGPAL), and twice with wash buffer (50 mM Tris pH 7.5, 150 mM NaCl, 5% glycerol). After the last wash, the beads were resuspended in 80 pl trypsin buffer (2 M Urea, 50 mM Tris pH 7.5, ImM DTT, 5 pg / ml trypsin) to digest the bound proteins at 25 °C for 1 h at 1200 rpm. The supernatant was collected, and the beads were washed twice with 60 pl Urea buffer (2 M Urea, 50 mM Tris pH 7.5) and the washes were combined with the supernatant. The combined elution was cleared of residual beads by a quick spin. 80ul of the elution was used and disulfide bonds were reduced with 5 mM dithiothreitol (DTT), and cysteines were subsequently alkylated with 10 mM iodoacetamide. Samples were further digested by adding 0.5 pg sequencing grade modified trypsin (Promega) at 25°C. After 16 h of digestion, samples were acidified with 1% formic acid. Tryptic peptides were desalted on C18 StageTips according to Rappsilber et al.99 and evaporated to dryness in a vacuum concentrator and reconstituted in 15 pl of 3% acetonitrile / 2% formic acid for LC-MS / MS.

[0269] LC-MS / MS analysis was performed on a Q-Exactive HF. 5uL of total peptides were analyzed on a Waters M-Class UPLC using a 15cm Ion-Optics column (1.7um, C18, 75um x 15cm) coupled to a benchtop Thermo Fisher Scientific Orbitrap Q Exactive HF mass spectrometer. Peptides were separated at a flow rate of 400 nL / min with a 90 min gradient, including sample loading and column equilibration times. Data was acquired in data-dependent mode. MSI spectra were measured with a resolution of 120,000, an AGC target of 3e6 and a mass range from 300 to 1800 m / z. MS2 spectra were measured with a resolution of 15,000, an AGC target of le5 and a mass range from 200 to 2000 m / z. MS2 isolation windows of 1.6 m / z were measured with a normalized collision energy of 25.

[0270] Proteomics raw data was analyzed by MaxQuant v2.0.3.084using a UniProt database (Homo sapiens, UP000005640), and MS / MS searches were performed under default settings with LFQ quantification. Data were further analyzed in R v3.6.3. Contaminants, and proteins only identified by site or reverse were removed, a pseudo count randomly taken from the bottom of the signal distributions was added to the LFQ intensity values and then the LFQ intensity values were log2 transformed. Proteins with a mean MS / MS count value for each IP -82- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760condition below 5 were removed from subsequent analysis. Interacting proteins were identified as those that passed a log2(Fold Change) sample IP over control IP cutoff of 1 and a E- value of 0.05, and any proteins that were enriched in FLAG control IPs were removed from the interacting proteins list. Functional group assignment was completed using Metascape85.[02711 Domain-specific RBP tethering assay using pPASPORT.

[0272] HEK293T cells were seeded into 96-well cell culture plates (Thermo Fisher Scientific, 152036) and cultured until approximately 80% confluency. Upstream 3xMS2 pPASPORT reporter and MCP-fused RBP domains (or domain deletion constructs) were cotransfected at a 1:3 ratio using Lipofectamine 3000 (Invitrogen, L3000015), following the manufacturer's instructions. After 48h, Firefly and Renilla luciferase activities were measured using the Dual -Luciferase® Reporter Assay System (Promega, El 910), according to the manufacturer’s protocol.

[0273] Co-immunoprecipitation and Western blot analysis.

[0274] FLA G-immunoprecipitation

[0275] HEK293T cells were seeded into 6-well cell culture plates (Genesee Scientific, 25-105) and cultured until they reached approximately 80% confluency. Applicant then transfected pcDNA3 mammalian expression vector encoding the N-terminal FLAG-tagged RBP of interest using Lipofectamine 3000 (Invitrogen, L3000015), following the manufacturer's protocol. After 48 hours, cells were washed with ice-cold PBS and harvested. The collected cells were briefly centrifuged and resuspended in a lysis buffer containing 10 mM HEPES-OH (pH 7.9), 150 mM NaCl, 1 mM MgCh, 0.5% NP-40, 1 mM PMSF, and a lx protease inhibitor cocktail (Thermo Fisher Scientific, PI78429). Cell lysis was performed by sonication, and the lysates were clarified by centrifugation at 16,000 x g for 30 minutes. The supernatant was collected and incubated with ANTI-FLAG M2 Affinity Gel (Millipore, A2220) at 4°C for 2 hours with gentle rotation. After binding, the affinity gel was recovered by centrifugation and washed three times with the lysis buffer. Elution of immunoprecipitated proteins was performed using the lysis buffer supplemented with 250 pg / mL FLAG peptide (Sigma-Aldrich, F3290). The eluates were precipitated with acetone overnight at -20°C. The-83- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760resulting protein pellets were resuspended in 1 * SDS loading buffer and prepared for Western blot analysis.

[0276] Endogenous Immunoprecipitation

[0277] HEK293T cells were seeded into 10-cm dishes (Falcon, 08772E) and grown to -90% confluency. Cells were lysed as described above. For each immunoprecipitation, 20 pL Dynabeads Protein A (Invitrogen, 10002D) was incubated with 2 pg of antibody (GRB2, RNPS1, FIP1, or CPSF7) for 1 h at 4 °C on a rotator, then mixed with cleared cell lysate and incubated overnight at 4 °C. For RNase treatment, 2 pL RNase A / Tl (Thermo Scientific, FEREN0551) was added during the overnight incubation. Beads were collected on a magnetic rack, washed three times with lysis buffer, and bound proteins were eluted in 1 * SDS sample buffer.

[0278] For Western blotting, protein samples were resolved on 4-20% Mini-PROTEAN® TGX Stain-Free™ Protein Gels (Bio-Rad, 4568096) and transferred onto 0.45 pm nitrocellulose membranes (Bio-Rad, 1620115) using the eBlot LI wet-transfer system (GenScript, L00686) with the standard program. Following transfer, membranes were blocked with 5% milk in PBST (0.1% Tween-20 in PBS) for 30 minutes at room temperature. Membranes were incubated overnight at 4°C with the appropriate primary antibody diluted in 5% milk in PBST. The following day, membranes were washed with PBST and incubated with an HRP-conjugated secondary antibody in 5% milk in PBST for 45 minutes at room temperature. After additional washes with PBST, chemiluminescent detection was performed using Radiance Q Chemiluminescent Substrate (Azure Biosystems, 10147-296). Images were acquired using the Bio-Rad ChemiDoc MP imaging system.

[0279] Protein purification

[0280] Cloning

[0281] Following the MacroBac protocol100, MacroBac438A (Addgene #55218) vectors expressing the following CPA subcomplexes were constructed: mPSF (FIP1, Strep-CPSF4, 6xHis-WDR331-572, CPSF1), mCF (CPSF2, CPSF3, Strep-SYMPK), CstF (6xHis-CstF77, CstF64, CstF50), and CFIIm (6xHis-PCFl I769-1555, 2xStrep-CLPl). For the CFIm complex, CPSF5, CPSF6, and CPSF7 were cloned into a modified 6xHis-TEV-SBP N-terminal tagged-84- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760MacroBac438A vector. Dual expression vectors containing 6*His-TEV-SBP-CPSF5 with CPSF6 and 6><His-TEV-SBP-CPSF5 with CPSF7 were cloned using the same protocol.MacroBac and pFastbac plasmids were transformed into DHIOBac cells, and recombinant bacmids for each construct were subsequently purified. For GST-RNPSlNterm, The N-terminal fragment of RNPS1 (1-160 aa) was cloned into the pGEX-2T vector and was transformed into BL21 E. coli cells.

[0282] Expression

[0283] Bacmids were transfected into Sf9 cells using Cellfectin II reagent (Gibco, 10362100). The virus harvested from transfection was designated Pl. Pl virus was amplified to generate P2, which was used to infect Sf9 cells. Cells were harvested three days after infection by centrifugation at l,000xg for 20 minutes. Pellets were flash frozen and stored at -80°C. For GST-RNPSlNterm, colonies were grown overnight (~12 h) in 25 mL LB supplemented with ampicillin. The following day, 5 mL of the overnight culture was inoculated into two 500 mL LB cultures (no antibiotic) and grown at 37 °C with shaking until OD600 ~ 0.6. Protein expression was induced with 1 mM IPTG for 2 h at 37 °C. Cells were harvested by centrifugation and the pellets were flash-frozen for subsequent purification.

[0284] Purification

[0285] CFIm: Frozen pellets were thawed and resuspended in Lysis Buffer (50 mM HEPES pH 8.0, 400 mM NaCl, 20 mM imidazole, 0.5 mM TCEP, 10% Glycerol) supplemented with Halt Protease Inhibitor Cocktail (Thermo Scientific, 87785). For CPSF5 & CPSF6 / 7, the buffer also contained 2 mM MgCL and Benzonase (Millipore Sigma, 101654). Cells were lysed by sonication, and for CPSF5 & CPSF6 / 7, samples were incubated for 30 min at 4°C before clarification. The lysate was clarified by centrifugation at 15,000*g (30 min for CPSF5, 45 min for CPSF5 & CPSF6 / 7). For CPSF5, the lysate was additionally filtered through a 1.1 pm filter. The clarified lysate was loaded onto a 1 mL HisTrap HP column (Cytiva, 29051021), washed, and eluted using a gradient up to 250 mM imidazole. Protein was concentrated and treated with TEV protease while being dialyzed into Lysis Buffer. The sample was directly injected into a Superose6 Increase 10 / 300 GL (Cytiva, 29091596) for size-exclusion chromatography (SEC). For CPSF5 & CPSF6 / 7, the TEV-cleaved sample was further dialyzed into SP Loading Buffer (50 mM Bicine pH 9.0, 150 mM NaCl, 10%-85- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760glycerol) for 3 hours, then injected onto Capto HiRes Q 5 / 50 (Cytiva, 29275878) and eluted using a gradient with SP Loading Buffer containing 1 M NaCl. Fractions were collected, concentrated, flash frozen, and stored at -80°C.

[0286] mCF: Frozen pellets were thawed and resuspended in Buffer A300 (50 mM HEPES pH 8.0, 300 mM NaCl, 2 mM MgC12, 10% glycerol) supplemented with Halt Protease Inhibitor Cocktail (Thermo Scientific, 87785) andBenzonase (Millipore Sigma, 101654). Cells were lysed by sonication and incubated at 4 °C for 30 mins. Lysate was clarified by centrifugation at 4 °C for 45 mins at 15,000xg and supernatant was filtered through a 1.1 pm filter. Supernatant was loaded onto a StrepTrap XT column (Cytiva, 29401320), column was washed with 10 CV of Buffer A75 (50 mM HEPES pH 8.0, 75 mM NaCl, 10% glycerol), and eluted with Buffer A75 supplemented with 50 mM biotin. Sample was concentrated and injected onto Capto HiRes Q 5 / 50 (Cytiva, 29275878) and eluted using a 40 CV gradient with Q Elution Buffer (50 mM HEPES pH 8.0, 1 M NaCl, 10% glycerol). Protein was concentrated and injected into a Superose6 Increase 10 / 300 GL (Cytiva, 29091596) for sizeexclusion chromatography. Fractions were collected, concentrated, flash frozen on dry ice, and stored at -80°C.

[0287] mPSF: Frozen pellets were thawed and resuspended in Buffer A300 (50 mM HEPES pH 8.0, 300 mM NaCl, 20 mM imidazole, 2 mM MgC12, 10% glycerol) supplemented with Halt Protease Inhibitor Cocktail (Thermo Scientific, 87785) and Benzonase (Millipore Sigma, 101654). Cells were lysed by sonication and incubated at 4 °C for 30 mins. Lysate was clarified by centrifugation at 4 °C for 45 mins at 15,000xg and supernatant was filtered through a 1.1 pm filter. Supernatant was loaded onto a StrepTrap XT column (Cytiva, 29401320), column was washed with 10 CV of Buffer A300, and eluted with Buffer A300+50 mM biotin onto a HisTrapTM HP column (Cytiva, 17524802). HisTrapTM HP column was eluted with Buffer Bl (50 mM HEPES pH 8.0, 75 mM NaCl, 250 mM imidazole, 10% glycerol) using a 10 CV gradient. Protein was concentrated and injected into a Superose6 Increase 10 / 300 GL (Cytiva, 29091596) for size-exclusion chromatography. Fractions were collected, concentrated, flash frozen on dry ice, and stored at -80°C.

[0288] CFIIm: CFII was purified using the same protocol with mPSF.-86- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[0289| CstF: Frozen pellets were thawed and resuspended in Buffer A300 (50 mM HEPES pH 8.0, 300 mM NaCl, 20 mM imidazole, 2 mM MgC12, 10% glycerol) supplemented with Halt Protease Inhibitor Cocktail (Thermo Scientific, 87785) and Benzonase (Millipore Sigma, 101654). Cells were lysed by sonication and incubated at 4 °C for 30 mins. Lysate was clarified by centrifugation at 4 °C for 45 mins at 15,000xg and supernatant was filtered through a 1.1 pm filter. Supernatant was loaded onto a HisTrapTM HP column (Cytiva, 17524802), column was washed with 10 CV of Buffer A75 (50 mM HEPES pH 8.0, 75 mM NaCl, 20 mM imidazole, 10% glycerol), and eluted with His Elution Buffer (50 mM HEPES pH 8.0, 75 mMNaCl, 250 mM imidazole, 10% glycerol). Sample was concentrated and injected onto Capto HiRes Q 5 / 50 (Cytiva, 29275878) and eluted using a 40 CV gradient with Q Elution Buffer (50 mM HEPES pH 8.0, 1 M NaCl, 10% glycerol). Protein was concentrated and injected into a Superose6 Increase 10 / 300 GL (Cytiva, 29091596) for sizeexclusion chromatography. Fractions were collected, concentrated, flash frozen on dry ice, and stored at -80°C.[0290| RBBP61-350aa: Cloning and protein purification of RBBP61-350aa was performed as previously described68.

[0291] GST-RNPSlNterm: Frozen pellets were thawed and resuspended in Buffer (20 mM HEPES pH 8.0, 250 mM NaCl, ImM DTT, ImM PMSF) supplemented with Halt Protease Inhibitor Cocktail (Thermo Scientific, 87785). Lysozyme was added to 1 mg / mL and the suspension was incubated on ice for 30 min prior to sonication. Cells were lysed by sonication. Lysates were clarified by centrifugation at 15,000 x g for 30 min at 4 °C, and the supernatant was passed through a 1.1 pm filter. The cleared lysate was incubated with 2 mL Glutathione Sepharose 4B resin (Cytiva, 17075601) for 2 h at 4 °C with end-over-end rotation. The resin was collected in a gravity -flow column and washed five times with 1 column volume (CV) of lysis buffer per wash. Bound GST-tagged protein was eluted with lysis buffer supplemented with 10 mM reduced glutathione, concentrated using an Amicon Ultra centrifugal filter (Millipore, UFC501008), and subjected to size-exclusion chromatography on a Superdex 200 Increase SEC column, 10 / 300 GL (Cytiva, 28990944). Fractions were collected, concentrated, flash frozen on dry ice, and stored at -80°C.[0292 | GST pull-down-87- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[02931 For the GST pull-down assay using purified recombinant proteins, 50 pL of Glutathione Sepharose 4B beads (Cytiva, 17075601) were washed three times with GST pulldown buffer (20 mM HEPES, pH 8.0, 150 mM NaCl, 1 mM MgCh, 0.1% NP-40) and resuspended in 500 pL of the same buffer. Recombinant GST (Rockland Immunochemicals Inc., 000-001-200), GRB2-GST (Rockland Immunochemicals Inc., 009-001-R99S), GST-RNPS1 Nterm was added to a final concentration of 40 nM, and the mixture was incubated with rotation for 1 hour at room temperature to allow protein binding to the beads. The beads were then washed three times with the buffer, resuspended in 500 pL of buffer supplemented with 5% BSA, and incubated overnight at 4°C with rotation for blocking. The following day, the beads were washed twice with the buffer, and equimolar amounts (final concentration 40 nM) of purified CPA subcomplexes and RBBP6 NTD1 -350 were added. The mixture was incubated with rotation for 2.5 hours at 4°C, followed by three washes with the buffer. The beads were then resuspended in 1 * SDS loading dye for subsequent Western blot analysis.

[0294] Preparation of RNPS1 / EIF4A3 / RBM8A overexpression samples

[0295] HEK293T cells were seeded into 6-well cell culture plates (Genesee Scientific, 25-105) and cultured until they reached 80% confluency. The cells were then transfected with the pcDNA3 mammalian expression vector, either empty (control) or containing FLAG-RNPS1, FLAG-EIF4A3, and FLAG-RBM8A using Lipofectamine 3000 (Invitrogen, L3000015). After 48 hours, cells were collected using trypsin. Total RNA was isolated using TRIzol (Invitrogen, 15596018). 2 pg of total RNA was used for PAS-seq library preparation.

[0296] TCGA analysis

[0297] The number of genes exhibiting transcript shortening and lengthening events across seven tumor types (bladder urothelial carcinoma, head and neck squamous cell carcinoma, lung adenocarcinoma, BRCA, kidney renal clear cell carcinoma, and uterine corpus endometrioid carcinoma) were obtained from a previous publication75. RNA expression data for each cancer type was obtained from the Genomic Data Commons (GDC) The Cancer Genome Atlas (TCGA) repository data release v41.0. Spearman’s R correlation was computed with the scipy package78 for RNPS1 using the expression of RNPS1 and the overall number of shortening events observed across all seven cancer types.

[0298] Experimental Summary-88- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760

[0299] Applicant’s disclosure provides a unique solution to the lack of a precise, programmable, and non-destructive method for modulating polyadenylation site (PAS) selection within target mRNA transcripts. The regulation of alternative polyadenylation (APA) is known to impact mRNA stability, translation, and localization, and it is often intimately tied to cell type, developmental stage, and disease state, including cancer and haploinsufficiency disorders. Current technologies for probing or reprogramming PAS usage, such as CRISPR-mediated PAS mutation, genome-wide knockdowns of RNA-binding proteins (RBPs), or use of generic reporter constructs, either create permanent DNA alterations, lack positional specificity relative to the PAS, or cannot isolate effects attributable to a specific PAS without perturbing global RNA processing networks. These limitations hinder mechanistic understanding of individual RBPs in PAS choice and obstruct the development of targeted interventions for transcriptome engineering or therapy.

[0300] In one aspect, Applicant provides a genetically encoded, nuclear-expressed antisense RNA that includes both a PAS-proximal antisense guide sequence and an aptamer that binds a PAS-modulating RBP with high affinity to address the limitations of the art. In use, the antisense portion hybridizes specifically to a defined region upstream or downstream of a target PAS, positioning the aptamer at an optimal distance to recruit the desired RBP either via direct binding or through a tethering system. This recruitment directs the RBP to act on the local mRNA 3' end processing machinery to either activate or inhibit cleavage and polyadenylation at that PAS. The system avoids genome editing, preserves endogenous mRNA sequence, and can be adapted for use with endogenous RBPs or exogenous RBP fusion proteins. Supporting elements include vectors to deliver antisense RNA and RBP coding sequences, promoter options for nuclear expression, and computational models trained to identify novel RBPs suitable for PAS modulation. The approach is validated experimentally using dual-luciferase MS2 tethering assays, knockdown PAS-seq and RNA-seq, biochemical interaction assays such as AP-MS and co-IP, and programmable U7smOPT-MS2 targeting to test positional impact.

[0301] Applicant’s disclosure enables the modularity and orthogonality of the disclosed components and methods. The antisense RNA design principles demonstrated in the examples, comprising an aptamer domain, an antisense guide domain, and optional stability / recruiting motifs, are not inherently limited to the MS2 aptamer, the U7 promoter, or -89- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760the particular RBPs CPSF5, CPSF6, GRB2, or RNPS1 exemplified in the experimental section. The functional mechanism relies on generalizable aptamer-RBP binding, which can be achieved with a variety of aptamer sequences derived from bacteriophages, viral proteins, or endogenous mammalian motifs, and with diverse RBPs that act on PAS usage either by RNA binding or through interactions with cleavage and polyadenylation (CPA) subcomplexes. Likewise, the PAS-proximal targeting principle is independent of the specific positioning coordinates used in the luciferase reporters; the data illustrate that certain distances (e.g., 60 nucleotides upstream) yield maximal activation for some RBPs, but the system inherently permits systematic repositioning for other RBPs or mRNA contexts. The host cell and vector embodiments are supported by standard molecular cloning and transfection methods, which are applicable across species and vector backbones. The computational prediction framework trained on tethered-function assay data is applicable to any RBP sequence input, regardless of family, because the features learned by the model are not restricted to the training species but capture conserved sequence-function relationships.[03021 Because the disclosed antisense RNA structure, aptamer-RBP recruitment architecture, PAS-proximal targeting logic, and validation assays are modular and platformagnostic, the skilled person can extend the exemplified constructs to other promoters, aptamer / RBP pairs, target PAS locations, host cell types, and disease-related transcripts without inventive step. The general principle of hybridizing an antisense guide near a PAS to deliver a selected PAS-modulating effector is transferable to any RNA context where the effector’s activity on PAS usage is advantageous, making the disclosed species an enabling foundation for broader claims covering a wide spectrum of antisense designs, RBP choices, targeting positions, vectors, cells, compositions, and prediction systems within the overall solution to the stated problem.

[0303] Tables

[0304] Table 3. Sequences of RNA hairpins and RNA-binding moi eties.

[0305] Hairpin sequences reproduced from Bos, T. J ., et al. (2016) Advances in experimental medicine and biology, 907, 61-88. https: / / doi.org / 10.1007 / 978-3-319-29073-7__3, the entire contents of which are incorporated herein.-90- 4918-7893-0315.1Atty. Dkt. No.: 114198-37604918-7893-0315.1Atty. Dkt. No.: 114198-37604918-7893-0315.1Atty. Dkt. No.: 114198-37604918-7893-0315.1Atty. Dkt. No.: 114198-3760-94- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-95- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-96- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760

[0306] Table 4 RBP sequence information. Included in the table are protein name, NCBI accession number, nucleotide sequence, and amino acid sequence.-97- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-98- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-99- 4918-7893-0315.1Atty. Dkt. No.: 114198-37604918-7893-0315.1Atty. Dkt. No.: 114198-3760-101- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-102- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-103- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-104- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-105- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-106- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-107- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-108- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-109- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-110- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760<<<-Ill- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-112- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-113- 4918-7893-0315.1Atty. Dkt. No.: 114198-37604918-7893-0315.1Atty. Dkt. No.: 114198-3760-115- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-116- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-117- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-118- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-119- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-120- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-121- 4918-7893-0315.1Atty. Dkt. No.: 114198-37604918-7893-0315.1Atty. Dkt. No.: 114198-37604918-7893-0315.1Atty. Dkt. No.: 114198-3760-124- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-125- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-126- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-127- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-128- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760<<<-129- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-130- 4918-7893-0315.1Atty. Dkt. No.: 114198-37604918-7893-0315.1Atty. Dkt. No.: 114198-3760-132- 4918-7893-0315.1Atty. Dkt. No.: 114198-37604918-7893-0315.1Atty. Dkt. No.: 114198-3760-134- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-135- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-136- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-137- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-138- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-139- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-140- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-141- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760-142- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760

[0307] Table 5 Antisense RNA sequence information. Features in the vector (promoter, hairpin, antisense guide to target gene of interest, snRNA backbone) can be found in FIG.13).-143- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760

[0308] Table 6. snRNA Targeting Guides for use in the antisense RNA.-144- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[0309| Polyadenylation (APA) is a widespread post-transcriptional mechanism that influences mRNA isoform diversity, stability, localization, and translation. Dysregulation of APA contributes to various diseases, including cancer and haploinsufficiency disorders. While some APA regulators have been identified, existing experimental technologies to study or therapeutically manipulate polyadenylation site (PAS) usage have significant limitations. Reporter assays can measure PAS strength but do not inform on PAS usage in the native transcriptome. CRISPR / Cas9-based PAS disruption creates double-stranded DNA breaks and removes PASs entirely, potentially causing off-target genomic alterations. Knockout or knockdown of RNA-binding proteins (RBPs) affect all sites regulated by that RBP, making it difficult to study PAS-specific functions. In addition, existing tethering systems lack a framework for modular, locus-specific action on endogenous PASs without global effects. Thus, there is a need for a genetically encoded, modular, and programmable RNA-guided system to selectively recruit APA-modulating factors to specific PASs in living cells without altering genomic DNA.[031(>| The present disclosure also provides antisense RNAs that base-pair with sequences proximal to a target PAS and include an aptamer region capable of high-affinity binding to an APA-modulating RBP or an RNA-binding domain fused to such an RBP. When expressed from a DNA vector in a host cell, the antisense RNA guides the bound RBP to the target PAS to modulate its usage, either activation or inhibition, without cutting DNA. The disclosure offers programmable locus-specific PAS recruitment via antisense complementarity, a modular design allowing different aptamer-RBP pairs to be inserted, the capacity to target either endogenous or reporter PASs, avoidance of permanent genomic changes, and the ability to use either RNA polymerase III or RNA polymerase II promoters for nuclear expression, enabling flexibility across experimental and therapeutic contexts.-145- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[03111 The written description, examples, and associated figures enable a person skilled in the art to make and use the disclosure with more than the specifically exemplified aptamers (e.g., MS2) and RBPs (e.g., GRB2, RNPS1, CPSF5, CPSF6). The specification provides general sequence requirements for aptamers, antisense regions, and RBP / RNP recruiting motifs, including length ranges, positions relative to PAS, and examples from multiple systems such as MS2, PP7, , QP, and IRE. The disclosure teaches design rules such as positioning antisense regions upstream or downstream of PAS, retaining functional aptamer hairpin structures, and using aptamers recognizable by endogenous or exogenous RBPs. The disclosure describes multiple vector systems and promoter types, including Ul, U6, and U7 for RNA polymerase III transcription and CMV or EFla for RNA polymerase II transcription, making it clear that various combinations of promoters and aptamer-RBP modules are operable.

[0312] Clauses

[0313] Clause 1. An antisense RNA comprising, from the 5' end to the 3' end: a) a nucleotide sequence encoding an aptamer, wherein the aptamer has a high affinity to an alternative polyadenylation (APA)-modulating RBP; and b) a nucleotide sequence antisense to a target region proximal to a polyA site (PAS).

[0314] Clause 2. The antisense RNA of clause 1, further comprising a nucleotide sequence comprising an RBP / RNP (ribonucleoprotein)-recruiting motif.

[0315] Clause 3. The antisense RNA of clause 1 or 2, wherein the aptamer is derived from a bacteriophage selected from: MS2, R17, , PP7, QP or GA; iron responsive protein (IRP); bovine immunodeficiency virus (BIV), or human Ul small nuclear ribonucleoprotein A.

[0316] Clause 4. The antisense RNA of clause 3, wherein the aptamer is derived from MS2.

[0317] Clause 5. The antisense RNA of clause 1, wherein: the RNA aptamer comprises a sequence consisting of SEQ ID NO: 1 (MS2 hairpin), SEQ ID NO: 2 (MS2 high affinity mutant), or an equivalent thereof having at least 90% sequence identity thereto; the aptamer is configured to recruit an RNA-binding protein selected from the group consisting of CPSF5 (SEQ ID NO: ), CPSF6 (SEQ ID NO: ), RNPS1 (SEQ ID NO: ), and GRB2 (SEQ ID NO: ); the antisense RNA hybridizes to a target region proximal to a polyadenylation site, wherein-146- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760"proximal" comprises a region upstream or downstream of the cleavage site sufficient to modulate PAS selection, without limitation to a fixed nucleotide distance; and optionally wherein the method is used for regulating PAS selection in an mRNA of a gene associated with a disease involving aberrant alternative polyadenylation, including cancers, haploinsufficiency disorders, and other APA-related pathological conditions.

[0318] Clause 6. The antisense RNA of any one of clauses 1-4, wherein the target region proximal to the PAS is approximately 20bp, approximately 40bp, approximately 60bp, approximately 80bp, approximately lOObp, approximately 120bp, approximately 140bp, approximately 160bp, approximately 180bp, or approximately 200bp upstream or downstream of the PAS.[0319| Clause 7. A vector comprising the antisense RNA of any one of clauses 1-6.

[0320] Clause 8. The vector of clause 7, wherein the vector is selected from a plasmid, a viral vector, a cosmid, or a phage, optionally wherein the viral vector is selected from a baculovirus, a retrovirus, or an adenovirus.

[0321] Clause 9. The vector of clause 7 or 8, further comprising a nucleotide sequence encoding a Polymerase II or Polymerase III promoter.[032 1 Clause 10. The vector of clause 9, wherein the Polymerase III promoter is selected from a Ul, U6, or U7 promoter.

[0323] Clause 11. The vector of any of clauses 7-10, wherein the antisense RNA is operably linked to a nucleotide sequence encoding an RBP, and wherein the RBP is a modulator of PAS selection.

[0324] Clause 12. The vector of clause 11, wherein the RBP is mammalian.

[0325] Clause 13. The vector of clause 11 or 12, wherein the RBP activates PAS selection, and wherein the RBP is selected from any of the RBPs in FIG. 2E.

[0326] Clause 14. The vector of clause 10 or 11, wherein the RBP is selected from CPSF5, CPSF6, RNPS1, and GRB2.-147- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760

[0327] Clause 15. An isolated host cell comprising one or more of the antisense RNA of any one of clauses 1-7 or the vector of any one of claims claim 8-13, wherein the host cell is a prokaryotic or a eukaryotic cell.

[0328] Clause 16. The isolated host cell of clause 14, further comprising a polynucleotide encoding the RBP or the RBP, optionally wherein the RBP is selected from any one of the RBPs identified in FIGS. 2A, 2E or Table 4.

[0329] Clause 17. The isolated host cell of clause 15, wherein the RBP is selected from CPSF5, CPSF6, RNPS1, and GRB2.

[0330] Clause 18. A composition comprising a pharmaceutically acceptable carrier and one or more of the antisense RNA of any one of clauses 1-6, the vector of any one of claims 7-14, and / or the isolated host cell of any of claims 15-17.[03311 Clause 19. A method of regulating polyA site (PAS) selection on mRNA in a cell, the method comprising: a) contacting the cell with the antisense RNA of any one of clauses 1-6 or the vector of any of claims 7-10, and b) contacting the cell with an RBP, or polynucleotide encoding an RBP, optionally wherein the RBP is selected from any one of the RBPs identified in FIG. 2A, wherein the antisense RNA binds to the target region proximal to the PAS, and wherein the RBP activates or inhibits PAS selection of the mRNA in the cell.

[0332] Clause 20. A method of regulating polyA site (PAS) selection on mRNA in a cell, the method comprising contacting the cell with the vector of any of clauses 11-14, wherein the antisense RNA binds to the target region proximal to the PAS, and wherein the RBP activates or inhibits PAS selection of the mRNA in the cell.[0333 [ Clause 21. A method of regulating polyA site (PAS) selection on mRNA in a cell, the method comprising contacting the cell with the antisense RNA of any one of clauses 1-6 or the vector of any of claims 7-10, wherein an endogenous RBP binds to the aptamer, and wherein the antisense RNA binds to a target region proximal to the PAS, and wherein the RBP activates or inhibits PAS selection of the mRNA in the cell.

[0334] Clause 22. The method of any of clauses 18-20, wherein the target mRNA is a gene implicated in haploinsufficiency diseases.-148- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760[0335[ Clause 23. The method of clause 22, wherein the haploinsufficiency disease is Dravet syndrome, autosomal dominant retinitis pigmentosa, or neurofibromatosis type 1.

[0336] Clause 24. The antisense RNA of Table 5.

[0337] Clause 25. A method of screening for RBPs that modulate PAS selection, comprising: a) providing a dual-luciferase bicistronic reporter comprising a first luciferase, a poly(A) signal (PAS) of interest, and a second luciferase; b) incorporating RNA hairpin aptamers upstream or downstream of the PAS for RBP tethering via a coat protein fusion; c) expressing in cells candidate RBP-coat protein fusions; and d) determining a luminescence ratio of the first luciferase to the second luciferase, wherein a change relative to a control indicates the RBP modulates PAS selection.[0338| Clause 26. The method of clause 25, wherein the reporter is an upstream reporter with the aptamer located about 15 nt upstream of the PAS, or a downstream reporter with the aptamer located about 56 nt downstream of the PAS.

[0339] Clause 27. The method of clause 25 or 26, wherein the PAS is an L3 adenovirus major late transcript PAS lacking UGUA motifs.

[0340] Clause 28. The method of any one of clauses 25-27, wherein the coat protein is MS2 coat protein (MCP).[0341 [ Clause 29. A method of identifying high-confidence PAS activators, comprising performing the method of any one of clauses 25-28 in a first screen of a library of RBP candidates, performing the method again in a second screen of candidates positive in the first screen, and designating as high-confidence those RBPs showing consistent activation in both screens with reporter-position specificity.[0342[ Clause 30. A pool of RBPs identified according to clause 29, wherein the pool comprises CPSF5, CPSF6, RNPS1, GRB2, MBNL1, MBNL2, YTHDF1, PCBP1, SCAF8, HNRNPF, and HNRNPH2.

[0343] Clause 31. A method of predicting RBP function in PAS selection, comprising: a) inputting amino acid sequences of RBPs into a fine-tuned protein language model trained on screening results of the method of any one of clauses 25-28; b) classifying the RBPs as-149- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760activators or non-activators based on a model score; and c) optionally outputting occlusion maps indicating sequence regions important for classification.

[0344] Clause 32. The method of clause 31, wherein the protein language model is ProteinBERT fine-tuned on tethered function assay data.

[0345] Clause 33. The method of clause 31 or 32, further comprising validating model predictions by MS2 tethering assays.

[0346] Clause 34. A method of mapping domain-level contributions to PAS activation in a candidate RBP, comprising: a) generating deletion variants or isolated domain constructs of the RBP fused to a coat protein; b) testing PAS activation by the method of any one of clauses 25-28; and c) associating changes in activity with the presence or absence of specific domains.[0347| Clause 35. The method of clause 34, wherein the RBP is GRB2 and the domains comprise N-terminal SH3, SH2, and C-terminal SH3 domains.

[0348] Clause 36. The method of clause 34, wherein the RBP is RNPS1 and the domains comprise an N-terminal intrinsically disordered region (IDR), an RNA recognition motif (RRM), and a C-terminal RS-rich region.

[0349] Clause 37. A method of determining direct interactions between an RBP and cleavage / polyadenylation (CPA) factor(s), comprising: a) purifying CPA subcomplexes; b) performing an in-vitro binding assay with the RBP or RBP fragment; and c) detecting specific subcomplex binding.

[0350] Clause 38. The method of clause 37, wherein GRB2 binds directly to CFIm complexes containing CPSF6 or CPSF7.

[0351] Clause 39. The method of clause 37, wherein RNPS1 binds directly to mPSF and CPSF6.

[0352] Clause 40. A method of modulating endogenous PAS selection, comprising delivering to a cell: a) an RBP-coat protein fusion; and b) a small nuclear RNA (snRNA) engineered to include an aptamer binding site for the coat protein and an antisense sequence targeting a PAS of interest.-150- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760

[0353] Clause 41. The method of clause 40, wherein the snRNA is a U7smOPT snRNA.

[0354] Clause 42. The method of clause 40 or 41, wherein the antisense sequence targets a region about 60 nt upstream of the PAS.

[0355] Clause 43. The method of any one of clauses 40-42, wherein the RBP is a high-confidence activator identified in claim 29.

[0356] Clause 44. A composition for modulating PAS selection in a target transcript, comprising: a) a minimal effector domain of an RBP identified as an activator by clause 29; and b) a programmable RNA targeting module configured to bind near a PAS.

[0357] Clause 45. The composition of clause 44, wherein the minimal effector domain is the C-terminal SH3 domain of GRB2 or the N-terminal IDR of RNPS1.

[0358] Clause 46. A method of treating a disease associated with aberrant APA, comprising administering to a subject an effective amount of the composition of clause 44 or 45 to restore or alter PAS usage in one or more target transcripts.

[0359] Clause 47. The method of clause 46, wherein the disease is selected from cancer, Dravet syndrome, autosomal dominant retinitis pigmentosa, and neurofibromatosis type 1.

[0360] Clause 48. A method of predicting protein regulators of polyadenylation site (PAS) selection, comprising: (a) obtaining amino acid sequences of a plurality of RNA-binding proteins (RBPs); (b) labeling each of the sequences as an activator or a non-activator of PAS selection based on experimental results of a tethered function assay; (c) inputting the sequences and corresponding labels into a sequence-based machine learning model; (d) training the machine learning model to classify proteins as activators or non-activators of PAS selection; and (e) outputting one or more predictions for an RBP of unknown PAS-activator status.

[0361] Clause 49. The method of clause 48, wherein the model comprises one or more of: a support vector machine, a convolutional neural network, or a protein language model.10362] Clause 50. The method of clause 49, wherein the protein language model is ProteinBERT.-151- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760

[0363] Clause 51. The method of clause 49, further comprising fine-tuning the model with class weighting and hyperparameter optimization.10364] Clause 52. The method of clause 49, wherein the sequences are trimmed at each terminus prior to training.

[0365] Clause 53. The method of clause 49, further comprising interpreting the trained model by analyzing attention weights and / or generating occlusion maps to identify amino acid regions contributory to classification.

[0366] Clause 54. The method of clause 53, wherein the contributory regions correspond to protein domains or low-complexity regions.

[0367] Clause 55. The method of clause 49, further comprising filtering the amino acid sequences to remove pairs of proteins having at least 80% sequence identity before splitting into training and test datasets.

[0368] Clause 56. The method of clause 55, wherein the machine learning model is validated on a set of zinc-finger proteins not included in the training dataset.

[0369] Clause 57. The method of clause 55, wherein the validation achieves a recall of at least 0.9 and an accuracy of at least 70%.

[0370] Clause 58. The method of clause 56 or 57, further comprising generating interpretive occlusion maps for a predicted activator from the zinc-finger protein set, wherein the occlusion maps identify functional domains associated with PAS regulation.

[0371] Clause 59. The method of clause 58, wherein the predicted activator is BRCA1 and the functional domains comprise a RING domain and a BRCT domain.

[0372] Equivalents

[0373] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this technology belongs.

[0374] The present technology illustratively described herein can suitably be practiced in the absence of any element or elements, limitation or limitations, not specifically disclosed herein. Thus, for example, the terms “comprising,” “including,” “containing,” etc. shall be-152- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760read expansively and without limitation. Additionally, the terms and expressions employed herein have been used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the present technology claimed.

[0375] Thus, it should be understood that the materials, methods, and examples provided here are representative of preferred aspects, are exemplary, and are not intended as limitations on the scope of the present technology.

[0376] The present technology has been described broadly and generically herein. Each of the narrower species and sub-generic groupings falling within the generic disclosure also form part of the present technology. This includes the generic description of the present technology with a proviso or negative limitation removing any subject matter from the genus, regardless of whether or not the excised material is specifically recited herein.

[0377] In addition, where features or aspects of the present technology are described in terms of Markush groups, those skilled in the art will recognize that the present technology is also thereby described in terms of any individual member or subgroup of members of the Markush group.

[0378] All publications, patent applications, patents, and other references mentioned herein are expressly incorporated by reference in their entirety, to the same extent as if each were incorporated by reference individually. In case of conflict, the present specification, including definitions, will control.

[0379] Other aspects are set forth within the following claims.-153- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760REFERENCES1. Gerstberger, S., Hafner, M., and Tuschl, T. (2014). A census of human RNA-binding proteins. Nat. Rev. Genet. 15, 829-845. https: / / doi.org / 10.1038 / nrg3813.2. Queiroz, R.M.L., Smith, T., Villanueva, E., Marti-Solano, M., Monti, M., Pizzinga, M., Mirea, D.-M., Ramakrishna, M., Harvey, R.F., Dezi, V., et al. (2019). Comprehensive identification of RNA-protein interactions in any organism using orthogonal organic phase separation (OOPS). Nat. Biotechnol. 37, 169-178. https: / / doi.org / 10.1038 / s41587- 018-0001-2.3. Shi, Y. (2012). Alternative polyadenylation: new insights from global analyses. RNA N. Y. N 18, 2105-2117. https: / / doi.org / 10.1261 / rna.035899.112.4. Tian, B., and Manley, J.L. (2017). Alternative polyadenylation of mRNA precursors. Nat. Rev. Mol. Cell Biol. 18, 18-30. https: / / doi.org / 10.1038 / nrm.2016.116.5. Derti, A., Garrett-Engele, P., Maclsaac, K.D., Stevens, R.C., Sriram, S., Chen, R., Rohl, C.A., Johnson, J.M., and Babak, T. (2012). A quantitative atlas of polyadenylation in five mammals. Genome Res. 22, 1173-1183. https: / / doi.org / 10.1101 / gr.132563.lll.6. Elkon, R., Ugalde, A.P., and Agami, R. (2013). Alternative cleavage and polyadenylation: extent, regulation and function. Nat. Rev. Genet. 14, 496-506. https: / / doi.org / 10.1038 / nrg3482.7. Gruber, A. J., and Zavolan, M. (2019). Alternative cleavage and polyadenylation in health and disease. Nat. Rev. Genet. 20, 599-614. https: / / doi.org / 10.1038 / s41576-019- 0145-z.8. Mitschka, S., and Mayr, C. (2022). Context-specific regulation and function of mRNA alternative polyadenylation | Nature Reviews Molecular Cell Biology. Nat. Rev. Mol. Cell Biol, https: / / doi.org / 10.1038 / s41580-022-00507-5.9. Kowalski, M.H., Wessels, H.-H., Linder, J., Choudhary, S., Hartman, A., Hao, Y., Mascio, I., Dalgamo, C., Kundaje, A., and Satija, R. (2023). CPA-Perturb-seq:Multiplexed single-cell characterization of alternative polyadenylation regulators (Genomics) https: / / doi.org / 10.1101 / 2023.02.09.527751.-154- 4918-7893-0315.1Atty. Dkt. No.: 114198-376010. Li, W., You, B., Hoque, M., Zheng, D., Luo, W., Ji, Z., Park, J.Y., Gunderson, S.I., Kalsotra, A., Manley, J.L., et al. (2015). Systematic Profiling of Poly(A)+ Transcripts Modulated by Core 3’ End Processing and Splicing Factors Reveals Regulatory Rules of Alternative Cleavage and Polyadenylation. PLOS Genet. 11, el005166. https: / / doi.org / 10.1371 / journal.pgen.1005166.11. Ogorodnikov, A., Levin, M., Tattikota, S., Tokalov, S., Hoque, M., Scherzinger, D., Marini, F., Poetsch, A., Binder, H., Macher-Goppinger, S., et al. (2018). Transcriptome 3 'end organization by PCF11 links alternative polyadenylation to formation and neuronal differentiation of neuroblastoma. Nat. Commun. 9, 5331. https: / / doi.org / 10.1038 / s41467-018-07580-5.12. Ni, Z., Ahmed, N., Nabeel-Shah, S., Guo, X., Pu, S., Song, J., Marcon, E., Burke, G.L., Tong, A.H.Y., Chan, K., et al. (2024). Identifying human pre-mRNA cleavage and polyadenylation factors by genome-wide CRISPR screens using a dual fluorescence readthrough reporter. Nucleic Acids Res. 52, 4483-4501. https: / / doi.org / 10.1093 / nar / gkae240.13. Schmok, J.C., Jain, M., Street, L.A., Tankka, A.T., Schafer, D., Her, H.-L., Elmsaouri, S., Gosztyla, M.L., Boyle, E.A., Jagannatha, P., et al. (2024). Large-scale evaluation of the ability of RNA-binding proteins to activate exon inclusion. Nat.Biotechnol., 1-13. https: / / doi.org / 10.1038 / s41587-023-02014-0.14. Erben, E.D., Fadda, A., Lueong, S., Hoheisel, J.D., and Clayton, C. (2014). A Genome-Wide Tethering Screen Reveals Novel Potential Post-Transcriptional Regulators in Trypanosoma brucei. PLoS Pathog. 10, el004178. https: / / doi.org / 10.1371 / journal.ppat.1004178.15. Luo, E.-C., Nathanson, J.L., Tan, F.E., Schwartz, J.L., Schmok, J.C., Shankar, A., Markmiller, S., Yee, B.A., Sathe, S., Pratt, G.A., et al. (2020). Large-scale tethered function assays identify factors that regulate mRNA stability and translation. Nat. Struct. Mol. Biol. 27, 989-1000. https: / / doi.org / 10.1038 / s41594-020-0477-6.16. Zhu, Y., Wang, X., Forouzmand, E., Jeong, J., Qiao, F., Sowd, G.A., Engelman, A.N., Xie, X., Hertel, K.J., and Shi, Y. (2018). Molecular Mechanisms for CFIm-Mediated-155- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760Regulation of mRNA Alternative Polyadenylation. Mol. Cell 69, 62-74. e4. https: / / doi.Org / 10.1016 / j.molcel.2017.l 1.031.17. Gruber, A.J., Schmidt, R., Gruber, A.R., Martin, G., Ghosh, S., Belmadani, M., Keller, W., and Zavolan, M. (2016). A comprehensive analysis of 3’ end sequencing data sets reveals novel polyadenylation signals and the repressive role of heterogeneous ribonucleoprotein C on cleavage and polyadenylation. Genome Res. 26, 1145-1159. https: / / doi.org / 10.1101 / gr.202432.115.18. Brown, K.M., and Gilmartin, G.M. (2003). A Mechanism for the Regulation of Pre-mRNA 3' Processing by Human Cleavage Factor Im. Mol. Cell 12, 1467-1476. https: / / doi.org / 10.1016 / S1097-2765(03)00453-2.19. Yang, Q., Gilmartin, G.M., and Doublie, S. (2010). Structural basis of UGUA recognition by the Nudix protein CFIm25 and implications for a regulatory role in mRNA 3' processing. Proc. Natl. Acad. Sci. 107, 10062-10067. https: / / doi.org / 10.1073 / pnas.1000848107.20. Yang, W., Hsu, P.L., Yang, F., Song, J.-E., and Varani, G. (2018). Reconstitution of the CstF complex unveils a regulatory role for CstF-50 in recognition of 3 '-end processing signals. Nucleic Acids Res. 46, 493-503. https: / / doi.org / 10.1093 / nar / gkxll77.21. Batra, R., Charizanis, K., Manchanda, M., Mohan, A., Li, M., Finn, D.J., Goodwin, M., Zhang, C., Sobczak, K., Thornton, C.A., et al. (2014). Loss of MBNL Leads to Disruption of Developmentally Regulated Alternative Polyadenylation in RNA-Mediated Disease. Mol. Cell 56, 311-322. https: / / doi.Org / 10.1016 / j.molcel.2014.08.027.22. Goodwin, M., Mohan, A., Batra, R., Lee, K.-Y., Charizanis, K., Gomez, F.J.F., Eddarkaoui, S., Sergeant, N., Buee, L., Kimura, T., et al. (2015). MBNL Sequestration by Toxic RNAs and RNA Misprocessing in the Myotonic Dystrophy Brain. Cell Rep. 12, 1159-1168. https: / / doi.Org / 10.1016 / j.celrep.2015.07.029.23. Mohan, N., Kumar, V., Kandala, D.T., Kartha, C.C., and Laishram, R.S. (2018). A Splicing-Independent Function of RBM10 Controls Specific 3' UTR Processing to Regulate Cardiac Hypertrophy. Cell Rep. 24, 3539-3553. https: / / doi.Org / 10.1016 / j.celrep.2018.08.077.-156- 4918-7893-0315.1h cottdpisn: / g / d troain.osrcgr / ip10ts..1 P10re1p / 2ri0n2t4 a.t0b6i.o1R2.x5i9v8, 7 h6tt6p.s: / / doi.org / 10.1101 / 2024.06.12.598766Atty. Dkt. No.: 114198-376024. Chen, L., Fu, Y., Hu, Z., Deng, K., Song, Z., Liu, S., Li, M., Ou, X., Wu, R., Liu, M., et al. (2022). Nuclear m6A reader YTHDC1 suppresses proximal alternative polyadenylation sites by interfering with the 3' processing machinery. EMBO Rep. 23, e54686. https: / / doi.org / 10.15252 / embr.202254686.25. Gregersen, L.H., Mitter, R., Ugalde, A.P., Nojima, T., Proudfoot, N.J., Agami, R., Stewart, A., and Svejstrup, J.Q. (2019). SCAF4 and SCAF8, mRNA Anti -Terminator Proteins. Cell 177, 1797-1813. el8. https: / / doi.Org / 10.1016 / j.cell.2019.04.038.26. Ji, X., Wan, J., Vishnu, M., Xing, Y., and Liebhaber, S.A. (2013). aCP Poly(C) binding proteins act as global regulators of alternative poly adenylation. Mol. Cell. Biol.33, 2560-2573. https: / / doi.org / 10.1128 / MCB.01380-12.27. Nazim, M., Masuda, A., Rahman, M.A., Nasrin, F., Takeda, J. -I., Ohe, K., Ohkawara, B., Ito, M., and Ohno, K. (2017). Competitive regulation of alternative splicing and alternative polyadenylation by hnRNP H and CstF64 determines acetylcholinesterase isoforms. Nucleic Acids Res. 45, 1455-1468. https: / / doi.org / 10.1093 / nar / gkw823.28. Gadgil, A., and Raczynska, K.D. (2021). U7 snRNA: A tool for gene therapy. J. Gene Med. 23, e3321. https: / / doi.org / 10.1002 / jgm.3321.29. Madocsai, C., Lim, S.R., Geib, T., Lam, B.J., and Hertel, K.J. (2005). Correction of SMN2 Pre-mRNA splicing by antisense U7 small nuclear RNAs. Mol. Ther. 12, 1013— 1022. https: / / doi.Org / 10.1016 / j.ymthe.2005.08.022.30. Hatch, S.T., Smargon, A. A., and Yeo, G.W. (2022). Engineered U1 snRNAs to modulate alternatively spliced exons. Methods San Diego Calif 205, 140-148. https: / / doi.Org / 10.1016 / j.ymeth.2022.06.008.31. Smargon, A.A., Pant, D., Glynne, S., Gomberg, T.A., and Yeo, G.W. (2024). Small nuclear RNAs enhance protein-free RNA-programmable base conversion on mammalian32. Jin, W., Brannan, K.W., Kapeli, K., Park, S.S., Tan, H.Q., Gosztyla, M.L., Mujumdar, M., Ahdout, J., Henroid, B., Rothamel, K., et al. (2023). HydRA: Deep-learning models for predicting RNA-binding capacity from protein interaction association context and -157- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760protein sequence. Mol. Cell 83, 2595-2611. ell. https: / / doi.Org / 10.1016 / j.molcel.2023.06.019.33. Brandes, N., Ofer, D., Peleg, Y., Rappoport, N., and Linial, M. (2022). ProteinBERT: a universal deep-learning model of protein sequence and function. Bioinformatics 38, 2102-2110. https: / / doi.org / 10.1093 / bioinformatics / btac020.34. Lemay, J.-F., Marguerat, S., Larochelle, M., Liu, X., van Nues, R., Hunyadkiirti, J., Hoque, M., Tian, B., Granneman, S., Bahler, J., et al. (2016). TheNrdl-like protein Sebl coordinates cotranscriptional 3’ end processing and polyadenylation site selection. Genes Dev. 30, 1558-1572. https: / / doi.org / 10.1101 / gad.280222.116.35. Meinhart, A., and Cramer, P. (2004). Recognition of RNA polymerase II carboxyterminal domain by 3’-RNA-processing factors. Nature 430, 223-226. https: / / doi.org / 10.1038 / nature02679.36. Gosztyla, M.L., Zhan, L., Olson, S., Wei, X., Naritomi, J., Nguyen, G., Street, L., Goda, G.A., Cavazos, F.F., Schmok, J.C., et al. (2024). Integrated multi-omics analysis of zinc-finger proteins uncovers roles in RNA regulation. Mol. Cell 84, 3826-3842. e8. https: / / doi.Org / 10.1016 / j.molcel.2024.08.010.37. Kleiman, F.E., and Manley, J.L. (2001). The BARDl-CstF-50 interaction links mRNA 3’ end formation to DNA damage and tumor suppression. Cell 104, 743-753. https: / / doi.org / 10.1016 / s0092-8674(01)00270-7.38. Kleiman, F.E., and Manley, J.L. (1999). Functional interaction of BRC Al -associated BARD1 with polyadenylation factor CstF-50. Science 285, 1576-1579. https: / / doi.Org / 10.l 126 / science.285.5433.1576.39. Shepard, P.J., Choi, E.-A., Lu, J., Flanagan, L.A., Hertel, K.J., and Shi, Y. (2011). Complex and dynamic landscape of RNA polyadenylation revealed by PAS-Seq. RNA 17, 761-772. https: / / doi.org / 10.1261 / rna.2581711.40. Yoon, Y., Soles, L.V., and Shi, Y. (2021). PAS-seq 2: A fast and sensitive method for global profiling of polyadenylated RNAs. Methods Enzymol. 655, 25-35. https: / / doi.org / 10.1016 / bs.mie.2021.03.013.-158- 4918-7893-0315.1Atty. Dkt. No.: 114198-376041. Veraldi, K.L., Arhin, G.K., Martincic, K., Chung-Ganster, L.H., Wilusz, J., and Milcarek, C. (2001). hnRNP F influences binding of a 64-kilodalton subunit of cleavage stimulation factor to mRNA precursors in mouse B cells. Mol. Cell. Biol. 21, 1228-1238. https: / / doi.Org / 10.1128 / MCB.21.4.1228-1238.2001.42. Bonnal, S., Martinez, C., Fbrch, P., Bachi, A., Wilm, M., and Valcarcel, J. (2008). RBM5 / Luca-15 / H37 Regulates Fas Alternative Splice Site Pairing after Exon Definition. Mol. Cell 32, 81-95. https: / / doi.Org / 10.1016 / j.molcel.2008.08.008.43. Fushimi, K., Ray, P., Kar, A., Wang, L., Sutherland, L.C., and Wu, J.Y. (2008). Upregulation of the proapoptotic caspase 2 splicing isoform by a candidate tumor suppressor, RBM5. Proc. Natl. Acad. Sci. U. S. A. 105, 15708-15713. https: / / doi.org / 10.1073 / pnas.0805569105.44. Hou, B., Xu, S., Xu, Y., Gao, Q., Zhang, C., Liu, L., Yang, H., Jiang, X., and Che, Y. (2019). Grb2 binds to PTEN and regulates its nuclear translocation to maintain the genomic stability in DNA damage response. Cell Death Dis. 10, 1-14. https: / / doi.org / 10.1038 / s41419-019-1762-3.45. Ye, Z., Xu, S., Shi, Y., Bacolla, A., Syed, A., Moiani, D., Tsai, C.-L., Shen, Q., Peng, G., Leonard, P.G., et al. (2021). GRB2 enforces homology-directed repair initiation by MREll. Sci. Adv. 7, eabe9254. https: / / doi.org / 10.1126 / sciadv.abe9254.46. Soles, L.V., Li, S., Liu, L., Sarkan, K.S.K., Alvstad, E.G., Tian, L., Yoon, Y., Valdez, M.C., Marazzi, I., and Shi, Y. (2025). The competition between splicing and 3' processing shapes the human transcriptome. Preprint at bioRxiv, https: / / doi.org / 10.1101 / 2025.08.19.671063 https: / / doi.org / 10.1101 / 2025.08.19.671063.47. Wang, E.T., Sandberg, R., Luo, S., Khrebtukova, I., Zhang, L., Mayr, C., Kingsmore, S.F., Schroth, G.P., and Burge, C.B. (2008). Alternative isoform regulation in human tissue transcriptomes. Nature 456, 470-476. https: / / doi.org / 10.1038 / nature07509.48. Zhang, Z., Bae, B., Cuddleston, W.H., and Miura, P. (2023). Coordination of alternative splicing and alternative polyadenylation revealed by targeted long read sequencing. Nat. Commun. 14, 5506. https: / / doi.org / 10.1038 / s41467-023-41207-8.-159- 4918-7893-0315.1Atty. Dkt. No.: 114198-376049. Szklarczyk, D., Kirsch, R., Koutrouli, M., Nastou, K., Mehryary, F., Hachilif, R., Gable, A.L., Fang, T., Doncheva, N.T., Pyysalo, S., et al. (2023). The STRING database in 2023 : protein-protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Res. 51, D638-D646. https: / / doi.org / 10.1093 / nar / gkacl000.50. Lowenstein, E.J., Daly, R.J., Batzer, A.G., Li, W., Margolis, B., Lammers, R., Ullrich, A., Skolnik, E.Y., Bar-Sagi, D., and Schlessinger, J. (1992). The SH2 and SH3 domaincontaining protein GRB2 links receptor tyrosine kinases to ras signaling. Cell 70, 431— 442. https: / / doi . org / 10.1016 / 0092-8674(92)90167-b .51. Belov, A. A., and Mohammadi, M. (2012). Grb2, a double-edged sword of receptor tyrosine kinase signaling. Sci. Signal. 5, pe49. https: / / doi.org / 10.1126 / scisignal.2003576.52. Tari, A.M., and Lopez-Berestein, G. (2001). GRB2: a pivotal protein in signal transduction. Semin. Oncol. 28, 142-147. https: / / doi.org / 10.1016 / s0093-7754(01)90291-x.53. Verbeek, B.S., Adriaansen-Slot, S.S., Rijksen, G., and Vroom, T.M. (1997). Grb2 overexpression in nuclei and cytoplasm of human breast cells: a histochemical and biochemical study of normal and neoplastic mammary tissue specimens. J. Pathol. 183, 195-203. https: / / doi.org / 10.1002 / (sici)1096-9896(199710)183:2<195::aid-path901>3.0.co;2-y.54. Yamazaki, T., Zaal, K., Hailey, D., Presley, J., Lippincott- Schwartz, J., and Samelson, L.E. (2002). Role of Grb2 in EGF-stimulated EGFR internalization. J. Cell Sci. 115, 1791-1802. https: / / doi.Org / 10.1242 / jcs.115.9.1791.55. Ye, Z., Xu, S., Shi, Y., Cheng, X., Zhang, Y., Roy, S., Namjoshi, S., Longo, M.A., Link, T.M., Schlacher, K., et al. (2024). GRB2 stabilizes RAD51 at reversed replication forks suppressing genomic instability and innate immunity against cancer. Nat. Commun.15, 2132. https: / / doi.org / 10.1038 / s41467-024-46283-y.56. Castello, A., Fischer, B., Eichelbaum, K., Horos, R., Beckmann, B.M., Strein, C., Davey, N.E., Humphreys, D.T., Preiss, T., Steinmetz, L.M., et al. (2012). Insights into-160- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760RNA biology from an atlas of mammalian mRNA-binding proteins. Cell 149, 1393-1406. https: / / doi.Org / 10.1016 / j.cell.2012.04.031.57. Schlautmann, L.P., Lackmann, J.-W., Altrmiller, J., Dieterich, C., Boehm, V., and Gehring, N.H. (2022). Exon junction complex-associated multi-adapter RNPS1 nucleates splicing regulatory complexes to maintain transcriptome surveillance. Nucleic Acids Res.50, 5899-5918. https: / / doi.org / 10.1093 / nar / gkac428.58. Murachelli, A.G., Ebert, J., Basquin, C., Le Hir, H., and Conti, E. (2012). The structure of the ASAP core complex reveals the existence of a Pinin-containing PSAP complex. Nat. Struct. Mol. Biol. 19, 378-386. https: / / doi.org / 10.1038 / nsmb.2242.59. Schwerk, C., Prasad, J., Degenhardt, K., Erdjument-Bromage, H., White, E., Tempst, P., Kidd, V.J., Manley, J.L., Lahti, J.M., and Reinberg, D. (2003). ASAP, a novel protein complex involved in RNA processing and apoptosis. Mol. Cell. Biol. 23, 2981-2990. https: / / doi.Org / 10.1128 / MCB.23.8.2981-2990.2003.60. Le Hir, H., Izaurralde, E., Maquat, L.E., and Moore, M.J. (2000). The spliceosome deposits multiple proteins 20-24 nucleotides upstream of mRNA exon-exon junctions. EMBO J. 19, 6860-6869. https: / / doi.org / 10.1093 / emboj / 19.24.6860.61. Sauliere, J., Murigneux, V., Wang, Z., Marquenet, E., Barbosa, I., Le Tonqueze, O., Audic, Y., Paillard, L., Roest Crollius, H., and Le Hir, H. (2012). CLIP-seq of eIF4AIII reveals transcriptome-wide mapping of the human exon junction complex. Nat. Struct. Mol. Biol. 19, 1124-1131. https: / / doi.org / 10.1038 / nsmb.2420.62. Hir, H.L., Sauliere, J., and Wang, Z. (2016). The exon junction complex as a node of post-transcriptional networks. Nat. Rev. Mol. Cell Biol. 17, 41-54. https: / / doi.Org / 10.1038 / nrm.2015.7.63. McCracken, S., Longman, D., Johnstone, I.L., Caceres, J.F., and Blencowe, B.J.(2003). An Evolutionarily Conserved Role for SRml60 in 3 '-End Processing That Functions Independently of Exon Junction Complex Formation*. J. Biol. Chem. 278, 44153-44160. https: / / doi.org / 10.1074 / jbc.M306856200.-161- 4918-7893-0315.1Atty. Dkt. No.: 114198-376064. Yoon, Y., Liu, L., Quan, C., and Shi, Y. (2025). Emerging Roles of Biomolecular Condensates in Pre-mRNA 3’ End Processing. Wiley Interdiscip. Rev. RNA 16, e70024. https: / / doi.org / 10.1002 / wrna.70024.65. Yoon, Y., Bournique, E., Soles, L.V., Yin, H., Chu, H.-F., Yin, C., Zhuang, Y., Liu, X., Liu, L., Jeong, J., et al. (2025). RBBP6 anchors pre-mRNA 3' end processing to nuclear speckles for efficient gene expression. Mol. Cell 85, 555-570. e8. https: / / doi.Org / 10.1016 / j.molcel.2024.12.016.66. Yang, S., Aelst, L.V., and Bar-Sagi, D. (1995). Differential Interactions of Human Sosl and Sos2 with Grb2 *. J. Biol. Chem. 270, 18212-18215. https: / / doi.org / 10.1074 / jbc.270.31.18212.67. Street, L.A., Rothamel, K.L., Brannan, K.W., Jin, W., Bokor, B.J., Dong, K., Rhine, K., Madrigal, A., Al-Azzam, N., Kim, J.K., et al. (2024). Large-scale map of RNA-binding protein interactomes across the mRNA life cycle. Mol. Cell 84, 3790-3809. e8. https: / / doi.Org / 10.1016 / j.molcel.2024.08.030.68. Yoon, Y., Bournique, E., Soles, L.V., Yin, H., Chu, H.-F., Yin, C., Zhuang, Y., Liu, X., Liu, L., Jeong, J., et al. (2025). RBBP6 anchors pre-mRNA 3' end processing to nuclear speckles for efficient gene expression. Mol. Cell 0. https: / / doi.Org / 10.1016 / j.molcel.2024.12.016.69. Schmidt, M., Kluge, F., Sandmeir, F., Kuhn, U., Schafer, P., Tilting, C., Ihling, C., Conti, E., and Wahle, E. (2022). Reconstitution of 3’ end processing of mammalian pre-mRNA reveals a central role of RBBP6. Genes Dev. 36, 195-209. https: / / doi.org / 10.1101 / gad.349217.121.70. Bisson, N., James, D.A., Ivosev, G., Tate, S.A., Bonner, R., Taylor, L., and Pawson, T. (2011). Selected reaction monitoring mass spectrometry reveals the dynamics of signaling through the GRB2 adaptor. Nat. Biotechnol. 29, 653-658. https: / / doi.org / 10.1038 / nbt.1905.71. Brumbaugh, J., Di Stefano, B., Wang, X., Borkent, M., Forouzmand, E., Clowers, K.J., Ji, F., Schwarz, B.A., Kalocsay, M., Elledge, S.J., et al. (2018). Nudt21 Controls-162- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760Cell Fate by Connecting Alternative Polyadenylation to Chromatin Signaling. Cell 172, 106-120. e21. https: / / doi.Org / 10.1016 / j.cell.2017.ll.023.72. Lee, S.-H., Singh, I., Tisdale, S., Abdel-Wahab, O., Leslie, C.S., and Mayr, C. (2018). Widespread intronic polyadenylation inactivates tumour suppressor genes in leukaemia. Nature 561, 127-131. https: / / doi.org / 10.1038 / s41586-018-0465-8.73. Insco, M.L., Abraham, B.J., Dubbury, S.J., Kaltheuner, I.H., Dust, S., Wu, C., Chen, K.Y., Liu, D., Bellaousov, S., Cox, A.M., et al. (2023). Oncogenic CDK13 mutations impede nuclear RNA surveillance. Science 380, eabn7625. https: / / doi.org / 10.1126 / science.abn7625.74. Mayr, C., and Bartel, D.P. (2009). Widespread shortening of 3'UTRs by alternative cleavage and polyadenylation activates oncogenes in cancer cells. Cell 138, 673. https: / / doi.Org / 10.1016 / j.cell.2009.06.016.75. Xia, Z., Donehower, L.A., Cooper, T.A., Neilson, J.R., Wheeler, D.A., Wagner, E.J., and Li, W. (2014). Dynamic analyses of alternative polyadenylation from RNA-seq reveal a 3'-UTR landscape across seven tumour types. Nat. Commun. 5, 5274. https: / / doi.org / 10.1038 / ncomms6274.76. Boehm, V., Britto-Borges, T., Steckelberg, A.-L., Singh, K.K., Gerbracht, J.V., Gueney, E., Blazquez, L., Altrmiller, J., Dieterich, C., and Gehring, N.H. (2018). Exon Junction Complexes Suppress Spurious Splice Sites to Safeguard Transcriptome Integrity. Mol. Cell 72, 482-495. e7. https: / / doi.Org / 10.1016 / j.molcel.2018.08.030.77. Soles, L.V., Liu, L., Zou, X., Yoon, Y., Li, S., Tian, L., Valdez, M., Yu, A.M., Yin, H., Li, W., et al. (2025). A nuclear RNA degradation code is recognized by PAXT for eukaryotic transcriptome surveillance. Mol. Cell 85, 1575-1588. e9. https: / / doi.Org / 10.1016 / j.molcel.2025.03.010.78. Virtanen, P., Gommers, R., Oliphant, T.E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., et al. (2020). SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat. Methods 17, 261-272. https: / / doi.org / 10.1038 / s41592-019-0686-2.-163- 4918-7893-0315.1Atty. Dkt. No.: 114198-376079. Badia-i-Mompel, P., Velez Santiago, J., Braunger, J., Geiss, C., Dimitrov, D., Miiller-Dott, S., Taus, P., Dugourd, A., Holland, C.H., Ramirez Flores, R.O., et al. (2022). decoupleR: ensemble of computational methods to infer biological activities from omics data. Bioinforma. Adv. 2, vbac016. https: / / doi.org / 10.1093 / bioadv / vbac016.80. Steinegger, M., and Sbding, J. (2017). MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat. Biotechnol. 35, 1026-1028. https: / / doi.org / 10.1038 / nbt.3988.81. Boyle, E.A., Her, H.-L., Mueller, J.R., Naritomi, J.T., Nguyen, G.G., and Yeo, G.W. (2023). Skipper analysis of eCLIP datasets enables sensitive detection of constrained translation factor binding sites. Cell Genomics 3, 100317. https: / / doi.Org / 10.1016 / j.xgen.2023.100317.82. Her, H.-L., Boyle, E., and Yeo, G.W. (2022). Metadensity: a background-aware python pipeline for summarizing CLIP signals on various transcriptomic sites.Bioinforma. Adv. 2, vbac083. https: / / doi.org / 10.1093 / bioadv / vbac083.83. Heinz, S., Benner, C., Spann, N., Bertolino, E., Lin, Y.C., Laslo, P., Cheng, J.X., Murre, C., Singh, H., and Glass, C.K. (2010). Simple combinations of lineagedetermining transcription factors prime cis-regulatory elements required for macrophage and B cell identities. Mol. Cell 38, 576-589. https: / / doi.Org / 10.1016 / j.molcel.2010.05.004.84. Cox, J., and Mann, M. (2008). MaxQuant enables high peptide identification rates, individualized p.p.b. -range mass accuracies and proteome-wide protein quantification. Nat. Biotechnol. 26, 1367-1372. https: / / doi.org / 10.1038 / nbt.1511.85. Zhou, Y., Zhou, B., Pache, L., Chang, M., Khodabakhshi, A.H., Tanaseichuk, O., Benner, C., and Chanda, S.K. (2019). Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. Nat. Commun. 10, 1523. https: / / doi.org / 10.1038 / s41467-019-09234-6.86. Quinlan, A.R., and Hall, I.M. (2010). BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841-842. https: / / doi.org / 10.1093 / bioinformatics / btq033.-164- 4918-7893-0315.1Atty. Dkt. No.: 114198-376081. Shen, S., Park, J.W., Lu, Z., Lin, L., Henry, M.D., Wu, Y.N., Zhou, Q., and Xing, Y. (2014). rMATS: Robust and flexible detection of differential alternative splicing from replicate RNA-Seq data. Proc. Natl. Acad. Sci. Ill, E5593-E5601. https: / / doi.org / 10.1073 / pnas.1419161111.88. Lackford, B., Yao, C., Charles, G.M., Weng, L., Zheng, X., Choi, E.-A., Xie, X., Wan, J., Xing, Y., Freudenberg, J.M., et al. (2014). Fipl regulates mRNA alternative polyadenylation to promote stem cell self-renewal. EMBO J. 33, 878-889. https: / / doi.org / 10.1002 / embj.201386537.89. Yao, C., Biesinger, J., Wan, J., Weng, L., Xing, Y., Xie, X., and Shi, Y. (2012). Transcriptome-wide analyses of CstF64-RNA interactions in global regulation of mRNA alternative polyadenylation. Proc. Natl. Acad. Sci. 109, 18773-18778. https: / / doi.org / 10.1073 / pnas.1211101109.90. HMMER http: / / hmmer.org / .91. Engler, C., Kandzia, R., and Marillonnet, S. (2008). A One Pot, One Step, Precision Cloning Method with High Throughput Capability. PLOS ONE 3, e3647. https: / / doi.org / 10.1371 / journal.pone.0003647.92. Sheets, M.D., Ogg, S.C., and Wickens, M.P. (1990). Point mutations in AAUAAA and the poly (A) addition site: effects on the accuracy and efficiency of cleavage and polyadenylation in vitro. Nucleic Acids Res. 18, 5799-5805.93. R Core Team (2023). R: A Language and Environment for Statistical Computing (R Foundation for Statistical Computing).94. The UniProt Consortium (2017). UniProt: the universal protein knowledgebase. Nucleic Acids Res. 45, D158-D169. https: / / doi.org / 10.1093 / nar / gkwl099.95. Frankish, A., Carbonell-Sala, S., Diekhans, M., Jungreis, I., Loveland, J.E., Mudge, J.M., Sisu, C., Wright, J.C., Arnan, C., Barnes, I., et al. (2023). GENCODE: reference annotation for the human and mouse genomes in 2023. Nucleic Acids Res. 51, D942-D949. https: / / doi.org / 10.1093 / nar / gkacl071.96. Wang, X., Liu, L., Whisnant, A.W., Hennig, T., Djakovic, L., Haque, N., Bach, C., Sandri-Goldin, R.M., Erhard, F., Friedel, C.C., et al. (2021). Mechanism and-165- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760consequences of herpes simplex virus 1 -mediated regulation of host mRNA alternative polyadenylation. PLoS Genet. 17, el 009263. https: / / doi.org / 10.1371 / journal.pgen.1009263.97. Waskom, M.L. (2021). seaborn: statistical data visualization. J. Open Source Softw.6, 3021. https: / / doi.org / 10.21105 / joss.03021.98. Blue, S.M., Yee, B.A., Pratt, G.A., Mueller, J.R., Park, S.S., Shishkin, A.A., Stamer, A.C., Van Nostrand, E.L., and Yeo, G.W. (2022). Transcriptome-wide identification of RNA-binding protein binding sites using seCLIP-seq. Nat. Protoc. 17, 1223-1265. https: / / doi.org / 10.1038 / s41596-022-00680-z.99. Rappsilber, J., Mann, M., and Ishihama, Y. (2007). Protocol for micro-purification, enrichment, pre-fractionation and storage of peptides for proteomics using StageTips. Nat. Protoc. 2, 1896-1906. https: / / doi.org / 10.1038 / nprot.2007.261.100. Gradia, S.D., Ishida, J.P., Tsai, M.-S., Jeans, C., Tainer, J. A., and Fuss, J.O. (2017). MacroBac: New Technologies for Robust and Efficient Large-Scale Production of Recombinant Multiprotein Complexes. Methods Enzymol. 592, 1-26. https: / / doi.org / 10.1016 / bs.mie.2017.03.008.4918-7893-0315.1

Claims

Atty. Dkt. No.: 114198-3760WHAT IS CLAIMED IS:

1. An antisense RNA comprising, from the 5 ’ end to the 3 ’ end:a) a nucleotide sequence encoding an aptamer, wherein the aptamer has a high affinity to an alternative polyadenylation (APA)-modulating RBP; andb) a nucleotide sequence antisense to a target region proximal to a polyA site (PAS).

2. The antisense RNA of claim 1, further comprising a nucleotide sequence comprising an RBP / RNP (ribonucleoprotein)-recruiting motif.

3. The antisense RNA of claim 1 or 2, wherein the aptamer is derived from a bacteriophage selected from: MS2, R17, , PP7, QP or GA; iron responsive protein (IRP); bovine immunodeficiency virus (BIV), or human U1 small nuclear ribonucleoprotein A.

4. The antisense RNA of claim 3, wherein the aptamer is derived from MS2.

5. The antisense RNA of claim 1, wherein: the RNA aptamer comprises a sequence consisting of SEQ ID NO: 1 (MS2 hairpin), SEQ ID NO: 2 (MS2 high affinity mutant), or an equivalent thereof having at least 90% sequence identity thereto; the aptamer is configured to recruit an RNA-binding protein selected from the group consisting of CPSF5 (SEQ ID NO: ), CPSF6 (SEQ ID NO: ), RNPS1 (SEQ ID NO: ), and GRB2 (SEQ ID NO: ); the antisense RNA hybridizes to a target region proximal to a polyadenylation site, wherein “proximal” comprises a region upstream or downstream of the cleavage site sufficient to modulate PAS selection, without limitation to a fixed nucleotide distance; and optionally wherein the method is used for regulating PAS selection in an mRNA of a gene associated with a disease involving aberrant alternative polyadenylation, including cancers, haploinsufficiency disorders, and other APA-related pathological conditions.

6. The antisense RNA of any one of claims 1-4, wherein the target region proximal to the PAS is approximately 20bp, approximately 40bp, approximately 60bp, approximately 80bp, approximately lOObp, approximately 120bp, approximately 140bp, approximately 160bp, approximately 180bp, or approximately 200bp upstream or downstream of the PAS.

7. A vector comprising the antisense RNA of any one of claims 1-6.-167- 4918-7893-0315.1Atty. Dkt. No.: 114198-37608. The vector of claim 7, wherein the vector is selected from a plasmid, a viral vector, a cosmid, or a phage, optionally wherein the viral vector is selected from a baculovirus, a retrovirus, or an adenovirus.

9. The vector of claim 7 or 8, further comprising a nucleotide sequence encoding a Polymerase II or Polymerase III promoter.

10. The vector of claim 9, wherein the Polymerase III promoter is selected from a Ul, U6, or U7 promoter.

11. The vector of any of claims 7-10, wherein the antisense RNA is operably linked to a nucleotide sequence encoding an RBP, and wherein the RBP is a modulator of PAS selection.

12. The vector of claim 11, wherein the RBP is mammalian.

13. The vector of claim 11 or 12, wherein the RBP activates PAS selection, and wherein the RBP is selected from any of the RBPs in FIG. 2E.

14. The vector of claim 10 or 11, wherein the RBP is selected from CPSF5, CPSF6, RNPS1, and GRB2.

15. An isolated host cell comprising one or more of the antisense RNA of any one of claims 1-7 or the vector of any one of claims claim 8-13, wherein the host cell is a prokaryotic or a eukaryotic cell.

16. The isolated host cell of claim 14, further comprising a polynucleotide encoding the RBP or the RBP, optionally wherein the RBP is selected from any one of the RBPs identified in FIGS. 2A or 2E or Table 4.

17. The isolated host cell of claim 15, wherein the RBP is selected from CPSF5, CPSF6, RNPS1, and GRB2.

18. A composition comprising a pharmaceutically acceptable carrier and one or more of the antisense RNA of any one of claims 1-6, the vector of any one of claims 7-14, and / or the isolated host cell of any of claims 15-17.

19. A method of regulating polyA site (PAS) selection on mRNA in a cell, the method comprising:-168- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760a) contacting the cell with the antisense RNA of any one of claims 1-6 or the vector of any of claims 7-10, andb) contacting the cell with an RBP, or polynucleotide encoding an RBP, optionally wherein the RBP is selected from any one of the RBPs identified in FIG.2A,wherein the antisense RNA binds to the target region proximal to the PAS, and wherein the RBP activates or inhibits PAS selection of the mRNA in the cell.

20. A method of regulating polyA site (PAS) selection on mRNA in a cell, the method comprising contacting the cell with the vector of any of claims 11-14, wherein the antisense RNA binds to the target region proximal to the PAS, and wherein the RBP activates or inhibits PAS selection of the mRNA in the cell.

21. A method of regulating polyA site (PAS) selection on mRNA in a cell, the method comprising contacting the cell with the antisense RNA of any one of claims 1-6 or the vector of any of claims 7-10, wherein an endogenous RBP binds to the aptamer, and wherein the antisense RNA binds to a target region proximal to the PAS, and wherein the RBP activates or inhibits PAS selection of the mRNA in the cell.

22. The method of any of claims 18-20, wherein the target mRNA is a gene implicated in haploinsufficiency diseases.

23. The method of claim 22, wherein the haploinsufficiency disease is Dravet syndrome, autosomal dominant retinitis pigmentosa, or neurofibromatosis type 1.

24. The antisense RNA of Table 5.

25. A method of screening for RBPs that modulate PAS selection, comprising:a) providing a dual-luciferase bicistronic reporter comprising a first luciferase, a poly(A) signal (PAS) of interest, and a second luciferase;b) incorporating RNA hairpin aptamers upstream or downstream of the PAS for RBP tethering via a coat protein fusion;c) expressing in cells candidate RBP-coat protein fusions; andd) determining a luminescence ratio of the first luciferase to the second luciferase, wherein a change relative to a control indicates the RBP modulates PAS selection.-169- 4918-7893-0315.1Atty. Dkt. No.: 114198-376026. The method of claim 25, wherein the reporter is an upstream reporter with the aptamer located about 15 nt upstream of the PAS, or a downstream reporter with the aptamer located about 56 nt downstream of the PAS.

27. The method of claim 25 or 26, wherein the PAS is an L3 adenovirus major late transcript PAS lacking UGUA motifs.

28. The method of any one of claims 25-27, wherein the coat protein is MS2 coat protein (MCP).

29. A method of identifying high-confidence PAS activators, comprising performing the method of any one of claims 25-28 in a first screen of a library of RBP candidates, performing the method again in a second screen of candidates positive in the first screen, and designating as high-confidence those RBPs showing consistent activation in both screens with reporter-position specificity.

30. A pool of RBPs identified according to claim 29, wherein the pool comprises CPSF5, CPSF6, RNPS1, GRB2, MBNL1, MBNL2, YTHDF1, PCBP1, SCAF8, HNRNPF, and HNRNPH2.

31. A method of predicting RBP function in PAS selection, comprising:a) inputting amino acid sequences of RBPs into a fine-tuned protein language model trained on screening results of the method of any one of claims 25-28;b) classifying the RBPs as activators or non-activators based on a model score; and c) optionally outputting occlusion maps indicating sequence regions important for classification.

32. The method of claim 31, wherein the protein language model is ProteinBERT fine-tuned on tethered function assay data.

33. The method of claim 31 or 32, further comprising validating model predictions by MS2 tethering assays.

34. A method of mapping domain-level contributions to PAS activation in a candidate RBP, comprising:a) generating deletion variants or isolated domain constructs of the RBP fused to a coat protein;-170- 4918-7893-0315.1Atty. Dkt. No.: 114198-3760b) testing PAS activation by the method of any one of claims 25-28; andc) associating changes in activity with the presence or absence of specific domains.

35. The method of claim 34, wherein the RBP is GRB2 and the domains comprise N-terminal SH3, SH2, and C-terminal SH3 domains.

36. The method of claim 34, wherein the RBP is RNPS1 and the domains comprise an N-terminal intrinsically disordered region (IDR), an RNA recognition motif (RRM), and a C-terminal RS-rich region.

37. A method of determining direct interactions between an RBP and cleavage / polyadenylation (CPA) factor(s), comprising:a) purifying CPA subcomplexes;b) performing an in-vitro binding assay with the RBP or RBP fragment; andc) detecting specific subcomplex binding.

38. The method of claim 37, wherein GRB2 binds directly to CFIm complexes containing CPSF6 or CPSF7.

39. The method of claim 37, wherein RNPS1 binds directly to mPSF and CPSF6.

40. A method of modulating endogenous PAS selection, comprising delivering to a cell: a) an RBP-coat protein fusion; andb) a small nuclear RNA (snRNA) engineered to include an aptamer binding site for the coat protein and an antisense sequence targeting a PAS of interest.

41. The method of claim 40, wherein the snRNA is a U7smOPT snRNA.

42. The method of claim 40 or 41, wherein the antisense sequence targets a region about 60 nt upstream of the PAS.

43. The method of any one of claims 40-42, wherein the RBP is a high-confidence activator identified in claim 29.

44. A composition for modulating PAS selection in a target transcript, comprising: a) a minimal effector domain of an RBP identified as an activator by claim 29; and b) a programmable RNA targeting module configured to bind near a PAS.-171- 4918-7893-0315.1Atty. Dkt. No.: 114198-376045. The composition of claim 44, wherein the minimal effector domain is the C-terminal SH3 domain of GRB2 or the N-terminal IDR of RNPS1.

46. A method of treating a disease associated with aberrant APA, comprising administering to a subject an effective amount of the composition of claim 44 or 45 to restore or alter PAS usage in one or more target transcripts.

47. The method of claim 46, wherein the disease is selected from cancer, Dravet syndrome, autosomal dominant retinitis pigmentosa, and neurofibromatosis type 1.

48. A method of predicting protein regulators of polyadenylation site (PAS) selection, comprising:(a) obtaining amino acid sequences of a plurality of RNA-binding proteins (RBPs);(b) labeling each of the sequences as an activator or a non-activator of PAS selection based on experimental results of a tethered function assay;(c) inputting the sequences and corresponding labels into a sequence-based machine learning model;(d) training the machine learning model to classify proteins as activators or non-activators of PAS selection; and(e) outputting one or more predictions for an RBP of unknown PAS-activator status.

49. The method of claim 48, wherein the model comprises one or more of: a support vector machine, a convolutional neural network, or a protein language model.

50. The method of claim 49, wherein the protein language model is ProteinBERT.

51. The method of claim 49, further comprising fine-tuning the model with class weighting and hyperparameter optimization.

52. The method of claim 49, wherein the sequences are trimmed at each terminus prior to training.

53. The method of claim 49, further comprising interpreting the trained model by analyzing attention weights and / or generating occlusion maps to identify amino acid regions contributory to classification.-172- 4918-7893-0315.1Atty. Dkt. No.: 114198-376054. The method of claim 53, wherein the contributory regions correspond to protein domains or low-complexity regions.

55. The method of claim 49, further comprising filtering the amino acid sequences to remove pairs of proteins having at least 80% sequence identity before splitting into training and test datasets.

56. The method of claim 55, wherein the machine learning model is validated on a set of zinc-finger proteins not included in the training dataset.

57. The method of claim 55, wherein the validation achieves a recall of at least 0.9 and an accuracy of at least 70%.

58. The method of claim 56 or 57, further comprising generating interpretive occlusion maps for a predicted activator from the zinc-finger protein set, wherein the occlusion maps identify functional domains associated with PAS regulation.

59. The method of claim 58, wherein the predicted activator is BRCA1 and the functional domains comprise a RING domain and a BRCT domain.-173- 4918-7893-0315.1