Splice-modulating antisense oligonucleotides

A systematic approach using deletion mutant expression cassettes and computational models predicts and validates the effects of mutations on splicing, addressing the inefficiencies in existing AON prediction methods and enabling precise modulation of splicing.

WO2025172378A1PCT designated stage Publication Date: 2025-08-21FUNDACIO CENTRE DE REGULACIO GEN MICA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/053759
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-13
Filing Date
2025-02-12
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing methods for predicting the impact of antisense oligonucleotides (AONs) on splicing are laborious and inefficient due to the challenge of identifying effective AONs for modulating splicing, as they require extensive testing of various designs, and the effects of insertions and deletions (indels) on splicing have not been systematically tested.

Method used

A systematic system and methods for predicting and validating the effects of mutations on splicing efficiency, involving the use of deletion mutant expression cassettes, computational models, and antisense oligonucleotides to modulate splicing, including the development of a database library and processor for generating and scoring deletion mutants.

Benefits of technology

Enables accurate prediction and identification of antisense oligonucleotides that modulate splicing efficiency, reducing the need for extensive testing and providing a comprehensive understanding of splicing perturbations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025053759_21082025_PF_FP_ABST
    Figure EP2025053759_21082025_PF_FP_ABST
Patent Text Reader

Abstract

The disclosure relates to systematic methods to enable the comprehensive analysis, assessment and / or quantification of the effects of mutations on splicing efficiency and perturbations. Also described are methods for the prediction of antisense oligonucleotide sequences that may affect splicing efficiency of target exons, which may enable the identification of antisense oligonucleotides having therapeutic or other beneficial activities in vivo and / or in vitro. The disclosure also relates to the identification of splicing silencing and enhancing sequences.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SPLICE-MODULATING ANTISENSE OLIGONUCLEOTIDES

[0002] Field of the invention

[0003] This invention relates to generation of oligonucleotides having desirable properties, and to methods of generating oligonucleotides. In particular, the invention relates to a method and system for generating antisense oligonucleotides that modulates splicing.

[0004] Background of the invention

[0005] Pre-mRNA splicing is the process by which introns are removed from transcripts and exons are joined together to form mature mRNAs that are then exported to the cytoplasm and translated. Altered splicing is an important mechanism by which genetic variants cause disease. Multiplex assays of variant effects (MAVEs) have revealed that random nucleotide substitutions in exons frequently affect splicing, with 60-70% of substitutions in over 90% of positions in alternatively- spliced exons and 5% of substitutions in constitutively-spliced exons altering exon inclusion. Comprehensive testing has also shown that ~10% of disease-causing missense variants, as well as 3% of common exonic substitutions, affect splicing. For example, in humans, FAS exon 6 encodes a transmembrane helix of the FAS / CD95 receptor. Inclusion of this exon generates a pro- apoptotic receptor whereas exon skipping produces an anti-apoptotic soluble inhibitor. A variant in exon 6 that causes exon 6 skipping causes autoimmune lymphoproliferative syndrome (ALPS).

[0006] The impact of genetic variation beyond substitutions on splicing have been far less studied. Insertions and deletions (indels) are abundant variants evident in 24% of Mendelian diseases, and disease-causing indels are enriched close to splice sites. However, the effects of indels on splicing have not been systematically tested.

[0007] The frequent disruption of splicing in human disease has led to extensive efforts to therapeutically modulate splicing. In particular, antisense oligonucleotides (AONs) that modulate splicing have been approved as therapies for spinal muscular atrophy and Duchenne muscular dystrophy. Indeed AONs - because of their programmable sequence specificity - may represent a general strategy to therapeutically modulate many splicing changes. However, predicting the impact of AONs on splicing is very challenging: AONs need to be long to achieve sequence specificity (typically 18 to 21 nts) whereas splicing regulatory elements are short and very poorly mapped genome-wide. Parameters that influence AON efficacy include length, proximity to splice sites and binding energy. In practice, however, identifying effective AONs requires laborious testing of many different designs.

[0008] Therefore, it would be beneficial to have a greater understanding of the effects of mutations on splicing. Hence, it would also be desirable to have a system for the accurate prediction of the effects of mutations on splicing; and new methods involving experimental (e.g. in vitro or in vivo) and / or in silico procedures. Summary of the Invention

[0009] The disclosure relates generally to a systematic system and methods to enable the comprehensive analysis of and, optionally, the assessment and / or quantification of the effects of mutations on splicing efficiency and perturbations. In particular, this disclosure relates to systems and methods for the prediction of antisense oligonucleotide sequences that may affect splicing efficiency of target exons, which may enable the identification of antisense oligonucleotides having research, therapeutic or other beneficial activities in vivo and / or in vitro. The disclosure also relates to the identification of splicing silencing and enhancing sequences.

[0010] In the first aspect, the present invention provides a method for predicting or validating a change in splicing efficiency caused by an antisense oligonucleotide that is adapted to bind to a predefined sequence of a target gene sequence comprising a target exon, and optionally one or more adjacent introns. The method comprises providing a cell or plurality of cells, each comprising one of a library of deletion mutant expression cassettes, each deletion mutant expression cassette adapted to express a deletion mutant comprising a sequence deletion compared to the target gene sequence, wherein the sequence deletion comprises or consists of the predefined sequence of the target gene sequence, and wherein the deletion mutant expression cassette is capable of expressing RNA encoded by the deletion mutant expression cassette in the cell or plurality of cells. The method also comprises incubating the cell or plurality of cells for a predetermined period of time to obtain a cell culture wherein cells of the cell culture express RNA encoded by the deletion mutant expression cassette; isolating mRNA expressed from the deletion mutant expression cassette from the cell culture; and assessing the level of the target exon that is incorporated into mRNA expressed from the deletion mutant expression cassette. The method further comprises identifying the predefined sequence that was deleted from the one or more deletion mutants as a target sequence for an antisense oligonucleotide for modulating splicing of the target exon. In embodiments, isolating mRNA may comprise isolating mRNA from a plurality of cells. In embodiments, isolating mRNA may comprise isolating mRNA from all libraries of deletion mutant expression cassettes.

[0011] In embodiments, the invention provides a method further comprising comparing the level of the target exon that is incorporated into mRNA expressed from the deletion mutant expression cassette with the level of the target exon that is incorporated into mRNA expressed from a control expression cassette comprising the wild-type or natural target gene sequence, and / or comparing the level of the target exon that is incorporated into mRNA expressed from the deletion mutant expression cassette with the level of the target exon that is incorporated into mRNA expressed from a control expression cassette comprising the wild-type or natural target gene sequence in the presence of an antisense oligonucleotide complementary to the predefined sequence deleted from the target gene sequence.

[0012] In another aspect, the invention provides a method for identifying an antisense oligonucleotide that modulates splicing of a target exon in a gene, the method comprising: identifying a gene sequence comprising a target exon and optionally one or more adjacent introns, generating a plurality of deletion mutants of the gene sequence, wherein each of the plurality of deletion mutants lacks a predefined sequence of one or more nucleotides of the gene sequence. The method also comprises scoring a plurality of the deletion mutants by applying a computational model to extract features indicative of disruptions on splicing acceptor and / or donor sites, applying the at least one extracted feature to score a change in splicing efficiency for each of the plurality of deletion mutants, where the splicing efficiency is defined as a degree to which the target exon is included in transcript molecules generated from the gene of interest. The method further comprises selecting one or more deletion mutants having a score above a predefined threshold score, and identifying the predefined sequence that was deleted from the one or more deletion mutants as a target sequence for an antisense oligonucleotide for modulating splicing of the target exon.

[0013] In another aspect, there is provided a system for generating antisense oligonucleotides that modulate splicing, comprising: a database library that includes one or more sets of data, where each set comprises a gene sequence comprising a target exon and / or intron adjacent to the target exon, and a plurality of predefined sequences to be deleted from the target exon and / or intron adjacent to the target exon, in order to define a plurality of deletion mutants of the gene sequence. The system also comprises of a processor in communication with the database, the processor being adapted to receive a set of data from the database, the set of data defining a target exon and / or intron adjacent to the target exon and the plurality of predefined sequences. The processor generates a plurality of deletion mutants of the target exon and / or intron adjacent to the target exon, wherein each of the plurality of deletion mutants lacks a predefined sequence of the target exon and / or intron adjacent to the target exon. The processor further generates a dataset comprising of the target exon and / or intron adjacent to the target exon and each of the deletion mutants. The processor also scores a plurality of the deletion mutants by optionally defining a training set comprising of the target exon and / or intron adjacent to the target exon and a subset of the plurality of deletion mutants, applying a computational model to extract features indicative of disruptions on splicing acceptor and / or donor sites within the target exon and / or adjacent intron, and applying the at least one extracted feature to score a change in splicing efficiency for each of the plurality of deletion mutants, where the splicing efficiency is defined as a degree to which the target exon is included in transcript molecules generated from the gene sequence. The processor may select one or more deletion mutants having a score above a predefined threshold score. The system may further comprise identifying the predefined sequence that was deleted from the one or more deletion mutants as a target sequence for an antisense oligonucleotide for modulating splicing of the target exon in the gene sequence.

[0014] Embodiments of any or these aspects may comprise generating a dataset comprising the gene sequence and each of the deletion mutants and / or a plurality of deletion mutants of a plurality of target exons and / or introns across the genome, wherein the dataset is obtained computationally or non-computationally; and providing the dataset to the computational model. In embodiments of any aspect, features extracted from the computational model comprise at least one of: splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain.

[0015] In embodiments of any aspect, applying the at least one extracted feature to score a change in splicing efficiency for each of the plurality of deletion mutants comprises taking the maximum score from the at least one extracted feature for each deletion mutant as a score for the change in splicing efficiency for that deletion mutant.

[0016] In embodiments of any aspect, DANGO score can be calculated for each deletion mutant, wherein the DANGO score is the maximum score taken from the individual scores for one or more of splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain, wherein the scores for splice site acceptor loss and splice site donor loss are first multiplied by -1 to generate a negative score. In embodiments, the DANGO score for each deletion mutant is the maximum score taken from the four individual scores for splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain, wherein the scores for splice site acceptor loss and splice site donor loss are first multiplied by -1 to generate a negative score.

[0017] In embodiments of any aspect, the number of nucleotides deleted from the target exon and / or intron of each of the deletion mutants is from about 1 to 60 nucleotides, from about 2 to 50 nucleotides, from about 3 to 40 nucleotides, from about 4 to 30 nucleotides, from about 5 to 25 nucleotides, from about 7 to 24 nucleotides, from about 10 to 23 nucleotides, from about 15 to 22 nucleotides, or from about 18 to 21 nucleotides. In embodiments the length of the deletion mutant includes at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 99% of a full wild-type exon and / or intron sequence of the gene or gene sequence. In embodiments, the target gene sequence includes at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 100% of a full wild-type exon and / or intron sequence of a gene of interest.

[0018] Embodiments of any aspect may comprise correlating a selected one or more deletion mutant with a corresponding antisense oligonucleotide that is capable of at least partially hybridising to the predefined sequence that was deleted from the selected one or more deletion mutants.

[0019] According to embodiments of any aspect, the methods and systems may comprise generating, providing or designing one or more antisense oligonucleotides that comprises or consists of a nucleotide sequence at least partially complementary to the predefined sequence that was deleted from one or more deletion mutants; and / or one or more antisense oligonucleotides that comprises a nucleotide sequence complementary to the predefined sequence or part thereof; and / or one or more antisense oligonucleotides that capable of hybridising to the predefined sequence or part thereof.

[0020] In embodiments of any aspect, a splicing efficiency is measured in percent-spliced-in-value.

[0021] In embodiments of any aspect, the location of the predefined sequence within the target exon and / or adjacent intron is unique for each of the plurality of deletion mutants.

[0022] In some embodiments, the predefined sequence includes at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 100% of the splicing silencer sequences; or the predefined sequence includes at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 100% of the splicing enhancer sequences.

[0023] It will be appreciated that any features of one aspect or embodiment of the invention may be combined with any combination of features in any other aspect or embodiment of the invention, unless otherwise stated, and such combinations are envisaged and are intended to be directly and unambiguously disclosed herein, and to fall within the scope of the present invention.

[0024] Brief Description of the Drawings

[0025] The invention is further illustrated by the accompanying drawings in which:

[0026] Figure 1 Deep indel mutagenesis of FAS exon 6. A. Plasmid library design. B. Experimental protocol of the massively-parallel splicing assay (MPSA). C. Correlation between PSI values from individual transfections and M PSA-derived PSI values. D. Distribution of PSI values for variants with 1 -nucleotide (nt) mutations. E. Heatmaps displaying inclusion levels of 1-nt substitutions, deletions, and insertions. F. Correlation between PSI values for 1-nt deletions and substitutions at the same position. G. Correlation between PSI values for 1-nt insertions after a given position and substitutions at that position. H. Correlation between PSI values for 1-nt insertions after a given position and 1-nt deletions at the same position.

[0027] Figure 2 The effects of short indels on the inclusion of FAS exon 6. A. Heatmap showing inclusion values for short deletions (1-9 nt in length). B. Heatmap showing inclusion values for 2- nt insertions. C. Heatmap showing inclusion values for 3-nt insertions. D. Correlation between 2 (or 3) nt sequences ranked by the median PSI of exon variants in the library, with this sequence inserted, and their rank based on median PSI of exons containing 30 (or 20) such 2mers (or3mers) in their sequence (GTEx adipose tissue). E. Distribution of PSI values in exons with a 3-nt insertion compared to the frequency of each nucleotide in the insertion. F. Distribution of PSI values in cassette exons (GTEx adipose tissue) relative to the percentage of each nucleotide present in the exon sequence.

[0028] Figure 3 Deep indel mutagenesis of FAS exon 6 reveals the origin of novel microexons. A. Relationship between exon length and inclusion (black curve represents a constrained Bspline fit to rolling median PSI values, and the yellow shaded area indicates the rolling interquartile range of PSI values). B. Distribution of cassette exon inclusion values versus exon length (GTEx adipose tissue). C. Sequence of FAS exon 6 and surrounding intronic sequences. D. Hypothetical mechanisms explaining how the experimental assay could result in the detection of a microexons in the mutant library. E. RT-PCR analysis of FAS exon 6 inclusion for the WT exonic sequence, two variants with insertions that introduce an AG after exon position 40 and one variant with an insertion after exon position 40 that does not introduce an AG dinucleotide. F. Illustration showing the impact of a weak or strong 3’ splice site before FAS exon 6, by the splicing machinery, of the exonic 3’ splice-site-like sequence in the central region of the exon (followed by an AG insertion).

[0029] G. RT-PCR analysis of FAS exon 6 inclusion in the presence of a weak or strong 3' splice site, for the WT exon and an exon with an insertion introducing an AG dinucleotide after exon position 40.

[0030] H, I. RT-PCR analysis of FAS exon 6 inclusion upon overexpression of PTB (WT sequence and variant with an insertion introducing an AG dinucleotide after exon position 40). J, K. RT-PCR analysis of FAS exon 6 inclusion upon overexpression of SRRM4 (WT sequence and variant with an insertion introducing an AG dinucleotide after exon position 40). L. Spliceosome assembly assays using the indicated fluorescently-labelled RNAs (wild type or mutant Fas exon 6 sequences) and HeLa nuclear extracts. The position of complexes assembling U2 snRNP (A complex) and hnRNP proteins (H complex) are indicated. M. Quantification of the ratio between A and H complexes for wild type (WT) and 3’ ss containing (AUG) mutant RNAs as in L. p-value corresponds to a 2-tailed t-test from 20 replicates with WT RNA, and 19 replicates with UAG- containing RNA. N. Distributions of PSI values for double-nucleotide substitutions targeting (i) neither of the putative exonic branchpoint adenines; (ii) the second putative exonic branchpoint adenine as well as another nucleotide in the exon; (iii) the first putative exonic branchpoint adenine as well as another nucleotide in the exon; (iv) both putative exonic branchpoint adenines. P values correspond to 2-tailed Wlcoxon tests with 16230 data points in the CTAACT group, 544 data points in the CTAXCT group, 549 data points in the CTXACT group, and 9 datapoints in the CTXXCT group. Boxplot boxes represent the median, the interquartile range, and the boxplot whiskers extend up to 1 .5 times the interquartile range. O. Distribution of PSI values for insertions that either introduce (left) or do not introduce (right) an AG dinucleotide after exonic position 40. p-value corresponds to a 2-tailed Wlcoxon test with 60 data points in the “no” group and 29 data points in the “yes” group. Boxplot boxes represent the median, the interquartile range, and the boxplot whiskers extend up to 1.5 times the interquartile range.

[0031] Figure 4 The inclusion of short alternative exons encoding one-pass transmembrane helices is regulated by exonic 3’ splice site-like sequences. A. Amino acid category encoded by each codon, categorized by the number of pyrimidines in the codon. B. Hydropathy score of exons categorized into different length groups, divided by whether they have an exonic sequence more or less similar to a 3’ splice site than FAS exon 6. C. Inclusion of exons categorized into different length groups, divided by whether they have an exonic sequence more or less similar to a 3’ splice site than FAS exon 6. D. Model for the regulation of alternative exons encoding a one-pass transmembrane helix. E. Hypothesis suggesting that 3’ splice site-like sequences in alternative exons encoding transmembrane helices promote exon skipping. F. RT-PCR analysis of the inclusion of two alternative exons (CHODL exon 5 and CXADR exon 6) that each encode a one- pass transmembrane helix, including wild-type (WT) sequences and sequence variants with the putative branchpoint adenine mutated.

[0032] Figure 5 Deep learning predicts the inclusion of variants in deep indel mutagenesis library.

[0033] A. Correlations between the inclusion levels of all variants in library and their inclusion levels according to five different predictors (Substitutions n = 189, Deletion n = 1985, Insertion n = 5744).

[0034] B. Lower triangle: Heatmap displaying inclusion levels of all deletion variants in the mutant library. Upper triangle: Heatmap showing inclusion levels of all deletion variants in the mutant library as predicted by SpliceAI. C. Left: Inclusion levels of all 1-6 nt long deletion variants along the sequence of FAS exon 6. The yellow line represents a loess fit with a 95% confidence band. Right: Inclusion levels of the same variants as predicted by SpliceAI.

[0035] Figure 6 A. Distribution of absolute effect sizes of all 4mer deletions in the exome, as predicted by SpliceAI and split by exon PSI groups. B. Hidden Markov model (HMM) with three states (enhancer, silencer, neutral) used to model the splicing regulatory architecture of exons across the genome. C. Predicted regulatory architecture of FAS exon 6 based on the HMM. D. Distribution of exonic splicing enhancer and silencer lengths across the exome, as predicted by the HMM. E. Distribution of the three states of the model in the first and last 20 4mers of all exons under 100 nucleotides long. F. Ternary plot illustrating the relative composition of E / N / S states along the sequences of all 18,551 exons in the dataset. The colour of each point corresponds to the exon’s length. G. Ternary plot illustrating the relative composition of E / N / S states along the sequences of 18,551 exons. The colour of each point corresponds to the exon’s PSI value.

[0036] Figure 7 A. The activity of a splicing regulatory element (SRE) could be modulated by using an antisense oligonucleotide (AON) to base pair with this region (therefore sterically blocking any proteins that might bind to the SRE) or alternatively by deleting the SRE altogether. B. Correlation between the PSI values of nine FAS exon 6 variants with 21 -nt deletions and the PSI values of WT FAS exon 6 with a 21 -nt AON base pairing to the corresponding regions. Horizontal error bars represent the standard deviation of three replicates. Vertical error bars represent the standard error of the mean in the deep insertion mutagenesis library. C. Correlation between the DANGO scores of the same nine FAS exon 6 variants as in panel B, and the PSI values of the WT FAS exon 6 with a 21-nt AON base pairing to the corresponding regions. Horizontal error bars represent the standard deviation of three replicates. D. Percentage of 21 -nt deletions with DANGO scores above the indicated thresholds as a function of exon length. E. Distribution of DANGO scores as a function of exon length. F. Custom genome browser track displaying the DANGO scores for FAS exon 6. The corresponding PSI values as measured experimentally in the deep mutagenesis assay are shown in the heatmap below, using the same colour scale as Figure 5B.

[0037] Figure 8 Pairwise correlations of enrichment scores for all exon variants in the library across nine experimental replicates.

[0038] Figure 9 Correlation between PSI values for 1-nt insertions before a given position and substitutions at that position.

[0039] Figure 10 Correlation between PSI values for 1-nt insertions before a given position and deletions at that position.

[0040] Figure 11 Distribution of inclusion values for exons with each possible 1-, 2-, and 3-nt insertion.

[0041] Figure 12 Correlation between 2 (or 3) nt sequences ranked by the median PSI of exon variants in the library, with this sequence inserted, and their rank based on median PSI of exons containing 30 (or 20) such 2mers (or 3mers) in their sequence (all GTEx tissues).

[0042] Figure 13 Distribution of PSI values in cassette exons (all GTEx tissues) relative to the percentage of each adenine present in the exon sequence.

[0043] Figure 14 Distribution of PSI values in cassette exons (all GTEx tissues) relative to the percentage of each cytosine present in the exon sequence.

[0044] Figure 15 Distribution of PSI values in cassette exons (all GTEx tissues) relative to the percentage of each guanine present in the exon sequence.

[0045] Figure 16 Distribution of PSI values in cassette exons (all GTEx tissues) relative to the percentage of each uracil present in the exon sequence.

[0046] Figure 17 Distribution of cassette exon inclusion values versus exon length (all GTEx tissues).

[0047] Figure 18 A. Sequences of microexons in the library with detectable levels of inclusion. B. RT-PCR analysis of the inclusion of exons with a sequence corresponding to the detectable microexons. C. RT-PCR analysis of the inclusion of FAS exon 6 with a UAA insertion after exon position 18.

[0048] Figure 19 Validation of spliceosome assembly assays and enhanced A (U2 snRNP containing) complex formation on Fas exon 6 upon inclusion of a 3’ splice site (TAG). Fluorescently-labelled RNAs corresponding to Fas exon 6 (wt), Fas exon 6 with a mutation that creates a 3’ splice site after the internal polypyrimidine tract (TAG) (see Figure 3) or Fas exon 6 flanked by intronic sequences (-68wt+25, which includes the 3’ 68 nucleotides of intron 5 and the 5’ 25 nucleotides of intron 6) were incubated with HeLa nuclear extracts under in vitro splicing conditions at 30°C (+) or on ice as a control (-) and the ribonucleoprotein complexes formed were fractionated by electrophoresis on a composite agarose-polyacrylamide gel. The electrophoretic positions of U2 snRNP-containing complexes (A complex) and hnRNP-containing complexes (H complex) are indicated. Controls of U1 I U2 snRNP inactivation by RNAse H-mediated digestion of the 5’ end of U1 snRNA or of the branch point recognition sequence of U2 snRNA are included. Inactivation of U2 snRNP (but not of U1 snRNP) reduces A complex formation on WT and TAG RNAs, while complexes formed on the -68wt+25 RNA are sensitive to inactivation of either U1 or U2 snRNPs because of exon definition-mediated effects (Izquierdo et al. (2005) Mol Cell 19: 475-484). Note the significant increase in A / H complex ratio upon inclusion of a 3’ splice site (TAG) in Fas exon 6 compared to the wild type sequence (WT).

[0049] Figure 20 A. E / N / S states in the first and last 20 4mers of all exons (100-150 nucleotides in length). B. E / N / S states in the first and last 20 4mers of all exons (over 150 nucleotides in length).

[0050] Figure 21 Graph showing absolute number of nucleotides in the E or S states against the length of the target exon. Data shows that the absolute number of nucleotides in the E or S states rises at a slower rate than would be expected if the proportion of nucleotides in these states remained constant regardless of exon length.

[0051] Figure 22 A. Distribution of the number of enhancer-silencer ("E" / "S") state transitions per 100 nucleotides in all exons shorterthan 101 nt, divided into groups based on exon inclusion levels.

[0052] B. Distribution of the number of enhancer-silencer ("E" / "S") state transitions per 100nucleotides in all exons with a length between 101 and 150 nt, divided into groups based on exon inclusion levels.

[0053] C. Distribution of the number of enhancer-silencer ("E" / "S") state transitions per 100 nucleotides in all exons longer than 150 nt, divided into groups based on exon inclusion levels.

[0054] Figure 23 A. Distribution of the number of enhancer-neutral ("E" / "N") state transitions per 100 nucleotides in all exons shorter than 101 nt, divided into groups based on exon inclusion levels. B. Distribution of the number of enhancer-neutral ("E" / "N") state transitions per 100 nucleotides in all exons with a length between 101 and 150 nt, divided into groups based on exon inclusion levels. C. Distribution of the number of enhancer-neutral ("E" / "N") state transitions per 100 nucleotides in all exons longer than 150 nt, divided into groups based on exon inclusion levels.

[0055] Detailed Description of the Invention

[0056] All references cited herein are incorporated by reference in their entirety. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs (e.g. in cell culture, molecular genetics, nucleic acid chemistry and biochemistry).

[0057] Unless otherwise indicated, the practice of the present invention employs conventional techniques of chemistry, molecular biology, microbiology, recombinant DNA technology, and chemical methods, which are within the capabilities of a person of ordinary skill in the art. Such techniques are also explained in the literature, for example, J. Sambrook, E. F. Fritsch, and T. Maniatis, 1989, Molecular Cloning: A Laboratory Manual, Second Edition, Books 1-3, Cold Spring Harbor Laboratory Press; Ausubel, F. M. et a / . (1995 and periodic supplements; Current Protocols in Molecular Biology, ch. 9, 13, and 16, John Wiley & Sons, New York, N. Y); B. Roe, J. Crabtree, and A. Kahn, 1996, DNA Isolation and Sequencing: Essential Techniques, John Wiley & Sons; J. M. Polak and James O'D. McGee, 1990, In Situ Hybridisation: Principles and Practice, Oxford University Press; M. J. Gait (Editor), 1984, Oligonucleotide Synthesis: A Practical Approach, IRL Press; and D. M. J. Lilley and J. E. Dahlberg, 1992, Methods of Enzymology: DNA Structure Part A: Synthesis and Physical Analysis of DNA Methods in Enzymology, Academic Press. Each of these general texts is herein incorporated by reference.

[0058] In order to assist with the understanding of the invention several terms are defined herein.

[0059] The terms “nucleic acid”, “nucleotide”, “polynucleotide”, and “oligonucleotide” are used interchangeably and refer to a deoxyribonucleotide (DNA) or ribonucleotide (RNA) polymer, in linear or circular conformation, and in either single- or double-stranded form. The shorthand “nt” denotes herein a “nucleotide”.

[0060] The term "sequence" in the context of nucleic acids refers to the specific order of nucleotide bases along the polymer chain. The sequence of adenine (A), thymine (T) (uracil (U) in RNA), cytosine (C), and guanine (G) carries the genetic information in the form of a code. A sequence in relation to an oligonucleotide can refer to one or more contiguous nucleotides or bases.

[0061] The term “nucleotide synthesis” means biochemical process by which nucleotide bases are joined to form a polynucleotide.

[0062] The term "exon" refers to a portion of a gene that is present in the mature form of mRNA. Exons include the open reading frame (ORF), i.e., the sequence which encodes protein, as well as the 5' and 3' untranslated regions (UTRs). The UTRs are important for translation of the protein.

[0063] The term "intron" refers to a portion of a gene that is not translated into protein and while present in genomic DNA and pre-mRNA, it is removed in the formation of mature mRNA.

[0064] The term “predefined sequence” in the context of this application refers to a sequence of nucleotide bases that have been selected. The term "messenger RNA" or "mRNA" refers to RNA that is transcribed from genomic DNA and that carries the coding sequence for protein synthesis. Precursor mRNA (Pre-mRNA) is transcribed from genomic DNA. mRNA typically comprises from 5' to 3'; a 5' cap (e.g. modified guanine nucleotide), a 5' UTR, the coding sequence (beginning with a start codon and ending with a stop codon), a 3' UTR, and a poly-adenosine (polyA) tail.

[0065] As used herein, the term "hybridisation" means the process in which, typically, a single-stranded sequence of nucleotides I oligonucleotide recognises and binds with another single-stranded nucleic acid molecule that is complementary thereto.

[0066] As used herein, the term “binding” in the context of the present disclosure refers to a non-covalent interaction between macromolecules. In some cases, binding will be sequence-specific, such as between one or more specific nucleotides and another one or more specific nucleotides.

[0067] The term “target exon”, as used herein, refers to an exon to which an antisense oligonucleotide of the disclosure is intended to bind, and / or which is the focus of the systems and methods described herein for assessing splicing of and inclusion or exclusion of the exon in an mRNA molecule.

[0068] A nucleic acid “target”, “target site” or “target sequence”, as used herein, is a nucleic acid sequence to which an antisense oligonucleotide of the disclosure is intended to bind, provided that conditions of the binding reaction are not prohibitive. A target site may be a nucleic acid molecule or a portion of a larger polynucleotide. The expression, amount, or activity of the target is capable of being modulated as a result of the interaction between the target and antisense oligonucleotide. As will be appreciated, the target site is typically within a target exon according to this disclosure.

[0069] The term “antisense oligonucleotide” (or AON) refers to a short oligonucleotide or modified oligonucleotide which is antisense to and binds to (i.e. hybridises to) a target region of a polynucleotide, such as a gene transcript, pre-mRNA, mRNA or RNA fragment. According to the disclosure, the target region for an antisense oligonucleotide is suitably in a pre-mRNA, i.e. the antisense oligonucleotide binds to a target region on a transcript before the splicing process takes place. The antisense oligonucleotides of the disclosure may comprise native or modified RNA, DNA, or mixtures thereof. Any modification may be naturally occurring or non-naturally occurring, as described below. In embodiments, examples of (naturally occurring) chemical modifications that may be incorporated on antisense oligonucleotides of the disclosure may include pseudouridine, methyl, N6-methyladenosine, 5-hydroxymethylcytosine and N1 -methyladenosine. Any references to an antisense oligonucleotide sequence provided as an RNA sequence is intended to also encompass the equivalent DNA sequence. In particular, any sequence comprising U bases is intended to refer equally to the corresponding sequence in which T bases are present in place of one or more (e.g. all) of the Us. In embodiments, the antisense oligonucleotides according to the disclosure may comprise nucleotides comprising inosine. In particular, any antisense oligonucleotide sequence disclosed herein is intended to encompass nucleotide sequences where one or more of the T, G or As are replaced with I, as far as this is compatible with the function of the antisense oligonucleotide. The antisense oligonucleotides according to the disclosure have an effect on the regulation of alternative splicing. In embodiments, the antisense oligonucleotides according to the disclosure advantageously promote the skipping of a specific (target) exon by base pairing at the splicing enhancer site, resulting in a decreased inclusion of the exon in the mature transcript resulting from splicing of the pre-mRNA. In embodiments, antisense oligonucleotides of the disclosure can promote inclusion of a specific (target) exon by base pairing at the splicing silencer site, resulting in an increased inclusion of the exon in the mature transcript resulting from splicing of the pre-mRNA. In embodiments, antisense oligonucleotides of the disclosure can promote suppression of the intron excision, which results in the retention of the entire intron. In embodiments, antisense oligonucleotides of the disclosure can promote extension or shortening of exons through the use of alternative 5’ or 3’ splice sites. As the skilled person would understand, the range of lengths of antisense oligonucleotides that are suitable will depend on the desired specificity and / or binding affinity (where longer AONs are expected to bind to their target sequence with higher affinity and / or specificity) and efficacy of the in vivo delivery (where longer AONs are expected to be more difficult to efficiently deliver to the target cells). The antisense oligonucleotides of the disclosure suitably have a length of at most 35 nucleotides, at most 25 nucleotides, at most 24 nucleotides, at most 21 nucleotides, or at most 18 nucleotides. An antisense oligonucleotide as used herein may have a length of at least 7 nucleotides, at least 10 nucleotides, at least 13 nucleotides, at least 14 nucleotides, or at least 17 nucleotides.

[0070] Suitably, an antisense oligonucleotide is complementary to a target site over a predefined contiguous stretch of nucleotides of the target, such as over a contiguous stretch of at least 7 nts, at least 10 nts, at least 12 nts, at least 15 nts, at least 18 nts or at least 21 nts; up to, for example, 30 nts. In embodiments, therefore, an antisense oligonucleotide is complementary to its target site over a contiguous stretch of from 7 to 30 nts, from 10 to 25 nts, from 12 to 24 nts, from 15 to 23 nts, from 18 to 22 nts, or from 19 to 21 nts. In particularly suitable embodiments, an antisense oligonucleotide is complementary to its target site over a stretch of 19, 20, 21 or 22 nts; and still more suitably over 21 nts. Antisense oligonucleotides may typically be of the same length as the region of complementarity with its target site, such that the entire sequence of the oligonucleotide is complementary to the target sequence.

[0071] As used herein, the term “exon variant” refers to an exon substantially similar to the target exon, but having one or more mutation to the wild-type or natural sequence. Said mutation(s) may comprise one or more of a deletion, a substitution and / or an insertion of one or more nucleotides. Suitably, an exon variant as used herein has either a substitution, a deletion or an insertion. The substitution, deletion or insertion is typically at only one location within the (target) exon; and said substitution, deletion or insertion may comprise one or more nucleotide at said location. In embodiments, a deletion may comprise or consist of from 1 to 300 nts, from 1 to 200 nts, from 1 to 100 nts, from 1 to 60 nts, from 2 to 50 nts, from 3 to 40 nts, from 4 to 30 nts, from 5 to 25 nts, from 7 to 24 nts, from 10 to 23 nts, from 15 to 22 nts, or from 18 to 21 nts. In embodiments, a deletion variant may have a deletion of 19, 20, 21 or 22 consecutive nucleotides. In embodiments, the length of the deletion mutant includes at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 100% of a full wild-type exon and / or intron sequence of the gene of interest. In embodiments, an insertion variant may have an insertion of from 1 to 30 nts, from 1 to 25 nts, from 1 to 21 nts, from 1 to 15 nts, from 1 to 12 nts, from 2 to 10 nts, from 3 to 8 nts, or from 4 to 7 nts. In embodiments, an insertion variant may have an insertion of 1 , 2, 3 or 4 nts. In embodiments, a substitution variant may include substitutions of from 1 to 30 nts, from 1 to 25 nts, from 1 to 21 consecutive nts, from 1 to 15 consecutive nts, from 1 to 12 consecutive nts, from 2 to 10 consecutive nts, from 3 to 8 consecutive nts, or from 4 to 7 consecutive nts. In embodiments, a substitution variant may have a substitution of 1 , 2, 3 or 4 consecutive nts.

[0072] “Indel” refers herein to insertions and deletions to a nucleic acid sequence.

[0073] “Modulation” is used herein to refer to a perturbation of function or activity. In embodiments, modulation of exon inclusion, for example, refers to an increase or decrease of inclusion of a target exon.

[0074] “Splicing” in the context of this application means a process wherein the non-coding regions (introns) are removed from a precursor messenger RNA (pre-mRNA), and the remaining coding regions (exons) are joined together.

[0075] The term “alternative splicing” means the process of different combinations of exons joining together, resulting in multiple mRNA variants from a single gene.

[0076] “Exon inclusion” refers to the process during RNA splicing where a particular exon (coding region), e.g. a ‘target exon’ is retained in the mature messenger RNA (mRNA) transcript.

[0077] “Exon skipping”, as used herein, refers to exclusion of exon sequences in the splicing process. Exon skipping can alter the reading frame of the resultant mRNA and can result in the production of a non-functional or truncated proteins.

[0078] “Machine learning”, as used herein, means use of computer systems that are able to learn and adapt without following explicit instructions, by using algorithms and statistical models to analyse and draw inferences from patterns in data.

[0079] A “machine learning model” refers herein to a computational system that is designed to make predictions or decisions based on data. A “training set” is a subset of data used to train a machine learning model.

[0080] A “feature” in the context of machine learning means an individual measurable property or data point relating to splicing of a target exon that can be used as input for a machine learning algorithm.

[0081] “Scoring” in the context of this application means generating values relating to the probability, predictability and / or efficiency of splicing of a target exon based on the trained machine learning model.

[0082] “Splicing efficiency” refers to the effectiveness with which a cell processes pre-messenger RNA (pre-mRNA) during RNA splicing and incorporates a target exon into the final RNA transcript.

[0083] As used herein, the term “percent spliced in (PSI)” refers to the relative rate of exon inclusion between a variant (e.g. deletion mutant) and the corresponding wild-type I natural target exon.

[0084] “Splice acceptor site” refers herein to the nucleotide sequence at the 3' end of an intron to be spliced. The acceptor site typically has a specific consensus sequence that signals the splicing machinery to recognise and accurately remove the intron. Loss or gain of the splice acceptor site sequences modulates splicing, and loss typically leads to disruption of the splicing process.

[0085] The term “splice donor site” refers to the nucleotide sequence at the 5’ end of an intron to be spliced. The splice donor site typically consists of a short, conserved sequence that often begins with the dinucleotide "GT". Loss or gain of the splice acceptor / donor site sequences modulates splicing, and loss typically leads to disruption of the splicing process.

[0086] As known in the art, the antisense oligonucleotides of the disclosure may be chemically modified, for example, to increase their stability, reduce their immunogenicity, increase their binding affinity to a target sequence, reduce their non-specific binding to unintended targets (see e.g. Mou & Gray, 2002) and / or enhance their delivery, etc. Chemical modification of any oligonucleotide refers to a chemical difference compared to the native form of ribo- or deoxyribonucelotides (i.e. nucleotides comprising the naturally occurring nucleobases of RNA or DNA, including (deoxy)adenosine, (deoxy)guanosine, (deoxy)thymidine, (deoxy)cytidine, 5-methyl (deoxy)cytidine and uridine; and a phosphate linking group). Chemical modifications may be applied to the sugar moiety of a nucleoside, the nucleobase moiety of a nucleoside and / or the phosphate backbone. Examples of chemical modification on the sugar moiety of nucleotides or analogues that may be used in the present disclosure include, N3’-P5’ phosphoroamidates (NPs), 2'-fluoro-2'-deoxyadenosine-5'- triphosphate, 2'-fluoro-2'-deoxycytidine-5'-triphosphate, 2'-fluoro-2'-deoxyguanosine-5'- triphosphate, 2'-fluoro-2'-deoxyuridine-5'-triphosphate, 2'-fluoro-2'-deoxythymidine-5'-triphosphate, 2'-0-methyladenosine-5'-triphosphate, 2'-0-methylcytidine-5'-triphosphate, 2'-O- methylguanosine-5'-triphosphate, 2'-0-methyluridine-5'-triphosphate, 2'-0-methylinosine-5'- triphosphate, 2'-0-methyl-2-aminoadenosine-5'-triphosphate, 2'-0-methylpseudouridine-5'- triphosphate, 2'-0-methyl-5-methyluridine-5'-triphosphate etc. Examples of chemically modified nucleotides or analogues that may be used in the present disclosure include locked nucleic acids, the use of 2’-O-methylated nucleotides, 2’-0-methoxyethyl modified (2’MOE) nucleotides, methylated cytosine, constrained ethyl (cET) nucleic acids, bridged nucleic acids (BNAs), phosphorodiamidate morpholino oligomers (PMOs), peptide nucleic aids (PNAs), cyclohexene nucleic acids (CeNAs), and tricycle-DNA (tcDNA). Examples of chemical modification on the phosphate backbone of nucleotides or analogues that maybe used in the present disclosure include the use of phosphorothioate backbone, methylphosphonate backbone, and Borano- phosphate backbone.

[0087] Chemical modifications may be applied to one or more nucleotides of an oligonucleotide and, as such, reference to a “modified (antisense) oligonucleotide” or “modified AON” relates to an oligonucleotide comprising at least one modified nucleotide. In embodiments, 2’-0 modified nucleotides may be used to surround (i.e. flank) a central sequence, thereby forming a “gapmer”, as known in the art. Such compounds may be particularly resistant to nuclease degradation. As the skilled person would understand, multiple modifications may be present on the same nucleotide. Further, different modifications may be present on different nucleotides of the same oligonucleotide. As known in the art, the optimal length of an antisense oligonucleotide may depend on the modifications (or absence thereof) that may be present on the antisense oligonucleotide.

[0088] A “locked nucleic acid” (LNA) as used herein refers to a ribonucleic acid where at least one of the nucleotides has the ribose moiety modified with a methylene bridge connecting the 2’ oxygen and the 4’ carbon. Without wishing to be bound by theory, it is believed that this locks the ribose ring in an ideal conformation for Watson-Crick base-pairing, making the pairing of a locked nucleotide with a complementary nucleotide strand more rapid and more stable. A locked nucleic acid may comprise a mixture of locked nucleotides and ribonucleic acids. In embodiments, all of the nucleotides of the oligonucleotides according to the disclosure are locked nucleotides. In embodiments where the antisense oligonucleotide is an LNA, the length of the antisense oligonucleotide may be particularly between about 1 to 60 nucleotides, between about 2 to 50 nucleotides, about between 3 to 40 nucleotides, between about 4 to 30 nucleotides, between about 5 to 25 nucleotides, between about 7 to 24 nucleotides, between about 10 to 23 nucleotides, between about 15 to 22 nucleotides, or between about 18 to 21 nucleotides. In some embodiments, the length of the antisense oligonucleotide may be between about 7 and 17 nucleotides. In some such embodiments, the length of the antisense oligonucleotide may be between about 10 and 15 nucleotides.

[0089] As used herein a “phosphorothioate (ribo)nucleic acid” (or oligonucleotide phosphorothioate) refers to a modified (ribo)nucleic acid in which one of the oxygen atoms in the phosphate moiety is replaced by sulphur, wherein the oxygen that is replaced by sulphur is at a non-bridging position. Without wishing to be bound by theory, it is believed that because phosphorothioate (ribo)nucleic acids are non-natural analogs of nucleic acids, oligonucleotide phosphorothioates are substantially more stable with respect to hydrolysis by nucleases, the class of enzymes that destroy nucleic acids by breaking the bridging P-0 bond of the phosphodiester moiety. An oligonucleotide phosphorothioate may comprise a mixture of modified and unmodified nucleotides. In embodiments, the whole backbone of the oligonucleotide of the disclosure is modified, and each nucleotide may be a phosphorothioate nucleotide.

[0090] Antisense oligonucleotides according to the disclosure may comprise 2'-O-methylated ribonucleotides. 2'-O-methylated ribonucleotides comprise a methyl group added to the 2' hydroxyl of the ribose moiety of the nucleotide, producing a methoxy group. In embodiments, Antisense oligonucleotides according to the disclosure may comprise one or more 2'-O-methylated ribonucleotides. In embodiments, all of the nucleotides of the antisense oligonucleotide may be 2'- O-methylated ribonucleotides.

[0091] Antisense oligonucleotides according to the disclosure may comprise 2'-0-methoxyethyl ribonucleotides. 2'-0-methoxyethyl ribonucleotides comprise a methoxyethyl group (- CH2CH2OCH3) added to the 2’ hydroxyl of the ribose moiety of the nucleotide. In embodiments, antisense oligonucleotides according to the disclosure may comprise one or more 2'-O- methoxyethyl ribonucleotides. In embodiments, all of the nucleotides of the antisense oligonucleotide may be 2'-O- methoxyethyl ribonucleotides.

[0092] Antisense oligonucleotides according to the disclosure may comprise phosphorodiamidate morpholino modification. Phosphorodiamidate morpholino comprises a morpholine ring in place of the furanose ring found in natural nucleic acids and a neutral phosphorodiamidate backbone in place of the negatively charged phosphodiester backbone. In embodiments, antisense oligonucleotides according to the disclosure may comprise one or more Phosphorodiamidate morpholinos oligonucleotides. In embodiments, all of the nucleotides of the antisense oligonucleotide may be phosphorodiamidate morpholino oligonucleotides.

[0093] In some embodiments the antisense oligonucleotides according to the disclosure may comprise a mixture of one or more ribonucleotides modified with a methyl group (e.g. 2'-O-methylated ribonucleotides) and one or more ribonucleotides modified with a methoxyethyl group (e.g. 2'-O- methoxylethyl ribonucleotides).

[0094] In embodiments, the antisense oligonucleotides according to the disclosure may comprise one or more conjugate groups, for example, in order to modify the properties of the oligonucleotide, such as the pharmacodynamics, distribution, stability, binding and / or absorption etc. of the oligonucleotide, as known in the art. In embodiments, the one or more conjugate groups may comprise a ligand that targets the oligonucleotide, for example to a specific type of cells. Thus, in embodiments, the disclosure encompasses compounds comprising one or more antisense oligonucleotide according to the disclosure fused to another chemical entity, such as a pharmaceutical drug. The conjugated entity (or group) may comprise nucleotides or may be a non- nucleotide-based molecule. In embodiments, the conjugated entity may comprise a splicing modifying drug. Examples of splicing modifying drugs are provided in Vigevani, L., & Valcarcel, J. (2012). In embodiments, the drug is selected from the group comprising sudemycins, spliceostatin, pladienolides and meayamycins, or derivatives thereof. In some embodiments, the conjugated entity may be a peptide conjugate, lipid conjugate, polymer conjugate, protein conjugate, small molecule conjugate, or aptamer conjugate. In some embodiments, the conjugated entity may comprise a G-quadruplexe. In some embodiments, the conjugated entity may comprise GalNAC.

[0095] Antisense Oligonucleotides (AONs) in Splicing Modulation

[0096] An antisense oligonucleotide (AON) is a short oligonucleotide or modified oligonucleotide which is antisense to and binds to (i.e. hybridises to) a target region of a polynucleotide, such as a gene transcript, pre-mRNA, mRNA or RNA fragment.

[0097] AONs can downregulate mRNA targets by a variety of mechanisms: induction of RNase H endonuclease activity that cleaves the RNA-DNA heteroduplex, 5' cap formation, modulating of splicing at a splicing site, and steric hindrance of ribosomal activity (Di Fusco et al., (2019) Front. Pharmacol. 10:305). AONs are highly target-specific due to their base-pairing requirements, and are capable of entry into most cell types in the body. AONs are also well tolerated, particularly in the CNS without significant toxicity, and benefically have a long-lasting effect in vivo (Sharma V.K et al., (2015) Future Med. Chem. 7:2221-2242).

[0098] According to the disclosure, the target region for an antisense oligonucleotide is suitably in a pre- mRNA, i.e. the antisense oligonucleotide suitably binds to a target region on a transcript before the splicing process is able to take place.

[0099] Conventionally, the length of AONs used for splicing modulation is typically in the range of about 15 to 30 nucleotides, but can be shorter or longer, depending on the design and other preferences. Suitably, an AON of this disclosure is designed to base-pair with a pre-mRNA and disrupt the normal splicing sites of the transcript, e.g. by blocking RNA-RNA base-pairing or protein-RNA binding interactions that occur between components of the splicing machinery and the pre-mRNA. AON base-pairing to a target RNA can alter the recognition of splice sites by the spliceosome, which in turn may lead to an alteration of normal splicing of the targeted transcript. In splicing, a target RNA region suitable for AON binding can be a sequence of an exon and / or a sequence which includes one or more splicing regulatory element. Nucleotides of an AON involved in splicing modulation may beneficially be chemically modified, as described herein, so that the RNA-cleaving enzyme RNase H is not recruited to degrade the pre-mRNA-AON complex. Thus, AONs may modify splicing without necessarily altering the abundance of the mRNA transcript. RNAse Fl- resistant features of AONs can be particularly beneficial where the intention is to employ an AON capable of altering splicing efficiency and not to cause the degradation of the bound pre-mRNA (Havens et al., 2016 Nucleic Acids Res. 2016 44(14): 6549-6563), which would also affect protein expression levels. Without wishing to be bound by theory, there are several ways in which modulation of splicing can be achieved by using AONs. In some embodiments, an AON may target and base-pair with an RNA sequence at a splicing enhancer or silencer site, resulting in interference of splicing regulatory protein interactions at the enhancer I silencer site, and inhibition I promotion of spliceosome assembly at the intronic splice sites flanking the exon and at a splicing enhancer or silencer site within an exon. In other embodiments, an AON may base-pair with splice site sequences to create a steric block to the binding of the spliceosome to the splicing site. Both of these forms of interaction are expected to modulate splicing outcomes, in the first case resulting in either inhibition or activation of splicing (i.e. exon skipping or inclusion), in the second resulting in inhibition of the inclusion of the target exon in the final mRNA product, i.e. exon skipping.

[0100] On the other hand, splicing I splicing efficiency can be enhanced I promoted by targeting and binding of an AON at a splicing silencer site. In this case, it is expected that preventing or inhibiting the action of a silencer sequence would allow the target exon to be included in the final mRNA product with greater efficiency. This process is known as exon inclusion.

[0101] Antisense Oligonucleotides (AONs) in Therapeutics

[0102] As described herein, AONs of the disclosure can be used to alter splicing efficiency to either increase or decrease the amount of a target exon that is incorporated into a gene of interest.

[0103] For example, in embodiments, inhibition of a splice site by an AON can offer a mechanism to block a splice site where it is desirable to reduce the amount of a particular target exon that is incorporated into a gene of interest. In embodiments, the splice site may be a cryptic splice site, e.g. which has been created by a genetic mutation, such that correct splicing or altered splicing is enabled, which in turn may restore the original version of a protein, or which may reduce the amount of a pathogenic protein. By way of example, thalassemia is caused by a splicing mutation in the human (3-globin gene that creates a cryptic 5' splice site, which is used preferentially over the natural, wild-type splice site. It has been shown that an AON designed to base-pair to the region encompassing the cryptic splice site can redirect splicing to the correct splice site (Dominski et al., (1993) Proc. Natl. Acad. Sci. U.S.A. 90:8673-867).

[0104] In embodiments, splicing inhibition can be used to restore the reading frame of an mRNA by skipping an exon that has a premature termination codon caused e.g. by a frameshift. In these circumstances, removing an exon from the mRNA may restore the wild-type mRNA reading frame. Duchenne Muscular Dystrophy is caused by exon skipping that disrupts the mRNA reading frame and creates premature termination codons that produce truncated and usually non-functional dystrophin protein. An AON designed to skip Exon 19 of dystrophin gene, and a cocktail of AONs designed to skip multiple exons for reading frame correction, have been shown to be effective at restoring wild-type gene expression in mice (Aoki et al., (2012) Proc. Natl. Acad. Sci. U.S.A.109: 13763-13768) . In embodiments, AONs can promote exon inclusion, and e.g. promote production of a target mRNA transcript with the result of increasing target protein production. Spinal Muscular Atrophy (SMA) is caused by insufficient production of the protein SMN, a highly conserved and ubiquitously expressed protein involved in pre-mRNA splicing. In humans, the SMN protein is produced from two different genes, SMN1 and SMN2. SMA patients do not have a functional copy of the SMN1 gene. The level of SMN2 protein production is low in both patients and non-patients, but a higher SMN2 copy number in SMA patients is correlated with a better prognosis. Skipping of exon 7 is the major cause of low SMN2 production. AONs may be designed to target exon 7 splicing and promote its inclusion to increase SMN2 production.

[0105] Antisense Oligonucleotide (AON) Design

[0106] Traditionally, potentially effective AONs have been developed using a trial-and-error approach, which involves speculative design followed by experimental AON activity validation. The approach is inefficient and sub-optimal, as on average, one out of ten 20nt AONs achieves modulation of the target mRNA level. The selected AON may then be chemically modified in a speculative manner (Bennett and Cowsert, (1999) CurrOpin Mol Ther 1 (3):359-71) to improve its efficacy in vivo. Other methods for identifying useful AONs include the random selection of oligomers out of all possible AON candidates. Generally, only a very small fraction of randomly selected AONs may be active. (Monia et al., (1996) J Biol Chem 271 (24): 14533-40). The traditional random selection method for AONs has relied on testing experimentally as many oligonucleotides as possible.

[0107] There are multiple factors that introduce variability and unpredictability to AON design. For instance, the length of an AON significantly impacts the binding affinity and functions of the AON. AONs longer than 20nts are considered to have a “length penalty”, which negatively impacts uptake efficiency and potency. Longer AONs are also more prone to self-dimerization or the formation of hairpin structures that can impede hybridization to their target polynucleotide. On the other hand, shorter AONs, e.g. of below about 10nts typically have low affinity for the target site. Furthermore, AONs under 18 nts may have increasing off-target effects. For example, 12nt AONs are, on average, predicted to have hundreds of perfectly matched off-target sites within a complex genome, some of which might cause unwanted off-target actions.

[0108] Another factor that should be taken into account when designing AONs is the secondary structure of the target polynucleotide (e.g. RNA), since secondary structure effects within the target can affect the accessibility of the AON binding site. If the AON is required to bind to a sequence within a complex secondary structure of a target RNA, AON design may require modifications or adjustments to ensure effective binding. Other factors include the thermodynamic stability of the AON, off-target effects, cellular uptake and intracellular localization, immunogenicity and stability of the AON. Advantageously, the AONs of the present disclosure are selected as those that are determined to provide a good likelihood of efficacy, based on any desired variables, including sequence length and context within a target polynucleotide or secondary structure, thereby improving over prior art ‘hit-and-miss’ processes for traditional AON design and selection, as described herein.

[0109] Splice-Modulating Antisense Oligonucleotide Design

[0110] The present discloses provides, in various aspects and embodiments, a novel method of identifying an AON that modulates splicing of a target exon in a gene of interest, by generating a dataset comprising of a plurality of deletion mutants of a plurality of target exons across the genome, scoring a plurality of the deletion mutants, identifying at least one target exon, selecting one or more deletion mutants having a score above a predefined threshold score, and identifying the predefined sequence that was deleted from the one or more deletion mutants as a target sequence for binding by an antisense oligonucleotide that may modulate splicing of the target exon. In aspects and embodiments, a target exon may be predetermined rather than being selected on the basis of a genome-wide (or partial genome) assessment. In various aspects and embodiments, the method may be an in silico method, in other aspects and embodiments the method may be an empirical I manual method. In other aspects and embodiments, the method may comprise a mixture of both in silico and empirical I manual processes. In aspects and embodiments, the method may be directed to a target exon. In other aspects and embodiments, the method may be directed to a target exon and one or more adjacent introns. In some alternative aspects and embodiments, the method may be directed to a target intron.

[0111] Further, the present disclosure relates to a method for validating a predicted change in splicing efficiency caused by an AON that is adapted to bind to a target sequence within a target exon and / or intron, the method comprising: providing a cell or cells comprising an AON that are designed for a target exon and / or intron, or an expression cassette comprising a deletion mutant, incubating the cell or cells for a predetermined period of time to obtain a cell culture wherein cells of the cell culture express RNA encoded by the deletion mutant of the target exon, isolating RNA comprising the deletion mutant of the target exon from the cell culture, assessing the level of the target exon that is incorporated into mRNA expressed from the expression cassette comprising the deletion mutant of the target exon and / or intron, and comparing the level of the target exon that is incorporated into mRNA expressed from the expression cassette comprising the deletion mutant of the target exon and / or intron with the level of the target exon that is incorporated into mRNA expressed from a control expression cassette comprising the wild-type or natural target exon and / or intron.

[0112] Thus, the present invention relates to a novel method of generating a plurality of deletion mutants of a target exon and / or intron, combined with validating a predicted change in splicing efficiency caused by an AON that is adapted to bind to a target sequence within the target exon and / or intron. In embodiments, the predicted change in splicing efficiency can be validated by transfecting AONs and / or corresponding deletion mutants into suitable cells in vitro or in vivo. Comprehensive Experimental Dataset Generated From Deep Indel Mutagenesis (DIM)

[0113] The present disclosure relates to the use of a comprehensive experimental dataset based on deep indel mutagenesis, which is adapted for evaluating splicing efficiency. The process of deep indel mutagenesis involves: (1) construction of a library of DNA variants (input I indel library), (2) providing a cell culture to express the library of RNA variants (output library), and (3) quantification of the variant abundances of input and output by, for instance, RNA sequencing or RT-qPCR (Figure 1A). DIM thus enables, forthe first time, the assessment and / or quantification of the impact of a comprehensive and methodical selection all deletion, substitution and / or insertion mutations in a target exon and / or intron on the splicing efficiency of a target exon (Example 1), such as an exon of a human gene.

[0114] The comprehensive dataset comprises of an input DNA variant library, output library of mRNA transcripts from a transfected culture, and the effect of each variant on splicing efficiency, calculated based on the proportion I quantity of wild-type target exons in input and output libraries. In some embodiments, the DNA variant input library comprises or consists of a plurality of mutation variants of a target exon. In some embodiments, the DNA variant input library consists of all possible variants of a particular type (e.g. single nucleotide deletions I insertions) from within a target exon.

[0115] Variants

[0116] The methods for assessment of exon inclusion or exclusion typically involve constructing a database of possible variants of a target exon and / or adjacent intron sequence. Beneficially, the database of variants is as comprehensive as necessary to provide valid datasets. In some aspects and embodiments the comprehensive dataset of variants may comprise all possible deletion mutants of a defined length (e.g. 1 or more nucleotide deletion, as described elsewhere herein). In some aspects and embodiments the comprehensive dataset of variants may comprise all possible insertion mutants of a defined length (e.g. 1 or more nucleotide insertion, as described elsewhere herein), wherein each possible nucleotide is inserted, in turn, in each position of the target sequence. An insertion may include 1 nucleotide only, or alternatively all possible insertions of 2, 3, or more nucleotides of each possible sequence may be inserted. In some aspects and embodiments the comprehensive dataset of variants may comprise all possible substitution mutants of a defined sequence length (e.g. 1 or more nucleotide may be substituted at a time, as described elsewhere herein). A substitution may include 1 nucleotide only, which may be substituted for each other of the possible nucleotides, or alternatively a substituted sequence may include 2, 3, or more adjacent nucleotides of each possible sequence, which may be substituted for each of the possible alternative nucleotides in turn.

[0117] By way of example, all possible single nucleotide substitutions may comprise: at each position of the target exon and / or adjacent intron, changing the nucleotide for each of the other 3 alternatives (e.g. if there is an A, then it will be changed for each of T, C, or G). Similarly, all possible deletions may comprise: for a deletion length of 1 , deleting the residue at position 1 , or position 2, or position 3, or position 4 etc along the entire length of the target exon and / or intron; for a deletion length of 2, deleting the residues as positions 1 and 2, or positions 2 and 3, or positions 3 and 4, etc along the entire length of the target exon and / or intron; and by way of example, a deletion length of 60 would mean the deletion of positions 1 to 60, or positions 2 to 61 , or positions 3 to 62 etc. along the length of the target exon and / or adjacent intron.

[0118] Comparatively, all possible 1 nucleotide insertions means that every possible nucleotide would be inserted after position 1 , after position 2, position 3 etc., in turn, along the length of the target exon and / or adjacent intron; and all possible 2 nucleotide insertions means that every possible dinucleotide would be inserted after position 1 , after position 2, etc., in turn, along the length of the target exon and / or adjacent intron.

[0119] In some embodiments, the database may incorporate all possible deletions of a certain length, or up to a predefined length; and / or all possible insertions of a certain length, or up to a predefined length, and / or all possible substitutions of a certain length, or up to a predefined length.

[0120] Splicing Efficiency and Exon Inclusion Level

[0121] In the context of alternative splicing, a change in splicing efficiency can be measured by calculating the percentage of mature transcripts that include the target exon. Splicing efficiency can, for example, be measured according to ‘inclusion level of a target exon’. Thus, the change in splicing efficiency can be defined by comparing the level of wild-type target exons from the variant to that of the control. Using the comprehensive dataset obtained by the methods of the present disclosure, the impact of a large number of variants can be comprehensively quantified. In embodiments, DIM calculates the PSI value (the percentage of mature transcripts that include the exon) of all designed exon variants based on enrichment score. This process is described in more detail in Example 1 , and illustrated in Figures 1A and 1 B.

[0122] In embodiments, the raw enrichment score (ES) for a single mutant x may be calculated as the ratio between the frequency in the output and the input libraries to the frequency of wild type (WT) target exon in the output and input libraries: fx, output

[0123] „ „ . , fx, input ,

[0124] ESX= In ( - )

[0125] JWT, output fwT, input

[0126] The PSI of the variant x can then be calculated by multiplying the PSI of WT target exon with the ratio of exponents of ES between the variant x and WT, as follows. Advantages of Using Comprehensive Dataset of Variants from DIM

[0127] The disclosure provides multiple advantages that are derived from generating a comprehensive dataset of variants through the use of DIM. In one embodiment, the effects of all selected single nucleotide substitutions (n = 189), insertions (n = 187) and deletions (n = 63) on the inclusion of FAS exon 6 was assessed (see Figures 1 D, 1 E). As described, in this example, the inventors identified the type of mutation that best contributes to any observed change in splicing efficiency I exon inclusion level.

[0128] The results revealed that nearly the entire range of inclusion values can be obtained as a consequence of at least one of the three classes of variant. Regardless of the type of mutation, just under two thirds of the mutations affected splicing by more than 10 PSI units (see Figures 1 D, 1E): 61.9% of substitutions changed the inclusion of FAS exon 6 by more than 10 PSI units (mean absolute APSI = 17.5 PSI units); comparatively, 58.7% of single nucleotide deletions (mean absolute APSI = 15.5 PSI units) and 61.3% of single nucleotide insertions (mean absolute APSI = 19.7 PSI units) had a relevant impact on exon inclusion. It was further observed that nucleotide base substitutions have a mode near to the wild type (WT) PSI (49.1%) and such substitutions appear to more often promote exon skipping than inclusion, with 47.6% of substitutions decreasing and 14.3% of substitutions increasing inclusion of the target exon by more than 10 PSI units. Similarly, deletions tend to promote skipping, with 15.9% of single nucleotide deletions promoting exon inclusion and 42.9% of single nucleotide deletions promoting skipping by more than 10 PSI units. In contrast, single nucleotide insertions more frequently promote target exon inclusion, with 41 .1% increasing the inclusion of FAS exon 6 and only 20.2% of nucleotide insertions decreasing inclusion by more than 10 PSI units (see Figure 1E).

[0129] The use of a large dataset based on DIM also allows to compare and correlate the effects of different mutation types in the same position along the exon. For example, the effects of single nucleotide substitutions and deletions in the same positions correlate moderately well (Spearman’s rho = 0.46, Figure 1F), consistent with some of these variants disrupting existing regulatory elements. Notably, based on the data obtained, it appears that positive regulatory elements cover a larger proportion of the target exon than negative regulatory elements (see Figures 1 D, 1E). However, the effects of substitutions and deletions correlate poorly with the effects of insertions immediately before or after the substituted I deleted position (for example, see Figures 1G, 1 H, Figure 9, 10). This data appears to suggest that nucleotide insertions tend to affect splicing of a target exon through a different mechanism to that of corresponding substitutions and deletions. Without wishing to be bound by theory, it is possible that the different mechanism may be based on the creation of new regulatory sequences I elements through the combination of original and inserted sequences (see below).

[0130] The use of a large, comprehensive dataset further allows to analyse the influence of short multinucleotide deletions (e.g. ranging from 2 to 9 nts) and insertions (e.g. either 2, 3 or 4 nts) on exon inclusion. In this regard, short multi-nucleotide deletions were found to have a similar effect on exon inclusion as single nucleotide deletions. Thus, 64.5% of 2-nt deletions affected splicing by more than 10 PSI units (mean absolute APSI = 17.0), as well as 62.3% of 3-nt deletions (mean absolute APSI = 17.9) and 73.3% of 4-nt deletions (mean absolute APSI = 20.8). Considering all short deletions up to 10 nts long, 66.0% altered target exon inclusion by more than 10 PSI units (mean absolute APSI = 18.7). Short multi-nucleotide deletions longer than 2nt more frequently promote target exon skipping, with 30.6% of 2-nt deletions, 39.3% of 3-nt deletions, 43.3% of 4-nt deletions, and 45.8% of all deletions spanning 10 or fewer nts decreasing exon inclusion by more than 10 PSI units. In contrast, 33.9% of 2-nt deletions, 23.0% of 3-nt deletions, 30% of 4-nt deletions and 20.2% of all deletions up to 10 nts long increase inclusion by more than 10 PSI units.

[0131] The effect of mutants generated from the input and output libraries on target exon inclusion (measured in PSI), is shown to be comparable to that of individual transfections (see Figure 1C). The resulting enrichment scores and percent-spliced-in (PSI) value of each variant were highly reproducible across nine experimental replicates (Pearson’s r between 0.97 and 0.98 for all pairs of replicates; see Figure 8)

[0132] Notably, longer multi-nucleotide deletions were found to be particularly effective at identifying splicing regulatory elements (Figure 2A). Thus, while 1 to 4 nt deletions at the 5’ end of an exon showed multiple, in some cases contradictory effects, whereas longer deletions delineated approx. 8 5’ terminal nucleotides whose collective deletion strongly promoted exon skipping. Similarly, deletions that cover exon positions 9 to 18 consistently increased exon inclusion. Interestingly, deletions that partially overlap this region (on either side) appear to have the inverse effect by promoting exon skipping. Systematic deletion scans thus delineate consistent discrete regulatory elements (enhancers and silencers), which are often adjacent to each other, and reveal patterns such as alternating elements with antagonistic effects (see Figure 2A). Short multi-nucleotide insertions were found to have a stronger effect on FAS exon 6 inclusion compared to single nucleotide insertions. For example, 74.5% of 2-nt insertions and 77.5% of 3-nt insertions changed inclusion by more than 10 PSI units (mean absolute APSI = 23.9 and APSI = 24.2, respectively). This library also contained a random selection of 585 4-nt insertions, of which 88.6% were found to change splicing by more than 10 PSI units (mean absolute APSI = 31.2). Similar to single nucleotide insertions, but in contrast with short multi-nucleotide deletions, short multi-nucleotide insertions tended to promote inclusion rather than skipping: 44.6% of 2-nt insertions increased exon inclusion by more than 10 PSI units (compared to 30.0% which decreased target exon inclusion by this same amount), as well as 42.4% of 3-nt insertions (35.1 % for exon skipping) and 57.5% of 4-nt insertions (c.f. 31 .2% for exon skipping).

[0133] In contrast to the regulatory landscape emerging from deletion analyses (Figure 2A), double or triple nucleotide insertions tended to show autonomous effects that were strongly influenced by the nature of the inserted nucleotides (see Figures 2B, 2C, 11). By way of example, insertion of GC, CG or GA dinucleotides generally resulted in increases in exon inclusion, almost independently of the position of the insertion, with the prominent exception of insertions at the 3’ end of the exon. In contrast, insertion of GG or CC dinucleotides showed markedly different effects depending on the site of insertion.

[0134] Further, use of the comprehensive dataset obtained from DIM allows to analyse the influence of variants in selected regions of the genome. For example, all CG-containing triplets enhanced exon inclusion in nearly all positions (similarly to the insertion of GC, CG or GA dinucleotides), with the notable exception of the five 3’ nucleotides of the target exon. Almost any insertion in this region (except those positioning an A at the 3’ end of the target exon) led to enhanced exon skipping (black vertical rectangles on the right of panels of Figures 2B, 2C), suggesting that it harbours a strong enhancer sequence that is very sensitive to mutations such as insertions or deletions, and the wild-type sequence seems to be important for activation of the adjacent 5’ splice site. Effects similar to those of CG-containing triplets were observed upon insertion of a variety of GC I GA I GG containing triplets (see lower rows of Figure 2C); some of these effects may be related to enrichment in purine residues which, together with other purines present in the insertion site, could function as purine-rich exonic enhancers, a known class of exonic regulatory elements. Indeed, these effects are also observed for other purine-rich triplets such as GAG or AAG, albeit not for all purine-rich triplets (e.g. AGG or GGG). Triplets containing AU / UA dinucleotides (e.g. UAG, UAA, UUA, AUU, CUA) were found to generally promote skipping when inserted at most exonic positions (cluster of blue I grey colour in mid-low rows of panel, Figure 2C, rows between and to right of dark grey triangles), which might be explained by enhanced binding of hnRNP proteins such as hnRNP A1 , which are known to mediate effects of exonic silencers. Pyrimidine-rich triplets (e.g. UCC, UUC, CUU), which could provide I reinforce binding sites for other repressive hnRNPs such as PTB / hnRNPI, have however very regional effects, promoting skipping mainly in already pyrimidine-rich regions like the previously described PTB silencer located in positions 28 to 39 (middle upper black rectangle in panel, Figure 2C). A variety of cytosine-containing triplets were found to systematically promote target exon inclusion when introduced between exon positions 17 and 24 (upper left black rectangle in panel, Figure 2C), which is a G-rich sequence, while insertion of 3 additional G nucleotides in this region was found to strongly inhibit inclusion, suggesting that a G-rich silencer in this region, possibly forming G quadruplexes recognized by hnRNP F / H factors, is disrupted by C-containing triplets. The latter effects could be in part linked to the creation of CG dinucleotides, which, as discussed above, appear to display strong enhancing effects.

[0135] The use of the comprehensive DIM dataset further allows systematic evaluation of exon mutations relative to a tissue specific transcriptomic dataset that annotates genome wide alternative splicing (GTEx data). As disclosed herein, it has been found that there is a strong positive correlation (Spearman’s rho between 0.72 and 0.87 for 2-mers, between 0.69 and 0.79 for 3-mers) between the APSI induced by kmer insertions in the deep indel mutagenesis library and the PSI of exons containing at least 20 (for 2-mers) or 10 (for 3-mers) such kmers in the GTEx database, either in adipose tissue (Figure 2D) or in all GTEx tissues (Figure 12). Furthermore, the relationship between the content of each individual nucleotide in the inserted 3-mers and the PSI of the mutated exon (Figure 2E) resembles the relationship between the content of each individual nucleotide in cassette exons and their PSI (Figure 2F). For example, increasing the number of uridines in an inserted triplet appears to correlate with increased target exon skipping (Figure 2E), similarly to how a larger uridine content is associated with increased skipping that is observed for exons throughout the genome (results for GTEx adipose tissue are shown in Figure 2F, results for all GTEx tissues are shown in Figures 13 to 16). These results suggest that the effects of systematic analysis of insertion mutations observed for one individual model exon have beneficially captured likely sequence features and patterns that may be relevant for the inclusion of alternatively-spliced exons genome-wide.

[0136] Example 9 (see below) further illustrates how systematic analysis using the data derived from DIM suggests a mechanism for the evolutionary origins of unusually short microexons and the repression of transmembrane domain-encoding exons, and reveals a checkerboard architecture of sequential enhancers and silencers in a model of alternatively spliced exons.

[0137] Identification of Antisense Oligonucleotide (AON) Target Binding Sites Based on Deep Indel Mutagenesis (DIM)

[0138] In addition to predicting the effects of individual exon and / or intron mutations using DIM, the present disclosure provides a novel method for identifying target RNA regions that may be bound by AONs in order to influence exon skipping. As demonstrated, for example, in Example 8 (described below), the present disclosure demonstrates a strong correlation between a change in splicing efficiency as a result of a particular deletion mutant and a change in splicing efficiency caused by an AON binding to the sequence of residues in a wild-type RNA that match the residues deleted from the deletion mutant. Thus, mutant regions which are identified as functionally important in a deletion analysis are generally found to provide advantageous target sequences for AON treatment for induction of the same effect.

[0139] In a clinical setting, AONs are generally longer than individual regulatory elements (e.g. typically 18 to 21 nts in length for an AON versus 5 to 10 nt regulatory motifs), making the prediction of AON effects challenging - for example, because an AON my affect more than one regulatory sequence, which may include exon inclusion enhancer and / or suppressor sequences. Deletion mutagenesis thus has been found herein to provide a rapid and effective method for predicting the effects of AONs that may bind to each different regions of an RNA transcript, since the effect of nucleotide deletions - which may remove one or more regulatory motif(s) - is expected to inhibit the assembly of cognate frans-acting regulatory factors at the site of deletion, which is likely also to be the mechanism of action by which an AON may cause its effect, i.e. by competing with the binding of frans-acting factors to the same sequences (see Figure 7A).

[0140] The effects on splicing of an array of partially overlapping AONs collectively covering the entire length of FAS exon 6 (known as an AON walk) with the effects of deletions of the same length along the same sequence region (i.e. deletion walk) were compared. Changes in exon inclusion (measured in PSI) correlated well for 21 -nt AONs and 21 -nt deletions spaced every 5 nucleotides along the exon (Spearman rho = 0.75, n=9, see Figure 7B). It was found that AONs modulate exon 6 inclusion over a wide dynamic range (from 20% to 50% inclusion, compared with the approximately 50% inclusion level of the WT exon). Thus, the strong correlation between exon inclusion (PSIs) of AONs and deletion mutants provides an efficient strategy to design AONs that would modulate the splicing in the gene of interest.

[0141] Computational Models in the Use of Deep Indel Mutagenesis (DIM)

[0142] There are a number of computational models known in the art that predict effects of different types of variants on alternative exon splicing (see Example 2).

[0143] For example, SMS score is an additive model which uses exonic 7-mers as an input feature (Ke et al., (2018) Genome Res. 28, 11-24). Some such models employ a deep neural network. HAL is another additive model using hexamers as input features with parameters learned from millions of random 50-nt-long exonic and intronic sequences (Rosenberg et al., (2015) Cell 163: 698-711). Some of the computational models employ deep learning. MMsplice is a modular neural network where modules were trained to predict the effects of mutations on splicing relevant regions, including the donor site, acceptor site, exon, 5’end and 3’end. SpliceAl predicts whether each position in a transcript is a splice donor site, a splice acceptor site, or neither, using as input only the genomic sequence of the pre-mRNA transcript (Jaganathan et al., (2019) Cell 176, 535- 548. e24). SpliceAl employs a network architecture consisting of dilated convolutional layers, and uses only primary sequence as an input to evaluate 10,000 nucleotides of the flanking context sequence to predict the splice function of each position in the pre-mRNA transcript. Pangolin is a deep learning model based on the SpliceAl architecture but also trained with data from three additional mammalian species (Zeng et al., (2022) Genome Biol. 23, 103).

[0144] In some aspects and embodiments of the invention, the effects of individual variants from at least one computational model are systematically evaluated, relative to a comprehensive experimental dataset of variants. In embodiments, the effects of individual variants on splicing are measured in “percent spliced in (PSI)”, or relative rate of target exon inclusion between a variant and the corresponding wild-type target exon. Beneficially, embodiments allow a comparison of results from computational models with the comprehensive experimental dataset in a systematic, high throughput manner, to enable the accurate prediction of the effects of variants (deletion, substitution or insertion mutants) or AON binding on splicing efficiency. In embodiments, the comprehensive experimental dataset consists of libraries and data on the effect of variants calculated from the libraries, which is suitably obtained from deep indel mutagenesis (DIM).

[0145] Aspect and embodiments of this disclosure provide the opportunity to test the computational models against variant exons and / or introns to predict the effect of mutations, as well as to perform independent evaluation of the accuracy of the model predictions of the effects of variants. In embodiments, a suitable systematic, high throughput comparison involves establishing correlations between the target exon inclusion levels of variants from at least one computational model and comprehensive experimental datasets. In embodiments, the Spearman correlation method can be used. Figure 5A illustrates correlations between the inclusion levels of variants in the library and their inclusion levels relative to each of the five computational models. In the example, the inclusion levels of the variants are calculated in “percent spliced in (PSI)”. The inventors evaluated the performance of five different models described: SMS score, HAL, MMSplice, SpliceAl, and Pangolin to predict the effects of single nucleotide substitutions. All models predicted the effects of single nucleotide substitutions at least moderately well (Figure 5A), with Pangolin showing the best performance (rho = 0.82), followed by SpliceAl (rho = 0.79), MMsplice (rho = 0.74), HAL (rho = 0.69) and SMS scores (rho = 0.50). These models had a similar range and order of performance when predicting the effects of all insertions (1-, 2-, 3- and 4-nts long) in the library of mutants I variants tested (rho = 0.80 for Pangolin, rho = 0.78 for SpliceAl, rho = 0.66 for MMsplice, rho = 0.67 for HAL and rho = 0.54 for SMS scores). Interestingly, the predictive performance for deletions was relatively low for some methods: rho = 0.18 for MMsplice, rho = - 0.41 for HAL, and rho = 0.17 for SMS, in comparison to the scores for SpliceAl (rho = 0.90) and Pangolin (rho = 0.89), which remained highly predictive.

[0146] Further, systematic evaluation of results from computational models relative to a comprehensive experimental dataset allowed the selection of a computational model that closely predicts the effects of the variants in interest. For example, SpliceAl as the best model for predicting the effects of single-nt deletions.

[0147] In addition, such aspects and embodiments allow computational reconstruction of a comprehensive experimental dataset. For example, SpliceAl predictions of inclusion levels of all 1 to 6 nucleotide deletion variants on FAS exon 6 closely replicated deletion maps from deep indel mutagenesis (see Figures 5B and 5C).

[0148] In Silico Deletion Mutagenesis Predicts Genome-Wide Variant Effects

[0149] In some aspects and embodiments, at least one computational model is applied to evaluate effects of a plurality of variants across multiple target exons in a genome-wide set. In embodiments, the genome-wide set can consist of an entire genome of an individual. In other embodiments, a genome-wide set may comprise genomes from different individuals, optionally including genetic variant information in respect of the different individuals.

[0150] In some aspects and embodiments, the computational model may perform systematic in silico deletion mutagenesis which can generate deletion mutants of thousands of potential target exons to create a genome-wide variant map. In embodiments, the genome-wide variant map is a map of deletion variants. In other embodiments, the genome-wide variant map is a map of substitution variants. In other embodiments, the genome-wide variant map is a map of insertion variants. In some embodiments, the genome-wide variant map is a map of different / multiple types of variant; for example, including two or more of deletion mutants, substitution mutants and / or insertion mutants. In an example, In silico deletion mutagenesis correctly predicts the effect of all 4-nt deletions in 18,551 exons expressed in at least 80% of GTEx tissues with a length between 50 and 200 nts. The impact of a 4 nucleotide deletion is predicted to depend on each individual exon, although highly-included exons (PSI > 90%) are predicted to be more robust to PSI changes compared to exons included at lower levels (see Figure 6A), which is an expected consequence of the scaling law that the effects of splicing mutations are predicted to follow (Baeza et al., 2019). Example 10 further illustrates how In silico deletion mutagenesis reveals the regulatory architecture of exons genome-wide.

[0151] DANGO: A Genome Wide Resource for Antisense Oligonucleotide (AON) Design

[0152] The present disclosure provides accurate prediction of the effects of deletions and other variants on exon inclusion using deep learning, that maps the splicing regulatory landscape of target exons and predicts effective splice-altering antisense oligonucleotides for altering splicing efficiency of those target exons.

[0153] In aspects and embodiments of the disclosure, a genome wide score (DANGO score) that indicates the effect of variants of each different target exon across the entire genome is encompassed. In embodiments, the effects of variants are annotated by change in splicing efficiency I exon inclusion level. In some embodiments, target gene sequences are limited to exomes. In other embodiments, target gene sequences incorporate a target exon and at least one adjacent intron. Such aspects and embodiments provide may provide an efficient, affordable strategy to identify regions in an exome that can be successfully targeted by AONs to achieve a range of desired splicing outcomes, e.g. for therapeutic or biotechnological applications.

[0154] In various aspects and embodiments of the disclosure, a genome wide score (DANGO score) may be generated by: defining an input set comprising of the target exons and a subset of the plurality of variants; selecting a computational model for prediction of the effects of the variants; applying the computational model to extract features indicative of disruptions on splicing acceptor and / or donor sites; and applying at least one extracted feature to score a change in splicing efficiency for at least one of the plurality of variants. In embodiments, the plurality of variants can be deletion mutants, substitution mutants, insertion mutants, or a combination of more than one type of mutant. In some embodiments, the input set can be derived from in silico mutagenesis. In some embodiments, the input set can be derived from the dataset of deep indel mutagenesis. In some embodiments, the computational model can be a deep learning model. In some embodiments, the features indicative of disruptions on splicing acceptor / donor sites can be measurements of splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain.

[0155] In aspects and embodiments of the disclosure, DANGO score is the maximum score taken from the individual scores for one or more of splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain, wherein the scores for splice site acceptor loss and splice site donor loss are first multiplied by -1 to generate a negative score; preferably wherein the DANGO score for each deletion mutant is the maximum score taken from the four individual scores for splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain, wherein the scores for splice site acceptor loss and splice site donor loss are first multiplied by -1 to generate a negative score.

[0156] For example, in one particular embodiment, to generate a genome-wide DANGO score, SpliceAl is used (although in other embodiments any of the alternative models, such as Pangolin, may be used or combined, as disclosed herein) to generate a comprehensive in silico mutagenesis data containing all possible 21 nt deletions across the target exome. Of course, it is not essential that the variants are based on 21 nt deletions, and other lengths of deletions, and / or suitable lengths of insertions or substitutions may alternatively be used. To compile the list and sequences of each exon, R package biomaRt (Dunrick et al., (2009) Nat. Protoc. 1141 4:1184-1191) was used with the hsapiens_gene_ensembl dataset. In this example, the analysis was limited to canonical exons only (referring to the transcript_is_canonical attribute of each exon). After compiling the list of exons and each of their 21 nt deletions, each sequence was assessed by SpliceAl using the default parameters. Four output scores from SpliceAl were measured: splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain. The maximum score was taken from these four output scores, and, if the score corresponded to splice site acceptor or donor loss, the number was multiplied by -1 to generate a negative score. The resulting score, i.e. the DANGO score for each 21 nt deletion mutant indicates the change in splicing efficiency I exon inclusion level predicted for that variant.

[0157] Importantly, based on experimental assays, a strong correlation was observed between the measured effect on splicing efficiency by physical AONs, and DANGO scores relating to the corresponding deletion mutants, i.e. deletion mutants of the 21 nucleotides targeted by the physical AON. The effects on splicing by an array of partially overlapping AONs collectively covering the entire length of FAS exon 6 (AON walk) was compared to corresponding DANGO scores predicted from in silico 21 nt deletions generated from SpliceAl. The correlation between the physical AON effects and the SpliceAI-predicted effects of the 21 nt deletions was found to be strong (rho = 0.74, Figure 7C). These results suggest that DANGO score can provide a reliable measure for the prediction of AON effects on splicing efficiency / inclusion level of target exons.

[0158] DANGO: A Genome Wide Resource for Antisense Oligonucleotide (AON) Selection

[0159] In aspects and embodiments of the disclosure, different thresholds can be applied to DANGO scores to select exon I intron variants that may be used to design particularly useful or advantageous AONs. In embodiments, the threshold level may be at least about 0.025, at least about 0.05, at least about 0.2, or at least about 0.4. In embodiments, the threshold level is less than about 0.5, less than about 0.5, less than about 0.1 , or less than about 0.04. For instance, in one example, the inventors found that 12.4% of all 21 nt deletions had an absolute DANGO score greater than 0.1 (mean absolute DANGO score across the exome = 0.05). Short exons (<100 nucleotides) were most vulnerable to these deletions (Figure 7D), with 44.8% of all 21 nt deletions in these short exons (or portions thereof) having an absolute DANGO score greater than 0.1 , compared to 10.2% for longer exons (above 100 nucleotides). Varying the DANGO score threshold (0.025, 0.05, 0.2, or 0.4) altered the proportion of deletions considered impactful, but it did not change the fundamental finding that short exons were more sensitive to the effects of the deletions. Regardless of the length of the exon, the proportion of negative DANGO scores is greater than the proportion of positive scores (Figure 7E). This suggests that AONs targeting exonic regions are more likely to reduce incorporation of the target exon, rather than to increase inclusion of the target exon.

[0160] The criteria for determining a suitable threshold for selection may depend on various factors and variables, such as the identity and context of the exon. For example, if an exon is naturally very highly included (e.g. PSI about 90% of higher), it is very difficult to get large positive DANGO scores, because it is only possible to slightly increase the proportion of inclusion of the exon; whereas it may be comparatively easy to observe larger negative scores, because there greater opportunities for the PSI to go down.

[0161] DANGO: A Genome Wide Resource for Target Exon Identification

[0162] In aspects and embodiments the disclosure also provides a method for identifying at least one target exon using DANGO scores of variants, such as deletion mutants, across a plurality of (e.g. all) possible exons. In embodiments, at least one target exon may be identified based on the average DANGO scores of the assessed variants (e.g. deletion mutants) across each possible target exon. In embodiments, at least one target exon may be identified from the distribution of DANGO scores of the assessed variants (e.g. deletion mutants) for each possible target exon. In embodiments, at least one target exon is identified from a dataset comprising of a plurality of variants (e.g. deletion mutants) from a plurality of target exons across the genome.

[0163] For example, DANGO scores for all 21 nt deletion mutants in the genome in an exon-by-exon basis can be explored. The exonic regions of interest that may be susceptible to splicing changes resulting from each of the 21 -nt deletions can be specifically explored. As DANGO score indicates an approximate change in splicing efficiency / exon inclusion level, exons with DANGO scores with large absolute values can be preferentially selected, if desired.

[0164] Experimental Validation of Antisense Oligonucleotides (AONs) Selected from the Scoring System

[0165] The present disclosure demonstrates that DIM can be used as an experimental method to validate the effect of an AON on splicing efficiency. In aspects and embodiments of the disclosure, therefore, there is provided a method for validating a predicted change in splicing efficiency caused by an AON, the method comprising: (a) providing a cell or cells comprising (i) one of a library of antisense oligonucleotides adapted to bind to a target exon sequence and / or adjacent intron sequence within an expression cassette capable of being expressed within the cell or cells, and / or (ii) any one of more of a library of expression cassettes, wherein each cell or cells comprises a variant of the target exon and / or adjacent intron (e.g. a deletion mutant), wherein the variant has a mutated sequence of the target exon and / or adjacent intron (e.g. wherein each deletion mutant lacks a predefined sequence of the target exon and / or intron), and wherein the expression cassette is capable of expressing RNA encoded by the variant (e.g. deletion mutant) of the target exon and / or adjacent intron within the cell or cells, and / or (iii) a wild-type or natural target exon and / or adjacent intron in the absence of an antisense oligonucleotide adapted to bind to the target exon and / or adjacent intron; (b) incubating the cell or cells of (a)(i) and / or (a)(ii) and / or (a)(iii) for a predetermined period of time to obtain a cell culture, wherein cells of the cell culture express RNA transcribed from the expression cassette comprising the target exon and / or adjacent intron; (c) isolating RNA transcribed from the expression cassette comprising the target exon and / or adjacent intron from the cell culture; (d) for any or each of (a)(i), (a)(ii) and / or (a)(iii) assessing the level of the target exon that is incorporated into mRNA expressed from the expression cassette; and (e)(i) comparing the level of the target exon that is incorporated into mRNA expressed from the expression cassette of the cell or cells of (a)(i) with the level of the target exon that is incorporated into mRNA expressed from the expression cassette comprising the wild-type or natural target exon and / or intron of the cell or cells of (a)(iii); and / or (e)(ii) comparing the level of the target exon that is incorporated into mRNA expressed from a cell or cells of (a)(ii) with the level of the target exon that is incorporated into mRNA expressed from the expression cassette comprising the wild-type or natural target exon and / or intron of the cell or cells of (a)(iii); and / or (e)(iii) comparing the level of the target exon that is incorporated into mRNA expressed from the expression cassette of the cell or cells of (a)(i) with the level of the target exon that is incorporated into mRNA expressed from a cell or cells of (a)(ii). In this way, the prediction of exon inclusion or exclusion can be validated against the experimental assay to directly influence exon inclusion or exclusion; or the predicted or measured amount of exon inclusion or exclusion can be readily compared to exon inclusion or exclusion in a wild-type or non-variant gene sequence.

[0166] Further details of the experimental validation and methods for expression of a deletion mutant or AON in a target cell or cells and exon inclusion I exclusion measurements related thereto, are outlined in Example 6 and Example 7, respectively.

[0167] EXAMPLES

[0168] Example 1 - Generating a Plurality of Deletion Mutants of the Target Exon and Generating a Dataset comprising of the Target Exon and each of the Deletion Mutants.

[0169] To quantify in parallel how single nucleotide (nt) substitutions, deletions and insertions affect alternative splicing of a human exon, a library containing all single-nt substitutions in the 63 nt-long FAS exon 6 (n = 189), as well as all deletions ranging in length from 1 to 60 nts (n = 2010), all possible 1-, 2-, and 3-nt insertions (n = 5208), and a random selection of 585 4-nt-long insertions were designed (Figure 1A). The library was cloned into a plasmid minigene vector spanning FAS exons 5-7, transfected into HEK293 cells; RNA was isolated 48h post-transfection and the exon 6 inclusion of each variant was quantified by counting how often it was present in the final exon inclusion product compared to every other variant in the library using deep sequencing of reverse transcription (RT)-PCR products (Figure 1B). The resulting enrichment scores allow the percent- spliced-in (PSI) value of each variant to be measured. The results were highly reproducible across nine experimental replicates (Pearson’s r between 0.97 and 0.98 for all pairs of replicates, Figure 8) and the estimated PSI values were very well correlated with PSI values determined by RT-PCR for 40 individual, independently transfected, mutant minigenes, which included indel and substitution mutants (Spearman’s rho = 0.91 , Figure 1C).

[0170] Input(indel) Library Construction

[0171] A sequence library was designed to include: the 63-nt-long wild-type sequence of FAS exon 6, all possible 189 single-nt substitutions, all 2010 possible deletions ranging in length from 1 to 60 nts, all 5208 possible 1-, 2-, and 3-nt-long insertions as well as 400 randomly-selected 4-nt insertions. The library was synthesised and purified by Twist Bioscience.

[0172] Input (Indel) Library Amplification

[0173] Oligo libraries were resuspended in 10 mM Tris buffer, pH 8.0 to a concentration of 20 ng / ul. 20 ng of template ssDNA Twist library was PCR amplified with Pfx Accuprime polymerase (Thermo scientific, 12344024) in a total reaction volume of 50 pl in quadruplicate, for 12 cycles (as recommended by Twist Bioscience for a 100-150 nt oligo pool) using the following flanking intronic primers: FAS_i5_GC_F (5’-TGTCCAATGTTCCAACCTACAG-3’; SEQ ID NO: 1) and FAS_i6_GC_R (5’-CTACTTCCCAAGTTATTTCAATCTG-3’; SEQ ID NO: 2). PCR reactions were combined and cleaned-up with the Quiaquick PCR purification kit, eluted with 50 pl elution buffer and dsDNA measured with a NanoDrop spectrophotometer.

[0174] Indel Library Subcloning

[0175] The amplified library was recombined with pCMV FAS wt minigene exon 5-6-7. Avectoninsert ratio of 1 :8 was used, using 150 ng of vector backbone and 20 ng of dsDNA amplified libraries and incubated at 50°C for two hours for DNA assembly, using a Gibson master mix developed at the CRG Protein Technologies Unit, which contains a mix of T5 exonuclease (T5E4111 K 1000U from Epicentre Biotech-Ecogen), Phusion polymerase (F530s 100U from VITRO and a Taq DNA ligase (Protein Technologies Unit CRG, homemade). After transformation into Stellar competent cells (Clontech, 636766), combining five replicates in order to maximise the number of individual transformants amplified, cells were grown for 18 hours in LB medium containing ampicillin. Approximately 4.29 million clones were obtained.

[0176] After bacterial transformation, the final plasmid library was purified using the Quiagen plasmid maxi kit (50912163, Quiagen) and quantified with a NanoDrop spectrophotometer. Transfection of Hek293 Cell Line to Generate Output Libraries

[0177] 750,000 Hek293 cells were plated on 100x20 mm petri dishes and transfected with 80 ng cloned libraries in 8 ml OPTIMEM Reduced Serum Medium with no phenol Red (Life technologies, 11058021) using Lipofectamine 2000 (Life technologies, 11668019) in nine biological replicates. 48h post transfection, cells were collected and RNA was prepared using Maxwell simplyRNA Tissue Kit (Promega, AS1280). cDNA was prepared with 400 ng total RNA using specific vector backbone PT2 primer (5’-AAGCTTGCATCGAATCAGTAG-3’; SEQ ID NO: 4) and Superscript III reverse transcriptase (Thermo Fisher, 18080085). PCR amplification of cDNA samples was performed with GoTaq flexi (PROMEGA, M7806) using distinct 8-mer barcoded oligos to distinguish the nine experimental replicates. PCR products were run on a 2% agarose gel and the smear corresponding to sizes of the amplification product expected from exon inclusion (full length and insertion and deletion mutants) was excised, purified using the Quiaquick Gel extraction kit (Qiagen, 50928704) and quantified with a NanoDrop spectrophotometer.

[0178] Input Indel Library

[0179] 20 ng of the plasmid library was amplified in triplicate using GoTaq flexi DNA polymerase (M7806, Promega) for 25 cycles with three different pairs of barcoded intronic primers. FAS_i5_TR_F and PT2 (“Indel library amplification primers” in Table 1). Since the insertions and deletions result in a library with exons of different length, this resulted in a PCR smear (corresponding to exons ranging in length from 3 to 67 nts), which was gel-purified and sequenced. Each pair of primers had a distinct 8-mer barcode sequence to discriminate between technical replicates (“Primers used for amplifying technical replicates (input library)” in Table 1).

[0180] Table 1 : Primer sequences

[0181] Sequencing

[0182] Equimolar quantities of three independent amplifications of the input library and equimolar quantities of the purified inclusion smear (output library) of each of the nine replicates were pooled and sequenced at the CRG Genomics Core Facility where Illumina Ampliseq PCR-free libraries were prepared and run on a single lane of an Illumina HiSeq2500. In total, 424 million paired-end reads were obtained (188 and 236 million for input and output respectively). The median sequencing coverage for all exon variants in the input was 2114 reads. In the output, the sequencing coverage was between 278 and 468 reads. Raw sequencing data has been submitted to GEO with accession number GSE244179.

[0183] Data Processing and Calculation of PSI Values

[0184] FastQ files from paired-end sequencing were processed with DiMSum v1 .2.764 using default settings with minor adjustments (https: / / github.com / lehner-lab / DiMSum). First, DiMSum was run in default paired-end mode to demultiplex reads into input and replicate output samples (Stage 0 only). Second, DiMSum Stages 1 to 5 were run in single-end mode ('--paired' = F) using only demultiplexed forward reads that have full coverage of the exon sequence.

[0185] Reverse reads, originally intended to cover a unique molecular identifier (UMI) and a 3' portion of the exon sequence, were discarded. The final stage estimates an enrichment score (ES) and associated error for each mutant variant based on its frequency in the input and output libraries, and relative to the wild type sequence in both libraries. Experimental design files and command- line options required for running DiMSum on this dataset are available on GitHub (https: / / github.com / lehner-lab / fas-indel-library).

[0186] The PSI of the wild type FAS exon 6 sequence has been experimentally shown to be 49.1% (Julien P et al., 2016). Therefore, the PSI of a variant of interest is estimated as follows:

[0187] Example 2 - Generating a Plurality of Deletion Mutants of the Target Exon and Generating a Dataset comprising of the Target Exon and each of the Deletion Mutants.

[0188] In Silica 4mer Deletions in Exons Genome-Wide

[0189] The SpliceAl developers created a file with annotations for all possible substitutions, 1 base insertions, and 1 to 4 base deletions across the genome. This file is available for download at https: / / basespace.illumina.eom / s / otSPW8hnhaZR. In this Example, 4mer deletion data for all exons across the genome was extracted and PSI values were calculated (see Estimating PSI values in the GTEx dataset). For each 4mer deletion in an exons, the researchers computed the SpliceAl score by taking the maximum value among the acceptor gain, acceptor loss, donor gain, and donor loss scores. If the maximum value was the acceptor loss or the donor loss score, the value was multiplied by -1 .

[0190] Example 3 - Scoring a Plurality of the Deletion Mutants using Different Computational Models

[0191] Predicting FAS Exon 6 Mutation Effects Using SMS Scores

[0192] Supplementary Table 7 from (Ke et al., (2018) Genome Res. 28, 11-24), which lists the SMS scores for all possible 7-mers was downloaded. To calculate the total SMS score for each exon in the present indel library, a sliding window analysis was performed by adding the SMS scores of consecutive 7-mers along its sequence. The final SMS score for each variant was obtained by subtracting the total SMS score for the wild type exon from that of the variant.

[0193] Predicting FAS Exon 6 Mutation Effects Using HAL

[0194] To predict the effects of exon variants on inclusion with HAL, a file containing the sequences in the present library was uploaded to http: / / splicing.cs.washington.edu / SE using 49.1% as the wild type levels of inclusion. An output file was returned that contained the predicted PSI values for each sequence in the input file.

[0195] Predicting FAS Exon 6 Mutation Effects Using Mmsplice

[0196] The Indel library design file was converted to VCF format, and this file was used as input for MMSplice. The algorithm was ran online, on the Google Colab notebook provided for this purpose (available at https: / / colab.research.qooqle.com / drive / 1 Kw5rHMXaxXXsmE3WecxbXvGQJma80Eq6) . This returned a CSV file with multiple columns containing different metrics for each exon variant. delta_logit_psi was selected as the predictor for the mutation effects.

[0197] Predicting FAS Exon 6 Mutation Effects Using Splice Al

[0198] Indel library design file is converted to VCF format, and used as input for SpliceAl. SpliceAl was run using GRCh38 as both the genome reference and gene annotation files. All other parameters were set to the default configuration. As SpliceAl outputs four scores for each mutant sequence (corresponding to splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain), the score associated with the highest absolute value in each case was selected.

[0199] Predicting FAS Exon 6 Mutation Effects Using Pangolin

[0200] The Indel library design file was converted to VCF format, and used as input for the Google Colab Notebook made available by the authors of Pangolin at https: / / colab.research.qooqle.com / qithub / tkzenq / Panqolin / blob / main / PanqolinColab.ipynb.

[0201] Pangolin was used with the default options chosen for the Colab Notebook, including GRCh37 as the genome reference. Like SpliceAl, Pangolin outputs four scores for each mutant sequence (corresponding to splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain). The Pangolin score selected for each mutation corresponded to that with the highest absolute value out of these four.

[0202] Example 3- Selecting one or more Deletion Mutant having a score above a Threshold Score For example, to generate a genome-wide DANGO score, the inventors first used SpliceAl to generate a comprehensive in silico mutagenesis dataset containing all possible 21 nt deletions across the exome. To compile the list and sequences of each exon, R package biomaRt (Dunrick etal., (2009) Nat. Protoc. 1141 4:1184-1191) was used with the hsapiens_gene_ensembl dataset. In this example, the analysis was limited to canonical exons only (referring to the transcript_is_canonical attribute of each exon). After compiling the list of all exons and each of their length minus 21 nucleotide deletions, each sequence was passed to SpliceAl using the default parameters. Four output scores from SpliceAl were measured: splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain. The maximum score was taken from these four individual scores, and, if the score corresponded to splice site acceptor or donor loss, the number was multiplied by -1 to generate a negative score. The resulting DANGO score for each 21 nt deletion mutant indicates the change in splicing efficiency I exon inclusion level upon deletion.

[0203] DANGO Scoring: 12.4% of all 21 nt deletions had an absolute DANGO score greater than 0.1 (mean absolute DANGO score across the exome = 0.05). Short exons (<100 nt) were most vulnerable to these deletions (Figure 7D), with 44.8% of all 21 nt deletions in these exons having an absolute DANGO score >0.1 , compared to 10.2% in longer exons. Varying the DANGO score threshold (0.025, 0.05, 0.2, or 0.4) altered the proportion of deletions considered impactful, but it did not change the fundamental finding that short exons are more sensitive to the effects of 21 nt deletions. Regardless of the length of the exon, the proportion of negative DANGO scores is greater than the proportion of positive scores (Figure 7E), which suggests that AONs targeting exonic regions are more likely to reduce inclusion of the exon, rather than increase its inclusion.

[0204] Identifying at Least One Target Exon from Clinical Knowledge

[0205] Human FAS Exon 6 was selected for further investigation. Skipping of FAS exon 6 is known to cause autoimmune lymphoproliferative syndrome (ALPS).

[0206] Identifying at Least One Target Exon from DANGO Score

[0207] DANGO scores for all 21 nt deletion mutants in the genome in an exon-by-exon basis were obtained. The inventors were then able to explore exonic regions of interest for sequences that may be susceptible to splicing changes upon 21 -nt deletions. By using DANGO score to predict a change in splicing efficiency I exon inclusion level, exons with DANGO scores with large absolute values can be selected as those that are most susceptible to changes in exon inclusion.

[0208] Interestingly, with a focus on FAS exon 6, it was revealed that DANGO scores cluster around the identified regulatory domains of the exon (Figure 7F), suggesting that these scores can also accurately reflect the regulatory architecture of exonic sequences. Furthermore, since FAS exon 6 is relatively short, the 21 nucleotide deletions push FAS exon 6 below the length threshold for exon definition (Figure 3A), nearly all 21 mer deletions in this exon are predicted to promote exon skipping (Figure 7F) and do, in fact, promote skipping as demonstrated in the experimental assay (Figure 5B). However, notably, deletions spanning silencer regions promote lower levels of exon skipping than deletions spanning enhancer regions.

[0209] Example 5- Generating Antisense Oligonucleotides (AONs)

[0210] Antisense oligonucleotides (AONs) for testing in vitro or in vivo can be designed and synthesised as is known in the art. The length of the AON can be selected based on preferences and the intended target sequence. In particular, it can be preferably to avoid AONs that have a sequence that has high identity I complementarity to other sequences found elsewhere in the target human or animal genome in order to avoid or minimise undesirable off target effects.

[0211] Typically, AONs ranging in length from about 7 nts to about 30 nts may be synthesised and used. Suitably, an AON is around 21 nts in length, for example, between about 18 and about 24 nts. As described elsewhere, AONs may contain one or more modified base, e.g. to improve delivery, stability etc. of the AON in use. Example 6. Validating a Change in Splicing Efficiency by Deletion Mutants.

[0212] Experimental Validation of Estimated PSI Values

[0213] To confirm the accuracy of PSI estimates on single-nt substitutions from DIM, previously experimentally-determined values of 24 exon variants (Julien P et al., 2016) are plotted against the experimentally-determined values (Figure 1C). For further validation of estimates of indel PSI values, 24 individual clones from the indel library were Sanger sequenced, of which 18 were also found in the output, and covered a wide range of estimated PSI. These variant clones were good for validation and checking correlation.

[0214] Individual variants I mutants were transfected into Hek293 cells in triplicate to quantify the ratio between FAS exon 6 inclusion and skipping. For RT-PCR, minigene-specific primers were used (“Primers used for amplifying biological replicates (Output library)”, Table 1). To avoid amplification of endogenous FAS RNAs, these primers (PT1 and PT2) are complementary to a plasmid backbone sequence distinct from endogenous DNA. RT-PCR products were fractionated by electrophoresis using 6% polyacrylamide gels in 1 x TBE and Sybr safe staining (ThermoFisher Scientific, S33102). The bands corresponding to exon inclusion or skipping were quantified using Imaged v1.47 (NIH, USA). PSI measurements are shown in Table 2.

[0215] Under the particular experimental conditions in which these indel mutants were tested, the wildtype exon was included with a PSI of 72% (compared to 49.1% in the experiment done to validate the single-nt substitutions). To visualise these results in the same plot as the single-nt substitutions (Figure 1C), splicing scaling law was used to adjust all experimentally-determined PSI values to their expected values had the wild type had a PSI of 49.1%.

[0216] A variant FAS exon 6 sequence library was designed to include: the 63 nucleotide long wild-type sequence of FAS exon 6, all possible 189 single nucleotide substitutions, all 2010 possible deletions ranging in length from 1 to 60 nts, all 5208 possible 1-, 2-, and 3 nt long insertions as well as 400 randomly-selected 4 nt insertions. The library of variant exons was synthesised and purified by Twist Bioscience.

[0217] Example 7. Validating a Change in Splicing Efficiency by Antisense Oligonucleotides (AONs) 100,000 HEK293 cells in a 6-well plate were transfected with antisense oligonucleotides harbouring 2’-OMe phosphorothioate modifications at each nucleotide position (Integrated technologies) using 3 pl of Lipofectamine 2000 (11668027, ThermoFisher Scientific) in 1 ml OPTIMEM I Reduced Serum Medium with no phenol red (11058021 , ThermoFisher Scientific) to a final concentration of 2.5 nM (exact sequences shown in Table 3).

[0218] Table 3: AON sequences and corresponding target exon sequences

[0219] Six hours post-transfection, the cell culture medium was replaced with DMEM Glutamax (61965059, ThermoFisher Scientific) containing 10% FBS and Pen / Strep antibiotics. 24 hours post-transfection, total RNA was isolated using the automated Maxwell LEV 16 simply RNA tissue kit (AS1280, Promega). cDNA was synthesised with 400 ng total RNA using Superscript III (18080085, Life Technologies) with a mix of random primers and oligodT. Effects on endogenous FAS exon 6 inclusion (or change in splicing efficiency) were determined by PCR using GoTaq flexi DNA polymerase (M7806, Promega) and the following primers:

[0220] FAS_e5_for: 5’-TGTGAACATGGAATCATCAAGG-3’ (SEQ ID NO: 66)

[0221] FAS_e7_endo_R 5’-AAAGTTGGAGATTCATGAGAACC-3’ (SEQ ID NO: 67)

[0222] Example 8. Correlating a Change in Splicing Efficiency by each Antisense Oligonucleotide to the Change in Splicing Efficiency of the Corresponding Deletion Mutant

[0223] A strong correlation between the measured effect of splicing efficiency by AONs in vitro, versus DANGO scores of the corresponding deletion mutants was observed.

[0224] The effects on splicing by an array of partially overlapping AONs collectively covering the entire length of FAS exon 6 (AON walk) was compared to DANGO scores predicted from in silico 21 nt deletions generated from SpliceAI. The correlation between the AON effects and the SpliceAI- predicted effects of 21 nt deletions was strong (rho = 0.74, Figure 7C), indicating that DANGO score can be a reliable measure to predict AON effect on the splicing efficiency / inclusion level of a target exon in vitro or in vivo.

[0225] The effects on splicing of an array of partially overlapping antisense oligonucleotides collectively covering the entire length of FAS exon 6 were compared to the effects of deletion mutations of the same length as assessed by DANGO score. Again, predicted changes in exon inclusion correlated well for 21 nt antisense oligonucleotides and 21 nt deletions spaced every 5 nt along the exon (Spearman rho = 0.75, n=9, Figure 1). Antisense oligonucleotides were found to modulate FAS exon 6 inclusion over a wide dynamic range (from 20% to 50% inclusion, compared with the approximately 50% inclusion level of the WT exon). The correlation between the antisense oligonucleotide effects and the scoring mechanism-predicted effects of 21 nt deletions was similarly strong (rho = 0.74, Figure 2). Example 9. Scientific Findings from Systematic Analysis Using Comprehensive Data from Deep Indel Mutagenesis (DIM)

[0226] Small Insertions Create Novel Microexons by Activating Cryptic Splice Sites Within FAS Exon 6 The relationship between exon length and inclusion in the input I indel library was analysed. The inventors have shown at 1 nt resolution, how short an exon can be while still being recognised by the splicing machinery. In this regard, no clear dependence of exon inclusion on length of the exon was found for exons longer than 50 nts (analysis based on deletions of up to 13 nt). However, it was found that expected inclusion of exons gradually decreased with increasing deletion length (i.e. shorter exons are less included), with almost no exons shorter than 30 nts showing detectable levels of inclusion in the input I indel library (Figure 3A). Consistent with this - and previous large deletions in constitutive exons (Berget et al., (1995) J. Biol. Chem. 270: 2411-2414) - exons shorter than 30 nts are more likely to be skipped genome-wide compared to longer exons (results for GTEx adipose tissue are shown in Figure 3B, results for all GTEx tissues are shown in Figure 17).

[0227] Interestingly, exons shorter than 27 nts are detected in multicellular animal but are categorised as a special class - microexons - whose recognition requires a dedicated set of regulatory sequences and factors (such as SRRM3 / 4) that enable their inclusion in specific tissues (e.g. the brain or endocrine pancreas (Irimia et al., (2014) Cell 159, 1511-1523). However, the present inventors have observed that a group of very short exons (microexons, length-wise) from the present libraries were detectably included (Figure 3A). These microexons appeared to correspond to large deletions at the 5’ or 3’ ends of FAS exon 6 (Figure 18A). However, these deletion mutants showed no evidence of exon inclusion when tested individually (Figure 18B). Without wishing to be bound by theory, it is possible that these apparently contradictory results might be explained if the clones detected by deep sequencing of the exon inclusion amplicon product (Figure 18A) were not the result of splicing of very short exons flanked by FAS exon 6 splice sites, but rather result from the use of cryptic splice sites within exon 6 that have been activated by another mutation that is no longer present in the spliced-in exon sequence. The central part of FAS exon 6 (positions 24 to 40) contains a pyrimidine-rich tract that resembles the polypyrimidine (Py)-tracts that precede 3’ splice sites, and the nucleotides in positions 10 to 15 contain a sequence that, strikingly, matches a branch point sequence (Figure 3C).

[0228] One possibility is that if an AG-containing kmer were introduced after the pyrimidine (Py)-rich segment, this would result in a 3’ splice site-like sequence arrangement that, if recognised by the spliceosome, could create a novel microexon spanning from this new 3’ splice site to the 3’ end of exon 6 (Figure 3D). To test this possibility, the inventors inserted AG-creating triplets (like CAG or CUA, since the next nucleotide is a G) after the Py-tract in the FAS exon 6 minigene and observed that exons corresponding to the final part of the exon were included to some extent in the mature RNA (Figure 3E). In contrast, insertion of GAC, that does not create a 3’ splice site, does not result in activation of a shorter exon, but rather enhances the inclusion of its full-length version (Figure 3E).

[0229] Interestingly, the exonic Py-tract is longer and more uridine-rich than the Py-tract associated with the natural 3’ splice site of intron 5 (Figure 3C). To assess whether the interplay between these 3’ splice sites plays a role in regulation, the inventors strengthened the Py-tract of the natural 3’ splice site (Figure 3F). In the presence of this mutation, inclusion of the full-length exon was enhanced in the wild type minigene and activation of the cryptic 3’ splice site in the AG-containing construct was greatly reduced (Figure 3G).

[0230] Previous work showed that the pyrimidine-rich sequence within FAS exon 6 functions as a silencer when bound by PTB (Izquierdo and Valcarcel (2007) J. Biol. Chem. 1078282: 1539-1543; Figure 3H). As expected, overexpression of PTB led to skipping of the wild type exon (Figure 3I, left panel) and also to reduced inclusion of the shorter exon in the AG-containing construct (Figure 3I, right panel), most likely due to direct competition between PTB and the Py-tract-binding splicing factor U2AF.

[0231] It was also investigated whether over-expression of SRRM4, which triggers inclusion of microexons in neurons, has any effect on inclusion of the shorter version of FAS exon 6 (Figure 3J). Surprisingly, SRRM4 overexpression was found to enhance inclusion of the shorter exon (Figure 3K) despite this exon not being flanked by c / s-acting sequences typically required for the inclusion of microexons (e.g. intronic UGC motifs). Interestingly, SRRM4 reduced, rather than enhanced, inclusion of the wild type full length exon 6 (Figure 3K). These results show that the short exon activated by a cryptic 3’ splice site in exon 6 is not only recognized by the splicing machinery but can be subject to splicing regulation by mechanisms similar to those operating on natural microexons.

[0232] The results thus far do not account for the detection of short exons spanning the first third of FAS exon 6 (Figure 3C). However, inserting sequences that mimic a 5’ splice site (e.g. UAA) after exon positions 18 or 19 induced the accumulation of spliced products containing these sequences (Figure 18C). Interestingly, the Py-tract in the central region of FAS exon 6 is similar to a Py-tract found downstream of the FAS exon 6 5’ splice site (Figure 3C), which is recognized by the protein TIA1 and enhances 5’ splice site recognition by U1 snRNP. The Py-tract in the central region of FAS exon 6 could, therefore, enhance recognition of upstream 5’ splice sites generated by exonic mutations.

[0233] These findings reveal that FAS exon 6 contains sequences that can be activated by simple mutations to function as bona fide 3’ and 5’-like splice sites, promoting the inclusion of very short exons even in the absence of regulatory sequence elements known to be involved in the activation of microexons. The evolutionary birth of new microexons is therefore likely to be simpler and more frequent than previously appreciated. Exonic Binding of U2 snRNP Promotes FAS Exon 6 Skipping

[0234] The presence of a relatively strong Py-tract preceded by a near-consensus branch point sequence within exon 6 (Figure 3C), and the inclusion of a shorter version of the exon when a mutation creates a functional 3’ splice site AG downstream of the Py-tract, opened the possibility that 3’ splice site recognizing factors assemble on FAS exon 6 (Figure 3L and 3M).

[0235] To directly assess whether U2 snRNP, the key ribonucleoprotein complex involved 3’ splice site recognition, can assemble on FAS exon 6 sequences, we incubated in vitro transcribed FAS exon 6 (wild type and mutants, all lacking the flanking splice sites) with HeLa nuclear extracts and measured the interaction by native gel electrophoresis. The results indicated that U2 snRNP can indeed assemble (complex A) on FAS exon 6, an interaction that was decreased upon mutation of the Py-tract or branch point sequences and enhanced upon introduction of an AG dinucleotide (Figure 3L and 3M and Figure 19).

[0236] It is conceivable that U2 snRNP assembly on the wild type exon, in the absence of a 3’ splice site, competes with recognition of the 3’ splice site of intron 5 by the splicing machinery and this contributes to modulate the levels of exon 6 inclusion. To test this possible mechanism, the saturation mutagenesis results were advantageously used. It was observed that exonic variants with an intact branchpoint-like sequence at positions 10 to 15 (CUAACU) displayed an average inclusion of 45%, whereas variants harbouring mutations at either of the two adenosines that could serve as branch sites in this sequence increased the levels of exon inclusion, and mutation of both adenosines further increased exon inclusion to an average of 75% (Figure 3N). Also consistent with this model, insertion of AG containing sequences after position 40 reduced full length exon inclusion, compared to insertion of non-AG-containing sequences, from an average of 75% to 25% inclusion (Figure 30).

[0237] Collectively, the results reveal a novel mechanism of exon skipping based upon assembly of U2 snRNP on exonic sequences that resemble (but cannot be active as) 3’ splice sites. This illustrates the value of saturation mutagenesis approaches to discover and test mechanistic hypotheses.

[0238] Cryptic 3’ Splice Sites Regulate Alternative Exons Encoding One Pass Transmembrane Helices It has previously been reported that the Py-tract binding protein U2AF2 binds to an exonic polypyrimidine tract in IL7R exon 6 (Schott et al., (2021) RAM 1087 27:571-583), promoting exon skipping. Like FAS exon 6, IL7R exon encodes a one-pass transmembrane helix. Interestingly, transmembrane helices are enriched in nonpolar amino acid residues that are encoded by codons with the highest number of pyrimidines (Figure 4A). Therefore transmembrane-encoding exons are expected to be rich in pyrimidines, allowing regulation by mechanisms similar to those described for IL7R exon 6 or FAS exon 6 (Figure 3). To investigate whether alternative exons, and particularly those coding for transmembrane domains, are generally regulated by this type of mechanism, SVM-BPfinder was used to scan exons throughout the genome harbouring a branchpoint motif followed by a polypyrimidine tract. For each input sequence, this tool returns a score (‘SVM score’) that reflects how strong a 3’ splice site is predicted to be. Interestingly, 23% of all exons had an SVM score greater than that of FAS exon 6 (1.19), suggesting that a significant proportion of exons across the genome may have cryptic 3’ splice sites or at least sequence elements that resemble 3’ splice site regions. We found that shorter exons (<100 nts) with SVM scores greater than 1.19 encoded the most hydrophobic amino acid sequences (Figure 4B), consistent with transmembrane domains. These exons had a lower average PSI compared to other exons (Figure 4C). These results are compatible with the existence of a category of (relatively) short exons containing 3’ splice site-like sequences and, in particular, those encoding individual transmembrane helices (Figure 4D), whose inclusion is decreased by exonic 3’ splice site-like sequences.

[0239] To experimentally validate this hypothesis, minigenes containing one-pass transmembrane domain-encoding exon 5 of CHODL and exon 6 of CXADR were built, and their exonic putative branchpoint adenosines were mutated (Figure 4E and 4F). In both cases, the mutations reduced exon skipping, suggesting that the levels of exon skipping were regulated by the recognition of 3’ splice site-like sequences. Cryptic 3’ splice sites in exons may therefore be a widespread mechanism to regulate the inclusion of alternative exons encoding transmembrane helices and thus modulate the balance between soluble and membrane bound protein isoforms.

[0240] Estimating PSI Values in the GTEx Dataset

[0241] The inventors estimated the PSI of exons in the GTEx dataset (GTEx Consortium, 2017) from the proportion of reads supporting exon inclusion in the GTEx junction read counts file (GTEx_Analysis_2016-01-15_v7_STARv2.4.2aJunctions.gct.gz; available for download at https: / / www.gtexportal.org / home / datasets). To do this, the quantifySplicing function from the Psichomics package in R66 was used. The minReads argument was set to 10 (such that a splicing event requires at least 10 reads for it to be quantified) and the eventType argument was set to ‘SE’ (instructing the quantifySplicing function to quantify alternative exon events). All estimates were based on the Psichomics hg19 / GRCh37 alternative splicing annotations.

[0242] Experimentally Validating Microexon Inclusion

[0243] Initially the microexon sequences observed in the output (i.e. those sequences corresponding to large deletions, Figure 18A) were cloned into plasmid vector backbone pCMV_FAS_exon4_exon6, and transfected. No inclusion band was found in the polyacrylamide gels (i.e. these sequences appear to have been skipped in essentially 100% of cases, Figure 18B).

[0244] Since the nucleotide composition of the PTB binding domain in the central region of FAS exon 6 is very similar to the polypyrimidine tract of a 3’ splice site, it was reasoned that an AG-containing insertion right after nucleotide 40 (e.g. CAG, AG, TA, A, CT A) could create a new 3’ splice site in this region of the exon. Such a splice site would be expected to result in the microexons detected in the deep mutagenesis experiment. These insertions (as well as the non-AG-containing GAC insertion as a negative control) were introduced into the vector backbone using site-directed mutagenesis (Agilent, 200523) using the relevant mutagenesis primers. The produced minigene constructs were then transfected into Hek293 cells in triplicate; and RT-PCR products were fractionated by electrophoresis on 6% native acrylamide gels.

[0245] Experimentally Testing Exon Inclusion with Different Splice Sites

[0246] To test the effects of mutations in the presence of different 3’ splice site strengths, partially complementary oligonucleotides were used in combination with TaqPlus precision (Agilent,600212) and PCR “around the world” (primers pointing in opposite directions from the mutagenesis site to amplify the full length of the plasmid) to replace the naturally weak 3’ splice site of FAS exon 6 (5’- UUUCAUAUAAAAUGUCCAAUGUUCCAACCUACAG-3’; SEQ ID NO: 68) with a strong 3’ splice site sequence (5’-UACUAACGGCUUUUUUUUCCUUUUUCAG-3’; SEQ ID NO: 69).

[0247] PTB / SRRM4 Overexpression Experiments

[0248] SRRM4 (or PTBP1) protein was overexpressed by co-transfecting minigenes containing a CAG or AG insertion after position 40 along with 1000 ng of pcDNA5_SRRM4_flag (or pcDNA5_PTBP1_T7) in lipofectamine 2000 for 24 hours. RT-PCR was then used to analyse the splicing ratios as described above.

[0249] Example 10. In Silico Deletion Mutagenesis Reveals the Regulatory Architecture of Exons Genome-Wide

[0250] To systematically analyse the regulatory architecture of exons throughout the genome, SpliceAl predictions were used to train a hidden Markov model with 3 states (Figure 6B): E (enhancer - corresponding to regions of the exon that promote skipping upon deletion), S (silencer - regions that promote inclusion upon deletion) and N (neutral - which have no consistent effect upon deletion). The model captured the regulatory architecture of FAS exon 6 as uncovered by deep indel mutagenesis experiments (as described herein): regions of the exon corresponding to inferred enhancers were predicted to be in state E, and regions corresponding to inferred silencers in state S (Figure 6C). These data suggest that the model can accurately detect splicing regulatory elements along an exon and / or intron sequence.

[0251] The trained hidden Markov model was first used to study the distribution of splicing regulatory element (SRE) lengths (i.e. stretches of nucleotides in the same E or S states within an exon) throughout the transcriptome. This data revealed that most exonic SREs are short, with a median length of 5 nts and a mean of 8.57 nts (compatible with the average binding site of various RNA- binding protein domains) and similar to what is found in FAS exon 6. Enhancers were predicted to be slightly shorter than silencers (median lengths = 4 vs 6 nts; mean lengths = 6.15 vs 10.70 nts; Figure 6D). The model also correctly interpreted exonic sequences that are part of splice site sequences and their immediate neighbourhood as inclusion-promoting (i.e. belonging to the E state, Figure 6E, Figure 20).

[0252] The model was then used to gain a comprehensive overview of the regulatory architecture of the entire exome. To do this, ternary plots are used to visualise all 18,551 exons in the dataset based on their predicted E / N / S states. These data revealed two strong trends. First, for exons shorter than 100 nts, the percentage of nts in the N state tends to be below 20%, irrespective of the proportion of nucleotides in the S or E states; while exons longer than 150 nts tend to have a higher than 20% percentage of nucleotides in the N state (Figure 6F). This suggests that the inclusion of short exons, such as FAS exon 6, may require a higher density of SREs compared to longer exons. Indeed, as exon length increases, the absolute number of nucleotides in the E or S states rises at a slower rate (solid grey line, Figure 21) than would be expected if the proportion of nucleotides in these states remained constant regardless of exon length (dotted line, Figure 21). Second, across highly-included exons, consistently fewer than 20% of nucleotides are in the S state, regardless of the proportion of nucleotides in the N or E states (Figure 6G). This suggests that maintaining exon inclusion relies more on a low proportion of silencers than on a high proportion of enhancers, as also evidenced by the high inclusion values of exons whose nucleotides predominantly fall into the N state (bottom right-hand corner in Figure 6G). Interestingly, FAS exon 6, which has intermediate inclusion levels, approximately aligns with this boundary, with 23% of its nucleotides predicted to be in the S state.

[0253] The deletion analysis revealed not only that nearly the entire sequence of FAS exon 6 is covered by SREs (as is apparently typical of short exons across the genome), but also that its enhancers and silencers alternate in a ‘checkerboard’ pattern along the exon. Genome-wide in silica deletion mutagenesis was used to evaluate if this ‘checkerboard’ pattern is likely to be common in other exons.

[0254] The first hypothesis was that exons encompassed entirely by SREs arranged in an alternating checkerboard pattern would distribute approximately 50% of their nucleotides in the S state and the remaining 50% in the E state. Ternary plots suggest that exons meeting this criterion are shorter than 100 nts (Figure 6F) and display relatively low levels of inclusion (Figure 6G). This hypothesis was then evaluated by counting the number of times the sequence of each exon in the dataset transitions from the E to the S state directly, and vice versa - without passing through the N state. For example, in the case of FAS exon 6, seven such transitions were found (Figure 6C), equivalent to 9.5 E / S state transitions per 100 nts of exon sequence. The sequences of short (<100 nt long) highly included (> 90%) exons were predicted to have very few E / S state transitions, with an average of 1 .5 E / S state transitions per 100 nts (Figure 22A, results for longer exons shown in Figure 22B and 22C). In contrast, short alternatively spliced exons included at lower levels (< 90%) had many more state changes (two-tailed Wilcoxon rank sum test p value < 2.2e-16), with an average of 3.7 E / S state transitions per 100 nts (Figure 22A, results for longer exons shown in Figure 22B and 22C). Repeating this analysis for E / N state transitions (i.e. where the sequence transitions from the E to the N state and vice versa, without passing through the S state) reveals that short highly-included exons have significantly more E / N state transitions per 100 nts compared to short exons with a PSI below 90% (median 2.4 vs 1.4, two-tailed Wilcoxon rank sum test p value < 2.2e-16, Figure 23A), in agreement with previous findings that constitutive exons are sustained by strong enhancers. Interestingly, this result did not hold true for exons longer than 100 nts (median 1.7 vs 1 .8, Wilcoxon rank sum test p value 0.57, Figure 23B and 23C).

[0255] Exons that are variably included are therefore predicted to have a high density of splicing regulatory elements, which suggests that their precise inclusion levels are tightly regulated and sensitive to mutation. The alternating pattern of enhancers and silencers further suggests that some of these regulatory domains likely act by modulating the function of a neighbouring domain (e.g. a silencer protein binding to its site might sterically prevent a neighbouring enhancer from being bound by an enhancer protein). On the other hand, the lower density of enhancer-silencer alternations in constitutive exons suggests that their high inclusion levels have not been achieved by fine-tuning binding of splicing regulatory machinery. Without wishing to be bound by theory, their high inclusion levels may be a function of their stronger splice sites.

[0256] Hidden Markov Model

[0257] The depmixS4 package in R was used to build a hidden Markov model that predicts the locations of exonic splicing enhancers (which promote exon inclusion) and silencers (which promote exon skipping) in each exon based on SpliceAl scores for 4mer deletions. Exons with lengths between 50 and 200 nts were used as input for the model, with each exon as a separate time series during training. The model has three hidden states: E (enhancer), N (neutral), and S (silencer), with mean scores of -0.15, 0, and 0.15, respectively. The standard deviations for the states were fixed at 0.1 , 0.025, and 0.1 to account for the variability of positive and negative values in the dataset. For example, although breaking a splice site has a much stronger effect than breaking a weak enhancer (resulting in a much more negative spliceAl score), both sequence elements should be classified together in the E state. In some cases a strong enhancer can be essential for exon inclusion and therefore have an equivalent effect to breaking a splice site.

[0258] CLAUSES

[0259] Aspects and embodiments of the disclosure are encompassed in the following clauses:

[0260] 1 . A method for identifying an antisense oligonucleotide that modulates splicing of an exon in a gene of interest, comprising: identifying a gene sequence comprising a target exon and optionally one or more adjacent introns; generating a plurality of deletion mutants of the gene sequence, wherein each of the plurality of deletion mutants lacks a predefined sequence of one or more nucleotides of the gene sequence; generating a dataset comprising of the gene sequence and each of the deletion mutants; scoring a plurality of the deletion mutants by: applying a computational model to extract features indicative of disruptions on splicing acceptor and / or donor sites; applying the at least one extracted feature to score a change in splicing efficiency for each of the plurality of deletion mutants, where the splicing efficiency is defined as a degree to which the target exon is included in transcript molecules generated from the gene of interest; selecting one or more deletion mutants having a score above a predefined threshold score; and identifying the predefined sequence that was deleted from the one or more deletion mutants as a target sequence for an antisense oligonucleotide for modulating splicing of the target exon.

[0261] 2. A method for identifying an antisense oligonucleotide that modulates splicing of an exon in a gene of interest, comprising: generating a dataset comprising of a plurality of deletion mutants of a plurality of target exons and optionally adjacent introns in a plurality of gene sequences across the genome, wherein each of the plurality of deletion mutants lacks a predefined sequence from one or plurality of the target exon and / or adjacent intron; scoring a plurality of the deletion mutants by: applying a computational model to extract features indicative of disruptions on splicing acceptor and / or donor sites; applying the at least one extracted feature to score a change in splicing efficiency for each of the plurality of deletion mutants, where the splicing efficiency is defined as a degree to which the target exon is included in transcript molecules generated from the gene; selecting at least one target exon and optional adjacent intron from the plurality of target exons and optionally one or more adjacent introns across the genome; selecting one or more deletion mutants of the at least one target exon and optional adjacent intron having a score above a predefined threshold score; and identifying the predefined sequence that was deleted from the one or more deletion mutants as a target sequence for an antisense oligonucleotide for modulating splicing of the target exon.

[0262] 2a. The method of Clause 1 or Clause 2, wherein scoring a plurality of the deletion mutants comprises defining a training or an input set comprising of the target exon and / or adjacent introns and a subset of the plurality of deletion mutants, and applying a computational model to extract features indicative of disruptions on splicing acceptor and / or donor sites in the subset.

[0263] 2b. The method of any of Clause 1 to 2a, wherein the gene sequence comprises or consists of the target exon or at least a part of the target exon.

[0264] 2c. The method of any of Clauses 1 to 2b, wherein the gene sequence comprises an intron adjacent to the target exon or at least a part of the intron adjacent to the target exon.

[0265] 3. The method of any of Clauses 1 to 2c, wherein the dataset comprising of a plurality of deletion mutants of one or more target exons and / or introns is generated in vitro or in silico.

[0266] 4. The method of any of Clauses 1 to 3, wherein a computational model generates a dataset comprising of a plurality of deletion mutants of one or more target exons and / or introns across the genome.

[0267] 5. The method of any of Clauses 1 to 4, wherein the predefined sequences deleted from the target exon and / or intron of each of the plurality of deletion mutants have the same number of nucleotides.

[0268] 6. The method of Clause 5, wherein the number of nucleotides deleted from the target exon and / or intron of each of the deletion mutants is from about 1 to 60 nucleotides, from about 2 to 50 nucleotides, from about 3 to 40 nucleotides, from about 4 to 30 nucleotides, from about 5 to 25 nucleotides, from about 7 to 24 nucleotides, from about 10 to 23 nucleotides, from about 15 to 22 nucleotides, or from about 18 to 21 nucleotides.

[0269] 7. The method of Clause 5 or Clause 6, wherein the number of nucleotides deleted from the target exon and / or intron of each of the deletion mutants is 18, 19, 20, 21 or 22 nucleotides; particularly 21 nucleotides.

[0270] 7a. The method of any of Clauses 5 to 7, wherein the length of the deletion mutant includes at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 99% of a full wild-type exon and / or intron sequence of the gene of interest.

[0271] 8. The method of any preceding clause, wherein applying a computational model involves selecting a computational model that closely predicts the effects of the deletion mutants.

[0272] 9. The method of Clause 8, wherein the computational model is selected by systematically evaluating at least one computational model relative to a comprehensive experimental dataset. 10. The method of Clause 9, wherein selecting a computational model involves comparing splicing efficiency of deletion mutants calculated from at least one computational model to that of an experimental dataset.

[0273] 11. The method of Clause 10, wherein comparing splicing efficiency involves comparing spearman correlation coefficient calculated from each computational method.

[0274] 12. The method of Clause 8 to 11 , wherein computational model includes SMS score, HAL, MMSplice, SpliceAl, and / or Pangolin.

[0275] 12a. The method of Clause 8 to 11 , wherein computational model includes SpliceAl and / or Pangolin.

[0276] 12b. The method of Clause 8 to 11 , wherein computational model includes SpliceAl.

[0277] 13. The method of any preceding clause, wherein a plurality of the deletion mutants refers to two or more, including up to all of the deletion mutants.

[0278] 14. The method of any preceding clause, comprising correlating the selected one or more deletion mutants to a corresponding antisense oligonucleotide that is capable of hybridising to the predefined sequence that was deleted from the selected one or more deletion mutants.

[0279] 15. The method of any preceding clause, comprising generating one or more antisense oligonucleotides that comprises or consists of a nucleotide sequence complementary to the predefined sequence that was deleted from the selected one or more deletion mutants.

[0280] 15a. The method of any preceding clause, comprising generating one or more antisense oligonucleotides that comprises or consists of a nucleotide sequence that is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90% or at least about 98% complementary to the predefined sequence that was deleted from the selected one or more deletion mutants.

[0281] 15b. The method of any preceding clause, comprising generating one or more antisense oligonucleotides that comprises or consists of a nucleotide sequence complementary to the predefined sequence that was deleted from the selected one or more deletion mutants over a sequence of between about 1 and 30 nucleotides, between about 2 and 28 nucleotides, between about 3 and 26 nucleotides, between about 4 and 24 nucleotides, between about 5 and 23 nucleotides, between about 7 and 22 nucleotides, between about 10 and 21 nucleotides, between about 12 and 20 nucleotides, or between about 15 and 19 nucleotides. 16. The method of any preceding clause, wherein the one or more antisense oligonucleotides comprises one or more nucleic acid modification, sugar modification, and / or backbone modification.

[0282] 17. The method of any preceding clause, wherein the target gene sequence or target exon and / or intron sequence includes at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 100% of a full wild-type exon and / or intron sequence of the gene of interest.

[0283] 18. The method of any preceding clause, wherein a splicing efficiency is measured in percent- spliced-in-value.

[0284] 19. The method of any preceding clause, wherein the location of the predefined sequence within the target exon and / or adjacent intron is unique for each of the plurality of deletion mutants. 19a. The method of any preceding clause, wherein the predefined sequence includes at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 100% of the splicing silencer sequences.

[0285] 19b. The method of any preceding clause, wherein the predefined sequence includes at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 100% of the splicing enhancer sequences.

[0286] 19c. The method of any preceding clause, wherein the predefined sequence is unique for each of the plurality of deletion mutants.

[0287] 20. The method of any preceding clause, wherein the length of each of the plurality of deletion mutants is shorter than the length of the target exon by the number of nucleotides in the predefined sequence that is deleted.

[0288] 21 . The method of Clause 2a or of any of Clauses 2b to 20, when dependent on Clause 2a, wherein the subset of the plurality of deletion mutants comprises at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 deletion mutants.

[0289] 22. The method of Clause 2a or of any of Clauses 2b to 21 , when dependent on Clause 2a, wherein a size of the subset is at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, or at least about 70% of the total number of deletion mutants. 23. The method of Clause 2a or of any of Clauses 2b to 22, when dependent on Clause 2a, wherein a size of the subset is less than about 95%, less than about 90%, less than about 85%, less than about 80%, less than about 75%, less than about 70%, less than about 65%, or less than about 60% of the total number of deletion mutants.

[0290] 24. The method of Clause 2a or of any of Clauses 2b to 23, when dependent on Clause 2a, wherein a size of the subset is between about 10% and 95%, between about 10% and 90%, between about 15% and 85%, between about 15% and 80%, between about 20% and 75%, between about 20% and 70%, between about 25% and 65%, or between about 25% and 60% of the total number of deletion mutants.

[0291] 25. The method of any preceding clause, wherein the features extracted from the computational model comprise at least one of: splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain.

[0292] 26. The method of any preceding clause, wherein the degree to which the target exon is included in transcript molecules generated from the gene is defined as the proportion of transcript molecules that incorporate the target exon.

[0293] 27. The method of any preceding clause, wherein the predefined threshold score is at least 0.025, at least 0.05, at least 0.2, or at least 0.4.

[0294] 28. The method of any preceding clause, wherein the predefined threshold score is less than 0.4, less than 0.2, less than 0.05, or less than 0.025.

[0295] 29. The method of any preceding clause, wherein at least one target exon is identified from a dataset comprising of a plurality of deletion mutants of a plurality of exons and / or introns across the genome.

[0296] 30. The method of any preceding clause, wherein applying the at least one extracted feature to score a change in splicing efficiency for each of the plurality of deletion mutants comprises taking the maximum score from the group of individual scores of the extracted feature or features for each deletion mutant as the score for the change in splicing efficiency for that deletion mutant.

[0297] 30a. The method of Clause 30, wherein at least one target exon is identified from the average of the maximum scores of the deletion mutants for each exon.

[0298] 30b. The method of Clause 25, or any of Clauses 26 to 30a when dependent on Clause 25, comprising determining the DANGO score for each deletion mutant, wherein the DANGO score is the maximum score taken from the individual scores for one or more of splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain, wherein the scores for splice site acceptor loss and splice site donor loss are first multiplied by -1 to generate a negative score; preferably wherein the DANGO score is the maximum score taken from the four individual scores for splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain, wherein the scores for splice site acceptor loss and splice site donor loss are first multiplied by -1 to generate a negative score.

[0299] 30c. The method of Clause 25, or any of Clauses 26 to 30b when dependent on Clause 25, wherein the scores are generated by SpliceAI.

[0300] 31 . The method of Clause 30b or Clause 31 , wherein at least one target exon is identified from the distribution of the DANGO scores for each exon and / or adjacent intron.

[0301] 32. The method of any preceding clause, wherein identifying the target gene sequence of a gene of interest involves selecting a target gene sequence based on screening from a plurality of candidate genes.

[0302] 33. A method for predicting or validating a change in splicing efficiency caused by an antisense oligonucleotide that is adapted to bind to a predefined sequence of a target gene sequence comprising a target exon and optionally one or more adjacent introns, the method comprising:

[0303] (a) providing a cell or plurality of cells, each comprising one of a library of deletion mutant expression cassettes, each deletion mutant expression cassette adapted to express a deletion mutant comprising a sequence deletion compared to the target gene sequence, wherein the sequence deletion comprises or consists of the predefined sequence of the target gene sequence, and wherein the deletion mutant expression cassette is capable of expressing RNA encoded by the deletion mutant expression cassette in the cell or plurality of cells;

[0304] (b) incubating the cell or plurality of cells for a predetermined period of time to obtain a cell culture wherein cells of the cell culture express RNA encoded by the deletion mutant expression cassette;

[0305] (c) isolating mRNA expressed from the deletion mutant expression cassette from the cell culture;

[0306] (d) assessing the level of the target exon that is incorporated into mRNA expressed from the deletion mutant expression cassette;

[0307] (e) comparing the level of the target exon that is incorporated into mRNA expressed from the deletion mutant expression cassette with the level of the target exon that is incorporated into mRNA expressed from a control expression cassette comprising the wild-type or natural target gene sequence.

[0308] 33a. A method for identifying an antisense oligonucleotide that modulates splicing of an exon in a gene of interest by binding to a predefined sequence of a target gene sequence comprising a target exon and optionally one or more adjacent introns, the method comprising: (a) providing a cell or plurality of cells, each comprising one of a library of deletion mutant expression cassettes, each deletion mutant expression cassette adapted to express a deletion mutant comprising a sequence deletion compared to a target wild-type or natural gene sequence, wherein the sequence deletion comprises or consists of the predefined sequence of the target gene sequence, and wherein the deletion mutant expression cassette is capable of expressing RNA encoded by the deletion mutant expression cassette in the cell or plurality of cells;

[0309] (b) incubating the cell or plurality of cells for a predetermined period of time to obtain a cell culture wherein cells of the cell culture express RNA encoded by the deletion mutant expression cassette;

[0310] (c) isolating mRNA expressed from the deletion mutant expression cassette from the cell culture;

[0311] (d) assessing the level of the target exon that is incorporated into mRNA expressed from the deletion mutant expression cassette;

[0312] (e) comparing the level of the target exon that is incorporated into mRNA expressed from the deletion mutant expression cassette with the level of the target exon that is incorporated into mRNA expressed from a control expression cassette comprising the wild-type or natural target gene sequence in the presence of an antisense oligonucleotide complementary to the predefined sequence of the target gene sequence.

[0313] 33b. A method for predicting or validating a change in splicing efficiency caused by an antisense oligonucleotide that is adapted to bind to a predefined sequence of a target gene sequence comprising a target exon and optionally one or more adjacent introns, the method comprising:

[0314] (a)(i) providing a cell or plurality of cells, each expressing the target gene sequence comprising the target exon;

[0315] (a)(ii) introducing into each cell or plurality of cells one of a library of antisense oligonucleotides, or one of a library of antisense oligonucleotide expression cassettes, each antisense oligonucleotide expression cassette capable of expressing in the cell or plurality of cells an antisense oligonucleotide, and wherein each of the antisense oligonucleotides is adapted to bind to the predefined sequence of the target gene sequence;

[0316] (b) incubating the cell or plurality of cells for a predetermined period of time to obtain a cell culture wherein cells of the cell culture are capable of expressing mRNA encoded by the target gene sequence and comprise or express one of the library of antisense oligonucleotides;

[0317] (c) isolating mRNA expressed from the target gene sequence from the cell culture;

[0318] (d) assessing the level of the target exon that is incorporated into mRNA expressed from the target gene sequence;

[0319] (e) comparing the level of the target exon that is incorporated into mRNA expressed from the target gene sequence in the presence of the antisense oligonucleotide with the level of the target exon that is incorporated into mRNA expressed from a control expression cassette comprising the wild-type or natural target gene sequence in cells in the absence of the antisense oligonucleotide, wherein mRNA expressed from the target gene sequence and lacking the target exon is considered a deletion mutant or variant. 34. The method of any of Clauses 33 to 33b, comprising: selecting a target exon and / or one or more adjacent introns; generating one or more deletion mutants of the target exon and / or one or more adjacent introns, wherein each of the one or more deletion mutants lacks a predefined sequence of one or more nucleotides of the target exon and / or one or more adjacent introns; and / or synthesizing an antisense oligonucleotide that comprises a sequence complementary to the predefined sequence and / or a sequence complementary to the sequence deletion of the target gene sequence.

[0320] 35. The method of any of Clauses 33 to 34, comprising, prior to step (a), transfecting or transforming a cell or plurality of cells with a vector comprising the expression cassette.

[0321] 36. The method of any of Clauses 33 to 35, wherein the cell or plurality of cells is HEK293.

[0322] 37. The method of any of Clauses 33 to 36, wherein the library of deletion mutant expression cassettes comprises of a plurality of variants of the target exon and optionally one or more adjacent introns.

[0323] 38. The method of any of Clauses 33 to 37, wherein the library of deletion mutant expression cassettes comprises all possible deletion mutants having a predefined sequence of a predetermined length.

[0324] 39. The method of any of Clauses 33 to 38, wherein the antisense oligonucleotide comprises one or more nucleic acid modification, sugar modification, and / or backbone modification at each nucleotide position.

[0325] 40. The method of any of Clauses 33 to 39, wherein the predetermined period of time is at least about 4 hours, at least about 6 hours, at least about 12 hours, at least about 24 hours, at least about 48 hours, or at least about 72 hours.

[0326] 41. The method of any of Clauses 33 to 39, where the predetermined period of time is less than about 4 hours, less than about 6 hours, less than about 12 hours, less than about 24 hours, less than about 48 hours, or less than about 72 hours.

[0327] 42. The method of any of Clauses 33 to 41 , wherein at least one antisense oligonucleotide sequence is selected based on the inclusion or incorporation level of the target exon, wherein the antisense oligonucleotide sequence selected corresponds to the sequence of an antisense oligonucleotide that is predicted or validated to modulate the splicing of an exon in a gene of interest by binding to the predefined sequence. 43. The method of Clause 42, wherein the selection of an antisense oligonucleotide sequence is based on the inclusion or incorporation level of the target exon in mRNA: (1) expressed from the deletion mutant expression cassette; or (2) expressed from the target gene sequence in the presence of the antisense oligonucleotide; and wherein the inclusion or incorporation level is less than about 95% of the control level, less than about 90% of the control level, less than about 80% of the control level, less than about 70% of the control level, less than about 60% of the control level, less than about 50% of the control level, less than about 40% of the control level, less than about 30% of the control level, less than about 20% of the control level, or less than about 10% of the control level.

[0328] 44. The method of any of Clauses 33 to 43, comprising synthesising cDNA from the isolated mRNA; and / or generating a cDNA library for sequencing.

[0329] 45. The method of Clause 44, wherein the step of synthesising cDNA from the isolated RNA comprises: designing primers to amplify cDNA from the expression cassette comprising the target exon or deletion mutant; and / or performing quantitative PCR to determine the amount or proportion of cDNA of the expression cassette that includes the target exon, and the amount or proportion of cDNA of the expression cassette that does not include the target exon.

[0330] 46. The method of any of Clauses 33 to 45, wherein the step of assessing the level of the target exon that is incorporated into mRNA expressed from: (1) the deletion mutant expression cassette; or (2) the target gene sequence in the presence of the antisense oligonucleotide, comprises quantifying the amount or proportion of the target exon that is incorporated into mRNA expressed from the deletion mutant expression cassette or the target gene sequence in the presence of the antisense oligonucleotide.

[0331] 47. The method of any of Clauses 33 to 46, comprising calculating a change in splicing efficiency by comparing the proportion of target exon incorporation in mRNA of cells expressing (1) the deletion mutant expression cassette; or (2) the target gene sequence in the presence of the antisense oligonucleotide, with the proportion of the target exon incorporation in mRNA of cells expressing the control expression cassette comprising the wild-type or natural target exon in the absence of antisense oligonucleotide.

[0332] 48. The method of Clause 47, wherein splicing efficiency of a deletion mutant or a variant is calculated in terms of the percentage of mature transcripts that include the exon (PSI).

[0333] 49. The method of Clause 47 or Clause 48, wherein PSI values of all possible deletion mutants or exon variants are calculated based on the raw enrichment score for each variant. 50. The method of Clause 49, wherein a raw enrichment score for a deletion mutant or variant x is calculated as the ratio between the frequency of the deletion mutant or variant x in the output and the input libraries to the frequency of wild type (WT) target exon in the output and input libraries based on the formula below, wherein an output library refers to a cDNA library from the mRNA encoded by the deletion mutant expression cassette, and an input library refers to a library of expression cassettes each comprising a variant or deletion mutant:

[0334] 51. The method of Clause 49 or Clause 50, wherein a PSI value fora deletion mutant or variant x is calculated by multiplying the PSI of WT with the ratio of exponents of ES between the deletion mutant or variant x and WT, using the formula:

[0335] 52. The method of any of Clauses 33 to 51 , comprising correlating a change in splicing efficiency with an antisense oligonucleotide designed to bind to the whole or part of the predefined sequence deleted from the target gene sequence of the corresponding deletion mutant expression cassette.

[0336] 53. The method of any of Clauses 33 to 52, wherein the method includes performing the method of any of Clauses 1 to 32 to identify a target sequence for an antisense oligonucleotide for modulating splicing of the target exon.

[0337] 54. A system for generating antisense oligonucleotides that modulate splicing, comprising: a database library that includes one or more sets of data, where each set comprises a gene sequence comprising a target exon and / or intron adjacent to the target exon, and a plurality of predefined sequences to be deleted from the target exon and / or intron adjacent to the target exon, in order to define a plurality of deletion mutants of the target exon and / or intron adjacent to the target exon; a processor in communication with the database, the processor being adapted to: receive a set of data from the database, the set of data defining a target exon and / or intron adjacent to the target exon and the plurality of predefined sequences; generate a plurality of deletion mutants of the target exon and / or intron adjacent to the target exon, wherein each of the plurality of deletion mutants lacks a predefined sequence of the target exon and / or intron adjacent to the target exon; generate a dataset comprising of the target exon and / or intron adjacent to the target exon and each of the deletion mutants; and score a plurality of the deletion mutants by: applying a computational model to extract features indicative of disruptions on splicing acceptor and / or donor sites; applying the at least one extracted feature to score a change in splicing efficiency for each of the plurality of deletion mutants, where the splicing efficiency is defined as a degree to which the target exon is included in transcript molecules generated from the corresponding gene sequence; select one or more deletion mutants having a score above a predefined threshold score; and identify the predefined sequence that was deleted from the one or more deletion mutants as a target sequence for an antisense oligonucleotide for modulating splicing of the target exon in the corresponding gene sequence.

[0338] 55. A system for generating antisense oligonucleotides that modulate splicing of a target exon, comprising: a database library that includes one or more sets of data, comprising a plurality of deletion mutants of a plurality of target exons and optionally adjacent introns in a plurality of gene sequences across the genome, and a plurality of predefined sequences to be deleted from one or a plurality of target exons and / or introns adjacent to the one or plurality of target exons, in order to define a plurality of deletion mutants of one or a plurality of gene sequences; a processor in communication with the database, the processor being adapted to: receive a set of data from the database, the set of data defining a gene sequence comprising a target exon and / or adjacent intron and the plurality of predefined sequences; generate a dataset comprising of a plurality of deletion mutants of a target exon and / or adjacent intron, wherein each of the plurality of deletion mutants lacks a predefined sequence of the target exon and / or adjacent intron; score a plurality of the deletion mutants by: applying a computational model to extract features indicative of disruptions on splicing acceptor and / or donor sites within the target exon and / or adjacent intron; applying the at least one extracted feature to score a change in splicing efficiency for each of the plurality of deletion mutants, where the splicing efficiency is defined as a degree to which the target exon is included in transcript molecules generated from the corresponding gene sequence; select one or more deletion mutants having a score above a predefined threshold score for the target exon; and identify the predefined sequence that was deleted from the one or more deletion mutants as a target sequence for an antisense oligonucleotide for modulating splicing of the target exon in the corresponding gene sequence. 55a. The system of Clause 55, comprising: generating a dataset comprising of a plurality of deletion mutants of a plurality of target exons and / or adjacent introns in a plurality of gene sequence; scoring a plurality of the deletion mutants of the plurality of gene sequences; and identifying a target exon and optionally an intron adjacent to a target exon based on the change in splicing efficiency score.

[0339] 55b. The system of any of Clauses 54 to 55a, wherein scoring a plurality of the deletion mutants comprises defining a training or an input set comprising of the target exon and / or adjacent introns and a subset of the plurality of deletion mutants, and applying a computational model to extract features indicative of disruptions on splicing acceptor and / or donor sites in the subset.

[0340] 56. The system of any of Clauses 54 to 55b, further comprising generating one or more antisense oligonucleotides that comprises a nucleotide sequence complementary to the identified predefined sequence or a partial sequence thereof.

[0341] 57. The system of any of Clauses 54 to 56, wherein the processor is configured to perform the steps of any of Clauses 1 to 32.

[0342] 58. One or more antisense oligonucleotide generated according to the method of any of Clauses 15 to 15b or any of Clauses 16 to 32 when dependent on any of Clauses 15 to 15b, or according to the system of any of Clauses 54 to 57.

[0343] 59. A method for identifying an antisense oligonucleotide that modulates splicing of an exon in a gene of interest, the method comprising: identifying a gene sequence comprising a target exon and optionally one or more adjacent introns; generating a plurality of variants of the gene sequence, wherein each of the plurality of variants includes one or more modification of the sequence of the target exon and / or adjacent introns; generating a dataset comprising of the gene sequence and each of the variants; scoring a plurality of the variants by: applying a computational model to extract features indicative of disruptions on splicing acceptor and / or donor sites; applying the at least one extracted feature to score a change in splicing efficiency for each of the plurality of variants, where the splicing efficiency is defined as a degree to which the target exon is included in transcript molecules generated from the gene of interest; selecting one or more variant having a score above a predefined threshold score. 60. A method for identifying an antisense oligonucleotide that modulates splicing of an exon in a gene of interest, comprising: generating a dataset comprising of a plurality of variants of a plurality of target exons and optionally adjacent introns, wherein each of the plurality of variants has a modified predefined sequence of the target exon and / or adjacent intron; scoring a plurality of the variants by: applying a computational model to extract features indicative of disruptions on splicing acceptor and / or donor sites; applying the at least one extracted feature to score a change in splicing efficiency for each of the plurality of deletion mutants, where the splicing efficiency is defined as a degree to which the target exon is included in transcript molecules generated from the gene; selecting a target exon from the plurality of target exons across the genome; selecting one or more variant having a score above a predefined threshold score for the selected target exon; and identifying the predefined sequence that was modified from the one or more variants as a target sequence for an antisense oligonucleotide for modulating splicing of the target exon in the gene of interest.

[0344] 61. The method of Clause 59 or Clause 60, wherein the variant is a deletion mutant, an insertion mutant or a substitution mutant.

[0345] 62. The method of any of Clauses 59 to 61 , wherein the variant is a deletion mutant, and the modification is the deletion of a predefined sequence from the target exon and / or adjacent intron of each of the plurality of deletion mutants.

[0346] 63. The method of Clause 61 or Clause 62, wherein the predefined sequences that are deleted from each of the plurality of deletion mutants have the same number of nucleotides.

[0347] 64. The method of Clause 62 or Clause 63, wherein the number of nucleotides deleted from the target exon of each of the deletion mutants is from about 1 to 60 nucleotides, from about 2 to 50 nucleotides, from about 3 to 40 nucleotides, from about 4 to 30 nucleotides, from about 5 to 25 nucleotides, from about 7 to 24 nucleotides, from about 10 to 23 nucleotides, from about 15 to 22 nucleotides, or from about 18 to 21 nucleotides.

[0348] 65. The method of any of Clauses 62 to 64, wherein the number of nucleotides deleted from the target exon of each of the deletion mutants is 18, 19, 20, 21 or 22 nucleotides; particularly 21 nucleotides.

[0349] 66. The method of any of Clauses 61 to 65, wherein the location of the predefined sequence within the target exon and / or adjacent intron is unique for each of the plurality of deletion mutants. 67. The method of any of Clauses 59 to 61 , wherein the variant is an insertion mutant, and the modification is the insertion of a predefined sequence into the target exon and / or adjacent intron of each of the plurality of insertion mutants.

[0350] 68. The method of Clause 67, wherein the length of the predefined sequence inserted into the target exon of each of the plurality of insertion mutants has a length of up to about 21 nucleotides, up to about 18 nucleotides, up to about 15 nucleotides, up to about 12 nucleotides or up to about 9 nucleotides.

[0351] 69. The method of Clause 67 or Clause 68, wherein the predefined sequence inserted into the target exon of each of the plurality of insertion mutants has a length of from about 1 to 21 nucleotides, from about 1 to 18 nucleotides, from about 1 to 15 nucleotides, from about 1 to 12 nucleotides, from about 1 to 9 nucleotides, from about 2 to 18 nucleotides, from about 2 to 15 nucleotides, from about 2 to 12 nucleotides, from about 2 to 9 nucleotides, from about 3 to 15 nucleotides, from about 3 to 12 nucleotides, from about 3 to 9 nucleotides, from about 4 to 12 nucleotides, from about 4 to 9 nucleotides, or from about 5 to 9 nucleotides.

[0352] 70. The method of any of Clauses 59 to 61 , wherein the variant is a substitution mutant, and the modification is the substitution of a predefined sequence of adjacent nucleotides from within the target exon and / or adjacent intron of each of the plurality of substitution mutants.

[0353] 71. The method of Clause 70, wherein the predefined sequence of adjacent nucleotides has from about 1 to 12 nucleotides, from about 1 to 9 nucleotides, from about 1 to 6 nucleotides, from about 1 to 3 nucleotides, from about 2 to 12 nucleotides, from about 2 to 9 nucleotides, from about 2 to 6 nucleotides, from about 3 to 12 nucleotides, from about 3 to 9 nucleotides, or from about 3 to 6 nucleotides.

[0354] 72. The method of Clause 70 or Clause 71 , wherein the predefined sequence of adjacent nucleotides has or consists of 1 , 2, 3, 4, 5 or 6 nucleotides.

[0355] 73. The method of any of Clauses 59 to 72, comprising computationally deleting, inserting or substituting the predefined sequence or predefined number of nucleotides from each of the plurality of variants of the target exon and / or adjacent intron.

[0356] 74. The method of any of Clauses 59 to 72, comprising experimentally deleting, inserting or substituting the predefined sequence or predefined number of nucleotides from each of the plurality of variants of the target exon and / or adjacent intron. 75. The system of any of Clauses 55 to 58, or the method of any of Clauses 59 to 74, comprising correlating the selected one or more variants to a corresponding antisense oligonucleotide

[0357] 75a. The system of any of Clauses 55 to 58, or the method of Clause 75, wherein the corresponding antisense oligonucleotide is an antisense oligonucleotide which has a sequence adapted to anneal to a sequence of the target exon and / or adjacent intron at a sequence region coterminous with, adjacent and / or overlapping the sequence comprising the one or more modification.

[0358] 76. The system of any of Clauses 55 to 58, or the method of Clause 75 or Clause 75a, wherein the antisense oligonucleotide is capable of: hybridising to the predefined sequence that was deleted from the selected one or more deletion mutants; or hybridising to a sequence adjacent to the position of insertion of the predefined sequence that was inserted into the selected one or more insertion mutants; or hybridising to, adjacent and / or overlapping the predefined sequence of the selected one or more substitution mutants.

[0359] 77. The system of any of Clauses 55 to 58, or the method of any of Clauses 59 to 72, comprising generating one or more antisense oligonucleotides that comprises or consists of: a nucleotide sequence complementary to the predefined sequence that was deleted from the selected one or more deletion mutants; a nucleotide sequence complementary to a sequence adjacent to the position of insertion of the predefined sequence that was inserted into the selected one or more insertion mutants; or a nucleotide sequence complementary to the predefined sequence that was substituted in the one or more substitution variants, and / or that is complementary to a sequence adjacent to the predefined sequence that was substituted in the one or more substitution variants.

[0360] 77a. The system or method of Clause 77, therein the sequence complementary is over a nucleotide length of between about 7 and 30 contiguous nucleotides, between about 8 and 28 contiguous nucleotides, between about 9 and 26 contiguous nucleotides, between about 11 and 24 contiguous nucleotides, between about 13 and 23 contiguous nucleotides, between about 14 and 22 contiguous nucleotides, between about 15 and 21 contiguous nucleotides, between about 16 and 20 contiguous nucleotides, or between about 17 and 19 contiguous nucleotides.

[0361] 77b. The system or method of Clause 77 or Clause 77a, therein the sequence complementary is over a nucleotide length of about 15 contiguous nucleotides, about 16 contiguous nucleotides, about 17 contiguous nucleotides, about 18 contiguous nucleotides, about 19 contiguous nucleotides, about 20 contiguous nucleotides, about 21 contiguous nucleotides, about 22 contiguous nucleotides, or about 23 contiguous nucleotides. 78. The system or method of any of Clauses 75 to 77b, wherein the one or more antisense oligonucleotides comprises one or more nucleic acid modification, sugar modification, and / or backbone modification.

[0362] 79. The method of any of Clauses 59 to 78, wherein the target exon and / or adjacent intron sequence includes at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 100% of a full wild-type exon and / or adjacent intron sequence of the gene of interest.

[0363] 80. The method of any of Clauses 59 to 79, wherein a splicing efficiency is measured in percent-spliced-in-value.

[0364] 81 . The method of any of Clauses 59 to 80, wherein a plurality of the variants refers to two or more variants, about 10 or more variants, about 30 or more variants, about 100 or more variants, about 500 or more variants, about 1 ,000 or more variants, about 10,000 or more variants, about 100,000 or more variants, about 1 ,000,000 or more variants, or up to all possible variants.

[0365] 82. The method of any of Clauses 59 to 81 , wherein the location of the one or more modification within the target exon and / or adjacent intron is unique for each of the plurality of variants.

[0366] 82a. The system of Clause 55b or the method of any of Clauses 59 to 82, wherein scoring a plurality of the deletion mutants comprises defining a training or an input set comprising of the target exon and / or adjacent introns and a subset of the plurality of deletion mutants, and applying a computational model to extract features indicative of disruptions on splicing acceptor and / or donor sites in the subset.

[0367] 83. The system or method of Clause 82a, wherein the subset of the plurality of deletion mutants comprises at least 10, at least 25, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 1 ,000, or at least 10,000 variants.

[0368] 84. The system of Clause 55b or the method of any of Clauses 59 to 83, wherein a size of the subset is less than about 95%, less than about 90%, less than about 85%, less than about 80%, less than about 75%, less than about 70%, less than about 65%, less than about 60%, less than about 50%, less than about 40%, less than about 30%, less than about 20%, or less than about 10% of the total number of variants.

[0369] 85. The system of Clause 55b or the method of any of Clauses 59 to 84, wherein a size of the subset is between about 10% and 95%, between about 10% and 90%, between about 15% and 85%, between about 15% and 80%, between about 20% and 75%, between about 20% and 70%, between about 25% and 65%, or between about 25% and 60% of the total number of variants.

[0370] 86. The system of any of Clauses 54 to 58 or the method of any of Clauses 59 to 85, wherein the features extracted from the computational model comprise at least one of: splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain.

[0371] 87. The system of any of Clauses 54 to 58 or the method of any of Clauses 59 to 86, wherein applying the at least one extracted feature to score a change in splicing efficiency for each of the plurality of deletion mutants comprises taking the maximum score from the group of individual scores of the extracted feature or features for each deletion mutant as the score for the change in splicing efficiency for that deletion mutant.

[0372] 88. The system or method of Clause 87, wherein at least one target exon is identified from the average of the maximum scores of the deletion mutants for each exon.

[0373] 89. The system or method of any of Clauses 86 to 88, comprising determining the DANGO score for each deletion mutant, wherein the DANGO score is the maximum score taken from the individual scores for one or more of splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain, wherein the scores for splice site acceptor loss and splice site donor loss are first multiplied by -1 to generate a negative score.

[0374] 90. The system or method of Clause 89, wherein the DANGO score for each deletion mutant is the maximum score taken from the four individual scores for splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain.

[0375] 91. The system or method of any of Clauses 86 to 90, wherein the scores are generated by SpliceAI.

[0376] 92. The system of any of Clauses 54 to 58 or 82a to 91 or the method of any of Clauses 59 to

[0377] 91 , wherein the degree to which the target exon is included in transcript molecules generated from the gene sequence is defined as the proportion of transcript molecules that incorporate the target exon.

[0378] 93. The system of any of Clauses 54 to 58 or 82a to 92 or the method of any of Clauses 59 to

[0379] 92, wherein identifying the target exon of a gene sequence or gene of interest involves screening from candidates of genes.

Claims

CLAIMS1 . A method for predicting or validating a change in splicing efficiency caused by an antisense oligonucleotide that is adapted to bind to a predefined sequence of a target gene sequence comprising a target exon and optionally one or more adjacent introns, the method comprising:(a) providing a cell or plurality of cells, each comprising one of a library of deletion mutant expression cassettes, each deletion mutant expression cassette adapted to express a deletion mutant comprising a sequence deletion compared to the target gene sequence, wherein the sequence deletion comprises or consists of the predefined sequence of the target gene sequence, and wherein the deletion mutant expression cassette is capable of expressing RNA encoded by the deletion mutant expression cassette in the cell or plurality of cells;(b) incubating the cell or plurality of cells for a predetermined period of time to obtain a cell culture wherein cells of the cell culture express RNA encoded by the deletion mutant expression cassette;(c) isolating mRNA expressed from the deletion mutant expression cassette from the cell culture;(d) assessing the level of the target exon that is incorporated into mRNA expressed from the deletion mutant expression cassette;(e) identifying the predefined sequence that was deleted from the one or more deletion mutants as a target sequence for an antisense oligonucleotide for modulating splicing of the target exon.

2. The method of claim 1 , comprising: comparing the level of the target exon that is incorporated into mRNA expressed from the deletion mutant expression cassette with the level of the target exon that is incorporated into mRNA expressed from a control expression cassette comprising the wild-type or natural target gene sequence; and / or comparing the level of the target exon that is incorporated into mRNA expressed from the deletion mutant expression cassette with the level of the target exon that is incorporated into mRNA expressed from a control expression cassette comprising the wild-type or natural target gene sequence in the presence of an antisense oligonucleotide complementary to the predefined sequence deleted from the target gene sequence.

3. A method for identifying an antisense oligonucleotide that modulates splicing of a target exon in a gene, the method comprising:(a) identifying a gene sequence comprising a target exon and optionally one or more adjacent introns;(b) generating a plurality of deletion mutants of the gene sequence, wherein each of the plurality of deletion mutants lacks a predefined sequence of one or more nucleotides of the gene sequence;(c) scoring a plurality of the deletion mutants by:applying a computational model to extract features indicative of disruptions on splicing acceptor and / or donor sites; applying the at least one extracted feature to score a change in splicing efficiency for each of the plurality of deletion mutants, where the splicing efficiency is defined as a degree to which the target exon is included in transcript molecules generated from the gene of interest;(d) selecting one or more deletion mutants having a score above a predefined threshold score; and(e) identifying the predefined sequence that was deleted from the one or more deletion mutants as a target sequence for an antisense oligonucleotide for modulating splicing of the target exon.

4. A system for generating antisense oligonucleotides that modulate splicing, comprising: a database library that includes one or more sets of data, where each set comprises a gene sequence comprising a target exon and / or intron adjacent to the target exon, and a plurality of predefined sequences to be deleted from the target exon and / or intron adjacent to the target exon, in order to define a plurality of deletion mutants of the gene sequence; a processor in communication with the database, the processor being adapted to: receive a set of data from the database, the set of data defining a target exon and / or intron adjacent to the target exon and the plurality of predefined sequences; generate a plurality of deletion mutants of the target exon and / or intron adjacent to the target exon, wherein each of the plurality of deletion mutants lacks a predefined sequence of the target exon and / or intron adjacent to the target exon; generate a dataset comprising of the target exon and / or intron adjacent to the target exon and each of the deletion mutants; and score a plurality of the deletion mutants by: optionally defining a training set comprising of the target exon and / or intron adjacent to the target exon and a subset of the plurality of deletion mutants; applying a computational model to extract features indicative of disruptions on splicing acceptor and / or donor sites within the target exon and / or adjacent intron; applying the at least one extracted feature to score a change in splicing efficiency for each of the plurality of deletion mutants, where the splicing efficiency is defined as a degree to which the target exon is included in transcript molecules generated from the gene sequence; select one or more deletion mutants having a score above a predefined threshold score; and identify the predefined sequence that was deleted from the one or more deletion mutants as a target sequence for an antisense oligonucleotide for modulating splicing of the target exon in the gene sequence.

5. The method of Claim 3 or the system of Claim 4, which comprises generating a dataset comprising the gene sequence and each of the deletion mutants and / or a plurality of deletion mutants of a plurality of target exons and / or introns across the genome; wherein the dataset is obtained computationally or non-computationally; and providing the dataset to the computational model.

6. The method of Claims 3 or 5 or the system of Claims 4 or 5, wherein the features extracted from the computational model comprise at least one of: splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain.

7. The method of any of Claims 3, 5 and 6 or the system of any of Claims 4 to 6, wherein applying the at least one extracted feature to score a change in splicing efficiency for each of the plurality of deletion mutants comprises taking the maximum score from the at least one extracted feature for each deletion mutant as a score for the change in splicing efficiency for that deletion mutant.

8. The method or system of Claim 7, comprising determining the DANGO score for each deletion mutant, wherein the DANGO score is the maximum score taken from the individual scores for one or more of splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain, wherein the scores for splice site acceptor loss and splice site donor loss are first multiplied by -1 to generate a negative score;9. The method or system of Claim 8, preferably wherein the DANGO score for each deletion mutant is the maximum score taken from the four individual scores for splice site acceptor loss, splice site acceptor gain, splice site donor loss, and splice site donor gain, wherein the scores for splice site acceptor loss and splice site donor loss are first multiplied by -1 to generate a negative score.

10. The method of any of Claims 1 to 3, and 5 to 9, or the system of any of Claims 4 to 9, wherein:(i) the number of nucleotides deleted from the target exon and / or intron of each of the deletion mutants is from about 1 to 60 nucleotides, from about 2 to 50 nucleotides, from about 3 to 40 nucleotides, from about 4 to 30 nucleotides, from about 5 to 25 nucleotides, from about 7 to 24 nucleotides, from about 10 to 23 nucleotides, from about 15 to 22 nucleotides, or from about 18 to 21 nucleotides; and / or(ii) the length of the deletion mutant includes at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 99% of a full wild-type exon and / or intron sequence of the gene or gene sequence; and / or(iii) the target gene sequence includes at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about70%, at least about 80%, at least about 90%, at least about 95%, or about 100% of a full wild-type exon and / or intron sequence of a gene of interest.11 . The method of Claim 3 or any of Claims 5 to 10 when dependent on Claim 3, or the system of Claim 4 or any of Claims 5 to 10 when dependent on Claim 4, comprising: correlating the selected one or more deletion mutant with a corresponding antisense oligonucleotide that is capable of at least partially hybridising to the predefined sequence that was deleted from the selected one or more deletion mutants.

12. The method of any of Claims 1 to 3 or 5 to 11 , or the system of any of Claims 4 to 11 , comprising generating, providing or designing one or more antisense oligonucleotides that comprises or consists of a nucleotide sequence at least partially complementary to the predefined sequence that was deleted from one or more deletion mutants; and / or one or more antisense oligonucleotides that comprises a nucleotide sequence complementary to the predefined sequence or part thereof; and / or one or more antisense oligonucleotides that capable of hybridising to the predefined sequence or part thereof.

13. The method of any of Claims 1 to 3 or 5 to 12, or the system of any of Claims 4 to 12, wherein a splicing efficiency is measured in percent-spliced-in-value.

14. The method of any of Claims 1 to 3 or 5 to 13, or the system of any of Claims 4 to 13, wherein the location of the predefined sequence within the target exon and / or adjacent intron is unique for each of the plurality of deletion mutants.

15. The method of any of Claims 1 to 3 or 5 to 14, or the system of any of Claims 4 to 14, wherein:(i) the predefined sequence includes at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 100% of the splicing silencer sequences; or(ii) the predefined sequence includes at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 100% of the splicing enhancer sequences.

Citation Information

Patent Citations

  • Methods for characterizing alternatively or aberrantly spliced mRNA isoforms

    US10724092B2

  • Method for screening splicing variants or events

    US20200392488A1