Gene system
Patent Information
- Application Number
- PCT/GB2026/050514
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-05-30
- Filing Date
- 2026-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure GB2026050514_01102026_PF_FP_ABST
Abstract
Description
[0001] GENE SYSTEM
[0002] BACKGROUND TO THE INVENTION
[0003] Alport syndrome is a genetic condition affecting approximately 1 in 5,000-10,000 of all individuals in continental Europe and the USA. Alport syndrome usually presents during childhood and is associated with a spectrum of phenotypes that include a progressive loss of kidney function, and which can also include hearing loss and eye abnormalities (see e.g. Savige, J., et al., 2021. European Journal of Human Genetics, 29(8), pp.1186-1197).
[0004] Alport syndrome is caused by pathogenic variants in the COL4A3, COL4A4 and COL4A5 genes, which result in abnormalities of the collagen IV a345 network of basement membranes. The condition can be transmitted in an X-linked, autosomal dominant, or autosomal recessive pattern, with X-linked being the common while autosomal recessive and autosomal dominant account for around 15% and 20% of cases respectively (see e.g. Warady, B.A., et al., 2020. Kidney medicine, 2(5), pp.639-649).
[0005] In the absence of treatment, renal disease progresses from microhematuria to proteinuria, progressive renal insufficiency and end-stage renal disease in all males with the X-linked form, and in all males and females with the autosomal recessive form. Alport syndrome can be diagnosed by genetic testing and current treatments include angiotensin converting enzyme (ACE) inhibitors or angiotensin receptor blockers (ARB) to delay onset of end-stage kidney disease. However, at present there is no way to prevent end-stage renal failure, with a renal transplant being the only option (see e.g. Kashtan, C.E et al., 2018. Kidney international, 93(5), pp.1045-1051).
[0006] There are significant challenges to overcome in developing a successful gene therapy for Alport syndrome. For example, whilst adeno-associated virus (AAV) vectors are a leading platform for gene delivery, the COL4A5, COL4A3 and COL4A4 polypeptides are each up to around 1600-1700 amino acids in length, rendering them challenging for delivery by an AAV vector, due to limited AAV cargo capacity.Thus, there are no approved gene therapies, or gene therapies in clinical development, for the treatment of Alport syndrome. Thus, there is a significant unmet need for gene therapies that can provide normally functioning COL4A3, COL4A4 or COL4A5 polypeptides in host cells and thereby provide a therapy for example for the prevention and / or treatment of Alport syndrome and related conditions (e.g. any condition in a subject who has a pathogenic variant of a COL4A3, COL4A4 or COL4A5 gene).
[0007] SUMMARY OF THE INVENTION
[0008] The invention concerns the surprising finding that a gene therapy vector system, which involves reconstitution of the coding sequence of a functional COL4A3, COL4A4 or COL4A5 polypeptide at the mRNA level is particularly effective for providing expression of these proteins in a host cell.
[0009] In a first aspect is provided a system for generating a coding sequence of a COL4A3, COL4A4 or COL4A5 polypeptide in a host cell comprising:
[0010] (a) a first nucleic acid molecule comprising a 5' coding sequence (CDS), wherein the 5' CDS encodes an N-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide, and a promoter upstream of the 5' CDS; and
[0011] (b) a second nucleic acid molecule comprising a 3' CDS, wherein the 3' CDS encodes a C-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide, and a promoter upstream of the 3' CDS,
[0012] wherein the coding sequence of the COL4A3, COL4A4 or COL4A5 polypeptide is formed via reconstitution at the mRNA level following delivery of the first nucleic acid molecule and the second nucleic acid molecule into the host cell.
[0013] In a second aspect is provided an isolated cell comprising the system according to the first aspect.
[0014] In a third aspect is provided a pharmaceutical composition comprising the system according to the first aspect or a cell according to the second aspect.In a fourth aspect is provided a use of a system according to the first aspect, the isolated cell of the second aspect, or the pharmaceutical composition of the third aspect for the manufacture of a medicament.
[0015] In a fifth aspect is provided a product comprising:
[0016] (a) a first nucleic acid molecule comprising a 5' coding sequence (CDS), wherein the 5' CDS encodes an N-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide, and a promoter upstream of the 5' CDS; and
[0017] (b) a second nucleic acid molecule comprising a 3' CDS, wherein the 3' CDS encodes a C-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide, and a promoter upstream of the 3' CDS,
[0018] as a combined preparation for simultaneous, separate or sequential use in therapy, wherein a CDS encoding the COL4A3, COL4A4 or COL4A5 polypeptide is reconstituted at the mRNA level and the COL4A3, COL4A4 or COL4A5 polypeptide is expressed following administration of the first nucleic acid molecule and the second nucleic acid molecule to a subject.
[0019] In a sixth aspect is provided a system according to the first aspect, an isolated cell according to the second aspect, a pharmaceutical composition according to the third aspect, or a product according to the fifth aspect for use in preventing and / or treating Alport Syndrome or for use in preventing and / or treating a condition in a subject who has a pathogenic variant in the COL4A3, COL4A4 or COL4A5 gene.
[0020] In a seventh aspect is provided a kit comprising a) the first nucleic acid molecule as defined in the first aspect; and b) the second nucleic acid molecule as defined in the first aspect.
[0021] In an eighth aspect is provided a method of transducing a cell ex vivo, wherein the cell is transduced with the system of the first aspect.
[0022] In a ninth aspect is provided a system for generating a coding sequence of a COL4A3, COL4A4 or COL4A5 polypeptide in a host cell comprising:(a) a first nucleic acid molecule comprising a 5' coding sequence (CDS), wherein the 5' CDS encodes an N-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide, and self-cleaving ribozyme sequence; and
[0023] (b) a second nucleic acid molecule comprising a 3' CDS, wherein the 3' CDS encodes a C-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide, and a self-cleaving ribozyme sequence,
[0024] wherein the coding sequence of the COL4A3, COL4A4 or COL4A5 polypeptide is formed via reconstitution at the mRNA level following delivery of the first nucleic acid molecule and the second nucleic acid molecule into the host cell.
[0025] In a tenth aspect is provided a product comprising:
[0026] (a) a first nucleic acid molecule comprising a 5' coding sequence (CDS), wherein the 5' CDS encodes an N-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide, and a self-cleaving ribozyme sequence; and
[0027] (b) a second nucleic acid molecule comprising a 3' CDS, wherein the 3' CDS encodes a C-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide, and a self-cleaving ribozyme sequence,
[0028] as a combined preparation for simultaneous, separate or sequential use in therapy, wherein a CDS encoding the COL4A3, COL4A4 or COL4A5 polypeptide is reconstituted at the mRNA level and the COL4A3, COL4A4 or COL4A5 polypeptide is expressed following administration of the first nucleic acid molecule and the second nucleic acid molecule to a subject.
[0029] DESCRIPTION OF THE FIGURES
[0030] Figure 1- Schematic of dual AAV vector systems encoding COL4A5 for reconstitution via mRNA trans-splicing using ribozymes: (A) Schematics showing example 5' AAV vectors encoding an N-terminal part of a wtCOL4A5 polypeptide utilising a TwistR ribozyme. (B) Schematics showing corresponding example 3' AAV vectors encoding a C-terminal part of a wtCOL4A5 polypeptide and utilising a TwistR or RzB ribozyme. CMV, promoter; TwistR, Twister selfcleaving ribozyme; RzB, RzB self-cleaving ribozyme; SD, splice donor; SA, splice acceptor; WPRE, enhancer.Figure 2 - Western blot showing protein expression in HEK293 cells transfected with different plasmid designs (as described in Examples 1 and 2) encoding human COL4A5 with and without mutations to remove cryptic splice donor sites: (A) Protein expression seen with anti-COL4A5 antibody (B) Protein expression seen with anti-V5 tag antibody.
[0031] Figure 3 - (A) Western blot showing protein expression seen with an anti-COL4A5 antibody in HEK293 cells transfected using a single plasmid expressing full-length COL4A5 as a positive control; (B) Western blot showing protein expression seen with an anti-COL4A5 antibody in human podocytes transduced with a dual AAV vector system with cryptic splice sites removed from the 5' CDS; and (C) Western blot showing protein expression seen with an anti-V5 antibody in human podocytes transduced with a dual AAV vector system with cryptic splice sites removed from the 5' CDS.
[0032] Figure 4 - Automated Capillary Western Blot showing protein expression seen with an anti-V5 antibody in human podocytes transduced with a dual AAV vector system with cryptic splice sites removed from the 5' CDS (PS005).
[0033] Figure 5 - (A) Western blot showing protein expression seen with an anti-COL4A5 antibody in AD-293 cells transduced with a dual AAV vector system with cryptic splice sites removed from the 5' CDS (PS005); and (B) Western blot showing protein expression seen with an anti-V5 antibody in AD-293 cells transduced with a dual AAV vector system with cryptic splice sites removed from the 5' CDS (PS005).
[0034] DETAILED DESCRIPTION OF THE INVENTION
[0035] In a first aspect is provided a system for generating a coding sequence of a COL4A3, COL4A4 or COL4A5 polypeptide in a host cell comprising:
[0036] (a) a first nucleic acid molecule comprising a 5' coding sequence (CDS), wherein the 5' CDS encodes an N-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide, and a promoter upstream of the 5' CDS; and(b) a second nucleic acid molecule comprising a 3' CDS, wherein the 3' CDS encodes a C-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide, and a promoter upstream of the 3' CDS,
[0037] wherein the coding sequence of the COL4A3, COL4A4 or COL4A5 polypeptide is formed via reconstitution at the mRNA level following delivery of the first nucleic acid molecule and the second nucleic acid molecule into the host cell.
[0038] The nucleic acid molecules described above may be vectors. Therefore, the system may be a dual vector system.
[0039] A dual vector system comprises a pair of gene therapy vectors, each delivering part of a nucleic acid to a cell. Such systems are particularly useful where the size of the nucleic acid needing to be delivered is too long for effective packaging and or delivery by a single vector. Evidently, there are added challenges with dual vector systems versus single vector systems, principally that the nucleic acids from each of the pair of gene therapy vectors must be reconstituted at some point downstream of transfection.
[0040] In the dual vector system provided herein, and in contrast to previously disclosed methods, the first and second vectors comprise promoters upstream of the 5' CDS and 3'CDS respectively, meaning both CDSs will be transcribed into mRNA in a transfected cell. The coding sequence of the COL4A3, COL4A4 or COL4A5 polypeptide is then reconstituted at the mRNA level by trans-ligation / trans-splicing, rather than at the DNA stage.
[0041] There are a number of technologies that may be adapted for use in facilitating the reconstitution of the COL4A3, A4 or A5 CDS at the RNA level, including the REVeRT technology and ribozyme mediated RNA-trans-splicing. REVeRT technology (as described for example in WO2020127831A1 and Riedmayr, L.M. et a / . mRNA trans-splicing dual AAV vectors for (epi)genome editing and gene therapy, Nat Commun 14, 6578 (2023), https: / / doi.org / 10.1038 / s41467-023-42386-0) involves the use of complementary binding domains in each of the first and second vector, these binding domains can hybridise to each other at the mRNA level and thus act as binding sites between the pre-mRNA encoded by the first vector and the pre-mRNA encoded by the second vector. The presence of these sites thusjoins the pre-mRNAs produced by first and second vector. The splicing machinery of the host cell can then act on this intermediate to remove the binding sites (and any other non-protein coding sequences) and join the 5'CDS to the 3'CDS and generate the reconstituted COL4 mRNA. The technology may also utilise a very strong splice acceptor site (SA) to enhance reconstitution efficiency.
[0042] Ribozyme mediated RNA-trans-splicing utilises the ribozyme mediated cleavage of RNA to facilitate trans-ligation of the pre-mRNAs produced from the first and second vectors. Each vector comprises a self-cleaving ribozyme sequence which, when transcribed to pre-mRNA, is capable of catalysing itself out of the mRNA creating ends suitable for joining for by the cellular machinery. Ribozyme mediated trans-splicing is described in WO 2021 / 158964 Al and WO2024 / 196855 Al.
[0043] The dual vector system provided herein may be for expressing a human COL4A5 polypeptide. The dual vector system may be forexpressing a full-length COL4A5 polypeptide, suitably a full-length human COL4A5 polypeptide. Such vector systems are of particular use in the treatment of diseases related to abnormal COL4A5 in humans.
[0044] The dual vector system provided herein may be for expressing a human COL4A3 polypeptide. The dual vector system may be forexpressing a full-length COL4A3 polypeptide, suitably a full-length human COL4A3 polypeptide. Such vector systems are of particular use in the treatment of diseases related to abnormal COL4A3 in humans.
[0045] The dual vector system provided herein may be for expressing a human COL4A4 polypeptide. The dual vector system may be for expressing a full-length COL4A4 polypeptide, suitably a full-length human COL4A4 polypeptide. Such vector systems are of particular use in the treatment of diseases related to abnormal COL4A4 in humans.
[0046] The CDS encoding the COL4A3, COL4A4 or COL4A5 polypeptide is not particularly limited. Any suitable CDS encoding a COL4A3, COL4A4 or COL4A5 polypeptide may be used. In some embodiments, the CDS encoding the COL4A3, COL4A4 or COL4A5 polypeptide comprises or consists of: (a) a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to SEQ ID NO: 27; (b) a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to SEQ ID NO: 28; or (c) a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to SEQ ID NO: 29. In some embodiments, the CDS encoding the COL4A3, COL4A4 or COL4A5 polypeptide comprises or consists of: (a) the nucleotide sequence of SEQ ID NO: 27; (b) the nucleotide sequence of SEQ ID NO: 28; or (c) the nucleotide sequence of SEQ ID NO: 29.
[0047] The COL4A3, COL4A4 or COL4A5 polypeptide is not particularly limited. Any suitable COL4A3, COL4A4 or COL4A5 polypeptide may be used. Suitably, the COL4A3, COL4A4 or COL4A5 polypeptide is a human COL4A3, COL4A4 or COL4A5 polypeptide. In some embodiments the COL4A3, COL4A4 or COL4A5 polypeptide comprises or consists of: (a) an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to any of SEQ ID NOs: 1 to 7; (b) an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to any of SEQ ID NOs: 8 to 20; or (c) an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to any of SEQ ID NOs: 21 to 26. In some embodiments the COL4A3, COL4A4 or COL4A5 polypeptide comprises or consists of: (a) an amino acid sequence having at least 70% identity to SEQ ID NO: 1; (b) an amino acid sequence having at least 70% identity to SEQ ID NO: 8; or (c) an amino acid sequence having at least 70% identity to SEQ ID NO: 21. In some embodiments, the COL4A3, COL4A4 or COL4A5 polypeptide comprises or consists of: (a) the amino acid sequence of SEQ ID NO: 1; (b) the amino acid sequence of SEQ ID NO: 8; or (c) the amino acid sequence of SEQ ID NO: 21.The 5' CDS encodes any suitable N-terminal part of a COL4A3, COL4A4orCOL4A5 polypeptide. Suitably, the 5' CDS encodes an N-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide having a length of from 400 amino acids to 1300 amino acids, from 500 amino acids to 1200 amino acids, from 600 amino acids to 1100 amino acids, or from 700 amino acids to 1000 amino acids. Suitably, the 5' CDS comprises exons 1-10 or more, exons 1-15 or more, exons 1-20 or more, or exons 1-25 or more of a CDS encoding a COL4A3, COL4A4 or COL4A5 polypeptide. Suitably, the 5' CDS comprises exons 1-45 or less, exons 1-40 or less, or exons 1-35 or less of a CDS encoding a COL4A3, COL4A4 or COL4A5 polypeptide. Suitably, the 5' CDS comprises or consists of: (a) at least exons 1-15 to at most exons 1-43 of a CDS encoding a COL4A3 polypeptide; (b) at least exons 1-14 to at most exons 1-39 of a CDS encoding a COL4A4 polypeptide; or (c) at least exons 1-15 to at most exons 1-41 of a CDS encoding a COL4A5 polypeptide. In some embodiments, the 5' CDS comprises or consists of: (a) exon 1-26, exons 1-27, exons 1-28, exons 1-29, exons 1-30, exons 1-31, exons 1-32, exons 1-33, exons 1-34, exons 1-35, exons 1-36, exons 1-37 or exons 1-38 of a CDS encoding a COL4A3 polypeptide; (b) exons 1-25, exons 1-26, exons 1-27, exons 1-28, exons 1-29, exons 1-30, exons 1-31, exons 1-32, exons 1-33, or exons 1-34 of a CDS encoding a COL4A4 polypeptide; or (c) exons 1-26, exons 1-27, exons 1-28, exons 1-29, exons 1-30, exons 1-31, exons 1-32, exons 1-33, exons 1-34, exons 1-35, exons 1-36, exons 1-37, exons 1-38 or exons 1-39 of a CDS encoding a COL4A5 polypeptide, or exons 1-40 of a CDS encoding a COL4A5 polypeptide. In some embodiments, the 5' CDS comprises or consists of: (a) a nucleotide sequence having at least 70% identity to any of SEQ ID NOs: 30 to 36; (b) a nucleotide sequence having at least 70% identity to any of SEQ ID NOs: 37 to 43; or (c) a nucleotide sequence having at least 70% identity to any of SEQ ID NOs: 44 to 50 or SEQ ID NO: 91. In some embodiments, the 5' CDS comprises or consists of: (a) a nucleotide sequence selected from any of SEQ ID NOs: 30 to 36; (b) a nucleotide sequence selected from any of SEQ ID NOs: 37 to 43; or (c) a nucleotide sequence selected from any of SEQ ID NOs: 44 to 50 or SEQ ID NO: 91.
[0048] The 3' CDS encodes any suitable C-terminal part of a COL4A3, COL4A4 or COL4A5 polypeptide. Suitably, the 3' CDS encodes ana C-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide having a length of from 500 amino acids to 1200 amino acids, from 600 amino acids to 1100 amino acids, or from 700 amino acids to 1000 amino acids. Suitably, the 3' CDS comprises: (a) exons 48-52 or more, exons 43-52 or more, exons 38-52 or more, or exons 33-52 or more of a CDS encoding a COL4A3 polypeptide; (b) exons 43-47 or more, exons 38-47 or more, exons 33-47 or more, or exons 28-47 or more of a CDS encoding a COL4A4 polypeptide; or (c) exons 47-51 or more, exons 42-51 or more, exons 37-51 or more, or exons 32-51 or more of a CDS encoding a COL4A5 polypeptide. Suitably, the 3' CDS comprises: (a) exons 18-52 or less, exons 23-52 or less, or exons 28-52 or less of a CDS encoding a COL4A3 polypeptide; (b) exons 13-47 or less, exons 18-47 or less, or exons 23-47 or less of a CDS encoding a COL4A4 polypeptide; or (c) exons 12-51 or less, exons 17-51 or less, or exons 22-51 or less of a CDS encoding a COL4A5 polypeptide. Suitably, the 3' CDS comprises or consists of: (a) at least exons 44-52 to at most exons 16-52 of a CDS encoding a COL4A3 polypeptide; (b) at least exons 40-47 to at most exons 15-47 of a CDS encoding a COL4A4 polypeptide; or (c) at least exons 42-51 to at most exons 16-51 of a CDS encoding a COL4A5 polypeptide. In some embodiments, the 3' CDS comprises or consists of: (a) exons 39-52, exons 38-52, exons 37-52, exons 36-52, exons 35-52, exons 34-52, exons 33-52, exons 32-52, exons 31-52, exons 30-52, exons 29-52, exons 28-52 or exons 27-52 of a CDS encoding a COL4A3 polypeptide; (b) exons 35-47, exons 34-47, exons 33-47, exons 32-47, exons 31-47, exons 30-47, exons 29-47, exons 28-47, exons l-M , or exons 26-47 of a CDS encoding a COL4A4 polypeptide; or (c) exons 40-51, exons 39-51, exons 38-51, exons 37-51, exons 36-51, exons 35-41, exons 34-51, exons 33-51, exons 32-51, exons 31-51, exons 30-51, exons 29-51, exons 28-51 or exons 27-51, of a CDS encoding a COL4A5 polypeptide. In some embodiments, the 3' CDS comprises or consists of: (a) a nucleotide sequence having at least 70% identity to any of SEQ ID NOs: 51 to 57; (b) a nucleotide sequence having at least 70% identity to any of SEQ ID NOs: 58 to 64; or (c) a nucleotide sequence having at least 70% identity to any of SEQ ID NOs: 92 to 94, SEQ ID NOs: 65 to 71 or SEQ ID NOs: 95 to 97. In some embodiments, the 3' CDS comprises or consists of: (a) a nucleotide sequence selected from any of SEQ ID NOs: 51 to 57; (b) a nucleotide sequence selected from any of SEQ ID NOs: 58 to 64; or (c) a nucleotide sequence selected from any of SEQ ID NOs: 92 to 94, SEQ ID NOs: 65 to 71 or SEQ ID NOs: 95 to 97.
[0049] The 5' CDS and the 3' CDS together provide the full-length CDS at the mRNA level upon codelivery of the first AAV vector and the second AAV vector and subsequent transcription. The 5' CDS and 3' CDS do not overlap, meaning there is no part of the full CDS which is common to both vectors. Suitably, the 5' CDS comprises exons 1-26, exons 1-27, exons 1-28, exons 1-29, exons 1-30, exons 1-31, exons 1-32, exons 1-33, exons 1-34 exons 1-35, exons 1-36, exons1-37, exons 1-38 or exons 1-39 of COL4A5 CDS; and the 3'CDS comprises exons 27-51, exons 28-51, exons 29-51, exons 30-51, exons 31-51, exons 32-51, exons 33-51, exons 34-51, exons 35-51 exons 36-51, exons 37-51, exons 38-51, exons 39-51 or exons 40-51 of COL4A5 CDS respectively. Suitably, the 5' CDS comprises exons 1-26, exons 1-27, exons 1-28, exons 1-29, exons 1-30, exons 1-31, exons 1-32, exons 1-33, exons 1-34, exons 1-35, exons 1-36, exons 1-37 or exons 1-38 of COL4A3 CDS; and the 3'CDS comprises exons 27-51, exons 28-52, exons 29-52, exons 30-52, exons 31-52, exons 32-52, exons 33-52, exons 34-52 exons 35-52, exons 36-52, exons 37-52, exons 38-52 or exons 39-52 of COL4A3 CDS, respectively. Suitably, the 5' CDS comprises exons 1-25, exons 1-26, exons 1-27, exons 1-28, exons 1-29, exons 1-30, exons 1-31, exons 1-32, exons 1-33 or exons 1-34 of COL4A4 CDS; and the 3'CDS comprises exons 26-47, exons l-M , exons 28-47, exons 29-47 exons 30-47, exons 31-47, exons 32-47, exons 33-47, exons 34-47 or exons 35-47 of COL4A4 CDS, respectively.
[0050] Suitably, the 5' CDS comprises or consists of exons 1-31 of COL4A5 CDS and the 3' CDS comprises or consists of exons 32-51 of COL4A5 CDS.
[0051] Suitably, the 5' CDS comprises or consists of exons 1-31 of COL4A3 CDS and the 3' CDS comprises or consists of exons 32-52 of COL4A3 CDS.
[0052] Suitably, the 5' CDS comprises or consists of exons 1-28 of COL4A4 CDS and the 3' CDS comprises or consists of exons 29-47 of COL4A4 CDS.
[0053] Suitably, the 5' CDS comprises or consists of exons 1-31 of COL4A4 CDS and the 3' CDS comprises or consists of exons 32-47 of COL4A4 CDS.
[0054] A suitable promoter for the promoter upstream of the 5' CDS and / or the promoter upstream of the 3' CDS are: a CMV promoter, a CBA promoter, a CAG promoter, a CB7 promoter, an EFla promoter, a CMV / EFla hybrid promoter, a NF-kB promoter, a pSE-7 promoter, a mPGK promoter, a mUla promoter, a U6 promoter, a U7 promoter, a MNDU3 promoter, a HLP promoter, an AAT promoter, an ALB promoter, a ApoE / AAT promoter, a EalbAAT promoter, a LP1 promoter, a TBG promoter, a TTR promoter, a SYN1 promoter, a NSE promoter, a tMCK promoter, a CK8 promoter, a MHCK7 promoter, a SMN promoter, a DES promoter, a RK promoter, a hRHO promoter, a GRK1 promoter, a hCAR promoter, a hRPE65p promoter, a P546promoter, a PR1.7 promoter, a hRSl promoter, a VMD2 promoter, an a-MHC promoter, a FREI promoter, a NPHS1 promoter, or a NPHS2 promoter.
[0055] Suitably one or both of the promoters are kidney-specific promoters. Suitably one or both of the promoters are podocyte specific promoters or tubular cell-specific promoters.
[0056] Suitably the promoter upstream of the 5' CDS and / or the 3'CDS is a NPHS1 promoter or a NPHS2 promoter, optionally wherein the promoter upstream of the 5' CDS and / or the 3'CDS is a minimal NPHS1 promoter, a minimal NPHS2 promoter or the NPHS1265 bp promoter.
[0057] The first and second vectors may be first and second AAV vectors respectively.
[0058] An AAV vector or AAV vector particle may comprise an AAV genome or a fragment or derivative thereof.
[0059] An AAV genome is a polynucleotide sequence, which may encode functions needed for production of an AAV particle. These functions include those operating in the replication and packaging cycle of AAV in a host cell, including encapsidation of the AAV genome into an AAV particle. Naturally occurring AAVs are replication-deficient and rely on the provision of helper functions in trans for completion of a replication and packaging cycle. Accordingly, the AAV genome of an AAV vector of the invention is typically replication deficient.
[0060] The AAV genome may be in single-stranded form (ssAAV), either positive or negative-sense, or alternatively in double-stranded form (dsAAV). The use of a double-stranded form allows bypass of the DNA replication step in the target cell and so can accelerate transgene expression. The maximum packaging capacity of the single-stranded form is larger than the double-stranded form. Suitably, the AAV genome is in single-stranded form.
[0061] AAVs occurring in nature may be classified according to various biological systems. The AAV genome may be from any naturally derived serotype, isolate or clade of AAV.AAV may be referred to in terms of their serotype. A serotype corresponds to a variant subspecies of AAV which, owing to its profile of expression of capsid surface antigens, has a distinctive reactivity which can be used to distinguish it from other variant subspecies. Typically, an AAV vector particle having a particular AAV serotype does not efficiently crossreact with neutralising antibodies specific for any other AAV serotype. AAV serotypes include AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10 and AAV11, and derivatives thereof.
[0062] The first vector genome and the second vector genome may each be the same or different serotypes. In some embodiments, the first vector genome and the second vector genome are the same serotype. Suitably, the AAV serotype is selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10 and AAV11, and derivatives thereof.
[0063] AAV may also be referred to in terms of clades or clones. This refers to the phylogenetic relationship of naturally derived AAVs, and typically to a phylogenetic group of AAVs which can be traced back to a common ancestor and includes all descendants thereof. Additionally, AAVs may be referred to in terms of a specific isolate, i.e. a genetic isolate of a specific AAV found in nature. The term genetic isolate describes a population of AAVs which has undergone limited genetic mixing with other naturally occurring AAVs, thereby defining a recognisably distinct population at a genetic level.
[0064] Typically, the AAV genome of a naturally derived serotype, isolate or clade of AAV comprises at least one inverted terminal repeat sequence (ITR). An ITR sequence acts in cis to provide a functional origin of replication and allows for integration and excision of the vector from the genome of a cell. ITRs may be the only sequences required in cis next to the therapeutic gene. An AAV genome may also comprise packaging genes, such as rep and / or cap genes which encode packaging functions for an AAV particle. A promoter may be operably linked to each of the packaging genes. Specific examples of such promoters include the p5, pl9 and p40 promoters. For example, the p5 and pl9 promoters are generally used to express the rep gene, while the p40 promoter is generally used to express the cap gene. The rep gene encodes one or more of the proteins Rep78, Rep68, Rep52 and Rep40 or variants thereof. The cap gene encodes one or more capsid proteins such as VP1, VP2 and VP3 or variants thereof. Theseproteins make up the capsid of an AAV particle, which determines the AAV serotype. VP1, VP2, and VP3 may be produced by alternate mRNA splicing (Trempe, J.P. and Carter, B.J., 1988. Journal of virology, 62(9), pp.3356-3363). Thus, VP1, VP2 and VP3 may have identical sequences, but wherein VP2 is truncated at the N-terminus relative to VP1, and VP3 is truncated at the N-terminus relative to VP2.
[0065] An AAV genome may be the full genome of a naturally occurring AAV. For example, a vector comprising a full AAV genome may be used to prepare an AAV vector or vector particle. Preferably, an AAV genome is derivatised for the purpose of administration to patients. Such derivatisation is standard in the art and the invention encompasses the use of any known derivative of an AAV genome, and derivatives which could be generated by applying techniques known in the art. An AAV genome may be a derivative of any naturally occurring AAV. Suitably, the AAV genome is a derivative of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, or AAV11. Suitably, the AAV genome is a derivative of AAV2.
[0066] The first vector genome and the second vector genome may each be derived from the same or different naturally occurring AAVs. In some embodiments, the first vector genome and the second vector genome are each derived from the same naturally occurring AAV. Suitably, the AAV genome is a derivative of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, or AAV11. Suitably, the AAV genome is a derivative of AAV2.
[0067] Derivatives of an AAV genome include any truncated or modified forms of an AAV genome which allow for expression of a transgene from an AAV vector of the invention in vivo. Typically, it is possible to truncate the AAV genome significantly to include minimal viral sequence yet retain the above function. This is preferred for safety reasons to reduce the risk of recombination of the vector with wild-type virus, and also to avoid triggering a cellular immune response by the presence of viral gene proteins in the target cell.
[0068] Typically, a derivative will include at least one inverted terminal repeat sequence (ITR), preferably more than one ITR, such as two ITRs or more. One or more of the ITRs may be derived from AAV genomes having different serotypes or may be a chimeric or mutant ITR. A preferred mutant ITR is one having a deletion of a trs (terminal resolution site). This deletionallows for continued replication of the genome to generate a single-stranded genome which contains both coding and complementary sequences, i.e. a self-complementary AAV (scAAV) genome. This allows for bypass of DNA replication in the target cell, and so enables accelerated transgene expression. However, the maximum packaging capacity of a scAAV is reduced. Suitably, the AAV genome is not a scAAV genome.
[0069] Suitably, the one or more ITRs flank the AAV genome at either end. The inclusion of one or more ITRs may aid concatamer formation of an AAV vector in the nucleus of a host cell, for example following the conversion of single-stranded vector DNA into double-stranded DNA by the action of host cell DNA polymerases. The formation of such episomal concatamers protects the AAV vector during the life of the host cell, thereby allowing for prolonged expression of the protein-coding sequence in vivo. Suitably, an AAV genome may comprise a 5' ITR and a 3' ITR.
[0070] The AAV genome may comprise one or more ITR sequences from any naturally derived serotype, isolate or clade of AAV or a variant thereof. The AAV genome may comprise at least one, such as two, AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, or AAV11 ITRs, or variants thereof. Suitably, the AAV genome may comprise at least one, such as two, AAV2 ITRs.
[0071] The first vector genome and the second vector genome may comprise the same or different ITRs. In some embodiments, the first vector genome and the second vector genome each comprise one or more ITR sequences which are the same. Suitably, the first vector genome and the second vector genome each comprise at least one, such as two, AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, or AAV11 ITRs, or variants thereof. Suitably, the first vector genome and the second vector genome each comprise at least one, such as two, AAV2 ITRs.
[0072] ITR sequences for use in the invention may be derived from the following accession numbers for AAV whole genome sequences: Adeno-associated virus 1 NC_002077; Adeno-associated virus 2 NC_001401; Adeno-associated virus 3 NC_001729; Adeno-associated virus 4 NC_001829; Adeno-associated virus 5 NC_006152; Adeno-associated virus 6 AF028704;Ade no-associated virus 7 NC_006260; and Adeno-associated virus 8 NC_006261. Other suitable ITRs are described in Wilmott, P., et al., 2019. Human gene therapy methods, 30(6), pp.206-213.
[0073] The inclusion of one or more ITRs is preferred to aid concatamer formation of the AAV vector in the nucleus of a host cell, for example following the conversion of single-stranded vector DNA into double-stranded DNA by the action of host cell DNA polymerases. The formation of such episomal concatamers protects the AAV vector during the life of the host cell, thereby allowing for prolonged expression of the transgene in vivo.
[0074] Suitably, ITR elements will be the only sequences retained from the native AAV genome in the derivative. A derivative will preferably not include the rep and / or cap genes of the native genome and any other sequences of the native genome. This is preferred for the reasons described above, and also to reduce the possibility of integration of the vector into the host cell genome. Additionally, reducing the size of the AAV genome allows for increased flexibility in incorporating other sequence elements (such as regulatory elements) within the vector in addition to the protein-coding sequence.
[0075] The following portions could therefore be removed in a derivative of the invention: one inverted terminal repeat (ITR) sequence, the replication (rep) and capsid (cap) genes. However, derivatives may additionally include one or more rep and / or cap genes or other viral sequences of an AAV genome. Naturally occurring AAV integrates with a high frequency at a specific site on human chromosome 19, and shows a negligible frequency of random integration, such that retention of an integrative capacity in the AAV vector may be tolerated in a therapeutic seting.
[0076] Suitably, the first vector comprises a 5' ITR and a 3' ITR. Suitably, the first vector comprises from 5' to 3': a 5' ITR, a promoter, optionally one or more regulatory sequences, the 5' CDS, and a 3' ITR. Suitably, the second vector comprises a 5' ITR and a 3' ITR. Suitably, the second vector comprises from 5' to 3': a 3' ITR, the 3' CDS, optionally one or more regulatory sequences, and a 3' ITR.The invention additionally encompasses the provision of sequences of an AAV genome in a different order and configuration to that of a native AAV genome. The invention also encompasses the replacement of one or more AAV sequences or genes with sequences from another virus or with chimeric genes composed of sequences from more than one virus. Such chimeric genes may be composed of sequences from two or more related viral proteins of different viral species.
[0077] AAV serotype and capsid proteins
[0078] An AAV vector particle may be encapsidated by capsid proteins. The serotype may facilitate the transduction of glomerular cells (e.g. podocytes), for example specific transduction of glomerular cells (e.g. podocytes).
[0079] The AAV vector particles may be a kidney-specific vector particle. Preferably, the AAV vector particles are glomerular-specific (e.g. podocyte-specific) vector particle. The AAV vector particles may be encapsidated by a glomerular-specific (e.g. podocyte-specific) capsid. The AAV vector particles may comprise a glomerular-specific (e.g. podocyte-specific) capsid protein.
[0080] Suitably, the AAV vector particles may be transcapsidated forms wherein an AAV genome or derivative having an ITR of one serotype is packaged in the capsid of a different serotype. The AAV vector particle also includes mosaic forms wherein a mixture of unmodified capsid proteins from two or more different serotypes makes up the viral capsid. The AAV vector particle also includes chemically modified forms bearing ligands adsorbed to the capsid surface. For example, such ligands may include antibodies for targeting a particular cell surface receptor.
[0081] Where a derivative comprises capsid proteins i.e. VP1, VP2 and / or VP3, the derivative may be a chimeric, shuffled or capsid-modified derivative of one or more naturally occurring AAVs. In particular, the invention encompasses the provision of capsid protein sequences from different serotypes, clades, clones, or isolates of AAV within the same vector (i.e. a pseudotyped vector). An AAV vector may be in the form of a pseudotyped AAV vector particle.Chimeric, shuffled or capsid-modified derivatives will be typically selected to provide one or more desired functionalities for the AAV vector. Thus, these derivatives may display increased efficiency of gene delivery, decreased immunogenicity (humoral or cellular), an altered tropism range and / or improved targeting of podocytes compared to an AAV vector comprising a naturally occurring AAV genome. Increased efficiency of gene delivery may be effected by improved receptor or co-receptor binding at the cell surface, improved internalisation, improved trafficking within the cell and into the nucleus, improved uncoating of the viral particle and improved conversion of a single-stranded genome to double-stranded form. Increased efficiency may also relate to an altered tropism range or targeting of podocytes, such that the vector dose is not diluted by administration to tissues where it is not needed.
[0082] Chimeric capsid proteins include those generated by recombination between two or more capsid coding sequences of naturally occurring AAV serotypes. This may be performed for example by a marker rescue approach in which non-infectious capsid sequences of one serotype are co-transfected with capsid sequences of a different serotype, and directed selection is used to select for capsid sequences having desired properties. The capsid sequences of the different serotypes can be altered by homologous recombination within the cell to produce novel chimeric capsid proteins.
[0083] Chimeric capsid proteins also include those generated by engineering of capsid protein sequences to transfer specific capsid protein domains, surface loops or specific amino acid residues between two or more capsid proteins, for example between two or more capsid proteins of different serotypes.
[0084] Shuffled or chimeric capsid proteins may also be generated by DNA shuffling or by error-prone PCR. Hybrid AAV capsid genes can be created by randomly fragmenting the sequences of related AAV genes e.g. those encoding capsid proteins of multiple different serotypes and then subsequently reassembling the fragments in a self-priming polymerase reaction, which may also cause crossovers in regions of sequence homology. A library of hybrid AAV genes created in this way by shuffling the capsid genes of several serotypes can be screened to identify viral clones having a desired functionality. Similarly, error prone PCR may be used to randomlymutate AAV capsid genes to create a diverse library of variants which may then be selected for a desired property.
[0085] The sequences of the capsid genes may also be genetically modified to introduce specific deletions, substitutions or insertions with respect to the native wild-type sequence. In particular, capsid genes may be modified by the insertion of a sequence of an unrelated protein or peptide within an open reading frame of a capsid coding sequence, or at the N-and / or C-terminus of a capsid coding sequence. The unrelated protein or peptide may advantageously be one which acts as a ligand for a particular cell type, thereby conferring improved binding to a target cell or improving the specificity of targeting of the vector to a particular cell population. The unrelated protein may also be one which assists purification of the viral particle as part of the production process, i.e. an epitope or affinity tag. The site of insertion will typically be selected so as not to interfere with other functions of the viral particle e.g. internalisation, trafficking of the viral particle.
[0086] The capsid protein may be an artificial or mutant capsid protein. The term "artificial capsid" as used herein means that the capsid particle comprises an amino acid sequence which does not occur in nature, or which comprises an amino acid sequence which has been engineered (e.g. modified) from a naturally occurring capsid amino acid sequence. In other words, the artificial capsid protein comprises a mutation or a variation in the amino acid sequence compared to the sequence of the parent capsid from which it is derived where the artificial capsid amino acid sequence and the parent capsid amino acid sequences are aligned.
[0087] The capsid protein may comprise a mutation or modification relative to the wild type capsid protein which improves the ability to transduce podocytes relative to an unmodified or wild type viral particle. Improved ability to transduce podocytes may be measured for example by measuring the expression of a reporter transgene, e.g. GFP, carried by the AAV vector particle, wherein expression of the transgene in podocytes correlates with the ability of the AAV vector particle to transduce podocytes.The first vector genome and the second vector genome may each comprises the same or different capsid proteins. In some embodiments, the first vector and the second vector comprise the same capsid proteins.
[0088] The first vector and / or the second vector may be an AAV3B, LK03, KPI, KP2, KP3, AAV9, ShHIO, AAV-DJ, AAV2, AAV6.2, AAV5, or AAV8 vector particle. In some embodiments, the first vector and the second vector is an AAV3B, LK03, KPI, KP2, KP3, AAV9, ShHIO, AAV-DJ, AAV2, AAV6.2, AAV5, or AAV8 vector particle. In some embodiments, the first vector and / or the second vector is an AAV3B vector particle or an LK03 vector particle. In one embodiment, the first vector and / or the second vector is an LK03 vector particle. In other embodiments, the first vector and / or the second vector is a ShHIO, AAV-DJ, AAV2, AAV6.2, KPI, KP2, KP3, or AAV5 vector particle. In some embodiments, the first vector and the second vector is an AAV3B vector particle or an LK03 vector particle. In one embodiment, the first vector and the second vector is an LK03 vector particle. In other embodiments, the first vector and the second vector is a ShHIO, AAV-DJ, AAV2, AAV6.2, KPI, KP2, KP3, or AAV5 vector particle.
[0089] The first vector particle and / or the second vector particle may comprise an AAV3B, LK03, AAV9, ShHIO, AAV-DJ, AAV2, AAV6.2, AAV5, KPI, KP2, KP3, or AAV8 capsid protein. In some embodiments, the first vector particle and / or the second vector particle each comprise an AAV3B, LK03, AAV9, ShHIO, AAV-DJ, AAV2, AAV6.2, AAV5, KPI, KP2, KP3, or AAV8 capsid protein. In some embodiments, the first vector particle and / or the second vector particle comprises an AAV3B capsid protein or an LK03 capsid protein. In one embodiment, the first vector particle and / or the second vector particle comprises an LK03 capsid protein. In other embodiments, the first vector particle and / or the second vector particle comprises a ShHIO, AAV-DJ, AAV2, AAV6.2, KPI, KP2, KP3, or AAV5 capsid protein. In some embodiments, the first vector particle and the second vector particle comprises an AAV3B capsid protein or an LK03 capsid protein. In one embodiment, the first vector particle and the second vector particle comprises an LK03 capsid protein. In other embodiments, the first vector particle and the second vector particle comprises a ShHIO, AAV-DJ, AAV2, AAV6.2, KPI, KP2, KP3, or AAV5 capsid protein.The first vector particle and / or the second vector particle may comprise AAV3B, LK03, AAV9, ShHIO, AAV-DJ, AAV2, AAV6.2, AAV5, KPI, KP2, KP3, or AAV8 capsid proteins VP1, VP2 and VP3. In some embodiments, the first vector particle and the second vector particle each comprise AAV3B, LK03, AAV9, ShHIO, AAV-DJ, AAV2, AAV6.2, AAV5, KPI, KP2, KP3, or AAV8 capsid proteins VP1, VP2 and VP3. In some embodiments, the first vector particle and / or the second vector particle comprises AAV3B or LK03 capsid proteins VP1, VP2 and VP3. In one embodiment, the first vector particle and / or the second vector particle comprises LK03 capsid proteins VP1, VP2 and VP3. In other embodiments, the first vector particle and / or the second vector particle comprises ShHIO, AAV-DJ, AAV2, AAV6.2, KPI, KP2, KP3, or AAV5 capsid proteins VP1, VP2 and VP3. In some embodiments, the first vector particle and the second vector particle comprise AAV3B or LK03 capsid proteins VP1, VP2 and VP3. In one embodiment, the first vector particle and the second vector particle comprises LK03 capsid proteins VP1, VP2 and VP3. In other embodiments, the first vector particle and the second vector particle comprise ShHIO, AAV-DJ, AAV2, AAV6.2, KPI, KP2, KP3, or AAV5 capsid proteins VP1, VP2 and VP3.
[0090] The first vector particle and / or the second vector particle may comprise one or more AAV2 ITR sequences and AAV3B capsid proteins, LK03 capsid proteins, AAV9 capsid proteins, ShHIO capsid proteins, AAV-DJ capsid proteins, AAV2 capsid proteins, AAV6.2 capsid proteins, AAV5 capsid proteins, KPI, KP2, KP3, or AAV8 capsid proteins. In some embodiments, the first vector particle and the second vector particle each comprise one or more AAV2 ITR sequences and AAV3B capsid proteins, LK03 capsid proteins, AAV9 capsid proteins, ShHIO capsid proteins, AAV-DJ capsid proteins, AAV2 capsid proteins, AAV6.2 capsid proteins, AAV5 capsid proteins, KPI, KP2, KP3, or AAV8 capsid proteins. In some embodiments, the first vector particle and / or the second vector particle comprises one or more AAV2 ITR sequences and AAV3B or LK03 capsid proteins. In other embodiments, the first vector particle and / or the second vector particle comprises one or more AAV2 ITR sequences and ShHIO, AAV-DJ, AAV2, AAV6.2, KPI, KP2, KP3, or AAV5 capsid proteins. In some embodiments, the first vector particle and the second vector particle comprise one or more AAV2 ITR sequences and AAV3B or LK03 capsid proteins. In other embodiments, the first vector particle and the second vector particle comprise one or more AAV2 ITR sequences and ShHIO, AAV-DJ, AAV2, AAV6.2, KPI, KP2, KP3, or AAV5 capsid proteins.The first vector particle and / or the second vector particle may have an AAV2 genome and AAV3B capsid proteins (AAV2 / 3B), an AAV2 genome and LK03 capsid proteins, an AAV2 genome and AAV9 capsid proteins (AAV2 / 9), an AAV2 genome and ShHIO capsid proteins, an AAV2 genome and AAV-DJ capsid proteins, an AAV2 genome and AAV2 capsid proteins, an AAV2 genome and AAV6.2 capsid proteins, an AAV2 genome and AAV5 capsid proteins, or an AAV2 genome and AAV8 capsid proteins (AAV2 / 8). In some embodiments, the first vector particle and the second vector particle each have an AAV2 genome and AAV3B capsid proteins (AAV2 / 3B), an AAV2 genome and LK03 capsid proteins, an AAV2 genome and AAV9 capsid proteins (AAV2 / 9), an AAV2 genome and ShHIO capsid proteins, an AAV2 genome and AAV-DJ capsid proteins, an AAV2 genome and AAV2 capsid proteins, an AAV2 genome and AAV6.2 capsid proteins, an AAV2 genome and AAV5 capsid proteins, an AAV2 genome and KPI capsid proteins, an AAV2 genome and KP2 capsid proteins, an AAV2 genome and KP3 capsid proteins, or an AAV2 genome and AAV8 capsid proteins (AAV2 / 8).
[0091] The nomenclature AAVX / Y may denote a pseudotyped AAV, for example where the ITR sequences are from AAVX and flank a cassette harbouring a payload which is encapsidated into serotype AAVY (i.e. with AAVY capsid proteins).
[0092] LK03 serotype
[0093] The first vector particle and / or the second vector particle may comprise a LK03 capsid protein. Suitably, the first vector particle and / or the second vector particle may be encapsidated by LK03 capsid proteins.
[0094] In some embodiments, the first vector particle and the second vector particle each comprise a LK03 capsid protein. In some embodiments, the first vector particle and the second vector particle are both encapsidated by LK03 capsid proteins.
[0095] The AAV-LK03 cap sequence consists of fragments from seven different wild-type serotypes (AAV1, 2, 3B, 4, 6, 8, 9) and is described in Lisowski, L., et al., 2014. Nature, 506(7488), pp.382-386. The present inventors have demonstrated that AAV-LK03 vectors can achieve high transduction in human podocytes in vitro.The first vector particle and / or the second vector particle may comprise an LK03 VPl capsid protein, an LK03 VP2 capsid protein, and / or an LK03 VP3 capsid protein. Suitably, the first vector particle and / or the second vector particle may be encapsidated by LK03 VPl capsid proteins, LK03 VP2 capsid proteins, and / or LK03 VP3 capsid proteins. Suitably, the first vector particle and / or the second vector particle may be encapsidated by LK03 VPl, VP2, and VP3 capsid proteins.
[0096] Suitably, the LK03 VPl capsid protein may comprise or consist of the amino acid sequence shown as SEQ ID NO: 81, or a variant which is at least 90% identical to SEQ ID NO: 81.
[0097] Suitably, the variant may be at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 81.
[0098] Suitably, the LK03 VP2 and VP3 capsid proteins may be N-terminal truncations of SEQ ID NO: 81, or N-terminal truncations of a variant which is at least 90% identical, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 81.
[0099] AAV3B serotype
[0100] The first vector particle and / or the second vector particle may comprise an AAV3B capsid protein. Suitably, the first vector particle and / or the second vector particle may be encapsidated by AAV3B capsid proteins.
[0101] In some embodiments, the first vector particle and the second vector particle each comprise an AAV3B capsid protein. In some embodiments, the first vector particle and the second vector particle are both encapsidated by AAV3B capsid proteins.
[0102] Two distinct AAV3 isolates (AAV3A and AAV3B) have been cloned. In comparison with vectors based on other AAV serotypes, it is thought that AAV3 vectors inefficiently transduce most cell types. However, AAV3B may efficiently transduce podocytes. AA3B has been described in Rutledge, E.A., et al., 1998. Journal of virology, 72(1), pp.309-319.The AAV vector particle may comprise an AAV3B VP1 capsid protein, an AAV3B VP2 capsid protein, and / or an AAV3B VP3 capsid protein. Suitably, the AAV vector particle may be encapsidated by AAV3B VP1 capsid proteins, AAV3B VP2 capsid proteins, and / or AAV3B VP3 capsid proteins. Suitably, the AAV vector particle may be encapsidated by AAV3B VP1, VP2, and VP3 capsid proteins.
[0103] Suitably, the AAV3B VP1 capsid protein may comprise or consist of the amino acid sequence shown as SEQ ID NO: 82, or a variant which is at least 90% identical to SEQ ID NO: 82.
[0104] Suitably, the variant may be at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 82.
[0105] Suitably, the AAV3B VP2 and VP3 capsid proteins may be N-terminal truncations of SEQ ID NO: 82, or N-terminal truncations of a variant which is at least 90% identical, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 82.
[0106] KPI, KP2 and KP3 capsids
[0107] Generation of the AAV-KP1, AAV-KP2 and AAV-KP3 vector particles has been described in Pekrun, K., et al., 2019. JCI Insight 2019;4(22):el31610. AAV3B is the parental sequence that is most closely related to AAV-KP1, AAV-KP2 and AAV-KP3 vector particles; AAV-KP1 and AAV-KP3 demonstrate 92% sequence identity to AAV3B and AAV-KP2 demonstrates 95% identity to AAV3B. The AAV-KP1 capsid sequence contains fragments from at least 7 of the 8 parental serotypes used in the capsid shuffling library whereas AAV-KP2 and AAV-KP3 contained fragments from at least 6 parental serotypes.
[0108] AAV-KP1, AAV-KP2 and AAV-KP3 vector particles may comprise an AAV-KP1, AAV-KP2 or AAV-KP3 capsid protein or a variant thereof, respectively. For example, AAV-KP1 vector particles may comprise AAV-KP1 VP1 proteins, AAV-KP1 VP2 proteins, and AAV-KP1 VP3 proteins, or variants thereof. For example, AAV-KP2 vector particles may comprise AAV-KP2 VP1 proteins, AAV-KP2 VP2 proteins, and AAV-KP2 VP3 proteins, or variants thereof. For example, AAV-KP3 vector particles may comprise AAV-KP3 VP1 proteins, AAV-KP3 VP2 proteins, and AAV-KP3 VP3 proteins, or variants thereof. AAV-KP1, AAV-KP2, and AAV-KP3 vector particles may comprisea total number of 60 VPl, VP2, and VP3 subunits per capsid. AAV- KPI, AAV-KP2, and AAV-KP3 vector particles may comprise KPI, KP2 or KP3 VPl, VP2 and VP3 capsid proteins, or variants thereof, in a ratio of 1:1:10 (VP1:VP2:VP3).
[0109] AAV-KP1, AAV-KP2 and AAV-KP3 variant capsids have been described in the art (see e.g. Pekrun, K., et al., 2019. JCI Insight 2019;4(22):el31610 and WO 2019 / 191701A1). Further AAV-KP1, AAV-KP2 and AAV-KP3 variant capsids may be generated by amino acid substitutions, deletions, or insertions.
[0110] An example AAV-KP1 VPl capsid sequence is provided below in SEQ ID NO: 104 (GenBank accession number MN428626; Protein accession number: QFR04619). AAV-KP1 VP2 and VP3 capsid proteins may be N-terminal truncations of SEQ ID NO: 104.
[0111] An example AAV-KP2 VPl capsid sequence is provided below in SEQ ID NO: 105 (GenBank accession number IVIN428627.1; Protein accession number QFR04620.1). AAV-KP2 VP2 and VP3 capsid proteins may be N-terminal truncations of SEQ ID NO: 105.
[0112] An example AAV-KP1 VP3 capsid sequence is provided below in SEQ ID NO: 106 (GenBank accession number IVIN428628.1; Protein accession number QFR04621). AAV-KP3 VP2 and VP3 capsid proteins may be N-terminal truncations of SEQ ID NO: 106.
[0113] AAV9 serotype
[0114] The first vector particle and / or the second vector particle may comprise an AAV9 capsid protein. Suitably, the first vector particle and / or the second vector particle may be encapsidated by AAV9 capsid proteins.
[0115] In some embodiments, the first vector particle and the second vector particle each comprise an AAV9 capsid protein. In some embodiments, the first vector particle and the second vector particle are both encapsidated by AAV9 capsid proteins.
[0116] The first vector particle and / or the second vector particle may comprise an AAV9 VPl capsid protein, an AAV9 VP2 capsid protein, and / or an AAV9 VP3 capsid protein. Suitably, the firstvector particle and / or the second vector particle may be encapsidated by AAV9 VPl capsid proteins, AAV9VP2 capsid proteins, and / or AAV9VP3 capsid proteins. Suitably, the first vector particle and / or the second vector particle may be encapsidated by AAV9 VPl, VP2, and VP3 capsid proteins.
[0117] Suitably, the AAV9 VPl capsid protein may comprise or consist of the amino acid sequence shown as SEQ ID NO: 83, or a variant which is at least 90% identical to SEQ ID NO: 83.
[0118] Suitably, the variant may be at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 83.
[0119] Suitably, the AAV9 VP2 and VP3 capsid proteins may be N-terminal truncations of SEQ ID NO: 83, or N-terminal truncations of a variant which is at least 90% identical, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 83.
[0120] ShHIO serotype
[0121] The first vector particle and / or the second vector particle may comprise a ShHIO capsid protein. Suitably, the first vector particle and / or the second vector particle may be encapsidated by ShHIO capsid proteins.
[0122] In some embodiments, the first vector particle and the second vector particle each comprise a ShHIO capsid protein. In some embodiments, the first vector particle and the second vector particle are both encapsidated by ShHIO capsid proteins.
[0123] The ShHIO variant is derived from AAV6 and has increased specificity and efficiency for Muller cells (see e.g. Klimczak, R.R., et al., 2009. PloS one, 4(10), p.e7467).
[0124] The first vector particle and / or the second vector particle may comprise a ShHIO VPl capsid protein, a ShHIO VP2 capsid protein, and / or a ShHIO VP3 capsid protein. Suitably, the first vector particle and / or the second vector particle may be encapsidated by ShHIO VPl capsid proteins, ShHIO VP2 capsid proteins, and / or ShHIO VP3 capsid proteins. Suitably, the firstvector particle and / or the second vector particle may be encapsidated by ShHIO VPl, VP2, and VP3 capsid proteins.
[0125] Suitably, the ShHIO VPl capsid protein may comprise or consist of the amino acid sequence shown as SEQ ID NO: 84, or a variant which is at least 90% identical to SEQ ID NO: 84.
[0126] Suitably, the variant may be at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 84.
[0127] Suitably, the ShHIO VP2 and VP3 capsid proteins may be N-terminal truncations of SEQ ID NO: 84, or N-terminal truncations of a variant which is at least 90% identical, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 84.
[0128] AAV-DJ serotype
[0129] The first vector particle and / or the second vector particle may comprise an AAV-DJ capsid protein. Suitably, the first vector particle and / or the second vector particle may be encapsidated by AAV-DJ capsid proteins.
[0130] In some embodiments, the first vector particle and the second vector particle each comprise an AAV-DJ capsid protein. In some embodiments, the first vector particle and the second vector particle are both encapsidated by AAV-DJ capsid proteins.
[0131] AAV-DJ is a hybrid vector created from DNA shuffling of eight AAV serotypes, which mediates efficient gene expression both in vitro and in vivo (see e.g. Mao, Y., et al., 2016. BMC biotechnology, 16, pp.1-8).
[0132] The first vector particle and / or the second vector particle may comprise an AAV-DJ VPl capsid protein, an AAV-DJ VP2 capsid protein, and / or an AAV-DJ VP3 capsid protein. Suitably, the first vector particle and / or the second vector particle may be encapsidated by AAV-DJ VPl capsid proteins, AAV-DJ VP2 capsid proteins, and / or AAV-DJ VP3 capsid proteins. Suitably, the first vector particle and / or the second vector particle may be encapsidated by AAV-DJ VPl, VP2, and VP3 capsid proteins.Suitably, the AAV-DJ VPl capsid protein may comprise or consist of the amino acid sequence shown as SEQ ID NO: 85, or a variant which is at least 90% identical to SEQ ID NO: 85.
[0133] Suitably, the variant may be at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 85.
[0134] Suitably, the AAV-DJ VP2 and VP3 capsid proteins may be N-terminal truncations of SEQ ID NO: 85, or N-terminal truncations of a variant which is at least 90% identical, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 85.
[0135] AAV2 serotype
[0136] The first vector particle and / or the second vector particle may comprise an AAV2 capsid protein. Suitably, the first vector particle and / or the second vector particle may be encapsidated by AAV2 capsid proteins.
[0137] In some embodiments, the first vector particle and the second vector particle each comprise an AAV2 capsid protein. In some embodiments, the first vector particle and the second vector particle are both encapsidated by AAV2 capsid proteins.
[0138] The first vector particle and / or the second vector particle may comprise an AAV2 VPl capsid protein, an AAV2 VP2 capsid protein, and / or an AAV2 VP3 capsid protein. Suitably, the first vector particle and / or the second vector particle may be encapsidated by AAV2 VPl capsid proteins, AAV2 VP2 capsid proteins, and / or AAV2VP3 capsid proteins. Suitably, the first vector particle and / or the second vector particle may be encapsidated by AAV2 VPl, VP2, and VP3 capsid proteins.
[0139] Suitably, the AAV2 VPl capsid protein may comprise or consist of the amino acid sequence shown as SEQ ID NO: 86, or a variant which is at least 90% identical to SEQ ID NO: 86.
[0140] Suitably, the variant may be at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 86.Suitably, the AAV2 VP2 and VP3 capsid proteins may be N-terminal truncations of SEQ ID NO: 86, or N-terminal truncations of a variant which is at least 90% identical, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 86.
[0141] AAV6.2 serotype
[0142] The first vector particle and / or the second vector particle may comprise an AAV6.2 capsid protein. Suitably, the first vector particle and / or the second vector particle may be encapsidated by AAV6.2 capsid proteins.
[0143] In some embodiments, the first vector particle and the second vector particle each comprise an AAV6.2 capsid protein. In some embodiments, the first vector particle and the second vector particle are both encapsidated by AAV6.2 capsid proteins.
[0144] The AAV6.2 vector mutant was created by mutating the phenylalanine (F) residue at position 129 in AAV6 to leucine (L) (see e.g. Limberis, M.P., et al., 2009. Molecular Therapy, 17(2), pp.294-301).
[0145] The first vector particle and / or the second vector particle may comprise an AAV6.2 VPl capsid protein, an AAV6.2 VP2 capsid protein, and / or an AAV6.2 VP3 capsid protein. Suitably, the first vector particle and / or the second vector particle may be encapsidated by AAV6.2 VPl capsid proteins, AAV6.2 VP2 capsid proteins, and / or AAV6.2 VP3 capsid proteins. Suitably, the first vector particle and / or the second vector particle may be encapsidated by AAV6.2 VPl, VP2, and VP3 capsid proteins.
[0146] Suitably, the AAV6.2 VPl capsid protein may comprise or consist of the amino acid sequence shown as SEQ ID NO: 87, or a variant which is at least 90% identical to SEQ ID NO: 87.
[0147] Suitably, the variant may be at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 87.Suitably, the AAV6.2 VP2 and VP3 capsid proteins may be N-terminal truncations of SEQ ID NO: 87, or N-terminal truncations of a variant which is at least 90% identical, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 87.
[0148] AAV5 serotype
[0149] The first vector particle and / or the second vector particle may comprise an AAV5 capsid protein. Suitably, the first vector particle and / or the second vector particle may be encapsidated by AAV5 capsid proteins.
[0150] In some embodiments, the first vector particle and the second vector particle each comprise an AAV5 capsid protein. In some embodiments, the first vector particle and the second vector particle are both encapsidated by AAV5 capsid proteins.
[0151] The first vector particle and / or the second vector particle may comprise an AAV5 VPl capsid protein, an AAV5 VP2 capsid protein, and / or an AAV5 VP3 capsid protein. Suitably, the first vector particle and / or the second vector particle may be encapsidated by AAV5 VPl capsid proteins, AAV5 VP2 capsid proteins, and / or AAV5 VP3 capsid proteins. Suitably, the first vector particle and / or the second vector particle may be encapsidated by AAV5 VPl, VP2, and VP3 capsid proteins.
[0152] Suitably, the AAV5 VPl capsid protein may comprise or consist of the amino acid sequence shown as SEQ ID NO: 88, or a variant which is at least 90% identical to SEQ ID NO: 88.
[0153] Suitably, the variant may be at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 88.
[0154] Suitably, the AAV5 VP2 and VP3 capsid proteins may be N-terminal truncations of SEQ ID NO: 88, or N-terminal truncations of a variant which is at least 90% identical, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 88.
[0155] Generation of the AAV-KP1, AAV-KP2 and AAV-KP3 vector particles has been described in Pekrun, K., et al., 2019. JCI Insight 2019;4(22):el31610. AAV3B is the parental sequence thatis most closely related to AAV-KPl, AAV-KP2 and AAV-KP3 vector particles; AAV-KPl and AAV-KP3 demonstrate 92% sequence identity to AAV3B and AAV-KP2 demonstrates 95% identity to AAV3B. The AAV-KPl capsid sequence contains fragments from at least 7 of the 8 parental serotypes used in the capsid shuffling library whereas AAV-KP2 and AAV-KP3 contained fragments from at least 6 parental serotypes.
[0156] AAV-KPl, AAV-KP2 and AAV-KP3 vector particles may comprise an AAV-KPl, AAV-KP2 or AAV-KP3 capsid protein or a variant thereof, respectively. For example, AAV-KPl vector particles may comprise AAV-KPl VP1 proteins, AAV-KPl VP2 proteins, and AAV-KPl VP3 proteins, or variants thereof. For example, AAV-KP2 vector particles may comprise AAV-KP2 VP1 proteins, AAV-KP2 VP2 proteins, and AAV-KP2 VP3 proteins, or variants thereof. For example, AAV-KP3 vector particles may comprise AAV-KP3 VP1 proteins, AAV-KP3 VP2 proteins, and AAV-KP3 VP3 proteins, or variants thereof. AAV-KPl, AAV-KP2, and AAV-KP3 vector particles may comprise a total number of 60 VP1, VP2, and VP3 subunits per capsid. AAV- KPI, AAV-KP2, and AAV-KP3 vector particles may comprise KPI, KP2 or KP3 VP1, VP2 and VP3 capsid proteins, or variants thereof, in a ratio of 1:1:10 (VP1:VP2:VP3).
[0157] AAV-KPl, AAV-KP2 and AAV-KP3 variant capsids have been described in the art (see e.g. Pekrun, K., et al., 2019. JCI Insight 2019;4(22):el31610 and WO 2019 / 191701A1). Further AAV-KPl, AAV-KP2 and AAV-KP3 variant capsids may be generated by amino acid substitutions, deletions, or insertions.
[0158] An example AAV-KPl VP1 capsid sequence is provided below in SEQ ID NO: 104 (GenBank accession number MN428626; Protein accession number: QFR04619). AAV-KPl VP2 and VP3 capsid proteins may be N-terminal truncations of SEQ ID NO: 104.
[0159] An example AAV-KP2 VP1 capsid sequence is provided below in SEQ ID NO: 105 (GenBank accession number IVIN428627.1; Protein accession number QFR04620.1). AAV-KP2 VP2 and VP3 capsid proteins may be N-terminal truncations of SEQ ID NO: 105.An example AAV-KPl VP3 capsid sequence is provided below in SEQ ID NO: 106 (GenBank accession number MN428628.1; Protein accession number QFR04621). AAV-KP3 VP2 and VP3 capsid proteins may be N-terminal truncations of SEQ ID NO: 106.
[0160] Ribozymes
[0161] The first nucleic acid molecule or vector may comprise a first self-cleaving ribozyme sequence; and the second nucleic acid molecule or vector may comprise a second self-cleaving ribozyme sequence;
[0162] wherein the first self-cleaving ribozyme sequence is positioned downstream of the 5' CDS; and wherein the second self-cleaving ribozyme sequence is positioned between the 3' CDS and the promoter.
[0163] A ribozyme sequence encodes for a ribozyme. This is a catalytic agent akin to an enzyme but comprised of RNA. Whilst many ribozyme families are capable of cleaving phosphodiester bonds, self-cleaving ribozymes are specifically capable of catalytic self-cleaving, that is they can independently catalyse cleavage of the phosphodiester bond which connects them to another sequence of RNA. Examples of self-cleaving ribozymes that may be utilised in the context of the present invention include the hammerhead (HH), hepatitis-delta virus (HDV), and twister ribozyme families. Exemplary ribozymes that may be used in the context of the present invention include, but is not limited to, members of the Hammerhead (HH), Hepatitis Delta Virus (HDV), Varkud Satellite (VS), Twister, Twister-sister, Hairpin, Hatchet and Pistol families of ribozymes. Further suitable ribozymes are disclosed in Huan Peng et al., RSC Chem. Biol., 2021,2, 1370-1383.
[0164] Ribozyme mediated cleavage of RNA results in the presence of 5'-hydroxyl (5'-OH) groups and 2',3'-cyclic phosphate (2',3'-cP) groups which are not found in standard cellular RNA processing where 5'-phosphate and 3'-hydroxyl groups dominate. 5'-OH and 2',3'-cP groups are recognised by RNA 2',3'-cP and 5'-OH ligase (RtcB) which naturally acts to repair RNAs which have been damaged or cleaved by ribotoxins and is also involved in the cis-splicing pathway. Without wishing to be bound by theory, we suggest that similar endogenous slicing machinery is used to trans-ligate the mRNA comprising the 5' CDS and 3' CDS after ribozymal self-cleavage, thereby reconstituting the CDS at the mRNA level.The first self-cleaving ribozyme may be a member of the Twister family of ribozymes. The Twister ribozyme may comprise one or more nucleotides in a Pl stem overhang. The number of nucleotides in the Pl stem overhang may be 1 or more, 2 or more, 3 or more, 4 or more, or 5 or more.
[0165] The Twister ribozyme (TwistR) may comprise or consist of SEQ ID NO: 98, as shown below:
[0166] 5' nnnnnuaacacugccaaugccggucccaagcccggauaaaaguggagggnnnnn 3' (SEQ ID NO: 98)
[0167] Nucleotides designated as n correspond to nucleotides that hybridize with nucleotides of the sequence downstream of said Twister ribozyme. The nucleotides designated as n may be A, C, G, T / U.
[0168] The first self-cleaving ribozyme may comprise or consist of SEQ ID NOs: 98 to 100.
[0169] TwistR:
[0170] tgggctaacactgccaatgccggtcccaagcccggataaaagtggaggggccca (SEQ ID NO: 99)
[0171] TwistR:
[0172] nnnnntaacactgccaatgccggtcccaagcccggataaaagtggagggnnnnn (SEQ ID NO: 100)
[0173] The second self-cleaving ribozyme may be a member of the HH family of ribozymes. The HH ribozyme may comprise one or more nucleotides in a stem 1 overhang that hybridize with nucleotides of the sequence upstream or downstream of said HH ribozyme. The number of nucleotides in the Stem 1 overhang may be 1 or more nucleotides, 2 or more nucleotides, 4 or more nucleotides, 6 or more nucleotides, 8 or more nucleotides, 10 or more nucleotide, 12 or more nucleotides, 14 or more nucleotides, 16 or more nucleotides, 18 or more nucleotides, or 20 or more nucleotides.
[0174] The HH ribozyme may be .modified in stem 1 to include a tertiary stabilizing motif (TSM). The HH ribozyme can be modified in the stem 2 loop and modified in stem 1 to include a tertiarystabilizing motif (TSM). A modified HH ribozyme cis- cleaves more efficiently than HH ribozyme. The modified HH ribozyme may be RzB (as described in Saksmerprome et al, RNA.
[0175] 2004 Dec;10(12):1916-1924. doi: 10.1261 / rna.7159504).
[0176] The RzB ribozyme may comprise or consist of SEQ ID NO: 101, as shown below:
[0177] 5' nnnnnnuaannnnncugaugagucgcugggaugcgacgaaacgccuucgggcguc 3' (SEQ ID NO: 101)
[0178] Nucleotides designated as n correspond to nucleotides that hybridize with nucleotides of the sequence downstream of said HH ribozyme. The nucleotides designated as n may be A, C, G, T / U.
[0179] The second self-cleaving ribozyme may comprise or consist of SEQ ID NOs: 101 to 103(RzB).
[0180] RzB:
[0181] ct tctta a cga ca ct a tga gtcgctggga t cga cga a a cgccttcgggcgtctgtcga ga ca g (SEQ ID NO: 102)
[0182] RzB:
[0183] nnnnnntaannnnn ctgatga gtcgctggga tgcga cga a a cgccttcgggcgtcn nnnnnnnnnn (SEQ ID NO: 103)
[0184] In an embodiment the first self-cleaving ribozyme is Twister and the second self-cleaving ribozyme is RzB.
[0185] The first self-cleaving ribozyme may be a member of the Twister family of ribozymes. The first self-cleaving ribozyme may comprise or consist of any one of SEQ ID NO. 98 to 100.
[0186] The second self-cleaving ribozyme may be a member of the Twister family of ribozymes. The second self-cleaving ribozyme may comprise or consist of any one of SEQ ID NO. 101 to 103.
[0187] In an embodiment the first self-cleaving ribozyme is Twister and the second self-cleaving ribozyme is Twister.Splice donor and acceptor sites
[0188] The first vector may comprise a splice donor sequence between the 5' CDS and the first selfcleaving ribozyme sequence; and the second vector may comprise a splice acceptor sequence between the second self-cleaving ribozyme and the 3' CDS.
[0189] Splice donor and acceptor sequences
[0190] Splice donor sequence
[0191] In some embodiments (e.g. in trans-splicing methods), the 5' AAV vector of the present invention comprises a splice donor sequence. Suitably, the splice donor sequence is downstream of the 5' CDS. Suitable, the splice donor sequence may be immediately downstream of the 5' CDA, i.e. without any dap between the 5' CDS and splice donor sequence.
[0192] Any suitable splice donor sequence may be used. Suitably, the splice donor sequence comprises a splice donor site at its 5' end. A splice donor site may include an almost invariant sequence GT at the 5' end. Suitably, the splice donor sequence comprises or consists of the nucleotide sequence GTNNNN(N)X, wherein x is from about 10 to about 1000 (e.g. from about 10 to about 500, from about 10 to about 200, or from about 10 to about 100).
[0193] Suitably, a consensus splice donor site comprises or consists of the nucleotide sequence GTRAGT (see e.g. Gao, K., et al., 2008. Nucleic acids research, 36(7), pp.2257-2267). Suitably, the splice donor sequence comprises or consists of the nucleotide sequence GTRAGT(N)X, wherein x is from about 10 to about 1000 (e.g. from about 10 to about 500, from about 10 to about 200, or from about 10 to about 100).
[0194] The splice donor sequence may be a synthetic or exogenous splice donor sequence. An example synthetic or exogenous splice donor sequence is provided in SEQ ID NO: 76.
[0195] In some embodiments, the splice donor sequence comprises or consists of a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least98%, at least 99%, or 100% identity to SEQ ID NO: 76. In some embodiments, the splice donor sequence comprises or consists of the nucleotide sequence of SEQ ID NO: 76.
[0196] The splice donor sequence may be an endogenous splice donor sequence, for example from a COL4A3, COL4A4, or COL4A5 gene.
[0197] In some embodiments, the endogenous splice donor sequence corresponds to a 5' fragment intronic region between the exons present in the 5' CDS and the 3' CDS. For example, if the 5' CDS comprises exons 1-32 and the 3' CDS comprises exons 33-51 of a CDS encoding a COL4A5 polypeptide, then the splice donor sequence may comprise a 5' fragment of intron 32 of the CDS encoding a COL4A5 polypeptide. An example endogenous splice donor sequence is provided in SEQ ID NO: 77.
[0198] In some embodiments, the splice donor sequence comprises or consists of a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to SEQ ID NO: 77. In some embodiments, the splice donor sequence comprises or consists of the nucleotide sequence of SEQ ID NO: 77.
[0199] In some embodiments (e.g. in mRNA trans-splicing methods), the 5' AAV vector of the present invention comprises a forward splice donor sequence. Suitably, the forward splice donor sequence is downstream of the 5' CDS. Suitably, the forward splice donor sequence may be immediately downstream of the 5' CDS, i.e. without any gap between the 5' CDS and forward splice donor sequence.
[0200] As used herein, a "forward" splice donor sequence may refer to a splice donor sequence which functions as a splice donor site when transcribed from the sense strand of an AAV vector. As used herein, a "reverse" splice donor sequence may refer to a splice donor sequence which functions as a splice donor site when transcribed from the anti-sense strand of an AAV vector. Any suitable forward splice donor sequence may be used, for example SEQ ID NO: 76 or a variant thereof, or SEQ ID NO: 77 or a variant thereof.Splice acceptor sequences
[0201] In some embodiments (e.g. in trans-splicing methods), the 3' AAV vector of the present invention comprises a splice acceptor sequence. Suitably, the splice acceptor sequence is upstream of the 3' CDS. Suitably, the splice acceptor sequence may be immediately upstream of the 3' CDS, i.e. without any gap between the 3' CDS and splice acceptor sequence.
[0202] Any suitable splice acceptor sequence may be used. Suitably, the splice acceptor sequence comprises a splice acceptor site at its 3' end. A splice acceptor site may include an almost invariant sequence AG at the 3' end. Suitably, the splice acceptor sequence comprises or consists of the nucleotide sequence (N)XNAG, wherein x is from about 10 to about 1000 (e.g. from about 10 to about 500, from about 10 to about 200, or from about 10 to about 100).
[0203] Suitably, a consensus splice acceptor site comprises or consists of the nucleotide sequence YAG. Suitably, the splice acceptor sequence comprises or consists of the nucleotide sequence (N)xYAG, wherein x is from about 10 to about 1000 (e.g. from about 10 to about 500, from about 10 to about 200, or from about 10 to about 100).
[0204] The splice acceptor sequence may be a synthetic or exogenous splice donor sequence. An example synthetic or exogenous splice acceptor sequence is provided in SEQ ID NO: 78.
[0205] In some embodiments, the splice acceptor sequence comprises or consists of a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to SEQ ID NO: 78. In some embodiments, the splice acceptor sequence comprises or consists of the nucleotide sequence of SEQ ID NO: 78.
[0206] The splice acceptor sequence may be an endogenous splice acceptor sequence, for example from a COL4A3, COL4A4, or COL4A5 gene.
[0207] In some embodiments, the endogenous splice acceptor sequence corresponds to a 3' fragment intronic region between the exons present in the 5' CDS and the 3' CDS. For example, if the 5' CDS comprises exons 1-32 and the 3' CDS comprises exons 33-51 of a CDS encoding aCOL4A5 polypeptide, then the splice acceptor sequence may comprise a 3' fragment of intron 32 of the CDS encoding a COL4A5 polypeptide. An example endogenous splice acceptor sequence is provided in SEQ ID NO: 79.
[0208] In some embodiments, the splice acceptor sequence comprises or consists of a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to SEQ ID NO: 79. In some embodiments, the splice acceptor sequence comprises or consists of the nucleotide sequence of SEQ ID NO: 79.
[0209] In some embodiments (e.g. in mRNA trans-splicing methods), the 3' AAV vector of the present invention comprises a reverse splice acceptor sequence. Suitably, the reverse splice acceptor sequence is upstream of the 3' CDS. Suitably, the reverse splice acceptor sequence may be immediately upstream of the 3' CDS, i.e. without any gap between the 3' CDS and reverse splice acceptor sequence.
[0210] As used herein, a "forward" splice acceptor sequence may refer to a splice acceptor sequence which functions as a splice acceptor site when transcribed from the sense strand of an AAV vector. As used herein, a "reverse" splice acceptor sequence may refer to a splice acceptor sequence which functions as a splice acceptor site when transcribed from the anti-sense strand of an AAV vector.
[0211] Any suitable reverse splice acceptor sequence may be used. An example reverse synthetic or exogenous splice acceptor sequence is provided in SEQ ID NO: 89, which is based on SEQ ID NO: 78.
[0212] In some embodiments, the reverse splice acceptor sequence comprises or consists of a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to SEQ ID NO: 89. In some embodiments, the reverse splice acceptor sequence comprises or consists of the nucleotide sequence of SEQ ID NO: 89.An example reverse endogenous splice acceptor sequence is provided in SEQ ID NO: 90, which is based on SEQ ID NO: 79.
[0213] In some embodiments, the reverse splice acceptor sequence comprises or consists of a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to SEQ ID NO: 89. In some embodiments, the reverse splice acceptor sequence comprises or consists of the nucleotide sequence of SEQ ID NO: 89.
[0214] Example combinations of splice donor and acceptor sequences
[0215] In some embodiments, the 5' AAV vector comprises a splice donor sequence downstream from the 5' CDS having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to SEQ ID NO: 76; and the 3' AAV vector comprises a splice acceptor sequence upstream from the 3' CDS having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to SEQ ID NO: 78. In some embodiments, the 5' AAV vector comprises a splice donor sequence downstream from the 5' CDS having at least 70% identity to SEQ ID NO: 76 and the 3' AAV vector comprises a splice acceptor sequence upstream from the 3' CDS having at least 70% identity to SEQ ID NO: 78.
[0216] In some embodiments, the 5' AAV vector comprises a splice donor sequence downstream from the 5' CDS having at least 80% identity to SEQ ID NO: 76 and the 3' AAV vector comprises a splice acceptor sequence upstream from the 3' CDS having at least 80% identity to SEQ ID NO: 78.
[0217] In some embodiments, the 5' AAV vector comprises a splice donor sequence downstream from the 5' CDS having at least 85% identity to SEQ ID NO: 76 and the 3' AAV vector comprisesa splice acceptor sequence upstream from the 3' CDS having at least 85% identity to SEQ ID NO: 78.
[0218] The inclusion of a splice donor sequence between the 5' CDS and the first self-cleaving ribozyme sequence of the first vector; and a splice acceptor sequence between the second self-cleaving ribozyme and the 3' CDS of the second vector enhances the likelihood of reconstitution of the mRNA.
[0219] Cryptic splice donor sites
[0220] It is evidently advantageous for a system which is formed via reconstitution of two nucleic acid molecules at the mRNA level that the system is spliced together only at the intended point such that the final sequence of the polypeptide is full length.
[0221] In some embodiments, the invention therefore also provides for at least one mutation of the 5' CDS to remove a cryptic splice donor site from the coding sequence. Thus the 5' CDS may be a mutant 5' CDS that is mutated version of the wild-type nucleotide sequence that encodes an N-terminal part of the COL4A3, COL4A4 or COL4A5 and comprises one or more mutations that remove one or more cryptic splice donor sites from the wild-type sequence. In some embodiments, the one or more mutations are silent mutations. In any embodiments where the 5'CDS encodes a COL4A5 polypeptide the 5'CDS nucleotide sequence may have at least 80% and less than 100% sequence identity to SEQ ID NO: 122. In any embodiments where the 5'CDS encodes a COL4A5 polypeptide the 5'CDS nucleotide sequence may comprise or consist of SEQ ID NO:123 or SEQ ID NO:124.
[0222] Sequences which can act as splice donor sites can vary considerably and, whilst they almost invariably comprise a GT dinucleotide, the surrounding nucleotide sequences can be quite variable. Splice sites are naturally found in introns and are targets for the splicing machinery which produces mature mRNA from immature mRNA. Cryptic splice donor sites are splice donor sites which are present within exons and which share features with regular splice donor sites but are not normally recognised by splicing machinery during native construction of mature mRNA, but which may be activated under certain conditions see: Roca X, Sachidanandam R, Krainer AR. Intrinsic differences between authentic and cryptic 5' splicesites. Nucleic Acids Res. 2003 Nov l;31(21):6321-33. doi: 10.1093 / nar / gkg830. PMID: 14576320; PMCID: PMC275472.. Cryptic splice donor sites may therefore lead to aberrant splicing patterns in the reconstitution of coding sequences, leading to the production of encoded protein(s) of abnormal length. Cryptic splice donor sites may be determined algorithmically using software such as Spliceator (https: / / www.lbgi.fr / spliceator / , https: / / doi.org / 10.1186 / sl2859-021-04471-3) although use of this tool should not be seen to limit the invention and other algorithms are available. Study of the rules governing cryptic splice donor sites is also ongoing and the invention should not be limited to the present understanding of these rules.
[0223] A mutation may comprise replacement of a single nucleotide with a different nucleotide. Alternatively, a mutation may comprise replacement of more than one nucleotide with other nucleotides in a localised area of the sequence which includes the cryptic splice donor site. A localised area of the sequence may cover, for example, 2-20 nucleotides, 5-15 nucleotides, or 7-10 nucleotides. To preserve the coding sequence, the mutation should not result in frameshifting. Insertion or deletion of nucleotides should only be considered where the reading frame of the 5' CDS is not impacted, i.e. where nucleotides are added or deleted in multiples of three. Thus, typically, the at least one mutation is a nucleotide replacement.
[0224] In some embodiments, the at least one mutation is silent. A silent mutation is a mutation of the coding sequence which does not result in a change to the amino acid sequence of the encoded protein. This is possible due to redundancy in the genetic code such that different codons in the coding sequence encode for the same amino acid. The third nucleotide in each codon has the most redundancy, therefore in some embodiments the mutation affects the third nucleotide in a codon. Because a silent mutation does not alter the amino acid sequence of the encoded protein, the expressed protein will be indistinguishable from wild-type protein which is advantageous for treatment of diseases.
[0225] In some embodiments of the system for generating a coding sequence of a COL4A5 polypeptide, the first nucleic acid molecule comprises exons 1-9 of COL4A5 and exon 8 and / or exon 9 comprises at least one mutation of a cryptic splice donor site in the coding sequence of COL4A5. Optionally, the at least one mutation is silent. Based on observed splicing patternsduring reconstitution of the coding sequence of COL4A5 by the system of the present invention, the region around exons 8 and 9 appears to be targeted by splicing machinery causing aberrant splicing and resulting in expression of different protein isoforms. Mutation of the cryptic splice donor sites identified in these exons results in increased expression of full length COL4A5 protein.
[0226] In some embodiments of the system for generating a coding sequence of a COL4A5 polypeptide, the first nucleic acid molecule comprises exons 1-31 of COL4A5, and at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or all ten of exons 6, 12-13, 15-16, 20, 21, 22, 24, 26, 29-30 and 31 comprise at least one mutation of a cryptic splice donor site in the coding sequence of COL4A5. Optionally, the at least one mutation is silent. Exons 6, 8, 9, 12-13, 15-16, 20, 21, 22, 24, 26, 29-30 and 31 were identified as comprising cryptic splice donor sites. Mutation of these cryptic splice donor sites results in increased expression of full length COL4A5 protein.
[0227] In some embodiments of the system for generating a coding sequence of a COL4A5 polypeptide, the cryptic splice donor site in the coding sequence of COL4A5, for example a wild-type sequence of COL4A5 such as provided in SEQ ID NO. 122, comprises an AGGT sequence and each mutation made comprises replacement of at least the thymine base. Mutation of the thymine in an AGGT sequence removes the GT dinucleotide which is almost invariable in splice donor sites. In some embodiments each mutation further comprises replacement of the adenine in the sequence AGGT of the coding sequence of COL4A5. The adenine and thymine are three nucleotides apart, meaning they may both fall as the third nucleotide in adjacent codons. As described above the third nucleotide in a codon is the most degenerate and thus silent mutations can be most easily achieved.
[0228] In some embodiments of the system for generating a coding sequence of a COL4A5 polypeptide, each mutation comprises replacement of at least 2, at least 3 or all 4 of the bases in the sequence AGGT of the coding sequence.
[0229] In some embodiments of the system for generating a coding sequence of a COL4A5 polypeptide, the AGGT sequences are present at:i) position 36-39 of exon 6;
[0230] ii) position 18-21 of exon 8;
[0231] iii) position 9-12 of exon 9;
[0232] iv) position 42 of exon 12 to position 3 of exon 13;
[0233] v) position 21-24 of exon 15;
[0234] vi) position 57 of exon 15 to position 3 of exon 16;
[0235] vii) position 59-62 of exon 20;
[0236] viii) position 47-50 of exon 21;
[0237] ix) position 8-11 of exon 22;
[0238] x) position 45-48 of exon 24;
[0239] xi) position 17-20 of exon 26;
[0240] xii) position 150 of exon 29 to position 2 of exon 30; and / or
[0241] xiii) position 149-152 of exon 31
[0242] of the coding sequence of COL4A5.
[0243] ATTG sequences at these positions have been identified as cryptic splice donor sites and mutation of these sites has been demonstrated to increase expression of full length COL4A5 protein.
[0244] In some embodiments of the system for generating a coding sequence of a COL4A5 polypeptide:
[0245] i) the sequence of exon 6 consists of the sequence GGAATGCCAGGCCACGATGGGGCCCCAGGACCTCAGGGAATTCCCGGATGCAATGGAACCA AG (SEQ ID NO: 107);
[0246] ii) the sequence of exon 8 consists of the sequence GGACCCCCTGGGATCCCCGGCATGAAG (SEQ ID NO: 108);
[0247] iii) the sequence of exon 9 consists of the sequence
[0248] G GTG A ACCTG G C AG C ATA ATTATGTC ATC ACTG CC AG G ACC A A AG G GTA ATCC AG G ATATCC A GGTCCTCCTGGAATACAA (SEQ ID NO: 109);
[0249] iv) the sequence of exon 12 consists of the sequence GGGAATATGGGCTTAAATTTCCAGGGACCCAAAGGTGAAAAG (SEQ ID NO: 110);
[0250] v) the sequence of exon 13 consists of the sequenceGGAGAGCAAGGTCTTCAGGGCCCACCTGGGCCACCTGGGCAGATCAGTGAACAGAAAAGA CCAATTGATGTAGAGTTTCAGAAAGGAGATCAG (SEQ ID NO: 111);
[0251] vi) the sequence of exon 15 consists of the sequence GGTCCCCCAGGTGGTGAGAAGGGAGAGAAGGGTGAGCAAGGAGAGCCAGGCAAACGG
[0252] (SEQ ID NO: 112);
[0253] vii) the sequence of exon 16 consists of the sequence GGAAAACCAGGCAAAGATGGAGAAAATGGCCAACCAGGAATTCCT (SEQ ID NO: 113); viii) the sequence of exon 20 consists of the sequence GGGCTGCAGTTATGGGTCCTCCTGGCCCTCCTGGATTTCCTGGAGAAAGGGGTCAGAAGGG AGATGAAGGACCACCTGGAATTTCCATTCCTGGACCTCCTGGACTTGACGGACAGCCTGGGGCTCCTG GGCTTCCAGGGCCTCCTGGCCCTGCTGGCCCTCACATTCCTCCTA (SEQ ID NO: 114);
[0254] ix) the sequence of exon 21 consists of the sequence GTGATGAGATATGTGAACCAGGCCCTCCAGGCCCCCCAGGATCTCCTGGAGATAAAGGACTC CAAGGAGAACAAGGAGTGAAAG (SEQ ID NO: 115);
[0255] x) the sequence of exon 22 consists of the sequence GTGACAAGGGAGACACTTGCTTCAACTGCATTGGAACTGGTATTTCAGGGCCTCCAGGTCAA CCTGGTTTGCCAGGTCTCCCAGGTCCTCCAG (SEQ ID NO: 116);
[0256] xi) the sequence of exon 24 consists of the sequence GGCATTCCAGGAGCTCCAGGTGCTCCAGGCTTTCCTGGATCTAAGGGAGAACCTGGTGATAT CCTCACTTTTCCAGGAATGAAGGGTGACAAAGGAGAGTTGGGTTCCCCTGGAGCTCCAGGGCTTCCT GGTTTACCTGGCACTCCTGGACAGGATGGATTGCCAGGGCTTCCTGGCCCGAAAGGAGAGCCT (SEQ ID NO: 117);
[0257] xii) the sequence of exon 26 consists of the sequence GTCCTAAAGGGGATCCTGGACAGACTATAACCCAGCCGGGGAAGCCTGGCTTGCCTGGTAAC CCAGGCAGAGATGGTGACGTAGGTCTTCCAG (SEQ ID NO: 118);
[0258] xiii) the sequence of exon 29 consists of the sequence GGTGAACCAGGATTTGCATTACCTGGGCCACCTGGGCCACCAGGACTTCCAGGTTTCAAAGG AGCACTTGGTCCAAAAGGTGATCGTGGTTTCCCAGGACCTCCGGGTCCTCCAGGACGCACTGGCTTA GATGGGCTCCCTGGACCAAAGG (SEQ ID NO: 119);
[0259] xiv) the sequence of exon 30 consists of the sequenceGAGATGTTGGACCAAATGGACAACCTGGACCAATGGGACCTCCTGGGCTGCCAGGAATAGG TGTTCAGGGACCACCAGGACCACCAGGGATTCCTGGGCCAATAGGTCAACCTG (SEQ ID NO: 120); and / or
[0260] xv) the sequence of exon 31 consists of the sequence GTTTACATGGAATACCAGGAGAGAAGGGGGATCCAGGACCTCCTGGACTTGATGTTCCAGGA CCCCCAGGTGAAAGAGGCAGTCCAGGGATCCCCGGAGCACCTGGTCCTATAGGACCTCCAGGATCAC CAGGGCTTCCAGGAAAAGCTGGAGCCTCTGGATTTCCAG (SEQ ID NO: 121).
[0261] These specific mutations have been found to increase expression of full length COL4A5
[0262] Cryptic splice acceptor sites
[0263] In some embodiments, the invention also provides for at least one mutation of the 3' CDS to remove a cryptic splice acceptor site from the coding sequence. Thus the 3' CDS may be a mutant 3' CDS that is a mutated version of the wild-type nucleotide sequence that encodes a C-terminal part of the COL4A3, COL4A4 or COL4A5 and comprises one or more mutations that remove one or more cryptic splice acceptor sites from the wild-type sequence.
[0264] Sequences which can act as splice acceptor sites can vary considerably and, whilst they commonly comprise an AG dinucleotide, the surrounding nucleotide sequences can be quite variable. Cryptic splice acceptor sites are splice acceptor sites which are present within exons and which share features with regular splice acceptor sites but are not normally recognised by splicing machinery during native construction of mature mRNA, but which may be activated under certain conditions see: Roca X, Sachidanandam R, Krainer AR. Intrinsic differences between authentic and cryptic 5' splice sites. Nucleic Acids Res. 2003 Nov l;31(21):6321-33. doi: 10.1093 / nar / gkg830. PMID: 14576320; PMCID: PMC275472.. Cryptic splice acceptor sites may therefore lead to aberrant splicing patterns in the reconstitution of coding sequences, leading to the production of encoded protein(s) of abnormal length. Cryptic splice acceptor sites may be determined algorithmically using software such as Spliceator (https: / / www.lbgi.fr / spliceator / , https: / / doi.org / 10.1186 / sl2859-021-04471-3) although use of this tool should not be seen to limit the invention and other algorithms are available. Study of the rules governing cryptic splice acceptor sites is also ongoing and the invention should not be limited to the present understanding of these rules.A mutation may comprise replacement of a single nucleotide with a different nucleotide. Alternatively, a mutation may comprise replacement of more than one nucleotide with other nucleotides in a localised area of the sequence which includes the cryptic splice donor site. A localised area of the sequence may cover, for example, 2-20 nucleotides, 5-15 nucleotides, or 7-10 nucleotides. To preserve the coding sequence, the mutation should not result in frameshifting. Insertion or deletion of nucleotides should only be considered where the reading frame of the 3' CDS is not impacted, i.e. where nucleotides are added or deleted in multiples of three. Thus, typically, the at least one mutation is a nucleotide replacement.
[0265] In some embodiments, the at least one mutation is silent. A silent mutation is a mutation of the coding sequence which does not result in a change to the amino acid sequence of the encoded protein. This is possible due to redundancy in the genetic code such that different codons in the coding sequence encode for the same amino acid. The third nucleotide in each codon has the most redundancy, therefore in some embodiments the mutation affects the third nucleotide in a codon. Because a silent mutation does not alter the amino acid sequence of the encoded protein, the expressed protein will be indistinguishable from wild-type protein which is advantageous for treatment of diseases.
[0266] Regulatory sequences
[0267] The second vector may further comprise one or more regulatory sequences downstream from the 3' CDS. Such regulatory sequences may comprise a Woodchuck hepatitis post-transcriptional regulatory element (WPRE). The WPRE may comprise or consist of a nucleotide sequence having 80% or more, 85% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% identity to SEQ ID NO: 80.
[0268] The second vector may further comprise a polyadenylation sequence downstream from the 3' CDS. The polyadenylation sequence may be a bovine growth hormone polyadenylation sequence (bGH), a soluble neuropilin-1 polyadenylation sequence, an early SV40 polyadenylation sequence (SV40pA), or a chicken beta-globin polyadenylation sequence; and / or wherein the polyadenylation sequence downstream from the 3' CDS comprises orconsists of a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to any of SEQ ID NOs: 72 to 75.
[0269] 5' and 3' CDS distribution
[0270] The 5' CDS may comprise or consist of exons 1 to x of a CDS encoding a COL4A5 polypeptide and the 3' CDS may comprise or consist of exons (x+1) to 51 of a CDS encoding a COL4A5 polypeptide, wherein x is from 27 to 33 (e.g. 27, 28, 29, 30, 31, 32 or 33). Splitting the CDS of COL4A5 into these portions has been found to provide good levels of protein expression, indicating an effective vector system.
[0271] In an embodiment the 5' CDS comprises or consists of exons 1 to 31 of a CDS encoding a COL4A5 polypeptide and the 3' CDS comprises or consists of exons 32 to 51 of a CDS encoding a COL4A5 polypeptide. Spliting the CDS of COL4A5 between exons 31 and 32 has proved to be particularly effective for protein expression.
[0272] Embodiments
[0273] In an embodiment the first AAV vector comprises in the 5'-3' direction:
[0274] a 5'-inverted terminal repeat (5'-ITR);
[0275] a promoter, optionally a kidney-specific promoter, optionally a podocyte-specific promoter;
[0276] the 5' CDS as defined as above, said 5' CDS being operably linked to and under control of the promoter, optionally the kidney-specific promoter, optionally the podocytespecific promoter;
[0277] optionally, a splice donor sequence;
[0278] a first self-cleaving ribozyme sequence, preferably a Twister ribozyme sequence; optionally a poly-adenylation sequence; and
[0279] a 3'-inverted terminal repeat (3'-ITR);
[0280] and the second AAV vector comprises in the 5' to 3' direction:
[0281] a 5'-ITR;
[0282] a promoter, optionally a kidney-specific promoter, optionally a podocyte-specific promoter;a second self-cleaving ribozyme sequence, preferably a HH ribozyme sequence, preferably a RzB sequence;
[0283] optionally a splice acceptor sequence;
[0284] a 3' CDS as defined above;
[0285] a WPRE;
[0286] a poly-adenylation sequence; and
[0287] a 3'-ITR, optionally wherein the first AAV vector comprises one or more regulatory elements.
[0288] In an embodiment the first AAV vector comprises in the 5'-3' direction:
[0289] a 5'-inverted terminal repeat (5'-ITR);
[0290] a promoter, optionally a kidney-specific promoter, optionally a podocyte-specific promoter;
[0291] the 5' CDS as defined as above, said 5' CDS being operably linked to and under control of the promoter, optionally the kidney-specific promoter, optionally the podocytespecific promoter;
[0292] optionally, a splice donor sequence;
[0293] a first self-cleaving ribozyme sequence, preferably a Twister ribozyme sequence; optionally a poly-adenylation sequence; and
[0294] a 3'-inverted terminal repeat (3'-ITR);
[0295] and the second AAV vector comprises in the 5' to 3' direction:
[0296] a 5'-ITR;
[0297] a promoter, optionally a kidney-specific promoter, optionally a podocyte-specific promoter;
[0298] a second self-cleaving ribozyme sequence, preferably a Twister ribozyme sequence; optionally a splice acceptor sequence;
[0299] a 3' CDS as defined above;
[0300] a WPRE;
[0301] a poly-adenylation sequence; and
[0302] a 3'-ITR, optionally wherein the first AAV vector comprises one or more regulatory elements.When the first and second vectors are AAV vectors, the first and / or second vector may be in the form of an AAV vector particle encapsidated by LK03, AAV3B, AAV9, ShHIO, AAV-DJ, AAV2, AAV6.2, KPI, KP2, KP3 orAAV5. Particularly the first AAV vector and / or the second AAV vector may be in the form of an AAV vector particle encapsidated by the LK03 capsid protein.
[0303] Cells
[0304] In a second aspect is provided an isolated cell comprising the system according to the first aspect.
[0305] The cell may be any cell type known in the prior art. The cell may be an isolated cell. The cell may be a human cell, suitably an isolated human cell.
[0306] Suitably, the cell may be a kidney cell or glomerular cell, for example a podocyte. Suitably, the cell may be an immortalized kidney cell or glomerular cell, for example an immortalized podocyte. Suitable podocyte cell lines will be well known to those of skill in the art, for example CIHP-1. Methods to generate immortalized podocytes will be well known to those of skill in the art. Suitable methods are described in Ni, L., et al., 2012. Nephrology, 17(6), pp.525-531.
[0307] Suitably, the cell may be a producer cell. The term "producer cell" includes a cell that produces viral particles, after transient transfection, stable transfection or vector transduction of all the elements necessary to produce the viral particles or any cell engineered to stably comprise the elements necessary to produce the viral particles. In some embodiments, the producer cell is an AAV producer cell. Suitable producer cells will be known to those of skill in the art (see e.g. Martin, J., et al. 2013. Human gene therapy methods, 24(4), pp.253-269) and may include HEK293, COS-1, COS-7, CV-1, HeLa, CHO, and A549 cell lines. In some embodiments, the producer cell is a HEK293 cell, or a derivative thereof (e.g. a HEK293T cell).
[0308] Suitably, the cell may be a packaging cell. The term "packaging cell" includes a cell which contains some or all of the elements necessary for packaging a recombinant virus genome. Typically, such packaging cells contain one or more vectors which are capable of expressing viral structural proteins (e.g. AAV rep and cap genes) and / or one or more genes encoding theviral structural proteins have been integrated into the genome of the packaging cell. Cells comprising only some of the elements required for the production of enveloped viral particles are useful as intermediate reagents in the generation of viral particle producer cell lines, through subsequent steps of transient transfection, transduction or stable integration of each additional required element. These intermediate reagents are encompassed by the term "packaging cell". In some embodiments, the packaging cell is an AAV packaging cell. Suitable packaging cells will be known to those of skill in the art (see e.g. Martin, J., et al. 2013. Human gene therapy methods, 24(4), pp.253-269).
[0309] Pharmaceutical Compositions
[0310] In a third aspect is provided a pharmaceutical composition comprising the system according to the first aspect or a cell according to the second aspect.
[0311] The system may be a dual vector system. The system may be a dual AAV vector system.
[0312] A pharmaceutical composition is a composition that comprises or consists of a therapeutically effective amount of a pharmaceutically active agent i.e. the dual AAV vector system. It preferably includes a pharmaceutically acceptable carrier, diluent or excipient (including combinations thereof).
[0313] By "pharmaceutically acceptable" is included that the formulation is sterile and pyrogen free. The carrier, diluent, and / or excipient must be "acceptable" in the sense of being compatible with the vector and not deleterious to the recipients thereof. Typically, the carriers, diluents, and excipients will be saline or infusion media which will be sterile and pyrogen free; however, other acceptable carriers, diluents, and excipients may be used.
[0314] Acceptable carriers, diluents, and excipients for therapeutic use are well known in the pharmaceutical art. The choice of pharmaceutical carrier, excipient or diluent can be selected with regard to the intended route of administration and standard pharmaceutical practice. The pharmaceutical compositions may comprise as - or in addition to - the carrier, excipient or diluent any suitable binder(s), lubricant(s), suspending agent(s), coating agent(s) or solubilising agent(s).Examples of pharmaceutically acceptable carriers include, for example, water, salt solutions, alcohol, silicone, waxes, petroleum jelly, vegetable oils, polyethylene glycols, propylene glycol, liposomes, sugars, gelatin, lactose, amylose, magnesium stearate, talc, surfactants, silicic acid, viscous paraffin, perfume oil, fatty acid monoglycerides and diglycerides, petroethral fatty acid esters, hydroxymethyl-cellulose, polyvinylpyrrolidone, and the like.
[0315] The system, cell, or pharmaceutical composition according to the present invention may be administered in a manner appropriate for treating and / or preventing the diseases described herein. The quantity and frequency of administration will be determined by such factors as the condition of the subject, and the type and severity of the subject's disease, although appropriate dosages may be determined by clinical trials. The pharmaceutical composition may be formulated accordingly.
[0316] The system, cell or pharmaceutical composition according to the present invention may be administered parenterally, for example, intravenously, or by infusion techniques. The system, cell or pharmaceutical composition may be administered in the form of a sterile aqueous solution which may contain other substances, for example, enough salts or glucose to make the solution isotonic with blood. The aqueous solution may be suitably buffered (preferably to a pH of from 3 to 9). The pharmaceutical composition may be formulated accordingly. The preparation of suitable parenteral formulations under sterile conditions is readily accomplished by standard pharmaceutical techniques well-known to those skilled in the art.
[0317] The system, cell or pharmaceutical composition according to the present invention may be administered systemically, for example by intravenous injection.
[0318] The system, cell or pharmaceutical composition according to the present invention may be administered locally, for example by targeting administration to the kidney. Suitably, the system, cell or pharmaceutical composition may be administered by injection into the renal artery, by intraparenchymal injection, by transparenchymal injection, by renal vein injection, or by ureteral or subcapsular injection. In some embodiments, the system, cell or pharmaceutical composition is administered by injection into the renal artery.The pharmaceutical compositions may comprise the system, or cell of the invention in infusion media, for example sterile isotonic solution. The pharmaceutical composition may be enclosed in ampoules, disposable syringes or multiple dose vials made of glass or plastic.
[0319] The system, cell or pharmaceutical composition may be administered in a single or in multiple doses. Particularly, the system, cell or pharmaceutical composition may be administered in a single, one-off dose. The pharmaceutical composition may be formulated accordingly.
[0320] The system, cell or pharmaceutical composition may be administered at varying doses (e.g. measured in vector genomes (vg) per kg). The physician in any event will determine the actual dosage which will be most suitable for any individual subject and it will vary with the age, weight and response of the particular subject. Typically, however, for AAV vectors, doses of IO10to 1014vg / kg, or 1011to 1013vg / kg may be administered. Suitably, the first and second vectors are administered at the same dose.
[0321] The system may be administered in any suitable dosage (e.g. measured in vector genomes (vg) or vg per kg). The dosage may be determined by such factors as the condition of the subject, age of the subject, weight of the subject, and the type and severity of the subject's disease, and appropriate dosages may be determined by a physician. The system may be formulated accordingly.
[0322] Suitably, each vector comprised in the dual vector system is administered in a dose of about lxlO6vg / kg or more, about lxlO7vg / kg or more, about lxlO8vg / kg or more, about lxlO9vg / kg or more, about lxlO10vg / kg or more, about lxlO11vg / kg or more, or about lxlO12vg / kg or more. Suitably, each vector comprised in the dual vector system is administered in a dose of about lxlO14vg / kg or less or about lxlO13vg / kg or less. Suitably, each vector comprised in the dual vector system is administered in a dose of from about lxlO6vg / kg to about lxlO14vg / kg. Suitably, each vector comprised in the dual vector system is administered in a dose of from about lxlO6vg / kg to about lxlO13vg / kg. In some embodiments, each vector comprised in the dual vector system is administered in a dose of from about lxlO9vg / kg to about lxlO12vg / kg.Suitably, each vector is administered in a dose of about lxlO8vg or more, about lxlO9vg or more, about lxlO10vg or more, about lxlO11vg or more, about lxlO12vg or more, about lxlO13vg or more, or about lxlO14vg or more. Suitably, each vector is administered in a dose of about lxlO15vg or less or about lxlO14vg or less. Suitably, each vector is administered in a dose of from about lxl08vg to about lxlO15vg. Suitably, each vector is administered in a dose of from about lxlO8vg to about 5xl014vg. In some embodiments, each vector is administered in a dose of from about lxlO11vg to about lxlO14vg.
[0323] The system may be administered in the form of a vector formulation. As used herein, a "vector formulation" may refer to a composition that comprises or consists of a therapeutically effective amount of one or more vector. It preferably includes a pharmaceutically acceptable carrier, diluent or excipient (including combinations thereof).
[0324] The system may be administered in the form of a sterile aqueous solution which may contain other substances, for example, enough salts or glucose to make the solution isotonic with blood. The aqueous solution may be suitably buffered (preferably to a pH of from 3 to 9). The preparation of suitable parenteral formulations under sterile conditions is readily accomplished by standard pharmaceutical techniques well-known to those skilled in the art. Suitably, the vector formulation comprises an isotonic buffer (e.g. at about pH 7.4).
[0325] The practice of the present invention will employ, unless otherwise indicated, conventional techniques of chemistry, biochemistry, molecular biology, microbiology and immunology, which are within the capabilities of a person of ordinary skill in the art. Such techniques are explained in the literature. See, for example: Skoog, D.A., et al. (2013) Fundamentals of Analytical Chemistry, 9th edition, Cengage learning; Walker J. M. (2009) The Protein Protocols Handbook, 3rd edition, Springer Nature; Green, M.R. and Sambrook, J. (2012) Molecular Cloning: A Laboratory Manual, 4th Edition, Cold Spring Harbor Laboratory Press; Ausubel, F.M., et al. (2003) Current Protocols in Molecular Biology, John Wiley & Sons; Hill, A. J. (2013) DNA Sequencing Protocols, Humana Press; Nielsen, B.S. and Jones, J. (2021) In Situ Hybridization Protocols, Springer US; Herdewijn, P. (2010) Oligonucleotide Synthesis: Methodsand Applications, Humana Press; and Luo, Y. (2019) CRISPR Gene Editing: Methods and Protocols, Springer New York. Each of these general texts is herein incorporated by reference. In a fourth aspect is provided a use of a system according to the first aspect, the isolated cell of the second aspect, or the pharmaceutical composition of the third aspect for the manufacture of a medicament.
[0326] In a fifth aspect is provided a product comprising:
[0327] (a) a first nucleic acid molecule comprising a 5' coding sequence (CDS), wherein the 5' CDS encodes an N-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide, and a promoter upstream of the 5' CDS; and
[0328] (b) a second nucleic acid molecule comprising a 3' CDS, wherein the 3' CDS encodes a C-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide, and a promoter upstream of the 3' CDS,
[0329] as a combined preparation for simultaneous, separate or sequential use in therapy, wherein a CDS encoding the COL4A3, COL4A4 or COL4A5 polypeptide is reconstituted at the mRNA level and the COL4A3, COL4A4 or COL4A5 polypeptide is expressed upon administration of the first nucleic acid molecule and the second nucleic acid molecule to a subject.
[0330] The 5' CDS may be a mutant 5' CDS that has been mutated to remove one or more cryptic splice donor sites that are present in the starting sequence, e.g. wild-type. Thus, in some embodiments, the 5' CDS comprises at least one mutation to remove a cryptic splice donor site from the coding sequence. The disclosures above regarding cryptic splice donor site mutations also apply to this aspect.
[0331] In some embodiments, the first nucleic acid molecule comprises a 5'CDS comprising exons 1-9 of COL4A5 and exon 8 and / or exon 9 comprise at least one mutation of a cryptic splice donor site in the coding sequence of COL4A5.
[0332] In some embodiments, where a COL4A5 polypeptide is reconstituted at the mRNA level, the 5' CDS comprises exons 1-31 of COL4A5, and wherein at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or all ten ofexons 6, 12-13, 15-16, 20, 21, 22, 24, 26, 29-30 and 31 comprise at least one mutation of a cryptic splice donor site in the coding sequence of COL4A5.
[0333] In some embodiments where a COL4A5 polypeptide is reconstituted at the mRNA level: i) the sequence of exon 6 consists of the sequence GGAATGCCAGGCCACGATGGGGCCCCAGGACCTCAGGGAATTCCCGGATGCAATGGAACCA AG (SEQ ID NO: 107);
[0334] ii) the sequence of exon 8 consists of the sequence GGACCCCCTGGGATCCCCGGCATGAAG (SEQ ID NO: 108);
[0335] iii) the sequence of exon 9 consists of the sequence
[0336] G GTG A ACCTG G C AG C ATA ATTATGTC ATC ACTG CC AG G ACC A A AG G GTA ATCC AG G ATATCC A GGTCCTCCTGGAATACAA (SEQ ID NO: 109);
[0337] iv) the sequence of exon 12 consists of the sequence GGGAATATGGGCTTAAATTTCCAGGGACCCAAAGGTGAAAAG (SEQ ID NO: 110);
[0338] v) the sequence of exon 13 consists of the sequence GGAGAGCAAGGTCTTCAGGGCCCACCTGGGCCACCTGGGCAGATCAGTGAACAGAAAAGA CCAATTGATGTAGAGTTTCAGAAAGGAGATCAG (SEQ ID NO: 111);
[0339] vi) the sequence of exon 15 consists of the sequence GGTCCCCCAGGTGGTGAGAAGGGAGAGAAGGGTGAGCAAGGAGAGCCAGGCAAACGG
[0340] (SEQ ID NO: 112);
[0341] vii) the sequence of exon 16 consists of the sequence GGAAAACCAGGCAAAGATGGAGAAAATGGCCAACCAGGAATTCCT (SEQ ID NO: 113); viii) the sequence of exon 20 consists of the sequence GGGCTGCAGTTATGGGTCCTCCTGGCCCTCCTGGATTTCCTGGAGAAAGGGGTCAGAAGGG AGATGAAGGACCACCTGGAATTTCCATTCCTGGACCTCCTGGACTTGACGGACAGCCTGGGGCTCCTG GGCTTCCAGGGCCTCCTGGCCCTGCTGGCCCTCACATTCCTCCTA (SEQ ID NO: 114);
[0342] ix) the sequence of exon 21 consists of the sequence GTGATGAGATATGTGAACCAGGCCCTCCAGGCCCCCCAGGATCTCCTGGAGATAAAGGACTC CAAGGAGAACAAGGAGTGAAAG (SEQ ID NO: 115);
[0343] x) the sequence of exon 22 consists of the sequence GTGACAAGGGAGACACTTGCTTCAACTGCATTGGAACTGGTATTTCAGGGCCTCCAGGTCAA CCTGGTTTGCCAGGTCTCCCAGGTCCTCCAG (SEQ ID NO: 116);xi) the sequence of exon 24 consists of the sequence GGCATTCCAGGAGCTCCAGGTGCTCCAGGCTTTCCTGGATCTAAGGGAGAACCTGGTGATAT CCTCACTTTTCCAGGAATGAAGGGTGACAAAGGAGAGTTGGGTTCCCCTGGAGCTCCAGGGCTTCCT GGTTTACCTGGCACTCCTGGACAGGATGGATTGCCAGGGCTTCCTGGCCCGAAAGGAGAGCCT (SEQ ID NO: 117);
[0344] xii) the sequence of exon 26 consists of the sequence GTCCTAAAGGGGATCCTGGACAGACTATAACCCAGCCGGGGAAGCCTGGCTTGCCTGGTAAC CCAGGCAGAGATGGTGACGTAGGTCTTCCAG (SEQ ID NO: 118);
[0345] xiii) the sequence of exon 29 consists of the sequence GGTGAACCAGGATTTGCATTACCTGGGCCACCTGGGCCACCAGGACTTCCAGGTTTCAAAGG AGCACTTGGTCCAAAAGGTGATCGTGGTTTCCCAGGACCTCCGGGTCCTCCAGGACGCACTGGCTTA GATGGGCTCCCTGGACCAAAGG (SEQ ID NO: 119);
[0346] xiv) the sequence of exon 30 consists of the sequence GAGATGTTGGACCAAATGGACAACCTGGACCAATGGGACCTCCTGGGCTGCCAGGAATAGG TGTTCAGGGACCACCAGGACCACCAGGGATTCCTGGGCCAATAGGTCAACCTG (SEQ ID NO: 120); and / or
[0347] xv) the sequence of exon 31 consists of the sequence GTTTACATGGAATACCAGGAGAGAAGGGGGATCCAGGACCTCCTGGACTTGATGTTCCAGGA CCCCCAGGTGAAAGAGGCAGTCCAGGGATCCCCGGAGCACCTGGTCCTATAGGACCTCCAGGATCAC CAGGGCTTCCAGGAAAAGCTGGAGCCTCTGGATTTCCAG (SEQ ID NO: 121).
[0348] In some embodiments where a COL4A5 polypeptide is reconstituted at the mRNA level, the 5' CDS comprises or consists of SEQ ID NO: 123 and the 3'CDS comprises or consists of SEQ ID NO: 125.
[0349] In some embodiments where a COL4A5 polypeptide is reconstituted at the mRNA level, the 5' CDS comprises or consists of SEQ ID NO: 124 and the 3' CDS comprises or consists of SEQ ID NO: 125.
[0350] In some embodiments where a COL4A5 polypeptide is reconstituted at the mRNA level, the first nucleic acid molecule comprises in the 5' to 3' direction:
[0351] a 5'-inverted terminal repeat (5'-ITR);a promoter, optionally a kidney-specific promoter, optionally a podocyte-specific promoter;
[0352] a 5' CDS according to SEQ ID NO: 123 or SEQ ID NO: 124, said 5' CDS being operably linked to and under control of the promoter, optionally the kidney-specific promoter, optionally the podocyte-specific promoter;
[0353] optionally, a splice donor sequence;
[0354] a 3' self-cleaving ribozyme sequence, preferably a Twister ribozyme sequence; optionally a poly-adenylation sequence;
[0355] a 3'-inverted terminal repeat (3'-ITR)
[0356] and the second nucleic acid molecule comprises in the 5' to 3' direction:
[0357] a 5'-ITR;
[0358] a promoter, optionally a kidney-specific promoter, optionally a podocyte-specific promoter;
[0359] a 5' self-cleaving ribozyme sequence, preferably a Twister or RzB ribozyme sequence; optionally a splice acceptor sequence;
[0360] a 3' CDS according to SEQ ID NO: 125;
[0361] a WPRE;
[0362] a poly-adenylation sequence; and
[0363] a 3'-ITR;
[0364] optionally wherein the first nucleic acid molecule further comprises one or more regulatory elements.
[0365] The nucleic acid molecules described above may be vectors. Therefore, the system may be a dual vector system. The system may be a dual AAV vector system.
[0366] The system, cell, pharmaceutical composition or product of the invention may be administered to any subject in need thereof.
[0367] The subject may be a mammal. In preferred embodiments, the subject is a human.The subject may be an adult, an adolescent, a child, a toddler, and infant, or a neonate. In some embodiments, the subject is a paediatric patient. In some embodiments, the subject is a child, an infant, or a neonate. In some embodiments, the subject is an infant or a neonate.
[0368] In some embodiments, the subject is aged about 18 years or older. In some embodiments, the subject is aged about 16 years or younger. In some embodiments, the subject is aged about 12 years or younger. In some embodiments, the subject is aged about 6 months or older. In some embodiments, the subject is aged about 1 years or older.
[0369] The subject may be male or female. The subject may be a male with X-chromosome linked Alport syndrome. In some embodiments, the subject is a male child, infant, or neonate with X-chromosome linked Alport syndrome.
[0370] Treatments
[0371] In a sixth aspect there is provided a system according to the first aspect, an isolated cell according to the second aspect, a pharmaceutical composition according to the third aspect, or a product according to the fifth aspect for use in preventing and / or treating Alport Syndrome or for use in preventing and / or treating a condition in a subject who has a pathogenic variant in the COL4A3, COL4A4 or COL4A5 gene.
[0372] The subject may have or may be at risk of Alport Syndrome. In some embodiments, the subject has or is at risk of Alport Syndrome.
[0373] Alport syndrome is a genetic condition characterised by kidney disease and extrarenal manifestations, such as loss of hearing and eye abnormalities. Alport syndrome is also known as familial nephritis, hereditary nephritis, thin basement membrane disease and thin basement membrane nephropathy. Alport syndrome is caused by pathogenic variants in the COL4A3, COL4A4 and COL4A5 genes, which result in abnormalities of the collagen IV a345 network of basement membranes. The condition can be transmitted in an X-linked, autosomal dominant, or autosomal recessive pattern, with X-linked being the common while autosomal recessive and autosomal dominant account for around 15% and 20% of cases respectively (see e.g. Warady, B.A., et al., 2020. Kidney medicine, 2(5), pp.639-649).The system, cell, pharmaceutical composition or product according to the present invention may be administered to a subject with Alport Syndrome in order to reverse the rejection or slow down progression of Alport Syndrome, or to lessen, reduce, or improve at least one symptom of Alport Syndrome.
[0374] Alport syndrome may be diagnosed by any method known in the art, including urinalysis, kidney biopsy and / or by genetic testing. For example, all males with X-linked Alport syndrome, as well as all males and females with autosomal recessive Alport syndrome, have persistent microhematuria (see e.g. Kashtan, C.E., 2021. American Journal of Kidney Diseases, 77(2), pp.272-279).
[0375] The subject may have haematuria, microalbuminuria, macroalbuminuria or nephrotic range proteinuria. Subjects with Alport syndrome typically present with haematuria, which may progress to proteinuria. Haematuria can be determined by the presence of erythrocytes in urine when viewed microscopically. A basal microalbuminuria level of less than 30 mg / day is usually considered non-pathological. Levels of about 30 mg / day to about 300 mg / day are termed microalbuminuria, which is considered pathologic. Albumin levels of over 300 mg / day are termed macroalbuminuria and levels of proteinuria over 3.5 g / day are considered to be nephrotic range proteinuria. Treating subjects prior to onset of proteinuria may slow or prevent progression of proteinuria and thereby delay or prevent end-stage renal failure. Subjects with nephrotic range proteinuria may also be treated. As the collagen IV a345 network of the glomerular basement membrane is changed and normalised or repaired by the transgene, proteinuria levels should be progressively reduced.
[0376] The subject may additionally or alternatively test positive for a pathogenic variant of COL4A3, COL4A4 or COL4A5 (see e.g. Savige, J., et al., 2019. Pediatric Nephrology, 34, pp.1175-1189). Pathogenic variants of COL4A3 and COL4A4 can be heterozygous (autosomal dominant), or biallelic (autosomal recessive). Pathogenic variants in COL4A5 are hemizygous or heterozygous (X-linked). The subject is preferably treated with a dual AAV vector system encoding a COL4A3, COL4A4, or COL4A5 polypeptide corresponding to the gene(s) for which the subject has a pathogenic variant. For example, a subject having a pathogenic COL4A5variant can be treated with a dual AAV vector system encoding a COL4A5 polypeptide. The subject may test positive for two or more pathogenic variants of COL4A3, COL4A4 or COL4A5. Such subjects may be treated with two or more dual AAV vector systems encoding different transgenes, i.e., each dual AAV vector system encoding a COL4A3, COL4A4, or COL4A5 polypeptide corresponding to the genes for which the subject has pathogenic variants.
[0377] In some embodiments, the pathogenic variant may be a mutation in the coding sequence or non-coding sequence of the COL4A3, COL4A4 or COL4A5 gene. In some embodiments, the pathogenic variant may be a mutation in the associated promoter or regulatory element(s) of the COL4A3, COL4A4 or COL4A5 gene.
[0378] In some embodiments, the Alport Syndrome is X-chromosome linked Alport syndrome or autosomal Alport syndrome.
[0379] In some embodiments, the Alport Syndrome is X-chromosome linked Alport syndrome. X-chromosome linked Alport syndrome (XLAS) is typically associated with a pathogenic variant of COL4A5 and accounts for about 85% of cases of Alport syndrome (see e.g. Hashimura, Y., et al., 2014. Kidney international, 85(5), pp.1208-1213). Over 300 mutations have been observed in the COL4A5 genes in families with XLAS, including splice-site mutations, missense mutations, and deletions of fewer than ten base pairs.
[0380] In some embodiments, the Alport Syndrome is autosomal Alport syndrome. Autosomal Alport syndrome is typically associated with pathogenic variants of COL4A3 or COL4A4. In some embodiments, the Alport Syndrome is autosomal recessive Alport syndrome. Autosomal recessive Alport syndrome (ARAS) accounts for about 10-15% of cases of Alport syndrome. In some embodiments, the Alport Syndrome is autosomal dominant Alport syndrome. Autosomal dominant Alport syndrome (ADAS) is a rare form of Alport syndrome. In patients with Alport syndrome, to date, only six mutations in the COL4A3 gene and twelve mutations in the COL4A4 gene are seen in patients with ARAS. The mutations include frameshift deletions, amino acid substitutions, missense mutations, splicing mutations, and in-frame deletions (see e.g. Watson S, et al. Alport Syndrome. In: StatPearls).Kits
[0381] In a seventh aspect is provided a kit comprising a) the first nucleic acid molecule as defined in the first aspect; and b) the second nucleic acid molecule as defined in the first aspect. The nucleic acid molecules described above may be vectors. Therefore, the kit may comprise a first vector and a second vector. The vector may be an AAV vector. Therefore, the kit may comprise a first AAV vector and a second AAV vector.
[0382] The kit may be a virus packaging kit or a virus production kit. As used herein, a "virus packaging kit or system" may comprise one or more components, and optionally instructions, for packaging the viral vector of the present invention. As used herein, a "virus production kit or system" may comprise one or more components, and optionally instructions, for producing the viral vector of the present invention.
[0383] The kit may comprise a transfer vector encoding the system of the invention and optionally one or more helper vectors. The kit may further comprise host cells (e.g. packaging cells or producer cells) and / or other reagents (e.g. transfection reagent, culture medium, etc.). The kit may further comprise any other suitable components, and optionally instructions for packaging and / or producing the system of the invention.
[0384] The system may be a dual vector system. The system may be a dual AAV vector system.
[0385] In some embodiments, the kit is for production of vector particles and comprises a transfer vector (e.g. plasmid) encoding the first vector of the present invention, and one or more helper plasmids encoding replication and capsid proteins. In some embodiments, the kit is for production of vector particles and comprises a transfer vector (e.g. plasmid) encoding the second vector of the present invention, and one or more helper plasmids encoding replication and capsid proteins. In some embodiments, the kit is for production of vector particles and comprises a transfer vector (e.g. plasmid) encoding the first vector of the present invention, a transfer vector (e.g. plasmid) encoding the second vector of the present invention, and one or more helper plasmids encoding replication and capsid proteins.In an eighth aspect is provided a method of transducing a cell ex vivo, wherein the cell is transduced with the system of the first aspect. In some embodiments, the cell is a glomerular cell. In some embodiments, the glomerular cell is a podocyte.
[0386] Variants, derivatives, homologues and fragments
[0387] In addition to the specific proteins and nucleotides mentioned herein, the invention also encompasses variants, derivatives, homologues and fragments thereof.
[0388] In the context of the invention, a "variant" of any given sequence is a sequence in which the specific sequence of residues (whether amino acid or nucleic acid residues) has been modified in such a manner that the polypeptide or polynucleotide in question retains at least one or all of its endogenous functions. A variant sequence can be obtained by addition, deletion, substitution, modification, replacement and / or variation of at least one residue present in the given sequence.
[0389] The term "derivative" as used herein in relation to proteins or polypeptides of the invention includes any substitution of, variation of, modification of, replacement of, deletion of and / or addition of one (or more) amino acid residues from or to the sequence, providing that the resultant protein or polypeptide retains at least one or all of its endogenous functions.
[0390] Typically, amino acid substitutions may be made, for example from 1, 2 or 3, to 10 or 20 substitutions, provided that the modified sequence retains the required activity or ability. Amino acid substitutions may include the use of non-naturally occurring analogues.
[0391] Polypeptides used in the invention may also have deletions, insertions or substitutions of amino acid residues which produce a silent change and result in a functionally equivalent protein. Deliberate amino acid substitutions may be made on the basis of similarity in polarity, charge, solubility, hydrophobicity, hydrophilicity and / or the amphipathic nature of the residues as long as the endogenous function is retained. For example, negatively charged amino acids include aspartic acid and glutamic acid; positively charged amino acids include lysine and arginine; and amino acids with uncharged polar head groups having similar hydrophilicity values include asparagine, glutamine, serine, threonine and tyrosine.Conservative substitutions may be made, for example according to the table below. Amino acids in the same block in the second column and preferably in the same line in the third column may be substituted for each other:
[0392] ALIPHATIC Non-polar G A P
[0393] I LV
[0394] Polar - uncharged C S T M
[0395] N Q
[0396] Polar - charged D E
[0397] K R H
[0398]
[0399] AROMATIC F W Y
[0400] The effect of additions, deletions, substitutions, modifications, replacements and / or variations may be predicted using any suitable prediction tool e.g. SIFT (Vaser, R., et al., 2016. Nature protocols, 11(1), pp.1-9), PolyPhen-2 (Adzhubei, I., et al., 2013. Current protocols in human genetics, 76(1), pp.7-20), CADD (Rentzsch, P., et al., 2021. Genome medicine, 13(1), pp.1-12), REVEL (loannidis, N.M., et al., 2016. The American Journal of Human Genetics, 99(4), pp.877-885), MetaLR (Dong, C., et al., 2015. Human molecular genetics, 24(8), pp.2125-2137), and / or MutationAssessor (Reva, B., et al., 2011. Nucleic acids research, 39(17), pp.ell8-ell8) or based on clinical data e.g. ClinVar (Landrum, M.J., et al., 2016. Nucleic acids research, 44(D1), pp.D862-D868). Suitable additions, deletions, substitutions, modifications, replacements and / or variations may be considered tolerated, benign, and / or likely benign.
[0401] In the present context, a variant sequence is taken to include an amino acid sequence which may be at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85% or at least 90% identical, suitably at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the subject sequence. Although a variant can also be considered in terms of similarity (i.e. amino acid residues having similar chemical properties / functions), in the context of the present invention it is preferred to express it in terms of sequence identity.
[0402] In the present context, a variant sequence is taken to include a nucleotide sequence which may be at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, atleast 80%, at least 85% or at least 90% identical, suitably at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the subject sequence. Although a variant can also be considered in terms of similarity, in the context of the present invention it is preferred to express it in terms of sequence identity.
[0403] The term "homologue" as used herein means a variant having a certain similarity with the wild type amino acid sequence or the wild type nucleotide sequence, e.g. having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85% or at least 90% similarity, suitably at least 95%, at least 96%, at least 97%, at least 98% or at least 99% similarity to the subject sequence.
[0404] Suitably, reference to a sequence which has a percent identity to any one of the SEQ ID NOs detailed herein refers to a sequence which has the stated percent identity over the entire length of the SEQ ID NO referred to.
[0405] Sequence identity comparisons can be conducted by eye, or more usually, with the aid of readily available sequence comparison programs. These commercially available computer programs can calculate percent identity between two or more sequences.
[0406] Percent identity may be calculated over contiguous sequences, i.e. one sequence is aligned with the other sequence and each amino acid or nucleotide in one sequence is directly compared with the corresponding amino acid or nucleotide in the other sequence, one residue at a time. This is called an "ungapped" alignment. Typically, such ungapped alignments are performed only over a relatively short number of residues.
[0407] Although this is a very simple and consistent method, it fails to take into consideration that, for example, in an otherwise identical pair of sequences, one insertion or deletion in the amino acid or nucleotide sequence may cause the following residues or codons to be put out of alignment, thus potentially resulting in a large reduction in percent identity when a global alignment is performed. Consequently, most sequence comparison methods are designed to produce optimal alignments that take into consideration possible insertions and deletionswithout penalising unduly the overall identity score. This is achieved by inserting "gaps" in the sequence alignment to try to maximise local identity.
[0408] However, these more complex methods assign "gap penalties" to each gap that occurs in the alignment so that, for the same number of identical amino acids or nucleotides, a sequence alignment with as few gaps as possible, reflecting higher relatedness between the two compared sequences, will achieve a higher score than one with many gaps. "Affine gap costs" are typically used that charge a relatively high cost for the existence of a gap and a smaller penalty for each subsequent residue in the gap. This is the most commonly used gap scoring system. High gap penalties will produce optimised alignments with fewer gaps. Most alignment programs allow the gap penalties to be modified. However, it is preferred to use the default values when using such software for sequence comparisons. For example, when using the GCG Wisconsin Bestfit package the default gap penalty for amino acid sequences is -12 for a gap and -4 for each extension.
[0409] Calculation of maximum percent identity therefore firstly requires the production of an optimal alignment, taking into consideration gap penalties. A suitable computer program for carrying out such an alignment is the GCG Wisconsin Bestfit package (see e.g. Devereux, J., et al., 1984. Nucleic acids research, 12(1), pp.387-395). Examples of other software that can perform sequence comparisons include, but are not limited to, the BLAST package (see e.g. Altschul, S.F., et al., 1990. Journal of molecular biology, 215(3), pp.403-410), BLAST 2 (see e.g. Tatusova, T.A. and Madden, T.L., 1999. FEMS microbiology letters, 174(2), pp.247-250), FASTA (see e.g. Pearson, W.R. and Lipman, D.J., 1988. PNAS, 85(8), pp.2444-2448.), EMBOSS Needle (Madeira, F., et al., 2019. Nucleic acids research, 47(W1), pp.W636-W641) and the GENEWORKS suite of comparison tools. For some applications, it is preferred to use EMBOSS Needle.
[0410] Although the final percent identity can be measured, the alignment process itself is typically not based on an all-or-nothing pair comparison. Instead, a scaled similarity score matrix is generally used that assigns scores to each pairwise comparison based on chemical similarity or evolutionary distance. An example of such a matrix commonly used is the BLOSUM62 matrix.Once the software has produced an optimal alignment, it is possible to calculate percent sequence identity. The software typically does this as part of the sequence comparison and generates a numerical result. The percent sequence identity may be calculated as the number of identical residues as a percentage of the total residues in the SEQ ID NO referred to.
[0411] The term "fragment" as used herein refers to a variant sequence that is a portion of a full-length polypeptide or polynucleotide. Fragments are typically selected regions of the polypeptide or polynucleotide that is of interest either functionally or, for example, in an assay. Such variants, derivatives, homologues and fragments may be prepared using standard recombinant DNA techniques such as site-directed mutagenesis. Where insertions are to be made, synthetic DNA encoding the insertion together with 5' and 3' flanking regions corresponding to the naturally-occurring sequence either side of the insertion site may be made. The flanking regions will contain convenient restriction sites corresponding to sites in the naturally-occurring sequence so that the sequence may be cut with the appropriate enzyme(s) and the synthetic DNA ligated into the cut. The DNA is then expressed in accordance with the invention to make the encoded protein. These methods are only illustrative of the numerous standard techniques known in the art for manipulation of DNA sequences and other known techniques may also be used.
[0412] EXAMPLES
[0413] The invention will now be further described by way of Examples, which are meant to serve to assist one of ordinary skill in the art in carrying out the invention and are not intended in any way to limit the scope of the invention.
[0414] WO 2021 / 158964 demonstrates that engineered ribozymes and RNA ligases can be used to "stitch" multiple RNA fragments together to generate a full-length functional RNA within cells, thereby overcoming viral vector size limitations and enabling delivery of large genes such as dystrophin, which exceeds 11 kb in coding length. However, this prior art contains no teaching or suggestion that applying this RNA recombination technology to the COL4A5 coding sequence would trigger aberrant splicing or lead to production of truncated COL4A5protein isoforms. On the contrary, the prior art consistently reports the generation of correctly spliced, full-length proteins only.
[0415] In contrast, the present inventors unexpectedly discovered that using the RNA recombination system on the wild-type COL4A5 coding sequence led to the production of truncated COL4A5 proteins. This was surprising because the wild-type COL4A5 sequence normally produces a full-length protein in human cells. Upon analysing the recombinant protein products, the inventors identified aberrant splicing at exon 8 of COL4A5. Initial efforts to correct this by mutating cryptic splice donor sites within exons 8 and 9 reduced the aberrant splicing at those positions but unexpectedly induced new aberrant splicing events elsewhere in the 5' coding sequence. Through systematic identification and mutation of all cryptic splice sites within the 5' portion of COL4A5, the inventors succeeded in eliminating aberrant splicing and restoring production of full-length COL4A5 protein.
[0416] The present inventors are the first to demonstrate that the COL4A5 coding sequence undergoes aberrant splicing when used in a dual-vector RNA recombination system, and the first to demonstrate that full-length COL4A5 can be successfully produced using a dual vector RNA recombination system.
[0417] Example 1 - Generation of AAV vector plasmids encoding a portion of wild-type human COL4A5 polypeptide
[0418] Plasmids were designed and generated encoding either an N-terminal portion of a wild type human COL4A5 polypeptide (exons 1-31) or a C-terminal portion of a wild type human COL4A5 polypeptide (exons 32-51) (see Figure 1).
[0419] Systems PS001 (wtCOL4A5 TwstR / TwstR) and PS002 (wtCOL4A5 TwstR / Rzb) were prepared featuring a split COL4A5 coding sequence. Plasmids were designed and generated encoding either an N-terminal portion of a wild type human COL4A5 polypeptide (exons 1-31) or a C-terminal portion of a wild type human COL4A5 polypeptide (exons 32-51) (as per Figure 1).
[0420] The first plasmid comprised the 5' CDS and further contained a CMV promoter upstream of the 5' CDS, a splice donor sequence, a TwstR ribozyme sequence and a poly A tail downstream of the 5' CDS (See Figure IB).The second plasmid comprised the 3' CDS and further contained a CMV promoter, a TwstR (PS001) or RzB (PS002) ribozyme sequence, and a splice acceptor sequence upstream of the 3' CDS, and a V5 peptide tag, WPRE and poly A tail downstream of the 3' CDS (See Figure IB).
[0421] HEK293 cells were transfected with system PS001 or PS002. Western blots revealed that some full-length COL4A5 protein was expressed (see lanes labelled wt: TwstR / TwstR and wt: TwstR / Rzb - Figure 2).
[0422] Example 2 - Transfection of AD-293 cells with plasmids encoding wild-type human COL4A5 with splice donor site mutations
[0423] Mutations of cryptic splice donor sites in exons 8 and 9 were made as sequencing analysis of demonstrated aberrant splicing occurring from a cryptic splice donor site in exon 8.
[0424] Systems PS003 (Exon 8 / 9 SD Correction: TwstR / TwstR) and PS004 (Exon 8 / 9 Correction: TwstR / RzB) were prepared as above with SEQ ID No: 108 in place of exon 8 and SEQ ID NO: 109 in place of exon 9 in the 5' CDS. These sequences comprise mutations to remove cryptic splice donor sites identified in those exons.
[0425] AD-293 cells were co-transfected with PS003 or PS004. Western blots using either the H53 antibody (an anti-COL4A5 antibody, (Chondrex, 7078)) or an anti-V5 antibody revealed that these mutations lead to more full-length protein being expressed (see Figure 2A and Figure 2B, respectively).
[0426] Further mutations were introduced to remove other cryptic splice donor sites identified in the 5' CDS so that exons 6, 8, 9, 12-13, 15-16, 20, 21, 22, 24, 26, 29-30 and 31 contained mutations, this sequence was used in systems PS005 (Full SD Correction: TwstR / TwstR) and PS006 (Full SD Correction: TwstR / RzB).
[0427] AD-293 cells were transfected with system PS005 or system PS006. Western blots revealed that these mutations lead to almost exclusively full-length protein being expressed (see Figure 2A and Figure 2B).Example 3 - Dual AAV vector system transduction with cryptic splice donor site correction produces full length COL4A5 in human podocytes.
[0428] The AAV vector plasmids of Example 2 were used to generate the dual vector system with all the cryptic splice sites removed from the 5' CDS, PS005 (Full SD Correction: TwstR / TwstR), in encapsidated form, i.e. as AAV viral particles.
[0429] Human podocytes were transduced with the dual AAV vectors of system PS005 at a range of multiplicity of infection (MOI). A single plasmid expressing full-length COL4A5 was used to transfect HEK293 cells (as a positive control) - see Figure 3A.
[0430] Western blots performed with either an anti-COL4A5 antibody or an anti-V5 antibody revealed that the mutations in the 5'CDS of the first vector of system P005 lead to almost exclusively full-length protein being expressed - see Figure 3B and Figure 3C, respectively.
[0431] Example 4 - Dual AAV vector system transduction with cryptic splice donor site correction produces full length COL4A5 in human podocytes.
[0432] Human podocytes were transduced with the dual AAV vectors of system PS005.
[0433] A western blot using the Jess™ Automated Western Blot System (Bio-Techne) performed with an anti-V5 antibody revealed that the mutations of the cryptic splice donor sites in the 5' CDS of the first vector of system PS005 lead to exclusively full-length protein being expressed - see Figure 4, lane 3.
[0434] Example 5 - Dual AAV vector system transduction with cryptic splice donor site correction produces full length COL4A5 in AD-293 cells.
[0435] AD-293 cells were transduced with the dual AAV vectors of system PS005 at a range of MOI. Western blots performed with either an anti-COL4A5 antibody or an anti-V5 antibody revealed that the mutations of the cryptic splice donor sites in the 5' CDS of the first vector of system PS005 lead to almost exclusively full-length protein being expressed - see Figure 5A and Figure 5B, respectively.Examples 1-5 clearly demonstrate that mutation of the cryptic splice donor sites in the 5' CDS of COL4A5 is required to enable a full-length protein to be produced in the absence of truncated COL4A5 protein fragments.
[0436] Materials and Methods for all Examples
[0437] Cell Culture
[0438] AD-293 cells were cultured in DMEM (Sigma, D5796) supplemented with 10% FBS (Gibco, Cat. No. 16629525) and sodium pyruvate. Cells were passaged at a ratio of 1:10 once they reached a confluency of ~80%. AD-293 cells were maintained at 37°C, 5% CO2.
[0439] Conditionally immortalised podocytes were cultured in RPMI (Sigma, R8758) supplemented with 10% FBS and IX insulin, transferrin and selenium (ITS) (Gibco, 12097549). Conditionally immortalised podocytes were maintained at 33°C, 5% CO2. Cells were passaged at ~80% confluency at a ratio of 1:4.
[0440] Viral production cells were cultured in viral production medium (Gibco, A4817901), supplemented with 1% Glutamax (ThermoFisher, 35050061). Cells were passaged every 3-4 days, seeding at a density of 0.6e6 cells / mL in shake flasks.
[0441] AAV Production & QC
[0442] For AAV production, suspension vector production cells were seeded at a density of 3e6 cells / mL and transfected with Helper, Rep / Cap and Genome plasmids at a 1:1:1 molar ratio using AAV Max transfection kit (Invitrogen, A50515). Cells were lysed using AAV Max Lysis Buffer (Gibco, 17331899), debris removed by filtration through 0.22um filtration unit, and concentrated via ultracentrifugation on an iodixanol gradient (Sigma). Concentrated AAV was washed 5 times in with PBS supplemented with 0.001% Pluronic (ThermoFisher, 24040032) in a 15ml amicon filtration unit (Millipore, UFC910024). Finally, the AAV preparation was filter sterilised using 0.22pm syringe filters. AAV particles were quantified using qPCR methods, with primers targeting either the promoter (CMV) or WPRE regions in the vectors.
[0443] AD-293 TransfectionsFor AD-293 transfection experiments, cells were seeded into 6-well plates at a density of 2e5 cells per well the day before transduction, and allowed to adhere overnight at 37°C, 5% CO2. Plasmids for DNA dual vector technologies needed to be linearised prior to transfection. 100g of each plasmid was digested overnight in rCutsmart (NEB), using Nrul-HF restriction enzyme (NEB) and incubated overnight at 37°C. The following day, restriction digests were cleaned up using Monarch PCR clean-up kit (NEB) and eluted in EB buffer, then quantified by nanodrop. An aliquot was taken and run on 1% Agarose gel with SYBRsafe to confirm linearisation. Transfection mixes were prepared in sterile tubes using PEIPro transfection reagent in plain DMEM, with a total of lpg DNA per well. 3OO0L transfection mix added dropwise to each well and cells replaced at 37°C, 5% CO2 overnight. Media was changed for 2mL complete DMEM 24 hours after transduction, then cells returned to 37°C, 5% CO2. 48 hours after transfection, cell culture supernatant was collected for western blot analysis. Cell Transduction and Harvest
[0444] For AD-293 transduction experiments, cells were seeded into 6-well plates at a density of 2e5 cells per well the day before transduction, and allowed to adhere overnight at 37°C, 5% CO2. The following day viral particles were added to the wells as necessary at a dose of le5 AAV vg per cell. Cells were incubated overnight at 37°C, 5% CO2. Media was changed 24 hours following transduction, and cells were incubated for a further 48-hours. Cell culture supernatant was collected for western blot analysis 72-hours after transduction.
[0445] For podocyte transduction experiments, cells were seeded into 6-well plates at a density of 7.3e4 cells per well the day before transduction, and allowed to adhere overnight at 33°C, 5% CO2. The following day viral particles were added to the wells as necessary at a dose of le5 AAV vg per cell. Cells were incubated overnight at 33°C, 5% CO2. Media was changed 24 hours following transduction, and cells were thermoswitched at 37°C. 48 hours later, cell culture supernatant was collected for analysis.
[0446] Western Blot
[0447] Samples were thawed on ice then mixed with IM DTT (Sigma-Aldrich, 10197777001) and LDS Sample Buffer (Thermo-Fisher, NP0008), then vortexed, spun down and then denatured. Samples were separated by running in tris-acetate NuPAGE (Invitrogen) and Tris / Acetate - SDS Running buffer (Biorad, 1610732) at 150V for 75 minutes. Gels wererinsed in diHzO, then transferred on to nitrocellulose membrane (Invitrogen, IB23001) using the i Blot 2 device. Membrane was rinsed in diHzO and blocked for 1 hour in 5% milk in TBS-T. Antibodies were prepared: Rb Anti-V5 (Abeam, AB206566) and Rat anti-COL4A5 H53 (Chondrex, 7078) in 5% milk in TBS-T, then incubated with membranes overnight on a roller at 4°C. Membranes were washed in TBS-T, then incubated in secondary antibodies in 5% milk in TBS-T for one hour, using Goat anti-Rb IgG HRP (Abeam, AB97057) and Goat anti-Rat IgG HRP (Abeam, AB97057). Membranes were washed again in TBS-T, then in TBS before adding Supersignal West Atto ECL Reagent (Thermo Scientific, 17181819), then imaged on Amersham ImageQuant 800.
[0448] Jess™ Automated Capillary Electrophoresis
[0449] DTT, 5x master mix and biotinylated ladder from 66-440kDa standard kit (Bio-Techne, SWPS-ST03-1L) were prepared according to manufacturer's instructions prior to sample preparation. 0.1X sample buffer prepared by diluting 10X sample buffer from 66-440kDa separation module (Bio-Techne, SM-W005) with diH2O. Supernatant samples were thawed on ice then mixed with 5x master mix in 0.5mL protein LoBind tubes (Eppendorf, 10316752), vortexed and de-natured. Antibodies were diluted in Antibody Diluent (Bio-Techne, DM-001): Rat anti-COL4A5 H53 (Chondrex, 7078), Rb anti-V5 (Abeam, AB206566) and Goat AntiRat HRP (Jackson, 112-035-167-JIR-0.5ml). After denaturation, samples were briefly vortexed and spun down, then loaded on to the 66-440kDa separation module plate (Bio-Techne, SM-W005), along with Antibody Diluent, primary antibodies, secondary antibodies: Anti-Rabbit HRP (Bio-Techne, DM-001) and Goat Anti-Rat HRP (Jackson, 112-035-167-JIR-0.5ml) as well as 1:1 luminol peroxide. Plate spun down briefly at 1000G for 10 minutes, then wash buffer (Bio-Techne, SM-W005) was added and plate inserted in Jess™, along with a 25 capillary cassette. Assay run under standard conditions and analysed using Compass for Simple Western Software.
[0450] Prophetic Example 6 - Dual AAV Vector System for COL4A3
[0451] To evaluate whether a dual AAV mRNA-reconstitution system constructed for human COL4A3 enables expression of full-length COL4A3 protein in human podocytes following transduction,and whether silent mutations to remove cryptic splice donor sites improve the fidelity of protein production, two vectors will be generated.
[0452] 5' COL4A3 CDS vector
[0453] - 5' ITR
[0454] promoter: CMV or a podocyte-specific promoter, such as NPHS1 or NPHS2 - FLAG tag
[0455] 5' CDS: Exons 1-31 of COL4A3 (either wild-type or cryptic splice site corrected) splice donor sequence
[0456] - Twister ribozyme (TwstR)
[0457] polyA sequence
[0458] - 3' ITR
[0459] 3' COL4A3 CDS vector
[0460] - 5' ITR
[0461] promoter: CMV or a podocyte-specific promoter, such as NPHS1 or NPHS2 twister ribozyme (TwstR) or RzB
[0462] splice acceptor sequence
[0463] 3' CDS: Exons 32-52 of COL4A3 (either wild-type or cryptic splice site corrected) V5 tag
[0464] - WPRE
[0465] polyA sequence
[0466] - 3' ITR
[0467] Silent mutations will be incorporated into cryptic splice donor sites within the 5' CDS, and into cryptic splice acceptor sites within the 3' CDS, identified using Spliceator or equivalent algorithms, to reduce aberrant mRNA splicing.
[0468] AAV vectors will be produced and encapsidated into AAV serotypes that display podocyte tropism. AAV particles will be purified and undergo quality control assays. Conditionally immortalised human podocytes will be transduced as described above.Western blot analysis using both an anti-COL4A3 antibody or an anti-V5 antibody is expected to demonstrate high levels of predominantly full-length COL4A3 protein where the cryptic splice donor sites were mutated with no truncated COL4A3 isoforms being produced.
[0469] This will demonstrate that the dual-AAV, ribozyme-enabled mRNA reconstitution system is generalizable across collagen IV family genes, and not limited to COL4A5.
[0470] Prophetic Example 7 - Dual AAV Vector System for COL4A4
[0471] To evaluate whether a dual AAV mRNA-reconstitution system constructed for human COL4A4 enables expression of full-length COL4A4 protein in human podocytes following transduction, and whether silent mutations to remove cryptic splice donor sites improve the fidelity of protein production, two vectors will be generated.
[0472] 5' COL4A4 CDS vector
[0473] - 5' ITR
[0474] promoter: CMV or a podocyte-specific promoter, such as NPHS1 or NPHS2 - FLAG tag
[0475] 5' CDS: Exons 1-31 of COL4A4 (either wild-type or cryptic splice site corrected) splice donor sequence
[0476] - Twister ribozyme (TwstR)
[0477] polyA sequence
[0478] - 3' ITR
[0479] 3' COL4A4 CDS vector
[0480] - 5' ITR
[0481] promoter: CMV or a podocyte-specific promoter, such as NPHS1 or NPHS2 twister ribozyme (TwstR) or RzB
[0482] splice acceptor sequence
[0483] 3' CDS: Exons 32-47 of COL4A4 (either wild-type or cryptic splice site corrected) V5 tag
[0484] WPREpolyA sequence
[0485] 3' ITR
[0486] Silent mutations will be incorporated into cryptic splice donor sites, and into cryptic splice acceptor sites within the 3' CDS, identified using Spliceator or equivalent algorithms, to reduce aberrant mRNA splicing.
[0487] AAV vectors will be produced and encapsidated into AAV serotypes that display podocyte tropism. AAV particles will be purified and undergo quality control assays. Conditionally immortalised human podocytes will be transduced as described above.
[0488] Western blot analysis using both an anti-COL4A4 antibody or an anti-V5 antibody is expected to demonstrate high levels of predominantly full-length COL4A4 protein where the cryptic splice donor sites were mutated with no truncated COL4A4 isoforms being produced.
[0489] This will demonstrate that the dual-AAV, ribozyme-enabled mRNA reconstitution system is generalizable across collagen IV family genes, and not limited to COL4A5.
Claims
1. CLAIMS1. A system for generating a coding sequence of a COL4A3, COL4A4 or COL4A5 polypeptide in a host cell comprising:(a) a first nucleic acid molecule comprising a 5' coding sequence (CDS), wherein the 5' CDS encodes an N-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide, and a promoter upstream of the 5' CDS; and(b) a second nucleic acid molecule comprising a 3' coding sequence (CDS), wherein the 3' CDS encodes a C-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide, and a promoter upstream of the 3'CDS,wherein the coding sequence of the COL4A3, COL4A4 or COL4A5 polypeptide is formed via reconstitution at the mRNA level following delivery of the first nucleic acid molecule and the second nucleic acid molecule into the host cell.
2. A system according to claim 1, wherein the first nucleic acid molecule is a vector and the second nucleic acid molecule is a vector.
3. A system according to claim 2, wherein the first and second vectors are AAV vectors.
4. A system according to any preceding claim, whereina) the first nucleic acid molecule comprises a first self-cleaving ribozyme sequence; andb) the second nucleic acid molecule comprises a second self-cleaving ribozyme sequence;wherein the first self-cleaving ribozyme sequence is positioned downstream of the 5' CDS; andwherein the second self-cleaving ribozyme sequence is positioned between the 3' CDS and the promoter.
5. The system according to claim 4, wherein the first and / or second self-cleaving ribozyme is selected from a group consisting of a member of the hammerhead (HH), hepatitisdelta virus (HDV), Twister, Varkud satellite, Twister-sister, hairpin, hatchet or pistol ribozyme families.
6. The system according to claim 4 or 5, wherein the first self-cleaving ribozyme is member of the Twister ribozyme family and the second self-cleaving ribozyme is a member of the Twister or HH family.
7. The system according to any one of claims 4 to 6, wherein the first self-cleaving ribozyme is a Twister ribozyme and the second self-cleaving ribozyme is a Twister ribozyme or an RzB ribozyme; optionally wherein the Twister ribozyme sequence comprises or consists of SEQ ID NO: 98 and the RzB ribozyme sequence comprises or consists of SEQ ID NO: 101.
8. The system according to any one of claims 4 to 7, wherein the first vector comprises a splice donor sequence between the 5' CDS and the first self-cleaving ribozyme sequence; and the second vector comprises a splice acceptor sequence between the second self-cleaving ribozyme and the 3' CDS.
9. A system according to any preceding claim, wherein the 5' CDS comprises at least one mutation to remove a cryptic splice donor site from the coding sequence.
10. A system according to claim 9 wherein the at least one mutation is silent.
11. A system according to claim 9 or 10 wherein the system is for generating a coding sequence of a COL4A5 polypeptide and wherein the 5' CDS comprises exons 1-9 of COL4A5 and wherein the at least one mutation is in exon 8 and / or exon 9.
12. A system according to any one of claims 9 to 11, wherein the 5' CDS comprises exons 1-31 of COL4A5, and wherein the at least one mutation removes one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, or all thirteen cryptic splice donor sites from the coding sequence of exons 6, 8, 9, 12-13, 15-16, 20, 21, 22, 24, 26, 29-30 and 31.
13. A system according to claim 11 or claim 12, wherein the cryptic splice donor site in the coding sequence of COL4A5 comprises an AGGT sequence and the at least one mutation comprises replacement of at least the thymine base.
14. A system according to claim 13 wherein the mutation further comprises replacement of the adenine in the sequence AGGT of the coding sequence of COL4A5.
15. A system according to claim 13 or claim 14 wherein the at least one mutation comprises replacement of at least 2, at least 3 or all 4 of the bases in the sequence AGGT of the coding sequence.
16. A system according to any one of claims 13 to 15 wherein the AGGT sequences are present at:i) position 36-39 of exon 6;ii) position 18-21 of exon 8;iii) position 9-12 of exon 9;iv) position 42 of exon 12 to position 3 of exon 13;v) position 21-24 of exon 15;vi) position 57 of exon 15 to position 3 of exon 16;vii) position 59-62 of exon 20;viii) position 47-50 of exon 21;ix) position 8-11 of exon 22;x) position 45-48 of exon 24;xi) position 17-20 of exon 26;xii) position 150 of exon 29 to position 2 of exon 30; and / orxiii) position 149-152 of exon 31of the coding sequence of COL4A5.
17. A system according to any one of claims 12 to 15 wherein:i) the sequence of exon 6 consists of the sequence GGAATGCCAGGCCACGATGGGGCCCCAGGACCTCAGGGAATTCCCGGATGCAATGGAACCA AG (SEQ ID NO: 107);ii) the sequence of exon 8 consists of the sequence GGACCCCCTGGGATCCCCGGCATGAAG (SEQ ID NO: 108);iii) the sequence of exon 9 consists of the sequenceG GTG A ACCTG G C AG C ATA ATTATGTC ATC ACTG CC AG G ACC A A AG G GTA ATCC AG G ATATCC A GGTCCTCCTGGAATACAA (SEQ ID NO: 109);iv) the sequence of exon 12 consists of the sequence GGGAATATGGGCTTAAATTTCCAGGGACCCAAAGGTGAAAAG (SEQ ID NO: 110); v) the sequence of exon 13 consists of the sequence GGAGAGCAAGGTCTTCAGGGCCCACCTGGGCCACCTGGGCAGATCAGTGAACAGAAAAGA CCAATTGATGTAGAGTTTCAGAAAGGAGATCAG (SEQ ID NO: 111);vi) the sequence of exon 15 consists of the sequence GGTCCCCCAGGTGGTGAGAAGGGAGAGAAGGGTGAGCAAGGAGAGCCAGGCAAACGG(SEQ ID NO: 112);vii) the sequence of exon 16 consists of the sequence GGAAAACCAGGCAAAGATGGAGAAAATGGCCAACCAGGAATTCCT (SEQ ID NO: 113); viii) the sequence of exon 20 consists of the sequence GGGCTGCAGTTATGGGTCCTCCTGGCCCTCCTGGATTTCCTGGAGAAAGGGGTCAGAAGGG AGATGAAGGACCACCTGGAATTTCCATTCCTGGACCTCCTGGACTTGACGGACAGCCTGGGGCTCCTG GGCTTCCAGGGCCTCCTGGCCCTGCTGGCCCTCACATTCCTCCTA (SEQ ID NO: 114);ix) the sequence of exon 21 consists of the sequence GTGATGAGATATGTGAACCAGGCCCTCCAGGCCCCCCAGGATCTCCTGGAGATAAAGGACTC CAAGGAGAACAAGGAGTGAAAG (SEQ ID NO: 115);x) the sequence of exon 22 consists of the sequence GTGACAAGGGAGACACTTGCTTCAACTGCATTGGAACTGGTATTTCAGGGCCTCCAGGTCAA CCTGGTTTGCCAGGTCTCCCAGGTCCTCCAG (SEQ ID NO: 116);xi) the sequence of exon 24 consists of the sequence GGCATTCCAGGAGCTCCAGGTGCTCCAGGCTTTCCTGGATCTAAGGGAGAACCTGGTGATAT CCTCACTTTTCCAGGAATGAAGGGTGACAAAGGAGAGTTGGGTTCCCCTGGAGCTCCAGGGCTTCCT GGTTTACCTGGCACTCCTGGACAGGATGGATTGCCAGGGCTTCCTGGCCCGAAAGGAGAGCCT (SEQ ID NO: 117);xii) the sequence of exon 26 consists of the sequenceGTCCTAAAGGGGATCCTGGACAGACTATAACCCAGCCGGGGAAGCCTGGCTTGCCTGGTAAC CCAGGCAGAGATGGTGACGTAGGTCTTCCAG (SEQ ID NO: 118);xiii) the sequence of exon 29 consists of the sequence GGTGAACCAGGATTTGCATTACCTGGGCCACCTGGGCCACCAGGACTTCCAGGTTTCAAAGG AGCACTTGGTCCAAAAGGTGATCGTGGTTTCCCAGGACCTCCGGGTCCTCCAGGACGCACTGGCTTA GATGGGCTCCCTGGACCAAAGG (SEQ ID NO: 119);xiv) the sequence of exon 30 consists of the sequence GAGATGTTGGACCAAATGGACAACCTGGACCAATGGGACCTCCTGGGCTGCCAGGAATAGG TGTTCAGGGACCACCAGGACCACCAGGGATTCCTGGGCCAATAGGTCAACCTG (SEQ ID NO: 120); and / orxv) the sequence of exon 31 consists of the sequence GTTTACATGGAATACCAGGAGAGAAGGGGGATCCAGGACCTCCTGGACTTGATGTTCCAGGA CCCCCAGGTGAAAGAGGCAGTCCAGGGATCCCCGGAGCACCTGGTCCTATAGGACCTCCAGGATCAC CAGGGCTTCCAGGAAAAGCTGGAGCCTCTGGATTTCCAG (SEQ ID NO: 121).
18. The system according to any preceding claim for expressing a human COL4A5 polypeptide.
19. The system according to any preceding claim for expressing a full-length COL4A5 polypeptide, preferably a full-length human COL4A5 polypeptide.
20. The system according to any preceding claim, wherein the promoter upstream of the 5' CDS and / orthe promoter upstream of the 3' CDS are selected from: a CMV promoter, a CBA promoter, a CAG promoter, a CB7 promoter, an EFla promoter, a CMV / EFla hybrid promoter, a NF-kB promoter, a pSE-7 promoter, a mPGK promoter, a mUla promoter, a U6 promoter, a U7 promoter, a MNDU3 promoter, a HLP promoter, an AAT promoter, an ALB promoter, a ApoE / AAT promoter, a EalbAAT promoter, a LP1 promoter, a TBG promoter, a TTR promoter, a SYN1 promoter, a NSE promoter, a tMCK promoter, a CK8 promoter, a MHCK7 promoter, a SMN promoter, a DES promoter, a RK promoter, a hRHO promoter, a GRK1 promoter, a hCAR promoter, a hRPE65p promoter, a P546 promoter, a PR1.7 promoter, a hRSl promoter, a VMD2 promoter, an a-MHC promoter, a FREI promoter, a NPHS1 promoter, and a NPHS2 promoter.
21. The system according to any preceding claim, wherein the promoter upstream of the 5' CDS and / or upstream the 3' CDS is a kidney-specific promoter, optionally wherein the promoter is a podocyte-specific promoter or a tubular cell specific promoter.
22. The system according to any preceding claim, wherein the promoter is a NPHS1 promoter or a NPHS2 promoter, optionally a minimal NPHS1 promoter, a minimal NPHS2 promoter or the NPHS1265bp promoter.
23. The system according to any preceding claim, wherein the second nucleic acid molecule comprises one or more regulatory sequences downstream from the 3' CDS.
24. The system according to claim 23, wherein the second nucleic acid molecule comprises a Woodchuck hepatitis post-transcriptional regulatory element (WPRE), downstream from the 3' CDS.
25. The system according to any one of the preceding claims, wherein the first and / or second nucleic acid molecule comprises a polyadenylation sequence downstream from the 3' CDS.
26. The system according to any of the preceding claims, wherein the first nucleic acid molecule comprises in the 5' to 3' direction:a 5'-inverted terminal repeat (5'-ITR);a promoter, optionally a kidney-specific promoter, optionally a podocyte-specific promoter;the 5' CDS as defined in any one of claims 1 to 25, said 5' CDS being operably linked to and under control of the promoter, optionally the kidney-specific promoter, optionally the podocyte-specific promoter;optionally, a splice donor sequence;a 3' self-cleaving ribozyme sequence, preferably a Twister ribozyme sequence; optionally a poly-adenylation sequence;a 3'-inverted terminal repeat (3'-ITR)and the second nucleic acid molecule comprises in the 5' to 3' direction:a 5'-ITR;a promoter, optionally a kidney-specific promoter, optionally a podocyte-specific promoter;a 5' self-cleaving ribozyme sequence, preferably a Twister or RzB ribozyme sequence; optionally a splice acceptor sequence;a 3' CDS as defined in any one of claims 1 to 25;a WPRE;a poly-adenylation sequence; anda 3'-ITR;optionally wherein the first nucleic acid molecule further comprises one or more regulatory elements.
27. The system according to any preceding claim, wherein the first and / or the second nucleic acid molecule is an AAV vector and is in the form of an AAV vector particle encapsidated by LKO3, AAV3B, AAV9, ShHIO, AAV-DJ, AAV2, AAV6.2, AAV5, KPI, KP2 or KP3 capsid proteins.
28. The system according to claim 27, wherein the first and / or the second nucleic acid molecule is an AAV vector and is in the form of an AAV vector particle encapsidated by the LKO3 capsid protein.
29. An isolated cell comprising the system according to any preceding claim.
30. A pharmaceutical composition comprising the system according to any of claims 1 to 28 or a cell according to claim 29.
31. A system according to any of claims 1 to 28, a cell according to claim 29 or a pharmaceutical composition according to claim 30, for use as a medicament.
32. Use of a system according to any of claims 1 to 28, a cell according to claim 29 or a pharmaceutical composition according to claim 30, for the manufacture of a medicament for the treatment or prevention of Alport syndrome.
33. A product comprising:(a) a first first nucleic acid molecule comprising a 5' coding sequence (CDS), wherein the 5' CDS encodes an N-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide consisting of exons 1-31, and a promoter upstream of the 5' CDS; and(b) a second nucleic acid molecule comprising a 3' coding sequence (CDS), wherein the 3' CDS encodes a C-terminal part of the COL4A3, COL4A4 or COL4A5 polypeptide consisting of exons 32-51, and a promoter upstream of the 3'CDS, as a combined preparation for simultaneous, separate or sequential use in therapy, wherein a CDS encoding the COL4A3, COL4A4 or COL4A5 polypeptide is reconstituted at the mRNA level and the COL4A3, COL4A4, or COL4A5 polypeptide is expressed upon administration of the first nucleic acid molecule and the second nucleic acid molecule to a subject; and optionally wherein the 5' CDS comprises at least one mutation to remove a cryptic splice donor site from the coding sequence.
34. A product according to claim 33 wherein the first and second nucleic acid molecules are AAV vectors.
35. A system according to any of claims 1 to 28, a cell according to claim 29, a pharmaceutical composition according to claim 30, or the product of claims 33 or 34 for use in preventing and / or treating Alport Syndrome or for use in preventing and / or treating a condition in a subject who has a pathogenic variant in the COL4A3, COL4A4 or COL4A5 gene.
36. A kit comprising: (a) the first nucleic acid molecule as defined in the system according to any one of claims 1 to 30; and (b) the second nucleic acid molecule as defined in the system according to any one of claims 1 to 28.
37. The kit according to claim 36, for use in preventing and / or treating Alport Syndrome or for use in preventing and / or treating a condition in a subject who has a pathogenic variant in the COL4A3, COL4A4 or COL4A5 gene, optionally wherein the first nucleic acid moleculeand the second nucleic acid molecule are administered simultaneously, separately or sequentially.
38. A method of transducing a cell ex vivo, wherein the cell is transduced with the system according to any one of claims 1 to 28, optionally wherein the cell is a glomerular cell, optionally wherein the glomerular cell is a podocyte.