Delivery of type iv collagen alpha chain using a split intein dual AAV vector

A split intein system delivered via AAV vectors splices together N- and C-terminal portions of the type IV collagen alpha chain to form functional collagen, addressing the inadequacies of current Alport syndrome treatments by potentially halting kidney disease progression.

WO2025193938A1PCT designated stage Publication Date: 2025-09-18OREGON HEALTH & SCI UNIV
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/019756
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-13
Filing Date
2025-03-13
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Current therapies for Alport syndrome, an inherited disorder caused by mutations in type IV collagen alpha chain genes, are inadequate in preventing the progression of kidney disease, leading to dialysis or kidney transplant in most patients, with a need for additional treatment options.

Method used

A split intein system is used, comprising nucleic acid molecules encoding a fusion protein that splices together N- and C-terminal portions of the type IV collagen alpha chain, delivered via adeno-associated viral (AAV) vectors to form the mature collagen chain, potentially treating Alport syndrome.

Benefits of technology

The split intein system effectively forms functional type IV collagen alpha chains in mammalian cells, demonstrating therapeutic potential in delaying or reversing kidney disease progression in Alport syndrome models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025019756_18092025_PF_FP_ABST
    Figure US2025019756_18092025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein is a split intein system including a first nucleic acid molecule encoding a fusion protein including an N-extein of an extein pair and an N-intein of an intein pair, where the N-extein includes a first signal sequence N-terminal to an N-terminal portion of a type IV collagen alpha chain. Also disclosed is a second nucleic acid molecule encoding a fusion protein including a second signal sequence, a C-intein of the intein pair, and a C-extein of the extein pair, where the C-extein includes a C-terminal portion of the type IV collagen alpha chain. Further disclosed are AAV vectors including the nucleic acid molecules, and methods of using the AAV vectors or nucleic acid molecules to treat a subject with Alport syndrome.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DELIVERY OF TYPE IV COLLAGEN ALPHA CHAIN USING A SPLIT INTEIN DUAL AAV VECTOR

[0002] CROSS REFERENCE TO RELATED APPLICATIONS

[0003] This application claims priority to U.S. Provisional Application No. 63 / 564,923, filed March 13, 2024, which is incorporated by reference in its entirety.

[0004] INCORPORATION OF ELETRONIC SEQUENCE LISTING

[0005] The Sequence Listing is submitted as an XML file in the form of the file named “899-111312- 02_Sequence ” (-621,127 bytes), which was created on February 19, 2025 which is incorporated by reference herein.

[0006] FIELD OF THE DISCLOSURE

[0007] This disclosure relates to gene therapy methods that include a split intein system, compositions, and systems for treating Alport syndrome, including the use of AAV vectors encoding the split intein system.

[0008] BACKGROUND

[0009] Alport syndrome, also known as hereditary nephritis, is an inherited disorder caused by mutations in type IV collagen alpha chain genes. The mutations prevent normal production and / or assembly of the type IV collagen network, which forms aspects of the kidneys, eyes, and inner ear. These mutations are associated with thinning, thickening, splitting, and lamellation of type IV collagen membranes. While there is currently no cure for Alport syndrome, therapies can delay the progression of kidney disease. However, in practice, most patients progress to requiring dialysis or a kidney transplant. In the United States, Alport syndrome accounts for about 2.2% of children and 0.2% of adults with end-stage renal disease. A need exists for additional therapies for the treatment of Alport syndrome.

[0010] SUMMARY OF THE DISCLOSURE

[0011] Disclosed herein is a split intein system including a first nucleic acid molecule encoding a fusion protein including in N to C terminal order, an N-extein of an extein pair and an N-intein of an intein pair, where the N-extein includes a first signal sequence N-terminal to an N-terminal portion of a type IV collagen alpha chain. The split intein system includes a second nucleic acid molecule encoding a fusion protein including in N to C terminal order, a second signal sequence, a C-intein of the intein pair, and a C- extein of the extein pair, where the C-extein includes a C-terminal portion of the type IV collagen alpha chain. In some examples, the type IV collagen alpha chain is COL4A5. The N- and C-terminal portions of the type IV collagen alpha chain together define the type IV collagen alpha chain sequence separated at a split point. The N- and C-exteins in the spit intein system are spliced together, when expressed in mammalian cells, to form the mature type IV collagen alpha chain. In some aspects, the system includes a first adeno-associated viral (AAV) vector including the first nucleic acid molecule and / or a second AAV vector including the second nucleic acid molecule.

[0012] Also disclosed is a method of treating Alport syndrome, including administering to the subject a therapeutically effective amount of the split intein system or a pharmaceutical composition including the split intein system.

[0013] The foregoing and other features and advantages of the invention will become more apparent from the following detailed description of several examples which proceeds with reference to the accompanying figures.

[0014] BRIEF DESCRIPTION OF THE FIGURES

[0015] FIGs. 1A-1C. Table of exemplary intein insertion sites in COL4A5 and their characteristics. In the column labeled “Structure”: SS, signal sequence; CD, collagenous domain; Int, interrupting non- collagenous region; NCI, non-collagenous C-terminal domain. In the column labeled “Fit in AAV?”, “Yes” indicates the construct can fit in AAV when the CAG promoter-driven expression cassettes are employed; “Possible” indicates the construct can fit in AAV when a combination of a shorter enhancer-promoter and a shorter polyadenylation signal the combined size of which is approximately 0.2 kb or less are employed.

[0016] FIG. 2. Receiver operating characteristic (ROC) curve analysis of the support vector machine (SVM) model-based prediction of similarity to the native extein splice sites. SVM scores of a total of 752 labeled data were determined by leave-one-out cross validation, and used for drawing an ROC curve to assess the model. The dashed 45-degree line represents random guessing.

[0017] FIG. 3. Table providing a confusion matrix for extein splice site prediction by the SVM model.

[0018] FIGs. 4A-4D. Table of exemplary intein insertion sites in COL4A3 and their characteristics. In the column labeled “Structure”: SS, signal sequence; CD, collagenous domain; Int, interrupting non- collagenous region; NCI, non-collagenous C-terminal domain. In the column labeled “Fit in AAV?”, “Yes” indicates the construct can fit in AAV when the CAG promoter-driven expression cassettes are employed; “Possible” indicates the construct can fit in AAV when a combination of a shorter enhancer-promoter and a shorter polyadenylation signal the combined size of which is approximately 0.2 kb or less are employed.

[0019] FIGs. 5A-5D. Table of exemplary intein insertion sites in COL4A4 and their characteristics. In the column labeled “Structure”: SS, signal sequence; CD, collagenous domain; Int, interrupting non- collagenous region; NCI, non-collagenous C-terminal domain. In the column labeled “Fit in AAV?”, “Yes” indicates the construct can fit in AAV when the CAG promoter-driven expression cassettes are employed; “Possible” indicates the construct can fit in AAV when a combination of a shorter enhancer-promoter and a shorter polyadenylation signal the combined size of which is approximately 0.2 kb or less are employed.

[0020] FIGs. 6A-6E. FIG. 6A: A map of the split intein COL4A5 constructs. Out of 27 split points, 12 split sites are indicated. The amino acid positions show +1 amino acid residues at the COL4A5 C-half in the split intein Col4A5 constructs. H53 indicates the position of the epitope of rat monoclonal anti-COL4A5 antibody (H53 clone) as used herein. The generic structure of the full-length COL4A5 construct (No. 1), split intein COL4A5 N-half constructs (No. 2) and split intein C-half constructs (Nos. 2-6) are shown. Asterisks (*) indicate the location of woodchuck hepatitis virus post-transcriptional regulatory element (WPRE) or woodchuck hepatitis virus post-transcriptional regulatory element 3 (WPRE3) insertion for the constructs that harbor WPRE or WPRE3. FIG. 6B: Role of a signal sequence in effective expression of split intein COL4A5 C-half constructs. HEK293 cells were transfected with plasmid DNA expressing natives S-Cint- Col4A5-FLAG (no.3 in Fig. 6A, split point 1), chymoSS-Cint-Col4A5-FLAG (no.4 in Fig. 6A, split point 1), or noSS-Cint-Col4A5-FLAG (no.5 in Fig. 6A, split point 1) and the expression levels of COL4A5 C-half proteins in cells were assessed by western blot with anti-FLAG antibodies. GAPDH was employed as a loading control. FIG. 6C: Efficacy of trans- splicing using split intein C-half constructs with a different signal sequence. HEK293 cells were transfected with plasmid DNA expressing the construct(s) as indicated above each lane. 1, full-length COL4A5-FLAG (no.l in FIG. 6A, split point 1); 2, Col4A5-Nint-FLAG (no.2 in FIG. 6A, split point 1); 3, nativeSS-Cint-Col4A5-FLAG (no.3 in FIG. 6A, split point 1); 4, chymoSS-Cint- Col4A5-FLAG (no.4 in FIG. 6A, split point 1); and 5, noSS-Cint-Col4A5-FLAG (no.5 in FIG. 6A, split point 1). Crude cell lysates were subjected to western blot analysis with anti-FLAG antibody. GAPDH was employed as a loading control. FIG. 6D: Split intein COL4A5 protein trans-splicing in HEK293 cells. HEK293 cells were transfected with plasmid DNA expressing the construct(s) as indicated above each lane. 1, full-length Col4A5-FLAG (no.l in FIG. 6A, split point 2); 2, nativeSS-Cint-Col4A5-FLAG (no.3 in FIG. 6A, split point 2); and 3, Col4A5-Nint-FLAG (no.2 in FIG. 6A, split point 2). Crude cell lysates were subjected to western blot analysis with anti-FLAG antibody. pAAV-CAG-FLAG-Empty plasmid was used for the negative control (Neg). FIG. 6E: Split intein dual AAV-Col4A5 vector-transduced human podocytes can secrete trans-spliced full-length COL4A5 protein. AAV-KP1 vectors indicated in the panel were produced, purified, and titered. Terminally differentiated human podocytes were infected with each AAV- KP1 vector at a multiplicity of infection (MOI) of I x 10\ Western blot analysis was performed to detect AAV vector-derived COL4A5 proteins produced in AAV vector-transduced podocytes and secreted into the culture media in a 48-hour period. Anti-HA antibody was used to differentiate endogenous COL4A5 proteins and AAV vector-derived HA-tagged COL4A5 proteins.

[0021] FIG. 7. Table showing the COL4A5 split points and the lengths of split intein dual AAV-CAG- Col4A5 N-half and C-half vector genomes.

[0022] FIGs. 8A-8D. Table of split intein dual AAV-CAG-Col4A5 vector plasmid constructs, n.a. indicates not applicable.

[0023] FIGs. 9A-9F. Functional validation of trans-splicing split point candidates in COL4A5. HEK293 cells were transfected with a plasmid DNA or a combination of plasmid DNAs expressing COL4A5 as listed in FIGs. 9E-9F. Forty-eight hours after transfection, cells were harvested and protein trans-splicing was assessed by western blot analysis with anti-FLAG and anti-human COL4A5 antibodies. Each panel represents independent experiments in all of which the SP2 constructs, the full-length control (ctrl), and an empty plasmid negative control were included. Samples indicated with N+C received both split intein COL4A5 N-half and C-half plasmid constructs. Lanes 4 and 10 samples received both N-half and C-half constructs, but the C-half constructs were devoid of a signal sequence. Lane 23 sample received both N-half and C-half constructs that were expected to mediate trans-splicing; however, the full-length COL4A5 constructs were devoid of a FLAG tag; and therefore, they were not detected with anti-FLAG antibody. The position of the full-length COL4A5 is indicated in each panel.

[0024] FIG. 9G. Protein trans-splicing (PTS) efficiencies of 11 viable split points determined by the method used in FIGs. 9A-9F.

[0025] FIG. 9H. COL4A5 amino acid motifs at the protein trans-splicing (PTS) sites, determined by the method used in FIGs. 9A-9F.

[0026] FIG. 10. Effects of split intein COL4A5 N-half and C-half construct ratios on protein trans-splicing. HEK293 cells were transfected with various ratios of split intein COL4A5 N-half and C-half plasmid constructs. Forty-eight hours after transfection, cells were harvested and protein trans-splicing was assessed by western blot analysis with anti-FLAG and anti-human COL4A5 antibodies. pAAV-CAG-Col4A5-FLAG- WPRE and pAAV-CAG-GFP plasmids were used for the full-length control (ctrl) and a negative control (ctrl), respectively. The position of the full-length COL4A5 is shown.

[0027] FIGs. 11A-11B. Split point (SP) effects on the split intein COL4A5 N-half and C-half constructs in HEK293 cells. HEK293 cells were transfected with plasmid DNA expressing split intein COL4A5 N-half constructs (FIG. 1 1 A) or C-half constructs (FIG. 1 1 B). The constructs expressed in each sample are indicated above the panels. Forty-eight hours after transfection, cells were harvested and steady-state levels of each construct were assessed by western blot analysis using anti-FLAG antibody. The bands near the bottom of Panel B are non-specific bands.

[0028] FIGs. 12A-12E. FIG. 12A: A genomic map of the LSL-Col4a5 knock-in allele. FIG. 12 A: A 1.1 -kb LSL cassette carrying 3 copies of SV40 polyadenylation signal floxed with 2 loxP sequences was inserted between mouse chromosome X positions 141,510,668 and 141,510,669 (mmlO). FIGs. 12B-12E: LSL- Col4a5 mouse phenotypes. Both hemizygous males (i.e., KO) and wild-types males were monitored for body weights (FIG. 12B), proteinuria (FIG. 12C), survival (FIG. 12D) and blood biomarkers (FIG. 12E), manifesting progressive body weight loss and chronic kidney disease, ultimately leading to end stage kidney disease. All affected animals reached death or the study end point by 34 weeks of age. Two-way ANOVA with repeated measures (FIGs. 12B-12C) and log-rank test (FIG. 12D) were used. ****, p <0.0001. Error bars represent SEM. Blood urea nitrogen (BUN). Hematocrit (het). Packed cell volume (PCV). Hemoglobin (Hb).

[0029] FIGs. 13A-13B. Enhancement of podocyte transductions by AAV9 IV, but not AAV-KP1 IV injection, in CKD1. FIG. 13A: A representative glomerulus in the kidneys of mice injected with each AAV vector expressing tdTomato (tdT). Podocyte transduction was assessed with anti-WTl antibody staining. Scale bar, 100 pm. FIG. 13B: Quantitative assessment of podocyte transduction efficiencies by a manual counting of both tdT+ and tdT- podocytes labeled with WT1. ****, p<0.0001; *, p<0.05; ns, not significant. Error bars represent SEM. A two-way ANOVA followed by Tukey’s post hoc test was used. P values are adjusted p values. FIG. 14A-14H. AAV9 IV gene therapy transduces podocytes and restores type IV collagen expression in the glomeruli in LSL Col4a5 XLAS mouse model. Three 17-week-old LSL Col4a5 hemizygous (He) male mice were treated with 6.8 x 10” vg of AAV9-CAG-Cre via IV administration. FIG. 14A: One AAV vector-treated mouse (He+AAV) was assessed for the expression of Col4a4 and Col4a5 proteins in the kidney 10 weeks post-injection using anti-COL4A4 (clone b42) and anti-COL4A5 (clone H53) antibodies. He and wild-type (WT) mice serve as negative and positive controls, respectively. * and ** indicate serial sections. Examples of the glomeruli that can be seen in both sections are indicated with arrowheads. Scale bar, 100 pm. Significantly more COL4A4 and COL4A5 expression is apparent in treated mice, relative to He controls, particularly in the glomeruli. FIG. 14B: The same glomerulus in the AAV vector-treated He mouse was stained with the indicated antibodies. Overlapping staining patterns for COL4A4 and COL4A5 are apparent. Scale bar, 50 Ltm. FIG. 14C: AAV9 IV gene therapy can mediate therapeutic effects in LSL Col4a5 XLAS mouse model. Three LSL Col4a5 hemizygous (He) male mice treated by IV injection of AAV9-CAG-Cre vector at a dose of 6.8 x 1011vg at the age of 17 weeks were monitored for their body weights up to 19 weeks post-injection (36 weeks of age). One mouse was euthanized for histological assessment 10 weeks post-injection. The body weights were normalized with their body weights at the time of injection (z.<?., 17 weeks of age). FIG. 14D: LSL-Col4a5 hemizygous (He) male mice (n=l 1) were treated with 3.0 x 10” vg of AAV9-CAG-Cre vector by retro-orbital injection at 4 weeks of age and monitored for their body weights up to 56 weeks of age. FIG. 14E: The AAV vector- treated mice shown in Panel D were monitored for albuminuria up to 17 weeks of age. The data equivocally indicated that AAV vector treatment effectively prevented albumin leakage into the urine. FIG. 14F: Blood biomarkers were measured in the AAV vector-treated mice (shown in FIG. 14D) at 31 weeks of age. FIG. 14G: Survival analysis of LSL-Col4a5 XLAS mice treated or untreated with AAV9 IV gene therapy at 4 weeks of age. The data indicated that the therapy increased the median survival time from 31.5 to 50.0 weeks. FIG. 14H: H&E-stained kidneys collected from mice at 31 to 32 weeks of age. This demonstrates that AAV9 IV gene therapy can mediate therapeutic effects in XLAS mice, leading to biomarker and survival improvement when injected at an earlier age. ****, p<0.0001. Two-way ANOVA with repeated measures was used. Error bars represent SEM.

[0030] FIGs. 15A-15C. Representative images of the glomeruli of Col4a5 G5X XLAS mice intravenously injected with split intein dual AAV9 vectors expressing COL4A5. Two 26-week-old Col4a5 G5X XLAS hemizygous male mice were intravenously injected with a 1 : 1 mixture of split intein dual AAV-CAG- Col4A5 N-half and C-half vectors. Three weeks after injection, the kidney tissue was stained with the antibodies indicated in each panel. HA, anti-HA antibody; LAMB2, anti-laminin f32 / yl antibody (a Glomerular Basement Membrane (GBM) marker); WT1, anti-WTl antibody (a podocyte marker); and H53, anti-COL4A5 antibody. Anti-COL4A5 H53 antibody detects the COL4A5 N-half while anti-HA antibody detects the C-half. FIG. 15A: An untreated mouse, no HA staining is apparent. FIGs. 15B-15C: Split intein dual AAV9 vector-treated mice. Arrowheads indicate COL4A5 -positive podocytes. Scale bar, 50 Ltm. HA staining pattern overlaps with LAMB2 staining patterns in FIG. 15B. HA staining pattern overlaps with H53 staining pattern in FIG. 15C.

[0031] FIGs. 16A-16D. Assessment of the percentage of Col4a5 deposition in the GBM (% deposition) in LSL-Col4a5 XLAS mice treated with IV injection of AAV9-CAG-Cre. Kidneys were harvested from two 31-week-old LSL-Col4a5 hemizygous male mice (#1 and #2) treated with IV injection of 3x 10' vg of AAV9-CAG-Cre at 4 weeks of age. These two showed therapeutic effects. FIG. 16A: Kidney sections were stained with anti-rat COL4A5 antibody (H53) and anti-rabbit agrin antibody, and subjected to confocal microscopy. Approximately 10 Col4a5-positive glomeruli were randomly selected and analyzed for the lengths of the membranous structures positive for agrin and Col4a5 inside the Bowman's capsule, which represent the total length of the GBM and Col4a5-deposited GBM, respectively. Using Image J software, the GBMs in each panel were traced with thin lines and the total length of the lines in each panel was determined. Scale bar, 20 Ltm. FIG. 16B: Percentage of Col4a5 deposition in the GBM in each glomerulus was calculated by the formula, (Col4a5 length) / (agrin lengthjx 100. A similar length-based approach was previously used for quantifying the degree of the foot process effacement observed in podocytopathy (Derewicz et al. Archives of the Balkan Medical Union 54:532-539, 2019). Error bars represent SD. FIG. 16C: The overall % deposition was determined based on the average % deposition in each Col4a5 -positive glomerulus and % of Col4a5 -positive glomeruli among the total glomeruli assessed. FIG. 16D: GBM positive for both AAV-derived COL4A5 and endogenous Col4a4 (indicated by arrows) in Col4a5 G5X XLAS mice treated with IV injection of a total of 4 x 1013vg of AAV9-CAG-COL4A5-Nint-F(SP2) and AAV9-CAG-COL4A5-Cint-H(SP2) mixed at a 1:1 ratio. Mice were treated at 18 weeks of age and euthanized at 32 weeks of age. COL4A5 reconstructed by PTS and endogenous Col4a4 were assessed with anti-HA and anti-COL4A4 (B42) antibodies. Overlap between Col4a4 and HA (COL4A5) staining patterns is apparent in merged view.

[0032] SEQUENCES

[0033] The nucleic and amino acid sequences listed in the accompanying sequence listing are shown using standard letter abbreviations for nucleotide bases, and single letter code for amino acids, as defined in 37 C.F.R. 1.822. Only one strand of each nucleic acid sequence is shown, but the complementary strand is understood as included by any reference to the displayed strand.

[0034] In the accompanying sequence listing:

[0035] SEQ ID NO: 1 is an exemplary CAG-WPRE-SV40pA expression cassette.

[0036] SEQ ID NO: 2 is an exemplary CAG-WPRE3-SV40pA expression cassette.

[0037] SEQ ID NO: 3 is an exemplary CAG-SV40pA expression cassette.

[0038] SEQ ID NO: 4 is an exemplary collagen signal sequence.

[0039] SEQ ID NO: 5 is an exemplary chymotrypsin signal sequence.

[0040] SEQ ID NO: 6 is an exemplary 22-amino-acid FLAG tag.

[0041] SEQ ID NO: 7 is an exemplary human influenza virus hemagglutinin tag. SEQ ID NO: 8 is an exemplary sgRNA, Col4a5_crRNAl.

[0042] SEQ ID NO: 9 is an exemplary sgRNA, Col4a5_crRNA2

[0043] SEQ ID NO: 10 is an exemplary sgRNA, Col4a5_crRNA3

[0044] SEQ ID NO: 11 is an exemplary forward primer originating from the LHA.

[0045] SEQ ID NO: 12 is an exemplary reverse primer originating from the KI region.

[0046] SEQ ID NO: 13 is an exemplary COL4A3 amino acid sequence.

[0047] SEQ ID NO: 14 is an exemplary COL4A4 amino acid sequence.

[0048] SEQ ID NO: 15 is an exemplary COL4A5 amino acid sequence.

[0049] SEQ ID NO: 16 is an exemplary COL4A5 amino acid sequence, NP_203699.1.

[0050] SEQ ID NO: 17 is an exemplary COL4A5 amino acid sequence, AAI51847.1.

[0051] SEQ ID NO: 18 is an exemplary COL4A5 amino acid sequence, AIW39922.1.

[0052] SEQ ID NO: 19 is an exemplary COL4A5 amino acid sequence, XP_016884749.1.

[0053] SEQ ID NO: 20 is an exemplary COL4A5 amino acid sequence, XP_016884748.1.

[0054] SEQ ID NO: 21 is an exemplary COL4A5 amino acid sequence, XP_011529151.2.

[0055] SEQ ID NO: 22 is an exemplary COL4A5 amino acid sequence, XP_047297766.1.

[0056] SEQ ID NO: 23 is an exemplary COL4A5 amino acid sequence, XP_016884750.1.

[0057] SEQ ID NO: 24 is an exemplary COL4A5 amino acid sequence, XP 047297767.1 .

[0058] SEQ ID NO: 25 is an exemplary COL4A5 amino acid sequence, XP 016884751.1.

[0059] SEQ ID NO: 26 is an exemplary N-intein of a Npu DnaE intein.

[0060] SEQ ID NO: 27 is an exemplary C-intein of a Npu DnaE intein.

[0061] SEQ ID NO: 28 is an exemplary nucleotide sequence encoding the open reading frame (ORF) of pAAV-CAG-Col4A5-FLAG-WPRE.

[0062] SEQ ID NO: 29 is an exemplary amino acid sequence of the ORF of pAAV-CAG-Col4A5-FLAG- WPRE.

[0063] SEQ ID NO: 30 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-SPl- Col4A5-Nint-FEAG-WPRE.

[0064] SEQ ID NO: 31 is an exemplary amino acid sequence of the ORF of pAAV-CAG-SPl-Col4A5- Nint-FLAG-WPRE

[0065] SEQ ID NO: 32 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-SP2- Col4A5-Nint-FEAG-noWPRE

[0066] SEQ ID NO: 33 is an exemplary amino acid sequence of the ORF of pAAV-CAG-SP2-Col4A5- Nint-FLAG-noWPRE

[0067] SEQ ID NO: 34 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-SP2- Col4A5-Nint-FEAG-WPRE3

[0068] SEQ ID NO: 35 is an exemplary amino acid sequence of the ORF of pAAV-CAG-SP2-Col4A5- Nint-FLAG-WPRE3 SEQ ID NO: 36 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-SP2- Col4A5-Nint-FLAG-WPRE

[0069] SEQ ID NO: 37 i: 5 an exemplary amino acid sequence of the ORF of pAAV-CAG-SP2-Col4A5- Nint-FLAG-WPRE

[0070] SEQ ID NO: 38 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-SP3- Col4A5-Nint-FLAG-WPRE

[0071] SEQ ID NO: 39 is an exemplary amino acid sequence of the ORF of pAAV-CAG-SP3-Col4A5- Nint-FLAG-WPRE

[0072] SEQ ID NO: 40 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-SP3- Col4A5-Nint-FLAG-WPRE3

[0073] SEQ ID NO: 41 is an exemplary amino acid sequence of the ORF of pAAV-CAG-SP3-Col4A5- Nint-FLAG-WPRE3

[0074] SEQ ID NO: 42 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-SP4-

[0075] Col4A5-Nint-FLAG-WPRE

[0076] SEQ ID NO: 43 is an exemplary amino acid sequence of the ORF of pAAV-CAG-SP4-Col4A5- Nint-FLAG-WPRE

[0077] SEQ ID NO: 44 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-SP4-

[0078] Col4A5-Nint-FLAG-WPRE3

[0079] SEQ ID NO: 45 is an exemplary amino acid sequence of the ORF of pAAV-CAG-SP4-Col4A5- Nint-FLAG-WPRE3

[0080] SEQ ID NO: 46 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-SP5-

[0081] Col4A5-Nint-FLAG-WPRE

[0082] SEQ ID NO: 47 is an exemplary amino acid sequence of the ORF of pAAV-CAG-SP5-Col4A5- Nint-FLAG-WPRE

[0083] SEQ ID NO: 48 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-SP5- Col4A5-Nint-FLAG-WPRE3

[0084] SEQ ID NO: 49 an exemplary amino acid sequence of the ORF of pAAV-CAG-SP5-Col4A5-Nint- FLAG-WPRE3

[0085] SEQ ID NO: 50 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-SP7- Col4A5-Nint-FLAG-WPRE

[0086] SEQ ID NO: 51 is an exemplary amino acid sequence of the ORF of pAAV-CAG-SP7-Col4A5- Nint-FLAG-WPRE

[0087] SEQ ID NO: 52 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-SP7- Col4A5-Nint-FLAG-noWPRE

[0088] SEQ ID NO: 53 is an exemplary amino acid sequence of the ORF of pAAV-CAG-SP7-Col4A5- Nint-FLAG-noWPRE SEQ ID NO: 54 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-SP9- Col4A5-Nint-FLAG-WPRE

[0089] SEQ ID NO: 55 i: 5 an exemplary amino acid sequence of the ORF of pAAV-CAG-SP9-Col4A5- Nint-FLAG-WPRE

[0090] SEQ ID NO: 56 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-SP9- Col4A5-Nint-FLAG-noWPRE

[0091] SEQ ID NO: 57 is an exemplary amino acid sequence of the ORF of pAAV-CAG-SP9-Col4A5- Nint-FLAG-noWPRE

[0092] SEQ ID NO: 58 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-SPl 1- Col4A5-Nint-FLAG-WPRE

[0093] SEQ ID NO: 59 is an exemplary amino acid sequence of the ORF of pAAV-CAG-SPll-Co!4A5- Nint-FLAG-WPRE

[0094] SEQ ID NO: 60 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-SP12-

[0095] Col4A5-Nint-FLAG-noWPRE

[0096] SEQ ID NO: 61 is an exemplary amino acid sequence of the ORF of pAAV-CAG-SP12-Col4A5- Nint-FLAG-noWPRE

[0097] SEQ ID NO: 62 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-noSS- Cint-SP 1 -Col4 A5-FLAG- WPRE

[0098] SEQ ID NO: 63 is an exemplary amino acid sequence of the ORF of pAAV-CAG-noSS-Cint-SPl-

[0099] Col4A5-FLAG-WPRE

[0100] SEQ ID NO: 64 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-nativeSS- Cint-SP 1 -Col4 A5-FLAG- WPRE

[0101] SEQ ID NO: 65 is an exemplary amino acid sequence of the ORF of pAAV-CAG-nativeSS-Cint- SPl-Col4A5-FLAG-WPRE

[0102] SEQ ID NO: 66 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-chymoSS- Cint-SP 1 -Col4 A5-FLAG- WPRE

[0103] SEQ ID NO: 67 is an exemplary amino acid sequence of the ORF of pAAV-CAG-chymoSS-Cint- SPl-Col4A5-FLAG-WPRE

[0104] SEQ ID NO: 68 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-nativeSS- Cint-SP2-Col4A5-HA-noWPRE

[0105] SEQ ID NO: 69 is an exemplary amino acid sequence of the ORF of pAAV-CAG-nativeSS-Cint- SP2-Col4A5-HA-noWPRE

[0106] SEQ ID NO: 70 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-nativeSS- Cint-SP2-Col4A5-FLAG-WPRE3

[0107] SEQ ID NO: 71 is an exemplary amino acid sequence of the ORF of pAAV-CAG-nativeSS-Cint- SP2-CO14A5-FLAG-WPRE3 SEQ ID NO: 72 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-nativeSS- Cint-SP2-Col4A5-FLAG-noWPRE

[0108] SEQ ID NO: 73 is an exemplary amino acid sequence of the ORF of pAAV-CAG-nativeSS-Cint- SP2-Col4A5-FLAG-noWPRE

[0109] SEQ ID NO: 74 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-chymoSS- Cint-SP2-Col4A5-FLAG-WPRE3

[0110] SEQ ID NO: 75 is an exemplary amino acid sequence of the ORF of pAAV-CAG-chymoSS-Cint- SP2-CO14A5-FLAG-WPRE3

[0111] SEQ ID NO: 76 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-nativeSS- Cint-SP2-Col4A5-FLAG-WPRE

[0112] SEQ ID NO: 77 is an exemplary amino acid sequence of the ORF of pAAV-CAG-nativeSS-Cint- SP2-CO14A5-FLAG-WPRE

[0113] SEQ ID NO: 78 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-chymoSS- Cint-SP2-Col4A5-FLAG-noWPRE

[0114] SEQ ID NO: 79 is an exemplary amino acid sequence of the ORF of pAAV-CAG-chymoSS-Cint- SP2-Col4A5-FLAG-noWPRE

[0115] SEQ ID NO: 80 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-noSS- Cint-SP2-Col4A5-FLAG-WPRE

[0116] SEQ ID NO: 81 is an exemplary amino acid sequence of the ORF of pAAV-CAG-noSS-Cint-SP2-

[0117] Col4A5-FLAG-WPRE

[0118] SEQ ID NO: 82 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-chymoSS- Cint-SP2-Col4A5-FLAG-WPRE

[0119] SEQ ID NO: 83 is an exemplary amino acid sequence of the ORF of pAAV-CAG-chymoSS-Cint- SP2-Col4A5-FLAG-WPRE

[0120] SEQ ID NO: 84 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-nativeSS- Cint-SP3-Col4A5-FLAG-WPRE

[0121] SEQ ID NO: 85 is an exemplary amino acid sequence of the ORF of pAAV-CAG-nativeSS-Cint- SP3-Col4A5-FLAG-WPRE

[0122] SEQ ID NO: 86 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-nativeSS- Cint-SP3-Col4A5-FLAG-WPRE3

[0123] SEQ ID NO: 87 is an exemplary amino acid sequence of the ORF of pAAV-CAG-nativeSS-Cint- SP3-CO14A5-FLAG-WPRE3

[0124] SEQ ID NO: 88 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-chymoSS- Cint-SP3-Col4A5-FLAG-WPRE3

[0125] SEQ ID NO: 89 is an exemplary amino acid sequence of the ORF of pAAV-CAG-chymoSS-Cint- SP3-CO14A5-FLAG-WPRE3 SEQ ID NO: 90 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-nativeSS- Cint-SP4-Col4A5-FLAG-WPRE

[0126] SEQ ID NO: 91 is an exemplary amino acid sequence of the ORF of pAAV-CAG-nativeSS-Cint- SP4-Col4A5-FLAG-WPRE

[0127] SEQ ID NO: 92 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-nativeSS- Cint-SP4-Col4A5-FLAG-WPRE3

[0128] SEQ ID NO: 93 is an exemplary amino acid sequence of the ORF of pAAV-CAG-nativeSS-Cint- SP4-CO14A5-FLAG-WPRE3

[0129] SEQ ID NO: 94 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-chymoSS- Cint-SP4-Col4A5-FLAG-WPRE3

[0130] SEQ ID NO: 95 is an exemplary amino acid sequence of the ORF of pAAV-CAG-chymoSS-Cint- SP4-CO14A5-FLAG-WPRE3

[0131] SEQ ID NO: 96 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-nativeSS- Cint-SP5-Col4A5-FLAG-WPRE

[0132] SEQ ID NO: 97 is an exemplary amino acid sequence of the ORF of pAAV-CAG-nativeSS-Cint- SP5-Col4A5-FLAG-WPRE

[0133] SEQ ID NO: 98 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-nativeSS- Cint-SP5-Col4A5-FLAG-WPRE3

[0134] SEQ ID NO: 99 is an exemplary amino acid sequence of the ORF of pAAV-CAG-nativeSS-Cint-

[0135] SP5-Col4A5-FLAG-WPRE3

[0136] SEQ ID NO: 100 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG- chymoSS-Cint-SP5-Col4A5-FLAG-WPRE3

[0137] SEQ ID NO: 101 is an exemplary amino acid sequence of the ORF of pAAV-CAG-chymoSS-Cint- SP5-Col4A5-FLAG-WPRE3

[0138] SEQ ID NO: 102 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-nativeSS- Cint-SP7-Col4A5-FLAG-noWPRE

[0139] SEQ ID NO: 103 is an exemplary amino acid sequence of the ORF of pAAV-CAG-nativeSS-Cint- SP7-Col4A5-FLAG-noWPRE

[0140] SEQ ID NO: 104 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-nativeSS- Cint-SP9-Col4A5-FLAG-noWPRE

[0141] SEQ ID NO: 105 is an exemplary amino acid sequence of the ORF of pAAV-CAG-nativeSS-Cint- SP9-Col4A5-FLAG-noWPRE

[0142] SEQ ID NO: 106 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-noSS- Cint-SPl 1-CO14A5-FLAG-WPRE

[0143] SEQ ID NO: 107 is an exemplary amino acid sequence of the ORF of pAAV-CAG-noSS-Cint-

[0144] SP11-C014A5-FLAG-WPRE SEQ ID NO: 108 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-nativeSS-

[0145] Cint-SPl 1-CO14A5-FLAG-WPRE

[0146] SEQ ID NO: 109 is an exemplary amino acid sequence of the ORF of pAAV-CAG-nativeSS-Cint-

[0147] SP1 l-Col4A5-FLAG-WPRE

[0148] SEQ ID NO: 110 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG- chymoSS-Cint-SPl 1 -Col4A5-FLAG-WPRE

[0149] SEQ ID NO: 111 is an exemplary amino acid sequence of the ORF of pAAV-CAG-chymoSS-Cint- SP11-CO14A5-FLAG-WPRE

[0150] SEQ ID NO: 112 is an exemplary nucleotide sequence encoding the ORF of pAAV-CAG-nativeSS- Cint-SP 12-Col4 A5-FL AG-noWPRE

[0151] SEQ ID NO: 113 is an exemplary amino acid sequence of the ORF of pAAV-CAG-nativeSS-Cint-

[0152] SP 12-Col4 A5-FL AG-noWPRE

[0153] SEQ ID NO: 114 i: s an exemplary 5’ aspect of SEQ ID NO: 1. It can exemplify a region of SEQ ID

[0154] NO: 1 5’ of an ORF.

[0155] SEQ ID NO: 115 i: s an exemplary 3’ aspect of SEQ ID NO: 1. It can exemplify a region of SEQ ID

[0156] NO: 1 3’ of an ORF.

[0157] SEQ ID NO: 116 is an exemplary 5’ aspect of SEQ ID NO: 2. It can exemplify a region of SEQ ID NO: 2 5’ of an ORF.

[0158] SEQ ID NO: 117 is an exemplary 3’ aspect of SEQ ID NO: 2. It can exemplify a region of SEQ ID NO: 2 3’ of an ORF.

[0159] SEQ ID NO: 118 is an exemplary 5’ aspect of SEQ ID NO: 3. It can exemplify a region of SEQ ID NO: 3 5’ of an ORF.

[0160] SEQ ID NO: 119 is an exemplary 3’ aspect of SEQ ID NO: 3. It can exemplify a region of SEQ ID NO: 3 3’ of an ORF.

[0161] SEQ ID NO: 120 is an exemplary nucleotide sequence encoding WPRE.

[0162] SEQ ID NO: 121 is an exemplary nucleotide sequence encoding WPRE3.

[0163] SEQ ID NO: 122 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG-

[0164] COL4A5-Nintl4-3xFLAG-WPRE.

[0165] SEQ ID NO: 123 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nintl4-3xFLAG-WPRE.

[0166] SEQ ID NO: 124 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- COL4A5-Nintl5-3xFLAG-WPRE.

[0167] SEQ ID NO: 125 i: s an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nintl5-3xFLAG-WPRE.

[0168] SEQ ID NO: 126 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- COL4A5-Nintl6-3xFLAG-WPRE. SEQ ID NO: 127 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nintl6-3xFLAG-WPRE.

[0169] SEQ ID NO: 128 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- COL4A5-Nintl7-3xFLAG-WPRE.

[0170] SEQ ID NO: 129 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nintl7-3xFLAG-WPRE.

[0171] SEQ ID NO: 130 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- COL4A5-Nintl8-3xFLAG-WPRE.

[0172] SEQ ID NO: 131 i: s an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nintl 8-3xFLAG-WPRE.

[0173] SEQ ID NO: 132 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- COL4A5-Nintl9-3xFLAG-WPRE.

[0174] SEQ ID NO: 133 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nintl9-3xFLAG-WPRE

[0175] SEQ ID NO: 134 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG-

[0176] COL4A5-Nint20-3xFLAG-WPRE

[0177] SEQ ID NO: 135 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nint20-3xFLAG-WPRE.

[0178] SEQ ID NO: 136 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- COL4 A5-Nint21 -3xFLAG-WPRE.

[0179] SEQ ID NO: 137 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nint21 -3xFLAG-WPRE.

[0180] SEQ ID NO: 138 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- COL4A5-Nint22-3xFLAG-WPRE.

[0181] SEQ ID NO: 139 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nint22-3xFLAG-WPRE.

[0182] SEQ ID NO: 140 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- COL4A5-Nint23-3xFLAG-WPRE.

[0183] SEQ ID NO: 141 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nint23-3xFLAG-WPRE.

[0184] SEQ ID NO: 142 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- COL4A5-Nint24-3xFLAG-WPRE.

[0185] SEQ ID NO: 143 i: s an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nint24-3xFLAG-WPRE.

[0186] SEQ ID NO: 144 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- COL4A5-Nint25-3xFLAG-WPRE. SEQ ID NO: 145 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nint25-3xFLAG-WPRE.

[0187] SEQ ID NO: 146 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- COL4A5-Nint26-3xFLAG-WPRE.

[0188] SEQ ID NO: 147 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nint26-3xFLAG-WPRE.

[0189] SEQ ID NO: 148 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- COL4A5-Nint27-3xFLAG-WPRE.

[0190] SEQ ID NO: 149 i: s an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nint27-3xFLAG-WPRE.

[0191] SEQ ID NO: 150 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- COL4A5-Nint28-3xFLAG-WPRE.

[0192] SEQ ID NO: 151 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nint28-3xFLAG-WPRE.

[0193] SEQ ID NO: 152 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- COL4A5-Nint6-3xFLAG-WPRE.

[0194] SEQ ID NO: 153 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nint6-3xFLAG-WPRE.

[0195] SEQ ID NO: 154 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- COL4A5-Nint8-3xFLAG-WPRE.

[0196] SEQ ID NO: 155 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nint8-3xFLAG-WPRE.

[0197] SEQ ID NO: 156 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- COL4A5-NintlO-3xFLAG-WPRE.

[0198] SEQ ID NO: 157 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-COL4A5- Nintl0-3xFLAG-WPRE.

[0199] SEQ ID NO: 158 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cintl4-COL4A5-HA-WPRE.

[0200] SEQ ID NO: 159 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cintl4-COL4A5-HA-WPRE.

[0201] SEQ ID NO: 160 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cintl5-COL4A5-HA-WPRE.

[0202] SEQ ID NO: 161 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cintl5-COL4A5-HA-WPRE.

[0203] SEQ ID NO: 162 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cintl6-COL4A5-HA-WPRE. SEQ ID NO: 163 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cintl6-COL4A5-HA-WPRE.

[0204] SEQ ID NO: 164 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cintl7-COL4A5-HA-WPRE.

[0205] SEQ ID NO: 165 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cintl7-COL4A5-HA-WPRE.

[0206] SEQ ID NO: 166 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cintl8-COL4A5-HA-WPRE.

[0207] SEQ ID NO: 167 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cintl8-COL4A5-HA-WPRE.

[0208] SEQ ID NO: 168 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cintl9-COL4A5-HA-WPRE.

[0209] SEQ ID NO: 169 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cintl9-COL4A5-HA-WPRE.

[0210] SEQ ID NO: 170 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cint20-COL4A5-HA-WPRE.

[0211] SEQ ID NO: 171 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cint20-COL4A5-HA-WPRE.

[0212] SEQ ID NO: 172 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cint21 -COL4A5-HA-WPRE.

[0213] SEQ ID NO: 173 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cint21 -COL4A5-HA-WPRE.

[0214] SEQ ID NO: 174 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cint22-COL4A5-HA-WPRE.

[0215] SEQ ID NO: 175 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cint22-COL4A5-HA-WPRE.

[0216] SEQ ID NO: 176 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cint23-COL4A5-HA-WPRE.

[0217] SEQ ID NO: 177 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cint23-COL4A5-HA-WPRE.

[0218] SEQ ID NO: 178 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cint24-COL4A5-HA-WPRE.

[0219] SEQ ID NO: 179 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cint24-COL4A5-HA-WPRE.

[0220] SEQ ID NO: 180 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cint25-COL4A5-HA-WPRE. SEQ ID NO: 181 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cint25-COL4A5-HA-WPRE.

[0221] SEQ ID NO: 182 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cint26-COL4A5-HA-WPRE.

[0222] SEQ ID NO: 183 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cint26-COL4A5-HA-WPRE.

[0223] SEQ ID NO: 184 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cint27-COL4A5-HA-WPRE.

[0224] SEQ ID NO: 185 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cint27-COL4A5-HA-WPRE.

[0225] SEQ ID NO: 186 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cint28-COL4A5-HA-WPRE.

[0226] SEQ ID NO: 187 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cint28-COL4A5-HA-WPRE.

[0227] SEQ ID NO: 188 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cint6-COL4A5-HA-WPRE.

[0228] SEQ ID NO: 189 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cint6-COL4A5-HA-WPRE.

[0229] SEQ ID NO: 190 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-Cint8-COL4A5-HA-WPRE.

[0230] SEQ ID NO: 191 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- Cint8-COL4A5-HA-WPRE.

[0231] SEQ ID NO: 192 is an exemplary nucleotide sequence encoding the ORF of pAAV-tCAG- nativeSP-CintlO-COL4A5-HA-WPRE.

[0232] SEQ ID NO: 193 is an exemplary amino acid sequence of the ORF of pAAV-tCAG-nativeSP- CintlO-COL4A5-HA-WPRE.

[0233] SEQ ID NO: 194 is an exemplary amino acid sequence used as negative SVM training data, corresponding to SP1.

[0234] SEQ ID NO: 195 is an exemplary amino acid sequence used as negative SVM training data, corresponding to SP14.

[0235] SEQ ID NO: 196 is an exemplary amino acid sequence used as negative SVM training data, corresponding to SP24.

[0236] SEQ ID NO: 197 is an exemplary amino acid sequence used as negative SVM training data, corresponding to SP25.

[0237] SEQ ID NO: 198 is an exemplary amino acid sequence used as negative SVM training data, corresponding to SP26. SEQ ID NO: 199 is an exemplary amino acid sequence used as negative SVM training data, corresponding to SP4.

[0238] SEQ ID NO: 200 is an exemplary amino acid sequence used as negative SVM training data, corresponding to SP15.

[0239] SEQ ID NO: 201 is an exemplary amino acid sequence used as negative SVM training data, corresponding to SP16.

[0240] SEQ ID NO: 202 is an exemplary amino acid sequence used as negative SVM training data, corresponding to SP17.

[0241] SEQ ID NO: 203 is an exemplary amino acid sequence used as negative SVM training data, corresponding to SP18.

[0242] SEQ ID NO: 204 is an exemplary amino acid sequence used as negative SVM training data, corresponding to SP19.

[0243] SEQ ID NO: 205 is an exemplary amino acid sequence used as negative SVM training data, corresponding to SP21.

[0244] SEQ ID NO: 206 is an exemplary amino acid sequence used as negative SVM training data, corresponding to SP22.

[0245] SEQ ID NO: 207 is an exemplary amino acid sequence used as negative SVM training data, corresponding to SP23.

[0246] SEQ ID NO: 208 is an exemplary amino acid sequence used as negative SVM training data, corresponding to SP28.

[0247] SEQ ID NO: 209 is an exemplary amino acid sequence used as negative SVM training data, corresponding to SP3.

[0248] SEQ ID NO: 210 is an exemplary amino acid sequence used as positive SVM training data, corresponding to SP10.

[0249] SEQ ID NO: 211 is an exemplary amino acid sequence used as positive SVM training data, corresponding to SP11.

[0250] SEQ ID NO: 212 is an exemplary amino acid sequence used as positive SVM training data, corresponding to SP12.

[0251] SEQ ID NO: 213 is an exemplary amino acid sequence used as positive SVM training data, corresponding to SP2.

[0252] SEQ ID NO: 214 is an exemplary amino acid sequence used as positive SVM training data, corresponding to SP20.

[0253] SEQ ID NO: 215 is an exemplary amino acid sequence used as positive SVM training data, corresponding to SP27.

[0254] SEQ ID NO: 216 is an exemplary amino acid sequence used as positive SVM training data, corresponding to SP5. SEQ ID NO: 217 is an exemplary amino acid sequence used as positive SVM training data, corresponding to SP6.

[0255] SEQ ID NO: 218 is an exemplary amino acid sequence used as positive SVM training data, corresponding to SP7.

[0256] SEQ ID NO: 219 is an exemplary amino acid sequence used as positive SVM training data, corresponding to SP8.

[0257] SEQ ID NO: 220 is an exemplary amino acid sequence used as positive SVM training data, corresponding to SP9.

[0258] SEQ ID NO: 221 is an exemplary C-terminal portion amino acid sequence.

[0259] SEQ ID NO: 222 is an exemplary C-terminal portion amino acid sequence.

[0260] SEQ ID NO: 223 is an exemplary C-terminal portion amino acid sequence.

[0261] SEQ ID NO: 224 is an exemplary C-terminal portion amino acid sequence.

[0262] SEQ ID NO: 225 is an exemplary C-terminal portion amino acid sequence.

[0263] SEQ ID NO: 226 is an exemplary C-terminal portion amino acid sequence.

[0264] SEQ ID NO: 227 is an exemplary C-terminal portion amino acid sequence.

[0265] SEQ ID NO: 228 is an exemplary C-terminal portion amino acid sequence.

[0266] SEQ ID NO: 229 is an exemplary C-terminal portion amino acid sequence.

[0267] SEQ ID NO: 230 is an exemplary C-terminal portion amino acid sequence.

[0268] SEQ ID NO: 231 is an exemplary C-terminal portion amino acid sequence.

[0269] Some of the amino acid sequences disclosed herein include an N-terminal methionine. In some examples, this methionine is optional. Thus, exemplary sequences are envisioned which lack an N-terminal methionine as well as sequences which include an N-terminal methionine.

[0270] DETAILED DESCRIPTION

[0271] I. Introduction

[0272] Alport syndrome can be caused by an inherited defect in the type IV collagen alpha chain, such as a mutation in the COL4A3, COL4A4, and / or COIAA5 gene. It is disclosed herein that complementation of the defective gene product can be used to treat Alport syndrome.

[0273] AAV vectors can be used to incorporate a complementary gene into a subject, however the payload of AAV vectors is limited to about 5.0kb, while the open reading frames of the type IV collagen alpha chains range from about 5.0-5. Ikb, making it infeasible to incorporate nucleic acid molecules encoding the full-length type IV collagen alpha chain ORF together with expression control sequences including a cis- regulatory element and a poly adenylation signal into a single AAV vector.

[0274] A split intein system can be used to express split gene products using trans-splicing, however protein trans-splicing using non-native exteins remains challenging. The process involves cutting a single protein into two distinct polypeptide sequences. This can cause issues with protein folding and stability. Further, cut site adjacent sequences which meet amino acid sequence criteria do not necessarily tolerate intein insertion. In addition, using a split intein system with secreted proteins could pose challenges because both the N-half and C-half proteins must be trafficked to the secretory pathway independently while undergoing successful trans-splicing.

[0275] These concerns are particularly pertinent to COL4A3, COL4A4, and COL4A5. It is disclosed herein that candidate split points fall within the collagenous domain of the type IV collagen alpha chain. This domain contains an evolutionary conserved G-X-Y motif, forming structurally conserved triple helices. The majority of the pathogenic missense mutations, which account for approximately 40% of COL4A5 mutations causing XLAS (Gubler et al., Nat Clin Pract Nephrol 4, 24-37 (2008)) involve the glycine residue in the G-X-Y motif in collagenous domain. Heterotrimerization takes place in the early stage of type IV collagen heterotrimer biogenesis in the endoplasmic reticulum (ER) before secretion. Thus, a collagenous domain split could have detrimental effects on protein stability, folding, heterotrimerization, and secretion.

[0276] However, a split intein system was produced, that includes a first nucleic acid molecule encoding a fusion protein including in N to C terminal order, an N-extein of an extein pair and an N-intein of an intein pair, where the N-extein includes a first signal sequence N-terminal to an N-terminal portion of a type IV collagen alpha chain. The split intein system also includes a second nucleic acid molecule encoding a fusion protein including in N to C terminal order, a second signal sequence, a C-intein of the intein pair, and a C- extein of the extein pair, where the C-extein includes a C-terminal portion of the type IV collagen alpha chain. In some examples, the N-terminal portion does not include a collagen signal sequence. It is disclosed herein that the order of the elements in the split intein system, and optionally the use of certain split points, enables production of full length COL4A3, COL4A4, and COL4A5 when the split intein system is introduced into mammalian cells.

[0277] II. Summary of Terms

[0278] Unless otherwise noted, technical terms are used according to conventional usage. Definitions of many common terms in molecular biology may be found in Krebs et al. (eds.), Lewin’s genes XII, published by Jones & Bartlett Learning, 2017. As used herein, the singular forms “a,” “an,” and “the,” refer to both the singular as well as plural, unless the context clearly indicates otherwise. For example, the term “a vector” includes singular or plural vectors and can be considered equivalent to the phrase “at least one vector.” As used herein, the term “comprises” means “includes.” The term “about” indicates within five percent. It is further to be understood that any and all base sizes or amino acid sizes, and all molecular weight or molecular mass values, given for nucleic acids or polypeptides are approximate, and are provided for descriptive purposes, unless otherwise indicated. Although many methods and materials similar or equivalent to those described herein can be used, particular suitable methods and materials are described herein. In case of conflict, the present specification, including explanations of terms, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting.

[0279] AAV Vector: Any vector that comprises or derives from components of AAV and is suitable to infect mammalian cells, including human cells, of any of a number of tissue types, such as brain, heart, lung, skeletal muscle, liver, kidney, spleen, or pancreas, whether in vitro or in vivo. The term “AAV vector” may be used to refer to an AAV type viral particle (or virion) comprising at least a nucleic acid molecule encoding a protein or nucleic acid molecule of interest.

[0280] Administration: Providing or giving a subject an agent by any effective route. Exemplary routes of administration include, but are not limited to, oral, injection (such as subcutaneous, intramuscular, intradermal, intraperitoneal, intravenous, intrathecal and intratumoral), sublingual, rectal, transdermal, intranasal, intraductal, vaginal and inhalation routes. In some aspects, administration is to a kidney of a subject.

[0281] Agent: Any vector, polypeptide, compound, small molecule, organic compound, salt, polynucleotide, or other molecule of interest. An agent can include a therapeutic agent, a diagnostic agent or a pharmaceutical agent. A “therapeutic agent” is a substance that demonstrates some therapeutic effect by restoring or maintaining health, such as by alleviating the symptoms associated with a disease or physiological disorder, or delaying (including preventing) progression or onset of a disease, such as a kidney disease, such as Alport syndrome. An agent can be an AAV vector.

[0282] Alport syndrome: A genetic disorder that can be caused by an inherited defect in type IV collagen, such as a mutation in the COL4A3, COL4 4, and / or COL4A5 gene. The mutations prevent normal production and / or assembly of the type IV collagen network (specifically the collagen 4 a345 network). This network is a part of tissues including the kidneys, eyes, and cochlea (inner ear).

[0283] Symptoms of Alport syndrome can include gross hematuria, microscopic hematuria, proteinuria, edema, nephrotic syndrome, chronic anemia, osteodystrophy, sensorineural hearing loss, anterior lenticonus, dot-and-fleck retinopathy, temporal macular thinning, leiomyomatosis of the tracheobronchial tree, leiomyomatosis of the esophagus. Alport syndrome can be diagnosed by a clinician, such as by examination of symptomology, or by genetic testing. Alport syndrome can be treated with ACE inhibitors, ARBs, and diuretics. However, many subjects with Alport syndrome will progress to dialysis or kidney transplantation.

[0284] Alport syndrome includes X-linked Alport syndrome (XLAS), which can be caused by a mutation in the COL4A5 gene. Alport syndrome also includes autosomal recessive (ARAS) or autosomal dominant Alport syndrome (ADAS), which can be caused by mutations in COL4A3 and / or COL4A4 gene.

[0285] Angiotensin-Converting-Enzyme Inhibitors (ACE Inhibitors): Agents which inhibit the activity of angiotensin-converting enzyme. Angiotensin-converting enzyme converts angiotensin I to angiotensin II and hydrolyses bradykinin. Exemplary ACE inhibitors include: benazepril, captopril, cilazapril, enalapril, fosinopril, lisinopril, moexipril, perindopril, ramipril, quinapril, and trandolapril.

[0286] Angiotensin II Receptor Blockers (ARBs): Agents which bind to and inhibit the angiotensin H type 1 receptor. Exemplary ARBs include: candesartan, eprosartan, irbesartan, losartan, olmesartan, telmisartan, and valsartan.

[0287] Conservative Substitutions: Such as of a polypeptide, involve the substitution of one or more amino acids for amino acids having similar biochemical properties that do not result in change or loss of a biological or biochemical function of the polypeptide are designated “conservative” substitutions. These conservative substitutions are likely to have minimal impact on the activity of the resultant protein. Table A shows amino acids that can be substituted for an original amino acid in a protein, and which are regarded as conservative substitutions.

[0288] TABLE A

[0289] Original Residue Conservative Substitutions

[0290] Ala ser

[0291] Arg lys

[0292] Asn gin; his

[0293] Asp glu

[0294] Cys ser

[0295] Gin asn

[0296] Glu asp

[0297] Gly pro

[0298] His asn; gin

[0299] He leu; val

[0300] Leu ile; val

[0301] Lys arg; gin; glu

[0302] Met leu; ile

[0303] Phe met; leu; tyr

[0304] Ser thr

[0305] Thr ser

[0306] Trp tyr

[0307] Tyr trp; phe

[0308] Val ile; leu

[0309] One or more conservative changes, or up to ten conservative changes (such as two substituted amino acids, three substituted amino acids, four substituted amino acids, or five substituted amino acids, etc.) can be made in the polypeptide without changing a biochemical function of the protein, such as a collagen or a Npu DnaE C-intein protein, such as the trans-splicing of the Npu DnaE intein pair.

[0310] Control: A reference standard. In some aspects, the control is a negative control sample obtained from a healthy patient. In other aspects, the control is a positive control, such as a sample obtained from a patient diagnosed with Alport syndrome or kidney disease. In still other aspects, the control is a historical control or standard reference value or range of values (such as a previously tested control sample, such as a group of Alport syndrome patients with known prognosis or outcome, or group of samples that represent baseline or normal values).

[0311] A “difference” between a test sample and a control can be an increase or conversely a decrease. The difference can be a qualitative difference or a quantitative difference, for example a statistically significant difference. In some examples, a difference is an increase or decrease, relative to a control, of at least about 5%, such as at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 68%, at least about 80%, at least about 90%, at least about 100%, at least about 150%, at least about 200%, at least about 250%, at least about 300%, at least about 350%, at least about 400%, at least about 500%, or greater than 500%.

[0312] Effective Amount: A quantity of a specified pharmaceutical or therapeutic agent (e.g., a recombinant AAV, first nucleic acid molecule, second nucleic acid molecule, or split intein system) sufficient to achieve a desired effect in a subject, or in a cell, being treated with the agent, such as increasing kidney function or another desired effect. The effective amount of the agent will be dependent on several factors, including, but not limited to, the subject or cells being treated, and the manner of administration of the therapeutic composition.

[0313] An effective amount of the split intein system is an amount sufficient to achieve a desired effect in a subject, or in a cell, being treated with the agent, such as increasing kidney function or another desired effect. An effective amount of a first nucleic acid molecule encoding an N-extein of an extein pair and an N-intein of an intein pair is an amount such that when administered with an effective amount of the second nucleic acid molecule encoding a C-intein of the intein pair, and a C-extein of the extein pair, the N- and C- exteins are spliced together to form a mature protein (such as COL4A3, COL4A4, or COL4A5) when the first and second nucleic acid molecules are expressed in mammalian cells, and an amount sufficient to achieve a desired effect in a subject, or in a cell, being treated with the agent, such as increasing kidney function or another desired effect.

[0314] Expression control sequences: Nucleic acid sequences that regulate the expression of a heterologous nucleic acid sequence to which it is operatively linked. Expression control sequences are operatively linked to a nucleic acid sequence when the expression control sequences control and regulate the transcription and, as appropriate, translation of the nucleic acid sequence. Thus, expression control sequences can include appropriate promoters, enhancers, transcription terminators, a start codon (i.e., ATG) in front of a protein-encoding gene, splicing signal for introns, maintenance of the correct reading frame of that gene to permit proper translation of mRNA, and stop codons. The term “control sequences” is intended to include, at a minimum, components whose presence can influence expression, and can also include additional components whose presence is advantageous, for example, leader sequences and fusion partner sequences. Expression control sequences can include a promoter, such as CAG. Expression control sequences can include a woodchuck hepatitis virus post-transcriptional regulatory element.

[0315] Heterologous: A heterologous protein or polypeptide refers to a protein or polypeptide derived from a different source or species. A heterologous nucleic acid molecule refers to a nucleic acid molecule derived from a different source or species. Thus, a heterologous protein, polypeptide, or nucleic acid molecule in a cell, refers to a protein, polypeptide, or nucleic acid molecule not naturally found in the cell in nature (e.g., an exogenous protein, polypeptide, or nucleic acid molecule). In some examples, a cell expressing a heterologous protein, polypeptide, or nucleic acid molecule is transgenic. Intein and Extein: Inteins are polypeptides capable of catalyzing their own excision from the N-extein and the C-extein that are adjacent to the intein sequence on its N-terminal and C-terminal side, respectively (the “exteins”). Inteins have been identified in organisms across all phylogenetic kingdoms, and for these naturally occurring inteins, the self -excision of the intein typically results in the linking of the N-extein to the C-extein through a peptide bond. The self-excision of inteins can result from interaction between intein sequences within the same polypeptide, or between intein polypeptides acting in trans (see split inteins, below).

[0316] “Split inteins” are intein polypeptides in which the N-terminal and C-terminal portions of the intein are expressed as separate polypeptides that can act in trans to accomplish the excision reaction.

[0317] The N-extein and the C-extein can be joined together at a “split point.” The split is the location in the mature polypeptide sequence at which the N-extein and the C-extein are linked via peptide bond. In this disclosure, the split point can be defined as the first amino acid residue of the C-extein. For example, in one exemplary sequence, the C-terminus of the N-extein includes the residues DEI, and the N-terminus of the C- extein includes the residues CEPG (SEQ ID NO: 230). Together, these can be represented by the sequence DEI|[C]EPG. In this example, the split point is the residue in brackets “[C]”. The ‘split’ between the N- extein and the C-extein occurs at the vertical bar character (“|”).

[0318] Isolated: An “isolated” biological component (such as a nucleic acid, peptide or protein) has been substantially separated, produced apart from, or purified away from other biological components in the cell of the organism in which the component naturally occurs, i.e., other chromosomal and extrachromosomal DNA and RNA, and proteins. Nucleic acids, peptides and proteins which have been “isolated” thus include nucleic acids and proteins purified by standard purification methods. The term also embraces nucleic acids, peptides and proteins prepared by recombinant expression in a host cell as well as chemically synthesized nucleic acids. An isolated cell type has been substantially separated from other cell types, such as a different cell type that occurs in an organ. A purified cell or component can be at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% pure.

[0319] Kidney: An organ found in vertebrates that participates in the control of the volume of various body fluids, fluid osmolality, acid-base balance, various electrolyte concentrations, and removal of toxins. The kidneys receive blood from the paired renal arteries; blood exits into the paired renal veins. Each kidney is attached to a ureter, a tube that carries excreted urine to the bladder. Filtration occurs in the glomerulus. One-fifth of the blood volume that enters the kidneys is filtered. Examples of substances reabsorbed are solute-free water, sodium, bicarbonate, glucose, and amino acids. Examples of substances secreted are hydrogen, ammonium, potassium and uric acid. The nephron is the structural and functional unit of the kidney. Each adult human kidney contains around 1 million nephrons, while a mouse kidney contains only about 12,500 nephrons. The kidneys also carry out functions independent of the nephrons. For example, they convert a precursor of vitamin D to its active form, calcitriol; and synthesize the hormones erythropoietin and renin. The kidney is comprised of the following components: (1) nephron, (2) interstitium and (3) vasculature. The (1) nephron comprises glomerular endothelium, mesangial cells, podocytes, renal tubules (proximal tubules, Loop of Henle cells, distal tubules, collecting duct cells, and parietal epithelial cells). The (2) interstitium comprises fibroblasts, myofibroblasts, and pericytes. The (3) vasculature comprises vascular endothelial cells, vascular smooth muscle cells. The kidney also contains blood cells or bone marrow-derived cells, including lymphocytes, monocytes, neutrophils, dendritic cells, mast cells and resident macrophages. The kidney also contains blood cells, including lymphocytes, monocytes, neutrophils and resident macrophages. The kidney also contains specialized endocrine cells that produce erythropoietin, renin and calcitriol. “Podocytes” are cells in Bowman's capsule in the kidneys that wrap around capillaries of the glomerulus. Podocytes make up the epithelial lining of Bowman's capsule of the kidney, which filters the blood. Although various viscera have epithelial layers, the name “visceral epithelial cells” usually refers specifically to podocytes, which are specialized epithelial cells that reside in the visceral layer of the capsule. The podocytes have long foot processes called pedicels. These pedicels wrap around the capillaries and leave slits between them. Blood is filtered through these slits, each known as a filtration slit, slit diaphragm, or slit pore. “Epithelial cells” form the lining of the urinary tract. “Renal tubular epithelial cells” line the collecting ducts and the distal and proximal tubules of the kidney. In the kidney, the “proximal tube” is the segment of the nephron in kidneys which begins from the renal pole of the Bowman's capsule to the beginning of loop of Henle. It can be further classified into the proximal convoluted tubule and the proximal straight tubule. “Proximal tubule cells” are the cells from the proximal tubule. The luminal surface of the epithelial cells of this segment of the nephron is covered with densely packed microvilli forming a border that increases the luminal surface area of the cells, which are also densely packed with mitochondria. Cuboidal epithelial cells lining the proximal tubule have extensive lateral interdigitations between neighboring cells, which lend an appearance of having no discrete cell margins when viewed with a light microscope. A “nephron” is a microscopic structural and functional unit of the kidney composed of a renal corpuscle and a renal tubule. The renal corpuscle consists of a tuft of capillaries called a glomerulus and a cup-shaped structure called Bowman's capsule. The renal tubule extends from the capsule. The capsule and tubule are connected and are composed of epithelial cells with a lumen. A healthy adult has 1 to 1.5 million nephrons in each kidney.

[0320] Nucleic Acid Molecule: A polymer composed of nucleotide units (ribonucleotides, deoxyribonucleotides, related naturally occurring structural variants, and synthetic non-naturally occurring analogs thereof) linked via phosphodiester bonds, related naturally occurring structural variants, and synthetic non-naturally occurring analogs thereof. Thus, the term includes nucleotide polymers in which the nucleotides and the linkages between them include non-naturally occurring synthetic analogs, such as, for example and without limitation, phosphorothioates, phosphoramidates, methyl phosphonates, chiral-methyl phosphonates, 2-O-methyl ribonucleotides, peptide-nucleic acids (PNAs), and the like. Such polynucleotides can be synthesized, for example, using an automated DNA synthesizer. The term “oligonucleotide” typically refers to short polynucleotides, generally no greater than about 50 nucleotides. It will be understood that when a nucleotide sequence is represented by a DNA sequence (i.e., A, T, G, C), this also includes an RNA sequence (i.e., A, U, G, C) in which “U” replaces “T.” Conventional notation is used herein to describe nucleotide sequences: the left-hand end of a singlestranded nucleotide sequence is the 5'-end; the left-hand direction of a double-stranded nucleotide sequence is referred to as the 5'-direction. The direction of 5' to 3' addition of nucleotides to nascent RNA transcripts is referred to as the transcription direction. The DNA strand having the same sequence as an mRNA is referred to as the “coding strand;” sequences on the DNA strand having the same sequence as an mRNA transcribed from that DNA and which are located 5' to the 5'-end of the RNA transcript are referred to as “upstream sequences;” sequences on the DNA strand having the same sequence as the RNA and which are 3' to the 3' end of the coding RNA transcript are referred to as “downstream sequences.” With regard to nucleotide molecules, the term “encoding” refers to the inherent property of specific sequences of nucleotides in a polynucleotide, such as a gene, a cDNA, or an mRNA, to serve as templates for synthesis of other polymers and macromolecules in biological processes having either a defined sequence of nucleotides (i.e., rRNA, tRNA and mRNA) or a defined sequence of amino acids and the biological properties resulting therefrom. Thus, a gene encodes a protein if transcription and translation of mRNA produced by that gene produces the protein in a cell or other biological system. Both the coding strand, the nucleotide sequence of which is identical to the mRNA sequence and is usually provided in sequence listings, and non-coding strand, used as the template for transcription, of a gene or cDNA can be referred to as encoding the protein or other product of that gene or cDNA. Unless otherwise specified, a "nucleotide sequence encoding an amino acid sequence" includes all nucleotide sequences that are degenerate versions of each other and that encode the same amino acid sequence. Nucleotide sequences that encode proteins and RNA may include introns.

[0321] In addition, a “recombinant nucleic acid” refers to a nucleic acid molecule that has a sequence that is not naturally occurring or has a sequence that is made by an artificial combination of two otherwise separated segments of sequence. This artificial combination is often accomplished by chemical synthesis or, more commonly, by the artificial manipulation of isolated segments of nucleic acids, such as by genetic engineering techniques. Similarly, a recombinant protein is one encoded for by a recombinant nucleic acid molecule. In addition, a recombinant virus is a virus comprising sequence (such as genomic sequence) that is non-naturally occurring or made by artificial combination of at least two sequences of different origin. The term “recombinant” also includes nucleic acids, proteins and viruses that have been altered solely by addition, substitution, or deletion of a portion of a natural nucleic acid molecule, protein or virus. As used herein, “recombinant AAV” refers to an AAV particle in which a recombinant nucleic acid molecule (such as a recombinant nucleic acid molecule encoding a therapeutic protein) has been packaged.

[0322] Nucleic acid sequences of the current disclosure can be defined in terms of particular identity and / or similarity with certain nucleic acid sequences described herein. The sequence identity can be greater than 60%, greater than 75%, greater than 80%, greater than 90%, and can be greater than 95%. The identity and / or similarity of a sequence can be 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% as compared to a sequence disclosed herein. The sequence identify can be about 95%, 96%, 97% 98%, or 99%.

[0323] Operably Linked: A first nucleic acid sequence is “operably linked” with a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For instance, a promoter is operably linked to a coding sequence if the promoter affects the transcription or expression of the coding sequence. Generally, operably linked DNA sequences are contiguous and, where necessary to join two protein coding regions, in the same reading frame.

[0324] Pharmaceutically Acceptable Carrier: A “pharmaceutically acceptable carrier” of use includes conventional excipients. Remington’s Pharmaceutical Sciences, 23rd Edition, Academic Press, Elsevier, (2020), describes compositions and formulations suitable for pharmaceutical delivery of the fusion proteins herein disclosed. In general, the nature of the carrier will depend on the particular mode of administration being employed. For instance, parenteral formulations usually comprise injectable fluids that include pharmaceutically and physiologically acceptable fluids such as water, physiological saline, balanced salt solutions, aqueous dextrose, glycerol or the like as a vehicle. In addition to biologically-neutral carriers, pharmaceutical compositions to be administered can contain minor amounts of non-toxic auxiliary substances, such as wetting or emulsifying agents, preservatives, and pH buffering agents and the like, for example sodium acetate or sorbitan monolaurate.

[0325] Polypeptide or Protein: a polymer in which the monomers are amino acid residues that are joined together through amide bonds. When the amino acids are alpha-amino acids, either the L-optical isomer or the D-optical isomer can be used, the L-isomers being preferred. The terms "polypeptide" or “protein” as used herein is intended to encompass any amino acid sequence and include modified sequences such as glycoproteins. The term “polypeptide” is specifically intended to cover naturally occurring proteins, as well as those that are recombinantly or synthetically produced. Conventional notation is used herein to describe polypeptides: the left-hand end of a polypeptide sequence is the N-terminal end, and the polypeptide progresses from N-terminus to C-terminus. Typically, the N-terminus includes a free amine group and typically the C-terminus includes a free carboxylic group. A “therapeutic protein” is a protein that, when expressed in a subject, results in an improvement of a sign or a symptom of a particular disorder in that subject, such as a kidney. In some aspects, administration of a “therapeutic protein” results in an improvement of a sign or a symptom of a kidney disorder in a subject. Administration of a therapeutic protein can result in improvement of disease in a subject, such as, Alport syndrome. A “fusion protein” is made from two heterologous proteins from different sources.

[0326] Polypeptide sequences of the current disclosure can be defined in terms of particular identity and / or similarity with certain polypeptides described herein. The sequence identity can be greater than 60%, greater than 75%, greater than 80%, greater than 90%, and can be greater than 95%. The identity and / or similarity of a sequence can be 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% as compared to a sequence disclosed herein. The sequence identify can be about 95%, 96%, 97% 98%, or 99%.

[0327] Preventing, Treating, and Ameliorating: “Preventing” a disease (such as a kidney disease) refers to inhibiting the full development of a disease. “Treating” refers to a therapeutic intervention that ameliorates a sign or symptom of a disease or pathological condition after it has begun to develop. “Ameliorating” refers to the reduction in the number or severity of signs or symptoms of a disease.

[0328] Promoter: An array of nucleic acid control sequences which direct transcription of a nucleic acid. A promoter includes necessary nucleic acid sequences near the start site of transcription, such as, in the case of an RNA polymerase II type promoter, a TATA element. A promoter also optionally includes distal enhancer or repressor elements which can be located as much as several thousand base pairs from the start site of transcription. Also included are those promoter elements which are sufficient to render promoterdependent gene expression controllable for cell type-specific, tissue-specific, or inducible by external signals or agents; such elements may be located in the 5' or 3' regions of the gene. Both constitutive and inducible promoters are included. The promoter can direct expression in the cells of the kidney.

[0329] Signal Sequence: A short amino acid sequence (e.g., approximately 18-30 amino acids in length) that directs newly synthesized secretory or membrane proteins to and through membranes (for example, the endoplasmic reticulum membrane). Signal sequences are typically located at the N terminus of a polypeptide and are removed by signal peptidases after the polypeptide has crossed the membrane. Signal sequences typically contain three common structural features: an N-terminal polar basic region (n-region), a hydrophobic core, and a hydrophilic c-region). Exemplary signal sequences are set forth in SEQ ID NOs: 4 and 5. A signal sequence is also known as a signal peptide.

[0330] Subject: Any mammal, such as humans, non-human primates, pigs, sheep, cows, rodents and the like which is to be the recipient of the particular treatment. In two non-limiting examples, a subject is a human subject or a non-human primate subject. The subject can have a kidney disease, such as Alport syndrome.

[0331] Transduces: A virus or vector “transduces” a cell when it transfers nucleic acid into the cell and results in functional consequences. A cell is “transformed” or “transfected” by a nucleic acid transduced into the cell when the DNA becomes stably maintained by the cell, either by incorporation of the nucleic acid into the cellular genome, or by episomal persistence.

[0332] Methods of transfection include: chemical methods (e.g., calcium-phosphate transfection), physical methods (e.g., electroporation, microinjection, particle bombardment), fusion (e.g., liposomes), receptor- mediated endocytosis (e.g., DNA-protein complexes, viral envelope / capsid-DNA complexes) and by biological infection by viruses such as recombinant viruses (Wolff, J. A., ed, Gene Therapeutics, Birkhauser, Boston, USA (1994)). In the case of infection by viruses, the infecting virus particles are absorbed by the target cells, resulting in integration of the viral genome into the cellular DNA or persistent presence of viral genetic components as episomes. Genetic modification of the target cell is an indicium of successful transfection. "Genetically modified cells" refers to cells whose genotypes have been altered as a result of cellular uptakes of exogenous nucleotide sequence by transfection. A reference to a transfected cell or a genetically modified cell includes both the particular cell into which a vector or polynucleotide is introduced and progeny of that cell.

[0333] Transgene: An exogenous gene supplied by a vector or a virus, such as an AAV.

[0334] Type IV Collagen: A type of collagen that is a structural component of basement membranes and found primarily in the basal lamina. In vivo, Type IV collagen aids in cell adhesion, migration, survival, expansion and differentiation of cells.

[0335] In the synthesis of type IV collagen, initially a trimer forms between three type IV collagen alpha chains to form a protomer. Two protomers dimerize to form a hexamer, these hexamers associate into a protein network. There are six human genes associated with type IV collagen: the C0L4A1, COL4A2, COL4A3, COL4A4, COL4A5, and COL4A6 genes. The structure of the protein network is affected by the constituent subunits that form the network, for example the type IV collagen a345 network includes the products of the COL4A3, COL4A4, and COL4A5 genes.

[0336] Type IV collagen proteins can include a 7S domain, a collagenous domain, and a non-collagenous C -terminal domain. The “7S domain” promotes the formation of heterotrimers and subsequent meshwork together with the NCI domain described below, and includes amino acid residues 27-41 relative to SEQ ID NO: 15. The “collagenous domain” (CD) forms a triple helical structure with the CDs of the other two alpha chains, contains regions that interact with other extracellular matrix components, and has a role in maintaining the structural integrity of the collagen meshwork. The CD includes amino acid residues 42-1456 relative to SEQ ID NO: 15. The non-collagenous C-terminal domain 1 (NCI) serves as the promoter of heterotrimerization and determines the combination of three alpha chains in the heterotrimers. Besides the NCI's role in maintaining the molecular architecture of the basement membranes, the NCI can be cleaved by proteolysis and serves as a soluble factor involved in various biological processes including morphogenesis, angiogenesis and tumorigenesis. The NCI domain includes amino acid residues 1457-1685 relative to SEQ ID NO: 15.

[0337] “COL4A3” a protein encoded by the COL4A3 gene, which encodes one possible subunit of type IV collagen. Exemplary COL4A3 amino acid sequences include SEQ ID NO: 13. NCBI Gene IDs: 1285 (human); 12828 (Mus musculus), as available on January 17, 2024. Exemplary mRNA encoding COL4A3 includes NCBI RefSeq NM 000091.5 (human) as available on March 11, 2024.

[0338] “COL4A4” a protein encoded by the COL4A4 gene, which encodes one possible subunit of type IV collagen. Exemplary COL4A4 amino acid sequences include SEQ ID NO: 14. NCBI Gene IDs: 1286 (human); 12829 (Mus musculus), as available on January 17, 2024. Exemplary mRNA encoding COL4A4 includes NCBI RefSeq NMJJ00092.5 (human) as available on March 11, 2024.

[0339] “COL4A5” a protein encoded by the COL4A5 gene, which encodes one of the subunits of type IV collagen. Exemplary COL4A5 amino acid sequences include SEQ ID NOs: 15-25. NCBI Gene IDs: 1287 (human); 12830 (Mus musculus), as available on January 17, 2024. Exemplary mRNA encoding COL4A5 include NCBI RefSeqs NM_000495.5, (human, transcript variant 1) and NM_033380.3 (human, transcript variant 2).

[0340] III. Overview

[0341] Disclosed herein is a split intein system including a first nucleic acid molecule encoding a fusion protein including in N to C terminal order, an N-extein of an extein pair and an N-intein of an intein pair, where the N-extein includes a first signal sequence N-terminal to an N-terminal portion of a type IV collagen alpha chain. The split intein system includes a second nucleic acid molecule encoding a fusion protein including in N to C terminal order, a second signal sequence, a C-intein of the intein pair, and a C- extein of the extein pair, where the C-extein includes a C-terminal portion of the type IV collagen alpha chain. In some examples, the type IV collagen alpha chain is COL4A5. In some examples, the N-terminal portion does not include a collagen signal sequence. In some examples, the N- and C-terminal portions of the type IV collagen alpha chain together define the type IV collagen alpha chain sequence separated at a split point. In some examples, the N- and C-exteins are spliced together to form the mature type IV collagen alpha chain when the first and second nucleic acid molecules are expressed in mammalian cells.

[0342] In some examples, the system includes a first adeno-associated viral (AAV) vector including the first nucleic acid molecule and / or a second AAV vector including the second nucleic acid molecule. In some examples, the first and / or the second AAV vector is AAV9. In some examples, the first and / or the second AAV vector is AAV-KP1. In some examples, the first and / or the second AAV vector is AAV1, AAVl_9mtl00, AAVl_9mt30, AAVl_9mt76, AAV2, AAV2G9, AAV2i8, AAV2retro, AAV2R585E, AAV2R585E9_2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV9AA272, AAV9AA22, AAV9W22A, AAV10, AAV11, AAV2.7m8, AAVAnc80, AAVbb.2, AAV-DJ, AAVHN1, AAVHN2, AAVHN3, AAVhu.ll, AAVhu.13, AAVhu.37, AAV-KP1, AAV-KP2, AAV-KP3, AAVLK03, AAVNP40, AAVNP59, AAVPHP.B, AAVPHP.eB, AAVPHP.S, AAVpol, AAVrh.8, AAVrh.10, AAVrh.20, AAVrh.43, and / or AAVShHIO.

[0343] In some examples, the system includes a first Dependoparvovirus viral vector including the first nucleic acid molecule and / or a second Dependoparvovirus viral vector including the second nucleic acid molecule. In some examples, the system includes a first Parvovirinae viral vector including the first nucleic acid molecule and / or a second Parvovirinae viral vector including the second nucleic acid molecule. In some examples, the system includes a first Parvoviridae viral vector including the first nucleic acid molecule and / or a second Parvoviridae viral vector including the second nucleic acid molecule.

[0344] In some examples, the intein pair is a Npu DnaE intein pair. In some examples, the N-intein includes the amino acid sequence of SEQ ID NO: 26 or an amino acid sequence at least 95% identical thereto. In some examples, the C-intein includes the amino acid sequence of SEQ ID NO: 27 or an amino acid sequence at least 95% identical thereto identical thereto. In some examples, the first signal sequence, the N-terminal portion of the type IV collagen alpha chain, and the N-intein are fused directly. In some examples, the second signal sequence, the C-intein, and the C-terminal portion of the type IV collagen alpha chain are fused directly.

[0345] In some examples, the split intein system includes the first nucleic acid molecule and the second nucleic acid molecule at a ratio of about 1:1 to about 0.010:1. In some examples, the split intein system includes the first nucleic acid molecule and the second nucleic acid molecule at a ratio of about 0.017:1 to about 0.010:1.

[0346] In some examples, the split point is within a collagenous domain of the type IV collagen alpha chain.

[0347] In some examples, the first signal sequence and / or the second signal sequence is a collagen signal sequence or a chymotrypsin signal sequence. In some examples, the first signal sequence and the second signal sequence are different. In some examples the collagen signal sequence includes the amino acid sequence of SEQ ID NO: 4. In some examples the chymotrypsin signal sequence includes the amino acid sequence of SEQ ID NO: 5.

[0348] In some examples, the type IV collagen alpha chain includes the amino acid sequence of any one of SEQ ID NOs: 15-25 or an amino acid sequence at least 95% identical thereto. In some examples, the type IV collagen alpha chain includes the amino acid sequence of SEQ ID NO: 15 or an amino acid sequence at least 95% identical thereto.

[0349] In some examples, the split point is at position 729 to position 1488 relative to SEQ ID NO: 15.

[0350] In some examples the N-terminal portion includes amino acids 1 - 450 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 451 to 1685 of the type IV collagen alpha chain (SP1). In some examples the N-terminal portion includes amino acids 1 - 696 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 697 to 1685 of the type IV collagen alpha chain (SP2). In some examples the N-terminal portion includes amino acids 1 - 863 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 864 to 1685 of the type IV collagen alpha chain (SP4). In some examples the N-terminal portion includes amino acids 1 - 878 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 879 to 1685 of the type IV collagen alpha chain (SP5) . In some examples the N-terminal portion includes amino acids 1 - 941 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 942 to 1685 of the type IV collagen alpha chain (SP7) . In some examples the N-terminal portion includes amino acids 1 - 977 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 978 to 1685 of the type IV collagen alpha chain (SP9). In some examples the N-terminal portion includes amino acids 1 - 1070 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 1071 to 1685 of the type IV collagen alpha chain (SP11). In some examples, the N-terminal portion comprises amino acids 1 - 1135 of the type IV collagen alpha chain, and / or the C-terminal portion comprises amino acids 1136 to 1685 of the type IV collagen alpha chain (SP12). In some examples the N-terminal portion includes amino acids 1 - 696 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 697 - 1685 of the type IV collagen alpha chain (SP2). In some examples the N-terminal portion includes amino acids 1 - 878 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 879 - 1685 of the type IV collagen alpha chain (SP5) . In some examples the N-terminal portion includes amino acids 1 - 888 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 889 - 1685 of the type IV collagen alpha chain (SP20). In some examples the N-terminal portion includes amino acids 1 - 915 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 916 - 1685 of the type IV collagen alpha chain (SP6). In some examples the N-terminal portion includes amino acids 1 - 941 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 942 - 1685 of the type IV collagen alpha chain (SP7). In some examples the N-terminal portion includes amino acids 1 - 961 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 962 - 1685 of the type IV collagen alpha chain (SP8). In some examples the N-terminal portion includes amino acids 1 - 977 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 978 - 1685 of the type IV collagen alpha chain (SP9). In some examples the N-terminal portion includes amino acids 1 - 995 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 996 - 1685 of the type IV collagen alpha chain (SP10). In some examples the N-terminal portion includes s amino acids 1 - 1070 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 1071 - 1685 of the type IV collagen alpha chain (SPll). In some examples the N-terminal portion includes amino acids 1 - 1098 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 1099 - 1685 of the type IV collagen alpha chain (SP27). In some examples the N-terminal portion includes amino acids 1 - 1135 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 1136 - 1685 of the type IV collagen alpha chain (SP12).

[0351] In some examples the N-terminal portion includes amino acids 27 - 696 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 697 - 1685 of the type IV collagen alpha chain (SP2). In some examples the N-terminal portion includes amino acids 27 - 878 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 879 - 1685 of the type IV collagen alpha chain (SP5). In some examples the N-terminal portion includes amino acids 27 - 888 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 889 - 1685 of the type IV collagen alpha chain (SP20). In some examples the N-terminal portion includes amino acids 27 - 915 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 916 - 1685 of the type IV collagen alpha chain (SP6). In some examples the N-terminal portion includes amino acids 27 - 941 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 942 - 1685 of the type IV collagen alpha chain (SP7). In some examples the N-terminal portion includes amino acids 27 - 961 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 962 - 1685 of the type IV collagen alpha chain (SP8). In some examples the N-terminal portion includes amino acids 27 - 977 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 978 - 1685 of the type IV collagen alpha chain (SP9). In some examples the N-terminal portion includes amino acids 27 - 995 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 996 - 1685 of the type IV collagen alpha chain (SP10). In some examples the N-terminal portion includes s amino acids 27- 1070 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 1071 - 1685 of the type IV collagen alpha chain (SP11). In some examples the N-terminal portion includes amino acids 27 - 1098 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 1099 - 1685 of the type IV collagen alpha chain (SP27). In some examples the N-terminal portion includes amino acids 27 - 1135 of the type IV collagen alpha chain, and / or the C-terminal portion includes amino acids 1136 - 1685 of the type IV collagen alpha chain (SP12).

[0352] In some examples the N terminal portion includes DEI at its C-terminus and / or the C-terminal portion includes CEPG (SEQ ID NO: 230) at its N-terminus (SP1). In some examples the N terminal portion includes IPG at its C-terminus and / or the C-terminal portion includes SKGE (SEQ ID NO: 221) at its N-terminus (SP2). In some examples the N terminal portion includes ERG at its C-terminus and / or the C-terminal portion includes SPGI (SEQ ID NO: 231) at its N-terminus (SP4). In some examples the N terminal portion includes PPG at its C-terminus and / or the C-terminal portion includes SPGL (SEQ ID NO: 222) at its N-terminus (SP5). In some examples the N terminal portion includes EKG at its C-terminus and / or the C-terminal portion includes SKGE (SEQ ID NO: 221) at its N-terminus (SP7). In some examples the N terminal portion includes PGV at its C-terminus and / or the C-terminal portion includes SGPK (SEQ ID NO: 225) at its N-terminus (SP9). In some examples the N terminal portion includes PGI at its C- terminus and / or the C-terminal portion includes SSIG (SEQ ID NO: 227) at its N-terminus (SP11). In some examples, the N terminal portion includes KGI at its C-terminus and / or the C-terminal portion includes SGPP (SEQ ID NO: 229) at its N-terminus (SP12).

[0353] In some examples the N-terminal portion comprises IPG at its C-terminus and / or the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP2). In some examples the N-terminal portion comprises PPG at its C-terminus and / or the C-terminal portion comprises SPGL (SEQ ID NO: 222) at its N-terminus (SP5). In some examples the N-terminal portion comprises AGA at its C-terminus and / or the C-terminal portion comprises SGFP (SEQ ID NO: 223) at its N-terminus (SP20). In some examples the N-terminal portion comprises PGR at its C-terminus and / or the C-terminal portion comprises SGVP (SEQ ID NO: 224) at its N-terminus (SP6). In some examples the N-terminal portion comprises EKG at its C- terminus and / or the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP7). In some examples the N-terminal portion comprises LLG at its C-terminus and / or the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP8). In some examples the N-terminal portion comprises PGV at its C-terminus and / or the C-terminal portion comprises SGPK (SEQ ID NO: 225) at its N- terminus (SP9). In some examples the N-terminal portion comprises PGL at its C-terminus and / or the C- terminal portion comprises SGQP (SEQ ID NO: 226) at its N-terminus (SP10). In some examples the N- terminal portion comprises PGI at its C-terminus and / or the C-terminal portion comprises SSIG (SEQ ID NO: 227) at its N-terminus (SP11). In some examples the N-terminal portion comprises IKG at its C- terminus and / or the C-terminal portion comprises SVGD (SEQ ID NO: 228) at its N-terminus (SP27). In some examples the N-terminal portion comprises KGI at its C-terminus and / or the C-terminal portion comprises SGPP (SEQ ID NO: 229) at its N-terminus (SP12).

[0354] In some examples, the site including the C-terminus of the N-extein and the N-terminus of the C- extein has at least 14%, at least 28%, at least 42%, at least 57%, at least 71%, at least 85%, or at least 100% identity to at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or at least 11 of SEQ ID NOs: 210-220. In some examples, the site including the C-terminus of the N- extein and the N-terminus of the C-extein shares at least 1, at least 2, at least 3, at least 4, at least 5, at least

[0355] 6, or at least 6 contiguous residues with at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least

[0356] 7, at least 8, at least 9, at least 10, or at least 11 of SEQ ID NOs: 210-220. In some examples, the site including the C-terminus of the N-extein and the N-terminus of the C-extein has at most 86%, at most 72%, at most 58%, at most 43%, at most 29%, or at most 15% identity to at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, or at least 16 of SEQ ID NOs: 194-209. In some examples, the site including the C-terminus of the N-extein and the N-terminus of the C-extein shares at most 1, at most 2, at most 3, at most 4, at most 5, or at most 6 contiguous residues with at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 1 1 , at least 12, at least 1 , at least 14, at least 15, or at least 16 of SEQ ID NOs: 194-209. In some examples, the C-terminal portion of the split intein system includes the amino acid S at its N-terminus.

[0357] In further examples, the first nucleic acid molecule and / or the second nucleic acid molecule is operably linked to a CAG promoter. In some examples, the first nucleic acid molecule and / or the second nucleic acid molecule further includes a woodchuck hepatitis virus post-transcriptional regulatory element. In specific examples, the first nucleic acid molecule and / or the second nucleic acid molecule is operably linked to the woodchuck hepatitis virus post-transcriptional regulatory element. In some examples, the first nucleic acid molecule and / or the second nucleic acid molecule further encodes a SV40 polyadenylation signal.

[0358] Further disclosed herein is a pharmaceutical composition including an effective amount of the split intein system and a pharmaceutically acceptable carrier, where the split intein system includes the first nucleic acid molecule and the second nucleic acid molecule at a ratio of about 1 : 1 to about 0.010: 1.

[0359] Also disclosed is herein is split intein system including a first adeno-associated viral (AAV) vector including a first nucleic acid molecule operably linked to a CAG promoter, where the first nucleic acid molecule encodes a fusion protein including in N to C terminal order, an N-extein of an extein pair and an N-intein of an Npu DnaE intein pair, where the N-extein includes a collagen signal sequence N-terminal to an N-terminal portion of a type IV collagen alpha 5 (COL4A5) chain, where the collagen signal sequence, the N-terminal portion of the COL4A5 chain, and the N-intein are fused directly. The split intein system can further include a second AAV vector including a second nucleic acid molecule operably linked to a CAG promoter, where the second nucleic acid molecule encodes a fusion protein including, in N to C terminal order, a signal sequence, the C-intein of the Npu DnaE intein pair, and the C-extein of the extein pair, where the C-extein includes a C-terminal portion of the COL4A5 chain, where the signal sequence, the C-intein, and the C-terminal portion of the COL4A5 chain are fused directly. In some examples, the N- and C- terminal portions of the COL4A5 chain together include the COL4A5 chain sequence separated at a split point. In further examples, the N- and C-exteins are spliced together to form the mature COL4A5 chain when the first and second nucleic acid molecules are expressed in mammalian cells.

[0360] In some examples the N-terminal portion includes amino acids 1 - 450 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 451 to 1685 of the COL4A5 chain (SP1). In some examples the N-terminal portion includes amino acids 1 - 696 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 697 to 1685 of the COL4A5 chain (SP2). In some examples the N-terminal portion includes amino acids 1 - 863 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 864 to 1685 of the COL4A5 chain (SP4). In some examples the N-terminal portion includes amino acids 1 - 878 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 879 to 1685 of the COL4A5 chain (SP5). In some examples the N-terminal portion includes amino acids 1 - 941 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 942 to 1685 of the COL4A5 chain (SP7). In some examples the N-terminal portion includes amino acids 1 - 977 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 978 to 1685 of the COL4A5 chain (SP9). In some examples the N-terminal portion includes amino acids 1 - 1070 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 1071 to 1685 of the COL4A5 chain (SP11). In some examples the N-terminal portion includes amino acids 1 - 1135 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 1136 to 1685 of the COL4A5 chain (SP12).

[0361] In some examples the N-terminal portion includes amino acids 1 - 696 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 697 - 1685 of the COL4A5 chain (SP2). In some examples the N-terminal portion includes amino acids 1 - 878 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 879 - 1685 of the COL4A5 chain (SP5). In some examples the N-terminal portion includes amino acids 1 - 888 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 889 - 1685 of the COL4A5 chain (SP20). In some examples the N-terminal portion includes amino acids 1 - 915 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 916 - 1685 of the COL4A5 chain (SP6). In some examples the N-terminal portion includes amino acids 1 - 941 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 942 - 1685 of the COL4A5 chain (SP7). In some examples the N-terminal portion includes amino acids 1 - 961 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 962 - 1685 of the COL4A5 chain (SP8). In some examples the N-terminal portion includes amino acids 1 - 977 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 978 - 1685 of the COL4A5 chain (SP9). In some examples the N-terminal portion includes amino acids 1 - 995 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 996 - 1685 of the COL4A5 chain (SP10). In some examples the N-terminal portion includes amino acids 1 - 1070 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 1071 - 1685 of the COL4A5 chain (SP11). In some examples the N-terminal portion includes amino acids 1 - 1098 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 1099 - 1685 of the COL4A5 chain (SP27). In some examples the N-terminal portion includes amino acids 1 - 1135 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 1136 - 1685 of the COL4A5 chain (SP12).

[0362] In some examples the N-terminal portion does not include an N-terminal methionine at amino acid 1. In some examples the N-terminal portion includes amino acids 2 - 450 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 451 to 1685 of the COL4A5 chain (SP1). In some examples the N- terminal portion includes amino acids 2 - 696 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 697 to 1685 of the COL4A5 chain (SP2). In some examples the N-terminal portion includes amino acids 2 - 863 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 864 to 1685 of the COL4A5 chain (SP4). In some examples the N-terminal portion includes amino acids 2 - 878 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 879 to 1685 of the COL4A5 chain (SP5). In some examples the N-terminal portion includes amino acids 2 - 941 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 942 to 1685 of the COL4A5 chain (SP7). In some examples the N-terminal portion includes amino acids 2 - 977 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 978 to 1685 of the COL4A5 chain (SP9). In some examples the N-terminal portion includes amino acids 2 - 1070 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 1071 to 1685 of the COL4A5 chain (SP11). In some examples the N-terminal portion includes amino acids 2 - 1135 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 1136 to 1685 of the COL4A5 chain (SP12).

[0363] In some further examples the N-terminal portion does not include an N-terminal methionine at amino acid 1. In some examples the N-terminal portion includes amino acids 2 - 696 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 697 - 1685 of the COL4A5 chain (SP2). In some examples the N-terminal portion includes amino acids 2 - 878 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 879 - 1685 of the COL4A5 chain (SP5). In some examples the N-terminal portion includes amino acids 2 - 888 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 889 - 1685 of the COL4A5 chain (SP20). In some examples the N-terminal portion includes amino acids 2 - 915 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 916 - 1685 of the COL4A5 chain (SP6). In some examples the N-terminal portion includes amino acids 2 - 941 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 942 - 1685 of the COL4A5 chain (SP7). In some examples the N-terminal portion includes amino acids 2 - 961 of the COL4A5 chain, and / or the C- terminal portion includes amino acids 962 - 1685 of the COL4A5 chain (SP8). In some examples the N- terminal portion includes amino acids 2 - 977 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 978 - 1685 of the COL4A5 chain (SP9). In some examples the N-terminal portion includes amino acids 2 - 995 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 996 - 1685 of the COL4A5 chain (SP10). In some examples the N-terminal portion includes amino acids 2 - 1070 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 1071 - 1685 of the COL4A5 chain (SP11). In some examples the N-terminal portion includes amino acids 2 - 1098 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 1099 - 1685 of the COL4A5 chain (SP27). In some examples the N-terminal portion includes amino acids 2 - 1135 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 1136 - 1685 of the COL4A5 chain (SP12).

[0364] In some examples the N-terminal portion does not include a collagen signal sequence, such as SEQ ID NO: 4. In some examples the N-terminal portion includes amino acids 27 - 450 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 451 to 1685 of the COL4A5 chain (SP1). In some examples the N-terminal portion includes amino acids 27 - 696 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 697 to 1685 of the COL4A5 chain (SP2). In some examples the N-terminal portion includes amino acids 27 - 863 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 864 to 1685 of the COL4A5 chain (SP4). In some examples the N-terminal portion includes amino acids 27 - 878 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 879 to 1685 of the COL4A5 chain (SP5). In some examples the N-terminal portion includes amino acids 27 - 941 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 942 to 1685 of the COL4A5 chain (SP7). In some examples the N-terminal portion includes amino acids 27 - 977 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 978 to 1685 of the COL4A5 chain (SP9). In some examples the N-terminal portion includes amino acids 27 - 1070 of the COL4A5 chain, and / or the C- terminal portion includes amino acids 1071 to 1685 of the COL4A5 chain (SP11). In some examples the N- terminal portion includes amino acids 27 - 1135 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 1136 to 1685 of the COL4A5 chain (SP12).

[0365] In some further examples the N-terminal portion does not include a collagen signal sequence, such as SEQ ID NO: 4. In some examples the N-terminal portion includes amino acids 27 - 696 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 697 - 1685 of the COL4A5 chain (SP2). In some examples the N-terminal portion includes amino acids 27 - 878 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 879 - 1685 of the COL4A5 chain (SP5). In some examples the N-terminal portion includes amino acids 27 - 888 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 889 - 1685 of the COL4A5 chain (SP20). In some examples the N-terminal portion includes amino acids 27 - 915 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 916 - 1685 of the COL4A5 chain (SP6). In some examples the N-terminal portion includes amino acids 27 - 941 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 942 - 1685 of the COL4A5 chain (SP7). In some examples the N-terminal portion includes amino acids 27 - 961 of the COL4A5 chain, and / or the C- terminal portion includes amino acids 962 - 1685 of the COL4A5 chain (SP8). In some examples the N- terminal portion includes amino acids 27 - 977 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 978 - 1685 of the COL4A5 chain (SP9). In some examples the N-terminal portion includes amino acids 27 - 995 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 996 - 1685 of the COL4A5 chain (SP10). In some examples the N-terminal portion includes amino acids 27 - 1070 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 1071 - 1685 of the COL4A5 chain (SP11). In some examples the N-terminal portion includes amino acids 27 - 1098 of the COL4A5 chain, and / or the C-terminal portion includes amino acids 1099 - 1685 of the COL4A5 chain (SP27). In some examples the N-terminal portion includes amino acids 27 - 1135 of the COL4A5 chain, and / or the C- terminal portion includes amino acids 1136 - 1685 of the COL4A5 chain (SP12).

[0366] In some examples the N terminal portion includes DEI at its C-terminus and / or the C-terminal portion includes CEPG (SEQ ID NO: 230) at its N-terminus (SP1). In some examples the N terminal portion includes IPG at its C-terminus and / or the C-terminal portion includes SKGE (SEQ ID NO: 221) at its N-terminus (SP2). In some examples the N terminal portion includes ERG at its C-terminus and / or the C-terminal portion includes SPGI (SEQ ID NO: 231) at its N-terminus (SP4). In some examples the N terminal portion includes PPG at its C-terminus and / or the C-terminal portion includes SPGL (SEQ ID NO: 222) at its N-terminus (SP5). In some examples the N terminal portion includes EKG at its C-terminus and / or the C-terminal portion includes SKGE (SEQ ID NO: 221) at its N-terminus (SP7). In some examples the N terminal portion includes PGV at its C-terminus and / or the C-terminal portion includes SGPK (SEQ ID NO: 225) at its N-terminus (SP9). In some examples the N terminal portion includes PGI at its C- terminus and / or the C-terminal portion includes SSIG (SEQ ID NO: 227) at its N-terminus (SP11). In some examples the N terminal portion includes KGI at its C-terminus and / or the C-terminal portion includes SGPP (SEQ ID NO: 229) at its N-terminus (SP12).

[0367] In some examples the N-terminal portion includes IPG at its C-terminus and / or the C-terminal portion includes SKGE (SEQ ID NO: 221) at its N-terminus (SP2). In some examples the N-terminal portion includes PPG at its C-terminus and / or the C-terminal portion includes SPGL (SEQ ID NO: 222) at its N-terminus (SP5). In some examples the N-terminal portion includes AGA at its C-terminus and / or the C-terminal portion includes SGFP (SEQ ID NO: 223) at its N-terminus (SP20). In some examples the N- terminal portion includes PGR at its C-terminus and / or the C-terminal portion includes SGVP (SEQ ID NO: 224) at its N-terminus (SP6). In some examples the N-terminal portion includes EKG at its C-terminus and / or the C-terminal portion includes SKGE (SEQ ID NO: 221) at its N-terminus (SP7). In some examples the N-terminal portion includes LLG at its C-terminus and / or the C-terminal portion includes SKGE (SEQ ID NO: 221) at its N-terminus (SP8). In some examples the N-terminal portion includes PGV at its C- terminus and / or the C-terminal portion includes SGPK (SEQ ID NO: 225) at its N-terminus (SP9). In some examples the N-terminal portion includes PGL at its C-terminus and / or the C-terminal portion includes SGQP (SEQ ID NO: 226) at its N-terminus (SP10). In some examples the N-terminal portion includes PGI at its C-terminus and / or the C-terminal portion includes SSIG (SEQ ID NO: 227) at its N-terminus (SP11). In some examples the N-terminal portion includes IKG at its C-terminus and / or the C-terminal portion includes SVGD (SEQ ID NO: 228) at its N-terminus (SP27). In some examples the N-terminal portion includes KGI at its C-terminus and / or the C-terminal portion includes SGPP (SEQ ID NO: 229) at its N- terminus (SP12).

[0368] Also disclosed is a pharmaceutical composition including an effective amount of the split intein systems described above, and a pharmaceutically acceptable carrier. Further disclosed is a method of treating Alport syndrome in a subject, including administering to the subject a therapeutically effective amount of the split intein system disclosed herein or the pharmaceutical composition disclosed herein, thereby treating the Alport syndrome in the subject. In some examples, the method includes delivering the split intein system to the kidney of the subject. In some examples, the method includes systemic administration, such as systemic injection or systemic infusion. In some examples, the method includes locally delivering the split intein system to the kidney of the subject. In some examples, locally delivering includes direct parenchymal injection, renal vein injection, and / or renal artery injection. In some examples, locally delivering includes direct pelvic injection or retrograde transureteral pelvic injection. In some examples, wherein the split intein system transduces at least one of mesangial cells, glomerular endothelial cells, parietal epithelial cells, podocytes, proximal tubule cells, Loop of Henle cells, distal tubule cells, collecting duct cells, fibroblasts, pericytes, or vascular smooth muscle cells.

[0369] In some examples, the method improves kidney function, delays onset of end stage renal disease, delays time to dialysis, delays time to renal transplant, and / or improves life expectancy of the subject. In some examples, the Alport syndrome is X-linked Alport syndrome. In some examples, the method further includes administering an effective amount of angiotensin II receptor blocker (ARB) and / or angiotensinconverting enzyme inhibitor. In some examples the method further includes administering one or more of candesartan, eprosartan, irbesartan, losartan, olmesartan, telmisartan, valsartan, benazepril, captopril, cilazapril, enalapril, fosinopril, lisinopril, moexipril, perindopril, ramipril, quinapril, and / or randolapril. In some examples, the ARB includes one or more of candesartan, captopril, eprosartan, irbesartan, losartan, moexipril, olmesartan, telmisartan, and / or valsartan. In some examples the angiotensin-converting enzyme inhibitor includes one or more of benazepril, cilazapril, enalapril, fosinopril, lisinopril, perindopril, ramipril, quinapril, and / or randolapril.

[0370] IV. Split Inteins and Systems

[0371] Interns can include naturally occurring or artificial polypeptide sequences capable of catalyzing a protein splicing reaction that excises the intein sequence from a precursor protein and joins the flanking sequences (N- and C-exteins) with a peptide bond. Some inteins include a homing endonuclease domain. Exemplary inteins can be found at inbase.ligsciss.com / iwai / . In a subset of inteins, the N- and C-terminal amino acid sequences are not directly linked via a peptide bond, and instead are separate fragments that can non-covalently re-associate, or reconstitute, into an intein that is functional for trans-splicing reactions. This is known as a split intein. For example, in cyanobacteria, DnaE, the catalytic subunit a of DNA polymerase HI, is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene is an exemplary N-intein, and the intein encoded by the dnaE-c gene is an exemplary C-intein.

[0372] In a particular example, an intein mutant is utilized in which the catalytic and second shell accelerator residues of the intein are maintained, and only non-catalytic residues, or residues outside the second shell are mutated. Second shell accelerator residues are those adjacent to the active site of the intein, which play a role in tuning the splicing activity of the inteins. Catalytic and accelerator residues of inteins of the DnaE family are described in Stevens et al., J. Am. Chem. Soc. (2016) 138:2162-65. Catalytic residues for other intein families, including DnaE, GyrA, GyrB, DnaB, TerL, gp41, IMPDH, are described in Shah & Muir, Chem. Sci (2014) 5:446-61. Methodologies to identify second shell accelerator residues are described in Stevens et al., J. Am. Chem. Soc. (2016) 138:2162-65. Non-limiting examples of intein pairs that may be used in accordance with the present disclosure include: Cfa DnaE intein, Gp41-1 intein, Ssp GyrB intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein and Cne Prp8 intein (e.g., as described in U.S. Pat. No. 8,394,604).

[0373] In some examples, the intein pair includes one or both members of a Npu DnaE intein pair. In some examples, the N-intein includes the amino acid sequence of SEQ ID NO: 26 or an amino acid sequence at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical thereto. In some examples, the C-intein includes the amino acid sequence of SEQ ID NO: 27 or an amino acid sequence at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical thereto identical thereto.

[0374] Disclosed herein is a fusion protein including an N-extein of an extein pair and an N-intein of an intein pair. This construct can have the architecture N-extein-N-intein (in N to C terminal order). Further disclosed herein is a fusion protein including a C-intein of the intein pair and a C-extein of the extein pair. This construct can have the architecture C-intein-C-extein (in N to C terminal order). In some examples, aspects of the first fusion protein and / or second fusion protein, such as the N-extein and N-intein are directly linked by peptide bonds. This may be described as being fused directly. In some examples, aspects of the first fusion protein and / or second fusion protein are linked by a linker sequence.

[0375] The nature of the linker will depend on the nature of the aspects to be linked. In a particular example, the linker is a peptide. In a particular example, the linker is a peptide having a length of 1, 2, 3, 4, 5, 10, 20, 50, 100 or more amino acid residues, such as 1 to 3 amino acid residues. In some examples, the linker is a flexible GS linker with stretches of Glycine and Serine residues. In one example, the N-terminus of the linker is linked to the C-terminus of the first aspect to be linked and the C-terminus of the linker is linked to the N-terminus of the second aspect to be linked through peptide bonds.

[0376] The protein to be reconstituted can be encoded by the N- and C-exteins. For example, the protein to be reconstituted could be a type IV collagenase alpha chain, such as COL4A3, COL4A4, or COL4A5. In a specific example, the protein to be reconstituted includes COL4A5.

[0377] Signal sequences can also be included. In some aspects, the first fusion protein includes, in N to C terminal order, a signal sequence, the N-terminal portion of a type IV collagenase alpha chain, and the N- intein. In more aspects, the second fusion protein includes, in N to C terminal order, a second signal sequence, the C-intein, and the C-terminal portion of the type IV collagenase alpha chain. In aspects, there is a linker between the first signal sequence and the N-terminal portion of the type IV collagen alpha chain, and / or between the N-terminal portion of the type IV collagen alpha chain and the N-intein. In more aspects, there is a linker between the second signal sequence and the C-intein, and / or between the C-intein and the C-terminal portion of the type IV collagen alpha chain. In some examples, the linker is a flexible GS linker with stretches of Glycine and Serine residues.

[0378] Disclosed herein are N- and C-exteins, that can be spliced together to form a mature type IV collagen alpha chain when the first and second nucleic acid molecules are expressed in mammalian cells. In some examples, the type IV collagen alpha chain includes the amino acid sequence of any one of SEQ ID NOs: 15-25 or an amino acid sequence at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical thereto. In a specific example, the type IV collagen alpha chain includes the amino acid sequence of SEQ ID NO: 15 or an amino acid sequence at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical thereto. In some examples, the type IV collagen alpha chain includes the amino acid sequence of SEQ ID NO: 13 or an amino acid sequence at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical thereto. In some examples, the type IV collagen alpha chain includes the amino acid sequence of SEQ ID NO: 14 or an amino acid sequence at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical thereto.

[0379] In some examples, the N- and C-terminal portions of the type IV collagen alpha chain together define the type IV collagen alpha chain sequence separated at a split point. In some examples, the split point can be within a collagenous domain of the type IV collagen alpha chain. In some examples the split point is within an interrupting non-collagenous region of the collagenous domain. In some examples, the split point is within a non-collagenous C-terminal domain. In some examples, the split point is at position 729 to position 1488 relative to SEQ ID NO: 15. In some examples, the split point is at position 697 to position 1488 relative to SEQ ID NO: 15. In some examples, the split point is at position 697 to position 1136 relative to SEQ ID NO: 15. In some examples, the split point is at position 879 to position 996 relative to SEQ ID NO: 15. In some examples, the split point is at position 942 to position 996 relative to SEQ ID NO: 15. In some examples, the split point is at position 565 to position 1111 relative to SEQ ID NO: 15. In some examples, the split point is at position 188 to position 1488 relative to SEQ ID NO: 15. In some examples, the split point is at position 573 to position 1108 relative to SEQ ID NO: 13. In some examples, the split point is at position 235 to position 1499 relative to SEQ ID NO: 13. In some examples, the split point is at position 585 to position 1116 relative to SEQ ID NO: 14. In some examples, the split point is at position 200 to position 1492 relative to SEQ ID NO: 14.

[0380] In some examples, the construct, the N-intein, the C-intein, the N-extein, or the C-extein include an epitope tag, such as a FLAG tag (such as SEQ ID NO: 6). In alternate examples, the construct, the N-intein, the C-intein, the N-extein, or the C-extein does not include an epitope tag, such as a FLAG tag (such as SEQ ID NO: 6). In some examples, the construct, the N-intein, the C-intein, the N-extein, or the C-extein does not include a 3xFLAG tag. In some examples, the construct, the N-intein, the C-intein, the N-extein, or the C-extein does not include an HA tag. In some examples, the first fusion protein and / or the second fusion protein include an epitope tag, such as a FLAG tag. In some examples, the first fusion protein and / or the second fusion protein include a 3x FLAG tag. In some examples, the first fusion protein and / or the second fusion protein include a HA tag. In alternate examples, the first fusion protein and / or the second fusion protein do not include a FLAG tag. In some examples, the first fusion protein and / or the second fusion protein do not include a 3xFLAG tag. In some examples, the first fusion protein and / or the second fusion protein do not include a HA tag. In some examples, the FLAG tag is C-terminal to an N-intein. In some examples, the 3xFLAG tag is C-terminal to an N-intein. In some examples, the HA tag is C-terminal to an N-intein. In some examples the FLAG tag is C-terminal to a C-extein. In some examples the 3x FLAG tag is C-terminal to a C-extein. In some examples the HA tag is C-terminal to a C-extein.

[0381] In some examples, the first nucleic acid molecule encodes a first signal sequence. In some examples the N-extein, includes the first signal sequence N-terminal to an N-terminal portion of a type IV collagen alpha chain. In other examples the N-extein, includes the first signal sequence C-terminal to an N-terminal portion of a type IV collagen alpha chain. In some examples, the first signal sequence falls within the N- terminal portion of a type IV collagen alpha chain. In some examples, the first nucleic acid molecule encodes a fusion protein including, in N to C terminal order, a first signal sequence, the N-extein of the extein pair and the N-intein of the intein pair. In some examples, the first nucleic acid molecule encodes a fusion protein including, in N to C terminal order, the N-extein of the extein pair, the first signal sequence, and the N-intein of the intein pair. In some examples, the first nucleic acid molecule encodes a fusion protein including, in N to C terminal order, the N-extein of the extein pair, the N-intein of the intein pair, and the first signal sequence.

[0382] In some examples, the second nucleic acid protein encodes a second signal sequence. In some examples the C-intein includes the second signal sequence at its N-terminal aspect. In some examples the C-intein includes the second signal sequence at its C-terminal aspect. In some examples, the second signal sequence is within the C-intein. In some examples the C-extein includes the second signal sequence at its N- terminal aspect. In some examples the C- extein includes the second signal sequence at its C-terminal aspect. In some examples, the second signal sequence is within the C- extein. In some examples, the second nucleic acid protein encodes a fusion protein including, in N to C terminal order, the second signal sequence, the C-intein of the intein pair, and the C-extein of the extein pair. In some examples, the second nucleic acid protein encodes a fusion protein including, in N to C terminal order the C-intein of the intein pair, the second signal sequence, and the C-extein of the extein pair. In some examples, the second nucleic acid protein encodes a fusion protein including, in N to C terminal order the C-intein of the intein pair, the C-extein of the extein pair, and the second signal sequence.

[0383] In some examples, the first signal sequence and / or the second signal sequence is a collagen signal sequence or a chymotrypsin signal sequence. In some examples, the first signal sequence and the second signal sequence are different. In some examples, the first signal sequence and the second signal sequence are the same. In some examples the collagen signal sequence includes the amino acid sequence of SEQ ID NO: 4 or an amino acid sequence at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical thereto. In some examples, the collagen signal sequence is the native COL4A5 signal sequence. In some examples, the collagen signal sequence is the native COL4A4 signal sequence. In some examples, the collagen signal sequence is the native COL4A3 signal sequence. In some examples the chymotrypsin signal sequence includes the amino acid sequence of SEQ ID NO: 5 or an amino acid sequence at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical thereto.

[0384] The inclusion of the WPRE element as a 3' UTR in transgenes can enhance transgene expression from gene delivery vectors. Although the mechanisms underlying the WPRE-mediated gene expressionenhancing effect are not fully understood, they include increased stabilization and nuclear export of mRNA, leading to elevated steady-state mRNA levels, reduction of readthrough transcription that improves transcriptional termination, and enhanced translation of mRNA into protein products. WPRE3, a truncated version of the traditional 0.6-kb WPRE, has been developed to maximize the AAV's packaging capacity for holding a transgene expression unit while leveraging the transgene expression-enhancing effect of WPRE (Choi et al., Mol Brain (2014) 7: 17). WPRE3 is 0.35 kb shorter than WPRE and comprises only minimal gamma and alpha elements among the three regulatory elements (alpha, beta, and gamma) found in WPRE. Despite this truncation, WPRE3 retains an enhancing effect comparable to that of WPRE

[0385] In some examples, the first nucleic acid molecule includes WPRE and / or WPRE3. In some examples, the WPRE and / or the WPRE3 is 3’ of the nucleotide sequence encoding a fusion protein including an N-extein of an extein pair and an N-intein of an intein pair. In some examples, the WPRE and / or the WPRE3 is 5’ of the nucleotide sequence encoding a fusion protein including an N-extein of an extein pair and an N-intein of an intein pair. In some examples, the second nucleic acid molecule includes WPRE and / or WPRE3. In some examples, the WPRE and / or the WPRE3 is 3’ of the nucleotide sequence encoding a fusion protein including a C-intein of the intein pair, and a C-extein of the extein pair. In some examples, the WPRE and / or the WPRE3 is 5’ of the nucleotide sequence encoding a fusion protein including a C-intein of the intein pair, and a C-extein of the extein pair.

[0386] In some examples, WPRE includes the nucleotide sequence of SEQ ID NO: 120, or a sequence at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% identical thereto. In some examples, WPRE3 includes the nucleotide sequence of SEQ ID NO: 121, or a sequence at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% identical thereto.

[0387] In some examples, the first fusion protein is encoded by a first nucleic acid molecule. In some examples, the second fusion protein is encoded by a second nucleic acid molecule. Further disclosed herein is a first AAV vector including the first nucleic acid molecule. Also disclosed herein is a second AAV vector including the second nucleic acid molecule. Administering one or both nucleic acid molecules, and / or administering one or both AAV vectors can result in reconstitution of the protein to be reconstituted. The goal of reconstituting these proteins may be to treat diseases associated with mutations in their encoding genes, by providing correct versions of such genes, for example treating Alport syndrome by providing a correct version of COL4A5.

[0388] In some examples, the split intein system includes a first nucleic acid molecule encoding the first fusion protein and a second nucleic acid molecule encoding the second fusion protein.

[0389] Disclosed herein is a split intein system including a first nucleic acid molecule encoding a fusion protein including an N-extein of an extein pair and an N-intein of an intein pair. The split intein system also includes a second nucleic acid molecule encoding a fusion protein including a C-intein of the intein pair, and a C-extein of the extein pair. It is disclosed herein that incorporating signal sequences into the first nucleic acid molecule or the second nucleic acid molecule can affect localization of their respective fusion proteins.

[0390] In some examples, the first fusion protein is encoded by a first nucleic acid molecule. In some examples, the first nucleic acid molecule encodes a fusion protein including a first signal sequence. In some examples the N-extein includes the first signal sequence N-terminal to an N-terminal portion of a type IV collagen alpha chain. In other examples, the N-extein includes the first signal sequence C-terminal to an N- terminal portion of a type IV collagen alpha chain. In some examples, the first signal sequence falls within the N-terminal portion of a type IV collagen alpha chain. In some examples, the first fusion protein includes in N to C terminal order, the first signal sequence, the N-extein of the extein pair and the N-intein of the intein pair. In some examples, the first fusion protein includes in N to C terminal order, the N-extein of the extein pair, the first signal sequence, and the N-intein of the intein pair. In some examples, the first fusion protein includes in N to C terminal order, the N-extein of the extein pair, the N-intein of the intein pair, and the first signal sequence.

[0391] In some examples, the second fusion protein is encoded by a second nucleic acid molecule. In some examples, the second nucleic acid molecule encodes a fusion protein including a second signal sequence. In some examples the C-intein includes the second signal sequence at its N-terminal aspect. In some examples the C-intein includes the second signal sequence at its C-terminal aspect. In some examples, the second signal sequence is within the C-intein. In some examples the C-extein includes the second signal sequence at its N-terminal aspect. In some examples the C- extein includes the second signal sequence at its C- terminal aspect. In some examples, the second signal sequence is within the C- extein. In some examples, the second fusion protein includes in N to C terminal order, the second signal sequence, the C-intein of the intein pair, and the C-extein of the extein pair. In some examples, the second fusion protein includes in N to C terminal order the C-intein of the intein pair, the second signal sequence, and the C-extein of the extein pair. In some examples, the second fusion protein includes in N to C terminal order the C-intein of the intein pair, the C-extein of the extein pair, and the second signal sequence. V. Recombinant Adeno-Associated Viral Vectors

[0392] AAV vectors can transduce cells of the kidney, such as of podocytes, mesangial cells, tubular epithelial cells, collecting duct cells, interstitial cells, and vascular cells. However, AAV vectors have size limitations such that full length collagen IV protein cannot be delivered in a single AAV vector. Disclosed herein are engineered, specific AAV vectors that deliver a split intein system, such that full length collagen is produced upon introduction of the AAV vectors into a mammalian cell. These AAV vectors are of use for delivering the disclosed split intein systems, and can be used for the treatment of Alport syndrome in a subject. In some aspects, the split intein systems includes (i) a first adeno-associated viral (AAV) vector comprising the first nucleic acid molecule; and (ii) a second AAV vector comprising the second nucleic acid molecule. The first AAV vector and the second AAV vector can be the same serotype. The first AAV vector and the second AAV vector can be different serotypes.

[0393] The AAVs of use in the methods disclosed herein may be derived from various serotypes, including combinations of capsid serotypes and viral genome serotypes (e.g., ''pseudotyped" AAV) or from various genomes (e.g., single-stranded or self-complementary).

[0394] AAV belongs to the family Parvoviridae and the genus Dependovirus. AAV is a small, nonenveloped virus that packages a linear, single-stranded DNA genome. Both sense and antisense strands of AAV DNA are packaged into AAV capsids with equal frequency. In some aspects the AAV DNA includes a first nucleic acid encoding a fusion protein including an N-extein of an extein pair and an N-intein of an intein pair. In some aspects the AAV DNA includes a second nucleic acid encoding a fusion protein a C- intein of the intein pair, and a C-extein of the extein pair. In some aspects, the N- and C-exteins are spliced together to form the mature type IV collagen alpha chain when the first and second nucleic acid molecules are expressed in mammalian cells, such as after a first AAV particle including the first nucleic acid and a second AAV particle including the second nucleic acid contact the mammalian cell.

[0395] The AAV genome is characterized by two inverted terminal repeats (ITRs) that flank two open reading frames (ORFs). In the AAV2 genome, for example, the first 125 nucleotides of the ITR are a palindrome, which folds upon itself to maximize base pairing and forms a T-shaped hairpin structure. The other 20 bases of the ITR, called the D sequence, remain unpaired. The ITRs are cis-acting sequences important for AAV DNA replication; the ITR is the origin of replication and serves as a primer for second- strand synthesis by DNA polymerase. The double-stranded DNA formed during this synthesis, which is called replicating-form monomer, is used for a second round of self-priming replication and forms a replicating-form dimer. These double-stranded intermediates are processed via a strand displacement mechanism, resulting in single-stranded DNA used for packaging and double-stranded DNA used for transcription. Located within the ITR are the Rep binding elements and a terminal resolution site (TRS). These features are used by the viral regulatory protein Rep during AAV replication to process the doublestranded intermediates. In addition to their role in AAV replication, the ITR is also involved in AAV genome packaging, transcription, negative regulation under non-permissive conditions, and site-specific integration (Daya and Berns, Clin Microbiol Rev 21(4):583-593, 2008). In some aspects, these elements are included in the AAV vector.

[0396] The left ORF of AAV contains the rep gene, which encodes four proteins - Rep78, Rep 68, Rep52 and Rep40. The right ORF contains the cap gene, which produces three viral capsid proteins (VP1, VP2 and VP3). The right ORF also contains two additional frame-shifted ORFs for the membrane-associated accessory protein (MAAP) and the assembly-activating protein (AAP). The AAV capsid contains 60 viral capsid proteins arranged into an icosahedral symmetry. VP1, VP2 and VP3 are present in a 1: 1: 10 molar ratio (Daya and Berns, Clin Microbiol Rev 21(4):583-593, 2008). In some aspects, these elements are included in the AAV vector. The AAV vector is packaged in the capsid proteins, which can target specific cell types, such as podocytes, mesangial cells, tubular epithelial cells, collecting duct cells, interstitial cells, and vascular cells.

[0397] In some aspects, a recombinant adeno-associated virus (rAAV) is generated having a capsid of interest, and can be used in the disclosed methods. In AAV, the capsid includes VP1, VP2, and VP3. In some non-limiting examples, to produce a vector, a host cell which can be cultured that contains a nucleic acid sequence encoding an adeno-associated virus (AAV) capsid protein of interest, or fragment thereof: a functional rep gene: a minigene composed of, at a minimum, AAV inverted terminal repeats (ITRs) and a transgene, such as a transgene encoding a therapeutic protein; and sufficient helper functions to permit packaging in the capsid protein. The components required to be cultured in the host cell to package an AAV minigene in an AAV capsid may be provided to the host cell in trans. Alternatively, any one or more of the components (e.g., minigene, rep gene sequences, cap gene sequences, and / or helper functions) may be provided by a stable host cell which has been engineered to contain one or more of the components. In some aspects, a stable host cell will contain the component(s) under the control of an inducible promoter. However, the component(s) can be under the control of a constitutive promoter. Promoters of use include, but are not limited to, the CMV promoter, the CAG promoter, the CB promoter, the SV40 promoter, the ubiquitin promoter, the EFloc promoter, the GAPDH promoter, the PGK promoter, the RSV promoter, and the P-actin promoter.

[0398] In still another alternative, a selected stable host cell may contain selected component(s) under the control of a constitutive promoter and other selected component(s) under the control of one or more inducible promoters. For example, a stable host cell may be generated which is derived from 293 cells (which contain El helper functions under the control of a constitutive promoter), but which contains the Rep and / or VP proteins under the control of inducible promoters.

[0399] The minigene, rep gene sequences, cap gene sequences, and helper functions required for producing a rAAV can be delivered to the packaging host cell in the form of any genetic element which transfer the sequences carried thereon. The selected genetic element may be delivered by any suitable method, including those described herein. The methods used to construct vectors include genetic engineering, recombinant engineering, and synthetic techniques. See, e.g., Sambrook et al, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Press, Cold Spring Harbor, N.Y. Similarly, methods of generating rAAV virions and the selection of a suitable method is not a limitation on the present invention. See, e.g.. K. Fisher et al, J. Virol., 70:520-532 (1993) and U.S. Pat. No. 5,478,745.

[0400] The AAV vector can include a promoter operably linked to a sequence encoding a fusion protein, such as a first fusion protein including an N-extein of an extein pair and an N-intein of an intein pair, or such as a second fusion protein encoding a C-intein of the intein pair and a C-extein of the extein pair. In some aspects, the promoter is a ubiquitous promoter such as the CMV promoter, the CAG promoter, the CB promoter, the SV40 promoter, the ubiquitin promoter, the EFla promoter, the GAPDH promoter, the PGK promoter, the RSV promoter, and the fl-actin promoter. In some aspects, the promoter is a kidney -specific or cell type-specific promoter, such as the podocin (NPHS2) promoter, the nephrin (NPHS1) promoter, the synaptopodin promoter, the WT 1 promoter, the podocalyxin prompter, the PAX8 promoter, the gammaglutamyltransferase (GGT) promoter, the sodium-glucose cotransporter-2 (SGLT2) promoter, the kidney androgen-regulated protein (KAP) promoter, the megalin promoter, the carbonic anhydrase II (CAII) promoter, the organic anion transporter 1 (OAT1, SLC22A6) promoter, the sodium-phosphate cotransporter 2a (NPT2a) promoter, the E-cadherin promoter, the kidney-specific cadherin promoter, the Na-K-Cl cotransporter 2 (NKCC2, SLC12A1) promoter, the WNK1 promoter, the WNK4 promoter, the v-type proton ATPase subunit Bl (ATP6V1B1) promoter, the thiazide-sensitive Na-Cl cotransporter (NCC) promoter, the calcium-sensing receptor (CaSR) promoter, the epithelial sodium channel (ENaC) promoter, the uromodulin promoter, the aquaporin 1 (AQP1) promoter, the aquaporin 2 (AQP2) promoter, the vasopressin receptor 2 (V2R) promoter, the transient receptor potential vanilloid 4 (TRPV4) promoter, the renin promoter, the parathyroid hormone receptor promoter, the smooth muscle alpha-actin (ACTA2) promoter, the COL4A1 promoter, the platelet-derived growth factor receptor beta (PDGFR[3) promoter, the smooth muscle protein 22-a (SM22a) promoter, the smooth muscle myosin heavy chain (SM-MHC) promoter, the neuron-glial antigen 2 (NG2, chondroitin sulphate proteoglycan 4 (CSPG4)) promoter, the desmin promoter, the regulator of G-protein signaling-5 (RGS5) promoter, the fibroblast-specific protein 1 (FSP1) promoter, the transcription factor 21 (TCF21) promoter, the promoter the COL1A2 promoter, the COL3A1 promoter, the vimentin promoter, the TIE2 promoter, the von Willebrand factor (vWF) promoter, the CD31 promoter and the VE-cadherin promoter. In some aspects, the first nucleic acid molecule and / or the second nucleic acid molecule is operably linked to a CAG promoter.

[0401] In some aspects, the AAV genome is modified to include a gene encoding a selectable marker, which includes, but are not limited to, a protein whose expression can be readily detected such as a fluorescent or luminescent protein or an enzyme that acts on a substrate to produce a colored, fluorescent, or luminescent substance ("detectable markers"). There are other genes of use, such as genes that encode drug resistance or provide a function that can be used to purify cells. Selectable markers include neomycin resistance gene (neo), puromycin resistance gene (puro), guanine phosphoribosyl transferase (gpt), dihydrofolate reductase (DHFR), adenosine deaminase (ada), puromycin-N-acetyltransferase (PAC), hygromycin resistance gene (hyg), multidrug resistance gene (mdr), thymidine kinase (TK), hypoxanthine - guanine phosphoribosyltransferase (HPRT), and hisD gene. Detectable markers include green fluorescent protein (GFP) blue, sapphire, yellow, red, orange, and cyan fluorescent proteins and variants of any of these. Luminescent proteins such as luciferase (e.g., firefly or Renilla luciferase) are also selectable makers.

[0402] In some aspects, the split intein system disclosed herein may include two or more AAV vectors. Different A A Vs may exhibit different tissue tropism, which refers to the preference or ability of an AAV to infect specific types of tissues or cells within an organism. Tissue tropism of an AAV is at least partially determined by the viral capsid proteins, which bind receptors on the surface of target cells. Therefore, issue tropism can therefore be programmed to target one or more desired tissues (such as the kidney tissue) by varying the sequences or structures of the capsid proteins.

[0403] To date over 100 AAV variants have been isolated from adenovirus stocks, or human or nonhuman primate tissues. Among them, AAV2, AAV3, AAV5, AAV6 were discovered in human cells, while AAV1, AAV4, AAV7, AAV8, AAV9, AAV10 (AAVrhlO), AAV11, AAV12 in nonhuman primate samples. Genome divergence among different serotypes is most concentrated on variable regions (VRs) of virus capsid, which might determine their tissue tropism. Besides virus capsid, tissue tropisms of AAV vectors are also influenced by cell surface receptors, cellular uptake, intracellular processing, nuclear delivery of the vector genome, uncoating, and second-strand DNA conversion.

[0404] AAV1 has been reported to exhibit natural tropism towards central nervous system (CNS), pancreas, heart, retina (e.g., retinal pigment epithelium (RPE)), and muscle. AAV2 has been reported to exhibit natural tropism towards neurons, vascular smooth muscle cells, hepatocytes, muscle, CNS, retina (e.g., RPE, photoreceptor cells), liver, and kidney. AAV3 has been reported to exhibit natural tropism towards retina (e.g., RPE), lung, liver, and heart. AAV4 has been reported to exhibit natural tropism towards CNS, retina (e.g., RPE), lung, and pancreas. AAV5 has been reported to exhibit natural tropism towards CNS, lung, retina (e.g., RPE, photoreceptor cells), and pancreas. AAV6 has been reported to exhibit natural tropism towards CNS, lung, liver, heart, and muscle. AAV7 has been reported to exhibit natural tropism towards liver, and muscle. AAV8 has been reported to exhibit natural tropism towards CNS, heart, retina (e.g., RPE, photoreceptor cells), liver, pancreas, and muscle. AAV9 has been reported to exhibit natural tropism towards CNS, lung, liver, heart, and muscle. The natural tropism of AAV may be utilized to concentrate AAV transduction to one or more specific tissues or cell types. To limit AAV tropism to specific tissues or cells in vivo, or to enhance transduction rate of specific tissues or cells in vivo, mosaic or chimeric AAV vectors can be developed by engineering AAV capsids to preferentially target or interact with specific cell types.

[0405] In some examples, AAV vectors that are capable of transducing, or preferentially transduce the kidney or kidney cells (e.g., mesangial cells, glomerular endothelial cells, parietal epithelial cells, podocytes, proximal tubule cells, Loop of Henle cells, distal tubule cells, collecting duct cells, fibroblasts, pericytes, vascular smooth muscle cells, or any combination thereof). In some examples, the split intein system transduces at least one, at least two at least three, or at least four of: podocytes, mesangial cells, glomerular endothelial cells, parietal epithelial cells, proximal tubule cells, Loop of Henle cells, distal tubule cells, collecting duct cells, fibroblasts, kidney fibroblasts, pericytes, kidney pericytes, vascular smooth muscle cells, and / or kidney vascular smooth muscle cells. In a specific example the split intein system transduces tubular epithelial cells.

[0406] For example, according to the disclosure in Provisional Patent Application No. 63 / 527,248, six AAV capsids (AAV-KP1, AAV-KP2, AAV-KP3, AAV-DJ, AAV2G9, AAV2.7m8) were identified that transduce the kidney more efficiently than AAV9 by renal pelvis injection. The present application discloses the surprising effect of AAV9 and AAV-KP1 capsids in transducing kidney cells when administered systemically.

[0407] The vectors disclosed herein can contain nucleic acid sequences encoding an intact AAV capsid which may be from a single AAV serotype. The vectors can include AAV1, AAV l_9mtl00, AAV l_9mt30, AAVl_9mt76, AAV2, AAV2G9, AAV2i8, AAV2retro, AAV2R585E, AAV2R585E9_2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV9AA272, AAV9AA22, AAV9W22A, AAV 10, AAV11, AAV2.7m8, AAVAnc80, AAVbb.2, AAV-DJ, AAVHN1, AAVHN2, AAVHN3, AAVhu.ll, AAVhu.13, AAVhu.37, AAV-KP1, AAV-KP2, AAV-KP3, AAVLK03, AAVNP40, AAVNP59, AAVPHP.B, AAVPHP.eB, AAVPHP.S, AAVpol, AAVrh.8, AAVrh.10, AAVrh.20, AAVrh.43, and AAVShHIO or a variant thereof exhibiting at least 70% identity thereto. The sequence identify can be about 95%, 96%, 97% 98%, or 99% (see Earley et al., J Virol (2017) 91:e01980-16; Shen et al., J Biol Chem (2013) 288:28814-23; Asokan et al., Nat Biotechnl (2010) 28:79-82; Tervo etal., Neuron (2016) 92: 372-82; Adachi et al., Nat Commun (2014) 5:3075; Dalkara et al., Sci Transl Med (2013) 5:89ral76; Zinn et al., Cell Rep (2015) 12:1056-1068; Gao et al., J Virol (2004) 78:6381-88; Grimm et al., J Virol (2008) 82:5887-11; Pekrun et al., JCI insight (2019) 4; Lisowski et al.. Nature (2014) 506:382-86; Paulk et al., Mol Ther (2018) 26:289-303; Deverman et al., Nat Biotechnol (2016) 34:204-09; Chan et al., Nat Neurosci (2017) 20:1172-79; Bello et al., Gene Therapy (2009) 16:1320-28; Klimczak et al., PLoS One (2009) 4: e7467; Issa et al., Cells (2023) 12:785; U.S. Pat. App. No. 63 / 527,248). In some examples the first and / or the second AAV vector includes AAV-KP2, AAV-KP3, AAV-DJ, AAV2G9, or AAV2.7m8. In some examples the first and / or the second AAV vector is AAV9. In other aspects, the first and / or the second AAV vector is AAV-KP1.

[0408] The AAV vector can transduce one or more kidney cell types, such as cells constituting a nephron, cells in the renal interstitium and cells constituting the vasculature in the kidney, These cells include, but are not limited to, mesangial cells, glomerular endothelial cells, podocytes, parietal epithelial cells in Bowman’ s capsule, proximal tubule cells, Loop of Henle cells, distal tubule cells, collecting duct cells, fibroblasts, myofibroblasts, pericytes, vascular endothelial cells, and vascular smooth muscle cells. In the method described herein. The AAV vector includes a capsid protein including a capsid sequence selected from the group of AAV1, AAVl_9mtl00, AAVl_9mt30, AAVl_9mt76, AAV2, AAV2G9, AAV2i8, AAV2retro, AAV2R585E, AAV2R585E9_2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV9AA272, AAV9AA22, AAV9W22A, AAV10, AAV11, AAV2.7m8, AAVAnc80, AAVbb.2, AAV-DJ, AAVHN1, AAVHN2, AAVHN3, AAVhu.l l, AAVhu.13, AAVhu.37, AAV-KP1, AAV-KP2, AAV-KP3, AAVLK03, AAVNP40, AAVNP59, AAVPHP.B, AAVPHP.eB, AAVPHP.S, AAVpol, AAVrh.8, AAVrh.10, AAVrh.20, AAVrh.43, and AAVShHIO capsid protein or a variant thereof exhibiting at least 70% identity thereto. The sequence identify can be about 95%, 96%, 97% 98%, or 99%. In some examples the first and / or the second AAV vector includes a capsid protein from AAV9. In other aspects, the first and / or the second AAV vector includes a capsid protein from AAV-KP1.

[0409] Disclosed here in the use of recombinant AAV vectors for the delivery of a split protein or into a cell, for example a split COL4A5. The N-terminal portion of the protein and the C-terminal portion of the protein are delivered by separate AAV vectors or particles into the same cell, if the full-length protein exceeds the packaging limit of AAV. Therefore, disclosed herein is a composition for delivering the split protein into a cell (e.g., a mammalian cell, a human cell). The AAV particles of the present disclosure can include a AAV vector (i.e., a recombinant genome of the AAV) encapsidated in the viral capsid proteins. The use of AAV vectors and particles is disclosed in U.S. Application No. 63 / 462,492, which is incorporated by reference in its entirety.

[0410] In some examples, the AAV vector includes: (1) a heterologous nucleic acid region including encoding a first or second fusion protein, (2) one or more nucleotide sequences including a sequence that facilitates expression of the heterologous nucleic acid region (e.g., a promoter), and (3) one or more nucleic acid regions including a sequence that facilitate integration of the heterologous nucleic acid region (optionally with the one or more nucleic acid regions including a sequence that facilitates expression) into the genome of a cell. In some examples, viral sequences that facilitate integration comprise Inverted Terminal Repeat (ITR) sequences.

[0411] In some examples, a composition including the AAV particles (in any form contemplated herein) further includes a pharmaceutically acceptable carrier. In some examples, the composition is formulated in appropriate pharmaceutical vehicles for administration to human or animal subjects.

[0412] Disclosed herein is a composition including a first AAV vector. In some examples, the first AAV vector includes a first signal sequence fused at its C-terminus to a first nucleotide sequence encoding an N- terminal portion of a type IV collagen alpha chain fused at its C-terminus to an N-intein (N-int). In some examples, the composition includes a second AAV vector. In some examples, the second AAV vector includes a second nucleotide sequence encoding a second signal sequence fused at its C-terminus to a C- intein (C-int), the C-int fused at its C-terminus to the N-terminus of a C-terminal portion of the type IV collagen alpha chain. In some examples, the type IV collagen alpha chain is COL4A3, COL4A4, or COL4A5. In some examples, the N-terminal portion and the C-terminal portion can be joined to form the complete type IV collagen alpha chain.

[0413] VI. Methods of Administration and Treatment

[0414] The split intein system disclosed herein can be delivered to any part of a body, such as any organ (e.g., the kidney, the eye, and the ear), any tissue, or any cell of the body, through any delivery method, such as systemic delivery and / or local delivery. The split intein system disclosed herein can also be delivered ex vivo. For example, cells may be isolated from a subject, modified or transduced by the system outside the body of the subject, and reintroduced into the subject. In some aspects, the delivery may not intentionally target a specific organ, tissue, or cell. For example, a systemic delivery delivers the split intein system through a circulatory system of a body, which can reach every part of the body.

[0415] In some other aspects, the delivery may target one or more specific types of cells, such as those within the same tissue or organ. In some examples, such targeted delivery may be achieved through local delivery. In some other examples, such targeted delivery may be achieved through systemic delivery, and choosing or engineering a vector that targets or transduces only the intended cell types. A targeted delivery may be advantageous in dose-sparing and reducing toxicity or side effects.

[0416] The split intein systems disclosed herein are administered in sufficient amounts to transduce the cells of one or more desired tissues and to provide sufficient levels of gene transfer and expression without undue adverse effects. Administration can be effected in one dose, continuously or intermittently throughout the course of a treatment.

[0417] Any administration route, systemic or local, can be used to deliver the split intein system. In some examples, the split intein system may be parenterally administered by injection, infusion or implantation. In some examples, the system may be administered intravenously, intraventricularly, intra-aurally, intra- ocularly, or peri-ocularly, orally, parenterally, subcutaneously, intramuscularly, intrapleurally, topically, intralymphatically; such administration may also be intra-arterial, intracardiac, subventricular, sub-retinal, intravitreal, intraarticular, intraparenchymal, intrapelvic, or any combination thereof.

[0418] In aspects, delivery is such that the first nucleic acid molecule and the second nucleic acid are provided any ratio, such as 10: 1, 5: 1, 2.5: 1, 2: 1, 1 : 1, 0.5: 1, 0.25:1, 0.17: 1, 0.1: 1, 0.050: 1, 0.025:1, 0.017: 1, 0.013: 1, 0.010: 1, 0.005:1, 0.001 : 1, or a range between any two of the preceding values. In some examples, delivery is such that that a first adeno-associated viral (AAV) vector including the first nucleic acid molecule and / or a second AAV vector including the second nucleic acid molecule is provided at a ratio of 10: 1, 5: 1, 2.5: 1, 2: 1, 1: 1, 0.5: 1, 0.25:1, 0.17: 1, 0.1: 1, 0.050: 1, 0.025:1, 0.017: 1, 0.013: 1, 0.010: 1, 0.005: 1, 0.001 : 1, or a range between any two of the preceding values.

[0419] Pharmaceutical forms suitable for administration include sterile aqueous solutions or dispersions, and sterile powders for the extemporaneous preparation of sterile injectable solutions or dispersions. Dispersions may also be prepared in glycerol, liquid polyethylene glycols, and mixtures thereof and in oils. Under ordinary conditions of storage and use, these preparations may contain a preservative to prevent the growth of microorganisms. In many cases the form is sterile and fluid to the extent that easy syringeability exists. The forms may also be stable under the conditions of manufacture and storage, and may be preserved against the contaminating action of microorganisms, such as bacteria and fungi.

[0420] The carrier may be a solvent or dispersion medium including, for example, water, ethanol, polyol (e.g., glycerol, propylene glycol, and liquid polyethylene glycol, and the like), suitable mixtures thereof, and / or vegetable oils. Proper fluidity may be maintained, for example, by the use of a coating, such as lecithin, by the maintenance of the required particle size in the case of dispersion and by the use of surfactants. The prevention of the action of microorganisms may be effected by various antibacterial and antifungal agents, for example, parabens, chlorobutanol, phenol, sorbic acid, thimerosal, and the like. In some examples, it will be preferable to include isotonic agents, such as sugars or sodium chloride. If prolonged absorption of the compositions is desired, agents delaying absorption, for example, aluminum monostearate and gelatin, may be included in the compositions.

[0421] For administration of an aqueous solution, the solution may be suitably buffered, if necessary, and the liquid diluent first rendered isotonic with sufficient saline or glucose. These particular aqueous solutions are especially suitable for intravenous, intramuscular, subcutaneous and intraperitoneal administration. For example, one dosage may be dissolved in 1 mL of isotonic NaCl solution and either added to 1000 mL of hypodermoclysis fluid or injected at the proposed site of infusion, (see for example, “Remington's Pharmaceutical Sciences” 15th Edition, pages 1035-1038 and 1570-1580). Some variation in dosage will necessarily occur depending on the condition of the host. The person responsible for administration will, in any event, determine the appropriate dose for the individual host. Sterile solutions for parenteral administration may be prepared by including the split intein system at an effective amount in the appropriate solvent with various other ingredients, as required or desired, followed by filtered sterilization. In aspects, a pharmaceutical composition includes an effective amount of the split intein system, as disclosed herein, and a pharmaceutically acceptable carrier. In other aspects, a pharmaceutical composition includes an effective amount of the first nucleic acid molecule, optionally in an AAV vector, as disclosed herein, and a pharmaceutically acceptable carrier. In more aspects, a pharmaceutical composition includes an effective amount of the second nucleic acid molecule, optionally in an AAV vector, as disclosed herein, and a pharmaceutically acceptable carrier.

[0422] Pharmaceutical compositions can be produced that contain the AAV and a pharmaceutically acceptable excipient. Such excipients include any pharmaceutical agent that does not itself induce the production of antibodies harmful to the individual receiving the composition, and which may be administered without undue toxicity. Pharmaceutically acceptable excipients include, but are not limited to, liquids such as water, saline, glycerol and ethanol. Pharmaceutically acceptable salts can be included therein, for example, mineral acid salts such as hydrochlorides, hydrobromides, phosphates, sulfates, and the like; and the salts of organic acids such as acetates, propionates, malonates, benzoates, and the like. Additionally, auxiliary substances, such as wetting or emulsifying agents, pH buffering substances, and the like, may be present in such vehicles. A thorough discussion of pharmaceutically acceptable excipients is available in Remington’s Pharmaceutical Sciences, 23rd Edition, Academic Press, Elsevier, (2020).

[0423] In some aspects, the excipients confer a protective effect on the AAV virion such that loss of AAV virions, as well as transduceability resulting from formulation procedures, packaging, storage, transport, and the like, is minimized. These excipient compositions are therefore considered virion-stabilizing in the sense that they provide higher AAV virion titers and higher transduceability levels than their non-protected counterparts, as measured using standard assays, see, for example, Published U.S. Application No. 2012 / 0219528. These compositions therefore demonstrate enhanced transduceability levels as compared to compositions lacking the particular excipients described herein, and are therefore more stable than their nonprotected counterparts.

[0424] Exemplary excipients that can used to protect an AAV virion from activity degradative conditions include, but are not limited to, detergents, proteins, e.g., ovalbumin and bovine serum albumin, amino acids, e.g., glycine, polyhydric and dihydric alcohols, such as but not limited to polyethylene glycols (PEG) of varying molecular weights, such as PEG-200, PEG-400, PEG-600, PEG-1000, PEG-1450, PEG-3350, PEG- 6000, PEG-8000 and any molecular weights in between these values, with molecular weights of 1500 to 6000 preferred, propylene glycols (PG), sugar alcohols, such as a carbohydrate, preferably, sorbitol. The detergent, when present, can be an anionic, a cationic, a zwitterionic or a nonionic detergent. An exemplary detergent is a nonionic detergent. One suitable type of nonionic detergent is a sorbitan ester, e.g., polyoxyethylenesorbitan monolaurate (TWEEN®-20), polyoxyethylenesorbitan monopalmitate (TWEEN®- 40), polyoxyethylenesorbitan monostearate (TWEEN®-60), polyoxyethylenesorbitan tristearate (TWEEN®- 65), polyoxyethylenesorbitan monooleate (TWEEN®-80), polyoxyethylenesorbitan trioleate (TWEEN®- 85), polyoxyethylene-polyoxypropylene block copolymer (Pluronic F68), such as TWEEN®-20 and / or TWEEN®-80. These excipients are commercially available from a number of vendors, such as Sigma, St. Louis, Mo.

[0425] The amount of the various excipients present in any of the disclosed compositions varies. For example, a protein excipient, such as BSA, if present, will can be present at a concentration of between 1.0 weight (wt.) % to about 20 wt. %, preferably 10 wt. %. If an amino acid such as glycine is used in the formulations, it can be present at a concentration of about 1 wt. % to about 5 wt. %. A carbohydrate, such as sorbitol, if present, can be present at a concentration of about 0.1 wt % to about 10 wt. %, such as between about 0.5 wt. % to about 15 wt. %, or about 1 wt. % to about 5 wt. %. If polyethylene glycol is present, it can generally be present on the order of about 2 wt. % to about 40 wt. %, such as about 10 wt. % top about 25 wt. %. If propylene glycol is used in the subject formulations, it will typically be present at a concentration of about 2 wt. % to about 60 wt. %, such as about 5 wt. % to about 30 wt. %. If a detergent such as a sorbitan ester (TWEEN®) is present, it can be present at a concentration of about 0.05 wt. % to about 5 wt. %, such as between about 0.1 wt. % and about 1 wt %, see U.S. Published Patent Application No. 2012 / 0219528. In one example, an aqueous virion-stabilizing formulation comprises a carbohydrate, such as sorbitol, at a concentration of between 0.1 wt. % to about 10 wt. %, such as between about 1 wt. % to about 5 wt. %, and a detergent, such as a sorbitan ester (TWEEN®) at a concentration of between about 0.05 wt. % and about 5 wt. %, such as between about 0.1 wt. % and about 1 wt. %. Virions are generally present in the composition in an amount sufficient to provide a therapeutic effect when given in one or more doses, as defined above.

[0426] Appropriate doses depend on the subject being treated (e.g., human or nonhuman primate or other mammal), age and general condition of the subject to be treated, the severity of the condition being treated, the mode of administration of the AAV vector / virion, among other factors. An appropriate effective amount can be readily determined by a clinician. Thus, an effective amount will fall in a relatively broad range that can be determined through clinical trials. The method can include measuring an outcome, such as kidney function. The method can include administering other agents, such as an AAV vector transductionenhancing agents including histone deacetylase inhibitors, proteasome inhibitors, immunosuppressive agents including corticosteroids, calcineurin inhibitors and sirolimus, and agents that degrade immunoglobulins including IdeS and IdeZ. The method can also combine with ultrasound-targeted microbubble destruction that can enhance AAV vector transduction. The method can also include having the subject make lifestyle modifications.

[0427] For example, for in vivo injection, i.e., injection directly to the subject, an effective dose can be on the order of from about 105to 1016of the AAV virions, such as 108to 1015AAV virions. The dose, of course, depends on the efficiency of transduction, promoter strength, the stability of the message and the protein encoded thereby, and clinical factors. Effective dosages can be readily established using dose response curves.

[0428] Dosage treatment may be a single dose schedule or a multiple dose schedule to ultimately deliver the amount specified above. Moreover, the subject may be administered as many doses as appropriate. Thus, the subject may be given, e.g., 105to 1016AAV virions in a single dose, or two, four, five, six or more doses that collectively result in delivery of, e.g., 105to 1016AAV virions.

[0429] In some aspects, the AAV is administered at a dose of about 1 x 1011to about 1 x 1014viral genomes (vg) per kilogram (kg). In some examples, the AAV is administered at a dose of about 1 x 1012to about 8 x 1013vg / kg. In other examples, the AAV is administered at a dose of about 1 x 1013to about 6 x 1013vg / kg. In specific non-limiting examples, the AAV is administered at a dose of at least about 1 x 1011, at least about 5 x 1011, at least about 1 x 1012, at least about 5 x 1012, at least about 1 x 1013, at least about 5 x 1013, or at least about 1 x 1014vg / kg. In other non-limiting examples, the AAV is administered at a dose of no more than about 5 x 1011, no more than about 1 x 1012, no more than about 5 x 1012, no more than about 1 x 1013, no more than about 5 x 1013, or no more than about 1 x 1014vg / kg. In one non-limiting example, the AAV is administered at a dose of about 1 x 1012vg / kg. In a specific example, the AAV is administered at a dose of about 1 x 1014viral genomes (vg) per kilogram (kg). The AAV can be administered in a single dose, or in multiple doses (such as 2, 3, 4, 5, 6, 7, 8, 9 or 10 doses) as needed for the desired therapeutic results.

[0430] In some examples, the system(s), composition(s), and / or vector(s) disclosed herein are delivered systemically. Systemic delivery refers to delivery of a substance, such as the split intein systems disclosed herein, so that the substance may reach various organs, tissues, and cells throughout the body of the subject, for example delivery to the circulatory system of the subject. Systemic delivery can be achieved through enteral administration (absorption of the drug through the gastrointestinal tract) or parenteral administration (generally injection, infusion, or implantation). Exemplary administration routes include oral administration, intravenous (IV) injection or infusion, intramuscular (IM) injection, or subcutaneous (SC) injection. The split intein system disclosed herein can be formulated for any systemic administration method and delivered systemically. In some aspects, the split intein system is administered parenterally, such as intravenously. In some aspects, the split intein systems delivered systemically may achieve systemic transfection or transduction throughout the body of a subject. In some other aspects, although the split intein systems are delivered systemically, they may exhibit a tendency to transfect or transduce only certain types of organs, tissues, or cells, such as the kidney, thereby achieving targeted delivery. AAV tropism that can assist targeted systemic delivery is discussed above.

[0431] In some examples, the system(s), composition(s), and / or vector(s) disclosed herein are delivered locally. Local delivery refers to delivery of a substance, such as the split intein systems disclosed herein, to a specific site or localized area within the body of a subject. Local delivery may be used to avoid systemic absorption and thereby minimizing systemic side effects. In some examples, the split intein system may be administered by direct parenchymal injection, renal vein injection, renal artery injection, pelvic injection, and / or retrograde transureteral pelvic injection. Pharmaceutical forms suitable for administration to nasal mucosa, eyes, or ear canal include drops, ointments, and sprays. In some other aspects, local delivery is achieved by injection to a vein or artery in a targeted tissue or organ, such as renal vein injection and renal artery injection. In such aspects, the substance (such as the split intein system) to be delivered can first reach the targeted organ and may enter the targeted cells before being circulated to other parts of the body.

[0432] In some other aspects, the split intein system can be locally delivered to the kidney of the subject, as described in Provisional Patent Application No. 63 / 527,248, which is incorporated herein by reference. Pelvic injection can include direct pelvic injection surgically or percutaneously and retrograde transureteral pelvic injection.

[0433] Local delivery may also be assisted with AAV tropism to enhance delivery to one or more targeted organs, tissues, or cell types. In aspects, the split intein systems disclosed herein can target (e.g., transfect or transduce) one or more specific organs, tissues, or cell types, such as the kidney, the eye, and / or the ear, with enhanced efficiency, by systemic or local delivery coupled with utilizing or engineering AAV tropism.

[0434] The systems, compositions, and methods disclosed herein can treat a subject with Alport syndrome. In some examples, the Alport syndrome is X-linked Alport syndrome. In some examples, the Alport syndrome is autosomal Alport syndrome, such as autosomal recessive Alport syndrome, or autosomal dominant Alport syndrome. In some examples the subject with Alport syndrome is male. In some examples the subject with Alport syndrome is female.

[0435] In some examples, treatment with systems, compositions, and methods disclosed herein improve kidney function by at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 1000% in comparison to a control. In some examples, treatment with systems, compositions, and methods disclosed herein delays onset of end stage renal disease by at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 1000% in comparison to a control. In some examples, treatment with systems, compositions, and methods disclosed herein delays time to dialysis by at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 1000% in comparison to a control. In some examples, treatment with systems, compositions, and methods disclosed herein delays time to renal transplant by at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 1000% in comparison to a control. In some examples, treatment with systems, compositions, and methods disclosed herein improves life expectancy by at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 1000% in comparison to a control.

[0436] The method can include the administration of other agents to the subject. In aspects, the subject can also be administered an effective amount of an angiotensin II receptor blocker and / or angiotensin-converting enzyme inhibitor, such as candesartan, eprosartan, irbesartan, losartan, olmesartan, telmisartan, valsartan, benazepril, captopril, cilazapril, enalapril, fosinopril, lisinopril, moexipril, perindopril, ramipril, quinapril, and / or randolapril.

[0437] VII. Systems and Kits

[0438] Also provided are systems and kits that can be used with the disclosed methods. In some examples, the system or kit includes the first nucleic acid molecule encoding a fusion protein including an N-extein of an extein pair and an N-intein of an intein pair, for example with a pharmaceutically acceptable carrier. In some examples, the system or kit includes the second nucleic acid molecule encoding a fusion protein including the C-intein of the intein pair and the C-extein of the extein pair, for example with a pharmaceutically acceptable carrier. In some examples, the system or kit includes a first AAV vector including the first nucleic acid molecule, and / or a second AAV vector including the second nucleic acid molecule, for example with a pharmaceutically acceptable carrier. The first nucleic acid molecule, second nucleic acid molecule, first AAV vector, and / or second AAV vector can be provided at any suitable ratio, such as a molar ratio of the first nucleic acid molecule and the second nucleic acid molecule of about 1 : 1 to about 0.010: 1.

[0439] The split intein system can provide the first nucleic acid molecule and the second nucleic acid at any ratio, such as 10: 1, 5: 1, 2.5: 1, 2: 1, 1: 1, 0.5: 1, 0.25: 1, 0.17:1, 0.1:1, 0.050: 1, 0.025: 1, 0.017: 1, 0.013: 1, 0.010: 1, 0.005: 1, 0.001 :1, or a range between any two of the preceding values. In some examples, the system includes a first adeno-associated viral (AAV) vector including the first nucleic acid molecule and / or a second AAV vector including the second nucleic acid molecule. The first AAV vector and the second AAV vector can be provided at any ratio, such as 10: 1, 5: 1, 2.5: 1, 2:1, 1 :1, 0.5:1, 0.25: 1, 0.17: 1, 0.1 :1, 0.050: 1, 0.025: 1, 0.017:1, 0.013: 1, 0.010: 1, 0.005: 1, 0.001 : 1, or a range between any two of the preceding values.

[0440] In some examples, the first nucleic acid molecule, second nucleic acid molecule, first AAV vector, and / or second AAV vector are provided in a single container. In some examples, one or more of the first nucleic acid molecule, second nucleic acid molecule, first AAV vector, and / or second AAV vector are provided in separate containers, suitable to be combined prior to administration to a subject. In some examples, the first nucleic acid molecule, second nucleic acid molecule, first AAV vector, and / or second AAV vector are provided in separate containers suitable to be administered concurrently, or subsequently to a subject. The container can be any suitable container, such as a syringe, IV bag, ampule, dropper, a vessel, or a vial.

[0441] In some examples, the kit includes an angiotensin II receptor blocker and / or angiotensin-converting enzyme inhibitor, such as candesartan, eprosartan, irbesartan, losartan, olmesartan, telmisartan, valsartan, benazepril, captopril, cilazapril, enalapril, fosinopril, lisinopril, moexipril, perindopril, ramipril, quinapril, and / or randolapril. These can be provided in a container with the first nucleic acid molecule, second nucleic acid molecule, first AAV vector, and / or second AAV vector, or they can be provided in a separate container suitable to be combined with the other container(s), administered concurrently, or administered subsequently to a subject.

[0442] In some examples, the kit includes a Dependoparvo virus viral particle, a Parvovirinae viral particle, a Parvoviridae viral particle, or an AAV particle. In some examples, the kit includes a Dependoparvovirus viral vector, Parvovirinae viral vector, a Parvoviridae viral vector, or an AAV particle.

[0443] VIII. Sequence Identity

[0444] For sequence comparison of amino acid or nucleic acid sequences, typically one sequence acts as a reference sequence, to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are entered into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. Default program parameters are used. Optimal alignment of sequences for comparison can be conducted, for example, by the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2:482, 1981, by the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443, 1970, by the search for similarity method of Pearson & Lipman, Proc. Naf 1. Acad. Sci. USA 85:2444, 1988, by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, WI), or by manual alignment and visual inspection (see for example, Current Protocols in Molecular Biology (Ausubel et al., eds 1995 supplement)).

[0445] One example of a useful algorithm is PILEUP. PILEUP uses a simplification of the progressive alignment method of Feng & Doolittle, J. Mol. Evol. 35:351-360, 1987. The method used is similar to the method described by Higgins & Sharp, CABIOS 5: 151-153, 1989. Using PILEUP, a reference sequence is compared to other test sequences to determine the percent sequence identity relationship using the following parameters: default gap weight (3.00), default gap length weight (0.10), and weighted end gaps. PILEUP can be obtained from the GCG sequence analysis software package, such as version 7.0 (Devereaux et al., Nuc. Acids Res. 12:387-395, 1984).

[0446] Another example of algorithms that are suitable for determining percent sequence identity and sequence similarity are the BLAST and the BLAST 2.0 algorithm, which are described in Altschul et al., J. Mol. Biol. 215:403-410, 1990 and Altschul et al., Nucleic Acids Res. 25:3389-3402, 1977. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (ncbi.nlm.nih.gov). The BLASTN program (for nucleotide sequences) uses as defaults a word length (W) of 11, alignments (B) of 50, expectation (E) of 10, M=5, N=-4, and a comparison of both strands. The BLASTP program (for amino acid sequences) uses as defaults a word length (W) of 3, and expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff & Henikoff, Proc. Natl. Acad. Sci. USA 89: 10915, 1989).

[0447] IX. Additional Embodiments

[0448] Clause 1 . A split intein system comprising:

[0449] (i) a first nucleic acid molecule encoding a fusion protein comprising in N to C terminal order, an N- extein of an extein pair and an N-intein of an intein pair, wherein the N-extein comprises a first signal sequence N-terminal to an N-terminal portion of a type IV collagen alpha chain; and

[0450] (ii) a second nucleic acid molecule encoding a fusion protein comprising in N to C terminal order, a second signal sequence, a C-intein of the intein pair, and a C-extein of the extein pair, wherein the C-extein comprises a C-terminal portion of the type IV collagen alpha chain; and wherein the type IV collagen alpha chain is COL4A5; the N- and C-terminal portions of the type IV collagen alpha chain together comprise the type IV collagen alpha chain sequence separated at a split point; and the N- and C-exteins are spliced together to form the mature type IV collagen alpha chain when the first and second nucleic acid molecules are expressed in mammalian cells.

[0451] Clause 2. The split intein system of clause 1, comprising:

[0452] (i) a first adeno-associated viral (AAV) vector comprising the first nucleic acid molecule; and

[0453] (ii) a second AAV vector comprising the second nucleic acid molecule.

[0454] Clause 3. The split intein system of clause 2, wherein the first and / or the second AAV vector is AAV9. Clause 4. The split intein system of clause 2, wherein the first and / or the second AAV vector is AAV- KP1.

[0455] Clause 5. The split intein system of clause 2, wherein the first and / or the second AAV vector is AAV1, AAVl_9mtl00, AAVl_9mt30, AAVl_9mt76, AAV2, AAV2G9, AAV2i8, AAV2retro, AAV2R585E, AAV2R585E9_2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV9AA272, AAV9AA22, AAV9W22A, AAV10, AAV11, AAV2.7m8, AAVAnc80, AAVbb.2, AAV-DJ, AAVHN1, AAVHN2, AAVHN3, AAVhu.l l, AAVhu.13, AAVhu.37, AAV-KP1, AAV-KP2, AAV-KP3, AAVLK03, AAVNP40, AAVNP59, AAVPHP.B, AAVPHP.eB, AAVPHP.S, AAVpol, AAVrh.8, AAVrh.10, AAVrh.20, AAVrh.43, or AAVShHIO.

[0456] Clause 6. The split intein system of any one of clauses 1-5, wherein the intein pair is a Npu DnaE intein pair.

[0457] Clause 7. The split intein system of clause 6, wherein the N-intein comprises the amino acid sequence of SEQ ID NO: 26 or an amino acid sequence at least 95% identical thereto; and the C-intein comprises the amino acid sequence of SEQ ID NO: 27 or an amino acid sequence at least 95% identical thereto.

[0458] Clause 8. The split intein system of any one of clauses 1-7, wherein, the first signal sequence, the N- terminal portion of the type IV collagen alpha chain, and the N-intein are fused directly.

[0459] Clause 9. The split intein system of any one of clauses 1-8, wherein, the second signal sequence, the C-intein, and the C-terminal portion of the type IV collagen alpha chain are fused directly.

[0460] Clause 10. The split intein system of any one of clauses 1-9, wherein the split intein system comprises the first nucleic acid molecule and the second nucleic acid molecule at a ratio of about 1 : 1 to about 0.010: 1.

[0461] Clause 11. The split intein system of clause 10, wherein the split intein system comprises the first nucleic acid molecule and the second nucleic acid molecule at a ratio of about 0.017: 1 to about 0.010: 1.

[0462] Clause 12. The split intein system of any one of clauses 1-11, wherein the split point is within a collagenous domain of the type IV collagen alpha chain.

[0463] Clause 13. The split intein system of any one of clauses 1-12, wherein the first signal sequence and / or the second signal sequence is a collagen signal sequence or a chymotrypsin signal sequence. Clause 14. The split intein system of clause 13, wherein the first signal sequence and the second signal sequence are different.

[0464] Clause 15. The split intein system of clause 13 or clause 14, wherein the collagen signal sequence comprises SEQ ID NO: 4 and / or wherein the chymotrypsin signal sequence comprises SEQ ID NO: 5.

[0465] Clause 16. The split intein system of any one of clauses 1-15, wherein the type IV collagen alpha chain comprises the amino acid sequence of any one of SEQ ID NOs: 15-25 or an amino acid sequence at least 95% identical thereto.

[0466] Clause 17. The split intein system of clause 16, wherein the type IV collagen alpha chain comprises the amino acid sequence of SEQ ID NO: 15 or an amino acid sequence at least 95% identical thereto.

[0467] Clause 18. The split intein system of any one of clauses 16-17, wherein the split point is at position 697-position 1488 relative to SEQ ID NO: 15.

[0468] Clause 19a. The split intein system of any one of clauses 16-18, wherein:

[0469] (a) the N-terminal portion comprises amino acids 1 - 450 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 451 to 1685 of the type IV collagen alpha chain (SP1);

[0470] (b) the N-terminal portion comprises amino acids 1 - 696 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 697 to 1685 of the type IV collagen alpha chain (SP2);

[0471] (c) the N-terminal portion comprises amino acids 1 - 863 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 864 to 1685 of the type IV collagen alpha chain (SP4);

[0472] (d) the N-terminal portion comprises amino acids 1 - 878 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 879 to 1685 of the type IV collagen alpha chain (SP5);

[0473] (e) the N-terminal portion comprises amino acids 1 - 941 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 942 to 1685 of the type IV collagen alpha chain (SP7);

[0474] (f) the N-terminal portion comprises amino acids 1 - 977 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 978 to 1685 of the type IV collagen alpha chain (SP9); (g) the N- terminal portion comprises amino acids 1 - 1070 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 1071 to 1685 of the type IV collagen alpha chain (SP11); or

[0475] (h) the N-terminal portion comprises amino acids 1 - 1135 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 1136 to 1685 of the type IV collagen alpha chain (SP12); and wherein amino acid numbering is relative to SEQ ID NO: 15.

[0476] Clause 19b. The split intein system of any one of clauses 16-18, wherein:

[0477] (a) the N-terminal portion comprises amino acids 27 - 696 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 697 - 1685 of the type IV collagen alpha chain (SP2);

[0478] (b) the N-terminal portion comprises amino acids 27 - 878 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 879 - 1685 of the type IV collagen alpha chain (SP5);

[0479] (c) the N-terminal portion comprises amino acids 27 - 888 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 889 - 1685 of the type IV collagen alpha chain (SP20);

[0480] (d) the N-terminal portion comprises amino acids 27 - 915 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 916 - 1685 of the type IV collagen alpha chain (SP6);

[0481] (e) the N-terminal portion comprises amino acids 27- 941 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 942 - 1685 of the type IV collagen alpha chain (SP7);

[0482] (f) the N-terminal portion comprises amino acids 27- 961 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 962 - 1685 of the type IV collagen alpha chain (SP8);

[0483] (g) the N-terminal portion comprises amino acids 27 - 977 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 978 - 1685 of the type IV collagen alpha chain (SP9);

[0484] (h) the N-terminal portion comprises amino acids 27 - 995 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 996 - 1685 of the type IV collagen alpha chain (SP10);

[0485] (i) the N-terminal portion comprises amino acids 27 - 1070 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 1071 - 1685 of the type IV collagen alpha chain (SP11);

[0486] (j) the N-terminal portion comprises amino acids 27 - 1098 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 1099 - 1685 of the type IV collagen alpha chain (SP27); or

[0487] (k) the N-terminal portion comprises amino acids 27 - 1135 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 1136 - 1685 of the type IV collagen alpha chain (SP12); and wherein amino acid numbering is relative to SEQ ID NO: 15. Clause 20a. The split intein system of any one of clauses 16-19, wherein:

[0488] (a) the N terminal portion comprises DEI at its C-terminus and the C-terminal portion comprises CEPG (SEQ ID NO: 230) at its N-terminus (SP1);

[0489] (b) the N terminal portion comprises IPG at its C-terminus and the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP2);

[0490] (c) the N terminal portion comprises ERG at its C-terminus and the C-terminal portion comprises SPGI (SEQ ID NO: 231) at its N-terminus (SP4);

[0491] (d) the N terminal portion comprises PPG at its C-terminus and the C-terminal portion comprises SPGL (SEQ ID NO: 222) at its N-terminus (SP5);

[0492] (e) the N terminal portion comprises EKG at its C-terminus and the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP7);

[0493] (f) the N terminal portion comprises PGV at its C-terminus and the C-terminal portion comprises SGPK (SEQ ID NO: 225)at its N-terminus (SP9);

[0494] (g) the N terminal portion comprises PGI at its C-terminus and the C-terminal portion comprises SSIG (SEQ ID NO: 227) at its N-terminus (SP11); or

[0495] (h) the N terminal portion comprises KGI at its C-terminus and the C-terminal portion comprises SGPP (SEQ ID NO: 229) at its N-terminus (SP12).

[0496] Clause 20b. The split intein system of any one of clauses 16-19, wherein:

[0497] (a) the N-terminal portion comprises IPG at its C-terminus and the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP2);

[0498] (b) the N-terminal portion comprises PPG at its C-terminus and the C-terminal portion comprises SPGL (SEQ ID NO: 222) at its N-terminus (SP5);

[0499] (c) the N-terminal portion comprises AGA at its C-terminus and the C-terminal portion comprises SGFP (SEQ ID NO: 223) at its N-terminus (SP20);

[0500] (d) the N-terminal portion comprises PGR at its C-terminus and the C-terminal portion comprises SGVP (SEQ ID NO: 224) at its N-terminus (SP6);

[0501] (e) the N-terminal portion comprises EKG at its C-terminus and the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP7);

[0502] (f) the N-terminal portion comprises LLG at its C-terminus and the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP8);

[0503] (g) the N-terminal portion comprises PGV at its C-terminus and the C-terminal portion comprises SGPK (SEQ ID NO: 225) at its N-terminus (SP9);

[0504] (h) the N-terminal portion comprises PGL at its C-terminus and the C-terminal portion comprises SGQP (SEQ ID NO: 226) at its N-terminus (SP10); (i) the N-terminal portion comprises PGI at its C-terminus and the C-terminal portion comprises SSIG (SEQ ID NO: 227) at its N-terminus (SP11);

[0505] (j) the N-terminal portion comprises IKG at its C-terminus and the C-terminal portion comprises SVGD (SEQ ID NO: 228) at its N-terminus (SP27); or

[0506] (k) the N-terminal portion comprises KGI at its C-terminus and the C-terminal portion comprises SGPP (SEQ ID NO: 229) at its N-terminus (SP12).

[0507] Clause 21a. The split intein system of any one of claims 1-20, wherein the N-terminal portion does not include a collagen signal sequence.

[0508] Clause 21b. The split intein system of any one of clauses 1-20, wherein the first nucleic acid molecule and / or the second nucleic acid molecule is operably linked to a CAG promoter.

[0509] Clause 22. The split intein system of any one of clauses 1-21, wherein the first nucleic acid molecule and / or the second nucleic acid molecule further comprises a woodchuck hepatitis virus post-transcriptional regulatory element; and / or wherein the first nucleic acid molecule and / or the second nucleic acid molecule is operably linked to the woodchuck hepatitis virus post-transcriptional regulatory element.

[0510] Clause 23. The split intein system of any one of clauses 1-22, wherein the first nucleic acid molecule and / or the second nucleic acid molecule further encodes a SV40 polyadenylation signal.

[0511] Clause 24. A pharmaceutical composition comprising an effective amount of the split intein system of any one of clauses 1-23 and a pharmaceutically acceptable carrier, wherein the split intein system comprises the first nucleic acid molecule and the second nucleic acid molecule at a ratio of about 1:1 to about 0.010:1.

[0512] Clause 25. A split intein system comprising:

[0513] (i) a first adeno-associated viral (AAV) vector comprising a first nucleic acid molecule operably linked to a CAG promoter, wherein the first nucleic acid molecule encodes a fusion protein comprising in N to C terminal order, an N-extein of an extein pair and an N-intein of an Npu DnaE intein pair, wherein the N- extein comprises a collagen signal sequence N-terminal to an N-terminal portion of a type IV collagen alpha 5 (COL4A5) chain, wherein the collagen signal sequence, the N-terminal portion of the COL4A5 chain, and the N-intein are fused directly;

[0514] (ii) a second AAV vector comprising a second nucleic acid molecule operably linked to a CAG promoter, wherein the second nucleic acid molecule encodes a fusion protein comprising, in N to C terminal order, a signal sequence, the C-intein of the Npu DnaE intein pair, and the C-extein of the extein pair, wherein the C-extein comprises a C-terminal portion of the COL4A5 chain, wherein the signal sequence, the C-intein, and the C-terminal portion of the COL4A5 chain are fused directly; and wherein the N- and C-terminal portions of the COL4A5 chain together comprise the COL4A5 chain sequence separated at a split point; and the N- and C-exteins are spliced together to form the mature COL4A5 chain when the first and second nucleic acid molecules are expressed in mammalian cells.

[0515] Clause 26a. The split intein system of clause 25, wherein:

[0516] (a) the N-terminal portion comprises amino acids 1 - 450 of the COL4A5 chain, and the C-terminal portion comprises amino acids 451 to 1685 of the COL4A5 chain fSPl);

[0517] (b) the N-terminal portion comprises amino acids 1 - 696 of the COL4A5 chain, and the C-terminal portion comprises amino acids 697 to 1685 of the COL4A5 chain (SP2);

[0518] (c) the N-terminal portion comprises amino acids 1 - 863 of the COL4A5 chain, and the C-terminal portion comprises amino acids 864 to 1685 of the COL4A5 chain (SP4);

[0519] (d) the N-terminal portion comprises amino acids 1 - 878 of the COL4A5 chain, and the C-terminal portion comprises amino acids 879 to 1685 of the COL4A5 chain (SP5);

[0520] (e) the N-terminal portion comprises amino acids 1 - 941 of the COL4A5 chain, and the C-terminal portion comprises amino acids 942 to 1685 of the COL4A5 chain (SP7);

[0521] (f) the N-terminal portion comprises amino acids 1 - 977 of the COL4A5 chain, and the C-terminal portion comprises amino acids 978 to 1685 of the COL4A5 chain (SP9);

[0522] (g) the N-terminal portion comprises amino acids 1 - 1070 of the COL4A5 chain, and the C- terminal portion comprises amino acids 1071 to 1685 of the COL4A5 chain (SP11); or

[0523] (h) the N-terminal portion comprises amino acids 1 - 1135 of the COL4A5 chain, and the C- terminal portion comprises amino acids 1136 to 1685 of the COL4A5 chain (SP12); and wherein amino acid numbering is relative to SEQ ID NO: 15.

[0524] Clause 26b. The split intein system of clause 25, wherein:

[0525] (a) the N-terminal portion comprises amino acids 27 - 696 of the COL4A5 chain, and the C- terminal portion comprises amino acids 697 - 1685 of the COL4A5 chain (SP2);

[0526] (b) the N-terminal portion comprises amino acids 27 - 878 of the COL4A5 chain, and the C- terminal portion comprises amino acids 879 - 1685 of the COL4A5 chain (SP5);

[0527] (c) the N-terminal portion comprises amino acids 27 - 888 of the COL4A5 chain, and the C- terminal portion comprises amino acids 889 - 1685 of the COL4A5 chain (SP20);

[0528] (d) the N-terminal portion comprises amino acids 27 - 915 of the COL4A5 chain, and the C- terminal portion comprises amino acids 916 - 1685 of the COL4A5 chain (SP6);

[0529] (e) the N-terminal portion comprises amino acids 27 - 941 of the COL4A5 chain, and the C- terminal portion comprises amino acids 942 - 1685 of the COL4A5 chain (SP7); (f) the N-terminal portion comprises amino acids 27 - 961 of the COL4A5 chain, and the C-terminal portion comprises amino acids 962 - 1685 of the COL4A5 chain (SP8);

[0530] (g) the N-terminal portion comprises amino acids 27 - 977 of the COL4A5 chain, and the C- terminal portion comprises amino acids 978 - 1685 of the COL4A5 chain (SP9);

[0531] (h) the N-terminal portion comprises amino acids 27 - 995 of the COL4A5 chain, and the C- terminal portion comprises amino acids 996 - 1685 of the COL4A5 chain (SP10);

[0532] (i) the N-terminal portion comprises amino acids 27 - 1070 of the COL4A5 chain, and the C- terminal portion comprises amino acids 1071 - 1685 of the COL4A5 chain (SP11);

[0533] (j) the N-terminal portion comprises amino acids 27 - 1098 of the COL4A5 chain, and the C- terminal portion comprises amino acids 1099 - 1685 of the COL4A5 chain (SP27); or

[0534] (k) the N-terminal portion comprises amino acids 27 - 1135 of the COL4A5 chain, and the C- terminal portion comprises amino acids 1136 - 1685 of the COL4A5 chain (SP12); and wherein amino acid numbering is relative to SEQ ID NO: 15.

[0535] Clause 27a. The split intein system of clause 25 or 26, wherein

[0536] (a) the N terminal portion comprises DEI at its C-terminus and the C-terminal portion comprises CEPG (SEQ ID NO: 230) at its N-terminus (SP1);

[0537] (b) the N terminal portion comprises IPG at its C-terminus and the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP2);

[0538] (c) the N terminal portion comprises ERG at its C-terminus and the C-terminal portion comprises SPGI (SEQ ID NO: 231) at its N-terminus (SP4);

[0539] (d) the N terminal portion comprises PPG at its C-terminus and the C-terminal portion comprises SPGL (SEQ ID NO: 222) at its N-terminus (SP5);

[0540] (e) the N terminal portion comprises EKG at its C-terminus and the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP7);

[0541] (f) the N terminal portion comprises PGV at its C-terminus and the C-terminal portion comprises SGPK (SEQ ID NO: 225) at its N-terminus (SP9);

[0542] (g) the N terminal portion comprises PGI at its C-terminus and the C-terminal portion comprises SSIG (SEQ ID NO: 227) at its N-terminus (SP11); or

[0543] (h) the N terminal portion comprises KGI at its C-terminus and the C-terminal portion comprises SGPP (SEQ ID NO: 229) at its N-terminus (SP12).

[0544] Clause 27b. The split intein system of clause 25 or 26, wherein

[0545] (a) the N-terminal portion comprises IPG at its C-terminus and the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP2);

[0546] (b) the N-terminal portion comprises PPG at its C-terminus and the C-terminal portion comprises SPGL (SEQ ID NO: 222) at its N-terminus (SP5); (c) the N-terminal portion comprises AGA at its C-terminus and the C-terminal portion comprises SGFP (SEQ ID NO: 223) at its N-terminus (SP20);

[0547] (d) the N-terminal portion comprises PGR at its C-terminus and the C-terminal portion comprises SGVP (SEQ ID NO: 224) at its N-terminus (SP6);

[0548] (e) the N-terminal portion comprises EKG at its C-terminus and the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP7);

[0549] (f) the N-terminal portion comprises LLG at its C-terminus and the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP8);

[0550] (g) the N-terminal portion comprises PGV at its C-terminus and the C-terminal portion comprises SGPK (SEQ ID NO: 225) at its N-terminus (SP9);

[0551] (h) the N-terminal portion comprises PGL at its C-terminus and the C-terminal portion comprises SGQP (SEQ ID NO: 226) at its N-terminus (SP10);

[0552] (i) the N-terminal portion comprises PGI at its C-terminus and the C-terminal portion comprises SSIG (SEQ ID NO: 227) at its N-terminus (SPl 1);

[0553] (j) the N-terminal portion comprises IKG at its C-terminus and the C-terminal portion comprises SVGD (SEQ ID NO: 228) at its N-terminus (SP27); or

[0554] (k) the N-terminal portion comprises KGI at its C-terminus and the C-terminal portion comprises SGPP (SEQ ID NO: 229) at its N-terminus (SP12).

[0555] Clause 28. A pharmaceutical composition comprising an effective amount of the split intein system of any one of clauses 25-27, and a pharmaceutically acceptable carrier.

[0556] Clause 29. A method of treating Alport syndrome in a subject, comprising administering to the subject an effective amount of the split intein system of any one of clauses 1-23 or 25-27, or the pharmaceutical composition of clause 24 or clause 28, thereby treating the Alport syndrome in the subject.

[0557] Clause 30. The method of clause 29, comprising delivering the split intein system to the kidney of the subject.

[0558] Clause 31. The method of clause 30, wherein delivering the split intein system to the kidney of the subject comprises systemic administration.

[0559] Clause 32. The method of clause 31, wherein the systemic administration comprises systemic injection or systemic infusion.

[0560] Clause 33. The method of clause 30 or clause 32, wherein delivering the split intein system to the kidney of the subject comprises locally delivering the split intein system to the kidney of the subject. Clause 34. The method of clause 33, wherein the locally delivering the split intein system to the kidney of the subject comprises direct parenchymal injection, renal vein injection, and / or renal artery injection.

[0561] Clause 35. The method of clause 33, wherein the locally delivering the split intein system to the kidney of the subject comprises direct pelvic injection or retrograde transureteral pelvic injection.

[0562] Clause 36. The method of any one of clauses 29-35, wherein the split intein system transduces at least one of mesangial cells, glomerular endothelial cells, parietal epithelial cells, podocytes, proximal tubule cells, Loop of Henle cells, distal tubule cells, collecting duct cells, fibroblasts, pericytes, or vascular smooth muscle cells.

[0563] Clause 37. The method of any one of clauses 29-36, wherein the method: a) improves kidney function; b) delays onset of end stage renal disease; c) delays time to dialysis; d) delays time to renal transplant; and / or e) improves life expectancy; of the subject.

[0564] Clause 38. The method of any one of clauses 29-37, wherein the Alport syndrome is X-linkcd Alport syndrome.

[0565] Clause 39. The method of any one of clauses 29-38, wherein the method further comprises administering an effective amount of angiotensin II receptor blocker (ARB) and / or angiotensin-converting enzyme inhibitor; wherein the ARB comprises candesartan, captopril, eprosartan, irbesartan, losartan, moexipril, olmesartan, telmisartan, and / or valsartan; and wherein the angiotensin-converting enzyme inhibitor comprises benazepril, cilazapril, enalapril, fosinopril, lisinopril, perindopril, ramipril, quinapril, and / or randolapril.

[0566] EXAMPLES

[0567] The following examples are provided to illustrate particular features of certain aspects of the disclosure, but the scope of the claims should not be limited to those features exemplified. Example 1: Materials and Methods

[0568] Vector support machine model: A vector support machine (VSM) similarity prediction model was established, as reported previously by Apgar et al. (Apgar et al., PLoS One 7, e37355 (2012)). With assistance of InBase, using the criteria used by Apgar et al., 188 unique exteins were identified and protein sequences were taken from GenBank (Perler et al., Nucleic Acids Res 30, 383-384 (2002)). Seven amino acid sequences were extracted termed “splice site cassette” including XXX[C / S / T]XXX from the splice sites found in the 188 native extein splice sites, in which X represents any amino acid residue while C / S / T represent the conserved amino acids found at +1 position of C-extein. They served as true positive extein splice sites. For true negative splice site cassette controls, three XXX[C / S / T]XXX sequences were randomly extracted from each of the 188 extein sequences that exclude true splice site cassettes and any replicated selections, generating a total 188x3=564 unique true negatives. Using a total of 188+564=752 labeled data as a training set, an S VM model was trained as described using Python’ s scikit-learn S VM with a linear kernel and a cost factor of 3 (Perler et al., Nucleic Acids Res 30, 383-384 (2002)). The model was trained 25 times using distinct random seeds per leave-one-out cross validation for each true positive and negative splice site cassette. The average of the 25 SVM scores obtained by the “predict_proba” method was then determined for each cassette. This SVM model exhibited a predictive power with the area under the receiver operating characteristic (ROC) curve of 0.881 (FIG. 2). The most optimal true or false binary clarification threshold that maximizes Youden's J was determined to be 0.305. With this threshold, the confusion matrix revealed that the sensitivity, specificity and overall accuracy of the similarity prediction by the SVM model were 0.761, 0.904 and 0.868, respectively (FIG. 3).

[0569] Design of the split intein COL4A5 constructs: The CAG-WPRE3-SV40pA sequence was constructed from the Addgene plasmid, pAAV-CAG-tdTomato (Addgene 59462). Using this plasmid as the starting material, the CAG-WPRE3-SV40pA (SEQ ID NO: 2) and CAG-SV40pA (SEQ ID NO: 3) plasmids were created by incorporating the WPRE3 sequence (Choi et al., Mol Brain 7, 17 (2014)).

[0570] SEQ ID NOs: 1-3 were incorporated between 130-bp AAV inverted terminal repeats (ITR) of AAV vector plasmid backbones (Addgene 59462).

[0571] The 26-amino-acid-long COL4A5 native signal sequence (MKLRGVSLAAGLFLLALSLWGQPAEA (SEQ ID NO: 4)) or the 18-amino-acid-long human chymotrypsin signal sequence (MAFLWLLSCWALLGTTFG (SEQ ID NO: 5)) was added to the N- terminus of the COL4A5 C-half (FIG. 6A). As a control, a COL4A5 C-half construct devoid of any signal peptide was created (FIG. 6A). All the constructs were tagged with either the 22-amino-acid-long FLAG tag (DYKDHDGDYKDHDIDYKDDDDK (SEQ ID NO: 6)) or the human influenza virus hemagglutinin tag (YPYDVPDYA (SEQ ID NO: 7)). The expected lengths of split intein dual AAV-CAG-Col4A5 vector genomes, including the two 145 nucleotide- long ITRs at both ends, for each of the split points (SPs) with or without WPRE or WPRE3 are summarized in FIG. 7.

[0572] In vitro assessment of protein trans-splicing using HEK293 cells'. HEK293 cells were seeded at a density of 5 x 1(F cells / well in 6-well plates and incubated at 37°C and 5% CO2. Twenty-four hours after seeding the cells, cells were transfected with 3 ig of plasmid DNA using lipofectamine 2000 (11668027, Invitrogen) and Opti-MEM (31985088, Gibco) as per the manufacturer's instructions. Forty-eight hours after transfection, cells were washed with ice-cold PBS, and subsequently lysed with lysis buffer (50 mM Tris- HC1, 150 mM NaCl, ImM EDTA, and 1% TritonX-100, pH7.5) supplemented with complete™ Mini protease inhibitor cocktail (11836170001, Roche). Lysis was conducted on ice for 30 min. After lysis, lysate was centrifuged at 16,000 x g and debris was removed. Total cell lysates were collected, combined with sample buffer, and denatured at 60°C for 20 min. Protein concentrations of the total cell lysates were determined using the DC Protein Assay Kit (Bio-Rad, Hercules, CA). For western blot analysis, cell lysates containing 100 pg of cellular proteins were separated by SDS-polyacrylamide gel electrophoresis (PAGE) using a 6% gel, and subsequently transferred onto polyvinylidene difluoride (PVDF) membrane using the Power Blotter-Semi-dry Transfer System (Invitrogen). PVDF membrane was probed with mouse monoclonal anti-FLAG M2 antibody (F1804, Sigma Aldrich), followed by incubation with anti-mouse IgG antibody conjugated to horseradish peroxidase (62-6520, Invitrogen) for signal detection. Signals of the PVDF blots were visualized using the Amersham ImageQuant™ 800 Western Blot Imaging System (Cytiva). Anti-glyceraldehyde-3-phosphate dehydrogenase (GAPDH) antibody was used to probe the quantify of GAPDH as a loading control to ensure consistent protein loading across the samples.

[0573] Human podocyte experiments: Differentiated human podocytes were seeded at a density of 4 x 105cells / well in a 6-well plate and incubated at 37°C under 5% CO2. Twenty-four hours after seeding, the culture medium was replaced with a fresh medium, and podocytes were infected with the AAV-KP1 vectors expressing split intein COL4A5 proteins (AAV-KPl-CAG-SP2-Col4A5-Nint-FLAG-WPRE and AAV- KPl-CAG-nativeSS-Cint-SP2-Col4A5-HA -noWPRE) or tdTomato (AAV-KPl-CAG-tdTomato) at a multiplicity of infection (MOI) of 105. Seventy-two hours post-injection, the culture medium was replaced with fresh medium again. These AAV vectors were produced in HEK293 cells by an adenovirus-free three plasmid transfection method (Matsushita et al., Gene Ther 5, 938-945 (1998)) using the following AAV vector plasmids, pAAV-CAG-SP2-Col4A5-Nint-FLAG-WPRE, pAAV-CAG-nativeSS-Cint-SP2-Col4A5- HA-noWPRE, and pAAV-CAG-tdTomato (Addgene 59462). The produced AAV vectors were then purified by two rounds of cesium chloride (CsCl) density -gradient ultracentrifugation followed by dialysis as previously described (Burton et al., Proc Natl Acad Sci U S A 96, 12725-12730 (1999)). The titers of the purified AAV vectors were determined by a quantitative dot blot assay (Powers et al., J Vis Exp (2018)). Five days after infection, 1 mL of culture media was collected and concentrated to 100 (tL with Amicon ultra centrifugal filter (UFC8010, Millipore) for western blot analysis. For western blot, 20 LIL of the concentrated media was loaded on a 7.5% polyacrylamide gel, separated by SDS-PAGE, and subsequently transferred onto polyvinylidene difluoride (PVDF) membrane using the Power Blotter Semi-dry Transfer System (Invitrogen). PVDF membrane was probed with rabbit monoclonal anti-HA antibody (3724, Cell Signaling), followed by incubation with anti-rabbit IgG antibody conjugated to horseradish peroxidase (65- 6120, Invitrogen). Signals of the PVDF blots were visualized using the Amersham ImageQuant™ 800 Western Blot Imaging System (Cytiva). Testing C-half ratios'. HEK293 cells were transfected as previously described. The following ratios of pAAV-CAG-SP2-Col4A5-Nint-FLAG-WPRE and pAAV-CAG-nativeSS-Cint-SP2-Col4A5-FLAG- WPRE were tested: 1:1, 0.050:1. 0.025:1, 0.017:1, 0.013: 1, 0.010:1, 0.005:1, and 0.001:1. A total of 3 pg of plasmid DNA, containing both plasmid DNAs at the specified ratios, was utilized for transfection done in 6 well plates. Transfected cells were harvested 48 hours post-transfection, and cell lysates containing 100 pg of cellular proteins were separated by SDS-PAGE using 6% gels, and subsequently transferred onto PVDF membranes. The membranes were then probed with mouse monoclonal anti-FLAG M2 antibody or rabbit polyclonal anti-human COL4A5 H53 antibody.

[0574] Assess functional viability of the COL4A5 splice split point candidates and identify split points'. HEK293 cells were seeded in 6-well plates at a density of 1 x 106cells / well and incubated at 37°C under 5% CO2. After 24-hour incubation, cells were transfected with 3 pg of plasmid DNA using lipofectamine 2000 (11668027, Invitrogen) and Opti-MEM® (31985088, Gibco) as per manufacturer's instructions. Forty-eight hours after transfection, cells were washed with ice-cold PBS, and subsequently lysed with lysis buffer (50 mM Tris-HCl, 150 mM NaCl, ImM EDTA, and 1% TritonX-100, pH7.5) supplemented with cOmplete™ Mini protease inhibitor cocktail (11836170001, Roche). Lysis was conducted on ice for 30 min. After lysis, lysate was centrifuged at 16,000 x g and debris was removed. Total cell lysates were collected, combined with sample buffer, and denatured at 60°C for 20 min. For western blot analysis, total cell lysates were separated by SDS-PAGE using 6% gels, and transferred onto a PVDF membrane using the Power Blotter Semi-dry Transfer System (Invitrogen). PVDF membrane was probed with mouse monoclonal anti-FLAG M2 antibody or rabbit polyclonal anti-human COL4A5 H53 antibody.

[0575] Split point effects on the steady-state level expression of split intein COL4A5 N -half and C-half constructs: HEK293 cells seeded in 6 -well plates were transfected with 3 pg of plasmid DNA expressing split intein N-half or C-half constructs of various lengths harboring different split points. Forty-eight hours post-transfection, cells were harvested and equal amounts of crude cell lysates were subjected to SDS-PAGE using a 6% gel for N-half constructs and a 10% gel for C-half constructs. This was followed by western blot analysis using mouse monoclonal anti-FLAG M2 antibody.

[0576] C57BL / 6-based LSL knock-in strain generation and breeding: C57BL / 6J was used as a background strain (JAX Registry 000664). A loxP-stop-loxP cassette, including three copies of the SV40 polyadenylation signals (i.e., SV40pA 3x sequence) flanked with a pair of loxP sites was introduced into intron 1, between mouse chromosome X positions 141,510,668 and 141,510,669 (mm 10), of the mouse Col4a5 gene (FIG. 12A). Three sgRNA (Col4a5_crRNAl: GCAGTGTAGAAGTCTCCTAA (SEQ ID NO: 8), Col4a5_crRNA2: CAGTGTAGAAGTCTCCTAAA (SEQ ID NO: 9) and Col4a5_crRNA3: GCCCTGTAAGCTGACCCTTT (SEQ ID NO: 10)) were designed to test target sequence cleavage in vitro and Col4a5_crRNAl was selected to introduce the loxP-stop-loxP cassette. A double-stranded DNA donor plasmid was constructed, which carried the loxP-stop -LoxP cassette within the left homology arm (LHA, 1.4 kb) and the right homology arm (RHA, 1.8 kb). Cas9 / sgRNA / donor plasmids were microinjected into 1- cell embryos and 300 micro-injected embryos were transferred to pseudo pregnant female recipients. DNA sequence analysis was performed with the founder generation mice to validate molecular genome changes in target genes. Selected founder knock-in (KI) mice were further crossed with C57BL / 6J mice in order to produce the N1 generation and mice from N2 generation. An inbred Fl generation was established from the N2 generation animals and used to establish a mouse colony.

[0577] To maintain the colony, male homozygous wild-type were crossed with female heterozygous KI mouse to obtain pups with 4 different genotypes (male homozygous wild-type, male hemizygous KI, female homozygous wild-type and female heterozygous KI). Genotyping primers were designed, with the forward primer originating from the LHA and the reverse primer derived from the KI region. Detection of a 324-bp PCR product signifies the presence of the knock-in (KI) allele, while its absence indicates the wild-type allele.

[0578] F primer: ACCCCATCAACTTGTCAACCT (SEQ ID NO: 11) R primer: AGAGTTTGTCCTCAACCGCG (SEQ ID NO: 12)

[0579] Phenotypes of LSL-Col4a5 XLAS mice: General appearance and body weight were monitored at least once a week. The development and progression of chronic kidney disease was assessed through the measurement of the albumin-to-creatinine (ACR) ratio in urine, as well as the determination of creatinine and blood urea nitrogen (BUN) levels in whole blood samples. To this end, spot urine samples were collected from mice, and albuminuria was quantified using Albuwell M (1011 , Ethos Biosciences, Logan Township, NJ) following the manufacturer’s protocol. Urine albumin levels were normalized by urine creatinine concentration measured with the Creatinine Reagent Assay (C75391250, Pointe Scientific, Canton, MI), expressed as g / g creatinine. Blood creatinine and BUN levels were measured using the i-STAT CHEM8+ Cartridge (09P31-26, Abbott Laboratories, IL.) in accordance with the manufacturer’s instructions. The resulting data were expressed as mg / dL for creatinine and mg / dL for BUN.

[0580] Intravenous (IV) injection ofAAV9 and kidney immunofluorescence in Col4a5 G5X mice: 25 to 30- wk-old Col4a5 G5X hemizygous male mice (Rheault et al., J Am Soc Nephrol 15, 1466-1474 (2004)), an XLAS mouse model, and age-matched wild-type (WT) male mice were injected with AAV9-CAG- tdTomato or AAV-KPl-CAG-tdTomato vector at IxlO13vg / kg (n=3 to 5 per group). For AAV vector injection, AAV vector stocks were diluted with 5% sorbitol / PBS and adjusted to a total volume of 300 pL. Two weeks post-injection, the kidney tissues were harvested following perfusion with 4% paraformaldehyde (PF A) (15710, Electron Microscopy Sciences, Hatfield, PA). Harvested kidneys were fixed in PFA and subsequently equilibrated in 30% sucrose (S5, Fisher Scientific, Hampton, NH) for cryoprotection. The fixed tissues were cryo-embedded in Tissue-Tek O.C.T. Compound (4583, Sakura Finetek, St. Torrance, CA) and cut into 5-pm-thick sections using a cryostat. Immunostaining was performed using rabbit anti- WT1 (1:100, ab89901, Abeam, Cambridge, UK) and goat anti-rabbit IgG Alexa Fluor 647 (1:500, 111-605- 144, Jackson ImmunoResearch, West Grove, PA). The sections were then counterstained with Hoechst (1:10000, H3570, ThermoFisher Scientific, Waltham, MA). Immunofluorescence images were obtained using LSM 900 Zeiss Laser-Scanning Confocal Microscope. To determine the percentage of tdTomato- positive podocytes, five glomeruli per kidney were selected from sections that were cut in close proximity to the middle of the glomeruli.

[0581] Intravenous injection of AAV9-CAG-Cre vector and kidney immunofluorescence in Col4a5 mice. Three 17-week-old LSL Col4a5 hemizygous male mice were treated with 6.8 xlO11vg of AAV9-CAG-Cre via intravenous (IV) injection. Affected mice at this age already manifest proteinuria. For this injection, the AAV vector stock was diluted and adjusted to 300 pL with 5% sorbitol / PBS. Kidneys were harvested from one mouse 10 weeks post-injection following perfusion with 4% paraformaldehyde (PFA) (15710, Electron Microscopy Sciences, Hatfield, PA). Kidneys from untreated hemizygous (He) and wild-type (WT) male mice were also harvested similarly, which served as negative and positive controls, respectively. Harvested kidneys were fixed in PFA and subsequently equilibrated in 30% sucrose (S5, Fisher Scientific, Hampton, NH) for cryoprotection. The fixed tissues were cryo-embedded in Tissue-Tek O.C.T. Compound (4583, Sakura Finetek, St. Torrance, CA) and cut into 5-pm-thick sections using a cryostat. Cryo-sections from the middle region of the left kidney were used to check the expression of COL4A5 and COL4A4 expression. Staining was performed using rat anti-COL4A5 (1: 100, 7078, Chondrex, Woodinville, WA), rat antiCOM A4 (1:50, b42 clone, received from Shigei Medical Research Institute, Japan. Kidney Int. 2004 Jul;66(l): 177-86), goat anti-rat IgG Cy3 (1:500, 112-165-167, Jackson ImmunoResearch, West Grove, PA). The sections were then counter stained with Hoechst (1 : 10000, H3570, ThermoFisher Scientific, Waltham, MA). Immunofluorescence images were obtained using BZ-X700 Fluorescence Microscope (Keyence, Osaka, Japan).

[0582] Intravenous injection of split intein AAV9-Col4A5 vectors, i.e., co-injection of AAV9-CAG-Nint2- Col4A5 and AAV9-CAG-Cint2-Col4A5. Twenty-six-week-old Col4a5 G5X XLAS hemizygous male mice were intravenously injected with AAV9-CAG-Nint2-Col4A5 and AAV9-CAG-Cint2-Col4A5 via that tail vein at a dose of 8.0 xlO11vg / mouse for each AAV vector (n=2). For vector injection, each AAV vector stock was diluted with 5% sorbitol / PBS and adjusted to a total volume of 300 pL. Three weeks postinjection, kidneys were harvested from mice following perfusion with 4% paraformaldehyde (PFA) (15710, Electron Microscopy Sciences, Hatfield, PA). Harvested kidneys were fixed in PFA and subsequently equilibrated in 30% sucrose (S5, Fisher Scientific, Hampton, NH) for cryoprotection. The fixed tissues were cryo-embedded in Tissue-Tek O.C.T. Compound (4583, Sakura Finetek, St. Torrance, CA) and cut into 5- pm-thick sections using a cryostat. Staining was performed using rabbit anti-HA (1:200, 3724, Cell Signaling, Danvers, MA), rat anti-COL4A5 (1:200, 7078, Chondrex, Woodinville, WA), rat anti-HA (1:200, 11867423001, Sigma, St. Louis, MO), rabbit anti-WTl (1:100, ab89901, Abeam, Cambridge, UK), rat antilaminin p2 / yl (1:100, MAI-06100, ThermoFisher Scientific, Waltham, MA), goat anti-rabbit IgG Alexa Fluor 647 (1:500, 111-605-144, Jackson ImmunoResearch, West Grove, PA), goat anti-rat IgG Cy3 (1:500, 112-165-167, Jackson ImmunoResearch, West Grove, PA) and Hoechst (1:10000, H3570, ThermoFisher Scientific, Waltham, MA). Immunofluorescence images were obtained using an BZ-X700 Fluorescence Microscope (Keyence, Osaka, Japan) Example 2: Selection of splice sites of the split intein COL4A5 constructs

[0583] The COL4A5 protein is coded by the COL4A5 gene comprising 51 coding exons. The major protein product of the COL4A5 gene is COL4A5 of 1685 amino acids in length while a protein product of 1691 amino acids in length has also been identified, which has additional 6 amino acids derived from two cryptic short exons. The COL4A5 of 1685 amino acids in length was chosen, coded by a 5,058 nucleotide-long open reading frame (ORF), and employed the Npu DnaE intein-based split intein system (Shah et al., J Am Chem Soc 135, 5839-5847 (2013); Padula et al., Mol Ther Methods Clin Dev 26, 495-504 (2022); Tornabene et al., Mol Ther Methods Clin Dev 23, 448-459 (2021); Truong et al., Nucleic Acids Res 43, 6450-6458 (2015); Iwai et al., FEES Lett 580, 1853-1858 (2006)), in which COL4A5 serves as a non-native extein. This system has a high degree of flexibility regarding the sequences split into N-halves and C-halves, except for the requirement that the first amino acid residue in the C-halves must be cysteine (Cys), serine (Ser) or threonine (Thr), which contains a thiol group, -SH, or a hydroxyl group, -OH, in their side chains (Apgar et al., PLoS One 7, e37355 (2012); Perler Nucleic Acids Res 30, 383-384 (2002)). Positions of Cys, Ser, and Thr in the 1685-amino-acid-long COL4A5 were identified (FIGs. 1A-1C). There were a total of 20, 64, and 37 positions for Cys, Ser, and Thr residues, respectively. Regarding gene expression cassettes used for COL4A5 expression, those that are driven by the 1 -kb C AG promoter with or without the woodchuck hepatitis virus post-transcriptional regulatory element (WPRE), and with the 0.2-kb SV40 polyadenyl tion signal were chosen. This cassette can drive robust gene expression in podocytes (Furusho et al., bioRxiv / 2023 / 548760 (2023) and Furusho et al. Nat Commun 15, 10728 (2024)). This therapeutic payload design accommodates a protein-coding sequence with a length of 3,360 base pairs (bps), which are equivalent to 1,120 amino acids, assuming that packaging size limit for AAV vectors is approximately 5 kilo bases (kbs). With these factors taken into account, a total of 25 candidates for potential split points were identified, as outlined in FIGs 1A-1C.

[0584] It remained challenging to identify intein insertion sites or extein split sites that successfully mediate protein splicing in non-native exteins, because the efficacy of protein splicing is significantly influenced by its neighboring extein sequences (Shah et al., J Am Chem Soc 135, 5839-5847 (2013); Iwai et al., FEES Lett 580, 1853-1858 (2006); Zettler et al., PLoS One 8, e72925 (2013); Amitai et al., Proc Natl Acad Sci U SA 106, 11005-11010 (2009); Cheriyan et al., The Journal of biological chemistry 288, 6202-6211 (2013); Lockless et al., Proc Natl Acad Sci USA 106, 10999-11004 (2009); Ho et al., Nat Commun 12, 2200 (2021); Lee et al., PLoS One 7, e43820 (2012)). In addition, the protein trans-splicing approach cuts a non- native extein into two pieces as separate protein molecules. This process increases the risk of unforeseen problems including issues in protein folding and stability. In addition, potential candidate split sites that fulfill the Cys / Ser / Thr requirement at +1 position in the N-half extein do not necessarily tolerate an intein insertion (Ho et al., Nat Commun 12, 2200 (2021); Lee et al., PLoS One 7, e43820 (2012)). Therefore, various efforts have been made to understand the mechanisms of protein splicing and develop new methods that can logically identify appropriate split sites. Several strategies have been employed successfully to identify functionally viable split points and create split intein trans-splicing proteins. These include: targeting structurally flexible and phylogenetic ally less conserved regions as split points (Ho et al., Nat Commun 12, 2200 (2021)), inclusion of functionally characterized in tein-extein junction sequences or modification of sequences around the split sites in exteins (Shah et al., J Am Cliem Soc 135, 5839-5847 (2013); Amitai et al., Proc Natl Acad Sci U SA 106, 11005-11010 (2009); Ho et al., Nat Commun 12, 2200 (2021); Palanisamy et al., Nat Commun 10, 4967 (2019)), and use of computational models utilizing split energy profiles or scoring similarities to the native extein splice site cassettes by machine learning (Apgar et al., PLoS One 7, e37355 (2012); Dagliyan et al., Nat Commun 9, 4042 (2018)). Due to the intricate nature of the effective protein splicing mechanisms, systematic screening has also been employed (Ho et al., Nat Commun 12, 2200 (2021); Dagliyan et al., Nat Commun 9, 4042 (2018)). Despite this, the mechanisms of protein splicing, and methods that can logically identify appropriate split sites remain elusive.

[0585] The logical design of split intein type IV collagen alpha chains for AAV vector-mediated gene therapy for Alport syndrome presents a considerable challenge, given the multifaceted constraints involved. First, split sites were selected from 25 amino acid positions to account for size limitations for the payloads for AAV vectors (FIGs. 1A-1C). Second, among the chosen sites, 20 were in the collagenous domain (CD) and 5 were in the non-collagenous regions that interrupt CD. The CD contains evolutionary conserved G-X- Y motif, forming structurally conserved triple helices. The majority of the pathogenic missense mutations, which account for approximately 40% of COL4A5 mutations causing XLAS (Gubler et al., Nat Clin Pract Nephrol 4, 24-37 (2008)) involve the glycine residue in the G-X-Y motif in CD. Given that heterotrimerization takes place in the early stage of type IV collagen heterotrimer biogenesis in the endoplasmic reticulum (ER) before secretion, it was not unreasonable to assume that a CD split by the split intein fusion has detrimental effects on protein stability, folding, heterotrimerization, and secretion. Third, in the discipline of gene delivery, the split intein approach has been exclusively used for non-secreted proteins (Padula et al., Mol Ther Methods Clin Dev 26, 495-504 (2022); Tornabene et al., Sci Transl Med 11 (2019); Li et al., Hum Gene Ther 19, 958-964 (2008)) and there has been no study that investigated how to design and create split intein constructs for secreted proteins. Coagulation factor 8 (F8, a secreted protein) has been reconstituted from N-half (heavy chain) and C-half (light chain) via a split intein dual AAV vector approach (Chen et al., Mol Ther 15, 1856-1862 (2007); Esposito et al., EMBO Mol Med 14, el5199 (2022); Zhu et al., Sci China Life Sci 56, 262-267 (2013)); however, the split intein is dispensable in the successful F8 secretion as reproducibly demonstrated by Burton et al., Proc Natl Acad Sci USA 96, 12725-12730 ( 1999) and Scallan et al., Blood 102, 3919-3926 (2003). This is because the covalent linkage of the heavy and light chains takes place during the natural process of F8 activation. By contrast, COL4A5 C-half exteins most likely require a signal sequence for correct intracellular trafficking and secretion. This necessitates additional protein engineering that would not be necessary for split intein constructs for non-secreted exteins.

[0586] With these difficulties posed in designing split intein constructs, potential COL4A5 splice sites based on the similarities to the native extein splice sites were selected, which could be an important feature identified in successfully constructed non-native exteins split by inteins (Apgar et al., PLoS One 7, e37355 (2012)). An SVM model was trained on 752 splice sites to determine extein splice site. This SVM model exhibited a predictive power with the area under the receiver operating characteristic (ROC) curve of 0.881 (FIG. 2). The confusion matrix revealed that the sensitivity, specificity and overall accuracy of the similarity prediction by the SVM model were 0.761, 0.904 and 0.868, respectively (FIG. 3). Using this machine learning-based prediction model, a total of 6 splice site cassettes similar to those of native exteins among the 25 candidate sites were identified.

[0587] 10 split sites from the 25 candidate sites, 9 from CD and 1 from the non-collagenous interrupting region, were selected (FIGs. 1A-1C). Among them, 3 split sites exhibited similarity to the native intein splice sites in amino acid sequence compositions. Two sites near but outside the cluster of the sites that can fit in the AAV-CAG vector cassette were also selected (FIG. 6A). The same approach can be employed for split intein dual AAV vector-mediated gene therapy for autosomal recessive Alport syndrome (ARAS) and autosomal dominant Alport syndrome (ADAS), which are caused by mutations of either the COL4A3 gene or the COL4A4 gene using FIGs. 4A-4D and 5A-5D.

[0588] Example 3: Design of the split intein COL4A5 constructs

[0589] After testing several different enhancers and promoters for AAV vector-mediated transgene expression in C57BL / 6 mice, it was found that the CAG promoter combined with the woodchuck hepatitis virus post-transcriptional regulatory element (WPRE) and the SV40 polyadenylation signal demonstrates the most robust gene expression in the kidney following intravenous injection of AAV9 and AAV-KP1 vector (Furusho et al., bioRxiv / 2023 / 548760 (2023) and Furusho et al. Nat Common 15, 10728 (2024)), better than the CMV enhancer-promoter with the human beta-globin intron sequence. With this observation, the CAG- WPRE-driven expression cassette was selected for the proof-of-concept studies aimed at developing split intein dual AAV vector-mediated gene therapy for Alport syndrome. The following alternative CAG promoter-driven expression cassettes were also used: CAG-WPRE3-SV40pA and CAG-SV40pA devoid of WPRE. These two gene expression cassettes were derived from pAAV-CAG-tdTomato (Addgene 59462), WPRE3 sequence was derived from Choi et al., Mol Brain 7, 17 (2014).

[0590] In the split intein COL4A5 approach, since both N-half and C-half of COL4A5 need to traffic to the ER, either the 26-amino-acid-long COL4A5 native signal sequence or the 18-amino-acid-long human chymotrypsin signal sequence were added to the N-terminus of the COL4A5 C-half (FIG. 6A). The functionality of the chymotrypsin signal sequence fused with a non-secreted protein was tested in the context of the human apoptosis inducing factor mitochondria associated l(AIFMl) protein in HEK293 cells in vitro, demonstrating effective secretion of AIFM1 protein into the culture media. As a control, a COL4A5 C-half construct devoid of any signal peptide was created (FIG. 6A). Example 4: In vitro assessment of protein trans-splicing using HEK293 cells

[0591] Plasmids carrying split intein COL4A5 N-half or C-half ORF between the two AAV2 ITRs listed in FIG. 8 were constructed. A plasmid carrying the full-length COL4A5 ORF, pAAV-CAG-Col4A5-FLAG- WPRE, as the full-length control was also created.

[0592] Split intein COL4A5 C-half signal peptide for effective expression. First, it was investigated whether a signal peptide can improve effective expression of split intein COL4A5 C-half protein fragment by an in vitro plasmid DNA transfection experiment using the following three split point 1 (SP 1) plasmids, pAAV- CAG-noSS-Cint-SPl-Col4A5-FLAG-WPRE, pAAV-CAG-nativeSS-Cint-SPl-Col4A5-FLAG-WPRE, and pAAV-CAG-chymoSS-Cint-SPl-Col4A5-FLAG-WPRE. To this end, HEK293 cells were transfected with the respective plasmids and protein levels were quantified by western blot. The result demonstrated that the presence of a signal sequence at the N-terminus of the split intein COL4A5 C-half construct dramatically improves the C-half expression (FIG. 6B). The same experiment using pAAV-CAG-noSS-Cint-SP2- Col4A5-FLAG-WPRE and pAAV-CAG-noSS-Cint-SPl l-Col4A5-FLAG-WPRE confirmed that no SP2 and SP11 split intein COL4A5 C-half products were detected at an appreciable level (FIGs. 9A-9D).

[0593] Both native COL4A5 and chymotrypsin signal peptides mediated effective expression of the split intein COL4A5 C-half protein, and induced protein trans-splicing when co-expressed with a respective split intein COL4A5 N-half protein. The above-described experiment also demonstrated that there was no appreciable difference in split intein COL4A5 C-half protein expression levels between the native COL4A5 signal peptide construct and the chymotrypsin signal peptide construct (FIG. 6B). To investigate whether protein trans-splicing takes place when split intein COL4A5 N-half and C-half constructs are co-expressed, the HEK293 cell experiment was performed in the same manner explained above using the following split point 1 (SP1) split intein COL4A5 plasmid combinations: (1) pAAV-CAG-Col4A5-FLAG-WPRE (the full- length control), (2) pAAV-CAG-SPl-Col4A5-Nint-FLAG-WPRE and pAAV-CAG-nativeSS-Cint-SPl- Col4A5-FLAG-WPRE, (3) pAAV-CAG-SPl-Col4A5-Nint-FLAG-WPRE and pAAV-CAG-chymoSS-Cint- SPl-Col4A5-FLAG-WPRE, and (4) pAAV-CAG-SPl-Col4A5-Nint-FLAG-WPRE and pAAV-CAG-noSS- Cint-SPl-Col4A5-FLAG-WPRE. Crude cell lysates prepared 48 hours post-transfection were subjected to western blot analysis using anti-FLAG antibody (FIG. 6C). The results demonstrated successful reconstitution of the full-length COL4A5 in cells by protein trans-splicing, using both the native COL4A5 and chymotrypsin signal sequence C-half constructs. Unspliced N-half products were observed in both C- half constructs. Split point 2 (SP2) constructs were investigated in the same manner using the following split intein COL4A5 plasmids: (1) pAAV-CAG-Col4A5-FLAG-WPRE (the full-length control), (2) pAAV- CAG-SP2-Col4A5-Nint-FLAG-WPRE and pAAV-CAG-nativeSS-Cint-SP2-Col4A5-FLAG-WPRE, (3) pAAV-CAG-SP2-Col4A5-Nint-FLAG-WPRE only, and (4) pAAV-CAG-nativeSS-Cint-SP2-Col4A5- FLAG-WPRE only. Crude cell lysates prepared 48 hours post-transfection were subjected to western blot analysis using anti-FLAG antibody (FIG. 6D). Both the Col4A5-Nint-FLAG and Cint-Col4A5-FLAG were expressed independently with the N-half construct being expressed at a level much higher than the full- length or the C-half constructs, which was also the case with SP1 constructs. Importantly, these results demonstrated successful reconstitution of the full-length COL4A5 in cells by protein trans-splicing when both the Col4A5-Nint-FLAG and Cint-Col4A5-FLAG constructs were co-expressed. The degree of reconstitution was excellent compared to the full-length construct although a substantial amount of N-half constructs remained unspliced presumably due to the excessive presence of the N-half proteins over the C- half proteins.

[0594] Demonstration of trafficking to the secretory pathway in human podocytes indicates the functional viability of the reconstituted protein by trans-splicing. To assess the secretion capability of the full-length COL4A5 reconstituted by trans-splicing, immortalized undifferentiated human podocyte cell line AB8 / 13 was differentiated into terminally differentiated podocytes as described (Saleem et al., J Am Soc Nephrol 13, 630-638 (2002)) and used to express the split intein COL4A5 constructs. To this end, the AAV-KP1 vector expressing split intein COL4A5 constructs of interest was produced. It has been shown that AAV-KP1 vector can effectively transduce terminally differentiated AB8 / 13 podocytes in vitro. Podocytes were infected with AAV-KP1 vectors expressing split intein COL4A5 proteins (AAV-KPl-CAG-SP2-Col4A5- Nint-FLAG-WPRE and AAV-KPl-CAG-nativeSS-Cint-SP2-Col4A5-HA -noWPRE) or tdTomato (AAV- KPl-CAG-tdTomato), and five days after infection, 1 mL of culture media was collected and concentrated for western blot analysis.

[0595] The results show that fully reconstituted COL4A5 expressed by the split intein dual AAV-CAG- Col4A5 vectors is capable of trafficking to the secretory pathway and is secreted from AAV vector- transduced cells effectively (FIG. 6E). These results also indicate that COL4A5 expressed by AAV vectors can form heterotrimers with COL4A3 and COL4A4 expressed in terminally differentiated human podocytes.

[0596] N-half vector dose can be significantly reduced for effective trans-splicing in the split intein dual AAV-Col4A5 vectors. It is disclosed herein that when the COL4A5 protein is split into N-half and C-half and fused with split inteins, the steady-state level of the N-half protein expression was substantially increased compared to that of the full-length COL4A5 and the COL4A5 C-half proteins (FIG. 6D). This unequal levels of split intein COL4A5 N-half and C-half expression, when their coding DNA templates are delivered into cells at a 1-to-l ratio, indicated that the amount of the C-half construct could be reduced to mediate effective trans-splicing. To explore this possibility, in vitro experiments were conducted in which HEK293 cells were transfected with split intein COL4A5 plasmids, as detailed earlier. In this experiment, the following ratios of pAAV-CAG-SP2-Col4A5-Nint-FLAG-WPRE and pAAV-CAG-nativeSS-Cint-SP2- Col4A5-FLAG-WPRE were tested: 1:1, 0.050:1. 0.025:1, 0.017:1, 0.013:1, 0.010:1, 0.005:1, and 0.001:1. Consequently, the result showed that the amount of the N-half plasmid DNA can be substantially reduced relative to the amount of the C-half plasmid up to 100 fold for effective protein trans-splicing (FIG. 10).

[0597] Experimental identification of split points in COL4A5 that mediate most effective protein trans- splicing. Efficiency of trans-splicing can vary depending on how exteins are split by inteins as detailed above. To experimentally assess functional viability of the COL4A5 splice split point candidates and identify the most effective split points, an in vitro plasmid transfection experiment using HEK293 cells primarily in the same manner as detailed earlier was conducted, total cell lysates were collected, and western blot was performed using anti-FLAG M2 antibody or rabbit polyclonal anti-human COL4A5 H53 antibody. Among 12 splice site candidates that were initially selected (FIGs. 1A-1C), 9 split points (SP1, SP2, SP3, SP4, SP5, SP7, SP9, SP11 and SP12) were initially assessed. Western blot analysis revealed that all constructs, except for SP3, SP4 and SP1, could mediate trans-splicing leading to the reconstruction of the full-length COL4A5 in HEK293 cells (FIGs. 9A-9D). The SP2 constructs consistently yielded robust trans- splicing in quadruplicated experiments. Although the results observed with some split points were somewhat inconsistent in replicated experiments, SP7 constructs consistently showed higher yields of full-length transspliced COL4A5 proteins than other split points, indicating that SP7 could be an effective split point.

[0598] Subsequently, an additional 17 pairs of the split intein COL4A5 constructs were constructed (SEQ ID NO: 122-193), creating a panel of 27 distinct SP constructs, SP1-SP12 and SP14-SP28, spanning from C451 to SI 136. They include 25 potential split points between S565 and SI 111, which lit in the dual vector system. These 27 pairs were then tested for PTS in HEK293 cells, followed by western blot. It was found that 16 (SP1, SP3, SP4, SP14-SP19, SP21-SP26, and SP28) failed to undergo PTS in HEK293 cells, while 11 (SP2, SP5-SP12, SP20, and SP27) successfully generated spliced products at varying efficiencies (FIG. 9G). Consequently, it was identified that SP6 and SP7, and potentially SP2 as well (with consideration for potential N-half dose reduction as indicated in FIG. 10), as the most optimal split point candidates. Additionally, the present disclosure provides insights into the motifs in COL4As that facilitate PTS (FIG. 9H). Using these 27 labeled data as a training set (SEQ ID NOs: 194-220), a support vector machine (SVM) model was trained. It was found that the model consistently achieved high accuracy in predicting the split point viability across all iterations of leave-one-out cross-validation. Thus, the present disclosure provides robust means to streamline the process of making novel split intein constructs for various collagen molecules.

[0599] Split point effects on the steady- state level expression of split intein COL4A5 N-half and C-half constructs. As presented earlier and shown in FIG. 6D, the split intein N-half constructs can be expressed at levels much higher than those of full-length COL4A5 or split intein C-half constructs in HEK293 cells when similar amounts of genetic payloads coding for respective proteins are delivered to the cells by plasmid DNA transfection. It could be that the N-half of COL4A5 is more resistant to protein degradation than the full-length COL4A5 or the C-half of COL4A5, and these attributes are molecular mass-dependent. To better understand the consequences (such as expression level in cells) of splitting COL4A5 into split intein N-half and C-half, an in vitro plasmid transfection experiments using HEK293 cells primarily in the same manner detailed above was conducted. In this experiment, a total of 9 split points (SP1, SP2, SP3, SP4, SP5, SP7, SP9, and SP11) were assessed for protein expression levels with pairs of COL4A5 N-half and C-half constructs (FIGs. 11A-11B). HEK293 cells were transfected with plasmid expressing split intein N-half or C-half constructs of various lengths harboring different split points, and 48 hours later, crude cell lysates were subjected to SDS-PAGE using a 6% gel for N-half constructs and a 10% gel for C-half constructs, followed by western blot using anti-FLAG M2 antibody. No significant differences were observed in the protein expression levels of split intein N-half constructs with varying split points. However, intriguingly, there was a molecular mass-dependent decrease in the steady-state levels of split intein COL4A5 C-half constructs. It may be that having split points more towards the C-terminal side could be advantageous in achieving a more balanced expression of N-half and C-half constructs, which might yield more trans-spliced products.

[0600] Example 5: AAV vector-mediated gene therapy for Alport syndrome using a loxP-stop-loxP (LSL)- Col4a5 XIAS mouse model.

[0601] Generation of LSL-Col4a5 XIAS mice. Therapeutic payloads for type IV collagen alpha (COL4A) chain gene replacement therapy for Alport syndrome do not fit in a single AAV vector because the ORFs of the human C0L4A genes by themselves are 5.01 to 5.06 kb in lengths, which exceed the packaging capacity of AAV vectors. In addition, AAV gene therapy for kidney diseases has remained elusive, in part because effective gene delivery to the kidney remains challenging and hampered by AAV’ s low transduction efficiency in target cells including podocytes, the factory for COL4A proteins (Zincarelli et al., Mol Ther 16, 1073-1080 (2008); Murata et al., Eur Radiol 18, 1464-1472 (2008); Saito et al., Clin Exp Nephrol 23, 1345- 1356 (2019); Rubin et al., Hum Gene Ther 30, 1559-1571 (2019); Asico et al., Biochem Biophys Res Commun 497, 19-24 (2018); Shen et al., Hum Gene Ther Methods 29, 251-258 (2018); Ikeda et al., J Am Soc Nephrol 29, 2287-2297 (2018); Konkalmatt et al., JCI Insight 1 (2016); Rocca et al., Gene Ther 21 , 618-628 (2014); Picconi et al., Mol Ther Methods Clin Dev 1, 14014 (2014); Chung et al., Nephron Extra 1, 217-223 (2011); Schievenbusch et al., Mol Ther 18, 1302-1309 (2010); Ito et al., BJU international 101, 376-381 (2008); Lipkowitz et al., J Am Soc Nephrol 10, 1908-1915 (1999); Woodard et al., J Vis Exp, 56324 (2018); Davis et al., Physiol Genomics 51, 449-461 (2019)). To demonstrate AAV vector-mediated gene therapy for Alport syndrome, the LSL Col4a5 knock-in XLAS mouse model was developed, and is disclosed herein. In this model, Col4a5 protein is restored under the control of the native Col4a5 gene enhancer-promoter upon Cre recombinase expression. The ORF of Cre recombinase is 1.03 kb, and AAV- Cre vectors can be used to restore expression of genes of interest in respective LSL mouse models. The LSL-Col4a5 XLAS mouse model was generated by introduced a loxP-stop-loxP cassette, including three copies of the SV40 polyadenylation signals (i.e., SV40pA 3x sequence) flanked with a pair of loxP sites, into intron 1, between mouse chromosome X positions 141,510,668 and 141,510,669 (mmlO), of the mouse Col4a5 gene (FIG. 12A).

[0602] FIGs. 12B-12E shows characteristics of this mouse model. Statistically significant differences in body weight, urine albumin levels, blood biomarkers, and survival probability with adjusted p<0.05 were observed from 14 weeks (Student's t-test with Bonferroni correction). Similar to Col4a5 G5X XLAS mice, LSL-Col4a5 XLAS mice also showed progressing chronic kidney diseases (CKDs) clinically and histologically. These observations confirm that LSL-Col4a5 XLAS mice are a good animal model of Alport syndrome which can be used to assess in vivo efficacy of novel therapeutic approaches.

[0603] Intravenous (IV) injection of AAV9 transduces podocytes in the XLAS-associated CKD kidney. To investigate how AAV9 transduces the kidney in healthy and XLAS-associated CKD kidneys via IV injection, Col4a5 G5X hemizygous male mice (Rheault et al., J Am Soc Nephrol 15, 1466-1474 (2004)), an established XLAS mouse model, and age-matched wild-type (WT) male mice were injected with AAV9- CAG-tdTomato or AAV-KPl-CAG-tdTomato vector. AAV9 and AAV-KP1 are contrastive capsids showing distinct renal transduction profiles and distinct pharmacokinetics (PK) in the blood. AAV9 and AAV-KP1 demonstrate very slow and very rapid blood clearance following IV injection, respectively (Kotchey et al., Mol Ther 19, 1079-1089 (2011); Furusho et al., bioRxiv / 2023 / 548760 (2023); Furusho et al. Nat Commun 15, 10728 (2024)). Urinary albumin levels at the time of injection were significantly higher in Col4a5 G5X than in WT, indicating increased glomerular filtration barrier permeability (Furusho et al., bioRxiv / 2023 / 548760 (2023) and Furusho et al. Nat Commun 15, 10728 (2024)). Two weeks following injection, the CKD kidney showed substantially enhanced podocyte transduction with AAV9 compared to the healthy kidney, achieving 35% podocyte transduction (FIGs. 13A-13B). Notably, the degree of enhancement was greater with AAV9 than that with AAV-KP1, indicating that the long blood circulation time can be a podocyte transduction-enhancing factor. (Furusho et al., bioRxiv / 2023 / 548760 (2023) and Furusho et al. Nat Commun 15, 10728 (2024)). This establishes that IV injection of AAV9 vector is a reasonable and logical approach for AAV vector-mediated gene therapy for Alport syndrome. Augmented podocyte transduction in CKD kidneys was demonstrated by Ding et al., Sci Transl Med 15, eabc8226 (2023).

[0604] Intravenous injection of AAV9-CAG-Cre vector transduces podocytes and restores type IV collagen expression in LSL-Col4a5 XLAS mice. LSL Col4a5 hemizygous male mice were treated by IV injection of AAV9-CAG-Cre. Immunofluorescence microscopic analysis revealed that many glomeruli were transduced with AAV9-CAG-Cre vector, resulting in the expression of tdTomato within the glomeruli. COL4A4 and COL4A5 signals, which were missing in untreated hemizygous male mice, colocalized within the membranous structure of the glomeruli, which most likely represents glomerular basement membrane (GBM) (FIGs. 14A-14B). This strongly indicates that AAV9 vector effectively transduced podocytes, leading to the restoration of heterotrimeric type IV collagen expression and its incorporation into the GBM.

[0605] Demonstration of therapeutic effects of AAV gene therapy for X-linked Alport syndrome (XLAS) in the LSL-Col4a5 XLAS mouse model. In the above-described experiment in which Col4a5-LSL hemizygous (He) male mice were treated with AAV9-CAG-Cre vector by intravenous (IV) injection at 17 weeks of age, two mice were kept under monitoring by weekly measurement of body weight. Surprisingly, the vector- treated XLAS mice survived with no loss of body weight at least until 36 weeks of age, surpassing the age by which all untreated XLAS mice would succumb or reach the endpoint (FIG. 14C). The macroscopic examination of the XLAS kidneys of these AAV vector-treated mice revealed unambiguous improvement, showing normal kidney size and shape, with surfaces appearing relatively smooth and lacking the evident granular appearance observed in untreated counterparts. This establishes that AAV9 IV gene therapy can mediate therapeutic effects in XLAS mice. Notably, the gene therapy has the capacity to halt the disease progression even when administered in the advanced stage of the established disease. In another experiment, LSL-Col4a5 hemizygous (He) male mice (n=l l) were subjected to treatment with AAV9-CAG-Cre vector by retro-orbital injection at 4 weeks of age. Under anesthesia, the mice received 50 uL of AAV vector solution (3.0 x 1011vg / mouse) injected into the right eyes. These mice were monitored with weekly measurements of body weight. The AAV vector-treated XLAS mice exhibited body weight trajectories similar to those of untreated XLAS mice until 20 weeks of age, (FIG 14D). To assess urine albumin, spot urine was collected and albumin-to-creatinine ratio (ACR) was determined. Urine albumin was quantified using Albuwell M (1011, Ethos Biosciences, Logan Township, NJ) and was normalized by urine creatinine concentration measured by Creatinine Reagent Assay (C75391250, Pointe Scientific, Canton, MI). The ACR value was expressed as grams per gram creatinine. All the assays were performed as per manufacturer’s instruction. The data indicated that AAV vector treatment effectively prevented albumin leakage into the urine in the given time points (FIG. 14E). This demonstrates that AAV9 gene therapy can mediate therapeutic effects in XLAS mice, leading to biomarker improvement when injected at an early age.

[0606] Therefore, using this mouse model, the present disclosure demonstrates that AAV9-CAG-Cre IV injection into LSL-Col4a5 hemizygous (He) male mice at 17 weeks of age could halt the progression of CKD (FIGs. 14F, 14H). It was confirmed that Col4a4 and Col4a5 expressions were restored in the glomeruli of all the mice euthanized at 10 or 22 weeks post-injection. Additionally, AAV treatment significantly reduced glomerulosclerosis in the kidneys, and tubules showed improvement in structure (FIG. 14H). When the LSL-Col4a5 XLAS mice were treated with IV injection of AAV9-CAG-Cre at 4 weeks of age, they showed more effective therapeutic effects with greater biomarker improvement (FIGs. 14D-14G). These observations unambiguously demonstrate that IV-injected AAV9 can cross the GFB and transduce podocytes before the onset of CKD in XLAS mice. These establish that AAV9 IV gene therapy can mediate therapeutic effects in LSL-Col4a5 XLAS mice irrespective of whether administered before or after the onset of CKD.

[0607] AAV9 IV gene therapy at 4 weeks of age extended the survival by 59% (FIG. 14G). Further followup of the mice, until all reached the experimental endpoint due to kidney failure, revealed that the therapy increased the median survival time from 31.5 to 50.0 weeks representing a 59% improvement (FIG. 14G). This further confirms the therapeutic potential of AAV9 IV gene therapy.

[0608] Example 6: Split intein AAV vector-mediated gene therapy for Alport syndrome using Col4a5 G5X XLAS mice.

[0609] Production of split intein AAV9-CAG-Col4A5 vectors using the split point 2 (SP2). Given that intravenous injection of AAV9 vector into mice affected with CKD can effectively transduce podocytes (as shown herein; Ding et al., Sci Transl Med 15, eabc8226 (2023); Furosho et al., bioRxiv / 2023 / 548760 (2023)), the AAV9 capsid was used for this study. As for the split point, SP2 was used due to that fact that the SP2 split intein constructs have consistently been shown to be very effective split sites. The AAV-CAG- SP2-Col4A5-Nint-FLAG-WPRE vector genome for the split intein N-half vector (4.6 kb) and the AAV- CAG-nativeSS-Cint-SP2-Col4A5-HA-noWPRE vector genome for the split intein C-half vector, (4.8 kb) were used. Two AAV9 vectors harboring these two vector genomes, namely AAV9-CAG-Nint2-Col4A5 and AAV9-CAG-Cint2-Col4A5, were produced in adherent HEK293 cells and purified by two rounds of cesium chloride (CsCl) density-gradient ultracentrifugation followed by dialysis as described earlier.

[0610] Intravenous injection of split intein AAV9-Col4A5 vectors, i.e., co-injection of AAV9-CAG-Nint2- Col4A5 and AAV9-CAG-Cint2-Col4A5, restored COL4A5 expression in podocytes and surrounding glomerular basement membrane ( GBM) in Col4a5 G5X XLAS mice. Animals were intravenously injected with AAV9-CAG-Nint2-Col4A5 and AAV9-CAG-Cint2-Col4A5, and immunofluorescence was performed on kidneys.

[0611] AAV vectors effectively transduced podocytes and expressed COL4A5 in the kidney of XLAS mice. The membranous tdTomato staining pattern associated with WT1 (a podocyte marker) and overlapped with laminin (a GBM marker) demonstrates restoration of type IV collagen expression in GBM by AAV vector-mediated gene delivery (FIGs. 15A-15C). In contrast to humans, wild-type mice express Col4a3, Col4a4, and Col4a5 in renal tubules. Consequently, the tubular basement membrane (TBM) in mice is positive for these collagen alpha chains (refer to FIGs. 14A-14B). Intravenously injected AAV9 vectors effectively transduced proximal tubules in CKD kidneys (see also Furusho et al., bioRxiv / 2023 / 548760 (2023) and Furusho et al. Nat Commun 15, 10728 (2024)). Therefore, the overlap of tdTomato or HA signal with the TBM structure indicates the restoration of type IV collagen expression in the TBM as well. This observation is a logical consequence of split intein AAV-Col4A5 vector transduction in renal tubules. Taken altogether, these observations strongly indicate that the split intein AAV-Col4A5 vectors transduced podocytes effectively in CKD kidney, and resulted in the expression of trans-spliced full-length functional COL4A5 protein, forming cross-species heterotrimeric type IV collagen with mouse Col4a3 and Col4a4 and integrating into the GBM (Heidet et al., Am J Pathol 163, 1633-1644 (2003)).

[0612] Example 7 : Relationship between % podocyte transduction and therapeutic effects in AAV gene therapy for X-linked Alport syndrome (XLAS).

[0613] The deposition of collagen IV protein in the glomerular basement membrane (GBM) is not restricted to distinct, exclusive segments assigned to individual podocytes. Instead, the GBM segments where each podocyte contributes to collagen IV deposition overlap extensively, creating a continuous and interconnected distribution of individual podocytes' segments for collagen IV deposition within the GBM. This overlapping nature makes it challenging to definitively demarcate specific GBM segments attributable to individual podocytes. Therefore, it is not feasible to accurately quantify AAV vector-mediated podocyte transduction efficiency using vector-mediated Col4a5 / COL4A5 expression as a marker. However, it is reasonable to propose that assessing the percentage of Col4a5 deposition in the GBM (% deposition) is a valid approach for understanding the level of podocyte transduction (% transduction) required to achieve therapeutic effects in XLAS mice. Although % deposition may slightly overestimate the true % transduction due to the interconnected and overlapping nature of Col4a5 deposition as described above, it is unlikely to result in underestimation. Thus, % deposition was determined in the kidneys harvested from two 32-week- old LSL-Col4a5 XLAS mice treated with IV injection of AAV9-CAG-Cre at 4 weeks of age showing therapeutic effects. The % depositions were 6.0% and 6.1% (FIGs. 16A-16C). This indicates that AAV gene therapy may achieve therapeutic effects with as low as ~6% or less podocyte transduction, further underscoring the significance of developing AAV gene therapy for Alport Syndrome.

[0614] Demonstration of the restoration of Col4a4 expression in XLAS mice treated with IV injection of the split intein dual AAV9-CAG-COL4A5. The results (FIG. 16D) strongly indicates that this split intein dual vector approach successfully delivers fully functional COL4A5 protein in podocytes in XLAS mice, restoring deposition of functional heterotrimers of collagen IV tx3a4tx5 chains in the GBM.

[0615] In view of the many possible examples to which the principles of our invention may be applied, it should be recognized that illustrated examples are only examples of the invention and should not be considered a limitation on the scope of the invention. Rather, the scope of the invention is defined by the following claims. We therefore claim as our invention all that comes within the scope and spirit of these claims.

Claims

We claim:

1. A split intein system comprising:(i) a first nucleic acid molecule encoding a fusion protein comprising in N to C terminal order, an N- extein of an extein pair and an N-intein of an intein pair, wherein the N-extein comprises a first signal sequence N-temiinal to an N-terminal portion of a type IV collagen alpha chain; and(ii) a second nucleic acid molecule encoding a fusion protein comprising in N to C terminal order, a second signal sequence, a C-intein of the intein pair, and a C-extein of the extein pair, wherein the C-extein comprises a C-terminal portion of the type IV collagen alpha chain; and wherein the type IV collagen alpha chain is COL4A5; the N- and C-terminal portions of the type IV collagen alpha chain together comprise the type IV collagen alpha chain sequence separated at a split point; and the N- and C-exteins are spliced together to form the mature type IV collagen alpha chain when the first and second nucleic acid molecules are expressed in mammalian cells.

2. The split intein system of claim 1, comprising:(i) a first adeno-associated viral (AAV) vector comprising the first nucleic acid molecule; and(ii) a second AAV vector comprising the second nucleic acid molecule.

3. The split intein system of claim 2, wherein the first and / or the second AAV vector is AAV9.

4. The split intein system of claim 2, wherein the first and / or the second AAV vector is AAV-KP1.

5. The split intein system of claim 2, wherein the first and / or the second AAV vector is AAV1,AAVl_9mtl00, AAVl_9mt30, AAVl_9mt76, AAV2, AAV2G9, AAV2i8, AAV2retro, AAV2R585E, AAV2R585E9_2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV9AA272, AAV9AA22, AAV9W22A, AAV10, AAV11, AAV2.7m8, AAVAnc80, AAVbb.2, AAV-DJ, AAVHN1, AAVHN2, AAVHN3, AAVhu.ll, AAVhu.13, AAVhu.37, AAV-KP1, AAV-KP2, AAV-KP3, AAVLK03, AAVNP40, AAVNP59, AAVPHP.B, AAVPHP.eB, AAVPHP.S, AAVpol, AAVrh.8, AAVrh.10, AAVrh.20, AAVrh.43, or AAVShHIO.

6. The split intein system of claim 1, wherein the intein pair is a Npu DnaE intein pair.

7. The split intein system of claim 6, wherein the N-intein comprises the amino acid sequence of SEQ ID NO: 26 or an amino acid sequence at least 95% identical thereto; and the C-intein comprises the amino acid sequence of SEQ ID NO: 27 or an amino acid sequence at least 95% identical thereto.

8. The split intein system of claim 1, wherein, the first signal sequence, the N-terminal portion of the type IV collagen alpha chain, and the N-intein are fused directly.

9. The split intein system of claim 1, wherein, the second signal sequence, the C-intein, and the C- terminal portion of the type IV collagen alpha chain are fused directly.

10. The split intein system of claim 1, wherein the split intein system comprises the first nucleic acid molecule and the second nucleic acid molecule at a ratio of about 1: 1 to about 0.010: 1.

11. The split intein system of claim 10, wherein the split intein system comprises the first nucleic acid molecule and the second nucleic acid molecule at a ratio of about 0.017: 1 to about 0.010: 1.

12. The split intein system of claim 1, wherein the split point is within a collagenous domain of the type IV collagen alpha chain.

13. The split intein system of claim 1 , wherein the first signal sequence and / or the second signal sequence is a collagen signal sequence or a chymotrypsin signal sequence.

14. The split intein system of claim 13, wherein the first signal sequence and the second signal sequence are different.

15. The split intein system of claim 13, wherein the collagen signal sequence comprises SEQ ID NO: 4 and / or wherein the chymotrypsin signal sequence comprises SEQ ID NO: 5.

16. The split intein system of claim 1, wherein the type IV collagen alpha chain comprises the amino acid sequence of any one of SEQ ID NOs: 15-25 or an amino acid sequence at least 95% identical thereto.

17. The split intein system of claim 16, wherein the type IV collagen alpha chain comprises the amino acid sequence of SEQ ID NO: 15 or an amino acid sequence at least 95% identical thereto.

18. The split intein system of claim 16, wherein the split point is at position 697-position 1488 relative to SEQ ID NO: 15.

19. The split intein system of claim 16, wherein:(a) the N-terminal portion comprises amino acids 27 - 696 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 697 - 1685 of the type IV collagen alpha chain (SP2);(b) the N-terminal portion comprises amino acids 27 - 878 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 879 - 1685 of the type IV collagen alpha chain (SP5);(c) the N-terminal portion comprises amino acids 27 - 888 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 889 - 1685 of the type IV collagen alpha chain (SP20);(d) the N-terminal portion comprises amino acids 27 - 915 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 916 - 1685 of the type IV collagen alpha chain (SP6);(e) the N-terminal portion comprises amino acids 27 - 941 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 942 - 1685 of the type IV collagen alpha chain (SP7);(f) the N-terminal portion comprises amino acids 27 - 961 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 962 - 1685 of the type IV collagen alpha chain (SP8);(g) the N-terminal portion comprises amino acids 27 - 977 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 978 - 1685 of the type IV collagen alpha chain (SP9);(h) the N-terminal portion comprises amino acids 27 - 995 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 996 - 1685 of the type IV collagen alpha chain (SP10);(i) the N-terminal portion comprises amino acids 27 - 1070 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 1071 - 1685 of the type IV collagen alpha chain (SP11);(j) the N-terminal portion comprises amino acids 27 - 1098 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 1099 - 1685 of the type IV collagen alpha chain (SP27); or(k) the N-terminal portion comprises amino acids 27 - 1135 of the type IV collagen alpha chain, and the C-terminal portion comprises amino acids 1136 - 1685 of the type IV collagen alpha chain (SP12); and wherein amino acid numbering is relative to SEQ ID NO: 15.

20. The split intein system of claim 16, wherein:(a) the N-terminal portion comprises IPG at its C-terminus and the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP2);(b) the N-terminal portion comprises PPG at its C-terminus and the C-terminal portion comprises SPGL (SEQ ID NO: 222) at its N-terminus (SP5);(c) the N-terminal portion comprises AGA at its C-terminus and the C-terminal portion comprises SGFP (SEQ ID NO: 223) at its N-terminus (SP20);(d) the N-terminal portion comprises PGR at its C-terminus and the C-terminal portion comprises SGVP (SEQ ID NO: 224) at its N-terminus (SP6);(e) the N-terminal portion comprises EKG at its C-terminus and the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP7);(f) the N-terminal portion comprises LLG at its C-terminus and the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP8);(g) the N-terminal portion comprises PGV at its C-terminus and the C-terminal portion comprises SGPK (SEQ ID NO: 225) at its N-terminus (SP9);(h) the N-terminal portion comprises PGL at its C-terminus and the C-terminal portion comprises SGQP (SEQ ID NO: 226) at its N-terminus (SP10);(i) the N-terminal portion comprises PGI at its C-terminus and the C-terminal portion comprises SSIG (SEQ ID NO: 227) at its N-terminus (SP11);(j) the N-terminal portion comprises IKG at its C-terminus and the C-terminal portion comprises SVGD (SEQ ID NO: 228) at its N-terminus (SP27); or(k) the N-terminal portion comprises KGI at its C-terminus and the C-terminal portion comprises SGPP (SEQ ID NO: 229) at its N-terminus (SP12).

21. The split intein system of claim 1, wherein the N-terminal portion does not include a collagen signal sequence.

22. The split intein system of claim 1, wherein the first nucleic acid molecule and / or the second nucleic acid molecule is operably linked to a CAG promoter.

23. The split intein system of claim 1, wherein the first nucleic acid molecule and / or the second nucleic acid molecule further comprises a woodchuck hepatitis virus post-transcriptional regulatory element; and / or wherein the first nucleic acid molecule and / or the second nucleic acid molecule is operably linked to the woodchuck hepatitis virus post-transcriptional regulatory element.

24. The split intein system of claim 1, wherein the first nucleic acid molecule and / or the second nucleic acid molecule further encodes a SV40 polyadenylation signal.

25. A pharmaceutical composition comprising an effective amount of the split intein system of claim 1 and a pharmaceutically acceptable carrier, wherein the split intein system comprises the first nucleic acid molecule and the second nucleic acid molecule at a ratio of about 1: 1 to about 0.010: 1.

26. A split intein system comprising:(i) a first adeno-associated viral (AAV) vector comprising a first nucleic acid molecule operably linked to a CAG promoter, wherein the first nucleic acid molecule encodes a fusion protein comprising in N to C terminal order, an N-extein of an extein pair and an N-intein of an Npu DnaE intein pair, wherein the N-extein comprises a collagen signal sequence N-terminal to an N-terminal portion of a type IV collagen alpha 5 (COL4A5) chain, wherein the collagen signal sequence, the N-terminal portion of the COL4A5 chain, and the N-intein are fused directly;(ii) a second AAV vector comprising a second nucleic acid molecule operably linked to a CAG promoter, wherein the second nucleic acid molecule encodes a fusion protein comprising, in N to C terminal order, a signal sequence, the C-intein of the Npu DnaE intein pair, and the C-extein of the extein pair, wherein the C-extein comprises a C-terminal portion of the COL4A5 chain, wherein the signal sequence, the C-intein, and the C-terminal portion of the COL4A5 chain are fused directly; and wherein the N- and C-terminal portions of the COL4A5 chain together comprise the COL4A5 chain sequence separated at a split point; and the N- and C-exteins are spliced together to form the mature COL4A5 chain when the first and second nucleic acid molecules are expressed in mammalian cells.

27. The split intein system of claim 26, wherein:(a) the N-terminal portion comprises amino acids 27 - 696 of the COL4A5 chain, and the C- terminal portion comprises amino acids 697 - 1685 of the COL4A5 chain (SP2);(b) the N-terminal portion comprises amino acids 27 - 878 of the COL4A5 chain, and the C- terminal portion comprises amino acids 879 - 1685 of the COL4A5 chain (SP5);(c) the N-terminal portion comprises amino acids 27 - 888 of the COL4A5 chain, and the C- terminal portion comprises amino acids 889 - 1685 of the COL4A5 chain (SP20);(d) the N-terminal portion comprises amino acids 27 - 915 of the COL4A5 chain, and the C- terminal portion comprises amino acids 916 - 1685 of the COL4A5 chain (SP6);(e) the N-terminal portion comprises amino acids 27 - 941 of the COL4A5 chain, and the C- terminal portion comprises amino acids 942 - 1685 of the COL4A5 chain (SP7);(f) the N-terminal portion comprises amino acids 27 - 961 of the COL4A5 chain, and the C-terminal portion comprises amino acids 962 - 1685 of the COL4A5 chain (SP8);(g) the N-terminal portion comprises amino acids 27 - 977 of the COL4A5 chain, and the C- terminal portion comprises amino acids 978 - 1685 of the COL4A5 chain (SP9);(h) the N-terminal portion comprises amino acids 27 - 995 of the COL4A5 chain, and the C- terminal portion comprises amino acids 996 - 1685 of the COL4A5 chain (SP10);(i) the N-terminal portion comprises amino acids 27 - 1070 of the COL4A5 chain, and the C- terminal portion comprises amino acids 1071 - 1685 of the COL4A5 chain (SP11);(j) the N-terminal portion comprises amino acids 27 - 1098 of the COL4A5 chain, and the C- terminal portion comprises amino acids 1099 - 1685 of the COL4A5 chain (SP27); or(k) the N-terminal portion comprises amino acids 27 - 1135 of the COL4A5 chain, and the C- terminal portion comprises amino acids 1136 - 1685 of the COL4A5 chain (SP12); and wherein amino acid numbering is relative to SEQ ID NO: 15.

28. The split intein system of claim 26, wherein(a) the N-terminal portion comprises IPG at its C-terminus and the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP2);(b) the N-terminal portion comprises PPG at its C-terminus and the C-terminal portion comprises SPGL (SEQ ID NO: 222) at its N-terminus (SP5);(c) the N-terminal portion comprises AGA at its C-terminus and the C-terminal portion comprises SGFP (SEQ ID NO: 223) at its N-terminus (SP20);(d) the N-terminal portion comprises PGR at its C-terminus and the C-terminal portion comprises SGVP (SEQ ID NO: 224) at its N-terminus (SP6);(e) the N-terminal portion comprises EKG at its C-terminus and the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP7);(f) the N-terminal portion comprises LLG at its C-terminus and the C-terminal portion comprises SKGE (SEQ ID NO: 221) at its N-terminus (SP8);(g) the N-terminal portion comprises PGV at its C-terminus and the C-terminal portion comprises SGPK (SEQ ID NO: 225) at its N-terminus (SP9);(h) the N-terminal portion comprises PGL at its C-terminus and the C-terminal portion comprises SGQP (SEQ ID NO: 226) at its N-terminus (SP10);(i) the N-terminal portion comprises PGI at its C-terminus and the C-terminal portion comprises SSIG (SEQ ID NO: 227) at its N-terminus (SP11);(j) the N-terminal portion comprises IKG at its C-terminus and the C-terminal portion comprises SVGD (SEQ ID NO: 228) at its N-terminus (SP27); or(k) the N-terminal portion comprises KGI at its C-terminus and the C-terminal portion comprises SGPP (SEQ ID NO: 229) at its N-terminus (SP12).

29. A pharmaceutical composition comprising an effective amount of the split intein system of claim 26, and a pharmaceutically acceptable carrier.

30. A method of treating Alport syndrome in a subject, comprising administering to the subject an effective amount of the pharmaceutical composition of claim 29, thereby treating the Alport syndrome in the subject.

31. The method of claim 30, comprising delivering the split intein system to the kidney of the subject.

32. The method of claim 31, wherein delivering the split intein system to the kidney of the subject comprises systemic administration.

33. The method of claim 32, wherein the systemic administration comprises systemic injection or systemic infusion.

34. The method of claim 31, wherein delivering the split intein system to the kidney of the subject comprises locally delivering the split intein system to the kidney of the subject.

35. The method of claim 34, wherein the locally delivering the split intein system to the kidney of the subject comprises direct parenchymal injection, renal vein injection, and / or renal artery injection.

36. The method of claim 34, wherein the locally delivering the split intein system to the kidney of the subject comprises direct pelvic injection or retrograde transureteral pelvic injection.

37. The method of claim 30, wherein the split intein system transduces at least one of mesangial cells, glomerular endothelial cells, parietal epithelial cells, podocytes, proximal tubule cells, Loop of Henle cells, distal tubule cells, collecting duct cells, fibroblasts, pericytes, or vascular smooth muscle cells.

38. The method of claim 30, wherein the method: a) improves kidney function; b) delays onset of end stage renal disease; c) delays time to dialysis; d) delays time to renal transplant; and / or e) improves life expectancy; of the subject.

39. The method of claim 30, wherein the Alport syndrome is X-linked Alport syndrome.

40. The method of claim 30, wherein the method further comprises administering an effective amount of angiotensin II receptor blocker (ARB) and / or angiotensin-converting enzyme inhibitor; wherein the ARB comprises candesartan, captopril, eprosartan, irbesartan, losartan, moexipril, olmesartan, telmisartan, and / or valsartan; and wherein the angiotensin-converting enzyme inhibitor comprises benazepril, cilazapril, enalapril, fosinopril, lisinopril, perindopril, ramipril, quinapril, and / or randolapril.

Citation Information

Patent Citations

  • Measurement of biosynthesis and breakdown rates of biological molecules that are inaccessible or not easily accessible to direct sampling, non-invasively, by label incorporation into metabolic derivatives and catabolitic products

    US20030228259A1

  • Genetic indicator and control system and method utilizing split Cas9 / CRISPR domains for transcriptional control in eukaryotic cell lines

    US20170233703A1

  • AAV delivery of nucleobase editors

    US20180127780A1

  • Collagen iv replacement

    US20180207240A1

  • Split inteins and their uses

    US20230116688A1