Large gene vectors and methods for their delivery and use

A dual vector system for expressing the STRC protein through intein-mediated splicing addresses the size limitations of AAV vectors, enabling effective treatment of non-symptomatic hearing loss by restoring protein function in inner ear cells.

JP7846626B2Active Publication Date: 2026-04-15CHILDRENS MEDICAL CENT CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-02-05
Publication Date
2026-04-15

AI Technical Summary

Technical Problem

Existing gene therapy methods, particularly using adeno-associated virus (AAV) vectors, are limited by their capacity to deliver and express large genes, such as the STRC gene, which is crucial for treating non-symptomatic hearing loss types like DFNB16, due to size constraints exceeding 4.5 kB.

Method used

A dual vector system is employed, comprising two vectors that express the N-terminal and C-terminal portions of the STRC protein through intein-mediated splicing, allowing the formation of a full-length protein, overcoming the size limitations of conventional AAV vectors.

Benefits of technology

This approach enables effective delivery and expression of the STRC protein, potentially treating non-symptomatic hearing loss by restoring normal protein function in inner ear cells, thereby improving auditory function.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007846626000068
    Figure 0007846626000068
  • Figure 0007846626000069
    Figure 0007846626000069
  • Figure 0007846626000070
    Figure 0007846626000070
Patent Text Reader

Abstract

The present disclosure provides a dual vector intein-mediated protein trans-splicing system, cells, compositions, and methods for gene therapy using them.In some embodiments, the present disclosure provides methods and compositions for treating autosomal recessive non-syndromic hearing loss DFNB16 by using the dual vector system described herein to deliver the STRC gene encoding STRC protein.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This PCT international application claims the benefits and priority of U.S. Provisional Application No. 62 / 971,555, filed on 7 February 2020, which is incorporated herein by effective reference in its entirety. [Background technology]

[0002] background Hearing loss is one of the most common neurological disorders in developed or industrialized countries, and is the most prevalent sensorineural disorder, accounting for over 466 million cases worldwide. Non-symptomatic hearing loss, or non-symptomatic hereditary hearing loss, is hearing loss that is not associated with other signs or symptoms. There are four types of non-symptomatic hearing loss: DFNA (autosomal dominant), DFNB (autosomal recessive), DFNX (X-linked), and mitochondrial. DFNB16, the 16th described type of autosomal recessive non-symptomatic hearing loss, is a monogenic, non-symptomatic, recessive hearing loss caused by a mutation in the STRC gene, which encodes an extracellular structural protein known as stereocillin. Normal expression of STRC in the inner ear is essential for auditory function. Stereocillin, found on the upper part of modified microvilli at the apical surface of sensory hair cells in the inner ear, is associated with the cilia-like structures known as immobile hairs that protrude from specialized cells in the inner ear. Mutations in STRC cause moderate to severe hearing loss, affecting an estimated 50,000 patients in the United States, and are therefore an attractive candidate for gene therapy. Stereocillins function to maintain the aggregated bundles of microvilli and anneal them to the tectum in the cochlea of ​​the inner ear. Global statistics suggest that DFNB16 accounts for a significant proportion of hereditary hearing loss, particularly in individuals with moderate hearing impairment. Based on data from Partners Laboratory of Molecular Medicine, 19% of patients with hereditary hearing loss tested in Boston had STRC mutations, making it the second most common type of hereditary hearing loss and the most common type affecting the sensory hair cells of the inner ear. Approximately 40 different mutations (mostly recessive) have been identified in the STRC gene, the majority of which result in the synthesis of defective stereocilin or completely prevent its synthesis. The absence of normal STRC protein separates the sensory cilia from the tectum necessary for proper sound-evoked stimuli. Patients with DFNB16 have moderate to severe hearing loss and are typically treated with hearing aids or cochlear implants. However, there are currently no biological treatments for DFNB16 hearing loss.

[0003] AAV offers an attractive vector system for gene therapy treatment of hereditary disorders, considering safety and gene delivery. Recombinant AAV (rAAV) is derived from non-pathogenic replication-deficient viruses and is non-cytotoxic to its host cells. Furthermore, rAAV lacks all viral DNA sequences except terminal inverted repeats (ITRs), exhibiting additional safety characteristics. ITRs are necessary for AAV DNA replication, packaging, chromosomal integration, and proviral rescue. AAV vectors have also proven to be a powerful tool for effective transgene delivery and sustained expression, for example, in inner ear cells. However, many proteins crucial for inner ear function have coding sequences that exceed the loading capacity of AAV vectors (approximately 4.5 kB), for example, STRC (approximately 5.8 kB). Therefore, methods for delivering and expressing proteins encoded by large (e.g., larger than 4 kB) genes, constructs of any of them, and vectors are needed as an effective form of gene therapy. [Overview of the Initiative]

[0004] overview Therefore, an object of this disclosure is to provide gene therapies for large genes (e.g., larger than 4 kB) for treating subjects affected by gene mutations. Gene therapy may enable, for example, the prevention and / or restoration of hearing in children and adults with DFNB16 hearing loss.

[0005] Another objective is to provide a method for delivering large gene sequences (e.g., larger than 4kB; STRC) that overcomes the size limitations of vectors, as well as vectors and constructs for delivering large gene sequences.

[0006] One aspect is providing a vector system (e.g., a dual vector system) for expressing a protein of interest in a cell, where the dual vector system is (a) In the direction from 5' to 3', - A signal sequence located at the 5' end of the subcoding sequence that codes for the amino-terminus (N-terminus) portion of the protein of interest (e.g., SEQ ID NO:9;SEQ ID NO:11); - A partial coding sequence that encodes the N-terminal portion of the protein of interest (e.g., N-STRC; SEQ ID NO: 15; SEQ ID NO: 16); - A sequence that codes for a splice donor sequence adjacent to the downstream of the subcode sequence (for example, the N-terminal fragment of an intein (N-intine), also known as split intein-N; SEQ ID NO:13 coding SEQ ID NO:14). A first vector containing a first nucleotide sequence (e.g., SEQ ID NO:5;SEQ ID NO:7), and (b) In the direction from 5' to 3', - A signal sequence located at the 5' end of the subcoding sequence that encodes the carboxyl terminus (C-terminus) of the protein of interest (e.g., SEQ ID NO:9;SEQ ID NO:11); - A sequence encoding a splice acceptor sequence (for example, SEQ ID NO: 21 encoding the C-terminal fragment of intein (C-intine), also known as split intein-C; SEQ ID NO: 22), wherein the splice acceptor sequence is sandwiched between a signal sequence and a subcoding sequence encoding the C-terminal portion of the protein of interest; - Partial coding sequences that encode the C-terminal portion of the protein of interest (e.g., C-STRC; SEQ ID NO:23; SEQ ID NO:24). A second vector containing a second nucleotide sequence (e.g., SEQ ID NO:17;SEQ ID NO:19) Includes.

[0007] Another aspect is a dual-vector system for expressing a protein of interest in a cell, (a) In the direction from 5' to 3', - 5'-terminal inverted repeat (5'ITR) sequence; - Promoter sequence; - A signal sequence that is functionally linked to the promoter and under the regulation of the promoter; - A partial coding sequence encoding the amino-terminal (N-terminal) portion of the protein of interest (e.g., N-STRC), which is functionally linked to the promoter and under the regulation of the promoter; - A sequence encoding the amino-terminal fragment (N-intein) of the intein, which is functionally linked to the promoter and under the regulation of the promoter; - Polyadenylation (polyA) signal sequence; - 3'-terminal inverted repeat (3'ITR) sequence A first vector comprising a first nucleotide sequence comprising the above, and (b) In the 5' to 3' direction, - 5'-terminal inverted repeat (5'ITR) sequence; - Promoter sequence; - A signal sequence that is functionally linked to the promoter and under the regulation of the promoter; - A sequence encoding the carboxy-terminal fragment (C-intein) of the intein, which is functionally linked to the promoter and under the regulation of the promoter; - A partial coding sequence encoding the carboxy-terminal (C-terminal) portion of the protein of interest (e.g., C-STRC), which is functionally linked to the promoter and under the regulation of the promoter; - Polyadenylation (polyA) signal sequence; - 3'-terminal inverted repeat (3'ITR) sequence A second vector comprising a second nucleotide sequence comprising the above A dual vector system comprising the above is provided.

[0008] In another aspect of the dual vector system, the first vector and the second vector are in a cell, (a) in the direction from the N-terminus to the C-terminus, - a signal peptide sequence linked to the N-terminal portion of the protein sequence of interest (e.g., STRC), wherein the protein sequence of interest is fused to the N-intein protein sequence at its C-terminus, the signal peptide sequence comprising a first protein sequence, and (b) in the direction from the N-terminus to the C-terminus, - a signal peptide sequence linked to the C-intein protein sequence, wherein the C-intein protein sequence is fused to the N-terminus of the C-terminal portion of the protein sequence of interest (e.g., STRC), the signal peptide sequence comprising a second protein sequence are each expressed.

[0009] In yet another aspect, the N-terminal portion (e.g., N-STRC) and the C-terminal portion (e.g., C-STRC) of the protein of interest form a full-length protein of interest (e.g., STRC). In some aspects, the signal peptide sequence of the first protein sequence and the signal peptide sequence of the second protein sequence are the same or different, or the signal peptide sequences of the first protein sequence and the signal peptide sequence of the second protein sequence are configured to transport the first protein sequence and the second protein sequence to the same cellular compartment. Further aspects may relate to signal sequences encoding a signal peptide sequence having a nucleic acid sequence having at least 80% identity with SEQ ID NO:9 or SEQ ID NO:11 and an amino acid sequence having at least 80% identity with SEQ ID NO:10 or SEQ ID NO:12. Other aspects provide vectors (e.g., the first and second vectors) which may be viral vectors, where the viral vectors may be adeno-associated virus (AAV) vectors or lentiviruses. One aspect may relate to viral vectors having the same or different serotypes. Another aspect of the dual vector system provides intein-mediated transsplicing of the protein of interest, where the N-terminal portion (e.g., N-STRC) and the C-terminal portion (e.g., C-STRC) of the protein of interest can form a full-length protein of interest (e.g., STRC) through a peptide bond, where the protein of interest may be the STRC protein encoded by the STRC gene.The nucleotide sequence encoding the N-terminal portion of the protein of interest (e.g., N-STRC; human SEQ ID NO:15 or mouse SEQ ID NO:16) is at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) or less than 54% (e.g., 53%, 52%, 51%, 50%, 45%, 43%, 41%) of the N-terminal portion of the full-length protein of interest (e.g., STRC; SEQ ID NO:25 or SEQ ID NO:26) and / or less than 54% identity with the N-terminal portion of the full-length protein of interest (e.g., STRC; SEQ ID NO:25 or SEQ ID NO:26), and / or less than 54% identity with the N-terminal portion of the full-length protein of interest (e.g., STRC; SEQ ID NO:25 or SEQ ID NO:26). The nucleic acid sequence comprises a nucleic acid sequence having at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with a nucleotide sequence of interest (e.g., STRC;SEQ ID NO:5 or SEQ ID NO:7) that encodes an amino acid sequence of interest (e.g., STRC;SEQ ID NO:6;SEQ ID NO:8;SEQ ID NO:15, or SEQ ID NO:16) having a length of less than 54% of the N-terminal portion of NO:26). Another aspect is to provide the N-terminal portion of the protein of interest (e.g., STRC) which contains an amino acid sequence having 41% or more (e.g., 42%, 43%, 44%, 45%, 50%, 51%, 52%, 53%) of the N-terminal portion of the full-length protein of interest (e.g., SEQ ID NO:25 or SEQ ID NO:26), and / or 41% or more identity with the N-terminal portion of the full-length protein of interest (e.g., SEQ ID NO:25 or SEQ ID NO:26), and / or a length of 41% or more of the N-terminal portion of the full-length protein of interest (e.g., SEQ ID NO:25 or SEQ ID NO:26).

[0010] A further aspect is provided which includes nucleotide sequences comprising signal sequences having nucleic acid sequences having at least 80% (e.g., 85%, 90%, 95%, 97%, 99%, 100%) identity with a desired signal sequence (e.g., SEQ ID NO:9;SEQ ID NO:11), which encodes a signal peptide sequence having an amino acid sequence having at least 80% (e.g., 85%, 90%, 95%, 97%, 99%, 100%) identity with a desired signal peptide sequence (e.g., SEQ ID NO:10;SEQ ID NO:12).

[0011] Another aspect may be to provide a desired N-intane sequence comprising a nucleic acid sequence having at least 80% (e.g., 85%, 90%, 95%, 97%, 99%, 100%) identity with a desired N-intane amino acid sequence (e.g., SEQ ID NO: 14) and encoding a desired N-intane amino acid sequence (e.g., SEQ ID NO: 13). A further aspect may relate to a desired C-intane sequence comprising a nucleic acid sequence having at least 80% (e.g., 85%, 90%, 95%, 97%, 99%, 100%) identity with a desired C-intane amino acid sequence (e.g., SEQ ID NO: 22), and encoding a desired C-intane amino acid sequence having at least 80% (e.g., 85%, 90%, 95%, 97%, 99%, 100%) identity with a desired C-intane nucleotide sequence (e.g., SEQ ID NO: 21).

[0012] In yet another aspect of the dual vector system of this disclosure, the C-terminal portion of the protein of interest (e.g., STRC) is at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) or 46% or more of the C-terminal portion of the full-length protein of interest (e.g., STRC;SEQ ID NO:25;SEQ ID NO:26) and / or has at least 46% identity with the C-terminal portion of the full-length protein of interest (e.g., STRC;SEQ ID NO:25;SEQ ID NO:26) and / or has at least 46% length of the C-terminal portion of the amino acid sequence of interest (e.g., STRC;SEQ ID NO:18;SEQ ID The present invention may include nucleic acid sequences having at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with the nucleotide sequence (e.g., STRC;SEQ ID NO:17;SEQ ID NO:19) encoding NO:20;SEQ ID NO:23;SEQ ID NO:24). Another aspect may provide a C-terminal portion of the protein of interest (e.g., STRC) having less than 60% identity with the C-terminal portion of the full-length protein of interest (e.g., STRC;SEQ ID NO:25;SEQ ID NO:26), and / or an amino acid sequence having less than 60% of the length of the C-terminal portion of the full-length protein of interest (e.g., STRC;SEQ ID NO:25;SEQ ID NO:26).

[0013] A further aspect may be provided to provide a vector system for expressing the coding sequence of the STRC gene in host cells, wherein the coding sequence comprises at least one vector containing, for example, the STRC nucleotide coding sequence of human STRC:SEQ ID NO:1 or SEQ ID NO:30 or mouse STRC:SEQ ID NO:32, and the STRC nucleotide coding sequence codes for, for example, the STRC protein of SEQ ID NO:2 or SEQ ID NO:25 or SEQ ID NO:4 or SEQ ID NO:26. Another aspect may relate to a nucleotide sequence that codes for a desired full-length protein, wherein the nucleotide sequence comprises, for example, human STRC:SEQ ID NO:1 or SEQ ID NO:33 or mouse STRC:SEQ ID NO:3 or SEQ ID NO:39 that codes for the desired protein, for example, human STRC:SEQ ID NO:2 or SEQ ID NO:25 or mouse STRC:SEQ ID NO:4 or SEQ ID NO:26.

[0014] Aspect of a vector system comprising a dual vector system for expressing the coding sequence of a STRC gene in host cells described herein, wherein the dual vector system provides a first vector comprising a first nucleotide sequence comprising a desired nucleic acid sequence having at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with a desired nucleic acid sequence of interest (e.g., SEQ ID NO: 5; SEQ ID NO: 7). In one aspect, the first vector (e.g., plasmid, transplicing plasmid, viral vector, adenovirus, AAV, AAV genome) includes a first nucleotide sequence (e.g., SEQ ID NO:5 encoding SEQ ID NO:6; SEQ ID NO:7 encoding SEQ ID NO:8) that includes a signal sequence (e.g., SEQ ID NO:9 encoding SEQ ID NO:10; or SEQ ID NO:11 encoding SEQ ID NO:12) at the 5' end of the subcoding sequence in the 5' to 3' direction (the subcoding sequence may be adjacent to a downstream sequence encoding a splice donor sequence (e.g., an N-terminal intein (also known as split intein-N, or N-intane); SEQ ID NO:13 encoding SEQ ID NO:14)); a subcoding sequence encoding the amino-terminal (N-terminal) portion of the protein of interest (e.g., STRC; SEQ ID NO:15; SEQ ID NO:16). Another aspect involves a first nucleotide sequence that includes the N-terminal portion of the protein of interest, also containing a signal sequence and a sequence encoding the desired N-intane protein, as well as terminal inversion repeats (ITRs), a promoter, and a polyadenylation (Poly-A) sequence.The first nucleotide sequence can encode an amino acid sequence of interest that has at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with a desired amino acid sequence (e.g., SEQ ID NO: 5; SEQ ID NO: 16) or a full-length amino acid sequence of interest (e.g., SEQ ID NO: 25; SEQ ID NO: 26).

[0015] Another aspect of the dual vector system of this disclosure also provides a second nucleotide sequence comprising the remainder of the desired nucleic acid sequence having at least 5% identity (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) with the desired nucleic acid sequence of interest (e.g., SEQ ID NO:17; SEQ ID NO:19). In one aspect, a second vector, such as a plasmid, transpricing plasmid, viral vector, adenovirus, AAV, or AAV genome, may contain a second nucleotide sequence (e.g., SEQ ID NO:17; SEQ ID NO:19; SEQ ID NO:20) in the 5' to 3' direction, where a signal sequence (e.g., SEQ ID NO:9 encoding SEQ ID NO:10; or SEQ ID NO:11 encoding SEQ ID NO:12) may be upstream of a splice acceptor sequence (e.g., C-terminal intein (C-intane); SEQ ID NO:21 encoding SEQ ID NO:22), which may be directly adjacent to a downstream subcoding sequence encoding the remainder of the full-length coding sequence of the protein of interest, i.e., the C-terminal portion of the protein of interest (e.g., STRC; SEQ ID NO:23; SEQ ID NO:24). Another aspect involves a second nucleotide sequence that includes the C-terminal portion of the protein of interest, the sequence encoding the signal sequence, and the sequence encoding the desired C-intane protein, and also includes terminal inversion repeats (ITRs), promoters, and polyadenylation (Poly-A) sequences. In some aspects, the second nucleotide sequence may also include linker sequences and myc-tag sequences.The second nucleotide sequence may encode an amino acid sequence that has at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with the desired amino acid sequence (e.g., SEQ ID NO:18;SEQ ID NO:20;SEQ ID NO:23;SEQ ID NO:24) or the full-length amino acid sequence of interest (e.g., SEQ ID NO:25;SEQ ID NO:26).

[0016] Another aspect may provide cells or host cells containing a vector system described herein (e.g., a dual vector system) for delivering a desired gene or a desired protein (e.g., STRC protein).

[0017] Further aspects may relate to a pharmaceutical composition comprising a vector system of the present disclosure (e.g., a dual vector system) for delivering a desired gene or a desired protein (e.g., an STRC protein) and a pharmaceutically acceptable medium (e.g., a diluent, an excipient).

[0018] In another aspect, a method for treating a disease or condition in a subject suffering from a gene mutation of a disease or condition, comprising administering an effective amount of the vector system of this disclosure (e.g., a dual vector system) to a subject in need thereof, wherein a method for treating a disease or condition in a subject suffering from a disease or condition caused by a gene mutation of the gene is provided, comprising delivering a desired wild-type or corrected gene or a desired wild-type or corrected protein (e.g., STRC) to a subject suffering from a disease or condition caused by a gene mutation of the said gene. In yet another aspect, a method for treating a disease or condition in a subject suffering from autosomal recessive deafness is provided, comprising administering an effective amount of the dual vector system described herein, which delivers a desired wild-type or corrected gene or a desired wild-type or corrected protein (e.g., STRC), to a subject in need thereof. A method for treating autosomal recessive deafness in a subject, comprising administering a cell or host cell or a pharmaceutical composition (including a pharmaceutically acceptable medium (e.g., diluent, excipient)) containing the dual vector system described herein for delivering a desired gene or a desired protein (e.g., STRC protein) to a subject requiring it. Another aspect may provide autosomal recessive deafness DFNB16.

[0019] In one aspect, the method of the present disclosure may involve contacting a cell of interest with a composition comprising the vector system of the present disclosure (e.g., a dual vector system) for delivering a desired gene or a desired protein thereof (e.g., an STRC protein) and a pharmaceutically acceptable medium (e.g., a diluent, an excipient), wherein the contact results in the delivery of a first nucleotide sequence and a second nucleotide sequence to the cell, and the cell may express the N-terminal portion and the C-terminal portion of the desired protein, which are linked by peptide bonds to form the desired full-length protein. In another aspect, the present disclosure provides a method for treating and / or preventing a pathology or disease characterized by hearing loss, comprising administering an effective amount of the dual vector system described herein for delivering a desired gene or a desired protein thereof (e.g., an STRC protein), or a cell containing the dual vector system described herein, or a pharmaceutically acceptable medium (e.g., a diluent, an excipient) to a subject in need thereof. In some aspects, the cells may be inner ear cells, inner hair cells, or outer hair cells of the ear, wherein the cells, or the method of administering them, may be in vivo, ex vivo, and / or in vitro. Further aspects of this disclosure include any of the methods described herein resulting in improvement or restoration of auditory function in a subject.

[0020] Other features and advantages of this disclosure will become apparent from the detailed description and the claims. [Invention 1001] A dual vector system for expressing a protein of interest in cells, (a) In the direction from 5' to 3', - A signal sequence located at the 5' end of the partial coding sequence that encodes the amino-terminus (N-terminus) portion of the protein of interest; - A partial coding sequence that encodes the N-terminal portion of the protein of interest; - A sequence that encodes the aminocarboxy terminal fragment (N-intene) of intein, located downstream and adjacent to the subcoding sequence. A first vector containing a first nucleotide sequence, and (b) In the direction from 5' to 3', - A signal sequence located at the 5' end of the partial coding sequence that encodes the carboxyl terminus (C-terminus) of the protein of interest; - A sequence encoding the carboxyl-terminal fragment (C-intane) of intein, wherein the C-intane is sandwiched between a signal sequence and a partial coding sequence encoding the C-terminal portion of the protein of interest; - Partial coding sequence that encodes the C-terminal portion of the protein of interest A second vector containing a second nucleotide sequence including A dual-vector system, including one. [Invention 1002] It is a dual vector system, (a) In the direction from 5' to 3', - 5'-terminal inverted repeat (5'ITR) sequence; - Promoter sequence; - A signal sequence that is functionally linked to and under the control of a promoter; - A partial coding sequence that encodes the amino-terminal (N-terminal) portion of a protein of interest, is functionally linked to a promoter, and is under the control of the promoter; - A sequence encoding split-intane-N, which is functionally linked to a promoter and under the control of the promoter; - Polyadenylated (poly-A) signal sequence; - 3'-terminal inverted repeat (3'ITR) sequence A first vector containing a first nucleotide sequence, and (b) In the direction from 5' to 3', - 5'-terminal inverted repeat (5'ITR) sequence; - Promoter sequence; - A signal sequence that is functionally linked to and under the control of a promoter; - A sequence encoding split-intane-C, which is functionally linked to a promoter and under the control of the promoter; - A partial coding sequence that encodes the carboxyl-terminal (C-terminal) portion of a protein of interest, is functionally linked to a promoter, and is under the control of the promoter; - Polyadenylated (poly-A) signal sequence; - 3'-terminal inverted repeat (3'ITR) sequence A second vector containing a second nucleotide sequence including A dual-vector system, including one. [Invention 1003] The first vector and the second vector, in the cell, (a) From the N-terminus towards the C-terminus, - A signal peptide sequence that is ligated to the N-terminal portion of a protein sequence of interest, and the protein sequence of interest is fused to an N-intane protein sequence at its C-terminus. The first protein sequence includes, and (b) From the N-terminus towards the C-terminus, - A signal peptide sequence that is linked to a C-intane protein sequence, wherein the C-intane protein sequence is fused to the N-terminus of the C-terminal portion of the protein sequence of interest. The second protein sequence includes A dual vector system of the present invention 1001 or 1002, which expresses each of the following: [Invention 1004] A dual vector system according to any one of the invention 1001 to 1003, wherein the N-terminal portion of the protein of interest and the C-terminal portion of the protein of interest are configured to form a full-length protein of interest. [Invention 1005] A dual vector system according to the present invention 1003, wherein the signal peptide sequence of the first protein sequence and the signal peptide sequence of the second protein sequence are the same. [Invention 1006] The dual vector system of the present invention 1003, wherein the signal peptide sequence of a first protein sequence and the signal peptide sequence of a second protein sequence are configured to transport the first protein sequence and the second protein sequence to the same cellular compartment of a cell. [Invention 1007] The dual vector system of the present invention 1003, wherein the signal peptide sequence of the first protein sequence and the signal peptide sequence of the second protein sequence are different, and each signal peptide sequence directs each protein sequence to the same cellular compartment of the cell. [Invention 1008] A dual vector system according to any one of invention 1001 to 1003, wherein the first vector and the second vector are each viral vectors. [Invention 1009] A dual-vector system according to the present invention 1008, wherein the viral vector is an adeno-associated virus (AAV) vector or a lentivirus. [Invention 1010] A dual vector system according to any of invention 1001 to 1004, wherein the protein of interest is an STRC protein. [Invention 1011] A dual vector system according to any one of the invention 1001 to 1004, wherein the signal sequence contains a nucleic acid sequence having at least 80% identity with SEQ ID NO:9 or SEQ ID NO:11, or the signal sequence encodes a signal peptide sequence having an amino acid sequence having at least 80% identity with SEQ ID NO:10 or SEQ ID NO:12. [Invention 1012] A dual vector system according to any one of the invention 1001 to 1004, wherein the N-terminal portion of the protein of interest contains a nucleic acid sequence having at least 70% identity with the nucleic acid sequence encoding SEQ ID NO:15 or SEQ ID NO:16, or encodes an amino acid sequence having at least 70% identity with SEQ ID NO:15 or SEQ ID NO:16. [Invention 1013] A dual vector system according to any one of the invention 1001 to 1004, wherein the N-intane sequence comprises a nucleic acid sequence having at least 80% identity with SEQ ID NO:13, or encodes an amino acid sequence having at least 80% identity with SEQ ID NO:14. [Invention 1014] A dual vector system according to any one of the invention 1001 to 1004, wherein the C-terminal portion of the protein of interest contains a nucleic acid sequence having at least 70% identity with the nucleic acid sequence encoding SEQ ID NO:23 or SEQ ID NO:24, or encodes an amino acid sequence having at least 70% identity with SEQ ID NO:23 or SEQ ID NO:24. [Invention 1015] A dual vector system according to any of invention 1001 to 1004, wherein the C-intane sequence contains a nucleic acid sequence having at least 80% identity with SEQ ID NO:21 or SEQ ID NO:46, or encodes an amino acid sequence having at least 80% identity with SEQ ID NO:22 or SEQ ID NO:49. [Invention 1016] A dual vector system according to any of invention 1001 to 1004, wherein the first nucleotide sequence comprises a nucleic acid sequence having at least 70% identity with SEQ ID NO:5 or SEQ ID NO:7, or encodes an amino acid sequence having at least 70% identity with SEQ ID NO:6 or SEQ ID NO:8. [Invention 1017] A dual vector system according to any of invention 1001 to 1004, wherein the second nucleotide sequence comprises a nucleic acid sequence having at least 70% identity with SEQ ID NO:17 or SEQ ID NO:19, or encodes an amino acid sequence having at least 70% identity with SEQ ID NO:18 or SEQ ID NO:20. [Invention 1018] A vector system for expressing the coding sequence of the STRC gene in a host cell, comprising at least one vector whose coding sequence is the STRC gene with SEQ ID NO:1, SEQ ID NO:3, SEQ ID NO:5, SEQ ID NO:7, SEQ ID NO:17, SEQ ID NO:19, SEQ ID NO:29, SEQ ID NO:31, SEQ ID NO:33, or SEQ ID NO:38, the mRNA sequence with SEQ ID NO:30 or SEQ ID NO:32, or a fragment thereof. [Invention 1019] The vector system of the present invention 1018, wherein the STRC gene encodes STRC proteins of SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:18, SEQ ID NO:20, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:36, SEQ ID NO:39, or combinations thereof. [Invention 1020] A vector system according to Invention 1018 or Invention 1019, comprising a dual vector system for expressing the coding sequence of the STRC gene in host cells, wherein the coding sequence comprises a 5' terminal fragment and a 3' terminal fragment, and the dual vector system is (a) In the direction from 5' to 3', - 5'-terminal inverted repeat (5'ITR) sequence; - Promoter sequence; - A signal sequence that is functionally linked to and under the control of a promoter; - A 5' terminal fragment of the STRC gene coding sequence, which is functionally ligated to a promoter and under the regulation of the promoter; - A sequence encoding the amino-terminal fragment (N-intene) of intein, which is functionally linked to a promoter and under the control of the promoter; - Polyadenylated (poly-A) signal sequences; and - 3'-terminal inverted repeat (3'ITR) sequence A first vector containing a first nucleotide sequence including, (b) In the direction from 5' to 3', - 5'-terminal inverted repeat (5'ITR) sequence; - Promoter sequence; - A signal sequence that is functionally linked to and under the control of a promoter; - A sequence encoding the carboxy-terminal fragment (C-intei) of intein, which is functionally linked to a promoter and under the control of the promoter; - A 3' terminal fragment of the STRC gene coding sequence, which is functionally ligated to a promoter and under the regulation of the promoter; - Polyadenylated (poly-A) signal sequences; and - 3'-terminal inverted repeat (3'ITR) sequence A second vector containing a second nucleotide sequence including A vector system, including [Invention 1021] The first vector and the second vector, in the cell, (a) From the N-terminus towards the C-terminus, - A signal peptide sequence that is ligated to the N-terminal portion of an STRC protein sequence, and in which the STRC protein sequence is fused to an N-intane protein sequence at its C-terminus. The first protein sequence includes, and (b) From the N-terminus towards the C-terminus, - A signal peptide sequence that is linked to a C-intane protein sequence, wherein the C-intane protein sequence is fused to the N-terminus of the C-terminal portion of the STRC protein sequence. The second protein sequence includes A dual vector system of the present invention 1020, which expresses each of the following. [Invention 1022] A dual vector system according to any one of the invention 1020 to 1021, wherein the N-terminal portion of the STRC protein and the C-terminal portion of the STRC protein are configured to form a full-length STRC protein. [Invention 1023] A dual vector system according to the present invention 1021, wherein the signal peptide sequence of the first protein sequence and the signal peptide sequence of the second protein sequence are the same. [Invention 1024] The dual vector system of the present invention 1021, wherein the signal peptide sequence of a first protein sequence and the signal peptide sequence of a second protein sequence are configured to transport the first protein sequence and the second protein sequence into the same cellular compartment. [Invention 1025] The dual vector system of the present invention 1021, wherein the signal peptide sequence of the first protein sequence and the signal peptide sequence of the second protein sequence are different, and each signal peptide sequence directs each protein sequence to the same cellular compartment. [Invention 1026] A dual-vector system according to any of invention 1020 to 1021, wherein the viral vector is an adeno-associated virus (AAV) vector. [Invention 1027] A dual vector system according to any of invention 1020 to 1021, wherein the signal sequence comprises a nucleic acid sequence having at least 80% identity with SEQ ID NO:9 or SEQ ID NO:11, or encodes an amino acid sequence having at least 80% identity with SEQ ID NO:10 or SEQ ID NO:12. [Invention 1028] A dual vector system according to any of invention 1020 to 1021, wherein the N-terminal portion of the STRC protein contains a nucleic acid sequence having at least 70% identity with the nucleic acid sequence encoding SEQ ID NO:15 or SEQ ID NO:16, or encodes an amino acid sequence having at least 70% identity with SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO:15, or SEQ ID NO:16. [Invention 1029] A dual vector system according to any one of the invention 1020 to 1021, wherein the N-terminal portion of the STRC protein contains less than 54% of the N-terminal portion of the full-length STRC protein. [Invention 1030] A dual vector system according to any one of the invention 1020 to 1021, wherein the N-intane sequence comprises a nucleic acid sequence having at least 80% identity with SEQ ID NO:13, or encodes an amino acid sequence having at least 80% identity with SEQ ID NO:14. [Invention 1031] A dual vector system according to any of invention 1020 to 1021, wherein the C-terminal portion of the STRC protein contains a nucleic acid sequence having at least 70% identity with the nucleic acid sequence encoding SEQ ID NO:18, SEQ ID NO:20, SEQ ID NO:23, or SEQ ID NO:24, or encodes an amino acid sequence having at least 70% identity with SEQ ID NO:18, SEQ ID NO:20, SEQ ID NO:23, or SEQ ID NO:24. [Invention 1032] A dual vector system according to any one of the invention 1020 to 1021, wherein the C-terminal portion of the STRC protein contains 46% or more of the C-terminal portion of the full-length STRC protein. [Invention 1033] A dual vector system according to any of invention 1020 to 1021, wherein the C-intane sequence contains a nucleic acid sequence that is at least 80% identical to SEQ ID NO:21, or encodes an amino acid sequence that is at least 80% identical to SEQ ID NO:22. [Invention 1034] A dual vector system according to any of invention 1020 to 1021, wherein the first nucleotide sequence comprises a nucleic acid sequence having at least 70% identity with SEQ ID NO:5 or SEQ ID NO:7, or encodes an amino acid sequence having at least 70% identity with SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO:15, or SEQ ID NO:16. [Invention 1035] A dual vector system according to any of invention 1020 to 1021, wherein the second nucleotide sequence contains a nucleic acid sequence having at least 70% identity with SEQ ID NO:17 or SEQ ID NO:19, or encodes an amino acid sequence having at least 70% identity with SEQ ID NO:18, SEQ ID NO:20, SEQ ID NO:23, or SEQ ID NO:24. [Invention 1036] A cell containing any of the vector systems described in invention 1001 to 1035. [Invention 1037] A pharmaceutical composition comprising a vector system according to any of invention 1001 to 1035 and a pharmaceutically acceptable medium. [Invention 1038] A method for treating autosomal recessive deafness in a subject, comprising the step of administering an effective amount of any dual vector system of the present invention 1001 to 1035 to a subject in need thereof. [Invention 1039] A method for treating autosomal recessive hearing loss in a subject, comprising the step of administering cells of the present invention 1036 to a subject in need thereof. [Invention 1040] A method for treating autosomal recessive hearing loss in a subject, comprising the step of administering the pharmaceutical composition of the present invention 1037 to a subject in need thereof. [Invention 1041] Any method of the present invention 1038 to 1040, wherein the autosomal recessive deafness is DFNB16. [Invention 1042] A method comprising the step of bringing at least one target cell into contact with the pharmaceutical composition of the present invention 1037, Contact delivers a vector system containing a first nucleotide sequence and a second nucleotide sequence to at least one target cell. A method comprising expressing the N-terminal and C-terminal portions of a protein, which are linked by peptide bonds to form a full-length protein, in contact with at least one cell. [Invention 1043] A method for treating and / or preventing a pathology or disease characterized by hearing loss, comprising the step of administering an effective amount of any vector system of Invention 1001 to 1035, at least one cell of Invention 1036, or a pharmaceutical composition of Invention 1037 to a subject in need thereof. [Invention 1044] A method according to any one of the present invention 1042 to 1043, wherein at least one cell is an inner ear cell. [Invention 1045] A method according to any one of the present invention 1042 to 1043, wherein at least one cell is an inner hair cell or an outer hair cell. [Invention 1046] A method according to any one of the present invention 1042 to 1045, wherein at least one cell is in vivo or in vitro. [Invention 1047] A method according to any of items 1042 to 1046 of the present invention for improving or restoring auditory function in a subject. [Brief explanation of the drawing]

[0021] [Figure 1]Figure 1 shows a schematic diagram of a construct for a single vector system for the expression of a desired full-length protein (e.g., STRC), which has an AAV2 terminal inverted repeat (ITR), a promoter, and a polyadenylated (PolyA) sequence. [Figure 2A] Figures 2A-2C show the human STRC protein, signal peptide sequence, linker sequence, sequence encoding the Myc tag, and nucleotide sequences encoding the start and stop codons (SEQ ID NO: 33). [Figure 2B] Refer to the explanation in Figure 2A. [Figure 2C] Refer to the explanation in Figure 2A. [Figure 3] Figure 3 shows the signal peptide sequence, human STRC protein sequence, linker sequence, and amino acid sequence containing the myc tag (SEQ ID NO: 36) encoded by the nucleotide sequences presented in Figures 2A to 2C. [Figure 4A] Figures 4A–4D show the mouse STRC protein, signal peptide sequence, linker sequence, sequence encoding the Myc tag, and nucleotide sequences encoding the start and stop codons (SEQ ID NO: 38). [Figure 4B] See the explanation in Figure 4A. [Figure 4C] See the explanation in Figure 4A. [Figure 4D] See the explanation in Figure 4A. [Figure 5] Figure 5 shows the signal peptide sequence, mouse STRC protein sequence, linker sequence, and amino acid sequence (SEQ ID NO: 39) containing the Myc tag, which are encoded by the nucleotide sequences presented in Figures 4A and 4D. [Figure 6] Figure 6 shows a schematic diagram of dual AAV intein-mediated stereocylinprotein transsplicing using an AAV2 terminal inverted repeat, promoter, and polyadenylated sequence. The intein fragment mediates protein recombination, cleaving itself and joining the remaining STRC fragment (extine) via a peptide bond. [Figure 7A] Figures 7A and 7B show the nucleotide sequence (SEQ ID NO: 5) (CFS-N-Strc-N-Int, construct 2, N-part) encoding the N-terminal portion, signal peptide sequence, and splice donor sequence (e.g., N-intane) of the human STRC protein. The nucleotide sequence contains the signal sequence, 5'Strc (the 5' fragment of the wild-type Strc coding sequence), and the N-intane sequence (encoding the N-terminal fragment of the intein protein). [Figure 7B] Refer to the explanation in Figure 7A. [Figure 8A] Figures 8A and 8B show the N-terminal portion, signal peptide sequence, and nucleotide sequence encoding the N-intane (SEQ ID NO: 7) of the mouse STRC protein (e.g., CFS-N-Strc-N-Int, construct 2, N-part). The nucleotide sequence contains the signal sequence, 5'Strc (the 5' fragment of the wild-type Strc coding sequence), and the N-intane sequence (encoding the N-terminal fragment of the intein protein). [Figure 8B] See the explanation in Figure 8A. [Figure 9] Figure 9 shows a schematic diagram of a construct used in the double AAV intein-mediated protein transsplicing described herein, which includes the signal peptide sequence, the N-terminal portion of the STRC protein, and the nucleotide sequences of Figures 8A and 8B that encode the N-intane. [Figure 10] Figure 10 shows the N-terminal portion, signal peptide sequence, and amino acid sequence containing N-intei (SEQ ID NO: 6) of the human STRC protein, encoded by the nucleotide sequences presented in Figures 7A and 7B. [Figure 11] Figure 11 shows the N-terminal portion, signal peptide sequence, and amino acid sequence containing N-intei (SEQ ID NO: 8) of the mouse STRC protein, which are encoded by the nucleotide sequences presented in Figures 8A and 8B. [Figure 12A]Figures 12A and 12B show the C-terminal portion, signal peptide sequence, and nucleotide sequence encoding the C-intane (SEQ ID NO: 17) of the human STRC protein (e.g., CFS-C-Strc-C-Int, construct 2, C-portion). The nucleotide sequence contains the signal sequence, the C-intane sequence (encoding the C-terminal fragment of the intein protein), 3'Strc (the 3' fragment of the wild-type Strc coding sequence), the linker sequence, and the myc tag sequence. [Figure 12B] Refer to the explanation in Figure 12A. [Figure 13A] Figures 13A and 13B show the C-terminal portion, signal peptide sequence, and nucleotide sequence encoding the C-intane (SEQ ID NO: 19) of the mouse STRC protein (e.g., CFS-C-Strc-C-Int, construct 2, C-portion). The nucleotide sequence contains the signal sequence, the C-intane sequence (encoding the C-terminal fragment of the intein protein), 3'Strc (the 3' fragment of the wild-type Strc coding sequence), the linker sequence, and the myc tag sequence. [Figure 13B] Refer to the explanation in Figure 13A. [Figure 14] Figure 14 shows a schematic diagram of a construct containing the signal peptide sequence, the C-intane, and the nucleotide sequences of Figures 13A and 13B that encode the C-terminal portion of the STRC protein, used in the double AAV intein-mediated protein transsplicing described herein. [Figure 15] Figure 15 shows the amino acid sequence (SEQ ID NO: 18) containing the signal peptide sequence, C-intane, the C-terminal portion of the human STRC protein, the linker sequence, and the myc tag, which are encoded by the nucleotide sequences presented in Figures 12A and 12B. [Figure 16] Figure 16 shows the amino acid sequence (SEQ ID NO: 20) containing the signal peptide sequence, C-intane, C-terminal portion of the mouse STRC protein, linker sequence, and myc tag, which are encoded by the nucleotide sequences presented in Figures 13A and 13B. [Figure 17] Figure 17 shows the predicted structure of a stereocillin (STRC) protein containing specified CFS cleavage sites with cysteine ​​(C;Cys), phenylalanine (F,Phe), and serine (S;Ser) as shown in Figures 11 and 16, produced by the sequences of Figures 8A, 8B, 13A, and 13B, as illustrated by the constructs in Figures 9 and 14. [Figure 18A] Figures 18A and 18B show the N-terminal portion of the STRC protein, the signal peptide sequence, and the nucleotide sequence encoding the N-intane (SEQ ID NO: 51) (e.g., CFS-N-Strc-N-Int, construct 1, N-part). The nucleotide sequence contains the signal sequence, 5'Strc (the 5' fragment of the wild-type Strc coding sequence), and the N-intane sequence (encoding the N-terminal fragment of the intein protein). [Figure 18B] See the explanation in Figure 18A. [Figure 19] Figure 19 shows a schematic diagram of a construct used in the double AAV intein-mediated protein transsplicing described herein, which includes a signal peptide sequence, the N-terminal portion of the STRC protein, and the nucleotide sequences of Figures 18A and 18B that encode the N-intane. [Figure 20] Figure 20 shows the N-terminal portion of the STRC protein, the signal peptide sequence, and the amino acid sequence containing the N-intane (SEQ ID NO: 52), which are encoded by the nucleotide sequences presented in Figures 18A and 18B. [Figure 21A] Figures 21A and 21B show the C-terminal portion of the STRC protein, the signal peptide sequence, and the nucleotide sequence encoding the C-intane (SEQ ID NO: 53) (e.g., CFS-C-Strc-C-Int, construct 1, C-portion). The nucleotide sequence contains the signal sequence, the C-intane sequence (encoding the C-terminal fragment of the intein protein), 3'Strc (the 3' fragment of the wild-type Strc coding sequence), the linker sequence, and the myc tag sequence. [Figure 21B] Refer to the explanation in Figure 21A. [Figure 22] Figure 22 shows a schematic diagram of a construct containing the signal peptide sequence, the C-intane, and the nucleotide sequences of Figures 21A and 21B that encode the C-terminal portion of the STRC protein, used in the double AAV intein-mediated protein transsplicing described herein. [Figure 23] Figure 23 shows the amino acid sequence (SEQ ID NO: 54) containing the signal peptide sequence, C-intane, C-terminal portion of the STRC protein, linker sequence, and myc tag, which are encoded by the nucleotide sequences presented in Figures 21A and 21B. [Figure 24] Figure 24 confirms the dual AAV intein-mediated protein trans-splicing and processing described herein, as demonstrated by a Western blot of isolated STRC protein in lane 4 using the sequences of Figures 8A, 8B, 13A, and 13B, illustrated by the construct (AAV2 / AAV9-Php.B-STRC-construct 2) in Figures 9 and 14. [Figure 25] Figure 25 confirms the usefulness of signal sequences in the dual AAV intein-mediated protein trans-splicing and processing described herein, as demonstrated by a Western blot of isolated STRC proteins in lane 6 using the sequences of Figures 8A, 8B, 13A, and 13B, illustrated by the constructs (AAV2 / AAV9-Php.B-STRC-construct 2) in Figures 9 and 14, which include signal sequences, in contrast to lane 4, which lacks signal sequences. [Figure 26] Figure 26 shows the recovery of hearing loss using the dual AAV in-mediated protein trans-splicing and processing described herein, as evidenced by the recovery of sound pressure levels (decibels, dB) in STRC knockout mice (Strc- / -) infected with the constructs in Figures 9 and 14, compared to wild-type (WT) mice (StrcWT / WT) and STRC knockout mice (Strc- / -). [Figure 27]Figure 27 shows the results of ABR and DPOAE in Strac knockout mice, demonstrating the restoration of auditory function following treatment with the dual AAV intein-mediated protein trans-splicing system described herein (construct 2: AAV2 / AAV9-PHP.B-CMV-Strc-N; AAV2 / AAV9-PHP.B-CMV-Strc-C). [Figure 28] Figure 28 shows ABR and DPOAE results demonstrating the lack of auditory function recovery in Strc knockout mice using only the construct encoding the N-terminal portion of the STRC protein shown in Figure 9. [Figure 29] Figure 29 shows the ABR and DPOAE results demonstrating the lack of auditory function recovery in Strc knockout mice using only the construct encoding the C-terminal portion of the STRC protein illustrated in Figure 14. [Figure 30] Figure 30 shows the in vivo time-course ABR and DPOAE results after treatment of Strc knockout mice with the dual AAV intein-mediated protein transsplicing system described herein (construction 2: AAV2 / AAV9-PHP.B-CMV-Strc-N; AAV2 / AAV9-PHP.B-CMV-Strc-C). [Figure 31A]Figures 31A–31C provide a dual-vector strategy using intein-mediated protein recombination. Figure 31A provides eight AAV2 plasmids, generated and contained in four different dual-vector variants. The N-terminal and C-terminal inteins were fused in-frame at the indicated sites for each of the four variants. Variants 1 and 2 differed in their splitting sites, with native cysteine ​​located at position 747 (Variant 1) and 970 (Variant 2). Variants 3 and 4 had identical splitting sites 1 and 2, respectively. Furthermore, variants 3 and 4 had a signal sequence upstream of the C-intane sequence, fused to the N-terminus of the STRC, as seen at the N-terminus of the C-terminal fragment. Intein-mediated protein recombination was predicted to yield full-length STRC and cleaved intein fragments. A Myc tag (not shown) was fused to the C-terminus of all C-terminal fragments. Figure 31B shows the splitting sites and surrounding amino acid sequences for the four variants. Figure 31C provides representative Western blot analyses of lysates derived from human embryonic kidney (HEK) 293 cells transfected with a plasmid encoding full-length STRC (lane 1), untransfected control (lane 2), HEK293 cells transfected with plasmids encoding the C-terminal fragments of variant 1 (lane 3) and variant 3 (lane 5), and HEK293 cells co-transfected with both the N- and C-fragments of variant 1 (lane 4) and variant 3 (lane 6). Anti-Myc antibodies were used to identify the C-terminal fragment (120kD) and full-length STRC (220kD). [Figure 31B] Refer to the explanation in Figure 31A. [Figure 31C] Refer to the explanation in Figure 31A. [Figure 32A]Figures 32A–32G show the generation and characterization of StrcΔ / Δ mice. Figure 32A illustrates the CRISPR / Cas9 strategy for disruption of WT Strc. Three guide RNAs (sgRNAs) were designed to target exon 4. The gene disruption strategy yielded a 249-nucleotide deletion, as well as two transpositions and inversions (947-1139 - purple and 1758-1835 - yellow), introducing a premature stop codon into the mutant Strc allele. Figure 32B provides the results of the PCR used to amplify genomic DNA, which gave clear bands for the WT (1kB) and mutant Strc (751bp) alleles when performed on a gel. Figures 32C–32D illustrate confocal images of WT and StrcΔ / Δ cochlea taken from tissue, stained with anti-STRC antibody, Alexa488 secondary, and phalloidin-Alexa555 for irradiation of hair follicles. Inner hair cells (IHCs) and 3-row outer hair cells (OHCs) are visible. Green STRC staining is seen in WT OHCs, and red actin staining is seen in both WT and StrcΔ / Δ IHCs. Scale bar = 10 mm. Figure 32E shows the mean ± SD of sensory conduction current amplitudes measured from IHCs and OHCs of StrcΔ / + mice (heterozygous - black circles) and StrcΔ / Δ mice (homozygous - red diamonds). Figure 32F shows the mean ± SD of DPOAE thresholds obtained from StrcD / D mice (n=6; red) and WT mice (n=5; black). Figure 32G shows the mean ± SD of ABR thresholds obtained from StrcΔ / Δ mice (n=6; red) and WT mice (n=5; black). [Figure 32B] Refer to the explanation in Figure 32A. [Figure 32C] Refer to the explanation in Figure 32A. [Figure 32D] Refer to the explanation in Figure 32A. [Figure 32E] Refer to the explanation in Figure 32A. [Figure 32F] Refer to the explanation in Figure 32A. [Figure 32G] Refer to the explanation in Figure 32A. [Figure 33A]Figures 33A-B demonstrate that dual AAV delivery restores STRC expression and hair follicle morphology. Figure 33A provides confocal images of a WT cochle (left), a StrcΔ / Δ cochle (center), and a dual-vector injected StrcΔ / Δ cochle (right) stained with anti-STRC antibody and Alexa488 conjugate secondary (green) and Alexa546-phalloidin (red). Scale bar = 10 mm. The top row of the image shows the two merged channels. The bottom row shows the localization of individual STRCs. Figure 33B shows scanning electron microscopy images of outer hair cell bundles in WT (top), StrcΔ / Δ (center), and dual-vector injected StrcΔ / Δ (bottom). Tissue was collected, fixed, and imaged at 4 or 12 weeks as described above. Scale bar = 5 mm (left and center) or 2 mm (right). [Figure 33B] Refer to the explanation in Figure 33A. [Figure 34A]Figures 34A–34D show that dual AAV delivery restores DPOAE and ABR thresholds. Figure 34A provides a Fourier analysis of the DPOAE waveform (top trace) revealing two frequency components at stimulation frequencies f1 (13.3 kHz) and f2 (16 kHz) in WT mouse cochlea, as well as a distortion component at the predicted frequency 2f1–f2 (10.6 kHz). The bottom trace shows the distortion component for sound pressure levels of 10–50 dB on extended frequency and amplitude scales for WT cochlea (left), StrcΔ / Δ cochlea (center), and dual-vector injected StrcΔ / Δ cochlea (right). The thick traces indicate the DPOAE threshold. Figure 34B provides the DPOAE threshold as a function of stimulation frequency (f2) for WT mice (black), as well as restored dual-vector injected StrcΔ / Δ mice (purple) and non-restored mice (StrcΔ / Δ; red; n=20; top horizontal line). The red and purple lines represent data from individual mice. Mean ± SE is shown for WT mice (black; n=5; bottom line), 20 recovered double-vector injected StrcΔ / Δ mice (purple; n=20; second line from the top), and the 5 best-recovered mice (green; n=5; third line from the top). Figure 34C illustrates a family of ABR traces recorded from WT cochlea (left), StrcΔ / Δ cochlea (center), and double-vector injected StrcΔ / Δ cochlea (right) induced by sound pressure levels of 25–110 dB. Thick traces indicate ABR thresholds. Figure 34D provides ABR thresholds plotted as a function of stimulus frequency for WT mice (black; n=5; bottom line), as well as recovered double-vector injected StrcΔ / Δ mice (purple; n=20; second line from the top) and non-recovered mice (StrcΔ / Δ; red; n=20; top line). The purple lines represent data from individual mice. The mean ± SE is shown for WT mice (black; n=5; bottom line), recovered double-vector injected StrcΔ / Δ mice (purple; n=20; second line from the top), non-recovered mice (StrcΔ / Δ; red; n=20; top line), and the best-recovered mice (green; n=5; third line from the top). [Figure 34B] Refer to the explanation in Figure 34A. [Figure 34C] Refer to the explanation in Figure 34A. [Figure 34D] Refer to the explanation in Figure 34A. [Modes for carrying out the invention]

[0022] Detailed explanation Detailed aspects of the present disclosure are disclosed herein; however, it should be understood that the disclosed aspects are merely illustrative of the various forms in which the present disclosure may be embodied. Furthermore, each of the examples given with respect to the various aspects of the present disclosure is illustrative and not limiting.

[0023] Compositions and methods for restoring hearing through the expression of stereocillin (STRC) are provided herein. Treatment with the gene of interest using two separate AAV particles or vectors, one comprising a signal sequence, a 5' terminal fragment of the gene coding sequence, and a sequence encoding the amino terminal fragment of intein (also known as N-intene, split-intene-N), and the other comprising a signal sequence, a sequence encoding the carboxyl terminal fragment of intein (also known as C-intene, split-intene-C), and a 3' terminal fragment of the gene coding sequence.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the meanings generally understood by those skilled in the art in the field to which this invention pertains. The following references provide general definitions of many of the terms used herein: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991).

[0025] definition All terms used herein shall have the ordinary meaning in the art unless otherwise provided. All concentrations are given as a weight percentage of the specified component relative to the total weight of the topical composition, unless otherwise defined.

[0026] As used herein, “a” or “an” means one or more. As used herein, when used in conjunction with the word “including,” “a” or “an” means one or more. As used herein, “another” means at least a second or more.

[0027] Adeno-associated virus (AAV) is a small virus that can infect humans and several other primate species. Vectors using AAV can infect dividing and quiescent cells without being integrated into the host cell's genome. Due to these characteristics, AAV is an attractive candidate for gene therapy viral vectors.

[0028] "AAV9-php.b vector" means a viral vector containing the AAV9-php.b polynucleotide or a fragment thereof that can transfect cells, for example, cells of the inner ear, adeno-associated virus serotype 9. In one embodiment, the AAV9-php.b vector transfects at least 70% or more of the cells (e.g., 75%, 80%, 85%, 90%, 95%, 97%, 99%, 100%). Another embodiment may relate to an AAV9-php.b vector containing the AAV9-php.b polynucleotide or a fragment thereof that can transfect at least 70% or more of the inner hair cells and / or outer hair cells (e.g., 75%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) after administration of the AAV9-php.b vector to the inner ear of a subject, or after contact of the AAV9-php.b vector with cells derived from the inner ear in vitro. In another embodiment, at least 85% (e.g., 90%, 95%, 100%) of inner hair cells and / or at least 85% (e.g., 90%, 95%, 100%) of outer hair cells are transfected with the AAV9-php.b vector. Transfection efficiency can be assessed in a mouse model using a label or tag (e.g., a gene encoding green fluorescent protein (GFP)).

[0029] One aspect of this disclosure may relate to at least one vector (e.g., plasmid, transpricing plasmid, viral vector (e.g., lentivirus), adenovirus, AAV, AAV genome) containing a nucleotide sequence encoding a desired protein (Figure 1). Figures 2A-2C show a nucleotide sequence (SEQ ID NO: 33) containing the human STRC gene coding sequence (SEQ ID NO: 1) (encoding the human STRC protein sequence (uppercase) in Figure 3, SEQ ID NO: 2) in the 5' to 3' direction: TIFF0007846626000001.tif97158TIFF0007846626000002.tif241158TIFF0007846626000003.tif241158TIFF0007846626000004.tif26158

[0030] Further embodiments may relate to at least one vector (e.g., plasmid, transpricing plasmid, viral vector, adenovirus, AAV, AAV genome) (see, for example, Figures 4A-4D; see SEQ ID NO:38) containing the mouse STRC gene coding sequence (SEQ ID NO:3) (encoding the mouse STRC protein sequence, SEQ ID NO:4 in Figure 5) in the 5' to 3' direction: TIFF0007846626000005.tif169158TIFF0007846626000006.tif241158TIFF0007846626000007.tif206158

[0031] Another aspect of this disclosure is a first vector (e.g., plasmid, transpricing plasmid, viral vector, adenovirus, AAV, AAV genome) comprising a first nucleotide sequence (e.g., SEQ ID NO:5 encoding SEQ ID NO:6; SEQ ID NO:7 encoding SEQ ID NO:8) including a signal sequence (e.g., SEQ ID NO:9 encoding SEQ ID NO:10; or SEQ ID NO:11 encoding SEQ ID NO:12) at the 5' end of the subcoding sequence in the 5' to 3' direction (the subcoding sequence may be adjacent to a downstream sequence encoding a splice donor sequence (e.g., N-terminal intein (N-intine); SEQ ID NO:13 encoding SEQ ID NO:14)); a subcoding sequence encoding the amino-terminal (N-terminal) portion of the protein of interest (e.g., STRC; SEQ ID NO:15; SEQ ID NO:16): and a first vector (e.g., plasmid, transpricing plasmid, viral vector, adenovirus, AAV, AAV genome) including a signal sequence (e.g., SEQ ID NO:11 encoding human SEQ ID NO:10) at the 5' end of the subcoding sequence in the 5' to 3' direction NO:9; or SEQ ID NO:11 encoding mouse SEQ ID NO:12) may be upstream of a splice acceptor sequence (e.g., C-terminal intein (C-intine); SEQ ID NO:21 encoding SEQ ID NO:22), which may provide a second vector (e.g., plasmid, transpricing plasmid, viral vector, adenovirus, AAV, AAV genome) containing a second nucleotide sequence (e.g., SEQ ID NO:17 encoding SEQ ID NO:18; SEQ ID NO:19 encoding SEQ ID NO:20) that is located directly adjacent to the downstream subcoding sequence encoding the remainder of the full-length coding sequence of the protein of interest, i.e., the C-terminal portion of the protein of interest (e.g., STRC; human SEQ ID NO:23; SEQ ID NO:24).When expressed in cells (e.g., mammalian host cells (e.g., human, dog, cat, horse, mouse)), the first and second vectors can each express their respective portions (e.g., N-STRC, C-STRC) of the protein of interest, which form the full-length protein of interest (e.g., STRC; SEQ ID NO: 25; SEQ ID NO: 26).

[0032] Another embodiment provides a first vector (e.g., plasmid, transplicating plasmid, viral vector, adenovirus, AAV, AAV genome) comprising a partial coding sequence encoding the amino-terminal (N-terminal) portion of a protein of interest (e.g., STRC), and a first nucleotide sequence containing a signal sequence at the 5' end of the partial coding sequence, wherein the partial coding sequence may be adjacent to a downstream sequence encoding a splice donor sequence (e.g., N-terminal intein (N-intane)), and the splice donor sequence is adjacent to a downstream 3' ITR sequence (e.g., AAV9-php.B-Prot-trans / donor). A second vector (e.g., plasmid, transpricing plasmid, viral vector, adenovirus, AAV, AAV genome) (e.g., AAV9-phpB-Prot-trans / acceptor) comprising a second nucleotide sequence containing: a 5'ITR sequence that may be upstream of a signal sequence that is directly adjacent to a downstream subcoding sequence encoding the remaining C-terminal portion of the protein of interest (e.g., STRC), wherein the second nucleotide sequence may further contain a C-terminal myc tag downstream of the subcoding sequence encoding the C-terminal portion of the protein of interest. In cells co-infected with two vectors (e.g., plasmids, transplicating plasmids, viral vectors, adenoviruses, AAVs, AAV genomes), head-to-tail recombination (5'-3'-5'-3'), transcription, and subsequent splicing at terminal inversion repeat (ITR) junctions of the two transplicating plasmids can form the full-length mRNA of interest (e.g., STRC mRNA).

[0033] "The sequence encoding the N-terminal portion or N-terminal fragment of the protein of interest" means, for example, if the protein of interest is a stereocillin (STRC) protein, a partial coding sequence of the N-terminal portion or N-terminal fragment of STRC, and in some cases, "the sequence encoding the N-terminal portion or N-terminal fragment of STRC" may include, upstream of the partial coding sequence (uppercase) of the N-terminal portion or N-terminal fragment of STRC, a sequence (lowercase, italicized, underlined) encoding a signal peptide coding sequence. In some embodiments, "the sequence encoding the N-terminal fragment of stereocillin (STRC)" may not include a nucleotide sequence encoding a signal peptide coding sequence. Another embodiment provides a nucleotide sequence comprising "the sequence encoding the N-terminal fragment or N-terminal portion of STRC" fused at its C-terminus to the N-terminal portion or N-terminal fragment of intein (N-intein) (bold, underlined), where the nucleotide sequence begins with an "ATG" start codon (bold) and ends with a stop codon (uppercase, italicized, underlined) at its 5' end.

[0034] An example human nucleotide sequence containing "a sequence encoding the N-terminal portion or N-terminal fragment of STRC" may be as follows: TIFF0007846626000008.tif104158TIFF0007846626000009.tif177158

[0035] Another exemplary mouse nucleotide sequence containing "a sequence encoding the N-terminal portion or fragment of STRC" could be as follows: TIFF0007846626000010.tif39158TIFF0007846626000011.tif242158TIFF0007846626000012.tif12158

[0036] "The N-terminal portion or N-terminal fragment of the protein of interest" means the amino acid sequence of the N-terminal portion or N-terminal fragment of a stereocillin (STRC) polypeptide, for example, if the protein of interest is the STRC protein. For example, the amino acid sequence of the N-terminal fragment of STRC may include a signal peptide sequence (lowercase, italicized, underlined) (e.g., 22 amino acids) at the N-terminus, beginning with methionine (M) (bold N-terminus) encoded by the ATG start codon. The amino acid sequence containing "the N-terminal portion or N-terminal fragment of the STRC protein" may further include, downstream and / or adjacent to, an intein N-terminal fragment (N-intine) (bold, underlined). In yet another embodiment, "the N-terminal portion or N-terminal fragment of the STRC protein" means the amino acid sequence of the N-terminal fragment of STRC that does not include a signal peptide sequence.

[0037] An exemplary amino acid sequence containing the N-terminal portion or fragment of a human STRC protein may be as follows: TIFF0007846626000013.tif98158

[0038] Another exemplary amino acid sequence containing the N-terminal portion or fragment of the mouse STRC protein may be as follows: TIFF0007846626000014.tif98158

[0039] Exemplary "N-terminal STRC polypeptide" sequences are provided below for human and mouse, respectively (from N-terminus to C-terminus).

[0040] Exemplary human N-terminal STRC polypeptide sequences (N-terminus to C-terminus direction) that do not contain methionine or signal peptide sequences encoded by the ATG start codon are provided below (N-terminus to C-terminus direction): TIFF0007846626000015.tif84158

[0041] Another exemplary mouse N-terminal STRC polypeptide sequence (N-terminus to C-terminus direction) that does not contain the methionine or signal peptide sequence encoded by the ATG start codon (17-residue hydrophobic region is underlined) (N-terminus to C-terminus direction): TIFF0007846626000016.tif83158

[0042] "A sequence encoding the C-terminal portion or C-terminal fragment of stereocillin (STRC)" means a partial coding sequence of the C-terminal portion or C-terminal fragment of STRC. One embodiment may provide a nucleotide sequence comprising a sequence encoding a signal peptide coding sequence (lowercase, italicized, underlined) adjacent to a sequence encoding the C-terminal fragment of intein (C-intene) (bold, underlined), which is adjacent to the upstream of the "nucleotide sequence encoding the C-terminal portion or C-terminal fragment of the STRC protein" (bold, underlined). Another embodiment may further provide a nucleotide sequence encoding the C-terminal portion or C-terminal fragment of STRC, comprising a downstream linker sequence (bold, italicized), a Myc tag (lowercase), and a stop codon (uppercase, italicized, underlined).

[0043] An exemplary human nucleotide sequence containing "a sequence encoding the C-terminal portion or C-terminal fragment of STRC" may be as follows: TIFF0007846626000017.tif46158TIFF0007846626000018.tif241158TIFF0007846626000019.tif105158

[0044] An exemplary mouse nucleotide sequence containing "a sequence encoding the C-terminal portion or C-terminal fragment of STRC" may be as follows: TIFF0007846626000020.tif111158TIFF0007846626000021.tif241158TIFF0007846626000022.tif41158

[0045] "The C-terminal portion or C-terminal fragment of the STRC protein" refers to the amino acid sequence of the C-terminal portion or C-terminal fragment of the stereocillin (STRC) polypeptide. The amino acid sequence containing "the C-terminal portion or C-terminal fragment of the STRC protein" may be preceded, from N-terminus to C-terminus, by methionine (M) (bold at the N-terminus) encoded by the ATG start codon, a signal peptide sequence (lowercase, italicized, underlined), and the C-terminal fragment of intein (C-intene) (bold, underlined). Downstream of "the C-terminal portion or C-terminal fragment of the STRC protein," the amino acid sequence may further include a linker sequence (bold, italicized) and a Myc tag (lowercase).

[0046] An exemplary amino acid sequence containing the C-terminal portion or fragment of a human STRC protein may be as follows: TIFF0007846626000023.tif134158

[0047] Another exemplary amino acid sequence containing the mouse "C-terminal portion or C-terminal fragment of the STRC protein" may be as follows: TIFF0007846626000024.tif134158

[0048] Exemplary "C-terminal STRC polypeptide" sequences are provided below for human and mouse, respectively (from N-terminus to C-terminus).

[0049] An exemplary human C-terminal region of a STRC protein that does not contain a signal peptide sequence may be as follows: TIFF0007846626000025.tif127158

[0050] An exemplary mouse C-terminal region of an STRC protein that does not contain a signal peptide sequence but includes a hydrophobic region of at least 16 underlined residues may be as follows: TIFF0007846626000026.tif128158

[0051] Ligation between the N-terminal portion of the STRC protein and the C-terminal portion of the STRC protein can occur, for example, through a peptide bond, thereby resulting in a full-length STRC protein.

[0052] An "intine" is a protein fragment that, in a process known as protein splicing, can cleave itself and be joined to the remaining fragment (extine) by a peptide bond. Inteins are also called "protein introns." The process by which an intein cleaves itself and joins to the remaining part of a protein is referred to herein as "protein splicing" or "intine-mediated protein splicing." In some embodiments, the intein of a precursor protein (an intein-containing protein before intein-mediated protein splicing) originates from two genes. Such inteins are referred herein as split inteins (e.g., split intein-N and split intein-C). For example, in cyanobacteria, DnaE, ​​the catalytic subunit a of DNA polymerase III, is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene may be referred herein as "intine-N." The intein encoded by the dnaE-c gene may be referred to herein as "intine-C".

[0053] Other intein systems may also be used. For example, synthetic inteins based on dnaE inteins, the intein pairs Cfa-N (e.g., split intein-N) and Cfa-C (e.g., split intein-C) are described (e.g., Stevens et al., J Am Chem Soc. 2016 Feb. 24;138(7):2162-5, incorporated herein by reference). Non-limiting examples of intein pairs that may be used pursuant to this disclosure include Cfa DnaE intein, Ssp GyrB intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein, and Cne Prp8 intein (e.g., those described in U.S. Patent No. 8,394,604, incorporated herein by reference).

[0054] The "myc tag" refers to a polypeptide protein derived from the c-myc gene, and its synthetic peptide sequence (i.e., EQKLISEEDL (SEQ ID NO:27)) corresponds to the C-terminal amino acids (410-419) of the human c-myc protein. This tag enables further research, such as, but not limited to, protein isolation (e.g., Western blotting, immunofluorescence, immunoprecipitation).

[0055] An example sequence of the AAV9-php.b vector is provided below. The AAV9-php.b vector (5' to 3') is, in some embodiments, This may provide nucleotide sequences that have at least 70% (for example, 75%, 80%, 85%, 90%, 95%, 97%, 100%) identity with TIFF0007846626000027.tif10158TIFF0007846626000028.tif241158TIFF0007846626000029.tif241158TIFF0007846626000030.tif241158TIFF0007846626000031.tif233158TIFF0007846626000032.tif12160.

[0056] "Administer" means providing one or more compositions, constructs, or viral vectors described herein to a target. For example, non-limitingly, administration may be carried out, for example, by injection into the cochlea. Other routes for delivering the composition to the mutation-affected cells (e.g., intravenous, direct injection, subcutaneous, vascular, and / or non-vascular intravenous) may be utilized. Administration may be, for example, by bolus injection or by slow perfusion over a period of time.

[0057] "Drug" means any small molecular weight chemical compound, antibody, nucleic acid molecule, polypeptide, or fragment thereof.

[0058] "Modification" means a change (increase or decrease) in the expression level or activity of a gene or polypeptide, as detected by standard known methods, such as those described herein. As used herein, modification may include changes in expression levels of 10% or more (e.g., 20%, 25%, 30%, 40%, 50%).

[0059] "To induce remission" means to reduce, decrease, suppress, weaken, shrink, halt, or stabilize the onset or progression of a disease or condition. Exemplary diseases or conditions may include any disease, e.g., those related to gene mutations. In one embodiment, a disease may be related to a dominant or recessive mutation, e.g., autosomal recessive hearing loss 1A (DFNB1A); autosomal recessive hearing loss 1B (DFNB1B); autosomal recessive hearing loss 2 (DFNB2); autosomal recessive hearing loss 4 (DFNB4); autosomal recessive hearing loss 6 (DFNB6); autosomal recessive hearing loss 7 (DFNB7); autosomal recessive hearing loss 8 (DFNB8); autosomal recessive hearing loss 9 (DFNB9); autosomal recessive hearing loss 10 (DFNB10); autosomal recessive hearing loss 16 (DFNB16); autosomal recessive hearing loss 24 (DFNB24); autosomal recessive hearing loss 25 (DFNB25); This could be autosomal recessive hearing loss 59 (DFNB59); autosomal recessive hearing loss 98 (DFNB98); autosomal recessive hearing loss 110 (DFNB110); autosomal recessive congenital keratosis disorder 5 (DKCB5); autosomal dominant hearing loss 2B (DFNA2B); autosomal dominant hearing loss 3A (DFNA3A); autosomal dominant hearing loss 9 (DFNA9); autosomal dominant hearing loss 30 (DFNA30); autosomal dominant hearing loss 36 (DFNA36); autosomal dominant isolated sensorineural hearing loss (DFNA); autosomal dominant congenital keratosis disorder 4; autosomal dominant non-syndromic sensorineural hearing loss (NSSHL); or autosomal dominant Wolfram-like syndrome (WFSL).

[0060] Further diseases or conditions that the dual-vector systems described herein may treat include, but are not limited to, dentinogenesis imperfecta (DGI)1; autosomal recessive hearing loss16; hearing loss infertility syndrome (DIS); CATSPER-associated male infertility; spermatogenesis imperfecta7 (SPGF7); Usher syndrome type I (USH1); Bloom syndrome (BLM); cloacal exstrophy; Pendred syndrome (PDS); gyroretinal choroidal atrophy (GACR); cataract41 (CTRCT41); prostate cancer; and breast cancer.

[0061] This disclosure may provide cells (e.g., host cells) which are any cells that hold or can hold the substance of interest. Often, host cells are mammalian cells (e.g., human, dog, cat, horse, mouse). Host cells can accept, for example, AAV constructs, AAV plasmids, helper constructs, auxiliary function vectors, etc. Host cells that may be used herein include offspring of a transfected initial cell. The term “host cell” in this disclosure may also mean a cell that has been transfected with an exogenous DNA sequence. Offspring of a single parent cell may not be morphologically or genomically or whole-DNA complementary to the original parent due to spontaneous mutation, unintended mutation, or intentional mutation. The term “transfection” may mean the uptake of exogenous DNA by a cell, as used herein, and a cell is “transfected” when the exogenous DNA has been introduced inside the cell membrane. Numerous transfection techniques are generally known in the art. See, for example, Graham et al. (1973) Virology, 52:456, Sambrook et al. (1989) Molecular Cloning, a laboratory manual, Cold Spring Harbor Laboratories, New York, Davis et al. (1986) Basic Methods in Molecular Biology, Elsevier, and Chu et al. (1981) Gene 13:197. Such techniques may be used to introduce one or more exogenous nucleic acids, such as nucleotide embedding vectors and other nucleic acid molecules, into suitable host cells.

[0062] "Nucleic acid," "nucleotide sequence," and "polynucleotide sequence" mean polymers of deoxyribonucleotides or ribonucleotides in either single-stranded or double-stranded form, and unless otherwise limited, include known analogues of natural nucleotides that can function in a manner similar to naturally occurring nucleotides.

[0063] The terms “substantially homologous,” “substantially identical,” or “substantially equivalent” refer to a characteristic of a nucleic acid sequence or amino acid sequence in which the selected nucleic acid sequence or amino acid sequence has at least 70% sequence identity with respect to the selected reference nucleic acid sequence or amino acid sequence. The selected sequence and the reference sequence may have at least 75% sequence identity (e.g., 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%).

[0064] Sequence identity or homology may be determined in terms of the entire length of the sequences being compared, or by a fragment of the sequence that may be 25% or less in total than that of a selected reference sequence (e.g., 20%, 15%, 10%, 5%). The reference sequence may be a portion of a larger sequence, e.g., a gene or a portion of an adjacent sequence, or a repeating portion of a chromosome. Two or more polynucleotide sequences may typically be compared to a reference sequence having at least 18–25 nucleotides, at least 26–35 nucleotides, or at least 40 (e.g., 50, 60, 70, 80, 90, 100, 500, 1000, 1500, 2000) nucleotides. Sequence identity or homology may be determined using well-known sequence comparison algorithms, e.g., the FASTA Biological Sequence Alignment / Comparison Software Program (see, e.g., Pearson and Lipman, 1985, 1988).

[0065] Polynucleotides, such as vectors and plasmids, for delivering a portion of a gene coding sequence of interest, for example, a portion of the STRC gene encoding a stereocillin protein, to a cell are provided herein. Some embodiments may provide a coding sequence derived from a human STRC gene containing 29 exons of 19kb (see, for example, NCBI gene ID: 161497; AC016135.3; NG_011636.1; NCBI nucleotide ID: NM_153700.2; AF375594). Other embodiments may provide a mouse (Mus musculus) STRC gene (see, for example, NCBI gene ID: 140476; AL845466; AF375593; AK144985; NM_080459). Stereocillin (STRC) protein or STRC precursor, which has 1,809 amino acids (see, for example, NCBI protein ID: NP_714544; AAL35321; BAE26168; NP_536707; UniProtKB / Swiss Prot: Q7RTU9), is a large extracellular structural protein found in the immobile ciliates of outer hair cells in the inner ear.

[0066] The STRC gene (e.g., NCBI accession number NG_011636; OMIM#603720; MIM 606440) encodes the stereocillin (STRC) protein, a large extracellular structural protein found in the immobile ciliates of outer hair cells in the inner ear, which are associated with horizontal top connectors and tectorial membrane attachment crowns, crucial for the proper aggregation and positioning of immobile ciliate tips. The STRC gene is located on chromosome 15q15, defines the autosomal recessive DFNB16 hearing loss locus, and contains 29 exons encompassing 19kb. "STRC polynucleotide" refers to a nucleic acid molecule encoding the STRC polypeptide or a fragment thereof.

[0067] The underlying cause of autosomal recessive non-syndromic hearing loss (DFNB16) is associated with a mutated STRC gene. The human STRC gene sequence (gene ID: 161497) contains the human nucleotide STRC coding sequence encoding the human stereocillin (STRC) protein (NCBI RefSeq: NP_714544). The human STRC coding sequence without the signal sequence (e.g., SEQ ID NO: 29) is as follows: TIFF0007846626000033.tif205158TIFF0007846626000034.tif241158TIFF0007846626000035.tif154158. The human STRC code sequence may be found in the constructs in Figures 2A-2C, and parts of the human STRC code sequence may be found in the constructs in Figures 7A-7B and 12A-12B.

[0068] The human mRNA sequence encoding the human STRC protein (NCBI RefSeq: NM153700) (SEQ ID NO: 30) is as follows: TIFF0007846626000036.tif39158TIFF0007846626000037.tif241158TIFF0007846626000038.tif241158TIFF0007846626000039.tif111158

[0069] Stereocillin expression is found only in sensory hair cells of the inner ear and is associated with immobile cilia, i.e., stiff microvilli that form structures for mechanoreception of auditory stimuli. The human STRC protein (1,775 amino acids; SEQ ID NO: 25), which contains the signal peptide sequence (amino acids 1-21, underlined), does not contain a linker sequence or Myc tag sequence, and whose splice site may be between Ala708 and Cys709 or between Ala933 and Cys934 (bold, underlined), is as follows: TIFF0007846626000040.tif205158

[0070] The mouse (Mus musculus) STRC gene (gene ID: 140476; CDS at base pairs 79-5508) encodes the mouse STRC protein (NCBI RefSeq: NP_536707), which contains 1,809 amino acids including a putative signal peptide and several hydrophobic moies. The mouse STRC coding sequence without the signal sequence (e.g., SEQ ID NO: 31) is as follows: TIFF0007846626000041.tif75158TIFF0007846626000042.tif241158TIFF0007846626000043.tif241158TIFF0007846626000044.tif54158. Mouse STRC code sequences may be found in the structures in Figures 4A-4D, and parts of the mouse STRC code sequences may be found in the structures in Figures 8A-8B and 13A-13B.

[0071] The mouse mRNA sequence encoding the mouse STRC protein (NCBI RefSeq: NM_080459; SEQ ID NO: 32) is as follows: TIFF0007846626000045.tif147158TIFF0007846626000046.tif241158TIFF0007846626000047.tif241158TIFF0007846626000048.tif25158

[0072] The mouse STRC protein (1,809 amino acids; SEQ ID NO: 26), which contains a signal peptide sequence (amino acids 1-22; underlined) and does not contain a linker sequence or Myc tag sequence, is obtained by cleaving a 22-amino acid signal peptide sequence, leaving a protein with 1,787 amino acids and a predicted molecular weight of 194 kD. The sequence is as follows, where the splice sites of Ser746 and Cys747 and Ala969 and Cys970 are in bold and underlined: TIFF0007846626000049.tif212158

[0073] "Stereocillin (STRC) protein" means a polypeptide or fragment thereof that has at least approximately 80% (e.g., 82%, 85%, 88%, 90%, 95%, 97%, 98%, 99%, 100%) amino acid identity with, for example, NCBI accession number NP_714544 or NP_536707; GenBank number AAL35321.

[0074] "Detecting" means identifying the presence, absence, or quantity of the analyte to be detected.

[0075] A “detectable label” means a composition that, when attached to a molecule of interest, makes it detectable by spectroscopic, photochemical, biochemical, immunochemical, or chemical means. Useful labels include, for example, radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes such as green fluorescent protein, high electron density reagents, enzymes (such as those commonly used in ELISA), biotin, digoxigenin, or haptens.

[0076] "Disease" means any condition or disorder that impairs or interferes with the normal function of cells, tissues, or organs. Examples of diseases include, but are not limited to, any pathology, such as hearing impairment, such as hearing impairment associated with recessive mutations, such as DFNB16.

[0077] The "effective dose" refers to the amount of medication needed to alleviate the symptoms of a disease compared to an untreated patient. The effective dose of the active compound used to carry out therapeutic treatment for a disease varies depending on the mode of administration, the age, weight, and overall health of the subject. Ultimately, the attending physician or veterinarian will determine the appropriate dose and administration plan. Such a dose is called the "effective" dose.

[0078] A “fragment” means a portion of a polypeptide or nucleic acid sequence or molecule. This portion may contain at least 10% (e.g., 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%) of the total length of the reference nucleic acid molecule or polypeptide. A fragment may contain 10 or more (e.g., 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500) nucleotides or amino acids.

[0079] "Hybridization" refers to hydrogen bonding between complementary nucleic acid bases, which can be Watson-Crick, Hoogsteen, or reverse Hoogsteen type hydrogen bonds. For example, adenine and thymine are complementary nucleic acid bases that form base pairs through the formation of hydrogen bonds.

[0080] "Identity" means the identity of the amino acid sequence or nucleic acid sequence between the sequence of interest and the reference sequence. "Substantially identical" means a polypeptide or nucleic acid molecule that exhibits at least 50% identity with a reference amino acid sequence (e.g., any of the amino acid sequences described herein) or nucleic acid sequence (e.g., any of the nucleic acid sequences described herein). Such a sequence may have at least 60% (e.g., 70%, 75%, 80%, 85%, 85%, 90%, 95%, 99%, 100%) identity at the amino acid level or nucleic acid level with respect to the sequence or reference sequence used for comparison.

[0081] Sequence identity can be measured using sequence analysis software (e.g., the sequence analysis software packages from Genetics Computer Group, University of Wisconsin Biotechnology Center (1710 University Avenue, Madison, Wis. 53705), such as BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In an exemplary approach to determining the degree of identity, e -3 ~e -100 A BLAST program can be used, where the probability scores indicate closely related sequences.

[0082] The terms “isolated,” “purified,” or “biologically pure” refer to material from which components normally associated with it, as found in its natural state, have been removed to varying degrees. “Isolating” means separation to some extent from its original source or surroundings. “Purifying” indicates a higher degree of separation than isolation. A “purified” or “biologically pure” protein is one from which other material has been sufficiently removed so that impurities do not substantially affect the biological properties of the protein and do not cause any other harmful consequences. That is, the nucleic acids or peptides of this disclosure are purified, for example, when produced by recombinant DNA technology, from which cellular material, viral material, and culture medium have been substantially removed, or when chemically synthesized, from which chemical precursors or other chemicals have been substantially removed. Purity and homogeneity are typically determined using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high-performance liquid chromatography. The term “purified” may indicate that the nucleic acid or protein produces essentially a single band in an electrophoretic gel. For proteins that can undergo modifications, such as phosphorylation or glycosylation, different modifications can result in different isolated proteins, which can then be purified separately.

[0083] "Isolated polynucleotide" means the nucleic acid (e.g., DNA) of the present disclosure, in which genes adjacent to the gene have been removed from the naturally occurring genome of the organism from which the nucleic acid molecule of the present invention originates. Accordingly, this term includes recombinant DNA that exists, for example, as a vector; an autonomously replicating plasmid or virus; or as something integrated into the genomic DNA of a prokaryotic or eukaryotic organism; or as another molecule independent of other sequences (e.g., cDNA or a fragment of genome or cDNA produced by PCR or restriction endonuclease digestion). Furthermore, this term includes RNA molecules transcribed from DNA molecules and recombinant DNA that is part of a hybrid gene encoding an additional polypeptide sequence.

[0084] "Isolated polypeptide" means the polypeptide of this disclosure that has been isolated from its naturally associated components. Typically, a polypeptide is isolated when at least 60% (e.g., 75%, 80%, 90%, 95%, 99%) by weight of the proteins and naturally occurring organic molecules associated with it have been removed. Isolated polypeptides of this disclosure may be obtained, for example, by extraction from natural sources, by expression of recombinant nucleic acids encoding such polypeptides, or by chemical synthesis of proteins. Purity may be measured by any suitable method, for example, by column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.

[0085] "Marker" means any protein or polynucleotide that has a modified expression level or activity associated with a disease or disorder.

[0086] "Mechanosensory perception" refers to responses to mechanical stimuli. Examples of the conversion of mechanical stimuli into neuronal signals include touch, hearing, and balance. Mechanosensory input is converted into responses to mechanical stimuli through a process called "mechanotransduction."

[0087] As used herein, "obtaining a drug" includes acquiring a drug by synthesis, purchase, or other means.

[0088] As used herein, terms such as “prevent,” “prevention,” “prevention,” and “preventive measures” refer to reducing the probability of developing a disorder or condition in subjects who do not currently have the disorder or condition but are at risk of developing it or are prone to developing it.

[0089] "Promoter" means a polynucleotide sufficient to direct the transcription of a downstream polynucleotide. In some embodiments, the polynucleotides described herein may comprise one or more regulatory elements. Those skilled in the art can select appropriate regulatory elements for use in cells, e.g., mammalian or human host cells. Non-limiting examples of regulatory elements include promoters, transcription termination sequences, translation termination sequences, enhancers, and polyadenylation elements. The polynucleotides described herein may comprise promoter sequences functionally linked to a nucleotide sequence encoding a desired polypeptide, e.g., stereocillin. Promoters intended for use in the present invention include, but are not limited to, the cytomegalovirus (CMV) promoter, the SV40 promoter, the Roussarcoma virus (RSV) promoter, the chimeric CMV / chicken β-actin promoter (CBA), and the truncated form of CBA (smCBA). In some embodiments, the promoter is the CMV promoter.

[0090] The phrase “pharmaceutically acceptable excipient” may include pharmaceutically acceptable materials, compositions, or media, such as liquid or solid fillers, diluents, carriers, solvents, or encapsulating materials, that are involved in the transport or delivery of the compound of the present invention from one organ or body part to another. Each carrier must be “acceptable” in the sense that it is compatible with the other components of the formulation and is not harmful to the patient. Some examples of materials that can function as pharmaceutically acceptable carriers include: (1) sugars, e.g., lactose, glucose, and sucrose; (2) starches, e.g., corn starch and potato starch; (3) cellulose and its derivatives, e.g., sodium carboxymethylcellulose, ethylcellulose, and cellulose acetate; (4) tragacanth powder; (5) malt; (6) gelatin; (7) talc; (8) excipients, e.g., cocoa butter and suppository wax; (9) oils, e.g., peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; (10) glycerides (11) Coals, e.g., propylene glycol; (12) Polyols, e.g., glycerin, sorbitol, mannitol, and polyethylene glycol; (13) Esters, e.g., ethyl oleate and ethyl laurate; (14) Agar; (15) Buffers, e.g., magnesium hydroxide and aluminum hydroxide; (16) Alginic acid; (17) Phenothermally hydrated substances; (18) Isotonic saline; (19) Ringer's solution; (10) Ethyl alcohol; (21) Phosphate-buffered aqueous solution; and (22) Other non-toxic, suitable substances used in pharmaceutical formulations.

[0091] Suitable additional carriers and their formulations are described, for example, in the latest edition of Remington's Pharmaceutical Sciences by E.W. Martin. The amount of therapeutic agent administered varies depending on the mode and method of administration, age, and disease state (e.g., the degree of hearing loss present before treatment).

[0092] "Stereocillin promoter" means a regulatory polynucleotide sequence derived from NCBI reference sequence:NG_011636.1 that is sufficient to direct the expression of a downstream polynucleotide in the inner hair cells (IHC) or outer hair cells (OHC) of a mature cochlear, in the horizontal upper coupler connecting the apical regions of adjacent immobile cilia within a hair bundle, and in the junction that attaches the longest immobile cilia to the upper tectorial membrane (TM). Stereocillin may also be expressed around the cilia of vestibular hair cells and immature OHCs. One aspect of this disclosure provides a stereocillin promoter comprising or comprising at least 350 base pairs or more (e.g., 500, 1000, 2000, 3000, 4000, 5000 base pairs) upstream of the stereocillin coding sequence.

[0093] "To reduce" means a negative modification of at least 5% or more (e.g., 10%, 15%, 20%, 25%, 50%, 75%, 100%).

[0094] "Reference" refers to a standard condition or a control condition.

[0095] A "reference sequence" is a clear sequence that can be used as the basis for sequence comparison. A reference sequence may be a part or all of a particular sequence, for example, a full-length cDNA or a fragment of a gene sequence, or a complete cDNA or gene sequence. For polypeptides, the length of a reference polypeptide sequence may be at least 10 amino acids or more (e.g., 15, 20, 25, 30, 35, 50, 100), or any integer close to or between those numbers. For nucleic acids, the length of a reference nucleic acid sequence may be at least 50 nucleotides or more (e.g., 55, 60, 75, 90, 100, 200, 300), or any integer close to or between those numbers.

[0096] "Specifically binding" means a compound or antibody that, in a sample naturally containing the polypeptide of the Disclosure, such as a biological sample, recognizes and binds to the polypeptide of the Disclosure, but substantially does not recognize or bind to any other molecule.

[0097] Nucleic acid molecules useful in the methods of the present disclosure include any nucleic acid molecules encoding the polypeptide of the present disclosure or a fragment or portion thereof. Such nucleic acid molecules do not need to be 100% identical to the endogenous nucleic acid sequence, but will typically exhibit substantial identity. Polynucleotides having "substantial identity" to the endogenous sequence can typically hybridize with at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecules encoding the polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules do not need to be 100% identical to the endogenous nucleic acid sequence, but will typically exhibit substantial identity. Polynucleotides having "substantial identity" to the endogenous sequence can typically hybridize with at least one strand of a double-stranded nucleic acid molecule. "Hybridizing" means a pair of complementary polynucleotide sequences (e.g., genes described herein) or portions thereof that form a double-stranded molecule under various stringency conditions (see, for example, Wahl, G.M. and S.B. Berger (1987) Methods Enzymol. 152:399; Kimmel, AR (1987) Methods Enzymol. 152:507).

[0098] Hybridization can occur under stringent salt concentrations, which may typically be less than 750 mM NaCl (e.g., 500 mM; 250 mM) and less than 75 mM trisodium citrate (e.g., 50 mM; 25 mM). Low-stringency hybridization can be achieved in the absence of organic solvents, such as formamide, while high-stringency hybridization can be achieved in the presence of at least 35% formamide (e.g., 50% formamide). Stringent temperature conditions may typically include temperatures of at least 30°C (e.g., 37°C, 42°C). Additional parameters that vary, such as hybridization time, the concentration of surfactants, such as sodium dodecyl sulfate (SDS), and the inclusion or exclusion of carrier DNA, are well known to those skilled in the art. Various levels of stringency can be achieved by combining these various conditions as needed. In one embodiment, hybridization may occur at 30°C with 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS. In another embodiment, hybridization may occur at 37°C with 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA). In yet another embodiment, hybridization may occur at 42°C with 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA. Useful variations of these conditions will be readily apparent to those skilled in the art.

[0099] For most applications, the stringency of the post-hybridization washing process can also vary. Washing stringency conditions can be defined by salt concentration and temperature. As mentioned above, washing stringency can be increased by decreasing salt concentration or increasing temperature. For example, a stringent salt concentration for the washing process may be less than 30 mM NaCl (e.g., 15 mM) and less than 3 mM trisodium citrate (e.g., 1.5 mM). A stringent temperature condition for the washing process can typically include a temperature of at least 25°C (e.g., 42°C, 68°C). In yet another embodiment, the washing process may be carried out at 25°C with 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. Another embodiment may provide a washing process carried out at 42°C with 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Further embodiments may provide a washing step carried out at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Further variations of these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.

[0100] "Subject" means mammals, including, but not limited to, human or non-human mammals such as cattle, horses, dogs, sheep, cats, or mice.

[0101] "STRC protein or polypeptide" means a polypeptide or fragment thereof that has sufficient activity to express STRC, which is essential for auditory function, and has at least approximately 85% amino acid sequence identity with NCBI accession number NP_714544 or GenBank No. AAL35321.

[0102] As used herein, terms such as “to treat,” “to treat,” and “treatment” refer to the reduction or remission of any disorder and / or associated symptoms. It will be understood, though not excluded, that treatment of a disorder or condition does not require the complete elimination of the disorder, condition, or associated symptoms.

[0103] Unless otherwise specified and evident from the context, the term “or” as used herein is understood to be inclusive. Unless otherwise specified and evident from the context, the terms “a,” “an,” and “the” as used herein are understood to be singular or plural.

[0104] In this disclosure, “comprises,” “comprising,” “contains,” and “have,” etc., may have the meanings given to them under U.S. patent law, and may mean “includes,” “including,” etc.; “essentially from” or “essentially,” likewise, may have the meanings given to them under U.S. patent law, and these terms are open-ended and allow for other existences, as long as the basic or novel features of what is described are not altered by the existence of other existences, but exclude aspects of the prior art. “Essentially” means, for example, that the components consist only of the components described, together with ordinary impurities present in commercially available materials and any other additives present at levels that do not affect the performance of the aspects disclosed herein, for example, less than 5% by weight, or 1% by weight, or even less than 0.5% by weight.

[0105] As used herein, all numerical ranges include the disclosed endpoints and all possible values ​​between the disclosed values. The exact values ​​of all half-integer values ​​are also construed as specifically disclosed and as limits for all subsets of the disclosed ranges. For example, the range 1 to 50 is understood to include any number, combination of numbers, or subrange from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50. Furthermore, the range 0.1% to 3% specifically discloses percentages of, for example, 0.1%, 1%, 1.5%, 2.0%, 2.5%, and 3%, or any other number in between. Additionally, the range 0.1% to 3% includes subsets of the first range, such as 0.5% to 2.5%, 1% to 3%, 0.1% to 2.5%, etc. It will be understood that the sum of all weight percent of the individual components will not exceed 100%. The ranges provided herein are understood to be abbreviations for all values ​​within that range.

[0106] Any enumeration of chemical groups in any definition of a variable herein includes the definition of that variable as any single group or combination of the listed groups. Any enumeration of aspects of a variable or aspect herein includes that aspect as any single aspect or in combination with any other aspect or part thereof.

[0107] Any composition or method provided herein may be combined with one or more of the other compositions and methods provided herein.

[0108] This disclosure relates to a system encoding a split protein of interest, e.g., a STRC protein (e.g., a vector, a recombinant virus), a nucleic acid sequence, or a construct, and to a method for delivering a protein of interest (e.g., STRC) to a host, host cell, or cell using the nucleic acid sequences or constructs described herein. A split protein, comprising an amino-terminal (N-terminal) fragment and a carboxy-terminal (C-terminal) fragment of a protein, is each encoded by a different, separate nucleic acid sequence or construct for delivery to a cell, for example, using a vector (e.g., a viral vector, e.g., adeno-associated virus (AAV) or lentivirus). The N-terminal polypeptide fragment and C-terminal polypeptide fragment of the protein of interest may be joined together to form a full-length protein of interest, for example, using intein-mediated protein splicing, in which the N-terminal polypeptide fragment and terminal polypeptide fragment of the protein of interest are linked by peptide bonds. The split sites may be located in similar or homologous regions across different species.

[0109] Vector System In some embodiments, a vector system (e.g., capsid, plasmid, transpricing plasmid, viral vector, adenovirus, AAV, AAV genome, lentivirus) for delivering a coding sequence of a desired full-length protein, comprising at least one vector containing the desired gene construct. Figure 1 shows a schematic diagram of such an exemplary AAV STRC construct. In other embodiments, the vector system may comprise the entire human STRC coding sequence (SEQ ID NO: 1), where the nucleotide sequence (Figures 2A-2C; SEQ ID NO: 33) begins with an "ATG" start codon (bold) at the 5' end and ends with a stop codon (uppercase, italicized, underlined), and the signal peptide coding sequence (lowercase, italicized, underlined; SEQ ID NO: 9) is located upstream of the STRC coding sequence. The linker sequence (bold, italic; SEQ ID NO:34) and the sequence encoding the myc tag (lowercase; SEQ ID NO:35) may optionally be included at the 3' end for use in subsequent studies, for example, but not limited to, protein isolation (e.g., Western blotting, immunofluorescence, immunoprecipitation). Figure 3 shows the amino acid sequence (SEQ ID NO:36) encoded by the human STRC nucleotide sequence of Figure 2, where the human STRC protein sequence includes methionine (M) (bold) encoded by the ATG codon, a signal peptide sequence (lowercase, italic, underlined; SEQ ID NO:10), an optional linker sequence (bold, italic) (N-terminus-TRTRPL-C-terminus; SEQ ID NO:37), and an optional Myc tag (lowercase; SEQ ID NO:27).

[0110] In another embodiment, the vector system may include the entire mouse STRC coding sequence (SEQ ID NO: 3), where the nucleotide sequence (Figures 4A-4D; SEQ ID NO: 38) begins with an "ATG" start codon (bold) and ends with a stop codon (uppercase, italic, underlined) at the 5' end, and the signal peptide coding sequence (lowercase, italic, underlined; SEQ ID NO: 11) is located upstream of the STRC coding sequence. Optionally, a linker sequence (bold, italic; SEQ ID NO: 34) and a sequence encoding the myc tag (lowercase; SEQ ID NO: 35) may be included at the 3' end for use in subsequent studies, such as, but not limited to, protein isolation (e.g., Western blotting, immunofluorescence, immunoprecipitation). Figure 5 shows the amino acid sequence encoded by the mouse STRC nucleotide sequence in Figure 4, where the mouse STRC protein sequence (SEQ ID NO:39) includes methionine (M) encoded by the ATG codon (bold), a signal peptide sequence (lowercase, italic, underlined), an optional linker sequence (bold, italic) (SEQ ID NO:37), and an optional Myc tag (lowercase) (SEQ ID NO:27).

[0111] Those skilled in the art will understand that it is quite possible to construct vectors (e.g., viruses (bacteriophages, animal viruses, and plant viruses, capsids), plasmids, cosmids, and artificial chromosomes (e.g., YACs)) through standard recombination techniques described in MR. Green and J. Sambrook, Molecular Cloning: A Laboratory Manual (2012), 4th Ed.; and Ausubel et al., Current Protocols in Molecular Biology (2003), both of which are incorporated herein by reference. Furthermore, methods for transferring or delivering DNA into cells are disclosed, for example, in Sung, Y., Kim, S. "Recent advances in the development of gene delivery systems." Biomator Res 23, 8 (2019); Jin, Lian et al. "Current progress in gene delivery technology based on chemical methods and nano-carriers." Theranostics 4(3): 240-255, 2014; Nayerossadat N, Maedeh T, Ali PA. "Viral and nonviral delivery systems for gene delivery." Adv Biomed Res. 1: 27, 2012; Machida, Curtis A. Viral Vectors for Gene Therapy Methods and Protocols. Humana Press, 2003; Heiser, William C. Gene Delivery to Mammalian Cells. Humana Press, 2004, respectively, which are incorporated herein by reference.

[0112] Intein-mediated protein transsplicing Other embodiments may provide a vector system, for example, a dual-vector system comprising two vectors for delivering different portions of the same desired protein of interest, each using an intein. An intein can be considered a protein intron, which is a portion of a protein that, by a process known as protein splicing, cuts itself out from an amino acid sequence and connects the remaining adjacent region (extine) by a peptide bond. For the intein-mediated protein trans-splicing process, see, for example, Mills et al. "Protein Splicing: How Inteins Escape from Precursor Protein" JBC.289(21):14498-14505, 2014, which is incorporated in its entirety herein by reference. Intein-mediated protein splicing occurs after the mRNA containing the intein has been translated into a protein. See Figure 6.

[0113] Inteins are a class of enzymes that catalyze a reaction in which they cleave themselves from a host protein-intene fusion, thereby yielding a mature host protein (extene) and isolated inteins, where a peptide bond ligates the splice site between the donor intein and the acceptor intein. Any known inteins, including but not limited to those identified by Perler, FB (InBase. The intein database. Nucleic Acids Res. 30:383-384, 2002), whose entirety may be incorporated herein by reference (for example, Npu-PCC73102 (DnaE-c intein (accession number ZP_00108882); DnaE-n intein (accession number ZP_00111398)) derived from Nostoc punctiforme PCC 73102 (ATCC® 29133®), whose entirety may be incorporated herein by reference), may be used in the manner of this disclosure.

[0114] In some aspects of this disclosure, the catalytic subunit of DNA polymerase III DnaE from the cyanobacterium Nostoc punctiforme (Npu) may be used in the split-intine-STRC dual vector system described herein. For example, the α subunit of DNA polymerase III DnaE from the cyanobacterium Nostoc punctiforme (Npu) is naturally split and located in two genes, dnaE-n and dnaE-c. The N-terminal and C-terminal portions of dnaE are encoded by two separate genes on opposite DNA strands in the genome. The protein encoded by dnaE-n contains an amino-terminal (N-terminal) dnaE fragment (e.g., N-extine) and an amino-terminal intein (N-intine), while dnaE-c encodes a protein containing a carboxy-terminal (C-intine) entity followed by a carboxy-terminal (C-terminal) dnaE fragment (e.g., C-extine). N-intines and C-intines recognize each other, splice themselves out of their amino acid sequences, and simultaneously ligate or fuse with adjacent N-terminal and C-terminal extensions of the target of interest via peptide bonds, thereby resulting in a fusion that forms the full-length protein of interest.

[0115] Intein activity can be environment-dependent, and specific peptide sequences may be required around the ligation or fusion junctions (called N-extrains and C-extrains) for efficient splicing to occur. For example, amino acids containing a thiol or hydroxyl group (e.g., cysteine ​​(Cys), serine (Ser), threonine (Thr)) as the first residue of the C-extrain may be useful for efficient splicing. Natural cysteine ​​can be used as a splice site. For example, the AAV vector can utilize natural cysteine ​​at positions 747 (Cys747; variants 1 and 3) and 970 (Cys970; variants 2 and 4) of SEQ ID NO:26, which can provide a splice site between Ser746 and Cys747 or between Ala969 and Cys970 (Figures 31A-31B), where the splitting site is the amino acid sequence: It may be present in TIFF0007846626000050.tif4128 or a portion thereof. Another example provides a splice site between Ala708 and Cys709 or between Ala933 and Cys934 in SEQ ID NO:25, where the splitting site is the amino acid sequence: It may be present in TIFF0007846626000051.tif4128 or a portion thereof.

[0116] One aspect of this disclosure provides an extain region containing each half of a protein of interest, e.g., a stereocillin. For example, the extain region may include an N-terminal and a C-terminal stereocillin that, when fused via a peptide bond, constitute a full-length stereocillin (e.g., NCBI:NM_153700;NP_714544;NM_080459;NP_536707;GenBank:BK000138;AF375594;DAA00085;AF375593;AAL35321). See Figure 6. Another aspect provides different splitting sites for a stereocillin, where the N-terminal and C-terminal portions of the protein of interest (e.g., a stereocillin) together constitute 100%. Furthermore, in a further aspect, the splitting site may be located inside, behind, before, or adjacent to the helix (e.g., α) of the protein secondary structure. One aspect may provide a splitting site that is not inside a β-chain or crosslink. Another embodiment provides a partitioning site within the helix (e.g., α) immediately upstream of the coil region. Yet another embodiment provides the protein of interest (e.g., stereocillin) and fragments of the protein of interest that are outside the transmembrane domain, i.e., not within the transmembrane region.

[0117] For example, splitting may occur such that the N-terminal portion of the protein of interest (e.g., stereocillin; human SEQ ID NO:15 or mouse SEQ ID NO:16) contains at least 10% (e.g., 15%, 20%, 25%, 30%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%) of the length of the N-terminus of the full-length protein of interest (e.g., full-length human SEQ ID NO:2 or SEQ ID NO:25 or mouse SEQ ID NO:4 or SEQ ID NO:26). Further embodiments provide the N-terminal portion of a protein of interest (e.g., stereocillin; human SEQ ID NO:15 or mouse SEQ ID NO:16) comprising a length of 100% or less of the N-terminus of the full-length protein of interest (e.g., full-length human SEQ ID NO:25 or mouse SEQ ID NO:26) (e.g., 99%, 97%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%). One other embodiment may provide the N-terminal portion of the protein of interest (e.g., stereocillin; human SEQ ID NO: 15 or mouse SEQ ID NO: 16) comprising 10% to 100% (e.g., 15% to 90%, 20% to 80%, 30% to 70%, 40% to 60%, 50%) of the length of the N-terminal portion of the full-length protein of interest (e.g., full-length human SEQ ID NO: 25 or mouse SEQ ID NO: 26).Further embodiments provide an N-terminal portion of a protein of interest (e.g., stereocillin; human SEQ ID NO:15 or mouse SEQ ID NO:16) comprising less than 54% (e.g., 53%, 52%, 51%, 50%, 45%, 43%, 41%, 40%) of the N-terminus of the full-length protein of interest (e.g., full-length human SEQ ID NO:25 or mouse SEQ ID NO:26), and / or less than 54% identity to the N-terminus of the full-length protein of interest (e.g., full-length human SEQ ID NO:25 or mouse SEQ ID NO:26), and / or less than 54% length of the N-terminus of the full-length protein of interest (e.g., full-length human SEQ ID NO:25 or mouse SEQ ID NO:16). Furthermore, a further embodiment may relate to the N-terminal portion of a protein of interest (e.g., stereocillin; human SEQ ID NO: 15 or mouse SEQ ID NO: 16), which includes a length less than 54% of the N-terminus of the full-length protein of interest (e.g., full-length human SEQ ID NO: 25 or mouse SEQ ID NO: 26). Another embodiment provides an N-terminal portion of a protein of interest (e.g., stereocillin; human SEQ ID NO:15 or mouse SEQ ID NO:16) that includes 40% or more of the N-terminus of the full-length protein of interest (e.g., full-length human SEQ ID NO:25 or mouse SEQ ID NO:26) (e.g., 41%, 42%, 43%, 44%, 45%, 50%, 51%, 52%, 53%), and / or 41% or more identity to the N-terminal portion of the full-length protein of interest (e.g., full-length human SEQ ID NO:25 or mouse SEQ ID NO:26), and / or 41% or more of the length of the N-terminal portion of the full-length protein of interest (e.g., full-length human SEQ ID NO:25 or mouse SEQ ID NO:26).

[0118] In yet another embodiment, the C-terminal portion of the protein of interest (e.g., stereocillin; human SEQ ID NO:23 or mouse SEQ ID NO:24) comprises at least 10% (e.g., 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%) of the length of the C-terminus of the full-length protein of interest (stereocillin; full-length human SEQ ID NO:25 or mouse SEQ ID NO:26). A further embodiment provides a C-terminal portion of a protein of interest (e.g., stereocillin; human SEQ ID NO:23 or mouse SEQ ID NO:24) comprising a length of 100% or less of the C-terminus of the full-length protein of interest (stereocillin; full-length human SEQ ID NO:25 or mouse SEQ ID NO:26) (e.g., 99%, 97%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%). One other embodiment may provide a C-terminal portion of a protein of interest (e.g., stereocillin; human SEQ ID NO: 23 or mouse SEQ ID NO: 24) comprising 10% to 99% (e.g., 15% to 90%, 20% to 80%, 30% to 70%, 40% to 60%, 50%) of the C-terminal portion of the full-length protein of interest (e.g., stereocillin; human SEQ ID NO: 23 or mouse SEQ ID NO: 24). A further embodiment provides a C-terminal portion of a protein of interest (e.g., stereocillin; human SEQ ID NO: 23 or mouse SEQ ID NO: 24) comprising 46% or more of the C-terminal portion of the full-length protein of interest (e.g., stereocillin; human SEQ ID NO: 25 or mouse SEQ ID NO: 26).Furthermore, a further embodiment may relate to a C-terminal portion of a protein of interest (e.g., stereocillin; human SEQ ID NO: 23 or mouse SEQ ID NO: 24) that includes 46% or more (e.g., 47%, 48%, 49%, 50%, 55%, 60%) of the C-terminal portion of the full-length protein of interest (e.g., stereocillin; full-length human SEQ ID NO: 25 or mouse SEQ ID NO: 26), and / or 46% or more identity with the C-terminal portion of the full-length protein of interest (e.g., stereocillin; full-length human SEQ ID NO: 25 or mouse SEQ ID NO: 26), and / or 46% or more of the length of the C-terminal portion of the C-terminal portion of the protein of interest (e.g., stereocillin; human SEQ ID NO: 23 or mouse SEQ ID NO: 24). Furthermore, a further embodiment may relate to the C-terminal portion of the protein of interest (e.g., stereocillin; human SEQ ID NO: 23 or mouse SEQ ID NO: 24) which includes a length of 46% or more (e.g., 47%, 48%, 49%, 50%, 55%, 60%) of the N-terminus of the full-length protein of interest (e.g., stereocillin; full-length human SEQ ID NO: 25 or mouse SEQ ID NO: 26). Another embodiment provides a C-terminal portion of a protein of interest (e.g., stereocillin; human SEQ ID NO: 23 or mouse SEQ ID NO: 24) comprising 60% or less of the C-terminus of the full-length protein of interest (e.g., stereocillin; full-length human SEQ ID NO: 25 or mouse SEQ ID NO: 26) (e.g., 59%, 58%, 57%, 56%, 55%, 50%, 45%) and / or 60% or less of identity with the C-terminus of the full-length protein of interest (e.g., stereocillin; full-length human SEQ ID NO: 25 or mouse SEQ ID NO: 26) and / or a length of 60% or less of the C-terminus of the full-length protein of interest (e.g., stereocillin; full-length human SEQ ID NO: 25 or mouse SEQ ID NO: 26).

[0119] A further aspect provides for splitting of the full-length wild-type stereocilin protein to form N-terminal and C-terminal fragments of stereocilin, for example, between Ala708 and Cys709 or between Ala933 and Cys934 of human SEQ ID NO:25, or between Ser746 and Cys747 or between Ala969 and Cys970 of mouse SEQ ID NO:26.

[0120] Dual vector system One aspect of the disclosure is a dual vector system for expressing a protein of interest in a cell, comprising: (a) In the 5' to 3' direction, - A signal sequence (e.g., SEQ ID NO:11 encoding SEQ ID NO:12) at the 5' end of a partial coding sequence; - A partial coding sequence encoding the amino-terminal (N-terminal) portion of the protein of interest (e.g., STRC; SEQ ID NO:15; SEQ ID NO:16); - A sequence encoding a splice donor sequence (e.g., the amino-terminal fragment of an intein (N-intein); SEQ ID NO:13 encoding SEQ ID NO:14) A first vector (e.g., a capsid, plasmid, trans-splicing plasmid, viral vector, adenovirus, AAV, AAV genome, lentivirus) comprising a first nucleotide sequence comprising: (b) In the 5' to 3' direction, - A signal sequence (e.g., SEQ ID NO:11 encoding SEQ ID NO:12) that may be upstream of a splice acceptor sequence; - A sequence encoding a splice acceptor sequence (e.g., the carboxy-terminal fragment of an intein (C-intein); SEQ ID NO:21 encoding SEQ ID NO:22); - A partial coding sequence encoding the carboxy-terminal (C-terminal) portion of the protein of interest (e.g., STRC; SEQ ID NO:23; SEQ ID NO:24) A second vector (e.g., capsid, plasmid, transpricing plasmid, viral vector, adenovirus, AAV, AAV genome, lentivirus) containing a second nucleotide sequence including We provide a dual vector system that includes this.

[0121] A further embodiment provides a dual vector system for expressing a protein of interest in a cell, comprising the following (see, for example, Figure 6): (a) In the direction from 5' to 3', - 5'-terminal inverted repeat (5'ITR) sequence; - Promoter sequence; - A signal sequence (e.g., SEQ ID NO:11 encoding SEQ ID NO:12) that is functionally linked to and under the control of the promoter; - A partial coding sequence that encodes the amino-terminal (N-terminal) portion of a protein of interest (e.g., STRC; SEQ ID NO: 15; SEQ ID NO: 16), which is functionally linked to a promoter and under the control of the promoter; - A sequence encoding a splice donor sequence (e.g., the amino-terminal fragment of intein (N-intine); SEQ ID NO:13 encoding SEQ ID NO:14), wherein the splice donor sequence (e.g., encoding N-intine) is functionally linked to and under the control of the promoter; - Polyadenylated (poly-A) signal sequence; - 3'-terminal inverted repeat (3'ITR) sequence A first vector containing a first nucleotide sequence, and (b) In the direction from 5' to 3', - 5'-terminal inverted repeat (5'ITR) sequence; - Promoter sequence; - A signal sequence (e.g., SEQ ID NO:11 encoding SEQ ID NO:12) that is functionally linked to and under the control of the promoter; - A sequence encoding a splice acceptor sequence (e.g., the carboxy-terminal fragment of intein (C-intein); SEQ ID NO:21 encoding SEQ ID NO:22), wherein the splice acceptor sequence (e.g., encoding C-intein) is functionally linked to and under the control of the promoter; - A partial coding sequence that encodes the carboxyl-terminal (C-terminal) portion of a protein of interest (e.g., STRC; SEQ ID NO: 23; SEQ ID NO: 24), which is functionally linked to a promoter and under the control of the promoter; - Polyadenylated (poly-A) signal sequence; - 3'-terminal inverted repeat (3'ITR) sequence A second vector containing a second nucleotide sequence including [the specified nucleotide sequence].

[0122] Another embodiment may relate to a dual vector system in which the first vector may contain a first nucleotide sequence containing the sequence of SEQ ID NO: 5 or 7 in the 5' to 3' direction. See, for example, Figures 7A, 7B, 8A, 8B, and 9. Figures 7A-7B (SEQ ID NO: 5) and 8A-8B (SEQ ID NO: 7) contain the start codon (ATG, bold), signal sequence (italicized, underlined) in the 5' to 3' direction (signal sequence: TIFF0007846626000052.tif19160), coding sequence of the N-terminal portion of the STRC gene (black), and N-intane (underlined) (N-intane sequence: Figure 9 shows the nucleotide sequence containing TIFF0007846626000053.tif48159), where the coding sequence encodes the stereocillin (STRC) protein. Figure 9 shows additional elements including the ITR, promoter, and poly-A tail, where mouse 5'STRC encodes the N-terminal portion (1-746 (Ser) amino acids; 79.7 kDa) of the full-length stereocillin (STRC) protein. Figures 10 and 11 show the amino acid sequence encoded by the first nucleotide sequence, where the amino acid sequence containing the N-terminal portion of the STRC protein (SEQ ID NO: 6; SEQ ID NO: 8, respectively) encodes methionine (M) (bold) encoded by the ATG codon, and the signal peptide sequence (lowercase, italicized, underlined) (signal peptide sequence: TIFF0007846626000054.tif12128), the N-terminal portion of the stereocillin protein (black), and N-intei (bold, underlined) Includes TIFF0007846626000055.tif19158.

[0123] A further embodiment may provide a dual vector system in which the second vector may contain a second nucleotide sequence containing the sequence of SEQ ID NO: 17 or 19 in the 5' to 3' direction. See, for example, Figures 12A, 12B, 13A, 13B, and 14. Figures 12A-12B (SEQ ID NO: 17) and 13A-13B (SEQ ID NO: 19) contain, in the 5' to 3' direction, a start codon (bold ATG), a signal sequence (lowercase, italicized, and underlined; SEQ ID NO: 9; SEQ ID NO: 11, respectively), and a C-intane sequence (bold and underlined). TIFF0007846626000056.tif19159, coding sequence of the C-terminal portion of the STRC gene (black) (the coding sequence codes for the stereocillin (STRC) protein), linker sequence (bold, italic). TIFF0007846626000057.tif4128, and Myc tag array (lowercase) Figure 14 shows the nucleotide sequence including TIFF0007846626000058.tif5146 and the stop codon (italicized, underlined). Figure 14 shows additional elements, e.g., the ITR, promoter, and poly-A tail, where mouse 3'-STRC encodes the C-terminal portion (747(Cys)~1,810 amino acids; 116.7 kDa) of the full-length stereocillin (STRC) protein. Figures 15 and 16 show the amino acid sequence encoded by the second nucleotide sequence, where the amino acid sequence of the C-terminal portion (SEQ ID NO: 18; SEQ ID NO: 20) is encoded by the ATG codon, methionine (M) (bold), signal peptide sequence (lowercase, italicized, underlined) (SEQ ID NO: 10; SEQ ID NO: 44), and C-intine (bold, underlined). TIFF0007846626000059.tif5138 includes the C-terminal portion of the stereosirin protein (black), the linker sequence (bold, italic) (N-TRTRPL-C; SEQ ID NO: 50), and the Myc tag (lowercase) (SEQ ID NO: 27). One embodiment may provide a full-length STRC protein in which the C-terminal portion of the stereosirin protein begins with cysteine ​​(C; Cys), phenylalanine (F; Phe), and serine (S; Ser) (Figure 17).

[0124] Furthermore, in a further embodiment, the dual vector system does not have to provide a first vector containing a first nucleotide sequence containing the sequence of SEQ ID NO: 51 in the 5' to 3' direction. See, for example, Figures 18A, 18B, and 19. Figures 18A-18B (SEQ ID NO: 51) illustrate a nucleotide sequence containing a start codon (ATG, bold), a signal sequence (lowercase, italicized, underlined) (SEQ ID NO: 11), a coding sequence for the N-terminal portion of the STRC gene (black), and an N-intane (bold, underlined) (SEQ ID NO: 42) (the coding sequence codes for the stereocillin (STRC) protein), as well as a stop codon (italicized, underlined), in the 5' to 3' direction. Figure 19 shows additional elements, such as the ITR, promoter, and poly-A tail, where 5'STRC encodes the N-terminal portion (1-969(Ala) amino acids; 104.8 kDa) of the full-length stereocillin (STRC) protein. Figure 20 shows the amino acid sequence encoded by the first nucleotide sequence, where the N-terminal amino acid sequence (SEQ ID NO: 52) includes methionine (M) (bold) encoded by the ATG codon (SEQ ID NO: 44), the signal peptide sequence (lowercase, italicized, underlined) (SEQ ID NO: 44), the N-terminal portion of the stereocillin protein (black), and N-intene (bold, underlined) (SEQ ID NO: 45).

[0125] Further embodiments may not provide a dual vector system having a second vector containing a second nucleotide sequence containing the sequence of SEQ ID NO: 53 in the 5' to 3' direction. See, for example, Figures 21A, 21B, and 22. Figures 21A-21B (SEQ ID NO: 53) illustrate a nucleotide sequence containing a start codon (bold ATG), a signal sequence (lowercase, italicized, underlined) (SEQ ID NO: 11), a C-intane sequence (bold, underlined) (SEQ ID NO: 21), a coding sequence for the C-terminal portion of the STRC gene (black) (the coding sequence codes for the stereocillin (STRC) protein), a linker sequence (bold, italicized) (SEQ ID NO: 47), a Myc tag sequence (lowercase) (SEQ ID NO: 48), and a stop codon (italicized, underlined) in the 5' to 3' direction. Figure 22 shows additional elements including the ITR, promoter, and poly(A) tail, where 3'-STRC encodes the C-terminal portion (970(Cys)~1,810 amino acids; 91.6 kDa) of the full-length stereocillin (STRC) protein. Figure 23 shows the amino acid sequence encoded by the second nucleotide sequence, where the C-terminal amino acid sequence (SEQ ID NO: 54) includes methionine (M) (bold) encoded by the ATG codon, signal peptide sequence (lowercase, italicized, underlined) (SEQ ID NO: 44), C-intine (bold, underlined) (SEQ ID NO: 49), the C-terminal portion of the stereocillin protein (black), linker sequence (bold, italicized) (SEQ ID NO: 50), and Myc tag (lowercase) (SEQ ID NO: 27).

[0126] One embodiment of a dual-vector system provides vectors for each segmented portion of a protein of interest (e.g., stereocillin), namely the N-terminal and C-terminal portions. The segmented portions of the protein of interest (e.g., stereocillin), and any additional regions necessary for the control, production, or expression of the protein of interest, are such that each portion and associated region does not exceed the load capacity of the respective vector (e.g., virus (e.g., viral vector, bacteriophage, phage, retrovirus), plasmid, cosmid, bacterial artificial chromosome, yeast artificial chromosome, human artificial chromosome). The segmented portions and additional regions of the same protein of interest must not exceed the size of the load capacity of the selected vector.

[0127] One aspect of the present disclosure provides a vector system (e.g., a dual-vector system) for delivering genes containing large coding sequences, for example, 4 kB or larger (e.g., 4.5 kB, 5 kB, 5.5 kB, 5.8 kB, 6 kB, 6.5 kB, 7 kB, 7.5 kB, 8 kB, 8.5 kB, 9 kB, 9.5 kB, 10 kB, 11 kB, 12 kB). The vectors of the present disclosure (e.g., the first vector and the second vector) may each be a viral vector (e.g., adenovirus, adeno-associated virus (AAV), lentivirus, herpes simplex virus I, vaccinia virus), and in some embodiments, the viral vector may be an AAV vector or a recombinant AAV vector. Another embodiment may provide viral vectors of the same or different serotypes (e.g., AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, synthetic serotypes). Useful AAV serotypes in the disclosures described herein may include, but are not limited to, AAV1, AAV2, AAV5, AAV6, AAV8, and AAV9.

[0128] In one embodiment, the first and second vectors are viral vectors lacking viral DNA, such as adeno-associated virus (AAV) or recombinant AAV (rAAV) as interchangeably used herein. The first vector of the dual vector system comprises a nucleotide sequence containing, in the 5' to 3' direction: a 5' inverted repeat (5'ITR) sequence; a promoter sequence capable of driving the transcription of a downstream polynucleotide of interest (e.g., STRC); a signal sequence functionally linked to and under the control of the promoter; a partial coding sequence encoding the amino-terminal (N-terminal) portion of the protein of interest (e.g., STRC), functionally linked to and under the control of the promoter; a sequence encoding the amino-terminal fragment (N-intane) of intein, functionally linked to and under the control of the promoter; a polyadenylation (Poly-A) signal sequence; and a 3'ITR sequence.

[0129] In another embodiment, the second vector comprises, in the 5' to 3' direction, a 5'-terminal inverted repeat (5'ITR) sequence; a promoter sequence capable of driving the transcription of a downstream polynucleotide of interest (e.g., STRC); a signal sequence functionally linked to and under the control of the promoter; a sequence encoding the carboxy-terminal fragment (C-intane) of intein, functionally linked to and under the control of the promoter; a subcoding sequence encoding the carboxy-terminal (C-terminal) portion of the protein of interest (e.g., STRC), functionally linked to and under the control of the promoter; a polyadenylation (Poly-A) signal sequence; and a 3'ITR sequence.

[0130] Further, in a further aspect, when the first vector and the second vector of the present disclosure are inserted into a cell (e.g., a host cell, a mammalian cell, a human cell, a bacterial cell) by any means including, but not limited to, viral transduction, bacterial transformation using calcium chloride, bacterial transformation or transduction by bacterial mating or conjugation, transfection (e.g., electroporation, calcium phosphate, liposome-based transfection), gene gun, etc., the vector includes, in the direction from the N-terminus to the C-terminus, a signal peptide sequence linked to the N-terminal portion of a protein sequence of interest (e.g., STRC) fused to the N-intein protein sequence at the C-terminus, a first protein sequence; and, in the direction from the N-terminus to the C-terminus, a second protein sequence including a signal peptide sequence linked to the C-intein protein sequence fused to the N-terminus of the C-terminal portion of the protein sequence of interest (e.g., STRC). Intein-mediated protein splicing of the N-terminal portion of the protein of interest (e.g., STRC) and the C-terminal portion of the same protein of interest (e.g., STRC) results in the expression of the full-length protein of interest (e.g., STRC).

[0131] Another aspect provides a signal peptide sequence of a dual vector system, which may be located at the N-terminus of each protein sequence encoded by the dual vector system. In the first protein sequence including the N-terminal portion of the protein of interest (e.g., STRC), the signal peptide sequence may be upstream of the coding region of the N-terminal portion of the protein of interest and the N-intein. In the second protein sequence including the C-terminal portion of the protein of interest (e.g., STRC), the signal peptide sequence may be upstream of the C-intein and the coding region of the C-terminal portion of the protein of interest (e.g., STRC).

[0132] Another aspect of this disclosure provides a dual-vector system in which a first vector and a second vector express, in a cell, a first protein sequence and a second protein sequence, respectively, each containing the same signal peptide sequence. Thus, the same signal peptide sequence allows the first protein sequence and the second protein sequence to be transported to the same cellular compartment. In a different embodiment, the signal peptide sequences of the first and second protein sequences may be different, but these signal peptide sequences direct each protein sequence to the same cellular compartment. The signal peptide sequences of the first and second protein sequences may be configured to transport the first and second protein sequences to the same cellular compartment. This allows each of the protein sequences to come close enough to each other so that intein-mediated protein fusion occurs, thereby forming a full-length protein of interest (e.g., STRC). The signal peptide sequence may be related to the protein of interest (e.g., STRC). Further embodiments may provide a signal sequence encoding a signal peptide sequence associated with a protein other than the protein of interest, where the signal sequence of a first nucleotide sequence and the signal sequence of a second nucleotide sequence are different, and the signal sequence also encodes a different signal peptide sequence, but the signal sequence is configured to transport the first protein sequence and the second protein sequence to the same cellular compartment. Another embodiment of the present disclosure provides a signal sequence that directs two fragments to the same cellular compartment or intracellular compartment without interfering with intein-mediated trans-splicing. The signal sequence may be particularly useful to ensure that the first protein sequence and the second protein sequence are sufficiently close together, since the N-terminal portion of the protein of interest (e.g., STRC) and the C-terminal portion of the protein of interest (e.g., STRC) are able to form a full-length protein of interest (e.g., STRC) through a peptide bond.

[0133] In one embodiment of this disclosure, the signal sequence may include a nucleic acid sequence having at least 80% (e.g., 85%, 90%, 95%, 97%, 99%, 100%) identity with a signal sequence encoding a signal peptide sequence of the protein of interest (e.g., STRC). For example, the signal sequence may include a nucleic acid sequence having at least 80% (e.g., 85%, 90%, 95%, 97%, 99%, 100%) identity with SEQ ID NO:11 of the gene sequence encoding the STRC protein signal peptide, or the signal sequence may include a nucleic acid sequence consisting of, for example, the protein of interest and any signal sequences directing each portion of the protein of interest to the same cellular compartment (e.g., SEQ ID NO:9 or SEQ ID NO:11). Another embodiment may provide a signal sequence encoding a signal peptide sequence having an amino acid sequence that is at least 80% (e.g., 85%, 90%, 95%, 97%, 99%, 100%) identical to the signal peptide sequence of the protein of interest (e.g., STRC; SEQ ID NO:10; SEQ ID NO:12). For example, the signal peptide sequence may include an amino acid sequence that is at least 80% (e.g., 85%, 90%, 95%, 97%, 99%, 100%) identical to SEQ ID NO:10 or SEQ ID NO:12 of the STRC protein signal peptide, or the signal peptide sequence may include an amino acid sequence consisting of SEQ ID NO:10 or SEQ ID NO:12.

[0134] Furthermore, a further embodiment can provide a partial coding sequence that codes for the N-terminal portion of a protein of interest (e.g., STRC), and which includes a nucleic acid sequence having at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with the coding sequence that codes for the N-terminal portion of the protein of interest (e.g., STRC; SEQ ID NO: 6, 8, 15, 16, 25, or 26). For example, a partial coding sequence encoding the N-terminal portion of a STRC protein (including a signal sequence that may be interchangeable with a different signal sequence) may contain nucleic acid sequences that have at least 5% identity (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) with the following sequences:

[0135] The partial coding sequence that encodes the N-terminal portion of the human STRC protein (including the start codon ATG (bold) and the signal sequence (lowercase, italicized, and underlined)) may be as follows: TIFF0007846626000060.tif249158

[0136] The partial coding sequence that encodes the N-terminal portion of the mouse STRC protein (including the start codon ATG (bold) and the signal sequence (lowercase, italicized, and underlined)) may be as follows: A partial coding sequence encoding the N-terminal portion of a STRC protein comprising a nucleic acid sequence consisting of TIFF0007846626000061.tif205158TIFF0007846626000062.tif55158, or SEQ ID NO: 55 or 56. Another embodiment provides the N-terminal portion of a protein of interest (e.g., STRC; SEQ ID NO: 25 or 26) having an amino acid sequence (including methionine and signal peptide sequences corresponding to the start codon ATG) that has at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with the protein of interest (e.g., STRC; SEQ ID NO: 25 or 26). For example, the N-terminal portion of an STRC protein (containing methionine corresponding to the start codon ATG and a potentially interchangeable STRC signal peptide sequence) may contain an amino acid sequence having at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with the sequence of SEQ ID NO: 15 or 16, or the N-terminal portion of an STRC protein containing an amino acid sequence consisting of SEQ ID NO: 15 or 16. The first nucleotide sequence of the first vector may include a subcoding sequence that encodes the N-terminal portion of the protein of interest (e.g., STRC; SEQ ID NO: 15 or 16), which includes its own signal sequence (e.g., SEQ ID NO: 10 or 12) or a different signal sequence, and a splice donor sequence (e.g., N-terminal intein (N-intane); SEQ ID NO: 14).

[0137] In one embodiment, a subcoding sequence encoding the C-terminal portion of a protein of interest (e.g., STRC) (including methionine corresponding to the start codon ATG and a potentially interchangeable signal sequence; optionally including a linker sequence and a Myc tag sequence), comprising a nucleic acid sequence having at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with the protein of interest (e.g., STRC; SEQ ID NO: 18, 20, 23, 24, 25, or 26). For example, a partial coding sequence encoding the C-terminal portion of a STRC protein may contain nucleic acid sequences that have at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with the following sequences.

[0138] The partial coding sequence encoding the C-terminal portion of the human STRC protein may be as follows (i.e., excluding the ATG start codon, signal sequence, or splice acceptor sequence): TIFF0007846626000063.tif197158TIFF0007846626000064.tif170158

[0139] The partial coding sequence encoding the C-terminal portion of the mouse STRC protein may be as follows: A partial coding sequence that encodes the C-terminal portion of a STRC protein, containing a nucleic acid sequence consisting of TIFF0007846626000065.tif46158, TIFF0007846626000066.tif241158, TIFF0007846626000067.tif76158, or SEQ ID NO: 57 or 58. Another embodiment may provide a C-terminal portion of the protein of interest (e.g., STRC; SEQ ID NO: 25 or 26) having an amino acid sequence that has at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with the protein of interest (e.g., STRC; SEQ ID NO: 25 or 26) (including methionine corresponding to the start codon ATG, a signal peptide sequence; optionally including a linker sequence and a Myc tag sequence). For example, the C-terminal portion of an STRC protein (containing methionine corresponding to the start codon ATG, an STRC signal peptide sequence; optionally containing a linker sequence and a Myc tag sequence) may contain an amino acid sequence having at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with the sequence of SEQ ID NO: 23 or 24, or the C-terminal portion of an STRC protein containing an amino acid sequence consisting of SEQ ID NO: 23 or 24. The second nucleotide sequence of the second vector may include a subcoding sequence that encodes the C-terminal portion of the protein of interest (e.g., STRC; SEQ ID NO: 23 or 24), which may include its own signal sequence (e.g., SEQ ID NO: 10 or 12) or a different signal sequence and splice acceptor sequence (e.g., C-terminal intein (C-intane); SEQ ID NO: 22).

[0140] One embodiment may provide an N-intane sequence comprising a nucleic acid sequence having at least 80% (e.g., 85%, 90%, 95%, 97%, 99%, 100%) identity with SEQ ID NO:42, or an N-intane sequence encoding an amino acid sequence having at least 80% (e.g., 85%, 90%, 95%, 97%, 99%, 100%) identity with SEQ ID NO:45. Another embodiment may relate to an N-intane sequence comprising the nucleic acid sequence of SEQ ID NO:42, or an N-intane sequence encoding an amino acid sequence comprising SEQ ID NO:45.

[0141] Another embodiment provides a C-intane sequence comprising a nucleic acid sequence having at least 80% (e.g., 85%, 90%, 95%, 97%, 99%, 100%) identity with SEQ ID NO:46, or a C-intane sequence encoding an amino acid sequence having at least 80% (e.g., 85%, 90%, 95%, 97%, 99%, 100%) identity with SEQ ID NO:49. Another embodiment may relate to a C-intane sequence comprising the nucleic acid sequence of SEQ ID NO:46, or a C-intane sequence encoding an amino acid sequence comprising SEQ ID NO:49.

[0142] In a further embodiment, the dual vector system of the present disclosure provides a first nucleotide sequence encoding the N-terminal portion of a protein of interest (e.g., STRC; SEQ ID NO: 15 or 16), wherein the first nucleotide sequence includes, but is not limited to, a signal sequence of the protein of interest, which may or may not form part of the partial coding sequence (5') of the N-terminal portion of the protein of interest (e.g., STRC), and a splice donor sequence (e.g., an N-intane sequence), which may or may not form part of the partial coding sequence of the N-terminal portion of the protein of interest (e.g., STRC). The nucleotide sequence may include, but is not limited to, an endogenous or exogenous signal sequence of the protein of interest, a partial coding sequence (5') of the N-terminal portion of the protein of interest, and a splice donor sequence (e.g., an N-intane sequence) (SEQ ID NO: 5 or 7). One embodiment can provide a first nucleotide sequence that includes a nucleic acid sequence having at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with SEQ ID NO: 5 or 7, or the first nucleotide sequence may include a nucleic acid sequence consisting of SEQ ID NO: 5 or 7.

[0143] Another embodiment provides a first nucleotide sequence encoding an amino acid sequence having at least 80% (e.g., 85%, 90%, 95%, 97%, 99%, 100%) identity with a sequence comprising the N-terminal portion of a protein of interest (e.g., STRC), wherein the first nucleotide sequence includes, but is not limited to, an endogenous or exogenous signal sequence of the protein of interest, which may or may not form part of the N-terminal partial coding sequence (5') of the protein of interest (e.g., STRC), and a splice donor sequence (e.g., an N-intane sequence), which may or may not form part of the N-terminal partial coding sequence of the protein of interest (e.g., STRC). A further embodiment may provide a first nucleotide sequence that codes for an amino acid sequence having at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with SEQ ID NO: 6 or 8, or the first nucleotide sequence may code for an amino acid sequence consisting of SEQ ID NO: 6 or 8.

[0144] Furthermore, a further embodiment of the dual vector system of the present disclosure provides a second nucleotide sequence encoding the C-terminal portion of a protein of interest (e.g., STRC; SEQ ID NO: 23 or 24), wherein the second nucleotide sequence includes, but is not limited to, an endogenous or exogenous signal sequence of the protein of interest, which may or may not form part of the partial coding sequence (3') of the C-terminal portion of the protein of interest (e.g., STRC), and a splice acceptor sequence (e.g., C-intane sequence), which may or may not form part of the partial coding sequence of the C-terminal portion of the protein of interest (e.g., STRC) (in some embodiments, the second nucleotide sequence includes a linker sequence and a Myc tag). The second nucleotide sequence may optionally include a sequence, and may have at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with the nucleotide sequence encoding the C-terminal portion of the protein of interest (e.g., STRC), and may include, but is not limited to, an endogenous or exogenous signal sequence of the protein of interest, a partial coding sequence (3') of the C-terminal portion of the protein of interest, and a splice acceptor sequence (e.g., a C-intane sequence) (optionally including a linker sequence and a Myc tag sequence) (SEQ ID NO: 17 or 19). One embodiment can provide a second nucleotide sequence that includes a nucleic acid sequence having at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with SEQ ID NO: 17 or 19, or the second nucleotide sequence may include a nucleic acid sequence consisting of SEQ ID NO: 17 or 19.

[0145] Another embodiment provides a second nucleotide sequence encoding an amino acid sequence having at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with a sequence comprising the C-terminal portion of a protein of interest (e.g., STRC), which includes, but is not limited to, a signal sequence, a C-intane sequence, and a partial coding sequence (3') of the C-terminal portion of the protein of interest (optionally including a linker sequence and a Myc tag sequence). A further embodiment may provide a second nucleotide sequence encoding an amino acid sequence having at least 5% (e.g., 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%) identity with SEQ ID NO: 18 or 20, or the second nucleotide sequence may encode an amino acid sequence consisting of SEQ ID NO: 18 or 20.

[0146] In a further embodiment, the dual vector system of the present disclosure provides a first nucleotide sequence (e.g., AAV vector 1) having an ITR, a promoter (e.g., CMV promoter), a subcoding sequence of interest (e.g., 5'Strc), a splice donor sequence (e.g., 5' intein), and an ITR in the 5' to 3' direction; and a second nucleotide sequence (e.g., AAV vector 2) having an ITR, a promoter (e.g., CMV promoter), a splice acceptor sequence (e.g., 3' intein), a subcoding sequence of interest (e.g., 3'Strc), and an ITR (Figure 31A). Upon translation, the first and second nucleotide sequences can generate multiple variants that, when spliced ​​using native cysteines at positions 747 and 970, i.e., Cys747 and Cys970 of SEQ ID NO:26 (or positions 709 and 934, i.e., Cys709 and Cys934 of SEQ ID NO:25), form the full-length STRC protein and cleaved inteins containing n-intane and c-intane. Figure 31A illustrates eight AAV2 plasmids containing four different dual-vector variants.

[0147] For example, variant 1 includes a first protein sequence comprising a signal peptide sequence, the N-terminal portion of the protein of interest (e.g., an approximately 80 kD n-STRC), a splice site (e.g., Ser746), and a splice donor sequence (n-intane), in the direction from the N-terminus to the C-terminus; and a second protein sequence comprising a splice acceptor sequence (c-intane), a splice site (e.g., Cys747), and the C-terminal portion of the protein of interest (e.g., an approximately 117 kD c-STRC), in the direction from the N-terminus to the C-terminus. Variant 2 includes, for example, a first protein sequence comprising a signal peptide sequence, an N-terminal portion of the protein of interest (e.g., an n-STRC of approximately 105 kD), a splice site (e.g., Ala969), and a splice donor sequence (n-intane) in the direction from the N-terminus to the C-terminus; and a second protein sequence comprising a splice acceptor sequence (c-intane), a splice site (e.g., Cys970), and a C-terminal portion of the protein of interest (e.g., a c-STRC of approximately 92 kD) in the direction from the N-terminus to the C-terminus. For example, variant 3 includes a first protein sequence comprising a signal peptide sequence, the N-terminal portion of the protein of interest (e.g., an approximately 80 kD n-STRC), a splice site (e.g., Ser746), and a splice donor sequence (n-intane), in the direction from the N-terminus to the C-terminus; and a second protein sequence comprising a signal peptide sequence, a splice acceptor sequence (c-intane), a splice site (e.g., Cys747), and the C-terminal portion of the protein of interest (e.g., an approximately 117 kD c-STRC), in the direction from the N-terminus to the C-terminus.Variant 4 includes, for example, a first protein sequence comprising a signal peptide sequence, an N-terminal portion of the protein of interest (e.g., an n-STRC of approximately 105 kD), a splice site (e.g., Ala969), and a splice donor sequence (n-intane) in the direction from the N-terminus to the C-terminus; and a second protein sequence comprising a signal peptide sequence, a splice acceptor sequence (c-intane), a splice site (e.g., Cys970), and a C-terminal portion of the protein of interest (e.g., a c-STRC of approximately 92 kD) in the direction from the N-terminus to the C-terminus.

[0148] Protein splicing of the translated variant sequence forms the full-length protein of interest (e.g., STRC). For example, the N-terminal portion of the protein of interest (n-STRC) can be ligated to the C-terminal portion of the protein of interest (e.g., c-STRC) in the direction from the N-terminus to the C-terminus. Splicing of the N-terminus and C-terminus results in the excision of splice donor and splice acceptor sequences. For example, the excised splice sequences can form the full-length splice sequence (e.g., excised inteins of n-intane and c-intane) (Figure 31A).

[0149] Figure 31C demonstrates that HEK cells transfected with both the N-terminal and C-terminal portions (c+n) of variant 3 resulted in the expression of full-length STRC, in contrast to variant 3 containing only the C-terminal portion (c), or variant 1 containing only the C-terminal portion (c) or both the C-terminal and N-terminal portions (c+n). The only difference between variant 1 and variant 3 is the presence of a signaling sequence in the C-terminal portion. Thus, Figure 31C demonstrates that the signaling sequences in both the N-terminal and C-terminal portions direct each portion to the same cellular compartment of the cell, thereby enabling the formation of full-length STRC by protein splicing.

[0150] Further embodiments provide cells (e.g., host cells, mammalian cells, human cells, bacterial cells) containing the dual vector system described herein, comprising a first vector and a second vector. In one embodiment, the cells may be inner ear cells, inner hair cells, or outer hair cells. Some embodiments may relate to cells that are mammalian cells (e.g., human, dog, cat, horse, mouse). Other embodiments may provide ear cells (e.g., inner ear cells, outer ear cells, inner hair cells, outer hair cells). The cells of this disclosure may be in vivo or in vitro. In some embodiments, the cells may be transfected or transformed by the first and second vectors of the dual vector system of this disclosure using any of a number of known transfection and transformation techniques generally known in the art. For example, see Graham et al. (1973) Virology, 52:456, Sambrook et al. (1989) Molecular Cloning, a laboratory manual, Cold Spring Harbor Laboratories, New York, Davis et al. (1986) Basic Methods in Molecular Biology, Elsevier, and Chu et al. (1981) Gene 13:197, all of which are incorporated herein in their entirety by reference. The first and second vectors of the dual vector system of this disclosure may be inserted into the cells described herein by any means, including, but not limited to, viral transduction, bacterial transformation using calcium chloride, bacterial transformation or transduction by bacterial mating or conjugation, transfection (e.g., electroporation, calcium phosphate, liposome-based transfection), gene guns, etc.

[0151] Another embodiment may relate to a composition or pharmaceutical composition comprising the dual vector system described herein and a pharmaceutically or physiologically acceptable medium (e.g., diluent, carrier, excipient). Compositions intended herein for the treatment of mutation-related diseases or conditions may include the dual vector system described herein. For therapeutic purposes, a composition comprising a polynucleotide of interest (e.g., STRC) or a fragment thereof, which, when appropriately processed according to the methods disclosed herein, results in the expression of a full-length protein of interest (e.g., STRC protein) in a genome containing mutations that cause or contribute to diseases or conditions described herein (e.g., autosomal recessive DFNB16 deafness), may be administered directly to a body area affected by the disease or condition (e.g., cochlea, inner ear). In some embodiments, the composition is formulated in a pharmaceutically acceptable buffer, e.g., saline. Non-limiting methods of administration include injection into the ear, inner ear, cochlear duct, or the perilymph-filled space around the cochlear duct (e.g., tympanic scala and vestibular scala). Injecting into the cochlear duct, which is filled with high-potassium endolymph, may provide direct access to hair cells. However, altering this delicate fluid environment may disrupt the cochlear potential and increase the risk of injection-related toxicity. The perilymphatic space surrounding the cochlear duct, the scala tympanicum and scala vestibular, can be accessed from the middle ear through either the oval or round window membrane. The round window membrane, the only non-bony opening to the inner ear, is relatively easily accessible in many animal models, and administration of viral vectors using this route is well tolerated. In humans, cochlear implantation routinely relies on surgical electrode insertion through the round window membrane.

[0152] How to use One embodiment may provide a method using a vector system described herein (e.g., a dual vector system; a capsid, a plasmid, a transpricing plasmid, a viral vector, an adenovirus, an AAV, an AAV genome, a lentivirus) that can treat and / or reduce and / or prevent a disease, condition, or symptom thereof resulting from a gene defect or mutation.

[0153] Another embodiment may involve a method comprising contacting cells (for example, of interest) with a composition comprising a vector system described herein (e.g., a dual vector system) and a pharmaceutically or physiologically acceptable medium (e.g., a carrier, diluent, excipient). The contact step with cells (e.g., of interest) may result in the delivery of a nucleotide sequence of a vector (e.g., a plasmid, transpricing plasmid, viral vector, adenovirus, AAV, AAV genome) comprising, for example, SEQ ID NO: 33 (Figures 2A-2C) or SEQ ID NO: 38 (Figures 4A-4D), where the cells may express a full-length protein of interest (e.g., STRC; human SEQ ID NO: 2 or 25; mouse SEQ ID NO: 4 or 26). The contact step with the target cell can result in the delivery of the first and second nucleotide sequences (of the first and second vectors, respectively), where the cell can express the N-terminal portion (e.g., STRC) and the C-terminal portion of the protein of interest, which are linked by peptide bonds to form the full-length protein of interest (e.g., STRC; human SEQ ID NO: 2 or 25; mouse SEQ ID NO: 4 or 26).

[0154] One embodiment is a vector system described herein (e.g., a dual vector system), or a composition described herein, or a cell containing a vector system of the disclosure (e.g., a dual vector system), or a composition of the disclosure, or a vector or composition comprising at least one nucleotide sequence (e.g., STRC;SEQ ID NO:5;SEQ ID NO:7;SEQ ID NO:17;SEQ ID NO:19;SEQ ID NO:30;SEQ ID NO:32) encoding a protein of interest (e.g., STRC;SEQ ID NO:25;SEQ ID NO:26), or the protein of interest (e.g., STRC;SEQ ID NO:25;SEQ ID NO:26) itself, or a portion of the protein of interest (e.g., SEQ ID NO:6~SEQ ID NO:16;SEQ ID NO:18~SEQ ID A method can be provided for treating autosomal recessive hearing loss in a subject, comprising administering an effective dose of NO:24) to a subject in need thereof, wherein the administration may result in a reduction or recovery of autosomal recessive hearing loss or its symptoms. In some embodiments, a subject in need thereof would be considered successfully treated if, after treatment, the subject has a hearing level of 69 dB or less (e.g., 60 dB, 55 dB, 50 dB, 45 dB, 40 dB, 35 dB, 30 dB, 26 dB, 25 dB, 20 dB, 15 dB, 10 dB, 5 dB, 0 dB). Generally, subjects with profound hearing loss cannot hear sounds below 95 dB; subjects with severe hearing loss cannot hear sounds between 70 dB and 94 dB; subjects with moderate hearing loss cannot hear sounds between 40 dB and 69 dB; and subjects with mild hearing loss cannot hear sounds between 26 dB and 40 dB. However, subjects suffering from hearing loss, such as autosomal recessive hearing loss, may be treated by any of the methods described herein, thereby resulting in hearing loss or reduction of its symptoms, and / or restoration or improvement of hearing (or auditory function in the subject), and / or maintenance of hearing. Subjects with normal hearing may be characterized by the ability to hear sounds below 25 dB (e.g., 20 dB, 15 dB, 10 dB, 5 dB, 0 dB).In some embodiments, autosomal recessive deafness is DFNB16.

[0155] Further embodiments may provide a method for treating and / or preventing a pathology or disease characterized by hearing loss, comprising administering an effective amount of the dual vector system described herein, the cells according to this disclosure, or the composition or pharmaceutical composition described herein to a subject in need. The cells according to this disclosure may be in vivo, in vitro inner ear cells, inner hair cells, outer hair cells, etc., or any combination thereof.

[0156] Polynucleotide delivery The success of these approaches largely depends on the safe and efficient delivery of exogenous gene constructs to relevant therapeutic cell targets within the organ of Corti in the cochlea. The organ of Corti contains two types of sensory hair cells: inner hair cells, which convert mechanical information carried by sound into electrical signals transmitted to neuronal structures, and outer hair cells, which help amplify and modulate the cochlear response, a process required for complex auditory function.

[0157] Methods for delivering nucleic acids to cells are generally known in the art, and a method for delivering a transgene-containing virus (also called a viral particle) to inner ear cells in vivo is described herein. As described herein, about 10 8 ~about 10 12 Individual virus particles can be administered to a target, and the virus can be suspended in an appropriate volume (e.g., 10 μL, 50 μL, 100 μL, 500 μL, or 1000 μL) of, for example, extraarterial lymphatic fluid.

[0158] Viruses described herein, comprising terminal inversion repeats (ITRs), promoters (e.g., Espin promoter, PCDH15 promoter, PTPRQ promoter, Myo6 promoter, KCNQ4 promoter, Myo7a promoter, synapsin promoter, GFAP promoter, CMV promoter, CAG promoter, CBH promoter, CBA promoter, U6 promoter, and TMHS(LHFPL5) promoter), signal sequences, polynucleotides encoding a protein of interest (e.g., STRC protein), and polyadenylation (Poly-A) sequences, and in some embodiments, containing linker sequences for ligating a c-myc tag, can be delivered to inner ear cells (e.g., cells in the cochlea) by a number of means. For example, a composition comprising viral particles containing the dual-vector intein-mediated protein trans-splicing system described herein can be injected in therapeutically effective amounts through a round window, oval window, or utricle, typically by a relatively simple (e.g., expat) procedure. In some embodiments, compositions containing a therapeutically effective number of viral particles, including a dual vector intein-mediated protein trans-splicing system (e.g., a dual AAV intein-mediated STRC protein system) as described herein, or one or more sets of different viral particles, can be delivered to an appropriate location within the ear during surgery (e.g., cochlear fenestration or canalostomy).

[0159] Furthermore, delivery media (e.g., polymers) that facilitate the transfer of drugs through the tympanic membrane and / or through the round window or utricle are available, and any such delivery media may be used to deliver the viruses described herein. For delivery media, see, for example, Arnold et al., 2005, Audio.Neurootol., 10:53-63, which is incorporated herein in its entirety by reference.

[0160] The compositions and methods described herein enable highly efficient delivery of nucleic acids to inner ear cells, e.g., cochlear cells. For example, a polynucleotide encoding a protein of interest (e.g., the STRC protein) or a fragment thereof can be cloned into a viral vector, and expression may be driven from its endogenous promoter, from viral terminal inverted repeats, or from a promoter specific to the target cell type of interest. Other viral vectors that may be used include, for example, vaccinia virus, bovine papillomavirus, or herpesvirus, e.g., Epstein-Barr virus. Viral vectors are used in clinical settings. In some embodiments, a viral vector (e.g., rAAV) may be used to deliver a large (e.g., Strc gene) polynucleotide in fragment form. In some embodiments, a viral vector may be used to deliver a Strc polynucleotide fragment to a specific area of ​​the body.

[0161] For example, the compositions and methods described herein enable the delivery and expression of a polynucleotide of interest (e.g., Strc) to at least 65% (e.g., 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) of inner hair cells and / or outer hair cells, or to at least 65% (e.g., 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) of outer hair cells.

[0162] Expression of STRC polynucleotides delivered using the dual-vector intein-mediated system described herein may result in improved structure and function of inner and outer hair cells, thereby restoring hearing over a long period (e.g., days, weeks, months, years, decades, or a lifetime). In one embodiment, hearing loss may be restored in subjects suffering from DFNB16, an autosomal recessive form of non-symptomatic hearing loss caused by mutations in the STRC gene. Normal expression of STRC and stereocillin (STRC) proteins is essential for auditory function.

[0163] As described herein, adeno-associated viruses (AAVs) are particularly efficient for delivering nucleic acids (e.g., polynucleotides encoding STRC polypeptides) to inner ear cells. The Anc80 vector is an example of an AAV targeting inner ear hair cells, favorably transducing more than 60% (e.g., 70%, 80%, 90%, 95%, 100%) of inner or outer hair cells. One embodiment may utilize an ancestral capsid protein belonging to the class of Anc80 ancestral capsid proteins, e.g., Anc80-0065, described in the international publication number WO2018 / 145111 (PCT / US2018 / 017104) relating to Anc80, which is incorporated herein by reference in its entirety. WO2015 / 054653, which is incorporated in its entirety herein by reference, describes a number of additional ancestral capsid proteins belonging to the class of Anc80 ancestral capsid proteins.

[0164] In specific embodiments, the adeno-associated virus (AAV) contains an ancestral AAV capsid protein having native or engineered directivity to hair cells. In some embodiments, the virus is an inner ear hair cell-targeting AAV that delivers a polynucleotide of interest encoding a polypeptide of interest (e.g., STRC protein) to the inner ear of a subject (e.g., a subject suffering from mutations in the DFNB16 and / or STRC gene). In some embodiments, the virus is an AAV containing a purified capsid polypeptide. In some embodiments, the virus is artificial. In some embodiments, the virus is an AAV having a lower seroprevalence than AAV2. In some embodiments, the virus is an exome-associated AAV. In some embodiments, the virus is an exome-associated AAV1. In some embodiments, the virus contains a capsid protein having at least 95% amino acid sequence identity or homology with the Anc80 capsid protein.

[0165] The expression of a polynucleotide of interest (e.g., STRC) may be directed by a heterologous promoter (e.g., CMV promoter, Espin promoter, PCDH15 promoter, PTPRQ promoter, TMHS(LHFPL5) promoter). As used herein, “heterologous promoter” means a promoter that does not naturally direct the expression of its sequence (i.e., is not naturally found in its sequence).

[0166] For example, methods for packaging transgenes into viruses containing the Anc80 capsid protein are known in the art and utilize conventional molecular biology and recombinant nucleic acid techniques. In one embodiment, a construct is provided comprising a nucleic acid sequence encoding the Anc80 capsid protein, as well as a construct holding a fragment of polynucleotide encoding the N-terminal and C-terminal portions of the STRC protein adjacent to a suitable terminal inversion repeat (ITR), which enables packaging into the Anc80 capsid protein.

[0167] The polynucleotide of interest (e.g., STRC) can be packaged into an AAV containing the Anc80 capsid protein, for example, using a packaging host cell. Components of the viral particle (e.g., rep sequence, cap sequence, terminal inversion repeat (ITR) sequence) can be introduced transiently or stably into a packaging host cell using one or more constructs described herein. The polynucleotide of interest may generally be a large gene (e.g., 4kB or larger) that may require splitting and packaging into multiple AAVs.

[0168] In some embodiments, AAVs containing the AAV9-php.b vector can be used to efficiently target inner ear cells. AAV9-php.b is described in international publication number WO2019 / 173367 (PCT / US2019 / 020794), the contents of which are incorporated herein by reference in their entirety. AAV-PHP.B encodes the 7-residue sequence TLAVPFK (SEQ ID NO: 59) and efficiently delivers the transgene to the cochlea, where it exhibits remarkably specific and robust expression in inner and outer hair cells. The AAV-PHP.B vector may contain, but is not limited to, any of the promoters described herein.

[0169] cDNA expression for use in polynucleotide therapies may be directed by any suitable promoter (e.g., human cytomegalovirus (CMV) promoter, monkey virus 40 (SV40) promoter, or metallothionein promoter) and controlled by any suitable mammalian regulatory element. For example, enhancers known to preferentially direct gene expression in a particular cell type may be used to direct nucleic acid expression, if desired. Enhancers used may include, but are not limited to, those characterized as tissue-specific or cell-specific enhancers. Alternatively, if a genomic clone is used as a therapeutic construct, control may be mediated by a congeneral regulatory sequence, or, if desired, by a regulatory sequence of heterologous origin, such as any of the aforementioned promoters or regulatory elements.

[0170] treatment method Another therapeutic approach included in this disclosure may involve administering recombinant therapeutic agents (e.g., recombinant STRC proteins, variants, or fragments thereof) directly or systemically to sites of potentially diseased or actually diseased tissue (e.g., by any conventional recombinant protein delivery technology). The dosage of the protein administered depends on a number of factors, including the size and health status of the individual patient. For any particular subject, a specific drug regimen should be adjusted over time according to individual needs and the professional judgment of the person administering or supervising the administration of the composition.

[0171] Some aspects of the present disclosure may provide a method for treating, preventing, or reducing a disease and / or disorder or its symptoms in a subject requiring it, comprising administering a therapeutically effective amount of a pharmaceutical composition comprising a vector system (e.g., a dual-vector intein-mediated system) containing a nucleotide sequence encoding a full-length protein of interest (e.g., STRC) to cells (e.g., of interest), wherein the cell genome may include mutations. In embodiments of the present disclosure, a method for treating, preventing, or reducing a disease and / or disorder or its symptoms in a subject requiring it may involve administering a therapeutically effective amount of a pharmaceutical composition comprising a dual-vector intein-mediated system containing a first nucleotide sequence encoding a portion of the protein of interest (e.g., N-STRC) and a second nucleotide sequence encoding the remainder of the protein of interest (e.g., C-STRC) to a mutant genome of a subject requiring it (e.g., a mammal, e.g., human). Thus, one aspect is a method for treating a subject suffering from or susceptible to a mutation-related disease or disorder or its symptoms. The method involves administering a composition according to this specification to a subject in a therapeutic amount sufficient to treat, prevent, or reduce a disease, disorder, or symptom. In some embodiments, the mutation is a recessive mutation.

[0172] The therapeutic methods of the present invention (including prophylactic treatment) generally involve administering a therapeutically effective amount of a compound or composition as provided herein, for example, a compound of a formulation as provided herein, to a subject in need (e.g., an animal, a human), for example, a mammal, for example, a human. Such treatment would be appropriately administered to a subject, specifically a human, who is suffering from, has, is susceptible to, or is at risk of suffering from, a disease, a disorder, or symptoms thereof. The determination of a subject "at risk" may be made by diagnostic tests or any objective or subjective determination by the subject or healthcare provider (e.g., genetic testing, enzyme or protein markers, markers (as defined herein), family history, etc.).

[0173] Treatment of human patients or non-human animals may be carried out using therapeutically effective amounts of combination therapeutic agents in a physiologically acceptable carrier. The term "pharmaceutically acceptable" means, within the bounds of reasonable medical judgment, that the compounds of this disclosure, compositions containing such compounds, and / or dosage forms are suitable for use in contact with human and animal tissues without excessive toxicity, irritation, allergic reactions, or other problems or complications, and that the benefit / risk ratio is reasonable.

[0174] The compositions may be conveniently presented in unit dosage forms and may be prepared by any method well known in the art. The amount of composition that can be combined with a carrier material to make a single dosage form (e.g., a vector containing a sequence encoding the N-terminal or C-terminal portion of the protein of interest (e.g., STRC)) will vary depending on the host being treated and the specific mode of administration. The amount of composition that can be combined with a carrier material to make a single dosage form will generally be the amount of composition that produces the therapeutic effect. Generally, this amount will be in the range of 1% to 99% of the composition (e.g., 5% to 70%, 10% to 30%) out of 100%.

[0175] In some cases, the composition may be administered in doses that control the clinical or physiological symptoms of a disease or condition, which may be determined by diagnostic methods known to those skilled in the art.

[0176] Therapeutic compositions and therapeutic combinations are administered in effective doses. For example, about 10 8 ~about 10 12 Individual virus particles can be administered to a target, and the virus can be suspended in an appropriate volume (e.g., 10 μL, 50 μL, 100 μL, 500 μL, or 1000 μL), for example, extraartificial lymphatic fluid.

[0177] Treatment methods for autosomal recessive hearing loss Compositions and methods for treating autosomal recessive hearing loss (e.g., autosomal recessive non-symptomatic hearing loss 16 (DFNB16)) are provided.

[0178] In short, DFNB16 is associated with mutations in the STRC gene in affected individuals. Normal expression of STRC, which encodes the extracellular structural protein stereocillin (STRC) in the inner ear, is essential for auditory function. To induce the restoration of hearing loss, the wild-type STRC gene is administered to subjects using the vector systems described herein (e.g., a dual AAV intein-mediated STRC protein trans-splicing system), i.e., by packaging the wild-type STRC gene sequence or a fragment thereof so that full-length mRNA and full-length STRC protein are expressed.

[0179] In some embodiments, a vector system encoding the STRC protein, such as a dual AAV intein-mediated STRC protein trans-splicing vector, may be administered to a subject with DFNB16 deafness by direct injection into the cochlea of ​​the subject, containing at least one vector encoding the stereocillin (STRC) protein (e.g., a 5'-STRC vector and a 3'-STRC vector). In some embodiments, one vector encodes only the N-STRC protein, and another vector encodes only the C-STRC protein. For therapeutic purposes, an STRC polypeptide, or a composition comprising an STRC polynucleotide encoding an STRC polypeptide, may be administered directly to a region of the body affected by the disease or condition (e.g., the cochlea), wherein the subject's genome contains an STRC mutation that causes or contributes to the deafness described herein (e.g., DFNB16).

[0180] One embodiment may provide a method for treating autosomal recessive deafness in a subject, comprising administering an effective amount of a pharmaceutical composition comprising a vector system described herein (e.g., a dual vector system); cells containing a vector system described herein (e.g., a dual vector system); or a pharmaceutically acceptable medium to a subject requiring such treatment. Some embodiments may be methods for treating autosomal recessive deafness, which is DFNB16.

[0181] Another aspect of the present disclosure provides a method comprising contacting a cell of interest with a composition comprising a vector system described herein (e.g., a dual vector system) and a pharmaceutically acceptable medium, wherein the contact results in the delivery to the cell of a first nucleotide sequence expressing the N-terminal portion of a protein and a second nucleotide sequence expressing the C-terminal portion of a protein, the cell expressing the N-terminal and C-terminal portions of the protein linked by peptide bonds to form a full-length protein.

[0182] Further aspects of this disclosure provide a method for treating and / or preventing a pathology or disease characterized by hearing loss, comprising administering an effective amount of a pharmaceutical composition comprising a vector system described herein (e.g., a dual vector system); cells containing a vector system described herein (e.g., a dual vector system); or a vector system described herein (e.g., a dual vector system) and a pharmaceutically acceptable medium to a subject in need thereof, wherein the administration step is carried out in at least one cell of the subject (e.g., inner ear cells, inner hair cells, outer hair cells). In one embodiment, the method of contacting cells with an effective amount of a pharmaceutical composition comprising a vector system described herein (e.g., a dual vector system); cells containing a vector system described herein (e.g., a dual vector system); or a vector system described herein (e.g., a dual vector system) and a pharmaceutically acceptable medium, or administering them to cells, is carried out in vivo, ex vivo, and / or in vitro. Another embodiment provides a method of improving or restoring auditory function in a subject, using any of the methods described herein.

[0183] Non-restrictive administration methods may include injection into the cochlear duct or the perilymphatic space surrounding the cochlear duct (e.g., the tympanic scala and vestibular scala). Injection into the cochlear duct, which is filled with high-potassium endolymph, may provide direct access to hair cells. However, altering this delicate fluid environment may interfere with the cochlear potential and increase the risk of injection-related toxicity. The perilymphatic space surrounding the cochlear duct, the tympanic scala and vestibular scala, can be accessed from the middle ear through either the oval or round window membrane. The round window membrane, the only non-bony opening to the inner ear, is relatively easily accessible in many animal models, and administration of viral vectors using this route is well-tolerated. In humans, cochlear implantation routinely relies on surgical electrode insertion through the round window membrane.

[0184] In some embodiments, the expression of the protein of interest (e.g., wild-type STRC) may restore auditory function in the subject. In some embodiments, the restored auditory function may be 10% or more (e.g., 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%) or less than 100% (e.g., 95%, 85%, 75%, 65%, 55%, 45%, 35%, 25%, 15%). [Examples]

[0185] The following embodiments illustrate specific aspects of the description of the present invention. These embodiments are intended only to provide a concrete understanding and implementation of the embodiments and their various aspects, and should not be construed as limiting.

[0186] Example 1: A dual-vector system for transducing large genes in vitro. AAV vector generation The dual vector or dual trans-splicing systems for the delivery of the protein of interest disclosed herein were demonstrated in the inner ear, for example, in two independent adeno-associated virus serotype 2 (AAV2) / adeno-associated virus serotype 9 (AAV9)-Php.B vectors, using the STRC coding sequence. A first AAV genome was constructed having a splice donor sequence (e.g., N-intane) located immediately after the first 2,247 base pairs of the STRC coding sequence (i.e., the N-terminal portion of STRC) followed by the 3' ITR (e.g., AAV2 / AAV9-Php.B-STRC-trans / donor). A second AAV genome was constructed having a corresponding splice acceptor sequence located after the 5' ITR and immediately before the rest of the STRC coding sequence (i.e., the C-terminal portion of STRC) containing the C-terminal myc tag (e.g., AAV2 / AAV9-Php.B-STRC-trans / acceptor). Since the AAV2 genome is known to form concatemers through terminal inversion repeats (ITRs) at each end of the viral genome, the STRC coding sequence was split into two fragments and packaged into separate AAV capsids (e.g., synthetic AAV:Anc80, which has been shown to efficiently transduce inner and outer hair cells).

[0187] To confirm intein-mediated trans-splicing and processing, HEK293 cells infected with the AAV2 / AAV9-Php.B vector encoding a portion of full-length STRC (5 days after infection) were analyzed by Western blotting (Figure 24). Specifically, Figure 24 shows the following: Lane 1: control or untransfected HEK293 T cells; Lane 2: full-length stereocillin (STRC) (196.4 kDa); Lane 3: pCFS-#2, C portion (vector containing construct #2 having only the C portion (116.7 kDa)); Lane 4: pCFS-#2, N+C portion (vector containing construct #2 having both the N and C portions); Lane 5: pCFS-#1, C portion (vector containing construct #1 having only the C portion (91.6 kDa)); Lane 6: pCFS-#1, N+C portion (vector containing construct #1 having both the N and C portions). The arrow on the right side of the Western blot points to the full-length STRC protein in lane 4, demonstrating that both the N-terminal and C-terminal portions of STRC were formed when two AAV2 / AAV9-Php.B vectors, each containing sequences encoding the N-terminal and C-terminal portions of STRC, were transfected into HEK293 cells.

[0188] Further Western blotting was performed to confirm the usefulness of the signal sequence. Figure 25 shows the following: Lane 1: Full-length stereocillin (STRC) (196.4 kDa); Lane 2: Control or untransfected HEK293 T cells; Lane 3: pCFS-#2, C portion, (-) signal (vector containing construct #2 having only the C portion (116.7 kDa) and no signal sequence); Lane 4: pCFS-#2, N+C portion, (-) signal (vector containing construct #2 having both the N and C portions and no signal sequence); Lane 5: pCFS-#2, C portion, (+) signal (vector containing construct #2 having only the C portion (91.6 kDa) containing the signal sequence); Lane 6: pCFS-#2, N+C portion, (+) signal (vector containing construct #2 having both the N and C portions containing the signal sequence). The arrow on the right side of the Western blot points to the full-length STRC protein in lane 6, demonstrating that a signal sequence was necessary for the formation of the full-length STRC protein.

[0189] The AAV vector was prepared by Boston Children's Hospital Viral Core (Boston, MA, USA). The plasmid containing STRC and intein was sequenced prior to packaging into AAV9-php.b-cmv (MGH DNA Core, fully plasmid sequenced). The vector titer, determined by qPCR specific to the terminal inversion repeat (AAV2) of the virus, was 4.8 × 10⁶. 14 The gc / ml result was positive.

[0190] Example 2: Analysis of a dual-vector system in in vivo Strc knockout mice animal All animals were housed and maintained in the facility. All studies involving animals were approved by the HMS Standing Committee on Animals (protocol number 03524) as well as the Boston Children's Hospital Institutional Animal Care and Use Committee (protocol numbers 2878 and 3396). All experiments were conducted in accordance with the animal protocols.

[0191] Null allele ("knockout") mice lacking stereocilin (STRC - / - ;Strc Δ / Δ ) were generated and used as a mouse model of the human DFNB16 phenotype of deafness caused by STRC mutations that result in the absence or non-functional stereocilin protein. Strc homozygous mutant mice (STRC#16 homozygous) showed profound deafness by 4 weeks of age as determined by auditory brainstem response (ABR), and by 6 weeks of age, the mutant mice were completely deaf. Strc homozygous mutant mice also lacked detectable distortion product otoacoustic emissions (DPOAE) up to 80 dB sound pressure level, which reflects the absence of normal outer hair cell (OHC) function.

[0192] Strc Δ / Δ mice were generated and characterized in Figures 32A - 32G. The wild-type (WT) protein of interest, for example, Strc, was disrupted using a CRISPR / Cas9 strategy by designing three guide RNAs (sgRNAs) targeting exon 4 of the Strc gene. The disruption resulted in a 249 nucleotide deletion (positions 1509 - 1758) as well as two translocations and inversions (positions 947 - 1139 near the 3' end; positions 1758 - 1835 near the 5' end).

[0193] Inner ear injection Strc - / - or Strc WT / WTOn postnatal day 1 (P1), 1 μl of AAV9-php.b-cmv-STRC intein virus was injected into the inner ear of pup mice at a rate of 60 nl / min. The pup mice were anesthetized using a 2-3 minute cold exposure in ice water. During anesthesia, a postauricular incision was made to expose the otic bulla and visualize the cochlea. The injection was administered manually using a glass micropipette. After injection, the skin incision was closed with sutures. The injected mice were then placed on a 42°C heating pad for recovery. The pup mice were returned to their mothers after complete recovery within approximately 10 minutes. Standard postoperative care was applied postoperatively. The sample size for the in vivo study was continuously determined to optimize sample size and reduce variance. In P5-P7, the organ of Corti was excised from the injected ear. The organ of Corti tissue was incubated at 37°C in 5% CO2 for 8-10 days, and the capsular membrane was removed immediately before electrophysiological recording.

[0194] Hearing test To determine whether a dual-vector system using AAV vectors containing sequences encoding N-STRC with N-intane and C-STRC with C-intane can restore hearing loss, and to what extent, auditory brainstem response (ABR) and distortion component otoacoustic emissions (DPOAE) were measured using stereocylin knockout (STRC - / -Measurements were taken in mice. ABR and DPOAE measurements were recorded using the EPL Acoustic system (Massachusetts Eye and Ear, Boston). Acoustic stimuli were generated by a 24-bit digital input / output card (National Instruments PXI-4461) in a PXI-1042Q chassis, amplified by an SA-1 speaker driver (Tucker-Davis Technologies, Inc.), and delivered by two electrostatic drivers (CUI CDMG15008-03A) in a custom acoustic system. External auditory canal sound pressure was monitored using an electret microphone (Knowles FG-23329-P07) at the end of a small probe tube. ABR and DPOAE were recorded from mice during the same session. ABR signals were collected using subcutaneous needle electrodes inserted in the auricle (active electrode), vertebral head (reference electrode), and buttocks (ground electrode). ABR potentials were amplified (10,000×), passed through a pass filter (0.3–10kHz), and digitized using custom data acquisition software (LabVIEW) from the Eaton-Peabody Laboratories Cochlear Function Test Suite. Acoustic stimuli and electrode potentials were sampled at 40μs intervals using a digital I / O board (National Instruments) and stored for offline analysis. The threshold was visually defined as the lowest decibel level at which peak 1 could be detected and reproduced by increasing sound intensity. ABR thresholds were averaged within each experimental group and used for statistical analysis. ABR and DPOAE measurements were performed by researchers who were unaware of the researchers' genotypes.

[0195] Mice were anesthetized by intraperitoneal (ip) injection of xylazine (5-10 mg / kg) and ketamine (60-100 mg / kg), and the base of the auricle was excised to expose the external auditory canal. Three subcutaneous needle electrodes were inserted into the skin at (a) the dorsal side between the two ears (reference electrode); (b) behind the left auricle (recording electrode); and (c) the dorsal side of the animal's rump (ground electrode). Additional aliquots of ketamine (60-100 mg / kg ip) were administered throughout the session to maintain anesthesia as needed. Prior to the ABR test, sound pressure at the entrance of the external auditory canal was calibrated for each individual test subject at all stimulus frequencies. ABR and DPOAE data were collected under the same conditions and during the same recording session.

[0196] DPOAE was recorded first. To generate DPOAE at 2f1-f2, the primary tone was created with a frequency ratio of 1.2 (frequency ratio of the primary tone f1 to f2 (f2 / f1=1.2)), and for each f2 / f1 pair, the f2 level was 10 dB lower in sound pressure level than the f1 level. The tone was presented by f2 and L1-L2 = 10 decibels in sound pressure level (dB SPL) varying between 5.6 and 32.0 kHz in half-octave steps. At each f2, L2 varied by 10 dB between 10 dB and 80 dB. The DPOAE threshold was defined as the L2 level that induces a DPOAE magnitude 5 dB higher than the noise floor, based on the average spectrum. The average noise floor level was less than 0 dB at all frequencies. At each level, waveform and spectral averaging was used to increase the signal-to-noise (S / N) ratio of the recorded auditory canal sound pressure. The DPOAE at 2f1–f2 had amplitudes extracted from the averaged spectrum and a noise floor at adjacent points in the spectrum. Interpolation from the plot of DPOAE amplitudes against sound levels yielded an iso-response curve. The threshold was defined as the f2 level required to generate a DPOAE above 0 dB.

[0197] Next, ABR experiments were conducted in a soundproof chamber at 32°C. To test auditory function, mice were presented with broadband "click" sounds and pure tones ranging from 5.6 to 32.0 kHz in half-octave increments, all presented as 5 ms tone pips. Responses were amplified (10,000x), filtered (0.1 to 3 kHz), and averaged using an analog-to-digital board on a PC-based data acquisition system (EPL, Cochlear function test suite, MEE, Boston). In various tests, the sound level was increased in 5-10 dB increments from 0 to 110 dB SPL. 512 responses were collected at each level, and after "artifact removal," they were averaged for each sound pressure level (by alternating stimulus polarity). The threshold was determined by visual inspection of the appearance of peak 1 against background noise. The data were analyzed and plotted using Origin-2015 (OriginLab Corporation, MA). Unless otherwise specified, the mean ± standard deviation of the threshold is presented. Most of these experiments were not conducted under blinded conditions.

[0198] Knockout mice lacking STRC were generated by disrupting the STRC gene coding sequence through Cas9-mediated cleavage via NHEJ. This resulted in a deletion of approximately 200 base pairs within exon 4 of the STRC gene, disrupting the synthesis of the functional protein. This mouse model was found to accurately reproduce the human deafness of the DFNB16 phenotype caused by STRC mutations resulting in the absence or non-functionality of stereocillin proteins. Four-week-old STRC homozygous mutant mice exhibited severe deafness based on auditory brainstem response (ABR), and by six weeks, these mice were completely deaf. The ABR threshold may be the lowest level at which a clear response (CR) is present. Distortion component otoacoustic emissions (DPOAEs) reflect outer hair cell integrity and cochlear function. These STRC homozygous mutant mice possessed detectable DPOAEs up to 80 decibel (dB) sound pressure levels, which reflects the absence of normal outer hair cell (OHC) function.

[0199] Figure 3 shows the sound pressure levels from the ABR waveform results. The waveforms in the center and on the left are from a single STRC-KO mouse, STRC - / - The waveform shows mice (Strc#16 homozygous) (n=5), and mice (n=9) injected with AAV vectors containing a signal sequence, N-intine or C-intine, and sequences encoding the N-terminal or C-terminal portion of the STRC protein, respectively (e.g., AAV2 / AAV9-Php.B-Cmv-Strc-N; AAV2 / AAV9-Php.B-Cmv-Strc-C). The waveform on the right shows wild-type (WT) mice with the complete STRC gene (STRC). WT / WTThis represents (n=6). WT mice show sound pressure levels in the range of 30dB to 100dB. The STRC KO mouse (Strc#16 homozygous) in the center shows a limited sound pressure level in the range of 70dB to 120dB, which demonstrates hearing loss below 70dB. However, the waveform on the left demonstrates that STRC KO mice injected with AAV2 / AAV9-Php.B-Cmv-Strc-N;AAV2 / AAV9-Php.B-Cmv-Strc-C were able to recover hearing loss to a level equivalent to that of WT STRC mice.

[0200] The mean auditory threshold or ABR was tested at all frequencies in 4-week-old mice. The ABR results in Figure 4A are shown for STRC knockout (KO) mice injected with the dual vector system described herein, which contains an AAV vector (e.g., AAV2 / AAV9-Php.B-Cmv-Strc-N; AAV2 / AAV9-Php.B-Cmv-Strc-C) containing, respectively, a signal sequence, an N-intane or C-intane, and a sequence encoding the N-terminal or C-terminal portion of the STRC protein. - / - The groups (n=9; center line) each showed recovery from hearing loss, with their minimum ABR thresholds decreasing to 30 dB at several frequencies. Wild-type mice (n=6; bottom line) had minimum ABR thresholds decreasing to 20 dB at several frequencies. However, STRC KO mice (n=5; top line) had severe hearing loss with thresholds exceeding 80 dB.

[0201] The ABR results in Figures 5A and 6A show that STRC KO mice (n=5; upper line with error bar) had severe hearing loss with a threshold exceeding 80 dB. STRC KO mice injected with either the intain AAV coding portion of wild-type STRC encoding N-STRC (n=4; AAV2 / AAV9-Php.B-Cmv-Strc-N) or C-STRC (n=3; AAV2 / AAV9-Php.B-Cmv-Strc-C) showed results similar to those of STRC KO mice, i.e., hearing thresholds exceeding 80 dB. However, the ABR response from wild-type mice (n=5; bottommost line) had a minimum ABR threshold that dropped to 20 dB at several frequencies.

[0202] The DPOAE results in Figure 4B (using the same STRC KO mice used in Figure 45A) demonstrate that STRC KO mice injected with the dual vector system described herein, containing AAV vectors (e.g., AAV2 / AAV9-Php.B-Cmv-Strc-N; AAV2 / AAV9-Php.B-Cmv-Strc-C; n=9; middle line) each containing a signal sequence, an N-intane or C-intane, and a sequence encoding the N-terminal or C-terminal portion of the STRC protein, respectively, show recovery of hearing loss, as evidenced by a DPOAE response reduced to 30 dB at several frequencies. Similarly, wild-type mice (n=6; bottom line) showed a DPOAE response with a minimum DPOAE threshold reduced to 30 dB at several frequencies. However, STRC KO mice (n=5; top line) did not show a DPOAE response (up to 80 dB) under the conditions tested.

[0203] The DPOAE results in Figures 5B and 6B (using the same STRC KO mice used in Figures 5A and 6A) demonstrated that the STRC KO mice (n=5; upper line) did not show a DPOAE response (up to 80 dB) under the tested conditions. Similarly, STRC KO mice injected with either the intein AAV coding portion of wild-type STRC encoding N-STRC (n=4; AAV2 / AAV9-Php.B-Cmv-Strc-N) or C-STRC (n=3; AAV2 / AAV9-Php.B-Cmv-Strc-C) did not show a DPOAE response. However, wild-type mice (n=5; bottom line) showed a DPOAE response with a minimum DPOAE threshold reduced to 30 dB at several frequencies.

[0204] Figures 7A and 7B show the results of monitoring over time three STRC KO mice injected with the dual vector system described herein, which contains an AAV vector (e.g., AAV2 / AAV9-Php.B-Cmv-Strc-N; AAV2 / AAV9-Php.B-Cmv-Strc-C, respectively) containing a signal sequence, an N-intane or C-intane, and a sequence encoding the N-terminal or C-terminal portion of the STRC protein, respectively. The ABR response for each mouse generally showed the lowest threshold at several frequencies for the mouse at week 4 (solid line). Generally, the threshold for each mouse increased over time, from that achieved at week 4 to that achieved at week 6 (dashed line) and week 8 (dashed / dotted line). For example, Figure 7A shows that at week 4 (solid line), mouse #3 had an ABR threshold in the range of 40 dB to 55 dB at frequencies below 10 kHz, and at week 6 (dashed line), mouse #3 had a larger decibel ABR threshold in the range of 50 dB to 70 dB at frequencies below 10 kHz than that observed at week 4. A shift from lower to higher frequencies was observed in the DPOAE response over time. In Figure 7B, the lowest DPOAE threshold response (50 dB) occurred at a frequency of 11 kHz at week 4 and at 16 kHz at week 6 for mouse #3.

[0205] Example 3: Morphological recovery using a dual-vector system The dual AAV delivery system restored STRC expression and tuft morphology, as demonstrated by visual observation of cochlea stained with anti-STRC antibody and Alex488 conjugate secondary antibody (green) and Alexa546-phalloidin (red). For example, Figure 33A shows wild-type (WT) Strc or Strc. Δ / Δ Cochlea injected with [the substance], and Strc injected with a dual AAV vector. Δ / Δ Confocal images of the cochlea are presented. When STRC and actin were stained, both were observed in WT (upper left), and Strc Δ / Δ +Partially observed in the double AAV vector sample (upper right), Strc Δ / Δ In the upper center, disrupted actin outer hair cell (OHC) bundles were observed. In WT, STRC localization was shown by green staining, and inverted V-shaped hair bundles were formed (lower left), and Strc Δ / Δ + In the dual AAV vector sample, partial presentation or recovery occurred (bottom right), Strc Δ / Δ The hair strands that were dyed green could not be shown (bottom center).

[0206] The dual AAV vector delivery system of this disclosure was observed in scanning electron microscopy images to restore the hair bundle morphology (lower panel) in Figure 33B to nearly WT level (upper panel). Δ / Δ Outer hair cell bundles injected with the agent result in disorganized or unorganized OHC bundles, in contrast to the organized OHC bundles of the wild type (center panel).

[0207] Example 4: Hearing function recovery using a dual-vector system The dual AAV vector system also recovers the thresholds of DPOAE and ABR, as demonstrated by Fourier analysis of DPOAE waveforms at sound pressure levels ranging from 10 dB to 50 dB, and Strc Δ / ΔThe dual AAV vector sample (Figure 34A, right) was observed to have an auditory function pattern similar to that of the wild type (Figure 34A, left). The DPOAE threshold was observed in Strc injected with the dual AAV vector. Δ / Δ This shows that the mice regained their auditory function (Figure 34B). Wild type (Figure 34C, left) and Strc Δ / Δ ABR traces recorded from cochlea injected with a dual AAV vector (Figure 34C, right) yielded similar sound pressure levels in the range of 25 dB to 110 dB, but Strc Δ / Δ The cochlea injected with the drug had a sound pressure level of 70 dB to 120 dB (Figure 34C, center). The ABR threshold was measured in the frequency range of 5 kHz to 30 kHz for a single strc Δ / Δ In comparison, Strc Δ / Δ +Dual AAV vectors were shown to demonstrate recovery when injected into mice (Figure 34D).

[0208] Specific manner The following are some non-limiting specific aspects that are considered to fall within the scope of this disclosure.

[0209] Specific embodiment 1. A dual vector system for expressing a protein of interest in a cell, (a) In the direction from 5' to 3', - A signal sequence located at the 5' end of the partial coding sequence that encodes the amino-terminus (N-terminus) portion of the protein of interest; - A partial coding sequence that encodes the N-terminal portion of the protein of interest; - An array that codes for a splice donor array, adjacent to the downstream of a subcode array. A first vector containing a first nucleotide sequence, and (b) In the direction from 5' to 3', - A signal sequence located at the 5' end of the partial coding sequence that encodes the carboxyl terminus (C-terminus) of the protein of interest; - A sequence encoding a splice acceptor sequence, wherein the splice acceptor sequence is sandwiched between a signal sequence and a partial coding sequence encoding the C-terminal portion of the protein of interest; - Partial coding sequence that encodes the C-terminal portion of the protein of interest A second vector containing a second nucleotide sequence including A dual-vector system, including one.

[0210] Specific embodiment 2. A dual vector system of specific embodiment 1, (a) In the direction from 5' to 3', - 5'-terminal inverted repeat (5'ITR) sequence; - Promoter sequence; - A signal sequence that is functionally linked to and under the control of a promoter; - A partial coding sequence that encodes the amino-terminal (N-terminal) portion of a protein of interest, is functionally linked to a promoter, and is under the control of the promoter; - A sequence encoding the amino-terminal fragment (N-intene) of intein, which is functionally linked to a promoter and under the control of the promoter; - Polyadenylated (poly-A) signal sequence; - 3'-terminal inverted repeat (3'ITR) sequence A first vector containing a first nucleotide sequence, and (b) In the direction from 5' to 3', - 5'-terminal inverted repeat (5'ITR) sequence; - Promoter sequence; - A signal sequence that is functionally linked to and under the control of a promoter; - A sequence encoding the carboxy-terminal fragment (C-intei) of intein, which is functionally linked to a promoter and under the control of the promoter; - A partial coding sequence that encodes the carboxyl-terminal (C-terminal) portion of a protein of interest, is functionally linked to a promoter, and is under the control of the promoter; - Polyadenylated (poly-A) signal sequence; - 3'-terminal inverted repeat (3'ITR) sequence A second vector containing a second nucleotide sequence including A dual-vector system, including one.

[0211] Specific embodiment 3. A dual vector system according to specific embodiment 1 or 2, wherein the first vector and the second vector are present in a cell, (a) From the N-terminus towards the C-terminus, - A signal peptide sequence that is ligated to the N-terminal portion of a protein sequence of interest, and the protein sequence of interest is fused to an N-intane protein sequence at its C-terminus. The first protein sequence includes, and (b) From the N-terminus towards the C-terminus, - A signal peptide sequence that is linked to a C-intane protein sequence, wherein the C-intane protein sequence is fused to the N-terminus of the C-terminal portion of the protein sequence of interest. The second protein sequence includes A dual vector system that expresses each of these.

[0212] Specific embodiment 4. A dual vector system according to any one of specific embodiments 1 to 3, wherein the N-terminal portion of the protein of interest and the C-terminal portion of the protein of interest are configured to form a full-length protein of interest.

[0213] Specific embodiment 5. A dual vector system according to any one of specific embodiments 1 to 4, wherein the signal peptide sequence of the first protein sequence and the signal peptide sequence of the second protein sequence are the same.

[0214] Specific Embodiment 6. A dual vector system according to any one of Specific Embodiments 1 to 4, wherein the signal peptide sequence of the first protein sequence and the signal peptide sequence of the second protein sequence are configured to transport the first protein sequence and the second protein sequence into the same cellular compartment.

[0215] Specific Embodiment 7. A dual vector system according to any one of Specific Embodiments 1 to 4, wherein the signal peptide sequence of the first protein sequence and the signal peptide sequence of the second protein sequence are different, and each signal peptide sequence directs its respective protein sequence to the same cellular compartment.

[0216] Specific embodiment 8. A dual vector system according to specific embodiments 1 to 7, wherein the first vector and the second vector are each viral vectors.

[0217] Specific embodiment 9. A dual-vector system of specific embodiment 8, wherein the viral vector is an adeno-associated virus (AAV) vector.

[0218] Specific embodiment 10. A dual vector system of specific embodiment 8 or specific embodiment 9, wherein the viral vectors are of the same or different serotypes.

[0219] Specific embodiment 11. A dual vector system from any one of specific embodiments 1 to 10, wherein the N-terminal and C-terminal portions are configured to form a full-length protein of interest through peptide bonds.

[0220] Specific embodiment 12. A dual vector system according to any one of specific embodiments 1 to 11, wherein the protein of interest is an STRC protein.

[0221] Specific embodiment 13. A dual vector system from any one of Specific Embodiment 12, wherein the STRC protein is encoded by the STRC gene.

[0222] Specific Embodiment 14. A dual vector system according to any one of Specific Embodiments 1 to 13, wherein the signal sequence includes a nucleic acid sequence having at least 80% (e.g., 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:9 or SEQ ID NO:11.

[0223] Specific Embodiment 15. A dual vector system according to any one of Specific Embodiments 1 to 14, wherein the signal sequence encodes a signal peptide sequence having an amino acid sequence that is at least 80% (e.g., 85%, 90%, 95%, 96%, 97%, 98%, 99%) identical to SEQ ID NO:10 or SEQ ID NO:12.

[0224] Specific Embodiment 16. A dual vector system according to any one of Specific Embodiments 1 to 15, wherein the N-terminal portion of the protein of interest includes a nucleic acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with the nucleic acid sequence encoding SEQ ID NO:5, SEQ ID NO:7, SEQ ID NO:15, or SEQ ID NO:16.

[0225] Specific Embodiment 17. A dual vector system according to any one of Specific Embodiments 1 to 16, wherein the N-terminal portion of the protein of interest encodes an amino acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:15 or SEQ ID NO:16.

[0226] Specific embodiment 18. A dual vector system according to any one of specific embodiments 1 to 17, wherein the N-intane sequence includes a nucleic acid sequence having at least 80% (e.g., 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:13.

[0227] Specific embodiment 19. A dual vector system according to any one of specific embodiments 1 to 18, wherein the N-intane sequence encodes an amino acid sequence having at least 80% (e.g., 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:14.

[0228] Specific Embodiment 20. A dual vector system according to any one of Specific Embodiments 1 to 19, wherein the C-terminal portion of the protein of interest includes a nucleic acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with the nucleic acid sequence encoding SEQ ID NO:23 or SEQ ID NO:24.

[0229] Specific Embodiment 21. A dual vector system according to any one of Specific Embodiments 1 to 20, wherein the C-terminal portion of the protein of interest encodes an amino acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:23 or SEQ ID NO:24.

[0230] Specific Embodiment 22. A dual vector system according to any one of Specific Embodiments 1 to 21, wherein the C-intane sequence includes a nucleic acid sequence having at least 80% (e.g., 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO: 21 or SEQ ID NO: 46.

[0231] Specific Embodiment 23. A dual vector system according to any one of Specific Embodiments 1 to 22, wherein the C-intane sequence encodes an amino acid sequence having at least 80% (e.g., 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:22 or SEQ ID NO:49.

[0232] Specific Embodiment 24. A dual vector system according to any one of Specific Embodiments 1 to 23, wherein the first nucleotide sequence includes a nucleic acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO: 5 or SEQ ID NO: 7.

[0233] Specific Embodiment 25. A dual vector system according to any one of Specific Embodiments 1 to 24, wherein the first nucleotide sequence encodes an amino acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:6 or SEQ ID NO:8.

[0234] Specific Embodiment 26. A dual vector system according to any one of Specific Embodiments 1 to 25, wherein the second nucleotide sequence includes a nucleic acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:17 or SEQ ID NO:19.

[0235] Specific Embodiment 27. A dual vector system according to any one of Specific Embodiments 1 to 26, wherein the second nucleotide sequence encodes an amino acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:18 or SEQ ID NO:20.

[0236] Specific embodiment 28. A vector system for expressing the coding sequence of the STRC gene in a host cell, comprising at least one vector whose coding sequence is the STRC gene of SEQ ID NO:1, SEQ ID NO:3, SEQ ID NO:5, SEQ ID NO:7, SEQ ID NO:17, SEQ ID NO:19, SEQ ID NO:29, SEQ ID NO:31, SEQ ID NO:33, or SEQ ID NO:38, the mRNA sequence of SEQ ID NO:30 or SEQ ID NO:32, or a fragment thereof.

[0237] Specific embodiment 29. A vector system of specific embodiment 28, wherein the STRC gene encodes STRC proteins of SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:18, SEQ ID NO:20, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:36, or SEQ ID NO:39, or combinations thereof.

[0238] Specific embodiment 30. A vector system of specific embodiment 28, comprising a dual vector system for expressing the coding sequence of the STRC gene in a host cell, wherein the coding sequence comprises a 5' terminal fragment and a 3' terminal fragment, and the dual vector system is (a) In the direction from 5' to 3', - 5'-terminal inverted repeat (5'ITR) sequence; - Promoter sequence; - A signal sequence that is functionally linked to and under the control of a promoter; - A 5' terminal fragment of the STRC gene coding sequence, which is functionally ligated to a promoter and under the regulation of the promoter; - A sequence encoding the amino-terminal fragment (N-intene) of intein, which is functionally linked to a promoter and under the control of the promoter; - Polyadenylated (poly-A) signal sequences; and - 3'-terminal inverted repeat (3'ITR) sequence A first vector containing a first nucleotide sequence including, (b) In the direction from 5' to 3', - 5'-terminal inverted repeat (5'ITR) sequence; - Promoter sequence; - A signal sequence that is functionally linked to and under the control of a promoter; - A sequence encoding the carboxy-terminal fragment (C-intei) of intein, which is functionally linked to a promoter and under the control of the promoter; - A 3' terminal fragment of the STRC gene coding sequence, which is functionally ligated to a promoter and under the regulation of the promoter; - Polyadenylated (poly-A) signal sequences; and - 3'-terminal inverted repeat (3'ITR) sequence A second vector containing a second nucleotide sequence including A vector system, including

[0239] Specific embodiment 31. The first vector and the second vector in a cell, (a) From the N-terminus towards the C-terminus, - A signal peptide sequence that is ligated to the N-terminal portion of an STRC protein sequence, and in which the STRC protein sequence is fused to an N-intane protein sequence at its C-terminus. The first protein sequence includes, and (b) From the N-terminus towards the C-terminus, - A signal peptide sequence that is linked to a C-intane protein sequence, wherein the C-intane protein sequence is fused to the N-terminus of the C-terminal portion of the STRC protein sequence. The second protein sequence includes A dual vector system according to specific embodiment 30, which expresses each of the following.

[0240] Specific embodiment 32. A dual vector system in any one of specific embodiments 30 to 31, wherein the N-terminal portion of the STRC protein and the C-terminal portion of the STRC protein form a full-length STRC protein.

[0241] Specific embodiment 33. A dual vector system according to any one of specific embodiments 30 to 32, wherein the signal peptide sequence of the first protein sequence and the signal peptide sequence of the second protein sequence are the same.

[0242] Specific embodiment 34. A dual vector system from any one of specific embodiments 30 to 33, wherein the signal peptide sequence of the first protein sequence and the signal peptide sequence of the second protein sequence are configured to transport the first protein sequence and the second protein sequence into the same cellular compartment.

[0243] Specific embodiment 35. A dual vector system according to any one of specific embodiments 30 to 34, wherein the signal peptide sequence of the first protein sequence and the signal peptide sequence of the second protein sequence are different, and each signal peptide sequence directs each protein sequence to the same cellular compartment.

[0244] Specific embodiment 36. A dual vector system according to any one of specific embodiments 30 to 35, wherein the first vector and the second vector are each viral vectors.

[0245] Specific embodiment 37. A dual-vector system of specific embodiment 36, wherein the viral vector is an adeno-associated virus (AAV) vector.

[0246] Specific embodiment 38. A dual vector system according to specific embodiment 36 or specific embodiment 37, wherein the viral vectors have the same serotype.

[0247] Specific embodiment 39. A dual vector system according to specific embodiment 36 or specific embodiment 37, wherein the viral vectors have different serotypes.

[0248] Specific embodiment 40. A dual vector system from any one of specific embodiments 30 to 39, wherein the N-terminal and C-terminal portions form a full-length STRC protein through peptide bonds.

[0249] Specific embodiment 41. A dual vector system of any one of Specific Embodiments 30 to 40, wherein the STRC protein is encoded by the STRC gene.

[0250] Specific embodiment 42. A dual vector system of any one of the specific embodiments 30 to 41, wherein the signal sequence includes a nucleic acid sequence having at least 80% (e.g., 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:9 or SEQ ID NO:11.

[0251] Specific embodiment 43. A dual vector system of any one of specific embodiments 30 to 42, wherein the signal sequence encodes an amino acid sequence having at least 80% (e.g., 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:10 or SEQ ID NO:12.

[0252] Specific embodiment 44. A dual vector system according to any one of the specific embodiments 30 to 43, wherein the N-terminal portion of the STRC protein contains a nucleic acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with the nucleic acid sequence encoding SEQ ID NO:5 or SEQ ID NO:7, or SEQ ID NO:15 or SEQ ID NO:16.

[0253] Specific embodiment 45. A dual vector system of any one of specific embodiments 30 to 44, wherein the N-terminal portion of the STRC protein encodes an amino acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO:15, or SEQ ID NO:16.

[0254] Specific embodiment 46. A dual vector system according to any one of specific embodiments 30 to 45, wherein the N-terminal portion of the STRC protein contains less than 54% of the N-terminal portion of the full-length STRC protein (e.g., 53.8%, 53.6%, 53.4%, 53.2%, 53%, 52%, 50%, 45%).

[0255] Specific embodiment 47. A dual vector system of any one of specific embodiments 30 to 46, wherein the N-intane sequence includes a nucleic acid sequence having at least 80% (e.g., 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:13.

[0256] Specific embodiment 48. A dual vector system of any one of specific embodiments 30 to 47, wherein the N-intane sequence encodes an amino acid sequence having at least 80% (e.g., 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:14.

[0257] Specific embodiment 49. A dual vector system of any one of the specific embodiments 30 to 48, wherein the C-terminal portion of the STRC protein contains a nucleic acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with the nucleic acid sequence encoding SEQ ID NO:17, SEQ ID NO:19, or SEQ ID NO:18, SEQ ID NO:20, SEQ ID NO:23, or SEQ ID NO:24.

[0258] Specific embodiment 50. A dual vector system of any one of the specific embodiments 30 to 49, wherein the C-terminal portion of the STRC protein encodes an amino acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:18, SEQ ID NO:20, SEQ ID NO:23, or SEQ ID NO:24.

[0259] Specific embodiment 51. A dual vector system according to any one of specific embodiments 30 to 50, wherein the C-terminal portion of the STRC protein contains 46% or more of the C-terminal portion of the full-length STRC protein (e.g., 46.2%, 46.4%, 46.6%, 46.8%, 47%, 48%, 50%, 55%).

[0260] Specific embodiment 52. A dual vector system of any one of specific embodiments 30 to 51, wherein the C-intane sequence contains a nucleic acid sequence that is at least 80% (e.g., 85%, 90%, 95%, 96%, 97%, 98%, 99%) identical to SEQ ID NO:21.

[0261] Specific embodiment 53. A dual vector system of any one of specific embodiments 30 to 52, wherein the C-intane sequence encodes an amino acid sequence that is at least 80% (e.g., 85%, 90%, 95%, 96%, 97%, 98%, 99%) identical to SEQ ID NO:22.

[0262] Specific embodiment 54. A dual vector system of any one of the specific embodiments 30 to 53, wherein the first nucleotide sequence includes a nucleic acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO: 5 or SEQ ID NO: 7.

[0263] Specific embodiment 55. A dual vector system of any one of specific embodiments 30 to 54, wherein the first nucleotide sequence encodes an amino acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO:15, or SEQ ID NO:16.

[0264] Specific embodiment 56. A dual vector system of any one of specific embodiments 30 to 55, wherein the second nucleotide sequence includes a nucleic acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:17 or SEQ ID NO:19.

[0265] Specific Embodiment 57. Any one of the Specific Embodiments 30 to 56, a dual vector system wherein the second nucleotide sequence encodes an amino acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%) identity with SEQ ID NO:18, SEQ ID NO:20, SEQ ID NO:23, or SEQ ID NO:24.

[0266] Specific Embodiment 58. At least one cell containing any one vector system of Specific Embodiments 1 to 57 (e.g., 10, 20, 50, 100, 200, 500, 1000, or any number of cells sufficient to successfully express a large biologically active protein), wherein at least one cell may be used to treat, inhibit, or reduce hearing loss in a subject, and the hearing loss may be autosomal recessive hearing loss.

[0267] Specific Embodiment 59. A pharmaceutical composition comprising one vector system from Specific Embodiments 1 to 57 and a pharmaceutically acceptable medium for treating, inhibiting, or reducing hearing loss in a subject, wherein the hearing loss may be autosomal recessive hearing loss.

[0268] Specific embodiment 60. A method for treating, inhibiting, or reducing autosomal recessive hearing loss in a subject, comprising administering an effective amount of one of the dual vector systems from Specific Embodiments 1 to 57 to a subject in need thereof.

[0269] Specific embodiment 61. A method for treating, inhibiting, or reducing autosomal recessive hearing loss in a subject, comprising administering an effective amount of at least one cell of Specific Embodiment 58 to a subject in need thereof.

[0270] Specific embodiment 62. A method for treating, inhibiting, or reducing autosomal recessive hearing loss in a subject, comprising administering an effective amount of the pharmaceutical composition of Specific Embodiment 59 to a subject in need thereof.

[0271] Specific embodiment 63. Any one of the methods in specific embodiments 60 to 62, wherein the autosomal recessive deafness is DFNB16.

[0272] Specific embodiment 64. A method comprising contacting at least one cell of a subject with the pharmaceutical composition of Specific Embodiment 59, wherein the contact delivers a vector system comprising a first nucleotide sequence and a second nucleotide sequence to at least one cell of a subject, and the contacted at least one cell expresses an N-terminal portion and a C-terminal portion of a protein linked by peptide bonds to form a full-length protein.

[0273] Specific embodiment 65. A method for treating and / or preventing a pathology or disease characterized by hearing loss, comprising administering to a subject in need of such treatment an effective amount of a vector system according to any of Specific Embodiments 1 to 57, at least one cell according to Specific Embodiment 58, or a pharmaceutical composition according to Specific Embodiment 59.

[0274] Specific embodiment 66. Any one of the methods of specific embodiments 64 to 65, wherein at least one cell is an inner ear cell.

[0275] Specific embodiment 67. Any one of the methods in specific embodiments 64 to 66, wherein at least one cell is an inner hair cell or an outer hair cell.

[0276] Specific embodiment 68. Any one of the methods in specific embodiments 60 to 67, wherein at least one cell is in vivo or in vitro.

[0277] Specific embodiment 69. One of the methods described in specific embodiments 60 to 68 for improving or restoring auditory function in a subject.

[0278] It will be apparent from the above description that the inventions described herein can be modified and adapted to various uses and conditions. Such embodiments are also within the scope of the following claims.

[0279] Any enumeration of elements in any definition of a variable in this specification includes the definition of that variable as any single element or as a combination (or partial combination) of the elements listed. Any enumeration of embodiments in this specification includes that embodiment as any single embodiment or as a combination with any other embodiment or a part thereof.

[0280] All patents and publications referenced herein are incorporated herein in whole by reference to the same extent that each individual patent and publication is specifically indicated as being incorporated by reference.

Claims

1. A dual vector system for expressing stereocillin (STRC) protein in cells, (a) In the direction from 5' to 3', - The first signal sequence located at the 5' end of the partial coding sequence that encodes the amino-terminus (N-terminus) portion of the STRC protein; - A partial coding sequence that encodes the N-terminal portion of the STRC protein; and - A sequence that encodes the amino-terminal fragment (N-intene) of intein, located downstream and adjacent to the subcoding sequence. A first vector containing a first nucleotide sequence, and (b) In the direction from 5' to 3', - Second signal array; - A sequence encoding the carboxy-terminal fragment (C-intane) of intein, flanked by a second signal sequence and a partial coding sequence encoding the carboxy-terminal (C-terminal) portion of the STRC protein, and adjacent to the downstream of the second signal sequence; and - Partial coding sequence that encodes the C-terminal portion of the STRC protein A second vector containing a second nucleotide sequence including Includes, The first and second signal sequences are configured to transport the first protein expressed by the first vector and the second protein expressed by the second vector to the same cellular compartment of the cell for intein-mediated STRC protein transsplicing. Dual vector system.

2. It is a dual vector system, (a) In the direction from 5' to 3', - 5'-terminal inverted repeat (5'ITR) sequence; - Promoter sequence; - A first signal sequence functionally linked to and under the control of the promoter; - A partial coding sequence that encodes the amino-terminal (N-terminal) portion of a STRC protein, which is functionally linked to a promoter and under the control of the promoter; - A sequence encoding a split-intane-N, which is functionally linked to a promoter, under the control of the promoter, and adjacent downstream of a subcoding sequence; - Polyadenylated (poly-A) signal sequences; and - 3' terminal inverted repeat (3'ITR) sequence A first vector containing a first nucleotide sequence, and (b) In the direction from 5' to 3', - 5'-terminal inverted repeat (5'ITR) sequence; - Promoter sequence; - A second signal sequence functionally linked to and under the control of the promoter; - A sequence encoding split-intein-C, which is functionally linked to and under the control of a promoter, and is adjacent downstream of a second signal sequence; - A partial coding sequence that encodes the carboxyl-terminal (C-terminal) portion of a STRC protein, which is functionally linked to a promoter and under the control of the promoter; - Polyadenylated (poly-A) signal sequences; and - 3' terminal inverted repeat (3'ITR) sequence A second vector containing a second nucleotide sequence including Includes, The first and second signal sequences are configured to transport the first protein expressed by the first vector and the second protein expressed by the second vector to the same cellular compartment of the cell for intein-mediated STRC protein transsplicing. Dual vector system.

3. (a) The first protein moves from the N-terminus to the C-terminus, - A signal peptide sequence comprising a signal peptide sequence which is ligated to the N-terminal portion of an STRC protein sequence, and the STRC protein sequence is fused to an N-intane protein sequence at its C-terminus, and (b) The second protein moves from the N-terminus to the C-terminus, - A dual vector system according to claim 1 or 2, comprising a signal peptide sequence, wherein the signal peptide sequence is linked to a C-intane protein sequence, and the C-intane protein sequence is fused to the N-terminus of the C-terminal portion of the STRC protein sequence.

4. The dual vector system according to any one of claims 1 to 3, wherein the N-terminal portion of the STRC protein and the C-terminal portion of the STRC protein are configured to form a full-length STRC protein.

5. The first vector and the second vector are both viral vectors; The signal sequence encodes either the nucleic acid sequence of SEQ ID NO:9 or SEQ ID NO:11, or the amino acid sequence of SEQ ID NO:10 or SEQ ID NO:12; The N-terminal sequence of the STRC protein is the amino acid sequence of SEQ ID NO:15 or SEQ ID NO:16; The sequence encoding the N-intei is either the nucleic acid sequence with SEQ ID NO:13 or the amino acid sequence with SEQ ID NO:14; The C-terminal sequence of the STRC protein is the amino acid sequence of SEQ ID NO:23 or SEQ ID NO:24; The sequence encoding C-intene is either the nucleic acid sequence of SEQ ID NO:21 or SEQ ID NO:46, or the amino acid sequence of SEQ ID NO:22 or SEQ ID NO:49; The first nucleotide sequence is either the nucleic acid sequence of SEQ ID NO:5 or SEQ ID NO:7, or it encodes the amino acid sequence of SEQ ID NO:6 or SEQ ID NO:8; or The dual vector system according to claim 3 or 4, wherein the second nucleotide sequence is a nucleic acid sequence of SEQ ID NO:17 or SEQ ID NO:19, or encodes an amino acid sequence of SEQ ID NO:18 or SEQ ID NO:

20.

6. The dual vector system according to claim 5, wherein the viral vector is an adeno-associated virus (AAV) vector or a lentivirus.

7. The STRC protein is encoded by the STRC gene. Each of the first and second vectors expresses the coding sequence of the STRC gene. A dual vector system according to any one of claims 1 to 6, wherein the coding sequence comprises the nucleic acid sequence of SEQ ID NO:1, SEQ ID NO:3, SEQ ID NO:5, SEQ ID NO:7, SEQ ID NO:17, SEQ ID NO:19, SEQ ID NO:29, SEQ ID NO:31, SEQ ID NO:33, or SEQ ID NO:

38.

8. The dual vector system according to claim 7, wherein the STRC gene encodes a STRC protein having the amino acid sequence of SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:25, or SEQ ID NO:

26.

9. The first vector and the second vector, in the cell, (a) From the N-terminus towards the C-terminus, - A signal peptide sequence that is ligated to the N-terminal portion of an STRC protein sequence, and in which the STRC protein sequence is fused to an N-intane protein sequence at its C-terminus. The first protein, and (b) From the N-terminus towards the C-terminus, - A signal peptide sequence that is linked to a C-intane protein sequence, wherein the C-intane protein sequence is fused to the N-terminus of the C-terminal portion of the STRC protein sequence. The second protein, which includes A dual vector system according to claim 7 or 8, each expressing a different of the following.

10. The N-terminal portion of the STRC protein and the C-terminal portion of the STRC protein are configured to form a full-length STRC protein; The signal peptide sequences of the first protein and the second protein are the same; The signal peptide sequences of the first protein and the second protein are configured to transport the first and second proteins to the same cellular compartment; The signal peptide sequences of the first protein and the second protein are different, and each signal peptide sequence directs the respective protein to the same cellular compartment; The viral vector is an adeno-associated virus (AAV) vector; The signal sequence either encodes a nucleic acid sequence with SEQ ID NO:9 or SEQ ID NO:11, or an amino acid sequence with SEQ ID NO:10 or SEQ ID NO:12; The N-terminal sequence of the STRC protein is the amino acid sequence of SEQ ID NO:15 or SEQ ID NO:16; The sequence encoding the N-intei is either the nucleic acid sequence with SEQ ID NO:13 or the amino acid sequence with SEQ ID NO:14; The C-terminal sequence of the STRC protein is the amino acid sequence of SEQ ID NO:18, SEQ ID NO:20, SEQ ID NO:23, or SEQ ID NO:24; The sequence encoding C-intene is either the nucleic acid sequence with SEQ ID NO:21 or the amino acid sequence with SEQ ID NO:22; The first nucleotide sequence is either the nucleic acid sequence of SEQ ID NO: 5 or SEQ ID NO: 7, or it encodes the amino acid sequence of SEQ ID NO: 6, SEQ ID NO: 8, SEQ ID NO: 15, or SEQ ID NO: 16; or The dual vector system according to any one of claims 7 to 9, wherein the second nucleotide sequence is the nucleic acid sequence of SEQ ID NO:17 or SEQ ID NO:19, or encodes the amino acid sequence of SEQ ID NO:18, SEQ ID NO:20, SEQ ID NO:23, or SEQ ID NO:

24.

11. A cell containing the vector system according to any one of claims 1 to 10, which is located in vitro.

12. A pharmaceutical composition for treating autosomal recessive hearing loss in a subject, comprising a vector system according to any one of claims 1 to 10 and a pharmaceutically acceptable medium, the pharmaceutical composition comprising an effective amount of the vector system.

13. A pharmaceutical composition for treating autosomal recessive deafness in a subject, comprising the cells described in claim 11, wherein the cells express full-length STRC protein.

14. The pharmaceutical composition according to claim 12 or 13, wherein the autosomal recessive deafness is DFNB16.

15. An in vitro method comprising the step of bringing at least one target cell into contact with the pharmaceutical composition according to claim 12, Contact delivers a vector system containing a first nucleotide sequence and a second nucleotide sequence to at least one target cell. An in vitro method in which at least one contacted cell expresses the N-terminal and C-terminal portions of an STRC protein, which are linked by peptide bonds to form a full-length STRC protein.

16. A pharmaceutical composition for treating and / or preventing a pathology or disease characterized by hearing loss in a subject, comprising an effective amount of the vector system according to any one of claims 1 to 10, or at least one cell according to claim 11.

17. The cell according to claim 11, which is an inner hair cell or an outer hair cell.

18. The pharmaceutical composition according to claim 13 or 16, wherein the cells are inner hair cells or outer hair cells.

19. The pharmaceutical composition according to claim 12, wherein the expression of full-length STRC protein improves or restores auditory function in a subject.

Citation Information

Patent Citations

  • Multi-vector system and its use

    JP2018512125A

  • Intein proteins and uses thereof

    JP2022512718A