Retroviral vector containing an RRE inserted within an intron

Inserting an intron into the retroviral genome by deleting the RRE and placing it near the splice acceptor site enhances transgene expression, addressing the inefficiencies of current gene therapies by increasing protein production and reducing vector doses.

JP2025533754APending Publication Date: 2025-10-09IMPERIAL COLLEGE INNVOATIONS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025517255
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-23
Filing Date
2023-09-22
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Current gene therapies face challenges in producing sufficient therapeutic proteins and require high doses, leading to costly production and potential immune responses, while existing methods to enhance protein production are not applicable to RNA viruses like retroviruses and lentiviruses.

Method used

Introduce an intron into the retroviral genome by deleting the endogenous Rev response element (RRE) and inserting it into the intron, specifically within 20 bp of the splice acceptor branch site, enhancing transgene expression by incorporating a chimeric intron such as β-globin/IgG or SV40 intron.

Benefits of technology

This approach significantly increases transgene expression, potentially by 686-fold or 501-fold, reducing the required dose of gene therapy vectors, making therapy safer and more cost-effective.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025533754000001_ABST
    Figure 2025533754000001_ABST
Patent Text Reader

Abstract

The present invention relates to retroviral vectors modified to improve transgene expression. In particular, the present invention relates to retroviral vectors that lack an endogenous Rev response element (RRE) and contain an intron, particularly a chimeric intron, into which an RRE has been inserted, as well as methods for producing and using the same.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to retroviral vectors modified to improve transgene expression. In particular, the present invention relates to retroviral vectors that lack an endogenous Rev response element (RRE) and contain an intron, particularly a chimeric intron, into which an RRE has been inserted, as well as methods for producing and using the same. [Background technology]

[0002] The use of nucleic acids as medicines, or gene therapy, is a promising new treatment. Many gene therapies currently in use or under development are ineffective in curing diseases because it is difficult to produce enough protein to reach the therapeutic threshold required to treat or cure the disease. Therefore, generating sufficient gene expression is a major obstacle to the success of many gene therapies.

[0003] The current approach to achieving the high doses required for successful gene therapy is to administer large doses of the virus to patients, at more than 1 trillion cells per kg of body weight. For example, Zolgensma delivers 1.1 x 10 cells per kg of body weight. 14The virus is delivered in a single viral genome. Producing such large amounts of virus is costly, contributing to the $1,000,000 USD gene therapy cost, and administering such large amounts of virus to humans could trigger immune responses that threaten the patient's health and therapeutic efficacy. To circumvent these issues, previous studies have focused on gain-of-function mutations that result in more potent proteins. Such approaches have previously been used in gene therapies for hemophilia B (the Padua mutation in factor IX) and lipoprotein lipase deficiency (the S447X variant in lipoprotein lipase). However, such gain-of-function mutations are unavailable in most gene therapies. Furthermore, even with gain-of-function mutations, high doses of gene therapy vectors are still required for effective treatment. Therefore, such gain-of-function mutations alone do not adequately address the existing problems associated with producing sufficient amounts of vector or the undesirable and clinically dangerous side effects associated with the required high doses.

[0004] The present inventors previously developed lentiviral vectors pseudotyped with hemagglutinin-neuraminidase (HN) and fusion (F) proteins from respiratory paramyxoviruses, containing a promoter and a transgene. Typically, the vector backbone is derived from a simian immunodeficiency virus (SIV), such as SIV1 or African green monkey SIV (SIV-AGM). Preferably, the backbone of the viral vector of the present invention is derived from SIV-AGM. The HN and F proteins each bind sialic acid and function to mediate cell fusion for vector entry into target cells. The present inventors discovered that this specifically F / HN-pseudotyped lentiviral vector can efficiently transduce airway epithelia, resulting in sustained transgene expression for a period exceeding the suggested lifespan of airway epithelial cells. Importantly, the present inventors also found that re-administration did not result in a loss of efficacy. These characteristics make the vectors of the present invention attractive candidates for treating diseases through their use in expressing therapeutic proteins (i) in airway cells; (ii) secreted into the airway lumen; and (iii) secreted into the circulatory system. However, even with this cutting-edge platform technology, the level of expressed transgene is below the predicted threshold required for clinical efficacy.

[0005] Thus, there is an unmet clinical need for new technologies to improve the efficacy of gene therapy. It is an object of the present invention to address one or more of these problems. In particular, it is an object of the present invention to provide new nucleic acid cassettes and gene therapy vectors that allow for increased production of therapeutic proteins and allow for smaller doses of vector to be administered to patients. Summary of the Invention [Problem to be solved by the invention]

[0006] Currently, there remains a pressing need for technologies to more efficiently produce therapeutic proteins for gene therapy, including those based on the inventors' unique lentiviral platform. To date, other groups have focused on increasing the amount of therapeutic protein produced by each copy of a gene. For example, research with adeno-associated viral vectors (AAV vectors) has introduced introns into the viral genome. Cellular splicing of the introns leads to mRNA stabilization, increasing the amount of mRNA in the cell and resulting in the production of more protein. However, this approach cannot be easily applied to viral vectors based on RNA viruses, such as retroviruses and lentiviruses. This is because the RNA genome undergoes the same intron removal step as mRNA, thereby removing introns from the RNA genome during production and reducing the amount of protein that the RNA viral vector can produce.

[0007] The present inventors have demonstrated for the first time that it is possible to introduce an intron into a lentiviral genome and that the introduction of such an intron can increase transgene expression. Specifically, the inventors found that removing the endogenous Rev response element of a simian immunodeficiency virus (SIV) vector pseudotyped with a VSV-G or F / HN envelope and introducing a β-globin / IgG intron with a precisely inserted SIV RRV into the SIV.VSV-G or SIV.F / HN genome increased AAT transgene expression by 686-fold or 501-fold, respectively, compared to the corresponding vector lacking the intron. Our innovative approach has the potential to provide several clinically significant advantages: (i) making gene therapy more effective and more easily reaching the required therapeutic window; (ii) reducing the dose of gene therapy agent required for patient administration, making gene therapy safer; and / or (iii) reducing production costs (because fewer vectors are required per patient), solving a major challenge for clinical trials, pharmaceutical companies, and medical professionals. [Means for solving the problem]

[0008] Thus, the present invention provides a retroviral vector containing an intron, wherein: (a) the endogenous Rev response element (RRE) of the retroviral genome is deleted; and (b) the retroviral RRE is inserted into the intron within 100 bp 5' of the splice acceptor branch site. The retroviral RRE is inserted within 20 bp 5' of the splice acceptor branch site. The intron may optionally be a chimeric intron selected from a β-globin / IgG chimeric intron or a chimeric intron from a CAGGS promoter. The intron may optionally be a viral intron selected from an SV40 intron, a CMV intron A, and an adenovirus tripartite leader sequence intron.

[0009] The present invention also preferably provides a retroviral vector comprising a chimeric intron; wherein: (a) the endogenous Rev response element (RRE) of the retroviral genome is deleted; and (b) the retroviral RRE is inserted into the chimeric intron.

[0010] According to any retroviral vector of the present invention, the retroviral RRE inserted into the intron may be an endogenous RRE of the retroviral genome. The RRE may be a simian immunodeficiency virus (SIV) RRE. The RRE may comprise or consist of a nucleic acid sequence having at least 90% identity to SEQ ID NO:1.

[0011] In accordance with any retroviral vector of the present invention, the intron may be less than 1000 bp in length, preferably less than 800 bp in length.

[0012] According to any retroviral vector of the present invention, the chimeric intron may be a β-globin / IgG chimeric intron or a chimeric intron from a CAGGS promoter. The chimeric intron may be a β-globin / IgG chimeric intron and an RRE inserted between (i) a splice donor site comprising or consisting of the nucleic acid sequence TGAGTTTAAGGTAAGT (SEQ ID NO: 2) and (ii) a splice acceptor site comprising or consisting of the nucleic acid sequence CTCTCCACAG (SEQ ID NO: 3). The β-globin / IgG chimeric intron may comprise or consist of a nucleic acid sequence having at least 90% identity to SEQ ID NO: 4. The intron may be a β-globin / IgG chimeric intron and an RRE SIV RRE, and optionally, the chimeric intron containing the RRE may comprise or consist of a nucleic acid sequence having at least 90% identity to SEQ ID NO: 5.

[0013] According to any retroviral vector of the present invention, an intron may be present between the promoter and the transgene operably linked to said promoter, optionally the promoter is selected from the group consisting of a cytomegalovirus (CMV) promoter, an elongation factor 1a (EF1a) promoter, and a hybrid human CMV enhancer / EF1a (hCEF) promoter, preferably an hCEF promoter. The transgene can encode a therapeutic protein, optionally selected from: (a) a secreted therapeutic protein, optionally selected from alpha-1 antitrypsin (AAT), factor VIII, surfactant protein B (SFTPB), ADAMTS13, factor VII, factor IX, factor X, factor XI, von Willebrand factor, granulocyte-macrophage colony-stimulating factor (GM-CSF), surfactant protein C (SP-C), decorin, an anti-inflammatory protein, and a monoclonal antibody against an infectious agent; or (b) CFTR, ABCA3, DNAH5, DNAH11, DNAI1, DNAI2, CSF2RA, CSF2RB, and TRIM-72.

[0014] Any retroviral vector of the present invention may be a lentiviral vector selected from the group consisting of a human immunodeficiency virus (HIV) vector, a simian immunodeficiency virus (SIV) vector, a feline immunodeficiency virus (FIV) vector, an equine infectious anemia virus (EIAV) vector, and a Visna / maedi virus vector.

[0015] Any retroviral vector of the invention is pseudotyped with the hemagglutinin-neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus or the G glycoprotein from vesicular stomatitis virus (G-VSV).

[0016] Any retroviral vector of the present invention may increase transgene expression by at least about 2-fold, preferably at least about 5-fold, and more preferably at least about 10-fold, compared to a corresponding vector lacking an intron into which a retroviral RRE has been inserted.

[0017] The present invention also provides a nucleic acid comprising or consisting of an intron into which a retroviral RRE has been inserted, optionally wherein (i) the intron; and / or (ii) the RRE is as defined herein.

[0018] The present invention further provides a plasmid comprising a nucleic acid of the present invention.

[0019] The present invention also provides retroviral vectors, nucleic acids and / or plasmids of the invention that are codon optimized.

[0020] The present invention further provides compositions comprising the retroviral vectors, nucleic acids and / or plasmids of the invention and a pharmaceutically acceptable carrier.

[0021] The present invention also provides host cells comprising the retroviral vectors, nucleic acids and / or plasmids of the invention.

[0022] The present invention also provides a retroviral vector, nucleic acid plasmid or composition in accordance with the description herein for use in a method of treatment.

[0023] The present invention also provides a method for producing a retroviral vector, the method comprising the steps of: (a) growing cells in suspension; (b) transfecting the cells with one or more plasmids; (c) adding a nuclease; (d) harvesting the lentivirus; (e) adding trypsin; and (f) purifying; the one or more plasmids comprise a nucleic acid of the invention and, optionally, a vector genome plasmid comprising (i) a promoter of the invention and / or (ii) a transgene of the invention.

[0024] The present invention also provides a method for distinguishing between a retroviral vector and a transgene expressed by the retroviral vector, the method comprising: (a1) transfecting cells with a retroviral vector of the present invention; (b1) culturing the cells to allow retroviral expression of the transgene; and (c1) quantifying RNA in the cells; or (a2) quantifying RNA in cells of a sample obtained from a patient treated with a retroviral vector, nucleic acid, plasmid, or composition of the present invention; (i) the amount of RNA containing a chimeric intron with a retroviral RRE inserted corresponds to the copy number of the retroviral vector; (ii) the amount of RNA lacking a chimeric intron with a retroviral RRE inserted corresponds to the amount of transgene mRNA; optionally, the RNA is quantified by a PCR-based or in situ hybridization-based assay. [Brief explanation of the drawings]

[0025] [Figure 1]Rev response element (RRE) introns created by inserting the rSIV RRE into chimeric introns. (A) Schematic of a chimeric intron composed of a splice donor (black) from an intron in hemoglobin subunit B and a splice acceptor (gray) from an intron in immunoglobulin gamma. (B) Schematic of an RRE intron with the rSIV RRE (light gray) inserted between the splice donor and splice acceptor. [Figure 2-1] A-H show schematic diagrams of exemplary plasmids used for the production of vectors of the invention. (A) A schematic diagram of an intron-containing lentiviral vector genome plasmid (pDNA1) encoding an alpha-1-antitrypsin transgene is shown. The intron is inserted between the promoters (hCEF), and the RRE is moved into the intron. (B) A schematic diagram of a plasmid (pDNA2a) encoding codon-optimized SIV Gag and Pol for lentiviral production. (C) A schematic diagram of a plasmid (pDNA2a) encoding SIV Gag and Pol for lentiviral production. (D) A schematic diagram of a plasmid (pDNA2b) encoding SIV Rev for lentiviral production. (E) A schematic diagram of a plasmid (pDNA3a) encoding a Sendai virus-derived fusion protein for lentiviral production. (F) A schematic diagram of a plasmid (pDNA3b) encoding a Sendai virus-derived hemagglutinin-neuraminidase protein for lentiviral production. (G) Schematic diagram of the intronless lentiviral vector genome plasmid (pDNA1) encoding the alpha-1-antitrypsin transgene. The RRE element is 5' to the promoter (hCEF) between the partial GAG sequence and the cPPT sequence. (H) Schematic diagram of the VSV glycoprotein-encoding plasmid (pDNA3) for lentivirus production. [Figure 2-2] Same as above. [Figure 2-3] Same as above. [Figure 2-4] Same as above. [Figure 2-5]Same as above. [Figure 2-6] Same as above. [Figure 2-7] Same as above. [Figure 2-8] Same as above. [Figure 3] The RRE intron enhances AAT expression by 10.6-fold. HEK293T cells were transfected with a plasmid encoding the alpha-1-antitrypsin (AAT) transgene without (intronless) or with (intron) the RRE intron. Inclusion of the intron significantly enhanced AAT expression (p=0.0003). Neg - negative control. Each dot represents a different well of transduced HEK293T cells. A Mann-Whitney test was used for statistical analysis. [Figure 4] Figure 1: The RRE intron is correctly spliced ​​in HEK293T cells. DNA (A) and RNA (B) were extracted from HEK293T cells transfected with a plasmid encoding the alpha-1-antitrypsin (AAT) transgene without (pGM407-intronless) or with (pGM991-intron) the RRE intron. (A) PCR of DNA extracted from transfected cells confirmed that pGM991 contains a 760-bp RRE intron. (B) Reverse transcriptase PCR of RNA extracted from transfected cells confirmed that the RRE intron was spliced ​​out during mRNA maturation. [Figure 5]Figure 1 shows that the RRE intron is packaged into rSIV.VSV-G lentivirus. DNA was extracted from HEK293T cells transduced with rSIV.VSV-G lentivirus expressing AAT without (vGM290) and with (vGM291) the RRE intron. Non-transduced cells (NTC) were used as controls. PCR of the resulting DNA using primers binding to either side of the intron revealed that the intron was packaged into vGM291, producing a 1127-bp product, whereas in the absence of the intron, vGM290 produced a smaller 371-bp product. A no-template control (n) and the lentiviral transfer plasmids vGM290 (pGM407) and vGM291 (pGM991) were included. [Figure 6] Figure 1 shows that the RRE intron is successfully spliced ​​by transduced HEK293T cells. RNA was extracted from HEK293T cells transduced with rSIV.VSV-G lentivirus expressing AAT without (vGM290) and with (vGM291) the RRE intron. Non-transduced cells (NTC) served as controls. RT-PCR of the resulting RNA using primers binding to both sides of the intron revealed that the intron was spliced ​​during expression of vGM291, producing a 277-bp fragment identical to that in vGM290-transduced producer cells. A no-template control (n) and lentiviral transfer plasmids for vGM290 (pGM407) and vGM291 (pGM991) were included. [Figure 7] Figure 1 shows that the RRE intron enhances AAT expression 686-fold in HEK293T cells. HEK293T cells were transduced with rSIV.VSV-G lentivirus expressing AAT without (intronless) and with (intron) the RRE intron. Compared to the leading-edge intronless virus, the RRE-containing virus significantly enhanced AAT expression (p=0.0037, Kruskal-Wallis test with Dunn's multiple comparisons correction). Non-transduced cells (NTC) were used as a control. [Figure 8]The RRE intron increases transgene transcription in HEK293T cells. (A) RT-ddPCR was performed on RNA extracted from transfected HEK293T cells. Inclusion of the intron increased the amount of mRNA (WPRE copies). WPRE transcript copies were normalized to the housekeeping gene beta-2-microglobulin (B2M). (B) Inclusion of the RRE intron increased the amount of mRNA produced per plasmid copy (quantified by ddPCR from DNA extracted from transfected HEK293T cells). Each dot represents a different well of transduced HEK293T cells. Statistical analysis was performed using the Mann-Whitney test. [Figure 9] Figure 1 shows that the RRE intron enhances AAT expression 501-fold in HEK293T cells. HEK293T cells were transduced with rSIV.F / HN lentivirus expressing AAT without (intronless) and with (intron) the RRE intron. Compared to the leading-edge intronless virus, the RRE-containing virus significantly enhanced AAT expression (p=0.0002, Mann-Whitney test). Non-transduced cells (NTC) were used as a control. DETAILED DESCRIPTION OF THE INVENTION

[0026] definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this disclosure pertains. Singleton et al., DICTIONARY OF MICROBIOLOGY AND MOLECULAR BIOLOGY, 20th edition, John Wiley & Sons, New York (1994), and Hale and Marham, THE HARPER COLLINS DICTIONARY OF BIOLOGY, Harper Perennial, NY (1991) provide those skilled in the art with a general dictionary for many of the terms used in this disclosure. The meaning and scope of the terms should be clear; in the event of any potential ambiguity, the definitions provided herein shall prevail over any dictionary or extraneous definitions. It should be understood that the present invention is not limited to the specific methods, protocols, and reagents described herein, and therefore may vary.

[0027] The present disclosure is not limited by the exemplary methods and materials disclosed herein, and any methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the present disclosure. The terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present invention, which is defined solely by the claims.

[0028] The description of the embodiments of the present disclosure is not intended to be exhaustive or to limit the disclosure to the precise form disclosed. Specific embodiments of, and examples of, the present disclosure are described herein for illustrative purposes; however, those skilled in the art will recognize that various equivalent modifications are possible within the scope of the present disclosure. For example, while method steps or functions are presented in a given order, alternative embodiments may perform the functions in a different order, or may perform the functions substantially simultaneously. The teachings of the present disclosure presented herein may be applied to other procedures or methods, as appropriate. The various embodiments described herein may be combined to present further embodiments. Aspects of the present disclosure may be modified, if necessary, to utilize the compositions, functions, and concepts of the above references and this application to yield still further embodiments of the present disclosure. Furthermore, considerations of biological functional equivalence may result in some changes in protein structure without affecting the type or amount of biological or chemical action. These and other changes may be made to the present disclosure in light of the detailed description. All such modifications are intended to be included within the scope of the appended claims.

[0029] Unless otherwise indicated, any nucleic acid sequence is written left to right in 5' to 3' orientation; an amino acid sequence is written left to right in amino to carboxy orientation, respectively.

[0030] The headings provided herein are not limitations of the various aspects or embodiments of the disclosure.

[0031] As used herein, the term "capable of" when used with a verb encompasses or refers to the action of the corresponding verb. For example, "capable of interacting" also means interacting, "capable of cleaving" also means cleaving, "capable of binding" also means binding, and "capable of specifically targeting" also means specifically targeting.

[0032] Definitions of other terms can be found throughout the specification. Before describing exemplary embodiments in more detail, it is to be understood that the disclosure is not limited to the particular embodiments described, and as such may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, as the scope of the disclosure will be defined only by the appended claims.

[0033] Numerical ranges are inclusive of the numbers defining the range. When a range of values ​​is presented, it is understood that each intermediate value, to the tenth of the unit of the lower limit, between the upper and lower limits of that range is also specifically disclosed, unless the context clearly dictates otherwise. Each subrange between any stated or intermediate value within a stated range and any other stated or intermediate value within that stated range is also encompassed within the disclosure. The upper and lower limits of these subranges may independently be included or excluded within the range, but each range in which one or both limits are included within the subrange, or each range in which neither limit is included, is also encompassed within the disclosure and subject to any specifically excluded limit within the stated range. When a stated range includes one or both of the limits, ranges excluding one or both of those included limits are also encompassed within the disclosure.

[0034] As used herein, the articles "a" and "an" may refer to one or to more than one (e.g., at least one) of the grammatical object of the article. Furthermore, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. In this application, the use of "or" means "and / or" unless otherwise stated. Furthermore, the use of the term "including," as well as other forms such as "includes" and "included," is not limiting.

[0035] "About" can generally refer to an acceptable degree of error for the quantity being measured, given the nature or precision of the measurement. Exemplary degrees of error are within 20 percent (%), typically within 10%, and more typically within 5% of a given value or range of values. Preferably, as used herein, the term "about" is understood as + or - (±) 5% from the numerical value of the number with which it is used, preferably ±4%, ±3%, ±2%, ±1%, ±0.5%, ±0.1%.

[0036] The term "consisting of" refers to compositions, methods, and their respective components, as described herein, excluding any element not recited in this description of the invention.

[0037] As used herein, the term "consisting essentially of" refers to elements required for a given invention. The term permits the presence of elements that do not materially affect the basic and novel characteristics or functional characteristics (i.e., inactive or non-immunogenic components) of the invention.

[0038] Any embodiment described herein as "comprising" one or more features is also considered a disclosure of corresponding embodiments "consisting of" and / or "consisting essentially of" such features.

[0039] Concentrations, amounts, volumes, percentages, and other numerical values ​​are presented herein in a range format. It should also be understood that such range format is used merely for convenience and brevity and should be interpreted flexibly to include not only the numerical values ​​explicitly recited as the limits of the range, but also all individual numerical values ​​or subranges encompassed within that range, as if each numerical value and subrange were explicitly recited.

[0040] A "vector" or "construct" (sometimes referred to as a gene delivery or gene transfer "vehicle") refers to a macromolecule or complex of molecules comprising a polynucleotide that is delivered to a host cell either in vitro or in vivo. A vector may be a linear or circular molecule. The vectors of the invention may be viral or non-viral. All disclosures herein regarding the vectors of the invention apply equally to viral and non-viral vectors, unless otherwise stated. All disclosures regarding the viral vectors of the invention apply equally and without reservation to lentiviral (e.g., SIV) vectors, particularly lentiviral (e.g., SIV) vectors pseudotyped with hemagglutinin-neuraminidase (HN) and fusion (F) proteins from respiratory paramyxoviruses (also referred to herein as SIV F / HN or SIV-FHN).

[0041] As used herein, the terms "viral vector," "retroviral vector," and "retroviral F / HN vector" are used interchangeably to refer to a retroviral vector containing retroviral RNA sequences and pseudotyped with the hemagglutinin-neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus, unless otherwise stated. The terms "lentiviral vector" and "lentiviral F / HN vector" are used interchangeably to refer to a lentiviral vector pseudotyped with the hemagglutinin-neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus, unless otherwise stated. All disclosures herein regarding the retroviral vectors of the present invention apply equally and without reservation to the lentiviral vectors of the present invention and to SIV vectors pseudotyped with the hemagglutinin-neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus (also referred to herein as SIV F / HN or SIV-FHN).

[0042] As used herein, the term "intron" refers to the nucleic acid sequence in a gene located between exons. Introns are transcribed together with exons, but are removed from the primary gene transcript by RNA splicing to form mature mRNA. Removal of introns typically leads to mRNA stabilization, increasing the amount of mRNA in cells.

[0043] Rev (viral regulatory factor) is a trans-acting nuclear protein whose functional expression is required for retroviral replication. Specifically, the rev gene product is required for the processing and translation of gag and env mRNAs, and thus rev regulates the expression of viral structural proteins.

[0044] The term "Rev response element" (RRE) refers to a cis-acting antirepressor sequence of env, which responds to the rev gene product. RRE-containing mRNAs are transported from the nucleus to the cytoplasm for translation and virion packaging. The terms "RRE" and "RRE sequence" are used interchangeably herein.

[0045] As used herein, the term "plasmid" refers to a general type of non-viral vector. A plasmid is an extrachromosomal DNA molecule separated from chromosomal DNA and can replicate independently of chromosomal DNA. Preferably, the plasmid is circular and may be double-stranded.

[0046] The terms "nucleic acid cassette," "nucleic acid construct," "expression cassette," and "nucleic acid expression cassette" are used interchangeably to refer to a nucleic acid molecule capable of directing transcription. A nucleic acid cassette contains at least a promoter or a structure functionally equivalent to a promoter and a nucleic acid sequence to be transcribed. Thus, a nucleic acid cassette contains at least a promoter or a structure functionally equivalent to a promoter and a nucleic acid sequence encoding a protein of interest. In the present invention, a nucleic acid cassette contains at least a promoter or a structure functionally equivalent to a promoter, a nucleic acid sequence encoding a signal peptide, and a nucleic acid encoding a therapeutic protein. A nucleic acid cassette may contain additional elements such as an enhancer and / or a transcription termination signal.

[0047] As used herein, the terms "signal peptide," "signal sequence," "targeting sequence," "leader sequence," and "secretion signal" are used interchangeably to refer to a heterologous peptide sequence found at the N-terminus of a secreted protein that is useful in initiating the secretory process. In particular, signal peptides are found in proteins that are targeted to the endoplasmic reticulum and ultimately secreted or retained in the plasma membrane of a cell, particularly as single-pass transmembrane proteins. Signal peptides are typically removed to produce the mature form of the protein. Signal peptides are usually short peptides, typically about 5 to about 35, or about 10 to about 35 amino acids in length, preferably about 5 to about 40 amino acids in length, such as about 10 to about 30 or about 15 to about 30 amino acids in length. Signal peptides may contain a core of hydrophobic amino acids, typically about 4 to about 20 amino acids in length, such as about 5 to about 20, about 5 to about 16, or about 5 to about 15 amino acids in length. If present, the signal peptide is typically at the N-terminus of the protein.

[0048] As used herein, the terms "transduced" and "modified" are used interchangeably to refer to cells that have been modified to express a transgene of interest. Typically, the modification occurs by transduction of the cell.

[0049] When used in reference to a Rev response element (RRE), the term "endogenous" refers to an RRE derived from the same retroviral / lentiviral vector as the retroviral / lentiviral vector of the present invention. A wild-type / unmodified vector contains an RRE in its genome, typically at a defined, standard location. A viral vector of the present invention typically has a genome lacking its endogenous RRE. The endogenous RRE is inserted into an intron and then introduced into the retroviral / lentiviral genome.

[0050] The exogenous RRE is derived from a virus different from the viral vector of the present invention. As a non-limiting example, the viral vector may be an HIV vector and the RRE may be an SIV RRE.

[0051] As used herein, the terms "titer" and "yield" are used interchangeably to refer to the amount of lentiviral (e.g., SIV) vector produced by the methods of the present invention. Titer is a primary criterion for characterizing production efficiency, with a higher titer generally indicating that more retroviral / lentiviral (e.g., SIV) vectors are produced (e.g., using the same amount of reagents). Titer or yield can refer to the number of vector genomes integrated into the genome of target cells (integration titer), which is a measure of "active" viral particles, i.e., the number of particles capable of transducing cells. Transducing units (TU / mL, also referred to as TTU / mL) are a biological readout for the number of host cells transduced under certain tissue culture / virus dilution conditions and are a measure of the number of "active" viral particles. The total number of viral particles (active + inactive) can also be determined using any appropriate means, such as by measuring how much Gag is present in the test solution or how many copies of viral RNA are present in the test solution. The lentiviral particle is then assumed to contain 2000 Gag molecules, or two viral RNA molecules. Once the total particle number and transduction titer / TU are determined, the particle:infectivity ratio is calculated. Amino acids are referred to herein using the amino acid name, three-letter abbreviation, or single-letter abbreviation.

[0052] As used herein, the terms "protein" and "polypeptide" are used interchangeably to refer to a series of amino acid residues joined together by peptide bonds between the alpha-amino and carboxy groups of adjacent residues. The terms "protein" and "polypeptide" refer to a polymer of amino acids, including modified amino acids (e.g., phosphorylated, glycosylated, glycosylated, etc.) and amino acid analogs, regardless of their size or function. Although "protein" and "polypeptide" are often used in reference to relatively large polypeptides, while the term "peptide" is often used in reference to small polypeptides, the use of these terms overlaps in the art. As used herein, the terms "protein" and "polypeptide" are used interchangeably when referring to gene products and fragments thereof. Thus, exemplary polypeptides or proteins include gene products, naturally occurring proteins, homologs, orthologs, paralogs, fragments, and other equivalents, variants, fragments, and analogs of the foregoing.

[0053] As used herein, the terms "polynucleotide," "nucleic acid," and "nucleic acid sequence" refer to any molecule, preferably a polymeric molecule, that incorporates ribonucleic acid, deoxyribonucleic acid, or analog units thereof. A nucleic acid may be single-stranded or double-stranded. A single-stranded nucleic acid may be one nucleic acid strand of a denatured double-stranded DNA. Alternatively, a nucleic acid may be a single-stranded nucleic acid that is not derived from double-stranded DNA. In one embodiment, a nucleic acid may be DNA. In another embodiment, a nucleic acid may be RNA. Suitable nucleic acid molecules are DNA, including genomic DNA or cDNA. Other suitable nucleic acid molecules are RNA, including siRNA, shRNA, and antisense oligonucleotides. The terms "transgene" and "gene" are also used interchangeably, and both terms encompass fragments or variants thereof that encode target proteins.

[0054] Transgenes of the present invention include nucleic acid sequences removed from their naturally occurring environment, recombinant or cloned DNA isolates, and chemically synthesized analogs or analogs biologically synthesized in heterologous systems.

[0055] Minor variations within the amino acid sequences of the present invention are contemplated as variations encompassed by the present invention, provided that the variations within the amino acid sequences maintain at least 60%, at least 70%, more preferably at least 80%, at least 85%, at least 90%, at least 95% sequence identity, and most preferably at least 97% or at least 99% sequence identity with the amino acid sequences of the present invention or fragments thereof, as defined anywhere herein. As used herein, the term "homology" is used to mean "identity." Thus, the sequences of variants or analogs of the amino acid sequences of the present invention may differ based on substitutions (typically conservative substitutions), deletions, or insertions. Proteins containing such variations are referred to herein as variants.

[0056] The proteins of the present invention may include variants in which amino acid residues from one species are substituted with the corresponding residue in another species, either at conserved or non-conserved positions. Variants of the protein molecules disclosed herein may be produced and used in the present invention. Following the precedent of computational chemistry in the application of multivariate data analysis techniques to structure / property-activity relationships [see, e.g., Wold et al., Multivariate data analysis in chemistry. Chemometrics-Mathematics and Statistics in Chemistry (ed.: B. Kowalski); D. Reidel Publishing Company, Dordrecht, Holland, 1984 (ISBN 90-277-1846-6)], quantitative activity-property relationships of proteins can be analyzed using well-known mathematical techniques such as statistical regression, pattern recognition, and pattern classification [see, e.g., Norman et al., Applied Regression Analysis. Wiley-Interscience; 3rd Edition (April 1998) ISBN: 0471170828; Kandel, Abraham et al., Computer-Assisted Reasoning in Cluster Analysis. Prentice Hall PTR, (May 11, 1995), ISBN: 0133418847; Krzanowski, Wojtek. Principles of Multivariate Analysis: A User's Perspective (Oxford Statistical Science Series, No. 22 (Paper)). Oxford University Press; (December 2000), ISBN: 0198507089; Witten, Ian H. et al., Data Mining: Practical Machine Learning Tools and Techniques with Java Implementations. Morgan Kaufmann; (October 11, 1999), ISBN: 1558605525; Denison David GT(eds.), Bayesian Methods for Nonlinear Classification and Regression (Wiley Series in Probability and Statistics). John Wiley & Sons; (July 2002), ISBN: 0471490369; Ghose, Arup K. et al., Combinatorial Library Design and Evaluation Principles, Software, Tools, and Applications in Drug Discovery. ISBN: 0-8247-0487-8. Protein properties are derived from empirical and theoretical models of protein sequence, functional structure, and three-dimensional structure (e.g., analysis of likely contact residues or calculation of physicochemical properties), and these properties are considered individually and in combination.

[0057] As used herein, amino acids are referred to using the amino acid name, three-letter abbreviation, or single-letter abbreviation. As used herein, the term "protein" includes proteins, polypeptides, and peptides. As used herein, the term "amino acid sequence" is synonymous with the term "polypeptide" and / or the term "protein." In some cases, the term "amino acid sequence" is synonymous with the term "peptide." As used herein, the terms "protein" and "polypeptide" are used interchangeably. In this disclosure and claims, conventional single-letter and three-letter codes for amino acid residues are used. The three-letter code for amino acids is defined in accordance with the IUPACIUB Joint Commission on Biochemical Nomenclature (JCBN). It is also understood that due to the degeneracy of the genetic code, a polypeptide may be coded for by more than one nucleotide sequence.

[0058] Amino acid residues at non-conserved positions may be substituted with either conservative or non-conservative residues. Conservative amino acid replacements are particularly contemplated.

[0059] " Conservative amino acid substitution " refers to the amino acid substitution in which an amino acid residue is replaced with an amino acid residue having a similar side chain.In the art, a family of amino acid residues with similar side chains has been defined, including basic side chains (e.g., lysine, arginine, or histidine), acidic side chains (e.g., aspartic acid or glutamic acid), uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, or cysteine), non-polar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, or tryptophan), beta-branched side chains (e.g., threonine, valine, isoleucine), and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, or histidine).Therefore, if an amino acid in a polypeptide is replaced with another amino acid from the same side chain family, the amino acid substitution is considered conservative. The incorporation of conservatively modified variants into the proteins of the present invention does not exclude other forms of variants, such as polymorphic variants, interspecies homologs, and alleles.

[0060] "Non-conservative amino acid substitutions" include (i) substitutions in which a residue having an electropositive side chain (e.g., Arg, His, or Lys) is substituted with or by an electronegative residue (e.g., Glu or Asp), (ii) substitutions in which a hydrophilic residue (e.g., Ser or Thr) is substituted with or by a hydrophobic residue (e.g., Ala, Leu, Ile, Phe, or Val), (iii) substitutions in which a cysteine ​​or proline is substituted with or by any other residue, or (iv) substitutions in which a residue having a bulky hydrophobic or aromatic side chain (e.g., Val, His, Ile, or Trp) is substituted with or by a residue having a small side chain (e.g., Ala or Ser) or no side chain (e.g., Gly).

[0061] "Insertions" or "deletions" are typically in the range of about 1, 2, or 3 amino acids. Permitted mutations are determined experimentally by using recombinant DNA techniques to systematically introduce amino acid insertions or deletions into a protein and assaying the resulting recombinant variants for activity. This does not require undue experimentation on the part of one of ordinary skill in the art.

[0062] A "fragment" of a polypeptide comprises at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, or more of the original polypeptide.

[0063] The polynucleotide of the present invention can be produced by any means known in the art.For example, large amounts of polynucleotides can be produced by replication in suitable host cells.The natural or synthetic DNA fragment encoding the desired fragment is incorporated into a recombinant nucleic acid construct, typically a DNA construct, which can be introduced into and replicated in prokaryotic or eukaryotic cells.Normally, the DNA construct will be suitable for self-replication in unicellular hosts such as yeast or bacteria, but can also be introduced into and integrated into the genome of cultured insect cell systems, mammalian cell systems, plant cell systems, or other eukaryotic cell systems.

[0064] Polynucleotides of the present invention can also be produced by chemical synthesis, for example, by the phosphoramidite or triester method, which can be performed on commercially available automated oligonucleotide synthesizers. Double-stranded fragments can be obtained from the single-stranded product of chemical synthesis by synthesizing the complementary strand and annealing the strands together under appropriate conditions, or by using DNA polymerase with appropriate primer sequences to add the complementary strand.

[0065] The term "isolated" in the context of the present invention, when applied to a nucleic acid sequence, indicates that the polynucleotide sequence has been removed from its natural genetic environment and, thus, is free of other exogenous or undesired coding sequences (but may include naturally occurring 5' and 3' untranslated regions, such as promoters and terminators), and is in a form suitable for use in genetically engineered protein production systems. Such isolated molecules are molecules that have been separated from their natural environment.

[0066] Given the degeneracy of the genetic code, considerable sequence variation is possible among the polynucleotides of the present invention. Degenerate codons, which include all possible codons for a given amino acid, are shown below:

[0067] [Table 1]

[0068] Those skilled in the art will understand that there is flexibility in determining degenerate codons, which represent all possible codons that encode each amino acid. For example, some polynucleotides encompassed by a degenerate sequence may encode variant amino acid sequences, and those skilled in the art can readily identify such variant sequences by reference to the amino acid sequences of the present invention.

[0069] A "variant" nucleic acid sequence has substantial homology or substantial similarity to a reference nucleic acid sequence (or a fragment thereof). A nucleic acid sequence or a fragment thereof is "substantially homologous" (or "substantially identical") to a reference sequence if, when optimally aligned (with appropriate nucleotide insertions or deletions) with another nucleic acid (or its complementary strand), there is nucleotide sequence identity in at least about 70%, 75%, 80%, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99%, or more of the nucleotide bases. Methods for determining nucleic acid sequence homology are known in the art.

[0070] Alternatively, a "variant" nucleic acid sequence is substantially homologous to (or substantially identical to) a reference sequence (or a fragment thereof) if the "variant" sequence and the reference sequence are capable of hybridizing under stringent (e.g., highly stringent) hybridization conditions. Hybridization of nucleic acid sequences will be affected by conditions such as salt concentration (e.g., NaCl), temperature, or organic solvents, in addition to base composition, length of complementary strands, and number of nucleotide base mismatches between hybridizing nucleic acids, as will be readily understood by those skilled in the art. Preferably, stringent temperature conditions are utilized, including temperatures generally above 30°C, typically above 37°C, and preferably above 45°C. Stringent salt conditions will usually be less than 1000 mM, typically less than 500 mM, and preferably less than 200 mM. pH is typically between 7.0 and 8.3. The combination of parameters is much more important than any single parameter.

[0071] The method for determining nucleic acid sequence identity percentage is known in the art.For example, when evaluating nucleic acid sequence identity, the sequence with a specified number of consecutive nucleotides is aligned with the nucleic acid sequence (with the same number of consecutive nucleotides) derived from the corresponding part of the nucleic acid sequence of the present invention.The tool known in the art for determining nucleic acid sequence identity percentage includes nucleotide BLAST (described below).

[0072] Those skilled in the art will understand that different species exhibit "preferential codon usage." As used herein, the term "preferential codon usage" refers to the codons that are most frequently used within the cells of a particular species, thus preferring one or a few codons representing the possible codons encoding each amino acid. For example, the amino acid threonine (Thr) is encoded by ACA, ACC, ACG, or ACT, but in mammalian host cells, ACC is the most commonly used codon; in other species, different codons may be preferred. Preferred codons for a particular host cell species are introduced into the polynucleotides of the present invention by various methods known in the art. Introduction of preferred codon sequences into recombinant DNA enhances protein production, for example, by making protein translation more efficient in a particular cell type or species. Thus, in accordance with the present invention, any nucleic acid sequence, in addition to the gag-pol gene, can be codon-optimized for expression in a host cell or target cell. In particular, the vector genome (or corresponding plasmid), the REV gene (or corresponding plasmid), the fusion protein (F) gene (or corresponding plasmid), and / or the hemagglutinin-neuraminidase (HN) gene (or corresponding plasmid), or any combination thereof, are codon-optimized.

[0073] A "fragment" of a polynucleotide of interest comprises a series of consecutive nucleotides derived from the sequence of said full-length polynucleotide. By way of example, a "fragment" of a polynucleotide of interest may comprise (or consist of) at least 30 consecutive nucleotides derived from the sequence of said polynucleotide (e.g., at least 35, 50, 75, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, or 1000 consecutive nucleic acid residues of said polynucleotide). A fragment may comprise at least one antigenic determinant and / or encode at least one antigenic epitope of the corresponding polypeptide of interest. Typically, a fragment as defined herein retains the same function as the full-length polynucleotide.

[0074] As used herein, the terms "reduce," "reduced," "reduction," or "inhibit" are all used to mean a statistically significant decrease. The terms "reduce," "reduction," "reducing," or "inhibiting" typically refer to a decrease of at least 10% compared to a reference level (e.g., in the absence of a given treatment), and can include, for example, a decrease of at least about 10%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, or more. As used herein, "reduction" or "inhibition" encompasses complete inhibition or reduction compared to a reference level. "Complete inhibition" is 100% inhibition (ie, abolition) compared to the reference level.

[0075] As used herein, the terms "increased," "increase," "enhance," or "activate" are all used to mean an increase by a statistically significant amount. The terms "increased," "increase," "enhance," or "activate" can mean an increase of at least 25%, at least 50% compared to a reference level, for example, an increase of at least about 50%, or at least about 75%, or at least about 80%, or at least about 90%, or at least about 100%, or at least about 150%, or at least about 200%, or at least about 250%, or more, compared to a reference level, or an increase of at least about 1.5-fold, or at least about 2-fold, or at least about 2.5-fold, or at least about 3-fold, or at least about 4-fold, or at least about 5-fold, or at least about 10-fold, or any increase between 1.5-fold and 10-fold or more, compared to a reference level. In the context of yield or titer, "increase" refers to an observable or statistically significant increase in such level.

[0076] As used herein, the terms "individual," "subject," and "patient" are used interchangeably to refer to a mammalian subject for whom diagnosis, prognosis, disease monitoring, treatment, therapy, and / or therapy optimization is desired. The mammal may be (without limitation) a human, non-human primate, mouse, rat, dog, cat, horse, or cow. In preferred embodiments, the individual, subject, or patient is human. An "individual" may be an adult, juvenile, or infant. An "individual" may be male or female.

[0077] A "subject in need" of treatment for a particular condition can be an individual who has the condition, has the condition, or has been diagnosed as being at risk of developing the condition.

[0078] The subject may be a subject who has already been diagnosed with, or who has been identified as suffering from, a condition requiring treatment, or one or more complications or symptoms associated with such a condition, and may optionally have already undergone treatment for the condition defined herein, or one or more complications or symptoms associated with said condition. Alternatively, the subject may also be a subject who has not yet been diagnosed with the condition defined herein, or one or more symptoms or complications associated with said condition. For example, the subject may be a subject who exhibits one or more risk factors for the condition, or one or more symptoms or complications associated with said condition, or may be a subject who does not exhibit risk factors.

[0079] As used herein, the term "healthy individual" refers to an individual or group of individuals in a healthy state, e.g., individuals who do not exhibit symptoms of disease, have not been diagnosed with a disease, and / or are unlikely to develop a disease, e.g., cystic fibrosis (CF), or any other disease described herein. Preferably, the healthy individual is not taking medication that affects CF and has not been diagnosed with any other disease. One or more healthy individuals may have similar gender, age, and / or body mass index (BMI) compared to the test individual. Application of standard statistical methods used in medicine allows for the determination of normal expression levels in healthy individuals and significant deviations from such normal levels.

[0080] As used herein, the terms "control" and "reference population" are used interchangeably.

[0081] As used herein, the term "pharmaceutically acceptable" means approved by a regulatory agency of the U.S. federal or state government or listed in the U.S. Pharmacopoeia, the European Pharmacopoeia, or other generally recognized pharmacopoeias.

[0082] The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein should be construed as an admission that such publications constitute prior art to the claims appended hereto.

[0083] Disclosure relating to various methods of the present invention is intended to apply equally to other methods, therapeutic uses or treatments, data storage media, or devices, computer program products, and vice versa.

[0084] Retroviral and lentiviral vectors The present invention relates to retroviral / lentiviral (e.g., SIV) vectors. The retroviral / lentiviral vectors of the present invention integrate into the genome of transduced cells, resulting in long-term persistent expression. The term "retrovirus" refers to any genus of RNA viruses in the Retroviridae family that encode the enzyme reverse transcriptase. The term "lentivirus" refers to the Retroviridae family. Accordingly, all references herein to retroviral vectors of the present invention apply equally and without reservation to lentiviral vectors. Furthermore, all references herein to lentiviral vectors of the present invention apply equally and without reservation to retroviral vectors.

[0085] Examples of retroviruses suitable for use in the present invention include gammaretroviruses such as murine leukemia virus (MLV) and feline leukemia virus (FLV). Examples of lentiviruses suitable for use in the present invention include simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), feline immunodeficiency virus (FIV), equine infectious anemia virus (EIAV), and visna / maedi virus. Preferably, the present invention relates to lentiviral vectors and their production. Particularly preferred lentiviral vectors are SIV vectors (including all strains and subtypes), such as SIV-AGM (originally isolated from African green monkeys, Cercopithecus aethiops). Alternatively, the present invention relates to HIV vectors.

[0086] Retroviral / lentiviral (e.g., SIV) vectors of the present invention are typically pseudotyped with the hemagglutinin-neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus or the G glycoprotein from vesicular stomatitis virus (referred to as VSV-G or G-VSV). Preferably, lentiviral (e.g., SIV) vectors of the present invention are pseudotyped with HN and F proteins from a respiratory paramyxovirus. Particularly preferably, the respiratory paramyxovirus is Sendai virus (murine parainfluenza virus type 1). Retroviral / lentiviral (e.g., SIV) vectors of the present invention are pseudotyped with proteins from another virus, provided that the pseudotyping proteins do not adversely affect (or even increase) the production titer of the vector and / or do not adversely affect (or even increase) the expression of the transgene. Non-limiting examples of other proteins that can be used to pseudotype the retroviral / lentiviral (e.g., SIV) vectors of the invention include the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) spike protein or modified versions thereof. The VSV-G and SARS-CoV2 spike proteins used for pseudotyping are those described in UK Patent Application No. 2118685.3 and International Application No. PCT / GB2022 / 050933, each of which is incorporated herein by reference in its entirety.

[0087] Retroviral / lentiviral (e.g., SIV) vectors for use in accordance with the present invention may be integrase-competent (IC). Alternatively, lentiviral (e.g., SIV) vectors may be integrase-deficient (ID).

[0088] The viral vectors of the present invention, particularly the retroviral / lentiviral (e.g., SIV) vectors described herein, are capable of transducing one or more of the cell types described herein to achieve long-term transgene expression.

[0089] The retroviral / lentiviral (e.g., SIV) vectors of the present invention allow for high-level transgene expression, and in particular, the retroviral / lentiviral (e.g., SIV) vectors of the present invention typically result in high-level (therapeutic) expression of therapeutic proteins.

[0090] The nucleic acid sequence encoding a therapeutic protein contained in the viral vector of the present invention, particularly the retroviral / lentiviral (e.g., SIV) vector of the present invention, is modified to facilitate expression.For example, the transgene sequence may be CpG-depleted (or CpG-free) and / or codon-optimized to facilitate gene expression.In the art, standard techniques for modifying the sequence of a transgene in this manner are known.The genome of the retroviral / lentiviral (e.g., SIV) vector is completely or partially CpG-depleted (or CpG-free) and / or codon-optimized.

[0091] Retroviral / lentiviral (e.g., SIV) vectors, such as those of the present invention, are suitable for transducing stem / progenitor cells because they integrate into the genome of transduced cells, resulting in long-term persistent expression. In the lung, several cell types with regenerative capacity have been identified that play a role in maintaining specific cell lineages in the conducting airways and alveoli. These include basal cells and submucosal gland duct cells in the upper respiratory tract, club cells and neuroendocrine cells in the bronchial airways, bronchoalveolar stem cells in the terminal bronchioles, and type II pneumocytes in the alveoli. Thus, and without being bound by theory, it is believed that the retroviral / lentiviral (e.g., HIV / SIV) vectors induce long-lasting gene expression of a transgene of interest by introducing the transgene into one or more long-lived airway epithelial cells or cell types, such as basal cells and submucosal gland duct cells in the upper respiratory tract, club cells and neuroendocrine cells in the bronchial airways, bronchoalveolar stem cells in the terminal bronchioles, type II alveolar epithelial cells in the alveoli, submucosal acinar cells, ionocytes, and type I pneumocytes. As demonstrated herein, integration of retroviral / lentiviral (e.g., SIV) vectors bearing the modified retroviral / lentiviral (e.g., SIV) RNA sequences of the present invention into target cell genomes unexpectedly does not have adverse effects and may even be increased.

[0092] Therefore, the retroviral / lentiviral (e.g., SIV) vectors of the present invention can transduce one or more cells or cell lines with regenerative potential in the lungs (including the airways and respiratory tract) to achieve long-term gene expression. For example, retroviral / lentiviral (e.g., SIV) vectors can transduce basal cells, such as basal cells, in the upper airways / respiratory tract. Basal cells play a central role in the process of epithelial maintenance and repair after injury. In addition, basal cells are widely distributed along the human respiratory epithelium, with their relative distribution ranging from 30% (large airways) to 6% (small airways).

[0093] The retroviral / lentiviral (e.g., SIV) vectors of the present invention are used to transduce stem / progenitor cells that have been isolated and expanded ex vivo before being administered to a patient. Preferably, the retroviral / lentiviral (e.g., SIV) vectors of the present invention are used to transduce cells in the lung (or airway / respiratory tract) in vivo.

[0094] The retroviral / lentiviral (e.g., SIV) vectors of the present invention exhibit remarkable resistance to shear forces when passed through clinically relevant delivery devices such as bronchoscopes, spray bottles, and nebulizers, with only a slight reduction in transduction capacity.

[0095] The retroviral / lentiviral (e.g., SIV) vector of the present invention allows high-level transgene expression, resulting in high-level (therapeutic) expression of therapeutic proteins. The retroviral / lentiviral (e.g., SIV) vector of the present invention typically results in high-level transgene expression when administered to a patient. The terms high expression and therapeutic expression are used interchangeably herein. Expression is measured by any suitable method (qualitative or quantitative, preferably quantitative), and the concentration is given in any suitable measurement unit, for example, ng / ml or nM.

[0096] The expression of the transgene of interest is given relative to the expression of the corresponding endogenous (deficient) gene in the patient. Expression is measured in terms of mRNA or protein expression. The expression of the transgene of the present invention, such as a functional CFTR gene, is quantified relative to the endogenous gene, such as an endogenous (dysfunctional) CFTR gene, in terms of mRNA copies per cell or any other suitable unit.

[0097] The expression level of the transgene and / or encoded therapeutic protein of the invention is measured in lung tissue, epithelial lining fluid, and / or serum / plasma, as appropriate. Thus, high expression levels and / or therapeutic expression levels may refer to concentrations in the lung, epithelial lining fluid, and / or serum / plasma.

[0098] The retroviral / lentiviral (e.g., SIV) vectors of the present invention allow for long-term transgene expression, resulting in long-term expression of a therapeutic protein. As described herein, the terms "long-term expression," "persistent expression," "long-term persistent expression," and "sustained expression" are used interchangeably. The retroviral / lentiviral (e.g., SIV) vectors of the present invention allow for long-term transgene expression, particularly by airway cells, as described herein, resulting in long-term expression of a therapeutic protein. Long-term expression according to the present invention refers to expression of a therapeutic gene and / or protein, preferably at therapeutic levels, for at least 45 days, at least 60 days, at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 730 days, or more. Preferably, long-term expression refers to expression for at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 720 days or more, more preferably at least 360 days, at least 450 days, at least 720 days or more, which long-term expression may be achieved by repeated administration or by a single administration.

[0099] In particular, the retroviral / lentiviral (e.g., SIV) vectors of the invention are capable of driving (increased) long-term, sustained expression of a therapeutic protein in airway cells in vivo in a patient. Preferably, the retroviral / lentiviral (e.g., SIV) vectors of the invention drive expression of a therapeutic protein in airway cells for at least 45 days, more preferably for at least 90 days.

[0100] Repeated doses are administered twice daily, daily, twice weekly, weekly, monthly, every two months, every three months, every four months, every six months, yearly, every two years or more, and are continued for as long as required, for example, at least six months, at least one year, two years, three years, four years, five years, ten years, fifteen years, twenty years or more, up to the life of the patient being treated.

[0101] The retroviral / lentiviral (e.g., SIV) vectors of the present invention exhibit enhanced expression of therapeutic proteins, and thus can provide long-lasting, repeatable, high-level expression, particularly in airway cells, without inducing excessive immune responses.

[0102] Preferably, the present invention relates to F / HN retroviral / lentiviral vectors, in particular SIV F / HN vectors, comprising a promoter and a transgene.

[0103] The viral vectors of the present invention are produced using any suitable process known in the art. In particular, the retroviral / lentiviral (e.g., SIV) vectors of the present invention are produced using the method disclosed in International Application No. PCT / GB2022 / 050524, the entire contents of which are incorporated herein by reference.

[0104] Viral vectors of the invention, particularly retroviral / lentiviral (e.g., SIV) vectors of the invention, may comprise a central polypurine tract (cPPT) and / or a woodchuck hepatitis virus posttranscriptional regulatory element (WPRE). An exemplary WPRE sequence is provided by SEQ ID NO:39.

[0105] The retroviral / lentiviral (e.g., SIV) vectors of the present invention have been modified to (i) delete the endogenous RRE as described herein; and (ii) introduce one or more introns into which a retroviral / lentiviral (e.g., SIV) RRE has been inserted. In other words, the retroviral / lentiviral (e.g., SIV) vectors of the present invention have a genome that has been modified to (i) delete the endogenous RRE as described herein; and (ii) introduce one or more introns into which a retroviral / lentiviral (e.g., SIV) RRE has been inserted. Any reference herein to a retroviral / lentiviral (e.g., SIV) vector of the present invention comprising an intron with an RRE inserted therein applies equally and without reservation to a retroviral / lentiviral (e.g., SIV) vector of the present invention whose genome comprises an intron with an RRE inserted therein.

[0106] As described herein, the RRE-containing intron typically introduced into the retroviral / lentiviral (e.g., SIV) vector of the present invention is properly spliced ​​by the target cell and aids in the maturation of a stable mRNA molecule. This results in increased expression of the coding region of the retroviral / lentiviral (e.g., SIV) genome containing the transgene. Thus, introduction of an RRE-containing intron into a retroviral / lentiviral (e.g., SIV) vector results in increased expression of a transgene that may encode a therapeutic protein. Introduction of an RRE-containing intron into a retroviral / lentiviral (e.g., SIV) vector can increase the expression of a therapeutic protein compared to the expression of a therapeutic protein from a corresponding retroviral / lentiviral (e.g., SIV) vector that does not have an RRE-containing intron. Thus, a retroviral / lentiviral (e.g., SIV) vector of the present invention that contains an RRE-containing intron typically exhibits increased transgene expression compared to the expression of a transgene from a corresponding retroviral / lentiviral (e.g., SIV) vector that lacks the RRE-containing intron. As a non-limiting example, a retroviral / lentiviral (e.g., SIV) vector of the present invention comprising an AAT transgene (SERPINA1) and a β-globulin / IgG chimeric intron with an inserted RRE (e.g., a β-globulin / IgG chimeric intron with an RRE of SEQ ID NO: 5) may increase AAT expression compared to a corresponding retroviral / lentiviral (e.g., SIV) vector comprising the AAT transgene but lacking the β-globulin / IgG chimeric intron with an inserted RRE (e.g., a β-globulin / IgG chimeric intron with an RRE of SEQ ID NO: 5).

[0107] The increase in therapeutic protein expression by a retroviral / lentiviral (e.g., SIV) vector of the present invention comprising an intron containing an RRE can be as defined herein. In particular, the increase in therapeutic protein expression by a retroviral / lentiviral (e.g., SIV) vector of the present invention comprising an intron containing an RRE can typically be at least about 5-fold, at least about 10-fold, at least about 50-fold, at least about 100-fold, at least about 200-fold, at least about 500-fold, at least about 600-fold, or more, compared to the expression of a therapeutic protein from a corresponding retroviral / lentiviral (e.g., SIV) vector without an intron containing an RRE. Preferably, the increase in therapeutic protein expression by a retroviral / lentiviral (e.g., SIV) vector of the present invention comprising an intron containing an RRE is typically at least about 10-fold, more preferably at least about 100-fold, and even more preferably at least about 500-fold, compared to the expression of a therapeutic protein from a corresponding retroviral / lentiviral (e.g., SIV) vector without an intron containing an RRE. As a non-limiting example, when the RRE-containing intron is an RRE-containing β-globulin / IgG chimeric intron, such as an intron containing the β-globulin / IgG chimeric RRE of SEQ ID NO: 5, the intron may increase transgene expression by a retroviral / lentiviral (e.g., SIV) vector of the present invention containing the β-globulin / IgG chimeric RRE-containing intron by at least 600-fold, e.g., about 686-fold, compared to transgene expression by a corresponding retroviral / lentiviral (e.g., SIV) vector that does not have the β-globin / IgG chimeric RRE-containing intron.

[0108] The increased expression of a therapeutic protein by a retroviral / lentiviral (e.g., SIV) vector of the present invention comprising an intron containing an RRE can be as defined herein. In particular, the increased expression of a therapeutic protein by a retroviral / lentiviral (e.g., SIV) vector of the present invention comprising an intron containing an RRE can typically be at least about 100%, at least about 500%, at least about 1000%, at least about 5000%, at least about 10000%, at least about 20000%, at least about 50000%, at least about 60000%, or more, compared to the expression of a therapeutic protein from a corresponding retroviral / lentiviral (e.g., SIV) vector without an intron containing an RRE. Preferably, the increase in therapeutic protein expression by a retroviral / lentiviral (e.g., SIV) vector of the present invention containing an intron comprising an RRE is typically at least about 1,000%, more preferably at least about 10,000%, and even more preferably at least about 50,000%, compared to the expression of a therapeutic protein from a corresponding retroviral / lentiviral (e.g., SIV) vector without an intron comprising an RRE. As a non-limiting example, when the RRE-containing intron is a β-globulin / IgG chimeric intron comprising an RRE, such as an intron comprising a β-globulin / IgG chimeric RRE of SEQ ID NO: 5, the intron may increase transgene expression by a retroviral / lentiviral (e.g., SIV) vector of the present invention containing an intron comprising the β-globulin / IgG chimeric RRE by at least 60,000%, e.g., about 68,600%, compared to the expression of a transgene by a corresponding retroviral / lentiviral (e.g., SIV) vector without the intron comprising the β-globulin / IgG chimeric RRE.

[0109] As described herein, the RRE-containing intron typically introduced into the retroviral / lentiviral (e.g., SIV) vector of the present invention is properly spliced ​​by the target cell and contributes to the maturation of stable mRNA molecules. This results in an increase in the number of mRNA molecules (i.e., an increase in mRNA copy number) produced by the retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid. Therefore, the introduction of an RRE-containing intron into a retroviral / lentiviral (e.g., SIV) vector results in increased expression of a transgene that may encode a therapeutic protein. The introduction of an RRE-containing intron into a retroviral / lentiviral (e.g., SIV) vector can increase the number of mRNA molecules produced by the retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid compared to the number of mRNA molecules produced by the corresponding retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid that does not have an RRE-containing intron. Thus, a retroviral / lentiviral (e.g., SIV) vector of the invention that includes an intron containing an RRE typically results in an increased number of mRNA molecules produced by the retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid compared to the number of mRNA molecules produced from a corresponding retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid lacking said RRE-containing intron.As a non-limiting example, a retroviral / lentiviral (e.g., SIV) vector of the invention comprising an AAT transgene (SERPINA1) and a β-globulin / IgG chimeric intron with an inserted RRE (e.g., a β-globulin / IgG chimeric intron with an RRE of SEQ ID NO: 5) may result in an increased number of mRNA molecules produced by the retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid compared to a corresponding retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid that comprises the AAT transgene but lacks the β-globulin / IgG chimeric intron with an inserted RRE (e.g., a β-globulin / IgG chimeric intron with an RRE of SEQ ID NO: 5).

[0110] The increase in the number of mRNA molecules produced by a retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid can be as defined herein. In particular, the increase in therapeutic protein expression by a retroviral / lentiviral (e.g., SIV) vector of the present invention containing an intron comprising an RRE can typically be at least about 2-fold, at least about 5-fold, at least about 7-fold, at least about 10-fold, at least about 12-fold, at least about 15-fold, or more, compared to the number of mRNA molecules produced by a corresponding retroviral / lentiviral (e.g., SIV) vector without an intron comprising an RRE. Preferably, the increase in the number of mRNA molecules produced by a retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid is typically at least about 10-fold, more preferably at least about 12-fold, and even more preferably at least about 13-fold, compared to the number of mRNA molecules produced by a corresponding retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid without an intron comprising an RRE. As a non-limiting example, when the RRE-containing intron is an RRE-containing β-globulin / IgG chimeric intron, such as an intron containing the β-globulin / IgG chimeric RRE of SEQ ID NO: 5, the intron may increase the number of mRNA molecules produced by a retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid by at least 12-fold, e.g., about 13.7-fold, compared to the number of mRNA molecules produced by a corresponding retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid that does not have the β-globulin / IgG chimeric RRE-containing intron.

[0111] As described herein, typically, the intron containing RRE introduced into the retroviral / lentiviral (e.g., SIV) vector of the present invention is properly spliced ​​by target cells, and contributes to the maturation of stable mRNA molecules.This results in an increase in the number of mRNA molecules produced per copy of retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid (i.e., an increase in mRNA copy number).Therefore, the introduction of an intron containing RRE into a retroviral / lentiviral (e.g., SIV) vector results in an increase in the expression of transgenes that can encode therapeutic proteins. Introduction of an RRE-containing intron into a retroviral / lentiviral (e.g., SIV) vector can increase the number of mRNA molecules produced per copy of the retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid compared to the number of mRNA molecules produced per copy of the retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid from a corresponding retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid that does not have an RRE-containing intron. Thus, a retroviral / lentiviral (e.g., SIV) vector of the present invention that includes an RRE-containing intron typically results in an increased number of mRNA molecules produced per copy of the retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid compared to the number of mRNA molecules produced from a corresponding retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid that lacks the RRE-containing intron.As a non-limiting example, a retroviral / lentiviral (e.g., SIV) vector of the invention comprising an AAT transgene (SERPINA1) and a β-globulin / IgG chimeric intron with an inserted RRE (e.g., a β-globulin / IgG chimeric intron with an RRE of SEQ ID NO: 5) may result in an increased number of mRNA molecules produced per copy of the retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid compared to a corresponding retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid that comprises the AAT transgene but lacks the β-globulin / IgG chimeric intron with an inserted RRE (e.g., a β-globulin / IgG chimeric intron with an RRE of SEQ ID NO: 5).

[0112] The increase in the number of mRNA molecules produced per copy of a retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid can be as defined herein. In particular, the increase in the number of mRNA molecules produced per copy of a retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid of the present invention comprising an intron containing an RRE can typically be at least about 10-fold, at least about 20-fold, at least about 30-fold, at least about 40-fold, at least about 45-fold, or more, compared to the number of mRNA molecules produced per copy of a retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid from a corresponding retroviral / lentiviral (e.g., SIV) vector that does not have an intron containing an RRE. Preferably, the increase in the number of mRNA molecules produced per copy of the retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid is typically at least about 20-fold, more preferably at least about 30-fold, and even more preferably at least about 40-fold, as compared to the number of mRNA molecules produced per copy of the retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid from a corresponding retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid that does not have an RRE-containing intron.As a non-limiting example, when the RRE-containing intron is an RRE-containing β-globulin / IgG chimeric intron, such as an intron containing the β-globulin / IgG chimeric RRE of SEQ ID NO: 5, the intron may increase the number of mRNA molecules produced per copy of a retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid by at least 40-fold, for example, about 42.2-fold, compared to the number of mRNA molecules produced per copy of a retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid by a corresponding retroviral / lentiviral (e.g., SIV) vector and / or vector genome plasmid that does not have the β-globulin / IgG chimeric RRE-containing intron.

[0113] The present invention also provides a host cell comprising the retroviral / lentiviral (e.g., SIV) vector of the present invention. Typically, the host cell is a mammalian cell, particularly a human cell or cell line. Non-limiting examples of host cells include HEK293 cells (such as HEK293F cells or HEK293T cells) and 293T / 17 cells. Commercially available cell lines suitable for virus production are also readily available (as described herein).

[0114] Rev response element (RRE) The retroviral / lentiviral (e.g., SIV) vectors of the present invention are designed such that the (RNA) genome of said retroviral / lentiviral (e.g., SIV) vector contains an intron that is not removed during vector manufacturing, such that the final retroviral / lentiviral (e.g., SIV) vector contains said intron, resulting in increased expression of the transgene upon transduction of target cells.

[0115] To facilitate retention of an intron within the genome of a retroviral / lentiviral (e.g., SIV) vector, the endogenous RRE of the retroviral / lentiviral (e.g., SIV) vector is deleted from its location in the wild-type / unmodified retroviral / lentiviral (e.g., SIV) genome and the RRE is inserted into the intron sequence.

[0116] The deletion of the endogenous RRE from its location within the wild-type / unmodified retroviral / lentiviral (e.g., SIV) genome may be a complete or partial deletion, provided that if the deletion is partial, the activity of the remaining RRE sequence is reduced or completely eliminated. Without being bound by theory, it is believed that a partial deletion of the RRE sequence is sufficient, provided that the activity of the remaining RRE sequence is insufficient for gene expression from the retroviral / lentiviral (e.g., SIV) genome and relies on the activity of the RRE within the intron, thus pressuring the retroviral / lentiviral (e.g., SIV) vector to retain the RRE inserted into the intron. Thus, reference herein to the deletion of an endogenous RRE encompasses both a complete and partial deletion of the endogenous RRE. Standard techniques for deleting nucleic acid sequences from nucleic acids (e.g., plasmids) are known in the art and can be readily used by those skilled in the art to delete endogenous RREs.

[0117] Any RRE can be inserted into an intron to be included in the genome of a retroviral / lentiviral (e.g., SIV) vector, provided that the RRE can promote the expression of a retroviral / lentiviral (e.g., SIV) gene in the absence of an endogenous RRE at the standard location in the wild-type / unmodified retroviral / lentiviral (e.g., SIV) genome. Typically, the inserted RRE is a viral RRE, particularly a retroviral RRE, and even more preferably a lentiviral RRE. Standard techniques for inserting a nucleic acid sequence into a nucleic acid (e.g., a plasmid) are known in the art and can be easily used by those skilled in the art to insert an RRE into an intron according to the present invention.

[0118] The RRE inserted into the intron may be the endogenous RRE of the retroviral / lentiviral (e.g., SIV) vector. Thus, the endogenous RRE is deleted from the wild-type / unmodified retroviral / lentiviral (e.g., SIV) genome and inserted into the intron itself introduced into the retroviral / lentiviral (e.g., SIV) vector of the present invention. In other words, the endogenous RRE is moved from its position in the wild-type / unmodified retroviral / lentiviral (e.g., SIV) genome and inserted into the intron. As a non-limiting example, in the SIV vector of the present invention, the endogenous SIV RRE is deleted from the wild-type / unmodified SIV genome, and the intron into which the endogenous SIV RRE is inserted is itself introduced into the SIV vector.

[0119] The RRE inserted into the intron may be an exogenous RRE. In the case of the retroviral vector of the present invention, the exogenous RRE may be an RRE derived from a different retrovirus. In the case of the lentiviral vector of the present invention, the exogenous RRE may be an RRE derived from a different lentivirus. As a non-limiting example, the HIV vector of the present invention has its endogenous HIV RRE deleted and an intron containing an SIV RRE introduced into the HIV genome.

[0120] Preferably, the RRE sequence inserted into an intron in a retroviral / lentiviral (e.g., SIV) vector of the present invention is the same as the endogenous RRE sequence deleted from the retroviral / lentiviral (e.g., SIV) genome. Thus, preferably, the RRE is derived from the same virus as the viral vector, but is inserted into a (chimeric) intron rather than in the canonical location within the viral genome. As a non-limiting example, the viral vector may be an SIV vector in which the SIV RRE has been deleted from the genome and a (chimeric) intron into which the SIV RRE has been inserted has been introduced. As a further non-limiting example, the viral vector may be an HIV vector, and the RRE may be an HIV RRE, but the HIV RRE is inserted into a (chimeric) intron rather than in the canonical location for the HIV RRE within the HIV genome.

[0121] The RRE inserted into the intron may be less than 1,000 bp, preferably less than 900 bp or less than 800 bp. Without being bound by theory, it is believed that an intron containing a smaller RRE allows greater flexibility in terms of additional elements to be included in the retroviral / lentiviral (e.g., SIV) genome and may be more suitable for general applicability. Particularly preferred is an RRE between about 750 bp and about 800 bp, e.g., about 760 bp, such as the exemplified SIV RRE of the present invention.

[0122] The RRE inserted into the intron may be an SIV RRE. The SIV RRE may comprise or consist of a nucleic acid sequence having at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or more, up to 100%, sequence identity to SEQ ID NO:1. Preferably, the SIV RRE comprises or consists of a nucleic acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO:1. More preferably, the SIV RRE consists of a nucleic acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO:1. Even more preferably, the SIV RRE comprises or consists of, and particularly consists of, the nucleic acid sequence of SEQ ID NO:1.

[0123] The RRE inserted into the intron may be an HIV RRE. The HIV RRE may comprise or consist of a nucleic acid sequence having at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or more, up to 100%, sequence identity to SEQ ID NO:50. Preferably, the HIV RRE comprises or consists of a nucleic acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO:50. More preferably, the HIV RRE consists of a nucleic acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO:50. Even more preferably, the HIV RRE comprises or consists of, and in particular consists of, the nucleic acid sequence of SEQ ID NO:50.

[0124] According to the present invention, an RRE inserted into an intron does not form part of the mature mRNA expressed in the host / target cell, as the RRE is typically spliced ​​out as part of the intron.

[0125] Introns The retroviral / lentiviral (e.g., SIV) vectors of the present invention have been modified to (i) delete the endogenous RRE as described herein; and (ii) introduce one or more introns into which the retroviral / lentiviral (e.g., SIV) RRE has been inserted.

[0126] The retroviral / lentiviral (e.g., SIV) vectors of the invention may contain one or more introns, such as one, two, three, four, five, or more introns. Typically, the retroviral / lentiviral (e.g., SIV) vectors of the invention contain one or two introns, preferably one intron.

[0127] A wild-type / unmodified retroviral / lentiviral (e.g., SIV) vector genome does not contain an intron. Thus, any reference herein to a retroviral / lentiviral (e.g., SIV) vector / vector genome containing one or more introns refers to a retroviral / lentiviral (e.g., SIV) vector / vector genome into which one or more introns have been introduced. In other words, any intron contained in a retroviral / lentiviral (e.g., SIV) vector / vector genome of the present invention is an intron that has been introduced into said retroviral / lentiviral (e.g., SIV) vector / vector genome as described herein.

[0128] The size of the one or more introns to be inserted is not particularly limited, provided that the retroviral / lentiviral (e.g., SIV) vector / vector genome containing one or more introns is within the packing limit of the retroviral / lentiviral (e.g., SIV) vector / vector genome. Because the retroviral / lentiviral (e.g., SIV) vector / vector genome contains other elements in addition to introns (including the genome backbone, transgene, and transgene promoter), the upper limit of the size of the one or more introns to be inserted is calculated as follows: UL i =VPL-GE In the formula, UL i is the upper size limit for introns, VPL is the packing limit of the retrovirus / lentivirus (e.g., SIV), and GE is the total size of other retrovirus / lentivirus (e.g., SIV) genomic elements.

[0129] Retroviral / lentiviral (e.g., SIV) vectors typically have a packing limit of about 10 kb, and therefore the size of the inserted intron(s) is typically less than about 5,000 bp, taking into account other elements that must be present in the retroviral / lentiviral (e.g., SIV) genome.

[0130] Introns of less than 1,000 bp, e.g., less than 900 bp or less than 800 bp, may be preferred. Without being bound by theory, it is believed that smaller introns of this type allow for greater flexibility with regard to the additional elements included in the retroviral / lentiviral (e.g., SIV) genome and may be more suitable for general applicability. Particularly preferred are introns between about 750 bp and about 800 bp, e.g., about 770 bp, such as the exemplified β-globulin / IgG chimeric introns of the present invention.

[0131] One or more introns may be introduced into any position within the retroviral / lentiviral (e.g., SIV) genome. Typically, one or more introns are introduced into a position within the retroviral / lentiviral (e.g., SIV) genome that does not disrupt the function of the retroviral / lentiviral (e.g., SIV) genome or any part thereof. As a non-limiting example, one or more introns may be inserted into any position within the retroviral / lentiviral (e.g., SIV) genome, provided that translation of the transgene is not reduced. Preferably, an intron is introduced between a transgene and a promoter operably linked to the transgene. Thus, preferably, the retroviral / lentiviral (e.g., SIV) vector / genome of the present invention comprises an intron between a transgene and a promoter operably linked to the transgene. According to the present invention, an intron containing an RRE does not comprise an expressed transgene.

[0132] The sequence of the introduced intron is not particularly limited. In fact, as exemplified herein, the inventors have shown that retroviral RREs remain functional in different contexts, thereby making the function of the RRE independent of a specific intron sequence. Furthermore, it has been known in the art for decades that non-chimeric introns can be split and joined to other DNA sequences (see, for example, Choi, T. et al. (1990) Molecular and Cellular Biology, 11:6, 3070-3074, incorporated herein by reference). Therefore, any suitable intron can be used in accordance with the present invention. In other words, any intron can have an inserted RRE sequence, and the intron / RRE can be introduced into the retroviral / lentiviral (e.g., SIV) vector / genome of the present invention. The intron can be naturally occurring, recombinant, or artificial, such as a chimeric intron. The intron can also be a viral intron. Non-limiting examples of introns include: SV40 intron, Ef1-α intron 1, CMB intron A, and adenovirus tripartite leader sequence intron. Introns may be chimeric or non-chimeric introns. Preferably, the intron is a chimeric intron, with the chimeric β-globulin / IgG intron exemplified herein being particularly preferred.

[0133] Typically, the intron is a chimeric intron, and thus the retroviral / lentiviral (e.g., SIV) vector of the present invention contains a chimeric intron. A chimeric intron is an artificial intron that contains or consists of sequences from two or more different introns. Non-limiting examples of chimeric introns include a β-globulin / IgG chimeric intron and a chimeric intron from the CAGGS promoter. The latter contains a splice donor derived from chicken β-actin and a splice acceptor derived from rabbit β-globulin. Preferably, the chimeric intron according to the present invention is a β-globulin / IgG chimeric intron as exemplified herein. Particularly preferred are β-globulin / IgG chimeric introns comprising or consisting of a nucleic acid sequence having at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or more, up to 100%, sequence identity to SEQ ID NO: 4. Preferably, the β-globulin / IgG chimeric intron comprises or consists of a nucleic acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 4. More preferably, the β-globulin / IgG chimeric intron consists of a nucleic acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 4. Even more preferably, the β-globulin / IgG chimeric intron comprises or consists of, and particularly consists of, the nucleic acid sequence of SEQ ID NO: 4.

[0134] A chimeric intron from a CAGGS promoter can comprise or consist of a nucleic acid sequence having at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or more, up to 100%, sequence identity to SEQ ID NO: 48. Preferably, a chimeric intron from a CAGGS promoter comprises or consists of a nucleic acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 48. More preferably, a chimeric intron from a CAGGS promoter consists of a nucleic acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 48. Even more preferably, a chimeric intron from a CAGGS promoter comprises or consists of, and in particular consists of, the nucleic acid sequence of SEQ ID NO: 48.

[0135] The RRE is introduced into the intron at a position that does not destroy the intron's splice donor and / or acceptor sites, allowing the intron to be accurately spliced ​​in the target cell. The splice donor and / or acceptor sites of a particular intron can be easily determined using conventional methods and techniques, for example, as described in Desmet et al. (Nucleic Acids Res. 2009 May; 37(9):e67) and Baten et al. (BMC Bioinformatics, Vol. 7, Paper No.: S15(2006)) (both of which are incorporated herein by reference in their entirety).

[0136] In particular, the inventors have found that positioning the RRE within 200 bp 5' of the splice acceptor branch site (and polypyrimidine tract), e.g., within 100 bp 5' of the splice acceptor branch site, within 50 bp 5' of the splice acceptor branch site, or within 20 bp 5' of the splice acceptor branch site, is advantageous because it allows for efficient splicing of the intron to produce retrovirus / lentivirus (e.g., SIV) while still allowing for the insertion of the RRE. In particular, and as exemplified herein, the RRE is preferably inserted within 20 bp 5' of the splice acceptor branch site (and polypyrimidine tract), e.g., 20 bp 5' of the splice acceptor branch site, 19 bp 5' of the splice acceptor branch site, 18 bp or less 5' of the splice acceptor branch site, with 18 bp 5' of the splice acceptor branch site being particularly preferred.

[0137] Typically, there is at least a 5-bp, 10-bp, or 15-bp gap between the RRE insertion site and the splice acceptor branch site to ensure that the branch site is not disrupted by the RRE insertion. Thus, the RRE is inserted about 5-20 bp 5′ from the splice acceptor branch site (and polypyrimidine tract), e.g., about 10-20 bp 5′ from the splice acceptor branch site, or about 15-20 bp 5′ from the splice acceptor branch site, with insertion of the RRE 18 bp 5′ from the splice acceptor branch site being preferred.

[0138] The RRE insertion site devised by the present inventors differs from insertion sites attempted in the art, which are typically closer to the splice donor site than the splice acceptor site.

[0139] The RRE insertion site of the present invention is also typically designed so that any other regulatory sequences within the intron are not disrupted by the RRE insertion.

[0140] If the intron is a chimeric intron, the RRE is inserted at the junction between sequences from two or more different introns. As a non-limiting example, if the chimeric intron is a β-globulin / IgG chimeric intron, the RRE is inserted at the junction between the β-globulin intron sequence (the 5' portion of the chimeric β-globulin / IgG chimeric intron) and the IgG intron sequence (the 3' portion of the β-globulin / IgG chimeric intron).

[0141] In a preferred embodiment involving the use of a β-globulin / IgG chimeric intron, an RRE is inserted between the splice donor site and the splice acceptor site, and (a) the splice donor site is at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or more, up to 100%, with SEQ ID NO:2 (derived from β-globulin). and / or (b) the splice acceptor site comprises or consists of a nucleic acid sequence having at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9% or more, up to 100%, sequence identity to SEQ ID NO:3 (IgG-derived). Preferably, (a) the splice donor site comprises or consists of a nucleic acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO:2; and / or (b) the splice acceptor site comprises or consists of a nucleic acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO:3. More preferably, (a) the splice donor site consists of a nucleic acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 2; and / or (b) the splice acceptor site consists of a nucleic acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 3. Even more preferably, (a) the splice donor site comprises or consists, in particular consists, of the nucleic acid sequence of SEQ ID NO: 2; and / or (b) the splice acceptor site comprises or consists, in particular consists, of the nucleic acid sequence of SEQ ID NO: 3.

[0142] The intron is introduced into the retroviral / lentiviral (e.g., SIV) vector / genome in either the forward or reverse orientation. Preferably, the intron is introduced into the retroviral / lentiviral (e.g., SIV) vector / genome in the forward orientation.

[0143] Preferably, and as exemplified herein, the intron is a β-globin / IgG chimeric intron and the RRE is a SIV RRE. A β-globin / IgG chimeric intron containing an SIV RRE can comprise or consist of a nucleic acid sequence having at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or more, up to 100%, sequence identity to SEQ ID NO: 5. Preferably, a β-globin / IgG chimeric intron containing an SIV RRE comprises or consists of a nucleic acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 5. More preferably, the SIV RRE-containing β-globin / IgG chimeric intron consists of a nucleic acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 5. Even more preferably, the SIV RRE-containing β-globin / IgG chimeric intron comprises or consists of, and in particular consists of, the nucleic acid sequence of SEQ ID NO: 5.

[0144] An intron into which an RRE has been inserted in accordance with the present invention is referred to interchangeably herein as an "RRE-comprising intron," an "intron comprising an RRE," an "intron with an inserted RRE," and an "intron with an introduced RRE."

[0145] Typically, the intron containing RRE in the retroviral / lentiviral (e.g., SIV) vector / genome is properly spliced ​​by the target cell and contributes to the maturation of stable mRNA molecules. This results in increased expression of the coding region of the retroviral / lentiviral (e.g., SIV) genome, including the transgene. Therefore, the introduction of the intron containing RRE into the retroviral / lentiviral (e.g., SIV) vector / genome results in increased expression of the transgene, which may encode a therapeutic protein.

[0146] When inserting an RRE into an intron according to the present invention, it is important that the position at which the RRE is inserted into the intron is carefully defined and controlled. This is because the RRE needs to be inserted at a position where it can still function and where the intron will still be recognized, such as during transcription of the retroviral / lentiviral (e.g., SIV) genome. The exact sequence of the intron is not limited, provided that the position at which the RRE is inserted into the intron is carefully defined and controlled as described herein.

[0147] Therefore, methods for designing and / or producing an intron containing an RRE according to the present invention are also provided. The methods may comprise or consist of (a) identifying splice donor and splice acceptor sequences within the intron; and (b) inserting an RRE into the intron such that the splice donor and split acceptor sequences remain intact. The methods may further comprise one or more steps of deleting an endogenous RRE from a retroviral / lentiviral (e.g., SIV) genome. Standard techniques for inserting and / or deleting nucleic acid sequences from nucleic acids (e.g., plasmids) are known in the art and may be used by those skilled in the art to insert an intron containing an RRE and / or delete an endogenous RRE according to the present invention.

[0148] Transgene and promoter The retroviral / lentiviral (e.g., SIV) vectors of the present invention typically contain a transgene encoding a therapeutic protein. A therapeutic protein is one that may be useful in the treatment or prevention of a disease or condition, such as those described herein. Thus, the retroviral / lentiviral (e.g., SIV) vectors of the present invention contain a transgene encoding a protein that has a therapeutic effect on the disease or condition being treated.

[0149] The retroviral / lentiviral (e.g., SIV) vector of the present invention can comprise a transgene encoding a therapeutic protein that is a functional or wild-type version of the protein present in the patient being treated, which is dysfunctional (whether the dysfunction is congenital or acquired).As used herein, the phrase "congenital dysfunction" refers to a protein that is dysfunctional at birth due to genetic factors, and the phrase "acquired dysfunction" refers to a protein that is dysfunctional due to postnatal environment or other factors.As a non-limiting example, CFTR is an example of a protein that is congenitally dysfunctional in patients with cystic fibrosis.

[0150] Thus, the retroviral / lentiviral (e.g., SIV) vectors of the invention can contain a transgene encoding a therapeutic protein that is a functional or wild-type form of a protein that is present in the patient but is dysfunctional due to a genetic disease, such as a genetic respiratory disease.

[0151] The retroviral / lentiviral (e.g., SIV) vectors of the present invention are useful in the treatment of disease through their use in expressing therapeutic proteins in target cells, which exert their therapeutic effect: (i) within the target cells; (ii) by secretion from said cells into surrounding tissues; or (iii) by secretion from said cells into the circulatory system.

[0152] The retroviral / lentiviral (e.g., SIV) vectors of the invention are pseudotyped (e.g., by pseudotyping with F and HN proteins from a respiratory paramyxovirus, such as Sendai virus) to target airway cells of the respiratory tract, as described herein. Such retroviral / lentiviral (e.g., SIV) vectors of the invention are useful for treating disease through their use in expressing therapeutic proteins in airway cells (i) within the respiratory tract; (ii) for secretion from said cells into the lumen of the respiratory tract; and (iii) for secretion from said cells into the circulatory system.

[0153] The therapeutic protein is selected from: (a) a secreted therapeutic protein, optionally selected from alpha-1-antitrypsin (AAT), factor VIII, surfactant protein B (SFTPB), factor VII, factor IX, factor X, factor XI, von Willebrand factor, granulocyte-macrophage colony-stimulating factor (GM-CSF), surfactant protein C (SP-C), an anti-inflammatory protein (e.g., IL-10 or TGGβ), or a monoclonal antibody, an anti-inflammatory decoy, and a monoclonal antibody against an infectious agent; or (b) CFTR, CSF2RA, CSF2RB, and ATP-binding cassette subfamily member A (ABCA3). Preferred examples of therapeutic proteins include AAT, GM-CSF, FVIII, CFTR, decorin, TRIM72, and ABCA3.

[0154] The transgene may encode: (i) a therapeutic protein secreted into epithelial lining fluid and / or blood; (ii) a therapeutic protein secreted into blood; or (iii) a therapeutic membrane protein. Preferred examples of these classes of transgenes include (i) AAT; (ii) FVIII; and (iii) CFTR.

[0155] In some embodiments, the therapeutic protein is not an antibody, particularly not a monoclonal antibody, and / or not a β-globin gene. In such embodiments, the therapeutic protein is selected from: (a) a secreted therapeutic protein, optionally selected from alpha-1-antitrypsin (AAT), factor VIII, surfactant protein B (SFTPB), factor VII, factor IX, factor X, factor XI, von Willebrand factor, granulocyte-macrophage colony-stimulating factor (GM-CSF), surfactant protein C (SP-C), an anti-inflammatory protein (e.g., IL-10, TGGβ, or TNF-alpha), and an anti-inflammatory decoy; or (b) CFTR, CSF2RA, CSF2RB, and ATP-binding cassette subfamily member A (ABCA3).

[0156] The retroviral / lentiviral (e.g., SIV) vectors of the invention are particularly efficient at driving the expression, secretion, and / or membrane insertion of proteins (e.g., therapeutic proteins described herein) by airway cells. This is particularly the case when the retroviral / lentiviral (e.g., SIV) vector is an F / HN-pseudotyped viral vector of the invention (described herein), which is efficient at targeting cells of the airway epithelium.

[0157] Thus, for therapeutic applications, the retroviral / lentiviral (e.g., SIV) vector of the present invention is typically delivered to cells of the respiratory tract, including cells of the airway epithelium. In other words, the retroviral / lentiviral (e.g., SIV) vector of the present invention is typically delivered to airway cells as described herein. Therefore, the retroviral / lentiviral (e.g., SIV) vector of the present invention is particularly suitable for treating diseases or disorders of the airway, respiratory tract, or lungs. Typically, the retroviral / lentiviral (e.g., SIV) vector of the present invention is used to treat genetic respiratory diseases.

[0158] The retroviral / lentiviral (e.g., SIV) vectors of the invention can contain a transgene encoding a polypeptide or protein that is therapeutic for the treatment of such diseases, particularly airway, respiratory, or lung diseases or disorders. The transgenes and therapeutic proteins of the invention are not limited, and one of skill in the art will be able to identify therapeutic proteins that are usefully delivered in accordance with the invention, particularly in the context of genetic diseases, particularly genetic respiratory diseases, and airway, respiratory, or lung diseases or disorders such as those described herein.

[0159] Thus, the retroviral / lentiviral (e.g., SIV) vectors of the invention can comprise: (a) a secreted therapeutic protein, optionally selected from alpha-1-antitrypsin (AAT), factor VIII, surfactant protein B (SFTPB), ADAMTS13, factor VII, factor IX, factor X, factor XI, von Willebrand factor, granulocyte-macrophage colony-stimulating factor (GM-CSF), surfactant protein C (SP-C), an anti-inflammatory protein (e.g., IL-10, TGGβ, or TNF-alpha), or a monoclonal antibody, an anti-inflammatory decoy, and a monoclonal antibody against an infectious agent; or (b) a nucleic acid sequence encoding a therapeutic protein selected from CFTR, CSF2RA, CSF2RB, ATP-binding cassette subfamily member A (ABCA3), DNAH5, DNAH11, DNAI1, and DNAI2. Other examples of therapeutic proteins encoded by transgenes contained in the retroviral / lentiviral (e.g., SIV) vectors of the present invention include genes related to or associated with other surfactant deficiencies. Preferred examples of therapeutic proteins include AAT, GM-CSF, FVIII, CFTR, ADAMTS13, SFTPB, decorin, TRIM72, and ABCA3.

[0160] The therapeutic protein encoded by the retroviral / lentiviral (e.g., SIV) vector of the present invention may be AAT. An example of an AAT therapeutic transgene (SERPINA1) is represented by SEQ ID NO: 6 or the complementary sequence of SEQ ID NO: 7. SEQ ID NO: 6 is a codon-optimized, CpG-depleted AAT transgene (SERPINA1) previously designed by the present inventors to enhance translation in human cells. Such optimization has been shown to enhance gene expression by up to 15-fold. Variants of the same sequence (as defined herein) having the same technical effect of enhancing translation compared to the unmodified (wild-type) AAT gene sequence are also encompassed by the present invention. The therapeutic protein encoded by the AAT transgene is exemplified by the polypeptide of SEQ ID NO: 8. Variants thereof (described therein) are also included, particularly variants having at least 90% (e.g., at least 90, 92, 94, 95, 96, 97, 98, 99, or 100%) identity with any one of SEQ ID NOs: 6, 7, or 8.

[0161] The therapeutic protein encoded by the retroviral / lentiviral (e.g., SIV) vector of the invention may be FVIII. Examples of FVIII therapeutic transgenes are provided by SEQ ID NOs: 9 and 10, or the complementary sequences of SEQ ID NOs: 11 and 12, respectively. Polypeptides encoded by FVIII transgenes are exemplified by the polypeptides of SEQ ID NOs: 13 or 14. Variants thereof (as described therein) are also included, particularly variants having at least 90% (e.g., at least 90, 92, 94, 95, 96, 97, 98, 99, or 100%) identity with any one of SEQ ID NOs: 9-14.

[0162] Preferably, the therapeutic protein encoded by the retroviral / lentiviral (e.g., SIV) vector of the invention is CFTR. An example of a CFTR transgene is provided by SEQ ID NO: 15. The polypeptide encoded by said CFTR transgene is exemplified by the polypeptide of SEQ ID NO: 16. Variants thereof (as described therein) are also included, particularly variants having at least 90% (e.g., at least 90, 92, 94, 95, 96, 97, 98, 99, or 100%) identity with SEQ ID NO: 15 or 16.

[0163] The therapeutic protein encoded by the retroviral / lentiviral (e.g., SIV) vector of the present invention may be GM-CSF. The GM-CSF transgene may comprise or consist of SEQ ID NO: 17 (human). The polypeptide encoded by the GM-CSF transgene is exemplified by the polypeptide of SEQ ID NO: 18 (human). Variants thereof (as described therein) are also included, particularly variants having at least 90% (e.g., at least 90, 92, 94, 95, 96, 97, 98, 99, or 100%) identity with any one of SEQ ID NOs: 17 and 18.

[0164] The transgene can encode decorin. An example of a DCN transgene is provided by SEQ ID NO: 21. The polypeptide encoded by the DCN transgene is exemplified by the polypeptide of SEQ ID NO: 22. Variants thereof (as described therein) are also included, particularly variants having at least 90% (e.g., at least 90, 92, 94, 95, 96, 97, 98, 99, or 100%) identity with SEQ ID NO: 21 or 22.

[0165] The transgene can encode TRIM72. An example of a TRIM72 transgene is provided by SEQ ID NO: 23. The polypeptide encoded by the TRIM72 transgene is exemplified by the polypeptide of SEQ ID NO: 24. Variants thereof (as described therein) are also included, particularly variants having at least 90% (e.g., at least 90, 92, 94, 95, 96, 97, 98, 99, or 100%) identity with SEQ ID NO: 23 or 24.

[0166] The transgene can encode ABCA3. An example of an ABCA3 transgene is provided by SEQ ID NO: 25. The polypeptide encoded by the ABCA3 transgene is exemplified by the polypeptide of SEQ ID NO: 26. Variants thereof (as described therein) are also included, particularly variants having at least 90% (e.g., at least 90, 92, 94, 95, 96, 97, 98, 99, or 100%) identity with SEQ ID NO: 25 or 26.

[0167] The transgene can encode SFTPB. An example of an SFTPB transgene is provided by SEQ ID NO: 40. The polypeptide encoded by said SFTPB transgene is exemplified by the polypeptide of SEQ ID NO: 41. Variants thereof (as described therein) are also included, particularly variants having at least 90% (e.g., at least 90, 92, 94, 95, 96, 97, 98, 99, or 100%) identity with SEQ ID NO: 40 or 41.

[0168] The transgene can encode ADAMTS13. An example of an ADAMTS13 transgene is provided by SEQ ID NO: 42. The polypeptide encoded by the ADAMTS13 transgene is exemplified by the polypeptide of SEQ ID NO: 43. Variants thereof (as described therein) are also included, particularly variants having at least 90% (e.g., at least 90, 92, 94, 95, 96, 97, 98, 99, or 100%) identity with SEQ ID NO: 42 or 43.

[0169] The therapeutic protein encoded by the retroviral / lentiviral (e.g., SIV) vectors of the present invention may be encoded by one of SFTPB, SFTPC, ADAMTS13, Factor V, Factor VII, Factor IX, Factor X and / or Factor XI, von Willebrand factor, GM-CSF, ABCA3, TRIM72 or DCN, or other known related genes.

[0170] When respiratory epithelial cells are targeted for delivery of the retroviral / lentiviral (e.g., SIV) vectors of the present invention, the therapeutic protein may be AAT, SFTPB, or GM-CSF. The therapeutic protein may be a monoclonal antibody (mAb) against an infectious agent (bacteria, fungi, or virus, e.g., SARS-CoV2 virus). The therapeutic protein may be anti-TNF alpha. The therapeutic protein may be involved in an inflammatory, immune, or metabolic condition.

[0171] The retroviral / lentiviral (e.g., SIV) vectors of the present invention are delivered to cells of the respiratory tract to enable the production of proteins that are secreted into the circulatory system. In such embodiments, the therapeutic protein can be any one of Factor VII, Factor VIII, Factor IX, Factor X, Factor XI, and / or von Willebrand factor. Such retroviral / lentiviral (e.g., SIV) vectors of the present invention are used to treat diseases, particularly cardiovascular diseases and blood disorders, preferably blood clotting disorders such as hemophilia. Again, the therapeutic protein can be a mAb against an infectious agent or a protein involved in an inflammatory, immune, or metabolic condition, such as a lysosomal storage disease.

[0172] Retroviral / lentiviral (e.g., SIV) vectors contain a promoter operably linked to a transgene to enable expression of the transgene. Typically, the promoter is a hybrid human CMV enhancer / EF1a (hCEF) promoter. This hCEF promoter may lack the intron corresponding to nucleotides 570-709 and the exon corresponding to nucleotides 728-733 of the hCEF promoter. A preferred example of an hCEF promoter sequence of the present invention is represented by SEQ ID NO: 27. The promoter may be a CMV promoter. An example of a CMV promoter sequence is represented by SEQ ID NO: 28. The promoter may be a human elongation factor 1a (EF1a) promoter. An example of an EF1a promoter is represented by SEQ ID NO: 29. Other promoters for transgene expression are known in the art, and their suitability for the retroviral / lentiviral (e.g., SIV) vectors of the present invention is determined using routine techniques known in the art. Non-limiting examples of other promoters include UBC and UCOE. As described herein, promoters are modified to further regulate expression of the transgenes of the present invention.

[0173] The promoter contained in the retroviral / lentiviral (e.g., SIV) vector of the present invention is specifically selected and / or modified to further refine the regulation of therapeutic gene expression. Again, suitable promoters and standard techniques for their modification are known in the art. As a non-limiting example, numerous suitable (CpG-free) promoters suitable for use in the present invention are described in Pringle et al. (J. Mol. Med. Berl. 2012, 90(12):1487-96), which is incorporated herein by reference in its entirety. Preferably, the retroviral / lentiviral vector of the present invention (particularly the SIV F / HN vector) comprises an hCEF promoter with low or no CpG dinucleotide content. In the hCEF promoter, all CG dinucleotides may be replaced with any one of AG, TG, or GT. Thus, the hCEF promoter may be CpG-free. A preferred example of a CpG-free hCEF promoter sequence of the present invention is represented by SEQ ID NO: 27. The absence of CpG dinucleotides typically further improves the performance of the retroviral / lentiviral (e.g., SIV) vectors of the invention, particularly in situations where it is undesirable to induce an immune response to the expressed antigen or an inflammatory response to the delivered expression construct. The absence of CpG dinucleotides reduces the occurrence of flu-like symptoms and inflammation that can result from administration of the construct, particularly when administered to the respiratory tract.

[0174] The retroviral / lentiviral (e.g., SIV) vector of the present invention is modified to allow gene expression to be silenced.Standard techniques for modifying vectors in this way are known in the art.As a non-limiting example, Tet-responsive promoters are widely used.

[0175] Retroviral / lentiviral (eg, SIV) vectors of the invention can comprise an hCEF promoter and a CFTR transgene, including those described herein.

[0176] Retroviral / lentiviral (eg, SIV) vectors of the invention can include the hCEF promoter and the AAT transgene (SERPINA1), including those described herein.

[0177] Retroviral / lentiviral (eg, SIV) vectors of the invention can comprise an hCEF or CMV promoter and a FVIII transgene, including those described herein.

[0178] Retroviral / lentiviral (eg, SIV) vectors of the invention can comprise an hCEF or CMV promoter and a DCN transgene, including those described herein.

[0179] Retroviral / lentiviral (eg, SIV) vectors of the invention can comprise an hCEF or CMV promoter and a TRIM72 transgene, including those described herein.

[0180] Retroviral / lentiviral (eg, SIV) vectors of the invention can comprise an hCEF or CMV promoter and an ABCA3 transgene, including those described herein.

[0181] Retroviral / lentiviral (eg, SIV) vectors of the invention can comprise an hCEF or CMV promoter and an SFTPB transgene, including those described herein.

[0182] Retroviral / lentiviral (eg, SIV) vectors of the invention can comprise an hCEF or CMV promoter and an ADAMTS13 transgene, including those described herein.

[0183] The retroviral / lentiviral (e.g., SIV) vectors of the present invention comprise a nucleic acid encoding a therapeutic protein (said nucleic acid is interchangeably referred to herein as a transgene). The nucleic acid sequence encodes a gene product, e.g., a protein, particularly a therapeutic protein.

[0184] For example, a retroviral / lentiviral (e.g., SIV) vector can contain a transgene encoding AAT, GM-CSF, FVIII, SFTPB, ADAMTS13, CFTR, decorin, TRIM72, or ABCA3, wherein the transgene comprises (or consists of) a nucleic acid sequence having at least 90% (e.g., at least 90, 92, 94, 95, 96, 97, 98, 99, or 100%) sequence identity to an AAT, GM-CSF, FVIII, SFTPB, ADAMTS13, CFTR, decorin, TRIM72, or ABCA3 transgene, respectively (examples of which are described herein). In further embodiments, the transgene encoding AAT, GM-CSF, FVIII, SFTPB, ADAMTS13, CFTR, decorin, TRIM72, or ABCA3 comprises (or consists of) a nucleic acid sequence having at least 95% (e.g., at least 95, 96, 97, 98, 99, or 100%) sequence identity to an AAT, GM-CSF, FVIII, SFTPB, ADAMTS13, CFTR, decorin, TRIM72, or ABCA3 nucleic acid sequence, examples of which are described herein, respectively. The nucleic acid sequence encoding CFTR is represented by SEQ ID NO: 15, the nucleic acid sequence encoding AAT is represented by SEQ ID NO: 6 or by the complementary sequence of SEQ ID NO: 7, and / or the nucleic acid sequence encoding FVIII is represented by SEQ ID NO: 11 or 12 or by the complementary sequence of SEQ ID NO: 13 or 14, respectively, and / or the nucleic acid sequence encoding SFTPB is represented by SEQ ID NO: 40, and / or the nucleic acid sequence encoding ADAMTS13 is represented by SEQ ID NO: 42, and / or the nucleic acid sequence encoding GM-CSF is represented by SEQ ID NO: 17, the nucleic acid sequence encoding decorin is represented by SEQ ID NO: 21, the nucleic acid sequence encoding TRIM72 is represented by SEQ ID NO: 23, and / or the nucleic acid sequence encoding ABCA3 is represented by SEQ ID NO: 25, or a variant thereof.

[0185] The amino acid sequence of a therapeutic protein may be a functional variant having at least 95% (e.g., at least 95, 96, 97, 98, 99, or 100%) sequence identity to the functional protein. For example, the AAT, FVIII, SFTPB, ADAMTS13, CFTR, GM-CSF, decorin, TRIM72, and / or ABCA3 polypeptides encoded by the respective AAT, FVIII, SFTPB, ADAMTS13, CFTR, CSF2, DCN, TRIM72, and / or ABCA3 transgenes may comprise (or consist of) an amino acid sequence having at least 95% (e.g., at least 95, 96, 97, 98, 99, or 100%) sequence identity to the functional AAT, FVIII, SFTPB, ADAMTS13, CFTR, GM-CSF, decorin, TRIM72, and / or ABCA3 polypeptide sequence, respectively.

[0186] A transgene encoding a therapeutic protein may or may not contain a nucleic acid sequence encoding the therapeutic protein's endogenous signal peptide. All disclosures herein relate to both transgenes and therapeutic proteins with and without endogenous signal peptides, unless expressly stated. As a non-limiting example, the sequence identity and / or fragment length of variants may be based on the sequence with or without the signal peptide.

[0187] The retroviral / lentiviral (e.g., SIV) vectors of the present invention typically further comprise a Rev protein. This Rev protein is typically provided by (encoded by) one of the plasmids used in the production of retroviral / lentiviral (e.g., SIV) vectors as described herein. As a non-limiting example, the Rev protein is provided by the Rev plasmid (pDNA2b), and when separate plasmids are used to provide the Gag-Pol and Rev proteins, or when a single plasmid is used to provide the Gag-Pol and Rev proteins, the Rev protein is provided by the Rev-Gag-Pol plasmid. An exemplary pDNA2b plasmid described herein is pGM299, as shown in FIG. 2D and having the sequence represented by SEQ ID NO: 33. An exemplary Rev protein is the rSIV Rev protein, which comprises or consists of the amino acid sequence of SEQ ID NO: 44. This Rev protein is encoded by the pGM299 plasmid.

[0188] nucleic acid The present invention also provides nucleic acids comprising or consisting of an intron (e.g., a chimeric intron) into which an RRE has been introduced, as described herein. Any intron described herein is included in the nucleic acids of the present invention. Similarly, any RRE described herein is included in the nucleic acids of the present invention. Each intron and RRE is independently selected, for example, from those described herein. Any disclosure herein relating to an intron and / or RRE in the context of a retroviral / lentiviral (e.g., SIV) vector of the present invention applies equally and without reservation to an intron and / or RRE in the context of a nucleic acid of the present invention.

[0189] As a non-limiting example, a nucleic acid of the invention can include any intron described herein, for example, a β-globin / IgG chimeric intron comprising or consisting of SEQ ID NO: 4 or a variant thereof, as described herein.

[0190] Alternatively, or additionally, as a further non-limiting example, a nucleic acid of the invention can comprise any RRE described herein, e.g., an SIV RRE comprising or consisting of SEQ ID NO: 1 or a variant thereof, as described herein.

[0191] Preferably, and as exemplified herein, the nucleic acids of the invention comprise or consist of a β-globin / IgG chimeric intron and an SIV RRE. Thus, the nucleic acids of the invention can comprise or consist of a β-globin / IgG chimeric intron containing an SIV RRE that comprises or consists of a nucleic acid sequence having at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or more, up to 100%, sequence identity to SEQ ID NO: 5. Preferably, the nucleic acids of the invention can comprise or consist of a β-globin / IgG chimeric intron containing an SIV RRE that comprises or consists of a nucleic acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 5. More preferably, the nucleic acids of the invention may comprise or consist of a β-globin / IgG chimeric intron containing an SIV RRE consisting of a nucleic acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 5. Even more preferably, the nucleic acids of the invention may comprise or consist of a β-globin / IgG chimeric intron containing an SIV RRE comprising or consisting of, and in particular consisting of, the nucleic acid sequence of SEQ ID NO: 5.

[0192] The nucleic acid of the present invention can typically further comprise a transgene encoding a therapeutic protein as described herein.All disclosures herein relating to the transgene in the context of the retroviral / lentiviral (e.g., SIV) vector of the present invention equally and without reservation apply to the transgene in the context of the nucleic acid of the present invention.In some preferred embodiments, the therapeutic protein encoded by the transgene is AAT or CFTR.

[0193] Nucleic acids of the invention that contain an intron containing an RRE may exhibit increased transgene expression compared to corresponding nucleic acids lacking said intron. The disclosures herein regarding increased transgene expression by retroviral / lentiviral (e.g., SIV) vectors of the invention that contain an intron containing an RRE apply equally and without reservation to increased transgene expression exhibited by nucleic acids of the invention.

[0194] As non-limiting examples, increased expression of a therapeutic protein by a nucleic acid of the invention containing an intron containing an RRE can typically be at least about a 5-fold increase, at least about a 10-fold increase, at least about a 50-fold increase, at least about a 100-fold increase, at least about a 200-fold increase, at least about a 500-fold increase, at least about a 600-fold increase, or more, compared to expression of a therapeutic protein from a corresponding nucleic acid that does not have an intron containing an RRE.

[0195] The nucleic acids of the present invention enable long-term transgene expression, resulting in long-term expression of a therapeutic protein. As used herein, the terms "long-term expression," "sustained expression," "long-term sustained expression," and "sustained expression" are used interchangeably. Long-term expression according to the present invention refers to expression of a therapeutic protein, preferably at therapeutic levels, for at least 45 days, at least 60 days, at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 730 days, or more. Preferably, long-term expression refers to expression for at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 720 days, or more, more preferably at least 360 days, at least 450 days, at least 720 days, or more.

[0196] In particular, the nucleic acids of the invention can drive (increased) long-term, sustained expression of a therapeutic protein in airway cells in vivo in a patient. Preferably, the nucleic acids of the invention drive expression of a therapeutic protein in airway cells for at least 45 days, more preferably for at least 90 days.

[0197] The nucleic acid of the nucleic acid may be as defined herein. The nucleic acid may include DNA and / or RNA. Preferably, the nucleic acid is DNA.

[0198] The nucleic acids of the present invention are optionally codon-optimized for expression in a particular cell type, for example, eukaryotic cells (e.g., mammalian cells, yeast cells, insect cells, or plant cells) or prokaryotic cells (e.g., E. coli). The term "codon-optimized" refers to the replacement of at least one codon in a base polynucleotide sequence with a codon that is preferentially used by the host organism in which the polynucleotide is expressed. Typically, the codon most frequently used in the host organism is used in the codon-optimized polynucleotide sequence. Methods of codon optimization are well known in the art.

[0199] It will be understood by those skilled in the art that, as a result of the degeneracy of the genetic code, many different polynucleotides can encode the same polypeptide. It will also be understood that those skilled in the art can, using conventional techniques, make nucleotide substitutions that do not affect the polypeptide sequence encoded by a nucleic acid molecule to reflect the codon usage of any particular host organism in which the polypeptide will be expressed. Thus, unless otherwise specified, nucleic acids encoding the Therapeutic proteins of the invention include all polynucleotide sequences that are degenerate versions of each other and encode the same amino acid sequence.

[0200] The nucleic acid cassette of the present invention typically comprises a promoter operably linked to a nucleic acid sequence encoding a therapeutic protein. By operably linked, we mean that the promoter is configured to express the nucleic acid sequence encoding the signal peptide and / or the nucleic acid sequence encoding the therapeutic protein. The disclosures herein regarding promoters in the context of the retroviral / lentiviral (e.g., SIV) vectors of the present invention equally apply to the nucleic acids of the present invention. In some preferred embodiments, the promoter is the hCEF promoter described herein.

[0201] The nucleic acid of the present invention may comprise at least a portion of a vector, particularly a regulatory element. As a non-limiting example, a promoter (e.g., the hCEFI promoter) in a nucleic acid cassette of the present invention is used to express more than one polypeptide, including one or more therapeutic proteins. Thus, the nucleic acid may contain a nucleic acid sequence that, when transcribed, produces multiple polypeptides; for example, the transcript may contain multiple open reading frames (ORFs) and one or more internal ribosome entry sites (IRES) to allow translation of ORFs after the first ORF. The transcript may be polycistronic; that is, the transcript is translated to produce a polypeptide and then cleaved to produce multiple polypeptides. Alternatively, the nucleic acid of the present invention may contain multiple promoters, thereby producing multiple transcripts and thus multiple polypeptides, including multiple therapeutic proteins. The nucleic acid may express, for example, one, two, three, four, or more polypeptides via a single promoter (e.g., the hCEFI promoter) or multiple promoters.

[0202] A nucleic acid may contain one or more translation initiation sequences (TIS). Translation initiation plays a key role in mRNA translation, typically by identifying the AUG start codon with a unique methionyl-tRNA (Met-tRNAi) to trigger downstream translation processes. Non-canonical start codons (e.g., CUG for valyl-tRNA) / TISs are also used.

[0203] The nucleic acids of the present invention may contain at least one termination signal. A "termination signal" or "terminator" is composed of a DNA sequence involved in the specific termination of an RNA transcript by an RNA polymerase. Thus, termination signals that terminate the production of an RNA transcript are contemplated according to the present invention. A terminator may be necessary in vivo to achieve a desired message level. In eukaryotic systems, the terminator region may also contain a specific DNA sequence that allows site-specific cleavage of the new transcript to expose a polyadenylation site. This signals a specialized endogenous polymerase to add a stretch of approximately 200 A residues (polyA) to the 3' end of the transcript. RNA molecules modified with this polyA tail appear to be more stable and are translated more efficiently. Therefore, when a nucleic acid is intended for expression in eukaryotes, the terminator typically contains a signal for RNA cleavage, and it is preferred that the terminator signal promotes message polyadenylation. Terminator and / or polyadenylation site elements may function to enhance message levels and minimize readthrough from the cassette into other sequences.

[0204] Terminators contemplated for use in the present invention include any known terminator of transcription described herein or known to those skilled in the art, including, but not limited to, gene termination sequences, such as the bovine growth hormone terminator, or viral termination sequences, such as the SV40 terminator. In certain embodiments, the termination signal can be the absence of a transcribable or translatable sequence, such as by sequence truncation.

[0205] The nucleic acid of the present invention can express therapeutic proteins in a given host cell. Any suitable host cell can be used, such as mammalian, bacterial, insect, yeast, and / or plant host cells. In addition, cell-free expression systems can be used. Such expression systems and host cells are standard in the art.

[0206] Typically, the nucleic acid cassettes and vectors of the invention are capable of expressing a therapeutic protein in airway cells, as described herein in connection with the retroviral / lentiviral (e.g., SIV) vectors of the invention.

[0207] The nucleic acid of the present invention can be prepared using any suitable method known in the art.Therefore, the nucleic acid can be prepared using chemical synthesis techniques.Alternatively, the nucleic acid of the present invention can be prepared using molecular biology techniques.

[0208] The nucleic acid of the present invention is used to produce a retroviral / lentiviral (e.g., SIV) vector as described herein.By way of non-limiting example, the nucleic acid of the present invention can be a plasmid used to produce a retroviral / lentiviral (e.g., SIV) vector.The nucleic acid of the present invention is contained in a retroviral / lentiviral (e.g., SIV) vector.

[0209] The nucleic acid of the present invention may be in the form of a DNA vector, such as a DNA plasmid. The vector may also be an RNA vector, such as an mRNA vector or a self-amplifying RNA vector. The DNA and / or RNA vector of the present invention may be expressible in eukaryotic and / or prokaryotic cells.

[0210] Typically, the DNA and / or RNA vectors are expressible in the cells of a subject, eg, the cells of a mammalian or avian subject that is to be immunized.

[0211] Typically, the nucleic acids of the invention are capable of expressing a therapeutic protein in airway cells (as described herein).

[0212] Non-viral vectors of the invention may be phage vectors, such as the AAV / phage hybrid vectors described in Hajitou et al., Cell 2006;125(2) 385-398, which is incorporated herein by reference.

[0213] Nucleic acids of the invention (eg, DNA or RNA vectors) can be designed in silico and then synthesized using conventional polynucleotide synthesis techniques.

[0214] Non-viral plasmids lack the viral genetic material to hijack the body's normal production machinery and therefore cannot replicate in the subject being treated, however, they are capable of replication in suitable host cells such as yeast or bacteria, including E. coli, and particularly airway cells as defined herein.

[0215] As used herein, the term "plasmid" refers to a construct composed of genetic material designed to direct transformation of a targeted cell. A plasmid comprises a plasmid backbone. As used herein, a "plasmid backbone" comprises multiple genetic elements that are positionally and sequentially oriented with other necessary genetic elements so that the nucleic acid within the nucleic acid is transcribed and, if necessary, translated in a transfected cell.

[0216] The plasmid backbone may contain one or more unique restriction sites within the backbone. The plasmid may be capable of autonomous replication in a defined host or organism, allowing the cloned sequences to replicate. The plasmid may confer some well-defined phenotype on the host organism that is selectable or easily detected. The plasmid or plasmid backbone may have a linear or circular configuration. Plasmid components may include, but are not limited to: (1) the plasmid backbone; (2) a sequence encoding a signal peptide; (3) a sequence encoding a therapeutic protein; and (4) a DNA molecule incorporating regulatory elements for transcription, translation, RNA stability, and replication.

[0217] The purpose of plasmids in human gene therapy is to efficiently deliver nucleic acid sequences to cells or tissues and express therapeutic proteins within the cells or tissues. In particular, the purpose of plasmids is to achieve high copy number, avoid potential sources of plasmid instability, and provide a means for plasmid selection. Regarding expression, the nucleic acids of the present invention contain elements necessary for expression of the transgene contained within the nucleic acid. Expression includes efficient transcription of the inserted gene, nucleic acid sequence, or nucleic acid within the plasmid.

[0218] The DNA plasmids can be CpG-free or optimized to reduce CpG dinucleotides as described herein.The DNA plasmids of the present invention are codon-optimized as described herein.

[0219] Methods for preparing plasmid DNA are well known in the art. Typically, they are capable of autonomous replication in a suitable host or producer cell.

[0220] The host cell containing the plasmid (e.g., transformed, transfected, or electroporated with the plasmid) may be prokaryotic or eukaryotic in nature, and may be stably or transiently transformed, transfected, or electroporated with the plasmid. Suitable host cells include bacteria, yeast, fungi, invertebrate, and mammalian cells. Preferably, the host cell is a bacterium; more preferably, it is E. coli.

[0221] The host cells are then used in a method for large-scale production of the plasmid. The cells are grown in an appropriate culture medium under suitable conditions, and the desired plasmid is isolated from the cells, or from the medium in which the cells are growing, by any purification technique well known to those of skill in the art; see, e.g., Sambrook et al., supra.

[0222] Any suitable delivery means can be used to deliver the non-viral vector (such as plasmid) of the present invention to target cells or patients.Suitable delivery means are known in the art and are within the scope of the ordinary skill of those skilled in the art.Non-limiting examples include the use of cationic lipids, polymers (such as polyethyleneimine and poly-L-lysine), and electroporation.

[0223] Preferably, cationic lipids are used to deliver the non-viral (e.g., plasmid) vectors of the present invention to target cells or patients. Non-limiting examples of cationic lipids suitable for use in accordance with the present invention are GL67A and lipofectamine.

[0224] The cationic lipid mixture GL67A is a mixture of three components: GL67 (cholest-5-en-3-ol(3β)-,3-[(3-aminopropyl)[4-[(3-aminopropyl)amino]butyl]carbamate], (CAS number: 179075-30-0)), DOPE (1,2-dioleoyl-sn-glycero-3-phosphoethanolamine), and DMPE-PEG5000 (1,2-dimyristoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)5000]). These components are combined in a molar ratio of 1:2:0.05 to form GL67A. The composition of GL67A and methods for its production, as well as methods for preparing mixtures of GL67A with exemplary non-viral vectors, are disclosed in WO2013 / 061091. The contents of WO2013 / 061091 are incorporated herein by reference in their entirety.

[0225] Lipofectamine consists of a 3:1 mixture of DOSPA (2,3-dioleoyloxy-N-[2(sperminecarboxamido)ethyl]-N,N-dimethyl-1-propaniminium trifluoroacetate) and DOPE.

[0226] The present invention also provides a host cell comprising a nucleic acid (e.g., a plasmid) of the present invention. Typically, the host cell is a mammalian cell, particularly a human cell or cell line. Non-limiting examples of host cells include HEK293 cells (e.g., HEK293F cells or HEK293T cells) and 293T / 17 cells. Commercially available cell lines suitable for virus production are also readily available (described herein).

[0227] Production method Also described herein are methods for producing retroviral / lentiviral (eg, SIV) vectors of the invention.

[0228] The present inventors have previously demonstrated that the use of a codon-optimized gal-pol gene from SIV does not adversely affect the production titer of SIV vectors pseudotyped with hemagglutinin-neuraminidase (HN) and fusion (F) proteins from respiratory paramyxoviruses, and can even result in increased vector titer. This is described in PCT / GB2022 / 050524, the entire contents of which are incorporated herein by reference. Furthermore, the present inventors have shown that retroviral vectors containing retroviral / lentiviral RNA sequences with (i) codon substitutions and (ii) a reduced number of modified retroviral / lentiviral open reading frames (ORFs) do not adversely affect the produced vector titer, transgene expression, and / or retroviral / lentiviral RNA sequence integration into the host / target cell genome, and can even result in increased vector titer, transgene expression, and / or retroviral / lentiviral RNA sequence integration. This is described in UK Application No. 2212472.1, which is incorporated herein by reference in its entirety.

[0229] Here, the present inventors have shown that it is possible to produce a retroviral / lentiviral (e.g., SIV) vector in which the endogenous RRE of the retroviral / lentiviral (e.g., SIV) genome has been deleted and a retroviral / lentiviral (e.g., SIV)-inserted intron has been introduced into the retroviral / lentiviral (e.g., SIV) genome, thereby increasing the expression of the transgene.

[0230] Thus, the present invention provides a method for producing a retroviral / lentiviral (e.g., SIV) vector in which (i) the endogenous RRE of the retroviral / lentiviral (e.g., SIV) genome has been deleted and (ii) an intron has been inserted into the retroviral / lentiviral (e.g., SIV) vector. The retroviral / lentiviral (e.g., SIV) vector is typically pseudotyped with the hemagglutinin-neuraminidase (HN) and fusion (F) proteins from respiratory paramyxoviruses or VSV-G, and contains a promoter and a transgene. Preferably, the retroviral / lentiviral (e.g., SIV) vector is a lentiviral vector, with simian immunodeficiency virus (SIV) vectors being particularly preferred.

[0231] The method of the present invention may be a scalable GMP-compliant method.

[0232] The methods of the present invention allow for the generation of retroviral / lentiviral (e.g., SIV) vectors described herein that exhibit high levels of transgene expression. Typically, the methods of the present invention produce retroviral / lentiviral (e.g., SIV) vectors described herein that exhibit increased transgene expression compared to corresponding retroviral / lentiviral (e.g., SIV) vectors that lack an intron containing an RRE according to the present invention.

[0233] Preferably, the increase in transgene expression by a retroviral / lentiviral (e.g., SIV) vector of the present invention can be at least about 5-fold, at least about 10-fold, at least about 50-fold, at least about 100-fold, at least about 200-fold, at least about 500-fold, at least about 600-fold, or more, compared to the expression of a transgene by a corresponding retroviral / lentiviral (e.g., SIV) vector lacking an RRE-containing intron according to the present invention and produced by the same method. More preferably, the increase in transgene expression by a retroviral / lentiviral (e.g., SIV) vector of the present invention can be at least about 10-fold, more preferably at least about 100-fold, and even more preferably at least about 500-fold, compared to the expression of a transgene by a corresponding retroviral / lentiviral (e.g., SIV) vector lacking an RRE-containing intron according to the present invention and produced by the same method. As a non-limiting example, when the RRE-containing intron is an RRE-containing β-globulin / IgG chimeric intron, such as an intron containing the β-globulin / IgG chimeric RRE of SEQ ID NO: 5, the intron may increase transgene expression by a retrovirus / lentivirus (e.g., SIV) of the present invention that contains the β-globulin / IgG chimeric RRE-containing intron by at least 600-fold, e.g., about 686-fold, compared to transgene expression by a corresponding retrovirus / lentivirus (e.g., SIV) vector that does not have the β-globulin / IgG chimeric RRE-containing intron.

[0234] The methods of the present invention typically allow for the production of retroviral / lentiviral (e.g., SIV) vectors containing modified retroviral / lentiviral (e.g., SIV) RNA sequences, with high levels of vector integration into the host / target cell genome. Alternatively, or in addition, the methods of the present invention may allow for the production of high-titer purified retroviral / lentiviral (e.g., SIV) vectors containing modified retroviral / lentiviral (e.g., SIV) RNA sequences. These advantageous properties of the vectors and methods of the present invention are as described in UK Application No. 2212472.1, the entire contents of which are incorporated herein by reference.

[0235] Production of retroviral / lentiviral (e.g., SIV) vectors typically utilizes one or more plasmids that provide the elements necessary for vector production: the genome, Gag-Pol, Rev, F, and HN for the retroviral / lentiviral vector. Multiple elements are provided on a single plasmid. Preferably, each element is provided on a separate plasmid, so that there are five plasmids, one each for the vector genome, Gag-Pol, Rev, F, and HN. Alternatively, a single plasmid may provide the Gag-Pol and Rev elements and is referred to as the packaging plasmid (pDNA2). The remaining elements (genome, F, and HN) are provided by separate plasmids (pDNA1, pDNA3a, and pDNA3b, respectively), so that four plasmids are used to produce retroviral / lentiviral (e.g., SIV) vectors according to the present invention. In the four-plasmid method, pDNA1, pDNA3a, and pDNA3b may be as described herein in the context of the five-plasmid method.

[0236] For retroviral / lentiviral (e.g., SIV) vectors pseudotyped with an alternative envelope protein, such as VSV-G, rather than F and HN proteins, the methods of the present invention again typically utilize one or more plasmids providing the elements necessary for vector production: the genome for the retroviral / lentiviral vector, Gag-Pol (pDNA2a), Rev (pDNA2b), and envelope (e.g., VSV-G) (pDNA3). Multiple elements are provided on a single plasmid. Preferably, each element is provided on a separate plasmid, so that there are four plasmids, one each for the vector genome, Gag-Pol, Rev, and envelope (e.g., VSV-G). In a four-plasmid method for VSV-G-pseudotyped retroviral / lentiviral vectors, pDNA1, pDNA2a, and pDNA2b may be as described herein in the context of the five-plasmid method for retroviral / lentiviral vectors pseudotyped with F and HN proteins. Alternatively, a single plasmid may provide the Gag-Pol and Rev elements and is referred to as the packaging plasmid (pDNA2). The remaining elements (genome and VSV-G) are provided by separate plasmids (pDNA1 and pDNA3, respectively), so that three plasmids are used for the production of retroviral / lentiviral (e.g., SIV) vectors according to the present invention. In the three-plasmid method, pDNA1 can be as described herein in the context of the five / four-plasmid method.

[0237] Preferably, the vector genome plasmid encodes all genetic material to be packaged into the final retroviral / lentiviral vector, including the transgene. The vector genome plasmid is designated herein as "pDNA1" and typically contains the transgene and the transgene promoter. As described herein, the intron containing the RRE is typically contained within the vector genome plasmid. The intron containing the RRE is typically integrated from the vector genome plasmid into the retroviral / lentiviral (e.g., SIV) RNA sequence.

[0238] The other four plasmids are production plasmids encoding the Gag-Pol, Rev, F, and HN proteins. These plasmids are designated "pDNA2a," "pDNA2b," "pDNA3a," and "pDNA3b," respectively.

[0239] Typically, the lentivirus is SIV, such as SIV1, preferably SIV-AGM. The F and HN proteins are derived from a respiratory paramyxovirus, preferably Sendai virus.

[0240] In a specific embodiment for AAT, the five plasmids are characterized by Figures 2A-2F, where pDNA1 is the pGM991 plasmid of Figure 2A, pDNA2a is the pGM691 plasmid of Figure 2B or the pGM297 plasmid of Figure 2C, pDNA2b is the pGM299 plasmid of Figure 2D, pDNA3a is the pGM301 plasmid of Figure 2E, and pDNA3b is the pGM303 plasmid of Figure 2F, or a variant thereof of any of these plasmids (described herein). pGM407 (shown in Figure 2G) is an unmodified version of the vector genome plasmid from which pGM991 is derived.

[0241] The plasmid defined in Figure 2A is represented by SEQ ID NO:30; the plasmid defined in Figure 2B is represented by SEQ ID NO:31; the plasmid defined in Figure 2C is represented by SEQ ID NO:32; the plasmid defined in Figure 2D is represented by SEQ ID NO:33; the plasmid defined in Figure 2E is represented by SEQ ID NO:34; the plasmid defined in Figure 2F is represented by SEQ ID NO:35; and the plasmid defined in Figure 2G is represented by SEQ ID NO:36. Variants of these plasmids (as defined herein) are also encompassed by the present invention. In particular, variants having at least 90% (e.g., at least 90, 92, 94, 95, 96, 97, 98, 99, 99.5, or 100%) sequence identity to any one of SEQ ID NOs:30-36 are encompassed.

[0242] In each of the three-, four-, or five-plasmid methods of the present invention, all of the plasmids contribute to the formation of the final retroviral / lentiviral (e.g., SIV) vector, but only the vector genome plasmid provides the nucleic acid sequences contained in the retroviral / lentiviral (e.g., SIV) RNA sequence. During the production of a retroviral / lentiviral (e.g., SIV) vector, the vector genome plasmid (pDNA1) provides the enhancer / promoter, Psi, intron including RRE, cPPT, mWPRE, SIN LTR, and SV40 polyA (see Figure 1A), which are important for viral production. Using pGM991 as a non-limiting example of pDNA1, the CMV enhancer / promoter, SV40 polyA, colE1 Ori, and KanR are involved in the production of the retroviral / lentiviral (e.g., SIV) vector of the present invention, but are not found in the final retroviral / lentiviral (e.g., SIV) vector. From pGM991, the cPPT (central polypurine tract), RRE-containing intron (inserted between the hCEF and AAT transgenes), hCEF, AAT (transgene), and mWPRE are found in the final retroviral / lentiviral (e.g., SIV) vector. The SIN LTR (long terminal repeat, SIN / IN self-inactivating), and Psi (packaging signal) are found in the final retroviral / lentiviral (e.g., SIV) vector. In contrast to pGM991, pGM407 (from which pGM991 is derived) lacks the RRE-containing intron but contains the endogenous SIV RRE located 5' to the hCEF promoter and between the partial Gag sequence and the cPPT sequence.

[0243] For other retroviral / lentiviral (e.g., SIV) vectors of the invention, corresponding elements from other vector genome plasmids (pDNA1) are required for production (but are not found in the final vector) or are present in the final retroviral / lentiviral (e.g., SIV) vector.

[0244] In a specific embodiment for pseudotyping with VSV-G, the four plasmids are characterized by Figures 2A-2F, whereby pDNA1 is the pGM991 plasmid of Figure 2A, pDNA2a is the pGM691 plasmid of Figure 2B or the pGM297 plasmid of Figure 2C, pDNA2b is the pGM299 plasmid of Figure 2D, and pDNA3 is the pMD2.G plasmid of Figure 2H, or variants of these plasmids (as described herein). The plasmid defined in Figure 2A is represented by SEQ ID NO:30; the plasmid defined in Figure 2B is represented by SEQ ID NO:31; the plasmid defined in Figure 2C is represented by SEQ ID NO:32; the plasmid defined in Figure 2D is represented by SEQ ID NO:33; and the plasmid defined in Figure 2H is represented by SEQ ID NO:49. Variants of these plasmids (as defined herein) are also encompassed by the present invention. In particular, variants having at least 90% (eg, at least 90, 92, 94, 95, 96, 97, 98, 99, 99.5, or 100%) sequence identity to any one of SEQ ID NOs: 30-33 and 49 are encompassed.

[0245] The F and HN proteins (preferably Sendai F and HN proteins) from pDNA3a and pDNA3b, or VSV-G from pDNA3, are important for the final retroviral / lentiviral (e.g., SIV) vector to infect target cells, i.e., for entry into the patient's epithelial cells (typically lung or nasal cells, as described herein). The products of the pDNA2a and pDNA2b plasmids (or pDNA2, when the Gag-Pol and Rev elements are combined in a single plasmid) are important for viral transduction, i.e., for inserting the retroviral / lentiviral (e.g., SIV) DNA into the host genome. The promoter, regulatory elements (e.g., WPRE), and transgene are important for transgene expression in target cells.

[0246] The methods of the present invention may comprise or consist of the following steps: (a) growing cells in suspension; (b) transfecting the cells with one or more plasmids; (c) adding a nuclease; (d) harvesting the retrovirus / lentivirus (e.g., SIV); (e) adding trypsin; and (f) purifying the retrovirus / lentivirus (e.g., SIV).

[0247] This method can use the three-, four-, or five-plasmid systems described herein. Thus, for the five-plasmid method, one or more plasmids can comprise or consist of: vector genome plasmid pDNA1; Gag-Pol plasmid (e.g., a codon-optimized Gag-Pol plasmid), pDNA2a; Rev plasmid, pDNA2b; fusion (F) protein plasmid, pDNA3a; and hemagglutinin-neuraminidase (HN) plasmid, pDNA3b. pDNA1 can be pGM991. pDNA2a can be pGM297 or pGM691, preferably pGM691. pDNA2b can be pGM299. pDNA3a can be pGM301. pDNA3b can be pGM303. Any combination of pDNA1, pDNA2a, pDNA2b, pDNA3a, and pDNA3b can be used. Preferably, pDNA1 is pGM991; pDNA2a is pGM691; pDNA2b is pGM299; pDNA3a is pGM301; and pDNA3b is pGM303.

[0248] For the four-plasmid method, the one or more plasmids may comprise or consist of: vector genome plasmid pDNA1; Gag-Pol plasmid (e.g., codon-optimized Gag-Pol plasmid), pDNA2a; Rev plasmid, pDNA2b; and VSV-G plasmid, pDNA3. pDNA1 may be pGM991. pDNA2a may be pGM297 or pGM691, preferably pGM691. pDNA2b may be pGM299. pDNA3 may be pMD2.G. Any combination of pDNA1, pDNA2a, pDNA2b, and pDNA3 may be used. Preferably, pDNA1 is pGM991; pDNA2a is pGM691; pDNA2b is pGM299; pDNA3a is pGM301; and pDNA3 is pMD2.G.

[0249] Any appropriate ratio of vector genome plasmid:Gag-Pol plasmid:Rev plasmid:F plasmid:HN plasmid can be used to further optimize (increase) the titer of the resulting retrovirus / lentivirus (e.g., SIV). As a non-limiting example, the ratio of vector genome plasmid:Gag-Pol plasmid:Rev plasmid:F plasmid:HN plasmid can range from 10:40:-4 to 20:3 to 12:3 to 12:3 to 12:12, typically 15:20:7 to 11:4 to 8:4 to 8:4 to 8, e.g., approximately 18:22:7 to 11:4 to 8:4 to 8:4 to 8, e.g., approximately 19:21:8 to 10:5 to 7:5 to 7:5 to 7:7. Preferably, the ratio of vector genome plasmid:Gag-Pol plasmid:Rev plasmid:F plasmid:HN plasmid is approximately 20:9:6:6:6. Preferably, the ratio of vector genome plasmid:Gag-Pol plasmid:Rev plasmid:VSV-G plasmid is about 20:9:6:12.

[0250] Steps (a) through (f) of the method are typically performed sequentially, starting with step (a) and continuing through step (f). The method may include one or more additional steps, such as an additional purification step, buffer exchange, concentration of the purified retroviral / lentiviral (e.g., SIV) vector, and / or formulation of the purified (or concentrated) retroviral / lentiviral (e.g., SIV) vector. Each step may include one or more substeps. For example, the harvesting step may include one or more steps or substeps, and / or the purification may include one or more steps or substeps.

[0251] To produce the retroviral / lentiviral (e.g., SIV) vector of the present invention, any suitable cell type is transfected with one or more plasmids (e.g., 5, 4, or 3 plasmids described herein).Typically, mammalian cells, particularly human cell lines, are used.Non-limiting examples of cells suitable for use in the method of the present invention are HEK293 cells (e.g., HEK293F cells or HEK293T cells) and 293T / 17 cells.Commercially available cell lines suitable for virus production are also readily available (e.g., Gibco Viral Production Cells from ThermoFisher Scientific - Catalog No. A35347).

[0252] The cells can be grown in animal component-free media, including serum-free media. The cells can be grown in media containing human components. The cells can be grown in defined media that contain or consist of synthetically produced components.

[0253] Any suitable transfection means can be used in accordance with the present invention. The selection of a suitable transfection means is within the routine practice of those skilled in the art. As a non-limiting example, transfection can be carried out using PEIPro™, Lipofectamine 2000™, or Lipofectamine 3000™.

[0254] Any suitable nuclease can be used in accordance with the present invention. The selection of a suitable nuclease is within the routine practice of those skilled in the art. Typically, the nuclease is an endonuclease. As a non-limiting example, the nuclease may be Benzonase® or Denarase®. The addition of the nuclease may be at a pre-collection stage, a post-collection stage, or during the collection step.

[0255] The gag-pol gene used to produce the retroviral / lentiviral (e.g., SIV) vectors of the present invention may be codon-optimized. Thus, the gag-pol gene in the pDNA2a plasmid may be codon-optimized. As a non-limiting example, the codon-optimized gag-pol gene may comprise or consist of the nucleic acid sequence of SEQ ID NO: 37, or a variant thereof (as defined herein). In particular, the codon-optimized gag-pol gene of the present invention may comprise or consist of a nucleic acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more sequence identity to SEQ ID NO: 37, preferably at least 95% identity to SEQ ID NO: 37. The codon-optimized gag-pol gene may consist of the nucleic acid sequence of SEQ ID NO: 37. A preferred pDNA2a, pGM691, contains the codon-optimized gag-pol gene of SEQ ID NO: 37.

[0256] Gag-pol genes (e.g., SIV gag-pol genes), including codon-optimized gag-pol genes, are typically operably linked to a promoter to promote expression of the gag-pol protein. Any suitable promoter can be used, including those described herein in the context of promoters for transgenes. Preferably, the promoter is a CAG promoter, such as that used in the exemplified pGM691 plasmid. An exemplary CAG promoter is set forth in SEQ ID NO: 38. The codon-optimized gag-pol gene of SEQ ID NO: 37 does not form a single conventional open reading frame because it contains translation slippage.

[0257] Codon-optimized gag-pol genes (or nucleic acids comprising or consisting of them) and plasmids comprising said genes or nucleic acids are advantageous in producing retroviral / lentiviral (e.g., SIV) vectors using the methods of the present invention, since they enable the production of high-titer retroviral / lentiviral (e.g., SIV) vectors. Typically, said codon-optimized gag-pol genes (or nucleic acids comprising or consisting of them) and plasmids comprising said genes or nucleic acids are used to produce retroviral / lentiviral (e.g., SIV) vectors with titers at least equivalent to those of retroviral / lentiviral (e.g., SIV) vectors produced by corresponding methods that do not use codon-optimized gag-pol genes, as described herein. Codon-optimized gag-pol genes are further disclosed in PCT / GB2022 / 050524, which is incorporated herein by reference in its entirety.

[0258] The present invention also provides retroviral / lentiviral (eg, SIV) vectors obtainable by the methods of the present invention.

[0259] Typically, retroviral / lentiviral (e.g., SIV) vectors obtainable by the methods of the present invention are produced at high titers, as described herein, where titers are measured in terms of transducing units, as defined herein.

[0260] Thus, retroviral / lentiviral (e.g., SIV) vectors of the invention, including those obtainable by the methods of the invention, optionally contain at least about 2.5 x 10 6 TU / mL, at least approximately 3.0 × 10 6 TU / mL, at least approximately 3.1 × 10 6 TU / mL, at least approximately 3.2 × 10 6 TU / mL, at least approximately 3.3 × 10 6 TU / mL, at least approximately 3.4 × 10 6 TU / mL, at least approximately 3.5 × 10 6 TU / mL, at least approximately 3.6 × 10 6 TU / mL, at least approximately 3.7 × 10 6 TU / mL, at least approximately 3.8 × 10 6 TU / mL, at least approximately 3.9 × 10 6 TU / mL, at least approximately 4.0 × 10 6 Preferably, the retroviral / lentiviral (e.g., SIV) vector has a titer of at least about 3.0 x 10 TU / mL. 6 TU / mL, or at least approximately 3.5 × 10 6 It is produced at a titer of TU / mL.

[0261] High-titer production of retroviral / lentiviral (e.g., SIV) vectors can confer other desirable properties to the resulting vector product. For example, without being bound by theory, high-titer production without the need for intensive enrichment by methods such as TFF is believed to result in higher-quality vector products than retroviral / lentiviral (e.g., SIV) vectors produced by corresponding methods without the use of codon-optimized gag-pol genes (and, optionally, modified vector genome plasmids). This is because the vectors are exposed to fewer shear forces that can damage viral particles and their RNA cargo.

[0262] Typically, the gag-pol gene (e.g., a codon-optimized gag-pol gene) used is adapted to the retroviral / lentiviral vector being produced. As a non-limiting example, if the lentiviral vector is an HIV vector, the codon-optimized gag-pol gene used is the HIV gag-pol gene. As a non-limiting example, if the lentiviral vector is an SIV vector, the codon-optimized gag-pol gene used is the SIV gag-pol gene.

[0263] Preferably, the codon-optimized gag-pol gene used is the SIV gag-pol gene.

[0264] As described herein, the retroviral / lentiviral (e.g., SIV) vectors of the present invention (i) lack an endogenous RRE; and (ii) contain an intron containing an RRE. Thus, the vector genome plasmid used to produce the retroviral / lentiviral (e.g., SIV) vectors of the present invention is modified to (i) delete the endogenous RRE and (ii) introduce an intron containing an RRE. Any disclosure herein relating to a retroviral / lentiviral (e.g., SIV) vector that (i) lacks an endogenous RRE and (ii) contains an intron containing an RRE applies equally and without reservation to the vector genome plasmid (pDNA1) described herein used to produce the retroviral / lentiviral (e.g., SIV) vectors of the present invention.

[0265] As used herein, the term "trypsin" refers to both trypsin and its equivalents. Equivalent enzymes are those that have the same or essentially the same cleavage specificity as trypsin. Trypsin cleavage activity is defined as cleavage at the C-terminal side of arginine or lysine residues, and typically cleaves only at the C-terminal side of arginine or lysine residues. Trypsin activity is preferably provided by a recombinant enzyme that is free of animal origin, such as TrypLE Select™. Addition of trypsin may occur at a pre-harvest stage, a post-harvest stage, or during the harvesting process.

[0266] Any suitable purification means can be used to purify retroviral / lentiviral (e.g., SIV) vectors. Non-limiting examples of suitable purification steps include depth / end filtration, tangential flow filtration (TFF), and chromatography. A purification step typically includes at least one chromatography step. Non-limiting examples of chromatography steps used in accordance with the present invention include mixed-mode size exclusion chromatography (SEC) and / or anion exchange chromatography. Elution can be performed with or without the use of a salt gradient, preferably without.

[0267] This method is used to produce retroviral / lentiviral (e.g., SIV) vectors of the invention, such as those containing the CFTR, AAT, and / or FVIII genes described herein. Alternatively, the retroviral / lentiviral (e.g., SIV) vectors of the invention contain any of the above-mentioned genes or genes encoding the above-mentioned proteins.

[0268] The method can use any combination of one or more of the specific plasmid constructs provided by Figures 2A-2F or 2H to provide the retroviral / lentiviral (e.g., SIV) vectors of the invention. In particular, the plasmid constructs of Figures 2A, 2B, and 2D-2F, or 2A, 2B, and 2D or 2H are used.

[0269] treatment index The retroviral / lentiviral (e.g., SIV) vectors and nucleic acids (e.g., plasmids) of the present invention enable high, sustained transgene expression. The retroviral / lentiviral (e.g., SIV) vectors and nucleic acids (e.g., plasmids) of the present invention, particularly F / HN-pseudotyped retroviral / lentiviral (e.g., SIV) vectors, enable: (i) airway transduction without disrupting epithelial integrity; (ii) sustained gene expression; (iii) lack of chronic toxicity; and (iv) efficient repeated administration. Preferably, long-term / sustained stable gene expression at therapeutically effective levels is achieved using repeated administration of the vectors of the present invention. Alternatively, a single administration is used to achieve the desired long-term expression.

[0270] Therefore, the retroviral / lentiviral (e.g., SIV) vectors and nucleic acids (e.g., plasmids) of the present invention can be advantageously used in gene therapy. For example, the efficient airway cell uptake properties of the retroviral / lentiviral (e.g., SIV) vectors and nucleic acids (e.g., plasmids) of the present invention make them highly suitable for treating respiratory diseases. The retroviral / lentiviral (e.g., SIV) vectors and nucleic acids (e.g., plasmids) of the present invention can also be used in gene therapy methods to promote the secretion of therapeutic proteins. As a further example, the present invention provides for the secretion of therapeutic proteins into the respiratory lumen or circulatory system. Thus, administration of the retroviral / lentiviral (e.g., SIV) vectors or nucleic acids (e.g., plasmids) of the present invention and their uptake by airway cells can enable the use of the lungs (or nasal cavity or airways) as a "factory" to produce therapeutic proteins, which are then secreted and enter the systemic circulation at therapeutic levels, where they can migrate to cells / tissues of interest to induce a therapeutic effect. In contrast to intracellular or membrane proteins, the production of such secreted proteins does not depend on the specific disease target cells to be transduced, which is a significant advantage, achieving high levels of protein expression. Therefore, other diseases that are not respiratory diseases, such as cardiovascular diseases and blood disorders, particularly blood coagulation disorders, can also be treated by the retroviral / lentiviral (e.g., SIV) vectors and nucleic acids (e.g., plasmids) of the present invention.

[0271] The retroviral / lentiviral (e.g., SIV) vectors and nucleic acids (e.g., plasmids) of the present invention can effectively treat diseases by providing transgenes to correct the disease. For example, inserting a functional copy of the CFTR gene to improve or prevent lung disease in CF patients, regardless of the underlying mutation. Thus, the retroviral / lentiviral (e.g., SIV) vectors and nucleic acids (e.g., plasmids) of the present invention are typically used to treat cystic fibrosis (CF) by gene therapy with the CFTR transgene described herein.

[0272] As another example, the retroviral / lentiviral (e.g., SIV) vectors and nucleic acids (e.g., plasmids) of the present invention are typically used to treat alpha-1 antitrypsin (AAT) deficiency by gene therapy with the AAT transgene described herein. AAT is a secreted antiprotease that is primarily produced in the liver and then transported in small amounts to the lungs, where it is also produced. The main function of AAT is to bind to and neutralize / inhibit neutrophil elastase. Gene therapy with AAT according to the present invention is applicable not only to patients with AAT deficiency, but also to other lung diseases such as CF or chronic obstructive pulmonary disease (COPD), and offers the opportunity to overcome some of the problems encountered with conventional enzyme replacement therapy (in which AAT is isolated from human blood and administered intravenously weekly), resulting in stable, long-term expression in target tissues (lung / nasal epithelium), ease of administration, and unlimited availability.

[0273] Transduction by retroviral / lentiviral (e.g., SIV) vector of the present invention or transfection by nucleic acid (e.g., plasmid) of the present invention can result in the secretion of recombinant protein into the lung lumen and circulation.One advantage of this is that therapeutic protein reaches the interstitium.Therefore, AAT gene therapy can also be beneficial in other disease indications, including, but not limited to, type 1 and type 2 diabetes, acute myocardial infarction, ischemic heart disease, rheumatoid arthritis, inflammatory bowel disease, transplant rejection, graft-versus-host (GvH) disease, multiple sclerosis, liver disease, cirrhosis, vasculitis, and infectious diseases such as bacterial and / or viral infections.

[0274] AAT has numerous other anti-inflammatory and tissue protective effects, for example, in preclinical models of diabetes, graft-versus-host disease, and inflammatory bowel disease. Thus, production of AAT in the lung and / or nasal cavity after transduction according to the present invention may be more broadly applicable, including to these indications.

[0275] Other examples of diseases that may be treated by secreted protein gene therapy according to the present invention include cardiovascular diseases and blood disorders, particularly blood clotting disorders such as hemophilia (A, B, or C), von Willebrand disease, and factor VII deficiency.

[0276] Other examples of diseases or disorders to be treated include inflammatory, infectious, immune, or metabolic conditions, such as primary ciliary dyskinesia (PCD), acute lung injury, surfactant protein B (SFTB) deficiency, pulmonary alveolar proteinosis (PAP), chronic obstructive pulmonary disease (COPD), and / or lysosomal storage diseases.

[0277] Therefore, the present invention provides a method for treating a disease, comprising administering to a subject a retroviral / lentiviral (e.g., SIV) vector or nucleic acid (e.g., a plasmid) of the present invention. Typically, the retroviral / lentiviral (e.g., SIV) vector or nucleic acid (e.g., a plasmid) is produced using the method of the present invention. Any disease described herein can be treated according to the present invention. In particular, the present invention provides a method for treating a lung disease using a retroviral / lentiviral (e.g., SIV) vector or nucleic acid (e.g., a plasmid) of the present invention. The disease to be treated can be a chronic disease. Preferably, a method for treating CF is provided.

[0278] The present invention also provides a retroviral / lentiviral (e.g., SIV) vector or nucleic acid (e.g., a plasmid) described herein for use in a method for treating a disease. Typically, the retroviral / lentiviral (e.g., SIV) vector or nucleic acid (e.g., a plasmid) is produced using the method of the present disclosure. Any disease described herein can be treated according to the present invention. In particular, the present invention provides a retroviral / lentiviral (e.g., SIV) vector or nucleic acid (e.g., a plasmid) of the present invention for use in a method for treating a lung disease. The disease to be treated can be a chronic disease. Preferably, a retroviral / lentiviral (e.g., SIV) vector or nucleic acid (e.g., a plasmid) is provided for use in treating CF.

[0279] The present invention also provides the use of a retroviral / lentiviral (e.g., SIV) vector or nucleic acid (e.g., a plasmid) described herein in the manufacture of a medicament for use in a method for treating a disease. Typically, the retroviral / lentiviral (e.g., SIV) vector or nucleic acid (e.g., a plasmid) is produced using the method of the present disclosure. Any disease described herein can be treated according to the present invention. In particular, the present invention provides the use of a retroviral / lentiviral (e.g., SIV) vector or nucleic acid (e.g., a plasmid) of the present invention in the manufacture of a medicament for use in a method for treating a lung disease. The disease to be treated can be a chronic disease. Preferably, the use of a retroviral / lentiviral (e.g., SIV) vector or nucleic acid (e.g., a plasmid) in the manufacture of a medicament for use in a method for treating CF is provided.

[0280] Formulation and Administration The retroviral / lentiviral (e.g., SIV) vectors and / or nucleic acids (e.g., plasmids) of the present invention are administered at any dosage appropriate to achieve the desired therapeutic effect. Appropriate dosages are determined by clinicians or other medical professionals using standard techniques within their normal course of practice. Non-limiting examples of suitable dosages for retroviral / lentiviral (e.g., SIV) vectors include 1 x 10 8 Transducing units (TU), 1 × 10 9 TU, 1×10 10 TU, 1×10 11 TU or higher.

[0281] The present invention also provides compositions comprising the above-described retroviral / lentiviral (e.g., SIV) vectors and / or nucleic acids (e.g., plasmids) and a pharmaceutically acceptable carrier. Non-limiting examples of pharmaceutically acceptable carriers include water, saline, and phosphate-buffered saline. However, in some embodiments, the compositions are in lyophilized form, in which case the compositions may contain stabilizers such as bovine serum albumin (BSA). In some embodiments, it may be desirable to formulate the compositions with preservatives such as thiomersal or sodium azide to facilitate long-term storage.

[0282] The retroviral / lentiviral (e.g., SIV) vectors and / or nucleic acids (e.g., plasmids) of the present invention are administered by any suitable route. It is desirable to direct the compositions of the present invention (as defined above) to the respiratory system of a subject. Efficient delivery of the therapeutic / prophylactic composition or medicament to the site of infection in the respiratory tract is achieved, for example, as an aerosol (e.g., nasal spray), by oral or intranasal administration, or by catheter. Typically, the retroviral / lentiviral (e.g., SIV) vectors and / or nucleic acids (e.g., plasmids) of the present invention are stable in clinically relevant nebulizers, inhalers (including metered-dose inhalers), catheters, aerosols, and the like. Thus, typically, the retroviral / lentiviral (e.g., SIV) vectors and / or nucleic acids (e.g., plasmids) of the present invention are formulated for pulmonary administration by any suitable means, for example, they are formulated for intratracheal administration, intranasal administration, aerosol delivery, or direct injection or delivery to the lungs (e.g., delivered by catheter). Other modes of delivery, such as intravenous delivery, are also encompassed by the present invention.

[0283] In some embodiments, the nasal cavity is a preferred production site for therapeutic proteins using the retroviral / lentiviral (e.g., SIV) vectors and / or nucleic acids (e.g., plasmids) of the present invention for at least one of the following reasons: (i) extracellular barriers such as inflammatory cells and sputum are less prominent in the nasal cavity; (ii) ease of vector administration; (iii) small amounts of vectors / nucleic acids required; and (iv) ethical considerations. Therefore, transduction of nasal epithelial cells with retroviral / lentiviral (e.g., SIV) vectors or transfection with nucleic acids (e.g., plasmids) of the present invention can result in efficient (high-level) and long-lasting expression of the therapeutic transgene of interest. Therefore, nasal administration of the retroviral / lentiviral (e.g., SIV) vectors or nucleic acids (e.g., plasmids) of the present invention may be preferred.

[0284] Formulations for intranasal administration may be in the form of nasal drops or nasal sprays. Nasal formulations may include drops having an approximate diameter of 100 to 5,000 μm, e.g., 500 to 4,000 μm, 1,000 to 3,000 μm, or 100 to 1,000 μm. Alternatively, the volume of the drops may be in the range of about 0.001 to 100 μl, e.g., 0.1 to 50 μl or 1.0 to 25 μl, or e.g., 0.001 to 1 μl.

[0285] Aerosol formulations can take the form of a powder, suspension, or solution. The size of the aerosol particles is related to the delivery capacity of the aerosol. Smaller particles may travel further down the respiratory tract toward the alveoli than larger particles. In one embodiment, the aerosol particles have a diameter distribution that facilitates delivery along the entire length of the bronchi, bronchioles, and alveoli. Alternatively, the particle size distribution can be selected to target specific compartments of the respiratory tract, such as the alveoli. For aerosol delivery of pharmaceuticals, the particles may have diameters ranging from approximately 0.1 to 50 μm, preferably 1 to 25 μm, and more preferably 1 to 5 μm.

[0286] The aerosol particles may be for delivery using a nebulizer (e.g., via the mouth) or a nasal spray. The aerosol formulation may optionally contain a propellant and / or a surfactant.

[0287] The formulation of pharmaceutical aerosols is routine for those skilled in the art; see, for example, Sciarra, J., Remington's Pharmaceutical Sciences, supra. The drug is formulated as a dry powder aerosol solution, aerosol dispersion or suspension, emulsion, or semisolid product. The aerosol is delivered using any propellant system known to those skilled in the art. The aerosol may be applied to the upper or lower respiratory tract, for example, via nasal inhalation, or both. The portion of the lung to which the drug is delivered is determined by the disorder. Particularly when intranasal delivery is used, compositions containing the vectors of the present invention may contain a humectant. This may help reduce or prevent mucosal drying and prevent membrane irritation. Suitable humectants include, for example, sorbitol, mineral oil, vegetable oil, and glycerol; soothing agents; membrane conditioners; sweeteners; and combinations thereof. The composition may contain a surfactant. Suitable surfactants include nonionic surfactants, anionic surfactants, and cationic surfactants. Examples of surfactants that may be used include polyoxyethylene derivatives of partial fatty acid esters of sorbitol anhydride, such as, for example, Tween 80, polyoxyl 40 stearate, polyoxyethylene 50 stearate, fusidates, bile salts, and octoxynol.

[0288] In some cases, subsequent administration of retroviral / lentiviral (e.g., SIV) vectors and / or nucleic acids (e.g., plasmids) is carried out after the initial administration. Administration can be, for example, at least one week, two weeks, one month, two months, six months, one year, or more after the initial administration. In some cases, the retroviral / lentiviral (e.g., SIV) vectors and / or nucleic acids (e.g., plasmids) of the present invention are administered at intervals of at least once weekly, once every two weeks, once monthly, every two months, every six months, every year, or more. Preferably, administration is every six months, more preferably every year. The retroviral / lentiviral (e.g., SIV) vectors and / or nucleic acids (e.g., plasmids) are administered at intervals determined, for example, by when the effectiveness of previous administrations is declining.

[0289] Any two or more retroviral / lentiviral (e.g., SIV) vectors and / or nucleic acids (e.g., plasmids) of the present invention may be administered separately, sequentially, or simultaneously. Thus, two or more retroviral / lentiviral (e.g., SIV) vectors and / or nucleic acids (e.g., plasmids) may be administered separately, simultaneously, or sequentially, where at least one retroviral / lentiviral (e.g., SIV) vector and / or nucleic acid (e.g., plasmid) is a retroviral / lentiviral (e.g., SIV) vector and / or nucleic acid (e.g., plasmid) of the present invention. In particular, two or more retroviral / lentiviral (e.g., SIV) vectors and / or nucleic acids (e.g., plasmids) of the present invention may be administered in this manner. The two may be administered in the same or different compositions. In preferred cases, the two retroviral / lentiviral (e.g., SIV) vectors and / or nucleic acids (e.g., plasmids) are delivered in the same composition.

[0290] How to distinguish between retroviral / lentiviral vectors and transgene mRNA Conventional methods for producing retroviral / lentiviral (e.g., SIV) vectors produce retroviral / lentiviral (e.g., SIV) vectors that, when transcribed, produce mRNA that is identical in sequence to the retroviral / lentiviral (e.g., SIV) genome. Therefore, it is not possible to distinguish between the retroviral / lentiviral (e.g., SIV) vector genome (the prodrug) and the transcribed mRNA (which is then translated to produce the therapeutic protein).

[0291] In contrast, the present invention relates to retroviral / lentiviral (e.g., SIV) vectors containing an intron containing an RRE, as described herein. Thus, the retroviral / lentiviral (e.g., SIV) genome (prodrug) contains the sequence of the RRE-containing intron. This RRE-containing intron is spliced ​​out and removed during transcription of the retroviral / lentiviral (e.g., SIV) genome, resulting in an mRNA lacking the RRE-containing intron and therefore having a different nucleic acid sequence compared to the retroviral / lentiviral (e.g., SIV) genome from which it was derived. The different sequences of the retroviral / lentiviral (e.g., SIV) genome and the transcribed mRNA allow for the development of specific PCR- and in situ hybridization-based assays to detect and quantify the different nucleic acid sequences, allowing for the quantification of the retroviral / lentiviral (e.g., SIV) genome and the transcribed mRNA, respectively. Thus, the present invention provides a means to distinguish between the retroviral / lentiviral (e.g., SIV) vector genome (prodrug) and the mRNA (the first step toward an active therapeutic). This is useful, for example, during the production of retroviral / lentiviral (e.g., SIV) vectors, during in vitro use, and / or for evaluating the clinical efficacy of retroviral / lentiviral (e.g., SIV) vectors.

[0292] Thus, the present invention provides a method for distinguishing between a retroviral / lentiviral (e.g., SIV) vector and a transgene expressed by said retroviral vector, the method comprising or consisting of: (a) transfecting cells with the retroviral / lentiviral (e.g., SIV) vector of the present invention; (b) culturing the cells to allow expression of the transgene by the retroviral / lentiviral (e.g., SIV) vector; and (c) quantifying RNA in the cells; (i) the amount of RNA containing an intron with a retroviral / lentiviral (e.g., SIV) RRE corresponds to the copy number of the retroviral / lentiviral (e.g., SIV) vector; and (ii) the amount of RNA lacking an intron with a retroviral / lentiviral (e.g., SIV) RRE corresponds to the amount of transgene mRNA.

[0293] The present invention provides a method for distinguishing between a retroviral / lentiviral (e.g., SIV) vector and mRNA transcribed from the retroviral / lentiviral (e.g., SIV) vector, the method comprising or consisting of: (a) transfecting a cell with the retroviral / lentiviral (e.g., SIV) vector of the present invention; (b) culturing the cell to allow transcription of the genome of the retroviral / lentiviral (e.g., SIV) vector; and (c) quantifying RNA in the cells; (i) the amount of RNA containing an intron into which a retroviral / lentiviral (e.g., SIV) RRE has been inserted corresponds to the copy number of the retroviral / lentiviral (e.g., SIV) vector; and (ii) the amount of RNA lacking an intron into which a retroviral / lentiviral (e.g., SIV) RRE has been inserted corresponds to the amount of mRNA transcribed from the retroviral / lentiviral (e.g., SIV) vector genome, and optionally the amount of transcribed mRNA corresponds to the expression level of the transgene.

[0294] The method can involve quantifying RNA by any suitable technique, examples of which are known in the art and can be selected by those skilled in the art without undue burden.Preferably, the method can involve quantifying RNA by PCR-based assay and / or in situ hybridization-based assay.Such a PCR-based method can involve the use of two pairs of primers.The first primer pair comprises one primer that binds to a sequence outside the intron and another primer that binds to a sequence inside the intron.This first primer pair can detect and quantify unspliced ​​retroviral / lentiviral (e.g., SIV) vectors.The second primer pair comprises two primers that bind to the outside of the intron on both sides, and therefore only quantifies spliced ​​retroviral / lentiviral (e.g., SIV) vectors. Since retroviral / lentiviral (e.g., SIV) mRNA is spliced, the second primer pair will be specific for mRNA transcribed from the retroviral / lentiviral (e.g., SIV) vector, while the first primer pair will detect the retroviral / lentiviral (e.g., SIV) vector genome and integrated retroviral / lentiviral (e.g., SIV) DNA.

[0295] sequence homology Any of a variety of sequence alignment methods can be used to determine percent identity, including, but not limited to, global methods, local methods, and hybrid methods such as segmental approaches. Protocols for determining percent identity are routine procedures within the skill of those in the art. Global methods align sequences from the beginning to the end of the molecule, summing the scores of each residue pair and applying gap penalties to determine the best alignment. Non-limiting methods include, for example, CLUSTAL W (see, e.g., Julie D. Thompson et al., CLUSTAL W: Improving the Sensitivity of Progressive Multiple Sequence Alignment Through Sequence Weighting, Position-Specific Gap Penalties and Weight Matrix Choice, 22(22) Nucleic Acids Research, 4673-4680 (1994)); and iterative refinement (see, e.g., Osamu Gotoh, Significant Improvement in Accuracy of Multiple Proteins. Sequence Alignments by Iterative Refinement as Assessed by Reference to Structural Alignments, 264(4) J. Mol. Biol. 823-838 (1996)). Local methods align sequences by identifying one or more conserved motifs shared by all of the input sequences.Non-limiting methods include, for example, Match-box (see, e.g., Eric Depiereux and Ernest Feytmans, Match-Box: A Fundamentally New Algorithm for the Simultaneous Alignment of Several Protein Sequences, 8(5) CABIOS 501-509 (1992)); Gibbs sampling (see, e.g., CELawrence et al., Detecting Subtle Sequence Signals: A Gibbs Sampling Strategy for Multiple Alignment, 262(5131) Science, 208-214 (1993)); Align-M (see, e.g., Ivo Van Walle et al., Align-M - A New Algorithm for Multiple Alignment of Highly Divergent Sequences, 20(9) Bioinformatics:1428-1435 (2004)).

[0296] Thus, percent sequence identity is determined by conventional methods. See, e.g., Altschul et al., Bull. Math. Bio. 48:603-16, 1986, and Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89:10915-19, 1992. Briefly, two amino acid sequences are aligned to optimize the alignment score, as shown below (amino acids are indicated by standard single-letter code), using a gap opening penalty of 10, a gap extension penalty of 1, and the "blosum 62" scoring matrix of Henikoff and Henikoff (ibid.).

[0297] The "percent sequence identity" between two or more nucleic acid or amino acid sequences is a function of the number of identical positions shared by the sequences. Thus, the percent identity is calculated as the number of identical nucleotides / amino acids divided by the total number of nucleotides / amino acids and multiplied by 100. The calculation of percent sequence identity can also take into account the number of gaps and the length of each gap that need to be introduced to optimize the alignment of two or more sequences. The comparison of sequences and the determination of percent identity between two or more sequences are performed using specific mathematical algorithms, such as BLAST, which are familiar to those skilled in the art.

[0298] Alignment scores for determining sequence identity ARNDCQEGHILKMFPSTWYV A4 R -1 5 N -2 0 6 D -2 -2 1 6 C 0 -3 -3 -3 9 Q -1 1 0 0 -3 5 E -1 0 0 2 -4 2 5 G 0 -2 0 -1 -3 -2 -2 6 H -2 0 1 -1 -3 0 0 -2 8 I -1 -3 -3 -3 -1 -3 -3 -4 -3 4 L -1 -2 -3 -4 -1 -2 -3 -4 -3 2 4 K -1 2 0 -1 -3 1 1 -2 -1 -3 -2 5 M -1 -1 -2 -3 -1 0 -2 -3 -2 1 2 -1 5 F -2 -3 -3 -3 -2 -3 -3 -3 -1 0 0 -3 0 6 P -1 -2 -2 -1 -3 -1 -1 -2 -2 -3 -3 -1 -2 -4 7 S 1 -1 1 0 -1 0 0 0 -1 -2 -2 0 -1 -2 -1 4 T 0 -1 0 -1 -1 -1 -1 -2 -2 -1 -1 -1 -1 -2 -1 1 5 W -3 -3 -4 -4 -2 -2 -3 -2 -2 -3 -2 -3 -1 1 -4 -3 -2 11 Y -2 -2 -2 -3 -2 -1 -2 -3 2 -1 -1 -2 -1 3 -3 -2 -2 2 7 V 0 -3 -3 -3 -1 -2 -2 -3 -3 3 1 -2 1 -1 -2 -2 0 -3 -1 4

[0299] Therefore, the percent identity is The total number of identical matches ________________________________________________________×100 [Length of the longer sequence + Alignment of the two sequences introduced into the longer sequence to Number of Gaps] It is calculated as:

[0300] Substantially homologous polypeptides are characterized as having one or more amino acid substitutions, deletions, or additions. These changes are preferably of a minor nature, being conservative amino acid substitutions (described herein) and other substitutions that do not significantly affect the folding or activity of the polypeptide; small deletions, typically from 1 to about 30 amino acids; and small amino- or carboxyl-terminal extensions such as an amino-terminal methionine residue, small linker peptides of up to about 20-25 residues, or affinity tags.

[0301] In addition to the 20 standard amino acids, non-standard amino acids (such as 4-hydroxyproline, 6-N-methyllysine, 2-aminoisobutyric acid, isovaline, and α-methylserine) are substituted for amino acid residues in the polypeptides of the invention. A limited number of non-conservative amino acids, amino acids not encoded by the genetic code, and unnatural amino acids are substituted for amino acid residues in the polypeptides. The polypeptides of the invention may also include non-naturally occurring amino acid residues.

[0302] Non-naturally occurring amino acids include, but are not limited to, trans-3-methylproline, 2,4-methano-proline, cis-4-hydroxyproline, trans-4-hydroxy-proline, N-methylglycine, allo-threonine, methyl-threonine, hydroxy-ethyl cysteine, hydroxyethyl homo-cysteine, nitro-glutamine, homoglutamine, pipecolic acid, tert-leucine, norvaline, 2-azaphenylalanine, 3-azaphenylalanine, 4-azaphenylalanine and 4-fluorophenylalanine.Several methods for incorporating non-naturally occurring amino acid residues into proteins are known in the art.For example, an in vitro system is utilized in which chemically aminoacylated suppressor tRNA is used to suppress nonsense mutations.Methods for synthesizing amino acids and aminoacylated tRNA are known in the art. Transcription and translation of plasmids containing nonsense mutations are carried out in a cell-free system containing Escherichia coli S30 extract and commercially available enzymes and other reagents. Proteins are purified by chromatography. See, for example, Robertson et al., J. Am. Chem. Soc. 113:2722, 1991; Ellman et al., Methods Enzymol. 202:301, 1991; Chung et al., Science 259:806-9, 1993; and Chung et al., Proc. Natl. Acad. Sci. USA 90:10145-9, 1993. In a second method, translation is carried out in Xenopus oocytes by microinjection of mutated mRNA and chemically aminoacylated suppressor tRNA (Turcatti et al., J. Biol. Chem. 271:19991-8, 1996). In a third method, E. coli cells are cultured in the absence of the natural amino acid to be replaced (e.g., phenylalanine) and in the presence of a desired unnatural amino acid (e.g., 2-azaphenylalanine, 3-azaphenylalanine, 4-azaphenylalanine, or 4-fluorophenylalanine), which is incorporated into the polypeptide in place of its natural counterpart.See Koide et al., Biochem. 33:7470-6, 1994. Naturally occurring amino acid residues are converted to non-naturally occurring species by in vitro chemical modification. Chemical modification is combined with site-directed mutagenesis to further expand the range of substitutions (Wynn and Richards, Protein Sci. 2:395-403, 1993).

[0303] A limited number of non-conservative amino acids, amino acids that are not encoded by the genetic code, non-naturally occurring amino acids, and unnatural amino acids may be substituted for amino acid residues in the polypeptides of the invention.

[0304] Essential amino acids within the polypeptides of the invention are identified according to procedures known in the art, such as site-directed mutagenesis or alanine scanning mutagenesis (Cunningham and Wells, Science 244:1081-5, 1989). Sites of biological interaction are also determined by physical analysis of structures determined by techniques such as nuclear magnetic resonance, crystallography, electron diffraction, or photoaffinity labeling, followed by mutation of amino acids at putative contact sites. See, e.g., de Vos et al., Science 255:306-12, 1992; Smith et al., J. Mol. Biol. 224:899-904, 1992; Wlodaver et al., FEBS Lett. 309:59-64, 1992. The identity of essential amino acids is also inferred from analysis of homology with related components of the polypeptides of the invention (e.g., translocation or protease components).

[0305] Multiple amino acid substitutions are made and tested using known mutagenesis and screening methods, such as those disclosed by Reidhaar-Olson and Sauer (Science 241:53-7, 1988) or Bowie and Sauer (Proc. Natl. Acad. Sci. USA, 86:2152-6, 1989). Briefly, these authors disclose methods for simultaneously randomizing two or more positions within a polypeptide, selecting for functional polypeptides, and then sequencing the mutated polypeptides to determine the spectrum of permissible substitutions at each position. Other methods used include phage display (e.g., Lowman et al., Biochem. 30:10832-7, 1991; Ladner et al., U.S. Pat. No. 5,223,409; Huse, WIPO Publication No. WO 92 / 06204), and region-specific mutagenesis (Derbyshire et al., Gene 46:145, 1986; Ner et al., DNA 7:127, 1988).

[0306] Multiple amino acid substitutions are made and tested using known mutagenesis and screening methods, such as those disclosed by Reidhaar-Olson and Sauer (Science 241:53-7, 1988) or Bowie and Sauer (Proc. Natl. Acad. Sci. USA, 86:2152-6, 1989). Briefly, these authors disclose methods for simultaneously randomizing two or more positions within a polypeptide, selecting for functional polypeptides, and then sequencing the mutated polypeptides to determine the spectrum of permissible substitutions at each position. Other methods used include phage display (e.g., Lowman et al., Biochem. 30:10832-7, 1991; Ladner et al., U.S. Pat. No. 5,223,409; Huse, WIPO Publication No. WO 92 / 06204), and region-specific mutagenesis (Derbyshire et al., Gene 46:145, 1986; Ner et al., DNA 7:127, 1988).

[0307] Sequence information SEQ ID NO: 1 SIV RRE sequence SEQ ID NO: 2 β-globulin / IgG chimeric intron splice donor site SEQ ID NO: 3 β-globulin / IgG chimeric intron splice acceptor site SEQ ID NO: 4 β-globulin / IgG chimeric intron SEQ ID NO: 5 β-globin / IgG chimeric intron containing SIV RRE SEQ ID NO: 6 Exemplary AAT transgene (SERPINA1) SEQ ID NO: 7 Complement to the exemplified AAT transgene SEQ ID NO: 8 Exemplary AAT Polypeptide SEQ ID NO: 9 Exemplary FVIII transgene (N6) SEQ ID NO: 10 Exemplary FVIII transgene (V3) SEQ ID NO: 11 Complementary to the exemplified FVIII transgene (N6) SEQ ID NO: 12 Complementary strand to the exemplified FVIII transgene (V3) SEQ ID NO: 13 Exemplary FVIII Polypeptide (N6) SEQ ID NO: 14 Exemplary FVIII Polypeptide (V3) SEQ ID NO: 15 Exemplary CFTR transgene (soCFTR2) SEQ ID NO: 16 Exemplary CFTR Polypeptide SEQ ID NO: 17 Exemplary hGM-CSF transgene SEQ ID NO: 18 Exemplary hGM-CSF Polypeptide SEQ ID NO: 19 Exemplary mGM-CSF transgene SEQ ID NO: 20 Exemplary mGM-CSF Polypeptide SEQ ID NO: 21 Exemplary human DCN (decorin) transgene SEQ ID NO: 22 Exemplary human decorin polypeptide SEQ ID NO: 23 Exemplary human TRIM72 transgene SEQ ID NO: 24 Exemplary human TRIM72 polypeptide SEQ ID NO: 25 Exemplary human ABCA3 (ABCA3) transgene SEQ ID NO: 26 Exemplary human ABCA3 polypeptide SEQ ID NO: 27 Exemplary hCEF promoter SEQ ID NO: 28 Exemplary CMV promoter SEQ ID NO: 29 Exemplary EF1a promoter SEQ ID NO: 30 Plasmid defined in Figure 2A (pDNA1 pGM991) SEQ ID NO: 31 Plasmid defined in Figure 2B (pDNA1 pGM691) SEQ ID NO: 32 Plasmid defined in Figure 2C (pDNA2a pGM297) SEQ ID NO: 33 Plasmid defined in Figure 2D (pDNA2b pGM299) SEQ ID NO: 34 Plasmid defined in Figure 2E (pDNA3a pGM301) SEQ ID NO: 35 Plasmid defined in Figure 2F (pDNA3b pGM303) SEQ ID NO: 36 Plasmid defined in Figure 2G (pDNA2a pGM407) SEQ ID NO: 37 Codon-optimized SIV gag-pol nucleic acid sequence (from pGM691) SEQ ID NO: 38 Exemplary CAG promoter SEQ ID NO: 39 Exemplary WPRE component (mWPRE) SEQ ID NO: 40 Exemplary human SFTPB transgene SEQ ID NO: 41 Exemplary human SFTPB polypeptide SEQ ID NO: 42 Exemplary human ADAMTS13 transgene SEQ ID NO: 43 Exemplary human ADAMTS13 polypeptide SEQ ID NO: 44 Exemplary rSIV Rev protein SEQ ID NO: 45 oIC001 DNA forward primer SEQ ID NO: 46 oIC002 DNA and RNA reverse primer SEQ ID NO: 47 oIC105 RNA forward primer SEQ ID NO: 48 CAGGS promoter containing chicken β-actin splice donor and rabbit β-globulin splice acceptor SEQ ID NO: 49 Plasmid defined in Figure 2H (pDNA3 pMD2.G) SEQ ID NO: 50 HIV RRE sequence

[0308] SEQ ID NO: 1 SIV RRE sequence CCGTTTGTGCTAGGGTTCTTAGGCTTCTTGGGGGCTGCTGGAACTGCAATGGGAGCAGCGGCGACAGCCCTGACGGTCCAGTCTCAGCATTTGCTTGCTGGGATACTGCAGCAGCAGAAGAATCTGCTGGCGGCTGTGGAGGCTCAACAGCA GATGTTGAAGCTGACCATTTGGGGTGTTAAAAACCTCAATGCCCGCGTCACAGCCCTTGAGAAGTACCTAGAGGATCAGGCACGACTAAAACTCCTGGGGGTGCGCATGGAAACAAGTATGTCATACCACAGTGGAGTGGCCCTGGACAAATCG GACTCCGGATTGGCAAAATATGACTTGGTTGGAGTGGGAAAGACAAATAGCTGATTTGGAAAGCAACATTACGAGACAATTAGTGAAGGCTAGAGAACAAGAGGAAAAGAATCTAGATGCCTATCAGAAGTTAACTAGTTGGTCAGATTTCTG GTCTTGGTTCGATTTCTCAAAATGGCTTAACATTTTAAAAAATGGGATTTTTAGTAATAGTAGGAATAATAGGGTTAAGATTACTTTACACAGTATATGGATGTATAGTGAGGGTTAGGCAGGGATATGTTCCTCTATCTCCAGATCCATAT

[0309] SEQ ID NO: 2 β-globulin / IgG chimeric intron splice donor site TGAGTTTAAGGTAAGT

[0310] SEQ ID NO: 3 β-globulin / IgG chimeric intron splice acceptor site CTCTCCACAG

[0311] SEQ ID NO: 4 β-globulin / IgG chimeric intron GTAAGTATCAAGGTTACAAGACAGGTTTAAGGAGACCAATAGAAACTGGGCTTGTCGAGACAGAGAAGACTCTTGCGTTTCTGATAGGCACCTATTGGTCTTACTGACATCCACTTTGCCTTTCTCTCCACAG

[0312] SEQ ID NO: 5 β-globin / IgG chimeric intron containing SIV RRE [ka] Underlined = β-globin / IgG chimeric intron Double underline = NotI restriction site Italics = SIV RRE sequence

[0313] SEQ ID NO: 6 Exemplary AAT transgene (SERPINA1) atgcccagct ctgtgtcctg gggcattctg ctgctggctg gcctgtgctg tctggtgcct 60 gtgtccctgg ctgaggaccc tcagggggat gctgcccaga aaacagacac ctcccaccat 120 gaccaggacc accccacctt caacaagatc acccccaacc tggcagagtt tgccttcagc 180 ctgtacagac agctggccca ccagagcaac agcaccaaca tctttttcag ccctgtgtcc 240 attgccacag ccttgccat gctgagcctg ggcaccaagg ctgacaccca tgatgagatc 300 ctggaaggcc tgaacttcaa cctgacagag atccctgagg cccagatcca tgagggcttc 360 caggaactgc tgagaaccct gaaccagcca gacagccagc tgcagctgac aacaggcaat 420 gggctgttcc tgtctgaggg cctgaagctg gtggacaagt ttctggaaga tgtgaagaag 480 ctgtaccact ctgaggcctt cacagtgaac tttggggaca cagaagaggc caagaaacag 540 atcaatgact atgtggaaaa gggcacccag ggcaagattg tggaccttgt gaaagagctg 600 gacagggaca ctgtgtttgc ccttgtgaac tacatcttct tcaagggcaa gtgggagagg 660 ccctttgaag tgaaggacac tgaggaagag gacttccatg tggaccaagt gaccacagtg 720 aaggtgccaa tgatgaagag actggggatg ttcaatatcc agcactgcaa gaaactgagc 780 agctgggtgc tgctgatgaa gtacctgggc aatgctacag ccatattctt tctgcctgat 840 gagggcaagc tgcagcacct ggaaaatgag ctgacccatg acatcatcac caaatttctg 900 gaaaatgagg acagaagatc tgccagcctg catctgccca agctgagcat cacaggcaca 960 tatgacctga agtctgtgct gggacagctg ggaatcacca aggtgttcag caatggggca 1020 gacctgagtg gagtgacaga ggaagcccct ctgaagctgt ccaaggctgt gcacaaggca 1080 gtgctgacca ttgatgagaa gggcacagag gctgctgggg ccatgtttct ggaagccatc 1140 cccatgtcca tccccccaga agtgaagttc aacaagccct ttgtgttcct gatgattgag 1200 cagaacacca agagccccct gttcatgggc aaggttgtga accccaccca gaaatga 1257

[0314] Complementary strand to the AAT transgene exemplified by SEQ ID NO: 7 tacgggtcga gacacaggac cccgtaagac gacgaccgac cggacacgac agaccacgga 60 cacagggacc gactcctggg agtcccccta cgacgggtct tttgtctgtg gagggtggta 120 ctggtcctgg tggggtggaa gttgttctag tgggggttgg accgtctcaa acggaagtcg 180 gacatgtctg tcgaccgggt ggtctcgttg tcgtggttgt agaaaaagtc gggacacagg 240 taacggtgtc ggaaacggta cgactcggac ccgtggttcc gactgtgggt actactctag 300 gaccttccgg acttgaagtt ggactgtctc tagggactcc gggtctaggt actcccgaag 360 gtccttgacg actcttggga cttggtcggt ctgtcggtcg acgtcgactg ttgtccgtta 420 cccgacaagg acagactccc ggacttcgac cacctgttca aagaccttct acacttcttc 480 gacatggtga gactccggaa gtgtcacttg aaacccctgt gtcttctccg gttctttgtc 540 tagttactga tacacctttt cccgtgggtc ccgttctaac acctggaaca ctttctcgac 600 ctgtccctgt gacacaaacg ggaacacttg atgtagaaga agttcccgtt caccctctcc 660 gggaaacttc acttcctgtg actccttctc ctgaaggtac acctggttca ctggtgtcac 720 ttccacggtt actacttctc tgacccctac aagttatagg tcgtgacgtt ctttgactcg 780 tcgacccacg acgactactt catggacccg ttacgatgtc ggtataagaa agacggacta 840 ctcccgttcg acgtcgtgga ccttttactc gactgggtac tgtagtagtg gtttaaagac 900 cttttactcc tgtcttctag acggtcggac gtagacgggt tcgactcgta gtgtccgtgt 960 atactggact tcagacacga ccctgtcgac ccttagtggt tccacaagtc gttaccccgt 1020 ctggactcac ctcactgtct ccttcgggga gacttcgaca ggttccgaca cgtgttccgt 1080 cacgactggt aactactctt cccgtgtctc cgacgacccc ggtacaaaga ccttcggtag 1140 gggtacaggt aggggggtct tcacttcaag ttgttcggga aacacaagga ctactaactc 1200 gtcttgtggt tctcggggga caagtacccg ttccaacact tggggtgggt ctttact 1257

[0315] The AAT polypeptide exemplified by SEQ ID NO: 8 Ala Glu Asp Pro Gln Gly Asp Ala Ala Gln Lys Thr Asp Thr Ser His 1 5 10 15 His Asp Gln Asp His Pro Thr Phe Ala Glu Asp Pro Gln Gly Asp Ala 20 25 30 Ala Gln Lys Thr Asp Thr Ser His His Asp Gln Asp His Pro Thr Phe 35 40 45 Asn Lys Ile Thr Pro Asn Leu Ala Glu Phe Ala Phe Ser Leu Tyr Arg 50 55 60 Gln Leu Ala His Gln Ser Asn Ser Thr Asn Ile Phe Phe Ser Pro Val 65 70 75 80 Ser Ile Ala Thr Ala Phe Ala Met Leu Ser Leu Gly Thr Lys Ala Asp 85 90 95 Thr His Asp Glu Ile Leu Glu Gly Leu Asn Phe Asn Leu Thr Glu Ile 100 105 110 Pro Glu Ala Gln Ile His Glu Gly Phe Gln Glu Leu Leu Arg Thr Leu 115 120 125 Asn Gln Pro Asp Ser Gln Leu Gln Leu Thr Thr Gly Asn Gly Leu Phe 130 135 140 Leu Ser Glu Gly Leu Lys Leu Val Asp Lys Phe Leu Glu Asp Val Lys 145 150 155 160 Lys Leu Tyr His Ser Glu Ala Phe Thr Val Asn Phe Gly Asp Thr Glu 165 170 175 Glu Ala Lys Lys Gln Ile Asn Asp Tyr Val Glu Lys Gly Thr Gln Gly 180 185 190 Lys Ile Val Asp Leu Val Lys Glu Leu Asp Arg Asp Thr Val Phe Ala 195 200 205 Leu Val Asn Tyr Ile Phe Phe Lys Gly Lys Trp Glu Arg Pro Phe Glu 210 215 220 Val Lys Asp Thr Glu Glu Glu Asp Phe His Val Asp Gln Val Thr Thr 225 230 235 240 Val Lys Val Pro Met Met Lys Arg Leu Gly Met Phe Asn Ile Gln His 245 250 255 Cys Lys Lys Leu Ser Ser Trp Val Leu Leu Met Lys Tyr Leu Gly Asn 260 265 270 Ala Thr Ala Ile Phe Phe Leu Pro Asp Glu Gly Lys Leu Gln His Leu 275 280 285 Glu Asn Glu Leu Thr His Asp Ile Ile Thr Lys Phe Leu Glu Asn Glu 290 295 300 Asp Arg Arg Ser Ala Ser Leu His Leu Pro Lys Leu Ser Ile Thr Gly 305 310 315 320 Thr Tyr Asp Leu Lys Ser Val Leu Gly Gln Leu Gly Ile Thr Lys Val 325 330 335 Phe Ser Asn Gly Ala Asp Leu Ser Gly Val Thr Glu Glu Ala Pro Leu 340 345 350 Lys Leu Ser Lys Ala Val His Lys Ala Val Leu Thr Ile Asp Glu Lys 355 360 365 Gly Thr Glu Ala Ala Gly Ala Met Phe Leu Glu Ala Ile Pro Met Ser 370 375 380 Ile Pro Pro Glu Val Lys Phe Asn Lys Pro Phe Val Phe Leu Met Ile 385,390,395,400 Glu Gln Asn Thr Lys Ser Pro Leu Phe Met Gly Lys Val Val Asn Pro 405 410 415 Thr Gln Lys

[0316] FVIII in the 9th century (N6) atgcagattg agctgagcac ctgctttcttc ctgtgcctgc tgaggttctg cttctctgcc 60 accaggagat actacctggg ggctgtggag ctgagctggg actacatgca gtctgacctg 120 ggggagctgc ctgtggatgc caggttcccc cccagagtgc ccaagagctt ccccttcaac 180 acctctgtgg tgtacaagaa gaccctgttt gtggagttca ctgaccacct gttcaacatt 240 gccaagccca ggcccccctg gatgggcctg ctgggcccca ccatccaggc tgaggtgtat 300 gacactgtgg tgatcaccct gaagaacatg gccagccacc ctgtgagcct gcatgctgtg 360 ggggtgagct actggaaggc ctctgagggg gctggattg atgaccagac cagccagagg 420 gagaaggagg atgacaaggt gttccctggg ggcagccaca cctatgtgtg gcaggtgctg 480 aaggaatg gccccatggc ctctgacccc ctgtgcctga cctacagcta cctgagccat 540 gtggacctgg tgaaggacct gaactctggc ctgattgggg ccctgctggt gtgcagggag 600 ggcagcctgg ccaaggagaa gacccagacc ctgcacaagt tcatcctgct gtttgctgtg 660 tttgatgagg gcaagagctg gcactctgaa accaagaaca gcctgatgca ggacagggat 720 gctgcccttg ccagggcctg gcccaagatg caacactgtga atggctatgt gaacaggagc 780 ctgcctggcc tgattggctg ccacaggaag tctgtgtact ggcatgtgat tggcatgggc 840 accacccctg aggtgcacag catcttcctg gagggccaca ccttcctggt caggaaccac 900 aggcaggcca gcctggagat cagccccatc accttcctga ctgcccagac cctgctgatg 960 gacctgggcc agttcctgct gttctgccac atcagcagcc accagcatga tggcatggag 1020 gcctatgtga aggtggacag ctgccctgag gagccccagc tgaggatgaa gaacaatgag 1080 gaggctgagg actatgatga tgacctgact gactctgaga tggatgtggt gaggtttgat 1140 gatgacaaca gccccagctt catccagatc aggtctgtgg ccaagaagca ccccaagacc 1200 tgggtgcact acattgctgc tgaggaggag gactgggact atgcccccct ggtgctggcc 1260 cctgatgaca ggagctacaa gagccagtac ctgaacaatg gcccccagag gattggcagg 1320 aagtacaaga aggtcaggtt catggcctac actgatgaaa ccttcaagac cagggaggcc 1380 atccagcatg agtctggcat cctgggcccc ctgctgtatg gggaggtggg ggacaccctg 1440 ctgatcatct tcaagaacca ggccagcagg ccctacaaca tctaccccca tggcatcact 1500 gatgtgaggc ccctgtacag caggaggctg cccaaggggg tgaagcacct gaaggacttc 1560 cccatcctgc ctggggagat cttcaagtac aagtggactg tgactgtgga ggatggcccc 1620 accaagtctg accccaggtg cctgaccaga tactacagca gctttgtgaa catggagagg 1680 gacctggcct ctggcctgat tggccccctg ctgatctgct acaaggagtc tgtggaccag 1740 aggggcaacc agatcatgtc tgacaagagg aatgtgatcc tgttctctgt gtttgatgag 1800 aacaggagct ggtacctgac tgagaacatc cagaggttcc tgcccaaccc tgctggggtg 1860 cagctggagg accctgagtt ccaggccagc aacatcatgc acagcatcaa tggctatgtg 1920 tttgacagcc tgcagctgtc tgtgtgcctg catgaggtgg cctactggta catcctgagc 1980 attggggccc agactgactt cctgtctgtg ttcttctctg gctacacctt caagcacaag 2040 atggtgtatg aggacaccct gaccctgttc cccttctctg gggagactgt gttcatgagc 2100 atggagaacc ctggcctgtg gattctgggc tgccacaact ctgacttcag gaacaggggc 2160 atgactgccc tgctgaaagt ctccagctgt gacaagaaca ctggggacta ctatgaggac 2220 agctatgagg acatctctgc ctacctgctg agcaagaaca atgccattga gcccaggagc 2280 ttcagccaga acagcaggca ccccagcacc aggcagaagc agttcaatgc caccaccatc 2340 cctgagaatg acatagagaa gacagaccca tggtttgccc accggacccc catgcccaag 2400 atccagaatg tgagcagctc tgacctgctg atgctgctga ggcagagccc caccccccat 2460 ggcctgagcc tgtctgacct gcaggaggcc aagtatgaaa ccttctctga tgaccccagc 2520 cctggggcca ttgacagcaa caacagcctg tctgagatga cccacttcag gccccagctg 2580 caccactctg gggacatggt gttcacccct gagtctggcc tgcagctgag gctgaatgag 2640 aagctgggca ccactgctgc cactgagctg aagaagctgg acttcaaagt ctccagcacc 2700 agcaacaacc tgatcagcac catcccctct gacaacctgg ctgctggcac tgacaacacc 2760 agcagcctgg gcccccccag catgcctgtg cactatgaca gccagctgga caccaccctg 2820 tttggcaaga agagcagccc cctgactgag tctgggggcc ccctgagcct gtctgaggag 2880 aacaatgaca gcaagctgct ggagtctggc ctgatgaaca gccaggagag cagctggggc 2940 aagaatgtga gcagcaggga gatcaccagg accaccctgc agtctgacca ggaggagatt 3000 gactatgatg acaccatctc tgtggagatg aagaaggagg actttgacat ctacgacgag 3060 gacgagaacc agagccccag gagcttccag aagaagacca ggcactactt cattgctgct 3120 gtggagaggc tgtgggacta tggcatgagc agcagccccc atgtgctgag gaacagggcc 3180 cagtctggct ctgtgcccca gttcaagaag gtggtgttcc aggagttcac tgatggcagc 3240 ttcacccagc ccctgtacag aggggagctg aatgagcacc tgggcctgct gggcccctac 3300 atcagggctg aggtggagga caacatcatg gtgaccttca ggaaccaggc cagcaggccc 3360 tacagcttt acagcagcct gatcagctat gaggaggacc agaggcaggg ggctgagccc 3420 aggaagaact ttgtgaagcc caatgaaacc aagacctact tctggaaggt gcagcaccac 3480 atggccccca ccaaggatga gtttgactgc aaggcctggg cctacttctc tgatgtggac 3540 ctggagaagg atgtgcactc tggcctgatt ggccccctgc tggtgtgcca caccaacacc 3600 ctgaaccctg cccatggcag gcaggtgact gtgcaggagt ttgccctgtt cttcaccatc 3660 tttgatgaaa ccagagctg gtactcact gagacatgg agaggactg caggggccccc 3720 tgcacatcc agatggagga ccccaccttc aaggagaact acaggttcca tgccatcaat 3780 ggctacatca tggaccct gcctggctg gtgatggccc aggaccagag gatcaggtgg 3840 tacctgctga gcatgggcag caatgagaac atccacagca tccactctc tggccatgtg 3900 ttcactgtga ggaagaagga ggagtacaag atggccctgt acaacctgta ccctggggtg 3960 tttgagactg tggagatgct gcccagcaag gctggcatct ggaggtgga gtgcctgatt 4020 ggggagcacc tgcatgctgg catgagcacc ctgttcctgg tgtacagcaa caagtgccag 4080 accccctgg gcatggcctc tggccacatc agggacttcc agatcactgc ctctggccag 4140 tatggccagt gggcccca gctggccagg ctgcactact ctggcagcat caatgcctgg 4200 agcaccaagg agccttcag ctggatcaag gtggacctgc tggccccat gatcatccat 4260 ggcatcaaga cccaggggggc caggcagaag ttcagcagcc tgtacatcag ccagttcatc 4320 atcatgtaca gcctggatgg cagaagtgg cagacctaca ggggcacag cactggcacc 4380 ctgatggtgt tctttggcaa tgtggacagc tctggcatca agcacaacat cttcaacccc 4440 cccatcattg ccagatacat caggctgcac cccacccact acagcatcag gagcaccctg 4500 aggatggagc tgatgggctg tgacctgaac agctgcagca tgcccctggg catggagagc 4560 aaggccatct ctgatgccca gatcactgcc agcagctact tcaccaacat gtttgccacc 4620 tggagcccca gcaaggccag gctgcacctg cagggcagga gcaatgcctg gaggccccag 4680 gtcaacaacc ccaaggagtg gctgcaggtg gacttccaga agaccatgaa ggtgactggg 4740 gtgaccaccc agggggtgaa gagcctgctg accagcatgt atgtgaagga gttcctgatc 4800 agcagcagcc aggatggcca ccagtggacc ctgttcttcc agaatggcaa ggtgaaggtg 4860 ttccagggca accaggacag cttcacccct gtggtgaaca gcctggaccc ccccctgctg 4920 accagatacc tgaggattca cccccagagc tgggtgcacc agattgccct gaggatggag 4980 gtgctgggct gtgaggccca ggacctgtac tga 5013

[0317] FVIII transgene (V3) exemplified by SEQ ID NO: 10 atgcagattg agctgagcac ctgcttcttc ctgtgcctgc tgaggttctg cttctctgcc 60 accaggagat actacctggg ggctgtggag ctgagctggg actacatgca gtctgacctg 120 ggggagctgc ctgtggatgc caggttcccc cccagagtgc ccaagagctt ccccttcaac 180 acctctgtgg tgtacaagaa gaccctgttt gtggagttca ctgaccacct gttcaacatt 240 gccaagccca ggcccccctg gatgggcctg ctgggcccca ccatccaggc tgaggtgtat 300 gacactgtgg tgatcaccct gaagaacatg gccagccacc ctgtgagcct gcatgctgtg 360 ggggtgagct actggaaggc ctctgagggg gctggattg atgaccagac cagccagagg 420 gagaaggagg atgacaaggt gttccctggg ggcagccaca cctatgtgtg gcaggtgctg 480 aaggaatg gccccatggc ctctgacccc ctgtgcctga cctacagcta cctgagccat 540 gtggacctgg tgaaggacct gaactctggc ctgattgggg ccctgctggt gtgcagggag 600 ggcagcctgg ccaaggagaa gacccagacc ctgcacaagt tcatcctgct gtttgctgtg 660 tttgatgagg gcaagagctg gcactctgaa accaagaaca gcctgatgca ggacagggat 720 gctgcctctg ccagggcctg gcccaagatg cacactgtga atggctatgt gaacaggagc 780 ctgcctggcc tgattggctg ccacaggaag tctgtgtact ggcatgtgat tggcatgggc 840 accacccctg aggtgcacag catcttcctg gagggccaca ccttcctggt caggaaccac 900 aggcaggcca gcctggagat cagccccatc accttcctga ctgcccagac cctgctgatg 960 gacctgggcc agttcctgct gttctgccac atcagcagcc accagcatga tggcatggag 1020 gcctatgtga aggtggacag ctgccctgag gagccccagc tgaggatgaa gaacaatgag 1080 gaggctgagg actatgatga tgacctgact gactctgaga tggatgtggt gaggtttgat 1140 gatgacaaca gccccagctt catccagatc aggtctgtgg ccaagaagca ccccaagacc 1200 tgggtgcact acattgctgc tgaggaggag gactgggact atgcccccct ggtgctggcc 1260 cctgatgaca ggagctacaa gagccagtac ctgaacaatg gcccccagag gattggcagg 1320 aagtacaaga aggtcaggtt catggcctac actgatgaaa ccttcaagac cagggaggcc 1380 atccagcatg agtctggcat cctgggcccc ctgctgtatg gggaggtggg ggacaccctg 1440 ctgatcatct tcaagaacca ggccagcagg ccctacaaca tctaccccca tggcatcact 1500 gatgtgaggc ccctgtacag caggaggctg cccaaggggg tgaagcacct gaaggacttc 1560 cccatcctgc ctggggagat cttcaagtac aagtggactg tgactgtgga ggatggcccc 1620 accaagtctg accccaggtg cctgaccaga tactacagca gctttgtgaa catggagagg 1680 gacctggcct ctggcctgat tggccccctg ctgatctgct acaaggagtc tgtggaccag 1740 aggggcaacc agatcatgtc tgacaagagg aatgtgatcc tgttctctgt gtttgatgag 1800 aacaggagct ggtacctgac tgagaacatc cagaggttcc tgcccaaccc tgctggggtg 1860 cagctggagg accctgagtt ccaggccagc aacatcatgc acagcatcaa tggctatgtg 1920 tttgacagcc tgcagctgtc tgtgtgcctg catgaggtgg cctactggta catcctgagc 1980 attggggccc agactgactt cctgtctgtg ttcttctctg gctacacctt caagcacaag 2040 atggtgtatg aggacaccct gaccctgttc cccttctctg gggagactgt gttcatgagc 2100 atggagaacc ctggcctgtg gattctgggc tgccacaact ctgacttcag gaacaggggc 2160 atgactgccc tgctgaaagt ctccagctgt gacaagaaca ctggggacta ctatgaggac 2220 agctatgagg acatctctgc ctacctgctg agcaagaaca atgccattga gcccaggagc 2280 ttcagccaga atgccactaa tgtgtctaac aacagcaaca ccagcaatga cagcaatgtg 2340 tctcccccag tgctgaagag gcaccagagg gagatcacca ggaccaccct gcagtctgac 2400 caggaggaga ttgactatga tgacaccatc tctgtggaga tgaagaagga ggactttgac 2460 atctacgacg aggacgagaa ccagagcccc aggagcttcc agaagaagac caggcactac 2520 ttcattgctg ctgtggagag gctgtgggac tatggcatga gcagcagccc ccatgtgctg 2580 aggaacaggg cccagtctgg ctctgtgccc cagttcaaga aggtggtgtt ccaggagttc 2640 actgatggca gcttcaccca gcccctgtac agaggggagc tgaatgagca cctgggcctg 2700 ctgggcccct acatcagggc tgaggtggag gacaacatca tggtgacctt caggaaccag 2760 gccagcaggc cctacagctt ctacagcagc ctgatcagct atgaggagga ccagaggcag 2820 ggggctgagc ccaggaagaa ctttgtgaag cccaatgaaa ccaagaccta cttctggaag 2880 gtgcagcacc acatggcccc caccaaggat gagtttgact gcaaggcctg ggcctacttc 2940 tctgatgtgg acctggagaa ggatgtgcac tctggcctga ttggccccct gctggtgtgc 3000 cacaccaaca ccctgaaccc tgcccatggc aggcaggtga ctgtgcagga gtttgccctg 3060 ttcttcacca tctttgatga aaccaagagc tggtacttca ctgagaacat ggagaggaac 3120 tgcagggccc cctgcaacat ccagatggag gaccccacct tcaaggagaa ctacaggttc 3180 catgccatca atggctacat catggacacc ctgcctggcc tggtgatggc ccaggaccag 3240 aggatcaggt ggtacctgct gagcatgggc agcaatgaga acatccacag catccacttc 3300 tctggccatg tgttcactgt gaggaagaag gaggagtaca agatggccct gtacaacctg 3360 taccctgggg tgtttgagac tgtggagatg ctgcccagca aggctggcat ctggagggtg 3420 gagtgcctga ttggggagca cctgcatgct ggcatgagca ccctgttcct ggtgtacagc 3480 aacaagtgcc agacccccct gggcatggcc tctggccaca tcagggactt ccagatcact 3540 gcctctggcc agtatggcca gtgggccccc aagctggcca ggctgcacta ctctggcagc 3600 atcaatgcct ggagcaccaa ggagcccttc agctggatca aggtggacct gctggccccc 3660 atgatcatcc atggcatcaa gacccagggg gccaggcaga agttcagcag cctgtacatc 3720 agccagttca tcatcatgta cagcctggat ggcaagaagt ggcagaccta caggggcaac 3780 agcactggca ccctgatggt gttctttggc aatgtggaca gctctggcat caagcacaac 3840 atcttcaacc cccccatcat tgccagatac atcaggctgc accccaccca ctacagcatc 3900 aggagcaccc tgaggatgga gctgatgggc tgtgacctga acagctgcag catgcccctg 3960 ggcatggaga gcaaggccat ctctgatgcc cagatcactg ccagcagcta cttcaccaac 4020 atgtttgcca cctggagccc cagcaaggcc aggctgcacc tgcagggcag gagcaatgcc 4080 tggaggcccc aggtcaacaa ccccaaggag tggctgcagg tggacttcca gaagaccatg 4140 aaggtgactg gggtgaccac ccagggggtg aagagcctgc tgaccagcat gtatgtgaag 4200 gagttcctga tcagcagcag ccaggatggc caccagtgga ccctgttctt ccagaatggc 4260 aaggtgaagg tgttccaggg caaccaggac agcttcaccc ctgtggtgaa cagcctggac 4320 ccccccctgc tgaccagata cctgaggatt cacccccaga gctgggtgca ccagattgcc 4380 ctgaggatgg aggtgctggg ctgtgaggcc caggacctgt actga 4425

[0318] Complementary strand to the FVIII transgene (N6) exemplified by SEQ ID NO: 11 tacgtctaac tcgactcgtg gacgaagaag gacacggacg actccaagac gaagagacgg 60 tggtcctcta tgatggaccc ccgacacctc gactcgaccc tgatgtacgt cagactggac 120 cccctcgacg gacacctacg gtccaagggg gggtctcacg ggttctcgaa ggggaagttg 180 tggagacacc acatgttctt ctgggacaaa cacctcaagt gactggtgga caagttgtaa 240 cggttcgggt ccggggggac ctacccggac gacccggggt ggtaggtccg actccacata 300 ctgtgacacc actagtggga cttcttgtac cggtcggtgg gacactcgga cgtacgacac 360 ccccactcga tgaccttccg gagactcccc cgactcatac tactggtctg gtcggtctcc 420 ctcttcctcc tactgttcca caagggaccc ccgtcggtgt ggatacacac cgtccacgac 480 ttcctcttac cggggtaccg gagactgggg gacacggact ggatgtcgat ggactcggta 540 cacctggacc acttcctgga cttgagaccg gactaacccc gggacgacca cacgtccctc 600 ccgtcggacc ggttcctctt ctgggtctgg gacgtgttca agtaggacca caaacgacac 660 aaactactcc cgttctcgac cgtgagactt tggttcttgt cggactacgt cctgtcccta 720 cgacggagac ggtcccggac cgggttctac gtgtgacact taccgataca cttgtcctcg 780 gacggaccgg actaaccgac ggtgtcctt agacacatga ccgtacacta accgtacccg 840 tggtggggac tccacgtgtc gtagaaggac ctcccggtgt ggaaggacca gtccttggtg 900 tccgtccggt cggacctcta gtcggggtag tggaaggact gacgggtctg ggacgactac 960 ctggacccgg tcaaggacga caagacggtg tagtcgtcgg tggtcgtact accgtacctc 1020 cggatacact tccacctgtc gacggactc ctcggggtcg actcctactt cttgttactc 1080 ctccgactcc tgatactact actggactga ctgagactct acctacacca ctccaaacta 1140 ctactgttgt cggggtcgaa gtaggtctag tccagacacc ggttcttcgt ggggttctgg 1200 acccacgtga tgtaacgacg actcctcctc ctgaccctga tacggggga ccacgaccgg 1260 ggactactgt cctcgatgtt ctcggtcatg gacttgttac cgggggtctc ctaaccgtcc 1320 ttcatgttct tccagtccaa gtaccggatg tgactacttt ggaagttctg gtccctccgg 1380 taggtcgtac tcagaccgta ggacccgggg gacgacatac ccctccaccc cctgtgggac 1440 gactagtaga agttcttggt ccggtcgtcc gggatgttgt agatgggggt accgtagtga 1500 ctacactccg gggacatgtc gtcctccgac gggttccccc acttcgtgga cttcctgaag 1560 gggtaggacg gacccctcta gaagttcatg ttcacctgac actgacacct cctaccgggg 1620 tggttcagac tggggtccac ggactggtct atgatgtcgt cgaaacactt gtacctctcc 1680 ctggaccgga gaccggacta accgggggac gactagacga tgttcctcag acacctggtc 1740 tccccgttgg tctagtacag actgttctcc ttacactagg acaagagaca caaactactc 1800 ttgtcctcga ccatggactg actcttgtag gtctccaagg acgggttggg acgaccccac 1860 gtcgacctcc tgggactcaa ggtccggtcg ttgtagtacg tgtcgtagtt accgatacac 1920 aaactgtcgg acgtcgacag acacacggac gtactccacc ggatgaccat gtaggactcg 1980 taaccccggg tctgactgaa ggacagacac aagaagagac cgatgtggaa gttcgtgttc 2040 taccacatac tcctgtggga ctgggacaag gggaagagac ccctctgaca caagtactcg 2100 tacctcttgg gaccggacac ctaagacccg acggtgttga gactgaagtc cttgtccccg 2160 tactgacggg acgactttca gaggtcgaca ctgttcttgt gacccctgat gatactcctg 2220 tcgatactcc tgtagagacg gatggacgac tcgttcttgt tacggtaact cgggtcctcg 2280 aagtcggtct tgtcgtccgt ggggtcgtgg tccgtcttcg tcaagttacg gtggtggtag 2340 ggactcttac tgtatctctt ctgtctgggt accaaacggg tggcctgggg gtacgggttc 2400 taggtcttac actcgtcgag actggacgac tacgacgact ccgtctcggg gtggggggta 2460 ccggactcgg acagactgga cgtcctccgg ttcatacttt ggaagagact actggggtcg 2520 ggaccccggt aactgtcgtt gttgtcggac agactctact gggtgaagtc cggggtcgac 2580 gtggtgagac ccctgtacca caagtgggga ctcagaccgg acgtcgactc cgacttactc 2640 ttcgacccgt ggtgacgacg gtgactcgac ttcttcgacc tgaagtttca gaggtcgtgg 2700 tcgttgttgg actagtcgtg gtaggggaga ctgttggacc gacgaccgtg actgttgtgg 2760 tcgtcggacc cgggggggtc gtacggacac gtgatactgt cggtcgacct gtggtgggac 2820 aaaccgttct tctcgtcggg ggactgactc agacccccgg gggactcgga cagactcctc 2880 ttgttactgt cgttcgacga cctcagaccg gactacttgt cggtcctctc gtcgaccccg 2940 ttcttacact cgtcgtccct ctagtggtcc tggtgggacg tcagactggt cctcctctaa 3000 ctgatactac tgtggtagag acacctctac ttcttcctcc tgaaactgta gatgctgctc 3060 ctgctcttgg tctcggggtc ctcgaaggtc ttcttctggt ccgtgatgaa gtaacgacga 3120 cacctctccg acaccctgat accgtactcg tcgtcggggg tacacgactc cttgtcccgg 3180 gtcagaccga gacacggggt caagttcttc caccacaagg tcctcaagtg actaccgtcg 3240 aagtgggtcg gggacatgtc tcccctcgac ttactcgtgg acccggacga cccggggatg 3300 tagtcccgac tccacctcct gttgtagtac cactggaagt ccttggtccg gtcgtccggg 3360 atgtcgaaga tgtcgtcgga ctagtcgata ctcctcctgg tctccgtccc ccgactcggg 3420 tccttcttga aacacttcgg gttactttgg ttctggatga agaccttcca cgtcgtggtg 3480 taccggggt ggttcctact caaactgacg ttccggaccc ggatgaagag actacacctg 3540 gacctcttcc tacacgtgag accggactaa ccgggggacg accacacggt gtggttgtgg 3600 gacttgggac gggtaccgtc cgtccactga cacgtcctca aacgggacaa gaagtggtag 3660 aaactacttt ggttctcgac catgaagtga ctcttgtacc tctccttgac gtcccggggg 3720 acgttgtagg tctacctcct ggggtggaag ttcctcttga tgtccaaggt acggtagtta 3780 ccgatgtagt acctgtggga cggaccggac cactaccggg tcctggtctc ctagtccacc 3840 atggacgact cgtacccgtc gttactcttg taggtgtcgt aggtgaagag accggtacac 3900 aagtgacact ccttcttcct cctcatgttc taccgggaca tgttggacat gggaccccac 3960 aaactctgac acctctacga cgggtcgttc cgaccgtaga cctcccacct cacggactaa 4020 cccctcgtgg acgtacgacc gtactcgtgg gacaaggacc acatgtcgtt gttcacggtc 4080 tggggggacc cgtaccggag accggtgtag tccctgaagg tctagtgacg gagaccggtc 4140 ataccggtca cccggggggtt cgaccggtcc gacgtgatga gaccgtcgta gttacggacc 4200 tcgtggttcc tcgggaagtc gacctagttc cacctggacg accggggta ctagtaggta 4260 ccgtagttct gggtcccccg gtccgtcttc aagtcgtcgg acatgtagtc ggtcaagtag 4320 tagtacatgt cggacctacc gttcttcacc gtctggatgt ccccgttgtc gtgaccgtgg 4380 gactaccaca agaaaccgtt acacctgtcg agaccgtagt tcgtgttgta gaagttgggg 4440 gggtagtaac ggtctatgta gtccgacgtg gggtgggtga tgtcgtagtc ctcgtgggac 4500 tcctacctcg actacccgac actggacttg tcgacgtcgt acggggaccc gtacctctcg 4560 ttccggtaga gactacgggt ctagtgacgg tcgtcgatga agtggttgta caaacggtgg 4620 acctcggggt cgttccggtc cgacgtggac gtcccgtcct cgttacggac ctccggggtc 4680 cagttgttgg ggttcctcac cgacgtccac ctgaaggtct tctggtactt ccactgaccc 4740 cactggtggg tccccccactt ctcggacgac tggtcgtaca tacacttcct caaggactag 4800 tcgtcgtcgg tcctaccggt ggtcacctgg gaagaagg tcttaccgtt ccacttccac 4860 aaggtcccgt tggtcctgtc gaagtgggga caccacttgt cggacctggg gggggacgac 4920 tggtctatgg actcctaagt gggggtctcg acccacgtgg tctaacggga ctcctacctc 4980 cacgacccga cactccgggt cctggacatg act 5013

[0319] Complementary strand to the FVIII transgene (V3) exemplified by SEQ ID NO: 12<着 tacgtctaac tcgactcgtg gacgaagaag gacacggacg actccaagac gaagagacgg 60 tggtcctcta tgatggaccc ccgacacctc gactcgaccc tgatgtacgt cagactggac 120 cccctcgacg gacacctacg gtccaagggg gggtctcacg ggttctcgaa ggggaagttg 180 tggagacacc acatgttctt ctgggacaaa cacctcaagt gactggtgga caagttgtaa 240 cggttcgggt ccggggggac ctacccggac gacccggggt ggtaggtccg actccacata 300 ctgtgacacc actagtggga cttcttgtac cggtcggtgg gacactcgga cgtacgacac 360 ccccactcga tgaccttccg gagactcccc cgactcatac tactggtctg gtcggtctcc 420 ctcttcctcc tactgttcca caagggaccc ccgtcggtgt ggatacacac cgtccacgac 480 It should be noted that the "<着 " in your original text seems to be an incorrect tag. I've translated it as is, but it might need to be corrected in the original context.ttcctcttac cggggtaccg gagactgggg gacacggact ggatgtcgat ggactcggta 540 cacctggacc acttcctgga cttgagaccg gactaacccc gggacgacca cacgtccctc 600 ccgtcggacc ggttcctctt ctgggtctgg gacgtgttca agtaggacca caaacgacac 660 aaactactcc cgttctcgac cgtgagactt tggttcttgt cggactacgt cctgtcccta 720 cgacggagac ggtcccggac cgggttctac gtgtgacact taccgataca cttgtcctcg 780 gacggaccgg actaaccgac ggtgtcctt agacacatga ccgtacacta accgtacccg 840 tggtggggac tccacgtgtc gtagaaggac ctcccggtgt ggaaggacca gtccttggtg 900 tccgtccggt cggacctcta gtcggggtag tggaaggact gacgggtctg ggacgactac 960 ctggacccgg tcaaggacga caagacggtg tagtcgtcgg tggtcgtact accgtacctc 1020 cggatacact tccacctgtc gacggactc ctcggggtcg actcctactt cttgttactc 1080 ctccgactcc tgatactact actggactga ctgagactct acctacacca ctccaaacta 1140 ctactgttgt cggggtcgaa gtaggtctag tccagacacc ggttcttcgt ggggttctgg 1200 acccacgtga tgtaacgacg actcctcctc ctgaccctga tacgggggga ccacgaccgg 1260 ggactactgt cctcgatgtt ctcggtcatg gacttgttac cgggggtctc ctaaccgtcc 1320 ttcatgttct tccagtccaa gtaccggatg tgactacttt ggaagttctg gtccctccgg 1380 taggtcgtac tcagaccgta ggacccgggg gacgacatac ccctccaccc cctgtgggac 1440 gactagtaga agttcttggt ccggtcgtcc gggatgttgt agatgggggt accgtagtga 1500 ctacactccg gggacatgtc gtcctccgac gggttccccc acttcgtgga cttcctgaag 1560 gggtaggacg gacccctcta gaagttcatg ttcacctgac actgacacct cctaccgggg 1620 tggttcagac tggggtccac ggactggtct atgatgtcgt cgaaacactt gtacctctcc 1680 ctggaccgga gaccggacta accgggggac gactagacga tgttcctcag acacctggtc 1740 tccccgttgg tctagtacag actgttctcc ttacactagg acaagagaca caaactactc 1800 ttgtcctcga ccatggactg actcttgtag gtctccaagg acgggttggg acgaccccac 1860 gtcgacctcc tgggactcaa ggtccggtcg ttgtagtacg tgtcgtagtt accgatacac 1920 aaactgtcgg acgtcgacag acacacggac gtactccacc ggatgaccat gtaggactcg 1980 taaccccggg tctgactgaa ggacagacac aagaagagac cgatgtggaa gttcgtgttc 2040 taccacatac tcctgtggga ctgggacaag gggaagagac ccctctgaca caagtactcg 2100 tacctcttgg gaccggacac ctaagacccg acggtgttga gactgaagtc cttgtccccg 2160 tactgacggg acgactttca gaggtcgaca ctgttcttgt gacccctgat gaactcctg 2220 tcgatactcc tgtagagacg gatggacgac tcgttcttgt tacggtaact cgggtcctcg 2280 aagtcggtct tacggtgatt acacagattg ttgtcgttgt ggtcgttact gtcgttacac 2340 agagggggtc acgactttc cgtggtctcc ctctagtggt cctggtggga cgtcagactg 2400 gtcctcctct aactgatact actgtggtag agacacctct acttcttcct cctgaaactg 2460 tagatgctgc tcctgctctt ggtctcgggg tcctcgaagg tcttcttctg gtccgtgatg 2520 aagtaacgac gacacctctc cgacaccctg ataccgtact cgtcgtcggg ggtacacgac 2580 tccttgtccc gggtcagacc gagacacggg gtcaagttct tccaccacaa ggtcctcaag 2640 tgactaccgt cgaagtgggt cggggacatg tctcccctcg acttactcgt ggacccggac 2700 gacccgggga tgtagtcccg actccacctc ctgttgtagt accactggaa gtccttggtc 2760 cggtcgtccg ggatgtcgaa gatgtcgtcg gactagtcga tactcctcct ggtctccgtc 2820 ccccgactcg ggtccttctt gaaacacttc gggttacttt ggttctggat gaagaccttc 2880 cacgtcgtgg tgtaccgggg gtggttccta ctcaaactga cgttccggac ccggatgaag 2940 agactacacc tggacctctt cctacacgtg agaccggact aaccggggga cgaccacacg 3000 gtgtggttgt gggacttggg acgggtaccg tccgtccact gacacgtcct caaacgggac 3060 aagaagtggt agaaactact ttggttctcg accatgaagt gactcttgta cctctccttg 3120 acgtcccggg ggacgttgta ggtctacctc ctggggtgga agttcctctt gatgtccaag 3180 gtacggtagt taccgatgta gtacctgtgg gacggaccgg accactaccg ggtcctggtc 3240 tcctagtcca ccatggacga ctcgtacccg tcgttactct tgtaggtgtc gtaggtgaag 3300 agaccggtac acaagtgaca ctccttcttc ctcctcatgt tctaccggga catgttggac 3360 atgggacccc acaaactctg acacctctac gacgggtcgt tccgaccgta gacctcccac 3420 ctcacggact aacccctcgt ggacgtacga ccgtactcgt gggacaagga ccacatgtcg 3480 ttgttcacgg tctggggga cccgtaccgg agaccggtgt agtccctgaa ggtctagtga 3540 cggagaccgg tcataccggt cacccgggg ttcgaccggt ccgacgtgat gagaccgtcg 3600 tagttacgga cctcgtggtt cctcgggaag tcgacctagt tccacctgga cgaccgggg 3660 tactagtagg taccgtagtt ctgggtcccc cggtccgtct tcaagtcgtc ggacatgtag 3720 tcggtcaagt agtagtacat gtcggaccta ccgttcttca ccgtctggat gtccccgttg 3780 tcgtgaccgt gggactacca caagaaaccg ttacacctgt cgagaccgta gttcgtgttg 3840 tagaagttgg gggggtagta acggtctatg tagtccgacg tggggtggt gatgtcgtag 3900 tcctcgtggg actctacct cgactacccg acactggact tgtcgacgtc gtacggggac 3960 ccgtacctct cgttccggta gagactacgg gtctagtgac ggtcgtcgat gaagtggttg 4020 tacaaacggt ggacctcggg gtcgttccgg tccgacgtgg acgtcccgtc ctcgttacgg 4080 acctccgggg tccagttgtt ggggttcctc accgacgtcc acctgaaggt cttctggtac 4140 ttccactgac cccactggtg ggtcccccac ttctcggacg actggtcgta catacacttc 4200 ctcaaggact agtcgtcgtc ggtcctaccg gtggtcacct gggacaagaa ggtcttaccg 4260 ttccacttcc acaaggtccc gttggtcctg tcgaagtggg gacaccactt gtcggacctg 4320 gggggggacg actggtctat ggactcctaa gtgggggtct cgacccacgt ggtctaacgg 4380 gactcctacc tccacgaccc gacactccgg gtcctggaca tgact 4425

[0320] FVIII polypeptide (N6) exemplified by SEQ ID NO: 13 Met Gln Ile Glu Leu Ser Thr Cys Phe Phe Leu Cys Leu Leu Arg Phe 1 5 10 15 Cys Phe Ser Ala Thr Arg Arg Tyr Tyr Leu Gly Ala Val Glu Leu Ser 20 25 30 Trp Asp Tyr Met Gln Ser Asp Leu Gly Glu Leu Pro Val Asp Ala Arg 35 40 45 Phe Pro Pro Arg Val Pro Lys Ser Phe Pro Phe Asn Thr Ser Val Val 50 55 60 Tyr Lys Lys Thr Leu Phe Val Glu Phe Thr Asp His Leu Phe Asn Ile 65 70 75 80 Ala Lys Pro Arg Pro Pro Trp Met Gly Leu Leu Gly Pro Thr Ile Gln 85 90 95 Ala Glu Val Tyr Asp Thr Val Val Ile Thr Leu Lys Asn Met Ala Ser 100 105 110 His Pro Val Ser Leu His Ala Val Gly Val Ser Tyr Trp Lys Ala Ser 115 120 125 Glu Gly Ala Glu Tyr Asp Asp Gln Thr Ser Gln Arg Glu Lys Glu Asp 130 135 140 Asp Lys Val Phe Pro Gly Gly Ser His Thr Tyr Val Trp Gln Val Leu 145 150 155 160 Lys Glu Asn Gly Pro Met Ala Ser Asp Pro Leu Cys Leu Thr Tyr Ser 165 170 175 Tyr Leu Ser His Val Asp Leu Val Lys Asp Leu Asn Ser Gly Leu Ile 180 185 190 Gly Ala Leu Leu Val Cys Arg Glu Gly Ser Leu Ala Lys Glu Lys Thr 195 200 205 Gln Thr Leu His Lys Phe Ile Leu Leu Phe Ala Val Phe Asp Glu Gly 210 215 220 Lys Ser Trp His Ser Glu Thr Lys Asn Ser Leu Met Gln Asp Arg Asp 225 230 235 240 Ala Ala Ser Ala Arg Ala Trp Pro Lys Met His Thr Val Asn Gly Tyr 245 250 255 Val Asn Arg Ser Leu Pro Gly Leu Ile Gly Cys His Arg Lys Ser Val 260 265 270 Tyr Trp His Val Ile Gly Met Gly Thr Thr Pro Glu Val His Ser Ile 275 280 285 Phe Leu Glu Gly His Thr Phe Leu Val Arg Asn His Arg Gln Ala Ser 290 295 300 Leu Glu Ile Ser Pro Ile Thr Phe Leu Thr Ala Gln Thr Leu Leu Met 305 310 315 320 Asp Leu Gly Gln Phe Leu Leu Phe Cys His Ile Ser Ser His Gln His 325 330 335 Asp Gly Met Glu Ala Tyr Val Lys Val Asp Ser Cys Pro Glu Glu Pro 340 345 350 Gln Leu Arg Met Lys Asn Asn Glu Glu Ala Glu Asp Tyr Asp Asp Asp 355 360 365 Leu Thr Asp Ser Glu Met Asp Val Val Arg Phe Asp Asp Asp Asn Ser 370 375 380 Pro Ser Phe Ile Gln Ile Arg Ser Val Ala Lys Lys His Pro Lys Thr 385 390 395 400 Trp Val His Tyr Ile Ala Ala Glu Glu Glu Asp Trp Asp Tyr Ala Pro 405 410 415 Leu Val Leu Ala Pro Asp Asp Arg Ser Tyr Lys Ser Gln Tyr Leu Asn 420 425 430 Asn Gly Pro Gln Arg Ile Gly Arg Lys Tyr Lys Lys Val Arg Phe Met 435 440 445 Ala Tyr Thr Asp Glu Thr Phe Lys Thr Arg Glu Ala Ile Gln His Glu 450 455 460 Ser Gly Ile Leu Gly Pro Leu Leu Tyr Gly Glu Val Gly Asp Thr Leu 465 470 475 480 Leu Ile Ile Phe Lys Asn Gln Ala Ser Arg Pro Tyr Asn Ile Tyr Pro 485 490 495 His Gly Ile Thr Asp Val Arg Pro Leu Tyr Ser Arg Arg Leu Pro Lys 500 505 510 Gly Val Lys His Leu Lys Asp Phe Pro Ile Leu Pro Gly Glu Ile Phe 515 520 525 Lys Tyr Lys Trp Thr Val Thr Val Glu Asp Gly Pro Thr Lys Ser Asp 530 535 540 Pro Arg Cys Leu Thr Arg Tyr Tyr Ser Ser Phe Val Asn Met Glu Arg 545 550 555 560 Asp Leu Ala Ser Gly Leu Ile Gly Pro Leu Leu Ile Cys Tyr Lys Glu 565 570 575 Ser Val Asp Gln Arg Gly Asn Gln Ile Met Ser Asp Lys Arg Asn Val 580 585 590 Ile Leu Phe Ser Val Phe Asp Glu Asn Arg Ser Trp Tyr Leu Thr Glu 595 600 605 Asn Ile Gln Arg Phe Leu Pro Asn Pro Ala Gly Val Gln Leu Glu Asp 610 615 620 Pro Glu Phe Gln Ala Ser Asn Ile Met His Ser Ile Asn Gly Tyr Val 625 630 635 640 Phe Asp Ser Leu Gln Leu Ser Val Cys Leu His Glu Val Ala Tyr Trp 645 650 655 Tyr Ile Leu Ser Ile Gly Ala Gln Thr Asp Phe Leu Ser Val Phe Phe 660 665 670 Ser Gly Tyr Thr Phe Lys His Lys Met Val Tyr Glu Asp Thr Leu Thr 675 680 685 Leu Phe Pro Phe Ser Gly Glu Thr Val Phe Met Ser Met Glu Asn Pro 690 695 700 Gly Leu Trp Ile Leu Gly Cys His Asn Ser Asp Phe Arg Asn Arg Gly 705 710 715 720 Met Thr Ala Leu Leu Lys Val Ser Ser Cys Asp Lys Asn Thr Gly Asp 725 730 735 Tyr Tyr Glu Asp Ser Tyr Glu Asp Ile Ser Ala Tyr Leu Leu Ser Lys 740 745 750 Asn Asn Ala Ile Glu Pro Arg Ser Phe Ser Gln Asn Ser Arg His Pro 755 760 765 Ser Thr Arg Gln Lys Gln Phe Asn Ala Thr Thr Ile Pro Glu Asn Asp 770 775 780 Ile Glu Lys Thr Asp Pro Trp Phe Ala His Arg Thr Pro Met Pro Lys 785 790 795 800 Ile Gln Asn Val Ser Ser Ser Asp Leu Leu Met Leu Leu Arg Gln Ser 805 810 815 Pro Thr Pro His Gly Leu Ser Leu Ser Asp Leu Gln Glu Ala Lys Tyr 820 825 830 Glu Thr Phe Ser Asp Asp Pro Ser Pro Gly Ala Ile Asp Ser Asn Asn 835 840 845 Ser Leu Ser Glu Met Thr His Phe Arg Pro Gln Leu His His Ser Gly 850 855 860 Asp Met Val Phe Thr Pro Glu Ser Gly Leu Gln Leu Arg Leu Asn Glu 865 870 875 880 Lys Leu Gly Thr Thr Ala Ala Thr Glu Leu Lys Lys Leu Asp Phe Lys 885 890 895 Val Ser Ser Thr Ser Asn Asn Leu Ile Ser Thr Ile Pro Ser Asp Asn 900 905 910 Leu Ala Ala Gly Thr Asp Asn Thr Ser Ser Leu Gly Pro Pro Ser Met 915 920 925 Pro Val His Tyr Asp Ser Gln Leu Asp Thr Thr Leu Phe Gly Lys Lys 930 935 940 Ser Ser Pro Leu Thr Glu Ser Gly Gly Pro Leu Ser Leu Ser Glu Glu 945 950 955 960 Asn Asn Asp Ser Lys Leu Leu Glu Ser Gly Leu Met Asn Ser Gln Glu 965 970 975 Ser Ser Trp Gly Lys Asn Val Ser Ser Arg Glu Ile Thr Arg Thr Thr 980 985 990 Leu Gln Ser Asp Gln Glu Glu Ile Asp Tyr Asp Asp Thr Ile Ser Val 995 1000 1005 Glu Met Lys Lys Glu Asp Phe Asp Ile Tyr Asp Glu Asp Glu Asn 1010 1015 1020 Gln Ser Pro Arg Ser Phe Gln Lys Lys Thr Arg His Tyr Phe Ile 1025 1030 1035 Ala Ala Val Glu Arg Leu Trp Asp Tyr Gly Met Ser Ser Ser Pro 1040 1045 1050 His Val Leu Arg Asn Arg Ala Gln Ser Gly Ser Val Pro Gln Phe 1055 1060 1065 Lys Lys Val Val Phe Gln Glu Phe Thr Asp Gly Ser Phe Thr Gln 1070 1075 1080 Pro Leu Tyr Arg Gly Glu Leu Asn Glu His Leu Gly Leu Leu Gly 1085 1090 1095 Pro Tyr Ile Arg Ala Glu Val Glu Asp Asn Ile Met Val Thr Phe 1100 1105 1110 Arg Asn Gln Ala Ser Arg Pro Tyr Ser Phe Tyr Ser Ser Leu Ile 1115 1120 1125 Ser Tyr Glu Glu Asp Gln Arg Gln Gly Ala Glu Pro Arg Lys Asn 1130 1135 1140 Phe Val Lys Pro Asn Glu Thr Lys Thr Tyr Phe Trp Lys Val Gln 1145 1150 1155 His His Met Ala Pro Thr Lys Asp Glu Phe Asp Cys Lys Ala Trp 1160 1165 1170 Ala Tyr Phe Ser Asp Val Asp Leu Glu Lys Asp Val His Ser Gly 1175 1180 1185 Leu Ile Gly Pro Leu Leu Val Cys His Thr Asn Thr Leu Asn Pro 1190 1195 1200 Ala His Gly Arg Gln Val Thr Val Gln Glu Phe Ala Leu Phe Phe 1205 1210 1215 Thr Ile Phe Asp Glu Thr Lys Ser Trp Tyr Phe Thr Glu Asn Met 1220 1225 1230 Glu Arg Asn Cys Arg Ala Pro Cys Asn Ile Gln Met Glu Asp Pro 1235 1240 1245 Thr Phe Lys Glu Asn Tyr Arg Phe His Ala Ile Asn Gly Tyr Ile 1250 1255 1260 Met Asp Thr Leu Pro Gly Leu Val Met Ala Gln Asp Gln Arg Ile 1265 1270 1275 Arg Trp Tyr Leu Leu Ser Met Gly Ser Asn Glu Asn Ile His Ser 1280 1285 1290 Ile His Phe Ser Gly His Val Phe Thr Val Arg Lys Lys Glu Glu 1295 1300 1305 Tyr Lys Met Ala Leu Tyr Asn Leu Tyr Pro Gly Val Phe Glu Thr 1310 1315 1320 Val Glu Met Leu Pro Ser Lys Ala Gly Ile Trp Arg Val Glu Cys 1325 1330 1335 Leu Ile Gly Glu His Leu His Ala Gly Met Ser Thr Leu Phe Leu 1340 1345 1350 Val Tyr Ser Asn Lys Cys Gln Thr Pro Leu Gly Met Ala Ser Gly 1355 1360 1365 His Ile Arg Asp Phe Gln Ile Thr Ala Ser Gly Gln Tyr Gly Gln 1370 1375 1380 Trp Ala Pro Lys Leu Ala Arg Leu His Tyr Ser Gly Ser Ile Asn 1385 1390 1395 Ala Trp Ser Thr Lys Glu Pro Phe Ser Trp Ile Lys Val Asp Leu 1400 1405 1410 Leu Ala Pro Met Ile Ile His Gly Ile Lys Thr Gln Gly Ala Arg 1415 1420 1425 Gln Lys Phe Ser Ser Leu Tyr Ile Ser Gln Phe Ile Ile Met Tyr 1430 1435 1440 Ser Leu Asp Gly Lys Lys Trp Gln Thr Tyr Arg Gly Asn Ser Thr 1445 1450 1455 Gly Thr Leu Met Val Phe Phe Gly Asn Val Asp Ser Ser Gly Ile 1460 1465 1470 Lys His Asn Ile Phe Asn Pro Pro Ile Ile Ala Arg Tyr Ile Arg 1475 1480 1485 Leu His Pro Thr His Tyr Ser Ile Arg Ser Thr Leu Arg Met Glu 1490 1495 1500 Leu Met Gly Cys Asp Leu Asn Ser Cys Ser Met Pro Leu Gly Met 1505 1510 1515 Glu Ser Lys Ala Ile Ser Asp Ala Gln Ile Thr Ala Ser Ser Tyr 1520 1525 1530 Phe Thr Asn Met Phe Ala Thr Trp Ser Pro Ser Lys Ala Arg Leu 1535 1540 1545 His Leu Gln Gly Arg Ser Asn Ala Trp Arg Pro Gln Val Asn Asn 1550 1555 1560 Pro Lys Glu Trp Leu Gln Val Asp Phe Gln Lys Thr Met Lys Val 1565 1570 1575 Thr Gly Val Thr Thr Gln Gly Val Lys Ser Leu Leu Thr Ser Met 1580 1585 1590 Tyr Val Lys Glu Phe Leu Ile Ser Ser Ser Gln Asp Gly His Gln 1595 1600 1605 Trp Thr Leu Phe Phe Gln Asn Gly Lys Val Lys Val Phe Gln Gly 1610 1615 1620 Asn Gln Asp Ser Phe Thr Pro Val Val Asn Ser Leu Asp Pro Pro 1625 1630 1635 Leu Leu Thr Arg Tyr Leu Arg Ile His Pro Gln Ser Trp Val His 1640 1645 1650 Gln Ile Ala Leu Arg Met Glu Val Leu Gly Cys Glu Ala Gln Asp 1655 1660 1665 Leu Tyr 1670

[0321] FVIII polypeptide (V3) exemplified by SEQ ID NO: 14 Met Gln Ile Glu Leu Ser Thr Cys Phe Phe Leu Cys Leu Leu Arg Phe 1 5 10 15 Cys Phe Ser Ala Thr Arg Arg Tyr Tyr Leu Gly Ala Val Glu Leu Ser 20 25 30 Trp Asp Tyr Met Gln Ser Asp Leu Gly Glu Leu Pro Val Asp Ala Arg 35 40 45 Phe Pro Pro Arg Val Pro Lys Ser Phe Pro Phe Asn Thr Ser Val Val 50 55 60 Tyr Lys Lys Thr Leu Phe Val Glu Phe Thr Asp His Leu Phe Asn Ile 65 70 75 80 Ala Lys Pro Arg Pro Pro Trp Met Gly Leu Leu Gly Pro Thr Ile Gln 85 90 95 Ala Glu Val Tyr Asp Thr Val Val Ile Thr Leu Lys Asn Met Ala Ser 100 105 110 His Pro Val Ser Leu His Ala Val Gly Val Ser Tyr Trp Lys Ala Ser 115 120 125 Glu Gly Ala Glu Tyr Asp Asp Gln Thr Ser Gln Arg Glu Lys Glu Asp 130 135 140 Asp Lys Val Phe Pro Gly Gly Ser His Thr Tyr Val Trp Gln Val Leu 145 150 155 160 Lys Glu Asn Gly Pro Met Ala Ser Asp Pro Leu Cys Leu Thr Tyr Ser 165 170 175 Tyr Leu Ser His Val Asp Leu Val Lys Asp Leu Asn Ser Gly Leu Ile 180 185 190 Gly Ala Leu Leu Val Cys Arg Glu Gly Ser Leu Ala Lys Glu Lys Thr 195 200 205 Gln Thr Leu His Lys Phe Ile Leu Leu Phe Ala Val Phe Asp Glu Gly 210 215 220 Lys Ser Trp His Ser Glu Thr Lys Asn Ser Leu Met Gln Asp Arg Asp 225 230 235 240 Ala Ala Ser Ala Arg Ala Trp Pro Lys Met His Thr Val Asn Gly Tyr 245 250 255 Val Asn Arg Ser Leu Pro Gly Leu Ile Gly Cys His Arg Lys Ser Val 260 265 270 Tyr Trp His Val Ile Gly Met Gly Thr Thr Pro Glu Val His Ser Ile 275 280 285 Phe Leu Glu Gly His Thr Phe Leu Val Arg Asn His Arg Gln Ala Ser 290 295 300 Leu Glu Ile Ser Pro Ile Thr Phe Leu Thr Ala Gln Thr Leu Leu Met 305 310 315 320 Asp Leu Gly Gln Phe Leu Leu Phe Cys His Ile Ser Ser His Gln His 325 330 335 Asp Gly Met Glu Ala Tyr Val Lys Val Asp Ser Cys Pro Glu Glu Pro 340 345 350 Gln Leu Arg Met Lys Asn Asn Glu Glu Ala Glu Asp Tyr Asp Asp Asp 355 360 365 Leu Thr Asp Ser Glu Met Asp Val Val Arg Phe Asp Asp Asp Asn Ser 370 375 380 Pro Ser Phe Ile Gln Ile Arg Ser Val Ala Lys Lys His Pro Lys Thr 385 390 395 400 Trp Val His Tyr Ile Ala Ala Glu Glu Glu Asp Trp Asp Tyr Ala Pro 405 410 415 Leu Val Leu Ala Pro Asp Asp Arg Ser Tyr Lys Ser Gln Tyr Leu Asn 420 425 430 Asn Gly Pro Gln Arg Ile Gly Arg Lys Tyr Lys Lys Val Arg Phe Met 435 440 445 Ala Tyr Thr Asp Glu Thr Phe Lys Thr Arg Glu Ala Ile Gln His Glu 450 455 460 Ser Gly Ile Leu Gly Pro Leu Leu Tyr Gly Glu Val Gly Asp Thr Leu 465 470 475 480 Leu Ile Ile Phe Lys Asn Gln Ala Ser Arg Pro Tyr Asn Ile Tyr Pro 485 490 495 His Gly Ile Thr Asp Val Arg Pro Leu Tyr Ser Arg Arg Leu Pro Lys 500 505 510 Gly Val Lys His Leu Lys Asp Phe Pro Ile Leu Pro Gly Glu Ile Phe 515 520 525 Lys Tyr Lys Trp Thr Val Thr Val Glu Asp Gly Pro Thr Lys Ser Asp 530 535 540 Pro Arg Cys Leu Thr Arg Tyr Tyr Ser Ser Phe Val Asn Met Glu Arg 545 550 555 560 Asp Leu Ala Ser Gly Leu Ile Gly Pro Leu Leu Ile Cys Tyr Lys Glu 565 570 575 Ser Val Asp Gln Arg Gly Asn Gln Ile Met Ser Asp Lys Arg Asn Val 580 585 590 Ile Leu Phe Ser Val Phe Asp Glu Asn Arg Ser Trp Tyr Leu Thr Glu 595 600 605 Asn Ile Gln Arg Phe Leu Pro Asn Pro Ala Gly Val Gln Leu Glu Asp 610 615 620 Pro Glu Phe Gln Ala Ser Asn Ile Met His Ser Ile Asn Gly Tyr Val 625 630 635 640 Phe Asp Ser Leu Gln Leu Ser Val Cys Leu His Glu Val Ala Tyr Trp 645 650 655 Tyr Ile Leu Ser Ile Gly Ala Gln Thr Asp Phe Leu Ser Val Phe Phe 660 665 670 Ser Gly Tyr Thr Phe Lys His Lys Met Val Tyr Glu Asp Thr Leu Thr 675 680 685 Leu Phe Pro Phe Ser Gly Glu Thr Val Phe Met Ser Met Glu Asn Pro 690 695 700 Gly Leu Trp Ile Leu Gly Cys His Asn Ser Asp Phe Arg Asn Arg Gly 705 710 715 720 Met Thr Ala Leu Leu Lys Val Ser Ser Cys Asp Lys Asn Thr Gly Asp 725 730 735 Tyr Tyr Glu Asp Ser Tyr Glu Asp Ile Ser Ala Tyr Leu Leu Ser Lys 740 745 750 Asn Asn Ala Ile Glu Pro Arg Ser Phe Ser Gln Asn Ala Thr Asn Val 755 760 765 Ser Asn Asn Ser Asn Thr Ser Asn Asp Ser Asn Val Ser Pro Pro Val 770 775 780 Leu Lys Arg His Gln Arg Glu Ile Thr Arg Thr Thr Leu Gln Ser Asp 785 790 795 800 Gln Glu Glu Ile Asp Tyr Asp Asp Thr Ile Ser Val Glu Met Lys Lys 805 810 815 Glu Asp Phe Asp Ile Tyr Asp Glu Asp Glu Asn Gln Ser Pro Arg Ser 820 825 830 Phe Gln Lys Lys Thr Arg His Tyr Phe Ile Ala Ala Val Glu Arg Leu 835 840 845 Trp Asp Tyr Gly Met Ser Ser Ser Pro His Val Leu Arg Asn Arg Ala 850 855 860 Gln Ser Gly Ser Val Pro Gln Phe Lys Lys Val Val Phe Gln Glu Phe 865 870 875 880 Thr Asp Gly Ser Phe Thr Gln Pro Leu Tyr Arg Gly Glu Leu Asn Glu 885 890 895 His Leu Gly Leu Leu Gly Pro Tyr Ile Arg Ala Glu Val Glu Asp Asn 900 905 910 Ile Met Val Thr Phe Arg Asn Gln Ala Ser Arg Pro Tyr Ser Phe Tyr 915 920 925 Ser Ser Leu Ile Ser Tyr Glu Glu Asp Gln Arg Gln Gly Ala Glu Pro 930 935 940 Arg Lys Asn Phe Val Lys Pro Asn Glu Thr Lys Thr Tyr Phe Trp Lys 945 950 955 960 Val Gln His His Met Ala Pro Thr Lys Asp Glu Phe Asp Cys Lys Ala 965 970 975 Trp Ala Tyr Phe Ser Asp Val Asp Leu Glu Lys Asp Val His Ser Gly 980 985 990 Leu Ile Gly Pro Leu Leu Val Cys His Thr Asn Thr Leu Asn Pro Ala 995 1000 1005 His Gly Arg Gln Val Thr Val Gln Glu Phe Ala Leu Phe Phe Thr 1010 1015 1020 Ile Phe Asp Glu Thr Lys Ser Trp Tyr Phe Thr Glu Asn Met Glu 1025 1030 1035 Arg Asn Cys Arg Ala Pro Cys Asn Ile Gln Met Glu Asp Pro Thr 1040 1045 1050 Phe Lys Glu Asn Tyr Arg Phe His Ala Ile Asn Gly Tyr Ile Met 1055 1060 1065 Asp Thr Leu Pro Gly Leu Val Met Ala Gln Asp Gln Arg Ile Arg 1070 1075 1080 Trp Tyr Leu Leu Ser Met Gly Ser Asn Glu Asn Ile His Ser Ile 1085 1090 1095 His Phe Ser Gly His Val Phe Thr Val Arg Lys Lys Glu Glu Tyr 1100 1105 1110 Lys Met Ala Leu Tyr Asn Leu Tyr Pro Gly Val Phe Glu Thr Val 1115 1120 1125 Glu Met Leu Pro Ser Lys Ala Gly Ile Trp Arg Val Glu Cys Leu 1130 1135 1140 Ile Gly Glu His Leu His Ala Gly Met Ser Thr Leu Phe Leu Val 1145 1150 1155 Tyr Ser Asn Lys Cys Gln Thr Pro Leu Gly Met Ala Ser Gly His 1160 1165 1170 Ile Arg Asp Phe Gln Ile Thr Ala Ser Gly Gln Tyr Gly Gln Trp 1175 1180 1185 Ala Pro Lys Leu Ala Arg Leu His Tyr Ser Gly Ser Ile Asn Ala 1190 1195 1200 Trp Ser Thr Lys Glu Pro Phe Ser Trp Ile Lys Val Asp Leu Leu 1205 1210 1215 Ala Pro Met Ile Ile His Gly Ile Lys Thr Gln Gly Ala Arg Gln 1220 1225 1230 Lys Phe Ser Ser Leu Tyr Ile Ser Gln Phe Ile Ile Met Tyr Ser 1235 1240 1245 Leu Asp Gly Lys Lys Trp Gln Thr Tyr Arg Gly Asn Ser Thr Gly 1250 1255 1260 Thr Leu Met Val Phe Phe Gly Asn Val Asp Ser Ser Gly Ile Lys 1265 1270 1275 His Asn Ile Phe Asn Pro Pro Ile Ile Ala Arg Tyr Ile Arg Leu 1280 1285 1290 His Pro Thr His Tyr Ser Ile Arg Ser Thr Leu Arg Met Glu Leu 1295 1300 1305 Met Gly Cys Asp Leu Asn Ser Cys Ser Met Pro Leu Gly Met Glu 1310 1315 1320 Ser Lys Ala Ile Ser Asp Ala Gln Ile Thr Ala Ser Ser Tyr Phe 1325 1330 1335 Thr Asn Met Phe Ala Thr Trp Ser Pro Ser Lys Ala Arg Leu His 1340 1345 1350 Leu Gln Gly Arg Ser Asn Ala Trp Arg Pro Gln Val Asn Asn Pro 1355 1360 1365 Lys Glu Trp Leu Gln Val Asp Phe Gln Lys Thr Met Lys Val Thr 1370 1375 1380 Gly Val Thr Thr Gln Gly Val Lys Ser Leu Leu Thr Ser Met Tyr 1385 1390 1395 Val Lys Glu Phe Leu Ile Ser Ser Ser Gln Asp Gly His Gln Trp 1400 1405 1410 Thr Leu Phe Phe Gln Asn Gly Lys Val Lys Val Phe Gln Gly Asn 1415 1420 1425 Gln Asp Ser Phe Thr Pro Val Val Asn Ser Leu Asp Pro Pro Leu 1430 1435 1440 Leu Thr Arg Tyr Leu Arg Ile His Pro Gln Ser Trp Val His Gln 1445 1450 1455 Ile Ala Leu Arg Met Glu Val Leu Gly Cys Glu Ala Gln Asp Leu 1460 1465 1470 Tyr

[0322] CFTR transgene (soCFTR2) exemplified by SEQ ID NO: 15 gctagccac atgcagagaa gccctctgga gaaggcctct gtggtgagca agctgttctt 60 cagctggacc aggcccatcc tgaggaaggg ctacaggcag agactggagc tgtctgacat 120 ctaccagatc ccctctgtgg actctgctga caacctgtct gagaagctgg agagggagtg 180 ggatagagag ctggccagca agaagaaccc caagctgatc aatgccctga ggagatgctt 240 cttctggaga ttcatgttct atggcatctt cctgtacctg ggggaagtga ccaaggctgt 300 gcagcctctg ctgctgggca gaatcattgc cagctatgac cctgacaaca aggaggagag 360 gagcattgcc atctacctgg gcattggcct gtgcctgctg ttcattgtga ggaccctgct 420 gctgcaccct gccatctttg gcctgcacca cattggcatg cagatgagga ttgccatgtt 480 cagcctgatc tacaagaaaa ccctgaagct gtccagcaga gtgctggaca agatcagcat 540 tggccagctg gtgagcctgc tgagcaacaa cctgaacaag tttgatgagg gcctggccct 600 ggcccacttt gtgtggattg cccctctgca ggtggccctg ctgatgggcc tgatttggga 660 gctgctgcag gcctctgcct tttgtggcct gggcttcctg attgtgctgg ccctgtttca 720 ggctggcctg ggcaggatga tgatgaagta cagggaccag agggcaggca agatcagtga 780 gaggctggtg atcacctctg agatgattga gaacatccag tctgtgaagg cctactgttg 840 ggaggaagct atggagaaga tgattgaaaa cctgaggcag acagagctga agctgaccag 900 gaaggctgcc tatgtgagat acttcaacag ctctgccttc ttcttctctg gcttctttgt 960 ggtgttcctg tctgtgctgc cctatgccct gatcaagggg atcatcctga gaaagatttt 1020 caccaccatc agcttctgca ttgtgctgag gatggctgtg accagacagt tcccctgggc 1080 tgtgcagacc tggtatgaca gcctgggggc catcaacaag atccaggact tcctgcagaa 1140 gcaggagtac aagaccctgg agtacaacct gaccaccaca gaagtggtga tggagaatgt 1200 gacagccttc tgggaggagg gctttgggga gctgtttgag aaggccaagc agaacaacaa 1260 caacagaaag accagcaatg gggatgactc cctgttcttc tccaacttct ccctgctggg 1320 cacacctgtg ctgaaggaca tcaacttcaa gattgagagg gggcagctgc tggctgtggc 1380 tggatctaca ggggctggca agaccagcct gctgatgatg atcatgggg agctggagcc 1440 ttctgagggc agatcaagc actctggcag gatcagcttt tgcagccag tcagctagc 1500 catgcctggc accatcagg agacatcat ctttggagtg agctatgatg agtacagata 1560 caggagtgtg atcaggcct gccagctgga ggaggacatc agcagtttg ctgagaagga 1620 siacattgtg ctgggggagg gaggcattac actgtctgggg ggccagagag ccagaatcag 1680 cctggccagg gctgtgtaca aggatgctga cctgtacctg ctggactccc cctttggcta 1740 cctggatgtg ctgacagaga aggatttt tgagagctgt gtgtgcaagc tgatggccaa 1800 cagaccaga atcctggtga caccacagat ggagcacctg aagaaggctg acagatcct 1860 gatcctgcat gagggcagca gctacttcta tgggaccttc tctgagctgc agaacctgca 1920 gcctgacttc agctctaagc tgatggctg tgacagcttt gaccagttct ctgctgagag 1980 gaggaacagc atcctgacag agaccctgca cagattcagc ctggagggag atgcccctgt 2040 gagctgaca gagachaga aggagctt gaggagaca ggggagtttg ggagagag 2100 gaagaactcc atcctgaacc ccatcaacag catcaggaag ttcagcattg tgcagaaaac 2160 ccccctgcag atgaatggca ttgaggaaga ttctgatgag cccctggaga ggagactgag 2220 cctggtgcct gattctgagc agggagaggc catcctgcct aggatctctg tgatcagcac 2280 aggccctaca ctgcaggcca gaaggaggca gtctgtgctg aacctgatga cccactctgt 2340 gaaccagggc cagaacatcc acaggaaaac cacagcctcc accaggaaag tgagcctggc 2400 ccctcaggcc aatctgacag agctggacat ctacagcagg aggctgtctc aggagacagg 2460 cctggagatt tctgaggaga tcaatgagga ggacctgaaa gagtgcttct ttgatgacat 2520 ggagagcatc cctgctgtga ccacctggaa cacctacctg agatacatca cagtgcacaa 2580 gagcctgatc tttgtgctga tctggtgcct ggtgatcttc ctggctgaag tggctgcctc 2640 tctggtggtg ctgtggctgc tgggaaacac cccactgcag gacaagggca acagcaccca 2700 cagcaggaac aacagctatg ctgtgatcat cacctccacc tccagctact atgtgttcta 2760 catctatgtg ggagtggctg ataccctgct ggctatgggc ttctttagag gcctgcccct 2820 ggtgcacaca ctgatcacag tgagcaagat cctccaccac aagatgctgc actctgtgct 2880 gcaggctcct atgagcaccc tgaataccct gaaggctggg ggcatcctga acagattctc 2940 caaggatatt gccatcctgg atgacctgct gcctctcacc atctttgact tcatccagct 3000 gctgctgatt gtgattgggg ccattgctgt ggtggcagtg ctgcagccct acatctttgt 3060 ggccacagtg cctgtgattg tggccttcat catgctgagg gcctactttc tgcagacctc 3120 ccagcagctg aagcagctgg agtctgaggg cagaagcccc atcttcaccc acctggtgac 3180 aagcctgaag ggcctgtgga ccctgagagc ctttggcagg cagccctact ttgagaccct 3240 gttccacaag gccctgaacc tgcacacagc caactggttc ctctacctgt ccaccctgag 3300 atggttccag atgagaattg agatgatctt tgtcatcttc ttcattgctg tgaccttcat 3360 cagcattctg accacaggag agggagaggg cagagtgggc attatcctga ccctggccat 3420 gaacatcatg agcacactgc agtgggcagt gaacagcagc attgatgtgg acagcctgat 3480 gaggagtgtg agcagagtgt tcaagttcat tgatatgccc acagagggca agcctaccaa 3540 gagcaccaag ccctacaga atggccagct gagcaaagtg atgatcattg agaacagcca 3600 tgtgaagaag gatgatatct ggcccagtgg agccagatg acagtgaagg acctgacagc 3660 caagtacaca gaggggggca atgctatcct ggagaacatc tccttcagca tctcccctgg 3720 ccagagagtg ggactgctgg gagaacagg ctctggcaag tctaccctgc tgtctgcctt 3780 cctgaggctg ctgaacacag agggagagat ccagattgat ggagtgtcct gggacagcat 3840 cacactgcag cagtggagga aggccttgg tgtgatcccc cagaagtgt tcatcttcag 3900 tggcaccttc aggagagacc tggacccta tgagcagtgg tctgaccagg agatttggaa 3960 agtggctgat gaagtgggcc tgagaagtgt gattgagcag ttccctggca agctggactt 4020 tgtcctggtg gatgggggct gtgtgctgag ccatggccac aagcagctga tgtgcctggc 4080 cagatcagtg ctgagcagg ccagatcct gctgctgat gagccttctg cccacctgga 4140 tcctgtgacc taccagatca tcaggac cctcaagcag gcctttgctg acctcacagt 4200 catcctgtgt gagcacagga tgaggccat gctggagtgc cagcagttcc tggtgattga 4260 ggagaacaaa gtgaggcagt atgacagcat ccagaagctg ctgaatgaga ggagcctgtt 4320 caggcaggcc atcagcccct ctgatagagt gaagctgttc ccccacagga acagctccaa 4380 gtgcaagagc aagccccaga ttgctgccct gaaggaggag acagaggagg aagtgcagga 4440 caccaggctg tgagggccc 4459

[0323] SEQ ID NO: 16 Exemplary CFTR Polypeptide

[0324] SEQ ID NO: 17 Exemplary hGM-CSF transgene ATGTGGCTGCAGAGCCTGCTGCTCTTGGGCACTGTGGCCTGCAGCATCTCTGCACCCGCCCGCTCGCCCAGCCCCAGCACGCAGCCCTGGGAGCATGTGAATGCCATCCAGGAGGCCCGGCGTCTCCTGAACCTGAGTAGAGACACTGCTGCTGAGATGAATGAAACAGTAGAAGTCATCTCAGAAATGTTTGACCTCCAGGAGCCGACCTGCCTAC AGACCCGCCTGGAGCTGTACAAGCAGGGCCTGCGGGGCAGCCTCACCAAGCTCAAGGGCCCCTTGACCATGATGGCCAGCCACTACAAGCAGCACTGCCCTCCAACCCCGGAAACTTCCTGTGCAACCCAGATTATCACCTTTGAAAGTTTCAAAGAGAACCTGAAGGACTTTCTGCTTGTCATCCCCTTTGACTGCTGGGAGCCAGTCCAGGAGTGA

[0325] SEQ ID NO: 18 Exemplary hGM-CSF Polypeptide MWLQSLLLLGTVACSISAPARSPSPSTQPWEHVNAIQEARRLLNLSRDTAAEMNETVEVISEMFDLQEPTCLQTRLELYKQGLRGSLTKLKGPLTMMASHYKQHCPPTPETSCATQIITFESFKENLKDFLLVIPFDCWEPVQE

[0326] SEQ ID NO: 19 Exemplary mGM-CSF transgene ATGTGGCTGCAGAACCTGCTGTTCCTGGGCATTGTGGTGTACAGCCTGTCTGCCCCTACAAGATCCCCTATCACAGTGACCAGACCTTGGAAACATGTGGAAGCCATCAAAGAGGCCCTGAATCTGCTGGATGACATGCCTGTGACACTGAATGAAGAGGTGGAAGTGGTGTCCAATGAGTTCAGCTTCAAGAAACTGACCTGTGTGCAGACC AGGCTGAAGATTTTTGAGCAGGGCCTGAGAGGCAACTTCACCAAGCTGAAAGGGGCTCTGAACATGACAGCCAGCTACTACCAGACCTACTGTCCTCCTACACCTGAGACAGACTGTGAAACCCAAGTGACCACCTATGCTGACTTCATTGACAGCCTCAAGACCTTCCTGACAGACATCCCCTTTGAGTGCAAGAAACCTGGCCAGAAGTGA

[0327] SEQ ID NO: 20 Exemplary mGM-CSF Polypeptide MWLQNLLFLGIVVYSLSAPTRSPITVTRPWKHVEAIKEALNLLDDMPVTLNEEVEVVSNEFSFKKLTCVQTRLKIFEQGLRGNFTKLKGALNMTASYYQTYCPPTPETDCETQVTTYADFIDSLKTFLTDIPFECKKPGQK

[0328] SEQ ID NO: 21 Exemplary human DCN (decorin) transgene

[0329] SEQ ID NO: 22 Exemplary human decorin polypeptide MKATIILLLLAQVSWAGPFQQRGLFDFMLEDEASGIGPEVPDDRDFEPSLGPVCPFRCQCHLRVVQCSDLGLDKVPKDLPPDTTLLDLQNNKITEIKDGDFKNLKNLHALILVNNKISKVSPGAFTPLVKLERLYLSKNQLKELPEKMPKTLQELRAHENEITKVRKVTFNGLNQMIVI ELGTNPLKSSGIENGAFQGMKKLSYIRIADTNITSIPQGLPPSLTELHLDGNKISRVDAASLKGLNNLAKLGLSFNSISAVDNGSLANTPHLRELHLDNNKLTRVPGGLAEHKYIQVVYLHNNNISVVGSSDFCPPGHNTKKASYSGVSLFSNPVQYWEIQPSTFRCVYVRSAIQLGNYK

[0330] SEQ ID NO: 23 Exemplary human TRIM72 transgene

[0331] SEQ ID NO: 24 Exemplary human TRIM72 polypeptide MSAAPGLLHQELSCPLCLQLFDAPVTAECGHSFCRACLGRVAGEPAADGTVLCPCCQAPTRPQALSTNLQLARLVEGLAQVPQGHCEEHLDPLSIYCEQDRALVCGVCASLGSHRGHRL LPAAEAHARLKTQLPQQKLQLQEACMRKEKSVAVLEHQLVEVEETVRQFRGAVGEQLGKMRVFLAALEGSLDREAERVRGEAGVALRRELGSLNSYLEQLRQMEKVLEEVADKPQTEFL MKYCLVTSRLQKILAESPPPARLDIQLPIISDDFKFQVWRKMFRALMPALEELTFDPSSAHPSLVVSSSGRRVECSEQKAPPAGEDPRQFDKAVAVVAHQQLSEGEHYWEVDVGDKPRW ALGVIAAEAPRRGRLHAVPSQGLWLLGLREGKILEAHVEAKEPRALRSPERRPTRIGLYLSFGDGVLSFYDASDADALVPLFAFHERLPRPVYPFFDVCWHDKGKNAQPLLLVGPEGAEA

[0332] SEQ ID NO: 25 Exemplary human ABCA3 (ABCA3) transgene

[0333] SEQ ID NO: 26 Exemplary human ABCA3 polypeptide

[0334] SEQ ID NO: 27 Exemplary hCEF promoter agatctgtta cataacttat ggtaaatggc ctgcctggct gactgcccaa tgacccctgc 60 ccaatgatgt caataatgat gtatgttccc atgtaatgcc aatagggact ttccattgat 120 gtcaatgggt ggagtattta tggtaactgc ccacttggca gtacatcaag tgtatcatat 180 gccaagtatg ccccctattg atgtcaatga tggtaaatgg cctgcctggc attatgccca 240 gtacatgacc ttatgggact ttcctacttg gcagtacatc tatgtattag tcattgctat 300 taccatggga attcactagt ggagaagagc atgcttgagg gctgagtgcc cctcagtggg 360 cagagagcac atggcccaca gtccctgaga agttgggggg aggggtggggc aattgaactg 420 gtgcctagag aaggtggggc ttgggtaaac tgggaaagtg atgtggtgta ctggctccac 480 ctttttcccc agggtggggg agaaccatat ataagtgcag tagtctctgt gaacattcaa 540 gcttctgcct tctccctcct gtgagtttgc tagc 574

[0335] SEQ ID NO: 28 Exemplary CMV promoter ccgcggagat ctcaatattg gccattagcc atattattca ttggttatat agcataaatc 60 atattggct attggccatt gcatacgttg tatctatatc atatatgta catttatatt 120 ggctcatgtc caatgacc gccatgttgg cattgatt tgactagtta ttaatagtaa 180 tcattacgg gtcattagt tcatagccca tatggag tccgcgttac attackacg 240 gtaaatggcc cgcctggctg acccccac gaccccgcc cattgacgtc ataatgacg 300 tatgttccca tagtaacgcc atagggact ttccattgac gtcaatgggt ggagtattta 360 cggtaaactg cccacttgc agtacatcaa gtgtatcata tgccaagtcc gcccctatt 420 gacgtcaatg acggtaaatg gcccgcctgg cattatgccc agtacatgac cttacgggac 480 tttcctactt ggcagtacat ctacgtatta gtcatcgcta ttaccatggt gatgcggttt 540 tggcagtaca ccaatgggcg tggatagcgg ttgactcac ggggatttcc aagtctccac 600 cccattgacg tcaatgggag ttgttttgg caccaaaatc aacgggactt tccaaaatgt 660 cgtaataacc ccgccccgtt gacgcaatg ggcggtaggc gtgtacggtg ggaggtctat 720 ataagcagag ctcgtttagt gaaccgtcag atcactagaa gctttattgc ggtagtttat 780 cacagttaaa ttgctaacgc agtcagtgct tctgacacaa cagtctcgaa cttaagctgc 840 agaagttggt cgtgaggcac tgggcaggct agc 873

[0336] SEQ ID NO: 29 Exemplary EF1a promoter agatccatat ccgcggcaat tttaaaagaa agggaggaat agggggacag acttcagcag 60 agagactaat taatataata acaacacaat tagaaataca acatttacaa accaaaattc 120 aaaaaatttt aaattttaga gccgcggaga tcccgtgagg ctccggtgcc cgtcagtggg 180 cagagcgcac atcgcccaca gtccccgaga agttgggggg aggggtcggc aattgaaccg 240 gtgcctagag aaggtggcgc ggggtaaact gggaaagtga tgtcgtgtac tggctccgcc 300 tttttcccga gggtggggga gaaccgtata taagtgcagt agtcgccgtg aacgttcttt 360 ttcgcaacgg gtttgccgcc agaacacagg ctagc 395

[0337] SEQ ID NO: 30 Plasmid defined in Figure 2A (pDNA1 pGM991) [ka] [ka] [ka] The hCEF promoter is underlined The β-globulin / IgG chimeric intron containing the SIV RRE intron is in italics SIV RRE is italicized and double underlined

[0338] SEQ ID NO: 31 Plasmid defined in Figure 2B (pDNA1 pGM691) attgattatt gactagttat taatagtaat caattacggg gtcattagtt catagcccat 60 atatggagtt ccgcgttaca taacttacgg taaatggccc gcctggctga ccgcccaacg 120 acccccgccc attgacgtca ataatgacgt atgttcccat agtaacgcca atagggactt 180 tccattgacg tcaatgggtg gagtatttac ggtaaactgc ccacttggca gtacatcaag 240 tgtatcatat gccaagtacg ccccctattg acgtcaatga cggtaaatgg ccgcctggc 300 attatgccca gtacatgacc ttatgggact ttcctacttg gcagtacatc tacgtattag 360 tcatcgctat taccatggtc gaggtgagcc ccacgttctg cttcactctc cccatctccc 420 ccccctcccc acccccaatt ttgtatttat ttattttta attattttgt gcagcgatgg 480 gggcgggggg gggggggggg cgcgcgccag gcggggcggg gcggggcgag gggcggggcg 540 gggcgaggcg gagaggtgcg gcggcagcca atcagagcgg cgcgctccga aagtttcctt 600 ttatggcgag gcggcggcgg cggcggccct ataaaaagcg aagcgcgcgg cgggcgggag 660 tcgctgcgcg ctgccttcgc cccgtgcccc gctccgccgc cgcctcgcgc cgcccgcccc 720 ggctctgact gaccgcgtta ctcccacagg tgagcgggcg ggacggccct tctcctccgg 780 gctgtaatta gcgcttggtt taatgacggc ttgtttcttt tctgtggctg cgtgaaagcc 840 ttgaggggct ccgggagggc cctttgtgcg gggggagcgg ctcggggggt gcgtgcgtgt 900 gtgtgtgcgt ggggagcgcc gcgtgcggct ccgcgctgcc cggcggctgt gagcgctgcg 960 ggcgcggcgc ggggctttgt gcgctccgca gtgtgcgcga ggggagcgcg gccgggggcg 1020 gtgccccgcg gtgcgggggg ggctgcgagg ggaacaaagg ctgcgtgcgg ggtgtgtgcg 1080 tgggggggtg agcagggggt gtgggcgcgt cggtcgggct gcaacccccc ctgcaccccc 1140 ctccccgagt tgctgagcac ggcccggctt cgggtgcggg gctccgtacg gggcgtggcg 1200 cggggctcgc cgtgccgggc ggggggtggc ggcaggtggg ggtgccgggc ggggcggggc 1260 cgcctcgggc cggggagggc tcgggggagg ggcgcggcgg cccccggagc gccggcggct 1320 gtcgaggcgc ggcgagccgc agccattgcc tttatggta atcgtgcgag agggcgcagg 1380 gacttcctt gtcccaaatc tgtgcggagc cgaaatctgg gaggcgccgc cgcaccccct 1440 ctagcggcg cggggcgaag cggtgcggcg ccggcaggaa ggaaatgggc gggagggcc 1500 ttcgtgcgtc gccgcgccgc cgtccccttc tccctctcca gcctcggggc tgtccgcggg 1560 gggacggctg ccttcgggg ggacggggca gggcggggtt cggcttctgg cgtgtgaccg 1620 gcggctctag agcctctgct aaccatgttc atgccttctt ctttttccta cagctcctgg 1680 gcaacgtgct ggttattgtg ctgtctcatc attttggcaa agaattgctc gagccaccat 1740 gggagctgcc acatctgccc tgaatagacg gcagctggac cagttcgaga agatcagact 1800 gcggcccaac ggcaagaaga agtaccagat caagcacctg atctgggccg gcaaagagat 1860 ggaaagattc ggcctgcacg agcggctgct ggaaaccgag gaaggctgca agaaattat 1920 cgaggtgctg taccctctgg aacctaccgg ctctgagggc ctgaagtccc tgttcaatct 1980 cgtgtgcgtg ctgtactgcc tgcacaaaga acagaaagtg aaggacaccg aagaggccgt 2040 ggccacagtt agacagcact gccacctggt ggaaaaagag aagtccgcca cagagacaag 2100 cagcggccag aagaagaacg acaagggaat tgctgcccct cctggcggca gccagaattt 2160 tcctgctcag cagcagggaa acgcctgggt gcacgttcca ctgagcccta gaacactgaa 2220 tgcctgggtc aaagccgtgg aagagaagaa gtttggcgcc gagatcgtgc ccatgttcca 2280 ggctctgtct gagggctgca ccccttacga catcaaccag atgctgaacg tgctgggaga 2340 tcaccagggc gctctgcaga tcgtgaaaga gatcatcaac gaagaggctg cccagtggga 2400 cgtgacacat ccattgcctg ctggacctct gccagccgga caactgagag atcctagagg 2460 ctctgatatc gccggcacca ccagctctgt gcaagagcag ctggaatgga tctacaccgc 2520 caatcctaga gtggacgtgg gcgccatcta cagaagatgg atcatcctgg gcctgcagaa 2580 atgcgtgaag atgtacaacc ccgtgtccgt gctggacatc agacagggac ccaaagagcc 2640 cttcaaggac tacgtggacc ggttctataa ggccattaga gccgagcagg ccagcggcga 2700 agtgaagcag tggatgacag agagcctgct gatccagaac gccaatccag actgcaaagt 2760 gatcctgaaa ggcctgggca tgcaccccac actggaagag atgctgacag cctgtcaagg 2820 cgttggcggc ccttcttaca aagccaaagt gatggccgag atgatgcaga ccatgcagaa 2880 ccagaacatg gtgcagcaag gcggccctaa gagagagagg cctcctctga gatgctacaa 2940 ctgcggcaag ttcggccaca tgcagagaca gtgtcctgag cctaggaaaa caaaatgtct 3000 aaaggtgga aaattgggac acctagcaaa agactgcagg ggacaggtga attttttagg 3060 gtatggacgg tggatgggg caaaaccgag aaattttccc gccgctactc ttggagcgga 3120 accgagtgcg cctctcccac cgagcggcac cacccatac gacccagcaa agaagctcct 3180 gcagcaatat gcagagaaag ggaaacaact gagggagcaa aaggaatc caccggcaat 3240 3300 accgtgtaca tcgagggcgt gcccatcaag gctctgctgg atacaggcgc cgacgacacc 3360 atcatcaaag agaacgacct gcagctgagc ggcccttgga ggcctaagat cattggagga 3420 atcggcggag gcctgaacgt caaagagtac aacgaccggg aagtgaagat cgagcaag 3480 atcctgaggg gcacaatcct gctgggcgcc acacctatca acatcatcgg cagaaatctg 3540 ctggcccctg ccggcgctag actggttatg ggacagctct ctgagaagat ccccgtgaca 3600 cccgtgaagc tgaaagaagg cgctagagga ccttgtgtgc gacagtggcc tctgagcaaa 3660 gagaattg aggccctgca agaaatctgt agccagctgg aacaagaggg caagatcagc 3720 agagttggcg gcgagaacgc ctacaatacc cctatcttct gcatcaagaa aaaggacaag 3780 agccagtggc ggatgctggt ggactttaga gagctgaaca aggctaccca ggacttcttc 3840 gaggtgcagc tgggaattcc tcatcctgcc ggcctgcgga agatgagaca gatcacagtg 3900 ctggatgtgg gcgacgccta ctacagcatc cctctggacc ccaacttcag aaagtacacc 3960 gccttcaa tccccaccgt gaaaatcaa ggccctggca tcagatacca gttcaactgc 4020 ctgcctcaag gctgggaaggg cagcccccacc attttcaga ataccgccgc cagcatcctg 4080 gaagaaatca agaaacct gcctgctctg accatcgtgc agtacatgga cgatctgtgg 4140 gtcggaagcc aagagaatga gcacacccac gacaagctgg tggaacagct gagacaaag 4200 ctgcaggcct ggggcctcga aacccctgag aagaaggtgc agaagaacc tccttacgag 4260 tggatgggct acagctgtg gcctcacaag tgggagctga gccggattca gctcgaagg 4320 aaggacgagt ggaccgtgaa cgacatccag aaactcgtgg gcaagctgaa ttggcagcc 4380 cagctgtatc ccggcctgag gaccagaac atctgcaagc tgatccgggg aaagagaac 4440 ctgctggaac tggtcacatg vakacctgag gccgaggccg atatgccga gatgccgaa 4500 atcctgaaaa ccgagcaaga ggggacctac tacaagcctg gcattccaat cagagctgcc 4560 gtgcagaaac tggaggcgg ccagtggtcc taccagtttta agcaagagg ccaggtcctg 4620 aaagtgggca agtacaccaa gcagagaac acccacca acgagctgag vakactggct 4680 ggcctgtcc agaaatctg caaagggcc ctggtcattt ggggcatct gcctgttctg 4740 gaactgccca ttgagcggga agtgtgggaa cagtggtggg ccgattactg gcaagtgtct 4800 tggatccccg agtgggactt cgtctacc cctcctctgc tgaactgtg gtacaccctg 4860 aaaagagc ccattcctaa agaggacgtc tactacgttg acggcgcctg siaccggaac 4920 tccaaagaag gcaaggccgg ctacatcagc cagtacggca agcagagagt ggaaaccctg 4980 gaaaacacca ccaaccagca ggccgagctg accgccatta agatggccct ggaagatagc 5040 ggccccaatg tgaacatcgt gaccgactct cagtacgcca tgggaatcct gacagcccag 5100 cctacacaga gcgatagccc tctggttgag cagatcattg ccctgatgat tcagaagcag 5160 caaatctacc tgcagtgggt gcccgctcac aaaggcatcg gcggaaacga agagatcgat 5220 aagctggtgt ccaagggaat cagacgggtg ctgttcctgg aaaagattga agaggcccaa 5280 gaggaacacg agcgctacca caacaactgg aagaatctgg ccgacaccta cggactgccc 5340 cagatcgtgg ccaaagaaat cgtggctatg tgccccaagt gtcagatcaa gggcgaacct 5400 gtgcacggcc aagtggatgc ttctcctggc acatggcaga tggactgtac ccacctggaa 5460 ggcaaagtgg tcatcgtggc tgtgcacgtg gcctccggct ttattgaggc cgaagtgatc 5520 cccagagaga caggcaaaga aaccgccaag ttcctgctga agatcctgtc cagatggccc 5580 atcacacagc tgcacaccga caacggccct aacttcacat ctcaagaggt ggccgccatc 5640 tgttggtggg gaaagattga gcacacaacc ggcattccct acaatccaca gagccagggc 5700 agcatcgagt ccatgaacaa gcagctcaaa gagattatcg gcaagatccg ggacgactgc 5760 footcacag aaacagccgt gctgatggcc tgtcacatcc aaacttcaa gcggaaaggc 5820 ggcatcggag gacagacatc tgccgagaga ctgatcaata tcatcaccac tcagctgggaa 5880 atccagcacc tccagaccaa gatccagaag attctgaact tccgggtgta ctaccgcgag 5940 ggcagagatc ctgtttggaa aggcccagca cagctgatct ggaaggcga aggtgccgtg 6000 gtgctgaagg atggctctga tctgaaggtg gtgcccagac ggaaggccaa gattatcaag 6060 gattacgagc ccaaacagcg cgtgggcaat gaaggcgacg ttgagggcac aagaggcagc 6120 gacaattgaa attcactcct caggtgcagg ctgcctatca gaaggtggtg gctggtgtgg 6180 ccaatgccct ggctcacaaa taccactgag atctttttcc ctctgccaaa aattatgggg 6240 acatcatgaa gccccttgag catctgactt ctggctaata aaggaaaattt attttcattg 6300 caatagtgtg ttggaatttt ttgtgtctct cactcggaag gacatatggg agggcaaatc 6360 atttaaaca tcagaatgag tatttggttt aggtttggc aacatatgcc catatgctgg ctgccatgaa caaaggttgg ctataaagag gtcatcagta tatgaaacag ccccctgctg tccattcctt attccataga aaagccttga cttgaggtta gatttttttt atttttgtt ttgtgttatt tttttcttta acatccctaa aattttcctt acatgtttta ctagccagat ttttcctcct ctcctgacta ctcccagtca tagctgtccc tcttctctta tggagatccc 6660 tcgacctgca gcccaagctt ggcgtaatca tggtcatagc tgtttcctgt gtgaaattgt 6720. tatccgctca caattccaca caacatacga gccggaagca taaagtgtaa agcctggggt gcctaatgag tgagctaact cacattaatt gcgttgcgct cactgcccgc tttccagtcg ggaaacctgt cgtgccagcg gatccgcatc tcaattagtc agcaaccata gtccccgcccc taactccgcc catccccgcc ctaactccgc ccagttccgc ccccatggct 6960. gactaatttt ttttatttat gcagaggccg aggccgcctc ggcctctgag ctattccaga 7020 agtagtgagg aggctttttt ggaggcctag gcttttgcaa aaagctaact tgtttattgc 7080. agcttataat ggttacaaat aaagcaatag catcacaaat ttcacaaata aagcattttt 7140 ttcactgcat tctagttgtg gtttgtccaa actcatcaat gtatcttatc atgtctgtcc 7200 gcttcctcgc tcactgactc gctgcgctcg gtcgttcggc tgcggcgagc ggtatcagct 7260 cactcaaagg cggtaatacg gttatccaca gaatcagggg ataacgcagg aaagaacatg 7320 tgagcaaaag gccagcaaaa ggccaggaac cgtaaaaagg ccgcgttgct ggcgtttttc 7380 cataggctcc gcccccctga cgagcatcac aaaaatcgac gctcaagtca gaggtggcga 7440 aacccgacag gactataaag ataccaggcg tttccccctg gaagctccct cgtgcgctct 7500 cctgttccga ccctgccgct taccggatac ctgtccgcct ttctcccttc gggaagcgtg 7560 gcgctttctc atagctcacg ctgtaggtat ctcagttcgg tgtaggtcgt tcgctccaag 7620 ctgggctgtg tgcacgaacc ccccgttcag cccgaccgct gcgccttatc cggtaactat 7680 cgtcttgagt ccaacccggt aagacacgac ttatcgccac tggcagcagc cactggtaac 7740 aggattagca gagcgaggta tgtaggcggt gctacagagt tcttgaagtg gtggcctaac 7800 tacggctaca ctagaagaac agtatttggt atctgcgctc tgctgaagcc agttaccttc 7860 ggaaaaagag ttggtagctc ttgatccggc aaacaaacca ccgctggtag cggtggtttt 7920 tttgtttgca agcagcagat tacgcgcaga aaaaaaggat ctcaagaga tcctttgatc 7980 tttctacgg gtctgacgc tcagtggaac gaaaaccc gttaaggat ttggtcatg 8040 agatttaca aaaggatct cacctagatc cttttaattt aaaaatgag ttttaatca 8100 atctaagta tatgagta aacttggtct gagttaga aaactcatc gagcatcaa 8160 tgaaactgca atttattcat atcaggatta tcaataccat attttgaa aagccgtttc 8220 tgtaatgaag gagaaaactc accgaggcag ttccatagga tggcagatc ctggtatcgg 8280 tctgcgattc cgactcgtcc aacatcaata caactatta atttcccctc gtcaaaata 8340 aggttatcaa gtgagaaatc accatgagtg acgactgaat ccggtgagaa tggcaacagc 8400 ttatgcattt ctttccagac ttgttcaca ggccagccat tacgctcgtc atcaaatca 8460 ctcgcatcaa ccaaccgtt attcattcgt gattgcgcct gagcgagacg aaatacgcga 8520 tcgctgttaa aaggacaatt acaaacagga atcgaatgca accggcgcag gaacactgcc 8580 agcgcatcaa caatattttc acctgaatca ggatattctt ctaatacctg gaatgctgtt 8640 tttccgggga tcgcagtggt gagtaaccat gcatcatcag gagtacggat aaaatgcttg 8700 atggtcggaa gaggcataaa ttccgtcagc cagtttagtc tgaccatctc atctgtaaca 8760 tcattggcaa cgctaccttt gccatgtttc agaaacaact ctggcgcatc gggcttccca 8820 tacaatcgat agattgtcgc acctgattgc ccgacattat cgcgagccca tttataccca 8880 tataaatcag catccatgtt ggaatttaat cgcggcctag agcaagacgt ttcccgttga 8940 atatggctca taacacccct tgtattactg tttatgtaag cagacagttt tattgttcat 9000 gatgatatat ttttatcttg tgcaatgtaa catcagagat tttgagacac aacaattggt 9060 cgac 9064

[0339] Plasmid (pDNA2a pGM297) defined in SEQ ID NO: 32, Figure 2C attgattatt gactagttat taatagtaat caattacggg gtcattagtt catagcccat 60 atatggagtt ccgcgttaca taacttacgg taaatggccc gcctggctga ccgcccaacg 120 acccccgccc attgacgtca ataatgacgt atgttcccat agtaacgcca atagggactt 180 tccattgacg tcaatgggtg gagtatttac ggtaaactgc ccacttggca gtacatcaag 240 tgtatcatat gccaagtacg ccccctattg acgtcaatga cggtaaatgg cccgcctggc 300 attatgccca gtacatgacc ttatgggact ttcctacttg gcagtacatc tacgtattag 360 tcatcgctat taccatggtc gaggtgagcc ccacgttctg cttcactctc cccatctccc 420 ccccctcccc acccccaatt ttgtatttat ttatttttta attattttgt gcagcgatgg 480 gggcgggggg gggggggggg cgcgcgccag gcggggcggg gcggggcgag gggcggggcg 540 gggcgaggcg gagaggtgcg gcggcagcca atcagagcgg cgcgctccga aagtttcctt 600 ttatggcgag gcggcggcgg cggcggccct ataaaaagcg aagcgcgcgg cgggcgggag 660 tcgctgcgcg ctgccttcgc cccgtgcccc gctccgccgc cgcctcgcgc cgcccgcccc 720 ggctctgact gaccgcgtta ctcccacagg tgagcgggcg ggacggccct tctcctccgg 780 gctgtaatta gcgcttggtt taatgacggc ttgtttcttt tctgtggctg cgtgaaagcc 840 ttgaggggct ccgggagggc cctttgtgcg gggggagcgg ctcggggggt gcgtgcgtgt 900 gtgtgtgcgt ggggagcgcc gcgtgcggct ccgcgctgcc cggcggctgt gagcgctgcg 960 ggcgcggcgc ggggctttgt gcgctccgca gtgtgcgcga ggggagcgcg gccgggggcg 1020 gtgccccgcg gtgcgggggg ggctgcgagg ggaacaaagg ctgcgtgcgg ggtgtgtgcg 1080 tgggggggtg agcagggggt gtgggcgcgt cggtcgggct gcaacccccc ctgcaccccc 1140 ctccccgagt tgctgagcac ggcccggctt cgggtgcggg gctccgtacg gggcgtggcg 1200 cggggctcgc cgtgccgggc ggggggtggc ggcaggtggg ggtgccgggc ggggcggggc 1260 cgcctcgggc cggggagggc tcgggggagg ggcgcggcgg cccccggagc gccggcggct 1320 gtcgaggcgc ggcgagccgc agccattgcc ttttatggta atcgtgcgag agggcgcagg 1380 gacttccttt gtcccaaatc tgtgcggagc cgaaatctgg gaggcgccgc cgcaccccct 1440 ctagcgggcg cggggcgaag cggtgcggcg ccggcaggaa ggaaatgggc ggggagggcc 1500 ttcgtgcgtc gccgcgccgc cgtccccttc tccctctcca gcctcggggc tgtccgcggg 1560 gggacggctg ccttcgggggg gggcggggtt cggctctgg cgtgtgaccg 1620 gcggctctag agcctctgct aaccatgttc atgccttctt ctttccta cagctcctgg 1680 gcaacgtgct gttattgtg ctgtctcatc attttgcaa agaattgctc gagactagtg 1740 acttggtgag taggcttcga gcctagttag aggactagga gaggccgtag ccgtactac 1800 tctgggcaag tagggcaggc ggtgggtacg caatggggc ggctacctca gcactaata 1860 gagacaatt agaccaattt gagaaaatac gactcgccc gaacgaag aaaagtacc 1920 aaattaaca ttatatatgg gcaggcagg agatggaggcg cttcggccctc catgagaggt 1980 tgttggagac agaggagggg tgtaaagaa tcatagaagt cctctacccc ctagaaccaa 2040 caggatcgga gggcttaaaa agtctgttca atcttgtgtg cgtactatat tgcttgcaca 2100 frequency agt frequency frequency cagtagcac agagquac cactgccatc 2160 tagtgaaaa agaaaaagt gcaacagaga helpctagtgg aaaaaaaaaaaaaaaaaaaagg 2220 gatagcagc gccacctggt ggcagtcaga attttccagc gcaacaaca ggaaatgcct 2280 gggtacatgt acccttgtca ccgcgcacct taaatgcgtg ggtaaaagca gtagaggaga 2340 aaaaatttgg agcagaaata gtacccatgt ttcaagccct atcagaaggc tgcacaccct 2400 atgacattaa tcagatgctt aatgtgctag gagatcatca aggggcatta caaatagtga 2460 aagagatcat taatgaagaa gcagcccagt gggatgtaac acaccacta cccgcaggac 2520 ccctaccagc aggacagctc agggaccctc gcggctcaga tatagcaggg accaccagct 2580 footcaaga acagttagaa tggatctata ctgctaaccc ccgggtagat gtaggtgcca 2640 tctaccggag atggattat ctaggacttc aaaagtgtgt caaaatgtac aacccagtat 2700 cagtcctaga cattaggcag ggacctaaag agcccttcaa ggattatgtg gacagatttt 2760 acaaggcaat tagagcagaa caagcctcag gggaagtgaa acaatggatg acagaatcat 2820 tactcattca aaatgctaat ccagattgta aggtcatcct gaagggccta ggaatgcacc 2880 ccacccttga agaaatgtta acggcttgtc aggggtagg aggcccaagc tacaaagcaa 2940 aagtaatggc agaaatgatg cagaccatgc aaaatcaaaa catggtgcag cagggaggtc 3000 CAAAAAGACA aagaccccca ctaagatgtt atattgtgg aaatttggc catatgcaaa 3060 gandaatgtcc ggaaccaagg aaaaaaaat gtctaagtg tggaaattg ggacacctag 3120 caaagactg caggggacag gtgaattttt taggtatgg acggtggatg ggggcaaac 3180 cgagaaattt tcccgccgct actctggag cggaccgag tgcgctcct cccaccgagcg 3240 gcaccacccc atacgaccca gcaagaagc tcctgcagca attgcagag aaagggaaac 3300 aactgaggga gcaaagagg aatccaccgg caatgaatcc ggattggacc gaggatatt 3360 ctttgaactc cctctttgga gagaccaat aagacagtg tatatagaag gggtccccat 3420 taaggcactg ctagacacag gggcagatga caccatatt aaagaaatg atttacaatt 3480 atcaggtcca tggagaccca aaattatagg gggcatagga ggaggcctta atgtaaaga 3540 ataacgac agggagta aaatagaga taaaattttg aggaggaaa tattgttagg 3600 agcaactccc attatata taggtagaaa ttgctggcc ccggcaggtg cccggttagt 3660 aatgggacaa ttatcagaa aaattcctgt cacacctgtc aaattgaagg aaggggctcg 3720 gggaccctgt gtaagacaat ggccctctc taagagaag attgaagctt tacaggaat 3780 atgttcccaa ttagagcagg aaggaaaaat cagtagagta ggaggagaa atgcataca 3840 taccccaata tttgcataa agagaagga caatcccag tggaggatgc tagtagactt 3900 taggagtta aaggcaa cccagattt ctttgaagtg cattaggga tacccaccc 3960 agcaggatta agaagatga agcagataac agttttagat gtaggagacg cctattattc 4020 cataccattg gatccaatt ttaggaata tactgctttt actattccca cagtgaataa 4080 tcagggaccc gggattaggt atcattcaa ctgtctcccg caagggtgga aaggatctcc 4140 tacaatcttc caaatacag cagcatccat tttggaggag aaaaagaa acttgccagc 4200 actaaccatt gtacaataca tggatgattt atgggtaggt tctcaagaaa atgacacac 4260 ccatgacaaa ttagtagaac agttagaac aaaattacaa gcctggggct tagaacccc 4320 agaaaagaag gtgcaaaag aaccacctta tgagtggatg ggatacaaac ttggcctca 4380 siaatgggaa ctaagcagaa tacaactgga ggaaaagat gatggactg tcaatgacat 4440 ccagaagtta gttgggaac taaattgggc agcacaatg tatccaggtc ttaggacca 4500 gatatatgc aagttatta gaggaaga aaatctgtta gagctagtga cttggacacc 4560 tgaggcagaa gctgaatg cagaaatgc agagattctt aaaacagaac aggaggaac 4620 ctattacaaa ccaggaatac ctattaggggc agcagtacag aaattggaag gaggacagtg 4680 gagttaccaa ttcaacaag aaggacaagt cttgaaagta ggaaataca ccaagcaaaa 4740 gaacaccat acaatgaac ttcgcacatt agctggttta gtgcagaaga tttgcaaga 4800 agctctagtt atttggggga tattaccagt tctgaacctc ccgatagaaa gagaggtag 4860 ggaacaatgg tgggcggatt actggcaggt aagctggatt cccgaatggg attttgtcag 4920 caccccacct ttgctcaac tatggtacac attaacaaa gaacccatac ccaaggagga 4980 cgtttactat gtagatgag catgcacag aaattcaaa gaaggaaag caggacaat 5040 ctcacaatac ggaaaacaga gagtagaaac attagaaac actaccaatc xxaaaaaa 5100 attackaaaatgg ctttggaga cagtggggcct atgtgaca tagtacaga 5160 ctctcaat gcaatgggaa ttttgacagc acaacccaca caagtgatt caccattagt 5220 agagcaatt atagccttaa tgatacaaaa gcaacaata tatttgcagt gggtaccagc 5280 acataagga atgagga atgaggagat agaataatta gtgagtaag gcattagag 5340 agttttattc ttagaaaaaa tagagaagc tchaagag catgaagat atcaataata 5400 ttggaaaaac ctagcagata catatgggct tccacaata gtagcaaag agatagtggc 5460 catgtgtcca aaatgtcaga taaagggaga accagtgcat ggacaagtgg atgcctcacc 5520 5580 tgtagccagt ggattcatag aagcagaagt catacctagg gaacaggaa aagaaacggc 5640 aaagtttcta ttaaaaatac tgagtagatg gcctataca cagttacaca cagacaatgg 5700 gcctaacttt acctcccaag aagtggcagc atatgttgg tggggaaaa ttgaacatac 5760 aacaggtata ccatataacc cccaatctca aggatcaata gaagcatga ashaacaat 5820 aaagagata attgggaaa taagagatga ttgccaatat acagagacag cagtactgat 5880 ggcttgccat attcacaatt ttaaaagaaa gggaggata gggggacaga cttcagcaga gagactaatt aataataata caacacaatt agaaataca catttacaaa ccaaaattca aaaaatttta aattttagg tctactacag agaagggaga gaccctgtgt ggaaaggacc 6060 agcacaatta atctggaag gggaaggagc agtggtcctc aaggacgga gtgacctaaa ggttgtacca agaaggaag ctaaaattat taggattat gaacccaaac aaagagtggg 6240. 6240. taatgagggt gacgtggaag gtaccagggg atctgataac taaatggcag ggaatagtca gatattggat gagacaaaga aatttgaat ggaactatta tatgcatcag ctggcggccg cgaattcact agtgattccc gtttgtgcta gggttcttag gcttcttggg ggctgctgga 6360. actgcaatgg gagcagcggc gacagccctg acggtccagt ctcagcattt gcttgctggg 6420 atactgcagc agcagaaga tctgctggcg gctgtggagg ctcaacagca gatgttgaag ctgaccattt ggggtgttaa aaacctcaat gcccgcgtca cagcccttga gaagtaccta gaggatcagg cacgactaaa ctcctggggg tgcgcatgga aacaagtatg tcataccaca gtggagtggc cctggacaaa tcggactccg gattggcaaa atatgacttg gttggagtgg 6660 gaagacaaa tagctgattt ggaagcaac attacgac attagtgaa ggctgagaa 6720 caaggaaaa agaatctaga tgctcag aagttaacta gttggtcaga ttctggtct 6780 tggttcgatt tctcaaatg gcttaacatt ttaaaatgg gatttttagt atagtagga 6840 Attaggg Window Tacacagta TGGTA TGGG 6900 tatgttcctc tatctccaca gatccatatc caatcgaatt cccgcggccg cattcactc 6960 ctcaggtgca ggctgcctat cagaaggtgg tggctgtgt ggccaatgcc ctggctcaca 7020 ataccactg agatcttt ccctctgcca aaaattatgg ggacatg aagccccttg 7080 agcatctgac ttctgctaa taaggaat ttattcat tgcaatagtg tgttgaatt 7140 ttttgtgtct ctcactcgga aggacatatg ggaggcaaa tcatttaaaa catcagaatg 7200 agtatttggt ttagagttg gcaacatatg cccatatgct ggctgccatg aacaaaggtt 7260 ggcttaaag agtcatcag tatatgaac agccccctgc tgtccattcc ttattccata 7320 gaaaagcctt gacttgaggt tagatttttt ttatattttg ttttgtgtta tttttttctt 7380 taacatccct aaaattttcc ttacatgttt tactagccag attttcctc ctctcctgac 7440 tactcccagt catagctgtc cctcttctct tatggagatc cctcgacctg cagcccaagc 7500 ttggcgtaat catggtcata gctgtttcct gtgtgaaatt gttatccgct cacaattcca 7560 caacacatac gagccggaag cataaagtgt aaagcctggg gtgcctaatg agtgagctaa 7620 ctcacatta ttgcgttgcg ctcactgccc gctttccagt cgggaaacct gtcgtgccag 7680 cggatccgca tctcaattag tcagcaacca tagtcccgcc cctaactccg cccatcccgc 7740 ccctaactcc gcccagttcc gcccattctc cgcccccatgg ctgactaatt ttttttattt 7800 atgcagaggc cgaggccgcc tcggcctctg agctattcca gaagtagtga ggaggctttt 7860 ttggaggcct aggcttttgc aaaaagctaa cttgtttatt gcagcttata atggttacaa 7920 ataaagcaat agcatcacaa atttcacaaa taaagcattt ttttcactgc attctagttg 7980 tggtttgtcc aaactcatca atgtatctta tcatgtctgt ccgcttctc gctcactgac 8040 tcgctgcgct cggtcgttcg gctgcggcga gcggtatcag ctcactcaaa ggcggtaata 8100 cggttatcca cagaatcagg ggataacgca ggaaagaaca tgtgagcaaa aggccagcaa 8160 aaggccagga accgtaaaaa ggccgcgttg ctggcgtttt tccataggct ccgcccccct 8220 gacgagcatc acaaaaatcg acgctcaagt cagaggtggc gaaacccgac aggactataa 8280 agataccagg cgtttccccc tggaagctcc ctcgtgcgct ctcctgttcc gaccctgccg 8340 cttaccggat acctgtccgc ctttctccct tcgggaagcg tggcgctttc tcatagctca 8400 cgctgtaggt atctcagttc ggtgtaggtc gttcgctcca agctgggctg tgtgcacgaa 8460 ccccccgttc agcccgaccg ctgcgcctta tccggtaact atcgtcttga gtccaacccg 8520 gtaagacacg acttatcgcc actggcagca gccactggta acaggattag cagagcgagg 8580 tatgtaggcg gtgctacaga gttcttgaag tggtggccta actacggcta cactagaaga 8640 acagtatttg gtatctgcgc tctgctgaag ccagttacct tcggaaaaag agttggtagc 8700 tcttgatccg gcaaacaaac caccgctggt agcggtggtt tttttgtttg caagcagcag 8760 attack gaaaaaaagg attack gatcctttga tcttttctac ggggtctgac gctcagtgga acgaaaactc acgttaaggg attttggtca tgagattatc aaaaaggatc ttcacctaga tccttttaaa ttaaaatga agttttaaat caatctaaag tatatatgag 9000. 9000. 9000. 9000. 9000. 9000. 9000. 9000. 9000. 9000 atatcaggat tatcaatacc atatttttga aaaagccgtt tctgtaatga aggagaaac tcaccgaggc agttccatag gatggcaaga tcctggtatc ggtctgcgat tccgactcgt 9180. snowflake 9180. snowflake 9180. snowflake 9240. tcaccatgag tgacgactga atccggtgag aatggcaaca gcttatgcat ttctttccag acttgttcaa caggccagcc attackcgctcg tcatcaaaat cactcgcatc aaccaaaccg ttattcattc gtgattgcgc ctgagcgaga cgaaatacgc gatcgctgtt aaaaggacaa ttacaaacag gaatcgaatg caaccggcgc aggaacactg ccagcgcatc aacaatattt tcacctgaat caggatattc ttctaatacc tggaatgctg tttttccggg gatcgcagtg gtgagtaacc atgcatcatc aggagtacgg ataaaatgct tgatggtcgg aagaggcata 9540 aattccgtca gccagtttag tctgaccatc tcatctgtaa catcattggc aacgctacct 9600 ttgccatgtt tcagaaacaa ctctggcgca tcgggcttcc catacaatcg atagattgtc 9660 gcacctgatt gcccgacatt atcgcgagcc catttatacc catataaatc agcatccatg 9720 ttggaattta atcgcggcct agagcaagac gtttcccgtt gaatatggct cataacaccc 9780 cttgtattac tgtttatgta agcagacagt tttattgttc atgatgatat atttttatct 9840 tgtgcaatgt aacatcagag attttgagac acaacaattg gtcgac 9886

[0340] Plasmid (pDNA2b pGM299) defined in FIG. 2D, SEQ ID NO: 33 tcaatattgg ccattagcca tattattcat tggttatata gcataaatca atattggcta 60 ttggccattg catacgttgt atctatatca taatatgtac atttatattg gctcatgtcc 120 aatatgaccg ccatgttggc attgattatt gactagttat taatagtaat caattacggg 180 gtcattagtt catagcccat atatggagtt ccgcgttaca taacttacgg taaatggccc 240 gcctggctga ccgcccaacg acccccgccc attgacgtca ataatgacgt atgttcccat 300 agtaacgcca atagggactt tccattgacg tcaatgggtg gagtatttac ggtaaactgc 360 ccacttggca gtacatcaag tgtatcatat gccaagtccg ccccctattg acgtcaatga 420 cggtaaatgg cccgcctggc attatgccca gtacatgacc ttacgggact ttcctacttg 480 gcagtacatc tacgtattag tcatcgctat taccatggtg atgcggtttt ggcagtacac 540 caatgggcgt ggatagcggt ttgactcacg gggatttcca agtctccacc ccattgacgt 600 caatgggagt ttgttttggc accaaaatca acgggacttt ccaaaatgtc gtaataaccc 660 cgccccgttg acgcaaatgg gcggtaggcg tgtacggtgg gaggtctata taagcagagc 720 tcgtttagtg aaccgtcaga tcactagaag ctttattgcg gtagtttatc acagttaaat 780 tgctaacgca gtcagtgctt ctgacacaac agtctcgaac ttaagctgca gaagttggtc 840 gtgaggcact gggcaggtaa gtatcaaggt tacaagacag gtttaaggag accaatagaa 900 actgggcttg tcgagacaga gaagactctt gcgtttctga taggcaccta ttggtcttac 960 tgacatccac ttgccttc tctccacagg tgtccactcc cagttcatt acagctctta 1020 aggtagagt acttaatacg actcactata ggctagccctc gagaattcga ttagcccct 1080 aggaccagaa gaaagaagat tgcttcgctt gatttggctc ctttacagca ccaatccata 1140 tccaccaagt ggggaaggga cggccagaca acgccgacga gccaggagaa ggtggagaca 1200 acagcaggat caattagag tcttggtaga aagactccaa gagcaggtgt atgcagttga 1260 ccgcctggct gacgaggctc aacacttggc tatacacag tgcctgacc ctcctcattc 1320 agcttagaat cactagtgaa ttcaccgtg gtaccttag agtcgaccg ggcggccgct 1380 tcgagcagac atgatagat acattgatga gtttggacaa accacacta gatgcagtg 1440 aaaaaaatgc tttattgtg aaatttgtga tgctattgct ttatttgtaa ccattataag 1500 ctgcaataaa caagttaaca acaacattg cattcatttt atgttcagg ttcaggggga 1560 gatgtgggag gttttttaaa gcaagtaaaa cctctacaa tgtggtaaaa tcgataagga 1620 tccgtcgacc aattgttgtg tctcaaatc tctgatgtta cattgcaca gataaaaata 1680 tatcatcatg aacaataaaa ctgtctgctt acataaacag taatacaagg ggtgttatga 1740 gccatattca acgggaaacg tcttgctcta ggccgcgatt aaattccaac atggatgctg 1800 atttatatgg gtataaatgg gctcgcgata atgtcgggca atcaggtgcg acaatctatc 1860 gattgtatgg gaagcccgat gcgccagagt tgtttctgaa acatggcaaa ggtagcgttg 1920 ccaatgatgt tacagatgag atggtcagac taaactggct gacggaattt atgcctcttc 1980 cgaccatcaa gcattttatc cgtactcctg atgatgcatg gttactcacc actgcgatcc 2040 ccggaaaaac agcattccag gtattagaag aatatcctga ttcaggtgaa aatattgttg 2100 atgcgctggc agtgttcctg cgccggttgc attcgattcc tgtttgtaat tgtcctttta 2160 acagcgatcg cgtatttcgt ctcgctcagg cgcaatcacg aatgaataac ggtttggttg 2220 atgcgagtga ttttgatgac gagcgtaatg gctggcctgt tgaacaagtc tggaaagaaa 2280 tgcataagct gttgccattc tcaccggatt cagtcgtcac tcatggtgat ttctcacttg 2340 ataaccttat ttttgacgag gggaaattaa taggttgtat tgatgttgga cgagtcggaa 2400 tcgcagaccg ataccaggat cttgccatcc tatggaactg cctcggtgag ttttctcctt 2460 cattacagaa acggcttttt caaaaatatg gtattgataa tcctgatatg aataaattgc 2520 agtttcattt gatgctcgat gagtttttct aactgtcaga ccaagtttac tcatatatac 2580 tttagattga tttaaaactt catttttaat ttaaaaggat ctaggtgaag atcctttttg 2640 ataatctcat gaccaaaatc ccttaacgtg agttttcgtt ccactgagcg tcagaccccg 2700 tagaaaagat caaaggatct tcttgagatc ctttttttct gcgcgtaatc tgctgcttgc 2760 aaacaaaaaa accaccgcta ccagcggtgg tttgtttgcc ggatcaagag ctaccaactc 2820 tttttccgaa ggtaactggc ttcagcagag cgcagatacc aaatactgtt cttctagtgt 2880 agccgtagtt aggccaccac ttcaagaact ctgtagcacc gcctacatac ctcgctctgc 2940 taatcctgtt accagtggct gctgccagtg gcgataagtc gtgtcttacc gggttggact 3000 caagacgata gttaccggat aaggcgcagc ggtcgggctg aacggggggt tcgtgcacac 3060 agcccagctt ggagcgaacg acctacaccg aactgagata cctacagcgt gagctatgag 3120 aaagcgccac gcttcccgaa gggagaaagg cggacaggta tccggtaagc ggcagggtcg 3180 gaacaggaga gcgcacgagg gagcttccag ggggaaacgc ctggtatctt tatagtcctg 3240 tcgggtttcg ccacctctga cttgagcgtc gatttttgtg atgctcgtca ggggggcgga 3300 gcctatggaa aaacgccagc aacgcggcct ttttacggtt cctggccttt tgctggcctt 3360 ttgctcacat ggctcgacag atct 3384

[0341] Plasmid (pDNA3a pGM301) defined in Figure 2E of SEQ ID NO: 34 attgattatt gactagttat taatagtaat caattacggg gtcattagtt catagcccat 60 atatggagtt ccgcgttaca taacttacgg taaatggccc gcctggctga ccgcccaacg 120 acccccgccc attgacgtca ataatgacgt atgttcccat agtaacgcca atagggactt 180 tccattgacg tcaatgggtg gagtatttac ggtaaactgc ccacttggca gtacatcaag 240 tgtatcatat gccaagtacg ccccctattg acgtcaatga cggtaaatgg cccgcctggc 300 attatgccca gtacatgacc ttatgggact ttcctacttg gcagtacatc tacgtattag 360 tcatcgctat taccatggtc gaggtgagcc ccacgttctg cttcactctc cccatctccc 420 ccccctcccc acccccaatt ttgtatttat ttatttttta attattttgt gcagcgatgg 480 gggcgggggg gggggggggg cgcgcgccag gcggggcggg gcggggcgag gggcggggcg 540 gggcgaggcg gagaggtgcg gcggcagcca atcagagcgg cgcgctccga aagtttcctt 600 ttatggcgag gcggcggcgg cggcggccct ataaaaagcg aagcgcgcgg cgggcgggag 660 tcgctgcgcg ctgccttcgc cccgtgcccc gctccgccgc cgcctcgcgc cgcccgcccc 720 ggctctgact gaccgcgtta ctcccacagg tgagcgggcg ggacggccct tctcctccgg 780 gctgtaatta gcgcttggtt taatgacggc ttgtttcttt tctgtggctg cgtgaaagcc 840 ttgaggggct ccgggagggc cctttgtgcg gggggagcgg ctcggggggt gcgtgcgtgt 900 gtgtgtgcgt ggggagcgcc gcgtgcggct ccgcgctgcc cggcggctgt gagcgctgcg 960 ggcgcggcgc ggggctttgt gcgctccgca gtgtgcgcga ggggagcgcg gccgggggcg 1020 gtgccccgcg gtgcgggggg ggctgcgagg ggaacaaagg ctgcgtgcgg ggtgtgtgcg 1080 tgggggggtg agcagggggt gtgggcgcgt cggtcgggct gcaacccccc ctgcaccccc 1140 ctccccgagt tgctgagcac ggcccggctt cgggtgcggg gctccgtacg gggcgtggcg 1200 cggggctcgc cgtgccgggc ggggggtggc ggcaggtggg ggtgccgggc ggggcggggc 1260 cgcctcgggc cggggagggc tcgggggagg ggcgcggcgg cccccggagc gccggcggct 1320 gtcgaggcgc ggcgagccgc agccattgcc ttttatggta atcgtgcgag agggcgcagg 1380 gacttccttt gtcccaaatc tgtgcggagc cgaaatctgg gaggcgccgc cgcaccccct 1440 ctagcgggcg cggggcgaag cggtgcggcg ccggcaggaa ggaaatgggc ggggagggcc 1500 ttcgtgcgtc gccgcgccgc cgtccccttc tccctctcca gcctcggggc tgtccgcggg 1560 gggacggctg ccttcggggg ggacggggca gggcggggtt cggcttctgg cgtgtgaccg 1620 gcggctctag agcctctgct aaccatgttc atgccttctt ctttttccta cagctcctgg 1680 gcaacgtgct ggttattgtg ctgtctcatc attttggcaa agaattcgat tgccatggca 1740 acatatatcc agagagtaca gtgcatctca acatcactac tggttgttct caccacattg 1800 gtctcgtgtc agatcccag ggataggctc tctaacatag gggtcatagt cgatgaaggg 1860 aaatcactga agatagctgg atcccacgaa tcgaggtaca tagtactgag tctagttccg 1920 ggggtagact ttgagaatgg gtgcggaaca gcccaggtta tccagtaca gagcctactg 1980 aacaggctgt taatcccatt gagggatgcc ttagatctc aggaggctct gataactgtc 2040 accaatgata cgacacaaaa tgccggtgct ccccagtcga gattctcgg tgctgtgatt 2100 ggtactatcg cacttggagt gcgacatca gcacaaatca ccgcagggat tgcactagcc 2160 gaagcgagggg aggcaaag agacatagcg ctcatcaag atcgatgac aaaaacacac 2220 aagtctatag aactgctgca aaacgctgtg ggggaacaa ttctgctct aagacactc 2280 caggatttcg tgaatgatga gatcaaccc gcataagcg aattaggctg tgagactgct 2340 gccttaagac tgggtataaa attgacac cattactccg agctgttaac tgcgttcggc 2400 tcgaatttcg gaaccatcgg agagagagc ctcacgctgc agggctgtc ttcactttac 2460 tctgctaaca ttactgagat tatgaccaca atcaggacag ggcagtctaa catctatgat 2520 gtcatttata cagaacagat caaaggaacg gtgatagatg tggatctaga gagatacatg 2580 gtcaccctgt ctgtgaagat ccctattctt tctgaagtcc caggtgtgct catacacaag 2640 gcatcatcta tttcttacaa catagacggg gaggaatggt atgtgactgt ccccagccat 2700 atactcagtc gtgcttcttt cttagggggt gcagacataa ccgattgtgt tgagtccaga 2760 ttgacctata tatgccccag ggatcccgca caactgatac ctgacagcca gcaaaagtgt 2820 atcctggggg acacaacaag gtgtcctgtc acaaaagttg tggacagcct tatccccaag 2880 tttgctttg tgaatggggg cgttgttgct aactgcatag catccacatg tacctgcggg 2940 acaggccgaa gaccaatcag tcaggatcgc tctaaaggtg tagtattcct aacccatgac 3000 aactgtggtc ttataggtgt caatggggta gaattgtatg ctaaccggag agggcacgat 3060 gccacttggg gggtccagaa cttgacagtc ggtcctgcaa ttgctatcag acccgttgat 3120 atttctctca accttgctga tgctacgaat ttcttgcaag actctaaggc tgagcttgag 3180 aaagcacgga aaatcctctc ggaggtaggt agatggtaca actcaagaga gactgtgatt 3240 acgatcatag tagttatggt cgtaatattg gtggtcatta tagtgatcat catcgtgctt 3300 tagatactca gaaggtgaaa tcactagtga attcactcct caggtgcagg ctgcctatca 3360 gaaggtggtg gctggtgtgg ccaatgccct ggctcacaaa taccactgag atctttttcc 3420 ctctgccaaa aattatgggg acatcatgaa gccccttgag catctgactt ctggctaata 3480 aggaaattt attttcattg caatagtgtg ttggaatttt ttgtgtctct cactcggaag 3540 gacatatggg agggcaaatc atttaaaaca tcagaatgag tatttggttt agagtttggc 3600 aacatatgcc catatgctgg ctgccatgaa caaaggttgg ctataaagag gtcatcagta 3660 tatgaaacag ccccctgctg tccattcctt attccataga aaagccttga cttgaggtta 3720 gattttttt atattttgtt ttgtgttatt tttttcttta acatccctaa aattttcctt 3780 acatgtttta ctagccagat ttttcctcct ctcctgacta ctcccagtca tagctgtccc 3840 tcttctctta tggagatccc tcgacctgca gcccaagctt ggcgtaatca tggtcatagc 3900 tgtttcctgt gtgaaattgt tatccgctca caattccaca caacatacga gccggaagca 3960 taaagtgtaa agcctggggt gcctaatgag tgagctaact cacattaatt gcgttgcgct 4080. cactcccgc tttccagtcg ggaaacctgt cgtgccagcg gatccgcatc tcaattagtc agcaaccata gtcccgcccc taactccgcc catcccgccc ctaactccgc ccagttccgc ccattctccg ccccatggct gactaatttt ttttatttt gcagaggccg aggccgcctc 4200. ggcctctgag ctattccaga agtagtgagg aggcttttttt ggaggcctag gcttttgcaa 4260 aaagctaact tgtttattgc agcttataat ggttacaaat aaagcaatag catcacaaat ttcacaaata aagcattttt ttcactgcat tctagttgtg gtttgtccaa actcatcaat gtatcttatc atgtctgtcc gcttcctcgc tcactgactc gctgcgctcg gtcgttcggc 4440 tgcggcgagc ggtatcagct cactcaaagg cggtaatacg gttatccaca gaatcagggg ataacgcagg aaagaacatg tgagcaaaag gccagcaaaa ggccaggac cgtaaaaagg ccgcgttgct ggcgtttttc cataggctcc gcccccctga cgagcatcac aaaaatcgac 4620 gctcaagtca gaggtggcga aacccgacag gactaaag ataccaggcg tttccccctg gaagctccct cgtgcgctct cctgttccga cctgccgct taccggatac ctgtccgcct 4740 ttctcccttc gggaagcgtg gcgctttctc atagctcacg ctgtaggtat ctcagttcgg 4800 tgtagtcgt tcgctccaag ctgggctgtg tgcacgaacc ccccgttcag cccgaccgct 4860 gcgccttatc cggtaactat cgtcttgagt ccaacccggt aagacacgac ttatcgccac 4920 tggcagcagc cactggtaac aggattagca gagcgaggta tgtaggcggt gctacagagt 4980 tcttgaagtg gtggcctaac tacggctaca ctagaagaac agtatttggt atctgcgctc 5040 tgctgaagcc agttaccttc ggaaaaagag ttggtagctc ttgatccggc aaacaaacca 5100 ccgctggtag cggtggtttt tttgtttgca agcagcagat tacgcgcaga aaaaaaggat 5160 ctcaagaaga tcctttgatc ttttctacgg ggtctgacgc tcagtggaac gaaaactcac 5220 gttaagggat tttggtcatg agattatcaa aaaggatctt cacctagatc cttttaaatt 5280 aaaaatgaag ttttaaatca atctaaagta tatatgagta aacttggtct gacagttaga 5340 aaaactcatc gagcatcaaa tgaaactgca atttattcat atcaggatta tcaataccat 5400 atttttgaaa aagccgtttc tgtaatgaag gagaaaactc accgaggcag ttccatagga 5460 tggcaagatc ctggtatcgg tctgcgattc cgactcgtcc aacatcaata caacctatta 5520 atttcccctc gtcaaaaata aggttatcaa gtgagaaatc accatgagtg acgactgaat 5580 ccggtgagaa tggcaacagc ttatgcattt ctttccagac ttgttcaaca ggccagccat 5640 tacgctcgtc atcaaaatca ctcgcatcaa ccaaaccgtt attcattcgt gattgcgcct 5700 gagcgagacg aaatacgcga tcgctgttaa aaggacaatt acaaacagga atcgaatgca 5760 accggcgcag gaacactgcc agcgcatcaa caatattttc acctgaatca ggatattctt 5820 ctaatacctg gaatgctgtt tttccgggga tcgcagtggt gagtaaccat gcatcatcag 5880 gagtacggat aaaatgcttg atggtcggaa gaggcataaa ttccgtcagc cagtttagtc 5940 tgaccatctc atctgtaaca tcattggcaa cgctaccttt gccatgtttc agaaacaact 6000 ctggcgcatc gggcttccca tacaatcgat agattgtcgc acctgattgc ccgacattat 6060 cgcgagccca tttataccca tataaatcag catccatgtt ggaatttaat cgcggcctag 6120 agcaagacgt ttcccgttga atatggctca taacacccct tgtattactg tttatgtaag 6180 cagacagttt tattgttcat gatgatatat ttttatcttg tgcaatgtaa catcagagat 6240 tttgagacac aacaattggt cgac 6264

[0342] Plasmid (pDNA3b pGM303) defined in Figure 2F of SEQ ID NO: 35 attgattatt gactagttat taatagtaat caattacggg gtcattagtt catagcccat 60 atatggagtt ccgcgttaca taacttacgg taaatggccc gcctggctga ccgcccaacg 120 acccccgccc attgacgtca ataatgacgt atgttcccat agtaacgcca atagggactt 180 tccattgacg tcaatgggtg gagtatttac ggtaaactgc ccacttggca gtacatcaag 240 tgtatcatat gccaagtacg ccccctattg acgtcaatga cggtaaatgg cccgcctggc 300 attatgccca gtacatgacc ttatgggact ttcctacttg gcagtacatc tacgtattag 360 tcatcgctat taccatggtc gaggtgagcc ccacgttctg cttcactctc cccatctccc 420 ccccctcccc acccccaatt ttgtatttat ttatttttta attattttgt gcagcgatgg 480 gggcgggggg gggggggggg cgcgcgccag gcggggcggg gcggggcgag gggcggggcg 540 gggcgaggcg gagaggtgcg gcggcagcca atcagagcgg cgcgctccga aagtttcctt 600 ttatggcgag gcggcggcgg cggcggccct ataaaaagcg aagcgcgcgg cgggcgggag 660 tcgctgcgcg ctgccttcgc cccgtgcccc gctccgccgc cgcctcgcgc cgcccgcccc 720 ggctctgact gaccgcgtta ctcccacagg tgagcgggcg ggacggccct tctcctccgg 780 gctgtaatta gcgcttggtt taatgacggc ttgtttcttt tctgtggctg cgtgaaagcc 840 ttgaggggct ccgggagggc cctttgtgcg gggggagcgg ctcggggggt gcgtgcgtgt 900 gtgtgtgcgt ggggagcgcc gcgtgcggct ccgcgctgcc cggcggctgt gagcgctgcg 960 ggcgcggcgc ggggctttgt gcgctccgca gtgtgcgcga ggggagcgcg gccgggggcg 1020 gtgccccgcg gtgcgggggg ggctgcgagg ggaacaaagg ctgcgtgcgg ggtgtgtgcg 1080 tgggggggtg agcagggggt gtgggcgcgt cggtcgggct gcaacccccc ctgcaccccc 1140 ctccccgagt tgctgagcac ggcccggctt cgggtgcggg gctccgtacg gggcgtggcg 1200 cggggctcgc cgtgccgggc ggggggtggc ggcaggtggg ggtgccgggc ggggcggggc 1260 cgcctcgggc cggggagggc tcgggggagg ggcgcggcgg cccccggagc gccggcggct 1320 gtcgaggcgc ggcgagccgc agccattgcc tttatggta atcgtgcgag agggcgcagg 1380 gacttcctt gtcccaaatc tgtgcggagc cgaaatctgg gaggcgccgc cgcaccccct 1440 ctagcgggcg cggggcgaag cggtgcggcg ccggcaggaa ggaaatgggc ggggagggcc 1500 ttcgtgcgtc gccgcgccgc cgtccccttc tccctctcca gcctcggggc tgtccgcggg 1560 gggacggggc agggcggggt tcggcttctg gcgtgtgacc ggcggctcta gagcctctgc 1620 taaccatgtt catgccttct tctttttcct acagctcctg ggcaacgtgc tggttattgt 1680 gctgtctcat cattttggca aagaattcct cgagcatgtg gtctgagtta aaaatcagga 1740 gcaacgacgg aggtgaagga ccagaggacg ccaacgaccc ccggggaaag ggggtgcaac 1800 acatccatat ccagccatct ctacctgttt atggacagag ggttagggat ggtgataggg 1860 gcaaacgtga ctcgtactgg tctacttctc ctagtggtag caccacaaaa ccagcatcag 1920 gttgggagag gtcaagtaaa gccgacacat ggttgctgat tctctcattc acccagtggg 1980 ctttgtcaat tgccacagtg atcatctgta tcataatttc tgctagacaa gggtatagta 2040 tgaaagagta ctcaatgact gtagaggcat tgaacatgag cagcagggag gtgaaagagt 2100 cacttaccag tctaataagg caagaggtta tagcaagggc tgtcaacatt cagagctctg 2160 tgcaaaccgg aatcccagtc ttgttgaaca aaaacagcag ggatgtcatc cagatgattg 2220 ataagtcgtg cagcagacaa gagctcactc agcactgtga gagtacgatc gcagtccacc 2280 atgccgatgg aattgcccca cttgagccac atagtttctg gagatgccct gtcggagaac 2340 cgtatcttag ctcagatcct gaaatctcat tgctgcctgg tccgagcttg ttatctggtt 2400 ctacaacgat ctctggatgt gttaggctcc cttcactctc aattggcgag gcaatctatg 2460 cctattcatc aaatctcatt acacaaggtt gtgctgacat agggaaatca tatcaggtcc 2520 tgcagctagg gtacatatca ctcaattcag atatgttccc tgatcttaac cccgtagtgt 2580 cccacactta tgacatcaac gacaatcgga aatcatgctc tgtggtggca accgggacta 2640 ggggtatca gctttgctcc atgccgactg tagacgaaag aaccgactac tctagtgatg 2700 gtattgagga tctggtcctt gatgtcctgg atctcaaagg gagaactaag tctcaccggt 2760 atcgcaacag cgaggtagat cttgatcacc cgttctctgc actatacccc agtgtaggca 2820 acggcattgc aacagaaggc tcattgatat ttcttgggta tggtggacta accacccctc 2880 tgcagggtga tacaaaatgt aggacccaag gatgccaaca ggtgtcgcaa gacacatgca 2940 atgaggctct gaaaattaca tggctaggag ggaaacaggt ggtcagcgtg atcatccagg 3000 tcaatgacta tctctcagag aggccaaaga taagagtcac aaccattcca atcactcaaa 3060 actatctcgg ggcggaaggt agattattaa aattgggtga tcgggtgtac atctatacaa 3120 gatcatcagg ctggcactct caactgcaga taggagtact tgatgtcagc caccctttga 3180 ctatcaactg gacacctcat gaagccttgt ctagaccagg aaataaagag tgcaattggt 3240 acaataagtg tccgaaggaa tgcatatcag gcgtatacac tgatgcttat ccattgccc 3300 ctgatgcagc taacgtcgct accgtcacgc tatatgccaa tacatcgcgt gtcaacccaa 3360 caatcatgta ttctaacact actaacatta taaatatgtt aaggataaag gatgttcaat 3420 tagaggctgc atataccacg acatcgtgta tcacgcattt tggtaaaggc tactgctttc 3480 acatcatcga gatcaatcag aagagcctga ataccttaca gccgatgctc tttaagacta 3540 gcatccctaa attatgcaag gccgagtctt aagcggccgc gcatgcgaat tcactcctca 3600 ggtgcaggct gcctatcaga aggtggtggc tggtgtggcc aatgccctgg ctcacaaata 3660 ccactgagat ctttttccct ctgccaaaaa ttatggggac atcatgaagc cccttgagca 3720 tctgacttct ggctaataaa ggaaatttat tttcattgca atagtgtgtt ggaatttttt 3780 gtgtctctca ctcggaagga catatgggag ggcaaatcat ttaaaacatc agaatgagta 3840 tttggtttag agtttggcaa catatgccca tatgctggct gccatgaaca aaggttggct 3900 ataaagaggt catcagtata tgaaacagcc ccctgctgtc tattccttat tccatagaaa 3960 agccttgact tgaggttaga ttttttttat attttgtttt gtgttatttt tttctttaac 4020 atccctaaaa ttttccttac atgttttact agccagattt ttcctcctct cctgactact 4080 cccagtcata gctgtccctc ttctcttatg gagatccctc gacctgcagc ccaagcttgg 4140 cgtaatcatg gtcatagctg tttcctgtgt gaaattgtta tccgctcaca attccacaca 4200 acatacgagc cggaagcata aagtgtaaag cctggggtgc ctaatgagtg agctaactca 4260 cattaattgc gttgcgctca ctgcccgctt tccagtcggg aaacctgtcg tgccagcgga 4320 tccgcatctc aattagtcag caaccatagt cccgccccta actccgccca tcccgcccct 4380 aactccgccc agttccgccc attctccgcc ccatggctga ctaatttttt ttatttatgc 4440 agaggccgag gccgcctcgg cctctgagct attccagaag tagtgaggag gcttttttgg 4500 aggcctaggc ttttgcaaaa agctaacttg tttattgcag cttataatgg ttacaaataa 4560 agcaatagca tcacaaattt cacaaataaa gcattttttt cactgcattc tagttgtggt 4620 ttgtccaaac tcatcaatgt atcttatcat gtctgtccgc ttcctcgctc actgactcgc 4680 tgcgctcggt cgttcggctg cggcgagcgg tatcagctca ctcaaaggcg gtaatacggt 4740 tatccacaga atcaggggat aacgcaggaa agaacatgtg agcaaaaggc cagcaaaagg 4800 ccaggaaccg taaaaaggcc gcgttgctgg cgtttttcca taggctccgc ccccctgacg 4860 agcatcacaa aaatcgacgc tcaagtcaga ggtggcgaaa cccgacagga ctataaagat 4920 accaggcgtt tccccctgga agctccctcg tgcgctctcc tgttccgacc ctgccgctta 4980 ccggatacct gtccgccttt ctcccttcgg gaagcgtggc gctttctcat agctcacgct 5040 gtaggtatct cagttcggtg tagtcgttc gctccaagct gggctgtgtg cacgaacccc 5100 ccgttcagcc cgaccgctgc gccttatccg gtaactatcg tcttgagtcc aacccggtaa 5160 gacacgactt atcgccactg gcagcagcca ctggtaacag gattagcaga gcgaggtatg 5220 taggcggtgc tacagagttc ttgaagtggt ggcctaacta cggctacact agaagaacag 5280 tatttggtat ctgcgctctg ctgaagccag ttaccttcgg aaaaagagtt ggtagctctt 5340 gatccggcaa acaaaccacc gctggtagcg gtggtttttt tgtttgcaag cagcagatta 5400 cgcgcagaaa aaaaggatct caagaagatc ctttgatcttt ttctacgggg tctgacgctc 5460 agtggaacga aaactcacgt taagggattt tggtcatgag attatcaaaa aggatcttca 5520 cctagatcct tttaaattaa aaatgaagtt ttaaatcaat ctaaagtata tatgagtaaa 5580 cttggtctga cagttagaa...

Claims

1. A retroviral vector containing an intron; (a) the endogenous Rev response element (RRE) of the retroviral genome is deleted; (b) The retroviral vector, wherein the retroviral RRE is inserted into the intron within 100 bp 5' of the branch site of the splice acceptor of the intron.

2. The retroviral vector of claim 1 , wherein the retroviral RRE is inserted within 20 bp 5′ of the branch site of the splice acceptor of the intron.

3. (a) the intron is optionally a chimeric intron selected from a β-globin / IgG chimeric intron or a chimeric intron from a CAGGS promoter, or (b) the intron is optionally a viral intron selected from an SV40 intron, a CMV intron A, and an adenovirus tripartite leader sequence intron; The retroviral vector according to claim 1 or 2.

4. A retroviral vector comprising a chimeric intron; (a) the endogenous RRE of the retroviral genome is deleted; (b) the retroviral vector, wherein a retroviral RRE is inserted into the chimeric intron.

5. The retroviral vector according to any one of claims 1 to 4, wherein the retroviral RRE inserted into the intron is the endogenous RRE of the retroviral genome.

6. The retroviral vector according to any one of claims 1 to 5, wherein the RRE is a simian immunodeficiency virus (SIV) RRE.

7. 7. The retroviral vector of claim 6, wherein the RRE comprises or consists of a nucleic acid sequence having at least 90% identity to SEQ ID NO:

1.

8. A retroviral vector according to any one of claims 1 to 7, wherein the intron is less than 1000 bp in length, preferably less than 800 bp in length.

9. The retroviral vector according to any one of claims 3 to 8, wherein the chimeric intron is a β-globin / IgG chimeric intron or a chimeric intron from a CAGGS promoter.

10. 10. The retroviral vector of claim 9, wherein the chimeric intron is a β-globin / IgG chimeric intron, and the RRE is inserted between (i) a splice donor site comprising or consisting of the nucleic acid sequence of TGAGTTTAAGGTAAGT (SEQ ID NO: 2) and (ii) a splice acceptor site comprising or consisting of the nucleic acid sequence of CTCTCCACAG (SEQ ID NO: 3).

11. The retroviral vector of claim 9 or 10, wherein the β-globin / IgG chimeric intron comprises or consists of a nucleic acid sequence having at least 90% identity to SEQ ID NO:

4.

12. 12. The retroviral vector of any one of claims 1 to 11, wherein the intron is a β-globin / IgG chimeric intron, the RRE is an SIV RRE, and optionally the chimeric intron including the RRE comprises or consists of a nucleic acid sequence having at least 90% identity to SEQ ID NO:

5.

13. 13. The retroviral vector of any one of claims 1 to 12, wherein the intron is between a promoter and a transgene operably linked to the promoter, and optionally the promoter is selected from the group consisting of a cytomegalovirus (CMV) promoter, an elongation factor 1a (EF1a) promoter, and a hybrid human CMV enhancer / EF1a (hCEF) promoter, preferably an hCEF promoter.

14. The transgene encodes a therapeutic protein, optionally the therapeutic protein comprising: (a) monoclonal antibodies against secreted therapeutic proteins, optionally alpha-1 antitrypsin (AAT), factor VIII, surfactant protein B (SFTPB), ADAMTS13, factor VII, factor IX, factor X, factor XI, von Willebrand factor, granulocyte-macrophage colony-stimulating factor (GM-CSF), decorin, surfactant protein C (SP-C), anti-inflammatory proteins, and infectious agents; or (b) CFTR, ABCA3, DNAH5, DNAH11, DNAI1, DNAI2, CSF2RA, CSF2RB, and TRIM-72 The retroviral vector of claim 13, wherein the retroviral vector is selected from the group consisting of:

15. The retroviral vector according to any one of claims 1 to 14, wherein the retroviral vector is a lentiviral vector.

16. 16. The retroviral vector of claim 15, wherein the lentiviral vector is selected from the group consisting of a human immunodeficiency virus (HIV) vector, a simian immunodeficiency virus (SIV) vector, a feline immunodeficiency virus (FIV) vector, an equine infectious anemia virus (EIAV) vector, and a Visna / Maedi virus vector.

17. 17. The retroviral vector of any one of claims 1 to 16, which is pseudotyped with hemagglutinin-neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus or G glycoprotein from vesicular stomatitis virus (G-VSV).

18. 18. The retroviral vector of any one of claims 13 to 17, which increases transgene expression by at least about 2-fold, preferably at least about 5-fold, and more preferably at least about 10-fold, compared to a corresponding vector lacking the retroviral RRE inserted intron.

19. 13. A nucleic acid comprising or consisting of an intron into which a retroviral RRE has been inserted, optionally wherein (i) the intron; and / or (ii) the RRE is as defined in any one of claims 1 to 12.

20. A plasmid comprising the nucleic acid of claim 19.

21. A retroviral vector according to any one of claims 1 to 18, a nucleic acid according to claim 19, or a plasmid according to claim 20, which is codon-optimized.

22. A composition comprising a retroviral vector according to any one of claims 1 to 18 or 21, a nucleic acid according to claim 19 or 21, or a plasmid according to claim 20 or 21, and a pharmaceutically acceptable carrier.

23. A host cell comprising a retroviral vector according to any one of claims 1 to 18 or 21, a nucleic acid according to claim 19 or 21, or a plasmid according to claim 20 or 21.

24. A retroviral vector according to any one of claims 1 to 18 or 21, a nucleic acid according to claim 19 or 21, a plasmid according to claim 20 or 21, or a composition according to claim 22 for use in a method of treatment.

25. 1. A method for producing a retroviral vector, comprising the steps of: (a) growing cells in suspension; (b) transfecting one or more plasmids into said cells; (c) adding a nuclease; (d) harvesting the lentivirus; (e) adding trypsin; and (f) Purification Including, The one or more plasmids comprise a nucleic acid as defined in claim 19, and optionally (i) a promoter as defined in claim 13 and / or (ii) 15. The method comprising a vector genome plasmid containing a transgene as defined in claim 14.

26. 1. A method for distinguishing between a retroviral vector and a transgene expressed by said retroviral vector, comprising: (a1) transducing a cell with a retroviral vector as defined in any one of claims 1 to 18; (b1) culturing the cells to allow expression of the transgene by the retrovirus; and (c1) quantifying the RNA in the cells; or (a2) quantifying intracellular RNA in a sample obtained from a patient treated with the retroviral vector of any one of claims 1 to 18 or 21, the nucleic acid of claim 19 or 21, the plasmid of claim 20 or 21, or the composition of claim 23 Including, (i) the amount of RNA containing the chimeric intron with the retroviral RRE inserted corresponds to the copy number of the retroviral vector; (ii) the amount of RNA lacking the chimeric intron into which the retroviral RRE has been inserted corresponds to the amount of transgene mRNA; Optionally, the RNA is quantified by a PCR-based assay or an in situ hybridization-based assay.