Acyclic lipids and methods of use thereof

JP2024534697A5Pending Publication Date: 2025-09-25RENEGADE THERAPEUTICS MANAGEMENT INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024540673
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-04-28
Filing Date
2022-09-14
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Current lipid-based delivery systems for nucleic acids and proteins lack targeted localization to specific tissues or organs, focusing primarily on cargo protection rather than precise delivery.

Method used

Development of optimized lipid systems for localized delivery of nucleic acid and protein therapeutics, utilizing a directed discovery platform to evaluate targeting systems and identify specific tropism profiles, incorporating cationic, neutral, anionic, and stealth lipids with polynucleotides to induce an immune response.

Benefits of technology

Achieves targeted delivery of nucleic acids and proteins to immune cells, inducing a significant immune response with inhibition or suppression of target expression up to 100% in specific cells, enhancing therapeutic efficacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2023044343000001
    Figure 2023044343000001
  • Figure 2023044343000002
    Figure 2023044343000002
  • Figure 2023044343000003
    Figure 2023044343000003
Patent Text Reader

Abstract

The present disclosure details various lipids, compositions, and / or methods of optimized systems and delivery vehicles for the delivery of nucleic acid sequences, polypeptides or peptides for use in immunizing against infectious agents.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of, and priority to, U.S. Provisional Patent Application Nos. 63 / 244,152, filed September 14, 2021; 63 / 293,284, filed December 23, 2021; and 63 / 336,018, filed April 28, 2022, the contents of each of which are hereby incorporated by reference in their entirety.

[0002] The present disclosure relates to optimized systems for the delivery of nucleic acid sequences, polypeptides or peptides and methods of using these optimized systems for the treatment of diseases, disorders and / or pathologies. [Background technology]

[0003] While proteins have become the standard for therapeutic agents, the use of nucleic acids as a therapeutic modality for a variety of diseases and therapeutic indications has gained attention over the past few years. Although various companies have shown that nucleic acids (e.g., siRNA, mRNA, circular RNA, DNA, ASO, etc.) can be more effective compared to protein-based therapies, targeted delivery systems for both nucleic acid and protein therapeutics are needed to ensure that the therapeutics are localized to the targeted cells, tissues, or organs.

[0004] Current delivery systems, including lipid-based delivery systems such as lipid nanoparticles, focus on protecting the cargo being delivered, but not on the lipid used for the delivery system, and often not on localized delivery of the cargo or the delivery system. There is a need in the art for improved lipid-based delivery systems. Summary of the Invention

[0005] The present disclosure provides a directed discovery platform for screening and developing new lipids that can be used in delivery vehicles of delivery systems and targeting systems for localized delivery of nucleic acid and protein therapeutics, e.g., to immune cells.

[0006] In an aspect of the disclosure, there is provided herein a lipid having the structure of any of formulas (VII-A), (VII-B), (VII-C), (IA), (II), (III-B), (III-C), (III-D), (III-E), (III-F), (VIII-B), (IV), (VI), and (X), or a pharma- ceutically acceptable salt thereof, or any lipid in Table (I), or a salt or solvate thereof (see below) (collectively referred to as the "Lipids of the Disclosure" and each individually referred to as a "Lipid of the Disclosure").

[0007] In an aspect of the disclosure, provided herein is a pharmaceutical composition comprising: a) a polynucleotide encoding at least one protein of interest, and b) A delivery vehicle comprising at least one lipid wherein the composition induces an immune response in a subject.

[0008] In one embodiment, the polynucleotide is DNA.

[0009] In one embodiment, the polynucleotide is RNA.

[0010] In one embodiment, the RNA is a small interfering RNA (siRNA).

[0011] In one embodiment, the siRNA inhibits or silences expression of a target of interest in a cell.

[0012] In one embodiment, the inhibition or suppression is about 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95% and 100%, or at least 20-30%, 20-40%, 20-50%, 20-60%, 20-70%, 20-80%, 20-90%, 20-95%, 20-100%, 30-40%, 30-50%, 30-60%, 30-70%, 30-80%, 30-90%, 30-95%, 30-100%, 40-50%, 40-50%, 40-60%, 40-70%, 40-80%, 40-90%, 40-95%, 40-10 ...100%, 40-100%, 40-100%, 40 60%, 40-70%, 40-80%, 40-90%, 40-95%, 40-100%, 50-60%, 50-70%, 50-80%, 50-90%, 50-95%, 50-100%, 60-70%, 60-80%, 60-90%, 60-95%, 60-100%, 70-80%, 70-90%, 70-95%, 70-100%, 80-90%, 80-95%, 80-100%, 90-95%, 90-100% or 95-100%.

[0013] In one aspect, the polynucleotide is substantially circular.

[0014] In one embodiment, the polynucleotide comprises an internal ribosome entry site (IRES) sequence operably linked to the payload sequence region.

[0015] In one embodiment, the IRES sequence comprises a sequence derived from a picornavirus complementary DNA, an encephalomyocarditis virus (EMCV) complementary DNA, a poliovirus complementary DNA, or the antennapedia gene from Drosophila melanogaster.

[0016] In one embodiment, the polynucleotide comprises a terminator element, the terminator element comprising at least one stop codon.

[0017] In one aspect, the polynucleotide comprises a regulatory element.

[0018] In one embodiment, the polynucleotide comprises at least one masking agent.

[0019] In one aspect, the substantially circular polynucleotide is generated using in vitro transcription.

[0020] In one aspect, the payload sequence region comprises a non-coding nucleic acid sequence.

[0021] In one aspect, the payload sequence region comprises a coding nucleic acid sequence.

[0022] In one embodiment, the coding nucleic acid sequence encodes a protein of interest for Campylobacter jejuni. In one embodiment, the coding nucleic acid sequence encodes a protein of interest for Clostridium difficile. In one embodiment, the coding nucleic acid sequence encodes a protein of interest for Entamoeba histolytica. In one embodiment, the coding nucleic acid sequence encodes a protein of interest for Enterotoxin B. In one embodiment, the coding nucleic acid sequence encodes a protein of interest for Norwalk virus or Norovirus. In one embodiment, the coding nucleic acid sequence encodes a protein of interest for Helicobacter pylori. In one embodiment, the coding nucleic acid sequence encodes a protein of interest for Rotavirus. In one embodiment, the coding nucleic acid sequence encodes a protein of interest for Candida yeast. In one embodiment, the coding nucleic acid sequence encodes a protein of interest for Coronavirus. In one embodiment, the coding nucleic acid sequence encodes a protein of interest for SARS-CoV. In one embodiment, the coding nucleic acid sequence encodes a protein of interest for SARS-CoV-2. In one embodiment, the coding nucleic acid sequence encodes a protein of interest for MERS-CoV. In one embodiment, the coding nucleic acid sequence encodes a protein of interest for Enterovirus 71. In one embodiment, the coding nucleic acid sequence encodes a protein of interest for Epstein-Barr Virus. In one embodiment, the coding nucleic acid sequence encodes a protein of interest for a Gram-negative bacterium. In one embodiment, the Gram-negative bacterium is Bordetella. In one embodiment, the coding nucleic acid sequence encodes a protein of interest for a Gram-positive bacterium. In one embodiment, the Gram-positive bacterium is Clostridium tetani. In one embodiment, the Gram-positive bacterium is Francisella tularensis. In one embodiment, the Gram-positive bacterium is a Streptococcus bacterium. In one embodiment, the Gram-positive bacterium is a Staphylococcus bacterium. In one embodiment, the coding nucleic acid sequence encodes a protein of interest for Hepatitis.In one aspect, the coding nucleic acid sequence encodes a protein of interest for human cytomegalovirus. In one aspect, the coding nucleic acid sequence encodes a protein of interest for human immunodeficiency virus. In one aspect, the coding nucleic acid sequence encodes a protein of interest for human papillomavirus. In one aspect, the coding nucleic acid sequence encodes a protein of interest for influenza. In one aspect, the coding nucleic acid sequence encodes a protein of interest for John Cunningham virus. In one aspect, the coding nucleic acid sequence encodes a protein of interest for Mycobacterium. In one aspect, the coding nucleic acid sequence encodes a protein of interest for poxvirus. In one aspect, the coding nucleic acid sequence encodes a protein of interest for Pseudomonas aeruginosa. In one aspect, the coding nucleic acid sequence encodes a protein of interest for respiratory syncytial virus. In one aspect, the coding nucleic acid sequence encodes a protein of interest for rubella virus. In one aspect, the coding nucleic acid sequence encodes a protein of interest for varicella zoster virus. In one aspect, the coding nucleic acid sequence encodes a protein of interest for chikungunya virus. In one aspect, the coding nucleic acid sequence encodes a protein of interest for Dengue virus. In one aspect, the coding nucleic acid sequence encodes a protein of interest for Rabies virus. In one aspect, the coding nucleic acid sequence encodes a protein of interest for Trypanosoma cruzi and / or Chagas disease. In one aspect, the coding nucleic acid sequence encodes a protein of interest for Ebola virus. In one aspect, the coding nucleic acid sequence encodes a protein of interest for Plasmodium falciparum. In one aspect, the coding nucleic acid sequence encodes a protein of interest for Marburg virus. In one aspect, the coding nucleic acid sequence encodes a protein of interest for Japanese encephalitis virus. In one aspect, the coding nucleic acid sequence encodes a protein of interest for St. Louis encephalitis virus. In one aspect, the coding nucleic acid sequence encodes a protein of interest for West Nile virus.In one aspect, the coding nucleic acid sequence encodes a protein of interest for a Yellow Fever Virus. In one aspect, the coding nucleic acid sequence encodes a protein of interest for a Bacillus anthracis. In one aspect, the coding nucleic acid sequence encodes a protein of interest for a botulinum toxin. In one aspect, the coding nucleic acid sequence encodes a protein of interest for a ricin. In one aspect, the coding nucleic acid sequence encodes a protein of interest for a Shiga toxin and / or a Shiga-like toxin. In one aspect, the polynucleotide comprises at least one modification.

[0023] In one embodiment, at least 20% of the bases are modified. In one embodiment, at least 30% of the bases are modified. In one embodiment, at least 40% of the bases are modified. In one embodiment, at least 50% of the bases are modified. In one embodiment, at least 60% of the bases are modified. In one embodiment, at least 70% of the bases are modified. In one embodiment, at least 80% of the bases are modified. In one embodiment, at least 90% of the bases are modified. In one embodiment, at least 100% of the bases are modified. In one embodiment, a particular base comprises at least one modification.

[0024] In one embodiment, the base is adenine. In one embodiment, at least 20% of the adenine bases are modified. In one embodiment, at least 30% of the adenine bases are modified. In one embodiment, at least 40% of the adenine bases are modified. In one embodiment, at least 50% of the adenine bases are modified. In one embodiment, at least 60% of the adenine bases are modified. In one embodiment, at least 70% of the adenine bases are modified. In one embodiment, at least 80% of the adenine bases are modified. In one embodiment, at least 90% of the adenine bases are modified. In one embodiment, at least 100% of the adenine bases are modified.

[0025] In one embodiment, the base is guanine. In one embodiment, at least 20% of the guanine bases are modified. In one embodiment, at least 30% of the guanine bases are modified. In one embodiment, at least 40% of the guanine bases are modified. In one embodiment, at least 50% of the guanine bases are modified. In one embodiment, at least 60% of the guanine bases are modified. In one embodiment, at least 70% of the guanine bases are modified. In one embodiment, at least 80% of the guanine bases are modified. In one embodiment, at least 90% of the guanine bases are modified. In one embodiment, at least 100% of the guanine bases are modified.

[0026] In one embodiment, the base is cytosine. In one embodiment, at least 20% of the cytosine bases are modified. In one embodiment, at least 30% of the cytosine bases are modified. In one embodiment, at least 40% of the cytosine bases are modified. In one embodiment, at least 50% of the cytosine bases are modified. In one embodiment, at least 60% of the cytosine bases are modified. In one embodiment, at least 70% of the cytosine bases are modified. In one embodiment, at least 80% of the cytosine bases are modified. In one embodiment, at least 90% of the cytosine bases are modified. In one embodiment, at least 100% of the cytosine bases are modified.

[0027] In one embodiment, the base is uracil. In one embodiment, at least 20% of the uracil bases are modified. In one embodiment, at least 30% of the uracil bases are modified. In one embodiment, at least 40% of the uracil bases are modified. In one embodiment, at least 50% of the uracil bases are modified. In one embodiment, at least 60% of the uracil bases are modified. In one embodiment, at least 70% of the uracil bases are modified. In one embodiment, at least 80% of the uracil bases are modified. In one embodiment, at least 90% of the uracil bases are modified. In one embodiment, at least 100% of the uracil bases are modified.

[0028] In one embodiment, the at least one modification is a pyridin-4-one ribonucleoside, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-taurinomethyluridine, 1-taurinomethyl-pseudouridine, 5-taurinomethyl-2-thio-uridine, 1 ... linomethyl-4-thio-uridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine, dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, and 4-methoxy-2-thio-pseudouridine. Douridine, 5-aza-cytidine, pseudoisocytidine, 3-methyl-cytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudoisocytidine, pyrrolo-cytidine, pyrrolo-pseudoisocytidine, 2-thio-cytidine, 2-thio-5-methyl-cytidine, 4-thio-pseudoisocytidine, 4-thio-1-methyl-pseudoisocytidine, 4-thio-1-methyl-1-deaza-pseudoisocytidine, 1-methyl-1-deaza-pseudoisocytidine, Zebularine, 5-aza-zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, and 4-methoxy-1-methyl-pseudoisocytidine, 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-Diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyladenosine, N6-(cis-hydroxyisopentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6-glycinylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methylthio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine, and 2-methoxy-adenine, inosine, 1-methyl-inosine, 7-isobutyric acid, and 1-isobutyric acid. 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, or N2,N2-dimethyl-6-thio-guanosine.

[0029] In one aspect, the pharmaceutical composition comprises at least one cationic lipid selected from the group consisting of any lipid in Table (I); any lipid having the structure of formula (VII-A), (VIII-A), (IX-A), (VII-B), (VII-C), (IA), (II), (III-B), (III-C), (III-D), (III-E), (III-F), (VIII-B), (IV), (VI), (X), or (XA); and combinations thereof.

[0030] In one embodiment, the pharmaceutical composition comprises an additional cationic lipid.

[0031] In one embodiment, the pharmaceutical composition comprises a neutral lipid.

[0032] In one embodiment, the pharmaceutical composition comprises an anionic lipid.

[0033] In one embodiment, the pharmaceutical composition comprises a helper lipid.

[0034] In one aspect, the pharmaceutical composition comprises a stealth lipid.

[0035] In one embodiment, the weight ratio of lipid to polynucleotide is about 100:1 to about 1:1.

[0036] In one embodiment, the pharmaceutical composition delivers cargo or payload to immune cells in a subject that needs it.Immune cells can be T cells, for example, CD8+ T cells, CD4+ T cells, or T regulatory cells.Immune cells can also be, for example, macrophages or dendritic cells.

[0037] In one aspect, the vaccine formulation comprises a pharmaceutical composition.

[0038] In one aspect, provided herein is a method of immunizing a subject against an infectious agent, the method comprising contacting the subject with a vaccine formulation or preparation and eliciting an immune response.

[0039] In one embodiment, the infectious agent is selected from the group consisting of Campylobacter jejuni, Clostridium difficile, Entamoeba histolytica, enterotoxin B, Norwalk virus or norovirus, Helicobacter pylori, rotavirus, Candida yeast, coronaviruses including SARS-CoV, SARS-CoV-2, and MERS-CoV, enterovirus 71, Epstein-Barr virus, gram-negative bacteria including Bordetella, gram-positive bacteria including Clostridium tetani, Francisella tularensis, Streptococcus, and Staphylococcus, and hepatitis, human cytomegalovirus, human immunodeficiency virus, human papillomavirus, influenza, John Cunningham virus, Mycobacterium, poxvirus, Pseudomonas aeruginosa, respiratory syncytial virus, rubella virus, varicella zoster virus, chikungunya virus, dengue virus, rabies virus, Trypanosoma cruzi and / or Chagas disease, Ebola virus, Plasmodium falciparum, Marburg virus, Japanese encephalitis virus, St. Louis encephalitis virus, West Nile virus, yellow fever virus, Bacillus anthracis, botulinum toxin, ricin, or Shiga toxin and / or Shiga-like toxin.

[0040] In one embodiment, the contacting can be enteral (into the gut), gastrointestinal, peridural (into the dura), oral (via the mouth), transdermal, intracerebral (into the brain), intraventricular (into the ventricle), epicutaneous (applied to the skin), intradermal (into the skin itself), subcutaneous (under the skin), intranasal administration (through the nose), intravenous (into a vein), intravenous bolus, intravenous infusion, intraarterial (into an artery), intramuscular (into a muscle), intracardiac (into the heart), intraosseous injection (into bone marrow), intrathecal (into the spinal canal), intraparenchymal (into brain tissue), intraperitoneal (infusion or injection into the abdominal cavity), intravesical infusion, intravitreal (through the eye), intracavernous, intraven ... Intravenous injection (into a pathological cavity), intracavitary (into the base of the penis), intravaginal administration, intrauterine, extra-amniotic administration, transdermal (diffusion through intact skin for systemic distribution), transmucosal (diffusion through mucous membranes), transvaginal, insufflation (aspiration), sublingual, sublabial, enema, ophthalmic (on the conjunctiva), ear drops, auricular (in or through the ear), buccal (towards the cheek), conjunctival, dermal, dental (into tooth(s)), electroosmosis, intracervical, intrasinus, intratracheal, extracorporeal, hemodialysis, penetration, intrainterstitial, intraperitoneal, intra-amniotic, intra-articular, intrabiliary, intra-bronchial, intra-capsular, intrachondral (into cartilage), intracaudal (into the cauda equina), intracisternal (into the cisterna magna of the cerebellum and medulla oblongata) (into the cornea), intracoronary (into the coronary artery), intracavernosal (into the expandable space of the corpus cavernosum of the penis), intradiscal (into the intervertebral disc), intraductal (into a glandular duct), intraduodenal (into the duodenum), intradural (in or under the dura), intraepidermal (into the epidermis), intraesophageal (into the esophagus), intragastric (into the stomach), intragingival (into the gums), intraileal (into the distal portion of the small intestine), intralesional (introduced directly into or to a localized lesion), intraluminal (into the lumen of a tube), intralymphatic (into the lymph), intramedullary (into the marrow cavity of a bone), intrameningeal (into the meninges), intramyocardial (into the heart) in muscle), intraocular (in the eye), intraovarian (in the ovary), intrapericardial (in the pericardium), intrapleural (in the pleura), intraprostatic (in the prostate), intrapulmonary (in the lungs or their bronchi), intrasinus (in the nasal or periorbital sinuses), intrathecal (in the spinal column), intraarticular (in the synovial cavities of a joint), intratendinous (in the tendons), intratesticular (in the testes), intraarachnoid (in the cerebrospinal fluid at any level of the neuraxis), intrathoracic (in the chest), intraductal (in the tubules of an organ), intratumoral (in a tumor), intratympanic (in the middle ear), intravascular (in the blood vessel(s)), intraventricular (in the ventricles of the brain),Iontophoresis (by electric current that transfers ions of soluble salts into body tissues), irrigation (bathing or flushing a wound or body cavity), laryngeal (directly into the larynx), nasogastric (through the nose into the stomach), occlusive dressing technique (administered topically and then covered with a dressing that occludes the area), ophthalmic (to the external eye), oropharyngeal (directly into the mouth and pharynx), parenteral, transdermal, periarticular, epidural, perineural, periodontal, rectal, respiratory (localized or systemic) into the airways by oral or nasal inhalation for therapeutic effect), retrobulbar (behind the pons or behind the eyeball), soft tissue, subarachnoid, subconjunctival, submucosal, topical, transplacental (through or through the placenta), transtracheal (through the wall of the trachea), transtympanic (through or through the tympanic cavity), ureteral (into the ureter), urethral (into the urethra), vaginal, sacral anesthesia, diagnostic, nerve block, bile perfusion, cardiac perfusion, photopheresis, or spinal cord. [Brief description of the drawings]

[0041] [Figure 1] 1 is a diagram illustrating one embodiment of a directional discovery platform of the present disclosure. [Diagram 2] FIG. 1 is a diagram illustrating originator polynucleotide constructs of the present disclosure, which can be linear or circular. [Figure 3A-1] 1 is a diagram illustrating a set of benchmark polynucleotide constructs of the present disclosure which may comprise at least one barcode region (BC) and / or reverse barcode region (CB) and a payload region (P). [Figure 3A-2] Same as above. [Figure 3B] FIG. 1 is a diagram illustrating a set of benchmark polynucleotide constructs of the present disclosure in which the barcode region (BC) or reverse barcode region (CB) may overlap with the payload region (P). [Figure 3C-1] 1 is a diagram illustrating a set of benchmark polynucleotide constructs of the present disclosure that may include at least one tag and / or label. [Figure 3C-2] Same as above. [Figure 4A-1]1 is a diagram illustrating a series of circular benchmark polynucleotide constructs of the present disclosure that may include at least one barcode region (BC) and / or reverse barcode region (CB) and a payload region (P). [Figure 4A-2] Same as above. [Figure 4A-3] Same as above. [Figure 4B] FIG. 1 is a diagram illustrating a series of circular benchmark polynucleotide constructs of the present disclosure in which the barcode region (BC) or reverse barcode region (CB) may overlap with the payload region (P). [Figure 4C-1] 1 is a diagram illustrating a series of circular benchmark polynucleotide constructs of the present disclosure that may include at least one tag and / or label. [Figure 4C-2] Same as above. [Diagram 5] 1 is a diagram illustrating a series of delivery vehicles of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0042] I. Introduction to Directed Delivery Systems Nucleic acid therapy has become known as a powerful method to treat various diseases and therapeutic indications, considering its versatility, lower immune response and higher efficacy when compared with conventional therapy.For example, nucleic acid therapy includes small interfering molecules (siRNA) to reduce the translation of messenger RNA (mRNA), mRNA as a method to generate targeted targets, circular RNA (oRNA) that can provide continuous production of polypeptides or peptides or can be a sponge to compete with other RNA molecules, and the use of viral vectors to provide continuous production of targeted targets.However, some nucleic acids are unstable and easily decomposed, so they need to be formulated to prevent degradation and to assist intracellular delivery of nucleic acids.

[0043] Current delivery vehicles, including lipid-based delivery vehicles such as lipid nanoparticles and liposomes, focus on protecting the cargo but do not address localizing the delivery of the cargo or delivery vehicle to a specific area in vivo.

[0044] Provided herein is a tropism discovery platform for evaluating targeting systems for localized delivery to specific target regions, cells or tissues. As shown in FIG. 1, the tropism discovery platform can be used to evaluate a lipid nanoparticle (LNP) library and / or a library of AAVs to determine the tropism or signature profile of the targeting system in the library. The library is administered to a subject (e.g., a non-human primate, rabbit, mouse, rat or another mammal), and the subject's organs and tissues can be scanned and / or harvested and analyzed to determine the location of identifiers (e.g., barcodes, labels, signals and / or tags) contained in or associated with the LNPs or AAVs in the library. This analysis provides a tropism signature or profile of each LNP and AAV in the library.

[0045] Originator construct structure The targeting system of the directed discovery platform may include an originator construct that encodes or includes a cargo or payload. An example of an originator polynucleotide construct 100, which may be linear or circular, is provided in FIG. 2. The originator polynucleotide construct 100 may include at least one payload region 10 that is or encodes a payload or cargo of interest. The originator polynucleotide construct 100 may contain one or two flanking regions 20, which may be located 5' to the payload region 10 or 3' to the payload region 10. In some examples, the originator polynucleotide construct 100 does not contain a flanking region 20. The flanking region 20 of the originator polynucleotide construct 100 may include at least one regulatory region 30. At least one flanking region 20 of the originator polynucleotide construct 100 may include at least one identifier region 40. The identifier region 40 may be, but is not limited to, a barcode, a label, a signal and / or a tag, and may be located within the payload region 10 or may be located in the payload region 10 and at least one adjacent region 20.

[0046] In some embodiments, the originator construct comprises about 5 to about 10,000 residues. As non-limiting examples, the length of the originator construct can be 5 to 30, 5 to 50, 5 to 100, 5 to 250, 5 to 500, 5 to 1,000, 5 to 1,500, 5 to 3,000, 5 to 5,000, 5 to 7,000, 5 to 10,000, 30 to 50, 30 to 100, 30 to 250, 30 to 500, 30 to 1,000, 30 to 1, 500, 30~3,000, 30~5,000, 30~7,000, 30~10,000, 100~250, 100~500, 100~1,000, 100~1,500, 100~3,000, 100~5,000, 100~7,000, 100~10,000, 500~1,000, 500~1,500, 500~2,000 , 500~3,000, 500~5,000, 500~7,000, 500~10,000, 1,000~1,500, 1,000~2,000, 1,000~3,000, 1,000~5,000, 1,000~7,000, 1,000~10,000, 1,500~3,000, 1,500~5,000, 1,500~7, 000, 1,500 to 10,000, 2,000 to 3,000, 2,000 to 5,000, 2,000 to 7,000, 2,000 to 10,000, 3,000 to 5,000, 3,000 to 7,000, 3,000 to 10,000, 5,000 to 7,000, 5,000 to 10,000, and 7,000 to 10,000.

[0047] In some embodiments, the length of the payload region is greater than about 5 residues in length, for example, but not limited to, at least about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120, 140, 160, 180, 20 0, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,500, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000 or more than 10,000 residues or more.

[0048] In some embodiments, the flanking regions are independently 0 to 10,000 residues in length, for example, but not limited to, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120, 140, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 7 The range can be 0, 160, 180, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,500, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, and 10,000.

[0049] In some embodiments, the regulatory regions are independently 0 to 3,000 residues in length, for example, but not limited to, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 40, 45, 50, 55, 60, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 120, 121, 122, 123, 124, 125, 126, 127, 128, 130, 131, 132, 133, 134, 135, 136, 137, 138, 1 The range can be 0, 70, 80, 90, 100, 120, 140, 160, 180, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,500, and 3,000.

[0050] In some embodiments, the originator construct may be circularized or concatamerized to generate molecules to aid in the interaction between the 3' and 5' ends of the originator construct.

[0051] Benchmark construct structure An originator construct that includes at least one identifier (e.g., a barcode, label, signal, and / or tag) is referred to as a benchmark construct. A benchmark polynucleotide construct may include 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more identifiers that may be the same or different throughout the benchmark polynucleotide construct.

[0052] In some embodiments, the identifier regions are independently 1 to 3,000 residues in length, for example, but not limited to, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 40, 45, 50, 55, 60, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 40, 45, 50, 55, 56, 57 The range can be 0, 70, 80, 90, 100, 120, 140, 160, 180, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,500, and 3,000.As non-limiting examples, the identifier region may be 1-5 residues, 2-5 residues, 3-5 residues, 2-7 residues, 3-7 residues, 1-10 residues, 2-10 residues, 3-10 residues, 5-10 residues, 7-10 residues, 1-15 residues, 2-15 residues, 3-15 residues, 5-15 residues, 7-15 residues, 10-15 residues, 12-15 residues, 1-20 residues, 2-20 residues, 3-20 residues, 5-20 residues, 7-20 residues, 10-20 residues, 12-20 residues, 15-20 residues, 17-20 residues, 1-25 residues, 2-25 residues, 3 ~25 residues, 5~25 residues, 7~25 residues, 10~25 residues, 12~25 residues, 15~25 residues, 17~25 residues, 20~25 residues, 1~30 residues, 2~30 residues, 3~30 residues, 5~30 residues, 7~30 residues, 10~30 residues, 12~30 residues, 15~30 residues, 17~30 residues, 20~30 residues, 25~30 residues, 1~35 residues, 2~35 residues, 3~35 residues, 5~35 residues, 7~35 residues, 10~35 residues, 12~35 residues, 15~35 residues, 17~35 residues, 20~35 residues, 25-35 residues, 30-35 residues, 1-35 residues, 2-35 residues, 3-35 residues, 5-35 residues, 7-35 residues, 10-35 residues, 12-35 residues, 15-35 residues, 17-35 residues, 20-35 residues, 25-35 residues, 30-35 residues, 1-40 residues, 2-40 residues, 3-40 residues, 5-40 residues, 7-40 residues, 10-40 residues, 12-40 residues, 15-40 residues, 17-40 residues, 20-40 residues, 25-40 residues, 30-40 residues, 35-40 residues, 1-45 residues, 2-45 residues The number of residues may be 1 to 50, 2 to 50, 3 to 50, 5 to 50, 7 to 45, 10 to 45, 12 to 45, 15 to 45, 17 to 45, 20 to 45, 25 to 45, 30 to 45, 35 to 45, 40 to 45, 1 to 50, 2 to 50, 3 to 50, 5 to 50, 7 to 50, 10 to 50, 12 to 50, 15 to 50, 17 to 50, 20 to 50, 25 to 50, 30 to 50, 35 to 50, 40 to 50, or 45 to 50 residues.

[0053] Non-limiting examples of benchmark polynucleotide constructs having at least one identifier, which may be linear or circular, are provided in Figures 3A, 3B, and 3C. Non-limiting examples of circular benchmark polynucleotide constructs having at least one identifier are provided in Figures 4A, 4B, and 4C. In Figures 3A, 3B, 4A, and 4B, the benchmark polynucleotide constructs include a payload region (referred to as "P" in the figures) and at least one identifier region (referred to as "BC" in the figures) and / or a reverse identifier region (referred to as "CB" in the figures). In Figures 3C and 4C, the benchmark polynucleotide constructs include a payload region (referred to as "P" in the figures) and at least one identifier site associated with the benchmark polynucleotide construct.

[0054] In some embodiments, the identifier region in the benchmark construct overlaps with the payload region. As used herein, "overlapping" means that at least one nucleotide of the identifier region extends into the payload region. In some embodiments, the identifier region overlaps with the payload region by 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides, 21 nucleotides, 22 nucleotides, 23 nucleotides, 24 nucleotides, 25 nucleotides, 26 nucleotides, 27 nucleotides, 28 nucleotides, 29 nucleotides, 30 nucleotides, 31 nucleotides, 32 nucleotides, 33 nucleotides, 34 nucleotides, 35 nucleotides, 36 nucleotides, 37 nucleotides, 38 nucleotides, 39 nucleotides, 40 nucleotides 41 nucleotides, 42 nucleotides, 43 nucleotides, 44 nucleotides, 45 nucleotides, 46 nucleotides, 47 nucleotides, 48 ​​nucleotides, 49 nucleotides, 50 nucleotides or more than 50 nucleotides. In some embodiments, the identifier region is 1-5 nucleotides, 2-5 nucleotides, 3-5 nucleotides, 2-7 nucleotides, 3-7 nucleotides, 1-10 nucleotides, 2-10 nucleotides, 3-10 nucleotides, 5-10 nucleotides, 7-10 nucleotides, 1-15 nucleotides, 2-15 nucleotides, 3-15 nucleotides, 5-15 nucleotides, 7-15 nucleotides, 10-15 nucleotides, 12-15 nucleotides, 1-20 nucleotides, 2-20 nucleotides, 3-20 nucleotides, 5-20 nucleotides, 7-20 nucleotides, 10-20 nucleotides, 12-20 nucleotides, 15-20 nucleotides, 17-20 nucleotides, 1-25 nucleotides, 2-25 nucleotides, 3-25 nucleotides, 5-25 nucleotides, 7-25 nucleotides, 10-25 nucleotides, 12-25 nucleotides, 15-25 nucleotides,17-25 nucleotides, 20-25 nucleotides, 1-30 nucleotides, 2-30 nucleotides, 3-30 nucleotides, 5-30 nucleotides, 7-30 nucleotides, 10-30 nucleotides, 12-30 nucleotides, 15-30 nucleotides, 17-30 nucleotides, 20-30 nucleotides, 25-30 nucleotides, 1-35 nucleotides, 2-35 nucleotides, 3-35 nucleotides, 5-35 nucleotides, 7-35 nucleotides, 10-35 nucleotides, 12-35 nucleotides, 1 5-35 nucleotides, 17-35 nucleotides, 20-35 nucleotides, 25-35 nucleotides, 30-35 nucleotides, 1-35 nucleotides, 2-35 nucleotides, 3-35 nucleotides, 5-35 nucleotides, 7-35 nucleotides, 10-35 nucleotides, 12-35 nucleotides, 15-35 nucleotides, 17-35 nucleotides, 20-35 nucleotides, 25-35 nucleotides, 30-35 nucleotides, 1-40 nucleotides, 2-40 nucleotides, 3-40 nucleotides, 5-40 nucleotides, 7-40 nucleotides, 10-40 nucleotides, 12-40 nucleotides, 15-40 nucleotides, 17-40 nucleotides, 20-40 nucleotides, 25-40 nucleotides, 30-40 nucleotides, 35-40 nucleotides, 1-45 nucleotides, 2-45 nucleotides, 3-45 nucleotides, 5-45 nucleotides, 7-45 nucleotides, 10-45 nucleotides, 12-45 nucleotides, 15-45 nucleotides, 17-45 nucleotides, 20-45 nucleotides overlapping by 25-45 nucleotides, 30-45 nucleotides, 35-45 nucleotides, 40-45 nucleotides, 1-50 nucleotides, 2-50 nucleotides, 3-50 nucleotides, 5-50 nucleotides, 7-50 nucleotides, 10-50 nucleotides, 12-50 nucleotides, 15-50 nucleotides, 17-50 nucleotides, 20-50 nucleotides, 25-50 nucleotides, 30-50 nucleotides, 35-50 nucleotides, 40-50 nucleotides, or 45-50 nucleotides.

[0055] In some embodiments, a benchmark polynucleotide construct comprises a payload region and an identifier region. The identifier region may be located 5' to the payload region, 3' to the payload region, or the identifier region may overlap the 5' or 3' end of the payload region.

[0056] In some embodiments, a benchmark polynucleotide construct comprises a payload region and two identifier regions, each of which may be independently located 5' to the payload region, 3' to the payload region, or the identifier region may overlap the 5' or 3' end of the payload region.

[0057] As a non-limiting example, the first identifier region is located 5' relative to the payload region and the second identifier region is located 3' relative to the payload region. As a non-limiting example, the first and second identifier regions are located 5' relative to the payload region. As a non-limiting example, the first and second identifier regions are located 3' relative to the payload region.

[0058] As a non-limiting example, the first identifier region is reversed and located 5' relative to the payload region and the second identifier region is located 3' relative to the payload region. As a non-limiting example, the first identifier region is reversed and located 5' relative to the payload region and the second identifier region is reversed and located 3' relative to the payload region. As a non-limiting example, the first identifier region is located 5' relative to the payload region and the second identifier region is reversed and located 3' relative to the payload region. As a non-limiting example, the first and second identifier regions are both reversed and located 5' relative to the payload region. As a non-limiting example, the first and second identifier regions are both 5' relative to the payload region and the first identifier region is reversed. As a non-limiting example, the first and second identifier regions are both 5' relative to the payload region and the second identifier region is reversed. As a non-limiting example, the first and second identifier regions are both reversed and located 3' relative to the payload region. As a non-limiting example, the first and second identifier regions are located 3' to the payload region and the first identifier region is reversed. As a non-limiting example, the first and second identifier regions are located 3' to the payload region and the second identifier region is reversed.

[0059] As a non-limiting example, the first identifier region is located 5' to the payload region and overlaps with the payload region, and the second identifier region is located 3' to the payload region. As a non-limiting example, the first identifier region is located 5' to the payload region and the second identifier region is located 3' to the payload region and overlaps with the payload region.

[0060] As a non-limiting example, the first and second identifier regions are located 5' to the payload region and the second identifier region overlaps with the payload region. As a non-limiting example, the first and second identifier regions are located 3' to the payload region and the first identifier region overlaps with the payload region.

[0061] As a non-limiting example, the first identifier region is reversed and located 5' to the payload region and overlaps with the payload region and the second identifier region is located 3' to the payload region. As a non-limiting example, the first identifier region is reversed and located 5' to the payload region and the second identifier region is located 3' to the payload region and overlaps with the payload region. As a non-limiting example, the first identifier region is reversed and located 5' to the payload region and the second identifier region is located 3' to the payload region and both the first and second identifier regions overlap with the payload region.

[0062] As a non-limiting example, the first identifier region is reversed and located 5' to the payload region and overlaps with the payload region, and the second identifier region is reversed and located 3' to the payload region. As a non-limiting example, the first identifier region is reversed and located 5' to the payload region and the second identifier region is reversed and located 3' to the payload region and overlaps with the payload region. As a non-limiting example, the first identifier region is reversed and located 5' to the payload region and the second identifier region is reversed and located 3' to the payload region and both the first and second identifier regions overlap with the payload region.

[0063] As a non-limiting example, the first identifier region is located 5' to the payload region and overlaps with the payload region, and the second identifier region is reversed and located 3' to the payload region. As a non-limiting example, the first identifier region is located 5' to the payload region and the second identifier region is reversed and located 3' to the payload region and overlaps with the payload region. As a non-limiting example, the first identifier region is located 5' to the payload region and the second identifier region is reversed and located 3' to the payload region and both the first and second identifier regions overlap with the payload region.

[0064] As a non-limiting example, the first and second identifier regions are inverted relative to each other and are located 5' relative to the payload region, and the second identifier region overlaps with the payload region. As a non-limiting example, the first and second identifier regions are located 5' relative to the payload region, the first identifier region is inverted, and the second identifier region overlaps with the payload region. As a non-limiting example, the first and second identifier regions are located 5' relative to the payload region, the second identifier region is inverted, and overlaps with the payload region. As a non-limiting example, the first and second identifier regions are inverted relative to each other and are located 3' relative to the payload region, and the first identifier region overlaps with the payload region. As a non-limiting example, the first and second identifier regions are located 3' relative to the payload region, and the first identifier region is inverted, and overlaps with the payload region. As a non-limiting example, the first and second identifier regions are located 3' to the payload region, the second identifier region is reversed, and the first payload region overlaps with the payload region.

[0065] In some embodiments, at least one identifier site may be associated with a benchmark polynucleotide construct. A benchmark polynucleotide construct may have 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more identifier sites associated with the benchmark polynucleotide construct, which may be the same site or different sites associated with the benchmark polynucleotide construct. Each identifier site may be independently located on the 5' flanking region to the payload region, on the 3' flanking region to the payload region, or the location of the identifier site may span the 5' or 3' end of the payload region and the flanking region. In some embodiments, the location of the identifier site may include one or more nucleotides of the payload region, such as, but not limited to, 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides, 21 nucleotides, 22 nucleotides, 23 nucleotides, 24 nucleotides, 25 nucleotides, 26 nucleotides, 27 nucleotides, 28 nucleotides, 29 nucleotides, 30 nucleotides, 31 nucleotides, 32 nucleotides, 33 nucleotides, 34 nucleotides, 35 nucleotides, 36 nucleotides, 37 nucleotides, 38 nucleotides, 39 nucleotides, 40 nucleotides 41 nucleotides, 42 nucleotides, 43 nucleotides, 44 nucleotides, 45 nucleotides, 46 nucleotides, 47 nucleotides, 48 ​​nucleotides, 49 nucleotides, 50 nucleotides or more than 50 nucleotides. In some embodiments, the location of the identifier moiety includes one or more nucleotides of the payload region, such as, but not limited to, 1-5 nucleotides, 2-5 nucleotides, 3-5 nucleotides, 2-7 nucleotides, 3-7 nucleotides, 1-10 nucleotides, 2-10 nucleotides, 3-10 nucleotides, 5-10 nucleotides, 7-10 nucleotides, 1-15 nucleotides, 2-15 nucleotides,3-15 nucleotides, 5-15 nucleotides, 7-15 nucleotides, 10-15 nucleotides, 12-15 nucleotides, 1-20 nucleotides, 2-20 nucleotides, 3-20 nucleotides, 5-20 nucleotides, 7-20 nucleotides, 10-20 nucleotides, 12-20 nucleotides, 15-20 nucleotides, 17-20 nucleotides, 1-25 nucleotides, 2-25 nucleotides, 3-25 nucleotides, 5-25 nucleotides, 7-25 nucleotides, 10-25 nucleotides, 12-25 nucleotides, 15-25 nucleotides tide, 17-25 nucleotides, 20-25 nucleotides, 1-30 nucleotides, 2-30 nucleotides, 3-30 nucleotides, 5-30 nucleotides, 7-30 nucleotides, 10-30 nucleotides, 12-30 nucleotides, 15-30 nucleotides, 17-30 nucleotides, 20-30 nucleotides, 25-30 nucleotides, 1-35 nucleotides, 2-35 nucleotides, 3-35 nucleotides, 5-35 nucleotides, 7-35 nucleotides, 10-35 nucleotides, 12-35 nucleotides, 15-35 nucleotides, 17-3 5 nucleotides, 20-35 nucleotides, 25-35 nucleotides, 30-35 nucleotides, 1-35 nucleotides, 2-35 nucleotides, 3-35 nucleotides, 5-35 nucleotides, 7-35 nucleotides, 10-35 nucleotides, 12-35 nucleotides, 15-35 nucleotides, 17-35 nucleotides, 20-35 nucleotides, 25-35 nucleotides, 30-35 nucleotides, 1-40 nucleotides, 2-40 nucleotides, 3-40 nucleotides, 5-40 nucleotides, 7-40 nucleotides, 10-40 nucleotides , 12-40 nucleotides, 15-40 nucleotides, 17-40 nucleotides, 20-40 nucleotides, 25-40 nucleotides, 30-40 nucleotides, 35-40 nucleotides, 1-45 nucleotides, 2-45 nucleotides, 3-45 nucleotides, 5-45 nucleotides, 7-45 nucleotides, 10-45 nucleotides, 12-45 nucleotides, 15-45 nucleotides, 17-45 nucleotides, 20-45 nucleotides, 25-45 nucleotides, 30-45 nucleotides, 35-45 nucleotides, 40-45 nucleotides,It may contain 1 to 50 nucleotides, 2 to 50 nucleotides, 3 to 50 nucleotides, 5 to 50 nucleotides, 7 to 50 nucleotides, 10 to 50 nucleotides, 12 to 50 nucleotides, 15 to 50 nucleotides, 17 to 50 nucleotides, 20 to 50 nucleotides, 25 to 50 nucleotides, 30 to 50 nucleotides, 35 to 50 nucleotides, 40 to 50 nucleotides, or 45 to 50 nucleotides.

[0066] In some embodiments, one identifier site may be associated with a benchmark polynucleotide construct. As a non-limiting example, the identifier site may be associated with a benchmark polynucleotide construct on the 5' end of the benchmark polynucleotide construct. As a non-limiting example, the identifier site may be associated with a benchmark polynucleotide construct on the 5' flanking region. As a non-limiting example, the identifier site may be associated with a benchmark polynucleotide construct on the 3' flanking region. As a non-limiting example, the identifier site may be associated with a benchmark polynucleotide construct on the 3' end of the benchmark polynucleotide construct. As a non-limiting example, the identifier site may be associated with a benchmark polynucleotide construct on the payload region. As a non-limiting example, the benchmark polynucleotide construct includes an identifier site, and the location of the identifier site spans the 5' end and the 5' flanking region of the payload region. As a non-limiting example, the benchmark polynucleotide construct includes an identifier site, and the location of the identifier site spans the 3' end and the 3' flanking region of the payload region.

[0067] In some embodiments, two identifier sites are associated with the benchmark polynucleotide construct. As a non-limiting example, the first identifier site and the second identifier site are located on the 5' flanking region. As a non-limiting example, the first identifier site and the second identifier site are located on the payload region. As a non-limiting example, the first identifier site and the second identifier site are located on the 3' flanking region. As a non-limiting example, the first identifier site and the second identifier site are located on the 5' end of the benchmark polynucleotide construct. As a non-limiting example, the first identifier site and the second identifier site are located on the 3' end of the benchmark polynucleotide construct.

[0068] As a non-limiting example, the first identifier site is located on the 5' end of the benchmark polynucleotide construct and the second identifier site is located on the 5' flanking region. As a non-limiting example, the first identifier site is located on the 5' end of the benchmark polynucleotide construct and the second identifier site is located on the payload region. As a non-limiting example, the first identifier site is located on the 5' end of the benchmark polynucleotide construct and the second identifier site is located on the 3' flanking region. As a non-limiting example, the first identifier site is located on the 5' end of the benchmark polynucleotide construct and the location of the second identifier site spans the 5' flanking region and the payload region. As a non-limiting example, the first identifier site is located on the 5' end of the benchmark polynucleotide construct and the location of the second identifier site spans the 3' flanking region and the payload region. As a non-limiting example, a first identifier site is located on the 5' end of the benchmark polynucleotide construct and a second identifier site is located on the 3' end of the benchmark polynucleotide construct.

[0069] As a non-limiting example, the first identifier site is located on the 5' flanking region and the second identifier site is located on the payload region. As a non-limiting example, the first identifier site is located on the 5' flanking region and the second identifier site is located on the 3' flanking region. As a non-limiting example, the first identifier site is located on the 5' flanking region and the position of the second identifier site spans the 5' flanking region and the payload region. As a non-limiting example, the first identifier site is located on the 5' flanking region and the position of the second identifier site spans the 3' flanking region and the payload region. As a non-limiting example, the first identifier site is located on the 5' flanking region and the second identifier site is located on the 5' end of the benchmark polynucleotide construct. As a non-limiting example, the first identifier site is located on the 5' flanking region and the second identifier site is located on the 3' end of the benchmark polynucleotide construct.

[0070] As a non-limiting example, the location of the first identifier site spans the 5' flanking region and the payload region, and the second identifier site is located on the 5' end of the benchmark polynucleotide construct. As a non-limiting example, the location of the first identifier site spans the 5' flanking region and the payload region, and the second identifier site is located on the 5' flanking region. As a non-limiting example, the location of the first identifier site spans the 5' flanking region and the payload region, and the second identifier site is located on the payload region. As a non-limiting example, the location of the first identifier site spans the 5' flanking region and the payload region, and the second identifier site is located on the 3' flanking region. As a non-limiting example, the location of the first identifier site spans the 5' flanking region and the payload region, and the second identifier site is located on the 3' flanking region. As a non-limiting example, the location of the first identifier site spans the 5' flanking region and the payload region, and the second identifier site is located on the 3' end of the benchmark polynucleotide construct.

[0071] As a non-limiting example, the first identifier site is located on the payload region and the second identifier site is located on the 5' end of the benchmark polynucleotide construct. As a non-limiting example, the first identifier site is located on the payload region and the second identifier site is located on the 5' flanking region. As a non-limiting example, the first identifier site is located on the payload region and the location of the second identifier site spans the 5' flanking region and the payload region. As a non-limiting example, the first identifier site is located on the payload region and the location of the second identifier site spans the 3' flanking region and the payload region. As a non-limiting example, the first identifier site is located on the payload region and the second identifier site is located on the 3' flanking region. As a non-limiting example, the first identifier site is located on the payload region and the second identifier site is located on the 3' end of the benchmark polynucleotide construct.

[0072] As a non-limiting example, the location of the first identifier site spans the 3' flanking region and the payload region, and the second identifier site is located on the 5' end of the benchmark polynucleotide construct. As a non-limiting example, the location of the first identifier site spans the 3' flanking region and the payload region, and the second identifier site is located on the 5' flanking region. As a non-limiting example, the location of the first identifier site spans the 3' flanking region and the payload region, and the second identifier site is located on the 5' flanking region. As a non-limiting example, the location of the first identifier site spans the 3' flanking region and the payload region, and the second identifier site is located on the payload region. As a non-limiting example, the location of the first identifier site spans the 3' flanking region and the payload region, and the second identifier site is located on the 3' flanking region. As a non-limiting example, the location of the first identifier site spans the 3' flanking region and the payload region, and the second identifier site is located on the 3' end of the benchmark polynucleotide construct.

[0073] As a non-limiting example, the location of the first identifier site spans the 3' flanking region and the payload region, and the second identifier site is located on the 5' flanking region. As a non-limiting example, the location of the first identifier site spans the 5' flanking region and the payload region, and the second identifier site is located on the payload region. As a non-limiting example, the location of the first identifier site spans the 5' flanking region and the payload region, and the second identifier site spans the 3' flanking region and the payload region. As a non-limiting example, the location of the first identifier site spans the 5' flanking region and the payload region, and the second identifier site is located on the 3' flanking region. As a non-limiting example, the location of the first identifier site spans the 5' flanking region and the payload region, and the second identifier site is located on the 3' end of the benchmark polynucleotide construct.

[0074] As a non-limiting example, the first identifier site is located on the 3' flanking region and the second identifier site is located on the 5' end of the benchmark polynucleotide construct. As a non-limiting example, the first identifier site is located on the 3' flanking region and the second identifier site is located on the 5' flanking region. As a non-limiting example, the first identifier site is located on the 3' flanking region and the location of the second identifier site spans the 5' flanking region and the payload region. As a non-limiting example, the first identifier site is located on the 3' flanking region and the second identifier site is located on the payload region. As a non-limiting example, the first identifier site is located on the 3' flanking region and the location of the second identifier site spans the 3' flanking region and the payload region. As a non-limiting example, the first identifier site is located on the 3' flanking region and the second identifier site is located on the 3' end of the benchmark polynucleotide construct.

[0075] As a non-limiting example, the first identifier site is located on the 3' end of the benchmark polynucleotide construct and the second identifier site is located on the 5' end of the benchmark polynucleotide construct. As a non-limiting example, the first identifier site is located on the 3' end of the benchmark polynucleotide construct and the second identifier site is located on the 5' flanking region. As a non-limiting example, the first identifier site is located on the 5' end of the benchmark polynucleotide construct and the location of the second identifier site spans the 5' flanking region and the payload region. As a non-limiting example, the first identifier site is located on the 3' end of the benchmark polynucleotide construct and the second identifier site is located on the payload region. As a non-limiting example, the first identifier site is located on the 5' end of the benchmark polynucleotide construct and the location of the second identifier site spans the 3' flanking region and the payload region. As a non-limiting example, a first identifier site is located on the 3' end of the benchmark polynucleotide construct and a second identifier site is located on the 3' adjacent region.

[0076] In some embodiments, three identifier sites are associated with a benchmark polynucleotide construct.

[0077] In some embodiments, four identifier sites are associated with a benchmark polynucleotide construct.

[0078] In some embodiments, five identifier sites are associated with the benchmark polynucleotide construct.

[0079] In some embodiments, six identifier sites are associated with a benchmark polynucleotide construct.

[0080] In some embodiments, seven identifier sites are associated with the benchmark polynucleotide construct.

[0081] In some embodiments, eight identifier sites are associated with the benchmark polynucleotide construct.

[0082] In some embodiments, nine identifier sites are associated with the benchmark polynucleotide construct.

[0083] In some embodiments, ten identifier sites are associated with a benchmark polynucleotide construct.

[0084] II. Cargo and Payload The originator and benchmark constructs of the present disclosure may contain, encode, or be conjugated to a cargo or payload. As used herein, the term "cargo" or "payload" may refer to one or more molecules or structures contained in a delivery vehicle for delivery to or into a cell or tissue. Non-limiting examples of cargo may include nucleic acids, polypeptides, peptides, proteins, liposomes, labels, tags, small chemical molecules, large biological molecules, and any combinations thereof or fragments thereof. In the originator and benchmark constructs, the region of the construct that contains or encodes the cargo or payload is referred to as the "cargo region" or "payload region."

[0085] In some embodiments, the cargo or payload is or codes for a biologically active molecule, such as, but not limited to, a therapeutic protein. As used herein, the term "biologically active" refers to the characteristic of any agent that has activity in a biological system, particularly an organism. For example, an agent that, when administered to an organism, has a biological effect on the organism, is considered biologically active. In some embodiments, the cargo or payload is or codes for one or more prophylactically or therapeutically active proteins, polypeptides, or other factors. As a non-limiting example, the cargo or payload may be or code for an agent that enhances tumor killing activity, such as, but not limited to, TRAIL or tumor necrosis factor (TNF), in cancer. As another non-limiting example, the cargo or payload may be a polypeptide that is capable of binding to or is associated with a muscular dystrophy (e.g., the cargo or payload is or encodes dystrophin), a cardiovascular disease (e.g., the cargo or payload is or encodes SERCA2a, GATA4, Tbx5, Mef2C, Hand2, Myocd, etc.), a neurodegenerative disease (e.g., the cargo or payload is or encodes NGF, BDNF, GDNF, NT-3, etc.), chronic pain (e.g., the cargo or payload is or encodes GlyRal), an enkephalin, or a glutamate. amino acid decarboxylase (e.g., the cargo or payload is or encodes GAD65, GAD67, or another isoform), pulmonary disease (e.g., the cargo or payload is or encodes CFTR), hemophilia (e.g., the cargo or payload is or encodes factor VIII or factor IX), neoplasia (e.g., the cargo or payload is PTEN, ATM, ATR, EGFR, ERBB2, ERBB3, ERBB4, Notchl, Notch2, Notch3, Notch4, AKT, AKT2, AKT3, HIF, HI Fla, HIF3a, Met, HRG, Bcl2, PPAR alpha, PPAR gamma, WT1 (Wilms' tumor), FGF receptor family members (five members: 1, 2, 3, 4, 5), CDKN2a, APC,RB (retinoblastoma), MEN1, VHL, BRCA1, BRCA2, AR (androgen receptor), TSG101, IGF, IGF receptor, Igf1 (4 variants), Igf2 (3 variants), Igf1 receptor, Igf2 receptor, Bax, Bcl2, caspase family (9 members: 1, 2, 3, 4, 6, 7, 8, 9, 12), Kras, Ape), age-related macular degeneration (e.g., the cargo or payload is or encodes Aber, Ccl2, Cc2, cp (ceruloplasmin), Timp3, cathepsin D, Vldlr), schizophrenia (e.g., neuregulin (Nrgl), Erb4 (neuregulin Receptor for DMPK (Myotonic Dystrophy), Complexin-l (Cplxl), Tphl tryptophan hydroxylase, Tph2 tryptophan hydroxylase 2, Neurexin 1, GSK3, GSK3a, GSK3b, 5-HIT (Slc6a4), COMT, DRD (Drdla), SLC6A3, DAOA, DTNBPI, Dao (Daol)), Trinucleotide repeat disorders (e.g., HTT (Huntington's Dx), SBMA / SMAXI / AR (Kennedy Dx), FXN / X25 (Friedreich's ataxia), ATX3 (Machado-Joseph Dx), ATXNI and ATXN2 (Spinocerebellar ataxia), DMPK (Myotonic dystrophy), Atrophin-1 and Atnl (DRPLA Dx), CBP (Creb-BP-global instability), VLDLR (Alzheimer's), Atxn7, Atxn10), Fragile X syndrome (e.g., the cargo or payload is or encodes FMR2, FXRI, FXR2, mGLUR5), secretase-associated disorders (e.g., the cargo or payload is or encodes APH-1 (alpha and beta), presenilin (Psenl), nicastrin (Ncstn), PEN-2), ALS (e.g., the cargo or payload is or encodes SOD1, ALS2, STEX, FUS, TARD BP, VEGF (VEGF-a, VEGF-b, VEGF-c)), autism (e.g., the cargo or payload is or encodes Mecp2, BZRAP1, MDGA2, Sema5A, Neurexin 1), Alzheimer's disease (e.g.,The cargo or payload is or encodes El, CHIP, UCH, UBB, Tau, LRP, PICALM, Clusterin, PS1, SORL1, CR1, Vldlr, Ubal, Uba3, CHIP28 (Aqpl, aquaporin 1), Uchll, Uchl3, APP), inflammation (e.g. the cargo or payload is or encodes IL-10, IL-1 (IL-Ia, IL-Ib), IL-13, IL-17 (IL-17a (CTLA8), IL-17b, IL-17c, IL-17d, IL-171), 11-23, Cx3crl, ptpn22, TNFa, NOD2 / CARD15 for IBD, IL-6, IL-12 (IL-12a, IL-12b), CTLA4, Cx3cll), Parkinson's disease (e.g. x-synuclein, DJ-1, LRRK2, Parkin, PINK1), blood and clotting disorders, e.g., anemia, bare lymphocyte syndrome, bleeding disorders, hemophagocytic lymphohistiocytosis disorder, hemophilia A, hemophilia B, bleeding disorders, white blood cell deficiencies and disorders, sickle cell anemia, and thalassemia (e.g., cargo or payload may be CRAN1, CDA1, RPS19, DBA, PKLR, PK1, NT5C3 , UMPH1, PSNI, RHAG, RH50A, NRAMP2, SPTB, ALAS2, ANH1, ASB, ABCB7, ABC7, ASAT, TAPBP, TPSN, TAP2, ABCB 3, PSF2, RING11, MHC2TA, C2TA, RFX5, RFXAP, RFX5, TBXA2R, P2RX1, P2X1, HF1, CFH, HUS, MCFD2, FANCA, FAC A, FA1, FA, FA A, FAAP95, FAAP90, FLJ34064, FANCB, FANCC, FACC, BRCA2, FANCDI, FANCD2, FANCD, FACD, FAD, FANCE, FACE, FANCF, XRCC9, FANCG, BR1PI, BACH1, FANCJ, PHF9, FANCL, FANCM, KIAA159 6, PRF1, HPLH2, UNC13D, MUNC13-4, HPLH3, HLH3, FHL3, F8, FSC, PI, ATT, F5, ITGB2, CD18, LCAMB, LAD, EIF2B1, EIF2BA, EIF2B2, EIF2B3, EIF2B5, LVWM, CACH, CLE, EIF2B4, HBB, HBA2,HBB, HBD, LCRB, HBA1), B-cell non-Hodgkin's lymphoma or leukemia (e.g., the cargo or payload is or encodes BCL7A, BCL7, ALI, TCL5, SCL, TAL2, FLT3, NBS1, NBS, ZNFN1AI, 1KI, LYF1, HOXD4, HOX4B, BCR, CML, PHL, ALL, ARNT, KRAS2, RASK2, GMPS, AFIO, ARHGEF12, LARG, KIAA0382, CALM, CLTH, CEBPA, CEBP, CHIC2, BTL, FLT3, KIT, PBT, LPP, NPMI, NUP214, D9S46E, CAN, CAIN, RUNXI, CBFA2, AML1, WHSC1LI, NSD3, FLT3, AF1Q, NPMI, NUMA1, ZNF145, PLZF, PML, MYL, STAT5B, AF1Q, CALM, CLTH, ARL11, ARLTS1, P2RX7, P2X7, BCR, CML, PHL, ALL, GRAF, NF1, VRNF, WSS, NFNS, PTPNII, PTP2C, SHP2, NS1, BCL2, CCND1, PRAD1, BCL1, TCRA, GATA1, GF 1, ERYF1, NFE1, ABLI, NQO1, DIA4, NMOR1, NUP214, D9S46E, CAN, CAIN), inflammatory and immune related diseases and disorders (e.g., the cargo or payload is or encodes KIR3DL1, NKAT3, NKB1, AMB11, K1R3DS1, IFNG, CXCL12, TNFRSF6, APT1, FAS, CD95, ALPS1A, IL2RG, SCIDX1, SCIDX, IMD4, CCL5, SCYA5, D17S136E, TCP228, IL10, CSIF, CMKBR2, C CR2, CMKBR5, CCCKR5 (CCR5), CD3E, CD3G, AICDA, AID, HIGM2, TNFRSF5, CD40, UNG, DGU, HIGM4, TNFSFS, CD40LG, HIGM1, IGM, FOXP3, IPEX, AIID, XPID, PIDX, TNFRSF14B, TACI), inflammation (e.g., the cargo or payload is or encodes IL-10, IL-1 (IL-IA, IL-IB), IL-13, IL-17 (IL-17a (CTLA8), IL-17b, IL-17c, IL-17d,IL-171), 11-23, Cx3crl, ptpn22, TNFa, NOD2 / CARD15 for IBD, IL-6, IL-12 (IL-12a, IL-12b), CTLA4, Cx3cII), JAK3, JAKL, DCLREIC, ARTEMIS, SCIDA, RAG1, RAG2, ADA, PTPRC, CD45, LCA, IL7R, CD3D, T3D, IL2RG, SCIDXI, SCIDX, IMD4), metabolic, liver, kidney and protein diseases and disorders (e.g., the cargo or payload is, or encodes, TT R, PALB, APOA1, APP, AAA, CVAP, ADI, GSN, FGA, LYZ, TTR, PALB, KRT18, KRT8, CIRH1A, NAIC, TEX292, KIAA1988, CFTR, ABCC7, CF, MRP7, SLC2A2, GLUT2, G6 PC, G6PT, G6PT1, GAA, LAMP2, LAMPB, AGL, GDE, GBE1, GYS2, PYGL, PFKM, TCF1, HNF1A, MODY3, SCOD1, SCOl, CTNNB1, PDGFRL, PDGRL, PRLTS, AX1NI, AXIN, CT NNB1, TP53, P53, LFS1, IGF2R, MPRI, MET, CASP8, MCH5, UMOD, HNFJ, FJHN, MCKD2, ADMCKD2, PAH, PKU1, QDPR, DHPR, PTS, FCYT, PKHD1, ARPKD, PKD1, PKD2, PKD4, PKDTS, PRKCSH, G19P1, PCLD, SEC63), muscular / skeletal diseases and disorders (e.g., the cargo or payload is or encodes DMD, BMD, MYF6, LMNA, LMN1, EMD2, FPLD, CMDIA, HGPS, L GMDIB, LMNA, LMNI, EMD2, FPLD, CMDIA, FSHMD1A, FSHD1A, FKRP, MDC1C, LGMD2I, LAMA2, LAMM, LARGE, KIAA0609, MDC1D, FCMD, TTID, MYOT, CAPN3, CANP3, D YSF, LGMD2B, SGCG, LGMD2C, DMDA1, SCG3, SGCA, ADL, DAG2, LGMD2D, DMDA2, SGCB, LGMD2E, SGCD, SGD, LGMD2F, CMD1L, TCAP, LGMD2G, CMD1N, TRIM32, HT2A,LGMD2H, FKRP, MDCIC, LGMD21, TTN, CMD1G, TMD, LGMD2J, POMT1, CAV3, LGMD1C, SEPN1, SELN, RSMD1, PLEC1, PLTN, EBS1, LRP5, BMNDl, LRP7, LR3, OPPG, VBCH2, CL, CN7, CLC7, OPTA2, OSTMI, GL, TCIRG1, TIRC7, OC116, OPTB1, VAPB, VAPC, ALS8, SMN1, SMA1, SMA2, SMA3, SMA4, BSCL2, SPG17, GARS, SMAD1, CMT2D, HEXB, IGHMBP2, SMUBP2, CATF1, SMARD1), neurological and neuronal diseases and disorders (e.g., the cargo or payload is or encodes SOD1, ALS2, STEX, FUS, TARDBP, VEGF (VEGF-a, VEGF-b, VEGF F-c), APP, AAA, CVAP, ADI, APOE, AD2, PSEN2, AD4, STM2, APBB2, FE65LI, NOS3, PLAU, URK, ACE, DCPI, ACEI, MPO, PAC1PI, PAXIPIL, PTIP, A2M, BLMH, BMH, PSEN1, AD3, Mecp2, BZRAP1, MDGA2, Sema5A, Neurexin 1, GLOl, MECP2, RTT, PPMX, MRX16, MRX79, NLGN3, NLGN4, KIAA1260, AUTSX2, FMR2, FXR1, FXR2, mGLUR 5, HD, IT15, PRNP, PRIP, JPH3, JP3, HDL2, TBP, SCA17, NR4A2, NURR1, NOT, TINUR, SNCAIP, TBP, SCA17, SNCA, NACP, PARK1, PARK4, DJI, PARK7, LRRK2, PAR K8, PINK1, PARK6, UCHL1, PARK5, SNCA, NACP, PARKl, PARK4, PRKN, PARK2, PDJ, DBH, NDUFV2, MECP2, RTT, PPMX, MRX16, MRX79, CDKL5, STK9, MECP2, RTT, P PMX, MRX16, MRX79, x-synuclein, DJ-1, Neuregulin-l (Nrgl), Erb4 (receptor for neuregulin), Complexin-l (Cplxl), Tphl tryptophan hydroxylase, Tph2, tryptophan hydroxylase 2, Neurexin 1, GSK3, GSK3a, GSK3b, 5-HTT (Slc6a4), CONT, DRD (Drdla), SLC6A, DAOA, DTNBP1, Dao (Daol), APH-l (alpha and beta), Presenilin (Psenl), Nicastrin,(Ncstn), PEN-2, Nosl, Parpl, Natl, Nat2, HTT, SBMA / SMAX1 / AR, FXN / X25, ATX3, TXN, ATXN2, DMPK, Atrophin-1, Atnl, CBP, VLDLR, Atxn7, and AtxnlO), and eye diseases and disorders (e.g., Aber, Ccl2, Cc2, cp (ceruloplasmin), Timp3, cathepsin-D, Vldlr, Ccr2, CRYAA, CRYA1, CRYBB2, CRYB2, PITX3, BFSP2, CP49, CP47, CRYAA, CRYAI, PAX6, AN2, MGDA, CRYBA1, CRYB1, CRYGC, CRYG3, CCL, LIM2, MP19, CRYGD, CRYG4 , BFSP2, CP49, CP47, HSF4, CTM, HSF4, CTM, MIP, AQPO, CRYAB, CRYA2, CTPP2, CRYBB1, CRYGD, CR YG4, CRYBB2, CRYB2, CRYGC, CRYG3, CCL, CRYAA, CRYAI, GJA8, CX50, CAE1, GJA3, CX46, CZP3, CAE 3, CCM1, CAM, KRIT1, APOA1, TGFBI, CSD2, CDGG1, CSD, BIGH3, CDG2, TACSTD2, TROP2, M1SI, VSX1, RINX, PPCD, PPD, KTCN, COL8A2, FECD, PPCD2, PIP5K3, CFD, KERA, CNA2, MYOC, TIGR, GLCIA, JO AG, GPOA, OPTN, GLC1E, FIP2, HYPL, NRP, CYP1BI, GLC3A, OPA1, NTG, NPG, CYP1BI, GLC3A, CRB1, RP12, CRX, CORD2, CRD, RPGRIPI, LCA6, CORD9, RPE65, RP20, AIPL1, LCA4, GUCY2D, GUC2D, LCA1, CORD6, RDH12, LCA3, ELOVL4, ADMD, STGD2, STGD3, RDS, RP7, PRPH2, PRPH, AVMD, AOFMD, and VMD2.

[0086] In some embodiments, the cargo or payload is or encodes a factor that can affect cell differentiation. As non-limiting examples, expression of one or more of Oct4, Klf4, Sox2, c-Myc, L-Myc, dominant-negative p53, Nanog, Glisl, Lin28, TFIID, mir-302 / 367, or other miRNAs can cause cells to become induced pluripotent stem (iPS) cells.

[0087] In some embodiments, the cargo or payload is or encodes a factor for transdifferentiating cells. Non-limiting examples of factors include one or more of GATA4, Tbx5, Mef2C, Myocd, Hand2, SRF, Mespl, SMARCD3 for cardiomyocytes; Ascii, Nurrl, LmxlA, Bm2, Mytll, NeuroDl, FoxA2 for neuronal cells; and Hnf4a, Foxal, Foxa2 or Foxa3 for liver cells.

[0088] Polypeptides, Proteins and Peptides The originator and benchmark constructs of the present disclosure may comprise, encode, or be conjugated to a cargo or payload that is a polypeptide, protein, or peptide. As used herein, the term "polypeptide" generally refers to a polymer of amino acids linked by peptide bonds and encompasses "proteins" and "peptides." For purposes of this disclosure, polypeptides include all polypeptides, proteins, and / or peptides known in the art. Non-limiting categories of polypeptides include antigens, antibodies, antibody fragments, cytokines, peptides, hormones, enzymes, oxidants, antioxidants, synthetic polypeptides, and chimeric polypeptides.

[0089] As used herein, the term "peptide" generally refers to a shorter polypeptide of about 50 amino acids or less. A peptide having only two amino acids may be referred to as a "dipeptide." A peptide having only three amino acids may be referred to as a "tripeptide." A polypeptide generally refers to a polypeptide having about 4 to about 50 amino acids. A peptide may be obtained via any method known to one of skill in the art. In some embodiments, a peptide may be expressed in culture. In some embodiments, a peptide may be obtained via chemical synthesis (e.g., solid phase peptide synthesis).

[0090] In some embodiments, the originator and benchmark constructs of the present disclosure may contain, encode, or be conjugated to a cargo or payload that is a simple protein that upon hydrolysis produces amino acids and optionally small carbohydrate compounds. Non-limiting examples of simple proteins include albumins, albuminoids, globulins, glutelins, histones, and protamines.

[0091] In some embodiments, the originator constructs and benchmark constructs of the present disclosure may comprise, encode, or be conjugated to a cargo or payload that is a conjugated protein, which may be a simple protein associated with a non-protein. Non-limiting examples of conjugated proteins include glycoproteins, hemoglobins, lecithoproteins, nucleoproteins, and phosphoproteins.

[0092] In some embodiments, the originator and benchmark constructs of the present disclosure may comprise, encode, or be conjugated to a cargo or payload that is a derived protein, which is a protein derived from a simple or conjugated protein by chemical or physical means. Non-limiting examples of derived proteins include denatured proteins and peptides.

[0093] In some embodiments, the polypeptide, protein or peptide may be unmodified.

[0094] In some embodiments, the polypeptide, protein or peptide may be modified. The types of modification include, but are not limited to, phosphorylation, glycosylation, acetylation, ubiquitination / sumoylation, methylation, palmitoylation, quinone, amidation, myristoylation, pyrrolidone carboxylic acid, hydroxylation, phosphopantetheine, prenylation, GPI anchoring, oxidation, ADP-ribosylation, sulfation, S-nitrosylation, citrullination, nitration, gamma-carboxyglutamic acid, formylation, hypusine, topaquinone (TPQ), bromination, lysine topaquinone (LTQ), tryptophan tryptophylquinone (TTQ), iodination, and cysteine ​​tryptophylquinone (CTQ). In some aspects, the polypeptide, protein or peptide may be modified by post-transcriptional modifications that may affect its structure, subcellular localization, and / or function.

[0095] In some embodiments, polypeptide, protein or peptide can be modified using phosphorylation.The phosphorylation of serine, threonine or tyrosine residue, or the addition of phosphate group, is one of the most common forms of protein modification.Protein phosphorylation plays an important role in fine-tuning the signal in intracellular signal transduction cascade.

[0096] In some embodiments, polypeptides, proteins or peptides can be modified using ubiquitination, which is the covalent attachment of ubiquitin to a target protein. Ubiquitination-mediated protein turnover has been shown to play a role in inducing cell cycle progression and in proteolysis-independent intracellular signaling pathways.

[0097] In some embodiments, polypeptides, proteins or peptides can be modified using acetylation and methylation, which can play a role in regulating gene expression. As a non-limiting example, acetylation and methylation can mediate the formation of chromatin domains (e.g., euchromatin and heterochromatin), which can have an impact on mediating gene silencing.

[0098] In some embodiments, polypeptides, proteins or peptides can be modified using glycosylation.Glycosylation is the attachment of one of many glycan groups, which occurs in about half of all proteins and is a modification that plays a role in biological processes, including but not limited to embryo development, cell division, and the control of protein structure.The two main types of protein glycosylation are N-glycosylation and O-glycosylation.In the case of N-glycosylation, glycans are attached to asparagine, and in the case of O-glycosylation, glycans are attached to serine or threonine.

[0099] In some embodiments, a polypeptide, protein or peptide may be modified using sumoylation, which is the addition of SUMO (small ubiquitin-like modifier) ​​to a protein, a post-translational modification similar to ubiquitination.

[0100] antibody As used herein, the term "antibody" is referred to in the broadest sense and specifically encompasses various embodiments including, but not limited to, monoclonal antibodies, polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies formed from at least two intact antibodies), and antibody fragments (e.g., diabodies) so long as they exhibit the desired biological activity (e.g., are "functional"). Antibodies are primarily amino acid-based molecules that are monomeric or multimeric polypeptides that contain at least one amino acid region derived from a known or parent antibody sequence and at least one amino acid region derived from a non-antibody sequence. Antibodies may contain one or more modifications, including, but not limited to, the addition of a sugar moiety, a fluorescent moiety, a chemical tag, and the like. For purposes herein, an "antibody" may include heavy and light variable domains as well as an Fc region.

[0101] The cargo or payload may comprise or encode one or more polypeptides that form a functional antibody.

[0102] In some embodiments, the cargo or payload may comprise or code for a polypeptide that forms or functions as any antibody, including, but not limited to, antibodies known in the art and / or commercially available antibodies that may be for therapeutic, diagnostic, or research purposes, and may comprise or code for a fragment or antibody of such an antibody, including, but not limited to, a variable domain or complementarity determining region (CDR).

[0103] As used herein, the term "native antibody" typically refers to a heterotetrameric glycoprotein of about 150,000 daltons composed of two identical light (L) chains and two identical heavy (H) chains. The genes encoding antibody heavy and light chains are known, and the segments that compose each have been well characterized and described (Matsuda, F. et al., 1998. The Journal of Experimental Medicine. 188(11); 2151-62 and Li, A. et al., 2004. Blood. 103(12): 4602-9, the contents of each of which are incorporated herein by reference in their entirety). Each light chain is linked to a heavy chain by one covalent disulfide bond, while the number of disulfide linkages varies among the heavy chains of different immunoglobulin isotypes. Each heavy and light chain also has regularly spaced intrachain disulfide bridges. Each heavy chain contains a variable domain (V) at one end and a variable domain (V) at the other end. H ) followed by several constant domains. Each light chain has a variable domain (V L ) at its other end and a constant domain; the constant domain of the light chain is aligned with the first constant domain of the heavy chain, and the variable domain of the light chain is aligned with the variable domain of the heavy chain. As used herein, the term "light chain" refers to a component of an antibody from any vertebrate species that is assigned to one of two clearly distinct types, called kappa and lambda, based on the amino acid sequence of the constant domain. Depending on the amino acid sequence of the constant domain of their heavy chain, antibodies can be assigned to different classes. There are five major classes of intact antibodies: IgA, IgD, IgE, IgG, and IgM, and several of these can be further divided into subclasses (isotypes), e.g., IgG1, IgG2, IgG3, IgG4, IgA, and IgA2.

[0104] As used herein, the term "variable domain" refers to a particular antibody domain found in both antibody heavy and light chains that varies widely in sequence between antibodies and is used in the binding and specificity of each particular antibody to its particular antigen. Variable domains include hypervariable regions. As used herein, the term "hypervariable region" refers to a region within a variable domain that contains the amino acid residues responsible for antigen binding. The amino acids present within the hypervariable region determine the structure of the complementarity determining region (CDR), which becomes part of the antigen binding site of the antibody. As used herein, the term "CDR" refers to the region of an antibody that contains the structure complementary to its target antigen or epitope. The other part of the variable domain that does not interact with the antigen is referred to as the framework (FW) region. The antigen binding site (also known as the antigen integration site or paratope) contains the amino acid residues necessary to interact with a particular antigen. The exact residues that make up the antigen-binding site are typically revealed by co-crystallography with bound antigen, but computational assessment can also be used based on comparison with other antibodies (Strohl, WR Therapeutic Antibody Engineering. Woodhead Publishing, Philadelphia PA. 2012. Ch. 3, p. 47-54, the contents of which are incorporated herein by reference in their entirety).Determining the residues that constitute the CDRs includes the methods of Kabat [Wu, TT et al., 1970, JEM, 132(2):211-50 and Johnson, G. et al., 2000, Nucleic Acids Res. 28(1):214-8 (the contents of each of which are incorporated herein by reference in their entirety)], Chothia [Chothia and Lesk, J. Mol. Biol. 196, 901(1987), Chothia et al., Nature 342, 877(1989) and Al-Lazikani, B. et al., 1997, J. Mol. Biol. 273(4):927-48 (the contents of each of which are incorporated herein by reference in their entirety)], Lefranc (Lefranc, MP et al., 2005, Immunome Res. 1:3) and Honegger (Honegger, A. and These may include the use of numbering schemes including, but not limited to, those taught by Pluckthun, A. 2001. J. Mol. Biol. 309(3):657-70, the contents of which are incorporated herein by reference in their entirety.

[0105] V H and V L Each domain has three CDRs. L The CDRs of are referred to herein as CDR-L1, CDR-L2 and CDR-L3, in the order in which they occur when moving from N-terminus to C-terminus along the variable domain polypeptide. H The CDRs are referred to herein as CDR-H1, CDR-H2, and CDR-H3 in the order in which they occur when moving from N-terminus to C-terminus along the variable domain polypeptide. Each of the CDRs has a preferred canonical structure, except for CDR-H3, which contains an amino acid sequence that can be highly variable in sequence and length between antibodies resulting in diverse three-dimensional structures in the antigen-binding domain. In some cases, CDR-H3 can be analyzed among a panel of related antibodies to assess antibody diversity.

[0106] Various methods of determining CDR sequences are known in the art and can be applied to known antibody sequences. The system described by Kabat, also referred to as "numbered according to Kabat," "Kabat numbering," "Kabat definition," and "Kabat designation," provides an unambiguous residue numbering system that is applicable to any variable domain of an antibody and provides precise residue boundaries that define the three CDRs in each chain. (Kabat et al., Sequences of Proteins of Immunological Interest, National Institutes of Health, Bethesda, Md. (1987) and (1991), the contents of which are incorporated by reference in their entirety. The Kabat CDRs include approximately residues 24-34 (CDR1), 50-56 (CDR2), and 89-97 (CDR3) in the light chain variable domain, and 31-35 (CDR1), 50-65 (CDR2), and 95-102 (CDR3) in the heavy chain variable domain. Chothia and colleagues discovered that certain subportions within the Kabat CDRs adopt nearly identical peptide backbone structures, despite a great deal of diversity at the level of amino acid sequence. (Chothia et al. (1987) J. Mol. Biol. 196:901-917; and Chothia et al. (1989) Nature 199:101-112; 342:877-883, the contents of each of which are incorporated herein by reference in their entirety. These CDRs may be referred to as "Chothia CDRs," "Chothia numbering," or "numbered according to Chothia," and include approximately residues 24-34 (CDR1), 50-56 (CDR2), and 89-97 (CDR3) in the light chain variable domain, and 26-32 (CDR1), 52-56 (CDR2), and 95-102 (CDR3) in the heavy chain variable domain. Mol. Biol. 196:901-917 (1987).The system described by MacCallum, also referred to as "numbered according to MacCallum" or "MacCallum numbering", includes approximately residues 30-36 (CDR1), 46-55 (CDR2) and 89-96 (CDR3) in the light chain variable domain, and 30-35 (CDR1), 47-58 (CDR2) and 93-101 (CDR3) in the heavy chain variable domain. (MacCallum et al. ((1996) J. Mol. Biol. 262(5):732-745), the contents of which are incorporated herein by reference in their entirety. The system described by AbM, also referred to as "numbering according to AbM" or "AbM numbering", includes approximately residues 24-34 (CDR1), 50-56 (CDR2) and 89-97 (CDR3) in the light chain variable domain, and 26-35 (CDR1), 50-58 (CDR2) and 95-102 (CDR3) in the heavy chain variable domain. The IMGT (INTERNATIONAL IMMUNOGENETICS INFORMATION SYSTEM) numbering of the variable regions may also be used, which is the numbering of residues in an immunoglobulin variable heavy or light chain according to the method of the IIMGT (Lefranc, M.-P., "The IMGT unique numbering for immunoglobulins, T cell Receptors and "Ig-like domains", The Immunologist, 7, 132-136 (1999) (incorporated herein by reference in its entirety). As used herein, "IMGT sequence numbering" or "numbered according to IMGT" refers to the numbering of sequences encoding variable regions according to IMGT. For heavy chain variable domains, when numbered according to IMGT, the hypervariable region ranges from amino acid positions 27-38 for CDR1, amino acid positions 56-65 for CDR2, and amino acid positions 105-117 for CDR3. For light chain variable domains, when numbered according to IMGT, the hypervariable region ranges from amino acid positions 27-38 for CDR1, amino acid positions 56-65 for CDR2, and amino acid positions 105-117 for CDR3.

[0107] In some embodiments, the cargo or payload may comprise or encode an antibody generated using methods known in the art, including, but not limited to, immuno- and display technologies (e.g., phage display, yeast display, and ribosome display), hybridoma technology, heavy and light chain variable region cDNA sequences selected from hybridomas or other sources.

[0108] In some embodiments, the cargo or payload may include or code for an antibody established using any naturally occurring or synthetic antigen. As used herein, an "antigen" is an entity that induces or evokes an immune response in an organism. An immune response is characterized by the reaction of cells, tissues and / or organs of an organism to the presence of a foreign entity. Such an immune response typically results in the organism's production of one or more antibodies against the foreign entity, e.g., an antigen or a portion of an antigen. As used herein, "antigen" also refers to a binding partner for a particular antibody or binder in a display library.

[0109] As used herein, the term "monoclonal antibody" refers to an antibody obtained from a population of substantially homogeneous cells (or clones), i.e., the individual antibodies comprising the population are identical and / or bind to the same epitope, except for possible variants that may arise during the generation of the monoclonal antibody (such variants are usually present in small amounts). In contrast to polyclonal antibody preparations, which typically include different antibodies directed against different determinants (epitopes), each monoclonal antibody is directed against a single determinant on the antigen.

[0110] The modifier "monoclonal" indicates the character of the antibody as being obtained from a substantially homogeneous population of antibodies, and is not to be construed as requiring production of the antibody by any particular method. Monoclonal antibodies herein include "chimeric" antibodies (immunoglobulins) in which a portion of the heavy and / or light chain is identical or homologous to corresponding sequences in antibodies from a particular species or belonging to a particular antibody class or subclass, while the remainder of the chain(s) is identical or homologous to corresponding sequences in antibodies from another species or belonging to another antibody class or subclass, and fragments of such antibodies.

[0111] As used herein, the term "humanized antibody" refers to a chimeric antibody that contains minimal portions from one or more non-human (e.g., murine) antibody source(s), with the remainder being derived from one or more human immunoglobulin sources. In most cases, humanized antibodies are human immunoglobulins (recipient antibody) in which residues from a hypervariable region from the recipient antibody are replaced by residues from a hypervariable region from an antibody of a non-human species (donor antibody), e.g., mouse, rat, rabbit or non-human primate, having the desired specificity, affinity, and / or capacity.

[0112] In some embodiments, the cargo or payload may include or code for an antibody mimic. As used herein, the term "antibody mimic" refers to any molecule that mimics the function or effect of an antibody and binds specifically and with high affinity to their molecular target. In some embodiments, the antibody mimic may be a monobody designed to incorporate a fibronectin type III domain (Fn3) as a protein scaffold. In some embodiments, the antibody mimic may be any known in the art, including, but not limited to, affibody molecules, affilins, affitins, anticalins, avimers, sentinels, DARPINs™, finomers, Kunitz domains, and domain peptides. In other embodiments, the antibody mimic may include one or more non-peptide regions.

[0113] Antibody Fragments and Variants In some embodiments, the cargo or payload may comprise or encode an antibody fragment that comprises an antigen-binding region from a full-length antibody. Non-limiting examples of antibody fragments include Fab, Fab', F(ab') 2 Antibody fragments include F(ab') fragments, Fv fragments, diabodies, linear antibodies, single-chain antibody molecules, and multispecific antibodies formed from antibody fragments. Papain digestion of antibodies produces two identical antigen-binding fragments called "Fab" fragments, each with a single antigen-binding site. A residual "Fc" fragment is also produced, the name reflecting its ability to crystallize readily. Pepsin treatment produces an F(ab') fragment that has two antigen-binding sites and is still capable of cross-linking antigen. 2 The compounds and / or compositions of the disclosure may include one or more of these fragments.

[0114] In some embodiments, the Fc region may be a modified Fc region, which may have a single amino acid substitution when compared to the corresponding sequence for a wild-type Fc region, which single amino acid substitution generates an Fc region with preferred properties relative to those of the wild-type Fc region. Non-limiting examples of Fc properties that may be altered by a single amino acid substitution include binding properties or response to pH conditions.

[0115] As used herein, the term "Fv" refers to an antibody fragment that contains the minimum fragment on an antibody required to form a complete antigen-binding site. These regions consist of a dimer of one heavy and one light chain variable domain in tight non-covalent association. Fv fragments can be generated by proteolytic cleavage, but are generally unstable. Recombinant methods are known in the art to generate stable Fv fragments, typically via insertion of a flexible linker between the light and heavy chain variable domains to form a single chain Fv (scFv), or via introduction of a disulfide bridge between the heavy and light chain variable domains.

[0116] As used herein, the term "single chain Fv" or "scFv" refers to a V Hand V L It refers to a fusion protein of antibody domains, in which these domains are linked together into a single polypeptide chain by a flexible peptide linker. In some embodiments, the Fv polypeptide linker allows the scFv to form the desired structure for antigen binding. In some embodiments, scFvs are utilized in conjunction with phage display, yeast display or other display methods, where they are expressed in association with surface members (e.g., phage coat proteins) and can be used in identifying high affinity peptides for a given antigen.

[0117] As used herein, the term "antibody variant" refers to modified antibodies or biomolecules (e.g., antibody mimetics) that resemble (relative to) a native or starting antibody in structure and / or function. Antibody variants may be altered in their amino acid sequence, composition, or structure when compared to a native antibody. Antibody variants include those with altered isotypes (e.g., IgA, IgD, IgE, IgG 1 , IgG 2 , IgG 3 , IgG 4 These may include, but are not limited to, antibodies having a specific IgM, IgA, IgB, IgC, IgD, IgE, IgF, IgH, IgIgIgM, IgF, IgH, IgIgM, IgF, IgH, IgIgM, IgH, I ...

[0118] multispecific antibodies In some embodiments, the cargo or payload may be or code for an antibody that binds multiple epitopes. As used herein, the term "multibody" or "multispecific antibody" refers to an antibody in which two or more variable regions bind to different epitopes. The epitopes may be on the same or different targets. In certain embodiments, the multispecific antibody is a "bispecific antibody" that recognizes two different epitopes on the same or different antigens.

[0119] In some embodiments, multispecific antibodies can be prepared by the methods used by BIOATLA® and described in International Patent Publication WO201109726, the contents of which are incorporated herein by reference in their entirety. First, a library of homologous naturally occurring antibodies is generated by any method known in the art (i.e., mammalian cell surface display) and then screened by FACSAria or another screening method for multispecific antibodies that specifically bind to two or more target antigens. In some embodiments, the identified multispecific antibodies are further evolved by any method known in the art to generate a set of modified multispecific antibodies. These modified multispecific antibodies are screened for binding to the target antigens. In some embodiments, the multispecific antibodies can be further optimized by screening the evolved modified multispecific antibodies for optimized or desired characteristics.

[0120] In some embodiments, multispecific antibodies can be prepared by the methods used by BIOATLA® and described in US Publication No. US20150252119, the contents of which are incorporated herein by reference in their entirety. In one approach, the variable domains of two parent antibodies, where the parent antibodies are monoclonal antibodies, are evolved using any method known in the art in a manner that allows a single light chain to functionally complement the heavy chains of two different parent antibodies. Another approach requires evolving the heavy chain of a single parent antibody to recognize a second target antigen. A third approach involves evolving the light chain of a parent antibody to recognize a second target antigen. Methods for polypeptide evolution are described in International Publication WO2012009026, the contents of which are incorporated herein by reference in their entirety, including, as non-limiting examples, global positional evolution (CPE), combinatorial protein synthesis (CPS), global positional insertion (CPI), global positional deletion (CPD), or any combination thereof. The Fc regions of the multispecific antibodies described in U.S. Publication No. US20150252119 can be generated using the knobs-in-holes approach, or any other method that allows the Fc domains to form heterodimers. The resulting multispecific antibodies can be further evolved for improved characteristics or properties, such as binding affinity to target antigens.

[0121] bispecific antibody In some embodiments, the cargo or payload may be or code for a bispecific antibody. As used herein, the term "bispecific antibody" refers to an antibody capable of binding to two different antigens. Such an antibody typically comprises regions from at least two different antibodies. Such an antibody typically comprises antigen-binding regions from at least two different antibodies. For example, a bispecific monoclonal antibody (BsMAb, BsAb) is an artificial protein composed of fragments of two different monoclonal antibodies, thus allowing the BsAb to bind to two different types of antigens.

[0122] In some cases, the cargo or payload may be or encode a bispecific antibody that contains antigen-binding regions from two different anti-tau antibodies. For example, such a bispecific antibody may contain binding regions from two different antibodies.

[0123] Bispecific antibody frameworks may include any of those described in Riethmuller, G., 2012. Cancer Immunity. 12:12-18; Marvin, JSet al., 2005. Acta Pharmacologica Sinica. 26(6):649-58; and Schaefer, W. et al., 2011. PNAS. 108(27):11187-92, the contents of each of which are incorporated herein by reference in their entirety.

[0124] A new generation of BsMAbs has been developed, called "trifunctional bispecific" antibodies, which consist of two heavy and two light chains, each derived from two different antibodies, with two Fab regions (arms) directed against two antigens and an Fc region (foot) that includes the two heavy chains and forms the third binding site.

[0125] Of the two paratopes that form the top of the variable domain of a bispecific antibody, one can be directed against the target antigen, and the other against the T-lymphocyte antigen-like CD3. In the case of a trifunctional antibody, the Fc region can additionally bind to cells expressing Fc receptors, such as macrophages, natural killer (NK) cells or dendritic cells. In essence, the targeted cell is connected to one or two cells of the immune system, which then destroy it.

[0126] Other types of bispecific antibodies have been designed to overcome certain problems caused by cytokine release, such as short half-life, immunogenicity and side effects. They include chemically linked Fabs consisting of only the Fab region, as well as various types of bivalent and trivalent single chain variable fragments (scFvs) (fusion proteins that mimic the variable domains of two antibodies). The most developed of these newer formats are bispecific T cell engagers (BiTEs) and mAb2 antibodies engineered to contain an Fcab antigen binding fragment instead of the Fc constant region.

[0127] Using molecular genetics, two scFvs can be engineered in tandem into a single polypeptide separated by a linker domain, called a "tandem scFv" (tascFv). TascFvs have been found to be poorly soluble and require refolding when produced in bacteria, or they can be manufactured in mammalian cell culture systems that circumvent the refolding requirement but may result in poor yields. Construction of a tascFv with genes for two different scFvs generates a "bispecific single-chain variable fragment" (bis-scFv). Only two tascFvs are in clinical development by commercial companies; both are bispecific agents under ongoing early-stage development by Micromet for cancer indications and are described as "bispecific T-cell engagers (BiTEs)". Blinatumomab is an anti-CD19 / anti-CD3 bispecific tascFv in phase 2 that promotes T-cell responses against B-cell non-Hodgkin's lymphoma. MT110 is an anti-EP-CAM / anti-CD3 bispecific tascFv that promotes T cell responses against solid tumors, in Phase 1. A bispecific tetravalent "TandAb" is also being investigated by Affimed.

[0128] In some embodiments, the cargo or payload may be or code for an antibody that contains a single antigen-binding domain. These molecules are extremely small, with molecular weights roughly one-tenth that observed in full-sized mAbs. Additional antibodies include those that contain the antigen-binding variable heavy chain region (VH) of heavy chain antibodies found in camels and llamas that lack light chains. HH ) may be included.

[0129] PCT Publication WO2014144573 from Memorial Sloan-Kettering Cancer Center, the contents of which are incorporated herein by reference in their entirety, discloses and claims multimerization techniques for generating dimeric polyspecific binding agents (e.g., fusion proteins comprising antibody components) that have improved properties over polyspecific binding agents that do not possess the ability to dimerize.

[0130] In some cases, the cargo or payload may be or encode a tetravalent bispecific antibody (TetBiAb, as disclosed and claimed in PCT Publication WO2014144357, the contents of which are incorporated herein in their entirety). TetBiAb features a second pair of Fab fragments with a second antigen specificity attached to the C-terminus of the antibody, thus providing a molecule that is bivalent for each of the two antigen specificities. Tetravalent antibodies are generated by genetic engineering methods by covalently linking an antibody heavy chain to a Fab light chain that associates with its cognate co-expressed Fab heavy chain.

[0131] In some embodiments, the cargo or payload may be or encode a biosynthetic antibody as described in U.S. Patent No. 5,091,513, the contents of which are incorporated herein by reference in their entirety. Such antibodies may include one or more sequences of amino acids that constitute a region that behaves as a biosynthetic antibody binding site (BABS). The sites include 1) non-covalently associated or disulfide-linked synthetic V H and V L Dimer, 2)V H -VL or V L -V H It is a single strand, V H and V L are linked by a polypeptide linker, or 3) each V H or V L The binding domain comprises a CDR and FR region linked together, which may be derived from separate immunoglobulins. Biosynthetic antibodies may also include other polypeptide sequences that function, for example, as enzymes, toxins, binding sites, or sites of attachment to immobilization media or radioactive atoms. Methods are disclosed for generating biosynthetic antibodies, for designing BABS with any specificity that can be elicited by in vivo generation of antibodies, and for generating analogs thereof.

[0132] In some embodiments, the cargo or payload may be or may code for an antibody having an antibody acceptor framework as taught in US Patent No. 8,399,625. Such an antibody acceptor framework may be a particularly well-suited acceptor CDR from the antibody of interest. In some cases, CDRs from an anti-tau antibody known in the art or developed according to the methods presented herein may be used.

[0133] miniaturized antibodies In some embodiments, the cargo or payload may be or code for a "miniaturized" antibody. The best examples of miniaturization of mAbs are the small modular immunopharmaceuticals (SMIPs) from Trubion Pharmaceuticals. These molecules, which may be monovalent or bivalent, are composed of one V L , 1 V HThey are recombinant single-chain molecules that contain an antigen-binding domain and one or two constant "effector" domains, all connected by a linker domain. Conceivably, such molecules could offer the advantage of increased tissue or tumor penetration sought by fragments, while maintaining the immune effector functions conferred by the constant domains. At least three "miniaturized" SMIPs are in clinical development. TRU-015, an anti-CD20 SMIP developed in collaboration with Wyeth, is the most advanced program, having progressed to phase 2 for rheumatoid arthritis (RA). Earlier attempts in systemic lupus erythematosus (SLE) and B-cell lymphoma were eventually discontinued. Trubion and Facet Biotechnology are collaborating in the development of TRU-016, an anti-CD37 SMIP for the treatment of CLL and other lymphoid neoplasms, a project that has reached phase 2. Wyeth has licensed the anti-CD20 SMIP SBI-087 for the treatment of autoimmune diseases, including RA, SLE, and potentially multiple sclerosis, but these projects remain in the earliest stages of clinical trials.

[0134] Diamond Body In some embodiments, the cargo or payload may be or may code for a diabody. As used herein, the term "diabody" refers to a small antibody fragment that has two antigen-binding sites. A diabody is a fragment of an antibody that contains a light chain variable domain V on the same polypeptide chain. L The heavy chain variable domain V H By using a linker that is too short to allow pairing between the two domains on the same chain, the domains are forced to pair with the complementary domains of another chain and create two antigen-binding sites.

[0135] Diabodies are functional bispecific single-chain antibodies (bscAbs). These bivalent antigen-binding molecules are composed of non-covalent dimers of scFvs and can be produced in mammalian cells using recombinant methods (see, for example, Mack et al., Proc. Natl. Acad. Sci., 92:7021-7025, 1995). Several diabodies are in clinical development. An iodine-123 labeled diabody version of the anti-CEA chimeric antibody cT84.66 is being evaluated for preoperative immunoscintigraphy detection of colon cancer in a study funded by the Beckman Research Institute of the City of Hope (Clinicaltrials.gov NCT00647153).

[0136] Unibody In some embodiments, the cargo or payload may be or code for a "unibody" in which the hinge region has been removed from an IgG4 molecule. Although IgG4 molecules are unstable and may exchange light-heavy chain heterodimers with each other, the deletion of the hinge region completely prevents heavy-heavy chain pairing, leaving highly specific monovalent light / heavy heterodimers while retaining the Fc region to ensure stability and half-life in vivo. Since IgG4 interacts poorly with FcRs and monovalent unibodies cannot promote intracellular signaling complex formation, this configuration may minimize the risk of immune activation or oncogenic growth. However, these arguments are largely supported by laboratory evidence rather than clinical. Other antibodies may be "miniaturized" antibodies, which are reduced 100 kDa antibodies.

[0137] Intrabody In some embodiments, the cargo or payload may be or may code for an intrabody. The term "intrabody" refers to a form of antibody that is not secreted from the cell in which it is produced, but instead targets one or more intracellular proteins. Intrabodies may be used to affect a number of cellular processes, including, but not limited to, intracellular trafficking, transcription, translation, metabolic processes, growth signaling, and cell division. In some embodiments, the methods of the disclosure may include intrabody-based therapy. In some such embodiments, the variable domain sequences and / or CDR sequences disclosed herein may be incorporated into one or more constructs for intrabody-based therapy. For example, intrabodies may target one or more glycosylated intracellular proteins or modulate the interaction between one or more glycosylated intracellular proteins and alternative proteins.

[0138] More than 20 years ago, intracellular antibodies against intracellular targets were first described (Biocca, Neuberger and Cattaneo EMBO J. 9:101-108, 1990, the contents of which are incorporated herein by reference in their entirety). Intracellular expression of intrabodies in different compartments of mammalian cells allows blocking or modulating the function of endogenous molecules (Biocca, et al., EMBO J. 9:101-108, 1990; Colby et al., Proc. Natl. Acad. Sci. USA 101:17616-21, 2004, the contents of which are incorporated herein by reference in their entirety). Intrabodies can alter protein folding, protein-protein, protein-DNA, protein-RNA interactions and protein modifications. They can induce phenotypic knockouts and act as neutralizing agents by direct binding to the target antigen, by bypassing its intracellular trafficking or by inhibiting its association with binding partners. They are primarily used as research tools and are emerging as therapeutic molecules for the treatment of human diseases such as viral pathologies, cancer and misfolding diseases. The burgeoning biomarket of recombinant antibodies offers intrabodies with improved binding specificity, stability and solubility, along with lower immunogenicity, for their use in therapy.

[0139] In some embodiments, intrabodies have advantages over interfering RNA (iRNA); for example, iRNA has been shown to exert multiple non-specific effects, whereas intrabodies have been shown to have high specificity and affinity for target antigens. Furthermore, as proteins, intrabodies possess a much longer half-life of activity than iRNA. Thus, if the intracellular target molecule has a long half-life of activity, iRNA-mediated gene silencing may be slow to produce an effect, whereas the effect of intrabody expression may be almost instantaneous. Finally, intrabodies can be designed to block certain binding interactions of specific target molecules while sparing others.

[0140] Intrabodies are often single chain variable fragments (scFv) that are expressed from recombinant nucleic acid molecules and engineered to be retained intracellularly (e.g., in the cytoplasm, endoplasmic reticulum, or periplasm). Intrabodies can be used, for example, to eliminate the function of a protein to which they bind. Expression of intrabodies can also be controlled through the use of inducible promoters in nucleic acid expression vectors that contain intrabodies. Intrabodies can be prepared using methods known in the art, e.g., Marasco et al., 1993 Proc. Natl. Acad. Sci. USA, 90:7889-7893; Chen et al., 1994, Hum. Gene Ther. 5:595-601; Chen et al., 1994, Proc. Natl. Acad. Sci. USA, 91:5932-5936; Maciejewski et al., 1995, Nature Med., 1:667-673; Marasco, 1995, Immunotech, 1:1-19; Mhashilkar, et al., 1995, EMBO J. 14:1542-51; Chen et al., 1996, Hum. Gene Therap., 7:1515-1525; Marasco, Gene Ther.4:11-15,1997;Rondon and Marasco,1997,Annu.Rev.Microbiol.51:257-283;Cohen, et al.,1998,Oncogene 17:2445-56;Proba et al.,1998,J.Mol.Biol.275:245-253;Cohen et al. al.,1998,Oncogene 17:2445-2456;Hassanzadeh,et al.,1998,FEBS Lett.437:81-6;Richardson et al.,1998,Gene Ther.5:635-44;Ohage and Steipe,1999,J.Mol.Biol.291:1119-1128;Ohage et al. al.,1999,J.Mol.Biol.291:1129-1134;Wirtz and Steipe,1999,Protein Sci.8:2245-2250;Zhu et al.,1999,J.Immunol.Methods 231:207-222; Arafat et al., 2000, Cancer Gene Ther. 7:1250-6; der Maur et al., 2002, J. Biol. Chem. 277:45075-85; Mhashilkar et al., 2002, Gene Ther. 9:307-19; and Wheeler et al., 2003, FASEB J. 17:1733-5 and references cited therein). In particular, CCR5 intrabodies have been generated by Steinberger et al., 2000, Proc. Natl. Acad. Sci. USA 97:805-810). See generally Marasco, WA, 1998, "Intrabodies: Basic Research and Clinical Gene Therapy Applications" Springer: New York; and for a review of scFvs, see Pluckthun in "The Pharmacology of Monoclonal Antibodies," 1994, vol. 113, Rosenburg and Moore eds. Springer-Verlag, New York, pp. 269-315, the contents of each of which are each incorporated by reference in their entirety.

[0141] Sequences from a donor antibody can be used to generate intrabodies. Intrabodies are often synthesized intracellularly as single domain fragments, e.g., isolated V H and V LThey are recombinantly expressed as domains or as single chain variable fragment (scFv) antibodies. For example, intrabodies are often expressed as a single polypeptide to form single chain antibodies comprising heavy and light chain variable domains connected by a flexible linker polypeptide. Intrabodies typically lack disulfide bonds and are capable of regulating the expression or activity of target genes through their specific binding activity. Single chain antibodies can also be expressed as a single chain variable region fragment connected to a light chain constant region.

[0142] As known in the art, intrabodies can be engineered into recombinant polynucleotide vectors to encode intracellular trafficking signals at their N- or C-terminus to allow expression at high concentrations in the intracellular compartment where the target protein is located. For example, intrabodies targeted to the endoplasmic reticulum (ER) are engineered to incorporate a leader peptide and, optionally, a C-terminal ER retention signal. Intrabodies intended to exert activity in the nucleus are engineered to include a nuclear localization signal. A lipid moiety is attached to the intrabody to tether it to the cytosolic side of the cell membrane. Intrabodies can also be targeted to exert a function in the cytosol. For example, cytosolic intrabodies are used to sequester factors in the cytosol, thereby preventing them from being transported to their natural cellular destination.

[0143] Intrabody expression presents certain technical challenges: in particular, protein structural folding and structural stability of newly synthesized intrabodies within cells are affected by the reducing conditions of the intracellular environment.

[0144] The intrabodies of the present disclosure may be promising therapeutic agents for the treatment of misfolding diseases including tauopathies, prion diseases, Alzheimer's, Parkinson's, and Huntington's diseases, due to the use of their virtually unlimited ability to specifically recognize different structures of proteins, including pathological isoforms, and because they can be targeted to potential sites of aggregation (both intracellular and extracellular sites). These molecules may act as neutralizing agents against amyloidogenic proteins by preventing their aggregation, and / or as molecular shunters of intracellular traffic by rerouting proteins from their potential aggregation sites.

[0145] Maxi Body In some embodiments, the cargo or payload may be or encode an IgG maxibody (Fc (a bivalent scFv fused to the amino terminus of the CH2-CH3 domains).

[0146] Chimeric antigen receptors (CARs) In some embodiments, the cargo or payload may be or may encode a chimeric antigen receptor (CAR) that, when transduced into immune cells (e.g., T cells and NK cells), can redirect the immune cell to a target (e.g., tumor cell) that expresses a molecule recognized by the extracellular targeting moiety of the CAR.

[0147] As used herein, the term "chimeric antigen receptor (CAR)" refers to a synthetic receptor that mimics the TCR on the surface of a T cell. Typically, a CAR is composed of an extracellular targeting domain, a transmembrane domain / region, and an intracellular signaling / activation domain. In a standard CAR receptor, the components: extracellular targeting domain, transmembrane domain, and intracellular signaling / activation domain are linearly assembled as a single fusion protein. The extracellular region contains a targeting domain / site (e.g., scFv) that recognizes a specific tumor antigen or other tumor cell-surface molecule. The intracellular region may contain a signaling domain of the TCR complex (e.g., the signal region of CD3ζ), and / or one or more costimulatory signaling domains, such as those from CD28, 4-1BB (CD137), and OX-40 (CD134). For example, while "first generation CARs" only have a CD3ζ signaling domain, in an effort to enhance T cell persistence and proliferation, a costimulatory intracellular domain has been added, resulting in the emergence of second generation CARs with one costimulatory signaling domain in addition to the CD3ζ signaling domain, and third generation CARs with two or more costimulatory signaling domains in addition to the CD3ζ signaling domain. When expressed by a T cell, a CAR results in a T cell with an antigen specificity determined by the extracellular targeting site of the CAR. In some embodiments, one or more elements, such as homing and suicide genes, may be added to develop a more potent and safer CAR structure (so-called fourth generation CARs).

[0148] In some embodiments, the extracellular targeting domain is connected to the intracellular signaling domain through a hinge (also referred to as a space domain or spacer) and a transmembrane region. The hinge connects the extracellular targeting domain to the transmembrane domain that crosses the cell membrane and connects to the intracellular signaling domain. Depending on the size of the target protein that the targeting moiety binds to, as well as the size and affinity of the targeting domain itself, the hinge may need to be altered to optimize the efficacy of the CAR transduced cells against cancer cells. Upon recognition and binding of the targeting moiety to the target cell, the intracellular signaling domain provides an activation signal to the CAR T cell, which is further amplified by a "second signal" from one or more intracellular costimulatory domains. Once activated, the CAR T cell can destroy the target cell.

[0149] In some embodiments, CARs can be split into two parts, each part linked to a dimerization domain, whereby dimerization-inducing inputs promote the assembly of an intact functional receptor. Wu and Lim report that the extracellular CD19-binding domain and intracellular signaling elements are separated, and that the FKBP domain and FRB domain heterodimerize in the presence of the rapamycin analog AP21967. * reported a split-CAR linked to the (T2089L mutant of FKBP-rapamycin binding) domain. The split receptor assembles with specific antigen binding in the presence of AP21967 and activates T cells (Wu et al., Science, 2015, 625(6258):aab4077, the contents of which are incorporated herein by reference in their entirety).

[0150] In some embodiments, the CAR can be designed as an inducible CAR with the incorporation of a Tet-On inducible system into the CD19 CAR construct. The CD19 CAR is activated only in the presence of doxycycline (Dox). Sakemura reported that Tet-CD19CAR T cells in the presence of Dox have comparable cytotoxicity against CD19+ cell lines and comparable cytokine production and proliferation upon CD19 stimulation compared to conventional CD19CAR T cells (Sakemura et al., Cancer Immuno. Res., 2016, Jun 21, Epub, the contents of which are incorporated herein by reference in their entirety). The dual system provides greater flexibility for turning on and off CAR expression in transduced T cells.

[0151] In some embodiments, the cargo or payload may be or code for a first generation CAR, or a second generation CAR, or a third generation CAR, or a fourth generation CAR. In some embodiments, the cargo or payload may be or code for a complete CAR construct consisting of an extracellular domain, a hinge and transmembrane domain, and an intracellular signaling region. In other embodiments, the cargo or payload may be or code for a component of a complete CAR construct including an extracellular targeting site, a hinge region, a transmembrane domain, an intracellular signaling domain, one or more costimulatory domains, and other additional elements that improve CAR structure and functionality, including, but not limited to, a leader sequence, a homing element, and a safety switch, or a combination of such components.

[0152] In some embodiments, the cargo or payload may be or encode a tunable CAR. A reversible on-off switch mechanism allows for management of acute toxicity caused by excessive CAR-T cell proliferation. Ligand-conferred CAR control may be effective in counteracting tumor escape induced by antigen loss, avoiding functional exhaustion caused by tonic signaling due to chronic antigen exposure, and improving persistence of CAR-expressing cells in vivo. Tunable CARs may be utilized to downregulate CAR expression to limit on-target to tissue toxicity caused by tumor lysis syndrome. Downregulating CAR expression after antitumor efficacy may prevent (1) on-target off-tumor toxicity caused by antigen expression in normal tissues; (2) antigen-independent activation in vivo.

[0153] Extracellular targeting domains / sites In some embodiments, the extracellular targeting moiety of the CAR can be any agent that recognizes and binds to a given target molecule, e.g., a neoantigen on a tumor cell, with high specificity and affinity. The targeting moiety can be an antibody and variants thereof that specifically binds to a target molecule on a tumor cell, or a peptide aptamer selected from a random sequence pool based on its ability to bind to a target molecule on a tumor cell, or a variant or fragment thereof that can bind to a target molecule on a tumor cell, or an antigen recognition domain from a native T cell receptor (TCR) (e.g., the CD4 extracellular domain that recognizes HIV-infected cells), or a foreign recognition component, e.g., a linked cytokine that results in recognition of a target cell bearing a cytokine receptor, or a natural ligand for the receptor.

[0154] In some embodiments, the targeting domain of the CAR can be an Ig NAR, a Fab fragment, a Fab' fragment, a F(ab)'2 fragment, a F(ab)'3 fragment, an Fv, a single chain variable fragment (scFv), a bis-scFv, (scFv)2, a minibody, a diabody, a triabody, a tetrabody, a disulfide stabilized Fv protein (dsFv), a unibody, a nanobody, or an antigen binding region derived from an antibody that specifically recognizes a target molecule, such as a tumor specific antigen (TSA). In one embodiment, the targeting moiety is an scFv antibody. The scFv domain is expressed on the surface of the CAR T cell and can then maintain the CAR T cell in the vicinity of the cancer cell and trigger T cell activation when bound to a target protein on the cancer cell. The scFv can be generated using conventional recombinant DNA technology techniques and are discussed in this disclosure.

[0155] In some embodiments, the targeting moiety of the CAR construct can be an aptamer, such as a peptide aptamer, that specifically binds to a target molecule of interest. The peptide aptamer can be selected from a random sequence pool based on its ability to bind to a target molecule of interest.

[0156] In some embodiments, the targeting moiety of the CAR construct can be the natural ligand of the target molecule, or a variant and / or fragment thereof capable of binding the target molecule. In some aspects, the targeting moiety of the CAR can be the receptor of the target molecule, e.g., full-length human CD27 as a CD70 receptor, fused in-frame to the signaling domain of CD3ζ forming a CD27 chimeric receptor as an immunotherapeutic agent for CD70-positive malignancies.

[0157] In some embodiments, the targeting moiety of the CAR can recognize a tumor-specific antigen (TSA), e.g., a cancer neoantigen that is exclusively expressed on tumor cells.

[0158] As non-limiting examples, CARs of the present disclosure include 5T4, 707-AP, A33, AFP (alpha-fetoprotein), AKAP-4 (A kinase anchoring protein 4), ALK, alpha5beta1-integrin, androgen receptor, annexin II, alpha-actinin-4, ART-4, B1, B7H3, B7H4, BAGE (B melanoma antigen), BCMA, BCR-ABL fusion protein, beta-catenin, BKT-antigen, BTAA, CA-I (carbonic anhydrase I), CA50 (cancer antigen 50), CA125, CA15-3, CA195, CA2 42, calretinin, CAIX (carbonic anhydrase), CAMEL (antigen recognized by cytotoxic T-lymphocytes on melanoma), CAM43, CAP-1, caspase-8 / m, CD4, CD5, CD7, CD19, CD20, CD22, CD23, CD25, CD27 / m, CD28, CD30, CD33, CD34, CD36, CD38, CD40 / CD154, CD41, CD44v6, CD44v7 / 8, CD45, CD49f, CD56, CD68\KP1, CD74, CD79a / CD79b, CD103, CD123, CD133, CD138 , CD171, cdc27 / m, CDK4 (cyclin-dependent kinase 4), CDKN2A, CDS, CEA (carcinoembryonic antigen), CEACAM5, CEACAM6, chromogranin, c-Met, c-Myc, coa-1, CSAp, CT7, CT10, cyclophilin B, cyclin B1, cytoplasmic tyrosine kinase, cytokeratin, DAM-10, DAM-6, dek-can fusion protein, desmin, DEPDC1 (DEP domain-containing 1), E2A-PRL, EBNA, EGF-R (epidermal growth factor receptor), EGP-1 (epithelial glycoprotein protein-1) (TROP-2), EGP-2, EGP-40, EGFR (epidermal growth factor receptor), EGFRvIII, EF-2, ELF2M, EMMPRIN, EpCAM (epithelial cell adhesion molecule), EphA2, Epstein-Barr virus antigen, Erb (ErbB1; ErbB3; ErbB4), ETA (epithelial tumor antigen), ETV6-AML1 fusion protein, FAP (fibroblast activation protein), FBP (folate binding protein), FGF-5, folate receptor, FOS-related antigen 1, fucosyl GM1, G250, GAGE ​​(GAGE-1;GAGE-2), galectins, GD2 (ganglioside), GD3, GFAP (glial fibrillary acidic protein), GM2 (carcinoembryonic antigen-immunogenic-1; OFA-I-1), GnT-V, Gp100, H4-RET, HAGE (helicase antigen), HER-2 / neu, HIF (hypoxia-inducible factor), HIF-1, HIF-2, HLA-A2, HLA-A; *0201-R170I, HLA-All, HMWMAA, Hom / Mel-40, HSP70-2M (heat shock protein 70), HST-2, HTgp-175, hTERT (or hTRT), human papillomavirus-E6 / human papillomavirus-E7 and E6, iCE (immunocapture EIA), IGF-1R, IGH-IGK, IL-2R, IL-5, ILK (integrin-linked kinase), IMP3 (insulin-like growth factor II mRNA-binding protein 3), IRF4 (interferon regulatory factor 4), KDR (kinase insert domain receptor), KIAA0205, KRAB-zinc finger protein (KID)-3; KID31, KSA(17-1A), K-ras, LAGE, LCK, LDLR / FUT (LDLR-fucosyltransferase AS fusion protein), LeY (Lewis Y), MAD-CT-1, MAGE (tyrosinase, melanoma-associated antigen) (MAGE-1; MAGE-3), melan-A tumor antigen (MART), MART-2 / Ski, MC 1R (melanocortin 1 receptor), MDM2, mesothelin, MPHOSPH1, MSA (muscle-specific actin), mTOR (mammalian target of rapamycin), MUC-1, MUC-2, MUM-1 (melanoma-associated antigen (mutated) 1), MUM-2, MUM-3, myosin / m, MYL-RAR, NA88-A, N-acetylglucosaminyltransferase, neo-PAP, NF-KB (nuclear factor-kappa B), neurofilament, NSE (neuron-specific enolase), notch receptor, NuMa, N-Ras, NY-BR-1, NY-CO-1, NY-ESO-1, oncostatin M, OS-9, OY-TES1, p53 mutant, p190 minor bcr-abl, pl5(58), pl85erbB2, pl80erbB-3, PAGE (prostate-associated gene), PAP (prostatic acid phosphatase), PAX3, PAX5, PDGFR (platelet-derived growth factor receptor), cytochrome P450 involved in piperidine and pyrrolidine utilization (PIPA), Pml-RAR alpha fusion protein, PR-3 (proteinase 3), PSA (prostate-specific antigen), PSM, PSMA (prostate stem cell antigen), PRAME (preferentially expressed antigen in melanoma), PTPRK, RAGE (kidney tumor antigen), Raf (A-Raf, B-Raf, and C-Raf), Ras, receptor tyrosine kinase, RCAS1, RGSS, ROR1 (receptor tyrosine kinase-like orphan receptor 1), RU1, RU2, SAGE, SART-1, SART-3, SCP-1, SDCCAG16, SP-17 (seminal protein 17), src-family, SSX (synovial sarcoma X breakpoint)-1, SSX-2 (HOM-MEL-40), SSX-3, SSX-4, SSX-5, STAT-3, STAT-5, STAT-6, STEAD, STn, survivin, syk-ZAP70, TA-90 (Mac-2 binding protein \ cyclophilin C-associated protein), TAAL6, TACSTD1 (tumor-associated calcium signaling transducer 1), TACSTD2, TAG-72-4, TAGE, TARP (T-cell receptor gamma alternate reading frame protein), TEL / AML1 fusion protein, TEM1, TEM8 (endosialin or CD248), TGFβ, TIE2, TLP, TMPRSS2 The extracellular targeting domain may comprise an extracellular targeting domain capable of binding to a tumor-specific antigen selected from an ETS fusion gene, a TNF-receptor (TNF-α receptor, TNF-β receptor; or TNF-γ receptor), a transferrin receptor, TPS, TRP-1 (tyrosine-based related protein 1), TRP-2, TRP-2 / INT2, TSP-180, a VEGF receptor, WNT, WT-1 (Wilms tumor antigen), and XAGE.

[0159] In some embodiments, the cargo or payload may be or may encode a CAR that comprises a universal immune receptor having a targeting site capable of binding to a labeled antigen.

[0160] In some embodiments, the cargo or payload may be or may encode a CAR that includes a targeting moiety capable of binding to a pathogen antigen.

[0161] In some embodiments, the cargo or payload may be or may encode a non-protein molecule, for example a CAR that includes a targeting moiety capable of binding to tumor-associated glycolipids and carbohydrates.

[0162] In some embodiments, the cargo or payload may be or may encode a CAR that contains a targeting moiety capable of binding to components within the tumor microenvironment, including proteins expressed in various tumor stromal cells, including tumor-associated macrophages (TAMs), immature monocytes, immature dendritic cells, immunosuppressive CD4+CD25+ regulatory T cells (Tregs), and MDSCs.

[0163] In some embodiments, the cargo or payload may be or encode a CAR that includes a targeting moiety capable of binding to a cell surface adhesion molecule, a surface molecule of an inflammatory cell that appears in autoimmune disease, or a TCR that triggers autoimmunity. As non-limiting examples, targeting moieties of the present disclosure can be scFv antibodies that recognize a tumor-specific antigen (TSA), such as the scFvs of antibodies SS, SS1 and HN1 that specifically recognize and bind to human mesothelin, the scFv of an antibody GD2, a CD19 antigen binding domain, an NKG2D ligand binding domain, a human anti-mesothelin scFv, an anti-CS1 binder, an anti-BCMA binding domain, an anti-CD19 scFv antibody, a GFRalpha4 antigen binding fragment, an anti-CLL-1 (C-type lectin-like molecule 1) binding domain, a CD33 binding domain, a GPC3 (glypican-3) binding domain, a GFRalpha4 (glycosyl-phosphatidylinositol (GPI)-linked GDNF family alpha-receptor 4 cell-surface receptor) binding domain, a CD123 binding domain, an anti-ROR1 antibody or fragment thereof, an scFv specific for GPC-3, an scFv against CSPG4, and an scFv against folate receptor alpha.

[0164] Intracellular signaling domains After binding to its target molecule, the intracellular domain of the CAR fusion polypeptide transmits a signal to the immune effector cell and activates at least one of the normal effector functions of the immune effector cell, including cytolytic activity (e.g., cytokine secretion) or helper activity. Thus, the intracellular domain comprises the "intracellular signaling domain" of the T cell receptor (TCR).

[0165] In some embodiments, the entire intracellular signaling domain can be used, while in other embodiments, truncated portions of the intracellular signaling domain can be used in place of the intact chain, so long as they transmit the effector function signal.

[0166] In some embodiments, the intracellular signaling domain may contain a signaling motif known as an immunoreceptor tyrosine-based activation motif (ITAM). Examples of ITAM-containing cytoplasmic signaling sequences include those derived from TCR CD3 zeta, FcR gamma, FcR beta, CD3 gamma, CD3 delta, CD3 epsilon, CD5, CD22, CD79a, CD79b, and CD66d. In one example, the intracellular signaling domain is a CD3 zeta (CD3ζ) signaling domain.

[0167] In some embodiments, the intracellular region further comprises one or more costimulatory signaling domains that provide additional signals to immune effector cells. These costimulatory signaling domains, in combination with the signaling domain, may further improve the proliferation, activation, memory, persistence, and tumor eradication efficiency of CAR-engineered immune cells (e.g., CAR T cells). In some cases, the costimulatory signaling region contains one, two, three, or four cytoplasmic domains of one or more intracellular signaling and / or costimulatory molecules. The costimulatory signaling domains may be the intracellular / cytoplasmic domains of costimulatory molecules, including but not limited to CD2, CD7, CD27, CD28, 4-1BB (CD137), OX40 (CD134), CD30, CD40, ICOS (CD278), GITR (glucocorticoid-induced tumor necrosis factor receptor), LFA-1 (lymphocyte function-associated antigen-1), LIGHT, NKG2C, B7-H3. In one example, the costimulatory signaling domain is derived from the cytoplasmic domain of CD28. In another example, the costimulatory signaling domain is derived from the cytoplasmic domain of 4-1BB (CD137). In another example, the costimulatory signaling domain can be the intracellular domain of GITR as taught in U.S. Patent No. 9,175,308, the contents of which are incorporated herein by reference in their entirety.

[0168] In some embodiments, the intracellular region is an MHC class I molecule, a TNF receptor protein, an immunoglobulin-like protein, a cytokine receptor, an integrin, a signaling lymphocytic activation protein (SLAM), e.g., CD48, CD229, 2B4, CD84, NTB-A, CRACC, BLAME, CD2F-10, SLAMF6, SLAMF7, an activating NK cell receptor, BTLA, a Toll ligand receptor, OX40, CD2, CD7, CD27, CD28, CD30, CD40, CDS, ICAM-1, LFA-1 (CD11a / CD18), 4-1BB (CD137), B7-H3, CDS, ICAM-1, ICOS (CD278) , GITR, BAFFR, LIGHT, HVEM (LIGHTR), SLAMF7, NKp80 (KLRF1), NKp44, NKp30, NKp46, CD19, CD4, CD8 alpha, CD8 beta, IL2R beta, IL2R gamma, IL7R alpha, IL-15Ra, ITGA4, VLA1, CD49a, ITGA4, IA4, CD49D, ITGA6, VLA-6, CD49f, ITGAD, CD11d, ITGAE, CD103, ITGAL, CD11a, LFA-1, ITGAM, CD11b, ITGAX, CD11c, ITGB1, CD29, ITGB2, CD18, LFA-1, ITGB7, NKG2D, NKG2C, NKD2CSLP76, TNFR2, TRANCE / RANKL, DNAM1 (CD226), SLAMF4 (CD244, 2B4), CD84, CD96 (Tactile), CEACAM1, CRTAM, Ly9 (CD229), CD160 (BY55), PSGL1, CD100 (SEMA 4D), CD69, SLAMF6 (NTB-A, Ly108), SLAM (SLAMF1, CD150, IPO-3), BLAME (SLAMF8), SELPLG (CD162), LTBR, ​​LAT, CD270 (HVEM), GADS, SLP-76, PAG / Cbp, CD19a , a ligand that specifically binds CD83, DAP10, TRIM, ZAP70, a killer immunoglobulin receptor (KIR), such as KIR2DL1, KIR2DL2 / L3, KIR2DL4, KIR2DL5A, KIR2DL5B, KIR2DS1, KIR2DS2, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DL1 / S1, KIR3DL2, KIR3DL3, and KIR2DP1; a lectin-associated NK cell receptor, such as Ly49, Ly49A, and Ly49C.

[0169] In some embodiments, the intracellular signaling domain of the present disclosure may contain a signaling domain derived from JAK-STAT. In other embodiments, the intracellular signaling domain of the present disclosure may contain a signaling domain derived from DAP-12 (death associated protein 12) (Topfer et al., Immunol., 2015, 194:3201-3212; and Wang et al., Cancer Immunol., 2015, 3:815-826). DAP-12 is a vital signaling receptor in NK cells. Activation signals mediated by DAP-12 play an important role in inducing NK cell cytotoxic responses against certain tumor cells and virus-infected cells. The cytoplasmic domain of DAP12 contains an immunoreceptor tyrosine-based activation motif (ITAM). Thus, CARs containing DAP12-derived signaling domains can be used for adoptive transfer of NK cells.

[0170] Transmembrane domain In some embodiments, CAR may comprise a transmembrane domain. As used herein, the term "transmembrane domain (TM)" broadly refers to an amino acid sequence that spans the cell membrane and is about 15 residues in length. The transmembrane domain may comprise at least 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 amino acid residues and spans the cell membrane. In some embodiments, the transmembrane domain may be derived from either natural or synthetic sources. The transmembrane domain of CAR may be derived from any natural membrane-associated or transmembrane protein. For example, the transmembrane region may be derived from (i.e., may include at least the transmembrane region(s) thereof) the alpha, beta, or zeta chain of the T cell receptor, CD3 epsilon, CD4, CD5, CD8, CD8α, CD9, CD16, CD22, CD33, CD28, CD37, CD45, CD64, CD80, CD86, CD134, CD137, CD152, or CD154.

[0171] Alternatively, the transmembrane domains of the present disclosure may be synthetic, hi some aspects, the synthetic sequences may comprise primarily hydrophobic residues, such as leucine and valine.

[0172] In some embodiments, the transmembrane domain may be selected from the group consisting of a CD8α transmembrane domain, a CD4 transmembrane domain, a CD28 transmembrane domain, a CTLA-4 transmembrane domain, a PD-1 transmembrane domain, and a human IgG4 Fc region.

[0173] In some embodiments, the CAR may include an optional hinge region (also called a spacer). The hinge sequence is a short sequence of amino acids that facilitates flexibility of the extracellular targeting domain to move the target binding domain away from the effector cell surface to allow proper cell / cell contact, target binding, and effector cell activation. The hinge sequence may be located between the targeting site and the transmembrane domain. The hinge sequence may be any suitable sequence derived from or obtained from any suitable molecule. The hinge sequence may be derived from an immunoglobulin (e.g., IgG1, IgG2, IgG3, IgG4) hinge region, i.e., the sequence that falls between the CHI and CH2 domains of the immunoglobulin, e.g., IgG4 Fc hinge, all or part of the extracellular region of type 1 membrane proteins such as CD8α CD4, CD28, and CD7, which may be a wild type sequence or a derivative. Some hinge regions include an immunoglobulin CH3 domain or both the CH3 and CH2 domains. In certain embodiments, the hinge region may be modified from IgG1, IgG2, IgG3, or IgG4 to include one or more amino acid residues, e.g., 1, 2, 3, 4 or 5 residues, substituted with an amino acid residue different from that present in the unmodified hinge.

[0174] In some embodiments, the CAR may include one or more linkers between any of the domains of the CAR. The linker may be 1 to 30 amino acids in length. In this regard, the linker may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids in length. In other embodiments, the linker may be flexible.

[0175] In some embodiments, components comprising a targeting moiety, a transmembrane domain, and an intracellular signaling domain may be assembled in a single fusion polypeptide, which may be the payload of an effector module of the present disclosure.

[0176] In some embodiments, the cargo or payload may be or encode a CD19-specific CAR that targets various B-cell malignancies and a HER2-specific CAR that targets sarcoma, glioblastoma, and advanced Her2-positive lung malignancies.

[0177] Tandem CAR (TanCAR) In some embodiments, the CAR can be a tandem chimeric antigen receptor (TanCAR) that can target two, three, four, or more tumor-specific antigens. In some aspects, the CAR is a bispecific TanCAR that includes two targeting domains that recognize two different TSAs on tumor cells. A bispecific TanCAR can be further defined as including an extracellular region that includes a targeting domain (e.g., an antigen recognition domain) specific for a first tumor antigen and a targeting domain (e.g., an antigen recognition domain) specific for a second tumor antigen. In other aspects, the CAR is a multispecific TanCAR that includes three or more targeting domains configured in a tandem arrangement. The space between the targeting domains in a TanCAR can be about 5 to about 30 amino acids in length, for example, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, and 30 amino acids.

[0178] Split CAR In some embodiments, the CAR components, including the targeting moiety, transmembrane domain, and intracellular signaling domain, can be split into two or more parts to rely on multiple inputs that drive the assembly of an intact functional receptor. As a non-limiting example, a split CAR consists of two parts that assemble in a small molecule-dependent manner; one part of the receptor features an extracellular antigen binding domain (e.g., scFv), and the other part has an intracellular signaling domain, e.g., the CD3ζ intracellular domain.

[0179] In other embodiments, the split part of the CAR system can be further modified to increase the signal. As a non-limiting example, the second part of the cytoplasmic fragment can be fixed to the cell membrane by incorporating a transmembrane domain (e.g., CD8α transmembrane domain) into the construct. Additional extracellular domains, such as the extracellular domain that mediates homodimerization, can also be added to the second part of the CAR system. These modifications can increase the receptor output activity, i.e., T cell activation.

[0180] In some embodiments, the two parts of the split CAR system contain heterodimerization domains that conditionally interact with each other upon binding of a heterodimerization small molecule. In this way, the receptor components are assembled in the presence of a small molecule to form an intact system, which can then be activated by antigen engagement. Any known heterodimerization components can be incorporated into the split CAR system. Other small molecule-dependent heterodimerization domains can also be used, including but not limited to the gibberellin-induced dimerization system (GID1-GAI), trimethoprim-SLF-induced ecDHFR and FKBP dimerization, and ABA (abscisic acid)-induced dimerization of PP2C and PYL domains. Dual control using inducible assembly (e.g., ligand-dependent dimerization) and degradation (e.g., destabilization domain-induced CAR degradation) of the split CAR system can provide greater flexibility to control the activity of CAR-modified T cells.

[0181] Switchable CAR In some embodiments, the CAR can be a switchable CAR, which is a controllable CAR that can be temporarily switched on in response to a stimulus (e.g., a small molecule). In this CAR design, the system is directly integrated into the hinge domain that separates the scFv domain from the cell membrane domain in the CAR. Such a system can split or combine different important functions of the CAR, such as activation and costimulation, within different chains of the receptor complex, mimicking the complexity of the native structure of the TCR. This integrated system can switch the interaction of the scFv and antigen between on / off states controlled by the absence / presence of the stimulus.

[0182] Reversible CAR In some embodiments, the CAR can be a reversible CAR system. In this CAR structure, a LID domain (ligand-induced degradation) is incorporated into the CAR system. The CAR can be temporarily downregulated by adding a ligand of the LID domain.

[0183] Inhibitory CARs (iCARs) In some embodiments, the CAR can be an inhibitory CAR. Inhibitory CAR (iCAR) refers to a bispecific CAR design in which a negative signal is used to improve tumor specificity and limit normal tissue toxicity. This design incorporates a second CAR with a surface antigen recognition domain combined with an inhibitory signal domain to limit T cell reactivity even with simultaneous engagement of an activating receptor. This antigen recognition domain is directed against a normal tissue specific antigen so that T cells can be activated in the presence of a first target protein, but in the presence of a second protein that binds to the iCAR, T cell activation is inhibited.

[0184] As a non-limiting example, iCARs against prostate-specific membrane antigen (PMSA) based on CTLA4 and PD1 inhibitory domains have demonstrated the ability to selectively limit cytokine secretion, cytotoxicity and proliferation induced by T cell activation.

[0185] Chimeric Switch Receptor In some embodiments, the cargo or payload may be or code for a chimeric switch receptor that can switch a negative signal to a positive signal. As used herein, the term "chimeric switch receptor" refers to a fusion protein that includes a first extracellular domain and a second transmembrane and intracellular domain, where the first domain includes a negative signal region and the second domain includes a positive intracellular signaling region. In some aspects, the fusion protein is a chimeric switch receptor that contains the extracellular domain of an inhibitory receptor on a T cell fused to the transmembrane and cytoplasmic domain of a costimulatory receptor. This chimeric switch receptor can convert a T cell inhibitory signal to a T cell stimulatory signal.

[0186] As a non-limiting example, a chimeric switch receptor can include the extracellular domain of PD-1 fused to the transmembrane and cytoplasmic domains of CD28. In some embodiments, the extracellular domains of other inhibitory receptors, such as CTLA-4, LAG-3, TIM-3, KIR, and BTLA, can also be fused to the transmembrane and cytoplasmic domains from costimulatory receptors, such as CD28, 4-1BB, CD27, OX40, CD40, GTIR, and ICOS.

[0187] In some embodiments, a chimeric switch receptor can include a recombinant receptor that includes the extracellular cytokine binding domain of an inhibitory cytokine receptor (e.g., IL-13 receptor alpha (IL-13Rα1), IL-10R, and IL-4Rα) fused to the intracellular signaling domain of a stimulatory cytokine receptor, such as IL-2R (IL-2Rα, IL-2Rβ, and IL-2R gamma) and IL-7Rα. One example of such a chimeric cytokine receptor is a recombinant receptor that contains the cytokine-binding extracellular domain of IL-4Rα linked to the intracellular signaling domain of IL-7Rα.

[0188] In some embodiments, the chimeric switch receptor can be a chimeric TGFβ receptor. The chimeric TGFβ receptor can include an extracellular domain derived from a TGFβ receptor, such as TGFβ receptor 1, TGFβ receptor 2, TGFβ receptor 3, or any other TGFβ receptor or variant thereof; and a non-TGFβ receptor intracellular domain. The non-TGFβ receptor intracellular domain can be an intracellular domain or a fragment thereof derived from TLR1, TLR2, TLR3, TLR4, TLR5, TLR6, TLR7, TLR8, TLR9, TLR10, CD28, 4-1BB (CD137), OX40 (CD134), CD3 zeta, CD40, CD27, or a combination thereof.

[0189] Activation conditional CAR In some embodiments, the cargo or payload can be or can encode an activation conditional chimeric antigen receptor that is expressed only in activated immune cells. Expression of the CAR can be coordinated with a sequence, e.g., an activation conditional control region, which refers to one or more nucleic acid sequences that induce transcription and / or expression of the CAR under its control. Such an activation conditional control region can be a promoter of a gene that is upregulated during activation of immune effector cells, e.g., an IL2 promoter or an NFAT binding site.

[0190] Targeting CAR to tumor cells with specific proteoglycan markers In some embodiments, the cargo or payload may be or encode a CAR that targets a specific type of cancer cell. Human cancer cells and metastases may express unique and otherwise abnormal proteoglycans, such as polysaccharide chains (e.g., chondroitin sulfate (CS), dermatan sulfate (DS or CSB), heparan sulfate (HS) and heparin). Thus, the CAR may be fused with a binding site that recognizes cancer-associated proteoglycans. In one example, the CAR may be fused with a VAR2CSA polypeptide that binds with high affinity to a specific type of chondroitin sulfate A (CSA) bound to proteoglycans (VAR2-CAR). The extracellular ScFv portion of the CAR may be replaced with a VAR2CSA variant that contains at least the minimal CSA binding domain, generating a CAR specific for the chondroitin sulfate A (CSA) modification. Alternatively, the CAR can be fused to a split protein binding system to generate a spy-CAR, in which the scFv portion of the CAR is replaced with a portion of the split protein binding system, such as SpyTag and Spy-catcher, and the cancer recognition molecule (e.g., scFv and / or VAR2-CSA) is attached to the CAR via the split protein binding system.

[0191] nucleic acid The originator constructs and benchmark constructs of the present disclosure may include a payload region (which may also be referred to as a cargo region) that is a nucleic acid. The term "nucleic acid" in its broadest sense includes any compound and / or substance that includes a polymer of nucleotides that may be referred to as a polynucleotide. Exemplary nucleic acids or polynucleotides include, but are not limited to, ribonucleic acid (RNA), deoxyribonucleic acid (DNA), threose nucleic acid (TNA), glycol nucleic acid (GNA), peptide nucleic acid (PNA), locked nucleic acid (LNA), or hybrids thereof.

[0192] In some embodiments, the payload region comprises a nucleic acid sequence encoding multiple cargoes or payloads.

[0193] In some embodiments, the payload region is or can encode a coding nucleic acid sequence.

[0194] In some embodiments, the payload region may be, or may encode, a non-coding nucleic acid sequence.

[0195] In some embodiments, the payload region can be, or encode, both coding and non-coding nucleic acid sequences.

[0196] DNA Deoxyribonucleic acid (DNA), which carries the genetic information for all living organisms, is a molecule consisting of two strands that wrap around each other to form a shape known as a double helix. Each strand has a backbone composed of alternating sugars (deoxyribose) and phosphate groups. One of four bases: adenine (A), cytosine (C), guanine (G), and thymine (T) is attached to each sugar. The two strands are held together by bonds between adenine and thymine or between cytosine and guanine. The sequence of bases along the backbone serves as the instructions for building proteins and RNA molecules.

[0197] In some embodiments, the payload region may be or encode coding DNA.

[0198] In some embodiments, the payload region may be non-coding DNA or may encode it.

[0199] In some embodiments, the payload region may be, or encode, both coding DNA and non-coding RNA.

[0200] In some embodiments, the DNA may be modified. Types of modifications include, but are not limited to, methylation, acetylation, phosphorylation, ubiquitination, and sumoylation.

[0201] vector In some embodiments, the originator constructs and / or benchmark constructs described herein may be or be encoded by a vector, such as a plasmid or a viral vector. In some embodiments, the originator constructs and / or benchmark constructs may be or be encoded by a viral vector. The viral vector may be, but is not limited to, a herpes virus (HSV) vector, a retroviral vector, an adenovirus vector, an adeno-associated virus (AAV) vector, a lentiviral vector, and the like. In some embodiments, the viral vector is an AAV vector. In some embodiments, the viral vector is a lentiviral vector. In some embodiments, the viral vector is a retroviral vector. In some embodiments, the viral vector is an adenoviral vector.

[0202] Adeno-associated viral (AAV) vectors Parvoviridae viruses are small, non-enveloped, icosahedral capsid viruses characterized by a single-stranded DNA genome. Parvoviridae viruses consist of two subfamilies: Parvovirinae, which infect vertebrates, and Densovirinae, which infect invertebrates. This family of viruses is useful as a biological tool due to its relatively simple structure, which is easily manipulated using standard molecular biology techniques. The genome of the virus can be modified to contain the minimal components for the construction of a functional recombinant virus, or virus particle, loaded with, or engineered to express or deliver the desired payload that can be delivered to a target cell, tissue, organ, or organism.

[0203] The family Parvoviridae includes the genus Dependovirus, which comprises the adeno-associated viruses (AAV) capable of replication in vertebrate hosts, including but not limited to human, primate, bovine, canine, equine, and ovine species.

[0204] AAV vector genomes are linear single-stranded DNA (ssDNA) molecules approximately 5,000 nucleotides (nt) in length. AAV vector genomes may contain a payload region and at least one terminal inverted repeat (ITR) or ITR region. The ITRs are usually flanked by coding nucleotide sequences for nonstructural proteins (encoded by Rep genes) and structural proteins (encoded by capsid or Cap genes). Without wishing to be bound by theory, AAV vector genomes typically contain two ITR sequences. AAV vector genomes contain a characteristic T-shaped hairpin structure defined by self-complementary terminal 145 nucleotides at the 5' and 3' ends of ssDNA that form an energetically stable double-stranded region. The double-stranded hairpin structure contains multiple functions, including but not limited to acting as an origin for DNA replication by serving as a primer for the endogenous DNA polymerase complex of the host viral replicating cell.

[0205] In addition to the encoded heterologous payload, the AAV vector genome may comprise, in whole or in part, any naturally occurring and / or recombinant AAV serotype nucleotide sequence or variant. AAV variants may have significant sequence homology at the nucleic acid (genome or capsid) and amino acid levels (capsid) to generate constructs that are usually physical and functional equivalents and that replicate and assemble by similar mechanisms. Chiorini et al., J. Vir. 71:6823-33 (1997); Srivastava et al., J. Vir. 45:555-64 (1983); Chiorini et al., J. Vir. 73:1309-1319 (1999); Rutledge et al., J. Vir. 72:309-319 (1998); and Wu et al., J. Vir. 74:8635-47 (2000), the contents of each of which are incorporated by reference in their entirety.

[0206] In some embodiments, the AAV vector genome comprises at least one control element that provides for the replication, transcription, and translation of the coding sequence encoded therein. Not all of the control elements need always be present, as long as the coding sequence can be replicated, transcribed, and / or translated in a suitable host cell. Non-limiting examples of expression control elements include sequences for transcription initiation and / or termination, promoter and / or enhancer sequences, efficient RNA processing signals, such as splicing and polyadenylation signals, sequences that stabilize cytoplasmic mRNA, sequences that improve translation efficiency (e.g., Kozak consensus sequences), sequences that improve protein stability, and / or sequences that improve protein processing and / or secretion.

[0207] The AAV vector genomes of the present disclosure can be recombinantly produced and can be based on an adeno-associated virus (AAV) parent or reference sequence. As used herein, a "vector genome" is any molecule or moiety that transports, transduces, or otherwise acts as a carrier of a heterologous molecule, such as a nucleic acid as described herein.

[0208] In addition to single-stranded AAV vector genomes (e.g., ssAAV), the present disclosure also provides self-complementary AAV (scAAV) vector genomes. scAAV vector genomes contain DNA strands that anneal together to form double-stranded DNA. By skipping second strand synthesis, scAAV allows for rapid expression in cells.

[0209] In some embodiments, the AAV vector genome is a scAAV.

[0210] In some embodiments, the AAV vector genome is a ssAAV.

[0211] In some embodiments, AAV vectors are part of AAV particles, AAV1, AAV2, AAV2G9, AAV3, AAV3a, AAV3b, AAV3-3, AAV4, AAV4-4, AAV5, AAV6, AAV6.1, AAV6.2. 、AAV6.1.2、AAV7、AAV7.2、AAV8、AAV9、AAV9.11、AAV9.13、AAV9.16、AAV9.24、AAV9.45、AAV9.47、AAV9.61、AAV9.68、AAV9.84、AAV9.9、AAV10、AAV11、 AAV12、AAV16.3、AAV24.1、AAV27.3、AAV42.12、AAV42-1b、AAV42-2、AAV42-3a、AAV42-3b、AAV42-4、AAV42-5a、AAV42-5b、AAV42-6b、AAV42-8、AAV42-8 10、AAV42-11、AAV42-12、AAV42-13、AAV42-15、AAV42-aa、AAV43-1、AAV43-12、AAV43-20、AAV43-21、AAV43-23、AAV43-25、AAV43-5、AAV44.1、AAV44.2 、AAV44.5、AAV223.1、AAV223.2、AAV223.4、AAV223.5、AAV223.6、AAV223.7、AAV1-7 / rh.48、AAV1-8 / rh.49、AAV2-15 / rh.62、AAV2-3 / rh.61、AAV2-4 / rh.50、AAV2-5 / rh.51、AAV3.1 / hu.6、AAV3.1 / hu.9、AAV3-9 / rh.52、AAV3-11 / rh.53、AAV4-8 / r11.64、AAV4-9 / rh.54、AAV4-19 / rh.55、AAV5-3 / rh.57 AAV5-22 / rh.58、AAV7.3 / hu.7、AAV16.8 / hu.10、AAV16.12 / hu.11、AAV29.3 / bb.1、AAV29.5 / bb.2、AAV106.1 / hu.37、AAV114.3 / hu.40、AAV127.2 / hu.41、AAV127.5 / hu.42、AAV128.3 / hu.44、AAV130.4 / hu.48、AAV145.1 / hu.53、AAV145.5 / hu.54、AAV145.6 / hu.55、AAV161.10 / hu.60、AAV161.6 / hu.61、AAV33.12 / hu.17、AAV33.4 / hu.15、AAV33.8 / hu.16、AAV52 / hu.19、AAV52.1 / hu.20、AAV58.2 / hu.25、AAVA3.3、AAVA3.4、AAVA3.5、AAVA3.7、AAVC1、AAV C2、AAVC5、AAV-DJ、AAV-DJ8、AAVF3、AAVF5、AAVH2、AAVrh.72、AAVhu.8、AAVrh.68、AAVrh.70、AAVpi.1、AAVpi.3、AAVpi.2、AAVrh.60、AAVrh.44、AAVrh. .65、AAVrh.55、AAVrh.47、AAVrh.69、AAVrh.45、AAVrh.59、AAVhu.12、AAVH6、AAVLK03、AAVH-1 / hu.1、AAVH-5 / hu.3、AAVLG-10 / rh.40、AAVLG-4 / rh.38、AAVLG-9 / hu.39、AAVN721-8 / rh.43、AAVCh.5、AAVCh.5R1、AAVcy.2、AAVcy.3、AAVcy.4、AAVcy.5、AAVCy.5R1、AAVCy.5R2、AAVCy.5R3、AAVCy.5R4、AAVc y.6、AAVhu.1、AAVhu.2、AAVhu.3、AAVhu.4、AAVhu.5、AAVhu.6、AAVhu.7、AAVhu.9、AAVhu.10、AAVhu.11、AAVhu.13、AAVhu.15、AAVhu.16、AAVhu.17、AA Vhu.18、AAVhu.20、AAVhu.21、AAVhu.22、AAVhu.23.2、AAVhu.24、AAVhu.25、AAVhu.27、AAVhu.28、AAVhu.29、AAVhu.29R、AAVhu.31、AAVhu.32、AAVhu. 34、AAVhu.35、AAVhu.37、AAVhu.39、AAVhu.40、AAVhu.41、AAVhu.42、AAVhu.43、AAVhu.44、AAVhu.44R1、AAVhu.44R2、AAVhu.44R3、AAVhu.45、AAVhu.4 6、AAVhu.47、AAVhu.48、AAVhu.48R1、AAVhu.48R2、AAVhu.48R3、AAVhu.49、AAVhu.51、AAVhu.52、AAVhu.54、AAVhu.55、AAVhu.56、AAVhu.57、AAVhu.58、AAVhu.60、AAVhu.61、AAVhu.63、AAVhu.64、AAVhu.66、AAVhu.67、AAVhu.14 / 9、AAVhu.t 19、AAVrh.2、AAVrh.2R、AAVrh.8、AAVrh.8R、AAVrh.10、AAVrh.12、AAVrh.13、AAVrh.13R、AAVrh.14、AAVrh.17、AAVrh.18、AAVrh.19、AAVrh.20 、AAVrh.21、AAVrh.22、AAVrh.23、AAVrh.24、AAVrh.25、AAVrh.31、AAVrh.32、AAVrh.33、AAVrh.34、AAVrh.35、AAVrh.36、AAVrh.37、AAVrh.37 2、AAVrh.38、AAVrh.39、AAVrh.40、AAVrh.46、AAVrh.48、AAVrh.48.1、AAVrh.48.1.2、AAVrh.48.2、AAVrh.49、AAVrh.51、AAVrh.52、AAVrh.53、 AAVrh.54、AAVrh.56、AAVrh.57、AAVrh.58、AAVrh.61、AAVrh.64、AAVrh.64R1、AAVrh.64R2、AAVrh.67、AAVrh.73、AAVrh.74、AAVrh.8R、AAVrh.8R. A586R mutation、AAVrh8R R533A mutation、AAAV、BAAV、ヤギAAV、ウシAAV、AAVhE1.1、AAVhEr1.5、AAVhER1.14、AAVhEr1.8、AAVhEr1.16、AAVhEr1.18、AAVhEr1.35、AAV hEr1.7、AAVhEr1.36、AAVhEr2.29、AAVhEr2.4、AAVhEr2.16、AAVhEr2.30、AAVhEr2.31、AAVhEr2.36、AAVhER1.23、AAVhEr3.1、AAV 2.5T、AAV-PAEC、AAV-LK01、AAV-LK02、AAV-LK03、AAV-LK04、AAV-LK05、AAV-LK06、AAV-LK07、AAV-LK08、AAV-LK09、AAV-LK10、AAV -LK11、AAV-LK12、AAV-LK13、AAV-LK14、AAV-LK15、AAV-LK16、AAV-LK17、AAV-LK18、AAV-LK19、AAV-PAEC2、AAV-PAEC4、AAV-PAEC6、AAV-PAEC7, AAV-PAEC8, AAV-PAEC11, AAV-PAEC12, AAV-2-pre-miRNA-101, AAV-8h, AAV-8b, AAV-h, AAV-b, AAV SM 10-2, AAV Shuffle 100-1, AAV Shuffle 100-3, AAV Shuffle 100-7, AAV Shuffle 10-2, AAV Shuffle 10-6, AAV Shuffle 10-8, AAV Shuffle 100-2, AAV SM 10-1, AAV SM 10-8, AAV SM 100-3, AAV SM 100-10, BNP61 AAV, BNP62 AAV, BNP63 AAV, AAVrh.50, AAVrh.43, AAVrh.62, AAVrh.48, AAVhu.19, AAVhu.11, AAVhu.53, AAV4-8 / rh.64, AAVLG-9 / hu.39, AAV54.5 / hu.23, AAV54.2 / hu.22, AAV54.7 / hu.24, AAV54.1 / hu.21, AAV54.4R / hu.27, AAV46.2 / hu.28, AAV46.6 / hu.29, AAV128.1 / hu.43, True type AAV (ttAAV), UPENN AAV 10, Japanese AAV10 serotype, AAV CBr-7.1, AAV CBr-7.10, AAV CBr-7.2, AAV CBr-7.3, AAV CBr-7.4, AAV CBr-7.5, AAV CBr-7.7, AAV CBr-7.8, AAV CBr-B7.3, AAV CBr-B7.4, AAV CBr-E1, AAV CBr-E2, AAV CBr-E3, AAV CBr-E4, AAV CBr-E5, AAV CBr-e5, AAV CBr-E6, AAV CBr-E7, AAV CBr-E8, AAV CHt-1, AAV CHt-2, AAV CHt-3, AAV CHt-6.1, AAV CHt-6.10, AAV CHt-6.5, AAV CHt-6.6, AAV CHt-6.7, AAV CHt-6.8, AAV CHt-P1, AAV CHt-P2, AAV CHt-P5, AAV CHt-P6, AAV CHt-P8, AAV CHt-P9, AAV CKd-1, AAV CKd-10, AAV CKd-2, AAV CKd-3, AAV CKd-4, AAV CKd-6,AAV CKd-7、AAV CKd-8、AAV CKd-B1、AAV CKd-B2、AAV CKd-B3、AAV CKd-B4、AAV CKd-B5、AAV CKd-B6、AAV CKd-B7、AAV CKd-B8、AAV CKd-H1、AAV CKd-H2、AAV CKd-H3、AAV CKd-H4、AAV CKd-H5、AAV CKd-H6、AAV CKd-N3、AAV CKd-N4、AAV CKd-N9、AAV CLg-F1、AAV CLg-F2、AAV CLg-F3、AAV CLg-F4、AAV CLg-F5、AAV CLg-F6、AAV CLg-F7、AAV CLg-F8、AAV CLv-1、AAV CLv1-1、AAV Clv1-10、AAV CLv1-2、AAV CLv-12、AAV CLv1-3、AAV CLv-13、AAV CLv1-4、AAV Clv1-7、AAV Clv1-8、AAV Clv1-9、AAV CLv-2、AAV CLv-3、AAV CLv-4、AAV CLv-6、AAV CLv-8、AAV CLv-D1、AAV CLv-D2、AAV CLv-D3、AAV CLv-D4、AAV CLv-D5、AAV CLv-D6、AAV CLv-D7、AAV CLv-D8、AAV CLv-E1、AAV CLv-K1、AAV CLv-K3、AAV CLv-K6、AAV CLv-L4、AAV CLv-L5、AAV CLv-L6、AAV CLv-M1、AAV CLv-M11、AAV CLv-M2、AAV CLv-M5、AAV CLv-M6、AAV CLv-M7、AAV CLv-M8、AAV CLv-M9、AAV CLv-R1、AAV CLv-R2、AAV CLv-R3、AAV CLv-R4、AAV CLv-R5、AAV CLv-R6、AAV CLv-R7、AAV CLv-R8、AAV CLv-R9、AAV CSp-1、AAV CSp-10、AAV CSp-11、AAV CSp-2、AAV CSp-3、AAV CSp-4、AAV CSp-6、AAV CSp-7、AAV CSp-8、AAV CSp-8.10、AAV CSp-8.2、AAV CSp-8.4、AAV CSp-8.5、AAV CSp-8.6、AAV CSp-8.7、AAV CSp-8.8、AAV CSp-8.9, AAV CSp-9, AAV.hu.48R3, AAV.VR-355, AAV3B, AAV4, AAV5, AAVF1 / HSC1, AAVF11 / HSC11, AAVF12 / HSC12, AAVF13 / HSC13, AAVF14 / HSC14, AAVF15 / HSC15, AAVF16 / HSC16, These may be, but are not limited to, AAVF17 / HSC17, AAVF2 / HSC2, AAVF3 / HSC3, AAVF4 / HSC4, AAVF5 / HSC5, AAVF6 / HSC6, AAVF7 / HSC7, AAVF8 / HSC8, AAVF9 / HSC9, PHP.B, PHP.A, G2B-26, G2B-13, TH1.1-32, and / or TH1.1-35 and variants thereof.

[0212] Inverted terminal repeat (ITR) In some embodiments, the AAV vector genome may comprise at least one ITR region and a payload region. In some embodiments, the vector genome has two ITRs. These two ITRs flank the payload region at the 5' and 3' ends. The ITRs function as origins of replication that contain recognition sites for replication. The ITRs contain sequence regions that can be complementary and symmetrically arranged. The ITRs incorporated into the vector genome of the present disclosure may comprise naturally occurring or recombinantly derived polynucleotide sequences.

[0213] The ITR may be from the same serotype as the capsid or its derivative. The ITR may be of a different serotype than the capsid. In some embodiments, the AAV particle has multiple ITRs. In a non-limiting example, the AAV particle has a vector genome that includes two ITRs. In some embodiments, the ITRs may be of the same serotype as each other. In another embodiment, the ITRs are of different serotypes. A non-limiting example includes none, one or both of the ITRs that have the same serotype as the capsid. In some embodiments, both of the ITRs of the vector genome of the AAV particle are AAV2 ITRs.

[0214] Independently, each ITR can be about 100 to about 150 nucleotides in length. The ITRs can be about 100-105 nucleotides in length, 106-110 nucleotides in length, 111-115 nucleotides in length, 116-120 nucleotides in length, 121-125 nucleotides in length, 126-130 nucleotides in length, 131-135 nucleotides in length, 136-140 nucleotides in length, 141-145 nucleotides in length, or 146-150 nucleotides in length. In some embodiments, the ITRs are 140-142 nucleotides in length. Non-limiting examples of ITR lengths are 102, 140, 141, 142, 145 nucleotides in length, and those having at least 95% identity thereto.

[0215] promoter In some embodiments, the payload region of the vector genome comprises at least one element for enhancing transgene target specificity and expression (see, e.g., Powell et al. Viral Expression Cassette Elements to Enhance Transgene Target Specificity and Expression in Gene Therapy, 2015, the contents of which are incorporated herein by reference in their entirety). Non-limiting examples of elements for enhancing transgene target specificity and expression include promoters, endogenous miRNAs, post-transcriptional regulatory elements (PREs), polyadenylation (polyA) signal sequences and upstream enhancers (USEs), CMV enhancers and introns.

[0216] In some embodiments, the promoter is efficient at inducing expression of a polypeptide(s) encoded in the payload region of the vector genome of an AAV particle.

[0217] In some embodiments, a promoter is considered efficient if it induces expression in targeted cells.

[0218] In some embodiments, the promoter induces expression of the payload in the targeted tissue for a predetermined period of time. The promoter-induced expression may be 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 23 hours, 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 8 days, 9 days, 10 days, 11 days, 12 days, 13 days, 2 weeks, 15 days, 16 days, 17 days, 18 days, 19 days, 20 days, 22 days, 23 days, 24 days, 25 days, 26 days, 27 days, 28 days, 29 days, 30 days, 31 days, 32 days, 33 days, 34 days, 35 days, 36 days, 37 days, 38 days, 39 days, 40 days, 41 days, 42 days, 43 days, 44 days, 45 days, 46 days, 47 days, 48 ​​days, 49 days, 50 days, 51 days, 52 days, 53 days, 54 days, 55 days, 56 days, 57 days, 58 days, 59 days, 60 days, 61 days, 62 days, 63 days, 64 days, 65 days, 66 days, 67 days, 68 days, 69 days, 70 days, 71 days, 72 days, 73 days, 74 days, 75 days, 76 days, 77 days, 78 days The period may be days, 3 weeks, 22 days, 23 days, 24 days, 25 days, 26 days, 27 days, 28 days, 29 days, 30 days, 31 days, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 1 year, 13 months, 14 months, 15 months, 16 months, 17 months, 18 months, 19 months, 20 months, 21 months, 22 months, 23 months, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, 10 years or more than 10 years. Onset can be 1-5 hours, 1-12 hours, 1-2 days, 1-5 days, 1-2 weeks, 1-3 weeks, 1-4 weeks, 1-2 months, 1-4 months, 1-6 months, 2-6 months, 3-6 months, 3-9 months, 4-8 months, 6-12 months, 1-2 years, 1-5 years, 2-5 years, 3-6 years, 3-8 years, 4-8 years, or 5-10 years.

[0219] In some embodiments, the promoter regulates expression of the payload for at least 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, 10 years, 11 years, 12 years, 13 years, 14 years, 15 years, 16 years, 17 years, 18 years, 19 years, 20 years , 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 65 years or more.

[0220] The promoter may be naturally occurring or non-naturally occurring. Non-limiting examples of promoters include viral promoters, plant promoters, and mammalian promoters. In some embodiments, the promoter may be a human promoter. In some embodiments, the promoter may be a truncated form.

[0221] Promoters that drive or enhance expression in most tissues include, but are not limited to, human elongation factor 1 alpha-subunit (EF1α), cytomegalovirus (CMV) immediate early enhancer and / or promoter, chicken beta-actin (CBA) and its derivatives CAG, beta glucuronidase (GUSB), or ubiquitin C (UBC). Tissue-specific expression elements can be used to restrict expression to certain cell types, such as, but not limited to, muscle-specific promoters, B-cell promoters, monocyte promoters, leukocyte promoters, macrophage promoters, pancreatic acinar cell promoters, endothelial cell promoters, lung tissue promoters, astrocyte promoters, or nervous system promoters that can be used to restrict expression to neurons, astrocytes, or oligodendrocytes.

[0222] Non-limiting examples of muscle-specific promoters include the mammalian muscle creatine kinase (MCK) promoter, the mammalian desmin (DES) promoter, the mammalian troponin I (TNNI2) promoter, and the mammalian skeletal alpha-actin (ASKA) promoter (see, e.g., U.S. Patent Publication US20110212529, the contents of which are incorporated herein by reference in their entirety).

[0223] Non-limiting examples of tissue-specific expression elements for neurons include neuron-specific enolase (NSE), platelet-derived growth factor (PDGF), platelet-derived growth factor B-chain (PDGF-β), synapsin (Syn), methyl-CpG binding protein 2 (MeCP2), Ca 2+ / Calmodulin-dependent protein kinase II (CaMKII), metabotropic glutamate receptor 2 (mGluR2), neurofilament light (NFL) or heavy (NFH), β-globin minigene nβ2, preproenkephalin (PPE), enkephalin (Enk) and excitatory amino acid transporter 2 (EAAT2) promoters. Non-limiting examples of tissue-specific expression elements for astrocytes include glial fibrillary acidic protein (GFAP) and EAAT2 promoters. Non-limiting examples of tissue-specific expression elements for oligodendrocytes include the myelin basic protein (MBP) promoter.

[0224] In some embodiments, the promoter may be less than 1 kb. In some embodiments, the nucleic acid sequence may be 0, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, or more than 800 nucleotides in length. The promoter may have a length of 200-300, 200-400, 200-500, 200-600, 200-700, 200-800, 300-400, 300-500, 300-600, 300-700, 300-800, 400-500, 400-600, 400-700, 400-800, 500-600, 500-700, 500-800, 600-700, 600-800, or 700-800.

[0225] In some embodiments, the promoter can be a combination of two or more components of the same or different initiating or parent promoters, such as, but not limited to, CMV and CBA. In some embodiments, the length may be 0, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, or greater than 800. Each component may have a length of 200-300, 200-400, 200-500, 200-600, 200-700, 200-800, 300-400, 300-500, 300-600, 300-700, 300-800, 400-500, 400-600, 400-700, 400-800, 500-600, 500-700, 500-800, 600-700, 600-800, or 700-800. In some embodiments, the promoter is a combination of a 382 nucleotide CMV-enhancer sequence and a 260 nucleotide CBA-promoter sequence.

[0226] In some embodiments, the vector genome comprises a ubiquitous promoter. Non-limiting examples of ubiquitous promoters include CMV, CBA (including derivatives CAG, CBh, etc.), EF-1α, PGK, UBC, GUSB (hGBp), and UCOE (promoter of HNRPA2B1-CBX3).

[0227] In some embodiments, the promoter is not cell-specific.

[0228] In some embodiments, the vector genome comprises an engineered promoter.

[0229] In some embodiments, the vector genome comprises a promoter from a naturally expressed protein.

[0230] Untranslated Regions (UTRs) By definition, the wild-type untranslated region (UTR) of a gene is transcribed but not translated. Usually, the 5'UTR begins at the transcription start site and ends at the start codon, and the 3'UTR begins immediately after the stop codon and continues until the termination signal for transcription.

[0231] Features typically found in abundantly expressed genes of a particular target organ can be engineered into the UTR to improve stability and protein production. As a non-limiting example, the 5'UTR from an mRNA normally expressed in the liver (e.g., albumin, serum amyloid A, apolipoprotein A / B / E, transferrin, alpha-fetoprotein, erythropoietin, or factor VIII) can be used in the vector genome of the AAV particles of the present disclosure to improve expression in hepatic cell lines or the liver.

[0232] Without wishing to be bound by theory, wild-type 5' untranslated regions (UTRs) contain features that play a role in translation initiation. The 5'UTR usually contains a Kozak sequence, which is commonly known to be involved in the process by which the ribosome initiates the translation of many genes. The Kozak sequence has the consensus CCR(A / G)CCAUGG, where R is a purine (adenine or guanine) three bases upstream of the start codon (ATG) (followed by another "G").

[0233] In some embodiments, the 5'UTR in the vector genome comprises a Kozak sequence.

[0234] In some embodiments, the 5'UTR in the vector genome does not include a Kozak sequence.

[0235] Without wishing to be bound by theory, it is known that wild-type 3'UTRs have stretches of adenosines and uridines embedded in them. These AU-rich signatures are particularly prevalent in genes with high turnover rates. Based on their sequence characteristics and functional properties, AU-rich elements (AREs) can be divided into three classes (Chen et al., 1995, the contents of which are incorporated herein by reference in their entirety): Class I AREs, such as but not limited to c-Myc and MyoD, contain several dispersed copies of the AUUUA motif within the U-rich region. Class II AREs, such as but not limited to GM-CSF and TNF-a, possess two or more overlapping UUAUUUA(U / A)(U / A) nonamers. Class III AREs, such as but not limited to c-Jun and myogenin, are less well defined. These U-rich regions do not contain the AUUUA motif. While most proteins that bind to AREs are known to destabilize messengers, members of the ELAV family, most notably HuR, have been documented to increase mRNA stability. HuR binds to all three classes of AREs. Engineering a HuR-specific binding site within the 3'UTR of a nucleic acid molecule results in HuR binding and thus stabilization of the message in vivo.

[0236] The introduction, removal or modification of 3'UTR AU-rich elements (AREs) can be used to regulate the stability of polynucleotides. When manipulating a particular polynucleotide, such as the payload region of a vector genome, one or more copies of AREs can be introduced to make the polynucleotide less stable, thereby reducing translation and decreasing the production of the resulting protein. Similarly, AREs can be identified and removed or mutated to increase intracellular stability, thereby increasing the translation and production of the resulting protein.

[0237] In some embodiments, the 3'UTR of the vector genome may contain an oligo(dT) sequence for templated addition of a poly-A tail.

[0238] In some embodiments, the vector genome may include at least one miRNA seed, binding site or complete sequence. MicroRNAs (or miRNAs or miRs) are 19-25 nucleotide non-coding RNAs that bind to sites in nucleic acid targets and downregulate gene expression either by reducing nucleic acid molecule stability or by inhibiting translation. The microRNA sequence includes a "seed" region, i.e., a sequence in the region of positions 2-8 of the mature microRNA, which has perfect Watson-Crick complementarity with the miRNA target sequence of the nucleic acid.

[0239] In some embodiments, the vector genome may be engineered to include, modify or remove at least one miRNA binding site, sequence, or seed region.

[0240] Any UTR from any gene known in the art can be incorporated into the vector genome of an AAV particle. These UTRs, or parts thereof, can be placed in the same orientation as the gene they are selected from, or they can be modified in orientation or position. In some embodiments, the UTRs used in the vector genome of an AAV particle can be reversed, shortened, lengthened, or made with one or more other 5'UTRs or 3'UTRs known in the art. As used herein, the term "modified" when referring to a UTR means that it is changed in some way relative to a reference sequence. For example, the 3' or 5'UTR can be modified compared to the wild-type or native UTR by changing orientation or position as taught above, or by including additional nucleotides, deleting nucleotides, exchanging or translocating nucleotides.

[0241] In some embodiments, the vector genome of the AAV particle comprises at least one artificial UTR that is not a variant of the wild-type UTR.

[0242] In some embodiments, the vector genome of the AAV particle comprises UTRs that are selected from a family of transcripts whose proteins share a common function, structure, feature or characteristic.

[0243] Polyadenylation sequence In some embodiments, the vector genome comprises at least one polyadenylation sequence between the 3' end of the payload coding sequence and the 5' end of the 3' ITR.

[0244] In some embodiments, polyadenylation (poly-A) sequences can range in length from absent to about 500 nucleotides. The polyadenylation sequences are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77 , 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 1 42, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259,260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290 , 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321 , 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 37 2, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 372, 373, 374, 375, 376, 377, 378, 379, 380, 381, 382, ​​383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 399, 400, 401, 4 3, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 4 14, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 443, 444, 445, 446, 447, 448, 449, 450, 451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, The length of the nucleic acid sequence may be, but is not limited to, 45, 446, 447, 448, 449, 450, 451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492, 493, 494, 495, 496, 497, 498, 499, and 500 nucleotides.

[0245] In some embodiments, the polyadenylation sequence is 50-100 nucleotides in length. In some embodiments, the polyadenylation sequence is 50-150 nucleotides in length. In some embodiments, the polyadenylation sequence is 50-160 nucleotides in length. In some embodiments, the polyadenylation sequence is 50-200 nucleotides in length. In some embodiments, the polyadenylation sequence is 60-100 nucleotides in length. In some embodiments, the polyadenylation sequence is 60-150 nucleotides in length. In some embodiments, the polyadenylation sequence is 60-160 nucleotides in length. In some embodiments, the polyadenylation sequence is 60-200 nucleotides in length. In some embodiments, the polyadenylation sequence is 70-100 nucleotides in length. In some embodiments, the polyadenylation sequence is 70-150 nucleotides in length. In some embodiments, the polyadenylation sequence is 70-160 nucleotides in length. In some embodiments, the polyadenylation sequence is 70-200 nucleotides in length. In some embodiments, the polyadenylation sequence is 80-100 nucleotides in length. In some embodiments, the polyadenylation sequence is 80-150 nucleotides in length. In some embodiments, the polyadenylation sequence is 80-160 nucleotides in length. In some embodiments, the polyadenylation sequence is 80-200 nucleotides in length. In some embodiments, the polyadenylation sequence is 90-100 nucleotides in length. In some embodiments, the polyadenylation sequence is 90-150 nucleotides in length. In some embodiments, the polyadenylation sequence is 90-160 nucleotides in length. In some embodiments, the polyadenylation sequence is 9 ...

[0246] Linker The vector genome may be engineered with one or more spacer or linker regions to separate coding and non-coding regions.

[0247] In some embodiments, the payload region of the vector genome may optionally encode one or more linker sequences. In some cases, the linker may be a peptide linker that can be used to connect the polypeptides encoded by the payload region (i.e., the light and heavy antibody chains during expression). Some peptide linkers may be cleaved after expression to separate the heavy and light chain domains, allowing the construction of a mature antibody or antibody fragment. Linker cleavage may be enzymatic. In some cases, the linker contains an enzymatic cleavage site to facilitate intracellular or extracellular cleavage. Some payload regions encode linkers that block polypeptide synthesis during translation of the linker sequence from an mRNA transcript. Such linkers may facilitate the translation of separate protein domains from a single transcript. In some cases, two or more linkers are encoded by the payload region of the vector genome.

[0248] An internal ribosome entry site (IRES) is a nucleotide sequence (>500 nucleotides) that allows initiation of translation in the middle of an mRNA sequence (Kim, JH et al., 2011. PLoS One 6(4):e18556, the contents of which are incorporated herein by reference in their entirety). The use of an IRES sequence ensures co-expression of the genes before and after the IRES, although the sequence after the IRES may be transcribed and translated at a lower level than the sequence preceding the IRES sequence.

[0249] 2A peptides are small "self-cleaving" peptides (18-22 amino acids) derived from viruses, e.g., foot and mouth disease virus (F2A), porcine teschovirus-1 (P2A), Thoseaasigna virus (T2A), or equine rhinitis A virus (E2A). The 2A designation specifically refers to a region of the picornavirus polyprotein that results in a ribosomal skip at a glycyl-prolyl bond at the C-terminus of the 2A peptide (Kim, J. Het al., 2011. PLoS One 6(4):e18556, the contents of which are incorporated herein by reference in their entirety). This skip results in cleavage between the 2A peptide and its immediately downstream peptide. In contrast to IRES linkers, 2A peptides generate stoichiometric expression of proteins adjacent to the 2A peptide, and their shorter length may be advantageous in generating viral expression vectors.

[0250] Some payload regions encode linkers that contain a furin cleavage site. Furin is a calcium-dependent serine endoprotease that cleaves proteins immediately downstream of a basic amino acid target sequence (Arg-X-(Arg / Lys)-Arg) (Thomas, G., 2002. Nature Reviews Molecular Cell Biology 3(10):753-66, the contents of which are incorporated herein by reference in their entirety). Furin is concentrated in the trans-Golgi network, where it is involved in the processing of cellular precursor proteins. Furin also plays a role in activating a number of pathogens. This activity can be utilized for the expression of the polypeptides of the present disclosure.

[0251] In some embodiments, the payload region may encode one or more linkers comprising cathepsin, matrix metalloproteinase or legumain cleavage sites. Such linkers are described, for example, by Cizeau and Macdonald in International Publication No. WO2008052322, the contents of which are incorporated herein in their entirety. Cathepsins are a family of proteases with unique mechanisms for cleaving specific proteins. Cathepsin B is a cysteine ​​protease and cathepsin D is an aspartyl protease. Matrix metalloproteinases are a family of calcium-dependent and zinc-containing endopeptidases. Legumain is an enzyme that catalyzes the hydrolysis of (-Asn-Xaa-) bonds in proteins and small molecule substrates.

[0252] In some embodiments, the payload region may encode a non-cleavable linker. Such a linker may include a single amino acid sequence, e.g., a glycine-rich sequence. In some cases, the linker may include a flexible peptide linker that includes glycine and serine residues. The linker may include flexible peptide linkers of different lengths, e.g., nxG4S, where n=1-10, with the encoded linker length varying from 5-50 amino acids. In a non-limiting example, the linker may be 5xG4S. These flexible linkers are small and have no side chains so that they tend not to affect secondary protein structure while providing a flexible linker between antibody segments (George, RA, et al., 2002. Protein Engineering 15(11):871-9; Huston, J Set al., 1988. PNAS 85:5879-83; and Shan, D. et al., 1999. Journal of Immunology. 162(11):6589-95, the contents of each of which are incorporated herein by reference in their entirety). Additionally, the polarity of the serine residues improves solubility and prevents aggregation problems.

[0253] In some embodiments, the payload regions of the present disclosure may encode small, unbranched, serine-rich peptide linkers, such as those described by Huston et al. in U.S. Patent No. US5525491, the contents of which are incorporated herein in their entirety. Polypeptides encoded by payload regions of the present disclosure linked by serine-rich linkers have increased solubility.

[0254] In some embodiments, the payload region of the present disclosure may encode an artificial linker, such as those described by Whitlow and Filpula in U.S. Patent No. US5856456 and by Ladner et al. in U.S. Patent No. US4946778, the contents of each of which are incorporated herein in their entirety.

[0255] Introns In some embodiments, the payload region includes at least one element for improving expression, such as one or more introns or portions thereof. Non-limiting examples of introns include MVM (67-97 bp), F.IX truncated intron 1 (300 bp), β-globin SD / immunoglobulin heavy chain splice acceptor (250 bp), adenovirus splice donor / immunoglobin splice acceptor (500 bp), SV40 late splice donor / splice acceptor (19S / 16S) (180 bp), and hybrid adenovirus splice donor / IgG splice acceptor (230 bp).

[0256] In some embodiments, an intron or intron portion may be 100 to 500 nucleotides in length. Introns may have a length of 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490 or 500. The intron may have a length of 80-100, 80-120, 80-140, 80-160, 80-180, 80-200, 80-250, 80-300, 80-350, 80-400, 80-450, 80-500, 200-300, 200-400, 200-500, 300-400, 300-500, or 400-500.

[0257] Lentiviral Vectors Lentiviral vectors are a kind of retrovirus, and can infect both dividing and non-dividing cells, because their viral shell can pass through the intact membrane of the nucleus of target cells.Lentiviral vectors have the ability to deliver transgenes in tissues where stable genetic manipulation has long been thought to be hopelessly ineffective.Lentiviral vectors also open up new perspectives for the genetic treatment of various inherited and acquired disorders, and the practical proposal for their clinical use seems imminent.

[0258] RNA Ribonucleic acid (RNA) is a molecule composed of nucleotides, which are ribose sugars linked to nitrogenous bases and phosphate groups. The nitrogenous bases include adenine (A), guanine (G), uracil (U), and cytosine (C). Usually, RNA mostly exists in single-stranded form, but in certain circumstances it can exist in double-stranded form. The length, form, and structure of RNA vary depending on the purpose of the RNA. For example, the length of RNA can vary from short sequences (e.g., siRNA) to long sequences (e.g., lncRNA), can be linear (e.g., mRNA) or circular (e.g., oRNA), and can be either coding (e.g., mRNA) or non-coding (e.g., lncRNA) sequences.

[0259] In some embodiments, the payload region may be or encode a coding RNA.

[0260] In some embodiments, the payload region may be or encode a non-coding RNA.

[0261] In some embodiments, the payload region may be, or encode, both coding and non-coding RNA.

[0262] In some embodiments, the payload region comprises a nucleic acid sequence encoding multiple cargoes or payloads.

[0263] In some embodiments, the payload region comprises a nucleic acid sequence for enhancing expression of a gene. As a non-limiting example, the nucleic acid sequence is messenger RNA (mRNA). As another non-limiting example, the nucleic acid sequence is circular RNA (oRNA).

[0264] In some embodiments, the payload region comprises a nucleic acid sequence for reducing or inhibiting expression of a gene. As non-limiting examples, the nucleic acid sequence is a small interfering RNA (siRNA) or a microRNA (miRNA).

[0265] Messenger RNA (mRNA) In some embodiments, the originator construct and / or the benchmark construct may be an mRNA. As used herein, the term "messenger RNA" (mRNA) refers to any polynucleotide that encodes a target of interest and can be translated in vitro, in vivo, in situ or ex vivo to generate the encoded target of interest.

[0266] Typically, an mRNA molecule includes at least a coding region, a 5' untranslated region (UTR), a 3' UTR, a 5' cap, and a poly-A tail. In some embodiments, one or more structural and / or chemical modifications or alterations may be included in the RNA that may reduce the innate immune response of a cell into which the mRNA is introduced. As used herein, a "structural" feature or modification is one in which two or more linked nucleotides are inserted, deleted, duplicated, inverted, or randomized in a nucleic acid without significant chemical modification to the nucleotides themselves. Since chemical bonds are necessarily broken and reformed to achieve structural modification, structural modifications are of a chemical nature and are therefore chemical modifications. However, structural modifications will result in a different nucleotide sequence. For example, the polynucleotide "ATCG" may be chemically modified to "AT-5meC-G".

[0267] Typically, the minimum length of a region of an originator construct and / or a benchmark construct may be a length of nucleic acid sequence sufficient to encode a dipeptide, tripeptide, tetrapeptide, pentapeptide, hexapeptide, heptapeptide, octapeptide, nonapeptide, or decapeptide. In another embodiment, the length may be sufficient to encode a peptide of 2-30 amino acids, e.g., 5-30, 10-30, 2-25, 5-25, 10-25, or 10-20 amino acids. The length may be sufficient to encode a peptide of at least 11, 12, 13, 14, 15, 17, 20, 25, or 30 amino acids, or a peptide of 40 amino acids or less, e.g., 35, 30, 25, 20, 17, 15, 14, 13, 12, 11, or 10 amino acids or less.

[0268] Typically, the length of the region of an mRNA encoding a target of interest is greater than about 30 nucleotides in length (e.g., at least about 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,500, and 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000 nucleotides or up to 100,000 nucleotides).

[0269] In some embodiments, the mRNA comprises from about 30 to about 100,000 nucleotides (e.g., 30 to 50, 30 to 100, 30 to 250, 30 to 500, 30 to 1,000, 30 to 1,500, 30 to 3,000, 30 to 5,000, 30 to 7,000, 30 to 10,000, 30 to 25,000, 30 to 50,000, 30 to 70,000, 100 to 250, 100 to 500, 100 to 1,000, 10 0~1,500, 100~3,000, 100~5,000, 100~7,000, 100~10,000, 100~25,000, 100~50,000, 100~70,000, 100~100,000, 500~1,000, 500~1,500, 500~2,000, 500~3,000, 500~5,000, 500~7,000, 500~10,000, 500~25,000, 500~50, 000, 500~70,000, 500~100,000, 1,000~1,500, 1,000~2,000, 1,000~3,000, 1,000~5,000, 1,000~7,000, 1,000~10,000, 1,000~25,000, 1,000~50,000, 1,000~70,000, 1,000~100,000, 1,500~3,000, 1,500~5,000, 1,500 ~7,000, 1,500~10,000, 1,500~25,000, 1,500~50,000, 1,500~70,000, 1,500~100,000, 2,000~3,000, 2,000~5,000, 2,000~7,000, 2,000~10,000, 2,000~25,000, 2,000~50,000, 2,000~70,000, and 2,000~100,000).

[0270] In some embodiments, the region encoding the target of interest or the regions adjacent to it can independently range from 15 to 1,000 nucleotides in length (e.g., greater than, or at least, 30, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, and 900 nucleotides).

[0271] In some embodiments, the mRNA comprises a tail sequence (e.g., at least 60, 70, 80, 90, 120, 140, 160, 180, 200, 250, 300, 350, 400, 450, or 500 nucleotides) that can range in length from none to 500 nucleotides. If the tail region is a polyA tail, the length can be determined in units of polyA binding protein binding or according to its function. In this embodiment, the polyA tail is long enough to bind at least 4 monomers of polyA binding protein. A polyA binding protein monomer binds to a stretch of approximately 38 nucleotides. Thus, polyA tails of about 80 nucleotides and 160 nucleotides have been observed to be functional.

[0272] In some embodiments, the mRNA comprises a cap sequence that includes a single cap or a series of nucleotides that form a cap. The cap sequence can be 1-10, e.g., 2-9, 3-8, 4-7, 1-5, 5-10, or at least 2, or no more than 10 nucleotides in length. In some embodiments, the cap sequence is absent.

[0273] In some embodiments, the mRNA includes a region that includes the start codon. The region that includes the start codon can range from 3 to 40, e.g., 5 to 30, 10 to 20, 15, or at least 4, or no more than 30 nucleotides in length.

[0274] In some embodiments, the mRNA includes a region that includes a stop codon. The region that includes the stop codon can range from 3 to 40, e.g., 5 to 30, 10 to 20, 15, or at least 4, or no more than 30 nucleotides in length.

[0275] In some embodiments, the mRNA includes a region that includes a restriction sequence. The region that includes a restriction sequence can range from 3 to 40, e.g., 5 to 30, 10 to 20, 15, or at least 4, or no more than 30 nucleotides in length.

[0276] Untranslated Regions (UTRs) In some embodiments, an mRNA comprises at least one untranslated region (UTR) adjacent to the region encoding the target of interest. UTRs are transcribed but not translated.

[0277] The 5'UTR begins at the transcription initiation site and continues up to, but not including, the start codon; whereas the 3'UTR begins just after the stop codon and continues up to the transcription termination signal. Without wishing to be bound by theory, UTRs may have a regulatory role in terms of translation and stability of the nucleic acid.

[0278] Natural 5'UTRs usually contain features that play a role in translation initiation, as they tend to contain Kozak sequences that are commonly known to be involved in the process by which ribosomes initiate the translation of many genes. Kozak sequences have the consensus CCR(A / G)CCAUGG, where R is a purine (adenine or guanine) three bases upstream of the start codon (AUG) (followed by another "G"). 5'UTRs are also known to form secondary structures involved in elongation factor binding.

[0279] 3'UTRs are known to have stretches of adenosines and uridines embedded in them. These AU-rich signatures are particularly prevalent in genes with high turnover rates. Based on their sequence features and functional properties, AU-rich elements (AREs) can be divided into three classes (Chen et al., 1995): Class I AREs contain several dispersed copies of the AUUUA motif within the U-rich region. C-Myc and MyoD contain Class I AREs. Class II AREs possess two or more overlapping UUAUUUA(U / A)(U / A) nonamers. Molecules containing this type of ARE include GM-CSF and TNF-a. Class III AREs are less well defined. These U-rich regions do not contain the AUUUA motif. c-Jun and myogenin are two well-studied examples of this class. While most proteins that bind to AREs are known to destabilize messengers, members of the ELAV family, most notably HuR, have been documented to increase mRNA stability. HuR binds to all three classes of AREs. Engineering a HuR-specific binding site into the 3'UTR of a nucleic acid molecule results in HuR binding and thus stabilization of the message in vivo. The introduction, removal or modification of 3'UTR AU-rich elements (AREs) can be used to modulate mRNA stability. For example, one or more copies of AREs can be introduced to render the mRNA less stable, thereby reducing translation and decreasing production of the resulting protein. Alternatively, AREs can be identified and removed or mutated to increase intracellular stability and thus increase translation and production of the resulting protein.

[0280] In some embodiments, mRNA stability and protein production may be improved in a particular organ and / or tissue by the introduction of a feature that is often expressed in genes of the target organ. As a non-limiting example, the feature may be a UTR. As another example, the feature may be an intron or a portion of an intron sequence.

[0281] 5' Capping The 5' cap structure of mRNA is involved in nuclear export, increases mRNA stability, and binds mRNA cap-binding protein (CBP), which is responsible for mRNA stability in the cell and translational competence through the association of CBP with poly(A)-binding protein to form mature circular mRNA species. The cap also assists in the removal of 5' proximal introns during mRNA splicing.

[0282] Endogenous mRNA molecules can be capped at the 5' end, generating a 5'-ppp-5'-triphosphate linkage between the terminal guanosine cap residue and the transcribed sense nucleotide at the 5' end of the mRNA molecule. This 5'-guanylate cap can then be methylated to generate an N7-methyl-guanylate residue. The ribose sugars of terminal and / or non-terminal transcribed nucleotides at the 5' end of the mRNA can also be optionally 2'-0-methylated. 5'-decapping via hydrolysis and cleavage of the guanylate cap structure can target nucleic acid molecules, e.g., mRNA molecules, for degradation.

[0283] Modifications to mRNA can generate non-hydrolyzable cap structures that prevent decapping, thus increasing mRNA half-life. Because hydrolysis of the cap structure requires cleavage of the 5'-ppp-5' phosphorodiester bond, modified nucleotides can be used during the capping reaction. For example, Vaccinia Capping Enzyme from New England Biolabs (Ipswich, MA) can be used with a-thio-guanosine nucleotides according to the manufacturer's instructions to generate phosphorothioate bonds in the 5'-ppp-5' cap.

[0284] Additional modified guanosine nucleotides, such as a-methyl-phosphonate and seleno-phosphate nucleotides, may be used.

[0285] Additional modifications include, but are not limited to, 2'-0-methylation of the ribose sugar of the 5'-terminus and / or 5'-non-terminal nucleotides of an mRNA on the 2'-hydroxyl group of the sugar ring (described above). A number of different 5'-cap structures can be used to generate the 5'-cap of a nucleic acid molecule, such as an mRNA molecule.

[0286] Cap analogs, also referred to herein as synthetic cap analogs, chemical caps, chemical cap analogs, or structural or functional cap analogs, differ in their chemical structure from the native (i.e., endogenous, wild-type or physiological) 5'-cap while retaining cap function. Cap analogs can be chemically (i.e., non-enzymatically) or enzymatically synthesized and / or linked to a nucleic acid molecule.

[0287] For example, the anti-reverse cap analog (ARCA) cap contains two guanines linked by a 5'-5'-triphosphate group, one of which has an N7 methyl group as well as a 3'-0-methyl group (i.e., N7,3'-0-dimethyl-guanosine-5'-triphosphate-5'-guanosine (m 7 The 3'-0 atom of the other unmodified guanine is linked to the 5' terminal nucleotide of the capped nucleic acid molecule (e.g., mRNA). The N7- and 3'-0-methylated guanine provides a terminal site for the capped nucleic acid molecule (e.g., mRNA).

[0288] Another exemplary cap is mCAP, which is similar to ARCA but has a 2'-O-methyl group on the guanosine (i.e., N7,2'-O-dimethyl-guanosine-5'-triphosphate-5'-guanosine, mCAP). 7 Gm-ppp-G).

[0289] Although cap analogs allow for concomitant capping of nucleic acid molecules in in vitro transcription reactions, up to 20% of the transcripts may remain uncapped. This, as well as the structural differences of cap analogs from the endogenous 5'-cap structures of nucleic acids produced by the endogenous cellular transcription machinery, may result in reduced translational competence and reduced cellular stability.

[0290] mRNA can also be post-transcriptionally capped using enzymes to generate a more authentic 5'-cap structure. As used herein, the phrase "more authentic" refers to characteristics that closely reflect or mimic endogenous or wild-type characteristics, either structurally or functionally. That is, a "more authentic" characteristic is more representative of an endogenous, wild-type, native or physiological cellular function and / or structure compared to prior art synthetic characteristics or analogs, etc., or outperforms the corresponding endogenous, wild-type, native or physiological characteristics in one or more respects. Non-limiting examples of more authentic 5' cap structures are those that have improved binding of cap-binding proteins, increased half-life, reduced susceptibility to 5' endonucleases and / or reduced 5' decapping, among others, when compared to synthetic 5' cap structures (or wild-type, native or physiological 5' cap structures) known in the art. For example, recombinant vaccinia virus capping enzyme and recombinant 2'-0-methyltransferase enzyme can generate a standard 5'-5'-triphosphate bond between the 5'-terminal nucleotide of the mRNA and the guanine cap nucleotide, where the cap guanine contains an N7 methylation and the 5'-terminal nucleotide of the mRNA contains a 2'-0-methyl. Such a structure is referred to as a Capl structure. This cap results in higher translational competence and cellular stability as well as reduced activation of cellular proinflammatory cytokines, for example, when compared to other 5' cap analog structures known in the art. The cap structure contains 7 mg (5 * )ppp(5 * )N,pN2p(cap 0),7mg(5 * )ppp(5 *)NlmpNp (cap 1), and 7 mg (5 * )-ppp(5')NlmpN2mp (Cap 2).

[0291] In some embodiments, the 5' end cap may include an endogenous cap or a cap analog.

[0292] In some embodiments, the 5' end cap may comprise a guanine analog. Useful guanine analogs include, but are not limited to, inosine, Nl-methyl-guanosine, 2'fluoro-guanosine, 7-deaza-guanosine, 8-oxo-guanosine, 2-amino-guanosine, LNA-guanosine, and 2-azido-guanosine.

[0293] IRES sequence In some embodiments, the mRNA may contain an internal ribosome entry site (IRES). IRES, first identified as a characteristic picornavirus RNA, plays an important role in initiating protein synthesis in the absence of a 5' cap structure. IRES may act as the only ribosome binding site or may function as one of multiple ribosome binding sites of the mRNA. An mRNA containing multiple functional ribosome binding sites may code for several peptides or polypeptides that are translated independently by ribosomes. Non-limiting examples of IRES sequences that may be used include, but are not limited to, those from picornaviruses (e.g., FMDV), plague viruses (CFFV), polioviruses (PV), encephalomyocarditis viruses (ECMV), foot and mouth disease viruses (FMDV), hepatitis C viruses (HCV), classical swine fever viruses (CSFV), murine leukemia viruses (MLV), simian immunodeficiency viruses (SIV) or cricket paralysis viruses (CrPV).

[0294] Poly-A tail During RNA processing, a long chain of adenine nucleotides (poly-A tail) can be added to polynucleotides, such as mRNA molecules, to increase stability. Immediately after transcription, the 3' end of the transcript can be cleaved to free the 3' hydroxyl. Poly-A polymerase then adds a chain of adenine nucleotides to RA. The process, called polyadenylation, adds a certain length of poly-A tail.

[0295] In some embodiments, the length of the poly-A tail is greater than 30 nucleotides in length, in other embodiments, the poly-A tail is greater than 35 nucleotides in length (e.g., at least about 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,500, and 3,000 nucleotides or more). In some embodiments, the mRNA comprises from about 30 to about 3,000 nucleotides (e.g., 30 to 50, 30 to 100, 30 to 250, 30 to 500, 30 to 750, 30 to 1,000, 30 to 1,500, 30 to 2,000, 30 to 2,500, 50 to 100, 50 to 250, 50 to 500, 50 to 750, 50 to 1,000, 50 to 1,500, 50 to 2,000, 50 to 2,500, 50 to 3,000, 100 to 500, 100 to 750, 100 to 1,000, 100 to 1,500, 10 Poly-A tails of 0-2,000, 100-2,500, 100-3,000, 500-750, 500-1,000, 500-1,500, 500-2,000, 500-2,500, 500-3,000, 1,000-1,500, 1,000-2,000, 1,000-2,500, 1,000-3,000, 1,500-2,000, 1,500-2,500, 1,500-3,000, 2,000-3,000, 2,000-2,500, and 2,500-3,000.

[0296] In some embodiments, the poly-A tail is designed relative to the length of the entire mRNA, which may be based on the length of the region encoding the target of interest, the length of a particular feature or region (e.g., a flanking region), or based on the length of the final product expressed from the mRNA.

[0297] In this context, the poly-A tail may be 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% longer in length than the mRNA or a feature thereof. The poly-A tail may also be designed as a fraction of the mRNA to which it belongs. In this context, the poly-A tail may be 10, 20, 30, 40, 50, 60, 70, 80, or 90% or more of the total length of the construct or the total length of the construct excluding the poly-A tail. Additionally, engineered binding sites and conjugation of the mRNA for poly-A binding proteins may improve expression.

[0298] Also, multiple different mRNAs can be linked together via their 3' ends to PABP (poly-A binding protein) using modified nucleotides at the 3' end of the poly-A tail. Transfection experiments can be performed in relevant cell lines and protein production can be assayed by ELISA at 12 hours, 24 hours, 48 ​​hours, 72 hours and 7 days after transfection.

[0299] In some embodiments, the mRNA is designed to contain a poly-AG quartet. A G-quartet is a cyclic hydrogen-bonded array of four guanine nucleotides that can be formed by G-rich sequences in both DNA and RNA. In this embodiment, the G-quartet is incorporated at the end of a poly-A tail.

[0300] Stop codon In some embodiments, the mRNA may include one stop codon. In some embodiments, the mRNA may include two stop codons. In some embodiments, the mRNA may include three stop codons. In some embodiments, the mRNA may include at least one stop codon. In some embodiments, the mRNA may include at least two stop codons. In some embodiments, the mRNA may include at least three stop codons. As non-limiting examples, the stop codons may be selected from TGA, TAA, and TAG.

[0301] In some embodiments, the mRNA includes the stop codon TGA and one additional stop codon. In further embodiments, the additional stop codon may be TAA.

[0302] Circular RNA (oRNA) In some embodiments, the originator construct and / or the benchmark construct is a circular RNA (oRNA). As used herein, the terms "oRNA" or "circular RNA" are used interchangeably and may refer to an RNA that forms a circular structure via covalent or non-covalent bonds.

[0303] In some embodiments, the oRNA may be non-immunogenic in mammals (eg, humans, non-human primates, rabbits, rats, and mice).

[0304] In some embodiments, the oRNA may be capable of replicating or does replicate in aquaculture animals (e.g., fish, crabs, shrimp, oysters, etc.), mammalian cells, cells from pet or zoo animals (e.g., cats, dogs, lizards, birds, lions, tigers, and bears, etc.), cells from farm or work animals (e.g., horses, cows, pigs, chickens, etc.), human cells, cultured cells, primary cells or cell lines, stem cells, progenitor cells, differentiated cells, embryonic cells, cancer cells (e.g., tumorigenic, metastatic), non-tumorigenic cells (e.g., normal cells), fetal cells, embryonic cells, adult cells, dividing cells, non-dividing cells, or any combination thereof.

[0305] In some embodiments, the oRNA has a half-life at least that of its linear counterpart. In some embodiments, the oRNA has an increased half-life over that of its linear counterpart. In some embodiments, the half-life is increased by about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more. In some embodiments, the oRNA has a half-life or persistence in cells of at least about 1 hour to about 30 days, or at least about 2 hours, 6 hours, 12 hours, 18 hours, 24 hours (1 day), 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 8 days, 9 days, 10 days, 11 days, 12 days, 13 days, 14 days, 15 days, 16 days, 17 days, 18 days, 19 days, 20 days, 21 days, 22 days, 23 days, 24 days, 25 days, 26 days, 27 days, 28 days, 29 days, 30 days, 60 days, or longer or any time in between. In some embodiments, the oRNA has a half-life or persistence in cells of about 10 minutes to about 7 days or less, or about 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 24 hours (1 day), 36 hours (1.5 days), 48 hours (2 days), 60 hours (2.5 days), 72 hours (3 days), 4 days, 5 days, 6 days, or 7 days or less.

[0306] In some embodiments, the oRNA has a half-life or persistence in a cell while the cell is dividing, hi some embodiments, the oRNA has a half-life or persistence in a cell after dividing. In certain embodiments, the oRNA has a half-life or persistence in dividing cells of about 10 minutes to greater than about 30 days, or at least about 10 minutes, 15 minutes, 30 minutes, 45 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 24 hours (1 day), 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 8 days, 9 days, 10 days, 11 days, 12 days, 13 days, 14 days, 15 days, 16 days, 17 days, 18 days, 19 days, 20 days, 21 days, 22 days, 23 days, 24 days, 25 days, 26 days, 27 days, 28 days, 29 days, 30 days, 60 days, or longer or any time in between.

[0307] In some embodiments, the oRNA modulates cellular function, e.g., transiently or long term. In certain embodiments, the cellular function is stably altered, such as modulation lasting for at least about 1 hour to about 30 days, or at least about 2 hours, 6 hours, 12 hours, 18 hours, 24 hours (1 day), 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 8 days, 9 days, 10 days, 11 days, 12 days, 13 days, 14 days, 15 days, 16 days, 17 days, 18 days, 19 days, 20 days, 21 days, 22 days, 23 days, 24 days, 25 days, 26 days, 27 days, 28 days, 29 days, 30 days, 60 days, or longer. In certain embodiments, the cellular function is temporarily altered, such as, for example, a modulation lasting for about 30 minutes to about 7 days, or for about 30 minutes, 45 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 23 hours, 24 hours (1 day), 36 hours (1.5 days), 48 hours (2 days), 60 hours (2.5 days), 72 hours (3 days), 4 days, 5 days, 6 days, or 7 days or less.

[0308] In some embodiments, the oRNA is at least about 20 nucleotides, at least about 30 nucleotides, at least about 40 nucleotides, at least about 50 nucleotides, at least about 75 nucleotides, at least about 100 nucleotides, at least about 200 nucleotides, at least about 300 nucleotides, at least about 400 nucleotides, at least about 500 nucleotides, at least about 1,000 nucleotides, at least about 2,000 nucleotides, at least about 5,000 nucleotides, at least about 6,000 nucleotides, at least about 7,000 nucleotides, at least about 8,000 nucleotides, at least about 9,000 nucleotides, at least about 10,000 nucleotides, at least about 12,000 nucleotides, at least about 14,000 nucleotides, at least about 15,000 nucleotides, at least about 16,000 nucleotides, at least about 17,000 nucleotides, at least about 18,000 nucleotides, at least about 19,000 nucleotides, or at least about 20,000 nucleotides. In some embodiments, the oRNA may be of sufficient size to provide a binding site for a ribosome.

[0309] In some embodiments, the maximum size of the oRNA may be limited by the ability to package and deliver the RNA to a target. In some embodiments, the size of the oRNA is sufficient to encode a polypeptide, and thus lengths of at least 20,000 nucleotides, at least 15,000 nucleotides, at least 10,000 nucleotides, at least 7,500 nucleotides, or at least 5,000 nucleotides, at least 4,000 nucleotides, at least 3,000 nucleotides, at least 2,000 nucleotides, at least 1,000 nucleotides, at least 500 nucleotides, at least 400 nucleotides, at least 300 nucleotides, at least 200 nucleotides, at least 100 nucleotides may be useful.

[0310] In some embodiments, the oRNA comprises one or more elements described elsewhere herein. In some embodiments, the elements may be separated from each other by a spacer sequence or linker. In some embodiments, the elements may be separated from each other by 1 nucleotide, 2 nucleotides, about 5 nucleotides, about 10 nucleotides, about 15 nucleotides, about 20 nucleotides, about 30 nucleotides, about 40 nucleotides, about 50 nucleotides, about 60 nucleotides, about 80 nucleotides, about 100 nucleotides, about 150 nucleotides, about 200 nucleotides, about 250 nucleotides, about 300 nucleotides, about 400 nucleotides, about 500 nucleotides, about 600 nucleotides, about 700 nucleotides, about 800 nucleotides, about 900 nucleotides, about 1000 nucleotides, up to about 1 kb, at least about 1000 nucleotides.

[0311] In some embodiments, one or more elements are contiguous to one another, eg, lacking a spacer element.

[0312] In some embodiments, one or more elements are structurally flexible, hi some embodiments, the structural flexibility is due to a sequence that is substantially free of secondary structure.

[0313] In some embodiments, the oRNA comprises a secondary or tertiary structure that provides binding sites for ribosomes, translation, or rolling circle translation.

[0314] In some embodiments, the oRNA comprises specific sequence features. For example, the oRNA may comprise a specific nucleotide composition. In some such embodiments, the oRNA may comprise one or more purine-rich regions (adenine or guanosine). In some such embodiments, the oRNA may comprise one or more purine-rich regions (adenine or guanosine). In some embodiments, the oRNA may comprise one or more AU-rich regions or elements (AREs). In some embodiments, the oRNA may comprise one or more adenine-rich regions.

[0315] In some embodiments, the oRNA comprises one or more of the modifications described elsewhere herein.

[0316] In some embodiments, the oRNA comprises one or more expression sequences and is configured for persistent expression in a cell of a subject in vivo. In some embodiments, the oRNA is configured such that expression of the one or more expression sequences in a cell at a later time point is equal to or faster than an earlier time point. In such embodiments, expression of the one or more expression sequences may be maintained at a relatively stable level or may increase over time. Expression of the expression sequences may be relatively stable for an extended period of time. For example, in some cases, expression of the one or more expression sequences in a cell does not decrease by 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% over a period of at least 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 23 or more days. In some cases, expression of one or more expression sequences in the cells is maintained at a level that does not change by more than 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% for at least 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 23 days or more.

[0317] Controllability In some embodiments, the oRNA comprises a regulatory element. As used herein, a "regulatory element" is a sequence that modifies the expression of an expression sequence. A regulatory element may include a sequence that is adjacent to a payload or cargo region. A regulatory element may be operably linked to a payload or cargo region.

[0318] In some embodiments, a regulatory element may increase the amount of payload or cargo expressed when compared to the amount expressed in the absence of the regulatory element. As a non-limiting example, a regulatory element may increase the amount of payload or cargo expressed for multiple payload or cargo sequences linked in tandem.

[0319] In some embodiments, a regulatory element may include a sequence for selectively initiating or activating translation of a payload or cargo.

[0320] In some embodiments, the regulatory element may include a sequence to initiate degradation of the oRNA or payload or cargo. Non-limiting examples of sequences to initiate degradation include, but are not limited to, riboswitches, aptazymes, and miRNA binding sites.

[0321] In some embodiments, the regulatory element may regulate the translation of the payload or cargo in the oRNA. Regulation may produce an increase (enhancer) or a decrease (suppressor) in the payload or cargo. The regulatory element may be located adjacent to the payload or cargo (e.g., on one or both sides of the payload or cargo).

[0322] In some embodiments, the translation initiation sequence functions as a regulatory element. In some embodiments, the translation initiation sequence comprises an AUG / ATG codon. In some embodiments, the translation initiation sequence comprises any eukaryotic initiation codon, including but not limited to, AUG / ATG, CUG / CTG, GUG / GTG, UUG / TTG, ACG, AUC / ATC, AUU, AAG, AUA / ATA, or AGG. In some embodiments, the translation initiation sequence comprises a Kozak sequence. In some embodiments, translation initiates under selective conditions, e.g., stress-inducing conditions, at an alternative translation initiation sequence, e.g., a translation initiation sequence other than an AUG / ATG codon. As a non-limiting example, circular polyribonucleotide translation may initiate at an alternative translation initiation sequence, e.g., ACG. As another non-limiting example, circular polyribonucleotide translation may initiate at an alternative translation initiation sequence, CUG / CTG. As another non-limiting example, translation may initiate at an alternative translation initiation sequence, GUG / GTG. As yet another non-limiting example, translation can initiate at repeat-associated non-AUG (RAN) sequences, such as alternative translation initiation sequences that contain short stretches of repetitive RNA, e.g., CGG, GGGGCC, CAG, CTG.

[0323] Masking Agent Masking any of the nucleotides adjacent to the codon that initiates translation can be used to modify the location of translation initiation, translation efficiency, length and / or structure of the oRNA. In some embodiments, a masking agent can be used near the start codon or alternative start codon to mask or hide the codon to reduce the likelihood of translation initiation at the masked start codon or alternative start codon. Non-limiting examples of masking agents include antisense locked nucleic acid (LNA) oligonucleotides and exon junction complexes (EJCs). In some embodiments, a masking agent can be used to mask the start codon of the oRNA to increase the likelihood that translation will initiate at the alternative start codon.

[0324] Translation initiation sequence In some embodiments, the oRNA encodes a polypeptide or peptide and may include a translation initiation sequence. The translation initiation sequence may include, but is not limited to, an initiation codon, a non-coding initiation codon, a Kozak sequence, or a Shine-Dalgarno sequence. The translation initiation sequence may be located adjacent to the payload or cargo (e.g., on one or both sides of the payload or cargo).

[0325] In some embodiments, the translation initiation sequence provides structural flexibility to the oRNA, hi some embodiments, the translation initiation sequence is within a substantially single-stranded region of the oRNA.

[0326] An oRNA can include more than one start codon, for example, but not limited to, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, or more than 15 start codons. Translation can begin at the first start codon or can begin downstream of the first start codon.

[0327] In some embodiments, the oRNA may initiate at the first start codon, e.g., a codon that is not AUG. Translation of the circular polyribonucleotide may initiate at an alternative translation initiation sequence, e.g., but not limited to, ACG, AGG, AAG, CUG / CTG, GUG / GTG, AUA / ATA, AUU / ATT, UUG / TTG. In some embodiments, translation may initiate at an alternative translation initiation sequence under selective conditions, e.g., stress-inducing conditions. As a non-limiting example, oRNA translation may initiate at an alternative translation initiation sequence, e.g., ACG. As another non-limiting example, oRNA translation may initiate at an alternative translation initiation sequence, CUG / CTG. As yet another non-limiting example, oRNA translation may initiate at an alternative translation initiation sequence, GTG / GUG. As yet another non-limiting example, an oRNA may initiate translation at a repeat-associated non-AUG (RAN) sequence, such as an alternative translation initiation sequence that contains short stretches of repeated RNA, e.g., CGG, GGGGCC, CAG, CTG.

[0328] IRES sequence In some embodiments, the oRNA described herein comprises an internal ribosome entry site (IRES) element capable of engaging a eukaryotic ribosome. In some embodiments, the IRES element is at least about 5 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 40 nucleotides, at least about 50 nucleotides, at least about 100 nucleotides, at least about 200 nucleotides, at least about 250 nucleotides, at least about 350 nucleotides, or at least about 500 nucleotides. In one embodiment, the IRES element is derived from DNA including, but not limited to, organisms, viruses, mammals, and Drosophila. Such viral DNA can be derived from, but is not limited to, picornavirus complementary DNA (cDNA), encephalomyocarditis virus (EMCV) cDNA, and poliovirus cDNA. In one embodiment, the Drosophila DNA from which the IRES element is derived includes, but is not limited to, the antennapedia gene from Drosophila melanogaster.

[0329] In some embodiments, the IRES element is derived at least in part from a virus, e.g., it is a viral IRES element, e.g., ABPV_IGRpred, AEV, ALPV_IGRpred, BQCV_IGRpred, BVDV1_1-385, BVDV1_29-391, CrPV_5NCR, CrPV_IGR, crTMV_IREScp, crTMV_IRESmp75, crTMV_IRESmp228, crTMV_IREScp, crTMV_IREScp, CSFV, CVB3, DCV_IGR, EMCV-R, EoPV_5NTR, ERAV 245-961, ERBV 162-920, EV71_1-748, FeLV-Notch2, FMDV_type_C, GBV-A, GBV-B, GBV-C, gypsy_env, gypsyD5, gypsyD2, HAV_HM175, HCV_ty pe_1a, HiPV_IGRpred, HIV-1, HoCV1_IGRpred, HRV-2, IAPV_IGRpred, idefix, KBV_IGRpred, LINE-1_ORF1_-101_~_-1, LINE- 1_ORF1-302_~_-202, LINE-1_ORF2-138_~_-86, LINE-1_ORF1_-44~_-1, PSIV_IGR, PV_type1_Mahoney, PV_type3_Leon, REV-A, RhPV_5NCR, RhPV_IGR, SINV1_IGRpred, SV40_661-830, TMEV, TMV_UI_IRESmp228, TRV_5NTR, TrV_IGR, or TSV_IGR.In some embodiments, the IRES element is a cellular IRES, such as AML1 / RUNX1, Antp-D, Antp-DE, Antp-CDE, Apaf-1, Apaf-1, AQP4, AT1R_var1, AT1R_var2, AT1R_var3, AT1R_var4, BAG1_p36 delta 236 nt, BAG1_p36, BCL2, BiP_-222_-3, c-IAP1_285-1399, c-IAP1_1313-1462, c-jun, c-my c, Cat-1224, CCND1, DAPS, eIF4G, eIF4GI-ext, eIF4GII, eIF4GII-long, ELG1, ELH, FGF1 A, FMR1, Gtx-133-141, Gtx-1-166, Gtx-1-120, Gtx-1-196, hairless, HAP4, HIF1a, hSN M1, Hsp101, hsp70, hsp70, Hsp90, IGF2_leader2, Kv1.4_1.2, L-myc, LamB1_-335_-1, LE derived at least in part from F1, MNT_75-267, MNT_36-160, MTG8a, MYB, MYT2_997-1152, n-MYC, NDST1, NDST2, NDST3, NDST4L, NDST4S, NRF_-653_-17, NtHSF1, ODC1, p27kip1, 03_128-269, PDGF2 / c-sis, Pim-1, PITSLRE_p58, Rbm3, reaper, Scamper, TFIID, TIF4631, Ubx_1-966, Ubx_373-961, UNR, Ure2, UtrA, VEGF-A-133-1, XIAP_5-464, XIAP_305-466, or YAP1.

[0330] terminating element In some embodiments, an oRNA includes one or more cargo or payload sequences (also referred to as expression sequences), each of which may or may not have a termination element.

[0331] In some embodiments, the oRNA comprises one or more cargo or payload sequences, and the sequences lack a termination element such that the oRNA is translated sequentially. Elimination of the termination element can result in rolling circle translation or sequential expression of the encoded peptide or polypeptide, as the ribosome does not stall or fall off. In such embodiments, rolling circle translation expresses sequential expression through each cargo or payload sequence.

[0332] In some embodiments, one or more of the cargo or payload sequences in the oRNA comprises a termination element.

[0333] In some embodiments, not all of the cargo or payload sequences in the oRNA contain a termination element. In such instances, the cargo or payload may leave the ribosome when the ribosome encounters a termination element and terminates translation. In some embodiments, translation is terminated but at least one region of the ribosome remains in contact with the oRNA.

[0334] Rolling Circle Translation In some embodiments, once translation of the oRNA is initiated, the ribosome bound to the oRNA does not leave the oRNA before completing at least one round of translation of the oRNA, hi some embodiments, the oRNAs described herein are suitable for rolling circle translation. In some embodiments, during rolling circle translation, once translation of the oRNA is initiated, a ribosome bound to the oRNA does not leave the oRNA before completing at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 500, at least 1000, at least 1500, at least 2000, at least 5000, at least 10000, at least 10 to the power of 5, or at least 10 to the power of 6 rounds of translation of the oRNA.

[0335] In some embodiments, rolling circle translation of the oRNA results in the production of a polypeptide that is translated from multiple rounds of translation of the oRNA, hi some embodiments, the oRNA comprises a stagger element and rolling circle translation of the oRNA results in the production of a polypeptide product that is generated from one round of translation or less than one round of translation of the oRNA.

[0336] Circularization In one embodiment, the linear RNA can be circularized or concatemerized. In some embodiments, the linear RNA can be circularized in vitro prior to formulation and / or delivery. In some embodiments, the linear RNA can be circularized intracellularly.

[0337] In some embodiments, the mechanism of circularization or concatemerization can occur via at least three different routes: 1) chemical, 2) enzymatic, and 3) ribozyme-catalyzed. The newly formed 5'- / 3'-bonds can be intramolecular or intermolecular.

[0338] In the first pathway, the 5' and 3' ends of a nucleic acid contain chemically reactive groups that, when brought close to each other, form a new covalent bond between the 5' and 3' ends of the molecule. The 5' end may contain an NHS-ester reactive group and the 3' end may contain a 3'-amino terminated nucleotide, such that in organic solvent the 3' amino terminated nucleotide on the 3' end of the synthetic mRNA molecule will undergo nucleophilic attack on the 5'-NHS-ester site to form a new 5' / 3'-amide bond.

[0339] In the second pathway, T4 RNA ligase can be used to enzymatically ligate a 5'-phosphorylated nucleic acid molecule to the 3'-hydroxyl group of a nucleic acid, forming a new phosphorodiester bond. In an exemplary reaction, a nucleic acid molecule is incubated with 1-10 units of T4 RNA ligase (New England Biolabs, Ipswich, MA) for 1 hour at 37°C according to the manufacturer's protocol. The ligation reaction can occur in the presence of a split oligonucleotide capable of base pairing with both the 5' and 3' regions in parallel to aid in the enzymatic ligation reaction.

[0340] In the third pathway, during in vitro transcription, either the 5' or 3' end of the cDNA template encodes a ligase ribozyme sequence, such that the resulting nucleic acid molecule may contain an active ribozyme sequence capable of ligating the 5' end of the nucleic acid molecule to the 3' end of the nucleic acid molecule. The ligase ribozyme may be derived from a group I intron, a group I intron, hepatitis delta virus, a hairpin ribozyme, or may be selected by SELEX (Systematic Evolution of Ligands by Exponential Enrichment). The ribozyme ligase reaction may be carried out at a temperature of 0-37°C for 1-24 hours.

[0341] In some embodiments, oRNA is generated via circularization of linear RNA.

[0342] Extracellular cyclization In some embodiments, linear RNA is circularized or concatemerized using chemical methods to form oRNA. In some chemical methods, the 5' and 3' ends of a nucleic acid (e.g., linear RNA) contain chemically reactive groups that, when brought close to each other, can form a new covalent bond between the 5' and 3' ends of the molecule. The 5' end can contain an NHS-ester reactive group and the 3' end can contain a 3'-amino terminated nucleotide, such that in organic solvents, the 3' amino terminated nucleotide on the 3' end of the linear RNA will undergo nucleophilic attack on the 5'-NHS-ester site to form a new 5' / 3'-amide bond.

[0343] In one embodiment, a DNA or RNA ligase can be used to enzymatically ligate a 5'-phosphorylated nucleic acid molecule (e.g., linear RNA) to the 3'-hydroxyl group of a nucleic acid (e.g., linear nucleic acid) forming a new phosphorodiester bond. In an exemplary reaction, linear RNA is incubated with 1-10 units of T4 RNA ligase for 1 hour at 37C according to the manufacturer's protocol. The ligation reaction can occur in the presence of a linear nucleic acid capable of base pairing with both the 5' and 3' regions in parallel to aid in the enzymatic ligation reaction. In one embodiment, the ligation is a splint ligation, where a single-stranded polynucleotide (splint), such as a single-stranded RNA, can be designed to hybridize to both ends of the linear RNA such that the two ends are juxtaposed by hybridization with the single-stranded splint. Thus, a splint ligase can catalyze the ligation of two ends of a linear RNA in parallel to generate an oRNA.

[0344] In one embodiment, a DNA or RNA ligase may be used in the synthesis of the oRNA. As a non-limiting example, the ligase may be a circ ligase or a circular ligase.

[0345] In one embodiment, either the 5' or 3' end of the linear RNA may encode a ligase ribozyme sequence such that during in vitro transcription, the resulting linear RNA contains an active ribozyme sequence capable of ligating the 5' end of the linear RNA to the 3' end of the linear RNA. The ligase ribozyme may be derived from a group I intron, hepatitis delta virus, a hairpin ribozyme, or may be selected by SELEX (Systematic Evolution of Ligands by Exponential Enrichment).

[0346] In one embodiment, the linear RNA can be circularized or concatemerized by using at least one non-nucleic acid site. In one aspect, the at least one non-nucleic acid site can react with a region or feature near the 5' end and / or near the 3' end of the linear RNA to circularize or link the linear RNA. In another aspect, the at least one non-nucleic acid site can be located or linked at or near the 5' end and / or 3' end of the linear RNA. The contemplated non-nucleic acid site can be homogeneous or heterogeneous. As a non-limiting example, the non-nucleic acid site can be a bond, such as a hydrophobic linkage, an ionic linkage, a biodegradable linkage, and / or a cleavable linkage. As another non-limiting example, the non-nucleic acid site is a ligation site. As yet another non-limiting example, the non-nucleic acid site can be an oligonucleotide or peptide site, such as an aptamer or a non-nucleic acid linker as described herein.

[0347] In one embodiment, linear RNAs can be circularized or concatemerized at the 5' and 3' ends of linear RNAs by atoms near or connected to them, non-nucleic acid moieties that cause attraction between molecular surfaces. As a non-limiting example, one or more linear RNAs can be circularized or concatemerized by intermolecular or intramolecular forces. Non-limiting examples of intermolecular forces include dipole-dipole forces, dipole-induced dipole forces, induced dipole-induced dipole forces, van der Waals forces, and London dispersion forces. Non-limiting examples of intramolecular forces include covalent bonds, metallic bonds, ionic bonds, resonance bonds, agnostic bonds, dipolar bonds, conjugation, hyperconjugation, and antibonds.

[0348] In one embodiment, the linear RNA may contain ribozyme RNA sequences near the 5' end and near the 3' end. The ribozyme RNA sequence may be covalently linked to a peptide when the sequence is exposed to the remainder of the ribozyme. In one aspect, the peptides covalently linked to the ribozyme RNA sequence near the 5' end and near the 3' end may associate with each other to cause the linear RNA to circularize or concatemerize. In another aspect, the peptides covalently linked to the ribozyme RNA sequence near the 5' end and near the 3' end may cause the linear RNA to circularize or concatemerize after being subjected to ligation using various methods known in the art, such as, but not limited to, protein ligation.

[0349] In some embodiments, the linear RNA may include a 5' triphosphate of a nucleic acid that has been converted to a 5' monophosphate, for example, by contacting the 5' triphosphate with RNA 5' pyrophosphohydrolase (RppH) or ATP diphosphohydrolase (apyrase). Alternatively, converting the 5' triphosphate of the linear RNA to a 5' monophosphate may occur in a two-step reaction that includes (a) contacting the 5' nucleotide of the linear RNA with a phosphatase (e.g., Antarctic phosphatase, shrimp alkaline phosphatase, or calf intestinal phosphatase) to remove all three phosphates; and (b) contacting the 5' nucleotide after step (a) with a kinase (e.g., polynucleotide kinase) that adds a single phosphate.

[0350] In some embodiments, RNA can be circularized using the methods described in WO2017222911 and WO2016197121, the contents of each of which are incorporated by reference herein in their entirety.

[0351] In some embodiments, the RNA can be circularized, for example, by backsplicing a non-mammalian exogenous intron or splint ligation of the 5' and 3' ends of a linear RNA. In one embodiment, the circular RNA is generated from a recombinant nucleic acid encoding the target RNA to be circularized. As a non-limiting example, the method includes: a) generating a recombinant nucleic acid encoding the target RNA to be circularized, the recombinant nucleic acid including, in 5' to 3' order, i) a 3' portion of the exogenous intron including a 3' splice site, ii) a nucleic acid sequence encoding the target RNA, and iii) a 5' portion of the exogenous intron including a 5' splice site; b) performing transcription, whereby an RNA is generated from the recombinant nucleic acid; and c) performing splicing of the RNA, whereby the RNA is circularized to generate an oRNA.

[0352] Without wishing to be bound by theory, circular RNAs generated with exogenous introns are recognized by the immune system as "non-self" and elicit an innate immune response, whereas circular RNAs generated with endogenous introns are recognized by the immune system as "self" and typically do not elicit an innate immune response, even if they harbor exons containing foreign RNA.

[0353] Thus, circular RNAs can be generated with either endogenous or exogenous introns to control immunological self / non-self discrimination as desired. Numerous intron sequences are known from a variety of organisms and viruses, including sequences derived from genes encoding proteins, ribosomal RNA (rRNA), or transfer RNA (tRNA).

[0354] Circular RNA can be generated from linear RNA in a number of ways. In some embodiments, circular RNA is generated from linear RNA by backsplicing of a downstream 5' splice site (splice donor) to an upstream 3' splice site (splice acceptor). Circular RNA can be generated in this manner by any non-mammalian splicing method. For example, linear RNA containing various types of introns can be circularized, including self-splicing group I introns, self-splicing group II introns, spliceosomal introns, and tRNA introns. In particular, group I and group II introns have the advantage that they can be easily used to generate circular RNA in vitro and in vivo due to their ability to undergo self-splicing by their autocatalytic ribozyme activity.

[0355] In some embodiments, circular RNA can be generated in vitro from linear RNA by chemical or enzymatic ligation of the 5' and 3' ends of RNA. Chemical ligation can be performed, for example, using cyanogen bromide (BrCN) or ethyl-3-(3'-dimethylaminopropyl)carbodiimide (EDC) for activation of nucleotide phosphomonoester groups to allow phosphodiester bond formation. See, for example, Sokolova (1988) FEBS Lett 232:153-155; Dolinnaya et al. (1991) Nucleic Acids Res., 19:3067-3072; Fedorova (1996) Nucleosides Nucleotides Nucleic Acids 15:1 137-1 147 (herein incorporated by reference). Alternatively, enzymatic ligation can be used to circularize RNA. Exemplary ligases that may be used include T4 DNA ligase (T4 Dnl), T4 RNA ligase 1 (T4 Rnl 1), and T4 RNA ligase 2 (T4 Rnl 2).

[0356] In some embodiments, splint ligation, using an oligonucleotide splint that hybridizes to the two ends of a linear RNA, can be used to ligate the ends of the linear RNA together. Hybridization of the splint, which can be either DNA or RNA, orients the 5'-phosphate and 3'-OH of the RNA end for ligation. Subsequent ligation can be performed using either chemical or enzymatic techniques, as described above. Enzymatic ligation can be performed, for example, with T4 DNA ligase (a DNA splint is required), T4 RNA ligase 1 (an RNA splint is required) or T4 RNA ligase 2 (a DNA or RNA splint). For example, chemical ligation with BrCN or EDC can be more efficient than enzymatic ligation in some cases, when the structure of the hybridized splint-RNA complex interferes with enzyme activity.

[0357] In some embodiments, the oRNA may further comprise an internal ribosome entry site (IRES) operably linked to the RNA sequence encoding the polypeptide. The inclusion of the IRES allows for translation of one or more open reading frames from the circular RNA. The IRES element attracts the eukaryotic ribosomal translation initiation complex and promotes translation initiation. See, e.g., Kaufman et al., Nuc. Acids Res. (1991) 19:4485-4490; Gurtu et al., Biochem. Biophys. Res. Comm. (1996) 229:295-298; Rees et al., BioTechniques (1996) 20:102-110; Kobayashi et al., BioTechniques (1996) 21:399-402; and Mosser et al., BioTechniques 1997 22 150-161).

[0358] In some embodiments, the cyclization efficiency of the cyclization methods provided herein is at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or 100%. In some embodiments, the cyclization efficiency of the cyclization methods provided herein is at least about 40%.

[0359] Splicing Elements In some embodiments, the oRNA includes at least one splicing element. The splicing element may be a complete splicing element capable of mediating splicing of the oRNA, or the splicing element may be a residual splicing element from a completed splicing event. For example, in some cases, the splicing element of a linear RNA may mediate a splicing event that results in circularization of the linear RNA, whereby the resulting oRNA includes a residual splicing element from such a splicing-mediated circularization event. In some cases, the residual splicing element is unable to mediate any splicing. In other cases, the residual splicing element may still mediate splicing under certain circumstances. In some embodiments, the splicing element flanks at least one expressed sequence. In some embodiments, the oRNA includes a splicing element flanking each expressed sequence. In some embodiments, the splicing element is on one or both sides of each expressed sequence, resulting in the separation of the expression product, e.g., peptide(s) and / or polypeptide(s).

[0360] In some embodiments, the oRNA includes an internal splicing element where the spliced ​​ends are joined together when duplicated. Some examples may include splice site sequences and short inverted repeats (30-40 nt), such as small introns (<100 nt) with motifs found in cis-sequence elements near backsplice events (suptable4 enriched motifs), such as AluSq2, AluJr, and AluSz, inverted sequences in flanking introns, Alu elements in flanking introns, and sequences 200 bp before (upstream) or after (downstream) the backsplice site with the flanking exon. In some embodiments, the oRNA includes at least one repeated nucleotide sequence described elsewhere herein as an internal splicing element. In such embodiments, the repeated nucleotide sequence may include repeated sequences from the Alu family of introns. See, e.g., U.S. Patent No. 11,058,706.

[0361] In some embodiments, the oRNA may include canonical splice sites adjacent to the head-to-tail junction of the oRNA.

[0362] In some embodiments, the oRNA may contain a bulge-helix-bulge motif that comprises a four base pair stem flanked by two three nucleotide bulges. Cleavage occurs at a site in the bulge region, generating a characteristic fragment with a terminal 5'-hydroxyl group and a 2',3'-cyclic phosphate. Circularization proceeds by nucleophilic attack of the 5'-OH group on the 2',3'-cyclic phosphate of the same molecule forming a 3',5'-phosphodiester bridge.

[0363] In some embodiments, the oRNA may include a sequence that mediates self-ligation. Non-limiting examples of sequences that may mediate self-ligation include a self-circularizing intron, such as a 5' and 3' slice junction, or a self-circularizing catalytic intron, such as a group I, group II, or group III intron. Non-limiting examples of group I intron self-splicing sequences may include the self-splicing permuted intron-exon sequence from the T4 bacteriophage gene td, and the Tetrahymena intervening sequence (IVS) rRNA.

[0364] Other Cyclization Methods In some embodiments, the linear RNA may include complementary sequences that include either repeat or non-repeated nucleic acid sequences within individual introns or across adjacent introns. In some embodiments, the oRNA includes a repeat nucleic acid sequence. In some embodiments, the repeat nucleotide sequence includes a polyCA or polyUG sequence. In some embodiments, the oRNA includes at least one repeat nucleic acid sequence that hybridizes to a complementary repeat nucleic acid sequence in another segment of the oRNA, the hybridized segment forming an internal duplex. In some embodiments, the repeat nucleic acid sequence and the complementary repeat nucleic acid sequence from two separate oRNAs hybridize to generate a single oRNA, the hybridized segment forming an internal duplex. In some embodiments, the complementary sequences are found at the 5' and 3' ends of the linear RNA. In some embodiments, the complementary sequence includes about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more paired nucleotides.

[0365] In some embodiments, chemical methods of circularization may be used to generate oRNA. Such methods may include, but are not limited to, click chemistry (e.g., alkyne and azide-based methods, or clickable bases), olefin metathesis, phosphoramidate ligation, hemiaminal-imine cross-linking, base modification, and any combination thereof.

[0366] In some embodiments, enzymatic methods of circularization may be used to generate the oRNA. In some embodiments, a ligation enzyme, such as a DNA or RNA ligase, may be used to generate the oRNA or complementary template, the complementary strand of the oRNA, or the oRNA.

[0367] Small interfering RNA (siRNA) In some embodiments, the payload region may be or encode an RNA interference (RNAi) sequence that can be used to reduce or inhibit expression of a gene. RNAi (also known as post-transcriptional gene silencing (PTGS), quelling, or co-suppression) is a post-transcriptional gene silencing process in which RNA molecules reduce or inhibit gene expression in a sequence-specific manner, typically by causing the destruction of specific mRNA molecules. The active components of RNAi are short / small double-stranded RNAs (dsRNAs) called small interfering RNAs (siRNAs), typically containing 15-30 nucleotides (e.g., 19-25, 19-24, or 19-21 nucleotides) and 2-nucleotide 3' overhangs, and matching the nucleic acid sequence of the target gene. These short RNA species can be naturally generated in vivo by Dicer-mediated cleavage of larger dsRNAs, and they are functional in mammalian cells.

[0368] Naturally expressed small RNA molecules, termed microRNAs (miRNAs), induce gene silencing by controlling the expression of mRNAs. miRNA-containing RNA-induced silencing complexes (RISCs) target mRNAs that exhibit perfect sequence complementarity with nucleotides 2-7 in the 5' region of the miRNA, called the seed region, and other base pairs with its 3' region. miRNA-mediated downregulation of gene expression can be caused by cleavage of the target mRNA, translational inhibition of the target mRNA, or mRNA decay inhibition. miRNA targeting sequences are usually located in the 3'-UTR of the target mRNA. A single miRNA can target more than 100 transcripts from various genes, and one mRNA can be targeted by different miRNAs.

[0369] siRNA duplexes or dsRNAs targeting specific mRNAs can be designed, synthesized in vitro, and introduced into cells to activate the RNAi process. It has been previously shown that 21-nucleotide siRNA duplexes (called small interfering RNA) were able to achieve strong and specific gene knockdown in mammalian cells without inducing immune responses. Now, post-transcriptional gene silencing by siRNA has quickly emerged as a powerful tool for genetic analysis in mammalian cells, and has the potential to generate novel therapeutic agents.

[0370] The siRNA sequence synthesized in vitro can be introduced into cells to activate RNAi. When exogenous siRNA duplexes are introduced into cells, they can be constructed to form a multi-unit complex, the RNA-induced silencing complex (RISC), which interacts with the RNA sequence complementary to one of the two strands of the siRNA duplex (i.e., the antisense strand), similar to endogenous dsRNA. During the process, the sense strand (or passenger strand) of the siRNA is lost from the complex, whereas the antisense strand (or guide strand) of the siRNA is matched with its complementary RNA. In particular, the target of the RISC complex containing the siRNA is the mRNA that presents perfect sequence complementarity. Then, siRNA-mediated gene silencing occurs by cleaving, releasing and degrading the target.

[0371] The siRNA duplex, which comprises a sense strand that is homologous to the target mRNA and an antisense strand that is complementary to the target mRNA, offers many advantages in terms of efficiency for target RNA destruction, compared with the use of single-stranded (ss)-siRNA (e.g., antisense strand RNA or antisense oligonucleotide).In many cases, it requires a higher concentration of ss-siRNA to achieve the effective gene silencing efficacy of the corresponding duplex.

[0372] Design and sequence of siRNA duplexes Several guidelines for designing siRNA have been proposed in the art. These guidelines usually recommend generating a 19-nucleotide double-stranded region targeting the region in the gene to be silenced, a symmetrical 2-3 nucleotide 3' overhang, a 5'-phosphate and a 3'-hydroxyl group. Other rules that may govern the preference of siRNA sequences include, but are not limited to, (i) A / U at the 5' end of the antisense strand; (ii) G / C at the 5' end of the sense strand; (iii) at least five A / U residues in the 1 / 3 of the 5' end of the antisense strand; and (iv) the absence of any GC stretches that are more than 9 nucleotides in length. Following such considerations, together with the specific sequence of the target gene, highly effective siRNA constructs essential for suppressing mammalian target gene expression can be easily designed.

[0373] In some embodiments, siRNA constructs (e.g., siRNA duplexes or coded dsRNAs) are designed to target specific genes. Such siRNA constructs can specifically suppress gene expression and protein production. In some aspects, siRNA constructs are designed and used to selectively "knock out" gene variants in cells, i.e., mutated transcripts that have been identified in patients or are responsible for various diseases and / or disorders. In some aspects, siRNA constructs are designed and used to selectively "knock down" gene variants in cells. In other aspects, siRNA constructs can inhibit or suppress both wild-type and mutated versions of genes.

[0374] In some embodiments, the siRNA sequence comprises a sense strand and a complementary antisense strand, both strands hybridize with each other to form a double-stranded structure. The antisense strand has sufficient complementarity to the mRNA sequence to induce target-specific RNAi, i.e., the siRNA sequence has sufficient sequence to induce destruction of the target mRNA by the RNAi mechanism or process.

[0375] In some embodiments, the siRNA sequence comprises a sense strand and a complementary antisense strand, both strands hybridize to each other to form a double-stranded structure, and the start site of hybridization to the mRNA is between nucleotides 100 and 10,000 on the mRNA sequence. As non-limiting examples, the initiation site may be located between nucleotides 100-150, 150-200, 200-250, 250-300, 300-350, 350-400, 400-450, 450-500, 500-550, 550-600, 600-650, 650-700, 700-70, 750-800, 800-850, 850-900, 900-950, 950-1000, 1000-1050, 1050-1100, 1100-1150, 1150-1200, 12 00~1250, 1250~1300, 1300~1350, 1350~1400, 1400~1450, 1450~1500, 1500~1550, 1550~1600, 1600~1650, 1650~1700, 1700~1750, 1750~1800, 1800~1850, 1850~1900, 1900~1950, 1950~2000, 2000~2050, 2050~2100, 2100~2150, 2150~2200, 2200~2250, 2250~230 0, 2300~2350, 2350~2400, 2400~2450, 2450~2500, 2500~2550, 2550~2600, 2600~2650, 2650~2700, 2700~2750, 2750~2800, 2800~2850, 2850~2900, 2900~2950, ​​2950~3000, 3000~3050, 3050~3100, 3100~3150, 3150~3200, 3200~3250, 3250~3300, 3300~3350, 3350 ~3400, 3400~3450, 3450~3500, 3500~3550, 3550~3600, 3600~3650, 3650~3700, 3700~3750, 3750~3800, 3800~3850, 3850~3900, 3900~3950, 3950~4000, 4000~4050, 4050~4100, 4100~4150, 4150~4200, 4200~4250, 4250~4300, 4300~4350, 4350~4400, 4400~4450,4450~4500、4500~4550、4550~4600、4600~4650、4650~4700、4700~4750、4750~4800、4800~4850、4850~4900、4900~4950、4950~5000、5000~5050、5050~5100、5100~5150、5150~5200、5200~5250、5250~5300、5300~5350、5350~5400、5400~5450、5450~5500、5500~5550、5550~5600、5600~5650、5650~5700、5700~5750、5750~5800、5800~5850、5850~5900、5900~5950、5950~6000、6000~6050、6050~6100、6100~6150、6150~6200、6200~6250、6250~6300、6300~6350、6350~6400、6400~6450、6450~6500、6500~6550、6550~6600、6600~6650、6650~6700、6700~6750、6750~6800、6800~6850、6850~6900、6900~6950、6950~7000、7000~7050、7050~7100、7100~7150、7150~7200、7200~7250、7250~7300、7300~7350、7350~7400、7400~7450、7450~7500、7500~7550、7550~7600、7600~7650、7650~7700、7700~7750、7750~7800、7800~7850、7850~7900、7900~7950、7950~8000、8000~8050、8050~8100、8100~8150、8150~8200、8200~8250、8250~8300、8300~8350、8350~8400、8400~8450、8450~8500、8500~8550、8550~8600、8600~8650、8650~8700、8700~8750、8750~8800、8800~8850、8850~8900、8900~8950、8950~9000、9000~9050、9050~9100、9100~9150、9150~9200、9200~9250、9250~9300、9300~9350、9350~9400、9400~9450、It can be 9450 to 9500, 9500 to 9550, 9550 to 9600, 9600 to 9650, 9650 to 9700, 9700 to 9750, 9750 to 9800, 9800 to 9850, 9850 to 9900, 9900 to 9950, or 9950 to 10000.

[0376] In some embodiments, the antisense strand and the target mRNA sequence are 100% complementary. The antisense strand can be complementary to any portion of the target mRNA sequence.

[0377] In other embodiments, the antisense strand and the target mRNA sequence contain at least one mismatch. As non-limiting examples, the antisense strand and the target mRNA sequence contain at least 30%, 40%, 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% or at least 20-30%, 20-40%, 20-50%, 20-60%, 20-70%, 20-80%, 20-90%, 20-95%, 20-99%, 30-40%, 30-50%, 30-60%, 30-70 ...30-60%, 30-70% or at least 20-30%, 20-40%, 20-50%, 30-60%, 30-70% or at least 20-30%, 20-40%, 20-50%, 30-60%, 30-70% or at least 20-30%, 20-40%, 20-50%, 30-60%, 30-70% or at least 20-30%, 20-40%, 20-50%, 30-60%, %, 30-80%, 30-90%, 30-95%, 30-99%, 40-50%, 40-60%, 40-70%, 40-80%, 40-90%, 40-95%, 40-99%, 50-60%, 50-70%, 50-80%, 50-90%, 50-95%, 50-99%, 60-70%, 60-80%, 60-90%, 60-95%, 60-99%, 70-80%, 70-90%, 70-95%, 70-99%, 80-90%, 80-95%, 80-99%, 90-95%, 90-99% or 95-99%.

[0378] In some embodiments, the siRNA sequence has a length of about 10-50 or more nucleotides, i.e., each strand contains 10-50 nucleotides (or nucleotide analogs). Preferably, the siRNA sequence has a length of about 15-30, e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in each strand, one of the strands being fully complementary to the target region. In some embodiments, the siRNA sequence has a length of about 19-25, 19-24, or 19-21 nucleotides.

[0379] In some embodiments, the siRNA sequence may be a synthetic RNA duplex comprising about 19 nucleotides to about 25 nucleotides and two overhanging nucleotides at the 3' end. In some aspects, the siRNA construct may be an unmodified RNA molecule. In other aspects, the siRNA construct may contain at least one modified nucleotide, such as a base, sugar or backbone modification.

[0380] In some embodiments, siRNA sequences can be encoded in plasmid vectors, viral vectors or other nucleic acid expression vectors for delivery to cells.DNA expression plasmids can be used to stably express siRNA duplexes or dsRNA in cells to achieve long-term inhibition of target gene expression.In one aspect, the sense and antisense strands of siRNA duplexes are typically linked by a short spacer sequence that results in the expression of a stem-loop structure called short hairpin RNA (shRNA).The hairpin is recognized and cleaved by Dicer, thus generating a mature siRNA construct.

[0381] In some embodiments, the sense and antisense strands of the siRNA duplex may be linked by a short spacer sequence, which may optionally be linked to additional flanking sequences, resulting in the expression of a flanking arm-stem-loop structure called a primary microRNA (pri-miRNA), which may be recognized and cleaved by Drosha and Dicer, thus generating a mature siRNA construct.

[0382] In some embodiments, the siRNA duplex or the encoded dsRNA suppresses (or degrades) the target mRNA. Thus, the siRNA duplex or the encoded dsRNA can be used to substantially inhibit gene expression in cells. In some aspects, the inhibition of gene expression is at least about 20%, preferably at least about 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95% and 100%, or at least 20-30%, 20-40%, 20-50%, 20-60%, 20-70%, 20-80%, 20-90%, 20-95%, 20-100%, 30-40%, 30-50%, 30-60%, 30-70%, 30-80%, 30-90%, 30-95%, 30-100%. , 40-50%, 40-60%, 40-70%, 40-80%, 40-90%, 40-95%, 40-100%, 50-60%, 50-70%, 50-80%, 50-90%, 50-95%, 50-100%, 60-70%, 60-80%, 60-90%, 60-95%, 60-100%, 70-80%, 70-90%, 70-95%, 70-100%, 80-90%, 80-95%, 80-100%, 90-95%, 90-100% or 95-100% inhibition. Thus, the protein product of the targeted gene is at least about 20%, preferably at least about 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95% and 100%, or at least 20-30%, 20-40%, 20-50%, 20-60%, 20-70%, 20-80%, 20-90%, 20-95%, 20-100%, 30-40%, 30-50%, 30-60%, 30-70%, 30-80%, 30-90%, 30-95%, 30-100%, 30-40%, 30-50%, 30-60%, 30-70%, 30-80%, 30-90%, 30-95%, 30-100%, 30-20%, 30-3 ... In some embodiments, the therapeutic effect may be inhibited by 00%, 40-50%, 40-60%, 40-70%, 40-80%, 40-90%, 40-95%, 40-100%, 50-60%, 50-70%, 50-80%, 50-90%, 50-95%, 50-100%, 60-70%, 60-80%, 60-90%, 60-95%, 60-100%, 70-80%, 70-90%, 70-95%, 70-100%, 80-90%, 80-95%, 80-100%, 90-95%, 90-100% or 95-100%.

[0383] In some embodiments, the siRNA construct comprises a miRNA seed match for the target located in the guide strand.In another embodiment, the siRNA construct comprises a miRNA seed match for the target located in the passenger strand.In yet another embodiment, the siRNA duplex or encoded dsRNA that targets a gene does not comprise a seed match for the target located in the guide or passenger strand.

[0384] In some embodiments, the siRNA duplex or encoded dsRNA targeting the gene may have little significant full-length off-target for the guide strand. In another embodiment, the siRNA duplex or encoded dsRNA targeting the gene may have little significant full-length off-target for the passenger strand. The siRNA duplex or encoded dsRNA targeting the gene may have less than 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 1-5%, 2-6%, 3-7%, 4-8%, 5-9%, 5-10%, 6-10%, 5-10%, 6-10%, 7-10%, 8-10%, 9-10%, 11-10%, 12-10%, 13-10%, 14-10%, 15-10%, 16-10%, 17-10%, 17-10%, 18-10%, 19-20%, 20-25%, 21-25%, 22-25%, 23-25%, 24-25%, 25-25%, 26-27%, 27-28%, 28-29%, 29-30%, 30-35%, 31-32%, 32-33%, 34-35%, 35-36%, 37-38%, 38-39%, 39-40%, 39-41%, 39-42%, 39-43%, 39-44%, 39-45%, 39-46%, 39-47%, 39-48%, 39-49%, 39-41%, 39-42%, 39-4 5%, 5-20%, 5-25%, 5-30%, 10-20%, 10-30%, 10-40%, 10-50%, 15-30%, 15-40%, 15-45%, 20-40%, 20-50%, 25-50%, 30-40%, 30-50%, 35-50%, 40-50%, 45-50% full length off-target effect for the passenger strand. In yet another embodiment, the siRNA duplex or encoded dsRNA targeting the gene may have little to no significant full length off-target for the guide strand or passenger strand. siRNA duplexes or encoded dsRNAs targeting genes were 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, less than 50%, 1-5%, 2-6%, 3-7%, 4-8%, 5-9%, 5-10%, 6-10%, 5-15%, 5-20%, 5-25% The full length off-target effects for the guide or passenger strand may be 5-30%, 10-20%, 10-30%, 10-40%, 10-50%, 15-30%, 15-40%, 15-45%, 20-40%, 20-50%, 25-50%, 30-40%, 30-50%, 35-50%, 40-50%, 45-50%.

[0385] In some embodiments, the siRNA duplex or encoded dsRNA targeting the gene may have high activity in vitro. In another embodiment, the siRNA construct may have low activity in vitro. In yet another embodiment, the siRNA duplex or dsRNA targeting the gene may have high guide strand activity and low passenger strand activity in vitro.

[0386] In some embodiments, the siRNA construct has high guide strand activity and low passenger strand activity in vitro. Target knockdown (KD) by the guide strand can be at least 40%, 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 99.5% or 100%. Target knockdown by guide strand is 40-50%, 45-50%, 50-55%, 50-60%, 60-65%, 60-70%, 60-75%, 60-80%, 60-85%, 60-90%, 60-95%, 60-99%, 60-99.5%, 60-100%, 65-70%, 65-75%, 65-80%, 65-85%, 65-90%, 65-95%, 65-99%, 65-99.5%, 65-100%, 70-75%, 70-80%, 70-85%, 70-90%, 70-95%, 70-99%, 70-99.5% , 70-100%, 75-80%, 75-85%, 75-90%, 75-95%, 75-99%, 75-99.5%, 75-100%, 80-85%, 80-90%, 80-95%, 80-99%, 80-99.5%, 80-100%, 85-90%, 85-95%, 85-99%, 85-99.5%, 85-100%, 90-95%, 90-99%, 90-99.5%, 90-100%, 95-99%, 95-99.5%, 95-100%, 99-99.5%, 99-100% or 99.5-100%. As a non-limiting example, the target knockdown (KD) by the guide strand is greater than 70%. As a non-limiting example, the target knockdown (KD) by the guide strand is greater than 60%.

[0387] In some embodiments, the expressed guide to passenger (G:P) (also referred to as antisense to sense) strand ratio is 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, 1;1, 2:10, 2:9, 2:8, 2:7, 2:6, 2:5, 2:4, 2:3, 2:2, 2:1, 3:10, 3:9, 3:8, 3:7, 3:6, 3:5, 3:4, 3:3, 3:2, 3:1, 4:10, 4:9, 4:10 ... :8, 4:7, 4:6, 4:5, 4:4, 4:3, 4:2, 4:1, 5:10, 5:9, 5:8, 5:7, 5:6, 5:5, 5:4, 5:3, 5:2, 5:1, 6:10, 6:9, 6:8, 6:7, 6:6, 6:5, 6:4, 6:3, 6:2, 6:1, 7:10, 7:9, 7:8, 7:7, 7:6, 7:5, 7 :4, 7:3, 7:2, 7:1, 8:10, 8:9, 8:8, 8:7, 8:6, 8:5, 8:4, 8:3, 8:2, 8:1, 9:10, 9:9, 9:8, 9:7, 9:6, 9:5, 9:4, 9:3, 9:2, 9:1, 10:10, 10:9, 10:8, 10:7, 10:6, 10:5, 10:4, 10:3, 1 It can be 0:2, 10:1, 1:99, 5:95, 10:90, 15:85, 20:80, 25:75, 30:70, 35:65, 40:60, 45:55, 50:50, 55:45, 60:40, 65:35, 70:30, 75:25, 80:20, 85:15, 90:10, 95:5, or 99:1. Guide to passenger ratio refers to the ratio of guide strand to passenger strand after intracellular processing of pri-microRNA. For example, a guide to passenger ratio of 80:20 would have 8 guide strands for every 2 passenger strands processed from the precursor. As a non-limiting example, the guide to passenger strand ratio is 8:2 in vitro. As a non-limiting example, the guide to passenger strand ratio is 8:2 in vivo. As a non-limiting example, the guide to passenger strand ratio is 9:1 in vitro. As a non-limiting example, the guide to passenger strand ratio is 9:1 in vivo.

[0388] In some embodiments, the expressed guide to passenger (G:P) (also referred to as antisense to sense) strand ratio is greater than 1. In some embodiments, the expressed guide to passenger (G:P) (also referred to as antisense to sense) strand ratio is greater than 2. In some embodiments, the expressed guide to passenger (G:P) (also referred to as antisense to sense) strand ratio is greater than 5. In some embodiments, the expressed guide to passenger (G:P) (also referred to as antisense to sense) strand ratio is greater than 10. In some embodiments, the expressed guide to passenger (G:P) (also referred to as antisense to sense) strand ratio is greater than 20. In some embodiments, the expressed guide to passenger (G:P) (also referred to as antisense to sense) strand ratio is greater than 50. In some embodiments, the expressed guide to passenger (G:P) (also referred to as antisense to sense) strand ratio is at least 3:1. In some embodiments, the ratio of expressed guide to passenger (G:P) (also referred to as antisense to sense) strands is at least 5:1. In some embodiments, the ratio of expressed guide to passenger (G:P) (also referred to as antisense to sense) strands is at least 10:1. In some embodiments, the ratio of expressed guide to passenger (G:P) (also referred to as antisense to sense) strands is at least 20:1. In some embodiments, the ratio of expressed guide to passenger (G:P) (also referred to as antisense to sense) strands is at least 50:1.

[0389] In some embodiments, the expressed passenger to guide (P:G) (also referred to as sense to antisense) strand ratio is 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, 1;1, 2:10, 2:9, 2:8, 2:7, 2:6, 2:5, 2:4, 2:3, 2:2, 2:1, 3:10, 3:9, 3:8, 3:7, 3:6, 3:5, 3:4, 3:3, 3:2, 3:1, 4:10, 4:9, 4:10 ... :8, 4:7, 4:6, 4:5, 4:4, 4:3, 4:2, 4:1, 5:10, 5:9, 5:8, 5:7, 5:6, 5:5, 5:4, 5:3, 5:2, 5:1, 6:10, 6:9, 6:8, 6:7, 6:6, 6:5, 6:4, 6:3, 6:2, 6:1, 7:10, 7:9, 7:8, 7:7, 7:6, 7:5, 7:4, 7:3, 7:2, 7:1, 8:10, 8:9, 8:8, 8:7, 8:6, 8:5, 8:4, 8:3, 8:2, 8:1, 9:10, 9:9, 9:8, 9:7, 9:6, 9:5, 9:4, 9:3, 9:2, 9:1, 10:10, 10:9, 10:8, 10:7, 10:6, 10:5, 10:4, 10:3 , 10:2, 10:1, 1:99, 5:95, 10:90, 15:85, 20:80, 25:75, 30:70, 35:65, 40:60, 45:55, 50:50, 55:45, 60:40, 65:35, 70:30, 75:25, 80:20, 85:15, 90:10, 95:5, or 99:1. Passenger to guide ratio refers to the ratio of passenger strand to guide strand after guide strand excision. For example, a passenger to guide ratio of 80:20 would have 8 passenger strands for every 2 guide strands processed from the precursor. As a non-limiting example, the passenger to guide strand ratio is 80:20 in vitro. As a non-limiting example, the passenger to guide strand ratio is 80:20 in vivo. As a non-limiting example, the passenger to guide strand ratio is 8:2 in vitro. As a non-limiting example, the passenger to guide strand ratio is 8:2 in vivo. As a non-limiting example, the passenger to guide strand ratio is 9:1 in vitro. As a non-limiting example, the passenger to guide strand ratio is 9:1 in vivo.

[0390] In some embodiments, the expressed passenger to guide (P:G) (also referred to as sense to antisense) strand ratio is greater than 1. In some embodiments, the expressed passenger to guide (P:G) (also referred to as sense to antisense) strand ratio is greater than 2. In some embodiments, the expressed passenger to guide (P:G) (also referred to as sense to antisense) strand ratio is greater than 5. In some embodiments, the expressed passenger to guide (P:G) (also referred to as sense to antisense) strand ratio is greater than 10. In some embodiments, the expressed passenger to guide (P:G) (also referred to as sense to antisense) strand ratio is greater than 20. In some embodiments, the expressed passenger to guide (P:G) (also referred to as sense to antisense) strand ratio is greater than 50. In some embodiments, the expressed passenger to guide (P:G) (also referred to as sense to antisense) strand ratio is at least 3:1. In some embodiments, the passenger to guide (P:G) (also referred to as sense to antisense) strand ratio expressed is at least 5:1. In some embodiments, the passenger to guide (P:G) (also referred to as sense to antisense) strand ratio expressed is at least 10:1. In some embodiments, the passenger to guide (P:G) (also referred to as sense to antisense) strand ratio expressed is at least 20:1. In some embodiments, the passenger to guide (P:G) (also referred to as sense to antisense) strand ratio expressed is at least 50:1.

[0391] In some embodiments, a passenger-guide strand duplex is considered effective if the pri- or pre-microRNA demonstrates a guide to passenger strand ratio of greater than 2-fold as measured by processing by methods known in the art and described herein. As non-limiting examples, the pri- or pre-microRNA demonstrates a guide to passenger strand ratio of 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, greater than 15x, or 2-5x, 2-10x, 2-15x, 3-5x, 3-10x, 3-15x, 4-5x, 4-10x, 4-15x, 5-10x, 5-15x, 6-10x, 6-15x, 7-10x, 7-15x, 8-10x, 8-15x, 9-10x, 9-15x, 10-15x, 11-15x, 12-15x, 13-15x, or 14-15x.

[0392] In some embodiments, the vector genome encoding the dsRNA comprises at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or more of the sequence of the entire length of the construct. As a non-limiting example, the vector genome comprises a sequence that is at least 80% of the entire length sequence of the construct.

[0393] In some embodiments, siRNA constructs may be used to silence a wild-type or mutant gene by targeting at least one exon on the sequence.

[0394] siRNA modification In some embodiments, when the siRNA construct is not delivered as a precursor or DNA, it can be chemically modified to adjust some characteristics of the RNA molecule, including but not limited to increasing the stability of the siRNA in vivo. The chemically modified siRNA construct can be used in human therapeutic applications, and improves the RNAi activity of the siRNA construct without compromising it. Non-limiting examples include siRNA constructs modified at both the 3' and 5' ends of both the sense and antisense strands.

[0395] In some embodiments, the modified nucleotides may be in the sense strand only.

[0396] In some embodiments, the modified nucleotides may be in the antisense strand only.

[0397] In some embodiments, modified nucleotides can be in both the sense and antisense strands.

[0398] In some embodiments, chemically modified nucleotides do not affect the ability of the antisense strand to pair with a target mRNA sequence.

[0399] MicroRNA (miR) backbone In some embodiments, the siRNA construct may be encoded in a polynucleotide sequence that also includes a microRNA (miR) backbone construct. As used herein, a "microRNA (miR) backbone construct" is a framework or starting molecule that forms the sequence or structural basis for designing or creating subsequent molecules.

[0400] In some embodiments, the miR backbone construct comprises at least one 5' flanking region. By way of non-limiting example, the 5' flanking region may be of any length and may comprise a 5' flanking sequence that may be derived in whole or in part from a wild-type microRNA sequence or may be a completely artificial sequence.

[0401] In some embodiments, the miR scaffold construct comprises at least one 3' flanking region. By way of non-limiting example, the 3' flanking region may be of any length and may comprise a 3' flanking sequence that may be derived in whole or in part from a wild-type microRNA sequence, or may be a completely artificial sequence.

[0402] In some embodiments, the miR scaffold construct comprises at least one loop motif region. As a non-limiting example, the loop motif region can comprise a sequence that can be of any length.

[0403] In some embodiments, the miR scaffold construct comprises a 5' flanking region, a loop motif region, and / or a 3' flanking region.

[0404] In some embodiments, at least one payload (e.g., siRNA, miRNA, or other RNAi described herein) may be encoded by a polynucleotide that may also include at least one miR backbone construct. The miR backbone construct may be of any length and may include a 5' flanking sequence that may be derived in whole or in part from a wild-type microRNA sequence, or may be completely artificial. The 3' flanking sequence may reflect the 5' flanking sequence and / or the 3' flanking sequence in size and origin. Either of the flanking sequences may be absent. The 3' flanking sequence may optionally contain one or more CNNC motifs ("N" represents any nucleotide).

[0405] In some embodiments, the 5' arm of the stem-loop structure of a polynucleotide comprising or encoding a miR scaffold construct comprises a sequence that encodes a sense sequence.

[0406] In some embodiments, the 3' arm of the stem-loop of a polynucleotide comprising or encoding a miR scaffold construct comprises a sequence encoding an antisense sequence, which in some instances comprises a "G" nucleotide at its 5'-most end.

[0407] In some embodiments, the sense sequence can be on the 3' arm while the antisense sequence is on the 5' arm of the stem of a stem-loop structure of a polynucleotide comprising or encoding a miR scaffold construct.

[0408] In some embodiments, the sense and antisense sequences can be completely complementary over a substantial portion of their length, while in other embodiments, the sense and antisense sequences can be, independently, at least 70, 80, 90, 95, or 99% complementary over at least 50, 60, 70, 80, 85, 90, 95, or 99% of the length of the strand.

[0409] Neither the identity of the sense sequence nor the homology of the antisense sequence need be 100% complementary to the target sequence.

[0410] In some embodiments, separating the sense and antisense sequences of the stem-loop structure of the polynucleotide is a loop sequence (also known as a loop motif, linker or linker motif). The loop sequence can be of any length from 4-30 nucleotides, 4-20 nucleotides, 4-15 nucleotides, 5-15 nucleotides, 6-12 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, and / or 15 nucleotides.

[0411] In some embodiments, the loop sequence comprises a nucleic acid sequence encoding at least one UGUG motif. In some embodiments, the nucleic acid sequence encoding a UGUG motif is located at the 5' end of the loop sequence.

[0412] In some embodiments, a spacer region may be present in the polynucleotide to separate one or more modules from each other (e.g., 5' flanking region, loop motif region, 3' flanking region, sense sequence, antisense sequence). One or more such spacer regions may be present.

[0413] In some embodiments, a spacer region of 8 to 20, i.e., 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides, may be present between the sense sequence and the flanking region sequence.

[0414] In some embodiments, the spacer region is 13 nucleotides in length and is located between the 5' end of the sense sequence and the 3' end of the flanking sequence. In some embodiments, the spacer is long enough to form approximately one helix of the sequence.

[0415] In some embodiments, a spacer region of 8 to 20, i.e., 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides, may be present between the antisense sequence and the flanking sequence.

[0416] In some embodiments, the spacer sequence is 10-13, i.e., 10, 11, 12 or 13 nucleotides, and is located between the 3' end of the antisense sequence and the 5' end of the flanking sequence. In some embodiments, the spacer is long enough to form approximately one helix of the sequence.

[0417] In some embodiments, the polynucleotide comprises, in the 5' to 3' direction, a 5' flanking sequence, a 5' arm, a loop motif, a 3' arm, and a 3' flanking sequence. As a non-limiting example, the 5' arm can comprise a sense sequence and the 3' arm comprises an antisense sequence. In another non-limiting example, the 5' arm comprises an antisense sequence and the 3' arm comprises a sense sequence.

[0418] In some embodiments, the 5' arm, the payload (e.g., sense and / or antisense sequences), the loop motif and / or the 3' arm sequence may be modified (e.g., by substituting one or more nucleotides, adding nucleotides and / or deleting nucleotides). Modifications may result in beneficial changes in the function of the construct (e.g., increasing knockdown of the target sequence, reducing degradation of the construct, reducing off-target effects, increasing payload efficiency, and reducing payload degradation).

[0419] In some embodiments, the miR backbone construct of the polynucleotide is aligned so that the excision rate of the guide strand is higher than the excision rate of the passenger strand. The excision rate of the guide or passenger strand can be independently 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or more than 99%. As a non-limiting example, the excision rate of the guide strand is at least 80%. As another non-limiting example, the excision rate of the guide strand is at least 90%.

[0420] In some embodiments, the excision rate of the guide strand is higher than the excision rate of the passenger strand, hi one aspect, the excision rate of the guide strand may be at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or greater than 99% higher than the passenger strand.

[0421] In some embodiments, the efficiency of excision of the guide strand is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or greater. As a non-limiting example, the efficiency of excision of the guide strand is greater than 80%.

[0422] In some embodiments, the efficiency of excision of the guide strand is higher than the excision of the passenger strand from the miR backbone construct. Excision of the guide strand can be 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 times more efficient than excision of the passenger strand from the miR backbone construct.

[0423] In some embodiments, the miR backbone construct comprises a dual-function targeting polynucleotide. As used herein, a "dual-function targeting" polynucleotide is one in which both the guide and passenger strands knock down the same target, or in which the guide and passenger strands knock down different targets.

[0424] In some embodiments, the miR backbone constructs of the polynucleotides described herein can comprise a 5' flanking region, a loop motif region, and a 3' flanking region.

[0425] In some embodiments, polynucleotides are designed with at least one of the following features: loop variants, seed mismatch / bulge / wobble variants, stem mismatches, loop variants and basal stem mismatch variants, seed mismatches and basal stem mismatch variants, stem mismatches and basal stem mismatch variants, seed wobble and basal stem wobble variants, or stem sequence variants.

[0426] In some embodiments, the miR backbone construct can be a natural pri-miRNA backbone.

[0427] In some embodiments, the selection of the miR backbone construct is determined by a method that compares the polynucleotides in the pri-miRNA.

[0428] In some embodiments, the selection of the miR backbone construct is determined by methods comparing the polynucleotides in natural and synthetic pri-miRNAs.

[0429] Transfer RNA (tRNA) Transfer RNA (tRNA) is an RNA molecule that translates mRNA into protein. tRNA contains a cloverleaf structure that includes a 3' acceptor site, a 5' terminal phosphate, a D arm, a T arm, and an anticodon arm. The primary purpose of tRNA is to carry an amino acid on its 3' acceptor site to the ribosomal complex with the help of aminoacyl-tRNA synthetase, an enzyme that loads the appropriate amino acid onto free tRNA to synthesize a protein. Once an amino acid is bound to tRNA, the tRNA is considered an aminoacyl-tRNA. The type of amino acid on tRNA is dependent on the mRNA codon. The anticodon arm of tRNA is complementary to the mRNA codon and is the site of the anticodon that directs which amino acid to carry. tRNA is also known to have a role in controlling apoptosis by acting as a cytochrome c scavenger.

[0430] In some embodiments, the originator construct and / or the benchmark construct comprises or encodes a tRNA.

[0431] Ribosomal RNA (rRNA) Ribosomal RNA (rRNA) is the RNA that forms the ribosome. Ribosomes are essential for protein synthesis and contain large and small ribosomal subunits. In prokaryotes, the small 30S and large 50S ribosomal subunits make up the 70S ribosome. In eukaryotes, the 40S and 60S subunits form the 80S ribosome. To bind aminoacyl-tRNA and link amino acids together to produce polypeptides, the ribosome contains three sites: the exit site (E), the peptidyl site (P), and the acceptor site (A).

[0432] In some embodiments, the originator construct and / or the benchmark construct comprises or encodes an rRNA.

[0433] MicroRNA (miRNA) MicroRNAs (or miRNAs) are 19-25 nucleotide long non-coding RNAs that bind to the 3'UTR of nucleic acid molecules and downregulate gene expression either by reducing nucleic acid molecule stability or by inhibiting translation. The originator construct and / or benchmark construct may include one or more microRNA target sequences, microRNA sequences, or microRNA seeds.

[0434] The microRNA sequence comprises a "seed" region, i.e., a sequence in the region of positions 2-8 of the mature microRNA, which has perfect Watson-Crick complementarity with the miRNA target sequence. The microRNA seed may comprise positions 2-8 or 2-7 of the mature microRNA. In some embodiments, the microRNA seed may comprise 7 nucleotides (e.g., nucleotides 2-8 of the mature microRNA) and the seed complementary site in the corresponding miRNA target is adjacent to an adenine (A) opposite microRNA position 1. In some embodiments, the microRNA seed may comprise 6 nucleotides (e.g., nucleotides 2-7 of the mature microRNA) and the seed complementary site in the corresponding miRNA target is adjacent to an adenine (A) opposite microRNA position 1. The bases of the microRNA seed have perfect complementarity with the target sequence. By engineering the microRNA target sequence into the 3'UTR of an mRNA, one may target the molecule for degradation or reduced translation, provided the microRNA of interest is available. This process would reduce the risk of off-target effects with nucleic acid molecule delivery.

[0435] As used herein, the term "microRNA site" refers to a microRNA target site or microRNA recognition site, or any nucleotide sequence to which a microRNA binds or associates. It should be understood that "binding" may follow conventional Watson-Crick hybridization rules or may reflect any stable association of the microRNA with a target sequence at or adjacent to the microRNA site.

[0436] Non-limiting examples of tissues in which microRNAs are known to regulate mRNA and thereby protein expression include, but are not limited to, liver (miR-122), muscle (miR-133, miR-206, miR-208), endothelial cells (miR-17-92, miR-126), bone marrow cells (miR-142-3p, miR-142-5p, miR-16, miR-21, miR-223, miR-24, miR-27), adipose tissue (let-7, miR-30c), heart (miR-ld, miR-149), kidney (miR-192, miR-194, miR-204), and lung epithelial cells (let-7, miR-133, miR-126). MicroRNAs can also regulate complex biological processes, such as angiogenesis (miR-132).

[0437] For example, if the nucleic acid molecule is an mRNA and is not intended to be delivered to the liver but ends up there, then miR-122, a microRNA abundant in the liver, can inhibit expression of a gene of interest if one or more target sites for miR-122 are engineered into the 3'UTR of the mRNA. Introduction of one or more binding sites for different microRNAs can be engineered to further reduce mRNA lifetime, stability, and protein translation.

[0438] Conversely, microRNA binding sites can be engineered out (i.e., removed) from the sequences in which they naturally occur to increase protein expression in specific tissues. For example, miR-122 binding sites can be removed to improve protein expression in the liver. Control of expression in multiple tissues can be achieved through the introduction or removal of one or several microRNA binding sites.

[0439] Long non-coding RNA (lncRNA) Long non-coding RNAs (lncRNAs) are regulatory RNA molecules that do not code for proteins but affect a wide range of biological processes. lncRNA designation can usually be restricted to non-coding transcripts longer than about 200 nucleotides. Length designation distinguishes lncRNAs from small regulatory RNAs such as small interfering RNAs (siRNAs) and microRNAs (miRNAs). In vertebrates, the number of lncRNA species is thought to greatly exceed the number of protein-coding species. lncRNAs are also thought to drive the biological complexity observed in vertebrates compared to invertebrates. Evidence of this complexity is seen in many cellular compartments of vertebrate organisms, for example, the T lymphocyte compartment of the adaptive immune system. Differences in lncRNA expression and function may be major contributing factors to human disease.

[0440] In some embodiments, the originator construct and / or the benchmark construct comprises a lncRNA.

[0441] RNA modification In some embodiments, an originator construct or a benchmark construct may contain one or more modified nucleotides, such as, but not limited to, sugar-modified nucleotides, nucleobase modifications, and / or backbone modifications. In some embodiments, an originator construct or a benchmark construct may contain combined modifications, such as combined nucleobase and backbone modifications.

[0442] In some embodiments, the modified nucleotide may be a sugar-modified nucleotide. Sugar-modified nucleotides include, but are not limited to, 2'-fluoro, 2'-amino and 2'-thio modified ribonucleotides, such as 2'-fluoro modified ribonucleotides. Modified nucleotides may be modified at the sugar moiety and nucleotides having sugars that are not ribosyl or analogs thereof. For example, the sugar moiety may be or be based on mannose, arabinose, glucopyranose, galactopyranose, 4'-thioribose, and other sugars, heterocycles, or carbocycles.

[0443] In some embodiments, the modified nucleotide may be a nucleobase-modified nucleotide.

[0444] In some embodiments, the modified nucleotide may be a backbone-modified nucleotide. In some embodiments, the originator construct or benchmark construct may further include other modifications on the backbone. Normal "backbone" as used herein refers to the repeating alternating sugar-phosphate sequence in a DNA or RNA molecule. The deoxyribose / ribose sugar is connected to phosphate groups at both the 3'-hydroxyl and 5'-hydroxyl groups in an ester bond, also known as a "phosphodiester" bond / linker (PO bond). The PO backbone may be modified as a "phosphorothioate backbone (PS bond)". In some cases, the natural phosphodiester bond may be replaced by an amide bond, but the four atoms between the two sugar units are maintained. Such amide modifications may facilitate solid-phase synthesis of oligonucleotides and increase the thermodynamic stability of duplexes formed with siRNA complements.

[0445] Modified bases refer to nucleotide bases modified by the replacement or addition of one or more atoms or groups, such as, but not limited to, adenine, guanine, cytosine, thymine, uracil, xanthine, inosine, and queosine. Some examples of modifications on the nucleobase moiety include, but are not limited to, alkylated, halogenated, thiolated, aminated, amidated, or acetylated bases, either individually or in combination. More specific examples include, for example, 5-propynyluridine, 5-propynylcytidine, 6-methyladenine, 6-methylguanine, N,N,-dimethyladenine, 2-propyladenine, 2-propylguanine, 2-aminoadenine, 1-methylinosine, 3-methyluridine, 5-methylcytidine, 5-methyluridine, and other nucleotides with modifications at the 5-position, such as 5-(2-amino)propyluridine, 5-halocyt ... uridine, 5-halouridine, 4-acetylcytidine, 1-methyladenosine, 2-methyladenosine, 3-methylcytidine, 6-methyluridine, 2-methylguanosine, 7-methylguanosine, 2,2-dimethylguanosine, 5-methylaminoethyluridine, 5-methyloxyuridine, deazanucleotides such as 7-deaza-adenosine, 6-azouridine, 6-azocytidine, 6-azothymidine, 5-methyloxyuridine, Included among these are aryl-2-thiouridine, other thio bases such as 2-thiouridine and 4-thiouridine and 2-thiocytidine, dihydrouridine, pseudouridine, queousine, archaeosine, naphthyl and substituted naphthyl groups, any O- and N-alkylated purines and pyrimidines such as N6-methyladenosine, 5-methylcarbonylmethyluridine, uridine 5-oxyacetic acid, pyridin-4-one, pyridin-2-one, phenyl and modified phenyl groups such as aminophenol or 2,4,6-trimethoxybenzene, modified cytosines which act as G-clamp nucleotides, 8-substituted adenines and guanines, 5-substituted uracils and thymines, azapyrimidines, carboxyhydroxyalkyl nucleotides, carboxyalkylaminoalkyl nucleotides, and alkylcarbonyl alkylated nucleotides.

[0446] The originator construct and / or benchmark construct may include one or more substitutions, insertions and / or additions, deletions, and covalent modifications relative to the reference sequence, particularly the parent RNA, and are within the scope of the present disclosure.

[0447] In some embodiments, the originator construct and / or the benchmark construct includes one or more post-transcriptional modifications (e.g., capping, cleavage, polyadenylation, splicing, poly-A sequences, methylation, acylation, phosphorylation, methylation of lysine and arginine residues, acetylation, and nitrosylation of thiol groups and tyrosine residues, etc.). The one or more post-transcriptional modifications can be any post-transcriptional modification, such as any of the hundreds of different nucleoside modifications that have been identified in RNA (Rozenski, J, Crain, P, and McCloskey, J. (1999). The RNA Modification Database: 1999 update. Nucl Acids Res 27:196-197). In some embodiments, the first isolated nucleic acid includes messenger RNA (mRNA). In some embodiments, the originator construct and / or the benchmark construct are selected from the group consisting of pyridin-4-one ribonucleosides, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-taurinomethyluridine, 1-taurinomethyl-pseudouridine, 5-taurinomethyl-2-thio-uridine, 1-taurinomethyl-pseudouridine, 5-taurinomethyl-2-thio-uridine, 1-taurinomethyl-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pse ...hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-hydroxyuridine, 3-methyluridine, 5-hydroxyuridine, 3-methyluridine, 5-hydroxyuridine, 3-methyluridine, 5-hydroxyuridine, 3-methyluridine The present invention further comprises at least one nucleoside selected from the group consisting of ethyl-4-thio-uridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine, dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, and 4-methoxy-2-thio-pseudouridine.In some embodiments, the mRNA is 5-aza-cytidine, pseudoisocytidine, 3-methyl-cytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudoisocytidine, pyrrolo-cytidine, pyrrolo-pseudoisocytidine, 2-thio-cytidine, 2-thio-5-methyl-cytidine, 4-thio-pseudoisocytidine, 4-thio-1-methyl-pseudoisocytidine, 4-thio The nucleoside comprises at least one nucleoside selected from the group consisting of o-1-methyl-1-deaza-pseudoisocytidine, 1-methyl-1-deaza-pseudoisocytidine, zebularine, 5-aza-zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, and 4-methoxy-1-methyl-pseudoisocytidine. In some embodiments, the mRNA is 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyladenosine, N6-(cis-hydroxyisopurine), N6-isopropyl adenosine, ... The nucleoside comprises at least one nucleoside selected from the group consisting of N6-(cis-hydroxyisopentenyl)adenosine, N6-glycinylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methylthio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine, and 2-methoxy-adenine.In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of inosine, 1-methyl-inosine, wyosine, wybutosine, 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, and N2,N2-dimethyl-6-thio-guanosine.

[0448] The originator construct and / or benchmark construct may include any useful modification, for example, to the sugar, nucleobase, or internucleoside linkage (e.g., linking phosphate / phosphodiester bond / phosphodiester backbone). One or more atoms of the pyrimidine nucleobase may be replaced or substituted with an optionally substituted amino, an optionally substituted thiol, an optionally substituted alkyl (e.g., methyl or ethyl), or a halo (e.g., chloro or fluoro). In certain embodiments, a modification (e.g., one or more modifications) is present in each of the sugar and the internucleoside linkage. The modification may be a modification from ribonucleic acid (RNA) to deoxyribonucleic acid (DNA), threose nucleic acid (TNA), glycol nucleic acid (GNA), peptide nucleic acid (PNA), locked nucleic acid (LNA), or a hybrid thereof). Additional modifications are described herein.

[0449] In some embodiments, the originator constructs and / or benchmark constructs include at least one N(6) methyl adenosine (m6A) modification to increase translation efficiency. In some embodiments, the N(6) methyl adenosine (m6A) modification may reduce the immunogenicity of the originator constructs and / or benchmark constructs.

[0450] In some embodiments, the modification may include chemical or cell-induced modifications. For example, some non-limiting examples of intracellular RNA modifications are described by Lewis and Pan in "RNA modifications and structures cooperate to guide RNA-protein interactions" from Nat.Reviews Mol.Cell Biol., 2017, 18:202-210.

[0451] In some embodiments, chemical modifications to RNA can improve immune evasion. RNA can be synthesized and / or modified by methods well established in the art, such as those described in "Current protocols in nucleic acid chemistry," Beaucage, SLet al. (Eds.), John Wiley & Sons, Inc., New York, NY, USA, which is hereby incorporated by reference. Modifications include, for example, end modifications, such as 5' end modifications (phosphorylation (mono, di and tri), conjugation, reverse linkage, etc.), 3' end modifications (conjugation, DNA nucleotides, reverse linkage, etc.), base modifications (e.g., replacement with a stabilizing base, a destabilizing base, or a base that base pairs with an extended repertoire of partners), removal of a base (abasic nucleotide), or conjugated base. Modified ribonucleotide bases can also include 5-methylcytidine and pseudouridine. In some embodiments, base modifications can modulate expression, immune response, stability, subcellular localization to specify some functional effect of RNA. In some embodiments, the modifications include biorthogonal nucleotides, e.g., unnatural bases. See, e.g., Kimoto et al., Chem Commun (Camb), 2017, 53:12309, DOI: 10.1039 / c7cc06661a, which is incorporated herein by reference.

[0452] In some embodiments, one or more RNA sugar modifications (e.g., at the 2' or 4' position) or sugar replacements may include modifications or replacements of phosphodiester bonds, as well as backbone modifications. Specific examples of modifications include modified backbones or non-natural internucleoside linkages, such as internucleoside modifications, including modifications or replacements of phosphodiester bonds. RNAs with modified backbones include, among others, those that do not have a phosphorus atom in the backbone. For the purposes of this application, and as sometimes referred to in the art, modified RNAs that do not have a phosphorus atom in their internucleoside backbone may also be considered oligonucleosides. In certain embodiments, RNAs will include ribonucleotides that have a phosphorus atom in their internucleoside backbone.

[0453] Modified RNA backbones can include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkyl phosphotriesters, methyl and other alkyl phosphonates, such as 3'-alkylene phosphonates and chiral phosphonates, phosphinates, phosphoramidates, such as 3'-amino phosphoramidates and aminoalkyl phosphoramidates, thionophosphoramidates, thionoalkyl phosphonates, thionoalkyl phosphotriesters, and boranophosphates with normal 3'-5' linkages, their 2-5' linkage analogs, and those with reversed polarity (adjacent pairs of nucleoside units are linked 3'-5' to 5'-3' or 2'-5' to 5'-2'). Also included are various salts, mixed salts, and free acid forms. In some embodiments, RNA can be negatively or positively charged.

[0454] Modified nucleotides may be modified on the internucleoside bond (e.g., the phosphate backbone). In the context of polynucleotide backbones, the terms "phosphate" and "phosphodiester" are used interchangeably herein. The backbone phosphate group may be modified by replacing one or more of the oxygen atoms with different substituents. In addition, modified nucleosides and nucleotides may include large-scale replacement of unmodified phosphate sites with alternative internucleoside linkages as described herein. Examples of modified phosphate groups include, but are not limited to, phosphorothioates, phosphoroselenates, boranophosphates, boranophosphate esters, hydrogen phosphonates, phosphoramidates, phosphorodiamidates, alkyl or aryl phosphonates, and phosphotriesters. Phosphorodithioates have both non-linked oxygens replaced by sulfur. Phosphate linkers can also be modified by replacement of the linking oxygen with nitrogen (bridging phosphoramidates), sulfur (bridging phosphorothioates), and carbon (bridging methylene-phosphonates).

[0455] The a-thio substituted phosphate moieties are provided to provide stability to RNA and DNA polymers through non-natural phosphorothioate backbone linkages. Phosphorothioate DNA and RNA have increased nuclease resistance and subsequently longer half-life in cellular environments. Phosphorothioates linked to RNA are expected to reduce innate immune responses through weaker binding / activation of cellular innate immune molecules.

[0456] In certain embodiments, modified nucleosides include alpha-thio-nucleosides (e.g., 5'-O-(1-thiophosphate)-adenosine, 5'-O-(1-thiophosphate)-cytidine (a-thio-cytidine), 5'-O-(1-thiophosphate)-guanosine, 5'-O-(1-thiophosphate)-uridine, or 5'-O-(1-thiophosphate)-pseudouridine).

[0457] Other internucleoside linkages that can be used in accordance with the present disclosure, including internucleoside linkages that do not contain a phosphorus atom, are described herein.

[0458] In some embodiments, the RNA can include one or more cytotoxic nucleosides, for example, cytotoxic nucleosides can be incorporated into the RNA, such as bifunctional modifications. Cytotoxic nucleosides may include, but are not limited to, adenosine arabinoside, 5-azacytidine, 4'-thio-aracytidine, cyclopentenylcytosine, cladribine, clofarabine, cytarabine, cytosine arabinoside, 1-(2-C-cyano-2-deoxy-beta-D-arabino-pentofuranosyl)-cytosine, decitabine, 5-fluorouracil, fludarabine, floxuridine, gemcitabine, a combination of tegafur and uracil, tegafur ((RS)-5-fluoro-1-(tetrahydrofuran-2-yl)pyrimidine-2,4(1H,3H)-dione), troxacitabine, tezacitabine, 2'...

Claims

1. Formula (X): 【Chemistry 80】 or a pharmaceutically acceptable salt thereof. (In the formula, each cc is independently selected from 3 to 9; R xx represents hydrogen and optionally substituted C 1 -C 6 alkyl; (i) ee is 1; each dd is independently selected from 1 to 4; Each R ww are independently 4 -C 14 Alkyl, branched C 4 -C 12 Alkenyl, C containing at least two double bonds 4 -C 12 Alkenyl, and C 9 -C 12 alkenyl, wherein C 4 -C 14 Any of -(CH 2 ) 2 - is C3-C 6 optionally substituted with cycloalkylenyl; or (ii) ee is 0; Each dd is 1; Each R ww is a linear C 4 -C 12 alkyl).

2. R xx is H, or a pharmaceutically acceptable salt thereof.

3. The compound of claim 1, or a pharmaceutically acceptable salt thereof, wherein ee is 1 and each dd is independently selected from 1 to 4; and each R ww is independently selected from the group consisting of C 4 -C 14 alkyl, branched C 4 -C 12 alkenyl, C 4 -C 12 alkenyl containing at least two double bonds, and C 9 -C 12 alkenyl, wherein any -(CH 2 ) 2 - in said C 4 -C 14 alkyl is optionally replaced by C 3 -C 6 cycloalkylenyl.

4. The compound of claim 1, or a pharmaceutically acceptable salt thereof, wherein dd is 1.

5. The compound of claim 1, or a pharmaceutically acceptable salt thereof, wherein dd is 3.

6. The compound of claim 1, or a pharmaceutically acceptable salt thereof, wherein cc is 3 to 7.

7. The compound of claim 1, or a pharmaceutically acceptable salt thereof, wherein each R ww is C 4 -C 14 alkyl, and any —(CH 2 ) 2 — in said C 4 -C 14 alkyl is optionally replaced by C 3 -C 6 cycloalkylenyl.

8. The compound of claim 1, or a pharmaceutically acceptable salt thereof, wherein each R ww is C 4 -C 14 alkyl, and any —(CH 2 ) 2 — in said C 4 -C 14 alkyl is optionally replaced with cyclopropylene.

9. The compound of claim 1, or a pharmaceutically acceptable salt thereof, wherein each R ww is C 9 -C 12 alkenyl.

10. Each R ww is a C containing at least two double bonds 8 -C 12 10. The compound of claim 1, or a pharmaceutically acceptable salt thereof, which is alkenyl.

11. The compound of claim 1, or a pharmaceutically acceptable salt thereof, wherein each R ww is branched C 4 -C 12 alkenyl.

12. The compound of claim 1, or a pharmaceutically acceptable salt thereof, wherein ee is 0, each dd is 1, and each R ww is linear C 4 -C 12 alkyl.

13. The compound of formula (X) is a compound of formula (X-A): 【Chemistry 81】 (In the formula, each cc is independently selected from 3 to 7; each dd is independently selected from 1 to 4; R xx represents hydrogen and optionally substituted C 1 -C 6 alkyl; Each R ww are independently 4 -C 14 Alkyl or (linear or branched C 3 -C 5 alkylenyl)-(branched C 5 -C 7 2. The compound of claim 1, wherein R is selected from the group consisting of: aryl, aryl(s), ... 【Request 14】 【Table 3-1】 【Table 3-2】 【Table 3-3】 【Table 3-4】 【Table 3-5】 【Table 3-6】 【Table 3-7】 【Table 3-8】 【Table 3-9】 2. The compound of claim 1, selected from the group consisting of: or a pharmaceutically acceptable salt thereof.

15. Formula (X): 【Chemistry 80】 or a pharmaceutically acceptable salt thereof. (In the formula, each cc is independently selected from 3 to 9; R xx is selected from hydrogen and optionally substituted C 1 -C 6 alkyl; (i) ee is 1; each dd is independently selected from 1 to 4; each R ww is independently selected from the group consisting of C 4 -C 14 alkyl, branched C 4 -C 12 alkenyl, C 4 -C 12 alkenyl containing at least two double bonds, and C 9 -C 12 alkenyl, wherein any —(CH 2 ) 2 — in said C 4 -C 14 alkyl is optionally replaced with C 2 -C 6 cycloalkylenyl; or (ii) ee is 0; Each dd is 1; Each R ww is a linear C 4 -C 12 alkyl.

16. The ionizable lipid of formula (X): 【Table 4-1】 【Table 4-2】 【Table 4-3】 【Table 4-4】 【Table 4-5】 【Table 4-6】 【Table 4-7】 【Table 4-8】 【Table 4-9】 or a pharmaceutically acceptable salt thereof.

17. (a) PEG-lipid (b) structured lipids; and (c) non-ionizable lipids and / or zwitterionic lipids The lipid nanoparticle of claim 15, further comprising:

18. 18. The LNP of claim 17, wherein the PEG-lipid is selected from the group consisting of PEG-c-DOMG, PEG-DMG, PEG-DLPE, PEG-DMPE, PEG-DPPC, and PEG-DSPE.

19. 18. The LNP of claim 17, wherein the structural lipid is selected from the group consisting of cholesterol, fecosterol, sitosterol, ergosterol, campesterol, stigmasterol, brassicasterol, tomatidine, ursolic acid, and alpha-tocopherol.

20. The non-ionizable lipids may be 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC), 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), 1,2-dilinoleoyl-sn-glycero-3-phosphocholine (DLPC), 1,2-dimyristoyl-sn-glycero-phosphocholine (DMPC), 1,2-dioleoyl-sn-glycero-3-phosphocholine (DOPC), 1,2-dipalmitoyl-sn-glycero-3-phosphocholine (DPPC), 1,2-diundecanoyl-sn-glycero-phosphocholine (DUPC), 1-palmitoyl-2-oleoyl-sn-glycero-3-phosphocholine (POPC), 1,2-di-O-octadecenyl-sn-glycero-3-phosphocholine (18:0 diether). PC), 1-oleoyl-2-cholesterylhemisuccinoyl-sn-glycero-3-phosphocholine (OChemsPC), 1-hexadecyl-sn-glycero-3-phosphocholine (C16 Lyso PC), 1,2-dilinolenoyl-sn-glycero-3-phosphocholine, 1,2-diarachidonoyl-sn-glycero-3-phosphocholine, 1,2-didocosahexaenoyl-sn-glycero-3-phosphocholine, 1,2-diphytanoyl-sn-glycero-3-phosphoethanolamine (ME 16.0 PE), 1,2-distearoyl-sn-glycero-3-phosphoethanolamine, 1,2-dilinoleoyl-sn-glycero-3-phosphoethanolamine, 1,2-dilinolenoyl-sn-glycero-3-phosphoethanolamine, 1,2-diarachidonoyl-sn-glycero-3-phosphoethanolamine, 1,2-didocosahexaenoyl-sn-glycero-3-phosphoethanolamine, 1,2-dioleoyl-sn-glycero-3-phospho-rac-(1-glycerol) sodium salt (DOPG), sodium (S)-2-ammonio-3-((((R)-2-(oleoyloxy)-3-(stearoyloxy)propoxy)oxidophosphoryl)oxy)propanoate (L-α-phosphatidylserine;Brain PS), dimyristoylphosphatidylcholine (DMPC), dimyristoylphosphoethanolamine (DMPE), dimyristoylphosphatidylglycerol (DMPG), dioleoyl-phosphatidylethanolamine 4-(N-maleimidomethyl)-cyclohexane-1-carboxylate (DOPE-mal), dioleoylphosphatidylglycerol (DOPG), 1,2-dioleoyl-sn-glycero-3-(phospho-L-serine) (DOPS), acetyl-phospholipid (ADP ... ), dipalmitoylphosphatidylethanolamine (DPPE), dipalmitoylphosphatidylglycerol (DPPG), dipalmitoylphosphatidylserine (DPPS), distearoyl-phosphatidyl-ethanolamine (DSPE), distearoylphosphoethanolamine imidazole (DSPEI), 1,2-diundecanoyl-sn-glycero-phosphocholine (DUPC), egg yolk phosphatidylcholine (EPC), 1,2-dioleoyl-sn-glycero-3-phosphate (18:1 PA; DOPA), ammonium bis((S)-2-hydroxy-3-(oleoyloxy)propyl)phosphate (18:1 DMP; LBPA), 1,2-dioleoyl-sn-glycero-3-phospho-(1'-myo-inositol) (DOPI; 18:1 PI), 1,2-distearoyl-sn-glycero-3-phospho-L-serine (18:0 PS), 1,2-dilinoleoyl-sn-glycero-3-phospho-L-serine (18:2 PS), 1-palmitoyl-2-oleoyl-sn-glycero-3-phospho-L-serine (16:0-18:1 PS; POPS), 1-stearoyl-2-oleoyl-sn-glycero-3-phospho-L-serine (18:0-18:1 18:1 Lyso PS), 1-stearoyl-2-linoleoyl-sn-glycero-3-phospho-L-serine (18:0-18:2 PS), 1-oleoyl-2-hydroxy-sn-glycero-3-phospho-L-serine (18:1 Lyso PS), 1-stearoyl-2-hydroxy-sn-glycero-3-phospho-L-serine (18:0 Lyso PS), and sphingomyelin.

21. The LNP of claim 17, comprising approximately 48.5 mol% ionizable lipids, approximately 10 mol% phospholipids, approximately 39 mol% structural lipids, and approximately 2.5 mol% PEG lipids.

22. The LNP of claim 17, comprising approximately 48.5 mol% ionizable lipids, approximately 10 mol% phospholipids, approximately 40 mol% structural lipids, and approximately 1.5 mol% PEG lipids.

23. (A) Formula (X): 【Chemistry 80】 or a pharmaceutically acceptable salt thereof. (In the formula, each cc is independently selected from 3 to 9; R xx is selected from hydrogen and optionally substituted C 1 -C 6 alkyl; (i) ee is 1; each dd is independently selected from 1 to 4; each R ww is independently selected from the group consisting of C 4 -C 14 alkyl, branched C 4 -C 12 alkenyl, C 4 -C 12 alkenyl containing at least two double bonds, and C 9 -C 12 alkenyl, wherein any —(CH 2 ) 2 — in said C 4 -C 14 alkyl is optionally replaced with C 2 -C 6 cycloalkylenyl; or (ii) ee is 0; Each dd is 1; each R ww is a linear C 4 -C 12 alkyl; and (B) Coding RNA A lipid nanoparticle (LNP) comprising:

24. The ionizable lipid of formula (X): 【Table 5-1】 【Table 5-2】 【Table 5-3】 【Table 5-4】 【Table 5-5】 【Table 5-6】 【Table 5-7】 【Table 5-8】 【Table 5-9】 or a pharmaceutically acceptable salt thereof.

25. (a) PEG-lipid (b) structured lipids; and (c) non-ionizable lipids and / or zwitterionic lipids The lipid nanoparticle of claim 23, further comprising:

26. 26. The LNP of claim 25, wherein the PEG-lipid is selected from the group consisting of PEG-c-DOMG, PEG-DMG, PEG-DLPE, PEG-DMPE, PEG-DPPC, and PEG-DSPE.

27. 26. The LNP of claim 25, wherein the structural lipid is selected from the group consisting of cholesterol, fecosterol, sitosterol, ergosterol, campesterol, stigmasterol, brassicasterol, tomatidine, ursolic acid, and alpha-tocopherol.

28. The non-ionizable lipids may be 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC), 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), 1,2-dilinoleoyl-sn-glycero-3-phosphocholine (DLPC), 1,2-dimyristoyl-sn-glycero-phosphocholine (DMPC), 1,2-dioleoyl-sn-glycero-3-phosphocholine (DOPC), 1,2-dipalmitoyl-sn-glycero-3-phosphocholine (DPPC), 1,2-diundecanoyl-sn-glycero-phosphocholine (DUPC), 1-palmitoyl-2-oleoyl-sn-glycero-3-phosphocholine (POPC), 1,2-di-O-octadecenyl-sn-glycero-3-phosphocholine (18:0 diether). PC), 1-oleoyl-2-cholesterylhemisuccinoyl-sn-glycero-3-phosphocholine (OChemsPC), 1-hexadecyl-sn-glycero-3-phosphocholine (C16 Lyso PC), 1,2-dilinolenoyl-sn-glycero-3-phosphocholine, 1,2-diarachidonoyl-sn-glycero-3-phosphocholine, 1,2-didocosahexaenoyl-sn-glycero-3-phosphocholine, 1,2-diphytanoyl-sn-glycero-3-phosphoethanolamine (ME 16.0 PE), 1,2-distearoyl-sn-glycero-3-phosphoethanolamine, 1,2-dilinoleoyl-sn-glycero-3-phosphoethanolamine, 1,2-dilinolenoyl-sn-glycero-3-phosphoethanolamine, 1,2-diarachidonoyl-sn-glycero-3-phosphoethanolamine, 1,2-didocosahexaenoyl-sn-glycero-3-phosphoethanolamine, 1,2-dioleoyl-sn-glycero-3-phospho-rac-(1-glycerol) sodium salt (DOPG), sodium (S)-2-ammonio-3-((((R)-2-(oleoyloxy)-3-(stearoyloxy)propoxy)oxidophosphoryl)oxy)propanoate (L-α-phosphatidylserine;Brain PS), dimyristoylphosphatidylcholine (DMPC), dimyristoylphosphoethanolamine (DMPE), dimyristoylphosphatidylglycerol (DMPG), dioleoyl-phosphatidylethanolamine 4-(N-maleimidomethyl)-cyclohexane-1-carboxylate (DOPE-mal), dioleoylphosphatidylglycerol (DOPG), 1,2-dioleoyl-sn-glycero-3-(phospho-L-serine) (DOPS), acetyl-phospholipid (ADP ... ), dipalmitoylphosphatidylethanolamine (DPPE), dipalmitoylphosphatidylglycerol (DPPG), dipalmitoylphosphatidylserine (DPPS), distearoyl-phosphatidyl-ethanolamine (DSPE), distearoylphosphoethanolamine imidazole (DSPEI), 1,2-diundecanoyl-sn-glycero-phosphocholine (DUPC), egg yolk phosphatidylcholine (EPC), 1,2-dioleoyl-sn-glycero-3-phosphate (18:1 PA; DOPA), ammonium bis((S)-2-hydroxy-3-(oleoyloxy)propyl)phosphate (18:1 DMP; LBPA), 1,2-dioleoyl-sn-glycero-3-phospho-(1'-myo-inositol) (DOPI; 18:1 PI), 1,2-distearoyl-sn-glycero-3-phospho-L-serine (18:0 PS), 1,2-dilinoleoyl-sn-glycero-3-phospho-L-serine (18:2 PS), 1-palmitoyl-2-oleoyl-sn-glycero-3-phospho-L-serine (16:0-18:1 PS; POPS), 1-stearoyl-2-oleoyl-sn-glycero-3-phospho-L-serine (18:0-18:1 26. The LNP of claim 25, wherein the phospholipid is selected from the group consisting of 1-stearoyl-2-linoleoyl-sn-glycero-3-phospho-L-serine (18:0-18:2 PS), 1-oleoyl-2-hydroxy-sn-glycero-3-phospho-L-serine (18:1 Lyso PS), 1-stearoyl-2-hydroxy-sn-glycero-3-phospho-L-serine (18:0 Lyso PS), and sphingomyelin;

29. The LNP described in Claim 25, wherein the coding RNA is mRNA.

30. The LNP described in claim 25, wherein the coding RNA is circRNA.