Cyclic lipids and methods of use thereof

TWI938369BActive Publication Date: 2026-09-11RENAGADE THERAPEUTICS MANAGEMENT INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
TW111134779
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-04-28
Filing Date
2022-09-14
Publication Date
2026-09-11
Estimated Expiration
2042-09-13

AI Technical Summary

Technical Problem

Current lipid-based delivery systems for nucleic acids and proteins lack targeted delivery capabilities, focusing primarily on cargo protection rather than localization and delivery to specific cells or tissues.

Method used

Development of novel lipids and a tropism discovery platform for evaluating targeting systems, enabling localized delivery of nucleic acid and protein therapeutics to immune cells using lipid nanoparticles and adeno-associated virus libraries.

Benefits of technology

The system achieves targeted delivery of nucleic acids and proteins to specific cells, enhancing therapeutic efficacy by up to 100% inhibition of target expression and eliciting an immune response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TB001909929_001
    Figure TWG2TB001909929_001
  • Figure TWG2TB001909929_002
    Figure TWG2TB001909929_002
  • Figure TWG2TB001909929_003
    Figure TWG2TB001909929_003
Patent Text Reader

Abstract

This invention details various lipids, compositions, and / or optimized systems and delivery vectors for delivering nucleic acid sequences, polypeptides, or peptides for vaccination against infectious agents.
Need to check novelty before this filing date? Find Prior Art

Description

technical field

[0001] The present invention relates to optimized systems for the delivery of nucleic acid sequences, polypeptides or peptides and methods of using such optimized systems to treat diseases, disorders and / or conditions. prior art

[0002] Proteins have always been the standard of care, but over the past few years, the use of nucleic acids as a therapeutic modality for a variety of diseases and therapeutic indications has gained prominence. Several companies have shown that nucleic acids (such as siRNA, mRNA, circular RNA, DNA, ASO, etc.) may be more effective when compared to protein-based therapies, but targeted delivery systems are required for both nucleic acid and protein-based therapies In order to ensure that the therapeutic agent is localized to the targeted cells, tissues or organs.

[0003] Current delivery systems, including lipid-based delivery systems such as lipid nanoparticles, focus on protecting the delivered cargo, but not on the lipids being used by the delivery system, and often not on the location of the cargo or the delivery system deliver. There is a need in the art for improved lipid-based delivery systems. Contents of the invention

[0004] The present invention provides novel lipids useful in delivery vehicles for delivery systems and a tropism discovery platform for screening and developing targeting systems for targeted delivery of nucleic acid and protein therapeutics, eg, to immune cells.

[0005] In one aspect of the invention, provided herein is a lipid having any of formulas (CY) and (CY-I)-(CY-IX), or a pharmaceutically acceptable salt or solvate thereof substances, or any of the lipids in Table (I) or their salts or solvates, see below, are collectively referred to as "lipids of the invention" and each individually as "lipids of the invention".

[0006] In one aspect of the present invention, a pharmaceutical composition is provided herein, comprising: a) a polynucleotide encoding at least one protein of interest, and b) a delivery vehicle comprising at least one lipid, wherein the composition elicits an immune response in the individual.

[0007] In one aspect, the polynucleotide is DNA.

[0008] In one aspect, the polynucleotide is RNA.

[0009] In one aspect, the RNA is short interfering RNA (siRNA).

[0010] In one aspect, the siRNA inhibits or suppresses the expression of a target of interest in a cell.

[0011] In one aspect, the inhibition or inhibition is about 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95% and 100%, or at least 20-30%, 20- 40%, 20-50%, 20-60%, 20-70%, 20-80%, 20-90%, 20-95%, 20-100%, 30-40%, 30-50%, 30- 60%, 30-70%, 30-80%, 30-90%, 30-95%, 30-100%, 40-50%, 40-60%, 40-70%, 40-80%, 40- 90%, 40-95%, 40-100%, 50-60%, 50-70%, 50-80%, 50-90%, 50-95%, 50-100%, 60-70%, 60- 80%, 60-90%, 60-95%, 60-100%, 70-80%, 70-90%, 70-95%, 70-100%, 80-90%, 80-95%, 80- 100%, 90-95%, 90-100%, or 95-100%.

[0012] In one aspect, the polynucleotides are substantially circular.

[0013] In one aspect, the polynucleotide comprises an internal ribosome entry site (IRES) sequence operably linked to the payload sequence region.

[0014] In one aspect, the IRES sequence comprises a sequence derived from picornavirus complementary DNA, encephalomyocarditis virus (EMCV) complementary DNA, poliovirus complementary DNA, or the antennapedia gene from Drosophila melanogaster.

[0015] In one aspect, the polynucleotide comprises a termination element, wherein the termination element comprises at least one stop codon.

[0016] In one aspect, a polynucleotide comprises regulatory elements.

[0017] In one aspect, the polynucleotide comprises at least one masking agent.

[0018] In one aspect, substantially circular polynucleotides are produced using in vitro transcription.

[0019] In one aspect, the payload sequence region comprises non-coding nucleic acid sequences.

[0020] In one aspect, the payload sequence region comprises a coding nucleic acid sequence.

[0021] In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Campylobacter jejuni. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Clostridium difficile. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Entamoeba histolytica. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for enterotoxin B. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Norwalk virus or norovirus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Helicobacter pylori. In one aspect, the encoding nucleic acid sequence encodes a rotavirus protein of interest. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Candida yeast. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for a coronavirus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for SARS-CoV. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for SARS-CoV-2. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for MERS-CoV. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for enterovirus 71. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Epstein-Barr virus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Gram-negative bacteria. In one aspect, the Gram-negative bacteria is Bordetella. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Gram-positive bacteria. In one aspect, the Gram-positive bacterium is Clostridium tetani. In one aspect, the Gram-positive bacterium is Francisella tularensis. In one aspect, the Gram-positive bacteria are Streptococcus bacteria. In one aspect, the Gram-positive bacteria are Staphylococcus bacteria. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for hepatitis. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Human Cytomegalovirus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest to Human Immunodeficiency Virus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Human Papilloma Virus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for influenza. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for John Cunningham Virus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Mycobacterium. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Poxviruses. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Pseudomonas aeruginosa. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Respiratory Syncytial Virus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Rubella virus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Varicella zoster virus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Chikungunya virus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Dengue virus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Rabies virus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Trypanosoma cruzi and / or Chagas disease. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Ebola virus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Plasmodium falciparum. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Marburg virus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Japanese encephalitis virus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for St. Louis encephalitis virus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for West Nile Virus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Yellow Fever virus. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Bacillus anthracis. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Botulinum toxin. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for Ricin. In one aspect, the encoding nucleic acid sequence encodes a protein of interest for a Shiga toxin and / or a Shiga-like toxin.

[0022] In one aspect, a polynucleotide comprises at least one modification.

[0023] In one aspect, at least 20% of the bases are modified. In one aspect, at least 30% of the bases are modified. In one aspect, at least 40% of the bases are modified. In one aspect, at least 50% of the bases are modified. In one aspect, at least 60% of the bases are modified. In one aspect, at least 70% of the bases are modified. In one aspect, at least 80% of the bases are modified. In one aspect, at least 90% of the bases are modified. In one aspect, at least 100% of the bases are modified. In one aspect, a particular base contains at least one modification.

[0024] In one aspect, the base is adenine. In one aspect, at least 20% of the adenine bases are modified. In one aspect, at least 30% of the adenine bases are modified. In one aspect, at least 40% of the adenine bases are modified. In one aspect, at least 50% of the adenine bases are modified. In one aspect, at least 60% of the adenine bases are modified. In one aspect, at least 70% of the adenine bases are modified. In one aspect, at least 80% of the adenine bases are modified. In one aspect, at least 90% of the adenine bases are modified. In one aspect, at least 100% of the adenine bases are modified.

[0025] In one aspect, the base is guanine. In one aspect, at least 20% of the guanine bases are modified. In one aspect, at least 30% of the guanine bases are modified. In one aspect, at least 40% of the guanine bases are modified. In one aspect, at least 50% of the guanine bases are modified. In one aspect, at least 60% of the guanine bases are modified. In one aspect, at least 70% of the guanine bases are modified. In one aspect, at least 80% of the guanine bases are modified. In one aspect, at least 90% of the guanine bases are modified. In one aspect, at least 100% of the guanine bases are modified.

[0026] In one aspect, the base is cytosine. In one aspect, at least 20% of the cytosine bases are modified. In one aspect, at least 30% of the cytosine bases are modified. In one aspect, at least 40% of the cytosine bases are modified. In one aspect, at least 50% of the cytosine bases are modified. In one aspect, at least 60% of the cytosine bases are modified. In one aspect, at least 70% of the cytosine bases are modified. In one aspect, at least 80% of the cytosine bases are modified. In one aspect, at least 90% of the cytosine bases are modified. In one aspect, at least 100% of the cytosine bases are modified.

[0027] In one aspect, the base is uracil. In one aspect, at least 20% of the uracil bases are modified. In one aspect, at least 30% of the uracil bases are modified. In one aspect, at least 40% of the uracil bases are modified. In one aspect, at least 50% of the uracil bases are modified. In one aspect, at least 60% of the uracil bases are modified. In one aspect, at least 70% of the uracil bases are modified. In one aspect, at least 80% of the uracil bases are modified. In one aspect, at least 90% of the uracil bases are modified. In one aspect, at least 100% of the uracil bases are modified.

[0028] In one aspect, the at least one modification is pyridin-4-ketoribonucleoside, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4- Thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5- Proynyl-uridine, 1-propynyl-pseudouridine, 5-taurine methyluridine, 1-taurine methyl-pseudouridine, 5-taurine methyl-2-thio Base-uridine, 1-taurine methyl-4-thio-uridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine Glycoside, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, two Hydrogenuridine, dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thio- Uridine, 4-methoxy-pseudouridine, and 4-methoxy-2-thio-pseudouridine, 5-aza-cytidine, pseudoisocytidine, 3-methyl-cytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudoisocytidine, pyrrolo-cytidine, pyrrolo- Pseudoisocytidine, 2-thio-cytidine, 2-thio-5-methyl-cytidine, 4-thio-pseudo-cytidine, 4-thio-1-methyl-pseudo-cytidine , 4-thio-1-methyl-1-deaza-pseudoisocytidine, 1-methyl-1-deaza-pseudoisocytidine, zebularine, 5-aza-ze Braline, 5-Methyl-Zebralin, 5-Aza-2-thio-Zebralin, 2-thio-Zebralin, 2-Methoxy-Cytidine, 2- Methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, and 4-methoxy-1-methyl-pseudoisocytidine, 2-aminopurine, 2,6- Diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-amine Base purine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-prenyl adenosine, N6-(cis-hydroxyprenyl)adenosine, 2-methylthio-N6-(cis-hydroxyprenyl)adenosine, N6-glycyl adenosine Aminoformyl adenosine, N6-threonylaminoformyl adenosine, 2-methylthio-N6-hydroxybutylaminoformyl adenosine, N6,N6-dimethyladenosine , 7-methyladenine, 2-methylthio-adenine, and 2-methoxy-adenine, inosine, 1-methyl-inosine, wyosine, wyotin ( wybutosine), 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-de Aza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methyl Guanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6- Thio-guanosine, N2-methyl-6-thio-guanosine or N2,N2-dimethyl-6-thio-guanosine.

[0029] In one aspect, the pharmaceutical composition comprises at least one cationic lipid selected from the group consisting of any lipid in Table (I), any lipid having the structure of formula (CY-I), any lipid having the formula Any lipid with the structure of (CY-II), any lipid with the structure of formula (CY-III), any lipid with the structure of formula (CY-IV), any lipid with the structure of formula (CY-V) Any lipid, any lipid having the structure of formula (CY-VI), and combinations thereof.

[0030] In one aspect, the cationic lipid is any lipid having the structure of formula (CY-I).

[0031] In one aspect, the cationic lipid is selected from the group consisting of compounds CY1, CY2, CY3, CY9, CY10, CY11, CY12, CY22, CY23, CY24, CY30, CY31, CY32, CY33, CY43, CY44, CY45 , CY50, CY51, CY52 and CY53.

[0032] In one aspect, the cationic lipid is any lipid having the structure of formula (CY-II).

[0033] In one aspect, the cationic lipid is selected from the group consisting of compounds CY4, CY5, CY16, CY17, CY18, CY25, CY26, CY37, CY38, CY39, CY46, CY56, and CY57.

[0034] In one aspect, the cationic lipid is any lipid having the structure of formula (CY-III).

[0035] In one aspect, the cationic lipid is selected from the group consisting of compounds CY6, CY14, CY27, CY35, CY47, and CY55.

[0036] In one aspect, the cationic lipid is any lipid having the structure of formula (CY-IV).

[0037] In one aspect, the cationic lipid is selected from the group consisting of compounds CY7, CY8, CY19, CY20, CY21, CY28, CY29, CY40, CY41, CY42, CY48, CY49, CY58, CY59, and CY60.

[0038] In one aspect, the cationic lipid is any lipid having the structure of formula (CY-V).

[0039] In one aspect, the cationic lipid is any lipid having the structure of formula (CY-VI).

[0040] In one aspect, the pharmaceutical composition includes additional cationic lipids.

[0041] In one aspect, the pharmaceutical composition comprises neutral lipids.

[0042] In one aspect, the pharmaceutical composition comprises anionic lipids.

[0043] In one aspect, the pharmaceutical composition includes a helper lipid.

[0044] In one aspect, the pharmaceutical composition comprises a stealth lipid.

[0045] In one aspect, the weight ratio of lipid to polynucleotide is from about 100:1 to about 1:1.

[0046] In one aspect, the pharmaceutical composition delivers a cargo or payload to the immune cells of an individual in need thereof. The immune cells can be T cells, such as CD8+ T cells, CD4+ T cells, or T regulatory cells. Immune cells can also be, for example, macrophages or dendritic cells.

[0047] In one aspect, the vaccine formulation comprises a pharmaceutical composition.

[0048] In one aspect, the vaccine is prepared to have any of formulas (I)-(VI).

[0049] In one aspect, provided herein is a method of vaccinating an individual against an infectious agent comprising contacting the individual with a vaccine formulation or preparation and eliciting an immune response.

[0050] In one aspect, the infectious agent is Campylobacter jejuni, Clostridium difficile, Entamoeba dysenteriae, enterotoxin B, Norwalk virus or norovirus, Helicobacter pylori, rotavirus, Candida yeast, coronavirus (including SARS-CoV, SARS-CoV-2 and MERS-CoV), enterovirus 71, Epstein-Barr virus, gram-negative bacteria (including Bordetella), gram-positive bacteria (including tetanus Clostridium, Francisella tularensis, Streptococcus and Staphylococcus), and hepatitis, human cytomegalovirus, human immunodeficiency virus, human papillomavirus, influenza, John Cunningham virus, mycobacteria , poxvirus, Pseudomonas aeruginosa, respiratory syncytial virus, rubella virus, varicella zoster virus, chikungunya virus, dengue virus, rabies virus, Trypanosoma cruzi and / or Chogers disease, Ebola Virus, Plasmodium falciparum, Marburg virus, Japanese encephalitis virus, St. Louis encephalitis virus, West Nile virus, yellow fever virus, Bacillus anthracis, botulinum toxin, ricin or Shiga toxin and / or Shiga-like toxin.

[0051] In one aspect, the exposure is enteral (into the intestine), gastrointestinal tract, epidural (into the dura mater), oral (by means of the mouth), transdermal, intracerebral (into the brain), ventricle Intradermal (into the ventricle), epidermal (applied onto the skin), intradermal (into the skin itself), subcutaneous (under the skin), nasal (through the nose), intravenous (into a vein), intravenous Intraarterial (into an artery), intramuscular (into a muscle), intracardiac (into the heart), intraosseous infusion (into the bone marrow), intrathecal (into the spinal canal) ), intraparenchymal (into brain tissue), intraperitoneal (infusion or injection into the peritoneum), intravesical infusion, intravitreal (via the eye), intracavernous injection (into pathological cavity), intracavity (into in the base of the penis), intravaginal administration, intrauterine, extraamniotic administration, transdermal (diffuse through intact skin for systemic dispersion), transmucosal (diffuse through mucous membranes), vaginal, insufflation (snorting), sublingual, Under the lips, enema, eye drops (onto the conjunctiva), ear drops, aural (in or by the ear), buccal (directed toward the cheek), transconjunctival, dermal, dental (to one or multiple teeth), electroosmosis, intracervical, intrasinus, intratracheal, in vitro, hemodialysis, infiltration, interstitial, intraabdominal, intraamniotic, intraarticular, intrabiliary, intrabronchial, intracystic, intrachondral ( In the cartilage), in the tail (in the cauda equina), in the cistern (in the cerebellomedullary cisterna), in the cornea (in the cornea), in the crown of the teeth, in the coronary arteries (in the coronary arteries), in the corpus cavernosum (in the In the expandable space of the corpus cavernosum), in the intervertebral disc (in the intervertebral disc), in the canal (in the gland duct), in the duodenum (in the duodenum), in the dura mater (in or below the dura mater) ), intraepidermal (to the epidermis), intraesophageal (to the esophagus), intragastric (in the stomach), intragingival (in the gums), intraileal (in the distal part of the small intestine), intralesional (in the localized lesion Intramedullary or directly introduced to localized lesions), intraluminal (in the lumen), intralymphatic (in the lymph), intramedullary (in the bone marrow cavity of the bone), intrameningeal (inside the meninges), intramyocardial (inside the myocardium), intraocular (inside the eye), intraocular (inside the ovary), intrapericardial (inside the pericardium), intrapleural (inside the pleura), intraprostatic (inside the prostate gland), lung Intra (in the lungs or their bronchi), intrasinus (in the nasal or periorbital sinuses), intraspinal (in the spinal column), intrasynovial (in the synovial cavity of the joint), intratendon (in the tendon) , intratesticular (within the testis), intrathecal (within the cerebrospinal fluid at any level of the cerebrospinal axis), intrathoracic (within the chest), intraductal (within the small ducts of the organ), intratumoral (within the tumor intratympanic (inside the middle ear (aurus media)), intravascular (inside one or more blood vessels), intraventricular (in a chamber), iontophoresis (in which ions of soluble salts are transported to the body by means of an electric current tissue), lavage (washing or flushing of open wounds or body cavities), translaryngeal (directly to the larynx), nasogastric tube (through the nose and into the stomach), occlusive dressing techniques (administered locally followed by occlusive area covered by dressings), ophthalmic (outside the eyes), oropharyngeal (directly to the mouth and pharynx), parenteral, transdermal, periarticular, epidural, perineural, periodontal, rectal, respiratory ( For local or systemic action by oral or nasal inhalation into the respiratory tract), retrobulbar (retropontine or retrobulbar), soft tissue, subarachnoid, subconjunctival, submucosa, topical, transplacental (via or transplacental), transtracheal (through tracheal wall), transtympanic (across or through tympanic cavity), ureteral (to ureter), urethral (to urethra), transvaginal, caudal block, diagnostic, Nerve block, biliary perfusion, cardiac perfusion, photoablation, or transspinal. Brief description of the diagram

[0052] [picture] [1] is a diagram illustrating an embodiment of the tropism exploration platform of the present invention.

[0053] [picture] [2] is a diagram showing the initial polynucleotide construct of the present invention, which can be linear or circular.

[0054] [picture] [3A] is a diagram illustrating a series of reference polynucleotide constructs of the present invention, which may include at least one barcode region (barcode region; BC) and / or reverse barcode region (CB) and a payload region (P ).

[0055] [picture] [3B] is a diagram depicting a series of reference polynucleotide constructs of the present invention, wherein the barcode region (BC) or reverse barcode region (CB) can overlap with the payload region (P).

[0056] [picture] [3C] is a diagram depicting a series of reference polynucleotide constructs of the present invention, which may include at least one tag and / or marker.

[0057] [picture] [4A] is a diagram depicting a series of circular reference polynucleotide constructs of the present invention, which may include at least one barcode region (BC) and / or reverse barcode region (CB) and a payload region (P) .

[0058] [picture] [4B] is a diagram showing a series of circular reference polynucleotide constructs of the present invention, wherein the barcode region (BC) or reverse barcode region (CB) can overlap with the payload region (P).

[0059] [picture] [4C] is a diagram depicting a series of reference polynucleotide constructs of the present invention, which may include at least one tag and / or marker.

[0060] [picture] [5] is a diagram showing a series of delivery vectors of the present invention. Implementation

[0061] Cross-references to related applications

[0062] This application asserts U.S. Provisional Patent Application Nos. 63 / 244,146, filed September 14, 2021, 63 / 293,286, filed December 23, 2021, and 63 / 336,008, filed April 28, 2022 The contents of each of these documents are incorporated herein by reference in their entirety. I. Introduction to the Tropical Delivery System

[0063] Compared with traditional therapies, nucleic acid therapy has emerged as a leading method for the treatment of various diseases and indications due to its versatility, lower immune response and higher efficacy. For example, nucleic acid therapy includes the use of small interfering RNA (siRNA) to reduce translation of messenger RNA (mRNA); mRNA, as a means of producing a target of interest; circular RNA (oRNA), which can provide Continuously produced polypeptides or peptides may be sponges that compete with other RNA molecules; and viral vectors for providing continuously produced targets of interest. However, some nucleic acids are unstable and prone to degradation, so they need to be formulated to prevent degradation and facilitate intracellular delivery of nucleic acids.

[0064] Current delivery vehicles, including lipid-based delivery vehicles, such as lipid nanoparticles and liposomes, focus on protecting the cargo, but not on targeting the delivery of the cargo or the delivery vehicle to a specific area in vivo.

[0065] Provided herein is a tropism discovery platform for evaluating targeting systems for localized delivery to specific target regions, cells or tissues. As shown in Figure 1, the tropism discovery platform can be used to evaluate lipid nanoparticle (LNP) libraries and / or AAV libraries in order to determine the tropism or signature profiles of targeting systems in the libraries. The library can be administered to an individual (e.g., a non-human primate, rabbit, mouse, rat, or another mammal), and organs and tissues of the individual are scanned and / or collected and analyzed to determine LNP or AAV in the library The location of an identifier (such as a barcode, mark, signal and / or label) contained in or associated with it. This analysis provides a signature or profile of the tropism of each LNP and AAV in the library. Initial Construct Architecture

[0066] The targeting system of the tropism discovery platform may include codes or initial constructs including cargo or payload. The initial polynucleotide construct can be linear or circular

[0100] Examples are provided in [picture] [2]. initial polynucleotide construct

[0100] May include at least one payload area that is or encodes the payload or cargo of interest

[10] . initial polynucleotide construct

[0100] Can contain 1 or 2 flanking regions 2 [0], and flanked by area 2 [0] can be located in the payload area

[10] 5' or payload area 3' of

[10] . In some cases, the initial polynucleotide construct

[0100] Does not contain flanking region 2 [0]. initial polynucleotide construct

[0100] The flanking area

[20] may include at least one conditioning zone

[30] . initial polynucleotide construct

[0100] at least one flanking region

[20] may include at least one identifier field

[40] . identifier field

[40] Can be, but is not limited to, barcodes, markers, signals and / or labels. Additionally, the identifier field

[40] can be located in the payload area

[10] or can be positioned within the payload area

[10] and at least one flanking area

[20] .

[0067] In some embodiments, the starting construct comprises about 5 to about 10,000 residues. As a non-limiting example, the length of the initial construct can be 5 to 30, 5 to 50, 5 to 100, 5 to 250, 5 to 500, 5 to 1,000, 5 to 1,500, 5 to 3,000, 5 to 5,000, 5 to 7,000, 5 to 10,000, 30 to 50, 30 to 100, 30 to 250, 30 to 500, 30 to 1,000, 30 to 1,500, 30 to 3,000, 30 to 5,000, 30 to 7,000, 30 to 10,000, 100 to 250 . 00 to 7,000, 500 to 10,000, 1,000 to 1,500, 1,000 to 2,000, 1,000 to 3,000, 1,000 to 5,000, 1,000 to 7,000, 1,000 to 10,000, 1,500 to 3,000, 1,500 to 5,000, 1,500 to 7,000, 1,500 to 10,000, 2,000 to 3,000 .

[0068] In some embodiments, the length of the payload region is greater than about 5 residues in length, such as but not limited to at least or greater than about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 residues in length. , 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 40, 45, 50, 55, 60, 70 . , 1,700, 1,800, 1,900, 2,000, 2,500, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, or more than 10,000 residues.

[0069] In some embodiments, the flanking regions may independently range from 0 to 10,000 residues in length, such as but not limited to at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,500, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000 and 10,000.

[0070] In some embodiments, the length of the regulatory region may independently range from 0 to 3,000 residues, such as but not limited to at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,500 and 3,000.

[0071] In some embodiments, the starting constructs can be circularized or concatenated to create molecules that facilitate interactions between the 3' and 5' ends of the starting constructs. Baseline Architecture

[0072] An initial construct that includes at least one identifier (eg, barcode, marker, signal, and / or label) is called a baseline construct. A reference polynucleotide construct may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more identifiers that may be present throughout the reference polynucleotide construct same or different.

[0073] In some embodiments, the length of the identifier region may independently range from 1 to 3,000 residues, such as but not limited to at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 40 ,45,50,55,60,70,80,90,100,120,140,160,180,200,250,300,350,400,450,500,600,700,800,900,1,000,1,100 , 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,500 and 3,000. As non-limiting examples, the length of the identifier region can be 1-5 residues, 2-5 residues, 3-5 residues, 2-7 residues, 3-7 residues, 1-5 residues, 10 residues, 2-10 residues, 3-10 residues, 5-10 residues, 7-10 residues, 1-15 residues, 2-15 residues, 3-15 residues, 5-15 residues, 7-15 residues, 10-15 residues, 12-15 residues, 1-20 residues, 2-20 residues, 3-20 residues residues, 5-20 residues, 7-20 residues, 10-20 residues, 12-20 residues, 15-20 residues, 17-20 residues, 1-25 residues base, 2-25 residues, 3-25 residues, 5-25 residues, 7-25 residues, 10-25 residues, 12-25 residues, 15-25 residues , 17-25 residues, 20-25 residues, 1-30 residues, 2-30 residues, 3-30 residues, 5-30 residues, 7-30 residues, 10-30 residues, 12-30 residues, 15-30 residues, 17-30 residues, 20-30 residues, 25-30 residues, 1-35 residues, 2 -35 residues, 3-35 residues, 5-35 residues, 7-35 residues, 10-35 residues, 12-35 residues, 15-35 residues, 17- 35 residues, 20-35 residues, 25-35 residues, 30-35 residues, 1-35 residues, 2-35 residues, 3-35 residues, 5-35 residues, 7-35 residues, 10-35 residues, 12-35 residues, 15-35 residues, 17-35 residues, 20-35 residues, 25-35 residues residues, 30-35 residues, 1-40 residues, 2-40 residues, 3-40 residues, 5-40 residues, 7-40 residues, 10-40 residues residues, 12-40 residues, 15-40 residues, 17-40 residues, 20-40 residues, 25-40 residues, 30-40 residues, 35-40 residues , 1-45 residues, 2-45 residues, 3-45 residues, 5-45 residues, 7-45 residues, 10-45 residues, 12-45 residues, 15-45 residues, 17-45 residues, 20-45 residues, 25-45 residues, 30-45 residues, 35-45 residues, 40-45 residues, 1 -50 residues, 2-50 residues, 3-50 residues, 5-50 residues, 7-50 residues, 10-50 residues, 12-50 residues, 15- 50 residues, 17-50 residues, 20-50 residues, 25-50 residues, 30-50 residues, 35-50 residues, 40-50 residues, or 45-50 residues.

[0074] Non-limiting examples of reference polynucleotide constructs, which may be linear or circular, having at least one identifier are provided at [picture] [3A] [,picture] [3B] and [picture] [3C]. Non-limiting examples of circular reference polynucleotide constructs having at least one identifier are provided at [picture] [4A] [,picture] [4B] and [picture] [4C]. exist [picture] [3A] [,picture] [3B] [,picture] [4A] and [picture] In [4B], the reference polynucleotide construct comprises a payload region (referred to as "P" in the figure) and at least one identifier region (referred to as "BC" in the figure) and / or a reverse identifier region (referred to as "CB" in the figure). exist [picture] [3C] and [picture] In [4C], the reference polynucleotide construct comprises a payload region (referred to as "P" in the figure) and at least one identifier moiety bound to the reference polynucleotide construct.

[0075] In some embodiments, the identifier region in the reference construct overlaps with the payload region. As used herein, "overlap" means that at least one nucleotide of the identifier region extends into the payload region. In some aspects, the identifier region overlaps the payload region by 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides Nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides, 21 nucleotides, 22 nucleotides, 23 nucleotides acid, 24 nucleotides, 25 nucleotides, 26 nucleotides, 27 nucleotides, 28 nucleotides, 29 nucleotides, 30 nucleotides, 31 nucleotides, 32 nucleotides, 33 nucleotides, 34 nucleotides, 35 nucleotides, 36 nucleotides, 37 nucleotides, 38 nucleotides, 39 nucleotides, 40 nucleotides Nucleotides, 41 nucleotides, 42 nucleotides, 43 nucleotides, 44 nucleotides, 45 nucleotides, 46 nucleotides, 47 nucleotides, 48 ​​nucleotides acid, 49 nucleotides, 50 nucleotides, or more than 50 nucleotides. In some aspects, the identifier region overlaps the payload region by 1-5 nucleotides, 2-5 nucleotides, 3-5 nucleotides, 2-7 nucleotides, 3-7 nucleotides Nucleotides, 1-10 nucleotides, 2-10 nucleotides, 3-10 nucleotides, 5-10 nucleotides, 7-10 nucleotides, 1-15 nucleotides acid, 2-15 nucleotides, 3-15 nucleotides, 5-15 nucleotides, 7-15 nucleotides, 10-15 nucleotides, 12-15 nucleotides, 1-20 nucleotides, 2-20 nucleotides, 3-20 nucleotides, 5-20 nucleotides, 7-20 nucleotides, 10-20 nucleotides, 12- 20 nucleotides, 15-20 nucleotides, 17-20 nucleotides, 1-25 nucleotides, 2-25 nucleotides, 3-25 nucleotides, 5-25 nucleotides Nucleotides, 7-25 nucleotides, 10-25 nucleotides, 12-25 nucleotides, 15-25 nucleotides, 17-25 nucleotides, 20-25 nucleotides acid, 1-30 nucleotides, 2-30 nucleotides, 3-30 nucleotides, 5-30 nucleotides, 7-30 nucleotides, 10-30 nucleotides, 12-30 nucleotides, 15-30 nucleotides, 17-30 nucleotides, 20-30 nucleotides, 25-30 nucleotides, 1-35 nucleotides, 2- 35 nucleotides, 3-35 nucleotides, 5-35 nucleotides, 7-35 nucleotides, 10-35 nucleotides, 12-35 nucleotides, 15-35 nucleotides Nucleotides, 17-35 nucleotides, 20-35 nucleotides, 25-35 nucleotides, 30-35 nucleotides, 1-35 nucleotides, 2-35 nucleotides acid, 3-35 nucleotides, 5-35 nucleotides, 7-35 nucleotides, 10-35 nucleotides, 12-35 nucleotides, 15-35 nucleotides, 17-35 nucleotides, 20-35 nucleotides, 25-35 nucleotides, 30-35 nucleotides, 1-40 nucleotides, 2-40 nucleotides, 3- 40 nucleotides, 5-40 nucleotides, 7-40 nucleotides, 10-40 nucleotides, 12-40 nucleotides, 15-40 nucleotides, 17-40 nucleotides Nucleotides, 20-40 nucleotides, 25-40 nucleotides, 30-40 nucleotides, 35-40 nucleotides, 1-45 nucleotides, 2-45 nucleotides acid, 3-45 nucleotides, 5-45 nucleotides, 7-45 nucleotides, 10-45 nucleotides, 12-45 nucleotides, 15-45 nucleotides, 17-45 nucleotides, 20-45 nucleotides, 25-45 nucleotides, 30-45 nucleotides, 35-45 nucleotides, 40-45 nucleotides, 1- 50 nucleotides, 2-50 nucleotides, 3-50 nucleotides, 5-50 nucleotides, 7-50 nucleotides, 10-50 nucleotides, 12-50 nucleotides Nucleotides, 15-50 nucleotides, 17-50 nucleotides, 20-50 nucleotides, 25-50 nucleotides, 30-50 nucleotides, 35-50 nucleotides acid, 40-50 nucleotides, or 45-50 nucleotides.

[0076] In some embodiments, a reference polynucleotide construct comprises a payload region and an identifier region. The identifier region can be located 5' of the payload region, 3' of the payload region, or the identifier region can overlap the 5' end or the 3' end of the payload region.

[0077] In some embodiments, a reference polynucleotide construct comprises a payload region and two identifier regions. Each identifier region can be independently located 5' of the payload region, 3' of the payload region, or the identifier region can overlap the 5' end or the 3' end of the payload region.

[0078] As a non-limiting example, the first identifier area is located 5' of the payload area and the second identifier area is located 3' of the payload area. As a non-limiting example, the first and second identifier fields are located 5' of the payload field. As a non-limiting example, the first and second identifier fields are located 3' of the payload field.

[0079] As a non-limiting example, the first identifier field is inverted and located 5' from the payload area, and the second identifier field is located 3' from the payload area. As a non-limiting example, the first identifier field is inverted and located 5' from the payload area, and the second identifier area is inverted and located 3' from the payload area. As a non-limiting example, the first identifier field is located 5' of the payload area and the second identifier field is reversed and located 3' of the payload area. As a non-limiting example, both the first and second identifier fields are reversed and located 5' of the payload field. As a non-limiting example, the first and second identifier fields are located 5' of the payload field, and the first identifier field is reversed. As a non-limiting example, the first and second identifier regions are located 5' of the payload region, and the second identifier region is reversed. As a non-limiting example, both the first and second identifier fields are reversed and located 3' of the payload field. As a non-limiting example, the first and second identifier fields are located 3' of the payload field, and the first identifier field is reversed. As a non-limiting example, the first and second identifier fields are located 3' of the payload field, and the second identifier field is reversed.

[0080] As a non-limiting example, the first identifier region is located 5' of the payload region and overlaps the payload region, and the second identifier region is located 3' of the payload region. As a non-limiting example, the first identifier region is located 5' of the payload region and the second identifier region is located 3' of the payload region and overlaps the payload region.

[0081] As a non-limiting example, the first and second identifier regions are located 5' of the payload region, and the second identifier region overlaps the payload region. As a non-limiting example, the first and second identifier regions are located 3' of the payload region, and the first identifier region overlaps the payload region.

[0082] As a non-limiting example, the first identifier region is inverted, located 5' from and overlapping the payload region, and the second identifier region is located 3' from the payload region. As a non-limiting example, the first identifier region is inverted and located 5' of the payload region, and the second identifier region is located 3' of the payload region and overlaps the payload region. As a non-limiting example, the first identifier region is reversed, located 5' from the payload region, the second identifier region is located 3' from the payload region, and both the first and second identifier regions are aligned with Payload areas overlap.

[0083] As a non-limiting example, the first identifier region is reversed, located 5' from and overlapping the payload region, and the second identifier region is reversed and located 3' from the payload region. As a non-limiting example, the first identifier region is inverted and located 5' from the payload region, and the second identifier region is inverted, located 3' from the payload region and overlaps the payload region. As a non-limiting example, the first identifier field is inverted and located 5' from the payload area, and the second identifier field is inverted and located 3' from the payload area, and the first and second identification The breaker area both overlaps the payload area.

[0084] As a non-limiting example, the first identifier region is located 5' of the payload region and overlaps the payload region, and the second identifier region is reversed and located 3' of the payload region. As a non-limiting example, the first identifier region is located 5' of the payload region and the second identifier region is reversed, located 3' of the payload region and overlaps the payload region. As a non-limiting example, the first identifier region is located 5' from the payload region, and the second identifier region is reversed and located 3' from the payload region, and both the first and second identifier regions are Overlaps the payload area.

[0085] As a non-limiting example, the first and second identifier regions are both inverted and located 5' of the payload region, and the second identifier region overlaps the payload region. As a non-limiting example, the first and second identifier regions are located 5' of the payload region, and the first identifier region is reversed, and the second identifier region overlaps the payload region. As a non-limiting example, the first and second identifier regions are located 5' of the payload region, and the second identifier region is inverted and overlaps the payload region. As a non-limiting example, the first and second identifier regions are both inverted and located 3' of the payload region, and the first identifier region overlaps the payload region. As a non-limiting example, the first and second identifier regions are located in the payload region 3', and the first identifier region is inverted and overlaps the payload region. As a non-limiting example, the first and second identifier regions are located 3' of the payload region, and the second identifier region is inverted and overlaps the payload region.

[0086] In some embodiments, at least one identifier portion can be associated with a reference polynucleotide construct. A reference polynucleotide construct may have 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more identifier moieties associated with the reference polynucleotide construct, the identifiers The portion may be the same portion or a different portion that binds to the reference polynucleotide construct. Each identifier portion can be independently positioned on the 5' side of the payload area, on the 3' side of the payload area, or the position of the identifier portion can span the 5' end or the 3' side of the payload area. 'end and flanking areas. In some aspects, the location of the identifier portion may include one or more nucleotides of the payload region, such as but not limited to 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 cores Nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides , 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides, 21 Nucleotides, 22 Nucleotides, 23 Nucleotides, 24 Nucleotides, 25 Nucleotides, 26 Nucleotides, 27 Nucleotides, 28 Nucleotides, 29 Nucleotides Nucleotides, 30 nucleotides, 31 nucleotides, 32 nucleotides, 33 nucleotides, 34 nucleotides, 35 nucleotides, 36 nucleotides, 37 nucleotides , 38 nucleotides, 39 nucleotides, 40 nucleotides, 41 nucleotides, 42 nucleotides, 43 nucleotides, 44 nucleotides, 45 nucleotides, 46 nucleotides nucleotides, 47 nucleotides, 48 ​​nucleotides, 49 nucleotides, 50 nucleotides or more than 50 nucleotides. In some aspects, the location of the identifier portion may include one or more nucleotides of the payload region, such as but not limited to 1-5 nucleotides, 2-5 nucleotides, 3-5 cores Nucleotide, 2-7 nucleotides, 3-7 nucleotides, 1-10 nucleotides, 2-10 nucleotides, 3-10 nucleotides, 5-10 nucleotides , 7-10 nucleotides, 1-15 nucleotides, 2-15 nucleotides, 3-15 nucleotides, 5-15 nucleotides, 7-15 nucleotides, 10 - 15 nucleotides, 12-15 nucleotides, 1-20 nucleotides, 2-20 nucleotides, 3-20 nucleotides, 5-20 nucleotides, 7-20 nucleotides, 10-20 nucleotides, 12-20 nucleotides, 15-20 nucleotides, 17-20 nucleotides, 1-25 nucleotides, 2-25 cores Nucleotides, 3-25 nucleotides, 5-25 nucleotides, 7-25 nucleotides, 10-25 nucleotides, 12-25 nucleotides, 15-25 nucleotides , 17-25 nucleotides, 20-25 nucleotides, 1-30 nucleotides, 2-30 nucleotides, 3-30 nucleotides, 5-30 nucleotides, 7 -30 nucleotides, 10-30 nucleotides, 12-30 nucleotides, 15-30 nucleotides, 17-30 nucleotides, 20-30 nucleotides, 25-30 nucleotides, 1-35 nucleotides, 2-35 nucleotides, 3-35 nucleotides, 5-35 nucleotides, 7-35 nucleotides, 10-35 nucleotides Nucleotides, 12-35 nucleotides, 15-35 nucleotides, 17-35 nucleotides, 20-35 nucleotides, 25-35 nucleotides, 30-35 nucleotides , 1-35 nucleotides, 2-35 nucleotides, 3-35 nucleotides, 5-35 nucleotides, 7-35 nucleotides, 10-35 nucleotides, 12 -35 nucleotides, 15-35 nucleotides, 17-35 nucleotides, 20-35 nucleotides, 25-35 nucleotides, 30-35 nucleotides, 1-40 nucleotides, 2-40 nucleotides, 3-40 nucleotides, 5-40 nucleotides, 7-40 nucleotides, 10-40 nucleotides, 12-40 cores Nucleotides, 15-40 nucleotides, 17-40 nucleotides, 20-40 nucleotides, 25-40 nucleotides, 30-40 nucleotides, 35-40 nucleotides , 1-45 nucleotides, 2-45 nucleotides, 3-45 nucleotides, 5-45 nucleotides, 7-45 nucleotides, 10-45 nucleotides, 12 - 45 nucleotides, 15-45 nucleotides, 17-45 nucleotides, 20-45 nucleotides, 25-45 nucleotides, 30-45 nucleotides, 35-45 nucleotides, 40-45 nucleotides, 1-50 nucleotides, 2-50 nucleotides, 3-50 nucleotides, 5-50 nucleotides, 7-50 cores Nucleotides, 10-50 nucleotides, 12-50 nucleotides, 15-50 nucleotides, 17-50 nucleotides, 20-50 nucleotides, 25-50 nucleotides , 30-50 nucleotides, 35-50 nucleotides, 40-50 nucleotides or 45-50 nucleotides.

[0087] In some embodiments, an identifier moiety can be associated with a reference polynucleotide construct. As a non-limiting example, an identifier moiety can be associated with a reference polynucleotide construct at the 5' end of the reference polynucleotide construct. As a non-limiting example, the identifier moiety can be bound to the reference polynucleotide construct at the 5' flanking region. As a non-limiting example, the identifier moiety can be bound to the reference polynucleotide construct at the 3' flanking region. As a non-limiting example, an identifier moiety can be associated with a reference polynucleotide construct at the 3' end of the reference polynucleotide construct. As a non-limiting example, an identifier moiety can be combined with a reference polynucleotide construct on the payload region. As a non-limiting example, the reference polynucleotide construct comprises an identifier portion, and the position of the identifier portion spans the 5' end of the payload region and the 5' flanking region. As a non-limiting example, the reference polynucleotide construct comprises an identifier portion, and the position of the identifier portion spans the 3' end of the payload region and the 3' flanking region.

[0088] In some embodiments, two identifier moieties are associated with a reference polynucleotide construct. As a non-limiting example, the first identifier portion and the second identifier portion are located on the 5' flanking region. As a non-limiting example, the first identifier portion and the second identifier portion are located on the payload area. As a non-limiting example, the first identifier portion and the second identifier portion are located on the 3′ flanking region. As a non-limiting example, the first identifier portion and the second identifier portion are located on the 5' end of the reference polynucleotide construct. As a non-limiting example, the first identifier portion and the second identifier portion are located on the 3' end of the reference polynucleotide construct.

[0089] As a non-limiting example, the first identifier portion is located on the 5' end of the reference polynucleotide construct, and the second identifier portion is located on the 5' flanking region. As a non-limiting example, the first identifier portion is located on the 5' end of the reference polynucleotide construct and the second identifier portion is located on the payload region. As a non-limiting example, the first identifier portion is located on the 5' end of the reference polynucleotide construct and the second identifier portion is located on the 3' flanking region. As a non-limiting example, the first identifier portion is located on the 5' end of the reference polynucleotide construct, and the second identifier portion is located across the 5' flanking region and the payload region. As a non-limiting example, the first identifier portion is located on the 5' end of the reference polynucleotide construct, and the second identifier portion is located across the 3' flanking region and the payload region. As a non-limiting example, the first identifier portion is located on the 5' end of the reference polynucleotide construct and the second identifier portion is located on the 3' end of the reference polynucleotide construct.

[0090] As a non-limiting example, the first identifier portion is located on the 5' flanking area and the second identifier portion is located on the payload area. As a non-limiting example, the first identifier portion is located on the 5' flanking region and the second identifier portion is located on the 3' flanking region. As a non-limiting example, the first identifier portion is located on the 5' flanking region, and the location of the second identifier portion spans the 5' flanking region and the payload region. As a non-limiting example, the first identifier portion is located on the 5' flanking region, and the location of the second identifier portion spans the 3' flanking region and the payload region. As a non-limiting example, the first identifier portion is located on the 5' flanking region and the second identifier portion is located on the 5' end of the reference polynucleotide construct. As a non-limiting example, the first identifier portion is located on the 5' flanking region and the second identifier portion is located on the 3' end of the reference polynucleotide construct.

[0091] As a non-limiting example, the location of the first identifier portion spans the 5' flanking region and the payload region, and the second identifier portion is located on the 5' end of the reference polynucleotide construct. As a non-limiting example, the location of the first identifier portion spans the 5' flanking region and the payload region, and the second identifier portion is located on the 5' flanking region. As a non-limiting example, the location of the first identifier portion spans the 5' flanking region and the payload region, and the second identifier portion is located on the payload region. As a non-limiting example, the location of the first identifier portion spans the 5′ flanking region and the payload region, and the location of the second identifier portion spans the 3′ flanking region and the payload region. As a non-limiting example, the location of the first identifier portion spans the 5' flanking region and the payload region, and the second identifier portion is located on the 3' flanking region. As a non-limiting example, the location of the first identifier portion spans the 5' flanking region and the payload region, and the second identifier portion is located on the 3' end of the reference polynucleotide construct.

[0092] As a non-limiting example, the first identifier portion is located on the payload region and the second identifier portion is located on the 5' end of the reference polynucleotide construct. As a non-limiting example, the first identifier portion is located on the payload area and the second identifier portion is located on the 5' flanking area. As a non-limiting example, the first identifier portion is located on the payload area, and the location of the second identifier portion spans the 5' flanking area and the payload area. As a non-limiting example, the first identifier portion is located on the payload area, and the location of the second identifier portion spans the 3' flanking area and the payload area. As a non-limiting example, the first identifier portion is located on the payload area and the second identifier portion is located on the 3' flanking area. As a non-limiting example, the first identifier portion is located on the payload region and the second identifier portion is located on the 3' end of the reference polynucleotide construct.

[0093] As a non-limiting example, the location of the first identifier portion spans the 3' flanking region and the payload region, and the second identifier portion is located on the 5' end of the reference polynucleotide construct. As a non-limiting example, the location of the first identifier portion spans the 3' flanking region and the payload region, and the second identifier portion is located on the 5' flanking region. As a non-limiting example, the location of the first identifier portion spans the 3′ flanking region and the payload region, and the location of the second identifier portion spans the 5′ flanking region and the payload region. As a non-limiting example, the location of the first identifier portion spans the 3' flanking region and the payload region, and the second identifier portion is located on the payload region. As a non-limiting example, the location of the first identifier portion spans the 3′ flanking region and the payload region, and the second identifier portion is located on the 3′ flanking region. As a non-limiting example, the location of the first identifier portion spans the 3' flanking region and the payload region, and the second identifier portion is located on the 3' end of the reference polynucleotide construct.

[0094] As a non-limiting example, the location of the first identifier portion spans the 3' flanking region and the payload region, and the second identifier portion is located on the 5' flanking region. As a non-limiting example, the location of the first identifier portion spans the 5' flanking region and the payload region, and the second identifier portion is located on the payload region. As a non-limiting example, the location of the first identifier portion spans the 5′ flanking region and the payload region, and the location of the second identifier portion spans the 3′ flanking region and the payload region. As a non-limiting example, the location of the first identifier portion spans the 5' flanking region and the payload region, and the second identifier portion is located on the 3' flanking region. As a non-limiting example, the location of the first identifier portion spans the 5' flanking region and the payload region, and the second identifier portion is located on the 3' end of the reference polynucleotide construct.

[0095] As a non-limiting example, the first identifier portion is located on the 3' flanking region and the second identifier portion is located on the 5' end of the reference polynucleotide construct. As a non-limiting example, the first identifier portion is located on the 3' flanking region and the second identifier portion is located on the 5' flanking region. As a non-limiting example, the first identifier portion is located on the 3' flanking region, and the location of the second identifier portion spans the 5' flanking region and the payload region. As a non-limiting example, the first identifier portion is located on the 3' flanking area and the second identifier portion is located on the payload area. As a non-limiting example, the first identifier portion is located on the 3' flanking region, and the location of the second identifier portion spans the 3' flanking region and the payload region. As a non-limiting example, the first identifier portion is located on the 3' flanking region and the second identifier portion is located on the 3' end of the reference polynucleotide construct.

[0096] As a non-limiting example, the first identifier portion is located on the 3' end of the reference polynucleotide construct and the second identifier portion is located on the 5' end of the reference polynucleotide construct. As a non-limiting example, the first identifier portion is located on the 3' end of the reference polynucleotide construct and the second identifier portion is located on the 5' flanking region. As a non-limiting example, the first identifier portion is located on the 5' end of the reference polynucleotide construct, and the second identifier portion is located across the 5' flanking region and the payload region. As a non-limiting example, the first identifier portion is located on the 3' end of the reference polynucleotide construct and the second identifier portion is located on the payload region. As a non-limiting example, the first identifier portion is located on the 5' end of the reference polynucleotide construct, and the second identifier portion is located across the 3' flanking region and the payload region. As a non-limiting example, the first identifier portion is located on the 3' end of the reference polynucleotide construct and the second identifier portion is located on the 3' flanking region.

[0097] In some embodiments, three identifier moieties are associated with a reference polynucleotide construct.

[0098] In some embodiments, four identifier moieties are associated with a reference polynucleotide construct.

[0099] In some embodiments, five identifier moieties are associated with a reference polynucleotide construct.

[0100] In some embodiments, six identifier segments are associated with a reference polynucleotide construct.

[0101] In some embodiments, seven identifier moieties are associated with a reference polynucleotide construct.

[0102] In some embodiments, eight identifier moieties are associated with a reference polynucleotide construct.

[0103] In some embodiments, nine identifier segments are combined with a reference polynucleotide construct.

[0104] In some embodiments, ten identifier segments are associated with a reference polynucleotide construct. II. Cargo and Payload

[0105] The initial constructs and reference constructs of the present invention may contain, encode, or be combined with a cargo or payload. As used herein, the term "cargo" or "payload" may refer to one or more molecules or structures contained in a delivery vehicle for delivery of a cell or tissue into a cell or tissue. Non-limiting examples of cargo can include nucleic acids, polypeptides, peptides, proteins, liposomes, markers, tags, small chemical molecules, large biomolecules, and any combination or fragment thereof. In the original structure and the reference structure, the area of ​​the structure that contains or encodes cargo or payload is called the "cargo area" or "payload area".

[0106] In some embodiments, the cargo or payload is or encodes a biologically active molecule, such as, but not limited to, a therapeutic protein. As used herein, the term "biological activity" refers to any property of an agent that is active in a biological system, and in particular an organism. For example, an agent that, when administered to an organism, has a biological effect on that organism is considered to be biologically active. In some embodiments, the cargo or payload is or encodes one or more prophylactically or therapeutically active proteins, polypeptides or other factors. As a non-limiting example, the cargo or payload can be or encode an agent that enhances tumor killing activity in cancer, such as but not limited to TRAIL or tumor necrosis factor (TNF). As another non-limiting example, the cargo or payload may encode an agent suitable for the treatment of conditions such as: muscular dystrophy (e.g. cargo or payload is or encodes myosin), cardiovascular disease (e.g. payload is or encodes SERCA2a, GATA4, Tbx5, Mef2C, Hand2, Myocd, etc.), neurodegenerative disease (eg, cargo or payload is or encodes NGF, BDNF, GDNF, NT-3, etc.), chronic pain (eg, cargo or payload cargo is or encodes GlyRal), enkephalin or glutamate decarboxylase (eg cargo or payload is or encodes GAD65, GAD67 or another isoform), lung disease (eg cargo or payload is or encodes CFTR), blood disease (e.g. cargo or payload is or encodes Factor VIII or Factor IX), neoplasia (e.g. cargo or payload is or encodes PTEN, ATM, ATR, EGFR, ERBB2, ERBB3, ERBB4, Notchl, Notch2, Notch3, Notch4, AKT, AKT2, AKT3, HIF, HI Fla, HIF3a, Met, HRG, Bcl2, PPARα, PPARγ, WT1 (Wilms Tumor), FGF receptor family members (5 members: 1, 2, 3, 4, 5), CDKN2a, APC, RB (retinoblastoma), MEN1, VHL, BRCA1, BRCA2, androgen receptor (Androgen Receptor; AR), TSG101, IGF, IGF receptor, Igfl ( 4 variants), Igf2 (3 variants), Igfl receptor, Igf2 receptor, Bax, Bcl2, caspase family (9 members: 1, 2, 3, 4, 6, 7, 8, 9 , 12), Kras, Ape), age-related macular degeneration (eg cargo or payload is or encodes Aber, Ccl2, Cc2, ceruloplasmin (cp), Timp3, cathepsin D, Vldlr), schizophrenia ( For example, neuregulin (Nrgl), Erb4 (neuregulin receptor), complexin (Complexin)-1 (Cplxl), Tphl tryptophan hydroxylase, Tph2 tryptophan hydroxylase 2, Neurexin 1, GSK3 , GSK3a, GSK3b, 5-HIT (Slc6a4), COMT, DRD (Drdla), SLC6A3, DAOA, DTNBPI, Dao (Daol)), trinucleotide repeat disorders (eg HTT (Huntington's (Huntington's) Dx), SBMA / SMAXI / AR (Kennedy's Dx), FXN / X25 (Friedrich's Ataxia), ATX3 (Machado-Joseph's Dx), ATXNI and ATXN2 (Spinocerebellar Ataxia), DMPK (Dystonic Dystrophy), Dystrophin-1 and Atnl (DRPLA Dx), CBP (Creb-BP-Global Instability), VLDLR (Alzheimer's )), Atxn7, Atxn10), X Fragile Syndrome (e.g. cargo or payload is or encodes FMR2, FXRI, FXR2, mGLUR5), secretase-associated disorders (e.g. cargo or payload is or encodes APH-1 (α and β ), presenilin (Psenl), nicastrin (nicastrin; Ncstn), PEN-2), ALS (e.g. cargo or payload is or encodes SOD1, ALS2, STEX, FUS, TARD BP, VEGF (VEGF-a , VEGF-b, VEGF-c)), autism (e.g. cargo or payload is or encodes Mecp2, BZRAP1, MDGA2, Sema5A, axon 1), Alzheimer's disease (e.g. cargo or payload is Or encode El, CHIP, UCH, UBB, Tau, LRP, PICALM, Clusterin (Clusterin), PS1, SORL1, CR1, Vldlr, Ubal, Uba3, CHIP28 (Aqpl, aquaporin 1), Uchll, Uchl3, APP) , inflammatory (e.g. cargo or payload is or encodes IL-10, IL-1 (IL-Ia, IL-Ib), IL-13, IL-17 (IL-17a (CTLA8), IL-17b, IL-17c , IL-17d, IL-171), 11-23, Cx3crl, ptpn22, TNFa, NOD2 / CARD15 for IBD, IL-6, IL-12 (IL-12a, IL-12b), CTLA4, Cx3cll), Parkinson's Disease (such as x-synuclein, DJ-1, LRRK2, Parkin, PINK1), blood and coagulation disorders (such as anemia, naked lymphocyte syndrome, bleeding disorders , hemophagocytic lymphohistiocytosis, hemophilia A, hemophilia B, bleeding disorders, leukocyte deficiency and disorders, sickle cell anemia and thalassemia) (e.g. cargo or payload is or encodes CRAN1, CDA1, RPS19, DBA, PKLR, PK1, NT5C3, UMPH1, PSNI, RHAG, RH50A, NRAMP2, SPTB, ALAS2, ANH1, ASB, ABCB7, ABC7, ASAT, TAPBP, TPSN, TAP2, ABCB3, PSF2, RING11, MHC2TA, C2TA, RFX5, RFXAP, RFX5, TBXA2R, P2RX1, P2X1, HF1, CFH, HUS, MCFD2, FANCA, FAC A, FA1, FA, FA A, FAAP95, FAAP90, FLJ34064, FANCB, FANCC, FACC, BRCA2, FANCDI, FANCD2, FANCD, FACD, FAD, FANCE, FACE, FANCF, XRCC9, FANCG, BR1PI, BACH1, FANCJ, PHF9, FANCL, FANCM, KIAA1596, PRF1, HPLH2, UNC13D, MUNC13-4, HPLH3, HLH3, FHL3, F8, FSC, PI, ATT, F5, ITGB2, CD18, LCAMB, LAD, EIF2B1, EIF2BA, EIF2B2, EIF2B3, EIF2B5, LVWM, CACH, CLE, EIF2B4, HBB, HBA2, HBB, HBD, LCRB, HBA1), B cell non Hodgkin lymphoma (B-cell non-Hodgkin lymphoma) or leukemia (e.g. cargo or payload is or encodes BCL7A, BCL7, ALI, TCL5, SCL, TAL2, FLT3, NBS1, NBS, ZNFN1AI, 1KI, LYF1, HOXD4 , HOX4B, BCR, CML, PHL, ALL, ARNT, KRAS2, RASK2, GMPS, AFIO, ARHGEF12, LARG, KIAA0382, CALM, CLTH, CEBPA, CEBP, CHIC2, BTL, FLT3, KIT, PBT, LPP, NPMI, NUP214 , D9S46E, CAN, CAIN, RUNXI, CBFA2, AML1, WHSC1LI, NSD3, FLT3, AF1Q, NPMI, NUMA1, ZNF145, PLZF, PML, MYL, STAT5B, AF1Q, CALM, CLTH, ARL11, ARLTS1, P2RX7, P2X7, BCR , CML, PHL, ALL, GRAF, NF1, VRNF, WSS, NFNS, PTPNII, PTP2C, SHP2, NS1, BCL2, CCND1, PRAD1, BCL1, TCRA, GATA1, GF1, ERYF1, NFE1, ABLI, NQO1, DIA4, NMOR1 , NUP214, D9S46E, CAN, CAIN), inflammatory and immune-related diseases and disorders (e.g. cargo or payload is or encodes KIR3DL1, NKAT3, NKB1, AMB11, K1R3DS1, IFNG, CXCL12, TNFRSF6, APT1, FAS, CD95, ALPS1A, IL2RG, SCIDX1, SCIDX, IMD4, CCL5, SCYA5, D17S136E, TCP228, IL10, CSIF, CMKBR2, CCR2, CMKBR5, CCCKR5 (CCR5), CD3E, CD3G, AICDA, AID, HIGM2, TNFRSF5, CD40, UNG, DGU, HIGM4 , TNFSFS, CD40LG, HIGM1, IGM, FOXP3, IPEX, AIID, XPID, PIDX, TNFRSF14B, TACI), inflammation (e.g. cargo or payload is or encodes IL-10, IL-1 (IL-IA, IL-IB ) , IL-13, IL-17 (IL-17a (CTLA8), IL-17b, IL-17c, IL-17d, IL-171), 11-23, Cx3crl, ptpn22, TNFa, NOD2 / CARD15 for IBD , IL-6, IL-12(IL-12a, IL-12b), CTLA4, Cx3cII), JAK3, JAKL, DCLREIC, ARTEMIS, SCIDA, RAG1, RAG2, ADA, PTPRC, CD45, LCA, IL7R, CD3D, T3D, IL2RG, SCIDXI, SCIDX, IMD4 ), metabolic, hepatic, renal and protein diseases and disorders (e.g. cargo or payload is or encodes TTR, PALB, APOA1, APP, AAA, CVAP, ADI, GSN, FGA, LYZ, TTR, PALB, KRT18, KRT8, CIRH1A , NAIC, TEX292, KIAA1988, CFTR, ABCC7, CF, MRP7, SLC2A2, GLUT2, G6PC, G6PT, G6PT1, GAA, LAMP2, LAMPB, AGL, GDE, GBE1, GYS2, PYGL, PFKM, TCF1, HNF1A, MODY3, SCOD1 , SCO1, CTNNB1, PDGFRL, PDGRL, PRLTS, AX1NI, AXIN, CTNNB1, TP53, P53, LFS1, IGF2R, MPRI, MET, CASP8, MCH5, UMOD, HNFJ, FJHN, MCKD2, ADMCKD2, PAH, PKU1, QDPR, DHPR , PTS, FCYT, PKHD1, ARPKD, PKD1, PKD2, PKD4, PKDTS, PRKCSH, G19P1, PCLD, SEC63), musculoskeletal diseases and conditions (e.g. cargo or payload is or encodes DMD, BMD, MYF6, LMNA, LMN1 , EMD2, FPLD, CMDIA, HGPS, LGMDIB, LMNA, LMNI, EMD2, FPLD, CMDIA, FSHMD1A, FSHD1A, FKRP, MDC1C, LGMD2I, LAMA2, LAMM, LARGE, KIAA0609, MDC1D, FCMD, TTID, MYOT, CAPN3, CANP3 , DYSF, LGMD2B, SGCG, LGMD2C, DMDA1, SCG3, SGCA, ADL, DAG2, LGMD2D, DMDA2, SGCB, LGMD2E, SGCD, SGD, LGMD2F, CMD1L, TCAP, LGMD2G, CMD1N, TRIM32, HT2A, LGMD2H, FKRP, MDCIC , LGMD21, TTN, CMD1G, TMD, LGMD2J, POMT1, CAV3, LGMD1C, SEPN1, SELN, RSMD1, PLEC1, PLTN, EBS1, LRP5, BMNDl, LRP7, LR3, OPPG, VBCH2, CLCN7, CLC7, OPTA2, OSTMI, GL , TCIRG1, TIRC7, OC116, OPTB1, VAPB, VAPC, ALS8, SMN1, SMA1, SMA2, SMA3, SMA4, BSCL2, SPG17, GARS, SMAD1, CMT2D, HEXB, IGHMBP2, SMUBP2, CATF1, SMARD1), nerves and neurons Diseases and disorders (e.g. cargo or payload is or encodes SOD1, ALS2, STEX, FUS, TARDBP, VEGF (VEGF-a, VEGF-b, VEGF-c), APP, AAA, CVAP, ADI, APOE, AD2, PSEN2 , AD4, STM2, APBB2, FE65LI, NOS3, PLAU, URK, ACE, DCPI, ACEI, MPO, PAC1PI, PAXIPIL, PTIP, A2M, BLMH, BMH, PSEN1, AD3, Mecp2, BZRAP1, MDGA2, Sema5A, Axon 1. GLO1, MECP2, RTT, PPMX, MRX16, MRX79, NLGN3, NLGN4, KIAA1260, AUTSX2, FMR2, FXR1, FXR2, mGLUR5, HD, IT15, PRNP, PRIP, JPH3, JP3, HDL2, TBP, SCA17, NR4A2, NURR1, NOT, TINUR, SNCAIP, TBP, SCA17, SNCA, NACP, PARK1, PARK4, DJI, PARK7, LRRK2, PARK8, PINK1, PARK6, UCHL1, PARK5, SNCA, NACP, PARK1, PARK4, PRKN, PARK2, PDJ, DBH, NDUFV2, MECP2, RTT, PPMX, MRX16, MRX79, CDKL5, STK9, MECP2, RTT, PPMX, MRX16, MRX79, x-synuclein, DJ-1, neuregulin-l (Nrgl), Erb4 (Neuroregulin Regulin receptor), Complexin-1 (Cplxl), Tphl Tryptophan Hydroxylase, Tph2, Tryptophan Hydroxylase 2, Axonin 1, GSK3, GSK3a, GSK3b, 5-HTT (Slc6a4) , CONT, DRD (Drdla), SLC6A, DAOA, DTNBP1, Dao (Daol), APH-l (α and β), Presenilin (Psenl), Nicastrin, (Ncstn), PEN-2, Nosl, Parpl, Natl, Nat2, HTT, SBMA / SMAX1 / AR, FXN / X25, ATX3, TXN, ATXN2, DMPK, Atrophin-1, Atnl, CBP, VLDLR, Atxn7, and AtxnlO) and eye diseases and disorders (such as Aber , Ccl2, Cc2, Ceruloplasmin (cp), Timp3, Cathepsin-D, Vldlr, Ccr2, CRYAA, CRYA1, CRYBB2, CRYB2, PITX3, BFSP2, CP49, CP47, CRYAA, CRYAI, PAX6, AN2, MGDA, CRYBA1 , CRYB1, CRYGC, CRYG3, CCL, LIM2, MP19, CRYGD, CRYG4, BFSP2, CP49, CP47, HSF4, CTM, HSF4, CTM, MIP, AQPO, CRYAB, CRYA2, CTPP2, CRYBB1, CRYGD, CRYG4, CRYBB2, CRYB2 , CRYGC, CRYG3, CCL, CRYAA, CRYAI, GJA8, CX50, CAE1, GJA3, CX46, CZP3, CAE3, CCM1, CAM, KRIT1, APOA1, TGFBI, CSD2, CDGG1, CSD, BIGH3, CDG2, TACSTD2, TROP2, M1SI , VSX1, RINX, PPCD, PPD, KTCN, COL8A2, FECD, PPCD2, PIP5K3, CFD, KERA, CNA2, MYOC, TIGR, GLCIA, JO AG, GPOA, OPTN, GLC1E, FIP2, HYPL, NRP, CYP1BI, GLC3A, OPA1, NTG, NPG, CYP1BI, GLC3A, CRB1, RP12, CRX, CORD2, CRD, RPGRIPI, LCA6, CORD9, RPE65, RP20, AIPL1, LCA4, GUCY2D, GUC2D, LCA1, CORD6, RDH12, LCA3, ELOVL4, ADMD, STGD2, STGD3, RDS, RP7, PRPH2, PRPH, AVMD, AOFMD, and VMD2).

[0107] In some embodiments, the cargo or payload is or encodes a factor that affects cell differentiation. As a non-limiting example, expression of one or more of Oct4, Klf4, Sox2, c-Myc, L-Myc, dominant negative p53, Nanog, Glisl, Lin28, TFIID, mir-302 / 367, or other miRNAs can be Make the cells into induced pluripotent stem (iPS) cells.

[0108] In some embodiments, the cargo or payload is or encodes a factor for transdifferentiating cells. Non-limiting examples of factors include: for cardiomyocytes, one or more of GATA4, Tbx5, Mef2C, Myocd, Hand2, SRF, Mespl, SMARCD3; for neurons, Ascii, Nurrl, Lmx1A, Bm2, Mytll, NeuroDl, FoxA2; and for hepatocytes, Hnf4a, Foxal, Foxa2 or Foxa3. Polypeptides, Proteins and Peptides

[0109] The starting and reference constructs of the invention may comprise, encode, or be associated with cargo or payload of polypeptides, proteins or peptides. As used herein, the term "polypeptide" generally refers to a polymer of amino acids linked by peptide bonds and encompasses both "proteins" and "peptides." The polypeptides used in the present invention include all polypeptides, proteins and / or peptides known in the art. Non-limiting classes of polypeptides include antigens, antibodies, antibody fragments, interkines, peptides, hormones, enzymes, oxidants, antioxidants, synthetic polypeptides, and chimeric polypeptides.

[0110] As used herein, the term "peptide" generally refers to shorter polypeptides of about 50 amino acids or less. Peptides with only two amino acids can be referred to as "dipeptides". Peptides with only three amino acids can be referred to as "tripeptides". Polypeptide generally refers to a polypeptide having about 4 to about 50 amino acids. Peptides can be obtained by any method known to those skilled in the art. In some embodiments, the peptides can be expressed in culture. In some embodiments, peptides can be obtained via chemical synthesis (eg, solid phase peptide synthesis).

[0111] In some embodiments, the starting and reference constructs of the invention may comprise, encode, or be associated with a cargo or payload that is a simple protein that upon hydrolysis yields amino acids and occasionally small carbohydrates. Non-limiting examples of simple proteins include albumins, acinoids, globulins, gluten, histones, and protamines.

[0112] In some embodiments, the starting and reference constructs of the invention may comprise, encode, or be associated with a cargo or payload that is a binding protein, which may be a simple protein that binds to a non-protein. Non-limiting examples of binding proteins include glycoproteins, hemoglobin, lecithin proteins, nucleoproteins, and phosphoproteins.

[0113] In some embodiments, the starting and reference constructs of the invention may comprise, encode, or be associated with a cargo or payload that is a derivatized protein that is derived from a simple or bound protein by chemical or physical means. derived protein. Non-limiting examples of derivatized proteins include denatured proteins and peptides.

[0114] In some embodiments, a polypeptide, protein or peptide may be unmodified.

[0115] In some embodiments, a polypeptide, protein or peptide can be modified. Types of modifications include, but are not limited to, phosphorylation, glycosylation, acetylation, ubiquitination / sucination, methylation, palmitoylation, quinone, amidation, myristylation, pyrrolidone carboxylic acid, hydroxyl Phosphopantetheine, prenylation, GPI anchoring, oxidation, ADP ribosylation, sulfation, S-nitrosylation, citrullination, nitration, γ-carboxyglutamate , formylation, hydroxyputresine lysine, topaquinone (Topaquinone; TPQ), bromination, lysine topaquinone (LTQ), tryptophan tryptopylquinone (Tryptophan tryptopylquinone; TTQ), iodide and cysteine ​​tryptophan quinone (CTQ). In some aspects, a polypeptide, protein or peptide can be modified by post-transcriptional modifications that can affect its structure, subcellular localization and / or function.

[0116] In some embodiments, a polypeptide, protein or peptide can be modified using phosphorylation. Phosphorylation or addition of phosphate groups to serine, threonine or tyrosine residues is one of the most common forms of protein modification. Protein phosphorylation plays an important role in fine-tuning signaling in intracellular signaling cascades.

[0117] In some embodiments, a polypeptide, protein or peptide can be modified using ubiquitination, which is the covalent attachment of ubiquitin to a protein of interest. Ubiquitination-mediated protein turnover has been shown to play a role in driving the cell cycle as well as protein-degradation-independent intracellular signaling pathways.

[0118] In some embodiments, polypeptides, proteins or peptides can be modified with acetylation and methylation which can play a role in regulating gene expression. As a non-limiting example, acetylation and methylation can mediate the formation of chromosomal domains such as euchromatin and heterochromatin, which may have effects in mediating gene silencing.

[0119] In some embodiments, polypeptides, proteins or peptides can be modified using glycosylation. Glycosylation is the attachment of one of a large number of glycan groups and is a modification that occurs in about half of all proteins and plays a role in biological processes including, but not limited to, embryonic development, cell division, and protein structure adjust. The two main types of protein glycosylation are N-glycosylation and O-glycosylation. For N-glycosylation, glycans are linked to asparagine; and for O-glycosylation, glycans are linked to serine or threonine.

[0120] In some embodiments, polypeptides, proteins or peptides can be modified using Susuylation. Susuylation is the addition of SUMO (small ubiquitin-like modulator) to proteins and is a post-translational modification similar to ubiquitinylation. Antibody

[0121] As used herein, the term "antibody" refers in its broadest sense and specifically covers various embodiments including, but not limited to, monoclonal antibodies, polyclonal antibodies, multispecific antibodies (such as biclonal antibodies formed from at least two whole antibodies) Specific antibodies) and antibody fragments (such as diabodies), as long as they exhibit the desired biological activity (such as "functionality"). Antibodies are primarily amino acid based molecules, which are monomeric or polymeric polypeptides, comprising at least one amino acid region derived from a known or parental antibody sequence and at least one amino acid region derived from a non-antibody sequence. Antibodies may contain one or more modifications (including but not limited to added sugar moieties, fluorescent moieties, chemical tags, etc.). For purposes herein, an "antibody" may comprise heavy and light chain variable domains and an Fc region.

[0122] The cargo or payload may comprise or may encode polypeptides that form one or more functional antibodies.

[0123] In some embodiments, the cargo or payload may comprise or may encode a polypeptide that forms or acts as any antibody, including but not limited to antibodies known in the art and / or commercially available as therapeutic, diagnostic or useful Antibodies for research purposes. Additionally, the cargo or payload may comprise or may encode fragments of such antibodies, such as, but not limited to, variable domains or complementarity determining regions (CDRs).

[0124] As used herein, the term "primary antibody" refers to a general heterotetrameric glycoprotein of about 150,000 Daltons, which is composed of two identical light (L) chains and two identical heavy (H) chains. The genes encoding antibody heavy and light chains are known and their respective constituent segments have been well characterized and described (Matsuda, F. et al., 1998. The Journal of Experimental Medicine. 188(11); 2151-62 and Li, A et al., 2004. Blood. 103(12): 4602-9, the content of each case is incorporated herein by reference in its entirety). Each light chain is linked to a heavy chain by one covalent disulfide bond, while the number of disulfide bonds varies among heavy chains of different immunoglobulin isotypes. Each heavy and light chain also has regularly spaced intrachain disulfide bridges. Each heavy chain has a variable domain (VH) at one end followed by constant domains. Each light chain has a variable domain at one end (VL) and a constant domain at its other end; the constant domain of the light chain is aligned with the first constant domain of the heavy chain, and the variable domain of the light chain is aligned with the variable domain of the heavy chain. alignment. As used herein, the term "light chain" refers to the components of antibodies from any vertebrate species that fall into one of two distinct classes based on the amino acid sequences of the constant domains, referred to as kappa and lambda. . Depending on the amino acid sequence of the constant domain of their heavy chains, antibodies can be assigned to different "classes." There are five major classes of intact antibodies: IgA, IgD, IgE, IgG, and IgM, and several of these antibodies can be further divided into subclasses (isotypes), such as IgGl, IgG2, IgG3, IgG4, IgA, and IgA2.

[0125] As used herein, the term "variable domain" refers to specific antibody domains found on both the heavy and light chains of antibodies, which vary widely among antibodies in sequence and are used for the binding and binding of each specific antibody to its specific antigen. specificity. Variable domains comprise hypervariable regions. As used herein, the term "hypervariable region" refers to the region within a variable domain comprising the amino acid residues responsible for antigen binding. Amino acids present within the hypervariable regions determine the structure of the complementarity determining regions (CDRs) which become part of the antigen binding site of an antibody. As used herein, the term "CDR" refers to the region of an antibody that contains structures complementary to its target antigen or epitope. The rest of the variable domains that do not interact with the antigen are called the framework (FW) regions. The antigen binding site (also known as the antigen combining site or paratope) comprises the amino acid residues necessary to interact with a particular antigen. The exact residues that make up the antigen-binding site are usually elucidated by co-crystallography with the bound antigen, however, computational assessment based on comparison with other antibodies can also be used (Strohl, W.R. Therapeutic Antibody Engineering. Woodhead Publishing, Philadelphia PA. 2012. Chapter 3, pp. 47-54, the content of which is incorporated herein by reference in its entirety). Determining the residues that make up a CDR may involve the use of numbering schemes, including but not limited to those taught by: Kabat [Wu, T.T. et al., 1970, JEM, 132(2):211-50 and Johnson, G. et al. , 2000, Nucleic Acids Res. 28(1): 214-8, the contents of which are each incorporated herein by reference in their entirety], Chothia [Chothia and Lesk, J. Mol. Biol. 196, 901 (1987), Chothia et al., Nature 342, 877 (1989) and Al-Lazikani, B. et al., 1997, J. Mol. Biol. 273(4):927-48, the contents of which are each incorporated herein by reference in their entirety] , Lefranc (Lefranc, M.P. et al., 2005, Immunome Res. 1:3) and Honegger (Honegger, A. and Pluckthun, A. 2001. J. Mol. Biol. 309(3):657-70, the contents of which are respectively incorporated herein by reference in its entirety).

[0126] The VH and VL domains each have three CDRs. The V LCDRs are referred to herein as CDR-L1, CDR-L2, and CDR-L3 in order of appearance as they travel from N-terminus to C-terminus along the variable domain polypeptide. The V HCDRs are referred to herein as CDR-H1, CDR-H2, and CDR-H3 in order of appearance as they travel from N-terminus to C-terminus along the variable domain polypeptide. Each of the CDRs has an advantageous typical structure, except for CDR-H3, which contains amino acid sequences that are highly variable in sequence and length between antibodies, resulting in a variety of three-dimensional structures in the antigen-binding domain. In some cases, CDR-H3 can be analyzed in a group of related antibodies to assess antibody diversity.

[0127] Various methods of determining CDR sequences are known in the art and can be applied to known antibody sequences. The system described by Kabat, also known as "numbering according to Kabat", "Kabat numbering", "Kabat definition" and "Kabat labeling", provides an unambiguous residue numbering system applicable to any variable domain of an antibody and provides a definition The exact residue boundaries of the three CDRs of each chain. (Kabat et al., Sequences of Proteins of Immunological Interest, National Institutes of Health, Bethesda, Md. (1987) and (1991), the contents of which are incorporated by reference in their entirety). The Kabat CDR comprises approximately residues 24-34 (CDR1), 50-56 (CDR2), and 89-97 (CDR3) in the variable domain of the light chain, and residues 31-35 (CDR1) in the variable domain of the heavy chain , 50-65 (CDR2) and 95-102 (CDR3). Chothia and colleagues found that certain subportions within the Kabat CDRs adopt nearly identical peptide backbone configurations despite enormous diversity at the amino acid sequence level. (Chothia et al. (1987) J. Mol. Biol. 196: 901-917; and Chothia et al. (1989) Nature 342: 877-883, the contents of each of which are incorporated herein by reference in their entirety). These CDRs may be referred to as "Chothia CDRs", "Chothia Numbering" or "According to Chothia Numbering" and comprise about residues 24-34 (CDR1), 50-56 (CDR2) and 89 in the light chain variable domain. -97 (CDR3) and residues 26-32 (CDR1), 52-56 (CDR2) and 95-102 (CDR3) in the heavy chain variable domain. Mol. Biol. 196:901-917 (1987). The system described by MacCallum, also known as "numbering according to MacCallum" or "MacCallum numbering", comprises approximately residues 30-36 (CDR1), 46-55 (CDR2), and 89-96 (CDR3) in the light chain variable domain. ) and residues 30-35 (CDR1), 47-58 (CDR2) and 93-101 (CDR3) in the heavy chain variable domain. (MacCallum et al. ((1996) J. Mol. Biol. 262(5):732-745), the contents of which are hereby incorporated by reference in their entirety). The system described by AbM, also known as "numbering according to AbM" or "AbM numbering", comprises about residues 24-34 (CDR1), 50-56 (CDR2) and 89-97 (CDR3) in the light chain variable domain. ) and residues 26-35 (CDR1), 50-58 (CDR2) and 95-102 (CDR3) in the heavy chain variable domain. The IMGT (International Immunogenetics Information System (INTERNATIONAL IMMUNOGENETICS INFORMATION SYSTEM)) numbering of the variable region can also be used, which is the numbering of residues in the variable heavy or light chain of an immunoglobulin according to the method of IIMGT (Lefranc, M .-P., "The IMGT unique numbering for immunoglobulins, T cell Receptors and Ig-like domains", The Immunologist, 7, 132-136 (1999), and is incorporated herein by reference in its entirety). As used herein, "IMGT sequence numbering" or "numbering according to IMTG" refers to the numbering of sequences encoding variable regions according to IMGT. For heavy chain variable domains, when numbered according to IMGT, the hypervariable region ranges from amino acid positions 27 to 38 for CDR1, 56 to 65 for CDR2, and for For CDR3 in the range of amino acid positions 105 to 117. For light chain variable domains, when numbered according to IMGT, the hypervariable region ranges from amino acid positions 27 to 38 for CDR1, amino acid positions 56 to 65 for CDR2, and for For CDR3 in the range of amino acid positions 105 to 117.

[0128] In some embodiments, the cargo or payload may comprise or may encode antibodies that have been produced using methods known in the art, such as, but not limited to, immunization and display techniques (e.g., phage display, yeast display, and ribosomal presented), fusionoma technology, heavy and light chain variable region cDNA sequences selected from fusionoma or other sources.

[0129] In some embodiments, the cargo or payload may comprise or may encode antibodies developed using any naturally occurring or synthetic antigen. As used herein, an "antigen" is an entity that induces or induces an immune response in an organism. An immune response is characterized by the response of the cells, tissues and / or organs of an organism to the presence of a foreign entity. Such an immune response typically causes the organism to produce one or more antibodies against the foreign entity (eg, an antigen or a portion of an antigen). As used herein, "antigen" also refers to a binding partner of a specific antibody or a binding agent present in a repertoire.

[0130] As used herein, the term "monoclonal antibody" refers to an antibody obtained from a population of cells (or colonies) that is substantially homogeneous, that is, excluding possible variants that may have arisen during the production of the monoclonal antibody (such variants Except usually in small amounts), the individual antibodies comprising the population are identical and / or bind to the same epitope. Each monoclonal antibody is directed against a single determinant on an antigen, in contrast to polyclonal antibody preparations, which often include different antibodies directed against different determinants (epitopes).

[0131] The modifier "monoclonal" indicates that the antibody is characterized as being obtained from a substantially homogeneous population of antibodies and should not be construed as requiring that the antibody be produced by any particular method. Monoclonal antibodies include herein "chimeric" antibodies (immunoglobulins) in which a portion of the heavy and / or light chain is identical to or identical to the corresponding sequence in an antibody derived from a particular species or belonging to a particular antibody class or subclass. Homologous, with the remainder of the chain being identical or homologous to the corresponding sequence in an antibody derived from another species or belonging to another antibody class or subclass, and fragments of such antibodies.

[0132] As used herein, the term "humanized antibody" refers to a chimeric antibody comprising minimal portions derived from one or more non-human (e.g., murine) antibody sources and the remainder derived from one or more human immune Source of globulin. For the most part, humanized antibodies are human immunoglobulins (recipient antibodies) in which residues from the hypervariable regions of the recipient antibody have been derived from a non-human species such as mouse, rat, rabbit, or nonhuman. Substitution of residues in the hypervariable region of an antibody (donor antibody) having the desired specificity, affinity and / or capacity of a human primate).

[0133] In some embodiments, the cargo or payload can comprise or can encode an antibody mimetic. As used herein, the term "antibody mimetic" refers to any molecule that mimics the function or action of an antibody and binds specifically and with high affinity to its molecular target. In some embodiments, an antibody mimetic may be a monofunctional antibody designed to incorporate a fibronectin type III domain (Fn3) as the protein backbone. In some embodiments, antibody mimetics may be those known in the art including, but not limited to, affinibody molecules, affilin, affitin, anticalin , high affinity multimers, Centyrin, DARPINS™, fynomers, Kunitz domains and domain peptides. In other embodiments, antibody mimetics may include one or more non-peptide regions. [Antibody Fragments and Variants] []

[0134] In some embodiments, the cargo or payload may comprise or may encode an antibody fragment comprising an antigen binding region from a full length antibody. Non-limiting examples of antibody fragments include Fab, Fab', F(ab')2, and Fv fragments, diabodies, linear antibodies, single chain antibody molecules, and multispecific antibodies formed from antibody fragments. Papain digestion of antibodies produces two identical antigen-binding fragments, termed "Fab" fragments, each with a single antigen-binding site. A residual "Fc" fragment is also produced, whose name reflects its ability to readily crystallize. Pepsin treatment yields an F(ab')2 fragment that has two antigen-combining sites and is still capable of cross-linking antigen. The compounds and / or compositions of the invention may comprise one or more of these fragments.

[0135] In some embodiments, the Fc region can be a modified Fc region, wherein the Fc region can have a single amino acid substitution compared to the corresponding sequence of a wild-type Fc region, wherein the single amino acid substitution results in a wild-type Fc region having The preferred properties of the Fc region are those of the Fc region. Non-limiting examples of Fc properties that can be altered by single amino acid substitutions include binding properties or response to pH conditions.

[0136] As used herein, the term "Fv" refers to an antibody fragment comprising the minimal fragment on the antibody required to form a complete antigen binding site. These regions consist of dimers of one heavy chain and one light chain variable domain in tight, non-covalent association. Fv fragments can be produced by proteolytic cleavage, but most are unstable. Recombinant methods for producing stable Fv fragments are known in the art, usually via insertion of a flexible linker between the light and heavy chain variable domains to form a single chain Fv (scFv), or By introducing disulfide bridges between the heavy and light chain variable domains.

[0137] As used herein, the term "single-chain Fv" or "scFv" refers to a fusion protein of VH and VL antibody domains, wherein these domains are linked together into a single polypeptide chain by a flexible peptide linker. In some embodiments, the Fv polypeptide linker enables the scFv to form a desired structure for antigen binding. In some embodiments, scFvs are used in conjunction with phage display, yeast display, or other display methods where they can be displayed in conjunction with surface members (eg, phage coat proteins) and used to identify high affinity peptides of a given antigen.

[0138] As used herein, the term "antibody variant" refers to a modified antibody (relative to a native or starting antibody) or biomolecule (eg, an antibody mimic) that is similar in structure and / or function to a native or starting antibody. Antibody variants may differ in their amino acid sequence, composition, or structure compared to the native antibody. Antibody variants may include, but are not limited to, antibodies with altered isotypes (e.g., IgA, IgD, IgE, IgG 1 , IgG 2, IgG 3, IgG 4, or IgM), humanized variants, optimized variants, polymorphic variants, Specific antibody variants (eg, bispecific variants) and antibody fragments. [Multispecific Antibody] []

[0139] In some embodiments, the cargo or payload can be or can encode an antibody that binds more than one epitope. As used herein, the term "multifunctional antibody" or "multispecific antibody" refers to an antibody in which two or more variable domains bind to different epitopes. Epitopes can be on the same or different targets. In certain embodiments, multispecific antibodies are "bispecific antibodies," which recognize two different epitopes on the same or different antigens.

[0140] In some embodiments, multispecific antibodies can be prepared by the methods used by BIOATLA® and described in International Patent Publication WO201109726, the contents of which are incorporated herein by reference in their entirety. First, a library of cognate, naturally occurring antibodies (i.e., displayed on the surface of mammalian cells) is generated by any method known in the art, followed by FACSAria or another screening method against two or more Screen for multispecific antibodies that specifically bind to a target antigen. In some embodiments, the identified multispecific antibodies are further evolved by any method known in the art to generate a panel of modified multispecific antibodies. These modified multispecific antibodies are screened for binding to the target antigen. In some embodiments, multispecific antibodies can be further optimized by screening evolved modified multispecific antibodies for optimized or desired characteristics.

[0141] In some embodiments, multispecific antibodies can be prepared by the methods used by BIOATLA® and described in US Publication No. US20150252119, the contents of which are incorporated herein by reference in their entirety. In one approach, the variable domains of two parental antibodies (where the parental antibodies are monoclonal antibodies) are made such that a single light chain is functionally identical to that of the two different parental antibodies, using any method known in the art. Evolution of heavy chain complementarity. Another approach entails evolving the heavy chain of a single parental antibody to recognize a second target antigen. A third approach involves evolving the light chains of the parental antibody to recognize a second target antigen. Methods for polypeptide evolution are described in International Publication WO2012009026, the contents of which are incorporated herein by reference in their entirety, and include, by way of non-limiting examples, comprehensive positional evolution (CPE), combinatorial protein synthesis (CPS), comprehensive positional Insertion (CPI), combined positional deletion (CPD) or any combination thereof. The Fc region of the multispecific antibody described in US Publication No. US20150252119 can be generated using the knob-in-hole approach or any other method that enables the Fc domain to form heterodimers. The resulting multispecific antibodies can be further evolved to acquire improved characteristics or properties, such as binding affinity for the target antigen. [Bispecific Antibody] []

[0142] In some embodiments, the cargo or payload can be or can encode a bispecific antibody. As used herein, the term "bispecific antibody" refers to an antibody capable of binding two different antigens. Such antibodies typically comprise regions from at least two different antibodies. Such antibodies typically comprise antigen binding regions from at least two different antibodies. For example, bispecific monoclonal antibodies (BsMAb, BsAb) are artificial proteins composed of fragments of two different monoclonal antibodies, thus enabling the BsAb to bind two different types of antigens.

[0143] In some cases, the cargo or payload can be or can encode a bispecific antibody comprising antigen binding regions from two different anti-tau antibodies. For example, such bispecific antibodies may comprise binding regions from two different antibodies.

[0144] The bispecific antibody framework may comprise any of those described in: Riethmuller, G., 2012. Cancer Immunity. 12:12-18; Marvin, J.S. et al., 2005. Acta Pharmacologica Sinica. 26( 6):649-58; and Schaefer, W. et al., 2011. PNAS. 108(27):11187-92, the contents of each of which are incorporated herein by reference in their entirety.

[0145] A new generation of BsMAbs, known as "trifunctional bispecific" antibodies, have been developed. These antibodies consist of two heavy chains and two light chains, each from two different antibodies, with two Fab regions (arms) directed against two antigens and an Fc region (foot) comprising the two heavy chains and forming the third binding site.

[0146] Of the two paratopes forming the top of the variable domain of the bispecific antibody, one may be directed against the target antigen and the other against a T-lymphocyte antigen, such as CD3. In the case of triantibodies, the Fc region can additionally bind to cells expressing Fc receptors, such as macrophages, natural killer (NK) cells or dendritic cells. In general, the targeted cell engages with one or two cells of the immune system, which then destroys the targeted cell.

[0147] Other types of bispecific antibodies have been designed to overcome certain problems, such as short half-life, immunogenicity, and side effects caused by release of cytokines. It includes chemically linked Fabs (consisting of Fab regions only), and various types of bivalent and trivalent single-chain variable fragments (scFv), fusion proteins that mimic the variable domains of two antibodies. These newer formats that have recently been developed are bispecific T cell engaging molecules (BiTEs) and mAb2, antibodies engineered to contain the Fcab antigen-binding fragment replacing the Fc constant region.

[0148] Using molecular genetics, two scFvs can be engineered in tandem into a single polypeptide, separated by a linker domain, termed "tandem scFv" (tascFv). TascFv has been found to be poorly soluble and requires refolding when produced in bacteria, or it can be produced in mammalian cell culture systems, avoiding the need for refolding, but possibly resulting in poor yields. The construction of a tascFv with the genes of two different scFvs results in a "bispecific single chain variable fragment" (dual scFv). There are only two tascFvs that have been developed clinically by commercial companies; both are bispecific agents for oncology indications that are actively developed by Micromet at an early stage, and are described as “bispecific T cell engaging molecules (BiTEs)”. Blinatumomab is an anti-CD19 / anti-CD3 bispecific tascFv that potentiates T cell responses to stage 2 B-cell non-Hodgkin lymphoma. MT110 is an anti-EP-CAM / anti-CD3 bispecific tascFv that potentiates T cell responses against stage 1 solid tumors. Affimed is also researching bispecific tetravalent "TandAb".

[0149] In some embodiments, the cargo or payload can be or can encode an antibody comprising a single antigen binding domain. These molecules are extremely small, with molecular weights roughly one-tenth those observed for full-sized mAbs. Other antibodies may include "nanobodies" derived from the antigen-binding variable heavy chain region (VHH) of heavy chain antibodies found in camels and llamas, which do not have light chains.

[0150] PCT Publication WO2014144573 of the Memorial Sloan-Kettering Cancer Center, the contents of which are incorporated herein by reference in its entirety, discloses and claims to be useful in the preparation of bispecific binding agents with improved properties relative to those without the ability to dimerize. Multimerization techniques for multispecific binding agents, such as fusion proteins comprising antibody components.

[0151] In some cases, the cargo or payload can be or can encode a tetravalent bispecific antibody (such as the TetBiAb disclosed and claimed in PCT Publication WO2014144357, the contents of which are incorporated herein in its entirety). TetBiAbs are characterized by a second pair of Fab fragments with a second antigen specificity linked to the C-terminus of the antibody, thus providing a molecule that is bivalent to each of the two antigen specificities. Tetravalent antibodies are produced by genetic engineering methods by covalently linking antibody heavy chains to Fab light chains that associate with their cognate, co-expressed Fab heavy chains.

[0152] In some aspects, the cargo or payload can be or can encode a biosynthetic antibody as described in US Patent No. 5,091,513, the contents of which are incorporated herein by reference in their entirety. Such antibodies may include one or more sequences of amino acids constituting a region serving as a biosynthetic antibody binding site (BABS). Sites include 1) non-covalently or disulfide-bonded synthetic VH and VL dimers, 2) VH-VL or VL-VH single chains, where VH and VL are linked by a polypeptide linker, or 3) Individual VH or VL domains. Binding domains comprise linked CDR and FR regions, which can be derived from individual immunoglobulins. Biosynthetic antibodies may also include other polypeptide sequences that serve, for example, as enzymes, toxins, binding sites, or attachment sites to immobilization media or radioactive atoms. Methods for making biosynthetic antibodies, for designing BABS with any specificity that can be elicited by in vivo production of antibodies, and for making analogs thereof are disclosed.

[0153] In some embodiments, the cargo or payload can be or can encode an antibody having the antibody receptor framework taught in US Pat. No. 8,399,625. Such antibody acceptor frameworks may be particularly suitable for accepting CDRs from an antibody of interest. In some cases, CDRs from anti-tau antibodies known in the art or developed according to the methods presented herein can be used. [Small] [Typed Antibody] []

[0154] In some embodiments, the cargo or payload can be or can encode a "miniature" antibody. Best examples of mAb miniaturization include Small Modular Immunopharmaceuticals (SMIP) from Trubion Pharmaceuticals. These molecules, which may be monovalent or bivalent, are recombinant single chain molecules containing one VL, one VH antigen binding domain and one or two constant "effector" domains, all linked by a linker domain. It is possible that such molecules may provide the enhanced tissue or tumor penetration claimed by the fragments while retaining the advantages of the immune effector functions conferred by the constant domains. At least three "miniature" SMIPs have entered clinical development. TRU-015, an anti-CD20 SMIP developed in collaboration with Wyeth, is the most advanced program in Phase 2 development for rheumatoid arthritis (RA). Earlier attempts in systemic lupus erythematosus (SLE) and B-cell lymphoma were eventually discontinued. Trubion and Facet Biotechnology are collaborating to develop TRU-016, an anti-CD37 SMIP, for the treatment of CLL and other lymphoid neoplasms, which is a Phase 2 project. Wyeth has approved the anti-CD20 SMIP SBI-087 for the treatment of autoimmune diseases including RA, SLE and possibly multiple sclerosis, but these programs are still in the earliest stages of clinical testing. [Bifunctional antibody] []

[0155] In some embodiments, the cargo or payload can be or can encode a diabody. As used herein, the term "diabodies" refers to small antibody fragments that have two antigen combining sites. A diabody comprises a heavy chain variable domain VH linked to a light chain variable domain VL in the same polypeptide chain. By using a linker that is too short to allow pairing between the two domains on the same chain, the domains are forced to pair with the complementary domains of another chain and two antigen-binding sites are created.

[0156] The bifunctional antibody is a functional bispecific single chain antibody (bscAb). These bivalent antigen binding molecules are composed of non-covalent dimers of scFv and can be produced in mammalian cells using recombinant methods. (See eg Mack et al., Proc. Natl. Acad. Sci., 92: 7021-7025, 1995). Few bifunctional antibodies have entered clinical development. In a study sponsored by the Beckman Research Institute of the City of Hope (ClinicalTRIals.gov NCT00647153), an iodine-123-labeled diabody version of the anti-CEA chimeric antibody cT84.66 was evaluated for preoperative immunization of colorectal cancer Blink scan detection. [Single antibody] []

[0157] In some embodiments, the cargo or payload can be or can encode a "unibody" in which the hinge region has been removed from the IgG4 molecule. Although IgG4 molecules are unstable and can exchange light chain-heavy chain heterodimers with each other, deletion of the hinge region completely prevents heavy chain-heavy chain pairing, leaving highly specific monovalent light chain / heavy chain heterodimers, At the same time, the Fc region is retained to ensure stability and half-life in vivo. This configuration minimizes the risk of immune activation or oncogenic growth, since IgG4 interacts poorly with FcRs and monovalent monoantibodies cannot promote intracellular signaling complex formation. However, such views are mainly supported by laboratory rather than clinical evidence. Other antibodies may be "miniature" antibodies, which are compacted 100 kDa antibodies. [Intrabodies] []

[0158] In some embodiments, the cargo or payload can be or can encode an intrabody. The term "intrabody" refers to a form of antibody that is not secreted by the cell in which it is produced, but instead targets one or more intracellular proteins. Intrabodies can be used to affect multiple cellular processes including, but not limited to, intracellular migration, transcription, translation, metabolic processes, proliferative signaling, and cell division. In some embodiments, the methods of the invention may include intrabody-based therapy. In some such embodiments, the variable domain sequences and / or CDR sequences disclosed herein can be incorporated into one or more constructs for intrabody-based therapy. For example, an intrabody can target one or more glycated intracellular proteins or can modulate the interaction between one or more glycated intracellular proteins and an alternative protein.

[0159] Intrabodies against intracellular targets were first described more than two decades ago (Biocca, Neuberger and Cattaneo EMBO J. 9: 101-108, 1990, the contents of which are hereby incorporated by reference in their entirety). Intracellular expression of intrabodies in different compartments of mammalian cells allows blocking or modulating the function of endogenous molecules (Biocca et al., EMBO J. 9: 101-108, 1990; Colby et al., Proc. Natl. Acad. Sci. U.S.A. 101: 17616-21, 2004, the contents of which are incorporated herein by reference in their entirety). Intrabodies can alter protein folding, protein-protein, protein-DNA, protein-RNA interactions, and protein modification. It can induce phenotypic gene knockout and act as a neutralizing agent by binding directly to the target antigen, by diverting its intracellular migration, or by inhibiting its association with a binding partner. It is primarily used as a research tool and has emerged as a therapeutic molecule for the treatment of human diseases such as viral lesions, cancer and misfolding diseases. With regard to their use in therapy, the rapidly growing biological market for recombinant antibodies provides intrabodies with enhanced binding specificity, stability and solubility, and lower immunogenicity.

[0160] In some embodiments, intrabodies have advantages over interfering RNA (iRNA); for example, iRNA has been shown to exert a variety of non-specific effects, while intrabodies have been shown to have high specificity and affinity for target antigens. Furthermore, as proteins, intrabodies have a much longer active half-life than iRNAs. Therefore, when the active half-life of the intracellular target molecule is long, the process of gene silencing by iRNA may be slower, whereas the effect of intrabody expression may be almost instantaneous. Ultimately, it may be possible to design intrabodies to block certain binding interactions of specific target molecules while retaining other effects.

[0161] Intrabodies are typically single-chain variable fragments (scFv) expressed by recombinant nucleic acid molecules and engineered to persist within the cell (eg, in the cytoplasm, endoplasmic reticulum, or extracellular cytoplasm). For example, intrabodies can be used to abolish the function of proteins to which the intrabody binds. Expression of intrabodies can also be modulated through the use of inducible promoters in nucleic acid expression vectors comprising intrabodies. Intrabodies for use in the viral gene bodies of the present invention can be generated using methods known in the art, such as those disclosed and reviewed in: Marasco et al., 1993 Proc. Natl. Acad. Sci. USA, 90: 7889 -7893; Chen et al., 1994, Hum. Gene Ther.5:595-601; Chen et al., 1994, Proc. Natl. Acad. Sci. USA, 91: 5932-5936; Maciejewski et al., 1995, Nature Med ., 1: 667-673; Marasco, 1995, Immunotech, 1: 1-19; Mhashilkar, et al., 1995, EMBO J. 14: 1542-51; Chen et al., 1996, Hum. Gene Therap., 7: 1515-1525; Marasco, Gene Ther. 4:11-15, 1997; Rondon and Marasco, 1997, Annu. Rev. Microbiol.51:257-283; Cohen, et al., 1998, Oncogene 17:2445-56; Proba et al. People, 1998, J. Mol. Biol. 275:245-253; Cohen et al., 1998, Oncogene 17:2445-2456; Hassanzadeh, et al., 1998, FEBS Lett. 437:81-6; Richardson et al., 1998, Gene Ther. 5:635-44; Ohage and Steipe, 1999, J. Mol. Biol. 291:1119-1128; Ohage et al., 1999, J. Mol. Biol.291:1129-1134; Wirtz and Steipe, 1999 , Protein Sci. 8:2245-2250; Zhu et al., 1999, J. Immunol. Methods231:207-222; Arafat et al., 2000, Cancer Gene Ther. 7:1250-6; der Maur et al., 2002, J 277:45075-85; Mhashilkar et al., 2002, Gene Ther. 9:307-19; and Wheeler et al., 2003, FASEB J. 17: 1733-5; and references cited therein). In particular, a CCR5 intrabody has been generated by Steinberger et al., 2000, Proc. Natl. Acad. Sci. USA 97:805-810. See generally Marasco, WA, 1998, "Intrabodies: Basic Research and Clinical Gene Therapy Applications" Springer: New York; and for a review of scFv, see Pluckthun, "The Pharmacology of Monoclonal Antibodies", 1994, Vol. 113, Rosenburg and Moore eds, Springer-Verlag, New York, pp. 269-315; the contents of each are each incorporated by reference in their entirety.

[0162] Sequences from donor antibodies can be used to generate intrabodies. Intrabodies are typically expressed recombinantly within cells as single domain fragments (such as isolated VH and VL domains) or as single chain variable fragment (scFv) antibodies. For example, intrabodies are often expressed as a single polypeptide to form a single chain antibody comprising the variable domains of the heavy and light chains joined by a flexible linker polypeptide. Intrabodies generally lack disulfide bonds and are capable of modulating the expression or activity of target genes through their specific binding activity. Single chain antibodies can also be expressed as a single chain variable region fragment joined to a light chain constant region.

[0163] As known in the art, intrabodies can be engineered into recombinant polynucleotide vectors to encode daughter cell migration signals at their N- or C-termini to allow expression at high concentrations in the daughter cell compartment where the protein of interest is located . For example, intrabodies targeting the endoplasmic reticulum (ER) are engineered to incorporate a leader peptide and optionally a C-terminal ER retention signal. Intrabodies intended to be active in the nucleus are engineered to include a nuclear localization signal. The lipid moiety is conjugated to the intrabody to tether the intrabody to the cytoplasmic side of the plasma membrane. Intrabodies can also be targeted to function in the cytosol. For example, cytoplasmic intrabodies are used to sequester factors within the cytosol, thereby preventing their transport to their natural cellular destination.

[0164] There are certain technical challenges with intrabody expression. In particular, the protein conformational folding and structural stability of newly synthesized intrabodies in cells are affected by the reducing conditions of the intracellular environment.

[0165] The intrabodies of the invention are due to their almost unlimited ability to specifically recognize different conformations of proteins, including pathological isoforms, and because they can target potential aggregation sites (both intracellular and extracellular sites) and may be a promising treatment for misfolding diseases, including tauopathies, prion diseases, Alzheimer's, Parkinson's, and Huntington's agent. These molecules may act as neutralizers for amyloid proteins by preventing their aggregation, and / or act as molecular dispatchers of intracellular traffic by rerouting proteins from their potential aggregation sites. [Maximum Antibody] []

[0166] In some embodiments, the cargo or payload can be or can encode a maxibody (a bivalent scFV fused to the amino terminus of the Fc (CH2-CH3 domain) of IgG). Chimeric Antigen Receptor (CAR)

[0167] In some embodiments, the cargo or payload can be or can encode a chimeric antigen receptor (CAR), which when transduced into immune cells (such as T cells and NK cells) can redirect immune cells to express The target of a molecule recognized by the extracellular target portion of the CAR (eg, a tumor cell).

[0168] As used herein, the term "chimeric antigen receptor (CAR)" refers to a synthetic receptor that mimics the TCR on the surface of T cells. In general, a CAR consists of an extracellular targeting domain, a transmembrane domain / region, and an intracellular signaling / activation domain. In standard CAR receptors, the components: extracellular targeting domain, transmembrane domain and intracellular signaling / activation domain are linearly constructed as a single fusion protein. The extracellular region contains targeting domains / moieties (eg scFv) that recognize specific tumor antigens or other tumor cell surface molecules. The intracellular region may contain the signaling domain of the TCR complex (eg, the signaling region of CD3ζ) and / or one or more co-stimulatory signaling domains, such as those from CD28, 4-1BB (CD137) and OX-40 (CD134) By. For example, the "first generation CAR" has only the CD3ζ signaling domain, and in order to increase T cell persistence and proliferation, a costimulatory intracellular domain is added, resulting in a second generation CAR with the CD3ζ signaling domain plus a costimulatory signaling domain , and third-generation CARs with a CD3ζ signaling domain plus two or more co-stimulatory signaling domains. When expressed by T cells, CARs confer on T cells antigen specificity determined by the extracellular targeting portion of the CAR. In some aspects, one or more elements, such as homing and suicide genes, can be added to generate more competent and safer CAR architectures (so-called fourth-generation CARs).

[0169] In some embodiments, the extracellular targeting domain is joined to the intracellular signaling domain via a hinge (also known as a spacer domain or spacer) and a transmembrane region. The hinge connects the extracellular targeting domain with the transmembrane domain, which crosses the cell membrane and connects to the intracellular signaling domain. Due to the size of the target protein to which the targeting moiety binds and the size and affinity of the targeting domain itself, the hinge may need to be altered to optimize the efficacy of CAR-transformed cells against cancer cells. After the targeting moiety recognizes and binds the target cell, the intracellular signaling domain generates an activation signal to the CAR T cell, which is further amplified by a "secondary signal" of one or more intracellular co-stimulatory domains. Once activated, CAR T cells can destroy target cells.

[0170] In some embodiments, the CAR can be split into two parts, each linking the dimerization domain, such that an input that triggers dimerization promotes assembly of a fully functional receptor. Wu and Lim report a split CAR in which the extracellular CD19-binding domain and intracellular signaling element are separated and heterodimeric with the FKBP domain and FRB in the presence of the rapamycin analog AP21967* (FKBP-rapamycin binding T2089L mutant) domain linkage. Split receptors are assembled in the presence of AP21967 and activate T cells together with specific antigen binding (Wu et al., Science, 2015, 625(6258): aab4077, the contents of which are incorporated herein by reference in their entirety middle).

[0171] In some embodiments, the CAR can be designed as an inducible CAR that has incorporated the Tet-On inducible system into the CD19 CAR construct. CD19 CAR is only activated in the presence of doxycycline (Dox). Sakemura reported that compared with conventional CD19CAR T cells, Tet-CD19CAR T cells in the presence of Dox were equivalently cytotoxic to CD19+ cell lines, and had equivalent cytokine production and proliferation after CD19 stimulation ((Sakemura et al., Cancer Immuno. Res., June 21, 2016, electronic version; the contents of which are incorporated herein by reference in their entirety). The dual system provides more information on switching CAR expression on and off in transduced T cells. of flexibility.

[0172] In some embodiments, the cargo or payload can be or can encode a first generation CAR, or a second generation CAR, or a third generation CAR, or a fourth generation CAR. In some embodiments, the cargo or payload can be or can encode a complete CAR construct consisting of an extracellular domain, hinge and transmembrane domains, and an intracellular signaling region. In other embodiments, the cargo or payload can be or can encode components of a complete CAR construct, including an extracellular targeting moiety, a hinge region, a transmembrane domain, an intracellular signaling domain, one or more co-stimulatory domains, And other additional elements that improve CAR architecture and functionality, including but not limited to leader sequences, homing elements, and safety switches, or combinations of such components.

[0173] In some embodiments, the cargo or payload may be or encode a tunable CAR. The reversible on-off switching mechanism allows the management of acute toxicity caused by excessive CAR-T cell expansion. Ligand-conferred CAR modulation may effectively counteract tumor escape induced by antigen deprivation, thereby avoiding functional exhaustion due to tonic signaling due to chronic antigen exposure and improving persistence of CAR-expressing cells in vivo. Tunable CARs can be used with the goal of downregulating CAR expression to limit tissue toxicity caused by tumor lysis syndrome. Downregulation of CAR expression after antitumor efficacy prevents (1) off-target tumor toxicity caused by antigen expression in normal tissues; (2) antigen-independent activation in vivo. [Extracellular targeting domain] [ / ] [part] []

[0174] In some embodiments, the extracellular target portion of the CAR can be any agent that recognizes and binds a given target molecule (eg, a neoantigen on a tumor cell) with high specificity and affinity. The target moiety can be: an antibody and its variants that specifically bind to the target molecule on the tumor cell; or an aptamer selected from a random sequence pool based on its ability to bind to the target molecule on the tumor cell; or can bind to the target molecule on the tumor cell. or a variant or fragment thereof to which the target molecule binds; or an antigen recognition domain from the native T cell receptor (TCR) (such as the CD4 ectodomain that recognizes HIV-infected cells); or foreign recognition components, such as those that cause recognition of cells carrying Cytokines that bind to target cells for receptors or natural ligands for the receptors.

[0175] In some embodiments, the targeting domain of CAR can be Ig NAR, Fab fragment, Fab' fragment, F(ab)'2 fragment, F(ab)'3 fragment, Fv, single chain variable fragment (scFv), Double scFv, (scFv)2, minibody, diabody, triabody, tetrabody, disulfide bond-stabilized Fv protein (dsFv), monobody, nanobody or derived from specific recognition target molecule (e.g. The antigen-binding region of an antibody to a tumor-specific antigen (TSA). In one embodiment, the targeting moiety is a scFv antibody. When expressed on the surface of CAR T cells and subsequently bound to target proteins on cancer cells, the scFv domain is able to maintain the proximity of CAR T cells to cancer cells and trigger T cell activation. scFv can be produced using conventional recombinant DNA technology techniques and are discussed herein.

[0176] In some embodiments, the targeting moiety of the CAR construct may be an aptamer, such as a peptide aptamer, that specifically binds a target molecule of interest. Peptide aptamers can be selected from a pool of random sequences based on their ability to bind a target molecule of interest.

[0177] In some embodiments, the targeting moiety of the CAR construct may be a natural ligand of the target molecule or a variant and / or fragment thereof capable of binding the target molecule. In some aspects, the targeting moiety of the CAR can be a receptor for the target molecule, for example, full-length human CD27, which acts as a CD70 receptor, can be fused in-frame with the signaling domain of CD3ζ to form a CD70-positive malignancy. CD27 chimeric receptor for immunotherapeutics.

[0178] In some embodiments, the targeting portion of the CAR can recognize a tumor specific antigen (TSA), such as a cancer neoantigen that is limitedly expressed on tumor cells.

[0179] As a non-limiting example, a CAR of the invention may comprise an extracellular targeting domain capable of binding to a tumor-specific antigen selected from: 5T4, 707-AP, A33, AFP (alpha-fetoprotein), AKAP-4 ( A kinase-anchored protein 4), ALK, α5β1-integrin, androgen receptor, phospholipid binding protein II, α-actinin-4, ART-4, B1, B7H3, B7H4, BAGE (B melanoma antigen ), BCMA, BCR-ABL fusion protein, β-catenin, BKT-antigen, BTAA, CA-I (carbonic anhydrase I), CA50 (cancer antigen 50), CA125, CA15-3, CA195, CA242, calcium retina Protein, CAIX (carbonic anhydrase), cytotoxic T-lymphocyte recognized antigen on melanoma (CAMEL), CAM43, CAP-1, caspase-8 / m, CD4, CD5 , CD7, CD19, CD20, CD22, CD23, CD25, CD27 / m, CD28, CD30, CD33, CD34, CD36, CD38, CD40 / CD154, CD41, CD44v6, CD44v7 / 8, CD45, CD49f, CD56, CD68\KP1 , CD74, CD79a / CD79b, CD103, CD123, CD133, CD138, CD171, cdc27 / m, CDK4 (cyclin-dependent kinase 4), CDKN2A, CDS, carcinoembryonic antigen (carcinoembryonic antigen; CEA), CEACAM5, CEACAM6, staining Granin, c-Met, c-Myc, coa-1, CSAp, CT7, CT10, cyclophilin B, cyclin B1, cytoplasmic tyrosine kinase, cytokeratin, DAM-10, DAM-6, dek- can fusion protein, desmin, DEP domain containing 1 (DEPDC1), E2A-PRL, EBNA, epidermal growth factor receptor (EGF-R), epithelial cells Glycoprotein-1 (epithelial glycoprotein -1; EGP-1) (TROP-2), EGP-2, EGP-40, epidermal growth factor receptor (EGFR), EGFRvIII, EF-2, ELF2M, EMMPRIN, epithelial cell adhesion molecule (EpCAM), EphA2, Epstein-Barr virus antigen, Erb (ErbB1; ErbB3; ErbB4), epithelial tumor antigen (ETA), ETV6-AML1 Fusion protein, fibroblast activation protein (FAP), folate-binding protein (FBP), FGF-5, folate receptor, FOS-related antigen 1, fucosyl GM1, G250, GAGE (GAGE-1; GAGE-2), galectin, GD2 (ganglioside), GD3, glial fibrillary acidic protein (GFAP), GM2 (carcinoembryonic antigen-immunogenicity- 1; OFA-I-1), GnT-V, Gp100, H4-RET, HAGE (helicase antigen), HER-2 / neu, hypoxia inducible factor (HIF), HIF-1, HIF - 2, HLA-A2, HLA-A*0201-R170I, HLA-A1 l, HMWMAA, Hom / Mel-40, HSP70-2M (heat shock protein 70), HST-2, HTgp-175, hTERT (or hTRT ), HPV-E6 / HPV-E7 and E6, iCE (immune capture EIA), IGF-1R, IGH-IGK, IL-2R, IL-5, integrin-linked kinase (integrin - linked kinase; ILK), IMP3 (insulin-like growth factor II mRNA-binding protein 3), interferon regulatory factor 4 (interferon regulatory factor 4; IRF4), KDR (kinase insertion domain receptor), KIAA0205, KRAB-zinc finger Protein (KID)-3; KID31, KSA (17-1A), K-ras, LAGE, LCK, LDLR / FUT (LDLR-fucosyltransferase AS fusion protein), LeY (Lewis Y), MAD-CT - 1. MAGE (tyrosinase, melanoma-associated antigen) (MAGE-1; MAGE-3), melanin-A tumor antigen (MART), MART-2 / Ski, MC1R (melanocortin 1 receptor), MDM2, mesothelin, MPHOSPH1, muscle-specific actin (muscle-specific actin; MSA), mammalian targets of rapamycin (mTOR), MUC-1, MUC-2, MUM- 1 (melanoma-associated antigen (mutant) 1), MUM-2, MUM-3, myosin / m, MYL-RAR, NA88-A, N-acetylglucosyltransferase, neo-PAP, NF - KB (nuclear factor-κB), neurofilament, neuron-specific enolase (NSE), Notch receptor, NuMa, N-Ras, NY-BR-1, NY-CO-1, NY -ESO-1, Oncostatin M, OS-9, OY-TES1, p53 mutant, p190 microbcr-abl, pl5(58), pl85erbB2, pl80erbB-3, PAGE (prostate-associated gene), prostaglandin phosphate Enzyme (prostatic acid phosphatase; PAP), PAX3, PAX5, platelet derived growth factor receptor (platelet derived growth factor receptor; PDGFR), cytochrome P450 involved in the use of piperidine and pyrrolidine (PIPA), Pml-RARα fusion protein, PR -3 (protease 3), prostate specific antigen (prostate specific antigen; PSA), PSM, prostate stem cell antigen (Prostate stem cell antigen; PSMA), PRAME (antigen of preferentially expressed melanoma), PTPRK, RAGE (renal tumor antigen), Raf (A-Raf, B-Raf, and C-Raf), Ras, receptor tyrosine kinase, RCAS1, RGSS, ROR1 (receptor tyrosine kinase-like orphan receptor 1), RU1, RU2 , SAGE, SART-1, SART-3, SCP-1, SDCCAG16, SP-17 (sperm protein 17), src family, SSX (synovial sarcoma X breakpoint)-1, SSX-2 (HOM-MEL-40 ), SSX-3, SSX-4, SSX-5, STAT-3, STAT-5, STAT-6, STEAD, STn, survivin, syk-ZAP70, TA-90 (Mac-2 binding protein\cyclophilin C-associated protein), TAAL6, TACSTD1 (tumor-associated calcium signal transducer 1), TACSTD2, TAG-72-4, TAGE, TARP (T cell receptor gamma alternative reading frame protein), TEL / AML1 fusion protein, TEM1 , TEM8 (endosialin or CD248), TGFβ, TIE2, TLP, TMPRSS2 ETS fusion gene, TNF-receptor (TNF-α receptor, TNF-β receptor; or TNF-γ receptor), transferrin Receptor, TPS, tyrosine related protein 1 (tyrosine related protein 1; TRP-1), TRP-2, TRP-2 / INT2, TSP-180, VEGF receptor, WNT, WT-1 (Wilm's Tumor (Wilm's tumor) antigen) and XAGE.

[0180] In some embodiments, the cargo or payload can be or can encode a universal immune receptor-containing CAR with a targeting moiety capable of binding to a labeled antigen.

[0181] In some embodiments, the cargo or payload can be or can encode a CAR comprising a targeting moiety capable of binding to a pathogen antigen.

[0182] In some embodiments, the cargo or payload can be or can encode a CAR comprising a targeting moiety capable of binding to non-protein molecules such as tumor-associated glycolipids and carbohydrates.

[0183] In some embodiments, the cargo or payload can be or can encode a CAR comprising a targeting moiety capable of binding to components within the tumor microenvironment, including proteins expressed in various tumor stromal cells, such cells Including tumor-associated macrophages (TAM), immature monocytes, immature dendritic cells, immunosuppressive CD4+CD25+ regulatory T cells (Treg) and MDSC.

[0184] In some embodiments, the cargo or payload can be or can encode a CAR comprising a targeting moiety capable of binding to a cell surface adhesion molecule, a surface molecule of an inflammatory cell that occurs in an autoimmune disease, or a TCR that elicits autoimmunity. As a non-limiting example, the targeting moiety of the present invention can be a scFv antibody that recognizes a tumor-specific antigen (TSA), such as an antibody that specifically recognizes and binds to human mesothelin, the scFv of antibodies SS, SS1 and HN1, and GD2. scFv, CD19 antigen-binding domain, NKG2D ligand-binding domain, human anti-mesothelin scFv, anti-CS1 binder, anti-BCMA binding domain, anti-CD19 scFv antibody, GFRα4 antigen-binding fragment, anti-CLL-1 (C-type lectin Like molecule 1) binding domain, CD33 binding domain, GPC3 (glypican-3) binding domain, GFRα4 (glycosyl-phosphatidylinositol (GPI)-linked GDNF family α-receptor 4 cell surface receptor body) binding domain, CD123 binding domain, anti-ROR1 antibody or fragment thereof, scFv specific for GPC-3, scFv for CSPG4 and scFv for folate receptor alpha. [Intracellular signaling domain] []

[0185] After binding to its target molecule, the intracellular domain of the CAR fusion polypeptide transmits a signal to the immune effector cell, thereby activating at least one of the normal effector functions of the immune effector cell, including cytolytic activity (such as cytokine secretion) or accessory activity . Thus, the intracellular domain comprises the "intracellular signaling domain" of the T cell receptor (TCR).

[0186] In some aspects, an entire intracellular signaling domain can be employed. In other aspects, a truncated portion of the intracellular signaling domain can be used in place of the full chain, so long as it transduces effector function signals.

[0187] In some embodiments, the intracellular signaling domain may contain a signaling motif known as an immunoreceptor tyrosine-based activation motif (ITAM). Examples of ITAM-containing cytoplasmic signaling sequences include those derived from the TCRs CD3ζ, FcRγ, FcRβ, CD3γ, CD3δ, CD3ε, CD5, CD22, CD79a, CD79b, and CD66d. In one example, the intracellular signaling domain is the CD3ζ (CD3 zeta) signaling domain.

[0188] In some embodiments, the intracellular domain further comprises one or more co-stimulatory signaling domains that provide additional signals to immune effector cells. These co-stimulatory signaling domains combined with signaling domains can further improve the expansion, activation, memory, persistence and tumor eradication efficiency of CAR-engineered immune cells such as CAR T cells. In some cases, the co-stimulatory signaling region contains 1, 2, 3 or 4 cytoplasmic domains with one or more intracellular signaling and / or co-stimulatory molecules. Costimulatory signaling domains can be intracellular / cytoplasmic domains of costimulatory molecules, including but not limited to CD2, CD7, CD27, CD28, 4-1BB (CD137), OX40 (CD134), CD30, CD40, ICOS (CD278), GITR (glucocorticoid-induced tumor necrosis factor receptor), LFA-1 (lymphocyte function-associated antigen-1), LIGHT, NKG2C, B7-H3. In one example, the co-stimulatory signaling domain is derived from the cytoplasmic domain of CD28. In another example, the co-stimulatory signaling domain is derived from the cytoplasmic domain of 4-1BB (CD137). In another example, the co-stimulatory signaling domain may be the intracellular domain of GITR as taught in US Patent No. 9,175,308; the contents of which are incorporated herein by reference in their entirety.

[0189] In some embodiments, the intracellular region may comprise a functional signaling domain from a protein selected from the group consisting of MHC class I molecules, TNF receptor proteins, immunoglobulin-like proteins, interleukin receptors, integrins, Signaling lymphocyte activating proteins (SLAM), such as CD48, CD229, 2B4, CD84, NTB-A, CRACC, BLAME, CD2F-10, SLAMF6, SLAMF7, activating NK cell receptor, BTLA, Toll ligand receptor, OX40 , CD2, CD7, CD27, CD28, CD30, CD40, CDS, ICAM-1, LFA-1 (CD11a / CD18), 4-1BB (CD137), B7-H3, CDS, ICAM-1, ICOS (CD278), GITR, BAFFR, LIGHT, HVEM (LIGHTR), SLAMF7, NKp80 (KLRF1), NKp44, NKp30, NKp46, CD19, CD4, CD8α, CD8β, IL2Rβ, IL2Rγ, IL7Rα, IL-15Ra, ITGA4, VLA1, CD49a, ITGA4, IA4, CD49D, ITGA6, VLA-6, CD49f, ITGAD, CD11d, ITGAE, CD103, ITGAL, CD11a, LFA-1, ITGAM, CD11b, ITGAX, CD11c, ITGB1, CD29, ITGB2, CD18, LFA-1, ITGB7, NKG2D, NKG2C, NKD2C SLP76, TNFR2, TRANCE / RANKL, DNAM1 (CD226), SLAMF4 (CD244, 2B4), CD84, CD96 (Tactile), CEACAM1, CRTAM, Ly9 (CD229), CD160 (BY55), PSGL1, CD100 ( SEMA4D), CD69, SLAMF6 (NTB-A, Ly108), SLAM (SLAMF1, CD150, IPO-3), BLAME (SLAMF8), SELPLG (CD162), LTBR, ​​LAT, CD270 (HVEM), GADS, SLP-76, PAG / Cbp, CD19a, ligands specifically binding to CD83, DAP 10, TRIM, ZAP70, killer immunoglobulin receptor (KIR), such as KIR2DL1, KIR2DL2 / L3, KIR2DL4, KIR2DL5A, KIR2DL5B, KIR2DS1 , KIR2DS2, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DL1 / S1, KIR3DL2, KIR3DL3, and KIR2DP1; lectin-related NK cell receptors, such as Ly49, Ly49A, and Ly49C.

[0190] In some embodiments, the intracellular signaling domains of the present invention may contain signaling domains derived from JAK-STAT. In other embodiments, the intracellular signaling domain of the present invention may contain a signaling domain derived from DAP-12 (death-associated protein 12) (Topfer et al., Immunol., 2015, 194: 3201-3212; and Wang et al., Cancer Immunol., 2015, 3: 815-826). DAP-12 is a key signaling receptor in NK cells. Activation signals mediated by DAP-12 play an important role in triggering NK cell cytotoxic responses against certain tumor cells and virus-infected cells. The cytoplasmic domain of DAP12 contains an immunoreceptor tyrosine-based activation motif (ITAM). Therefore, CARs containing DAP12-derived signaling domains can be used for recipient transfer of NK cells. [transmembrane domain] []

[0191] In some embodiments, a CAR may comprise a transmembrane domain. As used herein, the term "transmembrane domain (TM)" broadly refers to an amino acid sequence of about 15 residues in length that spans the plasma membrane. The transmembrane domain may comprise at least 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44 or 45 amino acid residues and span the plasma membrane. In some embodiments, transmembrane domains can be derived from natural or synthetic sources. The transmembrane domain of the CAR can be derived from any native membrane-bound or transmembrane protein. For example, the transmembrane region may be derived from (i.e. at least comprise its transmembrane region) the following alpha, beta or zeta chains: T cell receptor, CD3ε, CD4, CD5, CD8, CD8α, CD9, CD16, CD22 , CD33, CD28, CD37, CD45, CD64, CD80, CD86, CD134, CD137, CD152, or CD154.

[0192] Alternatively, the transmembrane domains of the invention may be synthetic. In some aspects, the synthetic sequence may consist primarily of hydrophobic residues, such as leucine and valine.

[0193] In some embodiments, the transmembrane domain can be selected from the group consisting of: CD8α transmembrane domain, CD4 transmembrane domain, CD28 transmembrane domain, CTLA-4 transmembrane domain, PD-1 transmembrane domain, and human IgG4 Fc region .

[0194] In some embodiments, a CAR may comprise an optional hinge region (also known as a spacer). The hinge sequence is a short amino acid sequence that facilitates the flexibility of the extracellular targeting domain to move the target binding domain away from the effector cell surface to enable proper cell / cell contact, target binding, and effector cell activation. A hinge sequence can be located between the targeting moiety and the transmembrane domain. The hinge sequence may be any suitable sequence derived or obtained from any suitable molecule. The hinge sequence can be derived from all or part of the hinge region of an immunoglobulin (e.g. IgG1, IgG2, IgG3, IgG4), i.e. the sequence falls between the CH1 and CH2 domains of an immunoglobulin, e.g. IgG4 Fc hinge, type 1 membrane protein (such as CD8α CD4, CD28 and CD7), which may be wild-type sequences or derived sequences. Some hinge regions include an immunoglobulin CH3 domain or both CH3 and CH2 domains. In certain embodiments, the hinge region may be modified relative to IgG1, IgG2, IgG3, or IgG4 by including one or more amino acid residues substituted with amino acid residues different from those present in the unmodified hinge. amino acid residues, for example 1, 2, 3, 4 or 5 residues.

[0195] In some embodiments, a CAR may comprise one or more linkers between any domains of the CAR. Linkers can be 1-30 amino acids in length. In this regard, the length of the linker can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 amino acids. In other embodiments, the linker can be flexible.

[0196] In some embodiments, components including a targeting moiety, a transmembrane domain, and an intracellular signaling domain can be constructed in a single fusion polypeptide. Fusion polypeptides can be the payload of the effector modules of the invention.

[0197] In some embodiments, the cargo or payload can be or can encode CD19-specific CARs targeting different B-cell malignancies and HER2-specific CARs targeting sarcoma, glioblastoma, and advanced Her2-positive lung malignancies. Tandem CAR (TanCAR)

[0198] In some embodiments, the CAR may be a tandem chimeric antigen receptor (TanCAR) capable of targeting two, three, four or more tumor-specific antigens. In some aspects, the CAR is a bispecific TanCAR comprising two targeting domains that recognize two different TSAs on tumor cells. A bispecific TanCAR can be further defined as comprising an extracellular protein containing a targeting domain specific for a first tumor antigen (e.g., an antigen recognition domain) and a targeting domain specific for a second tumor antigen (e.g., an antigen recognition domain). district. In other aspects, the CAR is a multispecific TanCAR comprising three or more targeting domains configured in a tandem arrangement. The space between targeting domains in a TanCAR can be about 5 to about 30 amino acids in length, e.g., 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18 , 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 and 30 amino acids. [split] [CAR]

[0199] In some embodiments, the CAR component, including the targeting moiety, transmembrane domain, and intracellular signaling domain, can be split into two or more parts such that it relies on multiple components that facilitate assembly of a fully functional receptor. input. As a non-limiting example, a split CAR consists of two parts that assemble in a small molecule-dependent manner; one part of the receptor is characterized by an extracellular antigen-binding domain (e.g., scFv), and the other part has an intracellular signaling domain, Such as the CD3ζ intracellular domain.

[0200] In other aspects, split portions of the CAR system can be further modified to increase signaling. As a non-limiting example, the second portion of the cytoplasmic fragment can be anchored to the plasma membrane by incorporating a transmembrane domain (eg, CD8α transmembrane domain) into the construct. Additional ectodomains can also be added to the second part of the CAR system, such as an ectodomain that mediates homodimerization. These modifications increase receptor export activity, ie T cell activation.

[0201] In some embodiments, the two parts of the split CAR system contain heterodimerization domains that conditionally interact upon binding of the heterodimeric small molecule. Thus, the receptor components assemble in the presence of small molecules to form a complete system that can then be activated by antigen engagement. Any known heterodimerization component can be incorporated into the split CAR system. Other small molecule-dependent heterodimerization domains can also be used, including but not limited to gibberellin-induced dimerization system (GID1-GAI), trimethoprim-SLF-induced ecDHFR and FKBP dimerization and ABA (abscisic acid)-induced dimerization of PP2C and PYL domains. Dual regulation of inducible assembly (e.g., ligand-dependent dimerization) and degradation (e.g., domain destabilization-induced CAR degradation) using split CAR systems may provide more flexibility to control CAR-modified T cells activity. [switchable] [CAR] [ , , ]

[0202] In some embodiments, the CAR can be a switchable CAR that can be temporarily switched on as a controllable CAR in response to a stimulus, such as a small molecule. In this CAR design, the system is integrated directly into the hinge domain that separates the scFv domain from the membrane domain in the CAR. Such systems have the potential to split or combine different key functions of CARs, such as activation and co-stimulation within different chains of the receptor complex, thereby mimicking the complexity of the native TCR architecture. This integrated system can switch scFv and antigen interactions between on / off states controlled by absence / presence of stimuli. [reversible] [CAR] [ , , ]

[0203] In some embodiments, the CAR can be a reversible CAR system. In this CAR architecture, a LID (ligand-induced degradation) domain is incorporated into the CAR system. CARs can be temporarily downregulated by ligands that add LID domains. [inhibitory] [CAR (iCAR)] []

[0204] In some embodiments, the CAR can be an inhibitory CAR. Inhibitory CAR (iCAR) refers to a bispecific CAR design in which a negative signal is used to enhance tumor specificity and limit normal tissue toxicity. This design incorporates a second CAR with a surface antigen recognition domain as well as an inhibitory signaling domain to limit T cell reactivity even while simultaneously engaging an activating receptor. This antigen recognition domain targets normal tissue-specific antigens so that T cells can be activated in the presence of the first target protein, but the presence of a second protein that binds to iCAR inhibits T cell activation.

[0205] As a non-limiting example, a CTLA4 and PD1 inhibitory domain-based iCAR against prostate-specific membrane antigen (PMSA) exhibits the ability to selectively limit cytokine secretion, cytotoxicity, and proliferation induced by T cell activation. [Chimeric Switch Receptor] []

[0206] In some embodiments, the cargo or payload can be or encode a chimeric switch receptor that can switch a negative signal to a positive signal. As used herein, the term "chimeric switch receptor" refers to a fusion protein comprising a first extracellular domain and a second transmembrane and intracellular domain, wherein the first domain comprises a negative signaling region and the second domain comprises a positive intracellular signaling domain. district. In some aspects, the fusion protein is a chimeric switch receptor containing the ectodomain of an inhibitory receptor on T cells fused to the transmembrane and cytoplasmic domains of a co-stimulatory receptor. The chimeric switch receptor converts T cell inhibitory signals into T cell stimulatory signals.

[0207] As a non-limiting example, a chimeric switch receptor can comprise the extracellular domain of PD-1 fused to the transmembrane and cytoplasmic domains of CD28. In some aspects, the ectodomains of other inhibitory receptors, such as CTLA-4, LAG-3, TIM-3, KIR, and BTLA, can also be combined with those derived from costimulatory receptors (such as CD28, 4-1BB, CD27, OX40, CD40, GTIR, and ICOS) fusion of the transmembrane and cytoplasmic domains.

[0208] In some embodiments, chimeric switch receptors may include recombinant receptors comprising: and stimulatory interleukin receptors such as IL-2R (IL-2Rα, IL-2Rβ, and IL-2Rγ) and IL-7Rα The extracellular interleukin-binding domain of an inhibitory interleukin receptor fused to the intracellular signaling domain of ), such as IL-13 receptor α (IL-13Rα1 ), IL-10R, and IL-4Rα. One example of such a chimeric interleukin receptor is a recombinant receptor containing the interleukin-binding extracellular domain of IL-4Rα linked to the intracellular signaling domain of IL-7Rα.

[0209] In some embodiments, the chimeric switch receptor can be a chimeric TGFβ receptor. A chimeric TGFβ receptor may comprise: an ectodomain derived from a TGFβ receptor, such as TGFβ receptor 1, TGFβ receptor 2, TGFβ receptor 3, or any other TGFβ receptor or variant thereof; and a non-TGFβ receptor intracellular area. The non-TGF beta receptor intracellular domain may be an intracellular domain or fragment thereof derived from: TLR1, TLR2, TLR3, TLR4, TLR5, TLR6, TLR7, TLR8, TLR9, TLR10, CD28, 4-1BB (CD137), OX40 (CD134), CD3ζ, CD40, CD27 or combinations thereof. [activation conditionality] [CAR] [ , , ]

[0210] In some embodiments, the cargo or payload can be or can encode an activating conditional chimeric antigen receptor that is expressed only in activated immune cells. Expression of the CAR can be coupled to an activation conditional control region that refers to a nucleic acid sequence under its control for the transcription and / or expression of one or more inducible sequences (eg, CAR). Such activation conditional control regions may be the promoters of genes that are upregulated during activation of immune effector cells, such as the IL2 promoter or NFAT binding sites. [Targeting tumor cells with specific proteoglycan markers] [CAR] [ , , ]

[0211] In some embodiments, the cargo or payload can be or can encode a CAR that targets a particular type of cancer cell. Human cancer cells and metastases can express unique and otherwise abnormal proteoglycans, such as polysaccharide chains (e.g., chondroitin sulfate (CS), dermatan sulfate (DS or CSB), heparan sulfate (heparan sulfate). sulfate; HS) and heparin). Thus, the CAR can be fused to a binding moiety that recognizes cancer-associated proteoglycans. In one example, a CAR can be fused to a VAR2CSA polypeptide that binds with high affinity to a specific type of chondroitin sulfate A (CSA) linked to proteoglycans (VAR2-CAR). The extracellular ScFv portion of the CAR can be substituted with a VAR2CSA variant containing at least a minimal CSA binding domain, resulting in a CAR specific for chondroitin sulfate A (CSA) modification. Alternatively, the CAR can be fused to a split protein binding system to generate a spy-CAR, wherein the scFv portion of the CAR is replaced by a portion of the split protein binding system (such as SpyTag and Spy Trap), and cancer recognition molecules (such as scFv and / or VAR2-CSA) is linked to CAR via a split protein binding system. nucleic acid

[0212] The initial and reference constructs of the invention may comprise a payload region (which may also be referred to as a cargo region), which is a nucleic acid. The term "nucleic acid" in its broadest sense includes any compound and / or substance comprising a polymer of nucleotides which may be referred to as a polynucleotide. Exemplary nucleic acids or polynucleotides include, but are not limited to, ribonucleic acid (RNA), deoxyribonucleic acid (DNA), threose nucleic acid (TNA), diol nucleic acid (GNA), peptide nucleic acid (PNA), locked nucleic acid (LNA) ) or a mixture thereof.

[0213] In some embodiments, the payload region comprises nucleic acid sequences encoding more than one cargo or payload.

[0214] In some embodiments, the payload region can be or encodes a coding nucleic acid sequence.

[0215] In some embodiments, the payload region can be or encode a non-coding nucleic acid sequence.

[0216] In some embodiments, the payload region can be or encode both coding and non-coding nucleic acid sequences. [DNA]

[0217] Deoxyribonucleic acid (DNA) is a molecule that carries the genetic information of all living things and consists of two strands wound around each other to form a shape called a double helix. Each strand has a backbone made up of alternating sugar (deoxyribose) and phosphate groups. Each sugar is linked to one of four bases: adenine (A), cytosine (C), guanine (G) and thymine (T). The two strands are held together by bonds between adenine and thymine or between cytosine and guanine. The sequence of bases along the backbone serves as instructions for assembling protein and RNA molecules.

[0218] In some embodiments, the payload region can be or encodes coding DNA.

[0219] In some embodiments, the payload region can be or encodes non-coding DNA.

[0220] In some embodiments, the payload region can be or encode both coding and non-coding DNA.

[0221] In some embodiments, DNA can be modified. Types of modifications include, but are not limited to, methylation, acetylation, phosphorylation, ubiquitination, and sumylation. carrier

[0222] In some embodiments, the initial constructs and / or reference constructs described herein can be or be encoded by vectors such as plastid or viral vectors. In some embodiments, the initial and / or reference constructs are or are encoded by viral vectors. Viral vectors may be, but are not limited to, herpes virus (HSV) vectors, retroviral vectors, adenoviral vectors, adeno-associated virus (AAV) vectors, lentiviral vectors, and the like. In some embodiments, the viral vector is an AAV vector. In some embodiments, the viral vector is a lentiviral vector. In some embodiments, the viral vector is a retroviral vector. In some embodiments, the viral vector is an adenoviral vector. [adeno-associated virus] [(AAV)] [carrier] []

[0223] Viruses of the Parvoviridae family are non-enveloped icosahedral small capsid viruses characterized by a single-stranded DNA genome. The Parvoviridae viruses consist of two subfamilies: the Parvovirinae, which infects vertebrates, and the Densovirinae, which infects invertebrates. Due to its relatively simple structure and ease of manipulation using standard molecular biology techniques, this family of viruses is suitable as a biological tool. The genome of the virus can be modified to contain the minimum components for assembly of a functional recombinant virus or virion loaded with or engineered to express or deliver the desired payload, the payload Can be delivered to target cells, tissues, organs or organisms.

[0224] The Parvoviridae family includes the Dependovirus genus, which includes adeno-associated viruses (AAV) capable of replicating in vertebrate hosts including, but not limited to, humans, primates, bovine, Canine, equine and sheep species.

[0225] AAV vector genomes are linear, single-stranded DNA (ssDNA) molecules approximately 5,000 nucleotides (nt) in length. The AAV vector gene body may comprise a payload region and at least one inverted terminal repeat (ITR) or ITR region. ITRs are traditionally flanked by nucleotide sequences encoding nonstructural proteins (encoded by the Rep gene) and structural proteins (encoded by the capsid gene or Cap gene). While not wishing to be bound by theory, AAV vector gene bodies typically contain two ITR sequences. The AAV vector gene body contains a characteristic T-shaped hairpin structure, which is limited by 145 nucleotides at the 5' end and 3' end of the ssDNA, which form an energetically stable double-stranded region. The double-stranded hairpin structure serves multiple functions, including but not limited to serving as an origin of DNA replication by serving as a primer for the endogenous DNA polymerase complex of the host virus replicating cell.

[0226] In addition to the encoded heterologous payload, the AAV vector gene body may comprise all or part of any naturally occurring and / or recombinant AAV serotype nucleotide sequence or variant. AAV variants may have significant homologous sequences at the nucleic acid (genome or capsid) and amino acid level (capsid) to generate substantially physically and functionally equivalent constructs that are identified by Similar mechanisms replicate and are assembled by similar mechanisms. Chiorini et al., J. Vir. 71: 6823-33 (1997); Srivastava et al., J. Vir. 45:555-64 (1983); Chiorini et al., J. Vir. 73:1309-1319 (1999) ; Rutledge et al., J. Vir. 72:309-319 (1998); and Wu et al., J. Vir. 74: 8635-47 (2000), the contents of each of which are incorporated herein by reference in their entirety.

[0227] In some embodiments, the AAV vector genome comprises at least one control element that provides for the replication, transcription and translation of the coding sequence encoded therein. Not all control elements need to be present at all times, as long as the coding sequence can be replicated, transcribed and / or translated in a suitable host cell. Non-limiting examples of expression control elements include sequences for transcription initiation and / or termination, promoter and / or enhancer sequences, efficient RNA processing signals (such as splicing and polyadenylation signals), factors that stabilize cytoplasmic mRNA, sequences, sequences that enhance translational efficacy (eg, Kozak consensus sequences), sequences that enhance protein stability, and / or sequences that enhance protein processing and / or secretion.

[0228] The AAV vector gene body of the present invention can be produced in a recombinant way, and can be based on adeno-associated virus (AAV) parent or reference sequence. As used herein, a "vector gene body" is any molecule or portion that transports, transduces, or otherwise acts as a vehicle for heterologous molecules such as the nucleic acids described herein.

[0229] In addition to single-stranded AAV vector gene bodies (eg, ssAAV), the present invention also provides self-complementary AAV (scAAV) vector gene bodies. The scAAV vector gene body contains DNA strands that anneal together to form double-stranded DNA. By skipping the second-strand synthesis, scAAV achieves rapid expression in cells.

[0230] In some embodiments, the AAV vector gene body is scAAV.

[0231] In some embodiments, the AAV vector gene body is ssAAV.

[0232] In some embodiments, the AAV vector gene body can be part of an AAV particle, wherein the serotype of the capsid can be, but not limited to, AAV1, AAV2, AAV2G9, AAV3, AAV3a, AAV3b, AAV3-3, AAV4, AAV4-4, AAV5, AAV6, AAV6.1, AAV6.2, AAV6.1.2, AAV7, AAV7.2, AAV8, AAV9, AAV9.11, AAV9.13, AAV9.16, AAV9.24, AAV9.45, AAV9.47, AAV9.61, AAV9.68, AAV9.84, AAV9.9, AAV10, AAV11, AAV12, AAV16.3, AAV24.1, AAV27.3, AAV42.12, AAV42-1b, AAV42-2, AAV42-3a, AAV42-3b, AAV42-4, AAV42-5a, AAV42-5b, AAV42-6b, AAV42-8, AAV42-10, AAV42-11, AAV42-12, AAV42-13, AAV42-15, AAV42-aa, AAV43- 1. AAV43-12, AAV43-20, AAV43-21, AAV43-23, AAV43-25, AAV43-5, AAV44.1, AAV44.2, AAV44.5, AAV223.1, AAV223.2, AAV223.4, AAV223.5, AAV223.6, AAV223.7, AAV1-7 / rh.48, AAV1-8 / rh.49, AAV2-15 / rh.62, AAV2-3 / rh.61, AAV2-4 / rh. 50. AAV2-5 / rh.51, AAV3.1 / hu.6, AAV3.1 / hu.9, AAV3-9 / rh.52, AAV3-11 / rh.53, AAV4-8 / r11.64, AAV4-9 / rh.54, AAV4-19 / rh.55, AAV5-3 / rh.57, AAV5-22 / rh.58, AAV7.3 / hu.7, AAV16.8 / hu.10, AAV16. 12 / hu.11, AAV29.3 / bb.1, AAV29.5 / bb.2, AAV106.1 / hu.37, AAV114.3 / hu.40, AAV127.2 / hu.41, AAV127.5 / hu.42, AAV128.3 / hu.44, AAV130.4 / hu.48, AAV145.1 / hu.53, AAV145.5 / hu.54, AAV145.6 / hu.55, AAV161.10 / hu. 60. AAV161.6 / hu.61, AAV33.12 / hu.17, AAV33.4 / hu.15, AAV33.8 / hu.16, AAV52 / hu.19, AAV52.1 / hu.20, AAV58. 2 / hu.25, AAVA3.3, AAVA3.4, AAVA3.5, AAVA3.7, AAVC1, AAVC2, AAVC5, AAV-DJ, AAV-DJ8, AAVF3, AAVF5, AAVH2, AAVrh.72, AAVhu.8, AAVrh.68, AAVrh.70, AAVpi.1, AAVpi.3, AAVpi.2, AAVrh.60, AAVrh.44, AAVrh.65, AAVrh.55, AAVrh.47, AAVrh.69, AAVrh.45, AAVrh. 59. AAVhu.12, AAVH6, AAVLK03, AAVH-1 / hu.1, AAVH-5 / hu.3, AAVLG-10 / rh.40, AAVLG-4 / rh.38, AAVLG-9 / hu.39, AAVN721-8 / rh.43, AAVCh.5, AAVCh.5R1, AAVcy.2, AAVcy.3, AAVcy.4, AAVcy.5, AAVCy.5R1, AAVCy.5R2, AAVCy.5R3, AAVCy.5R4, AAVcy. 6. AAVhu.1, AAVhu.2, AAVhu.3, AAVhu.4, AAVhu.5, AAVhu.6, AAVhu.7, AAVhu.9, AAVhu.10, AAVhu.11, AAVhu.13, AAVhu.15, AAVhu.16, AAVhu.17, AAVhu.18, AAVhu.20, AAVhu.21, AAVhu.22, AAVhu.23.2, AAVhu.24, AAVhu.25, AAVhu.27, AAVhu.28, AAVhu.29, AAVhu. 29R, AAVhu.31, AAVhu.32, AAVhu.34, AAVhu.35, AAVhu.37, AAVhu.39, AAVhu.40, AAVhu.41, AAVhu.42, AAVhu.43, AAVhu.44, AAVhu.44R1, AAVhu.44R2, AAVhu.44R3, AAVhu.45, AAVhu.46, AAVhu.47, AAVhu.48, AAVhu.48R1, AAVhu.48R2, AAVhu.48R3, AAVhu.49, AAVhu.51, AAVhu.52, AAVhu. 54, AAVhu.55, AAVhu.56, AAVhu.57, AAVhu.58, AAVhu.60, AAVhu.61, AAVhu.63, AAVhu.64, AAVhu.66, AAVhu.67, AAVhu.14 / 9, AAVhu. t 19, AAVrh.2, AAVrh.2R, AAVrh.8, AAVrh.8R, AAVrh.10, AAVrh.12, AAVrh.13, AAVrh.13R, AAVrh.14, AAVrh.17, AAVrh.18, AAVrh.19 , AAVrh.20, AAVrh.21, AAVrh.22, AAVrh.23, AAVrh.24, AAVrh.25, AAVrh.31, AAVrh.32, AAVrh.33, AAVrh.34, AAVrh.35, AAVrh.36, AAVrh .37, AAVrh.37R2, AAVrh.38, AAVrh.39, AAVrh.40, AAVrh.46, AAVrh.48, AAVrh.48.1, AAVrh.48.1.2, AAVrh.48.2, AAVrh.49, AAVrh.51, AAVrh .52, AAVrh.53, AAVrh.54, AAVrh.56, AAVrh.57, AAVrh.58, AAVrh.61, AAVrh.64, AAVrh.64R1, AAVrh.64R2, AAVrh.67, AAVrh.73, AAVrh.74 , AAVrh8R, AAVrh8R A586R mutant, AAVrh8R R533A mutant, AAAV, BAAV, goat AAV, bovine AAV, AAVhE1.1, AAVhEr1.5, AAVhER1.14, AAVhEr1.8, AAVhEr1.16, AAVhEr1.18, AAVhEr1.35 , AAVhEr1.7, AAVhEr1.36, AAVhEr2.29, AAVhEr2.4, AAVhEr2.16, AAVhEr2.30, AAVhEr2.31, AAVhEr2.36, AAVhER1.23, AAVhEr3.1, AAV2.5T, AAV-PAEC, AAV -LK01, AAV-LK02, AAV-LK03, AAV-LK04, AAV-LK05, AAV-LK06, AAV-LK07, AAV-LK08, AAV-LK09, AAV-LK10, AAV-LK11, AAV-LK12, AAV-LK13 , AAV-LK14, AAV-LK15, AAV-LK16, AAV-LK17, AAV-LK18, AAV-LK19, AAV-PAEC2, AAV-PAEC4, AAV-PAEC6, AAV-PAEC7, AAV-PAEC8, AAV-PAEC11, AAV -PAEC12, AAV-2-pre-miRNA-101, AAV-8h, AAV-8b, AAV-h, AAV-b, AAV SM 10-2, AAV Shuffle 100-1, AAV Shuffle 100- 3. AAV mix 100-7, AAV mix 10-2, AAV mix 10-6, AAV mix 10-8, AAV mix 100-2, AAV SM 10-1, AAV SM 10-8, AAV SM 100-3, AAV SM 100-10, BNP61 AAV, BNP62 AAV, BNP63 AAV, AAVrh.50, AAVrh.43, AAVrh.62, AAVrh.48, AAVhu.19, AAVhu.11, AAVhu.53, AAV4- 8 / rh.64, AAVLG-9 / hu.39, AAV54.5 / hu.23, AAV54.2 / hu.22, AAV54.7 / hu.24, AAV54.1 / hu.21, AAV54.4R / hu.27, AAV46.2 / hu.28, AAV46.6 / hu.29, AAV128.1 / hu.43, true AAV (ttAAV), UPENN AAV 10, Japanese AAV 10 serotype, AAV CBr-7.1, AAV CBr-7.10, AAV CBr-7.2, AAV CBr-7.3, AAV CBr-7.4, AAV CBr-7.5, AAV CBr-7.7, AAV CBr-7.8, AAV CBr-B7.3, AAV CBr-B7.4, AAV CBr-E1, AAV CBr-E2, AAV CBr-E3, AAV CBr-E4, AAV CBr-E5, AAV CBr-e5, AAV CBr-E6, AAV CBr-E7, AAV CBr-E8, AAV CHt-1, AAV CHt-2, AAV CHt-3, AAV CHt-6.1, AAV CHt-6.10, AAV CHt-6.5, AAV CHt-6.6, AAV CHt-6.7, AAV CHt-6.8. AAV CHt-P1, AAV CHt-P2, AAV CHt-P5, AAV CHt-P6, AAV CHt-P8, AAV CHt-P9, AAV CKd-1, AAV CKd-10, AAV CKd-2, AAV CKd- 3. AAV CKd-4, AAV CKd-6, AAV CKd-7, AAV CKd-8, AAV CKd-B1, AAV CKd-B2, AAV CKd-B3, AAV CKd-B4, AAV CKd-B5, AAV CKd- B6, AAV CKd-B7, AAV CKd-B8, AAV CKd-H1, AAV CKd-H2, AAV CKd-H3, AAV CKd-H4, AAV CKd-H5, AAV CKd-H6, AAV CKd-N3, AAV CKd- N4, AAV CKd-N9, AAV CLg-F1, AAV CLg-F2, AAV CLg-F3, AAV CLg-F4, AAV CLg-F5, AAV CLg-F6, AAV CLg-F7, AAV CLg-F8, AAV CLv- 1. AAV CLv1-1, AAV Clv1-10, AAV CLv1-2, AAV CLv-12, AAV CLv1-3, AAV CLv-13, AAV CLv1-4, AAV Clv1-7, AAV Clv1-8, AAV Clv1- 9. AAV CLv-2, AAV CLv-3, AAV CLv-4, AAV CLv-6, AAV CLv-8, AAV CLv-D1, AAV CLv-D2, AAV CLv-D3, AAV CLv-D4, AAV CLv- D5, AAV CLv-D6, AAV CLv-D7, AAV CLv-D8, AAV CLv-E1, AAV CLv-K1, AAV CLv-K3, AAV CLv-K6, AAV CLv-L4, AAV CLv-L5, AAV CLv- L6, AAV CLv-M1, AAV CLv-M11, AAV CLv-M2, AAV CLv-M5, AAV CLv-M6, AAV CLv-M7, AAV CLv-M8, AAV CLv-M9, AAV CLv-R1, AAV CLv- R2, AAV CLv-R3, AAV CLv-R4, AAV CLv-R5, AAV CLv-R6, AAV CLv-R7, AAV CLv-R8, AAV CLv-R9, AAV CSp-1, AAV CSp-10, AAV CSp- 11. AAV CSp-2, AAV CSp-3, AAV CSp-4, AAV CSp-6, AAV CSp-7, AAV CSp-8, AAV CSp-8.10, AAV CSp-8.2, AAV CSp-8.4, AAV CSp- 8.5, AAV CSp-8.6, AAV CSp-8.7, AAV CSp-8.8, AAV CSp-8.9, AAV CSp-9, AAV.hu.48R3, AAV.VR-355, AAV3B, AAV4, AAV5, AAVF1 / HSC1, AAVF11 / HSC11, AAVF12 / HSC12, AAVF13 / HSC13, AAVF14 / HSC14, AAVF15 / HSC15, AAVF16 / HSC16, AAVF17 / HSC17, AAVF2 / HSC2, AAVF3 / HSC3, AAVF4 / HSC4, AAVF5 / HSC5, AAVF6 / HSC6, AAVF7 / HSC7 , AAVF8 / HSC8, AAVF9 / HSC9, PHP.B, PHP.A, G2B-26, G2B-13, TH1.1-32 and / or TH1.1-35 and variants thereof. .

[0233] inverted terminal repeat (ITR)

[0234] In some embodiments, an AAV vector gene body may comprise at least one ITR region and a payload region. In some embodiments, the vector gene body has two ITRs. The two ITRs flank the payload area at the 5' and 3' ends. The ITR acts as an origin of replication, containing the recognition site for replication. The ITRs comprise sequence regions that may be complementary and symmetrically arranged. The ITR incorporated into the vector gene body of the present invention may consist of a naturally occurring polynucleotide sequence or a recombinantly derived polynucleotide sequence.

[0235] ITRs may be derived from the same serotype as the capsid or derivatives thereof. ITRs may be of a different serotype than capsids. In some embodiments, the AAV particle has more than one ITR. In a non-limiting example, an AAV particle has a vector gene body comprising two ITRs. In some embodiments, the ITRs are of the same serotype as each other. In another embodiment, the ITRs are of different serotypes. Non-limiting examples include zero, one or two ITRs with the same serotype as capsid. In some embodiments, both ITRs of the vector gene body of the AAV particle are AAV2 ITRs.

[0236] Independently, each ITR can be about 100 to about 150 nucleotides in length. The ITR can be about 100-105 nucleotides in length, 106-110 nucleotides in length, 111-115 nucleotides in length, 116-120 nucleotides in length, 121-125 nucleotides in length Nucleotides, 126-130 nucleotides in length, 131-135 nucleotides in length, 136-140 nucleotides in length, 141-145 nucleotides in length, or 146-150 nucleotides in length Nucleotides. In some embodiments, the ITR is 140-142 nucleotides in length. Non-limiting examples of ITR lengths are 102, 140, 141, 142, 145 nucleotides in length, and those with at least 95% identity thereto. [Promoter] []

[0237] In some embodiments, the payload region of the vector gene body comprises at least one element that enhances transgene target specificity and expression (see, e.g., Powell et al., Viral Expression Cassette Elements to Enhance Transgene Target Specificity and Expression in Gene Therapy, 2015; The content of which is incorporated herein by reference in its entirety). Non-limiting examples of elements that enhance target specificity and expression of transgenes include promoters, endogenous miRNAs, post-transcriptional regulatory elements (PREs), polyadenylation (PolyA) signal sequences, and upstream enhancers (USEs), CMV enhancers and introns.

[0238] In some embodiments, a promoter is considered effective when it drives the expression of a polypeptide encoded in the payload region of the vector gene body of an AAV particle.

[0239] In some embodiments, a promoter is considered effective when it drives expression in the targeted cell.

[0240] In some embodiments, the promoter drives the expression of the payload in the targeted tissue for a period of time. Expression driven by the promoter can last for 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 23 hours, 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 8 days , 9 days, 10 days, 11 days, 12 days, 13 days, 2 weeks, 15 days, 16 days, 17 days, 18 days, 19 days, 20 days, 3 weeks, 22 days, 23 days, 24 days, 25 days days, 26 days, 27 days, 28 days, 29 days, 30 days, 31 days, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months month, 9 months, 10 months, 11 months, 1 year, 13 months, 14 months, 15 months, 16 months, 17 months, 18 months, 19 months, 20 months, 21 months months, 22 months, 23 months, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, 10 years or more than 10 years. Performance lasts 1 to 5 hours, 1 to 12 hours, 1 to 2 days, 1 to 5 days, 1 to 2 weeks, 1 to 3 weeks, 1 to 4 weeks, 1 to 2 months, 1 to 4 months, 1 to 6 months, 2 to 6 months, 3 to 6 months, 3 to 9 months, 4 to 8 months, 6 to 12 months, 1 to 2 years, 1 to 5 years, 2 to 5 years , 3 to 6 years, 3 to 8 years, 4 to 8 years or 5 to 10 years.

[0241] In some embodiments, the expression of the promoter-driven payload persists for at least 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months , 10 months, 11 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, 10 years, 11 years, 12 years, 13 years, 14 years , 15 years, 16 years, 17 years, 18 years, 19 years, 20 years, 21 years, 22 years, 23 years, 24 years, 25 years, 26 years, 27 years, 28 years, 29 years, 30 years, 31 years Years, 32 years, 33 years, 34 years, 35 years, 36 years, 37 years, 38 years, 39 years, 40 years, 41 years, 42 years, 43 years, 44 years, 45 years, 46 years, 47 years, 48 years, 49 years, 50 years, 55 years, 60 years, 65 years or more than 65 years.

[0242] Promoters may be naturally occurring or non-naturally occurring. Non-limiting examples of promoters include viral promoters, plant promoters, and mammalian promoters. In some embodiments, the promoter can be a human promoter. In some embodiments, a promoter can be truncated.

[0243] Promoters that drive or promote expression in most tissues include, but are not limited to, human elongation factor 1α-subunit (EF1α), cytomegalovirus (CMV) immediate early enhancer and / or promoter, chicken β-actin ( CBA) and its derivatives CAG, β-glucuronidase (GUSB) or ubiquitin C (UBC). Tissue-specific expression elements can be used to restrict the expression of certain cell types, such as but not limited to muscle-specific promoters, B-cell promoters, monocyte promoters, leukocyte promoters, macrophage promoters, pancreatic acinar Cellular promoters, endothelial cell promoters, lung tissue promoters, astrocyte promoters, or nervous system promoters that can be used to limit neuronal, astrocyte, or oligodendrocyte expression.

[0244] Non-limiting examples of muscle-specific promoters include the mammalian muscle creatine kinase (MCK) promoter, the mammalian desmin (DES) promoter, the mammalian troponin I (TNNI2) promoter, and the mammalian skeletal Alpha-actin (ASKA) promoter (see, eg, US Patent Publication US20110212529, the contents of which are hereby incorporated by reference in their entirety).

[0245] Non-limiting examples of tissue-specific expression elements of neurons include neuron-specific enolase (NSE), platelet-derived growth factor (PDGF), platelet-derived growth factor B-chain (PDGF-beta), synapsin (Syn ), methyl-CpG binding protein 2 (MeCP2), Ca 2+ / calmodulin-dependent protein kinase II (CaMKII), metabotropic glutamate receptor 2 (mGluR2), neurofilament light chain (NFL) or heavy chain (NFH), β-hemoglobulin minigene nβ2, proenkephalin (PPE), enkephalin (Enk) and excitatory amino acid transporter 2 (EAAT2) promoters. Non-limiting examples of tissue-specific expression elements for astrocytes include glial fibrillary acidic protein (GFAP) and the EAAT2 promoter. A non-limiting example of a tissue-specific expression element for oligodendrocytes includes the myelin basic protein (MBP) promoter.

[0246] In some embodiments, a promoter can be less than 1 kb. The length of the promoter can be 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800 or more than 800 nucleotides. The length of the promoter can be 200-300, 200-400, 200-500, 200-600, 200-700, 200-800, 300-400, 300-500, 300-600, 300-700, 300-800, Between 400-500, 400-600, 400-700, 400-800, 500-600, 500-700, 500-800, 600-700, 600-800, or 700-800.

[0247] In some embodiments, the promoter can be a combination of two or more components of the same or different starting or parental promoters such as but not limited to CMV and CBA. The length of each component can be 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 381, 382, ​​383 ,384,385,386,387,388,389,390,400,410,420,430,440,450,460,470,480,490,500,510,520,530,540,550,560,570 ,580,590,600,610,620,630,640,650,660,670,680,690,700,710,720,730,740,750,760,770,780,790,800 or more than 800. The length of each component can be 200-300, 200-400, 200-500, 200-600, 200-700, 200-800, 300-400, 300-500, 300-600, 300-700, 300-800, Between 400-500, 400-600, 400-700, 400-800, 500-600, 500-700, 500-800, 600-700, 600-800, or 700-800. In some embodiments, the promoter is a combination of a 382 nucleotide CMV-enhancer sequence and a 260 nucleotide CBA-promoter sequence.

[0248] In some embodiments, the vector gene body comprises a ubiquitous promoter. Non-limiting examples of ubiquitous promoters include CMV, CBA (including derivatives CAG, CBh, etc.), EF-la, PGK, UBC, GUSB (hGBp), and UCOE (promoter for HNRPA2B1-CBX3).

[0249] In some embodiments, the promoter is not cell-specific.

[0250] In some embodiments, the vector gene body comprises an engineered promoter.

[0251] In some embodiments, the vector gene body comprises a promoter from a naturally expressed protein. [Non-translation area] [(] [UTR] [)] []

[0252] By definition, the wild-type untranslated region (UTR) of a gene is transcribed, but not translated. Typically, the 5'UTR starts at the transcription start site and ends at the start codon, and the 3'UTR starts immediately after the stop codon and continues until the transcription termination signal.

[0253] Features commonly found in genes that are well expressed in a particular target organ can be engineered into UTRs to enhance stability and protein production. As non-limiting examples, mRNAs from mRNAs normally expressed in the liver (e.g., albumin, serum amyloid A, lipoprotein A / B / E, transferrin, alpha-fetoprotein, erythropoietin, or factor VIII) The 5'UTR can be used in the vector gene body of the AAV particle of the present invention to enhance the expression in the liver cell line or liver.

[0254] While not wishing to be bound by theory, the wild-type 5' untranslated region (UTR) includes features that play a role in translation initiation. Kozak sequences are generally known to be involved in the process by the ribosome for initiating translation of various genes, usually included in the 5' UTR. Kozak sequences have a common CCR(A / G)CCAUGG, where R is a purine (adenine or guanine) three bases upstream of the start codon (ATG), followed by another 'G'.

[0255] In some embodiments, the 5'UTR in the vector gene body includes a Kozak sequence.

[0256] In some embodiments, the 5'UTR in the vector gene body does not include a Kozak sequence.

[0257] While not wishing to be bound by theory, it is known that there are stretches of adenosine and uridine embedded in the wild-type 3'UTR. Such AU-rich markers are especially prevalent in genes with high turnover rates. Based on their sequence characteristics and functional properties, AU-rich elements (AREs) can be divided into three classes (Chen et al., 1995, the contents of which are incorporated herein by reference in their entirety): Class I AREs, such as but not limited to c-Myc and MyoD, contains several scattered copies of the AUUUA motif within the U-rich region. Class II AREs, such as but not limited to GM-CSF and TNF-a, have two or more overlapping UUAUUUA(U / A)(U / A) nonamers. Class III AREs, such as but not limited to c-Jun and myogenin, are less well defined. These U-rich regions do not contain the AUUUA motif. Most proteins bound to AREs are known to destabilize the messenger, while members of the ELAV family (most notably, HuR) have been documented to increase the stability of mRNA. HuR binds to all three classes of AREs. Engineering a HuR-specific binding site into the 3'UTR of a nucleic acid molecule will result in HuR binding and thus message stabilization in vivo.

[0258] The introduction, removal or modification of 3'UTR AU-rich elements (AREs) can be used to modulate the stability of polynucleotides. When engineering a particular polynucleotide (e.g., the payload region of a vector gene body), one or more ARE copies can be introduced to reduce polynucleotide stability and thereby reduce translation of the resulting protein and reduce its output. Likewise, AREs can be identified and removed or mutated to increase intracellular stability and thus increase translation and production of the resulting protein.

[0259] In some embodiments, the 3' UTR of the vector gene body may include an oligo(dT) sequence for templated addition of a poly A tail.

[0260] In some embodiments, a vector gene body may include at least one miRNA seed, binding site, or full sequence. MicroRNAs (or miRNAs or miRs) are 19-25 nucleotide non-coding RNAs that bind to sites of nucleic acid targets and downregulate gene expression by reducing nucleic acid molecule stability or by inhibiting translation. The microRNA sequence comprises a "seed" region, i.e., a sequence in the region of positions 2-8 of the mature microRNA, which has perfect Watson-Crick complementarity to the miRNA target sequence of the nucleic acid .

[0261] In some embodiments, the vector genome can be engineered to include, alter or remove at least one miRNA binding site, sequence or seed region.

[0262] Any UTR from any gene known in the art can be incorporated into the vector gene body of the AAV particle. The UTRs or parts thereof may be placed in the same orientation as in the gene from which the UTRs are selected, or their orientation or position may vary. In some embodiments, the UTRs in the vector gene bodies used in AAV particles can be inverted, shortened, elongated or made with one or more other 5'UTRs or 3'UTRs known in the art. As used herein, the term "altered" in relation to a UTR means that the UTR has been altered in some way relative to a reference sequence. For example, the 3' or 5' UTR can be altered relative to the wild-type or native UTR by changes in orientation or position as taught above, or by inclusion of additional nucleotides, nucleotide deletions, nucleoside Altered by acid exchange or transposition.

[0263] In some embodiments, the vector gene body of the AAV particle comprises at least one artificial UTR that is not a variant of a wild-type UTR.

[0264] In some embodiments, the vector genome of the AAV particle comprises UTRs selected from a family of transcripts in which proteins share common functions, structures, characteristics or properties. [polyadenylation sequence] []

[0265] In some embodiments, the vector gene body comprises at least one polyadenylation sequence between the 3' end of the payload coding sequence and the 5' end of the 3'ITR.

[0266] In some embodiments, the polyadenylation (poly A) sequence can range in length from nonexistent to about 500 nucleotides. The length of the polyadenylation sequence can be, but is not limited to, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 372, 373, 374, 375, 376, 377, 378, 379, 380, 381, 382, ​​383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 443, 444, 445, 446, 447, 448, 449, 450, 451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492, 493, 494, 495, 496, 497, 498, 499 and 500 nucleotides.

[0267] In some embodiments, the polyadenylation sequence is 50-100 nucleotides in length. In some embodiments, the polyadenylation sequence is 50-150 nucleotides in length. In some embodiments, the polyadenylation sequence is 50-160 nucleotides in length. In some embodiments, the polyadenylation sequence is 50-200 nucleotides in length. In some embodiments, the polyadenylation sequence is 60-100 nucleotides in length. In some embodiments, the polyadenylation sequence is 60-150 nucleotides in length. In some embodiments, the polyadenylation sequence is 60-160 nucleotides in length. In some embodiments, the polyadenylation sequence is 60-200 nucleotides in length. In some embodiments, the polyadenylation sequence is 70-100 nucleotides in length. In some embodiments, the polyadenylation sequence is 70-150 nucleotides in length. In some embodiments, the polyadenylation sequence is 70-160 nucleotides in length. In some embodiments, the polyadenylation sequence is 70-200 nucleotides in length. In some embodiments, the polyadenylation sequence is 80-100 nucleotides in length. In some embodiments, the polyadenylation sequence is 80-150 nucleotides in length. In some embodiments, the polyadenylation sequence is 80-160 nucleotides in length. In some embodiments, the polyadenylation sequence is 80-200 nucleotides in length. In some embodiments, the polyadenylation sequence is 90-100 nucleotides in length. In some embodiments, the polyadenylation sequence is 90-150 nucleotides in length. In some embodiments, the polyadenylation sequence is 90-160 nucleotides in length. In some embodiments, the polyadenylation sequence is 90-200 nucleotides in length. [connector] []

[0268] Vector gene bodies can be engineered with one or more spacer or linker regions separating coding or non-coding regions.

[0269] In some embodiments, the payload region of the vector genome optionally encodes one or more linker sequences. In some cases, the linker can be a peptide linker that can be used to link the polypeptides encoded by the payload region (ie, during expression, the antibody light and heavy chains). Some peptide linkers can be cleaved into individual heavy and light chain domains after expression, enabling assembly of mature antibodies or antibody fragments. Linker cleavage can be enzymatic. In some cases, the linker contains an enzymatic cleavage site to facilitate intracellular or extracellular cleavage. Some payload regions encode linkers that interrupt polypeptide synthesis during translation of linker sequences from mRNA transcripts. Such linkers can facilitate translation of separate protein domains from a single transcript. In some cases, two or more linkers are encoded by the payload region of the vector gene body.

[0270] Internal ribosome entry sites (IRES) are nucleotide sequences (>500 nucleotides) that allow initiation of translation in the middle of the mRNA sequence (Kim, J.H. et al., 2011. PLoS One 6(4): e18556; The content of which is incorporated herein by reference in its entirety). The use of IRES sequences ensures co-expression of genes preceding and following the IRES, although sequences following the IRES may be transcribed and translated at a lower level than sequences preceding the IRES.

[0271] 2A peptides are small "self-cleaving" peptides (18-22 amino acids) derived from such species as foot-and-mouth disease virus (F2A), porcine teschovirus-1 (P2A), thoseaasigna ) virus (T2A) or equine rhinitis A virus (E2A) virus. The 2A designation refers in particular to the region of the picornavirus polymeric protein that causes ribosome skipping at the glycyl-prolinyl bond in the C-terminus of the 2A peptide (Kim, J.H. et al., 2011. PLoS One 6(4 ): e18556; the contents of which are incorporated herein by reference in their entirety). This jump causes a cleavage between the 2A peptide and its immediately downstream peptide. In contrast to the IRES linker, the 2A peptide produces a stoichiometric expression of the protein flanked by the 2A peptide, and its shorter length can be advantageous in generating viral expression vectors.

[0272] Some payload regions encode linkers that contain furin cleavage sites. Furin is a calcium-dependent serine endoprotease that cleaves proteins just downstream of the basic amino acid target sequence (Arg-X-(Arg / Lys)-Arg) (Thomas, G., 2002. Nature Reviews Molecular Cell Biology 3(10): 753-66; the contents of which are incorporated herein by reference in their entirety). Furin is enriched in the mature trans-golgi network of Golgi bodies, where it participates in the processing of cellular precursor proteins. Furin also plays a role in the activation of various pathogens. This activity can be used to express the polypeptides of the present invention.

[0273] In some embodiments, the payload region may encode one or more linkers comprising a cathepsin, matrix metalloproteinase, or pod protein cleavage site. Such linkers are described, for example, by Cizeau and Macdonald in International Publication No. WO2008052322, the contents of which are incorporated herein by reference in their entirety. Cathepsins are a family of proteases with a unique mechanism for cleaving specific proteins. Cathepsin B is a cysteine ​​protease, and cathepsin D is an aspartyl protease. Matrix metalloproteinases are a family of calcium-dependent and zinc-containing endopeptidases. Podin is an enzyme that catalyzes the hydrolysis of (-Asn-Xaa-) bonds in proteins and small molecule substrates.

[0274] In some embodiments, the payload region can encode an uncleaved linker. Such linkers may include simple amino acid sequences, such as glycine-rich sequences. In some cases, the linker can comprise a flexible peptide linker comprising glycine and serine residues. Linkers may comprise flexible peptide linkers of varying lengths, eg nxG4S, where n=1-10, and the length of the encoded linker varies between 5 and 50 amino acids. In a non-limiting example, the linker can be 5xG4S. These flexible linkers are small and have no side chains, so they tend not to affect secondary protein structure, while providing a flexible linker between antibody segments (George, R.A. et al., 2002. Protein Engineering 15(11): 871-9; Huston, J.S. et al., 1988. PNAS 85:5879-83; and Shan, D. et al., 1999. Journal of Immunology. 162(11):6589-95; their respective The content of which is incorporated herein by reference in its entirety). In addition, the polarity of serine residues improves solubility and prevents aggregation problems.

[0275] In some embodiments, the payload region of the present invention may encode a small and unbranched serine-rich peptide linker, such as that described by Huston et al. in U.S. Patent No. 5,525,491, the contents of which are incorporated by reference in their entirety incorporated into this article. Encoded by the payload region of the invention, polypeptides linked by serine-rich linkers have increased solubility.

[0276] In some embodiments, the payload region of the invention may encode an artificial linker, such as those described by Whitlow and Filpula in U.S. Patent No. 5,856,456 and by Ladner et al. in U.S. Patent No. 4,946,778 , their respective contents are incorporated herein by reference in their entirety. [intron] []

[0277] In some embodiments, the payload region comprises at least one element for enhancing performance, such as one or more introns or portions thereof. Non-limiting examples of introns include MVM (67-97 bp), F.IX truncated intron 1 (300 bp), β-hemoglobulin SD / immunoglobulin heavy chain splice acceptor (250 bp ), adenovirus splice donor / immunoglobulin splice acceptor (500 bp), SV40 late splice donor / splice acceptor (19S / 16S) (180 bp) and hybrid adenovirus splice donor / IgG splice acceptor ( 230 bp).

[0278] In some embodiments, an intron or portion of an intron may be 100-500 nucleotides in length. The length of the intron can be 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 190, 200 ,210,220,230,240,250,260,270,280,290,300,310,320,330,340,350,360,370,380,390,400,410,420,430,440,450 , 460, 470, 480, 490 or 500. The length of the intron can be 80-100, 80-120, 80-140, 80-160, 80-180, 80-200, 80-250, 80-300, 80-350, 80-400, 80-450 , 80-500, 200-300, 200-400, 200-500, 300-400, 300-500 or 400-500. [Lentiviral vector] []

[0279] Lentiviral vectors are a type of retrovirus that can infect both dividing and non-dividing cells because their viral capsids can penetrate the intact membrane of the target cell's nucleus. Lentiviral vectors have the ability to deliver transgenes to long-lasting tissues that cannot be treated with stable gene manipulation. Lentiviral vectors have also opened fresh perspectives for gene therapy of a wide range of genetic and acquired disorders, and real proposals for their clinical use appear to be on the horizon. RNA

[0280] Ribonucleic acid (RNA) is a molecule composed of nucleotides in the form of ribose sugar linked to a nitrogenous base and a phosphate group. Nitrogenous bases include adenine (A), guanine (G), uracil (U) and cytosine (C). Generally speaking, RNA mainly exists in single-stranded form, but in some cases, it can also exist in double-stranded form. The length, form and structure of the RNA vary depending on the purpose of the RNA. For example, RNAs can vary in length from short sequences (such as siRNA) to long sequences (such as lncRNA), can be linear (such as mRNA) or circular (such as oRNA), and can be coding (such as mRNA) or Non-coding (eg lncRNA) sequences.

[0281] In some embodiments, the payload region can be or encodes a coding RNA.

[0282] In some embodiments, the payload region can be or encodes a non-coding RNA.

[0283] In some embodiments, the payload region can be or encodes both coding and non-coding RNA.

[0284] In some embodiments, the payload region comprises nucleic acid sequences encoding more than one cargo or payload.

[0285] In some embodiments, the payload region comprises a nucleic acid sequence that enhances the expression of the gene. As a non-limiting example, the nucleic acid sequence is messenger RNA (mRNA). As another non-limiting example, the nucleic acid sequence is a circular RNA (oRNA).

[0286] In some embodiments, the payload region comprises a nucleic acid sequence that reduces or inhibits the expression of a gene. As a non-limiting example, the nucleic acid sequence is a small interfering RNA (siRNA) or a microRNA (miRNA). [messenger] [RNA (mRNA)]

[0287] In some embodiments, the initial construct and / or reference construct can be mRNA. As used herein, the term "messenger RNA" (mRNA) refers to any polynucleotide that encodes a target of interest and is capable of translation to produce the encoded target of interest in vitro, in vivo, in situ or ex vivo.

[0288] Generally speaking, an mRNA molecule at least includes a coding region, a 5' untranslated region (UTR), a 3' UTR, a 5' cap and a poly A tail. In some aspects, one or more structural and / or chemical modifications or alterations may be included in the RNA that reduce the innate immune response of the cell into which the mRNA is introduced. As used herein, a "structural" feature or modification is one in which two or more linked nucleotides are inserted, deleted, repeated, inverted, or randomized in a nucleic acid without significant chemical modification of the nucleotides themselves. . Structural modifications are chemical in nature and are therefore chemical modifications because chemical bonds will necessarily be broken and reformed to effectuate the structural modification. However, structural modifications will result in different nucleotide sequences. For example, the polynucleotide "ATCG" can be chemically modified into "AT-5meC-G".

[0289] In general, the minimum length of the region of the initial construct and / or reference construct may be sufficient to encode a nucleic acid for a dipeptide, tripeptide, tetrapeptide, pentapeptide, hexapeptide, heptapeptide, octapeptide, nonapeptide, or decapeptide The length of the sequence. In another embodiment, the length may be sufficient to encode a peptide of 2-30 amino acids, such as 5-30, 10-30, 2-25, 5-25, 10-25, or 10-20 amino acids . The length may be sufficient to encode a peptide of at least 11, 12, 13, 14, 15, 17, 20, 25 or 30 amino acids, or no greater than 40 amino acids, such as no greater than 35, 30, 25, 20, Peptides of 17, 15, 14, 13, 12, 11 or 10 amino acids.

[0290] Generally, the region of mRNA encoding a target of interest is greater than about 30 nucleotides in length (e.g., at least or greater than about 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120 , 140, 160, 180, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,00 0 . 000 nucleotides).

[0291] In some embodiments, the mRNA comprises about 30 to about 100,000 nucleotides (e.g., 30 to 50, 30 to 100, 30 to 250, 30 to 500, 30 to 1,000, 30 to 1,500, 30 to 3,000, 30 to 5,000 . to 10,000, 100 to 25,000, 100 to 50,000, 100 to 70,000, 100 to 100,000, 500 to 1,000, 500 to 1,500, 500 to 2,000, 500 to 3,000, 500 to 5,000, 500 to 7,000, 500 to 10,0 00, 500 to 25,000 , 500 to 50,000, 500 to 70,000, 500 to 100,000, 1,000 to 1,500, 1,000 to 2,000, 1,000 to 3,000, 1,000 to 5,000, 1,000 to 7,000, 1,000 to 10,000, 1,000 to 25,000 , 1,000 to 50,000, 1,000 to 70,000 , 1,000 to 100,000, 1,500 to 3,000, 1,500 to 5,000, 1,500 to 7,000, 1,500 to 10,000, 1,500 to 25,000, 1,500 to 50,000, 1,500 to 70,000, 1,500 to 100,000, 2,0 00 to 3,000, 2,000 to 5,000, 2,000 to 7,000 , 2,000 to 10,000, 2,000 to 25,000, 2,000 to 50,000, 2,000 to 70,000 and 2,000 to 100,000).

[0292] In some embodiments, one or more regions flanking the region encoding a target of interest may independently range in length from 15-1,000 nucleotides (e.g., greater than 30, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, and 900 nucleotides or at least 30, 40, 45, 50 , 55, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, and 1,000 nucleotides).

[0293] In some embodiments, the mRNA may range in length from nonexistent to 500 nucleotides (e.g., at least 60, 70, 80, 90, 120, 140, 160, 180, 200, 250, 300, 350, 400 , 450 or 500 nucleotides) tailed sequence. Where the tailed region is a polyA tail, the length can be determined in units of or as a function of polyA binding protein binding. In this example, the poly A tail is long enough to bind at least 4 monomers of the poly A binding protein. The poly A-binding protein monomer binds to an approximately 38 nucleotide extension. Thus, poly-A tails of approximately 80 nucleotides and 160 nucleotides have been observed to be functional.

[0294] In some embodiments, the mRNA comprises a capping sequence comprising a single cap or a series of nucleotides forming a cap. The capping sequence may be 1 to 10, eg 2-9, 3-8, 4-7, 1-5, 5-10 or at least 2, or 10 or fewer nucleotides in length. In some embodiments, capping sequences are absent.

[0295] In some embodiments, the mRNA comprises a region containing an initiation codon. The length of the region comprising the initiation codon may range from 3 to 40, such as 5-30, 10-20, 15, or at least 4, or 30 or fewer nucleotides.

[0296] In some embodiments, the mRNA comprises a region containing a stop codon. The length of the region comprising a stop codon may range from 3 to 40, such as 5-30, 10-20, 15, or at least 4, or 30 or fewer nucleotides.

[0297] In some embodiments, the mRNA comprises a region comprising restriction sequences. The region comprising the restriction sequence may range in length from 3 to 40, eg 5-30, 10-20, 15, or at least 4 or 30 or fewer nucleotides. [Non-translation area] [(] [UTR] [)] []

[0298] In some embodiments, the mRNA comprises at least one untranslated region (UTR) flanking a region encoding a target of interest. UTR is transcribed, but not translated.

[0299] The 5' UTR starts at the transcription start site and continues up to, but not including, the start codon; while the 3' UTR starts immediately at the stop codon and continues until a transcription termination signal. While not wishing to be bound by theory, UTRs may have regulatory roles in the translation of nucleic acids and their stability.

[0300] The translation initiation typically included in natural 5'UTRs is a useful feature because it tends to include Kozak sequences generally known to be involved in the process by which ribosomes initiate translation of many genes. Kozak sequences have a common CCR(A / G)CCAUGG, where R is a purine (adenine or guanine) three bases upstream of the start codon (AUG), followed by another 'G'. The 5'UTR is also known to form secondary structures involved in elongation factor binding.

[0301] It is known that adenosine and uridine extensions are embedded in the 3' UTR. Such AU-rich markers are especially prevalent in genes with high turnover rates. Based on their sequence features and functional properties, AU-rich elements (AREs) can be divided into three classes (Chen et al., 1995): Class I AREs contain several scattered copies of the AUUUA motif within the U-rich region. C-Myc and MyoD contain class I AREs. Class II AREs have two or more overlapping UUAUUUA(U / A)(U / A) nonamers. Molecules containing this type of ARE include GM-CSF and TNF-a. Class III AREs are less well defined. These U-rich regions do not contain the AUUUA motif. c-Jun and myogenin are two well-studied examples of this class. Most proteins bound to AREs are known to destabilize the messenger, while members of the ELAV family (most notably, HuR) have been documented to increase the stability of mRNA. HuR binds to all three classes of AREs. Engineering a HuR-specific binding site into the 3'UTR of a nucleic acid molecule will result in HuR binding and thus message stabilization in vivo. The introduction, removal or modification of 3'UTR AU-rich elements (AREs) can be used to regulate the stability of mRNA. For example, one or more copies of the ARE can be introduced to make the mRNA less stable and thus reduce translation and reduce the yield of the resulting protein. Alternatively, AREs can be identified and removed or mutated to increase intracellular stability and thus increase translation and production of the resulting protein.

[0302] In some embodiments, mRNA stability and protein production in a particular organ and / or tissue can be enhanced by introducing a signature that is normally expressed in the gene of the target organ. As a non-limiting example, this feature may be a UTR. As another example, the feature can be an intron or part of an intron sequence. [5'] [capped] []

[0303] The 5' cap structure of mRNA is involved in nuclear export, increases mRNA stability and binds mRNA cap-binding protein (CBP), which is responsible for the stabilization of mRNA in cells via the binding of CBP to poly(A)-binding protein to form mature circular mRNA species and translational abilities. The cap further aids in the removal of 5' proximal intron removal during mRNA splicing.

[0304] Endogenous mRNA molecules can be capped at the 5' end, thereby creating a 5'-ppp-5'-triphosphate bond between the terminal guanosine cap residue of the mRNA molecule and the transcribed sense nucleotide at the 5' end couplet. This 5'-guanylate cap can then be methylated to generate an N7-methyl-guanylate residue. The ribose sugar of the terminal and / or anteterminal transcribed nucleotides of the 5' end of the mRNA may also optionally be 2'-O-methylated. 5'-Decapping by hydrolysis and cleaving of the guanylate cap structure can target nucleic acid molecules, such as mRNA molecules, for degradation.

[0305] Modifications to mRNA can create a non-hydrolyzable cap structure, preventing uncapping and thus increasing mRNA half-life. Since hydrolysis of the cap structure requires cleavage of the 5'-ppp-5' phosphodiester linkage, modified nucleotides can be used during the capping reaction. For example, vaccinia capping enzyme from New England Biolabs (Ipswich, MA) and a-thio-guanosine nucleotides can be used according to the manufacturer's instructions to generate thio in the 5'-ppp-5' cap Phosphate linkage.

[0306] Additional modified guanosine nucleotides can be used, such as α-methyl-phosphonate and seleno-phosphate nucleotides.

[0307] Additional modifications include, but are not limited to, 2'-O-methylation of the ribose sugar on the 2'-hydroxyl of the sugar ring at the 5' end of the mRNA (as mentioned above) and / or the nucleotide preceding the 5' end. A number of different 5' cap structures can be used to create 5' caps for nucleic acid molecules such as mRNA molecules.

[0308] Cap analogs, also referred to herein as synthetic cap analogs, chemical caps, chemical cap analogs, or structural or functional cap analogs, differ in their chemical structure from natural (i.e., endogenous, wild-type, or physiological) cap analogs. ) 5' caps are different while retaining cap function. Cap analogs can be chemically (ie, non-enzymatically) or enzymatically synthesized and / or linked to nucleic acid molecules.

[0309] For example, the reverse cap analog (ARCA) cap contains two guanines linked by a 5'-5'-triphosphate group, one of which contains an N7 methyl group and a 3'-O-methyl ( That is, N7,3'-O-dimethyl-guanosine-5'-triphosphate-5'-guanosine (m 7G-3'mppp-G; which may equivalently be referred to as 3'O- Me-m7G(5')ppp(5')G). The 3'-O atom of another unmodified guanine becomes nucleotide-linked to the 5' end of a capped nucleic acid molecule (eg, mRNA). N7 - and 3'-O-methylated guanines provide a capped terminal portion of a nucleic acid molecule such as mRNA.

[0310] Another exemplary cap is mCAP, which is similar to ARCA but has a 2'-O-methyl group on the guanosine (i.e., N7,2'-O-dimethyl-guanosine-5'-triphosphate- 5'-guanosine, m 7Gm-ppp-G).

[0311] Although cap analogs allow concomitant capping of nucleic acid molecules in in vitro transcription reactions, up to 20% of transcripts can remain uncapped. This, as well as structural differences between cap analogs and the endogenous 5' cap structure of nucleic acids produced by the endogenous cellular transcription machinery, can lead to reduced translational capacity and reduced cellular stability.

[0312] mRNA can also be capped post-transcriptionally using enzymes to create a more realistic 5' cap structure. As used herein, the phrase "more authentic" refers to features that closely mirror or mimic endogenous or wild-type features, either structurally or functionally. That is, "true" features are better representative of endogenous, wild-type, native or physiological cell function and / or structure than synthetic features or analogs of the prior art, or in one or more respects Outperforms the corresponding endogenous, wild-type, native or physiological trait. Non-limiting examples of more authentic 5' cap structures are, inter alia, those that are comparable to synthetic 5' cap structures known in the art (or to wild-type, natural or physiological 5' cap structures ) have enhanced binding of cap binding proteins, increased half-life, reduced sensitivity to 5' endonucleases and / or reduced 5' decapping. For example, recombinant vaccinia virus capping enzymes and recombinant 2'-O-methyltransferases can generate a typical 5'-5'-tri-nucleotide between the 5' terminal nucleotide and the guanine cap nucleotide of the mRNA. Phosphate linkage, in which the cap guanine contains N7 methylation, and the 5' end nucleotide of the mRNA contains 2'-O-methyl. Such structures are called Capl structures. This cap results in higher translational capacity and cell stability and reduced activation of cellular pro-inflammatory cytokines compared to, for example, other 5' cap analog structures known in the art. Cap structures include, but are not limited to, 7mG(5*)ppp(5*)N,pN2p (cap 0), 7mG(5*)ppp(5*)NlmpNp (cap 1), and 7mG(5*)-ppp(5' )NlmpN2mp (cap 2).

[0313] In some embodiments, the 5' end cap can comprise an endogenous cap or a cap analog.

[0314] In some embodiments, the 5' end cap can comprise a guanine analog. Suitable guanine analogs include, but are not limited to, inosine, Nl-methyl-guanosine, 2'fluoro-guanosine, 7-deaza-guanosine, 8-oxo-guanosine, 2-amino- Guanosine, LNA-guanosine and 2-azido-guanosine. [IRES] [sequence] []

[0315] In some embodiments, the mRNA may contain an internal ribosome entry site (IRES). IRES were first identified as characteristic picornavirus RNAs that play an important role in initiating protein synthesis in the absence of a 5' cap structure. An IRES can serve as the only ribosome-binding site, or it can serve as one of multiple ribosome-binding sites for mRNA. An mRNA containing more than one functional ribosome binding site may encode several peptides or polypeptides that are translated independently by ribosomes. Non-limiting examples of IRES sequences that may be used include, but are not limited to, those from Picornaviruses (e.g. FMDV), Harmful Organism Viruses (CFFV), Poliovirus (PV), Encephalomyocarditis Virus (ECMV) ), foot-and-mouth disease virus (FMDV), hepatitis C virus (HCV), classical swine fever virus (CSFV), murine leukemia virus (MLV), simian immunodeficiency virus (SIV), or cricket paralysis virus (CrPV). [poly] [A] [tail] []

[0316] During RNA processing, long chains of adenine nucleotides (poly-A tails) can be added to polynucleotides such as mRNA molecules to increase stability. Immediately after transcription, the 3' end of the transcript can be cleaved to release the 3' hydroxyl. Next, poly A polymerase adds a series of adenine nucleotides to the RNA. This process, called polyadenylation, adds a polyA tail of a certain length.

[0317] In some embodiments, the length of the poly A tail is greater than 30 nucleotides in length. In another embodiment, the length of the poly A tail is greater than 35 nucleotides (e.g., at least or greater than about 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,500 and 3 ,000 Nucleotides). In some embodiments, the mRNA comprises about 30 to about 3,000 nucleotides (e.g., 30 to 50, 30 to 100, 30 to 250, 30 to 500, 30 to 750, 30 to 1,000, 30 to 1,500, 30 to 2,000 . , 100 to 1,000, 100 to 1,500, 100 to 2,000, 100 to 2,500, 100 to 3,000, 500 to 750, 500 to 1,000, 500 to 1,500, 500 to 2,000, 500 to 2,500, 500 to 3,000, 1,000 to 1,500, 1,000 to 2,000, 1,000 to 2,500, 1,000 to 3,000, 1,500 to 2,000, 1,500 to 2,500, 1,500 to 3,000, 2,000 to 3,000, 2,000 to 2,500 and 2,500 to 3,000).

[0318] In some embodiments, the poly A tail is designed relative to the length of the entire mRNA. This design can be based on the length of the region encoding the target of interest, the length of a particular feature or region such as a flanking region, or based on the length of the end product expressed from the mRNA.

[0319] In this context, the length of the poly A tail may be 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100% longer than the mRNA or its signature. The poly A tail can also be designed as part of the mRNA to which it belongs. In this context, the poly A tail can be the full length of the construct or 10, 20, 30, 40, 50, 60, 70, 80 or 90% or more of the full length of the construct minus the poly A tail. In addition, engineered binding sites of poly A binding proteins and mRNA binding can enhance performance.

[0320] In addition, multiple different mRNAs can be linked together at the 3' end of the poly A tail via 3' end with PABP (poly A binding protein) using modified nucleotides. Transfection experiments can be performed in relevant cell lines, and protein production can be analyzed by ELISA at 12 hours, 24 hours, 48 ​​hours, 72 hours and 7 days after transfection.

[0321] In some embodiments, the mRNA is designed to include a poly A-G quartet (Quartet). G-quadruplexes are circular hydrogen-bonded arrays of four guanine nucleotides that can be formed by G-rich sequences in DNA and RNA. In this example, a G-quartet is incorporated at the end of the poly A tail. [stop codon] []

[0322] In some embodiments, the mRNA can include a stop codon. In some embodiments, the mRNA can include two stop codons. In some embodiments, the mRNA can include three stop codons. In some embodiments, the mRNA can include at least one stop codon. In some embodiments, the mRNA can include at least two stop codons. In some embodiments, the mRNA can include at least three stop codons. As non-limiting examples, stop codons may be selected from TGA, TAA and TAG.

[0323] In some embodiments, the mRNA includes a stop codon TGA and one additional stop codon. In another embodiment, the addition of a stop codon may be TAA. [circular] [RNA] [] [(oRNA)] []

[0324] In some embodiments, the initial construct and / or reference construct is a circular RNA (oRNA). As used herein, the terms "oRNA" or "circular RNA" are used interchangeably and can refer to RNA that forms a circular structure via covalent or non-covalent bonds.

[0325] In some embodiments, oRNAs may be non-immunogenic in mammals (eg, humans, non-human primates, rabbits, rats, and mice).

[0326] In some embodiments, the oRNA is replicable or replicable in cells from aquatic animals (e.g., fish, crab, shrimp, oysters, etc.), mammalian cells, from pet or zoo animals (e.g., cats, dogs, lizards, etc.) , birds, lions, tigers and bears, etc.), cells from farm animals or draft animals (such as horses, cows, pigs, chickens, etc.), human cells, cultured cells, primary cells or cell lines, stem cells, precursors cells, differentiated cells, germ cells, cancer cells (e.g., tumorigenic, metastatic), non-tumorigenic cells (e.g., normal cells), fetal cells, embryonic cells, adult cells, mitotic cells, non-mitotic cells, or any combination thereof .

[0327] In some embodiments, the oRNA has a half-life that is at least that of the linear counterpart. In some embodiments, the oRNA has an increased half-life relative to that of the linear counterpart. In some embodiments, half-life is increased by about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or greater. In some embodiments, the half-life or persistence of the oRNA in the cell is at least about 1 hour to about 30 days, or at least about 2 hours, 6 hours, 12 hours, 18 hours, 24 hours (1 day), 2 days, 3, days, 4 days, 5 days, 6 days, 7 days, 8 days, 9 days, 10 days, 11 days, 12 days, 13 days, 14 days, 15 days, 16 days, 17 days, 18 days, 19 days days, 20 days, 21 days, 22 days, 23 days, 24 days, 25 days, 26 days, 27 days, 28 days, 29 days, 30 days, 60 days or longer or any period in between. In some embodiments, the half-life or persistence of the oRNA in the cell is no more than about 10 min to about 7 days, or no more than about 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 24 hours (1 day ), 36 hours (1.5 days), 48 hours (2 days), 60 hours (2.5 days), 72 hours (3 days), 4 days, 5 days, 6 days or 7 days.

[0328] In some embodiments, the oRNA has a half-life or persistence in the cell while the cell is dividing. In some embodiments, the oRNA has a half-life or persistence in post-dividing cells. In certain embodiments, the half-life or persistence of the oRNA in dividing cells is greater than about 10 minutes to about 30 days, or at least about 10 minutes, 15 minutes, 30 minutes, 45 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 24 hours (1 day ), 2 days, 3, days, 4 days, 5 days, 6 days, 7 days, 8 days, 9 days, 10 days, 11 days, 12 days, 13 days, 14 days, 15 days, 16 days, 17 days , 18 days, 19 days, 20 days, 21 days, 22 days, 23 days, 24 days, 25 days, 26 days, 27 days, 28 days, 29 days, 30 days, 60 days or longer or any period in between .

[0329] In some embodiments, the oRNA modulates cellular function, e.g., temporarily or long-term. In certain embodiments, cell function is stably altered, such as maintained for at least about 1 hour to about 30 days, or at least about 2 hours, 6 hours, 12 hours, 18 hours, 24 hours (1 day), 2 days, 3 days. days, 4 days, 5 days, 6 days, 7 days, 8 days, 9 days, 10 days, 11 days, 12 days, 13 days, 14 days, 15 days, 16 days, 17 days, 18 days, 19 days, 20 days, 21 days, 22 days, 23 days, 24 days, 25 days, 26 days, 27 days, 28 days, 29 days, 30 days, 60 days or longer. In certain embodiments, cell function is temporarily altered, e.g., maintained for no more than about 30 minutes to about 7 days, or for no more than about 30 minutes, 45 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours , 23 hours, 24 hours (1 day), 36 hours (1.5 days), 48 hours (2 days), 60 hours (2.5 days), 72 hours (3 days), 4 days, 5 days, 6 days, or 7 days .

[0330] In some embodiments, the oRNA is at least about 20 nucleotides, at least about 30 nucleotides, at least about 40 nucleotides, at least about 50 nucleotides, at least about 75 nucleotides, at least about 100 nucleotides, at least about 200 nucleotides, at least about 300 nucleotides, at least about 400 nucleotides, at least about 500 nucleotides, at least about 1,000 nucleotides, at least about 2,000 Nucleotides, at least about 5,000 nucleotides, at least about 6,000 nucleotides, at least about 7,000 nucleotides, at least about 8,000 nucleotides, at least about 9,000 nucleotides, at least about 10,000 nucleosides acid, at least about 12,000 nucleotides, at least about 14,000 nucleotides, at least about 15,000 nucleotides, at least about 16,000 nucleotides, at least about 17,000 nucleotides, at least about 18,000 nucleotides, At least about 19,000 nucleotides or at least about 20,000 nucleotides. In some embodiments, the oRNA can be of sufficient size to accommodate a ribosome binding site.

[0331] In some embodiments, the maximum size of the oRNA may be limited by the ability to encapsulate and deliver the RNA to the target. In some embodiments, the size of the oRNA is of sufficient length to encode a polypeptide, and thus at least 20,000 nucleotides, at least 15,000 nucleotides, at least 10,000 nucleotides, at least 7,500 nucleotides, or at least 5,000 nucleotides nucleotides, at least 4,000 nucleotides, at least 3,000 nucleotides, at least 2,000 nucleotides, at least 1,000 nucleotides, at least 500 nucleotides, at least 400 nucleotides, at least 300 nucleotides A length of nucleotides, at least 200 nucleotides, at least 100 nucleotides may be suitable.

[0332] In some embodiments, the oRNA comprises one or more elements described elsewhere herein. In some embodiments, elements may be separated from each other by spacer sequences or linkers. In some embodiments, an element can consist of 1 nucleotide, 2 nucleotides, about 5 nucleotides, about 10 nucleotides, about 15 nucleotides, about 20 nucleotides, about 30 nucleotides nucleotides, about 40 nucleotides, about 50 nucleotides, about 60 nucleotides, about 80 nucleotides, about 100 nucleotides, about 150 nucleotides, about 200 nucleotides Nucleotides, about 250 nucleotides, about 300 nucleotides, about 400 nucleotides, about 500 nucleotides, about 600 nucleotides, about 700 nucleotides, about 800 nucleotides The nucleotides are separated from each other by about 900 nucleotides, about 1000 nucleotides, up to about 1 kb, at least about 1000 nucleotides.

[0333] In some embodiments, one or more elements are adjacent to each other, eg, lacking spacer elements.

[0334] In some embodiments, one or more elements are configurationally flexible. In some embodiments, conformational flexibility is due to the sequence being substantially free of secondary structure.

[0335] In some embodiments, the oRNA comprises a secondary or tertiary structure that accommodates binding sites for ribosomes, translation, or rolling circle translation.

[0336] In some embodiments, oRNAs comprise specific sequence features. For example, an oRNA can comprise a specific nucleotide composition. In some such embodiments, the oRNA can include one or more purine-rich regions (adenine or guanosine). In some embodiments, an oRNA can include one or more AU-rich regions or elements (AREs). In some embodiments, an oRNA can include one or more adenine-rich regions.

[0337] In some embodiments, the oRNA comprises one or more modifications described elsewhere herein.

[0338] In some embodiments, the oRNA comprises one or more expression sequences and is configured for persistent expression in cells of an individual in vivo. In some embodiments, the oRNA is configured such that at a later time point the expression of one or more expressed sequences in the cell is equal to or higher than at an earlier time point. In such embodiments, the performance of one or more performance sequences may remain at a relatively constant level or may increase over time. The performance of the performance series can be relatively stable for an extended period of time. For example, in some cases, one or more expressed sequences are expressed in cells over a period of at least 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 23 or more days Performance is not reduced by 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10% or 5%. In some cases, the expression of one or more expressed sequences in the cells is maintained at a rate that does not vary by more than 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5%. at least 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 23 or more days. [adjustment element] []

[0339] In some embodiments, the oRNA comprises a regulatory element. As used herein, a "regulatory element" is a sequence that modulates the expression of an expressed sequence. Regulatory elements may include sequences located near the payload or cargo area. An adjustment element may be operatively connected to the payload or cargo area.

[0340] In some embodiments, the modulating element can increase the amount of payload or cargo represented compared to the amount represented when no modulating element is present. As a non-limiting example, an adjustment element can increase the amount of payload or cargo represented by multiple payload or cargo sequences connected in series.

[0341] In some embodiments, a regulatory element may comprise a sequence that selectively initiates or activates translation of a payload or cargo.

[0342] In some embodiments, a regulatory element may comprise a sequence that initiates degradation of an oRNA or payload or cargo. Non-limiting examples of sequences that initiate degradation include, but are not limited to, riboswitch aptazymes and miRNA binding sites.

[0343] In some embodiments, a regulatory element can regulate translation of a payload or cargo in an oRNA. This modulation can result in an increase (enhancer) or decrease (inhibitor) of the payload or cargo. The adjustment element may be located adjacent to the payload or cargo (eg, on one or both sides of the payload or cargo).

[0344] In some embodiments, translation initiation sequences serve as regulatory elements. In some embodiments, the translation initiation sequence comprises AUG / ATG codons. In some embodiments, the translation initiation sequence comprises any eukaryotic initiation codon, such as but not limited to AUG / ATG, CUG / CTG, GUG / GTG, UUG / TTG, ACG, AUC / ATC, AUU, AAG, AUA / ATA or AGG. In some embodiments, the translation initiation sequence comprises a Kozak sequence. In some embodiments, translation begins under selective conditions (eg, stress-inducing conditions) with an alternative translation initiation sequence, eg, a translation initiation sequence other than AUG / ATG. As a non-limiting example, the translation of a circular polyribonucleotide can begin with an alternative translation initiation sequence, such as ACG. As another non-limiting example, circular polyribonucleotide translation can begin with an alternative translation initiation sequence CUG / CTG. As another non-limiting example, translation can begin with an alternative translation initiation sequence GUG / GTG. As yet another non-limiting example, translation can begin with repeat-associated non-AUG (RAN) sequences, such as alternative translation initiation sequences comprising short stretches of repetitive RNA (eg, CGG, GGGGCC, CAG, CTG). [masking agent] []

[0345] Masking any nucleotides flanking the codon that initiates translation can be used to alter translation initiation location, translation efficiency, length and / or structure of the oRNA. In some embodiments, a masking agent may be used near the start codon or alternative start codon to mask or hide the codon to reduce the The translation initiation probability of . Non-limiting examples of masking agents include antisense locked nucleic acid (LNA) oligonucleotides and exon junction complexes (EJC). In some embodiments, a masking agent may be used to mask the start codon of the oRNA so as to increase the likelihood that translation will be initiated at an alternative start codon. [translation start sequence] []

[0346] In some embodiments, the oRNA encodes a polypeptide or peptide and may comprise a translation initiation sequence. Translation initiation sequences may include, but are not limited to, initiation codons, noncoding initiation codons, Kozak sequences, or Shine-Dalgarno sequences. The translation initiation sequence can be located near the payload or cargo (eg, on one or both sides of the payload or cargo).

[0347] In some embodiments, the translation initiation sequence provides conformational flexibility to the oRNA. In some embodiments, the translation initiation sequence is within a substantially single-stranded region of the oRNA.

[0348] The oRNA may include more than 1 initiation codon, such as but not limited to at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13 , at least 14, at least 15 or more than 15 initiation codons. Translation can be initiated at the first initiation codon or can be initiated downstream of the first initiation codon.

[0349] In some embodiments, the oRNA can start at a codon other than the first start codon (eg, AUG). The translation of circular polyribonucleotides can be initiated at alternative translation initiation sequences, such as but not limited to ACG, AGG, AAG, CUG / CTG, GUG / GTG, AUA / ATA, AUU / ATT, UUG / TTG . In some embodiments, translation begins at an alternative translation initiation sequence under selective conditions, such as stress-inducing conditions. As a non-limiting example, translation of the oRNA can begin with an alternative translation initiation sequence, such as ACG. As another non-limiting example, oRNA translation can begin with an alternative translation initiation sequence CUG / CTG. As yet another non-limiting example, oRNA translation can begin with an alternative translation initiation sequence GTG / GUG. As yet another non-limiting example, oRNAs can initiate translation at repeat-associated non-AUG (RAN) sequences, such as alternative translation initiation sequences that include short stretches of repetitive RNA (e.g., CGG, GGGGCC, CAG, CTG) . [IRES] [sequence] []

[0350] In some embodiments, an oRNA described herein comprises an internal ribosome entry site (IRES) element capable of engaging a eukaryotic ribosome. In some embodiments, the IRES element is at least about 5 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 40 nucleotides, at least about 50 nucleotides, at least about 100 nucleotides, at least about 200 nucleotides, at least about 250 nucleotides, at least about 350 nucleotides, or at least about 500 nucleotides. In one embodiment, the IRES element is derived from the DNA of organisms including, but not limited to, viruses, mammals, and Drosophila. Such viral DNA can be derived from, but is not limited to, picornavirus complementary DNA (cDNA) and encephalomyocarditis virus (EMCV) cDNA and poliovirus cDNA. In one embodiment, the Drosophila DNA from which the IRES element is derived includes, but is not limited to, the antennapedia gene from Drosophila melanogaster.

[0351] In some embodiments, the IRES element is at least partially derived from a virus, for example it may be derived from a viral IRES element such as ABPV_IGRpred, AEV, ALPV_IGRpred, BQCV_IGRpred, BVDV1_1-385, BVDV1_29-391, CrPV_5NCR, CrPV_IGR, crTMV_IREScp, crTMV_IRESmp75, crTMV_IRESmp228 , crTMV_IREScp, crTMV_IREScp, CSFV, CVB3, DCV_IGR, EMCV-R, EoPV_5NTR, ERAV 245-961, ERBV 162-920, EV71_1-748, FeLV-Notch2, FMDV_type_C, GBV-A, GBV-B, GBV-C, gypsy_env , gypsyD5, gypsyD2, HAV_HM175, HCV_type_1a, HiPV_IGRpred, HIV-1, HoCV1_IGRpred, HRV-2, IAPV_IGRpred, idefix, KBV_IGRpred, LINE-1_ORF1_-101_to_-1, LINE-1_ORF1-302_to_-202, LINE -1_ORF2-138_to_-86 , LINE-1_ORF1_-44to_-1, PSIV_IGR, PV_type1_Mahoney, PV_type3_Leon, REV-A, RhPV_5NCR, RhPV_IGR, SINV1_IGRpred, SV40_661-830, TMEV, TMV_UI_IRESmp228, TRV_5NTR, TrV_IGR, or TSV_IGR. In some embodiments, the IRES element is at least partially derived from a cellular IRES, such as AML1 / RUNX1, Antp-D, Antp-DE, Antp-CDE, Apaf-1, Apaf-1, AQP4, AT1R_var1, AT1R_var2, AT1R_var3, AT1R_var4 , BAG1_p36delta236 nt, BAG1_p36, BCL2, BiP_-222_-3, c-IAP1_285-1399, c-IAP1_1313-1462, c-jun, c-myc, Cat-1224, CCND1, DAPS, eIF4G, eIF4GI-ext, eIF4GII, eIF4GII-long, ELG1, ELH, FGF1A, FMR1, Gtx-133-141, Gtx-1-166, Gtx-1-120, Gtx-1-196, hairless, HAP4, HIF1a, hSNM1, Hsp101, hsp70, hsp70, Hsp90, IGF2_leader2, Kv1.4_1.2, L-myc, LamB1_-335_-1, LEF1, MNT_75-267, MNT_36-160, MTG8a, MYB, MYT2_997-1152, n-MYC, NDST1, NDST2, NDST3, NDST4L, NDST4S, NRF_-653_-17, NtHSF1, ODC1, p27kip1, 03_128-269, PDGF2 / c-sis, Pim-1, PITSLRE_p58, Rbm3, reaper, Scamper, TFIID, TIF4631, Ubx_1- 966, Ubx_373-961, UNR, Ure2, UtrA, VEGF-A-133-1, XIAP_5-464, XIAP_305-466 or YAP1. [termination element] []

[0352] In some embodiments, an oRNA includes one or more cargo or payload sequences (also referred to as presentation sequences), and each cargo or payload sequence may or may not have a termination element.

[0353] In some embodiments, the oRNA includes one or more cargo or payload sequences, and the sequences lack termination elements such that the oRNA is continuously translated. Exclusion of the termination element allows for inverted rolling circle translation or continuous expression of the encoded peptide or polypeptide because the ribosome will not stall or fall-off. In such embodiments, rolling circle translation exhibits continuous representation through each cargo or payload sequence.

[0354] In some embodiments, one or more cargo or payload sequences in the oRNA comprise a termination element.

[0355] In some embodiments, not all cargo or payload sequences in the oRNA include termination elements. In such cases, the cargo or payload can fall off the ribosome when the ribosome encounters the termination element and translation is terminated. In some embodiments, translation is terminated while at least one region of the ribosome remains in contact with the oRNA. [rolling ring translation] []

[0356] In some embodiments, once translation of the oRNA is initiated, ribosomes bound to the oRNA do not disengage from the oRNA until at least one round of oRNA translation is complete. In some embodiments, an oRNA as described herein is competent for rolling circle translation. In some embodiments, during rolling circle translation, once translation of the oRNA is initiated, at least 2 rounds, at least 3 rounds, at least 4 rounds, at least 5 rounds, at least 6 rounds, at least 7 rounds, at least 8 rounds are completed, At least 9 rounds, at least 10 rounds, at least 11 rounds, at least 12 rounds, at least 13 rounds, at least 14 rounds, at least 15 rounds, at least 20 rounds, at least 30 rounds, at least 40 rounds, at least 50 rounds, at least 60 rounds, at least 70 rounds rounds, at least 80 rounds, at least 90 rounds, at least 100 rounds, at least 150 rounds, at least 200 rounds, at least 250 rounds, at least 500 rounds, at least 1000 rounds, at least 1500 rounds, at least 2000 rounds, at least 5000 rounds, at least 10000 rounds, At least 10.sup.5 rounds or at least 10.sup.6 rounds of oRNA translation, ribosomes bound to oRNA will not detach from oRNA.

[0357] In some embodiments, rolling circle translation of an oRNA results in the production of translated polypeptides resulting from more than one round of oRNA translation. In some embodiments, the oRNA comprises a stagger element, and rolling circle translation of the oRNA results in a polypeptide product resulting from a single round or less of oRNA translation. [cyclization] []

[0358] In one embodiment, the linear RNA can be circularized or concatemerized. In some embodiments, linear RNA can be circularized in vitro prior to formulation and / or delivery. In some embodiments, the linear RNA can circularize within the cell.

[0359] In some embodiments, the mechanism of cyclization or concatenation can proceed via at least 3 different routes: 1) chemical, 2) enzymatic, and 3) ribonuclease catalyzed. The newly formed 5'- / 3'-linkages can be intramolecular or intermolecular.

[0360] In the first approach, the 5' and 3' ends of the nucleic acid contain chemically reactive groups that, when brought together, form new bonds between the 5' and 3' ends of the molecule. Covalently linked. The 5' end may contain an NHS-ester reactive group, and the 3' end may contain a 3'-amino-terminated nucleotide, so that in organic solvents, the 3'- Amine-terminated nucleotides will undergo nucleophilic attack on the 5'-NHS-ester moiety, forming a new 5'- / 3'-amide bond.

[0361] In the second pathway, T4 RNA ligase can be used to enzymatically connect 5'-phosphorylated nucleic acid molecules to the 3'-hydroxyl of nucleic acids, thereby forming new phosphodiester linkages. In one example reaction, 1 μg of nucleic acid molecules were incubated with 1-10 units of T4 RNA ligase (New England Biolabs, Ipswich, MA) at 37°C for 1 hour according to the manufacturer's protocol. The ligation reaction can be performed in the presence of a resolving oligonucleotide capable of base pairing with both the juxtaposed 5'- and 3'-regions that facilitate the enzymatic ligation reaction.

[0362] In the third approach, the 5' or 3' end of the cDNA template encodes a ligase ribonuclease sequence, so that during in vitro transcription, the resulting nucleic acid molecule may contain a nucleic acid molecule capable of ligating the 5' end of the nucleic acid molecule to the 3' end of the nucleic acid molecule. ' end of the active ribonuclease sequence. Ligase ribonucleases can be derived from group I introns, group I introns, hepatitis D virus, hairpin ribonucleases, or by systematic evolution of ligands by exponential enrichment; SELEX) to select. The ribonuclease ligase reaction can take from 1 to 24 hours at a temperature between 0 and 37°C.

[0363] In some embodiments, oRNAs are made by circularizing linear RNAs. [Extracellular cyclization] []

[0364] In some embodiments, linear RNA is chemically circularized or concatenated to form oRNA. In some chemistries, the 5' and 3' ends of nucleic acids (such as linear RNA) include chemically reactive groups that, when brought into close proximity, can bind between the 5' and 3' ends of the molecule. A new covalent bond is formed between the ends. The 5' end can contain an NHS-ester reactive group, and the 3' end can contain a 3'-amine-terminated nucleotide, so that in organic solvents, the 3'-amine on the 3' end of the linear RNA A group-terminated nucleotide will undergo nucleophilic attack on the 5'-NHS-ester moiety, forming a new 5'- / 3'-amide bond.

[0365] In one embodiment, a DNA or RNA ligase can be used to enzymatically ligate a 5'-phosphorylated nucleic acid molecule (e.g., linear RNA) to the 3'-hydroxyl of a nucleic acid (e.g., linear nucleic acid), thereby forming a new phosphodiester linkage . In one example reaction, linear RNA was incubated with 1-10 units of T4 RNA ligase at 37°C for 1 hour according to the manufacturer's protocol. The ligation reaction can be performed in the presence of a linear nucleic acid capable of base pairing with both the juxtaposed 5'- and 3'-regions that facilitate the enzymatic ligation reaction. In one embodiment, the ligation is a splint ligation, wherein a single-stranded polynucleotide (splint), such as a single-stranded RNA, can be designed to hybridize to both ends of the linear RNA so that the two ends can be separated. Juxtaposed after hybridization with single-strand splints. Therefore, splint ligase can catalyze the juxtaposed two-end joining of linear RNAs to generate oRNAs.

[0366] In one embodiment, DNA or RNA ligase can be used to synthesize oRNA. As a non-limiting example, the ligase can be a circ ligase or a circular ligase.

[0367] In one embodiment, the 5' or 3' end of the linear RNA may encode a ligase ribonuclease sequence such that during in vitro transcription, the resulting linear RNA includes a ligase capable of ligating the 5' end of the linear RNA to the 3' end of the linear RNA. end of the active ribonuclease sequence. Ligase ribozymes can be derived from group I introns, hepatitis D virus, hairpin ribozymes, or can be selected by systematic evolution by ligand exponential enrichment (SELEX).

[0368] In one embodiment, linear RNA can be circularized or concatenated by using at least one non-nucleic acid moiety. In one aspect, the at least one non-nucleic acid moiety can react with a region or feature near the 5' end and / or near the 3' end of the linear RNA to circularize or concatenate the linear RNA. In another aspect, the at least one non-nucleic acid moiety can be located in or attached to or near the 5' end and / or 3' end of the linear RNA. Contemplated non-nucleic acid moieties may be homologous or heterologous. As a non-limiting example, a non-nucleic acid moiety can be a linkage, such as a hydrophobic linkage, an ionic linkage, a biodegradable linkage, and / or a cleavable linkage. As another non-limiting example, the non-nucleic acid moiety is a linking moiety. As yet another non-limiting example, the non-nucleic acid moiety can be an oligonucleotide or peptide moiety, such as an aptamer or a non-nucleic acid linker as described herein.

[0369] In one embodiment, the linear RNA can be circularized or concatenated due to non-nucleic acid moieties that cause attraction between the 5' and 3' ends of the linear RNA, atoms near or attached thereto, molecular surfaces. As a non-limiting example, one or more linear RNAs can be circularized or concatenated by intermolecular or intramolecular forces. Non-limiting examples of intermolecular forces include dipole-dipole forces, dipole-induced dipole forces, induced dipole-induced dipole forces, Van der Waals forces, and London dispersion forces . Non-limiting examples of intramolecular forces include covalent, metallic, ionic, resonant, agnostic, dipolar, conjugation, hyperconjugation, and antibonding.

[0370] In one embodiment, the linear RNA can comprise a ribonuclease RNA sequence near the 5' end and near the 3' end. The ribonuclease RNA sequence can be covalently linked to the peptide while the sequence is exposed to the rest of the ribonuclease. In one aspect, peptides covalently linked to ribonuclease RNA sequences near the 5' and 3' ends can bind to each other, allowing circularization or concatenation of the linear RNA. In another aspect, the peptide covalently linked to the ribonuclease RNA near the 5' and 3' ends after being subjected to ligation using various methods known in the art, such as but not limited to protein ligation, can be Circularize or concatenate linear RNA.

[0371] In some embodiments, linear RNA can include nucleic acids that are converted to 5' monophosphates, for example, by contacting 5' triphosphates with RNA 5' pyrophosphohydrolase (RppH) or ATP diphosphohydrolase (apyrase). The 5' triphosphate. Alternatively, conversion of the 5' triphosphate of linear RNA to 5' monophosphate can be carried out by a two-step reaction comprising: (a) reacting the 5' nucleotide of the linear RNA with a phosphatase (e.g. Antarctic phosphatase, Shrimp alkaline phosphatase or calf intestinal phosphatase) to remove all three phosphates; and (b) subject the 5' nucleotide after step (a) to a monophosphate-adding kinase (e.g. poly nucleotide kinase) contact.

[0372] In some embodiments, the RNA can be circularized using the methods described in WO2017222911 and WO2016197121, the contents of each of which are incorporated herein by reference in their entirety.

[0373] In some embodiments, the RNA can be circularized, eg, by back-splicing a non-mammalian exogenous intron or splinting the 5' and 3' ends of the linear RNA. In one embodiment, a circular RNA is produced from a recombinant nucleic acid encoding a target RNA to be made circular. As a non-limiting example, the method comprises: a) producing a recombinant nucleic acid encoding a target RNA to be made circular, wherein the recombinant nucleic acid comprises, in 5' to 3' order: i) an exogenous nucleic acid comprising a 3' splice site the 3' portion of the intron, ii) the nucleic acid sequence encoding the RNA of interest, and iii) the 5' portion of the exogenous intron that includes the 5' splice site; b) undergoes transcription so that the RNA is produced from the recombinant nucleic acid; and c) RNA splicing whereby the RNA circularizes to produce oRNA.

[0374] While not wishing to be bound by theory, the resulting circular RNA with exogenous introns is recognized by the immune system as "non-self" and triggers an innate immune response. On the other hand, circRNAs produced with endogenous introns are recognized by the immune system as "self" and generally do not stimulate an innate immune response, even if carrying exons containing foreign RNAs.

[0375] Therefore, circRNAs with endogenous or exogenous introns can be generated as needed to control immunological self / non-self distinction. A variety of intron sequences are known from a wide variety of organisms and viruses, and include sequences derived from genes encoding proteins, ribosomal RNA (rRNA) or transfer RNA (tRNA).

[0376] Circular RNAs can be generated from linear RNAs in a variety of ways. In some embodiments, circular RNAs are produced from linear RNAs by back-splicing a downstream 5' splice site (splice donor) to an upstream 3' splice site (splice acceptor). Circular RNAs can be produced in this manner by any non-mammalian splicing method. For example, linear RNAs containing various types of introns, including self-splicing group I introns, self-splicing group II introns, splice introns, and tRNA introns, can be circularized. In particular, group I and II introns are advantageous because they can be readily used to generate circular RNAs in vitro as well as in vivo because they are capable of self-splicing due to their autocatalytic ribonuclease activity.

[0377] In some embodiments, circular RNA can be generated in vitro from linear RNA by chemically or enzymatically ligating the 5' and 3' ends of the RNA. Chemical ligation can be used, for example, using cyanogen bromide (BrCN) or ethyl-3-(3'-dimethylaminopropyl)carbodiimide (EDC) to activate nucleotide phosphate monoester groups to allow phosphate diphosphate ester bond formation to proceed. See for example Sokolova (1988) FEBS Lett 232: 153-155; Dolinnaya et al. (1991) Nucleic Acids Res., 19: 3067-3072; Fedorova (1996) Nucleosides Nucleotides Nucleic Acids 15: 1 137-1 147; way incorporated into this article. Alternatively, enzymatic ligation can be used to circularize RNA. Exemplary ligases that can be used include T4 DNA ligase (T4 Dnl), T4 RNA ligase 1 (T4 Rnl 1), and T4 RNA ligase 2 (T4 Rnl 2).

[0378] In some embodiments, splint ligation using oligonucleotide splints that hybridize to both ends of the linear RNA can be used to join the ends of the linear RNA together. Hybridization of the splint (which can be DNA or RNA) orients the 5'-phosphate and 3'-OH of the RNA ends for ligation. Subsequent ligation can be performed using chemical or enzymatic techniques, as described above. Enzymatic ligation can be performed, for example, with T4 DNA ligase (required for DNA splints), T4 RNA ligase 1 (required for RNA splints), or T4 RNA ligase 2 (for DNA or RNA splints). In some cases, chemical ligation, such as with BrCN or EDC, is more efficient than enzymatic ligation if the structure of the hybridized splint-RNA complex interferes with enzymatic activity.

[0379] In some embodiments, the oRNA may further comprise an internal ribosome entry site (IRES) operably linked to the RNA sequence encoding the polypeptide. Inclusion of an IRES permits translation of one or more open reading frames from the circular RNA. The IRES element attracts the eukaryotic ribosomal translation initiation complex and facilitates translation initiation. See, eg, Kaufman et al., Nuc. Acids Res. (1991) 19:4485-4490; Gurtu et al., Biochem. Biophys. Res. Comm. (1996) 229:295-298; Rees et al., BioTechniques (1996) 20 : 102-110; Kobayashi et al., BioTechniques (1996) 21 :399-402; and Mosser et al., BioTechniques 1997 22 150-161).

[0380] In some embodiments, the cyclization methods provided herein have a cyclization efficiency of at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40% %, at least about 45%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or 100%. In some embodiments, the cyclization methods provided herein have a cyclization efficiency of at least about 40%. [splice element] []

[0381] In some embodiments, the oRNA includes at least one splicing element. The splicing element can be an intact splicing element that can mediate splicing of the oRNA, or the splicing element can be a residual splicing element from a completed splicing event. For example, in some cases, a splicing element of a linear RNA may mediate a splicing event that causes circularization of the linear RNA, such that the resulting oRNA comprises residual splicing elements from such splicing-mediated circularization event. In some cases, residual splicing elements are unable to mediate any splicing. In other cases residual splicing elements can still mediate splicing in some cases. In some embodiments, a splicing element is adjacent to at least one expressed sequence. In some embodiments, the oRNA includes splicing elements adjacent to each expressed sequence. In some embodiments, splicing elements are on one or both sides of each expressed sequence, causing separation of expressed products, eg, peptides and / or polypeptides.

[0382] In some embodiments, the oRNA includes an internal splicing element where the spliced ​​ends join together upon replication. Some examples may include miniintrons (<100 nt) with splice site sequences and short inverted repeats (30-40 nt) such as AluSq2, AluJr, and AluSz, inverted sequences in flanking introns , Alu elements in flanking introns, and motifs found in cis-sequence elements close to back-splicing events (suptable4 enriched motifs), such as in reverse flanking exons Sequence within 200 bp before (upstream) or after (downstream) of the splice site. In some embodiments, the oRNA comprises at least one repeating nucleotide sequence described elsewhere herein as an internal splicing element. In such embodiments, the repetitive nucleotide sequence may comprise a repetitive sequence from an Alu family intron. See, eg, US Patent No. 11,058,706.

[0383] In some embodiments, the oRNA can include typical splice sites flanking head-to-tail junctions of the oRNA.

[0384] In some embodiments, the oRNA may comprise a bump-helix-knob motif comprising a 4 base pair stem flanked by two 3 nucleotide bumps. Cleavage occurs at one site in the knob region, resulting in a characteristic fragment with a terminal 5'-hydroxyl and 2',3'-cyclic phosphate. Cyclization proceeds by nucleophilic attack of a 5'-OH group onto a 2',3'-cyclic phosphate on the same molecule, thereby forming a 3',5'-phosphodiester bridge.

[0385] In some embodiments, the oRNA can include sequences that mediate self-ligation. Non-limiting examples of sequences that can mediate self-ligation include self-circularizing introns (such as 5' and 3' splice junctions (slice junctions)) or self-circularizing catalytic introns (such as group I, group II or group III intron). Non-limiting examples of group I intronic self-splicing sequences may include the self-splicing aligned intron-exon sequence derived from the T4 bacteriophage gene td, and the intermediate sequence (IVS) rRNA of Tetrahymena. [Other cyclization methods] []

[0386] In some embodiments, the linear RNA can include complementary sequences, including repetitive or non-repetitive nucleic acid sequences within individual introns or throughout flanking introns. In some embodiments, the oRNA comprises a repetitive nucleic acid sequence. In some embodiments, the repetitive nucleotide sequence includes polyCA or polyUG sequences. In some embodiments, the oRNA includes at least one repeat nucleic acid sequence that hybridizes to a complementary repeat nucleic acid sequence in another segment of the oRNA, wherein the hybridized segments form an internal duplex. In some embodiments, repeat nucleic acid sequences and complementary repeat nucleic acid sequences from two separate oRNAs are hybridized to produce a single oRNA, wherein the hybridized segments form an internal double-stranded. In some embodiments, complementary sequences are present at the 5' and 3' ends of the linear RNA. In some embodiments, the complementary sequence includes about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more pairs Nucleotides.

[0387] In some embodiments, circularization chemistry can be used to generate oRNA. Such methods may include, but are not limited to, click chemistry (e.g., alkyne and azide based methods or clickable bases), alkene metathesis, phosphoramidate linkage, semiamine aldehyde-imine crosslinking, base modification, and any combination.

[0388] In some embodiments, enzymatic methods of cyclization can be used to generate oRNA. In some embodiments, a ligase, such as a DNA or RNA ligase, can be used to generate an oRNA or a template for a complementary sequence, a complementary strand of an oRNA, or an oRNA. [little distraction] [RNA (siRNA)] []

[0389] In some embodiments, the payload region can be or encodes an RNA interference (RNAi) sequence that can be used to reduce or inhibit the expression of a gene. RNAi (also known as post-transcriptional gene silencing (PTGS), repression or co-suppression) is a method of post-transcriptional gene silencing in which RNA molecules reduce or inhibit gene expression in a sequence-specific manner, usually by disrupting specific mRNA molecules. The active component of RNAi is a short / small double-stranded RNA (dsRNA), called a small interfering RNA (siRNA), which typically contains 15-30 nucleotides (e.g., 19-25, 19-24, or 19-21 nucleotides) and a 2 nucleotide 3' overhang that matches the nucleic acid sequence of the target gene. These short RNA species can be produced naturally in vivo by Dicer-mediated cleavage of larger dsRNAs and are functional in mammalian cells.

[0390] Naturally expressed small RNA molecules, called microRNAs (miRNAs), cause gene silencing by modulating mRNA expression. The RNA-induced silencing complex (RISC) containing miRNAs target mRNAs that exhibit perfect sequence complementarity to nucleotides 2-7 in the 5' region of the miRNA (which is called the seed region), and the other bases to their 3' area for pairing. miRNA-mediated downregulation of gene expression may be caused by cleavage of target mRNA, translational repression of target mRNA, or mRNA breakdown. miRNA targeting sequences are usually located in the 3'-UTR of the target mRNA. A single miRNA can target more than 100 transcripts from various genes, and one mRNA can be targeted by different miRNAs.

[0391] siRNA duplexes or dsRNAs targeting specific mRNAs can be designed and synthesized in vitro, and introduced into cells to activate the RNAi process. It has previously been shown that 21-nucleotide siRNA duplexes, known as small interfering RNAs, are capable of potent and specific attenuation of gene expression in mammalian cells without inducing an immune response. Post-transcriptional gene silencing by siRNA is now rapidly emerging as a powerful tool for gene analysis in mammalian cells, with the potential to generate new therapeutic agents.

[0392] siRNA sequences synthesized in vitro can be introduced into cells to activate RNAi. When introduced into a cell, similar to endogenous dsRNA, exogenous siRNA duplexes can assemble to form the RNA-induced silencing complex (RISC), an RNA sequence that binds to one of the two strands of the siRNA duplex. Complementary (i.e., antisense strand)) interacting multiunit complexes. During this process, the sense strand (or passenger strand) of the siRNA is lost from the complex, while the antisense strand (or guide strand) of the siRNA is matched with its complementary DNA. In particular, RISC complexes containing siRNA target mRNAs that exhibit perfect sequence complementarity. Next, siRNA-mediated gene silencing occurs by cleavage, release and degradation of the target.

[0393] An siRNA duplex consisting of a sense strand homologous to the target mRNA and an antisense strand complementary to the target mRNA provides advantages in target RNA destruction efficiency over the use of single-stranded (ss)-siRNA (e.g., antisense RNA or antisense oligonucleotides) much more. In most cases, a higher concentration of ss-siRNA is required to achieve effective gene silencing performance corresponding to the double helix. [siRNA] [double helix] [Design and sequence] []

[0394] Some guidelines for designing siRNAs have been proposed in the art. These guidelines generally recommend generating a 19-nucleotide duplex region, a symmetrical 2-3 nucleotide 3' overhang, a 5'-phosphate, and targeting the 3'-hydroxyl of the region in the gene to be silenced. Other rules that can govern siRNA sequence bias include, but are not limited to: (i) A / U at the 5' end of the antisense strand; (ii) G / C at the 5' end of the sense strand; (iii) ) at least five A / U residues in the 5' terminal third of the antisense strand; and (iv) the absence of any GC extensions longer than 9 nucleotides. Based on such considerations, together with the specific sequence of the target gene, highly efficient siRNA constructs necessary to inhibit the expression of the target gene in mammals can be readily designed.

[0395] In some embodiments, siRNA constructs (eg, siRNA duplexes or encoded dsRNAs) are designed to target specific genes. Specifically, such siRNA constructs inhibit gene expression and protein production. In some aspects, siRNA constructs are designed and used to selectively "knock out" gene variants in cells, that is, mutant transcripts that are identified in patients or are the cause of various diseases and / or disorders . In some aspects, siRNA constructs are designed and used to selectively "down-repress" variants of a gene in cells. In other aspects, siRNA constructs are capable of inhibiting or suppressing both wild-type and mutant forms of a gene.

[0396] In some embodiments, the siRNA sequence comprises a sense strand and a complementary antisense strand, where the two strands hybridize together to form a double helix. The antisense strand has sufficient complementarity to the mRNA sequence to direct target-specific RNAi, that is, the siRNA sequence has a sequence sufficient to trigger the destruction of the target mRNA by the RNAi mechanism or process.

[0397] In some embodiments, the siRNA sequence comprises a sense strand and a complementary antisense strand, wherein the two strands hybridize together to form a double helix structure, and wherein the initiation site for hybridization to the mRNA is between nucleotides 100 and 100 of the mRNA sequence. Between 10,000. As a non-limiting example, the initiation site may be between: nucleotides 100-150, 150-200, 200-250, 250-300, 300-350, 350-400, 400-450 on the mRNA sequence , 450-500, 500-550, 550-600, 600-650, 650-700, 700-70, 750-800, 800-850, 850-900, 900-950, 950-1000, 1000-1050, 1050 -1100, 1100-1150, 1150-1200, 1200-1250, 1250-1300, 1300-1350, 1350-1400, 1400-1450, 1450-1500, 1500-1550, 1550-1600, 1600-1650, 1650-1 700 , 1700-1750, 1750-1800, 1800-1850, 1850-1900, 1900-1950, 1950-2000, 2000-2050, 2050-2100, 2100-2150, 2150-2200, 2200-2250, 2250-2300, 2300 -2350, 2350-2400, 2400-2450, 2450-2500, 2500-2550, 2550-2600, 2600-2650, 2650-2700, 2700-2750, 2750-2800, 2800-2850, 2850-2900, 2900-2 950 , 2950-3000, 3000-3050, 3050-3100, 3100-3150, 3150-3200, 3200-3250, 3250-3300, 3300-3350, 3350-3400, 3400-3450, 3450-3500, 3500-3550, 3550 -3600, 3600-3650, 3650-3700, 3700-3750, 3750-3800, 3800-3850, 3850-3900, 3900-3950, 3950-4000, 4000-4050, 4050-4100, 4100-4150, 4150-4 200 ,4200-4250,4250-4300,4300-4350,4350-4400,4400-4450,4450-4500,4500-4550,4550-4600,4600-4650,4650-4700,4700-4750,4750-4800, 4800 -4850, 4850-4900, 4900-4950, 4950-5000, 5000-5050, 5050-5100, 5100-5150, 5150-5200, 5200-5250, 5250-5300, 5300-5350, 5350-5400, 5400-5 450 ,5450-5500,5500-5550,5550-5600,5600-5650,5650-5700,5700-5750,5750-5800,5800-5850,5850-5900,5900-5950,5950-6000,6000-6050, 6050 -6100, 6100-6150, 6150-6200, 6200-6250, 6250-6300, 6300-6350, 6350-6400, 6400-6450, 6450-6500, 6500-6550, 6550-6600, 6600-6650, 6650-6 700 , 6700-6750, 6750-6800, 6800-6850, 6850-6900, 6900-6950, 6950-7000, 7000-7050, 7050-7100, 7100-7150, 7150-7200, 7200-7250, 7250-7300 7300 -7350, 7350-7400, 7400-7450, 7450-7500, 7500-7550, 7550-7600, 7600-7650, 7650-7700, 7700-7750, 7750-7800, 7800-7850, 7850-7900, 7900-7 950 , 7950-8000, 8000-8050, 8050-8100, 8100-8150, 8150-8200, 8200-8250, 8250-8300, 8300-8350, 8350-8400, 8400-8450, 8450-8500, 8500-8550, 8550 -8600, 8600-8650, 8650-8700, 8700-8750, 8750-8800, 8800-8850, 8850-8900, 8900-8950, 8950-9000, 9000-9050, 9050-9100, 9100-9150, 9150-9 200 , 9800 -9850, 9850-9900, 9900-9950, 9950-10000.

[0398] In some embodiments, the antisense strand is 100% complementary to the target mRNA sequence. The antisense strand can be complementary to any portion of the target mRNA sequence.

[0399] In other embodiments, the antisense strand and the target mRNA sequence contain at least one mismatch. As a non-limiting example, the antisense strand and the target mRNA sequence have at least 30%, 40%, 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86% , 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99%, or at least 20-30%, 20-40% %, 20-50%, 20-60%, 20-70%, 20-80%, 20-90%, 20-95%, 20-99%, 30-40%, 30-50%, 30-60% %, 30-70%, 30-80%, 30-90%, 30-95%, 30-99%, 40-50%, 40-60%, 40-70%, 40-80%, 40-90% %, 40-95%, 40-99%, 50-60%, 50-70%, 50-80%, 50-90%, 50-95%, 50-99%, 60-70%, 60-80% %, 60-90%, 60-95%, 60-99%, 70-80%, 70-90%, 70-95%, 70-99%, 80-90%, 80-95%, 80-99% %, 90-95%, 90-99%, or 95-99% complementarity.

[0400] In some embodiments, the siRNA sequences are about 10-50 or more nucleotides in length, ie, each strand comprises 10-50 nucleotides (or nucleotide analogs). Preferably, the siRNA sequences are about 15-30 in length in each strand, such as 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides, one of each strand is fully complementary to the target region. In some embodiments, the siRNA sequence is about 19-25, 19-24, or 19-21 nucleotides in length.

[0401] In some embodiments, the siRNA sequence can be a synthetic RNA duplex comprising about 19 nucleotides to about 25 nucleotides and two overhanging nucleotides at the 3' end. In some aspects, siRNA constructs can be unmodified RNA molecules. In other aspects, the siRNA construct may contain at least one modified nucleotide, such as a base, sugar or backbone modification.

[0402] In some embodiments, siRNA sequences can be encoded in plastid vectors, viral vectors, or other nucleic acid expression vectors for delivery to cells. DNA expression plastids can be used to stably express siRNA duplexes or dsRNA in cells and achieve long-term suppression of target gene expression. In one aspect, the sense and antisense strands of an siRNA duplex are typically linked by a short spacer sequence, resulting in a stem-loop structure known as short hairpin RNA (shRNA). The hairpin is recognized and cleaved by Dicer, thereby generating the mature siRNA construct.

[0403] In some embodiments, the sense and antisense strands of an siRNA duplex can be linked by a short spacer sequence, optionally linked with additional flanking sequences, resulting in the expression of a primary microRNA. (pri-miRNA) flanking arm-stem-loop structure. pri-miRNAs can be recognized and cleaved by Drosha and Dicer, and thus generate mature siRNA constructs.

[0404] In some embodiments, the siRNA duplex or encoded dsRNA inhibits (or degrades) a target mRNA. Thus, siRNA duplexes or encoded dsRNAs can be used to substantially inhibit gene expression in cells. In some aspects, inhibition of gene expression means inhibition of at least about 20%, preferably at least about 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, and 100%, or at least 20-30%, 20-40%, 20-50%, 20-60%, 20-70%, 20-80%, 20-90%, 20-95%, 20-100%, 30-40%, 30-50%, 30-60%, 30-70%, 30-80%, 30-90%, 30-95%, 30-100%, 40-50%, 40-60%, 40-70%, 40-80%, 40-90%, 40-95%, 40-100%, 50-60%, 50-70%, 50-80%, 50-90%, 50-95%, 50-100%, 60-70%, 60-80%, 60-90%, 60-95%, 60-100%, 70-80%, 70-90%, 70-95%, 70-100%, 80-90%, 80-95%, 80-100%, 90-95%, 90-100%, or 95-100%. Thus, the protein product of the targeted gene can be inhibited by at least about 20%, such as at least about 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95% and 100%, Or at least 20-30%, 20-40%, 20-50%, 20-60%, 20-70%, 20-80%, 20-90%, 20-95%, 20-100%, 30-40% %, 30-50%, 30-60%, 30-70%, 30-80%, 30-90%, 30-95%, 30-100%, 40-50%, 40-60%, 40-70% %, 40-80%, 40-90%, 40-95%, 40-100%, 50-60%, 50-70%, 50-80%, 50-90%, 50-95%, 50-100% %, 60-70%, 60-80%, 60-90%, 60-95%, 60-100%, 70-80%, 70-90%, 70-95%, 70-100%, 80-90% %, 80-95%, 80-100%, 90-95%, 90-100%, or 95-100%.

[0405] In some embodiments, the siRNA construct comprises a miRNA seed match of the target located in the guide strand. In another embodiment, the siRNA construct comprises a miRNA seed match of the target located in the passenger strand. In yet another embodiment, the siRNA duplex or encoded dsRNA targeting a gene does not contain a seed match for the target located in the guide or passenger strand.

[0406] In some embodiments, the gene-targeting siRNA duplex or encoded dsRNA may have little to no significant full-length off-target to the guide strand. In another example, the siRNA duplex or encoded dsRNA targeting a gene may have little significant full-length off-target effect on the passenger strand. The siRNA duplex or encoded dsRNA targeting a gene can have less than 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11% for the passenger strand , 12%, 13%, 14%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 1-5%, 2-6%, 3-7%, 4 -8%, 5-9%, 5-10%, 6-10%, 5-15%, 5-20%, 5-25%, 5-30%, 10-20%, 10-30%, 10% -40%, 10-50%, 15-30%, 15-40%, 15-45%, 20-40%, 20-50%, 25-50%, 30-40%, 30-50%, 35% -50%, 40-50%, 45-50% full-length off-target effects. In yet another embodiment, the gene-targeting siRNA duplex or encoded dsRNA may have few significant full-length off-targets to the guide or passenger strand. The siRNA duplex or encoded dsRNA targeting a gene can have less than 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 1-5%, 2-6%, 3-7% , 4-8%, 5-9%, 5-10%, 6-10%, 5-15%, 5-20%, 5-25%, 5-30%, 10-20%, 10-30% , 10-40%, 10-50%, 15-30%, 15-40%, 15-45%, 20-40%, 20-50%, 25-50%, 30-40%, 30-50% , 35-50%, 40-50%, 45-50% full-length off-target effects.

[0407] In some embodiments, the gene-targeting siRNA duplex or the encoded dsRNA may have high in vitro activity. In another example, the siRNA construct may have low in vitro activity. In yet another embodiment, the gene-targeting siRNA duplex or dsRNA may have high guide strand activity and low passenger strand activity in vitro.

[0408] In some embodiments, the siRNA constructs have high in vitro guide strand activity and low in vitro passenger strand activity. Knock-down (knock-down; KD) of the target gene can be at least 40%, 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 99.5%, or 100%. The reduction of target gene expression through the guide stock can be 40-50%, 45-50%, 50-55%, 50-60%, 60-65%, 60-70%, 60-75%, 60-80% %, 60-85%, 60-90%, 60-95%, 60-99%, 60-99.5%, 60-100%, 65-70%, 65-75%, 65-80%, 65-85% %, 65-90%, 65-95%, 65-99%, 65-99.5%, 65-100%, 70-75%, 70-80%, 70-85%, 70-90%, 70-95% %, 70-99%, 70-99.5%, 70-100%, 75-80%, 75-85%, 75-90%, 75-95%, 75-99%, 75-99.5%, 75-100 %, 80-85%, 80-90%, 80-95%, 80-99%, 80-99.5%, 80-100%, 85-90%, 85-95%, 85-99%, 85-99.5 %, 85-100%, 90-95%, 90-99%, 90-99.5%, 90-100%, 95-99%, 95-99.5%, 95-100%, 99-99.5%, 99-100 % or 99.5-100%. As a non-limiting example, the knockdown (KD) of target gene expression by the lead stock is greater than 70%. As a non-limiting example, the knockdown (KD) of target gene expression by the lead stock is greater than 60%.

[0409] In some embodiments, the guide-to-passenger (G:P) (also known as antisense-to-sense) strand ratio exhibited in vitro or in vivo is at least 1:10, 1:9, 1:8, 1:1: 7, 1:6, 1:5, 1:4, 1:3, 1:2, 1;1, 2:10, 2:9, 2:8, 2:7, 2:6, 2:5, 2:4, 2:3, 2:2, 2:1, 3:10, 3:9, 3:8, 3:7, 3:6, 3:5, 3:4, 3:3, 3: 2, 3:1, 4:10, 4:9, 4:8, 4:7, 4:6, 4:5, 4:4, 4:3, 4:2, 4:1, 5:10, 5:9, 5:8, 5:7, 5:6, 5:5, 5:4, 5:3, 5:2, 5:1, 6:10, 6:9, 6:8, 6: 7, 6:6, 6:5, 6:4, 6:3, 6:2, 6:1, 7:10, 7:9, 7:8, 7:7, 7:6, 7:5, 7:4, 7:3, 7:2, 7:1, 8:10, 8:9, 8:8, 8:7, 8:6, 8:5, 8:4, 8:3, 8: 2. 8:1, 9:10, 9:9, 9:8, 9:7, 9:6, 9:5, 9:4, 9:3, 9:2, 9:1, 10:10, 10:9, 10:8, 10:7, 10:6, 10:5, 10:4, 10:3, 10:2, 10:1, 1:99, 5:95, 10:90, 15: 85, 20:80, 25:75, 30:70, 35:65, 40:60, 45:55, 50:50, 55:45, 60:40, 65:35, 70:30, 75:25, 80:20, 85:15, 90:10, 95:5, or 99:1. Guide to passenger ratio refers to the ratio of guide strand to passenger strand after intracellular processing of pri-microRNA. For example, an 80:20 boot to passenger ratio would have 8 boot strands for every 2 passenger strands processed from the precursor. As a non-limiting example, the ex vivo guide-to-passenger-share ratio was 8:2. As a non-limiting example, the in vivo guide strand to passenger strand ratio is 8:2. As a non-limiting example, the ex vivo guide-to-passenger-share ratio was 9:1. As a non-limiting example, the in vivo guide strand to passenger strand ratio was 9:1.

[0410] In some embodiments, the represented guide-to-passenger (G:P) (also known as antisense-to-sense) share ratio is greater than 1. In some embodiments, the represented guide-to-passenger (G:P) (also known as antisense-to-sense) share ratio is greater than 2. In some embodiments, the represented guide-to-passenger (G:P) (also known as antisense-to-sense) share ratio is greater than 5. In some embodiments, the represented guide-to-passenger (G:P) (also known as antisense-to-sense) share ratio is greater than 10. In some embodiments, the represented guide-to-passenger (G:P) (also known as antisense-to-sense) share ratio is greater than 20. In some embodiments, the represented guide-to-passenger (G:P) (also known as antisense-to-sense) share ratio is greater than 50. In some embodiments, the represented guide-to-passenger (G:P) (also known as antisense-to-sense) share ratio is at least 3:1. In some embodiments, the expressed guide-to-passenger (G:P) (also known as antisense-to-sense) share ratio is at least 5:1. In some embodiments, the expressed guide-to-passenger (G:P) (also known as antisense-to-sense) share ratio is at least 10:1. In some embodiments, the expressed guide-to-passenger (G:P) (also known as antisense-to-sense) share ratio is at least 20:1. In some embodiments, the expressed guide-to-passenger (G:P) (also known as antisense-to-sense) share ratio is at least 50:1.

[0411] In some embodiments, the in vitro or in vivo expressed passenger to guide (P:G) (also known as sense to antisense) strand ratio is at least 1:10, 1:9, 1:8, 1:7 , 1:6, 1:5, 1:4, 1:3, 1:2, 1;1, 2:10, 2:9, 2:8, 2:7, 2:6, 2:5, 2 :4, 2:3, 2:2, 2:1, 3:10, 3:9, 3:8, 3:7, 3:6, 3:5, 3:4, 3:3, 3:2 , 3:1, 4:10, 4:9, 4:8, 4:7, 4:6, 4:5, 4:4, 4:3, 4:2, 4:1, 5:10, 5 :9, 5:8, 5:7, 5:6, 5:5, 5:4, 5:3, 5:2, 5:1, 6:10, 6:9, 6:8, 6:7 , 6:6, 6:5, 6:4, 6:3, 6:2, 6:1, 7:10, 7:9, 7:8, 7:7, 7:6, 7:5, 7 :4, 7:3, 7:2, 7:1, 8:10, 8:9, 8:8, 8:7, 8:6, 8:5, 8:4, 8:3, 8:2 , 8:1, 9:10, 9:9, 9:8, 9:7, 9:6, 9:5, 9:4, 9:3, 9:2, 9:1, 10:10, 10 :9, 10:8, 10:7, 10:6, 10:5, 10:4, 10:3, 10:2, 10:1, 1:99, 5:95, 10:90, 15:85 , 20:80, 25:75, 30:70, 35:65, 40:60, 45:55, 50:50, 55:45, 60:40, 65:35, 70:30, 75:25, 80 :20, 85:15, 90:10, 95:5, or 99:1. Passenger-to-lead ratio refers to the ratio of passenger shares to lead shares after cutting out the lead shares. For example, an 80:20 passenger to boot ratio would have 8 passenger strands for every 2 boot strands processed from the precursor. As a non-limiting example, the in vitro passenger to guide stock ratio was 80:20. As a non-limiting example, the ratio of in vivo passenger to boot shares was 80:20. As a non-limiting example, the in vitro passenger to guide stock ratio was 8:2. As a non-limiting example, the in vivo passenger to guide strand ratio was 8:2. As a non-limiting example, the in vitro passenger to guide stock ratio was 9:1. As a non-limiting example, the in vivo passenger to guide strand ratio was 9:1.

[0412] In some embodiments, the represented passenger-to-guide (P:G) (also known as sense-to-antisense) strand ratio is greater than 1. In some embodiments, the represented passenger-to-guide (P:G) (also known as sense-to-antisense) strand ratio is greater than 2. In some embodiments, the represented passenger-to-guide (P:G) (also known as sense-to-antisense) strand ratio is greater than 5. In some embodiments, the represented passenger-to-guide (P:G) (also known as sense-to-antisense) strand ratio is greater than 10. In some embodiments, the represented passenger-to-guide (P:G) (also known as sense-to-antisense) strand ratio is greater than 20. In some embodiments, the represented passenger-to-guide (P:G) (also known as sense-to-antisense) strand ratio is greater than 50. In some embodiments, the represented passenger-to-guide (P:G) (also known as sense-to-antisense) strand ratio is at least 3:1. In some embodiments, the represented passenger-to-guide (P:G) (also known as sense-to-antisense) strand ratio is at least 5:1. In some embodiments, the represented passenger-to-guide (P:G) (also known as sense-to-antisense) strand ratio is at least 10:1. In some embodiments, the represented passenger-to-guide (P:G) (also known as sense-to-antisense) strand ratio is at least 20:1. In some embodiments, the represented passenger-to-guide (P:G) (also known as sense-to-antisense) strand ratio is at least 50:1.

[0413] In some embodiments, when measuring processing, the passenger-guide strand duplex occurs when a pri- or pre-microRNA but known in the art and methods described herein exhibit a greater than 2-fold guide-to-passenger ratio. deemed valid. As a non-limiting example, when measuring processing, pri- or pre-microRNA exhibit greater than 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold , 12 times, 13 times, 14 times, 15 times or 2 to 5 times, 2 to 10 times, 2 to 15 times, 3 to 5 times, 3 to 10 times, 3 to 15 times, 4 to 5 times, 4 to 10 times, 4 to 15 times, 5 to 10 times, 5 to 15 times, 6 to 10 times, 6 to 15 times, 7 to 10 times, 7 to 15 times, 8 to 10 times, 8 to 15 times, 9 to 10x, 9x to 15x, 10x to 15x, 11x to 15x, 12x to 15x, 13x to 15x or 14x to 15x lead-to-passenger ratio.

[0414] In some embodiments, the dsRNA-encoding vector gene body comprises at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or more than 99% of the full length of the construct. The sequence of %. As a non-limiting example, the vector genome comprises a sequence that is at least 80% of the full-length sequence of the construct.

[0415] In some embodiments, siRNA constructs can be used to silence wild-type or mutant genes by targeting at least one exon on the sequence. [siRNA] [modification] []

[0416] In some embodiments, siRNA constructs, when not delivered as precursors or DNA, can be chemically modified to modulate some characteristic of the RNA molecule, such as, but not limited to, increasing the stability of the siRNA in vivo. The chemically modified siRNA constructs can be used for therapeutic applications in humans and are modified without compromising the RNAi activity of the siRNA constructs. As a non-limiting example, siRNA constructs are modified at both the 3' and 5' ends of the sense and antisense strands.

[0417] In some embodiments, modified nucleotides may be on the sense strand only.

[0418] In some embodiments, the modified nucleotides may be only on the antisense strand.

[0419] In some embodiments, modified nucleotides can be in both the sense and antisense.

[0420] In some embodiments, chemically modified nucleotides do not affect the ability of the antisense strand to pair with the target mRNA sequence. [tiny] [RNA (miR)] [skeleton] []

[0421] In some embodiments, the siRNA construct can be encoded in a polynucleotide sequence that also includes a microRNA (miR) backbone construct. As used herein, a "microRNA (miR) backbone construct" is a framework or starting molecule that forms the sequence or structural basis for designing or manufacturing subsequent molecules.

[0422] In some embodiments, the miR backbone construct comprises at least one 5' flanking region. As a non-limiting example, the 5' flanking region may comprise a 5' flanking sequence that may be of any length and may be derived in whole or in part from a wild-type microRNA sequence or entirely artificial.

[0423] In some embodiments, the miR backbone construct comprises at least one 3' flanking region. As a non-limiting example, the 3' flanking region may comprise a 3' flanking sequence that may be of any length and may be derived in whole or in part from a wild-type microRNA sequence or entirely artificial.

[0424] In some embodiments, the miR backbone construct comprises at least one loop motif region. As a non-limiting example, a loop motif region may comprise a sequence which may be of any length.

[0425] In some embodiments, the miR backbone construct comprises a 5' flanking region, a loop motif region and / or a 3' flanking region.

[0426] In some embodiments, at least one payload (eg, siRNA, miRNA, or other RNAi agent described herein) can be encoded by a polynucleotide that can also comprise at least one miR backbone construct. A miR backbone construct may comprise a 5' flanking sequence that may be of any length and may be derived in whole or in part from a wild-type microRNA sequence or entirely artificial. The 3' flanking sequence may mirror the 5' flanking sequence and / or the 3' flanking sequence in size and origin. There may be no flanking sequences of any kind. The 3' flanking sequence optionally contains one or more CNNC motifs, where "N" represents any nucleotide.

[0427] In some embodiments, the 5' arm of the stem-loop structure of a polynucleotide comprising or encoding a miR backbone construct comprises a sequence encoding a sense sequence.

[0428] In some embodiments, the 3' arm of the stem-loop of a polynucleotide comprising or encoding a miR backbone construct comprises a sequence encoding an antisense sequence. In some instances, the antisense sequence comprises a "G" nucleotide at the most 5' end.

[0429] In some embodiments, the sense sequence can reside on the 3' arm of the stem-loop structure of the polynucleotide comprising or encoding the miR backbone construct, while the antisense sequence resides on the 5' arm thereof.

[0430] In some embodiments, the sense and antisense sequences can be perfectly complementary for most of their lengths. In other embodiments, the sense and antisense sequences can be at least 70, 80, 90, 95 or 95% of the length of each strand independently at least 50, 60, 70, 80, 85, 90, 95 or 99% of the length of each strand. 99% complementary.

[0431] Neither sense sequence identity nor antisense sequence homology need be 100% complementary to the target sequence.

[0432] In some embodiments, separating the sense and antisense sequences of the stem-loop structure of a polynucleotide is a loop sequence (also known as a loop motif, linker or linker motif). The loop sequence can be of any length: between 4-30 nucleotides, between 4-20 nucleotides, between 4-15 nucleotides, between 5-15 nucleotides , between 6-12 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides nucleotides, 13 nucleotides, 14 nucleotides and / or 15 nucleotides.

[0433] In some embodiments, the loop sequence comprises a nucleic acid sequence encoding at least one UGUG motif. In some embodiments, the nucleic acid sequence encoding the UGUG motif is located 5' to the loop sequence.

[0434] In some embodiments, a spacer region may be present in a polynucleotide to bind one or more modules (e.g., 5' flanking region, loop motif region, 3' flanking region, sense sequence, reverse sequence) are separated from each other. There may be one or more such spacer regions.

[0435] In some embodiments, there is a spacer region of between 8 and 20, i.e., 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides It may be present between the sense sequence and the flanking region sequences.

[0436] In some embodiments, the spacer region is 13 nucleotides in length and is located between the 5' end of the sense sequence and the 3' end of the flanking sequence. In some embodiments, the spacer is of sufficient length to form approximately one helical turn of the sequence.

[0437] In some embodiments, there is a spacer region of between 8-20, i.e., 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides It may be present between the antisense sequence and the flanking sequence.

[0438] In some embodiments, the spacer sequence is between 10-13, i.e., 10, 11, 12 or 13 nucleotides, and is located between the 3' end of the antisense sequence and the 5' end of the flanking sequence . In some embodiments, the spacer is of sufficient length to form approximately one helical turn of the sequence.

[0439] In some embodiments, the polynucleotide comprises a 5' flanking sequence, a 5' arm, a loop motif, a 3' arm, and a 3' flanking sequence in the 5' to 3' direction. As a non-limiting example, the 5' arm may contain a sense sequence and the 3' arm an antisense sequence. In another non-limiting embodiment, the 5' arm comprises an antisense sequence and the 3' arm comprises a sense sequence.

[0440] In some embodiments, the 5' arm, payload (e.g., sense sequence and / or antisense sequence), loop motif, and / or 3' arm sequence can be altered (e.g., substitution of 1 or more nucleotides, addition of core nucleotides and / or missing nucleotides). Alterations can result in beneficial changes in the function of the construct (eg, increased attenuation of gene expression of the target sequence, reduced degradation of the construct, reduced off-target effects, increased efficiency of the payload, and reduced degradation of the payload).

[0441] In some embodiments, miR backbone constructs of polynucleotides are aligned in order to achieve a greater resection rate of the guide strand than the passenger strand. The excise rate of guide or passenger shares can be independently 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or more than 99%. As a non-limiting example, the resection rate of the leading strand is at least 80%. As another non-limiting example, the resection rate of the lead strand is at least 90%.

[0442] In some embodiments, the resection rate of the lead strand is greater than the resection rate of the passenger strand. In one aspect, the resection rate of the lead strand may be at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40% greater than that of the passenger strand %, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or more than 99%.

[0443] In some embodiments, the resection efficiency of the leading strand is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or more than 99%. As a non-limiting example, the resection efficiency of the guide strand is greater than 80%.

[0444] In some embodiments, the efficiency of ablation of the guide strand is greater than that of the passenger strand from the miR scaffold. Resection of the guide strand can be 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 times more efficient than resection of the passenger strand from the miR backbone construct.

[0445] In some embodiments, the miR backbone construct comprises a bifunctional targeting polynucleotide. As used herein, a "bifunctional targeting" polynucleotide is a polynucleotide in which the guide and passenger strand gene expression attenuation targets the same target or the guide and passenger strand gene attenuation targets are different.

[0446] In some embodiments, the miR backbone constructs of the polynucleotides described herein can comprise a 5' flanking region, a loop motif region, and a 3' flanking region.

[0447] In some embodiments, the polynucleotide is designed using at least one of the following properties: loop variant, seed mismatch / knob / wobble variant, stem mismatch, loop variant, and vassal stem mismatch Assortment variants, seed mismatch and basal stem mismatch variants, stem mismatch and basal stem mismatch variants, seed wobble and basal stem wobble variants, or stem sequence variants.

[0448] In some embodiments, the miR backbone construct can be a native pri-miRNA backbone.

[0449] In some embodiments, the choice of miR backbone construct is determined by comparing polynucleotides in pri-miRNAs.

[0450] In some embodiments, the choice of miR backbone construct is determined by comparing polynucleotides in natural pri-miRNAs and synthetic pri-miRNAs. [transfer] [RNA] [(tRNA)] []

[0451] Transfer RNA (tRNA) is an RNA molecule that translates mRNA into protein. A tRNA includes a cloverleaf structure comprising a 3' acceptor site, a 5' phosphate, a D arm, a T arm, and an anticodon arm. The main purpose of tRNA is to carry the amino acid on its 3' acceptor site to Ribosome complex. Once an amino acid is bound to a tRNA, the tRNA is considered an aminoacyl-tRNA. The type of amino acid on the tRNA depends on the mRNA codon. The anticodon arm of the tRNA is the site of the anticodon, which is complementary to the mRNA codon and specifies the amino acid to be carried. tRNAs are also known to play a role in regulating apoptosis by acting as cytochrome c scavengers.

[0452] In some embodiments, a starting construct and / or a reference construct comprises or encodes a tRNA. [Ribosome] [RNA] [(rRNA)] []

[0453] Ribosomal RNA (rRNA) is RNA that forms ribosomes. Ribosomes are essential for protein synthesis and contain large and small ribosomal subunits. In prokaryotes, the small 30S and large 50S ribosomal subunits constitute the 70S ribosome. In eukaryotes, 40S and 60S subunits form 80S ribosomes. In order to bind amino-tRNA and link amino acids together to produce polypeptides, ribosomes contain 3 sites: exit site (E), peptidyl site (P) and acceptor site (A) .

[0454] In some embodiments, the starting construct and / or the reference construct comprises or encodes rRNA. [tiny] [RNA (miRNA)] []

[0455] MicroRNA (or miRNA) is a 19-25 nucleotide long non-coding RNA that binds to the 3'UTR of a nucleic acid molecule and down-regulates gene expression by reducing the stability of the nucleic acid molecule or inhibiting translation. An initial construct and / or a reference construct can comprise one or more microRNA target sequences, microRNA sequences, or microRNA seeds.

[0456] The microRNA sequence comprises a "seed" region, ie, a sequence in the region of positions 2-8 of the mature microRNA, which has perfect Watson-Crick complementarity to the miRNA target sequence. A microRNA seed can comprise positions 2-8 or 2-7 of the mature microRNA. In some embodiments, a microRNA seed can comprise 7 nucleotides (e.g., nucleotides 2-8 of a mature microRNA), where the corresponding seed complementary site in the miRNA target consists of an adenine opposite position 1 of the microRNA (A) Side connection. In some embodiments, a microRNA seed may comprise 6 nucleotides (e.g., nucleotides 2-7 of a mature microRNA), where the corresponding seed complementary site in the miRNA target consists of an adenine opposite position 1 of the microRNA (A) Side connection. The bases of the microRNA seed are completely complementary to the target sequence. By engineering the microRNA target sequence into the 3'UTR of the mRNA, we can target molecules for degradation or reduced translation, provided that the microRNA in question is available. This process will reduce the risk of off-target effects in the delivery of nucleic acid molecules.

[0457] As used herein, the term "microRNA site" refers to a microRNA target site or a microRNA recognition site, or any nucleotide sequence that binds / associates with a microRNA. It is understood that "binding" may follow the traditional Watson-Crick hybridization rules, or may reflect any stable binding of the microRNA to the target sequence at or near the microRNA site.

[0458] Non-limiting examples of tissues in which microRNAs are known to regulate mRNA, and thus protein expression, include, but are not limited to, liver (miR-122), muscle (miR-133, miR-206, miR-208), endothelial cells (miR-17 -92, miR-126), bone marrow cells (miR-142-3p, miR-142-5p, miR-16, miR-21, miR-223, miR-24, miR-27), adipose tissue (let-7 , miR-30c), heart (miR-1d, miR-149), kidney (miR-192, miR-194, miR-204) and lung epithelial cells (let-7, miR-133, miR-126). MicroRNAs can also regulate complex biological processes, such as angiogenesis (miR-132).

[0459] For example, if the nucleic acid molecule is mRNA and is not intended to be delivered to the liver but ends up there, miR-122, a microRNA that is abundant in the liver, can be passed at one or more target sites of miR-122. Suppresses the expression of a gene of interest when engineered into the 3'UTR of an mRNA. The introduction of one or more binding sites for different microRNAs can be engineered to further reduce mRNA longevity, stability, and protein translation.

[0460] Conversely, microRNA binding sites can be engineered outside (ie, removed from) the sequence in which they naturally occur in order to increase protein expression in a particular tissue. For example, miR-122 binding sites can be removed to increase protein expression in the liver. Modulation of expression in multiple tissues can be achieved by introducing or removing one or several microRNA binding sites. [long non-encoded] [RNA (lncRNA)] []

[0461] Long non-coding RNAs (lncRNAs) are regulatory RNA molecules that do not code for proteins but affect a large number of biological processes. lncRNA names are generally restricted to non-coding transcripts longer than about 200 nucleotides. Length designations distinguish lncRNAs from small regulatory RNAs, such as short interfering RNAs (siRNAs) and microRNAs (miRNAs). In vertebrates, the number of lncRNA species is thought to greatly exceed the number of protein-coding species. It is also believed that lncRNAs drive the biological complexity observed in vertebrates compared to invertebrates. Evidence for this complexity can be found in many cellular compartments of vertebrate organisms, such as the T lymphocyte compartment of the adaptive immune system. Differences in the expression and function of lncRNAs may be the main cause of human diseases.

[0462] In some embodiments, the initial constructs and / or reference constructs comprise lncRNAs. [RNA] [modification] []

[0463] In some aspects, an initial or reference construct may contain one or more modified nucleotides, such as, but not limited to, sugar-modified nucleotides, nucleobase modifications, and / or backbone modifications. In some aspects, an initial or reference construct may contain combined modifications, such as combined nucleobase and backbone modifications.

[0464] In some embodiments, the modified nucleotides may be sugar-modified nucleotides. Sugar-modified nucleotides include, but are not limited to, 2'-fluoro, 2'-amine and 2'-thio-modified ribonucleotides, such as 2'-fluoro-modified ribonucleotides. Modified nucleotides can be modified on the sugar moiety, as can nucleotides with sugars or analogs thereof that are not ribosyl. For example, the sugar moiety can be or be based on mannose, arabinose, glucopyranose, galactopyranose, 4'-thioribose and other sugars, heterocycles or carbocycles.

[0465] In some embodiments, the modified nucleotides may be nucleobase modified nucleotides.

[0466] In some embodiments, the modified nucleotides may be backbone modified nucleotides. In some embodiments, the initial construct or reference construct may further comprise other modifications on the backbone. As used herein, the general "backbone" refers to the repeating sequence of alternating sugar phosphates in a DNA or RNA molecule. Deoxyribose / ribose is joined to the phosphate group at the 3'-hydroxyl and 5'-hydroxyl groups by an ester bond (also known as a "phosphodiester" bond / linker (PO linkage)). The PO backbone can be modified as a "phosphorothioate backbone (PS linkage)". In some cases, the natural phosphodiester linkage can be replaced by an amide linkage, but maintaining the four atoms between the two sugar units. Such amide modifications can facilitate solid-phase synthesis of oligonucleotides and increase the thermodynamic stability of the duplex formed with the siRNA complementary sequence.

[0467] Modified bases refer to nucleotide bases such as, but not limited to, adenine, guanine, cytosine, thymine, uracil, xanthine, inosine, and Q nucleoside, which have been modified by substitution or addition of a or multiple atoms or groups are modified. Some examples of modifications on nucleobase moieties include, but are not limited to, alkylated, halogenated, thiolated, aminated, amidated, or acetylated bases, individually or in combination. More specific examples include, for example, 5-propynyluridine, 5-propynylcytidine, 6-methyladenine, 6-methylguanine, N,N,-dimethyladenine, 2-propynyl Base adenine, 2-propylguanine, 2-amino adenine, 1-methylinosine, 3-methyluridine, 5-methylcytidine, 5-methyluridine and have Modified other nucleotides, 5-(2-amino)propyluridine, 5-halocytidine, 5-halouridine, 4-acetylcytidine, 1-methyladenosine, 2-methano Adenosine, 3-methylcytidine, 6-methyluridine, 2-methylguanosine, 7-methylguanosine, 2,2-dimethylguanosine, 5-methylaminoethylurea glycosides, 5-methoxyuridine, deazanucleotides (such as 7-deaza-adenosine, 6-azouridine, 6-azocytidine, 6-azothymidine), 5-methyl-2-thiouridine, other thiobases (such as 2-thiouridine and 4-thiouridine and 2-thiocytidine), dihydrouridine, pseudouridine, Q nucleosides, guanine, naphthyl and substituted naphthyl, any O-alkylated purines and pyrimidines and N-alkylated purines and pyrimidines (such as N6-methyladenosine, 5-methylcarbonylmethyl uridine, uridine 5-oxyacetic acid, pyridin-4-one, pyridin-2-one), phenyl and modified phenyl (such as aminophenol or 2,4,6-trimethoxybenzene) , modified cytosines (which act as G-clamp nucleotides), 8-substituted adenines and guanines, 5-substituted uracils and thymines, azapyrimidines, carboxyhydroxyalkyl nucleotides, carboxy Alkylamino nucleotides and alkylcarbonyl alkylated nucleotides.

[0468] Primary constructs and / or benchmark constructs that may include one or more substitutions, insertions and / or additions, deletions and covalent modifications relative to the reference sequence (in particular, the parental RNA) are included within the scope of the present invention .

[0469] In some embodiments, the initial and / or reference constructs include one or more post-transcriptional modifications (e.g., capping, cleavage, polyadenylation, splicing, polyA sequences, methylation, acylation, phosphorylation , methylation and acetylation of lysine and arginine residues, and nitrosylation of thiol and tyrosine residues, etc.). The one or more post-transcriptional modifications can be any post-transcriptional modification, such as any of over a hundred different nucleoside modifications that have been identified in RNA (Rozenski, J, Crain, P and McCloskey, J.( 1999). The RNA Modification Database: 1999 update. Nucl Acids Res 27: 196-197). In some embodiments, the first isolated nucleic acid comprises messenger RNA (mRNA). In some embodiments, the starting construct and / or the reference construct comprises at least one nucleoside selected from the group consisting of: pyridin-4-ketoribonucleoside, 5-aza-uridine, 2-thio- 5-Aza-uridine, 2-thiouridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl Base-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-taurine methyluridine, 1-taurine methyl base-pseudouridine, 5-taurine methyl-2-thio-uridine, 1-taurine methyl-4-thio-uridine, 5-methyl-uridine, 1-methyl -Pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio Base-1-methyl-1-deaza-pseudouridine, dihydrouridine, dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2- Methoxyuridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, and 4-methoxy-2-thio-pseudouridine. In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of 5-azacytidine, pseudoisocytidine, 3-methyl-cytidine, N4-acetylcytidine, 5 -Formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudoisocytidine, pyrrolo-cytidine, pyrrolo-pseudoisocytidine, 2-thio -cytidine, 2-thio-5-methyl-cytidine, 4-thio-pseudo-cytidine, 4-thio-1-methyl-pseudo-cytidine, 4-thio-1-methyl Base-1-deaza-pseudoisocytidine, 1-methyl-1-deaza-pseudoisocytidine, zebularine, 5-aza-zebraline, 5-methyl- Zebraline, 5-aza-2-thio-zebraline, 2-thio-zebraline, 2-methoxy-cytidine, 2-methoxy-5-methyl- Cytidine, 4-methoxy-pseudoisocytidine, and 4-methoxy-1-methyl-pseudoisocytidine. In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8- Aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza Aza-8-aza-2,6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-prenyladenosine, N6-(cis-hydroxyprenyl ) adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl) adenosine, N6-glycylaminoformyl adenosine, N6-threonylaminoformyl adenosine , 2-methylthio-N6-threonylaminoformyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine and 2-methoxy base - adenine. In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of inosine, 1-methyl-inosine, wyosine, wybutosine, 7-deaza -guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8- Aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio- Guanosine, N2-methyl-6-thio-guanosine and N2,N2-dimethyl-6-thio-guanosine.

[0470] The initial and / or reference constructs may include any applicable modifications such as sugars, nucleobases, or internucleoside linkages (eg, linking phosphate / phosphodiester linkages / phosphodiester backbone). One or more atoms of the pyrimidine nucleobase may be optionally substituted amine, optionally substituted thiol, optionally substituted alkyl (e.g., methyl or ethyl), or halo (e.g., chloro or Fluorine) replacement or substitution. In certain embodiments, the modification (eg, one or more modifications) is present in each of the sugar and the internucleoside linkage. The modification may be ribonucleic acid (RNA) to deoxyribonucleic acid (DNA), threose nucleic acid (TNA), diol nucleic acid (GNA), peptide nucleic acid (PNA), locked nucleic acid (LNA) or mixtures thereof. Additional modifications are described herein.

[0471] In some embodiments, the initial and / or reference constructs include at least one N(6)methyladenosine (m6A) modification that increases translation efficiency. In some embodiments, the N(6)methyladenosine (m6A) modification reduces the immunogenicity of the initial construct and / or the reference construct.

[0472] In some embodiments, modifications may include chemically or cell-induced modifications. For example, some non-limiting examples of intracellular RNA modifications are given by Lewis and Pan in "RNA modifications and structures cooperate to guide RNA-protein interactions", Nat. Reviews Mol. Cell Biol., 2017, 18:202-210 describe.

[0473] In some embodiments, chemical modification of RNA enhances immune evasion. RNA can be synthesized and / or modified by methods well established in the art, such as those described in: "Current protocols in nuclear acid chemistry", Beaucage, S. L. et al. (eds.), John Wiley & Sons , Inc., New York, N.Y., USA, which is incorporated herein by reference. Modifications include, for example, terminal modifications, such as 5' end modifications (phosphorylation (mono, di, and tri), binding, reverse linkage, etc.), 3' end modifications (binding, DNA nucleotides, reverse linkage, etc.), Base modification (eg, base substitution by a stabilized base, an unstable base, or a base pair with an expanded library of partners), removal of bases (abasic nucleotides), or incorporation of bases. Modified ribonucleotide bases may also include 5-methylcytidine and pseudouridine. In some embodiments, base modifications can modulate RNA expression, immune response, stability, subcellular localization, to name a few functional roles. In some embodiments, modifications include diorthogonal nucleotides, such as unnatural bases. See, eg, Kimoto et al., Chem Commun (Camb), 2017, 53:12309, DOI: 10.1039 / c7cc06661a, which is incorporated herein by reference.

[0474] In some embodiments, one or more RNA sugar modifications (eg, at the 2' position or 4' position) or sugar substitutions and backbone modifications can include modifications or substitutions of phosphodiester linkages. Specific examples of modifications include modified backbones or absence of natural internucleoside linkages, such as internucleoside modifications, including modification or replacement of phosphodiester linkages. RNAs with modified backbones especially include those that do not have phosphorus atoms in the backbone. For purposes of this application, and as sometimes referred to in the art, modified RNAs that do not have a phosphorus atom in their internucleoside backbone may also be considered oligonucleotides. In particular embodiments, RNA will include ribonucleotides with phosphorus atoms in their internucleoside backbone.

[0475] Modified RNA backbones can include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methyl and other alkyl phosphonates ( such as 3'-alkylene phosphonates), and chiral phosphonates, phosphonites, phosphoramidates (such as 3'-aminophosphoramidates and aminoalkylphosphoramidates) , thiocarbonyl phosphoramidate, thiocarbonyl alkyl phosphonate, thiocarbonyl alkyl phosphonate triester, borane phosphate with normal 3'-5' linkage, 2'-5' linked analogues of these , and those of reverse polarity (wherein adjacent pairs of nucleoside units are linked 3'-5' to 5'-3' or 2'-5' to 5'-2'). Also included are various salts, mixed salts and free acid forms. In some embodiments, RNA can be negatively or positively charged.

[0476] Modified nucleotides may be modified at internucleoside linkages (eg, phosphate backbone). Herein, the phrases "phosphate ester" and "phosphodiester" are used interchangeably in the context of polynucleotide backbones. The backbone phosphate groups can be modified by replacing one or more oxygen atoms with different substituents. In addition, modified nucleosides and nucleotides may comprise the bulk replacement of an unmodified phosphate moiety with another internucleoside linkage as described herein. Examples of modified phosphate groups include, but are not limited to, phosphorothioate, phosphoroselenoate, boranophosphate, boranophosphate ester, hydrophosphonate, phosphoramidate, Diamino phosphates, alkyl or aryl phosphonates and phosphate triesters. Both non-linking oxygens of the phosphorodithioate are replaced by sulfur. Phosphate linkers can also be modified by replacing the linking oxygen with nitrogen (bridged phosphoramidate), sulfur (bridged phosphorothioate), and carbon (bridged methylene-phosphonate).

[0477] The a-thio substituted phosphate moieties are provided to impart stability to RNA and DNA polymers via non-natural phosphorothioate backbone linkages. Phosphorothioate DNA and RNA have increased nuclease resistance and consequently longer half-lives in the cellular environment. Phosphorothioates linked to RNA are expected to reduce the innate immune response via weaker binding / activation of cellular innate immune molecules.

[0478] In particular embodiments, the modified nucleosides include α-thio-nucleosides (e.g., 5'-O-(1-phosphorothioate)-adenosine, 5'-O-(1-phosphorothioate) )-cytidine (a-thio-cytidine), 5'-O-(1-phosphorothioate)-guanosine, 5'-O-(1-phosphorothioate)-uridine or 5' -O-(1-phosphorothioate)-pseudouridine).

[0479] Other internucleoside linkages tha...

Claims

1. A compound of formula (CY-VI') or a pharmaceutically acceptable salt thereof, (CY-VI'), wherein: R1 is selected from the group consisting of: -OH, -OAc, R1a, and; Z1 is a substituted C1-C6 alkyl group, as appropriate; X1 is a substituted C2-C6 alkenyl group, as appropriate; X2 is -CH2CH2-; X4 and X5 are independently substituted C2-C14 alkenyl groups or substituted C2-C14 alkenyl groups, as appropriate; Y1 and Y2 are independently selected from the group consisting of: , , , , , , , , , , , or; wherein the bond marked with "*" is connected to X4 or X5; each Z2 is independently H or a substituted C1-C8 alkyl group, as appropriate; each Z3 is independently substituted C1-C6 alkenyl group, as appropriate; R1a is: , , or; R2a, R2b and R2c are independently hydrogen and C1-C6 alkyl groups; R3a, R3b and R3c are independently hydrogen and C1-C6 alkyl groups; R4a, R4b and R4c are independently hydrogen and C1-C6 alkyl; R5a, R5b and R5c are independently hydrogen and C1-C6 alkyl; R6, R7, R8 and R9 are independently, as appropriate, substituted C1-C14 alkyl, as appropriate, substituted C2-C14 alkenyl or -(CH2)mA-(CH2)nH; each A is independently C3-C8 cycloalkyl; each m is independently 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12; and each n is independently 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12.

2. The compound of claim 1 or its pharmaceutically acceptable salt, wherein Y1 is or and Y2 is or.

3. The compound of claim 1 or its pharmaceutically acceptable salt, wherein R1 is -OH.

4. The compound of claim 1 or a pharmaceutically acceptable salt thereof, wherein R1 is...

5. The compound of claim 1 or a pharmaceutically acceptable salt thereof, wherein X1 is a C2-C4 alkyl group.

6. The compound of claim 1 or a pharmaceutically acceptable salt thereof, wherein X4 is, as appropriate, a substituted C2-C6 alkyl group, and X5 is, as appropriate, a substituted C2-C6 alkyl group.

7. The compound of claim 1 or its pharmaceutically acceptable salt, wherein Y1 and Y2 are both -CH2CH2- and Z3 is -CH2CH2-.

8. The compound of claim 1 or a pharmaceutically acceptable salt thereof, wherein: R1 is -OH; X1 is a C2-C6 alkyl group; X2 is -CH2CH2-; X4 and X5 are independently C2-C6 alkyl groups; Y1 and Y2 are; wherein the bond marked with "*" is connected to X4 or X5; each Z3 is -CH2CH2-; R6, R7, R8 and R9 are independently substituted C1-C14 alkyl groups, substituted C2-C14 alkenyl groups or -(CH2)mA-(CH2)nH, as appropriate; A is a C3-C8 cycloalkyl group; each m is independently 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12; and each n is independently 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12.

9. The compound of claim 1 or its pharmaceutically acceptable salt is selected from the group consisting of: or its pharmaceutically acceptable salt.

10. A lipid nanoparticle (LNP) comprising an ionizable lipid having the structure of formula (CY-VI') or a medically acceptable salt thereof: (CY-VI'), wherein: R1 is selected from the group consisting of: -OH, -OAc, R1a, and; Z1 is a substituted C1-C6 alkyl group, as appropriate; X1 is a substituted C2-C6 alkenyl group, as appropriate; X2 is -CH2CH2-; X4 and X5 are independently substituted C2-C14 alkenyl groups or substituted C2-C14 alkenyl groups, as appropriate; Y1 and Y2 are independently selected from the group consisting of: , , , , , , , , , , , or; wherein the bond marked with "*" is connected to X4 or X5; each Z2 is independently H or a substituted C1-C8 alkyl group, as appropriate; each Z3 is independently substituted C1-C6 alkenyl group, as appropriate; R1a is: , , or; R2a, R2b and R2c are independently hydrogen and C1-C6 alkyl groups; R3a, R3b and R3c are independently hydrogen and C1-C6 alkyl groups; R4a, R4b and R4c are independently hydrogen and C1-C6 alkyl; R5a, R5b and R5c are independently hydrogen and C1-C6 alkyl; R6, R7, R8 and R9 are independently, as appropriate, substituted C1-C14 alkyl, as appropriate, substituted C2-C14 alkenyl or -(CH2)mA-(CH2)nH; each A is independently C3-C8 cycloalkyl; each m is independently 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12; and each n is independently 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12.

11. The LNP of claim 10, wherein the ionizable lipid of formula (CY-VI') is selected from the group consisting of: or a pharmaceutically acceptable salt thereof.

12. The LNP of claim 10 further comprises: (a) PEG-lipids; (b) structural lipids; and (c) non-ionizable lipids and / or zwitterionic lipids.

13. The LNP of claim 12, wherein the PEG-lipid is selected from the group consisting of: PEG-c-DOMG, PEG-DMG, PEG-DLPE, PEG-DMPE, PEG-DPPC and PEG-DSPE.

14. The LNP of claim 12, wherein the structural lipid is selected from the group consisting of: cholesterol, fecosterol, sitosterol, ergosterol, campesterol, stigmasterol, brassicasterol, tomatidine, ursolic acid, and α-tocopherol.

15. The LNP of claim 12, wherein the non-ionizable lipid is a phospholipid selected from the group consisting of: 1,2-distearyl-sn-glycero-3-phosphocholine (DSPC), 1,2-dioleyl-sn-glycero-3-phosphoethanolamine (DOPE), 1,2-dilinoleyl-sn-glycero-3-phosphocholine (DLPC), 1,2-dimyristyl-sn-glycero-3-phosphocholine (DMPC), 1,2-dioleyl-sn-glycero-3-phosphocholine (DOPC), 1,2- Di-palmaclopyr-sn-glycero-3-phosphocholine (DPPC), 1,2-diundecanoyl-sn-glycero-3-phosphocholine (DUPC), 1-palmaclopyr-2-oleyl-sn-glycero-3-phosphocholine (POPC), 1,2-di-O-octadecenyl-sn-glycero-3-phosphocholine (18:0 diether PC), 1-oleyl-2-cholesterolylhemisuccinyl-sn-glycero-3-phosphocholine (OChemsPC), 1-hexadecyl-sn-glycero-3-phosphocholine (C16) Lyso PC), 1,2-dilinalinyl-sn-glycero-3-phosphate choline, 1,2-diarachidonicyl-sn-glycero-3-phosphate choline, 1,2-bis(docosahexaenoyl)-sn-glycero-3-phosphate choline, 1,2-diphyliyl-sn-glycero-3-phosphate ethanolamine (ME 16.0) PE), 1,2-distearyl-sn-glycero-3-phosphate ethanolamine, 1,2-dilinoleyl-sn-glycero-3-phosphate ethanolamine, 1,2-dilinoleyl-sn-glycero-3-phosphate ethanolamine, 1,2-diarachidonicyl-sn-glycero-3-phosphate ethanolamine, 1,2-bis(docosahexaenoyl)-sn-glycero-3-phosphate ethanolamine, 1,2-dilinoleyl-sn-glycero-3-phosphate-rac-(1-glycerol) sodium salt (DOPG), (S)-2-ammonium-3-((((R)-2-(oleyloxy)-3-(stearyloxy)propoxy)oxyphosphatidyl)oxy)propionate sodium salt (L-α-phosphatidylinosine);Brain PS), dimyristylphosphatidylcholine (DMPC), dimyristylphosphatidylethanolamine (DMPE), dimyristylphosphatidylglycerol (DMPG), dioleoylphosphatidylethanolamine 4-(N-cis-butenediaminomethyl)cyclohexane-1-carboxylate (DOPE-mal), dioleoylphosphatidylglycerol (DOPG), 1,2-dioleoyl-sn-glycero-3-(phospho-L-serine) (DOPS), acell-fusogenic phospholipids. (DPhPE), Dipalmitoylphosphatidylethanolamine (DPPE), Dipalmitoylphosphatidylglycerol (DPPG), Dipalmitoylphosphatidylserine (DPPS), Distearatelphosphatidylcholine (DSPC), Distearatel-phosphatidyl-ethanolamine (DSPE), Distearatelphosphatidylphosphate ethanolamine imidazole (DSPEI), 1,2-Diundecanoyl-sn-glycero-phosphocholine (DUPC), Lecithin-acetylcholine (EPC), 1,2-Dioleoyl-sn-glycero-3-phosphate (18:1 PA; DOPA), Bis((S)-2-hydroxy-3-(oleoyloxy)propyl)ammonium phosphate (18:1 DMP; LBPA), 1,2-dioleyl-sn-glycero-3-phosphate-(1'-inositol) (DOPI; 18:1 PI), 1,2-distearyl-sn-glycero-3-phosphate-L-serine (18:0 PS), 1,2-dilinoleyl-sn-glycero-3-phosphate-L-serine (18:2 PS), 1-palmolyl-2-oleyl-sn-glycero-3-phosphate-L-serine (16:0-18:1 PS; POPS), 1-stearyl-2-oleyl-sn-glycero-3-phosphate-L-serine (18:0-18:1 PS; POPS) PS), 1-stearyl-2-linoleyl-sn-glycero-3-phosphate-L-serine (18:0-18:2 PS), 1-oleyl-2-hydroxy-sn-glycero-3-phosphate-L-serine (18:1 Lyso PS), 1-stearyl-2-hydroxy-sn-glycero-3-phosphate-L-serine (18:0 Lyso PS), and sphingomyelin.

16. The LNP of claim 12 comprises about 10 mol% phospholipids, about 39 mol% structural lipids, about 2.5 mol% PEG lipids and about 48.5 mol% ionizable lipids.

17. The LNP of claim 12, comprising about 10 mol% phospholipids, about 40 mol% structural lipids, about 1.5 mol% PEG lipids and about 48.5 mol% ionizable lipids.

18. A lipid nanoparticle (LNP) comprising: (A) an ionizable lipid having the structure of formula (CY-VI') or a medically acceptable salt thereof: (CY-VI'), wherein: R1 is selected from the group consisting of: -OH, -OAc, R1a, and; Z1 is a substituted C1-C6 alkyl group, as appropriate; X1 is a substituted C2-C6 alkenyl group, as appropriate; X2 is -CH2CH2-; X4 and X5 are independently substituted C2-C14 alkenyl groups or substituted C2-C14 alkenyl groups, as appropriate; Y1 and Y2 are independently selected from the group consisting of: , , , , , , , , , , , or; wherein the bond marked with "*" is connected to X4 or X5; each Z2 is independently H or a substituted C1-C8 alkyl group, as appropriate; each Z3 is independently substituted C1-C6 alkenyl group, as appropriate; R1a is: , , or; R2a, R2b and R2c are independently hydrogen and C1-C6 alkyl groups; R3a, R3b and R3c are independently hydrogen and C1-C6 alkyl groups; R4a, R4b and R4c are independently hydrogen and C1-C6 alkyl; R5a, R5b and R5c are independently hydrogen and C1-C6 alkyl; R6, R7, R8 and R9 are independently, as appropriate, substituted C1-C14 alkyl, as appropriate, substituted C2-C14 alkenyl or -(CH2)mA-(CH2)nH; each A is independently C3-C8 cycloalkyl; each m is independently 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12; and each n is independently 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12; and (B) encodes RNA.

19. The LNP of claim 7, wherein the ionizable lipid of formula (CY-VI') is selected from the group consisting of: or a pharmaceutically acceptable salt thereof.

20. The LNP of claim 18 further comprises: (a) PEG-lipids; (b) structural lipids; and (c) non-ionizable lipids and / or zwitterionic lipids.

21. The LNP of claim 20, wherein the PEG-lipid is selected from the group consisting of: PEG-c-DOMG, PEG-DMG, PEG-DLPE, PEG-DMPE, PEG-DPPC and PEG-DSPE.

22. The LNP of claim 20, wherein the structural lipid is selected from the group consisting of: cholesterol, fecosterol, sitosterol, ergosterol, campesterol, stigmasterol, brassicasterol, tomatidine, ursolic acid, and α-tocopherol.

23. The LNP of claim 20, wherein the non-ionizable lipid is a phospholipid selected from the group consisting of: 1,2-distearyl-sn-glycero-3-phosphocholine (DSPC), 1,2-dioleyl-sn-glycero-3-phosphoethanolamine (DOPE), 1,2-dilinoleyl-sn-glycero-3-phosphocholine (DLPC), 1,2-dimyristyl-sn-glycero-3-phosphocholine (DMPC), 1,2-dioleyl-sn-glycero-3-phosphocholine (DOPC), 1,2- Di-palmaclopyr-sn-glycero-3-phosphocholine (DPPC), 1,2-diundecanoyl-sn-glycero-3-phosphocholine (DUPC), 1-palmaclopyr-2-oleyl-sn-glycero-3-phosphocholine (POPC), 1,2-di-O-octadecenyl-sn-glycero-3-phosphocholine (18:0 diether PC), 1-oleyl-2-cholesterolylhemisuccinyl-sn-glycero-3-phosphocholine (OChemsPC), 1-hexadecyl-sn-glycero-3-phosphocholine (C16) Lyso PC), 1,2-dilinalinyl-sn-glycero-3-phosphate choline, 1,2-diarachidonicyl-sn-glycero-3-phosphate choline, 1,2-bis(docosahexaenoyl)-sn-glycero-3-phosphate choline, 1,2-diphyliyl-sn-glycero-3-phosphate ethanolamine (ME 16.0) PE), 1,2-distearyl-sn-glycero-3-phosphate ethanolamine, 1,2-dilinoleyl-sn-glycero-3-phosphate ethanolamine, 1,2-dilinoleyl-sn-glycero-3-phosphate ethanolamine, 1,2-diarachidonicyl-sn-glycero-3-phosphate ethanolamine, 1,2-bis(docosahexaenoyl)-sn-glycero-3-phosphate ethanolamine, 1,2-dilinoleyl-sn-glycero-3-phosphate-rac-(1-glycerol) sodium salt (DOPG), (S)-2-ammonium-3-((((R)-2-(oleyloxy)-3-(stearyloxy)propoxy)oxyphosphatidyl)oxy)propionate sodium salt (L-α-phosphatidylinosine);Brain PS), dimyristylphosphatidylcholine (DMPC), dimyristylphosphatidylethanolamine (DMPE), dimyristylphosphatidylglycerol (DMPG), dioleoylphosphatidylethanolamine 4-(N-cis-butenediaminomethyl)cyclohexane-1-carboxylate (DOPE-mal), dioleoylphosphatidylglycerol (DOPG), 1,2-dioleoyl-sn-glycero-3-(phospho-L-serine) (DOPS), acell-fusogenic phospholipids. (DPhPE), Dipalmitoylphosphatidylethanolamine (DPPE), Dipalmitoylphosphatidylglycerol (DPPG), Dipalmitoylphosphatidylserine (DPPS), Distearatelphosphatidylcholine (DSPC), Distearatel-phosphatidyl-ethanolamine (DSPE), Distearatelphosphatidylphosphate ethanolamine imidazole (DSPEI), 1,2-Diundecanoyl-sn-glycero-phosphocholine (DUPC), Lecithin-acetylcholine (EPC), 1,2-Dioleoyl-sn-glycero-3-phosphate (18:1 PA; DOPA), Bis((S)-2-hydroxy-3-(oleoyloxy)propyl)ammonium phosphate (18:1 DMP; LBPA), 1,2-dioleyl-sn-glycero-3-phosphate-(1'-inositol) (DOPI; 18:1 PI), 1,2-distearyl-sn-glycero-3-phosphate-L-serine (18:0 PS), 1,2-dilinoleyl-sn-glycero-3-phosphate-L-serine (18:2 PS), 1-palmolyl-2-oleyl-sn-glycero-3-phosphate-L-serine (16:0-18:1 PS; POPS), 1-stearyl-2-oleyl-sn-glycero-3-phosphate-L-serine (18:0-18:1 PS; POPS) PS), 1-stearyl-2-linoleyl-sn-glycero-3-phosphate-L-serine (18:0-18:2 PS), 1-oleyl-2-hydroxy-sn-glycero-3-phosphate-L-serine (18:1 Lyso PS), 1-stearyl-2-hydroxy-sn-glycero-3-phosphate-L-serine (18:0 Lyso PS), and sphingomyelin.

24. The LNP of request item 18, wherein the encoding RNA is mRNA.

25. The LNP of request item 18, wherein the coding RNA is a circular RNA (circRNA).

26. The compound of claim 1 or a pharmaceutically acceptable salt thereof, wherein: (a) The R2 group is selected from the following groups: , , and; (b) R3 is selected from the following groups: , , and.

27. As in request item 9, LNP, where: (a) R2 is selected from the following groups: , , and ; and (b) R3 is selected from the following groups: , , and .

28. As in request item 18, the LNP, where: (a) R2 is selected from the following groups: , , and ; and (b) R3 is selected from the following groups: , , and .

29. As in request item 10, LNP, where: R1 is -OH; X1 is a C2-C6 alkyl group; X2 is -CH2CH2-; X4 and X5 are independently C2-C6 alkyl groups; Y1 and Y2 are; wherein the bond marked with "*" is connected to X4 or X5; each Z3 is -CH2CH2-; R6, R7, R8 and R9 are independently substituted C1-C14 alkyl groups, substituted C2-C14 alkenyl groups or -(CH2)mA-(CH2)nH, as appropriate; A is a C3-C8 cycloalkyl group; each m is independently 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12; and each n is independently 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12.

30. As in request item 18, the LNP, wherein: R1 is -OH; X1 is a C2-C6 alkyl group; X2 is -CH2CH2-; X4 and X5 are independently C2-C6 alkyl groups; Y1 and Y2 are; wherein the bond marked with "*" is connected to X4 or X5; each Z3 is -CH2CH2-; R6, R7, R8 and R9 are independently substituted C1-C14 alkyl groups, substituted C2-C14 alkenyl groups or -(CH2)mA-(CH2)nH, as appropriate; A is a C3-C8 cycloalkyl group; each m is independently 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12; and each n is independently 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12.

Citation Information

Patent Citations

  • Lipid compounds and lipid nanoparticle compositions

    TW202214566A

  • Nucleoside-modified RNA for inducing an adaptive immune response

    WO2016176330A1

  • Compounds and compositions for intracellular delivery of agents

    WO2017112865A1