Genetically modified cells expressing GLP-1 receptor agonist
Genetically modified cells with inserted GLP-1R agonist sequences address the limitations of current treatments by providing sustained GLP-1R agonist secretion, reducing treatment frequency, and minimizing side effects, thus effectively managing diabetes and obesity.
Patent Information
- Application Number
- PCT/IB2025/053108
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-25
- Filing Date
- 2025-03-24
- Publication Date
- 2025-10-02
AI Technical Summary
Current GLP-1 receptor agonists for treating type 2 diabetes and obesity require frequent injections, cause weight gain upon cessation, and can lead to gastrointestinal discomfort and nausea.
Genetically modified cells expressing a GLP-1 receptor agonist through heterologous nucleotide sequences inserted into safe harbor loci, such as the INS gene locus, allowing for sustained expression and enzymatic cleavage of the GLP-1R agonist, reducing insulin production, and potentially combining with additional genetic modifications for enhanced efficacy.
The genetically modified cells provide sustained GLP-1R agonist secretion with reduced frequency of treatment, minimizing side effects and improving metabolic control.
Smart Images

Figure IB2025053108_02102025_PF_FP_ABST
Abstract
Description
GENETICALLY MODIFIED CELLS EXPRESSING GLP-1 RECEPTOR AGONISTRELATED APPLICATIONS
[0001] The present application claims priority to U.S. Provisional Application No. 63 / 569,424, filed on March 25, 2024; the content of this related application is incorporated herein by reference in its entirety for all purposes.REFERENCE TO SEQUENCE LISTING
[0002] The present application is being filed along with a Sequence Listing in electronic format. The Sequence Listing is provided as a file entitled 80EM-341787- WO SequenceListing, created March 21, 2025, which is 151 kb in size. The information in the electronic format of the Sequence Listing is incorporated herein by reference in its entirety.BACKGROUNDField
[0003] The present disclosure relates generally to the field of gene therapy. Description of the Related Art
[0004] Glucagon-like peptide- 1 (GLP-1) and gastric inhibitory peptide (GIP) are a group of metabolic hormones known as incretins. Incretins are released after a meal and modulate the secretion of insulin released from pancreatic beta cells through a blood-glucose-dependent mechanism. Incretin mimetics, GLP-1 receptor agonists, and DPP-4 inhibitors all aim to stimulate the function of the GLP-1 receptor as a treatment for type 2 diabetes and / or obesity. A decrease in cardiovascular risk has also been observed with these medicines. GLP-1 has also been shown to stimulate insulin production and to increase beta cell mass.
[0005] Although medicines utilizing the GLP-1 pathway have been successful in helping patients treat their type 2 diabetes or obesity, challenges remain. Significantly, these therapeutics must be injected weekly and the need for chronic treatment has become apparent as patients tend to gain weight and lose control of their diabetes upon cessation of treatment. Additionally, at higher pharmacological doses patients can experience gastrointestinal discomfort and nausea leading to some patients choosing to end their treatment. There is a need for methods and compositions for modulating the GLP-1 pathway that require fewer treatments and cause fewer side effects than those previously known in the art.SUMMARY
[0006] Disclosed herein include genetically modified cells. In some embodiments, the genetically modified cell comprises: a heterologous nucleotide sequence comprising a sequence encoding a glucagon-like peptide-1 (GLP-1) Receptor (GLP-1R) agonist; and wherein the cellexpresses the GLP-1R agonist.
[0007] In some embodiments, the heterologous nucleotide sequence is inserted into the genome of the cell. The cell can be, e.g., heterozygous or homozygous for the insertion. In some embodiments, the cell is capable of secreting the GLP-1R agonist. In some embodiments, the heterologous nucleotide sequence further comprises one or more of: a) a sequence encoding a first pro-protein convertase cleavage site; b) a sequence encoding an intervening peptide- 1 (IP-1) sequence of proglucagon; c) a sequence encoding a second pro-protein convertase cleavage site; and d) a sequence encoding at least a portion of an intervening peptide-2 (IP-2) sequence of proglucagon. In some embodiments, the heterologous nucleotide sequence comprises a sequence encoding a GLP-1R agonist protein precursor (pro-protein). In some embodiments, the GLP1-R agonist pro-protein comprises, from the N-terminus to the C-terminus: a) a first pro-protein convertase cleavage site; b) an intervening peptide- 1 (IP-1) sequence of proglucagon; c) the GLP- 1R agonist; d) a second pro-protein convertase cleavage site; and e) at least a portion of an intervening peptide-2 (IP-2) sequence of proglucagon.
[0008] The heterologous nucleotide sequence can be inserted into a safe harbor locus. In some embodiments, the safe harbor locus is selected from AAVS1 gene locus, HRPT gene locus, CCR5 gene locus, globin gene locus, TTR gene locus, TF gene locus, F9 gene locus, Alb gene locus, Gys2 gene locus and PCSK9 gene locus. In some embodiments, the heterologous nucleotide sequence is inserted into the insulin (INS) gene locus of the genome of the cell. In some embodiments, the insertion of the heterologous nucleotide sequence disrupts the INS gene, thereby the expression of insulin by the cell is reduced or eliminated. In some embodiments, the heterologous nucleotide sequence is inserted into the INS gene locus at a site downstream of the portion of the INS gene sequence that encodes insulin signal peptide.
[0009] In some embodiments, the heterologous nucleotide sequence is inserted into the INS gene locus downstream of the portion of the INS gene sequence that encodes insulin signal peptide, thereby the cell is capable of expressing a GLP-1R agonist fusion protein comprising, from N-terminus to C-terminus, (i) the insulin signal peptide and the GLP-1R agonist or (ii) the insulin signal peptide and the GLP-1R agonist pro-protein. In some embodiments, the GLP-1R agonist fusion protein comprises, from N-terminus to C-terminus: a) the insulin signal peptide; b) the first proprotein convertase cleavage site; c) the IP-1 sequence of proglucagon; d) the GLP-1R agonist; e) the second proprotein convertase cleavage site; and f) the at least a portion of the IP-2 sequence of proglucagon. In some embodiments, the genetically modified cell expresses the GLP- 1R agonist pro-protein, the GLP-1R agonist fusion protein, or both. In some embodiments, the GLP-1R agonist pro-protein or the GLP-1R agonist fusion protein is: i) capable of being enzymatically cleaved by one or more proprotein convertases in the cell to generate the GLP-1Ragonist, and optionally ii) the GLP-1R agonist is amidated at the C-terminal end.
[0010] In some embodiments, the GLP-1R agonist comprises a GLP-1 peptide or a functional fragment thereof, or a variant thereof. In some embodiments, the GLP-1R agonist comprises a chimeric peptide comprising at least a portion of a GLP-1 peptide and at least a portion of a gastric inhibitory peptide (GIP). In some embodiments, the heterologous nucleotide sequence comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 21 or SEQ ID NO: 23. In some embodiments, the heterologous nucleotide sequence comprises the sequence of SEQ ID NO: 21 or SEQ ID NO: 23. In some embodiments, the heterologous nucleotide sequence comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 8 or SEQ ID NO: 10. In some embodiments, the heterologous nucleotide sequence comprises the sequence of SEQ ID NO: 8 or SEQ ID NO: 10. In some embodiments, the heterologous nucleotide sequence comprises a stop codon at the 3’ end.
[0011] In some embodiments, the GLP-1R agonist comprises an amino acid sequence comprising the sequence of SEQ ID NO: 34, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 34, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 34. In some embodiments, the GLP-1R agonist comprises an amino acid sequence comprising the sequence of SEQ ID NO: 35, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 35, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 35. In some embodiments, the GLP-1R agonist pro-protein comprises an amino acid sequence comprising the sequence of SEQ ID NO: 29, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 29, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 29. In some embodiments, the GLP-1R agonist pro-protein comprises an amino acid sequence comprising the sequence of SEQ ID NO: 30, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 30, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 30. In some embodiments, the GLP-1R agonist fusion protein comprises an amino acid sequence comprising the sequence of SEQ ID NO: 32, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 32, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 32. In some embodiments, the GLP-1R agonist fusion protein comprises an amino acid sequence comprising the sequence of SEQ ID NO: 33, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 33, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 33.
[0012] The cell can be, e.g., a stem cell (e.g., an embryonic stem cell, an induced pluripotent stem cell, or an adult stem cell). In some embodiments, the stem cell is capable of being differentiated into a pancreatic endocrine cell, an intestinal enteroendocrine cell, a neuron,a definitive endoderm cell, a primitive gut tube cell, a posterior foregut cell, a pancreatic endoderm cell, a pancreatic endocrine precursor cell, an immature beta cell, and / or a pancreatic beta cell. In some embodiments, the stem cell is capable of being differentiated into a stage 6 (S6) pancreatic progenitor cell. In some embodiments, the S6 cell expresses one or markers each selected from Forkhead Box A2 (F0XA2), Chromogranin A (CHGA), Pancreatic and duodenal homeobox 1 (PDX1), NK6 Homeobox 1 (NKX6.1), Neurogenin 3 (NGN3), Insulin (INS), and ISL LIM homeobox 1 (ISL1). In some embodiments, the cell is a differentiated cell selected from the group consisting of a pancreatic endocrine cell, an intestinal enteroendocrine cell, and a neuron.
[0013] In some embodiments, the cell comprises one or more additional genetic modifications. In some embodiments, the one or more additional genetic modifications comprise: (1) a disrupted beta-2-microglobulin (B2M) gene and insertion of a nucleotide sequence encoding programmed death-ligand 1 (PD-L1) into the disrupted B2M gene, wherein the cell expresses PD- L1 and has reduced or eliminated expression of B2M; and / or (2) a disrupted thioredoxin interacting protein (TXNIP) gene and insertion of a nucleotide sequence encoding HLA class I histocompatibility antigen, alpha chain E (HLA-E) into the disrupted TXNIP gene; optionally the nucleotide sequence encoding HLA-E encodes an HLA-E trimer comprising a B2M signal peptide fused to an HLA-G presentation peptide fused to a B2M membrane protein fused to HLA-E without its signal peptide; wherein the cell expresses HLA-E or the HLA-E trimer and has reduced or eliminated expression of TXNIP. In some embodiments, the one or more additional genetic modifications comprise: (3) a disrupted B2M gene and insertion of a nucleotide sequence encoding PD-L1 and a nucleotide sequence encoding tumor necrosis factor alpha induced protein 3 (TNFAIP3) into the disrupted B2M gene, wherein the cell expresses PD-L1 and TNFAIP3 and has reduced or eliminated expression of B2M; and / or (4) a disrupted TXNIP gene and insertion of a nucleotide sequence encoding HLA-E and a nucleotide sequence mesencephalic astrocyte derived neurotrophic factor (MANF) into the disrupted TXNIP gene; optionally the nucleotide sequence encoding HLA-E encodes an HLA-E trimer comprising a B2M signal peptide fused to an HLA-G presentation peptide fused to a B2M membrane protein fused to HLA-E without its signal peptide; wherein the cell expresses HLA-E or the HLA-E trimer and MANF and has disrupted expression of TXNIP. In some embodiments, the nucleotide sequence encoding TNFAIP3 and the nucleotide sequence encoding PD-L1 are linked by a nucleotide sequence encoding a P2A peptide such that the insertion of (3) comprises a nucleotide sequence encoding TNFAIP3-P2A-PD-L1. In some embodiments, the nucleotide sequence encoding the HLA-E trimer and the nucleotide sequence encoding MANF are linked by a nucleotide sequence encoding a P2A peptide such that the insertion of (4) comprises a nucleotide sequence encoding MANF- P2A-HLA-E trimer. In some embodiments, the nucleotide sequence encoding PD-L1 comprisesa sequence that is at least 85% identical to the sequence of SEQ ID NO: 83. In some embodiments, the nucleotide sequence encoding the HLA-E trimer comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 84. In some embodiments, the nucleotide sequence encoding TNFAIP3 comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 85. In some embodiments, the nucleotide sequence encoding MANF comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 86. In some embodiments, the nucleotide sequence encoding TNFAIP3-P2A-PD-L1 comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 87. In some embodiments, the nucleotide sequence encoding MANF-P2A-HLA-E trimer comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 89. The heterologous nucleotide sequence can be, e.g., a DNA sequence or a RNA (e.g., mRNA) sequence.
[0014] Disclosed herein include methods for generating a genetically modified cell. In some embodiments, the method comprises delivering to a cell: (a) an RNA-guided nuclease and a gRNA targeting a target site within the genome of the cell; and (b) a nucleic acid comprising (i) a nucleotide sequence that is at least 85% identical to a region located upstream of the target site, (ii) a heterologous nucleotide sequence encoding a GLP-1 Receptor (GLP-1R) agonist, and (iii) a nucleotide sequence that is at least 85% identical to a region located downstream of the target site; thereby the target site in the genome of the cell is cleaved and the heterologous nucleotide sequence is inserted into the genome by homology directed repair (HDR), thereby generating a genetically modified cell, wherein the genetically modified cell expresses the GLP-1R agonist.
[0015] In some embodiments, the cell is heterozygous or homozygous for the insertion. In some embodiments, the cell is capable of secreting the GLP-1R agonist. In some embodiments, the heterologous nucleotide sequence further comprises one or more of: a) a sequence encoding a first pro-protein convertase cleavage site; b) a sequence encoding an intervening peptide- 1 (IP-1) sequence of proglucagon; c) a sequence encoding a second proprotein convertase cleavage site; and d) a sequence encoding at least a portion of an intervening peptide-2 (IP -2) sequence of proglucagon. In some embodiments, the heterologous nucleotide sequence comprises a sequence encoding a GLP-1 R agonist protein precursor (pro-protein). In some embodiments, the GLP-R agonist pro-protein comprises, from the N-terminus to the C- terminus: a) a first pro-protein convertase cleavage site; b) an intervening peptide- 1 (IP-1) sequence of proglucagon; c) the GLP-1R agonist; d) a second pro-protein convertase cleavage site; and e) at least a portion of an intervening peptide-2 (IP-2) sequence of proglucagon.
[0016] In some embodiments, the target site is within a gene locus. In some embodiments, the insertion of the heterologous sequence into the target site disrupts the gene, thereby reducing or eliminating expression of a gene product encoded by the gene. In someembodiments, the gene locus is a safe harbor locus. In some embodiments, the safe harbor locus is selected from the group consisting of AAVS1 gene locus, HRPT gene locus, CCR5 gene locus, globin gene locus, TTR gene locus, TF gene locus, F9 gene locus, Alb gene locus, Gys2 gene locus and PCSK9 gene locus. In some embodiments, the gene locus is the insulin (INS) gene locus. In some embodiments, the insertion of the heterologous nucleotide sequence disrupts the INS gene, thereby the expression of insulin by the cell is reduced or eliminated.
[0017] In some embodiments, the target site is a site downstream of the sequence of the INS gene that encodes insulin signal peptide. In some embodiments, the heterologous nucleotide sequence is inserted into the INS locus downstream of the portion of the INS gene sequence that encodes insulin signal peptide, thereby the cell is capable of expressing a GLP-1R agonist fusion protein comprising, from N-terminus to C-terminus, (i) the insulin signal peptide and the GLP-1R agonist or (ii) the insulin signal peptide and the GLP-1R agonist pro-protein. In some embodiments, the GLP-1R agonist fusion protein comprises, from N-terminus to C- terminus: a) the insulin signal peptide; b) the first proprotein convertase cleavage site; c) the IP-1 sequence of proglucagon; d) the GLP-1R agonist; e) the second proprotein convertase cleavage site; and f) the at least a portion of the IP-2 sequence of proglucagon. In some embodiments, the genetically modified cell expresses the GLP-1R agonist pro-protein, the GLP-1R agonist fusion protein, or both. In some embodiments, the GLP-1R agonist pro-protein or the GLP-1R agonist fusion protein is: i) capable of being enzymatically cleaved by one or more proprotein convertases in the cell to generate the GLP-1R agonist, and optionally ii) the GLP-1R agonist is amidated at the C-terminal end.
[0018] In some embodiments, the GLP-1R agonist comprises a GLP-1 peptide or a functional fragment thereof, or a variant thereof. In some embodiments, the GLP-1R agonist comprises a chimeric peptide comprising at least a portion of GLP-1 peptide and at least a portion of gastric inhibitory peptide (GIP). In some embodiments, the heterologous nucleotide sequence comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 21 or SEQ ID NO: 23. In some embodiments, the heterologous nucleotide sequence comprises the sequence of SEQ ID NO: 21 or SEQ ID NO: 23. In some embodiments, the heterologous nucleotide sequence comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 8 or SEQ ID NO: 10. In some embodiments, the heterologous nucleotide sequence comprises the sequence of SEQ ID NO: 8 or SEQ ID NO: 10. In some embodiments, the heterologous nucleotide sequence comprises a stop codon at the 3’ end. In some embodiments, the sequence of (b)(i) comprises the sequence of any one of SEQ ID NOs: 6, 19, and 96 or a sequence that is at least 85% identical to the sequence of any one of SEQ ID NOs: 6, 19, and 96. In some embodiments, the sequence of (b)(iii) comprises the sequence of SEQ ID NO: 22 or SEQ ID NO: 9 or a sequence that is at least85% identical to the sequence of SEQ ID NO: 22 or SEQ ID NO: 9. In some embodiments, the nucleic acid of (b) comprises the sequence of any one of SEQ ID NOs: 11-12 and 97-98 or a sequence that is at least 85% identical to the sequence of any one of SEQ ID NOs: 11-12 and 97- 98.
[0019] In some embodiments, the GLP-1R agonist comprises an amino acid sequence comprising the sequence of SEQ ID NO: 34, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 34, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 34. In some embodiments, the GLP-1R agonist comprises an amino acid sequence comprising the sequence of SEQ ID NO: 35, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 35, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 35. In some embodiments, the GLP-1R agonist pro-protein comprises an amino acid sequence comprising the sequence of SEQ ID NO: 29, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 29, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 29. In some embodiments, the GLP-1R agonist pro-protein comprises an amino acid sequence comprising the sequence of SEQ ID NO: 30, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 30, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 30. In some embodiments, the GLP-1R agonist fusion protein comprises an amino acid sequence comprising the sequence of SEQ ID NO: 32, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 32, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 32. In some embodiments, the GLP-1R agonist fusion protein comprises an amino acid sequence comprising the sequence of SEQ ID NO: 33, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 33, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 33.
[0020] In some embodiments, the RNA-guided nuclease is a CRISPR-associated nuclease. In some embodiments, the CRISPR-associated nuclease comprises an N-terminal nuclear localization signal (NLS), a C-terminal NLS, or both. In some embodiments, the CRISPR- associated nuclease comprises Cas9, Cpfl, or a variant thereof. In some embodiments, the Cas9 is S. pyogenes Cas9. In some embodiments, the gRNA and the Cas9 are delivered to the cell at a molar ratio of 5: 1. In some embodiments, the RNA-guided endonuclease and the gRNA are complexed as a ribonucleoprotein (RNP) complex prior to the delivering. In some embodiments, the gRNA comprises a spacer sequence comprising the sequence of SEQ ID NO: 1 or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 1.
[0021] In some embodiments, the cell is a stem cell. In some embodiments, the cell is an embryonic stem cell, an induced pluripotent stem cell, or an adult stem cell. In someembodiments, the stem cell is capable of being differentiated into a pancreatic endocrine cell, an intestinal enteroendocrine cell, a neuron, a definitive endoderm cell, a primitive gut tube cell, a posterior foregut cell, a pancreatic endoderm cell, a pancreatic endocrine precursor cells, an immature beta cell, and / or a pancreatic beta cell. In some embodiments, the stem cell is capable of being differentiated into a stage 6 (S6) pancreatic progenitor cell. In some embodiments, the S6 cell expresses one or markers each selected from Forkhead Box A2 (F0XA2), Chromogranin A (CHGA), Pancreatic and duodenal homeobox 1 (PDX1), NK6 Homeobox 1 (NKX6.1), Neurogenin 3 (NGN3), Insulin (INS), and ISL LIM homeobox 1 (ISL1). In some embodiments, the cell is a differentiated cell selected from a pancreatic endocrine cell, an intestinal enteroendocrine cell, and a neuron.
[0022] In some embodiments, the cell comprises one or more additional genetic modifications. In some embodiments, the one or more additional genetic modifications comprise: (1) a disrupted beta-2-microglobulin (B2M) gene and insertion of a nucleotide sequence encoding programmed death-ligand 1 (PD-L1) into the disrupted B2M gene, wherein the cell expresses PD- L1 and has reduced or eliminated expression of B2M; and / or (2) a disrupted thioredoxin interacting protein (TXNIP) gene and insertion of a nucleotide sequence encoding HLA class I histocompatibility antigen, alpha chain E (HLA-E); optionally the nucleotide sequence encoding HLA-E encodes an HLA-E trimer comprising a B2M signal peptide fused to an HLA-G presentation peptide fused to a B2M membrane protein fused to HLA-E without its signal peptide; wherein the cell expresses HLA-E or the HLA-E trimer and has reduced or eliminated expression of TXNIP. In some embodiments, the one or more additional genetic modifications comprise: (3) a disrupted beta-2-microglobulin (B2M) gene and insertion of a nucleotide sequence encoding PD-L1 and a nucleotide sequence encoding tumor necrosis factor alpha induced protein 3 (TNFAIP3) into the disrupted B2M gene, wherein the cell expresses PD-L1 and TNFAIP3 and has reduced or eliminated expression of B2M; and / or (4) a disrupted thioredoxin interacting protein (TXNIP) gene and insertion of a nucleotide sequence encoding HLA class I histocompatibility antigen, alpha chain E (HLA-E) and a nucleotide sequence mesencephalic astrocyte derived neurotrophic factor (MANF) into the disrupted TXNIP gene; optionally the nucleotide sequence encoding HLA-E encodes an HLA-E trimer comprising a B2M signal peptide fused to an HLA-G presentation peptide fused to a B2M membrane protein fused to HLA-E without its signal peptide; wherein the cell expresses HLA-E or the HLA-E trimer and MANF and has disrupted expression of TXNIP. In some embodiments, the nucleotide sequence encoding TNFAIP3 and the nucleotide sequence encoding PD-L1 are linked by a nucleotide sequence encoding a P2A peptide such that the insertion of (3) comprises a nucleotide sequence encoding TNFAIP3-P2A-PD-L1. In some embodiments, the nucleotide sequence encoding the HLA-Etrimer and the nucleotide sequence encoding MANF are linked by a nucleotide sequence encoding a P2A peptide such that the insertion of (4) comprises a nucleotide sequence encoding MANF- P2A-HLA-E trimer. In some embodiments, the nucleotide sequence encoding PD-L1 comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 83. In some embodiments, the nucleotide sequence encoding the HLA-E trimer comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 84. In some embodiments, the nucleotide sequence encoding TNFAIP3 comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 85. In some embodiments, the nucleotide sequence encoding MANF comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 86. In some embodiments, the nucleotide sequence encoding TNFAIP3-P2A-PD-L1 comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 87. In some embodiments, the nucleotide sequence encoding MANF-P2A-HLA-E trimer comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 89.
[0023] Disclosed herein include populations of cells comprising a plurality of any of the genetically modified cell of the disclosure or comprising a plurality of the genetically modified cell generating by the methods disclosed herein. Disclosed herein include compositions comprising a population of genetically modified cells described. Also disclosed herein are pharmaceutical compositions comprising the compositions of the disclosure and a pharmaceutically acceptable excipient. Provided herein are compositions and pharmaceutical compositions for use in treating a disease or disorder in a subject.
[0024] Disclosed herein include methods of treating a disease or disorder in a subject. In some embodiments, the method comprises administering to a subject in need thereof a therapeutically effective amount of any of the compositions disclosed herein, thereby treating the disease or disorder in the subject.
[0025] The disease or disorder can be, e.g., diabetes mellitus, dyslipidemia, fatty liver disease, metabolic syndrome, non-alcoholic steatohepatitis or obesity. In some embodiments, the diabetes mellitus is type 2 diabetes mellitus. In some embodiments, the disease or disorder is a weight-related comorbid condition selected from hypertension, dyslipidemia, prediabetes, type 2 diabetes mellitus, obstructive sleep apnea, cardiovascular disease, and any combination thereof.
[0026] Disclosed herein are methods for chronic weight management in a subject in need thereof. In some embodiments, the method comprises: administering to the subject a therapeutically effective amount of any of the compositions of the disclosure.
[0027] In some embodiments, the subject has a BMI > 27 kg / m2and at least one weight- related comorbidity. Each of the at least one weight-related comorbidity can be selected from, e.g., hypertension, dyslipidemia, type 2 diabetes mellitus, coronary heart disease, stroke,gallbladder disease, osteoarthritis, sleep apnea, cancer, mental illness, and body pain.BRIEF DESCRIPTION OF THE DRAWINGS
[0028] FIG. 1A-FIG. 1C display non-limiting exemplary schematics of the editing strategies and cells of the disclosure. FIG. 1A depicts a schematic displaying an exemplary strategy for knock-in (“KI”) of GLP-1R agonist as disclosed herein. Also shown are sequences of GLP-1 (7-37) and Tirzepatide. FIG. IB depicts a cartoon of an exemplary genetically modified cell of the disclosure. FIG. 1C displays the relative timing of insulin secretion relative to food intake.
[0029] FIG. 2 displays exemplary methods and for differentiation of stem cells into S6 progenitor cells.
[0030] FIG. 3 depicts an exemplary flow diagram for the characterization of nonlimiting examples of the genetically modified cells of the disclosure.
[0031] FIG. 4 depicts a non-limiting exemplary illustration of the sequence around exon 1 of the INS gene following insertion of nucleotide sequence encoding GLP-1 (7-37).
[0032] FIG. 5 depicts a non-limiting exemplary illustration of the sequence around exon 1 of the INS gene following insertion of nucleotide sequence encoding a GLP-1R agonist chimeric peptide (e.g., Tirzepatide).
[0033] FIG. 6 depicts a non-limiting exemplary illustration of the establishment of a mouse model of diabetes.
[0034] FIG. 7 depicts a non-limiting exemplary illustration of an experimental outline for characterization of edited cells in vivo.
[0035] FIG. 8 displays a diagram of timeline for differentiation of iPSCs (e.g., GLP1 / TIRZ edited iPSC lines) into stage 6 (S6) immature P cells.
[0036] FIG. 9A-FIG. 9B display exemplary images (FIG. 9A) and yield (FIG. 9B) of S6 cells comprising GLP1 or Tirzepatide KI.
[0037] FIG. 10 displays exemplary data of characterization of edited S6 cells by FACS.
[0038] FIG. 11 displays results of a GLP1 bioactivity assay. Using a HEK-GLP1R reporter line, it was found that the secreted peptide from the TIRZ-KI S6 produced a higher luminescence reading compared to the seed clone control. The data shows that secreted TIRZ peptide from the TIRZ-KI lines is biologically active.
[0039] FIG. 12 displays C-peptide levels detected following in vivo maturation of edited S6 cells.
[0040] FIG. 13A-FIG. 13B show graphs of body weight (FIG. 13 A) and glucose intolerance (FIG. 13B) in a mouse obesity model (high fat diet SCID-BEIGE mouse) to testtherapeutic efficacy of the disclosed genetically modified cells.DETAILED DESCRIPTION
[0041] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the Figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein and made part of the disclosure herein.
[0042] All patents, published patent applications, other publications, and sequences from GenBank, and other databases referred to herein are incorporated by reference in their entirety with respect to the related technology.
[0043] Disclosed herein include genetically modified cells. In some embodiments, the genetically modified cell comprises: a heterologous nucleotide sequence comprising a sequence encoding a glucagon-like peptide-1 (GLP-1) Receptor (GLP-1R) agonist; and wherein the cell expresses the GLP-1 R agonist.
[0044] Disclosed herein include methods for generating a genetically modified cell. In some embodiments, the method comprises delivering to a cell: (a) an RNA-guided nuclease and a gRNA targeting a target site within the genome of the cell; and (b) a nucleic acid comprising (i) a nucleotide sequence that is at least 85% identical to a region located upstream of the target site, (ii) a heterologous nucleotide sequence encoding a GLP-1 Receptor (GLP-1R) agonist, and (iii) a nucleotide sequence that is at least 85% identical to a region located downstream of the target site; thereby the target site in the genome of the cell is cleaved and the heterologous nucleotide sequence is inserted into the genome by homology directed repair (HDR), thereby generating a genetically modified cell, wherein the genetically modified cell expresses the GLP-1R agonist.
[0045] Disclosed herein include populations of cells comprising a plurality of any of the genetically modified cell of the disclosure or comprising a plurality of the genetically modified cell generating by the methods disclosed herein.
[0046] Disclosed herein include compositions comprising a population of genetically modified cells described. Also disclosed herein are pharmaceutical compositions comprising the compositions of the disclosure and a pharmaceutically acceptable excipient. Provided herein are compositions and pharmaceutical compositions for use in treating a disease or disorder in a subject.
[0047] Disclosed herein include methods of treating a disease or disorder in a subject. In some embodiments, the method comprises administering to a subject in need thereof a therapeutically effective amount of any of the compositions disclosed herein, thereby treating the disease or disorder in the subject. Disclosed herein are methods for chronic weight management in a subject in need thereof. In some embodiments, the method comprises: administering to the subject a therapeutically effective amount of any of the compositions of the disclosure. Definitions
[0048] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs. See, e.g., Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For purposes of the present disclosure, the following terms are defined below.
[0049] As used herein, the term “deletion,” which may be used interchangeably with the terms “genetic deletion” or “knock-out” can have their ordinary meaning and can also refer to a genetic modification wherein a site or region of genomic DNA is removed by any molecular biology method, e.g., methods described herein, e.g., by delivering to a site of genomic DNA an endonuclease and at least one gRNA. Any number of nucleotides can be deleted. In some embodiments, a deletion involves the removal of at least one, at least two, at least three, at least four, at least five, at least ten, at least fifteen, at least twenty, or at least 25 nucleotides. In some embodiments, a deletion involves the removal of 10-50, 25-75, 50-100, 50-200, or more than 100 nucleotides. In some embodiments, a deletion involves the removal of part or all of one target gene. In some embodiments, a deletion involves the removal of part or all of two target genes, three target gene, or four target genes. In some embodiments, the removal of part of a target gene refers to removal of all or part of a promoter and / or coding sequence of a gene. In some embodiments, a deletion involves the removal of a transcriptional regulator, e.g., a promoter region, of a target gene. In some embodiments, a deletion involves the removal of all or part of a coding region such that the product normally expressed by the coding region is no longer expressed, is expressed as a truncated form, or expressed at a reduced level. In some embodiments, a deletion leads to a decrease in expression of a gene relative to an unmodified cell. In some embodiments, a deletion leads to a loss of expression of a gene relative to an unmodified cell. In some embodiments, a knock-out of a gene can be achieved without deleting endogenous sequence. For example, an in-frame insertion of heterologous sequence into a genomic locus can knock-out the endogenous gene, for example, by introduction of a stop codon in the heterologous sequence. A genetic deletion or knock-out can be an indel mutation. An "indel", as used herein,refers to the insertion or deletion of a nucleotide base within a nucleic acid. Such insertions or deletions can lead to frame shift mutations within a coding region of a gene, resulting in loss-of- function in said gene.
[0050] As used herein the terms “disruption,” “disrupting,” or “disrupted” can have their ordinary meaning, and can also refer to genetic modifications that alter the level of expression of a target gene. In some aspects, the disruption can be due to a deletion of at least one nucleotide within or near the target gene or a deletion of part or all of a target gene, as described above. In other aspects, the disruption also can be due to a substitution of at least one nucleotide and / or an insertion of at least one nucleotide within or near the target gene. In further aspects, the disruption can be due to an insertion of one or more exogenous polynucleotides within or near the target gene. In general, as used herein, disrupted expression refers to reduced or eliminated expression of the target gene. In some embodiments, the disruption can be a reduced level of expression (e.g., express less than 30%, less than 25%, less than 20%, less than 10%, or less than 5% of the level of an unmodified cell). In some embodiments, the disruption can be eliminated expression (e.g., no expression or an undetectable level of RNA and / or protein expression). Expression can be measured using any standard RNA-based, protein-based, and / or antibody-based detection method (e.g., RT-PCR, ELISA, flow cytometry, immunocytochemistry, and the like). Detectable levels are defined as being higher than the limit of detection (LOD), which is the lowest concentration that can be measured (detected) with statistical significance by means of a given detection method.
[0051] As used herein, the term “endonuclease” generally refers to an enzyme that cleaves phosphodiester bonds within a polynucleotide. In some embodiments, an endonuclease specifically cleaves phosphodiester bonds within a DNA polynucleotide. In some embodiments, an endonuclease is a zinc finger nuclease (ZFN), transcription activator like effector nuclease (TALEN), homing endonuclease (HE), meganuclease, MegaTAL, or a CRISPR-associated endonuclease. In some embodiments, an endonuclease is an RNA-guided endonuclease. In certain aspects, the RNA-guided endonuclease is a CRISPR nuclease, e.g., a Type II CRISPR Cas9 endonuclease or a Type V CRISPR Cpfl endonuclease. In some embodiments, an endonuclease is a Casl, CaslB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csxl2), CaslOO, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4, or Cpfl endonuclease, or a homolog thereof, a recombination of the naturally occurring molecule thereof, a codon-optimized version thereof, or a modified version thereof, or combinations thereof. In some embodiments, an endonuclease may introduce one or more single-stranded breaks (SSBs) and / or one or more double-stranded breaks (DSBs).
[0052] As used herein, the terms “heterologous” and “exogenous” can be used interchangeably and refer to a nucleic acid molecule that originates from a foreign species, or, if from the same species, is modified from its native form in composition and / or genomic locus by deliberate human intervention.
[0053] As used herein, the term “genetic modification” generally refers to a site of genomic DNA that has been genetically edited or manipulated using any molecular biological method, e.g., methods described herein, e.g., by delivering to a site of genomic DNA an endonuclease and at least one gRNA. Example genetic modifications include insertions, deletions, duplications, inversions, and translocations, and combinations thereof. In some embodiments, a genetic modification is a deletion. In some embodiments, a genetic modification is an insertion. In other embodiments, a genetic modification is an insertion-deletion mutation (or indel), such that the reading frame of the target gene is shifted leading to an altered gene product or no gene product.
[0054] As used herein, the term “guide RNA” or “gRNA” generally refers to short ribonucleic acid (RNA) that can interact with, e.g., bind to, an endonuclease and bind, or hybridize to a specific (e.g., target) genomic site or region. In some embodiments, a gRNA is a singlemolecule guide RNA (sgRNA). In some embodiments, a gRNA may comprise a spacer extension region. In some embodiments, a gRNA may comprise a tracrRNA extension region. In some embodiments, a gRNA is single-stranded. In some embodiments, a gRNA comprises naturally occurring nucleotides. In some embodiments, a gRNA is a chemically modified gRNA. In some embodiments, a chemically modified gRNA is a gRNA that comprises at least one nucleotide with a chemical modification, e.g., a 2'-O-methyl sugar modification. In some embodiments, a chemically modified gRNA comprises a modified nucleic acid backbone. In some embodiments, a chemically modified gRNA comprises a 2'-O-methyl-phosphorothioate residue. In some embodiments, a gRNA may be pre-complexed with a DNA endonuclease.
[0055] As used herein, the term “protospacer adjacent sequence” or “PAM” refers to an approximately 2-6 base pair DNA sequence that is an important targeting component of a Cas9 nuclease. Typically, the PAM sequence is on either strand, and is downstream in the 5' to 3' direction of the Cas9 cut site. The canonical PAM sequence (z.e., the PAM sequence that is associated with the Cas9 nuclease of Streptococcus pyogenes or SpCas9) is 5'-NGG-3' wherein “N” is any nucleobase followed by two guanine (“G”) nucleobases. Different PAM sequences can be associated with different Cas9 nucleases or equivalent proteins from different organisms. In addition, any given Cas9 nuclease, e.g., SpCas9, may be modified to alter the PAM specificity of the nuclease such that the nuclease recognizes alternative PAM sequence.
[0056] The terms “polynucleotide” and “nucleic acid” are used interchangeably hereinand refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. A polynucleotide can be single-, double-, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids / triple helices, or a polymer including purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases. The term “nucleotide sequence” as used herein shall have its ordinary meaning, and can also refer to the order of nucleotides in a nucleic acid (DNA or RNA). As is understood by the skilled artisan, the nucleotide sequence determines the sequence of a gene product (e.g., an RNA or protein).
[0057] A “functional variant” or “functional mutant,” as used herein, refers to any variant or mutant of a reference protein (e.g., a wild-type protein) that encompasses one or more alterations to the amino acid sequence of the reference protein while retaining one or more of the functions, e.g., catalytic or binding functions. In some embodiments, the one or more alterations to the amino acid sequence comprises amino acid substitutions, insertions or deletions, or any combination thereof. In some embodiments, the one or more alterations to the amino acid sequence comprises amino acid substitutions. For example, a functional variant of a GLP-1 may comprise one or more amino acid substitutions compared to the amino acid sequence of wild-type GLP-1 but retains the ability under at least one set of conditions to bind and activate the GLP-1 receptor. When the reference protein is a fusion of multiple functional domains, a functional variant thereof may retain one or more of the functions of at least one of the functional domains. For example, in some embodiments, a functional fragment of a Cas9 may comprise one or more amino acid substitutions in a nuclease domain, e.g., an H840A amino acid substitution, compared to the amino acid sequence of a wild type Cas9, but retains the DNA binding ability and lacks the nuclease activity partially or completely.
[0058] As used herein, the term “binding” refers to a non-covalent interaction between macromolecules (e.g., between a protein and a nucleic acid). While in a state of non-covalent interaction, the macromolecules are said to be “associated” or “interacting” or “binding” (e.g., when a molecule X is said to interact with a molecule Y, it means that the molecule X binds to molecule Y in a non-covalent manner). Binding interactions can be characterized by a dissociation constant (Kd), for example a Kd of, or a Kd less than, 10'6M, 10'7M, 10'8M, 10'9M, IO'10M, 10"11M, 10'12M, 10'13M, 10'14M, 10'15M, or a number or a range between any two of these values. Kd can be dependent on environmental conditions, e.g., pH and temperature. “Affinity” refers to the strength of binding, and increased binding affinity is correlated with a lower Kd.
[0059] As used herein, the term “hybridizing” or “hybridize” refers to the pairing of substantially complementary or complementary nucleic acid sequences within two different molecules. Pairing can be achieved by any process in which a nucleic acid sequence joins with asubstantially or fully complementary sequence through base pairing to form a hybridization complex. “Hybridizing” or “hybridize” can comprise denaturing the molecules to disrupt the intramolecular structure(s) (e.g., secondary structure(s)) in the molecule. In some embodiments, denaturing the molecules comprises heating a solution comprising the molecules to a temperature sufficient to disrupt the intramolecular structures of the molecules. In some instances, denaturing the molecules comprises adjusting the pH of a solution comprising the molecules to a pH sufficient to disrupt the intramolecular structures of the molecules. For purposes of hybridization, two nucleic acid sequences or segments of sequences are “substantially complementary” if at least 80% of their individual bases are complementary to one another.
[0060] The terms “complementarity” and “complementary” mean that a nucleic acid can form hydrogen bond(s) with another nucleic acid based on traditional Watson-Crick base paring rule, that is, adenine (A) pairs with thymine (T, or uracil (U) in RNA) and guanine (G) pairs with cytosine (C). Complementarity can be perfect (e.g., complete complementarity) or imperfect (e.g. partial complementarity). Perfect or complete complementarity indicates that each and every nucleic acid base of one strand is capable of forming hydrogen bonds according to Watson-Crick canonical base pairing with a corresponding base in another, antiparallel nucleic acid sequence. Partial complementarity indicates that only a percentage of the contiguous residues of a nucleic acid sequence can form Watson-Crick base pairing with the same number of contiguous residues in another, antiparallel nucleic acid sequence. In some embodiments, the complementarity can be at least 70%, 80%, 90%, 100% or a number or a range between any two of these values. In some embodiments, the complementarity is perfect, i.e., 100%. For example, the complementary candidate sequence segment is perfectly complementary to the candidate sequence segment, whose sequence can be deduced from the candidate sequence segment using the Watson-Crick base pairing rules.
[0061] The terms “DNA editing efficiency,” or “editing efficiency” may be used interchangeably herein and can refer to the number or proportion of intended target sequences that are edited. For example, if a CRISPR-Cas9 system edits 10% of the intended target sequence (e.g., within a cell or within a population of cells), then the system can be described as being 10% efficient. In some embodiments, the efficiency can be reported as % indel, e.g., the proportion of insertions and / or deletions detected in the target sequence. Indels (e.g., insertion-deletions) can result from repair of double-stranded DNA breaks caused by Cas9 cleavage by processes including, but not limited to, non-homologous end joining (NHEJ) repair.
[0062] The term “off-target editing frequency,” as used herein, refers to the number or proportion of unintended DNA sequences that are edited or cut. On-target and off- target editing frequencies may be measured by the methods and assays described herein, furtherin view of techniques known in the art, including high-throughput sequencing reads. As used herein, high-throughput sequencing involves the hybridization of nucleic acid primers (e.g., DNA primers) with complementarity to nucleic acid (e.g., DNA) regions just upstream or downstream of the target sequence or off-target sequence of interest. Since many of the Cas9-dependent off-target sites have high sequence identity to the target site of interest, nucleic acid primers with sufficient complementarity to regions upstream or downstream of the Cas9- dependent off-target site may be designed using techniques and kits known in the art. These kits make use of polymerase chain reaction (PCR) amplification, which produces amplicons as intermediate products. The target and off-target sequences may comprise genomic loci that further comprise protospacers and PAMs. High-throughput sequencing techniques used herein may further include Sanger sequencing and / or whole genome sequencing (WGS).
[0063] As used herein, the terms “transfection” or “infection” refer to the introduction of a nucleic acid into a host cell, such as by contacting the cell with liposomes or nanoparticles (e.g., lipid nanoparticles) as described herein.
[0064] As used herein, “treatment” refers to a clinical intervention made in response to a disease, disorder or physiological condition manifested by a patient or to which a patient may be susceptible. The aim of treatment includes, but is not limited to, the alleviation or prevention of symptoms, slowing or stopping the progression or worsening of a disease, disorder, or condition and / or the remission of the disease, disorder or condition. “Treatments” refer to one or both of therapeutic treatment and prophylactic or preventative measures. Subjects in need of treatment include those already affected by a disease or disorder or undesired physiological condition as well as those in which the disease or disorder or undesired physiological condition is to be prevented.
[0065] As used herein, the terms “effective amount” or “pharmaceutically effective amount” or “therapeutically effective amount” refer to an amount sufficient to effect beneficial or desirable biological and / or clinical results.
[0066] The term “pharmaceutically acceptable excipient” as used herein refers to any suitable substance that provides a pharmaceutically acceptable carrier, additive or diluent for administration of a compound(s) of interest to a subject. Pharmaceutically acceptable excipients can encompass substances referred to as pharmaceutically acceptable diluents, pharmaceutically acceptable additives, and pharmaceutically acceptable carriers.
[0067] As used herein, a “subject” refers to an animal for whom a diagnosis, treatment, or therapy is desired. In some embodiments, the subject is a mammal. “Mammal,” as used herein, refers to an individual belonging to the class Mammalia and includes, but not limited to, humans,domestic and farm animals, zoo animals, sports and pet animals. Non-limiting examples of mammals include mice; rats; rabbits; guinea pigs; dogs; cats; sheep; goats; cows; horses; primates, such as monkeys, chimpanzees and apes, and, in particular, humans. In some embodiments, the mammal is a primate. In some embodiments, the mammal is a human. In some embodiments, the mammal is not a human. In some embodiments, the subject has or is suspected of having diabetes.
[0068] As used herein, the term “insertion” which may be used interchangeably with the terms “genetic insertion” or “knock-in”, generally refers to a genetic modification wherein a nucleotide sequence is introduced or added into a site or region of genomic DNA by any molecular biological method, e.g., methods described herein, e.g., by delivering to a site of genomic DNA an endonuclease and at least one gRNA. In some embodiments, an insertion of a heterologous nucleotide sequence occurs within or near a target gene. In some embodiments, an insertion of a heterologous nucleotide sequence may occur within or near a site of genomic DNA that has been the site of a prior genetic modification, e.g., a deletion or insertion-deletion mutation. In some embodiments, an insertion occurs at a site of genomic DNA that partially overlaps, completely overlaps, or is contained within a site of a prior genetic modification, e.g., a deletion or insertiondeletion mutation. In some embodiments, an insertion simultaneously leads to a disruption of the gene at the targeted site of the insertion. In some embodiments, an insertion involves the introduction of a polynucleotide that encodes a protein of interest. In some embodiments, an insertion involves the introduction of a polynucleotide that encodes a GLP-1 receptor agonist. In some embodiments, an insertion involves the introduction of an exogenous promoter, e.g., a constitutive promoter, e.g., a CAG or CAGGS promoter. In some embodiments, the insertion is designed such that the promoter of the target gene regulates the expression of the gene product encoded by the inserted sequence. In some embodiments, the insertion is designed such that the gene product encoded by the inserted sequence is inserted in-frame with the endogenous gene, thereby generating a fusion between the endogenous gene and the inserted sequence. In some embodiments, an insertion involves the introduction of a nucleotide sequence that encodes a noncoding gene. In general, the template sequence for the insertion (e.g., in a donor plasmid) is flanked by sequences (e.g., homology arms) having substantial sequence identity with genomic DNA at or near the site of insertion.
[0069] As used herein, the term “unmodified cell” refers to a cell that has not been subjected to a particular genetic modification. In some embodiments, the term “unmodified cell” can refer to a cell that has not be subjected to any genetic modification. In some embodiments, the term “unmodified cell” can refer to a cell that has not be subjected to a genetic modification comprising insertion of a nucleotide sequence encoding a GLP-1R agonist. In some embodiments, an unmodified cell may be a stem cell. In some embodiments, an unmodified cell may be anembryonic stem cell (ESC), an adult stem cell (ASC), an induced pluripotent stem cell (iPSC), or a hematopoietic stem or progenitor cell (HSPC) (also called a hematopoietic stem cell (HSC)). In some embodiments, an unmodified cell may be a differentiated cell. In some embodiments, an unmodified cell may be selected from somatic cells (e.g., pancreatic beta cells). If a genetically modified cell is compared “relative to an unmodified cell”, the genetically modified cell and the unmodified cell are the same cell type or share a common parent cell line, e.g., a GLP1KI iPSC is compared relative to an unmodified iPSC.
[0070] As used herein, the term “within or near a gene” refers to a site or region of genomic DNA that is an intronic or exonic component of said gene or is located proximal to said gene. In some embodiments, a site of genomic DNA is within a gene if it comprises at least a portion of an intron or exon of said gene. In some embodiments, a site of genomic DNA located near a gene may be at the 5’ or 3’ end of said gene (e.g., the 5’ or 3’ end of the coding region of said gene). In some embodiments, a site of genomic DNA located near a gene may be a promoter region or repressor region that modulates the expression of said gene. In some embodiments, a site of genomic DNA located near a gene may be on the same chromosome as said gene. In some embodiments, a site or region of genomic DNA is near a gene if it is within 50 kb, 40 kb, 30 kb, 20 kb, 10 kb, 5 kb, 1 kb, or closer to the 5’ or 3’ end of said gene (e.g., the 5’ or 3’ end of the coding region of said gene).
[0071] As used herein, the term “about” can mean plus or minus 5% of the provided value.Methods and compositions for treating diseases or disorders related to insulin signaling
[0072] Described herein are methods and compositions for providing a constitutive supply of GLP-1 or a GLP-l / GIP chimera fusion peptide to patients via a cell therapy that would continuously generate these moieties for the patient, and, in some embodiments, doing so in an insulin-controlled manner.
[0073] Described herein is a CRISPR-Cas9 gene editing approach is described herein to engineer a human pluripotent stem cell line where the GLP-1 or chimera peptide sequence is knocked into the insulin gene (INS) locus via homology directed repair (HDR). When engineered in this way, GLP-1 or the chimera peptide would be secreted into the bloodstream only upon a rising blood glucose level triggering the production of insulin. In some embodiments, this type of GLP-1 secretion can result in capturing the beneficial effects of GLP-1 while diminishing any negative gastrointestinal side effects as the peptide would not be present in circulation perpetually. Additionally, since the cell therapy would be “self-sustaining” and no additional GLP-1 injections would be necessary, this intervention can, in some embodiments, be a one-time treatment for patients.
[0074] Provided herein are multiple designs of genetic modifications for this cell therapy. A single guide RNA (sgRNA) to target the 5’ insulin (INS) gene was generated named INS-G13 (comprising a spacer sequence of SEQ ID NO: 1). This sgRNA targets a region near the signal peptide of the INS gene coding sequence to introduce a double strand break at this site. Additionally, plasmid templates have been generated for DNA repair by homology-directed repair (HDR). These plasmids contain a left and right homology arm around the INS-G13 cut site and encode a GLP-1 peptide sequence or a GLP-l / GIP chimera fusion peptide sequence based on the Tirzepatide peptide sequence (also referred to herein as TIRZ). The GLP-1 sequence is flanked by PC 1 / 3 Cleavage sites and parts of the IP-1 and IP -2 (intervening peptide) sequences that are native to the glucagon gene transcript. This design was described in Gene Therapy 17: 171-180, 2010. Additionally, in the design disclosed herein, in some embodiments, a stop codon follows the GLP-1 sequence, thus terminating any INS sequence translation following GLP-1. The TIRZ design is highly similar to chemically synthesized tirzepatide and aims to mimic this artificial chimera peptide as much as possible with certain amino acids substituting for chemical modifications. TIRZ is expected to be more stable in the bloodstream and potentially more potent than other GLP-1R agonists. Upon Cas9 and INS-G13 binding and induction of the double strand break, cells can use the HDR templates (e.g., donor plasmid) to repair the DNA. In this manner, the GLP-1 or TIRZ coding sequence is incorporated into the INS gene right after the INS signal peptide. In this configuration, the cells would not make the INS protein due to the presence of the stop codon but would instead make and secrete GLP-1 or TIRZ from the INS transcript.
[0075] In some embodiments, the genetic modifications are made in induced pluripotent stem cells (iPSCs). Then the cells would be differentiated to the pancreatic lineage to attain GLP-l / TIRZ production in an insulin gene dependent manner. Since, in some embodiments, insulin protein would not be produced by these cells (e.g., cells homozygous for the insertion), the target disease area for this cell therapy would be diabetes (e.g., type 2 diabetes) and obesity.
[0076] In some embodiments, the INS protein is not knocked out. In this configuration, a masking plasmid is introduced (INS-mask) carrying a homology donor template where the INS coding sequence is masked with synonymous amino acid changes. These synonymous changes can hide the DNA from the INS-G13 gRNAbut would not interfere with the expression of insulin. Upon introduction of this masking plasmid along with the GLP-1 or TIRZ plasmids, the cells can repair the DNA in such a way as to incorporate both templates into the genome, e.g., one allele becoming GLP-l / TIRZ positive and one allele becoming INS-mask positive. In this configuration, this cell line would be heterozygous for INS and GLP-l / TIRZ and would produce and secrete both proteins from the INS gene locus.
[0077] In some embodiments, the genetic modifications are made in iPSCs. The iPSCscan be differentiated to the pancreatic lineage to attain GLP-l / TIRZ production in an insulin gene dependent manner. Since insulin would still be produced by these cells, the target disease area for this cell therapy would be insulin requiring type 1 and type 2 diabetes.GLP-1 receptor agonist
[0078] Disclosed herein include genetically modified cells. In some embodiments, the genetically modified cell comprises: a heterologous nucleotide sequence comprising a sequence encoding a glucagon-like peptide-1 (GLP-1) Receptor (GLP-1R) agonist; and wherein the cell expresses the GLP-1 R agonist.
[0079] Disclosed herein include methods for generating a genetically modified cell. In some embodiments, the method comprises delivering to a cell: (a) an RNA-guided nuclease and a gRNA targeting a target site within the genome of the cell; and (b) a nucleic acid comprising (i) a nucleotide sequence that is at least 85% identical to a region located upstream of the target site, (ii) a heterologous nucleotide sequence encoding a GLP-1 Receptor (GLP-1R) agonist, and (iii) a nucleotide sequence that is at least 85% identical to a region located downstream of the target site; thereby the target site in the genome of the cell is cleaved and the heterologous nucleotide sequence is inserted into the genome by homology directed repair (HDR), thereby generating a genetically modified cell, wherein the genetically modified cell expresses the GLP-1R agonist.
[0080] The term "GLP-1 receptor agonist" as used herein refers to a molecule, which fully or partially activates the human GLP-1 receptor. In some embodiments, the GLP-1R agonist is also capable of binding to and activating the gastric inhibitory polypeptide (GIP) receptor. GIP can also be referred to as glucose-dependent insulinotropic polypeptide (also, GIP).
[0081] A GLP-1 receptor agonist can be a GLP-1 analogue or variant. The terms “analogue” or “variant” as used herein can be used interchangeably to refer to a GLP-1 peptide wherein at least one amino acid residue of the peptide has been substituted with another amino acid residue and / or wherein at least one amino acid residue has been deleted from the peptide and / or wherein at least one amino acid residue has been added to the peptide and / or wherein at least one amino acid residue of the peptide has been modified. Such addition or deletion of amino acid residues may take place at the N-terminal of the peptide and / or at the C-terminal of the peptide, at an internal position, or any combination thereof. The GLP-1 receptor agonist can comprise a maximum of twelve, such as a maximum of 10, 8 or 6, amino acids which have been altered, e.g., by substitution, deletion, insertion and / or modification, compared to e.g., GLP-1 (7- 37). The analogue can comprise up to 10 substitutions, deletions, additions and / or insertions, such as up to 9 substitutions, deletions, additions and / or insertions, up to 8 substitutions, deletions, additions and / or insertions, up to 7 substitutions, deletions, additions and / or insertions, up to 6substitutions, deletions, additions and / or insertions, up to 5 substitutions, deletions, additions and / or insertions, up to 4 substitutions, deletions, additions and / or insertions or up to 3 substitutions, deletions, additions and / or insertions, compared to e.g., GLP-l(7-37). Unless otherwise stated the GLP-1 comprises only L-amino acids. GLP-l(7-37) has the sequence HAEGTFTSDVSSYLEGQAAKEFIAWLVKGRG (SEQ ID NO: 34). The GLP-1R agonist can comprise an amino acid sequence comprising the sequence of SEQ ID NO: 34, a sequence that is at least 95% identical (e.g., 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values identical) to the sequence of SEQ ID NO: 34, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 34.
[0082] In vivo, GLP-1 active peptide is generated from proteolytic cleavage of proglucagon. The proglucagon gene is expressed in several organs including the pancreas (a-cells of the islets of Langerhans), gut (intestinal enteroendocrine L-cells) and brain (caudal brainstem and hypothalamus). Pancreatic proglucagon gene expression is promoted upon fasting and hypoglycemia induction and inhibited by insulin. Conversely, intestinal proglucagon gene expression is reduced during fasting and stimulated upon food consumption. In mammals, the transcription gives rise to identical mRNA in all three cell types, which is further translated to the 180 amino acid precursor called proglucagon. However, as a result of tissue-specific posttranslational processing mechanisms, different peptides are produced in the different cells.
[0083] In the pancreas (a-cells of the islets of Langerhans), proglucagon is cleaved by prohormone convertase (PC) 2 producing glicentin-related pancreatic peptide (GRPP), glucagon, intervening peptide-1 (IP-1) and major proglucagon fragment (MPGF). In the gut and brain, proglucagon is catalyzed by PC 1 / 3 giving rise to glicentin, which may be further processed to GRPP and oxyntomodulin, GLP-1, intervening peptide-2 (IP-2) and glucagon-like peptide-2 (GLP-2). Initially, GLP-1 was thought to correspond to proglucagon (72-108) suitable with the N-terminal of the MGPF, but sequencing experiments of endogenous GLP-1 revealed a structure corresponding to proglucagon (78-107) from which two discoveries were found. Firstly, the full- length GLP-1 (1-37) was found to be processed by endopeptidase to the biologically active GLP- 1 (7-37). Secondly, the glycine corresponding to proglucagon was found to serve as a substrate for amidation of the C-terminal arginine resulting in the equally potent GLP-1 (7-36) amide. In humans, almost all (>80%) secreted GLP-1 is amidated, whereas a considerable part remains GLP-1 (7-37) in other species.
[0084] In some embodiments, the heterologous nucleotide sequence further comprises one or more of: a) a sequence encoding a first pro-protein convertase cleavage site; b) a sequence encoding an intervening peptide-1 (IP-1) sequence of proglucagon; c) a sequence encoding a second pro-protein convertase cleavage site; and d) a sequence encoding at least a portion of anintervening peptide-2 (IP-2) sequence of proglucagon. In some embodiments, the heterologous nucleotide sequence comprises, from 5’ to 3’ : a) a sequence encoding a first pro-protein convertase cleavage site; b) a sequence encoding an intervening peptide- 1 (IP-1) sequence of proglucagon; c) a sequence encoding the GLP-1R agonist; d) a sequence encoding a second pro-protein convertase cleavage site; and e) a sequence encoding at least a portion of an intervening peptide- 2 (IP-2) sequence of proglucagon.
[0085] The heterologous nucleotide sequence can comprise a sequence encoding a GLP-1R agonist protein precursor (e.g., a pro-protein). The GLP1-R agonist pro-protein can comprise, from the N-terminus to the C-terminus: a) a first pro-protein convertase cleavage site; b) an intervening peptide-1 (IP-1) sequence of proglucagon; c) the GLP-1R agonist; d) a second pro-protein convertase cleavage site; and e) at least a portion of an intervening peptide-2 (IP-2) sequence of proglucagon. Thus, in some embodiments, the GLP-1R pro-protein (similar to, e.g., proglucagon) can be enzymatically processed within a cell by endogenous pro-protein convertases.
[0086] The GLP-1R agonist can, in some embodiments, comprise a GLP-1 peptide or a functional fragment thereof, or a variant thereof. The GLP-1R agonist can comprise a chimeric peptide comprising at least a portion of GLP-1 peptide and at least a portion of gastric inhibitory peptide (GIP). In some embodiments, the chimeric peptide comprises Tirzepatide or a variant thereof. Tirzepatide is a 39 amino acid linear peptide, similar in size to GIP and GLP-1, which are related hormones both belonging to the secretin family of gut peptides. GIP protein is important in maintaining glucose homeostasis as it is a potent stimulator of insulin secretion from pancreatic beta-cells following food ingestion and nutrient absorption. This gene stimulates insulin secretion via its G protein-coupled receptor activation of adenylyl cyclase and other signal transduction pathways. It is a relatively poor inhibitor of gastric acid secretion. The starting amino acid sequence of Tirzepatide was that of human GIP and tirzepatide retains 9 homologous amino acids from this peptide as well as 10 amino acids shared by GIP and GLP-1. Four amino acids correspond to the same position in the GLP-1 molecule, and 10 amino-terminal amino acids are identical to those in the sequence of exendin-4 (a peptide from the saliva of a lizard called Heloderma suspectum, later becoming exenatide, the first GLP-1 RA). In some embodiments, tirzepatide is a chemically synthesized peptide that can comprise the sequence of SEQ ID NO: 91. In some embodiments, the tirzepatide sequence has been modified for expression in a cell. Trizepatide can comprise the sequence ofYGEGTFTSDYSIYLDKIAQKAFVQWLIAGGPSSGAPPPS (SEQ ID NO: 35). The GLP-1R agonist can comprise an amino acid sequence comprising the sequence of SEQ ID NO: 35, a sequence that is at least 95% identical (e.g., 95%, 96%, 97%, 98%, 99%, 100%, or a number or arange between any two of these values identical) to the sequence of SEQ ID NO: 35, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 35.
[0087] In some embodiments, the GLP-1R agonists comprises a chimeric peptide. The chimeric peptide can comprise at least a portion of GLP-1, GIP, exendin-4 / exenatide, or any combination thereof. In some embodiments, the at least a portion of GLP-1 comprises one or more segments of GLP-1 peptide, wherein each of the one or more segments is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids in length. In some embodiments, the at least a portion of GIP comprises one or more segments of GIP peptide, wherein each of the one or more segments is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids in length. In some embodiments, the at least a portion of exendin-4 / exenatide comprises one or more segments of exendin-4 / exenatide, wherein each of the one or more segments is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids in length. In some embodiments, a particular residue or segment of the chimeric peptide may be shared by GLP-1 and GIP, GLP-1 and exendin- 4 / exenatide, GIP and exendin-4 / exenatide, or all of GLP-1, GIP, and exendin-4 / exenatide. In some embodiments, the chimeric peptide may comprise one or more segments comprising amino acid sequence that is 100% identical between GLP-1 and GIP, GLP-1 and exendin-4 / exenatide, GIP and exendin-4 / exenatide, or all of GLP-1, GIP, and exendin-4 / exenatide. For example, in some embodiments, positions 3-6 of the chimeric peptide may be identical between GLP-1, GIP, and GLP-1 and GIP, GLP-1 and exendin-4 / exenatide, GIP and exendin-4 / exenatide, or all of GLP-1, GIP, and exendin-4 / exenatide.
[0088] Each segment of the at least a portion of GLP-1 peptide, GIP peptide, and exendin-4 / exenatide can be in any order, from N-terminus to C-terminus, of the chimeric peptide. In some embodiments, a particular residue of the chimeric peptide may be unique to GLP-1, GIP, or exendin-4 / exenatide. In some embodiments, a particular amino acid residue of the chimeric peptide is not found in any of GLP-1, GIP and exendin-4 / exenatide. In some embodiments, a particular residue of the chimeric peptide is shared by GLP-1 and GIP, GLP-1 and exendin- 4 / exenatide, GIP and exendin-4 / exenatide, or all of GLP-1, GIP, and exendin-4 / exenatide.
[0089] The cell can be heterozygous or homozygous for the insertion. Thus, in some embodiments, a cell heterozygous for the insertion may comprise a wild-type allele of the insulin gene or a loss-of-function mutation in the insulin gene without integration of the heterologous nucleotide sequence. A cell comprising heterozygous insertion of the heterologous sequence and wild-type copy of INS gene can express both the GLP-1R agonist and insulin. A cell comprising heterozygous insertion of the heterologous sequence and a mutation (e.g., an indel mutation) in the INS gene, can express the GLP-1R agonist and does not express insulin. A cell homozygous for the insertion will lack a wild-type copy of the insulin gene, and therefore, not express insulin. The cell (either heterozygous or homozygous for the insertion) can be capable of secreting theGLP-1R agonist.
[0090] The heterologous nucleotide sequence can be inserted into any location in the genome. For example, the heterologous nucleotide sequence can be inserted into a safe harbor locus. As used herein, the term “safe harbor locus” generally refers to any location, site, or region of genomic DNA that may be able to accommodate a genetic insertion into said location, site, or region without adverse effects on a cell. In some embodiments, a safe harbor locus is an intragenic or extragenic region. In some embodiments, a safe harbor locus is a region of genomic DNA that is typically transcriptionally silent. In some embodiments, the safe harbor locus is selected from AAVS1 gene locus, HRPT gene locus, CCR5 gene locus, globin gene locus, TTR gene locus, TF gene locus, F9 gene locus, Alb gene locus, Gys2 gene locus and PCSK9 gene locus. The heterologous nucleotide sequence can be inserted into the insulin (INS) gene locus of the genome of the cell. In some embodiments, the insertion of the heterologous nucleotide sequence disrupts the INS gene, thereby the expression of insulin by the cell can be reduced or eliminated.
[0091] In some embodiments, the heterologous nucleotide sequence encoding the GLP-1R agonist is operably linked to a heterologous promoter sequence (e.g., a eukaryotic promoter that is active in mammalian cells). Non-limiting examples of suitable eukaryotic promoters (i.e., promoters functional in a eukaryotic cell) include those from cytomegalovirus (CMV) immediate early, herpes simplex virus (HSV) thymidine kinase, early and late SV40, long terminal repeats (LTRs) from retrovirus, human elongation factor- 1 a promoter (EFla), chicken beta-actin promoter (CBA), ubiquitin C promoter (UBC), a hybrid construct comprising the cytomegalovirus enhancer fused to the chicken beta-actin promoter, a hybrid construct comprising the cytomegalovirus enhancer fused to the promoter, the first exon, and the first intron of chicken beta-actin gene (CAG or CAGGS), murine stem cell virus promoter (MSCV), phosphoglycerate kinase- 1 locus promoter (PGK), and mouse metallothionein-I promoter.
[0092] A promoter can be an inducible promoter (e.g., a heat shock promoter, tetracycline-regulated promoter, steroid-regulated promoter, metal-regulated promoter, estrogen receptor-regulated promoter, etc.). The promoter can be a constitutive promoter (e.g., CMV promoter, UBC promoter, CAG or CAGGS promoter). In some cases, the promoter can be a spatially restricted and / or temporally restricted promoter (e.g., a tissue specific promoter, a cell type specific promoter, etc.).
[0093] The heterologous nucleotide sequence can be inserted into the INS gene locus at a site downstream of the portion of the INS gene sequence that encodes insulin signal peptide. The heterologous nucleotide sequence can be inserted into the INS gene locus downstream of the portion of the INS gene sequence that encodes insulin signal peptide, thereby the cell can be capable of expressing (e.g., when inserted “in-frame”) a GLP-1R agonist fusion proteincomprising, from N-terminus to C-terminus, (i) the insulin signal peptide and the GLP-1R agonist or (ii) the insulin signal peptide and the GLP-1R agonist pro-protein. The GLP-1R agonist fusion protein can comprise, from N-terminus to C-terminus: a) the insulin signal peptide; b) the first proprotein convertase cleavage site; c) the IP-1 sequence of proglucagon; d) the GLP-1R agonist; e) the second proprotein convertase cleavage site; and f) the at least a portion of the IP -2 sequence of proglucagon.
[0094] In some embodiments, the genetically modified cell expresses the GLP-1R agonist pro-protein, the GLP-1R agonist fusion protein, or both. In some embodiments, the GLP- 1R agonist pro-protein or the GLP-1R agonist fusion protein is: i) capable of being enzymatically cleaved by one or more proprotein convertases in the cell to generate the GLP-1R agonist, and optionally ii) the GLP-1R agonist is amidated at the C-terminal end.
[0095] Insulin is synthesized as an inactive precursor molecule, a 110 amino acid-long protein called "preproinsulin". Preproinsulin is translated directly into the rough endoplasmic reticulum (RER), where its signal peptide is removed by signal peptidase to form "proinsulin". As the proinsulin folds, opposite ends of the protein, called the " A-chain" and the "B-chain", are fused together with three disulfide bonds. Folded proinsulin then transits through the Golgi apparatus and is packaged into specialized secretory vesicles. In the granule, proinsulin is cleaved by proprotein convertase 1 / 3 and proprotein convertase 2, removing the middle part of the protein, called the "C-peptide". Finally, carboxypeptidase E removes two pairs of amino acids from the protein's ends, resulting in active insulin - the insulin A- and B- chains, now connected with two disulfide bonds. The resulting mature insulin is packaged inside mature granules waiting for metabolic signals (such as leucine, arginine, glucose and mannose) and vagal nerve stimulation to be exocytosed from the cell into the circulation.
[0096] In some embodiments, expression of the GLP-1R agonist fusion protein in the genetically modified cells of the disclosure is operably linked to endogenous cis-acting regulatory sequences of insulin, and the fusion protein (e.g., comprising the insulin signal peptide and GLP- 1R agonist proprotein) can be processed in the cell similar to endogenous hormones (e.g., preproinsulin and proglucagon) to generate the precursor protein (e.g., GLP-1R agonist proprotein) and finally, active GLP-1R agonist.
[0097] The GLP-1R agonist can comprise a GLP-1 peptide or a functional fragment thereof, or a variant thereof. The GLP-1R agonist can comprise a chimeric peptide comprising at least a portion of a GLP-1 peptide and at least a portion of a gastric inhibitory peptide (GIP). The heterologous nucleotide sequence can comprise or consist of a sequence that is at least 85% identical (e.g., to the sequence of SEQ ID NO: 21 or SEQ ID NO: 23. The heterologous nucleotide sequence can comprise or consist of the sequence of SEQ ID NO: 21 or SEQ ID NO: 23. Theheterologous nucleotide sequence can comprise or consist of a sequence that is at least 85% identical (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of SEQ ID NO: [GLP-1 insert] or SEQ ID NO: 10. The heterologous nucleotide sequence can comprise or consist of the sequence of SEQ ID NO: 8 or SEQ ID NO: 10. The heterologous nucleotide sequence can comprise a stop codon at the 3’ end. In some embodiments, the presence of the stop codon can prevent translation of endogenous sequence, e.g., insulin gene sequence downstream of the insertion.
[0098] The GLP-1R agonist can comprise or consist of an amino acid sequence comprising the sequence of SEQ ID NO: 34, a sequence that is at least 95% identical (e.g., 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of SEQ ID NO: 34, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 34. The GLP-1R agonist can comprise or consist of an amino acid sequence comprising the sequence of SEQ ID NO: 35, a sequence that is at least 95% identical (e.g., 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of SEQ ID NO: 35, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 35.
[0099] The GLP-1R agonist pro-protein can comprise or consist of an amino acid sequence comprising the sequence of SEQ ID NO: 29, a sequence that is at least 95% identical (e.g95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of SEQ ID NO: 29, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 29. The GLP-1R agonist pro-protein can comprise or consist of an amino acid sequence comprising the sequence of SEQ ID NO: 30, a sequence that is at least 95% identical (e.g., 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of SEQ ID NO: 30, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 30.
[0100] The GLP-1R agonist fusion protein can comprise or consist of an amino acid sequence comprising the sequence of SEQ ID NO: 32, a sequence that is at least 95% identical (e.g., 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of SEQ ID NO: 32, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 32. The GLP-1R agonist fusion protein can comprise or consist of an amino acid sequence comprising the sequence of SEQ ID NO: 33, a sequence that is at least 95% identical (e.g., 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of SEQ ID NO: 33, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 33.Genome Editing Methods
[0101] Also provided herein are methods for generating any of the genetically modified cells described herein. Genome editing generally refers to the process of modifying the nucleotide sequence of a genome, preferably in a precise or pre-determined manner. In some embodiments, genome editing methods as described herein, e.g., the CRISPR-endonuclease system, may be used to genetically modify a cell as described herein. Disclosed herein include methods for generating a genetically modified cell. In some embodiments, the method comprises delivering to a cell: (a) an RNA-guided nuclease and a gRNA targeting a target site within the genome of the cell; and (b) a nucleic acid comprising (i) a nucleotide sequence that is at least 85% identical to a region located upstream of the target site, (ii) a heterologous nucleotide sequence encoding a GLP-1 Receptor (GLP-1R) agonist, and (iii) a nucleotide sequence that is at least 85% identical to a region located downstream of the target site; thereby the target site in the genome of the cell is cleaved and the heterologous nucleotide sequence is inserted into the genome by homology directed repair (HDR), thereby generating a genetically modified cell, wherein the genetically modified cell expresses the GLP-1R agonist.
[0102] The cell can be heterozygous or homozygous for the insertion. Thus, in some embodiments, a cell heterozygous for the insertion may comprise a wild-type allele of the insulin gene or a loss-of-function mutation in the insulin gene without integration of the heterologous nucleotide sequence. A cell homozygous for the insertion will lack a wild-type copy of the insulin gene, and therefore, not express insulin. The cell (either heterozygous or homozygous for the insertion) can be capable of secreting the GLP-1R agonist.
[0103] In some embodiments, the heterologous nucleotide sequence further comprises one or more of: a) a sequence encoding a first pro-protein convertase cleavage site; b) a sequence encoding an intervening peptide- 1 (IP-1) sequence of proglucagon; c) a sequence encoding a second pro-protein convertase cleavage site; and d) a sequence encoding at least a portion of an intervening peptide-2 (IP-2) sequence of proglucagon. In some embodiments, the heterologous nucleotide sequence comprises, from 5’ to 3’ : a) a sequence encoding a first pro-protein convertase cleavage site; b) a sequence encoding an intervening peptide- 1 (IP-1) sequence of proglucagon; c)sequence encoding the GLP-1R agonist; d) a sequence encoding a second pro-protein convertase cleavage site; and e) a sequence encoding at least a portion of an intervening peptide-2 (IP-2) sequence of proglucagon.
[0104] The heterologous nucleotide sequence can comprise a sequence encoding a GLP-1 R agonist protein precursor (pro-protein). The GLP-R agonist pro-protein can comprise, from the N-terminus to the C-terminus: a) a first pro-protein convertase cleavage site; b) anintervening peptide-1 (IP-1) sequence of proglucagon; c) the GLP-1R agonist; d) a second proprotein convertase cleavage site; and e) at least a portion of an intervening peptide-2 (IP -2) sequence of proglucagon.
[0105] The target site can be within a gene locus. In some embodiments, the insertion of the heterologous sequence into the target site disrupts the gene, thereby reducing or eliminating expression of a gene product encoded by the gene. The gene locus can be a safe harbor locus. In some embodiments, the safe harbor locus is selected from the group consisting of AAVS1 gene locus, HRPT gene locus, CCR5 gene locus, globin gene locus, TTR gene locus, TF gene locus, F9 gene locus, Alb gene locus, Gys2 gene locus and PCSK9 gene locus.
[0106] The gene locus can be the insulin (INS) gene locus. In some embodiments, insertion of the heterologous nucleotide sequence disrupts the INS gene, thereby the expression of insulin by the cell can be reduced or eliminated. The target site can be a site downstream of the sequence of the INS gene that encodes insulin signal peptide.
[0107] The heterologous nucleotide sequence can be inserted into the INS locus downstream of the portion of the INS gene sequence that encodes insulin signal peptide, thereby the cell can be capable of expressing a GLP-1R agonist fusion protein comprising, from N- terminus to C-terminus, (i) the insulin signal peptide and the GLP-1R agonist or (ii) the insulin signal peptide and the GLP-1R agonist pro-protein. The GLP-1R agonist fusion protein can comprise, from N-terminus to C-terminus: a) the insulin signal peptide; b) the first proprotein convertase cleavage site; c) the IP-1 sequence of proglucagon; d) the GLP-1R agonist; e) the second proprotein convertase cleavage site; and f) the at least a portion of the IP -2 sequence of proglucagon.
[0108] In some embodiments, the genetically modified cell expresses the GLP-1R agonist pro-protein, the GLP-1R agonist fusion protein, or both. In some embodiments, the GLP- 1R agonist pro-protein or the GLP-1R agonist fusion protein is capable of being enzymatically cleaved by one or more proprotein convertases in the cell to generate the GLP-1R agonist, and the GLP-1R agonist can be optionally amidated at the C-terminal end.
[0109] The GLP-1R agonist can comprise a GLP-1 peptide or a functional fragment thereof, or a variant thereof. The GLP-1R agonist can comprise a chimeric peptide comprising at least a portion of GLP-1 peptide and at least a portion of gastric inhibitory peptide (GIP). The heterologous nucleotide sequence can comprise or consist of a sequence that is at least 85% identical (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of SEQ ID NO: 21 or SEQ ID NO: 23. The heterologous nucleotide sequence can comprise or consist of the sequence of SEQ ID NO: 21 or SEQ ID NO: 23. The heterologous nucleotide sequence cancomprise or consist of a sequence that is at least 85% identical (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of SEQ ID NO: 8 or SEQ ID NO: 10. The heterologous nucleotide sequence can comprise or consist of the sequence of SEQ ID NO: 8 or SEQ ID NO: 10. In some embodiments, the heterologous sequence is utilized as a template for homology- directed repair, thereby inserting the heterologous sequence into the genome of the cell. The heterologous nucleotide sequence can comprise a stop codon at the 3’ end, thus inhibiting translation of endogenous sequence downstream of the insertion.
[0110] The heterologous nucleotide sequence encoding a GLP-1 Receptor (GLP-1R) agonist can be flanked by sequence having homology (e.g., homology arms) to a region of the target gene locus (to facilitate homology directed repair). In some embodiments, each of the homology arms have at least 85% identity (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to a portion of the INS gene. The sequence of the upstream homology arm (e.g., left homology arm) can comprise or consist of the sequence of any one of SEQ ID NOs: 6, 19, and 96 or a sequence that is at least 85% identical (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of any one of SEQ ID NOs: 6, 19, and 96. The sequence of the downstream homology arm (e.g., the right homology arm) can comprise or consist of the sequence of SEQ ID NO: 22 or SEQ ID NO: 9 or a sequence that is at least 85% identical (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of SEQ ID NO: 22 or SEQ ID NO: 9.[oni] Provided herein are nucleic acids (e.g., donor plasmids) comprising, e.g., nucleotide sequence encoding a GLP-1R agonist, flanked by homology arms that have at least 85% sequence identity to a region in the genome (e.g., in the insulin gene). The nucleic acid can comprise or consist of the sequence of any one of SEQ ID NOs: 11-12 and 97-98 or a sequence that is at least 85% identical (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of any one of SEQ ID NOs: 11-12 and 97-98.
[0112] The GLP-1R agonist can comprise or consist of an amino acid sequence comprising the sequence of SEQ ID NO: 34, a sequence that is at least 95% identical (e.g., 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of SEQ ID NO: 34, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 34. The GLP-1R agonist can comprise or consist of an amino acid sequence comprising the sequence of SEQ ID NO: 35, a sequence that is at least 95% identical(e.g., 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of SEQ ID NO: 35, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 35.
[0113] In some embodiments, the heterologous nucleotide sequence encodes a GLP- 1R agonist precursor (e.g., a proprotein) that is capable of being enzymatically cleaved (e.g., processed) in a cell. The GLP-1R agonist pro-protein can comprise or consist of an amino acid sequence comprising the sequence of SEQ ID NO: 29, a sequence that is at least 95% identical (e.g., 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of SEQ ID NO: 29, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 29. The GLP-1R agonist pro-protein can comprise or consist of an amino acid sequence comprising the sequence of SEQ ID NO: 30, a sequence that is at least 95% identical (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of SEQ ID NO: 30, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 30.
[0114] In some embodiments, the heterologous nucleotide sequence encodes a GLP- 1R agonist precursor (e.g., a proprotein) that is capable of being enzymatically cleaved (e.g., processed) in a cell. In some embodiments, the heterologous nucleotide sequence is inserted into the INS gene “in-frame,” such that a fusion protein (e.g., a GLP-1R agonist fusion protein) is generated from the INS gene locus that comprises the insulin signal peptide (e.g., a prepropeptide). The GLP-1R agonist fusion protein can comprise or consist of an amino acid sequence comprising the sequence of SEQ ID NO: 32, a sequence that is at least 95% identical (e.g., 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of SEQ ID NO: 32, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 32. The GLP-1R agonist fusion protein can comprise or consist of an amino acid sequence comprising the sequence of SEQ ID NO: 33, a sequence that is at least 95% identical (e.g., 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to the sequence of SEQ ID NO: 33, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 33.
[0115] Examples of methods of genome editing described herein include methods of using site-directed nucleases to cut deoxyribonucleic acid (DNA) at precise target locations in the genome, thereby creating single-strand or double-strand DNA breaks at particular locations within the genome. Such breaks can be and regularly are repaired by natural, endogenous cellular processes, such as homology-directed repair (HDR) and non-homologous end joining (NHEJ), as described in Cox et al., “Therapeutic genome editing: prospects and challenges,” Nature Medicine,2015, 21(2), 121-31. These two main DNA repair processes consist of a family of alternative pathways. NHEJ directly joins the DNA ends resulting from a double-strand break, sometimes with the loss or addition of nucleotide sequence, which may disrupt or enhance gene expression. HDR utilizes a homologous sequence, or donor sequence, as a template for inserting a defined DNA sequence at the break point. The homologous sequence can be in the endogenous genome, such as a sister chromatid. Alternatively, the donor sequence can be an exogenous polynucleotide, such as a plasmid, a single-strand oligonucleotide, a double-stranded oligonucleotide, a duplex oligonucleotide or a virus, that has regions (e.g., left and right homology arms) of high homology with the nuclease-cleaved locus, but which can also contain additional sequence or sequence changes including deletions that can be incorporated into the cleaved target locus. A third repair mechanism can be microhomology-mediated end joining (MMEJ), also referred to as "Alternative NHEJ,” in which the genetic outcome is similar to NHEJ in that small deletions and insertions can occur at the cleavage site. MMEJ can make use of homologous sequences of a few base pairs flanking the DNA break site to drive a more favored DNA end joining repair outcome, and recent reports have further elucidated the molecular mechanism of this process; see, e.g., Cho and Greenberg, Nature, 2015, 518, 174-76; Kent et al., Nature Structural and Molecular Biology, 2015, 22(3):230-7; Mateos-Gomez et al., Nature, 2015, 518, 254-57; Ceccaldi et al., Nature, 2015, 528, 258-62. In some instances, it may be possible to predict likely repair outcomes based on analysis of potential microhomologies at the site of the DNA break.
[0116] Each of these genome editing mechanisms can be used to create desired genetic modifications. A step in the genome editing process can be to create one or two DNA breaks, the latter as double-strand breaks or as two single-stranded breaks, in the target locus as near the site of intended mutation. This can be achieved via the use of endonucleases, as described and illustrated herein. In general, the genome editing methods described herein can be in vitro or ex vivo methods. In some embodiments, the genome editing methods disclosed herein are not methods for treatment of the human or animal body by therapy and / or are not processes for modifying the germ line genetic identity of human beings.CRISPR Endonuclease System
[0117] The CRISPR-endonuclease system is a naturally occurring defense mechanism in prokaryotes that has been repurposed as an RNA-guided DNA-targeting platform used for gene editing. CRISPR systems include Types I, II, III, IV, V, and VI systems. In some aspects, the CRISPR system is a Type II CRISPR / Cas9 system. In other aspects, the CRISPR system is a Type V CRISPR / Cpf system. CRISPR systems rely on a DNA endonuclease, e.g., Cas9, and two noncoding RNAs - crisprRNA (crRNA) and trans-activating RNA (tracrRNA) - to target the cleavage of DNA.
[0118] The crRNA drives sequence recognition and specificity of the CRISPR- endonuclease complex through Watson-Crick base pairing, typically with a ~20 nucleotide (nt) sequence in the target DNA (e.g., a protospacer sequence). Changing the sequence of the 5’ 20 nt in the crRNA (e.g., the spacer sequence) allows targeting of the CRISPR-endonuclease complex to specific loci. The CRISPR-endonuclease complex only binds DNA sequences that contain a sequence match to the first 20 nt of the single-guide RNA (sgRNA) if the target sequence is followed by a specific short DNA motif (with the sequence NGG for, e.g., S. pyogenes Cas9) referred to as a protospacer adjacent motif (PAM). Complementarity between the spacer sequence and sequence on the non-PAM strand permits hybridization of the spacer to DNA sequence at the target site. TracrRNA hybridizes with the 3’ end of crRNA to form an RNA-duplex structure that is bound by the endonuclease to form the catalytically active CRISPR-endonuclease complex, which can then cleave the target DNA.
[0119] Once the CRISPR-endonuclease complex is bound to DNA at a target site, two independent nuclease domains within the endonuclease each cleave one of the DNA strands three bases upstream of the PAM site, leaving a double-strand break (DSB) where both strands of the DNA terminate in a blunt end.
[0120] In some embodiments, the endonuclease is a Cas9 (CRISPR associated protein 9). In some embodiments, the Cas9 endonuclease is from Streptococcus pyogenes, although other Cas9 homologs may be used, c.g, S. aureus Cas9, N. meningitidis Cas9, S. thermophilus CRISPR 1 Cas9, S. thermophilus CRISPR 3 Cas9, or T. denticola Cas9. In other instances, the CRISPR endonuclease is Cpfl, c.g, L. bacterium ND2006 Cpfl or Acidaminococcus sp. BV3L6 Cpfl. In some embodiments, the endonuclease is Casl, CaslB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csxl2), CaslOO, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4, or Cpfl endonuclease. In some embodiments, wild-type variants may be used. In some embodiments, modified versions (e.g., a homolog thereof, a recombination of the naturally occurring molecule thereof, codon-optimized thereof, or modified versions thereof) of the preceding endonucleases may be used. The CRISPR nuclease can be linked to at least one nuclear localization signal (NLS). The at least one NLS can be located at or within 50 amino acids of the amino-terminus of the CRISPR nuclease and / or at least one NLS can be located at or within 50 amino acids of the carboxy -terminus of the CRISPR nuclease.
[0121] Exemplary CRISPR / Cas polypeptides include the Cas9 polypeptides as published in Fonfara et al., “Phylogeny of Cas9 determines functional exchangeability of dual- RNA and Cas9 among orthologous type II CRISPR-Cas systems,” Nucleic Acids Research, 2014,42: 2577-2590. The CRISPR / Cas gene naming system has undergone extensive rewriting since the Cas genes were discovered. Fonfara et al. also provides PAM sequences for the Cas9 polypeptides from various species.Zinc Finger Nucleases
[0122] Zinc finger nucleases (ZFNs) are modular proteins comprised of an engineered zinc finger DNA binding domain linked to the catalytic domain of the type II endonuclease Fokl. Because Fokl functions only as a dimer, a pair of ZFNs must be engineered to bind to cognate target “half-site” sequences on opposite DNA strands and with precise spacing between them to enable the catalytically active Fokl dimer to form. Upon dimerization of the Fokl domain, which itself has no sequence specificity per se, a DNA double-strand break is generated between the ZFN half-sites as the initiating step in genome editing.
[0123] The DNA binding domain of each ZFN is typically comprised of 3-6 zinc fingers of the abundant Cys2-His2 architecture, with each finger primarily recognizing a triplet of nucleotides on one strand of the target DNA sequence, although cross-strand interaction with a fourth nucleotide also can be important. Alteration of the amino acids of a finger in positions that make key contacts with the DNA alters the sequence specificity of a given finger. Thus, a four- finger zinc finger protein will selectively recognize a 12 bp target sequence, where the target sequence is a composite of the triplet preferences contributed by each finger, although triplet preference can be influenced to varying degrees by neighboring fingers. An important aspect of ZFNs is that they can be readily re-targeted to almost any genomic address simply by modifying individual fingers. In most applications of ZFNs, proteins of 4-6 fingers are used, recognizing 12- 18 bp respectively. Hence, a pair of ZFNs will typically recognize a combined target sequence of 24-36 bp, not including the typical 5-7 bp spacer between half-sites. The binding sites can be separated further with larger spacers, including 15-17 bp. A target sequence of this length is likely to be unique in the human genome, assuming repetitive sequences or gene homologs are excluded during the design process. Nevertheless, the ZFN protein-DNA interactions are not absolute in their specificity so off-target binding and cleavage events do occur, either as a heterodimer between the two ZFNs, or as a homodimer of one or the other of the ZFNs. The latter possibility has been effectively eliminated by engineering the dimerization interface of the Fokl domain to create “plus” and “minus” variants, also known as obligate heterodimer variants, which can only dimerize with each other, and not with themselves. Forcing the obligate heterodimer prevents formation of the homodimer. This has greatly enhanced specificity of ZFNs, as well as any other nuclease that adopts these Fokl variants.
[0124] A variety of ZFN-based systems have been described in the art, modifications thereof are regularly reported, and numerous references describe rules and parameters that areused to guide the design of ZFNs; see, e.g., Segal et al., Proc Natl Acad Sci, 1999 96(6):2758-63; Dreier B et al., J Mol Biol., 2000, 303(4):489-502; Liu Q et al., J Biol Chem., 2002, 277(6):3850- 6; Dreier et al., J Biol Chem., 2005, 280(42):35588-97; and Dreier et al., J Biol Chem. 2001, 276(31):29466-78.Transcription Activator-Like Effector Nucleases (TALENs)
[0125] TALENs represent another format of modular nucleases whereby, as with ZFNs, an engineered DNA binding domain is linked to the FokI nuclease domain, and a pair of TALENs operate in tandem to achieve targeted DNA cleavage. The major difference from ZFNs is the nature of the DNA binding domain and the associated target DNA sequence recognition properties. The TALEN DNA binding domain derives from TALE proteins, which were originally described in the plant bacterial pathogen Xanthomonas sp. TALEs are comprised of tandem arrays of 33-35 amino acid repeats, with each repeat recognizing a single base pair in the target DNA sequence that is typically up to 20 bp in length, giving a total target sequence length of up to 40 bp. Nucleotide specificity of each repeat is determined by the repeat variable diresidue (RVD), which includes just two amino acids at positions 12 and 13. The bases guanine, adenine, cytosine and thymine are predominantly recognized by the four RVDs: Asn-Asn, Asn-Ile, His- Asp and Asn-Gly, respectively. This constitutes a much simpler recognition code than for zinc fingers, and thus represents an advantage over the latter for nuclease design. Nevertheless, as with ZFNs, the protein-DNA interactions of TALENs are not absolute in their specificity, and TALENs have also benefitted from the use of obligate heterodimer variants of the FokI domain to reduce off-target activity.
[0126] Additional variants of the FokI domain have been created that are deactivated in their catalytic function. If one half of either a TALEN or a ZFN pair contains an inactive FokI domain, then only single-strand DNA cleavage (nicking) will occur at the target site, rather than a DSB. The outcome is comparable to the use of CRISPR / Cas9 or CRISPR / Cpfl “nickase” mutants in which one of the Cas9 cleavage domains has been deactivated. DNA nicks can be used to drive genome editing by HDR, but at lower efficiency than with a DSB. The main benefit is that off-target nicks are quickly and accurately repaired, unlike the DSB, which is prone to NHEJ- mediated mis-repair.
[0127] A variety of TALEN-based systems have been described in the art, and modifications thereof are regularly reported; see, e.g., Boch, Science, 2009 326(5959): 1509-12; Mak et al., Science, 2012, 335(6069):716-9; and Moscou et al., Science, 2009, 326(5959): 1501. The use of TALENs based on the "Golden Gate" platform, or cloning scheme, has been described by multiple groups; see, e.g., Cermak et al., Nucleic Acids Res., 2011, 39(12):e82; Li et al., Nucleic Acids Res., 2011, 39( 14) :6315-25; Weber et al., PLoS One., 2011, 6(2):el6765; Wang etal., J Genet Genomics, 2014, 41(6):339-47.; and Cermak T et al., Methods Mol Biol., 2015 1239: 133-59.Homing Endonucleases
[0128] Homing endonucleases (HEs) are sequence-specific endonucleases that have long recognition sequences (14-44 base pairs) and cleave DNA with high specificity - often at sites unique in the genome. There are at least six known families of HEs as classified by their structure, including GIY-YIG, His-Cis box, H-N-H, PD-(DZE)xK, and Vsr-like that are derived from a broad range of hosts, including eukarya, protists, bacteria, archaea, cyanobacteria and phage. As with ZFNs and TALENs, HEs can be used to create a DSB at a target locus as the initial step in genome editing. In addition, some natural and engineered HEs cut only a single strand of DNA, thereby functioning as site-specific nickases. The large target sequence of HEs and the specificity that they offer have made them attractive candidates to create site-specific DSBs.
[0129] A variety of HE-based systems have been described in the art, and modifications thereof are regularly reported; see, e.g., the reviews by Steentoft et al., Glycobiology, 2014, 24(8):663-80; Belfort and Bonocora, Methods Mol Biol., 2014, 1123: 1-26; and Hafez and Hausner, Genome, 2012, 55(8): 553 -69.MegaTAL / Tev-mTALEN / MegaTev
[0130] As further examples of hybrid nucleases, the MegaTAL platform and Tev- mTALEN platform use a fusion of TALE DNA binding domains and catalytically active HEs, taking advantage of both the tunable DNA binding and specificity of the TALE, as well as the cleavage sequence specificity of the HE; see, e.g., Boissel et al., Nucleic Acids Res., 2014, 42: 2591-2601; Kleinstiver et al., G3, 2014, 4: 1155-65; and Boissel and Scharenberg, Methods Mol. Biol., 2015, 1239: 171-96.
[0131] In a further variation, the MegaTev architecture is the fusion of a meganuclease (Mega) with the nuclease domain derived from the GIY-YIG homing endonuclease LTevI (Tev). The two active sites are positioned ~30 bp apart on a DNA substrate and generate two DSBs with non-compatible cohesive ends; see, e.g., Wolfs et al., Nucleic Acids Res., 2014, 42, 8816-29. It is anticipated that other combinations of existing nuclease-based approaches will evolve and be useful in achieving the targeted genome modifications described herein. dCas9-FokI or dCpfl-Fokl and Other Nucleases
[0132] Combining the structural and functional properties of the nuclease platforms described above offers a further approach to genome editing that can potentially overcome some of the inherent deficiencies. As an example, the CRISPR genome editing system typically uses a single Cas9 endonuclease to create a DSB. The specificity of targeting is driven by a 20 or 24 nucleotide sequence in the guide RNA that undergoes Watson-Crick base-pairing with the targetDNA (plus an additional 2 bases in the adjacent NAG or NGG PAM sequence in the case of Cas9 from S. pyogenes). Such a sequence is long enough to be unique in the human genome, however, the specificity of the RNA / DNA interaction is not absolute, with significant promiscuity sometimes tolerated, particularly in the 5’ half of the target sequence, effectively reducing the number of bases that drive specificity. One solution to this has been to completely deactivate the Cas9 or Cpfl catalytic function - retaining only the RNA-guided DNA binding function - and instead fusing a FokI domain to the deactivated Cas9; see, e.g., Tsai et al., Nature Biotech, 2014, 32: 569-76; and Guilinger et al., Nature Biotech., 2014, 32: 577-82. Because FokI must dimerize to become catalytically active, two guide RNAs are required to tether two FokI fusions in close proximity to form the dimer and cleave DNA. This essentially doubles the number of bases in the combined target sites, thereby increasing the stringency of targeting by CRISPR-based systems.
[0133] As further example, fusion of the TALE DNA binding domain to a catalytically active HE, such as I-TevI, takes advantage of both the tunable DNA binding and specificity of the TALE, as well as the cleavage sequence specificity of I-TevI, with the expectation that off-target cleavage can be further reduced.RNA-Guided Endonucleases
[0134] The RNA-guided endonuclease systems as used herein can comprise an amino acid sequence having at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% amino acid sequence identity to a wild-type exemplary endonuclease, e.g., Cas9 from S. pyogenes, US2014 / 0068797 SEQ ID NO: 8 or Sapranauskas et al., Nucleic Acids Res, 39(21): 9275-9282 (2011). The endonuclease can comprise at least 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wild-type endonuclease (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids. The endonuclease can comprise at most: 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wild-type endonuclease (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids. The endonuclease can comprise at least: 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wild-type endonuclease (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids in a HNH nuclease domain of the endonuclease. The endonuclease can comprise at most: 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wild-type endonuclease (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids in a HNH nuclease domain of the endonuclease. The endonuclease can comprise at least: 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wild-type endonuclease (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids in a RuvC nuclease domain of the endonuclease. The endonuclease can comprise at most: 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wild-type endonuclease (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids in a RuvC nuclease domainof the endonuclease.
[0135] The endonuclease can comprise a modified form of a wild-type exemplary endonuclease. The modified form of the wild-type exemplary endonuclease can comprise a mutation that reduces the nucleic acid-cleaving activity of the endonuclease. The modified form of the wild-type exemplary endonuclease can have less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nucleic acid-cleaving activity of the wild-type exemplary endonuclease (e.g., Cas9 from S. pyogenes, supra). The modified form of the endonuclease can have no substantial nucleic acid-cleaving activity. When an endonuclease is a modified form that has no substantial nucleic acid-cleaving activity, it is referred to herein as "enzymatically inactive."
[0136] Mutations contemplated can include substitutions, additions, and deletions, or any combination thereof. The mutation converts the mutated amino acid to alanine. The mutation converts the mutated amino acid to another amino acid (e.g., glycine, serine, threonine, cysteine, valine, leucine, isoleucine, methionine, proline, phenylalanine, tyrosine, tryptophan, aspartic acid, glutamic acid, asparagine, glutamine, histidine, lysine, or arginine). The mutation converts the mutated amino acid to a non-natural amino acid (e.g., selenomethionine). The mutation converts the mutated amino acid to amino acid mimics (e.g., phosphomimics). The mutation can be a conservative mutation. For example, the mutation converts the mutated amino acid to amino acids that resemble the size, shape, charge, polarity, conformation, and / or rotamers of the mutated amino acids (e.g., cysteine / serine mutation, lysine / asparagine mutation, histidine / phenylalanine mutation). The mutation can cause a shift in reading frame and / or the creation of a premature stop codon. Mutations can cause changes to regulatory regions of genes or loci that affect expression of one or more genes.Guide RNAs
[0137] The present disclosure provides guide RNAs (gRNAs) that can direct the activities of an associated endonuclease to a specific target site within a polynucleotide. A guide RNA can comprise at least a spacer sequence that hybridizes to a target nucleic acid sequence of interest, and a CRISPR repeat sequence. In CRISPR Type II systems, the gRNA also comprises a second RNA called the tracrRNA sequence. In the CRISPR Type II guide RNA (gRNA), the CRISPR repeat sequence and tracrRNA sequence hybridize to each other to form a duplex. In CRISPR Type V systems, the gRNA comprises a crRNA that forms a duplex. In some embodiments, a gRNA can bind an endonuclease, such that the gRNA and endonuclease form a complex. The gRNA can provide target specificity to the complex by virtue of its association with the endonuclease. The genome-targeting nucleic acid thus can direct the activity of theendonuclease.
[0138] Exemplary guide RNAs include a spacer sequence that comprises 15-200 nucleotides wherein the gRNA targets a genome location based on the GRCh38 human genome assembly. As is understood by persons of skill in the art, each gRNA can be designed to include a spacer sequence complementary to its genomic target site or region. See Jinek et al., Science, 2012, 337, 816-821 and Deltcheva et al., Nature, 2011, 471, 602-607.
[0139] The gRNA can be a double-molecule guide RNA. The gRNA can be a singlemolecule guide RNA. A double-molecule guide RNA can comprise two strands of RNA. The first strand comprises in the 5' to 3' direction, an optional spacer extension sequence, a spacer sequence and a minimum CRISPR repeat sequence. The second strand can comprise a minimum tracrRNA sequence (complementary to the minimum CRISPR repeat sequence), a 3’ tracrRNA sequence and an optional tracrRNA extension sequence. A single-molecule guide RNA (sgRNA) can comprise, in the 5' to 3' direction, an optional spacer extension sequence, a spacer sequence, a minimum CRISPR repeat sequence, a single-molecule guide linker, a minimum tracrRNA sequence, a 3’ tracrRNA sequence and an optional tracrRNA extension sequence. The optional tracrRNA extension can comprise elements that contribute additional functionality (e.g., stability) to the guide RNA. The single-molecule guide linker can link the minimum CRISPR repeat and the minimum tracrRNA sequence to form a hairpin structure. The optional tracrRNA extension can comprise one or more hairpins.
[0140] In some embodiments, an sgRNA comprises a 20 nucleotide spacer sequence at the 5’ end of the sgRNA sequence. The spacer sequence can also be, for example, 16 nucleotides, 17 nucleotides, 18 nucleotides, or 19 nucleotides. In some embodiments, an sgRNA comprises a less than a 20 nucleotide spacer sequence at the 5’ end of the sgRNA sequence. In some embodiments, an sgRNA comprises a more than 20 nucleotide spacer sequence at the 5’ end of the sgRNA sequence. In some embodiments, an sgRNA comprises a variable length spacer sequence with 17-30 nucleotides at the 5’ end of the sgRNA sequence. In some embodiments, an sgRNA comprises a spacer extension sequence with a length of more than 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, or 200 nucleotides. In some embodiments, an sgRNA comprises a spacer extension sequence with a length of less than 3, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides.
[0141] In some embodiments, an sgRNA comprises a spacer extension sequence that comprises another moiety (e.g., a stability control sequence, an endoribonuclease binding sequence, or a ribozyme). The moiety can decrease or increase the stability of a nucleic acid targeting nucleic acid. The moiety can be a transcriptional terminator segment (i.e., a transcription termination sequence). The moiety can function in a eukaryotic cell. The moiety can function ina prokaryotic cell. The moiety can function in both eukaryotic and prokaryotic cells. Non-limiting examples of suitable moi eties include: a 5' cap (e.g., a 7-methylguanylate cap (m7 G)), a riboswitch sequence (e.g., to allow for regulated stability and / or regulated accessibility by proteins and protein complexes), a sequence that forms a dsRNA duplex (z.e., a hairpin), a sequence that targets the RNA to a subcellular location (e.g., nucleus, mitochondria, chloroplasts, and the like), a modification or sequence that provides for tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that allows for fluorescent detection, etc.), and / or a modification or sequence that provides a binding site for proteins (e.g., proteins that act on DNA, including transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, and the like).
[0142] In some embodiments, an sgRNA comprises a spacer sequence that hybridizes to a sequence in a target polynucleotide. The spacer of a gRNA can interact with a target polynucleotide in a sequence-specific manner via hybridization (i.e., base pairing). The nucleotide sequence of the spacer can vary depending on the sequence of the target nucleic acid of interest.
[0143] The target DNA can comprise a “PAM-strand” and a complementary “non- PAM” strand. One of skill in the art recognizes that the gRNA spacer sequence hybridizes to the complementary sequence located in the non-PAM strand of the target nucleic acid of interest. Thus, the gRNA spacer sequence can be the RNA equivalent of the target, “PAM-strand” sequence (e.g., protospacer sequence). The spacer of a gRNA interacts with a target nucleic acid of interest in a sequence-specific manner via hybridization (i.e., base pairing). The nucleotide sequence of the spacer thus varies depending on the target sequence of the target nucleic acid of interest. In some embodiments, the target sequence of the Insulin gene (INS) is within exon 1 of the INS gene.
[0144] In a CRISPR-endonuclease system, a spacer sequence can be designed to hybridize to a target polynucleotide that is complementary to a sequence located 5' of a PAM of the endonuclease used in the system. The spacer may perfectly match the target sequence or may have mismatches. Each endonuclease, e.g., Cas9 nuclease, has a particular PAM sequence that it recognizes in a target DNA. For example, S. pyogenes Cas9 recognizes a PAM that comprises the sequence 5'-NRG-3', where R comprises either A or G, where N is any nucleotide and N is immediately 3' of the target nucleic acid sequence targeted by the spacer sequence.
[0145] A target polynucleotide sequence can comprise 20 nucleotides. The target polynucleotide can comprise less than 20 nucleotides. The target polynucleotide can comprise more than 20 nucleotides. The target polynucleotide can comprise at least: 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 or more nucleotides. The target polynucleotide can comprise at most:5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 or more nucleotides. The target polynucleotide sequence can comprise 20 bases immediately 5' of the first nucleotide of the PAM.
[0146] A spacer sequence can have a length of at least about 6 nucleotides (nt). The spacer sequence can be at least about 6 nt, at least about 10 nt, at least about 15 nt, at least about18 nt, at least about 19 nt, at least about 20 nt, at least about 25 nt, at least about 30 nt, at least about 35 nt or at least about 40 nt, from about 6 nt to about 80 nt, from about 6 nt to about 50 nt, from about 6 nt to about 45 nt, from about 6 nt to about 40 nt, from about 6 nt to about 35 nt, from about 6 nt to about 30 nt, from about 6 nt to about 25 nt, from about 6 nt to about 20 nt, from about 6 nt to about 19 nt, from about 10 nt to about 50 nt, from about 10 nt to about 45 nt, from about 10 nt to about 40 nt, from about 10 nt to about 35 nt, from about 10 nt to about 30 nt, from about10 nt to about 25 nt, from about 10 nt to about 20 nt, from about 10 nt to about 19 nt, from about19 nt to about 25 nt, from about 19 nt to about 30 nt, from about 19 nt to about 35 nt, from about19 nt to about 40 nt, from about 19 nt to about 45 nt, from about 19 nt to about 50 nt, from about19 nt to about 60 nt, from about 20 nt to about 25 nt, from about 20 nt to about 30 nt, from about20 nt to about 35 nt, from about 20 nt to about 40 nt, from about 20 nt to about 45 nt, from about20 nt to about 50 nt, or from about 20 nt to about 60 nt. In some examples, the spacer sequence can comprise 20 nucleotides. In some examples, the spacer can comprise 22 nucleotides.
[0147] In some examples, the percent complementarity between the spacer sequence and the target nucleic acid is at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, or 100%. In some examples, the percent complementarity between the spacer sequence and the target nucleic acid is at most about 30%, at most about 40%, at most about 50%, at most about 60%, at most about 65%, at most about 70%, at most about 75%, at most about 80%, at most about 85%, at most about 90%, at most about 95%, at most about 97%, at most about 98%, at most about 99%, or 100%. In some examples, the percent complementarity between the spacer sequence and the target nucleic acid is 100% over the six contiguous 5'-most nucleotides of the target sequence of the complementary strand of the target nucleic acid. The percent complementarity between the spacer sequence and the target nucleic acid can be at least 60% over about 20 contiguous nucleotides. The length of the spacer sequence and the target nucleic acid can differ by 1 to 6 nucleotides, which may be thought of as a bulge or bulges.
[0148] A tracrRNA sequence can comprise nucleotides that hybridize to a minimum CRISPR repeat sequence in a cell. A minimum tracrRNA sequence and a minimum CRISPR repeat sequence may form a duplex, z.e., a base-paired double-stranded structure. Together, the minimum tracrRNA sequence and the minimum CRISPR repeat can bind to an RNA-guidedendonuclease. At least a part of the minimum tracrRNA sequence can hybridize to the minimum CRISPR repeat sequence. The minimum tracrRNA sequence can be at least about 30%, about 40%, about 50%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or 100% complementary to the minimum CRISPR repeat sequence.
[0149] The minimum tracrRNA sequence can have a length from about 7 nucleotides to about 100 nucleotides. For example, the minimum tracrRNA sequence can be from about 7 nucleotides (nt) to about 50 nt, from about 7 nt to about 40 nt, from about 7 nt to about 30 nt, from about 7 nt to about 25 nt, from about 7 nt to about 20 nt, from about 7 nt to about 15 nt, from about 8 nt to about 40 nt, from about 8 nt to about 30 nt, from about 8 nt to about 25 nt, from about 8 nt to about 20 nt, from about 8 nt to about 15 nt, from about 15 nt to about 100 nt, from about 15 nt to about 80 nt, from about 15 nt to about 50 nt, from about 15 nt to about 40 nt, from about 15 nt to about 30 nt or from about 15 nt to about 25 nt long. The minimum tracrRNA sequence can be approximately 9 nucleotides in length. The minimum tracrRNA sequence can be approximately 12 nucleotides. The minimum tracrRNA can consist of tracrRNA nt 23-48 described in Jinek et al., supra.
[0150] The minimum tracrRNA sequence can be at least about 60% identical to a reference minimum tracrRNA (e.g., wild type, tracrRNA from S. pyogenes) sequence over a stretch of at least 6, 7, or 8 contiguous nucleotides. For example, the minimum tracrRNA sequence can be at least about 65% identical, about 70% identical, about 75% identical, about 80% identical, about 85% identical, about 90% identical, about 95% identical, about 98% identical, about 99% identical or 100% identical to a reference minimum tracrRNA sequence over a stretch of at least 6, 7, or 8 contiguous nucleotides.
[0151] The duplex between the minimum CRISPR RNA and the minimum tracrRNA can comprise a double helix. The duplex between the minimum CRISPR RNA and the minimum tracrRNA can comprise at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotides. The duplex between the minimum CRISPR RNA and the minimum tracrRNA can comprise at most about 1,2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotides. The duplex can comprise a mismatch (z.e., the two strands of the duplex are not 100% complementary). The duplex can comprise at least about 1, 2,3, 4, or 5 or mismatches. The duplex can comprise at most about 1, 2, 3, 4, or 5 or mismatches. The duplex can comprise no more than 2 mismatches.
[0152] In some embodiments, a tracrRNA may be a 3' tracrRNA. In some embodiments, a 3’ tracrRNA sequence can comprise a sequence with at least about 30%, about 40%, about 50%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or 100% sequence identity to a reference tracrRNA sequence (e.g., a tracrRNA from S. pyogenes).
[0153] In some embodiments, a gRNA may comprise a tracrRNA extension sequence. A tracrRNA extension sequence can have a length from about 1 nucleotide to about 400 nucleotides. The tracrRNA extension sequence can have a length of more than 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, or 200 nucleotides. The tracrRNA extension sequence can have a length from about 20 to about 5000 or more nucleotides. The tracrRNA extension sequence can have a length of less than 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides. The tracrRNA extension sequence can comprise less than 10 nucleotides in length. The tracrRNA extension sequence can be 10-30 nucleotides in length. The tracrRNA extension sequence can be 30-70 nucleotides in length.
[0154] The tracrRNA extension sequence can comprise a functional moiety (e.g., a stability control sequence, ribozyme, endoribonuclease binding sequence). The functional moiety can comprise a transcriptional terminator segment (z.e., a transcription termination sequence). The functional moiety can have a total length from about 10 nucleotides (nt) to about 100 nucleotides, from about 10 nt to about 20 nt, from about 20 nt to about 30 nt, from about 30 nt to about 40 nt, from about 40 nt to about 50 nt, from about 50 nt to about 60 nt, from about 60 nt to about 70 nt, from about 70 nt to about 80 nt, from about 80 nt to about 90 nt, or from about 90 nt to about 100 nt, from about 15 nt to about 80 nt, from about 15 nt to about 50 nt, from about 15 nt to about 40 nt, from about 15 nt to about 30 nt, or from about 15 nt to about 25 nt.
[0155] In some embodiments, an sgRNA may comprise a linker sequence with a length from about 3 nucleotides to about 100 nucleotides. In Jinek et al., supra, for example, a simple 4 nucleotide "tetraloop" (-GAAA-) was used (Jinek et al., Science, 2012, 337(6096):816- 821). An illustrative linker has a length from about 3 nucleotides (nt) to about 90 nt, from about 3 nt to about 80 nt, from about 3 nt to about 70 nt, from about 3 nt to about 60 nt, from about 3 nt to about 50 nt, from about 3 nt to about 40 nt, from about 3 nt to about 30 nt, from about 3 nt to about 20 nt, from about 3 nt to about 10 nt. For example, the linker can have a length from about 3 nt to about 5 nt, from about 5 nt to about 10 nt, from about 10 nt to about 15 nt, from about 15 nt to about 20 nt, from about 20 nt to about 25 nt, from about 25 nt to about 30 nt, from about 30 nt to about 35 nt, from about 35 nt to about 40 nt, from about 40 nt to about 50 nt, from about 50 nt to about 60 nt, from about 60 nt to about 70 nt, from about 70 nt to about 80 nt, from about 80 nt to about 90 nt, or from about 90 nt to about 100 nt. The linker of a single-molecule guide nucleic acid can be between 4 and 40 nucleotides. The linker can be at least about 100, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, or 7000 or more nucleotides. The linker can be at most about 100, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, or 7000 or more nucleotides.
[0156] Linkers can comprise any of a variety of sequences, although in some examplesthe linker will not comprise sequences that have extensive regions of homology with other portions of the guide RNA, which might cause intramolecular binding that could interfere with other functional regions of the guide. In Jinek et al., supra, a simple 4 nucleotide sequence - GAAA- was used (Jinek et al., Science, 2012, 337(6096):816-821), but numerous other sequences, including longer sequences can likewise be used.
[0157] The linker sequence can comprise a functional moiety. For example, the linker sequence can comprise one or more features, including an aptamer, a ribozyme, a proteininteracting hairpin, a protein binding site, a CRISPR array, an intron, or an exon. The linker sequence can comprise at least about 1, 2, 3, 4, or 5 or more functional moieties. In some examples, the linker sequence can comprise at most about 1, 2, 3, 4, or 5 or more functional moieties.
[0158] In some embodiments, an sgRNA does not comprise a uracil, e.g., at the 3 ’end of the sgRNA sequence. In some embodiments, a sgRNA does comprise one or more uracils, e.g., at the 3’end of the sgRNA sequence. In some embodiments, a sgRNA comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 uracils (U) at the 3’ end of the sgRNA sequence.
[0159] An sgRNA may be chemically modified. In some embodiments, a chemically modified gRNA is a gRNA that comprises at least one nucleotide with a chemical modification, e.g., a 2'-O-methyl sugar modification. In some embodiments, a chemically modified gRNA comprises a modified nucleic acid backbone. In some embodiments, a chemically modified gRNA comprises a 2'-O-methyl-phosphorothioate residue. In some embodiments, chemical modifications enhance stability, reduce the likelihood or degree of innate immune response, and / or enhance other attributes, as described in the art. In some embodiments, a modified gRNA may comprise modified backbones, for example, phosphorothioates, phosphotriesters, morpholinos, methyl phosphonates, short chain alkyl or cycloalkyl intersugar linkages or short chain heteroatomic or heterocyclic intersugar linkages.
[0160] Morpholino-based compounds are described in Braasch and David Corey, Biochemistry, 2002, 41(14): 4503-4510; Genesis, 2001, Volume 30, Issue 3; Heasman, Dev. Biol., 2002, 243: 209-214; Nasevicius et al., Nat. Genet., 2000, 26:216-220; Lacerra et al., Proc. Natl. Acad. Sci., 2000, 97: 9591-9596.; and U.S. Pat. No. 5,034,506, issued Jul. 23, 1991. Cyclohexenyl nucleic acid oligonucleotide mimetics are described in Wang et al., J. Am. Chem. Soc., 2000, 122: 8595-8602.
[0161] In some embodiments, a modified gRNA may comprise one or more substituted sugar moieties, e.g., one of the following at the 2' position: OH, SH, SCH3, F, OCN, OCH3, OCH3, O(CH2)nCH3, O(CH2)nNH2, or O(CH2)nCH3, where n is from 1 to about 10; Cl to CIO lower alkyl, alkoxyalkoxy, substituted lower alkyl, alkaryl or aralkyl; Cl; Br; CN; CF3; OCF3;0-, S-, or N-alkyl; O-, S-, or N-alkenyl; SOCH3; SO2CH3; ONO2; NO2; N3; NH2; heterocycloalkyl; heterocycloalkaryl; aminoalkylamino; polyalkylamino; substituted silyl; an RNA cleaving group; a reporter group; an intercalator; 2'-O-(2-methoxyethyl); 2'-methoxy (2'-O- CH3); 2'-propoxy(2'-OCH2 CH2CH3); and 2'-fluoro (2'-F). Similar modifications may also be made at other positions on the gRNA, particularly the 3' position of the sugar on the 3' terminal nucleotide and the 5' position of 5' terminal nucleotide. In some examples, both a sugar and an internucleoside linkage, z.e., the backbone, of the nucleotide units can be replaced with novel groups.
[0162] Guide RNAs can also include, additionally or alternatively, nucleobase (often referred to in the art simply as "base") modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases include adenine (A), guanine (G), thymine (T), cytosine (C), and uracil (U). Modified nucleobases include nucleobases found only infrequently or transiently in natural nucleic acids, e.g., hypoxanthine, 6-methyladenine, 5-Me pyrimidines, particularly 5-methylcytosine (also referred to as 5-methyl-2' deoxycytosine and often referred to in the art as 5-Me-C), 5-hydroxymethylcytosine (HMC), glycosyl HMC and gentobiosyl HMC, as well as synthetic nucleobases, e.g., 2-aminoadenine, 2-(methylamino)adenine, 2- (imidazolylalkyl)adenine, 2-(aminoalklyamino)adenine or other heterosubstituted alkyladenines, 2-thiouracil, 2-thiothymine, 5-bromouracil, 5-hydroxymethyluracil, 8-azaguanine, 7- deazaguanine, N6 (6-aminohexyl)adenine, and 2,6-diaminopurine. Kornberg, A., DNA Replication, W. H. Freeman & Co., San Francisco, pp. 75-77, 1980; Gebeyehu et al., Nucl. Acids Res. 1997, 15:4513. A "universal" base known in the art, e.g., inosine, can also be included. 5- Me-C substitutions have been shown to increase nucleic acid duplex stability by 0.6-1.2 °C. (Sanghvi, Y. S., in Crooke, S. T. and Lebleu, B., eds., Antisense Research and Applications, CRC Press, Boca Raton, 1993, pp. 276-278) and are aspects of base substitutions.
[0163] Modified nucleobases can comprise other synthetic and natural nucleobases, such as 5-methylcytosine (5-me-C), 5 -hydroxymethyl cytosine, xanthine, hypoxanthine, 2- aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5- halouracil and cytosine, 5-propynyl uracil and cytosine, 6-azo uracil, cytosine and thymine, 5- uracil (pseudo-uracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8- thioalkyl, 8-hydroxyl and other 8- substituted adenines and guanines, 5-halo particularly 5- bromo, 5 -trifluoromethyl and other 5- substituted uracils and cytosines, 7-methylquanine and 7-methyladenine, 8-azaguanine and 8- azaadenine, 7-deazaguanine and 7-deazaadenine, and 3 -deazaguanine and 3 -deazaadenine.Complexes of a Genome-targeting Nucleic Acid and an Endonuclease
[0164] A gRNA interacts with an endonuclease (e.g., an RNA-guided nuclease suchas Cas9), thereby forming a complex. The gRNA guides the endonuclease to a target polynucleotide.
[0165] The endonuclease and gRNA can each be administered separately to a cell or a subject. In some embodiments, the endonuclease can be pre-complexed with one or more guide RNAs, or one or more crRNA together with a tracrRNA. The pre-complexed material can then be administered to a cell or a subject. Such pre-complexed material is known as a ribonucleoprotein particle (RNP). The endonuclease in the RNP can be, for example, a Cas9 endonuclease or a Cpfl endonuclease. The endonuclease can be flanked at the N-terminus, the C-terminus, or both the N-terminus and C-terminus by one or more nuclear localization signals (NLSs). For example, a Cas9 endonuclease can be flanked by two NLSs, one NLS located at the N-terminus and the second NLS located at the C-terminus. The NLS can be any NLS known in the art, such as a SV40 NLS. The molar ratio of genome-targeting nucleic acid to endonuclease in the RNP can range from about 1 : 1 to about 10: 1. For example, the molar ratio of sgRNA to Cas9 endonuclease in the RNP can be 3 : 1. The molar ratio of sgRNA to Cas9 endonuclease in the RNP can be 5: 1.Nucleic Acids Encoding System Components
[0166] The present disclosure provides a nucleic acid comprising a nucleotide sequence encoding a genome-targeting nucleic acid of the disclosure, an endonuclease of the disclosure, and / or any nucleic acid or proteinaceous molecule necessary to carry out the aspects of the methods of the disclosure. The encoding nucleic acids can be RNA, DNA, or a combination thereof.
[0167] The nucleic acid encoding a genome-targeting nucleic acid of the disclosure, an endonuclease of the disclosure, and / or any nucleic acid or proteinaceous molecule necessary to carry out the aspects of the methods of the disclosure can comprise a vector (e.g. , a recombinant expression vector).
[0168] The term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. One type of vector is a "plasmid", which refers to a circular double-stranded DNA loop into which additional nucleic acid segments can be ligated. Another type of vector is a viral vector, wherein additional nucleic acid segments can be ligated into the viral genome. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome.
[0169] In some examples, vectors can be capable of directing the expression of nucleicacids to which they are operatively linked. Such vectors are referred to herein as "recombinant expression vectors", or more simply "expression vectors", which serve equivalent functions.
[0170] The term "operably linked" means that the nucleotide sequence of interest is linked to regulatory sequence(s) in a manner that allows for expression of the nucleotide sequence. The term "regulatory sequence" is intended to include, for example, promoters, enhancers and other expression control elements (e.g., polyadenylation signals). Such regulatory sequences are well known in the art and are described, for example, in Goeddel; Gene Expression Technology: Methods in Enzymology, 1990, 185, Academic Press, San Diego, CA. Regulatory sequences include those that direct constitutive expression of a nucleotide sequence in many types of host cells, and those that direct expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). It will be appreciated by those skilled in the art that the design of the expression vector can depend on such factors as the choice of the target cell, the level of expression desired, and the like.
[0171] Expression vectors contemplated include, but are not limited to, viral vectors based on vaccinia virus, poliovirus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, human immunodeficiency virus, retrovirus (e.g., Murine Leukemia Virus, spleen necrosis virus, and vectors derived from retroviruses such as Rous Sarcoma Virus, Harvey Sarcoma Virus, avian leukosis virus, a lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus) and other recombinant vectors. Other vectors contemplated for eukaryotic target cells include, but are not limited to, the vectors pXTl, pSG5, pSVK3, pBPV, pMSG, and pSVLSV40 (Pharmacia). Other vectors can be used so long as they are compatible with the host cell.
[0172] In some examples, a vector can comprise one or more transcription and / or translation control elements. Depending on the host / vector system utilized, any of a number of suitable transcription and translation control elements, including constitutive and inducible promoters, transcription enhancer elements, transcription terminators, etc. can be used in the expression vector. The vector can be a self-inactivating vector that either inactivates the viral sequences or the components of the CRISPR machinery or other elements.
[0173] Non-limiting examples of suitable eukaryotic promoters (i.e., promoters functional in a eukaryotic cell) include those from cytomegalovirus (CMV) immediate early, herpes simplex virus (HSV) thymidine kinase, early and late SV40, long terminal repeats (LTRs) from retrovirus, human elongation factor- 1 a promoter (EFla), chicken beta-actin promoter (CBA), ubiquitin C promoter (UBC), a hybrid construct comprising the cytomegalovirus enhancer fused to the chicken beta-actin promoter, a hybrid construct comprising the cytomegalovirus enhancer fused to the promoter, the first exon, and the first intron of chicken beta-actin gene (CAGor CAGGS), murine stem cell virus promoter (MSCV), phosphoglycerate kinase- 1 locus promoter (PGK), and mouse metallothionein-I promoter.
[0174] A promoter can be an inducible promoter (e.g., a heat shock promoter, tetracycline-regulated promoter, steroid-regulated promoter, metal-regulated promoter, estrogen receptor-regulated promoter, etc.). The promoter can be a constitutive promoter (e.g., CMV promoter, UBC promoter, CAG or CAGGS promoter). In some cases, the promoter can be a spatially restricted and / or temporally restricted promoter (e.g., a tissue specific promoter, a cell type specific promoter, etc.).
[0175] Introduction of the complexes, polypeptides, and nucleic acids of the disclosure into cells can occur by viral or bacteriophage infection, transfection, conjugation, protoplast fusion, lipofection, electroporation, nucleofection, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran mediated transfection, liposome- mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery, and the like.
[0176] Provided herein are methods of insertion of heterologous nucleotide sequence encoding a GLP-1R agonist using, e.g., an RNA-guided endonuclease system. The RNA-guided nuclease can be a CRISPR-associated nuclease. The CRISPR-associated nuclease can comprise an N-terminal nuclear localization signal (NLS), a C-terminal NLS, or both. The CRISPR- associated nuclease can comprise Cas9, Cpfl, or a variant thereof. The Cas9 can be S. pyogenes Cas9. The gRNA and the Cas9 can be delivered to the cell at a molar ratio of 5: 1. The RNA- guided endonuclease and the gRNA can be complexed as a ribonucleoprotein (RNP) complex prior to the delivering. In some embodiments, the gRNA comprises a spacer sequence targeting a region or gene of interest (e.g. , the target gene or region). In some embodiments, the gRNA targets a sequence within a safe harbor gene locus. In some embodiments, the gRNA targets a sequence within the INS gene locus. The gRNA can comprise a spacer sequence comprising the sequence of SEQ ID NO: 1 or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 1.Cell Types
[0177] Cells as described herein, e.g., genetically modified cells (and corresponding unmodified cells) may belong to any possible class of cell type. In some embodiments, a cell may be a mammalian cell. In some embodiments, a cell may be a human cell. In some embodiments, a cell may be a stem cell. In some embodiments, a cell may be a pluripotent stem cell (PSC). In some embodiments, a cell may be an embryonic stem cell (ESC), an adult stem cell (ASC), an induced pluripotent stem cell (iPSC), or a hematopoietic stem or progenitor cell (HSPC) (also called a hematopoietic stem cell (HSC). The cell can be a stem cell. In some embodiments, thecell can be an embryonic stem cell, an induced pluripotent stem cell, or an adult stem cell.
[0178] The stem cell can be capable of being differentiated into a pancreatic endocrine cell, an intestinal enteroendocrine cell, a neuron, a definitive endoderm cell, a primitive gut tube cell, a posterior foregut cell, a pancreatic endoderm cell, a pancreatic endocrine precursor cells, an immature beta cell, and / or a pancreatic beta cell. In some embodiments, the cell is a stage 6 (S6) pancreatic progenitor cell.
[0179] In some embodiments, the S6 cell expresses one or markers each selected from Forkhead Box A2 (FOXA2), Chromogranin A (CHGA), Pancreatic and duodenal homeobox 1 (PDX1), NK6 Homeobox 1 (NKX6.1), Neurogenin 3 (NGN3), Insulin (INS), and ISL LIM homeobox 1 (ISL1). In some embodiments, the cell is a differentiated cell selected from the group consisting of a pancreatic endocrine cell, an intestinal enteroendocrine cell, and a neuron.
[0180] The genetically modified cells described herein may be differentiated into relevant cell types to assess HLA expression, as well as the evaluation of immunogenicity of the stem cell lines. In general, differentiation comprises maintaining the cells of interest for a period time and under conditions sufficient for the cells to differentiate into the differentiated cells of interest. For example, the genetically modified stem cells disclosed herein may be differentiated into mesenchymal progenitor cells (MPCs), hypoimmunogenic cardiomyocytes, muscle progenitor cells, blast cells, endothelial cells (ECs), macrophages, hepatocytes, beta cells (e.g., pancreatic beta cells), pancreatic endoderm progenitors, pancreatic endocrine progenitors, pancreatic endocrine cells, hematopoietic progenitor cells, or neural progenitor cells (NPCs). In some embodiments, the genetically modified cell may be differentiated into definitive endoderm cells, primitive gut tube cells, posterior foregut cells, pancreatic endoderm cells (PEC), pancreatic endocrine cells, immature beta cells, or maturing beta cells.
[0181] Stem cells are capable of both proliferation and giving rise to more progenitor cells, these in turn having the ability to generate a large number of mother cells that can in turn give rise to differentiated or differentiable daughter cells. The daughter cells themselves can be induced to proliferate and produce progeny that subsequently differentiate into one or more mature cell types, while also retaining one or more cells with parental developmental potential. The term "stem cell" refers then, to a cell with the capacity or potential, under particular circumstances, to differentiate to a more specialized or differentiated phenotype, and which retains the capacity, under certain circumstances, to proliferate without substantially differentiating. In one aspect, the term progenitor or stem cell refers to a generalized mother cell whose descendants (progeny) specialize, often in different directions, by differentiation, e.g., by acquiring completely individual characters, as occurs in progressive diversification of embryonic cells and tissues. Cellular differentiation is a complex process typically occurring through many cell divisions. Adifferentiated cell may derive from a multipotent cell that itself is derived from a multipotent cell, and so on. While each of these multipotent cells may be considered stem cells, the range of cell types that each can give rise to may vary considerably. Some differentiated cells also have the capacity to give rise to cells of greater developmental potential. Such capacity may be natural or may be induced artificially upon treatment with various factors.
[0182] For instance, the human embryonic stem cells (hESCs) can be differentiated artificially into insulin producing cells via a seven-stage process requiring the addition of specific growth factors and small molecules. These seven stages include 1) definitive endoderm, 2) primitive gut tube, 3) posterior foregut, 4) pancreatic endoderm, 5) pancreatic endoderm precursors, 6) immature beta cells, and 7) maturing beta cells. For example, human pluripotent stem cells can be differentiated into pancreatic lineages as described in Schulz et al. (2012) PLoS ONE 7(5): e37004, Rezania et al. (2014) Nat. Biotechnol. 32(11): 1121-1133, and / or US20200208116. In many biological instances, stem cells can also be "multipotent" because they can produce progeny of more than one distinct cell type, but this is not required for "stem-ness."
[0183] A "differentiated cell" is a cell that has progressed further down the developmental pathway than the cell to which it is being compared. Thus, stem cells can differentiate into lineage-restricted precursor cells (such as a myocyte progenitor cell), which in turn can differentiate into other types of precursor cells further down the pathway (such as a myocyte precursor), and then to an end-stage differentiated cell, such as a myocyte, which plays a characteristic role in a certain tissue type, and may or may not retain the capacity to proliferate further. In some embodiments, the differentiated cell may be a pancreatic beta cell.Embryonic Stem Cells
[0184] The cells described herein may be embryonic stem cells (ESCs). ESCs are derived from blastocytes of mammalian embryos and are able differentiate into any cell type and propagate rapidly. ESCs are also believed to have a normal karyotype, maintaining high telomerase activity, and exhibiting remarkable long-term proliferative potential.Adult Stem Cells
[0185] The cells described herein may be adult stem cells (ASCs). ASCs are undifferentiated cells that may be found in mammals, e.g., humans. ASCs are defined by their ability to self-renew, e.g., be passaged through several rounds of cell replication while maintaining their undifferentiated state, and ability to differentiate into several distinct cell types, e.g., glial cells. Adult stem cells are a broad class of stem cells that may encompass hematopoietic stem cells, mammary stem cells, intestinal stem cells, mesenchymal stem cells, endothelial stem cells, neural stem cells, olfactory adult stem cells, neural crest stem cells, and testicular cells.Induced Pluripotent Stem Cells
[0186] The cells described herein may be induced pluripotent stem cells (iPSCs). An iPSC may be generated directly from an adult human cell by introducing genes that encode critical transcription factors involved in pluripotency, e.g., OCT4, SOX2, cMYC, and KLF4. An iPSC may be derived from the same subject to which subsequent progenitor cells are to be administered. That is, a somatic cell can be obtained from a subject, reprogrammed to an induced pluripotent stem cell, and then re-differentiated into a progenitor cell to be administered to the subject (e.g., autologous cells). However, in the case of autologous cells, a risk of immune response and poor viability post-engraftment remain. In some embodiments, iPSCs are gene-edited before differentiation into lineage-restricted progenitor cells or fully differentiated somatic cells.Human Hematopoietic Stem and Progenitor Cells
[0187] The cells described herein may be human hematopoietic stem and progenitor cells (hHSPCs). This stem cell lineage gives rise to all blood cell types, including erythroid (erythrocytes or red blood cells (RBCs)), myeloid (monocytes and macrophages, neutrophils, basophils, eosinophils, megakaryocytes / platelets, and dendritic cells), and lymphoid (T-cells, B- cells, NK-cells). Blood cells are produced by the proliferation and differentiation of a very small population of pluripotent hematopoietic stem cells (HSCs) that also have the ability to replenish themselves by self-renewal. During differentiation, the progeny of HSCs progress through various intermediate maturational stages, generating multi-potential and lineage-committed progenitor cells prior to reaching maturity. Bone marrow (BM) is the major site of hematopoiesis in humans and, under normal conditions, only small numbers of hematopoietic stem and progenitor cells (HSPCs) can be found in the peripheral blood (PB). Treatment with cytokines, some myelosuppressive drugs used in cancer treatment, and compounds that disrupt the interaction between hematopoietic and BM stromal cells can rapidly mobilize large numbers of stem and progenitors into the circulation.Di fferentiation o f cells into other cell types
[0188] As disclosed herein, human iPSCs can be differentiated into definitive endoderm using various treatments, including activin and B27 supplement (Life Technologies). The definitive endoderm is further differentiated into hepatocyte, the treatment includes: FGF4, HGF, BMP2, BMP4, Oncostatin M, Dexamethasone, etc. (Duan et al, Stem Cells, 2010;28:674- 686; Ma et al, Stem Cells Translational Medicine, 2013;2:409-419). In some embodiments, the differentiating step may be performed according to Sawitza et al, Sci Rep. 2015; 5: 13320. A differentiated cell may be any somatic cell of a mammal, e.g., a human. In some embodiments, a somatic cell may be an exocrine secretory epithelial cells (e.g., salivary gland mucous cell, prostate gland cell), a hormone-secreting cell (e.g., anterior pituitary cell, gut tract cell, pancreatic islet), a keratinizing epithelial cell (e.g., epidermal keratinocyte), a wet stratified barrier epithelialcell, a sensory transducer cell (e.g., a photoreceptor), an autonomic neuron cells, a sense organ and peripheral neuron supporting cell (e.g., Schwann cell), a central nervous system neuron, a glial cell (e.g., astrocyte, oligodendrocyte), a lens cell, an adipocyte, a kidney cell, a barrier function cell (e.g., a duct cell), an extracellular matrix cell, a contractile cell (e.g., skeletal muscle cell, heart muscle cell, smooth muscle cell), a blood cell (e.g., erythrocyte), an immune system cell (e.g., megakaryocyte, microglial cell, neutrophil, Mast cell, a T cell, a B cell, a Natural Killer cell), a germ cell (e.g., spermatid), a nurse cell, or an interstitial cell.
[0189] Populations of the genetically modified cells disclosed herein can maintain expression of the inserted nucleotide sequence. For example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% of the genetically modified cells express the inserted GLP-1R agonist. Moreover, populations of lineage-restricted or fully differentiated cells derived from the genetically modified cells disclosed herein maintain expression of the inserted one or more nucleotide sequences over time. For example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% of the lineage-restricted or fully differentiated cells express the GLP-1R agonist.Additional genetic modifications
[0190] The cell can comprise one or more additional genetic modifications. In some embodiments, the cells that are edited to comprise, e.g., heterologous nucleotide sequence encoding a GLP-1R agonist, also comprise one or more additional genetic modifications. In some embodiments, the one or more additional genetic modifications enable genetically modified cells, to increase their survival or viability and / or evade immune response following engraftment into a subject. In some embodiments, these strategies enable the cells to survive and / or evade immune response at higher success rates than an unmodified cell. Methods and cells comprising one or more additional edits are disclosed herein and are also described in PCT publications W02020049535A1, WO2021044377A1, WO2021044379A1, and WO2022144855A1, which are hereby incorporated by reference in their entireties.
[0191] In some embodiments, genetically modified cells comprise the introduction of at least one genetic modification within or near at least one gene that encodes a survival factor. In some embodiments, genetically modified cells comprise the introduction of at least one genetic modification within or near at least one gene that encodes a survival factor, wherein the genetic modification comprises an insertion of a polynucleotide encoding a tolerogenic factor. The genetically modified cells may further comprise at least one genetic modification within or near agene that encodes one or more MHC-I or MHC-II human leukocyte antigens or a component or a transcriptional regulator of a MHC-I or MHC-II complex. The genetically modified cells may further comprise at least one genetic modification within or near a gene that encodes one or more MHC-I or MHC-II human leukocyte antigens or a component or a transcriptional regulator of a MHC-I or MHC-II complex, wherein said genetic modification comprises an insertion of a polynucleotide encoding a second tolerogenic factor.
[0192] As used herein, the term “survival factor” generally refers to a factor that, when increased or decreased in a cell, enables the cell, e.g., a genetically modified cell, to survive after transplantation or engraftment into a host subject at higher survival rates relative to an unmodified cell. In some embodiments, a survival factor is a human survival factor. In some embodiments, a survival factor is a member of a critical pathway involved in cell survival. In some embodiments, a critical pathway involved in cell survival has implications on hypoxia, reactive oxygen species, nutrient deprivation, and / or oxidative stress. In some embodiments, the genetic modification, e.g., deletion or insertion, of at least one survival factor enables a genetically modified cell to survive for a longer time period, e.g., at least 1.05, at least 1.1, at least 1.25, at least 1.5, at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, or at least 50 times longer time period, than an unmodified cell following engraftment. In some embodiments, a survival factor is MANF (NCBI Gene ID No: 7873), ZNF143 (NCBI Gene ID No: 7702), TXNIP (NCBI Gene ID No: 10628), FOXO1 (NCBI Gene ID No: 2308), or JNK (NCBI Gene ID No: 5599). In some embodiments, a survival factor is inserted into a cell. In some embodiments, a survival factor is deleted from a cell. In some embodiments, an insertion of a polynucleotide that encodes MANF enables a cell to survive after transplantation or engraftment into a host subject at higher survival rates relative to an unmodified cell. In some embodiments, a deletion or insertion-deletion mutation within or near a TXNIP gene enables a cell to survive after transplantation or engraftment into a host subject at higher survival rates relative to an unmodified cell.
[0193] As used herein, the term “tolerogenic factor” generally refers to a protein (e.g., expressed by a polynucleotide as described herein) that, when increased or decreased in a cell, enables the cell to inhibit or evade immune rejection after transplantation or engraftment into a host subject at higher rates relative to an unmodified cell. In some embodiments, a tolerogenic factor is a human tolerogenic factor. In some embodiments, the genetic modification of at least one tolerogenic factor (e.g., the insertion or deletion of at least one tolerogenic factor) enables a cell to inhibit or evade immune rejection with rates at least 1.05, at least 1.1, at least 1.25, at least 1.5, at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, or at least 50 times higher than an unmodified cell following engraftment. In some embodiments, a tolerogenic factor is TNFAIP3 (NCBI Gene ID No: 7128), CD39 (NCBI Gene ID No: 953), CD73 (NCBI Gene IDNo. 4907), PD-L-1 (NCBI Gene ID No: 29126), HLA-E (NCBI Gene ID No: 3133), HLA-G (NCBI Gene ID No: 3135), CTLA-4 (NCBI Gene ID No: 1493), or CD47 (NCBI Gene ID No: 961). In some embodiments, a tolerogenic factor is inserted into a cell. In some embodiments, a tolerogenic factor is deleted from a cell. In some embodiments, an insertion of a polynucleotide that encodes TNFAIP3, CD39, CD73, HLA-E, PD-L-1, HLA-G, CTLA-4, and / or CD47 enables a cell to inhibit or evade immune rejection after transplantation or engraftment into a host subject.
[0194] As used herein, the term “transcriptional regulator of MHC-I or MHC-II” generally refers to a biomolecule that modulates, e.g., increases or decreases, the expression of a MHC-I and / or MHC-II human leukocyte antigen. In some embodiments, a biomolecule is a polynucleotide, e.g., a gene, or a protein. In some embodiments, a transcriptional regulator of MHC-I or MHC-II will increase or decrease the cell surface expression of at least one MHC-I or MHC-II protein. In some embodiments, a transcriptional regulator of MHC-I or MHC-II will increase or decrease the expression of at least one MHC-I or MHC-II gene.
[0195] In some embodiments, genetically modified cells comprise the introduction of at least one genetic modification within or near at least one gene that decreases the expression of one or more MHC-I and MHC-II human leukocyte antigens relative to an unmodified cell; at least one genetic modification that increases the expression of at least one polynucleotide that encodes a tolerogenic factor relative to an unmodified cell; and at least one genetic modification that alters the expression of at least one gene that encodes a survival factor relative to an unmodified cell. In some embodiments, genetically modified cells comprise at least one deletion or insertion-deletion mutation within or near at least one gene that alters the expression of one or more MHC-I and MHC-II human leukocyte antigens relative to an unmodified cell. In some embodiments, genetically modified cells comprise at least one deletion or insertion-deletion mutation within or near at least one gene that alters the expression of one or more MHC-I and MHC-II human leukocyte antigens relative to an unmodified cell; and at least one insertion of a polynucleotide that encodes at least one tolerogenic factor at a site that partially overlaps, completely overlaps, or is contained within, the site of a deletion of a gene that alters the expression of one or more MHC-I and MHC-II HLAs. In some embodiments, genetically modified cells comprise at least one genetic modification that alters the expression of at least one gene that encodes a survival factor relative to an unmodified cell.
[0196] The genes that encode the major histocompatibility complex (MHC) are located on human Chr. 6p21. The resultant proteins coded by the MHC genes are a series of surface proteins that are essential in donor compatibility during cellular transplantation. MHC genes are divided into MHC class I (MHC-I) and MHC class II (MHC-II). MHC-I genes (HLA- A, HLA-B, and HLA-C) are expressed in almost all tissue cell types, presenting “non-self’antigen-processed peptides to CD8+ T cells, thereby promoting their activation to cytolytic CD8+ T cells. Transplanted or engrafted cells expressing “non-self’ MHC-I molecules will cause a robust cellular immune response directed at these cells and ultimately resulting in their demise by activated cytolytic CD8+ T cells. MHC-I proteins are intimately associated with B2M in the endoplasmic reticulum, which is essential for forming functional MHC-I molecules on the cell surface. In addition, there are three non-classical MHC-Ib molecules (HLA-E, HLA-F, and HLA- G), which have immune regulatory functions. MHC-II biomolecule include HLA-DP, HLA-DM, HLA-DOA, HLA-DOB, HLA-DQ, and HLA-DR. Due to their primary function in the immune response, MHC-I and MHC-II biomolecules contribute to immune rejection following cellular engraftment of non-host cells, e.g., cellular engraftment for purposes of regenerative medicine.
[0197] MHC-I cell surface molecules are composed of MHC-encoded heavy chains (HLA-A, HLA-B, or HLA-C) and the invariant subunit B2M. Thus, a reduction in the concentration of B2M within a cell allows for an effective method of reducing the cell surface expression of MHC-I cell surface molecules. In some embodiments, the genome of a cell has been modified to disrupt or decrease the expression of beta-2-microglobulin (B2M), also known as P2 microglobulin, B2 microglubulin, or IMD43. B2M is a non-polymorphic gene that encodes a common protein subunit required for surface expression of all polymorphic MHC class I heavy chains. HLA-I proteins are intimately associated with B2M in the endoplasmic reticulum, which is essential for forming functional, cell-surface expressed HLA-I molecules. In some embodiments, the genetic modification is generated using a CRISPR / Cas9 system, using e.g., a gRNA targeting B2M gene. In some embodiments, the gRNA targets a site within the B2M gene comprising a 5’ GCTACTCTCTCTTTCTGGCC 3’ sequence (SEQ ID NO: 50) and / or comprises a spacer sequence of SEQ ID NO: 37. In some embodiments, the gRNA targets a site within the B2M gene comprising a 5’ GGCCGAGATGTCTCGCTCCG 3’ sequence (SEQ ID NO: 51) and / or comprises a spacer sequence of SEQ ID NO: 38. In some embodiments, the gRNA targets a site within the B2M gene comprising a 5’ CGCGAGCACAGCTAAGGCCA 3’ sequence (SEQ ID NO: 52) and / or comprises a spacer sequence of SEQ ID NO: 39. In some embodiments, the gRNA targets a site within the B2M gene comprising any of the following sequences: 5’- TATAAGTGGAGGCGTCGCGC-3’ (SEQ ID NO: 53), 5’-GAGTAGCGCGAGCACAGCTA-3’ (SEQ ID NO: 54), 5’-ACTGGACGCGTCGCGCTGGC-3’ (SEQ ID NO: 55), 5’-AAGTGGAGGCGTCGCGCTGG-3’ (SEQ ID NO: 56), 5-GGCCACGGAGCGAGACATCT -3’ (SEQ ID NO: 57), 5’-GCCCGAATGCTGTCAGCTTC-3’ (SEQ ID NO: 58), 5’-CTCGCGCTACTCTCTCTTTC-3’ (SEQ ID NO: 59), 5’-TCCTGAAGCTGACAGCATTC-3’ (SEQ ID NO: 60), 5’-TTCCTGAAGCTGACAGCATT-3’ (SEQ ID NO: 61), or 5’- ACTCTCTCTTTCTGGCCTGG-3’ (SEQ ID NO: 62). In some embodiments, the gRNAcomprises a spacer comprising any of the following RNA sequences: 5’- TATAAGTGGAGGCGTCGCGC-3’ (SEQ ID NO: 40), 5’-GAGTAGCGCGAGCACAGCTA-3’ (SEQ ID NO: 41), 5’-ACTGGACGCGTCGCGCTGGC-3’ (SEQ ID NO: 42), 5’-AAGTGGAGGCGTCGCGCTGG-3’ (SEQ ID NO: 43), 5-GGCCACGGAGCGAGACATCT -3’ (SEQ ID NO: 44), 5’-GCCCGAATGCTGTCAGCTTC-3’ (SEQ ID NO: 45), 5’-CTCGCGCTACTCTCTCTTTC-3’ (SEQ ID NO: 46), 5’-TCCTGAAGCTGACAGCATTC-3’ (SEQ ID NO: 47), 5’-TTCCTGAAGCTGACAGCATT-3’ (SEQ ID NO: 48), or 5’- ACTCTCTCTTTCTGGCCTGG-3’ (SEQ ID NO: 49). In some embodiments, the gRNA comprises an RNA version of the polynucleotide sequence of SEQ ID NO: 51 (e.g., SEQ ID NO: 38). In some embodiments, the gRNA comprises an RNA version of any of SEQ ID NO: 50 or 52-62 (e.g., SEQ ID NO: 37 or any of 39-49). The gRNA / CRISPR nuclease complex targets and cleaves a target site in the B2M locus. Repair of a double-stranded break by NHEJ can result in a deletion of at least on nucleotide and / or an insertion of at least one nucleotide, thereby disrupting or eliminating expression of B2M. Alternatively, the B2M locus can be targeted by at least two CRISPR systems each comprising a different gRNA, such that cleavage at two sites in the B2M locus leads to a deletion of the sequence between the two cuts, thereby eliminating expression of B2M.
[0198] In some embodiments, genetically modified cells comprise at least one genetic modification that disrupts the expression of at least one gene that encodes a survival factor, such as TXNIP, relative to an unmodified cell. In some embodiments, the genome of a cell has been modified to decrease the expression of thioredoxin interacting protein (TXNIP), which is also known as EST01027, HHCPA78, THIF, VDUP1, or ARRDC6. TXNIP is metabolic gene involved in redox regulation that can also function as a tumor suppressor. Downregulation or knockout of TXNIP can protect cells from metabolic stress. In some embodiments, the genetic modification is generated using a CRISPR / Cas9 system, using e.g., a gRNA targeting TXNIP gene. In some embodiments, the gRNA targets a site within the TXNIP gene comprising a 5’- GAAGCGTGTCTTCATAGCGC-3’ sequence (SEQ ID NO: 73) and / or comprises a spacer sequence of SEQ ID NO: 63. In some embodiments, the gRNA targets a site within the TXNIP gene comprising a 5’-TTACTCGTGTCAAAGCCGTT-3’ sequence (SEQ ID NO: 74) and / or comprises a spacer sequence of SEQ ID NO: 64. In some embodiments, the gRNA targets a site within the TXNIP gene comprising a 5’-TGTCAAAGCCGTTAGGATCC-3’ sequence (SEQ ID NO: 75) and / or comprises a spacer sequence of SEQ ID NO: 65. In some embodiments, the gRNA targets a site within the TXNIP gene comprising a 5’-GCCGTTAGGATCCTGGCTTG-3’ sequence (SEQ ID NO: 76) and / or comprises a spacer sequence of SEQ ID NO: 66. In some embodiments, the gRNA targets a site within the TXNIP gene comprising a 5’-GCGGAGTGGCTAAAGTGCTT-3’ sequence (SEQ ID NO: 77) and / or comprises a spacer sequence of SEQ ID NO: 67. In some embodiments, the gRNA targets a site within the TXNIP gene comprising a 5’-TCCGCAAGCCAGGATCCTAA-3’ sequence (SEQ ID NO: 78) and / or comprises a spacer sequence of SEQ ID NO: 68. In some embodiments, the gRNA targets a site within the TXNIP gene comprising a 5’-GTTCGGCTTTGAGCTTCCTC-3’ sequence (SEQ ID NO: 79) and / or comprises a spacer sequence of SEQ ID NO: 69. In some embodiments, the gRNA targets site within the TXNIP gene comprising a 5’-GAGATGGTGATCATGAGACC-3’ sequence (SEQ ID NO: 80) and / or comprises a spacer sequence of SEQ ID NO: 70. In some embodiments, the gRNA targets a site within the TXNIP gene comprising a 5’- TTGTACTCATATTTGTTTCC-3’ sequence (SEQ ID NO: 81) and / or comprises a spacer sequence of SEQ ID NO: 71. In some embodiments, the gRNA targets a site within the TXNIP gene comprising a 5’-AACAAATATGAGTACAAGTT-3’ sequence (SEQ ID NO: 82) and / or comprises a spacer sequence of SEQ ID NO: 72. In some embodiments, the gRNA comprises an RNA version of the polynucleotide sequence of SEQ ID NO: 78 (e.g., SEQ ID NO: 68). In some embodiments, the gRNA comprises an RNA version of any one of SEQ ID NO: 73-77 or 79-82 (e.g., SEQ ID NOs: 63-67 or 69-72). The gRNA / CRISPR nuclease complex targets and cleaves a target site in the TXNIP gene locus. Repair of a double-stranded break by NHEJ can result in a deletion of at least one nucleotide and / or an insertion of at least one nucleotide, thereby disrupting or eliminating expression of TXNIP. Alternatively, insertion of a polynucleotide encoding an exogenous gene into the TXNIP gene locus can disrupt or eliminate expression of TXNIP.
[0199] In some embodiments, the genetic modification comprises insertion of one or more nucleotide sequences each encoding a tolerogenic factor. In some embodiments, the tolerogenic factor is PD-L-1 (programmed death ligand 1) also known as cluster of differentiation 274 (CD274), B7 homolog (B7-H, B7H1), PDCD1L1, PDCD1LG1, PD-L1, or PDL1. PD-L-1 appears to play a major role in suppressing the adaptive arm of immune system and is considered to be a co-inhibitory factor of the immune response. In some embodiments, the tolerogenic factor is HLA-E, also known as EA1.2, EA2.1, HLA-6.2, MHC, QA1, or major histocompatibility complex, class I, E. HLA-E is an important modulator of natural killer (NK) and cytotoxic T lymphocyte (CTL) activation and inhibitory function. In some embodiments, the tolerogenic factor is Tumor Necrosis Factor Alpha-Induced Protein 3 (TNFAIP3, also referred to herein as A20). TNFAIP3 was identified as a gene whose expression is rapidly induced by the tumor necrosis factor (TNF). The protein encoded by this gene is a zinc finger protein and ubiquitin- editing enzyme, and has been shown to inhibit NF-kappa B activation as well as TNF-mediated apoptosis. The encoded protein, which has both ubiquitin ligase and de-ubiquitinase activities, is involved in the cytokine-mediated immune and inflammatory responses. Several transcriptvariants encoding the same protein have been found for this gene.
[0200] In some embodiments, a polynucleotide encoding one or more survival factors, such as MANF, can be inserted into genetically modified or genetically unmodified cells to create genetically modified cells having increased survival. In some embodiments, the survival factor is MANF, which is also known as arginine-rich, mutated in early-stage tumors (ARMET), arginine- rich protein (ARP), or mesencephalic astrocyte derived neurotrophic factor. MANF is an endoplasmic reticulum (ER) stress-inducible neurotrophic factor that promotes proliferation and survival of pancreatic beta cells, as well as survival of dopaminergic neurons. In some embodiments, insertion of a polynucleotide encoding one or more survival factors, such as MANF, enables a genetically modified cell to survive after transplantation or engraftment into a host subject with rates at least 1.05, at least 1.1, at least 1.25, at least 1.5, at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, or at least 50 times higher than an unmodified cell following transplantation or engraftment.
[0201] In some embodiments, the genetic modification, e.g., insertion of at least one polynucleotide encoding at least one tolerogenic factor enables a genetically modified cell to inhibit or evade immune rejection with rates at least 1.05, at least 1.1, at least 1.25, at least 1.5, at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, or at least 50 times higher than an unmodified cell following engraftment. In some embodiments, a cell can comprise one or more of the genetic modifications described above.
[0202] In some embodiments, one or more nucleotide sequences each encoding a protein selected from PD-L1, MANF, TNFAIP3, and HLA-E (e.g., HLA-E trimer) is inserted into a gene locus. In some embodiments, the gene locus is B2M or TXNIP. In some embodiments, a polynucleotide encoding PD-L-1 is inserted at a site within or near a B2M gene locus or within or near a TXNIP gene locus. In some embodiments, a polynucleotide encoding PD-L-1 is inserted at a site within or near a B2M gene locus. In some embodiments, a polynucleotide encoding PD-L- 1 is inserted at a site within or near a B2M gene locus concurrent with or following a deletion of all or part of a B2M gene or promoter. In some embodiments, a polynucleotide encoding PD-L-1 is inserted at a site within or near a TXNIP gene locus concurrent with or following a deletion of all or part of a TXNIP gene or promoter. The polynucleotide encoding PD-L-1 can be operably linked to an exogenous promoter. The exogenous promoter can be a CAG or CAGGS promoter. In some embodiments, the polynucleotide encoding PD-L-1 comprises a nucleotide sequence of SEQ ID NO: 83, or nucleotide sequence having at least 85%, 90%, 95%, or 99% sequence identity with that of SEQ ID NO: 83.
[0203] In some embodiments, a polynucleotide encoding HLA-E is inserted at a site within or near a B2M gene locus or within or near a TXNIP gene locus. In some embodiments, apolynucleotide encoding HLA-E is inserted at a site within or near a TXNIP gene locus. In some embodiments, a polynucleotide encoding HLA-E is inserted at a site within or near a B2M gene locus concurrent with or following a deletion of all or part of a B2M gene or promoter. In some embodiments, a polynucleotide encoding HLA-E is inserted at a site within or near a TXNIP gene locus concurrent with or following a deletion of all or part of a TXNIP gene or promoter. In some embodiments, the polynucleotide encoding HLA-E comprises a sequence encoding an HLA-E trimer, the HLA-E trimer comprising a B2M signal peptide fused to an HLA-G presentation peptide fused to a B2M membrane protein fused to HLA-E without its signal peptide. In some embodiments, the polynucleotide encoding HLA-E is operably linked to an exogenous promoter. The exogenous promoter can be a CMV promoter. In some embodiments, the polynucleotide encoding HLA-E comprises a nucleotide sequence of SEQ ID NO: 84, or nucleotide sequence having at least 85%, 90%, 95%, or 99% sequence identity with that of SEQ ID NO: 84.
[0204] In some embodiments, a polynucleotide encoding TNFAIP3 is inserted at a site within or near a B2M gene locus or within or near a TXNIP gene locus. In some embodiments, a polynucleotide encoding TNFAIP3 is inserted at a site within or near a B2M gene locus. In some embodiments, a polynucleotide encoding TNFAIP3 is inserted at a site within or near a B2M gene locus concurrent with or following a deletion of all or part of a B2M gene or promoter. In some embodiments, a polynucleotide encoding TNFAIP3 is inserted at a site within or near a TXNIP gene locus concurrent with or following a deletion of all or part of a TXNIP gene or promoter. In some embodiments, the polynucleotide encoding TNFAIP3 is operably linked to an exogenous promoter. The exogenous promoter can be a CAG or CAGGS promoter. In some embodiments, the polynucleotide encoding TNFAIP3 comprises a nucleotide sequence of SEQ ID NO: 85, or nucleotide sequence having at least 85%, 90%, 95%, or 99% sequence identity with that of SEQ ID NO: 85.
[0205] In some embodiments, a polynucleotide encoding MANF is inserted at a site within or near a B2M gene locus or within or near a TXNIP gene locus. In some embodiments, a polynucleotide encoding MANF is inserted at a site within or near a B2M gene locus. In some embodiments, a polynucleotide encoding MANF is inserted at a site within or near a B2M gene locus concurrent with or following a deletion of all or part of a B2M gene or promoter. In some embodiments, a polynucleotide encoding MANF is inserted at a site within or near a TXNIP gene locus concurrent with or following a deletion of all or part of a TXNIP gene or promoter. The polynucleotide encoding MANF is operably linked to an exogenous promoter. The exogenous promoter can be a CAG or CAGGS promoter. In some embodiments, the polynucleotide encoding MANF comprises a nucleotide sequence of SEQ ID NO: 86, or nucleotide sequence having at least 85%, 90%, 95%, or 99% sequence identity with that of SEQ ID NO: 86.
[0206] In some embodiments, a polynucleotide encoding TNFAIP3 and PD-L-1 is inserted at a site within or near a B2M gene locus or within or near a TXNIP gene locus. In some embodiments, a polynucleotide encoding TNFAIP3 and PD-L-1 is inserted at a site within or near a B2M gene locus. In some embodiments, a polynucleotide encoding TNFAIP3 and PD-L-1 is inserted at a site within or near a B2M gene locus concurrent with or following a deletion of all or part of a B2M gene or promoter. The polynucleotide encoding TNFAIP3 and PD-L-1 comprises sequence encoding TNFAIP3 that is linked to sequence encoding a ribosome skip that is linked to sequence encoding PD-L-1. The ribosome skip can be a 2 A sequence family member, such as P2A. In some embodiments, the polynucleotide comprises TNFAIP3-P2A-PD-L-1 coding sequence. In some embodiments, the polynucleotide encoding TNFAIP3-P2A-PD-L-1 comprises or consists of a nucleotide sequence of SEQ ID NO: 87 or a nucleotide sequence having at least 85%, 90%, 95%, or 99% sequence identity with that of SEQ ID NO: 87. In some embodiments, the polynucleotide encoding TNFAIP3-P2A-PD-L-1 is operably linked to an exogenous promoter. The exogenous promoter can be a CAG or CAGGS promoter. In some embodiments, a donor plasmid encoding TNFAIP3-P2A-PD-L-1 and comprising B2M homology arms has a nucleotide sequence of SEQ ID NO: 88.
[0207] In some embodiments, a polynucleotide encoding MANF and HLA-E is inserted at a site within or near a B2M gene locus or within or near a TXNIP gene locus. In some embodiments, a polynucleotide encoding MANF and HLA-E is inserted at a site within or near a TXNIP gene locus. In some embodiments, a polynucleotide encoding MANF and HLA-E is inserted at a site within or near a TXNIP gene locus concurrent with, or following a deletion of all or part of a TXNIP gene or promoter. In some embodiments, the polynucleotide encoding MANF and HLA-E comprises sequence encoding MANF that is linked to sequence encoding a ribosome skip that is linked to sequence encoding HLA-E. The ribosome skip can be a 2A sequence family member, such as P2A. In some embodiments, the sequence encoding HLA-E comprises sequence encoding a HLA-E trimer, the HLA-E trimer comprising a B2M signal peptide fused to an HLA- G presentation peptide fused to a B2M membrane protein fused to HLA-E without its signal peptide. In some embodiments the polynucleotide comprises MANF-P2A-HLA-E coding sequence. In some embodiments, the polynucleotide encoding MANF-P2A-HLA-E comprises or consists of a nucleotide sequence of SEQ ID NO: 89 or a nucleotide sequence having at least 85%, 90%, 95%, or 99% sequence identity with that of SEQ ID NO: 89. In some embodiments, the polynucleotide encoding MANF-P2A-HLA-E is operably linked to an exogenous promoter. The exogenous promoter can be a CAG or CAGGS promoter. In some embodiments, a donor plasmid MANF-P2A-HLA-E and comprising TXNIP homology arms has a nucleotide sequence of SEQ ID NO: 90.
[0208] The one or more additional genetic modifications can comprise: (1) a disrupted beta-2-microglobulin (B2M) gene and insertion of a nucleotide sequence encoding programmed death-ligand 1 (PD-L1) into the disrupted B2M gene, wherein the cell expresses PD-L1 and has reduced or eliminated expression of B2M; and / or (2) a disrupted thioredoxin interacting protein (TXNIP) gene and insertion of a nucleotide sequence encoding HLA class I histocompatibility antigen, alpha chain E (HLA-E); optionally the nucleotide sequence encoding HLA-E encodes an HLA-E trimer comprising a B2M signal peptide fused to an HLA-G presentation peptide fused to a B2M membrane protein fused to HLA-E without its signal peptide; wherein the cell expresses HLA-E or the HLA-E trimer and has reduced or eliminated expression of TXNIP.
[0209] The one or more additional genetic modifications can comprise: (3) a disrupted beta-2-microglobulin (B2M) gene and insertion of a nucleotide sequence encoding PD-L1 and a nucleotide sequence encoding tumor necrosis factor alpha induced protein 3 (TNFAIP3) into the disrupted B2M gene, wherein the cell expresses PD-L1 and TNFAIP3 and has reduced or eliminated expression of B2M; and / or (4) a disrupted thioredoxin interacting protein (TXNIP) gene and insertion of a nucleotide sequence encoding HLA class I histocompatibility antigen, alpha chain E (HLA-E) and a nucleotide sequence mesencephalic astrocyte derived neurotrophic factor (MANF) into the disrupted TXNIP gene; optionally the nucleotide sequence encoding HLA-E encodes an HLA-E trimer comprising a B2M signal peptide fused to an HLA-G presentation peptide fused to a B2M membrane protein fused to HLA-E without its signal peptide; wherein the cell expresses HLA-E or the HLA-E trimer and MANF and has disrupted expression of TXNIP.
[0210] The nucleotide sequence encoding TNFAIP3 and the nucleotide sequence encoding PD-L1 can be linked by a nucleotide sequence encoding a P2A peptide such that the insertion of (3) comprises a nucleotide sequence encoding TNFAIP3-P2A-PD-L1. The nucleotide sequence encoding the HLA-E trimer and the nucleotide sequence encoding MANF can be linked by a nucleotide sequence encoding a P2A peptide such that the insertion of (4) comprises a nucleotide sequence encoding MANF-P2A-HLA-E trimer.
[0211] The nucleotide sequence encoding PD-L1 can comprise a sequence that is at least 85% identical to the sequence of SEQ ID NO: 83. The nucleotide sequence encoding the HLA-E trimer can comprise a sequence that is at least 85% identical to the sequence of SEQ ID NO: 84. The nucleotide sequence encoding TNFAIP3 can comprise a sequence that is at least 85% identical to the sequence of SEQ ID NO: 85. The nucleotide sequence encoding MANF can comprise a sequence that is at least 85% identical to the sequence of SEQ ID NO: 86. The nucleotide sequence encoding TNFAIP3-P2A-PD-L1 can comprise a sequence that is at least 85% identical to the sequence of SEQ ID NO: 87. The nucleotide sequence encoding MANF-P2A-HLA-E trimer can comprise a sequence that is at least 85% identical to the sequence of SEQ ID NO: 89.
[0212] The introduction of the genetic edits described herein can be performed in any order. In some embodiments, the nucleotide sequence encoding the GLP-1R agonist is inserted into the genome of a cell already comprising, e.g., the one or more additional edits. In some embodiments, the one or more additional edits are made to the cell following insertion of the nucleotide sequence encoding GLP-1R agonist into the genome of the cell.Formulation and administration of cells
[0213] Cells, e.g., genetically modified cells as described herein, can be formulated and administered to a subject by any manner known in the art.
[0214] The terms “administering,” “introducing,” “implanting,” “engrafting,” and "transplanting" are used interchangeably in the context of the placement of cells, e.g., progenitor cells, into a subject, by a method or route that results in at least partial localization of the introduced cells at a desired site. The cells e.g., progenitor cells, or their differentiated progeny can be administered by any appropriate route that results in delivery to a desired location in the subject where at least a portion of the implanted cells or components of the cells remain viable. The period of viability of the cells after administration to a subject can be as short as a few hours, e.g., twenty-four hours, to a few days, to as long as several years, or even the lifetime of the subject, i.e., long-term engraftment.
[0215] A genetically modified cell as described herein may be viable after administration to a subject for a period that is longer than that of an unmodified cell.
[0216] In some embodiments, a composition comprising cells as described herein may be administered by a suitable route, which may include intravenous administration, e.g., as a bolus or by continuous infusion over a period of time. In some embodiments, intravenous administration may be performed by intramuscular, intraperitoneal, intracerebrospinal, subcutaneous, intraarticular, intrasynovial, or intrathecal routes. In some embodiments, a composition may be in solid form, aqueous form, or a liquid form. In some embodiments, an aqueous or liquid form may be nebulized or lyophilized. In some embodiments, a nebulized or lyophilized form may be reconstituted with an aqueous or liquid solution.
[0217] A cell composition can also be emulsified or presented as a liposome composition, provided that the emulsification procedure does not adversely affect cell viability. The cells and any other active ingredient can be mixed with excipients that are pharmaceutically acceptable and compatible with the active ingredient, and in amounts suitable for use in the therapeutic methods described herein. In some embodiments, the cell composition is administered by a surgical route.
[0218] Additional agents included in a cell composition can include pharmaceutically acceptable salts of the components therein. Pharmaceutically acceptable salts include the acid addition salts (formed with the free amino groups of the polypeptide) that are formed with inorganic acids, such as, for example, hydrochloric or phosphoric acids, or such organic acids as acetic, tartaric, mandelic and the like. Salts formed with the free carboxyl groups can also be derived from inorganic bases, such as, for example, sodium, potassium, ammonium, calcium or ferric hydroxides, and such organic bases as isopropylamine, trimethylamine, 2-ethylamino ethanol, histidine, procaine and the like.
[0219] Physiologically tolerable carriers are well known in the art. Exemplary liquid carriers are sterile aqueous solutions that contain no materials in addition to the active ingredients and water, or contain a buffer such as sodium phosphate at physiological pH value, physiological saline or both, such as phosphate-buffered saline. Still further, aqueous carriers can contain more than one buffer salt, as well as salts such as sodium and potassium chlorides, dextrose, polyethylene glycol and other solutes. Liquid compositions can also contain liquid phases in addition to and to the exclusion of water. Exemplary of such additional liquid phases are glycerin, vegetable oils such as cottonseed oil, and water-oil emulsions. The amount of an active compound used in the cell compositions that is effective in the treatment of a particular disorder or condition can depend on the nature of the disorder or condition and can be determined by standard clinical techniques.
[0220] In some embodiments, the composition is administered two or more times to the subject. The length of time between the two or more administrations can vary. In some embodiments, the administration comprises a single administration of the composition to the subject. The frequency of administration of the compositions of the disclosure can be determined by, e.g., a medical professional.
[0221] In some embodiments, a composition comprising cells may be administered to a subject, e.g., a human subject, who has, is suspected of having, or is at risk for a disease or disorder. In some embodiments, a composition may be administered to a subject who does not have, is not suspected of having or is not at risk for a disease or disorder. In some embodiments, a subject is a healthy human. In some embodiments, a subject e.g., a human subject, who has, is suspected of having, or is at risk for a genetically inheritable disease or disorder. In some embodiments, the subject is suffering or is at risk of developing symptoms indicative of a disease or disorder. In some embodiments, the disease is diabetes (also referred to herein as diabetes mellitus), e.g., type 1 diabetes or type 2 diabetes.
[0222] Disclosed herein include populations of cells comprising a plurality of any of the genetically modified cell of the disclosure or comprising a plurality of the genetically modifiedcell generating by the methods disclosed herein. Disclosed herein include compositions comprising a population of genetically modified cells described. Also disclosed herein are pharmaceutical compositions comprising the compositions of the disclosure and a pharmaceutically acceptable excipient. Provided herein are compositions and pharmaceutical compositions for use in treating a disease or disorder in a subject.
[0223] Disclosed herein include methods of treating a disease or disorder in a subject. In some embodiments, the method comprises administering to a subject in need thereof a therapeutically effective amount of any of the compositions disclosed herein, thereby treating the disease or disorder in the subject.
[0224] The disease or disorder can be diabetes mellitus, dyslipidemia, fatty liver disease, metabolic syndrome, non-alcoholic steatohepatitis or obesity. The diabetes mellitus can be type 1 or type 2 diabetes mellitus. The disease or disorder can be a weight-related comorbid condition selected from the group consisting of hypertension, dyslipidemia, prediabetes, type 2 diabetes mellitus, obstructive sleep apnea, cardiovascular disease, and any combination thereof.
[0225] Disclosed herein are methods for chronic weight management in a subject in need thereof. In some embodiments, the method comprises: administering to the subject a therapeutically effective amount of any of the compositions of the disclosure.
[0226] In some embodiments, the subject has a BMI > 27 kg / m2and at least one weight- related comorbidity. Each of the at least one weight- related comorbidity can be selected from hypertension, dyslipidemia, type 2 diabetes mellitus, coronary heart disease, stroke, gallbladder disease, osteoarthritis, sleep apnea, cancer, mental illness, and body pain.
[0227] In some embodiments, administration of any of the compositions disclosed herein can result in decrease in glycolated hemoglobin (HbAlC or A1C) in the subject. In some embodiments, a decrease of at least 1% in A1C is observed in the subject after at least or at least about 6 months (e.g., 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 1 year, 13 months, 14 months, 15 months, 16 months, 17 months, 18 months, 19 months, 20 months, 21 months, 22 months, 23 months, 2 years or more) following treatment relative to the A1C level prior to the administration. In some embodiments, administration of any of the compositions disclosed herein can result in decrease in weight. For example, the weight of an adult subject may decrease by at least 2 kg about one year (e.g., 1 year, 13 months, 14 months, 15 months, 16 months, 17 months, 18 months, 19 months, 20 months, 21 months, 22 months, 23 months, 2 years or more) following the administration, as compared to the weight of the subject prior to the administration. In some embodiments, the body weight of a subject is reduced, for example, by between about 0.01% to about 0.1%, between about 0.1% to about 0.5%, between about 0.5% to about 1%, between about 1% to about 5%, between about 2% to about 3%, between about 5% to about 10%,between about 10% to about 15%, between about 15% to about 20%, between about 20% to about 25%, between about 25% to about 30%, between about 30% to about 35%, between about 35% to about 40%, between about 40% to about 45%, or between about 45% to about 50%, relative to the body weight of a subject prior to administration.
[0228] In some embodiments, the reduction in body weight is maintained for about 1 week, for about 2 weeks, for about 3 weeks, for about 1 month, for about 2 months, for about 3 months, for about 4 months, for about 5 months, for about 6 months, for about 7 months, for about 8 months, for about 9 months, for about 10 months, for about 11 months, for about 1 year, for about 1.5 years, for about 2 years, for about 2.5 years, for about 3 years, for about 3.5 years, for about 4 years, for about 4.5 years, for about 5 years, for about 6 years, for about 7 years, for about 8 years, for about 9 years, for about 10 years, for about 15 years, or for about 20 years, for example. In some embodiments, administration of any of the compositions disclosed herein can result in maintenance of weight of subject relative to the body weight of a subject prior to the administration. For example, the weight of an adult subject may be reduced by no more than 2 kg or increased by no more than 2 kg, relative to the weight of the subject prior to the administration, for about 1 week, for about 2 weeks, for about 3 weeks, for about 1 month, for about 2 months, for about 3 months, for about 4 months, for about 5 months, for about 6 months, for about 7 months, for about 8 months, for about 9 months, for about 10 months, for about 11 months, for about 1 year, for about 1.5 years, for about 2 years, for about 2.5 years, for about 3 years, for about 3.5 years, for about 4 years, for about 4.5 years, for about 5 years, for about 6 years, for about 7 years, for about 8 years, for about 9 years, for about 10 years, for about 15 years, or for about 20 years after the administration.EXAMPLES
[0229] Some aspects of the embodiments discussed above are disclosed in further detail in the following examples, which are not in any way intended to limit the scope of the present disclosure.Example 1Generation of genetically modified cells
[0230] Described in this Example are methods for editing cells (e.g., stem cells) to express a heterologous GLP-1R agonist.
[0231] Cells are generated in which a polynucleotide encoding GLP1 is inserted into Insulin gene locus. GLP1 is knocked into cells using GLPl-CD19-trunc mCherry donor plasmid (SEQ ID NO: 11) and the INS-G13 gRNA (SEQ ID NO: 1). The cells (for example cells comprising gene edits e.g., B2M KO / TNFAIP3-P2A-PD-L1 KI; TXNIP KO / MANF-P2A-HLA- E trimer KI, also referred to herein as “Modified Parent Cells” or “Modified Parent HumanPluripotent Stem Cells”) are passaged the two days before electroporation and seeded as 5 million cells per T75 flask. On the day of electroporation, the cells are split again and electroporated using the Neon Electroporator with the RNP mixture of Cas9 protein (Biomay) and guide RNA (Agilent) at a molar ratio of 5:1 (gRNA:Cas9) with absolute values of 125 pmol Cas9 and 625 pmol gRNA per 2 million cells. To form the RNP complex, gRNA and Cas9 are combined in one vessel with R-buffer (Neon Transfection Kit) to a total volume of 25-50 pL and incubated for 15 minutes at room temperature (RT). This mixture is combined with the cells and 4 pg plasmid to a total volume of ~115 pL using R-buffer. This mixture is electroporated with 2 pulses for 25 ms at 1400 V. Following electroporation, the cells are pipetted out into a 6 well plate filled with STEMFLEX™ media with REVITACELL™ Supplement (100X) and BIOLAMININ 521 CTG at 1 : 10 dilution. Cells are cultured in a normoxia incubator (37°C, 8% CO2). To create single cell clone, cells are FACS sorted for CD19 into a 96-well plate at day 2 after electroporation. After 12-14 days post sorting, sorted single cell clones are genotyped by PCR.Generation of TIRZ-KI into Modified Parent Human Pluripotent Stem Cells
[0232] Cells are generated in which a polynucleotide encoding TIRZ is inserted into Insulin gene locus. TIRZ is knocked into cells using TIRZ-CD19-trunc mCherry donor plasmid (SEQ ID NO: 12) and the INS-G13 gRNA (SEQ ID NO: 1). The cells (also comprising, e.g., B2M KO / TNFAIP3-P2A-PD-L1 KI; TXNIP KO / MANF-P2A-HLA-E trimer KI) are passaged the two days before electroporation and seeded as 5 million cells per T75 flask. On the day of electroporation, the cells are split again and electroporated using the Neon Electroporator with the RNP mixture of Cas9 protein (Biomay) and guide RNA (Agilent) at a molar ratio of 5: 1 (gRNA:Cas9) with absolute values of 125 pmol Cas9 and 625 pmol gRNA per 2 million cells. To form the RNP complex, gRNA and Cas9 are combined in one vessel with R-buffer (Neon Transfection Kit) to a total volume of 25-50 pL and incubated for 15 minutes at room temperature (RT). This mixture is combined with the cells and 4 pg plasmid to a total volume of -115 pL using R-buffer. This mixture is electroporated with 2 pulses for 25 ms at 1400 V. Following electroporation, the cells are pipetted out into a 6 well plate filled with STEMFLEX™ media with REVITACELL™ Supplement (100X) and BIOLAMININ 521 CTG at 1 :10 dilution. Cells are cultured in a normoxia incubator (37°C, 8% CO2). To create single cell clone, cells are FACS sorted for CD 19 into 96 well plate at day 2 after electroporation. After 12-14 days postelectroporation, sorted single cell clones are genotyped by PCR.
[0233] For genotyping GLP1 or TIRZ knock-in in the target INSULIN sequence, PCR for relevant regions is performed using a 2-step protocol with Phusion Flash High Fidelity PCR master mix (Thermo Fisher, cat#F548S). The sequences of the PCR primers are presented in Table 1 and the cycling conditions are provided in Table 2. The presence of a 753 bp band indicatessuccessful integration of the KI construct into the INSULIN gene locus, while the presence of a 588 bp band indicates a WT genotype.
[0234] For determining the presence of any unwanted bacterial plasmid elements from the KI plasmid, two PCRs are performed using Phusion Flash High Fidelity PCR master mix (Thermo Fisher, cat#F548S). The sequences of the PCR primers are presented in Table 3 and Table 5, and the cycling conditions are provided in Table 4 and Table 6. PCR for the genotyping of the edited clones (GLP1 knock-in, into INSULIN locus of the cells) is performed and the resulting amplified DNA assessed for cutting efficiency by TIDE analysis. PCR for the genotyping of the edited clones (TIRZ knock-in, into INSULIN locus of the cells) is performed and the resulting amplified DNA assessed for cutting efficiency by TIDE analysis.Table 1 : Plasmid #1 PrimersTable 2: Plasmid #1 PCR Cycling ParametersTable 3: Plasmid #1 PrimersTable 4: Plasmid #1 PCR Cycling ParametersTable 5: Plasmid #2 PrimersTable 6: Plasmid #2 PCR Cycling ParametersTable 7: Elements of GLP1 Donor PlasmidTable 8: Elements of TIRZ Donor PlasmidExample 2Differentiation and characterization of genetically modified cells
[0235] Provided in this Example are non-limiting exemplary methods related to differentiation and characterization of the edited cells described herein.Differentiation of Edited Human Embryonic Stem cells to Stage 6 cells (S6)Maintenance of edited human embryonic stem cells (ES)
[0236] The edited human pluripotent stem cells at various passages (P38-42) are maintained by seeding at 33,000 cells / cm2for a 4-day passage or 50,000 cells / cm2for a 3-day passage with hESM medium (DMEM / F12+10% KSR+ 10 ng / mL Activin A and 10 ng / mL Heregulin) and final 10% human AB serum.Aggregation of edited human embryonic stem cells for PECs differentiation
[0237] The edited cells are dissociated into single cells with ACCUTASE® and thencentrifuged and resuspended in 2% StemPro (Cat# A 1000701, Invitrogen, CA) in DMEM / F12 medium at 1 million cells per ml, and total 350-400 million of cells are seeded in one 850 cm2roller bottle (Cat#431198, Corning, NY) with rotation speed at 8RPM±0.5RPM for 18-20 hours before differentiation. Aggregates from edited human pluripotent stems cells are differentiated into pancreatic lineages as described in Rezania et al. (2014) Nat. Biotechnol. 32(11): 1121-1133 and US20200208116.
[0238] The differentiated aggregates are expected to show smooth and round surface. Yield of GLP1KI or TIRZKI are expected to be similar to un-edited cells.Gene Expression at PEC Stage and Stage 6
[0239] Targeted RNAseq for gene expression analysis is performed using Illumina TruSeq and a custom panel of oligos targeting 111 genes. The panel primarily contains genes that are markers of the developmental stages during pancreatic differentiation. At the end of PEC stage and Stage 6, 10 pL APV (aggregated pellet volume) is collected and RNA is extracted using the Qiagen RNeasy or RNeasy 96 spin column protocol, including on-column DNase treatment. Quantification and quality control are performed using either the TapeStation combined with Qubit, or by using the Qiagen QIAxcel. 50-200 ng of RNA is processed according to the Illumina TruSeq library preparation protocol, which consists of cDNA synthesis, hybridization of the custom oligo pool, washing, extension, ligation of the bound oligos, PCR amplification of the libraries, and clean-up of the libraries, prior to quantification and quality control of the resulting dsDNA libraries using either the TapeStation combined with Qubit, or by using the Qiagen QIAxcel. The libraries are subsequently diluted to a concentration of 4 nM and pooled, followed by denaturing, spike-in of PhiX control, and further dilution to 10-12 pM prior to loading on the Illumina MiSeq sequencer. Following the sequencing run, initial data analysis is performed automatically through BaseSpace, generating raw read counts for each of the custom probes. For each gene, these read counts are then summed for all probes corresponding to that gene, with the addition of 1 read count (to prevent downstream divisions by 0). Normalization is performed to the gene SF3B2, and the reads are visualized as fold change vs. Stage 0. When the data is processed for principal component analysis, normalization is performed using the DEseq method. In parallel, the purified RNA is submitted to CRO (MiRXES) for totalRNA sequencing.
[0240] The selective gene expression pattern and PCA plot from totalRNAseq is expected to show that GLP1KI or TIRZKI cell lines have good S6 differentiation.Flow Cytometry for CHGA. PDX1 and NKX6, 1 at PEC Stage and Stage 6
[0241] Stage 6 aggregates are washed with PBS and then enzymatically dissociated to single cell suspension at 37°C using ACCUMAX™ (Catalog# A7089, Sigma, MO). MACS Separation Buffer (Cat# 130-091-221, Miltenyi Biotec, North Rhine-Westphalia, Germany) isadded and the suspension passed through a 40 pm filter and pelleted. For intracellular marker staining, cells are fixed for 30 minutes in 4% (wt / v) paraformaldehyde, washed in FACS Buffer (PBS, 0.1% (wt / v) BSA, 0.1% (wt / v) NaN3) and then cells are permeabilized with Perm Buffer (PBS, 0.2% (v / v) Triton X-100 (Cat#A16046, Alfa Aesar, MA), 5% (v / v) normal donkey serum, 0.1% (wt / v) NaN3) for 30 minutes on ice and then washed with washing buffer (PBS, 1% (wt / v) BSA, 0.1% (wt / v) NaN3). Cells are incubated with primary antibodies (Table 9) diluted with Block Buffer (PBS, 0.1% (v / v) Triton X-100, 5% (v / v) normal donkey serum, 0.1% (wt / v) NaN3) overnight at 4°C. Cells are washed in IC buffer and then incubated with appropriate secondary antibodies for 60 minutes at 4°C. Cells are washed in IC buffer and then in FACS Buffer. Flow cytometry data are acquired with NovoCyte Flow Cytometer (ACEA Biosciences, Brussels). Data are analyzed using FlowJo software (Tree Star, Inc.). Intact cells are identified based on forward (low angle) and side (orthogonal, 90°) light scatter. Background is estimated using antibody controls and undifferentiated cells. In some embodiments, at least 40% of the cells of an S6 population differentiated from GLP1 KI or TIRZ KI are expected to display markers for endoderm or beta cells.Table 9: Antibodies for Flow CytometryGLP1 ELISA at Stage 6
[0242] 24 hours condition media are collected from 1 million of differentiated S6 cells from un-edited cells or GLP1KI or TIRZKI cells and applied to human GLP1 ELISA kit from BLI (cat#27201). Stage 6 (S6) cells differentiated from GLP1KI or TIRZKI cells are expected to show GLP1 expression above background (e.g., compared to cells that do not have the GLP1KI or the TIRZKI) as determined by ELISA.GLP-1 bioassay at Stage 6
[0243] HEK 293 cells stably expressing the human GLP-1 receptor (HEK-hGLP-lR) are purchased from BPS Bioscience (Cat#78176). These cells are transfected with pHTS-CRE plasmid (Biomyx, San Diego, CA) using Lipofectamine 2000 (Invitrogen) according to the manufacturer's instructions. These cells (HEK-hGLP-lR-Luc) are grown in HG-DMEM supplemented with 800 pg / ml each of hygromycin and geneticin (Invitrogen). For GLP-1 measurements, HEK-hGLP-lR-Luc cells are plated in a 96-well plate at a density of 5 * 104cells / well and incubated overnight at 37°C. The next day, the medium is aspirated and replaced with 100 pl of either GLP-1 standards, or 24 hours condition medium collected from differentiatedS6 cells from Modified Parent Cells (e.g., cells that do not comprise the GLP1KI or the TIRZ KI) or GLP1KI cells (e.g., cells comprising GLP1KI and all gene edits in Modified Parent Cells) or TIRZKI cells (e.g., cells comprising TIRZKI and all gene edits in Modified Parent Cells). Following a 5-h incubation at 37°C, a luciferase assay is performed using the Bright-Glo luciferase assay kit (Promega, Madison, MI) according to the manufacturer’s instructions. Luminescence is measured using a Synergy Hl microplate reader. Stage 6 (S6) cells differentiated from GLP1KI or TIRZKI cells are expected to display at least 2-fold increased bioactivity relative to un-edited cells.In vivo Maturation of Stage 6 (S6) cells differentiated from Modified Parent Cells or GLP1KI or TIRZKI cells in NSG mice
[0244] The S6 cells differentiated from GLP1KI and TIRZKI cells are transplanted under the kidney capsule of NSG immunodeficient mice to test their in vivo maturation. As summarized in Table 10, four groups of the NSG mice are transplanted with S6-GLP1KI and S6- TIRZKI cells which derived from four clonal lines (two GLP1 KI lines and two TIRZKI lines). Each mouse receives approximately 7 x 106S6 cells. These transplanted pancreatic progenitor cells mature in vivo into functional pancreatic endocrine cells including glucose-responsive, insulin-producing cells.
[0245] Starting at 4 weeks all surviving animals are subjected to efficacy evaluation through glucose stimulated insulin secretion (GSIS) testing. Blood samples are obtained from fasted animals after intraperitoneal administration of 3g / kg glucose. Serum concentrations of human C-peptide (pmol / L) are determined through standard enzyme linked immunosorbent assays.
[0246] GSIS testing is performed at 4, 8, 12, 16, 20, and 28 weeks. S6-TIRZKI and S6-TIRZKI groups are expected to present a similar range of C-peptide levels. The C-peptide levels are expected to be elevated to around 1300 ~ 2000 pmol / L at 28-week time point.
[0247] At 30 weeks, animals are euthanized and the kidney with S6 cells collected and fixed in neutral buffered formalin, processed to slides, and stained with H&E and by immunohistochemistry for insulin and cyst structure marker proteins (DLK1, SOX9, and CD9).Table 10: Study DesignIn vivo Obesity Model Establishment
[0248] An obesity model is established using SCID-BEIGE mice receiving normal chow or high fat (60% kcal) diet for around 5 weeks with weekly body weight monitor. Glucose intolerance test is performed to confirm the establishment of obesity prior to S6 cell transplantation. Blood glucose levels of fasted animals are obtained prior to and 15, 30, 60, 90, and 120 minutes after intraperitoneal administration of 3g / kg glucose. The area under curve is calculated. Animals receiving high fat diet are expected to have increased body weight and reduced glucose uptake capability compared to animals receiving normal chow.In vivo therapeutic efficacy study of S6-GLP1KI and S6-TIRZKI cells in high fat diet SCID- BEIGE mouse model
[0249] An obesity model is established using SCID-BEIGE mice receiving normal chow or high fat (60% kcal) diet for around 5 weeks with weekly body weight monitor. Glucose intolerance test is performed to confirm the establishment of obesity prior to S6 cell transplantation. As summarized in Table 11, animals are divided into four groups. One group is transplanted with S6-GLP1KI cells and the other group is transplanted with S6-TIRZKI cells. The other two groups are control groups. Each mouse receives approximately 7 * 106S6 cells. These transplanted pancreatic progenitor cells mature in vivo into functional pancreatic endocrine cells including glucose-responsive, insulin-producing cells.Table 11 : Study Design
[0250] Animal body weight is continuously monitored weekly till the end of the study.The results are expected to show that all groups that received high fat diet have increased body weight compared to the group that receives normal chow prior to S6 cell transplantation. After transplantation, however, animals transplanted with S6-GLP1KI-C5 and S6-TIRZKI-C12 cells are expected to show reversal of obesity compared to the group with high fat diet but no S6 transplantation.
[0251] The glucose intolerance test is performed at 15-weeks post-transplantation.Blood glucose levels of fasted animals are obtained prior to and 15, 30, 60, 90, and 120 minutes after intraperitoneal administration of 3g / kg glucose. The area under curve is also calculated. The impaired glucose uptake is expected to be ameliorated by the presence of both S6-GLP1KI-C5 and S6-TIRZKI-C12 cells.Table 12: Sequences of the DisclosureTable 13: Exemplary GLP-1R Agonist Amino Acid SequencesExample 3Generation and characterization of genetically modified cells
[0252] Described in this Example are methods for generating, e.g., genetically modified cells expressing GLP-1 or Tirzepatide and characterization of said genetically modified cells.Generation of GLP1-KI into G26 Human Pluripotent Stem Cells
[0253] Cells were generated in which a polynucleotide encoding GLP1 was inserted into the Insulin gene locus. GLPl-EFla-EGFP KI was knocked into the cells using the GLP1plasmid (SEQ ID NO: 98) and the INS-G13 gRNA (SEQ ID NO: 1). G26 cells were passaged the two days before electroporation and seeded as 5 million cells per T75 flask. On the day of electroporation, the cells were split again and electroporated using the Neon Electroporator with the RNP mixture of Cas9 protein (Biomay) and guide RNA (Agilent) at a molar ratio of 5: 1 (gRNA:Cas9) with absolute values of 125 pmol Cas9 and 625 pmol gRNA per 2 million cells. To form the RNP complex, gRNA and Cas9 was combined in one vessel with R-buffer (Neon Transfection Kit) to a total volume of 25-50 pL and incubated for 15 minutes at room temperature (RT). This mixture was then combined with the cells and 4 pg plasmid to a total volume of ~115 pL using R-buffer. This mixture was then electroporated with 2 pulses for 25 ms at 1400 V. Following electroporation, the cells were pipetted out into a 6 well plate filled with STEMFLEX™ media with REVITACELL™ Supplement (100X) and BIOLAMININ 521 CTG at 1 : 10 dilution. Cells were cultured in a normoxia incubator (37°C, 5% CO2). To create single cell clone, cells were FACS sorted for EGFP into A 96 well plate at day 4 after electroporation. After 12-14 days post, sorted single cell clones were genotyped by PCR. Information of single clone is shown below in Table 22.Generation of TIRZ-KI into G26 Human Pluripotent Stem Cells
[0254] Cells were generated in which a polynucleotide encoding GLP1 was inserted into the Insulin gene locus. TIRZ- EFla-EGFP KI was knocked into the cells using the TRIZ plasmid (SEQ ID NO: 99) and the INS-G13 gRNA (SEQ ID NO: 1). The G26 cells were passaged the two days before electroporation and seeded as 5 million cells per T75 flask. On the day of electroporation, the cells were split again and electroporated using the Neon Electroporator with the RNP mixture of Cas9 protein (Biomay) and guide RNA (Agilent) at a molar ratio of 5: 1 (gRNA:Cas9) with absolute values of 125 pmol Cas9 and 625 pmol gRNA per 2 million cells. To form the RNP complex, gRNA and Cas9 were combined in one vessel with R-buffer (Neon Transfection Kit) to a total volume of 25-50 pL and incubated for 15 minutes at room temperature (RT). This mixture was then combined with the cells and 4 pg plasmid to a total volume of ~115 pL using R-buffer. This mixture was then electroporated with 2 pulses for 25 ms at 1400 V. Following electroporation, the cells were pipetted out into a 6 well plate filled with STEMFLEX™ media with REVITACELL™ Supplement (100X) and BIOLAMININ 521 CTG at 1 : 10 dilution. Cells were cultured in a normoxia incubator (37°C, 5% CO2). To create single cell clone, cells were FACS sorted for EGFP into 96 well plate at day 2 after electroporation. After 12-14 days, sorted single cell clones were performed genotyping by PCR. The information of single clone is shown below in Table 23.
[0255] For determining GLP1 or TIRZ knock-in genotyping in the target INSULIN sequence, PCR for relevant regions was performed using a 2-step protocol with Phusion FlashHigh Fidelity PCR master mix (Thermo Fisher, cat#F548S). The sequences of the PCR primers are presented in Table 14; and the cycling conditions are provided in Table 15.Genotyping results of the transgene KI into INSULIN gene locus for various edited clones.
[0256] The presence of a 729 bp band indicated successful integration of the KI construct into the INSULIN gene locus, while the presence of a 588 bp band indicated a WT genotype.
[0257] For determining the presence of any unwanted bacterial plasmid elements from the KI plasmid, two PCRs were performed using Phusion Flash High Fidelity PCR master mix (Thermo Fisher, cat#F548S). The sequences of the PCR primers are presented in Table 16 and Table 18. The cycling conditions are provided in Table 17 and 19.
[0258] PCR for the genotyping of the edited clones (GLP1 knock-in, into INSULIN locus of G26 cells) was performed and the resulting amplified DNA was assessed for cutting efficiency by TIDE analysis. PCR for the genotyping of the edited clones (TIRZ knock-in, into INSULIN locus of G26 cells) was performed and the resulting amplified DNA was assessed for cutting efficiency by TIDE analysis.Table 14: Plasmid #1 PrimersTable 15: Plasmid #1 PCR Cycling ParametersTable 16: Plasmid #1 PrimersTable 17: Plasmid #1 PCR Cycling ParametersTable 18: Plasmid #2 PrimersTable 19: Plasmid #2 PCR Cycling ParametersTable 20: Elements of GLP1 Donor PlasmidTable 21 : Elements of TIRZ Donor PlasmidTable 22: Information of GLP1 Knock-in and INS Knock-outTable 23: Information of TIRZ Knock-in and INS Knock-outTable 24: Sequences of the DisclosureTable 25: Information of GLP1 or TIRZKI knock-in cell linesDifferentiation of Edited Human Induced Pluripotent Stem Cells to Stage 6 cellsMaintenance of edited human induced pluripotent cells (iPSC) in 2D culture
[0259] The edited human pluripotent stem cells were maintained by seeding at 27,000 cells per cm2for a 4-day passage with StemFlex media.Procedure of iPSC 3D culture in StemScale
[0260] For preparation of complete StemScale media, StemScale supplement was thawed and 100 mL of StemScale supplement was added to StemScale basal media (900 mL). Finally, 5 mL of Penicillin / Streptomycin was added.
[0261] Next, Accutase was pre-warmed in 37°C water bath and StemFlex media was allowed to come to room temperature. Next, iPSC cell culture in StemFlex was obtained from the incubator into biological safety cabinet (BSC). Media was removed from cells, and the cells were washed with proper amount of DPBS (CaCh-ZMgCh-), making sure not to aspirate cells. Warm Accutase was added, and the cell culture was placed back into the incubator for 6 minutes. Next, cells were checked for detachment and returned to incubator for an additional 2-3 minutes if needed.
[0262] Next, a proper volume of quenching media (complete StemFlex media) was added, and the cell suspension was transferred to a centrifuge tube. 500 pL of cell suspension was aliquoted for cell counting by NuclearCounter NC200. Based on the cell concentration, 100 x 106cells suspension was prepared and centrifuged at 1000 rpm for 4 minutes. Supernatant was aspirated and the cell pellet was flicked to break it up.
[0263] Next, the cell pellet was resuspended in complete StemScale media to a final concentration of 2 x 105cells / mL in 500 mL of StemScale media with 5 pM of Thiazovivin in 500 mL PBSmini. PBSmini vessel was placed onto the PBS base inside incubator and speed set at 45 rpm. This is Day 0 iPSC-3D culture.
[0264] On Day 1, some cell culture suspension was removed to check on the aggregate size under the microscope. No media changing is needed at this point. On Day 2, some cell culture suspension was removed to check the aggregate size under the microscope. All of the iPSC 3D aggregates were transferred into a 500 mL centrifuge bottle and the aggregates settled by gravity for 15-20 minutes. Half of the StemScale culture media was removed and 250 mL of fresh complete StemScale media added. On Day 3, some cell culture suspension was removed to check on the aggregate size under the microscope. Typically, the aggregate size reaches 300-400 pm on day 3; then it is ready for differentiation. The cell aggregates were differentiated into immature beta cells as described in the U.S. Provisional Application No. 63 / 776,284 filed March 24, 2025, the content of which is hereby incorporated by reference in its entirety. FIG. 8 displays differentiation scheme from induced pluripotent cells to immature P cells (stage 6). Shown in FIG. 9A-FIG. 9B show the yield of differentiated from TIRZ KI edit cells.Flow Cytometry for PDXL NKX6.L Insulin and ISL1 at Stage 6
[0265] Stage 6 aggregates were washed with PBS and then enzymatically dissociated to single cells suspension at 37°C using ACCUMAX™ (Catalog# A7089, Sigma, MO). MACS Separation Buffer (Cat# 130-091-221, Miltenyi Biotec, North Rhine-Westphalia, Germany) was added and the suspension was passed through a 40 pm filter and pelleted. For intracellular marker staining, cells were fixed for 30 minutes in 4% (wt / v) paraformaldehyde, washed in FACS Buffer (PBS, 0.1% (wt / v) BSA, 0.1% (wt / v) NaN3) and then cells were permeabilized with Perm Buffer (PBS, 0.2% (v / v) Triton X-100 (Cat#A16046, Alfa Aesar, MA), 5% (v / v) normal donkey serum, 0.1% (wt / v) NaN3) for 30 minutes on ice and then washed with washing buffer (PBS, 1% (wt / v) BSA, 0.1% (wt / v) NaN3). Cells were incubated with primary antibodies (Table 26) diluted with Block Buffer (PBS, 0.1% (v / v) Triton X-100, 5% (v / v) normal donkey serum, 0.1% (wt / v) NaN3) overnight at 4°C. Cells were washed in IC buffer and then incubated with appropriate secondary antibodies for 60 minutes at 4°C. Cells were washed in IC buffer and then in FACS Buffer. Flow cytometry data were acquired with NovoCyte Flow Cytometer (ACEA Biosciences, Brussels). Data were analyzed using NovoExpress software (ACEA Biosciences, Brussels). Intact cells were identified based on forward (low angle) and side (orthogonal, 90°) light scatter. Background was estimated using antibody controls and undifferentiated cells. In the figures, a representative flow cytometry plot is shown for one of the sub-populations. Numbers reported in the figures represent the percentage of total cells from the intact cells gate. FIG. 10 shows 41% of total cells are endoderm cells (PDX1+ / NKX6.1+) and 24% of total cells are beta cells (ISL1+ / NKX6.1) from Stage 6 (S6) cells differentiated from TIRZ KI edit cells.Table 26: Antibodies for Flow CytometryGLP-1 bioassay at Stage 6
[0266] HEK 293 cells stably expressing the human GLP-1 receptor (HEK-hGLP-lR) were purchased from BPS Bioscience (Cat#78176). These cells were transfected with pHTS-CRE plasmid (Biomyx, San Diego, CA) using Lipofectamine 2000 (Invitrogen) according to the manufacturer's instructions. These cells (HEK-hGLP-lR-Luc) were grown in HG-DMEM supplemented with 800 pg / mL each of hygromycin and geneticin (Invitrogen). For GLP-1 measurements, HEK-hGLP-lR-Luc cells were plated in a 96-well plate at a density of 5 * 104cells / well and incubated overnight at 37°C. The next day, the medium was aspirated and replaced with 100 pL of either GLP-1 standards, or 24 hours condition medium collected fromdifferentiated S6 cells from TIRZ KI edit cells. Following a 5h incubation at 37°C, a luciferase assay was performed using the Bright-Glo luciferase assay kit (Promega, Madison, MI) according to the manufacturer's instructions. Luminescence was measured using a Synergy Hl microplate reader.
[0267] FIG. 11 shows the GLP1 bioactivity in the condition media from Stage 6 (S6) cells differentiated from TIRZ KI edit cells. The data suggested that the secreted TIRZ peptide from TIRZ KI S6 cells is biologically active.In vivo Maturation of Stage 6 (S6) cells differentiated from TIRZ KI cells in NSG mice
[0268] The S6 cells (3.5 x io6) differentiated from TIRZKI cells were transplanted under the kidney capsule of NSG immunodeficient mice to test their in vivo maturation. These transplanted immature beta cells mature in vivo into functional pancreatic endocrine cells including glucose-responsive, insulin-producing cells. Starting at 8 weeks, all surviving animals were subjected to efficacy evaluation through glucose stimulated insulin secretion (GSIS) testing. Blood samples were obtained from fasted animals after intraperitoneal administration of 3g / kg glucose. Serum concentrations of human C-peptide (pmol / L) were determined through standard enzyme linked immunosorbent assays. GSIS testing was performed at 8, 12, 16, 20, 24 and 32 weeks. FIG. 12 showed that the C-peptide levels were elevated to 600 pmol / L from wkl6 to wk32 time-point.Example 4In vivo Therapeutic Efficacy study of S6-GLP1KI and S6-TIRZKI cells in high fat diet SCID- BEIGE mouse model
[0269] Described in this Example are in vivo studies of efficacy. An obesity model is established using SCID-BEIGE mice receiving normal chow or high fat (60% kcal) diet for around 5 weeks with weekly body weight monitor (FIG. 13 A). Glucose intolerance test is performed to confirm the establishment of obesity prior to S6 cell transplantation (FIG. 13B). As summarized in Table 27, animals are divided into four groups. One group is transplanted with S6-GLP1KI cells and the other group is transplanted with S6-TIRZKI cells. Each mouse received approximately 7 x 106S6 cells. These transplanted pancreatic progenitor cells mature in vivo into functional pancreatic endocrine cells including glucose-responsive, insulin-producing cells.Table 27: Study Design
[0270] It is expected that the genetically modified cells will be therapeutically effective in vivo, e.g., resulting in decreased weight and blood glucose in the mouse model.
[0271] In at least some of the previously described embodiments, one or more elements used in an embodiment can interchangeably be used in another embodiment unless such a replacement is not technically feasible. It will be appreciated by those skilled in the art that various other omissions, additions and modifications may be made to the methods and structures described above without departing from the scope of the claimed subject matter. All such modifications and changes are intended to fall within the scope of the subject matter, as defined by the appended claims.
[0272] With respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated.
[0273] It will be understood by those within the art that, in general, terms used herein, and especially in the appended claims (e.g., bodies of the appended claims) are generally intended as “open” terms (e.g., the term “including” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “having at least,” the term “includes” should be interpreted as “includes but is not limited to,” etc.). It will be further understood by those within the art that if a specific number of an introduced claim recitation is intended, such an intent will be explicitly recited in the claim, and in the absence of such recitation no such intent is present. For example, as an aid to understanding, the following appended claims may contain usage of the introductory phrases “at least one” and “one or more” to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles “a” or “an” limits any particular claim containing such introduced claim recitation to embodiments containing only one such recitation, even when the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “an” (e.g., “a” and / or “an” should be interpreted to mean “at least one” or “one or more”); the same holds true for the use of definite articles used to introduce claim recitations. In addition, even if a specific number of an introduced claim recitation is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number (e.g., thebare recitation of “two recitations,” without other modifiers, means at least two recitations, or two or more recitations). Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, etc.” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “ a system having at least one of A, B, and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). In those instances where a convention analogous to “at least one of A, B, or C, etc.” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “ a system having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms.
[0274] In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.
[0275] As will be understood by one skilled in the art, for any and all purposes, such as in terms of providing a written description, all ranges disclosed herein also encompass any and all possible sub-ranges and combinations of sub-ranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art all language such as “up to,” “at least,” “greater than,” “less than,” and the like include the number recited and refer to ranges which can be subsequently broken down into sub-ranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a group having 1-3 articles refers to groups having 1, 2, or 3 articles. Similarly, a group having 1-5 articles refers to groups having 1, 2, 3, 4, or 5 articles, and so forth.
[0276] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.
Claims
WHAT IS CLAIMED IS:
1. A genetically modified cell, comprising: a heterologous nucleotide sequence comprising a sequence encoding a glucagon- like peptide-1 (GLP-1) Receptor (GLP-1R) agonist; and wherein the cell expresses the GLP-1R agonist.
2. The genetically modified cell of claim 1, wherein the heterologous nucleotide sequence is inserted into the genome of the cell.
3. The genetically modified cell of claim 2, wherein the cell is heterozygous or homozygous for the insertion.
4. The genetically modified cell of any one of claims 1-3, wherein the cell is capable of secreting the GLP-1 R agonist.
5. The genetically modified cell of any one of claims 1-4, wherein the heterologous nucleotide sequence further comprises one or more of: a) a sequence encoding a first pro-protein convertase cleavage site; b) a sequence encoding an intervening peptide-1 (IP-1) sequence of proglucagon; c) a sequence encoding a second pro-protein convertase cleavage site; and d) a sequence encoding at least a portion of an intervening peptide-2 (IP-2) sequence of proglucagon.
6. The genetically modified cell of any one of claims 1-5, wherein the heterologous nucleotide sequence comprises a sequence encoding a GLP-1R agonist protein precursor (proprotein), wherein the GLP1-R agonist pro-protein comprises, from the N-terminus to the C- terminus: a) a first pro-protein convertase cleavage site; b) an intervening peptide-1 (IP-1) sequence of proglucagon; c) the GLP-1 R agonist; d) a second pro-protein convertase cleavage site; and e) at least a portion of an intervening peptide-2 (IP -2) sequence of proglucagon.
7. The genetically modified cell of any one of claims 2-6, wherein the heterologous nucleotide sequence is inserted into a safe harbor locus; optionally the safe harbor locus is selected from the group consisting of AAVS1 gene locus, HRPT gene locus, CCR5 gene locus, globin gene locus, TTR gene locus, TF gene locus, F9 gene locus, Alb gene locus, Gys2 gene locus and PCSK9 gene locus.
8. The genetically modified cell of any one of claims 2-6, wherein the heterologous nucleotide sequence is inserted into the insulin (INS) gene locus of the genome of the cell.
9. The genetically modified cell of claim 8, wherein the insertion of the heterologousnucleotide sequence disrupts the INS gene, thereby the expression of insulin by the cell is reduced or eliminated.
10. The genetically modified cell of any one of claims 8-9, wherein the heterologous nucleotide sequence is inserted into the INS gene locus at a site downstream of the portion of the INS gene sequence that encodes insulin signal peptide.
11. The genetically modified cell of any one of claims 8-10, wherein the heterologous nucleotide sequence is inserted into the INS gene locus downstream of the portion of the INS gene sequence that encodes insulin signal peptide, thereby the cell is capable of expressing a GLP-1R agonist fusion protein comprising, from N-terminus to C-terminus, (i) the insulin signal peptide and the GLP-1R agonist or (ii) the insulin signal peptide and the GLP-1R agonist pro-protein.
12. The genetically modified cell of claim 11, wherein the GLP-1R agonist fusion protein comprises, from N-terminus to C-terminus: a) the insulin signal peptide; b) the first proprotein convertase cleavage site; c) the IP-1 sequence of proglucagon; d) the GLP-1R agonist; e) the second proprotein convertase cleavage site; and f) the at least a portion of the IP-2 sequence of proglucagon.
13. The genetically modified cell of any one of claims 6-12, wherein the genetically modified cell expresses the GLP-1R agonist pro-protein, the GLP-1R agonist fusion protein, or both.
14. The genetically modified cell of claim 13, wherein the GLP-1R agonist pro-protein or the GLP-1R agonist fusion protein is: i) capable of being enzymatically cleaved by one or more proprotein convertases in the cell to generate the GLP-1R agonist, and optionally ii) the GLP-1R agonist is amidated at the C-terminal end.
15. The genetically modified cell of any one of claims 1-14, wherein the GLP-1R agonist comprises a GLP-1 peptide or a functional fragment thereof, or a variant thereof; or wherein the GLP-1R agonist comprises a chimeric peptide comprising at least a portion of a GLP- 1 peptide and at least a portion of a gastric inhibitory peptide (GIP).
16. The genetically modified cell of any one of claims 1-15, wherein the heterologous nucleotide sequence comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 21 or SEQ ID NO: 23.
17. The genetically modified cell of any one of claims 1-15, wherein the heterologous nucleotide sequence comprises the sequence of SEQ ID NO: 21 or SEQ ID NO: 23.
18. The genetically modified cell of any one of claims 1-15, wherein the heterologous nucleotide sequence comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 8 or SEQ ID NO: 10.
19. The genetically modified cell of any one of claims 1-15, wherein the heterologous nucleotide sequence comprises the sequence of SEQ ID NO: 8 or SEQ ID NO: 10.
20. The genetically modified cell of any one of claims 1-19, wherein the heterologous nucleotide sequence comprises a stop codon at the 3’ end.
21. The genetically modified cell of any one of claims 1-20, wherein the GLP-1R agonist comprises an amino acid sequence comprising the sequence of SEQ ID NO: 34, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 34, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 34.
22. The genetically modified cell of any one of claims 1-20, wherein the GLP-1R agonist comprises an amino acid sequence comprising the sequence of SEQ ID NO: 35, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 35, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 35.
23. The genetically modified cell of any one of claims 6-20, wherein the GLP-1R agonist pro-protein comprises an amino acid sequence comprising the sequence of SEQ ID NO:29, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 29, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 29.
24. The genetically modified cell of any one of claims 6-20, wherein the GLP-1R agonist pro-protein comprises an amino acid sequence comprising the sequence of SEQ ID NO:30, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 30, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 30.
25. The genetically modified cell of any one of claims 11-20, wherein the GLP-1R agonist fusion protein comprises an amino acid sequence comprising the sequence of SEQ ID NO:32, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 32, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 32.
26. The genetically modified cell of any one of claims 11-20, wherein the GLP-1R agonist fusion protein comprises an amino acid sequence comprising the sequence of SEQ ID NO:33, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 33, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 33.
27. The genetically modified cell of any one of claims 1-26, wherein the cell is a stem cell; optionally the cell is an embryonic stem cell, an induced pluripotent stem cell, or an adult stem cell.
28. The genetically modified cell of claim 27, wherein the stem cell is capable of beingdifferentiated into a pancreatic endocrine cell, an intestinal enteroendocrine cell, a neuron, a definitive endoderm cell, a primitive gut tube cell, a posterior foregut cell, a pancreatic endoderm cell, a pancreatic endocrine precursor cell, an immature beta cell, and / or a pancreatic beta cell.
29. The genetically modified cell of claim 27, wherein the stem cell is capable of being differentiated into a stage 6 (S6) pancreatic progenitor cell.
30. The genetically modified cell of claim 29, wherein the S6 cell expresses one or markers each selected from the group consisting of Forkhead Box A2 (FOXA2), Chromogranin A (CHGA), Pancreatic and duodenal homeobox 1 (PDX1), NK6 Homeobox 1 (NKX6.1), Neurogenin 3 (NGN3), Insulin (INS), and ISL LIM homeobox 1 (ISL1).
31. The genetically modified cell of any one of claims 1-26, wherein the cell is a differentiated cell selected from the group consisting of a pancreatic endocrine cell, an intestinal enteroendocrine cell, and a neuron.
32. The genetically modified cell of any one of claims 1-31, wherein the cell comprises one or more additional genetic modifications.
33. The genetically modified cell of claim 32, wherein the one or more additional genetic modifications comprise:(1) a disrupted beta-2-microglobulin (B2M) gene and insertion of a nucleotide sequence encoding programmed death-ligand 1 (PD-L1) into the disrupted B2M gene, wherein the cell expresses PD-L1 and has reduced or eliminated expression of B2M; and / or(2) a disrupted thioredoxin interacting protein (TXNIP) gene and insertion of a nucleotide sequence encoding HLA class I histocompatibility antigen, alpha chain E (HLA-E) into the disrupted TXNIP gene; optionally the nucleotide sequence encoding HLA-E encodes an HLA-E trimer comprising a B2M signal peptide fused to an HLA-G presentation peptide fused to a B2M membrane protein fused to HLA-E without its signal peptide; wherein the cell expresses HLA-E or the HLA-E trimer and has reduced or eliminated expression of TXNIP.
34. The genetically modified cell of claim 32, wherein the one or more additional genetic modifications comprise:(3) a disrupted B2M gene and insertion of a nucleotide sequence encoding PD-L1 and a nucleotide sequence encoding tumor necrosis factor alpha induced protein 3 (TNFAIP3) into the disrupted B2M gene, wherein the cell expresses PD-L1 and TNFAIP3 and has reduced or eliminated expression of B2M; and / or(4) a disrupted TXNIP gene and insertion of a nucleotide sequence encoding HLA- E and a nucleotide sequence mesencephalic astrocyte derived neurotrophic factor (MANF)into the disrupted TXNIP gene; optionally the nucleotide sequence encoding HLA-E encodes an HLA-E trimer comprising a B2M signal peptide fused to an HLA-G presentation peptide fused to a B2M membrane protein fused to HLA-E without its signal peptide; wherein the cell expresses HLA-E or the HLA-E trimer and MANF and has disrupted expression of TXNIP.
35. The genetically modified cell of claim 34, wherein the nucleotide sequence encoding TNFAIP3 and the nucleotide sequence encoding PD-L1 are linked by a nucleotide sequence encoding a P2A peptide such that the insertion of (3) comprises a nucleotide sequence encoding TNFAIP3-P2A-PD-LL36. The genetically modified cell of any one of claims 34-35, wherein the nucleotide sequence encoding the HLA-E trimer and the nucleotide sequence encoding MANF are linked by a nucleotide sequence encoding a P2A peptide such that the insertion of (4) comprises a nucleotide sequence encoding MANF-P2A-HLA-E trimer.
37. The genetically modified cell of any one of claims 33-36, wherein (a) the nucleotide sequence encoding PD-L1 comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 83; (b) the nucleotide sequence encoding the HLA-E trimer comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 84; (c) the nucleotide sequence encoding TNFAIP3 comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 85; (d) the nucleotide sequence encoding MANF comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 86; (e) the nucleotide sequence encoding TNFAIP3-P2A-PD-L1 comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 87; and / or (f) the nucleotide sequence encoding MANF-P2A-HLA-E trimer comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 89.
38. The genetically modified cell of claims 1-37, wherein the heterologous nucleotide sequence is a DNA sequence.
39. A method for generating a genetically modified cell, comprising delivering to a cell:(a) an RNA-guided nuclease and a gRNA targeting a target site within the genome of the cell; and(b) a nucleic acid comprising (i) a nucleotide sequence that is at least 85% identical to a region located upstream of the target site, (ii) a heterologous nucleotide sequence encoding a GLP-1 Receptor (GLP-1R) agonist, and (iii) a nucleotide sequence that is at least 85% identical to a region located downstream of the target site; thereby the target site in the genome of the cell is cleaved and the heterologous nucleotide sequence is inserted into the genome by homology directed repair (HDR), thereby generating agenetically modified cell, wherein the genetically modified cell expresses the GLP-1R agonist.
40. The method of claim 39, wherein the cell is heterozygous or homozygous for the insertion.
41. The method of any one of claims 39-40, wherein the cell is capable of secreting the GLP-1R agonist.
42. The method of any one of claims 39-41, wherein the heterologous nucleotide sequence further comprises one or more of a) a sequence encoding a first pro-protein convertase cleavage site; b) a sequence encoding an intervening peptide- 1 (IP-1) sequence of proglucagon; c) a sequence encoding a second pro-protein convertase cleavage site; and d) a sequence encoding at least a portion of an intervening peptide-2 (IP-2) sequence of proglucagon.
43. The method of any one of claims 39-42, wherein the heterologous nucleotide sequence comprises a sequence encoding a GLP-1R agonist protein precursor (pro-protein), wherein the GLP-R agonist pro-protein comprises, from the N-terminus to the C-terminus: a) a first pro-protein convertase cleavage site; b) an intervening peptide- 1 (IP-1) sequence of proglucagon; c) the GLP-1R agonist; d) a second pro-protein convertase cleavage site; and e) at least a portion of an intervening peptide-2 (IP -2) sequence of proglucagon.
44. The method of any one of claims 39-43, wherein the target site is within a gene locus; optionally wherein the insertion of the heterologous sequence into the target site disrupts the gene, thereby reducing or eliminating expression of a gene product encoded by the gene.
45. The method of claim 44, wherein the gene locus is a safe harbor locus; optionally the safe harbor locus is selected from the group consisting of AAVS1 gene locus, HRPT gene locus, CCR5 gene locus, globin gene locus, TTR gene locus, TF gene locus, F9 gene locus, Alb gene locus, Gys2 gene locus and PCSK9 gene locus.
46. The method of claim 44, wherein the gene locus is the insulin (INS) gene locus.
47. The method of claim 46, wherein the insertion of the heterologous nucleotide sequence disrupts the INS gene, thereby the expression of insulin by the cell is reduced or eliminated.
48. The method of any one of claims 46-47, wherein the target site is a site downstream of the sequence of the INS gene that encodes insulin signal peptide.
49. The method of any one of claims 46-48, wherein the heterologous nucleotide sequence is inserted into the INS locus downstream of the portion of the INS gene sequence thatencodes insulin signal peptide, thereby the cell is capable of expressing a GLP-1R agonist fusion protein comprising, from N-terminus to C-terminus, (i) the insulin signal peptide and the GLP-1R agonist or (ii) the insulin signal peptide and the GLP-1R agonist pro-protein.
50. The method of claim 49, wherein the GLP-1R agonist fusion protein comprises, from N-terminus to C-terminus: a) the insulin signal peptide; b) the first proprotein convertase cleavage site; c) the IP-1 sequence of proglucagon; d) the GLP-1R agonist; e) the second proprotein convertase cleavage site; and f) the at least a portion of the IP-2 sequence of proglucagon.
51. The method of any one of claims 43-50, wherein the genetically modified cell expresses the GLP-1R agonist pro-protein, the GLP-1R agonist fusion protein, or both.
52. The method of claim 51, wherein the GLP-1R agonist pro-protein or the GLP-1R agonist fusion protein is: i) capable of being enzymatically cleaved by one or more proprotein convertases in the cell to generate the GLP-1R agonist, and optionally ii) the GLP-1R agonist is amidated at the C-terminal end.
53. The method of any one of claims 39-52, wherein the GLP-1R agonist comprises a GLP-1 peptide or a functional fragment thereof, or a variant thereof; or wherein the GLP-1R agonist comprises a chimeric peptide comprising at least a portion of GLP-1 peptide and at least a portion of gastric inhibitory peptide (GIP).
54. The method of any one of claims 39-53, wherein the heterologous nucleotide sequence comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 21 or SEQ ID NO: 23; optionally the heterologous nucleotide sequence comprises the sequence of SEQ ID NO: 21 or SEQ ID NO: 23.
55. The method of any one of claims 39-54, wherein the heterologous nucleotide sequence comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 8 or SEQ ID NO: 10; optionally the heterologous nucleotide sequence comprises the sequence of SEQ ID NO: 8 or SEQ ID NO: 10.
56. The method of any one of claims 39-55, wherein the heterologous nucleotide sequence comprises a stop codon at the 3’ end.
57. The method of any one of claims 39-56, wherein the sequence of (b)(i) comprises the sequence of any one of SEQ ID NOs: 6, 19, and 96 or a sequence that is at least 85% identical to the sequence of any one of SEQ ID NOs: 6, 19, and 96.
58. The method of any one of claims 39-57, wherein the sequence of (b)(iii) comprises the sequence of SEQ ID NO: 22 or SEQ ID NO: 9 or a sequence that is at least 85% identical to the sequence of SEQ ID NO: 22 or SEQ ID NO: 9.
59. The method of any one of claims 39-58, wherein the nucleic acid of (b) comprises the sequence of any one of SEQ ID NOs: 11-12 and 97-98 or a sequence that is at least 85% identical to the sequence of any one of SEQ ID NOs: 11-12 and 97-98.
60. The method of any one of claims 39-59, wherein (a) the GLP-1R agonist comprises an amino acid sequence comprising the sequence of SEQ ID NO: 34, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 34, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 34; (b) the GLP-1R agonist comprises an amino acid sequence comprising the sequence of SEQ ID NO: 35, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 35, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 35; (c) the GLP-1R agonist pro-protein comprises an amino acid sequence comprising the sequence of SEQ ID NO: 29, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 29, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 29; (d) wherein the GLP-1R agonist pro-protein comprises an amino acid sequence comprising the sequence of SEQ ID NO: 30, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 30, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 30; (e) wherein the GLP-1R agonist fusion protein comprises an amino acid sequence comprising the sequence of SEQ ID NO: 32, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 32, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 32; or (f) wherein the GLP-1R agonist fusion protein comprises an amino acid sequence comprising the sequence of SEQ ID NO: 33, a sequence that is at least 95% identical to the sequence of SEQ ID NO: 33, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 33.
61. The method of any one of claims 39-60, wherein the RNA-guided nuclease is a CRISPR-associated nuclease; optionally the CRISPR-associated nuclease comprises an N- terminal nuclear localization signal (NLS), a C-terminal NLS, or both.
62. The method of claim 61, wherein the CRISPR-associated nuclease comprises Cas9, Cpfl, or a variant thereof; and optionally the Cas9 is S. pyogenes Cas9.
63. The method of claim 62, wherein the gRNA and the Cas9 are delivered to the cell at a molar ratio of 5 : 1.
64. The method of any one of claims 39-63, wherein the RNA-guided endonuclease and the gRNA are complexed as a ribonucleoprotein (RNP) complex prior to the delivering.
65. The method of any one of claims 39-64, wherein the gRNA comprises a spacersequence comprising the sequence of SEQ ID NO: 1, or a sequence having one, two, or three mismatches relative to the sequence of SEQ ID NO: 1.
66. The method of any one of claims 39-65, wherein the cell is a stem cell; optionally the cell is an embryonic stem cell, an induced pluripotent stem cell, or an adult stem cell.
67. The method of claim 66, wherein the stem cell is capable of being differentiated into a pancreatic endocrine cell, an intestinal enteroendocrine cell, a neuron, a definitive endoderm cell, a primitive gut tube cell, a posterior foregut cell, a pancreatic endoderm cell, a pancreatic endocrine precursor cells, an immature beta cell, and / or a pancreatic beta cell.
68. The method of claim 66, wherein the stem cell is capable of being differentiated into a stage 6 (S6) pancreatic progenitor cell; and optionally wherein the S6 cell expresses one or markers each selected from the group consisting of Forkhead Box A2 (FOXA2), Chromogranin A (CHGA), Pancreatic and duodenal homeobox 1 (PDX1), NK6 Homeobox 1 (NKX6.1), Neurogenin 3 (NGN3), Insulin (INS), and ISL LIM homeobox 1 (ISL1).
69. The method of any one of claims 39-65, wherein the cell is a differentiated cell selected from the group consisting of a pancreatic endocrine cell, an intestinal enteroendocrine cell, and a neuron.
70. The method of any one of claims 39-69, wherein the cell comprises one or more additional genetic modifications.
71. The method of claim 70, wherein the one or more additional genetic modifications comprise:(1) a disrupted beta-2-microglobulin (B2M) gene and insertion of a nucleotide sequence encoding programmed death-ligand 1 (PD-L1) into the disrupted B2M gene, wherein the cell expresses PD-L1 and has reduced or eliminated expression of B2M; and / or(2) a disrupted thioredoxin interacting protein (TXNIP) gene and insertion of a nucleotide sequence encoding HLA class I histocompatibility antigen, alpha chain E (HLA-E); optionally the nucleotide sequence encoding HLA-E encodes an HLA-E trimer comprising a B2M signal peptide fused to an HLA-G presentation peptide fused to a B2M membrane protein fused to HLA-E without its signal peptide; wherein the cell expresses HLA-E or the HLA-E trimer and has reduced or eliminated expression of TXNIP.
72. The method of claim 70, wherein the one or more additional genetic modifications comprise:(3) a disrupted beta-2-microglobulin (B2M) gene and insertion of a nucleotide sequence encoding PD-L1 and a nucleotide sequence encoding tumor necrosis factor alpha induced protein 3 (TNFAIP3) into the disrupted B2M gene, wherein the cell expresses PD-LI and TNFAIP3 and has reduced or eliminated expression of B2M; and / or(4) a disrupted thioredoxin interacting protein (TXNIP) gene and insertion of a nucleotide sequence encoding HLA class I histocompatibility antigen, alpha chain E (HLA-E) and a nucleotide sequence mesencephalic astrocyte derived neurotrophic factor (MANF) into the disrupted TXNIP gene; optionally the nucleotide sequence encoding HLA-E encodes an HLA-E trimer comprising a B2M signal peptide fused to an HLA-G presentation peptide fused to a B2M membrane protein fused to HLA-E without its signal peptide; wherein the cell expresses HLA-E or the HLA-E trimer and MANF and has disrupted expression of TXNIP.
73. The method of claim 72, wherein the nucleotide sequence encoding TNFAIP3 and the nucleotide sequence encoding PD-L1 are linked by a nucleotide sequence encoding a P2A peptide such that the insertion of (3) comprises a nucleotide sequence encoding TNFAIP3-P2A- PD-L1.
74. The method of any one of claims 72-73, wherein the nucleotide sequence encoding the HLA-E trimer and the nucleotide sequence encoding MANF are linked by a nucleotide sequence encoding a P2A peptide such that the insertion of (4) comprises a nucleotide sequence encoding MANF-P2A-HLA-E trimer.
75. The method of any one of claims 71-74, wherein (a) the nucleotide sequence encoding PD-L1 comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 83; (b) the nucleotide sequence encoding the HLA-E trimer comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 84; (c) the nucleotide sequence encoding TNFAIP3 comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 85; (d) the nucleotide sequence encoding MANF comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 86; (e) the nucleotide sequence encoding TNFAIP3-P2A-PD-L1 comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 87; and / or (f) the nucleotide sequence encoding MANF-P2A-HLA-E trimer comprises a sequence that is at least 85% identical to the sequence of SEQ ID NO: 89.
76. A population of cells comprising a plurality of the genetically modified cell of any one of claims 1-38 or comprising a plurality of the genetically modified cell generated by the method of any one of claims 39-75.
77. A composition comprising the population of cells of claim 76.
78. A pharmaceutical composition comprising the composition of claim 77 and a pharmaceutically acceptable excipient.
79. The composition of claim 77 or the pharmaceutical composition of claim 78, for use in treating a disease or disorder in a subject.
80. A method of treating a disease or disorder in a subject, comprising administering to a subject in need thereof a therapeutically effective amount of the composition of claim 77 or 78, thereby treating the disease or disorder in the subject.
81. The method of claim 80, wherein the disease or disorder is diabetes mellitus, dyslipidemia, fatty liver disease, metabolic syndrome, non-alcoholic steatohepatitis or obesity; and optionally wherein the diabetes mellitus is type 2 diabetes mellitus.
82. The method of claim 80, wherein the disease or disorder is a weight-related comorbid condition selected from the group consisting of hypertension, dyslipidemia, prediabetes, type 2 diabetes mellitus, obstructive sleep apnea, cardiovascular disease, and any combination thereof.
83. A method for chronic weight management in a subject in need thereof, comprising: administering to the subject a therapeutically effective amount of the composition of claim 77 or 78.
84. The method of any one of claims 80-83, wherein the subject has a BMI > 27 kg / m2and at least one weight- related comorbidity; and optionally wherein each of the at least one weight- related comorbidity is selected from the group consisting of hypertension, dyslipidemia, type 2 diabetes mellitus, coronary heart disease, stroke, gallbladder disease, osteoarthritis, sleep apnea, cancer, mental illness, and body pain.
Citation Information
Patent Citations
Methods and compositions for RNA-directed target DNA modification and for RNA-directed modulation of transcription
US20140068797A1
Generation of human pluripotent stem cell derived functional beta cells showing a glucose-dependent mitochondrial respiration and two-phase insulin secretion response
US20200208116A1
Uncharged morpholino-based polymers having achiral intersubunit linkages
US5034506A
Universal donor cells
WO2020049535A1
Universal donor cells
WO2021044377A1