Genome editing of Kozak sequences to treat disease
Genome editing of the Kozak sequence using CRISPR-Cas systems addresses the imbalance in protein levels of monogenic diseases by enhancing or suppressing translation efficiency, effectively treating haploinsufficiency and gene duplication disorders.
Patent Information
- Application Number
- JP2025505860
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-05
- Filing Date
- 2023-08-04
- Publication Date
- 2025-08-07
AI Technical Summary
Existing treatments for monogenic diseases caused by the loss or gain of function of one allele, such as haploinsufficiency and gene duplication, are inadequate in modulating translation efficiency to restore balanced protein levels without causing cellular imbalance.
Genome editing using CRISPR-Cas systems, including homology-directed repair, primed editing, and base editing, is employed to introduce specific nucleotide changes in the Kozak sequence to enhance or suppress translation efficiency, thereby correcting the functional imbalance of affected genes.
This approach allows for targeted modulation of protein production, restoring balanced gene expression levels without causing cellular imbalance, applicable to a wide range of monogenic diseases.
Smart Images

Figure 2025525882000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to the medical field of monogenic diseases caused by the loss or gain of function of one allele. The innovative approach developed is based on editing the human genome at the level of the Kozak sequence by means of CRISPR-Cas programmable nucleases. In particular, the present invention relates to variant Kozak sequences and related in vitro or in vivo methods for obtaining such variant Kozak sequences, which are applied for therapeutic use in the treatment of monogenic diseases caused by the loss or gain of function of one allele. These in vitro and in vivo methods include genome editing using CRISPR-Cas homology-directed repair, CRISPR-Cas primed editing, CRISPR-Cas base editing, or other programmable RNA-guided nucleases, and the introduction of specific nucleotide changes in the Kozak sequence of the disease-causing gene. These nucleotide changes promote or inhibit translation of the mRNA produced by the gene, compensating for the loss or gain of function of one allele in the disease. [Background technology]
[0002] Hundreds of disorders are caused by the loss or gain of function of one allele. Haploinsufficiency (HI), in particular, is a genetic condition in which mutational inactivation of one allele causes reduced protein levels, sufficient to result in a disease phenotype. Over 300 human diseases, ranging from cancer predisposition to developmental and neurological monogenic disorders, are caused by haploinsufficiency. Gene duplication (GD) instead results in the gain of one allele, resulting in elevated protein levels and a disease phenotype. Dozens of human diseases are caused by gene duplication.
[0003] The Kozak consensus sequence was first defined in the 1980s as the optimal nucleotide context around a protein start codon, and it is now widely accepted that the Kozak sequence plays a fundamental role in regulating translation by attracting specific interactions within the 48S initiation complex.
[0004] The impact of the Kozak sequence on disease is exemplified by hereditary and sporadic diseases in which point mutations and allelic variations near the AUG start codon affect translation efficiency, respectively. For example, a C-to-T mutation at position -1 relative to the AUG codon (c-1.C>T) in the Kozak sequence of the α-tocopherol transport protein gene reduces protein levels and causes the monogenic disorder AVED (ataxia due to vitamin E deficiency). As another example, a T / C polymorphism at position -1 of the CD40 gene (-1T>C, rs1883832) is associated with increased CD40 translation and therefore predisposes to Graves' disease and coronary heart disease.
[0005] Despite the seemingly rigid pattern of motifs describing the Kozak consensus sequence, sequence diversity exists around the AUG start codon in the human genome and in other vertebrate genomes, and recent studies aimed at measuring the strength of various AUG start codons as a function of surrounding bases have identified substantial variability. These results suggest that multiple genes may be translationally regulated by suboptimal Kozak sequences.
[0006] Given the citation data and its major role in regulating translation, Kozak sequences are also potential targets for modulating gene expression. Summary of the Invention
[0007] The present applicant has now discovered a genome editing method aimed at reversing diseases caused by loss or gain of function of one allele. Such a method can utilize any genome editing platform that can introduce single or multiple base substitutions within a short continuous region of DNA. Therefore, CRISPR-Cas homology-directed repair, CRISPR-Cas primed editing, CRISPR-Cas base editing, or any other programmable RNA-guided nuclease fused with an effector protein that allows the introduction of base substitutions into the genome are suitable.
[0008] In one embodiment of the present invention, Applicant focused on modifying the Kozak sequence for a HI gene of interest to enhance the translation efficiency of the wild-type allele. In other words, in this embodiment of the present invention, the procedure upregulates protein production through variation of the Kozak sequence for therapeutic purposes.
[0009] Specifically, as an example of this method, Applicants have now discovered specific nucleotide changes in the Kozak sequence of the HI gene that have been selected for use in the treatment of HI diseases, which fall into the categories of developmental disorders, metabolic syndromes, eye diseases, and hematopoietic disorders.
[0010] In another embodiment of the present invention, the applicant focused on modifying the Kozak sequence of a GD gene of interest to suppress the translation efficiency of the three alleles. In other words, in this embodiment of the present invention, the procedure downregulates protein production through variation of the Kozak sequence for therapeutic purposes.
[0011] A first object of the present invention therefore relates to a nucleotide sequence as defined in claim 1.
[0012] Advantageously, the present invention allows the use of targeted Kozak sequences by genome editing techniques aimed at modulating translation efficiency, which can be manipulated to up- or down-regulate protein production for therapeutic purposes. Other advantages of the approach on which the present invention is based are as follows: First, genome editing techniques can introduce permanent nucleotide changes into genomic DNA; second, nucleotide changes to Kozak sequences can be targeted to the HI gene by inducing a small, controlled increase in translation efficiency, which is crucial because the ultimate goal is to achieve protein levels that are sufficient to restore the HI phenotype but not so high that they are unbalanced and incompatible with cellular physiology; third, acting on cis elements that universally control translation can produce the desired effect in all cell types in the body, regardless of variability in gene expression, reducing noise in the system; fourth, multiple different Kozak variants can be tested for each gene, ensuring flexibility; and finally, because this approach does not rely on correcting a specific disease, any disease-causing mutant HI gene could, in principle, be restored in any patient.
[0013] Further objects, features, preferred embodiments and advantages of the present invention are described below, and the scope of the present invention is defined by the appended claims. [Brief explanation of the drawings]
[0014] [Figure 1]Enhancement of suboptimal Kozak sequences by base editing is shown. A. Sanger sequencing chromatograms representing wild-type (EGFP-1C) and mutant EGFP (EGFP-1T) with a single mutation at position -1 of the Kozak sequence. B. Western blot analysis of EGFP and mCherry expression in HEK293T cells transiently transfected with EGFP-1C or EGFP-1T plasmids. C. Representative FACS dot plots of HEK293T cells 3 days after transient transfection. D. FACS analysis of HEK293T cells transiently transfected with each plasmid. Data are normalized to EGFP-1C and shown as the mean ± standard deviation of n=3 biological replicates. Statistical significance was calculated by unpaired t-test. E. Representative Sanger sequencing chromatograms of HEK293T cells edited with the ABE7.10 base editor and sg-1. Comparison of ABE7.10 in combination with scrambled sgRNA (sg control). F. Percentage of correct T→C conversions analyzed using EditR software. G. Western blot analysis of EGFP and mCherry expression in HEK293T cells edited with ABE7.10 or ABEmax in combination with sg-1 or sg control. H. Representative FACS dot plot of cells edited with ABE7.10 and sg-1 3 days after transfection. Comparison of ABE7.10 in combination with scrambled sgRNA (sg control). I. FACS analysis of EGFP expression in cells transfected with base editors (ABE7.10 and ABEmax) in combination with sg control or sg-1. Data are means ± standard deviations from n=3 biological replicates. Statistically significant differences were calculated by unpaired t-test (p=0.0483). [Figure 2]This refers to high-throughput measurement of protein levels from Kozak sequence variants. A. mCherry expression in transduced cells in the first round of FACS-seq sorting. 5 x 106 mCherry-positive cells (23.1% of the total) were sorted. B. Second round of FACS-seq sorting. mCherry-positive cells from the gate set in C were divided into four gates according to EGFP / mCherry expression. Gates were set so that each bin contained 25% of the total population of interest. D. Heatmap showing the distribution of candidate HI genes and variants that passed statistical analysis. The top panel shows Kozak variants. The bottom panel shows the wild-type Kozak sequence of the HI gene. Each column corresponds to one of the four gates, while each row represents one of the Kozak variants. E. Logo representation of Kozak sequences extracted from each of the four gates. In each panel, the position along the Kozak sequence (A of ATG is the +1 position) is shown on the x-axis, and the probability of occurrence of each base is shown on the y-axis. Gate 1 (top panel) represents the lowest translation efficiency, while gate 4 (bottom panel) corresponds to the most functional Kozak sequence. Relevant positions (-3 and +5) are highlighted in yellow. F. Percentage of counts per million reads (CPM) for the wild type (WT) and each variant (Var) of the five selected genes across the four gates. [Figure 3] Validation of exploitable hit variants. A. Kozak sequences of wild-type (WT) and variant (Var) genes for selected hit genes. B. Translational enhancement analyzed by high-content image analysis as EGFP / mCherry expression. Violin plots show data distribution from n=3 biological replicates. The dashed line indicates the population median. C. Histograms show the population mean values analyzed by high-content image analysis. Data are means ± standard deviation from n=3 biological replicates. Numbers indicate the average increase of variants relative to WT. Statistically significant differences were calculated using an unpaired t-test between each variant and the corresponding WT. [Figure 4]Validation of exploitable hit variants for translational repression of PMP22. A. Kozak sequences of wild-type (WT) and variant (Var) PMP22 genes. B. Translational repression analyzed by high-content image analysis as EGFP / mCherry expression. Violin plots show data distribution from n=2, 3 biological replicates. Dashed lines indicate population medians. C. Histograms represent population means analyzed by high-content image analysis. Data are means ± standard deviations from n=2, 3 biological replicates. Statistically significant differences were calculated using unpaired t-tests between each variant and WT. [Figure 5]This refers to base editing of NCF1 to recapitulate the desired variant. A. Schematic of the Kozak sequences of NCF1 wild-type (WT), variant 2 (Var 2), and variant 4 (Var 4). The start codon is bolded blue; base changes in the variants are highlighted in pink. B. Editing efficiency at the target and bystander (red) guanine in the Raji bulk population 5 days after electroporation of AncBE4max and sgNCF1 or sg control was analyzed using EditR software. The percentage of correct G-to-A conversions (y-axis) at each position within the NCF1 Kozak sequence (x-axis, A of ATG is the +1 position) is shown. Data are means ± standard deviations from three independent experiments. C. Editing efficiency at the target and bystander (red) guanine in two clones (Var 2 and Var 4 cells) isolated from the bulk population. D. Chromatogram of Sanger sequencing of NCF1 Kozak sequences in Raji WT, Var 2, and Var 4 cells. E. Western blot analysis of p47phox protein in Raji cells (WT, Var 2, and Var 4). A representative blot result is shown. The arrow indicates the 47 KDa band corresponding to p47phox. F. Quantitation by Western blot. p47phox levels were normalized to a housekeeping protein, and fold changes relative to WT levels are shown. n = 3 biological replicates. G. qPCR of NCF1 in WT, Var 2, or Var 4 Raji cells. Data are means ± standard deviations from n = 3 independent experiments. H. Representative Western blot of two polysome markers (RPS6 and RPL26) in fractions isolated by sucrose gradient centrifugation. The input sample was cytoplasmic lysate loaded onto a sucrose gradient. tot = fraction corresponding to total RNA; pol = fraction selected as polysomes and used in I. I. Quantification of translation efficiency (TE) of NCF1 in Var 2 and Var 4 cells relative to WT cells. TE is the ratio between polysome mRNA levels (fractions 8–9) and total mRNA levels (fractions 4–9) (polysome fold change / total fold change), measured by qPCR.Data are means ± standard deviations from n = 3 independent experiments. Statistically significant differences were calculated by unpaired t-test between each variant and the WT. DETAILED DESCRIPTION OF THE INVENTION
[0015] With regard to the scope of the present invention, some terms and phrases used in this specification and the appended claims are defined below.
[0016] As used herein, the term "CRISPR" refers to a Clustered Regularly Interspaced Short Palindromic Repeat system or locus, i.e., a bacterial adaptive immune system that defends against invading mobile genetic elements. The term "Cas" (CRISPR-associated protein) refers to an RNA-guided programmable endonuclease that recognizes protospacer adjacent motifs (PAM sequences) and cleaves invading nucleic acids in the region complementary to the sequence encoded by the spacer (encoded in the CRISPR array). The term "Cas9n" refers to a partially inactive Cas9 endonuclease. Reference: Jinek, M., Chylinski, K., Fonfara, I., Hauer, M., Doudna, JA, & Charpentier, E. (2012). "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity". Science, 337(6096), 816-821.
[0017] As used herein, the term "base editor" refers to a tool that combines a deaminase enzyme with Cas9n and can perform a single-base conversion. There are three types of base editors: cytosine base editors (CBEs), which can convert CG to TA base pairs; adenine base editors (ABEs), which can convert AT to GC base pairs; and adenine to cytosine base editors, which can convert AT to CG. Reference: Komor, AC, Kim, YB, Packer, MS, Zuris, JA, & Liu, DR (2016). "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage". Nature, 533(7603), 420-424. Chen L, Hong M, Luan C, Gao H, Ru G, Guo X, Zhang D, Zhang S, Li C, Wu J, Randolph PB, Sousa AA, Qu C, Zhu Y, Guan Y, Wang L, Liu M, Feng B, Song G, Liu DR, Li D., Nat Biotechnol. 2023 Jun 15.
[0018] As intended herein, the term "gRNA" (single guide RNA) refers to an RNA molecule that acts as a guide to direct a Cas9 protein, or base editor, or prime editor, to a locus of interest (which is complementary to the spacer sequence of the gRNA and is adjacent to a PAM motif that is present within 30 nucleotides upstream or downstream of the ATG start codon).
[0019] As intended herein, the term "haploinsufficiency" refers to a molecular mechanism by which mutational inactivation of one allele of a gene is sufficient to produce a disease phenotype. "Haploinsufficient diseases" are disorders caused by such a process and characterized by insufficient amounts of a particular protein. The term "gene duplication disease" instead refers to a molecular mechanism by which gene duplication usually results in increased amounts of a particular protein, which in turn causes disease.
[0020] As used herein, the term "homology-directed repair" (HDR) refers to a repair pathway induced by double-strand breaks in DNA. Unlike non-homologous end joining repair, HDR allows for precise insertion of edits into DNA by providing a donor DNA molecule (used as a template for DNA repair) encoding the desired edit. In genome editing applications, HDR is used to insert genetic information encoded by the donor DNA after Cas9 cleaves the target locus. Reference: Doudna, JA, & Charpentier, E. (2014). "The new frontier of genome engineering with CRISPR-Cas9". Science, 346(6213), 1258096.
[0021] As intended herein, the phrase "Kozak consensus sequence" or "Kozak sequence" refers to a DNA sequence motif that serves as the translation initiation site in most eukaryotic mRNAs, which mediates ribosome reading of the AUG start codon and ensures that translation begins at the correct site on the mRNA.
[0022] As intended herein, the phrase "Kozak variant sequence" or "Kozak variant" refers to an alternative Kozak sequence designed by substituting some of the four nucleotides flanking the ATG codon of the wild-type Kozak sequence, where the ATG codon is kept constant (i.e., NNNN ATG NNNN). A variant may contain multiple alterations of the same type.
[0023] As used herein, the term "prime editor" refers to a complex characterized by Cas9n fused to reverse transcriptase. This complex can write new genetic information into a target locus using a programmable pegRNA (prime editing guide RNA), which encodes both the target site and a template containing the edit to be inserted. Reference: Anzalone, AV, Randolph, PB, Davis, JR, Sousa, AA, Koblan, LW, Levy, JM, ... & Liu, DR (2019). "Search-and-replace genome editing without double-strand breaks or donor DNA." Nature, 576(7785), 149-157.
[0024] As intended herein, the terms "translational modulator," "translational enhancer," or "translational repressor" refer to cis- and trans-elements capable of regulating, promoting, or suppressing the translation efficiency of a given mRNA, respectively. More specifically, in our context, this term refers to variant Kozak sequences with the same ability.
[0025] As intended herein, the term "vector" refers to a nucleic acid that is capable of entering, mutating, and replicating within a host cell, and then transferring the replicated form of the vector to another host cell.
[0026] Advantageously, in one embodiment of the present invention, a base editor is used to mutate the Kozak sequence. Furthermore, in one embodiment of the present invention, the targeted disease gene is an HI disease gene. In such cases, a change of one or several nucleotides in the Kozak sequence must be compatible with the action of the base editor and increase the amount of the encoded protein, thereby compensating for the deleterious effects of the loss of function of one allele. To this end, in the present invention, a gRNA is appropriately selected so that the base editor modifies one of several nucleotides that induce translational enhancement relative to the wild-type Kozak sequence. Thus, the present invention relates to a Kozak variant nucleotide sequence selected from SEQ ID NO: 1 to SEQ ID NO: 58 for use in treating HI disease.
[0027] According to a preferred embodiment, the HI disease is selected from the following disease areas: developmental disorders, metabolic syndromes, eye diseases, and hematopoietic diseases.
[0028] For each of the above HI disease regions, the corresponding specific sequences and further relevant details are described below.
[0029] According to a preferred embodiment, when the HI disease belongs to the category of developmental disorders, the nucleotide sequence is selected from SEQ ID NO: 1 to SEQ ID NO: 28. The HI disease gene (when mutated) causing the developmental disorder is selected from the following: DYRK1A, GRIN2B, NRXN1, and STXBP1. The developmental disorder disease is selected from the following: intellectual developmental disorder, autosomal dominant 7 (OMIM #614104) (caused by loss of function of one allele of the DYRK1A gene); intellectual developmental disorder, autosomal dominant 6, with or without seizures (OMIM #613970) (caused by loss of function of one allele of the GRIN2B gene); chromosome 2p16.3 deletion syndrome (OMIM #614332) (caused by loss of function of one allele of the NRXN1 gene); and developmental epileptic encephalopathy 4 (OMIM #612164) (caused by loss of function of one allele of the STXBP1 gene).
[0030] When the developmental disorder is intellectual developmental disorder, autosomal dominant7, the nucleotide sequence is selected from SEQ ID NO: 1 to SEQ ID NO: 3. When the developmental disorder is intellectual developmental disorder, autosomal dominant6, with or without seizures, the nucleotide sequence is selected from SEQ ID NO: 4 to SEQ ID NO: 8. When the developmental disorder is chromosome 2p16.3 deletion syndrome, the nucleotide sequence is selected from SEQ ID NO: 9 to SEQ ID NO: 14. When the developmental disorder is developmental epileptic encephalopathy4, the nucleotide sequence is selected from SEQ ID NO: 15 to SEQ ID NO: 28.
[0031] According to another preferred embodiment, when the HI disease belongs to the classification of metabolic syndrome, the nucleotide sequence is selected from SEQ ID NO: 29 to SEQ ID NO: 41. The HI disease genes causing metabolic syndrome are selected from the following: HNF1A, GHRL, and PROX1. The metabolic syndrome is selected from the following: maturity-onset diabetes of the young type 3 (OMIM#600496) (caused by loss of function of one allele of the HNF1A gene); susceptibility to obesity (OMIM#601665) (caused by loss of function of one allele of the GHRL gene); adult-onset obesity and lymphatic disease (caused by loss of function of one allele of the PROX1 gene).
[0032] When the metabolic syndrome disorder is maturity-onset diabetes of the young type 3, the nucleotide sequence is selected from SEQ ID NO: 29 to SEQ ID NO: 33. When the metabolic syndrome disorder is obesity susceptibility, the nucleotide sequence is selected from SEQ ID NO: 34 to SEQ ID NO: 38. When the metabolic syndrome disorder is adult-onset obesity and lymphatic vascular disease, the nucleotide sequence is selected from SEQ ID NO: 39 to SEQ ID NO: 41.
[0033] According to another preferred embodiment, when the HI disease belongs to the category of eye diseases, the nucleotide sequence is selected from SEQ ID NOs: 42 to 53. The HI disease gene causing the eye disease is selected from the following: EYA1, OPA1, and COL2A1. The eye disease is selected from the following: Branchio-Oto-Renal Syndrome 1 (OMIM#113650) (caused by loss of function of one allele of the EYA1 gene); Optic Atrophy 1 (OMIM#165500) (caused by loss of function of one allele of the OPA1 gene); Stickler Syndrome Type 1, Nonsyndromic Ocular (OMIM#609508) (caused by loss of function of one allele of the COL2A1 gene).
[0034] When the eye disease is branchio-oto-renal syndrome 1, the nucleotide sequence is selected from SEQ ID NO: 42 or SEQ ID NO: 43. When the eye disease is optic atrophy 1, the nucleotide sequence is selected from SEQ ID NO: 44 to SEQ ID NO: 50. When the eye disease is Stickler syndrome type 1, non-syndromic ophthalmic type, the nucleotide sequence is selected from SEQ ID NO: 51 to SEQ ID NO: 53.
[0035] According to another preferred embodiment, when the HI disease is a hematopoietic disease, the nucleotide sequence is selected from SEQ ID NOs: 54 to 58. The HI gene involved in the hematopoietic disease is NCF1. The hematopoietic disease is chronic granulomatous disease (OMIM#233700).
[0036] In another embodiment of the present invention, the targeted disease gene is a GD disease gene. In such cases, a change of one or a few nucleotides in the Kozak sequence must be compatible with the action of the base editor and reduce the amount of encoded protein, thereby compensating for the deleterious effect of the gain-of-function of one allele. To this end, in the present invention, a gRNA is appropriately selected to induce the base editor to modify one of a few nucleotides relative to the wild-type Kozak sequence to obtain translational repression. Thus, the present invention relates to a Kozak variant nucleotide sequence selected from SEQ ID NOs: 61 to 65 for use in treating HI disease.
[0037] According to a preferred embodiment, the HI disease is selected from the following disease areas: developmental disorders, metabolic syndromes, eye diseases, and hematopoietic diseases.
[0038] For each of the above HI disease regions, the corresponding specific sequences and further relevant details are described below.
[0039] According to a preferred embodiment, when the HI disease belongs to the category of developmental disorders, the nucleotide sequence is selected from SEQ ID NO: 1 to SEQ ID NO: 28. The HI disease gene (when mutated) causing the developmental disorder is selected from the following: DYRK1A, GRIN2B, NRXN1, and STXBP1. The developmental disorder disease is selected from the following: intellectual developmental disorder, autosomal dominant 7 (OMIM #614104) (caused by loss of function of one allele of the DYRK1A gene); intellectual developmental disorder, autosomal dominant 6, with or without seizures (OMIM #613970) (caused by loss of function of one allele of the GRIN2B gene); chromosome 2p16.3 deletion syndrome (OMIM #614332) (caused by loss of function of one allele of the NRXN1 gene); and developmental epileptic encephalopathy 4 (OMIM #612164) (caused by loss of function of one allele of the STXBP1 gene).
[0040] Furthermore, the present invention also relates to vectors suitable for genome editing, which contain any gRNA designed to edit a wild-type Kozak sequence to obtain any of the variant Kozak nucleotide sequences selected from the group of SEQ ID NO: 1 to SEQ ID NO: 58. These gRNAs are such that their targeting sequence corresponds to a target domain adjacent to a PAM sequence located within 30 nucleotides upstream or downstream of the ATG start codon embedded within the Kozak sequence.
[0041] According to a preferred embodiment, when the disease is chronic granulomatous disease, the gRNA that edits the wild-type Kozak sequence of NCF1 is encoded by SEQ ID NO: 59.
[0042] According to a preferred embodiment, when the disease is chronic optic atrophy, the gRNA that edits the wild-type Kozak sequence of OPA1 is encoded by SEQ ID NO: 60.
[0043] Furthermore, the present invention also relates to a pharmaceutical composition comprising at least one gRNA (designed to obtain any one of the variant Kozak sequences selected from SEQ ID NO:1 to SEQ ID NO:58), a vector suitable for genome editing, and a genome editing complex (selected from those required for CRISPR-Cas homology-directed repair, CRISPR-Cas primed editing, CRISPR-Cas base editing, or any other method that allows the introduction of one or more nucleotide alterations based on a programmable RNA-guided nuclease fused to an effector protein), and at least one pharmaceutically acceptable excipient. According to a preferred embodiment, the pharmaceutical composition is intravenously administrable.
[0044] An object of the present invention is therefore to provide a variant Kozak nucleotide sequence obtained by genome editing methods, which acts as a translation regulator of a gene encoding a protein in the treatment of a disease in which the expression of said protein is altered, wherein: the variant Kozak nucleotide sequence replaces the wild-type Kozak sequence in vitro or in vivo; Variant Kozak nucleotide sequence.
[0045] Preferably, the subject of the present invention is a variant Kozak nucleotide sequence according to claim 1, The genome editing method is selected from the group consisting of: CRISPR-Cas homology-directed repair, CRISPR-Cas primed editing, CRISPR-Cas base editing, or any other method that allows the introduction of one or more nucleotide alterations based on a programmable RNA-guided nuclease fused to an effector protein; Variant Kozak nucleotide sequence.
[0046] According to one embodiment, the variant Kozak nucleotide sequence described above comprises: The translation regulatory factor is a translation promoting factor or a translation repressing factor. Variant Kozak nucleotide sequences are preferred.
[0047] According to one embodiment, the variant Kozak nucleotide sequence described above comprises: The translation-enhancing factor is selected from the group consisting of SEQ ID NO: 1 to SEQ ID NO: 58, the variant Kozak nucleotide sequence replaces the wild-type Kozak sequence in vitro or in vivo; Variant Kozak nucleotide sequences are preferred.
[0048] Another object is the use of the variant Kozak nucleotide sequence described above, in Use of the compound in the treatment of haploinsufficiency disorders to improve the translation efficiency of a protein-encoding gene.
[0049] According to a preferred embodiment, the use of the variant Kozak nucleotide sequence described above, Preferred uses are those in which the haploinsufficiency disorder is selected from the following disease categories: developmental disorders, metabolic syndromes, eye disorders, and hematopoietic disorders.
[0050] According to one embodiment, there is provided a use of the variant Kozak nucleotide sequence described above, comprising: Preferably, the developmental disorder is selected from: intellectual developmental disorder, autosomal dominant7; intellectual developmental disorder6, with or without seizures; 2p16.3 deletion syndrome; developmental epileptic encephalopathy4.
[0051] According to one embodiment, there is provided the use of a variant Kozak nucleotide sequence according to any one of claims 6 to 7, wherein: When the developmental disorder is intellectual developmental disorder, autosomal dominant 7, the nucleotide sequence is selected from SEQ ID NO: 1 to SEQ ID NO: 3; When the developmental disorder is intellectual developmental disorder, autosomal dominant, with or without seizures, the nucleotide sequence is selected from SEQ ID NO: 4 to SEQ ID NO: 8; When the developmental disorder is chromosome 2p16.3 deletion syndrome, the nucleotide sequence is selected from SEQ ID NO: 9 to SEQ ID NO: 14; When the developmental disorder is developmental epileptic encephalopathy 4, the nucleotide sequence is preferably selected from SEQ ID NO: 15 to SEQ ID NO: 28.
[0052] According to one embodiment, there is provided a use of the variant Kozak nucleotide sequence described above, comprising: Preferred is the use, wherein said metabolic syndrome is selected from: maturity-onset diabetes of the young type 3; obesity susceptibility; lymphatic vascular defects and / or adult-onset obesity.
[0053] According to one embodiment, there is provided the use of the variant Kozak nucleotide sequence described above, wherein: when the metabolic syndrome is maturity-onset diabetes of the young type 3, the nucleotide sequence is selected from SEQ ID NO: 29 to SEQ ID NO: 33; when the metabolic syndrome is obesity susceptibility, the nucleotide sequence is selected from SEQ ID NO: 34 to SEQ ID NO: 38; When the metabolic syndrome is lymphatic vascular abnormality and / or adult-onset obesity, the nucleotide sequence is preferably selected from SEQ ID NO: 39 to SEQ ID NO: 41.
[0054] According to one embodiment, there is provided a use of the variant Kozak nucleotide sequence described above, comprising: Preferably, the eye disease is selected from branchio-oto-renal syndrome, optic atrophy, Stickler syndrome type 1, non-syndromic ophthalmic type.
[0055] Use of the variant Kozak nucleotide sequence described above, wherein: When the eye disease is branchio-oto-renal syndrome, the nucleotide sequence is selected from SEQ ID NO: 42 or SEQ ID NO: 43; When the eye disease is optic atrophy, the nucleotide sequence is selected from SEQ ID NO: 44 to SEQ ID NO: 50; When the eye disease is Stickler syndrome type 1, non-syndromic eye type, the nucleotide sequence is selected from SEQ ID NO: 51 to SEQ ID NO: 53.
[0056] According to one embodiment, there is provided a use of the variant Kozak nucleotide sequence described above, comprising: When the hematopoietic disease is chronic granulomatous disease, the nucleotide sequence is preferably selected from SEQ ID NO: 54 to SEQ ID NO: 58.
[0057] Another object is to provide a gRNA designed to edit a wild-type Kozak sequence to obtain any one of variant Kozak nucleotide sequences selected from the group of SEQ ID NO: 1 to SEQ ID NO: 58, The targeting sequence corresponds to a targeting domain adjacent to a PAM sequence located within 30 nucleotides upstream or downstream of the ATG start codon, gRNA.
[0058] According to one embodiment, the gRNA described above, The gRNA edits the wild-type Kozak sequence of NCF1, The gRNA has the nucleotide sequence of SEQ ID NO: 59. gRNA is preferred.
[0059] According to another embodiment, the gRNA described above, the gRNA edits the wild-type Kozak sequence of OPA1; The gRNA has the nucleotide sequence of SEQ ID NO: 60. gRNA is preferred.
[0060] Another object is a vector for genome editing, comprising any one of the gRNAs described above.
[0061] Another object is a pharmaceutical composition comprising the vector described above together with other components suitable for in vivo genome editing.
[0062] According to certain embodiments, the above-mentioned pharmaceutical compositions are preferably those that can be administered intravenously. [Example]
[0063] [Experimental part] [Base editing-mediated Kozak optimization enhances translation in reporter systems] To demonstrate the feasibility of the proposed method, experiments were conducted focusing on enhancing EGFP translation from a reporter vector. We constructed two bicistronic reporter vectors, pWPT-EGFP-IRESmCherry: one with the EGFP wild-type Kozak sequence (C at position -1, EGFP-1C) and one with a suboptimal motif with a T at position -1 (EGFP-1T) (Figure 1A). This single base change resulted in a 4- to 5-fold decrease in EGFP translation (Figure 1B, C, D). We then corrected this nucleotide mutation with a base editor and observed a significant increase in EGFP expression (Figure 1E-I). These data confirm that base editors can be used to selectively mutate a single nucleotide within a Kozak sequence.
[0064] [Design and construction of a library of exploitable Kozak variants] We screened the wild-type (WT) Kozak sequences of annotated HI genes and compared them with their respective variants to identify a specific set of exploitable changes that upregulated the translation efficiency of each WT Kozak sequence. We started with 230 haploinsufficient genes and generated an unbiased library of Kozak mutants from them. We obtained 5,539 variants, of which 4,838 were unique. As the destination vector, we used pWPT-EGFP-IRESmCherry. After recombination, the library of wild-type and variant Kozak sequences replaced the Kozak sequence of EGFP, directing EGFP expression. The resulting reporter containing the library was used to transduce HEK293T cells.
[0065] [Assessment of protein levels from Kozak sequences of wild-type and variants] To quantify the translation efficiency of the Kozak sequences of wild-type and variant HI genes, we sorted HEK293T cells transduced with the reporter library into four gates according to the normalized EGFP translation efficiency (EGFP / mCherry). In the first round, 5 x 10 mCherry-positive cells were selected to ensure 1000x coverage of the library (Figure 2A). In the second round, the resulting mCherry-positive cells were sorted into four bins with different fluorescence intensity ratios according to their EGFP / mCherry ratios (Figure 2B, C). The Kozak sequence regions from cells collected in each bin were PCR amplified. Deep sequencing of all fractions allowed us to compare the intensity of each HI wild-type Kozak sequence with its variants. 89 wild-type sequences and 403 variant sequences passed statistical analysis (Figure 2D). Next, we created motifs representing the nucleotide frequency at each position in the Kozak sequence for each of the four gates (Figure 2E). Aiming to select Kozak variants that upregulate their corresponding WT, we selected only variants with the greatest distance from the respective WT. We obtained 47 wild-type sequences and 149 variant sequences. From this list, we selected five HI genes and their corresponding variants for validation: PPARGC1B, FKBP6, GALR1, NRXN1, and NCF1 (Figure 2F).
[0066] [Validation of protein upregulation by selected hit Kozak sequence variants] To validate the selected hits, we cloned each Kozak sequence (wild-type and hit variants of the five selected genes) into a reporter vector in place of the Kozak sequence of EGFP, creating one new plasmid for each sequence. We transiently transfected HEK293T cells with each wild-type and hit Kozak variant and measured fluorescence by high-content image analysis 3 days after transfection (Figure 3). These analyses confirmed that 10 of the 11 Kozak variants tested increased translation efficiency compared to their respective wild-type sequences (Figure 3B, C).
[0067] [Enhancing NCF1 translation by base editing of Kozak sequences] We then recreated two variants that emerged from the screen (Var 2: SEQ ID NO: 55; Var 4: SEQ ID NO: 57) by base editing the endogenous locus in Raji cells (a Burkitt lymphoma-derived B lymphocyte cell line that constitutively expresses the gene of interest). We performed base editing by electroporating AncBE4max and the guide RNA sgNCF1 (SEQ ID NO: 59) (Figure 4). To improve the readout of the edits, we next decided to isolate cell clones. We identified and expanded clones that carried the desired base editor-mediated nucleotide changes and corresponded to Kozak NCF1 variants 2 and 4 (SEQ ID NO: 55 and SEQ ID NO: 57). Western blot analysis revealed that the protein encoded by NCF1, p47 phoxWe found that expression of NCF1 was increased in both variants compared to the wild type (Figure 4E, F). We also analyzed NCF1 mRNA levels in wild-type cells and clones and found that they were unchanged. These results strongly support the idea that increased gene expression is due to enhanced NCF1 translation by Kozak sequence editing (Figure 4G). Sucrose gradient fractionation in WT and edited cells showed that increased protein levels corresponded to increased mRNA loading into polysomes (Figure 4H, I). Taken together, these results demonstrated that this is a novel gene editing method targeting the Kozak sequence of a gene. The appropriate variant was introduced through base editing, resulting in translational upregulation of the target gene.
[0068] Table 1 below shows the variant Kozak sequences of the present invention (SEQ ID NOs: 1-58).
[0069] Table 2 below shows the gRNAs of the present invention (SEQ ID NOs: 59 and 60).
[0070] Table 3 below shows the variant Kozak of the present invention (SEQ ID NOs: 61-65).
[0071] JPEG2025525882000002.jpg203166JPEG2025525882000003.jpg209166
[0072] JPEG2025525882000004.jpg42166
[0073] JPEG2025525882000005.jpg69166
Claims
1. A variant Kozak nucleotide sequence obtained by genome editing that acts as a translation regulator of a gene encoding a protein in the treatment of a disease in which the expression of said protein is altered, the variant Kozak nucleotide sequence replaces the wild-type Kozak sequence in vitro or in vivo; Variant Kozak nucleotide sequence.
2. 2. A variant Kozak nucleotide sequence according to claim 1, The genome editing method is selected from the group consisting of: CRISPR-Cas homology-directed repair, CRISPR-Cas primed editing, CRISPR-Cas base editing, or any other method that allows the introduction of one or more nucleotide alterations based on a programmable RNA-guided nuclease fused to an effector protein. Selected from: Variant Kozak nucleotide sequence.
3. A variant Kozak nucleotide sequence according to any one of claims 1 to 2, The translation regulatory factor is a translation promoting factor or a translation repressing factor. Variant Kozak nucleotide sequence.
4. 4. A variant Kozak nucleotide sequence according to claim 3, The translation-enhancing factor is selected from the group consisting of SEQ ID NO: 1 to SEQ ID NO: 58; or The translation repressor is selected from the group consisting of SEQ ID NO: 61 to SEQ ID NO: 65; the variant Kozak nucleotide sequence replaces the wild-type Kozak sequence in vitro or in vivo; Variant Kozak nucleotide sequence.
5. 5. Use of a variant Kozak nucleotide sequence according to claim 4, comprising: To improve the translation efficiency of protein-coding genes in the treatment of haploinsufficiency disorders. use.
6. 6. Use of a variant Kozak nucleotide sequence according to claim 5, The haploinsufficiency disorder is selected from the following disease categories: developmental disorders, metabolic syndromes, eye diseases, and hematopoietic disorders. use.
7. 7. Use of a variant Kozak nucleotide sequence according to claim 6, comprising: The developmental disorder is selected from: intellectual developmental disorder, autosomal dominant 7; intellectual developmental disorder 6, with or without seizures; 2p16.3 deletion syndrome; developmental epileptic encephalopathy 4. use.
8. Use of a variant Kozak nucleotide sequence according to any one of claims 6 to 7, Where: - when the developmental disorder is intellectual developmental disorder, autosomal dominant 7, the nucleotide sequence is selected from SEQ ID NO: 1 to SEQ ID NO: 3; - when the developmental disorder is intellectual developmental disorder, autosomal dominant 6, with or without seizures, the nucleotide sequence is selected from SEQ ID NO: 4 to SEQ ID NO: 8; - when the developmental disorder is chromosome 2p16.3 deletion syndrome, the nucleotide sequence is selected from SEQ ID NO: 9 to SEQ ID NO: 14; - when the developmental disorder is developmental epileptic encephalopathy 4, the nucleotide sequence is selected from SEQ ID NO: 15 to SEQ ID NO: 28; use.
9. 7. Use of a variant Kozak nucleotide sequence according to claim 6, comprising: The metabolic syndrome is selected from: maturity-onset diabetes of the young type 3; susceptibility to obesity; lymphatic vascular defects and / or adult-onset obesity, use.
10. Use of a variant Kozak nucleotide sequence according to any one of claims 6 and 9, comprising: Where: - when the metabolic syndrome is maturity-onset diabetes of the young type 3, the nucleotide sequence is selected from SEQ ID NO: 29 to SEQ ID NO: 33; - when the metabolic syndrome is obesity susceptibility, the nucleotide sequence is selected from SEQ ID NO: 34 to SEQ ID NO: 38; - when the metabolic syndrome is lymphatic vascular abnormalities and / or adult-onset obesity, the nucleotide sequence is selected from SEQ ID NO: 39 to SEQ ID NO: 41; use.
11. 7. Use of a variant Kozak nucleotide sequence according to claim 6, comprising: The eye disease is selected from branchio-oto-renal syndrome, optic atrophy, Stickler syndrome type 1, and nonsyndromic ocular. use.
12. Use of a variant Kozak nucleotide sequence according to any one of claims 6 and 11, comprising: Where: - if the eye disease is branchio-oto-renal syndrome, the nucleotide sequence is selected from SEQ ID NO: 42 or SEQ ID NO: 43; - if the eye disease is optic atrophy, the nucleotide sequence is selected from SEQ ID NO: 44 to SEQ ID NO: 50; - if the eye disease is Stickler syndrome type 1, non-syndromic ocular type, the nucleotide sequence is selected from SEQ ID NO: 51 to SEQ ID NO: 53; use.
13. 7. Use of a variant Kozak nucleotide sequence according to claim 6, comprising: When the hematopoietic disease is chronic granulomatous disease, the nucleotide sequence is selected from SEQ ID NO: 54 to SEQ ID NO:
58. use.
14. A gRNA designed to edit a wild-type Kozak sequence to obtain any one of variant Kozak nucleotide sequences selected from the group of SEQ ID NO: 1 to SEQ ID NO: 58, The targeting sequence corresponds to a targeting domain adjacent to a PAM sequence located within 30 nucleotides upstream or downstream of the ATG start codon; gRNA.
15. 15. The gRNA of claim 14, The gRNA edits the wild-type Kozak sequence of NCF1; The gRNA has the nucleotide sequence of SEQ ID NO:
59. gRNA.
16. 15. The gRNA of claim 14, The gRNA edits the wild-type Kozak sequence of OPA1; The gRNA has the nucleotide sequence of SEQ ID NO:
60. gRNA.
17. A vector for genome editing comprising any one of the gRNAs described in any one of claims 14 to 16.
18. A pharmaceutical composition comprising the vector of claim 17 together with other components suitable for in vivo genome editing.
19. 19. The pharmaceutical composition of claim 18, It can be administered intravenously, Pharmaceutical compositions.
20. A variant Kozak nucleotide sequence according to any one of claims 1 to 4, For use in treating haploinsufficiency or gene duplication disorders, Variant Kozak nucleotide sequence.
21. 21. A variant Kozak nucleotide sequence for use according to claim 20, comprising: The haploinsufficiency disorder is selected from the group consisting of a developmental disorder, a metabolic syndrome, an eye disease, and a hematopoietic disorder. Variant Kozak nucleotide sequence.
22. 22. A variant Kozak nucleotide sequence for use according to claim 21, comprising: Where: - when the developmental disorder is intellectual developmental disorder, autosomal dominant 7, the nucleotide sequence is selected from SEQ ID NO: 1 to SEQ ID NO: 3; - when the developmental disorder is intellectual developmental disorder, autosomal dominant 6, with or without seizures, the nucleotide sequence is selected from SEQ ID NO: 4 to SEQ ID NO: 8; - when the developmental disorder is chromosome 2p16.3 deletion syndrome, the nucleotide sequence is selected from SEQ ID NO: 9 to SEQ ID NO: 14; - when the developmental disorder is developmental epileptic encephalopathy 4, the nucleotide sequence is selected from SEQ ID NO: 15 to SEQ ID NO: 28; Variant Kozak nucleotide sequence.
23. 22. A variant Kozak nucleotide sequence for use according to claim 21, comprising: Where: - when the metabolic syndrome is maturity-onset diabetes of the young type 3, the nucleotide sequence is selected from SEQ ID NO: 29 to SEQ ID NO: 33; - when the metabolic syndrome is obesity susceptibility, the nucleotide sequence is selected from SEQ ID NO: 34 to SEQ ID NO: 38; - when the metabolic syndrome is lymphatic vascular abnormalities and / or adult-onset obesity, the nucleotide sequence is selected from SEQ ID NO: 39 to SEQ ID NO: 41; Variant Kozak nucleotide sequence.
24. 22. A variant Kozak nucleotide sequence for use according to claim 21, comprising: Where: - if the eye disease is branchio-oto-renal syndrome, the nucleotide sequence is selected from SEQ ID NO: 42 or SEQ ID NO: 43; - if the eye disease is optic atrophy, the nucleotide sequence is selected from SEQ ID NO: 44 to SEQ ID NO: 50; - if the eye disease is Stickler syndrome type 1, non-syndromic ocular type, the nucleotide sequence is selected from SEQ ID NO: 51 to SEQ ID NO: 53; Variant Kozak nucleotide sequence.
25. 22. A variant Kozak nucleotide sequence for use according to claim 21, comprising: When the hematopoietic disease is chronic granulomatous disease, the nucleotide sequence is selected from SEQ ID NO: 54 to SEQ ID NO:
58. Variant Kozak nucleotide sequence.
26. 21. A variant Kozak nucleotide sequence for use according to claim 20, comprising: The gene duplication disease is selected from the group consisting of motor and sensory neuropathies. Variant Kozak nucleotide sequence.
27. 27. A variant Kozak nucleotide sequence for use according to claim 26, comprising: the motor and sensory neuropathy is Charcot-Marie-Tooth disease type 1A, and the nucleotide sequence is selected from SEQ ID NO:61 to SEQ ID NO:65; Variant Kozak nucleotide sequence.