HSD17B13 variants and uses thereof

The HSD17B13 rs72613567 variant gene and associated nucleic acids offer a means to develop therapeutic strategies for chronic liver diseases by altering gene expression and susceptibility, addressing the lack of effective treatments for alcoholic and non-alcoholic liver disease and cirrhosis.

JP7755636B2Active Publication Date: 2025-10-16REGENERON PHARMACEUTICALS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023217935
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-11-06
Filing Date
2023-12-25
Publication Date
2025-10-16
Estimated Expiration
2038-01-19

AI Technical Summary

Technical Problem

Current treatments for chronic liver diseases such as alcoholic and non-alcoholic liver disease and cirrhosis are lacking, despite advances in hepatitis C treatment, and there is a need for evidence-based therapeutic strategies given the high morbidity and mortality associated with these conditions.

Method used

The use of HSD17B13 rs72613567 variant gene, transcripts, and proteins, including specific nucleic acids and proteins, to develop methods for detecting and modifying the HSD17B13 gene, potentially altering its expression and susceptibility to chronic liver disease.

Benefits of technology

Provides a basis for novel therapeutic approaches by identifying protective genetic variants that could improve risk stratification and treatment of chronic liver diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007755636000037
    Figure 0007755636000037
  • Figure 0007755636000038
    Figure 0007755636000038
  • Figure 0007755636000039
    Figure 0007755636000039
Patent Text Reader

Abstract

To provide HSD17B13 variants and uses thereof.SOLUTION: Provided herein is an HSD17B13 variant discovered to be associated with: reduced alanine and aspartate transaminase levels; a reduced risk of chronic liver diseases; and reduced progression from simple steatosis to more clinically advanced stages of chronic liver disease. Also provided herein are isolated nucleic acids and proteins related to variants of HSD17B13, and cells comprising those nucleic acids and proteins. Also disclosed is a method for modifying a cell via the use of any combination of: a nuclease agent to express a recombinant HSD17B13 gene or a nucleic acid encoding an HSD17B13 protein; an exogenous donor sequence; a transcription activator; a transcription repressor; and an expression vector. Also disclosed is a therapeutic and prophylactic method for treating a subject having or at risk of developing chronic liver disease.SELECTED DRAWING: Figure 17
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Application No. 62 / 449,335, filed January 23, 2017, U.S. Application No. 62 / 472,972, filed March 17, 2017, and U.S. Application No. 62 / 581,918, filed November 6, 2017, the entirety of each of which is incorporated herein by reference for all purposes.

[0002] Reference to sequence listings submitted as text files via EFS WEB The sequence listing set forth in file 507242SEQLIST.txt is 507 kilobytes, was created on January 19, 2018, and is hereby incorporated by reference. [Background technology]

[0003] background Chronic liver disease and cirrhosis are leading causes of morbidity and mortality in the United States, accounting for 38,170 deaths (1.5% of all deaths) in 2014 (Kochanek et al. (2016) Natl Vital Stat Rep 65:1-122, incorporated herein by reference in its entirety for all purposes). The most common etiologies of cirrhosis in the United States are alcoholic liver disease, chronic hepatitis C, and nonalcoholic fatty liver disease (NAFLD), which together accounted for approximately 80% of patients awaiting liver transplantation between 2004 and 2013 (Wong et al. (2015) Gastroenterology 148:547-555, incorporated herein by reference in its entirety for all purposes). The estimated prevalence of NAFLD in the United States ranges from 19 to 46 percent (Browning et al. (2004) Hepatology 40:1387-1395; Lazo et al. (2013) Am J Epidemiol 178:38-45; and Williams et al. (2011) Gastroenterology 140:124-131, each of which is incorporated by reference in its entirety for all purposes), likely coupled with rising rates of obesity, its primary risk factor (Cohen et al. (2013)). 11) Science, Vol. 332: pp. 1519-1523, in its entirety for all purposes. The incidence of hepatitis C has increased over time (Younossi et al. (2011) Clin Gastroenterol Hepatol 9:524-530 e1; quiz e60 (2011), each of which is incorporated by reference in its entirety for all purposes). While there have been significant advances in the treatment of hepatitis C (Morgan et al. (2013) Ann Intern Med 158:329-337 and van der Meer et al. (2012) JAMA 308:2584-2593, each of which is incorporated by reference in its entirety for all purposes), there are currently no evidence-based treatments for alcoholic or non-alcoholic liver disease and cirrhosis.

[0004] Previous genome-wide association studies (GWAS) have identified a limited number of genes and variants associated with chronic liver disease. The most robustly validated genetic association to date involves a common missense variant in the patatin-like phospholipase domain-containing 3 gene (PNPLA3 p.Ile148Met, rs738409), which was initially found to be associated with an increased risk of nonalcoholic fatty liver disease (NAFLD) (Romeo et al. (2008) Nat. Genet. 40:1461-1466). 5 and Speliotes et al. (2011) PLoS Genet., 7:e1001324, (2009) J. Lipid Res. 50:2111-2116, each of which is incorporated by reference in its entirety for all purposes), followed by disease severity (Rotman et al. (2010) Hepatology 52:894-903 and Sookoian et al. (2009) J. Lipid Res. 50:2111-2116, respectively). (2016) J. Hepatol. doi:10.1016 / j.jhep.2016.03.011, which is incorporated by reference in its entirety for all purposes) and progression (Trepo et al. (2016) J. Hepatol. doi:10.1016 / j.jhep.2016.03.011, which is incorporated by reference in its entirety for all purposes). (The entire contents of which are incorporated herein by reference.) Variations in the transmembrane 6 superfamily member 2 (TM6SF2) gene have also been shown to confer an increased risk of NAFLD (Kozlitina et al. (2014) Nat. Genet. 46:352-356; Liu et al. (2014) Nat. Commun. 5:4309). (p. 113; and Sookoian et al. (2015) Hepatology 61:515-525, each of which is incorporated by reference in its entirety for all purposes). Although the normal functions of these two proteins are not fully understood, both have been proposed to be involved in hepatocyte lipid metabolism. How variants in PNPLA3 and TM6SF2 contribute to increased risk of liver disease remains to be elucidated. GWAS have also identified several genetic factors associated with serum alanine aminotransferase (ALT) and aspartate aminotransferase (AST), quantitative markers of hepatocellular injury and hepatic fat accumulation that are frequently measured clinically (Chambers et al. (2011) Nat. Genet. 43:131-1138 and Yuan et al. (2008) Am. J. Hum. Genet., 83:520-528, each of which is incorporated herein by reference in its entirety for all purposes. To date, no protective genetic variants have been described for chronic liver disease. The discovery of protective genetic variants in other contexts, such as loss-of-function variants of PCSK9 that reduce the risk of cardiovascular disease, has provided impetus for the development of new classes of therapeutic agents. Knowledge of the genetic factors underlying the development and progression of chronic liver disease could improve risk stratification and provide the basis for novel therapeutic strategies. To improve risk stratification and generate novel therapeutic approaches for liver disease, a better understanding of the underlying genetic factors is required. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Kochanek et al. (2016) Natl Vital Stat Rep 65:1-122 [Non-patent document 2] Wong et al. (2015) Gastroenterology, 148:547-555 [Non-patent document 3] Browning et al. (2004) Hepatology 40:1387-1395 [Non-patent document 4] Lazo et al. (2013) Am J Epidemiol 178:38-45 [Non-patent document 5] Williams et al. (2011) Gastroenterology 140:124-131 [Non-patent document 6] Cohen et al. (2011) Science 332:1519-1523 [Non-Patent Document 7] Younossi et al. (2011) Clin Gastroenterol Hepatol 9:524-530 [Non-patent document 8] Morgan et al. (2013) Ann Intern Med 158:329-337 [Non-Patent Document 9] van der Meer et al. (2012) JAMA 308:2584-2593 [Non-Patent Document 10] Romeo et al. (2008) Nat. Genet. 40:1461-1465 [Non-Patent Document 11] Speliotes et al. (2011) PLoS Genet., 7:e1001324 [Non-Patent Document 12] Rotman et al. (2010) Hepatology 52:894-903 [Non-Patent Document 13] Sookoian et al. (2009) J. Lipid Res., 50:2111-2116 [Non-Patent Document 14] Trepo et al. (2016) J. Hepatol. doi:10.1016 / j.jhep.2016.03.011 [Non-Patent Document 15] Kozlitina et al. (2014) Nat. Genet. 46:352-356 [Non-Patent Document 16] Liu et al. (2014) Nat. Commun. 5:4309 [Non-Patent Document 17] Sookoian et al. (2015) Hepatology, 61:515-525 [Non-Patent Document 18] Chambers et al. (2011) Nat. Genet. 43:131-1138 [Non-Patent Document 19] Yuan et al. (2008) Am. J. Hum. Genet. 83:520-528 Summary of the Invention [Means for solving the problem]

[0006] overview Methods and compositions relating to the HSD17B13 rs72613567 variant gene, variant HSD17B13 transcripts, and variant HSD17B13 protein isoforms are provided.

[0007] In one embodiment, an isolated nucleic acid is provided that comprises a mutant residue from the HSD17B13 rs72613567 variant gene. Such an isolated nucleic acid may comprise at least 15 contiguous nucleotides of the HSD17B13 gene, and when optimally aligned with SEQ ID NO: 1, a thymine is inserted between the nucleotide corresponding to position 12665 ​​and the nucleotide corresponding to position 12666 of SEQ ID NO: 1. Optionally, the contiguous nucleotides are at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the corresponding sequence of SEQ ID NO: 2, including position 12666 of SEQ ID NO: 2, when optimally aligned with SEQ ID NO: 2. Optionally, the HSD17B13 gene is a human HSD17B13 gene. Optionally, the isolated nucleic acid comprises at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least At least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10000, at least 11000, at least 12000, at least 13000, at least 14000, at least 15000, at least 16000, at least 17000, at least 18000, or at least 19000 consecutive nucleotides.

[0008] Some such isolated nucleic acids include an HSD17B13 minigene in which one or more non-essential segments of the gene are deleted relative to the corresponding wild-type HSD17B13 gene. Optionally, the deleted segments include one or more intron sequences. Optionally, the isolated nucleic acid further includes an intron corresponding to intron 6 of SEQ ID NO:2 when optimally aligned with SEQ ID NO:2. Optionally, the intron is intron 6 of SEQ ID NO:2.

[0009] In another aspect, isolated nucleic acids corresponding to different HSD17B13 mRNA transcripts or cDNAs are provided. Some such isolated nucleic acids comprise at least 15 contiguous nucleotides encoding all or part of the HSD17B13 protein, wherein the contiguous nucleic acid comprises a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment that is not present in SEQ ID NO:4 (HSD17B13 transcript A) but is present in SEQ ID NO:7 (HSD17B13 transcript D), SEQ ID NO:10 (HSD17B13 transcript G), and SEQ ID NO:11 (HSD17B13 transcript H). Optionally, the contiguous nucleotides further include a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment present in SEQ ID NO: 7 (HSD17B13 transcript D) that is not present in SEQ ID NO: 11 (HSD17B13 transcript H), and the contiguous nucleotides further include a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment present in SEQ ID NO: 7 (HSD17B13 transcript D) that is not present in SEQ ID NO: 10 (HSD17B13 transcript G). Optionally, the contiguous nucleotides further comprise a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment present in SEQ ID NO: 11 (HSD17B13 transcript H) that is not present in SEQ ID NO: 7 (HSD17B13 transcript D). Optionally, the contiguous nucleotides further comprise a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment present in SEQ ID NO: 10 (HSD17B13 transcript G) that is not present in SEQ ID NO: 7 (HSD17B13 transcript D).

[0010] Some such isolated nucleic acids comprise at least 15 contiguous nucleotides encoding all or a portion of an HSD17B13 protein, wherein the contiguous nucleotides include a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment present in SEQ ID NO:8 (HSD17B13 transcript E) that is not present in SEQ ID NO:4 (HSD17B13 transcript A). Optionally, the contiguous nucleotides further include a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment present in SEQ ID NO:8 (HSD17B13 transcript E) that is not present in SEQ ID NO:11 (HSD17B13 transcript H).

[0011] Some such isolated nucleic acids comprise at least 15 contiguous nucleotides encoding all or part of an HSD17B13 protein, wherein the contiguous nucleotides include a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment present in SEQ ID NO: 9 (HSD17B13 transcript F) that is not present in SEQ ID NO: 4 (HSD17B13 transcript A).

[0012] Some such isolated nucleic acids comprise at least 15 contiguous nucleotides encoding all or part of an HSD17B13 protein, wherein the contiguous nucleotides include a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment present in SEQ ID NO: 6 (HSD17B13 transcript C) that is not present in SEQ ID NO: 4 (HSD17B13 transcript A).

[0013] Optionally, the HSD17B13 protein is a human HSD17B13 protein. Optionally, the isolated nucleic acid comprises at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, or at least 2000 contiguous nucleotides encoding all or a portion of the HSD17B13 protein.

[0014] Some such isolated nucleic acids comprise a sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence set forth in SEQ ID NO: 6, 7, 8, 9, 10, or 11 (HSD17B13 transcript C, D, E, F, G, or H), and encode an HSD17B13 protein (HSD17B13 isoform C, D, E, F, G, or H) comprising the sequence set forth in SEQ ID NO: 14, 15, 16, 17, 18, or 19, respectively.

[0015] In any of the above nucleic acids, the contiguous nucleotides can optionally include sequences from at least two different exons of the HSD17B13 gene, without intervening introns.

[0016] In another aspect, there is provided a protein encoded by any of the above isolated nucleic acids.

[0017] In another embodiment, an isolated nucleic acid is provided that hybridizes to or approximates a mutant residue from the HSD17B13 rs72613567 variant gene. Such an isolated nucleic acid may comprise at least 15 contiguous nucleotides that hybridize to the HSD17B13 gene in a segment that includes or is within 1000, 500, 400, 300, 200, 100, 50, 45, 40, 35, 30, 25, 20, 15, 10, or 5 nucleotides of the position corresponding to SEQ ID NO:2 when optimally aligned with SEQ ID NO:2. Optionally, the segment is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the corresponding sequence of SEQ ID NO:2 when optimally aligned with SEQ ID NO:2. Optionally, the segment comprises at least 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or 2000 consecutive nucleotides of SEQ ID NO:2. Optionally, the segment comprises position 12666 of SEQ ID NO:2, or a position corresponding to position 12666 of SEQ ID NO:2 when optimally aligned with SEQ ID NO:2. Optionally, the HSD17B13 gene is a human HSD17B13 gene. Optionally, the isolated nucleic acid is up to about 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides in length. Optionally, the isolated nucleic acid is linked to a heterologous nucleic acid or comprises a heterologous label. Optionally, the heterologous label is a fluorescent label.

[0018] In another aspect, an isolated nucleic acid is provided that hybridizes with different HSD17B13 mRNA transcripts or cDNAs. Some such isolated nucleic acids hybridize with at least 15 consecutive nucleotides of a nucleic acid encoding an HSD17B13 protein, wherein the consecutive nucleotides include a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment that is not present in SEQ ID NO: 4 (HSD17B13 transcript A), but is present in SEQ ID NO: 7 (HSD17B13 transcript D), SEQ ID NO: 10 (HSD17B13 transcript G), and SEQ ID NO: 11 (HSD17B13 transcript H).

[0019] Some such isolated nucleic acids hybridize to at least 15 consecutive nucleotides of a nucleic acid encoding an HSD17B13 protein, wherein the consecutive nucleotides include a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a segment present in SEQ ID NO: 8 (HSD17B13 transcript E) and SEQ ID NO: 11 (HSD17B13 transcript H) that is not present in SEQ ID NO: 4 (HSD17B13 transcript A).

[0020] Some such isolated nucleic acids hybridize to at least 15 consecutive nucleotides of a nucleic acid encoding an HSD17B13 protein, wherein the consecutive nucleotides include a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a segment within SEQ ID NO: 9 (HSD17B13 transcript F) that is not present in SEQ ID NO: 4 (HSD17B13 transcript A).

[0021] Some such isolated nucleic acids hybridize to at least 15 consecutive nucleotides of a nucleic acid encoding an HSD17B13 protein, wherein the consecutive nucleotides include a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a segment present in SEQ ID NO: 6 (HSD17B13 transcript C) that is not present in SEQ ID NO: 4 (HSD17B13 transcript A).

[0022] Optionally, the HSD17B13 protein is human HSD17B13 protein.Optionally, the isolated nucleic acid is up to about 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides in length.Optionally, the isolated nucleic acid is linked to a heterologous nucleic acid or comprises a heterologous label.Optionally, the heterologous label is a fluorescent label.

[0023] Optionally, any of the above-mentioned isolated nucleic acids comprises DNA.Optionally, any of the above-mentioned isolated nucleic acids comprises RNA.Optionally, any of the above-mentioned isolated nucleic acids is antisense RNA, short hairpin RNA or short interfering RNA.Optionally, any of the above-mentioned isolated nucleic acids can comprise non-natural nucleotides.

[0024] In another aspect, a vector comprising any of the above isolated nucleic acids and a heterologous nucleic acid sequence and an exogenous donor sequence is provided.

[0025] In another aspect, there is provided use of any of the above isolated nucleic acids, vectors, or exogenous donor sequences in a method for detecting the HSD17B13 rs72613567 variant in a subject, a method for detecting the presence of HSD17B13 transcript C, D, E, F, G, or H in a subject, a method for determining a subject's susceptibility to developing chronic liver disease, a method for diagnosing a subject with fatty liver disease, or a method for modifying the HSD17B13 gene in a cell, or a method for altering expression of the HSD17B13 gene in a cell.

[0026] In another aspect, there is provided a guide RNA that targets HSD17B13 gene.Such guide RNA can be effective for directing Cas enzyme to bind to or cut HSD17B13 gene, and wherein guide RNA comprises a DNA targeting segment that hybridizes with guide RNA recognition sequence in HSD17B13 gene.That is, such guide RNA can be effective for directing Cas enzyme to bind to or cut HSD17B13 gene, and wherein guide RNA comprises a DNA targeting segment that targets guide RNA target sequence in HSD17B13 gene. Such a guide RNA can be effective to direct a Cas enzyme to bind to or cleave the HSD17B13 gene, wherein the guide RNA comprises a DNA-targeting segment that targets a guide RNA target sequence within the HSD17B13 gene that includes or is proximate to a position corresponding to position 12666 of SEQ ID NO:2 when the HSD17B13 gene is optimally aligned with SEQ ID NO:2. Optionally, the guide RNA target sequence comprises, consists essentially of, or consists of any one of SEQ ID NOs:226-239 and 264-268. Optionally, the DNA-targeting segment comprises, consists essentially of, or consists of any one of SEQ ID NOs:1629-1642 and 1648-1652. Optionally, the guide RNA comprises, consists essentially of, or consists of any one of SEQ ID NOS: 706-719; 936-949; 1166-1179, 1396-1409, 725-729, 955-959, 1185-1189, and 1415-1419. Optionally, the guide RNA target sequence is selected from SEQ ID NOS: 226-239 or 230 and 231. Optionally, the guide RNA target sequence is selected from SEQ ID NOS: 226-230 and 264-268. Optionally, the guide RNA target sequence is within a region corresponding to exon 6 and / or intron 6 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2.Optionally, the guide RNA target sequence is within a region corresponding to exon 6 and / or intron 6 and / or exon 7 of SEQ ID NO:2 when the HSD17B13 gene is optimally aligned with SEQ ID NO:2. Optionally, the guide RNA target sequence is within about 1000, 500, 400, 300, 200, 100, 50, 45, 40, 35, 30, 25, 20, 15, 10, or 5 nucleotides from a position corresponding to position 12666 of SEQ ID NO:2 when the HSD17B13 gene is optimally aligned with SEQ ID NO:2. Optionally, the guide RNA target sequence includes a position corresponding to position 12666 of SEQ ID NO:2 when the HSD17B13 gene is optimally aligned with SEQ ID NO:2.

[0027] Such a guide RNA can be effective for directing a Cas enzyme to bind to or cleave the HSD17B13 gene, wherein the guide RNA comprises a DNA-targeting segment that targets a guide RNA target sequence within the HSD17B13 gene that includes or is proximal to the start codon of the HSD17B13 gene. Optionally, the guide RNA target sequence comprises, consists essentially of, or consists of any one of SEQ ID NOs: 20-81 and 259-263. Optionally, the DNA-targeting segment comprises, consists essentially of, or consists of any one of SEQ ID NOs: 1423-1484 and 1643-1647. Optionally, the guide RNA comprises, consists essentially of, or consists of any one of SEQ ID NOs: 500-561, 730-791, 960-1021, 1190-1251, 720-724, 950-954, 1180-1184, and 1410-1414. Optionally, the guide RNA target sequence is selected from SEQ ID NOs: 20-81 and 259-263. Optionally, the guide RNA target sequence is selected from SEQ ID NOs: 21-23, 33, and 35. Optionally, the guide RNA target sequence is selected from SEQ ID NOs: 33 and 35. Optionally, the guide RNA target sequence is within a region corresponding to exon 1 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2. Optionally, the guide RNA target sequence is within about 1000, 500, 400, 300, 200, 100, 50, 45, 40, 35, 30, 25, 20, 15, 10, or 5 nucleotides from the start codon.

[0028] Such a guide RNA may be effective for directing a Cas enzyme to bind to or cleave the HSD17B13 gene, wherein the guide RNA comprises a DNA-targeting segment that targets a guide RNA target sequence within the HSD17B13 gene that includes or is adjacent to the stop codon of the HSD17B13 gene. Optionally, the guide RNA target sequence comprises, consists essentially of, or consists of any one of SEQ ID NOs: 82-225. Optionally, the DNA-targeting segment comprises, consists essentially of, or consists of any one of SEQ ID NOs: 1485-1628. Optionally, the guide RNA comprises, consists essentially of, or consists of any one of SEQ ID NOs: 562-705, 792-935, 1022-1165, and 1252-1395. Optionally, the guide RNA target sequence is selected from SEQ ID NOs: 82-225. Optionally, the guide RNA target sequence is within a region corresponding to exon 7 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2. Optionally, the guide RNA target sequence is within about 1000, 500, 400, 300, 200, 100, 50, 45, 40, 35, 30, 25, 20, 15, 10, or 5 nucleotides from the stop codon.

[0029] Optionally, the HSD17B13 gene is a human HSD17B13 gene. Optionally, the HSD17B13 gene comprises SEQ ID NO:2.

[0030] Some such guide RNAs include clustered regularly interspaced short palindromic repeats (CRISPR) RNA (crRNA) and trans-activating CRISPR RNA (tracrRNA) that contain a DNA-targeting segment. Optionally, the guide RNA is a modular guide RNA, in which the crRNA and tracrRNA are separate molecules that hybridize to each other. Optionally, the crRNA comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 1421, and the tracrRNA comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 1422. Optionally, the guide RNA is a single guide RNA in which the crRNA is fused to the tracrRNA via a linker. Optionally, the single guide RNA comprises, consists essentially of, or consists of the sequence set forth in any one of SEQ ID NOs: 1420 and 256-258.

[0031] In another aspect, antisense RNA, siRNA, or shRNA is provided that hybridizes with a sequence within the HSD17B13 transcript disclosed herein. Some such antisense RNA, siRNA, or shRNA hybridizes with a sequence within SEQ ID NO: 4 (HSD17B13 transcript A). Optionally, the antisense RNA, siRNA, or shRNA can reduce the expression of HSD17B13 transcript A in cells. Optionally, the antisense RNA, siRNA, or shRNA hybridizes with a sequence present in SEQ ID NO: 4 (HSD17B13 transcript A) that is not present in SEQ ID NO: 7 (HSD17B13 transcript D). Optionally, the antisense RNA, siRNA, or shRNA hybridizes with a sequence within exon 7 of SEQ ID NO: 4 (HSD17B13 transcript A) or a sequence spanning the boundary between exon 6 and exon 7. Some such antisense RNAs, siRNAs, or shRNAs hybridize with a sequence within SEQ ID NO: 7 (HSD17B13 transcript D). Optionally, the antisense RNAs, siRNAs, or shRNAs may reduce expression of HSD17B13 transcript D in cells. Optionally, the antisense RNAs, siRNAs, or shRNAs hybridize with a sequence present in SEQ ID NO: 7 (HSD17B13 transcript D) that is not present in SEQ ID NO: 4 (HSD17B13 transcript A). Optionally, the antisense RNAs, siRNAs, or shRNAs hybridize with a sequence within exon 7 of SEQ ID NO: 7 (HSD17B13 transcript D) or a sequence spanning the boundary between exons 6 and 7.

[0032] In another aspect, DNA is provided that encodes any of the above-mentioned guide RNA, antisense RNA, siRNA or shRNA.In another aspect, DNA is provided that encodes any of the above-mentioned guide RNA, antisense RNA, siRNA or shRNA and a vector that comprises heterologous nucleic acid.In another aspect, use of any of the above-mentioned guide RNA, antisense RNA, siRNA or shRNA, DNA that encodes guide RNA, antisense RNA, siRNA or shRNA, or the vector that comprises DNA that encodes guide RNA, antisense RNA, siRNA or shRNA is provided in the method for modifying HSD17B13 gene in cells or the method for changing the expression of HSD17B13 gene in cells.

[0033] In another aspect, a composition is provided comprising any of the above-described isolated nucleic acids, any of the above-described guide RNAs, any of the above-described isolated polypeptides, any of the above-described antisense RNAs, siRNAs, or shRNAs, any of the above-described vectors, or any of the above-described exogenous donor sequences. Optionally, the composition comprises any of the above-described guide RNAs and a Cas protein, such as a Cas9 protein. Optionally, such a composition comprises a carrier that increases the stability of the isolated polypeptide, guide RNA, antisense RNA, siRNA, shRNA, isolated nucleic acid, vector, or exogenous donor sequence. Optionally, the carrier comprises poly(lactic acid) (PLA) microspheres, poly(D,L-lactic acid-co-glycolic acid) (PLGA) microspheres, liposomes, micelles, reverse micelles, lipid cochleates, or lipid microtubules.

[0034] Also provided is a cell that comprises any of the above-mentioned isolated nucleic acids, any of the above-mentioned guide RNAs, any of the above-mentioned antisense RNAs, siRNAs or shRNAs, any of the above-mentioned isolated polypeptides, or any of the above-mentioned vectors.Optionally, the cell is human cell, rodent cell, mouse cell or rat cell.Optionally, any of the above-mentioned cell is hepatocyte or pluripotent cell.

[0035] Also provided is the use of any of the above guide RNAs in a method for modifying the HSD17B13 gene in a cell or a method for changing the expression of the HSD17B13 gene in a cell.Also provided is the use of any of the above antisense RNAs, siRNAs, or shRNAs in a method for changing the expression of the HSD17B13 gene in a cell.

[0036] Also provided are methods for modifying a cell, modifying the HSD17B13 gene, or altering expression of the HSD17B13 gene. Some such methods are for modifying the HSD17B13 gene in a cell, and include contacting the genome of the cell with (a) a Cas protein; and (b) a guide RNA that forms a complex with the Cas protein and targets a guide RNA target sequence in the HSD17B13 gene, where the guide RNA target sequence includes or is close to a position corresponding to position 12666 of SEQ ID NO:2 when the HSD17B13 gene is optimally aligned with SEQ ID NO:2, and the Cas protein cleaves the HSD17B13 gene. Optionally, the Cas protein is a Cas9 protein. Optionally, the guide RNA target sequence includes, consists essentially of, or consists of any one of SEQ ID NOs:226-239 and 264-268. Optionally, the DNA-targeting segment comprises, consists essentially of, or consists of any one of SEQ ID NOs: 1629-1642 and 1648-1652. Optionally, the guide RNA comprises, consists essentially of, or consists of any one of SEQ ID NOs: 706-719; 936-949; 1166-1179, 1396-1409, 725-729, 955-959, 1185-1189, and 1415-1419. Optionally, the guide RNA target sequence is selected from SEQ ID NOs: 226-239, or the guide RNA target sequence is selected from SEQ ID NOs: 230 and 231. Optionally, the guide RNA target sequence is selected from SEQ ID NOs: 226-239 and 264-268, or selected from SEQ ID NOs: 264-268. Optionally, the guide RNA target sequence is within a region corresponding to exon 6 and / or intron 6 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2. Optionally, the guide RNA target sequence is within a region corresponding to exon 6 and / or intron 6 and / or exon 7 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2.Optionally, the guide RNA target sequence is within about 1000, 500, 400, 300, 200, 100, 50, 45, 40, 35, 30, 25, 20, 15, 10, or 5 nucleotides from a position corresponding to position 12666 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2. Optionally, the guide RNA target sequence includes a position corresponding to position 12666 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2.

[0037] Some such methods further include contacting the genome with an exogenous donor sequence comprising a 5' homology arm that hybridizes to a target sequence 5' of the position corresponding to position 12666 of SEQ ID NO:2 and a 3' homology arm that hybridizes to a target sequence 3' of the position corresponding to position 12666 of SEQ ID NO:2, wherein the exogenous donor sequence is recombined with the HSD17B13 gene. Optionally, the exogenous donor sequence further comprises a nucleic acid insert flanked by the 5' homology arm and the 3' homology arm. Optionally, the nucleic acid insert comprises a thymine, such that when the exogenous donor sequence is recombined with the HSD17B13 gene, a thymine is inserted between the nucleotides corresponding to positions 12665 ​​and 12666 of SEQ ID NO:1 when the HSD17B13 gene is optimally aligned with SEQ ID NO:1. Optionally, the exogenous donor sequence is from about 50 nucleotides to about 1 kb in length, or from about 80 nucleotides to about 200 nucleotides in length. Optionally, the exogenous donor sequence is a single-stranded oligodeoxynucleotide.

[0038] Some such methods are for modifying the HSD17B13 gene in a cell, comprising contacting the genome of the cell with (a) a Cas protein; and (b) a first guide RNA that forms a complex with the Cas protein and targets a first guide RNA target sequence within the HSD17B13 gene, where the first guide RNA target sequence includes the start codon of the HSD17B13 gene, or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the start codon, or is selected from SEQ ID NOs: 20-81, or is selected from SEQ ID NOs: 20-81 and 259-263, wherein the Cas protein cleaves or alters expression of the HSD17B13 gene. Optionally, the first guide RNA target sequence comprises, consists essentially of, or consists of any one of SEQ ID NOs: 20-81 and 259-263. Optionally, the first guide RNA target sequence comprises, consists essentially of, or consists of any one of SEQ ID NOs: 20-41, any one of SEQ ID NOs: 21-23, 33, and 35, or any one of SEQ ID NOs: 33 and 35. Optionally, the first guide RNA comprises, consists essentially of, or consists of a DNA-targeting segment comprising any one of SEQ ID NOs: 1423-1484 and 1643-1647. Optionally, the first guide RNA comprises, consists essentially of, or consists of a DNA-targeting segment comprising any one of SEQ ID NOs: 1447-1468, any one of SEQ ID NOs: 1448-1450, 1460, and 1462; or any one of SEQ ID NOs: 1460 and 1462. Optionally, the first guide RNA comprises, consists essentially of, or consists of any one of SEQ ID NOs: 500-561, 730-791, 960-1021, 1190-1251, 720-724, 950-954, 1180-1184, and 1410-1414.Optionally, the first guide RNA comprises, consists essentially of, or consists of any one of SEQ ID NOs: 524-545, 754-775, 984-1005, and 1214-1235, or any one of SEQ ID NOs: 295-297, 525-527, 755-757, 985-987, 1215-1217, 307, 309, 537, 539, 767, 769, 997, 999, 1227, and 1229, or any one of SEQ ID NOs: 307, 309, 537, 539, 767, 769, 997, 999, 1227, and 1229. Optionally, the first guide RNA target sequence is selected from SEQ ID NOs: 20-41, or selected from SEQ ID NOs: 21-23, 33, and 35, or selected from SEQ ID NOs: 33 and 35. Optionally, the Cas protein is a Cas9 protein. Optionally, the Cas protein is a nuclease-active Cas protein. Optionally, the Cas protein is a nuclease-inactive Cas protein fused to a transcription activation domain or a nuclease-inactive Cas protein fused to a transcription repressor domain.

[0039] Some such methods further include contacting the genome of the cell with a second guide RNA that is complexed with a Cas protein and targets a second guide RNA target sequence in the HSD17B13 gene, where the second guide RNA target sequence includes a stop codon of the HSD17B13 gene or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the stop codon, or is selected from SEQ ID NOs: 82-225, wherein the cell has been modified to include a deletion between the first and second guide RNA target sequences. Optionally, the second guide RNA target sequence comprises, consists essentially of, or consists of any one of SEQ ID NOs: 82-225. Optionally, the second guide RNA comprises, consists essentially of, or consists of a DNA-targeting segment comprising any one of SEQ ID NOs: 1485-1628. Optionally, the second guide RNA comprises, consists essentially of, or consists of any one of SEQ ID NOs: 562-705, 792-935, 1022-1165, and 1252-1395.

[0040] Some such methods are for reducing the expression of the HSD17B13 gene in a cell or for reducing the expression of a specific HSD17B13 transcript (e.g., transcript A or transcript D) in a cell. Some such methods are for reducing the expression of the HSD17B13 gene in a cell, and include contacting the genome of the cell with an antisense RNA, siRNA, or shRNA that hybridizes to a sequence in exon 7 of SEQ ID NO: 4 (HSD17B13 transcript A) and reduces the expression of HSD17B13 transcript A. Some such methods are for reducing the expression of the HSD17B13 gene in a cell, and include contacting the genome of the cell with an antisense RNA, siRNA, or shRNA that hybridizes to a sequence in an HSD17B13 transcript disclosed herein. In some such methods, the antisense RNA, siRNA, or shRNA hybridizes to a sequence in SEQ ID NO: 4 (HSD17B13 transcript A). Optionally, the antisense RNA, siRNA, or shRNA may reduce expression of HSD17B13 transcript A in the cell. Optionally, the antisense RNA, siRNA, or shRNA hybridizes to a sequence present in SEQ ID NO: 4 (HSD17B13 transcript A) that is not present in SEQ ID NO: 7 (HSD17B13 transcript D). Optionally, the antisense RNA, siRNA, or shRNA hybridizes to a sequence within exon 7 of SEQ ID NO: 4 (HSD17B13 transcript A) or a sequence spanning the boundary between exons 6 and 7. In some such methods, the antisense RNA, siRNA, or shRNA hybridizes to a sequence within SEQ ID NO: 7 (HSD17B13 transcript D). Optionally, the antisense RNA, siRNA, or shRNA may reduce expression of HSD17B13 transcript D in the cell. Optionally, the antisense RNA, siRNA, or shRNA hybridizes to a sequence present in SEQ ID NO: 7 (HSD17B13 transcript D) that is not present in SEQ ID NO: 4 (HSD17B13 transcript A).Optionally, the antisense RNA, siRNA, or shRNA hybridizes to a sequence within exon 7 or spanning the boundary between exons 6 and 7 of SEQ ID NO: 7 (HSD17B13 transcript D).

[0041] In any of the above methods for modifying or altering the expression of the HSD17B13 gene, the method may further include introducing an expression vector into a cell, the expression vector comprising a recombinant HSD17B13 gene containing a thymine inserted between the nucleotides corresponding to positions 12665 ​​and 12666 of SEQ ID NO:1 when the recombinant HSD17B13 gene is optimally aligned with SEQ ID NO:1. Optionally, the recombinant HSD17B13 gene is a human gene. Optionally, the recombinant HSD17B13 gene is an HSD17B13 minigene in which one or more non-essential segments of the gene are deleted relative to the corresponding wild-type HSD17B13 gene. Optionally, the deleted segments comprise one or more intron sequences. Optionally, the HSD17B13 minigene comprises an intron corresponding to intron 6 of SEQ ID NO:2 when optimally aligned with SEQ ID NO:2.

[0042] In any of the above methods for modifying or altering expression of the HSD17B13 gene, the method may further comprise introducing into the cell an expression vector, wherein the expression vector comprises a nucleic acid encoding an HSD17B13 protein that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D). Optionally, the nucleic acid encoding the HSD17B13 protein is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 7 (HSD17B13 transcript D) when optimally aligned with SEQ ID NO: 7.

[0043] In any of the above methods for modifying the HSD17B13 gene or altering expression of the HSD17B13 gene, the method may further comprise introducing into a cell an HSD17B13 protein or a fragment thereof. Optionally, the HSD17B13 protein or fragment thereof is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D).

[0044] Some such methods are for modifying cells and include introducing an expression vector into the cell, wherein the expression vector comprises a recombinant HSD17B13 gene comprising a thymine inserted between the nucleotides corresponding to positions 12665 ​​and 12666 of SEQ ID NO:1 when the recombinant HSD17B13 gene is optimally aligned with SEQ ID NO:1. Optionally, the recombinant HSD17B13 gene is a human gene. Optionally, the recombinant HSD17B13 gene is an HSD17B13 minigene in which one or more non-essential segments of the gene are deleted relative to the corresponding wild-type HSD17B13 gene. Optionally, the deleted segments comprise one or more intron sequences. Optionally, the HSD17B13 minigene comprises an intron corresponding to intron 6 of SEQ ID NO:2 when optimally aligned with SEQ ID NO:2.

[0045] Some such methods are for modifying a cell and include introducing an expression vector into the cell, wherein the expression vector comprises a nucleic acid encoding an HSD17B13 protein that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D). Optionally, the nucleic acid encoding the HSD17B13 protein is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 7 (HSD17B13 transcript D) when optimally aligned with SEQ ID NO: 7.

[0046] Some such methods are for modifying a cell and include introducing into the cell an HSD17B13 protein or fragment thereof, optionally the HSD17B13 protein or fragment thereof is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D).

[0047] In any of the above methods of modifying cells, modifying the HSD17B13 gene, or altering expression of the HSD17B13 gene, the cells may be human cells, rodent cells, mouse cells, or rat cells. Any of the cells may be pluripotent cells or differentiated cells. Any of the cells may be hepatocytes. In any of the above methods of modifying cells, modifying the HSD17B13 gene, or altering expression of the HSD17B13 gene, the method or cells may be ex vivo or in vivo. The guide RNA used in any of the above methods may be a modular guide RNA comprising separate crRNA and tracrRNA molecules that hybridize to each other, or may be a single guide RNA in which the crRNA portion is fused to the tracrRNA portion (e.g., by a linker).

[0048] In another aspect, provided is a method for treating the subject who has or is prone to develop chronic liver disease.In another aspect, provided is a method for treating the subject who has or is prone to develop alcoholic or non-alcoholic liver disease.Such subject can be, for example, the subject who is not a carrier of HSD17B13 rs72613567 variant or the subject who is not a carrier of homozygous HSD17B13 rs72613567 variant.Some such methods include: A method of treating a subject who is not a carrier of the rs72613567 variant and has or is susceptible to developing chronic liver disease, comprising administering to the subject: (a) a Cas protein or a nucleic acid encoding the Cas protein; (b) a guide RNA or a nucleic acid encoding the guide RNA, wherein the guide RNA forms a complex with the Cas protein and targets a guide RNA target sequence in the HSD17B13 gene, the guide RNA target sequence including or adjacent to a position corresponding to position 12666 of SEQ ID NO:2 when the HSD17B13 gene is optimally aligned with SEQ ID NO:2; and (c) a nucleic acid encoding the guide RNA 5' to the position corresponding to position 12666 of SEQ ID NO:2. a 5' homologous arm that hybridizes to a target sequence of SEQ ID NO:2, a 3' homologous arm that hybridizes to a target sequence 3' to a position corresponding to position 12666 of SEQ ID NO:2, and a nucleic acid insert comprising a thymine flanked by the 5' homologous arm and the 3' homologous arm, wherein the Cas protein cleaves the HSD17B13 gene in hepatocytes of the subject, and the exogenous donor sequence recombines with the HSD17B13 gene in the hepatocytes, such that upon recombination of the exogenous donor sequence with the HSD17B13 gene, a thymine is inserted between the nucleotides corresponding to positions 12665 ​​and 12666 of SEQ ID NO:1 when the HSD17B13 gene is optimally aligned with SEQ ID NO:1.

[0049] Optionally, the guide RNA target sequence is selected from SEQ ID NOs: 226-239, or the guide RNA target sequence is selected from SEQ ID NOs: 230 and 231. Optionally, the guide RNA target sequence is selected from SEQ ID NOs: 226-239 and 264-268. Optionally, the guide RNA target sequence is within a region corresponding to exon 6 and / or intron 6 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2. Optionally, the guide RNA target sequence is within a region corresponding to exon 6 and / or intron 6 and / or exon 7 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2. Optionally, the guide RNA target sequence is within about 1000, 500, 400, 300, 200, 100, 50, 45, 40, 35, 30, 25, 20, 15, 10, or 5 nucleotides from a position corresponding to position 12666 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2. Optionally, the guide RNA target sequence includes a position corresponding to position 12666 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2.

[0050] Optionally, the exogenous donor sequence is about 50 nucleotides to about 1 kb in length. Optionally, the exogenous donor sequence is about 80 nucleotides to about 200 nucleotides in length. Optionally, the exogenous donor sequence is a single-stranded oligodeoxynucleotide.

[0051] Some such methods are methods of treating a subject who is not a carrier of the HSD17B13 rs72613567 variant and who has or is susceptible to developing chronic liver disease, comprising administering to the subject: (a) a Cas protein or a nucleic acid encoding a Cas protein; (b) a first guide RNA or a nucleic acid encoding the first guide RNA, wherein the first guide RNA forms a complex with the Cas protein and targets a first guide RNA target sequence in the HSD17B13 gene, and the first guide RNA target sequence includes the start codon of the HSD17B13 gene or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the start codon; or and (c) an expression vector containing a recombinant HSD17B13 gene comprising a thymine inserted between the nucleotide corresponding to positions 12665 ​​and 12666 of SEQ ID NO:1 when the recombinant HSD17B13 gene is optimally aligned with SEQ ID NO:1, wherein the Cas protein cleaves or alters expression of the HSD17B13 gene in hepatocytes of the subject, and the expression vector expresses the recombinant HSD17B13 gene in hepatocytes of the subject.Some such methods are methods of treating a subject who is not a carrier of the HSD17B13 rs72613567 variant and who has or is susceptible to developing chronic liver disease, comprising administering to the subject (a) a Cas protein or a nucleic acid encoding a Cas protein; and (b) a first guide RNA or a nucleic acid encoding the first guide RNA, wherein the first guide RNA forms a complex with the Cas protein and targets a first guide RNA target sequence in the HSD17B13 gene, and the first guide RNA target sequence includes the start codon of the HSD17B13 gene or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the start codon, or is within about SEQ ID NO: and optionally (c) introducing an expression vector containing a recombinant HSD17B13 gene comprising a thymine inserted between the nucleotides corresponding to positions 12665 ​​and 12666 of SEQ ID NO:1 when the recombinant HSD17B13 gene is optimally aligned with SEQ ID NO:1, wherein the Cas protein cleaves or alters expression of the HSD17B13 gene in hepatocytes of the subject, and the expression vector expresses the recombinant HSD17B13 gene in hepatocytes of the subject.

[0052] Optionally, the first guide RNA target sequence is selected from SEQ ID NOs: 20-41, or selected from SEQ ID NOs: 21-23, 33, and 35, or selected from SEQ ID NOs: 33 and 35. Optionally, the Cas protein is a nuclease-active Cas protein. Optionally, the Cas protein is a nuclease-inactive Cas protein fused to a transcriptional repressor domain.

[0053] Such methods may further include introducing into the subject a second guide RNA, wherein the second guide RNA forms a complex with a Cas protein and targets a second guide RNA target sequence within the HSD17B13 gene, wherein the second guide RNA target sequence includes the stop codon of the HSD17B13 gene or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the stop codon, or is selected from SEQ ID NOs: 82-225, and wherein the Cas protein cleaves the HSD17B13 gene in the hepatocytes both within the first guide RNA target sequence and within the second guide RNA target sequence, and the hepatocytes are modified to contain a deletion between the first guide RNA target sequence and the second guide RNA target sequence.

[0054] Optionally, the recombinant HSD17B13 gene is an HSD17B13 minigene in which one or more non-essential segments of the gene are deleted relative to the corresponding wild-type HSD17B13 gene.Optionally, the deleted segments include one or more intron sequences.Optionally, the HSD17B13 minigene includes an intron corresponding to intron 6 of SEQ ID NO:2 when optimally aligned with SEQ ID NO:2.

[0055] In any of the above-mentioned methods of treatment or prevention, the Cas protein may be a Cas9 protein. In any of the above-mentioned methods of treatment or prevention, the subject may be a human. In any of the above-mentioned methods of treatment or prevention, the chronic liver disease may be fatty liver disease, nonalcoholic fatty liver disease (NAFLD), alcoholic liver fatty liver disease, cirrhosis, or hepatocellular carcinoma. Similarly, in any of the above-mentioned methods, The method of treatment or prevention can be for a liver disease that is alcoholic liver disease or non-alcoholic liver disease.

[0056] Some such methods include a method of treating a subject who is not a carrier of the HSD17B13 rs72613567 variant and has or is susceptible to developing chronic liver disease, comprising introducing into the subject an antisense RNA, siRNA, or shRNA that hybridizes with a sequence within exon 7 or a sequence spanning the boundary between exons 6 and 7 of SEQ ID NO: 4 (HSD17B13 transcript A) and reduces expression of HSD17B13 transcript A in the subject's liver cells. Some such methods include a method of treating a subject who is not a carrier of the HSD17B13 rs72613567 variant and has or is susceptible to developing chronic liver disease, comprising introducing into the subject an antisense RNA, siRNA, or shRNA that hybridizes with a sequence within the HSD17B13 transcript disclosed herein. Optionally, the antisense RNA, siRNA, or shRNA hybridizes with a sequence within SEQ ID NO: 4 (HSD17B13 transcript A). Optionally, the antisense RNA, siRNA, or shRNA may reduce expression of HSD17B13 transcript A in the cell. Optionally, the antisense RNA, siRNA, or shRNA hybridizes to a sequence present in SEQ ID NO: 4 (HSD17B13 transcript A) that is not present in SEQ ID NO: 7 (HSD17B13 transcript D). Optionally, the antisense RNA, siRNA, or shRNA hybridizes to a sequence within exon 7 of SEQ ID NO: 4 (HSD17B13 transcript A) or a sequence spanning the boundary between exons 6 and 7.

[0057] If necessary, such a method further includes a step of introducing an expression vector into the subject, the expression vector comprising a recombinant HSD17B13 gene comprising a thymine inserted between the nucleotide corresponding to positions 12665 ​​and 12666 of SEQ ID NO: 1 when the recombinant HSD17B13 gene is optimally aligned with SEQ ID NO: 1, and the expression vector expresses the recombinant HSD17B13 gene in liver cells of the subject.

[0058] Optionally, such a method further comprises introducing into the subject an expression vector, wherein the expression vector comprises a nucleic acid encoding an HSD17B13 protein that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D), and wherein the expression vector expresses the nucleic acid encoding the HSD17B13 protein in liver cells of the subject. Optionally, the nucleic acid encoding the HSD17B13 protein is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 7 (HSD17B13 transcript D) when optimally aligned with SEQ ID NO: 7.

[0059] Optionally, such methods further include introducing messenger RNA into the subject, wherein the messenger RNA encodes an HSD17B13 protein that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D), and wherein the mRNA expresses the HSD17B13 protein in liver cells of the subject. Optionally, complementary DNA reverse-transcribed from the messenger RNA is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 7 (HSD17B13 transcript D) when optimally aligned with SEQ ID NO: 7.

[0060] Optionally, such methods further comprise introducing into the subject an HSD17B13 protein or fragment thereof, Optionally, the HSD17B13 protein or fragment thereof is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D).

[0061] Some such methods include methods for treating a subject who is not a carrier of the HSD17B13 rs72613567 variant and who has or is susceptible to developing chronic liver disease, comprising the step of introducing into the subject an expression vector, wherein the expression vector comprises a recombinant HSD17B13 gene comprising a thymine inserted between the nucleotides corresponding to positions 12665 ​​and 12666 of SEQ ID NO:1 when the recombinant HSD17B13 gene is optimally aligned with SEQ ID NO:1, and wherein the expression vector expresses the recombinant HSD17B13 gene in liver cells of the subject.

[0062] In any of the above methods, the recombinant HSD17B13 gene can be a human gene. In any of the above methods, the recombinant HSD17B13 gene can be at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:2 when optimally aligned with SEQ ID NO:2. In any of the above methods, the recombinant HSD17B13 gene can be an HSD17B13 minigene in which one or more non-essential segments of the gene are deleted relative to the corresponding wild-type HSD17B13 gene. Optionally, the deleted segments include one or more intron sequences. Optionally, the HSD17B13 minigene includes an intron corresponding to intron 6 of SEQ ID NO:2 when optimally aligned with SEQ ID NO:2.

[0063] Some such methods include a method of treating a subject who is not a carrier of the HSD17B13 rs72613567 variant and has or is susceptible to developing chronic liver disease, comprising the step of introducing an expression vector into the subject, wherein the expression vector comprises a nucleic acid encoding an HSD17B13 protein that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D), and the expression vector expresses the nucleic acid encoding the HSD17B13 protein in liver cells of the subject. Optionally, the nucleic acid encoding the HSD17B13 protein is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 7 (HSD17B13 transcript D) when optimally aligned with SEQ ID NO: 7.

[0064] Some such methods include a method of treating a subject who is not a carrier of the HSD17B13 rs72613567 variant and has or is susceptible to developing chronic liver disease, comprising introducing messenger RNA into the subject, wherein the messenger RNA encodes an HSD17B13 protein that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D), and the mRNA expresses the HSD17B13 protein in the subject's liver cells. Optionally, complementary DNA reverse-transcribed from the messenger RNA is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 7 (HSD17B13 transcript D) when optimally aligned with SEQ ID NO: 7.

[0065] Some such methods include methods of treating a subject who is not a carrier of the HSD17B13 rs72613567 variant and who has or is susceptible to developing chronic liver disease, comprising introducing an HSD17B13 protein or fragment thereof into the subject's liver. Optionally, the HSD17B13 protein or fragment thereof is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D).

[0066] In any of the above methods, the subject can be a human.In any of the above methods, the chronic liver disease can be non-alcoholic fatty liver disease (NAFLD), alcoholic fatty liver disease, cirrhosis, or hepatocellular carcinoma.Similarly, in any of the above methods, the treatment or prevention method can be for liver disease that is alcoholic liver disease or non-alcoholic liver disease.In any of the above methods, the step of introducing into the subject can include hydrodynamic delivery, virus-mediated delivery, lipid nanoparticle-mediated delivery, or intravenous injection. In certain embodiments, for example, the following items are provided: (Item 1) A guide RNA effective to direct a Cas enzyme to bind to or cleave an HSD17B13 gene, the guide RNA comprising a DNA targeting segment that targets a guide RNA target sequence within the HSD17B13 gene. (Item 2) 2. The guide RNA of item 1, wherein the guide RNA target sequence comprises or is adjacent to a position corresponding to position 12666 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2. (Item 3) (a) the guide RNA target sequence comprises any one of SEQ ID NOs: 226-239 and 264-268; and / or (b) the DNA targeting segment comprises any one of SEQ ID NOs: 1629-1642 and 1648-1652; and / or (c) The guide RNA according to Item 2, wherein the guide RNA comprises any one of SEQ ID NOs: 706 to 719; 936 to 949; 1166 to 1179, 1396 to 1409, 725 to 729, 955 to 959, 1185 to 1189, and 1415 to 1419. (Item 4) (a) the guide RNA target sequence is within a region corresponding to exon 6 and / or intron 6 and / or exon 7 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2; and / or (b) The guide RNA of item 2 or 3, wherein the guide RNA target sequence is within about 1000, 500, 400, 300, 200, 100, 50, 45, 40, 35, 30, 25, 20, 15, 10, or 5 nucleotides from a position corresponding to position 12666 of SEQ ID NO:2 when the HSD17B13 gene is optimally aligned with SEQ ID NO:2, and optionally the guide RNA target sequence comprises a position corresponding to position 12666 of SEQ ID NO:2 when the HSD17B13 gene is optimally aligned with SEQ ID NO:2. (Item 5) 2. The guide RNA of item 1, wherein the guide RNA target sequence includes or is adjacent to the start codon of the HSD17B13 gene. (Item 6) (a) the guide RNA target sequence comprises any one of SEQ ID NOs: 20 to 81 and 259 to 263; and / or (b) the DNA targeting segment comprises any one of SEQ ID NOs: 1423-1484 and 1643-1647; and / or (c) The guide RNA according to Item 5, wherein the guide RNA comprises any one of SEQ ID NOs: 500 to 561, 730 to 791, 960 to 1021, 1190 to 1251, 720 to 724, 950 to 954, 1180 to 1184, and 1410 to 1414. (Item 7) (a) the guide RNA target sequence is within a region corresponding to exon 1 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2; and / or (b) The guide RNA of item 5 or 6, wherein the guide RNA target sequence is within about 1000, 500, 400, 300, 200, 100, 50, 45, 40, 35, 30, 25, 20, 15, 10, or 5 nucleotides from the start codon. (Item 8) 2. The guide RNA of item 1, wherein the guide RNA target sequence includes or is adjacent to a stop codon of the HSD17B13 gene. (Item 9) (a) the guide RNA target sequence comprises any one of SEQ ID NOs: 82 to 225, and / or (b) the DNA targeting segment comprises any one of SEQ ID NOs: 1485-1628; and / or (c) The guide RNA according to Item 8, wherein the guide RNA comprises any one of SEQ ID NOs: 562 to 705, 792 to 935, 1022 to 1165, and 1252 to 1395. (Item 10) (a) the guide RNA target sequence is within a region corresponding to exon 7 of SEQ ID NO:2 when the HSD17B13 gene is optimally aligned with SEQ ID NO:2; and / or (b) The guide RNA of item 8 or 9, wherein the guide RNA target sequence is within about 1000, 500, 400, 300, 200, 100, 50, 45, 40, 35, 30, 25, 20, 15, 10, or 5 nucleotides from a stop codon. (Item 11) 11. The guide RNA of any one of items 1 to 10, wherein the HSD17B13 gene is the human HSD17B13 gene or the mouse Hsd17b13 gene, optionally wherein the HSD17B13 gene is the human HSD17B13 gene and comprises SEQ ID NO: 2. (Item 12) 12. The guide RNA of any one of items 1 to 11, comprising a clustered regularly interspaced short palindromic repeats (CRISPR) RNA (crRNA) and a trans-activating CRISPR RNA (tracrRNA) comprising the DNA-targeting segment. (Item 13) 13. The guide RNA of claim 12, wherein the crRNA and the tracrRNA are modular guide RNAs that are separate molecules that hybridize to each other, and optionally the crRNA comprises the sequence set forth in SEQ ID NO: 1421 and the tracrRNA comprises the sequence set forth in SEQ ID NO: 1422. (Item 14) Item 13. The guide RNA of Item 12, wherein the crRNA is a single guide RNA fused to the tracrRNA via a linker, and optionally comprises a sequence set forth in any one of SEQ ID NOs: 1420 and 256 to 258. (Item 15) 15. Use of the guide RNA according to any one of items 1 to 14 in a method for modifying the HSD17B13 gene in a cell or a method for altering the expression of the HSD17B13 gene in a cell. (Item 16) 15. An isolated nucleic acid comprising a DNA encoding the guide RNA of any one of items 1 to 14. (Item 17) An antisense RNA, siRNA, or shRNA that hybridizes to a sequence within SEQ ID NO: 4 (HSD17B13 transcript A) and reduces expression of HSD17B13 transcript A in a cell. (Item 18) (a) hybridizes to a sequence present in SEQ ID NO: 4 (HSD17B13 transcript A) that is not present in SEQ ID NO: 7 (HSD17B13 transcript D); and / or (b) The antisense RNA, siRNA, or shRNA of item 17, which hybridizes to a sequence spanning the boundary between exon 6 and exon 7 of SEQ ID NO: 4 (HSD17B13 transcript A). (Item 19) 19. Use of the antisense RNA, siRNA, or shRNA according to item 17 or 18 in a method for altering the expression of the HSD17B13 gene in a cell. (Item 20) 19. An isolated nucleic acid comprising DNA encoding the antisense RNA, siRNA, or shRNA of item 17 or 18. (Item 21) 21. A vector comprising the isolated nucleic acid and heterologous nucleic acid of item 16 or 20. (Item 22) 15. A composition comprising the guide RNA of any one of items 1 to 14 and a carrier that increases the stability of the guide RNA, optionally further comprising a Cas protein, optionally wherein the Cas protein is Cas9. (Item 23) 19. A composition comprising the antisense RNA, siRNA, or shRNA of item 17 or 18 and a carrier that increases the stability of the antisense RNA, siRNA, or shRNA. (Item 24) 15. A cell comprising the guide RNA of any one of items 1 to 14. (Item 25) 19. A cell comprising the antisense RNA, siRNA, or shRNA of item 17 or 18. (Item 26) 26. The cell according to item 24 or 25, which is a human cell, optionally a hepatocyte. (Item 27) 26. The cell of item 24 or 25, which is a rodent, mouse, or rat cell, optionally a pluripotent cell or a hepatocyte cell. (Item 28) 1. A method for modifying the HSD17B13 gene in a cell, comprising: (a) a Cas protein; and (b) a guide RNA that forms a complex with the Cas protein and targets a guide RNA target sequence in the HSD17B13 gene, the guide RNA target sequence including or adjacent to a position corresponding to position 12666 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2; contacting the The method, wherein the Cas protein cleaves the HSD17B13 gene. (Item 29) (a) the guide RNA target sequence comprises any one of SEQ ID NOs: 226-239 and 264-268; and / or (b) the DNA targeting segment comprises any one of SEQ ID NOs: 1629-1642 and 1648-1652; and / or (c) The method according to Item 28, wherein the guide RNA comprises any one of SEQ ID NOs: 706 to 719; 936 to 949; 1166 to 1179, 1396 to 1409, 725 to 729, 955 to 959, 1185 to 1189, and 1415 to 1419. (Item 30) (a) the guide RNA target sequence is within a region corresponding to exon 6 and / or intron 6 and / or exon 7 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2; and / or (b) The method of item 28 or 29, wherein the guide RNA target sequence is within about 1000, 500, 400, 300, 200, 100, 50, 45, 40, 35, 30, 25, 20, 15, 10, or 5 nucleotides from a position corresponding to position 12666 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2, and optionally the guide RNA target sequence comprises a position corresponding to position 12666 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2. 31. The method of any one of items 28 to 30, further comprising contacting the genome with an exogenous donor sequence comprising a 5' homologous arm that hybridizes to a target sequence 5' of the position corresponding to position 12666 of SEQ ID NO: 2 and a 3' homologous arm that hybridizes to a target sequence 3' of the position corresponding to position 12666 of SEQ ID NO: 2, wherein the exogenous donor sequence is recombined with the HSD17B13 gene. (Item 32) 32. The method of claim 31, wherein the exogenous donor sequence further comprises a nucleic acid insert flanked by the 5' homology arm and the 3' homology arm. (Item 33) 33. The method of claim 32, wherein the nucleic acid insert comprises a thymine, and when the exogenous donor sequence is recombined with the HSD17B13 gene, the thymine is inserted between the nucleotides corresponding to positions 12665 ​​and 12666 of SEQ ID NO: 1 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 1. (Item 34) (a) the exogenous donor sequence is from about 50 nucleotides to about 1 kb in length, and optionally, the exogenous donor sequence is from about 80 nucleotides to about 200 nucleotides in length; and / or (b) The method of any one of items 31 to 33, wherein the exogenous donor sequence is a single-stranded oligodeoxynucleotide. (Item 35) 1. A method for modifying the HSD17B13 gene in a cell, comprising: (a) a Cas protein; and (b) a first guide RNA that forms a complex with the Cas protein and targets a first guide RNA target sequence within the HSD17B13 gene, wherein the first guide RNA target sequence includes the start codon of the HSD17B13 gene or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the start codon. contacting the The method, wherein the Cas protein cleaves the HSD17B13 gene or alters the expression of the HSD17B13 gene. (Item 36) (a) the first guide RNA target sequence comprises any one of SEQ ID NOs: 20-81 and 259-263, and optionally the first guide RNA target sequence comprises any one of SEQ ID NOs: 20-41, any one of SEQ ID NOs: 21-23, 33, and 35, or any one of SEQ ID NOs: 33 and 35; and / or (b) the first guide RNA comprises a DNA-targeting segment comprising any one of SEQ ID NOs: 1423-1484 and 1643-1647, and optionally, the first guide RNA comprises a DNA-targeting segment comprising any one of SEQ ID NOs: 1447-1468, any one of SEQ ID NOs: 1448-1450, 1460, and 1462; or any one of SEQ ID NOs: 1460 and 1462; and / or (c) the first guide RNA comprises any one of SEQ ID NOs: 500 to 561, 730 to 791, 960 to 1021, 1190 to 1251, 720 to 724, 950 to 954, 1180 to 1184, and 1410 to 1414, and optionally, the first guide RNA comprises any one of SEQ ID NOs: 524 to 545, 754 to 775, 984 to 1005, and 1214 to 1235; or any one of SEQ ID NOs: 295 to 297, 525 to 527, 755 to 757, 985 to 987, 1215 to 1217, 307, 309, 537, 539, 767, 769, 997, 999, 1227, and 1229, or any one of SEQ ID NOs: 307, 309, 537, 539, 767, 769, 997, 999, 1227, and 1229. (Item 37) (a) the Cas protein is a nuclease-active Cas protein; or (b) The method of any of items 35 or 36, wherein the Cas protein is a nuclease-inactive Cas protein fused to a transcriptional activation domain or a transcriptional repressor domain. (Item 38) 38. The method of any one of items 35 to 37, further comprising contacting the genome of the cell with a second guide RNA that forms a complex with the Cas protein and targets a second guide RNA target sequence in the HSD17B13 gene, wherein the second guide RNA target sequence includes a stop codon of the HSD17B13 gene or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides of the stop codon, and wherein the cell has been modified to include a deletion between the first and second guide RNA target sequences. (Item 39) (a) the second guide RNA target sequence comprises any one of SEQ ID NOs: 82 to 225; and / or (b) the second guide RNA comprises a DNA-targeting segment comprising any one of SEQ ID NOs: 1485-1628; and / or (c) The method of Item 38, wherein the second guide RNA comprises any one of SEQ ID NOs: 562 to 705, 792 to 935, 1022 to 1165, and 1252 to 1395. (Item 40) A method for reducing the expression of the HSD17B13 gene in a cell, comprising contacting the genome of the cell with an antisense RNA, siRNA, or shRNA that hybridizes to an sequence within sequence number 4 (HSD17B13 transcript A) and reduces the expression of HSD17B13 transcript A. (Item 41) 41. The method of claim 40, wherein the antisense RNA, siRNA, or shRNA hybridizes to a sequence present in SEQ ID NO: 4 (HSD17B13 transcript A) that is not present in SEQ ID NO: 7 (HSD17B13 transcript D), and optionally the antisense RNA, siRNA, or shRNA hybridizes to a sequence spanning the boundary between exon 6 and exon 7 of SEQ ID NO: 4 (HSD17B13 transcript A). (Item 42) 42. The method of any one of items 35 to 41, further comprising the step of introducing an expression vector into the cell, wherein the expression vector comprises the recombinant HSD17B13 gene comprising a thymine inserted between the nucleotide corresponding to positions 12665 ​​and 12666 of SEQ ID NO: 1 when the recombinant HSD17B13 gene is optimally aligned with SEQ ID NO: 1, and optionally the recombinant HSD17B13 gene is a human gene. (Item 43) 43. The method of claim 42, wherein the recombinant HSD17B13 gene is an HSD17B13 minigene in which one or more non-essential segments of the gene are deleted relative to the corresponding wild-type HSD17B13 gene, and optionally the deleted segments include one or more intron sequences, and optionally the HSD17B13 minigene includes an intron corresponding to intron 6 of SEQ ID NO: 2 when optimally aligned with SEQ ID NO: 2. (Item 44) 42. The method of any one of items 35 to 41, further comprising the step of introducing an expression vector into the cell, wherein the expression vector comprises a nucleic acid encoding an HSD17B13 protein that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D), and optionally the nucleic acid encoding the HSD17B13 protein is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 7 (HSD17B13 transcript D) when optimally aligned with SEQ ID NO: 7. (Item 45) 42. The method of any one of items 35 to 41, further comprising the step of introducing into the cell an HSD17B13 protein or a fragment thereof, wherein the HSD17B13 protein or a fragment thereof is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D). (Item 46) 46. ​​The method of any one of items 28 to 45, wherein the Cas protein is Cas9. (Item 47) 48. A method for modifying a cell, comprising the step of introducing an expression vector into the cell, wherein the expression vector contains a recombinant HSD17B13 gene containing a thymine inserted between the nucleotides corresponding to positions 12665 ​​and 12666 of SEQ ID NO: 1 when the recombinant HSD17B13 gene is optimally aligned with SEQ ID NO: 1, and optionally the recombinant HSD17B13 gene is a human gene. 48. The method of claim 47, wherein the recombinant HSD17B13 gene is an HSD17B13 minigene in which one or more non-essential segments of the gene are deleted relative to the corresponding wild-type HSD17B13 gene, and optionally the deleted segments include one or more intron sequences, and optionally the HSD17B13 minigene includes an intron corresponding to intron 6 of SEQ ID NO: 2 when optimally aligned with SEQ ID NO: 2. (Item 49) 1. A method for modifying a cell, comprising the step of introducing an expression vector into the cell, wherein the expression vector comprises a nucleic acid encoding an HSD17B13 protein that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D), and optionally the nucleic acid encoding the HSD17B13 protein is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 7 (HSD17B13 transcript D) when optimally aligned with SEQ ID NO: 7. (Item 50) A method for modifying a cell, comprising the step of introducing into the cell an HSD17B13 protein or a fragment thereof, wherein the HSD17B13 protein or fragment thereof is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D). (Item 51) 51. The method of any one of paragraphs 28 to 50, wherein the cell is a rodent cell, a mouse cell, or a rat cell, and optionally the cell is a pluripotent cell or a hepatocyte cell. (Item 52) 51. The method of any one of items 28 to 50, wherein the cells are human cells, and optionally the cells are hepatocytes. (Item 53) 53. The method of any one of items 28 to 52, wherein the cell is ex vivo or in vivo. (Item 54) 1. A method of treating a subject who is not a carrier of the HSD17B13 rs72613567 variant and who has or is susceptible to developing chronic liver disease, comprising administering to the subject: (a) a Cas protein or a nucleic acid encoding the Cas protein; (b) a guide RNA or a nucleic acid encoding the guide RNA, wherein the guide RNA forms a complex with the Cas protein and targets a guide RNA target sequence within the HSD17B13 gene, the guide RNA target sequence comprising or adjacent to a position corresponding to position 12666 of SEQ ID NO:2 when the HSD17B13 gene is optimally aligned with SEQ ID NO:2; and (c) an exogenous donor sequence comprising a 5' homologous arm that hybridizes to a target sequence 5' to a position corresponding to position 12666 of SEQ ID NO:2, a 3' homologous arm that hybridizes to a target sequence 3' to a position corresponding to position 12666 of SEQ ID NO:2, and a nucleic acid insert comprising a thymine flanked by the 5' homologous arm and the 3' homologous arm. introducing The Cas protein cleaves the HSD17B13 gene in the subject's hepatocytes, and the exogenous donor sequence recombines with the HSD17B13 gene in the hepatocytes, such that when the exogenous donor sequence recombines with the HSD17B13 gene, the thymine is inserted between the nucleotides corresponding to positions 12665 ​​and 12666 of SEQ ID NO: 1 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 1. (Item 55) 1. A method of treating a subject who is not a carrier of the HSD17B13 rs72613567 variant and who has or is susceptible to developing chronic liver disease, comprising administering to the subject: (a) a Cas protein or a nucleic acid encoding the Cas protein; (b) a first guide RNA or a nucleic acid encoding the first guide RNA, wherein the first guide RNA forms a complex with the Cas protein and targets a first guide RNA target sequence within the HSD17B13 gene, and the first guide RNA target sequence includes the start codon of the HSD17B13 gene or is within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides from the start codon, or is selected from SEQ ID NOs: 20-81; and (c) an expression vector comprising a recombinant HSD17B13 gene, the recombinant HSD17B13 gene containing a thymine inserted between the nucleotide corresponding to positions 12665 ​​and 12666 of SEQ ID NO: 1 when the recombinant HSD17B13 gene is optimally aligned with SEQ ID NO: 1; introducing The method, wherein the Cas protein cleaves or alters the expression of the HSD17B13 gene in the subject's hepatocytes, and the expression vector expresses the recombinant HSD17B13 gene in the subject's hepatocytes. (Item 56) A method for treating a subject who is not a carrier of the HSD17B13 rs72613567 variant and who has or is susceptible to developing chronic liver disease, comprising the step of introducing into the subject an antisense RNA, siRNA, or shRNA that hybridizes to a sequence within sequence number 4 (HSD17B13 transcript A) in the subject's liver cells and reduces expression of HSD17B13 transcript A. (Item 57) A method for treating a subject who is not a carrier of the HSD17B13 rs72613567 variant and who has or is susceptible to developing chronic liver disease, comprising the step of introducing into the subject an expression vector, wherein the expression vector comprises a recombinant HSD17B13 gene comprising a thymine inserted between the nucleotides corresponding to positions 12665 ​​and 12666 of SEQ ID NO: 1 when the recombinant HSD17B13 gene is optimally aligned with SEQ ID NO: 1, and wherein the expression vector expresses the recombinant HSD17B13 gene in hepatocytes of the subject. (Item 58) A method for treating a subject who is not a carrier of the HSD17B13 rs72613567 variant and has or is prone to developing chronic liver disease, comprising the step of introducing into the subject an expression vector, wherein the expression vector comprises a nucleic acid encoding an HSD17B13 protein that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D), and wherein the expression vector expresses the nucleic acid encoding the HSD17B13 protein in hepatocytes of the subject. (Item 59) A method for treating a subject who is not a carrier of the HSD17B13 rs72613567 variant and who has or is susceptible to developing chronic liver disease, comprising the step of introducing messenger RNA into the subject, wherein the messenger RNA encodes an HSD17B13 protein that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D), and the mRNA expresses the HSD17B13 protein in hepatocytes of the subject. (Item 60) A method for treating a subject who is not a carrier of the HSD17B13 rs72613567 variant and who has or is prone to developing chronic liver disease, the method comprising the step of introducing an HSD17B13 protein or a fragment thereof into the liver of the subject, wherein the HSD17B13 protein or fragment thereof is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 15 (HSD17B13 isoform D). [Brief explanation of the drawings]

[0067] [Figure 1A] Figures 1A and 1B show Manhattan plots (left) and quantile-quantile plots (right) of the association of single nucleotide variants with median alanine aminotransferase (ALT; Figure 1A) and aspartate aminotransferase (AST; Figure 1B) levels in the GHS discovery cohort. Figure 1A shows that 31 variants were present in 16 genes that were significantly associated with ALT levels (N = 41,414) at P < 1.0 × 10-7. Figure 1B shows that 12 variants were present in 10 genes that were significantly associated with AST levels (N = 40,753) at P < 1.0 × 10-7. All significant associations are shown in Table 2. Thirteen variants were present in nine genes (referred to herein by their gene names), including HSD17B13, that remained significantly associated with ALT or AST in a replication meta-analysis of three separate European ancestry cohorts (Table 3). The association test was well calibrated, as shown by the exome-wide quantile-quantile plot and genomic control lambda values ​​(Figure 1A and Figure 1B). [Figure 1B]Figures 1A and 1B show Manhattan plots (left) and quantile-quantile plots (right) of the association of single nucleotide variants with median alanine aminotransferase (ALT; Figure 1A) and aspartate aminotransferase (AST; Figure 1B) levels in the GHS discovery cohort. Figure 1A shows that 31 variants were present in 16 genes that were significantly associated with ALT levels (N = 41,414) at P < 1.0 × 10-7. Figure 1B shows that 12 variants were present in 10 genes that were significantly associated with AST levels (N = 40,753) at P < 1.0 × 10-7. All significant associations are shown in Table 2. Thirteen variants were present in nine genes (referred to herein by their gene names), including HSD17B13, that remained significantly associated with ALT or AST in a replication meta-analysis of three separate European ancestry cohorts (Table 3). The association test was well calibrated, as shown by the exome-wide quantile-quantile plot and genomic control lambda values ​​(Figure 1A and Figure 1B).

[0068] [Figure 2A]Figures 2A and 2B show that HSD17B13 rs72613567:TA is associated with a reduced risk of alcoholic and nonalcoholic liver disease phenotypes in the discovery cohort (Figure 2A) and a reduced risk of progression from simple steatosis to steatohepatitis and fibrosis in the bariatric surgery cohort (Figure 2B). Odds ratios were calculated using logistic regression, adjusting for age, age 2, sex, BMI, and ancestry principal components. Genotype odds ratios for heterozygous (Het OR) and homozygous (Hom OR) carriers are also shown. In the GHS discovery cohort in Figure 2A, HSD17B13 variants were associated with a significantly reduced risk of nonalcoholic and alcoholic liver disease, cirrhosis, and hepatocellular carcinoma in an allele-dose-dependent manner. In the GHS bariatric surgery cohort in Figure 2B, HSD17B13 rs72613567 was associated with 13% and 52% lower odds of nonalcoholic steatohepatitis (NASH) and 13% and 61% lower odds of fibrosis in heterozygous and homozygous TA carriers, respectively. [Figure 2B]Figures 2A and 2B show that HSD17B13 rs72613567:TA is associated with a reduced risk of alcoholic and nonalcoholic liver disease phenotypes in the discovery cohort (Figure 2A) and a reduced risk of progression from simple steatosis to steatohepatitis and fibrosis in the bariatric surgery cohort (Figure 2B). Odds ratios were calculated using logistic regression, adjusting for age, age 2, sex, BMI, and ancestry principal components. Genotype odds ratios for heterozygous (Het OR) and homozygous (Hom OR) carriers are also shown. In the GHS discovery cohort in Figure 2A, HSD17B13 variants were associated with a significantly reduced risk of nonalcoholic and alcoholic liver disease, cirrhosis, and hepatocellular carcinoma in an allele-dose-dependent manner. In the GHS bariatric surgery cohort in Figure 2B, HSD17B13 rs72613567 was associated with 13% and 52% lower odds of nonalcoholic steatohepatitis (NASH) and 13% and 61% lower odds of fibrosis in heterozygous and homozygous TA carriers, respectively.

[0069] [Figure 3]Figures 3A-3D show the expression of four HSD17B13 transcripts (A-D) in homozygous reference (T / T), heterozygous (T / TA), and homozygous alternative (TA / TA) carriers of the HSD17B13 rs72613567 splice variant. Each transcript is illustrated with its corresponding gene model. Coding regions in the gene model are indicated by striped boxes, and untranslated regions are indicated by black boxes. Figure 3A shows a representation of transcript A and expression data for transcript A. Figure 3B shows a representation of transcript B and expression data for transcript B. In transcript B, exon 2 is skipped. Figure 3C shows a representation of transcript C and expression data for transcript C. In transcript C, exon 6 is skipped. Figure 3D shows a representation of transcript D and expression data for transcript D. The asterisk in transcript D indicates the insertion of G from rs72613567 at the 3' end of exon 6, resulting in premature truncation of the protein. Transcript D becomes the dominant transcript in homozygous carriers of the HSD17B13 splice variant. Gene expression is shown in FPKM (fragments per kilobase of transcript per million mapped reads). The insets in Figure 3B and Figure 3C show a larger view.

[0070] [Figure 4] Figure 4 shows that an RNA-Seq study of human liver reveals eight HSD17B13 transcripts, including six novel HSD17B13 transcripts (transcripts C-H). Transcript expression is shown in FPKM (fragments per kilobase of transcript per million mapped reads). Transcript structures are provided on the right side of the figure.

[0071] [Figure 5A]Figures 5A and 5B show locus zoom plots (region association plots in the region around HSD17B13) of HSD17B13 in the GHS discovery cohort for ALT and AST, respectively. No significant recombination was observed between the regions. Diamonds indicate the rs72613567 splice variant. Each circle represents a single nucleotide variant, with the circle color indicating the linkage disequilibrium (r2 calculated in the DiscovEHR cohort) between that variant and rs72613567. Lines indicate the estimated recombination rate in HapMap. The bottom panel shows the relative position of each gene within the locus and the transcribed strand. There was no significant association of ALT or AST with coding or splice region variants in the neighboring gene HSD17B11 (the most significant P values ​​for ALT and AST are 1.4 × 10-1 and 4.3 × 10-2, respectively). [Figure 5B] Figures 5A and 5B show locus zoom plots (region association plots in the region around HSD17B13) of HSD17B13 in the GHS discovery cohort for ALT and AST, respectively. No significant recombination was observed between the regions. Diamonds indicate the rs72613567 splice variant. Each circle represents a single nucleotide variant, with the circle color indicating the linkage disequilibrium (r2 calculated in the DiscovEHR cohort) between that variant and rs72613567. Lines indicate the estimated recombination rate in HapMap. The bottom panel shows the relative position of each gene within the locus and the transcribed strand. There was no significant association of ALT or AST with coding or splice region variants in the neighboring gene HSD17B11 (the most significant P values ​​for ALT and AST are 1.4 × 10-1 and 4.3 × 10-2, respectively).

[0072] [Figure 6]Figures 6A-6D show the mRNA expression of four additional novel HSD17B13 transcripts (E-H) in homozygous reference (T / T), heterozygous (T / TA), and homozygous alternative (TA / TA) carriers of HSD17B13 splice variants. Each transcript is illustrated with its corresponding gene model. Coding regions in the gene model are indicated by striped boxes, and untranslated regions are indicated by black boxes. Figures 6A and 6D show that transcripts E and H contain an additional exon between exons 3 and 4. Figure 6B shows that transcript F contains a readthrough from exon 6 to intron 6. Figure 6C shows that exon 2 is skipped in transcript G. The asterisks in transcripts G and H (Figures 6C and 6D, respectively) indicate the insertion of G from rs72613567 at the 3' end of exon 6, resulting in premature truncation of the protein. Transcripts are differentially expressed according to HSD17B13 genotype as shown in the box plots. mRNA expression is shown in FPKM (fragments per kilobase of transcript per million mapped reads).

[0073] [Figure 7A] 7A-7B show protein sequence alignments of HSD17B13 protein isoforms A-H. [Figure 7B] 7A-7B show protein sequence alignments of HSD17B13 protein isoforms A-H.

[0074] [Figure 8]Figure 8 shows that HSD17B13 rs72613567:TA is associated with a reduced risk of alcoholic and non-alcoholic liver disease phenotypes. Specifically, Figure 8 shows that in the Dallas Liver Study, HSD17B13 rs72613567 was associated with a lower odds of any liver disease in an allele-dose-dependent manner. Similar allele-dose-dependent effects were observed across liver disease subtypes. Odds ratios were calculated using logistic regression, adjusting for age, age 2, sex, BMI, and self-reported ethnicity.

[0075] [Figure 9] Figure 9 shows that HSD17B13 rs72613567 is associated with a reduced risk of progression from simple fatty liver to steatohepatitis and fibrosis. Specifically, it shows the prevalence of histopathologically characterized liver disease according to HSD17B13 rs72613567 genotype in 2,391 individuals with liver biopsies from the GHS bariatric surgery cohort. The prevalence of normal liver did not differ by genotype (P=0.5 by chi-square test for trend in proportions), but for each TA allele, the prevalence of NASH decreased (P=1.6×10-4) and the prevalence of simple fatty liver increased (P=1.1×10-3).

[0076] [Figure 10ABC]Figures 10A-10E show the expression, subcellular localization, and enzymatic activity of novel HSD17B13 transcripts. Figure 10A shows Western blots from HepG2 cells overexpressing HSD17B13 transcripts A and D, demonstrating that HSD17B13 transcript D was translated into a truncated protein with a lower molecular weight compared to HSD17B13 transcript A. Figure 10B shows HSD17B13 Western blots from fresh-frozen human liver and HEK293 cell samples. The human liver samples were from homozygous reference (T / T), heterozygous (T / TA), and homozygous alternative (TA / TA) carriers of the HSD17B13 rs72613567 splice variant. The cell samples were from HEK293 cells overexpressing untagged HSD17B13 transcripts A and D. HSD17B13 transcript D was translated into a truncated protein, IsoD, with a lower molecular weight than HSD17B13 IsoA. Figure 10C shows that HSD17B13 IsoD protein levels were lower than IsoA protein levels from both human liver (left) and cell (right) samples. Protein levels normalized to actin are shown in the bar columns; **P<0.001, *P<0.05. Figure 10D shows the enzymatic activity of HSD17B13 isoforms A and D toward 17-beta estradiol (estradiol), leukotriene B4 (LTB4), and 13-hydroxyoctadecadienoic acid (13(S)-HODE). HSD17B13 isoform D exhibits enzymatic activity less than 10% of the corresponding value for isoform A. FIG. 10E shows that when HSD17B13 isoform D was overexpressed in HEK293 cells, it showed poor conversion of estradiol (substrate) to estrone (product) as measured in the culture medium, whereas overexpressed HSD17B13 isoform A showed robust conversion. [Figure 10DE]Figures 10A-10E show the expression, subcellular localization, and enzymatic activity of novel HSD17B13 transcripts. Figure 10A shows Western blots from HepG2 cells overexpressing HSD17B13 transcripts A and D, demonstrating that HSD17B13 transcript D was translated into a truncated protein with a lower molecular weight compared to HSD17B13 transcript A. Figure 10B shows HSD17B13 Western blots from fresh-frozen human liver and HEK293 cell samples. The human liver samples were from homozygous reference (T / T), heterozygous (T / TA), and homozygous alternative (TA / TA) carriers of the HSD17B13 rs72613567 splice variant. The cell samples were from HEK293 cells overexpressing untagged HSD17B13 transcripts A and D. HSD17B13 transcript D was translated into a truncated protein, IsoD, with a lower molecular weight than HSD17B13 IsoA. Figure 10C shows that HSD17B13 IsoD protein levels were lower than IsoA protein levels from both human liver (left) and cell (right) samples. Protein levels normalized to actin are shown in the bar columns; **P<0.001, *P<0.05. Figure 10D shows the enzymatic activity of HSD17B13 isoforms A and D toward 17-beta estradiol (estradiol), leukotriene B4 (LTB4), and 13-hydroxyoctadecadienoic acid (13(S)-HODE). HSD17B13 isoform D exhibits enzymatic activity less than 10% of the corresponding value for isoform A. FIG. 10E shows that when HSD17B13 isoform D was overexpressed in HEK293 cells, it showed poor conversion of estradiol (substrate) to estrone (product) as measured in the culture medium, whereas overexpressed HSD17B13 isoform A showed robust conversion.

[0077] [Figure 11]Figures 11A-11C show that HSD17B13 isoform D protein has a lower molecular weight and is unstable when overexpressed in HEK293 cells. Figure 11A shows RT-PCR of HSD17B13 from HEK293 cells overexpressing HSD17B13 transcripts A (IsoA) and D (IsoD), demonstrating that HSD17B13 IsoD RNA levels were higher than IsoA RNA levels. Figure 11B shows Western blots from the same cell line, demonstrating that HSD17B13 transcript D was translated into a truncated protein with a lower molecular weight compared to HSD17B13 transcript A. Figure 11C shows that HSD17B13 IsoD protein levels were lower than IsoA protein levels, whereas RNA levels were higher. HSD17B13 protein levels were normalized to actin; *P<0.05.

[0078] [Figure 12] Figure 12 shows similar localization patterns of HSD17B13 isoform A and isoform D to isolated lipid droplets (LDs) from HepG2 stable cell lines. ADRP and TIP47 were used as lipid droplet markers. LAMP1, calreticulin, and COX IV were used as markers for lysosomes, endoplasmic reticulum, and mitochondrial compartments, respectively. GAPDH was included as a cytoplasmic marker, and actin was used as a cytoskeleton marker. This experiment was repeated twice in HepG2 cells, and the above is representative of both runs. PNS = postnuclear fraction; TM = total membrane.

[0079] [Figure 13]Figures 13A-13D show that oleic acid increased triglyceride content in HepG2 cells overexpressing HSD17B13 transcripts A or D. Figure 13A shows that treatment with high concentrations of oleic acid increased triglyceride (TG) content to a similar extent in control (cells overexpressing GFP) and HSD17B13 transcript A and D cell lines. Figure 13B shows that RNA levels of HSD17B13 transcripts A and D were similar in the cell lines. RNA levels are shown as reads per transcript kilobase per million mapped reads (RPKM). Figure 13C shows a Western blot from HepG2 cells overexpressing HSD17B13 transcripts A and D. HSD17B13 transcript D was translated into a truncated protein with a lower molecular weight compared to HSD17B13 transcript A. Figure 13D shows that HSD17B13 IsoD protein levels were lower than IsoA protein levels. Protein levels were normalized to actin; **P<0.01.

[0080] [Figure 14] Figure 14 shows the Km and Vmax values ​​for estradiol using purified recombinant HSD17B13 protein. For Km and Vmax determination, assays were performed using 500 μM NAD+ and 228 nM HSD17B13 with a dose range of 0.2 μM to 200 μM 17β-estradiol and time points from 5 minutes to 180 minutes. Vmax and Km were then determined using the Michaelis-Menten model and Prism software (GraphPad Software, USA).

[0081] [Figure 15]Figure 15 shows the genome editing rate (%) (total number of insertions or deletions observed within a 20 base pair window on either side of Cas9-induced DNA cleavage across the total number of sequence reads in PCR reactions derived from a pool of lysed cells) at the mouse Hsd17b13 locus as determined by next-generation sequencing (NGS) in primary hepatocytes isolated from hybrid wild-type mice (75% C57BL / 6NTac, 25% 129S6 / SvEvTac). The samples tested included hepatocytes treated with a ribonucleoprotein complex containing Cas9 and a guide RNA designed to target the mouse Hsd17b13 locus.

[0082] [Figure 16] Figure 16 shows the genome editing rate (%) (total number of insertions or deletions observed across the total number of sequence reads in PCR reactions derived from a pool of lysed cells) at the mouse Hsd17b13 locus as determined by next-generation sequencing (NGS) in samples isolated from mouse liver 3 weeks after injection of AAV8 containing an sgRNA expression cassette designed to target mouse Hsd17b13 into Cas9-ready mice. Wild-type mice that do not express any Cas9 were injected with AAV8 containing all sgRNA expression cassettes as a negative control.

[0083] [Figure 17] Figures 17A and 17B show the relative mRNA expression for mouse Hsd17b13 and non-targeted HSD family members, respectively, as determined by RT-qPCR in liver samples from Cas9-ready mice treated with AAV8 carrying a guide RNA expression cassette designed to target mouse Hsd17b13. Wild-type mice that do not express any Cas9 were injected with AAV8 carrying guide RNA expression cassettes for all guide RNAs and used as a negative control. DETAILED DESCRIPTION OF THE INVENTION

[0084] definition The terms "protein," "polypeptide," and "peptide" are used interchangeably herein and include polymeric forms of amino acids of any length, including coded and non-coded amino acids, and amino acids that are chemically or biochemically modified or derivatized. The terms also include modified polymers, such as polypeptides having modified peptide backbones.

[0085] Proteins are said to have an "N-terminus" and a "C-terminus". The term "N-terminus" refers to the beginning of a protein or polypeptide, which ends with the amino acid having a free amine group (-NH2). The term "C-terminus" refers to the end of a chain of amino acids (protein or polypeptide), which ends with a free carboxyl group (-COOH).

[0086] The terms "nucleic acid" and "polynucleotide" are used interchangeably herein and include polymeric forms of nucleotides of any length, including ribonucleotides, deoxyribonucleotides, or analogs or modified versions thereof. These terms include single-, double-, and multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, and polymers containing purine bases, pyrimidine bases, or other natural, chemically modified, biochemically modified, non-natural, or derivatized nucleotide bases.

[0087] Nucleic acids are said to have a "5' end" and a "3' end" because mononucleotides react to form oligonucleotides such that the 5' phosphate of one mononucleotide pentose ring is attached unidirectionally to the adjacent 3' oxygen by a phosphodiester linkage. An end of an oligonucleotide is referred to as the "5' end" if its 5' phosphate is not linked to the 3' oxygen of a mononucleotide pentose ring. An end of an oligonucleotide is referred to as the "3' end" if its 3' oxygen is not linked to the 5' phosphate of another mononucleotide pentose ring. Nucleic acid sequences, even those internal to a larger oligonucleotide, can also be said to have 5' and 3' ends. In either a linear or circular DNA molecule, distinct elements are referred to as being "upstream" or 5' of the "downstream" or 3' element.

[0088] The term "wild-type" encompasses entities having a structure and / or activity found in a normal state or context (as opposed to a mutant, diseased, altered state or context, etc.). Wild-type genes and polypeptides often exist in many different forms (e.g., alleles).

[0089] The term "isolated," with respect to proteins and nucleic acids, includes proteins and nucleic acids that are relatively purified with respect to other bacterial, viral, or cellular components that may normally be present in situ, up to and including substantially pure preparations of proteins and polynucleotides. The term "isolated" also includes proteins and nucleic acids that have no naturally occurring counterpart, that are chemically synthesized, and thus are substantially free from other protein or nucleic acid contaminants, or that have been separated or purified from the majority of other cellular components with which they are naturally associated (e.g., other cellular proteins, polynucleotides, or cellular components).

[0090] An "exogenous" molecule or sequence includes a molecule or sequence that is not normally present in a cell in that form. Normally present includes present with respect to a particular developmental stage and environmental conditions of the cell. An exogenous molecule or sequence may include, for example, a mutant version of a corresponding endogenous sequence in a cell, or may include a sequence that corresponds to an endogenous sequence in a cell but in a different form (i.e., not in a chromosome). In contrast, an endogenous molecule or sequence includes a molecule or sequence that is normally present in that form in a particular cell at a particular developmental stage under particular environmental conditions.

[0091] The term "heterologous," when used with respect to a nucleic acid or a protein, indicates that the nucleic acid or protein contains at least two moieties that are not naturally occurring together. Similarly, when the term "heterologous" is used with respect to a promoter operably linked to a nucleic acid encoding a protein, it indicates that the promoter and the nucleic acid encoding the protein are not naturally occurring together (i.e., not operably linked in nature). For example, the term "heterologous," when used with respect to a portion of a nucleic acid or a portion of a protein, indicates that the nucleic acid or protein contains two or more subsequences that are not found in the same relationship to each other (e.g., joined) in nature. As an example, a "heterologous" region of a nucleic acid vector is a segment of nucleic acid within or attached to another nucleic acid molecule that is not found in association with that molecule in nature. For example, a heterologous region of a nucleic acid vector can contain a coding sequence flanked by sequences not found in association with the coding sequence in nature. Similarly, a "heterologous" region of a protein is a segment of amino acids within or attached to another peptide molecule that is not found in association with that other peptide molecule in nature (e.g., a fusion protein, or a protein with a tag). Similarly, the nucleic acid or protein can include a heterologous tag or a heterologous secretion or localization sequence.

[0092] The term "label" refers to a chemical moiety or protein that is directly or indirectly detectable (e.g., due to its spectral properties, conformation, or activity) when attached to a target compound. Labels may be directly detectable (fluorophores) or indirectly detectable (haptens, enzymes, or fluorophore quenchers). Such labels may be detectable by spectroscopic, photochemical, biochemical, immunochemical, or chemical means. Examples of such labels include radiolabels, which can be measured using radiation-counting devices; pigments, dyes, or other chromogens, which can be visually observed or measured using a spectrophotometer; spin labels, which can be measured using spin-label analytical instruments; and fluorescent labels (fluorophores), in which an output signal is generated by excitation of an appropriate molecular adduct and can be visualized by excitation with light absorbed by the dye or measured using a standard fluorometer or imaging system. Labels may also be, for example, chemiluminescent substances, in which an output signal is generated by chemical modification of the signal compound; metal-containing substances; or enzymes, in which enzyme-dependent secondary signal generation occurs, such as the formation of a colored product from a colorless substrate. The term "label" can also refer to a "tag" or hapten that can be selectively attached to a conjugate molecule so that the conjugate molecule can be used to generate a detectable signal when subsequently added with a substrate. For example, biotin can be used as a tag, and then an avidin or streptavidin conjugate of horseradish peroxidase (HRP) can be used to bind to the tag, and then the presence of HRP can be detected using a calorimetric substrate (e.g., tetramethylbenzidine (TMB)) or a fluorogenic substrate. The term "label" can also refer to a tag that can be used, for example, to facilitate purification. Non-limiting examples of such tags include myc, HA, FLAG or 3xFLAG, 6xHis or polyhistidine, glutathione-S-transferase (GST), maltose-binding protein, epitope tags, or the Fc portion of an immunoglobulin.Many labels are known, including, for example, particles, fluorophores, haptens, enzymes and their calorimetric, fluorogenic and chemiluminescent substrates, and other labels.

[0093] "Codon optimization" refers to the process of utilizing codon degeneracy, as indicated by the multiplicity of three-base pair codon combinations that specify amino acids, to modify a nucleic acid sequence, generally for enhanced expression in a particular host cell, by replacing at least one codon of the native sequence with a codon more frequently or most frequently used in the host cell's genes while maintaining the native amino acid sequence. For example, a polynucleotide encoding a Cas9 protein can be modified to replace a codon with a codon more frequently used in a given prokaryotic or eukaryotic cell, including bacterial cells, yeast cells, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, hamster cells, or any other host cell, compared to the naturally occurring nucleic acid sequence. Codon usage tables are readily available, for example, in the "Codon Usage Database." These tables can be adapted in several ways. See Nakamura et al. (2000) Nucleic Acids Research 28:292, incorporated herein by reference in its entirety for all purposes. Computer algorithms are also available for codon optimization of a particular sequence for expression in a particular host (see, e.g., Gene See Forge).

[0094] The term "locus" refers to the specific location of a gene (or meaningful sequence), a DNA sequence, a sequence encoding a polypeptide, or the location on a chromosome of the genome of an organism.For example, "HSD17B13 locus" can refer to the specific location of the HSD17B13 gene, the HSD17B13 DNA sequence, the sequence encoding HSD17B13, or the location of HSD17B13 on a chromosome of the genome of an organism identified as having such a sequence.The "HSD17B13 locus" can include, for example, the regulatory elements of the HSD17B13 gene, including enhancer, promoter, 5' and / or 3' UTR, or a combination thereof.

[0095] The term "gene" refers to a DNA sequence within a chromosome that encodes a product (e.g., an RNA product and / or a polypeptide product), including the coding region interrupted by one or more non-coding introns and sequences located adjacent to the coding region at both the 5' and 3' ends; thus, a gene corresponds to the full-length mRNA (including 5' and 3' untranslated sequences). The term "gene" also includes other non-coding sequences, including regulatory sequences (e.g., promoters, enhancers, and transcription factor binding sites), polyadenylation signals, internal ribosome entry sites, silencers, insulating sequences, and matrix attachment regions. These sequences may be located near (e.g., within 10 kb) or distal to the coding region of a gene and affect the level or rate of transcription and translation of the gene. The term "gene" also encompasses "minigenes."

[0096] The term "minigene" refers to a gene in which one or more non-essential segments of the gene have been deleted relative to the corresponding naturally occurring germline gene, but at least one intron remains. The deleted segment may be an intron sequence. For example, the deleted segment may be at least about 500 base pairs to several kilobases of intron sequence. Generally, intron sequences that do not encompass essential regulatory elements may be deleted. Gene segments comprising a minigene are arranged in the same linear order as they are in the germline gene, but this is not necessarily the case. Some desired regulatory elements (e.g., enhancers, silencers) may be relatively position-insensitive; therefore, regulatory elements function correctly even if they are arranged differently in the minigene than in the corresponding germline gene. For example, an enhancer may be located at a different distance from the promoter, in a different orientation, and / or in a different linear order. For example, an enhancer located 3' to the promoter in the germline configuration may be located 5' to the promoter in the minigene. Similarly, some genes may have exons that are alternatively spliced ​​at the RNA level. Thus, a minigene can have fewer exons and / or exons in a different linear order than the corresponding germline gene and still encode a functional gene product. Minigenes can also be constructed using cDNAs encoding gene products (e.g., hybrid cDNA-genomic fusions).

[0097] The term "allele" refers to variant forms of a gene. Some genes have different forms located at the same chromosomal position, or locus. Diploid organisms have two alleles at each locus. Each pair of alleles represents a genotype at a particular locus. A genotype is described as homozygous when two identical alleles are present at a particular locus, and heterozygous when the two alleles are different.

[0098] The term "variant" or "genetic variant" refers to a nucleotide sequence that differs from the sequence most prevalent in a population (e.g., differs by one nucleotide). For example, a change or substitution in a portion of a nucleotide sequence alters a codon, thus encoding a different amino acid, resulting in a genetic variant polypeptide. The term "variant" can also refer to a gene whose sequence differs from the sequence most prevalent in a population at a position that does not change the amino acid sequence of the encoded polypeptide (i.e., a conservative change). Genetic variants can be risk-associated, protective, or neutral.

[0099] A "promoter" is a regulatory region of DNA, typically containing a TATA box that can direct RNA polymerase II to initiate RNA synthesis at the appropriate transcription start site for a particular polynucleotide sequence. A promoter may further contain other regions that affect the rate of transcription initiation. The promoter sequences disclosed herein modulate the transcription of an operably linked polynucleotide. The promoter may be active in one or more of the cell types disclosed herein (e.g., eukaryotic cells, non-human mammalian cells, human cells, rodent cells, pluripotent cells, differentiated cells, or combinations thereof). The promoter may be, for example, a constitutively active promoter, a conditional promoter, an inducible promoter, a temporally restricted promoter (e.g., a developmentally regulated promoter), or a spatially restricted promoter (e.g., a cell-specific or tissue-specific promoter). Examples of promoters can be found, for example, in WO2013 / 176772, which is incorporated herein by reference in its entirety for all purposes.

[0100] Examples of inducible promoters include, for example, chemically controlled promoters and physically controlled promoters. Chemically controlled promoters include, for example, alcohol-controlled promoters (e.g., alcohol dehydrogenase (alcA) gene promoter), tetracycline-controlled promoters (e.g., tetracycline-responsive promoter, tetracycline operator sequence (tetO), tet-On promoter, or tet-Off promoter), steroid-controlled promoters (e.g., rat glucocorticoid receptor, estrogen receptor promoter, or ecdysone receptor promoter), or metal-controlled promoters (e.g., metalloprotein promoters). Physically controlled promoters include, for example, temperature-controlled promoters (e.g., heat shock promoters) and light-controlled promoters (e.g., light-inducible promoters or light-repressible promoters).

[0101] The tissue-specific promoter can be, for example, a neuronal-specific promoter, a glial-specific promoter, a muscle cell-specific promoter, a cardiac cell-specific promoter, a kidney cell-specific promoter, a bone cell-specific promoter, an endothelial cell-specific promoter, or an immune cell-specific promoter (e.g., a B cell promoter or a T cell promoter).

[0102] Developmentally-regulated promoters include, for example, promoters that are active only during the embryonic stage of development or only in adult cells.

[0103] "Operable linkage" or "operably linked" includes proximity of two or more components (e.g., a promoter and another sequence element) that permits the possibility that both components can function normally and that at least one of the components can mediate a function exerted by at least one of the other components. For example, a promoter may be operably linked to a coding sequence if it regulates the level of transcription of the coding sequence in response to the presence or absence of one or more transcriptional regulatory factors. Operable linkage can include sequences that are contiguous or trans-acting with each other (e.g., regulatory sequences can act at a distance to regulate transcription of a coding sequence).

[0104] The term "primer" refers to an oligonucleotide capable of acting as a point of initiation for polynucleotide synthesis along a complementary strand when placed under conditions that catalyze the synthesis of a primer extension product complementary to the polynucleotide. Such conditions include the presence of four different nucleotide triphosphates or nucleoside analogs and one or more polymerization agents, such as DNA polymerase and / or reverse transcriptase, in a suitable buffer (containing cofactors or substituents that affect pH, ionic strength, etc.) at a suitable temperature. Sequence-specific primer extension can include, for example, methods of PCR, DNA sequencing, DNA extension, DNA polymerization, RNA transcription, or reverse transcription. A primer must be long enough to prime the synthesis of an extension product in the presence of an agent in lieu of polymerase. Typical primers are sequences at least about 5 nucleotides in length that are substantially complementary to the target sequence, although longer primers are preferred. Generally, primers are about 15-30 nucleotides in length, although longer primers can also be used. The primer sequence need not be exactly complementary to the template or target sequence, but must be sufficiently complementary to hybridize with the template or target sequence. The term "primer pair" refers to a set of primers comprising a 5' upstream primer that hybridizes with the 5' end of the DNA sequence to be amplified and a 3' downstream primer that hybridizes with the complement of the 3' end of the sequence to be amplified. Primer pairs can be used to amplify target polynucleotides (e.g., by polymerase chain reaction (PCR) or other conventional nucleic acid amplification methods). "PCR" or "polymerase chain reaction" is a technique used to amplify specific DNA segments (see U.S. Pat. Nos. 4,683,195 and 4,800,159, each of which is incorporated herein by reference in its entirety for all purposes).

[0105] The term "probe" refers to a molecule capable of detectably distinguishing between structurally different target molecules. Detection can be achieved in a variety of different ways depending on the type of probe and the type of target molecule used. Thus, for example, detection can be based on distinguishing the activity level of the target molecule, but is preferably based on detecting specific binding. Examples of such specific binding include antibody binding and nucleic acid probe hybridization. Thus, probes can include, for example, enzyme substrates, antibodies and antibody fragments, and nucleic acid hybridization probes. For example, a probe can be an isolated polynucleotide attached to a conventional detectable label or reporter molecule, such as a radioisotope, a ligand, a chemiluminescent agent, or an enzyme. Such a probe is complementary to a target polynucleotide strand, such as a polynucleotide containing the HSD17B13 rs72613567 variant or a specific HSD17B13 mRNA transcript. Deoxyribonucleic acid probes can be HSD17B13-mRNA / cDNA-specific primers or HSD17B13-rs72613567-specific primers, in These probes may include oligonucleotide probes synthesized in vitro or generated by PCR using DNA from bacterial artificial chromosome, fosmid, or cosmid libraries. Probes include not only deoxyribonucleic acid or ribonucleic acid, but also polyamides and other probe materials capable of specifically detecting the presence of target DNA sequences. For nucleic acid probes, detection reagents can include, for example, radiolabeled probes, enzyme-labeled probes (e.g., horseradish peroxidase and alkaline phosphatase), affinity-labeled probes (e.g., biotin, avidin, and streptavidin), and fluorescently labeled probes (e.g., 6-FAM, VIC, TAMRA, MGB, fluorescein, rhodamine, and Texas Red). The nucleic acid probes described herein can be easily incorporated into one of many well-known, established kit formats.

[0106] The term "antisense RNA" refers to a single-stranded RNA that is complementary to a messenger RNA strand transcribed in a cell.

[0107] The term "small interfering RNA (siRNA)" refers to a generally double-stranded RNA molecule that induces the RNA interference (RNAi) pathway. These molecules can vary in length (generally 18-30 base pairs) and contain varying degrees of complementarity to their target mRNA in the antisense strand. Some, but not all, siRNAs have unpaired overhanging bases at the 5' or 3' end of the sense and / or antisense strand. The term "siRNA" encompasses duplexes of two separate strands as well as single strands that can form hairpin structures containing the duplex region. The double-stranded structure can be, for example, less than 20, 25, 30, 35, 40, 45, or 50 nucleotides in length. For example, the double-stranded structure can be about 21-23 nucleotides in length, about 19-25 nucleotides in length, or about 19-23 nucleotides in length.

[0108] The term "short hairpin RNA (shRNA)" refers to a single strand of RNA base that can self-hybridize into a hairpin structure and induce the RNA interference (RNAi) pathway upon processing. These molecules can vary in length (generally about 50-90 nucleotides in length, or in some cases, for example, for microRNA-adapted shRNAs, up to more than 250 nucleotides in length). shRNA molecules are processed intracellularly to form siRNAs, which in turn can knock down gene expression. shRNAs can be incorporated into vectors. The term "shRNA" also refers to a DNA molecule from which a short hairpin RNA molecule can be transcribed.

[0109] "Complementarity" of nucleic acids refers to the ability of a nucleotide sequence in one strand of a nucleic acid to form hydrogen bonds with another sequence on an opposing strand due to the orientation of its nucleobase groups. In DNA, complementary bases are typically A-T and C-G. In RNA, complementary bases are typically C-G and U-A. Complementarity can be complete or substantial / sufficient. Complete complementarity between two nucleic acids means that the two nucleic acids can form a duplex in which every base in the duplex binds with a complementary base through Watson-Crick pairing. "Substantial" or "sufficient" complementarity means that the sequence in one strand is not completely and / or completely complementary to the sequence in the opposing strand, but sufficient binding occurs between the bases on the two strands under a set of hybridization conditions (e.g., salt concentration and temperature) to form a stable hybrid complex. Such conditions can be predicted by using sequences and standard mathematical calculations to predict the Tm (melting temperature) of hybridized strands, or by empirically determining the Tm using conventional methods. Tm comprises the temperature at which the population of hybridization complexes formed between two nucleic acid strands is 50% denatured (i.e., the population of double-stranded nucleic acid molecules is half-dissociated into single strands). Temperatures below Tm favor the formation of hybridization complexes, while temperatures above Tm favor melting or separation of strands within the hybridization complexes. For nucleic acids with known G+C content in aqueous 1M NaCl solution, Tm can be estimated, for example, by using Tm = 81.5 + 0.41 (% G+C), although other known Tm computations take into account the structural properties of nucleic acids.

[0110] "Hybridization conditions" include the cumulative environment in which one nucleic acid strand binds to a second nucleic acid strand through complementary strand interactions and hydrogen bonds to form a hybridization complex. Such conditions include the chemical components and their concentrations (e.g., salts, chelating agents, formamide) of the aqueous or organic solution containing the nucleic acid, as well as the temperature of the mixture. Other factors, such as the length of incubation time or the dimensions of the reaction chamber, may contribute to the environment. See, for example, Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd Edition, pp. 1.90-1.91, 9.47-1.98, which is incorporated herein by reference in its entirety for all purposes. See pages 9.51, 11.47-11.57 (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989).

[0111] Hybridization requires that two nucleic acids contain complementary sequences, although mismatches between bases are possible. Appropriate conditions for hybridization between two nucleic acids depend on the well-known variables of the length of the nucleic acids and the degree of complementarity. The greater the degree of complementarity between two nucleotide sequences, the greater the melting temperature (Tm) of hybrids of nucleic acids having those sequences. For hybridization between nucleic acids having short stretches of complementarity (e.g., complementarity over 35 nucleotides or less, 30 nucleotides or less, 25 nucleotides or less, 22 nucleotides or less, 20 nucleotides or less, or 18 nucleotides or less), the position of mismatches becomes important (see Sambrook et al., supra, pp. 11.7-11.8). Generally, the length of a hybridizable nucleic acid is at least about 10 nucleotides. Exemplary minimum lengths for a hybridizable nucleic acid include at least about 15 nucleotides, at least about 20 nucleotides, at least about 22 nucleotides, at least about 25 nucleotides, and at least about 30 nucleotides. Additionally, the temperature and salt concentration of the wash solutions can be adjusted as needed depending on factors such as the length of the complementary region and the degree of complementation.

[0112] The sequence of a polynucleotide need not be 100% complementary to the polynucleotide sequence of its target nucleic acid to be specifically hybridizable. Furthermore, a polynucleotide may hybridize across one or more segments such that intervening or adjacent segments are not involved in the hybridization event (e.g., a loop or hairpin structure). A polynucleotide (e.g., a gRNA) may contain at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence complementarity to a target region within the targeted target nucleic acid sequence. For example, a gRNA in which 18 of 20 nucleotides are complementary to a target region and therefore specifically hybridizes thereto would be 90% complementary. In this example, the remaining non-complementary nucleotides may be clustered or interspersed with complementary nucleotides and need not be contiguous with each other or with complementary nucleotides.

[0113] The percent complementarity between specific stretches of nucleic acid sequences within a nucleic acid can be determined using the BLAST program (basic local alignment search tool) and the Power BLAST program (Altschul et al. (1990) J. Mol. Biol. 215:403-410; Zhang and Madden (1997) Genome Res. 7:649-656), or by the algorithm of Smith and Waterman (Adv. Appl. Math. 1981, 2:482-489). It can be routinely determined by using the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.) using the default settings, which uses a nucleotide sequence analysis algorithm.

[0114] In the method and composition presented herein, various different components are used.Throughout the description, some components can have active variants and fragments.Such components include, for example, Cas9 protein, CRISPR RNA, tracrRNA and guide RNA.The biological activity of each of these components is described elsewhere herein.

[0115] "Sequence identity" or "identity," in the context of two polynucleotide or polypeptide sequences, refers to the residues in the two sequences that are the same when aligned for maximum correspondence over a specified comparison window. When percentage sequence identity is used with respect to proteins, non-identical residue positions often differ by conservative amino acid substitutions, in which an amino acid residue is replaced with another amino acid residue with similar chemical properties (e.g., charge or hydrophobicity), thereby not altering the functional properties of the molecule. When sequences differ by conservative substitutions, the percent sequence identity can be adjusted upward to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity." Methods for making this adjustment are well known. Generally, this involves scoring conservative substitutions as partial rather than complete mismatches, thereby increasing the percentage sequence identity. Thus, for example, where identical amino acids are assigned a score of 1 and non-conservative substitutions are assigned a score of zero, conservative substitutions are assigned a score between zero and 1. Conservative substitution scores are calculated, for example, as implemented in the program PC / GENE (Intelligenetics, Mountain View, California).

[0116] "Percentage of sequence identity" includes a value determined by comparing two optimally aligned sequences (maximum number of perfectly matched residues) over a comparison window, where a portion of the polynucleotide sequence within the comparison window may contain additions or deletions (i.e., gaps) compared to the reference sequence (which does not contain additions or deletions) due to optimal alignment of the two sequences. The percentage is calculated by determining the number of positions in both sequences where the same nucleic acid base or amino acid residue occurs to obtain the number of matching positions, dividing the number of matching positions by the total number of positions within the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. Unless otherwise specified (e.g., the shorter sequence includes a linked heterologous sequence), the comparison window is the entire length of the shorter of the two sequences being compared.

[0117] Unless otherwise specified, sequence identity / similarity values ​​include those obtained using GAP version 10 with the following parameters: % identity and % similarity for nucleotide sequences using a gap weight of 50 and a length weight of 3, and the nwsgapdna.cmp scoring matrix; % identity and % similarity for amino acid sequences using a gap weight of 8 and a length weight of 2, and the BLOSUM62 scoring matrix; or any equivalent program. "Equivalent program" includes any sequence comparison program that produces alignments with identical nucleotide or amino acid residue matches and identical percent sequence identity for any two sequences at issue compared to corresponding alignments produced by GAP version 10.

[0118] The term "conservative amino acid substitution" refers to the substitution of an amino acid normally present in a sequence with a different amino acid of similar size, charge, or polarity. Examples of conservative substitutions include the substitution of a non-polar (hydrophobic) residue such as isoleucine, valine, or leucine for another non-polar residue. Similarly, examples of conservative substitutions include the substitution of one polar (hydrophilic) residue for another polar (hydrophilic) residue, such as the substitution between arginine and lysine, the substitution between glutamine and asparagine, or the substitution between glycine and serine. Furthermore, the substitution of a basic residue such as lysine, arginine, or histidine for another basic residue, or the substitution of one acidic residue such as aspartic acid or glutamic acid for another acidic residue are further examples of conservative substitutions. Examples of non-conservative substitutions include substitutions of a polar (hydrophilic) residue such as cysteine, glutamine, glutamic acid or lysine with a nonpolar (hydrophobic) amino acid residue such as isoleucine, valine, leucine, alanine, or methionine, and / or substitutions of a nonpolar residue with a polar residue. Typical amino acid categorizations are summarized below. [Table A]

[0119] A subject nucleic acid, such as a primer or guide RNA, hybridizes to or targets or includes a position that is adjacent to or within about 1000, 500, 400, 300, 200, 100, 50, 45, 40, 35, 30, 25, 20, 15, 10, or 5 nucleotides from a specified nucleotide position in a reference nucleic acid.

[0120] The term "biological sample" refers to a sample of biological material within or available from a subject from which nucleic acids or proteins can be recovered. The term biological sample can also encompass any material obtained by processing the sample, such as cells or their progeny. Processing of a biological sample can involve one or more of filtration, distillation, extraction, concentration, fixation, inactivation of interfering components, etc. In some embodiments, a biological sample contains nucleic acids such as genomic DNA, cDNA, or mRNA. In some embodiments, a biological sample contains proteins. The subject can be any organism, including, for example, a human, a non-human mammal, a rodent, a mouse, or a rat. A biological sample can be derived from any cell, tissue, or biological fluid from a subject. The sample can include any clinically relevant tissue, such as a bone marrow sample, a tumor biopsy, a fine-needle aspirate, or a bodily fluid sample such as blood, plasma, serum, lymph, ascites, cyst fluid, or urine. In some cases, the sample includes a buccal swab. The samples used in the methods disclosed herein will vary based on the assay format, nature of the detection method and the tissues, cells, or extracts used as the sample.

[0121] The term " control sample " refers to the sample obtained from the subject who does not have HSD17B13 rs72613567 variant, preferably is homozygous for the wild-type allele of HSD17B13 gene.Such sample can be obtained at the same time as biological sample, or can be obtained at different times.Both biological sample and control sample can be obtained from the same tissue or body fluid.

[0122] A "homologous" sequence (e.g., a nucleic acid sequence) includes a sequence that is identical or substantially similar to a known reference sequence; thus, a homologous sequence is, for example, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the known reference sequence. Homologous sequences can include, for example, orthologous and paralogous sequences. Homologous genes, for example, are generally inherited from a common ancestral DNA sequence through either speciation events (orthologous genes) or gene duplication events (paralogous genes). "Orthologous" genes include genes from different species that have evolved from a common ancestral gene by speciation. Orthologs generally retain the same function during evolution. "Paralogous" genes include genes related by duplication within a genome. Paralogs can evolve new functions during evolution.

[0123] The term "in vitro" includes an artificial environment and processes or reactions that occur within an artificial environment (e.g., a test tube). The term "in vivo" includes a natural environment (e.g., a cell or organism or body, e.g., a cell within an organism or body) and processes or reactions that occur within a natural environment. The term "ex vivo" includes cells removed from an individual's body and processes or reactions that occur within such cells.

[0124] A composition or method that "comprising" or "including" one or more recited elements may include other elements not specifically recited. For example, a composition that "comprises" or "includes" a protein may contain the protein alone or in combination with other ingredients. The transitional phrase "consisting essentially of" means that the claims should be interpreted to include the specified elements recited in the claims as well as elements that do not materially affect the basic and novel characteristic(s) of the claimed invention. Thus, the term "consisting essentially of," when used in the claims of the present invention, is not intended to be interpreted as equivalent to "comprising."

[0125] "Optional" or "optionally" means that the subsequently described event or circumstance may or may not occur, and that the description includes instances in which the event or circumstance occurs and instances in which the event or circumstance does not occur.

[0126] The specification of a range of values ​​includes all integers within or defining the range, and all subranges defined by the integers within the range.

[0127] Unless the context makes clear otherwise, the term "about" encompasses values ​​within the standard measurement error limits (eg, SEM) of the stated value.

[0128] The term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items, as well as no combinations when interpreted as alternatives ("or").

[0129] The term "or" refers to any one member of a particular list and also includes any combination of members of that list.

[0130] The singular articles "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, the terms "a Cas9 protein" or "at least one Cas9 protein" can include multiple Cas9 proteins, including mixtures thereof.

[0131] Statistically significant means p≦0.05.

[0132] Detailed Description I. Overview The present paper presents HSD17B13 variant that is found to be associated with the reduction of alanine and aspartate transaminase levels; the reduction of the risk of chronic liver disease, including non-alcoholic and alcoholic fatty liver disease, cirrhosis and hepatocellular carcinoma; and the reduction of the progression from simple steatosis to more clinically advanced stage of chronic liver disease.The present paper also presents the HSD17B13 gene transcript that is associated with variant that has not been previously identified.

[0133] The present paper provides the nucleic acid and protein related to the variant of HSD17B13, and the cells containing these nucleic acids and proteins.Also provided is the method for modifying cells by using any combination of nuclease agent, exogenous donor sequence, transcription activator, transcription repressor and expression vector to express the recombinant HSD17B13 gene or nucleic acid that encodes HSD17B13 protein.Also provided is the therapeutic and prophylactic method for treating the subject who has chronic liver disease or is at risk of developing it.

[0134] II. HSD17B13 variants The present specification provides isolated nucleic acids and proteins related to variants of HSD17B13 (also known as hydroxysteroid 17-beta dehydrogenase 13, 17-beta-hydroxysteroid dehydrogenase 13, 17β-hydroxysteroid dehydrogenase-13, 17β-HSD13, short-chain dehydrogenase / reductase 9, SCDR9, HMFN0376, NIIL497, and SDR16C3).The human HSD17B13 gene is approximately 19 kb long and contains 7 exons and 6 introns, located at 4q22.1 in the genome. An exemplary human HSD17B13 protein sequence is assigned UniProt accession number Q7Z5P4 (SEQ ID NOs: 240 and 241; Q7Z5P4-1 and Q7Z5P4-2, respectively) and NCBI Reference Sequence numbers NP_835236 and NP_001129702 (SEQ ID NOs: 242 and 243, respectively). Exemplary human HSD17B13 mRNA is assigned NCBI Reference Sequence numbers NM_178135 and NM_001136230 (SEQ ID NOs: 244 and 245, respectively).

[0135] In particular, a splice HSD17B13 variant (rs72613567) is presented herein, in which an adenine is inserted adjacent to the donor splice site in intron 6. The adenine is inserted in the forward (plus) strand of the chromosome, which corresponds to a thymine inserted in the reverse (minus) strand of the chromosome. Because the human HSD17B13 gene is transcribed in reverse, this nucleotide insertion is reflected as a thymine insertion in the exemplary HSD17B13 rs72613567 variant sequence presented as SEQ ID NO: 2 relative to the exemplary wild-type HSD17B13 gene sequence presented as SEQ ID NO: 1. Therefore, this insertion is referred to herein as a thymine insertion between positions 12665 ​​and 12666 of SEQ ID NO: 1 or at position 12666 of SEQ ID NO: 2.

[0136] Two mRNA transcripts (A and B; SEQ ID NOS: 4 and 5, respectively) were previously identified as expressed in subjects with wild-type HSD17B13 genes. Transcript A contains all seven exons of the HSD17B13 gene, while transcript B skips exon 2. Transcript A is the predominant transcript in wild-type subjects. However, six additional, previously unidentified, expressed HSD17B13 transcripts (C-H; SEQ ID NOS: 6-11, respectively) are presented herein. These transcripts are shown in Figure 4. Transcript C skips exon 6 compared to transcript A. Transcript D contains a guanine inserted 3' of exon 6, resulting in a frameshift in exon 7 and a premature truncation of exon 7 compared to transcript A. Transcript E contains an additional exon between exons 3 and 4 compared to transcript A. Transcript F, which is expressed only in HSD17B13 rs72613567 variant carriers, has a readthrough from exon 6 to intron 6 compared to transcript A. Transcript G skips exon 2 and has a guanine inserted 3' of exon 6, resulting in a frameshift in exon 7 and a premature truncation of exon 7 compared to transcript A. Transcript H has an additional exon between exons 3 and 4 and a guanine inserted 3' of exon 6, resulting in a frameshift in exon 7 and a premature truncation of exon 7 compared to transcript A. Transcripts C, D, F, G, and H are predominant in HSD17B13 rs72613567 variant carriers, and transcript D is predominant in HSD17B13 This is the most abundant transcript in carriers of the rs72613567 variant. One additional previously unidentified HSD17B13 transcript (F', SEQ ID NO: 246) that is expressed at low levels is also presented herein. Similar to transcript F, transcript F' also contains a read-through from exon 6 to intron 6 compared to transcript A, but in contrast to transcript F, this read-through does not contain the thymine insertion present in the HSD17B13 rs72613567 variant gene. The nucleotide positions of the exons within the HSD17B13 gene for each transcript are presented below.

[0137] [Table B]

[0138] [Table C]

[0139] As described in more detail elsewhere herein, the HSD17B13 rs72613567 variant is associated with reduced alanine and aspartate transaminase levels and a reduced risk of chronic liver disease, including nonalcoholic and alcoholic fatty liver disease, cirrhosis, and hepatocellular carcinoma. The HSD17B13 rs72613567 variant is also associated with a reduced progression from simple steatosis to more clinically advanced stages of chronic liver disease.

[0140] A. Nucleic acid The present paper discloses the isolated nucleic acid related to HSD17B13 variant and variant HSD17B13 transcription product.Also disclosed is the isolated nucleic acid that hybridizes with any of the nucleic acids disclosed herein under stringent or moderate conditions.Such nucleic acid can be used for expressing HSD17B13 variant protein, or as primer, probe, exogenous donor sequence, guide RNA, antisense RNA, shRNA and siRNA, each of which is described in detail elsewhere herein.

[0141] Also disclosed are functional nucleic acids that can interact with the disclosed polynucleotides.Functional nucleic acids are nucleic acid molecules that have specific functions, such as binding to target molecules or catalyzing specific reactions.Examples of functional nucleic acids include antisense molecules, aptamers, ribozymes, triplex-forming molecules, and external guide sequences.Functional nucleic acid molecules can act as effectors, inhibitors, modulators, and stimulators of the specific activity of target molecules, or functional nucleic acid molecules can have new activities that are independent of any other molecules.

[0142] Antisense molecules are designed to interact with target nucleic acid molecules through either canonical or non-canonical base pairing. The interaction between the antisense molecule and the target molecule is designed to promote the destruction of the target molecule, for example, by RNase H-mediated RNA-DNA hybrid degradation. Alternatively, antisense molecules are designed to interfere with the processing functions normally performed on the target molecule, such as transcription or replication. Antisense molecules can be designed based on the sequence of the target molecule. There are many methods for optimizing antisense efficiency by finding the most accessible region of the target molecule. Exemplary methods are in vitro selection experiments and DNA modification tests using DMS and DEPC. Antisense molecules are generally designed to bind to the target molecule within 10 -6 Less than or equal to the dissociation constant (k d ), 10 -8 Less than or equal to the dissociation constant (kd ), 10 -10 Less than or equal to the dissociation constant (k d ), or 10 -12 Less than or equal to the dissociation constant (k d A representative sample of methods and techniques to aid in the design and use of antisense molecules can be found in the following non-limiting list of U.S. patents: U.S. Patent Nos. 5,135,917; 5,294,533; 5,627,158; 5,641,754; 5,691,317; 5,780,607; 5,786,138; 5,849,903; 5,856,103; 5,919,772; 5,955,590; 5,9 90,088; 5,994,320; 5,998,602; 6,005,095; 6,007,995; 6,013,522; 6,017,898; 6,018,042; 6,025,198; 6,033,910; 6,040,296; 6,046,004; 6,046,319; and 6,057,437, each of which is incorporated herein by reference in its entirety for all purposes.Examples of antisense molecules include antisense RNA, small interfering RNA (siRNA) and short hairpin RNA (shRNA), which are described in more detail elsewhere herein.

[0143] The isolated nucleic acid disclosed herein can comprise RNA, DNA, or both RNA and DNA.The isolated nucleic acid can also be linked or fused with a heterologous nucleic acid sequence, for example, in a vector, or with a heterologous label.For example, the isolated nucleic acid disclosed herein can be in a vector or exogenous donor sequence containing the isolated nucleic acid and a heterologous nucleic acid sequence.The isolated nucleic acid can also be linked or fused with a heterologous label, such as a fluorescent label.Other examples of labels are disclosed elsewhere herein.

[0144] The disclosed nucleic acid molecules may be composed of non-natural or modified nucleotides, such as nucleotides or nucleotide analogs or nucleotide substitutes. Such nucleotides include nucleotides containing modified bases, sugars, or phosphate groups, or nucleotides incorporating non-natural moieties into their structure. Examples of non-natural nucleotides include dideoxynucleotides, biotinylated nucleotides, aminated nucleotides, deaminated nucleotides, alkylated nucleotides, benzylated nucleotides, and fluorophore-labeled nucleotides.

[0145] The nucleic acid molecules disclosed herein may contain one or more nucleotide analogs or substitutions. A nucleotide analog is a nucleotide that contains some type of modification in either the base moiety, sugar moiety, or phosphate moiety. Modifications to the base moiety include natural and synthetic modifications of A, C, G, and T / U, as well as different purine or pyrimidine bases, such as pseudouridine, uracil-5-yl, hypoxanthin-9-yl (I), and 2-aminoadenin-9-yl. Modified bases include, for example, 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyluracil and cytosine, 6-azouracil, cytosine and thymine, 5-uracil, cytosine, and thymine. uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3-deazaguanine and 3-deazaadenine. Additional base modifications can be found, for example, in U.S. Pat. No. 3,687,808; Englisch et al. (1991) Angewandte Chemie, International Edition, Vol. 30:613; and Sanghvi, YS, Chapter 15, Antisense Research and Applications, pp. 289-302, Crooke, ST and Lebleu, B. (eds.), CRC Press, 1993, each of which is incorporated herein by reference in its entirety for all purposes.Certain nucleotide analogues, such as 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6 and O-6 substituted purines, including 2-aminopropyladenine, 5-propynyluracil, 5-propynylcytosine and 5-methylcytosine, can increase the stability of duplex formation.In many cases, base modification can be combined with sugar modification, such as 2'-O-methoxyethyl, to achieve unique properties such as increasing duplex stability. There are numerous U.S. patents, such as U.S. Patent Nos. 4,845,205; 5,130,302; 5,134,066; 5,175,273; 5,367,066; 5,432,272; 5,457,187; 5,459,255; 5,484,908; 5,502,177; 5,525,711; 5,552,540; 5,587,469; 5,594,121; 5,596,091; 5,614,617; and 5,681,941, which detail and describe a range of base modifications. each of which is incorporated herein by reference in its entirety for all purposes.

[0146] Nucleotide analogs may also contain modifications of the sugar moiety. Modifications to the sugar moiety can include, for example, natural and synthetic modifications to ribose and deoxyribose. Sugar modifications include, for example, the following modifications at the 2' position: OH; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S-, or N-alkynyl; or O-alkyl-O-alkyl, where the alkyl, alkenyl, and alkynyl can be substituted or unsubstituted C1-C10 alkyl or C2-C10 alkenyl and alkynyl. Exemplary 2' sugar modifications include, for example, -O[(CH2) n O]m CH3, -O(CH2) n OCH3, -O(CH2) n NH2, -O(CH2) n CH3, -O(CH2) n -ONH2, and -O(CH2) n ON[(CH2) nAlso included are aryl groups such as aryl, aryl, aryl- ...

[0147] Other modifications at the 2' position include, for example, C1 to C 10 Examples of suitable sugars include lower alkyl, substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleaving groups, reporter groups, intercalators, groups for improving the pharmacokinetic or pharmacodynamic properties of oligonucleotides, and other substituents with similar properties. Similar modifications can also be made to other positions on the sugar, particularly the 3' position of the sugar of the 3'-terminal nucleotide or 2'-5'-linked oligonucleotides and the 5' position of the 5'-terminal nucleotide. Modified sugars can also include sugars containing modifications to the bridging ring oxygen, such as CH2 and S. Nucleotide sugar analogs can also have sugar mimetics, such as a cyclobutyl moiety, in place of the pentofuranosyl sugar. Numerous United States patents teach the preparation of such modified sugar structures, for example, U.S. Pat. Nos. 4,981,957; 5,118,800; 5,319,080; 5,359,044; 5,393,878; 5,446,137; 5,466,786; 5,511,147; each of which is incorporated herein by reference in its entirety for all purposes. Nos. 4,785; 5,519,134; 5,567,811; 5,576,427; 5,591,722; 5,597,909; 5,610,300; 5,627,053; 5,639,873; 5,646,265; 5,658,873; 5,670,633; and 5,700,920.

[0148] Nucleotide analogues may be modified at the phosphate moiety.Modified phosphate moieties include, for example, phosphate moieties modified to contain phosphorothioate, chiral phosphorothioate, phosphorodithioate, phosphotriester, aminoalkylphosphotriester, methyl and other alkyl phosphonates, including 3'-alkylene phosphonate and chiral phosphonate, phosphinate, phosphoramidate, including 3'-aminophosphoramidate and aminoalkylphosphoramidate, thionophosphoramidate, thionoalkylphosphonate, thionoalkylphosphotriester, and boranophosphate.These phosphate or modified phosphate linkages between two nucleotides may be 3'-5' or 2'-5' linkages, and the linkage may contain reversed direction from 3'-5' to 5'-3' or from 2'-5' to 5'-2'.Various salts, mixed salts, and free acid forms are also included. Numerous United States patents teach how to make and use nucleotides containing modified phosphates, including, for example, U.S. Pat. Nos. 3,687,808; 4,469,863; 4,476,301; 5,023,243; 5,177,196; 5,188,897; 5,264,423; 5,276,019; 5,278,302; and 5,278,302, each of which is incorporated by reference in its entirety for all purposes. Nos. 5,286,717; 5,321,131; 5,399,676; 5,405,939; 5,453,496; 5,455,233; 5,466,677; 5,476,925; 5,519,126; 5,536,821; 5,541,306; 5,550,111; 5,563,253; 5,571,799; 5,587,361; and 5,625,050.

[0149] Nucleotide substitutes include molecules that have similar functional properties to nucleotides but do not contain a phosphate moiety, such as peptide nucleic acids (PNAs). Nucleotide substitutes also include molecules that recognize nucleic acids in a Watson-Crick or Hoogsteen manner but link through a moiety other than the phosphate moiety. Nucleotide substitutes can conform to a double-helix structure when interacting with the appropriate target nucleic acid.

[0150] Nucleotide substitutes also include nucleotides or nucleotide analogs in which the phosphate or sugar moiety is replaced. Nucleotide substitutes may not contain a standard phosphorus atom. Substitutes for phosphate may be, for example, short-chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short-chain heteroatom or heterocyclic internucleoside linkages. These include morpholino linkages (formed in part from the sugar portion of the nucleoside); siloxane backbones; sulfide, sulfoxide, and sulfone backbones; formacetyl and thioformacetyl backbones; methyleneformacetyl and thioformacetyl backbones; alkene-containing backbones; sulfamic acid backbones; methyleneimino and methylenehydrazino backbones; sulfonic acid and sulfonamide backbones; amide backbones; and others with mixed N, O, S, and CH2 moieties. Numerous United States patents disclose how to make and use these types of phosphate substitutions, including, but not limited to, U.S. Patent Nos. 5,034,506; 5,166,315; 5,185,444; 5,214,134; 5,216,141; 5,235,033; 5,264,562; 5,264,564; 5,405,938; and 5,434,257, each of which is incorporated by reference in its entirety for all purposes. ; 5,466,677; 5,470,967; 5,489,677; 5,541,307; 5,561,225; 5,596,086; 5,602,240; 5,610,289; 5,602,240; 5,608,046; 5,610,289; 5,618,704; 5,623,070; 5,663,312; 5,633,360; 5,677,437; and 5,677,439.

[0151] It is also understood that in nucleotide substitutes, both the sugar moiety and the phosphate moiety of nucleotide can be replaced, for example, by amide-type linkage (aminoethylglycine) (PNA).U.S. Patent No. 5,539,082; U.S. Patent No. 5,714,331; and U.S. Patent No. 5,719,262 teach how to make and use PNA molecules, and each is incorporated herein by reference in its entirety for all purposes.See also Nielsen et al. (1991) Science, 254:1497-1500, which is incorporated herein by reference in its entirety for all purposes.

[0152] Other types of molecules (conjugates) can also be linked to nucleotides or nucleotide analogs, for example, to enhance cellular uptake. The conjugates can be chemically linked to the nucleotide or nucleotide analog. Such conjugates include, for example, lipid moieties, such as cholesterol moieties (Letsinger et al., 2001). 989) Proc. Natl. Acad. Sci. USA 86:6553-6556, incorporated herein by reference in its entirety for all purposes), cholic acid (Manoharan et al. (1994) Bioorg. Med. Chem. Let. 4:1053-1060, incorporated herein by reference in its entirety for all purposes), thioethers such as hexyl-S-tritylthiol (Manoharan et al. (1992) Ann. NY Acad. Sci. 660:306-309; Manoharan et al. (1993) Bioorg. Med. Chem. Let. 3 2765-2770, incorporated herein by reference in its entirety for all purposes), thiocholesterol (Oberhauser et al. (1992) Nucl. Acids Res. 20:533-538, incorporated herein by reference in its entirety for all purposes), aliphatic chains such as dodecanediol or undecyl residues (Saison-Behmoaras et al. (1991) EMBO J. 10:1111-1118; Kabanov et al. (1990) FEBS Lett. 259:327-330; Svinarchuk et al. (1993) Biochimie 7 5:49-54, each of which is incorporated herein by reference in its entirety for all purposes), phospholipids such as di-hexadecyl-rac-glycerol or triethylammonium 1,2-di-O-hexadecyl-rac-glycero-3-H-phosphonate (Manoharan et al. (1995) Tetrahedron Lett. 36:3651-365 4; Shea et al. (1990) Nucl. Acids Res. 18:3777-3783, each of which is incorporated herein by reference in its entirety for all purposes), polyamine or polyethylene glycol chains (Manoharan et al. (1995) Nucleosides & Nucleotides 14:969-973, each of which is incorporated herein by reference in its entirety for all purposes), or adamantane acetic acid (Manoharan et al. (1995) Tetrahedron Lett. 36:3651-3654, each of which is incorporated herein by reference in its entirety for all purposes). , incorporated herein by reference), a palmityl moiety (Mishra et al. (1995) Biochim. Biophys. Acta 1264:229-237), or an octadecylamine or hexylamino-carbonyl-oxycholesterol moiety (Crooke et al. (1996) J. Pharmacol. Exp. Ther. 277:923-937, incorporated herein by reference in its entirety for all purposes). (which are incorporated herein by reference). Numerous U.S. patents teach the preparation of such conjugates, including, for example, U.S. Patent Nos. 4,828,979; 4,948,882; 5,218,105; 5,525,465; 5,541,313; 5,545,730; 5,552,538; 5,578,717; 5,580,731; 5,580,731; each of which is incorporated herein by reference in its entirety for all purposes. Same No. 5,591,584; Same No. 5,109,124; Same No. 5,118,802; Same No. 5,138,045; Same No. 5,414,077; Same No. 5,486,603; Same No. 5,512,439; Same No. 5,578,718; Same No. 5 , 608,046; 4,587,044; 4,605,735; 4,667,025; 4,762,779; 4,789,737; 4,824,941; 4,835,263; 4,876 ,335; 4,904,582; 4,958,013; 5,082,830; 5,112,963; 5,214,136; 5,082,830; 5,112,963; 5,214,13 No. 6; No. 5,245,022; No. 5,254,469; No. 5,258,506; No. 5,262,536; No. 5,272,250; No. 5,292,873; No. 5,317,098; No. 5,371,241; Nos. 5,391,723; 5,416,203, 5,451,463; 5,510,475; 5,512,667; 5,514,785; 5,565,552; 5,567,810; 5,574,142; 5,585,481; 5,587,371; 5,595,726; 5,597,696; 5,599,923; 5,599,928 and 5,688,941.

[0153] The isolated nucleic acid disclosed herein may comprise the nucleotide sequence of naturally occurring HSD17B13 gene or mRNA transcript, or may comprise a non-naturally occurring sequence.In one example, a non-naturally occurring sequence may differ from a non-naturally occurring sequence due to a synonymous mutation or a mutation that does not affect the encoded HSD17B13 protein.For example, a sequence may be identical except for a synonymous mutation or a mutation that does not affect the encoded HSD17B13 protein.A synonymous mutation or substitution is a substitution of one nucleotide for another nucleotide in an exon of a protein-encoding gene, so that the resulting amino acid sequence is not altered.This is possible due to the degeneracy of the genetic code, where some amino acids are coded by more than one three-base pair codon.Synonymous substitution is used, for example, in the process of codon optimization.

[0154] Also disclosed herein are proteins encoded by the nucleic acids disclosed herein, and isolated nucleic acids or proteins disclosed herein, and compositions comprising carriers that increase the stability of the isolated nucleic acids or proteins (e.g., extend the period during which degradation products remain below a threshold value, e.g., below 0.5% by weight of the starting nucleic acid or protein, under given storage conditions (e.g., −20° C., 4° C., or ambient temperature); or increase in vivo stability). Non-limiting examples of such carriers include poly(lactic acid) (PLA) microspheres, poly(D,L-lactic acid-co-glycolic acid) (PLGA) microspheres, liposomes, micelles, reverse micelles, lipid cocrystals, and lipid microtubules.

[0155] (1) A nucleic acid containing a mutant residue of the HSD17B13 rs72613567 variant. Disclosed herein is an isolated nucleic acid comprising at least 15 consecutive nucleotides of the HSD17B13 gene, which, when optimally aligned with the HSD17B13 rs72613567 variant, has a thymine at position 12666 of the HSD17B13 rs72613567 variant (or has a thymine at positions 12666 and 12667). That is, disclosed herein is an isolated nucleic acid comprising at least 15 consecutive nucleotides of the HSD17B13 gene, which, when optimally aligned with the wild-type HSD17B13 gene, includes a thymine inserted between the nucleotides corresponding to positions 12665 ​​and 12666 of the wild-type HSD17B13 gene (SEQ ID NO: 1). Such isolated nucleic acids may be useful, for example, for expressing HSD17B13 variant transcripts and proteins or as exogenous donor sequences. Such isolated nucleic acids may also be useful, for example, as guide RNAs, primers, and probes.

[0156] The HSD17B13 gene may be an HSD17B13 gene derived from any organism. For example, the HSD17B13 gene may be a human HSD17B13 gene or an ortholog derived from another organism, such as a non-human mammal, a rodent, a mouse, or a rat.

[0157] It is understood that gene sequences within a population may vary due to polymorphisms such as single nucleotide polymorphisms.The examples provided herein are merely exemplary sequences.Other sequences are also possible.For example, when optimally aligned with SEQ ID NO:2, at least 15 consecutive nucleotides can be at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the corresponding sequence of HSD17B13 rs72613567 variant (SEQ ID NO:2) comprising position 12666 or position 12666 and 12667 of SEQ ID NO:2.Optionally, the isolated nucleic acid comprises at least 15 consecutive nucleotides of SEQ ID NO:2 comprising position 12666 or position 12666 and 12667 of SEQ ID NO:2. As another example, the at least 15 contiguous nucleotides can be at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the corresponding sequence of the wild-type HSD17B13 gene (SEQ ID NO:1) including positions 12665 ​​and 12666 of SEQ ID NO:1 when optimally aligned with SEQ ID NO:1, wherein a thymine is present between positions corresponding to 12665 ​​and 12666 of SEQ ID NO:1. Optionally, the isolated nucleic acid comprises at least 15 contiguous nucleotides of SEQ ID NO:1, including positions 12665 ​​and 12666 of SEQ ID NO:1, wherein a thymine is present between positions corresponding to 12665 ​​and 12666 of SEQ ID NO:1.

[0158] The isolated nucleic acid can comprise, for example, at least 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 contiguous nucleotides of the HSD17B13 gene. Alternatively, the isolated nucleic acid can comprise, for example, at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, or 19000 contiguous nucleotides of the HSD17B13 gene.

[0159] In some cases, the isolated nucleic acid may comprise an HSD17B13 minigene in which one or more non-essential segments of the gene are deleted relative to the corresponding wild-type HSD17B13 gene. For example, the deleted segments include one or more intron sequences. For example, such an HSD17B13 minigene may comprise exons corresponding to exons 1-7 from HSD17B13 transcript D and an intron corresponding to intron 6 of SEQ ID NO:2 when optimally aligned with SEQ ID NO:2. For example, an HSD17B13 minigene may comprise exons 1-7 and intron 6 from SEQ ID NO:2. Minigenes are described in more detail elsewhere herein.

[0160] (2) a nucleic acid that hybridizes to a sequence adjacent to or including the mutant residue of the HSD17B13 rs72613567 variant; Also disclosed herein is an isolated nucleic acid that comprises at least 15 consecutive nucleotides that hybridize with HSD17B13 gene (for example, HSD17B13 minigene) at a segment that includes or is within 1000, 500, 400, 300, 200, 100, 50, 45, 40, 35, 30, 25, 20, 15, 10 or 5 nucleotides from the position corresponding to position 12666 or position 12666 and 12667 of HSD17B13 rs72613567 variant (SEQ ID NO: 2) when optimally aligned with HSD17B13 rs72613567 variant.Such isolated nucleic acid can be useful as, for example, guide RNA, primer, probe or exogenous donor sequence.

[0161] The HSD17B13 gene may be an HSD17B13 gene derived from any organism, for example, the HSD17B13 gene may be a human HSD17B13 gene or an ortholog derived from another organism, such as a non-human mammal, mouse, or rat.

[0162] As an example, at least 15 contiguous nucleotides may hybridize to a segment of the HSD17B13 gene or HSD17B13 minigene that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the corresponding sequence of the HSD17B13 rs72613567 variant (SEQ ID NO: 2) when optimally aligned with SEQ ID NO: 2. Optionally, the isolated nucleic acid may hybridize to at least 15 contiguous nucleotides of SEQ ID NO: 2. Optionally, the isolated nucleic acid hybridizes to a segment including positions 12666 or 12666 and 12667 of SEQ ID NO: 2, or positions corresponding to positions 12666 or 12666 and 12667 of SEQ ID NO: 2 when optimally aligned with SEQ ID NO: 2.

[0163] The segment to which the isolated nucleic acid can hybridize can include, for example, at least 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 75, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 contiguous nucleotides of the HSD17B13 gene. Alternatively, the isolated nucleic acid can comprise, for example, at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, or 19000 contiguous nucleotides of the HSD17B13 gene. Alternatively, the segment to which the isolated nucleic acid can hybridize can be, for example, up to 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 75, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 contiguous nucleotides of the HSD17B13 gene. For example, the segment can be about 15-100 nucleotides in length, or about 15-35 nucleotides in length.

[0164] (3) cDNA and variant transcripts resulting from the HSD17B13 rs72613567 variant Also provided are nucleic acids corresponding to all or a portion of the mRNA transcripts or cDNAs corresponding to any one of transcripts A-H (SEQ ID NOS: 4-11, respectively), particularly any one of transcripts C-H, when optimally aligned with any one of transcripts A-H. It is understood that gene sequences within a population and the mRNA sequences transcribed from such genes may vary due to polymorphisms, such as single nucleotide polymorphisms. The sequences presented herein for each transcript are merely exemplary sequences. Other sequences are possible. Specific, non-limiting examples are provided below. Such isolated nucleic acids may be useful, for example, for expressing HSD17B13 variant transcripts and proteins.

[0165] Isolated nucleic acid can be any length.For example, isolated nucleic acid can comprise at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or 2000 consecutive nucleotides that encode all or part of HSD17B13 protein.In some cases, isolated nucleic acid comprises consecutive nucleotides that encode all or part of HSD17B13 protein, and the consecutive nucleotides comprise the sequence from at least two different exons of HSD17B13 gene (for example, spanning at least one exon-exon boundary of HSD17B13 gene without intervening intron).

[0166] HSD17B13 transcript D (SEQ ID NO: 7), transcript G (SEQ ID NO: 10), and transcript H (SEQ ID NO: 11) contain a guanine insertion at the 3' end of exon 6, resulting in a frameshift of exon 7 and a truncation of the region of the HSD17B13 protein encoded by exon 7 compared to transcript A. Thus, presented herein are isolated nucleic acids containing segments (e.g., at least 15 contiguous nucleotides) present in transcripts D, G, and H (or fragments or homologs thereof) that are absent in transcript A (or fragments or homologs thereof). Such regions can be readily identified by comparing the sequences of the transcripts.For example, a segment of contiguous nucleotides (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides) that contains at least 15 contiguous nucleotides (e.g., at least 20 contiguous nucleotides or at least 30 contiguous nucleotides) encoding all or a portion of the HSD17B13 protein, when optimally aligned with SEQ ID NO: 7, 10, or 11, respectively, will have at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% similarity to the region spanning the boundary between exon 6 and exon 7 of SEQ ID NO: 7 (HSD17B13 transcript D), SEQ ID NO: 10 (HSD17B13 transcript G), or SEQ ID NO: 11 (HSD17B13 transcript H). or at least 99% identical, and wherein the segment contains a guanine at a residue corresponding to residue 878 at the 3' end of exon 6 of SEQ ID NO: 7 (i.e., a guanine inserted at the 3' end of exon 6 in addition to a guanine at the start of exon 7 compared to transcript A), a guanine at a residue corresponding to residue 770 at the 3' end of exon 6 of SEQ ID NO: 10 (i.e., a guanine inserted at the 3' end of exon 6 in addition to a guanine at the start of exon 7 compared to transcript B), or a guanine at a residue corresponding to residue 950 at the 3' end of exon 6 of SEQ ID NO: 11 (i.e., a guanine inserted at the 3' end of exon 6 in addition to a guanine at the start of exon 7 compared to transcript E). It is understood that such a nucleic acid would contain a sufficient number of nucleotides in each of exons 6 and 7 so that the guanine insertion can be distinguished from other features of the HSD17B13 transcript (e.g., a guanine at the start of exon 7, a read-through into intron 6 in transcript F, or a deletion of exon 6 in transcript C).

[0167] By way of example, the isolated nucleic acid may comprise at least 15 contiguous nucleotides (e.g., at least 20 contiguous nucleotides or at least 30 contiguous nucleotides) of SEQ ID NO:7 spanning the boundary between exon 6 and exon 7, optionally including exons 6 and 7 of SEQ ID NO:7, and optionally including the entire sequence of SEQ ID NO:7.

[0168] Optionally, the isolated nucleic acid further comprises a segment present in transcript D (or a fragment or homolog thereof) that is absent in transcript G (or a fragment or homolog thereof), and the isolated nucleic acid further comprises a segment present in transcript D (or a fragment or homolog thereof) that is absent in transcript H (or a fragment or homolog thereof). Such regions can be readily identified by comparing the sequences of the transcripts. For example, such an isolated nucleic acid can comprise a segment of contiguous nucleotides (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides) that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a region spanning the boundary between exons 3 and 4 of SEQ ID NO:7 (HSD17B13 transcript D) when optimally aligned with SEQ ID NO:7, as distinguished from transcript H. Similarly, such isolated nucleic acids can comprise a segment of contiguous nucleotides (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides) that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a region within exon 2 of SEQ ID NO:7 (HSD17B13 transcript D), a region spanning the boundary between exons 1 and 2 of SEQ ID NO:7, or a region spanning the boundary between exons 2 and 3 of SEQ ID NO:7, when optimally aligned with SEQ ID NO:7, so as to be distinguished from transcript G. Optionally, the isolated nucleic acid comprises a sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence set forth in SEQ ID NO:7 (HSD17B13 transcript D) and encodes an HSD17B13 protein comprising the sequence set forth in SEQ ID NO:15 (HSD17B13 isoform D). Similar to transcript D, transcript H (SEQ ID NO: 11) contains a guanine insertion 3' of exon 6 compared to transcript A.Transcript H further comprises an additional exon (exon 3') between exons 3 and 4 compared to transcript A and transcript D. Thus, as described with respect to the above, provided herein are isolated nucleic acids that include segments present in transcripts D, G, and H (or fragments or homologs thereof) that are absent in transcript A (or fragments or homologs thereof), but further comprise a segment (e.g., at least 15 contiguous nucleotides) of transcript H (or a fragment or homolog thereof) that is absent in transcript D (or a fragment or homolog thereof). Such regions can be readily identified by comparing the sequences of the transcripts. For example, provided herein are isolated nucleic acids described for transcript D that are at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a region within exon 3' of SEQ ID NO: 11 (HSD17B13 transcript H), a region spanning the boundary between exons 3 and 3' of SEQ ID NO: 11, or a region spanning the boundary between exons 3' and 4 of SEQ ID NO: 11, when a segment of contiguous nucleotides (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides) is optimally aligned with SEQ ID NO: 11. It is understood that such nucleic acids will contain a sufficient number of nucleotides in each of exons 3 and 3' or each of exons 3' and 4 to be distinct from other features of the HSD17B13 transcript (e.g., the boundary between exons 3 and 4). For example, the region of exon 3' can include the entire exon 3'. Optionally, the isolated nucleic acid comprises a sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence set forth in SEQ ID NO: 11 (HSD17B13 transcript H) and encodes an HSD17B13 protein comprising the sequence set forth in SEQ ID NO: 19 (HSD17B13 isoform H).

[0169] By way of example, the isolated nucleic acid may comprise at least 15 contiguous nucleotides (e.g., at least 20 contiguous nucleotides or at least 30 contiguous nucleotides) of SEQ ID NO:11, including a region within exon 3', a region spanning the boundary between exon 3 and exon 3', or a region spanning the boundary between exon 3' and exon 4, optionally including the entire exon 3' of SEQ ID NO:11, and optionally including the entire sequence of SEQ ID NO:11.

[0170] Like transcript D, transcript G (SEQ ID NO: 10) contains a guanine insertion 3' of exon 6 compared to transcript A. However, in addition, transcript G lacks exon 2 compared to transcripts A and D (i.e., transcript G contains the boundary between exons 1 and 3, which is absent in transcripts A and D). Thus, provided herein are isolated nucleic acids as described above that contain segments present in transcripts D, G, and H (or fragments or homologs thereof) that are absent in transcript A (or fragments or homologs thereof), but further contain a segment (e.g., at least 15 contiguous nucleotides) from transcript G (or a fragment or homolog thereof) that is absent in transcript D (or a fragment or homolog thereof). Such regions can be readily identified by comparing the sequences of the transcripts. For example, provided herein are isolated nucleic acids described for transcript D that are at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the region spanning the boundary between exons 1 and 3 of SEQ ID NO: 10 (HSD17B13 transcript G) when optimally aligned with SEQ ID NO: 10. It is understood that such nucleic acids will include a sufficient number of nucleotides in each of exons 1 and 3 to be distinct from other features of the HSD17B13 transcript (e.g., the boundary between exons 1 and 2 or the boundary between exons 2 and 3). For example, the region can include the entirety of exons 1 and 3 of SEQ ID NO: 10. Optionally, the isolated nucleic acid comprises a sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence set forth in SEQ ID NO: 10 (HSD17B13 transcript G) and encodes an HSD17B13 protein comprising the sequence set forth in SEQ ID NO: 18 (HSD17B13 isoform G).

[0171] By way of example, the isolated nucleic acid may comprise at least 15 contiguous nucleotides (e.g., at least 20 contiguous nucleotides or at least 30 contiguous nucleotides) of SEQ ID NO: 10, including the region spanning the boundary between exon 1 and exon 3, optionally including exons 1 and 3 of SEQ ID NO: 10, and optionally including the entire sequence of SEQ ID NO: 10.

[0172] Also provided herein is an isolated nucleic acid comprising a segment (e.g., at least 15 contiguous nucleotides) present in transcript E (or a fragment or homolog thereof) that is absent in transcript A (or a fragment or homolog thereof). Such regions can be readily identified by comparing the sequences of the transcripts. Transcript E (SEQ ID NO: 8) contains an additional exon between exons 3 and 4 compared to transcript A. Thus, provided herein are isolated nucleic acids comprising at least 15 consecutive nucleotides (e.g., at least 20 consecutive nucleotides or at least 30 consecutive nucleotides) encoding all or a portion of the HSD17B13 protein, wherein a segment of consecutive nucleotides (e.g., at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, or at least 15 consecutive nucleotides) is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a region within exon 3' of SEQ ID NO: 8 (HSD17B13 transcript E), a region spanning the boundary between exon 3 and exon 3' of SEQ ID NO: 8, or a region spanning the boundary between exon 3' and exon 4 of SEQ ID NO: 8, when optimally aligned with SEQ ID NO: 8. It is understood that such nucleic acids will contain a sufficient number of nucleotides in each of exons 3 and 3' or each of exons 3' and 4 to be distinguished from other features of the HSD17B13 transcript (e.g., the boundary between exons 3 and 4). For example, the region of exon 3' can include the entire exon 3'. Optionally, the isolated nucleic acid further includes a segment (e.g., at least 15 contiguous nucleotides) from transcript E (or a fragment or homolog thereof) that is not present in transcript H (or a fragment or homolog thereof). Such regions can be readily identified by comparing the sequences of the transcripts.For example, the above-described isolated nucleic acids are provided herein, in which a segment of consecutive nucleotides (e.g., at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, or at least 15 consecutive nucleotides) is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the region spanning the boundary between exon 6 and exon 7 of SEQ ID NO: 8 (HSD17B13 transcript E) when optimally aligned with SEQ ID NO: 8. It is understood that such nucleic acids will include a sufficient number of nucleotides in each of exons 6 and 7 to be distinct from other features of the HSD17B13 transcript (particularly the additional guanine at the 3' end of exon 6 of transcript H). For example, the region can include the entirety of exons 6 and 7 of SEQ ID NO: 8. Optionally, the isolated nucleic acid comprises a sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence set forth in SEQ ID NO: 8 (HSD17B13 transcript E) and encodes an HSD17B13 protein comprising the sequence set forth in SEQ ID NO: 16 (HSD17B13 isoform E).

[0173] By way of example, the isolated nucleic acid may comprise at least 15 contiguous nucleotides (e.g., at least 20 contiguous nucleotides or at least 30 contiguous nucleotides) of SEQ ID NO:8, including a region within exon 3', a region spanning the boundary between exon 3 and exon 3', or a region spanning the boundary between exon 3' and exon 4, optionally including the entire exon 3' of SEQ ID NO:8, and optionally including the entire sequence of SEQ ID NO:8.

[0174] Also provided herein is an isolated nucleic acid comprising a segment (e.g., at least 15 contiguous nucleotides) present in transcript F (or a fragment or homolog thereof) that is not present in transcript A (or a fragment or homolog thereof). Such a region can be readily identified by comparing the sequences of the transcripts. Compared to transcript A, transcript F (SEQ ID NO: 9) contains a readthrough from exon 6 to intron 6, which readthrough contains a thymine insertion present in the HSD17B13 rs72613567 variant gene. Thus, provided herein are isolated nucleic acids comprising at least 15 contiguous nucleotides (e.g., at least 20 contiguous nucleotides or at least 30 contiguous nucleotides) encoding all or a portion of an HSD17B13 protein, wherein a segment of contiguous nucleotides (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides) is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a region within the readthrough into intron 6 of SEQ ID NO:9 (HSD17B13 transcript F) or a region spanning the boundary between the readthrough into intron 6 of SEQ ID NO:9 and the remainder of exon 6. It is understood that such nucleic acids will include a sufficient number of nucleotides in the readthrough to distinguish it from other features of the HSD17B13 transcript (e.g., the boundary between exons 6 and 7 of other HSD17B13 transcripts). Optionally, the contiguous nucleotides include a sequence present in transcript F that is not present in transcript F' (SEQ ID NO: 246) (i.e., a thymine insertion). Transcript F' also includes a readthrough from exon 6 into intron 6 compared to transcript A, but the readthrough does not include the thymine insertion present in the HSD17B13 rs72613567 variant gene. For example, the region can be the entire readthrough into intron 6 of SEQ ID NO: 9.Optionally, the isolated nucleic acid comprises a sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence set forth in SEQ ID NO: 9 (HSD17B13 transcript F) and encodes an HSD17B13 protein comprising the sequence set forth in SEQ ID NO: 17 (HSD17B13 isoform F).

[0175] By way of example, the isolated nucleic acid may comprise at least 15 contiguous nucleotides (e.g., at least 20 contiguous nucleotides or at least 30 contiguous nucleotides) of SEQ ID NO: 9, including the region within the readthrough to intron 6 or the region spanning the boundary between the readthrough to intron 6 and the remainder of exon 6, optionally including the entire readthrough to intron 6, and optionally including the entire sequence of SEQ ID NO: 9.

[0176] Also provided herein is an isolated nucleic acid comprising a segment (e.g., at least 15 contiguous nucleotides) present in transcript F' (or a fragment or homolog thereof) that is not present in transcript A (or a fragment or homolog thereof). Such a region can be readily identified by comparing the sequences of the transcripts. Compared to transcript A, transcript F' (SEQ ID NO: 246) contains a read-through from exon 6 to intron 6, and the read-through does not contain the thymine insertion present in the HSD17B13 rs72613567 variant gene. Thus, provided herein are isolated nucleic acids comprising at least 15 contiguous nucleotides (e.g., at least 20 contiguous nucleotides or at least 30 contiguous nucleotides) encoding all or a portion of an HSD17B13 protein, wherein the segment of contiguous nucleotides (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides) is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a region within the readthrough into intron 6 of SEQ ID NO:246 (HSD17B13 transcript F') or to a region spanning the boundary between the readthrough into intron 6 of SEQ ID NO:246 and the remainder of exon 6. It is understood that such nucleic acids will include a sufficient number of nucleotides in the readthrough to distinguish it from other features of the HSD17B13 transcript (e.g., the boundary between exons 6 and 7 of other HSD17B13 transcripts). Optionally, the contiguous nucleotides include a sequence present in transcript F' that is not present in transcript F (SEQ ID NO: 9). The readthrough of transcript F includes the thymine insertion present in the HSD17B13 rs72613567 variant gene, but the readthrough of transcript F' does not. For example, the region can be the entire readthrough into intron 6 of SEQ ID NO: 246.Optionally, the isolated nucleic acid comprises a sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence set forth in SEQ ID NO: 246 (HSD17B13 transcript F') and encodes an HSD17B13 protein that comprises, consists essentially of, or consists of the sequence set forth in SEQ ID NO: 247 (HSD17B13 isoform F').

[0177] By way of example, the isolated nucleic acid may comprise at least 15 contiguous nucleotides (e.g., at least 20 contiguous nucleotides or at least 30 contiguous nucleotides) of SEQ ID NO:246, including the region within the readthrough to intron 6 or the region spanning the boundary between the readthrough to intron 6 and the remainder of exon 6, optionally including the entire readthrough to intron 6, and optionally including the entire sequence of SEQ ID NO:246.

[0178] Also provided herein is an isolated nucleic acid comprising a segment (e.g., at least 15 contiguous nucleotides) present in transcript C (or a fragment or homolog thereof) that is absent in transcript A (or a fragment or homolog thereof). Such a region can be readily identified by comparing the sequences of the transcripts. Transcript C (SEQ ID NO: 6) lacks exon 6 compared to transcript A (i.e., transcript C includes the boundary between exon 5 and exon 7, which is absent in transcript A). Thus, the present specification provides an isolated nucleic acid comprising at least 15 consecutive nucleotides (e.g., at least 20 consecutive nucleotides or at least 30 consecutive nucleotides) encoding all or part of the HSD17B13 protein, and wherein the segment of consecutive nucleotides (e.g., at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, or at least 15 consecutive nucleotides) is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the region spanning the boundary between exon 5 and exon 7 of SEQ ID NO: 6 (HSD17B13 transcript C) when optimally aligned with SEQ ID NO: 6. It is understood that such a nucleic acid will contain a sufficient number of nucleotides in each of exons 5 and 7 to be distinguished from other features of the HSD17B13 transcript (e.g., the boundary between exons 5 and 6 or the boundary between exons 6 and 7 of other HSD17B13 transcripts). For example, the region may comprise the entirety of exons 5 and 7 of SEQ ID NO: 6. Optionally, the isolated nucleic acid comprises a sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence set forth in SEQ ID NO: 6 (HSD17B13 transcript C) and encodes an HSD17B13 protein comprising the sequence set forth in SEQ ID NO: 14 (HSD17B13 isoform C).

[0179] By way of example, the isolated nucleic acid may comprise at least 15 contiguous nucleotides (e.g., at least 20 contiguous nucleotides or at least 30 contiguous nucleotides) of SEQ ID NO: 6, including the region spanning the boundary between exon 5 and exon 7, optionally including the entirety of exons 5 and 7 of SEQ ID NO: 6, and optionally including the entire sequence of SEQ ID NO: 6.

[0180] (4) cDNA and nucleic acids that hybridize with variant HSD17B13 transcripts Also provided are nucleic acids that hybridize to a segment of an mRNA transcript or cDNA corresponding to any one of transcripts A-H (SEQ ID NOS: 4-11, respectively), particularly transcripts C-H, when optimally aligned with any one of transcripts A-H. Specific, non-limiting examples are provided below. Such isolated nucleic acids may be useful, for example, as primers, probes, antisense RNA, siRNA, or shRNA.

[0181] The segment to which the isolated nucleic acid can hybridize can include, for example, at least 5, at least 10, or at least 15 consecutive nucleotides of a nucleic acid encoding an HSD17B13 protein. The segment to which the isolated nucleic acid can hybridize can include, for example, at least 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 75, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or 2000 consecutive nucleotides of a nucleic acid encoding an HSD17B13 protein. Alternatively, the segment to which the isolated nucleic acid can hybridize can be, for example, up to 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 75, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 contiguous nucleotides of the nucleic acid encoding the HSD17B13 protein. For example, the segment can be about 15-100 nucleotides in length, or about 15-35 nucleotides in length.

[0182] HSD17B13 transcript D (SEQ ID NO: 7), transcript G (SEQ ID NO: 10), and transcript H (SEQ ID NO: 11) contain a guanine insertion at the 3' end of exon 6, resulting in a frameshift in exon 7 and a premature truncation of exon 7 compared to transcript A. Thus, provided herein are isolated nucleic acids containing regions (e.g., at least 15 contiguous nucleotides) that hybridize to segments present in transcripts D, G, and H (or fragments or homologs thereof) that are not present in transcript A (or fragments or homologs thereof). Such regions can be readily identified by comparing the sequences of the transcripts. For example, provided herein is an isolated nucleic acid that hybridizes to at least 15 contiguous nucleotides of a nucleic acid encoding an HSD17B13 protein, wherein the contiguous nucleotides include a segment (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides) that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the region spanning the boundary between exon 6 and exon 7 of SEQ ID NO: 7 (HSD17B13 transcript D) when optimally aligned with SEQ ID NO: 7, and wherein the segment includes a guanine at the residue corresponding to residue 878 at the 3' end of exon 6 of SEQ ID NO: 7 (i.e., a guanine is inserted at the 3' end of exon 6 in addition to the guanine at the start of exon 7 compared to transcript A).Alternatively, provided herein is an isolated nucleic acid that hybridizes to a segment of at least 15 contiguous nucleotides of a nucleic acid encoding an HSD17B13 protein, wherein the contiguous nucleotides include a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides) to the region spanning the boundary between exon 6 and exon 7 of SEQ ID NO: 10 (HSD17B13 transcript G) when optimally aligned with SEQ ID NO: 10, and wherein the segment includes a guanine at the residue corresponding to residue 770 at the 3' end of exon 6 of SEQ ID NO: 10 (i.e., a guanine is inserted at the 3' end of exon 6 in addition to the guanine at the start of exon 7 compared to transcript B). Alternatively, provided herein is an isolated nucleic acid that hybridizes to at least 15 contiguous nucleotides of a nucleic acid encoding an HSD17B13 protein, wherein the contiguous nucleotides include a segment (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides) that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the region spanning the boundary between exon 6 and exon 7 of SEQ ID NO:11 (HSD17B13 transcript H) when optimally aligned with SEQ ID NO:11, and wherein the segment includes a guanine at the residue corresponding to residue 950 at the 3' end of exon 6 of SEQ ID NO:11 (i.e., a guanine is inserted at the 3' end of exon 6 in addition to the guanine at the start of exon 7 compared to transcript E). It is understood that such nucleic acids will be designed to hybridize with a sufficient number of nucleotides in each of exons 6 and 7 so that the guanine insertion can be distinguished from other features of the HSD17B13 transcript (e.g., read-through into intron 6 in transcript F or deletion of exon 6 in transcript C).

[0183] As one example, a segment may comprise the region spanning the boundary between exon 6 and exon 7 of SEQ ID NO:7 (i.e., including the guanine at residue 878 of SEQ ID NO:7). As another example, a segment may comprise the region spanning the boundary between exon 6 and exon 7 of SEQ ID NO:10 (i.e., including the guanine at residue 770 of SEQ ID NO:10). As another example, a segment may comprise the region spanning the boundary between exon 6 and exon 7 of SEQ ID NO:11 (i.e., including the guanine at residue 950 of SEQ ID NO:11).

[0184] Optionally, the isolated nucleic acid further comprises a region (e.g., 15 contiguous nucleotides) that hybridizes to a segment present in transcript D (or a fragment or homolog thereof) that is absent in transcript G (or a fragment or homolog thereof), and the isolated nucleic acid further comprises a region that hybridizes to a segment present in transcript D (or a fragment or homolog thereof) that is absent in transcript H (or a fragment or homolog thereof). Such segments can be readily identified by comparing the sequences of the transcripts. For example, a segment (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides) present in transcript D (or a fragment or homolog thereof) that is absent in transcript H (or a fragment or homolog thereof) can be at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a region spanning the boundary between exons 3 and 4 of SEQ ID NO:7 (HSD17B13 transcript D) when optimally aligned with SEQ ID NO:7, as distinguished from transcript H. Similarly, a segment (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides) present in transcript D (or a fragment or homolog thereof) that is not present in transcript G (or a fragment or homolog thereof) may be at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a region within exon 2 of SEQ ID NO: 7 (HSD17B13 transcript D), a region spanning the boundary between exons 1 and 2 of SEQ ID NO: 7, or a region spanning the boundary between exons 2 and 3 of SEQ ID NO: 7, when optimally aligned with SEQ ID NO: 7, so as to be distinguished from transcript G.

[0185] Like transcript D, transcript H (SEQ ID NO: 11) contains a guanine insertion at the 3' end of exon 6 compared to transcript A. Transcript H further contains an additional exon between exons 3 and 4 compared to transcripts A and D. Thus, provided herein are isolated nucleic acids as described above that contain regions that hybridize with segments present in transcripts D, G, and H (or fragments or homologs thereof) that are not present in transcript A (or fragments or homologs thereof), but further contain regions (e.g., at least 15 contiguous nucleotides) that hybridize with segments present in transcript H (or fragments or homologs thereof) but not in transcript D (or fragments or homologs thereof). Such regions can be readily identified by comparing the sequences of the transcripts. For example, a segment can be at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a region within exon 3' of SEQ ID NO: 11 (HSD17B13 transcript H) (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides), a region spanning the boundary between exon 3 and exon 3' of SEQ ID NO: 11, or a region spanning the boundary between exon 3' and exon 4 of SEQ ID NO: 11, when optimally aligned with SEQ ID NO: 11. It is understood that such nucleic acids will be designed to hybridize to a sufficient number of nucleotides of exons 3 and 3', respectively, or exons 3' and 4, respectively, to be distinguishable from other features of the HSD17B13 transcript (e.g., the boundary between exons 3 and 4).

[0186] As an example, a segment may include a region of SEQ ID NO: 11 within exon 3', spanning the boundary between exon 3 and exon 3', or spanning the boundary between exon 3' and exon 4.

[0187] Like transcript D, transcript G (SEQ ID NO: 10) contains a guanine insertion at the 3' end of exon 6 compared to transcript A. However, in addition, transcript G lacks exon 2 compared to transcripts A and D (i.e., transcript G contains the boundary between exons 1 and 3, which is absent in transcripts A and D). Thus, the above-described isolated nucleic acids are provided herein that contain a region that hybridizes with a segment present in transcripts D, G, and H (or a fragment or homolog thereof) that is absent in transcript A (or a fragment or homolog thereof), but further contain a region (e.g., at least 15 contiguous nucleotides) that hybridizes with a segment present in transcript G (or a fragment or homolog thereof) but absent in transcript D (or a fragment or homolog thereof). Such regions can be readily identified by comparing the sequences of the transcripts. For example, a segment may be at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a region spanning the boundary between exon 1 and exon 3 of SEQ ID NO: 10 (HSD17B13 transcript G) (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides) when optimally aligned with SEQ ID NO: 10. It is understood that such nucleic acids will be designed to hybridize to a sufficient number of nucleotides in exons 1 and 3, respectively, to be distinct from other features of the HSD17B13 transcript (e.g., the boundary between exons 1 and 2 or the boundary between exons 2 and 3).

[0188] As an example, a segment may include the region of SEQ ID NO: 10 that spans the boundary between exon 1 and exon 3.

[0189] Also provided is an isolated nucleic acid comprising a region (e.g., at least 15 contiguous nucleotides) that hybridizes with a segment of nucleic acid encoding the HSD17B13 protein that is present in transcript E (or a fragment or homolog thereof) but not in transcript A (or a fragment or homolog thereof). Such regions can be readily identified by comparing the sequences of the transcripts. Transcript E (SEQ ID NO: 8) contains an additional exon between exons 3 and 4 compared to transcript A. Thus, provided herein are isolated nucleic acids that hybridize to at least 15 consecutive nucleotides of a nucleic acid encoding an HSD17B13 protein, the consecutive nucleotides including a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a region within exon 3' of SEQ ID NO: 8 (HSD17B13 transcript E) (e.g., at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, or at least 15 consecutive nucleotides) when optimally aligned with SEQ ID NO: 8, a region spanning the boundary between exon 3 and exon 3' of SEQ ID NO: 8, or a region spanning the boundary between exon 3' and exon 4 of SEQ ID NO: 8. It is understood that such nucleic acids will be designed to hybridize to a sufficient number of nucleotides of exons 3 and 3', respectively, or exons 3' and 4, respectively, to be distinguishable from other features of the HSD17B13 transcript (e.g., the boundary between exons 3 and 4).

[0190] As an example, a segment may include a region of SEQ ID NO:8 that is within exon 3', that spans the boundary between exon 3 and exon 3' of SEQ ID NO:8, or that spans the boundary between exon 3' and exon 4.

[0191] Optionally, the isolated nucleic acid further comprises a region (e.g., 15 contiguous nucleotides) that hybridizes to a segment present in transcript E (or a fragment or homolog thereof) that is absent in transcript H (or a fragment or homolog thereof). Such segments can be readily identified by comparing the sequences of the transcripts. For example, a segment (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides) present in transcript E (or a fragment or homolog thereof) that is absent in transcript H (or a fragment or homolog thereof) can be at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a region spanning the boundary between exons 6 and 7 of SEQ ID NO:8 (HSD17B13 transcript E) when optimally aligned with SEQ ID NO:8, as distinguished from transcript G. It is understood that such nucleic acids will be designed to hybridize with a sufficient number of nucleotides in exons 6 and 7, respectively, to distinguish them from other features of the HSD17B13 transcript (particularly the additional guanine at the 3' end of exon 6 of transcript H).

[0192] Also provided is an isolated nucleic acid comprising a region (e.g., at least 15 contiguous nucleotides) that hybridizes with a segment of nucleic acid encoding the HSD17B13 protein that is present in transcript F (or a fragment or homolog thereof) but not in transcript A (or a fragment or homolog thereof). Such a region can be readily identified by comparing the sequences of the transcripts. Transcript F (SEQ ID NO: 9) contains a read-through from exon 6 to intron 6 compared to transcript A. Thus, provided herein are isolated nucleic acids that hybridize to at least 15 contiguous nucleotides of a nucleic acid encoding an HSD17B13 protein, wherein the contiguous nucleotides include a segment (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides) that, when optimally aligned with SEQ ID NO: 9, is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a region within the readthrough into intron 6 of SEQ ID NO: 9 (HSD17B13 transcript F) or to a region spanning the boundary between the readthrough into intron 6 of SEQ ID NO: 9 and the remainder of exon 6. It is understood that such nucleic acids will be designed to hybridize to a sufficient number of nucleotides within the readthrough to distinguish it from other features of the HSD17B13 transcript (e.g., the boundary between exons 6 and 7 of other HSD17B13 transcripts). Optionally, the contiguous nucleotides include a sequence present in transcript F that is not present in transcript F' (SEQ ID NO: 246) (i.e., a thymine insertion). Transcript F' also includes a readthrough from exon 6 to intron 6 compared to transcript A, but the readthrough does not include the thymine insertion present in the HSD17B13 rs72613567 variant gene.

[0193] As an example, the segment may include a region of SEQ ID NO:9 that is within the readthrough into intron 6 or that spans the boundary between the readthrough into intron 6 and the remainder of exon 6.

[0194] Also provided is an isolated nucleic acid comprising a region (e.g., at least 15 contiguous nucleotides) that hybridizes with a segment of nucleic acid encoding the HSD17B13 protein that is present in transcript F' (or a fragment or homolog thereof) but not in transcript A (or a fragment or homolog thereof). Such a region can be readily identified by comparing the sequences of the transcripts. Transcript F' (SEQ ID NO: 246) contains a read-through from exon 6 to intron 6 compared to transcript A. Thus, provided herein are isolated nucleic acids that hybridize to at least 15 consecutive nucleotides of a nucleic acid encoding an HSD17B13 protein, wherein the consecutive nucleotides include a segment (e.g., at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, or at least 15 consecutive nucleotides) that, when optimally aligned with SEQ ID NO: 246, is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a region within the readthrough into intron 6 of SEQ ID NO: 246 (HSD17B13 transcript F') or to a region spanning the boundary between the readthrough into intron 6 of SEQ ID NO: 246 and the remainder of exon 6. It is understood that such nucleic acids will be designed to hybridize to a sufficient number of nucleotides within the readthrough to distinguish it from other features of the HSD17B13 transcript (e.g., the boundary between exons 6 and 7 of other HSD17B13 transcripts). Optionally, the contiguous nucleotides include a sequence present in transcript F' that is not present in transcript F (SEQ ID NO: 9). The readthrough of transcript F includes the thymine insertion present in the HSD17B13 rs72613567 variant gene, but the readthrough of transcript F' does not.

[0195] As an example, the segment may include a region of SEQ ID NO: 246 that is within the readthrough into intron 6 or that spans the boundary between the readthrough into intron 6 and the remainder of exon 6.

[0196] Also provided is an isolated nucleic acid comprising a region (e.g., at least 15 contiguous nucleotides) that hybridizes with a segment of nucleic acid encoding the HSD17B13 protein that is present in transcript C (or a fragment or homolog thereof) but not in transcript A (or a fragment or homolog thereof). Such a region can be readily identified by comparing the sequences of the transcripts. Transcript C (SEQ ID NO: 6) lacks exon 6 compared to transcript A (i.e., transcript C includes the boundary between exon 5 and exon 7, which is absent in transcript A). Thus, provided herein are isolated nucleic acids that hybridize to at least 15 consecutive nucleotides of a nucleic acid encoding an HSD17B13 protein, wherein the consecutive nucleotides comprise a segment (e.g., at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, or at least 15 consecutive nucleotides) that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the region spanning the boundary between exon 5 and exon 7 of SEQ ID NO: 6 (HSD17B13 transcript C) when optimally aligned with SEQ ID NO: 6. It is understood that such nucleic acids will be designed to hybridize to a sufficient number of nucleotides of exons 5 and 7 to be distinguished from other features of the HSD17B13 transcript (e.g., the boundary between exons 5 and 6 or the boundary between exons 6 and 7 of other HSD17B13 transcripts).

[0197] As an example, a segment may include the region from SEQ ID NO:6 that spans the boundary between exon 5 and exon 7.

[0198] Also provided herein are isolated nucleic acids (e.g., antisense RNA, siRNA, or shRNA) that hybridize with at least 15 consecutive nucleotides of a nucleic acid encoding an HSD17B13 protein, wherein the consecutive nucleotides comprise a segment (e.g., at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, or at least 15 consecutive nucleotides) that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a region of HSD17B13 transcript D (SEQ ID NO: 7). The isolated nucleic acid may comprise a region (e.g., at least 15 consecutive nucleotides) that hybridizes with a segment present in transcript D (or a fragment or homolog thereof) that is not present in transcript A (or a fragment or homolog thereof). Such regions can be easily identified by comparing the sequences of the transcripts. HSD17B13 transcript D (SEQ ID NO: 7) contains a guanine insertion at the 3' end of exon 6, resulting in a frameshift in exon 7 and a premature truncation of exon 7 compared to transcript A (SEQ ID NO: 4). For example, provided herein are isolated nucleic acids that hybridize to at least 15 contiguous nucleotides of a nucleic acid encoding an HSD17B13 protein, wherein the contiguous nucleotides include a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides) to the region spanning the boundary between exon 6 and exon 7 of SEQ ID NO: 7 (HSD17B13 transcript D) when optimally aligned with SEQ ID NO: 7. The segment may include a guanine at the residue corresponding to residue 878 at the 3' end of exon 6 of SEQ ID NO: 7 (i.e., compared to transcript A, in addition to the guanine at the start of exon 7, a guanine is inserted at the 3' end of exon 6).It is understood that such nucleic acids will be designed to hybridize with a sufficient number of nucleotides in each of exons 6 and 7 so that the guanine insertion can be distinguished from other features of the HSD17B13 transcript (e.g., read-through into intron 6 in transcript F or deletion of exon 6 in transcript C).

[0199] Also provided herein are isolated nucleic acids (e.g., antisense RNA, siRNA, or shRNA) that hybridize with at least 15 consecutive nucleotides of a nucleic acid encoding an HSD17B13 protein, wherein the consecutive nucleotides comprise a segment (e.g., at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, or at least 15 consecutive nucleotides) that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a region of HSD17B13 transcript A (SEQ ID NO: 4). The isolated nucleic acid may comprise a region (e.g., at least 15 consecutive nucleotides) that hybridizes with a segment present in transcript A (or a fragment or homolog thereof) that is not present in transcript D (or a fragment or homolog thereof). Such regions can be easily identified by comparing the sequences of the transcripts. HSD17B13 transcript D (SEQ ID NO: 7) contains an insertion of a guanine at the 3' end of exon 6, resulting in a frameshift in exon 7 and a premature truncation of exon 7 compared to transcript A (SEQ ID NO: 4). For example, provided herein are isolated nucleic acids that hybridize to at least 15 contiguous nucleotides of a nucleic acid encoding an HSD17B13 protein, wherein the contiguous nucleotides include a segment that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical (e.g., at least 5 contiguous nucleotides, at least 10 contiguous nucleotides, or at least 15 contiguous nucleotides) to the region spanning the boundary between exon 6 and exon 7 of SEQ ID NO: 4 (HSD17B13 transcript A) when optimally aligned with SEQ ID NO: 4.

[0200] (5) Vector Also provided are vectors comprising any of the nucleic acids disclosed herein and heterologous nucleic acids. The vectors may be viral or non-viral vectors capable of transporting nucleic acids. In some cases, the vectors may be plasmids (e.g., circular double-stranded DNA into which additional DNA segments can be ligated). In some cases, the vectors may be viral vectors, and additional DNA segments can be ligated into the viral genome. In some cases, the vectors can replicate autonomously in a host cell into which they are introduced (e.g., bacterial vectors with a bacterial origin of replication and episomal mammalian vectors). In other cases, vectors (e.g., non-episomal mammalian vectors) can be integrated into the genome of a host cell upon introduction into the host cell, thereby replicating along with the host genome. Furthermore, certain vectors can direct the expression of genes to which they are operably linked. Such vectors may be referred to as "recombinant expression vectors" or "expression vectors." Such vectors may also be targeting vectors (i.e., exogenous donor sequences) as disclosed elsewhere herein.

[0201] In some cases, proteins encoded by the disclosed genetic variants are expressed by inserting nucleic acids encoding the disclosed genetic variants into an expression vector so that the genes are operably linked to necessary expression regulatory sequences, such as transcriptional and translational regulatory sequences. Expression vectors include, for example, plasmids, retroviruses, adenoviruses, adeno-associated viruses (AAVs), plant viruses such as cauliflower mosaic virus and tobacco mosaic virus, cosmids, YACs, and EBV-derived episomes. In some cases, nucleic acids containing the disclosed genetic variants can be ligated into a vector so that the transcriptional and translational regulatory sequences within the vector perform their intended function of controlling the transcription and translation of the genetic variant. The expression vector and expression regulatory sequences are selected to be compatible with the expression host cell used. Nucleic acid sequences containing the disclosed genetic variants can be inserted into separate vectors or into the same expression vector. Nucleic acid sequences containing the disclosed genetic variants can be inserted into an expression vector by standard methods (e.g., ligation of the vector with complementary restriction sites in the nucleic acid containing the disclosed genetic variant, or blunt-end ligation if no restriction sites are present).

[0202] In addition to the nucleic acid sequence containing the disclosed gene variants, the recombinant expression vector may have a control sequence that regulates the expression of the gene variant in a host cell. The design of the expression vector, including the selection of the control sequence, may depend on factors such as the choice of host cell to be transformed, the level of expression of the desired protein, etc. Preferred control sequences for expression in mammalian host cells include, for example, viral elements that direct high levels of protein expression in mammalian cells, such as retroviral LTRs, cytomegalovirus (CMV) (e.g., CMV promoter / enhancer), simian virus 40 (SV40) (e.g., SV40 promoter / enhancer), adenovirus (e.g., adenovirus major late promoter (AdMLP)), promoters and / or enhancers derived from polyoma, and strong mammalian promoters such as native immunoglobulin and actin promoters. Further description of viral regulatory elements, and sequences thereof, is provided in U.S. Patent Nos. 5,168,062; 4,510,245; and 4,968,615, each of which is incorporated by reference in its entirety for all purposes. Methods for expressing polypeptides in bacterial or fungal cells (e.g., yeast cells) are also well known.

[0203] In addition to the nucleic acid sequences and control sequences containing the disclosed gene variants, recombinant expression vectors may have additional sequences, such as sequences that control the replication of the vector in host cells (e.g., origins of replication) and selectable marker genes. The selectable marker gene can facilitate the selection of host cells into which the vector has been introduced (see, for example, U.S. Pat. Nos. 4,399,216; 4,634,665; and 5,179,017, each of which is incorporated herein by reference in its entirety for all purposes). For example, the selectable marker gene can confer resistance to drugs such as G418, hygromycin, or methotrexate to host cells into which the vector has been introduced. Exemplary selectable marker genes include the dihydrofolate reductase (DHFR) gene (for use with methotrexate selection / amplification in dhfr- host cells), the neo gene (for G418 selection), and the glutamate synthetase (GS) gene.

[0204] B. Protein Disclosed herein are isolated HSD17B13 proteins and fragments thereof, particularly HSD17B13 proteins and fragments thereof caused by the HSD17B13 rs72613567 variant.

[0205] The isolated protein disclosed herein may comprise the amino acid sequence of naturally occurring HSD17B13 protein, or may comprise a non-naturally occurring sequence.In one example, the non-naturally occurring sequence may differ from the non-naturally occurring sequence due to conservative amino acid substitution.For example, the sequence may be identical except for conservative amino acid substitution.

[0206] The isolated proteins disclosed herein can be linked or fused to heterologous polypeptides or heterologous molecules or labels, numerous examples of which are disclosed elsewhere herein. For example, proteins can be fused to heterologous polypeptides that increase or decrease stability. The fused domain or heterologous polypeptide can be located at the N-terminus, C-terminus, or internally of the protein. For example, the fusion partner can help provide a T-helper epitope (immunological fusion partner) or can help express the protein at a higher yield than the native recombinant protein (expression enhancer). Certain fusion partners enhance both immunity and expression. Other fusion partners can be selected to increase the solubility of the polypeptide or to enable targeting of the polypeptide to a desired intracellular compartment. Yet another fusion partner includes an affinity tag that facilitates polypeptide purification.

[0207] A fusion protein can be directly fused to a heterologous molecule or can be linked to a heterologous molecule via a linker, such as a peptide linker. Suitable peptide linker sequences can be selected based on, for example, the following factors: (1) the ability to adopt a flexible, extended conformation; (2) the inability to adopt secondary structures that may interact with functional epitopes on the first and second polypeptides; and (3) the absence of hydrophobic or charged residues that may react with functional epitopes on the polypeptides. For example, peptide linker sequences can contain Gly, Asn, and Ser residues. Other near-neutral amino acids, such as Thr and Ala, can also be used in linker sequences. Amino acid sequences that can be usefully used as linkers include those described in Maratea et al. (1985) Gene, 40:39-46; Murphy et al. (1986) Gene, 40:39-46; Murphy et al. (1987) Gene, 40:39-46; each of which is incorporated herein by reference in its entirety. Examples of suitable linker sequences include those disclosed in U.S. Patent No. 4,935,233 (1986) Proc. Natl. Acad. Sci. USA, 83:8258-8262; U.S. Patent No. 4,935,233; and U.S. Patent No. 4,751,180. Linker sequences generally may be, for example, from 1 to about 50 amino acids in length. Linker sequences are generally not required when the first and second polypeptides have non-essential N-terminal amino acid regions that can be used to separate functional domains and prevent steric interference.

[0208] Protein can also be operably linked to cell-penetrating domain.For example, cell-penetrating domain can be derived from HIV-1 TAT protein, TLM cell-penetrating motif from human hepatitis B virus, MPG, Pep-1, VP22, cell-penetrating peptide from herpes simplex virus, or polyarginine peptide sequence.For example, see WO2014 / 089290, the entirety of which is incorporated herein by reference for all purposes.Cell-penetrating domain can be located at the N-terminus, C-terminus, or anywhere within protein.

[0209] The protein can also be operably linked to a heterologous polypeptide to facilitate tracking or purification, such as a fluorescent protein, a purification tag, or an epitope tag. Examples of fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, Emerald, Azami Green, Monomeric Azami). Green, CopGFP, AceGFP, ZsGreenl), yellow fluorescent proteins (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, ZsYellowl), blue fluorescent proteins (e.g., eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire), cyan fluorescent proteins (e.g., eCFP, Cerulean, CyPet, AmCyanl, Midoriishi-Cyan), red fluorescent proteins (e.g., mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRedl, AsRed2, eqFP611, mRaspberry, mStrawberry, Jred), orange fluorescent proteins (e.g., mOrange, mKO, Kusabira-Orange, Monomeric Examples of tags include glutathione-S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, hemagglutinin (HA), nus, Softag1, Softag3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, histidine (His), biotin carboxyl carrier protein (BCCP), and calmodulin.

[0210] The isolated proteins herein may also include unnatural or modified amino acids or peptide analogs. For example, there are many D-amino acids or amino acids with functional substituents different from the naturally occurring amino acids. The opposite stereoisomers of naturally occurring peptides, as well as stereoisomers of peptide analogs, are disclosed. These amino acids can be easily incorporated into polypeptide chains by charging a tRNA molecule with the selected amino acid and engineering a genetic construct that utilizes, for example, an amber codon to site-specifically insert the analog amino acid into the peptide chain (Thorson et al. (1991) Methods Molec. Biol., 77:43-73; Zoller (1992) Current Opinion in Biotechnology 3:348-354; Ibba, (1995) Biotechnology & Genetic Engineering Reviews, 13:197-216; Cahill et al. (1989) TIBS, 14(10 Issue): pp. 400-403; Benner (1993) TIB Tech, Vol. 12: pp. 158-163 and Ibba and Hennecke (1994) Biotechnology 12:678-682 pages, each of which is incorporated herein by reference in its entirety for all purposes).

[0211] Molecules can be created that resemble peptides but are not connected by natural peptide linkages. For example, amino acid or amino acid analog linkages can include CHNH--, --CHS--, --CH----, --CH=CH-- (cis and trans), --COCH--, --CH(OH)CH--, and --CHHSO-- (see, e.g., Spatola, A.F., Chemistry and Biochemistry of Amino Acids, Peptides, and Proteins, B. Weinstein (ed.), Marcel Dekker, New York, p. 267 (1983), each of which is incorporated herein by reference in its entirety for all purposes. Spatola, AF, Vega Data (March 1983), Vol. 1, No. 3, Peptide Backbone Modifications (general review); Morley (1994) Trends Pharm Sci, 15(12):463-468; Hudson et al. (1979) Int J Pept Prot Res, 14:177-185; Spatola et al. (1986) Life Sci. 38:1243-1249; Hann (1982) Chem. Soc. Perkin Trans. 1 307-314; Almquist et al. (1980) J. Med. Chem. 23:1392-1398; Jennings-White et al. (1982) Tetrahedron Lett. 23:2533; Szelke et al., European Application EP45665CA (1982): 97:39405 (1982); Holladay et al. (1983) Tetrahedron. Lett. 24:4401-4404; and Hruby (1 (See Life Sci, 31:189-199, 1982). For example, b-ara Peptide analogs, such as thiamin, aminobutyric acid, etc., can have more than one atom between the bond atoms.

[0212] Amino acid analogs and peptide analogs often have enhanced or desirable properties, such as more economical production, greater chemical stability, enhanced pharmacological properties (half-life, absorption, potency, efficacy, etc.), altered specificity (e.g., broad spectrum biological activity), reduced antigenicity, and other desirable properties.

[0213] Because D-amino acids are not recognized by peptidases and the like, D-amino acids can be used to generate more stable peptides. Systematic substitution of one or more amino acids of a consensus sequence with a D-amino acid of the same type (e.g., D-lysine instead of L-lysine) can be used to generate more stable peptides. Cysteine ​​residues can be used to cyclize or attach two or more peptides together. This can be beneficial for constraining peptides to a particular conformation (see, e.g., Rizo and Gierasch (1992) Ann. Rev. Biochem., 61:387, incorporated herein by reference in its entirety for all purposes).

[0214] The nucleic acids encoding any of the proteins disclosed herein are also disclosed herein.This includes all degenerate sequences related to a specific polypeptide sequence (i.e., all nucleic acids having a sequence that encodes one specific polypeptide sequence, and all nucleic acids including degenerate nucleic acids that encode the disclosed variants and derivatives of protein sequences).Therefore, it is not possible to describe each and every specific nucleic acid sequence herein, but each and every sequence is actually disclosed and described herein through the disclosed polypeptide sequence.

[0215] Also disclosed herein are compositions comprising an isolated polypeptide or protein disclosed herein and a carrier that enhances the stability of the isolated polypeptide, including, but not limited to, poly(lactic acid) (PLA) microspheres, poly(D,L-lactic acid-co-glycolic acid) (PLGA) microspheres, liposomes, micelles, reverse micelles, lipid cocrystals, and lipid microtubules.

[0216] (1) HSD17B13 protein and fragments Disclosed herein are isolated HSD17B13 proteins and fragments thereof, particularly HSD17B13 proteins and fragments thereof caused by the HSD17B13 rs72613567 variant, or in particular HSD17B13 isoforms C, D, E, F, F', G, and H. Such proteins can include, for example, isolated polypeptides comprising at least 5, 6, 8, 10, 12, 14, 15, 16, 18, 20, 22, 24, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, or 300 consecutive amino acids of HSD17B13 isoforms C, D, E, F, F', G, or H, or fragments thereof. It is understood that gene sequences in a population and the proteins encoded by such genes may vary due to polymorphisms such as single nucleotide polymorphisms.The sequences presented herein for each HSD17B13 isoform are merely exemplary sequences.Other sequences are also possible.For example, when the isolated polypeptide is optimally aligned with each of isoforms C, D, E, F, F', G or H, it comprises an amino acid sequence (e.g., a sequence of consecutive amino acids) that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to HSD17B13 isoforms C, D, E, F, F', G or H.Optionally, the isolated polypeptide comprises an identical sequence to HSD17B13 isoforms C, D, E, F, F', G or H.

[0217] As an example, an isolated polypeptide can include a segment (e.g., at least 8 contiguous amino acids) present in isoforms D, G, and H (or fragments or homologs thereof) that is not present in isoform A (or fragments or homologs thereof). Such regions can be readily identified by comparing the sequences of the isoforms. The region encoded by exon 7 in isoforms D, G, and H is frameshifted and shortened compared to the region encoded by exon 7 in isoform A. Thus, such isolated polypeptides can comprise at least 5, 6, 8, 10, 12, 14, 15, 16, 18, 20, 22, 24, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, or 200 contiguous amino acids of an HSD17B13 protein (e.g., at least 8 contiguous amino acids, at least 10 contiguous amino acids, or at least 15 contiguous amino acids of an HSD17B13 protein), where contiguous amino acids (e.g., at least 3 contiguous amino acids, at least 5 contiguous amino acids, at least 8 contiguous amino acids, at least 10 contiguous amino acids, at least 15 contiguous amino acids, at least 10 ... The segment of at least 10 contiguous amino acids, at least 10 contiguous amino acids, or at least 15 contiguous amino acids is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment comprising at least a portion of the region encoded by exon 7 in SEQ ID NO: 15 (HSD17B13 isoform D), SEQ ID NO: 18 (HSD17B13 isoform G), or SEQ ID NO: 19 (HSD17B13 isoform H) when the isolated polypeptide is optimally aligned with SEQ ID NO: 15, 18, or 19, respectively.

[0218] Such isolated polypeptides may further comprise a segment present in isoform D (or a fragment or homolog thereof) that is absent in isoform G (or a fragment or homolog thereof), and may further comprise a segment present in isoform D (or a fragment or homolog thereof) that is absent in isoform H (or a fragment or homolog thereof). Such regions can be readily identified by comparing the sequences of the isoforms. For example, such isolated polypeptides may comprise a segment of contiguous amino acids (e.g., at least 3 contiguous amino acids, at least 5 contiguous amino acids, at least 8 contiguous amino acids, at least 10 contiguous amino acids, or at least 15 contiguous amino acids) that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a segment spanning the boundary between the region encoded by exon 3 and the region encoded by exon 4 of SEQ ID NO: 15 (HSD17B13 isoform D) when optimally aligned with SEQ ID NO: 15, as distinguished from isoform H. Similarly, such an isolated polypeptide may comprise a segment of consecutive amino acids (e.g., at least 3 consecutive amino acids, at least 5 consecutive amino acids, at least 8 consecutive amino acids, at least 10 consecutive amino acids, or at least 15 consecutive amino acids) that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a segment within the region encoded by exon 2 of SEQ ID NO: 15 (HSD17B13 isoform D), a segment spanning the boundary between the region encoded by exon 1 and the region encoded by exon 2 of SEQ ID NO: 15, or a segment spanning the boundary between the region encoded by exon 2 and the region encoded by exon 3 of SEQ ID NO: 15, when optimally aligned with SEQ ID NO: 15, as distinguished from isoform G.

[0219] Similar to isoform D, the region encoded by exon 7 in isoform H (SEQ ID NO: 19) is frameshifted and shortened compared to isoform A. However, in addition, isoform H contains a region encoded by an additional exon between exons 3 and 4 (exon 3') compared to isoforms A and D. Thus, such isolated polypeptides can be as described above, and include segments present in isoforms D, G, and H (or fragments or homologs thereof) that are absent in isoform A (or fragments or homologs thereof), but further contain a segment (e.g., at least 8 contiguous amino acids) from isoform H (or a fragment or homolog thereof) that is absent in isoform D (or a fragment or homolog thereof). Such regions can be readily identified by comparing the sequences of the isoforms. For example, such an isolated polypeptide may further comprise a segment of consecutive amino acids (e.g., at least 3 consecutive amino acids, at least 5 consecutive amino acids, at least 8 consecutive amino acids, at least 10 consecutive amino acids, or at least 15 consecutive amino acids) that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment comprising at least a portion of the region encoded by exon 3' of SEQ ID NO: 19 (HSD17B13 isoform H) when the isolated polypeptide is optimally aligned with SEQ ID NO: 19.

[0220] Similar to isoform D, the region encoded by exon 7 in isoform G (SEQ ID NO: 18) is frameshifted and shortened compared to isoform A. However, in addition, compared to isoforms A and D, isoform G lacks the region encoded by exon 2, and thus includes the boundary between exon 1 and exon 3, which is absent in isoforms A and D. Thus, such isolated polypeptides can be as described above, and include segments present in isoforms D, G, and H (or fragments or homologs thereof) that are absent in isoform A (or fragments or homologs thereof), but further include a segment (e.g., at least 8 contiguous amino acids) from isoform G (or a fragment or homolog thereof) that is absent in isoform D (or a fragment or homolog thereof). Such regions can be readily identified by comparing the sequences of the isoforms. For example, such an isolated polypeptide may further comprise a segment of consecutive amino acids (e.g., at least 3 consecutive amino acids, at least 5 consecutive amino acids, at least 8 consecutive amino acids, at least 10 consecutive amino acids, or at least 15 consecutive amino acids) that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment spanning the boundary between the region encoded by exon 1 and the region encoded by exon 3 of SEQ ID NO: 18 (HSD17B13 isoform G) when the isolated polypeptide is optimally aligned with SEQ ID NO: 18.

[0221] Also provided herein is an isolated polypeptide comprising a segment (e.g., at least 8 contiguous amino acids) present in isoform E (or a fragment or homolog thereof) that is not present in isoform A (or a fragment or homolog thereof). Isoform E contains a region encoded by an additional exon between exons 3 and 4 (exon 3') that is not present in isoform A. Such a region can be readily identified by comparing the sequences of the isoforms. Thus, an isolated polypeptide can comprise at least 5, 6, 8, 10, 12, 14, 15, 16, 18, 20, 22, 24, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, or 200 contiguous amino acids of an HSD17B13 protein (e.g., at least 8 contiguous amino acids, at least 10 contiguous amino acids, or at least 15 contiguous amino acids of an HSD17B13 protein), where contiguous amino acids (e.g., at least 3 contiguous amino acids, at least 5 contiguous amino acids, at least 6 contiguous amino acids, at least 7 contiguous amino acids, at least 8 contiguous amino acids, at least 9 contiguous amino acids, or at least 10 contiguous amino acids) are not contiguous. A segment of at least 8 contiguous amino acids, at least 8 contiguous amino acids, at least 10 contiguous amino acids, or at least 15 contiguous amino acids) is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment comprising at least a portion of the region encoded by exon 3' of SEQ ID NO: 16 (HSD17B13 isoform E) or SEQ ID NO: 19 (HSD17B13 isoform H) when the isolated polypeptide is optimally aligned with SEQ ID NO: 16 or 19, respectively. Optionally, such isolated polypeptides can further comprise a segment (e.g., at least 8 contiguous amino acids) from isoform E (or a fragment or homolog thereof) that is not present in isoform H (or a fragment or homolog thereof). Such regions can be readily identified by comparing the sequences of the isoforms.For example, such an isolated polypeptide may further comprise a segment of consecutive amino acids (e.g., at least 3 consecutive amino acids, at least 5 consecutive amino acids, at least 8 consecutive amino acids, at least 10 consecutive amino acids, or at least 15 consecutive amino acids) that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment spanning the boundary between the region encoded by exon 6 and the region encoded by exon 7 of SEQ ID NO: 16 (HSD17B13 isoform E) when the isolated polypeptide is optimally aligned with SEQ ID NO: 16.

[0222] Also provided is an isolated polypeptide comprising a segment (e.g., at least 8 contiguous amino acids) present in isoform F (or a fragment or homolog thereof) that is absent in isoform A (or a fragment or homolog thereof). Isoform F contains a region encoded by readthrough from exon 6 to intron 6 that is absent in isoform A. Such a region can be readily identified by comparing the sequences of the isoforms. Thus, an isolated polypeptide can comprise at least 5, 6, 8, 10, 12, 14, 15, 16, 18, 20, 22, 24, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, or 200 contiguous amino acids of an HSD17B13 protein (e.g., at least 8 contiguous amino acids, at least 10 contiguous amino acids, or at least 15 contiguous amino acids of an HSD17B13 protein), where contiguous amino acids (e.g., at least 3 contiguous amino acids) are not contiguous. The segment of SEQ ID NO: 17 (HSD17B13 isoform F) comprising at least a portion of the region encoded by the read-through into intron 6 of SEQ ID NO: 17 (HSD17B13 isoform F) is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment comprising at least a portion of the region encoded by the read-through into intron 6 of SEQ ID NO: 17 (HSD17B13 isoform F) when the isolated polypeptide is optimally aligned with SEQ ID NO: 17.

[0223] Also provided is an isolated polypeptide comprising a segment (e.g., at least 8 contiguous amino acids) present in isoform C (or a fragment or homolog thereof) that is absent in isoform A (or a fragment or homolog thereof). Compared to isoform A, isoform C lacks the region encoded by exon 6 and includes the boundary between exons 5 and 7, which is absent in isoform A. Such regions can be readily identified by comparing the sequences of the isoforms. Thus, an isolated polypeptide can comprise at least 5, 6, 8, 10, 12, 14, 15, 16, 18, 20, 22, 24, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, or 200 contiguous amino acids of an HSD17B13 protein (e.g., at least 8 contiguous amino acids, at least 10 contiguous amino acids, or at least 15 contiguous amino acids of an HSD17B13 protein), where contiguous amino acids (e.g., at least 3 contiguous amino acids) are not contiguous. The segment of SEQ ID NO: 14 (HSD17B13 isoform C) spanning the boundary between the region encoded by exon 5 and the region encoded by exon 7 of SEQ ID NO: 14 (HSD17B13 isoform C) is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a segment spanning the boundary between the region encoded by exon 5 and the region encoded by exon 7 of SEQ ID NO: 14 (HSD17B13 isoform C) when the isolated polypeptide is optimally aligned with SEQ ID NO: 14.

[0224] Any of the isolated polypeptides disclosed herein can be linked to a heterologous molecule or heterologous label. Examples of such heterologous molecules or labels are disclosed elsewhere herein. For example, the heterologous molecule can be an immunoglobulin Fc domain, a peptide tag disclosed elsewhere herein, poly(ethylene glycol), polysialic acid, or glycolic acid.

[0225] (2) Methods for producing HSD17B13 protein or fragments Also disclosed are methods for producing any of the HSD17B13 proteins or fragments thereof disclosed herein. Such HSD17B13 proteins or fragments thereof can be produced by any suitable method. For example, HSD17B13 proteins or fragments thereof can be produced from host cells containing nucleic acids (e.g., recombinant expression vectors) encoding such HSD17B13 proteins or fragments thereof. Such methods may include culturing host cells containing nucleic acids (e.g., recombinant expression vectors) encoding HSD17B13 proteins or fragments thereof, thereby producing HSD17B13 proteins or fragments thereof. The nucleic acid may be operably linked to a promoter active in the host cells, and culturing may be performed under conditions in which the nucleic acid is expressed. Such methods may further include recovering the expressed HSD17B13 proteins or fragments thereof. The recovering step may further include purifying the HSD17B13 proteins or fragments thereof.

[0226] Examples of suitable systems for protein expression include bacterial cell expression systems (e.g., Escherichia coli, Lactococcus lactis), yeast cell expression systems (e.g., Saccharomyces cerevisiae, Pichia pastoris), insect cell expression systems (e.g., baculovirus-mediated protein expression), and mammalian cell expression systems.

[0227] Examples of nucleic acids encoding HSD17B13 proteins or fragments thereof are disclosed in more detail elsewhere herein. If necessary, such nucleic acids are codon-optimized for expression in the host cell. If necessary, such nucleic acids are operably linked to a promoter active in the host cell. The promoter may be a heterologous promoter (i.e., a promoter other than the naturally occurring HSD17B13 promoter). Examples of promoters suitable for Escherichia coli include the arabinose, lac, tac, and T7 promoters. Examples of promoters suitable for Lactococcus lactis include the P170 and nisin promoters. Examples of promoters suitable for Saccharomyces cerevisiae include constitutive promoters such as the alcohol dehydrogenase (ADHI) or enolase (ENO) promoters, or inducible promoters such as PHO, CUP1, GAL1, and G10. Examples of promoters suitable for Pichia pastoris include the alcohol oxidase I (AOX I) promoter, the glyceraldehyde-3-phosphate dehydrogenase (GAP) promoter, and the glutathione-dependent formaldehyde dehydrogenase (FLDI) promoter. An example of a promoter suitable for baculovirus-mediated systems is the late virus strong polyhedrin promoter.

[0228] Optionally, the nucleic acid further encodes a tag in frame with the HSD17B13 protein or a fragment thereof to facilitate protein purification. Examples of tags are disclosed elsewhere herein. Such tags can, for example, bind to a partner ligand (e.g., immobilized on a resin), thereby allowing the tagged protein to be isolated from all other proteins (e.g., host cell proteins). Affinity chromatography, high-performance liquid chromatography (HPLC), and size-exclusion chromatography (SEC) are examples of methods that can be used to improve the purity of expressed proteins.

[0229] Other methods can also be used to produce HSD17B13 proteins or fragments thereof. For example, two or more peptides or polypeptides can be linked together using protein chemistry techniques. For example, peptides or polypeptides can be chemically synthesized using either Fmoc (9-fluorenylmethyloxycarbonyl) or Boc (tert-butyloxycarbonyl) chemistry. Such peptides or polypeptides can be synthesized by standard chemical reactions. For example, a peptide or polypeptide can be synthesized without cleavage from the synthesis resin, while the other fragment of the peptide or protein can be synthesized and then cleaved from the resin, thereby exposing a functionally blocked terminal group on the other fragment. A peptide condensation reaction can covalently join these two fragments via a peptide bond at their carboxyl and amino termini, respectively. (Grant GA (1999) (1992) Synthetic Peptides: A User Guide. W.H. Freeman and Co., NY (1992); and Bodansky M and Trost B., eds. (1993) Principles of Peptide Synthesis. Springer-Verlag Inc., NY, each of which is incorporated herein by reference in its entirety for all purposes. Alternatively, peptides or polypeptides can be independently synthesized in vivo as described herein. Once isolated, these independent peptides or polypeptides can be linked by similar peptide condensation reactions to form peptides or fragments thereof.

[0230] For example, enzymatic ligation of cloned or synthetic peptide segments allows for the joining of relatively short peptide fragments to create larger peptide fragments, polypeptides, or entire protein domains (Abrahmsen L et al. (1991) Biochemistry 30:4151, incorporated herein by reference in its entirety for all purposes). Alternatively, native chemical ligation of synthetic peptides can be utilized to synthetically construct large peptides or polypeptides from short peptide fragments. This method can consist of a two-step chemical reaction (Dawson et al. (1994) Science 266:776-779, incorporated herein by reference in its entirety for all purposes). The first step can be the chemoselective reaction of an unprotected synthetic peptide—a thioester—with another unprotected peptide segment containing an amino-terminal Cys residue, yielding a thioester-linked intermediate as the first covalently coupled product. If the reaction conditions are not changed, this intermediate spontaneously undergoes a rapid intramolecular reaction to form a native peptide bond at the ligation site (Baggiolini et al. (1992) FEBS Lett 307:97-101; Clark-Lewis et al. (1994) J Biol Chem 269:16075; Clark-Lewis et al. (1991) Biochemistry 30:3128). and Rajarathnam et al. (1994) Biochemistry 33:6623-6630 each of which is incorporated herein by reference in its entirety for all purposes).

[0231] Alternatively, unprotected peptide segments can be chemically linked, in which case the bond formed between the peptide segments as a result of chemical ligation is an unnatural (non-peptide) bond (Schnolzer et al. (1992) Science 256:221, incorporated herein by reference in its entirety for all purposes). This technique has been used to synthesize analogs of protein domains as well as large quantities of relatively pure, fully biologically active proteins (deLisle Milton RC et al., Techniques in Protein Chemistry IV. Academic Press, New York, pp. 257-267 (199 2 years), which is incorporated herein by reference in its entirety for all purposes).

[0232] C. Cell Also provided herein are cells (e.g., recombinant host cells) containing any of the nucleic acids and proteins disclosed herein. The cells can be in vitro, ex vivo, or in vivo. The nucleic acids can be linked to promoters and other control sequences, so that they can be expressed to produce the encoded proteins. Any cell type is provided.

[0233] The cells may be, for example, totipotent or pluripotent cells (e.g., embryonic stem (ES) cells such as rodent ES cells, mouse ES cells, or rat ES cells). Totipotent cells include undifferentiated cells that can give rise to any cell type, while pluripotent cells include undifferentiated cells that have the ability to develop into more than one differentiated cell type. Such pluripotent and / or totipotent cells may be, for example, ES cells or ES-like cells, such as induced pluripotent stem (iPS) cells. ES cells include embryo-derived totipotent or pluripotent cells that, when introduced into an embryo, can contribute to any tissue of the developing embryo. ES cells can be derived from the inner cell mass of a blastocyst and can be differentiated into cells of any of the three vertebrate germ layers (endoderm, ectoderm, and mesoderm).

[0234] The cells may be primary somatic cells or non-primary somatic cells. Somatic cells may include any cells that are not gametes, germ cells, gametocytes, or undifferentiated stem cells. The cells may also be primary cells. Primary cells include cells or cell cultures isolated directly from an organism, organ, or tissue. Primary cells include cells that are not transformed or immortal. Primary cells include any cells obtained from an organism, organ, or tissue that have not previously been passaged in tissue culture or that have previously been passaged in tissue culture but cannot be passaged indefinitely in tissue culture. Such cells can be isolated by conventional techniques and include, for example, somatic cells, hematopoietic cells, endothelial cells, epithelial cells, fibroblasts, mesenchymal cells, keratinocytes, melanocytes, monocytes, mononuclear cells, adipocytes, preadipocytes, neurons, glial cells, hepatocytes, skeletal myoblasts, and smooth muscle cells. For example, the primary cells may be derived from connective tissue, muscle tissue, nervous system tissue, or epithelial tissue.

[0235] Such cells do not usually proliferate indefinitely, but also include cells that can escape normal cellular senescence and continue to divide due to mutations or alterations. Such mutations or alterations can occur naturally or can be intentionally induced. Examples of immortalized cells include Chinese hamster ovary (CHO) cells, human embryonic kidney cells (e.g., HEK293 cells), and mouse embryonic fibroblasts (e.g., 3T3 cells). Many types of immortalized cells are well known. Immortalized or primary cells include cells that are commonly used for culturing or expressing recombinant genes or proteins.

[0236] The cell may also be a differentiated cell, such as a hepatocyte (eg, a human hepatocyte).

[0237] The cell may be derived from any source. For example, the cell may be a eukaryotic cell, an animal cell, a plant cell, or a fungal (e.g., yeast) cell. Such a cell may be a fish cell or an avian cell, and such a cell may be a mammalian cell, such as a human cell, a non-human mammalian cell, a rodent cell, a mouse cell, or a rat cell. Mammals include, for example, humans, non-human primates, monkeys, apes, cats, dogs, horses, oxen, deer, bison, sheep, rodents (e.g., mice, rats, hamsters, guinea pigs), and livestock animals (e.g., bovine species such as cows and steers; ovine species such as sheep and goats; and porcine species such as pigs and wild boars). Birds include, for example, chickens, turkeys, ostriches, geese, ducks, etc. Domestic animals and agricultural animals are also included. The term "non-human animal" excludes humans.

[0238] With respect to mouse cells, the mice can be of any strain, including, for example, 129, C57BL / 6, BALB / c, Swiss Webster, a mix of 129 and C57BL / 6, a mix of BALB / c and C57BL / 6, a mix of 129 and BALB / c, and a mix of BALB / c, C57BL / 6, and 129 strains. For example, the mice can be at least partially derived from the BALB / c strain (e.g., at least about 25%, at least about 50%, at least about 75% derived from the BALB / c strain, or about 25%, about 50%, about 75%, or about 100% derived from the BALB / c strain). In one example, the mice are a strain that includes 50% BALB / c, 25% C57BL / 6, and 25% 129. Alternatively, the mice include a strain or combination of strains that excludes BALB / c.

[0239] Examples of 129 lineages include 129P1, 129P2, 129P3, 129X1, 129S1 (e.g., 129S1 / SV, 129S1 / Svlm), 129S2, 129S4, 129S5, 129S9 / SvEvH, 129S6 (129 / SvEvTac), 129S7, 129S8, 129T1, and 129T2. See, e.g., Festing et al. (1999) Mammalian Genome 10(8):836, which is incorporated herein by reference in its entirety for all purposes. Examples of C57BL strains include C57BL / A, C57BL / An, C57BL / GrFa, C57BL / Kal_wN, C57BL / 6, C57BL / 6J, C57BL / 6ByJ, C57BL / 6NJ, C57BL / 10, C57BL / 10ScSn, C57BL / 10Cr, and C57BL / Ola. Mouse cells may also be derived from a mix of the above-mentioned 129 strain and the above-mentioned C57BL / 6 strain (e.g., 50% 129 and 50% C57BL / 6). Similarly, mouse cells may be derived from a mix of the above-mentioned 129 strains or a mix of the above-mentioned BL / 6 strains (e.g., 129S6 (129 / SvEvTac) strain).

[0240] For rat cells, rats may be, for example, ACI rat strain, Dark Agouti (DA) rat strain, Wistar rat strain, LEA rat strain, Sprague Dawley (SD) rat strain, or Fisher F344 or Fisher The rats may be any rat strain, including Fischer rat strains such as F6. The rats may also be derived from a strain derived from a mix of two or more of the above strains. For example, the rats may be derived from the DA strain or the ACI strain. The ACI rat strain has a black agouti color, a white abdomen and paws, and an RT1 av1 Such strains are available from a variety of sources, including Harlan Laboratories. The Dark Agouti (DA) rat strain possesses the agouti coat and RT1 haplotype. av1The rat is characterized as having haplotype.Such rats can be obtained from various sources, including Charles River and Harlan Laboratories.In some cases, rats are derived from inbred rat strains.See, for example, US2014 / 0235933A1, the entirety of which is incorporated herein by reference for all purposes.

[0241] III. Methods of modifying or altering the expression of HSD17B13 Various methods are provided for modifying cells by using any combination of nuclease agents, exogenous donor sequences, transcriptional activators, transcriptional repressors, antisense molecules such as antisense RNA, siRNA, and shRNA, HSD17B13 proteins or fragments thereof, and expression vectors for expressing recombinant HSD17B13 genes or nucleic acids encoding HSD17B13 proteins.The methods can be performed in vitro, ex vivo, or in vivo.The nuclease agents, exogenous donor sequences, transcriptional activators, transcriptional repressors, antisense molecules such as antisense RNA, siRNA, and shRNA, HSD17B13 proteins or fragments thereof, and expression vectors can be introduced into cells in any form by any means described elsewhere herein, and can be introduced in whole or in part in any combination simultaneously or sequentially.Some methods only involve modifying the endogenous HSD17B13 gene in cells. Some methods involve only altering the expression of endogenous HSD17B13 genes by using transcriptional activators or repressors, or by using antisense molecules such as antisense RNA, siRNA, and shRNA. Some methods involve only introducing recombinant HSD17B13 genes or nucleic acids encoding HSD17B13 proteins, or fragments thereof, into cells. Some methods involve only introducing HSD17B13 proteins or fragments thereof (e.g., any one or any combination of HSD17B13 proteins or fragments thereof disclosed herein, or any one or any combination of HSD17B13 isoforms A-H or fragments thereof disclosed herein) into cells. For example, such methods may involve introducing one or more of HSD17B13 isoforms C, D, F, G, and H (or fragments thereof) into cells, or introducing HSD17B13 isoform D (or fragments thereof) into cells.Alternatively, such methods may involve introducing one or more of HSD17B13 isoforms A, B, and E or isoforms A, B, E, and F' (or fragments thereof) into a cell, or introducing HSD17B13 isoform A (or a fragment thereof) into a cell. Other methods may involve both altering the endogenous HSD17B13 gene in a cell and introducing into a cell an HSD17B13 protein or fragment thereof, or a recombinant HSD17B13 gene or nucleic acid encoding an HSD17B13 protein, or fragment thereof. Still other methods may involve both altering the expression of an endogenous HSD17B13 gene in a cell and introducing into a cell an HSD17B13 protein or fragment thereof, or a recombinant HSD17B13 gene or nucleic acid encoding an HSD17B13 protein, or fragment thereof.

[0242] A. Methods of Modifying HSD17B13 Nucleic Acids Various methods are provided for modifying the HSD17B13 gene in the genome of a cell (e.g., a pluripotent cell or a differentiated cell such as a hepatocyte) by using a nuclease agent and / or an exogenous donor sequence. The methods can be performed in vitro, ex vivo, or in vivo. The nuclease agent can be used alone or in combination with the exogenous donor sequence. Alternatively, the exogenous donor sequence can be used alone or in combination with the nuclease agent.

[0243] Repair in response to double-strand breaks (DSBs) mainly occurs through two conservative DNA repair pathways: non-homologous end joining (NHEJ) and homologous recombination (HR).See Kasparek and Humphrey (2011) Seminars in Cell & Dev. Biol., vol. 22: 886-897, the entire contents of which are incorporated herein by reference for all purposes.NHEJ involves the repair of double-strand breaks in nucleic acids by directly ligating the cut ends with each other or with exogenous sequences without the need for a homologous template.The ligation of non-contiguous sequences by NHEJ can often result in deletions, insertions, or translocations near the site of double-strand breaks.

[0244] Repair of a target nucleic acid (e.g., the HSD17B13 gene) mediated by an exogenous donor sequence can involve any process of exchanging genetic information between two polynucleotides. For example, NHEJ can also result in targeted integration of the exogenous donor sequence by direct ligation of the cut ends with the ends of the exogenous donor sequence (i.e., NHEJ-based capture). Such NHEJ-mediated targeted integration may be preferable for inserting exogenous donor sequences when the homology-directed repair (HDR) pathway is not readily available (e.g., in non-dividing cells, primary cells, and cells where homology-based DNA repair is poorly performed). Furthermore, in contrast to homology-directed repair, knowledge of large regions of sequence identity flanking the cut site (beyond the overhang created by Cas-mediated cleavage) is not required, which may be beneficial when attempting targeted insertion into organisms with genomes with limited knowledge of the genomic sequence. Integration can proceed by blunt-end ligation between the exogenous donor sequence and the cleaved genomic sequence, or by sticky-end (i.e., with 5' or 3' overhangs) ligation using the exogenous donor sequence flanked by overhangs compatible with those generated by the Cas protein in the cleaved genomic sequence (see, e.g., US2011 / 020722, WO2014 / 033644, WO2014 / 089290, and Maresca et al. (2013) Genome Res. 23(3):539-54, each of which is incorporated by reference in its entirety for all purposes. See page 6. When ligating blunt ends, target and / or donor excision may be required in regions generating microhomologies necessary for fragment joining, which may result in undesired alterations to the target sequence.

[0245] Repair can also occur through homology-directed repair (HDR) or homologous recombination (HR). HDR or HR may require nucleotide sequence homology and include forms of nucleic acid repair that use a "donor" molecule as a template to repair a "target" molecule (i.e., a molecule in which a double-strand break has occurred), leading to the transfer of genetic information from the donor to the target. Without being bound by any particular theory, such transfer may involve mismatch correction of the heteroduplex DNA formed between the cut target and the donor, and / or synthesis-dependent strand annealing, in which the donor is used to resynthesize the genetic information that will become part of the target, and / or related processes. In some cases, the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide, or a portion of the copy of the donor polynucleotide is incorporated into the target DNA. See Wang et al. (2013) Cell 153:910-918; Mandalos et al. (2012) PLOS ONE 7:e45768:1-9; and Wang et al. (2013) Nat Biotechnol. 31:530-532, each of which is incorporated by reference in its entirety for all purposes.

[0246] The targeted genetic modification of HSD17B13 gene in genome can be caused by contacting cell with exogenous donor sequence, which comprises a 5' homologous arm that hybridizes with the 5' target sequence of the target genome locus in HSD17B13 gene and a 3' homologous arm that hybridizes with the 3' target sequence of the target genome locus in HSD17B13 gene.The exogenous donor sequence can be recombined with the target genome locus to cause the targeted genetic modification of HSD17B13 gene.For example, when HSD17B13 gene is optimally aligned with SEQ ID NO:2, the 5' homologous arm can be hybridized with the target sequence on the 5' side of the position corresponding to position 12666 of SEQ ID NO:2, and the 3' homologous arm can be hybridized with the target sequence on the 3' side of the position corresponding to position 12666 of SEQ ID NO:2. Such a method can result in, for example, an HSD17B13 gene containing a thymine inserted between the nucleotides corresponding to positions 12665 ​​and 12666 of SEQ ID NO: 1 (or an adenine inserted at the corresponding position on the opposite strand) when the HSD17B13 gene is optimally aligned with SEQ ID NO: 1. As another example, the 5' and 3' homology arms can hybridize to 5' and 3' target sequences, respectively, at positions corresponding to positions flanking exon 6 of SEQ ID NO: 1 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 1. Such a method can result in, for example, an HSD17B13 gene lacking the sequence corresponding to exon 6 of SEQ ID NO: 1 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 1. As another example, the 5' and 3' homology arms can be hybridized to 5' and 3' target sequences, respectively, at positions corresponding to positions flanking exon 2 of SEQ ID NO: 1 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 1. Such a method can result in, for example, an HSD17B13 gene lacking the sequence corresponding to exon 2 of SEQ ID NO: 1 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 1.As another example, the 5' and 3' homology arms can be hybridized to 5' and 3' target sequences, respectively, at positions corresponding to the exon 6 / intron 6 boundary of SEQ ID NO: 1 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 1. As another example, the 5' and 3' homology arms can be hybridized to 5' and 3' target sequences, respectively, at positions corresponding to exon 6 and exon 7 of SEQ ID NO: 1 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 1. Such a method can result in, for example, an HSD17B13 gene containing a thymine inserted between the nucleotides corresponding to positions 12665 ​​and 12666 of SEQ ID NO: 1 (or an adenine inserted at the corresponding position on the opposite strand) when the HSD17B13 gene is optimally aligned with SEQ ID NO: 1. As another example, the 5' and 3' homology arms can be hybridized to 5' and 3' target sequences, respectively, at positions corresponding to positions that flank or are within the region corresponding to the donor splice site in intron 6 of SEQ ID NO: 1 (i.e., the region at the 5' end of intron 6 of SEQ ID NO: 1). Such methods can result in, for example, an HSD17B13 gene in which the donor splice site within intron 6 has been disrupted. Examples of exogenous donor sequences are disclosed elsewhere herein.

[0247] Targeted gene modification of HSD17B13 gene in genome can also be caused by contacting cell with nuclease agent, which induces one or more nicks or double-strand breaks in the target sequence at the target genomic locus in HSD17B13 gene.For example, this method can result in HSD17B13 gene that the region corresponding to the donor splice site in intron 6 of SEQ ID NO: 1 (i.e., the region at the 5' end of intron 6 of SEQ ID NO: 1) is destroyed.The examples and changes of nuclease agent that can be used in this method are described elsewhere herein.

[0248] For example, targeted genetic modification of the HSD17B13 gene in genome can be produced by contacting a cell or the genome of a cell with a Cas protein and one or more guide RNAs that hybridize with one or more guide RNA recognition sequences in the target genome locus of the HSD17B13 gene.That is, targeted genetic modification of the HSD17B13 gene in genome can be produced by contacting a cell or the genome of a cell with a Cas protein and one or more guide RNAs that target one or more guide RNA target sequences in the target genome locus of the HSD17B13 gene.For example, such a method can include contacting a cell with a Cas protein and a guide RNA that targets the guide RNA target sequence in the HSD17B13 gene.As an example, when the HSD17B13 gene is optimally aligned with SEQ ID NO:2, the guide RNA target sequence is within the region corresponding to exon 6 and / or intron 6 of SEQ ID NO:2. As one example, the guide RNA target sequence is within a region corresponding to exon 6 and / or intron 6 and / or exon 7 (e.g., exon 6 and / or intron 6, or exon 6 and / or exon 7) of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2. As another example, the guide RNA target sequence may include or be adjacent to a position corresponding to position 12666 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2. For example, the guide RNA target sequence may be within about 1000, 500, 400, 300, 200, 100, 50, 45, 40, 35, 30, 25, 20, 15, 10, or 5 nucleotides from a position corresponding to position 12666 of SEQ ID NO: 2 when the HSD17B13 gene is optimally aligned with SEQ ID NO: 2. As yet another example, the guide RNA target sequence may include or be adjacent to the start codon of the HSD17B13 gene or the stop codon of the HSD17B13 gene.For example, the guide RNA target sequence can be within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides from the start codon or stop codon. The Cas protein and guide RNA form a complex, and the Cas protein cleaves the guide RNA target sequence. Cleavage by the Cas protein can create a double-stranded break or a single-stranded break (e.g., when the Cas protein is a nickase). Such a method can result in an HSD17B13 gene in which, for example, the region corresponding to the donor splice site in intron 6 of SEQ ID NO: 1 (i.e., the region at the 5' end of intron 6 of SEQ ID NO: 1) is disrupted, the start codon is disrupted, the stop codon is disrupted, or the coding sequence is deleted. Examples and variations of Cas (e.g., Cas9) proteins and guide RNAs that can be used in the method are described elsewhere herein.

[0249] In some methods, two or more nuclease agents can be used.For example, two nuclease agents can be used, each of which targets the nuclease target sequence in the region corresponding to exon 6 and / or intron 6, or exon 6 and / or exon 7 of SEQ ID NO:2 when HSD17B13 gene is optimally aligned with SEQ ID NO:2, or includes or is close to the position corresponding to position 12666 of SEQ ID NO:2 when HSD17B13 gene is optimally aligned with SEQ ID NO:2 (for example, within about 1000, 500, 400, 300, 200, 100, 50, 45, 40, 35, 30, 25, 20, 15, 10 or 5 nucleotides from the position corresponding to position 12666 of SEQ ID NO:2 when HSD17B13 gene is optimally aligned with SEQ ID NO:2). For example, two nuclease agents can be used, each targeting a nuclease target sequence within the region corresponding to exon 6 and / or intron 6 and / or exon 7 of SEQ ID NO:2 when the HSD17B13 gene is optimally aligned with SEQ ID NO:2. As another example, two or more nuclease agents can be used, each targeting a nuclease target sequence that includes or is adjacent to the start codon. As another example, two nuclease agents can be used, one targeting a nuclease target sequence that includes or is adjacent to the start codon and one targeting a nuclease target sequence that includes or is adjacent to the stop codon, and cleavage by the nuclease agents can result in deletion of the coding region between the two nuclease target sequences. As yet another example, three or more nuclease agents can be used, with one or more (e.g., two) targeting the sequence containing or adjacent to the start codon and one or more (e.g., two) targeting the sequence containing or adjacent to the stop codon, and cleavage by the nuclease agents can result in deletion of the coding region between the nuclease target sequence containing or adjacent to the start codon and the nuclease target sequence containing or adjacent to the stop codon.

[0250] Optionally, the cell can be further contacted with one or more additional guide RNAs that target additional guide RNA target sequences within the target genomic locus of the HSD17B13 gene. By contacting the cell with one or more additional guide RNAs (e.g., a second guide RNA that targets a second guide RNA target sequence), cleavage by the Cas protein can create two or more double-strand breaks or two or more single-strand breaks (e.g., when the Cas protein is a nickase).

[0251] Optionally, the cells can be further contacted with one or more exogenous donor sequences that recombine with the target genomic locus of the HSD17B13 gene to produce a targeted genetic modification. Examples of exogenous donor sequences and variations that can be used in the method are disclosed elsewhere herein.

[0252] The Cas protein, guide RNA(s), and exogenous donor sequence(s) can be introduced into the cell in any form and by any means described elsewhere herein, and all or part of the Cas protein, guide RNA(s), and exogenous donor sequence(s) can be introduced simultaneously or sequentially in any combination.

[0253] In some such methods, the repair of the target nucleic acid (e.g., HSD17B13 gene) by the exogenous donor sequence occurs by homology-directed repair (HDR). Homologous recombination repair can occur when a Cas protein cleaves both strands of the DNA of the HSD17B13 gene to create a double-stranded break, when a Cas protein is a nickase that cleaves one strand of the DNA of the target nucleic acid to create a single-stranded break, or when a Cas nickase is used to create a double-stranded break formed by two offset nicks. In such methods, the exogenous donor sequence includes 5' and 3' homology arms corresponding to the 5' and 3' target sequences. The guide RNA target sequence(s) or cleavage site(s) can be adjacent to the 5' target sequence, adjacent to the 3' target sequence, adjacent to both the 5' target sequence and the 3' target sequence, or not adjacent to either the 5' target sequence or the 3' target sequence. Optionally, the exogenous donor sequence may further comprise a nucleic acid insert flanked by 5' and 3' homology arms, wherein the nucleic acid insert is inserted between the 5' target sequence and the 3' target sequence. If the nucleic acid insert is not present, the exogenous donor sequence functions to delete the genomic sequence between the 5' target sequence and the 3' target sequence. Examples of exogenous donor sequences are disclosed elsewhere herein.

[0254] Alternatively, the repair of HSD17B13 gene mediated by exogenous donor sequence can be carried out by non-homologous end joining (NHEJ) mediated ligation.In this method, at least one end of exogenous donor sequence comprises a short single-stranded region that is complementary to at least one protrusion created by Cas-mediated cleavage in HSD17B13 gene.The complementary ends in exogenous donor sequence can flank nucleic acid insert.For example, each end of exogenous donor sequence can comprise a short single-stranded region that is complementary to the protrusion created by Cas-mediated cleavage in HSD17B13 gene, and these complementary regions in exogenous donor sequence can flank nucleic acid insert.

[0255] Overhangs (i.e., sticky ends) can be created by excising the blunt ends of the double-stranded breaks created by Cas-mediated cleavage. Such excision can result in microhomology regions necessary for fragment joining, which may create undesirable or uncontrollable changes to the HSD17B13 gene. Alternatively, such overhangs can be created by using paired Cas nickases. For example, cells can be contacted with a first and a second nickase that cleave opposite strands of DNA, thereby modifying the genome through double nicking. This can be achieved by contacting cells with a first Cas protein nickase, a first guide RNA that targets a first guide RNA target sequence within the target genomic locus of the HSD17B13 gene, a second Cas protein nickase, and a second guide RNA that targets a second guide RNA target sequence within the target genomic locus of the HSD17B13 gene. The first Cas protein and the first guide RNA form a first complex, the second Cas protein and the second guide RNA form a second complex, the first Cas protein nickase cleaves the first strand of genomic DNA within the first guide RNA target sequence, the second Cas protein nickase cleaves the second strand of genomic DNA within the second guide RNA target sequence, and optionally, the exogenous donor sequence recombines with the target genomic locus of the HSD17B13 gene to produce the targeted genetic modification.

[0256] The first nickase can cleave the first strand (i.e., the complementary strand) of genomic DNA, and the second nickase can cleave the second strand (i.e., the non-complementary strand) of genomic DNA. The first and second nickases can be created, for example, by mutating catalytic residues in the RuvC domain of Cas9 (e.g., the D10A mutation described elsewhere herein) or by mutating catalytic residues in the HNH domain of Cas9 (e.g., the H840A mutation described elsewhere herein). In such methods, double nicking can be used to create a double-stranded break with a sticky end (i.e., an overhang). The first and second guide RNA target sequences can be positioned such that the cleavage sites are created, resulting in a double-stranded break, due to the nicks created by the first and second nickases on the first and second strands of DNA. The overhang is created when the nicks in the first and second CRISPR RNA target sequences are offset. The offset window can be, for example, at least about 5 bp, 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, or larger, as described, for example, in Ran et al. (2013) Cell, 1999; 154:1380-1389; Mali et al. (2013) Nat. Biotech., 31:833-838; and Shen et al. (2014) Nat. Methods, 11:399-404 Please refer to.

[0257] (1) Type of targeted gene modification The methods described herein can be used to introduce various types of targeted gene modifications. Such targeted modifications can include, for example, the addition of one or more nucleotides, the deletion of one or more nucleotides, the substitution of one or more nucleotides, point mutations, or a combination thereof. For example, at least one, two, three, four, five, seven, eight, nine, ten, or more nucleotides can be changed (e.g., deleted, inserted, or substituted) to form targeted genome modifications. Deletions, insertions, or substitutions can be of any size, as disclosed elsewhere herein. For example, see Wang et al. (2013) Cell, 153:910-918; Mandalos et al. (2012) PLOS ONE, 7:e45768:1-9; and Wang et al. (2013) Nat Biotechnol., 31:530-532, each of which is incorporated herein by reference in its entirety for all purposes.

[0258] Such targeted gene modification can cause the destruction of the target genome locus.The destruction can include the change of control element (for example, promoter or enhancer), missense mutation, nonsense mutation, frameshift mutation, truncation mutation, null mutation, or the insertion or deletion of a small number of nucleotides (for example, causing frameshift mutation), and can result in the inactivation (i.e., loss of function) or loss of allele.For example, targeted modification can include the destruction of the start codon of HSD17B13 gene, so that the start codon is no longer functional.

[0259] In certain examples, the targeted modification may include a deletion between the first guide RNA target sequence and the second guide RNA target sequence or the Cas cleavage site. When an exogenous donor sequence (e.g., a repair template or a targeting vector) is used, the modification may include a deletion between the first guide RNA target sequence and the second guide RNA target sequence or the Cas cleavage site, as well as an insertion of a nucleic acid insert between the 5' and 3' target sequences.

[0260] Alternatively, when an exogenous donor sequence is used alone or in combination with a nuclease agent, the modification can include a deletion between the 5' and 3' target sequences in the pair of the first and second homologous chromosomes and an insertion of a nucleic acid insert between the 5' and 3' target sequences, resulting in a homozygous modified genome. Alternatively, when the exogenous donor sequence includes 5' and 3' homologous arms without a nucleic acid insert, the modification can include a deletion between the 5' and 3' target sequences.

[0261] The deletion between the first and second guide RNA target sequences or the deletion between the 5' and 3' target sequences can be a precise deletion, where the deleted nucleic acid consists only of the nucleic acid sequence between the first and second nuclease cleavage sites, or only of the nucleic acid sequence between the 5' and 3' target sequences, and therefore there is no additional deletion or insertion at the modified genomic target locus. The deletion between the first and second guide RNA target sequences can also be an imperfect deletion that extends beyond the first and second nuclease cleavage sites, consistent with imperfect repair by non-homologous end joining (NHEJ), resulting in an additional deletion and / or insertion at the modified genomic locus. For example, the deletion can extend beyond the first and second Cas protein cleavage sites by about 1 bp, about 2 bp, about 3 bp, about 4 bp, about 5 bp, about 10 bp, about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, or more. Similarly, the modified genomic locus may include additional insertions, such as about 1 bp, about 2 bp, about 3 bp, about 4 bp, about 5 bp, about 10 bp, about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, or more, consistent with non-strict repair by NHEJ.

[0262] Targeted gene modification can be, for example, biallelic modification or monoallelic modification. Biallelic modification includes the event that the same modification is made to the same locus on corresponding homologous chromosomes (for example, in diploid cells), or the event that different modifications are made to the same locus on corresponding homologous chromosomes. In some methods, targeted gene modification is monoallelic modification. Monoallelic modification includes the event that only one allele is modified (i.e., modification to only one of the two homologous chromosomes of the HSD17B13 gene). Homologous chromosomes include chromosomes (for example, chromosomes that pair during meiosis) that have the same gene at the same locus, but may have different alleles. The term allele includes any one or more alternative forms of a gene sequence. In diploid cells or organisms, the two alleles of a given sequence generally occupy corresponding loci on a pair of homologous chromosomes.

[0263] A single allele mutation can result in a cell that is heterozygous for the targeted HSD17B13 modification. Heterozygous includes situations in which only one allele of the HSD17B13 gene (i.e., corresponding alleles on both homologous chromosomes) has the targeted modification.

[0264] Biallelic modification can result in homozygotes for targeted modification. Homozygotes include the situation where both alleles of HSD17B13 gene (i.e., corresponding alleles on both homologous chromosomes) have targeted modification. Alternatively, biallelic modification can result in heterozygotes (e.g., hemizygotes) for targeted modification compounds. Compound heterozygotes include the situation where both alleles (i.e., alleles on both homologous chromosomes) of the target locus are modified, but modified differently (e.g., targeted modification in one allele and inactivation or destruction of the other allele). For example, in the allele without targeted modification, the double-strand break created by Cas protein may be repaired by non-homologous end joining (NHEJ)-mediated DNA repair, thereby generating mutant alleles containing insertion or deletion of nucleic acid sequence, thereby causing the destruction of the genomic locus. For example, biallelic modification can result in compound heterozygote when cell has one allele with targeted modification and another allele that cannot be expressed.Compound heterozygote includes hemizygote.Hemizygote includes the situation where only one allele of target locus exists (i.e., one allele of two homologous chromosomes).For example, biallelic modification can result in hemizygote for targeted modification when targeted modification occurs in one allele and the corresponding other allele is lost or deleted.

[0265] (2) Identification of cells with targeted gene modifications The method disclosed herein can further comprise the step of identifying the cell that has modified HSD17B13 gene.Various methods can be used to identify the cell that has targeted gene modification, such as deletion or insertion.Such method can comprise the step of identifying a cell that has targeted gene modification in HSD17B13 gene.Can carry out screening to identify such cell that has modified genomic locus.

[0266] The screening step may include a quantitative assay to evaluate the allelic modification (MOA) of the parent chromosome (e.g., loss of allele (LOA) and / or gain of allele (GOA) assay). For example, the quantitative assay can be performed by quantitative PCR, such as real-time PCR (qPCR). Real-time PCR can utilize a first primer set that recognizes the target genomic locus and a second primer set that recognizes a non-targeted reference locus. The primer set can include a fluorescent probe that recognizes the amplified sequence. The loss of allele (LOA) assay reverses the traditional screening logic and quantifies the number of copies of the native locus where the mutation is directed. In correctly targeted cell clones, the LOA assay detects one of the two native alleles (genes not on the X or Y chromosome), while the other allele is disrupted by the targeted modification. The same principle can be applied in reverse as a gain of allele (GOA) assay to quantify the copy number of the inserted targeting vector. For example, the combined use of GOA and LOA assays reveals that correctly targeted heterozygous clones have lost one copy of the native target gene and gained one copy of a drug resistance gene or other inserted marker.

[0267] For example, quantitative polymerase chain reaction (qPCR) can be used as a method for allele quantification, but any method that can reliably distinguish between zero, one, and two copies of a target gene or between zero, one, and two copies of a nucleic acid insert can be used to develop an MOA assay. For example, TAQMAN® can be used to quantify the number of copies of a DNA template in a genomic DNA sample, particularly by comparing it with a reference gene (see, for example, US6,596,541, the entire contents of which are incorporated herein by reference for all purposes). The reference gene is quantified as the target gene(s) or locus(s) within the same genomic DNA. Therefore, two TAQMAN® amplifications (each using its respective probe) are performed. One TAQMAN® probe determines the "Ct" (threshold cycle number) of the reference gene, and the other probe determines the Ct of the region of the targeted gene(s) or locus(s) that is displaced by successful targeting (i.e., LOA assay). Ct is a quantity that reflects the amount of starting DNA for each TAQMAN® probe; i.e., less abundant sequences require more cycles of PCR to reach the threshold cycle number. Reducing the number of copies of the template sequence for a TAQMAN® reaction by half results in an increase of approximately 1 Ct unit. TAQMAN® reactions in cells in which one allele of the target gene(s) or locus(s) has been replaced by homologous recombination will result in an increase of 1 Ct for the target TAQMAN® reaction, with no increase in Ct for the reference gene compared to DNA from non-targeted cells. For GOA assays, a separate TAQMAN® probe can be used to determine the Ct of the nucleic acid insert that replaced the targeted gene(s) or locus(s) by successful targeting.

[0268] Other examples of suitable quantitative assays include fluorescence-mediated in situ hybridization (FISH), comparative genomic hybridization, isothermal DNA amplification, quantitative hybridization to immobilized probe(s), INVADER® probe, TAQMAN® molecular beacon probe, or ECLIPSE™ probe technology (see, e.g., US2005 / 0144655, which is incorporated herein by reference in its entirety for all purposes). Conventional assays for screening for targeted modifications, such as long-range PCR, Southern blotting, or Sanger sequencing, can also be used. Such assays are generally used to obtain evidence of the linkage between the inserted targeting vector and the targeted genomic locus. For example, for long-range PCR assays, one primer can recognize a sequence within the inserted DNA, and the other recognizes a target genomic locus sequence beyond the end of the homologous arm of the targeting vector.

[0269] Next-generation sequencing (NGS) can also be used for screening.Next-generation sequencing can also be called "NGS" or "massively parallel sequencing" or "high-throughput sequencing".In the method disclosed herein, it is not necessary to use selection marker to screen targeted cells.For example, it can rely on MOA and NGS assay described herein without using selection cassette.

[0270] B. Methods of Altering Expression of HSD17B13 Nucleic Acids Various methods are provided for modifying the expression of nucleic acid encoding HSD17B13 protein.In some methods, as described in more detail elsewhere herein, expression is modified by cleavage using a nuclease agent to cause destruction of the nucleic acid encoding HSD17B13 protein.In some methods, expression is modified by using a DNA binding protein fused or linked with a transcription activation domain or a transcription repression domain.In some methods, expression is modified by using RNA interference compositions such as antisense RNA, shRNA, or siRNA.

[0271] In one example, the expression of the HSD17B13 gene or the nucleic acid encoding the HSD17B13 protein can be modified by contacting a cell or the genome in the cell with a nuclease agent or a nucleic acid encoding the HSD17B13 protein, which induces one or more nicks or double-strand breaks in the target sequence at the target genomic locus in the HSD17B13 gene. Such cleavage can result in the disruption of the expression of the HSD17B13 gene or the nucleic acid encoding the HSD17B13 protein. For example, the nuclease target sequence can include or be close to the start codon of the HSD17B13 gene. For example, the target sequence can be within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides from the start codon, and the cleavage by the nuclease agent can destroy the start codon. As another example, two or more nuclease agents can be used, each targeting a nuclease target sequence that includes or is adjacent to a start codon. As another example, two nuclease agents can be used, one targeting a nuclease target sequence that includes or is adjacent to a start codon and one targeting a nuclease target sequence that includes or is adjacent to a stop codon, and cleavage by the nuclease agents can result in deletion of the coding region between the two nuclease target sequences. As yet another example, three or more nuclease agents can be used, with one or more (e.g., two) targeting a nuclease target sequence containing or adjacent to the start codon and one or more (e.g., two) targeting a nuclease target sequence containing or adjacent to the stop codon, and cleavage by the nuclease agents can result in the deletion of the coding region between the nuclease target sequence containing or adjacent to the start codon and the nuclease target sequence containing or adjacent to the stop codon. Other examples of modifications of the HSD17B13 gene or nucleic acid encoding the HSD17B13 protein are disclosed elsewhere herein.

[0272] In another example, the expression of the HSD17B13 gene or the nucleic acid encoding the HSD17B13 protein can be modified by contacting a cell or a genome within the cell with a DNA-binding protein that binds to a target genomic locus within the HSD17B13 gene. The DNA-binding protein can be, for example, a nuclease-inactive Cas protein fused with a transcriptional activation domain or a transcriptional repressor domain. Other examples of DNA-binding proteins include zinc finger proteins fused with a transcriptional activation domain or a transcriptional repressor domain, or transcription activator-like effector (TALE) proteins fused with a transcriptional activation domain or a transcriptional repressor domain. Examples of such proteins are disclosed elsewhere herein. For example, in some methods, a transcriptional repressor can be used to reduce the expression of a wild-type HSD17B13 gene or an HSD17B13 gene that is not an rs72613567 variant (e.g., to reduce the expression of HSD17B13 transcripts or isoform A). Similarly, in some methods, a transcriptional activator can be used to increase expression of the HSD17B13 gene rs72613567 variant gene (e.g., to increase expression of the HSD17B13 transcript or isoform D).

[0273] The target sequence of DNA binding protein (for example, guide RNA target sequence) can be located anywhere in the HSD17B13 gene or the nucleic acid encoding HSD17B13 protein, which is suitable for modifying expression.For example, the target sequence can be located in or near a control element such as an enhancer or promoter.For example, the target sequence can include or be close to the start codon of HSD17B13 gene.For example, the target sequence can be within about 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1,000 nucleotides from the start codon.

[0274] In another example, antisense molecules can be used to alter the expression of HSD17B13 gene or the nucleic acid encoding HSD17B13 protein.Examples of antisense molecules include antisense RNA, small interfering RNA (siRNA), and short hairpin RNA (shRNA).Such antisense RNA, siRNA, or shRNA can be designed to target any region of mRNA.For example, antisense RNA, siRNA, or shRNA can be designed to target a region unique to one or more of the HSD17B13 transcripts disclosed herein, or a region common to one or more of the HSD17B13 transcripts disclosed herein.Examples of nucleic acids that hybridize with cDNA and variant HSD17B13 transcripts are disclosed in more detail elsewhere herein.For example, antisense RNA, siRNA, or shRNA can hybridize with the sequence in SEQ ID NO: 4 (HSD17B13 transcript A). Optionally, the antisense RNA, siRNA, or shRNA may reduce expression of HSD17B13 transcript A in the cell. Optionally, the antisense RNA, siRNA, or shRNA hybridizes to a sequence present in SEQ ID NO: 4 (HSD17B13 transcript A) that is not present in SEQ ID NO: 7 (HSD17B13 transcript D). Optionally, the antisense RNA, siRNA, or shRNA hybridizes to a sequence within exon 7 of SEQ ID NO: 4 (HSD17B13 transcript A) or a sequence spanning the boundary between exons 6 and 7.

[0275] As another example, the antisense RNA, siRNA, or shRNA may hybridize to a sequence within SEQ ID NO: 7 (HSD17B13 transcript D). Optionally, the antisense RNA, siRNA, or shRNA may reduce expression of HSD17B13 transcript D in the cell. Optionally, the antisense RNA, siRNA, or shRNA hybridizes to a sequence present in SEQ ID NO: 7 (HSD17B13 transcript D) that is not present in SEQ ID NO: 4 (HSD17B13 transcript A). Optionally, the antisense RNA, siRNA, or shRNA hybridizes to a sequence within exon 7 of SEQ ID NO: 7 (HSD17B13 transcript D) or a sequence spanning the boundary between exons 6 and 7.

[0276] C. Introduction of Nucleic Acids and Proteins into Cells The nucleic acids and proteins disclosed herein can be introduced into cells by any means. "Introducing" includes bringing a nucleic acid or protein into a cell so that the sequence enters the cell. Introduction can be achieved by any means, and one or more of the components (e.g., two of the components, or all of the components) can be introduced into a cell simultaneously or sequentially in any combination. For example, the exogenous donor sequence can be introduced before or after the introduction of the nuclease agent (e.g., the exogenous donor sequence can be administered about 1, 2, 3, 4, 8, 12, 24, 36, 48, or 72 hours before or about 1, 2, 3, 4, 8, 12, 24, 36, 48, or 72 hours after the introduction of the nuclease agent). See, for example, US2015 / 0240263 and US2015 / 0110762, each of which is incorporated by reference in its entirety for all purposes. Contacting the genome of a cell with a nuclease agent or exogenous donor sequence can include introducing one or more nuclease agents or nucleic acids encoding nuclease agents (e.g., one or more Cas proteins or nucleic acids encoding one or more Cas proteins, and one or more guide RNAs or nucleic acids encoding one or more guide RNAs (i.e., one or more CRISPR RNAs and one or more tracrRNAs)) and / or one or more exogenous donor sequences into the cell. Contacting the genome of a cell (i.e., contacting the cell) can include introducing only one of the above components, one or more of the components, or all of the components into the cell.

[0277] The nuclease agent can be introduced into the cell in the form of a protein or in the form of a nucleic acid encoding the nuclease agent, such as RNA (e.g., messenger RNA (mRNA)) or DNA. If introduced in the form of DNA, the DNA can be operably linked to a promoter active in the cell. Such DNA can be in one or more expression constructs.

[0278] For example, a Cas protein can be introduced into a cell in the form of a protein, such as a Cas protein complexed with a gRNA, or in the form of a nucleic acid encoding the Cas protein, such as RNA (e.g., messenger RNA (mRNA)) or DNA. A guide RNA can be introduced into a cell in the form of RNA or in the form of a DNA encoding the guide RNA. When introduced in the form of DNA, the DNA encoding the Cas protein and / or guide RNA can be operably linked to a promoter active in the cell. Such DNAs can be in one or more expression constructs. For example, such an expression construct can be a component of a single nucleic acid molecule. Alternatively, they can be separated in any combination in two or more nucleic acid molecules (i.e., DNA encoding one or more CRISPR RNAs, DNA encoding one or more tracrRNAs, and DNA encoding the Cas protein can be components of separate nucleic acid molecules).

[0279] In some methods, DNA encoding nuclease agents (e.g., Cas proteins and guide RNAs) and / or DNA encoding exogenous donor sequences can be introduced into cells by DNA minicircles.See, for example, WO2014 / 182700, the entire contents of which are incorporated herein by reference for all purposes.DNA minicircles are supercoiled DNA molecules that can be used for non-viral gene transfer, and do not have replication origins or antibiotic selection markers.Therefore, DNA minicircles are generally smaller in size than plasmid vectors.These DNAs lack bacterial DNA, and therefore lack the unmethylated CpG motifs found in bacterial DNA.

[0280] The method presented herein does not depend on the specific method for introducing nucleic acid or protein into cell, as long as nucleic acid or protein enters at least one cell.The method for introducing nucleic acid and protein into various cell types is known, and includes, for example, stable transfection method, transient transfection method and virus-mediated method.

[0281] Transfection protocols and protocols for introducing nucleic acids or proteins into cells can vary. Non-limiting transfection methods include liposomes; nanoparticles; calcium phosphate (Graham et al. (1973) Virology 52(2):456-67; Bacchetti et al. (1977) Proc. Natl. Acad. Sci. USA 74(1):101-104). (4): 1590-4, and Kriegler, M (1991), Transfer and Expression: A Laboratory Manual. New York: WH Freeman and Company, pp. 96-97); dendrimers; or chemical-based transfection methods using cationic polymers such as DEAE-dextran or polyethyleneimine. Non-chemical methods include electroporation, sonoporation, and optical transfection. Particle-based transfection includes the use of a gene gun or magnet-assisted transfection (Bertram (2006) Current Pharmaceutical Biotechnology, Vol. 7, pp. 277-28). Viral methods The method can also be used for transfection.

[0282] The introduction of nucleic acids or proteins into cells can also be mediated by electroporation, intracytoplasmic injection, viral infection, adenovirus, adeno-associated virus, lentivirus, retrovirus, transfection, lipid-mediated transfection, or nucleofection. Nucleofection is an improved electroporation technique that allows nucleic acid substrates to be delivered not only into the cytoplasm but also through the nuclear membrane and into the nucleus. Furthermore, the use of nucleofection in the methods disclosed herein generally requires far fewer cells than conventional electroporation (e.g., only about 2 million cells compared to 7 million cells in conventional electroporation). In one example, nucleofection is performed using the LONZA® NUCLEOFECTOR™ system.

[0283] Introduction of nucleic acids or proteins into cells can also be achieved by microinjection. Microinjection of mRNA is preferably into the cytoplasm (e.g., to deliver mRNA directly to the translation machinery), while microinjection of DNA encoding a protein or DNA encoding a Cas protein is preferably into the nucleus. Alternatively, microinjection can be performed by both intranuclear and intracytoplasmic injection: first, a needle can be introduced into the nucleus, a first amount can be injected, and a second amount can be injected into the cytoplasm while the needle is removed from the cell. When injecting a nuclease agent protein into the cytoplasm, the protein preferably contains a nuclear localization signal to ensure delivery to the nucleus / pronucleus. Methods for performing microinjection are well known. See, e.g., Nagy et al. (Nagy A, Gertsenstein M, Wintersten K, Behringer R. 2003. Manipulating the Mouse Embryo. Cold Spring Harbor, New York: Cold Spring Harbor Laboratory Press); Meyer et al. (2010) Proc. Natl. Acad. Sci. USA 107:15022-15026 and Meyer et al. (2012) Proc. Natl. Acad. Sci. USA 109:9354-93 Please refer to page 59.

[0284] Other methods for introducing nucleic acids or proteins into cells can include, for example, vector delivery, particle-mediated delivery, exosome-mediated delivery, lipid nanoparticle-mediated delivery, cell-penetrating peptide-mediated delivery, or implantable device-mediated delivery. Methods for administering nucleic acids or proteins to a subject to modify cells in vivo are disclosed elsewhere herein.

[0285] Introduction of nucleic acids and proteins into cells can also be achieved by hydrodynamic delivery (HDD). Hydrodynamic delivery has emerged as a nearly complete method for intracellular DNA delivery in vivo. For gene delivery to parenchymal cells, only the essential DNA sequence needs to be injected through selected blood vessels, eliminating the safety concerns associated with current viral and synthetic vectors. When injected into the bloodstream, DNA can reach cells in various tissues accessible to the blood. Hydrodynamic delivery uses the force generated by rapidly injecting a large volume of solution into the incompressible circulating blood to overcome the physical barriers of the endothelium and cell membranes that prevent large, membrane-impermeable compounds from entering parenchymal cells. In addition to DNA delivery, this method is useful for the efficient intracellular delivery of RNA, proteins, and other small compounds in vivo. See, for example, Bonamassa et al. (2011) Pharm. Res., 28(4):694-701, incorporated herein by reference in its entirety for all purposes.

[0286] Other methods for introducing nucleic acids or proteins into cells include, for example, vector delivery, particle-mediated delivery, exosome-mediated delivery, lipid nanoparticle-mediated delivery, cell-penetrating peptide-mediated delivery, or implantable device-mediated delivery. For example, nucleic acids or proteins can be introduced into cells in carriers such as poly(lactic acid) (PLA) microspheres, poly(D,L-lactic acid-co-glycolic acid) (PLGA) microspheres, liposomes, micelles, reverse micelles, lipid cocrystals, or lipid microtubules.

[0287] The nucleic acid or protein can be introduced into the cell once or multiple times over a period of time, for example, at least 2 times over a period of time, at least 3 times over a period of time, at least 4 times over a period of time, at least 5 times over a period of time, at least 6 times over a period of time, at least 7 times over a period of time, at least 8 times over a period of time, at least 9 times over a period of time, at least 10 times over a period of time, at least 11 times, at least 12 times over a period of time, at least 13 times over a period of time, at least 14 times over a period of time, at least 15 times over a period of time, at least 16 times over a period of time, at least 17 times over a period of time, at least 18 times over a period of time, at least 19 times over a period of time, or at least 20 times over a period of time.

[0288] In some cases, cells used in the methods and compositions have a DNA construct stably integrated into their genome. In such cases, contacting can include providing a cell with a construct that is already stably integrated into its genome. For example, a cell used in the methods disclosed herein can have an existing Cas-encoding gene stably integrated into its genome (i.e., a Cas-ready cell). "Stably integrated" or "stably introduced" or "stably integrated" includes introducing a polynucleotide into a cell such that the nucleotide sequence is integrated into the genome of the cell and can be inherited by its progeny. Any protocol can be used to stably integrate the DNA construct or various components of the targeted genome integration system.

[0289] D. Nuclease Agents and DNA-Binding Proteins Any nuclease agent that induces a nick or double-strand break within a desired target sequence or any DNA-binding protein that binds to a desired target sequence can be used in the methods and compositions disclosed herein. Naturally occurring or native nuclease agents can be used as long as the nuclease agent induces a nick or double-strand break within the desired target sequence. Similarly, naturally occurring or native DNA-binding proteins can be used as long as the DNA-binding protein binds to the desired target sequence. Alternatively, modified or engineered nuclease agents or DNA-binding proteins can be used. An "engineered nuclease agent or DNA-binding protein" includes a nuclease agent or DNA-binding protein that has been engineered (modified or derived) from its native form to specifically recognize a desired target sequence. Thus, an engineered nuclease agent or DNA-binding protein can be derived from a native, naturally occurring nuclease agent or DNA-binding protein, or it can be artificially created or synthesized. The engineered nuclease agent or DNA-binding protein can recognize a target sequence, for example, the target sequence is not a sequence recognized by the native (unengineered or unmodified) nuclease agent or DNA-binding protein. The modification of the nuclease agent or DNA-binding protein can be as small as one amino acid for a protein-cleaving agent or one nucleotide for a nucleic acid-cleaving agent. Creating a nick or double-strand break in a target sequence or other DNA may be referred to herein as "cutting" or "cleaving" the target sequence or other DNA.

[0290] Active variants and fragments of nuclease agents or DNA-binding proteins (i.e., engine...

Claims

1. A combination for modifying the HSD17B13 gene in a cell, comprising: (a) a Cas9 protein, or a nucleic acid encoding the Cas9 protein; (b) a guide RNA or a DNA encoding the guide RNA, wherein the guide RNA comprises a CRISPR RNA (crRNA) portion and a trans-activating CRISPR RNA (tracrRNA) portion, and the guide RNA forms a complex with the Cas9 protein and targets a guide RNA target sequence within the coding region of the HSD17B13 gene; Including, the Cas9 protein cleaves the guide RNA target sequence to produce a targeted gene modification in the HSD17B13 gene, and the targeted gene modification is generated by repair of the cleaved guide RNA target sequence by non-homologous end joining; the combination, when introduced into the cell, results in a loss of function of the HSD17B13 gene; and the cells are hepatocytes; Combination.

2. (a) the guide RNA target sequence comprises any one of SEQ ID NOs: 20-239 and 259-268; (b) the guide RNA comprises a DNA-targeting segment comprising any one of SEQ ID NOs: 1423-1652; or (c) the guide RNA comprises any one of SEQ ID NOs: 500 to 1419; The combination according to claim 1.

3. the cells are human cells, (a) the guide RNA target sequence comprises any one of SEQ ID NOs: 20-239; (b) the DNA-targeting segment comprises any one of SEQ ID NOs: 1423-1642; or (c) the guide RNA comprises any one of SEQ ID NOs: 500-719, 730-949, 960-1179, and 1190-1409; The combination according to claim 2.

4. the cell is a mouse cell, (a) the guide RNA target sequence comprises any one of SEQ ID NOs: 259-268; (b) the DNA-targeting segment comprises any one of SEQ ID NOs: 1643-1652; or (c) the guide RNA comprises any one of SEQ ID NOs: 720-729, 950-959, 1180-1189, and 1410-1419; The combination according to claim 2.

5. 5. The combination of claim 1, further comprising one or more additional guide RNAs or one or more DNAs encoding said one or more additional guide RNAs, wherein said one or more additional guide RNAs target one or more additional guide RNA target sequences in the HSD17B13 gene; the one or more additional guide RNAs form one or more complexes with the Cas9 protein and target the one or more additional guide RNA target sequences; Combination.

6. 5. The combination of claim 1, wherein the combination comprises the nucleic acid encoding the Cas9 protein.

7. 7. The combination of claim 6, wherein the nucleic acid encoding the Cas9 protein comprises DNA.

8. 7. The combination of claim 6, wherein the nucleic acid encoding the Cas9 protein comprises RNA.

9. The combination according to any one of claims 1 to 4, wherein the combination comprises the guide RNA in the form of RNA.

10. The combination according to any one of claims 1 to 4, wherein the combination comprises the DNA encoding the guide RNA.

11. The combination of any one of claims 1 to 4, wherein the Cas9 protein or the nucleic acid encoding the Cas9 protein, or the guide RNA or the DNA encoding the guide RNA, is in a lipid nanoparticle.

12. The combination of any one of claims 1 to 4, wherein the nucleic acid encoding the Cas9 protein or the DNA encoding the guide RNA is in an adeno-associated virus vector.

13. 5. The combination of claim 1, wherein the guide RNA is a single-molecule guide RNA in which the crRNA portion is linked to the tracrRNA portion.

14. The combination of claim 13, wherein the guide RNA comprises a sequence set forth in SEQ ID NO: 1420, 256, 257 or 258.

15. The combination of claim 1 , wherein the crRNA portion and the tracrRNA portion are separate RNA molecules.

16. The combination of claim 15, wherein the crRNA portion comprises the sequence set forth in SEQ ID NO: 1421 or the tracrRNA portion comprises the sequence set forth in SEQ ID NO: 1422.

17. 5. The combination of claim 1, wherein the guide RNA comprises a modification that confers modified or controlled stability.

18. The combination according to any one of claims 1 to 4, wherein the cell is in vivo.

19. The combination according to any one of claims 1 to 4, wherein the cells are human hepatocytes or mouse hepatocytes.

20. 20. The combination of claim 19, wherein the cells are human hepatocytes.

21. The combination of claim 20, wherein the human hepatocytes are in vivo.

22. 22. The combination according to claim 21, wherein the combination is introduced into the liver in vivo.

23. 22. The combination of claim 21, wherein the human hepatocytes are in a subject having or susceptible to developing chronic liver disease.

24. 24. The combination of claim 23, wherein the chronic liver disease is fatty liver disease, non-alcoholic fatty liver disease, alcoholic fatty liver disease, cirrhosis, or hepatocellular carcinoma.

Citation Information

Patent Citations

  • AM2012