Dnase 1-like protein 3 engineered to improve expression and use in therapy
By introducing linkers of N-terminal and C-terminal extensions into DNA enzymes, the production titer and homogeneity of DNA enzymes in the microbial expression system are solved, and efficient, non-immunogenic DNA enzyme variant production is achieved.
Patent Information
- Application Number
- CN202380088707.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-14
- Filing Date
- 2023-11-14
- Publication Date
- 2025-08-05
AI Technical Summary
The prior art is difficult to achieve efficient production of DNA enzymes, structural homogeneity and reduced immunogenicity in microbial expression systems, and there are problems of unnecessary secondary modification and invalid processing of signal sequences.
By introducing the N-terminal extension and the C-terminal extension into the DNA enzyme, a linker containing at least four amino acids, respectively, improves processing and expression titers, and controls post-translational modifications, ensuring complete processing and non-immunogenicity of the protein in the host cell.
It improves the expression titer and structural homogeneity of DNA enzymes, reduces immunogenicity, and achieves efficient production of non-immunogenic DNA enzyme variants in host cells.
Smart Images

Figure CN120435552A_ABST
Abstract
Description
[0001] priority
[0002] This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 425,106, filed on November 14, 2022, the contents of which are hereby incorporated by reference in their entirety. Technical Field
[0003] The present disclosure provides, in part, DNA enzymes engineered for improved production, structural homogeneity, and / or reduced immunogenicity for use in therapy.
[0004] Instructions for electronically submitted XML files
[0005] This application contains a sequence listing. The sequence listing has been submitted electronically via EFS-Web as an XML file titled "NTR-015PC_119604-5015_sequence_listing.xml". The sequence listing is 85,383 bytes in size and was created on November 13, 2023. The sequence listing is hereby incorporated by reference in its entirety. Background Art
[0006] Enzyme replacement therapy is a promising approach for the treatment of a variety of diseases and conditions. For example, inflammatory diseases such as systemic lupus erythematosus (SLE) can be treated with DNA enzymes. Lauková et al., Deoxyribonucleases and Their Applications in Biomedicine , Biomolecular cules 2020;10(7):1036. Enzymes are produced in microbial hosts such as Pichia pastoris or mammalian cells. However, not every protein of interest is produced or secreted to sufficiently high titers in the desired expression system such as Pichia pastoris, and there are other difficulties. For example, several studies have revealed unwanted secondary modifications, structural heterogeneity, and inefficient processing of N-terminal signal sequences. See, for example, Arbeitman et al., Structural and functional comparison of SARS-CoV-2- spike receptor binding domain produced in Pichia pastoris and mammalian cells .Sci Rep 2020;10:21779(2020);Reverter et al., Overexpression of Human Procarboxypeptidase A2 in Pichia pastoris and Detailed Chara cterization of Its Activation Pathway , Protein Chemistry and Structu re|1998;273(6):3535-3541; Katla et al., Novel glycosylated human interferon alpha 2b expressed in glycoengineered Pichia pastoris and its biological activity:N-linked glycoengineering approach, Enzyme and Microbial Technology 2019;128:49-58. Therefore, there is a need for improved expression systems to allow the production of recombinant therapeutic proteins with improved production titers, structural homogeneity and / or reduced immunogenicity. Summary of the Invention
[0007] In various aspects, the present disclosure provides recombinant proteins produced by secretion from an expression system with complete processing of the N-terminal signal sequence. In various aspects, the present disclosure provides methods for preparing such proteins, including DNA enzymes, and the use of the proteins in therapy.
[0008] In various aspects and embodiments, the present disclosure is based in part on the discovery that adding a linker of four or more amino acids between a secretion signal and a protein of interest can improve its processing and expression titer from a microbial expression system, as well as control post-translational modification profiles. The present disclosure is also based in part on the discovery that modifications to the C-terminus can improve the product homogeneity characteristics of proteins produced in an expression system.
[0009] In some aspects, the present disclosure provides a variant of a DNA enzyme 1-like protein 3 (D1L3 variant), the variant comprising an N-terminal extension. In some embodiments, the N-terminal extension comprises at least four amino acids. In some embodiments, the N-terminal extension does not undergo post-translational modification of the host cell and / or is non-immunogenic after administration to a human or animal subject. In some embodiments, the D1L3 variant produced by the host cell comprises a D1L3 enzyme lacking a signal peptide and comprises an amino acid sequence having at least 80% sequence identity with amino acids 21 to 282 of SEQ ID NO: 4 (isoform 1) or amino acids 21 to 252 of SEQ ID NO: 5 (isoform 2).
[0010] In some embodiments, the D1L3 variant is produced in a host cell (such as, but not limited to, Pichia pastoris) by cleavage of the N-terminal signal peptide. Any signal peptide that allows secretion of the D1L3 in the desired expression system can be used. In some embodiments, the signal peptide is the α mating factor (α MF) prepro secretory leader sequence from Saccharomyces cerevisiae (SEQ ID NO: 38).
[0011] In various embodiments, the length of the N-terminal extension is in the range of 4 to about 18 amino acids. In embodiments, the first N-terminal amino acid is not a Met residue and is not a Gly residue (which may undergo post-translational modification as disclosed herein). In some embodiments, the last amino acid residue of the N-terminal extension is not a Ser residue (which may increase the immunogenicity risk of D1L3 as disclosed herein). In some embodiments, the last amino acid residue of the N-terminal extension is not a polar or charged amino acid residue, such as an amino acid residue selected from Ser, Thr, Gln, Asn, Glu, Asp, Arg, His, and Lys. In some embodiments, the last amino acid residue of the N-terminal extension is a Gly residue. In some embodiments, the first amino acid residue of the N-terminal extension is not a Gly residue, and the last amino acid residue of the N-terminal extension is a Gly residue. In some embodiments, the N-terminal extension is mainly a Gly residue (i.e., more than 50% is a Gly residue). In some embodiments, the linker comprises at least one cysteine residue. In some embodiments, the N-terminal extension comprises or consists of an amino acid sequence of SGGGG (SEQ ID NO: 60). In some embodiments, the N-terminal extension comprises or consists of an amino acid sequence of CGGGG (SEQ ID NO: 74). In some embodiments, the N-terminal extension comprises or consists of an amino acid sequence of SGGSGGSGG (SEQ ID NO: 61). In some embodiments, the N-terminal extension comprises or consists of an amino acid sequence of SGGSGGSGGSGGSGG (SEQ ID NO: 62).
[0012] In some embodiments, the N-terminal extension improves the expression titer of the D1L3 variant compared to the same D1L3 sequence lacking the N-terminal extension, and with respect to the selected signal sequence and expression host. In some embodiments, the signal peptide is completely removed from the D1L3 variant after secretion from the host.
[0013] In some embodiments, the D1L3 variant further comprises a C-terminal extension that reduces heterogeneity. In some embodiments, the length of the C-terminal extension is at least two, or at least three, or at least four, or at least five amino acids. In some embodiments, the extension is a hydrophilic sequence of 2, 3 or 4 amino acids. In some embodiments, the C-terminal amino acid is lysine (Lys) or arginine (Arg). In some embodiments, the C-terminal extension comprises or consists of the amino acid sequence SSR, which in some embodiments can be used for D1L3 variants that completely or partially lack a C-terminal basic domain (as further described herein). In some embodiments, the D1L3 variant comprises a C-terminal basic domain having the amino acid sequence SSRAFTNSKKSVTLRKKTKSKRS (SEQ ID NO: 75) or a modified form thereof. When these D1L3 enzymes are expressed in Pichia pastoris, the resulting polypeptide will be cut after the SSR sequence at the beginning of the basic domain sequence. In still other embodiments, the D1L3 is encoded and expressed with a 20 amino acid deletion of the basic domain, thereby having the sequence SSR at the C-terminus.
[0014] In some embodiments, the D1L3 variant comprises a deletion of at least three, or at least five, or at least eight, or at least ten, or at least twelve, or at least fifteen, or at least eighteen, or at least twenty, or all 23 amino acids of the C-terminal basic domain (BD) defined by amino acids 283 to 305 of SEQ ID NO: 4. In some embodiments, the D1L3 variant comprises a deletion of 20 amino acids of the C-terminal basic domain (BD) defined by amino acids 283 to 305 of SEQ ID NO: 4, such that the D1L3 variant comprises a C-terminal extension of the sequence SSR.
[0015] In various embodiments, the D1L3 variant comprises an amino acid sequence having at least 70% sequence identity to D1L3 isoform 1 (SEQ ID NO: 4) or D1L3 isoform 2 (SEQ ID NO: 5) lacking the BD (i.e., having at least 80% sequence identity to amino acids 21 to 282 of SEQ ID NO: 4 or amino acids 21 to 252 of SEQ ID NO: 5), and wherein the D1L3 variant has a deletion of one or more amino acids in the BD. In embodiments, the D1L3 variant comprises an N-terminal extension of any embodiment disclosed herein, and / or the D1L3 variant comprises a C-terminal extension of any embodiment disclosed herein.
[0016] In some embodiments, the D1L3 variant has a substitution of the amino acid corresponding to C48 of SEQ ID NO: 4. These variants can remove the unpaired cysteine, thereby improving recombinant production and stability of the enzyme.
[0017] In some embodiments, the D1L3 variant includes one or more unpaired cysteines that are configured for and / or capable of dimerization, for example, relative to the amino acid sequence of SEQ ID NO:4 at an unpaired cysteine at position 48 (numbered in the absence of a signal peptide). In some embodiments, the disclosure provides a D1L3 dimer according to the disclosure, which is dimerized via a disulfide bridge at C48. In some embodiments, the D1L3 variant includes the introduction of non-natural cysteines that can promote dimerization, including the introduction of non-natural cysteines at the N-terminus. In such embodiments, the disclosure provides a D1L3 dimer that can be used in therapy according to the disclosure. In some embodiments, the D1L3 variant includes a substitution of cysteine (C48) at position 48 with respect to SEQ ID NO:4, to further control the dimerization position. For example, mutations can be selected from C48A, C48G, and C48S. In some embodiments, substitution of C48 (e.g., substitution to C48A or C48G) increases the enzymatic capacity of the D1L3 enzyme (e.g., chromatin degradation). In some embodiments, the D1L3 variant has a C-terminal extension comprising or consisting of an amino acid sequence SSR and a mutation with respect to C48A of SEQ ID NO: 4 (numbered without the signal peptide).
[0018] Exemplary D1L3 variants according to the present disclosure include SEQ ID NO:63, SEQ ID NO:66, SEQ ID NO:70, SEQ ID NO:71, SEQ ID NO:72, SEQ ID NO:73, or SEQ ID NO:77, wherein the signal peptide is absent and is fully processed by the host expression system.
[0019] In some embodiments, the D1L3 variant comprises a fusion or conjugation with a half-life extending portion. In some embodiments, the polymer is polyethylene glycol (PEG). In some embodiments, the PEG is conjugated to the N-terminus. In some embodiments, the PEG polymer is connected to the D1L3 variant molecule via the N-termini of two D1L3 variant molecules. In some embodiments, the PEG is conjugated to the C-terminus. In some embodiments, the PEG polymer is connected to the D1L3 variant molecule via the C-termini of two D1L3 variant molecules.
[0020] In some embodiments, the D1L3 variant comprises a fusion or conjugation with a half-life extending moiety. In some embodiments, the half-life extending moiety is a fusion partner. In some embodiments, the fusion partner is selected from albumin, transferrin, Fc, or elastin-like protein, XTEN sequence, or variants thereof. In some embodiments, the fusion partner is albumin. In some embodiments, the fusion partner is fused to a mature D1L3 enzyme lacking a signal peptide at the N-terminus (and via a linker sequence). In some embodiments, the D1L3 variant comprises the amino acid sequence of SEQ ID NO: 76 or SEQ ID NO: 78, or a variant thereof having at least about 99% sequence identity thereto. In some embodiments, the N-terminal extension of any embodiment disclosed herein connects the fusion partner to the N-terminus of the mature D1L3 enzyme. In some embodiments, the linker is about 5 to about 50 amino acids, or about 10 to about 35 amino acids, or about 15 to about 35 amino acids in length. In some embodiments, the linker comprises the amino acid sequence S(GGS)4GSS (SEQ ID NO: 23), S(GGS)9GSS (SEQ ID NO: 24), or (GGS)9GS (SEQ ID NO: 25).
[0021] In some aspects, the present disclosure provides a method for expressing the D1L3 variant of any embodiment disclosed herein. In some embodiments, the method comprises introducing a gene construct encoding the D1L3 variant of any embodiment disclosed herein and comprising a signal peptide into a yeast cell, and reclaiming the D1L3 variant. In some embodiments, the yeast cell is Pichia pastoris. In some embodiments, the signal peptide is an α mating factor (α MF) pre-pro-secretion leader sequence (SEQ ID NO: 38) from Saccharomyces cerevisiae. In some embodiments, the fusion protein is synthesized with an N-terminal signal peptide. The signal peptide can be completely removed during secretion from the host cell. Regarding expression in Pichia pastoris, an α mating factor (α MF) pre-pro-secretion leader sequence (SEQ ID NO: 38) from Saccharomyces cerevisiae can be used for expression. These elements are cut during expression and are not present in the D1L3 variant enzyme product.
[0022] In some aspects, the present disclosure provides an isolated polynucleotide encoding a D1L3 variant of any embodiment disclosed herein, and provides the advantages of in vitro or in vivo expression. In some aspects, the present disclosure provides a polynucleotide, the polynucleotide is mRNA or modified mRNA (mmRNA). In some aspects, the present disclosure provides a polynucleotide, the polynucleotide is DNA.
[0023] In some aspects, the present disclosure provides pharmaceutical compositions comprising the D1L3 enzyme described herein, or a polynucleotide encoding the D1L3 enzyme, or a transfection or expression vector comprising the polynucleotide, or a cell comprising the polynucleotide or vector, and a pharmaceutically acceptable carrier.
[0024] In some aspects, the present disclosure provides a pharmaceutical composition comprising an effective amount of a D1L3 variant according to any embodiment disclosed herein, a D1L3 variant produced according to a method according to any embodiment disclosed herein, a polynucleotide according to any embodiment disclosed herein, a vector according to any embodiment disclosed herein, or a host cell according to any embodiment disclosed herein, and a pharmaceutically acceptable carrier.
[0025] In some embodiments, the composition comprises a D1L3 variant of any embodiment disclosed herein and a pharmaceutically acceptable carrier for parenteral administration. In some embodiments, the pharmaceutical composition is formulated for topical, parenteral, or pulmonary administration. In some embodiments, the pharmaceutical composition is formulated for intradermal, intramuscular, intraperitoneal, intraarticular, intravenous, subcutaneous, intraarterial, ocular, oral, sublingual, pulmonary, or transdermal administration.
[0026] In embodiments, the method for the recombinant production of D1L3 enzyme variants adopts a non-mammalian expression system, for example, a eukaryotic non-mammalian expression system, such as Pichia pastoris. In embodiments, the Pichia pastoris encodes a DNA enzyme with a natural signal peptide, thereby allowing secretion from a host cell. In embodiments, the expression system is a mammalian cell expression system, such as Chinese hamster ovary (CHO) cells. In embodiments, the method for the recombinant production of D1L3 enzyme variants further comprises separating and / or purifying the D1L3 enzyme, and modifying the separated and / or purified D1L3 enzyme. In embodiments, the modification includes conjugating the separated and / or purified D1L3 enzyme to a polymer (for example, not limited to PEG). In embodiments, the polymer is added to a specific site using desired conjugation chemistry (for example, not limited to maleimide chemistry).
[0027] In other aspects, the present disclosure provides a method for treating a subject in need of extracellular chromatin degradation, extracellular trap (ET) degradation, and / or neutrophil extracellular trap (NET) degradation. The method comprises administering a therapeutically effective amount of a D1L3 enzyme or composition described herein.
[0028] In some aspects, the present disclosure provides an expression construct for improving the processing of a polypeptide precursor in a host, the expression construct comprising a signal peptide fused to the polypeptide via a joint. In some embodiments, the joint has a length of at least three amino acids. In some embodiments, the signal peptide is completely removed from the polypeptide. In some embodiments, the signal peptide is not removed from the polypeptide. In some embodiments, in the absence of the joint, the signal peptide is incompletely processed in the host. In some embodiments, the polypeptide is or includes an enzyme, a cytokine, a cytokine agonist, a cytokine antagonist, a hormone, a hormone agonist, a hormone antagonist, an antibody or its antigen-binding fragment, an antibody-like molecule or its antigen-binding fragment, an antigen, a component of a vaccine, a fusion protein, or a combination thereof.
[0029] Other aspects and embodiments of the present disclosure will become apparent from the following detailed description and working examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1A and Figure 1B Expression of D1L3 in Pichia pastoris using either the native secretion signal or the alpha mating factor (αMF) from Saccharomyces cerevisiae is described. Figure 1A The N terminus of D1L3 is shown, guided by the aMF secretion leader sequence from Saccharomyces cerevisiae. Figure 1B It was shown that secretion of the signal from αMF results in glycosylation and non-processing of the signal.
[0031] Figure 2 Structural model showing the N-terminal native secretion signal from alpha mating factor connected to the D1L3 enzyme via a linker. Without being bound by theory, it is believed that the flexible linker is positioned at the cleavage site of alpha mating factor to enable efficient processing.
[0032] Figure 3 Shown are mass spectrometry analyses of D1L3 enzyme variants comprising the BDD_D1L3 enzyme (S283_S305del) produced by a construct comprising an alpha mating factor + GGGGS linker (SEQ ID NO: 58). These data indicate that the secretion signal is properly removed, but the protein contains a post-translational modification (probably myristoylation) at the N-terminal glycine residue.
[0033] Figure 4 Shown are the results of a titration experiment performed to compare the chromatin degradation activity of the D1L3 variant of SEQ ID NO: 63 with that of DNase 1 (D1, SEQ ID NO: 1).
[0034] Figure 5AStructural heterogeneity of D1L3 variants of SEQ ID NO: 63 as demonstrated by mass spectrometry. Figure 5B The structural homogeneity of the D1L3 variants of SEQ ID NO: 66 as demonstrated by mass spectrometry is illustrated.
[0035] Figure 6 Show the western blot of the culture supernatant of the host expressing four different DNA enzyme 1L3 variants using anti-DNA enzyme 1L3 antibodies. In all variants, C48 is mutated to alanine or serine. Sample 1 (SEQ ID NO: 70) produced by the construct comprising α mating factor+SGGGG joint (SEQ ID NO: 60) is as follows D1L3 variant, and the D1L3 variant has C48A substitution and a C-terminal extension with sequence SSR. Samples 2-4 (respectively SEQ ID NO: 71 to 73) produced by the construct comprising α mating factor+CGGGG joint (SEQ ID NO: 74) are as follows D1L3 variant, and the D1L3 variant has a C-terminal extension with sequence SSR or a wild-type C-terminal basic domain SSRAFTNSKKSVTLRKKTKSKRS (SEQ ID NO: 75). Dimers were detected in DNA enzyme 1L3 variants containing N-terminal cysteine.
[0036] Figure 7 Western blot analysis of two DNase 1L3 variants characterized by either a wild-type C-terminal amino acid sequence (i.e., SSRAFTNSKKSVTLRKKTKSKRS (SEQ ID NO: 75)) or a modified C-terminal amino acid sequence (i.e., SSR) is shown. The variants were expressed in Pichia pastoris using the alpha mating factor as a signal sequence in conjunction with an N-terminal SGGGG linker (SEQ ID NO: 60). Western blot analysis of the supernatants did not detect the theoretical mass difference of 2.3 kDa between the two variants.
[0037] Figure 8 Figure 3 shows a high molecular weight (HMW) chromatin degradation assay comparing wild-type D1L3 and C48 variants. Wild-type D1L3 contains an unpaired cysteine at position 48 (e.g., C48). HMW chromatin (i.e., purified nuclei from HEK293 cells) incubated with equal amounts of D1L3 variants was used to characterize the effect of the amino acid substitution at C48 on enzyme activity. After incubation, DNA was isolated and its degradation was visualized by agarose gel electrophoresis (AGE). Mutations from C48 to C48A or C48G were associated with increased enzyme activity.
[0038] Figure 9Western blot analysis of Pichia pastoris supernatants is shown to assess dimerization in DNase 1L3 variants. Variants with unpaired C48 showed dimerization (signal at approximately 60 kDa), but variants with mutated C48 did not. DETAILED DESCRIPTION
[0039] In various aspects, the present disclosure provides recombinant proteins (e.g., recombinant proteins for human or animal therapy, including but not limited to DNA enzymes such as D1L3), produced by secretion from an expression system with a fully processed N-terminal signal sequence. In various aspects, the present disclosure provides methods for preparing such proteins and methods for their use in therapy.
[0040] In various aspects and embodiments, the present disclosure is based in part on the discovery that adding a linker of four or more amino acids (e.g., at least 5 amino acids) between a secretion signal and a protein of interest can improve its processing and expression titer from a microbial expression system, as well as control post-translational modification profiles. The present disclosure is also based in part on the discovery that modifications to the C-terminus can improve the product homogeneity characteristics of proteins produced in an expression system.
[0041] In some aspects, the present disclosure provides a variant of a DNA enzyme 1-like protein 3 (D1L3 variant), the variant comprising an N-terminal extension. In some embodiments, the N-terminal extension comprises at least four amino acids. As used herein, the term "N-terminal extension" refers to an amino acid sequence that is not a secretion signal and is therefore not removed / processed when secreted from the host. In some embodiments, the N-terminal extension does not undergo post-translational modification of the host and / or is non-immunogenic after administration to a human or animal subject. In an exemplary embodiment, the N-terminal extension is mainly Gly residues (i.e., more than 50% are Gly residues) and may have one or more Ser or Cys residues.
[0042] In some embodiments, the D1L3 variant produced by the host cell comprises a D1L3 enzyme lacking a signal peptide and comprises an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% sequence identity to amino acids 21 to 282 of SEQ ID NO: 4 (isoform 1) or amino acids 21 to 252 of SEQ ID NO: 5 (isoform 2). When referring to sequence identity to SEQ ID NO: 4 or 5, and unless otherwise stated, the sequence refers to the mature enzyme lacking the signal peptide. Furthermore, unless otherwise stated, amino acid positions are numbered relative to the native N-terminus of the enzyme in the absence of a signal peptide. Thus, for example, when referring to sequence identity to the enzyme of SEQ ID NO: 4 (human D1L3, isoform 1), the percent identity is to the mature enzyme having M21 at the N-terminus.
[0043] In some embodiments, the D1L3 variant is produced in a host cell (such as, but not limited to, Pichia pastoris) by cleavage of the N-terminal signal peptide. Any signal peptide that allows secretion of D1L3 in the desired expression system can be used. In some embodiments, the signal peptide is a signal peptide of a naturally secreted protein. In some embodiments, the signal peptide is a chimeric or synthetic signal peptide that enables protein secretion. In some embodiments, the signal peptide is a prokaryotic signal peptide. In some embodiments, the signal peptide is a microbial signal peptide. In some embodiments, the signal peptide is a eukaryotic signal peptide. In some embodiments, the signal peptide is a yeast signal peptide. In some embodiments, the signal peptide is a mammalian signal peptide. Signal peptides suitable for secretion are disclosed in the following documents: U.S. Patent Nos. 5,580,758; 6,107,057; 7,741,075; 10,435,694; 11,306,127; 11,370,815; U.S. Patent Application Publication Nos. 2007 / 0117186, 2010 / 0055125, and 2016 / 0168198, the disclosures of each of which are hereby incorporated by reference.
[0044] In some embodiments, the signal peptide is selected from E. coli OmpA signal peptide (SEQ ID NO: 56), E. coli DsbA signal peptide (SEQ ID NO: 67), E. coli ST-II signal peptide (SEQ ID NO: 68), E. coli FimD signal peptide (SEQ ID NO: 55), Salmonella enterica DsbA signal peptide (SEQ ID NO: 51), synthetic Bordetella pertussis signal peptide (SEQ ID NO: 57) and synthetic signal peptide sequences (e.g., SEQ ID NOs: 49, 50, 52, 53 and 54).
[0045] In some embodiments, the signal peptide is selected from the group consisting of DNase 1 L3 signal peptide (SEQ ID NO: 37), alpha mating factor (SEQ ID NO: 38), alpha mating factor presequence (SEQ ID NO: 39), human serum albumin signal peptide (SEQ ID NO: 40), bovine DNase 1 signal peptide (SEQ ID NO: 41), bovine DNase 1 signal peptide + Kex2 site (SEQ ID NO: 42), alpha amylase signal peptide (SEQ ID NO: 43), glucoamylase signal peptide (SEQ ID NO: 44), inulinase signal peptide (SEQ ID NO: 45), invertase signal peptide (SEQ ID NO: 46), killer protein signal peptide (SEQ ID NO: 47) and lysozyme signal peptide (SEQ ID NO: 48).
[0046] In some embodiments, the signal peptide is the alpha mating factor (aMF) prepro secretory leader from Saccharomyces cerevisiae (SEQ ID NO: 38).
[0047] In some embodiments, the length of the N-terminal extension is at least four, or at least five, or at least six, or at least seven, or at least eight, or at least nine, or at least ten, or at least 11, or at least 12, or at least 13, or at least 14, or at least 15 amino acids. In various embodiments, the length of the N-terminal extension is in the range of 4 to about 18 amino acids, or 4 to about 12 amino acids, or 4 to about 9 amino acids, or 4 to about 7 amino acids. For example, the length of the N-terminal extension can be 4, 5, 6, 7, 8, or 9 amino acids. In some embodiments, the N-terminal extension is a flexible or rigid sequence. In some embodiments, the first N-terminal amino acid is not a Met residue. In some embodiments, the first amino acid residue of the N-terminal extension is not a Gly residue (which can undergo post-translational modification as disclosed herein). In some embodiments, the last amino acid residue of the N-terminal extension is not a Ser residue (which can increase the risk of immunogenicity as disclosed herein). In some embodiments, the last amino acid residue of the N-terminal extension is not a polar or charged amino acid residue, such as an amino acid residue selected from Ser, Thr, Gln, Asn, Glu, Asp, Arg, His, and Lys. In some embodiments, the last amino acid residue of the N-terminal extension is an amino acid selected from Gly, Ala, and Val. In some embodiments, the last amino acid residue of the N-terminal extension is a Gly residue. In some embodiments, the first amino acid residue of the N-terminal extension is not a Gly residue, and the last amino acid residue of the N-terminal extension is a Gly residue. In some embodiments, the N-terminal extension is primarily Gly residues (i.e., more than 50% are Gly residues), essentially consists of Ser and Gly residues, or consists of Ser and Gly residues. In some embodiments, the linker comprises at least one cysteine residue. For example, a cysteine can be placed at the N-terminus to provide site-specific chemical conjugation (e.g., with polyethylene glycol) or disulfide dimerization. In some embodiments, the N-terminal extension comprises or consists of the amino acid sequence of SGGGG (SEQ ID NO: 60). In some embodiments, the N-terminal extension comprises or consists of the amino acid sequence CGGGG (SEQ ID NO: 74). In some embodiments, the N-terminal extension comprises or consists of the amino acid sequence SGGSGGSGG (SEQ ID NO: 61). In some embodiments, the N-terminal extension comprises or consists of the amino acid sequence SGGSGGSGGSGGSGGSGG (SEQ ID NO: 62). In some embodiments, the N-terminal extension comprises a protease cleavage site. In some embodiments, the N-terminal extension is cleavable by a coagulation pathway protease. In some embodiments, the protease is thrombin, or Factor XII, or a neutrophil protease.In some embodiments, the protease is thrombin.In some embodiments, the protease cleavage site comprises the amino acid sequence LVPRG (SEQ ID NO: 64), such as an N-terminal extension represented by the sequence SGGGGLVPRGSGGGG (SEQ ID NO: 65).
[0048] In some embodiments, the N-terminal extension does not include a consensus sequence for myristoylation. In some embodiments, the N-terminal extension does not include a consensus sequence for one or more of the following: protein acetylation, propionylation, methylation, myristoylation, palmitoylation, ubiquitination, and a protease cleavage site to avoid unwanted post-translational modifications. In alternative embodiments, the N-terminal extension includes a consensus sequence for one or more of the following: protein acetylation, propionylation, methylation, myristoylation, palmitoylation, ubiquitination, and a protease cleavage site, wherein these modifications are desired. In some embodiments, the N-terminal extension includes an amino acid with a chemical group suitable for chemical conjugation, the chemical group being selected from a thiol group, an amino group, an amide group, and a carboxyl group.
[0049] In some embodiments, the N-terminal extension is non-immunogenic. In some embodiments, the conjugate of the N-terminal extension with a sequence derived from a desired protein (e.g., but not limited to, D1L3) is non-immunogenic, as determined by a computer immunogenicity prediction algorithm (e.g., by Lonza Group AG). In silico and in vitro immunogenicity platforms). In some embodiments, the last amino acid residue of the N-terminal extension is not a Ser residue. In some embodiments, the last amino acid residue of the N-terminal extension is not a polar or charged amino acid residue selected from Ser, Thr, Gln, Asn, Glu, Asp, Arg, His, and Lys. In some embodiments, the last amino acid residue of the N-terminal extension is an amino acid selected from Gly, Ala, and Val. In some embodiments, the last amino acid residue of the N-terminal extension is a Gly residue.
[0050] In some embodiments, the N-terminal extension improves the expression titer of the D1L3 variant by at least 25%, or at least 50%, or at least 100%, or at least 150%, or at least 200%, or at least 250%, or at least 300%, or at least 350%, or at least 400%, or at least 500%, or more, compared to the same D1L3 sequence lacking the N-terminal extension, and with respect to the selected signal sequence and expression host. In some embodiments, the signal peptide is completely removed from the D1L3 variant after secretion from the host.
[0051] In some embodiments, the D1L3 variant further comprises a C-terminal extension that reduces heterogeneity. In some embodiments, the length of the C-terminal extension is at least two, or at least three, or at least four, or at least five amino acids. In some embodiments, the extension is a hydrophilic sequence of 2, 3, or 4 amino acids. In some embodiments, the C-terminal amino acid is lysine (Lys) or arginine (Arg). In some embodiments, the C-terminal extension comprises or consists of the amino acid sequence SSR, which in some embodiments can be used for the D1L3 variants that completely or partially lack the C-terminal basic domain (as further described herein). In some embodiments, the D1L3 variant comprises a C-terminal basic domain with the amino acid sequence SSRAFTNSKKSVTLRKKTKSKRS (SEQ ID NO: 75) or a modified form thereof with one to five amino acid modifications, the amino acid modifications being independently selected from amino acid substitutions, deletions, and insertions. When these D1L3 enzymes are expressed in Pichia pastoris, the resulting polypeptide will be cut after the SSR sequence at the beginning of the basic domain sequence.
[0052] In some aspects, the present disclosure provides a variant of DNA enzyme 1-like protein 3 (D1L3 variant), the variant comprising a C-terminal extension as described above. In some embodiments, the C-terminal extension reduces the heterogeneity observed for D1L3 variants lacking a basic domain. In some embodiments, the heterogeneity is caused by post-translational modification. In some aspects, the present disclosure provides a D1L3 variant, the D1L3 variant comprising the N-terminal extension already described and a C-terminal extension that reduces heterogeneity. In some embodiments, the C-terminal extension comprises or consists of the amino acid sequence SSR.
[0053] In some embodiments, the C-terminal extension comprises an amino acid having a chemical group suitable for chemical conjugation, the chemical group being selected from a thiol group, an amino group, an amide group, and a carboxyl group. In some embodiments, the C-terminal extension does not comprise a consensus sequence for one or more of the following: protein acetylation, propionylation, methylation, myristoylation, palmitoylation, ubiquitination, and a protease cleavage site to avoid undesirable post-translational modifications. In alternative embodiments, the C-terminal extension comprises a consensus sequence for one or more of the following: protein acetylation, propionylation, methylation, myristoylation, palmitoylation, ubiquitination, and a protease cleavage site, wherein post-translational modifications are desired. In some embodiments, the C-terminal extension is non-immunogenic. In some embodiments, the conjugate of the C-terminal extension with a sequence derived from D1L3 is non-immunogenic.
[0054] In some embodiments, the D1L3 variant comprises a deletion of at least three, or at least five, or at least eight, or at least ten, or at least twelve, or at least fifteen, or at least eighteen, or at least twenty, or at least twenty-one, or all 23 amino acids of the C-terminal basic domain (BD) defined by amino acids 283 to 305 of SEQ ID NO: 4. In some embodiments, the D1L3 variant comprises a deletion of at least three amino acids of the C-terminal basic domain (BD) defined by amino acids 283 to 305 of SEQ ID NO: 4, wherein the D1L3 variant comprises a C-terminal extension of any embodiment disclosed herein (for example, without limitation, having the amino acid sequence SSR). In such embodiments, the C-terminus having the SSR extension is equivalent to a deletion of 20 amino acids of the basic domain.
[0055] D1L3 is characterized by a 23-amino acid C-terminal tail defined by amino acids 283 to 305 of SEQ ID NO: 4, which contains 9 basic amino acids and is therefore referred to as a basic domain (BD). BD is unique to D1L3 and is absent in DNase 1 (D1). BD contains a nuclear localization signal (NLS), which is thought to target the enzyme to the nucleus during apoptosis. Although it has been widely considered that the BD is also critical for D1L3's activity in the extracellular space, the absence of the C-terminal tail actually stimulates D1L3's chromatinase activity. In some embodiments, the D1L3 variant comprises a BD having the sequence SSRAFTNSKKSVTLRKKTKSKRS (SEQ ID NO: 75) or a truncated derivative thereof having at least three, or at least five, or at least eight, or at least ten, or at least twelve, or at least fifteen, or at least eighteen amino acids, or about 20 amino acids.
[0056] In some embodiments, the D1L3 variant comprises mutations (e.g., substitutions, insertions, or deletions) of three sets of paired basic amino acids of unknown function, namely K291 / K292, R297 / K298 / K299, and K303 / R304 corresponding to SEQ ID NO: 4. In some embodiments, the D1L3 variant comprises a BD having the sequence SSRAFTNSKKSVTLRKKTKSKRS (SEQ ID NO: 75), or a derivative thereof having one to five amino acid modifications that alter the paired basic amino acids. In some embodiments, the D1L3 variant comprises a truncation of the BD that deletes at least one paired basic amino acid. In some embodiments, the D1L3 variant exhibits reduced proteolytic cleavage at these paired basic amino acids. In some embodiments, the D1L3 variant comprises an N-terminal extension of any embodiment disclosed herein (e.g., without limitation, comprising the amino acid sequence SGGGG (SEQ ID NO: 60) or CGGGG (SEQ ID NO: 74)), and / or the D1L3 variant comprises a C-terminal extension of any embodiment disclosed herein (e.g., without limitation, having the amino acid sequence SSR). In some embodiments, the C-terminal extension is added in place of the C-terminal basic domain.
[0057] In various embodiments, the D1L3 variant comprises an amino acid sequence having at least 70% sequence identity to D1L3 isoform 1 (SEQ ID NO: 4) or D1L3 isoform 2 (SEQ ID NO: 5) but lacking the BD (i.e., having at least 80% sequence identity to amino acids 21 to 282 of SEQ ID NO: 4 or amino acids 21 to 252 of SEQ ID NO: 5), and wherein the D1L3 variant has a deletion of one or more amino acids in the BD. Amino acid deletions in the basic domain of D1L3 improve its chromatin degradation activity. In addition, deletions of the BD that increase by 23 amino acids are directly associated with increased chromatin degradation activity. In some embodiments, D1L3 variants having a deletion of at least one amino acid, or at least 3, or at least 5, or at least 8, or at least 9, or at least 13, or at least 14, or at least 15, or at least 18, or about 20 C-terminal amino acids of the D1L3 basic domain have increased ability to degrade mononucleosomes. In some embodiments, the D1L3 variant comprises an N-terminal extension of any embodiment disclosed herein (for example, without limitation, comprising or consisting of the amino acid sequence SGGGG (SEQ ID NO: 60) or CGGGG (SEQ ID NO: 74)), and / or the D1L3 variant comprises a C-terminal extension of any embodiment disclosed herein (for example, without limitation, having the amino acid sequence SSR).
[0058] In various embodiments, the amino acid deletion from the BD is located at the C-terminus of the BD. For example, a D1L3 variant may have a deletion of at least five C-terminal amino acids of the BD. In some embodiments, a D1L3 variant has a deletion of at least eight C-terminal amino acids of the BD. In some embodiments, a D1L3 variant has a truncation of at least eight C-terminal amino acids of the BD. In some embodiments, a D1L3 variant has a deletion of at least ten C-terminal amino acids of the BD. In some embodiments, a D1L3 variant has a truncation of at least ten C-terminal amino acids of the BD. In some embodiments, a D1L3 variant has a deletion of at least twelve C-terminal amino acids of the BD. In some embodiments, a D1L3 variant has a truncation of at least twelve C-terminal amino acids of the BD. In some embodiments, a D1L3 variant has a deletion of at least fifteen C-terminal amino acids of the BD. In some embodiments, a D1L3 variant has a truncation of at least fifteen C-terminal amino acids of the BD. In some embodiments, a D1L3 variant has a deletion of at least eighteen C-terminal amino acids of the BD. In some embodiments, the D1L3 variant has a truncation of at least eighteen C-terminal amino acids of BD. In some embodiments, the D1L3 variant has a deletion of at least twenty C-terminal amino acids of BD. In some embodiments, the D1L3 variant has a truncation of at least (or about) twenty C-terminal amino acids of BD. In some embodiments, the D1L3 variant has a deletion of at least twenty-three C-terminal amino acids of BD. In some embodiments, the D1L3 variant has a truncation of at least twenty-three C-terminal amino acids of BD. In some embodiments, the D1L3 variant has a deletion of at least the C-terminal serine of BD. In some embodiments, the D1L3 variant has a truncation of the C-terminal serine of BD. In some embodiments, the D1L3 variant comprises an N-terminal extension of any embodiment disclosed herein (for example, but not limited to, comprising the amino acid sequence SGGGG (SEQ ID NO: 60) or CGGGG (SEQ ID NO: 74)), and / or the D1L3 variant comprises a C-terminal extension of any embodiment disclosed herein (for example, but not limited to, having the amino acid sequence SSR).
[0059] Alternatively, the deletion of the BD (e.g., three to 23 amino acids) can be located anywhere in the BD and is not necessarily from the C-terminus of the BD. For example, in various embodiments, the D1L3 variants are missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23 amino acids from the BD. In some embodiments, the D1L3 variants have a deletion of at least 1, or at least 3, or at least 5, or at least 8, or at least 12, or at least 15, or at least 18, or at least 21 amino acids from the BD. In some embodiments, the D1L3 variants have a truncation of at least 1, or at least 3, or at least 5, or at least 8, or at least 12, or at least 15, or at least 18, or at least 20 amino acids, or at least 21 amino acids from the BD. These deletions can be independently selected from the N-terminal side of the BD, the C-terminal side of the BD, and the interior of the BD. In some embodiments, one or more amino acid deletions are located within the NLS. In some embodiments, the deleted amino acid is the C-terminal serine of the BD. In some embodiments, the deletion is sufficient to remove all paired basic amino acids in the BD from the enzyme. In some embodiments, the D1L3 variant comprises an N-terminal extension of any embodiment disclosed herein (for example, not limited to, comprising the amino acid sequence SGGGG (SEQ ID NO: 60) or CGGGG (SEQ ID NO: 74)), and / or the D1L3 variant comprises a C-terminal extension of any embodiment disclosed herein (for example, not limited to, having the amino acid sequence SSR). In some embodiments, the C-terminal extension of the sequence SSR is equivalent to a truncation of 20 amino acids of the BD.
[0060] In addition to the deletion of one or more amino acids, BD may also include amino acid substitutions that may further affect chromatin degradation activity. For example, in addition to the deletion of at least three amino acids, D1L3 variants may have 1 to 20 amino acid substitutions of BD amino acids. In some embodiments, BD contains at least three amino acids, or at least five amino acids, or at least 10 amino acid substitutions. In some embodiments, at least two amino acid substitutions are located in the NLS of BD. In some embodiments, one or more paired basic amino acids in BD are substituted to prevent cutting. In such embodiments, more homogeneous enzymes can be expressed and secreted, for example, for recombinase production. In some embodiments, D1L3 variants include N-terminal extensions of any embodiment disclosed herein (for example, not limited to, including amino acid sequence SGGGG (SEQ ID NO: 60) or CGGGG (SEQ ID NO: 74)), and / or D1L3 variants include C-terminal extensions of any embodiment disclosed herein (for example, not limited to, with amino acid sequence SSR).
[0061] In some embodiments, in addition to the deletion of BD, the D1L3 variant has a deletion of one or more additional amino acids from the C-terminus. For example, in addition to the deletion of BD, the D1L3 variant can have a deletion of another one to fifty amino acids, or one to twenty amino acids, or one to ten amino acids, or one to five amino acids from the C-terminal amino acid of SEQ ID NO: 4 or SEQ ID NO: 5. In some embodiments, the D1L3 variant comprises an N-terminal extension of any embodiment disclosed herein (for example, not limited to, comprising the amino acid sequence SGGGG (SEQ ID NO: 60) or CGGGG (SEQ ID NO: 74)), and / or the D1L3 variant comprises a C-terminal extension of any embodiment disclosed herein (for example, not limited to, having the amino acid sequence SSR).
[0062] In some embodiments, after partial or complete deletion of the BD as described, 1 to 10 amino acids or 1 to 5 amino acids can be added to the C-terminus, and the amino acids do not affect the chromatin degradation activity. In some embodiments, the addition of the N-terminal extension of any embodiment disclosed herein (for example, without limitation, comprising the amino acid sequence SGGGG (SEQ ID NO: 60) or CGGGG (SEQ ID NO: 74)) and / or the C-terminal extension of any embodiment disclosed herein (for example, without limitation, having the amino acid sequence SSR) does not affect the chromatin degradation activity.
[0063] In some embodiments, the D1L3 variant comprises a substitution of C68 and / or C194 with respect to SEQ ID NO: 4 (C48 and C174 when numbered without a signal peptide). In some embodiments, the mutation is selected from C68S, C68A, C68G, C194S, C194A, and C194G with respect to SEQ ID NO: 4. In some embodiments, the D1L3 variant comprises one or more mutations that result in resistance to proteolysis by one or more of plasmin, thrombin, trypsin, and proteases produced by mammalian and non-mammalian cell lines. In some embodiments, the D1L3 variant has one or more mutations selected from amino acid residues K180, K200, K259, and R285 with respect to SEQ ID NO: 4. In some embodiments, the D1L3 variant has one or more mutations in amino acid residues selected from R22, R29, K45, K47, K74, R81, R92, K107, K176, R212, R226, R227, K250, K259, and K262 with respect to SEQ ID NO:4.
[0064] In some embodiments, D1L3 variants include one or more unpaired cysteines that are configured for and / or capable of dimerization, for example, relative to the amino acid sequence of SEQ ID NO:4 at an unpaired cysteine at position 48 (numbered in the absence of a signal peptide). In some embodiments, the disclosure provides a D1L3 dimer according to the disclosure that dimerizes via a disulfide bridge at C48. In some embodiments, D1L3 variants include the introduction of non-natural cysteines that can promote dimerization, including the introduction of non-natural cysteines at the N-terminus. In such embodiments, the disclosure provides a D1L3 dimer that can be used in therapy according to the disclosure. In some embodiments, D1L3 variants include substitutions with respect to SEQ ID NO:4 at position 48 cysteine (C48) to further control the dimerization position. For example, mutations can be selected from C48A, C48G, and C48S. In some embodiments, substitution of C48 (e.g., substitution to C48A or C48G) increases the enzymatic capacity of the D1L3 enzyme (e.g., chromatin degradation). In some embodiments, the D1L3 variant has a C-terminal extension comprising or consisting of the amino acid sequence SSR and a mutation with respect to C48A of SEQ ID NO: 4 (numbered without the signal peptide).
[0065] In some embodiments, the D1L3 variant comprises: (i) a D1L3 enzyme lacking a signal peptide and comprising an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% sequence identity to amino acids 21 to 282 of SEQ ID NO: 4 or amino acids 21 to 252 of SEQ ID NO: 5; (ii) a deletion of at least three amino acids of the C-terminal basic domain (BD) defined by amino acids 283 to 305 of SEQ ID NO: 4; and (iii) an N-terminal extension of at least four amino acids as described herein. In some embodiments, the D1L3 variant further comprises a substitution with respect to C48 of SEQ ID NO: 4. In some embodiments, the mutation is selected from C48S or C48A with respect to SEQ ID NO: 4 (numbering in the absence of a signal peptide). In some embodiments, the D1L3 variant further comprises modifications (e.g., without limitation, PEGylation) with respect to C48 and / or C174 of SEQ ID NO:4.
[0066] In some embodiments, the D1L3 variant comprises: (i) a D1L3 enzyme lacking a signal peptide and comprising an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% sequence identity to amino acids 21 to 282 of SEQ ID NO: 4 or amino acids 21 to 252 of SEQ ID NO: 5; (ii) a deletion of at least three amino acids of the C-terminal basic domain (BD) defined by amino acids 283 to 305 of SEQ ID NO: 4; and (iii) a C-terminal extension comprising the amino acid sequence SSR. In some embodiments, the D1L3 variant further comprises a substitution with respect to C48 and / or C174 of SEQ ID NO: 4 (numbering in the absence of a signal peptide). In some embodiments, the mutation is selected from C48S or C48A with respect to SEQ ID NO: 4. In some embodiments, the D1L3 variant further comprises modifications (e.g., without limitation, PEGylation) with respect to C48 and / or C174 of SEQ ID NO:4.
[0067] In some embodiments, the D1L3 variant comprises: (i) a mature D1L3 enzyme lacking a signal peptide and comprising an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% sequence identity to amino acids 21 to 282 of SEQ ID NO: 4 or amino acids 21 to 252 of SEQ ID NO: 5; (ii) a deletion of at least three amino acids of the C-terminal basic domain (BD) defined by amino acids 283 to 305 of SEQ ID NO: 4; (iii) an N-terminal extension as described herein; and (iv) a C-terminal extension comprising the amino acid sequence SSR. In some embodiments, the D1L3 variant further comprises a substitution with respect to C48 and / or C174 of SEQ ID NO: 4 (numbering in the absence of a signal peptide). In some embodiments, the mutation is selected from C48S and C48A with respect to SEQ ID NO: 4. In some embodiments, the D1L3 variant further comprises modifications (e.g., without limitation, PEGylation) with respect to C48 and / or C174 of SEQ ID NO:4.
[0068] In some embodiments, the D1L3 variant has one or more mutations of a serine residue, such as those selected from S91C, S131C, and S253C with respect to SEQ ID NO: 4 (numbering includes the signal peptide).
[0069] Exemplary D1L3 variants according to the present disclosure include SEQ ID NO: 63, SEQ ID NO: 66, SEQ ID NO: 70, SEQ ID NO: 71, SEQ ID NO: 72, SEQ ID NO: 73, or SEQ ID NO: 77, wherein the signal peptide is absent and fully processed by the host expression system. SEQ ID NO: 63 utilizes an N-terminal extension of SEQ ID NO: 60 to complete signal peptide processing and contains a complete basic domain deletion. SEQ ID NO: 66 also includes a C48S substitution and an SSR C-terminal extension (as compared to SEQ ID NO: 63). SEQ ID NO: 70 exemplifies a C48A substitution (but originally has the sequence of SEQ ID NO: 66). SEQ ID NO: 71 is similar to SEQ ID NO: 70, but has a CGGGG (SEQ ID NO: 74) linker before the signal peptide. SEQ ID NO: 72 exemplifies a C48S substitution and a CGGGG (SEQ ID NO: 74) linker. SEQ ID NO: 73 exemplifies a C48A substitution and a CGGGG (SEQ ID NO: 74) linker and an intact basic domain. When produced in Pichia pastoris, the polypeptide of SEQ ID NO: 73 will have an SSR at the C-terminus (e.g., 20 amino acids of the basic domain will be cleaved). SEQ ID NO: 77 utilizes a C48A substitution, an SGGGG (SEQ ID NO: 60) linker, and a C-terminal extension (SSR) added after an additional 12 amino acid deletion (in addition to the basic domain deletion).
[0070] In some embodiments, the D1L3 variant comprises a fusion or conjugation with a half-life extending portion. In some embodiments, the half-life extending portion is a polymer. In some embodiments, the polymer is polyethylene glycol (PEG). In some embodiments, PEG is conjugated to the N-terminus (optionally within the N-terminal extension of any embodiment disclosed herein, for example, the N-terminal extension comprises the amino acid sequence SGGGG (SEQ ID NO: 60) or CGGGG (SEQ ID NO: 74)). In some embodiments, the PEG polymer is connected to the D1L3 variant molecule via the N-terminus of two D1L3 variant molecules. In some embodiments, PEG is conjugated to the C-terminus (optionally within the C-terminal extension of any embodiment disclosed herein, for example, the C-terminal extension has the amino acid sequence SSR). In some embodiments, D1L3 comprises a basic domain, and PEG is conjugated to the basic domain. In some embodiments, the PEG polymer is connected to the D1L3 variant molecule via the C-terminus of two D1L3 variant molecules.
[0071] In some embodiments, the PEG polymer is conjugated to one or more amino acids within positions corresponding to R95 to V126 of SEQ ID NO: 4. In some embodiments, the one or more PEGylated amino acids are selected from lysine, cysteine, histidine, arginine, aspartic acid, glutamic acid, serine, threonine, and tyrosine, optionally wherein the one or more PEGylated amino acids are introduced by substitution of one or more amino acids between R95 and V126 relative to SEQ ID NO: 4.
[0072] In some embodiments, one or more amino acids are pegylated by: (a) pegylation of lysine (Lys or K) by amine conjugation; (b) pegylation of glutamine by transglutaminase (TGase)-mediated enzymatic conjugation; and / or (c) pegylation of cysteine (Cys or C) by thiol conjugation. In some embodiments, one or more pegylated amino acids are conjugated to a PEG moiety independently selected from a linear or branched PEG, the molecular weight of the PEG being independently selected and ranging from about 2kDa to about 60kDa or from about 5kDa to about 30kDa. Pegylation is disclosed in WO 2019 / 036719 and WO 2020 / 076817, both of which are hereby incorporated by reference in their entirety.
[0073] In these embodiments, the PEG moiety will provide half-life extension properties while avoiding disulfide scrambling and / or protein misfolding. In some embodiments, the PEG moiety is conjugated via maleimide chemistry, which can be performed under mild conditions. Other conjugation chemistries are known and can be used, such as vinyl sulfone, dihydropyridine, and iodoacetamide activation chemistries.
[0074] In some embodiments, the D1L3 variant comprises a fusion or conjugation to a half-life extending moiety. In some embodiments, the half-life extending moiety is a fusion partner. In some embodiments, the fusion partner is selected from albumin, transferrin, Fc, or an elastin-like protein, an XTEN sequence, or a variant thereof. In some embodiments, the fusion partner is albumin. In some embodiments, the fusion partner is human albumin comprising an amino acid sequence at least 80% identical to SEQ ID NO:26. In some embodiments, the human albumin comprises at least one of the substitutions E505Q, T527M, and K573P with respect to SEQ ID NO:26. In some embodiments, the human albumin comprises at least two of the substitutions E505Q, T527M, and K573P with respect to SEQ ID NO:26. In some embodiments, the human albumin comprises each of the substitutions E505Q, T527M, and K573P with respect to SEQ ID NO:26. In some embodiments, the fusion partner is fused at the N-terminus to a mature D1L3 enzyme lacking a signal peptide (and via a linker sequence), and the mature D1L3 enzyme comprises an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% sequence identity to amino acids 21 to 282 of SEQ ID NO: 4 or amino acids 21 to 252 of SEQ ID NO: 5, and has a deletion of at least three amino acids of the C-terminal basic domain (BD) defined by amino acids 283 to 305 of SEQ ID NO: 4. In some embodiments, the D1L3 variant comprises the amino acid sequence of SEQ ID NO: 76 or SEQ ID NO: 78, or a variant thereof having at least about 99% sequence identity thereto. In some embodiments, the N-terminal extension of any embodiment disclosed herein connects the fusion partner to the N-terminus of the mature D1L3 enzyme. In some embodiments, the fusion partner is fused to the mature D1L3 enzyme via a linker connecting the fusion partner and the mature D1L3 enzyme. In some embodiments, the N-terminal extension of any embodiment disclosed herein connects the fusion partner to the N-terminus of the mature D1L3 enzyme. In some embodiments, the linker is a flexible or rigid linker and / or comprises a protease cleavage site. In some embodiments, the linker can be cleaved by a coagulation pathway protease. In some embodiments, the protease is Factor XII or a neutrophil protease. In some embodiments, the protease is thrombin. In some embodiments, the protease cleavage site comprises the amino acid sequence LVPRG (SEQ ID NO: 64), such as a linker represented by the sequence SGGGGLVPRGSGGGG (SEQ ID NO: 65).In some embodiments, the length of the linker is about 5 to about 50 amino acids, or about 10 to about 35 amino acids, or about 15 to about 35 amino acids. In some embodiments, the linker comprises the amino acid sequence S(GGS)4GSS (SEQ ID NO: 23), S(GGS)9GSS (SEQ ID NO: 24), or (GGS)9GS (SEQ ID NO: 25). In some embodiments, the linker comprises the amino acid sequence (GGGGS)5GGGG (SEQ ID NO: 79), as shown by the fusion protein of SEQ ID NO: 76.
[0075] In some embodiments, the D1L3 variant comprises a flexible linker between the D1L3 sequence and the half-life extending moiety. The flexible linker is primarily or entirely composed of small non-polar or polar residues such as Gly, Ser, and Thr. Exemplary flexible linkers include (Gly y Ser) n S z Linkers wherein y is 1 to 10 (e.g., 1 to 5), n is 1 to about 10, and z is 0 or 1. In some embodiments, n is 3 to about 8 or 3 to about 6. In exemplary embodiments, y is 2 to 4 and n is 3 to 8. Due to their flexibility, these linkers are unstructured. More rigid linkers include polyproline or polyPro-Ala motifs and alpha helical linkers. An exemplary alpha helical linker is A(EAAAK) n A, wherein n is as defined above (e.g., 1 to 10 or 3 to 6). Typically, the linker can be primarily composed of an amino acid selected from Gly, Ser, Thr, Ala, and Pro. Exemplary linker sequences contain at least 10 amino acids and can be in the range of 10 to about 50 amino acids, or about 15 to about 40 amino acids, or about 15 to about 35 amino acids. Exemplary linker designs are provided as SEQ ID NOs: 18 to 25.
[0076] In some embodiments, the D1L3 variant with a fusion partner comprises a joint, wherein the amino acid sequence of the joint is mainly glycine and serine residues, or is essentially composed of glycine and serine residues, or is composed of glycine and serine residues. In some embodiments, the ratio of the joint to Ser and Gly is respectively about 1:1 to about 1:10, about 1:2 to about 1:6 or about 1:4. Exemplary joint sequences include S (GGS) 4GSS (SEQ ID NO:23), S (GGS) 9GSS (SEQ ID NO:24) or (GGS) 9GS (SEQ ID NO:25) or are composed of said sequence. In some embodiments, the joint has at least 10 amino acids, or at least 15 amino acids, or at least 20 amino acids, or at least 25 amino acids, or at least 30 amino acids. For example, the joint can have a length of 15 to 40 amino acids. In various embodiments, a longer joint of at least 15 amino acids can provide an improvement in the titer after expression in Pichia pastoris.
[0077] In some aspects, the disclosure provides a method for expressing the DNA enzyme 1-like protein 3 variant of any embodiment in the embodiments disclosed herein. In some embodiments, the method includes introducing a gene construct encoding the DNA enzyme 1-like protein 3 variant of any embodiment disclosed herein and comprising a signal peptide in a yeast cell, and reclaiming the DNA enzyme 1-like protein 3 variant. In some embodiments, the yeast cell is Pichia pastoris. In some embodiments, the signal peptide is the α mating factor (α MF) pre-pro-secretion leader sequence (SEQ ID NO:38) from saccharomyces cerevisiae. In some embodiments, the fusion protein is synthesized with an N-terminal signal peptide. The signal peptide can be completely removed during secretion from the host cell. Regarding expression in Pichia pastoris, the α mating factor (α MF) pre-pro-secretion leader sequence (SEQ ID NO:38) from saccharomyces cerevisiae can be used for expression. These elements are cut during expression and are not present in the D1L3 variant enzyme product.
[0078] In some aspects, the present disclosure provides an isolated polynucleotide encoding the D1L3 variant of any embodiment disclosed herein, and provides the advantages of in vitro or in vivo expression. In some aspects, the present disclosure provides a polynucleotide, which is mRNA or modified mRNA (mmRNA). In some aspects, the present disclosure provides a polynucleotide, which is DNA. The present disclosure provides a pharmaceutical composition in some aspects, which comprises a D1L3 enzyme as described herein, or a polynucleotide optionally encoding the D1L3 enzyme, or a transfection or expression vector comprising the polynucleotide, or a cell comprising the polynucleotide or vector, and a pharmaceutically acceptable carrier.
[0079] In some embodiments, the delivery of polynucleotides is used for therapy. Encoding polynucleotides can be delivered as mRNA or as a DNA construct using known procedures, for example, electroporation or cell extrusion, and / or carriers (including viral vectors). mRNA polynucleotides can include known modifications (mmRNA) to avoid activation of the innate immune system. Referring to WO 2014 / 028429, it is hereby incorporated by reference in its entirety. In some embodiments, polynucleotides are delivered to the body of the subject. In some embodiments, polynucleotides are delivered in vitro to cells, and the cells are delivered to the body of the subject. The cells can be, for example, leukocytes (for example, T cells, B cells or macrophages), endothelial cells, epithelial cells, hepatocytes, fibroblasts or stem cells (for example, hematopoietic stem cells).
[0080] In some embodiments, the polynucleotides used for therapy are modified mRNA (mmRNA). In some embodiments, mmRNA is administered to a subject in need of treatment. In some embodiments, the cells are transformed in vitro or ex vivo with modified mRNA (mmRNA), amplified before or after transfection, and used for therapy (cell therapy). In some embodiments, mmRNA can be modified uniformly along the entire length of the molecule. In alternative embodiments, mmRNA may not be modified uniformly along the entire length of the molecule. Different nucleotide modifications and / or main chain structures may be present at different positions in the nucleic acid. In some embodiments, nucleotide analogs or one or more other modifications may be located at any one or more positions of the nucleic acid so that the function of the nucleic acid is not significantly reduced. In some embodiments, mmRNA may include 5' or 3' end modifications.
[0081] In some embodiments, the mmRNA may contain at least about 5% modified nucleotides, or at least about 10% modified nucleotides, or at least about 20% modified nucleotides, or at least about 50% modified nucleotides, or at least about 80% modified nucleotides. In some embodiments, the mmRNA may contain less than about 10% modified nucleotides, or less than about 20% modified nucleotides, or less than about 50% modified nucleotides.
[0082] In some embodiments, mmRNA may include polynucleotide modifications, such as, but not limited to, nucleoside modifications. Nucleoside modifications may include, but are not limited to, pyridine-4-ketone ribonucleoside, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4-thio-pseudouridine, 2-thiouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-taurine methyluridine, 1-taurine methyl-pseudouridine, 5-taurine methyl-2-thiouridine, 1-taurine methyl-4-thiouridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thiol-1-methyl-pseudouridine , 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine, dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, 4-methoxy-2-thio-pseudouridine, 5-aza-cytidine, pseudoisocytidine, 3-methyl-cytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudoisocytidine, pyrrolo-cytidine, pyrrolo-pseudoisocytidine Cytidine, 2-thio-cytidine, 2-thio-5-methyl-cytidine, 4-thio-pseudoisocytidine, 4-thio-1-methyl-pseudoisocytidine, 4-thio-1-methyl-1-deaza-pseudoisocytidine, 1-methyl-1-deaza-pseudoisocytidine, zebulline, 5-aza-zebulline, 5-methyl-zebulline, 5-aza-2-thio-zebulline, 2-thio-zebulline, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, 4-methoxy-1-methyl-pseudoisocytidine, 2-aminopurine, 2,6-diaminopurine, 7-deaza-pseudoisocytidine Adenosine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyladenosine, N6-(cis-hydroxyisopentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6-glycylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methylthio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine, and 2-methoxy-adenine, inosine, 1-methyl-inosine, wyosine, wyobutosine, 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, and N2,N2-dimethyl-6-thio-guanosine, and combinations thereof. Suitable modifications are disclosed in US20190060458, the contents of which are hereby incorporated by reference in their entirety.
[0083] In some aspects, the present disclosure provides a vector for introducing a polynucleotide of any embodiment disclosed herein into a host cell. In some aspects, the present disclosure provides a host cell comprising a vector of any embodiment disclosed herein.
[0084] In some embodiments, the polynucleotides used for therapy are DNA molecules encoding wild-type D1L3 enzymes or any variants of D1L3 disclosed herein (i.e., gene therapy). In some embodiments, cells are transformed in vitro or ex vivo with DNA molecules encoding wild-type D1L3 enzymes or any variants of D1L3 disclosed herein, amplified, and used for therapy (i.e., cell therapy). In some embodiments, the DNA molecule is a vector. The vector typically comprises isolated nucleic acids, and it can be used to deliver the isolated nucleic acids to the interior of the cell. Various vectors are known in the art, including but not limited to linear polynucleotides, polynucleotides associated with ions or amphipathic compounds, plasmids, and viruses. In some embodiments, the vector is a viral vector. Exemplary vectors include autonomously replicating plasmids or viruses (e.g., AAV vectors). The term should also be considered to include non-plasmid and non-viral compounds that facilitate the transfer of nucleic acids into cells, such as, for example, polylysine compounds, liposomes, etc. The example of viral vectors includes but is not limited to adenoviral vectors, adeno-associated viral vectors, retroviral vectors, etc.
[0085] In some embodiments, polynucleotide or cell therapy can adopt expression vector, and the expression vector comprises the nucleic acid of coding chromatin enzyme (for example, D1L3) operably connected with the expression control region that plays a role in host cell.The expression control region can drive the expression of the coding nucleic acid that can be operably connected, so that the chromatin enzyme is produced in the human cell transformed with the expression vector.The expression control region is the regulatory polynucleotide (sometimes referred to as element herein) that affects the expression of the nucleic acid that can be operably connected, such as promoter and enhancer.The expression control region of expression vector can express the coding nucleic acid that can be operably connected in human cell.In embodiments, the expression control region gives the nucleic acid that can be operably connected with adjustable expression.The signal (sometimes referred to as stimulus) can increase or decrease the expression of the nucleic acid that can be operably connected with this expression control region.Such expression control regions that increase expression in response to a signal are generally referred to as inducible.Such expression control regions that reduce expression in response to a signal are generally referred to as repressible.In various embodiments, chromatin enzyme expression is inducible or repressible.Generally, the amount of increase or decrease given by such elements is proportional to the amount of the signal present; The greater the amount of the signal, the more increase or decrease in expression.
[0086] In some embodiments, the viral vector is an adeno-associated viral vector (AAV). In some embodiments, the ability of the AAV-based vector suitable in the present disclosure to induce an immune response in the human body is extremely limited. The AAV genome is generally constructed by a sense or negative single-stranded deoxyribonucleic acid (ssDNA) of about 4.7 kilobases in length. The AAV genome is included in terminal inverted repeats (ITRs) at both ends of the DNA chain, and two open reading frames (ORFs): rep and cap. The development of AAV as a gene therapy vector has eliminated the integration ability of the vector by removing rep and cap from the vector DNA. In some embodiments, a gene encoding a wild-type D1L3 enzyme or any variant of D1L3 disclosed herein that is operably connected to a promoter can be inserted between the terminal inverted repeats (ITRs). In some embodiments, after the single-stranded vector DNA is converted into double-stranded DNA by the host cell DNA polymerase complex, the AAV vector comprising the wild-type D1L3 enzyme or any variant D1L3 disclosed herein can form a concatemer in the cell nucleus. In some embodiments, AAV vectors comprising wild-type D1L3 enzyme or any variant D1L3 disclosed herein can thereby form episomal concatemers in the host cell nucleus. In some embodiments, the concatemers can remain intact during the life of non-dividing host cells. In some embodiments, the concatemers may be lost through cell division in dividing cells.
[0087] In an illustrative embodiment, an AAV serotype 8 (AAV2 / 8) vector is used. In some embodiments, the recombinant AAV serotype used for polynucleotide delivery is replication-defective, generally does not insert into the host genome, and shows a lack of pathogenicity and immune response in human subjects. Any AAV vector can be used, including but not limited to AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, and combinations thereof. In some cases, AAV comprises LTRs having a heterologous serotype compared to the capsid serotype (e.g., AAV2 ITRs with AAV5, AAV6, or AAV8 capsids).
[0088] Expression systems that function in human cells are well known in the art and include viral systems. Generally, the promoter that functions in human cells is any DNA sequence that can bind to mammalian RNA polymerase and initiate the downstream (3') transcription of the coding sequence into mRNA. The promoter will have a transcription initiation region generally near the 5' end of the coding sequence, and a TATA box generally located 25-30 base pairs upstream of the transcription start site. It is believed that the TATA box guides RNA polymerase II to start RNA synthesis at the correct site. The promoter will generally also contain an upstream promoter element (enhancer element), which is generally located within 100 to 200 base pairs upstream of the TATA box. The upstream promoter element determines the transcription initiation rate and can work in either direction. Particularly used as a promoter is a promoter from a mammalian viral gene, because the viral gene is generally highly expressed and has a wide host range. Examples include SV40 early promoter, mouse mammary tumor virus LTR promoter, adenovirus major late promoter, herpes simplex virus promoter, and CMV promoter.
[0089] Where appropriate, gene delivery agents such as, for example, integration sequences can also be used. Various integration sequences are known in the art (see, for example, Nunes-Duby et al., Nucleic Acids Res. 26: 391-406, 1998; Sadwoski, J. Bacteriol., 165: 341-357, 1986; Bestor, Cell, 122 (3): 322-325, 2005; Plasterk et al., TIG 15: 326-332, 1999; Kootstra et al., Ann. Rev. Pharm. Toxicol., 43: 413-439, 2003). These include recombinases and transposases. Examples include Cre (Sternberg and Hamilton, J. Mol. Biol., 150:467-486, 1981), λ (Nash, Nature, 247, 543-545, 1974), FIp (Broach et al., Cell, 29:227-234, 1982), R (Matsuzaki et al., J. Bacteriology, 172:610-618, 1990), cpC31 (see, e.g., Groth et al., J. Mol. Biol. 335:667-678, 2004), sleeping beauty ( Beauty) (transposase of the mariner family), and components for integrating viruses (such as AAV, retroviruses) and antiviruses (antivirus) (Kootstra et al., Ann. Rev. Pharm. Toxicol., 43: 413-439, 2003) with components that provide viral integration (such as LTR sequences of retroviruses or lentiviruses and ITR sequences of AAV). In addition, directed and targeted gene integration strategies can be used to insert nucleic acid sequences, including CRISPR / CAS9, zinc fingers, TALEN and meganuclease gene editing technologies.
[0090] Therefore, in some embodiments, the present invention provides mammalian host cells (e.g., human host cells), and methods for preparing and using the host cells. The host cell comprises a heterologous polynucleotide encoding a chromatin enzyme (as described). The host cell delivered to the subject expresses and secretes the encoded chromatin enzyme. In these aspects, the challenge in large-scale manufacturing of chromatin enzymes (e.g., D1L3) is avoided. In addition, by expressing and delivering D1L3 via heterologous expression in leukocytes (e.g., T cells, B cells, or macrophages) or fibroblasts, the D1L3 therapy can be partially positioned to the region of inflammation or tissue destruction or apoptosis or wound healing. In addition, since the circulating half-life of WT D1L3 is less than about 30 minutes, the cell therapy described herein provides sustained therapy, which is provided in some embodiments with as little as one, two, three, or four treatments. In some embodiments, the therapy is provided to the subject for treating cancer (e.g., leukemia) or viral infection (including lower respiratory tract infection). In some embodiments, host cells are produced from the cells of the subject to be treated or the donor of HLA matching. In some embodiments, the cells are HLA-null or are generated from HLA-matched source cells.
[0091] In some aspects, the present disclosure provides a pharmaceutical composition comprising an effective amount of a D1L3 variant according to any embodiment disclosed herein, a D1L3 variant produced according to a method according to any embodiment disclosed herein, a polynucleotide according to any embodiment disclosed herein, a vector according to any embodiment disclosed herein, or a host cell according to any embodiment disclosed herein, and a pharmaceutically acceptable carrier.
[0092] In some embodiments, the composition comprises a D1L3 variant of any embodiment disclosed herein and a pharmaceutically acceptable carrier for parenteral administration. In some embodiments, the pharmaceutical composition is formulated for topical, parenteral, or pulmonary administration. In some embodiments, the pharmaceutical composition is formulated for intradermal, intramuscular, intraperitoneal, intraarticular, intravenous, subcutaneous, intraarterial, ocular, oral, sublingual, pulmonary, or transdermal administration.
[0093] In embodiments, the method for the recombinant production of D1L3 enzyme variants adopts a non-mammalian expression system, for example, a eukaryotic non-mammalian expression system, such as Pichia pastoris. In embodiments, Pichia pastoris encodes a DNA enzyme with a natural signal peptide, thereby allowing secretion from a host cell. In embodiments, the expression system is a mammalian cell expression system, such as Chinese hamster ovary (CHO) cells. In embodiments, the method for the recombinant production of D1L3 enzyme variants further comprises separation and / or purification of the D1L3 enzyme, and modification of the separated and / or purified D1L3 enzyme. In embodiments, modification includes conjugation of the separated and / or purified D1L3 enzyme to a polymer (e.g., not limited to, PEG). In embodiments, the polymer is added to a specific site using desired conjugation chemistry (e.g., not limited to, maleimide chemistry).
[0094] In other aspects, the present disclosure provides a method for treating a subject in need of extracellular chromatin degradation, extracellular trap (ET) degradation, and / or neutrophil extracellular trap (NET) degradation. The method comprises administering a therapeutically effective amount of a D1L3 enzyme or composition described herein. Exemplary indications for subjects in need of extracellular chromatin degradation (including ET or NET degradation) are disclosed in PCT / US18 / 47084, the disclosure of which is hereby incorporated by reference.
[0095] Neutrophils (the main leukocytes in acute inflammation) generate neutrophil extracellular traps (NETs), a grid of high molecular weight chromatin filaments studded with bioactive proteins and peptides that lock bacteria in wounds. Systemic accumulation of NETs damages tissues and organs due to their cytotoxic, proinflammatory, and prothrombotic activities. In fact, NETs are often associated with inflammatory, ischemic, and autoimmune diseases, including systemic lupus erythematosus (SLE).
[0096] In embodiments, the present invention provides a method for treating, preventing, or managing a disease or condition characterized by the presence or accumulation of NETs. See Jiménez-Alcázar et al., "Host DNases prevent vascular occlusion by neutrophil extracellular traps." Science 358(6367):1202-1206(2017). A variety of stimuli (which sometimes promote inflammation and / or pathogenesis) induce NETs. These stimuli include phorbol 12-myristate 13-acetate (PMA, a potent mitogen), lipopolysaccharide (LPS), calcium ionophore A23187, the antibiotic nigericin (which also acts as a potassium ionophore), fungi such as Candida albicans, and bacteria such as Streptococcus agalactiae (a Group B Streptococcus), Klebsiella pneumoniae, and viruses such as SARS-CoV2. Leppkes et al. " Vascular occlusion by neutrophil extracellular traps in COVID-19 .EBioMedicine 58(2020)102925(2020); Claushuis et al., Role of peptidylargininedeiminase 4 in neutrophil extracell ular trap formation and host defense during Klebsiella pneumoniae-induced pneumonia- derived sepsis .” J Immunol. 201: 1241-1252 (2018); and Kenny et al., “ Diverse stimuli engage different neutrophil extracellular trap pathways .” Elife.6:e24437(2017). Diseases or conditions characterized by the presence or accumulation of NETs include, but are not limited to, diseases associated with chronic neutrophilia, neutrophil aggregation and / or leukostasis, thrombosis and vascular occlusion, ischemia-reperfusion injury, surgical and traumatic tissue injury, acute or chronic inflammatory reactions or diseases, autoimmune diseases, cardiovascular diseases, metabolic diseases, systemic inflammation, respiratory inflammatory diseases, renal inflammatory diseases, inflammatory diseases associated with transplanted tissue or hematopoietic stem cell transplantation (e.g., graft-versus-host disease), inflammation caused by viral infection (e.g., COVID-19), and cancer (including leukemia). In an embodiment, the present invention provides a method for treating complete or partial vascular or ductal obstruction, the obstruction involving extracellular chromatin and including NETs in an embodiment.
[0097] In embodiments, the method comprises administering a composition as described herein to a subject. In embodiments, the subject is at risk of vascular occlusion, particularly extracellular chromatin, including chromatin, released by cancer cells and damaged endothelial cells. Thus, in exemplary embodiments, the subject suffers from cancer (e.g., leukemia or solid tumors). In embodiments, the subject suffers from a hematological cancer selected from the group consisting of multiple myeloma (MM), Hodgkin's lymphoma (HL), non-Hodgkin's lymphoma (NHL), chronic lymphocytic leukemia (CLL), and acute lymphoblastic leukemia (ALL). In embodiments, the subject suffers from metastatic cancer.
[0098] Subjects receiving cancer therapy (including but not limited to T cell therapy) are at risk for tumor lysis syndrome and / or cytokine release syndrome, which occurs when tumor cells release their contents (including chromatin) into the bloodstream. Tumor lysis syndrome is a complication during cancer treatment in which a large number of tumor cells are killed simultaneously by the cancer treatment. Tumor lysis syndrome and / or cytokine release syndrome generally occur after treatment of lymphomas and leukemias. In embodiments, the therapies described herein treat, reduce or prevent tumor lysis syndrome.
[0099] In still other embodiments, the subject suffers from an inflammatory disease of the respiratory tract (such as the lower respiratory tract). Exemplary diseases include bacterial and viral infections. In embodiments, the subject suffers from acute respiratory distress syndrome (ARDS), acute lung injury (ALI), pneumonia, or asthma. Exemplary viral infections are RSV and coronavirus infections (such as SARS, or SARS-CoV-2, for example, COVID-19 and variants thereof).
[0100] In still other embodiments, the subject suffers from a disease or illness other than cancer. In embodiments, the disease or illness is an autoimmune or immunological disorder, such as selected from those of systemic lupus erythematosus (SLE), rheumatoid arthritis, psoriasis, inflammatory bowel disease, sprue sprue, pernicious anemia, scleroderma, Graves' disease, Sjögren's syndrome, autoimmune hemolytic anemia (AIHA), myasthenia gravis, cryoglobulinemia, thrombotic thrombocytopenic purpura (TTP), allograft rejection (e.g., transplant rejection of lung, kidney, heart, intestine, liver, pancreas, etc.), pemphigus vulgaris, vitiligo, Hashimoto's disease, Addison's disease, reactive arthritis, and type 1 diabetes.
[0101] In the embodiments, the subject suffers from SLE. The discovery of NETs has led to the following speculation: neutrophils may be the main source of autoantigens (i.e., dsDNA, chromatin) in SLE (Brinkmann et al. Neutrophil Extracellular Traps Kill Bacteria. Science, 303(5663): 1532-1545(2004)). In fact, autoantibodies such as anti-dsDNA, anti-histone and anti-nucleosome antibodies bind to NETs to form pathological ICs. See, for example, Hakkim et al., Impairment of neutrophil extracellular trap degradation is associated with lupus nephritis, Proceedings of the National Academy of Sciences 107: 9813-9818(2010). The accumulation of NET-ICs breaks immune tolerance by activating adaptive immune cells, and the activation of the adaptive immune cells leads to the production of autoantibodies against NET components, thereby forming a vicious cycle of inflammation and autoimmunity. See, e.g., Gupta and Kaplan, The role of neutrophils and NETosis in autoimmune and renal diseases. Nat Rev Nephrol. 12(7):402-13 (2016). Therefore, reducing the accumulation of NETs can break the cycle, thereby providing an attractive therapeutic strategy for SLE.
[0102] In embodiments, the present invention relates to the treatment of a disease or illness characterized by a lack of D1L3 or a lack of D1. In some cases, the subject has a mutation (e.g., a loss-of-function mutation) in the DNA enzyme 1L3 gene or the DNA enzyme 1 gene. Such subjects may exhibit autoimmune diseases such as systemic lupus erythematosus (SLE), lupus nephritis, scleroderma or systemic sclerosis, rheumatoid arthritis, inflammatory bowel disease, Crohn's disease, ulcerative colitis, and urticarial vasculitis. In some cases, the subject has an acquired inhibitor of D1 (e.g., an anti-DNA enzyme 1 antibody and actin) and / or an acquired inhibitor of D1L3 (e.g., an anti-DNA enzyme 1L3 antibody). Such subjects may also suffer from autoimmune or inflammatory diseases (e.g., SLE, systemic sclerosis). In some embodiments, the subject is treated with a D1L3 variant having a half-life extension portion.
[0103] In embodiments, the subject has or is at risk for a NET that blocks the ductal system. For example, a D1L3 enzyme or composition disclosed herein can be administered to a subject to treat pancreatitis, cholangitis, conjunctivitis, mastitis, dry eye, Stevens-Johnson syndrome, vas deferens obstruction, or nephropathy. For example, in embodiments, a D1L3 variant that does not contain any half-life extending moiety is administered as an eye drop to a subject with dry eye.
[0104] In embodiments, the subject has or is at risk for NETs that accumulate on endothelial surfaces (e.g., surgical adhesions), on the skin (e.g., wounds / scars), or in synovial joints (e.g., gout and arthritis, e.g., rheumatoid arthritis). The D1L3 enzymes and compositions described herein can be administered to a subject to treat a condition characterized by accumulation of NETs on endothelial surfaces, such as, but not limited to, surgical adhesions.
[0105] Other diseases and illnesses related to NET that can be treated or prevented using D1L3 enzymes disclosed herein or compositions include: ANCA-associated vasculitis, asthma, chronic obstructive pulmonary disease, neutrophilic dermatosis, dermatomyositis, burns, cellulitis, meningitis, encephalitis, otitis media, pharyngitis, tonsillitis, pneumonia, endocarditis, cystitis, pyelonephritis, appendicitis, cholecystitis, pancreatitis, uveitis, keratitis, disseminated intravascular coagulation, acute kidney injury, acute respiratory distress syndrome, liver shock, hepatorenal syndrome, myocardial infarction, stroke, intestinal ischemia, limb ischemia, testicular torsion, pre-eclampsia, eclampsia, and solid organ transplantation (e.g., kidney, heart, liver, and / or lung transplantation). In addition, D1L3 enzymes disclosed herein or compositions can be used to prevent scar or contracture (e.g., by topical application to skin) in an individual (e.g., with a surgical incision, laceration, or burn) at risk of scar or contracture.
[0106] In embodiments, a D1L3 variant of the present disclosure (e.g., without a half-life extending moiety) is administered to a subject having or at risk of ischemic stroke. In embodiments, the subject may further undergo therapy with tissue plasminogen activator (tPA).
[0107] In an embodiment, the subject has a disease that is or has been treated with wild-type DNA enzymes, including D1 and Streptococcus DNA enzymes. Such diseases or conditions include thrombosis, stroke, sepsis, lung injury, atherosclerosis, viral infection, sickle cell disease, myocardial infarction, ear infection, wound healing, liver injury, endocarditis, liver infection, pancreatitis, primary graft dysfunction, limb ischemia reperfusion, kidney injury, blood coagulation, alum-induced inflammation, liver and kidney injury, pleural effusion, hemothorax, blood clots in the biliary tract, post-inflation anemia, ulcers, otolaryngology disorders, oral Infections, minor injuries, sinusitis, post-operative rhinoplasty, infertility, urinary catheters, debridement, skin tests, pneumococcal meningitis, gout, leg ulcers, cystic fibrosis, Kartagena syndrome, asthma, atelectasis, chronic bronchitis, bronchiectasis, lupus, primary ciliary dyskinesia, bronchiolitis, empyema, pleural infection, cancer, dry eyes, lower respiratory tract infection, chronic hematomas, Alzheimer's disease, and obstructive pulmonary disease.
[0108] In embodiments, the subject has a loss-of-function mutation in one or both D1L3 genes and may exhibit symptoms of SLE, or may further be diagnosed with clinical SLE. In embodiments, the composition is administered no more than about once a week, or no more than about once every two or three weeks, or no more than about once a month.
[0109] In some aspects, the present disclosure provides an expression construct for improving the processing of a polypeptide precursor in a host, the expression construct comprising a signal peptide fused to the polypeptide via a linker. In some embodiments, the linker has a length of at least three amino acids. In some embodiments, the signal peptide is completely removed from the polypeptide. In some embodiments, the signal peptide is not removed from the polypeptide. In some embodiments, in the absence of a linker, the signal peptide is incompletely processed in the host.
[0110] In some embodiments, the polypeptide is or includes an enzyme, a cytokine, a cytokine agonist, a cytokine antagonist, a hormone, a hormone agonist, a hormone antagonist, an antibody or its antigen-binding fragment, an antibody-like molecule or its antigen-binding fragment, an antigen, a component of a vaccine, a fusion protein, or a combination thereof. In some embodiments, the enzyme is selected from DNase 1 (D1), DNase 1-like protein 1 (D1L1), DNase 1-like protein 2 (D1L2), DNase 1-like protein 3 isoform 1 (D1L3), DNase 1-like protein 3 isoform 2 (D1L3-2), DNase 2A (D2A) and DNase 2B (D2B) or variants thereof. In some embodiments, the cytokine is selected from IL-1, IL-2, IL-5, IL-6, IL-10 and IL-13, IL-12, CXCL8 (formerly known as IL-18), interferon-γ (IFN-γ) and tumor necrosis factor-β (TNF-β), TNF-α, G-CSF and GM-CSF.
[0111] In some embodiments, the hormone is selected from adrenocorticotropic hormone (ACTH), adropin, amylin, angiotensin, atrial natriuretic peptide (ANP), calcitonin, cholecystokinin (CCK), exenatide, gastrin, ghrelin, glucagon, GLP-1, growth hormone, GIP, EPO, follicle-stimulating hormone (FSH), insulin, leptin, luteinizing hormone (LH), melanocyte-stimulating hormone (MSH), oxytocin, parathyroid hormone (PTH), prolactin, renin, somatostatin, thyroid-stimulating hormone (TSH), thyrotropin-releasing hormone (TRH), vasopressin, and vasoactive intestinal peptide (VIP). In some embodiments, the hormone or polypeptide contains a half-life extending portion, such as an albumin or Fc fusion as described herein. In some embodiments, this fusion is located at the C-terminus.
[0112] In some embodiments, the linker is at least three, or at least four, or at least five, or at least six, or at least seven, or at least eight, or at least nine, or at least ten, or at least 11, or at least 12, or at least 13, or at least 14, or at least 15, or more amino acids in length. In some embodiments, the first amino acid residue of the N-terminal extension is not a Gly residue. In some embodiments, the last amino acid residue of the N-terminal extension is not a Ser residue. In some embodiments, the last amino acid residue of the N-terminal extension is not a polar or charged amino acid residue selected from Ser, Thr, Gln, Asn, Glu, Asp, Arg, His, and Lys. In some embodiments, the last amino acid residue of the N-terminal extension is an amino acid selected from Gly, Ala, and Val. In some embodiments, the last amino acid residue of the N-terminal extension is a Gly residue. In some embodiments, the first amino acid residue of the N-terminal extension is not a Gly residue, and the last amino acid residue of the N-terminal extension is a Gly residue. In some embodiments, the N-terminal extension is primarily Ser and Gly residues, consists essentially of Ser and Gly residues, or consists of Ser and Gly residues. In some embodiments, the linker comprises glycine, cysteine, and serine residues. In some embodiments, the linker comprises the amino acid sequence of SGGGG (SEQ ID NO: 60). In some embodiments, the linker comprises the amino acid sequence of CGGGG (SEQ ID NO: 74). In some embodiments, the linker comprises the amino acid sequence of SGGSGGSGG (SEQ ID NO: 61). In some embodiments, the linker comprises the amino acid sequence of SGGSGGSGGSGGSGGSGG (SEQ ID NO: 62). In some embodiments, the linker comprises the amino acid sequence of LVPRG (SEQ ID NO: 64). In some embodiments, the linker comprises the amino acid sequence of SGGGGLVPRGSGGGG (SEQ ID NO: 65).
[0113] In some embodiments, the signal peptide is a signal peptide of a naturally secreted protein. In some embodiments, the signal peptide is a chimeric or synthetic signal peptide that enables protein secretion. In some embodiments, the signal peptide is a prokaryotic signal peptide. In some embodiments, the signal peptide is a microbial signal peptide. In some embodiments, the signal peptide is a eukaryotic signal peptide. In some embodiments, the signal peptide is a yeast signal peptide. In some embodiments, the signal peptide is a mammalian signal peptide. Signal peptides suitable for secretion are disclosed in the following documents: U.S. Patent Nos. 5,580,758; 6,107,057; 7,741,075; 10,435,694; 11,306,127; 11,370,815; U.S. Patent Application Publication Nos. 2007 / 0117186, 2010 / 0055125, and 2016 / 0168198, the disclosures of each of which are hereby incorporated by reference.
[0114] In some embodiments, the N-terminal extension and / or the conjugate of the N-terminal extension and the sequence derived from the polypeptide is non-immunogenic. In some embodiments, the last amino acid of the N-terminal extension is not a polar or charged or aromatic amino acid. In some embodiments, the last two, or last three, or last four amino acids of the N-terminal extension are not polar, charged and / or aromatic amino acids. In some embodiments, the N-terminal residue is not methionine (Met). In some embodiments, the last amino acid of the N-terminal extension is Gly or Ala. In some embodiments, the N-terminal extension comprises an amino acid with a chemical group suitable for chemical conjugation, the chemical group being selected from a thiol group, an amino group, an amide group, and a carboxyl group. In some embodiments, the chemical modification is site-specific PEGylation, glycosylation, etc. In some embodiments, the signal peptide is selected from the group consisting of DNase 1L3 (SEQ ID NO:37), alpha mating factor (SEQ ID NO:38), alpha mating factor presequence (SEQ ID NO:39), human serum albumin (SEQ ID NO:40), bovine DNase 1 (SEQ ID NO:41), bovine DNase 1 + Kex2 site (SEQ ID NO:42), alpha amylase (SEQ ID NO:43), glucoamylase signal peptide (SEQ ID NO:44), inulinase (SEQ ID NO:45), invertase (SEQ ID NO:46), killer protein (SEQ ID NO:47), and lysozyme (SEQ ID NO:48).
[0115] In some embodiments, the signal peptide is selected from the group consisting of E. coli OmpA signal peptide (SEQ ID NO: 56), E. coli DsbA signal peptide (SEQ ID NO: 67), E. coli ST-II signal peptide (SEQ ID NO: 68), E. coli FimD signal peptide (SEQ ID NO: 55), Salmonella enteritidis DsbA signal peptide (SEQ ID NO: 51), synthetic Bordetella pertussis signal peptide (SEQ ID NO: 57), synthetic signal peptide sequences (e.g., SEQ ID NO: 49, 50, 52, 53, and 54). In some embodiments, the host is yeast, or a cell line selected from a mammalian cell line and an insect cell line. In some embodiments, the host is Pichia pastoris. In some embodiments, the signal peptide is the alpha mating factor (αMF) prepro secretory leader sequence from Saccharomyces cerevisiae (SEQ ID NO: 38).
[0116] Further aspects and embodiments of the invention will become apparent from the following examples.
[0117] Example
[0118] Example 1. Optimization of N-terminal secretion signal peptide
[0119] The expression of D1L3 enzyme and D1L3-albumin fusion protein in Pichia pastoris is disclosed in PCT International Application Publication No. WO2020076817, which is hereby incorporated by reference in its entirety. Briefly, the alpha mating factor (aMF) prepro secretory leader sequence from Saccharomyces cerevisiae (SEQ ID NO: 38) was used as the N-terminal secretion signal peptide. See Figure 1A aMF is a commonly used and effective tool for heterologous protein expression in Pichia pastoris. Figure 1B As shown in , the combination of aMF and human D1L3 resulted in unexpected non-processing of aMF with concomitant glycosylation.
[0120] exist Figure 1B In the present study, D1L3 was properly processed when N-terminally directed by its native secretory signal peptide. However, the native signal peptide of D1L3 resulted in a 3.5-fold decrease in expression titer when compared to aMF.
[0121] A series of secretion signal peptides were screened for expression in Pichia pastoris, including the following secretion signal peptides: serum albumin, α-amylase, glucoamylase, inulinase, invertase, killer protein, lysozyme, and bovine DNase 1. In addition, a variant of the bovine DNase 1 secretion signal peptide was tested, which contains a Kex2 cleavage site. See U.S. Patent No. 7,118,901, which is hereby incorporated by reference in its entirety. The experiment also included the signal peptide of human DNase 1L3 and two forms of the α mating factor (αMF) prepro secretion leader sequence from Saccharomyces cerevisiae.
[0122] Briefly, plasmids were synthesized using cDNA encoding various signal peptides of SEQ ID NOs: 37 to 48 at the N-terminus of human DNase 1L3. These plasmids were transformed into Pichia pastoris cells via electroporation. Clonal cells were cultured, and supernatants were analyzed by microfluidic capillary electrophoresis (mCE) to characterize target protein expression. As shown in Table 1, no or low levels of DNase 1L3 expression were detected using all tested secretion signal peptides. Optimal titers were observed using aMF (SEQ ID NO: 38), while DNase 1L3 levels were undetectable using SEQ ID NOs: 41 and 43-46.
[0123] Table 1: Titers of DNase 1L3 or DNase 1L3-BDD with various secretion signals.
[0124]
[0125]
[0126] Structural analysis of DNase 1L3 revealed that reducing access of the α-mating factor cleavage enzyme Kex2 to its consensus cleavage sequence, "LEKR," impairs proper processing into mature DNase 1L3.
[0127] A novel secretion signal peptide was designed, which contains the α mating factor (αMF) prepro secretion leader sequence from Saccharomyces cerevisiae (SEQ ID NO: 38) and a linker sequence to facilitate access to the Kex2 cleavage site ( Figure 2 ).
[0128] In a pilot study, flexible glycine-serine linker compositions with lengths ranging from 1 to 15 amino acids were tested. The α mating factor (α MF) prepro secretory leader sequence (SEQ ID NO: 38)-linker sequence from Saccharomyces cerevisiae was coupled to a DNA enzyme 1L3 variant with a basic domain deletion. As shown in Table 2, a significant increase in expression titer was observed starting with a linker length of 5 amino acids.
[0129] Table 2: Titers of DNase 1L3-BDD using the aMF secretion signal with or without a linker separating the linker from the DNase 1L3-BDD.
[0130]
[0131] DNase 1L3 was purified from culture supernatants using affinity chromatography. Analysis by mass spectrometry showed complete processing of the alpha mating factor (αMF) prepro secretory leader from Saccharomyces cerevisiae (SEQ ID NO: 38) and confirmed the predicted N-terminal amino acid sequence.
[0132] Accidentally, if Figure 3 As shown in Table 3, a mass shift of 210 Da was observed, which may be due to myristoylation of the N-terminal glycine residue. A new set of linker sequences was designed, which is characterized by an N-terminal serine residue. The linker length ranges from 1 to 18 amino acids. As shown in Table 3, a significant increase in expression titer was observed starting with a linker length greater than 3 amino acids.
[0133] Table 3: Titers of DNase 1L3-BDD using aMF secretion signals with additional linkers.
[0134]
[0135] As shown in Table 3, a significant increase in expression titer was observed starting at 3 amino acids. However, mCE and Western blot analysis showed that a linker length of 3 amino acids resulted in incomplete processing of the α-mating factor (αMF) pre-pro secretory leader sequence from Saccharomyces cerevisiae (SEQ ID NO: 38). The use of a 5-amino acid linker with the sequence SGGGG (SEQ ID NO: 60) resulted in both increased expression levels and complete processing of the α-mating factor (αMF) pre-pro secretory leader sequence from Saccharomyces cerevisiae (SEQ ID NO: 38). All linkers longer than 5 amino acids both increased expression levels and enabled complete processing of the α-mating factor (αMF) pre-pro secretory leader sequence, including linkers with the following sequences: SGGS GGSGG (SEQ ID NO: 61), SGGSGGSGSS (SEQ ID NO: 69), and SGGSGGSGGSGGSGGSGG (SEQ ID NO: 62).
[0136] The potential immunogenicity of the linker itself and the linker-D1L3 conjugates (e.g., SGGSGGSGSS-MRICSFNVRS (SEQ ID NO: 79); SGGGG-MRICSFNVRS (SEQ ID NO: 80); and SGGSGGSGG-MRIC SFNVRS (SEQ ID NO: 81)) was analyzed using a computer immunogenicity risk prediction algorithm. This analysis showed that the Ser residue of the SEQ ID NO: 79 conjugate contributed to a potential dominant epitope, which is indicated by underlined bold font. Other linkers with Gly residues at the terminal end did not produce similar dominant epitopes. These results indicate, among other things, that linkers used for expression of D1L3 should not end with a Ser residue and possibly other polar residues; and that linkers used for expression of D1L3 can end with a Gly residue to reduce the risk of immunogenicity.
[0137] The BDD_D1L3 enzyme (S283_S305del, SEQ ID NO: 13) was produced by a construct comprising an alpha mating factor + SGGGG linker (SEQ ID NO: 60), thereby generating SEQ ID NO: 63, and its enzymatic activity on chromatin substrates was analyzed. Briefly, DNase 1 (D1) and SEQ ID NO: 63 were produced in Pichia pastoris. The degradation of high molecular weight (HMW) chromatin (i.e., purified nuclei from HEK293 cells) was used as a readout to characterize the enzymatic activity in the culture supernatant. Briefly, HMW chromatin was incubated with equal amounts of D1 or SEQ ID NO: 63. After incubation, DNA was isolated and visualized by agarose gel electrophoresis (AGE). Figure 4 As shown in the results, it was observed that, unlike D1, D1L3 S283_S305del, which has an N-terminal SGGGG, specifically and efficiently degraded HMW chromatin (approximately 50-300,000,000 base pairs) into nucleosomes (approximately 180 base pairs), the basic unit of chromatin fiber, while this effect was not observed in samples using D1. These results indicate that the N-terminal extension does not affect the enzymatic activity of the D1L3 enzyme.
[0138] Example 2. Optimization of the C-terminus
[0139] SEQ ID NO: 63 was analyzed by mass spectrometry. Figure 5AAs shown in , the main elution peak of the resulting protein elution chromatogram was found to have a left shoulder, indicating sample heterogeneity. Various derivatives of SEQ ID NO: 63 were analyzed to address the issue of sample heterogeneity, and it was found that derivatives having the S283_S305delinsSSR mutation (i.e., having a modified C-terminus including a deletion of the basic domain, but characterized by a C-terminal addition of 3 amino acids (SSR; SEQ ID NO: 66)) did not produce such sample heterogeneity ( Figure 5B These results indicate that the C-terminal extension removes the source of the observed heterogeneity. Enzymes containing an N-terminal extension and / or a C-terminal extension do not affect the enzymatic activity of the D1L3 enzyme.
[0140] These results were confirmed by studying two DNA enzyme 1L3 variants characterized by either a wild-type C-terminal amino acid sequence (i.e., SSRAFTNSKKSVTLRKKTKSKRS (SEQ ID NO: 75)) or a modified C-terminal amino acid sequence (i.e., SSR). The variants were expressed in Pichia pastoris using the α mating factor as a signal sequence in conjunction with an N-terminal SGGGG linker (SEQ ID NO: 60). Western blot analysis of the supernatants did not detect a theoretical mass difference of 2.3 kDa between the two variants (e.g., Figure 7 ). In addition, intact mass analysis showed that the secreted DNase 1L3 variant had the same mass that matched the theoretical mass of the DNase 1L3 variant with the modified C-terminal amino acid sequence (i.e., SSR). In summary, these data indicate that the wild-type C-terminal amino acid sequence (i.e., SSRAFTNSKKSVTLRKKTKSKRS (SEQ ID NO: 75)) is cleaved by Pichia pastoris to the modified C-terminal amino acid sequence (i.e., SSR) during secretion. Pichia pastoris contains processing enzymes, e.g., Kex 1 and Kex 2, which have homologous human counterparts, e.g., furin. Processing of the DNase 1L3 C-terminus is also expected to occur naturally during secretion in humans.
[0141] Example 3: Cysteine mutation
[0142] Wild-type D1L3 contains an unpaired cysteine at position 48 (e.g., C48) (with respect to the mature protein sequence). Mutation of this cysteine stabilizes D1L3 and prevents cross-linking with plasma proteins. The effects of different amino acid substitutions at C48 on enzyme activity were tested. Enzyme activity was characterized using degradation of high molecular weight (HMW) chromatin (i.e., purified nuclei from HEK293 cells). Briefly, HMW chromatin was incubated with equal amounts of D1L3 variants. After incubation, DNA was isolated and degradation was visualized by agarose gel electrophoresis (AGE). Figure 8 As shown in , mutations of C48A or C48G were associated with increased enzyme activity compared to the D1L3 variant. These results suggest that removal of the unpaired cysteine increases the enzymatic activity of the D1L3 enzyme.
[0143] Example 4. Dimerization of DNase 1L3 via unpaired cysteines
[0144] Structural analysis of DNA enzyme 1L3 revealed that the unpaired C48 is located on the surface of the molecule. Western blot analysis of Pichia pastoris supernatants revealed that dimerization was observed in DNA enzyme 1L3 variants with unpaired C48, but not in variants with mutated C48 (e.g., Figure 9 These data suggest that dimerization of DNase 1L3 may occur physiologically.
[0145] Test the possibility of inserting new cysteine in the D1L3 variant of carrying the C48 of mutation, so that, for example, site-specific PEGylation can be performed. BDD-D1L3 enzyme (A286_S305del) and wild-type D1L3 are produced by the construct comprising α mating factor+CGGGG joint (SEQ ID NO:74). Four samples are compared. Sample 1 (SEQ ID NO:70) is produced by the construct comprising α mating factor+SGGGG joint, and is a D1L3 variant with C48A substitution and a C-terminal extension with sequence SSR. Samples 2-4 (respectively SEQ ID NO:71 to 73) are produced by the construct comprising α mating factor+CGGGG joint (SEQ ID NO:74), and are D1L3 variants with C48A / S substitution and a C-terminal extension. Dimers are detected in the DNA enzyme 1L3 variant containing N-terminal cysteine. In addition to D1L3 monomers, analysis of supernatants by Western blotting with anti-DNase 1L3 under non-reducing conditions revealed dimerized BDD-D1L3 variants (e.g., Figure 6 ). No dimers were observed under reducing conditions or with constructs containing the α mating factor + SGGGG linker. The data confirm that the introduction of the CGGGG linker enables dimerization via a disulfide bridge at the N-terminal cysteine residue. Similarly, N-terminal variants (e.g., to a cysteine residue) allow for increased functionality, such as site-specific PEGylation.
[0146] sequence
[0147] Wild-type human DNase
[0148] SEQ ID NO: 1
[0149] DNase 1 (NP_005212.2): signal peptide , mature protein:
[0150] MRGMKLLGALLALAALLQGAVS LKIAAFNIQTFGETKMSNATLVSYIVQILSRYDIALVQEVRDSHLTAVGKLLDNLNQDAPDTYHYVVSEPLGRNSYKERYLFVYRPDQVSAVDSYYYDDGCEPCGNDTFNREPAIVRFFSRFTEVREFAIVPLHAAPGDAVAEIDALYDVYLDVQEKWGLEDVMLMGDFNAGCSYVRPSQWSSIRLWTSPTFQWLIPDSADTTATPTHCAYDRIVVAGMLLRGAVVPDSALPFNFQAAYGLSDQLAQAISDHYPVEVMLK
[0151] SEQ ID NO:2
[0152] DNase 1-like protein 1 (NP_006721.1): signal peptide ; mature protein:
[0153] MHYPTALLFLILANGAQA FRICAFNAQRLTLAKVAREQVMDTLVRILARCDIMVLQEVVDSSGSAIPLLLRELNRFDGSGPYSTLSSPQLGRSTYMETYVYFYRSHKTQVLSSYVYNDEDDVFAREPFVAQFSLPSNVLPSLVLVPLHTTPKAVEKELNALYDVFLEVSQHWQSKDVILLGDFNADCASLTKKRLDKLELRTEPGFHWVIADGEDTTVRASTHCTYDRVVLHGERCRSLLHTAAAFDFPTSFQLTEEEALNISDHYPVEVELKLSQAHSVQPLSLTVLLLLSLLSPQLCPAA
[0154] SEQ ID NO:3
[0155] DNase 1-like protein 2 (NP_001365.1): signal peptide , mature protein:
[0156] MGGPRALLAALWALEAAGTAALRIGAFNIQSFGDSKVSDPACGSIIAKILAGYDLALVQEVRDPDLSAVSALMEQINSVSEHEYSFVSSQPLGRDQYKEMYLFVYRKDAVSVVDTYLYPDPEDVFSREPFVVKFSAPGTGERAPPLPSRRALTPPPLPAAAQNLVLIPLHAAPHQAVAEIDALYDVYLDVIDKWGTDDMLFLGDFNADCSYVRAQDWAAIRLRSSEVFKWLIPDSADTTVGNSDCAYDRIVACGARLRRSLKPQSATVHDFQEEFGLDQTQALAISDHFPVEVTLKFHR
[0157] SEQ ID NO:4
[0158] DNAse 1-like protein 3; Isoform 1 (NP_004935.1): signal peptide , mature protein:
[0159] MSRELAPLLLLLLSIHSALA MRICSFNVRSFGESKQEDKNAMDVIVKVIKRCDIILVMEIKDSNNRICPILMEKLNRNSRRGITYNYVISSRLGRNTYKEQYAFLYKEKLVSVKRSYHYHDYQDGDADVFSREPFVVWFQSPHTAVKDFVIIPLHTTPETSVKEIDELVEVYTDVKHRWKAENFIFMGDFNAGCSYVPKKAWKNIRLRTDPRFVWLIGDQEDTTVKKSTNCAYDRIVLRGQEIVSSVVPKSNSVFDFQKAYKLTEEEALDVSDHFPVEFKLQSSRAFTNSKKSVTLRKKTKSKRS
[0160] SEQ ID NO:5
[0161] DNAse 1-like protein 3, Isoform 2 (NP_001243489.1): signal peptide ; mature protein:
[0162] MSRELAPLLLLLLSIHSALAMRICSFNVRSFGESKQEDKNAMDVIVKVIKRCDIILVMEIKDSNNRICPILMEKLNREKLVSVKRSYHYHDYQDGDADVFSREPFVVWFQSPHTAVKDFVIIPLHTTPETSVKEIDELVEVYTDVKHRWKAENFIFMGDFNAGCSYVPKKAWKNIRLRTDPRFVWLIGDQEDTTVKKSTNCAYDRIVLRGQEIVSSVVPKSNSVFDFQKAYKLTEEEALDVSDHFPVEFKLQSSRAFTNSKKSVTLRKKTKSKRS
[0163] SEQ ID NO:6
[0164] DNAse 2A (O00115): signal peptide ; Mature protein:
[0165] MIPLLLAALLCVPAGALTCYGDSGQPVDWFVVYKLPALRGSGEAAQRGLQYKYLDESSGGWRDGRALINSPEGAVGRSLQPLYRSNTSQLAFLLYNDQPPQPSKAQDSSMRGHTKGVLLLDHDGGFWLVHSVPNFPPPASSAAYSWPHSACTYGQTLLCVSFPFAQFSKMGKQLTYTYPWVYNYQLEGIFAQEFPDLENVVKGHHVSQEPWNSSITLTSQAGAVFQSFAKFSKFGDDLYSGWLAAALGTNLQVQFWHKTVGILPSNCSDIWQVLNVNQIAFPGPAGPSFNSTEDHSKWCVSPKGPWTCVGDMNRNQGEEQRGGGTLCAQLPALWKAFQPLVKNYQPCNGMARKPSRAYKI
[0166] SEQ ID NO:7
[0167] DNAse 2B (Q8WZ79): signal peptide ; Mature protein:
[0168] MKQKMMARLLRTSFALLFLGLFGVLGAATISCRNEEGKAVDWFTFYKLPKRQNKESGETGLEYLYLDSTTRSWRKSEQLMNDTKSVLGRTLQQLYEAYASKSNNTAYLIYNDGVPKPVNYSRKYGHTKGLLLWNRVQGFWLIHSIPQFPPIPEEGYDYPPTGRRNGQSGICITFKYNQYEAIDSQLLVCNPNVYSCSIPATFHQELIHMPQLCTRASSSEIPGRLLTTLQSAQGQKFLHFAKSDSFLDDIFAAWMAQRLKTHLLTETWQRKRQELPSNCSLPYHVYNIKAIKLSRHSYFSSYQDHAKWCISQKGTKNRWTCIGDLNRSPHQAFRSGGFICTQNWQIYQAFQGLVLYYESCK
[0169] C-terminal deletion mutant of human DNAse1L3
[0170] SEQ ID NO:8
[0171] S305del: signal peptide ; Mature protein:
[0172] MSRELAPLLLLLLSIHSALA MRICSFNVRSFGESKQEDKNAMDVIVKVIKRCDIILVMEIKDSNNRICPILMEKLNRNSRRGITYNYVISSRLGRNTYKEQYAFLYKEKLVSVKRSYHYHDYQDGDADVFSREPFVVWFQSPHTAVKDFVIIPLHTTPETSVKEIDELVEVYTDVKHRWKAENFIFMGDFNAGCSYVPKKAWKNIRLRTDPRFVWLIGDQEDTTVKKSTNCAYDRIVLRGQEIVSSVVPKSNSVFDFQKAYKLTEEEALDVSDHFPVEFKLQSSRAFTNSKKSVTLRKKTKSKR
[0173] SEQ ID NO:9
[0174] K303_S305del: signal peptide ; Mature protein:
[0175] MSRELAPLLLLLLSIHSALAMRICSFNVRSFGESKQEDKNAMDVIVKVIKRCDIILVMEIKDSNNRICPILMEKLNRNSRRGITYNYVISSRLGRNTYKEQYAFLYKEKLVSVKRSYHYHDYQDGDADVFSREPFVVWFQSPHTAVKDFVIIPLHTTPETSVKEIDELVEVYTDVKHRWKAENFIFMGDFNAGCSYVPKKAWKNIRLRTDPRFVWLIGDQEDTTVKKSTNCAYDRIVLRGQEIVSSVVPKSNSVFDFQKAYKLTEEEALDVSDHFPVEFKLQSSRAFTNSKKSVTLRKKTKS
[0176] SEQ ID NO:10
[0177] V294_S305del: signal peptide ; mature protein:
[0178] MSRELAPLLLLLLSIHSALA MRICSFNVRSFGESKQEDKNAMDVIVKVIKRCDIILVMEIKDSNNRICPILMEKLNRNSRRGITYNYVISSRLGRNTYKEQYAFLYKEKLVSVKRSYHYHDYQDGDADVFSREPFVVWFQSPHTAVKDFVIIPLHTTPETSVKEIDELVEVYTDVKHRWKAENFIFMGDFNAGCSYVPKKAWKNIRLRTDPRFVWLIGDQEDTTVKKSTNCAYDRIVLRGQEIVSSVVPKSNSVFDFQKAYKLTEEEALDVSDHFPVEFKLQSSRAFTNSKKS
[0179] SEQ ID NO:11
[0180] K291_S305del: signal peptide ; mature protein:
[0181] MSRELAPLLLLLLSIHSALAMRICSFNVRSFGESKQEDKNAMDVIVKVIKRCDIILVMEIKDSNNRICPILMEKLNRNSRRGITYNYVISSRLGRNTYKEQYAFLYKEKLVSVKRSYHYHDYQDGDADVFSREPFVVWFQSPHTAVKDFVIIPLHTTPETSVKEIDELVEVYTDVKHRWKAENFIFMGDFNAGCSYVPKKAWKNIRLRTDPRFVWLIGDQEDTTVKKSTNCAYDRIVLRGQEIVSSVVPKSNSVFDFQKAYKLTEEEALDVSDHFPVEFKLQSSRAFTNS
[0182] SEQ ID NO:12
[0183] R285_S305del: signal peptide ; Mature protein:
[0184] MSRELAPLLLLLLSIHSALA MRICSFNVRSFGESKQEDKNAMDVIVKVIKRCDIILVMEIKDSNNRICPILMEKLNRNSRRGITYNYVISSRLGRNTYKEQYAFLYKEKLVSVKRSYHYHDYQDGDADVFSREPFVVWFQSPHTAVKDFVIIPLHTTPETSVKEIDELVEVYTDVKHRWKAENFIFMGDFNAGCSYVPKKAWKNIRLRTDPRFVWLIGDQEDTTVKKSTNCAYDRIVLRGQEIVSSVVPKSNSVFDFQKAYKLTEEEALDVSDHFPVEFKLQSS
[0185] SEQ ID NO:13
[0186] S283_S305del: signal peptide ; Mature protein:
[0187] MSRELAPLLLLLLSIHSALAMRICSFNVRSFGESKQEDKNAMDVIVKVIKRCDIILVMEIKDSNNRICPILMEKLNRNSRRGITYNYVISSRLGRNTYKEQYAFLYKEKLVSVKRSYHYHDYQDGDADVFSREPFVVWFQSPHTAVKDFVIIPLHTTPETSVKEIDELVEVYTDVKHRWKAENFIFMGDFNAGCSYVPKKAWKNIRLRTDPRFVWLIGDQEDTTVKKSTNCAYDRIVLRGQEIVSSVVPKSNSVFDFQKAYKLTEEEALDVSDHFPVEFKLQ
[0188] SEQ ID NO:14
[0189] K298_S305del: signal peptide ; Mature protein:
[0190] MSRELAPLLLLLLSIHSALA MRICSFNVRSFGESKQEDKNAMDVIVKVIKRCDIILVMEIKDSNNRICPILMEKLNRNSRRGITYNYVISSRLGRNTYKEQYAFLYKEKLVSVKRSYHYHDYQDGDADVFSREPFVVWFQSPHTAVKDFVIIPLHTTPETSVKEIDELVEVYTDVKHRWKAENFIFMGDFNAGCSYVPKKAWKNIRLRTDPRFVWLIGDQEDTTVKKSTNCAYDRIVLRGQEIVSSVVPKSNSVFDFQKAYKLTEEEALDVSDHFPVEFKLQSSRAFTNSKKSVTLR
[0191] SEQ ID NO:15
[0192] R297_S305del: signal peptide ; Mature protein:
[0193] MSRELAPLLLLLLSIHSALAMRICSFNVRSFGESKQEDKNAMDVIVKVIKRCDIILVMEIKDSNNRICPILMEKLNRNSRRGITYNYVISSRLGRNTYKEQYAFLYKEKLVSVKRSYHYHDYQDGDADVFSREPFVVWFQSPHTAVKDFVIIPLHTTPETSVKEIDELVEVYTDVKHRWKAENFIFMGDFNAGCSYVPKKAWKNIRLRTDPRFVWLIGDQEDTTVKKSTNCAYDRIVLRGQEIVSSVVPKSNSVFDFQKAYKLTEEEALDVSDHFPVEFKLQSSRAFTNSKKSVTL
[0194] SEQ ID NO:16
[0195] S293_S305del: signal peptide ; Mature protein:
[0196] MSRELAPLLLLLLSIHSALA MRICSFNVRSFGESKQEDKNAMDVIVKVIKRCDIILVMEIKDSNNRICPILMEKLNRNSRRGITYNYVISSRLGRNTYKEQYAFLYKEKLVSVKRSYHYHDYQDGDADVFSREPFVVWFQSPHTAVKDFVIIPLHTTPETSVKEIDELVEVYTDVKHRWKAENFIFMGDFNAGCSYVPKKAWKNIRLRTDPRFVWLIGDQEDTTVKKSTNCAYDRIVLRGQEIVSSVVPKSNSVFDFQKAYKLTEEEALDVSDHFPVEFKLQSSRAFTNSKK
[0197] SEQ ID NO:17
[0198] K292_S305del: signal peptide ; Mature protein:
[0199] MSRELAPLLLLLLSIHSALAMRICSFNVRSFGESKQEDKNAMDVIVKVIKRCDIILVMEIKDSNNRICPILMEKLNRNSRRGITYNYVISSRLGRNTYKEQYAFLYKEKLVSVKRSYHYHDYQDGDADVFSREPFVVWFQSPHTAVKDFVIIPLH TTPETSVKEIDELVEVYTDVKHRWKAENFIFMGDFNAGCSYVPKKAWKNIRLRTDPRFVWLIGDQEDTTVKKSTNCAYDRIVLRGQEIVSSVVPKSNSVFDFQKAYKLTEEEALDVSDHFPVEFKLQSSRAFTNSK
[0200] SEQ ID NO:63
[0201] α-mating factor signal peptide ; Linker; D1L3 S283_S305del mature protein (e.g., the enzyme produced by Pichia pastoris includes a SGGGG linker):
[0202]
[0203] SEQ ID NO:66
[0204] α-mating factor signal peptide ; Connector; with C48S replaces D1L3 S283_S305del mature protein; C-terminus Extension (e.g., the enzyme produced from Pichia pastoris includes a SGGGG linker and a C-terminal extension):
[0205]
[0206] SEQ ID NO:70
[0207] α-mating factor signal peptide ; Connector; with C48A replacement D1L3 S283_S305del mature protein; C-terminus Extension (e.g., the enzyme produced from Pichia pastoris includes a SGGGG linker and a C-terminal extension):
[0208]
[0209] SEQ ID NO:71
[0210] α-mating factor signal peptide ; Connector; with C48A replacementD1L3 S283_S305del mature protein; C-terminus Extension (e.g., the enzyme produced by Pichia pastoris includes a CGGGG linker and a C-terminal extension):
[0211]
[0212] SEQ ID NO:72
[0213] α-mating factor signal peptide ; Connector; with C48S replaces D1L3 S283_S305del mature protein; C-terminus Extension (e.g., the enzyme produced by Pichia pastoris includes a CGGGG linker and a C-terminal extension):
[0214]
[0215] SEQ ID NO:73
[0216] α-mating factor signal peptide ; Connector; with C48A replacement D1L3 variant; C-terminal basic domain (e.g., the enzyme produced by Pichia pastoris includes a CGGGG linker and a C-terminal extension):
[0217]
[0218] SEQ ID NO:77
[0219] N-terminal extension; having C48A replacement D1L3 S283_S305del mature protein; C-terminal extension (e.g., the enzyme produced from Pichia pastoris includes a SGGGG linker and a C-terminal extension):
[0220]
[0221] Adapter sequence
[0222] SEQ ID NO: 18
[0223] GGGGS
[0224] SEQ ID NO: 19
[0225] GGGGSGGGGSGGGGS
[0226] SEQ ID NO:20
[0227] APAPAPAPAPAPAP
[0228] SEQ ID NO:21
[0229] AEAAAKEAAAKA
[0230] SEQ ID NO:22
[0231] SGGSGSS
[0232] SEQ ID NO:23
[0233] SGGSGGSGGSGGSGSS
[0234] SEQ ID NO:24
[0235] SGGSGGSGGSGGSGGSGGSGGSGGSGGSGSS
[0236] SEQ ID NO:25
[0237] GGSGGSGGSGGSGGSGGSGGSGGSGGSGS
[0238] Other sequences
[0239] SEQ ID NO:26
[0240] Human serum albumin (mature protein):
[0241] DAHKSEVAHRFKDLGEENFKALVLIAFAQYLQQCPFEDHVKLVNEVTEFAKTCVADESAENCDKSLHTLFGDKLCTVATLRETYGEMADCCAKQEPERNECFLQHKDDNPNLPRLVRPEVDVMCTAFHDNEETFLKKYLYEIARRHPYFYAPELLFFAKRYKAAFTECCQAADKAACLLPKLDELRDEGKASSAKQRLKCASLQKFGERAFKAWAVARLSQRFPKAEFAEVSKLVTDLTKVHTECCHGDLLECADDRADLAKYICENQDSISSKLKECCEKPLLEKSHCIAEVENDEMPADLPSLAADFVESKDVCKNYAEAKDVFLGMFLYEYARRHPDYSVVLLLRLAKTYETTLEKCCAAADPHECYAKVFDEFKPLVEEPQNLIKQNCELFEQLGEYKFQNALLVRYTKKVPQVSTPTLVEVSRNLGKVGSKCCKHPEAKRMPCAEDYLSVVLNQLCVLHEKTPVSDRVTKCCTESLVNRRPCFSALEVDETYVPKEFNAETFTFHADICTLSEKERQIKKQTALVELVKHKPKATKEQLKAVMDDFAAFVEKCCKADDKETCFAEEGKKLVAASQAALGL
[0242] SEQ ID NO:27
[0243] Human Factor XI:
[0244] MIFLYQVVHFILFTSVSGECVTQLLKDTCFEGGDITTVFTPSAKYCQVVCTYHPRCLLFTFTAESPSEDPTRWFTCVLKDSVTETLPRVNRTAAISGYSFKQCSHQISACNKDIYVDLDMKGINYNSSVAKSAQECQERCTDDVHCHFFTYATRQFPSLEHRNICLLKHTQTGTPTRITKLDKVVSGFSLKSCALSNLACIRDIFPNTVFADSNIDSVMAPDAFVCGRICTHHPGCLFFTFFSQEWPKESQRNLCLLKTSESGLPSTRIKKSKALSGFSLQSCRHSIPVFCHSSFYHDTDFLGEELDIVAAKSHEACQKLCTNAVRCQFFTYTPAQASCNEGKGKCYLKLSSNGSPTKILHGRGGISGYTLRLCKMDNECTTKIKPRIVGGTASVRGEWPWQVTLHTTSPTQRHLCGGSIIGNQWILTAAHCFYGVESPKILRVYSGILNQSEIKEDTSFFGVQEIIIHDQYKMAESGYDIALLKLETTVNYTDSQRPICLPSKGDRNVIYTDCWVTGWGYRKLRDKIQNTLQKAKIPLVTNEECQKRYRGHKITHKMICAGYREGGKDACKGDSGGPLSCKHNEVWHLVGITSWGEGCAQRERPGVYTNVVEYVDWILEKTQAV
[0245] SEQ ID NO:28
[0246] Human kallikrein:
[0247] MILFKQATYFISLFATVSCGCLTQLYENAFFRGGDVASMYTPNAQYCQMRCTFHPRCLLFSFLPASSINDMEKRFGCFLKDSVTGTLPKVHRTGAVSGHSLKQCGHQISACHRDIYKGVDMRGVNFNVSKVSSVEECQKRCTNNIRCQFFSYATQTFHKAEYRNNCLLKYSPGGTPTAIKVLSNVESGFSLKPCALSEIGCHMNIFQHLAFSDVDVARVLTPDAFVCRTICTYHPNCLFFTFYTNVWKIESQRNVCLLKTSESGTPSSSTPQENTISGYSLLTCKRTLPEPCHSKIYPGVDFGGEELNVTFVKGVNVCQETCTKMIRCQFFTYSLLPEDCKEEKCKCFLRLSMDGSPTRIAYGTQGSSGYSLRLCNTGDNSVCTTKTSTRIVGGTNSSWGEWPWQVSLQVKLTAQRHLCGGSLIGHQWVLTAAHCFDGLPLQDVWRIYSGILNLSDITKDTPFSQIKEIIIHQNYKVSEGNHDIALIKLQAPLNYTEFQKPICLPSKGDTSTIYTNCWVTGWGFSKEKGEIQNILQKVNIPLVTNEECQKRYQDYKITQRMVCAGYKEGGKDACKGDSGGPLVCKHNGMWRLVGITSWGEGCARREQPGVYTKVAEYMDWILEKTQSSDGKAQMQSPA
[0248] Activatable linker sequence
[0249] SEQ ID NO:29
[0250] FXIIa-sensitive linker (Factor XI peptide):
[0251] CTTKIKPRIVGGTASVRGEWPWQVT
[0252] SEQ ID NO:30
[0253] FXIIa-sensitive linker
[0254] GGGGSPRIGGGGS
[0255] SEQ ID NO:31
[0256] FXIIa-sensitive linker (prekallikrein peptide):
[0257] VCTTKTSTRIVGGTNSSWGEWPWQVS
[0258] SEQ ID NO:32
[0259] FXIIa-sensitive linker (prekallikrein peptide):
[0260] STRIVGG
[0261] SEQ ID NO:64
[0262] Thrombin-sensitive linker 1:
[0263] LVPRG
[0264] SEQ ID NO:65
[0265] Thrombin-sensitive linker 2:
[0266] SGGGGLVPRGSGGGG
[0267] BD-deleted D1L3 fusion protein
[0268] SEQ ID NO:33
[0269] Albumin-DNase 1L3 variant-fusion protein. albumin , DNA enzyme 1L3 variant S283_S305del):
[0270]
[0271]
[0272] SEQ ID NO:34
[0273] BD-deleted D1L3–Fc fusion protein (signal peptide, DNAse 1L3 (S283_S305del), Fc fragment ):
[0274]
[0275] SEQ ID NO:76
[0276] Albumin-DNase 1L3 variant-fusion protein. albumin – Connector – has C48A Substituted DNase 1L3 variant S283_S305del– C-terminal extension ):
[0277]
[0278]
[0279] SEQ ID NO:78
[0280] Albumin-DNase 1L3 variant-fusion protein. albumin – Connector – has C48A Substituted DNase 1L3 variant S283_S305del– C-terminal extension ):
[0281]
[0282]
[0283] Wild-type D1L3 fusion protein
[0284] SEQ ID NO:35
[0285] Albumin–WT DNase 1L3–fusion protein.( albumin , DNase 1L3):
[0286]
[0287]
[0288] SEQ ID NO:36
[0289] WT DNA enzyme 1L3–Fc fusion protein. (Signal peptide, DNA enzyme 1L3, Fc fragment ):
[0290]
[0291] signal peptide
[0292] SEQ ID NO:37
[0293] DNase 1L3
[0294] MSRELAPLLLLLLSIHSALA
[0295] SEQ ID NO:38
[0296] α mating factor
[0297] MRFPSIFTAVLFAASSALAAPVNTTTEDETAQIPAEAVIGYSDLEGDFDVAVLPFSNSTNNGLLFINTTIASIAAKEEGVSLEKR
[0298] SEQ ID NO:39
[0299] α-mating factor presequence
[0300] EFETMRFPSIFTAVLFAASSALA
[0301] SEQ ID NO:40
[0302] Human serum albumin
[0303] MKWVTFISLLFLFSSAYS
[0304] SEQ ID NO:41
[0305] Bovine DNase 1
[0306] MRGTRLMGLLLALAGLLQLGLS
[0307] SEQ ID NO:42
[0308] Bovine DNA enzyme 1 + Kex2 site
[0309] MRGTRLMGLLLALAGLLQLGLSLEKR
[0310] SEQ ID NO:43
[0311] α-amylase
[0312] EFETMRFPSIFTAVLFAASSALA
[0313] SEQ ID NO:44
[0314] Glucoamylase signal peptide
[0315] EFETMSFRSLLALSGLVCSGLA
[0316] SEQ ID NO:45
[0317] Inulinase
[0318] EFETMKLAYSLLLPLAGVSA
[0319] SEQ ID NO:46
[0320] Invertase
[0321] EFETMLLQAFLFLLAGFAAKISA
[0322] SEQ ID NO:47
[0323] Killer protein
[0324] EFETMTKPTQVLVRSVSILFFITLLHLVVA
[0325] SEQ ID NO:48
[0326] Lysozyme
[0327] EFETMLGKNDPMCLVLVLLGLTALLGICQG
[0328] SEQ ID NO:49
[0329] Synthetic Escherichia coli signal peptide
[0330] MKKNIAFLLALMFVFSIATNAYA
[0331] SEQ ID NO:50
[0332] Synthetic Escherichia coli signal peptide
[0333] MKKNIAFLLAIMFVFSIATNAYA
[0334] SEQ ID NO:51
[0335] Salmonella Enteritidis DsbA signal peptide
[0336] MKKIWLALAGIVLAFSASA
[0337] SEQ ID NO:52
[0338] Synthetic Escherichia coli signal peptide
[0339] MKKIWLALAGLVLAFSAYA
[0340] SEQ ID NO:53
[0341] Synthetic Escherichia coli signal peptide
[0342] MKKNIAFLLAAMFVFSIATNAYA
[0343] SEQ ID NO:54
[0344] Synthetic Escherichia coli signal peptide
[0345] MKKNILFLLLLMFVFSIATNAYA
[0346] SEQ ID NO:55
[0347] Escherichia coli FimD signal peptide
[0348] MMTKIKLLMLIIFYLIISASAHA
[0349] SEQ ID NO:56
[0350] Escherichia coli OmpA signal peptide
[0351] MKKRARAIAIAVALAGFATVAHA
[0352] SEQ ID NO:57
[0353] Signal peptide from Bordetella pertussis
[0354] MKKWFVAAGIGAGLLMLSSAA
[0355] SEQ ID NO:67
[0356] Escherichia coli DsbA signal peptide
[0357] KKIWLALAGLVLAFSASA
[0358] SEQ ID NO:68
[0359] Escherichia coli ST-II signal peptide
[0360] MKKNIAFLLASMFVFSIATNAYA
[0361] A linker that separates the signal peptide from the secreted protein
[0362] SEQ ID NO:58
[0363] GGGGS
[0364] SEQ ID NO:59
[0365] GGGGSGGGGSGGGGS
[0366] SEQ ID NO:60
[0367] SGGGG
[0368] SEQ ID NO:74
[0369] CGGGG
[0370] SEQ ID NO:61
[0371] SGGSGGSGG
[0372] SEQ ID NO:62
[0373] SGGSGGSGGSGGSGGSGG
[0374] SEQ ID NO:69
[0375] SGGSGGSGSS
[0376] C-terminal basic domain
[0377] SEQ ID NO:75
[0378] SSRAFTNSKKSVTLRKKTKSKRS
Claims
1. A variant of a DNA enzyme 1-like protein 3 (D1L3 variant), comprising an N-terminal extension of at least 4 amino acids and no more than about 18 amino acids relative to SEQ ID NO:
4.
2. The D1L3 variant of claim 1, wherein the N-terminal amino acid is not Met.
3. The D1L3 variant of claim 1 or 2, wherein the N-terminal extension comprises primarily Gly residues.
4. The D1L3 variant of any one of claims 1 to 3, wherein the D1L3 variant comprises an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% sequence identity to amino acids 21 to 282 of SEQ ID NO: 4 or amino acids 21 to 252 of SEQ ID NO:
5.
5. The D1L3 variant of any one of claims 1 to 4, wherein the D1L3 variant is generated by cleavage of an N-terminal signal peptide selected from the group consisting of SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, and SEQ ID NO:
48.
6. The D1L3 variant of claim 5, wherein the signal peptide is the alpha mating factor (αMF) prepro secretory leader sequence from Saccharomyces cerevisiae (SEQ ID NO: 38).
7. The D1L3 variant of any one of claims 1 to 6, wherein the N-terminal extension is at least four, or at least five, or at least six, or at least seven, or at least eight, or at least nine, or at least ten, or at least 11, or at least 12, or at least 13, or at least 14, or at least 15 amino acids in length.
8. The D1L3 variant of claim 7, wherein the N-terminal extension consists essentially of or consists of Ser and Gly residues.
9. The D1L3 variant of claim 8, wherein the N-terminal extension and / or a conjugate of the N-terminal extension and a sequence derived from D1L3 is non-immunogenic.
10. The D1L3 variant of any one of claims 7 to 9, wherein the N-terminal extension comprises an amino acid having a chemical group suitable for chemical conjugation, the chemical group being selected from a thiol group, an amino group, an amide group, and a carboxyl group.
11. The D1L3 variant of any one of claims 7 to 10, wherein the N-terminal extension has an amino acid sequence of SGGGG (SEQ ID NO: 60) or CGGGG (SEQ ID NO: 74).
12. The D1L3 variant of any one of claims 1 to 11, wherein the N-terminal extension does not comprise a consensus sequence for myristoylation.
13. The D1L3 variant of any one of claims 1 to 12, wherein the D1L3 variant further comprises a C-terminal extension that reduces heterogeneity.
14. The D1L3 variant of claim 13, wherein the C-terminal extension is at least two, or at least three, or at least four, or at least five amino acids in length.
15. The D1L3 variant of any one of claims 1 to 14, wherein the C-terminal amino acid is Lys or Arg.
16. The D1L3 variant according to any one of claims 13 to 15, wherein the C-terminal extension comprises or consists of the amino acid sequence SSR.
17. The D1L3 variant of any one of claims 1 to 16, wherein the D1L3 variant comprises a basic domain.
18. The D1L3 variant of any one of claims 1 to 17, wherein the D1L3 variant comprises a deletion of at least three, or at least five, or at least eight, or at least ten, or at least twelve, or at least fifteen, or at least eighteen, or at least twenty, or all 23 amino acids of the C-terminal basic domain (BD) defined by amino acids 283 to 305 of SEQ ID NO:
4.
19. The D1L3 variant of any one of claims 1 to 18, wherein the D1L3 variant is configured to form a dimer via an unpaired Cys, wherein the unpaired Cys is optionally C48 with respect to SEQ ID NO:
4.
20. The D1L3 variant of any one of claims 1 to 18, wherein the D1L3 variant comprises a substitution with respect to C48 and / or C174 of SEQ ID NO:
4.
21. The D1L3 variant of claim 20, wherein the mutation is selected from C48S, C48G, C48A, C174S, C174G, and C174A with respect to SEQ ID NO:
4.
22. The D1L3 variant of claim 21, wherein the substitution is C48A with respect to SEQ ID NO:
4.
23. The D1L3 variant of any one of claims 1 to 22, wherein the C-terminal extension comprises or consists of the amino acid sequence SSR.
24. The D1L3 variant of any one of claims 1 to 23, wherein the D1L3 variant comprises one or more mutations that confer resistance to proteolysis by one or more of plasmin, thrombin, trypsin, and proteases produced by mammalian and non-mammalian cell lines.
25. The D1L3 variant of claim 24, wherein the D1L3 variant has one or more amino acid residue mutations selected from K180, K200, K259, and R285 with respect to SEQ ID NO:
4.
26. The D1L3 variant of claim 24 or claim 25, wherein the D1L3 variant has one or more amino acid residue mutations selected from R22, R29, K45, K47, K74, R81, R92, K107, K176, R212, R226, R227, K250, K259, and K262 with respect to SEQ ID NO:
4.
27. The D1L3 variant of claim 25 or claim 26, wherein the D1L3 variant has one or more amino acid residue mutations selected from S91C, S131C, and S253C with respect to SEQ ID NO:
4.
28. The D1L3 variant of claim 1, having the amino acid sequence of SEQ ID NO: 63, SEQ ID NO: 66, SEQ ID NO: 70, SEQ ID NO: 71, SEQ ID NO: 72, SEQ ID NO: 73, or SEQ ID NO:
77.
29. The D1L3 variant of claim 28, having the amino acid sequence of SEQ ID NO:
70.
30. The D1L3 variant of any one of claims 1 to 29, wherein the D1L3 variant comprises a fusion or conjugation to a half-life extending moiety.
31. The D1L3 variant of claim 30, wherein the half-life extending moiety is a polymer.
32. The D1L3 variant of claim 31 , wherein the polymer is polyethylene glycol (PEG).
33. The D1L3 variant of claim 32, wherein a PEG polymer is conjugated to the N-terminus.
34. The D1L3 variant of claim 33, wherein the PEG polymer is linked to the D1L3 variant molecules via the N-termini of the two D1L3 variant molecules.
35. The D1L3 variant of any one of claims 31-33, wherein a PEG polymer is conjugated to one or more amino acids within positions corresponding to R95 to V126 of SEQ ID NO:
4.
36. The D1L3 variant of claim 35, wherein the one or more PEGylated amino acids are selected from lysine, cysteine, histidine, arginine, aspartic acid, glutamic acid, serine, threonine and tyrosine, optionally wherein the one or more PEGylated amino acids are introduced by substitution of one or more amino acids between R95 and V126 relative to SEQ ID NO:
4.
37. The D1L3 variant of claim 36, wherein the one or more amino acids are pegylated by: (a) PEGylation of lysine (Lys or K) via amine conjugation; (b) PEGylation of glutamine by transglutaminase (TGase)-mediated enzymatic conjugation; and / or (c) PEGylation of cysteine (Cys or C) via thiol conjugation.
38. The D1L3 variant of any one of claims 32 to 37, wherein one or more PEGylated amino acids are conjugated to a PEG moiety independently selected from a linear or branched PEG, the molecular weight of the PEG being independently selected and ranging from about 2 kDa to about 60 kDa or from about 5 kDa to about 30 kDa.
39. The D1L3 variant of claim 30, wherein the half-life extending moiety is a fusion partner.
40. The D1L3 variant of claim 39, wherein the fusion partner is selected from albumin, transferrin, Fc, or an elastin-like protein, an XTEN sequence, or a variant thereof.
41. The D1L3 variant of claim 40, wherein the fusion partner is albumin.
42. The D1L3 variant of claim 41, wherein the fusion partner is human albumin comprising an amino acid sequence that is at least 80% identical to SEQ ID NO:
26.
43. The D1L3 variant of claim 42, wherein the human albumin comprises E505Q, T527M, K573P substitutions relative to SEQ ID NO:
26.
44. The D1L3 variant of any one of claims 39 to 43, wherein the fusion partner is fused to the mature D1L3 enzyme at the N-terminus.
45. The D1L3 variant of any one of claims 39 to 44, wherein the fusion partner is fused to the mature D1L3 enzyme via a linker that adjoins the fusion partner and the mature D1L3 enzyme.
46. The D1L3 variant of claim 45, wherein the linker has a length of about 5 to about 50 amino acids, or about 10 to about 35 amino acids, or about 15 to about 35 amino acids.
47. The D1L3 variant of claim 46, wherein the linker comprises the amino acid sequence S(GGS)4GSS (SEQ ID NO: 23), S(GGS)9GSS (SEQ ID NO: 24), or (GGS)9GS (SEQ ID NO: 25).
48. The D1L3 variant of claim 46 or 47, wherein the D1L3 variant comprises the amino acid sequence of SEQ ID NO: 76 or SEQ ID NO:
78.
49. The D1L3 variant of claim 45 or 46, wherein the linker is a flexible or rigid linker and / or comprises a protease cleavage site.
50. The D1L3 variant of claim 49, wherein the linker is cleavable by a coagulation pathway protease.
51. The D1L3 variant of claim 50, wherein the protease is Factor XII or neutrophil protease.
52. The D1L3 variant of claim 51, wherein the protease is thrombin.
53. The D1L3 variant of claim 50 or 51, wherein the protease cleavage site comprises the amino acid sequence LVPRG (SEQ ID NO: 64), and optionally SGGGGLVPRGSGGGG (SEQ ID NO: 65).
54. A method for expressing the DNA enzyme 1-like protein 3 variant according to any one of claims 1 to 53, the method comprising: introducing a gene construct encoding the DNA enzyme 1-like protein 3 variant comprising a signal peptide into a yeast cell, and The DNase 1-like protein 3 variant is recovered.
55. The method of claim 54, wherein the yeast cell is Pichia pastoris.
56. The method of claim 54 or claim 55, wherein the signal peptide is the alpha mating factor (aMF) prepro secretory leader from Saccharomyces cerevisiae (SEQ ID NO: 38).
57. An isolated polynucleotide encoding the D1L3 variant of any one of claims 1 to 53.
58. The isolated polynucleotide of claim 57, wherein the polynucleotide is mRNA or modified mRNA (mmRNA).
59. The isolated polynucleotide of claim 58, wherein the polynucleotide is DNA.
60. A vector for introducing the polynucleotide of any one of claims 57 to 59 into a host cell.
61. A host cell comprising the vector of claim 60.
62. A pharmaceutical composition comprising an effective amount of the D1L3 variant of any one of claims 1 to 53, or the D1L3 variant produced by the method of any one of claims 54 to 56, or the polynucleotide of any one of claims 57 to 59, or the vector of claim 60, or the host cell of claim 61, and a pharmaceutically acceptable carrier.
63. The pharmaceutical composition of claim 62, wherein the composition comprises the D1L3 variant of any one of claims 1 to 53 and a pharmaceutically acceptable carrier for parenteral administration.
64. The pharmaceutical composition of claim 62, formulated for ocular or pulmonary administration.
65. The pharmaceutical composition of claim 63, formulated for intradermal, intramuscular, intravenous, subcutaneous, or intraarterial administration.
66. A method for treating a subject in need of extracellular chromatin degradation, the method comprising administering the pharmaceutical composition of any one of claims 63 to 65 to a subject in need thereof.
67. The method of claim 66, wherein the subject is in need of extracellular trap (ET) degradation and / or neutrophil extracellular trap (NET) degradation.
68. The method of claim 66 or 67, wherein the subject has a loss-of-function mutation in one or both D1L3 genes.
69. The method of any one of claims 67 or 68, wherein the subject has SLE.
70. The method of any one of claims 66 to 69, wherein the subject suffers from a condition selected from the group consisting of chronic neutrophilia, neutrophil aggregation or leukostasis, thrombosis or vascular occlusion, ischemia-reperfusion injury, surgical or traumatic tissue injury, acute or chronic inflammatory reactions or diseases, autoimmune diseases, cardiovascular diseases, metabolic diseases, systemic inflammation, inflammatory diseases of the respiratory tract, inflammatory diseases of the kidneys, inflammatory diseases associated with transplanted tissue, and cancer.
71. The method of any one of claims 66 to 69, wherein the subject has or is at risk of having a NET obstructing the ductal system, wherein the condition is optionally selected from pancreatitis, cholangitis, conjunctivitis, mastitis, dry eyes, vas deferens obstruction, and nephropathy.
72. The method of any one of claims 66 to 71, wherein the subject has or is at risk of NETs accumulating on the endothelial surface.
73. An expression construct for improving the processing of a polypeptide precursor in a host, said expression construct comprising a signal peptide fused to said polypeptide via a linker of at least three amino acids in length.
74. The expression construct of claim 73, wherein in the absence of the linker, the signal peptide is incompletely processed in the host.
75. The expression construct of claim 73 or claim 74, wherein the polypeptide is selected from a component of an enzyme, a cytokine, a hormone, an antibody or its Fab, an antibody-like molecule or its Fab, an antigen, a vaccine, a fusion protein, and a combination thereof.
76. An expression construct as described in claim 75, wherein the enzyme is selected from DNase 1 (D1), DNase 1-like protein 1 (D1L1), DNase 1-like protein 2 (D1L2), DNase 1-like protein 3 isoform 1 (D1L3), DNase 1-like protein 3 isoform 2 (D1L3-2), DNase 2A (D2A) and DNase 2B (D2B) or variants thereof.
77. The expression construct of any one of claims 73 to 76, wherein the linker has a length of at least 5 amino acids, or a length of at least 9 amino acids, or a length of at least 12 amino acids, and is primarily serine and glycine residues.
78. The expression construct of any one of claims 73 to 77, wherein the N-terminal extension and / or a conjugate of the N-terminal extension and a sequence derived from D1L3 is non-immunogenic.
79. The expression construct of any one of claims 73 to 78, wherein the N-terminal extension comprises an amino acid having a chemical group suitable for chemical conjugation, the chemical group being selected from a thiol group, an amino group, an amide group, and a carboxyl group.
80. The expression construct of any one of claims 73 to 79, wherein the signal peptide is selected from the group consisting of DNase 1L3 (SEQ ID NO: 37), alpha mating factor (SEQ ID NO: 38), alpha mating factor presequence (SEQ ID NO: 39), human serum albumin (SEQ ID NO: 40), bovine DNase 1 (SEQ ID NO: 41), bovine DNase 1 + Kex2 site (SEQ ID NO: 42), alpha amylase (SEQ ID NO: 43), glucoamylase signal peptide (SEQ ID NO: 44), inulinase (SEQ ID NO: 45), invertase (SEQ ID NO: 46), killer protein (SEQ ID NO: 47), and lysozyme (SEQ ID NO: 48).
81. The expression construct of any one of claims 73 to 80, wherein the host is a yeast, or a cell line selected from a mammalian cell line and an insect cell line.
82. The expression construct of claim 81, wherein the host is Pichia pastoris.
83. The expression construct of claim 82, wherein the signal peptide is the alpha mating factor (aMF) prepro secretory leader from Saccharomyces cerevisiae (SEQ ID NO: 38).
Citation Information
Patent Citations
Methods and compositions for secretion of heterologous polypeptides
US10435694B2
Compositions and methods for producing high secreted yields of recombinant proteins
US11306127B2
Compositions and methods for producing high secreted yields of recombinant proteins
US11370815B2
Highly efficient secretory signal peptide and a protein expression system using the peptide thereof
US20070117186A1
Peptide Library
US20100055125A1
Cited By
Engineering of dnase enzymes for manufacturing and therapy
CN112996911A
Engineering of deoxyribonucleases for manufacturing and therapy
CN112996911B